instar 1.3.830 → 1.3.832

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (114) hide show
  1. package/dist/commands/server.d.ts.map +1 -1
  2. package/dist/commands/server.js +83 -12
  3. package/dist/commands/server.js.map +1 -1
  4. package/dist/core/AutonomousRealCheckAnnotator.d.ts +105 -0
  5. package/dist/core/AutonomousRealCheckAnnotator.d.ts.map +1 -0
  6. package/dist/core/AutonomousRealCheckAnnotator.js +122 -0
  7. package/dist/core/AutonomousRealCheckAnnotator.js.map +1 -0
  8. package/dist/core/AutonomousRunStore.d.ts +29 -0
  9. package/dist/core/AutonomousRunStore.d.ts.map +1 -1
  10. package/dist/core/AutonomousRunStore.js +38 -0
  11. package/dist/core/AutonomousRunStore.js.map +1 -1
  12. package/dist/core/BackupManager.d.ts.map +1 -1
  13. package/dist/core/BackupManager.js +47 -1
  14. package/dist/core/BackupManager.js.map +1 -1
  15. package/dist/core/CircuitBreakingIntelligenceProvider.d.ts +12 -0
  16. package/dist/core/CircuitBreakingIntelligenceProvider.d.ts.map +1 -1
  17. package/dist/core/CircuitBreakingIntelligenceProvider.js +56 -4
  18. package/dist/core/CircuitBreakingIntelligenceProvider.js.map +1 -1
  19. package/dist/core/CompletionEvaluator.d.ts +62 -2
  20. package/dist/core/CompletionEvaluator.d.ts.map +1 -1
  21. package/dist/core/CompletionEvaluator.js +124 -6
  22. package/dist/core/CompletionEvaluator.js.map +1 -1
  23. package/dist/core/DecisionQualityRecorderImpl.d.ts +192 -0
  24. package/dist/core/DecisionQualityRecorderImpl.d.ts.map +1 -0
  25. package/dist/core/DecisionQualityRecorderImpl.js +363 -0
  26. package/dist/core/DecisionQualityRecorderImpl.js.map +1 -0
  27. package/dist/core/IntelligenceRouter.d.ts +33 -0
  28. package/dist/core/IntelligenceRouter.d.ts.map +1 -1
  29. package/dist/core/IntelligenceRouter.js +240 -54
  30. package/dist/core/IntelligenceRouter.js.map +1 -1
  31. package/dist/core/JudgmentProvenanceLog.d.ts +102 -2
  32. package/dist/core/JudgmentProvenanceLog.d.ts.map +1 -1
  33. package/dist/core/JudgmentProvenanceLog.js +239 -11
  34. package/dist/core/JudgmentProvenanceLog.js.map +1 -1
  35. package/dist/core/MessagingToneGate.d.ts +30 -0
  36. package/dist/core/MessagingToneGate.d.ts.map +1 -1
  37. package/dist/core/MessagingToneGate.js +63 -0
  38. package/dist/core/MessagingToneGate.js.map +1 -1
  39. package/dist/core/PostUpdateMigrator.d.ts +11 -0
  40. package/dist/core/PostUpdateMigrator.d.ts.map +1 -1
  41. package/dist/core/PostUpdateMigrator.js +32 -0
  42. package/dist/core/PostUpdateMigrator.js.map +1 -1
  43. package/dist/core/WriteDomainRegistry.d.ts.map +1 -1
  44. package/dist/core/WriteDomainRegistry.js +24 -0
  45. package/dist/core/WriteDomainRegistry.js.map +1 -1
  46. package/dist/core/decisionGradingPass.d.ts +89 -0
  47. package/dist/core/decisionGradingPass.d.ts.map +1 -0
  48. package/dist/core/decisionGradingPass.js +188 -0
  49. package/dist/core/decisionGradingPass.js.map +1 -0
  50. package/dist/core/decisionQualityTypes.d.ts +170 -0
  51. package/dist/core/decisionQualityTypes.d.ts.map +1 -0
  52. package/dist/core/decisionQualityTypes.js +85 -0
  53. package/dist/core/decisionQualityTypes.js.map +1 -0
  54. package/dist/core/devGatedFeatures.d.ts.map +1 -1
  55. package/dist/core/devGatedFeatures.js +6 -0
  56. package/dist/core/devGatedFeatures.js.map +1 -1
  57. package/dist/core/machineCoherenceManifest.d.ts.map +1 -1
  58. package/dist/core/machineCoherenceManifest.js +4 -0
  59. package/dist/core/machineCoherenceManifest.js.map +1 -1
  60. package/dist/core/types.d.ts +69 -4
  61. package/dist/core/types.d.ts.map +1 -1
  62. package/dist/core/types.js.map +1 -1
  63. package/dist/data/provenanceCoverage.d.ts +156 -0
  64. package/dist/data/provenanceCoverage.d.ts.map +1 -0
  65. package/dist/data/provenanceCoverage.js +674 -0
  66. package/dist/data/provenanceCoverage.js.map +1 -0
  67. package/dist/monitoring/ExternalHogDecisionStore.d.ts +256 -0
  68. package/dist/monitoring/ExternalHogDecisionStore.d.ts.map +1 -0
  69. package/dist/monitoring/ExternalHogDecisionStore.js +481 -0
  70. package/dist/monitoring/ExternalHogDecisionStore.js.map +1 -0
  71. package/dist/monitoring/ExternalHogRealAdapters.d.ts +6 -2
  72. package/dist/monitoring/ExternalHogRealAdapters.d.ts.map +1 -1
  73. package/dist/monitoring/ExternalHogRealAdapters.js +2 -2
  74. package/dist/monitoring/ExternalHogRealAdapters.js.map +1 -1
  75. package/dist/monitoring/ExternalHogScanTick.d.ts +95 -3
  76. package/dist/monitoring/ExternalHogScanTick.d.ts.map +1 -1
  77. package/dist/monitoring/ExternalHogScanTick.js +102 -9
  78. package/dist/monitoring/ExternalHogScanTick.js.map +1 -1
  79. package/dist/monitoring/ExternalHogSentinel.d.ts +71 -3
  80. package/dist/monitoring/ExternalHogSentinel.d.ts.map +1 -1
  81. package/dist/monitoring/ExternalHogSentinel.js +126 -2
  82. package/dist/monitoring/ExternalHogSentinel.js.map +1 -1
  83. package/dist/monitoring/ExternalHogServerPrimitives.d.ts +10 -2
  84. package/dist/monitoring/ExternalHogServerPrimitives.d.ts.map +1 -1
  85. package/dist/monitoring/ExternalHogServerPrimitives.js +1 -1
  86. package/dist/monitoring/ExternalHogServerPrimitives.js.map +1 -1
  87. package/dist/monitoring/FeatureMetricsLedger.d.ts +304 -0
  88. package/dist/monitoring/FeatureMetricsLedger.d.ts.map +1 -1
  89. package/dist/monitoring/FeatureMetricsLedger.js +731 -0
  90. package/dist/monitoring/FeatureMetricsLedger.js.map +1 -1
  91. package/dist/scaffold/templates.d.ts.map +1 -1
  92. package/dist/scaffold/templates.js +2 -1
  93. package/dist/scaffold/templates.js.map +1 -1
  94. package/dist/server/AgentServer.d.ts.map +1 -1
  95. package/dist/server/AgentServer.js +55 -1
  96. package/dist/server/AgentServer.js.map +1 -1
  97. package/dist/server/CapabilityIndex.d.ts.map +1 -1
  98. package/dist/server/CapabilityIndex.js +15 -1
  99. package/dist/server/CapabilityIndex.js.map +1 -1
  100. package/dist/server/fileRoutes.d.ts.map +1 -1
  101. package/dist/server/fileRoutes.js +9 -0
  102. package/dist/server/fileRoutes.js.map +1 -1
  103. package/dist/server/routes.d.ts.map +1 -1
  104. package/dist/server/routes.js +334 -5
  105. package/dist/server/routes.js.map +1 -1
  106. package/package.json +1 -1
  107. package/src/data/builtin-manifest.json +65 -65
  108. package/src/data/provenanceCoverage.ts +850 -0
  109. package/src/scaffold/templates/jobs/instar/llm-decision-grading.md +32 -0
  110. package/src/scaffold/templates.ts +2 -1
  111. package/upgrades/1.3.831.md +30 -0
  112. package/upgrades/1.3.832.md +22 -0
  113. package/upgrades/side-effects/llm-decision-quality-meter.md +95 -0
  114. package/upgrades/side-effects/messaging-tone-gate-provenance-enrollment.md +68 -0
@@ -0,0 +1,32 @@
1
+ ---
2
+ name: LLM-Decision Grading Pass
3
+ description: "Hourly deterministic grading pass over the LLM-decision quality substrate. Runs POST /decision-quality/grade-pass: the endpoint walks NEW outcome evidence since the durable per-decision-point cursor (keyset (ts, correlation_id) — same-ms bursts cannot skip rows), applies the registered deterministic evidence rules, and upserts right/wrong/unknown grades idempotently (re-runs converge, never multiply; bounded per pass by provenance.quality.maxDecisionsPerPass). ZERO LLM spend in the pass itself — the grading ladder in this build is deterministic-only (FD11; the LLM evidence-interpreter rung ships NO code, ACT-1198). Ships enabled:false (cost-bearing job class); the operator read surface is GET /decision-quality. NEVER messages the user (FD5 — the meter is observe-only; grading produces rows, not messages). Runs per machine over that machine's local rows (the ratified machine-local data posture). Tier-1 supervised (this haiku job wraps the deterministic endpoint and sanity-checks the response shape). Spec docs/specs/llm-decision-quality-meter.md §5.5."
4
+ schedule: "0 * * * *"
5
+ priority: low
6
+ expectedDurationMinutes: 2
7
+ model: haiku
8
+ supervision: tier1
9
+ enabled: false
10
+ tags:
11
+ - cat:observability
12
+ - decision-quality
13
+ - role:worker
14
+ gate: curl -sf http://localhost:${INSTAR_PORT:-4042}/health >/dev/null 2>&1
15
+ toolAllowlist: "*"
16
+ unrestrictedTools: true
17
+ mcpAccess: none
18
+ perMachineIndependent: true
19
+ ---
20
+ Run one deterministic LLM-decision grading pass. This is a mechanical, near-silent cadence job — do NOT message the user (FD5: the quality meter is observe-only; grading writes grade rows, never messages, and never interprets the grades — interpretation belongs to the operator's read of GET /decision-quality). It exists because outcome evidence accrues continuously, but nothing grades it without a production trigger; this job IS that trigger on the hourly cadence the evidence windows are derived from.
21
+
22
+ AUTH="${INSTAR_AUTH_TOKEN:-$(python3 -c "import json; v=json.load(open('.instar/config.json')).get('authToken',''); print(v if isinstance(v, str) else '')" 2>/dev/null)}"
23
+ AGENT_ID="${INSTAR_AGENT_ID:-$(python3 -c "import json; print(json.load(open('.instar/config.json')).get('projectName',''))" 2>/dev/null)}"
24
+ PORT="${INSTAR_PORT:-4042}"
25
+
26
+ 1. Trigger one grading pass:
27
+ `curl -s -X POST -H "Authorization: Bearer $AUTH" -H "X-Instar-AgentId: $AGENT_ID" -H "Content-Type: application/json" -d '{}' http://localhost:$PORT/decision-quality/grade-pass`
28
+ The body is `{}` on purpose — every knob (pass bound, evidence windows, retention) comes from config, never from this job. A 503 means the quality substrate is dark for this agent (`provenance.uniformSeam` resolves off) — exit silently, there is nothing to do. On 200 the response is `{ graded, byRule, cursors }` where `cursors` maps decisionPoint → the advanced boundary. The endpoint does ALL the deterministic work: durable-cursor keyset walk, bounded per pass, idempotent grade upserts, P19 backoff on a stuck rule.
29
+
30
+ 2. **Tier-1 supervision (your job).** Sanity-check the response shape before concluding: `graded` should be a number ≥ 0, `byRule` an object, and `cursors` an object (possibly empty — an idle pass is healthy, not an error). If the response is malformed or the curl fails, do NOT retry-flood — note it once and exit; the next hourly tick re-attempts, and re-runs converge by the endpoint's idempotency.
31
+
32
+ 3. Exit silently. This job is just the cadence — it produces grade rows, not user messages. Do NOT relay anything to Telegram, do NOT summarize, and do NOT interpret or act on the grades yourself.
@@ -14,7 +14,7 @@
14
14
  // existing-agent migration there) so the two can never drift. Imported as a runtime
15
15
  // function call inside generateClaudeMd — no module-init cycle (PostUpdateMigrator
16
16
  // never imports templates).
17
- import { SESSION_LISTING_HYGIENE_CLAUDEMD_SECTION, PLAYWRIGHT_PROFILE_REGISTRY_CLAUDEMD_SECTION, MACHINE_LOAD_ASSESSMENT_CLAUDEMD_SECTION, DYNAMIC_MCP_CLAUDEMD_SECTION, SENDER_REJECTION_CLAUDEMD_SECTION, SCOPE_ACCRETION_CLAUDEMD_SECTION, MESH_SELF_HEALING_CLAUDEMD_SECTION, WRITE_ADMISSION_CLAUDEMD_SECTION, DOORWAY_REGISTRY_CLAUDEMD_SECTION, EXTERNAL_HOG_CLAUDEMD_SECTION, ROUTING_SPEND_CLAUDEMD_SECTION, DUPLICATE_RECONCILER_CLAUDEMD_SECTION, AUDIT_CONVERGENCE_CLAUDEMD_SECTION } from '../core/PostUpdateMigrator.js';
17
+ import { SESSION_LISTING_HYGIENE_CLAUDEMD_SECTION, PLAYWRIGHT_PROFILE_REGISTRY_CLAUDEMD_SECTION, MACHINE_LOAD_ASSESSMENT_CLAUDEMD_SECTION, DYNAMIC_MCP_CLAUDEMD_SECTION, SENDER_REJECTION_CLAUDEMD_SECTION, SCOPE_ACCRETION_CLAUDEMD_SECTION, MESH_SELF_HEALING_CLAUDEMD_SECTION, WRITE_ADMISSION_CLAUDEMD_SECTION, DOORWAY_REGISTRY_CLAUDEMD_SECTION, EXTERNAL_HOG_CLAUDEMD_SECTION, ROUTING_SPEND_CLAUDEMD_SECTION, DECISION_QUALITY_CLAUDEMD_SECTION, DUPLICATE_RECONCILER_CLAUDEMD_SECTION, AUDIT_CONVERGENCE_CLAUDEMD_SECTION } from '../core/PostUpdateMigrator.js';
18
18
 
19
19
  export interface AgentIdentity {
20
20
  name: string;
@@ -886,6 +886,7 @@ ${MESH_SELF_HEALING_CLAUDEMD_SECTION(port)}
886
886
  ${WRITE_ADMISSION_CLAUDEMD_SECTION(port)}
887
887
  ${DOORWAY_REGISTRY_CLAUDEMD_SECTION(port)}
888
888
  ${ROUTING_SPEND_CLAUDEMD_SECTION(port)}
889
+ ${DECISION_QUALITY_CLAUDEMD_SECTION(port)}
889
890
  ${SESSION_LISTING_HYGIENE_CLAUDEMD_SECTION(port)}
890
891
  ${AUDIT_CONVERGENCE_CLAUDEMD_SECTION(port)}
891
892
  ${DUPLICATE_RECONCILER_CLAUDEMD_SECTION(port)}
@@ -0,0 +1,30 @@
1
+ # Upgrade Guide — vNEXT
2
+
3
+ <!-- assembled-by: assemble-next-md -->
4
+ <!-- bump: patch -->
5
+
6
+ ## What Changed
7
+
8
+ The **LLM-Decision Quality Meter** (ACT-1193 uniform provenance + ACT-1194 outcome grading) — the remediation the LLM-decision audit named as the real gap. Instar has long had a *cost* meter for its AI decisions (tokens/latency per feature) but no *quality* meter: nothing recorded WHAT a high-stakes decision saw, and nothing checked afterward whether the call turned out right. This build lays that substrate:
9
+
10
+ - **Correlation spine + uniform provenance:** every enrolled high-stakes AI decision now mints a safe correlation id and records the context it saw and the choice it made (replay-proof; the caller's data is never mutated; a failed attempt's cost can never be double-counted). Provenance was wired to exactly 2 of ~59 callsites before; the census + ratchet now track the whole surface.
11
+ - **Outcome grading over time:** a deterministic, evidence-triggered grading pass stamps each recorded decision `right` / `wrong` / `unknown` as reality's evidence matures (did the killed process come back? did the "done" run actually finish?), exposed read-only at `GET /decision-quality` with a `POST /decision-quality/grade-pass` job.
12
+ - **Ships dark + dry-run (measure-only)** behind `provenance.uniformSeam` (dev-gated; `dryRun` default true → metadata-only would-writes, no durable persistence). It changes NO decision and grades nothing durably until the seam is deliberately flipped live after soak. First graded customer is process-kill decisions; the completion-judge grader is a tracked fast-follow (ACT-1202).
13
+ - **Two live production bug fixes ride along (these DO land live):** the dashboard file-browser could serve/edit the raw decision-provenance log (ACT-1200) — now denied; and the backup manager's per-file exclusion could be bypassed (ACT-1201) — now enforced.
14
+
15
+ ## What to Tell Your User
16
+
17
+ Two things. First, a small privacy/robustness fix that takes effect immediately: an internal audit log could previously be opened (and edited) through the file browser, and a backup exclusion could be sidestepped — both are now closed. Second, the bigger piece is measure-only for now: I've built the machinery to grade *my own* automatic decisions — not just count how many I make or what they cost, but check whether they actually turned out right over time. That's the foundation for deciding when a bigger model or a better prompt is genuinely warranted. It's running in a safe, off-by-default mode that records but doesn't act, until it's soaked and you choose to turn it on.
18
+
19
+ ## Summary of New Capabilities
20
+
21
+ - `GET /decision-quality` — read-only quality view (per-decision-point grade rates, strength-first, insufficient-evidence below sample floor, census debt, rejection counters; `?scope=pool` field-allowlisted). 503 when the seam is dark.
22
+ - `POST /decision-quality/grade-pass` — deterministic, idempotent, budget-bounded grading job (keyset cursor; returns `{graded, byRule, cursors}`).
23
+ - Correlation spine + `JudgmentProvenanceLog` uniform provenance seam (dark/dryRun behind `provenance.uniformSeam`; 59-point census + shrink-only coverage ratchet).
24
+ - Live fixes: file-browser exposure of the decision-provenance log denied (ACT-1200); backup per-file exclusion bypass closed (ACT-1201).
25
+
26
+ ## Evidence
27
+
28
+ - New unit/integration/e2e tiers green (read endpoint, grade-pass idempotency + rules, "feature-alive" 200-not-503, pool field-allowlist, dryRun would-write suppression, census/ratchet). Full CI sharded suite is the merge authority.
29
+ - Converged spec (`/spec-converge`, 7 rounds; stamp validator-earned, cross-model codex-cli:gpt-5.5 ok 7/7); `docs/specs/llm-decision-quality-meter.md` + `.eli16.md` companion.
30
+ - Follow-ups tracked: ACT-1202 (completion-judge realcheck grading — needs a Stop-hook protocol change), ACT-1203 (heavy-AgentServer test isolation), ACT-1195 (bench prompt-parity).
@@ -0,0 +1,22 @@
1
+ # Upgrade Guide — vNEXT
2
+
3
+ <!-- assembled-by: assemble-next-md -->
4
+ <!-- bump: patch -->
5
+
6
+ ## What Changed
7
+
8
+ The always-on outbound **MessagingToneGate** is now enrolled as a wired LLM-decision provenance point, executing the pending→wired expansion the LLM-Decision Quality Meter (#1458) already defined. The tone gate is the highest-volume decision point in the fleet, so it records at a hard `budget:500/day` count ceiling (never `full`) and stores its content as **identity only** — a `sha256` of the candidate plus byte/char bounds and code-derived features (channel, message kind, recent-message count, gate-signal kinds). The outbound message body and any plaintext slice of it are **never** stored. Enrolling a fourth wired point activated the per-point round-robin grading sub-budget so grading capacity is shared fairly across points. Recording is observability-only — the tone gate's PASS/BLOCK verdict is byte-identical whether the provenance write succeeds, fails, or is disabled — and is dark-gated behind `provenance.uniformSeam` (enabled on the development agent, dark on the fleet).
9
+
10
+ ## What to Tell Your User
11
+
12
+ None — internal, dev-gated observability with no user-facing surface. The change is dark on the fleet: it does not alter how the agent behaves, what it sends, or what a user sees. When enabled on a development agent it only records (machine-local, identity-only) what the always-on message-safety gate decided, so those decisions can later be graded. A user's messages, and the gate's block/allow behavior, are completely unaffected.
13
+
14
+ ## Summary of New Capabilities
15
+
16
+ None user-facing. Internally: the outbound tone/leak gate now contributes to the LLM-decision quality record (dev-agent only, dark on fleet), giving the periodic grading pass its highest-volume data source while storing message identity only, never content.
17
+
18
+ ## Evidence
19
+
20
+ - Side-effects review: `upgrades/side-effects/messaging-tone-gate-provenance-enrollment.md` (independent second-pass review concurred; verified critical-path inertness and no body-leak).
21
+ - Tests: `tests/unit/messaging-tone-gate-provenance-enrollment.test.ts`, `tests/integration/messaging-tone-gate-provenance.test.ts`, plus the updated `tests/unit/provenance-coverage-ratchet.test.ts` and `tests/unit/decision-grading-pass.test.ts` sub-budget cases — 49 targeted tests green; full lint chain (typecheck + dark-gate + attribution + ratchet) clean.
22
+ - Driven by the converged + approved spec `docs/specs/llm-decision-quality-meter.md` (§5.5 sub-budget, §5.6 volume valve + content classes).
@@ -0,0 +1,95 @@
1
+ # Side-Effects Review — LLM-Decision Quality Meter (uniform provenance + outcome grading)
2
+
3
+ **Version / slug:** `llm-decision-quality-meter`
4
+ **Date:** `2026-07-12`
5
+ **Author:** `echo`
6
+ **Second-pass reviewer:** `echo (dedicated reviewer subagent) — high-risk (touches gate/sentinel decision points + a provenance seam)`
7
+
8
+ ## Summary of the change
9
+
10
+ Adds the quality-meter substrate the LLM-decision audit named as the real gap (ACT-1193 uniform provenance, ACT-1194 outcome grading). Three parts: (1) a **correlation spine** — every enrolled high-stakes AI decision mints a replay-proof correlation id (`AutonomousRunStore`, `IntelligenceRouter`, `decisionQualityTypes`) so cost, provenance, and outcome all thread to one decision; (2) **uniform provenance** — `JudgmentProvenanceLog` recording extended from 2 callsites toward the full surface via `DecisionQualityRecorderImpl`, a 59-point census (`provenanceCoverage.ts`) and a shrink-only coverage ratchet; (3) **outcome grading** — a deterministic, evidence-triggered pass (`decisionGradingPass.ts`, `ExternalHogDecisionStore`, `AutonomousRealCheckAnnotator`) that stamps each recorded decision `right`/`wrong`/`unknown`, surfaced read-only at `GET /decision-quality` with a `POST /decision-quality/grade-pass` job, backed by 4 new `FeatureMetricsLedger` tables + a canonical view. The whole meter ships **dark + dryRun** behind `provenance.uniformSeam` (dev-gated; `dryRun` default true → metadata-only would-writes, nothing durable). Two live production bug fixes ride the same PR: `fileRoutes` no longer serves/edits the raw decision-provenance log (ACT-1200), and `BackupManager` per-file exclusion can no longer be bypassed (ACT-1201).
11
+
12
+ ## Decision-point inventory
13
+
14
+ - `ExternalHogSentinel` kill decision (`ExternalHogScanTick`/`ExternalHogDecisionStore`/`ExternalHogServerPrimitives`) — **pass-through + record** — the kill verdict is UNCHANGED; the change only *records* provenance and grades on-supersede, and the durable write is suppressed while `dryRun`.
15
+ - `CompletionEvaluator` (completion / P13-stop judge) — **pass-through + enroll** — the judge verdict is unchanged; it now mints/carries a correlation id (provenance enrollment). Realcheck-based grading of this customer is the tracked fast-follow ACT-1202, NOT in this build.
16
+ - `decisionGradingPass` grade-stamp — **add (signal only)** — deterministic evidence rules stamp right/wrong/unknown. Never gates, blocks, or acts.
17
+ - `GET /decision-quality` / `POST /decision-quality/grade-pass` — **add (read + job)** — 503 when the seam is dark; Bearer-authed; pool branch field-allowlisted.
18
+ - `fileRoutes` serve/edit of the JP log — **add block (ACT-1200)** — deterministic path-based deny of a known-sensitive internal log.
19
+ - `BackupManager` per-file exclusion — **modify (ACT-1201)** — closes a bypass so an excluded file stays excluded (incl. restore).
20
+
21
+ ---
22
+
23
+ ## 1. Over-block
24
+
25
+ The meter itself has no block/allow surface — it records and grades; it never rejects an input. Over-block is only in scope for the two ride-along fixes:
26
+
27
+ - **fileRoutes (ACT-1200):** the new deny is scoped to the decision-provenance log path(s), not a broad prefix. A legitimate project file that merely *contains* "provenance" in its name is not blocked — the deny matches the concrete JP-log location, verified by `tests/unit/fileRoutes-never-served.test.ts`. Risk of over-block: low, path-exact.
28
+ - **BackupManager (ACT-1201):** enforcing an exclusion cannot over-exclude a file the operator didn't list — the change makes the *configured* exclusion actually apply; it adds no new patterns.
29
+
30
+ ---
31
+
32
+ ## 2. Under-block
33
+
34
+ - **fileRoutes:** the deny covers the JP log; other internal `.jsonl` audit logs under `logs/` were already never in the file-browser's editable roots — but any FUTURE internal log added to a served root would need its own guard. Named residual, not introduced here.
35
+ - **Grading under-block N/A** (not a block surface). The honest under-grade risk instead: an outcome whose evidence never arrives stays `unknown` (correct, not a miss); an outcome graded before its evidence window closes could mis-stamp — mitigated by evidence-triggered grading (the pass only stamps when the deterministic signal is present) + idempotent re-reads.
36
+
37
+ ---
38
+
39
+ ## 3. Level-of-abstraction fit
40
+
41
+ Correct layer. Provenance recording lives at the decision recorder (`DecisionQualityRecorderImpl`), not smeared across each callsite; grading is a separate periodic consumer, not inlined into the gate that made the decision (so a gate never blocks on grading). The census/ratchet enforce coverage at the data layer where the contract lives. The meter feeds the existing `feature_metrics` cost surface rather than creating a parallel one — it extends, not duplicates.
42
+
43
+ ---
44
+
45
+ ## 4. Signal vs authority compliance
46
+
47
+ Compliant. The quality meter is a pure **signal producer** — it records what a decision saw and grades how it turned out; it holds NO blocking authority and changes no verdict. Grading is deterministic evidence-rule stamping, never an LLM re-judgment in this build (the LLM evidence-interpreter is explicitly deferred, FD12/ACT-1198). The two ride-along fixes DO add blocking authority, but each is **deterministic and narrow** (a path-exact file deny; an exclusion-list enforcement) — not brittle heuristics with broad authority. `docs/signal-vs-authority.md` respected: no brittle check gained blocking power.
48
+
49
+ ---
50
+
51
+ ## 5. Interactions
52
+
53
+ - Extends `feature_metrics` (adds tables + a view) — additive; existing cost rows/reads untouched (`FeatureMetricsLedger-quality.test.ts`, `CircuitBreaking-feature-metrics-tap.test.ts`).
54
+ - Retrofits the `/judgment-provenance` pool branch with the same credential guard used by the new `/decision-quality` pool branch — one shared guard, no double-fire.
55
+ - `grade-pass` classified in `WriteDomainRegistry` (`pool-scope-read-merge`, machine-local) so multi-machine write-admission doesn't misroute it.
56
+ - The dryRun gate suppresses the durable `persist()` in `ExternalHogDecisionStore` while keeping in-memory grade-on-supersede + would-write logging — so nothing double-writes, and enabling the seam later can't retro-corrupt.
57
+ - `evolutionActions.autoExpiry` cannot sweep the three deferred ACTs (they're pinned/critical-class per spec §5.6) — no race with queue cleanup.
58
+
59
+ ---
60
+
61
+ ## 6. External surfaces
62
+
63
+ - New API routes (`GET /decision-quality`, `POST /decision-quality/grade-pass`) — 503 when dark, so on the fleet (seam off) they are inert. CLAUDE.md template + `CapabilityIndex` updated (agent-awareness); `PostUpdateMigrator` migration for existing agents (`decision-quality-claudemd-migration.test.ts`).
64
+ - New job template `llm-decision-grading.md` ships `enabled:false`.
65
+ - No message-path, dispatch, or session-lifecycle surface changes. No timing/conversation-state dependence beyond the grading pass's own keyset cursor (idempotent).
66
+
67
+ ---
68
+
69
+ ## 7. Multi-machine posture (Cross-Machine Coherence)
70
+
71
+ **Machine-local BY DESIGN, proxied-on-read.** Provenance + grade rows are written machine-local (a decision happens on one machine; its evidence matures there). Reads merge across machines via the `?scope=pool` branch on `/decision-quality` (field-allowlisted so a hostile peer can't inject fields), matching the existing `/judgment-provenance` pool pattern. Cross-machine outcome ROUTING (a decision on machine A graded by evidence that lands on machine B) is explicitly deferred with honest-degradation shipping now — tracked **ACT-1199**. No user-facing notice surface here (API-only; dashboard is ACT-1197), so no one-voice gating needed. Durable state is per-machine and does not strand on topic transfer (it's not topic-keyed).
72
+
73
+ ---
74
+
75
+ ## 8. Rollback cost
76
+
77
+ Low. The meter is dark/dryRun behind `provenance.uniformSeam` — the back-out for the whole grading/provenance surface is to leave the flag off (fleet default) or set `dryRun:true`; nothing durable was written. The two live fixes are two-file, self-contained denials — rollback is a revert of `fileRoutes.ts` / `BackupManager.ts` via hot-fix release, no data migration, no agent-state repair. No schema destruction: the 4 new tables are additive and unused while dark.
78
+
79
+ ---
80
+
81
+ ## Second-pass review
82
+
83
+ _(Appended by the dedicated reviewer subagent — Phase 5, required: this change touches sentinel/gate decision points and a provenance seam.)_
84
+
85
+ **Concur with the review.**
86
+
87
+ Independently verified against the code, not the prose:
88
+
89
+ 1. **Signal-only holds end-to-end.** `decisionGradingPass.ts` only calls `annotate(...)` (never gates); the `POST /grade-pass` handler (routes.ts:15358) just runs `runDecisionGradingPass`; `ExternalHogSentinel.recordDecisions` (ExternalHogSentinel.ts:237-252) is pure observation — `store.record()` throws are swallowed into `dqStoreErrors`, so a persist failure can never alter the kill verdict.
90
+
91
+ 2. **dryRun is correct.** `this.dryRun = opts.dryRun !== false` (ExternalHogDecisionStore.ts:398, default TRUE); grade-on-supersede runs in-memory *before* the `if (this.dryRun) logWouldPersist()` else `persist()` branch (lines 528-583); `logWouldPersist` emits ledgerKey-count + byte-size only (line 473), no content. The pass treats a suppressed write as PENDING and does NOT advance the cursor (decisionGradingPass.ts:167-171) — a later `dryRun:false` flip won't skip ungraded rows.
92
+
93
+ 3. **Both fixes narrow.** `NEVER_SERVED_PREFIXES` is path-exact (fileRoutes.ts:97-117); `isDeniedForBackup` reuses three existing layers and is applied in dir-copy *and* restore (BackupManager.ts:353, 494).
94
+
95
+ 4. Both routes 503 when dark (15300/15359). 5. `pickDecisionQualityPointFields` is a real field allowlist applied instead of `{...row}` (15337); grade-pass classified machine-local in WriteDomainRegistry.ts:403. 6. Every deferral carries an ACT id; the "out of scope" mentions are non-goals, not dropped work. No missing failure mode.
@@ -0,0 +1,68 @@
1
+ # Side-Effects Review — MessagingToneGate Provenance Enrollment
2
+
3
+ **Version / slug:** `messaging-tone-gate-provenance-enrollment`
4
+ **Date:** `2026-07-12`
5
+ **Author:** `echo`
6
+ **Second-pass reviewer:** `independent reviewer subagent — CONCUR (verified all 6 axes; see Second-pass review section)`
7
+
8
+ ## Summary of the change
9
+
10
+ Enrolls the **MessagingToneGate outbound gate** as a WIRED provenance decision point, executing the pending→wired expansion path that PR #1458's LLM-Decision Quality Meter (`docs/specs/llm-decision-quality-meter.md`, §5.6) already defined and tracked. The census (`src/data/provenanceCoverage.ts`) already carried `messaging-tone-gate` as `status: 'pending:ACT-1193'`; this flips it to `wired` with the mandatory volume valve, and adds the `provenance: {...}` enrollment to the tone gate's single verdict-producing LLM call — mirroring the exact `options.provenance` pattern #1458 established for CompletionEvaluator/ExternalHog. The tone gate is the highest-volume decision point in the fleet (3,641 of 4,098 LLM calls/24h on the dev agent per §5.6), so it enrolls at `budget:500/day` (a hard count ceiling with a loud `droppedByBudget` counter — never `full`), and its content is stored as **identity only** (a sha256 of the candidate + byte/char bounds + code-derived features: channel, messageKind, recent-message count, gate-signal kinds) — **never the message body and never any plaintext slice of it**. Enrolling a 4th wired customer structurally required activating the §5.5 per-point round-robin grading sub-budget (`decisionGradingPass.ts`, `SUBBUDGET_IMPLEMENTED` false→true), so grading capacity is shared fairly across points instead of one point starving the others. Files: `provenanceCoverage.ts`, `MessagingToneGate.ts`, `decisionGradingPass.ts`, + ratchet floor test + 2 new tests. Recording is dev-gated behind the same `provenance.uniformSeam` flag #1458 uses (ENABLED on dev, DARK on fleet).
11
+
12
+ ## Decision-point inventory
13
+
14
+ - `MessagingToneGate` outbound-gate (`messaging-tone-gate`) — **pending→wired (observe)** — records the identity + verdict of the existing tone-gate authority. The gate's block/allow decision is untouched; the seam consumes the provenance block before any adapter and it never reaches the model.
15
+ - `decisionGradingPass` (per-point sub-budget) — **modify (mechanism)** — a 4th wired point activates the §5.5 round-robin sub-budget; no decision authority added, it only allocates grading capacity fairly across points.
16
+
17
+ ---
18
+
19
+ ## 1. Over-block
20
+
21
+ **No block/allow surface — over-block not applicable.** Observability-only. The `provenance` block is passed into the tone gate's LLM call options and consumed by the router-settlement seam *before any adapter*; it never reaches the model, never gates, holds, delays, or filters a message. The tone gate's PASS/BLOCK verdict is byte-identical whether the provenance write succeeds, fails, or is disabled (fail-open, inherited from #1458's recorder; asserted by test).
22
+
23
+ ## 2. Under-block
24
+
25
+ **No block/allow surface — under-block not applicable.** This records what the tone-gate authority already decided. The remaining ~49 pending census points stay tracked under `pending:ACT-1193` and the monotonic ratchet still forbids silent regression — this change *shrinks* the pending set by one (tone gate) and *grows* the wired set to 4.
26
+
27
+ ---
28
+
29
+ ## 3. Level-of-abstraction fit
30
+
31
+ Correct layer, and deliberately NOT a parallel mechanism. The enrollment sits at the tone gate's own verdict call (the only place the structured verdict exists), reusing #1458's census + recorder + envelope + ratchet rather than rebuilding any of it. The census entry was already authored as `pending:ACT-1193` for exactly this expansion; flipping it is the intended, ratchet-enforced growth path. (This increment is the corrective successor to a duplicate parallel build — PR #1460, closed — which is why it is scrupulously on #1458's seam.)
32
+
33
+ ## 4. Signal vs authority compliance
34
+
35
+ **Fully compliant — pure signal-side, zero authority added.** Per `docs/signal-vs-authority.md`: this records the existing tone-gate authority's verdict; it adds no detector-with-authority and no new authority. The one forbidden direction — a graded outcome feeding back into a decision input — is not crossed: grading only reads recorded rows and writes verdicts to the quality store; nothing here wires a feedback edge into model/door/prompt/floor selection (that remains the future benchmark increment's own spec).
36
+
37
+ ## 5. Interactions
38
+
39
+ - **Volume:** the tone gate fires on every drafted outbound message. The `budget:500/day` valve is the load control — a hard per-UTC-day count ceiling on the provenance JSONL archive with a loud `droppedByBudget` counter, deterministic rather than probabilistic. This is why `full` is forbidden for this point.
40
+ - **Grading fairness:** activating the §5.5 sub-budget (required by the 4th point) means the periodic grading pass round-robins capacity across all wired points; without it a high-volume point could monopolise the grading budget. Covered by 3 new sub-budget tests.
41
+ - **No shadow/double-fire:** one enrollment per verdict at the single `review()` call site; the deterministic availability-fallback paths are not the LLM decision and are not enrolled.
42
+ - **`/metrics/features` agreement:** model/door/tokens come from the existing usage/attribution path, so the provenance row and cost row agree by construction.
43
+
44
+ ## 6. External surfaces
45
+
46
+ - **`GET /judgment-provenance` / the decision-quality read surface** now also returns tone-gate rows (redacted, identity-only, Bearer-gated) — reusing #1458's existing envelope + `no-store` route. No new route.
47
+ - **Privacy — the load-bearing property:** content is stored as `sha256(candidate)` + bounds + code-derived features, **never the body and never a plaintext head**. This deviates from the initial "hash + bounded head" instruction: an integration test (written first) caught the literal outbound body landing on the served row because the credential-scrub does NOT strip non-credential PII/prose. The head was dropped entirely, matching the CompletionEvaluator content-bearing sibling's hash-only discipline — the stricter, correct reading for the very text this gate inspects for leaks.
48
+
49
+ ## 7. Multi-machine posture (Cross-Machine Coherence)
50
+
51
+ **machine-local write, proxied-on-read** — identical to #1458's provenance posture (inherited, not re-declared). Full rows are credential-scrubbed + machine-local (0700/0600, never-served-raw, short retention) and never replicated; the unified read is the existing redacted `?scope=pool` merge that #1458 built. Identity-only content further reduces the at-rest surface for this point. The ratchet + census are CI-level (machine-independent). No new machine-local surface is introduced.
52
+
53
+ ## 8. Rollback cost
54
+
55
+ **Cheap and reversible.** Recording is dev-gated behind `provenance.uniformSeam` (dark on fleet); flipping it off leaves the tone gate behaving exactly as before with zero rows. Full back-out is a plain revert: the census entry returns to `pending:ACT-1193`, the enrollment is removed, and `SUBBUDGET_IMPLEMENTED` reverts (the sub-budget is inert with <2 wired high-volume points anyway). No migration, no persisted state beyond the short-retention JSONL, no fleet coordination. Because recording never touched the tone verdict, a revert cannot regress any outbound-messaging behavior.
56
+
57
+ ## Second-pass review (independent)
58
+
59
+ **Concur with the review.** Verified against the real diff + code (not the artifact) on all six axes:
60
+
61
+ - **A / critical-path (holds):** In `review()` the `provenance` block is added to `opts` (`MessagingToneGate.ts:737-742`) and passed to `provider.evaluate(prompt, opts)` — but `IntelligenceRouter.mintDecision` clones the options and `delete internal.provenance` BEFORE any attempt (`IntelligenceRouter.ts:1138-1139`), so it never reaches the model. Settlement runs on the router's own path after the verdict (`:1110-1116`), and the single write is try/catch-contained — "the decision call is never failed or delayed by its audit trail" (`recordSettlement` :1288-1290); an errored exit re-throws the ORIGINAL error unchanged (`:1113-1116`), which the gate's own fail-closed/tier logic handles. `buildToneDecisionContext` is built synchronously outside the try block, but it is throw-proof: `crypto` (node:crypto, imported :21), `Buffer.byteLength`, `String().slice`, and `detectGateSignals` (itself fully try/catch-guarded, returns `[]` on any input — `GateSignalDetectors.ts:238-258`). A recorder failure/slow-write/throw cannot alter or hold the verdict — fail-open is real, not asserted.
62
+ - **B / body-leak (holds — load-bearing property intact):** `buildToneDecisionContext` (`MessagingToneGate.ts:128-160`) stores `candidate: {sha256, bytes, chars}` + `channel` (a 32-char slice of the CHANNEL name, not the body) + `messageKind` + `recentMessageCount` (a count) + `gateSignalKinds` (kind labels). Grepped the whole diff: the only `.slice` is on `context.channel`; no plaintext head/body field exists anywhere. The integration test drives a REAL gate→router→recorder→JPL→`GET /judgment-provenance` and asserts the distinctive body marker is absent from the served JSON AND `ctxParsed.candidate.head` is `undefined` (`messaging-tone-gate-provenance.test.ts:150-165`); the unit suite additionally proves a 1MB message yields a <512-byte envelope and an in-body secret never crosses. The router's `rawResponseHead` (`:1227-1230`) captures the MODEL's verdict JSON, never the candidate body.
63
+ - **C / volume valve (enforced, not cosmetic):** census is `budget:500` (`provenanceCoverage.ts:200`), and the recorder genuinely enforces it — `resolveJsonlDisposition` parses `budget:<N>`, counts today's rows, returns `'budget-dropped'` past the ceiling and bumps the `droppedByBudget` counter (`DecisionQualityRecorderImpl.ts:300-302, 338-346`) while still writing the ~250-byte SQLite row. The ratchet fails a `full` valve for this point (`not.toBe('full')` + `toMatch(/^budget:[1-9]\d*$/)`).
64
+ - **D / §5.5 sub-budget (no feedback edge):** `GRADE_PASS_POINTS = [DP_EXTERNAL_HOG_KILL_LEAVE]` (one point) → `perPointSubBudget(N,1)===N`, byte-identical to the prior single-point walk; the hog grading body was extracted verbatim into `gradeOnePoint`. The tone gate is explicitly NOT graded by this pass. The pass reads recorded rows + hog-store records and writes grades only — it touches no model/door/prompt/floor selection. Covered by 3 new sub-budget tests + the unchanged hog correctness tests.
65
+ - **E / ratchet monotonicity:** the `messaging-tone-gate::…::ACT-1193` line was removed from `PENDING_BASELINE`, so a regression to pending fails the shrink-only pending pin (a new unbaselined line) AND the dedicated `status==='wired'` assertion — two independent failures. No other census entry changed (4 wired = 3 first-customer + tone; 6 exempt unchanged).
66
+ - **F / dark-gate:** recording gates on the same `provenance.uniformSeam` flag #1458 uses (`DecisionQualityRecorderImpl.ts:195`, early-return `:212`) and the router no-ops when `getDecisionQualityRecorder()` is null. Gate-off = the inert `provenance` options block is never consumed; zero rows, zero behavior change.
67
+
68
+ Ran the three unit suites (47 passing, incl. the typed-import wired-verification) + the integration test (2 passing, body-marker-absent) + `tsc --noEmit` (clean). One cosmetic note (non-blocking): the summary line calls tone-gate the "4th wired customer" whereas the census/tests correctly call it the "third enrolled CUSTOMER" (customers key on `baseOf(component)`, so CompletionEvaluator + /P13 collapse to one) — the code and ratchet are internally consistent; only the artifact's prose is loose. No parallel mechanism; it rides #1458's seam exactly.