instar 1.3.1019 → 1.3.1021

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (62) hide show
  1. package/dist/commands/server.d.ts.map +1 -1
  2. package/dist/commands/server.js +81 -0
  3. package/dist/commands/server.js.map +1 -1
  4. package/dist/config/ConfigDefaults.d.ts.map +1 -1
  5. package/dist/config/ConfigDefaults.js +12 -0
  6. package/dist/config/ConfigDefaults.js.map +1 -1
  7. package/dist/core/componentCategories.d.ts.map +1 -1
  8. package/dist/core/componentCategories.js +2 -0
  9. package/dist/core/componentCategories.js.map +1 -1
  10. package/dist/core/devGatedFeatures.d.ts.map +1 -1
  11. package/dist/core/devGatedFeatures.js +6 -0
  12. package/dist/core/devGatedFeatures.js.map +1 -1
  13. package/dist/core/types.d.ts +16 -0
  14. package/dist/core/types.d.ts.map +1 -1
  15. package/dist/core/types.js.map +1 -1
  16. package/dist/data/llmBenchCoverage.d.ts.map +1 -1
  17. package/dist/data/llmBenchCoverage.js +10 -0
  18. package/dist/data/llmBenchCoverage.js.map +1 -1
  19. package/dist/data/provenanceCoverage.d.ts +4 -0
  20. package/dist/data/provenanceCoverage.d.ts.map +1 -1
  21. package/dist/data/provenanceCoverage.js +24 -0
  22. package/dist/data/provenanceCoverage.js.map +1 -1
  23. package/dist/memory/TopicMemory.d.ts +12 -0
  24. package/dist/memory/TopicMemory.d.ts.map +1 -1
  25. package/dist/memory/TopicMemory.js +49 -14
  26. package/dist/memory/TopicMemory.js.map +1 -1
  27. package/dist/messaging/TelegramAdapter.d.ts +1 -0
  28. package/dist/messaging/TelegramAdapter.d.ts.map +1 -1
  29. package/dist/messaging/TelegramAdapter.js +1 -0
  30. package/dist/messaging/TelegramAdapter.js.map +1 -1
  31. package/dist/messaging/shared/MessageLogger.d.ts +3 -0
  32. package/dist/messaging/shared/MessageLogger.d.ts.map +1 -1
  33. package/dist/messaging/shared/MessageLogger.js.map +1 -1
  34. package/dist/messaging/slack/SlackApiClient.d.ts.map +1 -1
  35. package/dist/messaging/slack/SlackApiClient.js.map +1 -1
  36. package/dist/monitoring/ClaimObservation.d.ts +35 -0
  37. package/dist/monitoring/ClaimObservation.d.ts.map +1 -1
  38. package/dist/monitoring/ClaimObservation.js +96 -9
  39. package/dist/monitoring/ClaimObservation.js.map +1 -1
  40. package/dist/monitoring/CompletionClaimVerifier.d.ts.map +1 -1
  41. package/dist/monitoring/CompletionClaimVerifier.js +4 -3
  42. package/dist/monitoring/CompletionClaimVerifier.js.map +1 -1
  43. package/dist/monitoring/GoalRealignment.d.ts +404 -0
  44. package/dist/monitoring/GoalRealignment.d.ts.map +1 -0
  45. package/dist/monitoring/GoalRealignment.js +1293 -0
  46. package/dist/monitoring/GoalRealignment.js.map +1 -0
  47. package/dist/server/CapabilityIndex.d.ts.map +1 -1
  48. package/dist/server/CapabilityIndex.js +9 -0
  49. package/dist/server/CapabilityIndex.js.map +1 -1
  50. package/dist/server/routes.d.ts.map +1 -1
  51. package/dist/server/routes.js +26 -0
  52. package/dist/server/routes.js.map +1 -1
  53. package/package.json +2 -1
  54. package/src/data/builtin-manifest.json +46 -46
  55. package/src/data/llmBenchCoverage.ts +10 -0
  56. package/src/data/provenanceCoverage.ts +30 -0
  57. package/src/data/state-coherence-registry.json +55 -0
  58. package/upgrades/1.3.1020.md +54 -0
  59. package/upgrades/1.3.1021.md +60 -0
  60. package/upgrades/periodic-goal-realignment-phase1.eli16.md +56 -0
  61. package/upgrades/side-effects/extraction-gap-signal-reason.md +254 -0
  62. package/upgrades/side-effects/periodic-goal-realignment-phase1.md +222 -0
@@ -0,0 +1,60 @@
1
+ # Upgrade Guide — vNEXT
2
+
3
+ <!-- assembled-by: assemble-next-md -->
4
+ <!-- bump: patch -->
5
+
6
+ ## What Changed
7
+
8
+ The claim-verification observer now records **why** a protected-cue span was booked as an extraction
9
+ gap, instead of only recording that one was. Per `docs/specs/claim-verification-sentinel.md` §2.1,
10
+ the deterministic protected-predicate lane is specified to emit
11
+ `ExtractionGapSignal {minimumCriticality, reason, span}`. That signal had never been built — the
12
+ shipped code returned a deduped list of bare cue-family names, so all three spec-distinguished causes
13
+ rendered as the same string.
14
+
15
+ Each gapped span now carries one of three reasons: `no-overlapping-claim` (the model extracted
16
+ nothing there), `unendorsed-overlap` (a claim *did* overlap, but only quoted/hedged/non-endorsed —
17
+ the extractor classified it correctly and was charged for a gap anyway), or `invalid-envelope`.
18
+
19
+ Two correctness fixes came with it:
20
+
21
+ - The cue scan stopped at the first matching candidate per family, so a later uncovered span was
22
+ never examined. It now scans every span, as §2.1 requires. Gap counts rise as a result — this is
23
+ a measurement fix, not an extractor regression.
24
+ - Audit rows written by this path are stamped `schemaVersion: 2`. **Gap rates must not be compared
25
+ across that boundary**: version-1 rows undercount and carry no reason.
26
+
27
+ ## What to Tell Your User
28
+
29
+ Nothing — there is no user-visible change, no new command, and no behavior difference in any
30
+ conversation. This is internal measurement plumbing for a feature that ships dark and runs in
31
+ dry-run: it observes and records, and holds no authority to block, delay, or alter any message.
32
+
33
+ If a user asks why it was worth doing: the observer had been running for days producing a number
34
+ nobody could interpret. About 89% of recorded gaps were the completion family, but with only a
35
+ family name stored there was no way to distinguish the observer genuinely missing a claim from the
36
+ observer being penalised for correctly marking something as hedged or quoted. The reason field
37
+ separates those two, which is what the accumulated data was missing.
38
+
39
+ ## Summary of New Capabilities
40
+
41
+ None for users. Internally, the claim observer's gap events now carry a reason, a byte span, and a
42
+ criticality floor per gapped span instead of a bare family name, and the audit rows that record them
43
+ are capped per family with an explicit marker whenever trimming occurs, so a trimmed sample can
44
+ never be mistaken for a complete one.
45
+
46
+ ## Evidence
47
+
48
+ - `tests/unit/extraction-gap-signal.test.ts` — 20 tests covering all three reasons, the no-gap case
49
+ (endorsed overlap), the criticality floor per cue family, span correctness, the every-span fix,
50
+ `gapKinds` back-compat, per-family truncation, and rejection of malformed caller payloads.
51
+ - Old-vs-new demonstrated directly: on `"The migration is done."` the previous algorithm emits
52
+ `["completion"]` for a hedged overlap **and** for a true no-overlap — identical output, different
53
+ cause. On `"The merge is done. The deploy is finished."` with the first sentence covered, it emits
54
+ `[]` and misses the genuine gap.
55
+ - Per-family truncation bug caught in second-pass review and reproduced before fixing: 72 signals on
56
+ a 24-candidate message, of which a flat head-slice kept only `capacity` and dropped `completion`
57
+ and `state` entirely.
58
+ - 36/36 green (20 new, 11 pre-existing `claim-observation-v1` untouched, 5 integration
59
+ `completion-claim-stats-route`).
60
+ - Side-effects review with second-pass: `upgrades/side-effects/extraction-gap-signal-reason.md`.
@@ -0,0 +1,56 @@
1
+ # ELI16 — A durable compass for long autonomous work (Phase 1)
2
+
3
+ ## What changed
4
+
5
+ Long autonomous runs now have an observation-only alignment reviewer. On development
6
+ agents, verified operator messages in an active run's topic feed a durable priority
7
+ ledger, and a periodic reviewer compares that ledger with the run's current goal and
8
+ unfinished tasks.
9
+
10
+ Phase 1 only sees and records. It does not inject advice into sessions, edit the run
11
+ plan, block work, create attention items, or notify the operator. The authenticated
12
+ `GET /goal-realignment` status surface shows the candidate inbox, active/closed
13
+ priorities, counters, and latest dry-run verdict.
14
+
15
+ ## Important lifetime rule
16
+
17
+ The recency window controls discovery of new priorities only. Once a verified
18
+ priority is recorded, it remains active until an explicit later operator message
19
+ supersedes it or clearly confirms it addressed. Silence and age never remove it.
20
+
21
+ The risky `diverged` verdict still requires validated evidence on both sides: an
22
+ authoritative priority citation and an exact contradictory or abandoning quote from
23
+ the current run focus. Missing, conflicting, or incomplete evidence becomes
24
+ `indeterminate`.
25
+
26
+ Incomplete is enforced before model review: truncated history, an unresolved
27
+ candidate extraction, or a digest projection that omits a live priority records
28
+ `indeterminate` at confidence zero and makes no reviewer call.
29
+
30
+ ## Durability and provenance
31
+
32
+ - Every deterministic instruction-shaped operator message enters a candidate inbox
33
+ before model classification.
34
+ - Extraction checkpoints persist the source cursor, exact raw provider output,
35
+ validated result, prompt ID, and model ID before any ledger event.
36
+ - A crash after the checkpoint reuses it, producing the same deterministic priority
37
+ ID without another model call.
38
+ - Telegram forwarded-message provenance now survives TopicMemory storage and
39
+ rebuilds. Legacy rows with unknown forwarded state remain ineligible.
40
+ - Both LLM decisions emit bounded, identity-only decision provenance; raw operator
41
+ text and run-focus bodies are not copied into the provenance archive.
42
+
43
+ ## Rollout
44
+
45
+ The feature is development-agent gated and structurally `dryRun: true`. Fleet agents
46
+ remain dark. Existing TopicMemory databases migrate automatically from schema 4 to
47
+ schema 5 by adding the nullable forwarded-provenance column. No operator action or
48
+ configuration change is required.
49
+
50
+ ## Validation
51
+
52
+ Refusal-first coverage pins the three acceptance cases: an old standing priority
53
+ survives the discovery window, quoted-only instructions require operator
54
+ confirmation, and crash replay reuses the checkpoint and priority ID. Unit,
55
+ integration, end-to-end, type, lint, and build results are recorded in the Phase 1
56
+ validation artifact.
@@ -0,0 +1,254 @@
1
+ # Side-Effects Review — ExtractionGapSignal reason + span (spec §2.1 conformance)
2
+
3
+ **Version / slug:** `extraction-gap-signal-reason`
4
+ **Date:** `2026-07-27`
5
+ **Author:** `echo`
6
+ **Second-pass reviewer:** `required — sentinel/observability surface (see Phase 5 section)`
7
+
8
+ ## Summary of the change
9
+
10
+ `docs/specs/claim-verification-sentinel.md` §2.1 specifies that the deterministic
11
+ protected-predicate lane emits `ExtractionGapSignal {minimumCriticality, reason, span}` when a
12
+ protected cue "has no overlapping endorsed claim, has only quoted/hedged/non-endorsed overlap, or
13
+ the envelope is invalid." That signal was never built: `git grep ExtractionGapSignal` returned zero
14
+ hits across `src/` and `tests/`. The shipped `protectedCueGaps()` returned a deduped list of bare
15
+ cue-family names, so all three spec-distinguished causes rendered as the same string.
16
+
17
+ This change implements the signal. `extractionGapSignals()` (in `src/monitoring/ClaimObservation.ts`)
18
+ scans every candidate span for every cue family and returns one signal per gapped span carrying the
19
+ reason, the byte span, and the §2.5 criticality floor. `protectedCueGaps()` is retained, now derived
20
+ from those signals, so existing counters and the `gapKinds` audit field are unchanged.
21
+ `CompletionClaimVerifier.observe()` writes the new `gapSignals` array alongside `gapKinds`, and
22
+ `ClaimObservationRecorder.recordEvent()` gains a strictly-clamped, per-family-capped allowlist
23
+ branch for it, stamping `schemaVersion: 2`.
24
+
25
+ `gapKinds` keeps its shape and its consumers, but its *values* are monotone-increasing — see §5.
26
+
27
+ Files touched: `src/monitoring/ClaimObservation.ts`, `src/monitoring/CompletionClaimVerifier.ts`,
28
+ `tests/unit/extraction-gap-signal.test.ts` (new).
29
+
30
+ **Why now:** the missing `reason` is the measurement blocker that held EVO-005 across eleven review
31
+ cycles. 89.1% of gap events (295/331 in the live audit) are the `completion` family, and with only a
32
+ family name recorded there is no way to tell a real extractor miss from the extractor being charged
33
+ for a span it correctly marked hedged or quoted. Demonstrated concretely: under the old algorithm a
34
+ hedged overlap and a true no-overlap both emit exactly `["completion"]`.
35
+
36
+ ## Decision-point inventory
37
+
38
+ - `protectedCueGaps` / `extractionGapSignals` (`src/monitoring/ClaimObservation.ts`) — **modify** —
39
+ deterministic cue detector. Produces observations only. Holds no blocking authority before or
40
+ after this change.
41
+ - `CompletionClaimVerifier.observe` audit emission — **modify** — adds a field to an existing
42
+ dark, dry-run audit row. No change to any return value that a caller branches on.
43
+ - `ClaimObservationRecorder.recordEvent` allowlist — **modify** — adds one clamped structural field.
44
+
45
+ ---
46
+
47
+ ## 1. Over-block
48
+
49
+ **No block/allow surface — over-block not applicable.** The signal never gates a message. The
50
+ feature is dark (`dryRun: true` in production today) and §2.1 states the shadow signal "never
51
+ creates a factual claim or verdict." `observe()`'s return contract is untouched.
52
+
53
+ The nearest analogue to an over-block is over-*counting*: the fixed scan now books gaps the old
54
+ first-match-then-break loop skipped, so gap counts rise. That is the correct count per §2.1 ("every
55
+ protected-cue span"), but see §5 for the baseline discontinuity it creates.
56
+
57
+ ## 2. Under-block
58
+
59
+ **No block/allow surface — under-block not applicable.**
60
+
61
+ What the change still does *not* measure, stated plainly:
62
+
63
+ - **Cue over-breadth is untouched.** The `completion` regex fires on the noun "fix" (983 occurrences
64
+ in the local message corpus, the single dominant token) and on instar's own term of art "ships
65
+ dark", which means deliberately DISABLED — the opposite of a completion claim. Those produce
66
+ `no-overlapping-claim` signals indistinguishable from genuine extractor misses. Separating them
67
+ requires recording the matched cue token, which is **out of scope here and deliberately so**: §2.6
68
+ fixes a closed audit field list and forbids raw claim text and plain content hashes, so a cue-token
69
+ field is a spec change needing its own convergence and approval.
70
+ <!-- tracked: ACT-1433 -->
71
+ - `reason` therefore quantifies the endorsement-filter share of the gap rate exactly, and leaves the
72
+ cue-over-breadth share unresolved. It narrows the ambiguity; it does not eliminate it.
73
+ - The `injection` family is assigned a `high` floor rather than its own tier. §2.1 names only
74
+ approval/capacity/completion/credential for `high` and consequential-action premises for
75
+ `irreversible-precondition`; injection markers are conservatively floored at `high` because
76
+ uncertainty rounds up per §2.3.
77
+ - **Per-family truncation.** A single turn emitting more than 8 gapped spans in one family keeps the
78
+ first 8 of that family and sets `gapSignalsTruncated: true`. Within-family reason shares are
79
+ therefore biased toward earlier spans on those (rare) rows. The truncation is recorded, so such
80
+ rows can be excluded from a reason-share computation rather than silently skewing it.
81
+
82
+ ## 3. Level-of-abstraction fit
83
+
84
+ Correct layer. The detector already lives in `ClaimObservation.ts` and already computed the overlap
85
+ relation this change reports — the information was being discarded one line after being derived, not
86
+ gathered somewhere new. No higher layer could reconstruct it, because only this function sees both
87
+ the cue span and the claim's endorsed/quoted/hedged flags at the same time.
88
+
89
+ This feeds the existing audit lane rather than creating a parallel one, per §2.6's "one bounded
90
+ origin-local append/read/rotation path."
91
+
92
+ ## 4. Signal vs authority compliance
93
+
94
+ **Compliant.** Reference: `docs/signal-vs-authority.md`.
95
+
96
+ This is a detector producing a strictly richer structured signal. It gains no blocking power, adds
97
+ no new code path that can refuse or alter a message, and does not become an input to any authority
98
+ in this change. Per §2.1 the gap signal "is the coarse non-LLM recall floor: a high-criticality
99
+ observation, not a normalized factual claim, and therefore cannot support/refute."
100
+
101
+ The `minimumCriticality` field is worth calling out specifically: it is a *label carried on an
102
+ observation*, not a threshold anything acts on. Nothing in this change reads it back.
103
+
104
+ ## 5. Interactions
105
+
106
+ - **`gapKinds` shape preserved; values monotone-increasing.** `protectedCueGaps()` keeps returning
107
+ deduped family names and still populates `gapKinds`. But old ⊆ new strictly: the old loop pushed a
108
+ family iff the *first* pattern-matching candidate was uncovered, the new one iff *any* is, so
109
+ `gapKinds` can gain a family and can never lose one. The existing `toContain('capacity')` assertion
110
+ is safe and all 11 pre-existing tests pass unmodified — but "unchanged" would be wrong, and callers
111
+ comparing counts across the boundary must not treat the two as the same measurement.
112
+ - **`coverageIncompleteTurns` shifts too.** `CompletionClaimVerifier` bumps it on `gaps.length > 0`,
113
+ and `gaps` can now be non-empty on turns where it previously was not. That counter moves in the
114
+ same direction and for the same reason, and is not a regression.
115
+ - **Baseline discontinuity — the real interaction risk.** Anyone comparing gap rates across this
116
+ deploy sees a step change that is a measurement fix, not an extractor regression. **Split on
117
+ `schemaVersion`**: pre-change rows are `1`, post-change rows are `2`. Key-absence was the first
118
+ proposed marker and is wrong — it only works via a coupling (the verifier happens to write the row
119
+ only when signals exist), and `recordEvent` will legitimately write a `gapSignals`-free row, which
120
+ a dedicated test asserts. The version stamp is the durable marker.
121
+ - **`schemaVersion` is per-row-TYPE, not per-file.** `recordEvent` rows are now `2`, while
122
+ `record()` and `recordAuthoritativeOutcome` keep writing `1` into the same
123
+ `claim-observation-audit-v2.jsonl`. Splitting gap rows on `schemaVersion >= 2` is correct because
124
+ gap rows are event rows; an analyst splitting the *whole file* will find claim-observation rows
125
+ stay `1` forever, which is expected and not a bug.
126
+ - **Truncation cannot silently bias the sample.** The per-family cap replaced a flat head-slice that
127
+ would have dropped whole trailing families. Because emission is family-outer, that slice kept only
128
+ `capacity` on a busy message and discarded `completion` (the family the statistic is about) and
129
+ `injection` (the adversarial family) entirely — reproduced at 72 signals → 1 surviving family
130
+ before the fix, and covered by a regression test now.
131
+ - **No double-fire.** The audit row is emitted once per turn under the existing `gaps.length > 0`
132
+ guard, which is unchanged.
133
+ - **Shared metric untouched.** `this.bump('protectedCueGaps')` still iterates deduped kinds, so the
134
+ `/metrics/features` magnitude is not redefined by the richer signal list.
135
+
136
+ ## 6. External surfaces
137
+
138
+ No route, no schema served to another agent, no user-visible surface. The audit file is local and
139
+ already exists. `/completion-claim/stats` is unaffected — its integration test passes unmodified.
140
+
141
+ Privacy: the new field carries a cue-family name, a reason enum, a criticality enum, and two integer
142
+ byte offsets. No message text.
143
+
144
+ **On whether byte offsets sit inside the §2.6 boundary — the honest argument.** An earlier draft of
145
+ this artifact claimed the offsets are "the same class already carried by
146
+ `sourceStartByte`/`sourceEndByte` throughout §2.6." That is false and has been removed: in §2.6 those
147
+ offsets appear only as HMAC inputs to `claimId`, never as a persisted plaintext field, so this is the
148
+ first plaintext persistence of that value. The real basis is threefold:
149
+
150
+ 1. The §2.6 closed field list is already pervasively exceeded by shipped code — `recordEvent` today
151
+ writes `ts, evaluated, flagged, event, verdict, actionKind, hadToolCalls, reason, gapKinds`, none
152
+ of which are in the list, and `record()` persists `actualLatencyMs`, a plaintext un-bucketed
153
+ numeric of exactly the class at issue. `gapSignals` introduces no *new* class of deviation.
154
+ 2. The forbidden categories are enumerated — raw text, paths, URLs, identities, secrets, credentials,
155
+ commands/results, free-form rationale, plain content hashes. A byte offset is none of them.
156
+ 3. The §2.6 re-identification concern is about *joining* shape/model/timing/pseudonym. These
157
+ `protected-cue-unextracted` rows carry no join key at all — no `messagePseudonym`,
158
+ `topicPseudonym`, or `claimId` — so the offsets are not joinable to a message or topic.
159
+
160
+ Residual disclosure is real and worth naming: an offset pair is a coarse sentence-length/position
161
+ fingerprint in mode-0600 local operational data. It sits strictly below the existing
162
+ `actualLatencyMs` bar. If a reviewer disagrees that (1) is acceptable, the correct response is a
163
+ spec-level cleanup of the closed list, not a carve-out for this field.
164
+
165
+ Clamping: `recordEvent` clamps every field, drops unknown keys, and caps per family — covered by a
166
+ test asserting a planted `secretField` never reaches disk. **This guarantee is conditional on
167
+ `opts.recorder` being set.** `CompletionClaimVerifier.appendAudit` has a fallback path that writes
168
+ the raw row to `logs/completion-claim-audit.jsonl` with no clamping. That is pre-existing (`gapKinds`
169
+ has identical exposure) and carries no live risk here because `gapSignals` is internally produced,
170
+ never caller-supplied — but the guarantee is not unconditional and should not be stated as such.
171
+
172
+ ## 7. Multi-machine posture (Cross-Machine Coherence)
173
+
174
+ **Machine-local BY DESIGN**, and strictly more so than the pooled-rows posture would require.
175
+
176
+ `gapSignals` never enters the pool projection at all: `readPoolAggregates`/`readPoolPage` read the
177
+ corpus path, not the audit path, and the peer merge in `routes.ts` is an explicit field allowlist
178
+ (`day, claimShapeId, modelDoor, verifierVersion, count, verdicts`). So there is no pool-side surface
179
+ to leak — a reviewer should not go looking for one. (An earlier draft said the field "inherits the
180
+ pooled untrusted-display posture," which invited exactly that wasted search.)
181
+
182
+ The consequence for measurement: a gap-reason figure is per-machine and cannot be assembled across
183
+ the pool through any existing read. No user-facing notice, no durable state that strands on topic
184
+ transfer, no generated URL.
185
+
186
+ ## 8. Rollback cost
187
+
188
+ **Low.** The change is additive and observation-only. Reverting the commit restores the prior
189
+ function body; no migration, no agent state repair, no data cleanup. Existing audit rows remain
190
+ readable either way because `gapSignals` is an optional key that consumers must already tolerate the
191
+ absence of (every historical row lacks it). No release is gated on it and nothing reads the field
192
+ back yet.
193
+
194
+ ---
195
+
196
+ ## Phase 5 — second-pass review
197
+
198
+ Required: the change touches a sentinel/observability surface in the claim-verification family.
199
+
200
+ **Reviewer verdict: "Concern raised" — one blocking finding, since resolved; several should-fixes,
201
+ all folded into the sections above.**
202
+
203
+ **Blocking (fixed).** The persisted `gapSignals` array used a flat `.slice(0, 24)`. Because
204
+ `extractionGapSignals` emits family-outer, that silently dropped whole trailing families —
205
+ reproduced on a 24-candidate message at 72 signals where only `capacity` survived and `completion`
206
+ and `state` were discarded entirely. The bias ran precisely against the change's purpose: the
207
+ motivating figure is that ~89% of gaps are the `completion` family, and `completion` was the first
208
+ thing dropped. Resolved with a per-family cap of 8 plus a `gapSignalsTruncated` marker, and a
209
+ regression test. **This is the finding that justifies the phase** — the change would otherwise have
210
+ shipped a measurement instrument biased in the one direction guaranteed to make it useless, and the
211
+ bias would not have shown up in any of the tests written before the review.
212
+
213
+ **Should-fixes, all applied:** the §2.6 span justification was factually wrong and has been rewritten
214
+ with the real argument (§6); `gapKinds` was described as "unchanged" when it is monotone-increasing
215
+ (§5, summary); the `coverageIncompleteTurns` shift was unlisted (§5); `schemaVersion` was not bumped
216
+ and the proposed key-absence split marker rested on a coupling (§5, now `schemaVersion: 2`); the
217
+ multi-machine section overstated exposure (§7); the clamping guarantee is conditional on
218
+ `opts.recorder` (§6); and two `§2.5` citations were actually `§2.3` (fixed in both the artifact and
219
+ the code comment).
220
+
221
+ **Second round — reviewer verdict: "Concur."** The blocking finding was independently re-verified
222
+ against the staged code (72 signals produced → 24 persisted as 8/8/8 across all three families,
223
+ `gapSignalsTruncated: true`, `schemaVersion: 2`). The reviewer agreed with capping per-family at 8
224
+ rather than raising to 144, on the grounds that the statistic is a per-family reason *share*, so
225
+ preserving the family distribution is the property that matters, and that the residual within-family
226
+ position bias is *recoverable* because truncated rows are marked and can be excluded.
227
+
228
+ **One new finding, introduced by the blocking fix, since resolved.** Replacing the flat
229
+ `.slice(0, 24)` removed the *unconditional* array bound: the per-family cap alone bounds the array
230
+ at 8 × distinct kinds, and `kind` is caller-supplied. Probed at 5,000 invented kinds → 5,000
231
+ persisted objects in a ~559KB row, with `gapSignalsTruncated` correctly false because nothing was
232
+ dropped per family. Not reachable from the in-tree producer (6 families, ≤48 objects), but
233
+ `recordEvent` is deliberately an untrusted-input boundary and `appendBounded` rotates-then-writes
234
+ instead of rejecting an oversized row, so an unbounded row would cost audit *history*. Resolved with
235
+ an unconditional `.slice(0, 48)` that also sets the truncation marker, plus a hostile-input test.
236
+
237
+ **Reviewer answers to the two questions posed:**
238
+ - **(a) Span inside the privacy boundary:** yes, but not for the reason originally given. Basis
239
+ rewritten in §6.
240
+ - **(b) Raising counts mid-soak:** correct; do not flag-preserve the old behavior. The old scan is a
241
+ recall bug, not a baseline — §2.7 sets deterministic protected-gap recall at 1.0 on closed cue
242
+ fixtures, which the broken scan structurally cannot meet. Preserving it would keep accumulating a
243
+ measurement known to be false in a known direction. Old ⊆ new is provable, so the discontinuity is
244
+ monotone and interpretable.
245
+
246
+ **Open question the review surfaced, not resolved here:** whether the §2.7 readiness clock (≥1,000
247
+ admitted messages / 200 settled T0 claims over 14 days) restarts for gap-rate metrics, given that
248
+ this changes the definition of the measured quantity mid-window. It plausibly should. That is a
249
+ judgment about the soak's validity, not about this code, and belongs to whoever reads the re-soak.
250
+ <!-- tracked: ACT-1433 -->
251
+
252
+ **Confirmed clean by the reviewer:** no blocking authority (nothing reads `gapSignals` or
253
+ `minimumCriticality` back — the sole consumer maps to `kind` and discards the rest);
254
+ level-of-abstraction fit; rollback cost; no double-fire; and the §2 cue-over-breadth disclosure.
@@ -0,0 +1,222 @@
1
+ # Side-Effects Review — Periodic Goal Re-Alignment Phase 1
2
+
3
+ **Version / slug:** `periodic-goal-realignment-phase1`
4
+ **Date:** `2026-07-27`
5
+ **Author:** `Instar-codey`
6
+ **Second-pass reviewer:** `not required`
7
+
8
+ ## Summary of the change
9
+
10
+ This change adds the observation-only first slice of periodic goal realignment:
11
+ Telegram operator provenance survives `TopicMemory`, verified instruction-shaped
12
+ messages enter a durable candidate inbox, semantic extraction is checkpointed before
13
+ append-only priority events, and a cadence-shaped reviewer writes dry-run verdicts
14
+ visible through authenticated `GET /goal-realignment`. It changes
15
+ `TopicMemory`, Telegram/shared logging, config/types/capability/route registration,
16
+ server wiring, a new `GoalRealignment` module, tests, and the approved spec's single
17
+ operator-amended lifetime decision. There is no injection or action path.
18
+
19
+ ## Decision-point inventory
20
+
21
+ - `detectCandidatePriority` — **add** — recall-biased deterministic detector decides
22
+ which verified messages enter the holding list; it cannot create authority or
23
+ retire a priority.
24
+ - `GoalRealignmentIntake` provenance eligibility — **add** — hard boundary requiring
25
+ exact authenticated operator UID plus explicit `forwarded:false`.
26
+ - `PriorityExtraction` — **add** — context-rich semantic classification proposes
27
+ priority/restatement/supersession/completion states, with exact authored grounding
28
+ and conservative confirmation thresholds.
29
+ - `AlignmentReviewer` — **add** — semantic signal producer labels current run focus;
30
+ `diverged` is mechanically downgraded to `indeterminate` without valid two-sided
31
+ evidence.
32
+ - `GoalDigestBuilder` — **add** — deterministic materialized projection excludes
33
+ only explicitly superseded or confirmed-addressed priorities; age never retires
34
+ authority.
35
+
36
+ ---
37
+
38
+ ## 1. Over-block
39
+
40
+ No message-delivery or work-action block surface exists. A legitimate operator
41
+ instruction that lacks every deterministic signal can fail to enter the Phase 1
42
+ candidate inbox; this is an observation false negative, not a blocked user action.
43
+ The detector intentionally recognizes broad imperative, priority, status, and
44
+ confirmation language, while the dry-run status counters and inbox make recognized
45
+ classification failures visible.
46
+
47
+ Legacy `TopicMemory` rows whose forwarded state is unknown are excluded from
48
+ authority. That can omit a legitimate historical operator message, but accepting an
49
+ unknown forwarded row would permit quoted third-party content to become authority.
50
+ New ingress persists explicit provenance.
51
+
52
+ ---
53
+
54
+ ## 2. Under-block
55
+
56
+ No block surface exists. Remaining observation misses include implicit priorities
57
+ without instruction-shaped language, paraphrased completion that the extractor
58
+ cannot ground to an exact authored substring, and active-run focus expressed outside
59
+ the registered condition/goal/task rows. These fail toward missing signal or
60
+ `indeterminate`, never action.
61
+
62
+ The source-history recovery read is bounded to 500 recent rows. If it reaches the
63
+ bound it returns `complete:false` and performs no partial reconciliation, preventing
64
+ false completeness but leaving the candidate backlog dependent on live intake until
65
+ the source volume is reduced or a later implementation provides pagination.
66
+
67
+ ---
68
+
69
+ ## 3. Level-of-abstraction fit
70
+
71
+ The deterministic matcher is correctly a low-level detector feeding a durable inbox,
72
+ not an authority. Semantic extraction is correctly delegated to the registered
73
+ `GoalPriorityExtractor` reflector with existing priorities and authored/quoted
74
+ separation. The `AlignmentReviewer` is also a registered reflector and produces only
75
+ a stored observation. Mechanical citation validation sits below the model because
76
+ exact identity, substring, enum, and size checks are enumerable invariants.
77
+
78
+ The coordinator composes existing `TopicMemory`, `TopicOperatorStore`,
79
+ `AutonomousRunStore`, shared `LlmQueue`, dev-agent gating, and capability routing. It
80
+ does not introduce a parallel operator identity store, run state writer, or general
81
+ workflow engine.
82
+
83
+ ---
84
+
85
+ ## 4. Signal vs authority compliance
86
+
87
+ **Required reference:** [docs/signal-vs-authority.md](../../docs/signal-vs-authority.md)
88
+
89
+ - [x] No — this change has no block/allow surface.
90
+
91
+ The candidate detector and both model-backed components emit structured evidence.
92
+ They cannot block work, mutate the planner/run file, inject session context, notify
93
+ the operator, or trigger recovery. `dryRun: true` is structural: the reviewer has no
94
+ injection dependency to call even if configuration is accidentally changed.
95
+
96
+ ---
97
+
98
+ ## 4b. Judgment-point check (Judgment Within Floors standard)
99
+
100
+ The candidate regex is not used at a competing-signals authority point; it only
101
+ admits additional rows into a holding list. Semantic priority meaning and alignment
102
+ remain reflector judgments. Exact sender/forwarded provenance, schema validation,
103
+ idempotency, and two-sided citation grounding are enumerable invariant floors named
104
+ in the driving spec's decision-point table.
105
+
106
+ ---
107
+
108
+ ## 5. Interactions
109
+
110
+ - **Shadowing:** the new `onMessageLogged` handler chains the pre-existing callback
111
+ before enqueueing intake. It does not replace TopicMemory, Presence Proxy, topic
112
+ intent capture, Usher, or correction learning callbacks.
113
+ - **Independent boot:** initial wiring accidentally sat inside Presence Proxy's
114
+ initialization scope. Review moved it to its own Telegram/dev-gated boot path and
115
+ a source-level test pins ordering before `let presenceProxy`.
116
+ - **Double-fire:** live intake and startup history reconciliation can see the same
117
+ message. The source-derived idempotency key and checkpoint/event IDs collapse both
118
+ paths to one extraction/event.
119
+ - **Races:** coordinator intake is serialized through `intakeTail`; review awaits the
120
+ tail and is guarded by `reviewInFlight`. Atomic runtime writes precede event
121
+ application.
122
+ - **Feedback loops:** reviewer output writes only its own audit/runtime status. It
123
+ never enters the autonomous state file or message history, so it cannot become its
124
+ own focus or source evidence.
125
+ - **LLM contention:** both reflectors use the existing shared background queue.
126
+ Unchanged digest+focus hashes reuse the prior verdict with zero model calls.
127
+
128
+ ---
129
+
130
+ ## 6. External surfaces
131
+
132
+ - New authenticated pull-only `GET /goal-realignment`; no mutation/action endpoint.
133
+ - `TopicMemory` schema moves from 4 to 5 with one nullable `forwarded` column and a
134
+ tested automatic migration. Existing unknown values remain safely unknown.
135
+ - New permission-restricted local stores:
136
+ candidate/checkpoint runtime, append-only priority events, and a scrubbed bounded
137
+ verdict log. The state registry declares resolution, compliance/archive, and
138
+ rotating-log retention respectively.
139
+ - Development agents resolve the omitted enable flag live; fleet agents resolve dark.
140
+ - No Telegram/Slack/attention notices, generated URLs, or operator actions are added.
141
+ Mobile-complete operator actions are therefore not applicable.
142
+ - LLM/provider timing can fail. Failures increment pull-visible counters and leave a
143
+ recognized candidate pending instead of silently classifying it away.
144
+
145
+ ## 6b. Operator-surface quality (Operator-Surface Quality standard)
146
+
147
+ No dashboard renderer, approval form, grant/revoke page, or other operator action
148
+ surface is touched. The only surface is an authenticated JSON status read, so this
149
+ section is not applicable.
150
+
151
+ ---
152
+
153
+ ## 7. Multi-machine posture (Cross-Machine Coherence)
154
+
155
+ **Machine-local BY DESIGN for Phase 1 dry-run observation.** The ledger is local to
156
+ the machine that owns the active autonomous run and has authenticated Telegram
157
+ history. Phase 1 neither speaks nor controls work, so it cannot double-notify or
158
+ create cross-machine authority. `GET /goal-realignment` reports this machine's
159
+ diagnostic state only.
160
+
161
+ This slice emits no user-facing notices and generates no URLs. Durable rows can stay
162
+ on the original machine if a topic transfers; they are not presented as a pool-wide
163
+ answer, and active transfer delivery is explicitly outside this observation-only
164
+ slice. The full approved design requires router-authoritative source reads,
165
+ owner-fenced replicated correctness metadata, and a transfer freshness barrier
166
+ before any later session-facing behavior is eligible. Phase 1's local evidence does
167
+ not grant that later authority.
168
+
169
+ ---
170
+
171
+ ## 8. Rollback cost
172
+
173
+ - **Hot-fix release:** disable explicitly with
174
+ `monitoring.goalRealignment.enabled:false` or revert the code and ship a patch.
175
+ - **Data migration:** the added nullable SQLite column is backward-compatible and
176
+ need not be removed. New goal-realignment state is isolated under its own directory
177
+ and log.
178
+ - **Agent state repair:** none required. Keeping the append-only evidence is safer
179
+ than deleting it; a later re-enable can reuse it.
180
+ - **User visibility:** none during rollback because Phase 1 sends nothing.
181
+
182
+ ---
183
+
184
+ ## Conclusion
185
+
186
+ The review found and fixed two substantive integration issues before commit: runtime
187
+ wiring was coupled to Presence Proxy, and production checkpoints persisted only a
188
+ re-serialized parsed result instead of exact provider output. It also added explicit
189
+ confirmed-addressed evidence, bounded resolved inbox/checkpoint state, archival
190
+ priority segments, and rotating verdict logs. With those corrections, the change is
191
+ a dev-gated, pull-visible signal producer with no action authority and is clear for
192
+ Tier 2 validation and ship.
193
+
194
+ ---
195
+
196
+ ## Second-pass review (if required)
197
+
198
+ **Reviewer:** not required
199
+ **Independent read of the artifact:** not required
200
+
201
+ The change does not touch outbound/inbound blocking, dispatch, session lifecycle,
202
+ recovery, a guard/gate/sentinel/watchdog, or other Phase-5 trigger. Its periodic loop
203
+ computes a bounded observation and cannot actuate.
204
+
205
+ ---
206
+
207
+ ## Evidence pointers
208
+
209
+ - `tests/unit/goal-realignment-phase1.test.ts`
210
+ - `tests/unit/topic-memory-forwarded-provenance.test.ts`
211
+ - `tests/integration/goal-realignment-routes.test.ts`
212
+ - `docs/specs/reports/periodic-goal-realignment-phase1-validation.md`
213
+
214
+ ---
215
+
216
+ ## Class-Closure Declaration (display-only mirror)
217
+
218
+ No agent-authored-artifact defect is being fixed. The cadence-shaped reviewer is not
219
+ an `unbounded-self-action` controller: it does not restart, swap, respawn, spawn,
220
+ notify, retry delivery, re-drive, or kill. Its only self-triggered effect is a
221
+ dry-run observation; unchanged semantic input spends zero calls, overlapping ticks
222
+ singleflight, and provider failure records a counter rather than self-retrying.