instar 1.3.990 → 1.3.992

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,178 @@
1
+ # Side-Effects Review — an alignment score that can say "not assessed"
2
+
3
+ **Version / slug:** `alignment-score-not-assessed`
4
+ **Date:** `2026-07-26`
5
+ **Author:** `Echo (instar-dev agent)`
6
+ **Second-pass reviewer:** `see Phase 5`
7
+
8
+ ## Summary of the change
9
+
10
+ `IntentDriftDetector.alignmentScore()` returned `score: 0, grade: 'F'` when the analysis window
11
+ contained no decisions. The `summary` field was honest ("No decisions logged — alignment cannot be
12
+ assessed") but no consumer read it; `score` and `grade` are what get rendered and compared. So
13
+ "nothing to assess" and "assessed, catastrophically bad" were identical on every field in use — on
14
+ the instrument whose purpose is honest alignment measurement.
15
+
16
+ Root cause is vocabulary, the same shape as the channel registry one increment earlier: the grade
17
+ union was `'A'|'B'|'C'|'D'|'F'` with no member meaning *no verdict*, so absence had to borrow the
18
+ worst real grade.
19
+
20
+ Adds `'N/A'` to the grade union, adds `assessable: boolean`, and makes `instar intent drift` print
21
+ the reason instead of a fabricated grade.
22
+
23
+ ## Refusal evidence (constraint 2)
24
+
25
+ ```
26
+ REFUSAL 1 — restore the fabricated 'F' on the unassessable case
27
+ × handles empty journal — not assessable, grade N/A → expected 'F' to be 'N/A'
28
+ × GET /intent/alignment reports NOT ASSESSED → expected 'F' to be 'N/A'
29
+ × an empty journal grades 'N/A', never 'F' → expected 'F' to be 'N/A'
30
+ Tests 3 failed | 25 passed (28)
31
+
32
+ REFUSAL 2 — always report assessable:true
33
+ × an empty journal is flagged unassessable → expected true to be false
34
+ × a real assessment is DISTINGUISHABLE from an empty one → expected true not to be true
35
+ × assessable tracks sampleSize exactly → expected true to be false
36
+ (+2) Tests 5 failed
37
+
38
+ REFUSAL 3 — disable the CLI's honest branch (`if (false && !alignment.assessable)`)
39
+ BEFORE the CLI test existed: Tests 28 passed (28) <-- the blindness, again
40
+ AFTER: × a STALE journal prints "not assessed", never a red F
41
+ Tests 1 failed | 4 passed (5)
42
+ ```
43
+
44
+ Restored: **42 passed** across the five affected files, `tsc --noEmit` exit 0.
45
+
46
+ **REFUSAL 3 is the finding, and it is the THIRD occurrence of this class tonight** (#1658 route
47
+ registry, #1659 route validator, now a CLI renderer). Each time the logic was thoroughly guarded and
48
+ the wiring to the surface a human or API client actually reads was not. This one is the sharpest:
49
+ the module returning `'N/A'` is worth precisely nothing if the renderer ignores it, and 28 green
50
+ tests said everything was fine while the renderer was disabled.
51
+
52
+ ## Two of my own claims were falsified during this work
53
+
54
+ Recorded because the corrections are the useful part, and because an artifact that hides them is the
55
+ failure mode this tier exists to remove.
56
+
57
+ 1. **"`instar intent drift` has been showing a red F for the journal's whole life."** FALSE. The
58
+ command returns early with a genuinely helpful message when the window holds no decisions; it
59
+ never reaches the scoring block. The empty case was already handled honestly there.
60
+ 2. **"Then it is reachable whenever the journal is merely stale."** ALSO FALSE. The early return
61
+ checks `windowDays` (default 14) and `alignmentScore()` is fixed at 30 — and 14 ⊂ 30, so anything
62
+ clearing the early return is inside the alignment window by construction.
63
+
64
+ **The actual reachable CLI case is narrow:** the operator must widen the window past 30
65
+ (`--window 60`), so a 40-day-old decision clears the early return and falls outside the fixed 30-day
66
+ alignment window. That is what the regression test constructs. I asserted twice before checking; the
67
+ test is what settled it.
68
+
69
+ ## Decision-point inventory
70
+
71
+ | point | classification | note |
72
+ |---|---|---|
73
+ | `sampleSize === 0` → `grade: 'N/A'`, `assessable: false` | `invariant` | Deterministic count check. No model. |
74
+ | `assessable` mirrors `sampleSize > 0` | `invariant` | Asserted by test; the two can never disagree. |
75
+ | CLI branches on `assessable` | `invariant` | Renders `summary` instead of a grade. |
76
+
77
+ No judgment points, no LLM, nothing gated or blocked.
78
+
79
+ ## 1. Over-block
80
+
81
+ Nothing is blocked — this is a read surface. The available harm is **misinforming a reader**, and
82
+ this change strictly reduces it in the direction that mattered (absence no longer reads as failure).
83
+
84
+ The mirror over-block is real and guarded: a genuinely-assessed period must never report `'N/A'` or
85
+ `assessable: false`, or a real alignment problem would be hidden as "no data" — strictly worse than
86
+ the original bug. Asserted by two tests (`a genuinely assessed period still reports a real letter
87
+ grade`, and the CLI's `a populated journal still prints a real graded score`).
88
+
89
+ **Caller sweep, run BEFORE writing this section** (the correction from #1659, where I wrote a
90
+ confident risk claim from the wrong measurement and CI falsified it): `alignmentScore()` has exactly
91
+ two production callers — `routes.ts:24305` (passes through verbatim) and `commands/intent.ts:431`
92
+ (now branches). Every `.grade` hit elsewhere in `src/` belongs to `DecisionQualityRecorder`'s
93
+ unrelated `right|wrong|unknown` grade, checked rather than assumed. Two tests asserted the old shape;
94
+ both updated.
95
+
96
+ ## 2. Under-block
97
+
98
+ **The mismatched windows are NOT fixed.** The early return uses `windowDays` while `alignmentScore()`
99
+ is hardcoded to 30. That divergence is what makes the CLI case reachable at all, and reconciling them
100
+ changes what the command reports for every user. Deliberately not folded in. <!-- tracked: CMT-1044 -->
101
+
102
+ **`score: 0` is retained on the unassessable case.** Changing it to `null` would be a breaking type
103
+ change for a field two consumers read; `assessable` is the additive signal instead. A consumer that
104
+ reads `.score` and ignores both `assessable` and `sampleSize` still sees a 0 — it can no longer see
105
+ an F, which is the part that read as a verdict.
106
+
107
+ **It does not improve alignment,** and it does not judge whether a cited principle was genuinely
108
+ consulted. Same honest limit as the increment before it.
109
+
110
+ ## 3. Level-of-abstraction fit
111
+
112
+ The honest state lives in the returned value, not in the renderer, so every consumer inherits it —
113
+ the route needed no change at all. The renderer's job is narrowed to *presenting* a state it no
114
+ longer has to infer. Had I fixed only the CLI, the API consumer (the one that is actually reachable
115
+ by default) would still have been lied to.
116
+
117
+ ## 4. Signal vs authority compliance
118
+
119
+ Pure signal. `docs/signal-vs-authority.md` is satisfied trivially: it produces a read-only score that
120
+ gates nothing, blocks nothing, and is consumed by one route and one command.
121
+
122
+ ## 4b. Judgment-point check (Judgment Within Floors standard)
123
+
124
+ None introduced. Two deterministic branches on a count.
125
+
126
+ ## 5. Interactions
127
+
128
+ - **`GET /intent/alignment`** — response gains `assessable`; `grade` may now be `'N/A'`. Additive plus
129
+ one widened union member.
130
+ - **`IntentDriftDetector.analyze()`** — untouched; drift scoring is a separate path.
131
+ - **`tests/unit/IntentDriftDetector.test.ts`** and **`tests/integration/drift-routes.test.ts`** — each
132
+ had a test asserting `grade === 'F'` on the empty case. **Both were encoding the defect**, not
133
+ merely stale. Updated with comments recording that, since a test that locks in a wrong answer is
134
+ the same instrument-honesty class as the defect itself.
135
+ - **`CapabilityIndex`** — unchanged; `intent` is already `INTERNAL_PREFIXES`.
136
+
137
+ ## 6. External surfaces
138
+
139
+ One API response shape change (additive field + widened union), one CLI rendering change. No config
140
+ key, no persisted state, no migration, no message to any user.
141
+
142
+ ## 6b. Operator-surface quality
143
+
144
+ The unassessable CLI output prints the reason (`No decisions logged — alignment cannot be assessed`)
145
+ in dim rather than a red grade, so it reads as an absence of data rather than an alarm. The component
146
+ breakdown is suppressed in that branch: four zeroed rows invite exactly the misreading being fixed.
147
+
148
+ ## 7. Multi-machine posture (Cross-Machine Coherence)
149
+
150
+ **Machine-local BY DESIGN.** The journal is a per-machine JSONL under `stateDir` and the score is
151
+ computed from it, so `assessable` answers "on this machine". An agent running on two machines has two
152
+ journals and two scores. That predates this change and is unaddressed here. <!-- tracked: CMT-1044 -->
153
+
154
+ ## 8. Rollback cost
155
+
156
+ Low. One union member, one boolean, one CLI branch, three test updates. No persisted state, no
157
+ migration; existing journal rows are read unchanged. Reverting restores the fabricated F.
158
+
159
+ ## Phase 5 — Second-pass review
160
+
161
+ Not a gate, sentinel, guard or watchdog; holds no block/allow authority; touches no session lifecycle
162
+ or trust level. The high-risk trigger list is not engaged. Author lenses, disclosed:
163
+
164
+ **Adversarial — "how would I make this useless?"** Three ways, all closed and asserted: report `'F'`
165
+ again (refusal 1), make `assessable` constant (refusal 2), or let the renderer ignore it entirely
166
+ (refusal 3 — the one that was genuinely open until I wrote the CLI test).
167
+
168
+ **"Would it have caught the incident?"** The incident here is my own: I read `topPrinciples: []` and
169
+ `score: 0 (F)` on a journal I already knew was empty, and had to reason my way to "that F is
170
+ meaningless" instead of being told. With this, the surface says it.
171
+
172
+ **"Symptom or cause?"** Cause, for the reporting defect — absence can no longer render as a verdict
173
+ because the type now has somewhere honest to put it. Symptom-level for the window mismatch, which is
174
+ named and left.
175
+
176
+ **Weakest point:** the CLI case is genuinely narrow (`--window > 30`), and I over-claimed its reach
177
+ twice before testing settled it. The API consumer is the one that matters by default. An artifact
178
+ claiming broad user impact here would be overstating it, so this one does not.
@@ -0,0 +1,149 @@
1
+ # Side-Effects Review — RecurrenceReader (Tier 2 core, read-only)
2
+
3
+ **Version / slug:** `recurrence-reader`
4
+ **Date:** `2026-07-27`
5
+ **Author:** `Echo (instar-dev agent)`
6
+ **Second-pass reviewer:** `see Phase 5`
7
+
8
+ ## Summary of the change
9
+
10
+ One pure module (`src/core/RecurrenceReader.ts`) that groups OPEN observations from the attention
11
+ queue, the evolution action queue and the sentinel log into recurrence clusters, plus a `coverage`
12
+ block naming every store it could not read.
13
+
14
+ Project `convergence-towards-coherence` Tier 2. The plan's diagnosis: instar notices constantly, in
15
+ three places, and nothing reads across them — so one problem is noticed dozens of times and closed
16
+ zero times (measured filing-to-completion ≈ 30:1).
17
+
18
+ **Measured on live data, 2026-07-27:**
19
+
20
+ ```
21
+ open observations across 3 stores : 2,068
22
+ distinct problems : 836
23
+ noticing ratio : 2.47
24
+ noticed repeatedly, NEVER tracked : 69 problems / 1,242 noticings
25
+ top clusters: 278x idle-timeout detection
26
+ 238x escalation-suppressed (telegramEscalation disabled)
27
+ 177x credential rebalancer ← 48% of the attention queue alone
28
+ ```
29
+
30
+ ## Refusal evidence (constraint 2)
31
+
32
+ The whole design risk is that a synthesiser becomes a MORE expensive version of the defect it
33
+ detects: reading 2 of 3 stores and reporting "nothing recurring" with the authority of having looked.
34
+
35
+ ```
36
+ REFUSAL — action store made unreadable, on REAL data
37
+ coverage : partial
38
+ could NOT read: [{"store":"actions","reason":"ENOENT: evolution store unreadable"}]
39
+ clusters : 59 ← still reports what it DID see
40
+ verdict : ABSENT — refuses to say no-recurrence
41
+
42
+ THE DISTINCTION THAT MATTERS
43
+ genuinely nothing there → "no-recurrence"
44
+ could not look → undefined (field absent, not hedged)
45
+ ```
46
+
47
+ Unit suite: **11 passed (11)**; `tsc --noEmit` exit 0.
48
+
49
+ ## Decision-point inventory
50
+
51
+ | point | classification | note |
52
+ |---|---|---|
53
+ | recurrence key (digits→N, hex→H) | `invariant` | Deterministic string normalization. No model. |
54
+ | cluster grouping | `invariant` | Map by key. |
55
+ | `verdict` emitted only on complete coverage | `invariant` | The load-bearing rule. |
56
+ | `significantClusters` minCount / untrackedOnly | `invariant` | Caller-supplied thresholds, defaulted, not inferred. |
57
+
58
+ No judgment points, no LLM, nothing gated. The module holds **no authority whatsoever** — it returns
59
+ a report.
60
+
61
+ ## 1. Over-block
62
+
63
+ Nothing is blocked; the module is read-only and returns data. The realistic over-*grouping* risk is
64
+ the blunt key: two genuinely different problems whose titles differ only by digits would merge. That
65
+ is a deliberate trade, stated in the source — surfacing shape is the goal, and an over-eager grouping
66
+ a human instantly recognises as one problem beats a precise grouping that preserves the illusion of
67
+ 371 separate things. `exemplar` and `sources` are carried on every cluster so a reader can spot a bad
68
+ merge immediately.
69
+
70
+ ## 2. Under-block
71
+
72
+ **Title-only keying.** Two reports of the same underlying problem with genuinely different wording
73
+ will not merge. Accepted: the alternative is semantic matching, which means an LLM, which means a
74
+ judgment point in something that currently has none.
75
+
76
+ **`open` is caller-supplied.** The module trusts the caller's open/closed determination per store.
77
+ That is the correct seam — each store knows its own status vocabulary — but it means a caller that
78
+ mis-maps status inflates or deflates the counts. The live harness maps `status === 'OPEN'`,
79
+ `pending|in_progress`, and treats sentinel events as open.
80
+
81
+ **No route yet, no action yet.** This increment is the reader only. Driving action is Tier 2 item 4,
82
+ deliberately separate because it carries authority this does not.
83
+
84
+ ## 3. Level-of-abstraction fit
85
+
86
+ A pure function over supplied observations, with I/O left entirely to the caller. That is
87
+ deliberate: it means the module **cannot** silently swallow a failed store read — the caller must
88
+ hand it a `coverage` block, so an unreadable store is structurally impossible to omit. Putting the
89
+ reads inside would have made "forgot to report the failure" a one-line mistake.
90
+
91
+ ## 4. Signal vs authority compliance
92
+
93
+ Textbook signal-producer. It returns a report and holds zero blocking, gating or notifying authority.
94
+ `docs/signal-vs-authority.md` satisfied — and this is the exact seam the operator flagged: synthesis
95
+ must drive action through EXISTING gated paths, never become a new notification channel. Keeping the
96
+ reader authority-free is what makes that possible later.
97
+
98
+ ## 5. Interactions
99
+
100
+ - **Attention queue / evolution actions / sentinel log** — read-only consumers, no writes, no schema
101
+ change. Nothing else observes this module yet.
102
+ - **Nothing shadows or is shadowed.** New module, no existing caller.
103
+
104
+ ## 6. External surfaces
105
+
106
+ **None in this increment.** No route, no config, no persisted state, no user-visible behaviour. A
107
+ route is the obvious next step and is deliberately not here.
108
+
109
+ ## 6b. Operator-surface quality
110
+
111
+ `coverage.unreadable[].reason` carries the actual failure text so a caller can say *why* it could not
112
+ look, not merely that it could not. `noticingRatio` is `null` — never `0` — when there is no
113
+ denominator, so a client that ignores the contract gets an obviously-missing value rather than a
114
+ plausible wrong one.
115
+
116
+ ## 7. Multi-machine posture
117
+
118
+ **Posture: `machine-local`.** `machine-local-justification: physical-credential-locality` — the three
119
+ stores are per-machine records of what THAT machine noticed, and observation titles routinely carry
120
+ machine ids, topic ids and account emails. Replicating them to synthesise centrally would multiply
121
+ at-rest exposure of that context across every machine. The correct cross-machine read is the existing
122
+ pool-scope fan-out (`?scope=pool`), which serves each machine's own data from that machine — a
123
+ follow-up for whoever adds the route, noted rather than assumed.
124
+
125
+ ## 8. Rollback cost
126
+
127
+ **Zero.** One new module and one new test file, with no callers. Deleting them removes the feature
128
+ entirely; nothing else changes. No persisted state, no migration, no config.
129
+
130
+ ## Phase 5 — Second-pass review
131
+
132
+ Not a gate, sentinel, guard or watchdog; no block/allow authority; no session lifecycle or trust
133
+ surface; no LLM. High-risk trigger list not engaged. Author lenses:
134
+
135
+ **Adversarial — "how would I make this useless?"** By letting it report a clean verdict over a
136
+ partial read. That is the one thing it structurally cannot do, asserted from both directions
137
+ (complete-and-empty → `no-recurrence`; incomplete → field absent) and demonstrated on real data.
138
+
139
+ **"Would it have caught the incident?"** The incident is the project's premise, and yes — 2,068
140
+ noticings collapsing to 836 problems with 69 untracked recurrers is precisely the shape nobody could
141
+ see. It found it on first run.
142
+
143
+ **"Symptom or cause?"** Cause, for the invisibility. NOT for the recurrence itself: this makes the
144
+ 69 untracked recurrers visible, it does not close them. Closing them is item 4, and claiming
145
+ otherwise would be the filing-as-progress failure the project exists to remove.
146
+
147
+ **Weakest point:** the blunt recurrence key. It will occasionally merge two things a human would
148
+ separate. Mitigated by carrying `exemplar` + `sources`, and preferable to under-grouping — but it is
149
+ the assumption most likely to need revisiting once a human reads a real report.