instar 1.3.990 → 1.3.992
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/commands/intent.d.ts.map +1 -1
- package/dist/commands/intent.js +11 -0
- package/dist/commands/intent.js.map +1 -1
- package/dist/core/IntentDriftDetector.d.ts +18 -2
- package/dist/core/IntentDriftDetector.d.ts.map +1 -1
- package/dist/core/IntentDriftDetector.js +4 -1
- package/dist/core/IntentDriftDetector.js.map +1 -1
- package/dist/core/PostUpdateMigrator.d.ts.map +1 -1
- package/dist/core/PostUpdateMigrator.js +6 -0
- package/dist/core/PostUpdateMigrator.js.map +1 -1
- package/dist/core/RecurrenceReader.d.ts +134 -0
- package/dist/core/RecurrenceReader.d.ts.map +1 -0
- package/dist/core/RecurrenceReader.js +111 -0
- package/dist/core/RecurrenceReader.js.map +1 -0
- package/dist/scaffold/templates.d.ts.map +1 -1
- package/dist/scaffold/templates.js +2 -0
- package/dist/scaffold/templates.js.map +1 -1
- package/package.json +1 -1
- package/src/data/builtin-manifest.json +19 -19
- package/src/scaffold/templates.ts +2 -0
- package/upgrades/1.3.991.md +57 -0
- package/upgrades/1.3.992.md +55 -0
- package/upgrades/side-effects/alignment-score-not-assessed.md +178 -0
- package/upgrades/side-effects/recurrence-reader.md +149 -0
|
@@ -0,0 +1,178 @@
|
|
|
1
|
+
# Side-Effects Review — an alignment score that can say "not assessed"
|
|
2
|
+
|
|
3
|
+
**Version / slug:** `alignment-score-not-assessed`
|
|
4
|
+
**Date:** `2026-07-26`
|
|
5
|
+
**Author:** `Echo (instar-dev agent)`
|
|
6
|
+
**Second-pass reviewer:** `see Phase 5`
|
|
7
|
+
|
|
8
|
+
## Summary of the change
|
|
9
|
+
|
|
10
|
+
`IntentDriftDetector.alignmentScore()` returned `score: 0, grade: 'F'` when the analysis window
|
|
11
|
+
contained no decisions. The `summary` field was honest ("No decisions logged — alignment cannot be
|
|
12
|
+
assessed") but no consumer read it; `score` and `grade` are what get rendered and compared. So
|
|
13
|
+
"nothing to assess" and "assessed, catastrophically bad" were identical on every field in use — on
|
|
14
|
+
the instrument whose purpose is honest alignment measurement.
|
|
15
|
+
|
|
16
|
+
Root cause is vocabulary, the same shape as the channel registry one increment earlier: the grade
|
|
17
|
+
union was `'A'|'B'|'C'|'D'|'F'` with no member meaning *no verdict*, so absence had to borrow the
|
|
18
|
+
worst real grade.
|
|
19
|
+
|
|
20
|
+
Adds `'N/A'` to the grade union, adds `assessable: boolean`, and makes `instar intent drift` print
|
|
21
|
+
the reason instead of a fabricated grade.
|
|
22
|
+
|
|
23
|
+
## Refusal evidence (constraint 2)
|
|
24
|
+
|
|
25
|
+
```
|
|
26
|
+
REFUSAL 1 — restore the fabricated 'F' on the unassessable case
|
|
27
|
+
× handles empty journal — not assessable, grade N/A → expected 'F' to be 'N/A'
|
|
28
|
+
× GET /intent/alignment reports NOT ASSESSED → expected 'F' to be 'N/A'
|
|
29
|
+
× an empty journal grades 'N/A', never 'F' → expected 'F' to be 'N/A'
|
|
30
|
+
Tests 3 failed | 25 passed (28)
|
|
31
|
+
|
|
32
|
+
REFUSAL 2 — always report assessable:true
|
|
33
|
+
× an empty journal is flagged unassessable → expected true to be false
|
|
34
|
+
× a real assessment is DISTINGUISHABLE from an empty one → expected true not to be true
|
|
35
|
+
× assessable tracks sampleSize exactly → expected true to be false
|
|
36
|
+
(+2) Tests 5 failed
|
|
37
|
+
|
|
38
|
+
REFUSAL 3 — disable the CLI's honest branch (`if (false && !alignment.assessable)`)
|
|
39
|
+
BEFORE the CLI test existed: Tests 28 passed (28) <-- the blindness, again
|
|
40
|
+
AFTER: × a STALE journal prints "not assessed", never a red F
|
|
41
|
+
Tests 1 failed | 4 passed (5)
|
|
42
|
+
```
|
|
43
|
+
|
|
44
|
+
Restored: **42 passed** across the five affected files, `tsc --noEmit` exit 0.
|
|
45
|
+
|
|
46
|
+
**REFUSAL 3 is the finding, and it is the THIRD occurrence of this class tonight** (#1658 route
|
|
47
|
+
registry, #1659 route validator, now a CLI renderer). Each time the logic was thoroughly guarded and
|
|
48
|
+
the wiring to the surface a human or API client actually reads was not. This one is the sharpest:
|
|
49
|
+
the module returning `'N/A'` is worth precisely nothing if the renderer ignores it, and 28 green
|
|
50
|
+
tests said everything was fine while the renderer was disabled.
|
|
51
|
+
|
|
52
|
+
## Two of my own claims were falsified during this work
|
|
53
|
+
|
|
54
|
+
Recorded because the corrections are the useful part, and because an artifact that hides them is the
|
|
55
|
+
failure mode this tier exists to remove.
|
|
56
|
+
|
|
57
|
+
1. **"`instar intent drift` has been showing a red F for the journal's whole life."** FALSE. The
|
|
58
|
+
command returns early with a genuinely helpful message when the window holds no decisions; it
|
|
59
|
+
never reaches the scoring block. The empty case was already handled honestly there.
|
|
60
|
+
2. **"Then it is reachable whenever the journal is merely stale."** ALSO FALSE. The early return
|
|
61
|
+
checks `windowDays` (default 14) and `alignmentScore()` is fixed at 30 — and 14 ⊂ 30, so anything
|
|
62
|
+
clearing the early return is inside the alignment window by construction.
|
|
63
|
+
|
|
64
|
+
**The actual reachable CLI case is narrow:** the operator must widen the window past 30
|
|
65
|
+
(`--window 60`), so a 40-day-old decision clears the early return and falls outside the fixed 30-day
|
|
66
|
+
alignment window. That is what the regression test constructs. I asserted twice before checking; the
|
|
67
|
+
test is what settled it.
|
|
68
|
+
|
|
69
|
+
## Decision-point inventory
|
|
70
|
+
|
|
71
|
+
| point | classification | note |
|
|
72
|
+
|---|---|---|
|
|
73
|
+
| `sampleSize === 0` → `grade: 'N/A'`, `assessable: false` | `invariant` | Deterministic count check. No model. |
|
|
74
|
+
| `assessable` mirrors `sampleSize > 0` | `invariant` | Asserted by test; the two can never disagree. |
|
|
75
|
+
| CLI branches on `assessable` | `invariant` | Renders `summary` instead of a grade. |
|
|
76
|
+
|
|
77
|
+
No judgment points, no LLM, nothing gated or blocked.
|
|
78
|
+
|
|
79
|
+
## 1. Over-block
|
|
80
|
+
|
|
81
|
+
Nothing is blocked — this is a read surface. The available harm is **misinforming a reader**, and
|
|
82
|
+
this change strictly reduces it in the direction that mattered (absence no longer reads as failure).
|
|
83
|
+
|
|
84
|
+
The mirror over-block is real and guarded: a genuinely-assessed period must never report `'N/A'` or
|
|
85
|
+
`assessable: false`, or a real alignment problem would be hidden as "no data" — strictly worse than
|
|
86
|
+
the original bug. Asserted by two tests (`a genuinely assessed period still reports a real letter
|
|
87
|
+
grade`, and the CLI's `a populated journal still prints a real graded score`).
|
|
88
|
+
|
|
89
|
+
**Caller sweep, run BEFORE writing this section** (the correction from #1659, where I wrote a
|
|
90
|
+
confident risk claim from the wrong measurement and CI falsified it): `alignmentScore()` has exactly
|
|
91
|
+
two production callers — `routes.ts:24305` (passes through verbatim) and `commands/intent.ts:431`
|
|
92
|
+
(now branches). Every `.grade` hit elsewhere in `src/` belongs to `DecisionQualityRecorder`'s
|
|
93
|
+
unrelated `right|wrong|unknown` grade, checked rather than assumed. Two tests asserted the old shape;
|
|
94
|
+
both updated.
|
|
95
|
+
|
|
96
|
+
## 2. Under-block
|
|
97
|
+
|
|
98
|
+
**The mismatched windows are NOT fixed.** The early return uses `windowDays` while `alignmentScore()`
|
|
99
|
+
is hardcoded to 30. That divergence is what makes the CLI case reachable at all, and reconciling them
|
|
100
|
+
changes what the command reports for every user. Deliberately not folded in. <!-- tracked: CMT-1044 -->
|
|
101
|
+
|
|
102
|
+
**`score: 0` is retained on the unassessable case.** Changing it to `null` would be a breaking type
|
|
103
|
+
change for a field two consumers read; `assessable` is the additive signal instead. A consumer that
|
|
104
|
+
reads `.score` and ignores both `assessable` and `sampleSize` still sees a 0 — it can no longer see
|
|
105
|
+
an F, which is the part that read as a verdict.
|
|
106
|
+
|
|
107
|
+
**It does not improve alignment,** and it does not judge whether a cited principle was genuinely
|
|
108
|
+
consulted. Same honest limit as the increment before it.
|
|
109
|
+
|
|
110
|
+
## 3. Level-of-abstraction fit
|
|
111
|
+
|
|
112
|
+
The honest state lives in the returned value, not in the renderer, so every consumer inherits it —
|
|
113
|
+
the route needed no change at all. The renderer's job is narrowed to *presenting* a state it no
|
|
114
|
+
longer has to infer. Had I fixed only the CLI, the API consumer (the one that is actually reachable
|
|
115
|
+
by default) would still have been lied to.
|
|
116
|
+
|
|
117
|
+
## 4. Signal vs authority compliance
|
|
118
|
+
|
|
119
|
+
Pure signal. `docs/signal-vs-authority.md` is satisfied trivially: it produces a read-only score that
|
|
120
|
+
gates nothing, blocks nothing, and is consumed by one route and one command.
|
|
121
|
+
|
|
122
|
+
## 4b. Judgment-point check (Judgment Within Floors standard)
|
|
123
|
+
|
|
124
|
+
None introduced. Two deterministic branches on a count.
|
|
125
|
+
|
|
126
|
+
## 5. Interactions
|
|
127
|
+
|
|
128
|
+
- **`GET /intent/alignment`** — response gains `assessable`; `grade` may now be `'N/A'`. Additive plus
|
|
129
|
+
one widened union member.
|
|
130
|
+
- **`IntentDriftDetector.analyze()`** — untouched; drift scoring is a separate path.
|
|
131
|
+
- **`tests/unit/IntentDriftDetector.test.ts`** and **`tests/integration/drift-routes.test.ts`** — each
|
|
132
|
+
had a test asserting `grade === 'F'` on the empty case. **Both were encoding the defect**, not
|
|
133
|
+
merely stale. Updated with comments recording that, since a test that locks in a wrong answer is
|
|
134
|
+
the same instrument-honesty class as the defect itself.
|
|
135
|
+
- **`CapabilityIndex`** — unchanged; `intent` is already `INTERNAL_PREFIXES`.
|
|
136
|
+
|
|
137
|
+
## 6. External surfaces
|
|
138
|
+
|
|
139
|
+
One API response shape change (additive field + widened union), one CLI rendering change. No config
|
|
140
|
+
key, no persisted state, no migration, no message to any user.
|
|
141
|
+
|
|
142
|
+
## 6b. Operator-surface quality
|
|
143
|
+
|
|
144
|
+
The unassessable CLI output prints the reason (`No decisions logged — alignment cannot be assessed`)
|
|
145
|
+
in dim rather than a red grade, so it reads as an absence of data rather than an alarm. The component
|
|
146
|
+
breakdown is suppressed in that branch: four zeroed rows invite exactly the misreading being fixed.
|
|
147
|
+
|
|
148
|
+
## 7. Multi-machine posture (Cross-Machine Coherence)
|
|
149
|
+
|
|
150
|
+
**Machine-local BY DESIGN.** The journal is a per-machine JSONL under `stateDir` and the score is
|
|
151
|
+
computed from it, so `assessable` answers "on this machine". An agent running on two machines has two
|
|
152
|
+
journals and two scores. That predates this change and is unaddressed here. <!-- tracked: CMT-1044 -->
|
|
153
|
+
|
|
154
|
+
## 8. Rollback cost
|
|
155
|
+
|
|
156
|
+
Low. One union member, one boolean, one CLI branch, three test updates. No persisted state, no
|
|
157
|
+
migration; existing journal rows are read unchanged. Reverting restores the fabricated F.
|
|
158
|
+
|
|
159
|
+
## Phase 5 — Second-pass review
|
|
160
|
+
|
|
161
|
+
Not a gate, sentinel, guard or watchdog; holds no block/allow authority; touches no session lifecycle
|
|
162
|
+
or trust level. The high-risk trigger list is not engaged. Author lenses, disclosed:
|
|
163
|
+
|
|
164
|
+
**Adversarial — "how would I make this useless?"** Three ways, all closed and asserted: report `'F'`
|
|
165
|
+
again (refusal 1), make `assessable` constant (refusal 2), or let the renderer ignore it entirely
|
|
166
|
+
(refusal 3 — the one that was genuinely open until I wrote the CLI test).
|
|
167
|
+
|
|
168
|
+
**"Would it have caught the incident?"** The incident here is my own: I read `topPrinciples: []` and
|
|
169
|
+
`score: 0 (F)` on a journal I already knew was empty, and had to reason my way to "that F is
|
|
170
|
+
meaningless" instead of being told. With this, the surface says it.
|
|
171
|
+
|
|
172
|
+
**"Symptom or cause?"** Cause, for the reporting defect — absence can no longer render as a verdict
|
|
173
|
+
because the type now has somewhere honest to put it. Symptom-level for the window mismatch, which is
|
|
174
|
+
named and left.
|
|
175
|
+
|
|
176
|
+
**Weakest point:** the CLI case is genuinely narrow (`--window > 30`), and I over-claimed its reach
|
|
177
|
+
twice before testing settled it. The API consumer is the one that matters by default. An artifact
|
|
178
|
+
claiming broad user impact here would be overstating it, so this one does not.
|
|
@@ -0,0 +1,149 @@
|
|
|
1
|
+
# Side-Effects Review — RecurrenceReader (Tier 2 core, read-only)
|
|
2
|
+
|
|
3
|
+
**Version / slug:** `recurrence-reader`
|
|
4
|
+
**Date:** `2026-07-27`
|
|
5
|
+
**Author:** `Echo (instar-dev agent)`
|
|
6
|
+
**Second-pass reviewer:** `see Phase 5`
|
|
7
|
+
|
|
8
|
+
## Summary of the change
|
|
9
|
+
|
|
10
|
+
One pure module (`src/core/RecurrenceReader.ts`) that groups OPEN observations from the attention
|
|
11
|
+
queue, the evolution action queue and the sentinel log into recurrence clusters, plus a `coverage`
|
|
12
|
+
block naming every store it could not read.
|
|
13
|
+
|
|
14
|
+
Project `convergence-towards-coherence` Tier 2. The plan's diagnosis: instar notices constantly, in
|
|
15
|
+
three places, and nothing reads across them — so one problem is noticed dozens of times and closed
|
|
16
|
+
zero times (measured filing-to-completion ≈ 30:1).
|
|
17
|
+
|
|
18
|
+
**Measured on live data, 2026-07-27:**
|
|
19
|
+
|
|
20
|
+
```
|
|
21
|
+
open observations across 3 stores : 2,068
|
|
22
|
+
distinct problems : 836
|
|
23
|
+
noticing ratio : 2.47
|
|
24
|
+
noticed repeatedly, NEVER tracked : 69 problems / 1,242 noticings
|
|
25
|
+
top clusters: 278x idle-timeout detection
|
|
26
|
+
238x escalation-suppressed (telegramEscalation disabled)
|
|
27
|
+
177x credential rebalancer ← 48% of the attention queue alone
|
|
28
|
+
```
|
|
29
|
+
|
|
30
|
+
## Refusal evidence (constraint 2)
|
|
31
|
+
|
|
32
|
+
The whole design risk is that a synthesiser becomes a MORE expensive version of the defect it
|
|
33
|
+
detects: reading 2 of 3 stores and reporting "nothing recurring" with the authority of having looked.
|
|
34
|
+
|
|
35
|
+
```
|
|
36
|
+
REFUSAL — action store made unreadable, on REAL data
|
|
37
|
+
coverage : partial
|
|
38
|
+
could NOT read: [{"store":"actions","reason":"ENOENT: evolution store unreadable"}]
|
|
39
|
+
clusters : 59 ← still reports what it DID see
|
|
40
|
+
verdict : ABSENT — refuses to say no-recurrence
|
|
41
|
+
|
|
42
|
+
THE DISTINCTION THAT MATTERS
|
|
43
|
+
genuinely nothing there → "no-recurrence"
|
|
44
|
+
could not look → undefined (field absent, not hedged)
|
|
45
|
+
```
|
|
46
|
+
|
|
47
|
+
Unit suite: **11 passed (11)**; `tsc --noEmit` exit 0.
|
|
48
|
+
|
|
49
|
+
## Decision-point inventory
|
|
50
|
+
|
|
51
|
+
| point | classification | note |
|
|
52
|
+
|---|---|---|
|
|
53
|
+
| recurrence key (digits→N, hex→H) | `invariant` | Deterministic string normalization. No model. |
|
|
54
|
+
| cluster grouping | `invariant` | Map by key. |
|
|
55
|
+
| `verdict` emitted only on complete coverage | `invariant` | The load-bearing rule. |
|
|
56
|
+
| `significantClusters` minCount / untrackedOnly | `invariant` | Caller-supplied thresholds, defaulted, not inferred. |
|
|
57
|
+
|
|
58
|
+
No judgment points, no LLM, nothing gated. The module holds **no authority whatsoever** — it returns
|
|
59
|
+
a report.
|
|
60
|
+
|
|
61
|
+
## 1. Over-block
|
|
62
|
+
|
|
63
|
+
Nothing is blocked; the module is read-only and returns data. The realistic over-*grouping* risk is
|
|
64
|
+
the blunt key: two genuinely different problems whose titles differ only by digits would merge. That
|
|
65
|
+
is a deliberate trade, stated in the source — surfacing shape is the goal, and an over-eager grouping
|
|
66
|
+
a human instantly recognises as one problem beats a precise grouping that preserves the illusion of
|
|
67
|
+
371 separate things. `exemplar` and `sources` are carried on every cluster so a reader can spot a bad
|
|
68
|
+
merge immediately.
|
|
69
|
+
|
|
70
|
+
## 2. Under-block
|
|
71
|
+
|
|
72
|
+
**Title-only keying.** Two reports of the same underlying problem with genuinely different wording
|
|
73
|
+
will not merge. Accepted: the alternative is semantic matching, which means an LLM, which means a
|
|
74
|
+
judgment point in something that currently has none.
|
|
75
|
+
|
|
76
|
+
**`open` is caller-supplied.** The module trusts the caller's open/closed determination per store.
|
|
77
|
+
That is the correct seam — each store knows its own status vocabulary — but it means a caller that
|
|
78
|
+
mis-maps status inflates or deflates the counts. The live harness maps `status === 'OPEN'`,
|
|
79
|
+
`pending|in_progress`, and treats sentinel events as open.
|
|
80
|
+
|
|
81
|
+
**No route yet, no action yet.** This increment is the reader only. Driving action is Tier 2 item 4,
|
|
82
|
+
deliberately separate because it carries authority this does not.
|
|
83
|
+
|
|
84
|
+
## 3. Level-of-abstraction fit
|
|
85
|
+
|
|
86
|
+
A pure function over supplied observations, with I/O left entirely to the caller. That is
|
|
87
|
+
deliberate: it means the module **cannot** silently swallow a failed store read — the caller must
|
|
88
|
+
hand it a `coverage` block, so an unreadable store is structurally impossible to omit. Putting the
|
|
89
|
+
reads inside would have made "forgot to report the failure" a one-line mistake.
|
|
90
|
+
|
|
91
|
+
## 4. Signal vs authority compliance
|
|
92
|
+
|
|
93
|
+
Textbook signal-producer. It returns a report and holds zero blocking, gating or notifying authority.
|
|
94
|
+
`docs/signal-vs-authority.md` satisfied — and this is the exact seam the operator flagged: synthesis
|
|
95
|
+
must drive action through EXISTING gated paths, never become a new notification channel. Keeping the
|
|
96
|
+
reader authority-free is what makes that possible later.
|
|
97
|
+
|
|
98
|
+
## 5. Interactions
|
|
99
|
+
|
|
100
|
+
- **Attention queue / evolution actions / sentinel log** — read-only consumers, no writes, no schema
|
|
101
|
+
change. Nothing else observes this module yet.
|
|
102
|
+
- **Nothing shadows or is shadowed.** New module, no existing caller.
|
|
103
|
+
|
|
104
|
+
## 6. External surfaces
|
|
105
|
+
|
|
106
|
+
**None in this increment.** No route, no config, no persisted state, no user-visible behaviour. A
|
|
107
|
+
route is the obvious next step and is deliberately not here.
|
|
108
|
+
|
|
109
|
+
## 6b. Operator-surface quality
|
|
110
|
+
|
|
111
|
+
`coverage.unreadable[].reason` carries the actual failure text so a caller can say *why* it could not
|
|
112
|
+
look, not merely that it could not. `noticingRatio` is `null` — never `0` — when there is no
|
|
113
|
+
denominator, so a client that ignores the contract gets an obviously-missing value rather than a
|
|
114
|
+
plausible wrong one.
|
|
115
|
+
|
|
116
|
+
## 7. Multi-machine posture
|
|
117
|
+
|
|
118
|
+
**Posture: `machine-local`.** `machine-local-justification: physical-credential-locality` — the three
|
|
119
|
+
stores are per-machine records of what THAT machine noticed, and observation titles routinely carry
|
|
120
|
+
machine ids, topic ids and account emails. Replicating them to synthesise centrally would multiply
|
|
121
|
+
at-rest exposure of that context across every machine. The correct cross-machine read is the existing
|
|
122
|
+
pool-scope fan-out (`?scope=pool`), which serves each machine's own data from that machine — a
|
|
123
|
+
follow-up for whoever adds the route, noted rather than assumed.
|
|
124
|
+
|
|
125
|
+
## 8. Rollback cost
|
|
126
|
+
|
|
127
|
+
**Zero.** One new module and one new test file, with no callers. Deleting them removes the feature
|
|
128
|
+
entirely; nothing else changes. No persisted state, no migration, no config.
|
|
129
|
+
|
|
130
|
+
## Phase 5 — Second-pass review
|
|
131
|
+
|
|
132
|
+
Not a gate, sentinel, guard or watchdog; no block/allow authority; no session lifecycle or trust
|
|
133
|
+
surface; no LLM. High-risk trigger list not engaged. Author lenses:
|
|
134
|
+
|
|
135
|
+
**Adversarial — "how would I make this useless?"** By letting it report a clean verdict over a
|
|
136
|
+
partial read. That is the one thing it structurally cannot do, asserted from both directions
|
|
137
|
+
(complete-and-empty → `no-recurrence`; incomplete → field absent) and demonstrated on real data.
|
|
138
|
+
|
|
139
|
+
**"Would it have caught the incident?"** The incident is the project's premise, and yes — 2,068
|
|
140
|
+
noticings collapsing to 836 problems with 69 untracked recurrers is precisely the shape nobody could
|
|
141
|
+
see. It found it on first run.
|
|
142
|
+
|
|
143
|
+
**"Symptom or cause?"** Cause, for the invisibility. NOT for the recurrence itself: this makes the
|
|
144
|
+
69 untracked recurrers visible, it does not close them. Closing them is item 4, and claiming
|
|
145
|
+
otherwise would be the filing-as-progress failure the project exists to remove.
|
|
146
|
+
|
|
147
|
+
**Weakest point:** the blunt recurrence key. It will occasionally merge two things a human would
|
|
148
|
+
separate. Mitigated by carrying `exemplar` + `sources`, and preferable to under-grouping — but it is
|
|
149
|
+
the assumption most likely to need revisiting once a human reads a real report.
|