instar 1.3.989 → 1.3.991

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,57 @@
1
+ # Upgrade Guide — vNEXT
2
+
3
+ <!-- assembled-by: assemble-next-md -->
4
+ <!-- bump: patch -->
5
+
6
+ ## What Changed
7
+
8
+ `IntentDriftDetector.alignmentScore()` reported `score: 0, grade: 'F'` when its analysis window held
9
+ no decisions. The `summary` field said "alignment cannot be assessed", but `score` and `grade` are
10
+ what consumers read — so "no data" and "assessed, catastrophically bad" were identical on every field
11
+ in use.
12
+
13
+ The grade union is now `'A'|'B'|'C'|'D'|'F'|'N/A'`, and the result carries `assessable: boolean`.
14
+ `GET /intent/alignment` returns both. `instar intent drift` prints the reason instead of a fabricated
15
+ grade when nothing was assessed.
16
+
17
+ Root cause was vocabulary: with no grade meaning *no verdict*, absence had to borrow the worst real
18
+ grade — the same shape as a status set with no way to say "undetermined".
19
+
20
+ ## What to Tell Your User
21
+
22
+ If you look at the alignment score, a flat zero with an F used to mean one of two very different
23
+ things: that your decisions genuinely scored badly, or that there was nothing to score. You could not
24
+ tell which from the number.
25
+
26
+ Now an unassessed period says so plainly, and the score carries a flag saying whether anything was
27
+ actually measured.
28
+
29
+ The honest scope: this changes what you are told, not how aligned anything is. It also does not
30
+ change the default command view, where the unassessable case was already handled well — that path
31
+ stops early with a helpful message and never reached the bad grade. The fabricated F is visible to
32
+ anything reading the score programmatically, and on the command line only if you widen the window
33
+ past thirty days by hand.
34
+
35
+ ## Summary of New Capabilities
36
+
37
+ No new endpoint or command. An existing score gained an honest "not assessed" state and a flag
38
+ distinguishing a placeholder from a measurement.
39
+
40
+ ## Evidence
41
+
42
+ - Restoring the fabricated `'F'`: 3 failed | 25 passed. Making `assessable` constant: 5 failed.
43
+ - Disabling the renderer's honest branch left all 28 module and route tests green — the wiring was
44
+ unguarded until a test was added that drives the real command and reads its output. With it, that
45
+ same change fails immediately.
46
+ - Restored: 42 passed across the five affected files; `tsc --noEmit` exit 0.
47
+ - Two existing tests asserted `grade === 'F'` on the empty case — they encoded the defect rather than
48
+ merely lagging it. Both updated, with the reason recorded in place.
49
+
50
+ ## Known limits
51
+
52
+ The early-return window and the scoring window differ (the command's `--window`, default 14, versus a
53
+ fixed 30), which is what makes the command-line case reachable at all; reconciling them changes what
54
+ every user sees and is deliberately not bundled. `score` stays `0` on the unassessable case rather
55
+ than becoming `null`, to avoid a breaking type change — `assessable` is the additive signal. Nothing
56
+ here improves alignment or judges whether a cited principle was genuinely consulted.
57
+ <!-- tracked: CMT-1044 -->
@@ -0,0 +1,178 @@
1
+ # Side-Effects Review — an alignment score that can say "not assessed"
2
+
3
+ **Version / slug:** `alignment-score-not-assessed`
4
+ **Date:** `2026-07-26`
5
+ **Author:** `Echo (instar-dev agent)`
6
+ **Second-pass reviewer:** `see Phase 5`
7
+
8
+ ## Summary of the change
9
+
10
+ `IntentDriftDetector.alignmentScore()` returned `score: 0, grade: 'F'` when the analysis window
11
+ contained no decisions. The `summary` field was honest ("No decisions logged — alignment cannot be
12
+ assessed") but no consumer read it; `score` and `grade` are what get rendered and compared. So
13
+ "nothing to assess" and "assessed, catastrophically bad" were identical on every field in use — on
14
+ the instrument whose purpose is honest alignment measurement.
15
+
16
+ Root cause is vocabulary, the same shape as the channel registry one increment earlier: the grade
17
+ union was `'A'|'B'|'C'|'D'|'F'` with no member meaning *no verdict*, so absence had to borrow the
18
+ worst real grade.
19
+
20
+ Adds `'N/A'` to the grade union, adds `assessable: boolean`, and makes `instar intent drift` print
21
+ the reason instead of a fabricated grade.
22
+
23
+ ## Refusal evidence (constraint 2)
24
+
25
+ ```
26
+ REFUSAL 1 — restore the fabricated 'F' on the unassessable case
27
+ × handles empty journal — not assessable, grade N/A → expected 'F' to be 'N/A'
28
+ × GET /intent/alignment reports NOT ASSESSED → expected 'F' to be 'N/A'
29
+ × an empty journal grades 'N/A', never 'F' → expected 'F' to be 'N/A'
30
+ Tests 3 failed | 25 passed (28)
31
+
32
+ REFUSAL 2 — always report assessable:true
33
+ × an empty journal is flagged unassessable → expected true to be false
34
+ × a real assessment is DISTINGUISHABLE from an empty one → expected true not to be true
35
+ × assessable tracks sampleSize exactly → expected true to be false
36
+ (+2) Tests 5 failed
37
+
38
+ REFUSAL 3 — disable the CLI's honest branch (`if (false && !alignment.assessable)`)
39
+ BEFORE the CLI test existed: Tests 28 passed (28) <-- the blindness, again
40
+ AFTER: × a STALE journal prints "not assessed", never a red F
41
+ Tests 1 failed | 4 passed (5)
42
+ ```
43
+
44
+ Restored: **42 passed** across the five affected files, `tsc --noEmit` exit 0.
45
+
46
+ **REFUSAL 3 is the finding, and it is the THIRD occurrence of this class tonight** (#1658 route
47
+ registry, #1659 route validator, now a CLI renderer). Each time the logic was thoroughly guarded and
48
+ the wiring to the surface a human or API client actually reads was not. This one is the sharpest:
49
+ the module returning `'N/A'` is worth precisely nothing if the renderer ignores it, and 28 green
50
+ tests said everything was fine while the renderer was disabled.
51
+
52
+ ## Two of my own claims were falsified during this work
53
+
54
+ Recorded because the corrections are the useful part, and because an artifact that hides them is the
55
+ failure mode this tier exists to remove.
56
+
57
+ 1. **"`instar intent drift` has been showing a red F for the journal's whole life."** FALSE. The
58
+ command returns early with a genuinely helpful message when the window holds no decisions; it
59
+ never reaches the scoring block. The empty case was already handled honestly there.
60
+ 2. **"Then it is reachable whenever the journal is merely stale."** ALSO FALSE. The early return
61
+ checks `windowDays` (default 14) and `alignmentScore()` is fixed at 30 — and 14 ⊂ 30, so anything
62
+ clearing the early return is inside the alignment window by construction.
63
+
64
+ **The actual reachable CLI case is narrow:** the operator must widen the window past 30
65
+ (`--window 60`), so a 40-day-old decision clears the early return and falls outside the fixed 30-day
66
+ alignment window. That is what the regression test constructs. I asserted twice before checking; the
67
+ test is what settled it.
68
+
69
+ ## Decision-point inventory
70
+
71
+ | point | classification | note |
72
+ |---|---|---|
73
+ | `sampleSize === 0` → `grade: 'N/A'`, `assessable: false` | `invariant` | Deterministic count check. No model. |
74
+ | `assessable` mirrors `sampleSize > 0` | `invariant` | Asserted by test; the two can never disagree. |
75
+ | CLI branches on `assessable` | `invariant` | Renders `summary` instead of a grade. |
76
+
77
+ No judgment points, no LLM, nothing gated or blocked.
78
+
79
+ ## 1. Over-block
80
+
81
+ Nothing is blocked — this is a read surface. The available harm is **misinforming a reader**, and
82
+ this change strictly reduces it in the direction that mattered (absence no longer reads as failure).
83
+
84
+ The mirror over-block is real and guarded: a genuinely-assessed period must never report `'N/A'` or
85
+ `assessable: false`, or a real alignment problem would be hidden as "no data" — strictly worse than
86
+ the original bug. Asserted by two tests (`a genuinely assessed period still reports a real letter
87
+ grade`, and the CLI's `a populated journal still prints a real graded score`).
88
+
89
+ **Caller sweep, run BEFORE writing this section** (the correction from #1659, where I wrote a
90
+ confident risk claim from the wrong measurement and CI falsified it): `alignmentScore()` has exactly
91
+ two production callers — `routes.ts:24305` (passes through verbatim) and `commands/intent.ts:431`
92
+ (now branches). Every `.grade` hit elsewhere in `src/` belongs to `DecisionQualityRecorder`'s
93
+ unrelated `right|wrong|unknown` grade, checked rather than assumed. Two tests asserted the old shape;
94
+ both updated.
95
+
96
+ ## 2. Under-block
97
+
98
+ **The mismatched windows are NOT fixed.** The early return uses `windowDays` while `alignmentScore()`
99
+ is hardcoded to 30. That divergence is what makes the CLI case reachable at all, and reconciling them
100
+ changes what the command reports for every user. Deliberately not folded in. <!-- tracked: CMT-1044 -->
101
+
102
+ **`score: 0` is retained on the unassessable case.** Changing it to `null` would be a breaking type
103
+ change for a field two consumers read; `assessable` is the additive signal instead. A consumer that
104
+ reads `.score` and ignores both `assessable` and `sampleSize` still sees a 0 — it can no longer see
105
+ an F, which is the part that read as a verdict.
106
+
107
+ **It does not improve alignment,** and it does not judge whether a cited principle was genuinely
108
+ consulted. Same honest limit as the increment before it.
109
+
110
+ ## 3. Level-of-abstraction fit
111
+
112
+ The honest state lives in the returned value, not in the renderer, so every consumer inherits it —
113
+ the route needed no change at all. The renderer's job is narrowed to *presenting* a state it no
114
+ longer has to infer. Had I fixed only the CLI, the API consumer (the one that is actually reachable
115
+ by default) would still have been lied to.
116
+
117
+ ## 4. Signal vs authority compliance
118
+
119
+ Pure signal. `docs/signal-vs-authority.md` is satisfied trivially: it produces a read-only score that
120
+ gates nothing, blocks nothing, and is consumed by one route and one command.
121
+
122
+ ## 4b. Judgment-point check (Judgment Within Floors standard)
123
+
124
+ None introduced. Two deterministic branches on a count.
125
+
126
+ ## 5. Interactions
127
+
128
+ - **`GET /intent/alignment`** — response gains `assessable`; `grade` may now be `'N/A'`. Additive plus
129
+ one widened union member.
130
+ - **`IntentDriftDetector.analyze()`** — untouched; drift scoring is a separate path.
131
+ - **`tests/unit/IntentDriftDetector.test.ts`** and **`tests/integration/drift-routes.test.ts`** — each
132
+ had a test asserting `grade === 'F'` on the empty case. **Both were encoding the defect**, not
133
+ merely stale. Updated with comments recording that, since a test that locks in a wrong answer is
134
+ the same instrument-honesty class as the defect itself.
135
+ - **`CapabilityIndex`** — unchanged; `intent` is already `INTERNAL_PREFIXES`.
136
+
137
+ ## 6. External surfaces
138
+
139
+ One API response shape change (additive field + widened union), one CLI rendering change. No config
140
+ key, no persisted state, no migration, no message to any user.
141
+
142
+ ## 6b. Operator-surface quality
143
+
144
+ The unassessable CLI output prints the reason (`No decisions logged — alignment cannot be assessed`)
145
+ in dim rather than a red grade, so it reads as an absence of data rather than an alarm. The component
146
+ breakdown is suppressed in that branch: four zeroed rows invite exactly the misreading being fixed.
147
+
148
+ ## 7. Multi-machine posture (Cross-Machine Coherence)
149
+
150
+ **Machine-local BY DESIGN.** The journal is a per-machine JSONL under `stateDir` and the score is
151
+ computed from it, so `assessable` answers "on this machine". An agent running on two machines has two
152
+ journals and two scores. That predates this change and is unaddressed here. <!-- tracked: CMT-1044 -->
153
+
154
+ ## 8. Rollback cost
155
+
156
+ Low. One union member, one boolean, one CLI branch, three test updates. No persisted state, no
157
+ migration; existing journal rows are read unchanged. Reverting restores the fabricated F.
158
+
159
+ ## Phase 5 — Second-pass review
160
+
161
+ Not a gate, sentinel, guard or watchdog; holds no block/allow authority; touches no session lifecycle
162
+ or trust level. The high-risk trigger list is not engaged. Author lenses, disclosed:
163
+
164
+ **Adversarial — "how would I make this useless?"** Three ways, all closed and asserted: report `'F'`
165
+ again (refusal 1), make `assessable` constant (refusal 2), or let the renderer ignore it entirely
166
+ (refusal 3 — the one that was genuinely open until I wrote the CLI test).
167
+
168
+ **"Would it have caught the incident?"** The incident here is my own: I read `topPrinciples: []` and
169
+ `score: 0 (F)` on a journal I already knew was empty, and had to reason my way to "that F is
170
+ meaningless" instead of being told. With this, the surface says it.
171
+
172
+ **"Symptom or cause?"** Cause, for the reporting defect — absence can no longer render as a verdict
173
+ because the type now has somewhere honest to put it. Symptom-level for the window mismatch, which is
174
+ named and left.
175
+
176
+ **Weakest point:** the CLI case is genuinely narrow (`--window > 30`), and I over-claimed its reach
177
+ twice before testing settled it. The API consumer is the one that matters by default. An artifact
178
+ claiming broad user impact here would be overstating it, so this one does not.
@@ -0,0 +1,227 @@
1
+ # Side-Effects Review — a decision log that refuses a decision which does not say why
2
+
3
+ **Version / slug:** `decision-journal-principle-required`
4
+ **Date:** `2026-07-26`
5
+ **Author:** `Echo (instar-dev agent)`
6
+ **Second-pass reviewer:** `see Phase 5`
7
+
8
+ ## Summary of the change
9
+
10
+ Operator directive (topic 29723, 20:4xZ): decide against the goal hierarchy instead of escalating,
11
+ log the reasoning for later review, and **enforce it via infrastructure rather than by remembering**.
12
+
13
+ Investigating what to build produced the finding that the machinery already exists — hierarchy
14
+ (`GET /intent/org`, `/intent/tradeoff-resolve`), recorder (`DecisionJournal`, `POST /intent/journal`),
15
+ and drift detector (`IntentDriftDetector`) — and had recorded **zero decisions, ever**. What is
16
+ missing is not machinery; it is anything that forces its use or reveals its disuse.
17
+
18
+ Then, using it, I produced the defect this PR fixes:
19
+
20
+ - I POSTed `reasoning` and `checkedAgainst`. **Neither is a field.** The route spread `...rest`
21
+ straight through with no validation, so both were persisted where no reader consumes them. The
22
+ write returned 201. I then told my operator the reasoning was recorded — believing it.
23
+ - `principle`, the typed field documented as *"Which AGENT.md principle or intent guided the choice"*,
24
+ was empty on all five entries.
25
+ - `stats()` therefore reported `topPrinciples: []` — **byte-identical to an empty journal.** The
26
+ instrument built to detect unreasoned decisions could not distinguish its own worst case from a
27
+ clean slate.
28
+
29
+ Adds `validateDecisionSubmission()` (pure) + `principledCount`/`unprincipledCount` on
30
+ `DecisionJournalStats`, wires the validator into `POST /intent/journal`, and registers the behaviour
31
+ change for new and existing agents.
32
+
33
+ ## Refusal evidence (constraint 2)
34
+
35
+ ```
36
+ REFUSAL 1 — unwire the validator from the ROUTE (`if (false && !verdict.ok)`)
37
+ UNIT: Tests 11 passed (11) <-- the blindness, reproduced deliberately
38
+ INTEGRATION: × the ROUTE refuses a decision that names no principle
39
+ × the ROUTE refuses fields no reader consumes
40
+ × a refused submission writes NOTHING
41
+ × the refusal message tells the caller where the content belongs
42
+ Tests 4 failed | 2 passed (6)
43
+
44
+ REFUSAL 2 — swallow unknown fields instead of refusing
45
+ × a field no reader consumes is REFUSED, not swallowed
46
+ × missing-required outranks unknown-fields, and still reports both
47
+ × (+3 integration) Tests 5 failed | 12 passed (17)
48
+
49
+ REFUSAL 3 — drop `principle` from the required set
50
+ × a decision naming no guiding principle is REFUSED
51
+ × a blank or whitespace principle does not satisfy the requirement
52
+ × (+3) Tests 5 failed | 12 passed (17)
53
+
54
+ REFUSAL 4 — revert stats to the ambiguous shape
55
+ × entries with no principle are COUNTED, not silently absent → expected +0 to be 2
56
+ × an empty journal is DISTINGUISHABLE from an unprincipled one → expected +0 not to be +0
57
+ × (+2) Tests 4 failed | 13 passed (17)
58
+ ```
59
+
60
+ Restored: **188 passed (188)** across the five affected files, `tsc --noEmit` exit 0.
61
+
62
+ **REFUSAL 1 is the finding, and it is the second occurrence tonight of the same class.** One feature
63
+ earlier (#1658) I emptied a route's registry and all 19 unit tests passed — module guarded, wiring
64
+ not. So here I went looking for it deliberately: unwiring the validator leaves **11/11 unit tests
65
+ green** while the route accepts everything it is supposed to refuse. The integration file exists for
66
+ exactly this assertion and nothing else.
67
+
68
+ ## Decision-point inventory
69
+
70
+ | point | classification | note |
71
+ |---|---|---|
72
+ | missing required field → refuse | `invariant` | Deterministic key/type check. No model call. |
73
+ | unknown field → refuse | `invariant` | Allowlist comparison; allowlist asserted against the documented field set by test. |
74
+ | missing-required outranks unknown-fields | `invariant` | Both are still reported, so one round trip surfaces both problems. |
75
+ | machine dispatch path exempt | `invariant` | Scoped by callsite, not by inspecting content. |
76
+
77
+ No judgment points. No LLM. Nothing is inferred about the *quality* of a cited principle — only that
78
+ one was cited. Judging whether a decision genuinely followed the principle it names is the drift
79
+ detector's job and is deliberately not attempted here.
80
+
81
+ ## 1. Over-block
82
+
83
+ **This is the section that shaped the design.** A blanket requirement would have been wrong.
84
+
85
+ `journal.log()` has two callers: the HTTP route (agent-authored) and `DispatchDecisionJournal
86
+ .logDispatchDecision` via `AutoDispatcher`, which writes auto-applied dispatch decisions whose own
87
+ documented shape is `{ dispatchDecision: 'accept', reasoning: 'auto-applied' }`. Those have no
88
+ principle to cite. **Enforcing inside `log()` would have broken automatic dispatch to buy nothing**,
89
+ so the refusal lives at the route and `log()` is untouched — asserted by a test that the module path
90
+ still accepts an entry with no principle.
91
+
92
+ Residual over-block risk on the agent path: a caller with a legitimate decision and genuinely no
93
+ guiding principle now gets a 400. I accept this deliberately — under the operator directive, a
94
+ decision made during operations without checking the stated goals is precisely what we are
95
+ eliminating. The failure is loud, names the missing field, and is fixed by adding one string.
96
+
97
+ Unknown-field rejection is a strict-schema change and could in principle break an existing caller.
98
+ The refusal names the offending keys and the correct destination, so a broken caller is told exactly
99
+ what to change rather than failing opaquely.
100
+
101
+ **CORRECTION — my first version of this section was too confident, and CI proved it.** I wrote that
102
+ "the journal had `count: 0` on this agent, so no caller has ever successfully written to it" and
103
+ treated that as evidence the back-compat risk was near-nil. That measurement was about the *live
104
+ agent's data file*; it says nothing about *callers in the codebase*. There was one:
105
+ `tests/integration/intent-routes.test.ts` POSTed three decisions with no `principle`, and my change
106
+ refused all three, so the journal file was never created and a later read failed `ENOENT`.
107
+
108
+ I did not catch it locally because I ran a targeted file set that did not include that test. CI did.
109
+ Two failures, both mine, neither a flake. Resolution: the tests now supply a `principle` — which is
110
+ the *correct* fix, since they were writing exactly the unreasoned decisions this gate exists to
111
+ refuse, and their failure is the refusal working on a real caller rather than a synthetic one.
112
+
113
+ The honest generalisation: "no rows in the data file" is not evidence of "no callers in the code".
114
+ Those are different questions and I conflated them. A repo-wide sweep afterwards found three files
115
+ POSTing to the route and six asserting on `stats()` shape; two needed changes, and both are fixed.
116
+
117
+ ## 2. Under-block
118
+
119
+ **It does not force the check, only the citation.** Nothing here fires at the moment a decision is
120
+ made and puts the hierarchy in front of the agent. An agent can still decide without consulting
121
+ `GET /intent/org` and then name a principle after the fact. This PR makes it impossible to *record*
122
+ a decision that claims no guiding intent; it does not make it impossible to *make* one. The
123
+ consult-side trigger is genuinely separate work and is not bundled. <!-- tracked: CMT-1044 -->
124
+
125
+ **It does not judge the principle.** Any non-empty string satisfies the requirement. A caller citing
126
+ "because I felt like it" passes. Judging alignment is `IntentDriftDetector`'s job.
127
+
128
+ **The five existing entries are not repaired.** They keep their unread `reasoning`/`checkedAgainst`
129
+ keys. The new counters will report them as unprincipled, which is accurate. Rewriting my own history
130
+ to look better than it was is the opposite of this tier's purpose.
131
+
132
+ **`GET /intent/journal` (read) is unchanged** and will still return the legacy rows with their
133
+ unread fields. Nothing marks those fields as unread on read.
134
+
135
+ ## 3. Level-of-abstraction fit
136
+
137
+ The validator is pure — no `fs`, no clock, no config, no server import — so the refusal is unit
138
+ testable without a server, and the route owns transport only. This mirrors the structure used one
139
+ feature earlier and, more importantly, is what made REFUSAL 1 possible to demonstrate: a pure
140
+ validator can be perfect while nothing calls it, which is precisely the failure mode being guarded.
141
+
142
+ The counters live on `stats()` rather than in a new surface, because the ambiguity being fixed is a
143
+ property of that existing return value. A new endpoint would have left the misleading one in place.
144
+
145
+ ## 4. Signal vs authority compliance
146
+
147
+ This **is** an authority — it blocks a write — so `docs/signal-vs-authority.md` applies directly
148
+ rather than trivially. It qualifies because the logic is deterministic and total: an allowlist
149
+ comparison and a set of required-key checks, no heuristics, no model, no inference about content.
150
+ That is the category the principle permits to hold blocking authority. Every uncertain input
151
+ (`null`, `undefined`, a string, a number) resolves to a refusal with a named reason rather than a
152
+ throw or a pass — asserted by test.
153
+
154
+ The blocked path writes nothing: a refused submission leaves `count: 0`, asserted by integration
155
+ test. A refusal that still recorded the row would be worse than no refusal, because the journal
156
+ would carry entries the caller was told were rejected.
157
+
158
+ ## 4b. Judgment-point check (Judgment Within Floors standard)
159
+
160
+ None introduced. Every branch is a deterministic key or type comparison.
161
+
162
+ ## 5. Interactions
163
+
164
+ - **`DispatchDecisionJournal` / `AutoDispatcher`** — deliberately NOT gated (see §1). Asserted.
165
+ - **`IntentDriftDetector`** — reads the journal; gains higher-quality input (entries now carry
166
+ `principle`) and is otherwise untouched. Its 16 tests pass unchanged.
167
+ - **`instar intent reflect` / `commands/intent.ts`** — calls `journal.log()` directly, module path,
168
+ unaffected by the route gate.
169
+ - **`tests/unit/DecisionJournal.test.ts`** — one test pinned the exact empty-`stats()` object shape
170
+ and now asserts the two new counters. This is a shape assertion updated to a new shape, not a
171
+ behavioural assertion weakened; the comment records why the zero values matter.
172
+ - **`CapabilityIndex`** — no change needed. `intent` is already classified under
173
+ `INTERNAL_PREFIXES` ("surfaced inside `evolution` subsystems"); this adds no new route prefix.
174
+
175
+ ## 6. External surfaces
176
+
177
+ `POST /intent/journal` changes response behaviour: submissions that previously returned 201 may now
178
+ return 400 with `{ error, reason, unknownFields, missingFields }`. `GET /intent/journal/stats` gains
179
+ two fields (additive). No config key, no new route, no persisted-state migration, no message to any
180
+ user. No credentials or content beyond what the caller submitted appear in any response.
181
+
182
+ ## 6b. Operator-surface quality
183
+
184
+ The refusal message names the offending fields AND the correct destination (`context` for reasoning,
185
+ `principle` for guiding intent) plus the full writable-field list. A refusal that only says "invalid"
186
+ relocates the failure rather than fixing it; asserted by a test that the message mentions `context`.
187
+
188
+ ## 7. Multi-machine posture (Cross-Machine Coherence)
189
+
190
+ **Machine-local BY DESIGN.** The journal is a per-machine JSONL under `stateDir`; the validator is
191
+ pure and stateless, so the refusal is identical on every machine with no coordination required.
192
+ There is no replication, no lease interaction, no generated URL, and no cross-machine read. Honest
193
+ limitation: an agent running on two machines keeps two separate decision journals, so
194
+ `principledCount` answers "on this machine". That predates this change and is not addressed here.
195
+ <!-- tracked: CMT-1044 -->
196
+
197
+ ## 8. Rollback cost
198
+
199
+ Low. One pure function, two counters, one route guard, plus doc/migration lines. No persisted-state
200
+ change and no data migration — existing rows are read unchanged. Reverting restores the permissive
201
+ route; already-written entries remain valid under both versions.
202
+
203
+ ## Phase 5 — Second-pass review
204
+
205
+ This change **does** hold block/allow authority on a write path, so the high-risk trigger list is
206
+ engaged. It does not touch session lifecycle, messaging, dispatch, trust levels, or recovery. Author
207
+ lenses, disclosed:
208
+
209
+ **Adversarial — "how would I make this useless?"** Three ways, all closed and asserted: let the
210
+ route skip the validator (REFUSAL 1 — the real one); accept unknown keys silently (REFUSAL 2); drop
211
+ the principle requirement (REFUSAL 3). A fourth — refuse but write the row anyway — is closed by the
212
+ `a refused submission writes NOTHING` test.
213
+
214
+ **"Would it have caught the incident?"** Yes, on the first attempt, which is the strongest thing I
215
+ can say for it. My literal opening submission is now a test fixture and returns 400 naming
216
+ `checkedAgainst` and `reasoning`. I would have lost seconds instead of finding out by reading my own
217
+ file afterwards and having already told my operator otherwise.
218
+
219
+ **"Symptom or cause?"** Partly cause, and I want this stated plainly rather than implied: it makes
220
+ an unreasoned decision unrecordable, which is a real structural gate. It does not make an unreasoned
221
+ decision unmakeable. The operator asked for decisions checked against the hierarchy; this delivers
222
+ the enforcement half of "log the reasoning" and none of the consult half.
223
+
224
+ **Weakest point:** the requirement is satisfiable by any non-empty string. Nothing distinguishes a
225
+ genuinely-consulted principle from a plausible-sounding one typed to clear the gate — and the agent
226
+ clearing it is the same one whose unreliability motivated the gate. That limit is inherent to a
227
+ deterministic check and is the reason the consult-side trigger is not optional future work.