tickmarkr 1.85.0 → 1.87.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (78) hide show
  1. package/README.md +4 -2
  2. package/dist/adapters/catalog-remote.d.ts +64 -0
  3. package/dist/adapters/catalog-remote.js +287 -0
  4. package/dist/adapters/catalog.d.ts +108 -0
  5. package/dist/adapters/catalog.js +189 -0
  6. package/dist/adapters/claude-code.js +5 -3
  7. package/dist/adapters/fake.js +33 -4
  8. package/dist/adapters/model-lints.d.ts +25 -5
  9. package/dist/adapters/model-lints.js +184 -50
  10. package/dist/adapters/model-windows.d.ts +31 -0
  11. package/dist/adapters/model-windows.js +69 -0
  12. package/dist/adapters/prompt.d.ts +5 -1
  13. package/dist/adapters/prompt.js +13 -4
  14. package/dist/adapters/registry.d.ts +34 -27
  15. package/dist/adapters/registry.js +215 -112
  16. package/dist/adapters/types.js +15 -3
  17. package/dist/brand.d.ts +5 -1
  18. package/dist/brand.js +18 -2
  19. package/dist/cli/commands/doctor.d.ts +3 -0
  20. package/dist/cli/commands/doctor.js +43 -21
  21. package/dist/cli/commands/fleet.d.ts +7 -0
  22. package/dist/cli/commands/fleet.js +94 -74
  23. package/dist/cli/commands/init.js +118 -5
  24. package/dist/cli/commands/plan.js +11 -1
  25. package/dist/cli/commands/resume.js +7 -1
  26. package/dist/cli/commands/status.js +44 -18
  27. package/dist/compile/gsd.d.ts +2 -1
  28. package/dist/compile/gsd.js +68 -2
  29. package/dist/compile/native.d.ts +14 -0
  30. package/dist/compile/native.js +168 -12
  31. package/dist/config/config.d.ts +20 -5
  32. package/dist/config/config.js +96 -64
  33. package/dist/config/fleet-overlay.d.ts +25 -20
  34. package/dist/config/fleet-overlay.js +195 -77
  35. package/dist/config/fleet-why.d.ts +23 -0
  36. package/dist/config/fleet-why.js +42 -0
  37. package/dist/drivers/herdr.d.ts +1 -0
  38. package/dist/drivers/herdr.js +70 -19
  39. package/dist/gates/acceptance.js +7 -2
  40. package/dist/gates/llm.d.ts +0 -1
  41. package/dist/gates/llm.js +5 -30
  42. package/dist/gates/review.d.ts +2 -1
  43. package/dist/gates/review.js +9 -7
  44. package/dist/gates/run-gates.d.ts +1 -0
  45. package/dist/gates/run-gates.js +21 -2
  46. package/dist/gates/verdict-cause.d.ts +4 -0
  47. package/dist/gates/verdict-cause.js +63 -0
  48. package/dist/graph/schema.d.ts +6 -0
  49. package/dist/graph/schema.js +8 -5
  50. package/dist/route/preference.d.ts +1 -1
  51. package/dist/route/preference.js +8 -1
  52. package/dist/route/router.d.ts +0 -5
  53. package/dist/route/router.js +16 -20
  54. package/dist/run/consult.d.ts +6 -0
  55. package/dist/run/consult.js +49 -26
  56. package/dist/run/daemon.js +81 -20
  57. package/dist/run/journal.js +87 -7
  58. package/dist/tui/cockpit/capture.d.ts +12 -0
  59. package/dist/tui/cockpit/capture.js +37 -1
  60. package/dist/tui/cockpit/components.js +8 -8
  61. package/dist/tui/cockpit/theme.d.ts +32 -26
  62. package/dist/tui/cockpit/theme.js +11 -5
  63. package/dist/tui/ink/components.d.ts +0 -15
  64. package/dist/tui/ink/components.js +0 -17
  65. package/dist/tui/ink/fleet-app.d.ts +4 -1
  66. package/dist/tui/ink/fleet-app.js +134 -13
  67. package/fixtures/gateway-models.json +1 -0
  68. package/fixtures/sample.native.md +1 -1
  69. package/package.json +1 -1
  70. package/skills/tickmarkr-overseer/SKILL.md +464 -34
  71. package/skills/tickmarkr-overseer/scripts/watch-artifacts.sh +86 -0
  72. package/skills/tickmarkr-overseer/scripts/watch-panes.sh +1 -1
  73. package/dist/tui/ink/studio-app.d.ts +0 -59
  74. package/dist/tui/ink/studio-app.js +0 -320
  75. package/dist/tui/save.d.ts +0 -38
  76. package/dist/tui/save.js +0 -96
  77. package/dist/tui/staging.d.ts +0 -29
  78. package/dist/tui/staging.js +0 -78
@@ -22,12 +22,26 @@ Requires `HERDR_ENV=1`; if unset, say so and stop.
22
22
  status, and either ADOPT the
23
23
  existing orchestrator (updated brief, re-armed watchers) or, if the old hierarchy is dead, archive the
24
24
  stale brief and build fresh.
25
+ 0a. **READ THE PROJECT MEMORY BEFORE YOU START — it already contains discipline you are about to re-earn.**
26
+ `~/.claude/projects/<cwd-slug>/memory/` (slug = the absolute cwd with `/` → `-`). Read `MEMORY.md`, then
27
+ `ls` the topic entries and open every one whose name concerns METHOD or DISCIPLINE rather than a shipped
28
+ milestone — names like `*-discipline`, `*-drill`, `*-parity`, `*-least-permission`, `context-reset-*`,
29
+ `consults-*`, `agent-*`.
30
+ **Earned 2026-08-04, expensively.** That directory held `…-falsification-drill-discipline.md`, written
31
+ three weeks earlier: *"a gate or grep-pin is assumed WRONG until a falsification drill proves it bites…
32
+ run the drill that should redden it and SEE the red before trusting green."* That is Evidence discipline
33
+ rule 11 below, verbatim in substance. Nothing surfaced it, so an orchestrator and an overseer re-derived
34
+ it independently, twice, inside one hour — and the overseer then filed it as a NEW rule into a
35
+ mission-scoped brief. **A memory that exists and is never opened costs more than one that was never
36
+ written, because everyone assumes the lesson is somewhere.** Entries may predate a project rename; search
37
+ by concept, not by the current product name.
25
38
  1. Load the `herdr` skill. `herdr pane list` to map the workspace — the focused pane is yours. Rename your
26
39
  tab OVERSEER; create ONE tab ORCHESTRATOR.
27
40
  **Live tab labels (standing operator rule, 2026-07-12):** on every decision or state change (role
28
41
  handoff, task done/merged, run end) rename the affected tabs — and keep labels SHORT: the role as the
29
42
  main name plus at most ONE hot-state token. Vocabulary: ORCH carries the milestone and progress
30
- fraction (`ORCH · v1.19 4/5`, updated on every task-done); WORKERS carries the task token (tickmarkr
43
+ fraction (`ORCH · v1.19 4/5`, updated on every task-done); tickmarkr opens ONE TAB PER TASK, labelled
44
+ with the task id and holding that task's worker plus its judge/review/consult panes (tickmarkr
31
45
  updates it). Never long context strings or ✓-chains.
32
46
  2. **Orchestrator**: Launch the orchestrator with your agent host. Spawning on current herdr is two-step — the one-shot `agent start --cwd` form was removed in the herdr CLI redesign and now fails with `unknown option` (OBS-138): first create the pane with `herdr tab create --workspace <ws> --cwd <repo> --label "ORCH · <version>"` and parse `result.root_pane.pane_id` from its JSON, then start the agent in it. For Claude Code, use `herdr agent start orchestrator --kind claude --pane <root-pane-id> -- --permission-mode bypassPermissions` (append `--model <m>` after the `--` if the operator has a policy). For Codex, use `herdr agent start orchestrator --kind codex --pane <root-pane-id> -- --dangerously-bypass-approvals-and-sandbox` (add `--model <m>` to specify the model). The unsandboxed flag is REQUIRED: codex's `workspace-write` sandbox keeps `.git` refs read-only, so a sandboxed orchestrator's `tickmarkr run` dies at integration-branch creation — do not downgrade it. Workers you never spawn — tickmarkr spawns its own visible worker panes. Auxiliary agents you do spawn (consultants, reviewers, scouts) follow the same forms: never launch a claude session in plan mode or default permission mode for autonomous work — both stall on per-command approval prompts nobody is watching; claude is always `--permission-mode bypassPermissions`, and a read-only codex consultant may use `--sandbox read-only`.
33
47
  3. **Standing instructions travel as a brief FILE, never as pane text** — PTY input truncates at ~1024B and a
@@ -35,32 +49,110 @@ Requires `HERDR_ENV=1`; if unset, say so and stop.
35
49
  (inside the tickmarkr state dir — already self-gitignored, no exclude step needed), then send one line:
36
50
  `herdr pane run <orch> "Read .tickmarkr/overseer/ORCH-BRIEF.md and follow it exactly."` The brief MUST contain: the
37
51
  mission, the pane mechanics below, rules 1–2, and require a verbatim one-sentence acknowledgment of the
38
- human-checkpoint rule before anything is dispatched. Delete the dir at mission end.
52
+ human-checkpoint rule before anything is dispatched.
53
+ **⚠ HARVEST BEFORE YOU DELETE.** At mission end the brief dir goes — but a long mission accumulates
54
+ *method guards* in that brief (how to know a thing, not what is true of this spec), and deleting them
55
+ re-earns each one at full price on the next mission. So before removing the dir: lift every durable,
56
+ mission-independent guard into **this skill** (Evidence discipline, below) or the project's `CLAUDE.md`,
57
+ and only then delete. A guard's home must outlive the mission that earned it. The project ledger does
58
+ NOT count as that home — `CLAUDE.md` itself says planning records are read-only archives and current
59
+ guidance belongs in the memory file or the shipped docs.
39
60
  4. Arm the watcher (Supervision). Report the hierarchy map (pane ids + names) to the user.
40
61
 
41
- ## Supervising tickmarkr as the executor
62
+ ## Supervising tickmarkr as the executor — WHO DOES WHAT
42
63
 
43
- When the mission runs `/tickmarkr-auto` (tickmarkr dispatches the workers), supervision changes shape:
64
+ When the mission runs `/tickmarkr-auto` (tickmarkr dispatches the workers), supervision changes shape —
65
+ and the first thing to get right is that **almost none of it is yours**.
44
66
 
45
- - **Give the run a live surface.** `tickmarkr run` is stdout-silent until run-end by design — split a pane in the
46
- ORCHESTRATOR tab running `tickmarkr status --watch`. Narration also arrives as herdr notifications.
47
- - **Watch the journal, not the panes.** The append-only journal
48
- (`.tickmarkr/runs/<runId>/journal.jsonl`) is the
49
- source of truth. Arm a background watcher on `run-end` / `task-human` / `task-failed` / `consult-verdict`
50
- events; never sleep-poll inside an agent turn.
67
+ **THE ORCHESTRATOR OWNS THE LOOP. You rule and record. You do not drive.**
68
+
69
+ | | ORCHESTRATOR | OVERSEER |
70
+ |---|---|---|
71
+ | `compile` · `plan` · `run` · `resume` | **owns** | never |
72
+ | journal watchers, live surface, dialog watchers | **owns** | watches the ORCHESTRATOR, not the run |
73
+ | orphan sweeps, worker pane hygiene | **owns** | — |
74
+ | reading a gate failure and assembling its evidence | **owns** | reads the file it writes |
75
+ | **deciding** a gate, spend, or ship | never | **owns** |
76
+ | executing `tickmarkr approve` after a ruling | **owns** | never |
77
+ | git writes, the ledger, records, the operator | never | **owns** |
78
+
79
+ **Why this is a rule and not a preference — it has a measured cost.** On 2026-08-04 an OVERSEER ran the
80
+ loop itself: compile, plan, run, resume, approve, sweeps, even source fixes. The operator caught it —
81
+ *"you are doing the job of the orchestrator, the orchestrator is the one should be taking care of the
82
+ gates"* — and the receipt was a pair of numbers: **the ORCHESTRATOR sat at 369K tokens while the OVERSEER
83
+ burned 747K.** Collapsing two tiers into one does not just waste a seat; it **burns the context of the
84
+ seat that cannot be replaced cheaply**, because the orchestrator can `/clear` against a brief while the
85
+ overseer holds the mission's judgment. A tier collapse is therefore a context leak with a delay fuse.
86
+
87
+ **The tell, and you will not notice it from inside:** if you are typing `tickmarkr resume`, or reading a
88
+ journal tail to decide what happens next, or sweeping orphans — you have taken the loop. Hand it back.
89
+
90
+ ### What the ORCHESTRATOR does, and what you require of it
91
+
92
+ - **A live surface.** `tickmarkr run` is stdout-silent until run-end by design, and the run spawns its own
93
+ watch board (`role: "watch"`, one per run) — **look for that pane before building anything.** Do NOT use
94
+ `tickmarkr status --watch` as the surface: `status <runId>` has reported the WRONG run, so a board built
95
+ on it shows a previous milestone's numbers under the current run's id.
96
+ - **The journal is the source of truth**, not panes. Watchers go on `run-end` / `task-human` /
97
+ `task-failed` / `consult-verdict`; never sleep-poll inside an agent turn. **Never key a watcher on an
98
+ agent's `done`** — that is turn end and fires the moment a seat finishes acknowledging you.
51
99
  - **Daemon liveness ≠ journal activity.** A dead daemon emits no events, so journal watchers sleep through
52
- its death. `tickmarkr status` prints last-event age + daemon pid liveness; check it before diagnosing a stall.
53
- Recovery is `tickmarkr resume <runId>` — crash-safe by design (journal replay restores attempt counts and
54
- consult channel bans).
55
- - **Gate quiet ≠ idle.** Between `worker-result` and the batched `gate-result`s tickmarkr runs shell gates plus a
56
- headless LLM judge/review with little visible signal — check the journal timestamps before intervening.
57
- - **Classify gate failures before reacting.** The same fingerprint failing across DIFFERENT workers, or a
58
- scope/test catch-22 (attempt N edits a file → scope gate fails; attempt N+1 leaves it → test gate fails),
59
- is a PLAN defect: widen `files_modified` in the phase PLAN, recompile the phase dir after the run ends or
60
- the task parks, release (`human → pending`), resume. A cross-vendor review rejection with concrete findings
61
- is a REAL defect — let the escalation ladder work.
62
- - **Dialog watchers go stale per attempt.** Every retry/escalation may spawn a new pane; re-arm dialog
63
- watchers on each `task-dispatch` journal event.
100
+ its death. Liveness comes from the lock's OWN pid (`kill -0`), never a command-name grep. Recovery is
101
+ `tickmarkr resume <runId>` — **the orchestrator's command, not yours** — and note that resume REPLAYS the
102
+ journal's `baseRef`, so a fix landed on the base branch is unreachable by the running run.
103
+ - **SWEEPING A WORKER ORPHANS ITS CLEANUP, NOT JUST ITS WORK — kill the process GROUP.** Measured
104
+ 2026-08-06: a worker running a legitimate load experiment had spawned CPU burners and held a trailing
105
+ `kill` line. It was SIGTERM'd as an orphan; **the cleanup never ran**, and **59 surviving burner shells
106
+ drove load to 243** — which then timed out a 2.4-second test at 20 seconds, failed the run's tip-verify,
107
+ and was initially blamed on an unrelated known defect. Seven workers were swept that day under a rule
108
+ that treats sweeping as pure hygiene, and nothing in it looks for pending cleanup.
109
+ **The fix is mechanical, not vigilance:** kill the process GROUP so forked children die with the parent,
110
+ and identify the target **by PID from a parse — never by name pattern.** A pattern matching a script path
111
+ also matches any supervisor carrying that path in its own argv, so `pgrep -f <script>` kills the watchdog
112
+ along with the watched (measured the same day, on the overseer's own dialog watcher).
113
+ **After any sweep, verify what SURVIVED, not just what died** — the supervisor, the daemon, and the
114
+ current attempt's worker.
115
+ - **Gate quiet ≠ idle.** Between `worker-result` and the batched `gate-result`s, shell gates plus a
116
+ headless judge/review run with little visible signal. Clock the CURRENT phase: a worker heartbeat is
117
+ stale by design once gates start, and clocking the wrong one makes a healthy gate read as a stalled
118
+ worker — the false alarm that gets a good run killed.
119
+ - **Classify gate failures before reacting.** The same fingerprint across DIFFERENT workers, or a
120
+ scope/test catch-22 (attempt N edits a file → scope fails; attempt N+1 leaves it → test fails), is a
121
+ PLAN defect. So is any blocker OUTSIDE the task's `files[]` — no worker can fix it and a retry is
122
+ knowingly wasted. A cross-vendor review rejection with concrete findings is a REAL defect; let the
123
+ ladder work.
124
+ - **Dialog watchers go stale per attempt.** Every retry may spawn a new pane; re-arm on each
125
+ `task-dispatch`.
126
+
127
+ ### What YOU do
128
+
129
+ Read the evidence file the orchestrator writes, rule on it against your pre-committed release criterion,
130
+ record the ruling with what it set aside, and hand the ruling back for execution. That is the whole job,
131
+ and it is the only work that cannot be delegated — which is exactly why nothing else should occupy you.
132
+
133
+ **And before you write the ruling: check that YOUR REMEDY is buildable inside the task's `files[]`.** The
134
+ orchestrator is told to classify a blocker outside a task's scope as a plan defect. Nothing tells the
135
+ OVERSEER that *its own instruction* can be that defect — so it arrives carrying your authority and is not
136
+ re-examined.
137
+
138
+ **Measured 2026-08-06.** A ruling directed a task to make a function parameter required. Verified
139
+ afterwards: two callers pass one argument, in a file **no task in the graph owned**. The worker would have
140
+ been trapped — edit it and fail `scope`, leave it and fail `build` — and the failure would have surfaced
141
+ as a worker defect at the next park. Worse, the unowned file was exactly the relation that task existed to
142
+ detect: **the ruling would have made the worker commit the violation the task was being built to catch.**
143
+
144
+ > Run the callers before you write the remedy: locate the symbol, **enumerate** every caller and asserting
145
+ > test, classify each against `files[]`, and resolve who owns the out-of-scope ones. A sweep for this class
146
+ > then found **five of six remaining tasks exposed**, so it is a shape, not an accident.
147
+
148
+ ### Context is a supervised resource, for BOTH tiers
149
+
150
+ Arm a context watcher on the orchestrator at spawn time and treat a threshold wake as a first-class event:
151
+ finish the step, write a handoff, `/clear` **plus a fresh brief — never `/compact`**, because a compaction
152
+ is a lossy summary nobody trusts while a clean session re-oriented from disk-verifiable state is reliable.
153
+ **Do the same for yourself before you are forced to**: write the handoff while your judgment is still
154
+ good, not after. If your own context cannot be read by the watcher, say so to the operator and ask for the
155
+ number — an unmeasured budget is not a small budget.
64
156
 
65
157
  ## Pane mechanics that bite
66
158
 
@@ -68,11 +160,36 @@ When the mission runs `/tickmarkr-auto` (tickmarkr dispatches the workers), supe
68
160
  by bracketed-paste on long payloads. Robust sequence: read the pane (bare prompt required) → send-text →
69
161
  sleep 2–3s → send-keys Enter → read back (input empty / agent `working`). Never report "briefed" without
70
162
  the read-back. Long content goes in a brief file, never pane text.
163
+ **PROBE THE READ-BACK WITH THE SHORTEST DISTINCTIVE TOKEN — a commit hash, a pid, an OBS id — NEVER a
164
+ sentence.** A long phrase crosses the pane's render wrap boundary, so grepping for it returns zero on a
165
+ message that arrived intact, and **a badly-probed successful send is byte-identical to a truncated one.**
166
+ Both natural reactions to that false negative are wrong: re-sending duplicates the message into the
167
+ target's queue, and escalating reports a delivery failure that never happened. Measured 2026-08-06
168
+ (OBS-396): a grep for the full sentence returned 0 while a grep for one word of the same sentence
169
+ returned 1. This trap lives *inside* the verification step above, which is why it survives — the rule
170
+ that is supposed to catch dropped sends is the rule that manufactures the phantom.
71
171
  - **Guard-before-Enter** (race-safe prompt answering): chain with `&&` — pane get shows `blocked` && pane
72
172
  read shows the expected option under the cursor && only then send-keys. If no longer `blocked`, someone
73
173
  already answered; do nothing.
74
- - `herdr wait agent-status` exits 1 on timeout, 0 on match — but ALSO 0 (with an error JSON) when the pane is
75
- GONE. Never chain `wait && act` without confirming the pane exists.
174
+ - **AGENT NAMES ARE GLOBAL ACROSS WORKSPACES — verify a seat you spawned by PANE ID, never by name.**
175
+ Names must be unique among live agents *everywhere*, not within your workspace, so another workspace can
176
+ already hold `opus`, `sol`, `reviewer` or `orch`. When it does, your `agent start` **fails**, your pane
177
+ is left a bare shell, and `agent list` / `agent read` / `agent prompt` for that name then resolve to the
178
+ **stranger's seat**. Measured 2026-08-06 (OBS-392): a spawn of `fable` collided with a live seat in
179
+ another workspace; `agent list` reported `fable -> blocked` and it was read as *this* seat coming up
180
+ blocked. It was an operator research session sitting on a *"Resume full session?"* prompt. One more
181
+ command would have submitted a brief into it. **Namespace every name you pick** (`fable-v187`, not
182
+ `fable`), **treat a failed `agent start` as fatal at the call site** rather than inferring it later from
183
+ a status read — the status read is exactly what the collision corrupts — and print the `workspace_id`
184
+ and `cwd` columns before dispatching to any name. Same class as the liveness rule: a matcher broader
185
+ than the thing it names finds things that are not it, and its output is shaped exactly like a right
186
+ answer.
187
+ - **A dead pane accepts your dispatch and reports success.** `herdr wait agent-status` exits 1 on timeout, 0
188
+ on match — but ALSO 0 (with an error JSON) when the pane is GONE. So does `pane run`: sending to a vanished
189
+ pane prints `{"error":{"code":"pane_not_found"}}` and **still exits 0**, so `pane run … >/dev/null && echo
190
+ sent` reports a delivery that never happened. Never chain `wait && act` or trust a send's exit status —
191
+ confirm the pane exists and read it back. An orchestrator's pane can vanish mid-mission without any event
192
+ reaching you; the first symptom is a dispatch into nothing.
76
193
  - Stale typed input is unclearable via CLI — supersede it:
77
194
  `pane run "<-- disregard everything before this arrow (stale draft). ACTUAL: <message>"`.
78
195
 
@@ -90,21 +207,334 @@ orchestrator gets a 90s grace window to handle worker blocks first. For long par
90
207
  `herdr wait agent-status <pane> --status <s> --timeout <ms>` beats the watcher. When parking a human
91
208
  checkpoint, also fire `herdr notification show "HUMAN CHECKPOINT: <gate>" --sound request`.
92
209
 
210
+ **⚠ THIS WATCHER KEYS ON `agent_status`, AND `agent_status` IS A PROXY THAT FAILS IN BOTH DIRECTIONS.**
211
+ Measured 2026-08-06 on ONE pane inside TEN MINUTES: a worker wedged behind a CLI's modal trust prompt
212
+ reported **`idle`** (not `blocked`), and the same pane minutes later reported **`done`** while demonstrably
213
+ mid-work — reading files, context climbing. So a status-keyed watcher can both **sleep through a wedged
214
+ worker** and **fire on a working one**, and neither failure announces itself. The bundled watcher inherits
215
+ this; so does any `herdr wait agent-status`. It is still worth arming — it catches vanished panes and real
216
+ blocks — but **never treat its silence as evidence a worker is healthy.**
217
+ Two keys that do not lie, in order of strength:
218
+ - **The daemon's own waiter.** What `herdr pane wait-output` is matching on tells you the phase from the
219
+ harness's state machine rather than from a status field: a `--match` on a readiness banner means the
220
+ worker has not launched; a `--regex` on the completion trailer means it is running. That is how a stuck
221
+ launch was distinguished from a slow one, and it beats reading the pane.
222
+ - **Pane CONTENT.** A rendered prompt pattern is the condition itself; `agent_status` is the harness's
223
+ opinion about the agent. The repo's own `trust-sweeper` scans content and caught a trust modal at
224
+ 04:40 that a status-keyed dialog watcher missed at 16:04 — same class, same day, same machine.
225
+ **The product fix for the modal case is adapter parity, not a sweeper:** the claude-code adapter passes
226
+ `--strict-mcp-config` precisely so MCP-trust modals cannot stall a worker. An adapter lacking that flag
227
+ will keep producing this stall, and a sweeper that has been running since 04:40 is evidence the gap was
228
+ visible and got swept instead of fixed.
229
+
230
+ **Every seat you spawn gets an ARTIFACT watcher armed in the SAME call that spawns it** — bundled, and
231
+ keyed on the deliverable rather than the seat:
232
+
233
+ ```bash
234
+ .claude/skills/tickmarkr-overseer/scripts/watch-artifacts.sh <MARKER> <cap-s> <poll-s> <file>...
235
+ ```
236
+
237
+ It wakes when every named file exists AND ends with its terminal marker, and on timeout it reports each
238
+ file as READY / PARTIAL / ABSENT so a quiet arm still proves the watcher was alive. Tell each seat, in its
239
+ brief, the exact marker its report must end with — you cannot watch for a marker you never demanded.
240
+
241
+ **Arm it in the same call as the spawn, not the next one.** A watcher armed "after I finish this step"
242
+ leaves a gap exactly as wide as however long you stay busy, and you will be busy — you just spawned work.
243
+ **Measured 2026-08-06 (OBS-369): two consult verdicts, 30KB and 12.8KB, sat COMPLETE with their markers
244
+ while the overseer hand-polled and reported them as still running. The operator had to ask.** Project
245
+ memory has carried this rule since before that session and the operator had already flagged it twice; it
246
+ was re-earned a third time because nothing in this skill made it a spawn-time step. It is one now.
247
+
248
+ **Never key a watcher on an agent's `done`.** `done` is TURN end, not mission end — a briefed seat flips to
249
+ `done` the moment it finishes acknowledging you, and a watcher waiting on "not working" wakes instantly and
250
+ reports a deliverable that does not exist. Key it on the artifact instead: the deliverable file existing **and
251
+ containing its terminal marker**, or `blocked`, or the pane being gone. Those are the three states that
252
+ actually require you. The same applies to a run: the journal's `run-end` event is the signal, never an
253
+ orchestrator turn boundary.
254
+
93
255
  ## Specialist pipeline rules
94
256
 
95
257
  - **Dedicated consultant tab**: Consultants (agents spawned to gather synthesis input for decisions like SCOPER analysis or architectural reviews) must run in a DEDICATED tab separate from the ORCHESTRATOR tab. When the orchestrator stands down, the consultant panes should persist so their assessments remain available for review and reference.
96
258
  - **Scoper worktree rule**: The SCOPER (or any worktree-based specialist synthesizing into the spec pipeline) must do ALL git operations in a dedicated worktree (e.g., `git worktree add /private/tmp/tkr-scoper-v155 -b spec/...`), never switching the main checkout's branch. This prevents race conditions between the specialist's branch operations and the orchestrator's shipping logic.
259
+ - **One fresh pane per consult ROUND** (operator rule, 2026-08-04): every consult round spawns a NEW pane in
260
+ the consult tab rather than re-prompting the seat that answered last round — unless there is a stated
261
+ reason to reuse. Two payoffs: each round starts with a clean context window instead of inheriting the
262
+ previous round's, which is what fills a long mission's seats and forces mid-mission `/clear`; and the
263
+ prior round's pane persists as a readable record of what that seat actually saw and said. Reuse only when
264
+ continuity of the seat's own reasoning is the point, and say so when you do.
265
+ - **CLOSE WHAT YOU SPAWNED** (operator observation, 2026-08-04: *"orch doesn't auto close the panes that he
266
+ created when no more needed"*). Panes accumulate silently — one mission reached **15 panes and 10 tabs**,
267
+ ten of them holding live agent sessions for work that had been on disk and fully consumed for hours.
268
+ Neither seat cleaned up, because neither had been told to.
269
+ - **Whoever spawns a pane owns closing it.** The orchestrator closes its workers; the overseer closes the
270
+ consultants and sweeps it spawned. The overseer sweeps whatever is left at mission end.
271
+ - **Verify the deliverable is ON DISK before closing** — a pane is the only place an unwritten finding
272
+ exists, and an agent that rendered "Done" without writing its artifact did not deliver (the
273
+ trust-disk-over-transcripts rule — cited by NAME, because a renumbered list orphans a "rule N").
274
+ **Non-empty is the FLOOR, not the test: completeness is the artifact's own TERMINAL MARKER.** One
275
+ consult verdict was 10KB on disk while its seat still read `working` — size proved it had started, and
276
+ only the closing `VERDICT:` line proved it had finished. Require the terminal marker a report is
277
+ supposed to end with, per seat, then close.
278
+ - **A pane is not an archive; the report is.** Once a worker's findings are written and synthesized, its
279
+ transcript adds nothing the report does not.
280
+ - **KEEP: the active seat, and the most recent SETTLED consult round.** That last one earns its place —
281
+ re-prompting the seat that found a defect to confirm its own fix is cheaper and stricter than briefing
282
+ a fresh one, which is the stated-reason exception above. Close consult rounds only once a later round
283
+ has re-derived their findings.
284
+ - Emptied tabs disappear on their own; do not close tabs by hand.
285
+ - **CLOSE IT AUTOMATICALLY, because remembering is what fails.** `watch-artifacts.sh` already fires on
286
+ the one signal that means a seat is finished — the artifact plus its terminal marker — so hand it the
287
+ panes too: `TKR_CLOSE_PANES="w1:p1,w1:p2" watch-artifacts.sh …`. It closes them on completion and
288
+ **never on timeout**, where the seats are still working. Closing on the marker cannot reap a seat
289
+ mid-write, which is exactly why `done` would be the wrong trigger.
290
+
291
+ - **MEASURE BEFORE EVERY SPLIT, AND JOIN *DOWN* WHEN A RIGHT-SPLIT WOULD GO UNDER THE FLOOR.**
292
+ **Operator-observed 2026-08-06, with a screenshot:** five consult panes in one tab rendered **14 columns
293
+ wide each** out of 220 — every one unreadable, including the two that had finished hours earlier.
294
+ **This rule's own earlier wording said to "split the newest pane; the tree stays balanced", and that
295
+ remedy is wrong** — an OVERSEER followed it the next session and measured `110/55/55`, which is the
296
+ exact split the old text cited as the *failure*. Corrected, with the measurement:
297
+ - **Direction is decided by arithmetic, not by which pane you pick.** The driver's floor is real and
298
+ derived from measurement — `TRAILER_SAFE_FLOOR_COLS = 108` (`src/drivers/herdr.ts:13`, *"narrowest
299
+ safe 53 → floor 108"*), and it splits right only while `paneWidth/2 ≥ 108 + 2` (`herdr.ts:494`),
300
+ otherwise **down**. Apply the same test by hand: `herdr pane layout --pane <id>`, halve the width,
301
+ and if the halves fall under the floor, split `--direction down`.
302
+ - **Binary splits cannot produce an even 3-column row at any width.** 220 goes to 110/55/55 whichever
303
+ pane you split. **At a 220-col terminal the width-derived cap is TWO side-by-side panes**; a third
304
+ seat goes below one of them, or into its own tab. "Three panes" is a *height* heuristic
305
+ (tickmarkr's own `workersPerTab: 3` assumes ~50 rows) and it does not authorise a third column.
306
+ - **A finished seat keeps its width.** Panes are a fixed budget — every seat you do not close is taken
307
+ out of the readability of the ones still working. There is no rebalance command, so the fix is
308
+ closing, not resizing.
309
+ **The general lesson, which is why this correction is worth its lines: a prose rule that restates a
310
+ measurement without carrying the number reproduces the defect at full price.** The floor lives in
311
+ `src/`; every seat that hand-splits panes is outside it and re-learns this by hand.
97
312
 
98
313
  ## Non-negotiable rules
99
314
 
100
- 1. **Takeover rule**: only act on a worker if it needs input AND the orchestrator is not `working`.
101
- 2. **Human checkpoints (absolute)**: any gate marked `autonomous: false` or asking for product/visual
102
- sign-off is NEVER auto-answered — regardless of how obviously correct the highlighted option looks. Leave
103
- it blocked and bring the user the decision WITH evidence. If the mission explicitly delegates authority,
104
- routine-class gates may be overseer-decided after polling the operator first — but spend and ship gates
105
- NEVER self-decide.
106
- 3. **Trust disk over transcripts**: verify artifacts on disk before building on them; a subagent killed
315
+ 1. **Do not drive the run.** The ORCHESTRATOR owns `compile`/`plan`/`run`/`resume`, the watchers, the
316
+ sweeps, and assembling gate evidence. You rule, record, and talk to the operator. Measured cost of
317
+ ignoring this: one OVERSEER at 747K tokens beside an idle ORCHESTRATOR at 369K, doing one tier's work
318
+ in the seat that cannot cheaply `/clear`. **The tell is your own hands** — typing `tickmarkr resume`,
319
+ tailing a journal to decide the next move, sweeping orphans. Hand it back.
320
+ 2. **Takeover rule**: only act on a worker if it needs input AND the orchestrator is not `working`.
321
+ 3. **Human checkpoints**: any gate marked `autonomous: false` or asking for product/visual sign-off is
322
+ NEVER auto-answered — regardless of how obviously correct the highlighted option looks. Leave it
323
+ blocked and bring the user the decision WITH evidence.
324
+ **When the mission delegates authority, the carve-out is the IRREVERSIBLE CREDENTIAL-BEARING ACT
325
+ ITSELF — not every decision upstream of it.** `npm publish`, `git tag`, a push to a public remote run
326
+ under the operator's name and account, and *"in charge" is not an npm token*. **Everything upstream is
327
+ yours**: whether a fix warrants a patch release, what rides which milestone, what to spend, what order
328
+ to ship in. Announce it with its cost basis and ACT.
329
+ **Earned 2026-08-06, at a measured price.** An OVERSEER holding a written delegation asked the operator
330
+ *"patch release ahead of the milestone, or fold it in?"* — a SCHEDULING question, no credential
331
+ anywhere near it, about a fix that did not yet exist. The operator was away **8.5 hours**. The run
332
+ ended, parked seven tasks and went unwatched; the context watcher fired and exited unread. Operator,
333
+ verbatim: *"why did you wait for me to decide earlier? you should have taken the decision your self ..
334
+ I delegated this to you, remember?"*
335
+ **The tell: if no credential, tag or public remote is touched by the ACTION you are about to take, it
336
+ is not the carve-out — decide it.** And a blocking ask is never the only option: route it through a
337
+ cross-vendor consult and rule, which is what a pre-committed release criterion already prescribes.
338
+ **A supervising seat that blocks is not neutral — it is unwatched.** Waiting has a running cost the
339
+ question never displays, and that cost lands on the run, not on the seat that waited.
340
+ 4. **Trust disk over transcripts**: verify artifacts on disk before building on them; a subagent killed
107
341
  mid-flight still renders "Done" without writing its artifact.
108
- 4. **Report concisely on every state change**: what happened, who handled it, what's next. Lead with the
109
- outcome. Surface product decisions; never make them.
110
- 5. **Log every abnormality** to `.planning/OBSERVATIONS.md` (or the project's ledger), even mid-run.
342
+ 5. **Report concisely on every state change**: what happened, who handled it, what's next. Lead with the
343
+ outcome. **Surface product decisions — and under a standing delegation, MAKE them and say you did.**
344
+ Without a delegation, surface and wait. With one, deciding IS the job; report the ruling and its basis
345
+ rather than the question. Say *"I decided"*, never *"you approved"* — a record implying a signature it
346
+ never received is this rule's own defect class running in the opposite direction.
347
+ 6. **Log every abnormality** to `.planning/OBSERVATIONS.md` (or the project's ledger), even mid-run.
348
+ 7. **Every fix is evaluated for shipping.** The tarball is `files: [dist, schema, skills, fixtures]` — so
349
+ `src/**` and `skills/**` reach users while `.overseer/**` and `.tickmarkr/**` reach nobody. Before
350
+ calling a fix done, ask where it lands: a local overlay or a scaffold script standing in for a source
351
+ fix helps ONE operator and leaves every other user with the defect. If an overlay is the interim, it
352
+ says so in writing and names its removal condition.
353
+
354
+ ---
355
+
356
+ ## Briefing a seat to audit a security-shaped check — phrasing matters
357
+
358
+ **Earned 2026-08-04.** A consult seat was asked to *"hunt one more forged pass"* on an authorization gate.
359
+ Its provider cut the session off mid-work — *"We take extra caution with cybersecurity requests"* — and the
360
+ report was never written. The seat had already found the real defect (a timezone-dependent clock) and that
361
+ finding was recovered only by reading its pane before closing it.
362
+
363
+ **The work is legitimate; the framing is what trips the filter.** Ask for completeness, not exploitation:
364
+
365
+ - ✗ "find a forged pass", "bypass this", "attack the gate", "how would you defeat it"
366
+ - ✓ **"enumerate every input this check depends on, and confirm each one is bound"**
367
+ - ✓ "which of these inputs can a caller still control?"
368
+ - ✓ "state what this check does NOT establish"
369
+
370
+ That phrasing produces the same findings — the timezone hole IS an unbound input — without asking a model to
371
+ generate an attack. **And read the pane before closing a seat that ended without its artifact:** a refusal
372
+ mid-work leaves real findings in the transcript and nowhere else, which is the one case where the pane, not
373
+ the report, is the deliverable.
374
+
375
+ ## Evidence discipline — the durable core
376
+
377
+ Distilled from a v1.86 spec-repair mission that produced 31 numbered rules, ~90 audit findings and nine
378
+ errors authored by the supervising seat itself. **Every line below was earned by a defect, most of them
379
+ twice.** They are mission-independent on purpose: nothing here names a task, a line number or a figure.
380
+
381
+ **Rot**
382
+
383
+ 1. **A quotation is exact bytes.** `grep` it before attributing it; if it does not hit, it is not a
384
+ quotation. A fabricated quotation is the only error that presents itself as primary evidence, so the
385
+ natural check is already answered on its face.
386
+ 2. **A verified quotation ROTS** — the source moves underneath it. Pin it (`as written at <sha/time>`) or
387
+ re-verify. Documenting a repair is the highest-risk case: the edit you describe is the edit that
388
+ falsifies your description.
389
+ 3. **A FINDING rots exactly like a quotation, and carries more authority while doing it** — a quotation
390
+ invites checking; *"the consultant found X"* invites action. Re-derive the premise before acting. A
391
+ *dissolved* finding gets marked SUPERSEDED, never silently dropped.
392
+ 4. **Before editing a line, sweep for the records that QUOTE it** — rule 2 used prospectively, which is the
393
+ only time it is cheap. **And the dual, from the mutating end: any repair to a CONDITION a document
394
+ DESCRIBES must sweep the descriptions in the same edit.** Fixing the world falsifies the prose about the
395
+ world, and that prose is nobody's assigned target — it is collateral. Four occurrences in one phase; twice
396
+ the fix was right and only the record was wrong, which is the version that survives review because the
397
+ change itself is defensible. **Keep the defect's record when you resolve it** — a passage that flagged a
398
+ risk which was then relied upon anyway is more instructive than a clean line saying "resolved".
399
+ 5. **Never freeze a moving number.** A figure describing anything under active edit is a quotation on a
400
+ timer; record the derivation command, not the value. A count over a population your own work adds to is
401
+ self-invalidating — state a floor.
402
+
403
+ **Scope of a result**
404
+
405
+ 6. **Every gate, tool and verdict states what it does NOT establish.** A green gate is a claim about form
406
+ until its negative scope says otherwise. This applies to a *seat's own verdict* as much as to a tool:
407
+ an unchecked cite in a task with no finding is unchecked, not confirmed.
408
+ **THE PRESENCE OF A ROW IS NOT EVIDENCE THAT THE WORK HAPPENED — READ ITS QUALIFYING FIELDS.** Three
409
+ instances in one run (2026-08-06), which is what makes it a law and not an anecdote: a `gate-result`
410
+ for `test` carrying `selectedTests` — a PASS over a 16-test subset, not the suite; a `phase-start` for
411
+ a gate with no result row at all, where *deferred* and *dropped* are indistinguishable; and a
412
+ `tip-verify` row with `cached: true`, whose own source comment says it *"keeps it honest about not
413
+ having re-run the command."* **In the first and third the product had already provided the qualifier
414
+ and the reader ignored it** — an OVERSEER read per-gate `tip-verify` rows as proof of a real verify
415
+ while the distinguishing field sat in its own tool output, and was corrected by the ORCHESTRATOR from
416
+ the same lines. So decompose the blame honestly, because the two halves ship to different places:
417
+ **rows that are never emitted are a PRODUCT defect; rows misread past their qualifiers are a READER
418
+ defect**, and no amount of product work fixes the second. Before quoting any row as evidence of an
419
+ action, ask what field on it would tell you the action was skipped, cached, subsetted or deferred —
420
+ and if you cannot name the field, you have not read the record, you have counted it.
421
+ 7. **Never aggregate per-axis PASSes into "it is clean."** Carrying the PASS and dropping the scope
422
+ manufactures a clean bill nobody issued.
423
+ 8. **Never exclude a path from a search whose purpose is to find a counterexample there** — and an
424
+ INCLUSION list excludes just as effectively, while being harder to see because every entry is
425
+ individually justified. Cite the line you actually verified.
426
+ 9. **A name-keyed sweep answers "is the name absent", not "is the concept absent."** Sweep the mechanism's
427
+ vocabulary, and one level further: sweep for what the mechanism *does to* its consumers, not who calls
428
+ it — a thing the harness *applies* to consumers is named by them in no vocabulary at all. Expect
429
+ over-return; discriminating hits is the cost of the method.
430
+ 10. **A hit proves BYTES, not attribution** — and N hits can be ONE origin copied N times.
431
+
432
+ **Instruments**
433
+
434
+ 11. **For any guard whose failure is SILENCE — detector, lint, watcher, gate, alarm branch — the acceptance
435
+ test is a POSITIVE CONTROL, not a clean run.** A zero cannot distinguish *nothing is broken* from *the
436
+ instrument is blind* from *the check does not exist*. Remove the condition it should catch and confirm
437
+ it FIRES; only then trust its quiet. **A comment asserting the check is enough to make its own author
438
+ believe it ran.**
439
+ **This rule is the oldest one here and the most re-earned.** Project memory has carried it since
440
+ 2026-07-14 as the *falsification drill* — *"eleven gate/pin defects in v1.7 alone, every one caught by a
441
+ drill rather than by a passing test"*, including a `grep -c` gate that exits 0 whether tests pass or
442
+ fail. Prefer a **compile-time guarantee** (a required parameter → a type error) over a grep-pin whenever
443
+ the choice exists: a grep-pin guarding a silent default is a hope. Read step 0a — this is what happens
444
+ when nobody opens the memory.
445
+ **Turn this rule on your own WATCHERS, because they are the guard you are least likely to aim it at.**
446
+ A journal watcher armed on `run-end`/`task-human`/`task-failed` is *supposed* to stay silent through a
447
+ clean merge — so a dead watcher and a correctly-quiet one emit byte-identical evidence from inside the
448
+ seat that owns it, and they diverge only at the first event the wake was actually for. **Watcher
449
+ liveness is proved by the PROCESS TABLE, never by its silence, and never by the report of the seat that
450
+ owns it** — "watchers alive" is the one claim a seat cannot verify about itself. Measured 2026-08-06:
451
+ an orchestrator sat `idle` through three merges and two dispatches with no journal watcher in the
452
+ process table, while its own last report read *"daemon, board, sweeper, watcher all alive"* (OBS-366).
453
+ Two corollaries: **re-arm a wake-and-exit watcher as the same turn's LAST act**, not the next turn's
454
+ first — the gap between them is unwatched and its width is however long the seat stays busy; and **a
455
+ handoff that re-arms one tier's watchers must say which tier's it did NOT re-arm.**
456
+ **That first corollary prescribes DISCIPLINE, and discipline is the wrong fix — measured 2026-08-06.**
457
+ One orchestrator lapsed its journal tier **31 minutes**, then, after diagnosing it and fully intending
458
+ to re-arm, lapsed it again for 3 minutes **while actively thinking about watchers**. Its own diagnosis
459
+ is the durable one: *"I still serialize re-arming behind whatever I am doing."* **A watcher whose
460
+ liveness depends on its owner being free is not armed, it is SCHEDULED.** The structural fix, which
461
+ then survived a wake with zero action from the seat: wrap every wake-and-exit watcher in a supervisor
462
+ that re-execs it, **detached (`ppid 1`) so it outlives the seat and not merely the seat's turn**, and
463
+ have it write a **heartbeat file** so the supervising tier proves liveness *from disk* instead of
464
+ asking the seat that owns it. Decouple **coverage** from **notification**: when the notifier later
465
+ broke, coverage held and nothing was lost — the failure the design was built for.
466
+ **And never convert instrument silence into a WORLD claim.** *"No watcher has fired since X"* is a
467
+ statement about your instrument; *"no state change"* is a statement about the run, and they have
468
+ different truth conditions. A terminal-event watcher is silent through every **non-terminal** change
469
+ **by design**, so its silence is evidence about a narrow event class and **never** about progress.
470
+ Measured the same day: an orchestrator reported *"no state change"* while five events, a completed
471
+ worker and a passing gate sat unread — it had asserted from memory one read-cycle behind a reading
472
+ that was about to arrive. Say the instrument sentence, or **re-read and then say the world one**.
473
+ **A watcher has TWO failure modes, and the second is invisible from inside: never armed, and
474
+ OUTLIVING ITS TRIGGER.** A watcher aimed at an event that can no longer occur **reads as coverage and
475
+ is worse than none** — the process table shows it alive and the seat that armed it remembers arming
476
+ it. When a decision cancels the event a watcher waits on, stand it down **in the same act**. (Earned
477
+ 2026-08-06: an orchestrator did exactly this, unprompted, the moment a ruling cancelled the recompile
478
+ its standby watcher was waiting for.)
479
+ **And ASK the negative, explicitly — it is the question that produces gaps.** *"Which tier's watchers
480
+ did you NOT re-arm?"* A handoff reporting what IS armed produces a list; a handoff required to name
481
+ what is not armed produced, in one answer: a run's live surface that had **died mid-run** and was
482
+ found only in the post-run audit, and a worker-liveness tier that had **never been armed**, leaving
483
+ the daemon both the supervised thing and the sole watcher of its own workers.
484
+ 12. **Check which QUANTIFIER the claim uses before quoting a derivation for it.** "The path is N" needs a
485
+ maximum; *"both chains"*, *"the only consumer"*, *"exactly one owner"* need an ENUMERATION — and a
486
+ max-with-tie-break silently answers the first question when you asked the second.
487
+ 13. **Derive mechanically; hand-derived sets are wrong.** Prefer the real parser's output over a
488
+ re-implementation, and a property over an enumeration. State what the mechanism cannot establish.
489
+ **And make every instrument PRINT WHAT IT ACTED ON.** A tool that does not name its target cannot be
490
+ caught answering about the wrong thing: a dry-compile helper that silently ignored its path argument
491
+ returned the same verdict for two different files, and a seat read one answer as two results and
492
+ concluded both forms were valid. The tell is unavailable unless the tool volunteers it. Corollary —
493
+ an instrument that takes an input must be handed a DELIBERATELY BAD one before its clean runs are
494
+ worth anything (rule 11 applied to tools, not just to gates).
495
+ 14. **A unit is not a measurement.** A configured timeout is a KILL CEILING, not a duration — never compare
496
+ it to a wall clock or quote it to an operator as an estimate.
497
+ 15. **Verify through the path that LOADS, not the path you edited.** Mirrored trees and symlinks mean your
498
+ check can confirm a shadow copy; `sed -i` on a tracked symlink silently replaces it with a regular file.
499
+ 16. **Never edit a script with a live instance** — bash reads by byte offset, so even a comment-only
500
+ insertion corrupts the running process. Cancel, edit, syntax-check, re-arm, in that order.
501
+
502
+ **Propagation**
503
+
504
+ 17. **A confirmed single-site or single-axis miss is a CLASS, not an instance.** Re-run the same sweep shape
505
+ on every sibling; ask what other dimension the fixtures hold constant. **A ruling that fixes one
506
+ instance of a class it just defined is incomplete by construction** — the sweep is part of the ruling.
507
+ 18. **When a boundary moves, every clause referencing the old boundary moves with it.** Neither clause ever
508
+ looks wrong alone, so single-clause review cannot catch this class. Disambiguate any term used at two
509
+ levels.
510
+ 19. **A METHOD GUARD found by one seat must be promoted to where every seat reads it, in the same pass that
511
+ reads it.** Otherwise it is re-earned at full price — and the second earning is worse, because by then
512
+ the wrong answer carries a citation.
513
+
514
+ **Authority**
515
+
516
+ 20. **Open the file the instruction is about, even when the instruction comes from above.** A ruling reads
517
+ as settled, and that is exactly when it goes unchecked. Overseer rulings are wrong at roughly the rate
518
+ of everyone else's.
519
+ 21. **State the verification standard alongside the instruction**, or the defect appears at the seam.
520
+ 22. **An overclaimed self-criticism is the least-audited sentence you will write** — a harsh line invites no
521
+ check, so it ships unverified. Including in a section like this one.
522
+ 23. **A pre-commitment needs a TRIGGER and a SUBJECT SET. Naming only the trigger is a live hazard.**
523
+ A bound was rewritten mid-mission to make its CONDITION mechanically checkable — and that rewrite was
524
+ already the product of one near-miss. It was **still** incomplete: a third task later parked with
525
+ **both halves of the condition present**, and the bound did not apply, because that task was not in
526
+ its subject set. A seat reading only the trigger would have closed the milestone on a task the
527
+ pre-commitment was never about. **State both: what fires it, and what it is ABOUT.** A correct trigger
528
+ with an unstated subject executes on the first thing matching its shape, carrying the authority of the
529
+ decision it was written for.
530
+ 24. **A REMEDIATION is believed where a guard would be drilled.** Rule 11 says a guard whose failure is
531
+ silence needs a positive control. **Nobody applies that to a FIX**, because a fix is not an
532
+ instrument — so a shipped remediation is remembered as coverage and never re-read. One was recalled as
533
+ *"the reaper shipped in v1.78"*; its own changelog said it reaps only what a helper **tracks**,
534
+ *"without a broad migration"*, and measurement found the untracked majority behaving exactly as it had
535
+ been left. **Ask of any remembered fix: what did it actually cover, in its own words, at the time?**
536
+ 25. **The fabrication lives in the INDEX line, not the body — and the index is what everyone loads.**
537
+ In that same case the memory body said *"expect regrowth until the product fix ships."* The one-line
538
+ summary said *"reaper shipped."* **The accurate body was never opened, because the index had already
539
+ answered the question.** Audit index and summary lines against the bodies they point at; a compression
540
+ that drops a qualifier is indistinguishable from a fact.