tickmarkr 2.1.7 → 2.1.9

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -41,8 +41,11 @@ through brief lineage. **An executor choice nobody made is still an executor cho
41
41
  is strictly WEAKER than the live seat's report rule 11 already forbids trusting — and it reads as
42
42
  settled fact. So at every adopt, walk the predecessor's watchers by class — **journal watchers,
43
43
  artifact watchers, dialog watchers and beat loops, which is the closed set a session owns** — probe
44
- each from the process table yourself (`pgrep -f <token>`, discriminated per rule 11), and re-arm every
45
- one the table does not show. Earned 2026-08-25 (OBS-622): a handoff recorded *"artifact watcher armed"*
44
+ each from the process table yourself **twice, a second apart, keeping only what appears in both**
45
+ (rule 11's liveness test — a loop of single `pgrep -f <token>` snapshots is what returned eight
46
+ phantom pids at an adopt on 2026-08-31, every one of them the probing shell itself), and re-arm every
47
+ one the table does not show. A zero here is not absence until the same probe has been run once against
48
+ a watcher you know is alive. Earned 2026-08-25 (OBS-622): a handoff recorded *"artifact watcher armed"*
46
49
  over two live consult verdicts; at adopt the only `watch-artifacts.sh` on the machine belonged to a
47
50
  different repository, and nothing had been watching either file.
48
51
  **An adopted seat ANNOUNCES itself, in the same act as re-arming:** tell the adopted orchestrator the
@@ -165,6 +168,10 @@ journal tail to decide what happens next, or sweeping orphans — you have taken
165
168
  - **The journal is the source of truth**, not panes. Watchers go on `run-end` / `task-human` /
166
169
  `task-failed` / `consult-verdict`; never sleep-poll inside an agent turn. **Never key a watcher on an
167
170
  agent's `done`** — that is turn end and fires the moment a seat finishes acknowledging you.
171
+ **All four are covered by one shipped instrument** — `scripts/watch-journal.sh <runs-dir> [poll] [cap]
172
+ [events-csv]` — which arms on a line baseline, wakes once, and grades a `run-end` against every green
173
+ clause. `scripts/watch-parks.sh` stays the park-specific wake for THIS seat (it counts parks and speaks
174
+ about rulings); the two overlap on `task-human` deliberately, and arming both is coverage, not a bug.
168
175
  - **Daemon liveness ≠ journal activity.** A dead daemon emits no events, so journal watchers sleep through
169
176
  its death. Liveness comes from the lock's OWN pid (`kill -0`), never a command-name grep. Recovery is
170
177
  `tickmarkr resume <runId>` — **the orchestrator's command, not yours** — and note that resume REPLAYS the
@@ -199,6 +206,58 @@ Read the evidence file the orchestrator writes, rule on it against your pre-comm
199
206
  record the ruling with what it set aside, and hand the ruling back for execution. That is the whole job,
200
207
  and it is the only work that cannot be delegated — which is exactly why nothing else should occupy you.
201
208
 
209
+ #### The release criterion, and the two clauses without which it grades work it never saw
210
+
211
+ **A pre-commitment is a file. A re-scope is a different file. Nothing joins them** — so a criterion
212
+ survives its own subject's rewrite and keeps grading, which is **worse than having no criterion at all**:
213
+ it carries the authority of a seal over work the seal never saw. Both clauses below, or the join stays
214
+ broken in the direction the missing half covers.
215
+
216
+ **Measured 2026-08-30 (OBS-804).** `RELEASE-CRITERION-v2.1.8.md` was sealed at 10:18:25, nine minutes
217
+ before its run, with zero gate results in existence — an honest pre-commitment that named its trigger AND
218
+ enumerated its subject set, exactly as rule 23 demands. **Two of its four named subjects were then
219
+ materially edited underneath it**: T3 narrowed at 11:49:58, T1 rewritten at 13:37:59 — *three hours and
220
+ nineteen minutes after sealing*. The run that actually delivered the work was, in the criterion's own
221
+ vocabulary, both *"a re-plan"* and *"a follow-on run"* — **two of its own named exclusions. The document
222
+ excluded the only run that ever ran it.** Three of five clauses had already graded MET before anyone asked
223
+ the subject question, and it was caught because a downstream seat refused to decide a ruling that was not
224
+ its own — **not because any instrument detected it.**
225
+
226
+ **Every sealed criterion carries a VOID CONDITIONS section. It is not optional and it is not boilerplate:**
227
+
228
+ ```markdown
229
+ ## SUBJECT
230
+ <what this criterion is ABOUT — see the identity rule below>
231
+
232
+ ## TRIGGER
233
+ <what fires the grading>
234
+
235
+ ## CLAUSES
236
+ 1. …
237
+
238
+ ## VOID CONDITIONS — this document is VOID, with no ruling required, if any of these occur
239
+ - any subject named above is re-scoped, narrowed, widened, split, merged or re-owned
240
+ - the work is delivered by a run this document's own exclusions would exclude
241
+ - <the specific things that would make these clauses grade something else>
242
+ ```
243
+
244
+ > **Naming a subject does not freeze it. A pre-commitment must name what VOIDS it, not only what it
245
+ > covers — and a re-scope of any named subject voids it AUTOMATICALLY, with no ruling required.**
246
+ > A void condition that needs a ruling to fire is not a void condition; it is a second thing to forget.
247
+
248
+ **⚡ IDENTIFY THE SUBJECT BY WHAT THE CLAIM IS ABOUT. Half of the failure above was a category error, and
249
+ it is the cheap half to fix:** a graph hash identifies a **PLAN**, and a plan is recompiled, re-cut and
250
+ re-owned as a matter of course. **A criterion about a SHIPPED TREE names the COMMIT** — or the tag, or the
251
+ export tree hash — **never a graph hash, never a run id, never a task list.** Ask what a reader would have
252
+ to hold in their hand to check the clause: if it is bytes, name the bytes.
253
+
254
+ **THE RECIPROCAL DUTY, and it is yours because you write both documents:** when you issue a ruling that
255
+ re-scopes, narrows, splits or re-owns anything, **the ruling must name every sealed document its subject
256
+ appears in** — and say, in the ruling, whether each one is now void. You are the only seat that can do
257
+ this: the criterion cannot watch for the ruling, and the product cannot know that an English document
258
+ elsewhere sealed a claim about a graph it is recompiling. **A re-scope ruling that names no sealed
259
+ documents is asserting there are none. Check before you assert it.**
260
+
202
261
  #### The one operational duty that IS yours: a verdict produced under starvation is not a verdict
203
262
 
204
263
  **Operator, 2026-08-07: *"that is the kind of job I need overseer to be vigilant about."*** Do not read the
@@ -536,6 +595,53 @@ they are left implicit:
536
595
 
537
596
  ## Supervision watcher
538
597
 
598
+ ### ⛔ EDITING A WATCHER WHILE WATCHERS ARE ARMED: REPLACE BY RENAME, NEVER IN PLACE
599
+
600
+ **`bash` reads a running script BY BYTE OFFSET.** Edit the file a live watcher is executing and every
601
+ offset after your edit shifts — a comment-only insertion is enough — and the process runs garbage from
602
+ wherever it happens to be. **The failure signature is SILENCE: a corrupted watcher and a correctly-quiet
603
+ one emit byte-identical evidence.**
604
+
605
+ > **Write a temp file, then `mv` it over the target.** `mv` swaps the directory entry and yields a NEW
606
+ > inode; the running `bash` keeps its old inode open and finishes unharmed on the old bytes. **An in-place
607
+ > `sed -i`, or a `>` truncate-and-rewrite, corrupts a running script mid-flight.** Re-arm afterwards to
608
+ > pick up the new bytes — the running process will not.
609
+
610
+ **And `skills/` (canonical) versus `.claude/skills/` (installed) is a MIXED tree — symlinks for some
611
+ files, independent copies for others — so a blanket rule in EITHER direction is wrong** (OBS-809). It
612
+ cuts both ways: for a **shared inode**, an edit in `skills/` reaches into the running process; for a
613
+ **copy**, a fix in `skills/` does **not** reach the running watcher at all, so a repaired watcher keeps
614
+ running the old bytes while the tree says it is fixed. Sync both trees or `skills-single-source.test.ts`
615
+ reds.
616
+
617
+ ⚠ **PROBE IT WITH `readlink` AND `stat -L`. NEVER BARE `stat -f %i`: on macOS that reports the SYMLINK'S
618
+ OWN inode, not its target's**, so every symlink reads as a separate file. Measured 2026-08-31 — one seat
619
+ probed **one** file that way, got differing inodes, and wrote *"they are all copies, editing them cannot
620
+ corrupt the running watchers"* into a seat brief as a blanket rule; following it would have silently
621
+ killed the watcher then handing a milestone's spec to its orchestrator. A second seat's first pass
622
+ returned **seven** false *"separate file"* verdicts before it re-probed.
623
+
624
+ ```bash
625
+ # Identity, correctly. It must compare the TWO TREES: pointed at either one alone it renders nothing,
626
+ # because the canonical files are all real files and the links live on the installed side.
627
+ cd <repo>
628
+ for f in $(cd skills && find . -type f -o -type l | sed 's|^\./||' | sort); do
629
+ a="skills/$f"; b=".claude/skills/$f"
630
+ [ -e "$b" ] || { printf '%-52s CANON-ONLY\n' "$f"; continue; }
631
+ ia=$(stat -L -f %i "$a"); ib=$(stat -L -f %i "$b") # -L: resolve, or symlinks read as separate files
632
+ [ "$ia" = "$ib" ] \
633
+ && printf '%-52s SHARED INODE %s -> mv-replace ONLY; a live reader is on these bytes\n' "$f" "$ia" \
634
+ || printf '%-52s copies (%s/%s) -> a fix in skills/ does NOT reach the running copy\n' "$f" "$ia" "$ib"
635
+ done
636
+ ```
637
+
638
+ **Two failures worth separating, and the second is the durable one:** *a probe that cannot render the
639
+ evidence cannot fail* — rule 11 aimed at your own instrument; and **a narrow verified fact was generalised
640
+ to a population it was never sampled over.** One file checked, seven ruled on. **The generalisation ran in
641
+ the REASSURING direction, which is worse than the alarming one: an alarming overclaim gets challenged, a
642
+ reassuring one gets acted on.**
643
+
644
+
539
645
  **Arm your OWN tier first, in the same call chain that arms everything else.** `status` derives each
540
646
  tier's state from a beat file the tier itself writes, so a seat that never beats reads `ABSENT` — and
541
647
  `ABSENT` means *never armed*, which is a lie about a seat that is working the run. Measured on the P99
@@ -575,7 +681,8 @@ one owned by an unrelated session. So:
575
681
  beat <tier> --seat <seat>` in this repo). Neither liveness claim is read from a recorded pid: a pid
576
682
  recorded earlier can be stale, reused, or detached from the beat now holding the tier green.
577
683
  - **At every adopt, clear, or re-brief, sweep for pre-existing loops on YOUR tier before arming one**
578
- (`pgrep -f "tickmarkr beat <tier>"`), trace each to its parent session, and kill the **loop only**
684
+ (`pgrep -f "tickmarkr beat <tier>"`, **read twice and intersected** — this exact probe returned its own
685
+ shell as pid 14680 on 2026-08-31), trace each survivor to its parent session, and kill the **loop only**
579
686
  — never the parent — then verify the parent survived.
580
687
  - **`ARMED (<seat>)` is an attributable claim, not proof that the named seat is still alive.** Before
581
688
  trusting it, ask whose session owns the beater; an orphan loop can keep naming a departed seat
@@ -996,27 +1103,56 @@ twice.** They are mission-independent on purpose: nothing here names a task, a l
996
1103
  the same root read the other way: a DETACHED loop outlives its seat and holds a tier `ARMED` with
997
1104
  nobody home (OBS-583). Neither direction may be assumed; the lifetime is a property of how the watcher
998
1105
  was launched, and it belongs in writing next to every claim that one is armed.
999
- **And the process-table probe has a standard idiom that DEFEATS it, so the rule above needs one more
1000
- line to be usable.** Never probe for a watcher with `ps … | grep <token> | grep -v grep`: a poll-grep
1001
- watcher carries the word `grep` in its own argv, so the filter whose job is removing the *probing* grep
1002
- removes the *watched* one. Measured 2026-08-06 against a positive control (OBS-415):
1003
- `ps -eo pid,ppid,etime,command | grep -F <token>` returned **4 matches**, and adding `| grep -v grep`
1004
- returned **0**. The seat concluded its watcher had died silently, reported that to the operator, filed
1005
- it as a defect — and was corrected forty minutes later when the watcher fired normally, having been
1006
- alive throughout. Two hypotheses (`ps` truncation; multi-column truncation) were formed and killed by
1007
- measurement first, and the first falsification was itself run against the wrong `ps` form. **Use
1008
- `pgrep -f <token>`, or read the lock's own pid.**
1009
- ⚠ **AND `pgrep -f` HAS ITS OWN INVERSE FAILURE, so the recommended fix is not free: it matches the
1010
- ARGV OF THE SHELL RUNNING IT.** A probe written as `pgrep -f "npm test"` is itself a process whose
1011
- command line contains `npm test`, so it returns its own shell — a PHANTOM that looks exactly like the
1012
- contamination you are hunting. Measured 2026-08-25: a load-ceiling wake was investigated, a second
1013
- `npm test` "outside the gate's worktree" was found, and it was the probe. It had vanished by the next
1014
- command, which is the tell — a real second suite does not exit between two reads. The escalation would
1015
- have been a false contamination alarm during a task's last attempt.
1016
- Discriminate before you believe a hit: **resolve each pid's `cwd` AND drop any whose own command
1017
- contains the probe** (`pgrep`/`bash -c`), or match a pattern the target has and the probe cannot —
1018
- the binary's real path rather than the words you typed. `grep -v grep` fails toward *not there*;
1019
- `pgrep -f` fails toward *there twice*, and this direction gets ACTED ON, which is worse.
1106
+ **And EVERY process-table probe has an idiom that defeats it, so the rule above needs the one test
1107
+ that survives all of them. Lead with this; it is not the last resort, it is the first move:**
1108
+
1109
+ > ### A REAL WATCHER DOES NOT EXIT BETWEEN TWO READS.
1110
+ > Read the process table twice, a second or two apart, and keep only what appears in both —
1111
+ > `kill -0 <pid>` on each candidate is the cheap form. Nothing else is needed to kill a phantom.
1112
+
1113
+ It works because it tests a **property of the thing you are hunting** (a watcher persists) rather than
1114
+ a **property of your probe** (how its argv happens to look). A probe-shaped exclusion has to enumerate
1115
+ every way a probe can look; the liveness test does not care what the probe looked like — which is why
1116
+ it is immune to both failures below instead of to one of them.
1117
+ **Measured 2026-08-31 (OBS-807), twice inside one hour, by two seats independently on one machine.**
1118
+ One seat swept for inherited watchers with a loop of `pgrep -f "<script>.sh"` and got **eight
1119
+ live-looking pids**; the other probed `pgrep -f "tickmarkr beat orchestrator"` and got **pid 14680**.
1120
+ **All nine were the probing shell's own argv, and one `kill -0` sweep one command later killed all
1121
+ nine at once** — no cwd resolution, no argv parsing, no per-hit judgement. Two supporting tells, both
1122
+ free: phantom pids arrive **sequential** (6781, 6786, 6791 … one per iteration of the probing loop),
1123
+ and a phantom's `argv` reads **empty** by the time you inspect it.
1124
+ ⛔ **Both seats were following this skill's own previous text correctly when they produced a phantom.**
1125
+ That text named `pgrep -f` as the *remedy* for `grep -v grep`. A remedy with that recurrence rate is
1126
+ not a remedy; it is a second trap wearing the first one's clothes.
1127
+
1128
+ **The two idioms, and the direction each one lies in — you still need to know these, because the
1129
+ liveness test tells you a hit is REAL, not that it is YOURS:**
1130
+ - `ps … | grep <token> | grep -v grep` fails toward ***not there***. A poll-grep watcher carries the
1131
+ word `grep` in its own argv, so the filter whose job is removing the *probing* grep removes the
1132
+ *watched* one. Measured 2026-08-06 against a positive control (OBS-415):
1133
+ `ps -eo pid,ppid,etime,command | grep -F <token>` returned **4 matches**, and adding `| grep -v grep`
1134
+ returned **0**. The seat concluded its watcher had died silently, reported that to the operator, and
1135
+ filed it as a defect — and was corrected forty minutes later when the watcher fired normally, having
1136
+ been alive throughout. Two hypotheses (`ps` truncation; multi-column truncation) were formed and
1137
+ killed by measurement first, and the first falsification was itself run against the wrong `ps` form.
1138
+ - `pgrep -f <token>` fails toward ***there twice***, because it matches the **argv of the shell running
1139
+ it**. Measured 2026-08-25: a load-ceiling wake was investigated, a second `npm test` "outside the
1140
+ gate's worktree" was found, and it was the probe; the escalation would have been a false
1141
+ contamination alarm during a task's last attempt. **This is the worse direction, because *there
1142
+ twice* gets ACTED ON** — reported to the operator as a defect, or swept.
1143
+
1144
+ **FALLBACK, for a hit that SURVIVES two reads and still might be yours:** resolve each surviving pid's
1145
+ own `cwd` (`lsof -a -p <pid> -d cwd`) and count only what belongs to the tree you are asking about, or
1146
+ match a pattern the target has and the probe cannot — the binary's real path rather than the words you
1147
+ typed. This costs one `lsof` per candidate plus a judgement call, which is why it is second and not
1148
+ first: after two reads there are usually no candidates left to spend it on.
1149
+
1150
+ ⚠ **AND A ZERO IS NOT ABSENCE UNTIL A POSITIVE CONTROL SAYS SO.** Rule 11 demands this of every guard
1151
+ whose failure is silence, and a process probe is exactly that guard: *no matches* and *my filter is
1152
+ broken* are byte-identical outputs. **Before you report a watcher dead, a tree clean, or a machine
1153
+ idle, run the same probe once against something you KNOW is alive** — a watcher you just armed, a
1154
+ `sleep 300` you just launched — and see it come back non-empty. The `grep -v grep` case above was
1155
+ caught by exactly this and by nothing else.
1020
1156
  The general rule: **an exclusion filter is exactly as
1021
1157
  dangerous as an over-broad inclusion filter, and it fails in the direction that reads as "not there" —
1022
1158
  which is the direction that gets acted on.**
@@ -1127,6 +1263,13 @@ twice.** They are mission-independent on purpose: nothing here names a task, a l
1127
1263
  pre-commitment was never about. **State both: what fires it, and what it is ABOUT.** A correct trigger
1128
1264
  with an unstated subject executes on the first thing matching its shape, carrying the authority of the
1129
1265
  decision it was written for.
1266
+ ⚠ **AND NAMING THE SUBJECT SET IS STILL NOT ENOUGH — the hazard also arrives from the opposite
1267
+ direction.** A criterion that named its subjects correctly, twice, by hash and by enumeration, was
1268
+ voided anyway when two of those subjects were re-scoped underneath it hours later (OBS-804): not a
1269
+ document reaching for new work, but **the work moving out from under a document that has no way to
1270
+ notice.** So a pre-commitment states a THIRD thing — **what VOIDS it** — and the seat that re-scopes a
1271
+ named subject names the sealed documents it just invalidated. Form and both duties:
1272
+ *The release criterion, and the two clauses without which it grades work it never saw*, above.
1130
1273
  24. **A REMEDIATION is believed where a guard would be drilled.** Rule 11 says a guard whose failure is
1131
1274
  silence needs a positive control. **Nobody applies that to a FIX**, because a fix is not an
1132
1275
  instrument — so a shipped remediation is remembered as coverage and never re-read. One was recalled as
@@ -1273,3 +1416,32 @@ twice.** They are mission-independent on purpose: nothing here names a task, a l
1273
1416
  identity**, so no record can later attribute it to a person. And **when an injected line agrees with
1274
1417
  what you were about to decide, that is the dangerous case, not the safe one** — a line that contradicts
1275
1418
  you gets caught; one that agrees gets executed and remembered as your own decision.
1419
+
1420
+ 37. **A TASK THAT CHANGES AN OBSERVABLE CONTRACT GETS SPIKED BEFORE ITS `files[]` IS SCOPED — AND THE
1421
+ SPIKE'S REDS ARE THE BLOCKER SET. A SWEEP IS NOT.** An observable contract is execution order,
1422
+ event-stream order, a diagnostic or output SET, a CLI surface, a serialised format, or a timing
1423
+ measurement. Implement it as a **throwaway spike**, run the **FULL** suite, read the reds, *then*
1424
+ scope. Adopted 2026-08-30 (RULING-219-11). **All four parts ship together or the rule is misapplied:**
1425
+ - **The trigger question:** *could a test this task does not own be asserting the thing I am changing?*
1426
+ **"I'd have to grep to know" is a YES.**
1427
+ - **The caveat:** a spike measures **ONE implementation**. It converts *unknown* → *measured for one
1428
+ specimen*, never *unknown* → *known*. **A worker taking a different route can still red on unowned
1429
+ collateral, and that is still a PLAN defect, never a retry.**
1430
+ - **The MEASURED cost: 518 s implement + 831 s suite = 1,348 s ≈ 22.5 min.** ⛔ ***"Far cheaper" is
1431
+ WITHDRAWN.*** Run 3131 died at ~20 min, so **on the direct leg the two are EQUAL**; the spike wins
1432
+ only on what it AVOIDS downstream — a halt, a sweep, a re-scope, five rulings, a second compile and
1433
+ plan. ⚡ **Therefore it pays only where a late plan defect is EXPENSIVE TO UNWIND.** Where a defect
1434
+ would surface and fix cheaply, **the spike is pure overhead: do not run it.** That the economics and
1435
+ the trigger name the same class, derived independently — one from a stopwatch, one from a taxonomy —
1436
+ is the best evidence this rule is real, and it is why neither half may be quoted without the other.
1437
+ - **The sweep matrix, as the REASON a better sweep is not the remedy** (RULING-219-09): concept **2/3**,
1438
+ symbol **1/3**, union **3/3 on files but only 2/3 on ACTIONABLE SIGNAL** — `gate-telemetry`'s single
1439
+ symbol hit is an order-INSENSITIVE sorted comparison at `:54` that correct triage *discards*, while
1440
+ the failing test at `:101` references it not at all. ⚠ **Rigorous triage makes that discard MORE
1441
+ likely, not less.** The sweep does not fail from sloppiness, so it cannot be fixed with care.
1442
+
1443
+ **Corroboration from the same halt:** of 13 stale-order carriers, the **9** answered from real
1444
+ execution were all proven; the only **2** labelled *"inspection, not execution"* were exactly the 2 the
1445
+ suite never reached. **Evidence quality tracked execution coverage with no exceptions**, while every
1446
+ sweep-based estimate — *"~15"*, *"13"*, *"union 3/3"* — was wrong, and one 512 s suite run gave the
1447
+ right answer (**3**) first time.
@@ -38,6 +38,8 @@
38
38
  # TKR_HANDOFF_MAX_AGE_S how fresh "fresh" is (default 900).
39
39
  # TKR_CLEAR_SETTLE_S seconds to let a cleared seat settle before the re-brief (default 6).
40
40
  # TKR_BLIND_ALARM_S seconds unreadable before CONTEXT_BLIND alarms (default 120).
41
+ # TKR_CONTEXT_WINDOW context window in tokens. Set it when the banner does not name one;
42
+ # NEVER guessed — a wrong denominator makes every warn and act line wrong.
41
43
 
42
44
  set -u
43
45
  ROLE="${1:?supervising seat role required: orchestrator|overseer}"
@@ -100,10 +102,96 @@ stand_down() { tickmarkr beat "$TIER" --stand-down --seat "$SEAT" >/dev/null 2>&
100
102
  # record the hand-off. A killed watcher never runs it, which is the one case that must read STALE.
101
103
  trap stand_down EXIT
102
104
 
103
- # The seat's rendered truth lives on the model banner. Select that line first, then read a percentage
104
- # only from it: a bare numeric search across the visible window can borrow an old N% from scrollback
105
- # when a long live-run segment pushes the real field past the pane's visible width (OBS-780).
106
- context_pct() {
105
+ # ── WHERE THE NUMBER COMES FROM, and this is the whole 2.1.7-series lesson ────────────────────────────
106
+ # The shipped version read a TERMINAL RENDERING and nothing else. Every context-measurement failure of
107
+ # that series is downstream of that one choice: a rendering can be truncated by pane width, can push the
108
+ # real field out of the visible window, and can leave a stale N% in scrollback for a bare numeric search
109
+ # to borrow (OBS-780). **The value is not on the screen. It is in the session JSONL**, which is exact,
110
+ # append-only, and indifferent to how wide the pane is.
111
+ #
112
+ # ⚠ NAME THE QUANTITY, because two numbers live in that file and they differ by two orders of magnitude:
113
+ # * CONTEXT FILL = the LAST usage-bearing record's input_tokens + cache_creation + cache_read.
114
+ # This is what is in the window right now. **This is the only one a clear threshold may use.**
115
+ # * CUMULATIVE CONSUMPTION = the same sum added up over every record, plus output.
116
+ # Measured 2026-08-31 on two live sessions in this workspace: **118.5x the fill over 175
117
+ # usage records, and 28.4x over 35.**
118
+ # ⛔ **THE FACTOR IS NOT A CONSTANT — IT GROWS WITH SESSION LENGTH**, because cumulative adds a
119
+ # term per request while fill is what ONE request holds. The two numbers above are two points on
120
+ # that curve, NOT a range and NOT a property of the quantity. **Never quote a single figure as
121
+ # "the" ratio**, and never average two: a longer session yields a larger one, without limit.
122
+ # The banner shows this one too, as `sum N tok`, right beside the fill percentage.
123
+ # **Quoting fill as consumption, or consumption as fill, is wrong by that factor. Say which you mean.**
124
+ #
125
+ # ⛔ AND THE DENOMINATOR IS NOT IN THE JSONL. Its `model` field reads `claude-opus-5` with no context-size
126
+ # suffix, so a 200k default would have read three live 1M seats here at 106%, 150% and 90% — the 150% seat
127
+ # would have been auto-cleared on the first tick. **A window is resolved explicitly or read once from the
128
+ # banner; it is never assumed.** Reading a per-session CONSTANT once from the fragile surface, and the
129
+ # per-tick VARIABLE from the robust one, is the trade this makes.
130
+ # Calibrated 2026-08-31 against three live seats: banner 21/30/18% vs JSONL fill 21.2/30.0/17.9%. Same
131
+ # quantity, same direction, same denominator — so the existing WARN/ACT thresholds carry over unchanged.
132
+
133
+ CTX_JSONL=""
134
+ CTX_WINDOW="${TKR_CONTEXT_WINDOW:-0}"
135
+
136
+ # TARGET is an agent name or a pane id; `herdr agent list` carries both, plus the session uuid that names
137
+ # the JSONL. The uuid is unique, so glob for it rather than reconstructing the project-dir slug — a slug
138
+ # rule is a second thing that can rot.
139
+ resolve_jsonl() {
140
+ local sid
141
+ sid=$(herdr agent list 2>/dev/null | python3 -c '
142
+ import sys, json
143
+ try: d = json.load(sys.stdin)
144
+ except Exception: sys.exit(0)
145
+ t = sys.argv[1]
146
+ for a in d.get("result", {}).get("agents", []):
147
+ if t in (a.get("pane_id"), a.get("name")):
148
+ print((a.get("agent_session") or {}).get("value") or "")
149
+ break
150
+ ' "$TARGET" 2>/dev/null) || return 1
151
+ [ -n "$sid" ] || return 1
152
+ CTX_JSONL=$(ls "$HOME"/.claude/projects/*/"$sid".jsonl 2>/dev/null | head -1)
153
+ [ -n "$CTX_JSONL" ]
154
+ }
155
+
156
+ # The banner names its own window ("Opus 5 (1M context)"). Read it ONCE; it cannot change mid-session.
157
+ resolve_window() {
158
+ [ "$CTX_WINDOW" -gt 0 ] 2>/dev/null && return 0
159
+ local w
160
+ w=$(herdr agent read "$TARGET" --source visible --lines 8 2>/dev/null |
161
+ grep -oEi '\(([0-9]+(\.[0-9]+)?)[KM] context\)' | tail -1 |
162
+ grep -oEi '[0-9]+(\.[0-9]+)?[KM]') || return 1
163
+ case "$w" in
164
+ *[Kk]) CTX_WINDOW=$(awk -v n="${w%[Kk]}" 'BEGIN{printf "%d", n*1000}') ;;
165
+ *[Mm]) CTX_WINDOW=$(awk -v n="${w%[Mm]}" 'BEGIN{printf "%d", n*1000000}') ;;
166
+ *) return 1 ;;
167
+ esac
168
+ [ "$CTX_WINDOW" -gt 0 ] 2>/dev/null
169
+ }
170
+
171
+ jsonl_pct() {
172
+ [ -n "$CTX_JSONL" ] && [ -f "$CTX_JSONL" ] && [ "$CTX_WINDOW" -gt 0 ] 2>/dev/null || return 1
173
+ python3 -c '
174
+ import sys, json
175
+ path, window = sys.argv[1], int(sys.argv[2])
176
+ fill = 0
177
+ with open(path, encoding="utf-8", errors="replace") as fh:
178
+ for line in fh:
179
+ try: rec = json.loads(line)
180
+ except Exception: continue
181
+ usage = (rec.get("message") or {}).get("usage")
182
+ if not isinstance(usage, dict): continue
183
+ # CONTEXT FILL — this request s window occupancy. Never a running total.
184
+ n = sum(int(usage.get(k) or 0) for k in
185
+ ("input_tokens", "cache_creation_input_tokens", "cache_read_input_tokens"))
186
+ if n: fill = n
187
+ if not fill: sys.exit(1)
188
+ print(int(round(fill * 100.0 / window)))
189
+ ' "$CTX_JSONL" "$CTX_WINDOW" 2>/dev/null
190
+ }
191
+
192
+ # The banner path stays as the FALLBACK, unchanged: select the model line first, then read a percentage
193
+ # only from it, so a bare numeric search cannot borrow an old N% from scrollback (OBS-780).
194
+ banner_pct() {
107
195
  local screen banner pct
108
196
  screen=$(herdr agent read "$TARGET" --source visible --lines 8 2>/dev/null) || return 1
109
197
  banner=$(printf '%s\n' "$screen" |
@@ -114,6 +202,12 @@ context_pct() {
114
202
  [ -n "$pct" ] && printf '%s\n' "$pct" || printf 'UNREADABLE\n'
115
203
  }
116
204
 
205
+ context_pct() {
206
+ local p
207
+ if p=$(jsonl_pct) && [ -n "$p" ]; then printf '%s\n' "$p"; return 0; fi
208
+ banner_pct
209
+ }
210
+
117
211
  handoff_fresh() {
118
212
  [ -n "$HANDOFF" ] || return 1
119
213
  [ -f "$HANDOFF" ] || return 1
@@ -144,6 +238,17 @@ act_on() {
144
238
  exit 0
145
239
  }
146
240
 
241
+ # Resolve the JSONL source ONCE, and SAY which surface is in use. A watcher that silently fell back to
242
+ # the fragile path looks identical to one on the robust path, and the difference is the whole point.
243
+ if resolve_jsonl && resolve_window; then
244
+ echo "CONTEXT_SOURCE jsonl $TARGET — CONTEXT FILL from $CTX_JSONL against a ${CTX_WINDOW}-token window"
245
+ echo " this is FILL (what is in the window now), never cumulative consumption — measured 28x and 118x"
246
+ echo " apart here, and that factor GROWS with session length; never quote one figure as the ratio"
247
+ else
248
+ echo "CONTEXT_SOURCE banner $TARGET — JSONL unresolved (session=${CTX_JSONL:-none} window=${CTX_WINDOW})"
249
+ echo " falling back to the rendered banner; set TKR_CONTEXT_WINDOW to use the JSONL"
250
+ fi
251
+
147
252
  warned=0
148
253
  elapsed=0
149
254
  blind=0 # seconds in the current failed-read spell (OBS-739)
@@ -0,0 +1,137 @@
1
+ #!/bin/bash
2
+ # watch-journal.sh — wake on a run's TERMINAL and DECISION events, print one reason, and exit.
3
+ #
4
+ # WHY THIS EXISTS, 2026-08-31 (OBS-808): `tickmarkr-auto` and `tickmarkr-loop` both tell the operator to
5
+ # *"watch the run journal for its terminal event rather than polling"* — and neither skill shipped a single
6
+ # script. Correct instruction, no means to follow it, so the reader polls or sleeps through the run's end.
7
+ # Worse, `task-failed` and `consult-verdict` appeared in NO shipped script at all, while the overseer skill
8
+ # instructs seats to arm watchers on all four events. This is that instrument, for the reader who has no
9
+ # supervising tier.
10
+ #
11
+ # ⚠ It is NOT a second implementation of the scoping idiom: the arm-time line baseline below is
12
+ # `watch-parks.sh`'s, kept deliberately identical, because that file paid for two traps that a fresh
13
+ # implementation re-earns (see ARM-TIME SCOPING). `watch-parks.sh` remains the AUTHORITY seat's park
14
+ # watcher — it counts parks and speaks about rulings; this one is the general four-event watcher.
15
+ #
16
+ # usage: watch-journal.sh <runs-dir> [poll-s] [cap-s] [events-csv]
17
+ # <runs-dir> e.g. .tickmarkr/runs — tracks the NEWEST run directory, so it follows a resume or a
18
+ # fresh run without being re-aimed.
19
+ # [events-csv] default `run-end,task-human,task-failed,consult-verdict` — the four the overseer skill
20
+ # names. Narrow it only when you have a reason; a watcher on fewer events is less coverage
21
+ # wearing the same name.
22
+ #
23
+ # Prints ONE wake reason and exits. RE-ARM AFTER EVERY WAKE — the gap between a wake and its re-arm is
24
+ # unwatched, and its width is however long the reader stays busy.
25
+
26
+ set -u
27
+ RUNS="${1:?runs dir required (e.g. .tickmarkr/runs)}"
28
+ POLL="${2:-20}"
29
+ CAP="${3:-28800}"
30
+ EVENTS="${4:-run-end,task-human,task-failed,consult-verdict}"
31
+
32
+ # Config flows into a regex, so validate the shape rather than trusting it (an unquoted or unchecked
33
+ # event list is a shell/regex injection and a silently-never-matching pattern at the same time).
34
+ #
35
+ # ⚠ VALIDATE THE RAW INPUT, AND NEVER NORMALISE FIRST. An earlier draft ran `tr -d '[:space:]'` before
36
+ # this check, so `run end` passed as the event name `runend` — a watcher armed on a name no journal will
37
+ # ever carry, which polls to its cap and reports "nothing happened". Caught by a control, not by use.
38
+ # **A cleanup that rescues a typo converts a loud exit into a silent never-match**, which is the exact
39
+ # failure this file exists to prevent, so whitespace is rejected rather than stripped.
40
+ case "$EVENTS" in
41
+ *[!a-z0-9,-]*|''|*,,*|,*|*,)
42
+ echo "watch-journal.sh: events must be a comma-separated list of [a-z0-9-] names with no spaces, got '$EVENTS'" >&2
43
+ exit 64 ;;
44
+ esac
45
+ ALT=$(printf '%s' "$EVENTS" | tr ',' '|')
46
+ PAT="\"event\":\"($ALT)\""
47
+
48
+ newest_journal() {
49
+ local d
50
+ d=$(ls -t "$RUNS" 2>/dev/null | grep '^run-' | head -1)
51
+ [ -n "$d" ] && [ -f "$RUNS/$d/journal.jsonl" ] && printf '%s' "$RUNS/$d/journal.jsonl"
52
+ }
53
+
54
+ # ── ARM-TIME SCOPING — the whole correctness argument, inherited from watch-parks.sh:57-64 ─────────────
55
+ # `run-end` is a HISTORICAL RECORD once written. A whole-file grep for it finds the PREVIOUS run's
56
+ # run-end on every resume, and on any re-arm after a run has ended — so the watcher exits in its first
57
+ # poll, a supervisor re-execs it into the same instant exit, and the process table shows coverage that
58
+ # does not exist. Only lines appended AFTER arming are evidence about now.
59
+ #
60
+ # Scoping by LINE POSITION rather than by an event COUNT also sidesteps the second trap that file
61
+ # documents: `grep -c` EXITS 1 WHEN THE COUNT IS ZERO while still printing `0`, so `n=$(grep -c P f ||
62
+ # echo 0)` yields the two-line string "0\n0" and every later `[ "$n" -gt … ]` dies with `integer
63
+ # expression expected` and evaluates FALSE — i.e. the armed-at-run-start watcher, the one case that
64
+ # matters, could never report its first event. This file never counts, so it cannot inherit that.
65
+ J=$(newest_journal)
66
+ base=0
67
+ [ -n "${J:-}" ] && base=$(wc -l < "$J" 2>/dev/null || echo 0)
68
+ seen_run="${J:-}"
69
+ since_arm() { tail -n +$((base + 1)) "$1" 2>/dev/null; }
70
+
71
+ field() { printf '%s' "$2" | sed -n "s/.*\"$1\":\"\([^\"]*\)\".*/\1/p" | head -1; }
72
+ bucket() { printf '%s' "$2" | sed -n "s/.*\"$1\":\[\([^]]*\)\].*/\1/p" | head -1; }
73
+
74
+ report() {
75
+ local line="$1" run="$2" ev
76
+ ev=$(field event "$line")
77
+ case "$ev" in
78
+ run-end)
79
+ local tv done failed human blocked pending verdict
80
+ tv=$(field tipVerify "$line")
81
+ done=$(bucket done "$line"); failed=$(bucket failed "$line")
82
+ human=$(bucket human "$line"); blocked=$(bucket blocked "$line")
83
+ pending=$(bucket pending "$line")
84
+ # GREEN IS A CONJUNCTION, AND THE SHORT FORM OF IT IS WRONG. "run-end plus tip verify" passes a run
85
+ # that ended `done=[T1,T3,T4] human=[T2]` — three delivered, one PARKED — and calling that green is
86
+ # how a park becomes invisible. Grade every clause here so the reader never has to remember to.
87
+ if [ "$tv" != "failed" ] && [ -z "$failed$human$blocked$pending" ]; then
88
+ verdict="GREEN"
89
+ else
90
+ verdict="NOT GREEN"
91
+ fi
92
+ echo "RUN_END $run — $verdict (tipVerify=${tv:-unknown})"
93
+ echo " done=[${done}] failed=[${failed}] human=[${human}] blocked=[${blocked}] pending=[${pending}]"
94
+ [ "$verdict" = "GREEN" ] \
95
+ && echo " all four buckets empty and tip verify is not failed — this run is green" \
96
+ || echo " a non-empty bucket above is the reason; name it, never report this run as green"
97
+ ;;
98
+ task-human)
99
+ echo "TASK_HUMAN $(field taskId "$line") — $run"
100
+ echo " $(printf '%s' "$line" | sed -n 's/.*"reason":"\([^"]\{0,160\}\).*/\1/p')"
101
+ echo " a park waits for a DECISION; read the gate evidence, then \`tickmarkr approve $run $(field taskId "$line")\` or re-scope"
102
+ ;;
103
+ task-failed)
104
+ echo "TASK_FAILED $(field taskId "$line") — $run"
105
+ echo " $(printf '%s' "$line" | sed -n 's/.*"error":"\([^"]\{0,160\}\).*/\1/p')"
106
+ echo " the run may continue on independent tasks; this task did not deliver"
107
+ ;;
108
+ consult-verdict)
109
+ echo "CONSULT_VERDICT $(field taskId "$line") action=$(field action "$line") — $run"
110
+ echo " $(printf '%s' "$line" | sed -n 's/.*"notes":"\([^"]\{0,160\}\).*/\1/p')"
111
+ ;;
112
+ *)
113
+ echo "JOURNAL_EVENT ${ev:-unparseable} — $run"
114
+ echo " ${line:0:200}"
115
+ ;;
116
+ esac
117
+ }
118
+
119
+ elapsed=0
120
+ while [ "$elapsed" -lt "$CAP" ]; do
121
+ sleep "$POLL"
122
+ elapsed=$((elapsed + POLL))
123
+
124
+ J=$(newest_journal)
125
+ [ -z "${J:-}" ] && continue
126
+
127
+ # A new run resets the baseline — its whole journal is unseen by definition.
128
+ if [ "$J" != "$seen_run" ]; then seen_run="$J"; base=0; fi
129
+
130
+ hit=$(since_arm "$J" | grep -E "$PAT" | head -1)
131
+ if [ -n "$hit" ]; then
132
+ report "$hit" "$(basename "$(dirname "$J")")"
133
+ exit 0
134
+ fi
135
+ done
136
+
137
+ echo "WATCH_CAP_REACHED — no ${EVENTS} in ${CAP}s (newest: $(basename "$(dirname "${J:-none/none}")"))"