@plot-pm/board 0.16.1 → 0.16.3

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1,705 +0,0 @@
1
- #!/usr/bin/env bash
2
- # Plot helper: the WorkerMonitor — watches the PROCESS a dispatched agent runs on.
3
- #
4
- # RUN, NOT SOURCED, and started by `start_worker()` in `plot-dispatch.sh` as a
5
- # child of the wrapper. It is never invoked by hand in normal operation: a
6
- # monitor an operator has to remember to start is one that will be missing on
7
- # the day it matters.
8
- #
9
- # ═══════════════════════════════════════════════════════════════════════════
10
- # TWO FINDINGS, AND ONLY TWO
11
- # ═══════════════════════════════════════════════════════════════════════════
12
- #
13
- # gone the agent pid names no live process
14
- # idle the pid lives, its TRANSCRIPT has been silent past the window with
15
- # no child process burning CPU behind it, across two consecutive
16
- # passes, the tree did not change between them, AND commits already
17
- # exist on the branch
18
- #
19
- # Anything else is `silent`, which is not a finding and is not published. The
20
- # distinction is the whole point: a monitor that reports every quiet moment
21
- # teaches an operator to ignore it, and then it is worse than absent.
22
- #
23
- # ═══════════════════════════════════════════════════════════════════════════
24
- # IT READS THE AGENT, NOT THE MACHINE
25
- # ═══════════════════════════════════════════════════════════════════════════
26
- #
27
- # Until 2026-09-02 the reading was a 0.4 s CPU sample of the agent's subtree,
28
- # taken twice ~30 s apart. An agent waiting on a model response burns no subtree
29
- # CPU, so a false zero was the COMMON reading rather than the rare one, and no
30
- # sampling interval closes that gap: a slow model response is indistinguishable
31
- # from a dead one by CPU alone. The rule ended eleven dispatched workers across
32
- # two days, several holding uncommitted work.
33
- #
34
- # A `claude -p` session appends a line to its transcript for every model turn,
35
- # tool call and tool result. Seconds since the newest line reads whether the
36
- # AGENT has produced anything — the question the CPU sample was standing in for.
37
- #
38
- # ═══════════════════════════════════════════════════════════════════════════
39
- # TWO READINGS, BECAUSE NEITHER ANSWERS IT ALONE
40
- # ═══════════════════════════════════════════════════════════════════════════
41
- #
42
- # A TRANSCRIPT IS EQUALLY QUIET IN TWO CASES that must end differently: an agent
43
- # waiting on a model, and an agent waiting on its own 20-minute test suite.
44
- # Wave 1 measured 7547 quiet stretches across 23 sessions on 2026-09-02, and 28
45
- # of the 37 that passed 30 s were the second kind — the four longest being this
46
- # repo's own gates, `gh pr checks --watch` at 600.8 s and `pnpm run test:board`
47
- # at 600.3 s.
48
- #
49
- # So the window (`PLOT_MONITOR_QUIET_SECONDS`, 900 s) is a GATE and the CPU is
50
- # the verdict beside it:
51
- #
52
- # transcript inside the window → busy (it just wrote something)
53
- # past the window, a child burning CPU → busy (its build is running)
54
- # past the window, no child on a core → quiet (it has stopped)
55
- # no transcript readable → unknown (see the fallback)
56
- #
57
- # THE CPU'S ROLE IS INVERTED FROM THE OLD RULE, and that is what makes it sound
58
- # here. The rejected rule read a FROZEN clock as a stall; this reads a MOVING
59
- # clock as life. A moving clock proves something is happening; a frozen one
60
- # proved nothing, which is precisely why it could not be trusted alone.
61
- #
62
- # ═══════════════════════════════════════════════════════════════════════════
63
- # WHY `idle` STILL CARRIES THE TREE AND COMMIT CONDITIONS
64
- # ═══════════════════════════════════════════════════════════════════════════
65
- #
66
- # What separated the three stalls measured on 2026-08-30 is that each had
67
- # already COMMITTED and then gone quiet:
68
- #
69
- # quiet, tree unchanged, commits present → idle
70
- # quiet, tree unchanged, no commits yet → silent (it may be thinking)
71
- # quiet, tree CHANGED between samples → silent (something is happening)
72
- #
73
- # THE MIDDLE ROW IS WHERE THE FALSE POSITIVES WOULD HAVE BEEN. An agent given a
74
- # hard first slice is quiet for a long time with nothing to show; calling that a
75
- # stall is the cry-wolf that costs the finding its readers. The extra two
76
- # conditions are not caution — they are what makes the word mean something.
77
- #
78
- # ═══════════════════════════════════════════════════════════════════════════
79
- # WHERE NO TRANSCRIPT CAN BE READ, IT SAYS SO
80
- # ═══════════════════════════════════════════════════════════════════════════
81
- #
82
- # The reading is UNAVAILABLE, not zero and not failed — the contract
83
- # `the-registry-supervises-its-agents` settled. The monitor then publishes
84
- # nothing and `Worker bound` is what ends the worker. The cost is stated rather
85
- # than hidden: a genuinely stuck agent holds a desk for up to 8 hours, which is
86
- # smaller than the measured cost of killing working ones.
87
- #
88
- # ═══════════════════════════════════════════════════════════════════════════
89
- # IT IS NOT CALLED `stalled`, AND THAT IS A CONTRACT
90
- # ═══════════════════════════════════════════════════════════════════════════
91
- #
92
- # The spec owns `stalled` for an AGENT fact — *"exited 0, unlanded work, no
93
- # PR"* (DESIGN-agent.md). A stalled agent has work to rescue; an idle worker may
94
- # just be waiting on the network. An earlier draft reused the name and put a
95
- # process fact on the agent side, which is the exact confusion CLAUDE.md's
96
- # Machine/Registry split exists to prevent: this monitor watches a PROCESS, so
97
- # its vocabulary is Worker-side.
98
- #
99
- # ═══════════════════════════════════════════════════════════════════════════
100
- # TWO SAMPLES, NEVER ONE
101
- # ═══════════════════════════════════════════════════════════════════════════
102
- #
103
- # A single idle reading is a process caught between syscalls. The COMPARISON is
104
- # the finding, so the monitor keeps the previous answer — one piece of state,
105
- # and derived rather than recorded: lose it (restart the monitor, say) and the
106
- # next pass rebuilds it, at one interval's delay. Nothing is persisted, because
107
- # nothing needs to be.
108
- #
109
- # ═══════════════════════════════════════════════════════════════════════════
110
- # IT MAKES NO HOST CALL AT ALL
111
- # ═══════════════════════════════════════════════════════════════════════════
112
- #
113
- # Not "few" — none. A monitor on a ~30s cadence that asks the host has become an
114
- # AgentMonitor with a fast loop, and the rate problem follows it: 127 git
115
- # processes per scan is what that costs in this repo. Every question here is
116
- # answered by the process table or by a local git ref. `commits present` is
117
- # counted against the LOCAL `origin/main` ref, never a fetch — and when that ref
118
- # is missing the question is unanswerable rather than answered zero, so `idle`
119
- # does not fire. A failure to observe is not evidence of something to see; the
120
- # same rule `plot_worker_task_state` reached the hard way after a fallback read
121
- # every clean branch as `stalled` in a repo with no remote.
122
- #
123
- # ═══════════════════════════════════════════════════════════════════════════
124
- # WHY IT IS THE WRAPPER'S CHILD
125
- # ═══════════════════════════════════════════════════════════════════════════
126
- #
127
- # `plot-dispatch.sh` does not spawn the agent directly — it spawns an `sh -c`
128
- # wrapper that backgrounds the agent, records its pid, `wait`s for it and writes
129
- # `.plot-worker.exit`. That wrapper ALREADY outlives its agent by construction,
130
- # because otherwise there would be no exit code to record; the comment at
131
- # `plot-dispatch.sh` states it outright: *"--stop kills the agent, the wrapper
132
- # survives to record the code."*
133
- #
134
- # A monitor that is its child inherits that survival. A SIBLING would not:
135
- # two processes started side by side are independently mortal, so the monitor
136
- # could be killed or crash with nothing noticing — which is the failure being
137
- # fixed, one level up.
138
- #
139
- # ═══════════════════════════════════════════════════════════════════════════
140
- # IT INHERITS THE STARTUP WINDOW RATHER THAN WIDENING IT
141
- # ═══════════════════════════════════════════════════════════════════════════
142
- #
143
- # There is a sub-millisecond gap after the wrapper starts and before it writes
144
- # `.plot-worker.pid`; a scan landing in it reads `none` — honest. This monitor
145
- # starts inside the same wrapper (the monitors are backgrounded BEFORE the
146
- # `printf > "$PLOT_PID_FILE"`), so its first pass can genuinely land in that
147
- # window. An ABSENT pid file therefore means *not yet*, never `gone`: reporting
148
- # a dead agent because its birth has not been recorded would make the monitor's
149
- # loudest finding also its least trustworthy.
150
- #
151
- # ═══════════════════════════════════════════════════════════════════════════
152
- # PUBLISHING, BEFORE THERE IS A CHANNEL
153
- # ═══════════════════════════════════════════════════════════════════════════
154
- #
155
- # The channel is a local socket under `.plot/`, and it is
156
- # `feature/the-channel-carries-the-findings`. Until it exists, a finding is
157
- # published by being APPENDED to a file the fleet already knows how to ignore.
158
- #
159
- # The name matters more than it looks. `plot-worker-state.sh` excludes Plot's
160
- # own records from both the dirty-tree filter and the marker search with ONE
161
- # pattern — `PLOT_WORKER_RECORD='\.plot-worker\.'` — after those two exclusions
162
- # had already drifted apart once. A finding file named `.plot-worker.monitor.*`
163
- # is covered by that pattern for free; anything else would make every monitored
164
- # worktree read as holding unlanded work, which is `stalled` for a fleet that
165
- # is perfectly healthy.
166
- #
167
- # THAT EXCLUSION IS ALSO WHY THE TREE FINGERPRINT CAN TRUST `git status`. This
168
- # monitor writes into the worktree it is watching, once per finding — so a
169
- # fingerprint over raw `git status` would see the monitor's own file appear and
170
- # read it as the tree changing, and the monitor would suppress `idle` forever on
171
- # the strength of its own output. `plot_worker_dirty_filter` drops exactly that
172
- # prefix, which is why the fingerprint goes through it rather than around it.
173
- set -uo pipefail
174
-
175
- usage() {
176
- cat >&2 <<'EOF'
177
- Usage: plot-worker-monitor.sh [--once]
178
-
179
- Started by plot-dispatch.sh inside the worker's wrapper. Reads its subject from
180
- the environment, exactly as the wrapper's other children do:
181
-
182
- PLOT_BRANCH the branch this worker is on
183
- PLOT_WORKTREE the desk it sits at
184
- PLOT_PID_FILE where the wrapper records the AGENT's pid
185
- PLOT_MONITOR_FILE where findings are published (default:
186
- $PLOT_WORKTREE/.plot-worker.monitor.worker.jsonl)
187
- PLOT_MONITOR_INTERVAL seconds between passes (default 30)
188
- PLOT_SESSION_ID the launch session id; the handle when the manifest
189
- carries no `resumeId`
190
- PLOT_MANIFEST_FILE the agent's manifest, whose `resumeId` names the current
191
- conversation. With neither set, `idle` is judged on the
192
- desk alone and one line on stderr says so.
193
-
194
- --once take one sample and exit, rather than looping. A single pass can
195
- never publish `idle` — that needs two — so this is how a test drives
196
- the `gone` arm and the "one sample says nothing" property directly.
197
- EOF
198
- }
199
-
200
- once=0
201
- while [ $# -gt 0 ]; do
202
- case "$1" in
203
- --once) once=1 ;;
204
- -h|--help) usage; exit 0 ;;
205
- *) echo "plot-worker-monitor: unknown argument '$1'" >&2; usage; exit 2 ;;
206
- esac
207
- shift
208
- done
209
-
210
- # THE MONITOR'S NAME IS ITS CONTRACT. It travels into every finding, and the
211
- # board will key on it — a WorkerMonitor `idle` and an AgentMonitor `owes a
212
- # review` must be distinguishable in the entry itself, which the plan's
213
- # attention slice requires and which a shared label would make impossible.
214
- monitor='WorkerMonitor'
215
-
216
- branch="${PLOT_BRANCH:-}"
217
- worktree="${PLOT_WORKTREE:-}"
218
- interval="${PLOT_MONITOR_INTERVAL:-30}"
219
- pid_file="${PLOT_PID_FILE:-${worktree:+$worktree/.plot-worker.pid}}"
220
-
221
- # ONE ANSWER TO "IS MY SUBJECT STILL THERE?", shared with the AgentMonitor
222
- # rather than written twice. `plot-worker-state.sh` carried five of six states
223
- # in duplicate until the copies drifted on the sixth; two monitors deciding
224
- # independently when to stop would drift the same way, and half a fix for a leak
225
- # looks exactly like a fix.
226
- #
227
- # Sourced from THIS script's directory, so a monitor started from a worktree's
228
- # own copy of the scripts finds that copy's helper — which is how every
229
- # dispatched worker runs.
230
- # shellcheck source=./plot-monitor-subject.sh
231
- . "$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)/plot-monitor-subject.sh"
232
-
233
- # THE DEFAULT PATH IS DERIVED, NOT REQUIRED. The wrapper passes
234
- # PLOT_MONITOR_FILE explicitly (one env var per path, so no quoting level inside
235
- # the single-quoted `sh -c` can mangle a path with spaces — the convention the
236
- # exit and pid files already use). The fallback exists so a hand-run monitor in
237
- # a worktree still writes somewhere sensible rather than refusing.
238
- findings="${PLOT_MONITOR_FILE:-${worktree:+$worktree/.plot-worker.monitor.worker.jsonl}}"
239
-
240
- # THE CPU SAMPLER IS BORROWED, NOT REBUILT. `plot_worker_activity` already sums
241
- # a pid's whole DESCENDANT subtree across a short interval and answers
242
- # `working`/`idle`/"" — including the awk that parses `[[HH:]MM:]SS.ss` from the
243
- # right so an hour of CPU does not wrap at 60, and the one-`ps`-snapshot walk
244
- # that avoids forking a process per descendant. Writing a second sampler beside
245
- # it would be two implementations of one measurement, drifting; this repo has
246
- # already paid for that once, in the classification `plot-worker-state.sh` was
247
- # extracted to hold.
248
- #
249
- # SOURCED WITH A GUARD because a monitor whose helper is missing must still say
250
- # so rather than die silently in a detached shell nobody is reading.
251
- plot_state_lib="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)/plot-worker-state.sh"
252
- # shellcheck source=plot-worker-state.sh
253
- if [ -r "$plot_state_lib" ]; then . "$plot_state_lib"; fi
254
-
255
- # THE TRANSCRIPT READER — the primary reading, sourced beside the CPU sampler
256
- # rather than replacing it. What each answers is different in kind: the
257
- # transcript says whether the AGENT has produced anything, the CPU says whether
258
- # a CHILD is on a core. `idle` now needs both to agree.
259
- plot_transcript_lib="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)/plot-transcript-quiet.sh"
260
- # shellcheck source=plot-transcript-quiet.sh
261
- if [ -r "$plot_transcript_lib" ]; then . "$plot_transcript_lib"; fi
262
-
263
- # THE CONVERSATION HANDLE — `session_handle`, the one the loop hands the prompt.
264
- # The manifest's `resumeId`, else `PLOT_SESSION_ID`, both read from the
265
- # environment the wrapper passes down. Never `plot_manifest_for_worktree`: it
266
- # resolves `--show-toplevel` to the desk and ignores `Agent registry`, so it
267
- # names a directory that does not exist (#1086).
268
- plot_manifest_lib="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)/plot-agent-manifest.sh"
269
- # shellcheck source=plot-agent-manifest.sh
270
- if [ -r "$plot_manifest_lib" ]; then . "$plot_manifest_lib"; fi
271
-
272
- # HOW LONG A TRANSCRIPT MUST BE QUIET BEFORE THE QUESTION IS EVEN ASKED.
273
- #
274
- # 900 s, and the number comes from wave 1's measurement rather than from taste.
275
- # `plot-quiet-stretch.sh` read 7547 quiet stretches across 23 sessions in 21
276
- # worktrees on 2026-09-02:
277
- #
278
- # p50 0s p90 2.6s p99 15.6s max 600.8s
279
- #
280
- # 900 s is 1.5x that maximum and 57x the p99. Every stretch ever measured on
281
- # this estate clears it with five minutes to spare.
282
- #
283
- # THE MAXIMUM IS NOT A DISTRIBUTION'S TAIL — IT IS A CEILING, and that is why
284
- # the threshold alone is not the answer. The four longest stretches are this
285
- # repo's own gates: `gh pr checks --watch` at 600.8s, `pnpm run test:board` at
286
- # 600.3s, `pnpm run test:reconcile` at 584.9s and 575.5s. They cluster at 600
287
- # because that is where a watch command and a test runner time out, not because
288
- # an agent's quiet naturally ends there. A project with a slower suite produces
289
- # a longer one, and any single number picked from this sample would kill its
290
- # workers on the day it adopted Plot.
291
- #
292
- # SO THE THRESHOLD IS A GATE, NOT THE VERDICT. Past it, the monitor still asks
293
- # whether a child process is on a core — see `sample_verdict`. The threshold
294
- # says *this has gone on long enough to be worth asking about*; the CPU reading
295
- # answers *and there is nothing running*. Together they separate the two cases
296
- # a transcript cannot tell apart on its own, both of which look identically
297
- # quiet: an agent waiting on its own 20-minute command, and an agent that has
298
- # stopped.
299
- : "${PLOT_MONITOR_QUIET_SECONDS:=900}"
300
-
301
- json_escape() { # $1 = raw → prints a JSON-safe string body
302
- printf '%s' "$1" | python3 -c 'import json,sys; sys.stdout.write(json.dumps(sys.stdin.read())[1:-1])' 2>/dev/null \
303
- || printf '%s' "$1" | sed 's/\\/\\\\/g; s/"/\\"/g'
304
- }
305
-
306
- # A FINDING CARRIES FOUR FIELDS: `finding`, `since`, `evidence`, `measuredAt`.
307
- #
308
- # `since` AND `measuredAt` ARE DIFFERENT TIMES, and this is the slice where they
309
- # start to differ. `measuredAt` is when this reading was taken; `since` is when
310
- # the finding first held. A finding that has held for twenty minutes and one
311
- # taken twenty minutes ago are not the same fact, and an operator triaging a
312
- # board needs the first — so `since` is carried forward across republishes and
313
- # only reset when the finding changes.
314
- publish() { # $1=finding $2=evidence $3=since
315
- local now
316
- now=$(date -u +%Y-%m-%dT%H:%M:%SZ)
317
- local line
318
- line=$(printf '{"monitor":"%s","branch":"%s","worktree":"%s","finding":"%s","since":"%s","evidence":"%s","measuredAt":"%s"}' \
319
- "$monitor" \
320
- "$(json_escape "$branch")" \
321
- "$(json_escape "$worktree")" \
322
- "$(json_escape "$1")" \
323
- "${3:-$now}" \
324
- "$(json_escape "$2")" \
325
- "$now")
326
- # Both destinations, deliberately. The file is what a test and a future
327
- # subscriber read; stdout lands in `.plot-worker.log` beside the agent's own
328
- # output, where an operator tailing a worker sees it without knowing a second
329
- # file exists.
330
- [ -n "$findings" ] && printf '%s\n' "$line" >> "$findings" 2>/dev/null
331
- printf 'plot-monitor %s\n' "$line"
332
- }
333
-
334
- # ---------------------------------------------------------------------------
335
- # THE PORTS — seven named seams, so every branch is reachable from a test
336
- # ---------------------------------------------------------------------------
337
- #
338
- # Each of these is one question against the machine, and each is a `monitor_*`
339
- # function for one reason: a test sources this file with `PLOT_MONITOR_NO_MAIN`
340
- # and REDEFINES them. That is what makes the interesting branches reachable at
341
- # all — a pid that dies between two samples, a tree that changes between two
342
- # readings, and a subtree whose CPU is frozen are all states a real machine will
343
- # not produce on demand, and a test that waits for one is a test that flakes.
344
- #
345
- # The seams are the ports; the sampler below is the logic. Nothing between them
346
- # touches the machine directly.
347
-
348
- # Does the agent pid name a live process?
349
- #
350
- # THREE ANSWERS, NOT TWO. `0` alive, `1` dead, `2` UNKNOWN — and the third is
351
- # the startup window. The wrapper backgrounds this monitor BEFORE it writes the
352
- # pid file, so an absent or empty file means the birth has not been recorded
353
- # yet. Collapsing that into `dead` would make `gone` fire on every worker's
354
- # first pass, which is the one moment it is guaranteed to be wrong.
355
- monitor_pid_alive() { # → 0 alive | 1 dead | 2 unknown (not recorded yet)
356
- local pid
357
- [ -n "$pid_file" ] && [ -s "$pid_file" ] || return 2
358
- pid=$(cat "$pid_file" 2>/dev/null | tr -d '[:space:]')
359
- [ -n "$pid" ] || return 2
360
- case "$pid" in *[!0-9]*) return 2 ;; esac
361
- kill -0 "$pid" 2>/dev/null && return 0
362
- return 1
363
- }
364
-
365
- # The agent pid as recorded, or "" when it has not been recorded.
366
- monitor_pid() {
367
- [ -n "$pid_file" ] && [ -s "$pid_file" ] || return 0
368
- cat "$pid_file" 2>/dev/null | tr -d '[:space:]'
369
- }
370
-
371
- # Is the agent's subtree burning CPU? `working` | `idle` | "" (nothing to
372
- # measure). Delegated wholesale to the borrowed sampler.
373
- monitor_activity() { # $1=pid → working | idle | ""
374
- command -v plot_worker_activity >/dev/null 2>&1 || return 0
375
- plot_worker_activity "$1"
376
- }
377
-
378
- # How long has the AGENT at this desk produced nothing? The primary reading.
379
- #
380
- # `unavailable` where no transcript can be read, and that word travels all the
381
- # way to the verdict rather than being collapsed into a number. Settled by
382
- # `the-registry-supervises-its-agents`: a capability the adopting project does
383
- # not provide is UNAVAILABLE, never failed and never zero. A missing helper
384
- # answers the same way — a monitor whose reader is absent must say it cannot
385
- # see, not that it saw nothing happen.
386
- monitor_transcript_quiet() { # → seconds | unavailable
387
- command -v plot_transcript_quiet_seconds >/dev/null 2>&1 || { printf 'unavailable'; return 0; }
388
- plot_transcript_quiet_seconds "$worktree"
389
- }
390
-
391
- # Has THIS worker's conversation written yet?
392
- #
393
- # THE DESK-WIDE NUMBER CANNOT SAY. After a hop to a new branch the loop mints a
394
- # fresh handle, and the new conversation has no transcript file until its first
395
- # line. Until then the desk's newest file is the PREVIOUS slice's, and its
396
- # silence is not this worker's. So the monitor asks the loop's own probe with
397
- # the loop's own handle: one probe, two readers, one answer.
398
- #
399
- # THREE ANSWERS, AND THE THIRD IS NOT THE SECOND. `0` the handle's file exists,
400
- # `1` it does not, `2` there is no handle to ask about. `plot_transcript_exists`
401
- # reads *no handle* as *no file*, which suits `session_flag`; here it would
402
- # make a hand-started monitor read every quiet worker as unspoken and disable
403
- # `idle` silently. So the handle is checked here, before the probe.
404
- monitor_conversation_spoken() { # → 0 spoken | 1 unspoken | 2 no handle
405
- command -v session_handle >/dev/null 2>&1 || return 2
406
- command -v plot_transcript_exists >/dev/null 2>&1 || return 2
407
- local handle
408
- handle=$(session_handle) || return 2
409
- [ -n "$handle" ] || return 2
410
- plot_transcript_exists "$worktree" "$handle" && return 0
411
- return 1
412
- }
413
-
414
- # A cheap stand-in for "the tree as it is right now", compared between passes.
415
- #
416
- # IT GOES THROUGH `plot_worker_dirty_filter`, which is not an optimisation — it
417
- # is what stops the monitor from watching itself. This script appends to
418
- # `.plot-worker.monitor.worker.jsonl` INSIDE the worktree it is watching, so a
419
- # raw `git status` fingerprint would change every time the monitor published and
420
- # `idle` could never hold for two passes. The filter drops the `.plot-worker.`
421
- # prefix (and editor leftovers, and tool scratch) for exactly the reasons
422
- # recorded where it is defined.
423
- #
424
- # THE FILTERED FILE LIST, NOT A CONTENT HASH. What is being asked is *did
425
- # anything happen here*, and an agent at work adds, removes and renames files
426
- # far more often than it rewrites one in place at byte-identical length. A
427
- # content hash over a large tree on a 30s loop would also be the one expensive
428
- # thing in an otherwise cheap monitor.
429
- monitor_tree_fingerprint() { # → an opaque string; unchanged means unchanged
430
- [ -n "$worktree" ] && [ -d "$worktree" ] || { printf 'no-tree'; return 0; }
431
- local status
432
- status=$(git -C "$worktree" status --porcelain 2>/dev/null)
433
- local head
434
- head=$(git -C "$worktree" rev-parse HEAD 2>/dev/null || printf 'no-head')
435
- # HEAD is part of the fingerprint too: an agent that COMMITS between two
436
- # passes has plainly done something, and its status output may well be
437
- # identical either side of the commit.
438
- if command -v plot_worker_dirty_filter >/dev/null 2>&1; then
439
- printf '%s\n%s' "$head" "$(plot_worker_dirty_filter "$status")"
440
- else
441
- printf '%s\n%s' "$head" "$status"
442
- fi
443
- }
444
-
445
- # Are there commits on this branch yet?
446
- #
447
- # THE THIRD CONDITION ON `idle`, and the one that separates a stall from an
448
- # agent still thinking about a hard first slice.
449
- #
450
- # COUNTED AGAINST THE LOCAL `origin/<default>` REF — never a fetch, because this
451
- # monitor makes no network call. And when there is no such ref the question is
452
- # UNANSWERABLE, so this returns 2 and `idle` does not fire: counting against
453
- # nothing would count the whole history from the root commit and read every
454
- # branch in a remote-less repo as having committed, which is the failure
455
- # `plot_worker_task_state` records having made in the other direction.
456
- monitor_has_commits() { # → 0 yes | 1 no | 2 unanswerable
457
- [ -n "$worktree" ] && [ -d "$worktree" ] || return 2
458
- local base n
459
- base=$(git -C "$worktree" symbolic-ref --quiet --short refs/remotes/origin/HEAD 2>/dev/null)
460
- [ -n "$base" ] || { git -C "$worktree" rev-parse --verify --quiet origin/main >/dev/null 2>&1 && base='origin/main'; }
461
- [ -n "$base" ] || return 2
462
- # COUNT THE AGENT'S WORK, NOT THE BRANCH'S COMMITS. `plot-dispatch.sh:2074`
463
- # writes `commit --allow-empty -m "plot: claim <branch>"` BEFORE the agent
464
- # starts, so `$base..HEAD` is never zero on a dispatched branch and this
465
- # condition could never refuse an `idle`. Measured 2026-08-30 (#538 red in CI):
466
- # a worker burning CPU in `yes > /dev/null` was reported idle, because the one
467
- # condition that could have saved it was satisfied by bookkeeping the agent did
468
- # not do.
469
- #
470
- # The `-- .` pathspec is what does it: `rev-list` with a pathspec keeps only
471
- # commits that TOUCHED A FILE, and the claim is empty by construction
472
- # (`--allow-empty`). That is a property rather than a message match — a claim
473
- # whose wording changes still reads as empty, and an agent committing an empty
474
- # marker of its own is correctly not counted as work either.
475
- n=$(git -C "$worktree" rev-list --count "$base..HEAD" -- . 2>/dev/null) || return 2
476
- case "$n" in ''|*[!0-9]*) return 2 ;; esac
477
- [ "$n" -gt 0 ] && return 0
478
- return 1
479
- }
480
-
481
- # ---------------------------------------------------------------------------
482
- # THE SAMPLER — one pass, using only the ports above
483
- # ---------------------------------------------------------------------------
484
- #
485
- # THE STATE IS TWO VARIABLES AND IT IS DERIVED. `prev_verdict` is the previous
486
- # pass's answer and `prev_tree` its fingerprint; `since` is when the CURRENT
487
- # published finding first held. Nothing is written down: kill the monitor and
488
- # the next one rebuilds all three, one interval late. That is the plan's "one
489
- # piece of state, derived rather than recorded", and it is why a monitor
490
- # restart costs an interval rather than a wrong answer.
491
- prev_verdict=''
492
- prev_tree=''
493
- published=''
494
- since=''
495
-
496
- # What this pass sees, before the two-sample rule is applied.
497
- #
498
- # THE ORDER IS LOAD-BEARING. `gone` is asked FIRST because a dead pid makes
499
- # every other question meaningless — you cannot measure the CPU of a subtree
500
- # that is not there, and `plot_worker_activity` would answer "" for it anyway,
501
- # which is indistinguishable from a live pid with no children.
502
- sample_verdict() { # → gone | quiet | busy | unknown | unspoken
503
- local alive
504
- monitor_pid_alive; alive=$?
505
- [ "$alive" = 1 ] && { printf 'gone'; return; }
506
- # `unknown` is the startup window: the wrapper has not recorded the pid yet.
507
- # Not a finding, and NOT `gone`.
508
- [ "$alive" = 2 ] && { printf 'unknown'; return; }
509
-
510
- # THE TRANSCRIPT IS ASKED FIRST, and it is asked instead of the CPU rather
511
- # than beside it. Until 2026-09-02 this read `plot_worker_activity` alone and
512
- # called a frozen 0.4 s CPU sample `quiet`; that rule ended eleven dispatched
513
- # workers across two days, several holding uncommitted work. An agent waiting
514
- # on a model response burns no subtree CPU, so a false zero was the COMMON
515
- # reading rather than the rare one, and no sampling interval closes that gap.
516
- #
517
- # A `claude -p` session appends to its transcript for every turn, tool call
518
- # and tool result. Seconds since the newest line is a direct reading of
519
- # whether the AGENT has done anything — which is the question the monitor was
520
- # always trying to ask.
521
- local quiet
522
- quiet=$(monitor_transcript_quiet)
523
-
524
- # UNAVAILABLE IS NOT A FINDING, and this is where the plan's fallback lands.
525
- # Where no transcript can be read there is no reading that distinguishes
526
- # thinking from stuck, so the monitor invents none: it reports `unknown`,
527
- # publishes nothing, and `Worker bound` is what ends the worker. The cost is
528
- # stated rather than hidden — a genuinely stuck agent then holds a desk for up
529
- # to 8 hours, which is smaller than the measured cost of the rule this
530
- # replaces.
531
- case "$quiet" in
532
- ''|unavailable) printf 'unknown'; return ;;
533
- *[!0-9]*) printf 'unknown'; return ;;
534
- esac
535
-
536
- # Inside the window, the agent has produced output recently. Nothing else
537
- # needs asking: no CPU sample can overturn a line written seconds ago.
538
- if [ "$quiet" -lt "$PLOT_MONITOR_QUIET_SECONDS" ]; then printf 'busy'; return; fi
539
-
540
- # PAST THE WINDOW, AND ONLY HERE, ASK WHETHER THIS CONVERSATION HAS WRITTEN.
541
- # The number is the desk's; a new conversation with no file yet has produced
542
- # none of its silence. `unspoken` is a reading that was made, and it is not
543
- # `unknown`, which is no reading. Only `1` answers it: with no handle (`2`)
544
- # the verdict is judged on the desk alone, as before this port existed.
545
- #
546
- # LAZY ON PURPOSE. `session_handle` starts one `node` (about 35 ms), so a
547
- # worker inside the window never pays it.
548
- local spoken
549
- monitor_conversation_spoken; spoken=$?
550
- [ "$spoken" = 1 ] && { printf 'unspoken'; return; }
551
-
552
- # PAST THE WINDOW, THE SECOND READING DECIDES — and it answers a question the
553
- # transcript cannot. A transcript is equally quiet whether the agent is
554
- # waiting on a model or waiting on its own 20-minute test suite. 28 of the 37
555
- # over-window stretches wave 1 measured were the latter.
556
- #
557
- # So the CPU is consulted for what it CAN say: `working` means a child is on a
558
- # core, and an agent whose build is running has not stopped. That is not the
559
- # rejected rule returning — the rejected rule read `idle` as a stall, and this
560
- # reads `working` as life. The asymmetry is the point: a moving clock proves
561
- # something is happening, while a frozen one proved nothing, which is exactly
562
- # why it could not be trusted alone.
563
- case "$(monitor_activity "$(monitor_pid)")" in
564
- working) printf 'busy' ;;
565
- # `idle` (frozen subtree clock) and "" (no child holding a clock at all)
566
- # agree here: fifteen minutes of transcript silence with nothing burning CPU
567
- # behind it. Unlike the old rule, "" is not refused — a live pid with no
568
- # child is precisely an agent that has stopped, and it only reaches this
569
- # line after the window has already elapsed.
570
- *) printf 'quiet' ;;
571
- esac
572
- }
573
-
574
- # One full pass: sample, apply the two-sample rule, publish only on a change.
575
- monitor_pass() {
576
- local verdict tree evidence finding
577
- tree=$(monitor_tree_fingerprint)
578
- verdict=$(sample_verdict)
579
-
580
- finding=''
581
- evidence=''
582
- case "$verdict" in
583
- gone)
584
- # ONE SAMPLE IS ENOUGH FOR `gone`, and only for `gone`. A dead pid is not
585
- # a transient reading the way a frozen CPU clock is — a process does not
586
- # come back. Requiring two passes here would delay the one finding that is
587
- # already certain by a whole interval, for no gain in confidence.
588
- finding='gone'
589
- evidence="the agent pid $(monitor_pid) names no live process; the worker's desk is unattended"
590
- ;;
591
- quiet)
592
- # THE TWO-SAMPLE RULE, and the two extra conditions with it. All four must
593
- # hold together: this pass quiet, the PREVIOUS pass quiet, the tree
594
- # unchanged between them, and commits already on the branch.
595
- if [ "$prev_verdict" = 'quiet' ] && [ "$tree" = "$prev_tree" ]; then
596
- local has rc2
597
- monitor_has_commits; rc2=$?
598
- if [ "$rc2" = 0 ]; then
599
- finding='idle'
600
- evidence="the agent pid $(monitor_pid) is alive but its transcript has been silent for over ${PLOT_MONITOR_QUIET_SECONDS}s with no child process burning CPU behind it, across two consecutive passes ~${interval}s apart, the tree is unchanged between them, and the branch already carries commits"
601
- fi
602
- # rc2 = 1 → no commits yet: the middle row. It may be thinking, and
603
- # calling that a stall is what teaches an operator to ignore the word.
604
- # rc2 = 2 → unanswerable: no ref to count against, so no finding. A
605
- # failure to observe is not evidence of something to see.
606
- fi
607
- ;;
608
- # `busy`, `unknown` and `unspoken` are not findings. Nothing is published,
609
- # which is the design: silence means healthy, and the AgentMonitor's slower
610
- # loop is what catches a worker that finished without saying so. `unspoken`
611
- # is recorded as `prev_verdict`, so `idle` needs two `quiet` passes after
612
- # the conversation's first line. No grace period bounds it: a prompt that
613
- # stays alive and never writes a line ends at `Worker bound`, the cost
614
- # `unknown` already carries.
615
- esac
616
-
617
- prev_verdict="$verdict"
618
- prev_tree="$tree"
619
-
620
- # PUBLISH ONLY ON A CHANGE — the plan's "it publishes the moment a finding
621
- # holds and publishes nothing when nothing changed". A monitor that
622
- # re-published `idle` every 30 seconds would fill the findings file with one
623
- # fact repeated, and a subscriber could not tell a NEW stall from an old one.
624
- #
625
- # The clearing case is a publish too: a finding that held and then stopped
626
- # holding is news, and a board that never hears it leaves a stale entry up
627
- # after the worker recovered.
628
- if [ "$finding" != "$published" ]; then
629
- if [ -n "$finding" ]; then
630
- since=$(date -u +%Y-%m-%dT%H:%M:%SZ)
631
- publish "$finding" "$evidence" "$since"
632
- elif [ -n "$published" ]; then
633
- since=$(date -u +%Y-%m-%dT%H:%M:%SZ)
634
- publish 'clear' "the ${published} finding no longer holds; the worker is measuring healthy again" "$since"
635
- fi
636
- published="$finding"
637
- fi
638
- }
639
-
640
- # SOURCEABLE FOR TESTS. A test that wants to drive `monitor_pass` against
641
- # redefined ports needs the functions without the loop; everything above this
642
- # line defines, and nothing below it runs when the guard is set.
643
- [ -n "${PLOT_MONITOR_NO_MAIN:-}" ] && return 0 2>/dev/null
644
-
645
- # ONE LINE AT START WHEN THERE IS NO HANDLE. Every wrapper-started monitor has
646
- # one, because `plot-dispatch.sh` sets `PLOT_SESSION_ID` on every launch; only a
647
- # monitor started by hand has none. It then behaves as it did before the
648
- # conversation probe, and says so rather than degrading silently.
649
- if [ -z "${PLOT_SESSION_ID:-}" ] && [ -z "${PLOT_MANIFEST_FILE:-}" ]; then
650
- echo 'plot-worker-monitor: no session handle (PLOT_SESSION_ID and PLOT_MANIFEST_FILE unset) — idle is judged on the desk alone' >&2
651
- fi
652
-
653
- monitor_pass
654
- [ "$once" = 1 ] && exit 0
655
-
656
- # THE LOOP IS WHERE THE COMPARISON LIVES. `idle` needs two readings, so a
657
- # monitor that ran once and exited could never report it — which is why `--once`
658
- # is a test affordance and not a mode anyone dispatches.
659
- #
660
- # SILENCE IS MEANINGFUL HERE, and it is the opposite of what the no-op slice
661
- # needed. That monitor published every pass so that an attached-but-blind
662
- # monitor could not be mistaken for a watching one; this one publishes only on a
663
- # change, because it HAS something to say and saying it repeatedly would bury
664
- # it. Telling a healthy silence from a dead monitor is the channel's job —
665
- # `feature/the-channel-carries-the-findings`, whose heartbeat is exactly that
666
- # distinction.
667
- #
668
- # AND IT ENDS WITH ITS AGENT. Until 2026-08-30 it did not, and the estate showed
669
- # it: 34 of 40 monitors on this machine were `ppid=1`, and the orphans cost half
670
- # the machine's spawn cost (23.3 ms per 100 forks against 4.8 ms quiet). The
671
- # wrapper `wait`s on the agent alone — correctly, since waiting on two infinite
672
- # loops would hang and `.plot-worker.exit` would never be written — so when the
673
- # wrapper exits, its monitors are re-parented to `init` and loop forever.
674
- # `docs/research/2026-08-30-what-ends-a-monitor.md` has the measurement and the
675
- # commands that show it, on both the ordinary path and the `Worker bound` one.
676
- #
677
- # PUBLISH FIRST, THEN ASK — the order is the lower bound, and this monitor is
678
- # exactly where it matters. `gone` is one of its two findings, so a monitor that
679
- # checked the subject BEFORE its pass would exit on a dead agent without ever
680
- # reporting the death — the loudest finding it has, lost to the mechanism meant
681
- # to bound it. `plot_monitor_wait` returns only after `monitor_pass` has run.
682
- #
683
- # IT IS A MEASUREMENT, NOT A TIMER, which the plan requires in as many words: a
684
- # monitor exiting after N seconds regardless would pass every visible assertion
685
- # and destroy the property the design rests on. This reads the process table —
686
- # the same source the `gone` finding above reads, asked for a different purpose.
687
- while plot_monitor_wait "$interval" "$pid_file"; do
688
- monitor_pass
689
- done
690
-
691
- # THE FINAL PASS, and for this monitor it is not a courtesy — it is the `gone`
692
- # finding itself.
693
- #
694
- # `plot_monitor_wait` returns non-zero the moment the agent's pid names no live
695
- # process, so control arrives here with the subject already dead and NOTHING yet
696
- # published about it. One more pass runs, `monitor_pass` measures the same dead
697
- # pid the wait just saw, and `gone` is published on the way out.
698
- #
699
- # Without this line the monitor would exit silently on exactly the event it
700
- # exists to report — the upper bound eating the finding rather than the lower
701
- # bound. It would still pass "no monitor remains", which is why the suite
702
- # asserts the last finding's `measuredAt` against the exit file rather than
703
- # asserting the exit alone.
704
- monitor_pass
705
- exit 0