@plot-pm/board 0.16.3 → 0.18.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1,480 +0,0 @@
1
- #!/usr/bin/env bash
2
- # Plot helper: the BuildMonitor — watches the RUN, and what it concluded about a sha.
3
- #
4
- # RUN, NOT SOURCED, and started by `start_worker()` in `plot-dispatch.sh` as a
5
- # child of the wrapper, beside the WorkerMonitor and the AgentMonitor.
6
- #
7
- # ═══════════════════════════════════════════════════════════════════════════
8
- # THREE MONITORS, BECAUSE THERE ARE THREE SUBJECTS
9
- # ═══════════════════════════════════════════════════════════════════════════
10
- #
11
- # Two monitors became three, and the reason is the one that split the first two:
12
- # a different subject on a different cadence.
13
- #
14
- # | monitor | subject | samples | asks |
15
- # |-------------------|-------------|------------------------------|----------------------------|
16
- # | **WorkerMonitor** | the process | ~30 s | is it doing anything? |
17
- # | **AgentMonitor** | the desk | ~5 min | what does this agent owe? |
18
- # | **BuildMonitor** | the run | ~30 s **while a run is live**| did the build change? |
19
- #
20
- # A Build is already an entity in the spec, with its own identity and its own
21
- # state: DESIGN-build.md — *"is the thing that RUNS … one RESULT of one run"*,
22
- # identified by its URL, holding a state, a start time and a duration. **A
23
- # monitor per entity is the pattern, not an exception to it.**
24
- #
25
- # ═══════════════════════════════════════════════════════════════════════════
26
- # FOUR FINDINGS, AND `head moved` IS THE ONE THAT EARNS THIS MONITOR
27
- # ═══════════════════════════════════════════════════════════════════════════
28
- #
29
- # | finding | measurement |
30
- # |--------------------------|----------------------------------------------------|
31
- # | **build failed** | a run for this branch's head reached a failing conclusion |
32
- # | **build passed** | it reached success |
33
- # | **build needs approval** | it is `action_required` |
34
- # | **head moved** | a newer sha exists, so the run in flight answers about the past |
35
- #
36
- # **`head moved` is why this cannot live in the AgentMonitor.** A build's
37
- # subject is a SHA, not a branch. A green result for code nobody will merge is
38
- # worse than no result — it invites a merge of the wrong thing. Measured
39
- # 2026-08-30: two merge waiters reported on superseded runs and had to be
40
- # stopped and re-armed.
41
- #
42
- # **`action_required` is a real state here, not an edge case.** Bot branches hit
43
- # it — the release PR's runs need manual approval before they start. A monitor
44
- # that folded it into "not passed yet" would report a build as pending forever
45
- # while it waits for a click nobody knows is needed.
46
- #
47
- # ═══════════════════════════════════════════════════════════════════════════
48
- # IT POLLS NOTHING WHEN NO RUN IS LIVE
49
- # ═══════════════════════════════════════════════════════════════════════════
50
- #
51
- # That is what makes a 30-second cadence against a host affordable, and it is
52
- # the property that separates this monitor's budget from the AgentMonitor's.
53
- # The AgentMonitor's five-minute budget exists because it asks ON EVERY PASS;
54
- # this one asks only while there is something to ask about.
55
- #
56
- # STRUCTURALLY, NOT INCIDENTALLY: `monitor_head_sha` is a LOCAL git read and it
57
- # gates the host call. `sample_finding` returns before `monitor_run_for_sha` is
58
- # ever reached whenever there is no head to ask about, and once a sha has
59
- # reached a terminal conclusion it is never asked about again. A monitor that
60
- # kept questioning an idle host is the rate problem this whole design avoids,
61
- # so the silence is asserted by the tests rather than assumed.
62
- #
63
- # ═══════════════════════════════════════════════════════════════════════════
64
- # THE FINDINGS ARE TRANSITIONS, NOT CONDITIONS
65
- # ═══════════════════════════════════════════════════════════════════════════
66
- #
67
- # The other monitors report states that PERSIST — `owes a review` holds until a
68
- # PR exists, and republishing it would be repeating one fact. A build's answer
69
- # CHANGES ONCE AND STAYS: a run that failed has failed, and it will still have
70
- # failed in thirty seconds.
71
- #
72
- # So the state this monitor carries is keyed by SHA as well as by finding.
73
- # Publishing `build passed` for one sha does not suppress `build passed` for the
74
- # next one — that would silence the answer an operator is actually waiting for,
75
- # on the very push they pushed to get it. And once a sha's build is terminal,
76
- # the sha is not asked about again: the answer cannot change, so continuing to
77
- # poll would be spending a host round trip to re-learn a fact already published.
78
- #
79
- # ═══════════════════════════════════════════════════════════════════════════
80
- # IT OBSERVES; IT DOES NOT ACT
81
- # ═══════════════════════════════════════════════════════════════════════════
82
- #
83
- # It does not rerun a workflow, approve a run that is `action_required`, merge a
84
- # PR that went green, or push a fix for one that went red. Every one of those is
85
- # a judgement with a blast radius. Approving a run in particular is a human's
86
- # call by construction — `action_required` EXISTS because a person is meant to
87
- # look — and a monitor that clicked it would defeat the gate it is reporting.
88
- #
89
- # PUBLISHING IS ITS ONLY OUTPUT, as next door: no state file, no cache, nothing
90
- # written into the repository it watches, not even a record of what it last
91
- # published. The variables that make "publish on change" work live in memory and
92
- # die with the process, so a restarted monitor re-derives them one interval late
93
- # rather than reading a stale one.
94
- set -uo pipefail
95
-
96
- usage() {
97
- cat >&2 <<'EOF'
98
- Usage: plot-build-monitor.sh [--once]
99
-
100
- Started by plot-dispatch.sh inside the worker's wrapper. Reads its subject from
101
- the environment, exactly as the wrapper's other children do:
102
-
103
- PLOT_BRANCH the branch whose builds this monitor will report
104
- PLOT_WORKTREE the desk it STARTS on; it follows the manifest's
105
- `worktree` from then on (PLOT_MANIFEST_FILE below)
106
- PLOT_MANIFEST_FILE the manifest this agent is named in, re-read each pass;
107
- absent or unreadable keeps watching PLOT_WORKTREE
108
- PLOT_MONITOR_FILE where findings are published (default:
109
- $PLOT_WORKTREE/.plot-worker.monitor.build.jsonl). Set
110
- explicitly, it WINS and does not follow a hop.
111
- PLOT_MONITOR_INTERVAL seconds between passes (default 30)
112
-
113
- --once take one sample and exit, rather than looping. A test
114
- affordance: nothing dispatches this mode.
115
- EOF
116
- }
117
-
118
- once=0
119
- while [ $# -gt 0 ]; do
120
- case "$1" in
121
- --once) once=1 ;;
122
- -h|--help) usage; exit 0 ;;
123
- *) echo "plot-build-monitor: unknown argument '$1'" >&2; usage; exit 2 ;;
124
- esac
125
- shift
126
- done
127
-
128
- monitor='BuildMonitor'
129
-
130
- branch="${PLOT_BRANCH:-}"
131
- # The branch at start, kept for a detached or unreadable desk; `monitor_branch`
132
- # re-reads the desk on every pass.
133
- start_branch="$branch"
134
- # THE DESK THIS MONITOR WAS LAUNCHED ON, FIXED FOR ITS WHOLE LIFE. `worktree`
135
- # itself is no longer fixed: `monitor_pass` reassigns it every pass by asking
136
- # `plot_watched_desk`, which follows a hop to the manifest's new `worktree`
137
- # field. This is the launch desk `plot_watched_desk` falls back to when the
138
- # manifest carries no override — never read directly after startup.
139
- launched_worktree="${PLOT_WORKTREE:-}"
140
- worktree="$launched_worktree"
141
- manifest_file="${PLOT_MANIFEST_FILE:-}"
142
- # THIRTY SECONDS IS AFFORDABLE ONLY BECAUSE OF THE SILENCE RULE. This cadence
143
- # matches the WorkerMonitor's rather than the AgentMonitor's, and it asks a HOST
144
- # — which would be the rate problem the AgentMonitor's 300 s exists to avoid,
145
- # were it asking on every pass. It is not: no live run, no question. The budget
146
- # is bounded by how long a build takes, not by how long a worker lives.
147
- interval="${PLOT_MONITOR_INTERVAL:-30}"
148
-
149
- # IF `PLOT_MONITOR_FILE` IS SET IT WINS AND DOES NOT FOLLOW THE DESK — the
150
- # existing contract this usage text already states. Otherwise `findings`
151
- # follows `worktree`, reassigned alongside it at the top of `monitor_pass`.
152
- monitor_file_override="${PLOT_MONITOR_FILE:-}"
153
- findings="${monitor_file_override:-${worktree:+$worktree/.plot-worker.monitor.build.jsonl}}"
154
-
155
- # THE SUBJECT, read the same way the other two monitors read it.
156
- pid_file="${PLOT_PID_FILE:-${worktree:+$worktree/.plot-worker.pid}}"
157
-
158
- # ONE ANSWER TO "IS MY SUBJECT STILL THERE?", shared with both siblings, for the
159
- # reason `plot-monitor-subject.sh` documents: three monitors deciding
160
- # independently when to stop would drift, and the failure would be silent.
161
- # shellcheck source=./plot-monitor-subject.sh
162
- . "$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)/plot-monitor-subject.sh"
163
-
164
- # THE HOST ADAPTER, sourced for its path rather than its functions: the one
165
- # operation this monitor asks is `plot-host.sh run-for-sha`, and `plot-host.sh`
166
- # is the ONE place that talks to the host CLI. A monitor calling `gh` directly
167
- # would be a second adapter, and the backend split (github/bitbucket) would have
168
- # to be decided twice.
169
- host_script="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)/plot-host.sh"
170
-
171
- json_escape() { # $1 = raw → prints a JSON-safe string body
172
- printf '%s' "$1" | python3 -c 'import json,sys; sys.stdout.write(json.dumps(sys.stdin.read())[1:-1])' 2>/dev/null \
173
- || printf '%s' "$1" | sed 's/\\/\\\\/g; s/"/\\"/g'
174
- }
175
-
176
- # A FINDING CARRIES THE SAME FOUR FIELDS the other two monitors publish —
177
- # `finding`, `since`, `evidence`, `measuredAt` — for the same reason: one
178
- # subscriber will read all three files and must not need a third parser to do
179
- # it.
180
- #
181
- # `since` HERE IS WHEN THE BUILD REACHED THIS ANSWER, as far as this monitor
182
- # can tell — the pass that first saw it. On a 30 s cadence the gap to
183
- # `measuredAt` is small by construction, which is the opposite of the
184
- # AgentMonitor's case and follows from the same field meaning the same thing.
185
- publish() { # $1=finding $2=evidence $3=since
186
- local now
187
- now=$(date -u +%Y-%m-%dT%H:%M:%SZ)
188
- local line
189
- line=$(printf '{"monitor":"%s","branch":"%s","worktree":"%s","finding":"%s","since":"%s","evidence":"%s","measuredAt":"%s"}' \
190
- "$monitor" \
191
- "$(json_escape "$branch")" \
192
- "$(json_escape "$worktree")" \
193
- "$(json_escape "$1")" \
194
- "${3:-$now}" \
195
- "$(json_escape "$2")" \
196
- "$now")
197
- # Both destinations, for the reason the siblings give: the file is what a
198
- # subscriber reads, and stdout lands in `.plot-worker.log` beside the agent's
199
- # own output where an operator tailing a worker sees it.
200
- [ -n "$findings" ] && printf '%s\n' "$line" >> "$findings" 2>/dev/null
201
- printf 'plot-monitor %s\n' "$line"
202
- }
203
-
204
- # ---------------------------------------------------------------------------
205
- # THE PORTS — two named seams, so every branch is reachable from a test
206
- # ---------------------------------------------------------------------------
207
- #
208
- # Same convention as the four next door and the five beyond it: a test sources
209
- # this file with `PLOT_MONITOR_NO_MAIN=1` and REDEFINES them. Here there are
210
- # only two, and one of them is the host round trip — which is exactly the seam a
211
- # test can least afford to drive for real. You cannot ask GitHub for a run that
212
- # is `action_required` on demand, you cannot make a run vanish to order, and you
213
- # certainly cannot arrange two runs for two shas at the instant a test needs
214
- # them. Every one of those is a branch this monitor must get right.
215
-
216
- # What sha is this branch's head, right now?
217
- #
218
- # → prints the sha, or nothing when it cannot be read
219
- #
220
- # LOCAL, AND IT GATES THE HOST CALL. This is the cheap reading that makes the
221
- # silence rule structural: no head, no question. It reads the WORKTREE's HEAD
222
- # rather than a remote ref, because the head a build should be about is the one
223
- # the agent has actually produced — and because a fetch per pass would be a
224
- # second network call on a 30-second loop.
225
- monitor_head_sha() { # → prints a sha, or nothing
226
- [ -n "$worktree" ] && [ -d "$worktree" ] || return 0
227
- git -C "$worktree" rev-parse --verify --quiet HEAD 2>/dev/null || true
228
- }
229
-
230
- # Which branch does the desk hold, right now?
231
- #
232
- # → prints the branch, or the branch the monitor started with when the desk is
233
- # detached or unreadable
234
- #
235
- # READ ON EVERY PASS, NOT ONCE. A free agent starts with an empty `PLOT_BRANCH`
236
- # and is handed its slices later, in the same desk; a dispatched agent hops to
237
- # its next slice in the same desk too. A branch fixed at start made this monitor
238
- # unaskable for every free agent (`monitor_run_for_sha` returns 2 on an empty
239
- # branch) and wrong after every hop, so the correction path in
240
- # `plot-worker-loop.sh` received no finding for most slices.
241
- monitor_branch() { # → prints a branch, or nothing
242
- local current=''
243
- if [ -n "$worktree" ] && [ -d "$worktree" ]; then
244
- current=$(git -C "$worktree" branch --show-current 2>/dev/null || true)
245
- fi
246
- printf '%s' "${current:-$start_branch}"
247
- }
248
-
249
- # What does the host say about the run for ONE sha?
250
- #
251
- # → prints the run JSON, or nothing when there is no run for it
252
- # → returns 2 when the host could not be asked at all
253
- #
254
- # THE DISTINCTION BETWEEN "NO RUN" AND "COULD NOT ASK" IS THE POINT, and it is
255
- # the same discipline `monitor_pr_state` keeps next door. An unreachable host is
256
- # NOT evidence that a build is absent. A monitor that read a `gh` failure as "no
257
- # run" would go silent about every build on the estate the moment a token
258
- # expired — and silence is this monitor's healthy signal, so the failure would
259
- # be invisible by construction.
260
- #
261
- # An EMPTY result with a reachable host is a real and common answer: the run has
262
- # not been created yet. A monitor polling a fresh push sees exactly this until
263
- # CI wakes up, and it is not a finding.
264
- monitor_run_for_sha() { # $1 = sha → prints run JSON | rc 2 = unaskable
265
- [ -n "$branch" ] || return 2
266
- [ -x "$host_script" ] || return 2
267
- local out
268
- out=$("$host_script" run-for-sha "$branch" "$1" 2>/dev/null) || return 2
269
- printf '%s' "$out"
270
- return 0
271
- }
272
-
273
- # ---------------------------------------------------------------------------
274
- # READING ONE FIELD OUT OF THE RUN
275
- # ---------------------------------------------------------------------------
276
- #
277
- # `jq` where it exists and `sed` where it does not, the same fallback shape
278
- # `json_escape` uses above. A monitor that died because `jq` was missing would
279
- # be a monitor that reported nothing on exactly the machines least likely to
280
- # have anyone watching.
281
- run_field() { # $1 = json, $2 = key → prints the value, or nothing for null
282
- local v
283
- if command -v jq >/dev/null 2>&1; then
284
- v=$(printf '%s' "$1" | jq -r --arg k "$2" '.[$k] // empty' 2>/dev/null)
285
- else
286
- v=$(printf '%s' "$1" | sed -n 's/.*"'"$2"'":"\([^"]*\)".*/\1/p')
287
- fi
288
- printf '%s' "$v"
289
- }
290
-
291
- # ---------------------------------------------------------------------------
292
- # THE SAMPLER — one pass, using only the ports above
293
- # ---------------------------------------------------------------------------
294
- #
295
- # THE STATE IS TWO VARIABLES AND IT IS DERIVED, as next door — but `published`
296
- # is keyed by SHA here, because the findings are transitions rather than
297
- # conditions. `published_sha` records which commit the standing answer is about,
298
- # so the same word about a different commit is still news.
299
- published=''
300
- published_sha=''
301
- since=''
302
- # The shas whose builds have reached a terminal answer. Once a run has failed,
303
- # passed, or been superseded, asking again spends a host round trip to re-learn
304
- # a fact already published — so it is not asked. THIS is the second half of "it
305
- # polls nothing when no run is live": the first half is having no head at all,
306
- # and this is having no OPEN question about the head there is.
307
- settled_shas=''
308
-
309
- sha_is_settled() { # $1 = sha → 0 settled | 1 not
310
- case " $settled_shas " in *" $1 "*) return 0 ;; esac
311
- return 1
312
- }
313
-
314
- # ---------------------------------------------------------------------------
315
- # ONE FINDING PER PASS, AND `head moved` COMES FIRST
316
- # ---------------------------------------------------------------------------
317
- #
318
- # Unlike the AgentMonitor's four, these are near-exclusive by construction — a
319
- # run has one status. The one genuine overlap is the one that matters: a run
320
- # that concluded `success` for a sha the branch has since moved past is BOTH
321
- # "passed" and "superseded", and reporting it as passed is precisely the failure
322
- # this monitor exists to prevent.
323
- #
324
- # So `head moved` is decided FIRST and about the run's own sha, not about the
325
- # branch: if the answer in hand describes a commit that is no longer the head,
326
- # the answer is about the past whatever it says.
327
- sample_finding() { # → prints "finding\tevidence", or nothing
328
- local head
329
- head=$(monitor_head_sha)
330
-
331
- # NO HEAD, NO QUESTION. The cheap local reading refuses before anything
332
- # reaches the host — the structural form of "it polls nothing when no run is
333
- # live". A worktree that is gone, or a branch with no commit yet, asks
334
- # nothing at all.
335
- [ -n "$head" ] || return 0
336
-
337
- # ALREADY ANSWERED. The head's build reached a terminal conclusion on an
338
- # earlier pass and a build's answer does not change back. Asking again would
339
- # be the idle polling this design refuses.
340
- sha_is_settled "$head" && return 0
341
-
342
- local run rc
343
- run=$(monitor_run_for_sha "$head"); rc=$?
344
-
345
- # UNASKABLE — no finding. The host could not be asked, so nothing about the
346
- # build is known. Reporting anything here would be inventing an answer out of
347
- # a failure to observe.
348
- [ "$rc" = 2 ] && return 0
349
-
350
- # NO RUN YET — no finding, and not an error. The commit exists and CI has not
351
- # created a run for it. This is the ordinary state of a freshly pushed sha,
352
- # and it is what the monitor sees on every pass until the run appears.
353
- [ -n "$run" ] || return 0
354
-
355
- local run_sha status conclusion url
356
- run_sha=$(run_field "$run" sha)
357
- status=$(run_field "$run" status)
358
- conclusion=$(run_field "$run" conclusion)
359
- url=$(run_field "$run" url)
360
-
361
- # 1. HEAD MOVED — the run in hand is about a commit that is no longer the
362
- # head. Decided BEFORE the conclusion is read, because a green run for
363
- # superseded code is the specific wrong answer this monitor was built to
364
- # avoid: it invites a merge of the wrong thing. Measured 2026-08-30 — two
365
- # merge waiters reported on superseded runs and had to be stopped and
366
- # re-armed.
367
- #
368
- # This fires when the host answered about a DIFFERENT sha than the one asked
369
- # about, which is the shape a race actually takes: the head moved between the
370
- # local read and the host's reply.
371
- if [ -n "$run_sha" ] && [ "$run_sha" != "$head" ]; then
372
- printf 'head moved\tthe run at %s is for %s, but the branch head is now %s; its answer is about the past\n' \
373
- "${url:-an unknown url}" "$run_sha" "$head"
374
- return 0
375
- fi
376
-
377
- # 2. BUILD NEEDS APPROVAL — a real state, not an edge case. Bot branches hit
378
- # it: the release PR's runs need a manual click before they start. GitHub
379
- # reports it as a `status` of `waiting`/`action_required` and as a
380
- # `conclusion` of `action_required`, depending on where the run is, so both
381
- # are read. Folding it into "not passed yet" would report the build pending
382
- # forever while it waits for a click nobody knows is needed.
383
- case "$status:$conclusion" in
384
- *action_required*|waiting:*)
385
- printf 'build needs approval\tthe run at %s for %s is waiting for a manual approval before it can start\n' \
386
- "${url:-an unknown url}" "$head"
387
- return 0
388
- ;;
389
- esac
390
-
391
- # 3 & 4. THE TERMINAL CONCLUSIONS. An empty conclusion means the run is still
392
- # going — queued or in progress — and a monitor whose subject is a transition
393
- # says nothing about a state that has not changed yet.
394
- [ -n "$conclusion" ] || return 0
395
-
396
- case "$conclusion" in
397
- success)
398
- printf 'build passed\tthe run at %s for %s concluded success\n' "${url:-an unknown url}" "$head"
399
- return 0
400
- ;;
401
- # EVERY OTHER TERMINAL CONCLUSION IS A FAILURE TO A READER WAITING ON GREEN.
402
- # `failure`, `timed_out`, `cancelled` and `startup_failure` differ in cause
403
- # and not in consequence: none of them is a build somebody may merge on. The
404
- # cause is not thrown away — it rides in the evidence, where a reader
405
- # deciding whether to rerun can see it.
406
- *)
407
- printf 'build failed\tthe run at %s for %s concluded %s\n' "${url:-an unknown url}" "$head" "$conclusion"
408
- return 0
409
- ;;
410
- esac
411
- }
412
-
413
- # One full pass: sample, publish only on a change of ANSWER-ABOUT-A-COMMIT.
414
- monitor_pass() {
415
- local row finding evidence head
416
- # REASSIGNED ONCE, HERE, RATHER THAN THREADED THROUGH EACH READER. `worktree`
417
- # feeds `monitor_head_sha`, `monitor_branch` and `publish()`'s own field, so a
418
- # hop is picked up in one place rather than three. `PLOT_MONITOR_FILE`, when
419
- # set, WINS and does not follow the desk — the existing contract the usage
420
- # text states.
421
- worktree=$(plot_watched_desk "$manifest_file" "$launched_worktree")
422
- findings="${monitor_file_override:-${worktree:+$worktree/.plot-worker.monitor.build.jsonl}}"
423
- branch=$(monitor_branch)
424
- head=$(monitor_head_sha)
425
- row=$(sample_finding)
426
- finding="${row%% *}"
427
- evidence=''
428
- case "$row" in *" "*) evidence="${row#* }" ;; esac
429
- [ -z "$row" ] && finding=''
430
-
431
- # PUBLISH ON A CHANGE OF EITHER THE FINDING OR THE COMMIT IT IS ABOUT. The
432
- # second half is what makes these transitions rather than conditions: `build
433
- # passed` for a new sha is news even though the word is the same as last
434
- # time, and suppressing it would silence exactly the answer an operator
435
- # pushed in order to get.
436
- if [ "$finding" != "$published" ] || { [ -n "$finding" ] && [ "$head" != "$published_sha" ]; }; then
437
- if [ -n "$finding" ]; then
438
- since=$(date -u +%Y-%m-%dT%H:%M:%SZ)
439
- publish "$finding" "$evidence" "$since"
440
- # SETTLED, so it is never asked about again. Every finding this monitor
441
- # publishes is terminal for its sha: a failure stays failed, a pass stays
442
- # passed, and a superseded run does not become current. Recording it here
443
- # rather than in `sample_finding` keeps the decision beside the publish it
444
- # follows from.
445
- [ -n "$head" ] && settled_shas="$settled_shas $head"
446
- fi
447
- # NO CLEARING PUBLISH, and that is the difference from the AgentMonitor. A
448
- # debt is cleared when it is paid — that is news. A build's answer is never
449
- # withdrawn: `build failed` does not stop being true about that sha, and the
450
- # next answer is a new finding about a new commit, which the sha key already
451
- # carries. Publishing `clear` here would tell a subscriber a failure had
452
- # been resolved when all that happened is the branch moved on.
453
- published="$finding"
454
- published_sha="$head"
455
- fi
456
- }
457
-
458
- # SOURCEABLE FOR TESTS, the same guard the siblings carry. A test that wants to
459
- # drive `monitor_pass` against redefined ports needs the functions without the
460
- # loop; everything above this line defines, and nothing below it runs when the
461
- # guard is set.
462
- [ -n "${PLOT_MONITOR_NO_MAIN:-}" ] && return 0 2>/dev/null
463
-
464
- monitor_pass
465
- [ "$once" = 1 ] && exit 0
466
-
467
- # IT ENDS WITH ITS AGENT, by the mechanism `plot-monitor-subject.sh` documents
468
- # and for the reason `docs/research/2026-08-30-what-ends-a-monitor.md` measured.
469
- #
470
- # PUBLISH FIRST, THEN LEAVE. The final pass runs with the agent already gone,
471
- # and it matters here for a reason of its own: an agent that pushes and exits
472
- # leaves a run still going, and the answer arrives after there is nobody left to
473
- # see it. The last pass is the one chance to catch a build that concluded during
474
- # the shutdown.
475
- while plot_monitor_wait "$interval" "$pid_file"; do
476
- monitor_pass
477
- done
478
-
479
- monitor_pass
480
- exit 0
@@ -1,172 +0,0 @@
1
- #!/usr/bin/env bash
2
- # The ONE answer to "how long has this worktree's agent been quiet?" — sourced,
3
- # not run, by `plot-worker-loop.sh` for its own watcher (`plot-worker-state.sh`'s
4
- # `plot_worker_idle_watch_pass`) and for its own ending message. Sourced by the
5
- # now-deleted `plot-worker-monitor.sh` until `bug/the-loop-reports-idle`.
6
- #
7
- # It reads the AGENT rather than the machine. A `claude -p` session appends a
8
- # timestamped line to its transcript for every model turn, tool call and tool
9
- # result; the seconds since the newest of those lines is how long the agent has
10
- # produced nothing. A CPU sample answers *is this process on a core right now?*,
11
- # which is a different question and is zero for most of a working agent's life.
12
- #
13
- # ═══════════════════════════════════════════════════════════════════════════
14
- # THREE ANSWERS, AND THE THIRD IS NOT A FAILURE
15
- # ═══════════════════════════════════════════════════════════════════════════
16
- #
17
- # <seconds> the newest transcript line is that many seconds old
18
- # unavailable no transcript can be read for this worktree, INCLUDING the
19
- # case where the arguments name no worktree at all
20
- #
21
- # `unavailable` is a first-class answer, settled by
22
- # `the-registry-supervises-its-agents`: a capability the adopting project does
23
- # not provide is UNAVAILABLE, never failed and never zero. A caller that read
24
- # it as "quiet for 0 seconds" would report every unreadable agent healthy; one
25
- # that read it as an error would refuse to run where Plot's own contract says
26
- # it should degrade. So it is a word, and callers match on it.
27
- #
28
- # ═══════════════════════════════════════════════════════════════════════════
29
- # THE TRANSCRIPT IS FOUND BY PATH, NOT BY SESSION ID — AND THAT IS A CHOICE
30
- # ═══════════════════════════════════════════════════════════════════════════
31
- #
32
- # `.plot/worker-prompt.sh:29` DOES pass `--session-id` now (2026-09-04), so an
33
- # exact join is available here and this deliberately does not take it. The
34
- # reason is the question, not the capability: this asks *is anything happening
35
- # at this desk*, and an operator's own session at the same worktree is a true
36
- # answer to it. Joining on the worker's id alone would report a desk quiet
37
- # while somebody is visibly working at it.
38
- #
39
- # THE OPPOSITE JOIN IS RIGHT FOR THE OPPOSITE QUESTION. *What has THIS agent
40
- # spent* is per session and never per worktree — one project directory measured
41
- # 2026-09-03 held 45 session files, 30 of them subagents, and a sum across them
42
- # belongs to no one. `rules/spend.ts` states that side; the two read the same
43
- # files and must not be made to share a join.
44
- #
45
- # A THIRD QUESTION USES BOTH KEYS: *has this worker's conversation written?* It
46
- # is per worktree AND per conversation handle, and it is asked as a file's
47
- # existence, never as a time: `plot_transcript_exists` below. The loop asks it
48
- # to choose `--session-id` or `--resume`; the worker monitor asks it before it
49
- # calls a quiet desk idle, because until the new conversation writes its first
50
- # line the desk's newest file belongs to the previous one. It does not change
51
- # the quiet number, which stays about the desk.
52
- #
53
- # So the join here is the one `plot-quiet-stretch.mjs` already made and proved
54
- # on 23 real sessions: the runtime stores a session under
55
- # `$HOME/.claude/projects/<slug>` where the slug is the WORKTREE PATH with `/`
56
- # and `.` replaced by `-`. A dispatched worker has its own worktree, so the
57
- # path identifies the DESK without any id being passed anywhere.
58
- #
59
- # `agent-` PREFIXED FILES ARE SKIPPED, for wave 1's reason: a subagent's
60
- # transcript is a true statement about the wrong process. A worker whose
61
- # subagent is chatting while the worker itself has stopped must read as quiet.
62
- #
63
- # ═══════════════════════════════════════════════════════════════════════════
64
- # THE NEWEST LINE ACROSS ALL OF A WORKTREE'S SESSIONS
65
- # ═══════════════════════════════════════════════════════════════════════════
66
- #
67
- # A worktree can hold several sessions: a worker that hopped waves, or an
68
- # operator who opened a session at the same desk. Any of them producing output
69
- # means SOMEBODY is working there, and the monitor's question is about the desk.
70
- # Taking the maximum timestamp is the answer that never ends a live session
71
- # because a stale sibling exists beside it.
72
-
73
- # The runtime's project-slug derivation. Duplicated from `plot-quiet-stretch.mjs`
74
- # — which duplicates it from the board's `projectSlug` — because a monitor on a
75
- # 30s loop must not depend on a build step. One line, pinned by a test on each
76
- # side.
77
- plot_transcript_slug() { # $1=worktree path → slug
78
- printf '%s' "$1" | tr '/.' '--'
79
- }
80
-
81
- # Where the runtime keeps this worktree's transcripts, or "" if nowhere.
82
- plot_transcript_dir() { # $1=worktree → directory path (may not exist)
83
- local wt="$1" home="${PLOT_TRANSCRIPT_HOME:-${HOME:-}}"
84
- [ -n "$wt" ] && [ -n "$home" ] || return 0
85
- printf '%s/.claude/projects/%s' "$home" "$(plot_transcript_slug "$wt")"
86
- }
87
-
88
- # Seconds since this worktree's agent last wrote anything.
89
- #
90
- # THE MTIME IS THE READING, not the file's contents. The runtime appends as it
91
- # works, so the file's modification time IS the timestamp of its newest line —
92
- # and reading it costs one `stat` rather than parsing a transcript that reaches
93
- # tens of megabytes on a long session. Wave 1 parsed timestamps because it was
94
- # measuring a DISTRIBUTION of past gaps; this needs only the newest, on a 30s
95
- # loop, for as many workers as the machine holds.
96
- #
97
- # A CLOCK IS THE RIGHT INSTRUMENT HERE, and that is worth stating because Plot
98
- # usually refuses one. `plot-estate-changed.sh` hashes content rather than
99
- # reading mtime, because its question is *did this change?* and a checkout moves
100
- # mtime without changing anything. This question is *how long since output?* —
101
- # which is a question about elapsed time, and mtime is the measurement of it.
102
- plot_transcript_quiet_seconds() { # $1=worktree → seconds | unavailable
103
- local wt="$1" dir newest now mtime
104
- dir=$(plot_transcript_dir "$wt")
105
- [ -n "$dir" ] && [ -d "$dir" ] || { printf 'unavailable'; return 0; }
106
-
107
- # The newest mtime across every non-`agent-` session file in the directory.
108
- # `find -print0` and a while-read keep paths with spaces intact; `stat` is
109
- # asked once per file, and a worktree holds one to eight.
110
- newest=''
111
- while IFS= read -r -d '' f; do
112
- case "$(basename "$f")" in agent-*) continue ;; esac
113
- mtime=$(plot_transcript_mtime "$f") || continue
114
- [ -n "$mtime" ] || continue
115
- if [ -z "$newest" ] || [ "$mtime" -gt "$newest" ] 2>/dev/null; then newest="$mtime"; fi
116
- done < <(find "$dir" -maxdepth 1 -type f -name '*.jsonl' -print0 2>/dev/null)
117
-
118
- # A directory that exists but holds no session file is still UNAVAILABLE. The
119
- # runtime creates the directory when the project is first opened, so an empty
120
- # one means no session has written here — which is exactly "no transcript can
121
- # be read", not "quiet for a very long time".
122
- [ -n "$newest" ] || { printf 'unavailable'; return 0; }
123
-
124
- now=$(date +%s)
125
- local quiet=$(( now - newest ))
126
- # A transcript written in the future — a clock skew across a mounted volume —
127
- # reads as zero rather than negative. An end condition comparing a negative
128
- # against a threshold would behave correctly by accident here and not
129
- # elsewhere; clamping says what is meant.
130
- [ "$quiet" -lt 0 ] && quiet=0
131
- printf '%s' "$quiet"
132
- }
133
-
134
- # Does the conversation `id` have a transcript at this worktree?
135
- #
136
- # EXISTENCE, NOT A TIMESTAMP. The runtime creates `<id>.jsonl` with its first
137
- # line and appends to it after, under both `--session-id` and `--resume`. So the
138
- # file's presence says the conversation has written, and no comparison of
139
- # clocks is made, so a file created in the same second as a manifest write
140
- # reads as present.
141
- #
142
- # NO HANDLE AND NO FILE ARE ONE ANSWER HERE, and a caller that must tell them
143
- # apart checks the handle first. `session_flag` reads both as *create*; the
144
- # monitor's port does not, because a monitor with no handle must not read every
145
- # quiet worker as unspoken.
146
- plot_transcript_exists() { # $1=worktree $2=id → 0 found | 1 not
147
- local wt="$1" id="$2" dir
148
- [ -n "$wt" ] && [ -n "$id" ] || return 1
149
- dir=$(plot_transcript_dir "$wt" 2>/dev/null) || return 1
150
- [ -n "$dir" ] || return 1
151
- [ -f "$dir/$id.jsonl" ]
152
- }
153
-
154
- # A file's modification time as a unix epoch. BSD and GNU `stat` disagree on the
155
- # flag, and a monitor that works on the author's laptop and not in CI is a
156
- # monitor nobody trusts.
157
- #
158
- # THE `||` IS NOT ENOUGH, AND CI MEASURED WHY. On Linux `stat -f` is not an
159
- # unknown flag — it means FILESYSTEM info, and it SUCCEEDS. So the BSD form
160
- # never falls through: it printed `Namelen: 255 Type: ext2/ext3` and the
161
- # caller subtracted that from a clock. Two of this branch's own tests failed on
162
- # it, 2026-09-02, having passed on macOS.
163
- #
164
- # So the answer is validated rather than trusted. Each form must yield digits;
165
- # anything else is treated as that form not being available here.
166
- plot_transcript_mtime() { # $1=file → epoch seconds
167
- local m
168
- m=$(stat -c '%Y' "$1" 2>/dev/null)
169
- case "$m" in ''|*[!0-9]*) m=$(stat -f '%m' "$1" 2>/dev/null) ;; esac
170
- case "$m" in ''|*[!0-9]*) return 1 ;; esac
171
- printf '%s' "$m"
172
- }