@plot-pm/board 0.16.3 → 0.18.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/board-server.mjs +252 -153
- package/package.json +3 -4
- package/plot-agent-manifest.sh +9 -44
- package/plot-config.sh +16 -14
- package/plot-default-branch.sh +15 -19
- package/plot-dispatch.sh +29 -26
- package/plot-fleet-scan.sh +154 -141
- package/plot-host.sh +51 -59
- package/plot-monitor-subject.sh +7 -8
- package/plot-worker-state.sh +10 -163
- package/plot-build-monitor.sh +0 -480
- package/plot-transcript-quiet.sh +0 -172
package/plot-build-monitor.sh
DELETED
|
@@ -1,480 +0,0 @@
|
|
|
1
|
-
#!/usr/bin/env bash
|
|
2
|
-
# Plot helper: the BuildMonitor — watches the RUN, and what it concluded about a sha.
|
|
3
|
-
#
|
|
4
|
-
# RUN, NOT SOURCED, and started by `start_worker()` in `plot-dispatch.sh` as a
|
|
5
|
-
# child of the wrapper, beside the WorkerMonitor and the AgentMonitor.
|
|
6
|
-
#
|
|
7
|
-
# ═══════════════════════════════════════════════════════════════════════════
|
|
8
|
-
# THREE MONITORS, BECAUSE THERE ARE THREE SUBJECTS
|
|
9
|
-
# ═══════════════════════════════════════════════════════════════════════════
|
|
10
|
-
#
|
|
11
|
-
# Two monitors became three, and the reason is the one that split the first two:
|
|
12
|
-
# a different subject on a different cadence.
|
|
13
|
-
#
|
|
14
|
-
# | monitor | subject | samples | asks |
|
|
15
|
-
# |-------------------|-------------|------------------------------|----------------------------|
|
|
16
|
-
# | **WorkerMonitor** | the process | ~30 s | is it doing anything? |
|
|
17
|
-
# | **AgentMonitor** | the desk | ~5 min | what does this agent owe? |
|
|
18
|
-
# | **BuildMonitor** | the run | ~30 s **while a run is live**| did the build change? |
|
|
19
|
-
#
|
|
20
|
-
# A Build is already an entity in the spec, with its own identity and its own
|
|
21
|
-
# state: DESIGN-build.md — *"is the thing that RUNS … one RESULT of one run"*,
|
|
22
|
-
# identified by its URL, holding a state, a start time and a duration. **A
|
|
23
|
-
# monitor per entity is the pattern, not an exception to it.**
|
|
24
|
-
#
|
|
25
|
-
# ═══════════════════════════════════════════════════════════════════════════
|
|
26
|
-
# FOUR FINDINGS, AND `head moved` IS THE ONE THAT EARNS THIS MONITOR
|
|
27
|
-
# ═══════════════════════════════════════════════════════════════════════════
|
|
28
|
-
#
|
|
29
|
-
# | finding | measurement |
|
|
30
|
-
# |--------------------------|----------------------------------------------------|
|
|
31
|
-
# | **build failed** | a run for this branch's head reached a failing conclusion |
|
|
32
|
-
# | **build passed** | it reached success |
|
|
33
|
-
# | **build needs approval** | it is `action_required` |
|
|
34
|
-
# | **head moved** | a newer sha exists, so the run in flight answers about the past |
|
|
35
|
-
#
|
|
36
|
-
# **`head moved` is why this cannot live in the AgentMonitor.** A build's
|
|
37
|
-
# subject is a SHA, not a branch. A green result for code nobody will merge is
|
|
38
|
-
# worse than no result — it invites a merge of the wrong thing. Measured
|
|
39
|
-
# 2026-08-30: two merge waiters reported on superseded runs and had to be
|
|
40
|
-
# stopped and re-armed.
|
|
41
|
-
#
|
|
42
|
-
# **`action_required` is a real state here, not an edge case.** Bot branches hit
|
|
43
|
-
# it — the release PR's runs need manual approval before they start. A monitor
|
|
44
|
-
# that folded it into "not passed yet" would report a build as pending forever
|
|
45
|
-
# while it waits for a click nobody knows is needed.
|
|
46
|
-
#
|
|
47
|
-
# ═══════════════════════════════════════════════════════════════════════════
|
|
48
|
-
# IT POLLS NOTHING WHEN NO RUN IS LIVE
|
|
49
|
-
# ═══════════════════════════════════════════════════════════════════════════
|
|
50
|
-
#
|
|
51
|
-
# That is what makes a 30-second cadence against a host affordable, and it is
|
|
52
|
-
# the property that separates this monitor's budget from the AgentMonitor's.
|
|
53
|
-
# The AgentMonitor's five-minute budget exists because it asks ON EVERY PASS;
|
|
54
|
-
# this one asks only while there is something to ask about.
|
|
55
|
-
#
|
|
56
|
-
# STRUCTURALLY, NOT INCIDENTALLY: `monitor_head_sha` is a LOCAL git read and it
|
|
57
|
-
# gates the host call. `sample_finding` returns before `monitor_run_for_sha` is
|
|
58
|
-
# ever reached whenever there is no head to ask about, and once a sha has
|
|
59
|
-
# reached a terminal conclusion it is never asked about again. A monitor that
|
|
60
|
-
# kept questioning an idle host is the rate problem this whole design avoids,
|
|
61
|
-
# so the silence is asserted by the tests rather than assumed.
|
|
62
|
-
#
|
|
63
|
-
# ═══════════════════════════════════════════════════════════════════════════
|
|
64
|
-
# THE FINDINGS ARE TRANSITIONS, NOT CONDITIONS
|
|
65
|
-
# ═══════════════════════════════════════════════════════════════════════════
|
|
66
|
-
#
|
|
67
|
-
# The other monitors report states that PERSIST — `owes a review` holds until a
|
|
68
|
-
# PR exists, and republishing it would be repeating one fact. A build's answer
|
|
69
|
-
# CHANGES ONCE AND STAYS: a run that failed has failed, and it will still have
|
|
70
|
-
# failed in thirty seconds.
|
|
71
|
-
#
|
|
72
|
-
# So the state this monitor carries is keyed by SHA as well as by finding.
|
|
73
|
-
# Publishing `build passed` for one sha does not suppress `build passed` for the
|
|
74
|
-
# next one — that would silence the answer an operator is actually waiting for,
|
|
75
|
-
# on the very push they pushed to get it. And once a sha's build is terminal,
|
|
76
|
-
# the sha is not asked about again: the answer cannot change, so continuing to
|
|
77
|
-
# poll would be spending a host round trip to re-learn a fact already published.
|
|
78
|
-
#
|
|
79
|
-
# ═══════════════════════════════════════════════════════════════════════════
|
|
80
|
-
# IT OBSERVES; IT DOES NOT ACT
|
|
81
|
-
# ═══════════════════════════════════════════════════════════════════════════
|
|
82
|
-
#
|
|
83
|
-
# It does not rerun a workflow, approve a run that is `action_required`, merge a
|
|
84
|
-
# PR that went green, or push a fix for one that went red. Every one of those is
|
|
85
|
-
# a judgement with a blast radius. Approving a run in particular is a human's
|
|
86
|
-
# call by construction — `action_required` EXISTS because a person is meant to
|
|
87
|
-
# look — and a monitor that clicked it would defeat the gate it is reporting.
|
|
88
|
-
#
|
|
89
|
-
# PUBLISHING IS ITS ONLY OUTPUT, as next door: no state file, no cache, nothing
|
|
90
|
-
# written into the repository it watches, not even a record of what it last
|
|
91
|
-
# published. The variables that make "publish on change" work live in memory and
|
|
92
|
-
# die with the process, so a restarted monitor re-derives them one interval late
|
|
93
|
-
# rather than reading a stale one.
|
|
94
|
-
set -uo pipefail
|
|
95
|
-
|
|
96
|
-
usage() {
|
|
97
|
-
cat >&2 <<'EOF'
|
|
98
|
-
Usage: plot-build-monitor.sh [--once]
|
|
99
|
-
|
|
100
|
-
Started by plot-dispatch.sh inside the worker's wrapper. Reads its subject from
|
|
101
|
-
the environment, exactly as the wrapper's other children do:
|
|
102
|
-
|
|
103
|
-
PLOT_BRANCH the branch whose builds this monitor will report
|
|
104
|
-
PLOT_WORKTREE the desk it STARTS on; it follows the manifest's
|
|
105
|
-
`worktree` from then on (PLOT_MANIFEST_FILE below)
|
|
106
|
-
PLOT_MANIFEST_FILE the manifest this agent is named in, re-read each pass;
|
|
107
|
-
absent or unreadable keeps watching PLOT_WORKTREE
|
|
108
|
-
PLOT_MONITOR_FILE where findings are published (default:
|
|
109
|
-
$PLOT_WORKTREE/.plot-worker.monitor.build.jsonl). Set
|
|
110
|
-
explicitly, it WINS and does not follow a hop.
|
|
111
|
-
PLOT_MONITOR_INTERVAL seconds between passes (default 30)
|
|
112
|
-
|
|
113
|
-
--once take one sample and exit, rather than looping. A test
|
|
114
|
-
affordance: nothing dispatches this mode.
|
|
115
|
-
EOF
|
|
116
|
-
}
|
|
117
|
-
|
|
118
|
-
once=0
|
|
119
|
-
while [ $# -gt 0 ]; do
|
|
120
|
-
case "$1" in
|
|
121
|
-
--once) once=1 ;;
|
|
122
|
-
-h|--help) usage; exit 0 ;;
|
|
123
|
-
*) echo "plot-build-monitor: unknown argument '$1'" >&2; usage; exit 2 ;;
|
|
124
|
-
esac
|
|
125
|
-
shift
|
|
126
|
-
done
|
|
127
|
-
|
|
128
|
-
monitor='BuildMonitor'
|
|
129
|
-
|
|
130
|
-
branch="${PLOT_BRANCH:-}"
|
|
131
|
-
# The branch at start, kept for a detached or unreadable desk; `monitor_branch`
|
|
132
|
-
# re-reads the desk on every pass.
|
|
133
|
-
start_branch="$branch"
|
|
134
|
-
# THE DESK THIS MONITOR WAS LAUNCHED ON, FIXED FOR ITS WHOLE LIFE. `worktree`
|
|
135
|
-
# itself is no longer fixed: `monitor_pass` reassigns it every pass by asking
|
|
136
|
-
# `plot_watched_desk`, which follows a hop to the manifest's new `worktree`
|
|
137
|
-
# field. This is the launch desk `plot_watched_desk` falls back to when the
|
|
138
|
-
# manifest carries no override — never read directly after startup.
|
|
139
|
-
launched_worktree="${PLOT_WORKTREE:-}"
|
|
140
|
-
worktree="$launched_worktree"
|
|
141
|
-
manifest_file="${PLOT_MANIFEST_FILE:-}"
|
|
142
|
-
# THIRTY SECONDS IS AFFORDABLE ONLY BECAUSE OF THE SILENCE RULE. This cadence
|
|
143
|
-
# matches the WorkerMonitor's rather than the AgentMonitor's, and it asks a HOST
|
|
144
|
-
# — which would be the rate problem the AgentMonitor's 300 s exists to avoid,
|
|
145
|
-
# were it asking on every pass. It is not: no live run, no question. The budget
|
|
146
|
-
# is bounded by how long a build takes, not by how long a worker lives.
|
|
147
|
-
interval="${PLOT_MONITOR_INTERVAL:-30}"
|
|
148
|
-
|
|
149
|
-
# IF `PLOT_MONITOR_FILE` IS SET IT WINS AND DOES NOT FOLLOW THE DESK — the
|
|
150
|
-
# existing contract this usage text already states. Otherwise `findings`
|
|
151
|
-
# follows `worktree`, reassigned alongside it at the top of `monitor_pass`.
|
|
152
|
-
monitor_file_override="${PLOT_MONITOR_FILE:-}"
|
|
153
|
-
findings="${monitor_file_override:-${worktree:+$worktree/.plot-worker.monitor.build.jsonl}}"
|
|
154
|
-
|
|
155
|
-
# THE SUBJECT, read the same way the other two monitors read it.
|
|
156
|
-
pid_file="${PLOT_PID_FILE:-${worktree:+$worktree/.plot-worker.pid}}"
|
|
157
|
-
|
|
158
|
-
# ONE ANSWER TO "IS MY SUBJECT STILL THERE?", shared with both siblings, for the
|
|
159
|
-
# reason `plot-monitor-subject.sh` documents: three monitors deciding
|
|
160
|
-
# independently when to stop would drift, and the failure would be silent.
|
|
161
|
-
# shellcheck source=./plot-monitor-subject.sh
|
|
162
|
-
. "$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)/plot-monitor-subject.sh"
|
|
163
|
-
|
|
164
|
-
# THE HOST ADAPTER, sourced for its path rather than its functions: the one
|
|
165
|
-
# operation this monitor asks is `plot-host.sh run-for-sha`, and `plot-host.sh`
|
|
166
|
-
# is the ONE place that talks to the host CLI. A monitor calling `gh` directly
|
|
167
|
-
# would be a second adapter, and the backend split (github/bitbucket) would have
|
|
168
|
-
# to be decided twice.
|
|
169
|
-
host_script="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)/plot-host.sh"
|
|
170
|
-
|
|
171
|
-
json_escape() { # $1 = raw → prints a JSON-safe string body
|
|
172
|
-
printf '%s' "$1" | python3 -c 'import json,sys; sys.stdout.write(json.dumps(sys.stdin.read())[1:-1])' 2>/dev/null \
|
|
173
|
-
|| printf '%s' "$1" | sed 's/\\/\\\\/g; s/"/\\"/g'
|
|
174
|
-
}
|
|
175
|
-
|
|
176
|
-
# A FINDING CARRIES THE SAME FOUR FIELDS the other two monitors publish —
|
|
177
|
-
# `finding`, `since`, `evidence`, `measuredAt` — for the same reason: one
|
|
178
|
-
# subscriber will read all three files and must not need a third parser to do
|
|
179
|
-
# it.
|
|
180
|
-
#
|
|
181
|
-
# `since` HERE IS WHEN THE BUILD REACHED THIS ANSWER, as far as this monitor
|
|
182
|
-
# can tell — the pass that first saw it. On a 30 s cadence the gap to
|
|
183
|
-
# `measuredAt` is small by construction, which is the opposite of the
|
|
184
|
-
# AgentMonitor's case and follows from the same field meaning the same thing.
|
|
185
|
-
publish() { # $1=finding $2=evidence $3=since
|
|
186
|
-
local now
|
|
187
|
-
now=$(date -u +%Y-%m-%dT%H:%M:%SZ)
|
|
188
|
-
local line
|
|
189
|
-
line=$(printf '{"monitor":"%s","branch":"%s","worktree":"%s","finding":"%s","since":"%s","evidence":"%s","measuredAt":"%s"}' \
|
|
190
|
-
"$monitor" \
|
|
191
|
-
"$(json_escape "$branch")" \
|
|
192
|
-
"$(json_escape "$worktree")" \
|
|
193
|
-
"$(json_escape "$1")" \
|
|
194
|
-
"${3:-$now}" \
|
|
195
|
-
"$(json_escape "$2")" \
|
|
196
|
-
"$now")
|
|
197
|
-
# Both destinations, for the reason the siblings give: the file is what a
|
|
198
|
-
# subscriber reads, and stdout lands in `.plot-worker.log` beside the agent's
|
|
199
|
-
# own output where an operator tailing a worker sees it.
|
|
200
|
-
[ -n "$findings" ] && printf '%s\n' "$line" >> "$findings" 2>/dev/null
|
|
201
|
-
printf 'plot-monitor %s\n' "$line"
|
|
202
|
-
}
|
|
203
|
-
|
|
204
|
-
# ---------------------------------------------------------------------------
|
|
205
|
-
# THE PORTS — two named seams, so every branch is reachable from a test
|
|
206
|
-
# ---------------------------------------------------------------------------
|
|
207
|
-
#
|
|
208
|
-
# Same convention as the four next door and the five beyond it: a test sources
|
|
209
|
-
# this file with `PLOT_MONITOR_NO_MAIN=1` and REDEFINES them. Here there are
|
|
210
|
-
# only two, and one of them is the host round trip — which is exactly the seam a
|
|
211
|
-
# test can least afford to drive for real. You cannot ask GitHub for a run that
|
|
212
|
-
# is `action_required` on demand, you cannot make a run vanish to order, and you
|
|
213
|
-
# certainly cannot arrange two runs for two shas at the instant a test needs
|
|
214
|
-
# them. Every one of those is a branch this monitor must get right.
|
|
215
|
-
|
|
216
|
-
# What sha is this branch's head, right now?
|
|
217
|
-
#
|
|
218
|
-
# → prints the sha, or nothing when it cannot be read
|
|
219
|
-
#
|
|
220
|
-
# LOCAL, AND IT GATES THE HOST CALL. This is the cheap reading that makes the
|
|
221
|
-
# silence rule structural: no head, no question. It reads the WORKTREE's HEAD
|
|
222
|
-
# rather than a remote ref, because the head a build should be about is the one
|
|
223
|
-
# the agent has actually produced — and because a fetch per pass would be a
|
|
224
|
-
# second network call on a 30-second loop.
|
|
225
|
-
monitor_head_sha() { # → prints a sha, or nothing
|
|
226
|
-
[ -n "$worktree" ] && [ -d "$worktree" ] || return 0
|
|
227
|
-
git -C "$worktree" rev-parse --verify --quiet HEAD 2>/dev/null || true
|
|
228
|
-
}
|
|
229
|
-
|
|
230
|
-
# Which branch does the desk hold, right now?
|
|
231
|
-
#
|
|
232
|
-
# → prints the branch, or the branch the monitor started with when the desk is
|
|
233
|
-
# detached or unreadable
|
|
234
|
-
#
|
|
235
|
-
# READ ON EVERY PASS, NOT ONCE. A free agent starts with an empty `PLOT_BRANCH`
|
|
236
|
-
# and is handed its slices later, in the same desk; a dispatched agent hops to
|
|
237
|
-
# its next slice in the same desk too. A branch fixed at start made this monitor
|
|
238
|
-
# unaskable for every free agent (`monitor_run_for_sha` returns 2 on an empty
|
|
239
|
-
# branch) and wrong after every hop, so the correction path in
|
|
240
|
-
# `plot-worker-loop.sh` received no finding for most slices.
|
|
241
|
-
monitor_branch() { # → prints a branch, or nothing
|
|
242
|
-
local current=''
|
|
243
|
-
if [ -n "$worktree" ] && [ -d "$worktree" ]; then
|
|
244
|
-
current=$(git -C "$worktree" branch --show-current 2>/dev/null || true)
|
|
245
|
-
fi
|
|
246
|
-
printf '%s' "${current:-$start_branch}"
|
|
247
|
-
}
|
|
248
|
-
|
|
249
|
-
# What does the host say about the run for ONE sha?
|
|
250
|
-
#
|
|
251
|
-
# → prints the run JSON, or nothing when there is no run for it
|
|
252
|
-
# → returns 2 when the host could not be asked at all
|
|
253
|
-
#
|
|
254
|
-
# THE DISTINCTION BETWEEN "NO RUN" AND "COULD NOT ASK" IS THE POINT, and it is
|
|
255
|
-
# the same discipline `monitor_pr_state` keeps next door. An unreachable host is
|
|
256
|
-
# NOT evidence that a build is absent. A monitor that read a `gh` failure as "no
|
|
257
|
-
# run" would go silent about every build on the estate the moment a token
|
|
258
|
-
# expired — and silence is this monitor's healthy signal, so the failure would
|
|
259
|
-
# be invisible by construction.
|
|
260
|
-
#
|
|
261
|
-
# An EMPTY result with a reachable host is a real and common answer: the run has
|
|
262
|
-
# not been created yet. A monitor polling a fresh push sees exactly this until
|
|
263
|
-
# CI wakes up, and it is not a finding.
|
|
264
|
-
monitor_run_for_sha() { # $1 = sha → prints run JSON | rc 2 = unaskable
|
|
265
|
-
[ -n "$branch" ] || return 2
|
|
266
|
-
[ -x "$host_script" ] || return 2
|
|
267
|
-
local out
|
|
268
|
-
out=$("$host_script" run-for-sha "$branch" "$1" 2>/dev/null) || return 2
|
|
269
|
-
printf '%s' "$out"
|
|
270
|
-
return 0
|
|
271
|
-
}
|
|
272
|
-
|
|
273
|
-
# ---------------------------------------------------------------------------
|
|
274
|
-
# READING ONE FIELD OUT OF THE RUN
|
|
275
|
-
# ---------------------------------------------------------------------------
|
|
276
|
-
#
|
|
277
|
-
# `jq` where it exists and `sed` where it does not, the same fallback shape
|
|
278
|
-
# `json_escape` uses above. A monitor that died because `jq` was missing would
|
|
279
|
-
# be a monitor that reported nothing on exactly the machines least likely to
|
|
280
|
-
# have anyone watching.
|
|
281
|
-
run_field() { # $1 = json, $2 = key → prints the value, or nothing for null
|
|
282
|
-
local v
|
|
283
|
-
if command -v jq >/dev/null 2>&1; then
|
|
284
|
-
v=$(printf '%s' "$1" | jq -r --arg k "$2" '.[$k] // empty' 2>/dev/null)
|
|
285
|
-
else
|
|
286
|
-
v=$(printf '%s' "$1" | sed -n 's/.*"'"$2"'":"\([^"]*\)".*/\1/p')
|
|
287
|
-
fi
|
|
288
|
-
printf '%s' "$v"
|
|
289
|
-
}
|
|
290
|
-
|
|
291
|
-
# ---------------------------------------------------------------------------
|
|
292
|
-
# THE SAMPLER — one pass, using only the ports above
|
|
293
|
-
# ---------------------------------------------------------------------------
|
|
294
|
-
#
|
|
295
|
-
# THE STATE IS TWO VARIABLES AND IT IS DERIVED, as next door — but `published`
|
|
296
|
-
# is keyed by SHA here, because the findings are transitions rather than
|
|
297
|
-
# conditions. `published_sha` records which commit the standing answer is about,
|
|
298
|
-
# so the same word about a different commit is still news.
|
|
299
|
-
published=''
|
|
300
|
-
published_sha=''
|
|
301
|
-
since=''
|
|
302
|
-
# The shas whose builds have reached a terminal answer. Once a run has failed,
|
|
303
|
-
# passed, or been superseded, asking again spends a host round trip to re-learn
|
|
304
|
-
# a fact already published — so it is not asked. THIS is the second half of "it
|
|
305
|
-
# polls nothing when no run is live": the first half is having no head at all,
|
|
306
|
-
# and this is having no OPEN question about the head there is.
|
|
307
|
-
settled_shas=''
|
|
308
|
-
|
|
309
|
-
sha_is_settled() { # $1 = sha → 0 settled | 1 not
|
|
310
|
-
case " $settled_shas " in *" $1 "*) return 0 ;; esac
|
|
311
|
-
return 1
|
|
312
|
-
}
|
|
313
|
-
|
|
314
|
-
# ---------------------------------------------------------------------------
|
|
315
|
-
# ONE FINDING PER PASS, AND `head moved` COMES FIRST
|
|
316
|
-
# ---------------------------------------------------------------------------
|
|
317
|
-
#
|
|
318
|
-
# Unlike the AgentMonitor's four, these are near-exclusive by construction — a
|
|
319
|
-
# run has one status. The one genuine overlap is the one that matters: a run
|
|
320
|
-
# that concluded `success` for a sha the branch has since moved past is BOTH
|
|
321
|
-
# "passed" and "superseded", and reporting it as passed is precisely the failure
|
|
322
|
-
# this monitor exists to prevent.
|
|
323
|
-
#
|
|
324
|
-
# So `head moved` is decided FIRST and about the run's own sha, not about the
|
|
325
|
-
# branch: if the answer in hand describes a commit that is no longer the head,
|
|
326
|
-
# the answer is about the past whatever it says.
|
|
327
|
-
sample_finding() { # → prints "finding\tevidence", or nothing
|
|
328
|
-
local head
|
|
329
|
-
head=$(monitor_head_sha)
|
|
330
|
-
|
|
331
|
-
# NO HEAD, NO QUESTION. The cheap local reading refuses before anything
|
|
332
|
-
# reaches the host — the structural form of "it polls nothing when no run is
|
|
333
|
-
# live". A worktree that is gone, or a branch with no commit yet, asks
|
|
334
|
-
# nothing at all.
|
|
335
|
-
[ -n "$head" ] || return 0
|
|
336
|
-
|
|
337
|
-
# ALREADY ANSWERED. The head's build reached a terminal conclusion on an
|
|
338
|
-
# earlier pass and a build's answer does not change back. Asking again would
|
|
339
|
-
# be the idle polling this design refuses.
|
|
340
|
-
sha_is_settled "$head" && return 0
|
|
341
|
-
|
|
342
|
-
local run rc
|
|
343
|
-
run=$(monitor_run_for_sha "$head"); rc=$?
|
|
344
|
-
|
|
345
|
-
# UNASKABLE — no finding. The host could not be asked, so nothing about the
|
|
346
|
-
# build is known. Reporting anything here would be inventing an answer out of
|
|
347
|
-
# a failure to observe.
|
|
348
|
-
[ "$rc" = 2 ] && return 0
|
|
349
|
-
|
|
350
|
-
# NO RUN YET — no finding, and not an error. The commit exists and CI has not
|
|
351
|
-
# created a run for it. This is the ordinary state of a freshly pushed sha,
|
|
352
|
-
# and it is what the monitor sees on every pass until the run appears.
|
|
353
|
-
[ -n "$run" ] || return 0
|
|
354
|
-
|
|
355
|
-
local run_sha status conclusion url
|
|
356
|
-
run_sha=$(run_field "$run" sha)
|
|
357
|
-
status=$(run_field "$run" status)
|
|
358
|
-
conclusion=$(run_field "$run" conclusion)
|
|
359
|
-
url=$(run_field "$run" url)
|
|
360
|
-
|
|
361
|
-
# 1. HEAD MOVED — the run in hand is about a commit that is no longer the
|
|
362
|
-
# head. Decided BEFORE the conclusion is read, because a green run for
|
|
363
|
-
# superseded code is the specific wrong answer this monitor was built to
|
|
364
|
-
# avoid: it invites a merge of the wrong thing. Measured 2026-08-30 — two
|
|
365
|
-
# merge waiters reported on superseded runs and had to be stopped and
|
|
366
|
-
# re-armed.
|
|
367
|
-
#
|
|
368
|
-
# This fires when the host answered about a DIFFERENT sha than the one asked
|
|
369
|
-
# about, which is the shape a race actually takes: the head moved between the
|
|
370
|
-
# local read and the host's reply.
|
|
371
|
-
if [ -n "$run_sha" ] && [ "$run_sha" != "$head" ]; then
|
|
372
|
-
printf 'head moved\tthe run at %s is for %s, but the branch head is now %s; its answer is about the past\n' \
|
|
373
|
-
"${url:-an unknown url}" "$run_sha" "$head"
|
|
374
|
-
return 0
|
|
375
|
-
fi
|
|
376
|
-
|
|
377
|
-
# 2. BUILD NEEDS APPROVAL — a real state, not an edge case. Bot branches hit
|
|
378
|
-
# it: the release PR's runs need a manual click before they start. GitHub
|
|
379
|
-
# reports it as a `status` of `waiting`/`action_required` and as a
|
|
380
|
-
# `conclusion` of `action_required`, depending on where the run is, so both
|
|
381
|
-
# are read. Folding it into "not passed yet" would report the build pending
|
|
382
|
-
# forever while it waits for a click nobody knows is needed.
|
|
383
|
-
case "$status:$conclusion" in
|
|
384
|
-
*action_required*|waiting:*)
|
|
385
|
-
printf 'build needs approval\tthe run at %s for %s is waiting for a manual approval before it can start\n' \
|
|
386
|
-
"${url:-an unknown url}" "$head"
|
|
387
|
-
return 0
|
|
388
|
-
;;
|
|
389
|
-
esac
|
|
390
|
-
|
|
391
|
-
# 3 & 4. THE TERMINAL CONCLUSIONS. An empty conclusion means the run is still
|
|
392
|
-
# going — queued or in progress — and a monitor whose subject is a transition
|
|
393
|
-
# says nothing about a state that has not changed yet.
|
|
394
|
-
[ -n "$conclusion" ] || return 0
|
|
395
|
-
|
|
396
|
-
case "$conclusion" in
|
|
397
|
-
success)
|
|
398
|
-
printf 'build passed\tthe run at %s for %s concluded success\n' "${url:-an unknown url}" "$head"
|
|
399
|
-
return 0
|
|
400
|
-
;;
|
|
401
|
-
# EVERY OTHER TERMINAL CONCLUSION IS A FAILURE TO A READER WAITING ON GREEN.
|
|
402
|
-
# `failure`, `timed_out`, `cancelled` and `startup_failure` differ in cause
|
|
403
|
-
# and not in consequence: none of them is a build somebody may merge on. The
|
|
404
|
-
# cause is not thrown away — it rides in the evidence, where a reader
|
|
405
|
-
# deciding whether to rerun can see it.
|
|
406
|
-
*)
|
|
407
|
-
printf 'build failed\tthe run at %s for %s concluded %s\n' "${url:-an unknown url}" "$head" "$conclusion"
|
|
408
|
-
return 0
|
|
409
|
-
;;
|
|
410
|
-
esac
|
|
411
|
-
}
|
|
412
|
-
|
|
413
|
-
# One full pass: sample, publish only on a change of ANSWER-ABOUT-A-COMMIT.
|
|
414
|
-
monitor_pass() {
|
|
415
|
-
local row finding evidence head
|
|
416
|
-
# REASSIGNED ONCE, HERE, RATHER THAN THREADED THROUGH EACH READER. `worktree`
|
|
417
|
-
# feeds `monitor_head_sha`, `monitor_branch` and `publish()`'s own field, so a
|
|
418
|
-
# hop is picked up in one place rather than three. `PLOT_MONITOR_FILE`, when
|
|
419
|
-
# set, WINS and does not follow the desk — the existing contract the usage
|
|
420
|
-
# text states.
|
|
421
|
-
worktree=$(plot_watched_desk "$manifest_file" "$launched_worktree")
|
|
422
|
-
findings="${monitor_file_override:-${worktree:+$worktree/.plot-worker.monitor.build.jsonl}}"
|
|
423
|
-
branch=$(monitor_branch)
|
|
424
|
-
head=$(monitor_head_sha)
|
|
425
|
-
row=$(sample_finding)
|
|
426
|
-
finding="${row%% *}"
|
|
427
|
-
evidence=''
|
|
428
|
-
case "$row" in *" "*) evidence="${row#* }" ;; esac
|
|
429
|
-
[ -z "$row" ] && finding=''
|
|
430
|
-
|
|
431
|
-
# PUBLISH ON A CHANGE OF EITHER THE FINDING OR THE COMMIT IT IS ABOUT. The
|
|
432
|
-
# second half is what makes these transitions rather than conditions: `build
|
|
433
|
-
# passed` for a new sha is news even though the word is the same as last
|
|
434
|
-
# time, and suppressing it would silence exactly the answer an operator
|
|
435
|
-
# pushed in order to get.
|
|
436
|
-
if [ "$finding" != "$published" ] || { [ -n "$finding" ] && [ "$head" != "$published_sha" ]; }; then
|
|
437
|
-
if [ -n "$finding" ]; then
|
|
438
|
-
since=$(date -u +%Y-%m-%dT%H:%M:%SZ)
|
|
439
|
-
publish "$finding" "$evidence" "$since"
|
|
440
|
-
# SETTLED, so it is never asked about again. Every finding this monitor
|
|
441
|
-
# publishes is terminal for its sha: a failure stays failed, a pass stays
|
|
442
|
-
# passed, and a superseded run does not become current. Recording it here
|
|
443
|
-
# rather than in `sample_finding` keeps the decision beside the publish it
|
|
444
|
-
# follows from.
|
|
445
|
-
[ -n "$head" ] && settled_shas="$settled_shas $head"
|
|
446
|
-
fi
|
|
447
|
-
# NO CLEARING PUBLISH, and that is the difference from the AgentMonitor. A
|
|
448
|
-
# debt is cleared when it is paid — that is news. A build's answer is never
|
|
449
|
-
# withdrawn: `build failed` does not stop being true about that sha, and the
|
|
450
|
-
# next answer is a new finding about a new commit, which the sha key already
|
|
451
|
-
# carries. Publishing `clear` here would tell a subscriber a failure had
|
|
452
|
-
# been resolved when all that happened is the branch moved on.
|
|
453
|
-
published="$finding"
|
|
454
|
-
published_sha="$head"
|
|
455
|
-
fi
|
|
456
|
-
}
|
|
457
|
-
|
|
458
|
-
# SOURCEABLE FOR TESTS, the same guard the siblings carry. A test that wants to
|
|
459
|
-
# drive `monitor_pass` against redefined ports needs the functions without the
|
|
460
|
-
# loop; everything above this line defines, and nothing below it runs when the
|
|
461
|
-
# guard is set.
|
|
462
|
-
[ -n "${PLOT_MONITOR_NO_MAIN:-}" ] && return 0 2>/dev/null
|
|
463
|
-
|
|
464
|
-
monitor_pass
|
|
465
|
-
[ "$once" = 1 ] && exit 0
|
|
466
|
-
|
|
467
|
-
# IT ENDS WITH ITS AGENT, by the mechanism `plot-monitor-subject.sh` documents
|
|
468
|
-
# and for the reason `docs/research/2026-08-30-what-ends-a-monitor.md` measured.
|
|
469
|
-
#
|
|
470
|
-
# PUBLISH FIRST, THEN LEAVE. The final pass runs with the agent already gone,
|
|
471
|
-
# and it matters here for a reason of its own: an agent that pushes and exits
|
|
472
|
-
# leaves a run still going, and the answer arrives after there is nobody left to
|
|
473
|
-
# see it. The last pass is the one chance to catch a build that concluded during
|
|
474
|
-
# the shutdown.
|
|
475
|
-
while plot_monitor_wait "$interval" "$pid_file"; do
|
|
476
|
-
monitor_pass
|
|
477
|
-
done
|
|
478
|
-
|
|
479
|
-
monitor_pass
|
|
480
|
-
exit 0
|
package/plot-transcript-quiet.sh
DELETED
|
@@ -1,172 +0,0 @@
|
|
|
1
|
-
#!/usr/bin/env bash
|
|
2
|
-
# The ONE answer to "how long has this worktree's agent been quiet?" — sourced,
|
|
3
|
-
# not run, by `plot-worker-loop.sh` for its own watcher (`plot-worker-state.sh`'s
|
|
4
|
-
# `plot_worker_idle_watch_pass`) and for its own ending message. Sourced by the
|
|
5
|
-
# now-deleted `plot-worker-monitor.sh` until `bug/the-loop-reports-idle`.
|
|
6
|
-
#
|
|
7
|
-
# It reads the AGENT rather than the machine. A `claude -p` session appends a
|
|
8
|
-
# timestamped line to its transcript for every model turn, tool call and tool
|
|
9
|
-
# result; the seconds since the newest of those lines is how long the agent has
|
|
10
|
-
# produced nothing. A CPU sample answers *is this process on a core right now?*,
|
|
11
|
-
# which is a different question and is zero for most of a working agent's life.
|
|
12
|
-
#
|
|
13
|
-
# ═══════════════════════════════════════════════════════════════════════════
|
|
14
|
-
# THREE ANSWERS, AND THE THIRD IS NOT A FAILURE
|
|
15
|
-
# ═══════════════════════════════════════════════════════════════════════════
|
|
16
|
-
#
|
|
17
|
-
# <seconds> the newest transcript line is that many seconds old
|
|
18
|
-
# unavailable no transcript can be read for this worktree, INCLUDING the
|
|
19
|
-
# case where the arguments name no worktree at all
|
|
20
|
-
#
|
|
21
|
-
# `unavailable` is a first-class answer, settled by
|
|
22
|
-
# `the-registry-supervises-its-agents`: a capability the adopting project does
|
|
23
|
-
# not provide is UNAVAILABLE, never failed and never zero. A caller that read
|
|
24
|
-
# it as "quiet for 0 seconds" would report every unreadable agent healthy; one
|
|
25
|
-
# that read it as an error would refuse to run where Plot's own contract says
|
|
26
|
-
# it should degrade. So it is a word, and callers match on it.
|
|
27
|
-
#
|
|
28
|
-
# ═══════════════════════════════════════════════════════════════════════════
|
|
29
|
-
# THE TRANSCRIPT IS FOUND BY PATH, NOT BY SESSION ID — AND THAT IS A CHOICE
|
|
30
|
-
# ═══════════════════════════════════════════════════════════════════════════
|
|
31
|
-
#
|
|
32
|
-
# `.plot/worker-prompt.sh:29` DOES pass `--session-id` now (2026-09-04), so an
|
|
33
|
-
# exact join is available here and this deliberately does not take it. The
|
|
34
|
-
# reason is the question, not the capability: this asks *is anything happening
|
|
35
|
-
# at this desk*, and an operator's own session at the same worktree is a true
|
|
36
|
-
# answer to it. Joining on the worker's id alone would report a desk quiet
|
|
37
|
-
# while somebody is visibly working at it.
|
|
38
|
-
#
|
|
39
|
-
# THE OPPOSITE JOIN IS RIGHT FOR THE OPPOSITE QUESTION. *What has THIS agent
|
|
40
|
-
# spent* is per session and never per worktree — one project directory measured
|
|
41
|
-
# 2026-09-03 held 45 session files, 30 of them subagents, and a sum across them
|
|
42
|
-
# belongs to no one. `rules/spend.ts` states that side; the two read the same
|
|
43
|
-
# files and must not be made to share a join.
|
|
44
|
-
#
|
|
45
|
-
# A THIRD QUESTION USES BOTH KEYS: *has this worker's conversation written?* It
|
|
46
|
-
# is per worktree AND per conversation handle, and it is asked as a file's
|
|
47
|
-
# existence, never as a time: `plot_transcript_exists` below. The loop asks it
|
|
48
|
-
# to choose `--session-id` or `--resume`; the worker monitor asks it before it
|
|
49
|
-
# calls a quiet desk idle, because until the new conversation writes its first
|
|
50
|
-
# line the desk's newest file belongs to the previous one. It does not change
|
|
51
|
-
# the quiet number, which stays about the desk.
|
|
52
|
-
#
|
|
53
|
-
# So the join here is the one `plot-quiet-stretch.mjs` already made and proved
|
|
54
|
-
# on 23 real sessions: the runtime stores a session under
|
|
55
|
-
# `$HOME/.claude/projects/<slug>` where the slug is the WORKTREE PATH with `/`
|
|
56
|
-
# and `.` replaced by `-`. A dispatched worker has its own worktree, so the
|
|
57
|
-
# path identifies the DESK without any id being passed anywhere.
|
|
58
|
-
#
|
|
59
|
-
# `agent-` PREFIXED FILES ARE SKIPPED, for wave 1's reason: a subagent's
|
|
60
|
-
# transcript is a true statement about the wrong process. A worker whose
|
|
61
|
-
# subagent is chatting while the worker itself has stopped must read as quiet.
|
|
62
|
-
#
|
|
63
|
-
# ═══════════════════════════════════════════════════════════════════════════
|
|
64
|
-
# THE NEWEST LINE ACROSS ALL OF A WORKTREE'S SESSIONS
|
|
65
|
-
# ═══════════════════════════════════════════════════════════════════════════
|
|
66
|
-
#
|
|
67
|
-
# A worktree can hold several sessions: a worker that hopped waves, or an
|
|
68
|
-
# operator who opened a session at the same desk. Any of them producing output
|
|
69
|
-
# means SOMEBODY is working there, and the monitor's question is about the desk.
|
|
70
|
-
# Taking the maximum timestamp is the answer that never ends a live session
|
|
71
|
-
# because a stale sibling exists beside it.
|
|
72
|
-
|
|
73
|
-
# The runtime's project-slug derivation. Duplicated from `plot-quiet-stretch.mjs`
|
|
74
|
-
# — which duplicates it from the board's `projectSlug` — because a monitor on a
|
|
75
|
-
# 30s loop must not depend on a build step. One line, pinned by a test on each
|
|
76
|
-
# side.
|
|
77
|
-
plot_transcript_slug() { # $1=worktree path → slug
|
|
78
|
-
printf '%s' "$1" | tr '/.' '--'
|
|
79
|
-
}
|
|
80
|
-
|
|
81
|
-
# Where the runtime keeps this worktree's transcripts, or "" if nowhere.
|
|
82
|
-
plot_transcript_dir() { # $1=worktree → directory path (may not exist)
|
|
83
|
-
local wt="$1" home="${PLOT_TRANSCRIPT_HOME:-${HOME:-}}"
|
|
84
|
-
[ -n "$wt" ] && [ -n "$home" ] || return 0
|
|
85
|
-
printf '%s/.claude/projects/%s' "$home" "$(plot_transcript_slug "$wt")"
|
|
86
|
-
}
|
|
87
|
-
|
|
88
|
-
# Seconds since this worktree's agent last wrote anything.
|
|
89
|
-
#
|
|
90
|
-
# THE MTIME IS THE READING, not the file's contents. The runtime appends as it
|
|
91
|
-
# works, so the file's modification time IS the timestamp of its newest line —
|
|
92
|
-
# and reading it costs one `stat` rather than parsing a transcript that reaches
|
|
93
|
-
# tens of megabytes on a long session. Wave 1 parsed timestamps because it was
|
|
94
|
-
# measuring a DISTRIBUTION of past gaps; this needs only the newest, on a 30s
|
|
95
|
-
# loop, for as many workers as the machine holds.
|
|
96
|
-
#
|
|
97
|
-
# A CLOCK IS THE RIGHT INSTRUMENT HERE, and that is worth stating because Plot
|
|
98
|
-
# usually refuses one. `plot-estate-changed.sh` hashes content rather than
|
|
99
|
-
# reading mtime, because its question is *did this change?* and a checkout moves
|
|
100
|
-
# mtime without changing anything. This question is *how long since output?* —
|
|
101
|
-
# which is a question about elapsed time, and mtime is the measurement of it.
|
|
102
|
-
plot_transcript_quiet_seconds() { # $1=worktree → seconds | unavailable
|
|
103
|
-
local wt="$1" dir newest now mtime
|
|
104
|
-
dir=$(plot_transcript_dir "$wt")
|
|
105
|
-
[ -n "$dir" ] && [ -d "$dir" ] || { printf 'unavailable'; return 0; }
|
|
106
|
-
|
|
107
|
-
# The newest mtime across every non-`agent-` session file in the directory.
|
|
108
|
-
# `find -print0` and a while-read keep paths with spaces intact; `stat` is
|
|
109
|
-
# asked once per file, and a worktree holds one to eight.
|
|
110
|
-
newest=''
|
|
111
|
-
while IFS= read -r -d '' f; do
|
|
112
|
-
case "$(basename "$f")" in agent-*) continue ;; esac
|
|
113
|
-
mtime=$(plot_transcript_mtime "$f") || continue
|
|
114
|
-
[ -n "$mtime" ] || continue
|
|
115
|
-
if [ -z "$newest" ] || [ "$mtime" -gt "$newest" ] 2>/dev/null; then newest="$mtime"; fi
|
|
116
|
-
done < <(find "$dir" -maxdepth 1 -type f -name '*.jsonl' -print0 2>/dev/null)
|
|
117
|
-
|
|
118
|
-
# A directory that exists but holds no session file is still UNAVAILABLE. The
|
|
119
|
-
# runtime creates the directory when the project is first opened, so an empty
|
|
120
|
-
# one means no session has written here — which is exactly "no transcript can
|
|
121
|
-
# be read", not "quiet for a very long time".
|
|
122
|
-
[ -n "$newest" ] || { printf 'unavailable'; return 0; }
|
|
123
|
-
|
|
124
|
-
now=$(date +%s)
|
|
125
|
-
local quiet=$(( now - newest ))
|
|
126
|
-
# A transcript written in the future — a clock skew across a mounted volume —
|
|
127
|
-
# reads as zero rather than negative. An end condition comparing a negative
|
|
128
|
-
# against a threshold would behave correctly by accident here and not
|
|
129
|
-
# elsewhere; clamping says what is meant.
|
|
130
|
-
[ "$quiet" -lt 0 ] && quiet=0
|
|
131
|
-
printf '%s' "$quiet"
|
|
132
|
-
}
|
|
133
|
-
|
|
134
|
-
# Does the conversation `id` have a transcript at this worktree?
|
|
135
|
-
#
|
|
136
|
-
# EXISTENCE, NOT A TIMESTAMP. The runtime creates `<id>.jsonl` with its first
|
|
137
|
-
# line and appends to it after, under both `--session-id` and `--resume`. So the
|
|
138
|
-
# file's presence says the conversation has written, and no comparison of
|
|
139
|
-
# clocks is made, so a file created in the same second as a manifest write
|
|
140
|
-
# reads as present.
|
|
141
|
-
#
|
|
142
|
-
# NO HANDLE AND NO FILE ARE ONE ANSWER HERE, and a caller that must tell them
|
|
143
|
-
# apart checks the handle first. `session_flag` reads both as *create*; the
|
|
144
|
-
# monitor's port does not, because a monitor with no handle must not read every
|
|
145
|
-
# quiet worker as unspoken.
|
|
146
|
-
plot_transcript_exists() { # $1=worktree $2=id → 0 found | 1 not
|
|
147
|
-
local wt="$1" id="$2" dir
|
|
148
|
-
[ -n "$wt" ] && [ -n "$id" ] || return 1
|
|
149
|
-
dir=$(plot_transcript_dir "$wt" 2>/dev/null) || return 1
|
|
150
|
-
[ -n "$dir" ] || return 1
|
|
151
|
-
[ -f "$dir/$id.jsonl" ]
|
|
152
|
-
}
|
|
153
|
-
|
|
154
|
-
# A file's modification time as a unix epoch. BSD and GNU `stat` disagree on the
|
|
155
|
-
# flag, and a monitor that works on the author's laptop and not in CI is a
|
|
156
|
-
# monitor nobody trusts.
|
|
157
|
-
#
|
|
158
|
-
# THE `||` IS NOT ENOUGH, AND CI MEASURED WHY. On Linux `stat -f` is not an
|
|
159
|
-
# unknown flag — it means FILESYSTEM info, and it SUCCEEDS. So the BSD form
|
|
160
|
-
# never falls through: it printed `Namelen: 255 Type: ext2/ext3` and the
|
|
161
|
-
# caller subtracted that from a clock. Two of this branch's own tests failed on
|
|
162
|
-
# it, 2026-09-02, having passed on macOS.
|
|
163
|
-
#
|
|
164
|
-
# So the answer is validated rather than trusted. Each form must yield digits;
|
|
165
|
-
# anything else is treated as that form not being available here.
|
|
166
|
-
plot_transcript_mtime() { # $1=file → epoch seconds
|
|
167
|
-
local m
|
|
168
|
-
m=$(stat -c '%Y' "$1" 2>/dev/null)
|
|
169
|
-
case "$m" in ''|*[!0-9]*) m=$(stat -f '%m' "$1" 2>/dev/null) ;; esac
|
|
170
|
-
case "$m" in ''|*[!0-9]*) return 1 ;; esac
|
|
171
|
-
printf '%s' "$m"
|
|
172
|
-
}
|