@plot-pm/board 0.14.2 → 0.14.4
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/board-server.mjs +144 -142
- package/package.json +1 -1
- package/plot-budget.sh +166 -1
- package/plot-deliver.sh +14 -8
- package/plot-dispatch.sh +54 -46
- package/plot-fleet-scan.sh +134 -7
- package/plot-host.sh +728 -38
- package/plot-plan-meta.sh +23 -7
- package/plot-reap.sh +37 -10
package/package.json
CHANGED
package/plot-budget.sh
CHANGED
|
@@ -221,7 +221,7 @@ BUDGET_FALLBACK_WINDOW_MS=3600000
|
|
|
221
221
|
# which is a reading about whichever pool was spent last — so a caller deciding
|
|
222
222
|
# whether a bucket is spent must name that bucket. `graphql_budget_spent` does,
|
|
223
223
|
# and this is why.
|
|
224
|
-
|
|
224
|
+
budget_rate_read() {
|
|
225
225
|
local connector="${1:-}" account="${2:-}" bucket="${3:-}" now="${4:-}"
|
|
226
226
|
local path
|
|
227
227
|
[ -n "$now" ] || now="$(budget_now_ms)"
|
|
@@ -321,6 +321,171 @@ budget_rate() {
|
|
|
321
321
|
' "$path" 2>/dev/null || echo '{"spent":0,"spanMs":0,"perHour":null,"lines":0,"unreadable":0,"limit":null,"remaining":null,"resetAt":null,"basis":"unknown"}'
|
|
322
322
|
}
|
|
323
323
|
|
|
324
|
+
# THE SAME ANSWER IS SCANNED FOR ONCE PER PROCESS, and that is the whole of this
|
|
325
|
+
# memo. Measured on `quatico/quaweb-website` at Plot 2.19.0: `budget.tsv` holds
|
|
326
|
+
# 312589 lines across 17.6 MB, one read of it takes **516 ms** against 5 ms over
|
|
327
|
+
# fifty lines, and a single `pr-list` calls this three times for the same answer.
|
|
328
|
+
# The file is append-only and a process that asks twice is asking about the same
|
|
329
|
+
# window, so the second scan buys nothing a caller can observe.
|
|
330
|
+
#
|
|
331
|
+
# THE MEMO IS A FILE, AND A VARIABLE ALONE COULD NOT HAVE WORKED. Every caller
|
|
332
|
+
# in `plot-host.sh` writes `rate="$(budget_rate ...)"`, and a command
|
|
333
|
+
# substitution is a SUBSHELL: a variable the function sets inside one dies when
|
|
334
|
+
# the substitution closes. Measured on this branch with a counter — three
|
|
335
|
+
# substituted calls on one key scanned the ledger 3 times with a variable memo
|
|
336
|
+
# in place, against 1 for three direct calls. A variable memo is not slow at the
|
|
337
|
+
# call sites that exist, it is ABSENT from them, and every behavioural test
|
|
338
|
+
# still passes because the three answers are identical.
|
|
339
|
+
#
|
|
340
|
+
# `$$`, NOT `BASHPID`, IS THE SCOPE. `$$` is the invoking shell's pid and is
|
|
341
|
+
# deliberately NOT updated inside a subshell, so it names the process TREE —
|
|
342
|
+
# which is exactly "the life of the process" this memo is bounded by. `BASHPID`
|
|
343
|
+
# tracks the fork and would give every substitution its own empty cache, which
|
|
344
|
+
# is the variable memo's defect with an extra file write.
|
|
345
|
+
#
|
|
346
|
+
# THE COST IS PAID BACK ABOUT TWO THOUSAND TIMES. Measured here: a builtin
|
|
347
|
+
# `read` of one memo line is 82 us and a write 153 us, against 188 ms for one
|
|
348
|
+
# scan of a 40000-line ledger and the 516 ms reported above. The read is a
|
|
349
|
+
# builtin with no fork, which is what keeps it in that range.
|
|
350
|
+
#
|
|
351
|
+
# PER PROCESS IS THE WHOLE BOUND. There is no TTL, no `stat` of the record and
|
|
352
|
+
# no invalidation hook, because a `plot-host.sh` invocation is short-lived and a
|
|
353
|
+
# memo that tried to notice the file moving would re-`stat` on every call — the
|
|
354
|
+
# cost this exists to remove, re-introduced in a smaller form. `budget_append`
|
|
355
|
+
# writes between two reads and the second read is deliberately served the first
|
|
356
|
+
# one's answer.
|
|
357
|
+
#
|
|
358
|
+
# THE KEY IS THE TRIPLE, NEVER ONE FIELD OF IT. An EMPTY bucket means *every
|
|
359
|
+
# bucket* and is a different question from any named one — `budget_rate_read`
|
|
360
|
+
# says so in its own words above, and the two live in one process: measured with
|
|
361
|
+
# a probe, one `pr-list` asks `(github,jwloka,'')` for the concurrency bound and
|
|
362
|
+
# `(github,jwloka,graphql)` for the transport choice. A memo keyed on less than
|
|
363
|
+
# the triple answers the first with the second's reading.
|
|
364
|
+
#
|
|
365
|
+
# AN EXPLICIT `now` BYPASSES THE MEMO ENTIRELY, and a reader will ask why. It is
|
|
366
|
+
# a FOURTH question rather than a fourth key: a caller naming a moment is asking
|
|
367
|
+
# what the record looked like THEN, and every window boundary above is computed
|
|
368
|
+
# from it. Keying on it would make every lookup a miss, because the callers that
|
|
369
|
+
# omit it get `budget_now_ms()` and differ by milliseconds — a memo that is dead
|
|
370
|
+
# code and still passes every behavioural test. Serving a memo across different
|
|
371
|
+
# `now` values would answer a named moment with another one's window. Only the
|
|
372
|
+
# callers that omit it — every one in `plot-host.sh` — are memoised.
|
|
373
|
+
#
|
|
374
|
+
# NOT-YET-COMPUTED AND COMPUTED-TO-EMPTY ARE TWO STATES, and the FILE'S
|
|
375
|
+
# EXISTENCE is what separates them. The zero object is a well-formed answer for
|
|
376
|
+
# a missing record and for a `budget_path` that failed, so it is CACHED like any
|
|
377
|
+
# other — a memo that treated it as no-answer-worth-keeping would restore the
|
|
378
|
+
# scan on exactly the machines with nothing to scan. The entry is never empty:
|
|
379
|
+
# `budget_rate_read` prints a JSON object on every path, so a zero-byte entry
|
|
380
|
+
# means a torn write and is re-read rather than served.
|
|
381
|
+
#
|
|
382
|
+
# PUBLISHED BY `mv`, NEVER BY `>`, for `budget_slot_acquire`'s reason one
|
|
383
|
+
# paragraph down: a redirect creates the NAME before the CONTENT, so a sibling
|
|
384
|
+
# subshell can open the file and read half an answer. The rename is atomic
|
|
385
|
+
# within a directory, so the name and the answer arrive together.
|
|
386
|
+
budget_rate() {
|
|
387
|
+
local connector="${1:-}" account="${2:-}" bucket="${3:-}" now="${4:-}"
|
|
388
|
+
|
|
389
|
+
# A caller that named a moment is asking a different question. Straight
|
|
390
|
+
# through, neither read nor written.
|
|
391
|
+
if [ -n "$now" ]; then
|
|
392
|
+
budget_rate_read "$connector" "$account" "$bucket" "$now"
|
|
393
|
+
return $?
|
|
394
|
+
fi
|
|
395
|
+
|
|
396
|
+
local key slot
|
|
397
|
+
key="${connector}|${account}|${bucket}"
|
|
398
|
+
# Every character a variable name may not hold becomes `_`. Two distinct
|
|
399
|
+
# triples could collide only by differing in punctuation alone, which no
|
|
400
|
+
# connector, account or bucket name does.
|
|
401
|
+
slot="_budget_rate_memo_$(printf '%s' "$key" | LC_ALL=C tr -c '[:alnum:]_' '_')"
|
|
402
|
+
|
|
403
|
+
# TIER ONE, FREE: the same shell asking twice. SET, NOT NON-EMPTY — the zero
|
|
404
|
+
# object is a real answer and an empty one is a state this memo never stores.
|
|
405
|
+
if eval "[ -n \"\${${slot}+set}\" ]"; then
|
|
406
|
+
eval "printf '%s\\n' \"\${${slot}}\""
|
|
407
|
+
return 0
|
|
408
|
+
fi
|
|
409
|
+
|
|
410
|
+
# TIER TWO, 82 us: a subshell asking what its parent already asked. This is
|
|
411
|
+
# the tier that fires at every call site `plot-host.sh` actually has.
|
|
412
|
+
local dir file answer
|
|
413
|
+
dir="$(budget_memo_dir)" || dir=''
|
|
414
|
+
if [ -n "$dir" ]; then
|
|
415
|
+
file="$dir/$slot"
|
|
416
|
+
# `-s`, NOT `-f`. A zero-byte entry is a torn write, never an answer.
|
|
417
|
+
if [ -s "$file" ]; then
|
|
418
|
+
IFS= read -r answer < "$file" 2>/dev/null || answer=''
|
|
419
|
+
if [ -n "$answer" ]; then
|
|
420
|
+
eval "${slot}=\$answer"
|
|
421
|
+
printf '%s\n' "$answer"
|
|
422
|
+
return 0
|
|
423
|
+
fi
|
|
424
|
+
fi
|
|
425
|
+
fi
|
|
426
|
+
|
|
427
|
+
local rc
|
|
428
|
+
# THE EXIT CODE IS THE READ'S, never this wrapper's. `budget_rate_read`
|
|
429
|
+
# returns 0 on every path today and its callers test the CONTENT, so a memo
|
|
430
|
+
# that invented an exit status would be the one place the two disagree.
|
|
431
|
+
answer="$(budget_rate_read "$connector" "$account" "$bucket")"; rc=$?
|
|
432
|
+
if [ "$rc" -ne 0 ]; then
|
|
433
|
+
# A read that failed is not an answer, so nothing is remembered: the next
|
|
434
|
+
# caller asks again rather than inheriting a failure for the process life.
|
|
435
|
+
printf '%s\n' "$answer"
|
|
436
|
+
return "$rc"
|
|
437
|
+
fi
|
|
438
|
+
eval "${slot}=\$answer"
|
|
439
|
+
# NEVER FAILS ITS CALLER, `budget_append`'s rule for the same reason: the memo
|
|
440
|
+
# is an optimisation beside an answer that is already correct, so a cache that
|
|
441
|
+
# cannot be written must not turn a good reading into a failed one.
|
|
442
|
+
if [ -n "$dir" ] && mkdir -p "$dir" 2>/dev/null; then
|
|
443
|
+
if printf '%s\n' "$answer" >"$file.$BASHPID.tmp" 2>/dev/null; then
|
|
444
|
+
mv -f "$file.$BASHPID.tmp" "$file" 2>/dev/null || rm -f "$file.$BASHPID.tmp" 2>/dev/null || true
|
|
445
|
+
fi
|
|
446
|
+
fi
|
|
447
|
+
printf '%s\n' "$answer"
|
|
448
|
+
return 0
|
|
449
|
+
}
|
|
450
|
+
|
|
451
|
+
# Where THIS PROCESS's memoised rates live, and it is deliberately not beside
|
|
452
|
+
# the record. `$PLOT_BUDGET_HOME/memo/<pid>` — same override as the record and
|
|
453
|
+
# the slots, so a test pointing `PLOT_BUDGET_HOME` at a sandbox gets a sandboxed
|
|
454
|
+
# cache too, and a `$PLOT_BUDGET_HOME` that cannot be resolved means no cache
|
|
455
|
+
# rather than a cache in the wrong place.
|
|
456
|
+
#
|
|
457
|
+
# `$$` IS THE DIRECTORY NAME because it is the one identifier that is stable
|
|
458
|
+
# across a command substitution — see `budget_rate` above, where that property
|
|
459
|
+
# is the whole reason this file exists. A pid is reused by the kernel after the
|
|
460
|
+
# process ends, so an entry could in principle be inherited by a later,
|
|
461
|
+
# unrelated process holding the same pid. That is bounded by `budget_memo_clear`
|
|
462
|
+
# below, which the adapter calls on exit, and it is why the cache holds a
|
|
463
|
+
# DERIVED reading rather than anything a caller could act on irreversibly: the
|
|
464
|
+
# worst case is one stale rate, which is the same staleness the memo grants
|
|
465
|
+
# within a process by design.
|
|
466
|
+
budget_memo_dir() {
|
|
467
|
+
local home="${PLOT_BUDGET_HOME:-}"
|
|
468
|
+
if [ -z "$home" ]; then
|
|
469
|
+
[ -n "${HOME:-}" ] || return 1
|
|
470
|
+
home="$HOME/.plot/state"
|
|
471
|
+
fi
|
|
472
|
+
printf '%s\n' "$home/memo/$$"
|
|
473
|
+
}
|
|
474
|
+
|
|
475
|
+
# Removes this process's memo directory. Called on exit by the adapter, so a
|
|
476
|
+
# long-lived machine does not accumulate one directory per `plot-host.sh` call.
|
|
477
|
+
#
|
|
478
|
+
# NEVER FAILS ITS CALLER. It runs in a trap beside work that has already
|
|
479
|
+
# happened, and a cache that cannot be cleared must not change an exit status.
|
|
480
|
+
budget_memo_clear() {
|
|
481
|
+
local dir
|
|
482
|
+
dir="$(budget_memo_dir)" || return 0
|
|
483
|
+
case "$dir" in
|
|
484
|
+
*/memo/[0-9]*) rm -rf "$dir" 2>/dev/null || true ;;
|
|
485
|
+
esac
|
|
486
|
+
return 0
|
|
487
|
+
}
|
|
488
|
+
|
|
324
489
|
# ── The concurrency bound ────────────────────────────────────────────────────
|
|
325
490
|
#
|
|
326
491
|
# HOW MANY CALLS THIS ACCOUNT HAS OPEN AT ONCE, bounded across PROCESSES. The
|
package/plot-deliver.sh
CHANGED
|
@@ -425,9 +425,15 @@ decide_transition() { # $1=file → prints "<Phase>\t<record>\t<write|already>"
|
|
|
425
425
|
# carrying BOTH front matter and a `## Status` block was delivered in a project
|
|
426
426
|
# repo: `flip_phase` wrote `Delivered` into the block, `mv` landed it, the
|
|
427
427
|
# summary said `phase=flipped`, and `plot-plan-meta.sh` went on answering
|
|
428
|
-
# `approved` — because it
|
|
429
|
-
#
|
|
430
|
-
#
|
|
428
|
+
# `approved` — because it preferred front matter wherever it existed. The write
|
|
429
|
+
# took effect on bytes nobody reads.
|
|
430
|
+
#
|
|
431
|
+
# THAT PRECEDENCE INVERTED IN #933, so the two-record plan now delivers rather
|
|
432
|
+
# than refusing here: the parser reads the `## Status` block, which is the field
|
|
433
|
+
# every lifecycle script writes. This gate is unchanged and is not softened —
|
|
434
|
+
# its condition simply stops holding for that shape. It still fires on a scratch
|
|
435
|
+
# copy the parser cannot read, and it still asks the parser rather than trusting
|
|
436
|
+
# that awk changed a line.
|
|
431
437
|
#
|
|
432
438
|
# `flip_phase`'s awk matches only inside `section == "status"`. That one guard
|
|
433
439
|
# IS the defect: on a front-matter plan it edits the block and leaves the front
|
|
@@ -468,11 +474,11 @@ phase_would_read() { # $1=scratch file $2=expected phase (lowercase) → 0 agree
|
|
|
468
474
|
# holding two records of one fact is the thing to fix — and which format ought
|
|
469
475
|
# to win is a decision this gate deliberately leaves to a person.
|
|
470
476
|
echo "plot-deliver: $rel — wrote phase '$want', but the parser still reads '$got'." >&2
|
|
471
|
-
echo " The
|
|
472
|
-
echo "
|
|
473
|
-
echo "
|
|
474
|
-
echo "
|
|
475
|
-
echo "
|
|
477
|
+
echo " The write landed and the parser reads something else, so the delivery" >&2
|
|
478
|
+
echo " would have reported a success it did not achieve." >&2
|
|
479
|
+
echo " Nothing was written — the plan is unchanged. Check what the plan says" >&2
|
|
480
|
+
echo " its phase is, and where: a plan stating it in two places reports the" >&2
|
|
481
|
+
echo " '## Status' block, which is the field every lifecycle script writes." >&2
|
|
476
482
|
echo " See what the parser reads: $script_dir/plot-plan-meta.sh $rel" >&2
|
|
477
483
|
return 1
|
|
478
484
|
}
|
package/plot-dispatch.sh
CHANGED
|
@@ -1904,62 +1904,70 @@ EOF
|
|
|
1904
1904
|
esac
|
|
1905
1905
|
fi
|
|
1906
1906
|
|
|
1907
|
-
# THE COUNT IS THE RULE'S, and the rule is
|
|
1908
|
-
# fleet-size.
|
|
1909
|
-
#
|
|
1910
|
-
#
|
|
1911
|
-
#
|
|
1912
|
-
#
|
|
1913
|
-
#
|
|
1914
|
-
#
|
|
1915
|
-
#
|
|
1916
|
-
#
|
|
1907
|
+
# THE COUNT IS THE RULE'S, and the rule is asked through its BUNDLE —
|
|
1908
|
+
# `board/plot-fleet-size.mjs`, tracked in git beside the other 24. There is no
|
|
1909
|
+
# second copy of the default, the subtraction or the machine's veto living in
|
|
1910
|
+
# shell.
|
|
1911
|
+
#
|
|
1912
|
+
# A SOURCE IMPORT CANNOT REACH A PLUGIN INSTALL, and this block used to be
|
|
1913
|
+
# one. It imported `rules/fleet-size.ts` and `entities/machine.ts` as `file://`
|
|
1914
|
+
# sources. Node 24 strips types, so the TypeScript was never the obstacle — the
|
|
1915
|
+
# SECOND import is: `machine.ts` opens with `import { z } from 'zod'`, and an
|
|
1916
|
+
# install carrying no `node_modules` cannot resolve it. Measured 2026-09-17
|
|
1917
|
+
# against a copy with no `node_modules` on the path:
|
|
1918
|
+
#
|
|
1919
|
+
# machine.ts FAILED: Cannot find package 'zod'
|
|
1920
|
+
# fleet-size.ts: imported
|
|
1921
|
+
#
|
|
1922
|
+
# So the bundle carries BOTH rules with `zod` bundled in. `a-shell-script-asks
|
|
1923
|
+
# -the-domain` settled the shape: a bundle under `skills/plot/scripts/board/`
|
|
1924
|
+
# is how a shell script reaches a rule, and a skill's own script directory is
|
|
1925
|
+
# what a plugin ships.
|
|
1917
1926
|
#
|
|
1918
1927
|
# TWO MODULES, BECAUSE THE VERDICT AND THE COUNT ARE TWO RULES. `headroomFor`
|
|
1919
1928
|
# owns what a fork cost MEANS and `fleetSize` owns what to do about it; the
|
|
1920
1929
|
# count rule takes the verdict as a reading rather than deriving it, so the
|
|
1921
|
-
# thresholds have exactly one home and
|
|
1922
|
-
#
|
|
1923
|
-
#
|
|
1924
|
-
#
|
|
1925
|
-
#
|
|
1926
|
-
#
|
|
1927
|
-
|
|
1928
|
-
|
|
1929
|
-
|
|
1930
|
-
PLOT_COST="$start_cost" PLOT_RULE="$start_rule" \
|
|
1931
|
-
PLOT_MACHINE="file://$start_domain/entities/machine.ts" \
|
|
1932
|
-
node --input-type=module - <<'NODE_EOF' 2>/dev/null
|
|
1933
|
-
const { fleetSize, DEFAULT_FLEET_SIZE } = await import(process.env.PLOT_RULE);
|
|
1934
|
-
const { headroomFor } = await import(process.env.PLOT_MACHINE);
|
|
1935
|
-
|
|
1936
|
-
// AN ABSENT COUNT IS THE RULE'S DEFAULT, resolved here rather than in the
|
|
1937
|
-
// shell: the number and the argument for it have one home.
|
|
1938
|
-
const requested =
|
|
1939
|
-
process.env.PLOT_REQUESTED === "" ? DEFAULT_FLEET_SIZE : Number(process.env.PLOT_REQUESTED);
|
|
1940
|
-
|
|
1941
|
-
// An UNMEASURED cost is null, never zero: zero is the fastest fork there is and
|
|
1942
|
-
// would read as the clearest possible machine.
|
|
1943
|
-
const spawnCostMs = process.env.PLOT_COST === "" ? null : Number(process.env.PLOT_COST);
|
|
1944
|
-
|
|
1945
|
-
const answer = fleetSize({
|
|
1946
|
-
requested,
|
|
1947
|
-
running: Number(process.env.PLOT_RUNNING),
|
|
1948
|
-
spawnCostMs,
|
|
1949
|
-
headroom: headroomFor(spawnCostMs),
|
|
1950
|
-
});
|
|
1951
|
-
|
|
1952
|
-
process.stdout.write(`${answer.start}\t${answer.headroom}\t${answer.shortfall}`);
|
|
1953
|
-
NODE_EOF
|
|
1954
|
-
)
|
|
1930
|
+
# thresholds have exactly one home and the bundle's entry is the join.
|
|
1931
|
+
#
|
|
1932
|
+
# A RULE THAT CANNOT BE ASKED STARTS NOTHING AND SAYS SO. A missing bundle, a
|
|
1933
|
+
# missing node, a module that throws all leave the answer empty. The direction
|
|
1934
|
+
# is the reaper's: silence is never permission, and here permission would spawn
|
|
1935
|
+
# detached processes.
|
|
1936
|
+
start_bundle="$script_dir/board/plot-fleet-size.mjs"
|
|
1937
|
+
start_answer=$(printf '%s\t%s\t%s' "$start_count" "$start_running" "$start_cost" \
|
|
1938
|
+
| node "$start_bundle" 2>/dev/null)
|
|
1955
1939
|
|
|
1956
1940
|
if [ -z "$start_answer" ]; then
|
|
1941
|
+
# THE REFUSAL NAMES THE CONDITION THAT FAILED, never two that hold. The
|
|
1942
|
+
# message this replaced said *"it needs node 24 and a readable checkout of
|
|
1943
|
+
# packages/domain"* to an operator whose node was 24.4.1 and whose checkout
|
|
1944
|
+
# was readable — the import it could not resolve was named nowhere. A
|
|
1945
|
+
# refusal that names the wrong condition costs more than one that says
|
|
1946
|
+
# nothing, because it looks actionable. So the two causes are separated and
|
|
1947
|
+
# tested in the order that distinguishes them: an absent bundle is a broken
|
|
1948
|
+
# or partial installation, and a present bundle that answered nothing is the
|
|
1949
|
+
# runtime underneath it.
|
|
1957
1950
|
echo "plot-dispatch: --start could not ask how many agents to start — starting none." >&2
|
|
1958
|
-
|
|
1959
|
-
|
|
1951
|
+
if [ ! -f "$start_bundle" ]; then
|
|
1952
|
+
echo " The rule's bundle is missing: $start_bundle" >&2
|
|
1953
|
+
echo " Every bundle is tracked in git, so this is a broken or partial installation." >&2
|
|
1954
|
+
echo " In a development checkout, run 'pnpm build:board'." >&2
|
|
1955
|
+
else
|
|
1956
|
+
start_node_v="$(node --version 2>/dev/null)" || start_node_v=""
|
|
1957
|
+
echo " The rule's bundle is $start_bundle" >&2
|
|
1958
|
+
if [ -z "$start_node_v" ]; then
|
|
1959
|
+
echo " No usable 'node' was found on PATH. The bundle needs node 20 or newer." >&2
|
|
1960
|
+
else
|
|
1961
|
+
echo " The bundle is present but answered nothing under node $start_node_v." >&2
|
|
1962
|
+
echo " Run it directly to see why: printf '%s\\t%s\\t%s' '$start_count' '$start_running' '$start_cost' | node '$start_bundle'" >&2
|
|
1963
|
+
fi
|
|
1964
|
+
fi
|
|
1960
1965
|
exit 1
|
|
1961
1966
|
fi
|
|
1962
1967
|
|
|
1968
|
+
# THREE FIELDS, AND THE SENTENCE IS LAST. `start_why` is printed to the
|
|
1969
|
+
# operator below, so it travels; taking it as the whole remainder means a
|
|
1970
|
+
# shortfall can never be truncated by its own punctuation.
|
|
1963
1971
|
start_n=${start_answer%%$'\t'*}
|
|
1964
1972
|
start_rest=${start_answer#*$'\t'}
|
|
1965
1973
|
start_headroom=${start_rest%%$'\t'*}
|
package/plot-fleet-scan.sh
CHANGED
|
@@ -89,6 +89,9 @@
|
|
|
89
89
|
# Output: per-plan wave report on stdout, terminated by a machine-countable
|
|
90
90
|
# summary line:
|
|
91
91
|
# summary: plans=1 waves=3 branches=5 claimed=1 eligible=2 blocked=1 deferred=1 waiting=1 prereq_missing=0 merge_detect=pr-merge host=ok main=main
|
|
92
|
+
# `host` is one of ok, partial, throttled, secondary, failed, unasked —
|
|
93
|
+
# `partial` means some of the host's states answered and some did not,
|
|
94
|
+
# so the PR readings below are incomplete rather than absent.
|
|
92
95
|
# `blocked` counts WAVES an earlier wave holds; `waiting` and
|
|
93
96
|
# `prereq_missing` count BRANCHES their `waits:` annotation holds.
|
|
94
97
|
# merge_detect names how merged-and-deleted branches were detected:
|
|
@@ -631,6 +634,19 @@ PR_LIST_LIMIT="${PLOT_PR_LIST_LIMIT:-1000}"
|
|
|
631
634
|
# secondary — a burst refusal (`plot-host.sh` exit 6). Nothing is broken
|
|
632
635
|
# either, and it clears in seconds rather than minutes.
|
|
633
636
|
# failed — any other failure (exit 3, or anything unclassified).
|
|
637
|
+
# partial — SOME of the host's states answered and some did not
|
|
638
|
+
# (`plot-host.sh` exit 7). The rows that arrived are real and
|
|
639
|
+
# are parsed; what is missing is a whole state, so the reading
|
|
640
|
+
# is incomplete rather than absent. Only Bitbucket can produce
|
|
641
|
+
# it: `bb pr list` has no `all` state, so the arm asks once per
|
|
642
|
+
# state, while GitHub takes `--state all` in one call.
|
|
643
|
+
#
|
|
644
|
+
# IT IS NOT `ok` AND IT IS NOT `failed`. Reporting `ok` would
|
|
645
|
+
# serve a page missing a state as a complete answer — #912, where
|
|
646
|
+
# nine branches read `commits, no PR ever opened` and two had
|
|
647
|
+
# live PRs. Reporting `failed` would throw away rows that
|
|
648
|
+
# arrived and make every branch `unknown`, which is not
|
|
649
|
+
# startable — the right refusal about the wrong thing.
|
|
634
650
|
# unasked — no host to ask, or --offline WITHOUT `--next`. Not a
|
|
635
651
|
# degradation: the scan was never going to ask, and saying
|
|
636
652
|
# `failed` would report a fault where there is a configuration.
|
|
@@ -644,6 +660,43 @@ PR_LIST_LIMIT="${PLOT_PR_LIST_LIMIT:-1000}"
|
|
|
644
660
|
# secondary limit waits minutes for a ceiling that cleared in seconds.
|
|
645
661
|
HOST_VERDICT=unasked
|
|
646
662
|
|
|
663
|
+
# The remote branches this scan tracks, read once for every question that asks.
|
|
664
|
+
#
|
|
665
|
+
# MOVED UP FROM ITS OLD POSITION (#333) so the host call below can be told which
|
|
666
|
+
# branches to ask about. It is the same single `for-each-ref` it always was —
|
|
667
|
+
# see the commentary at its old site — and the reasons it exists are unchanged:
|
|
668
|
+
# `git show-ref --verify` was asked once per branch from two places, and
|
|
669
|
+
# `%(objectname)` rides along free for the commit walk.
|
|
670
|
+
REMOTE_REFS=$(git for-each-ref --format='%(refname:strip=3)%09%(objectname)' \
|
|
671
|
+
"refs/remotes/origin" </dev/null 2>/dev/null)
|
|
672
|
+
|
|
673
|
+
# The branch names alone, space-separated — what the sweep asks the host about.
|
|
674
|
+
#
|
|
675
|
+
# THE JOIN'S OWN KEYS, AND NOTHING WIDER. `prefill_pr_states` indexes the host's
|
|
676
|
+
# reply by branch and every row it cannot key is discarded, so the set this asks
|
|
677
|
+
# about is exactly the set that could ever be used. Measured 2026-09-20 on
|
|
678
|
+
# `quatico/quaweb-website`: 11 remote branches against 902 pull requests, of
|
|
679
|
+
# which a listing hands over 50 — the sweep asks 11 questions and gets 11
|
|
680
|
+
# answers, where the listing asked one and answered for 5%.
|
|
681
|
+
#
|
|
682
|
+
# `HEAD` IS DROPPED. `refs/remotes/origin/HEAD` is a symbolic ref naming the
|
|
683
|
+
# default branch, not a branch of its own; asking the host about a branch called
|
|
684
|
+
# `HEAD` spends a query to learn that nothing is named that.
|
|
685
|
+
#
|
|
686
|
+
# SPACE-SEPARATED, AND GIT IS WHAT MAKES THAT SAFE. Both this list and the
|
|
687
|
+
# adapter's `PR_LIST_BRANCHES` are read by an unquoted `for`, so a name carrying
|
|
688
|
+
# whitespace would split into two branches that do not exist. `git
|
|
689
|
+
# check-ref-format` REFUSES a ref name containing a space or a tab — verified
|
|
690
|
+
# 2026-09-20, both exit non-zero — so no such branch can reach this, and the
|
|
691
|
+
# separator is git's guarantee rather than a hopeful convention.
|
|
692
|
+
#
|
|
693
|
+
# EMPTY IS A REAL ANSWER AND IT DISABLES THE SWEEP. A checkout with no remote
|
|
694
|
+
# refs has no branches to ask about, and a sweep of nothing would state that
|
|
695
|
+
# every tracked branch answered — a completeness claim over an empty set, which
|
|
696
|
+
# would license `NONE` for branches nobody asked about. `prefill_pr_states`
|
|
697
|
+
# falls back to the listing there, which is what it has always done.
|
|
698
|
+
TRACKED_BRANCHES=$(printf '%s\n' "$REMOTE_REFS" | cut -f1 | grep -v '^HEAD$' | grep -v '^$' | tr '\n' ' ')
|
|
699
|
+
|
|
647
700
|
prefill_pr_states() {
|
|
648
701
|
[ "$HOST_LOOKUP_OK" = 1 ] || return 0
|
|
649
702
|
[ -n "$HOST_STATE_CACHE" ] || return 0
|
|
@@ -672,10 +725,34 @@ prefill_pr_states() {
|
|
|
672
725
|
# and /dev/null keeps the call working with the text simply unavailable.
|
|
673
726
|
host_list_out="${HOST_STATE_CACHE:+$HOST_STATE_CACHE/pr-list.json}"
|
|
674
727
|
host_list_out="${host_list_out:-/dev/null}"
|
|
728
|
+
# THE BRANCHES THIS SCAN TRACKS, HANDED TO THE HOST (#333). The adapter uses
|
|
729
|
+
# them only where it can — the Bitbucket arm sweeps its REST endpoint once per
|
|
730
|
+
# branch per state — and ignores them everywhere else, so the GitHub arm makes
|
|
731
|
+
# the single call it always made. Passing them unconditionally keeps one call
|
|
732
|
+
# shape here rather than a backend test this script has no business making.
|
|
733
|
+
#
|
|
734
|
+
# AN EMPTY SET PASSES NOTHING and the adapter lists as before. See
|
|
735
|
+
# `TRACKED_BRANCHES`: a completeness claim over an empty set would license
|
|
736
|
+
# `NONE` for branches nobody asked about.
|
|
737
|
+
_branch_args=()
|
|
738
|
+
for _tb in $TRACKED_BRANCHES; do _branch_args+=(--branch "$_tb"); done
|
|
675
739
|
host_err=$("$script_dir/plot-host.sh" pr-list --state all --limit "$PR_LIST_LIMIT" --rich \
|
|
740
|
+
${_branch_args[@]+"${_branch_args[@]}"} \
|
|
676
741
|
</dev/null 2>&1 >"$host_list_out"); rc=$?
|
|
677
742
|
js=$(cat "$host_list_out" 2>/dev/null)
|
|
678
|
-
|
|
743
|
+
# A PARTIAL ANSWER TAKES THE PARSE PATH AND STILL DEGRADES THE VERDICT, which
|
|
744
|
+
# is a control-flow change rather than another `case` arm below: every other
|
|
745
|
+
# non-zero rc sets a verdict and returns BEFORE `$js` is read, because there
|
|
746
|
+
# is nothing to read. Here there is — the rows of the states that answered are
|
|
747
|
+
# already in `host_list_out`, since stdout is redirected to a file and stderr
|
|
748
|
+
# captured separately.
|
|
749
|
+
#
|
|
750
|
+
# BOTH HALVES ARE REQUIRED. Falling through without setting the verdict would
|
|
751
|
+
# report `ok` over a page missing a whole state, which is #912; returning
|
|
752
|
+
# early would throw away rows the host did answer with.
|
|
753
|
+
if [ "$rc" -eq 7 ]; then
|
|
754
|
+
HOST_VERDICT=partial
|
|
755
|
+
elif [ "$rc" -ne 0 ]; then
|
|
679
756
|
# THREE OUTCOMES, NOT TWO. `unasked` already means "the question was
|
|
680
757
|
# never put" (see HOST_VERDICT above: *not a degradation, the scan was
|
|
681
758
|
# never asking*), and a host that cannot be ASKED AT ALL belongs there
|
|
@@ -743,7 +820,12 @@ prefill_pr_states() {
|
|
|
743
820
|
return 0
|
|
744
821
|
fi
|
|
745
822
|
# The list arrived. An empty one arrived too — that is the whole distinction.
|
|
746
|
-
|
|
823
|
+
#
|
|
824
|
+
# A PARTIAL VERDICT IS NOT OVERWRITTEN HERE. Exit 7 reaches this line
|
|
825
|
+
# deliberately, because its rows must be parsed; an unguarded `ok` would
|
|
826
|
+
# undo the one thing that distinguishes an incomplete page from a whole one
|
|
827
|
+
# and report #912 as a healthy reading.
|
|
828
|
+
[ "$HOST_VERDICT" = partial ] || HOST_VERDICT=ok
|
|
747
829
|
# `pr-list` emits one compact JSON object per line. PARSED IN ONE PASS, and
|
|
748
830
|
# that is a correctness-of-cost property rather than a style preference:
|
|
749
831
|
# measured 2026-08-18 on this repo's 221 PRs, a `sed` per field per row —
|
|
@@ -864,9 +946,39 @@ EOF
|
|
|
864
946
|
# A repository genuinely holding zero PRs loses nothing by being asked: it has
|
|
865
947
|
# no branches with PRs for the join to serve either, so the cost is zero calls
|
|
866
948
|
# in both readings.
|
|
867
|
-
|
|
868
|
-
|
|
869
|
-
|
|
949
|
+
#
|
|
950
|
+
# A SWEEP STATES ITS COMPLETENESS; A PAGE ONLY EVER IMPLIED IT (#333). The
|
|
951
|
+
# test above reads completeness off one page's size, which is the only
|
|
952
|
+
# evidence a listing offers — and on Bitbucket it is evidence the listing
|
|
953
|
+
# cannot give at all, since `bb pr list` returns a fixed 50 whether or not
|
|
954
|
+
# more exist. Where the adapter swept per branch it says so on stderr, naming
|
|
955
|
+
# both counts, and that sentence is a stronger claim than any row count: every
|
|
956
|
+
# tracked branch was asked and each one answered.
|
|
957
|
+
#
|
|
958
|
+
# THE ROW COUNT IS NOT CONSULTED ON THAT PATH, and it must not be. A sweep
|
|
959
|
+
# over 11 branches of which 2 have pull requests emits 2 rows — a true and
|
|
960
|
+
# complete answer that `0 < rows < PR_LIST_LIMIT` would also accept, but for
|
|
961
|
+
# the wrong reason, and which a sweep of 0 matches would fail outright despite
|
|
962
|
+
# being equally complete. Reading the claim the adapter made is exact where
|
|
963
|
+
# re-deriving it from the output is a coincidence.
|
|
964
|
+
#
|
|
965
|
+
# THE WORDING IS A CONTRACT between this script and `plot-host.sh`'s
|
|
966
|
+
# `pr_sweep_report`, pinned on both sides. A partial sweep never prints it, so
|
|
967
|
+
# a match is licence and a miss is silence — never a guess.
|
|
968
|
+
#
|
|
969
|
+
# WHAT IS LOST BY GETTING THIS WRONG IS COST, NOT CORRECTNESS. Without the
|
|
970
|
+
# marker, `host_pr_state --ask` falls through to one `pr-state` call per
|
|
971
|
+
# unjoined branch and still answers correctly — the per-branch N+1 that #216
|
|
972
|
+
# removed. That is why withholding the marker is always the safe direction and
|
|
973
|
+
# is what every failure path here does.
|
|
974
|
+
case "$host_err" in
|
|
975
|
+
*"pr-list sweep complete"*)
|
|
976
|
+
printf '1' > "$HOST_STATE_CACHE/.list-complete" 2>/dev/null || true ;;
|
|
977
|
+
*)
|
|
978
|
+
if [ "$_pr_rows" -gt 0 ] && [ "$_pr_rows" -lt "$PR_LIST_LIMIT" ] 2>/dev/null; then
|
|
979
|
+
printf '1' > "$HOST_STATE_CACHE/.list-complete" 2>/dev/null || true
|
|
980
|
+
fi ;;
|
|
981
|
+
esac
|
|
870
982
|
}
|
|
871
983
|
prefill_pr_states
|
|
872
984
|
|
|
@@ -1487,8 +1599,12 @@ worktree_locked() { # $1=worktree path → 0 when a lock is held there
|
|
|
1487
1599
|
# on the population that must stay free. The walk here is a subject/emptiness
|
|
1488
1600
|
# question rather than a timestamp read, but the guard is deliberately broad and
|
|
1489
1601
|
# loosening it to fit this change is how a guard rots.
|
|
1490
|
-
|
|
1491
|
-
|
|
1602
|
+
# READ ABOVE `prefill_pr_states`, not here. The per-branch sweep (#333) hands
|
|
1603
|
+
# the host the branches this scan tracks, and that list is exactly what this
|
|
1604
|
+
# batch already answers — so the assignment moved up rather than a second
|
|
1605
|
+
# `for-each-ref` being added beside it. Everything documented above still
|
|
1606
|
+
# describes it; only the line's position changed, and it depends on nothing but
|
|
1607
|
+
# git, so nothing between the two points can read a different answer.
|
|
1492
1608
|
|
|
1493
1609
|
# Whether `origin/$1` exists, answered from the batch rather than by spawning.
|
|
1494
1610
|
#
|
|
@@ -4265,6 +4381,17 @@ elif [ "$HOST_VERDICT" = failed ]; then
|
|
|
4265
4381
|
echo " branch below reads from local evidence alone, and a branch whose"
|
|
4266
4382
|
echo " PR is unknown reads 'unknown' rather than 'open'. This is not a"
|
|
4267
4383
|
echo " rate limit — waiting will not clear it; check the host and auth."
|
|
4384
|
+
# THE PAGE IS SHORT, NOT ABSENT, and that is a different instruction to a
|
|
4385
|
+
# reader. The three notes above all say *no PR could be read*; here some were,
|
|
4386
|
+
# so the branches below are a MIXTURE — a branch shown without a PR may have one
|
|
4387
|
+
# in the state that failed. Telling a reader to treat this as an outage would
|
|
4388
|
+
# discard the rows that arrived; telling them nothing is #912, where nine
|
|
4389
|
+
# branches read as having no PR and two had live ones.
|
|
4390
|
+
elif [ "$HOST_VERDICT" = partial ]; then
|
|
4391
|
+
echo " note: the git host answered for some states and not others, so the PR"
|
|
4392
|
+
echo " list below is INCOMPLETE. A branch shown without a PR may have one"
|
|
4393
|
+
echo " in the state that failed — do not read this as evidence that a"
|
|
4394
|
+
echo " branch is unreviewed. Re-run to get the whole list."
|
|
4268
4395
|
fi
|
|
4269
4396
|
# A STALE PULSE SAYS SO. The fetch used to fail silently, which made a scan of
|
|
4270
4397
|
# hour-old refs read exactly like a scan of current ones — the same
|