@feigi/fleet-ctl 3.23.23 → 3.23.25

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -10,7 +10,7 @@ argument-hint: [pr-number]
10
10
  See **Specialists** — those constraints are not optional.
11
11
  2. Plan the actions. Split explicitly into **apply now** / **defer**. A finding is deferred, never dropped. **Every `suggestion` — the review's `severity`, the plugin's *Suggestions* bucket — arrives unchecked: the review budgets that band 0 refuters, so nothing has looked at it yet.** Checking it is now your job, not nobody's. So split it by SCOPE first. Out of the PR's ticket scope → defer and file. In scope → dispatch one refuter against it, biased to refuse, and hand it this instruction verbatim with one substitution: resolve `<scratch>/pr<N>/<finding>/` in it to an absolute path under your own run root — `mktemp -d` once per round for `<scratch>/pr<N>/fix-XXXXXXXX/`, never bare `pr<N>`, since finding ids restart at 1 each round too — and write that absolute path into the refuter's prompt in its place; a refuter never derives its own path: *Try to REFUTE this finding. Default to refuted=true if uncertain. Verify by RUNNING something — compile it, run the test, apply the mutation. Do not reason your way to agreement. Observe that run synchronously — run the command, wait for it, read its exit code. Never poll a log file for a completion marker: prefer ONE blocking run to a poll loop, and treat its return as permission to look, never as the answer. Reading a log the run has already finished writing is fine; waiting on one is not. If you match a test reporter's own output, accepting both `ℹ` and `#` is necessary but NOT sufficient — strip SGR escapes first as well. node's prefix moves with the node version and with whether stdout is a TTY, and color wraps the whole line so it begins with ESC and no prefix anchor matches at all, which returns empty at exit 0 — indistinguishable from a hung run and from a run of zero tests. For an uncolored baseline use `env -u FORCE_COLOR`; `FORCE_COLOR=` empty still enables color, so it is not a control. State your search scope AND what your pattern would have missed. A grep over one ref does not support a claim about history; a pattern built from the token a diff removed does not support a claim that the category is empty. A failure injection with no positive control has produced NO result, never a negative one. Before you read an injected fault — an env var, an argv word, a mutant — as having had no effect, prove the injection reached the child: one run whose output differs with it present versus absent, or the child echoing the injected value back. Uncontrolled, the cell is unrun — say so in your verdict instead of reporting a no-effect result. Build such an invocation as an array expanded braced and quoted — `cfg=(SETB=1 BADJ=1); env "${cfg[@]}" sh ./probe.sh` — or inline the assignments literally — `env SETB=1 BADJ=1 sh ./probe.sh`; NEVER from an unquoted scalar — `cfg="SETB=1 BADJ=1"; env $cfg sh ./probe.sh` — which under zsh passes ONE argument, sets a variable literally named `SETB` to `1 BADJ=1`, never sets `BADJ` at all, and still exits 0. `env $cfg[@]` is not the portable spelling either: measured, bash word-splits it into `SETB=1` and `BADJ=1[@]`, so the injection variable is set to a corrupted value, while zsh behaves exactly as with the bare `$cfg` — `BADJ` never set, exit 0. Everything you write — mutants, fixtures, scratch repos — goes under `<scratch>/pr<N>/<finding>/` and nowhere else; the checkout and any worktree are never write targets, though `git show`/`git archive` at a pinned ref read fine anywhere. Chain the directory change into the command, `cd "$D" && git …`, never `cd "$D"; git …`, so a failed `cd` cannot leave a `git` command running in the checkout — and bracket a fixture's own git with `git rev-parse --show-toplevel`: before `git init` it must NOT resolve to the repository, and a fresh scratch dir's `fatal: not a git repository` (exit 128) is the pass, not a failure; before any `git commit` it must resolve to your scratch path — compare resolved forms (`realpath`), since `--show-toplevel` can report `/private/tmp/…` for a `/tmp` scratch dir on macOS.* That last clause is the whole mechanism — a refuter left to reason its way to agreement rubber-stamps, and "apply only what survives" silently degrades into "apply everything in scope". Apply only what survives; a refuted one defers with the refutation in its body. "In scope" means the ticket's scope, not merely the same files — and applying one still never licenses editing a file the ticket has no business in. **A finding in `unverified` whose refuters ran and crashed always defers, whatever its severity** — and which of the two it is, you read off the finding's **`refutersDispatched`**, never off severity and never off an empty vote list: above zero with nothing surviving means every refuter dispatched against it died, and at `critical` that is every one of them. Only a `survived` finding has been checked; severity says how much a finding would matter if true, never whether anything looked. **Deferring a crash is your move, and re-running it is not** — the review already re-dispatched every crashed refuter pair once, in-run, so a crash that reaches you crashed twice, and the result's `resume` field names the one move left, which is not yours: nothing replays a review that ran in a runner's kernel — it is reported, not acted on. So say *in the deferral* that its refuters crashed rather than that it went unchecked, and re-run nothing for it. **A `suggestion` is also in `unverified`, for a different reason — `refutersDispatched` of zero, the 0-refuter budget the workflow gives that band by policy, so nothing looked *there* — and the scope split above, not this rule, covers it.**
12
12
  **Write every ruling to `<scratch>/dispositions-<pr>.json`, beside the review result file — the disposition record `dispositions-check.mjs` judges against it before any finisher is dispatched.** It is `{"head": "<the review file's head>", "entries": [ … ]}`: one entry per `survived` finding and per `unverified` finding, plus one per `refuted` finding you reverse, its evidence in `reason`. Each entry carries `bucket` (`survived|unverified|refuted`), `index` (its position in that bucket of the review file — stable, since a round's file never changes), `scope` (`in|out`), `claimKind` (`behavior|shape`) and `disposition` (`apply|defer`), and where they apply `reason`, `issue` (the number you filed it to or commented on), `verdictPath` (an in-scope suggestion's refuter verdict file) and `remedyFiles` (the files its remedy names). The check reads scope off the diff before it reads your entry: a finding with no `line`, or on a line `git diff` touches from the PR's merge-base with `origin/main` to the review's `head`, is in scope whatever `scope` says, and only otherwise does the declared one stand. An in-scope `survived` finding deferred passes only with `reason` `false-rationale`, `mutual-exclusion`, `remedy-worse` or `remedy-outside-diff`. `remedy-outside-diff` — the remedy would edit a file the PR's diff does not touch — also needs `remedyFiles` to name at least one file absent from `git diff` from that merge-base to the review's `head`: none named, or every one in that diff, is a mismatch. A `critical` or `important` finding deferred for it goes to a human instead of passing: the script exits 1 with verdict `escalate`, `ledger.mjs dispatch` refuses the PR's finisher, and nothing retries it; a `suggestion` deferred for it passes. Any other reason, none, or a `survived` or `unverified` finding with no entry is a mismatch: the script exits 1 naming each violating entry's bucket, index and rule, and `ledger.mjs dispatch` refuses the PR's finisher.
13
- 3. **Run `testCmd` before you commit**, from the worktree, copied verbatim — not a runner you picked. `testCmd` is the command your dispatcher handed you; **standalone, with no dispatcher to hand you one, it is the repo's own test command**, the Test entrypoint in the repo's Recipe cache (`derive-testcmd.sh <repo> test` prints it) — here `node --test plugin/scripts/*.test.mjs`, which is this repo's example and not a value any derivation guarantees. Read it off the repo, never invent one: "not a runner you picked" is what this step forbids, and on the standalone path nothing else defines the referent. With no usable Recipe cache (`derive-testcmd.sh` refuses), run the Recipe derivation step first — dispatch a `task` member, agent `fleet-recipe-deriver`, its prompt naming the main checkout's absolute path as `<repo>` and a scratch directory of its own; it proves both commands in a throwaway worktree and writes the cache only on proof, and its `RECIPE NOT PROVEN` is a stop, reported with its reason. Never fall back to a runner of your own choosing. `tests 0` is a FAILED run, not a pass: a glob matching nothing still exits 0. **Zero is not the only no-work count.** A run reporting **0 passes with no failures** skipped everything it collected, and a count **materially below the whole suite** means you ran a partial tree — this suite reported `tests 937` when this line was written, so compare what you get against a full run rather than against zero. The number drifts; that is the safe direction, because a stale one produces a loud false red while its absence produces a silent false green. Red, zero-test, or either no-work count → fix it or move the finding to defer; never commit over it. Then commit and push — `git push origin HEAD:<branch>`, the explicit refspec: this step also runs from a `--detach` worktree (`run-team/SKILL.md`'s worktree-recreation form), where a bare `git push` exits 128 (`fatal: You are not currently on a branch`) and any push whose destination is `HEAD` rather than a full or branch refname (`-u` or not, forced or not) exits 1 (`error: The destination you provided is not a full refname`) — and report the test result alongside the SHA. **Then reconcile the PR body with what you actually did.** It merges into the permanent record and nothing downstream re-derives it: your commit messages are yours to fix and usually get fixed, the body is neither. Removing a change, or reversing one, leaves the body asserting work the diff no longer contains. Measured twice in one run — one body kept asserting a calibration its own fix had just deleted as false, another kept claiming a hunk the fix had removed as "this diff owns it" — both caught by a finisher's reading rather than by any rule, and both would otherwise have merged intact. `gh pr edit <pr> --body-file <path>`, and re-read the body against the final diff rather than against your memory of it. **Never `--no-verify`.** The fleet runs in repos it knows nothing about: it must not assume a pre-commit hook exists, and must never bypass one that does — a hook that fails is a finding to report, not an obstacle. Nothing re-reviews what you commit here, so this run is the only gate between your edit and a merge.
13
+ 3. **Run `testCmd` before you commit**, from the worktree, copied verbatim — not a runner you picked. `testCmd` is the command your dispatcher handed you; **standalone, with no dispatcher to hand you one, it is the repo's own test command**, the Test entrypoint in the repo's Recipe cache (`derive-testcmd.sh <repo> test` prints it). Read it off the repo, never invent one: "not a runner you picked" is what this step forbids, and on the standalone path nothing else defines the referent. With no usable Recipe cache (`derive-testcmd.sh` refuses), run the Recipe derivation step first — dispatch a `task` member, agent `fleet-recipe-deriver`, its prompt naming the main checkout's absolute path as `<repo>` and a scratch directory of its own; it proves both commands in a throwaway worktree and writes the cache only on proof, and its `RECIPE NOT PROVEN` is a stop, reported with its reason. Never fall back to a runner of your own choosing. `tests 0` is a FAILED run, not a pass: a glob matching nothing still exits 0. **Zero is not the only no-work count.** A run reporting **0 passes with no failures** skipped everything it collected, and a count **materially below the whole suite** means you ran a partial tree — this suite reported `tests 937` when this line was written, so compare what you get against a full run rather than against zero. The number drifts; that is the safe direction, because a stale one produces a loud false red while its absence produces a silent false green. Red, zero-test, or either no-work count → fix it or move the finding to defer; never commit over it. Then commit and push — `git push origin HEAD:<branch>`, the explicit refspec: this step also runs from a `--detach` worktree (`run-team/SKILL.md`'s worktree-recreation form), where a bare `git push` exits 128 (`fatal: You are not currently on a branch`) and any push whose destination is `HEAD` rather than a full or branch refname (`-u` or not, forced or not) exits 1 (`error: The destination you provided is not a full refname`) — and report the test result alongside the SHA. **Then reconcile the PR body with what you actually did.** It merges into the permanent record and nothing downstream re-derives it: your commit messages are yours to fix and usually get fixed, the body is neither. Removing a change, or reversing one, leaves the body asserting work the diff no longer contains. Measured twice in one run — one body kept asserting a calibration its own fix had just deleted as false, another kept claiming a hunk the fix had removed as "this diff owns it" — both caught by a finisher's reading rather than by any rule, and both would otherwise have merged intact. `gh pr edit <pr> --body-file <path>`, and re-read the body against the final diff rather than against your memory of it. **Never `--no-verify`.** The fleet runs in repos it knows nothing about: it must not assume a pre-commit hook exists, and must never bypass one that does — a hook that fails is a finding to report, not an obstacle. Nothing re-reviews what you commit here, so this run is the only gate between your edit and a merge.
14
14
  4. **Under the fleet, do not hold this wait.** The controller owns a persistent CI Monitor and dispatches a finisher once your report has arrived and CI is green — so push, then report to the controller a **final verdict** (every specialist you dispatched + its outcome + your apply/defer adjudication + the pushed SHA) and stop. A specialist that has neither reported nor been relayed to you is not a verdict line — ask the controller for it by name before ruling; see **Specialists**. The controller keys "review complete" on that verdict, **never on you merely going idle** — a slow specialist can still be running when your turn ends, and finishing off idle-alone ships a premature label (observed: a PR labelled while its contract-test specialist was still running and about to surface a vacuous pin). A turn-based reviewer re-reading `gh pr checks` each idle cycle rebuilds its whole context for nothing the Monitor lacks. **Standalone (no controller): watch** the run until the diff-validating jobs are green (bind via `ci-state.mjs` below); fix a **genuine** failure and repeat from 3. **Staleness is not a failure:** a `rebase-check` red or heavy jobs `skipped` **solely from a non-zero behind-count** never clears by waiting and must not be rebased away here — see **`skipped` is not `passed`** and **Do NOT rebase to label** below for what it looks like and why. Stop watching and go to 6. A turn-based reviewer that "waits for green" on a behind PR stalls forever and ends its turn without labelling. **`ci-state.mjs` returns `verdict: "no-ci"`** → there is no run to watch, so don't: go straight to 6 and let that step's `--declare-no-ci` branch decide.
15
15
  5. File each deferred finding: `gh issue create --label <ready-for-agent|needs-triage>` with the finding, `file:line`, why deferred, `Deferred from PR #<pr> review`, and **the finding's `dimension` in the issue body** — the review returns it on every finding, and a backlog that does not carry it cannot be attributed to the specialist that generated it. **The label follows the finding's state, not this step** — but it is never omitted; an unlabelled issue is invisible to both `candidates.mjs --require-label` and the cockpit's pool query, reachable only by a human. **The bar here is whether the DEFECT is confirmed — never whether the remedy is settled.** Confirmed defect with the remedy still open → `ready-for-agent`, and say in the body that the remedy is open. Torn → `needs-triage`, narrowed: torn means **unsure anything is broken**, not unsure how to fix it. **A finding a refuter CHECKED and did not survive is neither** — it is not a confirmed defect, and you are not unsure whether anything is broken: you measured that nothing is. Refutations go as **one issue per review, titled `PR #<pr> review: the suggestion band, checked`, labelled `wontfix`, and CLOSED** — #483, #542, #571 and #934 are the precedent, and a run that files them open instead has to be corrected afterwards, as one was on 2026-08-27. Left open they read as queued agent work and rot into filing artifacts; a measured triage pass found a quarter of its `needs-triage` backlog was refuted records, bulk-closable by title shape. Closed is still a durable home — `check` reads the tracker at `--state all`, so a future run re-deriving the finding lands on the evidence rather than re-litigating it. Split the band only where an entry is **not** a refutation: a residual the refuter surfaced, or a finding refused as out of scope and never checked, is real deferred work — and real deferred work earns its own open issue only past a second bar. **Worth a claim: a finding whose own claim is that correct code could be shaped better — style, naming, layout, redundancy, a micro-simplification — never gets its own open issue.** Its only open question is "is this worth doing?", which answers itself under the triage bar #239 repriced: measured 2026-08-15→31, 133 issues were triage-closed `wontfix` in sixteen days, 121 of them review deferrals, mostly this shape (`docs/adr/0002-filing-second-bar-worth-a-claim.md`). Append it to the same closed record issue under a `Below the claim bar` heading — `file:line`, dimension, the claim, and a re-open trigger per entry. Closed is a durable home for these too: `check` reads `--state all`, so a run re-deriving the finding lands on the record instead of re-filing — and a later review returning the same finding as `survived`, or a maintainer hitting it, is the promotion signal: file it open then, citing the record. **An in-scope rediscovery step 2 applied is that same promotion signal, and it files nothing** — so the record carries no mark unless you write one: run `check` over each finding step 2 applied as well, and where its tracker rows put one on a `Below the claim bar` entry, comment on that record issue, opening `Promoted — applied, not filed`, and name the entry, that its refuter let it through, and the commit carrying the fix. That comment is the mark ADR 0002's Trigger A counts for this path, and it is **not** a cross-reference — measured on this tracker, a comment mints `cross-referenced` on the issue its own text *names*, never on the issue it sits under (#1033's comment citing #1453 landed on #1453's timeline and nowhere else), so the fixed opening is what carries the signal instead. **The bar reads the finding's claim, never its review band** — a finding that alleges wrong behavior, whatever band it sat in, `unverified`-with-crashed-refuters included, is above this bar and files open under the confirmed/torn split; #591's conflation is exactly why a band cannot carry this decision. Do **not** reach for phase 0's **decided?** test here. That is a *consumption* gate, priced for a seat where implementer divergence costs a claim, a worktree and a dispatch; at filing time divergence costs one PR review the maintainer already runs. Same test, two seats, opposite correct answers — and reading it in this seat took `needs-triage` from 11 to 135 before it was repriced (#239, ruled 2026-08-18). **Dedupe through `~/.fleet/bin/fleet-run ledger.mjs check "<subject>"`, never a bare `gh issue list --search`** — the duplicate class that guard exists to catch was filed from *here*: the ledger row for the finding rediscovered five times in one measured run carries a reviewer's own source tag, `(review-pr-108)`. `check` reads this run's filed list *and* the tracker at `--state all`, so a closed duplicate still counts. Exit **1** is already filed in this run; exit **3** is a clean ledger with matching tracker rows that score against the subject — read the ones it prints and comment on the existing follow-up rather than filing a second issue. **Exit 0 is not automatically "safe to file"**: it covers `clean`, `soft-hit` and `unverified` alike, and an unreachable `gh` degrades to a ledger-only answer that says `TRACKER NOT CHECKED`, so read the stdout `verdict` to tell a tracker that was searched and clean from one nobody read — under `unverified`, search it yourself before filing — or, for the applied-promotion comment above, before treating a tracker nobody could read as a reason to skip it. **`soft-hit` is the answer that costs most to misread**: it means related rows were found and none of them was established as a duplicate — a near-miss above the floor, or tracker rows the scorer rates 0.00 — so read the rows it printed and decide, and never treat that exit 0 as permission. Exit **2** is a usage error, meaning nothing was checked at all: fix the invocation and re-run, never read it as a duplicate. Full exit-code and `verdict` semantics live in `$(~/.fleet/bin/fleet-run --root)/skills/run-team/SKILL.md`'s **Run ledger** section — read them there rather than restating them here. **The guard dedupes, it never drops:** a duplicate becomes a comment on the issue that already covers it, never a dropped finding. **Record every issue you create, at once:** immediately after each `gh issue create`, run `~/.fleet/bin/fleet-run ledger.mjs filed <N> "<subject>"` — `<N>` the number `gh issue create` printed, and **the same subject string you passed to `check`**, so the next `check` for that finding meets it in the filed list instead of scoring a rewording. Without it the guard's ledger half never holds a row, and its exit **1** can never fire. If `filed` fails, retry it once; if it fails again, leave the issue alone — it exists and is correct — and report it as `unrecorded: #N <subject>`, so the controller runs `filed` for it. Report `filed: #N <subject>` for every issue you created — under the fleet, in your report to the controller. **The one window the guard cannot see:** `check` → `gh issue create` → `filed` is three commands, and the ledger's lock covers each alone, so two filers can both read `clean` in the seconds before either records. That gap is accepted, never a reason to skip `filed`. Post the numbers as one PR comment.
16
16
  6. Diff-check green (run-bound per `ci-state.mjs`; under staleness see **`skipped` is not `passed`** for what the label then rests on) **and** deferrals filed **and** the PR carrying exactly one release label (`patch`/`minor`/`major`, counting only those three; **required on every PR** — `release-label.yml` auto-adds `patch` when none is present and hard-fails on more than one, so this never needs to check whether the repo defines the gate at all, and `gh label list` settles nothing here any more) **and** `dispositions-check.mjs --member fix-pr-<pr> --scratch <scratch>` run over your own record exiting 0 (standalone that is yours to run, from the worktree, after step 2 writes `<scratch>/dispositions-<pr>.json`; it reads the review file beside it, which you write yourself from the findings step 1 collected, in the shape **The review result file** describes — `pr`, `head`, `snapshot`, `counts` and the `survived`, `refuted` and `unverified` buckets; standalone you pass `--no-ledger`, so it writes no token and the exit status is the verdict — a `.fleet/ledger.md` the repository already holds is a previous run's, and without the flag it exits 2 for a ledger with no row for you; under the fleet `ledger.mjs dispatch` has already refused a finisher without it) → `gh pr edit <pr> --add-label ready-to-merge`. Do not merge. A non-zero dispositions exit halts you before the label, and you report the violations it printed. Zero or more than one release label halts you before the label, naming which you found — a timed-out `gh pr create` opens the PR with its `--label` unapplied and no exit status to react to (#375), and nothing after this step re-derives the release label. **Under the fleet a finisher does this** — a fresh small agent the controller dispatches (after your final verdict, never on your bare idle). Gate it on the diff-validating **`check`** job, **not** on `ci-state --quiet` exit 0, which a behind PR — the normal case once any sibling merges — never reaches: a `rebase-check` red or heavy jobs `skipped` off the behind-count hold it short of full green, so an exit-0 gate strands every behind PR unlabelled. Label when `check` is present and green **and no heavy job (the diff-validating suites — *not* the `rebase-check` currency gate) is in `failure`** — a `skipped` heavy job is behind-count staleness and fine — **solely** off that count, which is a condition to establish rather than infer, see the five-condition note in `$(~/.fleet/bin/fleet-run --root)/skills/run-team/SKILL.md`'s fix-applier prompt block. A `failure` is the diff's own fault and blocks the label, else a genuinely broken PR gets labelled off a green `check` and burns a merge-bot rebase+CI cycle before the failure resurfaces. Read per-job state for this: run `ci-state.mjs` **without** `--quiet`, or read its `jobs` — the `--quiet` payload drops the per-job `jobs`/`missing` and the `dropped` list, keeping `reasons` (which still distinguishes `is failure` from `is skipped`). You reach step 6 yourself only when green is already in hand at push time.
@@ -77,9 +77,9 @@ It cannot ask the user anything and cannot wait on CI — rebasing, watching che
77
77
  - **Say read-only.** No `fleet-review-*` definition's frontmatter restricts tool access — a `.agent.md` prompt saying "report only" is instruction, not enforcement — so an agent hand-dispatched here, named or general-purpose, has write tools it must not use: one committed and pushed during a report-only dispatch. No commits, no pushes, no worktree edits.
78
78
  - **Cut TWO trees, not one — a shared snapshot is not enough.** Mutation probing is a *write*, so a mutating specialist and read-only specialists cannot share a copy: the readers then observe the mutant exactly as if it were real code. Give them `snap-ro`, pristine and **never written**, and one private copy **per mutating specialist**. Keep every one out of the worktree. Observed with a single shared snapshot: three specialists read two *different* in-flight mutations, one reporting the PR's own bug as still present; elsewhere three watched their file go clean → `M` mid-analysis; twice a reviewer came one step from filing a false finding. Reverting a probe does not help — it leaves a window where every concurrent reader sees a lie, and serializing does not close it because readers are concurrent with the *mutator*.
79
79
  - **Do not trust the obvious contamination check.** `diff -rq` and `md5` against the mutating tree both returned *clean* — because the probe had been reverted between the two reads. A clean diff against a live tree is not evidence in either direction; only `snap-ro` and the object store settle it.
80
- - **`git archive` carries tracked files only.** No `node_modules`, no gitignored runner or config. Provision in the same step — symlink `node_modules`, copy in what the runner needs — or specialists silently have no runnable suite and reason from source instead of measuring. That failure is invisible: you get confident prose where you asked for a measurement.
80
+ - **`git archive` carries tracked files only.** No installed dependencies, no gitignored runner or config. Provision in the same step — run the Install step in the copy (read it off the worktree or main checkout, `derive-testcmd.sh <worktree or main checkout> install`; pointed at the copy it refuses, since a copy holds no Recipe cache) or symlink in what it produced in the worktree, copy in what the runner needs — or specialists silently have no runnable suite and reason from source instead of measuring. That failure is invisible: you get confident prose where you asked for a measurement.
81
81
  - **Initialize the copy before anyone measures in it — a bare `git archive` tree cannot validly run every suite.** It is not a git repository, so anything that asks git about the checkout breaks: `derive_workspace_id` shells out to `git rev-parse`, falls back to the directory basename and the hook tests fail on the *name* — naming the dir after the repo makes those pass **by coincidence, not correctness** — while every sweep that asks git what ships declines instead of running, which is the silent half (measured on this repo at `ffa9026`: 2401 tests, 2382 pass, 19 skipped in an extraction against 2401/2401/0 in a worktree, the totals identical). `git init && git add -A -f && git commit` in the extracted copy is what makes a count mean what it looks like; the review's snapshot agent does exactly that and verifies it by comparing the new commit's tree hash against the reviewed commit's (#1056), and a hand-cut tree needs the same before you read a count off it. Suites that need the real worktree for some *other* reason — the stack below — still do.
82
- - **Give specialists a stack-free test command, not the worktree's runner, and their own scratch dir** — `<scratch>/pr<N>/fleet-review-<key>/`, resolved to an absolute path that you write into that specialist's prompt; a specialist never derives its own. Specialists inherit no environment, and one falling back to the default config runs a `globalSetup` that brings the shared compose stack up and tears it down, recreating the DB mid-run for every sibling. **The snapshot does not cover this** — the compose project name comes from the environment, not the working directory, so three agents on three copies still collide. But `./agent-test` does not fix it either: it exports **one** `TEST_COMPOSE_PROJECT`/port triple per *worktree* (`ab-<issue>`), so N concurrent specialists sharing it tear down each other's postgres mid-run. Observed: a probe returned `No test files found / No such container` and would have read as a test failure. **That war story is from another repo, and the config-swap escape it calls for does not exist here** — `feigi/claude-config` has no compose file, no `globalSetup`, and no vitest, so there is no `vitest.ci.config.ts` to point anyone at. (It does now have CI — `.github/workflows/ci.yml` — but that runs `node --test`, not a stack.) Naming one sends every specialist into `Cannot find module`, which reads as a broken tree rather than a bad instruction. In *this* repo the collision is inert for a different reason: the `export TEST_COMPOSE_PROJECT=…` line `claim-ticket.sh` writes into the runner sets `TEST_COMPOSE_PROJECT`/`TEST_POSTGRES_PORT`/`TEST_OLLAMA_PORT` and **nothing reads them** — the suite is `node --test` throughout. So tell specialists this repo's test command directly — read off the Recipe cache the way step 3 reads it (`derive-testcmd.sh <worktree or main checkout> test` prints it; the snapshot carries no `.fleet/` cache, so pointing it at the snapshot refuses), here in glob form `node --test plugin/scripts/*.test.mjs` as this repo's example — and keep `./agent-test` for your own verification. **Hand that command with its reading rule: `tests 0` is a failed run, not a pass.** A glob is not self-checking. Where it matches nothing, bash passes the pattern through literally, node globs it itself and finds no files, and you get `ℹ tests 0` / exit 0 — a green having run nothing. Read the count under either reporter: `node --test` prints `ℹ tests 0` to a terminal and `# tests 0` to a pipe on older node, so a rule that greps only for `ℹ` finds nothing in a CI log and cannot tell a zero-test run from a full one. Do not count on the shell to catch it: zsh errors (`no matches found`, exit 1) and bash does not, so the guard has to be the reading rule. Node v26.5.0 has no flag that fails a zero-test run. **And zero is not the only no-work count.** **0 passes with no failures** is everything skipped, and a count **materially below the full suite** is a partial copy of the tree — which is the *documented* behavior, since the prompt tells specialists to run their probe work inside their own copy of the snapshot and nothing anchors an expected count for them. `review-core.mjs` classifies the first itself, in `unrunReason`; the second it structurally cannot, because that function is kept pure and so never learns how many tests the tree has. Handing the rule over with the command is the only guard the partial-copy case has. Carry the rule, not the command: in a repo that *does* have a stack, N concurrent specialists sharing one runner still collide, and the escape is whatever that repo's stack-free config is.
82
+ - **Give specialists a stack-free test command, not the worktree's runner, and their own scratch dir** — `<scratch>/pr<N>/fleet-review-<key>/`, resolved to an absolute path that you write into that specialist's prompt; a specialist never derives its own. Specialists inherit no environment, and one falling back to the default config runs a `globalSetup` that brings the shared compose stack up and tears it down, recreating the DB mid-run for every sibling. **The snapshot does not cover this** — the compose project name comes from the environment, not the working directory, so three agents on three copies still collide. But `./agent-test` does not fix it either: it exports **one** `TEST_COMPOSE_PROJECT`/port triple per *worktree* (`ab-<issue>`), so N concurrent specialists sharing it tear down each other's postgres mid-run. Observed: a probe returned `No test files found / No such container` and would have read as a test failure. So tell specialists the repo's test command directly — the Test entrypoint, read off the Recipe cache the way step 3 reads it (`derive-testcmd.sh <worktree or main checkout> test` prints it; the snapshot carries no `.fleet/` cache, so pointing it at the snapshot refuses) — when it brings up no shared stack, and the repo's own stack-free config when it does; never name a config the repo does not have, which sends every specialist into a missing-file error that reads as a broken tree rather than a bad instruction. Keep `./agent-test` for your own verification. **Hand that command with its reading rule: `tests 0` is a failed run, not a pass.** A command is not self-checking: a path glob or filter that matches nothing can exit 0 having run nothing. Measured with `node --test` over a glob matching no file: bash passes the pattern through literally, node globs it itself and finds no files, and you get `ℹ tests 0` / exit 0 — and `# tests 0` to a pipe on older node, so a rule that greps only for `ℹ` finds nothing in a CI log and cannot tell a zero-test run from a full one. Do not count on the shell to catch it: zsh errors (`no matches found`, exit 1) and bash does not, so the guard has to be the reading rule. **And zero is not the only no-work count.** **0 passes with no failures** is everything skipped, and a count **materially below the full suite** is a partial copy of the tree — which is the *documented* behavior, since the prompt tells specialists to run their probe work inside their own copy of the snapshot and nothing anchors an expected count for them. `review-core.mjs` classifies the first itself, in `unrunReason`; the second it structurally cannot, because that function is kept pure and so never learns how many tests the tree has. Handing the rule over with the command is the only guard the partial-copy case has. Carry the rule, not the command: wherever a repo has a stack, N concurrent specialists sharing one runner collide, and the escape is whatever that repo's stack-free config is.
83
83
  - **Completion is not delivery — but you can reach a specialist, so never wait idle.** A specialist you dispatched reports when it settles; if a report is missing, `hub send` it by name and ask for the report. Judge delivery by whether a report came out. **Do not settle a dimension you hold no report for**: ask the specialist, then the controller by name; only when neither answers, rule and name that dimension **unrun** — a killed specialist never reports, so waiting on one is unbounded. Never ship "dispatched, never delivered" as settled; a ruling citing a report you do not hold gets verified from source, never applied on trust.
84
84
  - **Apply only once the fan-out is settled — every report either in hand or accounted for.** Collection is not guaranteed; a report that never arrived is an open dimension, not a smaller set. And editing while they read is the same defect as probing — one specialist reviewed uncommitted code that was never in the PR diff.
85
85
 
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@feigi/fleet-ctl",
3
- "version": "3.23.23",
3
+ "version": "3.23.25",
4
4
  "description": "Agent fleet: run-team controller, merge bot, PR reviewer, and the ticket pipeline they share",
5
5
  "license": "Apache-2.0",
6
6
  "repository": {
@@ -0,0 +1,126 @@
1
+ import { test } from "node:test";
2
+ import assert from "node:assert/strict";
3
+ import { readFileSync, readdirSync } from "node:fs";
4
+ import { join } from "node:path";
5
+ import { between, bullet, paragraph, phrase } from "./prose-pin.mjs";
6
+
7
+ // The prose that describes the consumer repo as "whatever the Recipe says"
8
+ // rather than as a Node project: every shipped runbook and the requirement
9
+ // docs name the Recipe's Install step and Test entrypoint, and none of them
10
+ // names a Node package manager, a lockfile or a Node-only tool as the way a
11
+ // consumer is installed or tested. Each pin below holds ONE contiguous clause
12
+ // of that vocabulary inside the bounded slice that carries it.
13
+ const PLUGIN = join(import.meta.dirname, "..");
14
+ const ROOT = join(PLUGIN, "..");
15
+ const readPlugin = (rel) => readFileSync(join(PLUGIN, rel), "utf8");
16
+ const readRoot = (rel) => readFileSync(join(ROOT, rel), "utf8");
17
+ const RUN_TEAM = readPlugin("skills/run-team/SKILL.md");
18
+ const REAPING = readPlugin("skills/run-team/references/reaping.md");
19
+ const ISOLATION = readPlugin("skills/run-team/references/isolation.md");
20
+ const REVIEW_AND_FIX = readPlugin("commands/review-and-fix.md");
21
+ const REQUIREMENTS = readRoot("docs/requirements.md");
22
+ const PULL_AND_CLAIM = readRoot("docs/components/pull-and-claim.md");
23
+ const EXTERNAL = readRoot("docs/research/external-assumptions.md");
24
+
25
+ // A Node-consumer assumption: a lockfile or `node_modules` named as the
26
+ // consumer's install state, or a Node package-manager / runner invocation
27
+ // named as how a consumer is installed or checked.
28
+ const NODE_CONSUMER = /npm ci|node_modules|package-lock|pnpm-lock|yarn\.lock|npm install|npx (?:vitest|tsc)/;
29
+
30
+ const markdownUnder = (dir) =>
31
+ readdirSync(join(PLUGIN, dir), { recursive: true })
32
+ .filter((rel) => rel.endsWith(".md"))
33
+ .map((rel) => join(PLUGIN, dir, rel));
34
+
35
+ test("the sweep's own pattern matches each Node-consumer form it exists to catch", () => {
36
+ for (const sample of ["npm ci", "its node_modules still on disk", "package-lock.json", "pnpm-lock.yaml", "yarn.lock", "run npm install", "npx tsc --noEmit", "npx vitest run"]) {
37
+ assert.match(sample, NODE_CONSUMER, `the pattern no longer matches "${sample}"`);
38
+ }
39
+ });
40
+
41
+ test("no runbook, command, agent definition or requirement doc names a Node-consumer install state or tool", () => {
42
+ const files = [
43
+ ...markdownUnder("skills"),
44
+ ...markdownUnder("commands"),
45
+ ...markdownUnder("agents"),
46
+ join(ROOT, "docs", "requirements.md"),
47
+ join(ROOT, "docs", "components", "pull-and-claim.md"),
48
+ join(ROOT, "README.md"),
49
+ ];
50
+ assert.ok(files.includes(join(PLUGIN, "skills", "run-team", "SKILL.md")), "the sweep no longer reaches run-team/SKILL.md");
51
+ const offenders = files.filter((f) => NODE_CONSUMER.test(readFileSync(f, "utf8")));
52
+ assert.deepEqual(offenders, [], `Node-consumer wording is back in: ${offenders.join(", ")}`);
53
+ });
54
+
55
+ test("the claim step runs the Recipe's Install step in the worktree, read from the cache", () => {
56
+ const step = bullet(PULL_AND_CLAIM, "4. **Claim.**", "5. **Dispatch**", "pull-and-claim.md step 4");
57
+ assert.match(step, phrase("runs the [Recipe](recipe.md)'s Install step there — read from the Recipe cache, refusing if it changed the tree"));
58
+ });
59
+
60
+ test("requirements 2.3 describes the Recipe derivation: any technology, what it reads, what proves it, what it caches", () => {
61
+ const section = between(REQUIREMENTS, "### 2.3 Installable and testable — HARD", "### 2.4", "requirements.md section 2.3");
62
+ assert.match(section, phrase("Any technology. Fleet derives your repo's Recipe (Install step + Test entrypoint) by agent reasoning"));
63
+ const reads = bullet(section, "- **What the derivation reads:**", "- **What the proof requires:**", "requirements.md 2.3 derivation bullet");
64
+ assert.match(reads, phrase("nothing checks for a particular manifest or lockfile."));
65
+ const proof = bullet(section, "- **What the proof requires:**", "- **The cache:**", "requirements.md 2.3 proof bullet");
66
+ assert.match(proof, phrase("`RECIPE NOT PROVEN` writes no cache and halts the run."));
67
+ const cache = bullet(section, "- **The cache:**", "Hard rules for any technology:", "requirements.md 2.3 cache bullet");
68
+ assert.match(cache, phrase("Every claim reads both commands from it and every review its Test entrypoint; a missing or unproven cache is a refusal naming the derivation step, never a guess."));
69
+ assert.match(section, phrase("~/.fleet/bin/fleet-run derive-testcmd.sh . install && ~/.fleet/bin/fleet-run derive-testcmd.sh . test"));
70
+ });
71
+
72
+ test("requirements 1.1 adds whatever the Recipe runs to the tool floor", () => {
73
+ assert.match(
74
+ between(REQUIREMENTS, "| `shasum` | any |", "### 1.2", "requirements.md section 1.1 table tail"),
75
+ phrase("Plus whatever your Recipe runs (§2.3): `mvn`, `go`, `cargo`, …"),
76
+ );
77
+ });
78
+
79
+ test("the research note marks the Node-only consumer rows superseded by the Recipe ruling", () => {
80
+ const row = bullet(EXTERNAL, "- ~~Contains `package.json`", "\n- `node` on PATH", "external-assumptions.md consumer-repo row");
81
+ assert.match(row, phrase("Superseded by ADR 0015: the claim reads the Install step and Test entrypoint from the Recipe cache, and nothing checks for a manifest, a test-file name or a lockfile."));
82
+ const implied = bullet(EXTERNAL, "- Consumer must be a Node project", "\n- `python3` and `shasum` on PATH", "external-assumptions.md implied-only row");
83
+ assert.match(implied, phrase("Ruled 2026-09-28: ADR 0015 — any technology, Recipe by agent reasoning"));
84
+ });
85
+
86
+ test("review-and-fix step 3 reads the standalone testCmd off the Recipe cache and derives when there is none", () => {
87
+ const step = bullet(REVIEW_AND_FIX, "3. **Run `testCmd` before you commit**", "\n4. **", "review-and-fix.md step 3");
88
+ assert.match(step, phrase("the Test entrypoint in the repo's Recipe cache (`derive-testcmd.sh <repo> test` prints it)"));
89
+ assert.match(step, phrase("run the Recipe derivation step first"));
90
+ assert.doesNotMatch(step, /node --test/, "step 3 names this repository's own test command again");
91
+ });
92
+
93
+ test("review-and-fix's git archive bullet runs the Install step in the copy, read from a tree that holds the cache", () => {
94
+ const archive = bullet(REVIEW_AND_FIX, "- **`git archive` carries tracked files only.**", "\n- **Initialize the copy", "review-and-fix.md git archive bullet");
95
+ assert.match(archive, phrase("run the Install step in the copy (read it off the worktree or main checkout, `derive-testcmd.sh <worktree or main checkout> install`; pointed at the copy it refuses, since a copy holds no Recipe cache)"));
96
+ });
97
+
98
+ test("run-team: a tree with no runner gets the Recipe's Test entrypoint only when it brings up no shared stack", () => {
99
+ const p = paragraph(RUN_TEAM, "**A reused worktree may lack the runner** — no longer here", "run-team/SKILL.md reused-worktree runner paragraph");
100
+ assert.match(p, phrase("the Recipe's Test entrypoint (`~/.fleet/bin/fleet-run derive-testcmd.sh <main checkout> test` prints it) only when it brings up no shared stack to collide on, else that repo's own stack-free command."));
101
+ });
102
+
103
+ test("run-team: testCmd comes from the Recipe cache and names no repository's own command", () => {
104
+ const p = paragraph(RUN_TEAM, "**Where `testCmd` comes from:**", "run-team/SKILL.md testCmd source paragraph");
105
+ assert.match(p, phrase("the repository's Test entrypoint, out of the Recipe cache phase 0's Recipe derivation step proved — `~/.fleet/bin/fleet-run derive-testcmd.sh . test` prints it — and the one you hand specialists"));
106
+ assert.doesNotMatch(p, /node --test/, "the testCmd source paragraph names this repository's own test command again");
107
+ });
108
+
109
+ test("run-team: the quick-install red flag points at the Recipe's Install step, not at an install command", () => {
110
+ const flag = bullet(RUN_TEAM, "- \"A quick install to set up the worktree\"", "- \"I'm on my own copy, so I'm isolated\"", "run-team/SKILL.md quick-install red flag");
111
+ assert.match(flag, phrase("the claim already ran the Recipe's Install step; any other install can rewrite the lockfile for the whole repo."));
112
+ });
113
+
114
+ test("a merged branch's stale worktree is said to keep its installed dependencies, in SKILL.md and in reaping.md", () => {
115
+ const skill = paragraph(RUN_TEAM, "The merge bot deletes the remote branch after each merge", "run-team/SKILL.md reap paragraph");
116
+ assert.match(skill, phrase("with its worktree — and its installed dependencies — still on disk."));
117
+ const evidence = paragraph(REAPING, "Merge deletes remote branch, leaves local branch", "reaping.md reap-after-each-merge paragraph");
118
+ assert.match(evidence, phrase("with worktree — and its installed dependencies — still on disk."));
119
+ });
120
+
121
+ test("the isolation reference says symlinking installed dependencies in does not isolate the stack, and reproduces a diagnostic by re-running the check", () => {
122
+ const stack = paragraph(ISOLATION, "Private copy and test command solve different problems", "isolation.md stack-isolation paragraph");
123
+ assert.match(stack, phrase("Symlinking installed dependencies in does not help."));
124
+ const diag = paragraph(ISOLATION, "Probe copies carry same filenames as real tree", "isolation.md diagnostics paragraph");
125
+ assert.match(diag, phrase("(re-run the check that raised it — the typechecker, the linter — from there)"));
126
+ });
@@ -71,11 +71,12 @@ for (const [name, text, from, to] of slices) {
71
71
  // target's presence needs its own pin or the pointer orphans in silence. The end
72
72
  // anchor is the generic next-bullet marker, not the following bullet's wording,
73
73
  // which would make an unrelated rewrite of that bullet a boundary failure here.
74
- // Bounding is not optional: the glob occurs twice in that file, so an unbounded
75
- // pin would stay satisfied by the other occurrence with this bullet deleted.
74
+ // Bounded to the bullet so that a copy of the phrase anywhere else in the file
75
+ // cannot keep the pin satisfied with this bullet deleted.
76
76
  test("review-and-fix.md still hands specialists the command both guards point at", () => {
77
77
  const bullet = between(REVIEW_AND_FIX, "Give specialists a stack-free test command", "\n- **", "review-and-fix.md");
78
- assert.match(bullet, phrase("node --test plugin/scripts/*.test.mjs"));
78
+ assert.match(bullet, phrase("the Test entrypoint, read off the Recipe cache"));
79
+ assert.match(bullet, phrase("`derive-testcmd.sh <worktree or main checkout> test` prints it"));
79
80
  });
80
81
 
81
82
  // #1150: any member that dispatches a child writes the child's absolute scratch
@@ -824,9 +824,10 @@ prior-run worktree). #55 tracked the runner, so any worktree checked out from
824
824
  `origin/main` now carries it. A tree that is NOT a checkout — a `git archive`
825
825
  snapshot, a `cp -R` subset — still has whatever was copied into it, and a repo
826
826
  that tracks no runner never had one: there, tell the member the runner is absent
827
- and to run docker-free suites directly (`npx vitest run --config
828
- vitest.ci.config.ts <file>` — the CI unit config has no `globalSetup`, so there
829
- is no stack to collide on).
827
+ and to run docker-free suites directly — the Recipe's Test entrypoint
828
+ (`~/.fleet/bin/fleet-run derive-testcmd.sh <main checkout> test` prints it) only
829
+ when it brings up no shared stack to collide on, else that repo's own stack-free
830
+ command.
830
831
 
831
832
  **A reused worktree may also be on the wrong COMMIT.** `git worktree add <path>
832
833
  <branch>` checks out the existing LOCAL branch and never consults the remote, so
@@ -2212,10 +2213,10 @@ reads gates and never injects a fault.
2212
2213
 
2213
2214
  **Where `testCmd` comes from:** the repository's Test entrypoint, out of the
2214
2215
  Recipe cache phase 0's Recipe derivation step proved —
2215
- `~/.fleet/bin/fleet-run derive-testcmd.sh . test` prints it, in this repo
2216
- `node --test plugin/scripts/*.test.mjs` — and the one you hand specialists per
2217
- **Give specialists a stack-free test command** above. Pass the same string to
2218
- the review, to the fix-applier, and to the finisher — whose duty-2 mutation
2216
+ `~/.fleet/bin/fleet-run derive-testcmd.sh . test` prints it — and the one you
2217
+ hand specialists per **Give specialists a stack-free test command** above. Pass
2218
+ the same string to the review, to the fix-applier, and to the finisher — whose
2219
+ duty-2 mutation
2219
2220
  gate runs it too — so every gate runs one command. Omit it from the review args
2220
2221
  and the review (`review-core.mjs`, in a runner) reads the same cache itself,
2221
2222
  refusing outright when there is none; the fix-applier has no such fallback, so
@@ -3095,7 +3096,7 @@ ancestry proof, post-rebase red triage. Do not restate them here.
3095
3096
 
3096
3097
  The merge bot deletes the remote branch after each merge (`delete-merged-branch.sh`,
3097
3098
  in `run-merge-bot.md` step 4), leaving the local branch `[gone]` with its
3098
- worktree — and its `node_modules` — still on disk. Reap after **each** merge pass,
3099
+ worktree — and its installed dependencies — still on disk. Reap after **each** merge pass,
3099
3100
  not once at the end: a stale worktree still answers `git worktree list`, so the
3100
3101
  in-flight probe (`inflight.sh`, run by the Shortlist and by every Pull) reads an
3101
3102
  already-merged ticket as taken and the queue quietly shrinks. The trigger is the
@@ -3523,8 +3524,9 @@ failures arrive as *wrong findings*, not errors:
3523
3524
  under your own path. See references/isolation.md.
3524
3525
  - **IDE/harness diagnostics attribute by bare filename, with no path.** **Never
3525
3526
  relay a diagnostic without reproducing it in that member's specific worktree**
3526
- (`npx tsc --noEmit` from there): probe copies carry the real tree's filenames,
3527
- so a sibling's throwaway mutation reads exactly like a live worktree's error.
3527
+ (re-run the check that raised it from there): probe copies carry the real
3528
+ tree's filenames, so a sibling's throwaway mutation reads exactly like a live
3529
+ worktree's error.
3528
3530
  See references/isolation.md.
3529
3531
 
3530
3532
  Suspect a neighbour before a member's own diff — for unexplained failures, and
@@ -3845,8 +3847,8 @@ Plus a queue-depth line: shortlist, supply, whether triage was suggested.
3845
3847
  Ask what scope it searched.
3846
3848
  - "Five implementers = five times throughput" → reviews are 3-5x longer. It means
3847
3849
  a backlog.
3848
- - "`npm install` to set up the worktree" → the wrong install mutates the lockfile
3849
- for the whole repo.
3850
+ - "A quick install to set up the worktree" → the claim already ran the Recipe's
3851
+ Install step; any other install can rewrite the lockfile for the whole repo.
3850
3852
  - "I'm on my own copy, so I'm isolated" → not from the docker stack.
3851
3853
  - "`commit-commands:clean_gone` printed nothing, the tree is clean" → its
3852
3854
  `[gone]` detection works fine; it just runs `-D`/`--force` with no merged
@@ -10,7 +10,7 @@ The runner is re-materialized on every invocation, so no copy of it survives a r
10
10
 
11
11
  ## Filesystem isolation is not stack isolation
12
12
 
13
- Private copy and test command solve different problems; conflating them is how second gets skipped: compose project name comes from environment, not working directory, so three agents on three snapshots still collide on one postgres. Symlinking `node_modules` does not help. "I'm on my own copy" is exactly the intuition that skips command — say both, every time. Which command depends on audience: member in worktree uses `./agent-test`; specialist on snapshot does not — the runner that bootstrap (#55) materializes derives its isolation triple from the directory name, and a snapshot is not a claimed worktree, so every snapshot gets the same fixed ports (measured in a snapshot: `ports derive from the issue number: postgres=16000 ollama=22000`) — and takes one `review-and-fix.md` hands out. The bootstrap no longer FAILS there: since #1056 the snapshot is a git repository, so `claim-ticket.sh --write-runner` succeeds in it, and the reason to hand a specialist a different command is the stack rather than a missing repo.
13
+ Private copy and test command solve different problems; conflating them is how second gets skipped: compose project name comes from environment, not working directory, so three agents on three snapshots still collide on one postgres. Symlinking installed dependencies in does not help. "I'm on my own copy" is exactly the intuition that skips command — say both, every time. Which command depends on audience: member in worktree uses `./agent-test`; specialist on snapshot does not — the runner that bootstrap (#55) materializes derives its isolation triple from the directory name, and a snapshot is not a claimed worktree, so every snapshot gets the same fixed ports (measured in a snapshot: `ports derive from the issue number: postgres=16000 ollama=22000`) — and takes one `review-and-fix.md` hands out. The bootstrap no longer FAILS there: since #1056 the snapshot is a git repository, so `claim-ticket.sh --write-runner` succeeds in it, and the reason to hand a specialist a different command is the stack rather than a missing repo.
14
14
 
15
15
  ## Scratchpad paths need two levels — the member's own partition, then one directory per child it dispatches
16
16
 
@@ -20,4 +20,4 @@ The axis is two-sided: the member partitions, and the parent assigns each child'
20
20
 
21
21
  ## IDE/harness diagnostics attribute by bare filename, with no path
22
22
 
23
- Probe copies carry same filenames as real tree, so specialist's throwaway mutation surfaces as errors that read exactly like live worktree's — and line numbers can plausibly line up with real in-flight edits. Never relay diagnostic without reproducing it in that member's specific worktree (`npx tsc --noEmit` from there). Ran twice in one session: clean first time (sibling's probe), genuinely broken second. Telling implementer to chase phantom in file it is mid-rewrite on is expensive failure.
23
+ Probe copies carry same filenames as real tree, so specialist's throwaway mutation surfaces as errors that read exactly like live worktree's — and line numbers can plausibly line up with real in-flight edits. Never relay diagnostic without reproducing it in that member's specific worktree (re-run the check that raised it — the typechecker, the linter — from there). Ran twice in one session: clean first time (sibling's probe), genuinely broken second. Telling implementer to chase phantom in file it is mid-rewrite on is expensive failure.
@@ -4,7 +4,7 @@ Why reap runs after each merge pass, why `commit-commands:clean_gone` disqualifi
4
4
 
5
5
  ## Reap after each merge pass, not once at the end
6
6
 
7
- Merge deletes remote branch, leaves local branch `[gone]` with worktree — and its `node_modules` — still on disk. Stale worktree still answers `git worktree list`, so the in-flight probe (`inflight.sh`, run by the Shortlist and again by every Pull) reads already-merged ticket as taken and queue quietly shrinks as run goes on.
7
+ Merge deletes remote branch, leaves local branch `[gone]` with worktree — and its installed dependencies — still on disk. Stale worktree still answers `git worktree list`, so the in-flight probe (`inflight.sh`, run by the Shortlist and again by every Pull) reads already-merged ticket as taken and queue quietly shrinks as run goes on.
8
8
 
9
9
  ## Why `commit-commands:clean_gone` is disqualified
10
10