@feigi/fleet-ctl 3.23.22 → 3.23.24
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/commands/review-and-fix.md +3 -3
- package/package.json +1 -1
- package/scripts/citation-sweep-prose.test.mjs +6 -2
- package/scripts/consumer-technology-prose.test.mjs +126 -0
- package/scripts/fleet-heartbeat.mjs +4 -4
- package/scripts/snapshot-runner-audience-prose.test.mjs +4 -3
- package/skills/run-team/SKILL.md +14 -12
- package/skills/run-team/references/isolation.md +2 -2
- package/skills/run-team/references/reaping.md +1 -1
|
@@ -10,7 +10,7 @@ argument-hint: [pr-number]
|
|
|
10
10
|
See **Specialists** — those constraints are not optional.
|
|
11
11
|
2. Plan the actions. Split explicitly into **apply now** / **defer**. A finding is deferred, never dropped. **Every `suggestion` — the review's `severity`, the plugin's *Suggestions* bucket — arrives unchecked: the review budgets that band 0 refuters, so nothing has looked at it yet.** Checking it is now your job, not nobody's. So split it by SCOPE first. Out of the PR's ticket scope → defer and file. In scope → dispatch one refuter against it, biased to refuse, and hand it this instruction verbatim with one substitution: resolve `<scratch>/pr<N>/<finding>/` in it to an absolute path under your own run root — `mktemp -d` once per round for `<scratch>/pr<N>/fix-XXXXXXXX/`, never bare `pr<N>`, since finding ids restart at 1 each round too — and write that absolute path into the refuter's prompt in its place; a refuter never derives its own path: *Try to REFUTE this finding. Default to refuted=true if uncertain. Verify by RUNNING something — compile it, run the test, apply the mutation. Do not reason your way to agreement. Observe that run synchronously — run the command, wait for it, read its exit code. Never poll a log file for a completion marker: prefer ONE blocking run to a poll loop, and treat its return as permission to look, never as the answer. Reading a log the run has already finished writing is fine; waiting on one is not. If you match a test reporter's own output, accepting both `ℹ` and `#` is necessary but NOT sufficient — strip SGR escapes first as well. node's prefix moves with the node version and with whether stdout is a TTY, and color wraps the whole line so it begins with ESC and no prefix anchor matches at all, which returns empty at exit 0 — indistinguishable from a hung run and from a run of zero tests. For an uncolored baseline use `env -u FORCE_COLOR`; `FORCE_COLOR=` empty still enables color, so it is not a control. State your search scope AND what your pattern would have missed. A grep over one ref does not support a claim about history; a pattern built from the token a diff removed does not support a claim that the category is empty. A failure injection with no positive control has produced NO result, never a negative one. Before you read an injected fault — an env var, an argv word, a mutant — as having had no effect, prove the injection reached the child: one run whose output differs with it present versus absent, or the child echoing the injected value back. Uncontrolled, the cell is unrun — say so in your verdict instead of reporting a no-effect result. Build such an invocation as an array expanded braced and quoted — `cfg=(SETB=1 BADJ=1); env "${cfg[@]}" sh ./probe.sh` — or inline the assignments literally — `env SETB=1 BADJ=1 sh ./probe.sh`; NEVER from an unquoted scalar — `cfg="SETB=1 BADJ=1"; env $cfg sh ./probe.sh` — which under zsh passes ONE argument, sets a variable literally named `SETB` to `1 BADJ=1`, never sets `BADJ` at all, and still exits 0. `env $cfg[@]` is not the portable spelling either: measured, bash word-splits it into `SETB=1` and `BADJ=1[@]`, so the injection variable is set to a corrupted value, while zsh behaves exactly as with the bare `$cfg` — `BADJ` never set, exit 0. Everything you write — mutants, fixtures, scratch repos — goes under `<scratch>/pr<N>/<finding>/` and nowhere else; the checkout and any worktree are never write targets, though `git show`/`git archive` at a pinned ref read fine anywhere. Chain the directory change into the command, `cd "$D" && git …`, never `cd "$D"; git …`, so a failed `cd` cannot leave a `git` command running in the checkout — and bracket a fixture's own git with `git rev-parse --show-toplevel`: before `git init` it must NOT resolve to the repository, and a fresh scratch dir's `fatal: not a git repository` (exit 128) is the pass, not a failure; before any `git commit` it must resolve to your scratch path — compare resolved forms (`realpath`), since `--show-toplevel` can report `/private/tmp/…` for a `/tmp` scratch dir on macOS.* That last clause is the whole mechanism — a refuter left to reason its way to agreement rubber-stamps, and "apply only what survives" silently degrades into "apply everything in scope". Apply only what survives; a refuted one defers with the refutation in its body. "In scope" means the ticket's scope, not merely the same files — and applying one still never licenses editing a file the ticket has no business in. **A finding in `unverified` whose refuters ran and crashed always defers, whatever its severity** — and which of the two it is, you read off the finding's **`refutersDispatched`**, never off severity and never off an empty vote list: above zero with nothing surviving means every refuter dispatched against it died, and at `critical` that is every one of them. Only a `survived` finding has been checked; severity says how much a finding would matter if true, never whether anything looked. **Deferring a crash is your move, and re-running it is not** — the review already re-dispatched every crashed refuter pair once, in-run, so a crash that reaches you crashed twice, and the result's `resume` field names the one move left, which is not yours: nothing replays a review that ran in a runner's kernel — it is reported, not acted on. So say *in the deferral* that its refuters crashed rather than that it went unchecked, and re-run nothing for it. **A `suggestion` is also in `unverified`, for a different reason — `refutersDispatched` of zero, the 0-refuter budget the workflow gives that band by policy, so nothing looked *there* — and the scope split above, not this rule, covers it.**
|
|
12
12
|
**Write every ruling to `<scratch>/dispositions-<pr>.json`, beside the review result file — the disposition record `dispositions-check.mjs` judges against it before any finisher is dispatched.** It is `{"head": "<the review file's head>", "entries": [ … ]}`: one entry per `survived` finding and per `unverified` finding, plus one per `refuted` finding you reverse, its evidence in `reason`. Each entry carries `bucket` (`survived|unverified|refuted`), `index` (its position in that bucket of the review file — stable, since a round's file never changes), `scope` (`in|out`), `claimKind` (`behavior|shape`) and `disposition` (`apply|defer`), and where they apply `reason`, `issue` (the number you filed it to or commented on), `verdictPath` (an in-scope suggestion's refuter verdict file) and `remedyFiles` (the files its remedy names). The check reads scope off the diff before it reads your entry: a finding with no `line`, or on a line `git diff` touches from the PR's merge-base with `origin/main` to the review's `head`, is in scope whatever `scope` says, and only otherwise does the declared one stand. An in-scope `survived` finding deferred passes only with `reason` `false-rationale`, `mutual-exclusion`, `remedy-worse` or `remedy-outside-diff`. `remedy-outside-diff` — the remedy would edit a file the PR's diff does not touch — also needs `remedyFiles` to name at least one file absent from `git diff` from that merge-base to the review's `head`: none named, or every one in that diff, is a mismatch. A `critical` or `important` finding deferred for it goes to a human instead of passing: the script exits 1 with verdict `escalate`, `ledger.mjs dispatch` refuses the PR's finisher, and nothing retries it; a `suggestion` deferred for it passes. Any other reason, none, or a `survived` or `unverified` finding with no entry is a mismatch: the script exits 1 naming each violating entry's bucket, index and rule, and `ledger.mjs dispatch` refuses the PR's finisher.
|
|
13
|
-
3. **Run `testCmd` before you commit**, from the worktree, copied verbatim — not a runner you picked. `testCmd` is the command your dispatcher handed you; **standalone, with no dispatcher to hand you one, it is the repo's own test command**, the Test entrypoint in the repo's Recipe cache (`derive-testcmd.sh <repo> test` prints it)
|
|
13
|
+
3. **Run `testCmd` before you commit**, from the worktree, copied verbatim — not a runner you picked. `testCmd` is the command your dispatcher handed you; **standalone, with no dispatcher to hand you one, it is the repo's own test command**, the Test entrypoint in the repo's Recipe cache (`derive-testcmd.sh <repo> test` prints it). Read it off the repo, never invent one: "not a runner you picked" is what this step forbids, and on the standalone path nothing else defines the referent. With no usable Recipe cache (`derive-testcmd.sh` refuses), run the Recipe derivation step first — dispatch a `task` member, agent `fleet-recipe-deriver`, its prompt naming the main checkout's absolute path as `<repo>` and a scratch directory of its own; it proves both commands in a throwaway worktree and writes the cache only on proof, and its `RECIPE NOT PROVEN` is a stop, reported with its reason. Never fall back to a runner of your own choosing. `tests 0` is a FAILED run, not a pass: a glob matching nothing still exits 0. **Zero is not the only no-work count.** A run reporting **0 passes with no failures** skipped everything it collected, and a count **materially below the whole suite** means you ran a partial tree — this suite reported `tests 937` when this line was written, so compare what you get against a full run rather than against zero. The number drifts; that is the safe direction, because a stale one produces a loud false red while its absence produces a silent false green. Red, zero-test, or either no-work count → fix it or move the finding to defer; never commit over it. Then commit and push — `git push origin HEAD:<branch>`, the explicit refspec: this step also runs from a `--detach` worktree (`run-team/SKILL.md`'s worktree-recreation form), where a bare `git push` exits 128 (`fatal: You are not currently on a branch`) and any push whose destination is `HEAD` rather than a full or branch refname (`-u` or not, forced or not) exits 1 (`error: The destination you provided is not a full refname`) — and report the test result alongside the SHA. **Then reconcile the PR body with what you actually did.** It merges into the permanent record and nothing downstream re-derives it: your commit messages are yours to fix and usually get fixed, the body is neither. Removing a change, or reversing one, leaves the body asserting work the diff no longer contains. Measured twice in one run — one body kept asserting a calibration its own fix had just deleted as false, another kept claiming a hunk the fix had removed as "this diff owns it" — both caught by a finisher's reading rather than by any rule, and both would otherwise have merged intact. `gh pr edit <pr> --body-file <path>`, and re-read the body against the final diff rather than against your memory of it. **Never `--no-verify`.** The fleet runs in repos it knows nothing about: it must not assume a pre-commit hook exists, and must never bypass one that does — a hook that fails is a finding to report, not an obstacle. Nothing re-reviews what you commit here, so this run is the only gate between your edit and a merge.
|
|
14
14
|
4. **Under the fleet, do not hold this wait.** The controller owns a persistent CI Monitor and dispatches a finisher once your report has arrived and CI is green — so push, then report to the controller a **final verdict** (every specialist you dispatched + its outcome + your apply/defer adjudication + the pushed SHA) and stop. A specialist that has neither reported nor been relayed to you is not a verdict line — ask the controller for it by name before ruling; see **Specialists**. The controller keys "review complete" on that verdict, **never on you merely going idle** — a slow specialist can still be running when your turn ends, and finishing off idle-alone ships a premature label (observed: a PR labelled while its contract-test specialist was still running and about to surface a vacuous pin). A turn-based reviewer re-reading `gh pr checks` each idle cycle rebuilds its whole context for nothing the Monitor lacks. **Standalone (no controller): watch** the run until the diff-validating jobs are green (bind via `ci-state.mjs` below); fix a **genuine** failure and repeat from 3. **Staleness is not a failure:** a `rebase-check` red or heavy jobs `skipped` **solely from a non-zero behind-count** never clears by waiting and must not be rebased away here — see **`skipped` is not `passed`** and **Do NOT rebase to label** below for what it looks like and why. Stop watching and go to 6. A turn-based reviewer that "waits for green" on a behind PR stalls forever and ends its turn without labelling. **`ci-state.mjs` returns `verdict: "no-ci"`** → there is no run to watch, so don't: go straight to 6 and let that step's `--declare-no-ci` branch decide.
|
|
15
15
|
5. File each deferred finding: `gh issue create --label <ready-for-agent|needs-triage>` with the finding, `file:line`, why deferred, `Deferred from PR #<pr> review`, and **the finding's `dimension` in the issue body** — the review returns it on every finding, and a backlog that does not carry it cannot be attributed to the specialist that generated it. **The label follows the finding's state, not this step** — but it is never omitted; an unlabelled issue is invisible to both `candidates.mjs --require-label` and the cockpit's pool query, reachable only by a human. **The bar here is whether the DEFECT is confirmed — never whether the remedy is settled.** Confirmed defect with the remedy still open → `ready-for-agent`, and say in the body that the remedy is open. Torn → `needs-triage`, narrowed: torn means **unsure anything is broken**, not unsure how to fix it. **A finding a refuter CHECKED and did not survive is neither** — it is not a confirmed defect, and you are not unsure whether anything is broken: you measured that nothing is. Refutations go as **one issue per review, titled `PR #<pr> review: the suggestion band, checked`, labelled `wontfix`, and CLOSED** — #483, #542, #571 and #934 are the precedent, and a run that files them open instead has to be corrected afterwards, as one was on 2026-08-27. Left open they read as queued agent work and rot into filing artifacts; a measured triage pass found a quarter of its `needs-triage` backlog was refuted records, bulk-closable by title shape. Closed is still a durable home — `check` reads the tracker at `--state all`, so a future run re-deriving the finding lands on the evidence rather than re-litigating it. Split the band only where an entry is **not** a refutation: a residual the refuter surfaced, or a finding refused as out of scope and never checked, is real deferred work — and real deferred work earns its own open issue only past a second bar. **Worth a claim: a finding whose own claim is that correct code could be shaped better — style, naming, layout, redundancy, a micro-simplification — never gets its own open issue.** Its only open question is "is this worth doing?", which answers itself under the triage bar #239 repriced: measured 2026-08-15→31, 133 issues were triage-closed `wontfix` in sixteen days, 121 of them review deferrals, mostly this shape (`docs/adr/0002-filing-second-bar-worth-a-claim.md`). Append it to the same closed record issue under a `Below the claim bar` heading — `file:line`, dimension, the claim, and a re-open trigger per entry. Closed is a durable home for these too: `check` reads `--state all`, so a run re-deriving the finding lands on the record instead of re-filing — and a later review returning the same finding as `survived`, or a maintainer hitting it, is the promotion signal: file it open then, citing the record. **An in-scope rediscovery step 2 applied is that same promotion signal, and it files nothing** — so the record carries no mark unless you write one: run `check` over each finding step 2 applied as well, and where its tracker rows put one on a `Below the claim bar` entry, comment on that record issue, opening `Promoted — applied, not filed`, and name the entry, that its refuter let it through, and the commit carrying the fix. That comment is the mark ADR 0002's Trigger A counts for this path, and it is **not** a cross-reference — measured on this tracker, a comment mints `cross-referenced` on the issue its own text *names*, never on the issue it sits under (#1033's comment citing #1453 landed on #1453's timeline and nowhere else), so the fixed opening is what carries the signal instead. **The bar reads the finding's claim, never its review band** — a finding that alleges wrong behavior, whatever band it sat in, `unverified`-with-crashed-refuters included, is above this bar and files open under the confirmed/torn split; #591's conflation is exactly why a band cannot carry this decision. Do **not** reach for phase 0's **decided?** test here. That is a *consumption* gate, priced for a seat where implementer divergence costs a claim, a worktree and a dispatch; at filing time divergence costs one PR review the maintainer already runs. Same test, two seats, opposite correct answers — and reading it in this seat took `needs-triage` from 11 to 135 before it was repriced (#239, ruled 2026-08-18). **Dedupe through `~/.fleet/bin/fleet-run ledger.mjs check "<subject>"`, never a bare `gh issue list --search`** — the duplicate class that guard exists to catch was filed from *here*: the ledger row for the finding rediscovered five times in one measured run carries a reviewer's own source tag, `(review-pr-108)`. `check` reads this run's filed list *and* the tracker at `--state all`, so a closed duplicate still counts. Exit **1** is already filed in this run; exit **3** is a clean ledger with matching tracker rows that score against the subject — read the ones it prints and comment on the existing follow-up rather than filing a second issue. **Exit 0 is not automatically "safe to file"**: it covers `clean`, `soft-hit` and `unverified` alike, and an unreachable `gh` degrades to a ledger-only answer that says `TRACKER NOT CHECKED`, so read the stdout `verdict` to tell a tracker that was searched and clean from one nobody read — under `unverified`, search it yourself before filing — or, for the applied-promotion comment above, before treating a tracker nobody could read as a reason to skip it. **`soft-hit` is the answer that costs most to misread**: it means related rows were found and none of them was established as a duplicate — a near-miss above the floor, or tracker rows the scorer rates 0.00 — so read the rows it printed and decide, and never treat that exit 0 as permission. Exit **2** is a usage error, meaning nothing was checked at all: fix the invocation and re-run, never read it as a duplicate. Full exit-code and `verdict` semantics live in `$(~/.fleet/bin/fleet-run --root)/skills/run-team/SKILL.md`'s **Run ledger** section — read them there rather than restating them here. **The guard dedupes, it never drops:** a duplicate becomes a comment on the issue that already covers it, never a dropped finding. **Record every issue you create, at once:** immediately after each `gh issue create`, run `~/.fleet/bin/fleet-run ledger.mjs filed <N> "<subject>"` — `<N>` the number `gh issue create` printed, and **the same subject string you passed to `check`**, so the next `check` for that finding meets it in the filed list instead of scoring a rewording. Without it the guard's ledger half never holds a row, and its exit **1** can never fire. If `filed` fails, retry it once; if it fails again, leave the issue alone — it exists and is correct — and report it as `unrecorded: #N <subject>`, so the controller runs `filed` for it. Report `filed: #N <subject>` for every issue you created — under the fleet, in your report to the controller. **The one window the guard cannot see:** `check` → `gh issue create` → `filed` is three commands, and the ledger's lock covers each alone, so two filers can both read `clean` in the seconds before either records. That gap is accepted, never a reason to skip `filed`. Post the numbers as one PR comment.
|
|
16
16
|
6. Diff-check green (run-bound per `ci-state.mjs`; under staleness see **`skipped` is not `passed`** for what the label then rests on) **and** deferrals filed **and** the PR carrying exactly one release label (`patch`/`minor`/`major`, counting only those three; **required on every PR** — `release-label.yml` auto-adds `patch` when none is present and hard-fails on more than one, so this never needs to check whether the repo defines the gate at all, and `gh label list` settles nothing here any more) **and** `dispositions-check.mjs --member fix-pr-<pr> --scratch <scratch>` run over your own record exiting 0 (standalone that is yours to run, from the worktree, after step 2 writes `<scratch>/dispositions-<pr>.json`; it reads the review file beside it, which you write yourself from the findings step 1 collected, in the shape **The review result file** describes — `pr`, `head`, `snapshot`, `counts` and the `survived`, `refuted` and `unverified` buckets; standalone you pass `--no-ledger`, so it writes no token and the exit status is the verdict — a `.fleet/ledger.md` the repository already holds is a previous run's, and without the flag it exits 2 for a ledger with no row for you; under the fleet `ledger.mjs dispatch` has already refused a finisher without it) → `gh pr edit <pr> --add-label ready-to-merge`. Do not merge. A non-zero dispositions exit halts you before the label, and you report the violations it printed. Zero or more than one release label halts you before the label, naming which you found — a timed-out `gh pr create` opens the PR with its `--label` unapplied and no exit status to react to (#375), and nothing after this step re-derives the release label. **Under the fleet a finisher does this** — a fresh small agent the controller dispatches (after your final verdict, never on your bare idle). Gate it on the diff-validating **`check`** job, **not** on `ci-state --quiet` exit 0, which a behind PR — the normal case once any sibling merges — never reaches: a `rebase-check` red or heavy jobs `skipped` off the behind-count hold it short of full green, so an exit-0 gate strands every behind PR unlabelled. Label when `check` is present and green **and no heavy job (the diff-validating suites — *not* the `rebase-check` currency gate) is in `failure`** — a `skipped` heavy job is behind-count staleness and fine — **solely** off that count, which is a condition to establish rather than infer, see the five-condition note in `$(~/.fleet/bin/fleet-run --root)/skills/run-team/SKILL.md`'s fix-applier prompt block. A `failure` is the diff's own fault and blocks the label, else a genuinely broken PR gets labelled off a green `check` and burns a merge-bot rebase+CI cycle before the failure resurfaces. Read per-job state for this: run `ci-state.mjs` **without** `--quiet`, or read its `jobs` — the `--quiet` payload drops the per-job `jobs`/`missing` and the `dropped` list, keeping `reasons` (which still distinguishes `is failure` from `is skipped`). You reach step 6 yourself only when green is already in hand at push time.
|
|
@@ -77,9 +77,9 @@ It cannot ask the user anything and cannot wait on CI — rebasing, watching che
|
|
|
77
77
|
- **Say read-only.** No `fleet-review-*` definition's frontmatter restricts tool access — a `.agent.md` prompt saying "report only" is instruction, not enforcement — so an agent hand-dispatched here, named or general-purpose, has write tools it must not use: one committed and pushed during a report-only dispatch. No commits, no pushes, no worktree edits.
|
|
78
78
|
- **Cut TWO trees, not one — a shared snapshot is not enough.** Mutation probing is a *write*, so a mutating specialist and read-only specialists cannot share a copy: the readers then observe the mutant exactly as if it were real code. Give them `snap-ro`, pristine and **never written**, and one private copy **per mutating specialist**. Keep every one out of the worktree. Observed with a single shared snapshot: three specialists read two *different* in-flight mutations, one reporting the PR's own bug as still present; elsewhere three watched their file go clean → `M` mid-analysis; twice a reviewer came one step from filing a false finding. Reverting a probe does not help — it leaves a window where every concurrent reader sees a lie, and serializing does not close it because readers are concurrent with the *mutator*.
|
|
79
79
|
- **Do not trust the obvious contamination check.** `diff -rq` and `md5` against the mutating tree both returned *clean* — because the probe had been reverted between the two reads. A clean diff against a live tree is not evidence in either direction; only `snap-ro` and the object store settle it.
|
|
80
|
-
- **`git archive` carries tracked files only.** No
|
|
80
|
+
- **`git archive` carries tracked files only.** No installed dependencies, no gitignored runner or config. Provision in the same step — run the Install step in the copy (read it off the worktree or main checkout, `derive-testcmd.sh <worktree or main checkout> install`; pointed at the copy it refuses, since a copy holds no Recipe cache) or symlink in what it produced in the worktree, copy in what the runner needs — or specialists silently have no runnable suite and reason from source instead of measuring. That failure is invisible: you get confident prose where you asked for a measurement.
|
|
81
81
|
- **Initialize the copy before anyone measures in it — a bare `git archive` tree cannot validly run every suite.** It is not a git repository, so anything that asks git about the checkout breaks: `derive_workspace_id` shells out to `git rev-parse`, falls back to the directory basename and the hook tests fail on the *name* — naming the dir after the repo makes those pass **by coincidence, not correctness** — while every sweep that asks git what ships declines instead of running, which is the silent half (measured on this repo at `ffa9026`: 2401 tests, 2382 pass, 19 skipped in an extraction against 2401/2401/0 in a worktree, the totals identical). `git init && git add -A -f && git commit` in the extracted copy is what makes a count mean what it looks like; the review's snapshot agent does exactly that and verifies it by comparing the new commit's tree hash against the reviewed commit's (#1056), and a hand-cut tree needs the same before you read a count off it. Suites that need the real worktree for some *other* reason — the stack below — still do.
|
|
82
|
-
- **Give specialists a stack-free test command, not the worktree's runner, and their own scratch dir** — `<scratch>/pr<N>/fleet-review-<key>/`, resolved to an absolute path that you write into that specialist's prompt; a specialist never derives its own. Specialists inherit no environment, and one falling back to the default config runs a `globalSetup` that brings the shared compose stack up and tears it down, recreating the DB mid-run for every sibling. **The snapshot does not cover this** — the compose project name comes from the environment, not the working directory, so three agents on three copies still collide. But `./agent-test` does not fix it either: it exports **one** `TEST_COMPOSE_PROJECT`/port triple per *worktree* (`ab-<issue>`), so N concurrent specialists sharing it tear down each other's postgres mid-run. Observed: a probe returned `No test files found / No such container` and would have read as a test failure.
|
|
82
|
+
- **Give specialists a stack-free test command, not the worktree's runner, and their own scratch dir** — `<scratch>/pr<N>/fleet-review-<key>/`, resolved to an absolute path that you write into that specialist's prompt; a specialist never derives its own. Specialists inherit no environment, and one falling back to the default config runs a `globalSetup` that brings the shared compose stack up and tears it down, recreating the DB mid-run for every sibling. **The snapshot does not cover this** — the compose project name comes from the environment, not the working directory, so three agents on three copies still collide. But `./agent-test` does not fix it either: it exports **one** `TEST_COMPOSE_PROJECT`/port triple per *worktree* (`ab-<issue>`), so N concurrent specialists sharing it tear down each other's postgres mid-run. Observed: a probe returned `No test files found / No such container` and would have read as a test failure. So tell specialists the repo's test command directly — the Test entrypoint, read off the Recipe cache the way step 3 reads it (`derive-testcmd.sh <worktree or main checkout> test` prints it; the snapshot carries no `.fleet/` cache, so pointing it at the snapshot refuses) — when it brings up no shared stack, and the repo's own stack-free config when it does; never name a config the repo does not have, which sends every specialist into a missing-file error that reads as a broken tree rather than a bad instruction. Keep `./agent-test` for your own verification. **Hand that command with its reading rule: `tests 0` is a failed run, not a pass.** A command is not self-checking: a path glob or filter that matches nothing can exit 0 having run nothing. Measured with `node --test` over a glob matching no file: bash passes the pattern through literally, node globs it itself and finds no files, and you get `ℹ tests 0` / exit 0 — and `# tests 0` to a pipe on older node, so a rule that greps only for `ℹ` finds nothing in a CI log and cannot tell a zero-test run from a full one. Do not count on the shell to catch it: zsh errors (`no matches found`, exit 1) and bash does not, so the guard has to be the reading rule. **And zero is not the only no-work count.** **0 passes with no failures** is everything skipped, and a count **materially below the full suite** is a partial copy of the tree — which is the *documented* behavior, since the prompt tells specialists to run their probe work inside their own copy of the snapshot and nothing anchors an expected count for them. `review-core.mjs` classifies the first itself, in `unrunReason`; the second it structurally cannot, because that function is kept pure and so never learns how many tests the tree has. Handing the rule over with the command is the only guard the partial-copy case has. Carry the rule, not the command: wherever a repo has a stack, N concurrent specialists sharing one runner collide, and the escape is whatever that repo's stack-free config is.
|
|
83
83
|
- **Completion is not delivery — but you can reach a specialist, so never wait idle.** A specialist you dispatched reports when it settles; if a report is missing, `hub send` it by name and ask for the report. Judge delivery by whether a report came out. **Do not settle a dimension you hold no report for**: ask the specialist, then the controller by name; only when neither answers, rule and name that dimension **unrun** — a killed specialist never reports, so waiting on one is unbounded. Never ship "dispatched, never delivered" as settled; a ruling citing a report you do not hold gets verified from source, never applied on trust.
|
|
84
84
|
- **Apply only once the fan-out is settled — every report either in hand or accounted for.** Collection is not guaranteed; a report that never arrived is an open dimension, not a smaller set. And editing while they read is the same defect as probing — one specialist reviewed uncommitted code that was never in the PR diff.
|
|
85
85
|
|
package/package.json
CHANGED
|
@@ -509,10 +509,14 @@ const FILES = [
|
|
|
509
509
|
/re-issuing\s+(?:\/\/\s*)?it\b/,
|
|
510
510
|
/same remedy the\s+(?:\/\/\s*)?CI gate/,
|
|
511
511
|
],
|
|
512
|
+
// #2677. Both comments once attributed omp's 300s command deadline to
|
|
513
|
+
// run-merge-bot.md's CI wait rule, which names no 300s — it names the ~60s
|
|
514
|
+
// auto-background and the 5-6 minute CI cycle. The 300s is stated as omp's
|
|
515
|
+
// own default now, and only the CI cycle still cites the rule.
|
|
512
516
|
// One needle per SITE — either can revert independently of the other.
|
|
513
517
|
live: [
|
|
514
|
-
"
|
|
515
|
-
"run-merge-bot.md's CI wait rule
|
|
518
|
+
"its command deadline defaults to 300s (the bash tool's default `timeout`)",
|
|
519
|
+
"the 5-6 minute CI cycle run-merge-bot.md's CI wait rule states",
|
|
516
520
|
],
|
|
517
521
|
},
|
|
518
522
|
{
|
|
@@ -0,0 +1,126 @@
|
|
|
1
|
+
import { test } from "node:test";
|
|
2
|
+
import assert from "node:assert/strict";
|
|
3
|
+
import { readFileSync, readdirSync } from "node:fs";
|
|
4
|
+
import { join } from "node:path";
|
|
5
|
+
import { between, bullet, paragraph, phrase } from "./prose-pin.mjs";
|
|
6
|
+
|
|
7
|
+
// The prose that describes the consumer repo as "whatever the Recipe says"
|
|
8
|
+
// rather than as a Node project: every shipped runbook and the requirement
|
|
9
|
+
// docs name the Recipe's Install step and Test entrypoint, and none of them
|
|
10
|
+
// names a Node package manager, a lockfile or a Node-only tool as the way a
|
|
11
|
+
// consumer is installed or tested. Each pin below holds ONE contiguous clause
|
|
12
|
+
// of that vocabulary inside the bounded slice that carries it.
|
|
13
|
+
const PLUGIN = join(import.meta.dirname, "..");
|
|
14
|
+
const ROOT = join(PLUGIN, "..");
|
|
15
|
+
const readPlugin = (rel) => readFileSync(join(PLUGIN, rel), "utf8");
|
|
16
|
+
const readRoot = (rel) => readFileSync(join(ROOT, rel), "utf8");
|
|
17
|
+
const RUN_TEAM = readPlugin("skills/run-team/SKILL.md");
|
|
18
|
+
const REAPING = readPlugin("skills/run-team/references/reaping.md");
|
|
19
|
+
const ISOLATION = readPlugin("skills/run-team/references/isolation.md");
|
|
20
|
+
const REVIEW_AND_FIX = readPlugin("commands/review-and-fix.md");
|
|
21
|
+
const REQUIREMENTS = readRoot("docs/requirements.md");
|
|
22
|
+
const PULL_AND_CLAIM = readRoot("docs/components/pull-and-claim.md");
|
|
23
|
+
const EXTERNAL = readRoot("docs/research/external-assumptions.md");
|
|
24
|
+
|
|
25
|
+
// A Node-consumer assumption: a lockfile or `node_modules` named as the
|
|
26
|
+
// consumer's install state, or a Node package-manager / runner invocation
|
|
27
|
+
// named as how a consumer is installed or checked.
|
|
28
|
+
const NODE_CONSUMER = /npm ci|node_modules|package-lock|pnpm-lock|yarn\.lock|npm install|npx (?:vitest|tsc)/;
|
|
29
|
+
|
|
30
|
+
const markdownUnder = (dir) =>
|
|
31
|
+
readdirSync(join(PLUGIN, dir), { recursive: true })
|
|
32
|
+
.filter((rel) => rel.endsWith(".md"))
|
|
33
|
+
.map((rel) => join(PLUGIN, dir, rel));
|
|
34
|
+
|
|
35
|
+
test("the sweep's own pattern matches each Node-consumer form it exists to catch", () => {
|
|
36
|
+
for (const sample of ["npm ci", "its node_modules still on disk", "package-lock.json", "pnpm-lock.yaml", "yarn.lock", "run npm install", "npx tsc --noEmit", "npx vitest run"]) {
|
|
37
|
+
assert.match(sample, NODE_CONSUMER, `the pattern no longer matches "${sample}"`);
|
|
38
|
+
}
|
|
39
|
+
});
|
|
40
|
+
|
|
41
|
+
test("no runbook, command, agent definition or requirement doc names a Node-consumer install state or tool", () => {
|
|
42
|
+
const files = [
|
|
43
|
+
...markdownUnder("skills"),
|
|
44
|
+
...markdownUnder("commands"),
|
|
45
|
+
...markdownUnder("agents"),
|
|
46
|
+
join(ROOT, "docs", "requirements.md"),
|
|
47
|
+
join(ROOT, "docs", "components", "pull-and-claim.md"),
|
|
48
|
+
join(ROOT, "README.md"),
|
|
49
|
+
];
|
|
50
|
+
assert.ok(files.includes(join(PLUGIN, "skills", "run-team", "SKILL.md")), "the sweep no longer reaches run-team/SKILL.md");
|
|
51
|
+
const offenders = files.filter((f) => NODE_CONSUMER.test(readFileSync(f, "utf8")));
|
|
52
|
+
assert.deepEqual(offenders, [], `Node-consumer wording is back in: ${offenders.join(", ")}`);
|
|
53
|
+
});
|
|
54
|
+
|
|
55
|
+
test("the claim step runs the Recipe's Install step in the worktree, read from the cache", () => {
|
|
56
|
+
const step = bullet(PULL_AND_CLAIM, "4. **Claim.**", "5. **Dispatch**", "pull-and-claim.md step 4");
|
|
57
|
+
assert.match(step, phrase("runs the [Recipe](recipe.md)'s Install step there — read from the Recipe cache, refusing if it changed the tree"));
|
|
58
|
+
});
|
|
59
|
+
|
|
60
|
+
test("requirements 2.3 describes the Recipe derivation: any technology, what it reads, what proves it, what it caches", () => {
|
|
61
|
+
const section = between(REQUIREMENTS, "### 2.3 Installable and testable — HARD", "### 2.4", "requirements.md section 2.3");
|
|
62
|
+
assert.match(section, phrase("Any technology. Fleet derives your repo's Recipe (Install step + Test entrypoint) by agent reasoning"));
|
|
63
|
+
const reads = bullet(section, "- **What the derivation reads:**", "- **What the proof requires:**", "requirements.md 2.3 derivation bullet");
|
|
64
|
+
assert.match(reads, phrase("nothing checks for a particular manifest or lockfile."));
|
|
65
|
+
const proof = bullet(section, "- **What the proof requires:**", "- **The cache:**", "requirements.md 2.3 proof bullet");
|
|
66
|
+
assert.match(proof, phrase("`RECIPE NOT PROVEN` writes no cache and halts the run."));
|
|
67
|
+
const cache = bullet(section, "- **The cache:**", "Hard rules for any technology:", "requirements.md 2.3 cache bullet");
|
|
68
|
+
assert.match(cache, phrase("Every claim reads both commands from it and every review its Test entrypoint; a missing or unproven cache is a refusal naming the derivation step, never a guess."));
|
|
69
|
+
assert.match(section, phrase("~/.fleet/bin/fleet-run derive-testcmd.sh . install && ~/.fleet/bin/fleet-run derive-testcmd.sh . test"));
|
|
70
|
+
});
|
|
71
|
+
|
|
72
|
+
test("requirements 1.1 adds whatever the Recipe runs to the tool floor", () => {
|
|
73
|
+
assert.match(
|
|
74
|
+
between(REQUIREMENTS, "| `shasum` | any |", "### 1.2", "requirements.md section 1.1 table tail"),
|
|
75
|
+
phrase("Plus whatever your Recipe runs (§2.3): `mvn`, `go`, `cargo`, …"),
|
|
76
|
+
);
|
|
77
|
+
});
|
|
78
|
+
|
|
79
|
+
test("the research note marks the Node-only consumer rows superseded by the Recipe ruling", () => {
|
|
80
|
+
const row = bullet(EXTERNAL, "- ~~Contains `package.json`", "\n- `node` on PATH", "external-assumptions.md consumer-repo row");
|
|
81
|
+
assert.match(row, phrase("Superseded by ADR 0015: the claim reads the Install step and Test entrypoint from the Recipe cache, and nothing checks for a manifest, a test-file name or a lockfile."));
|
|
82
|
+
const implied = bullet(EXTERNAL, "- Consumer must be a Node project", "\n- `python3` and `shasum` on PATH", "external-assumptions.md implied-only row");
|
|
83
|
+
assert.match(implied, phrase("Ruled 2026-09-28: ADR 0015 — any technology, Recipe by agent reasoning"));
|
|
84
|
+
});
|
|
85
|
+
|
|
86
|
+
test("review-and-fix step 3 reads the standalone testCmd off the Recipe cache and derives when there is none", () => {
|
|
87
|
+
const step = bullet(REVIEW_AND_FIX, "3. **Run `testCmd` before you commit**", "\n4. **", "review-and-fix.md step 3");
|
|
88
|
+
assert.match(step, phrase("the Test entrypoint in the repo's Recipe cache (`derive-testcmd.sh <repo> test` prints it)"));
|
|
89
|
+
assert.match(step, phrase("run the Recipe derivation step first"));
|
|
90
|
+
assert.doesNotMatch(step, /node --test/, "step 3 names this repository's own test command again");
|
|
91
|
+
});
|
|
92
|
+
|
|
93
|
+
test("review-and-fix's git archive bullet runs the Install step in the copy, read from a tree that holds the cache", () => {
|
|
94
|
+
const archive = bullet(REVIEW_AND_FIX, "- **`git archive` carries tracked files only.**", "\n- **Initialize the copy", "review-and-fix.md git archive bullet");
|
|
95
|
+
assert.match(archive, phrase("run the Install step in the copy (read it off the worktree or main checkout, `derive-testcmd.sh <worktree or main checkout> install`; pointed at the copy it refuses, since a copy holds no Recipe cache)"));
|
|
96
|
+
});
|
|
97
|
+
|
|
98
|
+
test("run-team: a tree with no runner gets the Recipe's Test entrypoint only when it brings up no shared stack", () => {
|
|
99
|
+
const p = paragraph(RUN_TEAM, "**A reused worktree may lack the runner** — no longer here", "run-team/SKILL.md reused-worktree runner paragraph");
|
|
100
|
+
assert.match(p, phrase("the Recipe's Test entrypoint (`~/.fleet/bin/fleet-run derive-testcmd.sh <main checkout> test` prints it) only when it brings up no shared stack to collide on, else that repo's own stack-free command."));
|
|
101
|
+
});
|
|
102
|
+
|
|
103
|
+
test("run-team: testCmd comes from the Recipe cache and names no repository's own command", () => {
|
|
104
|
+
const p = paragraph(RUN_TEAM, "**Where `testCmd` comes from:**", "run-team/SKILL.md testCmd source paragraph");
|
|
105
|
+
assert.match(p, phrase("the repository's Test entrypoint, out of the Recipe cache phase 0's Recipe derivation step proved — `~/.fleet/bin/fleet-run derive-testcmd.sh . test` prints it — and the one you hand specialists"));
|
|
106
|
+
assert.doesNotMatch(p, /node --test/, "the testCmd source paragraph names this repository's own test command again");
|
|
107
|
+
});
|
|
108
|
+
|
|
109
|
+
test("run-team: the quick-install red flag points at the Recipe's Install step, not at an install command", () => {
|
|
110
|
+
const flag = bullet(RUN_TEAM, "- \"A quick install to set up the worktree\"", "- \"I'm on my own copy, so I'm isolated\"", "run-team/SKILL.md quick-install red flag");
|
|
111
|
+
assert.match(flag, phrase("the claim already ran the Recipe's Install step; any other install can rewrite the lockfile for the whole repo."));
|
|
112
|
+
});
|
|
113
|
+
|
|
114
|
+
test("a merged branch's stale worktree is said to keep its installed dependencies, in SKILL.md and in reaping.md", () => {
|
|
115
|
+
const skill = paragraph(RUN_TEAM, "The merge bot deletes the remote branch after each merge", "run-team/SKILL.md reap paragraph");
|
|
116
|
+
assert.match(skill, phrase("with its worktree — and its installed dependencies — still on disk."));
|
|
117
|
+
const evidence = paragraph(REAPING, "Merge deletes remote branch, leaves local branch", "reaping.md reap-after-each-merge paragraph");
|
|
118
|
+
assert.match(evidence, phrase("with worktree — and its installed dependencies — still on disk."));
|
|
119
|
+
});
|
|
120
|
+
|
|
121
|
+
test("the isolation reference says symlinking installed dependencies in does not isolate the stack, and reproduces a diagnostic by re-running the check", () => {
|
|
122
|
+
const stack = paragraph(ISOLATION, "Private copy and test command solve different problems", "isolation.md stack-isolation paragraph");
|
|
123
|
+
assert.match(stack, phrase("Symlinking installed dependencies in does not help."));
|
|
124
|
+
const diag = paragraph(ISOLATION, "Probe copies carry same filenames as real tree", "isolation.md diagnostics paragraph");
|
|
125
|
+
assert.match(diag, phrase("(re-run the check that raised it — the typechecker, the linter — from there)"));
|
|
126
|
+
});
|
|
@@ -62,8 +62,8 @@ export function interval({ quiet, base, ceiling, multiplier }) {
|
|
|
62
62
|
//
|
|
63
63
|
// Nothing lets one command block for twenty minutes. Measured: omp
|
|
64
64
|
// auto-backgrounds any foreground command at 60s (`bash.autoBackground`), and
|
|
65
|
-
// its command deadline defaults to 300s
|
|
66
|
-
//
|
|
65
|
+
// its command deadline defaults to 300s (the bash tool's default
|
|
66
|
+
// `timeout`).
|
|
67
67
|
//
|
|
68
68
|
// So a long interval is served by SEVERAL holds, and the elapsed total is
|
|
69
69
|
// persisted rather than counted in prose: the script says how much remains, so
|
|
@@ -96,8 +96,8 @@ const OPTIONS = {
|
|
|
96
96
|
ceiling: { type: "string", default: String(DEFAULT_CEILING_S) },
|
|
97
97
|
multiplier: { type: "string", default: "2" },
|
|
98
98
|
// Per-invocation blocking budget, under omp's defaults. 240s leaves
|
|
99
|
-
// headroom below omp's 300s deadline,
|
|
100
|
-
//
|
|
99
|
+
// headroom below omp's 300s command deadline, which is itself shorter than
|
|
100
|
+
// the 5-6 minute CI cycle run-merge-bot.md's CI wait rule states. Raise
|
|
101
101
|
// it only together with the tool call's own timeout.
|
|
102
102
|
hold: { type: "string", default: "240" },
|
|
103
103
|
// The deliberate stop, #1597. No default and no boolean spelling: the flag
|
|
@@ -71,11 +71,12 @@ for (const [name, text, from, to] of slices) {
|
|
|
71
71
|
// target's presence needs its own pin or the pointer orphans in silence. The end
|
|
72
72
|
// anchor is the generic next-bullet marker, not the following bullet's wording,
|
|
73
73
|
// which would make an unrelated rewrite of that bullet a boundary failure here.
|
|
74
|
-
//
|
|
75
|
-
//
|
|
74
|
+
// Bounded to the bullet so that a copy of the phrase anywhere else in the file
|
|
75
|
+
// cannot keep the pin satisfied with this bullet deleted.
|
|
76
76
|
test("review-and-fix.md still hands specialists the command both guards point at", () => {
|
|
77
77
|
const bullet = between(REVIEW_AND_FIX, "Give specialists a stack-free test command", "\n- **", "review-and-fix.md");
|
|
78
|
-
assert.match(bullet, phrase("
|
|
78
|
+
assert.match(bullet, phrase("the Test entrypoint, read off the Recipe cache"));
|
|
79
|
+
assert.match(bullet, phrase("`derive-testcmd.sh <worktree or main checkout> test` prints it"));
|
|
79
80
|
});
|
|
80
81
|
|
|
81
82
|
// #1150: any member that dispatches a child writes the child's absolute scratch
|
package/skills/run-team/SKILL.md
CHANGED
|
@@ -824,9 +824,10 @@ prior-run worktree). #55 tracked the runner, so any worktree checked out from
|
|
|
824
824
|
`origin/main` now carries it. A tree that is NOT a checkout — a `git archive`
|
|
825
825
|
snapshot, a `cp -R` subset — still has whatever was copied into it, and a repo
|
|
826
826
|
that tracks no runner never had one: there, tell the member the runner is absent
|
|
827
|
-
and to run docker-free suites directly
|
|
828
|
-
|
|
829
|
-
|
|
827
|
+
and to run docker-free suites directly — the Recipe's Test entrypoint
|
|
828
|
+
(`~/.fleet/bin/fleet-run derive-testcmd.sh <main checkout> test` prints it) only
|
|
829
|
+
when it brings up no shared stack to collide on, else that repo's own stack-free
|
|
830
|
+
command.
|
|
830
831
|
|
|
831
832
|
**A reused worktree may also be on the wrong COMMIT.** `git worktree add <path>
|
|
832
833
|
<branch>` checks out the existing LOCAL branch and never consults the remote, so
|
|
@@ -2212,10 +2213,10 @@ reads gates and never injects a fault.
|
|
|
2212
2213
|
|
|
2213
2214
|
**Where `testCmd` comes from:** the repository's Test entrypoint, out of the
|
|
2214
2215
|
Recipe cache phase 0's Recipe derivation step proved —
|
|
2215
|
-
`~/.fleet/bin/fleet-run derive-testcmd.sh . test` prints it
|
|
2216
|
-
|
|
2217
|
-
|
|
2218
|
-
|
|
2216
|
+
`~/.fleet/bin/fleet-run derive-testcmd.sh . test` prints it — and the one you
|
|
2217
|
+
hand specialists per **Give specialists a stack-free test command** above. Pass
|
|
2218
|
+
the same string to the review, to the fix-applier, and to the finisher — whose
|
|
2219
|
+
duty-2 mutation
|
|
2219
2220
|
gate runs it too — so every gate runs one command. Omit it from the review args
|
|
2220
2221
|
and the review (`review-core.mjs`, in a runner) reads the same cache itself,
|
|
2221
2222
|
refusing outright when there is none; the fix-applier has no such fallback, so
|
|
@@ -3095,7 +3096,7 @@ ancestry proof, post-rebase red triage. Do not restate them here.
|
|
|
3095
3096
|
|
|
3096
3097
|
The merge bot deletes the remote branch after each merge (`delete-merged-branch.sh`,
|
|
3097
3098
|
in `run-merge-bot.md` step 4), leaving the local branch `[gone]` with its
|
|
3098
|
-
worktree — and its
|
|
3099
|
+
worktree — and its installed dependencies — still on disk. Reap after **each** merge pass,
|
|
3099
3100
|
not once at the end: a stale worktree still answers `git worktree list`, so the
|
|
3100
3101
|
in-flight probe (`inflight.sh`, run by the Shortlist and by every Pull) reads an
|
|
3101
3102
|
already-merged ticket as taken and the queue quietly shrinks. The trigger is the
|
|
@@ -3523,8 +3524,9 @@ failures arrive as *wrong findings*, not errors:
|
|
|
3523
3524
|
under your own path. See references/isolation.md.
|
|
3524
3525
|
- **IDE/harness diagnostics attribute by bare filename, with no path.** **Never
|
|
3525
3526
|
relay a diagnostic without reproducing it in that member's specific worktree**
|
|
3526
|
-
(
|
|
3527
|
-
so a sibling's throwaway mutation reads exactly like a live
|
|
3527
|
+
(re-run the check that raised it from there): probe copies carry the real
|
|
3528
|
+
tree's filenames, so a sibling's throwaway mutation reads exactly like a live
|
|
3529
|
+
worktree's error.
|
|
3528
3530
|
See references/isolation.md.
|
|
3529
3531
|
|
|
3530
3532
|
Suspect a neighbour before a member's own diff — for unexplained failures, and
|
|
@@ -3845,8 +3847,8 @@ Plus a queue-depth line: shortlist, supply, whether triage was suggested.
|
|
|
3845
3847
|
Ask what scope it searched.
|
|
3846
3848
|
- "Five implementers = five times throughput" → reviews are 3-5x longer. It means
|
|
3847
3849
|
a backlog.
|
|
3848
|
-
- "
|
|
3849
|
-
for the whole repo.
|
|
3850
|
+
- "A quick install to set up the worktree" → the claim already ran the Recipe's
|
|
3851
|
+
Install step; any other install can rewrite the lockfile for the whole repo.
|
|
3850
3852
|
- "I'm on my own copy, so I'm isolated" → not from the docker stack.
|
|
3851
3853
|
- "`commit-commands:clean_gone` printed nothing, the tree is clean" → its
|
|
3852
3854
|
`[gone]` detection works fine; it just runs `-D`/`--force` with no merged
|
|
@@ -10,7 +10,7 @@ The runner is re-materialized on every invocation, so no copy of it survives a r
|
|
|
10
10
|
|
|
11
11
|
## Filesystem isolation is not stack isolation
|
|
12
12
|
|
|
13
|
-
Private copy and test command solve different problems; conflating them is how second gets skipped: compose project name comes from environment, not working directory, so three agents on three snapshots still collide on one postgres. Symlinking
|
|
13
|
+
Private copy and test command solve different problems; conflating them is how second gets skipped: compose project name comes from environment, not working directory, so three agents on three snapshots still collide on one postgres. Symlinking installed dependencies in does not help. "I'm on my own copy" is exactly the intuition that skips command — say both, every time. Which command depends on audience: member in worktree uses `./agent-test`; specialist on snapshot does not — the runner that bootstrap (#55) materializes derives its isolation triple from the directory name, and a snapshot is not a claimed worktree, so every snapshot gets the same fixed ports (measured in a snapshot: `ports derive from the issue number: postgres=16000 ollama=22000`) — and takes one `review-and-fix.md` hands out. The bootstrap no longer FAILS there: since #1056 the snapshot is a git repository, so `claim-ticket.sh --write-runner` succeeds in it, and the reason to hand a specialist a different command is the stack rather than a missing repo.
|
|
14
14
|
|
|
15
15
|
## Scratchpad paths need two levels — the member's own partition, then one directory per child it dispatches
|
|
16
16
|
|
|
@@ -20,4 +20,4 @@ The axis is two-sided: the member partitions, and the parent assigns each child'
|
|
|
20
20
|
|
|
21
21
|
## IDE/harness diagnostics attribute by bare filename, with no path
|
|
22
22
|
|
|
23
|
-
Probe copies carry same filenames as real tree, so specialist's throwaway mutation surfaces as errors that read exactly like live worktree's — and line numbers can plausibly line up with real in-flight edits. Never relay diagnostic without reproducing it in that member's specific worktree (
|
|
23
|
+
Probe copies carry same filenames as real tree, so specialist's throwaway mutation surfaces as errors that read exactly like live worktree's — and line numbers can plausibly line up with real in-flight edits. Never relay diagnostic without reproducing it in that member's specific worktree (re-run the check that raised it — the typechecker, the linter — from there). Ran twice in one session: clean first time (sibling's probe), genuinely broken second. Telling implementer to chase phantom in file it is mid-rewrite on is expensive failure.
|
|
@@ -4,7 +4,7 @@ Why reap runs after each merge pass, why `commit-commands:clean_gone` disqualifi
|
|
|
4
4
|
|
|
5
5
|
## Reap after each merge pass, not once at the end
|
|
6
6
|
|
|
7
|
-
Merge deletes remote branch, leaves local branch `[gone]` with worktree — and its
|
|
7
|
+
Merge deletes remote branch, leaves local branch `[gone]` with worktree — and its installed dependencies — still on disk. Stale worktree still answers `git worktree list`, so the in-flight probe (`inflight.sh`, run by the Shortlist and again by every Pull) reads already-merged ticket as taken and queue quietly shrinks as run goes on.
|
|
8
8
|
|
|
9
9
|
## Why `commit-commands:clean_gone` is disqualified
|
|
10
10
|
|