@feigi/fleet-ctl 3.21.13 → 3.22.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/package.json +1 -1
- package/scripts/ambient-git-vars-mjs-prose.test.mjs +7 -3
- package/scripts/cell-readout.mjs +148 -0
- package/scripts/cell-readout.test.mjs +341 -0
- package/scripts/fleet-tick.mjs +46 -5
- package/scripts/fleet-tick.test.mjs +132 -4
- package/scripts/member-outcomes-header.test.mjs +45 -85
- package/scripts/member-outcomes.mjs +2 -2
- package/scripts/member-outcomes.test.mjs +3 -3
- package/scripts/tier-outcomes-header.test.mjs +37 -62
- package/skills/run-team/SKILL.md +39 -34
package/skills/run-team/SKILL.md
CHANGED
|
@@ -1125,31 +1125,29 @@ Pull is one member, so a rate counted per Pull is a rate counted per member —
|
|
|
1125
1125
|
one implementer in five, however the Pulls fall across the run — and the
|
|
1126
1126
|
ledger, not your memory, is what still holds the count after a compaction.
|
|
1127
1127
|
|
|
1128
|
-
**Do not label it anywhere — the dispatch record already carries it.**
|
|
1129
|
-
|
|
1130
|
-
|
|
1131
|
-
|
|
1132
|
-
|
|
1133
|
-
definition
|
|
1134
|
-
|
|
1135
|
-
|
|
1136
|
-
|
|
1137
|
-
|
|
1138
|
-
|
|
1139
|
-
|
|
1140
|
-
|
|
1141
|
-
|
|
1142
|
-
|
|
1143
|
-
|
|
1144
|
-
|
|
1145
|
-
|
|
1146
|
-
|
|
1147
|
-
|
|
1148
|
-
|
|
1149
|
-
|
|
1150
|
-
|
|
1151
|
-
to the same model (a `modelRoles` override, measured on omp 2026-09-12)
|
|
1152
|
-
controls nothing.
|
|
1128
|
+
**Do not label it anywhere — the dispatch record already carries it.** A
|
|
1129
|
+
cell's comparisons are a query, and `cell-readout.mjs` is that query: it joins
|
|
1130
|
+
`docs/metrics/ticket-features.tsv`, whose `chosen_cell` names the cell the
|
|
1131
|
+
router chose, to `docs/metrics/member-outcomes.tsv` on `session`+`agent`, and
|
|
1132
|
+
admits a row only when its `subagent_type` is that cell's
|
|
1133
|
+
`fleet-implementer-<cell>` definition and its `effort` is the cell's level.
|
|
1134
|
+
That column is the record of what each member was dispatched AS, scraped like
|
|
1135
|
+
every other column, so it survives the file's regeneration and mislabels no
|
|
1136
|
+
historical row — the pre-cutover `fleet-implementer` and
|
|
1137
|
+
`fleet-implementer-alt` rows are no cell's, and the readout never admits them.
|
|
1138
|
+
A hand-set column would fail both of those tests, and one derived against
|
|
1139
|
+
today's declared tiers would mislabel every historical row.
|
|
1140
|
+
|
|
1141
|
+
**The readout counts DELIBERATE comparisons.** An earlier rule asked
|
|
1142
|
+
only for a `session`+`role` carrying more than one distinct `model`, which any
|
|
1143
|
+
session that happened to stage a routine ticket beside a correction one
|
|
1144
|
+
satisfies. Measured 2026-09-12 on one corpus: 198 keys as that query was
|
|
1145
|
+
literally written, 35 sessions across 21 `run_date`s once restricted to
|
|
1146
|
+
implementers, against 17 real pairs across 8 — so the gate below read itself
|
|
1147
|
+
past its floor on sessions where nothing had been controlled. Dispatch alone
|
|
1148
|
+
is not enough either: the two rows must actually have RUN a different model or
|
|
1149
|
+
effort, because a deliberate dispatch whose two sides resolve to the same model
|
|
1150
|
+
(a `modelRoles` override, measured on omp 2026-09-12) controls nothing.
|
|
1153
1151
|
|
|
1154
1152
|
**Why one in five within the run, not a week of one tier then a week of the other:**
|
|
1155
1153
|
tier would then be confounded with calendar date and therefore with prompt
|
|
@@ -1397,11 +1395,14 @@ pooled across cells, and stop until
|
|
|
1397
1395
|
there are at least ten of them across five or more distinct `run_date`s; below
|
|
1398
1396
|
that, a comparison count is a number, not evidence, and the last guard fired
|
|
1399
1397
|
with n=1 on the control side. **Report the count the per-cell readout prints,
|
|
1400
|
-
never one from any other query:**
|
|
1401
|
-
|
|
1402
|
-
|
|
1403
|
-
|
|
1404
|
-
|
|
1398
|
+
never one from any other query:** `~/.fleet/bin/fleet-run cell-readout.mjs`
|
|
1399
|
+
prints `<cell> <comparisons> <run_dates> <resolved models>` for each cell past
|
|
1400
|
+
that floor and a cell below it only as a count on stderr; a `mixed (…)` models
|
|
1401
|
+
column means the cell's history spans more than one model behind its role.
|
|
1402
|
+
The pre-cell pair query counted pairs that are no cell's comparisons, and the
|
|
1403
|
+
query before that one counted every session whose implementers merely
|
|
1404
|
+
differed, so it cleared this floor by an order of magnitude while the
|
|
1405
|
+
controlled comparison did not exist yet.
|
|
1405
1406
|
|
|
1406
1407
|
Why the agent body carries what it does — read this before editing any
|
|
1407
1408
|
`fleet-implementer-<cell>` file, and keep every cell's body byte-identical. Each rule in the body's
|
|
@@ -2031,9 +2032,12 @@ is never something to assume.
|
|
|
2031
2032
|
`fix-pr-<pr#>` is one; a finisher and a CI wait are zero; and a PR holds at most
|
|
2032
2033
|
one slot at a time. In-flight reviews are derived, never remembered: a row with
|
|
2033
2034
|
`review=` and no `reviewed=` is one, which is why the token goes on before
|
|
2034
|
-
anything else.
|
|
2035
|
-
|
|
2036
|
-
|
|
2035
|
+
anything else. A review whose PR has left the open list is probed: when
|
|
2036
|
+
`gh pr view` reports that PR MERGED or CLOSED, the tick drops the review,
|
|
2037
|
+
since none runs against a finished PR, and a probe that
|
|
2038
|
+
cannot answer keeps the slot. The defaults are 2 implementers and 6 reviewers,
|
|
2039
|
+
and neither is a hard cap. The cap bounds units, not the agents a review fans
|
|
2040
|
+
out to — 1 snapshot + up to 6 specialists + 2 refuters per critical/important finding — and
|
|
2037
2041
|
how many reviews may be in flight at once, inside it, is `--max-reviews`. The
|
|
2038
2042
|
slots a review does not hold still serve fix-appliers.
|
|
2039
2043
|
|
|
@@ -3703,7 +3707,8 @@ written with `row` (**Reviewers**): `review=wf:<runId>` |
|
|
|
3703
3707
|
`review=member:review-pr-<n>` | `review=fallback:review-pr-<n>[-b]` at launch,
|
|
3704
3708
|
settled dead as `…=failed`, then `reviewed=<head>:<survived>/<refuted>/<unverified>`
|
|
3705
3709
|
when the result lands. A row with `review=` and no `reviewed=` is a review in
|
|
3706
|
-
flight, and the tick counts it against the reviewer cap
|
|
3710
|
+
flight, and the tick counts it against the reviewer cap unless gh reports its PR
|
|
3711
|
+
MERGED or CLOSED. `ci=<run-id>:<attempt>:<conclusion>`
|
|
3707
3712
|
and `held-behind:#<lower>` are row tokens the same way (Phase 3), and so is
|
|
3708
3713
|
`conflict-hold:#<pr>` — the one row token the merge bot writes itself, naming
|
|
3709
3714
|
the row's own PR (`run-merge-bot.md` step 1). It is not an Exclusion: that
|