@feigi/fleet-ctl 3.21.13 → 3.22.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1125,31 +1125,29 @@ Pull is one member, so a rate counted per Pull is a rate counted per member —
1125
1125
  one implementer in five, however the Pulls fall across the run — and the
1126
1126
  ledger, not your memory, is what still holds the count after a compaction.
1127
1127
 
1128
- **Do not label it anywhere — the dispatch record already carries it.** The
1129
- pairing is a query over `docs/metrics/member-outcomes.tsv` (the exact awk sits
1130
- in that file's header): a `session` that ran BOTH pre-cutover implementer
1131
- definitions at DIFFERENT `model`s — one member whose `subagent_type` is
1132
- `fleet-implementer-alt`, another whose is `fleet-implementer`. Neither
1133
- definition is dispatched any more, so the query counts pre-cutover history
1134
- only: a session whose implementers ran as `fleet-implementer-<cell>` is not in
1135
- it until #2133 replaces it with the per-cell readout. That column is
1136
- the record of what each member was dispatched AS, scraped like every other
1137
- column, so it survives the file's regeneration and mislabels no historical row
1138
- — a pre-rule session carries no such dispatch to find. A hand-set column would
1139
- fail both of those tests, and one derived against today's declared tiers would
1140
- mislabel every historical row.
1141
-
1142
- **The query counts DELIBERATE pairs, and that is the whole of #1066.** The
1143
- rule this replaces asked only for a `session`+`role` carrying more than one
1144
- distinct `model`, which any session that happened to stage a routine ticket
1145
- beside a correction one satisfies. Measured 2026-09-12 on one corpus: 198
1146
- keys as that query is literally written, 35 sessions across 21 `run_date`s
1147
- once restricted to implementers, against 17 real pairs across 8 — so the gate
1148
- below read itself past its floor on sessions where nothing had been
1149
- controlled. Dispatch alone is not enough either: both arms must actually have
1150
- RUN different models, because a deliberate alt dispatch whose two arms resolve
1151
- to the same model (a `modelRoles` override, measured on omp 2026-09-12)
1152
- controls nothing.
1128
+ **Do not label it anywhere — the dispatch record already carries it.** A
1129
+ cell's comparisons are a query, and `cell-readout.mjs` is that query: it joins
1130
+ `docs/metrics/ticket-features.tsv`, whose `chosen_cell` names the cell the
1131
+ router chose, to `docs/metrics/member-outcomes.tsv` on `session`+`agent`, and
1132
+ admits a row only when its `subagent_type` is that cell's
1133
+ `fleet-implementer-<cell>` definition and its `effort` is the cell's level.
1134
+ That column is the record of what each member was dispatched AS, scraped like
1135
+ every other column, so it survives the file's regeneration and mislabels no
1136
+ historical row — the pre-cutover `fleet-implementer` and
1137
+ `fleet-implementer-alt` rows are no cell's, and the readout never admits them.
1138
+ A hand-set column would fail both of those tests, and one derived against
1139
+ today's declared tiers would mislabel every historical row.
1140
+
1141
+ **The readout counts DELIBERATE comparisons.** An earlier rule asked
1142
+ only for a `session`+`role` carrying more than one distinct `model`, which any
1143
+ session that happened to stage a routine ticket beside a correction one
1144
+ satisfies. Measured 2026-09-12 on one corpus: 198 keys as that query was
1145
+ literally written, 35 sessions across 21 `run_date`s once restricted to
1146
+ implementers, against 17 real pairs across 8 — so the gate below read itself
1147
+ past its floor on sessions where nothing had been controlled. Dispatch alone
1148
+ is not enough either: the two rows must actually have RUN a different model or
1149
+ effort, because a deliberate dispatch whose two sides resolve to the same model
1150
+ (a `modelRoles` override, measured on omp 2026-09-12) controls nothing.
1153
1151
 
1154
1152
  **Why one in five within the run, not a week of one tier then a week of the other:**
1155
1153
  tier would then be confounded with calendar date and therefore with prompt
@@ -1397,11 +1395,14 @@ pooled across cells, and stop until
1397
1395
  there are at least ten of them across five or more distinct `run_date`s; below
1398
1396
  that, a comparison count is a number, not evidence, and the last guard fired
1399
1397
  with n=1 on the control side. **Report the count the per-cell readout prints,
1400
- never one from any other query:** until that readout exists there is no count
1401
- to report, the pair query in `member-outcomes.tsv`'s header counts pre-cutover
1402
- pairs only, which are no cell's comparisons, and the query before that one
1403
- counted every session whose implementers merely differed, so it cleared this
1404
- floor by an order of magnitude while the controlled comparison did not exist yet.
1398
+ never one from any other query:** `~/.fleet/bin/fleet-run cell-readout.mjs`
1399
+ prints `<cell> <comparisons> <run_dates> <resolved models>` for each cell past
1400
+ that floor and a cell below it only as a count on stderr; a `mixed (…)` models
1401
+ column means the cell's history spans more than one model behind its role.
1402
+ The pre-cell pair query counted pairs that are no cell's comparisons, and the
1403
+ query before that one counted every session whose implementers merely
1404
+ differed, so it cleared this floor by an order of magnitude while the
1405
+ controlled comparison did not exist yet.
1405
1406
 
1406
1407
  Why the agent body carries what it does — read this before editing any
1407
1408
  `fleet-implementer-<cell>` file, and keep every cell's body byte-identical. Each rule in the body's
@@ -2031,9 +2032,12 @@ is never something to assume.
2031
2032
  `fix-pr-<pr#>` is one; a finisher and a CI wait are zero; and a PR holds at most
2032
2033
  one slot at a time. In-flight reviews are derived, never remembered: a row with
2033
2034
  `review=` and no `reviewed=` is one, which is why the token goes on before
2034
- anything else. The defaults are 2 implementers and 6 reviewers, and neither is
2035
- a hard cap. The cap bounds units, not the agents a review fans out to — 1
2036
- snapshot + up to 6 specialists + 2 refuters per critical/important finding — and
2035
+ anything else. A review whose PR has left the open list is probed: when
2036
+ `gh pr view` reports that PR MERGED or CLOSED, the tick drops the review,
2037
+ since none runs against a finished PR, and a probe that
2038
+ cannot answer keeps the slot. The defaults are 2 implementers and 6 reviewers,
2039
+ and neither is a hard cap. The cap bounds units, not the agents a review fans
2040
+ out to — 1 snapshot + up to 6 specialists + 2 refuters per critical/important finding — and
2037
2041
  how many reviews may be in flight at once, inside it, is `--max-reviews`. The
2038
2042
  slots a review does not hold still serve fix-appliers.
2039
2043
 
@@ -3703,7 +3707,8 @@ written with `row` (**Reviewers**): `review=wf:<runId>` |
3703
3707
  `review=member:review-pr-<n>` | `review=fallback:review-pr-<n>[-b]` at launch,
3704
3708
  settled dead as `…=failed`, then `reviewed=<head>:<survived>/<refuted>/<unverified>`
3705
3709
  when the result lands. A row with `review=` and no `reviewed=` is a review in
3706
- flight, and the tick counts it against the reviewer cap. `ci=<run-id>:<attempt>:<conclusion>`
3710
+ flight, and the tick counts it against the reviewer cap unless gh reports its PR
3711
+ MERGED or CLOSED. `ci=<run-id>:<attempt>:<conclusion>`
3707
3712
  and `held-behind:#<lower>` are row tokens the same way (Phase 3), and so is
3708
3713
  `conflict-hold:#<pr>` — the one row token the merge bot writes itself, naming
3709
3714
  the row's own PR (`run-merge-bot.md` step 1). It is not an Exclusion: that