pi-gauntlet 5.20.0 → 5.21.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,15 @@
1
1
  # Changelog
2
2
 
3
+ ## v5.21.0 - 2026-09-26
4
+
5
+ ### Added
6
+
7
+ - Council telemetry records chair activity and per-member severity/disposition counts in `derived.council`; `gauntlet-performance` groups council results by roster, reports chair summaries and contribution flags after five runs, and exposes the aggregates in JSON. The performance skill includes a council assessment in its reply.
8
+
9
+ ### Changed
10
+
11
+ - `roasting-the-spec` audit lines and brainstorming's gate template carry `[<severity>]` and `raised-by: [...]` on every `Applied:`/`Deferred:`/`Rejected:` item, one item per line, `none` for an empty list. `gauntlet-performance` is split into section modules under `src/bins/performance/`; existing `runs` and `by version` output is unchanged.
12
+
3
13
  ## v5.20.0 - 2026-09-26
4
14
 
5
15
  ### Changed
package/README.md CHANGED
@@ -69,7 +69,7 @@ Everything between gate 1 and gate 2 - task breakdown, implementation, both revi
69
69
 
70
70
  pi-gauntlet ships three kinds of pieces, layered on top of pi-cohort's dispatch:
71
71
 
72
- - **20 skills** - the workflow logic. Thirteen activate automatically when pi sees the matching kind of task, and each one gates the next: `brainstorming`, `writing-plans`, `roasting-the-spec`, `test-driven-development`, `subagent-driven-development`, `dispatching-parallel-agents`, `verification-before-completion`, `requesting-code-review`, `receiving-code-review`, `using-git-worktrees`, `finishing-a-development-branch`, `writing-skills`, `linear` (reads/searches/comments on/manages Linear tickets via the `linearis` CLI; owns all linearis mechanics and the `## Issue tracker` overrides schema; tracker-facing skills route to it). Seven more are explicit-invocation-only (`disable-model-invocation: true`): `shape-ticket` creates or repairs one tracker issue per run against a Context/Problem/Idea/Acceptance-Criteria template, gated by an AC integrity check, a cheap council roast, and a single human-confirmed write - run it with `/skill:shape-ticket`. `gatekeep-pr` is consent-gated pre-merge verification of a PR against its issue - read-only gathering, verification evidence resolved CI-first (green checks on the exact assessed head count as evidence; the project's verification command runs only as fallback), a rubric-based review, then a deterministic authorship-aware menu with stable finding IDs (P#/L#/C#/F#) and numbered pre-composed courses (fixes execute as a single parallel-safe wave: one gate run, one re-review, one push); nothing mutates (fixes, pushes, reviews, merges) until you pick a row - run it with `/skill:gatekeep-pr <pr>`. `check-delivery` is a post-merge detective control: proves an issue actually shipped (default-branch landing, delivery target, per-AC evidence) before its tracker status advances; it never writes a terminal status - run it with `/skill:check-delivery <ref>`. `chase-bug` is human-only bug triage: read-only root-cause discovery to an evidenced verdict menu (real bug -> ticket/brainstorm/hotfix/respond; five negative verdicts), then a gated response to the reporter for addressable origins (GitHub issue / tracker ticket) and a rendered verdict summary otherwise - it never fixes during triage; the hotfix row hands off to `skills/chase-bug/hotfix.md` after the menu - run it with `/skill:chase-bug`. `gauntlet-performance` reads the committed run telemetry (current repo by default, `--dir <path>` adds others) through the parse-only `gauntlet-performance` CLI and answers with one example-led recommendation, 3-5 cornerstone numbers, and a menu of at most three actions (render a report file, open the recommendation as a ticket via `shape-ticket`, drill into one run); it writes nothing unless you pick render - run it with `/skill:gauntlet-performance [--dir <path>]... [--since <version>]`. `gauntlet-handoff` ends a session whose flow a fresh session will continue: it invokes pi-cohort's `handoff` skill (`--out <path>` or `--key <name>`, default key = the run worktree's branch with `/` flattened to `-`, default file `<tmpdir>/pi-handoff/<key>.md`), then appends the gauntlet process-state section (phase/plan tracker status) per the shared contract `skills/gauntlet-resume/reference/brief-contract.md` - run it with `/skill:gauntlet-handoff [--out <path> | --key <name>]`. `gauntlet-resume` is the only way back into an interrupted flow from a fresh session: it takes a `gauntlet-handoff` brief (a file path, pasted text, or - with no arguments - a pick from the briefs under `<tmpdir>/pi-handoff/`), a bare worktree that already holds a spec, or a spec tracked on main under a configured spec dir (a spec seed: it creates or reuses the worktree through `using-git-worktrees` and pins the spec), restores phase/plan tracker state through the legal arming sequence (`start brainstorm`, `skip` with `resume:` reasons, `plan_check` before implement-or-later), and never infers approval from artifacts - run it with `/skill:gauntlet-resume [<brief-file> | <spec>.md | <worktree-name-or-path>]`.
72
+ - **20 skills** - the workflow logic. Thirteen activate automatically when pi sees the matching kind of task, and each one gates the next: `brainstorming`, `writing-plans`, `roasting-the-spec`, `test-driven-development`, `subagent-driven-development`, `dispatching-parallel-agents`, `verification-before-completion`, `requesting-code-review`, `receiving-code-review`, `using-git-worktrees`, `finishing-a-development-branch`, `writing-skills`, `linear` (reads/searches/comments on/manages Linear tickets via the `linearis` CLI; owns all linearis mechanics and the `## Issue tracker` overrides schema; tracker-facing skills route to it). Seven more are explicit-invocation-only (`disable-model-invocation: true`): `shape-ticket` creates or repairs one tracker issue per run against a Context/Problem/Idea/Acceptance-Criteria template, gated by an AC integrity check, a cheap council roast, and a single human-confirmed write - run it with `/skill:shape-ticket`. `gatekeep-pr` is consent-gated pre-merge verification of a PR against its issue - read-only gathering, verification evidence resolved CI-first (green checks on the exact assessed head count as evidence; the project's verification command runs only as fallback), a rubric-based review, then a deterministic authorship-aware menu with stable finding IDs (P#/L#/C#/F#) and numbered pre-composed courses (fixes execute as a single parallel-safe wave: one gate run, one re-review, one push); nothing mutates (fixes, pushes, reviews, merges) until you pick a row - run it with `/skill:gatekeep-pr <pr>`. `check-delivery` is a post-merge detective control: proves an issue actually shipped (default-branch landing, delivery target, per-AC evidence) before its tracker status advances; it never writes a terminal status - run it with `/skill:check-delivery <ref>`. `chase-bug` is human-only bug triage: read-only root-cause discovery to an evidenced verdict menu (real bug -> ticket/brainstorm/hotfix/respond; five negative verdicts), then a gated response to the reporter for addressable origins (GitHub issue / tracker ticket) and a rendered verdict summary otherwise - it never fixes during triage; the hotfix row hands off to `skills/chase-bug/hotfix.md` after the menu - run it with `/skill:chase-bug`. `gauntlet-performance` reads the committed run telemetry (current repo by default, `--dir <path>` adds others) through the parse-only `gauntlet-performance` CLI and answers with one example-led recommendation, 3-5 cornerstone numbers, a council assessment when available, and a menu of at most three actions (render a report file, open the recommendation as a ticket via `shape-ticket`, drill into one run); it writes nothing unless you pick render - run it with `/skill:gauntlet-performance [--dir <path>]... [--since <version>]`. `gauntlet-handoff` ends a session whose flow a fresh session will continue: it invokes pi-cohort's `handoff` skill (`--out <path>` or `--key <name>`, default key = the run worktree's branch with `/` flattened to `-`, default file `<tmpdir>/pi-handoff/<key>.md`), then appends the gauntlet process-state section (phase/plan tracker status) per the shared contract `skills/gauntlet-resume/reference/brief-contract.md` - run it with `/skill:gauntlet-handoff [--out <path> | --key <name>]`. `gauntlet-resume` is the only way back into an interrupted flow from a fresh session: it takes a `gauntlet-handoff` brief (a file path, pasted text, or - with no arguments - a pick from the briefs under `<tmpdir>/pi-handoff/`), a bare worktree that already holds a spec, or a spec tracked on main under a configured spec dir (a spec seed: it creates or reuses the worktree through `using-git-worktrees` and pins the spec), restores phase/plan tracker state through the legal arming sequence (`start brainstorm`, `skip` with `resume:` reasons, `plan_check` before implement-or-later), and never infers approval from artifacts - run it with `/skill:gauntlet-resume [<brief-file> | <spec>.md | <worktree-name-or-path>]`.
73
73
  - **7 subagent personas** - the specialized child agents the skills dispatch via pi-cohort: `implementer`, `code-reviewer`, `spec-reviewer`, `conformance-reviewer`, `spec-summarizer`, `spec-council-member`, `spec-council-synthesizer`. See [doc/personas.md](./doc/personas.md) for what each one does and why its permissions are scoped the way they are.
74
74
  - **4 runtime extensions** - the enforcement layer. `plan-tracker` and `phase-tracker` are tools skills call to track progress (with a TUI widget); `verify-before-ship` is a hook that warns if you push or open a PR without a passing test run since your last edit; a phase-tracker flow guard reminds on implement-phase commits missing spec/code review. In a brainstorming-entered flow, phase-tracker rejects `implement` or `verify` completion while tracker tasks remain pending or in progress; see [its configuration reference](./doc/configuration.md#phase-tracker). phase-tracker also registers `plan_check`, which verifies a plan against its spec and against the grammar in [skills/writing-plans/reference/plan-contract.md](./skills/writing-plans/reference/plan-contract.md), including that each task's `Tests:` commands are selective and never the full suite; a pass stamps the plan for implementation. `telemetry` records one YAML record per gauntlet run (phase timing, models, personas, gate/fix rounds, diff at ship) that ships in the squash beside the spec - it is written only inside a brainstorming-entered flow, kept untracked and git-excluded until `finishing-a-development-branch` seals and commits it once with `gauntlet-telemetry-seal` - and blocks a brainstorm `write` into an already-shipped spec; see [its configuration reference](./doc/configuration.md#telemetry). See [doc/configuration.md](./doc/configuration.md) for the settings each one reads.
75
75
 
@@ -141,7 +141,7 @@ Pin an exact release with `npm:pi-gauntlet@X.Y.Z`. See [doc/install-internals.md
141
141
 
142
142
  ## Performance digest
143
143
 
144
- `gauntlet-performance` digests telemetry records without judging them: `node <pi-gauntlet-package>/bin/gauntlet-performance.mjs [--dir <repo root or telemetry dir>]... [--since <version>] [--json]` prints a `corpus:` line, one row per run (`run_id`, repo, spec, version, status - `shipped*` when the record has no ship phase - wall, per-phase minutes, tokens, cost, models, dispatches, `fix_round_grants`, reopens, conformance loops, reviewer findings, council dispatches), a `by version` block of p50/max over shipped, non-truncated runs, and `skipped:` lines for files it could not use. Default corpus is the current repo's telemetry dir; each `--dir` adds a repo root (its own `telemetry.dir` setting is honoured) or a bare telemetry dir. `--since` compares versions numerically. Exit 0 always except usage errors. `/skill:gauntlet-performance` is the reasoning layer on top.
144
+ `gauntlet-performance` digests telemetry records without judging them: `node <pi-gauntlet-package>/bin/gauntlet-performance.mjs [--dir <repo root or telemetry dir>]... [--since <version>] [--json]` prints a `corpus:` line, one row per run (`run_id`, repo, spec, version, status - `shipped*` when the record has no ship phase - wall, per-phase minutes, tokens, cost, models, dispatches, `fix_round_grants`, reopens, conformance loops, reviewer findings, council dispatches), a `by version` block of p50/max over shipped, non-truncated runs, a `council` block grouped by member roster, and `skipped:` lines for files it could not use. Each council table shows member `total`, `unique`, `applied`, `uniq_appl`, `deferred`, `rejected` blocker/major/minor counts and `uniq_appl_nonminor/run`; `chair:` shows runs, average clusters and members reported, and retried runs. Roster flags (JSON `flags[].rule`: `low_unique_applied` - fewer than 0.2 unique-and-applied blocker/major findings per run; `high_rejection` - over 50% rejected while another member is below 25%) require five runs and print as `flag: <member> - <detail>`. Smaller groups print `not enough runs to assess (N of 5)`; absent council data prints `no council data in selected records`. JSON includes `council: [{ roster, runs, members, chairs, flags }]`. Default corpus is the current repo's telemetry dir; each `--dir` adds a repo root (its own `telemetry.dir` setting is honoured) or a bare telemetry dir. `--since` compares versions numerically. Exit 0 always except usage errors. `/skill:gauntlet-performance` is the reasoning layer on top.
145
145
 
146
146
  For local development against a checkout instead of npm:
147
147
 
@@ -1,4 +1,9 @@
1
1
  #!/usr/bin/env node
2
+ var __defProp = Object.defineProperty;
3
+ var __export = (target, all) => {
4
+ for (var name in all)
5
+ __defProp(target, name, { get: all[name], enumerable: true });
6
+ };
2
7
 
3
8
  // src/bins/gauntlet-performance.mjs
4
9
  import { existsSync, readdirSync, readFileSync, realpathSync, statSync } from "node:fs";
@@ -43,20 +48,227 @@ function resolveTelemetry(g) {
43
48
  return { enabled, dir, buckets, warning: joinWarn(warnings) };
44
49
  }
45
50
 
46
- // src/bins/gauntlet-performance.mjs
51
+ // src/bins/performance/shared.mjs
47
52
  var PHASES = ["brainstorm", "plan", "implement", "verify", "ship"];
48
53
  var TOKEN_KEYS = ["input", "output", "cache_read", "cache_write"];
49
- var RUN_HEADER = ["run_id", "repo", "spec", "version", "status", "wall", "b/p/i/v/s min", "tokens", "cost", "models", "disp", "grants", "reopens", "loops", "findings", "council"];
50
- var VERSION_HEADER = ["version", "n", "shipped", "truncated", "wall p50/max", "tokens p50/max", "cost p50/max", "disp p50", "grants p50/max", "reopens p50/max", "loops p50/max", "findings p50 b/M/m", "models"];
51
- var usage = () => {
52
- process.stderr.write("usage: gauntlet-performance [--dir <repo root or telemetry dir>]... [--since <version>] [--json]\n");
53
- process.exit(1);
54
- };
55
54
  var semver = (s) => {
56
55
  const m = /^(0|[1-9]\d*)\.(0|[1-9]\d*)\.(0|[1-9]\d*)$/.exec(typeof s === "string" ? s.trim() : "");
57
56
  return m ? [Number(m[1]), Number(m[2]), Number(m[3])] : void 0;
58
57
  };
59
58
  var cmpSemver = (a, b) => a[0] - b[0] || a[1] - b[1] || a[2] - b[2];
59
+ var num = (v) => typeof v === "number" && Number.isFinite(v) ? v : null;
60
+ var str = (v) => typeof v === "string" ? v : null;
61
+ var obj = (v) => v && typeof v === "object" && !Array.isArray(v) ? v : null;
62
+ var sumOrNull = (vals) => {
63
+ const xs = vals.filter((v) => v !== null);
64
+ return xs.length ? xs.reduce((a, b) => a + b, 0) : null;
65
+ };
66
+ var p50 = (xs) => {
67
+ const s = [...xs].sort((a, b) => a - b);
68
+ const m = s.length >> 1;
69
+ return s.length % 2 ? s[m] : (s[m - 1] + s[m]) / 2;
70
+ };
71
+ var stat = (rows, pick) => {
72
+ const xs = rows.map(pick).filter((v) => v !== null && v !== void 0);
73
+ return xs.length ? { p50: p50(xs), max: Math.max(...xs) } : { p50: null, max: null };
74
+ };
75
+ var dash = (v) => v === null || v === void 0 ? "-" : String(v);
76
+ var fmtMin = (s) => s === null ? "-" : `${Math.round(s / 60)}m`;
77
+ var fmtCount = (n) => n === null ? "-" : n >= 1e6 ? `${(n / 1e6).toFixed(1)}M` : n >= 1e3 ? `${(n / 1e3).toFixed(1)}k` : String(n);
78
+ var fmtCost = (c) => c === null ? "-" : c.toFixed(2);
79
+ var pair = (s, f) => `${f(s.p50)}/${f(s.max)}`;
80
+ var findingsCell = (f) => f === null ? "-" : `${dash(f.blocker)}/${dash(f.major)}/${dash(f.minor)}`;
81
+ var table = (header, rows) => {
82
+ const w = header.map((h, i) => Math.max(h.length, ...rows.map((r) => r[i].length)));
83
+ return [header, ...rows].map((r) => r.map((c, i) => c.padEnd(w[i])).join(" ").trimEnd()).join("\n");
84
+ };
85
+
86
+ // src/bins/performance/runs.mjs
87
+ var runs_exports = {};
88
+ __export(runs_exports, {
89
+ aggregate: () => aggregate,
90
+ key: () => key,
91
+ renderText: () => renderText
92
+ });
93
+ var key = "runs";
94
+ var HEADER = ["run_id", "repo", "spec", "version", "status", "wall", "b/p/i/v/s min", "tokens", "cost", "models", "disp", "grants", "reopens", "loops", "findings", "council"];
95
+ var runRow = (r) => [
96
+ r.run_id.slice(0, 8),
97
+ r.repo,
98
+ r.spec,
99
+ r.version,
100
+ `${dash(r.status)}${r.truncated ? "*" : ""}`,
101
+ fmtMin(r.wall_s),
102
+ PHASES.map((p) => r.phase_min[p] === null ? "-" : String(Math.round(r.phase_min[p]))).join("/"),
103
+ fmtCount(r.tokens),
104
+ fmtCost(r.cost),
105
+ r.models.join(",") || "-",
106
+ dash(r.dispatches),
107
+ dash(r.grants),
108
+ dash(r.reopens),
109
+ dash(r.loops),
110
+ findingsCell(r.findings),
111
+ dash(r.council)
112
+ ];
113
+ var aggregate = (entries) => entries.map((e) => e.row);
114
+ var renderText = (rows) => [`runs (${rows.length})`, table(HEADER, rows.map(runRow))].join("\n");
115
+
116
+ // src/bins/performance/versions.mjs
117
+ var versions_exports = {};
118
+ __export(versions_exports, {
119
+ aggregate: () => aggregate2,
120
+ key: () => key2,
121
+ renderText: () => renderText2
122
+ });
123
+ var key2 = "by_version";
124
+ var HEADER2 = ["version", "n", "shipped", "truncated", "wall p50/max", "tokens p50/max", "cost p50/max", "disp p50", "grants p50/max", "reopens p50/max", "loops p50/max", "findings p50 b/M/m", "models"];
125
+ var versionOrder = (a, b) => {
126
+ const x = semver(a), y = semver(b);
127
+ return x && y ? cmpSemver(x, y) : x ? -1 : y ? 1 : 0;
128
+ };
129
+ function aggregate2(entries) {
130
+ const rows = entries.map((e) => e.row);
131
+ const groups = /* @__PURE__ */ new Map();
132
+ for (const r of rows) {
133
+ if (!groups.has(r.version)) groups.set(r.version, []);
134
+ groups.get(r.version).push(r);
135
+ }
136
+ return [...groups.entries()].sort(([a], [b]) => versionOrder(a, b)).map(([version, all]) => {
137
+ const shipped = all.filter((r) => r.status === "shipped" && !r.truncated);
138
+ const models = {};
139
+ for (const r of shipped) for (const m of r.models) models[m] = (models[m] ?? 0) + 1;
140
+ return {
141
+ version,
142
+ n: all.length,
143
+ shipped: shipped.length,
144
+ truncated: all.filter((r) => r.truncated).length,
145
+ wall_s: stat(shipped, (r) => r.wall_s),
146
+ tokens: stat(shipped, (r) => r.tokens),
147
+ cost: stat(shipped, (r) => r.cost),
148
+ dispatches: { p50: stat(shipped, (r) => r.dispatches).p50 },
149
+ grants: stat(shipped, (r) => r.grants),
150
+ reopens: stat(shipped, (r) => r.reopens),
151
+ loops: stat(shipped, (r) => r.loops),
152
+ findings: {
153
+ blocker: { p50: stat(shipped, (r) => r.findings?.blocker ?? null).p50 },
154
+ major: { p50: stat(shipped, (r) => r.findings?.major ?? null).p50 },
155
+ minor: { p50: stat(shipped, (r) => r.findings?.minor ?? null).p50 }
156
+ },
157
+ models
158
+ };
159
+ });
160
+ }
161
+ var versionRow = (g) => [
162
+ g.version,
163
+ String(g.n),
164
+ String(g.shipped),
165
+ String(g.truncated),
166
+ pair(g.wall_s, fmtMin),
167
+ pair(g.tokens, fmtCount),
168
+ pair(g.cost, fmtCost),
169
+ dash(g.dispatches.p50),
170
+ pair(g.grants, dash),
171
+ pair(g.reopens, dash),
172
+ pair(g.loops, dash),
173
+ `${dash(g.findings.blocker.p50)}/${dash(g.findings.major.p50)}/${dash(g.findings.minor.p50)}`,
174
+ Object.entries(g.models).map(([m, n]) => `${m}:${n}`).join(",") || "-"
175
+ ];
176
+ var renderText2 = (agg) => ["by version", table(HEADER2, agg.map(versionRow))].join("\n");
177
+
178
+ // src/bins/performance/council.mjs
179
+ var council_exports = {};
180
+ __export(council_exports, {
181
+ aggregate: () => aggregate3,
182
+ key: () => key3,
183
+ renderText: () => renderText3
184
+ });
185
+ var key3 = "council";
186
+ var SEVERITIES = ["blocker", "major", "minor"];
187
+ var COUNTERS = ["total", "unique", "applied", "unique_applied", "deferred", "rejected"];
188
+ var MIN_RUNS = 5;
189
+ var HEADER3 = ["member", "total", "unique", "applied", "uniq_appl", "deferred", "rejected", "uniq_appl_nonminor/run"];
190
+ var zero = () => ({ blocker: 0, major: 0, minor: 0 });
191
+ var addInto = (acc, src) => {
192
+ for (const s of SEVERITIES) acc[s] += num(obj(src)?.[s]) ?? 0;
193
+ };
194
+ var sum = (c) => c.blocker + c.major + c.minor;
195
+ var cell = (c) => `${c.blocker}/${c.major}/${c.minor}`;
196
+ var pct = (x) => `${Math.round(x * 100)}%`;
197
+ function flagsFor(members, runs) {
198
+ const share = (m) => sum(m.total) === 0 ? null : sum(m.rejected) / sum(m.total);
199
+ const quiet = Object.entries(members).filter(([, m]) => share(m) !== null && share(m) < 0.25);
200
+ const out = [];
201
+ for (const [name, m] of Object.entries(members)) {
202
+ if (m.unique_applied_nonminor_per_run < 0.2) {
203
+ out.push({ member: name, rule: "low_unique_applied", detail: `${m.unique_applied_nonminor_per_run.toFixed(2)} unique-and-applied blocker/major per run in ${runs} runs (total ${cell(m.total)})` });
204
+ }
205
+ const s = share(m);
206
+ const other = quiet[0];
207
+ if (s !== null && s > 0.5 && other) {
208
+ out.push({ member: name, rule: "high_rejection", detail: `rejected ${pct(s)} vs ${other[0]} ${pct(share(other[1]))} (total ${cell(m.total)})` });
209
+ }
210
+ }
211
+ return out;
212
+ }
213
+ function aggregate3(entries) {
214
+ const groups = /* @__PURE__ */ new Map();
215
+ for (const [i, e] of entries.entries()) {
216
+ const members = obj(obj(e.council)?.members);
217
+ if (!members) continue;
218
+ const roster = Object.keys(members).sort();
219
+ const k = roster.join(", ");
220
+ if (!groups.has(k)) groups.set(k, { roster, entries: [] });
221
+ groups.get(k).entries.push(e);
222
+ groups.get(k).last = i;
223
+ }
224
+ return [...groups.values()].sort((a, b) => b.last - a.last).map((g) => {
225
+ const runs = g.entries.length;
226
+ const members = {};
227
+ for (const m of g.roster) members[m] = { dispatches: 0, ...Object.fromEntries(COUNTERS.map((c) => [c, zero()])) };
228
+ const chairs = /* @__PURE__ */ new Map();
229
+ for (const e of g.entries) {
230
+ for (const m of g.roster) {
231
+ const src = obj(e.council.members[m]);
232
+ members[m].dispatches += num(src?.dispatches) ?? 0;
233
+ for (const c2 of COUNTERS) addInto(members[m][c2], src?.[c2]);
234
+ }
235
+ const ch = obj(e.council.chair);
236
+ const model = str(ch?.model) ?? "unknown";
237
+ const c = chairs.get(model) ?? { model, runs: 0, clusters: 0, reported: 0, retried_runs: 0 };
238
+ c.runs += 1;
239
+ c.clusters += num(ch?.clusters) ?? 0;
240
+ c.reported += num(ch?.members_reported) ?? 0;
241
+ if ((num(ch?.dispatches) ?? 1) > 1) c.retried_runs += 1;
242
+ chairs.set(model, c);
243
+ }
244
+ for (const m of g.roster) members[m].unique_applied_nonminor_per_run = (members[m].unique_applied.blocker + members[m].unique_applied.major) / runs;
245
+ const chairList = [...chairs.values()].sort((a, b) => b.runs - a.runs || a.model.localeCompare(b.model)).map((c) => ({ model: c.model, runs: c.runs, avg_clusters: c.clusters / c.runs, avg_members_reported: c.reported / c.runs, retried_runs: c.retried_runs }));
246
+ return { roster: g.roster, runs, members, chairs: chairList, flags: runs >= MIN_RUNS ? flagsFor(members, runs) : [] };
247
+ });
248
+ }
249
+ function renderText3(agg) {
250
+ if (!agg.length) return "council\nno council data in selected records";
251
+ const out = ["council", "severity is the chair's consolidated severity (highest any co-raiser assigned), not the member's own"];
252
+ for (const g of agg) {
253
+ out.push(`roster (${g.runs} runs): ${g.roster.join(", ")}`);
254
+ const rows = g.roster.map((m) => {
255
+ const x = g.members[m];
256
+ return [m, cell(x.total), cell(x.unique), cell(x.applied), cell(x.unique_applied), cell(x.deferred), cell(x.rejected), x.unique_applied_nonminor_per_run.toFixed(2)];
257
+ });
258
+ for (const line of table(HEADER3, rows).split("\n")) out.push(` ${line}`);
259
+ for (const c of g.chairs) out.push(` chair: ${c.model} - ${c.runs} runs, avg ${c.avg_clusters.toFixed(1)} clusters from ${c.avg_members_reported.toFixed(1)} members reported, retried in ${c.retried_runs}`);
260
+ if (g.runs < MIN_RUNS) out.push(` not enough runs to assess (${g.runs} of ${MIN_RUNS})`);
261
+ for (const f of g.flags) out.push(` flag: ${f.member} - ${f.detail}`);
262
+ }
263
+ return out.join("\n");
264
+ }
265
+
266
+ // src/bins/gauntlet-performance.mjs
267
+ var SECTIONS = [runs_exports, versions_exports, council_exports];
268
+ var usage = () => {
269
+ process.stderr.write("usage: gauntlet-performance [--dir <repo root or telemetry dir>]... [--since <version>] [--json]\n");
270
+ process.exit(1);
271
+ };
60
272
  function parseArgs(argv) {
61
273
  const opts = { dirs: [], since: void 0, json: false };
62
274
  for (let i = 0; i < argv.length; i++) {
@@ -112,13 +324,6 @@ function* yamlFiles(dir) {
112
324
  else if (e.isFile() && e.name.endsWith(".yaml")) yield p;
113
325
  }
114
326
  }
115
- var num = (v) => typeof v === "number" && Number.isFinite(v) ? v : null;
116
- var str = (v) => typeof v === "string" ? v : null;
117
- var obj = (v) => v && typeof v === "object" && !Array.isArray(v) ? v : null;
118
- var sumOrNull = (vals) => {
119
- const xs = vals.filter((v) => v !== null);
120
- return xs.length ? xs.reduce((a, b) => a + b, 0) : null;
121
- };
122
327
  function loadRecord(file, label) {
123
328
  let doc;
124
329
  try {
@@ -179,96 +384,10 @@ function loadRecord(file, label) {
179
384
  council: num(obj(personas["spec-council-member"])?.dispatches),
180
385
  ship_option: str(gates.ship_option),
181
386
  tests: str(obj(derived.tests)?.result)
182
- }
387
+ },
388
+ council: derived.council ?? null
183
389
  };
184
390
  }
185
- var p50 = (xs) => {
186
- const s = [...xs].sort((a, b) => a - b);
187
- const m = s.length >> 1;
188
- return s.length % 2 ? s[m] : (s[m - 1] + s[m]) / 2;
189
- };
190
- var stat = (rows, pick) => {
191
- const xs = rows.map(pick).filter((v) => v !== null && v !== void 0);
192
- return xs.length ? { p50: p50(xs), max: Math.max(...xs) } : { p50: null, max: null };
193
- };
194
- var versionOrder = (a, b) => {
195
- const x = semver(a), y = semver(b);
196
- return x && y ? cmpSemver(x, y) : x ? -1 : y ? 1 : 0;
197
- };
198
- function aggregate(rows) {
199
- const groups = /* @__PURE__ */ new Map();
200
- for (const r of rows) {
201
- if (!groups.has(r.version)) groups.set(r.version, []);
202
- groups.get(r.version).push(r);
203
- }
204
- return [...groups.entries()].sort(([a], [b]) => versionOrder(a, b)).map(([version, all]) => {
205
- const shipped = all.filter((r) => r.status === "shipped" && !r.truncated);
206
- const models = {};
207
- for (const r of shipped) for (const m of r.models) models[m] = (models[m] ?? 0) + 1;
208
- return {
209
- version,
210
- n: all.length,
211
- shipped: shipped.length,
212
- truncated: all.filter((r) => r.truncated).length,
213
- wall_s: stat(shipped, (r) => r.wall_s),
214
- tokens: stat(shipped, (r) => r.tokens),
215
- cost: stat(shipped, (r) => r.cost),
216
- dispatches: { p50: stat(shipped, (r) => r.dispatches).p50 },
217
- grants: stat(shipped, (r) => r.grants),
218
- reopens: stat(shipped, (r) => r.reopens),
219
- loops: stat(shipped, (r) => r.loops),
220
- findings: {
221
- blocker: { p50: stat(shipped, (r) => r.findings?.blocker ?? null).p50 },
222
- major: { p50: stat(shipped, (r) => r.findings?.major ?? null).p50 },
223
- minor: { p50: stat(shipped, (r) => r.findings?.minor ?? null).p50 }
224
- },
225
- models
226
- };
227
- });
228
- }
229
- var dash = (v) => v === null || v === void 0 ? "-" : String(v);
230
- var fmtMin = (s) => s === null ? "-" : `${Math.round(s / 60)}m`;
231
- var fmtCount = (n) => n === null ? "-" : n >= 1e6 ? `${(n / 1e6).toFixed(1)}M` : n >= 1e3 ? `${(n / 1e3).toFixed(1)}k` : String(n);
232
- var fmtCost = (c) => c === null ? "-" : c.toFixed(2);
233
- var pair = (s, f) => `${f(s.p50)}/${f(s.max)}`;
234
- var findingsCell = (f) => f === null ? "-" : `${dash(f.blocker)}/${dash(f.major)}/${dash(f.minor)}`;
235
- var table = (header, rows) => {
236
- const w = header.map((h, i) => Math.max(h.length, ...rows.map((r) => r[i].length)));
237
- return [header, ...rows].map((r) => r.map((c, i) => c.padEnd(w[i])).join(" ").trimEnd()).join("\n");
238
- };
239
- var runRow = (r) => [
240
- r.run_id.slice(0, 8),
241
- r.repo,
242
- r.spec,
243
- r.version,
244
- `${dash(r.status)}${r.truncated ? "*" : ""}`,
245
- fmtMin(r.wall_s),
246
- PHASES.map((p) => r.phase_min[p] === null ? "-" : String(Math.round(r.phase_min[p]))).join("/"),
247
- fmtCount(r.tokens),
248
- fmtCost(r.cost),
249
- r.models.join(",") || "-",
250
- dash(r.dispatches),
251
- dash(r.grants),
252
- dash(r.reopens),
253
- dash(r.loops),
254
- findingsCell(r.findings),
255
- dash(r.council)
256
- ];
257
- var versionRow = (g) => [
258
- g.version,
259
- String(g.n),
260
- String(g.shipped),
261
- String(g.truncated),
262
- pair(g.wall_s, fmtMin),
263
- pair(g.tokens, fmtCount),
264
- pair(g.cost, fmtCost),
265
- dash(g.dispatches.p50),
266
- pair(g.grants, dash),
267
- pair(g.reopens, dash),
268
- pair(g.loops, dash),
269
- `${dash(g.findings.blocker.p50)}/${dash(g.findings.major.p50)}/${dash(g.findings.minor.p50)}`,
270
- Object.entries(g.models).map(([m, n]) => `${m}:${n}`).join(",") || "-"
271
- ];
272
391
  function main() {
273
392
  const opts = parseArgs(process.argv.slice(2));
274
393
  const corpora = [];
@@ -285,7 +404,7 @@ function main() {
285
404
  const since = opts.since === void 0 ? void 0 : semver(opts.since);
286
405
  const seen = /* @__PURE__ */ new Set();
287
406
  const corpus = {};
288
- const rows = [];
407
+ const entries = [];
289
408
  for (const c of corpora) {
290
409
  corpus[c.label] = corpus[c.label] ?? 0;
291
410
  for (const file of yamlFiles(c.dir)) {
@@ -302,26 +421,25 @@ function main() {
302
421
  const v = semver(res.row.version);
303
422
  if (since && (!v || cmpSemver(v, since) < 0)) continue;
304
423
  corpus[c.label]++;
305
- rows.push(res.row);
424
+ entries.push(res);
306
425
  }
307
426
  }
308
- rows.sort((a, b) => (a.created_at ?? "").localeCompare(b.created_at ?? "") || a.run_id.localeCompare(b.run_id));
309
- const byVersion = rows.length ? aggregate(rows) : [];
427
+ entries.sort((a, b) => (a.row.created_at ?? "").localeCompare(b.row.created_at ?? "") || a.row.run_id.localeCompare(b.row.run_id));
428
+ const aggregates = SECTIONS.map((s) => [s, s.aggregate(entries)]);
310
429
  if (opts.json) {
311
- console.log(JSON.stringify({ corpus, since: opts.since ?? null, runs: rows, by_version: byVersion, skipped }, null, 2));
430
+ const out = { corpus, since: opts.since ?? null };
431
+ for (const [s, agg] of aggregates) out[s.key] = agg;
432
+ out.skipped = skipped;
433
+ console.log(JSON.stringify(out, null, 2));
312
434
  return;
313
435
  }
314
436
  console.log(`corpus: ${Object.entries(corpus).map(([k, v]) => `${k}=${v}`).join(", ")}${opts.since === void 0 ? "" : ` since: ${opts.since}`}`);
315
- if (rows.length === 0) {
437
+ if (entries.length === 0) {
316
438
  console.log(`no records found in ${corpora.map((c) => c.dir).join(", ") || process.cwd()}`);
317
439
  } else {
318
- console.log(`runs (${rows.length})`);
319
- console.log(table(RUN_HEADER, rows.map(runRow)));
320
- console.log("");
321
- console.log("by version");
322
- console.log(table(VERSION_HEADER, byVersion.map(versionRow)));
440
+ console.log(aggregates.map(([s, agg]) => s.renderText(agg)).join("\n\n"));
323
441
  }
324
- if (rows.length && skipped.length) console.log("");
325
- if (rows.length) for (const s of skipped) console.log(`skipped: ${s.file}: ${s.reason}`);
442
+ if (entries.length && skipped.length) console.log("");
443
+ if (entries.length) for (const s of skipped) console.log(`skipped: ${s.file}: ${s.reason}`);
326
444
  }
327
445
  main();
@@ -348,3 +348,103 @@ test("usage error exits 1", (t) => {
348
348
  assert.equal(invoke(root, ["--since", "abc"]).status, 1);
349
349
  assert.equal(invoke(root, ["--since", "5.09.1"]).status, 1);
350
350
  });
351
+
352
+ const sev = (b, M, m) => `{ blocker: ${b}, major: ${M}, minor: ${m} }`;
353
+ const member = (dispatches, total, unique, applied, uniqApplied, deferred, rejected) =>
354
+ `{ dispatches: ${dispatches}, total: ${total}, unique: ${unique}, applied: ${applied}, unique_applied: ${uniqApplied}, deferred: ${deferred}, rejected: ${rejected} }`;
355
+ const ALPHA = member(1, sev(1, 3, 1), sev(0, 1, 1), sev(1, 3, 0), sev(0, 1, 0), sev(0, 0, 0), sev(0, 0, 1));
356
+ const BETA = member(1, sev(0, 2, 2), sev(0, 1, 0), sev(0, 1, 1), sev(0, 1, 0), sev(0, 1, 0), sev(0, 0, 1));
357
+ const GAMMA = member(2, sev(0, 1, 4), sev(0, 0, 3), sev(0, 0, 1), sev(0, 0, 1), sev(0, 0, 0), sev(0, 1, 3));
358
+ const councilRecord = (id, createdAt, version, chair, chairDispatches, members) => `schema: 1
359
+ spec: doc/specs/c-${id}.md
360
+ run_id: ${id}-0000-0000-0000-000000000000
361
+ status: shipped
362
+ created_at: ${createdAt}
363
+ shipped_at: ${createdAt}
364
+ versions: { pi-gauntlet: ${version} }
365
+ derived:
366
+ duration_s: 60
367
+ phases:
368
+ ship: ${phase("model: m/alpha")}
369
+ personas: {}
370
+ conformance_loops: 0
371
+ gates: { spec_rounds: 1, plan_rounds: 1, fix_round_grants: 0, task_reopens: 0 }
372
+ council:
373
+ chair: { model: ${chair}, dispatches: ${chairDispatches}, clusters: 6, members_reported: 3 }
374
+ members:
375
+ ${Object.entries(members).map(([k, v]) => ` ${k}: ${v}`).join("\n")}
376
+ amendments: 0
377
+ spec_edits_after_ship: 0
378
+ spec_writes: {}
379
+ events_dropped: 0
380
+ accumulators: {}
381
+ events: []
382
+ `;
383
+ const ZERO = member(1, ...Array(6).fill(sev(0, 0, 0)));
384
+ const BIG = { "p/alpha:xhigh": ALPHA, "p/beta:high": BETA, "p/gamma:high": GAMMA, "p/zero:high": ZERO };
385
+ const SMALL = { "p/alpha:xhigh": ALPHA, "p/delta:high": BETA };
386
+ const councilCorpus = (root) => {
387
+ for (let i = 0; i < 6; i++) write(root, `${TDIR}/big-${i}.yaml`, councilRecord(`c000000${i}`, `2026-09-1${i}T00:00:00Z`, "5.20.0", i < 4 ? "p/chair:medium" : "p/other:high", i === 2 ? 2 : 1, BIG));
388
+ for (let i = 0; i < 2; i++) write(root, `${TDIR}/small-${i}.yaml`, councilRecord(`d000000${i}`, `2026-09-2${i}T00:00:00Z`, "5.21.0", "p/chair:medium", 1, SMALL));
389
+ write(root, `${TDIR}/plain.yaml`, COMPLETE);
390
+ return root;
391
+ };
392
+
393
+ test("council section: rosters, sums, chairs and flags", (t) => {
394
+ const root = councilCorpus(gitRepo()); cleanup(t, root);
395
+ const d = json(root);
396
+ assert.equal(d.council.length, 2);
397
+ assert.deepEqual(d.council[0].roster, ["p/alpha:xhigh", "p/delta:high"]);
398
+ assert.equal(d.council[0].runs, 2);
399
+ assert.deepEqual(d.council[0].flags, []);
400
+ const big = d.council[1];
401
+ assert.deepEqual(big.roster, ["p/alpha:xhigh", "p/beta:high", "p/gamma:high", "p/zero:high"]);
402
+ assert.equal(big.runs, 6);
403
+ assert.deepEqual(big.members["p/alpha:xhigh"].total, { blocker: 6, major: 18, minor: 6 });
404
+ assert.equal(big.members["p/alpha:xhigh"].dispatches, 6);
405
+ assert.equal(big.members["p/gamma:high"].dispatches, 12);
406
+ assert.equal(big.members["p/alpha:xhigh"].unique_applied_nonminor_per_run, 1);
407
+ assert.equal(big.members["p/gamma:high"].unique_applied_nonminor_per_run, 0);
408
+ assert.deepEqual(big.chairs, [
409
+ { model: "p/chair:medium", runs: 4, avg_clusters: 6, avg_members_reported: 3, retried_runs: 1 },
410
+ { model: "p/other:high", runs: 2, avg_clusters: 6, avg_members_reported: 3, retried_runs: 0 },
411
+ ]);
412
+ assert.deepEqual(big.flags.map((f) => [f.member, f.rule]), [
413
+ ["p/gamma:high", "low_unique_applied"], ["p/gamma:high", "high_rejection"],
414
+ ["p/zero:high", "low_unique_applied"],
415
+ ]);
416
+ assert.equal(big.flags[0].detail, "0.00 unique-and-applied blocker/major per run in 6 runs (total 0/6/24)");
417
+ assert.equal(big.flags[1].detail, "rejected 80% vs p/alpha:xhigh 20% (total 0/6/24)");
418
+ assert.equal(big.flags[2].detail, "0.00 unique-and-applied blocker/major per run in 6 runs (total 0/0/0)");
419
+ const r = invoke(root);
420
+ assert.equal(r.status, 0, r.stderr);
421
+ const at = r.lines.indexOf("council");
422
+ assert.ok(at > 0);
423
+ assert.equal(r.lines[at + 1], "severity is the chair's consolidated severity (highest any co-raiser assigned), not the member's own");
424
+ assert.equal(r.lines[at + 2], "roster (2 runs): p/alpha:xhigh, p/delta:high");
425
+ assert.match(r.lines[at + 3], /^ member\s+total\s+unique\s+applied\s+uniq_appl\s+deferred\s+rejected\s+uniq_appl_nonminor\/run$/);
426
+ assert.ok(r.lines.includes(" not enough runs to assess (2 of 5)"));
427
+ assert.ok(r.lines.includes("roster (6 runs): p/alpha:xhigh, p/beta:high, p/gamma:high, p/zero:high"));
428
+ assert.match(r.lines.find((l) => /^ p\/alpha:xhigh\s+6\/18\/6\s/.test(l)), /^ p\/alpha:xhigh\s+6\/18\/6\s+0\/6\/6\s+6\/18\/0\s+0\/6\/0\s+0\/0\/0\s+0\/0\/6\s+1\.00$/);
429
+ assert.ok(r.lines.includes(" chair: p/chair:medium - 4 runs, avg 6.0 clusters from 3.0 members reported, retried in 1"));
430
+ assert.ok(r.lines.includes(" chair: p/other:high - 2 runs, avg 6.0 clusters from 3.0 members reported, retried in 0"));
431
+ assert.ok(r.lines.includes(" flag: p/gamma:high - 0.00 unique-and-applied blocker/major per run in 6 runs (total 0/6/24)"));
432
+ assert.ok(r.lines.includes(" flag: p/gamma:high - rejected 80% vs p/alpha:xhigh 20% (total 0/6/24)"));
433
+ });
434
+
435
+ test("council section: --since and no council data", (t) => {
436
+ const root = councilCorpus(gitRepo()); cleanup(t, root);
437
+ const d = json(root, ["--since", "5.21.0"]);
438
+ assert.equal(d.council.length, 1);
439
+ assert.deepEqual(d.council[0].roster, ["p/alpha:xhigh", "p/delta:high"]);
440
+ const plain = corpus(gitRepo()); cleanup(t, plain);
441
+ assert.deepEqual(json(plain).council, []);
442
+ const r = invoke(plain);
443
+ const at = r.lines.indexOf("council");
444
+ assert.ok(at > 0);
445
+ assert.equal(r.lines[at + 1], "no council data in selected records");
446
+ const empty = gitRepo(); cleanup(t, empty);
447
+ const none = invoke(empty);
448
+ assert.match(none.lines[1], /^no records found/);
449
+ assert.equal(none.lines.length, 2);
450
+ });
@@ -168,7 +168,7 @@ function derive(rec, now) {
168
168
  }
169
169
  for (const k of Object.keys(acc.phases)) phases[k] = compact({ ...phases[k] ?? {}, ...acc.phases[k] });
170
170
  const live = liveShipEvent(rec.events);
171
- return compact({ duration_s: seconds(rec.created_at, rec.shipped_at ?? rec.abandoned_at ?? now), phases, personas: acc.personas, reviews: Object.keys(acc.reviews).length ? acc.reviews : void 0, conformance_loops: acc.conformance_loops, conformance_open_gaps: acc.conformance_open_gaps, gates: compact({ ...acc.gates, ship_option: live?.option }), plan: rec.derived.plan, tests: rec.derived.tests, amendments: acc.amendments, spec_edits_after_ship: acc.spec_edits_after_ship, diff: rec.derived.diff, modified_files: rec.derived.modified_files, spec_writes: acc.spec_writes, events_dropped: acc.events_dropped });
171
+ return compact({ duration_s: seconds(rec.created_at, rec.shipped_at ?? rec.abandoned_at ?? now), phases, personas: acc.personas, reviews: Object.keys(acc.reviews).length ? acc.reviews : void 0, conformance_loops: acc.conformance_loops, conformance_open_gaps: acc.conformance_open_gaps, gates: compact({ ...acc.gates, ship_option: live?.option }), plan: rec.derived.plan, tests: rec.derived.tests, amendments: acc.amendments, spec_edits_after_ship: acc.spec_edits_after_ship, diff: rec.derived.diff, modified_files: rec.derived.modified_files, council: rec.derived.council, spec_writes: acc.spec_writes, events_dropped: acc.events_dropped });
172
172
  }
173
173
  var stripUndefined = (v) => {
174
174
  if (Array.isArray(v)) return v.map(stripUndefined);
@@ -202,7 +202,11 @@ function serializeRecord(rec) {
202
202
  const body = stripUndefined({ ...head, derived, accumulators });
203
203
  const doc = new Document(body);
204
204
  const derivedNode = doc.get("derived", true);
205
- if (isMap(derivedNode)) flowLeafChildren(derivedNode, /* @__PURE__ */ new Set(["modified_files"]));
205
+ if (isMap(derivedNode)) {
206
+ flowLeafChildren(derivedNode, /* @__PURE__ */ new Set(["modified_files", "council"]));
207
+ const council = derivedNode.get("council", true);
208
+ if (isMap(council)) flowLeafChildren(council);
209
+ }
206
210
  const accumulatorNode = doc.get("accumulators", true);
207
211
  if (isMap(accumulatorNode)) {
208
212
  for (const session of accumulatorNode.items) if (isMap(session.value)) flowLeafChildren(session.value);
@@ -100,6 +100,16 @@ test("seals an untracked in_progress record: stamp, ship event, diff, re-derived
100
100
  assert.equal(porcelain(root), "");
101
101
  });
102
102
 
103
+ test("seal preserves derived.council through the shipped stamp", (t) => {
104
+ const council = { chair: { model: "p/chair:medium", dispatches: 1, clusters: 0, members_reported: 0 }, members: {} };
105
+ const root = repo({ record: RECORD.replace(" spec_writes: {}", ` council: ${JSON.stringify(council)}\n spec_writes: {}`) }); cleanup(t, root);
106
+ const r = run(root, ["--spec", SPEC]);
107
+ assert.equal(r.status, 0, r.stderr);
108
+ assert.equal(record(root).status, "shipped");
109
+ assert.deepEqual(record(root).derived.council, council);
110
+ assert.deepEqual(parseYaml(out(root, ["show", `HEAD:${REC}`])).derived.council, council);
111
+ });
112
+
103
113
  test("--spec accepts an absolute path within the checkout", (t) => {
104
114
  const root = repo(); cleanup(t, root);
105
115
  const r = run(root, ["--spec", join(realpathSync(root), SPEC)]);
@@ -0,0 +1,91 @@
1
+ import assert from "node:assert/strict";
2
+ import { test } from "node:test";
3
+ import { buildCouncil, memberSlugOf, normalizeRaisedBy, parseAudit, parseChairReport, resolveMembers, slug } from "./telemetry-council.ts";
4
+
5
+ const zero = { blocker: 0, major: 0, minor: 0 };
6
+ test("slug and member filename normalization", () => {
7
+ assert.equal(slug("p/alpha:xhigh"), "p-alpha-xhigh");
8
+ assert.equal(memberSlugOf("/tmp/retry/member-2-p-beta-high.md"), "p-beta-high");
9
+ assert.equal(memberSlugOf("/tmp/member-0-p-alpha"), "p-alpha");
10
+ assert.equal(memberSlugOf(undefined), undefined);
11
+ assert.equal(memberSlugOf("chair.md"), undefined);
12
+ assert.equal(normalizeRaisedBy("member-0-p-alpha"), "p-alpha");
13
+ assert.equal(normalizeRaisedBy("p-alpha.md"), "p-alpha");
14
+ assert.equal(normalizeRaisedBy(" p/alpha:xhigh "), "p-alpha-xhigh");
15
+ });
16
+ const chairText = [
17
+ "consensus: needs-work",
18
+ "clusters:",
19
+ "- [blocker] issue — raised-by: [p-alpha] — fix",
20
+ "- [Major] issue - raised-by: [p-alpha, p-beta-high, p-alpha] - fix",
21
+ "- [minor] issue -- raised-by: [p-beta-high] -- cut",
22
+ "resolved:",
23
+ "- [blocker] ignored — raised-by: [nobody]",
24
+ ].join("\n");
25
+ test("chair clusters parse only before resolved and require attributes", () => {
26
+ assert.deepEqual(parseChairReport(chairText), { clusters: [
27
+ { severity: "blocker", raisedBy: ["p-alpha"] },
28
+ { severity: "major", raisedBy: ["p-alpha", "p-beta-high"] },
29
+ { severity: "minor", raisedBy: ["p-beta-high"] },
30
+ ] });
31
+ assert.equal(parseChairReport("clusters:\n"), null);
32
+ assert.equal(parseChairReport("consensus: x\n"), null);
33
+ assert.deepEqual(parseChairReport("consensus: sound\nclusters:\nresolved:\n"), { clusters: [] });
34
+ assert.equal(parseChairReport("consensus: x\nclusters:\n- [major] no raiser"), null);
35
+ });
36
+ const auditText = [
37
+ "Applied: [blocker] issue — raised-by: [p-alpha] -> fix",
38
+ "- Applied: [major] over-spec — raised-by: [p-alpha, p-beta-high] -> cut",
39
+ "Applied: [minor] issue — raised-by: [p-beta-high] -> open question (missing)",
40
+ "Deferred: [major] x — raised-by: [p-alpha] -> later",
41
+ "Rejected: [minor] issue — raised-by: [p-beta-high] -> no",
42
+ ].join("\n");
43
+ test("audit reads dispositions, none and cut/open question", () => {
44
+ assert.deepEqual(parseAudit(auditText), [
45
+ { disposition: "applied", severity: "blocker", raisedBy: ["p-alpha"] },
46
+ { disposition: "applied", severity: "major", raisedBy: ["p-alpha", "p-beta-high"] },
47
+ { disposition: "applied", severity: "minor", raisedBy: ["p-beta-high"] },
48
+ { disposition: "deferred", severity: "major", raisedBy: ["p-alpha"] },
49
+ { disposition: "rejected", severity: "minor", raisedBy: ["p-beta-high"] },
50
+ ]);
51
+ assert.equal(parseAudit("Applied: [major] x — raised-by: [p-alpha]\nDeferred: none"), null);
52
+ assert.equal(parseAudit("Applied: x — raised-by: [p-alpha]\nDeferred: none\nRejected: none"), null);
53
+ assert.equal(parseAudit("Applied: [major] x\nDeferred: none\nRejected: none"), null);
54
+ assert.deepEqual(parseAudit("Applied: none\nDeferred: none\nRejected: none"), []);
55
+ });
56
+ const results = [
57
+ { model: "p/alpha:xhigh", exitCode: 0, savedOutputPath: "/tmp/member-0-p-alpha.md" },
58
+ { model: "p/beta:high", exitCode: 1 },
59
+ { model: "p/beta:high", exitCode: 0, savedOutputPath: "/tmp/retry/member-1-p-beta-high.md" },
60
+ ];
61
+ test("member retry counts dispatches and ambiguous slug fails", () => {
62
+ const members = resolveMembers(results)!;
63
+ assert.deepEqual([...members.dispatches], [["p/alpha:xhigh", 1], ["p/beta:high", 2]]);
64
+ assert.deepEqual([...members.slugToModel], [["p-alpha", "p/alpha:xhigh"], ["p-beta-high", "p/beta:high"]]);
65
+ assert.equal(resolveMembers([...results, { model: "p/other", exitCode: 0, savedOutputPath: "/tmp/member-3-p-alpha.md" }]), null);
66
+ assert.deepEqual([...resolveMembers([{ exitCode: 0, savedOutputPath: "/tmp/member-0-x.md" }])!.dispatches], []);
67
+ });
68
+ const clusters = [{ severity: "blocker" as const, raisedBy: ["p-alpha"] }, { severity: "major" as const, raisedBy: ["p-alpha", "p-beta-high"] }, { severity: "minor" as const, raisedBy: ["p-beta-high"] }];
69
+ const audit = [{ disposition: "applied" as const, severity: "blocker" as const, raisedBy: ["p-alpha"] }, { disposition: "deferred" as const, severity: "major" as const, raisedBy: ["p-beta-high", "p-alpha"] }, { disposition: "rejected" as const, severity: "minor" as const, raisedBy: ["p-beta-high"] }];
70
+ const input = () => ({ chair: { model: "p/chair:medium", dispatches: 2, clusters }, members: resolveMembers(results)!, audit });
71
+ test("council joins multiset and counts each severity per member", () => {
72
+ const block = buildCouncil(input())!;
73
+ assert.deepEqual(block.chair, { model: "p/chair:medium", dispatches: 2, clusters: 3, members_reported: 2 });
74
+ assert.deepEqual(block.members["p/alpha:xhigh"], { dispatches: 1, total: { blocker: 1, major: 1, minor: 0 }, unique: { blocker: 1, major: 0, minor: 0 }, applied: { blocker: 1, major: 0, minor: 0 }, unique_applied: { blocker: 1, major: 0, minor: 0 }, deferred: { blocker: 0, major: 1, minor: 0 }, rejected: zero });
75
+ assert.deepEqual(block.members["p/beta:high"], { dispatches: 2, total: { blocker: 0, major: 1, minor: 1 }, unique: { blocker: 0, major: 0, minor: 1 }, applied: zero, unique_applied: zero, deferred: { blocker: 0, major: 1, minor: 0 }, rejected: { blocker: 0, major: 0, minor: 1 } });
76
+ for (const member of Object.values(block.members)) for (const severity of ["blocker", "major", "minor"] as const) assert.equal(member.total[severity], member.applied[severity] + member.deferred[severity] + member.rejected[severity]);
77
+ });
78
+ test("zero-member and zero-finding completeness", () => {
79
+ const members = resolveMembers([...results, { model: "p/gamma:high", exitCode: 0, savedOutputPath: "/tmp/member-2-p-gamma-high.md" }])!;
80
+ const block = buildCouncil({ ...input(), members })!;
81
+ assert.deepEqual(block.members["p/gamma:high"], { dispatches: 1, total: zero, unique: zero, applied: zero, unique_applied: zero, deferred: zero, rejected: zero });
82
+ assert.equal(block.chair.members_reported, 2);
83
+ assert.deepEqual(buildCouncil({ ...input(), chair: { model: "p/chair:medium", dispatches: 1, clusters: [] }, audit: [] })!.chair, { model: "p/chair:medium", dispatches: 1, clusters: 0, members_reported: 0 });
84
+ });
85
+ test("incomplete batches return null", () => {
86
+ assert.equal(buildCouncil({ ...input(), audit: [...audit, { disposition: "rejected", severity: "minor", raisedBy: ["p-nobody"] }] }), null);
87
+ assert.equal(buildCouncil({ ...input(), chair: { ...input().chair, clusters: [...clusters, { severity: "minor", raisedBy: ["p-nobody"] }] } }), null);
88
+ assert.equal(buildCouncil({ ...input(), audit: [...audit, { disposition: "rejected", severity: "minor", raisedBy: ["p-alpha"] }] }), null);
89
+ assert.equal(buildCouncil({ ...input(), audit: audit.slice(0, 2) }), null);
90
+ assert.equal(buildCouncil({ ...input(), audit: [{ ...audit[0], severity: "major" }, audit[1], audit[2]] }), null);
91
+ });
@@ -0,0 +1,121 @@
1
+ export type Severity = "blocker" | "major" | "minor";
2
+ export type Disposition = "applied" | "deferred" | "rejected";
3
+ export type SeverityCounts = Record<Severity, number>;
4
+ export interface Cluster { severity: Severity; raisedBy: string[] }
5
+ export interface AuditItem extends Cluster { disposition: Disposition }
6
+ export interface CouncilMember {
7
+ dispatches: number; total: SeverityCounts; unique: SeverityCounts; applied: SeverityCounts;
8
+ unique_applied: SeverityCounts; deferred: SeverityCounts; rejected: SeverityCounts;
9
+ }
10
+ export interface CouncilBlock {
11
+ chair: { model: string; dispatches: number; clusters: number; members_reported: number };
12
+ members: Record<string, CouncilMember>;
13
+ }
14
+ export interface MemberResult { model?: string; exitCode?: number; savedOutputPath?: string }
15
+ export interface ResolvedMembers { dispatches: Map<string, number>; slugToModel: Map<string, string> }
16
+ const zero = (): SeverityCounts => ({ blocker: 0, major: 0, minor: 0 });
17
+ export const slug = (model: string): string => model.replace(/[^A-Za-z0-9]/g, "-");
18
+ export function memberSlugOf(path: string | undefined): string | undefined {
19
+ if (!path) return undefined;
20
+ return /^member-\d+-(.+?)(\.md)?$/.exec(path.split(/[\\/]/).pop() ?? "")?.[1];
21
+ }
22
+ export const normalizeRaisedBy = (token: string): string => slug(token.trim().replace(/^member-\d+-/, "").replace(/\.md$/, ""));
23
+ const attributes = (line: string): Cluster | undefined => {
24
+ const severity = /\[(blocker|major|minor)\]/i.exec(line);
25
+ const raisers = /raised-by:\s*\[([^\]]*)\]/i.exec(line);
26
+ if (!severity || !raisers) return undefined;
27
+ const raisedBy = [...new Set(raisers[1].split(",").map(normalizeRaisedBy).filter((x) => x && x !== "-"))];
28
+ return raisedBy.length ? { severity: severity[1].toLowerCase() as Severity, raisedBy } : undefined;
29
+ };
30
+ export function parseChairReport(text: string): { clusters: Cluster[] } | null {
31
+ if (!/^consensus:/m.test(text)) return null;
32
+ const lines = text.split("\n");
33
+ const start = lines.findIndex((line) => /^clusters:\s*$/.test(line));
34
+ if (start < 0) return null;
35
+ const clusters: Cluster[] = [];
36
+ for (const line of lines.slice(start + 1)) {
37
+ if (/^resolved:/.test(line)) break;
38
+ if (!/^\s*[-*]\s+/.test(line)) continue;
39
+ const item = attributes(line);
40
+ if (!item) return null;
41
+ clusters.push(item);
42
+ }
43
+ return { clusters };
44
+ }
45
+ export function parseAudit(text: string): AuditItem[] | null {
46
+ const seen = new Set<string>();
47
+ const items: AuditItem[] = [];
48
+ for (const line of text.split("\n")) {
49
+ const match = /^\s*(?:[-*]\s*)?(Applied|Deferred|Rejected):\s*(.*)$/.exec(line);
50
+ if (!match) continue;
51
+ const disposition = match[1].toLowerCase() as Disposition;
52
+ seen.add(disposition);
53
+ if (match[2].trim().toLowerCase() === "none") continue;
54
+ const item = attributes(match[2]);
55
+ if (!item) return null;
56
+ items.push({ disposition, ...item });
57
+ }
58
+ return seen.size === 3 ? items : null;
59
+ }
60
+ export function resolveMembers(results: MemberResult[]): ResolvedMembers | null {
61
+ const dispatches = new Map<string, number>();
62
+ const slugToModel = new Map<string, string>();
63
+ for (const result of results) {
64
+ if (typeof result.model !== "string") continue;
65
+ dispatches.set(result.model, (dispatches.get(result.model) ?? 0) + 1);
66
+ const memberSlug = result.exitCode === 0 ? memberSlugOf(result.savedOutputPath) : undefined;
67
+ if (!memberSlug) continue;
68
+ const previous = slugToModel.get(memberSlug);
69
+ if (previous !== undefined && previous !== result.model) return null;
70
+ slugToModel.set(memberSlug, result.model);
71
+ }
72
+ return { dispatches, slugToModel };
73
+ }
74
+ export interface CouncilInput {
75
+ chair: { model: string; dispatches: number; clusters: Cluster[] };
76
+ members: ResolvedMembers;
77
+ audit: AuditItem[];
78
+ }
79
+ export function buildCouncil(input: CouncilInput): CouncilBlock | null {
80
+ const resolve = (slugs: string[]): string[] | undefined => {
81
+ const models = new Set<string>();
82
+ for (const memberSlug of slugs) {
83
+ const model = input.members.slugToModel.get(memberSlug);
84
+ if (model === undefined) return undefined;
85
+ models.add(model);
86
+ }
87
+ return [...models].sort();
88
+ };
89
+ const key = (severity: Severity, models: string[]) => JSON.stringify([severity, models]);
90
+ const reported = new Set<string>();
91
+ const clusterKeys: string[] = [];
92
+ for (const cluster of input.chair.clusters) {
93
+ const models = resolve(cluster.raisedBy);
94
+ if (!models) return null;
95
+ for (const model of models) reported.add(model);
96
+ clusterKeys.push(key(cluster.severity, models));
97
+ }
98
+ const resolvedAudit: { item: AuditItem; models: string[] }[] = [];
99
+ for (const item of input.audit) {
100
+ const models = resolve(item.raisedBy);
101
+ if (!models) return null;
102
+ resolvedAudit.push({ item, models });
103
+ }
104
+ if (JSON.stringify(clusterKeys.sort()) !== JSON.stringify(resolvedAudit.map(({ item, models }) => key(item.severity, models)).sort())) return null;
105
+ const members: Record<string, CouncilMember> = {};
106
+ for (const [model, dispatches] of [...input.members.dispatches].sort(([a], [b]) => a.localeCompare(b))) {
107
+ members[model] = { dispatches, total: zero(), unique: zero(), applied: zero(), unique_applied: zero(), deferred: zero(), rejected: zero() };
108
+ }
109
+ for (const { item, models } of resolvedAudit) {
110
+ for (const model of models) {
111
+ const member = members[model];
112
+ member.total[item.severity]++;
113
+ member[item.disposition][item.severity]++;
114
+ if (models.length === 1) {
115
+ member.unique[item.severity]++;
116
+ if (item.disposition === "applied") member.unique_applied[item.severity]++;
117
+ }
118
+ }
119
+ }
120
+ return { chair: { model: input.chair.model, dispatches: input.chair.dispatches, clusters: input.chair.clusters.length, members_reported: reported.size }, members };
121
+ }
@@ -4,6 +4,7 @@ import {
4
4
  capEvents, derive, diffPhases, emptyAccumulators, emptyPhases, foldAccumulators, implementAutoCompletes,
5
5
  newRecord, parseRecord, serializeRecord, type Accumulators, type TelemetryEvent,
6
6
  } from "./telemetry-record.ts";
7
+ import type { CouncilBlock } from "./telemetry-council.ts";
7
8
 
8
9
  const ev = (over: Partial<TelemetryEvent> & { kind: string }): TelemetryEvent =>
9
10
  ({ ts: "2026-09-17T10:00:00Z", session: "s1", phase: "unphased", ...over }) as TelemetryEvent;
@@ -165,3 +166,28 @@ test("parseRecord drops malformed accumulator blocks and fills missing fields",
165
166
  assert.equal(derive(parsed, "2026-09-17T10:01:00Z").amendments, 2);
166
167
  });
167
168
 
169
+ const COUNCIL: CouncilBlock = {
170
+ chair: { model: "p/chair:medium", dispatches: 1, clusters: 2, members_reported: 2 },
171
+ members: {
172
+ "p/alpha:xhigh": { dispatches: 1, total: { blocker: 0, major: 1, minor: 1 }, unique: { blocker: 0, major: 1, minor: 0 }, applied: { blocker: 0, major: 1, minor: 0 }, unique_applied: { blocker: 0, major: 1, minor: 0 }, deferred: { blocker: 0, major: 0, minor: 0 }, rejected: { blocker: 0, major: 0, minor: 1 } },
173
+ "p/beta:high": { dispatches: 2, total: { blocker: 0, major: 0, minor: 1 }, unique: { blocker: 0, major: 0, minor: 0 }, applied: { blocker: 0, major: 0, minor: 0 }, unique_applied: { blocker: 0, major: 0, minor: 0 }, deferred: { blocker: 0, major: 0, minor: 0 }, rejected: { blocker: 0, major: 0, minor: 1 } },
174
+ },
175
+ };
176
+
177
+ test("derive carries derived.council through like tests/diff; absent stays absent", () => {
178
+ const rec = newRecord({ spec: "doc/specs/a.md", session: "s1", now: "2026-09-17T10:00:00Z", runId: "r1" });
179
+ assert.equal(derive(rec, "2026-09-17T10:01:00Z").council, undefined);
180
+ rec.derived.council = COUNCIL;
181
+ assert.deepEqual(derive(rec, "2026-09-17T10:01:00Z").council, COUNCIL);
182
+ });
183
+
184
+ test("serializeRecord writes council.chair and each member as flow maps under block keys; parseRecord round-trips", () => {
185
+ const rec = newRecord({ spec: "doc/specs/a.md", session: "s1", now: "2026-09-17T10:00:00Z", runId: "r1" });
186
+ rec.derived.council = COUNCIL;
187
+ const text = serializeRecord(rec);
188
+ assert.match(text, /^ council:\n chair: \{ model: p\/chair:medium, dispatches: 1, clusters: 2, members_reported: 2 \}\n members:\n p\/alpha:xhigh: \{ dispatches: 1, total: \{ blocker: 0, major: 1, minor: 1 \}/m);
189
+ assert.match(text, /^ p\/beta:high: \{ dispatches: 2, /m);
190
+ assert.deepEqual(parseRecord(text)!.derived.council, COUNCIL);
191
+ const plain = newRecord({ spec: "doc/specs/a.md", session: "s1", now: "2026-09-17T10:00:00Z", runId: "r2" });
192
+ assert.equal(parseRecord(serializeRecord(plain))!.derived.council, undefined);
193
+ });
@@ -1,6 +1,7 @@
1
1
  // Telemetry record model (#33): phases, events, accumulators, derive, YAML.
2
2
  import { Document, isCollection, isMap, isSeq, parse as parseYaml, stringify as stringifyYaml, type YAMLMap } from "yaml";
3
3
  import type { BucketStat, ShipOption } from "./telemetry-paths.ts";
4
+ import type { CouncilBlock } from "./telemetry-council.ts";
4
5
 
5
6
  export const PHASES = ["brainstorm", "plan", "implement", "verify", "ship"] as const;
6
7
  export type Phase = (typeof PHASES)[number];
@@ -91,7 +92,7 @@ export interface Derived {
91
92
  duration_s: number; phases: Partial<Record<PhaseKey, PhaseAcc & { started_at?: string; completed_at?: string; duration_s?: number; model?: string; thinking?: string }>>;
92
93
  personas: Record<string, PersonaAcc>; reviews?: Record<string, ReviewAcc>; conformance_loops: number; conformance_open_gaps?: number; gates: Gates & { ship_option?: string };
93
94
  plan?: { tasks: number; complete: number; failed: number; skipped: number }; tests?: { command: string; result: "pass" | "fail" }; amendments: number;
94
- spec_edits_after_ship: number; diff?: DiffSummary; modified_files?: string[]; spec_writes: Accumulators["spec_writes"]; events_dropped: number;
95
+ spec_edits_after_ship: number; diff?: DiffSummary; modified_files?: string[]; council?: CouncilBlock; spec_writes: Accumulators["spec_writes"]; events_dropped: number;
95
96
  }
96
97
  export type RecordStatus = "in_progress" | "shipped" | "abandoned";
97
98
  export interface TelemetryRecord {
@@ -121,7 +122,7 @@ export function derive(rec: TelemetryRecord, now: string): Derived {
121
122
  }
122
123
  for (const k of Object.keys(acc.phases) as PhaseKey[]) phases[k] = compact({ ...(phases[k] ?? {}), ...acc.phases[k] });
123
124
  const live = liveShipEvent(rec.events);
124
- return compact({ duration_s: seconds(rec.created_at, rec.shipped_at ?? rec.abandoned_at ?? now), phases, personas: acc.personas, reviews: Object.keys(acc.reviews).length ? acc.reviews : undefined, conformance_loops: acc.conformance_loops, conformance_open_gaps: acc.conformance_open_gaps, gates: compact({ ...acc.gates, ship_option: live?.option }), plan: rec.derived.plan, tests: rec.derived.tests, amendments: acc.amendments, spec_edits_after_ship: acc.spec_edits_after_ship, diff: rec.derived.diff, modified_files: rec.derived.modified_files, spec_writes: acc.spec_writes, events_dropped: acc.events_dropped }) as Derived;
125
+ return compact({ duration_s: seconds(rec.created_at, rec.shipped_at ?? rec.abandoned_at ?? now), phases, personas: acc.personas, reviews: Object.keys(acc.reviews).length ? acc.reviews : undefined, conformance_loops: acc.conformance_loops, conformance_open_gaps: acc.conformance_open_gaps, gates: compact({ ...acc.gates, ship_option: live?.option }), plan: rec.derived.plan, tests: rec.derived.tests, amendments: acc.amendments, spec_edits_after_ship: acc.spec_edits_after_ship, diff: rec.derived.diff, modified_files: rec.derived.modified_files, council: rec.derived.council, spec_writes: acc.spec_writes, events_dropped: acc.events_dropped }) as Derived;
125
126
  }
126
127
 
127
128
  const stripUndefined = (v: unknown): unknown => {
@@ -156,7 +157,11 @@ export function serializeRecord(rec: TelemetryRecord): string {
156
157
  const body = stripUndefined({ ...head, derived, accumulators }) as Record<string, unknown>;
157
158
  const doc = new Document(body);
158
159
  const derivedNode = doc.get("derived", true);
159
- if (isMap(derivedNode)) flowLeafChildren(derivedNode, new Set(["modified_files"]));
160
+ if (isMap(derivedNode)) {
161
+ flowLeafChildren(derivedNode, new Set(["modified_files", "council"]));
162
+ const council = derivedNode.get("council", true);
163
+ if (isMap(council)) flowLeafChildren(council);
164
+ }
160
165
  const accumulatorNode = doc.get("accumulators", true);
161
166
  if (isMap(accumulatorNode)) {
162
167
  for (const session of accumulatorNode.items) if (isMap(session.value)) flowLeafChildren(session.value);
@@ -511,6 +511,96 @@
511
511
 
512
512
  const subagentResult = (results: unknown[], content = "") => ({ toolName: "subagent", toolCallId: "sa", input: {}, content: [{ type: "text", text: content }], isError: false, details: { results } });
513
513
 
514
+ async function boundInBrainstorm() {
515
+ const h = harness();
516
+ await h.phaseResult("start", P({ brainstorm: "in_progress" }));
517
+ await h.writeSpec("doc/specs/a.md");
518
+ return h;
519
+ }
520
+ const memberBatch = (isError = false) => ({ ...subagentResult([
521
+ { agent: "spec-council-member", exitCode: 0, model: "p/alpha:xhigh", savedOutputPath: "/tmp/c/member-0-p-alpha.md", usage: usage(1) },
522
+ { agent: "spec-council-member", exitCode: 1, model: "p/beta:high", usage: usage(1) },
523
+ { agent: "spec-council-member", exitCode: 0, model: "p/beta:high", savedOutputPath: "/tmp/c/retry/member-1-p-beta-high.md", usage: usage(1) },
524
+ ]), isError });
525
+ const CHAIR_TEXT = "consensus: needs-work\nclusters:\n- [major] a - raised-by: [p-alpha] - x\n- [minor] b - raised-by: [p-alpha, p-beta-high] - x\nresolved:\n";
526
+ const chairResult = (text = CHAIR_TEXT, exitCode = 0) => ({ ...subagentResult([{ agent: "spec-council-synthesizer", exitCode, model: "p/chair:medium", usage: usage(1) }], text), isError: exitCode !== 0 });
527
+ const AUDIT_TEXT = "Applied: [major] a - raised-by: [p-alpha] -> edit\nDeferred: none\nRejected: [minor] b - raised-by: [p-alpha, p-beta-high] -> reason";
528
+ const assistant = (text: string) => ({ type: "message_end", message: { role: "assistant", usage: usage(1), content: [{ type: "text", text }] } });
529
+ const counts = (major = 0, minor = 0) => ({ blocker: 0, major, minor });
530
+ const EXPECTED_COUNCIL = {
531
+ chair: { model: "p/chair:medium", dispatches: 1, clusters: 2, members_reported: 2 },
532
+ members: {
533
+ "p/alpha:xhigh": { dispatches: 1, total: counts(1, 1), unique: counts(1), applied: counts(1), unique_applied: counts(1), deferred: counts(), rejected: counts(0, 1) },
534
+ "p/beta:high": { dispatches: 2, total: counts(0, 1), unique: counts(), applied: counts(), unique_applied: counts(), deferred: counts(), rejected: counts(0, 1) },
535
+ },
536
+ };
537
+
538
+ test("council captures failed member batch, chair, audit and persists on later flush", async () => {
539
+ const h = await boundInBrainstorm();
540
+ await h.emit("tool_result", memberBatch(true));
541
+ await h.emit("tool_result", chairResult());
542
+ assert.equal(h.readRecord().derived.council, undefined);
543
+ await h.emit("message_end", assistant(AUDIT_TEXT));
544
+ assert.deepEqual(h.readRecord().derived.council, EXPECTED_COUNCIL);
545
+ await h.phaseResult("complete", P({ brainstorm: "complete" }));
546
+ assert.deepEqual(h.readRecord().derived.council, EXPECTED_COUNCIL);
547
+ assert.equal(h.readRecord().derived.personas["spec-council-member"], undefined);
548
+ });
549
+ test("council defaults missing chair model and survives unrelated result before audit", async () => {
550
+ const h = await boundInBrainstorm();
551
+ await h.emit("tool_result", memberBatch());
552
+ await h.emit("tool_result", subagentResult([{ agent: "spec-council-synthesizer", exitCode: 0, usage: usage(1) }], CHAIR_TEXT));
553
+ await h.emit("tool_result", subagentResult([{ agent: "spec-summarizer", exitCode: 0, model: "p/s", usage: usage(1) }]));
554
+ await h.emit("message_end", assistant(AUDIT_TEXT));
555
+ assert.deepEqual(h.readRecord().derived.council, { ...EXPECTED_COUNCIL, chair: { ...EXPECTED_COUNCIL.chair, model: "unknown" } });
556
+ });
557
+ test("council ignores plan phase", async () => {
558
+ const h = await boundInPlan();
559
+ await h.emit("tool_result", memberBatch());
560
+ await h.emit("tool_result", chairResult());
561
+ await h.emit("message_end", assistant(AUDIT_TEXT));
562
+ assert.equal(h.readRecord().derived.council, undefined);
563
+ });
564
+ test("council counts failed chair retries and retains last usable clusters", async () => {
565
+ const h = await boundInBrainstorm();
566
+ await h.emit("tool_result", memberBatch());
567
+ await h.emit("tool_result", chairResult("Failed", 1));
568
+ await h.emit("tool_result", chairResult());
569
+ await h.emit("tool_result", chairResult("no consensus here", 1));
570
+ await h.emit("message_end", assistant(AUDIT_TEXT));
571
+ assert.deepEqual(h.readRecord().derived.council!.chair, { ...EXPECTED_COUNCIL.chair, dispatches: 3 });
572
+ });
573
+ test("council retains pending on incomplete audit and ignores audit before chair", async () => {
574
+ const h = await boundInBrainstorm();
575
+ await h.emit("message_end", assistant(AUDIT_TEXT));
576
+ await h.emit("tool_result", memberBatch());
577
+ await h.emit("tool_result", chairResult());
578
+ await h.emit("message_end", assistant("Applied: [major] a - raised-by: [p-alpha] -> edit\nDeferred: none\nRejected: none"));
579
+ assert.equal(h.readRecord().derived.council, undefined);
580
+ await h.emit("message_end", assistant("Applied: a -> edit\nDeferred: none\nRejected: none"));
581
+ assert.equal(h.readRecord().derived.council, undefined);
582
+ await h.emit("message_end", assistant(AUDIT_TEXT));
583
+ assert.deepEqual(h.readRecord().derived.council, EXPECTED_COUNCIL);
584
+ });
585
+ test("council reset clears pending and new member batch supersedes held chair", async () => {
586
+ const h = await boundInBrainstorm();
587
+ await h.emit("tool_result", memberBatch());
588
+ await h.emit("tool_result", chairResult());
589
+ await h.phaseResult("reset", P());
590
+ await h.phaseResult("start", P({ brainstorm: "in_progress" }));
591
+ await h.writeSpec("doc/specs/a.md");
592
+ await h.emit("message_end", assistant(AUDIT_TEXT));
593
+ assert.equal(h.readRecord().derived.council, undefined);
594
+ await h.emit("tool_result", memberBatch());
595
+ await h.emit("tool_result", chairResult());
596
+ await h.emit("tool_result", memberBatch());
597
+ await h.emit("message_end", assistant(AUDIT_TEXT));
598
+ assert.equal(h.readRecord().derived.council, undefined);
599
+ await h.emit("tool_result", chairResult());
600
+ await h.emit("message_end", assistant(AUDIT_TEXT));
601
+ assert.deepEqual(h.readRecord().derived.council, EXPECTED_COUNCIL);
602
+ });
603
+
514
604
  test("subagent results: dispatch events, persona tokens, spec_rounds, reviews, conformance_loops; async results: [] contribute nothing", async () => {
515
605
  const h = await boundInPlan();
516
606
  await h.emit("tool_result", subagentResult([
@@ -24,6 +24,7 @@
24
24
  import { COUNTED_USER_PHASES, REVIEWER_AGENTS, countFindings, countOpenGaps, hasReopen, insertedText, planTotals, textOf, usageToTokens } from "./lib/telemetry-collect.ts";
25
25
  import { addTokens, capEvents, compact, currentPhase, derive, diffPhases, emptyAccumulators, emptyPhases, foldAccumulators, implementAutoCompletes, liveShipEvent, newRecord, parseRecord, serializeRecord, type Accumulators, type BaseEvent, type PhaseAcc, type PhaseKey, type PhaseMap, type TelemetryEvent, type TelemetryRecord, type Tokens } from "./lib/telemetry-record.ts";
26
26
  import { guardReason } from "./lib/telemetry-ship.ts";
27
+ import { buildCouncil, parseAudit, parseChairReport, resolveMembers, type Cluster, type MemberResult } from "./lib/telemetry-council.ts";
27
28
 
28
29
  // ---- deps seam -------------------------------------------------------------------
29
30
 
@@ -204,6 +205,7 @@
204
205
  let settingsWarned = false;
205
206
  let sessionId = "";
206
207
  let currentDir = "";
208
+ let councilPending: { memberResults: MemberResult[]; chair?: { model: string; dispatches: number; clusters?: Cluster[] } } | undefined;
207
209
 
208
210
  // Per-event settings read; undefined => this event is a no-op.
209
211
  const enabledSettings = (ctx: ExtensionContext): SettingsSnapshot | undefined => {
@@ -364,6 +366,7 @@
364
366
  record = undefined;
365
367
  recordRel = undefined;
366
368
  frozen = false;
369
+ councilPending = undefined;
367
370
  };
368
371
 
369
372
  const isSafeSpecLink = (link: string): boolean => {
@@ -432,6 +435,7 @@
432
435
  ? { kind: "phase", action: "reset" }
433
436
  : { kind: "phase", action: t.action, name: t.name },
434
437
  );
438
+ if (t.action === "start" && t.name === "brainstorm") councilPending = undefined;
435
439
  phases = details.phases;
436
440
  if (record && t.name === "brainstorm" && (t.action === "complete" || t.action === "skip")) record.approved_at ??= e.ts;
437
441
  if (record && t.name === "ship" && t.action === "complete" && checkoutVia === "jj") await onShipKeep(snap);
@@ -538,8 +542,9 @@
538
542
  }
539
543
  return onSpecInteraction(event.toolName, loc, event, snap);
540
544
  }
541
- if (event.toolName === "subagent" && !event.isError) {
542
- await onSubagentResult(event.details, event.content);
545
+ if (event.toolName === "subagent") {
546
+ if (phaseNow() === "brainstorm") onCouncilResults(event.details, event.content);
547
+ if (!event.isError) await onSubagentResult(event.details, event.content);
543
548
  return undefined;
544
549
  }
545
550
  if (event.toolName === "plan_tracker" && !event.isError) {
@@ -625,6 +630,42 @@
625
630
  acc.tokens = addTokens(acc.tokens, t);
626
631
  };
627
632
 
633
+ const onCouncilResults = (details: unknown, content: unknown) => {
634
+ const results = (details as { results?: unknown[] } | undefined)?.results;
635
+ if (!Array.isArray(results) || !results.length) return;
636
+ const typed = results.map((raw) => raw as { agent?: unknown; exitCode?: unknown; model?: unknown; savedOutputPath?: unknown });
637
+ if (councilPending?.chair && typed.some((r) => r.agent === "spec-council-member")) councilPending = undefined;
638
+ for (const r of typed) {
639
+ if (r.agent === "spec-council-member") {
640
+ (councilPending ??= { memberResults: [] }).memberResults.push({
641
+ model: typeof r.model === "string" ? r.model : undefined,
642
+ exitCode: typeof r.exitCode === "number" ? r.exitCode : undefined,
643
+ savedOutputPath: typeof r.savedOutputPath === "string" ? r.savedOutputPath : undefined,
644
+ });
645
+ } else if (r.agent === "spec-council-synthesizer") {
646
+ const chair = ((councilPending ??= { memberResults: [] }).chair ??= { model: "unknown", dispatches: 0 });
647
+ if (typeof r.model === "string") chair.model = r.model;
648
+ chair.dispatches += 1;
649
+ const parsed = parseChairReport(textOf(content));
650
+ if (parsed) chair.clusters = parsed.clusters;
651
+ }
652
+ }
653
+ };
654
+
655
+ const onCouncilAudit = (content: unknown) => {
656
+ const pending = councilPending;
657
+ const clusters = pending?.chair?.clusters;
658
+ if (!record || !pending?.chair || !clusters) return;
659
+ const audit = parseAudit(textOf(content));
660
+ if (!audit) return;
661
+ const members = resolveMembers(pending.memberResults);
662
+ const block = members && buildCouncil({ chair: { ...pending.chair, clusters }, members, audit });
663
+ if (!block) return;
664
+ record.derived.council = block;
665
+ councilPending = undefined;
666
+ flush();
667
+ };
668
+
628
669
  const onSubagentResult = async (details: unknown, content: unknown) => {
629
670
  const results = (details as { results?: unknown[] } | undefined)?.results;
630
671
  if (!Array.isArray(results) || results.length === 0) return;
@@ -723,9 +764,10 @@
723
764
  pi.on("message_end", async (event, ctx) => {
724
765
  lastCtx = ctx;
725
766
  if (!enabledSettings(ctx)) return;
726
- const msg = event.message as { role?: string; usage?: unknown };
767
+ const msg = event.message as { role?: string; usage?: unknown; content?: unknown };
727
768
  if (msg.role !== "assistant") return;
728
769
  addPhaseTokens(usageToTokens(msg.usage));
770
+ if (phaseNow() === "brainstorm") onCouncilAudit(msg.content);
729
771
  });
730
772
 
731
773
  pi.on("turn_end", async (_event, ctx) => {
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "pi-gauntlet",
3
- "version": "5.20.0",
3
+ "version": "5.21.0",
4
4
  "description": "Opinionated, gated workflow skills, subagent personas, and runtime extensions for the pi coding agent.",
5
5
  "author": "Jacek Juraszek",
6
6
  "type": "module",
@@ -238,10 +238,10 @@ Paste the summary verbatim, unedited in the template below; use adjacent lines f
238
238
  Spec written and committed to <project>/doc/specs/<filename>.md (worktree: <path>).
239
239
 
240
240
  Coverage: <N> of <M> members reported; <slug>: <reason> (line present only when coverage was partial)
241
- Applied: <cluster -> edit>, ...
242
- Deferred: <cluster -> where it belongs>, ...
243
- Rejected: <cluster -> one-line reason>, ...
244
- (omit the audit lines above when the worker path ran, not the council)
241
+ Applied: [<severity>] <cluster> — raised-by: [<slugs>] -> <edit>
242
+ Deferred: [<severity>] <cluster> — raised-by: [<slugs>] -> <where it belongs>
243
+ Rejected: [<severity>] <cluster> — raised-by: [<slugs>] -> <one-line reason>
244
+ (one line per item, exactly as returned by roasting-the-spec - `Applied: none` / `Deferred: none` / `Rejected: none` when a list is empty; omit the audit lines when the worker path ran, not the council)
245
245
 
246
246
  <unresolved ambiguities; every gap-footer entry from the summary>
247
247
 
@@ -18,13 +18,14 @@ The CLI parses and aggregates; you reason. Never open a telemetry YAML record yo
18
18
 
19
19
  ## Read the digest
20
20
 
21
- `runs` has one row per record; `shipped*` marks a truncated run (no ship phase recorded - an older salvage stamped it at landing; its `wall` is `-` and it feeds no aggregate). `by version` groups by pi-gauntlet version: `n` counts every row, `shipped` counts the rows behind the p50/max columns. `grants` is `fix_round_grants` - the fix-round proxy (human-granted extra review rounds); schema 1 has no code-review round count.
21
+ `runs` has one row per record; `shipped*` marks a truncated run (no ship phase recorded - an older salvage stamped it at landing; its `wall` is `-` and it feeds no aggregate). `by version` groups by pi-gauntlet version: `n` counts every row, `shipped` counts the rows behind the p50/max columns. `grants` is `fix_round_grants` - the fix-round proxy (human-granted extra review rounds); schema 1 has no code-review round count. Read `council` by roster: counts are blocker/major/minor, `uniq_appl_nonminor/run` counts unique-and-applied blocker/major findings per run, and `chair:` summarizes synthesis and retries. Read severity as the chair's consolidated severity (the highest any co-raiser assigned), not the member's own. A roster with fewer than five runs is not assessed, and older records can have no council data.
22
22
 
23
23
  ## Reply - exactly this, in this order
24
24
 
25
25
  1. **Recommendation** (2-4 sentences). One claim, led by the run that exemplifies it: quote its `spec` slug, `run_id`, and the 1-3 numbers that carry the claim. When `grants` is the evidence, call it the fix-round proxy. If no version group has `shipped >= 2`, the recommendation is "sample too small" with the `n`/`shipped` counts per version.
26
26
  2. **Cornerstones**: 3-5 bullets of aggregate facts from `by version` - corpus size, truncated count, the p50s and model tallies that moved between versions.
27
- 3. **Menu**, numbered, at most 3 items, rendered exactly as:
27
+ 3. **Council assessment**: Render this block independently of the version sample gate, even when item 1 says "sample too small". If the digest prints `no council data in selected records`, render only `council: no data`. If the most recent roster has fewer than 5 runs, render only `council: <roster> - not enough runs to assess (N of 5)`. Otherwise render one line per `flag:` of the most recent roster as the roster-change recommendations; when an older roster exists, add one line comparing the two rosters' `uniq_appl_nonminor/run` and rejected share as groups. Compare rosters as groups, never a member across rosters.
28
+ 4. **Menu**, numbered, at most 3 items, rendered exactly as:
28
29
  - `1. render report` - ask for a target path; write markdown there: the digest verbatim, then the recommendation and cornerstones above. If the file exists, ask before overwriting. Write nothing unless this item is chosen.
29
30
  - `2. open recommendation as ticket` - hand the claim and its numbers to `/skill:shape-ticket`; never create a ticket directly.
30
31
  - `3. drill into <slug>` - re-run the CLI with `--json` and show that run's fields.
@@ -126,7 +126,7 @@ For each cluster in the chair's report, decide one of:
126
126
 
127
127
  Also inline any `external-ref:` cluster you have context for (e.g. a ticket fetched during brainstorming) as part of the apply-set — this is your call, same as any other cluster.
128
128
 
129
- An `over-spec:` cluster decided **apply** is executed as deletion or shrink of the quoted clause **and** any acceptance-criteria or testing-approach line that exists only for it. Its audit line reads `Applied: over-spec: <clause> -> cut (was adds: M files / N tests / K ACs)` or `Applied: over-spec: <clause> -> shrunk to <replacement> (was adds: ...)`, so the gate shows what was removed. `defer`/`reject` are unchanged.
129
+ An `over-spec:` cluster decided **apply** is executed as deletion or shrink of the quoted clause **and** any acceptance-criteria or testing-approach line that exists only for it. Its audit line reads `Applied: [<severity>] over-spec: <clause> — raised-by: [<slugs>] -> cut (was adds: M files / N tests / K ACs)` or `Applied: [<severity>] over-spec: <clause> — raised-by: [<slugs>] -> shrunk to <replacement> (was adds: ...)`, so the gate shows what was removed. `defer`/`reject` are unchanged.
130
130
 
131
131
  You are the advocate — decide on scope grounds — and, unlike a dispatched subagent, also the executor: you hold `edit`/`write` tools directly, so apply the edit yourself instead of proposing it for someone else to make. Do this **before** returning to brainstorming.
132
132
 
@@ -135,9 +135,9 @@ You are the advocate — decide on scope grounds — and, unlike a dispatched su
135
135
  Return a structured audit, gate-only (not a committed spec section) — a coverage line plus three labelled lists:
136
136
 
137
137
  - `Coverage:` — `N of M members reported; <slug>: <reason>` — present only when member coverage was partial; omitted at full coverage.
138
- - `Applied:` — one of `Applied: <cluster> -> <edit> (grounded: <member probe>)`, `Applied: <cluster> -> <edit> (probed: <check> - <result>)` for a confirmed hypothesis, `Applied: <cluster> -> open question (<not found | inconclusive: <check> | contradicted: <result>>)`. The probe rides on the audit line because member files are removed in section 5.
139
- - `Deferred:` — cluster -> where it belongs.
140
- - `Rejected:` — cluster -> one-line reason.
138
+ - `Applied:` — one line per applied cluster: `Applied: [<severity>] <cluster> — raised-by: [<slugs>] -> <edit> (grounded: <member probe>)`, `... -> <edit> (probed: <check> - <result>)` for a confirmed hypothesis, or `... -> open question (<not found | inconclusive: <check> | contradicted: <result>>)`. Copy `[<severity>]` and `raised-by: [...]` verbatim from the cluster line - telemetry joins on them. No applied cluster -> the single line `Applied: none`. The probe rides on the audit line because member files are removed in section 5.
139
+ - `Deferred:` — one line per deferred cluster, same `[<severity>] <cluster> — raised-by: [<slugs>]` prefix, then `-> <where it belongs>`. None -> `Deferred: none`.
140
+ - `Rejected:` — one line per rejected cluster, same prefix, then `-> <one-line reason>`. None -> `Rejected: none`.
141
141
 
142
142
  Hand this audit to brainstorming along with the now-final spec. brainstorming writes it into the **spec commit message body** (git-native, readable pre-squash) so it survives for finish-time revert visibility, then shows it to the user alongside the final spec at its one review gate. The user can revert any applied edit there — that gate, not this skill, is where ratification happens.
143
143