pi-gauntlet 5.20.0 → 5.22.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +20 -0
- package/README.md +3 -3
- package/bin/gauntlet-performance.mjs +233 -115
- package/bin/gauntlet-performance.test.mjs +100 -0
- package/bin/gauntlet-spec-index.mjs +83 -22
- package/bin/gauntlet-spec-index.test.mjs +181 -61
- package/bin/gauntlet-telemetry-seal.mjs +6 -2
- package/bin/gauntlet-telemetry-seal.test.mjs +10 -0
- package/extensions/lib/telemetry-council.test.ts +91 -0
- package/extensions/lib/telemetry-council.ts +121 -0
- package/extensions/lib/telemetry-record.test.ts +26 -0
- package/extensions/lib/telemetry-record.ts +8 -3
- package/extensions/telemetry.test.ts +90 -0
- package/extensions/telemetry.ts +45 -3
- package/package.json +1 -1
- package/skills/brainstorming/SKILL.md +22 -6
- package/skills/brainstorming/gatherer.md +6 -3
- package/skills/gauntlet-performance/SKILL.md +3 -2
- package/skills/roasting-the-spec/SKILL.md +4 -4
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,25 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
+
## v5.22.0 - 2026-09-26
|
|
4
|
+
|
|
5
|
+
### Added
|
|
6
|
+
|
|
7
|
+
- `gauntlet-spec-index` gains a `state` column after `shipped_at` (`superseded` for a `- fully` banner directly under the title, `live` otherwise) and a repeatable `--exclude <path>` flag. `/skill:brainstorming` runs a second predecessor query at spec self-review from the finished spec's title, goal, and H2 headings and lists new `live` candidates at the review gate; the scout composes its first query from the ticket title and body.
|
|
8
|
+
|
|
9
|
+
### Changed
|
|
10
|
+
|
|
11
|
+
- `gauntlet-spec-index` applies a confidence rule before `--limit`: only terms in fewer than half the specs count as evidence, a row needs two distinct evidence terms in title or goal, and rows under half the best live score are dropped, so queries with no real match return zero rows. Live rows sort before superseded rows. The cache schema is version 2 and rebuilds on first use.
|
|
12
|
+
|
|
13
|
+
## v5.21.0 - 2026-09-26
|
|
14
|
+
|
|
15
|
+
### Added
|
|
16
|
+
|
|
17
|
+
- Council telemetry records chair activity and per-member severity/disposition counts in `derived.council`; `gauntlet-performance` groups council results by roster, reports chair summaries and contribution flags after five runs, and exposes the aggregates in JSON. The performance skill includes a council assessment in its reply.
|
|
18
|
+
|
|
19
|
+
### Changed
|
|
20
|
+
|
|
21
|
+
- `roasting-the-spec` audit lines and brainstorming's gate template carry `[<severity>]` and `raised-by: [...]` on every `Applied:`/`Deferred:`/`Rejected:` item, one item per line, `none` for an empty list. `gauntlet-performance` is split into section modules under `src/bins/performance/`; existing `runs` and `by version` output is unchanged.
|
|
22
|
+
|
|
3
23
|
## v5.20.0 - 2026-09-26
|
|
4
24
|
|
|
5
25
|
### Changed
|
package/README.md
CHANGED
|
@@ -69,7 +69,7 @@ Everything between gate 1 and gate 2 - task breakdown, implementation, both revi
|
|
|
69
69
|
|
|
70
70
|
pi-gauntlet ships three kinds of pieces, layered on top of pi-cohort's dispatch:
|
|
71
71
|
|
|
72
|
-
- **20 skills** - the workflow logic. Thirteen activate automatically when pi sees the matching kind of task, and each one gates the next: `brainstorming`, `writing-plans`, `roasting-the-spec`, `test-driven-development`, `subagent-driven-development`, `dispatching-parallel-agents`, `verification-before-completion`, `requesting-code-review`, `receiving-code-review`, `using-git-worktrees`, `finishing-a-development-branch`, `writing-skills`, `linear` (reads/searches/comments on/manages Linear tickets via the `linearis` CLI; owns all linearis mechanics and the `## Issue tracker` overrides schema; tracker-facing skills route to it). Seven more are explicit-invocation-only (`disable-model-invocation: true`): `shape-ticket` creates or repairs one tracker issue per run against a Context/Problem/Idea/Acceptance-Criteria template, gated by an AC integrity check, a cheap council roast, and a single human-confirmed write - run it with `/skill:shape-ticket`. `gatekeep-pr` is consent-gated pre-merge verification of a PR against its issue - read-only gathering, verification evidence resolved CI-first (green checks on the exact assessed head count as evidence; the project's verification command runs only as fallback), a rubric-based review, then a deterministic authorship-aware menu with stable finding IDs (P#/L#/C#/F#) and numbered pre-composed courses (fixes execute as a single parallel-safe wave: one gate run, one re-review, one push); nothing mutates (fixes, pushes, reviews, merges) until you pick a row - run it with `/skill:gatekeep-pr <pr>`. `check-delivery` is a post-merge detective control: proves an issue actually shipped (default-branch landing, delivery target, per-AC evidence) before its tracker status advances; it never writes a terminal status - run it with `/skill:check-delivery <ref>`. `chase-bug` is human-only bug triage: read-only root-cause discovery to an evidenced verdict menu (real bug -> ticket/brainstorm/hotfix/respond; five negative verdicts), then a gated response to the reporter for addressable origins (GitHub issue / tracker ticket) and a rendered verdict summary otherwise - it never fixes during triage; the hotfix row hands off to `skills/chase-bug/hotfix.md` after the menu - run it with `/skill:chase-bug`. `gauntlet-performance` reads the committed run telemetry (current repo by default, `--dir <path>` adds others) through the parse-only `gauntlet-performance` CLI and answers with one example-led recommendation, 3-5 cornerstone numbers, and a menu of at most three actions (render a report file, open the recommendation as a ticket via `shape-ticket`, drill into one run); it writes nothing unless you pick render - run it with `/skill:gauntlet-performance [--dir <path>]... [--since <version>]`. `gauntlet-handoff` ends a session whose flow a fresh session will continue: it invokes pi-cohort's `handoff` skill (`--out <path>` or `--key <name>`, default key = the run worktree's branch with `/` flattened to `-`, default file `<tmpdir>/pi-handoff/<key>.md`), then appends the gauntlet process-state section (phase/plan tracker status) per the shared contract `skills/gauntlet-resume/reference/brief-contract.md` - run it with `/skill:gauntlet-handoff [--out <path> | --key <name>]`. `gauntlet-resume` is the only way back into an interrupted flow from a fresh session: it takes a `gauntlet-handoff` brief (a file path, pasted text, or - with no arguments - a pick from the briefs under `<tmpdir>/pi-handoff/`), a bare worktree that already holds a spec, or a spec tracked on main under a configured spec dir (a spec seed: it creates or reuses the worktree through `using-git-worktrees` and pins the spec), restores phase/plan tracker state through the legal arming sequence (`start brainstorm`, `skip` with `resume:` reasons, `plan_check` before implement-or-later), and never infers approval from artifacts - run it with `/skill:gauntlet-resume [<brief-file> | <spec>.md | <worktree-name-or-path>]`.
|
|
72
|
+
- **20 skills** - the workflow logic. Thirteen activate automatically when pi sees the matching kind of task, and each one gates the next: `brainstorming`, `writing-plans`, `roasting-the-spec`, `test-driven-development`, `subagent-driven-development`, `dispatching-parallel-agents`, `verification-before-completion`, `requesting-code-review`, `receiving-code-review`, `using-git-worktrees`, `finishing-a-development-branch`, `writing-skills`, `linear` (reads/searches/comments on/manages Linear tickets via the `linearis` CLI; owns all linearis mechanics and the `## Issue tracker` overrides schema; tracker-facing skills route to it). Seven more are explicit-invocation-only (`disable-model-invocation: true`): `shape-ticket` creates or repairs one tracker issue per run against a Context/Problem/Idea/Acceptance-Criteria template, gated by an AC integrity check, a cheap council roast, and a single human-confirmed write - run it with `/skill:shape-ticket`. `gatekeep-pr` is consent-gated pre-merge verification of a PR against its issue - read-only gathering, verification evidence resolved CI-first (green checks on the exact assessed head count as evidence; the project's verification command runs only as fallback), a rubric-based review, then a deterministic authorship-aware menu with stable finding IDs (P#/L#/C#/F#) and numbered pre-composed courses (fixes execute as a single parallel-safe wave: one gate run, one re-review, one push); nothing mutates (fixes, pushes, reviews, merges) until you pick a row - run it with `/skill:gatekeep-pr <pr>`. `check-delivery` is a post-merge detective control: proves an issue actually shipped (default-branch landing, delivery target, per-AC evidence) before its tracker status advances; it never writes a terminal status - run it with `/skill:check-delivery <ref>`. `chase-bug` is human-only bug triage: read-only root-cause discovery to an evidenced verdict menu (real bug -> ticket/brainstorm/hotfix/respond; five negative verdicts), then a gated response to the reporter for addressable origins (GitHub issue / tracker ticket) and a rendered verdict summary otherwise - it never fixes during triage; the hotfix row hands off to `skills/chase-bug/hotfix.md` after the menu - run it with `/skill:chase-bug`. `gauntlet-performance` reads the committed run telemetry (current repo by default, `--dir <path>` adds others) through the parse-only `gauntlet-performance` CLI and answers with one example-led recommendation, 3-5 cornerstone numbers, a council assessment when available, and a menu of at most three actions (render a report file, open the recommendation as a ticket via `shape-ticket`, drill into one run); it writes nothing unless you pick render - run it with `/skill:gauntlet-performance [--dir <path>]... [--since <version>]`. `gauntlet-handoff` ends a session whose flow a fresh session will continue: it invokes pi-cohort's `handoff` skill (`--out <path>` or `--key <name>`, default key = the run worktree's branch with `/` flattened to `-`, default file `<tmpdir>/pi-handoff/<key>.md`), then appends the gauntlet process-state section (phase/plan tracker status) per the shared contract `skills/gauntlet-resume/reference/brief-contract.md` - run it with `/skill:gauntlet-handoff [--out <path> | --key <name>]`. `gauntlet-resume` is the only way back into an interrupted flow from a fresh session: it takes a `gauntlet-handoff` brief (a file path, pasted text, or - with no arguments - a pick from the briefs under `<tmpdir>/pi-handoff/`), a bare worktree that already holds a spec, or a spec tracked on main under a configured spec dir (a spec seed: it creates or reuses the worktree through `using-git-worktrees` and pins the spec), restores phase/plan tracker state through the legal arming sequence (`start brainstorm`, `skip` with `resume:` reasons, `plan_check` before implement-or-later), and never infers approval from artifacts - run it with `/skill:gauntlet-resume [<brief-file> | <spec>.md | <worktree-name-or-path>]`.
|
|
73
73
|
- **7 subagent personas** - the specialized child agents the skills dispatch via pi-cohort: `implementer`, `code-reviewer`, `spec-reviewer`, `conformance-reviewer`, `spec-summarizer`, `spec-council-member`, `spec-council-synthesizer`. See [doc/personas.md](./doc/personas.md) for what each one does and why its permissions are scoped the way they are.
|
|
74
74
|
- **4 runtime extensions** - the enforcement layer. `plan-tracker` and `phase-tracker` are tools skills call to track progress (with a TUI widget); `verify-before-ship` is a hook that warns if you push or open a PR without a passing test run since your last edit; a phase-tracker flow guard reminds on implement-phase commits missing spec/code review. In a brainstorming-entered flow, phase-tracker rejects `implement` or `verify` completion while tracker tasks remain pending or in progress; see [its configuration reference](./doc/configuration.md#phase-tracker). phase-tracker also registers `plan_check`, which verifies a plan against its spec and against the grammar in [skills/writing-plans/reference/plan-contract.md](./skills/writing-plans/reference/plan-contract.md), including that each task's `Tests:` commands are selective and never the full suite; a pass stamps the plan for implementation. `telemetry` records one YAML record per gauntlet run (phase timing, models, personas, gate/fix rounds, diff at ship) that ships in the squash beside the spec - it is written only inside a brainstorming-entered flow, kept untracked and git-excluded until `finishing-a-development-branch` seals and commits it once with `gauntlet-telemetry-seal` - and blocks a brainstorm `write` into an already-shipped spec; see [its configuration reference](./doc/configuration.md#telemetry). See [doc/configuration.md](./doc/configuration.md) for the settings each one reads.
|
|
75
75
|
|
|
@@ -137,11 +137,11 @@ Pin an exact release with `npm:pi-gauntlet@X.Y.Z`. See [doc/install-internals.md
|
|
|
137
137
|
|
|
138
138
|
## Spec search index
|
|
139
139
|
|
|
140
|
-
`gauntlet-spec-index` provides lexical search across `doc/specs/*.md` at the repository root and one service level down. From a repository worktree, run `node <pi-gauntlet-package>/bin/gauntlet-spec-index.mjs --query "<text>" [--limit N]
|
|
140
|
+
`gauntlet-spec-index` provides lexical search across `doc/specs/*.md` at the repository root and one service level down. From a repository worktree, run `node <pi-gauntlet-package>/bin/gauntlet-spec-index.mjs --query "<text>" [--limit N] [--exclude <repo-relative path>]...`; it requires Node >=24.15.0, refreshes its FTS5 index on every query, and prints tab-separated `score`, `path`, `service`, `title`, `status`, `shipped_at`, `state`, `files`, and `snippet` columns. Rows pass a confidence rule before `--limit` applies: a query term is evidence when it occurs in fewer than half the indexed specs (stem variants of one word count once); query tokens split on word boundaries, so `gauntlet-spec-index` searches as three words; a row is returned only when it matches at least two distinct evidence terms in its `title` or `goal`, and rows scoring below half the best live row's score are dropped. A query with no real match therefore prints the header only and exits 0, and a corpus of two or fewer specs never returns rows. `--exclude` (repeatable) removes a path before the rule runs, so the spec under review never sets the bar. `state` is `superseded` when a `> **Superseded by:** ... - fully` banner sits in the block directly under the title and `live` otherwise; named-section banners and consumer-defined banner syntaxes both read as `live`. Live rows sort before superseded rows, then by score. The `files` column is a `;`-separated list of repo-relative paths the spec's shipped change modified and that still exist in the repository, the literal `missing` when the spec's telemetry record has no `derived.modified_files` list, or blank when there is no readable record or no recorded path remains. The per-worktree cache lives at `.pi/gauntlet/index.sqlite`, and its first creation adds `/.pi/gauntlet/index.sqlite*` to Git's `info/exclude` so the database and SQLite sidecars stay out of `git status`. `/skill:brainstorming` queries the index twice: the scout at gather time from the request, and the main loop at spec-writing from the finished spec's title, goal, and headings, surfacing new `live` candidates at the review gate.
|
|
141
141
|
|
|
142
142
|
## Performance digest
|
|
143
143
|
|
|
144
|
-
`gauntlet-performance` digests telemetry records without judging them: `node <pi-gauntlet-package>/bin/gauntlet-performance.mjs [--dir <repo root or telemetry dir>]... [--since <version>] [--json]` prints a `corpus:` line, one row per run (`run_id`, repo, spec, version, status - `shipped*` when the record has no ship phase - wall, per-phase minutes, tokens, cost, models, dispatches, `fix_round_grants`, reopens, conformance loops, reviewer findings, council dispatches), a `by version` block of p50/max over shipped, non-truncated runs, and `skipped:` lines for files it could not use. Default corpus is the current repo's telemetry dir; each `--dir` adds a repo root (its own `telemetry.dir` setting is honoured) or a bare telemetry dir. `--since` compares versions numerically. Exit 0 always except usage errors. `/skill:gauntlet-performance` is the reasoning layer on top.
|
|
144
|
+
`gauntlet-performance` digests telemetry records without judging them: `node <pi-gauntlet-package>/bin/gauntlet-performance.mjs [--dir <repo root or telemetry dir>]... [--since <version>] [--json]` prints a `corpus:` line, one row per run (`run_id`, repo, spec, version, status - `shipped*` when the record has no ship phase - wall, per-phase minutes, tokens, cost, models, dispatches, `fix_round_grants`, reopens, conformance loops, reviewer findings, council dispatches), a `by version` block of p50/max over shipped, non-truncated runs, a `council` block grouped by member roster, and `skipped:` lines for files it could not use. Each council table shows member `total`, `unique`, `applied`, `uniq_appl`, `deferred`, `rejected` blocker/major/minor counts and `uniq_appl_nonminor/run`; `chair:` shows runs, average clusters and members reported, and retried runs. Roster flags (JSON `flags[].rule`: `low_unique_applied` - fewer than 0.2 unique-and-applied blocker/major findings per run; `high_rejection` - over 50% rejected while another member is below 25%) require five runs and print as `flag: <member> - <detail>`. Smaller groups print `not enough runs to assess (N of 5)`; absent council data prints `no council data in selected records`. JSON includes `council: [{ roster, runs, members, chairs, flags }]`. Default corpus is the current repo's telemetry dir; each `--dir` adds a repo root (its own `telemetry.dir` setting is honoured) or a bare telemetry dir. `--since` compares versions numerically. Exit 0 always except usage errors. `/skill:gauntlet-performance` is the reasoning layer on top.
|
|
145
145
|
|
|
146
146
|
For local development against a checkout instead of npm:
|
|
147
147
|
|
|
@@ -1,4 +1,9 @@
|
|
|
1
1
|
#!/usr/bin/env node
|
|
2
|
+
var __defProp = Object.defineProperty;
|
|
3
|
+
var __export = (target, all) => {
|
|
4
|
+
for (var name in all)
|
|
5
|
+
__defProp(target, name, { get: all[name], enumerable: true });
|
|
6
|
+
};
|
|
2
7
|
|
|
3
8
|
// src/bins/gauntlet-performance.mjs
|
|
4
9
|
import { existsSync, readdirSync, readFileSync, realpathSync, statSync } from "node:fs";
|
|
@@ -43,20 +48,227 @@ function resolveTelemetry(g) {
|
|
|
43
48
|
return { enabled, dir, buckets, warning: joinWarn(warnings) };
|
|
44
49
|
}
|
|
45
50
|
|
|
46
|
-
// src/bins/
|
|
51
|
+
// src/bins/performance/shared.mjs
|
|
47
52
|
var PHASES = ["brainstorm", "plan", "implement", "verify", "ship"];
|
|
48
53
|
var TOKEN_KEYS = ["input", "output", "cache_read", "cache_write"];
|
|
49
|
-
var RUN_HEADER = ["run_id", "repo", "spec", "version", "status", "wall", "b/p/i/v/s min", "tokens", "cost", "models", "disp", "grants", "reopens", "loops", "findings", "council"];
|
|
50
|
-
var VERSION_HEADER = ["version", "n", "shipped", "truncated", "wall p50/max", "tokens p50/max", "cost p50/max", "disp p50", "grants p50/max", "reopens p50/max", "loops p50/max", "findings p50 b/M/m", "models"];
|
|
51
|
-
var usage = () => {
|
|
52
|
-
process.stderr.write("usage: gauntlet-performance [--dir <repo root or telemetry dir>]... [--since <version>] [--json]\n");
|
|
53
|
-
process.exit(1);
|
|
54
|
-
};
|
|
55
54
|
var semver = (s) => {
|
|
56
55
|
const m = /^(0|[1-9]\d*)\.(0|[1-9]\d*)\.(0|[1-9]\d*)$/.exec(typeof s === "string" ? s.trim() : "");
|
|
57
56
|
return m ? [Number(m[1]), Number(m[2]), Number(m[3])] : void 0;
|
|
58
57
|
};
|
|
59
58
|
var cmpSemver = (a, b) => a[0] - b[0] || a[1] - b[1] || a[2] - b[2];
|
|
59
|
+
var num = (v) => typeof v === "number" && Number.isFinite(v) ? v : null;
|
|
60
|
+
var str = (v) => typeof v === "string" ? v : null;
|
|
61
|
+
var obj = (v) => v && typeof v === "object" && !Array.isArray(v) ? v : null;
|
|
62
|
+
var sumOrNull = (vals) => {
|
|
63
|
+
const xs = vals.filter((v) => v !== null);
|
|
64
|
+
return xs.length ? xs.reduce((a, b) => a + b, 0) : null;
|
|
65
|
+
};
|
|
66
|
+
var p50 = (xs) => {
|
|
67
|
+
const s = [...xs].sort((a, b) => a - b);
|
|
68
|
+
const m = s.length >> 1;
|
|
69
|
+
return s.length % 2 ? s[m] : (s[m - 1] + s[m]) / 2;
|
|
70
|
+
};
|
|
71
|
+
var stat = (rows, pick) => {
|
|
72
|
+
const xs = rows.map(pick).filter((v) => v !== null && v !== void 0);
|
|
73
|
+
return xs.length ? { p50: p50(xs), max: Math.max(...xs) } : { p50: null, max: null };
|
|
74
|
+
};
|
|
75
|
+
var dash = (v) => v === null || v === void 0 ? "-" : String(v);
|
|
76
|
+
var fmtMin = (s) => s === null ? "-" : `${Math.round(s / 60)}m`;
|
|
77
|
+
var fmtCount = (n) => n === null ? "-" : n >= 1e6 ? `${(n / 1e6).toFixed(1)}M` : n >= 1e3 ? `${(n / 1e3).toFixed(1)}k` : String(n);
|
|
78
|
+
var fmtCost = (c) => c === null ? "-" : c.toFixed(2);
|
|
79
|
+
var pair = (s, f) => `${f(s.p50)}/${f(s.max)}`;
|
|
80
|
+
var findingsCell = (f) => f === null ? "-" : `${dash(f.blocker)}/${dash(f.major)}/${dash(f.minor)}`;
|
|
81
|
+
var table = (header, rows) => {
|
|
82
|
+
const w = header.map((h, i) => Math.max(h.length, ...rows.map((r) => r[i].length)));
|
|
83
|
+
return [header, ...rows].map((r) => r.map((c, i) => c.padEnd(w[i])).join(" ").trimEnd()).join("\n");
|
|
84
|
+
};
|
|
85
|
+
|
|
86
|
+
// src/bins/performance/runs.mjs
|
|
87
|
+
var runs_exports = {};
|
|
88
|
+
__export(runs_exports, {
|
|
89
|
+
aggregate: () => aggregate,
|
|
90
|
+
key: () => key,
|
|
91
|
+
renderText: () => renderText
|
|
92
|
+
});
|
|
93
|
+
var key = "runs";
|
|
94
|
+
var HEADER = ["run_id", "repo", "spec", "version", "status", "wall", "b/p/i/v/s min", "tokens", "cost", "models", "disp", "grants", "reopens", "loops", "findings", "council"];
|
|
95
|
+
var runRow = (r) => [
|
|
96
|
+
r.run_id.slice(0, 8),
|
|
97
|
+
r.repo,
|
|
98
|
+
r.spec,
|
|
99
|
+
r.version,
|
|
100
|
+
`${dash(r.status)}${r.truncated ? "*" : ""}`,
|
|
101
|
+
fmtMin(r.wall_s),
|
|
102
|
+
PHASES.map((p) => r.phase_min[p] === null ? "-" : String(Math.round(r.phase_min[p]))).join("/"),
|
|
103
|
+
fmtCount(r.tokens),
|
|
104
|
+
fmtCost(r.cost),
|
|
105
|
+
r.models.join(",") || "-",
|
|
106
|
+
dash(r.dispatches),
|
|
107
|
+
dash(r.grants),
|
|
108
|
+
dash(r.reopens),
|
|
109
|
+
dash(r.loops),
|
|
110
|
+
findingsCell(r.findings),
|
|
111
|
+
dash(r.council)
|
|
112
|
+
];
|
|
113
|
+
var aggregate = (entries) => entries.map((e) => e.row);
|
|
114
|
+
var renderText = (rows) => [`runs (${rows.length})`, table(HEADER, rows.map(runRow))].join("\n");
|
|
115
|
+
|
|
116
|
+
// src/bins/performance/versions.mjs
|
|
117
|
+
var versions_exports = {};
|
|
118
|
+
__export(versions_exports, {
|
|
119
|
+
aggregate: () => aggregate2,
|
|
120
|
+
key: () => key2,
|
|
121
|
+
renderText: () => renderText2
|
|
122
|
+
});
|
|
123
|
+
var key2 = "by_version";
|
|
124
|
+
var HEADER2 = ["version", "n", "shipped", "truncated", "wall p50/max", "tokens p50/max", "cost p50/max", "disp p50", "grants p50/max", "reopens p50/max", "loops p50/max", "findings p50 b/M/m", "models"];
|
|
125
|
+
var versionOrder = (a, b) => {
|
|
126
|
+
const x = semver(a), y = semver(b);
|
|
127
|
+
return x && y ? cmpSemver(x, y) : x ? -1 : y ? 1 : 0;
|
|
128
|
+
};
|
|
129
|
+
function aggregate2(entries) {
|
|
130
|
+
const rows = entries.map((e) => e.row);
|
|
131
|
+
const groups = /* @__PURE__ */ new Map();
|
|
132
|
+
for (const r of rows) {
|
|
133
|
+
if (!groups.has(r.version)) groups.set(r.version, []);
|
|
134
|
+
groups.get(r.version).push(r);
|
|
135
|
+
}
|
|
136
|
+
return [...groups.entries()].sort(([a], [b]) => versionOrder(a, b)).map(([version, all]) => {
|
|
137
|
+
const shipped = all.filter((r) => r.status === "shipped" && !r.truncated);
|
|
138
|
+
const models = {};
|
|
139
|
+
for (const r of shipped) for (const m of r.models) models[m] = (models[m] ?? 0) + 1;
|
|
140
|
+
return {
|
|
141
|
+
version,
|
|
142
|
+
n: all.length,
|
|
143
|
+
shipped: shipped.length,
|
|
144
|
+
truncated: all.filter((r) => r.truncated).length,
|
|
145
|
+
wall_s: stat(shipped, (r) => r.wall_s),
|
|
146
|
+
tokens: stat(shipped, (r) => r.tokens),
|
|
147
|
+
cost: stat(shipped, (r) => r.cost),
|
|
148
|
+
dispatches: { p50: stat(shipped, (r) => r.dispatches).p50 },
|
|
149
|
+
grants: stat(shipped, (r) => r.grants),
|
|
150
|
+
reopens: stat(shipped, (r) => r.reopens),
|
|
151
|
+
loops: stat(shipped, (r) => r.loops),
|
|
152
|
+
findings: {
|
|
153
|
+
blocker: { p50: stat(shipped, (r) => r.findings?.blocker ?? null).p50 },
|
|
154
|
+
major: { p50: stat(shipped, (r) => r.findings?.major ?? null).p50 },
|
|
155
|
+
minor: { p50: stat(shipped, (r) => r.findings?.minor ?? null).p50 }
|
|
156
|
+
},
|
|
157
|
+
models
|
|
158
|
+
};
|
|
159
|
+
});
|
|
160
|
+
}
|
|
161
|
+
var versionRow = (g) => [
|
|
162
|
+
g.version,
|
|
163
|
+
String(g.n),
|
|
164
|
+
String(g.shipped),
|
|
165
|
+
String(g.truncated),
|
|
166
|
+
pair(g.wall_s, fmtMin),
|
|
167
|
+
pair(g.tokens, fmtCount),
|
|
168
|
+
pair(g.cost, fmtCost),
|
|
169
|
+
dash(g.dispatches.p50),
|
|
170
|
+
pair(g.grants, dash),
|
|
171
|
+
pair(g.reopens, dash),
|
|
172
|
+
pair(g.loops, dash),
|
|
173
|
+
`${dash(g.findings.blocker.p50)}/${dash(g.findings.major.p50)}/${dash(g.findings.minor.p50)}`,
|
|
174
|
+
Object.entries(g.models).map(([m, n]) => `${m}:${n}`).join(",") || "-"
|
|
175
|
+
];
|
|
176
|
+
var renderText2 = (agg) => ["by version", table(HEADER2, agg.map(versionRow))].join("\n");
|
|
177
|
+
|
|
178
|
+
// src/bins/performance/council.mjs
|
|
179
|
+
var council_exports = {};
|
|
180
|
+
__export(council_exports, {
|
|
181
|
+
aggregate: () => aggregate3,
|
|
182
|
+
key: () => key3,
|
|
183
|
+
renderText: () => renderText3
|
|
184
|
+
});
|
|
185
|
+
var key3 = "council";
|
|
186
|
+
var SEVERITIES = ["blocker", "major", "minor"];
|
|
187
|
+
var COUNTERS = ["total", "unique", "applied", "unique_applied", "deferred", "rejected"];
|
|
188
|
+
var MIN_RUNS = 5;
|
|
189
|
+
var HEADER3 = ["member", "total", "unique", "applied", "uniq_appl", "deferred", "rejected", "uniq_appl_nonminor/run"];
|
|
190
|
+
var zero = () => ({ blocker: 0, major: 0, minor: 0 });
|
|
191
|
+
var addInto = (acc, src) => {
|
|
192
|
+
for (const s of SEVERITIES) acc[s] += num(obj(src)?.[s]) ?? 0;
|
|
193
|
+
};
|
|
194
|
+
var sum = (c) => c.blocker + c.major + c.minor;
|
|
195
|
+
var cell = (c) => `${c.blocker}/${c.major}/${c.minor}`;
|
|
196
|
+
var pct = (x) => `${Math.round(x * 100)}%`;
|
|
197
|
+
function flagsFor(members, runs) {
|
|
198
|
+
const share = (m) => sum(m.total) === 0 ? null : sum(m.rejected) / sum(m.total);
|
|
199
|
+
const quiet = Object.entries(members).filter(([, m]) => share(m) !== null && share(m) < 0.25);
|
|
200
|
+
const out = [];
|
|
201
|
+
for (const [name, m] of Object.entries(members)) {
|
|
202
|
+
if (m.unique_applied_nonminor_per_run < 0.2) {
|
|
203
|
+
out.push({ member: name, rule: "low_unique_applied", detail: `${m.unique_applied_nonminor_per_run.toFixed(2)} unique-and-applied blocker/major per run in ${runs} runs (total ${cell(m.total)})` });
|
|
204
|
+
}
|
|
205
|
+
const s = share(m);
|
|
206
|
+
const other = quiet[0];
|
|
207
|
+
if (s !== null && s > 0.5 && other) {
|
|
208
|
+
out.push({ member: name, rule: "high_rejection", detail: `rejected ${pct(s)} vs ${other[0]} ${pct(share(other[1]))} (total ${cell(m.total)})` });
|
|
209
|
+
}
|
|
210
|
+
}
|
|
211
|
+
return out;
|
|
212
|
+
}
|
|
213
|
+
function aggregate3(entries) {
|
|
214
|
+
const groups = /* @__PURE__ */ new Map();
|
|
215
|
+
for (const [i, e] of entries.entries()) {
|
|
216
|
+
const members = obj(obj(e.council)?.members);
|
|
217
|
+
if (!members) continue;
|
|
218
|
+
const roster = Object.keys(members).sort();
|
|
219
|
+
const k = roster.join(", ");
|
|
220
|
+
if (!groups.has(k)) groups.set(k, { roster, entries: [] });
|
|
221
|
+
groups.get(k).entries.push(e);
|
|
222
|
+
groups.get(k).last = i;
|
|
223
|
+
}
|
|
224
|
+
return [...groups.values()].sort((a, b) => b.last - a.last).map((g) => {
|
|
225
|
+
const runs = g.entries.length;
|
|
226
|
+
const members = {};
|
|
227
|
+
for (const m of g.roster) members[m] = { dispatches: 0, ...Object.fromEntries(COUNTERS.map((c) => [c, zero()])) };
|
|
228
|
+
const chairs = /* @__PURE__ */ new Map();
|
|
229
|
+
for (const e of g.entries) {
|
|
230
|
+
for (const m of g.roster) {
|
|
231
|
+
const src = obj(e.council.members[m]);
|
|
232
|
+
members[m].dispatches += num(src?.dispatches) ?? 0;
|
|
233
|
+
for (const c2 of COUNTERS) addInto(members[m][c2], src?.[c2]);
|
|
234
|
+
}
|
|
235
|
+
const ch = obj(e.council.chair);
|
|
236
|
+
const model = str(ch?.model) ?? "unknown";
|
|
237
|
+
const c = chairs.get(model) ?? { model, runs: 0, clusters: 0, reported: 0, retried_runs: 0 };
|
|
238
|
+
c.runs += 1;
|
|
239
|
+
c.clusters += num(ch?.clusters) ?? 0;
|
|
240
|
+
c.reported += num(ch?.members_reported) ?? 0;
|
|
241
|
+
if ((num(ch?.dispatches) ?? 1) > 1) c.retried_runs += 1;
|
|
242
|
+
chairs.set(model, c);
|
|
243
|
+
}
|
|
244
|
+
for (const m of g.roster) members[m].unique_applied_nonminor_per_run = (members[m].unique_applied.blocker + members[m].unique_applied.major) / runs;
|
|
245
|
+
const chairList = [...chairs.values()].sort((a, b) => b.runs - a.runs || a.model.localeCompare(b.model)).map((c) => ({ model: c.model, runs: c.runs, avg_clusters: c.clusters / c.runs, avg_members_reported: c.reported / c.runs, retried_runs: c.retried_runs }));
|
|
246
|
+
return { roster: g.roster, runs, members, chairs: chairList, flags: runs >= MIN_RUNS ? flagsFor(members, runs) : [] };
|
|
247
|
+
});
|
|
248
|
+
}
|
|
249
|
+
function renderText3(agg) {
|
|
250
|
+
if (!agg.length) return "council\nno council data in selected records";
|
|
251
|
+
const out = ["council", "severity is the chair's consolidated severity (highest any co-raiser assigned), not the member's own"];
|
|
252
|
+
for (const g of agg) {
|
|
253
|
+
out.push(`roster (${g.runs} runs): ${g.roster.join(", ")}`);
|
|
254
|
+
const rows = g.roster.map((m) => {
|
|
255
|
+
const x = g.members[m];
|
|
256
|
+
return [m, cell(x.total), cell(x.unique), cell(x.applied), cell(x.unique_applied), cell(x.deferred), cell(x.rejected), x.unique_applied_nonminor_per_run.toFixed(2)];
|
|
257
|
+
});
|
|
258
|
+
for (const line of table(HEADER3, rows).split("\n")) out.push(` ${line}`);
|
|
259
|
+
for (const c of g.chairs) out.push(` chair: ${c.model} - ${c.runs} runs, avg ${c.avg_clusters.toFixed(1)} clusters from ${c.avg_members_reported.toFixed(1)} members reported, retried in ${c.retried_runs}`);
|
|
260
|
+
if (g.runs < MIN_RUNS) out.push(` not enough runs to assess (${g.runs} of ${MIN_RUNS})`);
|
|
261
|
+
for (const f of g.flags) out.push(` flag: ${f.member} - ${f.detail}`);
|
|
262
|
+
}
|
|
263
|
+
return out.join("\n");
|
|
264
|
+
}
|
|
265
|
+
|
|
266
|
+
// src/bins/gauntlet-performance.mjs
|
|
267
|
+
var SECTIONS = [runs_exports, versions_exports, council_exports];
|
|
268
|
+
var usage = () => {
|
|
269
|
+
process.stderr.write("usage: gauntlet-performance [--dir <repo root or telemetry dir>]... [--since <version>] [--json]\n");
|
|
270
|
+
process.exit(1);
|
|
271
|
+
};
|
|
60
272
|
function parseArgs(argv) {
|
|
61
273
|
const opts = { dirs: [], since: void 0, json: false };
|
|
62
274
|
for (let i = 0; i < argv.length; i++) {
|
|
@@ -112,13 +324,6 @@ function* yamlFiles(dir) {
|
|
|
112
324
|
else if (e.isFile() && e.name.endsWith(".yaml")) yield p;
|
|
113
325
|
}
|
|
114
326
|
}
|
|
115
|
-
var num = (v) => typeof v === "number" && Number.isFinite(v) ? v : null;
|
|
116
|
-
var str = (v) => typeof v === "string" ? v : null;
|
|
117
|
-
var obj = (v) => v && typeof v === "object" && !Array.isArray(v) ? v : null;
|
|
118
|
-
var sumOrNull = (vals) => {
|
|
119
|
-
const xs = vals.filter((v) => v !== null);
|
|
120
|
-
return xs.length ? xs.reduce((a, b) => a + b, 0) : null;
|
|
121
|
-
};
|
|
122
327
|
function loadRecord(file, label) {
|
|
123
328
|
let doc;
|
|
124
329
|
try {
|
|
@@ -179,96 +384,10 @@ function loadRecord(file, label) {
|
|
|
179
384
|
council: num(obj(personas["spec-council-member"])?.dispatches),
|
|
180
385
|
ship_option: str(gates.ship_option),
|
|
181
386
|
tests: str(obj(derived.tests)?.result)
|
|
182
|
-
}
|
|
387
|
+
},
|
|
388
|
+
council: derived.council ?? null
|
|
183
389
|
};
|
|
184
390
|
}
|
|
185
|
-
var p50 = (xs) => {
|
|
186
|
-
const s = [...xs].sort((a, b) => a - b);
|
|
187
|
-
const m = s.length >> 1;
|
|
188
|
-
return s.length % 2 ? s[m] : (s[m - 1] + s[m]) / 2;
|
|
189
|
-
};
|
|
190
|
-
var stat = (rows, pick) => {
|
|
191
|
-
const xs = rows.map(pick).filter((v) => v !== null && v !== void 0);
|
|
192
|
-
return xs.length ? { p50: p50(xs), max: Math.max(...xs) } : { p50: null, max: null };
|
|
193
|
-
};
|
|
194
|
-
var versionOrder = (a, b) => {
|
|
195
|
-
const x = semver(a), y = semver(b);
|
|
196
|
-
return x && y ? cmpSemver(x, y) : x ? -1 : y ? 1 : 0;
|
|
197
|
-
};
|
|
198
|
-
function aggregate(rows) {
|
|
199
|
-
const groups = /* @__PURE__ */ new Map();
|
|
200
|
-
for (const r of rows) {
|
|
201
|
-
if (!groups.has(r.version)) groups.set(r.version, []);
|
|
202
|
-
groups.get(r.version).push(r);
|
|
203
|
-
}
|
|
204
|
-
return [...groups.entries()].sort(([a], [b]) => versionOrder(a, b)).map(([version, all]) => {
|
|
205
|
-
const shipped = all.filter((r) => r.status === "shipped" && !r.truncated);
|
|
206
|
-
const models = {};
|
|
207
|
-
for (const r of shipped) for (const m of r.models) models[m] = (models[m] ?? 0) + 1;
|
|
208
|
-
return {
|
|
209
|
-
version,
|
|
210
|
-
n: all.length,
|
|
211
|
-
shipped: shipped.length,
|
|
212
|
-
truncated: all.filter((r) => r.truncated).length,
|
|
213
|
-
wall_s: stat(shipped, (r) => r.wall_s),
|
|
214
|
-
tokens: stat(shipped, (r) => r.tokens),
|
|
215
|
-
cost: stat(shipped, (r) => r.cost),
|
|
216
|
-
dispatches: { p50: stat(shipped, (r) => r.dispatches).p50 },
|
|
217
|
-
grants: stat(shipped, (r) => r.grants),
|
|
218
|
-
reopens: stat(shipped, (r) => r.reopens),
|
|
219
|
-
loops: stat(shipped, (r) => r.loops),
|
|
220
|
-
findings: {
|
|
221
|
-
blocker: { p50: stat(shipped, (r) => r.findings?.blocker ?? null).p50 },
|
|
222
|
-
major: { p50: stat(shipped, (r) => r.findings?.major ?? null).p50 },
|
|
223
|
-
minor: { p50: stat(shipped, (r) => r.findings?.minor ?? null).p50 }
|
|
224
|
-
},
|
|
225
|
-
models
|
|
226
|
-
};
|
|
227
|
-
});
|
|
228
|
-
}
|
|
229
|
-
var dash = (v) => v === null || v === void 0 ? "-" : String(v);
|
|
230
|
-
var fmtMin = (s) => s === null ? "-" : `${Math.round(s / 60)}m`;
|
|
231
|
-
var fmtCount = (n) => n === null ? "-" : n >= 1e6 ? `${(n / 1e6).toFixed(1)}M` : n >= 1e3 ? `${(n / 1e3).toFixed(1)}k` : String(n);
|
|
232
|
-
var fmtCost = (c) => c === null ? "-" : c.toFixed(2);
|
|
233
|
-
var pair = (s, f) => `${f(s.p50)}/${f(s.max)}`;
|
|
234
|
-
var findingsCell = (f) => f === null ? "-" : `${dash(f.blocker)}/${dash(f.major)}/${dash(f.minor)}`;
|
|
235
|
-
var table = (header, rows) => {
|
|
236
|
-
const w = header.map((h, i) => Math.max(h.length, ...rows.map((r) => r[i].length)));
|
|
237
|
-
return [header, ...rows].map((r) => r.map((c, i) => c.padEnd(w[i])).join(" ").trimEnd()).join("\n");
|
|
238
|
-
};
|
|
239
|
-
var runRow = (r) => [
|
|
240
|
-
r.run_id.slice(0, 8),
|
|
241
|
-
r.repo,
|
|
242
|
-
r.spec,
|
|
243
|
-
r.version,
|
|
244
|
-
`${dash(r.status)}${r.truncated ? "*" : ""}`,
|
|
245
|
-
fmtMin(r.wall_s),
|
|
246
|
-
PHASES.map((p) => r.phase_min[p] === null ? "-" : String(Math.round(r.phase_min[p]))).join("/"),
|
|
247
|
-
fmtCount(r.tokens),
|
|
248
|
-
fmtCost(r.cost),
|
|
249
|
-
r.models.join(",") || "-",
|
|
250
|
-
dash(r.dispatches),
|
|
251
|
-
dash(r.grants),
|
|
252
|
-
dash(r.reopens),
|
|
253
|
-
dash(r.loops),
|
|
254
|
-
findingsCell(r.findings),
|
|
255
|
-
dash(r.council)
|
|
256
|
-
];
|
|
257
|
-
var versionRow = (g) => [
|
|
258
|
-
g.version,
|
|
259
|
-
String(g.n),
|
|
260
|
-
String(g.shipped),
|
|
261
|
-
String(g.truncated),
|
|
262
|
-
pair(g.wall_s, fmtMin),
|
|
263
|
-
pair(g.tokens, fmtCount),
|
|
264
|
-
pair(g.cost, fmtCost),
|
|
265
|
-
dash(g.dispatches.p50),
|
|
266
|
-
pair(g.grants, dash),
|
|
267
|
-
pair(g.reopens, dash),
|
|
268
|
-
pair(g.loops, dash),
|
|
269
|
-
`${dash(g.findings.blocker.p50)}/${dash(g.findings.major.p50)}/${dash(g.findings.minor.p50)}`,
|
|
270
|
-
Object.entries(g.models).map(([m, n]) => `${m}:${n}`).join(",") || "-"
|
|
271
|
-
];
|
|
272
391
|
function main() {
|
|
273
392
|
const opts = parseArgs(process.argv.slice(2));
|
|
274
393
|
const corpora = [];
|
|
@@ -285,7 +404,7 @@ function main() {
|
|
|
285
404
|
const since = opts.since === void 0 ? void 0 : semver(opts.since);
|
|
286
405
|
const seen = /* @__PURE__ */ new Set();
|
|
287
406
|
const corpus = {};
|
|
288
|
-
const
|
|
407
|
+
const entries = [];
|
|
289
408
|
for (const c of corpora) {
|
|
290
409
|
corpus[c.label] = corpus[c.label] ?? 0;
|
|
291
410
|
for (const file of yamlFiles(c.dir)) {
|
|
@@ -302,26 +421,25 @@ function main() {
|
|
|
302
421
|
const v = semver(res.row.version);
|
|
303
422
|
if (since && (!v || cmpSemver(v, since) < 0)) continue;
|
|
304
423
|
corpus[c.label]++;
|
|
305
|
-
|
|
424
|
+
entries.push(res);
|
|
306
425
|
}
|
|
307
426
|
}
|
|
308
|
-
|
|
309
|
-
const
|
|
427
|
+
entries.sort((a, b) => (a.row.created_at ?? "").localeCompare(b.row.created_at ?? "") || a.row.run_id.localeCompare(b.row.run_id));
|
|
428
|
+
const aggregates = SECTIONS.map((s) => [s, s.aggregate(entries)]);
|
|
310
429
|
if (opts.json) {
|
|
311
|
-
|
|
430
|
+
const out = { corpus, since: opts.since ?? null };
|
|
431
|
+
for (const [s, agg] of aggregates) out[s.key] = agg;
|
|
432
|
+
out.skipped = skipped;
|
|
433
|
+
console.log(JSON.stringify(out, null, 2));
|
|
312
434
|
return;
|
|
313
435
|
}
|
|
314
436
|
console.log(`corpus: ${Object.entries(corpus).map(([k, v]) => `${k}=${v}`).join(", ")}${opts.since === void 0 ? "" : ` since: ${opts.since}`}`);
|
|
315
|
-
if (
|
|
437
|
+
if (entries.length === 0) {
|
|
316
438
|
console.log(`no records found in ${corpora.map((c) => c.dir).join(", ") || process.cwd()}`);
|
|
317
439
|
} else {
|
|
318
|
-
console.log(
|
|
319
|
-
console.log(table(RUN_HEADER, rows.map(runRow)));
|
|
320
|
-
console.log("");
|
|
321
|
-
console.log("by version");
|
|
322
|
-
console.log(table(VERSION_HEADER, byVersion.map(versionRow)));
|
|
440
|
+
console.log(aggregates.map(([s, agg]) => s.renderText(agg)).join("\n\n"));
|
|
323
441
|
}
|
|
324
|
-
if (
|
|
325
|
-
if (
|
|
442
|
+
if (entries.length && skipped.length) console.log("");
|
|
443
|
+
if (entries.length) for (const s of skipped) console.log(`skipped: ${s.file}: ${s.reason}`);
|
|
326
444
|
}
|
|
327
445
|
main();
|
|
@@ -348,3 +348,103 @@ test("usage error exits 1", (t) => {
|
|
|
348
348
|
assert.equal(invoke(root, ["--since", "abc"]).status, 1);
|
|
349
349
|
assert.equal(invoke(root, ["--since", "5.09.1"]).status, 1);
|
|
350
350
|
});
|
|
351
|
+
|
|
352
|
+
const sev = (b, M, m) => `{ blocker: ${b}, major: ${M}, minor: ${m} }`;
|
|
353
|
+
const member = (dispatches, total, unique, applied, uniqApplied, deferred, rejected) =>
|
|
354
|
+
`{ dispatches: ${dispatches}, total: ${total}, unique: ${unique}, applied: ${applied}, unique_applied: ${uniqApplied}, deferred: ${deferred}, rejected: ${rejected} }`;
|
|
355
|
+
const ALPHA = member(1, sev(1, 3, 1), sev(0, 1, 1), sev(1, 3, 0), sev(0, 1, 0), sev(0, 0, 0), sev(0, 0, 1));
|
|
356
|
+
const BETA = member(1, sev(0, 2, 2), sev(0, 1, 0), sev(0, 1, 1), sev(0, 1, 0), sev(0, 1, 0), sev(0, 0, 1));
|
|
357
|
+
const GAMMA = member(2, sev(0, 1, 4), sev(0, 0, 3), sev(0, 0, 1), sev(0, 0, 1), sev(0, 0, 0), sev(0, 1, 3));
|
|
358
|
+
const councilRecord = (id, createdAt, version, chair, chairDispatches, members) => `schema: 1
|
|
359
|
+
spec: doc/specs/c-${id}.md
|
|
360
|
+
run_id: ${id}-0000-0000-0000-000000000000
|
|
361
|
+
status: shipped
|
|
362
|
+
created_at: ${createdAt}
|
|
363
|
+
shipped_at: ${createdAt}
|
|
364
|
+
versions: { pi-gauntlet: ${version} }
|
|
365
|
+
derived:
|
|
366
|
+
duration_s: 60
|
|
367
|
+
phases:
|
|
368
|
+
ship: ${phase("model: m/alpha")}
|
|
369
|
+
personas: {}
|
|
370
|
+
conformance_loops: 0
|
|
371
|
+
gates: { spec_rounds: 1, plan_rounds: 1, fix_round_grants: 0, task_reopens: 0 }
|
|
372
|
+
council:
|
|
373
|
+
chair: { model: ${chair}, dispatches: ${chairDispatches}, clusters: 6, members_reported: 3 }
|
|
374
|
+
members:
|
|
375
|
+
${Object.entries(members).map(([k, v]) => ` ${k}: ${v}`).join("\n")}
|
|
376
|
+
amendments: 0
|
|
377
|
+
spec_edits_after_ship: 0
|
|
378
|
+
spec_writes: {}
|
|
379
|
+
events_dropped: 0
|
|
380
|
+
accumulators: {}
|
|
381
|
+
events: []
|
|
382
|
+
`;
|
|
383
|
+
const ZERO = member(1, ...Array(6).fill(sev(0, 0, 0)));
|
|
384
|
+
const BIG = { "p/alpha:xhigh": ALPHA, "p/beta:high": BETA, "p/gamma:high": GAMMA, "p/zero:high": ZERO };
|
|
385
|
+
const SMALL = { "p/alpha:xhigh": ALPHA, "p/delta:high": BETA };
|
|
386
|
+
const councilCorpus = (root) => {
|
|
387
|
+
for (let i = 0; i < 6; i++) write(root, `${TDIR}/big-${i}.yaml`, councilRecord(`c000000${i}`, `2026-09-1${i}T00:00:00Z`, "5.20.0", i < 4 ? "p/chair:medium" : "p/other:high", i === 2 ? 2 : 1, BIG));
|
|
388
|
+
for (let i = 0; i < 2; i++) write(root, `${TDIR}/small-${i}.yaml`, councilRecord(`d000000${i}`, `2026-09-2${i}T00:00:00Z`, "5.21.0", "p/chair:medium", 1, SMALL));
|
|
389
|
+
write(root, `${TDIR}/plain.yaml`, COMPLETE);
|
|
390
|
+
return root;
|
|
391
|
+
};
|
|
392
|
+
|
|
393
|
+
test("council section: rosters, sums, chairs and flags", (t) => {
|
|
394
|
+
const root = councilCorpus(gitRepo()); cleanup(t, root);
|
|
395
|
+
const d = json(root);
|
|
396
|
+
assert.equal(d.council.length, 2);
|
|
397
|
+
assert.deepEqual(d.council[0].roster, ["p/alpha:xhigh", "p/delta:high"]);
|
|
398
|
+
assert.equal(d.council[0].runs, 2);
|
|
399
|
+
assert.deepEqual(d.council[0].flags, []);
|
|
400
|
+
const big = d.council[1];
|
|
401
|
+
assert.deepEqual(big.roster, ["p/alpha:xhigh", "p/beta:high", "p/gamma:high", "p/zero:high"]);
|
|
402
|
+
assert.equal(big.runs, 6);
|
|
403
|
+
assert.deepEqual(big.members["p/alpha:xhigh"].total, { blocker: 6, major: 18, minor: 6 });
|
|
404
|
+
assert.equal(big.members["p/alpha:xhigh"].dispatches, 6);
|
|
405
|
+
assert.equal(big.members["p/gamma:high"].dispatches, 12);
|
|
406
|
+
assert.equal(big.members["p/alpha:xhigh"].unique_applied_nonminor_per_run, 1);
|
|
407
|
+
assert.equal(big.members["p/gamma:high"].unique_applied_nonminor_per_run, 0);
|
|
408
|
+
assert.deepEqual(big.chairs, [
|
|
409
|
+
{ model: "p/chair:medium", runs: 4, avg_clusters: 6, avg_members_reported: 3, retried_runs: 1 },
|
|
410
|
+
{ model: "p/other:high", runs: 2, avg_clusters: 6, avg_members_reported: 3, retried_runs: 0 },
|
|
411
|
+
]);
|
|
412
|
+
assert.deepEqual(big.flags.map((f) => [f.member, f.rule]), [
|
|
413
|
+
["p/gamma:high", "low_unique_applied"], ["p/gamma:high", "high_rejection"],
|
|
414
|
+
["p/zero:high", "low_unique_applied"],
|
|
415
|
+
]);
|
|
416
|
+
assert.equal(big.flags[0].detail, "0.00 unique-and-applied blocker/major per run in 6 runs (total 0/6/24)");
|
|
417
|
+
assert.equal(big.flags[1].detail, "rejected 80% vs p/alpha:xhigh 20% (total 0/6/24)");
|
|
418
|
+
assert.equal(big.flags[2].detail, "0.00 unique-and-applied blocker/major per run in 6 runs (total 0/0/0)");
|
|
419
|
+
const r = invoke(root);
|
|
420
|
+
assert.equal(r.status, 0, r.stderr);
|
|
421
|
+
const at = r.lines.indexOf("council");
|
|
422
|
+
assert.ok(at > 0);
|
|
423
|
+
assert.equal(r.lines[at + 1], "severity is the chair's consolidated severity (highest any co-raiser assigned), not the member's own");
|
|
424
|
+
assert.equal(r.lines[at + 2], "roster (2 runs): p/alpha:xhigh, p/delta:high");
|
|
425
|
+
assert.match(r.lines[at + 3], /^ member\s+total\s+unique\s+applied\s+uniq_appl\s+deferred\s+rejected\s+uniq_appl_nonminor\/run$/);
|
|
426
|
+
assert.ok(r.lines.includes(" not enough runs to assess (2 of 5)"));
|
|
427
|
+
assert.ok(r.lines.includes("roster (6 runs): p/alpha:xhigh, p/beta:high, p/gamma:high, p/zero:high"));
|
|
428
|
+
assert.match(r.lines.find((l) => /^ p\/alpha:xhigh\s+6\/18\/6\s/.test(l)), /^ p\/alpha:xhigh\s+6\/18\/6\s+0\/6\/6\s+6\/18\/0\s+0\/6\/0\s+0\/0\/0\s+0\/0\/6\s+1\.00$/);
|
|
429
|
+
assert.ok(r.lines.includes(" chair: p/chair:medium - 4 runs, avg 6.0 clusters from 3.0 members reported, retried in 1"));
|
|
430
|
+
assert.ok(r.lines.includes(" chair: p/other:high - 2 runs, avg 6.0 clusters from 3.0 members reported, retried in 0"));
|
|
431
|
+
assert.ok(r.lines.includes(" flag: p/gamma:high - 0.00 unique-and-applied blocker/major per run in 6 runs (total 0/6/24)"));
|
|
432
|
+
assert.ok(r.lines.includes(" flag: p/gamma:high - rejected 80% vs p/alpha:xhigh 20% (total 0/6/24)"));
|
|
433
|
+
});
|
|
434
|
+
|
|
435
|
+
test("council section: --since and no council data", (t) => {
|
|
436
|
+
const root = councilCorpus(gitRepo()); cleanup(t, root);
|
|
437
|
+
const d = json(root, ["--since", "5.21.0"]);
|
|
438
|
+
assert.equal(d.council.length, 1);
|
|
439
|
+
assert.deepEqual(d.council[0].roster, ["p/alpha:xhigh", "p/delta:high"]);
|
|
440
|
+
const plain = corpus(gitRepo()); cleanup(t, plain);
|
|
441
|
+
assert.deepEqual(json(plain).council, []);
|
|
442
|
+
const r = invoke(plain);
|
|
443
|
+
const at = r.lines.indexOf("council");
|
|
444
|
+
assert.ok(at > 0);
|
|
445
|
+
assert.equal(r.lines[at + 1], "no council data in selected records");
|
|
446
|
+
const empty = gitRepo(); cleanup(t, empty);
|
|
447
|
+
const none = invoke(empty);
|
|
448
|
+
assert.match(none.lines[1], /^no records found/);
|
|
449
|
+
assert.equal(none.lines.length, 2);
|
|
450
|
+
});
|