canary-test-cli 7.1.0 → 8.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (128) hide show
  1. package/agents/skills/README.md +327 -0
  2. package/agents/skills/canary:generate.md +49 -0
  3. package/agents/skills/canary:init.md +37 -0
  4. package/agents/skills/canary:migrate.md +66 -0
  5. package/agents/skills/claude-code/canary-add-framework/SKILL.md +248 -0
  6. package/agents/skills/claude-code/canary-batwoman/SKILL.md +119 -0
  7. package/agents/skills/claude-code/canary-blackhawk/SKILL.md +170 -0
  8. package/agents/skills/claude-code/canary-blackhawk/scripts/cli.mjs +188 -0
  9. package/agents/skills/claude-code/canary-blackhawk/scripts/rules.mjs +120 -0
  10. package/agents/skills/claude-code/canary-blackhawk/scripts/scanner.mjs +244 -0
  11. package/agents/skills/claude-code/canary-blackhawk/scripts/string-literals.mjs +116 -0
  12. package/agents/skills/claude-code/canary-cassandra/SKILL.md +187 -0
  13. package/agents/skills/claude-code/canary-cassandra/scripts/cli.mjs +270 -0
  14. package/agents/skills/claude-code/canary-cassandra/scripts/engine.mjs +95 -0
  15. package/agents/skills/claude-code/canary-ci-ready/SKILL.md +178 -0
  16. package/agents/skills/claude-code/canary-ci-ready/skill.yaml +14 -0
  17. package/agents/skills/claude-code/canary-company-knowledge/SKILL.md +196 -0
  18. package/agents/skills/claude-code/canary-critical-areas/SKILL.md +142 -0
  19. package/agents/skills/claude-code/canary-critical-areas/skill.yaml +16 -0
  20. package/agents/skills/claude-code/canary-edge-case-discovery/SKILL.md +160 -0
  21. package/agents/skills/claude-code/canary-edge-case-discovery/skill.yaml +16 -0
  22. package/agents/skills/claude-code/canary-fail-fast/SKILL.md +75 -0
  23. package/agents/skills/claude-code/canary-fail-fast/scripts/cli.mjs +118 -0
  24. package/agents/skills/claude-code/canary-fail-fast/scripts/digest.mjs +69 -0
  25. package/agents/skills/claude-code/canary-fail-fast/scripts/failures.mjs +60 -0
  26. package/agents/skills/claude-code/canary-fail-fast/scripts/fastfail_check.mjs +43 -0
  27. package/agents/skills/claude-code/canary-fail-fast/scripts/parse.mjs +149 -0
  28. package/agents/skills/claude-code/canary-failure-impact/SKILL.md +153 -0
  29. package/agents/skills/claude-code/canary-failure-impact/skill.yaml +15 -0
  30. package/agents/skills/claude-code/canary-fleet-health/SKILL.md +197 -0
  31. package/agents/skills/claude-code/canary-generate-test/SKILL.md +185 -0
  32. package/agents/skills/claude-code/canary-instrument/SKILL.md +157 -0
  33. package/agents/skills/claude-code/canary-instrument/scripts/cli.mjs +178 -0
  34. package/agents/skills/claude-code/canary-instrument/scripts/otel_bootstrap/instrument.mjs +96 -0
  35. package/agents/skills/claude-code/canary-instrument/scripts/otel_bootstrap/playwright-fixture.ts +44 -0
  36. package/agents/skills/claude-code/canary-instrument/scripts/run_types.mjs +81 -0
  37. package/agents/skills/claude-code/canary-instrument/scripts/span_reader.mjs +187 -0
  38. package/agents/skills/claude-code/canary-katana/SKILL.md +243 -0
  39. package/agents/skills/claude-code/canary-katana/scripts/alarm.mjs +296 -0
  40. package/agents/skills/claude-code/canary-katana/scripts/cli.mjs +247 -0
  41. package/agents/skills/claude-code/canary-katana/scripts/diffscan.mjs +0 -0
  42. package/agents/skills/claude-code/canary-katana/scripts/ledger.mjs +183 -0
  43. package/agents/skills/claude-code/canary-pr-guardian/SKILL.md +144 -0
  44. package/agents/skills/claude-code/canary-pr-guardian/skill.yaml +17 -0
  45. package/agents/skills/claude-code/canary-promote-test/SKILL.md +228 -0
  46. package/agents/skills/claude-code/canary-savant/SKILL.md +233 -0
  47. package/agents/skills/claude-code/canary-savant/scripts/cli.mjs +274 -0
  48. package/agents/skills/claude-code/canary-savant/scripts/restoration.mjs +274 -0
  49. package/agents/skills/claude-code/canary-savant/scripts/rules.mjs +168 -0
  50. package/agents/skills/claude-code/canary-savant/scripts/runner.mjs +572 -0
  51. package/agents/skills/claude-code/canary-savant/scripts/scanner.mjs +374 -0
  52. package/agents/skills/claude-code/canary-savant/scripts/string-literals.mjs +116 -0
  53. package/agents/skills/claude-code/canary-screech/SKILL.md +109 -0
  54. package/agents/skills/claude-code/canary-screech/scripts/blast.mjs +125 -0
  55. package/agents/skills/claude-code/canary-screech/scripts/cli.mjs +128 -0
  56. package/agents/skills/claude-code/canary-screech/scripts/cluster.mjs +97 -0
  57. package/agents/skills/claude-code/canary-screech/scripts/history.mjs +73 -0
  58. package/agents/skills/claude-code/canary-screech/scripts/redness.mjs +94 -0
  59. package/agents/skills/claude-code/canary-setup-harness/SKILL.md +263 -0
  60. package/agents/skills/claude-code/canary-shadow/SKILL.md +131 -0
  61. package/agents/skills/claude-code/canary-shadow/scripts/cases.example.json +32 -0
  62. package/agents/skills/claude-code/canary-shadow/scripts/cli.mjs +195 -0
  63. package/agents/skills/claude-code/canary-ship/SKILL.md +177 -0
  64. package/agents/skills/claude-code/canary-ship/skill.yaml +16 -0
  65. package/agents/skills/claude-code/canary-strix/SKILL.md +130 -0
  66. package/agents/skills/claude-code/canary-strix/scripts/cli.mjs +255 -0
  67. package/agents/skills/claude-code/canary-strix/scripts/scanner.mjs +252 -0
  68. package/agents/skills/claude-code/canary-strix/scripts/terms.mjs +132 -0
  69. package/agents/skills/claude-code/canary-test-pipeline/SKILL.md +159 -0
  70. package/agents/skills/claude-code/canary-test-pipeline/skill.yaml +19 -0
  71. package/agents/skills/claude-code/canary-test-reporter/SKILL.md +138 -0
  72. package/agents/skills/claude-code/canary-test-reporter/scripts/cli.mjs +98 -0
  73. package/agents/skills/claude-code/canary-test-reporter/scripts/json_report.mjs +58 -0
  74. package/agents/skills/claude-code/canary-test-reporter/scripts/parse.mjs +216 -0
  75. package/agents/skills/claude-code/canary-test-reporter/scripts/render.mjs +114 -0
  76. package/agents/skills/lib/parse-args.mjs +275 -0
  77. package/dist/engine/analysis/batwoman/audit.js +39 -0
  78. package/dist/engine/analysis/batwoman/closure.js +159 -0
  79. package/dist/engine/analysis/batwoman/gh-history.js +119 -0
  80. package/dist/engine/analysis/batwoman/probes.js +195 -0
  81. package/dist/engine/analysis/batwoman/registry.js +142 -0
  82. package/dist/engine/analysis/batwoman/render.js +194 -0
  83. package/dist/engine/analysis/batwoman/run-window.js +122 -0
  84. package/dist/engine/analysis/batwoman/text.js +84 -0
  85. package/dist/engine/analysis/batwoman/triggers.js +122 -0
  86. package/dist/engine/analysis/batwoman/verdict.js +64 -0
  87. package/dist/engine/analysis/cli.js +47 -14
  88. package/dist/engine/analysis/gh-flaky/gh-run-attempts.js +206 -0
  89. package/dist/engine/batwoman-cli.js +119 -0
  90. package/dist/engine/ci-ready-cli.js +71 -0
  91. package/dist/engine/cli-commands.js +49 -72
  92. package/dist/engine/cli.core.js +16 -0
  93. package/dist/engine/company-knowledge-cli.js +10 -2
  94. package/dist/engine/core/ci-ready.js +112 -0
  95. package/dist/engine/core/company-knowledge.js +8 -0
  96. package/dist/engine/core/migrator.js +147 -20
  97. package/dist/engine/core/permission-matrix.js +219 -0
  98. package/dist/engine/core/quality-scorer.js +27 -19
  99. package/dist/engine/core/scaling-curve.js +143 -0
  100. package/dist/engine/core/skill-dispatch.js +115 -0
  101. package/dist/engine/core/skill-examples.js +103 -3
  102. package/dist/engine/core/skill-registry.js +59 -4
  103. package/dist/engine/core/string-literals.js +3 -1
  104. package/dist/engine/core/test-files.js +77 -0
  105. package/dist/engine/core/vacuity-scanner.js +330 -15
  106. package/dist/engine/core/workflow-discovery.js +41 -23
  107. package/dist/engine/guardian/adjudication-github.js +136 -0
  108. package/dist/engine/guardian/adjudication.js +119 -340
  109. package/dist/engine/guardian/analysis-emit.js +7 -2
  110. package/dist/engine/guardian/cli.js +277 -249
  111. package/dist/engine/guardian/coverage.js +2 -1
  112. package/dist/engine/guardian/diff-coverage/coverage-delta.js +162 -0
  113. package/dist/engine/guardian/diff-coverage/formats/cobertura.js +45 -1
  114. package/dist/engine/guardian/diff-coverage/orchestrator.js +25 -21
  115. package/dist/engine/guardian/diff-coverage/paths.js +5 -9
  116. package/dist/engine/guardian/diff-coverage/report-tier.js +88 -12
  117. package/dist/engine/guardian/diff-extractor.js +31 -32
  118. package/dist/engine/guardian/pr-check.js +354 -223
  119. package/dist/engine/guardian/pr-comment.js +35 -58
  120. package/dist/engine/guardian/weak-test.js +236 -0
  121. package/dist/engine/mcp-server.js +67 -4
  122. package/dist/engine/permission-matrix-cli.js +51 -0
  123. package/dist/engine/scaling-curve-cli.js +147 -0
  124. package/dist/engine/skills-cli.js +171 -51
  125. package/dist/engine/workflow-cli.js +85 -65
  126. package/dist/reporters/testtracker.d.ts +1 -1
  127. package/dist/reporters/testtracker.js +1 -1
  128. package/package.json +3 -2
@@ -0,0 +1,183 @@
1
+ // ledger -- the quarantine record for tests that are out of the suite.
2
+ //
3
+ // Every captured deletion is written with its provenance (who, when, what
4
+ // commit, why) so a test that vanishes leaves a trail instead of a silent gap.
5
+ // The ledger is de-duplicated: re-running the capture on the same change adds
6
+ // nothing, and a batch of new entries is sorted for a stable on-disk order.
7
+ //
8
+ // SCHEMA v2 (#771) adds the WHY a row is out, alongside the provenance of how it
9
+ // left: `cause`, `issue`, `expiry`. Two fields deserve their separation --
10
+ // `reason` is DERIVED (the commit subject, "what change did this"), `cause` is
11
+ // ASSERTED (a judgement someone made, "why is it out"). Collapsing them would
12
+ // dress an auto-derived string up as a claim a person stands behind.
13
+ //
14
+ // v2 also makes the ledger no longer purely append-only, for one narrow case
15
+ // documented at `appendEntries`: a row that states a cause SUPERSEDES a causeless
16
+ // row for the same (test, file). See the comment there for why leaving both is
17
+ // worse than replacing one.
18
+
19
+ import fs from 'node:fs';
20
+ import path from 'node:path';
21
+
22
+ export const SCHEMA_VERSION = 2;
23
+
24
+ // The fields, in the order they define a row's identity for de-duplication.
25
+ const FIELDS = [
26
+ 'test',
27
+ 'file',
28
+ 'kind',
29
+ 'marker',
30
+ 'commit',
31
+ 'author',
32
+ 'date',
33
+ 'reason',
34
+ 'cause',
35
+ 'issue',
36
+ 'expiry',
37
+ ];
38
+
39
+ /**
40
+ * Why a test is out of the suite. The two middle values are the ones where the
41
+ * TEST IS CORRECT and someone else owns the fix, so a producer must require an
42
+ * `issue` for them -- an untracked "the product is broken" note decays into an
43
+ * unexplained skip within a release.
44
+ */
45
+ export const CAUSES = ['flaky', 'product-defect', 'blocked-data', 'obsolete'];
46
+
47
+ /** Causes for which a row without an `issue` is not a real record. */
48
+ export const CAUSES_REQUIRING_ISSUE = ['product-defect', 'blocked-data'];
49
+
50
+ /**
51
+ * @typedef {{test: string, file: string, kind: string, marker: string,
52
+ * commit: string, author: string, date: string, reason: string,
53
+ * cause: string, issue: string, expiry: string}} LedgerRow
54
+ */
55
+
56
+ /** Normalize an entry-like object into a row with the canonical field order. */
57
+ export function LedgerEntry(fields) {
58
+ const row = {};
59
+ for (const f of FIELDS) row[f] = fields[f] ?? '';
60
+ // v1 wrote the tracker link as `ticket` (#781). It is the same fact under a
61
+ // different name, so it migrates onto `issue` rather than being dropped or
62
+ // kept alongside -- two fields answering "what is this waiting on" is how a
63
+ // consumer ends up reading the empty one. `Ticket:` survives as the name of
64
+ // the COMMIT TRAILER that populates it, which is a mechanism, not a schema.
65
+ if (!row.issue && typeof fields.ticket === 'string')
66
+ row.issue = fields.ticket;
67
+ return row;
68
+ }
69
+
70
+ // The subset that defines a row's IDENTITY for de-duplication.
71
+ //
72
+ // `issue` and `expiry` are excluded, for the reason #781 excluded `ticket`:
73
+ // identity is what happened -- which test, in which file, muted how, by which
74
+ // commit and why -- and a tracker link is an attribute of that event rather
75
+ // than part of it. Including them would also break the append-only guarantee
76
+ // across this schema change, since rows written before the fields existed key
77
+ // as `''` and the same capture re-run with an issue present would hash
78
+ // differently and append a duplicate of a row already on disk.
79
+ //
80
+ // `cause` IS identity, and deliberately so: `appendEntries` decides supersede
81
+ // on whether a row states one, so a caused and a causeless row for the same
82
+ // test must remain distinguishable here for that decision to have anything to
83
+ // act on.
84
+ const IDENTITY_FIELDS = FIELDS.filter((f) => f !== 'issue' && f !== 'expiry');
85
+
86
+ const key = (row) => IDENTITY_FIELDS.map((f) => row[f] ?? '').join('\u0000');
87
+
88
+ /** The (test, file) pair a supersede decision is made on. */
89
+ const pair = (row) => `${row.test ?? ''}\u0000${row.file ?? ''}`;
90
+
91
+ /**
92
+ * Load the ledger document, or an empty one when the file is absent.
93
+ * Throws on unparseable JSON or a non-object top level -- a corrupt ledger is a
94
+ * hard error the caller must surface, not silently overwrite.
95
+ *
96
+ * Every row is normalized through `LedgerEntry`, so a v1 file read here comes
97
+ * back with the v2 fields present and empty. That is what makes writing
98
+ * `schema_version: 2` honest: the version claims "these rows have these fields",
99
+ * and after normalization they do. Stamping the version over un-migrated rows
100
+ * would make it a promise the file does not keep.
101
+ * @returns {{schema_version: number, entries: LedgerRow[]}}
102
+ */
103
+ export function load(filePath) {
104
+ if (!fs.existsSync(filePath)) {
105
+ return { schema_version: SCHEMA_VERSION, entries: [] };
106
+ }
107
+ let data;
108
+ try {
109
+ data = JSON.parse(fs.readFileSync(filePath, 'utf8'));
110
+ } catch (exc) {
111
+ throw new Error(`ledger is not valid JSON: ${filePath}: ${exc.message}`);
112
+ }
113
+ if (data === null || typeof data !== 'object' || Array.isArray(data)) {
114
+ throw new Error(`ledger top level must be an object: ${filePath}`);
115
+ }
116
+ if (!('schema_version' in data)) data.schema_version = SCHEMA_VERSION;
117
+ if (!('entries' in data)) data.entries = [];
118
+ if (!Array.isArray(data.entries)) {
119
+ throw new Error(`ledger entries must be an array: ${filePath}`);
120
+ }
121
+ data.entries = data.entries.map((e) => LedgerEntry(e));
122
+ return data;
123
+ }
124
+
125
+ /**
126
+ * Append `entries` to the ledger at `filePath` and persist it. New entries are
127
+ * sorted by (file, test) and de-duplicated against the batch and disk.
128
+ *
129
+ * ONE ROW PER (test, file) MAY STATE A CAUSE, and a caused row wins:
130
+ *
131
+ * - a new row WITH a cause replaces a causeless row for the same pair
132
+ * - a new row WITHOUT a cause is dropped when a caused row already exists
133
+ *
134
+ * Without this, katana recording `{kind: 'skipped', cause: ''}` and a quarantine
135
+ * writer recording `{kind: 'skipped', cause: 'product-defect', issue: '#123'}`
136
+ * differ in `key()` and BOTH persist. A consumer that fails on an unlinked
137
+ * quarantine (canary-ci-ready does) then fails on the causeless row while the
138
+ * linked row sits beside it -- the ledger contradicting itself about one test.
139
+ *
140
+ * History is not lost to this: rows that differ in cause-bearing state are the
141
+ * only ones that collapse. Two caused rows, or two causeless rows, keep the
142
+ * full-field identity and both remain.
143
+ */
144
+ export function appendEntries(filePath, entries) {
145
+ const doc = load(filePath);
146
+ const existing = doc.entries;
147
+ const seen = new Set(existing.map(key));
148
+ const causedPairs = new Set(existing.filter((r) => r.cause).map(pair));
149
+
150
+ const newRows = entries
151
+ .map((e) => LedgerEntry(e))
152
+ .sort(
153
+ (a, b) => a.file.localeCompare(b.file) || a.test.localeCompare(b.test),
154
+ );
155
+
156
+ for (const row of newRows) {
157
+ const k = key(row);
158
+ if (seen.has(k)) continue;
159
+ const p = pair(row);
160
+
161
+ if (row.cause) {
162
+ // Supersede: drop any causeless row for this pair, then take its place.
163
+ for (let i = existing.length - 1; i >= 0; i--) {
164
+ if (!existing[i].cause && pair(existing[i]) === p) {
165
+ seen.delete(key(existing[i]));
166
+ existing.splice(i, 1);
167
+ }
168
+ }
169
+ causedPairs.add(p);
170
+ } else if (causedPairs.has(p)) {
171
+ // A causeless row must never sit next to a caused one for the same test.
172
+ continue;
173
+ }
174
+
175
+ seen.add(k);
176
+ existing.push(row);
177
+ }
178
+
179
+ doc.schema_version = SCHEMA_VERSION;
180
+ fs.mkdirSync(path.dirname(path.resolve(filePath)), { recursive: true });
181
+ fs.writeFileSync(filePath, `${JSON.stringify(doc, null, 2)}\n`, 'utf8');
182
+ return doc;
183
+ }
@@ -0,0 +1,144 @@
1
+ ---
2
+ name: canary-pr-guardian
3
+ description: >
4
+ PR/pre-commit test-guardian orchestrator: runs the deterministic Tier-0
5
+ diff-coverage pass, audits affected tests via canary-test-reviewer (Tier 1),
6
+ and — at the desk with authorTests opt-in — authors missing tests via
7
+ canary-test-author (Tier 2), staging them and blocking the commit once for
8
+ human review. Use to guard a change's test quality before it lands.
9
+ ---
10
+
11
+ # Canary: PR Guardian
12
+
13
+ Guards a change's test quality before it lands. Composes the deterministic
14
+ Tier-0 diff-coverage engine with two native agents — `canary-test-reviewer`
15
+ (read-only audit) and `canary-test-author` (authoring) — under a strict
16
+ write-safety model. This is the **Option A** driver: the engine never calls an
17
+ LLM; **this skill** invokes the agents in-session and enforces
18
+ stage-and-block-once.
19
+
20
+ **On the tier numbers.** Here they are the values of
21
+ `canary guardian pr-check --tier 0|1|2`, not a repo-wide capability scale — Tier
22
+ 1 means "the `--tier 1` pass," which is the agent audit. `Tier-0` is the one
23
+ number with a repo-wide meaning (deterministic, no network, no agent), and
24
+ `Tier-1`/`Tier-2` are guardian-local by
25
+ [ADR 0015](../../../../docs/knowledge/decisions/0015-skill-capability-vocabulary.md).
26
+ Do not carry them into other skills.
27
+
28
+ ## When to Use
29
+
30
+ - Before opening or updating a PR, to check that new/changed code is tested.
31
+ - As a pre-commit companion when `preCommit.authorTests: true` is set and you
32
+ want the guardian to author the missing tests for you (at the desk only).
33
+ - NOT in CI for Tier-2 write-back — that is a NON-GOAL. CI runs Tier-0 only (the
34
+ `CANARY_GUARDIAN_AGENT` env is unset there).
35
+
36
+ ## Safety model (non-negotiable)
37
+
38
+ - **NEVER commit or push.** This skill only authors and `git add`s. The human
39
+ reviews the staged tests and re-commits.
40
+ - **Honor every `skipped` reason** from `author-plan` verbatim (opt-in-off /
41
+ tier / fork / collision / loop-guard). Never override a skip.
42
+ - **Block once.** When `block.block == true`, print the block message and stop —
43
+ leave the staged tests for the human. The loop-guard sentinel you write in
44
+ Phase 3 is what stops the guardian re-authoring over its own output on the
45
+ next run. It is stamped with the current `HEAD` and expires by itself once the
46
+ human's review commit moves `HEAD` — never delete it yourself.
47
+ - **Authoring is opt-in.** No `preCommit.authorTests: true` ⇒ no writes, ever.
48
+
49
+ ## Phases
50
+
51
+ ### Phase 0 — Deterministic scope
52
+
53
+ Run the Tier-0 pass and read its findings:
54
+
55
+ ```bash
56
+ canary guardian pr-check --format json
57
+ ```
58
+
59
+ Findings are `untested-new-code` gaps. If there are none, report clean and stop.
60
+
61
+ **Coverage regressions (#606).** `--coverage` alone answers only "is this
62
+ changed unit covered at all?". To also catch a unit whose coverage _fell_
63
+ against the base branch, pass the base ref's report as well:
64
+
65
+ ```bash
66
+ canary guardian pr-check --coverage lcov.info --base-coverage base-lcov.info --format json
67
+ ```
68
+
69
+ That adds `coverage-regression` findings, graded by how many percentage points
70
+ were lost. Most CI never uploads a base-branch artifact; without
71
+ `--base-coverage` the run degrades **loudly** to "delta unavailable — head-only"
72
+ and `coverage_delta.status` reports `unavailable`. Read that field before
73
+ treating an empty finding list as "no regressions" — a run that compared nothing
74
+ has abstained, not passed.
75
+
76
+ ### Phase 1 — Quality audit (Tier ≥ 1, read-only)
77
+
78
+ Export the availability signal so the probe reports the ceiling, then audit the
79
+ affected tests with `canary-test-reviewer`:
80
+
81
+ ```bash
82
+ export CANARY_GUARDIAN_AGENT=1 # or 2 when authoring is enabled
83
+ ```
84
+
85
+ If this checkout is a **fork** (an untrusted/read-only context — detect it, e.g.
86
+ `git config --get remote.origin.url` pointing at a fork, or a CI fork PR), arm
87
+ the fork guard so the engine's safety layer never authors on it:
88
+
89
+ ```bash
90
+ export CANARY_GUARDIAN_IS_FORK=1 # any value other than "0"/unset means fork
91
+ ```
92
+
93
+ Leave `CANARY_GUARDIAN_IS_FORK` unset (or `0`) at your own desk on a trusted
94
+ checkout. The guard fails CLOSED: any ambiguous value is treated as a fork and
95
+ authoring is skipped.
96
+
97
+ Use the `canary-test-reviewer` agent to review the affected tests. This is
98
+ **read-only** — surface weak-test findings; write nothing.
99
+
100
+ ### Phase 2 — Authoring plan (Tier 2 + opt-in)
101
+
102
+ Ask the engine's safety layer for the plan (intents + block decision):
103
+
104
+ ```bash
105
+ canary guardian author-plan --json
106
+ ```
107
+
108
+ The JSON is `{"intents": [...], "block": {...}}`. Each intent carries `status`
109
+ (`planned` | `authored` | `skipped`), `target_path`, `requirement`, and a
110
+ `skip_reason` when skipped. **Do not author anything the plan skipped** — the
111
+ Engine guards (opt-in, fork, collision, loop-guard) are authoritative.
112
+
113
+ ### Phase 3 — Author, stage, block once
114
+
115
+ For each intent with `status: "planned"`, use the `canary-test-author` agent
116
+ with the intent's `requirement` as the task and `target_path` as the
117
+ destination. **Never overwrite an existing file at `target_path`.** The engine's
118
+ collision guard runs at plan time, so between planning and writing another
119
+ PR/session may have created the file (a TOCTOU window). If the target already
120
+ exists at write time, skip that intent and report it — do not clobber it. Then
121
+ stage the authored files:
122
+
123
+ ```bash
124
+ git add <target_path>
125
+ ```
126
+
127
+ When `block.block == true`, record the guardian-authored paths in the loop-guard
128
+ sentinel BEFORE blocking, by running the deterministic producer command once per
129
+ authored path:
130
+
131
+ ```bash
132
+ canary guardian mark-authored --path <target_path> [--path <target_path> ...]
133
+ ```
134
+
135
+ This is a real CLI step (not something you `touch` yourself): it writes the
136
+ sentinel inside the real git dir and records exactly the paths you authored,
137
+ stamped with the current `HEAD`. `author-plan` reads it on the next run and
138
+ returns a `loop-guard` skip **while `HEAD` is unchanged**, so the guardian never
139
+ authors on top of its own output — and once the human commits the reviewed
140
+ tests, `HEAD` moves and authoring re-enables itself.
141
+
142
+ Then print `block.message` (the "N test(s) authored & staged — review and
143
+ re-commit" notice) and **stop**. Do not commit. The human reviews the staged
144
+ tests and re-commits.
@@ -0,0 +1,17 @@
1
+ name: canary-pr-guardian
2
+ version: '1.0.0'
3
+ description:
4
+ PR/pre-commit test-guardian orchestrator — Tier-0 diff-coverage, Tier-1
5
+ quality audit via canary-test-reviewer, and at-desk Tier-2 authoring via
6
+ canary-test-author with stage-and-block-once review.
7
+ stability: static
8
+ triggers:
9
+ - manual
10
+ platforms:
11
+ - claude-code
12
+ type: rigid
13
+ tools: []
14
+ tier: 1
15
+ depends_on:
16
+ - canary-test-reviewer
17
+ - canary-test-author
@@ -0,0 +1,228 @@
1
+ ---
2
+ name: canary-promote-test
3
+ description: >
4
+ Move a generated test from `tests/generated/` into the committed test suite —
5
+ reviews it for correctness, drops the generation header, relocates it to the
6
+ matching suite directory, and confirms it runs in the project's normal test
7
+ flow. Use for "promote this test", "commit this generated test", "move this
8
+ test into the suite", or "keep this test" — always after the generated test
9
+ has been validated against the SUT, never before. Not for tests needing
10
+ substantial rewriting (regenerate instead) or throwaway investigation tests
11
+ (leave in `tests/generated/`).
12
+ ---
13
+
14
+ # Canary: Promote Test
15
+
16
+ > Move a generated test from `tests/generated/` into the committed test suite.
17
+ > Reviews the test for correctness, drops the generation header, relocates it
18
+ > under the appropriate suite directory, and confirms it runs in the project's
19
+ > normal test flow.
20
+
21
+ ## When to Use
22
+
23
+ - After `canary-generate-test` produced a file that has been validated against
24
+ the SUT
25
+ - When a user explicitly asks to "commit", "save", or "keep" a generated test
26
+ - When extending an existing suite with a new case the team has agreed to
27
+ maintain
28
+ - NOT before the generated test has been executed and reviewed — promotion is
29
+ the _last_ step, not the first
30
+ - NOT for tests that need substantial rewriting — regenerate with a better
31
+ prompt instead of hand-patching
32
+ - NOT for tests targeting throwaway investigations (perf spikes, ad-hoc bug
33
+ triage) — leave those in `tests/generated/`
34
+
35
+ ## Process
36
+
37
+ ### Phase 0: GATE — Get the Structured Verdict First
38
+
39
+ Run this before reading the test. It is deterministic, takes no API key, and it
40
+ will refuse the drafts that are not worth your review time.
41
+
42
+ ```bash
43
+ canary promote-check tests/generated/api/orders_post.py --json
44
+ ```
45
+
46
+ | Exit | Verdict | What it means |
47
+ | ---- | --------- | ------------------------------------------------------------ |
48
+ | `0` | `promote` | No gating defect. Advisory findings may remain — your call. |
49
+ | `1` | `block` | A gating defect. **Do not promote.** Regenerate or fix. |
50
+ | `3` | `abstain` | No verdict could be produced. Promotion is **not** approved. |
51
+
52
+ **Which axes gate, and why only those.** Gating on all eight axes of a quality
53
+ critique would block nearly every promotion, so only deterministic defects do:
54
+
55
+ | Axis | Rules | Gates | Reason |
56
+ | ----------------- | ------------------------------ | ----- | ------------------------------------------------------ |
57
+ | `soundness` | `SOUND-001/002/003` | yes | Pins a value no correct implementation must produce |
58
+ | `assertions` | `LINT-006` | yes | A test that asserts nothing always passes |
59
+ | `flakiness` | `FLAKE-001/002` | yes | A hardcoded sleep in a committed suite is a future red |
60
+ | `vacuity` | `VAC-001/003`, annotated `002` | yes | Cannot fail, or contradicts a declared `@covers` |
61
+ | `selectors` | `LINT-001/002/003` | no | Brittle, not wrong — a reviewer's call |
62
+ | `maintainability` | `LINT-005`, `FLAKE-003/004` | no | Style and softer signals |
63
+
64
+ **`VAC-002` gates only at `annotated` fidelity.** At `import-inferred` it is an
65
+ inference about which symbol the test meant to exercise, and a heuristic must
66
+ not be load-bearing on a promotion gate. If you want the gate to check the real
67
+ target, add the annotation to the generated test:
68
+
69
+ ```ts
70
+ // @covers resolveOverlay
71
+ ```
72
+
73
+ **An `abstain` is not a pass.** Exit 3 means the checker had no subject — an
74
+ unparseable extension, or a file with no test declarations. Promotion falls back
75
+ to the manual review below and is **not** approved by silence. Do not read a
76
+ missing verdict as either stricter or looser than one.
77
+
78
+ **An LLM judgement never gates.** `harness:test-craft` runs an 8-axis per-test
79
+ critique and remains exactly what step 5 below calls it: an optional deeper
80
+ audit for a human. Everything that blocks in this repo is deterministic, and
81
+ `promote-check` keeps it that way — the verdict has no field an LLM opinion
82
+ could arrive in.
83
+
84
+ ### Phase 1: REVIEW — Confirm the Test Is Worth Keeping
85
+
86
+ 1. **Read the test end-to-end.** Treat it like any other code review — naming,
87
+ assertions, hardcoded values, missing edge cases. Generated tests are drafts,
88
+ not finished artifacts.
89
+ 2. **Run the test against the real SUT** (not just the env it was generated
90
+ against). A test that only passes in one environment is a fixture-bound test,
91
+ not a regression test.
92
+ 3. **Check assertion strength.** A test that only asserts "status code 200" is
93
+ weak; promote it only if that's genuinely the contract. If the SUT returns
94
+ structured data, the test should assert on shape.
95
+ 4. **Confirm no hardcoded secrets or environment-specific URLs.** Replace with
96
+ fixtures or env-driven config before promoting.
97
+ 5. **Optional: run `harness:test-craft` for a deeper quality audit.**
98
+ `test-craft` runs an 8-axis per-test LLM critique (assertion density,
99
+ flakiness risk, contract vs implementation, etc.). Use it when the generated
100
+ test is substantial or when the team wants a second opinion before
101
+ committing. Not required for simple happy-path tests, and **never a blocker**
102
+ — see Phase 0.
103
+ 6. **Triage the advisory findings** `promote-check` reported. They did not
104
+ block; deciding whether they matter here is the review's job.
105
+ 7. **Decide: promote, regenerate, or discard.** If review reveals more than ~3
106
+ small fixes, regenerate with a sharper prompt instead.
107
+
108
+ ### Phase 2: RELOCATE — Move into the Suite
109
+
110
+ 1. **Identify the destination directory.** Mirror the suite's structure:
111
+ - `tests/generated/api/foo.py` → `tests/api/foo.py`
112
+ - `tests/generated/e2e/checkout.spec.ts` → `tests/e2e/checkout.spec.ts`
113
+ - `tests/generated/unit/validator.py` → `tests/unit/validator.py`
114
+ 2. **Match suite conventions.** Look at neighboring files for:
115
+ - Import style (relative vs absolute)
116
+ - Fixture/setup imports (most suites have a `conftest.py` or shared setup)
117
+ - Naming conventions (`test_<feature>_<case>.py`, `<feature>.spec.ts`)
118
+ 3. **Move with `git mv`** so the history is preserved if anyone later runs
119
+ `git log --follow`.
120
+ 4. **Update imports** if the file referenced anything by relative path from
121
+ `tests/generated/`.
122
+
123
+ ### Phase 3: CLEAN — Drop Generation Artifacts
124
+
125
+ 1. **Remove the timestamped generation header.** The "Generated by Canary on
126
+ [date]" comment is useful in scratch space — meaningless in a committed test
127
+ and rots immediately.
128
+ 2. **Remove any placeholder TODOs.** The generator sometimes emits
129
+ `# TODO: adjust selector` or similar — either resolve them or stop the
130
+ promotion and regenerate.
131
+ 3. **Tighten formatting.** Run the project's formatter (`black`, `prettier`,
132
+ etc.) so the file matches surrounding style.
133
+ 4. **Strip dead code.** Imports the generator added "just in case" but the test
134
+ doesn't use.
135
+
136
+ ### Phase 4: VERIFY — Run in the Project's Normal Flow
137
+
138
+ 1. **Run the suite that owns this test:**
139
+ - Python: `pytest tests/api/test_orders_post.py`
140
+ - Playwright: `npx playwright test tests/e2e/checkout.spec.ts`
141
+ - Match whatever CI runs.
142
+ 2. **Run the full suite** to confirm no collateral failure (shared fixtures,
143
+ port conflicts, ordering issues).
144
+ 3. **Confirm CI configuration picks it up.** If the suite has a glob in CI
145
+ config, verify the new path matches. If not, add it.
146
+ 4. **Log the promotion.** Append a one-line entry to `docs/CANARY_STATE.md` so
147
+ the project ledger tracks which generated tests have been promoted.
148
+
149
+ ## Canary Integration
150
+
151
+ - **`tests/generated/`** — Source for promotion. Gitignored; nothing here is
152
+ ever a final artifact.
153
+ - **`tests/<suite>/`** — Destination. Each suite has its own conventions; never
154
+ invent a new top-level dir during promotion.
155
+ - **`docs/CANARY_STATE.md`** — Append a promotion entry: requirement, generated
156
+ path, promoted path, date.
157
+
158
+ ## Success Criteria
159
+
160
+ - The promoted test passes when run via the project's normal test command
161
+ - The full suite passes (no collateral breakage)
162
+ - The file matches the surrounding code style (formatter clean, lint clean)
163
+ - No generation artifacts remain (timestamp header, placeholder TODOs, unused
164
+ imports)
165
+ - CI picks up the new test on the next push
166
+
167
+ ## Rationalizations to Reject
168
+
169
+ | Rationalization | Why It Is Wrong |
170
+ | ------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
171
+ | "`promote-check` exited 3, so nothing was wrong" | Exit 3 is an abstention — the checker had no subject at all. It is the one outcome that proves nothing, so it can never stand in for approval. |
172
+ | "`promote-check` blocked on VAC-002, but the test is fine" | Only an `annotated` VAC-002 blocks, which means the file's own `@covers` names a symbol the test never touches. Either the annotation is wrong or the test is. |
173
+ | "I'll wire `harness:test-craft` into the gate for better coverage" | An 8-axis LLM critique gating a promotion would block nearly everything and would make a judgement load-bearing. Deterministic verdicts gate; critiques inform. |
174
+ | "The test works, I'll skip the review and just commit it" | Generated code looks plausible but commonly has weak assertions or hardcoded values. Review is the whole point — promotion without review just imports debt into the suite. |
175
+ | "I'll fix the 6 issues I found in the review by editing the test" | Six issues means the prompt was wrong. Regenerate; hand-edits won't transfer to the next similar test. |
176
+ | "I'll leave the timestamp header so we know when it was generated" | Git history already records when the file was committed. The header rots and creates noise. |
177
+ | "It passes against staging, that's good enough" | If the test only passes in one environment, it's a fixture-bound smoke check, not a regression test. Either parametrize the env or don't promote it. |
178
+ | "I'll commit it without running the full suite — I only changed one test" | New tests can break shared fixtures, conflict on ports, leak state. Always run the full suite once before commit. |
179
+
180
+ ## Examples
181
+
182
+ ### Example: Clean promotion
183
+
184
+ **Source:** `tests/generated/api/orders_post_201.py`, validated, single status
185
+ assertion is exactly the contract.
186
+
187
+ **Action:**
188
+
189
+ 1. `git mv tests/generated/api/orders_post_201.py tests/api/test_orders_post_201.py`
190
+ 2. Drop the `# Generated by Canary...` header.
191
+ 3. Adjust import to match `tests/api/conftest.py` fixtures.
192
+ 4. Run `pytest tests/api/` → passes.
193
+ 5. Append to `CANARY_STATE.md`: promoted `orders_post_201.py` →
194
+ `tests/api/test_orders_post_201.py` on [date].
195
+
196
+ ### Example: Promotion abandoned — regenerate instead
197
+
198
+ **Source:** `tests/generated/e2e/checkout.spec.ts`. Review finds: hardcoded test
199
+ user creds, no wait for navigation, weak assertion (`expect(true).toBe(true)`),
200
+ TODO comments left in three places.
201
+
202
+ **Action:** Stop. Regenerate with a sharper prompt that names the real fixture
203
+ user, specifies the wait condition, and asserts the actual checkout success
204
+ state. Do not commit the broken draft.
205
+
206
+ ### Example: Test for throwaway investigation
207
+
208
+ **Source:** `tests/generated/performance/spike_search_50rps.js`. Used once to
209
+ confirm a single perf hypothesis. No ongoing value.
210
+
211
+ **Action:** Do not promote. Leave in `tests/generated/`. If the team wants
212
+ ongoing perf monitoring, generate a _new_ test with a sustainable RPS profile
213
+ and promote that.
214
+
215
+ ## Escalation
216
+
217
+ - **When the review reveals the SUT itself is broken:** Don't promote the test
218
+ that passes against the broken SUT. File a bug; only promote the test once the
219
+ SUT is fixed and the assertion has a real contract behind it.
220
+ - **When the suite has no existing convention to match:** This usually means the
221
+ test belongs in a new sub-suite. Ask the user to confirm the suite structure
222
+ before inventing one.
223
+ - **When CI doesn't pick up the new path:** Don't merge until CI config matches.
224
+ A test that exists but isn't run is worse than no test — it implies coverage
225
+ that doesn't exist.
226
+ - **When the promoted test starts flaking after merge:** Treat as a real
227
+ regression in the test (or the SUT). Don't quarantine in `tests/generated/` —
228
+ fix or delete.