canary-test-cli 7.1.0 → 8.0.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/agents/skills/README.md +327 -0
- package/agents/skills/canary:generate.md +49 -0
- package/agents/skills/canary:init.md +37 -0
- package/agents/skills/canary:migrate.md +66 -0
- package/agents/skills/claude-code/canary-add-framework/SKILL.md +248 -0
- package/agents/skills/claude-code/canary-batwoman/SKILL.md +119 -0
- package/agents/skills/claude-code/canary-blackhawk/SKILL.md +170 -0
- package/agents/skills/claude-code/canary-blackhawk/scripts/cli.mjs +188 -0
- package/agents/skills/claude-code/canary-blackhawk/scripts/rules.mjs +120 -0
- package/agents/skills/claude-code/canary-blackhawk/scripts/scanner.mjs +244 -0
- package/agents/skills/claude-code/canary-blackhawk/scripts/string-literals.mjs +116 -0
- package/agents/skills/claude-code/canary-cassandra/SKILL.md +187 -0
- package/agents/skills/claude-code/canary-cassandra/scripts/cli.mjs +270 -0
- package/agents/skills/claude-code/canary-cassandra/scripts/engine.mjs +95 -0
- package/agents/skills/claude-code/canary-ci-ready/SKILL.md +178 -0
- package/agents/skills/claude-code/canary-ci-ready/skill.yaml +14 -0
- package/agents/skills/claude-code/canary-company-knowledge/SKILL.md +196 -0
- package/agents/skills/claude-code/canary-critical-areas/SKILL.md +142 -0
- package/agents/skills/claude-code/canary-critical-areas/skill.yaml +16 -0
- package/agents/skills/claude-code/canary-edge-case-discovery/SKILL.md +160 -0
- package/agents/skills/claude-code/canary-edge-case-discovery/skill.yaml +16 -0
- package/agents/skills/claude-code/canary-fail-fast/SKILL.md +75 -0
- package/agents/skills/claude-code/canary-fail-fast/scripts/cli.mjs +118 -0
- package/agents/skills/claude-code/canary-fail-fast/scripts/digest.mjs +69 -0
- package/agents/skills/claude-code/canary-fail-fast/scripts/failures.mjs +60 -0
- package/agents/skills/claude-code/canary-fail-fast/scripts/fastfail_check.mjs +43 -0
- package/agents/skills/claude-code/canary-fail-fast/scripts/parse.mjs +149 -0
- package/agents/skills/claude-code/canary-failure-impact/SKILL.md +153 -0
- package/agents/skills/claude-code/canary-failure-impact/skill.yaml +15 -0
- package/agents/skills/claude-code/canary-fleet-health/SKILL.md +197 -0
- package/agents/skills/claude-code/canary-generate-test/SKILL.md +185 -0
- package/agents/skills/claude-code/canary-instrument/SKILL.md +157 -0
- package/agents/skills/claude-code/canary-instrument/scripts/cli.mjs +178 -0
- package/agents/skills/claude-code/canary-instrument/scripts/otel_bootstrap/instrument.mjs +96 -0
- package/agents/skills/claude-code/canary-instrument/scripts/otel_bootstrap/playwright-fixture.ts +44 -0
- package/agents/skills/claude-code/canary-instrument/scripts/run_types.mjs +81 -0
- package/agents/skills/claude-code/canary-instrument/scripts/span_reader.mjs +187 -0
- package/agents/skills/claude-code/canary-katana/SKILL.md +243 -0
- package/agents/skills/claude-code/canary-katana/scripts/alarm.mjs +296 -0
- package/agents/skills/claude-code/canary-katana/scripts/cli.mjs +247 -0
- package/agents/skills/claude-code/canary-katana/scripts/diffscan.mjs +0 -0
- package/agents/skills/claude-code/canary-katana/scripts/ledger.mjs +183 -0
- package/agents/skills/claude-code/canary-pr-guardian/SKILL.md +144 -0
- package/agents/skills/claude-code/canary-pr-guardian/skill.yaml +17 -0
- package/agents/skills/claude-code/canary-promote-test/SKILL.md +228 -0
- package/agents/skills/claude-code/canary-savant/SKILL.md +233 -0
- package/agents/skills/claude-code/canary-savant/scripts/cli.mjs +274 -0
- package/agents/skills/claude-code/canary-savant/scripts/restoration.mjs +274 -0
- package/agents/skills/claude-code/canary-savant/scripts/rules.mjs +168 -0
- package/agents/skills/claude-code/canary-savant/scripts/runner.mjs +572 -0
- package/agents/skills/claude-code/canary-savant/scripts/scanner.mjs +374 -0
- package/agents/skills/claude-code/canary-savant/scripts/string-literals.mjs +116 -0
- package/agents/skills/claude-code/canary-screech/SKILL.md +109 -0
- package/agents/skills/claude-code/canary-screech/scripts/blast.mjs +125 -0
- package/agents/skills/claude-code/canary-screech/scripts/cli.mjs +128 -0
- package/agents/skills/claude-code/canary-screech/scripts/cluster.mjs +97 -0
- package/agents/skills/claude-code/canary-screech/scripts/history.mjs +73 -0
- package/agents/skills/claude-code/canary-screech/scripts/redness.mjs +94 -0
- package/agents/skills/claude-code/canary-setup-harness/SKILL.md +263 -0
- package/agents/skills/claude-code/canary-shadow/SKILL.md +131 -0
- package/agents/skills/claude-code/canary-shadow/scripts/cases.example.json +32 -0
- package/agents/skills/claude-code/canary-shadow/scripts/cli.mjs +195 -0
- package/agents/skills/claude-code/canary-ship/SKILL.md +177 -0
- package/agents/skills/claude-code/canary-ship/skill.yaml +16 -0
- package/agents/skills/claude-code/canary-strix/SKILL.md +130 -0
- package/agents/skills/claude-code/canary-strix/scripts/cli.mjs +255 -0
- package/agents/skills/claude-code/canary-strix/scripts/scanner.mjs +252 -0
- package/agents/skills/claude-code/canary-strix/scripts/terms.mjs +132 -0
- package/agents/skills/claude-code/canary-test-pipeline/SKILL.md +159 -0
- package/agents/skills/claude-code/canary-test-pipeline/skill.yaml +19 -0
- package/agents/skills/claude-code/canary-test-reporter/SKILL.md +138 -0
- package/agents/skills/claude-code/canary-test-reporter/scripts/cli.mjs +98 -0
- package/agents/skills/claude-code/canary-test-reporter/scripts/json_report.mjs +58 -0
- package/agents/skills/claude-code/canary-test-reporter/scripts/parse.mjs +216 -0
- package/agents/skills/claude-code/canary-test-reporter/scripts/render.mjs +114 -0
- package/agents/skills/lib/parse-args.mjs +275 -0
- package/dist/engine/analysis/batwoman/audit.js +39 -0
- package/dist/engine/analysis/batwoman/closure.js +159 -0
- package/dist/engine/analysis/batwoman/gh-history.js +119 -0
- package/dist/engine/analysis/batwoman/probes.js +195 -0
- package/dist/engine/analysis/batwoman/registry.js +142 -0
- package/dist/engine/analysis/batwoman/render.js +194 -0
- package/dist/engine/analysis/batwoman/run-window.js +122 -0
- package/dist/engine/analysis/batwoman/text.js +84 -0
- package/dist/engine/analysis/batwoman/triggers.js +122 -0
- package/dist/engine/analysis/batwoman/verdict.js +64 -0
- package/dist/engine/analysis/cli.js +47 -14
- package/dist/engine/analysis/gh-flaky/gh-run-attempts.js +206 -0
- package/dist/engine/batwoman-cli.js +119 -0
- package/dist/engine/ci-ready-cli.js +71 -0
- package/dist/engine/cli-commands.js +49 -72
- package/dist/engine/cli.core.js +16 -0
- package/dist/engine/company-knowledge-cli.js +10 -2
- package/dist/engine/core/ci-ready.js +112 -0
- package/dist/engine/core/company-knowledge.js +8 -0
- package/dist/engine/core/migrator.js +147 -20
- package/dist/engine/core/permission-matrix.js +219 -0
- package/dist/engine/core/quality-scorer.js +27 -19
- package/dist/engine/core/scaling-curve.js +143 -0
- package/dist/engine/core/skill-dispatch.js +115 -0
- package/dist/engine/core/skill-examples.js +103 -3
- package/dist/engine/core/skill-registry.js +59 -4
- package/dist/engine/core/string-literals.js +3 -1
- package/dist/engine/core/test-files.js +77 -0
- package/dist/engine/core/vacuity-scanner.js +330 -15
- package/dist/engine/core/workflow-discovery.js +41 -23
- package/dist/engine/guardian/adjudication-github.js +136 -0
- package/dist/engine/guardian/adjudication.js +119 -340
- package/dist/engine/guardian/analysis-emit.js +7 -2
- package/dist/engine/guardian/cli.js +277 -249
- package/dist/engine/guardian/coverage.js +2 -1
- package/dist/engine/guardian/diff-coverage/coverage-delta.js +162 -0
- package/dist/engine/guardian/diff-coverage/formats/cobertura.js +45 -1
- package/dist/engine/guardian/diff-coverage/orchestrator.js +25 -21
- package/dist/engine/guardian/diff-coverage/paths.js +5 -9
- package/dist/engine/guardian/diff-coverage/report-tier.js +88 -12
- package/dist/engine/guardian/diff-extractor.js +31 -32
- package/dist/engine/guardian/pr-check.js +354 -223
- package/dist/engine/guardian/pr-comment.js +35 -58
- package/dist/engine/guardian/weak-test.js +236 -0
- package/dist/engine/mcp-server.js +67 -4
- package/dist/engine/permission-matrix-cli.js +51 -0
- package/dist/engine/scaling-curve-cli.js +147 -0
- package/dist/engine/skills-cli.js +171 -51
- package/dist/engine/workflow-cli.js +85 -65
- package/dist/reporters/testtracker.d.ts +1 -1
- package/dist/reporters/testtracker.js +1 -1
- package/package.json +3 -2
|
@@ -0,0 +1,183 @@
|
|
|
1
|
+
// ledger -- the quarantine record for tests that are out of the suite.
|
|
2
|
+
//
|
|
3
|
+
// Every captured deletion is written with its provenance (who, when, what
|
|
4
|
+
// commit, why) so a test that vanishes leaves a trail instead of a silent gap.
|
|
5
|
+
// The ledger is de-duplicated: re-running the capture on the same change adds
|
|
6
|
+
// nothing, and a batch of new entries is sorted for a stable on-disk order.
|
|
7
|
+
//
|
|
8
|
+
// SCHEMA v2 (#771) adds the WHY a row is out, alongside the provenance of how it
|
|
9
|
+
// left: `cause`, `issue`, `expiry`. Two fields deserve their separation --
|
|
10
|
+
// `reason` is DERIVED (the commit subject, "what change did this"), `cause` is
|
|
11
|
+
// ASSERTED (a judgement someone made, "why is it out"). Collapsing them would
|
|
12
|
+
// dress an auto-derived string up as a claim a person stands behind.
|
|
13
|
+
//
|
|
14
|
+
// v2 also makes the ledger no longer purely append-only, for one narrow case
|
|
15
|
+
// documented at `appendEntries`: a row that states a cause SUPERSEDES a causeless
|
|
16
|
+
// row for the same (test, file). See the comment there for why leaving both is
|
|
17
|
+
// worse than replacing one.
|
|
18
|
+
|
|
19
|
+
import fs from 'node:fs';
|
|
20
|
+
import path from 'node:path';
|
|
21
|
+
|
|
22
|
+
export const SCHEMA_VERSION = 2;
|
|
23
|
+
|
|
24
|
+
// The fields, in the order they define a row's identity for de-duplication.
|
|
25
|
+
const FIELDS = [
|
|
26
|
+
'test',
|
|
27
|
+
'file',
|
|
28
|
+
'kind',
|
|
29
|
+
'marker',
|
|
30
|
+
'commit',
|
|
31
|
+
'author',
|
|
32
|
+
'date',
|
|
33
|
+
'reason',
|
|
34
|
+
'cause',
|
|
35
|
+
'issue',
|
|
36
|
+
'expiry',
|
|
37
|
+
];
|
|
38
|
+
|
|
39
|
+
/**
|
|
40
|
+
* Why a test is out of the suite. The two middle values are the ones where the
|
|
41
|
+
* TEST IS CORRECT and someone else owns the fix, so a producer must require an
|
|
42
|
+
* `issue` for them -- an untracked "the product is broken" note decays into an
|
|
43
|
+
* unexplained skip within a release.
|
|
44
|
+
*/
|
|
45
|
+
export const CAUSES = ['flaky', 'product-defect', 'blocked-data', 'obsolete'];
|
|
46
|
+
|
|
47
|
+
/** Causes for which a row without an `issue` is not a real record. */
|
|
48
|
+
export const CAUSES_REQUIRING_ISSUE = ['product-defect', 'blocked-data'];
|
|
49
|
+
|
|
50
|
+
/**
|
|
51
|
+
* @typedef {{test: string, file: string, kind: string, marker: string,
|
|
52
|
+
* commit: string, author: string, date: string, reason: string,
|
|
53
|
+
* cause: string, issue: string, expiry: string}} LedgerRow
|
|
54
|
+
*/
|
|
55
|
+
|
|
56
|
+
/** Normalize an entry-like object into a row with the canonical field order. */
|
|
57
|
+
export function LedgerEntry(fields) {
|
|
58
|
+
const row = {};
|
|
59
|
+
for (const f of FIELDS) row[f] = fields[f] ?? '';
|
|
60
|
+
// v1 wrote the tracker link as `ticket` (#781). It is the same fact under a
|
|
61
|
+
// different name, so it migrates onto `issue` rather than being dropped or
|
|
62
|
+
// kept alongside -- two fields answering "what is this waiting on" is how a
|
|
63
|
+
// consumer ends up reading the empty one. `Ticket:` survives as the name of
|
|
64
|
+
// the COMMIT TRAILER that populates it, which is a mechanism, not a schema.
|
|
65
|
+
if (!row.issue && typeof fields.ticket === 'string')
|
|
66
|
+
row.issue = fields.ticket;
|
|
67
|
+
return row;
|
|
68
|
+
}
|
|
69
|
+
|
|
70
|
+
// The subset that defines a row's IDENTITY for de-duplication.
|
|
71
|
+
//
|
|
72
|
+
// `issue` and `expiry` are excluded, for the reason #781 excluded `ticket`:
|
|
73
|
+
// identity is what happened -- which test, in which file, muted how, by which
|
|
74
|
+
// commit and why -- and a tracker link is an attribute of that event rather
|
|
75
|
+
// than part of it. Including them would also break the append-only guarantee
|
|
76
|
+
// across this schema change, since rows written before the fields existed key
|
|
77
|
+
// as `''` and the same capture re-run with an issue present would hash
|
|
78
|
+
// differently and append a duplicate of a row already on disk.
|
|
79
|
+
//
|
|
80
|
+
// `cause` IS identity, and deliberately so: `appendEntries` decides supersede
|
|
81
|
+
// on whether a row states one, so a caused and a causeless row for the same
|
|
82
|
+
// test must remain distinguishable here for that decision to have anything to
|
|
83
|
+
// act on.
|
|
84
|
+
const IDENTITY_FIELDS = FIELDS.filter((f) => f !== 'issue' && f !== 'expiry');
|
|
85
|
+
|
|
86
|
+
const key = (row) => IDENTITY_FIELDS.map((f) => row[f] ?? '').join('\u0000');
|
|
87
|
+
|
|
88
|
+
/** The (test, file) pair a supersede decision is made on. */
|
|
89
|
+
const pair = (row) => `${row.test ?? ''}\u0000${row.file ?? ''}`;
|
|
90
|
+
|
|
91
|
+
/**
|
|
92
|
+
* Load the ledger document, or an empty one when the file is absent.
|
|
93
|
+
* Throws on unparseable JSON or a non-object top level -- a corrupt ledger is a
|
|
94
|
+
* hard error the caller must surface, not silently overwrite.
|
|
95
|
+
*
|
|
96
|
+
* Every row is normalized through `LedgerEntry`, so a v1 file read here comes
|
|
97
|
+
* back with the v2 fields present and empty. That is what makes writing
|
|
98
|
+
* `schema_version: 2` honest: the version claims "these rows have these fields",
|
|
99
|
+
* and after normalization they do. Stamping the version over un-migrated rows
|
|
100
|
+
* would make it a promise the file does not keep.
|
|
101
|
+
* @returns {{schema_version: number, entries: LedgerRow[]}}
|
|
102
|
+
*/
|
|
103
|
+
export function load(filePath) {
|
|
104
|
+
if (!fs.existsSync(filePath)) {
|
|
105
|
+
return { schema_version: SCHEMA_VERSION, entries: [] };
|
|
106
|
+
}
|
|
107
|
+
let data;
|
|
108
|
+
try {
|
|
109
|
+
data = JSON.parse(fs.readFileSync(filePath, 'utf8'));
|
|
110
|
+
} catch (exc) {
|
|
111
|
+
throw new Error(`ledger is not valid JSON: ${filePath}: ${exc.message}`);
|
|
112
|
+
}
|
|
113
|
+
if (data === null || typeof data !== 'object' || Array.isArray(data)) {
|
|
114
|
+
throw new Error(`ledger top level must be an object: ${filePath}`);
|
|
115
|
+
}
|
|
116
|
+
if (!('schema_version' in data)) data.schema_version = SCHEMA_VERSION;
|
|
117
|
+
if (!('entries' in data)) data.entries = [];
|
|
118
|
+
if (!Array.isArray(data.entries)) {
|
|
119
|
+
throw new Error(`ledger entries must be an array: ${filePath}`);
|
|
120
|
+
}
|
|
121
|
+
data.entries = data.entries.map((e) => LedgerEntry(e));
|
|
122
|
+
return data;
|
|
123
|
+
}
|
|
124
|
+
|
|
125
|
+
/**
|
|
126
|
+
* Append `entries` to the ledger at `filePath` and persist it. New entries are
|
|
127
|
+
* sorted by (file, test) and de-duplicated against the batch and disk.
|
|
128
|
+
*
|
|
129
|
+
* ONE ROW PER (test, file) MAY STATE A CAUSE, and a caused row wins:
|
|
130
|
+
*
|
|
131
|
+
* - a new row WITH a cause replaces a causeless row for the same pair
|
|
132
|
+
* - a new row WITHOUT a cause is dropped when a caused row already exists
|
|
133
|
+
*
|
|
134
|
+
* Without this, katana recording `{kind: 'skipped', cause: ''}` and a quarantine
|
|
135
|
+
* writer recording `{kind: 'skipped', cause: 'product-defect', issue: '#123'}`
|
|
136
|
+
* differ in `key()` and BOTH persist. A consumer that fails on an unlinked
|
|
137
|
+
* quarantine (canary-ci-ready does) then fails on the causeless row while the
|
|
138
|
+
* linked row sits beside it -- the ledger contradicting itself about one test.
|
|
139
|
+
*
|
|
140
|
+
* History is not lost to this: rows that differ in cause-bearing state are the
|
|
141
|
+
* only ones that collapse. Two caused rows, or two causeless rows, keep the
|
|
142
|
+
* full-field identity and both remain.
|
|
143
|
+
*/
|
|
144
|
+
export function appendEntries(filePath, entries) {
|
|
145
|
+
const doc = load(filePath);
|
|
146
|
+
const existing = doc.entries;
|
|
147
|
+
const seen = new Set(existing.map(key));
|
|
148
|
+
const causedPairs = new Set(existing.filter((r) => r.cause).map(pair));
|
|
149
|
+
|
|
150
|
+
const newRows = entries
|
|
151
|
+
.map((e) => LedgerEntry(e))
|
|
152
|
+
.sort(
|
|
153
|
+
(a, b) => a.file.localeCompare(b.file) || a.test.localeCompare(b.test),
|
|
154
|
+
);
|
|
155
|
+
|
|
156
|
+
for (const row of newRows) {
|
|
157
|
+
const k = key(row);
|
|
158
|
+
if (seen.has(k)) continue;
|
|
159
|
+
const p = pair(row);
|
|
160
|
+
|
|
161
|
+
if (row.cause) {
|
|
162
|
+
// Supersede: drop any causeless row for this pair, then take its place.
|
|
163
|
+
for (let i = existing.length - 1; i >= 0; i--) {
|
|
164
|
+
if (!existing[i].cause && pair(existing[i]) === p) {
|
|
165
|
+
seen.delete(key(existing[i]));
|
|
166
|
+
existing.splice(i, 1);
|
|
167
|
+
}
|
|
168
|
+
}
|
|
169
|
+
causedPairs.add(p);
|
|
170
|
+
} else if (causedPairs.has(p)) {
|
|
171
|
+
// A causeless row must never sit next to a caused one for the same test.
|
|
172
|
+
continue;
|
|
173
|
+
}
|
|
174
|
+
|
|
175
|
+
seen.add(k);
|
|
176
|
+
existing.push(row);
|
|
177
|
+
}
|
|
178
|
+
|
|
179
|
+
doc.schema_version = SCHEMA_VERSION;
|
|
180
|
+
fs.mkdirSync(path.dirname(path.resolve(filePath)), { recursive: true });
|
|
181
|
+
fs.writeFileSync(filePath, `${JSON.stringify(doc, null, 2)}\n`, 'utf8');
|
|
182
|
+
return doc;
|
|
183
|
+
}
|
|
@@ -0,0 +1,144 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: canary-pr-guardian
|
|
3
|
+
description: >
|
|
4
|
+
PR/pre-commit test-guardian orchestrator: runs the deterministic Tier-0
|
|
5
|
+
diff-coverage pass, audits affected tests via canary-test-reviewer (Tier 1),
|
|
6
|
+
and — at the desk with authorTests opt-in — authors missing tests via
|
|
7
|
+
canary-test-author (Tier 2), staging them and blocking the commit once for
|
|
8
|
+
human review. Use to guard a change's test quality before it lands.
|
|
9
|
+
---
|
|
10
|
+
|
|
11
|
+
# Canary: PR Guardian
|
|
12
|
+
|
|
13
|
+
Guards a change's test quality before it lands. Composes the deterministic
|
|
14
|
+
Tier-0 diff-coverage engine with two native agents — `canary-test-reviewer`
|
|
15
|
+
(read-only audit) and `canary-test-author` (authoring) — under a strict
|
|
16
|
+
write-safety model. This is the **Option A** driver: the engine never calls an
|
|
17
|
+
LLM; **this skill** invokes the agents in-session and enforces
|
|
18
|
+
stage-and-block-once.
|
|
19
|
+
|
|
20
|
+
**On the tier numbers.** Here they are the values of
|
|
21
|
+
`canary guardian pr-check --tier 0|1|2`, not a repo-wide capability scale — Tier
|
|
22
|
+
1 means "the `--tier 1` pass," which is the agent audit. `Tier-0` is the one
|
|
23
|
+
number with a repo-wide meaning (deterministic, no network, no agent), and
|
|
24
|
+
`Tier-1`/`Tier-2` are guardian-local by
|
|
25
|
+
[ADR 0015](../../../../docs/knowledge/decisions/0015-skill-capability-vocabulary.md).
|
|
26
|
+
Do not carry them into other skills.
|
|
27
|
+
|
|
28
|
+
## When to Use
|
|
29
|
+
|
|
30
|
+
- Before opening or updating a PR, to check that new/changed code is tested.
|
|
31
|
+
- As a pre-commit companion when `preCommit.authorTests: true` is set and you
|
|
32
|
+
want the guardian to author the missing tests for you (at the desk only).
|
|
33
|
+
- NOT in CI for Tier-2 write-back — that is a NON-GOAL. CI runs Tier-0 only (the
|
|
34
|
+
`CANARY_GUARDIAN_AGENT` env is unset there).
|
|
35
|
+
|
|
36
|
+
## Safety model (non-negotiable)
|
|
37
|
+
|
|
38
|
+
- **NEVER commit or push.** This skill only authors and `git add`s. The human
|
|
39
|
+
reviews the staged tests and re-commits.
|
|
40
|
+
- **Honor every `skipped` reason** from `author-plan` verbatim (opt-in-off /
|
|
41
|
+
tier / fork / collision / loop-guard). Never override a skip.
|
|
42
|
+
- **Block once.** When `block.block == true`, print the block message and stop —
|
|
43
|
+
leave the staged tests for the human. The loop-guard sentinel you write in
|
|
44
|
+
Phase 3 is what stops the guardian re-authoring over its own output on the
|
|
45
|
+
next run. It is stamped with the current `HEAD` and expires by itself once the
|
|
46
|
+
human's review commit moves `HEAD` — never delete it yourself.
|
|
47
|
+
- **Authoring is opt-in.** No `preCommit.authorTests: true` ⇒ no writes, ever.
|
|
48
|
+
|
|
49
|
+
## Phases
|
|
50
|
+
|
|
51
|
+
### Phase 0 — Deterministic scope
|
|
52
|
+
|
|
53
|
+
Run the Tier-0 pass and read its findings:
|
|
54
|
+
|
|
55
|
+
```bash
|
|
56
|
+
canary guardian pr-check --format json
|
|
57
|
+
```
|
|
58
|
+
|
|
59
|
+
Findings are `untested-new-code` gaps. If there are none, report clean and stop.
|
|
60
|
+
|
|
61
|
+
**Coverage regressions (#606).** `--coverage` alone answers only "is this
|
|
62
|
+
changed unit covered at all?". To also catch a unit whose coverage _fell_
|
|
63
|
+
against the base branch, pass the base ref's report as well:
|
|
64
|
+
|
|
65
|
+
```bash
|
|
66
|
+
canary guardian pr-check --coverage lcov.info --base-coverage base-lcov.info --format json
|
|
67
|
+
```
|
|
68
|
+
|
|
69
|
+
That adds `coverage-regression` findings, graded by how many percentage points
|
|
70
|
+
were lost. Most CI never uploads a base-branch artifact; without
|
|
71
|
+
`--base-coverage` the run degrades **loudly** to "delta unavailable — head-only"
|
|
72
|
+
and `coverage_delta.status` reports `unavailable`. Read that field before
|
|
73
|
+
treating an empty finding list as "no regressions" — a run that compared nothing
|
|
74
|
+
has abstained, not passed.
|
|
75
|
+
|
|
76
|
+
### Phase 1 — Quality audit (Tier ≥ 1, read-only)
|
|
77
|
+
|
|
78
|
+
Export the availability signal so the probe reports the ceiling, then audit the
|
|
79
|
+
affected tests with `canary-test-reviewer`:
|
|
80
|
+
|
|
81
|
+
```bash
|
|
82
|
+
export CANARY_GUARDIAN_AGENT=1 # or 2 when authoring is enabled
|
|
83
|
+
```
|
|
84
|
+
|
|
85
|
+
If this checkout is a **fork** (an untrusted/read-only context — detect it, e.g.
|
|
86
|
+
`git config --get remote.origin.url` pointing at a fork, or a CI fork PR), arm
|
|
87
|
+
the fork guard so the engine's safety layer never authors on it:
|
|
88
|
+
|
|
89
|
+
```bash
|
|
90
|
+
export CANARY_GUARDIAN_IS_FORK=1 # any value other than "0"/unset means fork
|
|
91
|
+
```
|
|
92
|
+
|
|
93
|
+
Leave `CANARY_GUARDIAN_IS_FORK` unset (or `0`) at your own desk on a trusted
|
|
94
|
+
checkout. The guard fails CLOSED: any ambiguous value is treated as a fork and
|
|
95
|
+
authoring is skipped.
|
|
96
|
+
|
|
97
|
+
Use the `canary-test-reviewer` agent to review the affected tests. This is
|
|
98
|
+
**read-only** — surface weak-test findings; write nothing.
|
|
99
|
+
|
|
100
|
+
### Phase 2 — Authoring plan (Tier 2 + opt-in)
|
|
101
|
+
|
|
102
|
+
Ask the engine's safety layer for the plan (intents + block decision):
|
|
103
|
+
|
|
104
|
+
```bash
|
|
105
|
+
canary guardian author-plan --json
|
|
106
|
+
```
|
|
107
|
+
|
|
108
|
+
The JSON is `{"intents": [...], "block": {...}}`. Each intent carries `status`
|
|
109
|
+
(`planned` | `authored` | `skipped`), `target_path`, `requirement`, and a
|
|
110
|
+
`skip_reason` when skipped. **Do not author anything the plan skipped** — the
|
|
111
|
+
Engine guards (opt-in, fork, collision, loop-guard) are authoritative.
|
|
112
|
+
|
|
113
|
+
### Phase 3 — Author, stage, block once
|
|
114
|
+
|
|
115
|
+
For each intent with `status: "planned"`, use the `canary-test-author` agent
|
|
116
|
+
with the intent's `requirement` as the task and `target_path` as the
|
|
117
|
+
destination. **Never overwrite an existing file at `target_path`.** The engine's
|
|
118
|
+
collision guard runs at plan time, so between planning and writing another
|
|
119
|
+
PR/session may have created the file (a TOCTOU window). If the target already
|
|
120
|
+
exists at write time, skip that intent and report it — do not clobber it. Then
|
|
121
|
+
stage the authored files:
|
|
122
|
+
|
|
123
|
+
```bash
|
|
124
|
+
git add <target_path>
|
|
125
|
+
```
|
|
126
|
+
|
|
127
|
+
When `block.block == true`, record the guardian-authored paths in the loop-guard
|
|
128
|
+
sentinel BEFORE blocking, by running the deterministic producer command once per
|
|
129
|
+
authored path:
|
|
130
|
+
|
|
131
|
+
```bash
|
|
132
|
+
canary guardian mark-authored --path <target_path> [--path <target_path> ...]
|
|
133
|
+
```
|
|
134
|
+
|
|
135
|
+
This is a real CLI step (not something you `touch` yourself): it writes the
|
|
136
|
+
sentinel inside the real git dir and records exactly the paths you authored,
|
|
137
|
+
stamped with the current `HEAD`. `author-plan` reads it on the next run and
|
|
138
|
+
returns a `loop-guard` skip **while `HEAD` is unchanged**, so the guardian never
|
|
139
|
+
authors on top of its own output — and once the human commits the reviewed
|
|
140
|
+
tests, `HEAD` moves and authoring re-enables itself.
|
|
141
|
+
|
|
142
|
+
Then print `block.message` (the "N test(s) authored & staged — review and
|
|
143
|
+
re-commit" notice) and **stop**. Do not commit. The human reviews the staged
|
|
144
|
+
tests and re-commits.
|
|
@@ -0,0 +1,17 @@
|
|
|
1
|
+
name: canary-pr-guardian
|
|
2
|
+
version: '1.0.0'
|
|
3
|
+
description:
|
|
4
|
+
PR/pre-commit test-guardian orchestrator — Tier-0 diff-coverage, Tier-1
|
|
5
|
+
quality audit via canary-test-reviewer, and at-desk Tier-2 authoring via
|
|
6
|
+
canary-test-author with stage-and-block-once review.
|
|
7
|
+
stability: static
|
|
8
|
+
triggers:
|
|
9
|
+
- manual
|
|
10
|
+
platforms:
|
|
11
|
+
- claude-code
|
|
12
|
+
type: rigid
|
|
13
|
+
tools: []
|
|
14
|
+
tier: 1
|
|
15
|
+
depends_on:
|
|
16
|
+
- canary-test-reviewer
|
|
17
|
+
- canary-test-author
|
|
@@ -0,0 +1,228 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: canary-promote-test
|
|
3
|
+
description: >
|
|
4
|
+
Move a generated test from `tests/generated/` into the committed test suite —
|
|
5
|
+
reviews it for correctness, drops the generation header, relocates it to the
|
|
6
|
+
matching suite directory, and confirms it runs in the project's normal test
|
|
7
|
+
flow. Use for "promote this test", "commit this generated test", "move this
|
|
8
|
+
test into the suite", or "keep this test" — always after the generated test
|
|
9
|
+
has been validated against the SUT, never before. Not for tests needing
|
|
10
|
+
substantial rewriting (regenerate instead) or throwaway investigation tests
|
|
11
|
+
(leave in `tests/generated/`).
|
|
12
|
+
---
|
|
13
|
+
|
|
14
|
+
# Canary: Promote Test
|
|
15
|
+
|
|
16
|
+
> Move a generated test from `tests/generated/` into the committed test suite.
|
|
17
|
+
> Reviews the test for correctness, drops the generation header, relocates it
|
|
18
|
+
> under the appropriate suite directory, and confirms it runs in the project's
|
|
19
|
+
> normal test flow.
|
|
20
|
+
|
|
21
|
+
## When to Use
|
|
22
|
+
|
|
23
|
+
- After `canary-generate-test` produced a file that has been validated against
|
|
24
|
+
the SUT
|
|
25
|
+
- When a user explicitly asks to "commit", "save", or "keep" a generated test
|
|
26
|
+
- When extending an existing suite with a new case the team has agreed to
|
|
27
|
+
maintain
|
|
28
|
+
- NOT before the generated test has been executed and reviewed — promotion is
|
|
29
|
+
the _last_ step, not the first
|
|
30
|
+
- NOT for tests that need substantial rewriting — regenerate with a better
|
|
31
|
+
prompt instead of hand-patching
|
|
32
|
+
- NOT for tests targeting throwaway investigations (perf spikes, ad-hoc bug
|
|
33
|
+
triage) — leave those in `tests/generated/`
|
|
34
|
+
|
|
35
|
+
## Process
|
|
36
|
+
|
|
37
|
+
### Phase 0: GATE — Get the Structured Verdict First
|
|
38
|
+
|
|
39
|
+
Run this before reading the test. It is deterministic, takes no API key, and it
|
|
40
|
+
will refuse the drafts that are not worth your review time.
|
|
41
|
+
|
|
42
|
+
```bash
|
|
43
|
+
canary promote-check tests/generated/api/orders_post.py --json
|
|
44
|
+
```
|
|
45
|
+
|
|
46
|
+
| Exit | Verdict | What it means |
|
|
47
|
+
| ---- | --------- | ------------------------------------------------------------ |
|
|
48
|
+
| `0` | `promote` | No gating defect. Advisory findings may remain — your call. |
|
|
49
|
+
| `1` | `block` | A gating defect. **Do not promote.** Regenerate or fix. |
|
|
50
|
+
| `3` | `abstain` | No verdict could be produced. Promotion is **not** approved. |
|
|
51
|
+
|
|
52
|
+
**Which axes gate, and why only those.** Gating on all eight axes of a quality
|
|
53
|
+
critique would block nearly every promotion, so only deterministic defects do:
|
|
54
|
+
|
|
55
|
+
| Axis | Rules | Gates | Reason |
|
|
56
|
+
| ----------------- | ------------------------------ | ----- | ------------------------------------------------------ |
|
|
57
|
+
| `soundness` | `SOUND-001/002/003` | yes | Pins a value no correct implementation must produce |
|
|
58
|
+
| `assertions` | `LINT-006` | yes | A test that asserts nothing always passes |
|
|
59
|
+
| `flakiness` | `FLAKE-001/002` | yes | A hardcoded sleep in a committed suite is a future red |
|
|
60
|
+
| `vacuity` | `VAC-001/003`, annotated `002` | yes | Cannot fail, or contradicts a declared `@covers` |
|
|
61
|
+
| `selectors` | `LINT-001/002/003` | no | Brittle, not wrong — a reviewer's call |
|
|
62
|
+
| `maintainability` | `LINT-005`, `FLAKE-003/004` | no | Style and softer signals |
|
|
63
|
+
|
|
64
|
+
**`VAC-002` gates only at `annotated` fidelity.** At `import-inferred` it is an
|
|
65
|
+
inference about which symbol the test meant to exercise, and a heuristic must
|
|
66
|
+
not be load-bearing on a promotion gate. If you want the gate to check the real
|
|
67
|
+
target, add the annotation to the generated test:
|
|
68
|
+
|
|
69
|
+
```ts
|
|
70
|
+
// @covers resolveOverlay
|
|
71
|
+
```
|
|
72
|
+
|
|
73
|
+
**An `abstain` is not a pass.** Exit 3 means the checker had no subject — an
|
|
74
|
+
unparseable extension, or a file with no test declarations. Promotion falls back
|
|
75
|
+
to the manual review below and is **not** approved by silence. Do not read a
|
|
76
|
+
missing verdict as either stricter or looser than one.
|
|
77
|
+
|
|
78
|
+
**An LLM judgement never gates.** `harness:test-craft` runs an 8-axis per-test
|
|
79
|
+
critique and remains exactly what step 5 below calls it: an optional deeper
|
|
80
|
+
audit for a human. Everything that blocks in this repo is deterministic, and
|
|
81
|
+
`promote-check` keeps it that way — the verdict has no field an LLM opinion
|
|
82
|
+
could arrive in.
|
|
83
|
+
|
|
84
|
+
### Phase 1: REVIEW — Confirm the Test Is Worth Keeping
|
|
85
|
+
|
|
86
|
+
1. **Read the test end-to-end.** Treat it like any other code review — naming,
|
|
87
|
+
assertions, hardcoded values, missing edge cases. Generated tests are drafts,
|
|
88
|
+
not finished artifacts.
|
|
89
|
+
2. **Run the test against the real SUT** (not just the env it was generated
|
|
90
|
+
against). A test that only passes in one environment is a fixture-bound test,
|
|
91
|
+
not a regression test.
|
|
92
|
+
3. **Check assertion strength.** A test that only asserts "status code 200" is
|
|
93
|
+
weak; promote it only if that's genuinely the contract. If the SUT returns
|
|
94
|
+
structured data, the test should assert on shape.
|
|
95
|
+
4. **Confirm no hardcoded secrets or environment-specific URLs.** Replace with
|
|
96
|
+
fixtures or env-driven config before promoting.
|
|
97
|
+
5. **Optional: run `harness:test-craft` for a deeper quality audit.**
|
|
98
|
+
`test-craft` runs an 8-axis per-test LLM critique (assertion density,
|
|
99
|
+
flakiness risk, contract vs implementation, etc.). Use it when the generated
|
|
100
|
+
test is substantial or when the team wants a second opinion before
|
|
101
|
+
committing. Not required for simple happy-path tests, and **never a blocker**
|
|
102
|
+
— see Phase 0.
|
|
103
|
+
6. **Triage the advisory findings** `promote-check` reported. They did not
|
|
104
|
+
block; deciding whether they matter here is the review's job.
|
|
105
|
+
7. **Decide: promote, regenerate, or discard.** If review reveals more than ~3
|
|
106
|
+
small fixes, regenerate with a sharper prompt instead.
|
|
107
|
+
|
|
108
|
+
### Phase 2: RELOCATE — Move into the Suite
|
|
109
|
+
|
|
110
|
+
1. **Identify the destination directory.** Mirror the suite's structure:
|
|
111
|
+
- `tests/generated/api/foo.py` → `tests/api/foo.py`
|
|
112
|
+
- `tests/generated/e2e/checkout.spec.ts` → `tests/e2e/checkout.spec.ts`
|
|
113
|
+
- `tests/generated/unit/validator.py` → `tests/unit/validator.py`
|
|
114
|
+
2. **Match suite conventions.** Look at neighboring files for:
|
|
115
|
+
- Import style (relative vs absolute)
|
|
116
|
+
- Fixture/setup imports (most suites have a `conftest.py` or shared setup)
|
|
117
|
+
- Naming conventions (`test_<feature>_<case>.py`, `<feature>.spec.ts`)
|
|
118
|
+
3. **Move with `git mv`** so the history is preserved if anyone later runs
|
|
119
|
+
`git log --follow`.
|
|
120
|
+
4. **Update imports** if the file referenced anything by relative path from
|
|
121
|
+
`tests/generated/`.
|
|
122
|
+
|
|
123
|
+
### Phase 3: CLEAN — Drop Generation Artifacts
|
|
124
|
+
|
|
125
|
+
1. **Remove the timestamped generation header.** The "Generated by Canary on
|
|
126
|
+
[date]" comment is useful in scratch space — meaningless in a committed test
|
|
127
|
+
and rots immediately.
|
|
128
|
+
2. **Remove any placeholder TODOs.** The generator sometimes emits
|
|
129
|
+
`# TODO: adjust selector` or similar — either resolve them or stop the
|
|
130
|
+
promotion and regenerate.
|
|
131
|
+
3. **Tighten formatting.** Run the project's formatter (`black`, `prettier`,
|
|
132
|
+
etc.) so the file matches surrounding style.
|
|
133
|
+
4. **Strip dead code.** Imports the generator added "just in case" but the test
|
|
134
|
+
doesn't use.
|
|
135
|
+
|
|
136
|
+
### Phase 4: VERIFY — Run in the Project's Normal Flow
|
|
137
|
+
|
|
138
|
+
1. **Run the suite that owns this test:**
|
|
139
|
+
- Python: `pytest tests/api/test_orders_post.py`
|
|
140
|
+
- Playwright: `npx playwright test tests/e2e/checkout.spec.ts`
|
|
141
|
+
- Match whatever CI runs.
|
|
142
|
+
2. **Run the full suite** to confirm no collateral failure (shared fixtures,
|
|
143
|
+
port conflicts, ordering issues).
|
|
144
|
+
3. **Confirm CI configuration picks it up.** If the suite has a glob in CI
|
|
145
|
+
config, verify the new path matches. If not, add it.
|
|
146
|
+
4. **Log the promotion.** Append a one-line entry to `docs/CANARY_STATE.md` so
|
|
147
|
+
the project ledger tracks which generated tests have been promoted.
|
|
148
|
+
|
|
149
|
+
## Canary Integration
|
|
150
|
+
|
|
151
|
+
- **`tests/generated/`** — Source for promotion. Gitignored; nothing here is
|
|
152
|
+
ever a final artifact.
|
|
153
|
+
- **`tests/<suite>/`** — Destination. Each suite has its own conventions; never
|
|
154
|
+
invent a new top-level dir during promotion.
|
|
155
|
+
- **`docs/CANARY_STATE.md`** — Append a promotion entry: requirement, generated
|
|
156
|
+
path, promoted path, date.
|
|
157
|
+
|
|
158
|
+
## Success Criteria
|
|
159
|
+
|
|
160
|
+
- The promoted test passes when run via the project's normal test command
|
|
161
|
+
- The full suite passes (no collateral breakage)
|
|
162
|
+
- The file matches the surrounding code style (formatter clean, lint clean)
|
|
163
|
+
- No generation artifacts remain (timestamp header, placeholder TODOs, unused
|
|
164
|
+
imports)
|
|
165
|
+
- CI picks up the new test on the next push
|
|
166
|
+
|
|
167
|
+
## Rationalizations to Reject
|
|
168
|
+
|
|
169
|
+
| Rationalization | Why It Is Wrong |
|
|
170
|
+
| ------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
171
|
+
| "`promote-check` exited 3, so nothing was wrong" | Exit 3 is an abstention — the checker had no subject at all. It is the one outcome that proves nothing, so it can never stand in for approval. |
|
|
172
|
+
| "`promote-check` blocked on VAC-002, but the test is fine" | Only an `annotated` VAC-002 blocks, which means the file's own `@covers` names a symbol the test never touches. Either the annotation is wrong or the test is. |
|
|
173
|
+
| "I'll wire `harness:test-craft` into the gate for better coverage" | An 8-axis LLM critique gating a promotion would block nearly everything and would make a judgement load-bearing. Deterministic verdicts gate; critiques inform. |
|
|
174
|
+
| "The test works, I'll skip the review and just commit it" | Generated code looks plausible but commonly has weak assertions or hardcoded values. Review is the whole point — promotion without review just imports debt into the suite. |
|
|
175
|
+
| "I'll fix the 6 issues I found in the review by editing the test" | Six issues means the prompt was wrong. Regenerate; hand-edits won't transfer to the next similar test. |
|
|
176
|
+
| "I'll leave the timestamp header so we know when it was generated" | Git history already records when the file was committed. The header rots and creates noise. |
|
|
177
|
+
| "It passes against staging, that's good enough" | If the test only passes in one environment, it's a fixture-bound smoke check, not a regression test. Either parametrize the env or don't promote it. |
|
|
178
|
+
| "I'll commit it without running the full suite — I only changed one test" | New tests can break shared fixtures, conflict on ports, leak state. Always run the full suite once before commit. |
|
|
179
|
+
|
|
180
|
+
## Examples
|
|
181
|
+
|
|
182
|
+
### Example: Clean promotion
|
|
183
|
+
|
|
184
|
+
**Source:** `tests/generated/api/orders_post_201.py`, validated, single status
|
|
185
|
+
assertion is exactly the contract.
|
|
186
|
+
|
|
187
|
+
**Action:**
|
|
188
|
+
|
|
189
|
+
1. `git mv tests/generated/api/orders_post_201.py tests/api/test_orders_post_201.py`
|
|
190
|
+
2. Drop the `# Generated by Canary...` header.
|
|
191
|
+
3. Adjust import to match `tests/api/conftest.py` fixtures.
|
|
192
|
+
4. Run `pytest tests/api/` → passes.
|
|
193
|
+
5. Append to `CANARY_STATE.md`: promoted `orders_post_201.py` →
|
|
194
|
+
`tests/api/test_orders_post_201.py` on [date].
|
|
195
|
+
|
|
196
|
+
### Example: Promotion abandoned — regenerate instead
|
|
197
|
+
|
|
198
|
+
**Source:** `tests/generated/e2e/checkout.spec.ts`. Review finds: hardcoded test
|
|
199
|
+
user creds, no wait for navigation, weak assertion (`expect(true).toBe(true)`),
|
|
200
|
+
TODO comments left in three places.
|
|
201
|
+
|
|
202
|
+
**Action:** Stop. Regenerate with a sharper prompt that names the real fixture
|
|
203
|
+
user, specifies the wait condition, and asserts the actual checkout success
|
|
204
|
+
state. Do not commit the broken draft.
|
|
205
|
+
|
|
206
|
+
### Example: Test for throwaway investigation
|
|
207
|
+
|
|
208
|
+
**Source:** `tests/generated/performance/spike_search_50rps.js`. Used once to
|
|
209
|
+
confirm a single perf hypothesis. No ongoing value.
|
|
210
|
+
|
|
211
|
+
**Action:** Do not promote. Leave in `tests/generated/`. If the team wants
|
|
212
|
+
ongoing perf monitoring, generate a _new_ test with a sustainable RPS profile
|
|
213
|
+
and promote that.
|
|
214
|
+
|
|
215
|
+
## Escalation
|
|
216
|
+
|
|
217
|
+
- **When the review reveals the SUT itself is broken:** Don't promote the test
|
|
218
|
+
that passes against the broken SUT. File a bug; only promote the test once the
|
|
219
|
+
SUT is fixed and the assertion has a real contract behind it.
|
|
220
|
+
- **When the suite has no existing convention to match:** This usually means the
|
|
221
|
+
test belongs in a new sub-suite. Ask the user to confirm the suite structure
|
|
222
|
+
before inventing one.
|
|
223
|
+
- **When CI doesn't pick up the new path:** Don't merge until CI config matches.
|
|
224
|
+
A test that exists but isn't run is worse than no test — it implies coverage
|
|
225
|
+
that doesn't exist.
|
|
226
|
+
- **When the promoted test starts flaking after merge:** Treat as a real
|
|
227
|
+
regression in the test (or the SUT). Don't quarantine in `tests/generated/` —
|
|
228
|
+
fix or delete.
|