canary-test-cli 7.1.0 → 8.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (128) hide show
  1. package/agents/skills/README.md +327 -0
  2. package/agents/skills/canary:generate.md +49 -0
  3. package/agents/skills/canary:init.md +37 -0
  4. package/agents/skills/canary:migrate.md +66 -0
  5. package/agents/skills/claude-code/canary-add-framework/SKILL.md +248 -0
  6. package/agents/skills/claude-code/canary-batwoman/SKILL.md +119 -0
  7. package/agents/skills/claude-code/canary-blackhawk/SKILL.md +170 -0
  8. package/agents/skills/claude-code/canary-blackhawk/scripts/cli.mjs +188 -0
  9. package/agents/skills/claude-code/canary-blackhawk/scripts/rules.mjs +120 -0
  10. package/agents/skills/claude-code/canary-blackhawk/scripts/scanner.mjs +244 -0
  11. package/agents/skills/claude-code/canary-blackhawk/scripts/string-literals.mjs +116 -0
  12. package/agents/skills/claude-code/canary-cassandra/SKILL.md +187 -0
  13. package/agents/skills/claude-code/canary-cassandra/scripts/cli.mjs +270 -0
  14. package/agents/skills/claude-code/canary-cassandra/scripts/engine.mjs +95 -0
  15. package/agents/skills/claude-code/canary-ci-ready/SKILL.md +178 -0
  16. package/agents/skills/claude-code/canary-ci-ready/skill.yaml +14 -0
  17. package/agents/skills/claude-code/canary-company-knowledge/SKILL.md +196 -0
  18. package/agents/skills/claude-code/canary-critical-areas/SKILL.md +142 -0
  19. package/agents/skills/claude-code/canary-critical-areas/skill.yaml +16 -0
  20. package/agents/skills/claude-code/canary-edge-case-discovery/SKILL.md +160 -0
  21. package/agents/skills/claude-code/canary-edge-case-discovery/skill.yaml +16 -0
  22. package/agents/skills/claude-code/canary-fail-fast/SKILL.md +75 -0
  23. package/agents/skills/claude-code/canary-fail-fast/scripts/cli.mjs +118 -0
  24. package/agents/skills/claude-code/canary-fail-fast/scripts/digest.mjs +69 -0
  25. package/agents/skills/claude-code/canary-fail-fast/scripts/failures.mjs +60 -0
  26. package/agents/skills/claude-code/canary-fail-fast/scripts/fastfail_check.mjs +43 -0
  27. package/agents/skills/claude-code/canary-fail-fast/scripts/parse.mjs +149 -0
  28. package/agents/skills/claude-code/canary-failure-impact/SKILL.md +153 -0
  29. package/agents/skills/claude-code/canary-failure-impact/skill.yaml +15 -0
  30. package/agents/skills/claude-code/canary-fleet-health/SKILL.md +197 -0
  31. package/agents/skills/claude-code/canary-generate-test/SKILL.md +185 -0
  32. package/agents/skills/claude-code/canary-instrument/SKILL.md +157 -0
  33. package/agents/skills/claude-code/canary-instrument/scripts/cli.mjs +178 -0
  34. package/agents/skills/claude-code/canary-instrument/scripts/otel_bootstrap/instrument.mjs +96 -0
  35. package/agents/skills/claude-code/canary-instrument/scripts/otel_bootstrap/playwright-fixture.ts +44 -0
  36. package/agents/skills/claude-code/canary-instrument/scripts/run_types.mjs +81 -0
  37. package/agents/skills/claude-code/canary-instrument/scripts/span_reader.mjs +187 -0
  38. package/agents/skills/claude-code/canary-katana/SKILL.md +243 -0
  39. package/agents/skills/claude-code/canary-katana/scripts/alarm.mjs +296 -0
  40. package/agents/skills/claude-code/canary-katana/scripts/cli.mjs +247 -0
  41. package/agents/skills/claude-code/canary-katana/scripts/diffscan.mjs +0 -0
  42. package/agents/skills/claude-code/canary-katana/scripts/ledger.mjs +183 -0
  43. package/agents/skills/claude-code/canary-pr-guardian/SKILL.md +144 -0
  44. package/agents/skills/claude-code/canary-pr-guardian/skill.yaml +17 -0
  45. package/agents/skills/claude-code/canary-promote-test/SKILL.md +228 -0
  46. package/agents/skills/claude-code/canary-savant/SKILL.md +233 -0
  47. package/agents/skills/claude-code/canary-savant/scripts/cli.mjs +274 -0
  48. package/agents/skills/claude-code/canary-savant/scripts/restoration.mjs +274 -0
  49. package/agents/skills/claude-code/canary-savant/scripts/rules.mjs +168 -0
  50. package/agents/skills/claude-code/canary-savant/scripts/runner.mjs +572 -0
  51. package/agents/skills/claude-code/canary-savant/scripts/scanner.mjs +374 -0
  52. package/agents/skills/claude-code/canary-savant/scripts/string-literals.mjs +116 -0
  53. package/agents/skills/claude-code/canary-screech/SKILL.md +109 -0
  54. package/agents/skills/claude-code/canary-screech/scripts/blast.mjs +125 -0
  55. package/agents/skills/claude-code/canary-screech/scripts/cli.mjs +128 -0
  56. package/agents/skills/claude-code/canary-screech/scripts/cluster.mjs +97 -0
  57. package/agents/skills/claude-code/canary-screech/scripts/history.mjs +73 -0
  58. package/agents/skills/claude-code/canary-screech/scripts/redness.mjs +94 -0
  59. package/agents/skills/claude-code/canary-setup-harness/SKILL.md +263 -0
  60. package/agents/skills/claude-code/canary-shadow/SKILL.md +131 -0
  61. package/agents/skills/claude-code/canary-shadow/scripts/cases.example.json +32 -0
  62. package/agents/skills/claude-code/canary-shadow/scripts/cli.mjs +195 -0
  63. package/agents/skills/claude-code/canary-ship/SKILL.md +177 -0
  64. package/agents/skills/claude-code/canary-ship/skill.yaml +16 -0
  65. package/agents/skills/claude-code/canary-strix/SKILL.md +130 -0
  66. package/agents/skills/claude-code/canary-strix/scripts/cli.mjs +255 -0
  67. package/agents/skills/claude-code/canary-strix/scripts/scanner.mjs +252 -0
  68. package/agents/skills/claude-code/canary-strix/scripts/terms.mjs +132 -0
  69. package/agents/skills/claude-code/canary-test-pipeline/SKILL.md +159 -0
  70. package/agents/skills/claude-code/canary-test-pipeline/skill.yaml +19 -0
  71. package/agents/skills/claude-code/canary-test-reporter/SKILL.md +138 -0
  72. package/agents/skills/claude-code/canary-test-reporter/scripts/cli.mjs +98 -0
  73. package/agents/skills/claude-code/canary-test-reporter/scripts/json_report.mjs +58 -0
  74. package/agents/skills/claude-code/canary-test-reporter/scripts/parse.mjs +216 -0
  75. package/agents/skills/claude-code/canary-test-reporter/scripts/render.mjs +114 -0
  76. package/agents/skills/lib/parse-args.mjs +275 -0
  77. package/dist/engine/analysis/batwoman/audit.js +39 -0
  78. package/dist/engine/analysis/batwoman/closure.js +159 -0
  79. package/dist/engine/analysis/batwoman/gh-history.js +119 -0
  80. package/dist/engine/analysis/batwoman/probes.js +195 -0
  81. package/dist/engine/analysis/batwoman/registry.js +142 -0
  82. package/dist/engine/analysis/batwoman/render.js +194 -0
  83. package/dist/engine/analysis/batwoman/run-window.js +122 -0
  84. package/dist/engine/analysis/batwoman/text.js +84 -0
  85. package/dist/engine/analysis/batwoman/triggers.js +122 -0
  86. package/dist/engine/analysis/batwoman/verdict.js +64 -0
  87. package/dist/engine/analysis/cli.js +47 -14
  88. package/dist/engine/analysis/gh-flaky/gh-run-attempts.js +206 -0
  89. package/dist/engine/batwoman-cli.js +119 -0
  90. package/dist/engine/ci-ready-cli.js +71 -0
  91. package/dist/engine/cli-commands.js +49 -72
  92. package/dist/engine/cli.core.js +16 -0
  93. package/dist/engine/company-knowledge-cli.js +10 -2
  94. package/dist/engine/core/ci-ready.js +112 -0
  95. package/dist/engine/core/company-knowledge.js +8 -0
  96. package/dist/engine/core/migrator.js +147 -20
  97. package/dist/engine/core/permission-matrix.js +219 -0
  98. package/dist/engine/core/quality-scorer.js +27 -19
  99. package/dist/engine/core/scaling-curve.js +143 -0
  100. package/dist/engine/core/skill-dispatch.js +115 -0
  101. package/dist/engine/core/skill-examples.js +103 -3
  102. package/dist/engine/core/skill-registry.js +59 -4
  103. package/dist/engine/core/string-literals.js +3 -1
  104. package/dist/engine/core/test-files.js +77 -0
  105. package/dist/engine/core/vacuity-scanner.js +330 -15
  106. package/dist/engine/core/workflow-discovery.js +41 -23
  107. package/dist/engine/guardian/adjudication-github.js +136 -0
  108. package/dist/engine/guardian/adjudication.js +119 -340
  109. package/dist/engine/guardian/analysis-emit.js +7 -2
  110. package/dist/engine/guardian/cli.js +277 -249
  111. package/dist/engine/guardian/coverage.js +2 -1
  112. package/dist/engine/guardian/diff-coverage/coverage-delta.js +162 -0
  113. package/dist/engine/guardian/diff-coverage/formats/cobertura.js +45 -1
  114. package/dist/engine/guardian/diff-coverage/orchestrator.js +25 -21
  115. package/dist/engine/guardian/diff-coverage/paths.js +5 -9
  116. package/dist/engine/guardian/diff-coverage/report-tier.js +88 -12
  117. package/dist/engine/guardian/diff-extractor.js +31 -32
  118. package/dist/engine/guardian/pr-check.js +354 -223
  119. package/dist/engine/guardian/pr-comment.js +35 -58
  120. package/dist/engine/guardian/weak-test.js +236 -0
  121. package/dist/engine/mcp-server.js +67 -4
  122. package/dist/engine/permission-matrix-cli.js +51 -0
  123. package/dist/engine/scaling-curve-cli.js +147 -0
  124. package/dist/engine/skills-cli.js +171 -51
  125. package/dist/engine/workflow-cli.js +85 -65
  126. package/dist/reporters/testtracker.d.ts +1 -1
  127. package/dist/reporters/testtracker.js +1 -1
  128. package/package.json +3 -2
@@ -0,0 +1,153 @@
1
+ ---
2
+ name: canary-failure-impact
3
+ description: >
4
+ For a given test, function, or code path, traces downstream effects and
5
+ produces a severity label. Investigates config/auth failures using the
6
+ consuming repo's declared user_catalog_skill. Optionally focuses on critical
7
+ paths when critical-areas.json is present.
8
+ ---
9
+
10
+ # Canary: Failure Impact
11
+
12
+ Answers "what actually breaks if this code fails and no test catches it?"
13
+ Produces a severity label and a concrete description of downstream effects to
14
+ help prioritise where to invest test coverage.
15
+
16
+ ## When to Use
17
+
18
+ - Before deciding which gap to close first: "which of these matters most?"
19
+
20
+ - When a test fails and you need to understand the blast radius
21
+
22
+ - As Phase 3 of `/canary-test-pipeline`
23
+
24
+ - When asked "what's the impact if this breaks?"
25
+
26
+ ## This skill vs. `canary guardian analyze`
27
+
28
+ This skill discovers downstream dependents with harness's `compute_blast_radius`
29
+ primitive when the MCP is present (degrading to plain `grep -r` when it is not),
30
+ then applies a **domain-keyword heuristic** (Steps 3–4 below) to turn that
31
+ dependent set into a severity label. The severity labeling is the heuristic part
32
+ — keyword matching over dependent file and function names.
33
+ `canary guardian analyze` (`ts/src/guardian/`, wired to
34
+ `canary guardian analyze` in `ts/src/cli.ts`) is a **real OpenAPI-diff
35
+ blast-radius engine** — it diffs two OpenAPI specs (`--spec-before` /
36
+ `--spec-after`), extracts the actual added/removed/changed endpoints, and maps
37
+ each to coverage gaps against a `coverage-report.json`. For the class of change
38
+ it covers, guardian is strictly higher-fidelity than the heuristics here.
39
+
40
+ - **Use `canary guardian analyze`** when the change is an API/schema change and
41
+ you have (or can generate) before/after OpenAPI specs — it gives exact
42
+ endpoint-level impact and coverage-gap data instead of a keyword guess.
43
+ - **Use this skill** for everything guardian doesn't cover: non-API code paths
44
+ (services, UI components, internal functions), impact tracing where no OpenAPI
45
+ spec exists, or when you need the broader billing/auth/compliance
46
+ domain-severity labeling in Step 3 rather than a strict API diff.
47
+
48
+ They are complementary, not competing — do not duplicate guardian's spec-diff
49
+ logic here if an OpenAPI change is in scope; delegate to
50
+ `canary guardian analyze` instead.
51
+
52
+ ## Input
53
+
54
+ Provide one of:
55
+
56
+ - A test file path: `tests/loyalty/points.spec.ts`
57
+
58
+ - A function name: `accruePoints`
59
+
60
+ - A code path: `src/loyalty/points.service.ts`
61
+
62
+ If `.canary/critical-areas.json` is present, focus tracing on paths with
63
+ `risk_score ≥ 0.7`.
64
+
65
+ ## Tracing Logic
66
+
67
+ ### Step 1 — Identify the code path
68
+
69
+ Resolve the input to a specific file and function. If ambiguous, ask before
70
+ proceeding.
71
+
72
+ ### Step 2 — Walk downstream dependents
73
+
74
+ **With harness MCP available:** call `compute_blast_radius` for the target file
75
+ (`file`, `mode: "detailed"`). It simulates cascading failure with a
76
+ probability-weighted BFS and returns each affected node with a cumulative
77
+ failure probability — this is the purpose-built blast-radius primitive, so use
78
+ it instead of hand-walking `get_relationships` hop-by-hop. Feed the returned
79
+ node set into Step 3, and let the cumulative probability weight the severity
80
+ (high-probability nodes dominate). When you additionally need the affected set
81
+ grouped by kind (tests vs docs vs code), call `get_impact` for the same target.
82
+
83
+ **Fallback:** use `grep -r` to find files that import or call the target. Limit
84
+ to direct dependents (1 hop) when MCP is unavailable.
85
+
86
+ ### Step 3 — Classify each dependent by domain
87
+
88
+ Apply these heuristics to the dependent paths and function names:
89
+
90
+ | Domain signal | Severity modifier |
91
+ | --------------------------------------- | ----------------------------------- |
92
+ | billing / payment / charge / invoice | +2 (financial impact) |
93
+ | auth / session / token / permission | +2 (security/access) |
94
+ | compliance / audit / PHI / PII / HIPAA | +2 (regulatory) |
95
+ | data / persist / write / store / commit | +1 (data integrity) |
96
+ | UI / render / display / format / label | −1 (user-facing only, no data risk) |
97
+
98
+ ### Step 4 — Aggregate to severity label
99
+
100
+ Base score starts at 2 (Medium). Sum modifiers from step 3. Cap at 4 (Critical).
101
+
102
+ | Score | Label |
103
+ | ----- | -------- |
104
+ | 5+ | Critical |
105
+ | 3–4 | High |
106
+ | 2 | Medium |
107
+ | 0–1 | Low |
108
+
109
+ ### Step 5 — User catalog investigation
110
+
111
+ When a test failure in the target path involves an auth, permission, or
112
+ configuration error:
113
+
114
+ 1. Read `user_catalog_skill` from `.canary/company.json`
115
+ 2. If present: invoke `canary skills run <user_catalog_skill>` with the required
116
+ attributes from the error context
117
+ 3. If a matching user/config is found: surface it as a suggestion
118
+ 4. If absent or no match: present constructively —
119
+
120
+ > "This failure may be a test user or test data configuration issue. Check
121
+ > your user catalog if you have one, or set up the required test data before
122
+ > re-running."
123
+
124
+ ## Output Format
125
+
126
+ ```text
127
+ Failure impact — src/loyalty/points.service.ts::accruePoints
128
+
129
+ Severity: HIGH
130
+
131
+ If this breaks undetected:
132
+ · Members see incorrect balance in the partner portal (user-facing)
133
+ · Points journal diverges from the ledger (data integrity)
134
+ · Downstream: redemption.service.ts · tier-upgrade.service.ts ·
135
+ reporting.service.ts (3 dependents)
136
+
137
+ Priority: write failure-path tests before next release
138
+ Suggested: /canary-write-test "test failure paths for accruePoints"
139
+ ```
140
+
141
+ ## Related skills
142
+
143
+ - `/canary-critical-areas` — produces `critical-areas.json` used for focus
144
+
145
+ - `/canary-ci-ready` — uses the same user-catalog investigation pattern
146
+
147
+ - `/canary-write-test` — generates tests for the identified high-impact gaps
148
+
149
+ - `/canary-test-pipeline` — Phase 3
150
+
151
+ - `canary guardian analyze` (CLI, `ts/src/guardian/`) — higher-fidelity
152
+ OpenAPI-diff blast-radius engine; use instead of this skill's heuristics when
153
+ the change is an API/schema change with before/after specs available
@@ -0,0 +1,15 @@
1
+ name: canary-failure-impact
2
+ version: '1.0.0'
3
+ description:
4
+ Trace the downstream blast radius of an undetected failure for a given test,
5
+ function, or code path and assign a severity label.
6
+ stability: static
7
+ triggers:
8
+ - manual
9
+ platforms:
10
+ - claude-code
11
+ type: rigid
12
+ tools: []
13
+ tier: 1
14
+ depends_on:
15
+ - canary-ci-ready
@@ -0,0 +1,197 @@
1
+ ---
2
+ name: canary-fleet-health
3
+ description: >
4
+ Fleet-wide test health summary across suites — flaky tests, failure spikes,
5
+ cross-suite common failures, and regression candidates from the run-history
6
+ store. Use when the user asks "how healthy is our test fleet", "fleet-wide
7
+ flake report", "any regressions this week", "failure spikes across suites", or
8
+ "canary analyze". Produces one compact, scannable summary — not a dashboard.
9
+ NOT for diagnosing a single known-flaky test (canary-flake-hunter) or scoring
10
+ one suite's CI readiness (canary-ci-ready).
11
+ ---
12
+
13
+ # Canary: Fleet Health
14
+
15
+ Wraps `canary analyze` (fleet-wide flake/spike/regression analytics) and the
16
+ run-history store to answer "how's the whole fleet doing?" in one chat-turn
17
+ summary. Today, [`canary-flake-hunter`](../../../canary-flake-hunter.md) only
18
+ diagnoses a single test you already suspect is flaky — this skill is the
19
+ fleet-wide counterpart: it tells you _where to look_ before you reach for the
20
+ hunter.
21
+
22
+ Per the adoption audit this skill implements (candidate #10), this is
23
+ deliberately a **low-cost validation step**: a compact text summary a human can
24
+ scan in one turn, not a dashboard or visual surface. If fleet-wide analytics
25
+ prove valuable, a richer surface is a separate, larger investment — don't
26
+ over-build this one.
27
+
28
+ ## When to Use
29
+
30
+ - Weekly/periodic health check: "how's the test fleet looking?"
31
+ - Before a release: "any regressions or spikes we should know about?"
32
+ - Triaging where to spend test-maintenance effort across many suites
33
+ - NOT for a single test you already know is flaky — use
34
+ [`canary-flake-hunter`](../../../canary-flake-hunter.md) to diagnose root
35
+ cause and propose a fix
36
+ - NOT for scoring one suite's CI readiness — use
37
+ [`canary-ci-ready`](../canary-ci-ready/SKILL.md) (coverage depth, assertion
38
+ quality, runtime for _one_ suite)
39
+ - NOT a substitute for `.canary/critical-areas.json` risk ranking — use
40
+ [`canary-critical-areas`](../canary-critical-areas/SKILL.md) for code-level
41
+ risk, this skill is history-data-level health
42
+
43
+ ## Process
44
+
45
+ ### Phase 1: RESOLVE THE STORE
46
+
47
+ Fleet analytics read from the run-history store, not from a live test run.
48
+ Before running anything, confirm data exists:
49
+
50
+ ```bash
51
+ canary history summary <suite> --runs 1
52
+ ```
53
+
54
+ - **Configured store:** `CANARY_HISTORY_DB_URL` env var (Supabase-backed) or
55
+ falls back to the local NDJSON file at
56
+ `test-results/reports/history-v2.jsonl`.
57
+ - **No local file and no `CANARY_HISTORY_DB_URL`:** there is nothing to analyze
58
+ yet. Say so plainly — "No run history found. Push results with
59
+ `canary history push` after a CI run, or run `canary history migrate` if you
60
+ have v1 history.jsonl data." Do not fabricate a health summary from nothing.
61
+
62
+ ### Phase 2: RUN THE RELEVANT ANALYSES
63
+
64
+ Default to the combined digest unless the user asked about one specific
65
+ dimension:
66
+
67
+ ```bash
68
+ canary analyze digest --json
69
+ ```
70
+
71
+ If the user asked about one thing specifically, run only that subcommand instead
72
+ of the full digest — cheaper and more focused:
73
+
74
+ | User asks about | Command |
75
+ | --------------------------------------- | -------------------------------------------------------------------- |
76
+ | Flaky tests fleet-wide | `canary analyze flaky --window-runs 30 --min-rate-pct 10 --json` |
77
+ | Failure spikes | `canary analyze spikes --delta-pp 20 --json` |
78
+ | CI-run flakes (reruns to green) | `canary analyze gh-flaky --repo <owner/name> --json` |
79
+ | Cross-suite common failures | `canary analyze common-failures --min-suites 2 --json` |
80
+ | Newly broken tests after a green streak | `canary analyze regression-candidates --json` |
81
+ | "Area health" / degrading areas | See the caveat below — this dimension does not currently return data |
82
+
83
+ **Known limitation — be upfront about it:** `canary analyze area-health` (and
84
+ the `area_health` section of `digest`) is currently wired to an empty data set
85
+ in `ts/src/analysis/cli.ts` / `ts/src/analysis/engine.ts` — it always reports
86
+ "No area health data available," regardless of history. Don't present this as a
87
+ working check; tell the user area-degradation tracking isn't implemented yet
88
+ rather than silently omitting it.
89
+
90
+ **Store-type caveat:** `flaky` queries the store directly and works with either
91
+ backend. `spikes`, `common-failures`, and `regression-candidates` currently only
92
+ populate fully when the backing store is the local NDJSON file (`AnalysisEngine`
93
+ special-cases `LocalHistoryStore` for suite discovery and per-test aggregation)
94
+ — with a Supabase-backed store (`CANARY_HISTORY_DB_URL` set), those three may
95
+ come back empty even with real history. If the digest shows all-zero
96
+ spikes/common-failures/ regressions _and_ `CANARY_HISTORY_DB_URL` is set, flag
97
+ this as a likely store-support gap, not a clean bill of health.
98
+
99
+ ### Phase 3: CONDENSE TO ONE SCREEN
100
+
101
+ Don't paste raw Markdown tables from the CLI — they're built for file artifacts,
102
+ not chat. Pull the top 3–5 rows per section and compress to the Output Format
103
+ below. If a section is empty, say "none" in one line; don't render an empty
104
+ table.
105
+
106
+ ### Phase 4: SURFACE THE ONE ACTIONABLE THING
107
+
108
+ Look across sections for correlation — e.g., a suite with both a recent spike
109
+ and several tests over the flake threshold is a stronger signal than either
110
+ alone. Call that out explicitly as the suggested next step, and name the
111
+ specific downstream skill:
112
+
113
+ - Single suspicious test → point at
114
+ [`canary-flake-hunter`](../../../canary-flake-hunter.md)
115
+ - Whole suite trending down → point at
116
+ [`canary-ci-ready`](../canary-ci-ready/SKILL.md) for that suite
117
+ - Systemic cross-suite pattern (e.g. the same connection error in 3 suites) →
118
+ this is infrastructure/environment, not a test bug; say so instead of
119
+ suggesting a test fix
120
+
121
+ ## Output Format
122
+
123
+ ```text
124
+ Fleet Health — window: 30 runs
125
+
126
+ Flaky (≥10%): 3 tests top: checkout_retry_test (32%, api suite)
127
+ Spikes (≥20pp): 1 suite e2e_ui +25pp since 2026-07-10
128
+ Area health: not available (not yet implemented — always empty)
129
+ Common failures: 1 pattern "ECONNREFUSED 127.0.0.1:5432" across 2 suites
130
+ Regressions: 2 tests orders_post_201 broke after 12-run green streak
131
+
132
+ Suggested next step: e2e_ui's spike + 2 of the 3 flaky tests are in that
133
+ suite — investigate the suite before chasing individual flakes.
134
+ Run /canary-ci-ready on e2e_ui, or /canary-flake-hunter on
135
+ checkout_retry_test for a root-cause fix.
136
+ ```
137
+
138
+ Keep it to one screen. Omit a line entirely rather than padding with "no data"
139
+ noise for sections that were never requested.
140
+
141
+ ## Flags
142
+
143
+ - `--window-runs <runs>` — rolling window measured in RUNS, not days (default:
144
+ 30, passed through to `canary analyze`)
145
+ - `--min-rate-pct <percent>` — minimum flake rate to report, on a 0-100 percent
146
+ scale (default: 10, so `10` means 10%, not 0.1)
147
+ - `--delta-pp <points>` — spike threshold as a PERCENTAGE-POINT rise in failure
148
+ rate (default: 20, so `20` means 5% → 25%, not 5% → 6%)
149
+ - `--suite <name>` — scope to one suite instead of the whole fleet
150
+ - `--json` — pass through the underlying CLI's `--json` when the caller wants
151
+ structured data instead of a chat summary
152
+
153
+ The unitless spellings `--window`, `--delta`, and `--min-rate` are deprecated
154
+ aliases: they still work and take the same values, but they print a note on
155
+ stderr naming the replacement. Prefer the unit-bearing names in anything you
156
+ write down.
157
+
158
+ ## Error Handling
159
+
160
+ | Situation | What To Do |
161
+ | ------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
162
+ | No history data at all (fresh repo) | Say so; point at `canary history push` / `canary history migrate`. Don't run the analysis commands against nothing. |
163
+ | `CANARY_HISTORY_DB_URL` set but Supabase unreachable | `make_store` only guards against a missing `agent.history.supabase_store` import, not a live connection failure — a query error will surface as a CLI exception. Report the exception text, suggest checking connectivity/credentials, don't retry silently. |
164
+ | `area-health` requested explicitly | Explain the known limitation (Phase 2) rather than showing an empty table with no context. |
165
+ | Digest looks suspiciously all-zero with `CANARY_HISTORY_DB_URL` set | Flag the store-type caveat (Phase 2) before concluding the fleet is healthy. |
166
+ | User wants a specific suite that has no history | `canary history summary <suite>` returns `total_runs: 0` — report that directly instead of running the full fleet analysis. |
167
+
168
+ ## Examples
169
+
170
+ ### Example: Weekly health check with a correlated signal
171
+
172
+ **Prompt:** "How's the fleet looking this week?"
173
+
174
+ **Action:** Confirm local history file exists. Run `analyze digest --json`.
175
+ Condense: 3 flaky tests (2 in `e2e_ui`), 1 spike (`e2e_ui`, +25pp), no common
176
+ failures, 2 regression candidates. Correlate: `e2e_ui` shows up in both flaky
177
+ and spikes — call it out as the priority, suggest `canary-ci-ready` on `e2e_ui`
178
+ before chasing the individual flaky tests.
179
+
180
+ ### Example: No history yet
181
+
182
+ **Prompt:** "Give me a fleet health summary."
183
+
184
+ **Action:** `history summary api --runs 1` returns `total_runs: 0` and no local
185
+ NDJSON file exists. Report plainly: no run history is available yet; point at
186
+ `canary history push` after the next CI run. Do not run `analyze digest` against
187
+ an empty store and present an empty report as "all clear" — absence of data is
188
+ not evidence of health.
189
+
190
+ ## Related Skills
191
+
192
+ - [`canary-flake-hunter`](../../../canary-flake-hunter.md) — single-test
193
+ root-cause diagnosis once fleet health points at a candidate
194
+ - [`canary-ci-ready`](../canary-ci-ready/SKILL.md) — single-suite CI readiness
195
+ scoring (coverage, assertions, runtime)
196
+ - [`canary-critical-areas`](../canary-critical-areas/SKILL.md) — code-level risk
197
+ ranking, a different signal from history-based health
@@ -0,0 +1,185 @@
1
+ ---
2
+ name: canary-generate-test
3
+ description: >
4
+ Generate a framework-appropriate test from a natural-language requirement by
5
+ routing through Canary's classify → recommend → generate pipeline (the
6
+ `/canary-write-test` slash command), writing the test under `tests/generated/`
7
+ and optionally executing it. Use for "write a test for X", "I need an API test
8
+ that does Y", "scaffold a new test from this description", or triaging a bug
9
+ report into a regression test — when a CLI batch pipeline run (not an
10
+ interactive session-generated file) is what's wanted. See
11
+ `agents/canary-test-author.md` (interactive, session-generated, wired to
12
+ `/canary-write-test`) and `agents/canary-test-generator.md` (MCP
13
+ write_test_file retry loop) for the other two "write a test" paths. Not for
14
+ editing existing tests or choosing between frameworks abstractly.
15
+ ---
16
+
17
+ # Canary: Generate Test
18
+
19
+ > Generate a framework-appropriate test from a natural-language requirement.
20
+ > Routes through Canary's classify → recommend → generate pipeline, writes the
21
+ > test under `tests/generated/`, and optionally executes it.
22
+
23
+ ## When to Use
24
+
25
+ - When the user asks to scaffold a new test from a natural-language description
26
+ ("write a test for X", "I need an API test that does Y")
27
+ - When triaging a bug report into a regression test
28
+ - When extending an existing test suite with a new case but the framework choice
29
+ is ambiguous
30
+ - NOT for editing existing tests — use the project's normal edit flow
31
+ - NOT for choosing between frameworks abstractly — use the framework-registry
32
+ docs, not this skill
33
+ - NOT for running existing tests — use the framework's CLI directly
34
+
35
+ ### Relative to the other "write a test" paths
36
+
37
+ This skill is the **batch generation path** (the `/canary-write-test` slash
38
+ command) — use it for scripted/CI-driven generation runs consumed
39
+ programmatically from `tests/generated/`. For an interactive, session-generated
40
+ test with a human reviewing framework and code before it lands, use
41
+ `agents/canary-test-author.md` (wired to `/canary-write-test`). For a
42
+ single-file, automatic write-run-revise loop via MCP tools, use
43
+ `agents/canary-test-generator.md`.
44
+
45
+ ## Process
46
+
47
+ ### Phase 1: CLARIFY — Resolve the Requirement
48
+
49
+ 1. **Confirm test type if ambiguous.** If the prompt doesn't clearly indicate
50
+ `unit | api | e2e | performance`, ask one targeted question before invoking
51
+ the pipeline. The classifier will guess, but a wrong guess costs a
52
+ regeneration.
53
+ 2. **Confirm target framework if the user has a preference.** The recommender
54
+ will pick by category from the registry; if the user explicitly wants
55
+ Playwright but the registry would pick Cypress for the category, surface that
56
+ mismatch now.
57
+ 3. **Capture concrete inputs.** Endpoint URL, payload shape, expected status,
58
+ selectors, performance thresholds — whatever the test actually needs. Vague
59
+ prompts produce vague tests.
60
+
61
+ ### Phase 2: GENERATE — Run the Pipeline
62
+
63
+ 1. **Invoke generation** by running the `/canary-write-test` slash command in
64
+ Claude Code with the requirement as its prompt.
65
+
66
+ 2. **Read the printed classification + recommendation.** Verify the resolved
67
+ `test_type` and `framework` match intent. If they don't, refine the prompt
68
+ and re-run — do not hand-edit the generated file to compensate for a
69
+ misclassification.
70
+ 3. **Locate the output.** The CLI prints an absolute path under
71
+ `tests/generated/<category>/`. The orchestrator return dict's `output_path`
72
+ is authoritative.
73
+
74
+ ### Phase 3: VALIDATE — Execute or Dry-Run
75
+
76
+ 1. **Run the generated test** with the framework's CLI (or pass `--execute` to
77
+ the generator). For api/e2e tests, run against a known-good environment
78
+ first.
79
+ 2. **If execution fails, classify the failure:**
80
+ - **Generation error** (syntax, wrong API shape) → regenerate with a more
81
+ specific prompt; don't hand-fix unless trivial
82
+ - **Environment error** (missing creds, wrong base URL) → fix the env, rerun
83
+ - **Real assertion failure** (the SUT behaves differently than the prompt
84
+ asserted) → this is a useful signal; review with the requester before
85
+ changing the test
86
+ 3. **Log the run.** Append the requirement, classification, framework, and
87
+ pass/fail to `docs/CANARY_STATE.md` so downstream sessions can pick up
88
+ context.
89
+
90
+ ### Phase 4: PROMOTE — Move from Generated to Committed
91
+
92
+ If the generated test passes review and belongs in the committed suite, use the
93
+ [`canary-promote-test`](../canary-promote-test/SKILL.md) skill. Promotion is its
94
+ own workflow — don't collapse it into this one.
95
+
96
+ ## Canary Integration
97
+
98
+ - **`/canary-write-test "<prompt>"`** — Primary entry. Runs the full classify →
99
+ recommend → generate pipeline in the Claude Code session and can immediately
100
+ run the generated test.
101
+ - **`CanaryOrchestrator.run(prompt, execute=False)`** — Programmatic entry.
102
+ Returns the structured pipeline-result dict.
103
+ - **`ts/src/data/frameworks/registry.json`** — Maps `test_type` → framework.
104
+ Edit here when adding a new framework; never hard-code framework choices in
105
+ callers.
106
+ - **Generation runs in your Claude Code session** — there is no LLM provider
107
+ layer or API key to configure (that layer was removed in v3.0).
108
+
109
+ ## Success Criteria
110
+
111
+ - The generated file parses and runs under its framework's CLI
112
+ - The classifier's `test_type` matches the user's actual intent
113
+ - The recommender's framework choice resolves from the registry (no nulls)
114
+ - The execution result is captured in the return dict (when `execute=True`)
115
+ - A promoted test passes review and runs cleanly in CI
116
+
117
+ ## Rationalizations to Reject
118
+
119
+ | Rationalization | Why It Is Wrong |
120
+ | ----------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------- |
121
+ | "The classifier picked the wrong type but I'll hand-fix the output" | The hand-fix masks a real classifier gap. Refine the prompt or file a classifier issue — don't paper over routing bugs in the generated file. |
122
+ | "I'll commit the generated test as-is, it's good enough" | Generated tests live in `tests/generated/` for a reason — they're unreviewed scratch. Promote intentionally. |
123
+ | "The registry doesn't have an entry for this test_type, I'll add `framework: null`" | Null breaks the contract. Every `test_type` must map to a framework. Add a real entry or change the classifier output. |
124
+ | "I'll skip the validation phase, the test looks right" | LLM output that looks right and runs are different things. Always execute (or dry-run) before promoting. |
125
+
126
+ ## Examples
127
+
128
+ ### Example: API test for a known endpoint
129
+
130
+ **Prompt:** `Test that POST /v1/orders returns 201 with a valid payload`
131
+
132
+ **Pipeline trace:**
133
+
134
+ ```text
135
+ Classification: intent=generate_tests, test_type=api, confidence=0.85
136
+ Recommendation: framework=requests-pytest, ext=.py, category=api
137
+ Output: tests/generated/api/orders_post_201.py
138
+ Execution: returncode=0, 1 passed in 0.42s
139
+ ```
140
+
141
+ **Action:** Review the generated file, promote to
142
+ `tests/api/orders_post_201.py`, drop the timestamped header.
143
+
144
+ ### Example: Ambiguous prompt — clarification first
145
+
146
+ **Prompt:** `Test the new orders feature`
147
+
148
+ **Action:** Do NOT invoke the pipeline yet. Ask: "Is this an end-to-end UI test
149
+ of the checkout flow, an API contract test for `/v1/orders`, or a unit test of
150
+ the order-validation function?" Only after the user picks should you run
151
+ `generate`.
152
+
153
+ ### Example: Performance test with thresholds
154
+
155
+ **Prompt:** `Load test /v1/search at 200 RPS for 5 minutes, p95 latency < 300ms`
156
+
157
+ **Pipeline trace:**
158
+
159
+ ```text
160
+ Classification: test_type=performance, confidence=0.95
161
+ Recommendation: framework=k6, ext=.js, category=performance
162
+ Output: tests/generated/performance/search_load.js
163
+ ```
164
+
165
+ **Validation:** Run against a staging environment, not prod. Compare p95 to the
166
+ threshold; if the test passes locally but the threshold was unrealistic, surface
167
+ that to the requester before promoting.
168
+
169
+ ## Escalation
170
+
171
+ - **When the registry has no entry for the classified `test_type`:** Stop and
172
+ file a registry update. Do not invent a framework name.
173
+ - **When the generated test repeatedly fails to parse:** This usually means the
174
+ prompt is under-specified. Narrow it to a single behavior and regenerate
175
+ before assuming a bug.
176
+ - **When execution requires creds you don't have:** Surface the missing-cred
177
+ error to the user; never embed dummy creds in a generated test to make it
178
+ "run".
179
+ - **When the classifier's confidence is below 0.7:** Treat as a clarification
180
+ trigger, not a generation trigger. Loop back to Phase 1. Note: `confidence` is
181
+ a hand-calibrated heuristic prior per keyword-match branch in
182
+ `ts/src/core/classifier.ts`, not a statistically calibrated probability — 0.7
183
+ is a coarse "does this branch's signal look weak" cutoff, not P(correct
184
+ classification) ≥ 0.7. Don't read finer distinctions (e.g. 0.85 vs. 0.95) as
185
+ meaningfully different confidence levels.