blastproof 0.3.0 → 0.4.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -4,74 +4,51 @@
4
4
  [![Dogfood](https://github.com/hamc/blastproof/actions/workflows/dogfood.yml/badge.svg)](https://github.com/hamc/blastproof/actions/workflows/dogfood.yml)
5
5
  [![License: MIT](https://img.shields.io/badge/License-MIT-blue.svg)](./LICENSE)
6
6
 
7
- **Open-source AI testing agent for pull requests.** Diff in, confidence out — no test scripts to write or maintain.
8
-
9
- `blastproof` is an open-source AI QA agent for pull requests: it reads your PR diff, maps the blast radius, generates end-to-end tests in plain English, executes them on a real browser with self-healing, and scores the result before merge. 100% local, MIT licensed, bring your own LLM key.
7
+ **Open-source AI testing agent for pull requests.** Write end-to-end tests as plain English. An agent drives a real browser to run them, selects only the ones your diff can affect, and scores the result before merge.
10
8
 
11
9
  ```
12
10
  git diff → impact mapping → test generation → agentic execution → report + score
13
11
  ```
14
12
 
15
- ## How it works
16
-
17
- 1. **Reads the diff** — `blastproof test --base main` parses the branch diff and maps it to affected routes.
18
- 2. **Maps the blast radius** — matches changed files against the `routes:` map you maintain, to select the journeys worth running. Deterministic and free; see [Closing the coverage hole](#closing-the-coverage-hole) for what that catches and what it does not.
19
- 3. **Writes the missing tests** — drafts plain-English YAML for affected routes nothing covers. Drafts are yours to review; existing tests are never rewritten.
20
- 4. **Executes agentically** — an LLM-driven loop over Playwright resolves elements via the accessibility tree on every step. No static selectors to rot: the agent re-resolves when the UI shifts. Being model-driven it is not deterministic, which is why the free, deterministic impact analysis is what runs on every pull request.
21
- 5. **Reports & scores** — console, JUnit XML and HTML reports, plus a priority-weighted score that fails the run below `--min-score`.
22
-
23
- ## Does this fit your application?
24
-
25
- Read this before installing. The agent drives your app the way a screen reader would, and that has consequences worth knowing in advance rather than discovering on a failing run.
26
-
27
- **Your markup has to be accessible. This is a hard requirement, not a nice-to-have.** Every element is found by role, label or visible text, resolved from the page's accessibility tree on each attempt — that is what removes selectors and what lets the agent survive a redesign. The cost is that an interface the accessibility tree cannot describe is one the agent cannot drive at all. Icon-only buttons with no accessible name, custom dropdowns and comboboxes without ARIA roles, and `div`-based controls do not degrade into something slower; they simply cannot be targeted, and there is deliberately no CSS or XPath escape hatch to fall back on. If you are unsure, run your app through any accessibility checker first — the result predicts how well blastproof will do better than anything else.
28
-
29
- **Some interactions are not supported yet:**
30
-
31
- | Not supported | Consequence |
32
- |---|---|
33
- | `iframe` content | Anything inside an embedded frame is invisible to the agent, including the page snapshot. **Hosted payment widgets such as Stripe Elements or PayPal buttons fall in here** — if your checkout embeds one, the checkout journey cannot be driven end to end today |
34
- | Hover, scroll-to, drag and drop | Menus that open only on hover, and content that loads only once scrolled into view, are out of reach |
35
- | File upload | No action exists for it |
36
- | Multiple tabs or windows | The agent works in one page; a flow that opens a popup loses it |
37
- | Native `alert` / `confirm` dialogs | Not handled |
38
-
39
- **Large pages are truncated.** The accessibility snapshot sent to the model is capped at 200 lines. A dense admin table or a long product listing will have its lower content invisible on every step.
40
-
41
- **The deterministic half stands on its own.** If none of the above fits, `blastproof run --impacted --dry-run` still works: it reads the diff, reports which routes a change touches, which files are classified by nothing, and which affected routes no test covers. That needs no API key, no browser and no model, costs nothing per pull request, and is useful purely as a coverage-gap report — see [Closing the coverage hole](#closing-the-coverage-hole).
13
+ 100% local. MIT. Bring your own LLM key.
42
14
 
43
15
  ## Quick start
44
16
 
45
17
  ```bash
46
- npm install -g blastproof # requires Node.js >= 20.19
47
- npx playwright install --with-deps chromium # one-time browser download
48
-
18
+ npm install -g blastproof # Node.js >= 20.19
49
19
  cd your-project
50
- blastproof init # scaffolds .blastproof/
51
- # start your app, then point base_url at it in .blastproof/config.yaml
52
- export ANTHROPIC_API_KEY=... # or OPENAI_API_KEY, or a local Ollama model
53
- blastproof run # runs .blastproof/tests/**/*.yaml agentically
20
+ blastproof init # scaffolds .blastproof/
54
21
  ```
55
22
 
56
- `init` scaffolds one smoke test that works against any app, plus a commented login template you rename and edit once it describes *your* login. Nothing is written that assumes anything about your application.
23
+ Point `base_url` at your running app in `.blastproof/config.yaml`, then check the setup — this needs **no API key and no browser**:
57
24
 
58
- Published with [npm provenance](https://docs.npmjs.com/generating-provenance-statements), so the tarball is verifiably built from this repository.
25
+ ```bash
26
+ blastproof run --dry-run
27
+ ```
59
28
 
60
- Try it without your own app — the demo shop lives in this repository, so clone it:
29
+ To actually execute tests you need a browser and a model:
61
30
 
62
31
  ```bash
63
- git clone https://github.com/hamc/blastproof && cd blastproof
64
- npm install && npm run build
65
- node examples/demo-app/serve.mjs 4173 & # home, login, cart, checkout, orders
66
- export ANTHROPIC_API_KEY=...
67
- node dist/cli.js run
32
+ npx playwright install --with-deps chromium # NEEDS SUDO — see below
33
+ export ANTHROPIC_API_KEY=... # or OPENAI_API_KEY, or local Ollama
34
+ blastproof run
68
35
  ```
69
36
 
70
- > **Status:** the pipeline is complete and published. `init`, `run` (including `--impacted`), `plan` and `test`, with authentication, JUnit and HTML reports, a merge gate, and a GitHub Action. Pre-1.0: the command surface may still change.
37
+ **No sudo?** `--with-deps` installs system libraries as root. Without it, run `npx playwright install chromium` and obtain `libnspr4`, `libnss3`, `libnssutil3` and `libasound2` however you can. Note that a useful half of blastproof needs neither browser nor key — see [Without a browser or a key](#without-a-browser-or-a-key).
38
+
39
+ Before `run`, `plan` or `test` do anything, they check what they are about to spend — the browser can launch, the model provider is reachable, `base_url` responds — and report every unmet one together, so a stopped app or a missing browser is never a wall you hit one crash at a time. A missing system library names the exact install command and says it needs root; nothing is installed on your behalf. Silent when everything is fine, and skipped entirely by `--dry-run`, which needs none of it.
40
+
41
+ ## Does this fit your application?
42
+
43
+ **Your markup must be accessible — a hard requirement.** Elements are found by role, label or visible text from the accessibility tree. That is what removes selectors and survives redesigns; the cost is that an interface the accessibility tree cannot describe cannot be driven at all, and there is deliberately no CSS or XPath fallback. Icon-only buttons without accessible names, `div`-based controls and ARIA-less dropdowns simply cannot be targeted. Run an accessibility checker first — the result predicts how well this will work better than anything else.
44
+
45
+ **Not supported yet:** `iframe` content (so hosted payment widgets like Stripe Elements are invisible — an embedded checkout cannot be driven end to end), hover, scroll-to, drag and drop, file upload, multiple tabs, native `alert`/`confirm` dialogs. Page snapshots are capped at 200 lines by default, so very dense pages are truncated — raise it with `browser.max_snapshot_lines` if your pages need more; truncation is always marked in the snapshot so the model is never misled into thinking it saw the whole page.
46
+
47
+ `browser.timeout_ms` bounds every wait — resolving a target element from the accessibility tree, and navigation — not only the click or fill performed afterwards. Raise it for an application that is merely slow to hydrate; the trade-off is that a genuinely missing element then takes longer to fail. It never changes how many self-healing retries a step gets — waiting and retrying are deliberately separate.
71
48
 
72
- ## Test format
49
+ ## Writing tests
73
50
 
74
- Tests live in `.blastproof/tests/` as plain-English YAML — no selectors, no framework lock-in:
51
+ Tests live in `.blastproof/tests/` as plain-English YAML — no selectors:
75
52
 
76
53
  ```yaml
77
54
  summary: Checkout with discount
@@ -85,179 +62,36 @@ steps:
85
62
  - complete checkout
86
63
  ```
87
64
 
88
- `priority` is P0–P2 (default P1); `tags`, `setup` steps and `auth` are optional (`auth: false` runs the test signed out — a login test needs it). `routes` (optional) declares the URLs a test covers: `blastproof run --impacted` runs only tests whose `routes:` intersect the routes affected by your PR diff (mapped from changed files via the `routes:` globs in `.blastproof/config.yaml`). Route strings compare by exact equality (`/cart` ≠ `/cart/` — write them consistently). Tests without `routes:` are skipped and reported under `--impacted`.
89
-
90
- ## CLI
91
-
92
- | Command | Description |
93
- | --- | --- |
94
- | `blastproof init` | Scaffold `.blastproof/` config and sample tests (idempotent) |
95
- | `blastproof run [--tag smoke] [--priority P0] [--query checkout]` | Run tests only — exit 0 pass, 1 fail, 2 usage/config error |
96
- | `blastproof run --impacted [--base <ref>]` | Run only tests impacted by the diff vs the base ref (default `main`). Unrouted tests are skipped and reported; affected-but-uncovered routes are reported without failing the run |
97
- | `blastproof run --dry-run` | Print the selection plan (affected routes, unmapped files, selected/skipped tests) and exit 0 — no browser launched, no LLM key needed |
98
- | `blastproof run --url <url>` | Override `base_url` for this run only (e.g. a PR preview environment); the config file is never mutated |
99
- | `blastproof plan [--base <ref>]` | Generate plain-English tests for affected routes no test covers yet. Prints drafts; nothing is written without `--write` |
100
- | `blastproof plan --route <route>` | Generate for a route explicitly, skipping the diff (repeatable) — how you bootstrap coverage on an app with no suite yet |
101
- | `blastproof plan --write` | Persist drafts to `.blastproof/tests/<route-slug>.yaml`. Never overwrites: a colliding filename fails that route |
102
- | `blastproof run --min-score <n>` | Require a weighted score of at least `n` (0–100). **Replaces** the all-must-pass rule — see below |
103
- | `blastproof run --junit [path]` | Write a JUnit XML report; without a path it lands in `.blastproof/reports/<session>/junit.xml` |
104
- | `blastproof run --html [path]` | Write a self-contained HTML report with failure screenshots embedded inline |
105
- | `--max-llm-calls <n>` \| `--max-tokens <n>` \| `--max-duration <seconds>` | Bound `run`, `plan` or `test` by model calls, tokens or wall-clock time — every command that can call the model accepts these. Overrides `.blastproof/config.yaml`'s `budget:` section — see [Bounding a run](#bounding-a-run-budget-and-deadline) |
106
- | `blastproof test [--base <ref>]` | The full pipeline: run the tests covering the diff, then draft tests for the gaps |
107
-
108
- ### The full pipeline: `blastproof test`
109
-
110
- One command for the whole loop — map the blast radius, run what covers it, draft what doesn't exist yet:
111
-
112
- ```bash
113
- blastproof test --base main --min-score 80 --junit junit.xml --html report.html
114
- ```
115
-
116
- It does two things and reports them separately:
65
+ `priority` is P0–P2 (default P1). `tags`, `setup` steps and `auth` are optional — `auth: false` runs the test signed out, which a login test needs. `routes` declares the URLs a test covers, which is what `--impacted` selects on; write route strings consistently, since they compare by exact equality (`/cart` ≠ `/cart/`).
117
66
 
118
- 1. **Verify** — executes the tests covering the affected routes, scores them, applies the gate
119
- 2. **Draft** — generates tests for affected routes no test covers, and prints them
120
-
121
- **Generated drafts are never executed, and never affect the score.** That is deliberate. An unreviewed, model-written test in the merge path fails in two directions: a hallucinated expectation blocks a correct pull request, and a credulous one waves a broken change through while looking like coverage. Either costs more trust than the automation saves.
122
-
123
- So read the result honestly: `test` does not make an uncovered route safe. It makes the gap visible, with a draft ready for you to review, and the score keeps describing only what was actually verified. Add `--write` to persist the drafts (never overwriting an existing file) and commit them once you have read them.
124
-
125
- Exit codes: 2 usage/config/diff, 1 when the gate fails **or** a draft could not be generated, 0 otherwise.
126
-
127
- ### Reports
128
-
129
- `--junit` is for CI; `--html` is for humans. The HTML report is a single self-contained file — inline CSS, screenshots embedded as data URIs, no scripts — so it opens offline, survives being moved, and uploads as one artifact. It leads with the score and gate verdict, sorts failures above passes, and expands each failure to its failing step, reason and screenshot.
130
-
131
- ### Generating tests with `plan`
132
-
133
- `plan` closes the gap `run --impacted` reports. It takes the affected routes no test covers, loads each one in the browser, and asks the model to write a test from the page's real accessibility tree plus the changed files that made the route impacted — so the generated steps name controls that actually exist:
67
+ ## Commands
134
68
 
135
69
  ```bash
136
- blastproof plan --base main # preview drafts for uncovered routes
137
- blastproof plan --base main --write # persist them, then review and commit
138
- blastproof plan --route /checkout # bootstrap a route without a diff
70
+ blastproof init # scaffold .blastproof/
71
+ blastproof run # run every test
72
+ blastproof run --impacted --base main # run only what the diff can affect
73
+ blastproof plan --base main # draft tests for uncovered routes
74
+ blastproof test --base main # run what covers the diff, then draft the gaps
139
75
  ```
140
76
 
141
- Drafts are **previews by default** — nothing touches disk until `--write`, and `--write` never overwrites an existing file, so a regeneration can't silently replace a test you edited by hand. Each written file carries a header recording its route, base ref and generation date. Review before committing: the steps are model-written and meant to be edited.
142
-
143
- Exit codes: 0 when every route generated (or nothing needed coverage), 1 when a route failed or the budget/deadline stopped generation early, 2 on usage/config/diff errors. A route that fails to load never aborts the others; a budget or deadline stop does — see [Bounding a run](#bounding-a-run-budget-and-deadline).
144
-
145
- `plan` uses the same `auth:` recipe as `run`, so a route behind a login is drafted from the real page rather than from the login wall.
146
-
147
- ### Closing the coverage hole
77
+ Common flags — `blastproof <command> --help` has the full list:
148
78
 
149
- Impact mapping has a failure mode worth understanding. A changed file that matches no `routes:` glob contributes no affected routes — so a diff touching only a shared module selects nothing, scores 100 because nothing executed, and merges green. The information is printed, but nobody reads a passing run.
150
-
151
- Each changed file is classified three ways:
152
-
153
- | a changed file | means |
79
+ | flag | |
154
80
  | --- | --- |
155
- | matches a `routes:` glob | contributes its routes |
156
- | matches an `ignore:` glob | knowingly irrelevant to any page |
157
- | matches neither | **nobody has said what this affects** |
158
-
159
- ```yaml
160
- routes:
161
- "src/cart/**": ["/cart", "/checkout"]
162
- ignore:
163
- - "**/*.md"
164
- - ".github/**"
165
- ```
166
-
167
- ```bash
168
- blastproof run --impacted --fail-on-unmapped
169
- ```
170
-
171
- The flag fails the run on the third case only, naming the files and both ways to resolve them. `ignore:` is what makes that signal survivable — without it the flag would fire on every README edit and get switched off within a day, and a disabled gate protects nothing.
172
-
173
- Nothing is ignored by default, on purpose: a file nobody has classified is exactly the risk the flag exists to surface, and a default that guesses on your behalf would hide the first files worth thinking about. This flag is **additive** — a run can meet its `--min-score` and still be blocked here, because "the tests I ran passed" and "something changed that nobody has classified" are different claims.
174
-
175
- Be clear on its limit: it catches files that are *unclassified*, not files that are *misclassified*. A shared module mapped to one route when it can break five will still slip through. Mapping by import graph is the answer to that, and blastproof does not do it yet.
176
-
177
- ### Score and merge gating
178
-
179
- Every run ends with a score: the percentage of executed test **weight** that passed, where a test weighs 3 at P0, 2 at P1 and 1 at P2. Weighting is the point — a failing checkout costs three times a failing tooltip, so a pile of trivial passes can't hide a broken critical journey.
180
-
181
- ```bash
182
- blastproof run # any failure exits 1 (strict, the default)
183
- blastproof run --min-score 80 # passes at 80+, so one failing P2 is tolerated
184
- blastproof run --min-score 100 # identical to the default strict behaviour
185
- ```
81
+ | `--dry-run` | Print the selection (or, for `plan`, the routes it would draft) and exit. No browser, no API key |
82
+ | `--tag` · `--priority` · `--query` | Select a subset of tests |
83
+ | `--url <url>` | Override `base_url` for this run (e.g. a PR preview) |
84
+ | `--min-score <n>` | Gate on a weighted score instead of all-must-pass |
85
+ | `--fail-on-unmapped` | Fail when a changed file matches no `routes:` or `ignore:` glob |
86
+ | `--junit [path]` · `--html [path]` | Write reports |
87
+ | `--write` | `plan` only — persist drafts instead of previewing |
88
+ | `--max-llm-calls` · `--max-tokens` · `--max-duration` | Bound what a run may spend |
186
89
 
187
- `--min-score` **replaces** the all-must-pass rule rather than adding to it. Without it, any failure exits 1. With it, the score alone decides — which is what lets you say "a P2 may break, a P0 may not" in one number. Only executed tests count: tests removed by `--tag`/`--priority`/`--query`, and tests skipped as unrouted under `--impacted`, are neither numerator nor denominator. A run that executed nothing scores 100 (the output says so explicitly), so a docs-only PR is never blocked.
90
+ Exit codes: **0** pass, **1** the gate failed, **2** usage or config error.
188
91
 
189
- For CI:
92
+ **Generated drafts are never executed and never affect the score.** An unreviewed model-written test in the merge path fails in two directions: a hallucinated expectation blocks a correct PR, and a credulous one waves a broken change through while looking like coverage. `plan` makes the gap visible with a draft to review; it does not make an uncovered route safe.
190
93
 
191
- ```bash
192
- blastproof run --impacted --base "$BASE_REF" --min-score 80 --junit junit.xml
193
- ```
194
-
195
- Exit 0 merge-able, 1 blocked, 2 usage/config error. The JUnit report carries the score as a `<property name="score">` so a parser can read it without scraping stdout, and tests skipped for having no `routes:` appear as `<skipped/>` cases — the coverage gap shows up in CI instead of vanishing.
196
-
197
- ## Bounding a run: budget and deadline
198
-
199
- Nothing stops a run by default — a suite runs to completion or the provider refuses. `budget:` puts a ceiling on it, in `.blastproof/config.yaml`. It applies to `run`, `plan` and `test` alike: every model call any of them makes — agent action, assert judgment, or test planning — is counted, because a budget that only covered `run` would leave `plan`'s calls unbounded.
200
-
201
- ```yaml
202
- budget:
203
- max_llm_calls: 500 # stop after this many model calls
204
- max_tokens: 2000000 # stop after this many tokens spent across all calls
205
- max_duration_s: 900 # stop after this many seconds of wall-clock time
206
- ```
207
-
208
- Each limit is independent and optional; a config with no `budget:` section — or a run with none of `--max-llm-calls` / `--max-tokens` / `--max-duration` — behaves exactly as before. All three are counted in **calls and tokens, not currency**: a price table keyed by model and provider goes stale the day a provider reprices, and a limit that silently stops meaning what it says is worse than no limit, because it is trusted. Calls and tokens are exact, already reported by every provider, and yours to convert to a dollar figure with your own rates if you want one.
209
-
210
- `--dry-run` reports the ceiling before you spend anything — the worst case a selection could cost, computed from step counts alone, no provider contacted:
211
-
212
- ```
213
- Dry run: 12 test(s) selected, base_url=http://localhost:4173
214
- Worst case: up to 216 model call(s) for this selection (a maximum, not a prediction).
215
- ```
216
-
217
- That number is a ceiling, not a forecast — a real run almost always finishes in a fraction of it, because most steps complete long before either cap runs out. Per step it is the iteration cap **plus** `max_retries_per_step` (read from your config, not assumed): a malformed model response is retried without spending an iteration, and a failing `assert` spends a retry *and* an iteration in the same call — so the two caps are added, not one doubled and the other ignored. If `auth.steps` is configured, the login journey's steps are counted too — it runs once before any test and spends model calls through the same loop, so a ceiling that excluded it could be exceeded by the very first run that logs in.
218
-
219
- **Exhausting a budget stops the run — it does not fail a test.** Running out of quota says nothing about the application under test, so recording it as a failure would manufacture a defect that does not exist. Tests the run never reached are reported as **not run**, a third state distinct from passed and failed, and excluded from the score's denominator entirely — counting them as failures would just be a quieter version of the same lie a false pass would have been.
220
-
221
- An interrupted run is unmistakably incomplete: the process **exits 1 unconditionally**, even when `--min-score` is given and the tests that did execute would have satisfied it. The tests that finished are whichever ones happened to run first, not a representative sample, so nothing about them is a verdict:
222
-
223
- ```
224
- Run incomplete: model call budget exhausted: reached the configured maximum of 500 call(s)
225
- Score over executed tests: 92 (not a verdict — exit code 1 regardless of --min-score)
226
- ```
227
-
228
- Both the JUnit and HTML reports carry the same signal: unexecuted tests appear as `<skipped/>` cases naming the limit, distinct from `<failure>` cases, and the HTML report leads with a banner stating the run was stopped and why.
229
-
230
- Like the LLM provider settings, every field overrides from the environment (`BLASTPROOF_MAX_LLM_CALLS`, `BLASTPROOF_MAX_TOKENS`, `BLASTPROOF_MAX_DURATION_S`), and a CLI flag beats both:
231
-
232
- ```bash
233
- blastproof run --impacted --max-llm-calls 200 --max-duration 300
234
- blastproof plan --max-llm-calls 200
235
- ```
236
-
237
- `blastproof test` composes `run` then `plan` in one process. It resolves the budget once and hands the same instance to both phases, so the pipeline stays within the configured maximum overall — not up to double it, which is what each phase resolving its own budget would silently allow.
238
-
239
- ## blastproof tests itself
240
-
241
- The **Dogfood** badge above is blastproof running against the demo app in this repo: real Chromium, real LLM, plain-English tests, scored and gated. The run logs are public — the agent's reasoning, step by step, is there to read.
242
-
243
- It catches real regressions rather than diffing strings. Changing the demo app's discount from 20% to 5%, while leaving the on-screen message still claiming *"Promo code SAVE20 applied: 20% off"*, produces:
244
-
245
- ```
246
- FAIL P0 Promo code SAVE20 applies a 20% discount in the cart
247
- failing step: verify a 20% discount of $24.00 is shown
248
- reason: the discount is currently -$6.00, but a 20% discount on
249
- $120.00 should be -$24.00
250
- Score: 50 — min-score 80: FAIL (below threshold)
251
- ```
252
-
253
- No selector was updated and no assertion was rewritten to catch that. The agent read the rendered value, did the arithmetic, and disagreed with the page.
254
-
255
- Two workflows, split by what they cost:
256
-
257
- - **Impact** — runs on every pull request, including forks. Deterministic and keyless: it reports the blast radius of the diff and which tests cover it, before anyone spends a token.
258
- - **Dogfood** — runs daily and on demand. The agentic run needs an API key, so it stays out of the merge path: a non-deterministic model answer should never block a merge.
259
-
260
- ## GitHub Action
94
+ ## In CI
261
95
 
262
96
  ```yaml
263
97
  name: blastproof
@@ -269,89 +103,87 @@ jobs:
269
103
  steps:
270
104
  - uses: actions/checkout@v4
271
105
  with:
272
- fetch-depth: 0 # required: the diff needs a merge-base
106
+ fetch-depth: 0 # required — the diff needs a merge-base
273
107
 
274
108
  - run: npm start & # however your app boots
275
109
 
276
- - uses: hamc/blastproof@v0.2.2
110
+ - uses: hamc/blastproof@v0.4.0
277
111
  with:
112
+ version: '0.4.0' # pin both when this gates merges
278
113
  api-key: ${{ secrets.ANTHROPIC_API_KEY }}
279
114
  base: ${{ github.event.pull_request.base.ref }}
280
115
  min-score: '80'
281
- fail-on-unmapped: 'true' # a change nobody classified must not pass silently
116
+ fail-on-unmapped: 'true'
282
117
  ```
283
118
 
284
- Exit non-zero blocks the merge. Use the score in a later step:
119
+ A non-zero exit blocks the merge. The action outputs `score` (0–100, empty when no report was produced) for later steps. Full input list: [`action.yml`](./action.yml).
285
120
 
286
- ```yaml
287
- - id: bp
288
- uses: hamc/blastproof@v0.2.2
289
- with:
290
- version: '0.2.2' # pin both when the result gates merges
291
- api-key: ${{ secrets.ANTHROPIC_API_KEY }}
292
- base: ${{ github.event.pull_request.base.ref }}
293
- min-score: '80'
294
- fail-on-unmapped: 'true'
295
- html: report.html
121
+ `fetch-depth: 0` is not optional — the default checkout is shallow and has no merge-base. The action detects this and fails immediately rather than letting it surface as a git error mid-run.
122
+
123
+ ## Impact mapping
296
124
 
297
- - if: always()
298
- run: echo "scored ${{ steps.bp.outputs.score }}"
125
+ `--impacted` runs only the tests whose `routes:` intersect the routes your diff can affect, mapped from changed files by globs you maintain in `.blastproof/config.yaml`:
299
126
 
300
- - if: always()
301
- uses: actions/upload-artifact@v4
302
- with: { name: blastproof-report, path: report.html }
127
+ ```yaml
128
+ routes:
129
+ "src/cart/**": ["/cart", "/checkout"]
130
+ ignore:
131
+ - "**/*.md"
303
132
  ```
304
133
 
305
- | input | |
306
- | --- | --- |
307
- | `api-key` | Your provider key, from a secret. Not needed for ollama |
308
- | `provider` · `model` · `llm-base-url` | Override the committed config without editing it |
309
- | `command` | `test` (default), `run` or `plan` |
310
- | `base` | Base ref for the diff |
311
- | `url` | Base URL of the app under test |
312
- | `min-score` | Weighted score threshold — replaces the all-must-pass rule |
313
- | `fail-on-unmapped` | Fail on a changed file matching no `routes:` or `ignore:` glob |
314
- | `write` | Persist generated drafts instead of previewing |
315
- | `junit` · `html` | Where to write the reports |
316
- | `version` | Which blastproof release to install (default `latest`) |
317
- | `install-browser` | Set `false` if the workflow already installed Chromium |
318
- | `working-directory` | Directory containing `.blastproof/` |
134
+ Every changed file lands in one of three buckets: it matches `routes:` and contributes them, matches `ignore:` and is knowingly irrelevant, or **matches neither — nobody has said what it affects**. `--fail-on-unmapped` blocks on that third case, naming the files and both ways to resolve them.
319
135
 
320
- Output: **`score`** — 0–100, empty when no report was produced, so "no score" stays distinguishable from "scored zero".
136
+ **Nothing is ignored by default**, on purpose: a default that guesses on your behalf would hide the first files worth thinking about. The flag is additive — a run can meet `--min-score` and still be blocked here, because "the tests I ran passed" and "something changed that nobody classified" are different claims.
321
137
 
322
- **Pin both the tag and `version` when the result gates merges.** `latest` is convenient for trying it out, but a pipeline that blocks merges should not change behaviour without you changing something.
138
+ Its limit is worth knowing: it catches files that are *unclassified*, not *misclassified*. A shared module mapped to one route when it can break five still slips through. Impact by import graph is the fix, and blastproof does not do it yet.
323
139
 
324
- **`fetch-depth: 0` is not optional** for `test` and `plan`. The default checkout is shallow and has no merge-base; the action detects this and fails immediately rather than letting it surface as a git error mid-run.
140
+ ## Score and merge gating
325
141
 
326
- ## Trust boundaries
142
+ Each run scores the percentage of executed test **weight** that passed, weighing 3 at P0, 2 at P1, 1 at P2 — so a failing checkout costs three times a failing tooltip.
327
143
 
328
- For what the agent can and cannot *drive*, see [Does this fit your application?](#does-this-fit-your-application). This section is about what it is allowed to *reach*.
144
+ ```bash
145
+ blastproof run # any failure exits 1 (strict, the default)
146
+ blastproof run --min-score 80 # one failing P2 is tolerated
147
+ ```
329
148
 
330
- The application under test is not trusted input. Its page content reaches the model — that is how the agent knows what is on screen — so a page able to influence its own accessible text can try to influence the agent. Two things constrain that:
149
+ `--min-score` **replaces** the all-must-pass rule rather than adding to it. Only executed tests count: filtered and unrouted tests are neither numerator nor denominator, and a run that executed nothing scores 100, so a docs-only PR is never blocked. JUnit carries the score as a `<property name="score">`, and unrouted tests appear as `<skipped/>` so the coverage gap shows up in CI rather than vanishing.
331
150
 
332
- **The agent cannot leave your application.** `navigate` is bounded by `base_url`'s origin. An app that legitimately spans hosts declares them:
151
+ ## Without a browser or a key
333
152
 
334
- ```yaml
335
- allowed_origins:
336
- - https://auth.example.com
153
+ Half of blastproof is deterministic and free. These need no model, no browser and no network:
154
+
155
+ ```bash
156
+ blastproof run --dry-run # what would run
157
+ blastproof run --impacted --dry-run # + which routes the diff touches
158
+ blastproof run --impacted --fail-on-unmapped --dry-run # + gate on unclassified files
159
+ blastproof plan --base main --dry-run # affected routes no test covers, no key needed
337
160
  ```
338
161
 
339
- This is enforced by comparison, not by asking the model nicely, so it holds regardless of what the page says.
162
+ They report affected routes, files nobody has classified, and affected routes no test covers — a coverage-gap report with an exit code, useful even on a repo whose suite is Playwright or Cypress.
340
163
 
341
- **Your secrets stay out of the model's prompts.** `{{env.*}}` placeholders stay intact all the way through the prompt and are substituted at the moment of typing. Every value any test or the auth recipe references is also redacted from anything else crossing into a prompt — page snapshots included, since your app may render the credential itself — in both literal and percent-encoded form. This matters because blastproof encourages pointing `llm.base_url` at a gateway you do not run.
164
+ ## Bounding a run
342
165
 
343
- Redaction matches known values, so it cannot anticipate every way a page might transform one before rendering it. Treat it as a strong default, not a guarantee against a hostile application.
166
+ Nothing stops a run by default. `budget:` puts a ceiling on `run`, `plan` and `test` alike — every model call any of them makes is counted:
344
167
 
345
- The system prompt also tells the model that page content is data under test and never an instruction to obey. That raises the cost of a casual injection and is **not** a security boundary — a determined one will get past prompt wording. The origin constraint is the boundary; treat the rest as hygiene, and do not point blastproof at an application you would not run locally.
168
+ ```yaml
169
+ budget:
170
+ max_llm_calls: 500
171
+ max_tokens: 2000000
172
+ max_duration_s: 900
173
+ ```
346
174
 
347
- ## Testing behind a login
175
+ Each limit is optional; with none set, nothing binds. They count **calls and tokens, not currency** — a price table keyed by model and provider goes stale the day a provider reprices, and a limit that quietly stops meaning what it says is worse than none, because it is trusted.
176
+
177
+ Exhausting a budget **stops the run; it does not fail a test.** Running out of quota says nothing about the code under review. Unreached tests are reported as `not run`, a third state excluded from the score entirely, and the process exits 1 unconditionally — `--min-score` cannot rescue it, because the tests that finished are whichever ran first, not a representative sample.
348
178
 
349
- Most of a product lives behind authentication. Declare a recipe once and blastproof signs in a single time per run, then reuses that session for every test **and** for `plan` — so generated drafts describe the actual feature instead of the login wall.
179
+ `--dry-run` reports the ceiling before you spend anything.
350
180
 
351
- Pick exactly one strategy:
181
+ ## Testing behind a login
182
+
183
+ Declare a recipe once; blastproof signs in one time per run and reuses that session for every test and for `plan`. Pick exactly one strategy:
352
184
 
353
185
  ```yaml
354
- # 1) A plain-English login journey — form login, or anything a person can click through
186
+ # 1) A plain-English journey — form login, or anything a person can click through
355
187
  auth:
356
188
  steps:
357
189
  - navigate to /login
@@ -370,32 +202,13 @@ auth:
370
202
  Authorization: "Bearer {{env.API_TOKEN}}"
371
203
  ```
372
204
 
373
- A test that exercises the login itself must start signed out:
374
-
375
- ```yaml
376
- summary: Login with valid credentials succeeds
377
- auth: false
378
- ```
379
-
380
- **`verify` is worth the one extra call.** Without it, a wrong password surfaces as every test failing on a login wall — N failures, none naming the cause. With it, the run stops before the first test and says what happened. Authentication failure exits 2 and never reports as failing tests: a login you cannot complete says nothing about the code under review, so it must not produce a score.
381
-
382
- Each test still gets its own browser context; it simply starts from the shared session rather than empty, so isolation is unchanged. Set `auth.cache: true` to reuse a session across runs — off by default, because an expired session produces failures at random points with nothing pointing at the cause.
383
-
384
- **A captured session is a credential.** The file holds live cookies: whoever has it is signed in as that user. `init` git-ignores it; never commit one.
205
+ **`verify` is worth the extra call.** Without it a wrong password surfaces as every test failing on a login wall — N failures, none naming the cause. Authentication failure exits 2 and never reports as failing tests, because a login you cannot complete says nothing about the code under review.
385
206
 
386
- > **Note on self-healing:** the executor recovers from failed steps by re-reading the page, which means it can complete a login using credentials the page itself displays — some apps show demo credentials on the sign-in form. That is the self-healing loop working as designed, but it does mean a deliberately-wrong password is not a reliable way to test your auth failure path.
207
+ **A captured session is a credential** — the file holds live cookies. `init` git-ignores it; never commit one.
387
208
 
388
209
  ## LLM providers (BYOK)
389
210
 
390
- Bring your own key — runs 100% locally:
391
-
392
- - **Anthropic** (`ANTHROPIC_API_KEY`)
393
- - **OpenAI** (`OPENAI_API_KEY`)
394
- - **Ollama** (local, no key needed)
395
-
396
- ### Configuring from the environment
397
-
398
- You never have to commit a provider choice just to configure a pipeline. These variables override `.blastproof/config.yaml`, and precedence is **CLI flag > environment > file**:
211
+ **Anthropic** (`ANTHROPIC_API_KEY`), **OpenAI** (`OPENAI_API_KEY`), or **Ollama** (local, no key). Any setting can be overridden from the environment, with precedence **CLI flag > environment > file**:
399
212
 
400
213
  | variable | overrides |
401
214
  | --- | --- |
@@ -404,11 +217,9 @@ You never have to commit a provider choice just to configure a pipeline. These v
404
217
  | `BLASTPROOF_LLM_MODEL` | the model name |
405
218
  | `BLASTPROOF_LLM_BASE_URL` | the provider endpoint — *not* the app |
406
219
  | `BLASTPROOF_LLM_API_KEY_ENV` | the **name** of the variable holding your key |
407
- | `BLASTPROOF_MAX_LLM_CALLS` | `budget.max_llm_calls` — see [Bounding a run](#bounding-a-run-budget-and-deadline) |
408
- | `BLASTPROOF_MAX_TOKENS` | `budget.max_tokens` |
409
- | `BLASTPROOF_MAX_DURATION_S` | `budget.max_duration_s`, in seconds |
220
+ | `BLASTPROOF_MAX_LLM_CALLS` · `BLASTPROOF_MAX_TOKENS` · `BLASTPROOF_MAX_DURATION_S` | the budget fields |
410
221
 
411
- Running the committed config against an OpenAI-compatible gateway, without editing a file:
222
+ So an OpenAI-compatible gateway needs no file edit:
412
223
 
413
224
  ```bash
414
225
  export BLASTPROOF_LLM_PROVIDER=openai
@@ -418,33 +229,46 @@ export BLASTPROOF_LLM_API_KEY_ENV=OPENROUTER_API_KEY
418
229
  blastproof run --impacted --min-score 80
419
230
  ```
420
231
 
421
- Note the last one names *which variable* holds your key — the key itself is never read from a `BLASTPROOF_*` variable, so error messages can keep naming the variable you chose. An empty value counts as unset, so `FOO=` in a CI matrix will not blank a configured setting.
232
+ The key itself is never read from a `BLASTPROOF_*` variable — you name *which* variable holds it, so errors can keep naming the one you chose.
422
233
 
423
- ## Roadmap
234
+ ## Trust boundaries
424
235
 
425
- - [x] Repository & spec-driven development setup
426
- - [x] **M1** — `init` + `run`: YAML test runner with agentic LLM executor
427
- - [x] **M2** — diff analysis, impact mapping (`run --impacted`) and test generation (`plan`)
428
- - [x] **M3** — Reports (JUnit + HTML), priority-weighted score, `--min-score` gate, `blastproof test`
429
- - [x] **M4** — [published to npm](https://www.npmjs.com/package/blastproof) and a consumable GitHub Action
430
- - [ ] Next — impact by import graph, so a shared module's blast radius stops depending on hand-curated globs
431
- - [ ] Post-MVP — VS Code extension, session replay, worker parallelism, PR comments
236
+ The application under test is not trusted input: its page content reaches the model, so a page that controls its own accessible text can try to influence the agent. Two things constrain that.
432
237
 
433
- ## Development
238
+ **The agent cannot leave your application.** `navigate` is bounded by `base_url`'s origin; an app spanning hosts declares them in `allowed_origins:`. This is enforced by comparison, not by asking the model nicely.
239
+
240
+ **Your secrets stay out of prompts.** `{{env.*}}` placeholders survive intact and are substituted at the moment of typing. Every value your tests or auth recipe reference is redacted from everything else crossing into a prompt — page snapshots included — in literal and percent-encoded form. Redaction matches known values, so treat it as a strong default rather than a guarantee against a hostile app.
241
+
242
+ The system prompt also tells the model that page content is data, never instruction. That raises the cost of casual injection and is **not** a boundary — the origin constraint is. Do not point blastproof at an application you would not run locally.
434
243
 
435
- Built with AI assistance, using spec-driven development throughout: every change began as a written
436
- proposal with its design rationale, and those documents are kept rather than discarded — `openspec/`
437
- holds the reasoning behind each decision, including the alternatives that were rejected and why. The
438
- `.claude/`, `.cursor/` and `.opencode/` directories are agent configuration for the tools used to
439
- build it; ignore them unless you are contributing with one.
244
+ ## blastproof tests itself
245
+
246
+ The **Dogfood** badge is blastproof running against the demo app in this repo — real Chromium, real model, scored and gated, with public logs. It catches real regressions rather than diffing strings: change the demo discount from 20% to 5% while the page still claims *"20% off"* and it reports
247
+
248
+ ```
249
+ FAIL P0 Promo code SAVE20 applies a 20% discount in the cart
250
+ reason: the discount is currently -$6.00, but a 20% discount on
251
+ $120.00 should be -$24.00
252
+ ```
253
+
254
+ No selector was updated to catch that. The agent read the value, did the arithmetic, and disagreed with the page.
255
+
256
+ Try it yourself:
257
+
258
+ ```bash
259
+ git clone https://github.com/hamc/blastproof && cd blastproof
260
+ npm install && npm run build
261
+ node examples/demo-app/serve.mjs 4173 &
262
+ export ANTHROPIC_API_KEY=...
263
+ node dist/cli.js run
264
+ ```
265
+
266
+ ## Development
440
267
 
441
- This project uses **spec-driven development** via [OpenSpec](https://github.com/Fission-AI/OpenSpec). See [`AGENTS.md`](./AGENTS.md) for architecture, conventions and the contribution workflow — every change starts with an OpenSpec proposal.
268
+ Built with AI assistance using spec-driven development: every change began as a written proposal with its design rationale, and those documents are kept rather than discarded. `openspec/` holds the reasoning behind each decision, including the alternatives that were rejected and why — start at [`AGENTS.md`](./AGENTS.md) for architecture, conventions and the contribution workflow. Open work lives in [issues](https://github.com/hamc/blastproof/issues).
442
269
 
443
270
  ```bash
444
- # requires Node.js >= 20.19 (see engines in package.json)
445
- npm install
446
- npm run build
447
- npm test
271
+ npm install && npm run build && npm test
448
272
  ```
449
273
 
450
274
  ## License