blastproof 0.3.0 → 0.5.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +132 -308
- package/dist/cli.js +261 -27
- package/dist/cli.js.map +1 -1
- package/package.json +1 -1
package/README.md
CHANGED
|
@@ -4,74 +4,51 @@
|
|
|
4
4
|
[](https://github.com/hamc/blastproof/actions/workflows/dogfood.yml)
|
|
5
5
|
[](./LICENSE)
|
|
6
6
|
|
|
7
|
-
**Open-source AI testing agent for pull requests.**
|
|
8
|
-
|
|
9
|
-
`blastproof` is an open-source AI QA agent for pull requests: it reads your PR diff, maps the blast radius, generates end-to-end tests in plain English, executes them on a real browser with self-healing, and scores the result before merge. 100% local, MIT licensed, bring your own LLM key.
|
|
7
|
+
**Open-source AI testing agent for pull requests.** Write end-to-end tests as plain English. An agent drives a real browser to run them, selects only the ones your diff can affect, and scores the result before merge.
|
|
10
8
|
|
|
11
9
|
```
|
|
12
10
|
git diff → impact mapping → test generation → agentic execution → report + score
|
|
13
11
|
```
|
|
14
12
|
|
|
15
|
-
|
|
16
|
-
|
|
17
|
-
1. **Reads the diff** — `blastproof test --base main` parses the branch diff and maps it to affected routes.
|
|
18
|
-
2. **Maps the blast radius** — matches changed files against the `routes:` map you maintain, to select the journeys worth running. Deterministic and free; see [Closing the coverage hole](#closing-the-coverage-hole) for what that catches and what it does not.
|
|
19
|
-
3. **Writes the missing tests** — drafts plain-English YAML for affected routes nothing covers. Drafts are yours to review; existing tests are never rewritten.
|
|
20
|
-
4. **Executes agentically** — an LLM-driven loop over Playwright resolves elements via the accessibility tree on every step. No static selectors to rot: the agent re-resolves when the UI shifts. Being model-driven it is not deterministic, which is why the free, deterministic impact analysis is what runs on every pull request.
|
|
21
|
-
5. **Reports & scores** — console, JUnit XML and HTML reports, plus a priority-weighted score that fails the run below `--min-score`.
|
|
22
|
-
|
|
23
|
-
## Does this fit your application?
|
|
24
|
-
|
|
25
|
-
Read this before installing. The agent drives your app the way a screen reader would, and that has consequences worth knowing in advance rather than discovering on a failing run.
|
|
26
|
-
|
|
27
|
-
**Your markup has to be accessible. This is a hard requirement, not a nice-to-have.** Every element is found by role, label or visible text, resolved from the page's accessibility tree on each attempt — that is what removes selectors and what lets the agent survive a redesign. The cost is that an interface the accessibility tree cannot describe is one the agent cannot drive at all. Icon-only buttons with no accessible name, custom dropdowns and comboboxes without ARIA roles, and `div`-based controls do not degrade into something slower; they simply cannot be targeted, and there is deliberately no CSS or XPath escape hatch to fall back on. If you are unsure, run your app through any accessibility checker first — the result predicts how well blastproof will do better than anything else.
|
|
28
|
-
|
|
29
|
-
**Some interactions are not supported yet:**
|
|
30
|
-
|
|
31
|
-
| Not supported | Consequence |
|
|
32
|
-
|---|---|
|
|
33
|
-
| `iframe` content | Anything inside an embedded frame is invisible to the agent, including the page snapshot. **Hosted payment widgets such as Stripe Elements or PayPal buttons fall in here** — if your checkout embeds one, the checkout journey cannot be driven end to end today |
|
|
34
|
-
| Hover, scroll-to, drag and drop | Menus that open only on hover, and content that loads only once scrolled into view, are out of reach |
|
|
35
|
-
| File upload | No action exists for it |
|
|
36
|
-
| Multiple tabs or windows | The agent works in one page; a flow that opens a popup loses it |
|
|
37
|
-
| Native `alert` / `confirm` dialogs | Not handled |
|
|
38
|
-
|
|
39
|
-
**Large pages are truncated.** The accessibility snapshot sent to the model is capped at 200 lines. A dense admin table or a long product listing will have its lower content invisible on every step.
|
|
40
|
-
|
|
41
|
-
**The deterministic half stands on its own.** If none of the above fits, `blastproof run --impacted --dry-run` still works: it reads the diff, reports which routes a change touches, which files are classified by nothing, and which affected routes no test covers. That needs no API key, no browser and no model, costs nothing per pull request, and is useful purely as a coverage-gap report — see [Closing the coverage hole](#closing-the-coverage-hole).
|
|
13
|
+
100% local. MIT. Bring your own LLM key.
|
|
42
14
|
|
|
43
15
|
## Quick start
|
|
44
16
|
|
|
45
17
|
```bash
|
|
46
|
-
npm install -g blastproof
|
|
47
|
-
npx playwright install --with-deps chromium # one-time browser download
|
|
48
|
-
|
|
18
|
+
npm install -g blastproof # Node.js >= 20.19
|
|
49
19
|
cd your-project
|
|
50
|
-
blastproof init
|
|
51
|
-
# start your app, then point base_url at it in .blastproof/config.yaml
|
|
52
|
-
export ANTHROPIC_API_KEY=... # or OPENAI_API_KEY, or a local Ollama model
|
|
53
|
-
blastproof run # runs .blastproof/tests/**/*.yaml agentically
|
|
20
|
+
blastproof init # scaffolds .blastproof/
|
|
54
21
|
```
|
|
55
22
|
|
|
56
|
-
`
|
|
23
|
+
Point `base_url` at your running app in `.blastproof/config.yaml`, then check the setup — this needs **no API key and no browser**:
|
|
57
24
|
|
|
58
|
-
|
|
25
|
+
```bash
|
|
26
|
+
blastproof run --dry-run
|
|
27
|
+
```
|
|
59
28
|
|
|
60
|
-
|
|
29
|
+
To actually execute tests you need a browser and a model:
|
|
61
30
|
|
|
62
31
|
```bash
|
|
63
|
-
|
|
64
|
-
|
|
65
|
-
|
|
66
|
-
export ANTHROPIC_API_KEY=...
|
|
67
|
-
node dist/cli.js run
|
|
32
|
+
npx playwright install --with-deps chromium # NEEDS SUDO — see below
|
|
33
|
+
export ANTHROPIC_API_KEY=... # or OPENAI_API_KEY, or local Ollama
|
|
34
|
+
blastproof run
|
|
68
35
|
```
|
|
69
36
|
|
|
70
|
-
|
|
37
|
+
**No sudo?** `--with-deps` installs system libraries as root. Without it, run `npx playwright install chromium` and obtain `libnspr4`, `libnss3`, `libnssutil3` and `libasound2` however you can. Note that a useful half of blastproof needs neither browser nor key — see [Without a browser or a key](#without-a-browser-or-a-key).
|
|
38
|
+
|
|
39
|
+
Before `run`, `plan` or `test` do anything, they check what they are about to spend — the browser can launch, the model provider is reachable, `base_url` responds — and report every unmet one together, so a stopped app or a missing browser is never a wall you hit one crash at a time. A missing system library names the exact install command and says it needs root; nothing is installed on your behalf. Silent when everything is fine, and skipped entirely by `--dry-run`, which needs none of it.
|
|
40
|
+
|
|
41
|
+
## Does this fit your application?
|
|
42
|
+
|
|
43
|
+
**Your markup must be accessible — a hard requirement.** Elements are found by role, label or visible text from the accessibility tree. That is what removes selectors and survives redesigns; the cost is that an interface the accessibility tree cannot describe cannot be driven at all, and there is deliberately no CSS or XPath fallback. Icon-only buttons without accessible names, `div`-based controls and ARIA-less dropdowns simply cannot be targeted. Run an accessibility checker first — the result predicts how well this will work better than anything else.
|
|
44
|
+
|
|
45
|
+
**Not supported yet:** `iframe` content (so hosted payment widgets like Stripe Elements are invisible — an embedded checkout cannot be driven end to end), hover, scroll-to, drag and drop, file upload, multiple tabs, native `alert`/`confirm` dialogs. Page snapshots are capped at 200 lines by default, so very dense pages are truncated — raise it with `browser.max_snapshot_lines` if your pages need more; truncation is always marked in the snapshot so the model is never misled into thinking it saw the whole page.
|
|
46
|
+
|
|
47
|
+
`browser.timeout_ms` bounds every wait — resolving a target element from the accessibility tree, and navigation — not only the click or fill performed afterwards. Raise it for an application that is merely slow to hydrate; the trade-off is that a genuinely missing element then takes longer to fail. It never changes how many self-healing retries a step gets — waiting and retrying are deliberately separate.
|
|
71
48
|
|
|
72
|
-
##
|
|
49
|
+
## Writing tests
|
|
73
50
|
|
|
74
|
-
Tests live in `.blastproof/tests/` as plain-English YAML — no selectors
|
|
51
|
+
Tests live in `.blastproof/tests/` as plain-English YAML — no selectors:
|
|
75
52
|
|
|
76
53
|
```yaml
|
|
77
54
|
summary: Checkout with discount
|
|
@@ -85,179 +62,36 @@ steps:
|
|
|
85
62
|
- complete checkout
|
|
86
63
|
```
|
|
87
64
|
|
|
88
|
-
`priority` is P0–P2 (default P1)
|
|
89
|
-
|
|
90
|
-
## CLI
|
|
91
|
-
|
|
92
|
-
| Command | Description |
|
|
93
|
-
| --- | --- |
|
|
94
|
-
| `blastproof init` | Scaffold `.blastproof/` config and sample tests (idempotent) |
|
|
95
|
-
| `blastproof run [--tag smoke] [--priority P0] [--query checkout]` | Run tests only — exit 0 pass, 1 fail, 2 usage/config error |
|
|
96
|
-
| `blastproof run --impacted [--base <ref>]` | Run only tests impacted by the diff vs the base ref (default `main`). Unrouted tests are skipped and reported; affected-but-uncovered routes are reported without failing the run |
|
|
97
|
-
| `blastproof run --dry-run` | Print the selection plan (affected routes, unmapped files, selected/skipped tests) and exit 0 — no browser launched, no LLM key needed |
|
|
98
|
-
| `blastproof run --url <url>` | Override `base_url` for this run only (e.g. a PR preview environment); the config file is never mutated |
|
|
99
|
-
| `blastproof plan [--base <ref>]` | Generate plain-English tests for affected routes no test covers yet. Prints drafts; nothing is written without `--write` |
|
|
100
|
-
| `blastproof plan --route <route>` | Generate for a route explicitly, skipping the diff (repeatable) — how you bootstrap coverage on an app with no suite yet |
|
|
101
|
-
| `blastproof plan --write` | Persist drafts to `.blastproof/tests/<route-slug>.yaml`. Never overwrites: a colliding filename fails that route |
|
|
102
|
-
| `blastproof run --min-score <n>` | Require a weighted score of at least `n` (0–100). **Replaces** the all-must-pass rule — see below |
|
|
103
|
-
| `blastproof run --junit [path]` | Write a JUnit XML report; without a path it lands in `.blastproof/reports/<session>/junit.xml` |
|
|
104
|
-
| `blastproof run --html [path]` | Write a self-contained HTML report with failure screenshots embedded inline |
|
|
105
|
-
| `--max-llm-calls <n>` \| `--max-tokens <n>` \| `--max-duration <seconds>` | Bound `run`, `plan` or `test` by model calls, tokens or wall-clock time — every command that can call the model accepts these. Overrides `.blastproof/config.yaml`'s `budget:` section — see [Bounding a run](#bounding-a-run-budget-and-deadline) |
|
|
106
|
-
| `blastproof test [--base <ref>]` | The full pipeline: run the tests covering the diff, then draft tests for the gaps |
|
|
107
|
-
|
|
108
|
-
### The full pipeline: `blastproof test`
|
|
109
|
-
|
|
110
|
-
One command for the whole loop — map the blast radius, run what covers it, draft what doesn't exist yet:
|
|
111
|
-
|
|
112
|
-
```bash
|
|
113
|
-
blastproof test --base main --min-score 80 --junit junit.xml --html report.html
|
|
114
|
-
```
|
|
115
|
-
|
|
116
|
-
It does two things and reports them separately:
|
|
65
|
+
`priority` is P0–P2 (default P1). `tags`, `setup` steps and `auth` are optional — `auth: false` runs the test signed out, which a login test needs. `routes` declares the URLs a test covers, which is what `--impacted` selects on; write route strings consistently, since they compare by exact equality (`/cart` ≠ `/cart/`).
|
|
117
66
|
|
|
118
|
-
|
|
119
|
-
2. **Draft** — generates tests for affected routes no test covers, and prints them
|
|
120
|
-
|
|
121
|
-
**Generated drafts are never executed, and never affect the score.** That is deliberate. An unreviewed, model-written test in the merge path fails in two directions: a hallucinated expectation blocks a correct pull request, and a credulous one waves a broken change through while looking like coverage. Either costs more trust than the automation saves.
|
|
122
|
-
|
|
123
|
-
So read the result honestly: `test` does not make an uncovered route safe. It makes the gap visible, with a draft ready for you to review, and the score keeps describing only what was actually verified. Add `--write` to persist the drafts (never overwriting an existing file) and commit them once you have read them.
|
|
124
|
-
|
|
125
|
-
Exit codes: 2 usage/config/diff, 1 when the gate fails **or** a draft could not be generated, 0 otherwise.
|
|
126
|
-
|
|
127
|
-
### Reports
|
|
128
|
-
|
|
129
|
-
`--junit` is for CI; `--html` is for humans. The HTML report is a single self-contained file — inline CSS, screenshots embedded as data URIs, no scripts — so it opens offline, survives being moved, and uploads as one artifact. It leads with the score and gate verdict, sorts failures above passes, and expands each failure to its failing step, reason and screenshot.
|
|
130
|
-
|
|
131
|
-
### Generating tests with `plan`
|
|
132
|
-
|
|
133
|
-
`plan` closes the gap `run --impacted` reports. It takes the affected routes no test covers, loads each one in the browser, and asks the model to write a test from the page's real accessibility tree plus the changed files that made the route impacted — so the generated steps name controls that actually exist:
|
|
67
|
+
## Commands
|
|
134
68
|
|
|
135
69
|
```bash
|
|
136
|
-
blastproof
|
|
137
|
-
blastproof
|
|
138
|
-
blastproof
|
|
70
|
+
blastproof init # scaffold .blastproof/
|
|
71
|
+
blastproof run # run every test
|
|
72
|
+
blastproof run --impacted --base main # run only what the diff can affect
|
|
73
|
+
blastproof plan --base main # draft tests for uncovered routes
|
|
74
|
+
blastproof test --base main # run what covers the diff, then draft the gaps
|
|
139
75
|
```
|
|
140
76
|
|
|
141
|
-
|
|
142
|
-
|
|
143
|
-
Exit codes: 0 when every route generated (or nothing needed coverage), 1 when a route failed or the budget/deadline stopped generation early, 2 on usage/config/diff errors. A route that fails to load never aborts the others; a budget or deadline stop does — see [Bounding a run](#bounding-a-run-budget-and-deadline).
|
|
144
|
-
|
|
145
|
-
`plan` uses the same `auth:` recipe as `run`, so a route behind a login is drafted from the real page rather than from the login wall.
|
|
146
|
-
|
|
147
|
-
### Closing the coverage hole
|
|
77
|
+
Common flags — `blastproof <command> --help` has the full list:
|
|
148
78
|
|
|
149
|
-
|
|
150
|
-
|
|
151
|
-
Each changed file is classified three ways:
|
|
152
|
-
|
|
153
|
-
| a changed file | means |
|
|
79
|
+
| flag | |
|
|
154
80
|
| --- | --- |
|
|
155
|
-
|
|
|
156
|
-
|
|
|
157
|
-
|
|
|
158
|
-
|
|
159
|
-
|
|
160
|
-
|
|
161
|
-
|
|
162
|
-
|
|
163
|
-
- "**/*.md"
|
|
164
|
-
- ".github/**"
|
|
165
|
-
```
|
|
166
|
-
|
|
167
|
-
```bash
|
|
168
|
-
blastproof run --impacted --fail-on-unmapped
|
|
169
|
-
```
|
|
170
|
-
|
|
171
|
-
The flag fails the run on the third case only, naming the files and both ways to resolve them. `ignore:` is what makes that signal survivable — without it the flag would fire on every README edit and get switched off within a day, and a disabled gate protects nothing.
|
|
172
|
-
|
|
173
|
-
Nothing is ignored by default, on purpose: a file nobody has classified is exactly the risk the flag exists to surface, and a default that guesses on your behalf would hide the first files worth thinking about. This flag is **additive** — a run can meet its `--min-score` and still be blocked here, because "the tests I ran passed" and "something changed that nobody has classified" are different claims.
|
|
174
|
-
|
|
175
|
-
Be clear on its limit: it catches files that are *unclassified*, not files that are *misclassified*. A shared module mapped to one route when it can break five will still slip through. Mapping by import graph is the answer to that, and blastproof does not do it yet.
|
|
176
|
-
|
|
177
|
-
### Score and merge gating
|
|
178
|
-
|
|
179
|
-
Every run ends with a score: the percentage of executed test **weight** that passed, where a test weighs 3 at P0, 2 at P1 and 1 at P2. Weighting is the point — a failing checkout costs three times a failing tooltip, so a pile of trivial passes can't hide a broken critical journey.
|
|
180
|
-
|
|
181
|
-
```bash
|
|
182
|
-
blastproof run # any failure exits 1 (strict, the default)
|
|
183
|
-
blastproof run --min-score 80 # passes at 80+, so one failing P2 is tolerated
|
|
184
|
-
blastproof run --min-score 100 # identical to the default strict behaviour
|
|
185
|
-
```
|
|
81
|
+
| `--dry-run` | Print the selection (or, for `plan`, the routes it would draft) and exit. No browser, no API key |
|
|
82
|
+
| `--tag` · `--priority` · `--query` | Select a subset of tests |
|
|
83
|
+
| `--url <url>` | Override `base_url` for this run (e.g. a PR preview) |
|
|
84
|
+
| `--min-score <n>` | Gate on a weighted score instead of all-must-pass |
|
|
85
|
+
| `--fail-on-unmapped` | Fail when a changed file matches no `routes:` or `ignore:` glob |
|
|
86
|
+
| `--junit [path]` · `--html [path]` | Write reports |
|
|
87
|
+
| `--write` | `plan` only — persist drafts instead of previewing |
|
|
88
|
+
| `--max-llm-calls` · `--max-tokens` · `--max-duration` | Bound what a run may spend |
|
|
186
89
|
|
|
187
|
-
|
|
90
|
+
Exit codes: **0** pass, **1** the gate failed, **2** usage or config error.
|
|
188
91
|
|
|
189
|
-
|
|
92
|
+
**Generated drafts are never executed and never affect the score.** An unreviewed model-written test in the merge path fails in two directions: a hallucinated expectation blocks a correct PR, and a credulous one waves a broken change through while looking like coverage. `plan` makes the gap visible with a draft to review; it does not make an uncovered route safe.
|
|
190
93
|
|
|
191
|
-
|
|
192
|
-
blastproof run --impacted --base "$BASE_REF" --min-score 80 --junit junit.xml
|
|
193
|
-
```
|
|
194
|
-
|
|
195
|
-
Exit 0 merge-able, 1 blocked, 2 usage/config error. The JUnit report carries the score as a `<property name="score">` so a parser can read it without scraping stdout, and tests skipped for having no `routes:` appear as `<skipped/>` cases — the coverage gap shows up in CI instead of vanishing.
|
|
196
|
-
|
|
197
|
-
## Bounding a run: budget and deadline
|
|
198
|
-
|
|
199
|
-
Nothing stops a run by default — a suite runs to completion or the provider refuses. `budget:` puts a ceiling on it, in `.blastproof/config.yaml`. It applies to `run`, `plan` and `test` alike: every model call any of them makes — agent action, assert judgment, or test planning — is counted, because a budget that only covered `run` would leave `plan`'s calls unbounded.
|
|
200
|
-
|
|
201
|
-
```yaml
|
|
202
|
-
budget:
|
|
203
|
-
max_llm_calls: 500 # stop after this many model calls
|
|
204
|
-
max_tokens: 2000000 # stop after this many tokens spent across all calls
|
|
205
|
-
max_duration_s: 900 # stop after this many seconds of wall-clock time
|
|
206
|
-
```
|
|
207
|
-
|
|
208
|
-
Each limit is independent and optional; a config with no `budget:` section — or a run with none of `--max-llm-calls` / `--max-tokens` / `--max-duration` — behaves exactly as before. All three are counted in **calls and tokens, not currency**: a price table keyed by model and provider goes stale the day a provider reprices, and a limit that silently stops meaning what it says is worse than no limit, because it is trusted. Calls and tokens are exact, already reported by every provider, and yours to convert to a dollar figure with your own rates if you want one.
|
|
209
|
-
|
|
210
|
-
`--dry-run` reports the ceiling before you spend anything — the worst case a selection could cost, computed from step counts alone, no provider contacted:
|
|
211
|
-
|
|
212
|
-
```
|
|
213
|
-
Dry run: 12 test(s) selected, base_url=http://localhost:4173
|
|
214
|
-
Worst case: up to 216 model call(s) for this selection (a maximum, not a prediction).
|
|
215
|
-
```
|
|
216
|
-
|
|
217
|
-
That number is a ceiling, not a forecast — a real run almost always finishes in a fraction of it, because most steps complete long before either cap runs out. Per step it is the iteration cap **plus** `max_retries_per_step` (read from your config, not assumed): a malformed model response is retried without spending an iteration, and a failing `assert` spends a retry *and* an iteration in the same call — so the two caps are added, not one doubled and the other ignored. If `auth.steps` is configured, the login journey's steps are counted too — it runs once before any test and spends model calls through the same loop, so a ceiling that excluded it could be exceeded by the very first run that logs in.
|
|
218
|
-
|
|
219
|
-
**Exhausting a budget stops the run — it does not fail a test.** Running out of quota says nothing about the application under test, so recording it as a failure would manufacture a defect that does not exist. Tests the run never reached are reported as **not run**, a third state distinct from passed and failed, and excluded from the score's denominator entirely — counting them as failures would just be a quieter version of the same lie a false pass would have been.
|
|
220
|
-
|
|
221
|
-
An interrupted run is unmistakably incomplete: the process **exits 1 unconditionally**, even when `--min-score` is given and the tests that did execute would have satisfied it. The tests that finished are whichever ones happened to run first, not a representative sample, so nothing about them is a verdict:
|
|
222
|
-
|
|
223
|
-
```
|
|
224
|
-
Run incomplete: model call budget exhausted: reached the configured maximum of 500 call(s)
|
|
225
|
-
Score over executed tests: 92 (not a verdict — exit code 1 regardless of --min-score)
|
|
226
|
-
```
|
|
227
|
-
|
|
228
|
-
Both the JUnit and HTML reports carry the same signal: unexecuted tests appear as `<skipped/>` cases naming the limit, distinct from `<failure>` cases, and the HTML report leads with a banner stating the run was stopped and why.
|
|
229
|
-
|
|
230
|
-
Like the LLM provider settings, every field overrides from the environment (`BLASTPROOF_MAX_LLM_CALLS`, `BLASTPROOF_MAX_TOKENS`, `BLASTPROOF_MAX_DURATION_S`), and a CLI flag beats both:
|
|
231
|
-
|
|
232
|
-
```bash
|
|
233
|
-
blastproof run --impacted --max-llm-calls 200 --max-duration 300
|
|
234
|
-
blastproof plan --max-llm-calls 200
|
|
235
|
-
```
|
|
236
|
-
|
|
237
|
-
`blastproof test` composes `run` then `plan` in one process. It resolves the budget once and hands the same instance to both phases, so the pipeline stays within the configured maximum overall — not up to double it, which is what each phase resolving its own budget would silently allow.
|
|
238
|
-
|
|
239
|
-
## blastproof tests itself
|
|
240
|
-
|
|
241
|
-
The **Dogfood** badge above is blastproof running against the demo app in this repo: real Chromium, real LLM, plain-English tests, scored and gated. The run logs are public — the agent's reasoning, step by step, is there to read.
|
|
242
|
-
|
|
243
|
-
It catches real regressions rather than diffing strings. Changing the demo app's discount from 20% to 5%, while leaving the on-screen message still claiming *"Promo code SAVE20 applied: 20% off"*, produces:
|
|
244
|
-
|
|
245
|
-
```
|
|
246
|
-
FAIL P0 Promo code SAVE20 applies a 20% discount in the cart
|
|
247
|
-
failing step: verify a 20% discount of $24.00 is shown
|
|
248
|
-
reason: the discount is currently -$6.00, but a 20% discount on
|
|
249
|
-
$120.00 should be -$24.00
|
|
250
|
-
Score: 50 — min-score 80: FAIL (below threshold)
|
|
251
|
-
```
|
|
252
|
-
|
|
253
|
-
No selector was updated and no assertion was rewritten to catch that. The agent read the rendered value, did the arithmetic, and disagreed with the page.
|
|
254
|
-
|
|
255
|
-
Two workflows, split by what they cost:
|
|
256
|
-
|
|
257
|
-
- **Impact** — runs on every pull request, including forks. Deterministic and keyless: it reports the blast radius of the diff and which tests cover it, before anyone spends a token.
|
|
258
|
-
- **Dogfood** — runs daily and on demand. The agentic run needs an API key, so it stays out of the merge path: a non-deterministic model answer should never block a merge.
|
|
259
|
-
|
|
260
|
-
## GitHub Action
|
|
94
|
+
## In CI
|
|
261
95
|
|
|
262
96
|
```yaml
|
|
263
97
|
name: blastproof
|
|
@@ -269,89 +103,87 @@ jobs:
|
|
|
269
103
|
steps:
|
|
270
104
|
- uses: actions/checkout@v4
|
|
271
105
|
with:
|
|
272
|
-
fetch-depth: 0 # required
|
|
106
|
+
fetch-depth: 0 # required — the diff needs a merge-base
|
|
273
107
|
|
|
274
108
|
- run: npm start & # however your app boots
|
|
275
109
|
|
|
276
|
-
- uses: hamc/blastproof@v0.
|
|
110
|
+
- uses: hamc/blastproof@v0.5.0
|
|
277
111
|
with:
|
|
112
|
+
version: '0.5.0' # pin both when this gates merges
|
|
278
113
|
api-key: ${{ secrets.ANTHROPIC_API_KEY }}
|
|
279
114
|
base: ${{ github.event.pull_request.base.ref }}
|
|
280
115
|
min-score: '80'
|
|
281
|
-
fail-on-unmapped: 'true'
|
|
116
|
+
fail-on-unmapped: 'true'
|
|
282
117
|
```
|
|
283
118
|
|
|
284
|
-
|
|
119
|
+
A non-zero exit blocks the merge. The action outputs `score` (0–100, empty when no report was produced) for later steps. Full input list: [`action.yml`](./action.yml).
|
|
285
120
|
|
|
286
|
-
|
|
287
|
-
|
|
288
|
-
|
|
289
|
-
with:
|
|
290
|
-
version: '0.2.2' # pin both when the result gates merges
|
|
291
|
-
api-key: ${{ secrets.ANTHROPIC_API_KEY }}
|
|
292
|
-
base: ${{ github.event.pull_request.base.ref }}
|
|
293
|
-
min-score: '80'
|
|
294
|
-
fail-on-unmapped: 'true'
|
|
295
|
-
html: report.html
|
|
121
|
+
`fetch-depth: 0` is not optional — the default checkout is shallow and has no merge-base. The action detects this and fails immediately rather than letting it surface as a git error mid-run.
|
|
122
|
+
|
|
123
|
+
## Impact mapping
|
|
296
124
|
|
|
297
|
-
|
|
298
|
-
run: echo "scored ${{ steps.bp.outputs.score }}"
|
|
125
|
+
`--impacted` runs only the tests whose `routes:` intersect the routes your diff can affect, mapped from changed files by globs you maintain in `.blastproof/config.yaml`:
|
|
299
126
|
|
|
300
|
-
|
|
301
|
-
|
|
302
|
-
|
|
127
|
+
```yaml
|
|
128
|
+
routes:
|
|
129
|
+
"src/cart/**": ["/cart", "/checkout"]
|
|
130
|
+
ignore:
|
|
131
|
+
- "**/*.md"
|
|
303
132
|
```
|
|
304
133
|
|
|
305
|
-
|
|
306
|
-
| --- | --- |
|
|
307
|
-
| `api-key` | Your provider key, from a secret. Not needed for ollama |
|
|
308
|
-
| `provider` · `model` · `llm-base-url` | Override the committed config without editing it |
|
|
309
|
-
| `command` | `test` (default), `run` or `plan` |
|
|
310
|
-
| `base` | Base ref for the diff |
|
|
311
|
-
| `url` | Base URL of the app under test |
|
|
312
|
-
| `min-score` | Weighted score threshold — replaces the all-must-pass rule |
|
|
313
|
-
| `fail-on-unmapped` | Fail on a changed file matching no `routes:` or `ignore:` glob |
|
|
314
|
-
| `write` | Persist generated drafts instead of previewing |
|
|
315
|
-
| `junit` · `html` | Where to write the reports |
|
|
316
|
-
| `version` | Which blastproof release to install (default `latest`) |
|
|
317
|
-
| `install-browser` | Set `false` if the workflow already installed Chromium |
|
|
318
|
-
| `working-directory` | Directory containing `.blastproof/` |
|
|
134
|
+
Every changed file lands in one of three buckets: it matches `routes:` and contributes them, matches `ignore:` and is knowingly irrelevant, or **matches neither — nobody has said what it affects**. `--fail-on-unmapped` blocks on that third case, naming the files and both ways to resolve them.
|
|
319
135
|
|
|
320
|
-
|
|
136
|
+
**Nothing is ignored by default**, on purpose: a default that guesses on your behalf would hide the first files worth thinking about. The flag is additive — a run can meet `--min-score` and still be blocked here, because "the tests I ran passed" and "something changed that nobody classified" are different claims.
|
|
321
137
|
|
|
322
|
-
|
|
138
|
+
Its limit is worth knowing: it catches files that are *unclassified*, not *misclassified*. A shared module mapped to one route when it can break five still slips through. Impact by import graph is the fix, and blastproof does not do it yet.
|
|
323
139
|
|
|
324
|
-
|
|
140
|
+
## Score and merge gating
|
|
325
141
|
|
|
326
|
-
|
|
142
|
+
Each run scores the percentage of executed test **weight** that passed, weighing 3 at P0, 2 at P1, 1 at P2 — so a failing checkout costs three times a failing tooltip.
|
|
327
143
|
|
|
328
|
-
|
|
144
|
+
```bash
|
|
145
|
+
blastproof run # any failure exits 1 (strict, the default)
|
|
146
|
+
blastproof run --min-score 80 # one failing P2 is tolerated
|
|
147
|
+
```
|
|
329
148
|
|
|
330
|
-
|
|
149
|
+
`--min-score` **replaces** the all-must-pass rule rather than adding to it. Only executed tests count: filtered and unrouted tests are neither numerator nor denominator, and a run that executed nothing scores 100, so a docs-only PR is never blocked. JUnit carries the score as a `<property name="score">`, and unrouted tests appear as `<skipped/>` so the coverage gap shows up in CI rather than vanishing.
|
|
331
150
|
|
|
332
|
-
|
|
151
|
+
## Without a browser or a key
|
|
333
152
|
|
|
334
|
-
|
|
335
|
-
|
|
336
|
-
|
|
153
|
+
Half of blastproof is deterministic and free. These need no model, no browser and no network:
|
|
154
|
+
|
|
155
|
+
```bash
|
|
156
|
+
blastproof run --dry-run # what would run
|
|
157
|
+
blastproof run --impacted --dry-run # + which routes the diff touches
|
|
158
|
+
blastproof run --impacted --fail-on-unmapped --dry-run # + gate on unclassified files
|
|
159
|
+
blastproof plan --base main --dry-run # affected routes no test covers, no key needed
|
|
337
160
|
```
|
|
338
161
|
|
|
339
|
-
|
|
162
|
+
They report affected routes, files nobody has classified, and affected routes no test covers — a coverage-gap report with an exit code, useful even on a repo whose suite is Playwright or Cypress.
|
|
340
163
|
|
|
341
|
-
|
|
164
|
+
## Bounding a run
|
|
342
165
|
|
|
343
|
-
|
|
166
|
+
Nothing stops a run by default. `budget:` puts a ceiling on `run`, `plan` and `test` alike — every model call any of them makes is counted:
|
|
344
167
|
|
|
345
|
-
|
|
168
|
+
```yaml
|
|
169
|
+
budget:
|
|
170
|
+
max_llm_calls: 500
|
|
171
|
+
max_tokens: 2000000
|
|
172
|
+
max_duration_s: 900
|
|
173
|
+
```
|
|
346
174
|
|
|
347
|
-
|
|
175
|
+
Each limit is optional; with none set, nothing binds. They count **calls and tokens, not currency** — a price table keyed by model and provider goes stale the day a provider reprices, and a limit that quietly stops meaning what it says is worse than none, because it is trusted.
|
|
176
|
+
|
|
177
|
+
Exhausting a budget **stops the run; it does not fail a test.** Running out of quota says nothing about the code under review. Unreached tests are reported as `not run`, a third state excluded from the score entirely, and the process exits 1 unconditionally — `--min-score` cannot rescue it, because the tests that finished are whichever ran first, not a representative sample.
|
|
348
178
|
|
|
349
|
-
|
|
179
|
+
`--dry-run` reports the ceiling before you spend anything.
|
|
350
180
|
|
|
351
|
-
|
|
181
|
+
## Testing behind a login
|
|
182
|
+
|
|
183
|
+
Declare a recipe once; blastproof signs in one time per run and reuses that session for every test and for `plan`. Pick exactly one strategy:
|
|
352
184
|
|
|
353
185
|
```yaml
|
|
354
|
-
# 1) A plain-English
|
|
186
|
+
# 1) A plain-English journey — form login, or anything a person can click through
|
|
355
187
|
auth:
|
|
356
188
|
steps:
|
|
357
189
|
- navigate to /login
|
|
@@ -370,32 +202,13 @@ auth:
|
|
|
370
202
|
Authorization: "Bearer {{env.API_TOKEN}}"
|
|
371
203
|
```
|
|
372
204
|
|
|
373
|
-
|
|
374
|
-
|
|
375
|
-
```yaml
|
|
376
|
-
summary: Login with valid credentials succeeds
|
|
377
|
-
auth: false
|
|
378
|
-
```
|
|
379
|
-
|
|
380
|
-
**`verify` is worth the one extra call.** Without it, a wrong password surfaces as every test failing on a login wall — N failures, none naming the cause. With it, the run stops before the first test and says what happened. Authentication failure exits 2 and never reports as failing tests: a login you cannot complete says nothing about the code under review, so it must not produce a score.
|
|
381
|
-
|
|
382
|
-
Each test still gets its own browser context; it simply starts from the shared session rather than empty, so isolation is unchanged. Set `auth.cache: true` to reuse a session across runs — off by default, because an expired session produces failures at random points with nothing pointing at the cause.
|
|
383
|
-
|
|
384
|
-
**A captured session is a credential.** The file holds live cookies: whoever has it is signed in as that user. `init` git-ignores it; never commit one.
|
|
205
|
+
**`verify` is worth the extra call.** Without it a wrong password surfaces as every test failing on a login wall — N failures, none naming the cause. Authentication failure exits 2 and never reports as failing tests, because a login you cannot complete says nothing about the code under review.
|
|
385
206
|
|
|
386
|
-
|
|
207
|
+
**A captured session is a credential** — the file holds live cookies. `init` git-ignores it; never commit one.
|
|
387
208
|
|
|
388
209
|
## LLM providers (BYOK)
|
|
389
210
|
|
|
390
|
-
|
|
391
|
-
|
|
392
|
-
- **Anthropic** (`ANTHROPIC_API_KEY`)
|
|
393
|
-
- **OpenAI** (`OPENAI_API_KEY`)
|
|
394
|
-
- **Ollama** (local, no key needed)
|
|
395
|
-
|
|
396
|
-
### Configuring from the environment
|
|
397
|
-
|
|
398
|
-
You never have to commit a provider choice just to configure a pipeline. These variables override `.blastproof/config.yaml`, and precedence is **CLI flag > environment > file**:
|
|
211
|
+
**Anthropic** (`ANTHROPIC_API_KEY`), **OpenAI** (`OPENAI_API_KEY`), or **Ollama** (local, no key). Any setting can be overridden from the environment, with precedence **CLI flag > environment > file**:
|
|
399
212
|
|
|
400
213
|
| variable | overrides |
|
|
401
214
|
| --- | --- |
|
|
@@ -404,11 +217,9 @@ You never have to commit a provider choice just to configure a pipeline. These v
|
|
|
404
217
|
| `BLASTPROOF_LLM_MODEL` | the model name |
|
|
405
218
|
| `BLASTPROOF_LLM_BASE_URL` | the provider endpoint — *not* the app |
|
|
406
219
|
| `BLASTPROOF_LLM_API_KEY_ENV` | the **name** of the variable holding your key |
|
|
407
|
-
| `BLASTPROOF_MAX_LLM_CALLS`
|
|
408
|
-
| `BLASTPROOF_MAX_TOKENS` | `budget.max_tokens` |
|
|
409
|
-
| `BLASTPROOF_MAX_DURATION_S` | `budget.max_duration_s`, in seconds |
|
|
220
|
+
| `BLASTPROOF_MAX_LLM_CALLS` · `BLASTPROOF_MAX_TOKENS` · `BLASTPROOF_MAX_DURATION_S` | the budget fields |
|
|
410
221
|
|
|
411
|
-
|
|
222
|
+
So an OpenAI-compatible gateway needs no file edit:
|
|
412
223
|
|
|
413
224
|
```bash
|
|
414
225
|
export BLASTPROOF_LLM_PROVIDER=openai
|
|
@@ -418,33 +229,46 @@ export BLASTPROOF_LLM_API_KEY_ENV=OPENROUTER_API_KEY
|
|
|
418
229
|
blastproof run --impacted --min-score 80
|
|
419
230
|
```
|
|
420
231
|
|
|
421
|
-
|
|
232
|
+
The key itself is never read from a `BLASTPROOF_*` variable — you name *which* variable holds it, so errors can keep naming the one you chose.
|
|
422
233
|
|
|
423
|
-
##
|
|
234
|
+
## Trust boundaries
|
|
424
235
|
|
|
425
|
-
|
|
426
|
-
- [x] **M1** — `init` + `run`: YAML test runner with agentic LLM executor
|
|
427
|
-
- [x] **M2** — diff analysis, impact mapping (`run --impacted`) and test generation (`plan`)
|
|
428
|
-
- [x] **M3** — Reports (JUnit + HTML), priority-weighted score, `--min-score` gate, `blastproof test`
|
|
429
|
-
- [x] **M4** — [published to npm](https://www.npmjs.com/package/blastproof) and a consumable GitHub Action
|
|
430
|
-
- [ ] Next — impact by import graph, so a shared module's blast radius stops depending on hand-curated globs
|
|
431
|
-
- [ ] Post-MVP — VS Code extension, session replay, worker parallelism, PR comments
|
|
236
|
+
The application under test is not trusted input: its page content reaches the model, so a page that controls its own accessible text can try to influence the agent. Two things constrain that.
|
|
432
237
|
|
|
433
|
-
|
|
238
|
+
**The agent cannot leave your application.** `navigate` is bounded by `base_url`'s origin; an app spanning hosts declares them in `allowed_origins:`. This is enforced by comparison, not by asking the model nicely.
|
|
239
|
+
|
|
240
|
+
**Your secrets stay out of prompts.** `{{env.*}}` placeholders survive intact and are substituted at the moment of typing. Every value your tests or auth recipe reference is redacted from everything else crossing into a prompt — page snapshots included — in literal and percent-encoded form. Redaction matches known values, so treat it as a strong default rather than a guarantee against a hostile app.
|
|
241
|
+
|
|
242
|
+
The system prompt also tells the model that page content is data, never instruction. That raises the cost of casual injection and is **not** a boundary — the origin constraint is. Do not point blastproof at an application you would not run locally.
|
|
434
243
|
|
|
435
|
-
|
|
436
|
-
|
|
437
|
-
|
|
438
|
-
|
|
439
|
-
|
|
244
|
+
## blastproof tests itself
|
|
245
|
+
|
|
246
|
+
The **Dogfood** badge is blastproof running against the demo app in this repo — real Chromium, real model, scored and gated, with public logs. It catches real regressions rather than diffing strings: change the demo discount from 20% to 5% while the page still claims *"20% off"* and it reports
|
|
247
|
+
|
|
248
|
+
```
|
|
249
|
+
FAIL P0 Promo code SAVE20 applies a 20% discount in the cart
|
|
250
|
+
reason: the discount is currently -$6.00, but a 20% discount on
|
|
251
|
+
$120.00 should be -$24.00
|
|
252
|
+
```
|
|
253
|
+
|
|
254
|
+
No selector was updated to catch that. The agent read the value, did the arithmetic, and disagreed with the page.
|
|
255
|
+
|
|
256
|
+
Try it yourself:
|
|
257
|
+
|
|
258
|
+
```bash
|
|
259
|
+
git clone https://github.com/hamc/blastproof && cd blastproof
|
|
260
|
+
npm install && npm run build
|
|
261
|
+
node examples/demo-app/serve.mjs 4173 &
|
|
262
|
+
export ANTHROPIC_API_KEY=...
|
|
263
|
+
node dist/cli.js run
|
|
264
|
+
```
|
|
265
|
+
|
|
266
|
+
## Development
|
|
440
267
|
|
|
441
|
-
|
|
268
|
+
Built with AI assistance using spec-driven development: every change began as a written proposal with its design rationale, and those documents are kept rather than discarded. `openspec/` holds the reasoning behind each decision, including the alternatives that were rejected and why — start at [`AGENTS.md`](./AGENTS.md) for architecture, conventions and the contribution workflow. Open work lives in [issues](https://github.com/hamc/blastproof/issues).
|
|
442
269
|
|
|
443
270
|
```bash
|
|
444
|
-
|
|
445
|
-
npm install
|
|
446
|
-
npm run build
|
|
447
|
-
npm test
|
|
271
|
+
npm install && npm run build && npm test
|
|
448
272
|
```
|
|
449
273
|
|
|
450
274
|
## License
|