blastproof 0.2.2 → 0.4.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +132 -240
- package/dist/cli.js +604 -83
- package/dist/cli.js.map +1 -1
- package/package.json +1 -1
package/README.md
CHANGED
|
@@ -4,54 +4,51 @@
|
|
|
4
4
|
[](https://github.com/hamc/blastproof/actions/workflows/dogfood.yml)
|
|
5
5
|
[](./LICENSE)
|
|
6
6
|
|
|
7
|
-
**Open-source AI testing agent for pull requests.**
|
|
8
|
-
|
|
9
|
-
`blastproof` is an open-source AI QA agent for pull requests: it reads your PR diff, maps the blast radius, generates end-to-end tests in plain English, executes them on a real browser with self-healing, and scores the result before merge. 100% local, MIT licensed, bring your own LLM key.
|
|
7
|
+
**Open-source AI testing agent for pull requests.** Write end-to-end tests as plain English. An agent drives a real browser to run them, selects only the ones your diff can affect, and scores the result before merge.
|
|
10
8
|
|
|
11
9
|
```
|
|
12
10
|
git diff → impact mapping → test generation → agentic execution → report + score
|
|
13
11
|
```
|
|
14
12
|
|
|
15
|
-
|
|
16
|
-
|
|
17
|
-
1. **Reads the diff** — `blastproof test --base main` parses the branch diff and maps it to affected routes.
|
|
18
|
-
2. **Maps the blast radius** — traces changed files to the user journeys and routes most likely affected.
|
|
19
|
-
3. **Writes the missing tests** — drafts plain-English YAML for affected routes nothing covers. Drafts are yours to review; existing tests are never rewritten.
|
|
20
|
-
4. **Executes agentically** — an LLM-driven loop over Playwright resolves elements via the accessibility tree on every step. No static selectors to rot: the agent re-resolves when the UI shifts. Being model-driven it is not deterministic, which is why the free, deterministic impact analysis is what runs on every pull request.
|
|
21
|
-
5. **Reports & scores** — console, JUnit XML and HTML reports, plus a priority-weighted score that fails the run below `--min-score`.
|
|
13
|
+
100% local. MIT. Bring your own LLM key.
|
|
22
14
|
|
|
23
15
|
## Quick start
|
|
24
16
|
|
|
25
17
|
```bash
|
|
26
|
-
npm install -g blastproof
|
|
27
|
-
npx playwright install --with-deps chromium # one-time browser download
|
|
28
|
-
|
|
18
|
+
npm install -g blastproof # Node.js >= 20.19
|
|
29
19
|
cd your-project
|
|
30
|
-
blastproof init
|
|
31
|
-
# start your app, then point base_url at it in .blastproof/config.yaml
|
|
32
|
-
export ANTHROPIC_API_KEY=... # or OPENAI_API_KEY, or a local Ollama model
|
|
33
|
-
blastproof run # runs .blastproof/tests/**/*.yaml agentically
|
|
20
|
+
blastproof init # scaffolds .blastproof/
|
|
34
21
|
```
|
|
35
22
|
|
|
36
|
-
`
|
|
23
|
+
Point `base_url` at your running app in `.blastproof/config.yaml`, then check the setup — this needs **no API key and no browser**:
|
|
37
24
|
|
|
38
|
-
|
|
25
|
+
```bash
|
|
26
|
+
blastproof run --dry-run
|
|
27
|
+
```
|
|
39
28
|
|
|
40
|
-
|
|
29
|
+
To actually execute tests you need a browser and a model:
|
|
41
30
|
|
|
42
31
|
```bash
|
|
43
|
-
|
|
44
|
-
|
|
45
|
-
|
|
46
|
-
export ANTHROPIC_API_KEY=...
|
|
47
|
-
node dist/cli.js run
|
|
32
|
+
npx playwright install --with-deps chromium # NEEDS SUDO — see below
|
|
33
|
+
export ANTHROPIC_API_KEY=... # or OPENAI_API_KEY, or local Ollama
|
|
34
|
+
blastproof run
|
|
48
35
|
```
|
|
49
36
|
|
|
50
|
-
|
|
37
|
+
**No sudo?** `--with-deps` installs system libraries as root. Without it, run `npx playwright install chromium` and obtain `libnspr4`, `libnss3`, `libnssutil3` and `libasound2` however you can. Note that a useful half of blastproof needs neither browser nor key — see [Without a browser or a key](#without-a-browser-or-a-key).
|
|
38
|
+
|
|
39
|
+
Before `run`, `plan` or `test` do anything, they check what they are about to spend — the browser can launch, the model provider is reachable, `base_url` responds — and report every unmet one together, so a stopped app or a missing browser is never a wall you hit one crash at a time. A missing system library names the exact install command and says it needs root; nothing is installed on your behalf. Silent when everything is fine, and skipped entirely by `--dry-run`, which needs none of it.
|
|
40
|
+
|
|
41
|
+
## Does this fit your application?
|
|
42
|
+
|
|
43
|
+
**Your markup must be accessible — a hard requirement.** Elements are found by role, label or visible text from the accessibility tree. That is what removes selectors and survives redesigns; the cost is that an interface the accessibility tree cannot describe cannot be driven at all, and there is deliberately no CSS or XPath fallback. Icon-only buttons without accessible names, `div`-based controls and ARIA-less dropdowns simply cannot be targeted. Run an accessibility checker first — the result predicts how well this will work better than anything else.
|
|
51
44
|
|
|
52
|
-
|
|
45
|
+
**Not supported yet:** `iframe` content (so hosted payment widgets like Stripe Elements are invisible — an embedded checkout cannot be driven end to end), hover, scroll-to, drag and drop, file upload, multiple tabs, native `alert`/`confirm` dialogs. Page snapshots are capped at 200 lines by default, so very dense pages are truncated — raise it with `browser.max_snapshot_lines` if your pages need more; truncation is always marked in the snapshot so the model is never misled into thinking it saw the whole page.
|
|
53
46
|
|
|
54
|
-
|
|
47
|
+
`browser.timeout_ms` bounds every wait — resolving a target element from the accessibility tree, and navigation — not only the click or fill performed afterwards. Raise it for an application that is merely slow to hydrate; the trade-off is that a genuinely missing element then takes longer to fail. It never changes how many self-healing retries a step gets — waiting and retrying are deliberately separate.
|
|
48
|
+
|
|
49
|
+
## Writing tests
|
|
50
|
+
|
|
51
|
+
Tests live in `.blastproof/tests/` as plain-English YAML — no selectors:
|
|
55
52
|
|
|
56
53
|
```yaml
|
|
57
54
|
summary: Checkout with discount
|
|
@@ -65,136 +62,36 @@ steps:
|
|
|
65
62
|
- complete checkout
|
|
66
63
|
```
|
|
67
64
|
|
|
68
|
-
`priority` is P0–P2 (default P1)
|
|
65
|
+
`priority` is P0–P2 (default P1). `tags`, `setup` steps and `auth` are optional — `auth: false` runs the test signed out, which a login test needs. `routes` declares the URLs a test covers, which is what `--impacted` selects on; write route strings consistently, since they compare by exact equality (`/cart` ≠ `/cart/`).
|
|
69
66
|
|
|
70
|
-
##
|
|
71
|
-
|
|
72
|
-
| Command | Description |
|
|
73
|
-
| --- | --- |
|
|
74
|
-
| `blastproof init` | Scaffold `.blastproof/` config and sample tests (idempotent) |
|
|
75
|
-
| `blastproof run [--tag smoke] [--priority P0] [--query checkout]` | Run tests only — exit 0 pass, 1 fail, 2 usage/config error |
|
|
76
|
-
| `blastproof run --impacted [--base <ref>]` | Run only tests impacted by the diff vs the base ref (default `main`). Unrouted tests are skipped and reported; affected-but-uncovered routes are reported without failing the run |
|
|
77
|
-
| `blastproof run --dry-run` | Print the selection plan (affected routes, unmapped files, selected/skipped tests) and exit 0 — no browser launched, no LLM key needed |
|
|
78
|
-
| `blastproof run --url <url>` | Override `base_url` for this run only (e.g. a PR preview environment); the config file is never mutated |
|
|
79
|
-
| `blastproof plan [--base <ref>]` | Generate plain-English tests for affected routes no test covers yet. Prints drafts; nothing is written without `--write` |
|
|
80
|
-
| `blastproof plan --route <route>` | Generate for a route explicitly, skipping the diff (repeatable) — how you bootstrap coverage on an app with no suite yet |
|
|
81
|
-
| `blastproof plan --write` | Persist drafts to `.blastproof/tests/<route-slug>.yaml`. Never overwrites: a colliding filename fails that route |
|
|
82
|
-
| `blastproof run --min-score <n>` | Require a weighted score of at least `n` (0–100). **Replaces** the all-must-pass rule — see below |
|
|
83
|
-
| `blastproof run --junit [path]` | Write a JUnit XML report; without a path it lands in `.blastproof/reports/<session>/junit.xml` |
|
|
84
|
-
| `blastproof run --html [path]` | Write a self-contained HTML report with failure screenshots embedded inline |
|
|
85
|
-
| `blastproof test [--base <ref>]` | The full pipeline: run the tests covering the diff, then draft tests for the gaps |
|
|
86
|
-
|
|
87
|
-
### The full pipeline: `blastproof test`
|
|
88
|
-
|
|
89
|
-
One command for the whole loop — map the blast radius, run what covers it, draft what doesn't exist yet:
|
|
67
|
+
## Commands
|
|
90
68
|
|
|
91
69
|
```bash
|
|
92
|
-
blastproof
|
|
70
|
+
blastproof init # scaffold .blastproof/
|
|
71
|
+
blastproof run # run every test
|
|
72
|
+
blastproof run --impacted --base main # run only what the diff can affect
|
|
73
|
+
blastproof plan --base main # draft tests for uncovered routes
|
|
74
|
+
blastproof test --base main # run what covers the diff, then draft the gaps
|
|
93
75
|
```
|
|
94
76
|
|
|
95
|
-
|
|
96
|
-
|
|
97
|
-
1. **Verify** — executes the tests covering the affected routes, scores them, applies the gate
|
|
98
|
-
2. **Draft** — generates tests for affected routes no test covers, and prints them
|
|
99
|
-
|
|
100
|
-
**Generated drafts are never executed, and never affect the score.** That is deliberate. An unreviewed, model-written test in the merge path fails in two directions: a hallucinated expectation blocks a correct pull request, and a credulous one waves a broken change through while looking like coverage. Either costs more trust than the automation saves.
|
|
101
|
-
|
|
102
|
-
So read the result honestly: `test` does not make an uncovered route safe. It makes the gap visible, with a draft ready for you to review, and the score keeps describing only what was actually verified. Add `--write` to persist the drafts (never overwriting an existing file) and commit them once you have read them.
|
|
103
|
-
|
|
104
|
-
Exit codes: 2 usage/config/diff, 1 when the gate fails **or** a draft could not be generated, 0 otherwise.
|
|
105
|
-
|
|
106
|
-
### Reports
|
|
107
|
-
|
|
108
|
-
`--junit` is for CI; `--html` is for humans. The HTML report is a single self-contained file — inline CSS, screenshots embedded as data URIs, no scripts — so it opens offline, survives being moved, and uploads as one artifact. It leads with the score and gate verdict, sorts failures above passes, and expands each failure to its failing step, reason and screenshot.
|
|
109
|
-
|
|
110
|
-
### Generating tests with `plan`
|
|
111
|
-
|
|
112
|
-
`plan` closes the gap `run --impacted` reports. It takes the affected routes no test covers, loads each one in the browser, and asks the model to write a test from the page's real accessibility tree plus the changed files that made the route impacted — so the generated steps name controls that actually exist:
|
|
113
|
-
|
|
114
|
-
```bash
|
|
115
|
-
blastproof plan --base main # preview drafts for uncovered routes
|
|
116
|
-
blastproof plan --base main --write # persist them, then review and commit
|
|
117
|
-
blastproof plan --route /checkout # bootstrap a route without a diff
|
|
118
|
-
```
|
|
119
|
-
|
|
120
|
-
Drafts are **previews by default** — nothing touches disk until `--write`, and `--write` never overwrites an existing file, so a regeneration can't silently replace a test you edited by hand. Each written file carries a header recording its route, base ref and generation date. Review before committing: the steps are model-written and meant to be edited.
|
|
121
|
-
|
|
122
|
-
Exit codes: 0 when every route generated (or nothing needed coverage), 1 when a route failed, 2 on usage/config/diff errors. A route that fails to load never aborts the others.
|
|
123
|
-
|
|
124
|
-
`plan` uses the same `auth:` recipe as `run`, so a route behind a login is drafted from the real page rather than from the login wall.
|
|
125
|
-
|
|
126
|
-
### Closing the coverage hole
|
|
77
|
+
Common flags — `blastproof <command> --help` has the full list:
|
|
127
78
|
|
|
128
|
-
|
|
129
|
-
|
|
130
|
-
Each changed file is classified three ways:
|
|
131
|
-
|
|
132
|
-
| a changed file | means |
|
|
79
|
+
| flag | |
|
|
133
80
|
| --- | --- |
|
|
134
|
-
|
|
|
135
|
-
|
|
|
136
|
-
|
|
|
137
|
-
|
|
138
|
-
|
|
139
|
-
|
|
140
|
-
|
|
141
|
-
|
|
142
|
-
- "**/*.md"
|
|
143
|
-
- ".github/**"
|
|
144
|
-
```
|
|
145
|
-
|
|
146
|
-
```bash
|
|
147
|
-
blastproof run --impacted --fail-on-unmapped
|
|
148
|
-
```
|
|
81
|
+
| `--dry-run` | Print the selection (or, for `plan`, the routes it would draft) and exit. No browser, no API key |
|
|
82
|
+
| `--tag` · `--priority` · `--query` | Select a subset of tests |
|
|
83
|
+
| `--url <url>` | Override `base_url` for this run (e.g. a PR preview) |
|
|
84
|
+
| `--min-score <n>` | Gate on a weighted score instead of all-must-pass |
|
|
85
|
+
| `--fail-on-unmapped` | Fail when a changed file matches no `routes:` or `ignore:` glob |
|
|
86
|
+
| `--junit [path]` · `--html [path]` | Write reports |
|
|
87
|
+
| `--write` | `plan` only — persist drafts instead of previewing |
|
|
88
|
+
| `--max-llm-calls` · `--max-tokens` · `--max-duration` | Bound what a run may spend |
|
|
149
89
|
|
|
150
|
-
|
|
90
|
+
Exit codes: **0** pass, **1** the gate failed, **2** usage or config error.
|
|
151
91
|
|
|
152
|
-
|
|
92
|
+
**Generated drafts are never executed and never affect the score.** An unreviewed model-written test in the merge path fails in two directions: a hallucinated expectation blocks a correct PR, and a credulous one waves a broken change through while looking like coverage. `plan` makes the gap visible with a draft to review; it does not make an uncovered route safe.
|
|
153
93
|
|
|
154
|
-
|
|
155
|
-
|
|
156
|
-
### Score and merge gating
|
|
157
|
-
|
|
158
|
-
Every run ends with a score: the percentage of executed test **weight** that passed, where a test weighs 3 at P0, 2 at P1 and 1 at P2. Weighting is the point — a failing checkout costs three times a failing tooltip, so a pile of trivial passes can't hide a broken critical journey.
|
|
159
|
-
|
|
160
|
-
```bash
|
|
161
|
-
blastproof run # any failure exits 1 (strict, the default)
|
|
162
|
-
blastproof run --min-score 80 # passes at 80+, so one failing P2 is tolerated
|
|
163
|
-
blastproof run --min-score 100 # identical to the default strict behaviour
|
|
164
|
-
```
|
|
165
|
-
|
|
166
|
-
`--min-score` **replaces** the all-must-pass rule rather than adding to it. Without it, any failure exits 1. With it, the score alone decides — which is what lets you say "a P2 may break, a P0 may not" in one number. Only executed tests count: tests removed by `--tag`/`--priority`/`--query`, and tests skipped as unrouted under `--impacted`, are neither numerator nor denominator. A run that executed nothing scores 100 (the output says so explicitly), so a docs-only PR is never blocked.
|
|
167
|
-
|
|
168
|
-
For CI:
|
|
169
|
-
|
|
170
|
-
```bash
|
|
171
|
-
blastproof run --impacted --base "$BASE_REF" --min-score 80 --junit junit.xml
|
|
172
|
-
```
|
|
173
|
-
|
|
174
|
-
Exit 0 merge-able, 1 blocked, 2 usage/config error. The JUnit report carries the score as a `<property name="score">` so a parser can read it without scraping stdout, and tests skipped for having no `routes:` appear as `<skipped/>` cases — the coverage gap shows up in CI instead of vanishing.
|
|
175
|
-
|
|
176
|
-
## blastproof tests itself
|
|
177
|
-
|
|
178
|
-
The **Dogfood** badge above is blastproof running against the demo app in this repo: real Chromium, real LLM, plain-English tests, scored and gated. The run logs are public — the agent's reasoning, step by step, is there to read.
|
|
179
|
-
|
|
180
|
-
It catches real regressions rather than diffing strings. Changing the demo app's discount from 20% to 5%, while leaving the on-screen message still claiming *"Promo code SAVE20 applied: 20% off"*, produces:
|
|
181
|
-
|
|
182
|
-
```
|
|
183
|
-
FAIL P0 Promo code SAVE20 applies a 20% discount in the cart
|
|
184
|
-
failing step: verify a 20% discount of $24.00 is shown
|
|
185
|
-
reason: the discount is currently -$6.00, but a 20% discount on
|
|
186
|
-
$120.00 should be -$24.00
|
|
187
|
-
Score: 50 — min-score 80: FAIL (below threshold)
|
|
188
|
-
```
|
|
189
|
-
|
|
190
|
-
No selector was updated and no assertion was rewritten to catch that. The agent read the rendered value, did the arithmetic, and disagreed with the page.
|
|
191
|
-
|
|
192
|
-
Two workflows, split by what they cost:
|
|
193
|
-
|
|
194
|
-
- **Impact** — runs on every pull request, including forks. Deterministic and keyless: it reports the blast radius of the diff and which tests cover it, before anyone spends a token.
|
|
195
|
-
- **Dogfood** — runs daily and on demand. The agentic run needs an API key, so it stays out of the merge path: a non-deterministic model answer should never block a merge.
|
|
196
|
-
|
|
197
|
-
## GitHub Action
|
|
94
|
+
## In CI
|
|
198
95
|
|
|
199
96
|
```yaml
|
|
200
97
|
name: blastproof
|
|
@@ -206,87 +103,87 @@ jobs:
|
|
|
206
103
|
steps:
|
|
207
104
|
- uses: actions/checkout@v4
|
|
208
105
|
with:
|
|
209
|
-
fetch-depth: 0 # required
|
|
106
|
+
fetch-depth: 0 # required — the diff needs a merge-base
|
|
210
107
|
|
|
211
108
|
- run: npm start & # however your app boots
|
|
212
109
|
|
|
213
|
-
- uses: hamc/blastproof@v0.
|
|
110
|
+
- uses: hamc/blastproof@v0.4.0
|
|
214
111
|
with:
|
|
112
|
+
version: '0.4.0' # pin both when this gates merges
|
|
215
113
|
api-key: ${{ secrets.ANTHROPIC_API_KEY }}
|
|
216
114
|
base: ${{ github.event.pull_request.base.ref }}
|
|
217
115
|
min-score: '80'
|
|
218
|
-
fail-on-unmapped: 'true'
|
|
116
|
+
fail-on-unmapped: 'true'
|
|
219
117
|
```
|
|
220
118
|
|
|
221
|
-
|
|
119
|
+
A non-zero exit blocks the merge. The action outputs `score` (0–100, empty when no report was produced) for later steps. Full input list: [`action.yml`](./action.yml).
|
|
120
|
+
|
|
121
|
+
`fetch-depth: 0` is not optional — the default checkout is shallow and has no merge-base. The action detects this and fails immediately rather than letting it surface as a git error mid-run.
|
|
122
|
+
|
|
123
|
+
## Impact mapping
|
|
124
|
+
|
|
125
|
+
`--impacted` runs only the tests whose `routes:` intersect the routes your diff can affect, mapped from changed files by globs you maintain in `.blastproof/config.yaml`:
|
|
222
126
|
|
|
223
127
|
```yaml
|
|
224
|
-
|
|
225
|
-
|
|
226
|
-
|
|
227
|
-
|
|
228
|
-
|
|
229
|
-
|
|
230
|
-
|
|
231
|
-
|
|
232
|
-
|
|
128
|
+
routes:
|
|
129
|
+
"src/cart/**": ["/cart", "/checkout"]
|
|
130
|
+
ignore:
|
|
131
|
+
- "**/*.md"
|
|
132
|
+
```
|
|
133
|
+
|
|
134
|
+
Every changed file lands in one of three buckets: it matches `routes:` and contributes them, matches `ignore:` and is knowingly irrelevant, or **matches neither — nobody has said what it affects**. `--fail-on-unmapped` blocks on that third case, naming the files and both ways to resolve them.
|
|
135
|
+
|
|
136
|
+
**Nothing is ignored by default**, on purpose: a default that guesses on your behalf would hide the first files worth thinking about. The flag is additive — a run can meet `--min-score` and still be blocked here, because "the tests I ran passed" and "something changed that nobody classified" are different claims.
|
|
137
|
+
|
|
138
|
+
Its limit is worth knowing: it catches files that are *unclassified*, not *misclassified*. A shared module mapped to one route when it can break five still slips through. Impact by import graph is the fix, and blastproof does not do it yet.
|
|
139
|
+
|
|
140
|
+
## Score and merge gating
|
|
233
141
|
|
|
234
|
-
|
|
235
|
-
run: echo "scored ${{ steps.bp.outputs.score }}"
|
|
142
|
+
Each run scores the percentage of executed test **weight** that passed, weighing 3 at P0, 2 at P1, 1 at P2 — so a failing checkout costs three times a failing tooltip.
|
|
236
143
|
|
|
237
|
-
|
|
238
|
-
|
|
239
|
-
|
|
144
|
+
```bash
|
|
145
|
+
blastproof run # any failure exits 1 (strict, the default)
|
|
146
|
+
blastproof run --min-score 80 # one failing P2 is tolerated
|
|
240
147
|
```
|
|
241
148
|
|
|
242
|
-
|
|
243
|
-
| --- | --- |
|
|
244
|
-
| `api-key` | Your provider key, from a secret. Not needed for ollama |
|
|
245
|
-
| `provider` · `model` · `llm-base-url` | Override the committed config without editing it |
|
|
246
|
-
| `command` | `test` (default), `run` or `plan` |
|
|
247
|
-
| `base` | Base ref for the diff |
|
|
248
|
-
| `url` | Base URL of the app under test |
|
|
249
|
-
| `min-score` | Weighted score threshold — replaces the all-must-pass rule |
|
|
250
|
-
| `fail-on-unmapped` | Fail on a changed file matching no `routes:` or `ignore:` glob |
|
|
251
|
-
| `write` | Persist generated drafts instead of previewing |
|
|
252
|
-
| `junit` · `html` | Where to write the reports |
|
|
253
|
-
| `version` | Which blastproof release to install (default `latest`) |
|
|
254
|
-
| `install-browser` | Set `false` if the workflow already installed Chromium |
|
|
255
|
-
| `working-directory` | Directory containing `.blastproof/` |
|
|
149
|
+
`--min-score` **replaces** the all-must-pass rule rather than adding to it. Only executed tests count: filtered and unrouted tests are neither numerator nor denominator, and a run that executed nothing scores 100, so a docs-only PR is never blocked. JUnit carries the score as a `<property name="score">`, and unrouted tests appear as `<skipped/>` so the coverage gap shows up in CI rather than vanishing.
|
|
256
150
|
|
|
257
|
-
|
|
151
|
+
## Without a browser or a key
|
|
258
152
|
|
|
259
|
-
|
|
153
|
+
Half of blastproof is deterministic and free. These need no model, no browser and no network:
|
|
260
154
|
|
|
261
|
-
|
|
155
|
+
```bash
|
|
156
|
+
blastproof run --dry-run # what would run
|
|
157
|
+
blastproof run --impacted --dry-run # + which routes the diff touches
|
|
158
|
+
blastproof run --impacted --fail-on-unmapped --dry-run # + gate on unclassified files
|
|
159
|
+
blastproof plan --base main --dry-run # affected routes no test covers, no key needed
|
|
160
|
+
```
|
|
262
161
|
|
|
263
|
-
|
|
162
|
+
They report affected routes, files nobody has classified, and affected routes no test covers — a coverage-gap report with an exit code, useful even on a repo whose suite is Playwright or Cypress.
|
|
264
163
|
|
|
265
|
-
|
|
164
|
+
## Bounding a run
|
|
266
165
|
|
|
267
|
-
|
|
166
|
+
Nothing stops a run by default. `budget:` puts a ceiling on `run`, `plan` and `test` alike — every model call any of them makes is counted:
|
|
268
167
|
|
|
269
168
|
```yaml
|
|
270
|
-
|
|
271
|
-
|
|
169
|
+
budget:
|
|
170
|
+
max_llm_calls: 500
|
|
171
|
+
max_tokens: 2000000
|
|
172
|
+
max_duration_s: 900
|
|
272
173
|
```
|
|
273
174
|
|
|
274
|
-
|
|
175
|
+
Each limit is optional; with none set, nothing binds. They count **calls and tokens, not currency** — a price table keyed by model and provider goes stale the day a provider reprices, and a limit that quietly stops meaning what it says is worse than none, because it is trusted.
|
|
275
176
|
|
|
276
|
-
|
|
177
|
+
Exhausting a budget **stops the run; it does not fail a test.** Running out of quota says nothing about the code under review. Unreached tests are reported as `not run`, a third state excluded from the score entirely, and the process exits 1 unconditionally — `--min-score` cannot rescue it, because the tests that finished are whichever ran first, not a representative sample.
|
|
277
178
|
|
|
278
|
-
|
|
279
|
-
|
|
280
|
-
The system prompt also tells the model that page content is data under test and never an instruction to obey. That raises the cost of a casual injection and is **not** a security boundary — a determined one will get past prompt wording. The origin constraint is the boundary; treat the rest as hygiene, and do not point blastproof at an application you would not run locally.
|
|
179
|
+
`--dry-run` reports the ceiling before you spend anything.
|
|
281
180
|
|
|
282
181
|
## Testing behind a login
|
|
283
182
|
|
|
284
|
-
|
|
285
|
-
|
|
286
|
-
Pick exactly one strategy:
|
|
183
|
+
Declare a recipe once; blastproof signs in one time per run and reuses that session for every test and for `plan`. Pick exactly one strategy:
|
|
287
184
|
|
|
288
185
|
```yaml
|
|
289
|
-
# 1) A plain-English
|
|
186
|
+
# 1) A plain-English journey — form login, or anything a person can click through
|
|
290
187
|
auth:
|
|
291
188
|
steps:
|
|
292
189
|
- navigate to /login
|
|
@@ -305,32 +202,13 @@ auth:
|
|
|
305
202
|
Authorization: "Bearer {{env.API_TOKEN}}"
|
|
306
203
|
```
|
|
307
204
|
|
|
308
|
-
|
|
205
|
+
**`verify` is worth the extra call.** Without it a wrong password surfaces as every test failing on a login wall — N failures, none naming the cause. Authentication failure exits 2 and never reports as failing tests, because a login you cannot complete says nothing about the code under review.
|
|
309
206
|
|
|
310
|
-
|
|
311
|
-
summary: Login with valid credentials succeeds
|
|
312
|
-
auth: false
|
|
313
|
-
```
|
|
314
|
-
|
|
315
|
-
**`verify` is worth the one extra call.** Without it, a wrong password surfaces as every test failing on a login wall — N failures, none naming the cause. With it, the run stops before the first test and says what happened. Authentication failure exits 2 and never reports as failing tests: a login you cannot complete says nothing about the code under review, so it must not produce a score.
|
|
316
|
-
|
|
317
|
-
Each test still gets its own browser context; it simply starts from the shared session rather than empty, so isolation is unchanged. Set `auth.cache: true` to reuse a session across runs — off by default, because an expired session produces failures at random points with nothing pointing at the cause.
|
|
318
|
-
|
|
319
|
-
**A captured session is a credential.** The file holds live cookies: whoever has it is signed in as that user. `init` git-ignores it; never commit one.
|
|
320
|
-
|
|
321
|
-
> **Note on self-healing:** the executor recovers from failed steps by re-reading the page, which means it can complete a login using credentials the page itself displays — some apps show demo credentials on the sign-in form. That is the self-healing loop working as designed, but it does mean a deliberately-wrong password is not a reliable way to test your auth failure path.
|
|
207
|
+
**A captured session is a credential** — the file holds live cookies. `init` git-ignores it; never commit one.
|
|
322
208
|
|
|
323
209
|
## LLM providers (BYOK)
|
|
324
210
|
|
|
325
|
-
|
|
326
|
-
|
|
327
|
-
- **Anthropic** (`ANTHROPIC_API_KEY`)
|
|
328
|
-
- **OpenAI** (`OPENAI_API_KEY`)
|
|
329
|
-
- **Ollama** (local, no key needed)
|
|
330
|
-
|
|
331
|
-
### Configuring from the environment
|
|
332
|
-
|
|
333
|
-
You never have to commit a provider choice just to configure a pipeline. These variables override `.blastproof/config.yaml`, and precedence is **CLI flag > environment > file**:
|
|
211
|
+
**Anthropic** (`ANTHROPIC_API_KEY`), **OpenAI** (`OPENAI_API_KEY`), or **Ollama** (local, no key). Any setting can be overridden from the environment, with precedence **CLI flag > environment > file**:
|
|
334
212
|
|
|
335
213
|
| variable | overrides |
|
|
336
214
|
| --- | --- |
|
|
@@ -339,8 +217,9 @@ You never have to commit a provider choice just to configure a pipeline. These v
|
|
|
339
217
|
| `BLASTPROOF_LLM_MODEL` | the model name |
|
|
340
218
|
| `BLASTPROOF_LLM_BASE_URL` | the provider endpoint — *not* the app |
|
|
341
219
|
| `BLASTPROOF_LLM_API_KEY_ENV` | the **name** of the variable holding your key |
|
|
220
|
+
| `BLASTPROOF_MAX_LLM_CALLS` · `BLASTPROOF_MAX_TOKENS` · `BLASTPROOF_MAX_DURATION_S` | the budget fields |
|
|
342
221
|
|
|
343
|
-
|
|
222
|
+
So an OpenAI-compatible gateway needs no file edit:
|
|
344
223
|
|
|
345
224
|
```bash
|
|
346
225
|
export BLASTPROOF_LLM_PROVIDER=openai
|
|
@@ -350,33 +229,46 @@ export BLASTPROOF_LLM_API_KEY_ENV=OPENROUTER_API_KEY
|
|
|
350
229
|
blastproof run --impacted --min-score 80
|
|
351
230
|
```
|
|
352
231
|
|
|
353
|
-
|
|
232
|
+
The key itself is never read from a `BLASTPROOF_*` variable — you name *which* variable holds it, so errors can keep naming the one you chose.
|
|
354
233
|
|
|
355
|
-
##
|
|
234
|
+
## Trust boundaries
|
|
356
235
|
|
|
357
|
-
|
|
358
|
-
- [x] **M1** — `init` + `run`: YAML test runner with agentic LLM executor
|
|
359
|
-
- [x] **M2** — diff analysis, impact mapping (`run --impacted`) and test generation (`plan`)
|
|
360
|
-
- [x] **M3** — Reports (JUnit + HTML), priority-weighted score, `--min-score` gate, `blastproof test`
|
|
361
|
-
- [x] **M4** — [published to npm](https://www.npmjs.com/package/blastproof) and a consumable GitHub Action
|
|
362
|
-
- [ ] Next — impact by import graph, so a shared module's blast radius stops depending on hand-curated globs
|
|
363
|
-
- [ ] Post-MVP — VS Code extension, session replay, worker parallelism, PR comments
|
|
236
|
+
The application under test is not trusted input: its page content reaches the model, so a page that controls its own accessible text can try to influence the agent. Two things constrain that.
|
|
364
237
|
|
|
365
|
-
|
|
238
|
+
**The agent cannot leave your application.** `navigate` is bounded by `base_url`'s origin; an app spanning hosts declares them in `allowed_origins:`. This is enforced by comparison, not by asking the model nicely.
|
|
239
|
+
|
|
240
|
+
**Your secrets stay out of prompts.** `{{env.*}}` placeholders survive intact and are substituted at the moment of typing. Every value your tests or auth recipe reference is redacted from everything else crossing into a prompt — page snapshots included — in literal and percent-encoded form. Redaction matches known values, so treat it as a strong default rather than a guarantee against a hostile app.
|
|
241
|
+
|
|
242
|
+
The system prompt also tells the model that page content is data, never instruction. That raises the cost of casual injection and is **not** a boundary — the origin constraint is. Do not point blastproof at an application you would not run locally.
|
|
243
|
+
|
|
244
|
+
## blastproof tests itself
|
|
245
|
+
|
|
246
|
+
The **Dogfood** badge is blastproof running against the demo app in this repo — real Chromium, real model, scored and gated, with public logs. It catches real regressions rather than diffing strings: change the demo discount from 20% to 5% while the page still claims *"20% off"* and it reports
|
|
247
|
+
|
|
248
|
+
```
|
|
249
|
+
FAIL P0 Promo code SAVE20 applies a 20% discount in the cart
|
|
250
|
+
reason: the discount is currently -$6.00, but a 20% discount on
|
|
251
|
+
$120.00 should be -$24.00
|
|
252
|
+
```
|
|
253
|
+
|
|
254
|
+
No selector was updated to catch that. The agent read the value, did the arithmetic, and disagreed with the page.
|
|
366
255
|
|
|
367
|
-
|
|
368
|
-
|
|
369
|
-
|
|
370
|
-
|
|
371
|
-
|
|
256
|
+
Try it yourself:
|
|
257
|
+
|
|
258
|
+
```bash
|
|
259
|
+
git clone https://github.com/hamc/blastproof && cd blastproof
|
|
260
|
+
npm install && npm run build
|
|
261
|
+
node examples/demo-app/serve.mjs 4173 &
|
|
262
|
+
export ANTHROPIC_API_KEY=...
|
|
263
|
+
node dist/cli.js run
|
|
264
|
+
```
|
|
265
|
+
|
|
266
|
+
## Development
|
|
372
267
|
|
|
373
|
-
|
|
268
|
+
Built with AI assistance using spec-driven development: every change began as a written proposal with its design rationale, and those documents are kept rather than discarded. `openspec/` holds the reasoning behind each decision, including the alternatives that were rejected and why — start at [`AGENTS.md`](./AGENTS.md) for architecture, conventions and the contribution workflow. Open work lives in [issues](https://github.com/hamc/blastproof/issues).
|
|
374
269
|
|
|
375
270
|
```bash
|
|
376
|
-
|
|
377
|
-
npm install
|
|
378
|
-
npm run build
|
|
379
|
-
npm test
|
|
271
|
+
npm install && npm run build && npm test
|
|
380
272
|
```
|
|
381
273
|
|
|
382
274
|
## License
|