blastproof 0.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/LICENSE +21 -0
- package/README.md +287 -0
- package/dist/cli.js +2067 -0
- package/dist/cli.js.map +1 -0
- package/dist/snapshot-CAIB2OHX.js +43 -0
- package/dist/snapshot-CAIB2OHX.js.map +1 -0
- package/package.json +67 -0
package/LICENSE
ADDED
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
MIT License
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2026 blastproof contributors
|
|
4
|
+
|
|
5
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
6
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
7
|
+
in the Software without restriction, including without limitation the rights
|
|
8
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
9
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
10
|
+
furnished to do so, subject to the following conditions:
|
|
11
|
+
|
|
12
|
+
The above copyright notice and this permission notice shall be included in all
|
|
13
|
+
copies or substantial portions of the Software.
|
|
14
|
+
|
|
15
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
16
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
17
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
18
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
19
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
20
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
21
|
+
SOFTWARE.
|
package/README.md
ADDED
|
@@ -0,0 +1,287 @@
|
|
|
1
|
+
# blastproof
|
|
2
|
+
|
|
3
|
+
[](https://github.com/hamc/blastproof/actions/workflows/ci.yml)
|
|
4
|
+
[](https://github.com/hamc/blastproof/actions/workflows/dogfood.yml)
|
|
5
|
+
[](./LICENSE)
|
|
6
|
+
|
|
7
|
+
**Open-source AI testing agent for pull requests.** Diff in, confidence out — no test scripts to write or maintain.
|
|
8
|
+
|
|
9
|
+
`blastproof` is an open-source AI QA agent for pull requests: it reads your PR diff, maps the blast radius, generates end-to-end tests in plain English, executes them on a real browser with self-healing, and scores the result before merge. 100% local, MIT licensed, bring your own LLM key.
|
|
10
|
+
|
|
11
|
+
```
|
|
12
|
+
git diff → impact mapping → test generation → agentic execution → report + score
|
|
13
|
+
```
|
|
14
|
+
|
|
15
|
+
## How it works
|
|
16
|
+
|
|
17
|
+
1. **Reads the diff** — `blastproof test --base main` parses the branch diff and maps it to affected routes.
|
|
18
|
+
2. **Maps the blast radius** — traces changed files to the user journeys and routes most likely affected.
|
|
19
|
+
3. **Generates tests** — writes/updates plain-English YAML tests in `.blastproof/tests/`.
|
|
20
|
+
4. **Executes agentically** — an LLM-driven loop over Playwright resolves elements via the accessibility tree on every step. No static selectors, so no flakiness: the agent re-resolves when the UI shifts.
|
|
21
|
+
5. **Reports & scores** — console, JUnit XML and HTML reports, plus a priority-weighted score that fails the run below `--min-score`.
|
|
22
|
+
|
|
23
|
+
## Quick start
|
|
24
|
+
|
|
25
|
+
```bash
|
|
26
|
+
npm install -g blastproof # requires Node.js >= 20.19
|
|
27
|
+
npx playwright install --with-deps chromium # one-time browser download
|
|
28
|
+
cd your-project
|
|
29
|
+
blastproof init
|
|
30
|
+
export ANTHROPIC_API_KEY=... # or OPENAI_API_KEY, or a local Ollama model
|
|
31
|
+
blastproof run # runs .blastproof/tests/**/*.yaml agentically
|
|
32
|
+
```
|
|
33
|
+
|
|
34
|
+
Published with [npm provenance](https://docs.npmjs.com/generating-provenance-statements), so the tarball is verifiably built from this repository.
|
|
35
|
+
|
|
36
|
+
Try it locally without your own app — this repo ships a demo shop:
|
|
37
|
+
|
|
38
|
+
```bash
|
|
39
|
+
node examples/demo-app/serve.mjs 4173 & # home, login and cart + promo-code pages
|
|
40
|
+
blastproof init
|
|
41
|
+
export ANTHROPIC_API_KEY=...
|
|
42
|
+
blastproof run
|
|
43
|
+
```
|
|
44
|
+
|
|
45
|
+
> **Status:** the full pipeline works today — `init`, `run` (including `--impacted`), `plan` and `test`, with JUnit and HTML reports. The consumable GitHub Action is next.
|
|
46
|
+
|
|
47
|
+
## Test format
|
|
48
|
+
|
|
49
|
+
Tests live in `.blastproof/tests/` as plain-English YAML — no selectors, no framework lock-in:
|
|
50
|
+
|
|
51
|
+
```yaml
|
|
52
|
+
summary: Checkout with discount
|
|
53
|
+
priority: P0
|
|
54
|
+
tags: [checkout, discount]
|
|
55
|
+
routes: ["/cart", "/checkout"]
|
|
56
|
+
steps:
|
|
57
|
+
- add item to cart
|
|
58
|
+
- apply promo code SAVE20
|
|
59
|
+
- verify a 20% discount is applied
|
|
60
|
+
- complete checkout
|
|
61
|
+
```
|
|
62
|
+
|
|
63
|
+
`priority` is P0–P2 (default P1); `tags` and `setup` steps are optional. `routes` (optional) declares the URLs a test covers: `blastproof run --impacted` runs only tests whose `routes:` intersect the routes affected by your PR diff (mapped from changed files via the `routes:` globs in `.blastproof/config.yaml`). Route strings compare by exact equality (`/cart` ≠ `/cart/` — write them consistently). Tests without `routes:` are skipped and reported under `--impacted`.
|
|
64
|
+
|
|
65
|
+
## CLI
|
|
66
|
+
|
|
67
|
+
| Command | Description |
|
|
68
|
+
| --- | --- |
|
|
69
|
+
| `blastproof init` | Scaffold `.blastproof/` config and sample tests (idempotent) |
|
|
70
|
+
| `blastproof run [--tag smoke] [--priority P0] [--query checkout]` | Run tests only — exit 0 pass, 1 fail, 2 usage/config error |
|
|
71
|
+
| `blastproof run --impacted [--base <ref>]` | Run only tests impacted by the diff vs the base ref (default `main`). Unrouted tests are skipped and reported; affected-but-uncovered routes are reported without failing the run |
|
|
72
|
+
| `blastproof run --dry-run` | Print the selection plan (affected routes, unmapped files, selected/skipped tests) and exit 0 — no browser launched, no LLM key needed |
|
|
73
|
+
| `blastproof run --url <url>` | Override `base_url` for this run only (e.g. a PR preview environment); the config file is never mutated |
|
|
74
|
+
| `blastproof plan [--base <ref>]` | Generate plain-English tests for affected routes no test covers yet. Prints drafts; nothing is written without `--write` |
|
|
75
|
+
| `blastproof plan --route <route>` | Generate for a route explicitly, skipping the diff (repeatable) — how you bootstrap coverage on an app with no suite yet |
|
|
76
|
+
| `blastproof plan --write` | Persist drafts to `.blastproof/tests/<route-slug>.yaml`. Never overwrites: a colliding filename fails that route |
|
|
77
|
+
| `blastproof run --min-score <n>` | Require a weighted score of at least `n` (0–100). **Replaces** the all-must-pass rule — see below |
|
|
78
|
+
| `blastproof run --junit [path]` | Write a JUnit XML report; without a path it lands in `.blastproof/reports/<session>/junit.xml` |
|
|
79
|
+
| `blastproof run --html [path]` | Write a self-contained HTML report with failure screenshots embedded inline |
|
|
80
|
+
| `blastproof test [--base <ref>]` | The full pipeline: run the tests covering the diff, then draft tests for the gaps |
|
|
81
|
+
|
|
82
|
+
### The full pipeline: `blastproof test`
|
|
83
|
+
|
|
84
|
+
One command for the whole loop — map the blast radius, run what covers it, draft what doesn't exist yet:
|
|
85
|
+
|
|
86
|
+
```bash
|
|
87
|
+
blastproof test --base main --min-score 80 --junit junit.xml --html report.html
|
|
88
|
+
```
|
|
89
|
+
|
|
90
|
+
It does two things and reports them separately:
|
|
91
|
+
|
|
92
|
+
1. **Verify** — executes the tests covering the affected routes, scores them, applies the gate
|
|
93
|
+
2. **Draft** — generates tests for affected routes no test covers, and prints them
|
|
94
|
+
|
|
95
|
+
**Generated drafts are never executed, and never affect the score.** That is deliberate. An unreviewed, model-written test in the merge path fails in two directions: a hallucinated expectation blocks a correct pull request, and a credulous one waves a broken change through while looking like coverage. Either costs more trust than the automation saves.
|
|
96
|
+
|
|
97
|
+
So read the result honestly: `test` does not make an uncovered route safe. It makes the gap visible, with a draft ready for you to review, and the score keeps describing only what was actually verified. Add `--write` to persist the drafts (never overwriting an existing file) and commit them once you have read them.
|
|
98
|
+
|
|
99
|
+
Exit codes: 2 usage/config/diff, 1 when the gate fails **or** a draft could not be generated, 0 otherwise.
|
|
100
|
+
|
|
101
|
+
### Reports
|
|
102
|
+
|
|
103
|
+
`--junit` is for CI; `--html` is for humans. The HTML report is a single self-contained file — inline CSS, screenshots embedded as data URIs, no scripts — so it opens offline, survives being moved, and uploads as one artifact. It leads with the score and gate verdict, sorts failures above passes, and expands each failure to its failing step, reason and screenshot.
|
|
104
|
+
|
|
105
|
+
### Generating tests with `plan`
|
|
106
|
+
|
|
107
|
+
`plan` closes the gap `run --impacted` reports. It takes the affected routes no test covers, loads each one in the browser, and asks the model to write a test from the page's real accessibility tree plus the changed files that made the route impacted — so the generated steps name controls that actually exist:
|
|
108
|
+
|
|
109
|
+
```bash
|
|
110
|
+
blastproof plan --base main # preview drafts for uncovered routes
|
|
111
|
+
blastproof plan --base main --write # persist them, then review and commit
|
|
112
|
+
blastproof plan --route /checkout # bootstrap a route without a diff
|
|
113
|
+
```
|
|
114
|
+
|
|
115
|
+
Drafts are **previews by default** — nothing touches disk until `--write`, and `--write` never overwrites an existing file, so a regeneration can't silently replace a test you edited by hand. Each written file carries a header recording its route, base ref and generation date. Review before committing: the steps are model-written and meant to be edited.
|
|
116
|
+
|
|
117
|
+
Exit codes: 0 when every route generated (or nothing needed coverage), 1 when a route failed, 2 on usage/config/diff errors. A route that fails to load never aborts the others.
|
|
118
|
+
|
|
119
|
+
**Known limitation:** a route behind authentication snapshots as the login wall, so its draft describes logging in rather than the feature. The `auth` config recipe is not applied by the planner yet — generate those routes after an auth session lands, or write them by hand.
|
|
120
|
+
|
|
121
|
+
### Closing the coverage hole
|
|
122
|
+
|
|
123
|
+
Impact mapping has a failure mode worth understanding. A changed file that matches no `routes:` glob contributes no affected routes — so a diff touching only a shared module selects nothing, scores 100 because nothing executed, and merges green. The information is printed, but nobody reads a passing run.
|
|
124
|
+
|
|
125
|
+
Each changed file is classified three ways:
|
|
126
|
+
|
|
127
|
+
| a changed file | means |
|
|
128
|
+
| --- | --- |
|
|
129
|
+
| matches a `routes:` glob | contributes its routes |
|
|
130
|
+
| matches an `ignore:` glob | knowingly irrelevant to any page |
|
|
131
|
+
| matches neither | **nobody has said what this affects** |
|
|
132
|
+
|
|
133
|
+
```yaml
|
|
134
|
+
routes:
|
|
135
|
+
"src/cart/**": ["/cart", "/checkout"]
|
|
136
|
+
ignore:
|
|
137
|
+
- "**/*.md"
|
|
138
|
+
- ".github/**"
|
|
139
|
+
```
|
|
140
|
+
|
|
141
|
+
```bash
|
|
142
|
+
blastproof run --impacted --fail-on-unmapped
|
|
143
|
+
```
|
|
144
|
+
|
|
145
|
+
The flag fails the run on the third case only, naming the files and both ways to resolve them. `ignore:` is what makes that signal survivable — without it the flag would fire on every README edit and get switched off within a day, and a disabled gate protects nothing.
|
|
146
|
+
|
|
147
|
+
Nothing is ignored by default, on purpose: a file nobody has classified is exactly the risk the flag exists to surface, and a default that guesses on your behalf would hide the first files worth thinking about. This flag is **additive** — a run can meet its `--min-score` and still be blocked here, because "the tests I ran passed" and "something changed that nobody has classified" are different claims.
|
|
148
|
+
|
|
149
|
+
Be clear on its limit: it catches files that are *unclassified*, not files that are *misclassified*. A shared module mapped to one route when it can break five will still slip through. Mapping by import graph is the answer to that, and blastproof does not do it yet.
|
|
150
|
+
|
|
151
|
+
### Score and merge gating
|
|
152
|
+
|
|
153
|
+
Every run ends with a score: the percentage of executed test **weight** that passed, where a test weighs 3 at P0, 2 at P1 and 1 at P2. Weighting is the point — a failing checkout costs three times a failing tooltip, so a pile of trivial passes can't hide a broken critical journey.
|
|
154
|
+
|
|
155
|
+
```bash
|
|
156
|
+
blastproof run # any failure exits 1 (strict, the default)
|
|
157
|
+
blastproof run --min-score 80 # passes at 80+, so one failing P2 is tolerated
|
|
158
|
+
blastproof run --min-score 100 # identical to the default strict behaviour
|
|
159
|
+
```
|
|
160
|
+
|
|
161
|
+
`--min-score` **replaces** the all-must-pass rule rather than adding to it. Without it, any failure exits 1. With it, the score alone decides — which is what lets you say "a P2 may break, a P0 may not" in one number. Only executed tests count: tests removed by `--tag`/`--priority`/`--query`, and tests skipped as unrouted under `--impacted`, are neither numerator nor denominator. A run that executed nothing scores 100 (the output says so explicitly), so a docs-only PR is never blocked.
|
|
162
|
+
|
|
163
|
+
For CI:
|
|
164
|
+
|
|
165
|
+
```bash
|
|
166
|
+
blastproof run --impacted --base "$BASE_REF" --min-score 80 --junit junit.xml
|
|
167
|
+
```
|
|
168
|
+
|
|
169
|
+
Exit 0 merge-able, 1 blocked, 2 usage/config error. The JUnit report carries the score as a `<property name="score">` so a parser can read it without scraping stdout, and tests skipped for having no `routes:` appear as `<skipped/>` cases — the coverage gap shows up in CI instead of vanishing.
|
|
170
|
+
|
|
171
|
+
## blastproof tests itself
|
|
172
|
+
|
|
173
|
+
The **Dogfood** badge above is blastproof running against the demo app in this repo: real Chromium, real LLM, plain-English tests, scored and gated. The run logs are public — the agent's reasoning, step by step, is there to read.
|
|
174
|
+
|
|
175
|
+
It catches real regressions rather than diffing strings. Changing the demo app's discount from 20% to 5%, while leaving the on-screen message still claiming *"Promo code SAVE20 applied: 20% off"*, produces:
|
|
176
|
+
|
|
177
|
+
```
|
|
178
|
+
FAIL P0 Promo code SAVE20 applies a 20% discount in the cart
|
|
179
|
+
failing step: verify a 20% discount of $24.00 is shown
|
|
180
|
+
reason: the discount is currently -$6.00, but a 20% discount on
|
|
181
|
+
$120.00 should be -$24.00
|
|
182
|
+
Score: 50 — min-score 80: FAIL (below threshold)
|
|
183
|
+
```
|
|
184
|
+
|
|
185
|
+
No selector was updated and no assertion was rewritten to catch that. The agent read the rendered value, did the arithmetic, and disagreed with the page.
|
|
186
|
+
|
|
187
|
+
Two workflows, split by what they cost:
|
|
188
|
+
|
|
189
|
+
- **Impact** — runs on every pull request, including forks. Deterministic and keyless: it reports the blast radius of the diff and which tests cover it, before anyone spends a token.
|
|
190
|
+
- **Dogfood** — runs daily and on demand. The agentic run needs an API key, so it stays out of the merge path: a non-deterministic model answer should never block a merge.
|
|
191
|
+
|
|
192
|
+
## Testing behind a login
|
|
193
|
+
|
|
194
|
+
Most of a product lives behind authentication. Declare a recipe once and blastproof signs in a single time per run, then reuses that session for every test **and** for `plan` — so generated drafts describe the actual feature instead of the login wall.
|
|
195
|
+
|
|
196
|
+
Pick exactly one strategy:
|
|
197
|
+
|
|
198
|
+
```yaml
|
|
199
|
+
# 1) A plain-English login journey — form login, or anything a person can click through
|
|
200
|
+
auth:
|
|
201
|
+
steps:
|
|
202
|
+
- navigate to /login
|
|
203
|
+
- fill the email field with {{env.TEST_EMAIL}}
|
|
204
|
+
- fill the password field with {{env.TEST_PASSWORD}}
|
|
205
|
+
- submit the login form
|
|
206
|
+
verify: a signed-in indicator is visible # optional, strongly recommended
|
|
207
|
+
|
|
208
|
+
# 2) A session captured by hand — for SSO, MFA or magic links
|
|
209
|
+
auth:
|
|
210
|
+
storage_state: .blastproof/auth.json
|
|
211
|
+
|
|
212
|
+
# 3) Static values — for token-based apps
|
|
213
|
+
auth:
|
|
214
|
+
headers:
|
|
215
|
+
Authorization: "Bearer {{env.API_TOKEN}}"
|
|
216
|
+
```
|
|
217
|
+
|
|
218
|
+
A test that exercises the login itself must start signed out:
|
|
219
|
+
|
|
220
|
+
```yaml
|
|
221
|
+
summary: Login with valid credentials succeeds
|
|
222
|
+
auth: false
|
|
223
|
+
```
|
|
224
|
+
|
|
225
|
+
**`verify` is worth the one extra call.** Without it, a wrong password surfaces as every test failing on a login wall — N failures, none naming the cause. With it, the run stops before the first test and says what happened. Authentication failure exits 2 and never reports as failing tests: a login you cannot complete says nothing about the code under review, so it must not produce a score.
|
|
226
|
+
|
|
227
|
+
Each test still gets its own browser context; it simply starts from the shared session rather than empty, so isolation is unchanged. Set `auth.cache: true` to reuse a session across runs — off by default, because an expired session produces failures at random points with nothing pointing at the cause.
|
|
228
|
+
|
|
229
|
+
**A captured session is a credential.** The file holds live cookies: whoever has it is signed in as that user. `init` git-ignores it; never commit one.
|
|
230
|
+
|
|
231
|
+
> **Note on self-healing:** the executor recovers from failed steps by re-reading the page, which means it can complete a login using credentials the page itself displays — some apps show demo credentials on the sign-in form. That is the self-healing loop working as designed, but it does mean a deliberately-wrong password is not a reliable way to test your auth failure path.
|
|
232
|
+
|
|
233
|
+
## LLM providers (BYOK)
|
|
234
|
+
|
|
235
|
+
Bring your own key — runs 100% locally:
|
|
236
|
+
|
|
237
|
+
- **Anthropic** (`ANTHROPIC_API_KEY`)
|
|
238
|
+
- **OpenAI** (`OPENAI_API_KEY`)
|
|
239
|
+
- **Ollama** (local, no key needed)
|
|
240
|
+
|
|
241
|
+
### Configuring from the environment
|
|
242
|
+
|
|
243
|
+
You never have to commit a provider choice just to configure a pipeline. These variables override `.blastproof/config.yaml`, and precedence is **CLI flag > environment > file**:
|
|
244
|
+
|
|
245
|
+
| variable | overrides |
|
|
246
|
+
| --- | --- |
|
|
247
|
+
| `BLASTPROOF_BASE_URL` | `base_url` — the app under test |
|
|
248
|
+
| `BLASTPROOF_LLM_PROVIDER` | `anthropic` \| `openai` \| `ollama` |
|
|
249
|
+
| `BLASTPROOF_LLM_MODEL` | the model name |
|
|
250
|
+
| `BLASTPROOF_LLM_BASE_URL` | the provider endpoint — *not* the app |
|
|
251
|
+
| `BLASTPROOF_LLM_API_KEY_ENV` | the **name** of the variable holding your key |
|
|
252
|
+
|
|
253
|
+
Running the committed config against an OpenAI-compatible gateway, without editing a file:
|
|
254
|
+
|
|
255
|
+
```bash
|
|
256
|
+
export BLASTPROOF_LLM_PROVIDER=openai
|
|
257
|
+
export BLASTPROOF_LLM_MODEL=anthropic/claude-haiku-4.5
|
|
258
|
+
export BLASTPROOF_LLM_BASE_URL=https://openrouter.ai/api/v1
|
|
259
|
+
export BLASTPROOF_LLM_API_KEY_ENV=OPENROUTER_API_KEY
|
|
260
|
+
blastproof run --impacted --min-score 80
|
|
261
|
+
```
|
|
262
|
+
|
|
263
|
+
Note the last one names *which variable* holds your key — the key itself is never read from a `BLASTPROOF_*` variable, so error messages can keep naming the variable you chose. An empty value counts as unset, so `FOO=` in a CI matrix will not blank a configured setting.
|
|
264
|
+
|
|
265
|
+
## Roadmap
|
|
266
|
+
|
|
267
|
+
- [x] Repository & spec-driven development setup
|
|
268
|
+
- [x] **M1** — `init` + `run`: YAML test runner with agentic LLM executor
|
|
269
|
+
- [x] **M2** — diff analysis, impact mapping (`run --impacted`) and test generation (`plan`)
|
|
270
|
+
- [x] **M3** — Reports (JUnit + HTML), priority-weighted score, `--min-score` gate, `blastproof test`
|
|
271
|
+
- [ ] **M4** — GitHub Action, npm publish
|
|
272
|
+
- [ ] Post-MVP — VS Code extension, session replay, worker parallelism
|
|
273
|
+
|
|
274
|
+
## Development
|
|
275
|
+
|
|
276
|
+
This project uses **spec-driven development** via [OpenSpec](https://github.com/Fission-AI/OpenSpec). See [`AGENTS.md`](./AGENTS.md) for architecture, conventions and the contribution workflow — every change starts with an OpenSpec proposal.
|
|
277
|
+
|
|
278
|
+
```bash
|
|
279
|
+
# requires Node.js >= 20.19 (see engines in package.json)
|
|
280
|
+
npm install
|
|
281
|
+
npm run build
|
|
282
|
+
npm test
|
|
283
|
+
```
|
|
284
|
+
|
|
285
|
+
## License
|
|
286
|
+
|
|
287
|
+
[MIT](./LICENSE)
|