blastproof 0.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/LICENSE ADDED
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 blastproof contributors
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
package/README.md ADDED
@@ -0,0 +1,287 @@
1
+ # blastproof
2
+
3
+ [![CI](https://github.com/hamc/blastproof/actions/workflows/ci.yml/badge.svg)](https://github.com/hamc/blastproof/actions/workflows/ci.yml)
4
+ [![Dogfood](https://github.com/hamc/blastproof/actions/workflows/dogfood.yml/badge.svg)](https://github.com/hamc/blastproof/actions/workflows/dogfood.yml)
5
+ [![License: MIT](https://img.shields.io/badge/License-MIT-blue.svg)](./LICENSE)
6
+
7
+ **Open-source AI testing agent for pull requests.** Diff in, confidence out — no test scripts to write or maintain.
8
+
9
+ `blastproof` is an open-source AI QA agent for pull requests: it reads your PR diff, maps the blast radius, generates end-to-end tests in plain English, executes them on a real browser with self-healing, and scores the result before merge. 100% local, MIT licensed, bring your own LLM key.
10
+
11
+ ```
12
+ git diff → impact mapping → test generation → agentic execution → report + score
13
+ ```
14
+
15
+ ## How it works
16
+
17
+ 1. **Reads the diff** — `blastproof test --base main` parses the branch diff and maps it to affected routes.
18
+ 2. **Maps the blast radius** — traces changed files to the user journeys and routes most likely affected.
19
+ 3. **Generates tests** — writes/updates plain-English YAML tests in `.blastproof/tests/`.
20
+ 4. **Executes agentically** — an LLM-driven loop over Playwright resolves elements via the accessibility tree on every step. No static selectors, so no flakiness: the agent re-resolves when the UI shifts.
21
+ 5. **Reports & scores** — console, JUnit XML and HTML reports, plus a priority-weighted score that fails the run below `--min-score`.
22
+
23
+ ## Quick start
24
+
25
+ ```bash
26
+ npm install -g blastproof # requires Node.js >= 20.19
27
+ npx playwright install --with-deps chromium # one-time browser download
28
+ cd your-project
29
+ blastproof init
30
+ export ANTHROPIC_API_KEY=... # or OPENAI_API_KEY, or a local Ollama model
31
+ blastproof run # runs .blastproof/tests/**/*.yaml agentically
32
+ ```
33
+
34
+ Published with [npm provenance](https://docs.npmjs.com/generating-provenance-statements), so the tarball is verifiably built from this repository.
35
+
36
+ Try it locally without your own app — this repo ships a demo shop:
37
+
38
+ ```bash
39
+ node examples/demo-app/serve.mjs 4173 & # home, login and cart + promo-code pages
40
+ blastproof init
41
+ export ANTHROPIC_API_KEY=...
42
+ blastproof run
43
+ ```
44
+
45
+ > **Status:** the full pipeline works today — `init`, `run` (including `--impacted`), `plan` and `test`, with JUnit and HTML reports. The consumable GitHub Action is next.
46
+
47
+ ## Test format
48
+
49
+ Tests live in `.blastproof/tests/` as plain-English YAML — no selectors, no framework lock-in:
50
+
51
+ ```yaml
52
+ summary: Checkout with discount
53
+ priority: P0
54
+ tags: [checkout, discount]
55
+ routes: ["/cart", "/checkout"]
56
+ steps:
57
+ - add item to cart
58
+ - apply promo code SAVE20
59
+ - verify a 20% discount is applied
60
+ - complete checkout
61
+ ```
62
+
63
+ `priority` is P0–P2 (default P1); `tags` and `setup` steps are optional. `routes` (optional) declares the URLs a test covers: `blastproof run --impacted` runs only tests whose `routes:` intersect the routes affected by your PR diff (mapped from changed files via the `routes:` globs in `.blastproof/config.yaml`). Route strings compare by exact equality (`/cart` ≠ `/cart/` — write them consistently). Tests without `routes:` are skipped and reported under `--impacted`.
64
+
65
+ ## CLI
66
+
67
+ | Command | Description |
68
+ | --- | --- |
69
+ | `blastproof init` | Scaffold `.blastproof/` config and sample tests (idempotent) |
70
+ | `blastproof run [--tag smoke] [--priority P0] [--query checkout]` | Run tests only — exit 0 pass, 1 fail, 2 usage/config error |
71
+ | `blastproof run --impacted [--base <ref>]` | Run only tests impacted by the diff vs the base ref (default `main`). Unrouted tests are skipped and reported; affected-but-uncovered routes are reported without failing the run |
72
+ | `blastproof run --dry-run` | Print the selection plan (affected routes, unmapped files, selected/skipped tests) and exit 0 — no browser launched, no LLM key needed |
73
+ | `blastproof run --url <url>` | Override `base_url` for this run only (e.g. a PR preview environment); the config file is never mutated |
74
+ | `blastproof plan [--base <ref>]` | Generate plain-English tests for affected routes no test covers yet. Prints drafts; nothing is written without `--write` |
75
+ | `blastproof plan --route <route>` | Generate for a route explicitly, skipping the diff (repeatable) — how you bootstrap coverage on an app with no suite yet |
76
+ | `blastproof plan --write` | Persist drafts to `.blastproof/tests/<route-slug>.yaml`. Never overwrites: a colliding filename fails that route |
77
+ | `blastproof run --min-score <n>` | Require a weighted score of at least `n` (0–100). **Replaces** the all-must-pass rule — see below |
78
+ | `blastproof run --junit [path]` | Write a JUnit XML report; without a path it lands in `.blastproof/reports/<session>/junit.xml` |
79
+ | `blastproof run --html [path]` | Write a self-contained HTML report with failure screenshots embedded inline |
80
+ | `blastproof test [--base <ref>]` | The full pipeline: run the tests covering the diff, then draft tests for the gaps |
81
+
82
+ ### The full pipeline: `blastproof test`
83
+
84
+ One command for the whole loop — map the blast radius, run what covers it, draft what doesn't exist yet:
85
+
86
+ ```bash
87
+ blastproof test --base main --min-score 80 --junit junit.xml --html report.html
88
+ ```
89
+
90
+ It does two things and reports them separately:
91
+
92
+ 1. **Verify** — executes the tests covering the affected routes, scores them, applies the gate
93
+ 2. **Draft** — generates tests for affected routes no test covers, and prints them
94
+
95
+ **Generated drafts are never executed, and never affect the score.** That is deliberate. An unreviewed, model-written test in the merge path fails in two directions: a hallucinated expectation blocks a correct pull request, and a credulous one waves a broken change through while looking like coverage. Either costs more trust than the automation saves.
96
+
97
+ So read the result honestly: `test` does not make an uncovered route safe. It makes the gap visible, with a draft ready for you to review, and the score keeps describing only what was actually verified. Add `--write` to persist the drafts (never overwriting an existing file) and commit them once you have read them.
98
+
99
+ Exit codes: 2 usage/config/diff, 1 when the gate fails **or** a draft could not be generated, 0 otherwise.
100
+
101
+ ### Reports
102
+
103
+ `--junit` is for CI; `--html` is for humans. The HTML report is a single self-contained file — inline CSS, screenshots embedded as data URIs, no scripts — so it opens offline, survives being moved, and uploads as one artifact. It leads with the score and gate verdict, sorts failures above passes, and expands each failure to its failing step, reason and screenshot.
104
+
105
+ ### Generating tests with `plan`
106
+
107
+ `plan` closes the gap `run --impacted` reports. It takes the affected routes no test covers, loads each one in the browser, and asks the model to write a test from the page's real accessibility tree plus the changed files that made the route impacted — so the generated steps name controls that actually exist:
108
+
109
+ ```bash
110
+ blastproof plan --base main # preview drafts for uncovered routes
111
+ blastproof plan --base main --write # persist them, then review and commit
112
+ blastproof plan --route /checkout # bootstrap a route without a diff
113
+ ```
114
+
115
+ Drafts are **previews by default** — nothing touches disk until `--write`, and `--write` never overwrites an existing file, so a regeneration can't silently replace a test you edited by hand. Each written file carries a header recording its route, base ref and generation date. Review before committing: the steps are model-written and meant to be edited.
116
+
117
+ Exit codes: 0 when every route generated (or nothing needed coverage), 1 when a route failed, 2 on usage/config/diff errors. A route that fails to load never aborts the others.
118
+
119
+ **Known limitation:** a route behind authentication snapshots as the login wall, so its draft describes logging in rather than the feature. The `auth` config recipe is not applied by the planner yet — generate those routes after an auth session lands, or write them by hand.
120
+
121
+ ### Closing the coverage hole
122
+
123
+ Impact mapping has a failure mode worth understanding. A changed file that matches no `routes:` glob contributes no affected routes — so a diff touching only a shared module selects nothing, scores 100 because nothing executed, and merges green. The information is printed, but nobody reads a passing run.
124
+
125
+ Each changed file is classified three ways:
126
+
127
+ | a changed file | means |
128
+ | --- | --- |
129
+ | matches a `routes:` glob | contributes its routes |
130
+ | matches an `ignore:` glob | knowingly irrelevant to any page |
131
+ | matches neither | **nobody has said what this affects** |
132
+
133
+ ```yaml
134
+ routes:
135
+ "src/cart/**": ["/cart", "/checkout"]
136
+ ignore:
137
+ - "**/*.md"
138
+ - ".github/**"
139
+ ```
140
+
141
+ ```bash
142
+ blastproof run --impacted --fail-on-unmapped
143
+ ```
144
+
145
+ The flag fails the run on the third case only, naming the files and both ways to resolve them. `ignore:` is what makes that signal survivable — without it the flag would fire on every README edit and get switched off within a day, and a disabled gate protects nothing.
146
+
147
+ Nothing is ignored by default, on purpose: a file nobody has classified is exactly the risk the flag exists to surface, and a default that guesses on your behalf would hide the first files worth thinking about. This flag is **additive** — a run can meet its `--min-score` and still be blocked here, because "the tests I ran passed" and "something changed that nobody has classified" are different claims.
148
+
149
+ Be clear on its limit: it catches files that are *unclassified*, not files that are *misclassified*. A shared module mapped to one route when it can break five will still slip through. Mapping by import graph is the answer to that, and blastproof does not do it yet.
150
+
151
+ ### Score and merge gating
152
+
153
+ Every run ends with a score: the percentage of executed test **weight** that passed, where a test weighs 3 at P0, 2 at P1 and 1 at P2. Weighting is the point — a failing checkout costs three times a failing tooltip, so a pile of trivial passes can't hide a broken critical journey.
154
+
155
+ ```bash
156
+ blastproof run # any failure exits 1 (strict, the default)
157
+ blastproof run --min-score 80 # passes at 80+, so one failing P2 is tolerated
158
+ blastproof run --min-score 100 # identical to the default strict behaviour
159
+ ```
160
+
161
+ `--min-score` **replaces** the all-must-pass rule rather than adding to it. Without it, any failure exits 1. With it, the score alone decides — which is what lets you say "a P2 may break, a P0 may not" in one number. Only executed tests count: tests removed by `--tag`/`--priority`/`--query`, and tests skipped as unrouted under `--impacted`, are neither numerator nor denominator. A run that executed nothing scores 100 (the output says so explicitly), so a docs-only PR is never blocked.
162
+
163
+ For CI:
164
+
165
+ ```bash
166
+ blastproof run --impacted --base "$BASE_REF" --min-score 80 --junit junit.xml
167
+ ```
168
+
169
+ Exit 0 merge-able, 1 blocked, 2 usage/config error. The JUnit report carries the score as a `<property name="score">` so a parser can read it without scraping stdout, and tests skipped for having no `routes:` appear as `<skipped/>` cases — the coverage gap shows up in CI instead of vanishing.
170
+
171
+ ## blastproof tests itself
172
+
173
+ The **Dogfood** badge above is blastproof running against the demo app in this repo: real Chromium, real LLM, plain-English tests, scored and gated. The run logs are public — the agent's reasoning, step by step, is there to read.
174
+
175
+ It catches real regressions rather than diffing strings. Changing the demo app's discount from 20% to 5%, while leaving the on-screen message still claiming *"Promo code SAVE20 applied: 20% off"*, produces:
176
+
177
+ ```
178
+ FAIL P0 Promo code SAVE20 applies a 20% discount in the cart
179
+ failing step: verify a 20% discount of $24.00 is shown
180
+ reason: the discount is currently -$6.00, but a 20% discount on
181
+ $120.00 should be -$24.00
182
+ Score: 50 — min-score 80: FAIL (below threshold)
183
+ ```
184
+
185
+ No selector was updated and no assertion was rewritten to catch that. The agent read the rendered value, did the arithmetic, and disagreed with the page.
186
+
187
+ Two workflows, split by what they cost:
188
+
189
+ - **Impact** — runs on every pull request, including forks. Deterministic and keyless: it reports the blast radius of the diff and which tests cover it, before anyone spends a token.
190
+ - **Dogfood** — runs daily and on demand. The agentic run needs an API key, so it stays out of the merge path: a non-deterministic model answer should never block a merge.
191
+
192
+ ## Testing behind a login
193
+
194
+ Most of a product lives behind authentication. Declare a recipe once and blastproof signs in a single time per run, then reuses that session for every test **and** for `plan` — so generated drafts describe the actual feature instead of the login wall.
195
+
196
+ Pick exactly one strategy:
197
+
198
+ ```yaml
199
+ # 1) A plain-English login journey — form login, or anything a person can click through
200
+ auth:
201
+ steps:
202
+ - navigate to /login
203
+ - fill the email field with {{env.TEST_EMAIL}}
204
+ - fill the password field with {{env.TEST_PASSWORD}}
205
+ - submit the login form
206
+ verify: a signed-in indicator is visible # optional, strongly recommended
207
+
208
+ # 2) A session captured by hand — for SSO, MFA or magic links
209
+ auth:
210
+ storage_state: .blastproof/auth.json
211
+
212
+ # 3) Static values — for token-based apps
213
+ auth:
214
+ headers:
215
+ Authorization: "Bearer {{env.API_TOKEN}}"
216
+ ```
217
+
218
+ A test that exercises the login itself must start signed out:
219
+
220
+ ```yaml
221
+ summary: Login with valid credentials succeeds
222
+ auth: false
223
+ ```
224
+
225
+ **`verify` is worth the one extra call.** Without it, a wrong password surfaces as every test failing on a login wall — N failures, none naming the cause. With it, the run stops before the first test and says what happened. Authentication failure exits 2 and never reports as failing tests: a login you cannot complete says nothing about the code under review, so it must not produce a score.
226
+
227
+ Each test still gets its own browser context; it simply starts from the shared session rather than empty, so isolation is unchanged. Set `auth.cache: true` to reuse a session across runs — off by default, because an expired session produces failures at random points with nothing pointing at the cause.
228
+
229
+ **A captured session is a credential.** The file holds live cookies: whoever has it is signed in as that user. `init` git-ignores it; never commit one.
230
+
231
+ > **Note on self-healing:** the executor recovers from failed steps by re-reading the page, which means it can complete a login using credentials the page itself displays — some apps show demo credentials on the sign-in form. That is the self-healing loop working as designed, but it does mean a deliberately-wrong password is not a reliable way to test your auth failure path.
232
+
233
+ ## LLM providers (BYOK)
234
+
235
+ Bring your own key — runs 100% locally:
236
+
237
+ - **Anthropic** (`ANTHROPIC_API_KEY`)
238
+ - **OpenAI** (`OPENAI_API_KEY`)
239
+ - **Ollama** (local, no key needed)
240
+
241
+ ### Configuring from the environment
242
+
243
+ You never have to commit a provider choice just to configure a pipeline. These variables override `.blastproof/config.yaml`, and precedence is **CLI flag > environment > file**:
244
+
245
+ | variable | overrides |
246
+ | --- | --- |
247
+ | `BLASTPROOF_BASE_URL` | `base_url` — the app under test |
248
+ | `BLASTPROOF_LLM_PROVIDER` | `anthropic` \| `openai` \| `ollama` |
249
+ | `BLASTPROOF_LLM_MODEL` | the model name |
250
+ | `BLASTPROOF_LLM_BASE_URL` | the provider endpoint — *not* the app |
251
+ | `BLASTPROOF_LLM_API_KEY_ENV` | the **name** of the variable holding your key |
252
+
253
+ Running the committed config against an OpenAI-compatible gateway, without editing a file:
254
+
255
+ ```bash
256
+ export BLASTPROOF_LLM_PROVIDER=openai
257
+ export BLASTPROOF_LLM_MODEL=anthropic/claude-haiku-4.5
258
+ export BLASTPROOF_LLM_BASE_URL=https://openrouter.ai/api/v1
259
+ export BLASTPROOF_LLM_API_KEY_ENV=OPENROUTER_API_KEY
260
+ blastproof run --impacted --min-score 80
261
+ ```
262
+
263
+ Note the last one names *which variable* holds your key — the key itself is never read from a `BLASTPROOF_*` variable, so error messages can keep naming the variable you chose. An empty value counts as unset, so `FOO=` in a CI matrix will not blank a configured setting.
264
+
265
+ ## Roadmap
266
+
267
+ - [x] Repository & spec-driven development setup
268
+ - [x] **M1** — `init` + `run`: YAML test runner with agentic LLM executor
269
+ - [x] **M2** — diff analysis, impact mapping (`run --impacted`) and test generation (`plan`)
270
+ - [x] **M3** — Reports (JUnit + HTML), priority-weighted score, `--min-score` gate, `blastproof test`
271
+ - [ ] **M4** — GitHub Action, npm publish
272
+ - [ ] Post-MVP — VS Code extension, session replay, worker parallelism
273
+
274
+ ## Development
275
+
276
+ This project uses **spec-driven development** via [OpenSpec](https://github.com/Fission-AI/OpenSpec). See [`AGENTS.md`](./AGENTS.md) for architecture, conventions and the contribution workflow — every change starts with an OpenSpec proposal.
277
+
278
+ ```bash
279
+ # requires Node.js >= 20.19 (see engines in package.json)
280
+ npm install
281
+ npm run build
282
+ npm test
283
+ ```
284
+
285
+ ## License
286
+
287
+ [MIT](./LICENSE)