blastproof 0.2.1 → 0.3.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -15,11 +15,31 @@ git diff → impact mapping → test generation → agentic execution → report
15
15
  ## How it works
16
16
 
17
17
  1. **Reads the diff** — `blastproof test --base main` parses the branch diff and maps it to affected routes.
18
- 2. **Maps the blast radius** — traces changed files to the user journeys and routes most likely affected.
18
+ 2. **Maps the blast radius** — matches changed files against the `routes:` map you maintain, to select the journeys worth running. Deterministic and free; see [Closing the coverage hole](#closing-the-coverage-hole) for what that catches and what it does not.
19
19
  3. **Writes the missing tests** — drafts plain-English YAML for affected routes nothing covers. Drafts are yours to review; existing tests are never rewritten.
20
20
  4. **Executes agentically** — an LLM-driven loop over Playwright resolves elements via the accessibility tree on every step. No static selectors to rot: the agent re-resolves when the UI shifts. Being model-driven it is not deterministic, which is why the free, deterministic impact analysis is what runs on every pull request.
21
21
  5. **Reports & scores** — console, JUnit XML and HTML reports, plus a priority-weighted score that fails the run below `--min-score`.
22
22
 
23
+ ## Does this fit your application?
24
+
25
+ Read this before installing. The agent drives your app the way a screen reader would, and that has consequences worth knowing in advance rather than discovering on a failing run.
26
+
27
+ **Your markup has to be accessible. This is a hard requirement, not a nice-to-have.** Every element is found by role, label or visible text, resolved from the page's accessibility tree on each attempt — that is what removes selectors and what lets the agent survive a redesign. The cost is that an interface the accessibility tree cannot describe is one the agent cannot drive at all. Icon-only buttons with no accessible name, custom dropdowns and comboboxes without ARIA roles, and `div`-based controls do not degrade into something slower; they simply cannot be targeted, and there is deliberately no CSS or XPath escape hatch to fall back on. If you are unsure, run your app through any accessibility checker first — the result predicts how well blastproof will do better than anything else.
28
+
29
+ **Some interactions are not supported yet:**
30
+
31
+ | Not supported | Consequence |
32
+ |---|---|
33
+ | `iframe` content | Anything inside an embedded frame is invisible to the agent, including the page snapshot. **Hosted payment widgets such as Stripe Elements or PayPal buttons fall in here** — if your checkout embeds one, the checkout journey cannot be driven end to end today |
34
+ | Hover, scroll-to, drag and drop | Menus that open only on hover, and content that loads only once scrolled into view, are out of reach |
35
+ | File upload | No action exists for it |
36
+ | Multiple tabs or windows | The agent works in one page; a flow that opens a popup loses it |
37
+ | Native `alert` / `confirm` dialogs | Not handled |
38
+
39
+ **Large pages are truncated.** The accessibility snapshot sent to the model is capped at 200 lines. A dense admin table or a long product listing will have its lower content invisible on every step.
40
+
41
+ **The deterministic half stands on its own.** If none of the above fits, `blastproof run --impacted --dry-run` still works: it reads the diff, reports which routes a change touches, which files are classified by nothing, and which affected routes no test covers. That needs no API key, no browser and no model, costs nothing per pull request, and is useful purely as a coverage-gap report — see [Closing the coverage hole](#closing-the-coverage-hole).
42
+
23
43
  ## Quick start
24
44
 
25
45
  ```bash
@@ -82,6 +102,7 @@ steps:
82
102
  | `blastproof run --min-score <n>` | Require a weighted score of at least `n` (0–100). **Replaces** the all-must-pass rule — see below |
83
103
  | `blastproof run --junit [path]` | Write a JUnit XML report; without a path it lands in `.blastproof/reports/<session>/junit.xml` |
84
104
  | `blastproof run --html [path]` | Write a self-contained HTML report with failure screenshots embedded inline |
105
+ | `--max-llm-calls <n>` \| `--max-tokens <n>` \| `--max-duration <seconds>` | Bound `run`, `plan` or `test` by model calls, tokens or wall-clock time — every command that can call the model accepts these. Overrides `.blastproof/config.yaml`'s `budget:` section — see [Bounding a run](#bounding-a-run-budget-and-deadline) |
85
106
  | `blastproof test [--base <ref>]` | The full pipeline: run the tests covering the diff, then draft tests for the gaps |
86
107
 
87
108
  ### The full pipeline: `blastproof test`
@@ -119,9 +140,9 @@ blastproof plan --route /checkout # bootstrap a route without a diff
119
140
 
120
141
  Drafts are **previews by default** — nothing touches disk until `--write`, and `--write` never overwrites an existing file, so a regeneration can't silently replace a test you edited by hand. Each written file carries a header recording its route, base ref and generation date. Review before committing: the steps are model-written and meant to be edited.
121
142
 
122
- Exit codes: 0 when every route generated (or nothing needed coverage), 1 when a route failed, 2 on usage/config/diff errors. A route that fails to load never aborts the others.
143
+ Exit codes: 0 when every route generated (or nothing needed coverage), 1 when a route failed or the budget/deadline stopped generation early, 2 on usage/config/diff errors. A route that fails to load never aborts the others; a budget or deadline stop does — see [Bounding a run](#bounding-a-run-budget-and-deadline).
123
144
 
124
- **Known limitation:** a route behind authentication snapshots as the login wall, so its draft describes logging in rather than the feature. The `auth` config recipe is not applied by the planner yet — generate those routes after an auth session lands, or write them by hand.
145
+ `plan` uses the same `auth:` recipe as `run`, so a route behind a login is drafted from the real page rather than from the login wall.
125
146
 
126
147
  ### Closing the coverage hole
127
148
 
@@ -173,6 +194,48 @@ blastproof run --impacted --base "$BASE_REF" --min-score 80 --junit junit.xml
173
194
 
174
195
  Exit 0 merge-able, 1 blocked, 2 usage/config error. The JUnit report carries the score as a `<property name="score">` so a parser can read it without scraping stdout, and tests skipped for having no `routes:` appear as `<skipped/>` cases — the coverage gap shows up in CI instead of vanishing.
175
196
 
197
+ ## Bounding a run: budget and deadline
198
+
199
+ Nothing stops a run by default — a suite runs to completion or the provider refuses. `budget:` puts a ceiling on it, in `.blastproof/config.yaml`. It applies to `run`, `plan` and `test` alike: every model call any of them makes — agent action, assert judgment, or test planning — is counted, because a budget that only covered `run` would leave `plan`'s calls unbounded.
200
+
201
+ ```yaml
202
+ budget:
203
+ max_llm_calls: 500 # stop after this many model calls
204
+ max_tokens: 2000000 # stop after this many tokens spent across all calls
205
+ max_duration_s: 900 # stop after this many seconds of wall-clock time
206
+ ```
207
+
208
+ Each limit is independent and optional; a config with no `budget:` section — or a run with none of `--max-llm-calls` / `--max-tokens` / `--max-duration` — behaves exactly as before. All three are counted in **calls and tokens, not currency**: a price table keyed by model and provider goes stale the day a provider reprices, and a limit that silently stops meaning what it says is worse than no limit, because it is trusted. Calls and tokens are exact, already reported by every provider, and yours to convert to a dollar figure with your own rates if you want one.
209
+
210
+ `--dry-run` reports the ceiling before you spend anything — the worst case a selection could cost, computed from step counts alone, no provider contacted:
211
+
212
+ ```
213
+ Dry run: 12 test(s) selected, base_url=http://localhost:4173
214
+ Worst case: up to 216 model call(s) for this selection (a maximum, not a prediction).
215
+ ```
216
+
217
+ That number is a ceiling, not a forecast — a real run almost always finishes in a fraction of it, because most steps complete long before either cap runs out. Per step it is the iteration cap **plus** `max_retries_per_step` (read from your config, not assumed): a malformed model response is retried without spending an iteration, and a failing `assert` spends a retry *and* an iteration in the same call — so the two caps are added, not one doubled and the other ignored. If `auth.steps` is configured, the login journey's steps are counted too — it runs once before any test and spends model calls through the same loop, so a ceiling that excluded it could be exceeded by the very first run that logs in.
218
+
219
+ **Exhausting a budget stops the run — it does not fail a test.** Running out of quota says nothing about the application under test, so recording it as a failure would manufacture a defect that does not exist. Tests the run never reached are reported as **not run**, a third state distinct from passed and failed, and excluded from the score's denominator entirely — counting them as failures would just be a quieter version of the same lie a false pass would have been.
220
+
221
+ An interrupted run is unmistakably incomplete: the process **exits 1 unconditionally**, even when `--min-score` is given and the tests that did execute would have satisfied it. The tests that finished are whichever ones happened to run first, not a representative sample, so nothing about them is a verdict:
222
+
223
+ ```
224
+ Run incomplete: model call budget exhausted: reached the configured maximum of 500 call(s)
225
+ Score over executed tests: 92 (not a verdict — exit code 1 regardless of --min-score)
226
+ ```
227
+
228
+ Both the JUnit and HTML reports carry the same signal: unexecuted tests appear as `<skipped/>` cases naming the limit, distinct from `<failure>` cases, and the HTML report leads with a banner stating the run was stopped and why.
229
+
230
+ Like the LLM provider settings, every field overrides from the environment (`BLASTPROOF_MAX_LLM_CALLS`, `BLASTPROOF_MAX_TOKENS`, `BLASTPROOF_MAX_DURATION_S`), and a CLI flag beats both:
231
+
232
+ ```bash
233
+ blastproof run --impacted --max-llm-calls 200 --max-duration 300
234
+ blastproof plan --max-llm-calls 200
235
+ ```
236
+
237
+ `blastproof test` composes `run` then `plan` in one process. It resolves the budget once and hands the same instance to both phases, so the pipeline stays within the configured maximum overall — not up to double it, which is what each phase resolving its own budget would silently allow.
238
+
176
239
  ## blastproof tests itself
177
240
 
178
241
  The **Dogfood** badge above is blastproof running against the demo app in this repo: real Chromium, real LLM, plain-English tests, scored and gated. The run logs are public — the agent's reasoning, step by step, is there to read.
@@ -210,7 +273,7 @@ jobs:
210
273
 
211
274
  - run: npm start & # however your app boots
212
275
 
213
- - uses: hamc/blastproof@v0.2.1
276
+ - uses: hamc/blastproof@v0.2.2
214
277
  with:
215
278
  api-key: ${{ secrets.ANTHROPIC_API_KEY }}
216
279
  base: ${{ github.event.pull_request.base.ref }}
@@ -222,9 +285,9 @@ Exit non-zero blocks the merge. Use the score in a later step:
222
285
 
223
286
  ```yaml
224
287
  - id: bp
225
- uses: hamc/blastproof@v0.2.1
288
+ uses: hamc/blastproof@v0.2.2
226
289
  with:
227
- version: '0.2.1' # pin both when the result gates merges
290
+ version: '0.2.2' # pin both when the result gates merges
228
291
  api-key: ${{ secrets.ANTHROPIC_API_KEY }}
229
292
  base: ${{ github.event.pull_request.base.ref }}
230
293
  min-score: '80'
@@ -260,7 +323,9 @@ Output: **`score`** — 0–100, empty when no report was produced, so "no score
260
323
 
261
324
  **`fetch-depth: 0` is not optional** for `test` and `plan`. The default checkout is shallow and has no merge-base; the action detects this and fails immediately rather than letting it surface as a git error mid-run.
262
325
 
263
- ## What the agent can and cannot do
326
+ ## Trust boundaries
327
+
328
+ For what the agent can and cannot *drive*, see [Does this fit your application?](#does-this-fit-your-application). This section is about what it is allowed to *reach*.
264
329
 
265
330
  The application under test is not trusted input. Its page content reaches the model — that is how the agent knows what is on screen — so a page able to influence its own accessible text can try to influence the agent. Two things constrain that:
266
331
 
@@ -273,7 +338,9 @@ allowed_origins:
273
338
 
274
339
  This is enforced by comparison, not by asking the model nicely, so it holds regardless of what the page says.
275
340
 
276
- **Your secrets never reach the model.** `{{env.*}}` placeholders stay intact all the way through the prompt and are substituted at the moment of typing. The model is told to pass them through unchanged. This matters because blastproof encourages pointing `llm.base_url` at a gateway you do not run — the credential now stays on your machine either way.
341
+ **Your secrets stay out of the model's prompts.** `{{env.*}}` placeholders stay intact all the way through the prompt and are substituted at the moment of typing. Every value any test or the auth recipe references is also redacted from anything else crossing into a prompt — page snapshots included, since your app may render the credential itself — in both literal and percent-encoded form. This matters because blastproof encourages pointing `llm.base_url` at a gateway you do not run.
342
+
343
+ Redaction matches known values, so it cannot anticipate every way a page might transform one before rendering it. Treat it as a strong default, not a guarantee against a hostile application.
277
344
 
278
345
  The system prompt also tells the model that page content is data under test and never an instruction to obey. That raises the cost of a casual injection and is **not** a security boundary — a determined one will get past prompt wording. The origin constraint is the boundary; treat the rest as hygiene, and do not point blastproof at an application you would not run locally.
279
346
 
@@ -337,6 +404,9 @@ You never have to commit a provider choice just to configure a pipeline. These v
337
404
  | `BLASTPROOF_LLM_MODEL` | the model name |
338
405
  | `BLASTPROOF_LLM_BASE_URL` | the provider endpoint — *not* the app |
339
406
  | `BLASTPROOF_LLM_API_KEY_ENV` | the **name** of the variable holding your key |
407
+ | `BLASTPROOF_MAX_LLM_CALLS` | `budget.max_llm_calls` — see [Bounding a run](#bounding-a-run-budget-and-deadline) |
408
+ | `BLASTPROOF_MAX_TOKENS` | `budget.max_tokens` |
409
+ | `BLASTPROOF_MAX_DURATION_S` | `budget.max_duration_s`, in seconds |
340
410
 
341
411
  Running the committed config against an OpenAI-compatible gateway, without editing a file:
342
412