blastproof 0.10.0 → 0.12.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -1,4 +1,10 @@
1
- # blastproof
1
+ <h1>
2
+ <picture>
3
+ <source media="(prefers-color-scheme: dark)" srcset="./.github/logo-dark.svg">
4
+ <img src="./.github/logo-light.svg" alt="" width="30" height="30">
5
+ </picture>
6
+ blastproof
7
+ </h1>
2
8
 
3
9
  [![CI](https://github.com/hamc/blastproof/actions/workflows/ci.yml/badge.svg)](https://github.com/hamc/blastproof/actions/workflows/ci.yml)
4
10
  [![Dogfood](https://github.com/hamc/blastproof/actions/workflows/dogfood.yml/badge.svg)](https://github.com/hamc/blastproof/actions/workflows/dogfood.yml)
@@ -12,6 +18,10 @@ git diff → impact mapping → test generation → agentic execution → report
12
18
 
13
19
  100% local. MIT. Bring your own LLM key.
14
20
 
21
+ ▶ **[Watch the introduction](https://www.youtube.com/shorts/miqN5FzMF_k)** — what it does, in a minute.
22
+
23
+ **Documentation:** [Configuration](./docs/configuration.md) · [Testing behind a login](./docs/auth.md) · [Running in CI](./docs/ci.md) · [Contributing](./CONTRIBUTING.md) · [Architecture](./AGENTS.md)
24
+
15
25
  ## Quick start
16
26
 
17
27
  ```bash
@@ -38,21 +48,53 @@ blastproof run
38
48
 
39
49
  Before `run`, `plan` or `test` do anything, they check what they are about to spend — the browser can launch, the model provider is reachable, `base_url` responds — and report every unmet one together, so a stopped app or a missing browser is never a wall you hit one crash at a time. A missing system library names the exact install command and says it needs root; nothing is installed on your behalf. Silent when everything is fine, and skipped entirely by `--dry-run`, which needs none of it.
40
50
 
51
+ Provider options, budgets and browser tuning: [Configuration](./docs/configuration.md).
52
+
41
53
  ## Does this fit your application?
42
54
 
43
- **Your markup must be accessible — a hard requirement.** Elements are found by role, label or visible text from the accessibility tree. That is what removes selectors and survives redesigns; the cost is that an interface the accessibility tree cannot describe cannot be driven at all, and there is deliberately no CSS or XPath fallback. Icon-only buttons without accessible names, `div`-based controls and ARIA-less dropdowns simply cannot be targeted. Run an accessibility checker first — the result predicts how well this will work better than anything else.
55
+ Three questions. The first one decides most cases.
56
+
57
+ ### 1. Is your markup accessible?
58
+
59
+ **A hard requirement, not a preference.** blastproof finds elements the way a screen reader does — by role, by label, by visible text. That is what removes selectors and survives redesigns. The cost is that there is deliberately no CSS or XPath fallback, so anything the accessibility tree cannot describe cannot be driven at all.
60
+
61
+ | works | cannot be driven |
62
+ | --- | --- |
63
+ | `<button>Add to cart</button>` | a `<div>` with a click handler |
64
+ | `<label for="email">` + `<input>` | an input with no label |
65
+ | `<button aria-label="Delete note">` | an icon-only button with no name |
66
+ | `<select>` with `<option>`s | an ARIA-less custom dropdown |
67
+
68
+ **Run an accessibility checker on your app before installing anything.** The result predicts how well this will work better than anything else you could measure — and the fixes it suggests are worth making regardless of whether you adopt this tool.
69
+
70
+ ### 2. Does your journey need anything on this list?
44
71
 
45
- **Windows is untested.** Development and CI run on Linux and macOS. Nothing is known to be broken and reports are welcome (#8).
72
+ Not supported yet:
46
73
 
47
- **Not supported yet:** `iframe` content (so hosted payment widgets like Stripe Elements are invisible — an embedded checkout cannot be driven end to end), hover, scroll-to, drag and drop, file upload, multiple tabs, native `alert`/`confirm` dialogs. Page snapshots are capped at 200 lines by default, so very dense pages are truncated — raise it with `browser.max_snapshot_lines` if your pages need more; truncation is always marked in the snapshot so the model is never misled into thinking it saw the whole page.
74
+ - **`iframe` content** — a hosted payment widget is invisible, so an embedded checkout cannot be driven end to end
75
+ - **hover, scroll-to, drag and drop, file upload**
76
+ - **multiple tabs**, and native `alert` / `confirm` dialogs
48
77
 
49
- **Point it at disposable data.** Within a step, an action that commits — a click, or pressing Enter — is never performed twice: the runner refuses the repeat and tells the agent it already did that. This closes the case that used to produce duplicate records, where a submit answered by a redirect came back to a reset form and the agent, seeing no evidence of its own work, submitted again. It is not a guarantee of zero duplicate writes: an agent that reaches the same effect by a genuinely different route — another control with the same effect — is not caught. Use a seeded database, a staging environment you can reset, or a throwaway account; do not gate on a run against production data.
78
+ **Windows is untested.** Development and CI run on Linux and macOS. Nothing is known to be broken and reports are welcome ([#8](https://github.com/hamc/blastproof/issues/8)).
50
79
 
51
- `browser.timeout_ms` bounds every wait — resolving a target element from the accessibility tree, and navigation — not only the click or fill performed afterwards. Raise it for an application that is merely slow to hydrate; the trade-off is that a genuinely missing element then takes longer to fail. It never changes how many self-healing retries a step gets — waiting and retrying are deliberately separate.
80
+ If a critical journey needs one of these, that journey stays with your existing test suite. The two can coexist — nothing here replaces what you already have.
52
81
 
53
- Writing each step so it states its own outcome helps here as well as everywhere else: `submit the form, then verify the confirmation shows the reference number` gives the agent something to check, where `click the submit button` leaves it to invent an expectation — and a poor invented expectation is what turns one submission into three.
82
+ ### 3. Can you point it at data you can throw away?
54
83
 
55
- **This is now load-bearing, not just advisable.** The judge decides whether a step's own outcome holds, using the model's expectation only as the claim offered in support of it — a step that never says what its outcome is gives the judge nothing to anchor on beyond whatever the model happened to check that turn. `verify the confirmation shows the reference number` gives the judge a real question; `verify it worked` does not, and may now fail where a looser judge previously let a true-but-unrelated claim pass it.
84
+ **Use a seeded database, a staging environment you can reset, or a throwaway account. Do not gate on a run against production data.**
85
+
86
+ Within a step, an action that commits — a click, or pressing Enter — is never performed twice: the runner refuses the repeat and tells the agent it already did that. This closes the case that used to produce duplicate records, where a submit answered by a redirect came back to a reset form and the agent, seeing no evidence of its own work, submitted again.
87
+
88
+ It is **not** a guarantee of zero duplicate writes. An agent that reaches the same effect by a genuinely different route — another control that does the same thing — is not caught.
89
+
90
+ ## How it works
91
+
92
+ <img src="./.github/coverage-flow.svg" width="100%"
93
+ alt="How a diff becomes a merge decision. In CI, unattended: changed files are matched against the routes: and ignore: globs; matched files contribute affected routes, files matching neither are reported as unclassified and fail the run only under --fail-on-unmapped. Tests declaring an affected route are executed and produce a weighted score, which --min-score gates on. An affected route no test declares is reported as a coverage gap and never fails the run. Separately and manually, outside CI: blastproof plan loads such a route in Chromium, makes one model call, and produces a YAML draft you review, edit and run before committing it.">
94
+
95
+ **The boundary in the middle is the point.** Everything above it runs unattended on every pull request and ends in an exit code. Everything below it is something you choose to run, on your machine, and review before it lands.
96
+
97
+ A route no test covers is *reported*, never failed — blocking on it would punish you for an incomplete map instead of teaching you to complete it. Turning that report into a test is the manual half, and the draft it produces is not trusted until a person has read it.
56
98
 
57
99
  ## Writing tests
58
100
 
@@ -70,7 +112,7 @@ steps:
70
112
  - complete checkout
71
113
  ```
72
114
 
73
- `priority` is P0–P2 (default P1). `tags`, `setup` steps and `auth` are optional — `auth: false` runs the test signed out, which a login test needs. `routes` declares the URLs a test covers, which is what `--impacted` selects on; write route strings consistently, since they compare by exact equality (`/cart` ≠ `/cart/`).
115
+ `priority` is P0–P2 (default P1). `tags`, `setup` steps and `auth` are optional — `auth: false` runs the test signed out, which a login test needs. `routes` declares the URLs a test covers, which is what `--impacted` selects on; write route strings consistently, since they compare by exact equality (`/cart` ≠ `/cart/`). `run` warns to stderr — non-fatal — when a test declares a route no `routes:` mapping declares, since that route contributes nothing to `--impacted` selection.
74
116
 
75
117
  ### Say what each step should produce
76
118
 
@@ -92,6 +134,28 @@ Write steps that end in an observable result — text on the page, a count, a st
92
134
 
93
135
  **Inline error messages should be plain visible text.** `role="alert"` is read correctly from the accessibility tree and needs no special handling, but note that an alert your page has cleared shows up as an empty element: if a verdict says an alert exists whose content is missing, the message was emptied, not hidden.
94
136
 
137
+ ### A step that enters a value writes the value
138
+
139
+ ```yaml
140
+ - fill the note field with Order not received # runs
141
+ - fill the note field # cannot run
142
+ ```
143
+
144
+ The agent is **forbidden from inventing values** — one it types must come from the step, from the page, or from an `{{env.*}}` placeholder. So the second step is not merely vague, it is impossible, and the run discovers that a minute in, with a failure reason about page state that reads as though your application is broken.
145
+
146
+ `run` warns about it first, on every path, before launching a browser or asking for a key:
147
+
148
+ ```
149
+ Authoring (a step enters a value but names none):
150
+ Add a note (.blastproof/tests/notes.yaml) step 2:
151
+ fill the note field
152
+ → fill the note field with <value>
153
+ ```
154
+
155
+ Non-fatal by default — `--fail-on-authoring` turns it into exit 1 for teams enforcing it in CI. Taking the value from the page is fine and is not flagged: `fill the recipient field with the address shown on the confirmation page`.
156
+
157
+ **The check reads English only.** A suite written in another language runs exactly as well but is not inspected, and prints no warning saying so — silence from this check means "nothing found in English", never "this suite is clean".
158
+
95
159
  ## Commands
96
160
 
97
161
  ```bash
@@ -111,43 +175,16 @@ Common flags — `blastproof <command> --help` has the full list:
111
175
  | `--url <url>` | Override `base_url` for this run (e.g. a PR preview) |
112
176
  | `--min-score <n>` | Gate on a weighted score instead of all-must-pass |
113
177
  | `--fail-on-unmapped` | Fail when a changed file matches no `routes:` or `ignore:` glob |
178
+ | `--fail-on-authoring` | Fail when a step enters a value but names none (warns by default) |
114
179
  | `--junit [path]` · `--html [path]` | Write reports |
180
+ | `--concurrency <n>` | Run tests at once — [when that is safe](./docs/configuration.md#concurrency--running-tests-at-once) |
115
181
  | `--write` | `plan` only — persist drafts instead of previewing |
116
- | `--max-llm-calls` · `--max-tokens` · `--max-duration` | Bound what a run may spend |
182
+ | `--max-llm-calls` · `--max-tokens` · `--max-duration` | [Bound what a run may spend](./docs/configuration.md#budget--bounding-what-a-run-spends) |
117
183
 
118
184
  Exit codes: **0** pass, **1** the gate failed, **2** usage or config error.
119
185
 
120
186
  **Generated drafts are never executed and never affect the score.** An unreviewed model-written test in the merge path fails in two directions: a hallucinated expectation blocks a correct PR, and a credulous one waves a broken change through while looking like coverage. `plan` makes the gap visible with a draft to review; it does not make an uncovered route safe.
121
187
 
122
- ## In CI
123
-
124
- ```yaml
125
- name: blastproof
126
- on: pull_request
127
-
128
- jobs:
129
- test:
130
- runs-on: ubuntu-latest
131
- steps:
132
- - uses: actions/checkout@v4
133
- with:
134
- fetch-depth: 0 # required — the diff needs a merge-base
135
-
136
- - run: npm start & # however your app boots
137
-
138
- - uses: hamc/blastproof@v0.10.0
139
- with:
140
- version: '0.6.0' # pin both when this gates merges
141
- api-key: ${{ secrets.ANTHROPIC_API_KEY }}
142
- base: ${{ github.event.pull_request.base.ref }}
143
- min-score: '80'
144
- fail-on-unmapped: 'true'
145
- ```
146
-
147
- A non-zero exit blocks the merge. The action outputs `score` (0–100, empty when no report was produced) for later steps. Full input list: [`action.yml`](./action.yml).
148
-
149
- `fetch-depth: 0` is not optional — the default checkout is shallow and has no merge-base. The action detects this and fails immediately rather than letting it surface as a git error mid-run.
150
-
151
188
  ## Impact mapping
152
189
 
153
190
  `--impacted` runs only the tests whose `routes:` intersect the routes your diff can affect, mapped from changed files by globs you maintain in `.blastproof/config.yaml`:
@@ -176,6 +213,8 @@ blastproof run --min-score 80 # one failing P2 is tolerated
176
213
 
177
214
  `--min-score` **replaces** the all-must-pass rule rather than adding to it. Only executed tests count: filtered and unrouted tests are neither numerator nor denominator, and a run that executed nothing scores 100, so a docs-only PR is never blocked. JUnit carries the score as a `<property name="score">`, and unrouted tests appear as `<skipped/>` so the coverage gap shows up in CI rather than vanishing.
178
215
 
216
+ Wiring this into a pipeline, with the gating patterns worth knowing: [Running in CI](./docs/ci.md).
217
+
179
218
  ## Without a browser or a key
180
219
 
181
220
  Half of blastproof is deterministic and free. These need no model, no browser and no network:
@@ -189,103 +228,6 @@ blastproof plan --base main --dry-run # affected routes no test
189
228
 
190
229
  They report affected routes, files nobody has classified, and affected routes no test covers — a coverage-gap report with an exit code, useful even on a repo whose suite is Playwright or Cypress.
191
230
 
192
- ## Running tests at once
193
-
194
- Tests run one at a time by default. Raise it when your tests can stand it:
195
-
196
- ```yaml
197
- concurrency: 4
198
- ```
199
-
200
- or `blastproof run --concurrency 4` for a single invocation. On this repository's own suite that takes a run from 156s to 68s — **2.3× faster, for the same 81 model calls.** Parallelism buys wall-clock, not spend.
201
-
202
- **The default is 1 on purpose, and raising it is your call to make.** Other test runners default to parallel because their tests are isolated by construction — separate processes, separate fixtures. These are journeys driven against **one running application**, so two tests can see each other's data. A suite is safe to parallelise when its tests do not write state that another test reads.
203
-
204
- The test in this repository's own suite that could not run beside itself is a good shape to recognise: it adds a note and then asserts *"one note on file"*. It writes shared server state, and it asserts on a global count. Either alone is a warning; together they mean the test's verdict depends on nothing else touching the application at that moment.
205
-
206
- Two practical notes. Four concurrent journeys are four times the traffic against whatever you pointed at — usually fine for a development instance, worth knowing for a shared one. And with several model calls in flight, a `budget:` limit can overshoot by up to the concurrency rather than by a single call, since the calls already sent are allowed to finish.
207
-
208
- ## Bounding a run
209
-
210
- Nothing stops a run by default. `budget:` puts a ceiling on `run`, `plan` and `test` alike — every model call any of them makes is counted:
211
-
212
- ```yaml
213
- budget:
214
- max_llm_calls: 500
215
- max_tokens: 2000000
216
- max_duration_s: 900
217
- ```
218
-
219
- Each limit is optional; with none set, nothing binds. They count **calls and tokens, not currency** — a price table keyed by model and provider goes stale the day a provider reprices, and a limit that quietly stops meaning what it says is worse than none, because it is trusted.
220
-
221
- Exhausting a budget **stops the run; it does not fail a test.** Running out of quota says nothing about the code under review. Unreached tests are reported as `not run`, a third state excluded from the score entirely, and the process exits 1 unconditionally — `--min-score` cannot rescue it, because the tests that finished are whichever ran first, not a representative sample.
222
-
223
- **Every run reports what it spent**, so you can size a limit from experience rather than guesswork:
224
-
225
- ```
226
- Spent: 82 model call(s), 115407 token(s)
227
- Score: 100
228
- ```
229
-
230
- The figures also land in the JUnit report as `llm_calls` and `llm_tokens`, beside `score`, so a pipeline can trend cost without scraping output. A run stopped by its own budget reports the spend too — that is the case where the number is least guessable. Where a provider reports no token usage, the line says so rather than showing zero.
231
-
232
- `--dry-run` reports the ceiling before you spend anything. Read it as a maximum and nothing more: for this repository's own suite it says 735 calls where a real run spends 82. Size a budget from what your runs actually report, not from the ceiling.
233
-
234
- For an order of magnitude, this repository's own suite — 7 tests, 31 steps, an authenticated demo shop, `anthropic/claude-haiku-4.5` — spends **about 82 model calls and 115k tokens**, taking 156s serially or 68s at `--concurrency 4`. That is a number you can reproduce (`node examples/demo-app/serve.mjs 4173` and `blastproof run`), not a forecast for your suite: cost scales with steps, page density and how often the agent has to retry. Run yours once and read the `Spent:` line.
235
-
236
- ## Testing behind a login
237
-
238
- Declare a recipe once; blastproof signs in one time per run and reuses that session for every test and for `plan`. Pick exactly one strategy:
239
-
240
- ```yaml
241
- # 1) A plain-English journey — form login, or anything a person can click through
242
- auth:
243
- steps:
244
- - navigate to /login
245
- - fill the email field with {{env.TEST_EMAIL}}
246
- - fill the password field with {{env.TEST_PASSWORD}}
247
- - submit the login form
248
- verify: a signed-in indicator is visible # optional, strongly recommended
249
-
250
- # 2) A session captured by hand — for SSO, MFA or magic links
251
- auth:
252
- storage_state: .blastproof/auth.json
253
-
254
- # 3) Static values — for token-based apps
255
- auth:
256
- headers:
257
- Authorization: "Bearer {{env.API_TOKEN}}"
258
- ```
259
-
260
- **`verify` is worth the extra call.** Without it a wrong password surfaces as every test failing on a login wall — N failures, none naming the cause. Authentication failure exits 2 and never reports as failing tests, because a login you cannot complete says nothing about the code under review.
261
-
262
- **A captured session is a credential** — the file holds live cookies. `init` git-ignores it; never commit one.
263
-
264
- ## LLM providers (BYOK)
265
-
266
- **Anthropic** (`ANTHROPIC_API_KEY`), **OpenAI** (`OPENAI_API_KEY`), or **Ollama** (local, no key). Any setting can be overridden from the environment, with precedence **CLI flag > environment > file**:
267
-
268
- | variable | overrides |
269
- | --- | --- |
270
- | `BLASTPROOF_BASE_URL` | `base_url` — the app under test |
271
- | `BLASTPROOF_LLM_PROVIDER` | `anthropic` \| `openai` \| `ollama` |
272
- | `BLASTPROOF_LLM_MODEL` | the model name |
273
- | `BLASTPROOF_LLM_BASE_URL` | the provider endpoint — *not* the app |
274
- | `BLASTPROOF_LLM_API_KEY_ENV` | the **name** of the variable holding your key |
275
- | `BLASTPROOF_MAX_LLM_CALLS` · `BLASTPROOF_MAX_TOKENS` · `BLASTPROOF_MAX_DURATION_S` | the budget fields |
276
-
277
- So an OpenAI-compatible gateway needs no file edit:
278
-
279
- ```bash
280
- export BLASTPROOF_LLM_PROVIDER=openai
281
- export BLASTPROOF_LLM_MODEL=anthropic/claude-haiku-4.5
282
- export BLASTPROOF_LLM_BASE_URL=https://openrouter.ai/api/v1
283
- export BLASTPROOF_LLM_API_KEY_ENV=OPENROUTER_API_KEY
284
- blastproof run --impacted --min-score 80
285
- ```
286
-
287
- The key itself is never read from a `BLASTPROOF_*` variable — you name *which* variable holds it, so errors can keep naming the one you chose.
288
-
289
231
  ## Trust boundaries
290
232
 
291
233
  The application under test is not trusted input: its page content reaches the model, so a page that controls its own accessible text can try to influence the agent. Two things constrain that.
package/dist/cli.js CHANGED
@@ -7,6 +7,20 @@ import { Command, InvalidArgumentError } from "commander";
7
7
  import { mkdir, writeFile } from "fs/promises";
8
8
  import path from "path";
9
9
  import { access } from "fs/promises";
10
+
11
+ // src/report/errors.ts
12
+ var ReportError = class extends Error {
13
+ constructor(message) {
14
+ super(message);
15
+ this.name = "ReportError";
16
+ }
17
+ };
18
+ function fsReason(error) {
19
+ if (!(error instanceof Error)) return String(error);
20
+ return error.message.replace(/^E[A-Z]+:\s*/, "").replace(/,\s+\w+\s+'[^']*'$/, "").trim();
21
+ }
22
+
23
+ // src/commands/init.ts
10
24
  var DEFAULT_CONFIG = `# blastproof configuration
11
25
  # Docs: https://github.com/hamc/blastproof
12
26
 
@@ -124,14 +138,26 @@ async function exists(file) {
124
138
  return false;
125
139
  }
126
140
  }
141
+ var InitError = class extends Error {
142
+ constructor(message) {
143
+ super(message);
144
+ this.name = "InitError";
145
+ }
146
+ };
127
147
  async function writeIfAbsent(file, content, result) {
128
148
  if (await exists(file)) {
129
149
  result.kept.push(file);
130
150
  return;
131
151
  }
132
- await mkdir(path.dirname(file), { recursive: true });
133
- await writeFile(file, content);
134
- result.created.push(file);
152
+ try {
153
+ await mkdir(path.dirname(file), { recursive: true });
154
+ await writeFile(file, content);
155
+ result.created.push(file);
156
+ } catch (error) {
157
+ throw new InitError(
158
+ `Cannot scaffold ${file}: ${fsReason(error)}. Check that ${path.dirname(file)} is a directory you can write to, not a file.`
159
+ );
160
+ }
135
161
  }
136
162
  async function initProject(cwd = process.cwd()) {
137
163
  const result = { created: [], kept: [] };
@@ -1248,11 +1274,13 @@ function plannerSystemPrompt() {
1248
1274
  You receive a YAML accessibility snapshot of the page (roles and accessible names, exactly what a user perceives) and the list of source files a pull request changed in the area this page covers.
1249
1275
 
1250
1276
  Rules:
1251
- - Write steps a human tester could follow without looking at the code. One action or check per step.
1277
+ - Write steps a human tester could follow without looking at the code. One move per step \u2014 a single action together with what it should produce, or a single check. Never two unrelated actions in one step.
1252
1278
  - Refer to controls by the accessible name shown in the snapshot, spelled exactly. Never invent buttons, fields or links that are not in the snapshot.
1253
1279
  - Never write CSS selectors, XPath, IDs or any code \u2014 the runner resolves elements live from the accessibility tree.
1254
1280
  - Prefer the journey the changed files touch over a generic tour of the page. The changed files tell you which part of the page matters.
1255
- - End with at least one step that verifies an observable outcome (visible text, a count, a state change).
1281
+ - **The test starts at the application's base URL, not at this route.** Begin with a step that navigates to the route and says what should be visible once it loads \u2014 "navigate to /support and verify the heading "Contact support" is shown". Without it the run opens the home page and every later step looks for controls that are not there.
1282
+ - **Every step says what it should produce.** Name what must be true once the step has been carried out, not the action alone: "submit the support form and verify the confirmation page shows the ticket number", never "submit the support form". A step that names an action without an outcome asks the runner to judge whether something happened while looking at the page that succeeding produces \u2014 a submitted form comes back empty, a redirect moves the URL \u2014 and that is the shape behind several real failures.
1283
+ - **A step that enters a value writes the value.** "fill the subject field with Order not received", never "enter a subject". The runner is forbidden from inventing values, so a step that does not supply one cannot be carried out.
1256
1284
  - If a step needs a credential or any secret, write it as a placeholder like {{env.TEST_PASSWORD}}. Never write a real or invented password, token or key.
1257
1285
  - Keep the whole test to a handful of steps: one journey, not an exhaustive suite.`;
1258
1286
  }
@@ -1445,7 +1473,7 @@ var TestFileError = class extends Error {
1445
1473
  this.name = "TestFileError";
1446
1474
  }
1447
1475
  };
1448
- function formatIssues2(file, error) {
1476
+ function formatIssues2(error) {
1449
1477
  return error.issues.map((issue) => ` - ${issue.path.join(".") || "(root)"}: ${issue.message}`).join("\n");
1450
1478
  }
1451
1479
  async function parseTestFile(filePath) {
@@ -1466,7 +1494,7 @@ async function parseTestFile(filePath) {
1466
1494
  const result = testFileSchema.safeParse(data ?? {});
1467
1495
  if (!result.success) {
1468
1496
  throw new TestFileError(`Invalid test file ${filePath}:
1469
- ${formatIssues2(filePath, result.error)}`);
1497
+ ${formatIssues2(result.error)}`);
1470
1498
  }
1471
1499
  return { path: filePath, ...result.data };
1472
1500
  }
@@ -1880,9 +1908,15 @@ ${detail}
1880
1908
  `;
1881
1909
  }
1882
1910
  async function writeHtml(file, html) {
1883
- await mkdir4(path7.dirname(file), { recursive: true });
1884
- await writeFile4(file, html, "utf8");
1885
- return file;
1911
+ try {
1912
+ await mkdir4(path7.dirname(file), { recursive: true });
1913
+ await writeFile4(file, html, "utf8");
1914
+ return file;
1915
+ } catch (error) {
1916
+ throw new ReportError(
1917
+ `Cannot write HTML report to ${file}: ${fsReason(error)}. Check that the path is writable and is not a directory.`
1918
+ );
1919
+ }
1886
1920
  }
1887
1921
 
1888
1922
  // src/report/junit.ts
@@ -1953,9 +1987,45 @@ ${reason}` : reason;
1953
1987
  return lines.join("\n");
1954
1988
  }
1955
1989
  async function writeJUnit(file, xml) {
1956
- await mkdir5(path8.dirname(file), { recursive: true });
1957
- await writeFile5(file, xml, "utf8");
1958
- return file;
1990
+ try {
1991
+ await mkdir5(path8.dirname(file), { recursive: true });
1992
+ await writeFile5(file, xml, "utf8");
1993
+ return file;
1994
+ } catch (error) {
1995
+ throw new ReportError(
1996
+ `Cannot write JUnit report to ${file}: ${fsReason(error)}. Check that the path is writable and is not a directory.`
1997
+ );
1998
+ }
1999
+ }
2000
+
2001
+ // src/runner/authoring.ts
2002
+ var VALUE_VERBS = ["fill", "enter", "type", "input", "set"];
2003
+ var LEADING_VALUE_VERB = new RegExp(`^\\s*(?:${VALUE_VERBS.join("|")})\\b`, "i");
2004
+ var CONNECTOR_WORDS = ["with", "to", "as", "using", "into", "from", "in"];
2005
+ var CONNECTOR_WORD = new RegExp(`\\b(?:${CONNECTOR_WORDS.join("|")})\\b`, "i");
2006
+ var CONNECTOR_SYMBOL = /["'`:=]|\{\{env\./i;
2007
+ var PHRASAL_IN = new RegExp(`^(\\s*(?:${VALUE_VERBS.join("|")}))\\s+in\\b`, "i");
2008
+ function namesNoValue(step) {
2009
+ if (!LEADING_VALUE_VERB.test(step)) return false;
2010
+ const withoutPhrasal = step.replace(PHRASAL_IN, "$1");
2011
+ return !CONNECTOR_WORD.test(withoutPhrasal) && !CONNECTOR_SYMBOL.test(withoutPhrasal);
2012
+ }
2013
+ function detectMissingValues(tests) {
2014
+ const findings = [];
2015
+ for (const test of tests) {
2016
+ for (const origin of ["setup", "steps"]) {
2017
+ const steps = origin === "setup" ? test.setup ?? [] : test.steps;
2018
+ steps.forEach((step, position) => {
2019
+ if (namesNoValue(step)) {
2020
+ findings.push({ test, origin, index: position + 1, step });
2021
+ }
2022
+ });
2023
+ }
2024
+ }
2025
+ return { findings };
2026
+ }
2027
+ function suggestValueClause(step) {
2028
+ return `${step.trimEnd()} with <value>`;
1959
2029
  }
1960
2030
 
1961
2031
  // src/runner/pool.ts
@@ -2013,6 +2083,16 @@ function selectImpactedTests(tests, affectedRoutes, filters = { tags: [] }) {
2013
2083
  uncoveredRoutes: affectedRoutes.filter((route) => !covered.has(route))
2014
2084
  };
2015
2085
  }
2086
+ function detectRouteDrift(tests, declaredRoutes) {
2087
+ const known = new Set(declaredRoutes);
2088
+ if (known.size === 0) return { drifted: [] };
2089
+ const drifted = [];
2090
+ for (const test of tests) {
2091
+ const unknown = [...new Set(test.routes)].filter((route) => !known.has(route)).sort();
2092
+ if (unknown.length > 0) drifted.push({ test, routes: unknown });
2093
+ }
2094
+ return { drifted };
2095
+ }
2016
2096
 
2017
2097
  // src/commands/run.ts
2018
2098
  var EXIT_OK = 0;
@@ -2242,6 +2322,34 @@ function reportUnclassified(options, impact) {
2242
2322
  );
2243
2323
  return true;
2244
2324
  }
2325
+ function declaredConfigRoutes(config) {
2326
+ const routes = config.routes ? Object.values(config.routes) : [];
2327
+ return [...new Set(routes.flat())].sort();
2328
+ }
2329
+ function printRouteDrift(drift, cwd) {
2330
+ if (drift.drifted.length === 0) return;
2331
+ console.error(
2332
+ "Route drift (test routes declared by no routes: mapping \u2014 contribute nothing to --impacted selection):"
2333
+ );
2334
+ for (const { test, routes } of drift.drifted) {
2335
+ console.error(` ${test.summary} (${path9.relative(cwd, test.path)}): ${routes.join(", ")}`);
2336
+ }
2337
+ console.error("Fix the route typo, or add the route to a routes: mapping in .blastproof/config.yaml.");
2338
+ }
2339
+ function printAuthoring(result, cwd) {
2340
+ if (result.findings.length === 0) return;
2341
+ console.error("\nAuthoring (a step enters a value but names none):");
2342
+ for (const { test, origin, index, step } of result.findings) {
2343
+ const where = origin === "setup" ? `setup step ${index}` : `step ${index}`;
2344
+ console.error(` ${test.summary} (${path9.relative(cwd, test.path)}) ${where}:`);
2345
+ console.error(` ${step}`);
2346
+ console.error(` \u2192 ${suggestValueClause(step)}`);
2347
+ }
2348
+ console.error(
2349
+ "The runner is forbidden from inventing values, so a step that supplies none cannot be carried out."
2350
+ );
2351
+ console.error("Steps are inspected in English only; steps in other languages are not checked.");
2352
+ }
2245
2353
  function printDryRun(selected, config, cwd) {
2246
2354
  console.log(`
2247
2355
  Dry run: ${selected.length} test(s) selected, base_url=${config.base_url}`);
@@ -2324,6 +2432,16 @@ async function runCommand(options) {
2324
2432
  throw error;
2325
2433
  }
2326
2434
  }
2435
+ const drift = detectRouteDrift(parsed, declaredConfigRoutes(config));
2436
+ printRouteDrift(drift, options.cwd);
2437
+ const authoring = detectMissingValues(parsed);
2438
+ printAuthoring(authoring, options.cwd);
2439
+ if (options.failOnAuthoring && authoring.findings.length > 0) {
2440
+ console.error(
2441
+ `error: --fail-on-authoring: ${authoring.findings.length} step(s) enter a value but name none.`
2442
+ );
2443
+ return EXIT_FAILED;
2444
+ }
2327
2445
  const selection = impact ? selectImpactedTests(parsed, impact.affectedRoutes, options) : {
2328
2446
  selected: parsed.filter((test) => matchesFilters(test, options)),
2329
2447
  unroutedSkipped: [],
@@ -2718,6 +2836,7 @@ async function testCommand(options) {
2718
2836
  junit: options.junit,
2719
2837
  html: options.html,
2720
2838
  failOnUnmapped: options.failOnUnmapped,
2839
+ failOnAuthoring: options.failOnAuthoring,
2721
2840
  budget
2722
2841
  });
2723
2842
  if (runCode === EXIT_USAGE) return EXIT_USAGE;
@@ -2774,7 +2893,7 @@ function parsePositiveNumber(flag) {
2774
2893
  };
2775
2894
  }
2776
2895
  var program = new Command();
2777
- program.name("blastproof").description("Open-source AI testing agent: plain-English YAML tests executed agentically on a real browser.").version("0.10.0");
2896
+ program.name("blastproof").description("Open-source AI testing agent: plain-English YAML tests executed agentically on a real browser.").version("0.12.0");
2778
2897
  program.command("init").description("Scaffold .blastproof/ (config, tests, sample tests) in the current directory").action(async () => {
2779
2898
  try {
2780
2899
  const result = await initProject(process.cwd());
@@ -2789,7 +2908,7 @@ program.command("run").description("Discover and run all tests under .blastproof
2789
2908
  "--min-score <n>",
2790
2909
  "require a weighted score of at least n (0-100); replaces the all-must-pass rule",
2791
2910
  parseMinScore
2792
- ).option("--junit [path]", "write a JUnit XML report (default: .blastproof/reports/<session>/junit.xml)").option("--html [path]", "write a self-contained HTML report (default: .blastproof/reports/<session>/report.html)").option("--fail-on-unmapped", "fail when a changed file matches no routes: or ignore: glob").option(
2911
+ ).option("--junit [path]", "write a JUnit XML report (default: .blastproof/reports/<session>/junit.xml)").option("--html [path]", "write a self-contained HTML report (default: .blastproof/reports/<session>/report.html)").option("--fail-on-unmapped", "fail when a changed file matches no routes: or ignore: glob").option("--fail-on-authoring", "fail when a step enters a value but names none (warns by default)").option(
2793
2912
  "--concurrency <n>",
2794
2913
  "run this many tests at once (overrides config; default 1 \u2014 see the README on when this is safe)",
2795
2914
  parsePositiveInt("--concurrency")
@@ -2821,6 +2940,7 @@ program.command("run").description("Discover and run all tests under .blastproof
2821
2940
  junit: options.junit,
2822
2941
  html: options.html,
2823
2942
  failOnUnmapped: options.failOnUnmapped,
2943
+ failOnAuthoring: options.failOnAuthoring,
2824
2944
  concurrency: options.concurrency,
2825
2945
  maxLlmCalls: options.maxLlmCalls,
2826
2946
  maxTokens: options.maxTokens,
@@ -2871,7 +2991,7 @@ program.command("test").description("Full PR pipeline: run the tests covering th
2871
2991
  "--min-score <n>",
2872
2992
  "require a weighted score of at least n (0-100); replaces the all-must-pass rule",
2873
2993
  parseMinScore
2874
- ).option("--junit [path]", "write a JUnit XML report (default: .blastproof/reports/<session>/junit.xml)").option("--html [path]", "write a self-contained HTML report (default: .blastproof/reports/<session>/report.html)").option("--write", "persist generated drafts under .blastproof/tests/ instead of previewing them").option("--fail-on-unmapped", "fail when a changed file matches no routes: or ignore: glob").option(
2994
+ ).option("--junit [path]", "write a JUnit XML report (default: .blastproof/reports/<session>/junit.xml)").option("--html [path]", "write a self-contained HTML report (default: .blastproof/reports/<session>/report.html)").option("--write", "persist generated drafts under .blastproof/tests/ instead of previewing them").option("--fail-on-unmapped", "fail when a changed file matches no routes: or ignore: glob").option("--fail-on-authoring", "fail when a step enters a value but names none (warns by default)").option(
2875
2995
  "--max-llm-calls <n>",
2876
2996
  "stop the pipeline after this many model calls, shared by both phases (overrides config)",
2877
2997
  parsePositiveInt("--max-llm-calls")
@@ -2895,6 +3015,7 @@ program.command("test").description("Full PR pipeline: run the tests covering th
2895
3015
  html: options.html,
2896
3016
  write: options.write,
2897
3017
  failOnUnmapped: options.failOnUnmapped,
3018
+ failOnAuthoring: options.failOnAuthoring,
2898
3019
  maxLlmCalls: options.maxLlmCalls,
2899
3020
  maxTokens: options.maxTokens,
2900
3021
  maxDuration: options.maxDuration