blastproof 0.11.0 → 0.12.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -1,4 +1,10 @@
1
- # blastproof
1
+ <h1>
2
+ <picture>
3
+ <source media="(prefers-color-scheme: dark)" srcset="./.github/logo-dark.svg">
4
+ <img src="./.github/logo-light.svg" alt="" width="30" height="30">
5
+ </picture>
6
+ blastproof
7
+ </h1>
2
8
 
3
9
  [![CI](https://github.com/hamc/blastproof/actions/workflows/ci.yml/badge.svg)](https://github.com/hamc/blastproof/actions/workflows/ci.yml)
4
10
  [![Dogfood](https://github.com/hamc/blastproof/actions/workflows/dogfood.yml/badge.svg)](https://github.com/hamc/blastproof/actions/workflows/dogfood.yml)
@@ -12,6 +18,10 @@ git diff → impact mapping → test generation → agentic execution → report
12
18
 
13
19
  100% local. MIT. Bring your own LLM key.
14
20
 
21
+ ▶ **[Watch the introduction](https://www.youtube.com/shorts/miqN5FzMF_k)** — what it does, in a minute.
22
+
23
+ **Documentation:** [Configuration](./docs/configuration.md) · [Testing behind a login](./docs/auth.md) · [Running in CI](./docs/ci.md) · [Contributing](./CONTRIBUTING.md) · [Architecture](./AGENTS.md)
24
+
15
25
  ## Quick start
16
26
 
17
27
  ```bash
@@ -38,21 +48,53 @@ blastproof run
38
48
 
39
49
  Before `run`, `plan` or `test` do anything, they check what they are about to spend — the browser can launch, the model provider is reachable, `base_url` responds — and report every unmet one together, so a stopped app or a missing browser is never a wall you hit one crash at a time. A missing system library names the exact install command and says it needs root; nothing is installed on your behalf. Silent when everything is fine, and skipped entirely by `--dry-run`, which needs none of it.
40
50
 
51
+ Provider options, budgets and browser tuning: [Configuration](./docs/configuration.md).
52
+
41
53
  ## Does this fit your application?
42
54
 
43
- **Your markup must be accessible — a hard requirement.** Elements are found by role, label or visible text from the accessibility tree. That is what removes selectors and survives redesigns; the cost is that an interface the accessibility tree cannot describe cannot be driven at all, and there is deliberately no CSS or XPath fallback. Icon-only buttons without accessible names, `div`-based controls and ARIA-less dropdowns simply cannot be targeted. Run an accessibility checker first — the result predicts how well this will work better than anything else.
55
+ Three questions. The first one decides most cases.
56
+
57
+ ### 1. Is your markup accessible?
58
+
59
+ **A hard requirement, not a preference.** blastproof finds elements the way a screen reader does — by role, by label, by visible text. That is what removes selectors and survives redesigns. The cost is that there is deliberately no CSS or XPath fallback, so anything the accessibility tree cannot describe cannot be driven at all.
60
+
61
+ | works | cannot be driven |
62
+ | --- | --- |
63
+ | `<button>Add to cart</button>` | a `<div>` with a click handler |
64
+ | `<label for="email">` + `<input>` | an input with no label |
65
+ | `<button aria-label="Delete note">` | an icon-only button with no name |
66
+ | `<select>` with `<option>`s | an ARIA-less custom dropdown |
67
+
68
+ **Run an accessibility checker on your app before installing anything.** The result predicts how well this will work better than anything else you could measure — and the fixes it suggests are worth making regardless of whether you adopt this tool.
69
+
70
+ ### 2. Does your journey need anything on this list?
71
+
72
+ Not supported yet:
44
73
 
45
- **Windows is untested.** Development and CI run on Linux and macOS. Nothing is known to be broken and reports are welcome (#8).
74
+ - **`iframe` content** — a hosted payment widget is invisible, so an embedded checkout cannot be driven end to end
75
+ - **hover, scroll-to, drag and drop, file upload**
76
+ - **multiple tabs**, and native `alert` / `confirm` dialogs
46
77
 
47
- **Not supported yet:** `iframe` content (so hosted payment widgets like Stripe Elements are invisible — an embedded checkout cannot be driven end to end), hover, scroll-to, drag and drop, file upload, multiple tabs, native `alert`/`confirm` dialogs. Page snapshots are capped at 200 lines by default, so very dense pages are truncated — raise it with `browser.max_snapshot_lines` if your pages need more; truncation is always marked in the snapshot so the model is never misled into thinking it saw the whole page.
78
+ **Windows is untested.** Development and CI run on Linux and macOS. Nothing is known to be broken and reports are welcome ([#8](https://github.com/hamc/blastproof/issues/8)).
48
79
 
49
- **Point it at disposable data.** Within a step, an action that commits — a click, or pressing Enter — is never performed twice: the runner refuses the repeat and tells the agent it already did that. This closes the case that used to produce duplicate records, where a submit answered by a redirect came back to a reset form and the agent, seeing no evidence of its own work, submitted again. It is not a guarantee of zero duplicate writes: an agent that reaches the same effect by a genuinely different route — another control with the same effect — is not caught. Use a seeded database, a staging environment you can reset, or a throwaway account; do not gate on a run against production data.
80
+ If a critical journey needs one of these, that journey stays with your existing test suite. The two can coexist — nothing here replaces what you already have.
50
81
 
51
- `browser.timeout_ms` bounds every wait — resolving a target element from the accessibility tree, and navigation — not only the click or fill performed afterwards. Raise it for an application that is merely slow to hydrate; the trade-off is that a genuinely missing element then takes longer to fail. It never changes how many self-healing retries a step gets — waiting and retrying are deliberately separate.
82
+ ### 3. Can you point it at data you can throw away?
52
83
 
53
- Writing each step so it states its own outcome helps here as well as everywhere else: `submit the form, then verify the confirmation shows the reference number` gives the agent something to check, where `click the submit button` leaves it to invent an expectation — and a poor invented expectation is what turns one submission into three.
84
+ **Use a seeded database, a staging environment you can reset, or a throwaway account. Do not gate on a run against production data.**
54
85
 
55
- **This is now load-bearing, not just advisable.** The judge decides whether a step's own outcome holds, using the model's expectation only as the claim offered in support of it — a step that never says what its outcome is gives the judge nothing to anchor on beyond whatever the model happened to check that turn. `verify the confirmation shows the reference number` gives the judge a real question; `verify it worked` does not, and may now fail where a looser judge previously let a true-but-unrelated claim pass it.
86
+ Within a step, an action that commits — a click, or pressing Enter — is never performed twice: the runner refuses the repeat and tells the agent it already did that. This closes the case that used to produce duplicate records, where a submit answered by a redirect came back to a reset form and the agent, seeing no evidence of its own work, submitted again.
87
+
88
+ It is **not** a guarantee of zero duplicate writes. An agent that reaches the same effect by a genuinely different route — another control that does the same thing — is not caught.
89
+
90
+ ## How it works
91
+
92
+ <img src="./.github/coverage-flow.svg" width="100%"
93
+ alt="How a diff becomes a merge decision. In CI, unattended: changed files are matched against the routes: and ignore: globs; matched files contribute affected routes, files matching neither are reported as unclassified and fail the run only under --fail-on-unmapped. Tests declaring an affected route are executed and produce a weighted score, which --min-score gates on. An affected route no test declares is reported as a coverage gap and never fails the run. Separately and manually, outside CI: blastproof plan loads such a route in Chromium, makes one model call, and produces a YAML draft you review, edit and run before committing it.">
94
+
95
+ **The boundary in the middle is the point.** Everything above it runs unattended on every pull request and ends in an exit code. Everything below it is something you choose to run, on your machine, and review before it lands.
96
+
97
+ A route no test covers is *reported*, never failed — blocking on it would punish you for an incomplete map instead of teaching you to complete it. Turning that report into a test is the manual half, and the draft it produces is not trusted until a person has read it.
56
98
 
57
99
  ## Writing tests
58
100
 
@@ -70,7 +112,7 @@ steps:
70
112
  - complete checkout
71
113
  ```
72
114
 
73
- `priority` is P0–P2 (default P1). `tags`, `setup` steps and `auth` are optional — `auth: false` runs the test signed out, which a login test needs. `routes` declares the URLs a test covers, which is what `--impacted` selects on; write route strings consistently, since they compare by exact equality (`/cart` ≠ `/cart/`).
115
+ `priority` is P0–P2 (default P1). `tags`, `setup` steps and `auth` are optional — `auth: false` runs the test signed out, which a login test needs. `routes` declares the URLs a test covers, which is what `--impacted` selects on; write route strings consistently, since they compare by exact equality (`/cart` ≠ `/cart/`). `run` warns to stderr — non-fatal — when a test declares a route no `routes:` mapping declares, since that route contributes nothing to `--impacted` selection.
74
116
 
75
117
  ### Say what each step should produce
76
118
 
@@ -92,6 +134,30 @@ Write steps that end in an observable result — text on the page, a count, a st
92
134
 
93
135
  **Inline error messages should be plain visible text.** `role="alert"` is read correctly from the accessibility tree and needs no special handling, but note that an alert your page has cleared shows up as an empty element: if a verdict says an alert exists whose content is missing, the message was emptied, not hidden.
94
136
 
137
+ ### A step that enters a value writes the value
138
+
139
+ ```yaml
140
+ - fill the note field with Order not received # runs
141
+ - fill the note field # cannot run
142
+ ```
143
+
144
+ The agent is **forbidden from inventing values** — one it types must come from the step, from the page, or from an `{{env.*}}` placeholder. But that rule lives in a prompt, and a prompt instructs rather than enforces. Run against a real model, the second step does not fail: the agent makes a value up, fills it, and the step **passes**. `fill the note field` produced "This is a test note." on two runs and "This is a new note" on a third.
145
+
146
+ That is worse than a failure. The test goes green having verified a value nobody wrote, differing between runs — a passing check over an unspecified input. Closing the gap in the runner is [#57](https://github.com/hamc/blastproof/issues/57); until then, this warning is what stands between you and a green test that means nothing.
147
+
148
+ `run` warns about it first, on every path, before launching a browser or asking for a key:
149
+
150
+ ```
151
+ Authoring (a step enters a value but names none):
152
+ Add a note (.blastproof/tests/notes.yaml) step 2:
153
+ fill the note field
154
+ → fill the note field with <value>
155
+ ```
156
+
157
+ Non-fatal by default — `--fail-on-authoring` turns it into exit 1 for teams enforcing it in CI. Taking the value from the page is fine and is not flagged: `fill the recipient field with the address shown on the confirmation page`.
158
+
159
+ **The check reads English only.** A suite written in another language runs exactly as well but is not inspected, and prints no warning saying so — silence from this check means "nothing found in English", never "this suite is clean".
160
+
95
161
  ## Commands
96
162
 
97
163
  ```bash
@@ -111,43 +177,16 @@ Common flags — `blastproof <command> --help` has the full list:
111
177
  | `--url <url>` | Override `base_url` for this run (e.g. a PR preview) |
112
178
  | `--min-score <n>` | Gate on a weighted score instead of all-must-pass |
113
179
  | `--fail-on-unmapped` | Fail when a changed file matches no `routes:` or `ignore:` glob |
180
+ | `--fail-on-authoring` | Fail when a step enters a value but names none (warns by default) |
114
181
  | `--junit [path]` · `--html [path]` | Write reports |
182
+ | `--concurrency <n>` | Run tests at once — [when that is safe](./docs/configuration.md#concurrency--running-tests-at-once) |
115
183
  | `--write` | `plan` only — persist drafts instead of previewing |
116
- | `--max-llm-calls` · `--max-tokens` · `--max-duration` | Bound what a run may spend |
184
+ | `--max-llm-calls` · `--max-tokens` · `--max-duration` | [Bound what a run may spend](./docs/configuration.md#budget--bounding-what-a-run-spends) |
117
185
 
118
186
  Exit codes: **0** pass, **1** the gate failed, **2** usage or config error.
119
187
 
120
188
  **Generated drafts are never executed and never affect the score.** An unreviewed model-written test in the merge path fails in two directions: a hallucinated expectation blocks a correct PR, and a credulous one waves a broken change through while looking like coverage. `plan` makes the gap visible with a draft to review; it does not make an uncovered route safe.
121
189
 
122
- ## In CI
123
-
124
- ```yaml
125
- name: blastproof
126
- on: pull_request
127
-
128
- jobs:
129
- test:
130
- runs-on: ubuntu-latest
131
- steps:
132
- - uses: actions/checkout@v4
133
- with:
134
- fetch-depth: 0 # required — the diff needs a merge-base
135
-
136
- - run: npm start & # however your app boots
137
-
138
- - uses: hamc/blastproof@v0.11.0
139
- with:
140
- version: '0.6.0' # pin both when this gates merges
141
- api-key: ${{ secrets.ANTHROPIC_API_KEY }}
142
- base: ${{ github.event.pull_request.base.ref }}
143
- min-score: '80'
144
- fail-on-unmapped: 'true'
145
- ```
146
-
147
- A non-zero exit blocks the merge. The action outputs `score` (0–100, empty when no report was produced) for later steps. Full input list: [`action.yml`](./action.yml).
148
-
149
- `fetch-depth: 0` is not optional — the default checkout is shallow and has no merge-base. The action detects this and fails immediately rather than letting it surface as a git error mid-run.
150
-
151
190
  ## Impact mapping
152
191
 
153
192
  `--impacted` runs only the tests whose `routes:` intersect the routes your diff can affect, mapped from changed files by globs you maintain in `.blastproof/config.yaml`:
@@ -176,6 +215,8 @@ blastproof run --min-score 80 # one failing P2 is tolerated
176
215
 
177
216
  `--min-score` **replaces** the all-must-pass rule rather than adding to it. Only executed tests count: filtered and unrouted tests are neither numerator nor denominator, and a run that executed nothing scores 100, so a docs-only PR is never blocked. JUnit carries the score as a `<property name="score">`, and unrouted tests appear as `<skipped/>` so the coverage gap shows up in CI rather than vanishing.
178
217
 
218
+ Wiring this into a pipeline, with the gating patterns worth knowing: [Running in CI](./docs/ci.md).
219
+
179
220
  ## Without a browser or a key
180
221
 
181
222
  Half of blastproof is deterministic and free. These need no model, no browser and no network:
@@ -189,103 +230,6 @@ blastproof plan --base main --dry-run # affected routes no test
189
230
 
190
231
  They report affected routes, files nobody has classified, and affected routes no test covers — a coverage-gap report with an exit code, useful even on a repo whose suite is Playwright or Cypress.
191
232
 
192
- ## Running tests at once
193
-
194
- Tests run one at a time by default. Raise it when your tests can stand it:
195
-
196
- ```yaml
197
- concurrency: 4
198
- ```
199
-
200
- or `blastproof run --concurrency 4` for a single invocation. On this repository's own suite that takes a run from 156s to 68s — **2.3× faster, for the same 81 model calls.** Parallelism buys wall-clock, not spend.
201
-
202
- **The default is 1 on purpose, and raising it is your call to make.** Other test runners default to parallel because their tests are isolated by construction — separate processes, separate fixtures. These are journeys driven against **one running application**, so two tests can see each other's data. A suite is safe to parallelise when its tests do not write state that another test reads.
203
-
204
- The test in this repository's own suite that could not run beside itself is a good shape to recognise: it adds a note and then asserts *"one note on file"*. It writes shared server state, and it asserts on a global count. Either alone is a warning; together they mean the test's verdict depends on nothing else touching the application at that moment.
205
-
206
- Two practical notes. Four concurrent journeys are four times the traffic against whatever you pointed at — usually fine for a development instance, worth knowing for a shared one. And with several model calls in flight, a `budget:` limit can overshoot by up to the concurrency rather than by a single call, since the calls already sent are allowed to finish.
207
-
208
- ## Bounding a run
209
-
210
- Nothing stops a run by default. `budget:` puts a ceiling on `run`, `plan` and `test` alike — every model call any of them makes is counted:
211
-
212
- ```yaml
213
- budget:
214
- max_llm_calls: 500
215
- max_tokens: 2000000
216
- max_duration_s: 900
217
- ```
218
-
219
- Each limit is optional; with none set, nothing binds. They count **calls and tokens, not currency** — a price table keyed by model and provider goes stale the day a provider reprices, and a limit that quietly stops meaning what it says is worse than none, because it is trusted.
220
-
221
- Exhausting a budget **stops the run; it does not fail a test.** Running out of quota says nothing about the code under review. Unreached tests are reported as `not run`, a third state excluded from the score entirely, and the process exits 1 unconditionally — `--min-score` cannot rescue it, because the tests that finished are whichever ran first, not a representative sample.
222
-
223
- **Every run reports what it spent**, so you can size a limit from experience rather than guesswork:
224
-
225
- ```
226
- Spent: 82 model call(s), 115407 token(s)
227
- Score: 100
228
- ```
229
-
230
- The figures also land in the JUnit report as `llm_calls` and `llm_tokens`, beside `score`, so a pipeline can trend cost without scraping output. A run stopped by its own budget reports the spend too — that is the case where the number is least guessable. Where a provider reports no token usage, the line says so rather than showing zero.
231
-
232
- `--dry-run` reports the ceiling before you spend anything. Read it as a maximum and nothing more: for this repository's own suite it says 735 calls where a real run spends 82. Size a budget from what your runs actually report, not from the ceiling.
233
-
234
- For an order of magnitude, this repository's own suite — 7 tests, 31 steps, an authenticated demo shop, `anthropic/claude-haiku-4.5` — spends **about 82 model calls and 115k tokens**, taking 156s serially or 68s at `--concurrency 4`. That is a number you can reproduce (`node examples/demo-app/serve.mjs 4173` and `blastproof run`), not a forecast for your suite: cost scales with steps, page density and how often the agent has to retry. Run yours once and read the `Spent:` line.
235
-
236
- ## Testing behind a login
237
-
238
- Declare a recipe once; blastproof signs in one time per run and reuses that session for every test and for `plan`. Pick exactly one strategy:
239
-
240
- ```yaml
241
- # 1) A plain-English journey — form login, or anything a person can click through
242
- auth:
243
- steps:
244
- - navigate to /login
245
- - fill the email field with {{env.TEST_EMAIL}}
246
- - fill the password field with {{env.TEST_PASSWORD}}
247
- - submit the login form
248
- verify: a signed-in indicator is visible # optional, strongly recommended
249
-
250
- # 2) A session captured by hand — for SSO, MFA or magic links
251
- auth:
252
- storage_state: .blastproof/auth.json
253
-
254
- # 3) Static values — for token-based apps
255
- auth:
256
- headers:
257
- Authorization: "Bearer {{env.API_TOKEN}}"
258
- ```
259
-
260
- **`verify` is worth the extra call.** Without it a wrong password surfaces as every test failing on a login wall — N failures, none naming the cause. Authentication failure exits 2 and never reports as failing tests, because a login you cannot complete says nothing about the code under review.
261
-
262
- **A captured session is a credential** — the file holds live cookies. `init` git-ignores it; never commit one.
263
-
264
- ## LLM providers (BYOK)
265
-
266
- **Anthropic** (`ANTHROPIC_API_KEY`), **OpenAI** (`OPENAI_API_KEY`), or **Ollama** (local, no key). Any setting can be overridden from the environment, with precedence **CLI flag > environment > file**:
267
-
268
- | variable | overrides |
269
- | --- | --- |
270
- | `BLASTPROOF_BASE_URL` | `base_url` — the app under test |
271
- | `BLASTPROOF_LLM_PROVIDER` | `anthropic` \| `openai` \| `ollama` |
272
- | `BLASTPROOF_LLM_MODEL` | the model name |
273
- | `BLASTPROOF_LLM_BASE_URL` | the provider endpoint — *not* the app |
274
- | `BLASTPROOF_LLM_API_KEY_ENV` | the **name** of the variable holding your key |
275
- | `BLASTPROOF_MAX_LLM_CALLS` · `BLASTPROOF_MAX_TOKENS` · `BLASTPROOF_MAX_DURATION_S` | the budget fields |
276
-
277
- So an OpenAI-compatible gateway needs no file edit:
278
-
279
- ```bash
280
- export BLASTPROOF_LLM_PROVIDER=openai
281
- export BLASTPROOF_LLM_MODEL=anthropic/claude-haiku-4.5
282
- export BLASTPROOF_LLM_BASE_URL=https://openrouter.ai/api/v1
283
- export BLASTPROOF_LLM_API_KEY_ENV=OPENROUTER_API_KEY
284
- blastproof run --impacted --min-score 80
285
- ```
286
-
287
- The key itself is never read from a `BLASTPROOF_*` variable — you name *which* variable holds it, so errors can keep naming the one you chose.
288
-
289
233
  ## Trust boundaries
290
234
 
291
235
  The application under test is not trusted input: its page content reaches the model, so a page that controls its own accessible text can try to influence the agent. Two things constrain that.
package/dist/cli.js CHANGED
@@ -7,6 +7,20 @@ import { Command, InvalidArgumentError } from "commander";
7
7
  import { mkdir, writeFile } from "fs/promises";
8
8
  import path from "path";
9
9
  import { access } from "fs/promises";
10
+
11
+ // src/report/errors.ts
12
+ var ReportError = class extends Error {
13
+ constructor(message) {
14
+ super(message);
15
+ this.name = "ReportError";
16
+ }
17
+ };
18
+ function fsReason(error) {
19
+ if (!(error instanceof Error)) return String(error);
20
+ return error.message.replace(/^E[A-Z]+:\s*/, "").replace(/,\s+\w+\s+'[^']*'$/, "").trim();
21
+ }
22
+
23
+ // src/commands/init.ts
10
24
  var DEFAULT_CONFIG = `# blastproof configuration
11
25
  # Docs: https://github.com/hamc/blastproof
12
26
 
@@ -124,14 +138,26 @@ async function exists(file) {
124
138
  return false;
125
139
  }
126
140
  }
141
+ var InitError = class extends Error {
142
+ constructor(message) {
143
+ super(message);
144
+ this.name = "InitError";
145
+ }
146
+ };
127
147
  async function writeIfAbsent(file, content, result) {
128
148
  if (await exists(file)) {
129
149
  result.kept.push(file);
130
150
  return;
131
151
  }
132
- await mkdir(path.dirname(file), { recursive: true });
133
- await writeFile(file, content);
134
- result.created.push(file);
152
+ try {
153
+ await mkdir(path.dirname(file), { recursive: true });
154
+ await writeFile(file, content);
155
+ result.created.push(file);
156
+ } catch (error) {
157
+ throw new InitError(
158
+ `Cannot scaffold ${file}: ${fsReason(error)}. Check that ${path.dirname(file)} is a directory you can write to, not a file.`
159
+ );
160
+ }
135
161
  }
136
162
  async function initProject(cwd = process.cwd()) {
137
163
  const result = { created: [], kept: [] };
@@ -1254,7 +1280,7 @@ Rules:
1254
1280
  - Prefer the journey the changed files touch over a generic tour of the page. The changed files tell you which part of the page matters.
1255
1281
  - **The test starts at the application's base URL, not at this route.** Begin with a step that navigates to the route and says what should be visible once it loads \u2014 "navigate to /support and verify the heading "Contact support" is shown". Without it the run opens the home page and every later step looks for controls that are not there.
1256
1282
  - **Every step says what it should produce.** Name what must be true once the step has been carried out, not the action alone: "submit the support form and verify the confirmation page shows the ticket number", never "submit the support form". A step that names an action without an outcome asks the runner to judge whether something happened while looking at the page that succeeding produces \u2014 a submitted form comes back empty, a redirect moves the URL \u2014 and that is the shape behind several real failures.
1257
- - **A step that enters a value writes the value.** "fill the subject field with Order not received", never "enter a subject". The runner is forbidden from inventing values, so a step that does not supply one cannot be carried out.
1283
+ - **A step that enters a value writes the value.** "fill the subject field with Order not received", never "enter a subject". The runner is forbidden from inventing values, but a step that supplies none does not fail \u2014 the value gets made up, the step passes, and the test then verifies something nobody wrote.
1258
1284
  - If a step needs a credential or any secret, write it as a placeholder like {{env.TEST_PASSWORD}}. Never write a real or invented password, token or key.
1259
1285
  - Keep the whole test to a handful of steps: one journey, not an exhaustive suite.`;
1260
1286
  }
@@ -1447,7 +1473,7 @@ var TestFileError = class extends Error {
1447
1473
  this.name = "TestFileError";
1448
1474
  }
1449
1475
  };
1450
- function formatIssues2(file, error) {
1476
+ function formatIssues2(error) {
1451
1477
  return error.issues.map((issue) => ` - ${issue.path.join(".") || "(root)"}: ${issue.message}`).join("\n");
1452
1478
  }
1453
1479
  async function parseTestFile(filePath) {
@@ -1468,7 +1494,7 @@ async function parseTestFile(filePath) {
1468
1494
  const result = testFileSchema.safeParse(data ?? {});
1469
1495
  if (!result.success) {
1470
1496
  throw new TestFileError(`Invalid test file ${filePath}:
1471
- ${formatIssues2(filePath, result.error)}`);
1497
+ ${formatIssues2(result.error)}`);
1472
1498
  }
1473
1499
  return { path: filePath, ...result.data };
1474
1500
  }
@@ -1882,9 +1908,15 @@ ${detail}
1882
1908
  `;
1883
1909
  }
1884
1910
  async function writeHtml(file, html) {
1885
- await mkdir4(path7.dirname(file), { recursive: true });
1886
- await writeFile4(file, html, "utf8");
1887
- return file;
1911
+ try {
1912
+ await mkdir4(path7.dirname(file), { recursive: true });
1913
+ await writeFile4(file, html, "utf8");
1914
+ return file;
1915
+ } catch (error) {
1916
+ throw new ReportError(
1917
+ `Cannot write HTML report to ${file}: ${fsReason(error)}. Check that the path is writable and is not a directory.`
1918
+ );
1919
+ }
1888
1920
  }
1889
1921
 
1890
1922
  // src/report/junit.ts
@@ -1955,9 +1987,45 @@ ${reason}` : reason;
1955
1987
  return lines.join("\n");
1956
1988
  }
1957
1989
  async function writeJUnit(file, xml) {
1958
- await mkdir5(path8.dirname(file), { recursive: true });
1959
- await writeFile5(file, xml, "utf8");
1960
- return file;
1990
+ try {
1991
+ await mkdir5(path8.dirname(file), { recursive: true });
1992
+ await writeFile5(file, xml, "utf8");
1993
+ return file;
1994
+ } catch (error) {
1995
+ throw new ReportError(
1996
+ `Cannot write JUnit report to ${file}: ${fsReason(error)}. Check that the path is writable and is not a directory.`
1997
+ );
1998
+ }
1999
+ }
2000
+
2001
+ // src/runner/authoring.ts
2002
+ var VALUE_VERBS = ["fill", "enter", "type", "input", "set"];
2003
+ var LEADING_VALUE_VERB = new RegExp(`^\\s*(?:${VALUE_VERBS.join("|")})\\b`, "i");
2004
+ var CONNECTOR_WORDS = ["with", "to", "as", "using", "into", "from", "in"];
2005
+ var CONNECTOR_WORD = new RegExp(`\\b(?:${CONNECTOR_WORDS.join("|")})\\b`, "i");
2006
+ var CONNECTOR_SYMBOL = /["'`:=]|\{\{env\./i;
2007
+ var PHRASAL_IN = new RegExp(`^(\\s*(?:${VALUE_VERBS.join("|")}))\\s+in\\b`, "i");
2008
+ function namesNoValue(step) {
2009
+ if (!LEADING_VALUE_VERB.test(step)) return false;
2010
+ const withoutPhrasal = step.replace(PHRASAL_IN, "$1");
2011
+ return !CONNECTOR_WORD.test(withoutPhrasal) && !CONNECTOR_SYMBOL.test(withoutPhrasal);
2012
+ }
2013
+ function detectMissingValues(tests) {
2014
+ const findings = [];
2015
+ for (const test of tests) {
2016
+ for (const origin of ["setup", "steps"]) {
2017
+ const steps = origin === "setup" ? test.setup ?? [] : test.steps;
2018
+ steps.forEach((step, position) => {
2019
+ if (namesNoValue(step)) {
2020
+ findings.push({ test, origin, index: position + 1, step });
2021
+ }
2022
+ });
2023
+ }
2024
+ }
2025
+ return { findings };
2026
+ }
2027
+ function suggestValueClause(step) {
2028
+ return `${step.trimEnd()} with <value>`;
1961
2029
  }
1962
2030
 
1963
2031
  // src/runner/pool.ts
@@ -2015,6 +2083,16 @@ function selectImpactedTests(tests, affectedRoutes, filters = { tags: [] }) {
2015
2083
  uncoveredRoutes: affectedRoutes.filter((route) => !covered.has(route))
2016
2084
  };
2017
2085
  }
2086
+ function detectRouteDrift(tests, declaredRoutes) {
2087
+ const known = new Set(declaredRoutes);
2088
+ if (known.size === 0) return { drifted: [] };
2089
+ const drifted = [];
2090
+ for (const test of tests) {
2091
+ const unknown = [...new Set(test.routes)].filter((route) => !known.has(route)).sort();
2092
+ if (unknown.length > 0) drifted.push({ test, routes: unknown });
2093
+ }
2094
+ return { drifted };
2095
+ }
2018
2096
 
2019
2097
  // src/commands/run.ts
2020
2098
  var EXIT_OK = 0;
@@ -2244,6 +2322,34 @@ function reportUnclassified(options, impact) {
2244
2322
  );
2245
2323
  return true;
2246
2324
  }
2325
+ function declaredConfigRoutes(config) {
2326
+ const routes = config.routes ? Object.values(config.routes) : [];
2327
+ return [...new Set(routes.flat())].sort();
2328
+ }
2329
+ function printRouteDrift(drift, cwd) {
2330
+ if (drift.drifted.length === 0) return;
2331
+ console.error(
2332
+ "Route drift (test routes declared by no routes: mapping \u2014 contribute nothing to --impacted selection):"
2333
+ );
2334
+ for (const { test, routes } of drift.drifted) {
2335
+ console.error(` ${test.summary} (${path9.relative(cwd, test.path)}): ${routes.join(", ")}`);
2336
+ }
2337
+ console.error("Fix the route typo, or add the route to a routes: mapping in .blastproof/config.yaml.");
2338
+ }
2339
+ function printAuthoring(result, cwd) {
2340
+ if (result.findings.length === 0) return;
2341
+ console.error("\nAuthoring (a step enters a value but names none):");
2342
+ for (const { test, origin, index, step } of result.findings) {
2343
+ const where = origin === "setup" ? `setup step ${index}` : `step ${index}`;
2344
+ console.error(` ${test.summary} (${path9.relative(cwd, test.path)}) ${where}:`);
2345
+ console.error(` ${step}`);
2346
+ console.error(` \u2192 ${suggestValueClause(step)}`);
2347
+ }
2348
+ console.error(
2349
+ "The runner is forbidden from inventing values, but only by instruction: in practice the model supplies one anyway, the step passes, and the test verifies a value nobody wrote \u2014 a different one on each run."
2350
+ );
2351
+ console.error("Steps are inspected in English only; steps in other languages are not checked.");
2352
+ }
2247
2353
  function printDryRun(selected, config, cwd) {
2248
2354
  console.log(`
2249
2355
  Dry run: ${selected.length} test(s) selected, base_url=${config.base_url}`);
@@ -2326,6 +2432,16 @@ async function runCommand(options) {
2326
2432
  throw error;
2327
2433
  }
2328
2434
  }
2435
+ const drift = detectRouteDrift(parsed, declaredConfigRoutes(config));
2436
+ printRouteDrift(drift, options.cwd);
2437
+ const authoring = detectMissingValues(parsed);
2438
+ printAuthoring(authoring, options.cwd);
2439
+ if (options.failOnAuthoring && authoring.findings.length > 0) {
2440
+ console.error(
2441
+ `error: --fail-on-authoring: ${authoring.findings.length} step(s) enter a value but name none.`
2442
+ );
2443
+ return EXIT_FAILED;
2444
+ }
2329
2445
  const selection = impact ? selectImpactedTests(parsed, impact.affectedRoutes, options) : {
2330
2446
  selected: parsed.filter((test) => matchesFilters(test, options)),
2331
2447
  unroutedSkipped: [],
@@ -2720,6 +2836,7 @@ async function testCommand(options) {
2720
2836
  junit: options.junit,
2721
2837
  html: options.html,
2722
2838
  failOnUnmapped: options.failOnUnmapped,
2839
+ failOnAuthoring: options.failOnAuthoring,
2723
2840
  budget
2724
2841
  });
2725
2842
  if (runCode === EXIT_USAGE) return EXIT_USAGE;
@@ -2776,7 +2893,7 @@ function parsePositiveNumber(flag) {
2776
2893
  };
2777
2894
  }
2778
2895
  var program = new Command();
2779
- program.name("blastproof").description("Open-source AI testing agent: plain-English YAML tests executed agentically on a real browser.").version("0.11.0");
2896
+ program.name("blastproof").description("Open-source AI testing agent: plain-English YAML tests executed agentically on a real browser.").version("0.12.1");
2780
2897
  program.command("init").description("Scaffold .blastproof/ (config, tests, sample tests) in the current directory").action(async () => {
2781
2898
  try {
2782
2899
  const result = await initProject(process.cwd());
@@ -2791,7 +2908,7 @@ program.command("run").description("Discover and run all tests under .blastproof
2791
2908
  "--min-score <n>",
2792
2909
  "require a weighted score of at least n (0-100); replaces the all-must-pass rule",
2793
2910
  parseMinScore
2794
- ).option("--junit [path]", "write a JUnit XML report (default: .blastproof/reports/<session>/junit.xml)").option("--html [path]", "write a self-contained HTML report (default: .blastproof/reports/<session>/report.html)").option("--fail-on-unmapped", "fail when a changed file matches no routes: or ignore: glob").option(
2911
+ ).option("--junit [path]", "write a JUnit XML report (default: .blastproof/reports/<session>/junit.xml)").option("--html [path]", "write a self-contained HTML report (default: .blastproof/reports/<session>/report.html)").option("--fail-on-unmapped", "fail when a changed file matches no routes: or ignore: glob").option("--fail-on-authoring", "fail when a step enters a value but names none (warns by default)").option(
2795
2912
  "--concurrency <n>",
2796
2913
  "run this many tests at once (overrides config; default 1 \u2014 see the README on when this is safe)",
2797
2914
  parsePositiveInt("--concurrency")
@@ -2823,6 +2940,7 @@ program.command("run").description("Discover and run all tests under .blastproof
2823
2940
  junit: options.junit,
2824
2941
  html: options.html,
2825
2942
  failOnUnmapped: options.failOnUnmapped,
2943
+ failOnAuthoring: options.failOnAuthoring,
2826
2944
  concurrency: options.concurrency,
2827
2945
  maxLlmCalls: options.maxLlmCalls,
2828
2946
  maxTokens: options.maxTokens,
@@ -2873,7 +2991,7 @@ program.command("test").description("Full PR pipeline: run the tests covering th
2873
2991
  "--min-score <n>",
2874
2992
  "require a weighted score of at least n (0-100); replaces the all-must-pass rule",
2875
2993
  parseMinScore
2876
- ).option("--junit [path]", "write a JUnit XML report (default: .blastproof/reports/<session>/junit.xml)").option("--html [path]", "write a self-contained HTML report (default: .blastproof/reports/<session>/report.html)").option("--write", "persist generated drafts under .blastproof/tests/ instead of previewing them").option("--fail-on-unmapped", "fail when a changed file matches no routes: or ignore: glob").option(
2994
+ ).option("--junit [path]", "write a JUnit XML report (default: .blastproof/reports/<session>/junit.xml)").option("--html [path]", "write a self-contained HTML report (default: .blastproof/reports/<session>/report.html)").option("--write", "persist generated drafts under .blastproof/tests/ instead of previewing them").option("--fail-on-unmapped", "fail when a changed file matches no routes: or ignore: glob").option("--fail-on-authoring", "fail when a step enters a value but names none (warns by default)").option(
2877
2995
  "--max-llm-calls <n>",
2878
2996
  "stop the pipeline after this many model calls, shared by both phases (overrides config)",
2879
2997
  parsePositiveInt("--max-llm-calls")
@@ -2897,6 +3015,7 @@ program.command("test").description("Full PR pipeline: run the tests covering th
2897
3015
  html: options.html,
2898
3016
  write: options.write,
2899
3017
  failOnUnmapped: options.failOnUnmapped,
3018
+ failOnAuthoring: options.failOnAuthoring,
2900
3019
  maxLlmCalls: options.maxLlmCalls,
2901
3020
  maxTokens: options.maxTokens,
2902
3021
  maxDuration: options.maxDuration