blastproof 0.8.0 → 0.10.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -42,6 +42,8 @@ Before `run`, `plan` or `test` do anything, they check what they are about to sp
42
42
 
43
43
  **Your markup must be accessible — a hard requirement.** Elements are found by role, label or visible text from the accessibility tree. That is what removes selectors and survives redesigns; the cost is that an interface the accessibility tree cannot describe cannot be driven at all, and there is deliberately no CSS or XPath fallback. Icon-only buttons without accessible names, `div`-based controls and ARIA-less dropdowns simply cannot be targeted. Run an accessibility checker first — the result predicts how well this will work better than anything else.
44
44
 
45
+ **Windows is untested.** Development and CI run on Linux and macOS. Nothing is known to be broken and reports are welcome (#8).
46
+
45
47
  **Not supported yet:** `iframe` content (so hosted payment widgets like Stripe Elements are invisible — an embedded checkout cannot be driven end to end), hover, scroll-to, drag and drop, file upload, multiple tabs, native `alert`/`confirm` dialogs. Page snapshots are capped at 200 lines by default, so very dense pages are truncated — raise it with `browser.max_snapshot_lines` if your pages need more; truncation is always marked in the snapshot so the model is never misled into thinking it saw the whole page.
46
48
 
47
49
  **Point it at disposable data.** Within a step, an action that commits — a click, or pressing Enter — is never performed twice: the runner refuses the repeat and tells the agent it already did that. This closes the case that used to produce duplicate records, where a submit answered by a redirect came back to a reset form and the agent, seeing no evidence of its own work, submitted again. It is not a guarantee of zero duplicate writes: an agent that reaches the same effect by a genuinely different route — another control with the same effect — is not caught. Use a seeded database, a staging environment you can reset, or a throwaway account; do not gate on a run against production data.
@@ -70,6 +72,26 @@ steps:
70
72
 
71
73
  `priority` is P0–P2 (default P1). `tags`, `setup` steps and `auth` are optional — `auth: false` runs the test signed out, which a login test needs. `routes` declares the URLs a test covers, which is what `--impacted` selects on; write route strings consistently, since they compare by exact equality (`/cart` ≠ `/cart/`).
72
74
 
75
+ ### Say what each step should produce
76
+
77
+ This is the one rule that decides whether a suite works. An outside evaluation took the **same application, same suite, same version from Score 64 to Score 100 by rewriting two steps** — nothing else changed:
78
+
79
+ ```yaml
80
+ # Fragile: a bare action. Nothing says what should be true afterwards.
81
+ - submit the add-task form
82
+ - verify the new task "Fix the flaky test" appears in the task list
83
+
84
+ # Robust: the step carries its own outcome.
85
+ - submit the add-task form and verify the task "Fix the flaky test" appears
86
+ in the task list with priority High and status Open
87
+ ```
88
+
89
+ The bare version fails on a shape that is everywhere: the form POSTs, the server redirects back to the same page, and the form comes back empty. The agent is asked whether "submit the add-task form" happened, and is looking at a page that is indistinguishable from one where nothing did. A step that names the outcome gives it something to check that survives the action.
90
+
91
+ Write steps that end in an observable result — text on the page, a count, a state change — and this class of failure does not arise.
92
+
93
+ **Inline error messages should be plain visible text.** `role="alert"` is read correctly from the accessibility tree and needs no special handling, but note that an alert your page has cleared shows up as an empty element: if a verdict says an alert exists whose content is missing, the message was emptied, not hidden.
94
+
73
95
  ## Commands
74
96
 
75
97
  ```bash
@@ -113,7 +135,7 @@ jobs:
113
135
 
114
136
  - run: npm start & # however your app boots
115
137
 
116
- - uses: hamc/blastproof@v0.8.0
138
+ - uses: hamc/blastproof@v0.10.0
117
139
  with:
118
140
  version: '0.6.0' # pin both when this gates merges
119
141
  api-key: ${{ secrets.ANTHROPIC_API_KEY }}
@@ -167,6 +189,22 @@ blastproof plan --base main --dry-run # affected routes no test
167
189
 
168
190
  They report affected routes, files nobody has classified, and affected routes no test covers — a coverage-gap report with an exit code, useful even on a repo whose suite is Playwright or Cypress.
169
191
 
192
+ ## Running tests at once
193
+
194
+ Tests run one at a time by default. Raise it when your tests can stand it:
195
+
196
+ ```yaml
197
+ concurrency: 4
198
+ ```
199
+
200
+ or `blastproof run --concurrency 4` for a single invocation. On this repository's own suite that takes a run from 156s to 68s — **2.3× faster, for the same 81 model calls.** Parallelism buys wall-clock, not spend.
201
+
202
+ **The default is 1 on purpose, and raising it is your call to make.** Other test runners default to parallel because their tests are isolated by construction — separate processes, separate fixtures. These are journeys driven against **one running application**, so two tests can see each other's data. A suite is safe to parallelise when its tests do not write state that another test reads.
203
+
204
+ The test in this repository's own suite that could not run beside itself is a good shape to recognise: it adds a note and then asserts *"one note on file"*. It writes shared server state, and it asserts on a global count. Either alone is a warning; together they mean the test's verdict depends on nothing else touching the application at that moment.
205
+
206
+ Two practical notes. Four concurrent journeys are four times the traffic against whatever you pointed at — usually fine for a development instance, worth knowing for a shared one. And with several model calls in flight, a `budget:` limit can overshoot by up to the concurrency rather than by a single call, since the calls already sent are allowed to finish.
207
+
170
208
  ## Bounding a run
171
209
 
172
210
  Nothing stops a run by default. `budget:` puts a ceiling on `run`, `plan` and `test` alike — every model call any of them makes is counted:
@@ -182,7 +220,18 @@ Each limit is optional; with none set, nothing binds. They count **calls and tok
182
220
 
183
221
  Exhausting a budget **stops the run; it does not fail a test.** Running out of quota says nothing about the code under review. Unreached tests are reported as `not run`, a third state excluded from the score entirely, and the process exits 1 unconditionally — `--min-score` cannot rescue it, because the tests that finished are whichever ran first, not a representative sample.
184
222
 
185
- `--dry-run` reports the ceiling before you spend anything.
223
+ **Every run reports what it spent**, so you can size a limit from experience rather than guesswork:
224
+
225
+ ```
226
+ Spent: 82 model call(s), 115407 token(s)
227
+ Score: 100
228
+ ```
229
+
230
+ The figures also land in the JUnit report as `llm_calls` and `llm_tokens`, beside `score`, so a pipeline can trend cost without scraping output. A run stopped by its own budget reports the spend too — that is the case where the number is least guessable. Where a provider reports no token usage, the line says so rather than showing zero.
231
+
232
+ `--dry-run` reports the ceiling before you spend anything. Read it as a maximum and nothing more: for this repository's own suite it says 735 calls where a real run spends 82. Size a budget from what your runs actually report, not from the ceiling.
233
+
234
+ For an order of magnitude, this repository's own suite — 7 tests, 31 steps, an authenticated demo shop, `anthropic/claude-haiku-4.5` — spends **about 82 model calls and 115k tokens**, taking 156s serially or 68s at `--concurrency 4`. That is a number you can reproduce (`node examples/demo-app/serve.mjs 4173` and `blastproof run`), not a forecast for your suite: cost scales with steps, page density and how often the agent has to retry. Run yours once and read the `Spent:` line.
186
235
 
187
236
  ## Testing behind a login
188
237
 
package/dist/cli.js CHANGED
@@ -210,6 +210,7 @@ var RunBudget = class {
210
210
  startedAt;
211
211
  calls = 0;
212
212
  tokens = 0;
213
+ callsWithUsage = 0;
213
214
  constructor(options = {}) {
214
215
  this.maxCalls = options.maxCalls;
215
216
  this.maxTokens = options.maxTokens;
@@ -248,7 +249,25 @@ var RunBudget = class {
248
249
  /** Records what a completed model call spent. */
249
250
  record(usage) {
250
251
  this.calls += 1;
251
- if (usage?.totalTokens !== void 0) this.tokens += usage.totalTokens;
252
+ if (usage?.totalTokens !== void 0) {
253
+ this.tokens += usage.totalTokens;
254
+ this.callsWithUsage += 1;
255
+ }
256
+ }
257
+ /**
258
+ * What this budget has been spent on, for every surface that reports it
259
+ * (design report-what-it-spent, D1). One method rather than four getters read
260
+ * separately: the console, the JUnit report and the HTML report take the same
261
+ * numbers from the same object, so they cannot disagree about what a run cost.
262
+ */
263
+ spend() {
264
+ return {
265
+ calls: this.calls,
266
+ maxCalls: this.maxCalls,
267
+ tokens: this.tokens,
268
+ callsWithUsage: this.callsWithUsage,
269
+ maxTokens: this.maxTokens
270
+ };
252
271
  }
253
272
  };
254
273
  function estimateMaxModelCalls(tests, maxIterationsPerStep, maxRetriesPerStep) {
@@ -417,7 +436,8 @@ async function performAction(page, action, ctx) {
417
436
  const url = new URL(value, ctx.baseUrl);
418
437
  assertAllowedOrigin(url, ctx);
419
438
  await page.goto(url.toString(), { timeout: ctx.resolveTimeoutMs ?? 3e4 });
420
- return `ok: navigated to ${url.toString()}`;
439
+ const landed = page.url();
440
+ return landed === url.toString() ? `ok: navigated to ${url.toString()}` : `ok: navigated to ${url.toString()}, which redirected to ${landed}`;
421
441
  }
422
442
  case "click": {
423
443
  const target = requireTarget(action);
@@ -623,11 +643,21 @@ async function executeTest(page, test, options) {
623
643
  }
624
644
  if (action.action === "assert") {
625
645
  const expectation = action.expectation ?? action.reasoning;
626
- let judgment = await brain.judge(mask(step), mask(expectation), mask(snap));
646
+ let judgment = await brain.judge(
647
+ mask(step),
648
+ mask(expectation),
649
+ mask(snap),
650
+ recovery.stepHistory()
651
+ );
627
652
  if (!judgment.pass) {
628
653
  await waitForSettled(page);
629
654
  const freshSnap = await takeSnapshot(page);
630
- judgment = await brain.judge(mask(step), mask(expectation), mask(freshSnap));
655
+ judgment = await brain.judge(
656
+ mask(step),
657
+ mask(expectation),
658
+ mask(freshSnap),
659
+ recovery.stepHistory()
660
+ );
631
661
  }
632
662
  const result = judgment.pass ? `ok: assertion passed: ${judgment.reason}` : `assertion failed: ${judgment.reason}`;
633
663
  emitAction(index, action, result);
@@ -947,6 +977,13 @@ var configSchema = z.object({
947
977
  allowed_origins: z.array(z.string().url()).optional(),
948
978
  auth: authSchema.optional(),
949
979
  max_retries_per_step: z.number().int().min(1).default(3),
980
+ /**
981
+ * How many tests may run at once (design tests-in-parallel, D1). Defaults to
982
+ * 1 — tests are journeys driven against one running application, and whether
983
+ * two of them can run at the same time is a property of that application and
984
+ * those tests, not of the runner. Opted into by the person who knows.
985
+ */
986
+ concurrency: z.number().int().min(1, "concurrency must be at least 1").default(1),
950
987
  budget: budgetSchema.optional()
951
988
  });
952
989
  var ENV_OVERRIDES = {
@@ -1177,13 +1214,24 @@ A value sitting in a control that was just typed into \u2014 an open dialog's te
1177
1214
 
1178
1215
  A step that names an ACTION (submit, click, create, add, ...) is satisfied by evidence the action took effect, not by the action's own control still being on the page. A successful action ordinarily replaces or moves past exactly the form, button or field the step names, so that control's absence is normal evidence of success, not evidence the step is unverifiable \u2014 do not fail such a step only because you can no longer see the thing it names. Fail it instead when the snapshot shows the action did NOT take effect: an error message, a validation warning, or the very same pre-action page still in front of you with nothing changed. A different page, a new state, or the result the action was meant to produce counts as evidence it worked.
1179
1216
 
1217
+ You may also be shown the actions already performed in this step, with their results. That record tells you what was ATTEMPTED and what it produced \u2014 for instance that a navigation was performed and which URL the server ultimately served, or that a form was submitted. Use it to avoid concluding that something never happened when the page simply cannot show it any more: a navigation the server redirected does not leave the browser at the path that was requested, and that is what success looks like, not failure.
1218
+
1219
+ The record is not evidence that the step's outcome holds. An action reported as \`ok\` establishes that it ran and what it returned; whether the thing the step describes is now TRUE is still decided by the snapshot alone. Never pass a step because the record shows an action succeeded while the snapshot does not show the outcome.
1220
+
1180
1221
  Be strict about what the step asks, not about withholding a pass you can plainly see is earned. Answer with pass=true/false and a one-sentence reason.
1181
1222
 
1182
1223
  \`***\` marks a secret deliberately withheld from you \u2014 a password, token or key. Seeing it is expected. A field holding \`***\` is filled, not empty, so do not fail a step on the grounds that a value was redacted. This applies only to the redaction itself: everything else the step asks for must still be visibly satisfied by the snapshot, and a step you genuinely cannot check against what you were shown still fails.`;
1183
1224
  }
1184
- function assertUserPrompt(step, expectation, snapshot) {
1185
- return [
1186
- `Step under test: ${step}`,
1225
+ function assertUserPrompt(step, expectation, snapshot, stepHistory) {
1226
+ const parts = [`Step under test: ${step}`];
1227
+ if (stepHistory && stepHistory.length > 0) {
1228
+ parts.push(
1229
+ "",
1230
+ "Actions already performed in this step, with their results (what was DONE \u2014 not evidence of what is now true):",
1231
+ ...stepHistory.map((entry, i) => `${i + 1}. ${entry.action} -> ${entry.result}`)
1232
+ );
1233
+ }
1234
+ parts.push(
1187
1235
  "",
1188
1236
  `Model's expectation (the claim offered in support of the step, not the question itself): ${expectation}`,
1189
1237
  "",
@@ -1191,7 +1239,8 @@ function assertUserPrompt(step, expectation, snapshot) {
1191
1239
  snapshot,
1192
1240
  "",
1193
1241
  "Does the snapshot establish that the step's own outcome holds?"
1194
- ].join("\n");
1242
+ );
1243
+ return parts.join("\n");
1195
1244
  }
1196
1245
  function plannerSystemPrompt() {
1197
1246
  return `You are a QA engineer writing one end-to-end test for a web page, in plain English.
@@ -1280,12 +1329,12 @@ function createBrain(model, generate = generateObject, budget) {
1280
1329
  }
1281
1330
  return parsed.data;
1282
1331
  },
1283
- async judge(step, expectation, snapshot) {
1332
+ async judge(step, expectation, snapshot, stepHistory) {
1284
1333
  const result = await countedGenerate(generate, budget, {
1285
1334
  model,
1286
1335
  schema: assertJudgmentSchema,
1287
1336
  system: assertSystemPrompt(),
1288
- prompt: assertUserPrompt(step, expectation, snapshot)
1337
+ prompt: assertUserPrompt(step, expectation, snapshot, stepHistory)
1289
1338
  });
1290
1339
  const parsed = assertJudgmentSchema.safeParse(result.object);
1291
1340
  if (!parsed.success) {
@@ -1626,6 +1675,43 @@ function printPreflightFailures(failures) {
1626
1675
  for (const failure of failures) console.error(` - ${failure}`);
1627
1676
  }
1628
1677
 
1678
+ // src/report/score.ts
1679
+ var WEIGHTS = { P0: 3, P1: 2, P2: 1 };
1680
+ var DEFAULT_WEIGHT = WEIGHTS.P1;
1681
+ function computeScore(results) {
1682
+ let total = 0;
1683
+ let passed = 0;
1684
+ for (const result of results) {
1685
+ if (result.status === "not-run") continue;
1686
+ const weight = WEIGHTS[result.priority] ?? DEFAULT_WEIGHT;
1687
+ total += weight;
1688
+ if (result.status === "passed") passed += weight;
1689
+ }
1690
+ if (total === 0) return 100;
1691
+ return Math.round(100 * passed / total);
1692
+ }
1693
+ function formatScoreLine(score, results, threshold) {
1694
+ if (results.length === 0) {
1695
+ const base = "Score: 100 (no tests executed)";
1696
+ return threshold === void 0 ? base : `${base} \u2014 min-score ${threshold}: pass`;
1697
+ }
1698
+ if (threshold === void 0) return `Score: ${score}`;
1699
+ return score >= threshold ? `Score: ${score} \u2014 min-score ${threshold}: pass` : `Score: ${score} \u2014 min-score ${threshold}: FAIL (below threshold)`;
1700
+ }
1701
+ function formatIncompleteLine(score, reason) {
1702
+ return `Run incomplete: ${reason}
1703
+ Score over executed tests: ${score} (not a verdict \u2014 exit code 1 regardless of --min-score)`;
1704
+ }
1705
+ function formatSpendLine(spend) {
1706
+ const calls = spend.maxCalls === void 0 ? `${spend.calls} model call(s)` : `${spend.calls} of ${spend.maxCalls} model call(s)`;
1707
+ if (spend.callsWithUsage === 0) {
1708
+ return `Spent: ${calls}; token usage not reported by the provider`;
1709
+ }
1710
+ const tokens = spend.maxTokens === void 0 ? `${spend.tokens} token(s)` : `${spend.tokens} of ${spend.maxTokens} token(s)`;
1711
+ const coverage = spend.callsWithUsage < spend.calls ? ` (tokens reported by ${spend.callsWithUsage} of ${spend.calls} call(s))` : "";
1712
+ return `Spent: ${calls}, ${tokens}${coverage}`;
1713
+ }
1714
+
1629
1715
  // src/commands/run.ts
1630
1716
  import path9 from "path";
1631
1717
 
@@ -1770,7 +1856,7 @@ async function renderHtml(results, skipped, meta) {
1770
1856
  <body>
1771
1857
  <main>
1772
1858
  <h1>blastproof report</h1>
1773
- <p class="sub">${escapeHtml(generatedAt)} \xB7 ${seconds(meta.durationMs)}</p>
1859
+ <p class="sub">${escapeHtml(generatedAt)} \xB7 ${seconds(meta.durationMs)}${meta.spend ? ` \xB7 ${escapeHtml(formatSpendLine(meta.spend))}` : ""}</p>
1774
1860
 
1775
1861
  ${banner} <section class="score">
1776
1862
  <b>${meta.score}</b>
@@ -1824,6 +1910,11 @@ function renderJUnit(results, skipped, meta) {
1824
1910
  `<testsuite name="blastproof" tests="${results.length + skipped.length}" failures="${failures}" skipped="${skipped.length + notRun.length}" time="${seconds2(meta.durationMs)}">`,
1825
1911
  " <properties>",
1826
1912
  ` <property name="score" value="${meta.score}"/>`,
1913
+ ...meta.spend ? [` <property name="llm_calls" value="${meta.spend.calls}"/>`] : [],
1914
+ // Omitted rather than emitted as zero when no call reported usage: a
1915
+ // property carrying 0 would be read by a pipeline as "this run spent no
1916
+ // tokens", which is a different claim from "the provider did not say".
1917
+ ...meta.spend && meta.spend.callsWithUsage > 0 ? [` <property name="llm_tokens" value="${meta.spend.tokens}"/>`] : [],
1827
1918
  ...meta.incomplete !== void 0 ? [
1828
1919
  ' <property name="incomplete" value="true"/>',
1829
1920
  ` <property name="incomplete_reason" value="${escapeXml(meta.incomplete)}"/>`
@@ -1867,32 +1958,26 @@ async function writeJUnit(file, xml) {
1867
1958
  return file;
1868
1959
  }
1869
1960
 
1870
- // src/report/score.ts
1871
- var WEIGHTS = { P0: 3, P1: 2, P2: 1 };
1872
- var DEFAULT_WEIGHT = WEIGHTS.P1;
1873
- function computeScore(results) {
1874
- let total = 0;
1875
- let passed = 0;
1876
- for (const result of results) {
1877
- if (result.status === "not-run") continue;
1878
- const weight = WEIGHTS[result.priority] ?? DEFAULT_WEIGHT;
1879
- total += weight;
1880
- if (result.status === "passed") passed += weight;
1881
- }
1882
- if (total === 0) return 100;
1883
- return Math.round(100 * passed / total);
1884
- }
1885
- function formatScoreLine(score, results, threshold) {
1886
- if (results.length === 0) {
1887
- const base = "Score: 100 (no tests executed)";
1888
- return threshold === void 0 ? base : `${base} \u2014 min-score ${threshold}: pass`;
1961
+ // src/runner/pool.ts
1962
+ async function runWithConcurrency(items, concurrency, run) {
1963
+ if (concurrency < 1) throw new RangeError(`concurrency must be at least 1, got ${concurrency}`);
1964
+ const results = new Array(items.length);
1965
+ let next = 0;
1966
+ const worker = async () => {
1967
+ while (true) {
1968
+ const index = next++;
1969
+ if (index >= items.length) return;
1970
+ results[index] = await run(items[index], index);
1971
+ }
1972
+ };
1973
+ const workers = [];
1974
+ for (let i = 0; i < Math.min(concurrency, items.length); i++) {
1975
+ workers.push(worker());
1889
1976
  }
1890
- if (threshold === void 0) return `Score: ${score}`;
1891
- return score >= threshold ? `Score: ${score} \u2014 min-score ${threshold}: pass` : `Score: ${score} \u2014 min-score ${threshold}: FAIL (below threshold)`;
1892
- }
1893
- function formatIncompleteLine(score, reason) {
1894
- return `Run incomplete: ${reason}
1895
- Score over executed tests: ${score} (not a verdict \u2014 exit code 1 regardless of --min-score)`;
1977
+ const settled = await Promise.allSettled(workers);
1978
+ const failure = settled.find((outcome) => outcome.status === "rejected");
1979
+ if (failure && failure.status === "rejected") throw failure.reason;
1980
+ return results;
1896
1981
  }
1897
1982
 
1898
1983
  // src/runner/selection.ts
@@ -1983,19 +2068,19 @@ function notRunResults(tests, reason) {
1983
2068
  durationMs: 0
1984
2069
  }));
1985
2070
  }
1986
- function printEvent(event) {
2071
+ function printEvent(event, write = console.log) {
1987
2072
  switch (event.type) {
1988
2073
  case "step-start":
1989
- console.log(` ${event.setup ? "(setup) " : ""}step ${event.index + 1}/${event.total}: ${event.step}`);
2074
+ write(` ${event.setup ? "(setup) " : ""}step ${event.index + 1}/${event.total}: ${event.step}`);
1990
2075
  break;
1991
2076
  case "action": {
1992
2077
  const { action, result } = event;
1993
- console.log(` -> ${describeAction(action)} :: ${result}`);
2078
+ write(` -> ${describeAction(action)} :: ${result}`);
1994
2079
  break;
1995
2080
  }
1996
2081
  case "step-end":
1997
2082
  if (event.status === "failed") {
1998
- console.log(` X step failed: ${event.reason ?? "unknown reason"}`);
2083
+ write(` X step failed: ${event.reason ?? "unknown reason"}`);
1999
2084
  }
2000
2085
  break;
2001
2086
  }
@@ -2040,7 +2125,7 @@ ${notRun.length} test(s) not run (run stopped by its budget or deadline):`);
2040
2125
  for (const r of notRun) console.log(` - ${r.summary} (${r.file})`);
2041
2126
  }
2042
2127
  }
2043
- async function runOne(browser, test, config, sessionDir, session, runMask, budget) {
2128
+ async function runOne(browser, test, config, sessionDir, session, runMask, budget, write) {
2044
2129
  const brain = createBrain(createModel(config.llm).model, void 0, budget);
2045
2130
  let resolved;
2046
2131
  const mask = runMask;
@@ -2075,7 +2160,7 @@ async function runOne(browser, test, config, sessionDir, session, runMask, budge
2075
2160
  timeoutMs: config.browser.timeout_ms,
2076
2161
  maxSnapshotLines: config.browser.max_snapshot_lines,
2077
2162
  mask: (text) => mask.mask(text),
2078
- onEvent: printEvent
2163
+ onEvent: (event) => printEvent(event, write)
2079
2164
  });
2080
2165
  } finally {
2081
2166
  await context.close();
@@ -2106,8 +2191,9 @@ function printImpactReport(impact, selection, cwd) {
2106
2191
  }
2107
2192
  console.log("---------------------------------------------------------------");
2108
2193
  }
2109
- async function finalize(results, skipped, options, sessionDir, durationMs, impact, incomplete) {
2194
+ async function finalize(results, skipped, options, sessionDir, durationMs, impact, incomplete, spend) {
2110
2195
  if (results.length > 0) printSummary(results);
2196
+ if (spend) console.log(formatSpendLine(spend));
2111
2197
  const score = computeScore(results);
2112
2198
  console.log(
2113
2199
  incomplete ? formatIncompleteLine(score, incomplete.message) : formatScoreLine(score, results, options.minScore)
@@ -2118,7 +2204,8 @@ async function finalize(results, skipped, options, sessionDir, durationMs, impac
2118
2204
  score,
2119
2205
  durationMs,
2120
2206
  cwd: options.cwd,
2121
- incomplete: incomplete?.message
2207
+ incomplete: incomplete?.message,
2208
+ spend
2122
2209
  });
2123
2210
  await writeJUnit(target, xml);
2124
2211
  console.log(`JUnit report: ${path9.relative(options.cwd, target)}`);
@@ -2130,7 +2217,8 @@ async function finalize(results, skipped, options, sessionDir, durationMs, impac
2130
2217
  durationMs,
2131
2218
  minScore: options.minScore,
2132
2219
  cwd: options.cwd,
2133
- incomplete: incomplete?.message
2220
+ incomplete: incomplete?.message,
2221
+ spend
2134
2222
  });
2135
2223
  await writeHtml(target, html);
2136
2224
  console.log(`HTML report: ${path9.relative(options.cwd, target)}`);
@@ -2317,21 +2405,38 @@ ${results.length} test file(s) could not be parsed:`);
2317
2405
  if (incomplete) {
2318
2406
  results.push(...notRunResults(selected, incomplete.message));
2319
2407
  } else {
2320
- for (let i = 0; i < selected.length; i++) {
2321
- const test = selected[i];
2322
- console.log(`
2323
- > ${test.summary} [${test.priority}] (${path9.relative(options.cwd, test.path)})`);
2408
+ const concurrency = options.concurrency ?? config.concurrency;
2409
+ const streaming = concurrency === 1;
2410
+ let stoppedBy;
2411
+ const outcomes = await runWithConcurrency(selected, concurrency, async (test, index) => {
2412
+ if (stoppedBy) return { index, stopped: stoppedBy };
2413
+ const lines = [];
2414
+ const header = `
2415
+ > ${test.summary} [${test.priority}] (${path9.relative(options.cwd, test.path)})`;
2416
+ const write = streaming ? console.log : (line) => lines.push(line);
2417
+ if (streaming) console.log(header);
2418
+ else lines.push(header);
2324
2419
  try {
2325
2420
  budget.check();
2326
- results.push(await runOne(browser, test, config, sessionDir, session, runMask, budget));
2421
+ const result = await runOne(browser, test, config, sessionDir, session, runMask, budget, write);
2422
+ if (!streaming) for (const line of lines) console.log(line);
2423
+ return { index, result };
2327
2424
  } catch (error) {
2328
2425
  if (error instanceof BudgetExhaustedError) {
2329
- incomplete = error;
2330
- results.push(...notRunResults(selected.slice(i), error.message));
2331
- break;
2426
+ stoppedBy ??= error;
2427
+ if (!streaming) for (const line of lines) console.log(line);
2428
+ return { index, stopped: error };
2332
2429
  }
2333
2430
  throw error;
2334
2431
  }
2432
+ });
2433
+ for (const outcome of outcomes) {
2434
+ if (outcome.result) {
2435
+ results.push(outcome.result);
2436
+ continue;
2437
+ }
2438
+ incomplete ??= outcome.stopped;
2439
+ results.push(...notRunResults([selected[outcome.index]], outcome.stopped.message));
2335
2440
  }
2336
2441
  }
2337
2442
  } finally {
@@ -2344,7 +2449,13 @@ ${results.length} test file(s) could not be parsed:`);
2344
2449
  sessionDir,
2345
2450
  Date.now() - startedAt,
2346
2451
  impact,
2347
- incomplete
2452
+ incomplete,
2453
+ // The command that constructed the budget is the one that reports it
2454
+ // (design report-what-it-spent, D2). `test` hands one budget to both of its
2455
+ // phases deliberately, so reporting at the point of use rather than the
2456
+ // point of ownership would print the running total twice for a single
2457
+ // allowance, the second line silently including the first.
2458
+ options.budget === void 0 ? budget.spend() : void 0
2348
2459
  );
2349
2460
  }
2350
2461
 
@@ -2574,6 +2685,7 @@ ${renderTestYaml(draft, { route, base })}`);
2574
2685
  for (const route of notAttempted) console.log(` ${route}`);
2575
2686
  }
2576
2687
  }
2688
+ if (options.budget === void 0) console.log(formatSpendLine(budget.spend()));
2577
2689
  console.log("---------------------------------------------------------------");
2578
2690
  return incomplete || failed.length > 0 ? EXIT_FAILED : EXIT_OK;
2579
2691
  }
@@ -2618,6 +2730,8 @@ async function testCommand(options) {
2618
2730
  budget
2619
2731
  });
2620
2732
  if (planCode === EXIT_USAGE) return EXIT_USAGE;
2733
+ console.log(`
2734
+ ${formatSpendLine(budget.spend())}`);
2621
2735
  console.log(
2622
2736
  "\nDrafts are not executed and do not affect the score \u2014 review them before they join the suite."
2623
2737
  );
@@ -2660,7 +2774,7 @@ function parsePositiveNumber(flag) {
2660
2774
  };
2661
2775
  }
2662
2776
  var program = new Command();
2663
- program.name("blastproof").description("Open-source AI testing agent: plain-English YAML tests executed agentically on a real browser.").version("0.8.0");
2777
+ program.name("blastproof").description("Open-source AI testing agent: plain-English YAML tests executed agentically on a real browser.").version("0.10.0");
2664
2778
  program.command("init").description("Scaffold .blastproof/ (config, tests, sample tests) in the current directory").action(async () => {
2665
2779
  try {
2666
2780
  const result = await initProject(process.cwd());
@@ -2676,6 +2790,10 @@ program.command("run").description("Discover and run all tests under .blastproof
2676
2790
  "require a weighted score of at least n (0-100); replaces the all-must-pass rule",
2677
2791
  parseMinScore
2678
2792
  ).option("--junit [path]", "write a JUnit XML report (default: .blastproof/reports/<session>/junit.xml)").option("--html [path]", "write a self-contained HTML report (default: .blastproof/reports/<session>/report.html)").option("--fail-on-unmapped", "fail when a changed file matches no routes: or ignore: glob").option(
2793
+ "--concurrency <n>",
2794
+ "run this many tests at once (overrides config; default 1 \u2014 see the README on when this is safe)",
2795
+ parsePositiveInt("--concurrency")
2796
+ ).option(
2679
2797
  "--max-llm-calls <n>",
2680
2798
  "stop the run after this many model calls, reported as incomplete (overrides config)",
2681
2799
  parsePositiveInt("--max-llm-calls")
@@ -2703,6 +2821,7 @@ program.command("run").description("Discover and run all tests under .blastproof
2703
2821
  junit: options.junit,
2704
2822
  html: options.html,
2705
2823
  failOnUnmapped: options.failOnUnmapped,
2824
+ concurrency: options.concurrency,
2706
2825
  maxLlmCalls: options.maxLlmCalls,
2707
2826
  maxTokens: options.maxTokens,
2708
2827
  maxDuration: options.maxDuration