blastproof 0.8.0 → 0.9.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -113,7 +113,7 @@ jobs:
113
113
 
114
114
  - run: npm start & # however your app boots
115
115
 
116
- - uses: hamc/blastproof@v0.8.0
116
+ - uses: hamc/blastproof@v0.9.0
117
117
  with:
118
118
  version: '0.6.0' # pin both when this gates merges
119
119
  api-key: ${{ secrets.ANTHROPIC_API_KEY }}
@@ -167,6 +167,22 @@ blastproof plan --base main --dry-run # affected routes no test
167
167
 
168
168
  They report affected routes, files nobody has classified, and affected routes no test covers — a coverage-gap report with an exit code, useful even on a repo whose suite is Playwright or Cypress.
169
169
 
170
+ ## Running tests at once
171
+
172
+ Tests run one at a time by default. Raise it when your tests can stand it:
173
+
174
+ ```yaml
175
+ concurrency: 4
176
+ ```
177
+
178
+ or `blastproof run --concurrency 4` for a single invocation. On this repository's own suite that takes a run from 156s to 68s — **2.3× faster, for the same 81 model calls.** Parallelism buys wall-clock, not spend.
179
+
180
+ **The default is 1 on purpose, and raising it is your call to make.** Other test runners default to parallel because their tests are isolated by construction — separate processes, separate fixtures. These are journeys driven against **one running application**, so two tests can see each other's data. A suite is safe to parallelise when its tests do not write state that another test reads.
181
+
182
+ The test in this repository's own suite that could not run beside itself is a good shape to recognise: it adds a note and then asserts *"one note on file"*. It writes shared server state, and it asserts on a global count. Either alone is a warning; together they mean the test's verdict depends on nothing else touching the application at that moment.
183
+
184
+ Two practical notes. Four concurrent journeys are four times the traffic against whatever you pointed at — usually fine for a development instance, worth knowing for a shared one. And with several model calls in flight, a `budget:` limit can overshoot by up to the concurrency rather than by a single call, since the calls already sent are allowed to finish.
185
+
170
186
  ## Bounding a run
171
187
 
172
188
  Nothing stops a run by default. `budget:` puts a ceiling on `run`, `plan` and `test` alike — every model call any of them makes is counted:
@@ -182,7 +198,16 @@ Each limit is optional; with none set, nothing binds. They count **calls and tok
182
198
 
183
199
  Exhausting a budget **stops the run; it does not fail a test.** Running out of quota says nothing about the code under review. Unreached tests are reported as `not run`, a third state excluded from the score entirely, and the process exits 1 unconditionally — `--min-score` cannot rescue it, because the tests that finished are whichever ran first, not a representative sample.
184
200
 
185
- `--dry-run` reports the ceiling before you spend anything.
201
+ **Every run reports what it spent**, so you can size a limit from experience rather than guesswork:
202
+
203
+ ```
204
+ Spent: 82 model call(s), 115407 token(s)
205
+ Score: 100
206
+ ```
207
+
208
+ The figures also land in the JUnit report as `llm_calls` and `llm_tokens`, beside `score`, so a pipeline can trend cost without scraping output. A run stopped by its own budget reports the spend too — that is the case where the number is least guessable. Where a provider reports no token usage, the line says so rather than showing zero.
209
+
210
+ `--dry-run` reports the ceiling before you spend anything. Read it as a maximum and nothing more: for this repository's own suite it says 735 calls where a real run spends 82. Size a budget from what your runs actually report, not from the ceiling.
186
211
 
187
212
  ## Testing behind a login
188
213
 
package/dist/cli.js CHANGED
@@ -210,6 +210,7 @@ var RunBudget = class {
210
210
  startedAt;
211
211
  calls = 0;
212
212
  tokens = 0;
213
+ callsWithUsage = 0;
213
214
  constructor(options = {}) {
214
215
  this.maxCalls = options.maxCalls;
215
216
  this.maxTokens = options.maxTokens;
@@ -248,7 +249,25 @@ var RunBudget = class {
248
249
  /** Records what a completed model call spent. */
249
250
  record(usage) {
250
251
  this.calls += 1;
251
- if (usage?.totalTokens !== void 0) this.tokens += usage.totalTokens;
252
+ if (usage?.totalTokens !== void 0) {
253
+ this.tokens += usage.totalTokens;
254
+ this.callsWithUsage += 1;
255
+ }
256
+ }
257
+ /**
258
+ * What this budget has been spent on, for every surface that reports it
259
+ * (design report-what-it-spent, D1). One method rather than four getters read
260
+ * separately: the console, the JUnit report and the HTML report take the same
261
+ * numbers from the same object, so they cannot disagree about what a run cost.
262
+ */
263
+ spend() {
264
+ return {
265
+ calls: this.calls,
266
+ maxCalls: this.maxCalls,
267
+ tokens: this.tokens,
268
+ callsWithUsage: this.callsWithUsage,
269
+ maxTokens: this.maxTokens
270
+ };
252
271
  }
253
272
  };
254
273
  function estimateMaxModelCalls(tests, maxIterationsPerStep, maxRetriesPerStep) {
@@ -947,6 +966,13 @@ var configSchema = z.object({
947
966
  allowed_origins: z.array(z.string().url()).optional(),
948
967
  auth: authSchema.optional(),
949
968
  max_retries_per_step: z.number().int().min(1).default(3),
969
+ /**
970
+ * How many tests may run at once (design tests-in-parallel, D1). Defaults to
971
+ * 1 — tests are journeys driven against one running application, and whether
972
+ * two of them can run at the same time is a property of that application and
973
+ * those tests, not of the runner. Opted into by the person who knows.
974
+ */
975
+ concurrency: z.number().int().min(1, "concurrency must be at least 1").default(1),
950
976
  budget: budgetSchema.optional()
951
977
  });
952
978
  var ENV_OVERRIDES = {
@@ -1626,6 +1652,43 @@ function printPreflightFailures(failures) {
1626
1652
  for (const failure of failures) console.error(` - ${failure}`);
1627
1653
  }
1628
1654
 
1655
+ // src/report/score.ts
1656
+ var WEIGHTS = { P0: 3, P1: 2, P2: 1 };
1657
+ var DEFAULT_WEIGHT = WEIGHTS.P1;
1658
+ function computeScore(results) {
1659
+ let total = 0;
1660
+ let passed = 0;
1661
+ for (const result of results) {
1662
+ if (result.status === "not-run") continue;
1663
+ const weight = WEIGHTS[result.priority] ?? DEFAULT_WEIGHT;
1664
+ total += weight;
1665
+ if (result.status === "passed") passed += weight;
1666
+ }
1667
+ if (total === 0) return 100;
1668
+ return Math.round(100 * passed / total);
1669
+ }
1670
+ function formatScoreLine(score, results, threshold) {
1671
+ if (results.length === 0) {
1672
+ const base = "Score: 100 (no tests executed)";
1673
+ return threshold === void 0 ? base : `${base} \u2014 min-score ${threshold}: pass`;
1674
+ }
1675
+ if (threshold === void 0) return `Score: ${score}`;
1676
+ return score >= threshold ? `Score: ${score} \u2014 min-score ${threshold}: pass` : `Score: ${score} \u2014 min-score ${threshold}: FAIL (below threshold)`;
1677
+ }
1678
+ function formatIncompleteLine(score, reason) {
1679
+ return `Run incomplete: ${reason}
1680
+ Score over executed tests: ${score} (not a verdict \u2014 exit code 1 regardless of --min-score)`;
1681
+ }
1682
+ function formatSpendLine(spend) {
1683
+ const calls = spend.maxCalls === void 0 ? `${spend.calls} model call(s)` : `${spend.calls} of ${spend.maxCalls} model call(s)`;
1684
+ if (spend.callsWithUsage === 0) {
1685
+ return `Spent: ${calls}; token usage not reported by the provider`;
1686
+ }
1687
+ const tokens = spend.maxTokens === void 0 ? `${spend.tokens} token(s)` : `${spend.tokens} of ${spend.maxTokens} token(s)`;
1688
+ const coverage = spend.callsWithUsage < spend.calls ? ` (tokens reported by ${spend.callsWithUsage} of ${spend.calls} call(s))` : "";
1689
+ return `Spent: ${calls}, ${tokens}${coverage}`;
1690
+ }
1691
+
1629
1692
  // src/commands/run.ts
1630
1693
  import path9 from "path";
1631
1694
 
@@ -1770,7 +1833,7 @@ async function renderHtml(results, skipped, meta) {
1770
1833
  <body>
1771
1834
  <main>
1772
1835
  <h1>blastproof report</h1>
1773
- <p class="sub">${escapeHtml(generatedAt)} \xB7 ${seconds(meta.durationMs)}</p>
1836
+ <p class="sub">${escapeHtml(generatedAt)} \xB7 ${seconds(meta.durationMs)}${meta.spend ? ` \xB7 ${escapeHtml(formatSpendLine(meta.spend))}` : ""}</p>
1774
1837
 
1775
1838
  ${banner} <section class="score">
1776
1839
  <b>${meta.score}</b>
@@ -1824,6 +1887,11 @@ function renderJUnit(results, skipped, meta) {
1824
1887
  `<testsuite name="blastproof" tests="${results.length + skipped.length}" failures="${failures}" skipped="${skipped.length + notRun.length}" time="${seconds2(meta.durationMs)}">`,
1825
1888
  " <properties>",
1826
1889
  ` <property name="score" value="${meta.score}"/>`,
1890
+ ...meta.spend ? [` <property name="llm_calls" value="${meta.spend.calls}"/>`] : [],
1891
+ // Omitted rather than emitted as zero when no call reported usage: a
1892
+ // property carrying 0 would be read by a pipeline as "this run spent no
1893
+ // tokens", which is a different claim from "the provider did not say".
1894
+ ...meta.spend && meta.spend.callsWithUsage > 0 ? [` <property name="llm_tokens" value="${meta.spend.tokens}"/>`] : [],
1827
1895
  ...meta.incomplete !== void 0 ? [
1828
1896
  ' <property name="incomplete" value="true"/>',
1829
1897
  ` <property name="incomplete_reason" value="${escapeXml(meta.incomplete)}"/>`
@@ -1867,32 +1935,26 @@ async function writeJUnit(file, xml) {
1867
1935
  return file;
1868
1936
  }
1869
1937
 
1870
- // src/report/score.ts
1871
- var WEIGHTS = { P0: 3, P1: 2, P2: 1 };
1872
- var DEFAULT_WEIGHT = WEIGHTS.P1;
1873
- function computeScore(results) {
1874
- let total = 0;
1875
- let passed = 0;
1876
- for (const result of results) {
1877
- if (result.status === "not-run") continue;
1878
- const weight = WEIGHTS[result.priority] ?? DEFAULT_WEIGHT;
1879
- total += weight;
1880
- if (result.status === "passed") passed += weight;
1881
- }
1882
- if (total === 0) return 100;
1883
- return Math.round(100 * passed / total);
1884
- }
1885
- function formatScoreLine(score, results, threshold) {
1886
- if (results.length === 0) {
1887
- const base = "Score: 100 (no tests executed)";
1888
- return threshold === void 0 ? base : `${base} \u2014 min-score ${threshold}: pass`;
1938
+ // src/runner/pool.ts
1939
+ async function runWithConcurrency(items, concurrency, run) {
1940
+ if (concurrency < 1) throw new RangeError(`concurrency must be at least 1, got ${concurrency}`);
1941
+ const results = new Array(items.length);
1942
+ let next = 0;
1943
+ const worker = async () => {
1944
+ while (true) {
1945
+ const index = next++;
1946
+ if (index >= items.length) return;
1947
+ results[index] = await run(items[index], index);
1948
+ }
1949
+ };
1950
+ const workers = [];
1951
+ for (let i = 0; i < Math.min(concurrency, items.length); i++) {
1952
+ workers.push(worker());
1889
1953
  }
1890
- if (threshold === void 0) return `Score: ${score}`;
1891
- return score >= threshold ? `Score: ${score} \u2014 min-score ${threshold}: pass` : `Score: ${score} \u2014 min-score ${threshold}: FAIL (below threshold)`;
1892
- }
1893
- function formatIncompleteLine(score, reason) {
1894
- return `Run incomplete: ${reason}
1895
- Score over executed tests: ${score} (not a verdict \u2014 exit code 1 regardless of --min-score)`;
1954
+ const settled = await Promise.allSettled(workers);
1955
+ const failure = settled.find((outcome) => outcome.status === "rejected");
1956
+ if (failure && failure.status === "rejected") throw failure.reason;
1957
+ return results;
1896
1958
  }
1897
1959
 
1898
1960
  // src/runner/selection.ts
@@ -1983,19 +2045,19 @@ function notRunResults(tests, reason) {
1983
2045
  durationMs: 0
1984
2046
  }));
1985
2047
  }
1986
- function printEvent(event) {
2048
+ function printEvent(event, write = console.log) {
1987
2049
  switch (event.type) {
1988
2050
  case "step-start":
1989
- console.log(` ${event.setup ? "(setup) " : ""}step ${event.index + 1}/${event.total}: ${event.step}`);
2051
+ write(` ${event.setup ? "(setup) " : ""}step ${event.index + 1}/${event.total}: ${event.step}`);
1990
2052
  break;
1991
2053
  case "action": {
1992
2054
  const { action, result } = event;
1993
- console.log(` -> ${describeAction(action)} :: ${result}`);
2055
+ write(` -> ${describeAction(action)} :: ${result}`);
1994
2056
  break;
1995
2057
  }
1996
2058
  case "step-end":
1997
2059
  if (event.status === "failed") {
1998
- console.log(` X step failed: ${event.reason ?? "unknown reason"}`);
2060
+ write(` X step failed: ${event.reason ?? "unknown reason"}`);
1999
2061
  }
2000
2062
  break;
2001
2063
  }
@@ -2040,7 +2102,7 @@ ${notRun.length} test(s) not run (run stopped by its budget or deadline):`);
2040
2102
  for (const r of notRun) console.log(` - ${r.summary} (${r.file})`);
2041
2103
  }
2042
2104
  }
2043
- async function runOne(browser, test, config, sessionDir, session, runMask, budget) {
2105
+ async function runOne(browser, test, config, sessionDir, session, runMask, budget, write) {
2044
2106
  const brain = createBrain(createModel(config.llm).model, void 0, budget);
2045
2107
  let resolved;
2046
2108
  const mask = runMask;
@@ -2075,7 +2137,7 @@ async function runOne(browser, test, config, sessionDir, session, runMask, budge
2075
2137
  timeoutMs: config.browser.timeout_ms,
2076
2138
  maxSnapshotLines: config.browser.max_snapshot_lines,
2077
2139
  mask: (text) => mask.mask(text),
2078
- onEvent: printEvent
2140
+ onEvent: (event) => printEvent(event, write)
2079
2141
  });
2080
2142
  } finally {
2081
2143
  await context.close();
@@ -2106,8 +2168,9 @@ function printImpactReport(impact, selection, cwd) {
2106
2168
  }
2107
2169
  console.log("---------------------------------------------------------------");
2108
2170
  }
2109
- async function finalize(results, skipped, options, sessionDir, durationMs, impact, incomplete) {
2171
+ async function finalize(results, skipped, options, sessionDir, durationMs, impact, incomplete, spend) {
2110
2172
  if (results.length > 0) printSummary(results);
2173
+ if (spend) console.log(formatSpendLine(spend));
2111
2174
  const score = computeScore(results);
2112
2175
  console.log(
2113
2176
  incomplete ? formatIncompleteLine(score, incomplete.message) : formatScoreLine(score, results, options.minScore)
@@ -2118,7 +2181,8 @@ async function finalize(results, skipped, options, sessionDir, durationMs, impac
2118
2181
  score,
2119
2182
  durationMs,
2120
2183
  cwd: options.cwd,
2121
- incomplete: incomplete?.message
2184
+ incomplete: incomplete?.message,
2185
+ spend
2122
2186
  });
2123
2187
  await writeJUnit(target, xml);
2124
2188
  console.log(`JUnit report: ${path9.relative(options.cwd, target)}`);
@@ -2130,7 +2194,8 @@ async function finalize(results, skipped, options, sessionDir, durationMs, impac
2130
2194
  durationMs,
2131
2195
  minScore: options.minScore,
2132
2196
  cwd: options.cwd,
2133
- incomplete: incomplete?.message
2197
+ incomplete: incomplete?.message,
2198
+ spend
2134
2199
  });
2135
2200
  await writeHtml(target, html);
2136
2201
  console.log(`HTML report: ${path9.relative(options.cwd, target)}`);
@@ -2317,21 +2382,38 @@ ${results.length} test file(s) could not be parsed:`);
2317
2382
  if (incomplete) {
2318
2383
  results.push(...notRunResults(selected, incomplete.message));
2319
2384
  } else {
2320
- for (let i = 0; i < selected.length; i++) {
2321
- const test = selected[i];
2322
- console.log(`
2323
- > ${test.summary} [${test.priority}] (${path9.relative(options.cwd, test.path)})`);
2385
+ const concurrency = options.concurrency ?? config.concurrency;
2386
+ const streaming = concurrency === 1;
2387
+ let stoppedBy;
2388
+ const outcomes = await runWithConcurrency(selected, concurrency, async (test, index) => {
2389
+ if (stoppedBy) return { index, stopped: stoppedBy };
2390
+ const lines = [];
2391
+ const header = `
2392
+ > ${test.summary} [${test.priority}] (${path9.relative(options.cwd, test.path)})`;
2393
+ const write = streaming ? console.log : (line) => lines.push(line);
2394
+ if (streaming) console.log(header);
2395
+ else lines.push(header);
2324
2396
  try {
2325
2397
  budget.check();
2326
- results.push(await runOne(browser, test, config, sessionDir, session, runMask, budget));
2398
+ const result = await runOne(browser, test, config, sessionDir, session, runMask, budget, write);
2399
+ if (!streaming) for (const line of lines) console.log(line);
2400
+ return { index, result };
2327
2401
  } catch (error) {
2328
2402
  if (error instanceof BudgetExhaustedError) {
2329
- incomplete = error;
2330
- results.push(...notRunResults(selected.slice(i), error.message));
2331
- break;
2403
+ stoppedBy ??= error;
2404
+ if (!streaming) for (const line of lines) console.log(line);
2405
+ return { index, stopped: error };
2332
2406
  }
2333
2407
  throw error;
2334
2408
  }
2409
+ });
2410
+ for (const outcome of outcomes) {
2411
+ if (outcome.result) {
2412
+ results.push(outcome.result);
2413
+ continue;
2414
+ }
2415
+ incomplete ??= outcome.stopped;
2416
+ results.push(...notRunResults([selected[outcome.index]], outcome.stopped.message));
2335
2417
  }
2336
2418
  }
2337
2419
  } finally {
@@ -2344,7 +2426,13 @@ ${results.length} test file(s) could not be parsed:`);
2344
2426
  sessionDir,
2345
2427
  Date.now() - startedAt,
2346
2428
  impact,
2347
- incomplete
2429
+ incomplete,
2430
+ // The command that constructed the budget is the one that reports it
2431
+ // (design report-what-it-spent, D2). `test` hands one budget to both of its
2432
+ // phases deliberately, so reporting at the point of use rather than the
2433
+ // point of ownership would print the running total twice for a single
2434
+ // allowance, the second line silently including the first.
2435
+ options.budget === void 0 ? budget.spend() : void 0
2348
2436
  );
2349
2437
  }
2350
2438
 
@@ -2574,6 +2662,7 @@ ${renderTestYaml(draft, { route, base })}`);
2574
2662
  for (const route of notAttempted) console.log(` ${route}`);
2575
2663
  }
2576
2664
  }
2665
+ if (options.budget === void 0) console.log(formatSpendLine(budget.spend()));
2577
2666
  console.log("---------------------------------------------------------------");
2578
2667
  return incomplete || failed.length > 0 ? EXIT_FAILED : EXIT_OK;
2579
2668
  }
@@ -2618,6 +2707,8 @@ async function testCommand(options) {
2618
2707
  budget
2619
2708
  });
2620
2709
  if (planCode === EXIT_USAGE) return EXIT_USAGE;
2710
+ console.log(`
2711
+ ${formatSpendLine(budget.spend())}`);
2621
2712
  console.log(
2622
2713
  "\nDrafts are not executed and do not affect the score \u2014 review them before they join the suite."
2623
2714
  );
@@ -2660,7 +2751,7 @@ function parsePositiveNumber(flag) {
2660
2751
  };
2661
2752
  }
2662
2753
  var program = new Command();
2663
- program.name("blastproof").description("Open-source AI testing agent: plain-English YAML tests executed agentically on a real browser.").version("0.8.0");
2754
+ program.name("blastproof").description("Open-source AI testing agent: plain-English YAML tests executed agentically on a real browser.").version("0.9.0");
2664
2755
  program.command("init").description("Scaffold .blastproof/ (config, tests, sample tests) in the current directory").action(async () => {
2665
2756
  try {
2666
2757
  const result = await initProject(process.cwd());
@@ -2676,6 +2767,10 @@ program.command("run").description("Discover and run all tests under .blastproof
2676
2767
  "require a weighted score of at least n (0-100); replaces the all-must-pass rule",
2677
2768
  parseMinScore
2678
2769
  ).option("--junit [path]", "write a JUnit XML report (default: .blastproof/reports/<session>/junit.xml)").option("--html [path]", "write a self-contained HTML report (default: .blastproof/reports/<session>/report.html)").option("--fail-on-unmapped", "fail when a changed file matches no routes: or ignore: glob").option(
2770
+ "--concurrency <n>",
2771
+ "run this many tests at once (overrides config; default 1 \u2014 see the README on when this is safe)",
2772
+ parsePositiveInt("--concurrency")
2773
+ ).option(
2679
2774
  "--max-llm-calls <n>",
2680
2775
  "stop the run after this many model calls, reported as incomplete (overrides config)",
2681
2776
  parsePositiveInt("--max-llm-calls")
@@ -2703,6 +2798,7 @@ program.command("run").description("Discover and run all tests under .blastproof
2703
2798
  junit: options.junit,
2704
2799
  html: options.html,
2705
2800
  failOnUnmapped: options.failOnUnmapped,
2801
+ concurrency: options.concurrency,
2706
2802
  maxLlmCalls: options.maxLlmCalls,
2707
2803
  maxTokens: options.maxTokens,
2708
2804
  maxDuration: options.maxDuration