blastproof 0.7.0 → 0.9.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +30 -3
- package/dist/cli.js +172 -51
- package/dist/cli.js.map +1 -1
- package/package.json +1 -1
package/README.md
CHANGED
|
@@ -113,7 +113,7 @@ jobs:
|
|
|
113
113
|
|
|
114
114
|
- run: npm start & # however your app boots
|
|
115
115
|
|
|
116
|
-
- uses: hamc/blastproof@v0.
|
|
116
|
+
- uses: hamc/blastproof@v0.9.0
|
|
117
117
|
with:
|
|
118
118
|
version: '0.6.0' # pin both when this gates merges
|
|
119
119
|
api-key: ${{ secrets.ANTHROPIC_API_KEY }}
|
|
@@ -167,6 +167,22 @@ blastproof plan --base main --dry-run # affected routes no test
|
|
|
167
167
|
|
|
168
168
|
They report affected routes, files nobody has classified, and affected routes no test covers — a coverage-gap report with an exit code, useful even on a repo whose suite is Playwright or Cypress.
|
|
169
169
|
|
|
170
|
+
## Running tests at once
|
|
171
|
+
|
|
172
|
+
Tests run one at a time by default. Raise it when your tests can stand it:
|
|
173
|
+
|
|
174
|
+
```yaml
|
|
175
|
+
concurrency: 4
|
|
176
|
+
```
|
|
177
|
+
|
|
178
|
+
or `blastproof run --concurrency 4` for a single invocation. On this repository's own suite that takes a run from 156s to 68s — **2.3× faster, for the same 81 model calls.** Parallelism buys wall-clock, not spend.
|
|
179
|
+
|
|
180
|
+
**The default is 1 on purpose, and raising it is your call to make.** Other test runners default to parallel because their tests are isolated by construction — separate processes, separate fixtures. These are journeys driven against **one running application**, so two tests can see each other's data. A suite is safe to parallelise when its tests do not write state that another test reads.
|
|
181
|
+
|
|
182
|
+
The test in this repository's own suite that could not run beside itself is a good shape to recognise: it adds a note and then asserts *"one note on file"*. It writes shared server state, and it asserts on a global count. Either alone is a warning; together they mean the test's verdict depends on nothing else touching the application at that moment.
|
|
183
|
+
|
|
184
|
+
Two practical notes. Four concurrent journeys are four times the traffic against whatever you pointed at — usually fine for a development instance, worth knowing for a shared one. And with several model calls in flight, a `budget:` limit can overshoot by up to the concurrency rather than by a single call, since the calls already sent are allowed to finish.
|
|
185
|
+
|
|
170
186
|
## Bounding a run
|
|
171
187
|
|
|
172
188
|
Nothing stops a run by default. `budget:` puts a ceiling on `run`, `plan` and `test` alike — every model call any of them makes is counted:
|
|
@@ -182,7 +198,16 @@ Each limit is optional; with none set, nothing binds. They count **calls and tok
|
|
|
182
198
|
|
|
183
199
|
Exhausting a budget **stops the run; it does not fail a test.** Running out of quota says nothing about the code under review. Unreached tests are reported as `not run`, a third state excluded from the score entirely, and the process exits 1 unconditionally — `--min-score` cannot rescue it, because the tests that finished are whichever ran first, not a representative sample.
|
|
184
200
|
|
|
185
|
-
|
|
201
|
+
**Every run reports what it spent**, so you can size a limit from experience rather than guesswork:
|
|
202
|
+
|
|
203
|
+
```
|
|
204
|
+
Spent: 82 model call(s), 115407 token(s)
|
|
205
|
+
Score: 100
|
|
206
|
+
```
|
|
207
|
+
|
|
208
|
+
The figures also land in the JUnit report as `llm_calls` and `llm_tokens`, beside `score`, so a pipeline can trend cost without scraping output. A run stopped by its own budget reports the spend too — that is the case where the number is least guessable. Where a provider reports no token usage, the line says so rather than showing zero.
|
|
209
|
+
|
|
210
|
+
`--dry-run` reports the ceiling before you spend anything. Read it as a maximum and nothing more: for this repository's own suite it says 735 calls where a real run spends 82. Size a budget from what your runs actually report, not from the ceiling.
|
|
186
211
|
|
|
187
212
|
## Testing behind a login
|
|
188
213
|
|
|
@@ -241,7 +266,9 @@ The key itself is never read from a `BLASTPROOF_*` variable — you name *which*
|
|
|
241
266
|
|
|
242
267
|
The application under test is not trusted input: its page content reaches the model, so a page that controls its own accessible text can try to influence the agent. Two things constrain that.
|
|
243
268
|
|
|
244
|
-
**The agent cannot leave your application.**
|
|
269
|
+
**The agent cannot leave your application.** The boundary is `base_url`'s origin plus whatever `allowed_origins:` declares, and it constrains where the page **is**, not only where an action asked to go. A `navigate` outside it is refused before the request; a page that ends up outside it any other way — a redirect, a link to another host, a script setting the location — fails the step, and its content is never sent to the model. Enforced by comparison, not by asking the model nicely.
|
|
270
|
+
|
|
271
|
+
If your application legitimately spans hosts (an identity provider, a hosted payment step), declare them. A suite that was quietly walking onto a foreign page will now fail and name the origin to add.
|
|
245
272
|
|
|
246
273
|
**Your secrets stay out of prompts.** `{{env.*}}` placeholders survive intact and are substituted at the moment of typing. Every value your tests or auth recipe reference is redacted from everything else crossing into a prompt — page snapshots included — in literal and percent-encoded form. Redaction matches known values, so treat it as a strong default rather than a guarantee against a hostile app.
|
|
247
274
|
|
package/dist/cli.js
CHANGED
|
@@ -210,6 +210,7 @@ var RunBudget = class {
|
|
|
210
210
|
startedAt;
|
|
211
211
|
calls = 0;
|
|
212
212
|
tokens = 0;
|
|
213
|
+
callsWithUsage = 0;
|
|
213
214
|
constructor(options = {}) {
|
|
214
215
|
this.maxCalls = options.maxCalls;
|
|
215
216
|
this.maxTokens = options.maxTokens;
|
|
@@ -248,7 +249,25 @@ var RunBudget = class {
|
|
|
248
249
|
/** Records what a completed model call spent. */
|
|
249
250
|
record(usage) {
|
|
250
251
|
this.calls += 1;
|
|
251
|
-
if (usage?.totalTokens !== void 0)
|
|
252
|
+
if (usage?.totalTokens !== void 0) {
|
|
253
|
+
this.tokens += usage.totalTokens;
|
|
254
|
+
this.callsWithUsage += 1;
|
|
255
|
+
}
|
|
256
|
+
}
|
|
257
|
+
/**
|
|
258
|
+
* What this budget has been spent on, for every surface that reports it
|
|
259
|
+
* (design report-what-it-spent, D1). One method rather than four getters read
|
|
260
|
+
* separately: the console, the JUnit report and the HTML report take the same
|
|
261
|
+
* numbers from the same object, so they cannot disagree about what a run cost.
|
|
262
|
+
*/
|
|
263
|
+
spend() {
|
|
264
|
+
return {
|
|
265
|
+
calls: this.calls,
|
|
266
|
+
maxCalls: this.maxCalls,
|
|
267
|
+
tokens: this.tokens,
|
|
268
|
+
callsWithUsage: this.callsWithUsage,
|
|
269
|
+
maxTokens: this.maxTokens
|
|
270
|
+
};
|
|
252
271
|
}
|
|
253
272
|
};
|
|
254
273
|
function estimateMaxModelCalls(tests, maxIterationsPerStep, maxRetriesPerStep) {
|
|
@@ -336,14 +355,33 @@ var ActionError = class extends Error {
|
|
|
336
355
|
this.name = "ActionError";
|
|
337
356
|
}
|
|
338
357
|
};
|
|
339
|
-
|
|
340
|
-
|
|
341
|
-
|
|
358
|
+
var BLANK_PAGE = "about:blank";
|
|
359
|
+
function allowedOriginsFor(baseUrl, allowedOrigins) {
|
|
360
|
+
const allowed = /* @__PURE__ */ new Set([new URL(baseUrl).origin]);
|
|
361
|
+
for (const origin of allowedOrigins ?? []) {
|
|
342
362
|
allowed.add(new URL(origin).origin);
|
|
343
363
|
}
|
|
344
|
-
|
|
364
|
+
return allowed;
|
|
365
|
+
}
|
|
366
|
+
function isOriginAllowed(url, allowed) {
|
|
367
|
+
if (url === BLANK_PAGE) return true;
|
|
368
|
+
let origin;
|
|
369
|
+
try {
|
|
370
|
+
origin = new URL(url).origin;
|
|
371
|
+
} catch {
|
|
372
|
+
return false;
|
|
373
|
+
}
|
|
374
|
+
if (origin === "null") return false;
|
|
375
|
+
return allowed.has(origin);
|
|
376
|
+
}
|
|
377
|
+
function describeBoundary(allowed) {
|
|
378
|
+
return [...allowed].join(" or ");
|
|
379
|
+
}
|
|
380
|
+
function assertAllowedOrigin(url, ctx) {
|
|
381
|
+
const allowed = allowedOriginsFor(ctx.baseUrl, ctx.allowedOrigins);
|
|
382
|
+
if (!isOriginAllowed(url.toString(), allowed)) {
|
|
345
383
|
throw new ActionError(
|
|
346
|
-
`Refusing to navigate outside the application: ${url.origin} is not ${
|
|
384
|
+
`Refusing to navigate outside the application: ${url.origin} is not ${describeBoundary(allowed)}. Add it to allowed_origins in .blastproof/config.yaml if the app legitimately spans hosts.`
|
|
347
385
|
);
|
|
348
386
|
}
|
|
349
387
|
}
|
|
@@ -548,6 +586,7 @@ async function executeTest(page, test, options) {
|
|
|
548
586
|
onEvent({ type: "action", index, action: maskedAction, result: mask(result) });
|
|
549
587
|
};
|
|
550
588
|
let failure;
|
|
589
|
+
const boundary = allowedOriginsFor(baseUrl, allowedOrigins);
|
|
551
590
|
await page.goto(new URL(baseUrl).toString());
|
|
552
591
|
for (let index = 0; index < allSteps.length; index++) {
|
|
553
592
|
const { step, setup } = allSteps[index];
|
|
@@ -564,6 +603,11 @@ async function executeTest(page, test, options) {
|
|
|
564
603
|
throw new StepFailure(`step exceeded ${maxIterationsPerStep} actions without completing`);
|
|
565
604
|
}
|
|
566
605
|
await waitForSettled(page);
|
|
606
|
+
if (!isOriginAllowed(page.url(), boundary)) {
|
|
607
|
+
throw new StepFailure(
|
|
608
|
+
`The page left the application: ${page.url()} is outside ${describeBoundary(boundary)}. Add it to allowed_origins in .blastproof/config.yaml if the app legitimately spans hosts.`
|
|
609
|
+
);
|
|
610
|
+
}
|
|
567
611
|
const snap = await takeSnapshot(page);
|
|
568
612
|
let action;
|
|
569
613
|
try {
|
|
@@ -922,6 +966,13 @@ var configSchema = z.object({
|
|
|
922
966
|
allowed_origins: z.array(z.string().url()).optional(),
|
|
923
967
|
auth: authSchema.optional(),
|
|
924
968
|
max_retries_per_step: z.number().int().min(1).default(3),
|
|
969
|
+
/**
|
|
970
|
+
* How many tests may run at once (design tests-in-parallel, D1). Defaults to
|
|
971
|
+
* 1 — tests are journeys driven against one running application, and whether
|
|
972
|
+
* two of them can run at the same time is a property of that application and
|
|
973
|
+
* those tests, not of the runner. Opted into by the person who knows.
|
|
974
|
+
*/
|
|
975
|
+
concurrency: z.number().int().min(1, "concurrency must be at least 1").default(1),
|
|
925
976
|
budget: budgetSchema.optional()
|
|
926
977
|
});
|
|
927
978
|
var ENV_OVERRIDES = {
|
|
@@ -1601,6 +1652,43 @@ function printPreflightFailures(failures) {
|
|
|
1601
1652
|
for (const failure of failures) console.error(` - ${failure}`);
|
|
1602
1653
|
}
|
|
1603
1654
|
|
|
1655
|
+
// src/report/score.ts
|
|
1656
|
+
var WEIGHTS = { P0: 3, P1: 2, P2: 1 };
|
|
1657
|
+
var DEFAULT_WEIGHT = WEIGHTS.P1;
|
|
1658
|
+
function computeScore(results) {
|
|
1659
|
+
let total = 0;
|
|
1660
|
+
let passed = 0;
|
|
1661
|
+
for (const result of results) {
|
|
1662
|
+
if (result.status === "not-run") continue;
|
|
1663
|
+
const weight = WEIGHTS[result.priority] ?? DEFAULT_WEIGHT;
|
|
1664
|
+
total += weight;
|
|
1665
|
+
if (result.status === "passed") passed += weight;
|
|
1666
|
+
}
|
|
1667
|
+
if (total === 0) return 100;
|
|
1668
|
+
return Math.round(100 * passed / total);
|
|
1669
|
+
}
|
|
1670
|
+
function formatScoreLine(score, results, threshold) {
|
|
1671
|
+
if (results.length === 0) {
|
|
1672
|
+
const base = "Score: 100 (no tests executed)";
|
|
1673
|
+
return threshold === void 0 ? base : `${base} \u2014 min-score ${threshold}: pass`;
|
|
1674
|
+
}
|
|
1675
|
+
if (threshold === void 0) return `Score: ${score}`;
|
|
1676
|
+
return score >= threshold ? `Score: ${score} \u2014 min-score ${threshold}: pass` : `Score: ${score} \u2014 min-score ${threshold}: FAIL (below threshold)`;
|
|
1677
|
+
}
|
|
1678
|
+
function formatIncompleteLine(score, reason) {
|
|
1679
|
+
return `Run incomplete: ${reason}
|
|
1680
|
+
Score over executed tests: ${score} (not a verdict \u2014 exit code 1 regardless of --min-score)`;
|
|
1681
|
+
}
|
|
1682
|
+
function formatSpendLine(spend) {
|
|
1683
|
+
const calls = spend.maxCalls === void 0 ? `${spend.calls} model call(s)` : `${spend.calls} of ${spend.maxCalls} model call(s)`;
|
|
1684
|
+
if (spend.callsWithUsage === 0) {
|
|
1685
|
+
return `Spent: ${calls}; token usage not reported by the provider`;
|
|
1686
|
+
}
|
|
1687
|
+
const tokens = spend.maxTokens === void 0 ? `${spend.tokens} token(s)` : `${spend.tokens} of ${spend.maxTokens} token(s)`;
|
|
1688
|
+
const coverage = spend.callsWithUsage < spend.calls ? ` (tokens reported by ${spend.callsWithUsage} of ${spend.calls} call(s))` : "";
|
|
1689
|
+
return `Spent: ${calls}, ${tokens}${coverage}`;
|
|
1690
|
+
}
|
|
1691
|
+
|
|
1604
1692
|
// src/commands/run.ts
|
|
1605
1693
|
import path9 from "path";
|
|
1606
1694
|
|
|
@@ -1745,7 +1833,7 @@ async function renderHtml(results, skipped, meta) {
|
|
|
1745
1833
|
<body>
|
|
1746
1834
|
<main>
|
|
1747
1835
|
<h1>blastproof report</h1>
|
|
1748
|
-
<p class="sub">${escapeHtml(generatedAt)} \xB7 ${seconds(meta.durationMs)}</p>
|
|
1836
|
+
<p class="sub">${escapeHtml(generatedAt)} \xB7 ${seconds(meta.durationMs)}${meta.spend ? ` \xB7 ${escapeHtml(formatSpendLine(meta.spend))}` : ""}</p>
|
|
1749
1837
|
|
|
1750
1838
|
${banner} <section class="score">
|
|
1751
1839
|
<b>${meta.score}</b>
|
|
@@ -1799,6 +1887,11 @@ function renderJUnit(results, skipped, meta) {
|
|
|
1799
1887
|
`<testsuite name="blastproof" tests="${results.length + skipped.length}" failures="${failures}" skipped="${skipped.length + notRun.length}" time="${seconds2(meta.durationMs)}">`,
|
|
1800
1888
|
" <properties>",
|
|
1801
1889
|
` <property name="score" value="${meta.score}"/>`,
|
|
1890
|
+
...meta.spend ? [` <property name="llm_calls" value="${meta.spend.calls}"/>`] : [],
|
|
1891
|
+
// Omitted rather than emitted as zero when no call reported usage: a
|
|
1892
|
+
// property carrying 0 would be read by a pipeline as "this run spent no
|
|
1893
|
+
// tokens", which is a different claim from "the provider did not say".
|
|
1894
|
+
...meta.spend && meta.spend.callsWithUsage > 0 ? [` <property name="llm_tokens" value="${meta.spend.tokens}"/>`] : [],
|
|
1802
1895
|
...meta.incomplete !== void 0 ? [
|
|
1803
1896
|
' <property name="incomplete" value="true"/>',
|
|
1804
1897
|
` <property name="incomplete_reason" value="${escapeXml(meta.incomplete)}"/>`
|
|
@@ -1842,32 +1935,26 @@ async function writeJUnit(file, xml) {
|
|
|
1842
1935
|
return file;
|
|
1843
1936
|
}
|
|
1844
1937
|
|
|
1845
|
-
// src/
|
|
1846
|
-
|
|
1847
|
-
|
|
1848
|
-
|
|
1849
|
-
let
|
|
1850
|
-
|
|
1851
|
-
|
|
1852
|
-
|
|
1853
|
-
|
|
1854
|
-
|
|
1855
|
-
|
|
1856
|
-
}
|
|
1857
|
-
|
|
1858
|
-
|
|
1859
|
-
|
|
1860
|
-
function formatScoreLine(score, results, threshold) {
|
|
1861
|
-
if (results.length === 0) {
|
|
1862
|
-
const base = "Score: 100 (no tests executed)";
|
|
1863
|
-
return threshold === void 0 ? base : `${base} \u2014 min-score ${threshold}: pass`;
|
|
1938
|
+
// src/runner/pool.ts
|
|
1939
|
+
async function runWithConcurrency(items, concurrency, run) {
|
|
1940
|
+
if (concurrency < 1) throw new RangeError(`concurrency must be at least 1, got ${concurrency}`);
|
|
1941
|
+
const results = new Array(items.length);
|
|
1942
|
+
let next = 0;
|
|
1943
|
+
const worker = async () => {
|
|
1944
|
+
while (true) {
|
|
1945
|
+
const index = next++;
|
|
1946
|
+
if (index >= items.length) return;
|
|
1947
|
+
results[index] = await run(items[index], index);
|
|
1948
|
+
}
|
|
1949
|
+
};
|
|
1950
|
+
const workers = [];
|
|
1951
|
+
for (let i = 0; i < Math.min(concurrency, items.length); i++) {
|
|
1952
|
+
workers.push(worker());
|
|
1864
1953
|
}
|
|
1865
|
-
|
|
1866
|
-
|
|
1867
|
-
|
|
1868
|
-
|
|
1869
|
-
return `Run incomplete: ${reason}
|
|
1870
|
-
Score over executed tests: ${score} (not a verdict \u2014 exit code 1 regardless of --min-score)`;
|
|
1954
|
+
const settled = await Promise.allSettled(workers);
|
|
1955
|
+
const failure = settled.find((outcome) => outcome.status === "rejected");
|
|
1956
|
+
if (failure && failure.status === "rejected") throw failure.reason;
|
|
1957
|
+
return results;
|
|
1871
1958
|
}
|
|
1872
1959
|
|
|
1873
1960
|
// src/runner/selection.ts
|
|
@@ -1958,19 +2045,19 @@ function notRunResults(tests, reason) {
|
|
|
1958
2045
|
durationMs: 0
|
|
1959
2046
|
}));
|
|
1960
2047
|
}
|
|
1961
|
-
function printEvent(event) {
|
|
2048
|
+
function printEvent(event, write = console.log) {
|
|
1962
2049
|
switch (event.type) {
|
|
1963
2050
|
case "step-start":
|
|
1964
|
-
|
|
2051
|
+
write(` ${event.setup ? "(setup) " : ""}step ${event.index + 1}/${event.total}: ${event.step}`);
|
|
1965
2052
|
break;
|
|
1966
2053
|
case "action": {
|
|
1967
2054
|
const { action, result } = event;
|
|
1968
|
-
|
|
2055
|
+
write(` -> ${describeAction(action)} :: ${result}`);
|
|
1969
2056
|
break;
|
|
1970
2057
|
}
|
|
1971
2058
|
case "step-end":
|
|
1972
2059
|
if (event.status === "failed") {
|
|
1973
|
-
|
|
2060
|
+
write(` X step failed: ${event.reason ?? "unknown reason"}`);
|
|
1974
2061
|
}
|
|
1975
2062
|
break;
|
|
1976
2063
|
}
|
|
@@ -2015,7 +2102,7 @@ ${notRun.length} test(s) not run (run stopped by its budget or deadline):`);
|
|
|
2015
2102
|
for (const r of notRun) console.log(` - ${r.summary} (${r.file})`);
|
|
2016
2103
|
}
|
|
2017
2104
|
}
|
|
2018
|
-
async function runOne(browser, test, config, sessionDir, session, runMask, budget) {
|
|
2105
|
+
async function runOne(browser, test, config, sessionDir, session, runMask, budget, write) {
|
|
2019
2106
|
const brain = createBrain(createModel(config.llm).model, void 0, budget);
|
|
2020
2107
|
let resolved;
|
|
2021
2108
|
const mask = runMask;
|
|
@@ -2050,7 +2137,7 @@ async function runOne(browser, test, config, sessionDir, session, runMask, budge
|
|
|
2050
2137
|
timeoutMs: config.browser.timeout_ms,
|
|
2051
2138
|
maxSnapshotLines: config.browser.max_snapshot_lines,
|
|
2052
2139
|
mask: (text) => mask.mask(text),
|
|
2053
|
-
onEvent: printEvent
|
|
2140
|
+
onEvent: (event) => printEvent(event, write)
|
|
2054
2141
|
});
|
|
2055
2142
|
} finally {
|
|
2056
2143
|
await context.close();
|
|
@@ -2081,8 +2168,9 @@ function printImpactReport(impact, selection, cwd) {
|
|
|
2081
2168
|
}
|
|
2082
2169
|
console.log("---------------------------------------------------------------");
|
|
2083
2170
|
}
|
|
2084
|
-
async function finalize(results, skipped, options, sessionDir, durationMs, impact, incomplete) {
|
|
2171
|
+
async function finalize(results, skipped, options, sessionDir, durationMs, impact, incomplete, spend) {
|
|
2085
2172
|
if (results.length > 0) printSummary(results);
|
|
2173
|
+
if (spend) console.log(formatSpendLine(spend));
|
|
2086
2174
|
const score = computeScore(results);
|
|
2087
2175
|
console.log(
|
|
2088
2176
|
incomplete ? formatIncompleteLine(score, incomplete.message) : formatScoreLine(score, results, options.minScore)
|
|
@@ -2093,7 +2181,8 @@ async function finalize(results, skipped, options, sessionDir, durationMs, impac
|
|
|
2093
2181
|
score,
|
|
2094
2182
|
durationMs,
|
|
2095
2183
|
cwd: options.cwd,
|
|
2096
|
-
incomplete: incomplete?.message
|
|
2184
|
+
incomplete: incomplete?.message,
|
|
2185
|
+
spend
|
|
2097
2186
|
});
|
|
2098
2187
|
await writeJUnit(target, xml);
|
|
2099
2188
|
console.log(`JUnit report: ${path9.relative(options.cwd, target)}`);
|
|
@@ -2105,7 +2194,8 @@ async function finalize(results, skipped, options, sessionDir, durationMs, impac
|
|
|
2105
2194
|
durationMs,
|
|
2106
2195
|
minScore: options.minScore,
|
|
2107
2196
|
cwd: options.cwd,
|
|
2108
|
-
incomplete: incomplete?.message
|
|
2197
|
+
incomplete: incomplete?.message,
|
|
2198
|
+
spend
|
|
2109
2199
|
});
|
|
2110
2200
|
await writeHtml(target, html);
|
|
2111
2201
|
console.log(`HTML report: ${path9.relative(options.cwd, target)}`);
|
|
@@ -2292,21 +2382,38 @@ ${results.length} test file(s) could not be parsed:`);
|
|
|
2292
2382
|
if (incomplete) {
|
|
2293
2383
|
results.push(...notRunResults(selected, incomplete.message));
|
|
2294
2384
|
} else {
|
|
2295
|
-
|
|
2296
|
-
|
|
2297
|
-
|
|
2298
|
-
|
|
2385
|
+
const concurrency = options.concurrency ?? config.concurrency;
|
|
2386
|
+
const streaming = concurrency === 1;
|
|
2387
|
+
let stoppedBy;
|
|
2388
|
+
const outcomes = await runWithConcurrency(selected, concurrency, async (test, index) => {
|
|
2389
|
+
if (stoppedBy) return { index, stopped: stoppedBy };
|
|
2390
|
+
const lines = [];
|
|
2391
|
+
const header = `
|
|
2392
|
+
> ${test.summary} [${test.priority}] (${path9.relative(options.cwd, test.path)})`;
|
|
2393
|
+
const write = streaming ? console.log : (line) => lines.push(line);
|
|
2394
|
+
if (streaming) console.log(header);
|
|
2395
|
+
else lines.push(header);
|
|
2299
2396
|
try {
|
|
2300
2397
|
budget.check();
|
|
2301
|
-
|
|
2398
|
+
const result = await runOne(browser, test, config, sessionDir, session, runMask, budget, write);
|
|
2399
|
+
if (!streaming) for (const line of lines) console.log(line);
|
|
2400
|
+
return { index, result };
|
|
2302
2401
|
} catch (error) {
|
|
2303
2402
|
if (error instanceof BudgetExhaustedError) {
|
|
2304
|
-
|
|
2305
|
-
|
|
2306
|
-
|
|
2403
|
+
stoppedBy ??= error;
|
|
2404
|
+
if (!streaming) for (const line of lines) console.log(line);
|
|
2405
|
+
return { index, stopped: error };
|
|
2307
2406
|
}
|
|
2308
2407
|
throw error;
|
|
2309
2408
|
}
|
|
2409
|
+
});
|
|
2410
|
+
for (const outcome of outcomes) {
|
|
2411
|
+
if (outcome.result) {
|
|
2412
|
+
results.push(outcome.result);
|
|
2413
|
+
continue;
|
|
2414
|
+
}
|
|
2415
|
+
incomplete ??= outcome.stopped;
|
|
2416
|
+
results.push(...notRunResults([selected[outcome.index]], outcome.stopped.message));
|
|
2310
2417
|
}
|
|
2311
2418
|
}
|
|
2312
2419
|
} finally {
|
|
@@ -2319,7 +2426,13 @@ ${results.length} test file(s) could not be parsed:`);
|
|
|
2319
2426
|
sessionDir,
|
|
2320
2427
|
Date.now() - startedAt,
|
|
2321
2428
|
impact,
|
|
2322
|
-
incomplete
|
|
2429
|
+
incomplete,
|
|
2430
|
+
// The command that constructed the budget is the one that reports it
|
|
2431
|
+
// (design report-what-it-spent, D2). `test` hands one budget to both of its
|
|
2432
|
+
// phases deliberately, so reporting at the point of use rather than the
|
|
2433
|
+
// point of ownership would print the running total twice for a single
|
|
2434
|
+
// allowance, the second line silently including the first.
|
|
2435
|
+
options.budget === void 0 ? budget.spend() : void 0
|
|
2323
2436
|
);
|
|
2324
2437
|
}
|
|
2325
2438
|
|
|
@@ -2549,6 +2662,7 @@ ${renderTestYaml(draft, { route, base })}`);
|
|
|
2549
2662
|
for (const route of notAttempted) console.log(` ${route}`);
|
|
2550
2663
|
}
|
|
2551
2664
|
}
|
|
2665
|
+
if (options.budget === void 0) console.log(formatSpendLine(budget.spend()));
|
|
2552
2666
|
console.log("---------------------------------------------------------------");
|
|
2553
2667
|
return incomplete || failed.length > 0 ? EXIT_FAILED : EXIT_OK;
|
|
2554
2668
|
}
|
|
@@ -2593,6 +2707,8 @@ async function testCommand(options) {
|
|
|
2593
2707
|
budget
|
|
2594
2708
|
});
|
|
2595
2709
|
if (planCode === EXIT_USAGE) return EXIT_USAGE;
|
|
2710
|
+
console.log(`
|
|
2711
|
+
${formatSpendLine(budget.spend())}`);
|
|
2596
2712
|
console.log(
|
|
2597
2713
|
"\nDrafts are not executed and do not affect the score \u2014 review them before they join the suite."
|
|
2598
2714
|
);
|
|
@@ -2635,7 +2751,7 @@ function parsePositiveNumber(flag) {
|
|
|
2635
2751
|
};
|
|
2636
2752
|
}
|
|
2637
2753
|
var program = new Command();
|
|
2638
|
-
program.name("blastproof").description("Open-source AI testing agent: plain-English YAML tests executed agentically on a real browser.").version("0.
|
|
2754
|
+
program.name("blastproof").description("Open-source AI testing agent: plain-English YAML tests executed agentically on a real browser.").version("0.9.0");
|
|
2639
2755
|
program.command("init").description("Scaffold .blastproof/ (config, tests, sample tests) in the current directory").action(async () => {
|
|
2640
2756
|
try {
|
|
2641
2757
|
const result = await initProject(process.cwd());
|
|
@@ -2651,6 +2767,10 @@ program.command("run").description("Discover and run all tests under .blastproof
|
|
|
2651
2767
|
"require a weighted score of at least n (0-100); replaces the all-must-pass rule",
|
|
2652
2768
|
parseMinScore
|
|
2653
2769
|
).option("--junit [path]", "write a JUnit XML report (default: .blastproof/reports/<session>/junit.xml)").option("--html [path]", "write a self-contained HTML report (default: .blastproof/reports/<session>/report.html)").option("--fail-on-unmapped", "fail when a changed file matches no routes: or ignore: glob").option(
|
|
2770
|
+
"--concurrency <n>",
|
|
2771
|
+
"run this many tests at once (overrides config; default 1 \u2014 see the README on when this is safe)",
|
|
2772
|
+
parsePositiveInt("--concurrency")
|
|
2773
|
+
).option(
|
|
2654
2774
|
"--max-llm-calls <n>",
|
|
2655
2775
|
"stop the run after this many model calls, reported as incomplete (overrides config)",
|
|
2656
2776
|
parsePositiveInt("--max-llm-calls")
|
|
@@ -2678,6 +2798,7 @@ program.command("run").description("Discover and run all tests under .blastproof
|
|
|
2678
2798
|
junit: options.junit,
|
|
2679
2799
|
html: options.html,
|
|
2680
2800
|
failOnUnmapped: options.failOnUnmapped,
|
|
2801
|
+
concurrency: options.concurrency,
|
|
2681
2802
|
maxLlmCalls: options.maxLlmCalls,
|
|
2682
2803
|
maxTokens: options.maxTokens,
|
|
2683
2804
|
maxDuration: options.maxDuration
|