blastproof 0.8.0 → 0.10.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +51 -2
- package/dist/cli.js +174 -55
- package/dist/cli.js.map +1 -1
- package/package.json +2 -2
package/README.md
CHANGED
|
@@ -42,6 +42,8 @@ Before `run`, `plan` or `test` do anything, they check what they are about to sp
|
|
|
42
42
|
|
|
43
43
|
**Your markup must be accessible — a hard requirement.** Elements are found by role, label or visible text from the accessibility tree. That is what removes selectors and survives redesigns; the cost is that an interface the accessibility tree cannot describe cannot be driven at all, and there is deliberately no CSS or XPath fallback. Icon-only buttons without accessible names, `div`-based controls and ARIA-less dropdowns simply cannot be targeted. Run an accessibility checker first — the result predicts how well this will work better than anything else.
|
|
44
44
|
|
|
45
|
+
**Windows is untested.** Development and CI run on Linux and macOS. Nothing is known to be broken and reports are welcome (#8).
|
|
46
|
+
|
|
45
47
|
**Not supported yet:** `iframe` content (so hosted payment widgets like Stripe Elements are invisible — an embedded checkout cannot be driven end to end), hover, scroll-to, drag and drop, file upload, multiple tabs, native `alert`/`confirm` dialogs. Page snapshots are capped at 200 lines by default, so very dense pages are truncated — raise it with `browser.max_snapshot_lines` if your pages need more; truncation is always marked in the snapshot so the model is never misled into thinking it saw the whole page.
|
|
46
48
|
|
|
47
49
|
**Point it at disposable data.** Within a step, an action that commits — a click, or pressing Enter — is never performed twice: the runner refuses the repeat and tells the agent it already did that. This closes the case that used to produce duplicate records, where a submit answered by a redirect came back to a reset form and the agent, seeing no evidence of its own work, submitted again. It is not a guarantee of zero duplicate writes: an agent that reaches the same effect by a genuinely different route — another control with the same effect — is not caught. Use a seeded database, a staging environment you can reset, or a throwaway account; do not gate on a run against production data.
|
|
@@ -70,6 +72,26 @@ steps:
|
|
|
70
72
|
|
|
71
73
|
`priority` is P0–P2 (default P1). `tags`, `setup` steps and `auth` are optional — `auth: false` runs the test signed out, which a login test needs. `routes` declares the URLs a test covers, which is what `--impacted` selects on; write route strings consistently, since they compare by exact equality (`/cart` ≠ `/cart/`).
|
|
72
74
|
|
|
75
|
+
### Say what each step should produce
|
|
76
|
+
|
|
77
|
+
This is the one rule that decides whether a suite works. An outside evaluation took the **same application, same suite, same version from Score 64 to Score 100 by rewriting two steps** — nothing else changed:
|
|
78
|
+
|
|
79
|
+
```yaml
|
|
80
|
+
# Fragile: a bare action. Nothing says what should be true afterwards.
|
|
81
|
+
- submit the add-task form
|
|
82
|
+
- verify the new task "Fix the flaky test" appears in the task list
|
|
83
|
+
|
|
84
|
+
# Robust: the step carries its own outcome.
|
|
85
|
+
- submit the add-task form and verify the task "Fix the flaky test" appears
|
|
86
|
+
in the task list with priority High and status Open
|
|
87
|
+
```
|
|
88
|
+
|
|
89
|
+
The bare version fails on a shape that is everywhere: the form POSTs, the server redirects back to the same page, and the form comes back empty. The agent is asked whether "submit the add-task form" happened, and is looking at a page that is indistinguishable from one where nothing did. A step that names the outcome gives it something to check that survives the action.
|
|
90
|
+
|
|
91
|
+
Write steps that end in an observable result — text on the page, a count, a state change — and this class of failure does not arise.
|
|
92
|
+
|
|
93
|
+
**Inline error messages should be plain visible text.** `role="alert"` is read correctly from the accessibility tree and needs no special handling, but note that an alert your page has cleared shows up as an empty element: if a verdict says an alert exists whose content is missing, the message was emptied, not hidden.
|
|
94
|
+
|
|
73
95
|
## Commands
|
|
74
96
|
|
|
75
97
|
```bash
|
|
@@ -113,7 +135,7 @@ jobs:
|
|
|
113
135
|
|
|
114
136
|
- run: npm start & # however your app boots
|
|
115
137
|
|
|
116
|
-
- uses: hamc/blastproof@v0.
|
|
138
|
+
- uses: hamc/blastproof@v0.10.0
|
|
117
139
|
with:
|
|
118
140
|
version: '0.6.0' # pin both when this gates merges
|
|
119
141
|
api-key: ${{ secrets.ANTHROPIC_API_KEY }}
|
|
@@ -167,6 +189,22 @@ blastproof plan --base main --dry-run # affected routes no test
|
|
|
167
189
|
|
|
168
190
|
They report affected routes, files nobody has classified, and affected routes no test covers — a coverage-gap report with an exit code, useful even on a repo whose suite is Playwright or Cypress.
|
|
169
191
|
|
|
192
|
+
## Running tests at once
|
|
193
|
+
|
|
194
|
+
Tests run one at a time by default. Raise it when your tests can stand it:
|
|
195
|
+
|
|
196
|
+
```yaml
|
|
197
|
+
concurrency: 4
|
|
198
|
+
```
|
|
199
|
+
|
|
200
|
+
or `blastproof run --concurrency 4` for a single invocation. On this repository's own suite that takes a run from 156s to 68s — **2.3× faster, for the same 81 model calls.** Parallelism buys wall-clock, not spend.
|
|
201
|
+
|
|
202
|
+
**The default is 1 on purpose, and raising it is your call to make.** Other test runners default to parallel because their tests are isolated by construction — separate processes, separate fixtures. These are journeys driven against **one running application**, so two tests can see each other's data. A suite is safe to parallelise when its tests do not write state that another test reads.
|
|
203
|
+
|
|
204
|
+
The test in this repository's own suite that could not run beside itself is a good shape to recognise: it adds a note and then asserts *"one note on file"*. It writes shared server state, and it asserts on a global count. Either alone is a warning; together they mean the test's verdict depends on nothing else touching the application at that moment.
|
|
205
|
+
|
|
206
|
+
Two practical notes. Four concurrent journeys are four times the traffic against whatever you pointed at — usually fine for a development instance, worth knowing for a shared one. And with several model calls in flight, a `budget:` limit can overshoot by up to the concurrency rather than by a single call, since the calls already sent are allowed to finish.
|
|
207
|
+
|
|
170
208
|
## Bounding a run
|
|
171
209
|
|
|
172
210
|
Nothing stops a run by default. `budget:` puts a ceiling on `run`, `plan` and `test` alike — every model call any of them makes is counted:
|
|
@@ -182,7 +220,18 @@ Each limit is optional; with none set, nothing binds. They count **calls and tok
|
|
|
182
220
|
|
|
183
221
|
Exhausting a budget **stops the run; it does not fail a test.** Running out of quota says nothing about the code under review. Unreached tests are reported as `not run`, a third state excluded from the score entirely, and the process exits 1 unconditionally — `--min-score` cannot rescue it, because the tests that finished are whichever ran first, not a representative sample.
|
|
184
222
|
|
|
185
|
-
|
|
223
|
+
**Every run reports what it spent**, so you can size a limit from experience rather than guesswork:
|
|
224
|
+
|
|
225
|
+
```
|
|
226
|
+
Spent: 82 model call(s), 115407 token(s)
|
|
227
|
+
Score: 100
|
|
228
|
+
```
|
|
229
|
+
|
|
230
|
+
The figures also land in the JUnit report as `llm_calls` and `llm_tokens`, beside `score`, so a pipeline can trend cost without scraping output. A run stopped by its own budget reports the spend too — that is the case where the number is least guessable. Where a provider reports no token usage, the line says so rather than showing zero.
|
|
231
|
+
|
|
232
|
+
`--dry-run` reports the ceiling before you spend anything. Read it as a maximum and nothing more: for this repository's own suite it says 735 calls where a real run spends 82. Size a budget from what your runs actually report, not from the ceiling.
|
|
233
|
+
|
|
234
|
+
For an order of magnitude, this repository's own suite — 7 tests, 31 steps, an authenticated demo shop, `anthropic/claude-haiku-4.5` — spends **about 82 model calls and 115k tokens**, taking 156s serially or 68s at `--concurrency 4`. That is a number you can reproduce (`node examples/demo-app/serve.mjs 4173` and `blastproof run`), not a forecast for your suite: cost scales with steps, page density and how often the agent has to retry. Run yours once and read the `Spent:` line.
|
|
186
235
|
|
|
187
236
|
## Testing behind a login
|
|
188
237
|
|
package/dist/cli.js
CHANGED
|
@@ -210,6 +210,7 @@ var RunBudget = class {
|
|
|
210
210
|
startedAt;
|
|
211
211
|
calls = 0;
|
|
212
212
|
tokens = 0;
|
|
213
|
+
callsWithUsage = 0;
|
|
213
214
|
constructor(options = {}) {
|
|
214
215
|
this.maxCalls = options.maxCalls;
|
|
215
216
|
this.maxTokens = options.maxTokens;
|
|
@@ -248,7 +249,25 @@ var RunBudget = class {
|
|
|
248
249
|
/** Records what a completed model call spent. */
|
|
249
250
|
record(usage) {
|
|
250
251
|
this.calls += 1;
|
|
251
|
-
if (usage?.totalTokens !== void 0)
|
|
252
|
+
if (usage?.totalTokens !== void 0) {
|
|
253
|
+
this.tokens += usage.totalTokens;
|
|
254
|
+
this.callsWithUsage += 1;
|
|
255
|
+
}
|
|
256
|
+
}
|
|
257
|
+
/**
|
|
258
|
+
* What this budget has been spent on, for every surface that reports it
|
|
259
|
+
* (design report-what-it-spent, D1). One method rather than four getters read
|
|
260
|
+
* separately: the console, the JUnit report and the HTML report take the same
|
|
261
|
+
* numbers from the same object, so they cannot disagree about what a run cost.
|
|
262
|
+
*/
|
|
263
|
+
spend() {
|
|
264
|
+
return {
|
|
265
|
+
calls: this.calls,
|
|
266
|
+
maxCalls: this.maxCalls,
|
|
267
|
+
tokens: this.tokens,
|
|
268
|
+
callsWithUsage: this.callsWithUsage,
|
|
269
|
+
maxTokens: this.maxTokens
|
|
270
|
+
};
|
|
252
271
|
}
|
|
253
272
|
};
|
|
254
273
|
function estimateMaxModelCalls(tests, maxIterationsPerStep, maxRetriesPerStep) {
|
|
@@ -417,7 +436,8 @@ async function performAction(page, action, ctx) {
|
|
|
417
436
|
const url = new URL(value, ctx.baseUrl);
|
|
418
437
|
assertAllowedOrigin(url, ctx);
|
|
419
438
|
await page.goto(url.toString(), { timeout: ctx.resolveTimeoutMs ?? 3e4 });
|
|
420
|
-
|
|
439
|
+
const landed = page.url();
|
|
440
|
+
return landed === url.toString() ? `ok: navigated to ${url.toString()}` : `ok: navigated to ${url.toString()}, which redirected to ${landed}`;
|
|
421
441
|
}
|
|
422
442
|
case "click": {
|
|
423
443
|
const target = requireTarget(action);
|
|
@@ -623,11 +643,21 @@ async function executeTest(page, test, options) {
|
|
|
623
643
|
}
|
|
624
644
|
if (action.action === "assert") {
|
|
625
645
|
const expectation = action.expectation ?? action.reasoning;
|
|
626
|
-
let judgment = await brain.judge(
|
|
646
|
+
let judgment = await brain.judge(
|
|
647
|
+
mask(step),
|
|
648
|
+
mask(expectation),
|
|
649
|
+
mask(snap),
|
|
650
|
+
recovery.stepHistory()
|
|
651
|
+
);
|
|
627
652
|
if (!judgment.pass) {
|
|
628
653
|
await waitForSettled(page);
|
|
629
654
|
const freshSnap = await takeSnapshot(page);
|
|
630
|
-
judgment = await brain.judge(
|
|
655
|
+
judgment = await brain.judge(
|
|
656
|
+
mask(step),
|
|
657
|
+
mask(expectation),
|
|
658
|
+
mask(freshSnap),
|
|
659
|
+
recovery.stepHistory()
|
|
660
|
+
);
|
|
631
661
|
}
|
|
632
662
|
const result = judgment.pass ? `ok: assertion passed: ${judgment.reason}` : `assertion failed: ${judgment.reason}`;
|
|
633
663
|
emitAction(index, action, result);
|
|
@@ -947,6 +977,13 @@ var configSchema = z.object({
|
|
|
947
977
|
allowed_origins: z.array(z.string().url()).optional(),
|
|
948
978
|
auth: authSchema.optional(),
|
|
949
979
|
max_retries_per_step: z.number().int().min(1).default(3),
|
|
980
|
+
/**
|
|
981
|
+
* How many tests may run at once (design tests-in-parallel, D1). Defaults to
|
|
982
|
+
* 1 — tests are journeys driven against one running application, and whether
|
|
983
|
+
* two of them can run at the same time is a property of that application and
|
|
984
|
+
* those tests, not of the runner. Opted into by the person who knows.
|
|
985
|
+
*/
|
|
986
|
+
concurrency: z.number().int().min(1, "concurrency must be at least 1").default(1),
|
|
950
987
|
budget: budgetSchema.optional()
|
|
951
988
|
});
|
|
952
989
|
var ENV_OVERRIDES = {
|
|
@@ -1177,13 +1214,24 @@ A value sitting in a control that was just typed into \u2014 an open dialog's te
|
|
|
1177
1214
|
|
|
1178
1215
|
A step that names an ACTION (submit, click, create, add, ...) is satisfied by evidence the action took effect, not by the action's own control still being on the page. A successful action ordinarily replaces or moves past exactly the form, button or field the step names, so that control's absence is normal evidence of success, not evidence the step is unverifiable \u2014 do not fail such a step only because you can no longer see the thing it names. Fail it instead when the snapshot shows the action did NOT take effect: an error message, a validation warning, or the very same pre-action page still in front of you with nothing changed. A different page, a new state, or the result the action was meant to produce counts as evidence it worked.
|
|
1179
1216
|
|
|
1217
|
+
You may also be shown the actions already performed in this step, with their results. That record tells you what was ATTEMPTED and what it produced \u2014 for instance that a navigation was performed and which URL the server ultimately served, or that a form was submitted. Use it to avoid concluding that something never happened when the page simply cannot show it any more: a navigation the server redirected does not leave the browser at the path that was requested, and that is what success looks like, not failure.
|
|
1218
|
+
|
|
1219
|
+
The record is not evidence that the step's outcome holds. An action reported as \`ok\` establishes that it ran and what it returned; whether the thing the step describes is now TRUE is still decided by the snapshot alone. Never pass a step because the record shows an action succeeded while the snapshot does not show the outcome.
|
|
1220
|
+
|
|
1180
1221
|
Be strict about what the step asks, not about withholding a pass you can plainly see is earned. Answer with pass=true/false and a one-sentence reason.
|
|
1181
1222
|
|
|
1182
1223
|
\`***\` marks a secret deliberately withheld from you \u2014 a password, token or key. Seeing it is expected. A field holding \`***\` is filled, not empty, so do not fail a step on the grounds that a value was redacted. This applies only to the redaction itself: everything else the step asks for must still be visibly satisfied by the snapshot, and a step you genuinely cannot check against what you were shown still fails.`;
|
|
1183
1224
|
}
|
|
1184
|
-
function assertUserPrompt(step, expectation, snapshot) {
|
|
1185
|
-
|
|
1186
|
-
|
|
1225
|
+
function assertUserPrompt(step, expectation, snapshot, stepHistory) {
|
|
1226
|
+
const parts = [`Step under test: ${step}`];
|
|
1227
|
+
if (stepHistory && stepHistory.length > 0) {
|
|
1228
|
+
parts.push(
|
|
1229
|
+
"",
|
|
1230
|
+
"Actions already performed in this step, with their results (what was DONE \u2014 not evidence of what is now true):",
|
|
1231
|
+
...stepHistory.map((entry, i) => `${i + 1}. ${entry.action} -> ${entry.result}`)
|
|
1232
|
+
);
|
|
1233
|
+
}
|
|
1234
|
+
parts.push(
|
|
1187
1235
|
"",
|
|
1188
1236
|
`Model's expectation (the claim offered in support of the step, not the question itself): ${expectation}`,
|
|
1189
1237
|
"",
|
|
@@ -1191,7 +1239,8 @@ function assertUserPrompt(step, expectation, snapshot) {
|
|
|
1191
1239
|
snapshot,
|
|
1192
1240
|
"",
|
|
1193
1241
|
"Does the snapshot establish that the step's own outcome holds?"
|
|
1194
|
-
|
|
1242
|
+
);
|
|
1243
|
+
return parts.join("\n");
|
|
1195
1244
|
}
|
|
1196
1245
|
function plannerSystemPrompt() {
|
|
1197
1246
|
return `You are a QA engineer writing one end-to-end test for a web page, in plain English.
|
|
@@ -1280,12 +1329,12 @@ function createBrain(model, generate = generateObject, budget) {
|
|
|
1280
1329
|
}
|
|
1281
1330
|
return parsed.data;
|
|
1282
1331
|
},
|
|
1283
|
-
async judge(step, expectation, snapshot) {
|
|
1332
|
+
async judge(step, expectation, snapshot, stepHistory) {
|
|
1284
1333
|
const result = await countedGenerate(generate, budget, {
|
|
1285
1334
|
model,
|
|
1286
1335
|
schema: assertJudgmentSchema,
|
|
1287
1336
|
system: assertSystemPrompt(),
|
|
1288
|
-
prompt: assertUserPrompt(step, expectation, snapshot)
|
|
1337
|
+
prompt: assertUserPrompt(step, expectation, snapshot, stepHistory)
|
|
1289
1338
|
});
|
|
1290
1339
|
const parsed = assertJudgmentSchema.safeParse(result.object);
|
|
1291
1340
|
if (!parsed.success) {
|
|
@@ -1626,6 +1675,43 @@ function printPreflightFailures(failures) {
|
|
|
1626
1675
|
for (const failure of failures) console.error(` - ${failure}`);
|
|
1627
1676
|
}
|
|
1628
1677
|
|
|
1678
|
+
// src/report/score.ts
|
|
1679
|
+
var WEIGHTS = { P0: 3, P1: 2, P2: 1 };
|
|
1680
|
+
var DEFAULT_WEIGHT = WEIGHTS.P1;
|
|
1681
|
+
function computeScore(results) {
|
|
1682
|
+
let total = 0;
|
|
1683
|
+
let passed = 0;
|
|
1684
|
+
for (const result of results) {
|
|
1685
|
+
if (result.status === "not-run") continue;
|
|
1686
|
+
const weight = WEIGHTS[result.priority] ?? DEFAULT_WEIGHT;
|
|
1687
|
+
total += weight;
|
|
1688
|
+
if (result.status === "passed") passed += weight;
|
|
1689
|
+
}
|
|
1690
|
+
if (total === 0) return 100;
|
|
1691
|
+
return Math.round(100 * passed / total);
|
|
1692
|
+
}
|
|
1693
|
+
function formatScoreLine(score, results, threshold) {
|
|
1694
|
+
if (results.length === 0) {
|
|
1695
|
+
const base = "Score: 100 (no tests executed)";
|
|
1696
|
+
return threshold === void 0 ? base : `${base} \u2014 min-score ${threshold}: pass`;
|
|
1697
|
+
}
|
|
1698
|
+
if (threshold === void 0) return `Score: ${score}`;
|
|
1699
|
+
return score >= threshold ? `Score: ${score} \u2014 min-score ${threshold}: pass` : `Score: ${score} \u2014 min-score ${threshold}: FAIL (below threshold)`;
|
|
1700
|
+
}
|
|
1701
|
+
function formatIncompleteLine(score, reason) {
|
|
1702
|
+
return `Run incomplete: ${reason}
|
|
1703
|
+
Score over executed tests: ${score} (not a verdict \u2014 exit code 1 regardless of --min-score)`;
|
|
1704
|
+
}
|
|
1705
|
+
function formatSpendLine(spend) {
|
|
1706
|
+
const calls = spend.maxCalls === void 0 ? `${spend.calls} model call(s)` : `${spend.calls} of ${spend.maxCalls} model call(s)`;
|
|
1707
|
+
if (spend.callsWithUsage === 0) {
|
|
1708
|
+
return `Spent: ${calls}; token usage not reported by the provider`;
|
|
1709
|
+
}
|
|
1710
|
+
const tokens = spend.maxTokens === void 0 ? `${spend.tokens} token(s)` : `${spend.tokens} of ${spend.maxTokens} token(s)`;
|
|
1711
|
+
const coverage = spend.callsWithUsage < spend.calls ? ` (tokens reported by ${spend.callsWithUsage} of ${spend.calls} call(s))` : "";
|
|
1712
|
+
return `Spent: ${calls}, ${tokens}${coverage}`;
|
|
1713
|
+
}
|
|
1714
|
+
|
|
1629
1715
|
// src/commands/run.ts
|
|
1630
1716
|
import path9 from "path";
|
|
1631
1717
|
|
|
@@ -1770,7 +1856,7 @@ async function renderHtml(results, skipped, meta) {
|
|
|
1770
1856
|
<body>
|
|
1771
1857
|
<main>
|
|
1772
1858
|
<h1>blastproof report</h1>
|
|
1773
|
-
<p class="sub">${escapeHtml(generatedAt)} \xB7 ${seconds(meta.durationMs)}</p>
|
|
1859
|
+
<p class="sub">${escapeHtml(generatedAt)} \xB7 ${seconds(meta.durationMs)}${meta.spend ? ` \xB7 ${escapeHtml(formatSpendLine(meta.spend))}` : ""}</p>
|
|
1774
1860
|
|
|
1775
1861
|
${banner} <section class="score">
|
|
1776
1862
|
<b>${meta.score}</b>
|
|
@@ -1824,6 +1910,11 @@ function renderJUnit(results, skipped, meta) {
|
|
|
1824
1910
|
`<testsuite name="blastproof" tests="${results.length + skipped.length}" failures="${failures}" skipped="${skipped.length + notRun.length}" time="${seconds2(meta.durationMs)}">`,
|
|
1825
1911
|
" <properties>",
|
|
1826
1912
|
` <property name="score" value="${meta.score}"/>`,
|
|
1913
|
+
...meta.spend ? [` <property name="llm_calls" value="${meta.spend.calls}"/>`] : [],
|
|
1914
|
+
// Omitted rather than emitted as zero when no call reported usage: a
|
|
1915
|
+
// property carrying 0 would be read by a pipeline as "this run spent no
|
|
1916
|
+
// tokens", which is a different claim from "the provider did not say".
|
|
1917
|
+
...meta.spend && meta.spend.callsWithUsage > 0 ? [` <property name="llm_tokens" value="${meta.spend.tokens}"/>`] : [],
|
|
1827
1918
|
...meta.incomplete !== void 0 ? [
|
|
1828
1919
|
' <property name="incomplete" value="true"/>',
|
|
1829
1920
|
` <property name="incomplete_reason" value="${escapeXml(meta.incomplete)}"/>`
|
|
@@ -1867,32 +1958,26 @@ async function writeJUnit(file, xml) {
|
|
|
1867
1958
|
return file;
|
|
1868
1959
|
}
|
|
1869
1960
|
|
|
1870
|
-
// src/
|
|
1871
|
-
|
|
1872
|
-
|
|
1873
|
-
|
|
1874
|
-
let
|
|
1875
|
-
|
|
1876
|
-
|
|
1877
|
-
|
|
1878
|
-
|
|
1879
|
-
|
|
1880
|
-
|
|
1881
|
-
}
|
|
1882
|
-
|
|
1883
|
-
|
|
1884
|
-
|
|
1885
|
-
function formatScoreLine(score, results, threshold) {
|
|
1886
|
-
if (results.length === 0) {
|
|
1887
|
-
const base = "Score: 100 (no tests executed)";
|
|
1888
|
-
return threshold === void 0 ? base : `${base} \u2014 min-score ${threshold}: pass`;
|
|
1961
|
+
// src/runner/pool.ts
|
|
1962
|
+
async function runWithConcurrency(items, concurrency, run) {
|
|
1963
|
+
if (concurrency < 1) throw new RangeError(`concurrency must be at least 1, got ${concurrency}`);
|
|
1964
|
+
const results = new Array(items.length);
|
|
1965
|
+
let next = 0;
|
|
1966
|
+
const worker = async () => {
|
|
1967
|
+
while (true) {
|
|
1968
|
+
const index = next++;
|
|
1969
|
+
if (index >= items.length) return;
|
|
1970
|
+
results[index] = await run(items[index], index);
|
|
1971
|
+
}
|
|
1972
|
+
};
|
|
1973
|
+
const workers = [];
|
|
1974
|
+
for (let i = 0; i < Math.min(concurrency, items.length); i++) {
|
|
1975
|
+
workers.push(worker());
|
|
1889
1976
|
}
|
|
1890
|
-
|
|
1891
|
-
|
|
1892
|
-
|
|
1893
|
-
|
|
1894
|
-
return `Run incomplete: ${reason}
|
|
1895
|
-
Score over executed tests: ${score} (not a verdict \u2014 exit code 1 regardless of --min-score)`;
|
|
1977
|
+
const settled = await Promise.allSettled(workers);
|
|
1978
|
+
const failure = settled.find((outcome) => outcome.status === "rejected");
|
|
1979
|
+
if (failure && failure.status === "rejected") throw failure.reason;
|
|
1980
|
+
return results;
|
|
1896
1981
|
}
|
|
1897
1982
|
|
|
1898
1983
|
// src/runner/selection.ts
|
|
@@ -1983,19 +2068,19 @@ function notRunResults(tests, reason) {
|
|
|
1983
2068
|
durationMs: 0
|
|
1984
2069
|
}));
|
|
1985
2070
|
}
|
|
1986
|
-
function printEvent(event) {
|
|
2071
|
+
function printEvent(event, write = console.log) {
|
|
1987
2072
|
switch (event.type) {
|
|
1988
2073
|
case "step-start":
|
|
1989
|
-
|
|
2074
|
+
write(` ${event.setup ? "(setup) " : ""}step ${event.index + 1}/${event.total}: ${event.step}`);
|
|
1990
2075
|
break;
|
|
1991
2076
|
case "action": {
|
|
1992
2077
|
const { action, result } = event;
|
|
1993
|
-
|
|
2078
|
+
write(` -> ${describeAction(action)} :: ${result}`);
|
|
1994
2079
|
break;
|
|
1995
2080
|
}
|
|
1996
2081
|
case "step-end":
|
|
1997
2082
|
if (event.status === "failed") {
|
|
1998
|
-
|
|
2083
|
+
write(` X step failed: ${event.reason ?? "unknown reason"}`);
|
|
1999
2084
|
}
|
|
2000
2085
|
break;
|
|
2001
2086
|
}
|
|
@@ -2040,7 +2125,7 @@ ${notRun.length} test(s) not run (run stopped by its budget or deadline):`);
|
|
|
2040
2125
|
for (const r of notRun) console.log(` - ${r.summary} (${r.file})`);
|
|
2041
2126
|
}
|
|
2042
2127
|
}
|
|
2043
|
-
async function runOne(browser, test, config, sessionDir, session, runMask, budget) {
|
|
2128
|
+
async function runOne(browser, test, config, sessionDir, session, runMask, budget, write) {
|
|
2044
2129
|
const brain = createBrain(createModel(config.llm).model, void 0, budget);
|
|
2045
2130
|
let resolved;
|
|
2046
2131
|
const mask = runMask;
|
|
@@ -2075,7 +2160,7 @@ async function runOne(browser, test, config, sessionDir, session, runMask, budge
|
|
|
2075
2160
|
timeoutMs: config.browser.timeout_ms,
|
|
2076
2161
|
maxSnapshotLines: config.browser.max_snapshot_lines,
|
|
2077
2162
|
mask: (text) => mask.mask(text),
|
|
2078
|
-
onEvent: printEvent
|
|
2163
|
+
onEvent: (event) => printEvent(event, write)
|
|
2079
2164
|
});
|
|
2080
2165
|
} finally {
|
|
2081
2166
|
await context.close();
|
|
@@ -2106,8 +2191,9 @@ function printImpactReport(impact, selection, cwd) {
|
|
|
2106
2191
|
}
|
|
2107
2192
|
console.log("---------------------------------------------------------------");
|
|
2108
2193
|
}
|
|
2109
|
-
async function finalize(results, skipped, options, sessionDir, durationMs, impact, incomplete) {
|
|
2194
|
+
async function finalize(results, skipped, options, sessionDir, durationMs, impact, incomplete, spend) {
|
|
2110
2195
|
if (results.length > 0) printSummary(results);
|
|
2196
|
+
if (spend) console.log(formatSpendLine(spend));
|
|
2111
2197
|
const score = computeScore(results);
|
|
2112
2198
|
console.log(
|
|
2113
2199
|
incomplete ? formatIncompleteLine(score, incomplete.message) : formatScoreLine(score, results, options.minScore)
|
|
@@ -2118,7 +2204,8 @@ async function finalize(results, skipped, options, sessionDir, durationMs, impac
|
|
|
2118
2204
|
score,
|
|
2119
2205
|
durationMs,
|
|
2120
2206
|
cwd: options.cwd,
|
|
2121
|
-
incomplete: incomplete?.message
|
|
2207
|
+
incomplete: incomplete?.message,
|
|
2208
|
+
spend
|
|
2122
2209
|
});
|
|
2123
2210
|
await writeJUnit(target, xml);
|
|
2124
2211
|
console.log(`JUnit report: ${path9.relative(options.cwd, target)}`);
|
|
@@ -2130,7 +2217,8 @@ async function finalize(results, skipped, options, sessionDir, durationMs, impac
|
|
|
2130
2217
|
durationMs,
|
|
2131
2218
|
minScore: options.minScore,
|
|
2132
2219
|
cwd: options.cwd,
|
|
2133
|
-
incomplete: incomplete?.message
|
|
2220
|
+
incomplete: incomplete?.message,
|
|
2221
|
+
spend
|
|
2134
2222
|
});
|
|
2135
2223
|
await writeHtml(target, html);
|
|
2136
2224
|
console.log(`HTML report: ${path9.relative(options.cwd, target)}`);
|
|
@@ -2317,21 +2405,38 @@ ${results.length} test file(s) could not be parsed:`);
|
|
|
2317
2405
|
if (incomplete) {
|
|
2318
2406
|
results.push(...notRunResults(selected, incomplete.message));
|
|
2319
2407
|
} else {
|
|
2320
|
-
|
|
2321
|
-
|
|
2322
|
-
|
|
2323
|
-
|
|
2408
|
+
const concurrency = options.concurrency ?? config.concurrency;
|
|
2409
|
+
const streaming = concurrency === 1;
|
|
2410
|
+
let stoppedBy;
|
|
2411
|
+
const outcomes = await runWithConcurrency(selected, concurrency, async (test, index) => {
|
|
2412
|
+
if (stoppedBy) return { index, stopped: stoppedBy };
|
|
2413
|
+
const lines = [];
|
|
2414
|
+
const header = `
|
|
2415
|
+
> ${test.summary} [${test.priority}] (${path9.relative(options.cwd, test.path)})`;
|
|
2416
|
+
const write = streaming ? console.log : (line) => lines.push(line);
|
|
2417
|
+
if (streaming) console.log(header);
|
|
2418
|
+
else lines.push(header);
|
|
2324
2419
|
try {
|
|
2325
2420
|
budget.check();
|
|
2326
|
-
|
|
2421
|
+
const result = await runOne(browser, test, config, sessionDir, session, runMask, budget, write);
|
|
2422
|
+
if (!streaming) for (const line of lines) console.log(line);
|
|
2423
|
+
return { index, result };
|
|
2327
2424
|
} catch (error) {
|
|
2328
2425
|
if (error instanceof BudgetExhaustedError) {
|
|
2329
|
-
|
|
2330
|
-
|
|
2331
|
-
|
|
2426
|
+
stoppedBy ??= error;
|
|
2427
|
+
if (!streaming) for (const line of lines) console.log(line);
|
|
2428
|
+
return { index, stopped: error };
|
|
2332
2429
|
}
|
|
2333
2430
|
throw error;
|
|
2334
2431
|
}
|
|
2432
|
+
});
|
|
2433
|
+
for (const outcome of outcomes) {
|
|
2434
|
+
if (outcome.result) {
|
|
2435
|
+
results.push(outcome.result);
|
|
2436
|
+
continue;
|
|
2437
|
+
}
|
|
2438
|
+
incomplete ??= outcome.stopped;
|
|
2439
|
+
results.push(...notRunResults([selected[outcome.index]], outcome.stopped.message));
|
|
2335
2440
|
}
|
|
2336
2441
|
}
|
|
2337
2442
|
} finally {
|
|
@@ -2344,7 +2449,13 @@ ${results.length} test file(s) could not be parsed:`);
|
|
|
2344
2449
|
sessionDir,
|
|
2345
2450
|
Date.now() - startedAt,
|
|
2346
2451
|
impact,
|
|
2347
|
-
incomplete
|
|
2452
|
+
incomplete,
|
|
2453
|
+
// The command that constructed the budget is the one that reports it
|
|
2454
|
+
// (design report-what-it-spent, D2). `test` hands one budget to both of its
|
|
2455
|
+
// phases deliberately, so reporting at the point of use rather than the
|
|
2456
|
+
// point of ownership would print the running total twice for a single
|
|
2457
|
+
// allowance, the second line silently including the first.
|
|
2458
|
+
options.budget === void 0 ? budget.spend() : void 0
|
|
2348
2459
|
);
|
|
2349
2460
|
}
|
|
2350
2461
|
|
|
@@ -2574,6 +2685,7 @@ ${renderTestYaml(draft, { route, base })}`);
|
|
|
2574
2685
|
for (const route of notAttempted) console.log(` ${route}`);
|
|
2575
2686
|
}
|
|
2576
2687
|
}
|
|
2688
|
+
if (options.budget === void 0) console.log(formatSpendLine(budget.spend()));
|
|
2577
2689
|
console.log("---------------------------------------------------------------");
|
|
2578
2690
|
return incomplete || failed.length > 0 ? EXIT_FAILED : EXIT_OK;
|
|
2579
2691
|
}
|
|
@@ -2618,6 +2730,8 @@ async function testCommand(options) {
|
|
|
2618
2730
|
budget
|
|
2619
2731
|
});
|
|
2620
2732
|
if (planCode === EXIT_USAGE) return EXIT_USAGE;
|
|
2733
|
+
console.log(`
|
|
2734
|
+
${formatSpendLine(budget.spend())}`);
|
|
2621
2735
|
console.log(
|
|
2622
2736
|
"\nDrafts are not executed and do not affect the score \u2014 review them before they join the suite."
|
|
2623
2737
|
);
|
|
@@ -2660,7 +2774,7 @@ function parsePositiveNumber(flag) {
|
|
|
2660
2774
|
};
|
|
2661
2775
|
}
|
|
2662
2776
|
var program = new Command();
|
|
2663
|
-
program.name("blastproof").description("Open-source AI testing agent: plain-English YAML tests executed agentically on a real browser.").version("0.
|
|
2777
|
+
program.name("blastproof").description("Open-source AI testing agent: plain-English YAML tests executed agentically on a real browser.").version("0.10.0");
|
|
2664
2778
|
program.command("init").description("Scaffold .blastproof/ (config, tests, sample tests) in the current directory").action(async () => {
|
|
2665
2779
|
try {
|
|
2666
2780
|
const result = await initProject(process.cwd());
|
|
@@ -2676,6 +2790,10 @@ program.command("run").description("Discover and run all tests under .blastproof
|
|
|
2676
2790
|
"require a weighted score of at least n (0-100); replaces the all-must-pass rule",
|
|
2677
2791
|
parseMinScore
|
|
2678
2792
|
).option("--junit [path]", "write a JUnit XML report (default: .blastproof/reports/<session>/junit.xml)").option("--html [path]", "write a self-contained HTML report (default: .blastproof/reports/<session>/report.html)").option("--fail-on-unmapped", "fail when a changed file matches no routes: or ignore: glob").option(
|
|
2793
|
+
"--concurrency <n>",
|
|
2794
|
+
"run this many tests at once (overrides config; default 1 \u2014 see the README on when this is safe)",
|
|
2795
|
+
parsePositiveInt("--concurrency")
|
|
2796
|
+
).option(
|
|
2679
2797
|
"--max-llm-calls <n>",
|
|
2680
2798
|
"stop the run after this many model calls, reported as incomplete (overrides config)",
|
|
2681
2799
|
parsePositiveInt("--max-llm-calls")
|
|
@@ -2703,6 +2821,7 @@ program.command("run").description("Discover and run all tests under .blastproof
|
|
|
2703
2821
|
junit: options.junit,
|
|
2704
2822
|
html: options.html,
|
|
2705
2823
|
failOnUnmapped: options.failOnUnmapped,
|
|
2824
|
+
concurrency: options.concurrency,
|
|
2706
2825
|
maxLlmCalls: options.maxLlmCalls,
|
|
2707
2826
|
maxTokens: options.maxTokens,
|
|
2708
2827
|
maxDuration: options.maxDuration
|