blastproof 0.15.0 → 0.16.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +29 -10
- package/dist/cli.js +94 -21
- package/dist/cli.js.map +1 -1
- package/package.json +1 -1
package/README.md
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
<h1>
|
|
2
2
|
<picture>
|
|
3
|
-
<source media="(prefers-color-scheme: dark)" srcset="
|
|
4
|
-
<img src="
|
|
3
|
+
<source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/hamc/blastproof/main/.github/logo-dark.svg">
|
|
4
|
+
<img src="https://raw.githubusercontent.com/hamc/blastproof/main/.github/logo-light.svg" alt="" width="30" height="30">
|
|
5
5
|
</picture>
|
|
6
6
|
blastproof
|
|
7
7
|
</h1>
|
|
@@ -18,10 +18,29 @@ git diff → impact mapping → test generation → agentic execution → report
|
|
|
18
18
|
|
|
19
19
|
100% local. MIT. Bring your own LLM key.
|
|
20
20
|
|
|
21
|
+
<img src="https://raw.githubusercontent.com/hamc/blastproof/main/.github/dogfood.gif" width="100%"
|
|
22
|
+
alt="A terminal running blastproof against this repo's demo app. It passes four steps, then fails
|
|
23
|
+
step 5: the page says the SAVE20 promo gave 20% off, but the discount shown is -$6.00 where
|
|
24
|
+
20% of $120.00 would be -$24.00. Score: 0.">
|
|
25
|
+
|
|
21
26
|
▶ **[Watch the introduction](https://www.youtube.com/shorts/miqN5FzMF_k)** — what it does, in a minute.
|
|
22
27
|
|
|
23
28
|
**Documentation:** [Configuration](./docs/configuration.md) · [Testing behind a login](./docs/auth.md) · [Running in CI](./docs/ci.md) · [Contributing](./CONTRIBUTING.md) · [Architecture](./AGENTS.md)
|
|
24
29
|
|
|
30
|
+
## Set it up with your coding agent
|
|
31
|
+
|
|
32
|
+
Most of what the setup asks — which stack, which port, which journeys matter — is a question your coding agent can already answer by looking at your project. So let it answer them:
|
|
33
|
+
|
|
34
|
+
```bash
|
|
35
|
+
npx skills add hamc/blastproof
|
|
36
|
+
```
|
|
37
|
+
|
|
38
|
+
Then tell your agent: **"set up e2e tests"**. The skill installs for whichever agents you have — Claude Code, Cursor, Codex and others — and walks the whole path: check whether your markup is reachable at all, pick a provider, scaffold, generate drafts from your *running* app, curate them, run them, and write down the accessibility constraints that keep the suite alive as the project grows.
|
|
39
|
+
|
|
40
|
+
Two things it deliberately will not do, so their absence does not read as a defect: it does not configure [authentication](./docs/auth.md), and it does not wire up [CI](./docs/ci.md). Both are worth doing after you have seen a green run on your own machine.
|
|
41
|
+
|
|
42
|
+
The skill lives in [`skills/blastproof/`](./skills/blastproof/) and is worth reading even if you never install it. [`references/authoring.md`](./skills/blastproof/references/authoring.md) sets out what separates a test that detects something from one that reports Score 100 and detects nothing — including two drafts this tool generated that passed unedited and were worth nothing.
|
|
43
|
+
|
|
25
44
|
## Quick start
|
|
26
45
|
|
|
27
46
|
```bash
|
|
@@ -89,7 +108,7 @@ It is **not** a guarantee of zero duplicate writes. An agent that reaches the sa
|
|
|
89
108
|
|
|
90
109
|
## How it works
|
|
91
110
|
|
|
92
|
-
<img src="
|
|
111
|
+
<img src="https://raw.githubusercontent.com/hamc/blastproof/main/.github/coverage-flow.svg" width="100%"
|
|
93
112
|
alt="How a diff becomes a merge decision. In CI, unattended: changed files are matched against the routes: and ignore: globs; matched files contribute affected routes, files matching neither are reported as unclassified and fail the run only under --fail-on-unmapped. Tests declaring an affected route are executed and produce a weighted score, which --min-score gates on. An affected route no test declares is reported as a coverage gap and never fails the run. Separately and manually, outside CI: blastproof plan loads such a route in Chromium, makes one model call, and produces a YAML draft you review, edit and run before committing it.">
|
|
94
113
|
|
|
95
114
|
**The boundary in the middle is the point.** Everything above it runs unattended on every pull request and ends in an exit code. Everything below it is something you choose to run, on your machine, and review before it lands.
|
|
@@ -221,6 +240,10 @@ blastproof run --min-score 80 # one failing P2 is tolerated
|
|
|
221
240
|
|
|
222
241
|
`--min-score` **replaces** the all-must-pass rule rather than adding to it. Only executed tests count: filtered and unrouted tests are neither numerator nor denominator, and a run that executed nothing scores 100, so a docs-only PR is never blocked. JUnit carries the score as a `<property name="score">`, and unrouted tests appear as `<skipped/>` so the coverage gap shows up in CI rather than vanishing.
|
|
223
242
|
|
|
243
|
+
**The verdict behind that number is pinned, and that is not the same as reproducible.** The model call that decides whether a step passed runs at temperature 0, so the same page and the same expectation are not re-sampled into a different answer — a gate that flips on identical input teaches people to re-run until green, which is worse than no gate. What pinning does not survive: provider-side batching, floating point, and a gateway routing two calls to different providers or quantizations. Expect far less variance, not a guarantee of the same output twice.
|
|
244
|
+
|
|
245
|
+
The calls that *choose* an action, and the one that drafts a test, are deliberately left free — that latitude is what re-resolves a control after a redesign instead of failing on it.
|
|
246
|
+
|
|
224
247
|
Wiring this into a pipeline, with the gating patterns worth knowing: [Running in CI](./docs/ci.md).
|
|
225
248
|
|
|
226
249
|
## Without a browser or a key
|
|
@@ -246,17 +269,13 @@ If your application legitimately spans hosts (an identity provider, a hosted pay
|
|
|
246
269
|
|
|
247
270
|
**Your secrets stay out of prompts.** `{{env.*}}` placeholders survive intact and are substituted at the moment of typing. Every value your tests or auth recipe reference is redacted from everything else crossing into a prompt — page snapshots included — in literal and percent-encoded form. Redaction matches known values, so treat it as a strong default rather than a guarantee against a hostile app.
|
|
248
271
|
|
|
272
|
+
**A redacted value cannot be asserted on.** The masking is thorough by design, and page snapshots are not exempt — so a step that verifies text which happens to equal an `{{env.*}}` value can never pass, because the judge is shown `***` where the page shows the thing. The failure is the most misleading shape available: the test is right, the application is right, and the report blames the application. Put a value in `{{env.*}}` because it is a secret or because it varies by environment, but do not then write a step that asserts on it ([#87](https://github.com/hamc/blastproof/issues/87)).
|
|
273
|
+
|
|
249
274
|
The system prompt also tells the model that page content is data, never instruction. That raises the cost of casual injection and is **not** a boundary — the origin constraint is. Do not point blastproof at an application you would not run locally.
|
|
250
275
|
|
|
251
276
|
## blastproof tests itself
|
|
252
277
|
|
|
253
|
-
The **Dogfood** badge is blastproof running against the demo app in this repo — real Chromium, real model, scored and gated, with public logs. It catches real regressions rather than diffing strings: change the demo discount from 20% to 5% while the page still claims *"20% off"* and it
|
|
254
|
-
|
|
255
|
-
```
|
|
256
|
-
FAIL P0 Promo code SAVE20 applies a 20% discount in the cart
|
|
257
|
-
reason: the discount is currently -$6.00, but a 20% discount on
|
|
258
|
-
$120.00 should be -$24.00
|
|
259
|
-
```
|
|
278
|
+
The **Dogfood** badge is blastproof running against the demo app in this repo — real Chromium, real model, scored and gated, with public logs. It catches real regressions rather than diffing strings: change the demo discount from 20% to 5% while the page still claims *"20% off"* and it fails the step — that is the run in the GIF at the top of this page, verbatim.
|
|
260
279
|
|
|
261
280
|
No selector was updated to catch that. The agent read the value, did the arithmetic, and disagreed with the page.
|
|
262
281
|
|
package/dist/cli.js
CHANGED
|
@@ -458,7 +458,28 @@ function requireValue(action) {
|
|
|
458
458
|
}
|
|
459
459
|
return action.value;
|
|
460
460
|
}
|
|
461
|
+
var INTERCEPTS_POINTER = /(<[^>\n]{1,400}>)[^\n]{0,400}?intercepts pointer events/;
|
|
462
|
+
var MAX_BLOCKER_LENGTH = 80;
|
|
463
|
+
function truncateBlocker(tag) {
|
|
464
|
+
return tag.length <= MAX_BLOCKER_LENGTH ? tag : `${tag.slice(0, MAX_BLOCKER_LENGTH)}\u2026>`;
|
|
465
|
+
}
|
|
466
|
+
function obstructionFor(error, action) {
|
|
467
|
+
const message = error instanceof Error ? error.message : String(error);
|
|
468
|
+
const match = INTERCEPTS_POINTER.exec(message);
|
|
469
|
+
if (!match) return void 0;
|
|
470
|
+
const where = action.target ? ` on ${describeTarget(action.target)}` : "";
|
|
471
|
+
return new ActionError(
|
|
472
|
+
`blocked: the ${action.action}${where} was NOT performed. The target was found and is visible, enabled and stable \u2014 nothing about it is wrong. ${truncateBlocker(match[1])} is on top of it and received the pointer event instead. Something is covering the page: find it in the snapshot \u2014 a dialog, a cookie banner, an onboarding overlay \u2014 and dismiss it first, using its own close or accept control, or by pressing Escape with no target. Then act on this target again. Choosing a different name for the same target cannot help.`
|
|
473
|
+
);
|
|
474
|
+
}
|
|
461
475
|
async function performAction(page, action, ctx) {
|
|
476
|
+
try {
|
|
477
|
+
return await performResolvedAction(page, action, ctx);
|
|
478
|
+
} catch (error) {
|
|
479
|
+
throw obstructionFor(error, action) ?? error;
|
|
480
|
+
}
|
|
481
|
+
}
|
|
482
|
+
async function performResolvedAction(page, action, ctx) {
|
|
462
483
|
switch (action.action) {
|
|
463
484
|
case "navigate": {
|
|
464
485
|
const value = resolve(requireValue(action), ctx);
|
|
@@ -1321,6 +1342,7 @@ Rules:
|
|
|
1321
1342
|
- Return "done" when the current step's outcome holds \u2014 including when it already held before you acted, or was achieved by your previous action. "Already true" is done, never failure. Do not return "done" for work belonging to later steps.
|
|
1322
1343
|
- Return "fail" only when the step's outcome cannot be reached: the element is still absent after retries, the page cannot support the step, or an error blocks progress. Never return "fail" because the work appears to have been done already.
|
|
1323
1344
|
- If your previous action errored, re-read the fresh snapshot and choose an alternative element or approach. Do not repeat the exact same failing action.
|
|
1345
|
+
- An action reported as "blocked" is the exception to that rule: it means another element is on top of your target, not that you picked the wrong target. Re-targeting cannot fix it. Whatever is covering the page is in the snapshot \u2014 a dialog, a cookie banner, an onboarding overlay \u2014 so dismiss that first, with its own close or accept control, or by pressing Escape with no target, and then act on your original target again. Overlays can be stacked: clearing one may reveal another, and that is progress, not failure.
|
|
1324
1346
|
- Never invent a value. A value you type must come from the step, from the page, or from an {{env.*}} placeholder. This one is enforced, not merely asked: a fill or select whose value is in none of those is refused and not performed. If a step needs a value it does not give you, that is a failing step, not a gap for you to fill in.
|
|
1325
1347
|
- A record of the actions you already performed in this step may be shown to you. It is the ground truth about what happened, even when the page no longer shows it: a form that submitted successfully and came back empty looks exactly like one you never submitted. Do not redo work that record says you already did.
|
|
1326
1348
|
- \`***\` in a snapshot is a redacted secret \u2014 a password, token or key deliberately withheld from you. Seeing it is expected and is not a problem. A field showing \`***\` after you filled it from an {{env.VAR}} placeholder means the fill worked; treat that as success and move on. Never retry a fill because its value is redacted, and never report failure because a value was withheld.
|
|
@@ -1414,19 +1436,49 @@ function plannerUserPrompt(input) {
|
|
|
1414
1436
|
|
|
1415
1437
|
// src/llm/schemas.ts
|
|
1416
1438
|
import { z as z2 } from "zod";
|
|
1439
|
+
function absentAsNull(schema) {
|
|
1440
|
+
return schema.nullable().transform((value) => value ?? void 0);
|
|
1441
|
+
}
|
|
1442
|
+
var ACTION_NAMES = ["navigate", "click", "fill", "press", "select", "assert", "done", "fail"];
|
|
1417
1443
|
var agentActionSchema = z2.object({
|
|
1418
|
-
action: z2.enum(
|
|
1419
|
-
target:
|
|
1420
|
-
|
|
1421
|
-
|
|
1422
|
-
|
|
1423
|
-
|
|
1424
|
-
|
|
1425
|
-
|
|
1444
|
+
action: z2.enum(ACTION_NAMES).describe("The next browser action to perform for the current step."),
|
|
1445
|
+
target: absentAsNull(
|
|
1446
|
+
z2.object({
|
|
1447
|
+
role: absentAsNull(
|
|
1448
|
+
z2.string().describe('ARIA role of the target element, e.g. "button", "link", "textbox".')
|
|
1449
|
+
),
|
|
1450
|
+
name: absentAsNull(
|
|
1451
|
+
z2.string().describe("Accessible name of the target element, exactly as shown in the snapshot.")
|
|
1452
|
+
),
|
|
1453
|
+
text: absentAsNull(
|
|
1454
|
+
z2.string().describe("Visible text of the target element, used as a fallback when no role matches.")
|
|
1455
|
+
)
|
|
1456
|
+
})
|
|
1457
|
+
).describe("Element to act on, resolved from the accessibility snapshot. Null for navigate/done/fail."),
|
|
1458
|
+
value: absentAsNull(
|
|
1459
|
+
z2.string().describe(
|
|
1460
|
+
'Action payload: URL/path for navigate, text for fill, key for press (e.g. "Enter"), option label for select.'
|
|
1461
|
+
)
|
|
1426
1462
|
),
|
|
1427
1463
|
reasoning: z2.string().describe("One sentence explaining why this action moves the step forward."),
|
|
1428
|
-
expectation:
|
|
1464
|
+
expectation: absentAsNull(
|
|
1465
|
+
z2.string().describe("For assert: the condition the current page snapshot must satisfy.")
|
|
1466
|
+
)
|
|
1429
1467
|
});
|
|
1468
|
+
var parsedAgentActionSchema = z2.object({
|
|
1469
|
+
action: z2.enum(ACTION_NAMES),
|
|
1470
|
+
target: z2.object({
|
|
1471
|
+
role: z2.string().optional(),
|
|
1472
|
+
name: z2.string().optional(),
|
|
1473
|
+
text: z2.string().optional()
|
|
1474
|
+
}).optional(),
|
|
1475
|
+
value: z2.string().optional(),
|
|
1476
|
+
reasoning: z2.string(),
|
|
1477
|
+
expectation: z2.string().optional()
|
|
1478
|
+
});
|
|
1479
|
+
function parseAgentAction(value) {
|
|
1480
|
+
return parsedAgentActionSchema.safeParse(value);
|
|
1481
|
+
}
|
|
1430
1482
|
var assertJudgmentSchema = z2.object({
|
|
1431
1483
|
pass: z2.boolean().describe("Whether the snapshot satisfies the expectation."),
|
|
1432
1484
|
reason: z2.string().describe("One sentence explaining the judgment.")
|
|
@@ -1447,9 +1499,24 @@ var MalformedModelOutputError = class extends Error {
|
|
|
1447
1499
|
this.name = "MalformedModelOutputError";
|
|
1448
1500
|
}
|
|
1449
1501
|
};
|
|
1502
|
+
var MAX_PROVIDER_DETAIL = 300;
|
|
1503
|
+
function withProviderDetail(error) {
|
|
1504
|
+
if (!(error instanceof Error)) return error;
|
|
1505
|
+
const body = error.responseBody;
|
|
1506
|
+
if (typeof body !== "string" || body.length === 0) return error;
|
|
1507
|
+
const detail = body.replace(/\s+/g, " ").trim().slice(0, MAX_PROVIDER_DETAIL);
|
|
1508
|
+
if (error.message.includes(detail)) return error;
|
|
1509
|
+
error.message = `${error.message} \u2014 provider said: ${detail}`;
|
|
1510
|
+
return error;
|
|
1511
|
+
}
|
|
1450
1512
|
async function countedGenerate(generate, budget, options) {
|
|
1451
1513
|
budget.check();
|
|
1452
|
-
|
|
1514
|
+
let result;
|
|
1515
|
+
try {
|
|
1516
|
+
result = await generate(options);
|
|
1517
|
+
} catch (error) {
|
|
1518
|
+
throw withProviderDetail(error);
|
|
1519
|
+
}
|
|
1453
1520
|
budget.record(result.usage);
|
|
1454
1521
|
return result;
|
|
1455
1522
|
}
|
|
@@ -1462,7 +1529,7 @@ function createBrain(model, generate = generateObject, budget) {
|
|
|
1462
1529
|
system: agentSystemPrompt(),
|
|
1463
1530
|
prompt: agentUserPrompt(input)
|
|
1464
1531
|
});
|
|
1465
|
-
const parsed =
|
|
1532
|
+
const parsed = parseAgentAction(result.object);
|
|
1466
1533
|
if (!parsed.success) {
|
|
1467
1534
|
throw new MalformedModelOutputError(
|
|
1468
1535
|
`Model returned an invalid action: ${parsed.error.issues[0]?.message ?? "unknown"}`
|
|
@@ -1475,7 +1542,8 @@ function createBrain(model, generate = generateObject, budget) {
|
|
|
1475
1542
|
model,
|
|
1476
1543
|
schema: assertJudgmentSchema,
|
|
1477
1544
|
system: assertSystemPrompt(),
|
|
1478
|
-
prompt: assertUserPrompt(step, expectation, snapshot, stepHistory)
|
|
1545
|
+
prompt: assertUserPrompt(step, expectation, snapshot, stepHistory),
|
|
1546
|
+
temperature: 0
|
|
1479
1547
|
});
|
|
1480
1548
|
const parsed = assertJudgmentSchema.safeParse(result.object);
|
|
1481
1549
|
if (!parsed.success) {
|
|
@@ -1560,7 +1628,7 @@ function createModel(llm, env = process.env) {
|
|
|
1560
1628
|
}
|
|
1561
1629
|
|
|
1562
1630
|
// src/planner.ts
|
|
1563
|
-
import { writeFile as writeFile3 } from "fs/promises";
|
|
1631
|
+
import { mkdir as mkdir4, writeFile as writeFile3 } from "fs/promises";
|
|
1564
1632
|
import path6 from "path";
|
|
1565
1633
|
import { stringify } from "yaml";
|
|
1566
1634
|
|
|
@@ -1687,16 +1755,21 @@ async function generateForRoute(page, options) {
|
|
|
1687
1755
|
return { ...generated, routes: [route] };
|
|
1688
1756
|
}
|
|
1689
1757
|
async function writeDraft(cwd, draft, meta) {
|
|
1690
|
-
const
|
|
1758
|
+
const dir = path6.join(cwd, TESTS_RELATIVE_DIR);
|
|
1759
|
+
const file = path6.join(dir, `${routeToSlug(meta.route)}.yaml`);
|
|
1691
1760
|
try {
|
|
1761
|
+
await mkdir4(dir, { recursive: true });
|
|
1692
1762
|
await writeFile3(file, renderTestYaml(draft, meta), { flag: "wx" });
|
|
1693
1763
|
} catch (error) {
|
|
1694
|
-
|
|
1764
|
+
const err = error;
|
|
1765
|
+
if (err.code === "EEXIST" && err.syscall !== "mkdir") {
|
|
1695
1766
|
throw new PlannerError(
|
|
1696
1767
|
`Refusing to overwrite ${path6.relative(cwd, file)} (route ${meta.route}). Delete or rename it to regenerate.`
|
|
1697
1768
|
);
|
|
1698
1769
|
}
|
|
1699
|
-
throw
|
|
1770
|
+
throw new PlannerError(
|
|
1771
|
+
`Cannot write draft to ${file}: ${fsReason(error)}. Check that ${path6.dirname(file)} is a directory you can write to, not a file.`
|
|
1772
|
+
);
|
|
1700
1773
|
}
|
|
1701
1774
|
return file;
|
|
1702
1775
|
}
|
|
@@ -1857,7 +1930,7 @@ function formatSpendLine(spend) {
|
|
|
1857
1930
|
import path9 from "path";
|
|
1858
1931
|
|
|
1859
1932
|
// src/report/html.ts
|
|
1860
|
-
import { mkdir as
|
|
1933
|
+
import { mkdir as mkdir5, readFile as readFile4, writeFile as writeFile4 } from "fs/promises";
|
|
1861
1934
|
import path7 from "path";
|
|
1862
1935
|
var ESCAPES = {
|
|
1863
1936
|
"&": "&",
|
|
@@ -2022,7 +2095,7 @@ ${detail}
|
|
|
2022
2095
|
}
|
|
2023
2096
|
async function writeHtml(file, html) {
|
|
2024
2097
|
try {
|
|
2025
|
-
await
|
|
2098
|
+
await mkdir5(path7.dirname(file), { recursive: true });
|
|
2026
2099
|
await writeFile4(file, html, "utf8");
|
|
2027
2100
|
return file;
|
|
2028
2101
|
} catch (error) {
|
|
@@ -2033,7 +2106,7 @@ async function writeHtml(file, html) {
|
|
|
2033
2106
|
}
|
|
2034
2107
|
|
|
2035
2108
|
// src/report/junit.ts
|
|
2036
|
-
import { mkdir as
|
|
2109
|
+
import { mkdir as mkdir6, writeFile as writeFile5 } from "fs/promises";
|
|
2037
2110
|
import path8 from "path";
|
|
2038
2111
|
var ESCAPES2 = {
|
|
2039
2112
|
"&": "&",
|
|
@@ -2101,7 +2174,7 @@ ${reason}` : reason;
|
|
|
2101
2174
|
}
|
|
2102
2175
|
async function writeJUnit(file, xml) {
|
|
2103
2176
|
try {
|
|
2104
|
-
await
|
|
2177
|
+
await mkdir6(path8.dirname(file), { recursive: true });
|
|
2105
2178
|
await writeFile5(file, xml, "utf8");
|
|
2106
2179
|
return file;
|
|
2107
2180
|
} catch (error) {
|
|
@@ -3006,7 +3079,7 @@ function parsePositiveNumber(flag) {
|
|
|
3006
3079
|
};
|
|
3007
3080
|
}
|
|
3008
3081
|
var program = new Command();
|
|
3009
|
-
program.name("blastproof").description("Open-source AI testing agent: plain-English YAML tests executed agentically on a real browser.").version("0.
|
|
3082
|
+
program.name("blastproof").description("Open-source AI testing agent: plain-English YAML tests executed agentically on a real browser.").version("0.16.0");
|
|
3010
3083
|
program.command("init").description("Scaffold .blastproof/ (config, tests, sample tests) in the current directory").action(async () => {
|
|
3011
3084
|
try {
|
|
3012
3085
|
const result = await initProject(process.cwd());
|