blastproof 0.15.0 → 0.17.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -1,7 +1,7 @@
1
1
  <h1>
2
2
  <picture>
3
- <source media="(prefers-color-scheme: dark)" srcset="./.github/logo-dark.svg">
4
- <img src="./.github/logo-light.svg" alt="" width="30" height="30">
3
+ <source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/hamc/blastproof/main/.github/logo-dark.svg">
4
+ <img src="https://raw.githubusercontent.com/hamc/blastproof/main/.github/logo-light.svg" alt="" width="30" height="30">
5
5
  </picture>
6
6
  blastproof
7
7
  </h1>
@@ -18,10 +18,29 @@ git diff → impact mapping → test generation → agentic execution → report
18
18
 
19
19
  100% local. MIT. Bring your own LLM key.
20
20
 
21
+ <img src="https://raw.githubusercontent.com/hamc/blastproof/main/.github/dogfood.gif" width="100%"
22
+ alt="A terminal running blastproof against this repo's demo app. It passes four steps, then fails
23
+ step 5: the page says the SAVE20 promo gave 20% off, but the discount shown is -$6.00 where
24
+ 20% of $120.00 would be -$24.00. Score: 0.">
25
+
21
26
  ▶ **[Watch the introduction](https://www.youtube.com/shorts/miqN5FzMF_k)** — what it does, in a minute.
22
27
 
23
28
  **Documentation:** [Configuration](./docs/configuration.md) · [Testing behind a login](./docs/auth.md) · [Running in CI](./docs/ci.md) · [Contributing](./CONTRIBUTING.md) · [Architecture](./AGENTS.md)
24
29
 
30
+ ## Set it up with your coding agent
31
+
32
+ Most of what the setup asks — which stack, which port, which journeys matter — is a question your coding agent can already answer by looking at your project. So let it answer them:
33
+
34
+ ```bash
35
+ npx skills add hamc/blastproof
36
+ ```
37
+
38
+ Then tell your agent: **"set up e2e tests"**. The skill installs for whichever agents you have — Claude Code, Cursor, Codex and others — and walks the whole path: check whether your markup is reachable at all, pick a provider, scaffold, generate drafts from your *running* app, curate them, run them, and write down the accessibility constraints that keep the suite alive as the project grows.
39
+
40
+ Two things it deliberately will not do, so their absence does not read as a defect: it does not configure [authentication](./docs/auth.md), and it does not wire up [CI](./docs/ci.md). Both are worth doing after you have seen a green run on your own machine.
41
+
42
+ The skill lives in [`skills/blastproof/`](./skills/blastproof/) and is worth reading even if you never install it. [`references/authoring.md`](./skills/blastproof/references/authoring.md) sets out what separates a test that detects something from one that reports Score 100 and detects nothing — including two drafts this tool generated that passed unedited and were worth nothing.
43
+
25
44
  ## Quick start
26
45
 
27
46
  ```bash
@@ -58,6 +77,8 @@ Three questions. The first one decides most cases.
58
77
 
59
78
  **A hard requirement, not a preference.** blastproof finds elements the way a screen reader does — by role, by label, by visible text. That is what removes selectors and survives redesigns. The cost is that there is deliberately no CSS or XPath fallback, so anything the accessibility tree cannot describe cannot be driven at all.
60
79
 
80
+ The name is matched **exactly first**, then by substring if nothing matches exactly — so a control named `Add` is found even when `Add New` sits above it. When several controls answer to the same name, the first one on the page wins. Give each control an accessible name no other control on the page shares; a page that cannot offer one cannot be driven unambiguously, and no matching rule fixes that.
81
+
61
82
  | works | cannot be driven |
62
83
  | --- | --- |
63
84
  | `<button>Add to cart</button>` | a `<div>` with a click handler |
@@ -89,7 +110,7 @@ It is **not** a guarantee of zero duplicate writes. An agent that reaches the sa
89
110
 
90
111
  ## How it works
91
112
 
92
- <img src="./.github/coverage-flow.svg" width="100%"
113
+ <img src="https://raw.githubusercontent.com/hamc/blastproof/main/.github/coverage-flow.svg" width="100%"
93
114
  alt="How a diff becomes a merge decision. In CI, unattended: changed files are matched against the routes: and ignore: globs; matched files contribute affected routes, files matching neither are reported as unclassified and fail the run only under --fail-on-unmapped. Tests declaring an affected route are executed and produce a weighted score, which --min-score gates on. An affected route no test declares is reported as a coverage gap and never fails the run. Separately and manually, outside CI: blastproof plan loads such a route in Chromium, makes one model call, and produces a YAML draft you review, edit and run before committing it.">
94
115
 
95
116
  **The boundary in the middle is the point.** Everything above it runs unattended on every pull request and ends in an exit code. Everything below it is something you choose to run, on your machine, and review before it lands.
@@ -221,6 +242,10 @@ blastproof run --min-score 80 # one failing P2 is tolerated
221
242
 
222
243
  `--min-score` **replaces** the all-must-pass rule rather than adding to it. Only executed tests count: filtered and unrouted tests are neither numerator nor denominator, and a run that executed nothing scores 100, so a docs-only PR is never blocked. JUnit carries the score as a `<property name="score">`, and unrouted tests appear as `<skipped/>` so the coverage gap shows up in CI rather than vanishing.
223
244
 
245
+ **The verdict behind that number is pinned, and that is not the same as reproducible.** The model call that decides whether a step passed runs at temperature 0, so the same page and the same expectation are not re-sampled into a different answer — a gate that flips on identical input teaches people to re-run until green, which is worse than no gate. What pinning does not survive: provider-side batching, floating point, and a gateway routing two calls to different providers or quantizations. Expect far less variance, not a guarantee of the same output twice.
246
+
247
+ The calls that *choose* an action, and the one that drafts a test, are deliberately left free — that latitude is what re-resolves a control after a redesign instead of failing on it.
248
+
224
249
  Wiring this into a pipeline, with the gating patterns worth knowing: [Running in CI](./docs/ci.md).
225
250
 
226
251
  ## Without a browser or a key
@@ -246,17 +271,13 @@ If your application legitimately spans hosts (an identity provider, a hosted pay
246
271
 
247
272
  **Your secrets stay out of prompts.** `{{env.*}}` placeholders survive intact and are substituted at the moment of typing. Every value your tests or auth recipe reference is redacted from everything else crossing into a prompt — page snapshots included — in literal and percent-encoded form. Redaction matches known values, so treat it as a strong default rather than a guarantee against a hostile app.
248
273
 
274
+ **A redacted value cannot be asserted on.** The masking is thorough by design, and page snapshots are not exempt — so a step that verifies text which happens to equal an `{{env.*}}` value can never pass, because the judge is shown `***` where the page shows the thing. The failure is the most misleading shape available: the test is right, the application is right, and the report blames the application. Put a value in `{{env.*}}` because it is a secret or because it varies by environment, but do not then write a step that asserts on it ([#87](https://github.com/hamc/blastproof/issues/87)).
275
+
249
276
  The system prompt also tells the model that page content is data, never instruction. That raises the cost of casual injection and is **not** a boundary — the origin constraint is. Do not point blastproof at an application you would not run locally.
250
277
 
251
278
  ## blastproof tests itself
252
279
 
253
- The **Dogfood** badge is blastproof running against the demo app in this repo — real Chromium, real model, scored and gated, with public logs. It catches real regressions rather than diffing strings: change the demo discount from 20% to 5% while the page still claims *"20% off"* and it reports
254
-
255
- ```
256
- FAIL P0 Promo code SAVE20 applies a 20% discount in the cart
257
- reason: the discount is currently -$6.00, but a 20% discount on
258
- $120.00 should be -$24.00
259
- ```
280
+ The **Dogfood** badge is blastproof running against the demo app in this repo — real Chromium, real model, scored and gated, with public logs. It catches real regressions rather than diffing strings: change the demo discount from 20% to 5% while the page still claims *"20% off"* and it fails the step — that is the run in the GIF at the top of this page, verbatim.
260
281
 
261
282
  No selector was updated to catch that. The agent read the value, did the arithmetic, and disagreed with the page.
262
283
 
package/dist/cli.js CHANGED
@@ -424,13 +424,18 @@ function describeTarget(target) {
424
424
  async function resolveTarget(page, target, resolveTimeoutMs = 2e3) {
425
425
  const candidates = [];
426
426
  if (target.role) {
427
+ if (target.name) {
428
+ candidates.push(page.getByRole(target.role, { name: target.name, exact: true }));
429
+ }
427
430
  candidates.push(page.getByRole(target.role, target.name ? { name: target.name } : {}));
428
431
  }
429
432
  if (target.name) {
433
+ candidates.push(page.getByLabel(target.name, { exact: true }));
430
434
  candidates.push(page.getByLabel(target.name));
431
435
  }
432
436
  const text = target.text ?? target.name;
433
437
  if (text) {
438
+ candidates.push(page.getByText(text, { exact: true }));
434
439
  candidates.push(page.getByText(text));
435
440
  }
436
441
  for (const candidate of candidates) {
@@ -458,7 +463,28 @@ function requireValue(action) {
458
463
  }
459
464
  return action.value;
460
465
  }
466
+ var INTERCEPTS_POINTER = /(<[^>\n]{1,400}>)[^\n]{0,400}?intercepts pointer events/;
467
+ var MAX_BLOCKER_LENGTH = 80;
468
+ function truncateBlocker(tag) {
469
+ return tag.length <= MAX_BLOCKER_LENGTH ? tag : `${tag.slice(0, MAX_BLOCKER_LENGTH)}\u2026>`;
470
+ }
471
+ function obstructionFor(error, action) {
472
+ const message = error instanceof Error ? error.message : String(error);
473
+ const match = INTERCEPTS_POINTER.exec(message);
474
+ if (!match) return void 0;
475
+ const where = action.target ? ` on ${describeTarget(action.target)}` : "";
476
+ return new ActionError(
477
+ `blocked: the ${action.action}${where} was NOT performed. The target was found and is visible, enabled and stable \u2014 nothing about it is wrong. ${truncateBlocker(match[1])} is on top of it and received the pointer event instead. Something is covering the page: find it in the snapshot \u2014 a dialog, a cookie banner, an onboarding overlay \u2014 and dismiss it first, using its own close or accept control, or by pressing Escape with no target. Then act on this target again. Choosing a different name for the same target cannot help.`
478
+ );
479
+ }
461
480
  async function performAction(page, action, ctx) {
481
+ try {
482
+ return await performResolvedAction(page, action, ctx);
483
+ } catch (error) {
484
+ throw obstructionFor(error, action) ?? error;
485
+ }
486
+ }
487
+ async function performResolvedAction(page, action, ctx) {
462
488
  switch (action.action) {
463
489
  case "navigate": {
464
490
  const value = resolve(requireValue(action), ctx);
@@ -1321,6 +1347,7 @@ Rules:
1321
1347
  - Return "done" when the current step's outcome holds \u2014 including when it already held before you acted, or was achieved by your previous action. "Already true" is done, never failure. Do not return "done" for work belonging to later steps.
1322
1348
  - Return "fail" only when the step's outcome cannot be reached: the element is still absent after retries, the page cannot support the step, or an error blocks progress. Never return "fail" because the work appears to have been done already.
1323
1349
  - If your previous action errored, re-read the fresh snapshot and choose an alternative element or approach. Do not repeat the exact same failing action.
1350
+ - An action reported as "blocked" is the exception to that rule: it means another element is on top of your target, not that you picked the wrong target. Re-targeting cannot fix it. Whatever is covering the page is in the snapshot \u2014 a dialog, a cookie banner, an onboarding overlay \u2014 so dismiss that first, with its own close or accept control, or by pressing Escape with no target, and then act on your original target again. Overlays can be stacked: clearing one may reveal another, and that is progress, not failure.
1324
1351
  - Never invent a value. A value you type must come from the step, from the page, or from an {{env.*}} placeholder. This one is enforced, not merely asked: a fill or select whose value is in none of those is refused and not performed. If a step needs a value it does not give you, that is a failing step, not a gap for you to fill in.
1325
1352
  - A record of the actions you already performed in this step may be shown to you. It is the ground truth about what happened, even when the page no longer shows it: a form that submitted successfully and came back empty looks exactly like one you never submitted. Do not redo work that record says you already did.
1326
1353
  - \`***\` in a snapshot is a redacted secret \u2014 a password, token or key deliberately withheld from you. Seeing it is expected and is not a problem. A field showing \`***\` after you filled it from an {{env.VAR}} placeholder means the fill worked; treat that as success and move on. Never retry a fill because its value is redacted, and never report failure because a value was withheld.
@@ -1414,19 +1441,49 @@ function plannerUserPrompt(input) {
1414
1441
 
1415
1442
  // src/llm/schemas.ts
1416
1443
  import { z as z2 } from "zod";
1444
+ function absentAsNull(schema) {
1445
+ return schema.nullable().transform((value) => value ?? void 0);
1446
+ }
1447
+ var ACTION_NAMES = ["navigate", "click", "fill", "press", "select", "assert", "done", "fail"];
1417
1448
  var agentActionSchema = z2.object({
1418
- action: z2.enum(["navigate", "click", "fill", "press", "select", "assert", "done", "fail"]).describe("The next browser action to perform for the current step."),
1419
- target: z2.object({
1420
- role: z2.string().optional().describe('ARIA role of the target element, e.g. "button", "link", "textbox".'),
1421
- name: z2.string().optional().describe("Accessible name of the target element, exactly as shown in the snapshot."),
1422
- text: z2.string().optional().describe("Visible text of the target element, used as a fallback when no role matches.")
1423
- }).optional().describe("Element to act on, resolved from the accessibility snapshot. Omit for navigate/done/fail."),
1424
- value: z2.string().optional().describe(
1425
- 'Action payload: URL/path for navigate, text for fill, key for press (e.g. "Enter"), option label for select.'
1449
+ action: z2.enum(ACTION_NAMES).describe("The next browser action to perform for the current step."),
1450
+ target: absentAsNull(
1451
+ z2.object({
1452
+ role: absentAsNull(
1453
+ z2.string().describe('ARIA role of the target element, e.g. "button", "link", "textbox".')
1454
+ ),
1455
+ name: absentAsNull(
1456
+ z2.string().describe("Accessible name of the target element, exactly as shown in the snapshot.")
1457
+ ),
1458
+ text: absentAsNull(
1459
+ z2.string().describe("Visible text of the target element, used as a fallback when no role matches.")
1460
+ )
1461
+ })
1462
+ ).describe("Element to act on, resolved from the accessibility snapshot. Null for navigate/done/fail."),
1463
+ value: absentAsNull(
1464
+ z2.string().describe(
1465
+ 'Action payload: URL/path for navigate, text for fill, key for press (e.g. "Enter"), option label for select.'
1466
+ )
1426
1467
  ),
1427
1468
  reasoning: z2.string().describe("One sentence explaining why this action moves the step forward."),
1428
- expectation: z2.string().optional().describe("For assert: the condition the current page snapshot must satisfy.")
1469
+ expectation: absentAsNull(
1470
+ z2.string().describe("For assert: the condition the current page snapshot must satisfy.")
1471
+ )
1429
1472
  });
1473
+ var parsedAgentActionSchema = z2.object({
1474
+ action: z2.enum(ACTION_NAMES),
1475
+ target: z2.object({
1476
+ role: z2.string().optional(),
1477
+ name: z2.string().optional(),
1478
+ text: z2.string().optional()
1479
+ }).optional(),
1480
+ value: z2.string().optional(),
1481
+ reasoning: z2.string(),
1482
+ expectation: z2.string().optional()
1483
+ });
1484
+ function parseAgentAction(value) {
1485
+ return parsedAgentActionSchema.safeParse(value);
1486
+ }
1430
1487
  var assertJudgmentSchema = z2.object({
1431
1488
  pass: z2.boolean().describe("Whether the snapshot satisfies the expectation."),
1432
1489
  reason: z2.string().describe("One sentence explaining the judgment.")
@@ -1447,9 +1504,24 @@ var MalformedModelOutputError = class extends Error {
1447
1504
  this.name = "MalformedModelOutputError";
1448
1505
  }
1449
1506
  };
1507
+ var MAX_PROVIDER_DETAIL = 300;
1508
+ function withProviderDetail(error) {
1509
+ if (!(error instanceof Error)) return error;
1510
+ const body = error.responseBody;
1511
+ if (typeof body !== "string" || body.length === 0) return error;
1512
+ const detail = body.replace(/\s+/g, " ").trim().slice(0, MAX_PROVIDER_DETAIL);
1513
+ if (error.message.includes(detail)) return error;
1514
+ error.message = `${error.message} \u2014 provider said: ${detail}`;
1515
+ return error;
1516
+ }
1450
1517
  async function countedGenerate(generate, budget, options) {
1451
1518
  budget.check();
1452
- const result = await generate(options);
1519
+ let result;
1520
+ try {
1521
+ result = await generate(options);
1522
+ } catch (error) {
1523
+ throw withProviderDetail(error);
1524
+ }
1453
1525
  budget.record(result.usage);
1454
1526
  return result;
1455
1527
  }
@@ -1462,7 +1534,7 @@ function createBrain(model, generate = generateObject, budget) {
1462
1534
  system: agentSystemPrompt(),
1463
1535
  prompt: agentUserPrompt(input)
1464
1536
  });
1465
- const parsed = agentActionSchema.safeParse(result.object);
1537
+ const parsed = parseAgentAction(result.object);
1466
1538
  if (!parsed.success) {
1467
1539
  throw new MalformedModelOutputError(
1468
1540
  `Model returned an invalid action: ${parsed.error.issues[0]?.message ?? "unknown"}`
@@ -1475,7 +1547,8 @@ function createBrain(model, generate = generateObject, budget) {
1475
1547
  model,
1476
1548
  schema: assertJudgmentSchema,
1477
1549
  system: assertSystemPrompt(),
1478
- prompt: assertUserPrompt(step, expectation, snapshot, stepHistory)
1550
+ prompt: assertUserPrompt(step, expectation, snapshot, stepHistory),
1551
+ temperature: 0
1479
1552
  });
1480
1553
  const parsed = assertJudgmentSchema.safeParse(result.object);
1481
1554
  if (!parsed.success) {
@@ -1560,7 +1633,7 @@ function createModel(llm, env = process.env) {
1560
1633
  }
1561
1634
 
1562
1635
  // src/planner.ts
1563
- import { writeFile as writeFile3 } from "fs/promises";
1636
+ import { mkdir as mkdir4, writeFile as writeFile3 } from "fs/promises";
1564
1637
  import path6 from "path";
1565
1638
  import { stringify } from "yaml";
1566
1639
 
@@ -1687,16 +1760,21 @@ async function generateForRoute(page, options) {
1687
1760
  return { ...generated, routes: [route] };
1688
1761
  }
1689
1762
  async function writeDraft(cwd, draft, meta) {
1690
- const file = path6.join(cwd, TESTS_RELATIVE_DIR, `${routeToSlug(meta.route)}.yaml`);
1763
+ const dir = path6.join(cwd, TESTS_RELATIVE_DIR);
1764
+ const file = path6.join(dir, `${routeToSlug(meta.route)}.yaml`);
1691
1765
  try {
1766
+ await mkdir4(dir, { recursive: true });
1692
1767
  await writeFile3(file, renderTestYaml(draft, meta), { flag: "wx" });
1693
1768
  } catch (error) {
1694
- if (error.code === "EEXIST") {
1769
+ const err = error;
1770
+ if (err.code === "EEXIST" && err.syscall !== "mkdir") {
1695
1771
  throw new PlannerError(
1696
1772
  `Refusing to overwrite ${path6.relative(cwd, file)} (route ${meta.route}). Delete or rename it to regenerate.`
1697
1773
  );
1698
1774
  }
1699
- throw error;
1775
+ throw new PlannerError(
1776
+ `Cannot write draft to ${file}: ${fsReason(error)}. Check that ${path6.dirname(file)} is a directory you can write to, not a file.`
1777
+ );
1700
1778
  }
1701
1779
  return file;
1702
1780
  }
@@ -1857,7 +1935,7 @@ function formatSpendLine(spend) {
1857
1935
  import path9 from "path";
1858
1936
 
1859
1937
  // src/report/html.ts
1860
- import { mkdir as mkdir4, readFile as readFile4, writeFile as writeFile4 } from "fs/promises";
1938
+ import { mkdir as mkdir5, readFile as readFile4, writeFile as writeFile4 } from "fs/promises";
1861
1939
  import path7 from "path";
1862
1940
  var ESCAPES = {
1863
1941
  "&": "&amp;",
@@ -2022,7 +2100,7 @@ ${detail}
2022
2100
  }
2023
2101
  async function writeHtml(file, html) {
2024
2102
  try {
2025
- await mkdir4(path7.dirname(file), { recursive: true });
2103
+ await mkdir5(path7.dirname(file), { recursive: true });
2026
2104
  await writeFile4(file, html, "utf8");
2027
2105
  return file;
2028
2106
  } catch (error) {
@@ -2033,7 +2111,7 @@ async function writeHtml(file, html) {
2033
2111
  }
2034
2112
 
2035
2113
  // src/report/junit.ts
2036
- import { mkdir as mkdir5, writeFile as writeFile5 } from "fs/promises";
2114
+ import { mkdir as mkdir6, writeFile as writeFile5 } from "fs/promises";
2037
2115
  import path8 from "path";
2038
2116
  var ESCAPES2 = {
2039
2117
  "&": "&amp;",
@@ -2101,7 +2179,7 @@ ${reason}` : reason;
2101
2179
  }
2102
2180
  async function writeJUnit(file, xml) {
2103
2181
  try {
2104
- await mkdir5(path8.dirname(file), { recursive: true });
2182
+ await mkdir6(path8.dirname(file), { recursive: true });
2105
2183
  await writeFile5(file, xml, "utf8");
2106
2184
  return file;
2107
2185
  } catch (error) {
@@ -2879,13 +2957,11 @@ async function planCommand(options) {
2879
2957
  written.push(path10.relative(options.cwd, file));
2880
2958
  console.log(` wrote ${path10.relative(options.cwd, file)}`);
2881
2959
  } catch (error) {
2882
- if (error instanceof PlannerError) {
2883
- generated.pop();
2884
- failed.push({ route, reason: error.message });
2885
- console.log(` X failed: ${error.message}`);
2886
- continue;
2887
- }
2888
- throw error;
2960
+ generated.pop();
2961
+ const reason = error instanceof Error ? error.message : String(error);
2962
+ failed.push({ route, reason });
2963
+ console.log(` X failed: ${reason}`);
2964
+ continue;
2889
2965
  }
2890
2966
  } else {
2891
2967
  console.log(`
@@ -3006,7 +3082,7 @@ function parsePositiveNumber(flag) {
3006
3082
  };
3007
3083
  }
3008
3084
  var program = new Command();
3009
- program.name("blastproof").description("Open-source AI testing agent: plain-English YAML tests executed agentically on a real browser.").version("0.15.0");
3085
+ program.name("blastproof").description("Open-source AI testing agent: plain-English YAML tests executed agentically on a real browser.").version("0.17.0");
3010
3086
  program.command("init").description("Scaffold .blastproof/ (config, tests, sample tests) in the current directory").action(async () => {
3011
3087
  try {
3012
3088
  const result = await initProject(process.cwd());