blastproof 0.9.0 → 0.11.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -42,6 +42,8 @@ Before `run`, `plan` or `test` do anything, they check what they are about to sp
42
42
 
43
43
  **Your markup must be accessible — a hard requirement.** Elements are found by role, label or visible text from the accessibility tree. That is what removes selectors and survives redesigns; the cost is that an interface the accessibility tree cannot describe cannot be driven at all, and there is deliberately no CSS or XPath fallback. Icon-only buttons without accessible names, `div`-based controls and ARIA-less dropdowns simply cannot be targeted. Run an accessibility checker first — the result predicts how well this will work better than anything else.
44
44
 
45
+ **Windows is untested.** Development and CI run on Linux and macOS. Nothing is known to be broken and reports are welcome (#8).
46
+
45
47
  **Not supported yet:** `iframe` content (so hosted payment widgets like Stripe Elements are invisible — an embedded checkout cannot be driven end to end), hover, scroll-to, drag and drop, file upload, multiple tabs, native `alert`/`confirm` dialogs. Page snapshots are capped at 200 lines by default, so very dense pages are truncated — raise it with `browser.max_snapshot_lines` if your pages need more; truncation is always marked in the snapshot so the model is never misled into thinking it saw the whole page.
46
48
 
47
49
  **Point it at disposable data.** Within a step, an action that commits — a click, or pressing Enter — is never performed twice: the runner refuses the repeat and tells the agent it already did that. This closes the case that used to produce duplicate records, where a submit answered by a redirect came back to a reset form and the agent, seeing no evidence of its own work, submitted again. It is not a guarantee of zero duplicate writes: an agent that reaches the same effect by a genuinely different route — another control with the same effect — is not caught. Use a seeded database, a staging environment you can reset, or a throwaway account; do not gate on a run against production data.
@@ -70,6 +72,26 @@ steps:
70
72
 
71
73
  `priority` is P0–P2 (default P1). `tags`, `setup` steps and `auth` are optional — `auth: false` runs the test signed out, which a login test needs. `routes` declares the URLs a test covers, which is what `--impacted` selects on; write route strings consistently, since they compare by exact equality (`/cart` ≠ `/cart/`).
72
74
 
75
+ ### Say what each step should produce
76
+
77
+ This is the one rule that decides whether a suite works. An outside evaluation took the **same application, same suite, same version from Score 64 to Score 100 by rewriting two steps** — nothing else changed:
78
+
79
+ ```yaml
80
+ # Fragile: a bare action. Nothing says what should be true afterwards.
81
+ - submit the add-task form
82
+ - verify the new task "Fix the flaky test" appears in the task list
83
+
84
+ # Robust: the step carries its own outcome.
85
+ - submit the add-task form and verify the task "Fix the flaky test" appears
86
+ in the task list with priority High and status Open
87
+ ```
88
+
89
+ The bare version fails on a shape that is everywhere: the form POSTs, the server redirects back to the same page, and the form comes back empty. The agent is asked whether "submit the add-task form" happened, and is looking at a page that is indistinguishable from one where nothing did. A step that names the outcome gives it something to check that survives the action.
90
+
91
+ Write steps that end in an observable result — text on the page, a count, a state change — and this class of failure does not arise.
92
+
93
+ **Inline error messages should be plain visible text.** `role="alert"` is read correctly from the accessibility tree and needs no special handling, but note that an alert your page has cleared shows up as an empty element: if a verdict says an alert exists whose content is missing, the message was emptied, not hidden.
94
+
73
95
  ## Commands
74
96
 
75
97
  ```bash
@@ -113,7 +135,7 @@ jobs:
113
135
 
114
136
  - run: npm start & # however your app boots
115
137
 
116
- - uses: hamc/blastproof@v0.9.0
138
+ - uses: hamc/blastproof@v0.11.0
117
139
  with:
118
140
  version: '0.6.0' # pin both when this gates merges
119
141
  api-key: ${{ secrets.ANTHROPIC_API_KEY }}
@@ -209,6 +231,8 @@ The figures also land in the JUnit report as `llm_calls` and `llm_tokens`, besid
209
231
 
210
232
  `--dry-run` reports the ceiling before you spend anything. Read it as a maximum and nothing more: for this repository's own suite it says 735 calls where a real run spends 82. Size a budget from what your runs actually report, not from the ceiling.
211
233
 
234
+ For an order of magnitude, this repository's own suite — 7 tests, 31 steps, an authenticated demo shop, `anthropic/claude-haiku-4.5` — spends **about 82 model calls and 115k tokens**, taking 156s serially or 68s at `--concurrency 4`. That is a number you can reproduce (`node examples/demo-app/serve.mjs 4173` and `blastproof run`), not a forecast for your suite: cost scales with steps, page density and how often the agent has to retry. Run yours once and read the `Spent:` line.
235
+
212
236
  ## Testing behind a login
213
237
 
214
238
  Declare a recipe once; blastproof signs in one time per run and reuses that session for every test and for `plan`. Pick exactly one strategy:
package/dist/cli.js CHANGED
@@ -436,7 +436,8 @@ async function performAction(page, action, ctx) {
436
436
  const url = new URL(value, ctx.baseUrl);
437
437
  assertAllowedOrigin(url, ctx);
438
438
  await page.goto(url.toString(), { timeout: ctx.resolveTimeoutMs ?? 3e4 });
439
- return `ok: navigated to ${url.toString()}`;
439
+ const landed = page.url();
440
+ return landed === url.toString() ? `ok: navigated to ${url.toString()}` : `ok: navigated to ${url.toString()}, which redirected to ${landed}`;
440
441
  }
441
442
  case "click": {
442
443
  const target = requireTarget(action);
@@ -642,11 +643,21 @@ async function executeTest(page, test, options) {
642
643
  }
643
644
  if (action.action === "assert") {
644
645
  const expectation = action.expectation ?? action.reasoning;
645
- let judgment = await brain.judge(mask(step), mask(expectation), mask(snap));
646
+ let judgment = await brain.judge(
647
+ mask(step),
648
+ mask(expectation),
649
+ mask(snap),
650
+ recovery.stepHistory()
651
+ );
646
652
  if (!judgment.pass) {
647
653
  await waitForSettled(page);
648
654
  const freshSnap = await takeSnapshot(page);
649
- judgment = await brain.judge(mask(step), mask(expectation), mask(freshSnap));
655
+ judgment = await brain.judge(
656
+ mask(step),
657
+ mask(expectation),
658
+ mask(freshSnap),
659
+ recovery.stepHistory()
660
+ );
650
661
  }
651
662
  const result = judgment.pass ? `ok: assertion passed: ${judgment.reason}` : `assertion failed: ${judgment.reason}`;
652
663
  emitAction(index, action, result);
@@ -1203,13 +1214,24 @@ A value sitting in a control that was just typed into \u2014 an open dialog's te
1203
1214
 
1204
1215
  A step that names an ACTION (submit, click, create, add, ...) is satisfied by evidence the action took effect, not by the action's own control still being on the page. A successful action ordinarily replaces or moves past exactly the form, button or field the step names, so that control's absence is normal evidence of success, not evidence the step is unverifiable \u2014 do not fail such a step only because you can no longer see the thing it names. Fail it instead when the snapshot shows the action did NOT take effect: an error message, a validation warning, or the very same pre-action page still in front of you with nothing changed. A different page, a new state, or the result the action was meant to produce counts as evidence it worked.
1205
1216
 
1217
+ You may also be shown the actions already performed in this step, with their results. That record tells you what was ATTEMPTED and what it produced \u2014 for instance that a navigation was performed and which URL the server ultimately served, or that a form was submitted. Use it to avoid concluding that something never happened when the page simply cannot show it any more: a navigation the server redirected does not leave the browser at the path that was requested, and that is what success looks like, not failure.
1218
+
1219
+ The record is not evidence that the step's outcome holds. An action reported as \`ok\` establishes that it ran and what it returned; whether the thing the step describes is now TRUE is still decided by the snapshot alone. Never pass a step because the record shows an action succeeded while the snapshot does not show the outcome.
1220
+
1206
1221
  Be strict about what the step asks, not about withholding a pass you can plainly see is earned. Answer with pass=true/false and a one-sentence reason.
1207
1222
 
1208
1223
  \`***\` marks a secret deliberately withheld from you \u2014 a password, token or key. Seeing it is expected. A field holding \`***\` is filled, not empty, so do not fail a step on the grounds that a value was redacted. This applies only to the redaction itself: everything else the step asks for must still be visibly satisfied by the snapshot, and a step you genuinely cannot check against what you were shown still fails.`;
1209
1224
  }
1210
- function assertUserPrompt(step, expectation, snapshot) {
1211
- return [
1212
- `Step under test: ${step}`,
1225
+ function assertUserPrompt(step, expectation, snapshot, stepHistory) {
1226
+ const parts = [`Step under test: ${step}`];
1227
+ if (stepHistory && stepHistory.length > 0) {
1228
+ parts.push(
1229
+ "",
1230
+ "Actions already performed in this step, with their results (what was DONE \u2014 not evidence of what is now true):",
1231
+ ...stepHistory.map((entry, i) => `${i + 1}. ${entry.action} -> ${entry.result}`)
1232
+ );
1233
+ }
1234
+ parts.push(
1213
1235
  "",
1214
1236
  `Model's expectation (the claim offered in support of the step, not the question itself): ${expectation}`,
1215
1237
  "",
@@ -1217,7 +1239,8 @@ function assertUserPrompt(step, expectation, snapshot) {
1217
1239
  snapshot,
1218
1240
  "",
1219
1241
  "Does the snapshot establish that the step's own outcome holds?"
1220
- ].join("\n");
1242
+ );
1243
+ return parts.join("\n");
1221
1244
  }
1222
1245
  function plannerSystemPrompt() {
1223
1246
  return `You are a QA engineer writing one end-to-end test for a web page, in plain English.
@@ -1225,11 +1248,13 @@ function plannerSystemPrompt() {
1225
1248
  You receive a YAML accessibility snapshot of the page (roles and accessible names, exactly what a user perceives) and the list of source files a pull request changed in the area this page covers.
1226
1249
 
1227
1250
  Rules:
1228
- - Write steps a human tester could follow without looking at the code. One action or check per step.
1251
+ - Write steps a human tester could follow without looking at the code. One move per step \u2014 a single action together with what it should produce, or a single check. Never two unrelated actions in one step.
1229
1252
  - Refer to controls by the accessible name shown in the snapshot, spelled exactly. Never invent buttons, fields or links that are not in the snapshot.
1230
1253
  - Never write CSS selectors, XPath, IDs or any code \u2014 the runner resolves elements live from the accessibility tree.
1231
1254
  - Prefer the journey the changed files touch over a generic tour of the page. The changed files tell you which part of the page matters.
1232
- - End with at least one step that verifies an observable outcome (visible text, a count, a state change).
1255
+ - **The test starts at the application's base URL, not at this route.** Begin with a step that navigates to the route and says what should be visible once it loads \u2014 "navigate to /support and verify the heading "Contact support" is shown". Without it the run opens the home page and every later step looks for controls that are not there.
1256
+ - **Every step says what it should produce.** Name what must be true once the step has been carried out, not the action alone: "submit the support form and verify the confirmation page shows the ticket number", never "submit the support form". A step that names an action without an outcome asks the runner to judge whether something happened while looking at the page that succeeding produces \u2014 a submitted form comes back empty, a redirect moves the URL \u2014 and that is the shape behind several real failures.
1257
+ - **A step that enters a value writes the value.** "fill the subject field with Order not received", never "enter a subject". The runner is forbidden from inventing values, so a step that does not supply one cannot be carried out.
1233
1258
  - If a step needs a credential or any secret, write it as a placeholder like {{env.TEST_PASSWORD}}. Never write a real or invented password, token or key.
1234
1259
  - Keep the whole test to a handful of steps: one journey, not an exhaustive suite.`;
1235
1260
  }
@@ -1306,12 +1331,12 @@ function createBrain(model, generate = generateObject, budget) {
1306
1331
  }
1307
1332
  return parsed.data;
1308
1333
  },
1309
- async judge(step, expectation, snapshot) {
1334
+ async judge(step, expectation, snapshot, stepHistory) {
1310
1335
  const result = await countedGenerate(generate, budget, {
1311
1336
  model,
1312
1337
  schema: assertJudgmentSchema,
1313
1338
  system: assertSystemPrompt(),
1314
- prompt: assertUserPrompt(step, expectation, snapshot)
1339
+ prompt: assertUserPrompt(step, expectation, snapshot, stepHistory)
1315
1340
  });
1316
1341
  const parsed = assertJudgmentSchema.safeParse(result.object);
1317
1342
  if (!parsed.success) {
@@ -2751,7 +2776,7 @@ function parsePositiveNumber(flag) {
2751
2776
  };
2752
2777
  }
2753
2778
  var program = new Command();
2754
- program.name("blastproof").description("Open-source AI testing agent: plain-English YAML tests executed agentically on a real browser.").version("0.9.0");
2779
+ program.name("blastproof").description("Open-source AI testing agent: plain-English YAML tests executed agentically on a real browser.").version("0.11.0");
2755
2780
  program.command("init").description("Scaffold .blastproof/ (config, tests, sample tests) in the current directory").action(async () => {
2756
2781
  try {
2757
2782
  const result = await initProject(process.cwd());