blastproof 0.12.1 → 0.14.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -141,11 +141,13 @@ Write steps that end in an observable result — text on the page, a count, a st
141
141
  - fill the note field # cannot run
142
142
  ```
143
143
 
144
- The agent is **forbidden from inventing values** — one it types must come from the step, from the page, or from an `{{env.*}}` placeholder. But that rule lives in a prompt, and a prompt instructs rather than enforces. Run against a real model, the second step does not fail: the agent makes a value up, fills it, and the step **passes**. `fill the note field` produced "This is a test note." on two runs and "This is a new note" on a third.
144
+ The agent is **forbidden from inventing values** — one it types must come from the step, from the page, or from an `{{env.*}}` placeholder — and the runner enforces it rather than asking. A `fill` or `select` whose value appears in none of those is refused: it is not typed, the agent is told which sources it may draw from, and the step fails on the retry budget if it keeps insisting.
145
145
 
146
- That is worse than a failure. The test goes green having verified a value nobody wrote, differing between runs — a passing check over an unspecified input. Closing the gap in the runner is [#57](https://github.com/hamc/blastproof/issues/57); until then, this warning is what stands between you and a green test that means nothing.
146
+ The rule used to live only in the prompt, and a prompt instructs rather than enforces. Run against a real model, `fill the note field` did not fail — the agent made a value up, filled it, and the step **passed**, producing "This is a test note." on two runs and "This is a new note" on a third. A test going green over a value nobody wrote, differing between runs, is worse than a failure.
147
147
 
148
- `run` warns about it first, on every path, before launching a browser or asking for a key:
148
+ Two limits worth knowing. A value the page shows in one format and the field wants in another — `1234` in the step, `1,234.00` in the box — is refused, and the fix is to write the value the way it is typed. And a very short value (`3`) appears somewhere in almost any page, so it will pass; this closes fabricated content, not every fabricated character.
149
+
150
+ `run` also warns about it first — the same rule caught earlier, from the test file, on every path, before launching a browser or asking for a key:
149
151
 
150
152
  ```
151
153
  Authoring (a step enters a value but names none):
@@ -156,7 +158,7 @@ Authoring (a step enters a value but names none):
156
158
 
157
159
  Non-fatal by default — `--fail-on-authoring` turns it into exit 1 for teams enforcing it in CI. Taking the value from the page is fine and is not flagged: `fill the recipient field with the address shown on the confirmation page`.
158
160
 
159
- **The check reads English only.** A suite written in another language runs exactly as well but is not inspected, and prints no warning saying so — silence from this check means "nothing found in English", never "this suite is clean".
161
+ **The warning reads English only.** A suite written in another language runs exactly as well but is not inspected, and prints no warning saying so — silence from this check means "nothing found in English", never "this suite is clean". The runner's refusal has no such limit: it compares text rather than parsing grammar, so a suite in any language is still held to the rule at run time.
160
162
 
161
163
  ## Commands
162
164
 
@@ -198,6 +200,8 @@ ignore:
198
200
  - "**/*.md"
199
201
  ```
200
202
 
203
+ The key is the file glob and the value is the routes it can affect — the opposite way round from a test file's own `routes:`, which is a plain list of the routes that test covers. Inverted, the map matches nothing at all and every run reports a diff that affected no page, so `blastproof` refuses a map written that way rather than running green against it.
204
+
201
205
  Every changed file lands in one of three buckets: it matches `routes:` and contributes them, matches `ignore:` and is knowingly irrelevant, or **matches neither — nobody has said what it affects**. `--fail-on-unmapped` blocks on that third case, naming the files and both ways to resolve them.
202
206
 
203
207
  **Nothing is ignored by default**, on purpose: a default that guesses on your behalf would hide the first files worth thinking about. The flag is additive — a run can meet `--min-score` and still be blocked here, because "the tests I ran passed" and "something changed that nobody classified" are different claims.
package/dist/cli.js CHANGED
@@ -49,6 +49,9 @@ browser:
49
49
  # max_retries_per_step: 3
50
50
 
51
51
  # Impact hints mapping file globs to routes (used by \`blastproof run --impacted\`).
52
+ # Read each line as "if this file changed, these pages are at risk" \u2014 the key is the
53
+ # file glob, the value is the routes. Not the other way round: a test file's own
54
+ # \`routes:\` is a list of routes, and this one is a map keyed by source path.
52
55
  routes:
53
56
  "src/auth/**": ["/login"]
54
57
  "src/cart/**": ["/cart", "/checkout"]
@@ -505,6 +508,11 @@ async function performAction(page, action, ctx) {
505
508
  // src/runner/recovery.ts
506
509
  var COMMIT_ACTIONS = /* @__PURE__ */ new Set(["click", "press"]);
507
510
  var COMMIT_KEYS = /* @__PURE__ */ new Set(["Enter", "NumpadEnter", " ", "Space", "Spacebar"]);
511
+ var SOURCED_VALUE_ACTIONS = /* @__PURE__ */ new Set(["fill", "select"]);
512
+ var ENV_PLACEHOLDER2 = /^\s*\{\{env\.[A-Za-z_][A-Za-z0-9_]*\}\}\s*$/;
513
+ function normalise(text) {
514
+ return text.replace(/\s+/g, " ").trim().toLowerCase();
515
+ }
508
516
  function describeAction(action) {
509
517
  const target = action.target ? ` ${action.target.role ?? ""} "${action.target.name ?? action.target.text ?? ""}"` : "";
510
518
  const value = action.value ? ` [${action.value}]` : "";
@@ -522,10 +530,41 @@ function identity(action) {
522
530
  var StepRecovery = class {
523
531
  performed = /* @__PURE__ */ new Set();
524
532
  history = [];
533
+ /**
534
+ * Everything the model has been allowed to read this step, normalised: the
535
+ * step's own text, plus every snapshot it has been shown, plus every value it
536
+ * has already typed successfully.
537
+ *
538
+ * One accumulating haystack rather than the snapshot in hand (design D2). The
539
+ * model may read an order number on one page, navigate away, and type it into
540
+ * a box on another — at that moment the current snapshot does not contain it,
541
+ * and a check against only that snapshot would refuse a legitimate action.
542
+ *
543
+ * Bounded by construction: `max_snapshot_lines` caps each snapshot,
544
+ * `maxIterationsPerStep` caps how many arrive, and the instance dies with the
545
+ * step — which is also the whole of the "does not cross steps" rule.
546
+ */
547
+ readable;
548
+ constructor(step) {
549
+ this.readable = [normalise(step)];
550
+ }
551
+ /**
552
+ * Records a snapshot the model was shown, as a source it may quote from.
553
+ *
554
+ * Takes the **masked** text (design D4). `executor.ts` shows the model
555
+ * `mask(snap)` and never the raw tree, so a secret rendered on the page is
556
+ * `***` to the model and cannot have been copied by it. Accepting the raw
557
+ * snapshot here would credit the model with access it never had, and would
558
+ * leave the executor and the model disagreeing about what the page said.
559
+ */
560
+ observe(maskedSnapshot) {
561
+ this.readable.push(normalise(maskedSnapshot));
562
+ }
525
563
  /** Records an action that was actually performed and succeeded. */
526
564
  record(action, description, result) {
527
565
  this.performed.add(identity(action));
528
566
  this.history.push({ action: description, result });
567
+ if (action.value) this.readable.push(normalise(action.value));
529
568
  }
530
569
  /**
531
570
  * The reason to refuse `action`, or `undefined` when it may be performed.
@@ -544,11 +583,43 @@ var StepRecovery = class {
544
583
  * and #28 has now produced one on three applications.
545
584
  */
546
585
  refusalFor(action) {
586
+ return this.repeatedCommitRefusal(action) ?? this.unsourcedValueRefusal(action);
587
+ }
588
+ repeatedCommitRefusal(action) {
547
589
  if (!COMMIT_ACTIONS.has(action.action)) return void 0;
548
590
  if (action.action === "press" && !COMMIT_KEYS.has(action.value ?? "")) return void 0;
549
591
  if (!this.performed.has(identity(action))) return void 0;
550
592
  return `refused: this exact action already succeeded earlier in this step, so it was NOT performed again. Repeating something that commits repeats whatever it changed in the application. If the page no longer shows that it worked, that is normal for a submit answered by a redirect \u2014 check the record of what you have already done. Verify the step's outcome another way, or fail the step.`;
551
593
  }
594
+ /**
595
+ * Refuses a typed value that came from nowhere the model was entitled to read
596
+ * (design refuse-an-invented-value).
597
+ *
598
+ * `prompts.ts` has forbidden inventing a value since 0.7.0, and measured
599
+ * against a real model the rule simply does not hold: given `fill the note
600
+ * field`, the model supplied "This is a test note." twice and "This is a new
601
+ * note" once, and the step passed all three times. A prompt instructs; it does
602
+ * not enforce. The result is the false negative this project exists to remove
603
+ * — a green test over an input nobody wrote, differing run to run, with
604
+ * nothing in the report to say so because nothing knew.
605
+ *
606
+ * The comparison is `includes`, not equality: a value is legitimately a
607
+ * fragment of a step that also names the field, and of a snapshot that also
608
+ * holds the rest of the page.
609
+ *
610
+ * Unlike the authoring warning that predicts this before a run, nothing here
611
+ * parses English, so the guarantee holds for a suite written in any language.
612
+ */
613
+ unsourcedValueRefusal(action) {
614
+ if (!SOURCED_VALUE_ACTIONS.has(action.action)) return void 0;
615
+ const value = action.value;
616
+ if (!value) return void 0;
617
+ if (ENV_PLACEHOLDER2.test(value)) return void 0;
618
+ const needle = normalise(value);
619
+ if (needle === "") return void 0;
620
+ if (this.readable.some((source) => source.includes(needle))) return void 0;
621
+ return `refused: the value "${value}" was NOT typed, because it appears neither in this step nor anywhere on the pages you have been shown. A value you type must come from the step, from the page, or from an {{env.*}} placeholder \u2014 one you make up would put the test's verdict on an input nobody wrote. Use a value the step or the page gives you, or fail the step and say it supplies none.`;
622
+ }
552
623
  /** The step's history so far, oldest first, for the model's prompt. */
553
624
  stepHistory() {
554
625
  return this.history;
@@ -623,7 +694,7 @@ async function executeTest(page, test, options) {
623
694
  let failedAttempts = 0;
624
695
  let lastResult;
625
696
  let stepFailedReason;
626
- const recovery = new StepRecovery();
697
+ const recovery = new StepRecovery(step);
627
698
  try {
628
699
  while (true) {
629
700
  if (iterations >= maxIterationsPerStep) {
@@ -636,12 +707,14 @@ async function executeTest(page, test, options) {
636
707
  );
637
708
  }
638
709
  const snap = await takeSnapshot(page);
710
+ const maskedSnap = mask(snap);
711
+ recovery.observe(maskedSnap);
639
712
  let action;
640
713
  try {
641
714
  action = await brain.nextAction({
642
715
  step,
643
716
  isSetup: setup,
644
- snapshot: mask(snap),
717
+ snapshot: maskedSnap,
645
718
  lastResult: lastResult === void 0 ? void 0 : mask(lastResult),
646
719
  // Already masked when recorded, on the same boundary as everything
647
720
  // else crossing into a prompt (design contained-recovery, D2).
@@ -672,16 +745,18 @@ async function executeTest(page, test, options) {
672
745
  let judgment = await brain.judge(
673
746
  mask(step),
674
747
  mask(expectation),
675
- mask(snap),
748
+ maskedSnap,
676
749
  recovery.stepHistory()
677
750
  );
678
751
  if (!judgment.pass) {
679
752
  await waitForSettled(page);
680
753
  const freshSnap = await takeSnapshot(page);
754
+ const maskedFresh = mask(freshSnap);
755
+ recovery.observe(maskedFresh);
681
756
  judgment = await brain.judge(
682
757
  mask(step),
683
758
  mask(expectation),
684
- mask(freshSnap),
759
+ maskedFresh,
685
760
  recovery.stepHistory()
686
761
  );
687
762
  }
@@ -988,11 +1063,33 @@ var authSchema = z.object({
988
1063
  });
989
1064
  }
990
1065
  });
1066
+ var ROUTE_SHAPED = /^\/[^*?]*$/;
1067
+ var FILE_SHAPED = /[*?]|\.[a-z0-9]{1,8}$/i;
1068
+ function findInvertedRouteEntries(routes) {
1069
+ const inverted = [];
1070
+ for (const [key, values] of Object.entries(routes)) {
1071
+ if (!ROUTE_SHAPED.test(key)) continue;
1072
+ const file = values.find((value) => FILE_SHAPED.test(value));
1073
+ if (file !== void 0) inverted.push({ key, file });
1074
+ }
1075
+ return inverted;
1076
+ }
991
1077
  var configSchema = z.object({
992
1078
  base_url: z.string().url(),
993
1079
  llm: llmSchema.default({}),
994
1080
  browser: browserSchema.default({}),
995
- routes: z.record(z.array(z.string())).optional(),
1081
+ routes: z.record(z.array(z.string())).superRefine((value, ctx) => {
1082
+ const inverted = findInvertedRouteEntries(value);
1083
+ const first = inverted[0];
1084
+ if (!first) return;
1085
+ const others = inverted.length > 1 ? ` (and ${inverted.length - 1} more)` : "";
1086
+ ctx.addIssue({
1087
+ code: z.ZodIssueCode.custom,
1088
+ message: `is the wrong way round${others} \u2014 the key is the file glob, the value is the routes it affects.
1089
+ found: "${first.key}": ["${first.file}"]
1090
+ expected: "${first.file}": ["${first.key}"]`
1091
+ });
1092
+ }).optional(),
996
1093
  /** Globs for files knowingly irrelevant to any route (docs, CI config, licences). */
997
1094
  ignore: z.array(z.string()).optional(),
998
1095
  /**
@@ -1208,7 +1305,7 @@ Rules:
1208
1305
  - Return "done" when the current step's outcome holds \u2014 including when it already held before you acted, or was achieved by your previous action. "Already true" is done, never failure. Do not return "done" for work belonging to later steps.
1209
1306
  - Return "fail" only when the step's outcome cannot be reached: the element is still absent after retries, the page cannot support the step, or an error blocks progress. Never return "fail" because the work appears to have been done already.
1210
1307
  - If your previous action errored, re-read the fresh snapshot and choose an alternative element or approach. Do not repeat the exact same failing action.
1211
- - Never invent a value. A value you type must come from the step, from the page, or from an {{env.*}} placeholder. If a step needs a value it does not give you, that is a failing step, not a gap for you to fill in.
1308
+ - Never invent a value. A value you type must come from the step, from the page, or from an {{env.*}} placeholder. This one is enforced, not merely asked: a fill or select whose value is in none of those is refused and not performed. If a step needs a value it does not give you, that is a failing step, not a gap for you to fill in.
1212
1309
  - A record of the actions you already performed in this step may be shown to you. It is the ground truth about what happened, even when the page no longer shows it: a form that submitted successfully and came back empty looks exactly like one you never submitted. Do not redo work that record says you already did.
1213
1310
  - \`***\` in a snapshot is a redacted secret \u2014 a password, token or key deliberately withheld from you. Seeing it is expected and is not a problem. A field showing \`***\` after you filled it from an {{env.VAR}} placeholder means the fill worked; treat that as success and move on. Never retry a fill because its value is redacted, and never report failure because a value was withheld.
1214
1311
  - Keep reasoning to one short sentence.`;
@@ -1280,7 +1377,7 @@ Rules:
1280
1377
  - Prefer the journey the changed files touch over a generic tour of the page. The changed files tell you which part of the page matters.
1281
1378
  - **The test starts at the application's base URL, not at this route.** Begin with a step that navigates to the route and says what should be visible once it loads \u2014 "navigate to /support and verify the heading "Contact support" is shown". Without it the run opens the home page and every later step looks for controls that are not there.
1282
1379
  - **Every step says what it should produce.** Name what must be true once the step has been carried out, not the action alone: "submit the support form and verify the confirmation page shows the ticket number", never "submit the support form". A step that names an action without an outcome asks the runner to judge whether something happened while looking at the page that succeeding produces \u2014 a submitted form comes back empty, a redirect moves the URL \u2014 and that is the shape behind several real failures.
1283
- - **A step that enters a value writes the value.** "fill the subject field with Order not received", never "enter a subject". The runner is forbidden from inventing values, but a step that supplies none does not fail \u2014 the value gets made up, the step passes, and the test then verifies something nobody wrote.
1380
+ - **A step that enters a value writes the value.** "fill the subject field with Order not received", never "enter a subject". The runner is forbidden from inventing values, and enforces it: a fill whose value is in neither the step nor the page is refused, so a step that supplies none cannot be relied on to run.
1284
1381
  - If a step needs a credential or any secret, write it as a placeholder like {{env.TEST_PASSWORD}}. Never write a real or invented password, token or key.
1285
1382
  - Keep the whole test to a handful of steps: one journey, not an exhaustive suite.`;
1286
1383
  }
@@ -2346,7 +2443,7 @@ function printAuthoring(result, cwd) {
2346
2443
  console.error(` \u2192 ${suggestValueClause(step)}`);
2347
2444
  }
2348
2445
  console.error(
2349
- "The runner is forbidden from inventing values, but only by instruction: in practice the model supplies one anyway, the step passes, and the test verifies a value nobody wrote \u2014 a different one on each run."
2446
+ "The runner is forbidden from inventing values and enforces it: one it cannot trace to the step, the page or an {{env.*}} placeholder is refused, so a step supplying none fails at run time. Naming the value here is what makes it run."
2350
2447
  );
2351
2448
  console.error("Steps are inspected in English only; steps in other languages are not checked.");
2352
2449
  }
@@ -2893,7 +2990,7 @@ function parsePositiveNumber(flag) {
2893
2990
  };
2894
2991
  }
2895
2992
  var program = new Command();
2896
- program.name("blastproof").description("Open-source AI testing agent: plain-English YAML tests executed agentically on a real browser.").version("0.12.1");
2993
+ program.name("blastproof").description("Open-source AI testing agent: plain-English YAML tests executed agentically on a real browser.").version("0.14.0");
2897
2994
  program.command("init").description("Scaffold .blastproof/ (config, tests, sample tests) in the current directory").action(async () => {
2898
2995
  try {
2899
2996
  const result = await initProject(process.cwd());