blastproof 0.13.0 → 0.15.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -141,11 +141,15 @@ Write steps that end in an observable result — text on the page, a count, a st
141
141
  - fill the note field # cannot run
142
142
  ```
143
143
 
144
- The agent is **forbidden from inventing values** — one it types must come from the step, from the page, or from an `{{env.*}}` placeholder. But that rule lives in a prompt, and a prompt instructs rather than enforces. Run against a real model, the second step does not fail: the agent makes a value up, fills it, and the step **passes**. `fill the note field` produced "This is a test note." on two runs and "This is a new note" on a third.
144
+ The agent is **forbidden from inventing values** — one it types must come from the step, from the page, or from an `{{env.*}}` placeholder — and the runner enforces it rather than asking. A `fill` or `select` whose value appears in none of those is refused: it is not typed, the agent is told which sources it may draw from, and the step fails on the retry budget if it keeps insisting.
145
145
 
146
- That is worse than a failure. The test goes green having verified a value nobody wrote, differing between runs — a passing check over an unspecified input. Closing the gap in the runner is [#57](https://github.com/hamc/blastproof/issues/57); until then, this warning is what stands between you and a green test that means nothing.
146
+ The rule used to live only in the prompt, and a prompt instructs rather than enforces. Run against a real model, `fill the note field` did not fail — the agent made a value up, filled it, and the step **passed**, producing "This is a test note." on two runs and "This is a new note" on a third. A test going green over a value nobody wrote, differing between runs, is worse than a failure.
147
147
 
148
- `run` warns about it first, on every path, before launching a browser or asking for a key:
148
+ A placeholder counts as a source **only when the step names that variable**. `fill the password field with {{env.TEST_PASSWORD}}` works; `fill the password field` does not become valid because the agent supplies `{{env.SOMETHING}}` itself. An agent cannot know the name of a variable nobody showed it, so one it produces is a guess — and a guessed variable would put a live credential into a field your test never pointed one at, in output that cannot redact a secret it was never told about.
149
+
150
+ Two limits worth knowing. A value the page shows in one format and the field wants in another — `1234` in the step, `1,234.00` in the box — is refused, and the fix is to write the value the way it is typed. And a very short value (`3`) appears somewhere in almost any page, so it will pass; this closes fabricated content, not every fabricated character.
151
+
152
+ `run` also warns about it first — the same rule caught earlier, from the test file, on every path, before launching a browser or asking for a key:
149
153
 
150
154
  ```
151
155
  Authoring (a step enters a value but names none):
@@ -156,7 +160,7 @@ Authoring (a step enters a value but names none):
156
160
 
157
161
  Non-fatal by default — `--fail-on-authoring` turns it into exit 1 for teams enforcing it in CI. Taking the value from the page is fine and is not flagged: `fill the recipient field with the address shown on the confirmation page`.
158
162
 
159
- **The check reads English only.** A suite written in another language runs exactly as well but is not inspected, and prints no warning saying so — silence from this check means "nothing found in English", never "this suite is clean".
163
+ **The warning reads English only.** A suite written in another language runs exactly as well but is not inspected, and prints no warning saying so — silence from this check means "nothing found in English", never "this suite is clean". The runner's refusal has no such limit: it compares text rather than parsing grammar, so a suite in any language is still held to the rule at run time.
160
164
 
161
165
  ## Commands
162
166
 
package/dist/cli.js CHANGED
@@ -508,6 +508,14 @@ async function performAction(page, action, ctx) {
508
508
  // src/runner/recovery.ts
509
509
  var COMMIT_ACTIONS = /* @__PURE__ */ new Set(["click", "press"]);
510
510
  var COMMIT_KEYS = /* @__PURE__ */ new Set(["Enter", "NumpadEnter", " ", "Space", "Spacebar"]);
511
+ var SOURCED_VALUE_ACTIONS = /* @__PURE__ */ new Set(["fill", "select"]);
512
+ function referencesOnlyNamedVars(value, step) {
513
+ const named = new Set(referencedEnvVars(step));
514
+ return referencedEnvVars(value).every((name) => named.has(name));
515
+ }
516
+ function normalise(text) {
517
+ return text.replace(/\s+/g, " ").trim().toLowerCase();
518
+ }
511
519
  function describeAction(action) {
512
520
  const target = action.target ? ` ${action.target.role ?? ""} "${action.target.name ?? action.target.text ?? ""}"` : "";
513
521
  const value = action.value ? ` [${action.value}]` : "";
@@ -525,10 +533,50 @@ function identity(action) {
525
533
  var StepRecovery = class {
526
534
  performed = /* @__PURE__ */ new Set();
527
535
  history = [];
536
+ /**
537
+ * Everything the model has been allowed to read this step, normalised: the
538
+ * step's own text, plus every snapshot it has been shown, plus every value it
539
+ * has already typed successfully.
540
+ *
541
+ * One accumulating haystack rather than the snapshot in hand (design D2). The
542
+ * model may read an order number on one page, navigate away, and type it into
543
+ * a box on another — at that moment the current snapshot does not contain it,
544
+ * and a check against only that snapshot would refuse a legitimate action.
545
+ *
546
+ * Bounded by construction: `max_snapshot_lines` caps each snapshot,
547
+ * `maxIterationsPerStep` caps how many arrive, and the instance dies with the
548
+ * step — which is also the whole of the "does not cross steps" rule.
549
+ */
550
+ readable;
551
+ /**
552
+ * The step exactly as written, for deciding which `{{env.*}}` variables it
553
+ * names (design env-placeholder-must-be-named, D3). Kept alongside the
554
+ * normalised copy rather than derived from it: `readable` is lowercased, and
555
+ * environment variable names are case-sensitive, so `TOKEN` and `token` must
556
+ * stay distinguishable in the one place that decides which secret gets typed.
557
+ */
558
+ step;
559
+ constructor(step) {
560
+ this.step = step;
561
+ this.readable = [normalise(step)];
562
+ }
563
+ /**
564
+ * Records a snapshot the model was shown, as a source it may quote from.
565
+ *
566
+ * Takes the **masked** text (design D4). `executor.ts` shows the model
567
+ * `mask(snap)` and never the raw tree, so a secret rendered on the page is
568
+ * `***` to the model and cannot have been copied by it. Accepting the raw
569
+ * snapshot here would credit the model with access it never had, and would
570
+ * leave the executor and the model disagreeing about what the page said.
571
+ */
572
+ observe(maskedSnapshot) {
573
+ this.readable.push(normalise(maskedSnapshot));
574
+ }
528
575
  /** Records an action that was actually performed and succeeded. */
529
576
  record(action, description, result) {
530
577
  this.performed.add(identity(action));
531
578
  this.history.push({ action: description, result });
579
+ if (action.value) this.readable.push(normalise(action.value));
532
580
  }
533
581
  /**
534
582
  * The reason to refuse `action`, or `undefined` when it may be performed.
@@ -547,11 +595,47 @@ var StepRecovery = class {
547
595
  * and #28 has now produced one on three applications.
548
596
  */
549
597
  refusalFor(action) {
598
+ return this.repeatedCommitRefusal(action) ?? this.unsourcedValueRefusal(action);
599
+ }
600
+ repeatedCommitRefusal(action) {
550
601
  if (!COMMIT_ACTIONS.has(action.action)) return void 0;
551
602
  if (action.action === "press" && !COMMIT_KEYS.has(action.value ?? "")) return void 0;
552
603
  if (!this.performed.has(identity(action))) return void 0;
553
604
  return `refused: this exact action already succeeded earlier in this step, so it was NOT performed again. Repeating something that commits repeats whatever it changed in the application. If the page no longer shows that it worked, that is normal for a submit answered by a redirect \u2014 check the record of what you have already done. Verify the step's outcome another way, or fail the step.`;
554
605
  }
606
+ /**
607
+ * Refuses a typed value that came from nowhere the model was entitled to read
608
+ * (design refuse-an-invented-value).
609
+ *
610
+ * `prompts.ts` has forbidden inventing a value since 0.7.0, and measured
611
+ * against a real model the rule simply does not hold: given `fill the note
612
+ * field`, the model supplied "This is a test note." twice and "This is a new
613
+ * note" once, and the step passed all three times. A prompt instructs; it does
614
+ * not enforce. The result is the false negative this project exists to remove
615
+ * — a green test over an input nobody wrote, differing run to run, with
616
+ * nothing in the report to say so because nothing knew.
617
+ *
618
+ * The comparison is `includes`, not equality: a value is legitimately a
619
+ * fragment of a step that also names the field, and of a snapshot that also
620
+ * holds the rest of the page.
621
+ *
622
+ * Unlike the authoring warning that predicts this before a run, nothing here
623
+ * parses English, so the guarantee holds for a suite written in any language.
624
+ */
625
+ unsourcedValueRefusal(action) {
626
+ if (!SOURCED_VALUE_ACTIONS.has(action.action)) return void 0;
627
+ const value = action.value;
628
+ if (!value) return void 0;
629
+ if (!referencesOnlyNamedVars(value, this.step)) {
630
+ const named = referencedEnvVars(this.step);
631
+ return `refused: the value "${value}" was NOT typed, because this step does not reference that {{env.*}} variable. ${named.length > 0 ? `This step references ${named.map((n) => `{{env.${n}}}`).join(", ")}.` : "This step references no environment variable."} A placeholder is a source only when the step names it \u2014 otherwise the test never asked for that secret to be typed here. Use a value this step supplies, or fail the step and say it supplies none.`;
632
+ }
633
+ if (referencedEnvVars(value).length > 0) return void 0;
634
+ const needle = normalise(value);
635
+ if (needle === "") return void 0;
636
+ if (this.readable.some((source) => source.includes(needle))) return void 0;
637
+ return `refused: the value "${value}" was NOT typed, because it appears neither in this step nor anywhere on the pages you have been shown. A value you type must come from the step, from the page, or from an {{env.*}} placeholder \u2014 one you make up would put the test's verdict on an input nobody wrote. Use a value the step or the page gives you, or fail the step and say it supplies none.`;
638
+ }
555
639
  /** The step's history so far, oldest first, for the model's prompt. */
556
640
  stepHistory() {
557
641
  return this.history;
@@ -626,7 +710,7 @@ async function executeTest(page, test, options) {
626
710
  let failedAttempts = 0;
627
711
  let lastResult;
628
712
  let stepFailedReason;
629
- const recovery = new StepRecovery();
713
+ const recovery = new StepRecovery(step);
630
714
  try {
631
715
  while (true) {
632
716
  if (iterations >= maxIterationsPerStep) {
@@ -639,12 +723,14 @@ async function executeTest(page, test, options) {
639
723
  );
640
724
  }
641
725
  const snap = await takeSnapshot(page);
726
+ const maskedSnap = mask(snap);
727
+ recovery.observe(maskedSnap);
642
728
  let action;
643
729
  try {
644
730
  action = await brain.nextAction({
645
731
  step,
646
732
  isSetup: setup,
647
- snapshot: mask(snap),
733
+ snapshot: maskedSnap,
648
734
  lastResult: lastResult === void 0 ? void 0 : mask(lastResult),
649
735
  // Already masked when recorded, on the same boundary as everything
650
736
  // else crossing into a prompt (design contained-recovery, D2).
@@ -675,16 +761,18 @@ async function executeTest(page, test, options) {
675
761
  let judgment = await brain.judge(
676
762
  mask(step),
677
763
  mask(expectation),
678
- mask(snap),
764
+ maskedSnap,
679
765
  recovery.stepHistory()
680
766
  );
681
767
  if (!judgment.pass) {
682
768
  await waitForSettled(page);
683
769
  const freshSnap = await takeSnapshot(page);
770
+ const maskedFresh = mask(freshSnap);
771
+ recovery.observe(maskedFresh);
684
772
  judgment = await brain.judge(
685
773
  mask(step),
686
774
  mask(expectation),
687
- mask(freshSnap),
775
+ maskedFresh,
688
776
  recovery.stepHistory()
689
777
  );
690
778
  }
@@ -1233,7 +1321,7 @@ Rules:
1233
1321
  - Return "done" when the current step's outcome holds \u2014 including when it already held before you acted, or was achieved by your previous action. "Already true" is done, never failure. Do not return "done" for work belonging to later steps.
1234
1322
  - Return "fail" only when the step's outcome cannot be reached: the element is still absent after retries, the page cannot support the step, or an error blocks progress. Never return "fail" because the work appears to have been done already.
1235
1323
  - If your previous action errored, re-read the fresh snapshot and choose an alternative element or approach. Do not repeat the exact same failing action.
1236
- - Never invent a value. A value you type must come from the step, from the page, or from an {{env.*}} placeholder. If a step needs a value it does not give you, that is a failing step, not a gap for you to fill in.
1324
+ - Never invent a value. A value you type must come from the step, from the page, or from an {{env.*}} placeholder. This one is enforced, not merely asked: a fill or select whose value is in none of those is refused and not performed. If a step needs a value it does not give you, that is a failing step, not a gap for you to fill in.
1237
1325
  - A record of the actions you already performed in this step may be shown to you. It is the ground truth about what happened, even when the page no longer shows it: a form that submitted successfully and came back empty looks exactly like one you never submitted. Do not redo work that record says you already did.
1238
1326
  - \`***\` in a snapshot is a redacted secret \u2014 a password, token or key deliberately withheld from you. Seeing it is expected and is not a problem. A field showing \`***\` after you filled it from an {{env.VAR}} placeholder means the fill worked; treat that as success and move on. Never retry a fill because its value is redacted, and never report failure because a value was withheld.
1239
1327
  - Keep reasoning to one short sentence.`;
@@ -1305,7 +1393,7 @@ Rules:
1305
1393
  - Prefer the journey the changed files touch over a generic tour of the page. The changed files tell you which part of the page matters.
1306
1394
  - **The test starts at the application's base URL, not at this route.** Begin with a step that navigates to the route and says what should be visible once it loads \u2014 "navigate to /support and verify the heading "Contact support" is shown". Without it the run opens the home page and every later step looks for controls that are not there.
1307
1395
  - **Every step says what it should produce.** Name what must be true once the step has been carried out, not the action alone: "submit the support form and verify the confirmation page shows the ticket number", never "submit the support form". A step that names an action without an outcome asks the runner to judge whether something happened while looking at the page that succeeding produces \u2014 a submitted form comes back empty, a redirect moves the URL \u2014 and that is the shape behind several real failures.
1308
- - **A step that enters a value writes the value.** "fill the subject field with Order not received", never "enter a subject". The runner is forbidden from inventing values, but a step that supplies none does not fail \u2014 the value gets made up, the step passes, and the test then verifies something nobody wrote.
1396
+ - **A step that enters a value writes the value.** "fill the subject field with Order not received", never "enter a subject". The runner is forbidden from inventing values, and enforces it: a fill whose value is in neither the step nor the page is refused, so a step that supplies none cannot be relied on to run.
1309
1397
  - If a step needs a credential or any secret, write it as a placeholder like {{env.TEST_PASSWORD}}. Never write a real or invented password, token or key.
1310
1398
  - Keep the whole test to a handful of steps: one journey, not an exhaustive suite.`;
1311
1399
  }
@@ -2371,7 +2459,7 @@ function printAuthoring(result, cwd) {
2371
2459
  console.error(` \u2192 ${suggestValueClause(step)}`);
2372
2460
  }
2373
2461
  console.error(
2374
- "The runner is forbidden from inventing values, but only by instruction: in practice the model supplies one anyway, the step passes, and the test verifies a value nobody wrote \u2014 a different one on each run."
2462
+ "The runner is forbidden from inventing values and enforces it: one it cannot trace to the step, the page or an {{env.*}} placeholder is refused, so a step supplying none fails at run time. Naming the value here is what makes it run."
2375
2463
  );
2376
2464
  console.error("Steps are inspected in English only; steps in other languages are not checked.");
2377
2465
  }
@@ -2918,7 +3006,7 @@ function parsePositiveNumber(flag) {
2918
3006
  };
2919
3007
  }
2920
3008
  var program = new Command();
2921
- program.name("blastproof").description("Open-source AI testing agent: plain-English YAML tests executed agentically on a real browser.").version("0.13.0");
3009
+ program.name("blastproof").description("Open-source AI testing agent: plain-English YAML tests executed agentically on a real browser.").version("0.15.0");
2922
3010
  program.command("init").description("Scaffold .blastproof/ (config, tests, sample tests) in the current directory").action(async () => {
2923
3011
  try {
2924
3012
  const result = await initProject(process.cwd());