blastproof 0.13.0 → 0.14.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -141,11 +141,13 @@ Write steps that end in an observable result — text on the page, a count, a st
141
141
  - fill the note field # cannot run
142
142
  ```
143
143
 
144
- The agent is **forbidden from inventing values** — one it types must come from the step, from the page, or from an `{{env.*}}` placeholder. But that rule lives in a prompt, and a prompt instructs rather than enforces. Run against a real model, the second step does not fail: the agent makes a value up, fills it, and the step **passes**. `fill the note field` produced "This is a test note." on two runs and "This is a new note" on a third.
144
+ The agent is **forbidden from inventing values** — one it types must come from the step, from the page, or from an `{{env.*}}` placeholder — and the runner enforces it rather than asking. A `fill` or `select` whose value appears in none of those is refused: it is not typed, the agent is told which sources it may draw from, and the step fails on the retry budget if it keeps insisting.
145
145
 
146
- That is worse than a failure. The test goes green having verified a value nobody wrote, differing between runs — a passing check over an unspecified input. Closing the gap in the runner is [#57](https://github.com/hamc/blastproof/issues/57); until then, this warning is what stands between you and a green test that means nothing.
146
+ The rule used to live only in the prompt, and a prompt instructs rather than enforces. Run against a real model, `fill the note field` did not fail — the agent made a value up, filled it, and the step **passed**, producing "This is a test note." on two runs and "This is a new note" on a third. A test going green over a value nobody wrote, differing between runs, is worse than a failure.
147
147
 
148
- `run` warns about it first, on every path, before launching a browser or asking for a key:
148
+ Two limits worth knowing. A value the page shows in one format and the field wants in another — `1234` in the step, `1,234.00` in the box — is refused, and the fix is to write the value the way it is typed. And a very short value (`3`) appears somewhere in almost any page, so it will pass; this closes fabricated content, not every fabricated character.
149
+
150
+ `run` also warns about it first — the same rule caught earlier, from the test file, on every path, before launching a browser or asking for a key:
149
151
 
150
152
  ```
151
153
  Authoring (a step enters a value but names none):
@@ -156,7 +158,7 @@ Authoring (a step enters a value but names none):
156
158
 
157
159
  Non-fatal by default — `--fail-on-authoring` turns it into exit 1 for teams enforcing it in CI. Taking the value from the page is fine and is not flagged: `fill the recipient field with the address shown on the confirmation page`.
158
160
 
159
- **The check reads English only.** A suite written in another language runs exactly as well but is not inspected, and prints no warning saying so — silence from this check means "nothing found in English", never "this suite is clean".
161
+ **The warning reads English only.** A suite written in another language runs exactly as well but is not inspected, and prints no warning saying so — silence from this check means "nothing found in English", never "this suite is clean". The runner's refusal has no such limit: it compares text rather than parsing grammar, so a suite in any language is still held to the rule at run time.
160
162
 
161
163
  ## Commands
162
164
 
package/dist/cli.js CHANGED
@@ -508,6 +508,11 @@ async function performAction(page, action, ctx) {
508
508
  // src/runner/recovery.ts
509
509
  var COMMIT_ACTIONS = /* @__PURE__ */ new Set(["click", "press"]);
510
510
  var COMMIT_KEYS = /* @__PURE__ */ new Set(["Enter", "NumpadEnter", " ", "Space", "Spacebar"]);
511
+ var SOURCED_VALUE_ACTIONS = /* @__PURE__ */ new Set(["fill", "select"]);
512
+ var ENV_PLACEHOLDER2 = /^\s*\{\{env\.[A-Za-z_][A-Za-z0-9_]*\}\}\s*$/;
513
+ function normalise(text) {
514
+ return text.replace(/\s+/g, " ").trim().toLowerCase();
515
+ }
511
516
  function describeAction(action) {
512
517
  const target = action.target ? ` ${action.target.role ?? ""} "${action.target.name ?? action.target.text ?? ""}"` : "";
513
518
  const value = action.value ? ` [${action.value}]` : "";
@@ -525,10 +530,41 @@ function identity(action) {
525
530
  var StepRecovery = class {
526
531
  performed = /* @__PURE__ */ new Set();
527
532
  history = [];
533
+ /**
534
+ * Everything the model has been allowed to read this step, normalised: the
535
+ * step's own text, plus every snapshot it has been shown, plus every value it
536
+ * has already typed successfully.
537
+ *
538
+ * One accumulating haystack rather than the snapshot in hand (design D2). The
539
+ * model may read an order number on one page, navigate away, and type it into
540
+ * a box on another — at that moment the current snapshot does not contain it,
541
+ * and a check against only that snapshot would refuse a legitimate action.
542
+ *
543
+ * Bounded by construction: `max_snapshot_lines` caps each snapshot,
544
+ * `maxIterationsPerStep` caps how many arrive, and the instance dies with the
545
+ * step — which is also the whole of the "does not cross steps" rule.
546
+ */
547
+ readable;
548
+ constructor(step) {
549
+ this.readable = [normalise(step)];
550
+ }
551
+ /**
552
+ * Records a snapshot the model was shown, as a source it may quote from.
553
+ *
554
+ * Takes the **masked** text (design D4). `executor.ts` shows the model
555
+ * `mask(snap)` and never the raw tree, so a secret rendered on the page is
556
+ * `***` to the model and cannot have been copied by it. Accepting the raw
557
+ * snapshot here would credit the model with access it never had, and would
558
+ * leave the executor and the model disagreeing about what the page said.
559
+ */
560
+ observe(maskedSnapshot) {
561
+ this.readable.push(normalise(maskedSnapshot));
562
+ }
528
563
  /** Records an action that was actually performed and succeeded. */
529
564
  record(action, description, result) {
530
565
  this.performed.add(identity(action));
531
566
  this.history.push({ action: description, result });
567
+ if (action.value) this.readable.push(normalise(action.value));
532
568
  }
533
569
  /**
534
570
  * The reason to refuse `action`, or `undefined` when it may be performed.
@@ -547,11 +583,43 @@ var StepRecovery = class {
547
583
  * and #28 has now produced one on three applications.
548
584
  */
549
585
  refusalFor(action) {
586
+ return this.repeatedCommitRefusal(action) ?? this.unsourcedValueRefusal(action);
587
+ }
588
+ repeatedCommitRefusal(action) {
550
589
  if (!COMMIT_ACTIONS.has(action.action)) return void 0;
551
590
  if (action.action === "press" && !COMMIT_KEYS.has(action.value ?? "")) return void 0;
552
591
  if (!this.performed.has(identity(action))) return void 0;
553
592
  return `refused: this exact action already succeeded earlier in this step, so it was NOT performed again. Repeating something that commits repeats whatever it changed in the application. If the page no longer shows that it worked, that is normal for a submit answered by a redirect \u2014 check the record of what you have already done. Verify the step's outcome another way, or fail the step.`;
554
593
  }
594
+ /**
595
+ * Refuses a typed value that came from nowhere the model was entitled to read
596
+ * (design refuse-an-invented-value).
597
+ *
598
+ * `prompts.ts` has forbidden inventing a value since 0.7.0, and measured
599
+ * against a real model the rule simply does not hold: given `fill the note
600
+ * field`, the model supplied "This is a test note." twice and "This is a new
601
+ * note" once, and the step passed all three times. A prompt instructs; it does
602
+ * not enforce. The result is the false negative this project exists to remove
603
+ * — a green test over an input nobody wrote, differing run to run, with
604
+ * nothing in the report to say so because nothing knew.
605
+ *
606
+ * The comparison is `includes`, not equality: a value is legitimately a
607
+ * fragment of a step that also names the field, and of a snapshot that also
608
+ * holds the rest of the page.
609
+ *
610
+ * Unlike the authoring warning that predicts this before a run, nothing here
611
+ * parses English, so the guarantee holds for a suite written in any language.
612
+ */
613
+ unsourcedValueRefusal(action) {
614
+ if (!SOURCED_VALUE_ACTIONS.has(action.action)) return void 0;
615
+ const value = action.value;
616
+ if (!value) return void 0;
617
+ if (ENV_PLACEHOLDER2.test(value)) return void 0;
618
+ const needle = normalise(value);
619
+ if (needle === "") return void 0;
620
+ if (this.readable.some((source) => source.includes(needle))) return void 0;
621
+ return `refused: the value "${value}" was NOT typed, because it appears neither in this step nor anywhere on the pages you have been shown. A value you type must come from the step, from the page, or from an {{env.*}} placeholder \u2014 one you make up would put the test's verdict on an input nobody wrote. Use a value the step or the page gives you, or fail the step and say it supplies none.`;
622
+ }
555
623
  /** The step's history so far, oldest first, for the model's prompt. */
556
624
  stepHistory() {
557
625
  return this.history;
@@ -626,7 +694,7 @@ async function executeTest(page, test, options) {
626
694
  let failedAttempts = 0;
627
695
  let lastResult;
628
696
  let stepFailedReason;
629
- const recovery = new StepRecovery();
697
+ const recovery = new StepRecovery(step);
630
698
  try {
631
699
  while (true) {
632
700
  if (iterations >= maxIterationsPerStep) {
@@ -639,12 +707,14 @@ async function executeTest(page, test, options) {
639
707
  );
640
708
  }
641
709
  const snap = await takeSnapshot(page);
710
+ const maskedSnap = mask(snap);
711
+ recovery.observe(maskedSnap);
642
712
  let action;
643
713
  try {
644
714
  action = await brain.nextAction({
645
715
  step,
646
716
  isSetup: setup,
647
- snapshot: mask(snap),
717
+ snapshot: maskedSnap,
648
718
  lastResult: lastResult === void 0 ? void 0 : mask(lastResult),
649
719
  // Already masked when recorded, on the same boundary as everything
650
720
  // else crossing into a prompt (design contained-recovery, D2).
@@ -675,16 +745,18 @@ async function executeTest(page, test, options) {
675
745
  let judgment = await brain.judge(
676
746
  mask(step),
677
747
  mask(expectation),
678
- mask(snap),
748
+ maskedSnap,
679
749
  recovery.stepHistory()
680
750
  );
681
751
  if (!judgment.pass) {
682
752
  await waitForSettled(page);
683
753
  const freshSnap = await takeSnapshot(page);
754
+ const maskedFresh = mask(freshSnap);
755
+ recovery.observe(maskedFresh);
684
756
  judgment = await brain.judge(
685
757
  mask(step),
686
758
  mask(expectation),
687
- mask(freshSnap),
759
+ maskedFresh,
688
760
  recovery.stepHistory()
689
761
  );
690
762
  }
@@ -1233,7 +1305,7 @@ Rules:
1233
1305
  - Return "done" when the current step's outcome holds \u2014 including when it already held before you acted, or was achieved by your previous action. "Already true" is done, never failure. Do not return "done" for work belonging to later steps.
1234
1306
  - Return "fail" only when the step's outcome cannot be reached: the element is still absent after retries, the page cannot support the step, or an error blocks progress. Never return "fail" because the work appears to have been done already.
1235
1307
  - If your previous action errored, re-read the fresh snapshot and choose an alternative element or approach. Do not repeat the exact same failing action.
1236
- - Never invent a value. A value you type must come from the step, from the page, or from an {{env.*}} placeholder. If a step needs a value it does not give you, that is a failing step, not a gap for you to fill in.
1308
+ - Never invent a value. A value you type must come from the step, from the page, or from an {{env.*}} placeholder. This one is enforced, not merely asked: a fill or select whose value is in none of those is refused and not performed. If a step needs a value it does not give you, that is a failing step, not a gap for you to fill in.
1237
1309
  - A record of the actions you already performed in this step may be shown to you. It is the ground truth about what happened, even when the page no longer shows it: a form that submitted successfully and came back empty looks exactly like one you never submitted. Do not redo work that record says you already did.
1238
1310
  - \`***\` in a snapshot is a redacted secret \u2014 a password, token or key deliberately withheld from you. Seeing it is expected and is not a problem. A field showing \`***\` after you filled it from an {{env.VAR}} placeholder means the fill worked; treat that as success and move on. Never retry a fill because its value is redacted, and never report failure because a value was withheld.
1239
1311
  - Keep reasoning to one short sentence.`;
@@ -1305,7 +1377,7 @@ Rules:
1305
1377
  - Prefer the journey the changed files touch over a generic tour of the page. The changed files tell you which part of the page matters.
1306
1378
  - **The test starts at the application's base URL, not at this route.** Begin with a step that navigates to the route and says what should be visible once it loads \u2014 "navigate to /support and verify the heading "Contact support" is shown". Without it the run opens the home page and every later step looks for controls that are not there.
1307
1379
  - **Every step says what it should produce.** Name what must be true once the step has been carried out, not the action alone: "submit the support form and verify the confirmation page shows the ticket number", never "submit the support form". A step that names an action without an outcome asks the runner to judge whether something happened while looking at the page that succeeding produces \u2014 a submitted form comes back empty, a redirect moves the URL \u2014 and that is the shape behind several real failures.
1308
- - **A step that enters a value writes the value.** "fill the subject field with Order not received", never "enter a subject". The runner is forbidden from inventing values, but a step that supplies none does not fail \u2014 the value gets made up, the step passes, and the test then verifies something nobody wrote.
1380
+ - **A step that enters a value writes the value.** "fill the subject field with Order not received", never "enter a subject". The runner is forbidden from inventing values, and enforces it: a fill whose value is in neither the step nor the page is refused, so a step that supplies none cannot be relied on to run.
1309
1381
  - If a step needs a credential or any secret, write it as a placeholder like {{env.TEST_PASSWORD}}. Never write a real or invented password, token or key.
1310
1382
  - Keep the whole test to a handful of steps: one journey, not an exhaustive suite.`;
1311
1383
  }
@@ -2371,7 +2443,7 @@ function printAuthoring(result, cwd) {
2371
2443
  console.error(` \u2192 ${suggestValueClause(step)}`);
2372
2444
  }
2373
2445
  console.error(
2374
- "The runner is forbidden from inventing values, but only by instruction: in practice the model supplies one anyway, the step passes, and the test verifies a value nobody wrote \u2014 a different one on each run."
2446
+ "The runner is forbidden from inventing values and enforces it: one it cannot trace to the step, the page or an {{env.*}} placeholder is refused, so a step supplying none fails at run time. Naming the value here is what makes it run."
2375
2447
  );
2376
2448
  console.error("Steps are inspected in English only; steps in other languages are not checked.");
2377
2449
  }
@@ -2918,7 +2990,7 @@ function parsePositiveNumber(flag) {
2918
2990
  };
2919
2991
  }
2920
2992
  var program = new Command();
2921
- program.name("blastproof").description("Open-source AI testing agent: plain-English YAML tests executed agentically on a real browser.").version("0.13.0");
2993
+ program.name("blastproof").description("Open-source AI testing agent: plain-English YAML tests executed agentically on a real browser.").version("0.14.0");
2922
2994
  program.command("init").description("Scaffold .blastproof/ (config, tests, sample tests) in the current directory").action(async () => {
2923
2995
  try {
2924
2996
  const result = await initProject(process.cwd());