blastproof 0.13.0 → 0.15.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +8 -4
- package/dist/cli.js +96 -8
- package/dist/cli.js.map +1 -1
- package/package.json +1 -1
package/README.md
CHANGED
|
@@ -141,11 +141,15 @@ Write steps that end in an observable result — text on the page, a count, a st
|
|
|
141
141
|
- fill the note field # cannot run
|
|
142
142
|
```
|
|
143
143
|
|
|
144
|
-
The agent is **forbidden from inventing values** — one it types must come from the step, from the page, or from an `{{env.*}}` placeholder
|
|
144
|
+
The agent is **forbidden from inventing values** — one it types must come from the step, from the page, or from an `{{env.*}}` placeholder — and the runner enforces it rather than asking. A `fill` or `select` whose value appears in none of those is refused: it is not typed, the agent is told which sources it may draw from, and the step fails on the retry budget if it keeps insisting.
|
|
145
145
|
|
|
146
|
-
|
|
146
|
+
The rule used to live only in the prompt, and a prompt instructs rather than enforces. Run against a real model, `fill the note field` did not fail — the agent made a value up, filled it, and the step **passed**, producing "This is a test note." on two runs and "This is a new note" on a third. A test going green over a value nobody wrote, differing between runs, is worse than a failure.
|
|
147
147
|
|
|
148
|
-
`
|
|
148
|
+
A placeholder counts as a source **only when the step names that variable**. `fill the password field with {{env.TEST_PASSWORD}}` works; `fill the password field` does not become valid because the agent supplies `{{env.SOMETHING}}` itself. An agent cannot know the name of a variable nobody showed it, so one it produces is a guess — and a guessed variable would put a live credential into a field your test never pointed one at, in output that cannot redact a secret it was never told about.
|
|
149
|
+
|
|
150
|
+
Two limits worth knowing. A value the page shows in one format and the field wants in another — `1234` in the step, `1,234.00` in the box — is refused, and the fix is to write the value the way it is typed. And a very short value (`3`) appears somewhere in almost any page, so it will pass; this closes fabricated content, not every fabricated character.
|
|
151
|
+
|
|
152
|
+
`run` also warns about it first — the same rule caught earlier, from the test file, on every path, before launching a browser or asking for a key:
|
|
149
153
|
|
|
150
154
|
```
|
|
151
155
|
Authoring (a step enters a value but names none):
|
|
@@ -156,7 +160,7 @@ Authoring (a step enters a value but names none):
|
|
|
156
160
|
|
|
157
161
|
Non-fatal by default — `--fail-on-authoring` turns it into exit 1 for teams enforcing it in CI. Taking the value from the page is fine and is not flagged: `fill the recipient field with the address shown on the confirmation page`.
|
|
158
162
|
|
|
159
|
-
**The
|
|
163
|
+
**The warning reads English only.** A suite written in another language runs exactly as well but is not inspected, and prints no warning saying so — silence from this check means "nothing found in English", never "this suite is clean". The runner's refusal has no such limit: it compares text rather than parsing grammar, so a suite in any language is still held to the rule at run time.
|
|
160
164
|
|
|
161
165
|
## Commands
|
|
162
166
|
|
package/dist/cli.js
CHANGED
|
@@ -508,6 +508,14 @@ async function performAction(page, action, ctx) {
|
|
|
508
508
|
// src/runner/recovery.ts
|
|
509
509
|
var COMMIT_ACTIONS = /* @__PURE__ */ new Set(["click", "press"]);
|
|
510
510
|
var COMMIT_KEYS = /* @__PURE__ */ new Set(["Enter", "NumpadEnter", " ", "Space", "Spacebar"]);
|
|
511
|
+
var SOURCED_VALUE_ACTIONS = /* @__PURE__ */ new Set(["fill", "select"]);
|
|
512
|
+
function referencesOnlyNamedVars(value, step) {
|
|
513
|
+
const named = new Set(referencedEnvVars(step));
|
|
514
|
+
return referencedEnvVars(value).every((name) => named.has(name));
|
|
515
|
+
}
|
|
516
|
+
function normalise(text) {
|
|
517
|
+
return text.replace(/\s+/g, " ").trim().toLowerCase();
|
|
518
|
+
}
|
|
511
519
|
function describeAction(action) {
|
|
512
520
|
const target = action.target ? ` ${action.target.role ?? ""} "${action.target.name ?? action.target.text ?? ""}"` : "";
|
|
513
521
|
const value = action.value ? ` [${action.value}]` : "";
|
|
@@ -525,10 +533,50 @@ function identity(action) {
|
|
|
525
533
|
var StepRecovery = class {
|
|
526
534
|
performed = /* @__PURE__ */ new Set();
|
|
527
535
|
history = [];
|
|
536
|
+
/**
|
|
537
|
+
* Everything the model has been allowed to read this step, normalised: the
|
|
538
|
+
* step's own text, plus every snapshot it has been shown, plus every value it
|
|
539
|
+
* has already typed successfully.
|
|
540
|
+
*
|
|
541
|
+
* One accumulating haystack rather than the snapshot in hand (design D2). The
|
|
542
|
+
* model may read an order number on one page, navigate away, and type it into
|
|
543
|
+
* a box on another — at that moment the current snapshot does not contain it,
|
|
544
|
+
* and a check against only that snapshot would refuse a legitimate action.
|
|
545
|
+
*
|
|
546
|
+
* Bounded by construction: `max_snapshot_lines` caps each snapshot,
|
|
547
|
+
* `maxIterationsPerStep` caps how many arrive, and the instance dies with the
|
|
548
|
+
* step — which is also the whole of the "does not cross steps" rule.
|
|
549
|
+
*/
|
|
550
|
+
readable;
|
|
551
|
+
/**
|
|
552
|
+
* The step exactly as written, for deciding which `{{env.*}}` variables it
|
|
553
|
+
* names (design env-placeholder-must-be-named, D3). Kept alongside the
|
|
554
|
+
* normalised copy rather than derived from it: `readable` is lowercased, and
|
|
555
|
+
* environment variable names are case-sensitive, so `TOKEN` and `token` must
|
|
556
|
+
* stay distinguishable in the one place that decides which secret gets typed.
|
|
557
|
+
*/
|
|
558
|
+
step;
|
|
559
|
+
constructor(step) {
|
|
560
|
+
this.step = step;
|
|
561
|
+
this.readable = [normalise(step)];
|
|
562
|
+
}
|
|
563
|
+
/**
|
|
564
|
+
* Records a snapshot the model was shown, as a source it may quote from.
|
|
565
|
+
*
|
|
566
|
+
* Takes the **masked** text (design D4). `executor.ts` shows the model
|
|
567
|
+
* `mask(snap)` and never the raw tree, so a secret rendered on the page is
|
|
568
|
+
* `***` to the model and cannot have been copied by it. Accepting the raw
|
|
569
|
+
* snapshot here would credit the model with access it never had, and would
|
|
570
|
+
* leave the executor and the model disagreeing about what the page said.
|
|
571
|
+
*/
|
|
572
|
+
observe(maskedSnapshot) {
|
|
573
|
+
this.readable.push(normalise(maskedSnapshot));
|
|
574
|
+
}
|
|
528
575
|
/** Records an action that was actually performed and succeeded. */
|
|
529
576
|
record(action, description, result) {
|
|
530
577
|
this.performed.add(identity(action));
|
|
531
578
|
this.history.push({ action: description, result });
|
|
579
|
+
if (action.value) this.readable.push(normalise(action.value));
|
|
532
580
|
}
|
|
533
581
|
/**
|
|
534
582
|
* The reason to refuse `action`, or `undefined` when it may be performed.
|
|
@@ -547,11 +595,47 @@ var StepRecovery = class {
|
|
|
547
595
|
* and #28 has now produced one on three applications.
|
|
548
596
|
*/
|
|
549
597
|
refusalFor(action) {
|
|
598
|
+
return this.repeatedCommitRefusal(action) ?? this.unsourcedValueRefusal(action);
|
|
599
|
+
}
|
|
600
|
+
repeatedCommitRefusal(action) {
|
|
550
601
|
if (!COMMIT_ACTIONS.has(action.action)) return void 0;
|
|
551
602
|
if (action.action === "press" && !COMMIT_KEYS.has(action.value ?? "")) return void 0;
|
|
552
603
|
if (!this.performed.has(identity(action))) return void 0;
|
|
553
604
|
return `refused: this exact action already succeeded earlier in this step, so it was NOT performed again. Repeating something that commits repeats whatever it changed in the application. If the page no longer shows that it worked, that is normal for a submit answered by a redirect \u2014 check the record of what you have already done. Verify the step's outcome another way, or fail the step.`;
|
|
554
605
|
}
|
|
606
|
+
/**
|
|
607
|
+
* Refuses a typed value that came from nowhere the model was entitled to read
|
|
608
|
+
* (design refuse-an-invented-value).
|
|
609
|
+
*
|
|
610
|
+
* `prompts.ts` has forbidden inventing a value since 0.7.0, and measured
|
|
611
|
+
* against a real model the rule simply does not hold: given `fill the note
|
|
612
|
+
* field`, the model supplied "This is a test note." twice and "This is a new
|
|
613
|
+
* note" once, and the step passed all three times. A prompt instructs; it does
|
|
614
|
+
* not enforce. The result is the false negative this project exists to remove
|
|
615
|
+
* — a green test over an input nobody wrote, differing run to run, with
|
|
616
|
+
* nothing in the report to say so because nothing knew.
|
|
617
|
+
*
|
|
618
|
+
* The comparison is `includes`, not equality: a value is legitimately a
|
|
619
|
+
* fragment of a step that also names the field, and of a snapshot that also
|
|
620
|
+
* holds the rest of the page.
|
|
621
|
+
*
|
|
622
|
+
* Unlike the authoring warning that predicts this before a run, nothing here
|
|
623
|
+
* parses English, so the guarantee holds for a suite written in any language.
|
|
624
|
+
*/
|
|
625
|
+
unsourcedValueRefusal(action) {
|
|
626
|
+
if (!SOURCED_VALUE_ACTIONS.has(action.action)) return void 0;
|
|
627
|
+
const value = action.value;
|
|
628
|
+
if (!value) return void 0;
|
|
629
|
+
if (!referencesOnlyNamedVars(value, this.step)) {
|
|
630
|
+
const named = referencedEnvVars(this.step);
|
|
631
|
+
return `refused: the value "${value}" was NOT typed, because this step does not reference that {{env.*}} variable. ${named.length > 0 ? `This step references ${named.map((n) => `{{env.${n}}}`).join(", ")}.` : "This step references no environment variable."} A placeholder is a source only when the step names it \u2014 otherwise the test never asked for that secret to be typed here. Use a value this step supplies, or fail the step and say it supplies none.`;
|
|
632
|
+
}
|
|
633
|
+
if (referencedEnvVars(value).length > 0) return void 0;
|
|
634
|
+
const needle = normalise(value);
|
|
635
|
+
if (needle === "") return void 0;
|
|
636
|
+
if (this.readable.some((source) => source.includes(needle))) return void 0;
|
|
637
|
+
return `refused: the value "${value}" was NOT typed, because it appears neither in this step nor anywhere on the pages you have been shown. A value you type must come from the step, from the page, or from an {{env.*}} placeholder \u2014 one you make up would put the test's verdict on an input nobody wrote. Use a value the step or the page gives you, or fail the step and say it supplies none.`;
|
|
638
|
+
}
|
|
555
639
|
/** The step's history so far, oldest first, for the model's prompt. */
|
|
556
640
|
stepHistory() {
|
|
557
641
|
return this.history;
|
|
@@ -626,7 +710,7 @@ async function executeTest(page, test, options) {
|
|
|
626
710
|
let failedAttempts = 0;
|
|
627
711
|
let lastResult;
|
|
628
712
|
let stepFailedReason;
|
|
629
|
-
const recovery = new StepRecovery();
|
|
713
|
+
const recovery = new StepRecovery(step);
|
|
630
714
|
try {
|
|
631
715
|
while (true) {
|
|
632
716
|
if (iterations >= maxIterationsPerStep) {
|
|
@@ -639,12 +723,14 @@ async function executeTest(page, test, options) {
|
|
|
639
723
|
);
|
|
640
724
|
}
|
|
641
725
|
const snap = await takeSnapshot(page);
|
|
726
|
+
const maskedSnap = mask(snap);
|
|
727
|
+
recovery.observe(maskedSnap);
|
|
642
728
|
let action;
|
|
643
729
|
try {
|
|
644
730
|
action = await brain.nextAction({
|
|
645
731
|
step,
|
|
646
732
|
isSetup: setup,
|
|
647
|
-
snapshot:
|
|
733
|
+
snapshot: maskedSnap,
|
|
648
734
|
lastResult: lastResult === void 0 ? void 0 : mask(lastResult),
|
|
649
735
|
// Already masked when recorded, on the same boundary as everything
|
|
650
736
|
// else crossing into a prompt (design contained-recovery, D2).
|
|
@@ -675,16 +761,18 @@ async function executeTest(page, test, options) {
|
|
|
675
761
|
let judgment = await brain.judge(
|
|
676
762
|
mask(step),
|
|
677
763
|
mask(expectation),
|
|
678
|
-
|
|
764
|
+
maskedSnap,
|
|
679
765
|
recovery.stepHistory()
|
|
680
766
|
);
|
|
681
767
|
if (!judgment.pass) {
|
|
682
768
|
await waitForSettled(page);
|
|
683
769
|
const freshSnap = await takeSnapshot(page);
|
|
770
|
+
const maskedFresh = mask(freshSnap);
|
|
771
|
+
recovery.observe(maskedFresh);
|
|
684
772
|
judgment = await brain.judge(
|
|
685
773
|
mask(step),
|
|
686
774
|
mask(expectation),
|
|
687
|
-
|
|
775
|
+
maskedFresh,
|
|
688
776
|
recovery.stepHistory()
|
|
689
777
|
);
|
|
690
778
|
}
|
|
@@ -1233,7 +1321,7 @@ Rules:
|
|
|
1233
1321
|
- Return "done" when the current step's outcome holds \u2014 including when it already held before you acted, or was achieved by your previous action. "Already true" is done, never failure. Do not return "done" for work belonging to later steps.
|
|
1234
1322
|
- Return "fail" only when the step's outcome cannot be reached: the element is still absent after retries, the page cannot support the step, or an error blocks progress. Never return "fail" because the work appears to have been done already.
|
|
1235
1323
|
- If your previous action errored, re-read the fresh snapshot and choose an alternative element or approach. Do not repeat the exact same failing action.
|
|
1236
|
-
- Never invent a value. A value you type must come from the step, from the page, or from an {{env.*}} placeholder. If a step needs a value it does not give you, that is a failing step, not a gap for you to fill in.
|
|
1324
|
+
- Never invent a value. A value you type must come from the step, from the page, or from an {{env.*}} placeholder. This one is enforced, not merely asked: a fill or select whose value is in none of those is refused and not performed. If a step needs a value it does not give you, that is a failing step, not a gap for you to fill in.
|
|
1237
1325
|
- A record of the actions you already performed in this step may be shown to you. It is the ground truth about what happened, even when the page no longer shows it: a form that submitted successfully and came back empty looks exactly like one you never submitted. Do not redo work that record says you already did.
|
|
1238
1326
|
- \`***\` in a snapshot is a redacted secret \u2014 a password, token or key deliberately withheld from you. Seeing it is expected and is not a problem. A field showing \`***\` after you filled it from an {{env.VAR}} placeholder means the fill worked; treat that as success and move on. Never retry a fill because its value is redacted, and never report failure because a value was withheld.
|
|
1239
1327
|
- Keep reasoning to one short sentence.`;
|
|
@@ -1305,7 +1393,7 @@ Rules:
|
|
|
1305
1393
|
- Prefer the journey the changed files touch over a generic tour of the page. The changed files tell you which part of the page matters.
|
|
1306
1394
|
- **The test starts at the application's base URL, not at this route.** Begin with a step that navigates to the route and says what should be visible once it loads \u2014 "navigate to /support and verify the heading "Contact support" is shown". Without it the run opens the home page and every later step looks for controls that are not there.
|
|
1307
1395
|
- **Every step says what it should produce.** Name what must be true once the step has been carried out, not the action alone: "submit the support form and verify the confirmation page shows the ticket number", never "submit the support form". A step that names an action without an outcome asks the runner to judge whether something happened while looking at the page that succeeding produces \u2014 a submitted form comes back empty, a redirect moves the URL \u2014 and that is the shape behind several real failures.
|
|
1308
|
-
- **A step that enters a value writes the value.** "fill the subject field with Order not received", never "enter a subject". The runner is forbidden from inventing values,
|
|
1396
|
+
- **A step that enters a value writes the value.** "fill the subject field with Order not received", never "enter a subject". The runner is forbidden from inventing values, and enforces it: a fill whose value is in neither the step nor the page is refused, so a step that supplies none cannot be relied on to run.
|
|
1309
1397
|
- If a step needs a credential or any secret, write it as a placeholder like {{env.TEST_PASSWORD}}. Never write a real or invented password, token or key.
|
|
1310
1398
|
- Keep the whole test to a handful of steps: one journey, not an exhaustive suite.`;
|
|
1311
1399
|
}
|
|
@@ -2371,7 +2459,7 @@ function printAuthoring(result, cwd) {
|
|
|
2371
2459
|
console.error(` \u2192 ${suggestValueClause(step)}`);
|
|
2372
2460
|
}
|
|
2373
2461
|
console.error(
|
|
2374
|
-
"The runner is forbidden from inventing values
|
|
2462
|
+
"The runner is forbidden from inventing values and enforces it: one it cannot trace to the step, the page or an {{env.*}} placeholder is refused, so a step supplying none fails at run time. Naming the value here is what makes it run."
|
|
2375
2463
|
);
|
|
2376
2464
|
console.error("Steps are inspected in English only; steps in other languages are not checked.");
|
|
2377
2465
|
}
|
|
@@ -2918,7 +3006,7 @@ function parsePositiveNumber(flag) {
|
|
|
2918
3006
|
};
|
|
2919
3007
|
}
|
|
2920
3008
|
var program = new Command();
|
|
2921
|
-
program.name("blastproof").description("Open-source AI testing agent: plain-English YAML tests executed agentically on a real browser.").version("0.
|
|
3009
|
+
program.name("blastproof").description("Open-source AI testing agent: plain-English YAML tests executed agentically on a real browser.").version("0.15.0");
|
|
2922
3010
|
program.command("init").description("Scaffold .blastproof/ (config, tests, sample tests) in the current directory").action(async () => {
|
|
2923
3011
|
try {
|
|
2924
3012
|
const result = await initProject(process.cwd());
|