blastproof 0.12.1 → 0.14.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +8 -4
- package/dist/cli.js +106 -9
- package/dist/cli.js.map +1 -1
- package/package.json +1 -1
package/README.md
CHANGED
|
@@ -141,11 +141,13 @@ Write steps that end in an observable result — text on the page, a count, a st
|
|
|
141
141
|
- fill the note field # cannot run
|
|
142
142
|
```
|
|
143
143
|
|
|
144
|
-
The agent is **forbidden from inventing values** — one it types must come from the step, from the page, or from an `{{env.*}}` placeholder
|
|
144
|
+
The agent is **forbidden from inventing values** — one it types must come from the step, from the page, or from an `{{env.*}}` placeholder — and the runner enforces it rather than asking. A `fill` or `select` whose value appears in none of those is refused: it is not typed, the agent is told which sources it may draw from, and the step fails on the retry budget if it keeps insisting.
|
|
145
145
|
|
|
146
|
-
|
|
146
|
+
The rule used to live only in the prompt, and a prompt instructs rather than enforces. Run against a real model, `fill the note field` did not fail — the agent made a value up, filled it, and the step **passed**, producing "This is a test note." on two runs and "This is a new note" on a third. A test going green over a value nobody wrote, differing between runs, is worse than a failure.
|
|
147
147
|
|
|
148
|
-
`
|
|
148
|
+
Two limits worth knowing. A value the page shows in one format and the field wants in another — `1234` in the step, `1,234.00` in the box — is refused, and the fix is to write the value the way it is typed. And a very short value (`3`) appears somewhere in almost any page, so it will pass; this closes fabricated content, not every fabricated character.
|
|
149
|
+
|
|
150
|
+
`run` also warns about it first — the same rule caught earlier, from the test file, on every path, before launching a browser or asking for a key:
|
|
149
151
|
|
|
150
152
|
```
|
|
151
153
|
Authoring (a step enters a value but names none):
|
|
@@ -156,7 +158,7 @@ Authoring (a step enters a value but names none):
|
|
|
156
158
|
|
|
157
159
|
Non-fatal by default — `--fail-on-authoring` turns it into exit 1 for teams enforcing it in CI. Taking the value from the page is fine and is not flagged: `fill the recipient field with the address shown on the confirmation page`.
|
|
158
160
|
|
|
159
|
-
**The
|
|
161
|
+
**The warning reads English only.** A suite written in another language runs exactly as well but is not inspected, and prints no warning saying so — silence from this check means "nothing found in English", never "this suite is clean". The runner's refusal has no such limit: it compares text rather than parsing grammar, so a suite in any language is still held to the rule at run time.
|
|
160
162
|
|
|
161
163
|
## Commands
|
|
162
164
|
|
|
@@ -198,6 +200,8 @@ ignore:
|
|
|
198
200
|
- "**/*.md"
|
|
199
201
|
```
|
|
200
202
|
|
|
203
|
+
The key is the file glob and the value is the routes it can affect — the opposite way round from a test file's own `routes:`, which is a plain list of the routes that test covers. Inverted, the map matches nothing at all and every run reports a diff that affected no page, so `blastproof` refuses a map written that way rather than running green against it.
|
|
204
|
+
|
|
201
205
|
Every changed file lands in one of three buckets: it matches `routes:` and contributes them, matches `ignore:` and is knowingly irrelevant, or **matches neither — nobody has said what it affects**. `--fail-on-unmapped` blocks on that third case, naming the files and both ways to resolve them.
|
|
202
206
|
|
|
203
207
|
**Nothing is ignored by default**, on purpose: a default that guesses on your behalf would hide the first files worth thinking about. The flag is additive — a run can meet `--min-score` and still be blocked here, because "the tests I ran passed" and "something changed that nobody classified" are different claims.
|
package/dist/cli.js
CHANGED
|
@@ -49,6 +49,9 @@ browser:
|
|
|
49
49
|
# max_retries_per_step: 3
|
|
50
50
|
|
|
51
51
|
# Impact hints mapping file globs to routes (used by \`blastproof run --impacted\`).
|
|
52
|
+
# Read each line as "if this file changed, these pages are at risk" \u2014 the key is the
|
|
53
|
+
# file glob, the value is the routes. Not the other way round: a test file's own
|
|
54
|
+
# \`routes:\` is a list of routes, and this one is a map keyed by source path.
|
|
52
55
|
routes:
|
|
53
56
|
"src/auth/**": ["/login"]
|
|
54
57
|
"src/cart/**": ["/cart", "/checkout"]
|
|
@@ -505,6 +508,11 @@ async function performAction(page, action, ctx) {
|
|
|
505
508
|
// src/runner/recovery.ts
|
|
506
509
|
var COMMIT_ACTIONS = /* @__PURE__ */ new Set(["click", "press"]);
|
|
507
510
|
var COMMIT_KEYS = /* @__PURE__ */ new Set(["Enter", "NumpadEnter", " ", "Space", "Spacebar"]);
|
|
511
|
+
var SOURCED_VALUE_ACTIONS = /* @__PURE__ */ new Set(["fill", "select"]);
|
|
512
|
+
var ENV_PLACEHOLDER2 = /^\s*\{\{env\.[A-Za-z_][A-Za-z0-9_]*\}\}\s*$/;
|
|
513
|
+
function normalise(text) {
|
|
514
|
+
return text.replace(/\s+/g, " ").trim().toLowerCase();
|
|
515
|
+
}
|
|
508
516
|
function describeAction(action) {
|
|
509
517
|
const target = action.target ? ` ${action.target.role ?? ""} "${action.target.name ?? action.target.text ?? ""}"` : "";
|
|
510
518
|
const value = action.value ? ` [${action.value}]` : "";
|
|
@@ -522,10 +530,41 @@ function identity(action) {
|
|
|
522
530
|
var StepRecovery = class {
|
|
523
531
|
performed = /* @__PURE__ */ new Set();
|
|
524
532
|
history = [];
|
|
533
|
+
/**
|
|
534
|
+
* Everything the model has been allowed to read this step, normalised: the
|
|
535
|
+
* step's own text, plus every snapshot it has been shown, plus every value it
|
|
536
|
+
* has already typed successfully.
|
|
537
|
+
*
|
|
538
|
+
* One accumulating haystack rather than the snapshot in hand (design D2). The
|
|
539
|
+
* model may read an order number on one page, navigate away, and type it into
|
|
540
|
+
* a box on another — at that moment the current snapshot does not contain it,
|
|
541
|
+
* and a check against only that snapshot would refuse a legitimate action.
|
|
542
|
+
*
|
|
543
|
+
* Bounded by construction: `max_snapshot_lines` caps each snapshot,
|
|
544
|
+
* `maxIterationsPerStep` caps how many arrive, and the instance dies with the
|
|
545
|
+
* step — which is also the whole of the "does not cross steps" rule.
|
|
546
|
+
*/
|
|
547
|
+
readable;
|
|
548
|
+
constructor(step) {
|
|
549
|
+
this.readable = [normalise(step)];
|
|
550
|
+
}
|
|
551
|
+
/**
|
|
552
|
+
* Records a snapshot the model was shown, as a source it may quote from.
|
|
553
|
+
*
|
|
554
|
+
* Takes the **masked** text (design D4). `executor.ts` shows the model
|
|
555
|
+
* `mask(snap)` and never the raw tree, so a secret rendered on the page is
|
|
556
|
+
* `***` to the model and cannot have been copied by it. Accepting the raw
|
|
557
|
+
* snapshot here would credit the model with access it never had, and would
|
|
558
|
+
* leave the executor and the model disagreeing about what the page said.
|
|
559
|
+
*/
|
|
560
|
+
observe(maskedSnapshot) {
|
|
561
|
+
this.readable.push(normalise(maskedSnapshot));
|
|
562
|
+
}
|
|
525
563
|
/** Records an action that was actually performed and succeeded. */
|
|
526
564
|
record(action, description, result) {
|
|
527
565
|
this.performed.add(identity(action));
|
|
528
566
|
this.history.push({ action: description, result });
|
|
567
|
+
if (action.value) this.readable.push(normalise(action.value));
|
|
529
568
|
}
|
|
530
569
|
/**
|
|
531
570
|
* The reason to refuse `action`, or `undefined` when it may be performed.
|
|
@@ -544,11 +583,43 @@ var StepRecovery = class {
|
|
|
544
583
|
* and #28 has now produced one on three applications.
|
|
545
584
|
*/
|
|
546
585
|
refusalFor(action) {
|
|
586
|
+
return this.repeatedCommitRefusal(action) ?? this.unsourcedValueRefusal(action);
|
|
587
|
+
}
|
|
588
|
+
repeatedCommitRefusal(action) {
|
|
547
589
|
if (!COMMIT_ACTIONS.has(action.action)) return void 0;
|
|
548
590
|
if (action.action === "press" && !COMMIT_KEYS.has(action.value ?? "")) return void 0;
|
|
549
591
|
if (!this.performed.has(identity(action))) return void 0;
|
|
550
592
|
return `refused: this exact action already succeeded earlier in this step, so it was NOT performed again. Repeating something that commits repeats whatever it changed in the application. If the page no longer shows that it worked, that is normal for a submit answered by a redirect \u2014 check the record of what you have already done. Verify the step's outcome another way, or fail the step.`;
|
|
551
593
|
}
|
|
594
|
+
/**
|
|
595
|
+
* Refuses a typed value that came from nowhere the model was entitled to read
|
|
596
|
+
* (design refuse-an-invented-value).
|
|
597
|
+
*
|
|
598
|
+
* `prompts.ts` has forbidden inventing a value since 0.7.0, and measured
|
|
599
|
+
* against a real model the rule simply does not hold: given `fill the note
|
|
600
|
+
* field`, the model supplied "This is a test note." twice and "This is a new
|
|
601
|
+
* note" once, and the step passed all three times. A prompt instructs; it does
|
|
602
|
+
* not enforce. The result is the false negative this project exists to remove
|
|
603
|
+
* — a green test over an input nobody wrote, differing run to run, with
|
|
604
|
+
* nothing in the report to say so because nothing knew.
|
|
605
|
+
*
|
|
606
|
+
* The comparison is `includes`, not equality: a value is legitimately a
|
|
607
|
+
* fragment of a step that also names the field, and of a snapshot that also
|
|
608
|
+
* holds the rest of the page.
|
|
609
|
+
*
|
|
610
|
+
* Unlike the authoring warning that predicts this before a run, nothing here
|
|
611
|
+
* parses English, so the guarantee holds for a suite written in any language.
|
|
612
|
+
*/
|
|
613
|
+
unsourcedValueRefusal(action) {
|
|
614
|
+
if (!SOURCED_VALUE_ACTIONS.has(action.action)) return void 0;
|
|
615
|
+
const value = action.value;
|
|
616
|
+
if (!value) return void 0;
|
|
617
|
+
if (ENV_PLACEHOLDER2.test(value)) return void 0;
|
|
618
|
+
const needle = normalise(value);
|
|
619
|
+
if (needle === "") return void 0;
|
|
620
|
+
if (this.readable.some((source) => source.includes(needle))) return void 0;
|
|
621
|
+
return `refused: the value "${value}" was NOT typed, because it appears neither in this step nor anywhere on the pages you have been shown. A value you type must come from the step, from the page, or from an {{env.*}} placeholder \u2014 one you make up would put the test's verdict on an input nobody wrote. Use a value the step or the page gives you, or fail the step and say it supplies none.`;
|
|
622
|
+
}
|
|
552
623
|
/** The step's history so far, oldest first, for the model's prompt. */
|
|
553
624
|
stepHistory() {
|
|
554
625
|
return this.history;
|
|
@@ -623,7 +694,7 @@ async function executeTest(page, test, options) {
|
|
|
623
694
|
let failedAttempts = 0;
|
|
624
695
|
let lastResult;
|
|
625
696
|
let stepFailedReason;
|
|
626
|
-
const recovery = new StepRecovery();
|
|
697
|
+
const recovery = new StepRecovery(step);
|
|
627
698
|
try {
|
|
628
699
|
while (true) {
|
|
629
700
|
if (iterations >= maxIterationsPerStep) {
|
|
@@ -636,12 +707,14 @@ async function executeTest(page, test, options) {
|
|
|
636
707
|
);
|
|
637
708
|
}
|
|
638
709
|
const snap = await takeSnapshot(page);
|
|
710
|
+
const maskedSnap = mask(snap);
|
|
711
|
+
recovery.observe(maskedSnap);
|
|
639
712
|
let action;
|
|
640
713
|
try {
|
|
641
714
|
action = await brain.nextAction({
|
|
642
715
|
step,
|
|
643
716
|
isSetup: setup,
|
|
644
|
-
snapshot:
|
|
717
|
+
snapshot: maskedSnap,
|
|
645
718
|
lastResult: lastResult === void 0 ? void 0 : mask(lastResult),
|
|
646
719
|
// Already masked when recorded, on the same boundary as everything
|
|
647
720
|
// else crossing into a prompt (design contained-recovery, D2).
|
|
@@ -672,16 +745,18 @@ async function executeTest(page, test, options) {
|
|
|
672
745
|
let judgment = await brain.judge(
|
|
673
746
|
mask(step),
|
|
674
747
|
mask(expectation),
|
|
675
|
-
|
|
748
|
+
maskedSnap,
|
|
676
749
|
recovery.stepHistory()
|
|
677
750
|
);
|
|
678
751
|
if (!judgment.pass) {
|
|
679
752
|
await waitForSettled(page);
|
|
680
753
|
const freshSnap = await takeSnapshot(page);
|
|
754
|
+
const maskedFresh = mask(freshSnap);
|
|
755
|
+
recovery.observe(maskedFresh);
|
|
681
756
|
judgment = await brain.judge(
|
|
682
757
|
mask(step),
|
|
683
758
|
mask(expectation),
|
|
684
|
-
|
|
759
|
+
maskedFresh,
|
|
685
760
|
recovery.stepHistory()
|
|
686
761
|
);
|
|
687
762
|
}
|
|
@@ -988,11 +1063,33 @@ var authSchema = z.object({
|
|
|
988
1063
|
});
|
|
989
1064
|
}
|
|
990
1065
|
});
|
|
1066
|
+
var ROUTE_SHAPED = /^\/[^*?]*$/;
|
|
1067
|
+
var FILE_SHAPED = /[*?]|\.[a-z0-9]{1,8}$/i;
|
|
1068
|
+
function findInvertedRouteEntries(routes) {
|
|
1069
|
+
const inverted = [];
|
|
1070
|
+
for (const [key, values] of Object.entries(routes)) {
|
|
1071
|
+
if (!ROUTE_SHAPED.test(key)) continue;
|
|
1072
|
+
const file = values.find((value) => FILE_SHAPED.test(value));
|
|
1073
|
+
if (file !== void 0) inverted.push({ key, file });
|
|
1074
|
+
}
|
|
1075
|
+
return inverted;
|
|
1076
|
+
}
|
|
991
1077
|
var configSchema = z.object({
|
|
992
1078
|
base_url: z.string().url(),
|
|
993
1079
|
llm: llmSchema.default({}),
|
|
994
1080
|
browser: browserSchema.default({}),
|
|
995
|
-
routes: z.record(z.array(z.string())).
|
|
1081
|
+
routes: z.record(z.array(z.string())).superRefine((value, ctx) => {
|
|
1082
|
+
const inverted = findInvertedRouteEntries(value);
|
|
1083
|
+
const first = inverted[0];
|
|
1084
|
+
if (!first) return;
|
|
1085
|
+
const others = inverted.length > 1 ? ` (and ${inverted.length - 1} more)` : "";
|
|
1086
|
+
ctx.addIssue({
|
|
1087
|
+
code: z.ZodIssueCode.custom,
|
|
1088
|
+
message: `is the wrong way round${others} \u2014 the key is the file glob, the value is the routes it affects.
|
|
1089
|
+
found: "${first.key}": ["${first.file}"]
|
|
1090
|
+
expected: "${first.file}": ["${first.key}"]`
|
|
1091
|
+
});
|
|
1092
|
+
}).optional(),
|
|
996
1093
|
/** Globs for files knowingly irrelevant to any route (docs, CI config, licences). */
|
|
997
1094
|
ignore: z.array(z.string()).optional(),
|
|
998
1095
|
/**
|
|
@@ -1208,7 +1305,7 @@ Rules:
|
|
|
1208
1305
|
- Return "done" when the current step's outcome holds \u2014 including when it already held before you acted, or was achieved by your previous action. "Already true" is done, never failure. Do not return "done" for work belonging to later steps.
|
|
1209
1306
|
- Return "fail" only when the step's outcome cannot be reached: the element is still absent after retries, the page cannot support the step, or an error blocks progress. Never return "fail" because the work appears to have been done already.
|
|
1210
1307
|
- If your previous action errored, re-read the fresh snapshot and choose an alternative element or approach. Do not repeat the exact same failing action.
|
|
1211
|
-
- Never invent a value. A value you type must come from the step, from the page, or from an {{env.*}} placeholder. If a step needs a value it does not give you, that is a failing step, not a gap for you to fill in.
|
|
1308
|
+
- Never invent a value. A value you type must come from the step, from the page, or from an {{env.*}} placeholder. This one is enforced, not merely asked: a fill or select whose value is in none of those is refused and not performed. If a step needs a value it does not give you, that is a failing step, not a gap for you to fill in.
|
|
1212
1309
|
- A record of the actions you already performed in this step may be shown to you. It is the ground truth about what happened, even when the page no longer shows it: a form that submitted successfully and came back empty looks exactly like one you never submitted. Do not redo work that record says you already did.
|
|
1213
1310
|
- \`***\` in a snapshot is a redacted secret \u2014 a password, token or key deliberately withheld from you. Seeing it is expected and is not a problem. A field showing \`***\` after you filled it from an {{env.VAR}} placeholder means the fill worked; treat that as success and move on. Never retry a fill because its value is redacted, and never report failure because a value was withheld.
|
|
1214
1311
|
- Keep reasoning to one short sentence.`;
|
|
@@ -1280,7 +1377,7 @@ Rules:
|
|
|
1280
1377
|
- Prefer the journey the changed files touch over a generic tour of the page. The changed files tell you which part of the page matters.
|
|
1281
1378
|
- **The test starts at the application's base URL, not at this route.** Begin with a step that navigates to the route and says what should be visible once it loads \u2014 "navigate to /support and verify the heading "Contact support" is shown". Without it the run opens the home page and every later step looks for controls that are not there.
|
|
1282
1379
|
- **Every step says what it should produce.** Name what must be true once the step has been carried out, not the action alone: "submit the support form and verify the confirmation page shows the ticket number", never "submit the support form". A step that names an action without an outcome asks the runner to judge whether something happened while looking at the page that succeeding produces \u2014 a submitted form comes back empty, a redirect moves the URL \u2014 and that is the shape behind several real failures.
|
|
1283
|
-
- **A step that enters a value writes the value.** "fill the subject field with Order not received", never "enter a subject". The runner is forbidden from inventing values,
|
|
1380
|
+
- **A step that enters a value writes the value.** "fill the subject field with Order not received", never "enter a subject". The runner is forbidden from inventing values, and enforces it: a fill whose value is in neither the step nor the page is refused, so a step that supplies none cannot be relied on to run.
|
|
1284
1381
|
- If a step needs a credential or any secret, write it as a placeholder like {{env.TEST_PASSWORD}}. Never write a real or invented password, token or key.
|
|
1285
1382
|
- Keep the whole test to a handful of steps: one journey, not an exhaustive suite.`;
|
|
1286
1383
|
}
|
|
@@ -2346,7 +2443,7 @@ function printAuthoring(result, cwd) {
|
|
|
2346
2443
|
console.error(` \u2192 ${suggestValueClause(step)}`);
|
|
2347
2444
|
}
|
|
2348
2445
|
console.error(
|
|
2349
|
-
"The runner is forbidden from inventing values
|
|
2446
|
+
"The runner is forbidden from inventing values and enforces it: one it cannot trace to the step, the page or an {{env.*}} placeholder is refused, so a step supplying none fails at run time. Naming the value here is what makes it run."
|
|
2350
2447
|
);
|
|
2351
2448
|
console.error("Steps are inspected in English only; steps in other languages are not checked.");
|
|
2352
2449
|
}
|
|
@@ -2893,7 +2990,7 @@ function parsePositiveNumber(flag) {
|
|
|
2893
2990
|
};
|
|
2894
2991
|
}
|
|
2895
2992
|
var program = new Command();
|
|
2896
|
-
program.name("blastproof").description("Open-source AI testing agent: plain-English YAML tests executed agentically on a real browser.").version("0.
|
|
2993
|
+
program.name("blastproof").description("Open-source AI testing agent: plain-English YAML tests executed agentically on a real browser.").version("0.14.0");
|
|
2897
2994
|
program.command("init").description("Scaffold .blastproof/ (config, tests, sample tests) in the current directory").action(async () => {
|
|
2898
2995
|
try {
|
|
2899
2996
|
const result = await initProject(process.cwd());
|