scenescout 3.1.0 → 3.2.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +28 -0
- package/dist/engine/lane.js +169 -0
- package/dist/engine/memory.js +24 -0
- package/dist/engine/policy.js +53 -2
- package/dist/engine/replay.js +64 -2
- package/dist/mcp-server.js +27 -22
- package/package.json +3 -2
- package/skills/scenescout/SKILL.md +2 -1
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,33 @@
|
|
|
1
1
|
# scenescout
|
|
2
2
|
|
|
3
|
+
## 3.2.0
|
|
4
|
+
|
|
5
|
+
### Minor Changes
|
|
6
|
+
|
|
7
|
+
- a3c7ed4: Parallel lanes hand their results back as one typed lane report, not prose. A new `scout_lane_report` tool serves both halves: without a reply it returns the paragraph to put in a lane's prompt (verdict, severity and category from closed sets, a calibrated confidence per decision, a bounded evidence signature, routes covered, what blocked the lane), and with one it parses what the lane handed back and returns the one-line fold or the reason the reply was refused. The instruction is generated from the same constants the parser checks and states every limit the parser enforces. The finding categories now live in one list shared by `scout_finding`, the lane report and the skill text. The skill's parallel-agents section tells the planner to use the tool and to run lanes at medium effort, the pick from a benchmark under `scripts/bench/` that measured both reply shapes at every effort level the CLI accepts.
|
|
8
|
+
|
|
9
|
+
### Patch Changes
|
|
10
|
+
|
|
11
|
+
- 0b99401: The read-only write policy no longer refuses a control because of a destructive word in its description. A card or tile that is a button carries prose in its accessible name, a title then a sentence about it, and a word in that sentence describes what the thing is for rather than what the click does. A label that is prose, longer than six words and containing a sentence, is now judged by its first six words, where the verb lives; every other label is still judged whole, so a long confirm button stays refused. The pattern still sees the whole label, so exemptions that look ahead ("reset filters") keep working. This is what refused the manager card on a sign-in page whose sentence mentioned orders that need sign-off. Separately, "sign off" is now read as the noun when the word before it says so ("Needs sign-off", "Awaiting sign-off", "Send for sign-off"), so those short labels are no longer refused either.
|
|
12
|
+
|
|
13
|
+
## 3.1.1
|
|
14
|
+
|
|
15
|
+
### Patch Changes
|
|
16
|
+
|
|
17
|
+
- 99562e7: The run's page now notices when the engine behind it has exited.
|
|
18
|
+
|
|
19
|
+
That address is served by a process. Once the process is gone, reloading it
|
|
20
|
+
gets the browser's own "site can't be reached" and the tab is lost, although
|
|
21
|
+
everything on the page was still readable a moment earlier. The served copy
|
|
22
|
+
now watches the engine, and when it goes says so in place: the page is still
|
|
23
|
+
good, reloading will not reach anything, and the copy that survives — with its
|
|
24
|
+
frames — is at the path it names, ready to copy. Leaving the page from then on
|
|
25
|
+
asks first, so a reflexive refresh cannot throw it away.
|
|
26
|
+
|
|
27
|
+
A browser will not follow a `file://` link from a served page, so the saved
|
|
28
|
+
copy cannot be opened from there; the path is offered instead. The copy on
|
|
29
|
+
disk carries none of this and stays a plain document with no script in it.
|
|
30
|
+
|
|
3
31
|
## 3.1.0
|
|
4
32
|
|
|
5
33
|
### Minor Changes
|
|
@@ -0,0 +1,169 @@
|
|
|
1
|
+
/**
|
|
2
|
+
* The lane report: what an agent exploring one area in parallel hands back to
|
|
3
|
+
* the agent that split the work, as typed decisions rather than prose.
|
|
4
|
+
*
|
|
5
|
+
* A parallel run has one planning agent and several lanes, each driving its
|
|
6
|
+
* own browser. A lane used to come back with a page of markdown, which the
|
|
7
|
+
* planner then had to read, re-judge and re-type into findings: a model pass
|
|
8
|
+
* over every lane's output, with nothing stopping a lane from inventing a
|
|
9
|
+
* category or writing "Low-Medium" as a severity.
|
|
10
|
+
*
|
|
11
|
+
* This module is the alternative, borrowed from decision-only models: every
|
|
12
|
+
* answer is a value from a closed set, carries a calibrated confidence, and
|
|
13
|
+
* is checked by a schema on arrival. The planner folds a report with a
|
|
14
|
+
* function call, not a reading.
|
|
15
|
+
*
|
|
16
|
+
* Both halves are reached through the scout_lane_report tool: with no reply
|
|
17
|
+
* it returns the instruction the planner puts in a lane's prompt, with one it
|
|
18
|
+
* parses it. The instruction is generated from the same constants the parser
|
|
19
|
+
* checks, and a table test asserts every value and every cap the parser
|
|
20
|
+
* enforces is in the text, because a limit the lane is not told refuses good
|
|
21
|
+
* replies: three of the first ten real replies were lost that way.
|
|
22
|
+
*
|
|
23
|
+
* Pure logic, no browser, so every rule here is table-tested.
|
|
24
|
+
*/
|
|
25
|
+
import { z } from "zod";
|
|
26
|
+
import { FINDING_CATEGORIES } from "./memory.js";
|
|
27
|
+
/** Upper bound on the decisions and routes arrays, so a report is always small enough to hold whole. */
|
|
28
|
+
export const LANE_MAX_ITEMS = 255;
|
|
29
|
+
/** A signature such as "GET /api/things 500", not a sentence. */
|
|
30
|
+
export const LANE_EVIDENCE_MAX = 160;
|
|
31
|
+
/** A lane names its own observations, and a name that says what was seen needs more than forty characters. */
|
|
32
|
+
export const LANE_OBSERVATION_MAX = 64;
|
|
33
|
+
/** One line saying what blocked the lane; the detail belongs in a finding. */
|
|
34
|
+
export const LANE_BLOCKED_BY_MAX = 200;
|
|
35
|
+
/** The planner chooses the lane name, so this cap is on the planner's side; it is stated all the same. */
|
|
36
|
+
export const LANE_NAME_MAX = 40;
|
|
37
|
+
/** A route as the lane saw it, query string included. */
|
|
38
|
+
export const LANE_ROUTE_MAX = 200;
|
|
39
|
+
export const LANE_SEVERITIES = ["high", "medium", "low"];
|
|
40
|
+
/** The same set scout_finding accepts, so a lane can report every finding it filed. */
|
|
41
|
+
export const LANE_CATEGORIES = FINDING_CATEGORIES;
|
|
42
|
+
export const LANE_VERDICTS = ["defect", "not_a_defect", "unsure"];
|
|
43
|
+
export const LANE_STATUSES = ["complete", "partial", "blocked"];
|
|
44
|
+
const Confidence = z.number().min(0).max(1);
|
|
45
|
+
/**
|
|
46
|
+
* One observation judged: is it a defect, how bad, of what kind, and how sure
|
|
47
|
+
* the lane is. A non-defect may still carry a category and a signature: an
|
|
48
|
+
* "unsure" with a tentative category tells the planner where to look.
|
|
49
|
+
*/
|
|
50
|
+
export const LaneDecision = z
|
|
51
|
+
.object({
|
|
52
|
+
/** A short id the lane assigns, unique within the report, so a decision can be traced back to what the lane saw. */
|
|
53
|
+
observation: z.string().min(1).max(LANE_OBSERVATION_MAX),
|
|
54
|
+
verdict: z.enum(LANE_VERDICTS),
|
|
55
|
+
severity: z.enum(LANE_SEVERITIES).nullable(),
|
|
56
|
+
category: z.enum(LANE_CATEGORIES).nullable(),
|
|
57
|
+
confidence: Confidence,
|
|
58
|
+
/** The same machine signature scout_finding takes as its cross-run dedup key. */
|
|
59
|
+
evidence: z.string().min(1).max(LANE_EVIDENCE_MAX).nullable(),
|
|
60
|
+
})
|
|
61
|
+
.strict()
|
|
62
|
+
.superRefine((d, ctx) => {
|
|
63
|
+
if (d.verdict === "defect" && (d.severity === null || d.category === null)) {
|
|
64
|
+
ctx.addIssue({ code: z.ZodIssueCode.custom, message: "a defect needs a severity and a category" });
|
|
65
|
+
}
|
|
66
|
+
});
|
|
67
|
+
export const LaneReport = z
|
|
68
|
+
.object({
|
|
69
|
+
lane: z.string().min(1).max(LANE_NAME_MAX),
|
|
70
|
+
status: z.enum(LANE_STATUSES),
|
|
71
|
+
decisions: z.array(LaneDecision).max(LANE_MAX_ITEMS),
|
|
72
|
+
/** The routes the lane covered, as it saw them; the planner normalises when it compares them with the split. */
|
|
73
|
+
routes: z.array(z.string().min(1).max(LANE_ROUTE_MAX)).max(LANE_MAX_ITEMS),
|
|
74
|
+
/** Required when the status is "blocked"; a "partial" lane may use it to say what cut it short; a "complete" one has nothing to put there. */
|
|
75
|
+
blocked_by: z.string().min(1).max(LANE_BLOCKED_BY_MAX).nullable(),
|
|
76
|
+
})
|
|
77
|
+
.strict()
|
|
78
|
+
.superRefine((r, ctx) => {
|
|
79
|
+
if (r.status === "blocked" && r.blocked_by === null) {
|
|
80
|
+
ctx.addIssue({ code: z.ZodIssueCode.custom, message: "a blocked lane must say what blocked it", path: ["blocked_by"] });
|
|
81
|
+
}
|
|
82
|
+
if (r.status === "complete" && r.blocked_by !== null) {
|
|
83
|
+
ctx.addIssue({ code: z.ZodIssueCode.custom, message: "a complete lane cannot also have been blocked", path: ["blocked_by"] });
|
|
84
|
+
}
|
|
85
|
+
const seen = new Set();
|
|
86
|
+
r.decisions.forEach((d, i) => {
|
|
87
|
+
if (seen.has(d.observation)) {
|
|
88
|
+
ctx.addIssue({ code: z.ZodIssueCode.custom, message: `observation "${d.observation}" is judged twice`, path: ["decisions", i, "observation"] });
|
|
89
|
+
}
|
|
90
|
+
seen.add(d.observation);
|
|
91
|
+
});
|
|
92
|
+
});
|
|
93
|
+
const FENCE_OPEN = /^```[a-z]*\s*\n/i;
|
|
94
|
+
const FENCE_CLOSE = /\n?```\s*$/;
|
|
95
|
+
/**
|
|
96
|
+
* The reply is one JSON object and nothing else. A fenced block around it is
|
|
97
|
+
* tolerated, because every model has been trained to add one; text before or
|
|
98
|
+
* after it is not, because the point is that no one has to read the reply.
|
|
99
|
+
* The reason names what was wrong, since the planner relays it to the lane.
|
|
100
|
+
*
|
|
101
|
+
* With `expectedLane`, a report that names another lane is refused: the
|
|
102
|
+
* planner joins reports to the split by name, and a reply that carries the
|
|
103
|
+
* wrong one would be folded into the wrong lane.
|
|
104
|
+
*/
|
|
105
|
+
export function parseLaneReport(text, expectedLane) {
|
|
106
|
+
let body = text.trim();
|
|
107
|
+
if (FENCE_OPEN.test(body)) {
|
|
108
|
+
if (!FENCE_CLOSE.test(body))
|
|
109
|
+
return { ok: false, reason: "the closing ``` must end the reply, with no text after it" };
|
|
110
|
+
body = body.replace(FENCE_OPEN, "").replace(FENCE_CLOSE, "").trim();
|
|
111
|
+
}
|
|
112
|
+
if (!body.startsWith("{"))
|
|
113
|
+
return { ok: false, reason: "the reply must be one JSON object, with no text before it" };
|
|
114
|
+
let raw;
|
|
115
|
+
try {
|
|
116
|
+
raw = JSON.parse(body);
|
|
117
|
+
}
|
|
118
|
+
catch (e) {
|
|
119
|
+
const message = e.message;
|
|
120
|
+
const trailing = /after JSON|Unexpected non-whitespace/i.test(message);
|
|
121
|
+
return { ok: false, reason: trailing ? "the reply must be one JSON object, with no text after it" : `not valid JSON: ${message}` };
|
|
122
|
+
}
|
|
123
|
+
const result = LaneReport.safeParse(raw);
|
|
124
|
+
if (!result.success) {
|
|
125
|
+
const issue = result.error.issues[0];
|
|
126
|
+
const at = issue.path.length ? ` at ${issue.path.join(".")}` : "";
|
|
127
|
+
return { ok: false, reason: `${issue.message}${at}` };
|
|
128
|
+
}
|
|
129
|
+
if (expectedLane !== undefined && result.data.lane !== expectedLane) {
|
|
130
|
+
return { ok: false, reason: `the report names lane "${result.data.lane}", but this reply was asked of lane "${expectedLane}"` };
|
|
131
|
+
}
|
|
132
|
+
return { ok: true, report: result.data };
|
|
133
|
+
}
|
|
134
|
+
/**
|
|
135
|
+
* The paragraph the planner puts in every lane's prompt. It is generated from
|
|
136
|
+
* the same constants the parser checks, so it can never list a value the
|
|
137
|
+
* parser would refuse, and it states every cap the parser enforces.
|
|
138
|
+
*/
|
|
139
|
+
export function laneReportInstruction(lane) {
|
|
140
|
+
return [
|
|
141
|
+
"Reply with ONE JSON object and nothing else — no prose before or after it, no explanation, no headings. A fenced ```json block is fine.",
|
|
142
|
+
`Shape: {"lane":${JSON.stringify(lane)},"status":<${quoteAll(LANE_STATUSES)}>,"decisions":[…],"routes":[…],"blocked_by":<string or null>}.`,
|
|
143
|
+
`Each decision: {"observation":<a short id for what was observed, unique in the report, at most ${LANE_OBSERVATION_MAX} characters>,"verdict":<${quoteAll(LANE_VERDICTS)}>,"severity":<${quoteAll(LANE_SEVERITIES)} or null>,"category":<${quoteAll(LANE_CATEGORIES)} or null>,"confidence":<0..1>,"evidence":<machine signature such as "GET /api/things 500", or null>}.`,
|
|
144
|
+
`A "defect" must carry a severity and a category. "evidence" is a signature, not a sentence: at most ${LANE_EVIDENCE_MAX} characters. "confidence" is how sure you are of the verdict, calibrated: 0.5 means a coin flip, 0.95 means you would bet on it.`,
|
|
145
|
+
`"routes" lists the routes you covered, each at most ${LANE_ROUTE_MAX} characters. "blocked_by" is one line of at most ${LANE_BLOCKED_BY_MAX} characters saying what stopped you: required when the status is "blocked", allowed with "partial", null with "complete"; the detail belongs in a finding. The lane name is at most ${LANE_NAME_MAX} characters. At most ${LANE_MAX_ITEMS} decisions and ${LANE_MAX_ITEMS} routes. Unknown keys are refused.`,
|
|
146
|
+
"The object IS your final report: whatever hands it back must hand back the object verbatim, not a summary of it.",
|
|
147
|
+
].join(" ");
|
|
148
|
+
}
|
|
149
|
+
function quoteAll(values) {
|
|
150
|
+
return values.map((v) => `"${v}"`).join("|");
|
|
151
|
+
}
|
|
152
|
+
/** The planner's fold: one line per lane, with the numbers it decides on. */
|
|
153
|
+
export function summarizeLaneReport(r) {
|
|
154
|
+
const defects = r.decisions.filter((d) => d.verdict === "defect");
|
|
155
|
+
const high = defects.filter((d) => d.severity === "high").length;
|
|
156
|
+
const unsure = r.decisions.filter((d) => d.verdict === "unsure").length;
|
|
157
|
+
const mean = r.decisions.length ? r.decisions.reduce((s, d) => s + d.confidence, 0) / r.decisions.length : 0;
|
|
158
|
+
const parts = [
|
|
159
|
+
r.status,
|
|
160
|
+
`${r.decisions.length} judged`,
|
|
161
|
+
`${defects.length} defects (${high} high)`,
|
|
162
|
+
`${unsure} unsure`,
|
|
163
|
+
`mean confidence ${mean.toFixed(2)}`,
|
|
164
|
+
`${r.routes.length} routes`,
|
|
165
|
+
];
|
|
166
|
+
if (r.blocked_by)
|
|
167
|
+
parts.push(`blocked by ${r.blocked_by}`);
|
|
168
|
+
return `${r.lane}: ${parts.join(", ")}`;
|
|
169
|
+
}
|
package/dist/engine/memory.js
CHANGED
|
@@ -1,6 +1,30 @@
|
|
|
1
1
|
import fs from "node:fs";
|
|
2
2
|
import path from "node:path";
|
|
3
3
|
import { normalizePath, shortHash } from "./fingerprint.js";
|
|
4
|
+
/**
|
|
5
|
+
* The kinds a finding can be. One list: scout_finding's input schema, the lane
|
|
6
|
+
* report a parallel agent hands back, and the skill text all read it from
|
|
7
|
+
* here, so a category cannot exist in one and be refused by another.
|
|
8
|
+
*/
|
|
9
|
+
export const FINDING_CATEGORIES = [
|
|
10
|
+
"console-error",
|
|
11
|
+
"page-error",
|
|
12
|
+
"http-error",
|
|
13
|
+
"network",
|
|
14
|
+
"dead-end",
|
|
15
|
+
"ux-confusing",
|
|
16
|
+
"ux-polish",
|
|
17
|
+
"visual",
|
|
18
|
+
"a11y",
|
|
19
|
+
"permission-leak",
|
|
20
|
+
"data-inconsistency",
|
|
21
|
+
"stale-state",
|
|
22
|
+
"data-loss",
|
|
23
|
+
"performance",
|
|
24
|
+
"security",
|
|
25
|
+
"missing-testid",
|
|
26
|
+
"other",
|
|
27
|
+
];
|
|
4
28
|
/** An element class must appear on this many routes at minimum before it can count as shared chrome. */
|
|
5
29
|
const CHROME_MIN_ROUTES = 4;
|
|
6
30
|
/** …and on at least this share of all visited routes (a majority — "it's on every page"). */
|
package/dist/engine/policy.js
CHANGED
|
@@ -22,7 +22,11 @@ const DESTRUCTIVE_PATTERNS = [
|
|
|
22
22
|
/\breset\b(?!\s+(filters?|zoom|search|view|sort|order|form|password|layout|columns?|selection|preferences?|defaults?))/i,
|
|
23
23
|
/\bdiscard\b/i,
|
|
24
24
|
/\bcancel subscription\b/i,
|
|
25
|
-
|
|
25
|
+
// The verb: a control that signs the user off, alone or after another verb
|
|
26
|
+
// ("Save and sign off"). The noun is how approval apps label things, and it
|
|
27
|
+
// is told by the word before it ("Needs sign-off", "Awaiting sign-off",
|
|
28
|
+
// "Send for sign-off"); those destroy nothing. Same shape as the reset rule.
|
|
29
|
+
/(?<!\b(?:needs?|awaiting|awaits|pending|requires?|required|for|before|after|without|of|the|a|an|its|their|your|my)\s)\bsign(?: |-)?off\b/i,
|
|
26
30
|
// The tool is generic — destructive labels come in many languages.
|
|
27
31
|
/\b(eliminar|borrar|suprimir)\b/i, // es
|
|
28
32
|
/\b(excluir|apagar|remover)\b/i, // pt
|
|
@@ -74,8 +78,55 @@ export function isDestructiveWire(url, body) {
|
|
|
74
78
|
}
|
|
75
79
|
/** Auth/session flows must work even under strict write policies (login, token refresh, logout). */
|
|
76
80
|
export const AUTH_FLOW_RE = /\/(auth|login|logout|signin|sign-in|signup|sign-up|session|token|verify|oauth|sso|password)\b/i;
|
|
81
|
+
/**
|
|
82
|
+
* A control's label is a verb phrase: "Delete", "Remove user", "Archive the
|
|
83
|
+
* project". A card or tile that is a button carries prose in its accessible
|
|
84
|
+
* name, the title then a sentence about it ("Manager Approves or rejects
|
|
85
|
+
* orders that need sign-off."), and a destructive word in that sentence
|
|
86
|
+
* describes what the thing is for; it is not the command the click sends, the
|
|
87
|
+
* same distinction the wire policy draws for words inside a POST body.
|
|
88
|
+
*
|
|
89
|
+
* So a label that is prose, longer than this many words AND containing a
|
|
90
|
+
* sentence, is judged by its head, the first few words, where the imperative
|
|
91
|
+
* lives: a match counts only when it STARTS there. The pattern still sees the
|
|
92
|
+
* whole label, so "reset filters" straddling the boundary keeps the context
|
|
93
|
+
* its exemption looks ahead at. Everything else is judged whole: a short
|
|
94
|
+
* label, a long one with no sentence in it ("Yes, I am sure I want to delete
|
|
95
|
+
* this"), a script written without spaces.
|
|
96
|
+
*
|
|
97
|
+
* A confirm button whose long label ends in a full stop and puts the verb
|
|
98
|
+
* late is the residue this lets through. The wire policy below still blocks
|
|
99
|
+
* the DELETE, PUT, PATCH or verb-path POST it would send; a plain POST to a
|
|
100
|
+
* neutral path passes in read-only, as it does for any control this list does
|
|
101
|
+
* not name.
|
|
102
|
+
*/
|
|
103
|
+
export const LABEL_HEAD_WORDS = 6;
|
|
104
|
+
/** The head: up to LABEL_HEAD_WORDS whitespace-separated words from the start. */
|
|
105
|
+
const HEAD_RE = new RegExp(`^\\s*(?:\\S+\\s+){0,${LABEL_HEAD_WORDS - 1}}\\S+`);
|
|
106
|
+
/** A sentence ends in the label: a full stop, question or exclamation mark followed by space or the end. */
|
|
107
|
+
const SENTENCE_RE = /[.!?](?:\s|$)/;
|
|
108
|
+
/**
|
|
109
|
+
* Where a match must start to count: the end of the head for prose, the end
|
|
110
|
+
* of the label for everything else.
|
|
111
|
+
*/
|
|
112
|
+
function matchLimit(label) {
|
|
113
|
+
if (!SENTENCE_RE.test(label))
|
|
114
|
+
return label.length;
|
|
115
|
+
const head = HEAD_RE.exec(label);
|
|
116
|
+
if (!head || head[0].length === label.trimEnd().length)
|
|
117
|
+
return label.length;
|
|
118
|
+
return head[0].length;
|
|
119
|
+
}
|
|
77
120
|
export function isDestructive(...labels) {
|
|
78
|
-
return labels.some((label) =>
|
|
121
|
+
return labels.some((label) => {
|
|
122
|
+
if (typeof label !== "string" || label.length === 0)
|
|
123
|
+
return false;
|
|
124
|
+
const limit = matchLimit(label);
|
|
125
|
+
return DESTRUCTIVE_PATTERNS.some((re) => {
|
|
126
|
+
const m = re.exec(label);
|
|
127
|
+
return m !== null && m.index < limit;
|
|
128
|
+
});
|
|
129
|
+
});
|
|
79
130
|
}
|
|
80
131
|
export function destructiveRefusal(label, mode = "read-only") {
|
|
81
132
|
return (`REFUSED by ${mode} policy: "${label}" matches a destructive-action pattern. ` +
|
package/dist/engine/replay.js
CHANGED
|
@@ -254,8 +254,8 @@ export function renderMarkdown(md, evidence = [], framePrefix = "", savedAt = ""
|
|
|
254
254
|
return out.join("\n");
|
|
255
255
|
}
|
|
256
256
|
const STYLE = `
|
|
257
|
-
:root { color-scheme: light dark; --bg:#f6f7f9; --panel:#fff; --line:#d9dde3; --text:#15181d; --muted:#5d6673; --accent:#2563eb; --bad:#b91c1c; }
|
|
258
|
-
@media (prefers-color-scheme: dark) { :root { --bg:#0e1116; --panel:#161a21; --line:#2a303a; --text:#e6e9ee; --muted:#98a2b3; --accent:#7aa2ff; --bad:#fca5a5; } }
|
|
257
|
+
:root { color-scheme: light dark; --bg:#f6f7f9; --panel:#fff; --line:#d9dde3; --text:#15181d; --muted:#5d6673; --accent:#2563eb; --bad:#b91c1c; --warn:#7c4a03; --warn-bg:#fef3c7; }
|
|
258
|
+
@media (prefers-color-scheme: dark) { :root { --bg:#0e1116; --panel:#161a21; --line:#2a303a; --text:#e6e9ee; --muted:#98a2b3; --accent:#7aa2ff; --bad:#fca5a5; --warn:#fbbf24; --warn-bg:#3a2a08; } }
|
|
259
259
|
* { box-sizing: border-box; }
|
|
260
260
|
body { margin:0; background:var(--bg); color:var(--text); font:15px/1.55 system-ui,-apple-system,"Segoe UI",sans-serif; }
|
|
261
261
|
header { position:sticky; top:0; z-index:2; display:flex; flex-wrap:wrap; gap:8px 16px; align-items:baseline; padding:14px 20px; background:var(--panel); border-bottom:1px solid var(--line); }
|
|
@@ -303,6 +303,13 @@ details.evidence > summary { cursor:pointer; color:var(--muted); font-size:13px;
|
|
|
303
303
|
details.evidence .shots { display:flex; flex-wrap:wrap; gap:14px; margin-top:12px; }
|
|
304
304
|
details.evidence figure { margin:0; max-width:min(100%,460px); }
|
|
305
305
|
details.evidence img { width:100%; max-height:300px; object-fit:cover; object-position:top; border:1px solid var(--line); border-radius:6px; display:block; background:var(--panel); }
|
|
306
|
+
#gone-bar { position:sticky; top:0; z-index:3; margin:0; padding:14px 22px; background:var(--warn-bg); color:var(--warn);
|
|
307
|
+
border-bottom:1px solid var(--line); }
|
|
308
|
+
#gone-bar p { margin:0 0 6px; max-width:68ch; }
|
|
309
|
+
#gone-bar .path { display:flex; flex-wrap:wrap; gap:10px; align-items:center; }
|
|
310
|
+
#gone-bar code { font:12px/1.6 ui-monospace,SFMono-Regular,Menlo,Consolas,monospace; overflow-wrap:anywhere; user-select:all; }
|
|
311
|
+
#gone-bar button { font:inherit; padding:4px 10px; border:1px solid currentColor; border-radius:6px; background:transparent; color:inherit; cursor:pointer; }
|
|
312
|
+
#gone-bar .why { margin:0; font-size:12px; opacity:.8; }
|
|
306
313
|
.onDisk { display:block; margin-top:4px; color:var(--muted); font:11px/1.5 ui-monospace,SFMono-Regular,Menlo,Consolas,monospace; overflow-wrap:anywhere; }
|
|
307
314
|
header .served { flex:1 1 100%; margin:6px 0 0; color:var(--muted); font-size:12px; }
|
|
308
315
|
header .served code { font:11px/1.5 ui-monospace,SFMono-Regular,Menlo,Consolas,monospace; overflow-wrap:anywhere; }
|
|
@@ -312,6 +319,53 @@ figure.gone img, a.gone img { display:none; }
|
|
|
312
319
|
a.frame.gone { display:block; max-width:min(100%,720px); }
|
|
313
320
|
details.evidence figcaption { margin-top:4px; color:var(--muted); font:11px/1.5 ui-monospace,SFMono-Regular,Menlo,Consolas,monospace; overflow-wrap:anywhere; }
|
|
314
321
|
`;
|
|
322
|
+
/**
|
|
323
|
+
* Only the SERVED copy carries this. It watches the engine, and the moment the
|
|
324
|
+
* engine is gone it says so, because from then on a refresh gets the browser's
|
|
325
|
+
* "site can't be reached" and the tab is lost for nothing — this page is still
|
|
326
|
+
* perfectly readable, and the copy on disk is the one that survives.
|
|
327
|
+
*
|
|
328
|
+
* A browser refuses to follow a `file://` link from an `http://` page, so the
|
|
329
|
+
* saved copy cannot be opened from here. The path is offered to copy instead.
|
|
330
|
+
*
|
|
331
|
+
* No template literals and no backticks: this whole module is one.
|
|
332
|
+
*/
|
|
333
|
+
function exitWatch(savedFile) {
|
|
334
|
+
return ("<script>(function () {\n" +
|
|
335
|
+
" var saved = " +
|
|
336
|
+
JSON.stringify(savedFile) +
|
|
337
|
+
";\n" +
|
|
338
|
+
" var gone = false;\n" +
|
|
339
|
+
" var bar = document.getElementById('gone-bar');\n" +
|
|
340
|
+
" function ended() {\n" +
|
|
341
|
+
" if (gone) return;\n" +
|
|
342
|
+
" gone = true;\n" +
|
|
343
|
+
" bar.hidden = false;\n" +
|
|
344
|
+
" document.body.classList.add('engine-gone');\n" +
|
|
345
|
+
" }\n" +
|
|
346
|
+
" function check() {\n" +
|
|
347
|
+
" fetch('api/status', { cache: 'no-store' })\n" +
|
|
348
|
+
" .then(function (r) { if (!r.ok && r.status >= 500) ended(); })\n" +
|
|
349
|
+
" .catch(ended);\n" +
|
|
350
|
+
" }\n" +
|
|
351
|
+
" setInterval(check, 5000);\n" +
|
|
352
|
+
" check();\n" +
|
|
353
|
+
" document.getElementById('gone-copy').addEventListener('click', function () {\n" +
|
|
354
|
+
" var self = this;\n" +
|
|
355
|
+
" function said(text) { self.textContent = text; setTimeout(function () { self.textContent = 'Copy the path'; }, 2000); }\n" +
|
|
356
|
+
" if (navigator.clipboard && navigator.clipboard.writeText) {\n" +
|
|
357
|
+
" navigator.clipboard.writeText(saved).then(function () { said('Copied'); }, function () { said('Select it and copy'); });\n" +
|
|
358
|
+
" } else { said('Select it and copy'); }\n" +
|
|
359
|
+
" });\n" +
|
|
360
|
+
" // Refresh, close, back — all of them lose this page once the engine has\n" +
|
|
361
|
+
" // exited, and none of them can be told apart here. Ask first.\n" +
|
|
362
|
+
" window.addEventListener('beforeunload', function (e) {\n" +
|
|
363
|
+
" if (!gone) return;\n" +
|
|
364
|
+
" e.preventDefault();\n" +
|
|
365
|
+
" e.returnValue = '';\n" +
|
|
366
|
+
" });\n" +
|
|
367
|
+
"})();</script>");
|
|
368
|
+
}
|
|
315
369
|
/** The whole document: one file, no external assets, opens from the file system. */
|
|
316
370
|
export function buildReplayHtml(input) {
|
|
317
371
|
const framed = input.sessions.some((s) => s.steps.some((x) => x.frame));
|
|
@@ -333,6 +387,13 @@ export function buildReplayHtml(input) {
|
|
|
333
387
|
${savedAt ? `<p class="served">This page is served by the engine and goes when it does. The copy that stays is <code>${escapeHtml(savedAt)}/report.html</code>, beside the frames it shows.</p>` : ""}
|
|
334
388
|
<nav><a href="#report">Report</a><a href="#steps">Steps</a></nav>
|
|
335
389
|
</header>
|
|
390
|
+
${savedAt
|
|
391
|
+
? `<div id="gone-bar" role="alert" data-testid="run-engine-gone" hidden>
|
|
392
|
+
<p><b>The engine has exited.</b> Everything here is still readable, but reloading this address will not reach anything. The copy that survives, with its frames, is on disk:</p>
|
|
393
|
+
<p class="path"><code>${escapeHtml(savedAt)}/report.html</code> <button type="button" id="gone-copy" data-testid="run-copy-path">Copy the path</button></p>
|
|
394
|
+
<p class="why">A browser will not open a file from a served page, so it has to be opened from there.</p>
|
|
395
|
+
</div>`
|
|
396
|
+
: ""}
|
|
336
397
|
<main>
|
|
337
398
|
<h2 id="report">Report</h2>
|
|
338
399
|
${renderMarkdown(input.markdown, input.evidence ?? [], prefix, savedAt)}
|
|
@@ -342,6 +403,7 @@ ${renderMarkdown(input.markdown, input.evidence ?? [], prefix, savedAt)}
|
|
|
342
403
|
: "This run was not recorded, so there are no frames — attach with record:true to keep them."}</p>
|
|
343
404
|
${sessions || "<p>No session recorded any action.</p>"}
|
|
344
405
|
</main>
|
|
406
|
+
${savedAt ? exitWatch(`${savedAt}/report.html`) : ""}
|
|
345
407
|
</body>
|
|
346
408
|
</html>
|
|
347
409
|
`;
|
package/dist/mcp-server.js
CHANGED
|
@@ -34,7 +34,8 @@ import { ErrorCode, GetPromptRequestSchema, ListPromptsRequestSchema, McpError }
|
|
|
34
34
|
import { z } from "zod";
|
|
35
35
|
import { BrowserEngine } from "./engine/browser.js";
|
|
36
36
|
import { reapOrphanBrowsers } from "./engine/reaper.js";
|
|
37
|
-
import { MemoryStore, redactSecrets } from "./engine/memory.js";
|
|
37
|
+
import { FINDING_CATEGORIES, MemoryStore, redactSecrets } from "./engine/memory.js";
|
|
38
|
+
import { LANE_NAME_MAX, laneReportInstruction, parseLaneReport, summarizeLaneReport } from "./engine/lane.js";
|
|
38
39
|
import { SessionQueue, withWatchdog } from "./engine/dispatch.js";
|
|
39
40
|
import { FIXTURE_KINDS } from "./engine/fixtures.js";
|
|
40
41
|
import { feedForSession, LIVE_ENV, writeStatusFile, LIVE_TOKEN_FILE, LiveServer, StatusBoard, } from "./engine/live.js";
|
|
@@ -483,6 +484,30 @@ server.server.setRequestHandler(GetPromptRequestSchema, (request) => {
|
|
|
483
484
|
}
|
|
484
485
|
return { messages: [{ role: "user", content: { type: "text", text: message } }] };
|
|
485
486
|
});
|
|
487
|
+
// The lane report: how a parallel agent hands its results back to the planner
|
|
488
|
+
// as typed decisions. One tool for both halves, so the instruction a lane is
|
|
489
|
+
// given and the parser its reply meets are the same code.
|
|
490
|
+
server.registerTool("scout_lane_report", {
|
|
491
|
+
description: "For a run split across parallel agents (lanes). Without `reply`: returns the paragraph to put in a lane's prompt, telling it to hand back ONE typed JSON object (verdicts, severities and categories from closed sets, a calibrated confidence per decision, routes covered, what blocked it). " +
|
|
492
|
+
"With `reply`: parses what the lane handed back and returns the one-line fold (defects, highs, unsure, mean confidence, routes) or the reason it was refused, to relay to the lane once. Touches no browser.",
|
|
493
|
+
inputSchema: {
|
|
494
|
+
lane: z.string().min(1).max(LANE_NAME_MAX).describe("The lane's name, as used in its session"),
|
|
495
|
+
reply: z.string().optional().describe("The text the lane handed back; omit to get the instruction instead"),
|
|
496
|
+
},
|
|
497
|
+
}, async ({ lane, reply }) => {
|
|
498
|
+
try {
|
|
499
|
+
if (reply === undefined)
|
|
500
|
+
return { content: [{ type: "text", text: laneReportInstruction(lane) }] };
|
|
501
|
+
const parsed = parseLaneReport(reply, lane);
|
|
502
|
+
const out = parsed.ok
|
|
503
|
+
? `Lane report accepted — ${summarizeLaneReport(parsed.report)}`
|
|
504
|
+
: `Lane report REFUSED: ${parsed.reason}. Ask the lane once for the corrected object; do not re-judge its prose.`;
|
|
505
|
+
return { content: [{ type: "text", text: out }] };
|
|
506
|
+
}
|
|
507
|
+
catch (err) {
|
|
508
|
+
return errorText(err);
|
|
509
|
+
}
|
|
510
|
+
});
|
|
486
511
|
server.registerTool("scout_scan", {
|
|
487
512
|
description: "Scan a project directory to discover the frontend workspace, framework, routes, dev command, Playwright auth storage states, and testid conventions. Run this first.",
|
|
488
513
|
inputSchema: { projectPath: z.string().describe("Absolute path to the project root") },
|
|
@@ -955,27 +980,7 @@ server.registerTool("scout_finding", {
|
|
|
955
980
|
description: "Record a structured finding (bug, UX issue, or improvement). Deduplicates across runs; automatically captures the recent action trace as the repro. Use for anything worth reporting: crashes, oracle violations you confirmed, dead ends, confusing UX, permission leaks, missing testids — and design-audit improvement opportunities (ux-polish) with their concrete measurements.",
|
|
956
981
|
inputSchema: {
|
|
957
982
|
severity: z.enum(["high", "medium", "low"]),
|
|
958
|
-
category: z
|
|
959
|
-
.enum([
|
|
960
|
-
"console-error",
|
|
961
|
-
"page-error",
|
|
962
|
-
"http-error",
|
|
963
|
-
"network",
|
|
964
|
-
"dead-end",
|
|
965
|
-
"ux-confusing",
|
|
966
|
-
"ux-polish",
|
|
967
|
-
"visual",
|
|
968
|
-
"a11y",
|
|
969
|
-
"permission-leak",
|
|
970
|
-
"data-inconsistency",
|
|
971
|
-
"stale-state",
|
|
972
|
-
"data-loss",
|
|
973
|
-
"performance",
|
|
974
|
-
"security",
|
|
975
|
-
"missing-testid",
|
|
976
|
-
"other",
|
|
977
|
-
])
|
|
978
|
-
.describe("Pick the closest — use 'other' only when nothing fits"),
|
|
983
|
+
category: z.enum(FINDING_CATEGORIES).describe("Pick the closest — use 'other' only when nothing fits"),
|
|
979
984
|
title: z.string().describe("One-line summary of the defect"),
|
|
980
985
|
detail: z.string().describe("What happened, what was expected, and the evidence"),
|
|
981
986
|
evidence: z
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "scenescout",
|
|
3
|
-
"version": "3.
|
|
3
|
+
"version": "3.2.0",
|
|
4
4
|
"description": "SceneScout — exploratory UI testing for AI coding agents. An MCP server that gives any agent (Claude Code, Cursor, VS Code Copilot, Codex, Gemini CLI and others) a structured view of a running web app, always-on oracles, a network-level write policy, memory across runs and a gap-checked report.",
|
|
5
5
|
"license": "MIT",
|
|
6
6
|
"author": "brunoboto96",
|
|
@@ -71,7 +71,7 @@
|
|
|
71
71
|
"mcp-check": "npm run build && npm run mcp-check:run",
|
|
72
72
|
"mcp-check:run": "tsx scripts/mcp-check.ts",
|
|
73
73
|
"test": "npm run build && npm run test:unit && npm run smoke:run && npm run mcp-check:run",
|
|
74
|
-
"test:unit": "npm run scan-test && npm run oracle-test && npm run policy-test && npm run fixture-test && npm run dispatch-test && npm run design-test && npm run contract-test && npm run memory-test && npm run install-test && npm run live-test && npm run demo-test && npm run hygiene-test",
|
|
74
|
+
"test:unit": "npm run scan-test && npm run oracle-test && npm run policy-test && npm run fixture-test && npm run dispatch-test && npm run design-test && npm run contract-test && npm run lane-test && npm run memory-test && npm run install-test && npm run live-test && npm run demo-test && npm run hygiene-test",
|
|
75
75
|
"scan-test": "tsx scripts/scan-test.ts",
|
|
76
76
|
"oracle-test": "tsx --test scripts/oracle-test.ts",
|
|
77
77
|
"policy-test": "tsx --test scripts/policy-test.ts",
|
|
@@ -79,6 +79,7 @@
|
|
|
79
79
|
"dispatch-test": "tsx --test scripts/dispatch-test.ts",
|
|
80
80
|
"design-test": "tsx --test scripts/design-test.ts",
|
|
81
81
|
"contract-test": "tsx --test scripts/contract-test.ts",
|
|
82
|
+
"lane-test": "tsx --test scripts/lane-test.ts",
|
|
82
83
|
"memory-test": "tsx --test scripts/memory-test.ts",
|
|
83
84
|
"install-test": "tsx --test scripts/install-test.ts",
|
|
84
85
|
"live-test": "tsx --test scripts/live-test.ts",
|
|
@@ -71,6 +71,7 @@ Some flows need a TEAM — a document one role submits and another approves, a r
|
|
|
71
71
|
- **Infer the PERSONA behind each role, and write it down.** From what a role can see and do (its nav, its dashboard, the capability matrix in the report), state what this person is FOR: "qa = reviewer — approves orders, assigns reviewers, no admin" / "user = front-line user — reads documents, completes reviews, raises orders". Record it with `scout_note {section:'roles'}`. Then test the persona's WORLD, not just the permissions: does the operator's landing page serve an operator? Is anything they need N clicks deep? The capability matrix's divergent rows are questions, not verdicts — each is either a correct boundary or a gap ("should this role be able to do this?"); say which you believe it is and why.
|
|
72
72
|
- `scout_close {all: true}` at the end of a multi-role run; `scout_close {session}` to drop one role early.
|
|
73
73
|
- **Several agents in parallel** (subagents or a workflow, each driving its own session): each agent attaches its session when it STARTS and closes it by name when it FINISHES. Never open sessions ahead for agents that have not started, and never hand an open session from one agent to the next: an agent waiting for its turn should hold no browser. Give each session an `objective` when you attach it (`scout_attach {session, objective:"Approve and reject orders as a manager"}`) and wrap each goal in `scout_journey`: the live view shows that objective and the task it is on beside the session's feed, which is how the person watching knows what every agent is for. Keep that goal TRUE: one journey per goal, one goal per thing you are checking ("Save a settings change as the auditor", not "Check every page"), ended the moment it is decided and the next one started before you move on. A journey that outlives its goal shows the viewer an objective the session left behind minutes ago. Run roughly as many agents at once as the machine has cores, less two, since each drives a real browser; beyond that they only queue. Exploring one area is well within a mid-tier model, so run these agents on one (Sonnet or its equivalent in your client) unless the user names a model; keep the larger model for the agent that plans the split and writes the report. No agent may call `scout_close {all: true}` while others run — only the last step, once every agent has finished.
|
|
74
|
+
- **Lanes report in typed decisions, not prose.** Each lane files its findings with `scout_finding` as it goes, so the report already has them; what the lane hands BACK to the planner is one JSON object, the lane report, and nothing else. Get the paragraph to put in a lane's prompt from `scout_lane_report {lane}`: it names the shape (`status`, one decision per observation with `verdict`, `severity`, `category`, a calibrated `confidence` and a machine-signature `evidence`, the `routes` covered, `blocked_by`), every value's closed set (the categories are `scout_finding`'s) and every length limit, because a limit a lane is not told refuses good replies. When the lane hands back, pass its text to `scout_lane_report {lane, reply}`: it returns the one-line fold (defects, highs, unsure, mean confidence, routes, what blocked it) or the reason the reply was refused. Relay a refusal to the lane once and ask for the corrected object; never re-judge its prose yourself, and tell every lane that the object IS its final report, since a lane that writes the object and then hands back a summary of it has failed. The planner then folds lanes by counting, not by reading: a lane's reply is a few hundred tokens whatever it found, and a value the schema refuses is caught at the boundary instead of becoming a severity like "Low-Medium" in the report. Where the client lets you set a lane's reasoning effort, use **medium**: in this project's benchmark it judged as well as high in three-quarters of the time and under half of xhigh, low inflated severities, and max took three minutes per lane for no better agreement.
|
|
74
75
|
|
|
75
76
|
## The impatient-user pass (extensive)
|
|
76
77
|
|
|
@@ -91,7 +92,7 @@ The engine is self-healing (orphaned browsers reaped, wedged calls time out with
|
|
|
91
92
|
- Tabs and views (`?tab=…`) are distinct states in coverage — prefer clicking the actual tab controls when testing a flow; deep-link only to re-verify.
|
|
92
93
|
- **Lifecycle:** if the user says something is fixed, or the evidence no longer reproduces on a route you re-tested, mark it with `scout_resolve {id}` (spelled `findingId` if you prefer the long form). Historical findings in the report are unverified — don't re-report them as new.
|
|
93
94
|
- **A `SHARED CHROME` block in a design audit is ONE finding, not one per page.** Those elements (sidebar, header, breadcrumb bar) are the app shell; the audit already excludes them from the page's score and reports them once. File a single finding for the shell and move on — filing per-page produced five separate tickets for two CSS declarations in a real run.
|
|
94
|
-
- **`scout_finding` takes `severity`, `category`, `title`, `detail` and `evidence`.** `category` is the finding's kind — `
|
|
95
|
+
- **`scout_finding` takes `severity`, `category`, `title`, `detail` and `evidence`.** `category` is the finding's kind — `console-error`, `page-error`, `http-error`, `network`, `dead-end`, `ux-confusing`, `ux-polish`, `visual`, `a11y`, `permission-leak`, `data-inconsistency`, `stale-state`, `data-loss`, `performance`, `security`, `missing-testid`, `other` — and is what groups the report.
|
|
95
96
|
- **Always pass `evidence`** to `scout_finding` — a canonical machine signature like `GET /api/reports/dashboard 403` or `widget dashboard-summary-widget shows 0`. It's what deduplicates the same bug across runs when titles get rephrased.
|
|
96
97
|
- File judgment findings too: confusing flows, no-feedback actions, state lost on refresh, permission leaks (low-privilege role reaching admin surface), `missing-testid` (low), unnamed interactables (a11y, low).
|
|
97
98
|
- Respect refusals — never retry or route around a policy refusal; note it and move on.
|