scenescout 3.2.0 → 3.4.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -26,7 +26,7 @@ You are the brain of an exploratory UI tester. The SceneScout MCP server gives y
26
26
  - `safe-write` (`--safe-write`, or the user asks to test creating/editing things): create freely — **prioritize testing CREATE flows** — then edit/delete ONLY the records you created (the engine tracks your creations and blocks mutations on anything else). Never attempt to clean up or modify pre-existing data.
27
27
  - `destructive` (`--allow-destructive` only): everything allowed. Requires the user to explicitly confirm the environment is disposable/seeded. Never decide this yourself.
28
28
  A `🛡 WRITE-POLICY blocked` notice in a tool result is the engine's safety net, NOT an app bug — never file a finding for the error UI it causes; note it and move on (or suggest the user re-run with a laxer mode if that flow matters).
29
- **`⚠ AUTH FAILED` on attach means the storage state is stale** — its token has expired and the browser is sitting on a login page. Stop and ask the user to regenerate it (usually their Playwright auth-setup project); do not explore, and do not file findings from a logged-out session. Mid-run, `⚠ SESSION AUTH LOST` means the same thing happened after N navigations: everything since is meaningless, so re-attach rather than pressing on. Routes bounced this way are recorded as NOT covered, so the gap ledger will still show them.
29
+ **`⚠ AUTH FAILED` on attach means the storage state is stale** — its token has expired and the browser is sitting on a login page. Stop and ask the user to regenerate it (usually their Playwright auth-setup project); do not explore, and do not file findings from a logged-out session. The message quotes whatever an earlier run wrote down about regenerating it, so relay that first rather than asking from scratch — and the moment the user tells you what actually worked, write it down with `scout_note {section:'setup', note}` (the exact command, what has to be on PATH, how long a token lasts). A login state expires on a timer nobody remembers, and this is the single learning a run most reliably pays for and then forgets. Mid-run, `⚠ SESSION AUTH LOST` means the same thing happened after N navigations: everything since is meaningless, so re-attach rather than pressing on. Routes bounced this way are recorded as NOT covered, so the gap ledger will still show them.
30
30
  6. **`scout_note {action:'read'}` immediately after attaching.** ASSUMPTIONS.md is the run-over-run written memory: what the app is, who each role is FOR, constraints discovered the hard way ("an order can only ship once approved"). Starting without reading it means re-learning what a previous run already paid to find out. Throughout the run, `scout_note {action:'add'}` every DURABLE learning the moment you confirm it — app model, role personas, conventions, constraints, risks, glossary. Not session facts (ids, counts); knowledge a future run should start with.
31
31
 
32
32
  ## The efficient loop (this is the core method)
@@ -58,6 +58,8 @@ You are the brain of an exploratory UI tester. The SceneScout MCP server gives y
58
58
 
59
59
  Route knowledge is generic: routes read from source (file-based routing, or React Router / Vue Router / Angular configuration) ∪ links harvested from every snapshot (including `?tab=` screens) form the contract — it works on any app, not just Next. For a responsive pass, re-attach with `viewportWidth: 390, viewportHeight: 844` and re-run the design audit on key pages.
60
60
 
61
+ **Slowing a run down so a person can follow it.** By default every session runs as fast as its page allows: the engine waits on the requests the action actually fired, not on a fixed sleep. When the user is watching and wants to keep up — taking notes alongside the agent, demonstrating a flow to someone — `scout_attach {paceMs: 5000}` makes that session pause at least five seconds between actions, and `scout_session {paceMs: 5000}` changes it mid-run (with no `name`, it applies to every attached session at once). `paceMs: 0` puts it back to full speed. Set it only when the user asks to slow down; it is a floor on every action, so a run at 5000 takes as long as it sounds.
62
+
61
63
  ## Multi-role collaboration (named sessions)
62
64
 
63
65
  Some flows need a TEAM — a document one role submits and another approves, a review one role assigns and another completes, permission checks that only mean something side-by-side. `scout_attach {session: "admin", storageStatePath: …}` then `scout_attach {session: "qa", storageStatePath: …}` keeps BOTH browsers live and authenticated. Coverage and findings from all roles merge into one project memory.
@@ -71,7 +73,8 @@ Some flows need a TEAM — a document one role submits and another approves, a r
71
73
  - **Infer the PERSONA behind each role, and write it down.** From what a role can see and do (its nav, its dashboard, the capability matrix in the report), state what this person is FOR: "qa = reviewer — approves orders, assigns reviewers, no admin" / "user = front-line user — reads documents, completes reviews, raises orders". Record it with `scout_note {section:'roles'}`. Then test the persona's WORLD, not just the permissions: does the operator's landing page serve an operator? Is anything they need N clicks deep? The capability matrix's divergent rows are questions, not verdicts — each is either a correct boundary or a gap ("should this role be able to do this?"); say which you believe it is and why.
72
74
  - `scout_close {all: true}` at the end of a multi-role run; `scout_close {session}` to drop one role early.
73
75
  - **Several agents in parallel** (subagents or a workflow, each driving its own session): each agent attaches its session when it STARTS and closes it by name when it FINISHES. Never open sessions ahead for agents that have not started, and never hand an open session from one agent to the next: an agent waiting for its turn should hold no browser. Give each session an `objective` when you attach it (`scout_attach {session, objective:"Approve and reject orders as a manager"}`) and wrap each goal in `scout_journey`: the live view shows that objective and the task it is on beside the session's feed, which is how the person watching knows what every agent is for. Keep that goal TRUE: one journey per goal, one goal per thing you are checking ("Save a settings change as the auditor", not "Check every page"), ended the moment it is decided and the next one started before you move on. A journey that outlives its goal shows the viewer an objective the session left behind minutes ago. Run roughly as many agents at once as the machine has cores, less two, since each drives a real browser; beyond that they only queue. Exploring one area is well within a mid-tier model, so run these agents on one (Sonnet or its equivalent in your client) unless the user names a model; keep the larger model for the agent that plans the split and writes the report. No agent may call `scout_close {all: true}` while others run — only the last step, once every agent has finished.
74
- - **Lanes report in typed decisions, not prose.** Each lane files its findings with `scout_finding` as it goes, so the report already has them; what the lane hands BACK to the planner is one JSON object, the lane report, and nothing else. Get the paragraph to put in a lane's prompt from `scout_lane_report {lane}`: it names the shape (`status`, one decision per observation with `verdict`, `severity`, `category`, a calibrated `confidence` and a machine-signature `evidence`, the `routes` covered, `blocked_by`), every value's closed set (the categories are `scout_finding`'s) and every length limit, because a limit a lane is not told refuses good replies. When the lane hands back, pass its text to `scout_lane_report {lane, reply}`: it returns the one-line fold (defects, highs, unsure, mean confidence, routes, what blocked it) or the reason the reply was refused. Relay a refusal to the lane once and ask for the corrected object; never re-judge its prose yourself, and tell every lane that the object IS its final report, since a lane that writes the object and then hands back a summary of it has failed. The planner then folds lanes by counting, not by reading: a lane's reply is a few hundred tokens whatever it found, and a value the schema refuses is caught at the boundary instead of becoming a severity like "Low-Medium" in the report. Where the client lets you set a lane's reasoning effort, use **medium**: in this project's benchmark it judged as well as high in three-quarters of the time and under half of xhigh, low inflated severities, and max took three minutes per lane for no better agreement.
76
+ - **Let the engine divide the app: `scout_lane_brief {lanes, goal}`.** Call it after the first crawl, when route knowledge is complete. It splits the known routes into whole modules — everything under `/orders` goes to one lane, so that lane carries state between its own steps instead of re-learning the app on every route — deals the modules out so the lanes come out within a route or two of each other, and returns each lane's session name, the `objective` to attach it with, and the routes it owns. Pass `routes` to split a subset instead. Hand each brief to its agent verbatim. Dividing by hand fails in two ways that a finished run cannot tell apart from success: two lanes audit the same register while a third module is never opened, and route coverage reads complete either way; and lanes launch without an objective, so the person watching sees browsers clicking through their app with nothing to say why.
77
+ - **Lanes report in typed decisions, not prose.** Each lane files its findings with `scout_finding` as it goes, so the report already has them; what the lane hands BACK to the planner is one JSON object, the lane report, and nothing else. Get the paragraph to put in a lane's prompt from `scout_lane_report {lane}`: it names the shape (`status`, one decision per observation with `verdict`, `severity`, `category`, a calibrated `confidence` and a machine-signature `evidence`, the `routes` covered, `blocked_by`), every value's closed set (the categories are `scout_finding`'s) and every length limit, because a limit a lane is not told refuses good replies. When the lane hands back, pass its text to `scout_lane_report {lane, reply}`: it returns the one-line fold (defects, highs, unsure, mean confidence, routes, what blocked it) or the reason the reply was refused, and keeps each decision so the confidence the lane stated can be checked against what the run went on to file — the report's calibration section is built from that, and it only appears once enough decisions exist to mean something. **So the confidence is worth stating honestly**: a lane that writes 0.95 on everything makes the section say so. Relay a refusal to the lane once and ask for the corrected object; never re-judge its prose yourself, and tell every lane that the object IS its final report, since a lane that writes the object and then hands back a summary of it has failed. The planner then folds lanes by counting, not by reading: a lane's reply is a few hundred tokens whatever it found, and a value the schema refuses is caught at the boundary instead of becoming a severity like "Low-Medium" in the report. Where the client lets you set a lane's reasoning effort, use **medium**: in this project's benchmark it judged as well as high in three-quarters of the time and under half of xhigh, low inflated severities, and max took three minutes per lane for no better agreement.
75
78
 
76
79
  ## The impatient-user pass (extensive)
77
80
 
@@ -86,11 +89,14 @@ The engine is self-healing (orphaned browsers reaped, wedged calls time out with
86
89
 
87
90
  ## Judgment (what the engine can't do)
88
91
 
92
+ - **A hidden or disabled control proves nothing: `scout_request` is how a refusal becomes evidence.** Any finding that a role CANNOT do something must show the server refusing it, not just the UI withholding the button — a control hidden client-side is a decision the browser made, and the endpoint behind it may answer 200 to anyone who asks. `scout_request {session, method, path}` calls the app's own API as that session, with the UI bypassed: it runs inside the page, so it carries the session's cookies, replays the Authorization header the app itself last sent, and meets the same write policy a click does (in safe-write a mutation on a record the session did not create is refused here too, which is the engine's safety net and never a finding). It returns the signature to quote (`GET /api/admin/users 403`), the timing, the headers that decide whether two responses are genuinely identical, and the body. Use it for every permission boundary, for comparing one route across roles, and for confirming a 402/403 the UI turned into a friendly empty state. Paths are fenced to the attached origin. Unlike a shell call, every request is recorded in the run's trail and shows up in the report.
89
93
  - **Ground absence-claims in the source before filing.** Black-box behaviour tells you what happened, never *why*. Any finding asserting a control is MISSING ("no permission check", "no validation", "no segregation of duties", "not implemented") must be checked against the codebase first — you have Read/Grep/Bash, use them. Three outcomes, all valuable: (1) genuinely absent → file it with the file:line you checked, which makes it actionable instead of speculative; (2) implemented but not reached on your path → you found a *routing/config* bug, a different and often better finding; (3) implemented elsewhere in the same codebase but not here → **the strongest kind of finding**, an inconsistency against the team's own established pattern, with the reference implementation attached as the fix. A grep costs seconds and is the difference between "I think this is wrong" and "here is the line". Behaviour-only findings (a crash, a 500, a dead end) need no such check — the evidence is self-contained.
90
94
  - Oracle violations are *evidence* — reproduce/contextualize, then `scout_finding`. **Severity anchors:** crash, data loss, security/permission breach, or a flow the user cannot complete or escape (infinite spinner, no-way-out page, broken primary journey) = **high** — a dead end is high even without a crash. Broken but recoverable = medium. Polish/friction = low.
95
+ - **A `refused_empty` or `false_success` violation is the highest-value thing the engine hands you — file it, always.** Both mean the page CONTRADICTED the server. `refused_empty`: a list request was refused (401/403/5xx) and the page showed its empty state with no error, so the user is told they have nothing when the truth is that nothing could be loaded — this is how a permission regression ships without anyone noticing, and it is a `data-inconsistency` or `security` finding at **high**, not a cosmetic one. `false_success`: a save was refused and the page said it worked, so the user walks away believing their change was kept — **high**, `data-inconsistency`. The engine only raises either when a request genuinely failed AND the page admitted nothing, so do not re-litigate whether the refusal happened; reproduce it, name the endpoint and the message the user saw, and read the handler in the source to say which branch swallowed the response. A page that shows an error is correct behaviour and never raises these — quote that contrast in the finding, since it is usually the fix.
91
96
  - **Provenance for deep-link findings:** if you reached a state by direct URL (crawl or scout_navigate) and it looks broken (no app shell, dead end), check whether real users can reach it through the UI before filing — many apps render tab content at internal URLs that users only ever reach via `?tab=` links. Always state in the finding how the state was reached.
92
97
  - Tabs and views (`?tab=…`) are distinct states in coverage — prefer clicking the actual tab controls when testing a flow; deep-link only to re-verify.
93
98
  - **Lifecycle:** if the user says something is fixed, or the evidence no longer reproduces on a route you re-tested, mark it with `scout_resolve {id}` (spelled `findingId` if you prefer the long form). Historical findings in the report are unverified — don't re-report them as new.
99
+ - **Re-testing what earlier runs left open: `scout_verify`.** Against an app this project has tested before, run it right after the crawl, and run it again whenever the user says a wave of fixes has landed. Called bare it returns the open findings in the order to re-test them — worst route first, grouped so you walk a route once instead of once per finding — each with the evidence that identifies it and the steps that produced it; `scout_verify {ids}` narrows it to specific ones, and names any that are not open rather than quietly shortening the list. Re-test a finding, then record what you saw with `scout_verify {id, verdict, note}`: `"gone"` resolves it, `"present"` stamps it confirmed so the report dates the confirmation instead of calling it unverified, `"changed"` keeps it open and says the behaviour differs — and if it is now a different bug, file that as its own finding. A campaign is far cheaper than re-deriving the same list from the report by hand, and it is the only way the historical section stops being a pile nobody trusts.
94
100
  - **A `SHARED CHROME` block in a design audit is ONE finding, not one per page.** Those elements (sidebar, header, breadcrumb bar) are the app shell; the audit already excludes them from the page's score and reports them once. File a single finding for the shell and move on — filing per-page produced five separate tickets for two CSS declarations in a real run.
95
101
  - **`scout_finding` takes `severity`, `category`, `title`, `detail` and `evidence`.** `category` is the finding's kind — `console-error`, `page-error`, `http-error`, `network`, `dead-end`, `ux-confusing`, `ux-polish`, `visual`, `a11y`, `permission-leak`, `data-inconsistency`, `stale-state`, `data-loss`, `performance`, `security`, `missing-testid`, `other` — and is what groups the report.
96
102
  - **Always pass `evidence`** to `scout_finding` — a canonical machine signature like `GET /api/reports/dashboard 403` or `widget dashboard-summary-widget shows 0`. It's what deduplicates the same bug across runs when titles get rephrased.
@@ -100,4 +106,6 @@ The engine is self-healing (orphaned browsers reaped, wedged calls time out with
100
106
 
101
107
  ## Finishing
102
108
 
109
+ `scout_report` prints the findings THIS run made in full, and lists earlier ones — and resolved ones — as an index: a row each with id, severity, age and title. That is the default because a project with hundreds of remembered findings otherwise produces a document nobody opens, in which the handful the run actually made are buried; on one real project it was 1.75 MB against 113 KB, and nearly half of it was findings already fixed. Nothing is lost: the ids are there, and `scout_report {history:'full'}` prints every one as before — use it when handing the document to someone who cannot read the project's memory. Age is what decides whether an unverified old finding is worth re-testing, so it is on every row.
110
+
103
111
  `scout_report` (satisfy the contract first) → `scout_close` → summarize in chat: worst findings first, coverage numbers, report path (`.scenescout/report.md`), and suggest promoting high findings to real Playwright regression tests (skeletons are in the report).