scenescout 1.1.0 → 1.3.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +29 -0
- package/README.md +126 -27
- package/dist/browsers.js +196 -0
- package/dist/cli.js +308 -70
- package/dist/clients.js +213 -0
- package/dist/engine/browser.js +148 -20
- package/dist/engine/hover.js +42 -0
- package/dist/engine/launch.js +12 -5
- package/dist/engine/live-page.js +644 -0
- package/dist/engine/live.js +549 -0
- package/dist/engine/memory.js +3 -0
- package/dist/engine/probes.js +4 -1
- package/dist/engine/report.js +8 -2
- package/dist/installer.js +106 -12
- package/dist/mcp-server.js +268 -28
- package/dist/playbook.js +83 -0
- package/package.json +17 -5
- package/skills/scenescout/SKILL.md +5 -4
package/package.json
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "scenescout",
|
|
3
|
-
"version": "1.
|
|
4
|
-
"description": "SceneScout —
|
|
3
|
+
"version": "1.3.0",
|
|
4
|
+
"description": "SceneScout — exploratory UI testing for AI coding agents. An MCP server that gives any agent (Claude Code, Cursor, VS Code Copilot, Codex, Gemini CLI and others) a structured view of a running web app, always-on oracles, a network-level write policy, memory across runs and a gap-checked report.",
|
|
5
5
|
"license": "MIT",
|
|
6
6
|
"author": "brunoboto96",
|
|
7
7
|
"repository": {
|
|
@@ -14,13 +14,24 @@
|
|
|
14
14
|
},
|
|
15
15
|
"keywords": [
|
|
16
16
|
"mcp",
|
|
17
|
+
"mcp-server",
|
|
17
18
|
"model-context-protocol",
|
|
18
19
|
"exploratory-testing",
|
|
19
20
|
"ui-testing",
|
|
21
|
+
"e2e-testing",
|
|
22
|
+
"qa",
|
|
20
23
|
"playwright",
|
|
24
|
+
"browser-automation",
|
|
25
|
+
"ai-agent",
|
|
26
|
+
"ai-testing",
|
|
27
|
+
"accessibility",
|
|
21
28
|
"claude-code",
|
|
22
|
-
"
|
|
23
|
-
"
|
|
29
|
+
"cursor",
|
|
30
|
+
"vscode",
|
|
31
|
+
"github-copilot",
|
|
32
|
+
"codex",
|
|
33
|
+
"gemini-cli",
|
|
34
|
+
"windsurf"
|
|
24
35
|
],
|
|
25
36
|
"type": "module",
|
|
26
37
|
"bin": {
|
|
@@ -60,7 +71,7 @@
|
|
|
60
71
|
"mcp-check": "npm run build && npm run mcp-check:run",
|
|
61
72
|
"mcp-check:run": "tsx scripts/mcp-check.ts",
|
|
62
73
|
"test": "npm run build && npm run test:unit && npm run smoke:run && npm run mcp-check:run",
|
|
63
|
-
"test:unit": "npm run scan-test && npm run oracle-test && npm run policy-test && npm run fixture-test && npm run dispatch-test && npm run design-test && npm run contract-test && npm run memory-test && npm run install-test && npm run hygiene-test",
|
|
74
|
+
"test:unit": "npm run scan-test && npm run oracle-test && npm run policy-test && npm run fixture-test && npm run dispatch-test && npm run design-test && npm run contract-test && npm run memory-test && npm run install-test && npm run live-test && npm run hygiene-test",
|
|
64
75
|
"scan-test": "tsx scripts/scan-test.ts",
|
|
65
76
|
"oracle-test": "tsx --test scripts/oracle-test.ts",
|
|
66
77
|
"policy-test": "tsx --test scripts/policy-test.ts",
|
|
@@ -70,6 +81,7 @@
|
|
|
70
81
|
"contract-test": "tsx --test scripts/contract-test.ts",
|
|
71
82
|
"memory-test": "tsx --test scripts/memory-test.ts",
|
|
72
83
|
"install-test": "tsx --test scripts/install-test.ts",
|
|
84
|
+
"live-test": "tsx --test scripts/live-test.ts",
|
|
73
85
|
"hygiene-test": "tsx --test scripts/hygiene-test.ts"
|
|
74
86
|
},
|
|
75
87
|
"dependencies": {
|
|
@@ -5,16 +5,16 @@ description: AI exploratory UI testing — drive the SceneScout MCP browser tool
|
|
|
5
5
|
|
|
6
6
|
# SceneScout — exploratory UI testing agent
|
|
7
7
|
|
|
8
|
-
You are the brain of an exploratory UI tester. The SceneScout MCP server gives you deterministic browser tools (the `scout_*` tools —
|
|
8
|
+
You are the brain of an exploratory UI tester. The SceneScout MCP server gives you deterministic browser tools (the `scout_*` tools — some clients show them under a prefix: Claude Code lists `mcp__scenescout__scout_*`, or `mcp__plugin_scenescout_scenescout__scout_*` when installed as a plugin); you provide intent, judgment, and curiosity. The engine gives you structured render-state (elements, geometry, oracles) — never parse pixels when text will do. Argument hint: `[--level minimal|medium|extensive] [--url URL] [--role NAME] [--safe-write | --allow-destructive]`.
|
|
9
9
|
|
|
10
10
|
**The mission is wider than pass/fail.** Scripted e2e suites answer "does it still work?" as a binary and say nothing about what they don't cover; a human can't manually exercise a large app. You cover both gaps: find what's broken (oracles, dead ends, permission leaks) AND report how the product could be *better* — confusing flows, weak hierarchy, design-system drift, friction. Improvement feedback with concrete measurements is a first-class deliverable, not garnish; a run that finds no crashes but produces sharp `ux-polish`/`visual` suggestions is a successful run.
|
|
11
11
|
|
|
12
12
|
## Setup (in order)
|
|
13
13
|
|
|
14
|
-
1. **Check the tools exist.**
|
|
14
|
+
1. **Check the tools exist.** Reading this as the result of `scout_playbook` or of the `explore` prompt? Then they do: go to step 2. Otherwise look for a `scout_scan` tool under either prefix above. If there is none, stop and tell the user how to get it, then to start a fresh session (these are Claude Code's commands; the README has the config for other clients):
|
|
15
15
|
- as a plugin: `/plugin marketplace add brunoboto96/SceneScout` then `/plugin install scenescout@scenescout-marketplace`
|
|
16
16
|
- or by hand: `claude mcp add --scope user scenescout -- npx -y scenescout serve` (from a source checkout, register with an **absolute node path** instead — a bare `node` fails with "Executable not found in $PATH" under nvm/fnm: `claude mcp add --scope user scenescout -- "$(which node)" <checkout>/dist/mcp-server.js`)
|
|
17
|
-
If attach later reports that
|
|
17
|
+
If attach later reports that a browser build has not been downloaded, relay the one-time command it names. Attach drives Chromium unless you pass `browser: "firefox"` or `"webkit"`; do that only when the user asks for a cross-browser pass, and say in each finding which browser showed it.
|
|
18
18
|
2. **`scout_scan`** the project's absolute path. Read routes, framework, auth states, notes.
|
|
19
19
|
- **No source here?** When the target is a remote URL (staging, a deployed site) and the scan reports no frontend workspace, that is a supported mode, not an error: you are a black-box QA tester. Skip step 3's launch logic and keep the current directory as `projectPath` (memory and the report still need a home). Link harvesting builds the route list, and `scout_crawl` has nothing to crawl until it does — so snapshot the landing page and main navigation first, then `scout_crawl`, and `scout_crawl` again to pick up what those pages linked to. With no scanned auth states, `--role` is a path to a Playwright storage-state JSON. A remote target is far more likely to hold real data: confirm the user is authorized to test it if that is not evident, attach in `observe` mode unless the user says form submissions are acceptable on this target (then `read-only`), and never go past `read-only` unless they say the environment is disposable. In `observe` you may fill and submit forms freely: the engine blocks the request, so you still see client-side validation, and nothing is created. Claims that something is *absent* cannot be source-checked here: file them as behaviour-only and say so. `extensive` still needs ≥2 login states. The report cannot record where routes came from, so say it in your summary to the user and in an `scout_note`: routes were discovered from same-origin links only, and pages nothing links to are outside the contract.
|
|
20
20
|
- **Source available?** Use it beyond the scan: when you file a finding, read the component or handler behind it and name the file and the likely fix — that is the difference between "the save button does nothing" and a finding a developer can act on in one step.
|
|
@@ -68,6 +68,7 @@ Some flows need a TEAM — a document one role submits and another approves, a r
|
|
|
68
68
|
- While roles are NOT collaborating, use each one productively where its permissions matter (admin in /admin surfaces, low-privilege probing for permission leaks) — same coverage contract, different vantage points.
|
|
69
69
|
- **Infer the PERSONA behind each role, and write it down.** From what a role can see and do (its nav, its dashboard, the capability matrix in the report), state what this person is FOR: "qa = reviewer — approves orders, assigns reviewers, no admin" / "user = front-line user — reads documents, completes reviews, raises orders". Record it with `scout_note {section:'roles'}`. Then test the persona's WORLD, not just the permissions: does the operator's landing page serve an operator? Is anything they need N clicks deep? The capability matrix's divergent rows are questions, not verdicts — each is either a correct boundary or a gap ("should this role be able to do this?"); say which you believe it is and why.
|
|
70
70
|
- `scout_close {all: true}` at the end of a multi-role run; `scout_close {session}` to drop one role early.
|
|
71
|
+
- **Several agents in parallel** (subagents or a workflow, each driving its own session): each agent attaches its session when it STARTS and closes it by name when it FINISHES. Never open sessions ahead for agents that have not started, and never hand an open session from one agent to the next: an agent waiting for its turn should hold no browser. Give each session a `task` when you attach it (`scout_attach {session, task:"Approve and reject orders as a manager"}`) and wrap each goal in `scout_journey`: the live view shows the task and the active journey's goal beside the session's feed, which is how the person watching knows what every agent is for. Keep that goal TRUE: one journey per goal, one goal per thing you are checking ("Save a settings change as the auditor", not "Check every page"), ended the moment it is decided and the next one started before you move on. A journey that outlives its goal shows the viewer an objective the session left behind minutes ago. Run roughly as many agents at once as the machine has cores, less two, since each drives a real browser; beyond that they only queue. Exploring one area is well within a mid-tier model, so run these agents on one (Sonnet or its equivalent in your client) unless the user names a model; keep the larger model for the agent that plans the split and writes the report. No agent may call `scout_close {all: true}` while others run — only the last step, once every agent has finished.
|
|
71
72
|
|
|
72
73
|
## The impatient-user pass (extensive)
|
|
73
74
|
|
|
@@ -78,7 +79,7 @@ Polite, precise testing misses how real users behave. Once per module's key flow
|
|
|
78
79
|
- **Wrong-order behaviour:** press Enter mid-form before required fields are filled; go `scout_back` mid-wizard and return; submit, then immediately back-button. State should survive all three without data loss or duplicate records.
|
|
79
80
|
- Keep attribution honest: these are deliberate probes — say so in findings ("under rapid double-click…"), so a developer can reproduce exactly.
|
|
80
81
|
|
|
81
|
-
The engine is self-healing (orphaned browsers reaped, wedged calls time out with guidance instead of hanging) and observable: `.scenescout/status.json` + `scenescout status <project>` show what
|
|
82
|
+
The engine is self-healing (orphaned browsers reaped, wedged calls time out with guidance instead of hanging) and observable: `.scenescout/status.json` + `scenescout status <project>` show what every session is doing right now, and there is a live view for the person running you: one card per session with its current tool, how long it has been there, a thumbnail of its page, a feed of the actions it just took, and a stream they can switch on. It works for headless runs too. **`scout_attach` returns its address on a `Live view:` line: pass that address to the user in your next message, once, so they can watch.** `scout_session` with no arguments repeats it if they ask again, and `scenescout watch <project>` opens it from a terminal. It is for the person watching: you do not need to open it, and a session shown as stuck there is one to re-attach.
|
|
82
83
|
|
|
83
84
|
## Judgment (what the engine can't do)
|
|
84
85
|
|