scenescout 3.6.1 → 3.7.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,17 @@
1
1
  # scenescout
2
2
 
3
+ ## 3.7.0
4
+
5
+ ### Minor Changes
6
+
7
+ - 1b4a688: The live view's close-up has its own Stream button, the same one as the session's card: switch streaming off or on without leaving the close-up. Opening a close-up still streams its session, and closing it without touching the button leaves the card as it was.
8
+ - a11c96e: `scout_coverage` lists the options of each dropdown used in the run that no session has chosen, since a filter counts as exercised after one choice. Disabled and hidden options, and an empty-value placeholder or "all" option, are never listed; dropdowns with more than 20 options are pickers and are not listed either.
9
+
10
+ ### Patch Changes
11
+
12
+ - b726af1: Two findings that quote the same control no longer merge when their evidence names different requests, and a filing merged into an existing finding now names it, with its severity and title.
13
+ - 56382db: The lane-report fold no longer lists a judged defect as unfiled when its evidence appears word for word inside a filed finding's evidence.
14
+
3
15
  ## 3.6.1
4
16
 
5
17
  ### Patch Changes
@@ -56,6 +56,31 @@ const BENIGN_MUTATION_RE = /\/auth\/(refresh|token|session)|refresh[-_]?token|\/
56
56
  function xpathLookup(xpath) {
57
57
  return `document.evaluate(${JSON.stringify(xpath)}, document, null, XPathResult.FIRST_ORDERED_NODE_TYPE, null).singleNodeValue`;
58
58
  }
59
+ /**
60
+ * Runs in the page against one dropdown, BEFORE a choice: its options' values
61
+ * and labels. Read first because a dropdown may reset itself (a bulk-action or
62
+ * "jump to" menu) or remove itself on change. Skips disabled and hidden
63
+ * options, and a placeholder or "all" option with an empty value — the state
64
+ * the page loads in, which is not an option anyone owes a choice.
65
+ */
66
+ function describeSelect(node) {
67
+ // A plan may target the dropdown by its <label>; selectOption follows a label to its control, so this does too.
68
+ const target = node instanceof HTMLLabelElement ? (node.control ?? node.querySelector("select")) : node;
69
+ const select = target;
70
+ if (!select || !select.options)
71
+ return null;
72
+ return Array.from(select.options)
73
+ .filter((o) => !o.disabled && !o.hidden && o.value !== "")
74
+ .map((o) => ({ value: o.value, label: (o.label || o.textContent || "").trim().slice(0, 80) }))
75
+ .filter((o) => o.label !== "");
76
+ }
77
+ /**
78
+ * A dropdown's options, read without waiting: a select that is not there to
79
+ * read is not worth stalling the action for.
80
+ */
81
+ async function readSelectOptions(loc) {
82
+ return loc.evaluate(describeSelect, undefined, { timeout: 1000 }).catch(() => null);
83
+ }
59
84
  /** Runs in the page against one file input (or the one a chooser belongs to). */
60
85
  function describeFileInput(node) {
61
86
  const input = node;
@@ -1783,6 +1808,15 @@ export class BrowserEngine {
1783
1808
  ? `\nRevealed on hover:\n${notes.map((t) => ` · ${t}`).join("\n")}${caveat}`
1784
1809
  : `\n(no tooltip, overlay, or new page text appeared within ${HOVER_REVEAL_WINDOW_MS / 1000}s — this element reveals nothing on hover${churning ? "; page content was changing on its own, so the text-diff fallback was suppressed" : ""}${this.headed ? ". NOTE: in headed mode the PHYSICAL mouse cursor competes with the synthetic pointer — if it is resting over the browser window, hover warm-ups are cancelled; ask the user to move it off the window and retry" : ""})`));
1785
1810
  }
1811
+ /** Record a dropdown's options and the ones picked, by the values selectOption reported. */
1812
+ recordSelectChoice(fingerprint, key, options, picked) {
1813
+ const labels = options.map((o) => o.label);
1814
+ const chosen = picked.map((v) => options.find((o) => o.value === v)?.label).filter((l) => !!l);
1815
+ if (chosen.length === 0)
1816
+ this.memory.recordSelectChoice(fingerprint, key, labels, "");
1817
+ for (const label of chosen)
1818
+ this.memory.recordSelectChoice(fingerprint, key, labels, label);
1819
+ }
1786
1820
  async select(ref, value) {
1787
1821
  const page = this.requirePage();
1788
1822
  const { el, liveLabel } = await this.resolveForAction(ref);
@@ -1802,8 +1836,12 @@ export class BrowserEngine {
1802
1836
  return destructiveRefusal(optionLabel || value, this.mode);
1803
1837
  }
1804
1838
  }
1805
- await page.locator(`xpath=${el.xpath}`).selectOption(value, { timeout: ACTION_TIMEOUT_MS });
1839
+ const loc = page.locator(`xpath=${el.xpath}`);
1840
+ const options = el.tag === "select" ? await readSelectOptions(loc) : null;
1841
+ const picked = await loc.selectOption(value, { timeout: ACTION_TIMEOUT_MS });
1806
1842
  this.memory.markExercised(this.currentFingerprint, el.key, "select");
1843
+ if (options)
1844
+ this.recordSelectChoice(this.currentFingerprint, el.key, options, picked);
1807
1845
  return this.afterAction("select", `${el.role} "${el.name}" = ${value}`);
1808
1846
  }
1809
1847
  /**
@@ -2131,6 +2169,8 @@ export class BrowserEngine {
2131
2169
  // What a type step has to say about the field it typed into; it goes on
2132
2170
  // the step's own line, so it cannot read as the previous step's.
2133
2171
  let note = "";
2172
+ // A select step's options and choice, recorded against the dropdown the bookkeeping below finds.
2173
+ let chose = null;
2134
2174
  let preState = null;
2135
2175
  try {
2136
2176
  if (step.action === "navigate") {
@@ -2215,8 +2255,12 @@ export class BrowserEngine {
2215
2255
  await loc.press("Enter", { timeout: ACTION_TIMEOUT_MS });
2216
2256
  }
2217
2257
  }
2218
- else if (step.action === "select")
2219
- await loc.selectOption(step.value ?? "", { timeout: ACTION_TIMEOUT_MS });
2258
+ else if (step.action === "select") {
2259
+ const options = await readSelectOptions(loc);
2260
+ const picked = await loc.selectOption(step.value ?? "", { timeout: ACTION_TIMEOUT_MS });
2261
+ if (options)
2262
+ chose = { options, picked };
2263
+ }
2220
2264
  else if (step.action === "upload") {
2221
2265
  const r = await this.performUpload(loc, planUploadOptions(step.value));
2222
2266
  if (r.refused) {
@@ -2263,12 +2307,13 @@ export class BrowserEngine {
2263
2307
  // state so a target the pre-capture missed (or a failed capture)
2264
2308
  // still records something rather than nothing.
2265
2309
  const preHit = preState ? findIn(preState.elements) : undefined;
2266
- if (preHit)
2267
- this.memory.markExercised(preState.fp, preHit.key, `plan:${step.action}`);
2268
- else {
2269
- const postHit = findIn(elements);
2270
- if (postHit)
2271
- this.memory.markExercised(fp, postHit.key, `plan:${step.action}`);
2310
+ const postHit = preHit ? undefined : findIn(elements);
2311
+ const hit = preHit ? { fp: preState.fp, el: preHit } : postHit ? { fp, el: postHit } : undefined;
2312
+ if (hit) {
2313
+ this.memory.markExercised(hit.fp, hit.el.key, `plan:${step.action}`);
2314
+ // The lookup can fall back to a name match; options belong only to a dropdown.
2315
+ if (chose && hit.el.tag === "select")
2316
+ this.recordSelectChoice(hit.fp, hit.el.key, chose.options, chose.picked);
2272
2317
  }
2273
2318
  }
2274
2319
  catch {
@@ -258,13 +258,46 @@ export function unfiledDefects(decisions, findings) {
258
258
  const ids = identifiers(d.evidence);
259
259
  const w = words(d.evidence);
260
260
  const text = squash(d.evidence);
261
- if (filed.some((f) => (text !== "" && f.evidence === text) || covers(ids, w, f)))
261
+ if (filed.some((f) => (text !== "" && f.evidence === text) || restates(text, f) || covers(ids, w, f)))
262
262
  continue;
263
263
  }
264
264
  out.push(d.evidence ? `${d.observation} — ${d.evidence}` : d.observation);
265
265
  }
266
266
  return out;
267
267
  }
268
+ /**
269
+ * Shortest reported evidence that counts as restating a filing it is part of.
270
+ */
271
+ const MIN_RESTATED = 24;
272
+ /** A test id, captured without its attribute; and a kebab-case id of three or more parts. Lowercase text only. */
273
+ const TESTID_RE = /testid=["']?([a-z0-9_-]+)/g;
274
+ const KEBAB_ID_RE = /\b[a-z][a-z0-9]*(?:-[a-z0-9]+){2,}\b/g;
275
+ /**
276
+ * Whether the reported evidence appears word for word inside a finding's
277
+ * evidence: a lane that filed "testid=row-1..6 non-interactive; GET /api/x 200
278
+ * with only three fields" and reported the first half. It must say something
279
+ * beyond naming a control — two words once test ids and kebab-case ids of
280
+ * three or more parts are taken out — or a bare "testid=inventory-sort-qty" would be found inside
281
+ * every finding about that control, which is the one-shared-id match covers()
282
+ * refuses. One direction only: a short FILED evidence found inside a longer
283
+ * report proves nothing about the rest of the report.
284
+ */
285
+ function restates(reported, f) {
286
+ if (reported.length < MIN_RESTATED || !f.evidence.includes(reported))
287
+ return false;
288
+ const rest = reported.replace(TESTID_RE, " ").replace(KEBAB_ID_RE, " ");
289
+ return words(rest).size >= 2;
290
+ }
291
+ /**
292
+ * Whether a finding covers a decision. Lanes reword evidence between filing it
293
+ * and reporting it — an arrow for a hyphen, quoted JSON for bare, "8 links"
294
+ * for "8 link(s)" — but keep the identifiers: test ids and contrast ratios.
295
+ * Two shared identifiers, or a near-identical wording, is the same
296
+ * observation. One shared test id is never enough, with or without some
297
+ * words in common: two different defects on one control share both (sort by
298
+ * quantity sorting as text, and the same sort showing no active state), and a
299
+ * real miss must not hide behind its neighbour.
300
+ */
268
301
  function covers(ids, w, f) {
269
302
  let shared = 0;
270
303
  for (const id of ids)
@@ -284,9 +317,9 @@ function covers(ids, w, f) {
284
317
  function identifiers(text) {
285
318
  const out = new Set();
286
319
  const t = text.toLowerCase();
287
- for (const m of t.matchAll(/testid=["']?([a-z0-9_-]+)/g))
320
+ for (const m of t.matchAll(TESTID_RE))
288
321
  out.add(m[1]);
289
- for (const m of t.matchAll(/\b[a-z][a-z0-9]*(?:-[a-z0-9]+){2,}\b/g))
322
+ for (const m of t.matchAll(KEBAB_ID_RE))
290
323
  out.add(m[0]);
291
324
  for (const m of t.matchAll(/\b\d+(?:\.\d+)?:1\b/g))
292
325
  out.add(m[0]);
@@ -225,6 +225,7 @@ export const LIVE_PAGE = `<!doctype html>
225
225
  <div class="bar">
226
226
  <strong id="focus-name"></strong>
227
227
  <span class="line" id="focus-line"></span>
228
+ <button type="button" id="focus-stream" aria-pressed="false" data-testid="live-focus-stream-toggle">Stream</button>
228
229
  <button type="button" id="focus-close" data-testid="live-focus-close">Close</button>
229
230
  </div>
230
231
  <div class="stage" id="focus-stage">
@@ -253,6 +254,11 @@ export const LIVE_PAGE = `<!doctype html>
253
254
  /** The header filter, lower-cased. Hides cards; never stops a session running. */
254
255
  var filter = '';
255
256
  var focused = null;
257
+ // Whether opening the close-up is what switched its session's stream on. The
258
+ // close-up streams while it is open; closing it hands the card back as it
259
+ // was, unless the viewer used the close-up's own Stream button, whose choice
260
+ // stands.
261
+ var focusStartedStream = false;
256
262
  var skew = 0;
257
263
  var latest = {};
258
264
  var focusTick = 0;
@@ -309,7 +315,6 @@ export const LIVE_PAGE = `<!doctype html>
309
315
  function syncEvents() {
310
316
  var want = {};
311
317
  Object.keys(cards).forEach(function (name) { if (cards[name].live) want[name] = true; });
312
- if (focused && cards[focused]) want[focused] = true;
313
318
  var key = Object.keys(want).sort().join(',');
314
319
  if (key === eventsKey) return;
315
320
  eventsKey = key;
@@ -353,8 +358,20 @@ export const LIVE_PAGE = `<!doctype html>
353
358
  // Frames for a live card arrive over the shared connection; the thumbnail poll takes over again when it is switched off.
354
359
  if (on && frames[card.name]) card.img.src = frames[card.name];
355
360
  if (!on) card.img.src = shotUrl(card.name);
361
+ if (focused === card.name) paintFocusStream();
356
362
  syncEvents();
357
363
  }
364
+ // The close-up's Stream button is the card's, shown where the viewer is looking.
365
+ function paintFocusStream() {
366
+ var card = focused && cards[focused];
367
+ var button = document.getElementById('focus-stream');
368
+ button.hidden = !card;
369
+ if (!card) return;
370
+ button.setAttribute('aria-pressed', card.live ? 'true' : 'false');
371
+ button.textContent = card.live ? 'Streaming' : 'Stream';
372
+ // Off, the picture is a still that the status poll refreshes.
373
+ if (!card.live && !scrubbed) document.getElementById('focus-img').src = shotUrl(card.name);
374
+ }
358
375
  function refreshThumb(card) {
359
376
  if (card.live || document.hidden) return;
360
377
  var next = new Image();
@@ -697,7 +714,7 @@ export const LIVE_PAGE = `<!doctype html>
697
714
  scrubbed = null;
698
715
  document.getElementById('scrub-where').textContent = 'Live';
699
716
  document.getElementById('scrub-live').hidden = true;
700
- if (focused) document.getElementById('focus-img').src = frames[focused] || shotUrl(focused);
717
+ if (focused) document.getElementById('focus-img').src = (cards[focused] && cards[focused].live && frames[focused]) || shotUrl(focused);
701
718
  // Otherwise the step just left keeps its outline until the next full feed.
702
719
  renderTimeline(null);
703
720
  }
@@ -794,6 +811,10 @@ export const LIVE_PAGE = `<!doctype html>
794
811
  document.getElementById('scrub-where').textContent = 'Live';
795
812
  document.getElementById('scrub-live').hidden = true;
796
813
  document.getElementById('focus').classList.add('open');
814
+ var card = cards[name];
815
+ focusStartedStream = !!card && !card.live;
816
+ if (focusStartedStream) setLive(card, true);
817
+ paintFocusStream();
797
818
  hoverTask = null;
798
819
  renderFeed(document.getElementById('focus-feed'), (latest[name] || {}).feed, showTask);
799
820
  syncEvents();
@@ -806,6 +827,8 @@ export const LIVE_PAGE = `<!doctype html>
806
827
  focused = null;
807
828
  document.getElementById('focus').classList.remove('open');
808
829
  document.getElementById('focus-img').removeAttribute('src');
830
+ if (focusStartedStream && was && cards[was]) setLive(cards[was], false);
831
+ focusStartedStream = false;
809
832
  syncEvents();
810
833
  if (was && cards[was]) cards[was].shot.focus();
811
834
  }
@@ -865,6 +888,7 @@ export const LIVE_PAGE = `<!doctype html>
865
888
  if (!s) { line.textContent = focused ? 'This session has closed.' : ''; return; }
866
889
  var d = describe(s);
867
890
  line.textContent = d.badge + ' · ' + d.tool + ' · ' + (s.url || '');
891
+ paintFocusStream();
868
892
  if (focusTick % 3 === 0) loadFullFeed(focused);
869
893
  focusTick += 1;
870
894
  }
@@ -981,6 +1005,13 @@ export const LIVE_PAGE = `<!doctype html>
981
1005
  });
982
1006
  document.getElementById('focus-img').addEventListener('load', function () { document.getElementById('focus-stage').classList.remove('empty'); });
983
1007
  document.getElementById('focus-close').addEventListener('click', closeFocus);
1008
+ document.getElementById('focus-stream').addEventListener('click', function () {
1009
+ var card = focused && cards[focused];
1010
+ if (!card) return;
1011
+ // The viewer chose: closing the close-up no longer undoes it.
1012
+ focusStartedStream = false;
1013
+ setLive(card, !card.live);
1014
+ });
984
1015
  // A re-rendered feed replaces the group under the pointer without a mouseleave; leaving the feed itself still resets.
985
1016
  document.getElementById('focus-feed').addEventListener('mouseleave', function () { showTask(null); });
986
1017
  document.getElementById('focus').addEventListener('click', function (e) { if (e.target === this) closeFocus(); });
@@ -283,6 +283,13 @@ function decisionKey(d) {
283
283
  * them on each save, which is what made an old history slow to open.
284
284
  */
285
285
  export const MAX_LANE_DECISIONS = 1000;
286
+ /**
287
+ * Most options a dropdown may have and still be tracked for unchosen options.
288
+ * A status or sort filter has a handful, each of which can change what the
289
+ * page asks the server for; a country or time-zone picker has hundreds, and
290
+ * nobody owes the page a choice of each. Larger dropdowns are not tracked.
291
+ */
292
+ export const MAX_SELECT_OPTIONS = 20;
286
293
  const MAX_DISCOVERED_ROUTES = 300;
287
294
  /** Shared finding-similarity helpers (used by live dedup and retro-merge). */
288
295
  function findingTokens(s) {
@@ -412,6 +419,23 @@ function sameEndpointBug(existing, incoming) {
412
419
  return false;
413
420
  return sharesEndpointSignature(existing, incoming);
414
421
  }
422
+ /**
423
+ * Whether two pieces of evidence each name a request and share none — compared
424
+ * by method and normalised path, ignoring the status, so `GET /api/x` and
425
+ * `GET /api/x/ 404` are one request. Evidence that names no request (a test id,
426
+ * a toast's text) disagrees with nothing.
427
+ */
428
+ export function requestsDisagree(a, b) {
429
+ const requests = (ev) => new Set([...endpointSignatures(ev)].map((sig) => sig.split(" ").slice(0, 2).join(" ")));
430
+ const aReq = requests(a);
431
+ const bReq = requests(b);
432
+ if (aReq.size === 0 || bReq.size === 0)
433
+ return false;
434
+ for (const r of aReq)
435
+ if (bReq.has(r))
436
+ return false;
437
+ return true;
438
+ }
415
439
  /**
416
440
  * Families of finding kinds that one bug is plausibly filed under by two
417
441
  * sessions: a crash is a page-error to one and a console-error to another, a
@@ -474,7 +498,19 @@ function sameFinding(a, b) {
474
498
  // detail or its own title. Merged, the layout defect was filed and then lost
475
499
  // from the report on most runs of a benchmark. Within a family, one bug
476
500
  // filed twice under neighbouring categories still merges.
477
- if (sameFamily(a.category, b.category)) {
501
+ //
502
+ // And never when both findings' evidence names requests with no endpoint in
503
+ // common. A quoted control name bridges two findings about that control
504
+ // within one family too: "an unknown order id still offers its actions"
505
+ // (GET /api/orders/9999 404) mentions the "Request manager approval" button
506
+ // that "Request manager approval stays enabled on a pending order" (POST
507
+ // …/request-approval 409) quotes in its title, and was merged into it.
508
+ // Evidence that names a request is the finding's own statement of where it
509
+ // happened; two findings naming different requests are two bugs. A request
510
+ // that answered 2xx counts too — a false success names one — so the same bug
511
+ // described once by its page load and once by its failing call stays as two
512
+ // findings: a visible duplicate, the direction ADR 4 accepts.
513
+ if (sameFamily(a.category, b.category) && !requestsDisagree(a.evidence, b.evidence)) {
478
514
  const aTitleLits = findingLiterals(a.title);
479
515
  const bTitleLits = findingLiterals(b.title);
480
516
  if (aTitleLits.size > 0 || bTitleLits.size > 0) {
@@ -645,6 +681,44 @@ export class MemoryStore {
645
681
  this.probes = [];
646
682
  this.injectionsReported.clear();
647
683
  this.auditsThisRun = 0;
684
+ this.selectChoices.clear();
685
+ }
686
+ /**
687
+ * Each dropdown's options and the ones chosen in THIS run, by any session,
688
+ * keyed by route and element. A select counts as exercised after one choice,
689
+ * so a lane that tried four of a filter's seven options — and reported having
690
+ * tried them all — left the one that failed untried with nothing to say so.
691
+ * Per run, like the probes: whether an earlier run chose an option says
692
+ * nothing about whether this one looked. Keyed without the role: when two
693
+ * roles see different options in one dropdown, the list read last is the
694
+ * one reported.
695
+ */
696
+ selectChoices = new Map();
697
+ recordSelectChoice(fingerprint, key, options, chosen) {
698
+ const route = fingerprint.split("#")[0];
699
+ const id = `${route}\u0000${key}`;
700
+ const distinct = [...new Set(options)];
701
+ if (distinct.length > MAX_SELECT_OPTIONS) {
702
+ this.selectChoices.delete(id);
703
+ return;
704
+ }
705
+ const entry = this.selectChoices.get(id) ?? { route, key, options: [], chosen: new Set() };
706
+ // The latest list wins: options a page added or removed since are not owed.
707
+ if (distinct.length > 0)
708
+ entry.options = distinct;
709
+ if (chosen)
710
+ entry.chosen.add(chosen);
711
+ this.selectChoices.set(id, entry);
712
+ }
713
+ /** Dropdowns with options no session chose this run, in the order they were first used. */
714
+ unchosenOptions() {
715
+ const out = [];
716
+ for (const { route, key, options, chosen } of this.selectChoices.values()) {
717
+ const unchosen = options.filter((o) => !chosen.has(o));
718
+ if (unchosen.length > 0)
719
+ out.push({ route, key, unchosen });
720
+ }
721
+ return out;
648
722
  }
649
723
  constructor(projectDir) {
650
724
  this.dir = path.join(projectDir, MEMORY_DIRNAME);
@@ -318,6 +318,21 @@ export function formatRouteCoverage(allRoutes, unvisited) {
318
318
  ? ` — UNVISITED: ${unvisited.slice(0, 25).join(", ")}${unvisited.length > 25 ? " …" : ""} (scout_crawl covers these in one call)`
319
319
  : " ✓"));
320
320
  }
321
+ /**
322
+ * The dropdowns this run used without trying every option. A select counts as
323
+ * exercised after one choice, so the unexercised list above never shows these:
324
+ * a status filter whose one failing option nobody chose looks fully covered.
325
+ * Empty when there is nothing to say.
326
+ */
327
+ export function formatUnchosenOptions(dropdowns) {
328
+ if (dropdowns.length === 0)
329
+ return [];
330
+ return [
331
+ "Dropdown options never chosen this run (each can change what the page asks the server for):",
332
+ ...dropdowns.slice(0, 15).map((d) => ` ${d.route} ${d.key}: ${d.unchosen.map((o) => JSON.stringify(o)).join(", ")}`),
333
+ ...(dropdowns.length > 15 ? [` … +${dropdowns.length - 15} more`] : []),
334
+ ];
335
+ }
321
336
  /**
322
337
  * The GAP LEDGER — an explicit enumeration of what was NOT tested. This is
323
338
  * what turns "extensive" from a vibe into a verifiable claim: a run is only
@@ -41,7 +41,7 @@ import { SessionQueue, withWatchdog } from "./engine/dispatch.js";
41
41
  import { FIXTURE_KINDS } from "./engine/fixtures.js";
42
42
  import { feedForSession, LIVE_ENV, writeStatusFile, LIVE_TOKEN_FILE, liveEngines, liveTokenFileName, pidAlive, statusFileName, LiveServer, StatusBoard, } from "./engine/live.js";
43
43
  import { formatBriefs, MAX_LANES, planLanes } from "./engine/brief.js";
44
- import { computeGaps, formatRouteCoverage, generateReport, replayDocument, reportEvidence } from "./engine/report.js";
44
+ import { computeGaps, formatRouteCoverage, formatUnchosenOptions, generateReport, replayDocument, reportEvidence } from "./engine/report.js";
45
45
  import { describeVerdict, formatWorklist, unknownIds, VERDICTS, verifyWorklist } from "./engine/verify.js";
46
46
  import { RECORD_MAX_FRAMES, resolveFrame } from "./engine/replay.js";
47
47
  import { describePace, normalizePace } from "./engine/settle.js";
@@ -1147,14 +1147,14 @@ server.registerTool("scout_finding", {
1147
1147
  ? `Finding recorded: [${finding.severity}] ${finding.title} (id ${finding.id})`
1148
1148
  : finding.regressedAt
1149
1149
  ? `⟳ REOPENED as a REGRESSION: finding ${finding.id} was previously resolved but the evidence reproduces again (seen in ${finding.runs} runs). Worth calling out to the user.`
1150
- : `Duplicate of existing finding ${finding.id} (seen in ${finding.runs} runs) — already known, keep exploring.`, session);
1150
+ : `Not recorded as new: merged into existing finding ${finding.id} — [${finding.severity}] ${finding.title}${finding.evidence ? ` (evidence: ${finding.evidence.slice(0, 160)})` : " (no evidence)"}, seen in ${finding.runs} runs. If yours is a different bug, file it again with evidence naming the request that failed for you (method and path): two findings are kept apart when both name requests and none is shared.`, session);
1151
1151
  }
1152
1152
  catch (err) {
1153
1153
  return errorText(err);
1154
1154
  }
1155
1155
  }));
1156
1156
  server.registerTool("scout_coverage", {
1157
- description: "Show exploration coverage: states visited across all runs and which elements remain unexercised. Use to decide where to explore next and when the level's budget is satisfied.",
1157
+ description: "Show exploration coverage: states visited across all runs, which elements remain unexercised, and which options of a dropdown used this run no session has chosen yet. Use to decide where to explore next and when the level's budget is satisfied.",
1158
1158
  inputSchema: { session: sessionParam },
1159
1159
  }, serializedPerSession("scout_coverage", async (_args, session) => {
1160
1160
  try {
@@ -1173,6 +1173,7 @@ server.registerTool("scout_coverage", {
1173
1173
  formatRouteCoverage(eng.allKnownRoutes(), unvisited),
1174
1174
  `Unexercised elements by route:`,
1175
1175
  ...cov.unexercised.slice(0, 25).map((u) => ` ${u.state}: ${u.keys.slice(0, 6).join(", ")}${u.keys.length > 6 ? ` … +${u.keys.length - 6}` : ""}`),
1176
+ ...formatUnchosenOptions(eng.memory.unchosenOptions()),
1176
1177
  ];
1177
1178
  return text(lines.join("\n"), session);
1178
1179
  }
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "scenescout",
3
- "version": "3.6.1",
3
+ "version": "3.7.0",
4
4
  "description": "SceneScout — exploratory UI testing for AI coding agents. An MCP server that gives any agent (Claude Code, Cursor, VS Code Copilot, Codex, Gemini CLI and others) a structured view of a running web app, always-on oracles, a network-level write policy, memory across runs and a gap-checked report.",
5
5
  "license": "MIT",
6
6
  "author": "brunoboto96",
@@ -42,7 +42,7 @@ You are the brain of an exploratory UI tester. The SceneScout MCP server gives y
42
42
  8. **Design-connoisseur pass without pixels: `scout_design_audit`.** Run it once per representative page (dashboard, a form, a detail view, a data table). Its output has two tiers: **⚠ measurable defects** (WCAG contrast, tiny targets, clipped text, aspect-distorted images, horizontal overflow, keyboard tab stops with no visible focus indicator — sampled with real Tab presses) and **→ craft suggestions** (line measure and line-height rhythm, spacing-grid adherence, typography entropy, gray census and accent-hue count, pure-#000 body text, elevation/control consistency, heading structure, indistinguishable links, AI-slop tells like gradient text/glassmorphism/side-stripe borders/identical card grids), closing with a SYSTEM SUMMARY of design-system coherence. Judge every line with product context (dense tables legitimately have small targets; a chart page legitimately uses many hues). File ⚠ defects as `visual`/`a11y`, and genuine → opportunities as `ux-polish` findings **quoting the concrete numbers** — "~142 characters per line (65–75 ideal)" beats "text feels wide". Every audit ends with a **PAGE SCORE** (0–100 overall + a11y/craft/consistency/task-clarity subscores) persisted per route — the report ranks pages worst-first, so re-runs show whether pages got better or worse. Separately, every `scout_snapshot` runs an **overlay/modal probe** automatically: an empty dialog over a grayed page, a backdrop with no dialog, a far-off-centre dialog leaving a blank band, or a dialog extending unreachably below the viewport appear as OVERLAY lines in GEOMETRY issues — treat these as high-value findings (the user is visually stuck). This is where "how could this page be better" gets answered, not just "is it broken".
43
43
  9. **Measure task EASE with `scout_journey`, not just correctness.** Wrap each module's primary task (`{action:"start", goal:"Create an order"}` → do it → `{action:"end", completed:…}`). Navigate by CLICKING like a first-time user — typing a known deep URL shortcuts the very thing being measured (a route you can only reach by editing the address bar is itself a finding). The result gives interaction count, distinct screens, the path taken, and BACKTRACKS — returning to a screen already left is the clearest evidence the next step wasn't discoverable. An abandoned journey (`completed:false`) is a high-severity finding: the task is blocked or undiscoverable, which no passing e2e suite would ever reveal.
44
44
  10. **Walk the auth surface too — anonymously.** Attach a second session WITHOUT a storage-state file (a fresh logged-out profile) and exercise signup, login failure states, and forgot/reset-password **as far as they physically go**. The mailbox wall is expected — reaching "check your email" IS the success condition; everything before it is what you're testing: does submit actually fire (a dead signup button is a high finding), are errors specific and actionable, can the user resend or recover from a typo, does the flow dead-end. Use plausible synthetic identities only (invent `qa-<runid>@example.com`-style addresses, never a real person's), submit each form valid AND invalid, and judge the feedback. Two classic findings live here: a forgot-password that answers "no account with that email" is an **account-enumeration leak** (file as security; "if an account exists, we sent a link" is the correct shape), and a signup that accepts the form then lands on a blank or logged-out page with no guidance is a **journey dead-end**. Signup creates a record, so what this pass may do depends on the mode. In `observe`, fill and submit the auth forms for their CLIENT-SIDE behaviour only: the engine blocks signup, password change and reset, and lets only a login itself go out. Disclose the server-side half as a gap. Actually creating an account needs the user's explicit okay and safe-write mode; the engine tracks the created account like any other creation.
45
- 11. **`scout_coverage` decides what's next** — it lists unvisited routes and unexercised elements. Trust it over your memory. Prefer reaching routes by clicking real navigation; fall back to direct URLs for coverage completeness and re-verification, and say which you used when it affects the finding (see the provenance rule below).
45
+ 11. **`scout_coverage` decides what's next** — it lists unvisited routes, unexercised elements, and the options of each dropdown you used that no session has chosen this run (a filter counts as exercised after one choice, and the option you skipped can be the one whose request fails). Trust it over your memory. Prefer reaching routes by clicking real navigation; fall back to direct URLs for coverage completeness and re-verification, and say which you used when it affects the finding (see the provenance rule below).
46
46
 
47
47
  ## Levels (completion contracts — the engine ENFORCES them via `scout_report {level}`)
48
48