design-playbook 0.22.2 → 0.24.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/codex/AGENTS.md CHANGED
@@ -1,4 +1,4 @@
1
- <!-- generated-by design-playbook v0.22.2 -->
1
+ <!-- generated-by design-playbook v0.24.0 -->
2
2
  # design-playbook for Codex
3
3
 
4
4
  ## Install (path of record)
@@ -13,7 +13,7 @@ Scan user-side `.scratch/<run>/` (not monorepo `dogfood/*` globs). **Include** o
13
13
  1. **Inclusion manifest** first: `path | status` (`included` / `skipped` + reason). Note: hash match ≠ honest transcription of `observed`.
14
14
  2. **Per-run table** — mandatory **run-path** column; other columns as needed; **gate** from real `validate_run.py` exit when the script is present (plugin install: `packages/design-playbook/scripts/validate_run.py` per run); else literal `not checked`. **Never** infer ok from "artifacts look complete".
15
15
  3. **Repeat blockers** — pure frequency table `count | runs | observed text` (verbatim first-seen text). A **repeat blocker** is the same normalized `observed` text recurring across runs (**counting, not judging**). Rows only where ledger `result != pass`. Grouping key = `observed` **casefold + whitespace-collapsed**, then **char-for-char** equality only; `count ≥ 2`. Literal differences stay separate; optional `similar:` pointer line, never merge counts. **`_none_` is normal** when nothing qualifies (do not loosen normalization to manufacture repeats).
16
- 4. **Rule candidate queue** (derived view, protocol — vNext S5) — derived from the point-back **findings** (not the ledger): group findings by normalized `issue` text (same normalization as repeat blockers), then a candidate enters the queue when **distinct runs ≥ 3 AND distinct task contexts ≥ 2 AND unexplained false positives = 0**. Task context per occurrence comes from the run's contract / spec / manifest method-semantics keys (user / task / environment / method) — **repeats with different contexts are never merged**; a corpus without readable contexts reports the context gap instead of qualifying. Show qualifying candidates (`candidate id | runs | contexts | occurrences`), plus below-threshold signals with their gap list (e.g. `distinct_runs 2 < 3`) — the queue reports distance to qualification, never silently drops it. **Report only**: candidates are never written back to the registry, the governance log, or the baseline; promotion is a user decision recorded in `<project>/rules-governance.jsonl` (append-only; agent may not write adjudication events).
16
+ 4. **Rule candidate queue** (derived view, protocol — vNext S5) — derived from the point-back **findings** (not the ledger): group findings by normalized `issue` text (same normalization as repeat blockers), then a candidate enters the queue when **distinct runs ≥ 3 AND distinct task contexts ≥ 2 AND unexplained false positives = 0**. Task context per occurrence comes from the run's contract / spec / manifest method-semantics keys (user / task / environment / method) — **repeats with different contexts are never merged**; a corpus without readable contexts reports the context gap instead of qualifying. Show qualifying candidates (`candidate id | runs | contexts | occurrences`), plus below-threshold signals with their gap list (e.g. `distinct_runs 2 < 3`) — the queue reports distance to qualification, never silently drops it. Frozen pre-v0.20 history may carry legacy severity spellings (`high (blocking)|high|med|low`); the derivation folds them mechanically onto the axis (`S3|S2|S1|S1`, the review-prototype Q1 table) for counting and reports the folded count per candidate — new findings still require the `S3|S2|S1|S0` axis (G2 rejects legacy spellings). The queue also reports `context_coverage` (occurrences with/without a task context) so a supply gap is visible, never silent. **Report only**: candidates are never written back to the registry, the governance log, or the baseline; promotion is a user decision recorded in `<project>/rules-governance.jsonl` (append-only; agent may not write adjudication events).
17
17
  5. **Point-back** cites: path + verbatim `observed:` quote; no line numbers.
18
18
  6. Rollup numbers derived **row-by-row** from the tables above — no "overall it seems".
19
19
 
@@ -0,0 +1,101 @@
1
+ """Pure action-parameter facts shared by preflight and capture runtime.
2
+
3
+ Mirrors execute_capture_plan handler contracts: empty fill/type/select values
4
+ and wait ms=0 are legal; missing required keys and wrong types are not.
5
+ No Playwright and no filesystem I/O.
6
+
7
+ Action verb dialect is owned HERE, not duplicated by the callers: ``do`` is
8
+ canonicalized via :func:`normalize_action_do` (strip + lower) so ``"Click"``,
9
+ ``" click "`` and ``"FILL"`` resolve to the same handler as ``"click"``. Both
10
+ preflight and runtime feed the raw action through this one normalizer before
11
+ checking ``KNOWN_DOS`` or parameter rules — no second dialect (FIX-03).
12
+
13
+ The ``index`` argument is 0-based everywhere (``actions[0]``); preflight and
14
+ runtime emit identical labels for the same position (FIX-03).
15
+ """
16
+ from __future__ import annotations
17
+
18
+ from typing import Any
19
+
20
+ KNOWN_DOS = frozenset(
21
+ {
22
+ "click",
23
+ "fill",
24
+ "type",
25
+ "press",
26
+ "wait_for_selector",
27
+ "wait_for_state",
28
+ "wait",
29
+ "sleep",
30
+ "select_option",
31
+ }
32
+ )
33
+
34
+
35
+ def normalize_action_do(do: object) -> str:
36
+ """Canonical action verb: strip surrounding whitespace and lowercase.
37
+
38
+ Shared by preflight and runtime so ``"Click"``, ``" click "`` and
39
+ ``"FILL"`` all hit the same handler as ``"click"``. Returns ``""`` for a
40
+ non-string or blank value so callers can detect the missing/empty-do case
41
+ through one falsy check instead of re-implementing strip/lower.
42
+ """
43
+ if not isinstance(do, str):
44
+ return ""
45
+ return do.strip().lower()
46
+
47
+
48
+ def _is_number(value: object) -> bool:
49
+ return isinstance(value, (int, float)) and not isinstance(value, bool)
50
+
51
+
52
+ def action_param_errors(action: dict[str, Any], index: int) -> list[str]:
53
+ """Human details for parameter violations. Empty list = would run.
54
+
55
+ ``index`` is 0-based (``actions[0]``) and matches the runtime's
56
+ ``_run_actions`` labeling so preflight and provider messages align for the
57
+ same position. ``do`` is normalized internally — callers pass the raw
58
+ action object.
59
+ """
60
+ do = normalize_action_do(action.get("do"))
61
+ if not do or do not in KNOWN_DOS:
62
+ return [f"actions[{index}].do must be one of the v1 actions"]
63
+ label = f"actions[{index}]"
64
+ errors: list[str] = []
65
+ if do in {"click", "fill", "type", "wait_for_selector", "select_option"}:
66
+ selector = action.get("selector")
67
+ if not isinstance(selector, str) or not selector:
68
+ errors.append(f"{label}.selector required for {do}")
69
+ if do in {"fill", "type"}:
70
+ if "value" in action:
71
+ if not isinstance(action.get("value"), str):
72
+ errors.append(f"{label}.value must be a string")
73
+ elif "text" in action and not isinstance(action.get("text"), str):
74
+ errors.append(f"{label}.text must be a string")
75
+ if do == "press":
76
+ key = action.get("key") or action.get("value")
77
+ if not isinstance(key, str) or not key:
78
+ errors.append(f"{label}.key required for press")
79
+ if do == "wait_for_state":
80
+ state = action.get("state")
81
+ if not isinstance(state, str) or not state:
82
+ errors.append(f"{label}.state required for wait_for_state")
83
+ if do in {"wait", "sleep"}:
84
+ if "ms" in action:
85
+ ms = action.get("ms")
86
+ if not _is_number(ms) or ms < 0: # type: ignore[operator]
87
+ errors.append(f"{label}.ms must be a non-negative number")
88
+ elif "timeout_ms" in action:
89
+ ms = action.get("timeout_ms")
90
+ if not _is_number(ms) or ms < 0: # type: ignore[operator]
91
+ errors.append(f"{label}.timeout_ms must be a non-negative number")
92
+ if do == "select_option":
93
+ value = action.get("value")
94
+ label_value = action.get("label")
95
+ if value is None and label_value is None:
96
+ errors.append(f"{label}.value or label required for select_option")
97
+ if value is not None and not isinstance(value, str):
98
+ errors.append(f"{label}.value must be a string")
99
+ if label_value is not None and not isinstance(label_value, str):
100
+ errors.append(f"{label}.label must be a string")
101
+ return errors
@@ -9,12 +9,21 @@ misconfig is visible to the orchestrator.
9
9
  from __future__ import annotations
10
10
 
11
11
  import json
12
+ import math
12
13
  import os
13
14
  from pathlib import Path
14
15
  from typing import Any, Protocol
15
16
 
16
17
  from design_playbook.mcp.evidence import containment
18
+ from design_playbook.mcp.evidence.action_params import (
19
+ action_param_errors,
20
+ normalize_action_do,
21
+ )
17
22
  from design_playbook.mcp.evidence.capture_contract import parse_capture_contract
23
+ from design_playbook.mcp.evidence.path_syntax import (
24
+ probe_sidecar_rel,
25
+ trimmed_relpath,
26
+ )
18
27
  from design_playbook.mcp.evidence.disclosure import (
19
28
  LAYOUT_PROBE_JS,
20
29
  VIEWPORTS,
@@ -22,6 +31,10 @@ from design_playbook.mcp.evidence.disclosure import (
22
31
  metric_payload,
23
32
  probe_layout,
24
33
  )
34
+ from design_playbook.mcp.evidence.page_defects import (
35
+ PROBE_SCHEMA,
36
+ probe_defects,
37
+ )
25
38
  from design_playbook.mcp.util import log as _log
26
39
 
27
40
  CAPTURE_TYPES = frozenset({"screenshot", "a11y tree", "interaction trace"})
@@ -36,6 +49,7 @@ ALLOWED_ARGUMENTS = frozenset(
36
49
  "overwrite",
37
50
  "viewport",
38
51
  "freeze",
52
+ "storage_state",
39
53
  }
40
54
  )
41
55
  RUN_ROOT_ENV = "DESIGN_PLAYBOOK_RUN_ROOT"
@@ -104,6 +118,8 @@ def _captured(
104
118
  observed_state: str,
105
119
  written_path: str,
106
120
  request: dict[str, Any],
121
+ *,
122
+ probe_artifact: str = "",
107
123
  ) -> dict[str, Any]:
108
124
  """Successful capture payload.
109
125
 
@@ -111,8 +127,10 @@ def _captured(
111
127
  (resolved under DESIGN_PLAYBOOK_RUN_ROOT or process cwd). Relative
112
128
  ``artifact`` stays the run-root-relative path for manifest binding.
113
129
  ``request`` echoes the normalized capture contract for manifest embedding.
130
+ ``probe_artifact`` is the sibling page-probe JSON when a screenshot
131
+ capture produced one (empty otherwise). Facts, not a judgment.
114
132
  """
115
- return {
133
+ payload = {
116
134
  "artifact": artifact,
117
135
  "observed_state": observed_state,
118
136
  "result": "captured",
@@ -120,6 +138,115 @@ def _captured(
120
138
  "written_path": written_path,
121
139
  "request": request,
122
140
  }
141
+ if probe_artifact:
142
+ payload["probe_artifact"] = probe_artifact
143
+ return payload
144
+
145
+
146
+ _MEASUREMENT_STATUSES = frozenset({"measured", "blocked", "unmeasured"})
147
+
148
+
149
+ def _measurement_meta(
150
+ status: object, error: object, *, missing: str
151
+ ) -> dict[str, str]:
152
+ """Coerce one measurement face. measured → empty error; other states keep a reason."""
153
+ text_status = str(status) if status else "unmeasured"
154
+ if text_status not in _MEASUREMENT_STATUSES:
155
+ text_status = "unmeasured"
156
+ text_error = str(error or "")
157
+ if text_status == "measured":
158
+ return {"measurement_status": "measured", "measurement_error": ""}
159
+ if not text_error:
160
+ text_error = (
161
+ missing if text_status == "unmeasured" else "probe measurement blocked"
162
+ )
163
+ return {
164
+ "measurement_status": text_status,
165
+ "measurement_error": text_error,
166
+ }
167
+
168
+
169
+ def _console_face(probed: dict[str, Any]) -> tuple[list[str], dict[str, str]]:
170
+ """Map capture_and_probe console_errors to rows + measurement meta.
171
+
172
+ Missing key → unmeasured. None → unmeasured. Non-list → blocked.
173
+ A list (including empty) → measured. Never coerce None/bad shape to a
174
+ clean zero-hit.
175
+ """
176
+ missing = "console probe not returned"
177
+ if "console_errors" not in probed:
178
+ return [], _measurement_meta(None, None, missing=missing)
179
+ raw = probed.get("console_errors")
180
+ if raw is None:
181
+ return [], _measurement_meta(
182
+ "unmeasured", "console probe returned no list", missing=missing
183
+ )
184
+ if not isinstance(raw, (list, tuple)):
185
+ return [], _measurement_meta(
186
+ "blocked",
187
+ "console probe output is not a list",
188
+ missing=missing,
189
+ )
190
+ rows = [str(item) for item in raw if item]
191
+ return rows, _measurement_meta("measured", "", missing=missing)
192
+
193
+
194
+ def _write_probe_sidecar(probe_rel: str, probed: dict[str, Any]) -> str:
195
+ """Write page-probe/v1 JSON next to a screenshot.
196
+
197
+ Path containment failures raise ValueError so a probing capture cannot
198
+ report success without a sidecar.
199
+ """
200
+ out_path = _resolve_artifact_path(probe_rel)
201
+ metrics = probed.get("metrics")
202
+ defects = probed.get("defects")
203
+ layout: dict[str, Any] = {
204
+ "sw": 0,
205
+ "innerH": 0,
206
+ "hOverflow": 0,
207
+ "inFold": False,
208
+ **_measurement_meta(None, None, missing="layout probe not returned"),
209
+ }
210
+ if metrics is not None:
211
+ layout = {
212
+ "sw": getattr(metrics, "sw", 0),
213
+ "innerH": getattr(metrics, "innerH", 0),
214
+ "hOverflow": getattr(metrics, "hOverflow", 0),
215
+ "inFold": getattr(metrics, "inFold", False),
216
+ **_measurement_meta(
217
+ getattr(metrics, "measurement_status", "unmeasured"),
218
+ getattr(metrics, "measurement_error", ""),
219
+ missing="layout probe not returned",
220
+ ),
221
+ }
222
+ leak_rows = list(getattr(defects, "leaks", ()) or ()) if defects is not None else []
223
+ tap_rows = list(getattr(defects, "tap_fails", ()) or ()) if defects is not None else []
224
+ if defects is None:
225
+ defects_meta = _measurement_meta(
226
+ None, None, missing="defect probe not returned"
227
+ )
228
+ else:
229
+ defects_meta = _measurement_meta(
230
+ getattr(defects, "measurement_status", "unmeasured"),
231
+ getattr(defects, "measurement_error", ""),
232
+ missing="defect probe not returned",
233
+ )
234
+ console_rows, console_meta = _console_face(probed)
235
+ payload = {
236
+ "schema": PROBE_SCHEMA,
237
+ "layout": layout,
238
+ "leaks": leak_rows,
239
+ "tapFails": tap_rows,
240
+ "consoleErrors": console_rows,
241
+ "defects": defects_meta,
242
+ "console": console_meta,
243
+ }
244
+ out_path.parent.mkdir(parents=True, exist_ok=True)
245
+ out_path.write_text(
246
+ json.dumps(payload, ensure_ascii=False, indent=2),
247
+ encoding="utf-8",
248
+ )
249
+ return probe_rel
123
250
 
124
251
 
125
252
  def _apply_freeze(page: Any, freeze: dict[str, Any]) -> None:
@@ -188,10 +315,75 @@ def _resolve_artifact_path(artifact_path: str) -> Path:
188
315
  result = containment.write_target(artifact_path, _run_root())
189
316
  if result.ok:
190
317
  assert result.path is not None # ok implies path is set
318
+ _refuse_reserved_write(result.path)
191
319
  return result.path
192
320
  raise ValueError(_reason_message(result.reason))
193
321
 
194
322
 
323
+ def _refuse_reserved_write(path: Path) -> None:
324
+ """Provider never writes the manifest SSOT, including via sidecar aliases."""
325
+ if path.name.casefold() == "manifest.jsonl":
326
+ raise ValueError("provider never writes manifest.jsonl")
327
+
328
+
329
+ _JS_MAX_SAFE_INTEGER = 2**53
330
+
331
+
332
+ def _load_storage_state_object(text: str) -> dict[str, Any]:
333
+ """Parse Playwright storage_state JSON without nonstandard constants."""
334
+
335
+ def reject_constant(name: str) -> object:
336
+ del name
337
+ raise ValueError("storage_state JSON contains a nonstandard constant")
338
+
339
+ def parse_int(raw: str) -> int:
340
+ value = int(raw)
341
+ if abs(value) > _JS_MAX_SAFE_INTEGER:
342
+ raise ValueError("storage_state JSON contains an overflowing integer")
343
+ return value
344
+
345
+ def parse_float(raw: str) -> float:
346
+ value = float(raw)
347
+ if not math.isfinite(value):
348
+ raise ValueError("storage_state JSON contains a non-finite number")
349
+ return value
350
+
351
+ try:
352
+ session = json.loads(
353
+ text,
354
+ parse_constant=reject_constant,
355
+ parse_int=parse_int,
356
+ parse_float=parse_float,
357
+ )
358
+ except json.JSONDecodeError as exc:
359
+ raise ValueError("storage_state is not readable JSON") from exc
360
+ if not isinstance(session, dict):
361
+ raise ValueError("storage_state must be a JSON object")
362
+ return session
363
+
364
+
365
+ def _safe_capture_failure(
366
+ exc: BaseException, *, operation: str, session: dict[str, Any] | None
367
+ ) -> str:
368
+ """Log/return one diagnostic. Session context never echoes exception text.
369
+
370
+ When a session was loaded we suppress the raw exception string (it may
371
+ contain ``Call log`` or other operator-sensitive detail). The diagnostic
372
+ keeps the failure type and the operation, plus a neutral recovery cue that
373
+ does NOT assume the failure is a session problem — a navigation
374
+ ``TimeoutError`` after a valid session load is not an auth failure
375
+ (A4-007), so we point at the run log and a supported path rather than
376
+ telling the operator to refresh credentials.
377
+ """
378
+ if session is None:
379
+ return str(exc)
380
+ kind = type(exc).__name__
381
+ return (
382
+ f"{operation} failed ({kind}); operator: check the run log or retry "
383
+ "with a supported path, then recapture"
384
+ )
385
+
386
+
195
387
  def _reason_message(reason: str) -> str:
196
388
  """Provider message for a containment reason code (ADR-0026).
197
389
 
@@ -321,13 +513,18 @@ def _run_actions(page: Any, actions: list[dict[str, Any]]) -> None:
321
513
  for i, action in enumerate(actions):
322
514
  if not isinstance(action, dict):
323
515
  raise ValueError(f"actions[{i}] must be an object")
324
- do = action.get("do")
325
- if not isinstance(do, str) or not do.strip():
516
+ raw_do = action.get("do")
517
+ do = normalize_action_do(raw_do)
518
+ if not do:
326
519
  raise ValueError(f"actions[{i}].do is required")
327
- do = do.strip().lower()
328
520
  handler = _ACTION_HANDLERS.get(do)
329
521
  if handler is None:
330
522
  raise ValueError(f"actions[{i}]: unsupported do={do!r}")
523
+ checked = dict(action)
524
+ checked["do"] = do
525
+ param_errors = action_param_errors(checked, i)
526
+ if param_errors:
527
+ raise ValueError(param_errors[0])
331
528
  handler(page, action, i, do)
332
529
 
333
530
 
@@ -417,26 +614,41 @@ class PlaywrightBrowserAdapter:
417
614
  viewport: dict[str, Any],
418
615
  freeze: dict[str, Any],
419
616
  probe: bool,
420
- ) -> tuple[str, ViewportMetrics | None]:
617
+ storage_state: str | None = None,
618
+ ) -> dict[str, Any]:
421
619
  """Single Playwright capture path shared by the two adapter seams.
422
620
 
423
621
  ``probe=False`` stops after the observed state (the plain
424
622
  :class:`BrowserAdapter` contract); ``probe=True`` additionally
425
- evaluates the layout probe on the same page before teardown so the
426
- screenshot and its metrics share one browser pass.
623
+ evaluates layout and defect probes on the same page before teardown
624
+ so the screenshot and its sidecar share one browser pass.
427
625
  """
428
626
  with self._sync_playwright() as playwright:
429
627
  browser = playwright.chromium.launch(headless=True)
430
628
  try:
431
- context = browser.new_context(
432
- viewport={
629
+ context_kwargs: dict[str, Any] = {
630
+ "viewport": {
433
631
  "width": viewport["width"],
434
632
  "height": viewport["height"],
435
633
  },
436
- device_scale_factor=viewport["devicePixelRatio"],
437
- color_scheme=viewport["colorScheme"],
438
- )
634
+ "device_scale_factor": viewport["devicePixelRatio"],
635
+ "color_scheme": viewport["colorScheme"],
636
+ }
637
+ if storage_state:
638
+ context_kwargs["storage_state"] = storage_state
639
+ context = browser.new_context(**context_kwargs)
439
640
  page = context.new_page()
641
+ console_errors: list[str] = []
642
+ page.on(
643
+ "console",
644
+ lambda msg: console_errors.append(msg.text)
645
+ if msg.type == "error" and msg.text
646
+ else None,
647
+ )
648
+ page.on(
649
+ "pageerror",
650
+ lambda exc: console_errors.append(str(exc)),
651
+ )
440
652
  if viewport.get("media"):
441
653
  page.emulate_media(media=viewport["media"])
442
654
  wait_until = (
@@ -458,10 +670,16 @@ class PlaywrightBrowserAdapter:
458
670
 
459
671
  observed = _read_observed_state(page)
460
672
  if not probe:
461
- return observed, None
673
+ return {"observed_state": observed}
462
674
  raw = page.evaluate(LAYOUT_PROBE_JS)
463
675
  metrics = probe_layout(lambda _js: raw)
464
- return observed, metrics
676
+ defects = probe_defects(page.evaluate)
677
+ return {
678
+ "observed_state": observed,
679
+ "metrics": metrics,
680
+ "defects": defects,
681
+ "console_errors": console_errors,
682
+ }
465
683
  finally:
466
684
  browser.close()
467
685
 
@@ -474,8 +692,9 @@ class PlaywrightBrowserAdapter:
474
692
  out_path: Path,
475
693
  viewport: dict[str, Any],
476
694
  freeze: dict[str, Any],
695
+ storage_state: str | None = None,
477
696
  ) -> str:
478
- observed, _metrics = self._capture_page(
697
+ result = self._capture_page(
479
698
  url=url,
480
699
  capture_type=capture_type,
481
700
  actions=actions,
@@ -483,8 +702,9 @@ class PlaywrightBrowserAdapter:
483
702
  viewport=viewport,
484
703
  freeze=freeze,
485
704
  probe=False,
705
+ storage_state=storage_state,
486
706
  )
487
- return observed
707
+ return str(result["observed_state"])
488
708
 
489
709
  def capture_and_probe(
490
710
  self,
@@ -495,9 +715,10 @@ class PlaywrightBrowserAdapter:
495
715
  out_path: Path,
496
716
  viewport: dict[str, Any],
497
717
  freeze: dict[str, Any],
718
+ storage_state: str | None = None,
498
719
  ) -> dict[str, Any]:
499
720
  """Capture and probe the same page before closing its browser."""
500
- observed, metrics = self._capture_page(
721
+ return self._capture_page(
501
722
  url=url,
502
723
  capture_type=capture_type,
503
724
  actions=actions,
@@ -505,8 +726,8 @@ class PlaywrightBrowserAdapter:
505
726
  viewport=viewport,
506
727
  freeze=freeze,
507
728
  probe=True,
729
+ storage_state=storage_state,
508
730
  )
509
- return {"observed_state": observed, "metrics": metrics}
510
731
 
511
732
 
512
733
  def _validate_runtime_object(
@@ -564,26 +785,64 @@ def execute_capture_plan(
564
785
  url, cap_type, state, actions = _validate_runtime_object(args)
565
786
  artifact_path = args.get("artifact_path")
566
787
  overwrite = args.get("overwrite", False)
788
+ storage_state_rel = args.get("storage_state")
567
789
 
568
790
  if not isinstance(artifact_path, str) or not artifact_path.strip():
569
791
  raise ValueError("artifact_path is required")
570
792
  if not isinstance(overwrite, bool):
571
793
  raise ValueError("overwrite must be a boolean")
794
+ if storage_state_rel is None:
795
+ storage_state_rel = ""
796
+ if storage_state_rel != "" and not isinstance(storage_state_rel, str):
797
+ raise ValueError("storage_state must be a string path when provided")
572
798
 
573
799
  rel = artifact_path.strip()
574
800
  try:
575
801
  out_path = _resolve_artifact_path(rel)
576
802
  except ValueError as exc:
577
803
  return _failed(rel, str(exc), request=request)
804
+ storage_state_path = ""
805
+ session_obj: dict[str, Any] | None = None
806
+ if storage_state_rel:
807
+ session_rel = trimmed_relpath(storage_state_rel)
808
+ resolved = containment.read_under(_run_root(), session_rel)
809
+ if not resolved.ok or resolved.path is None:
810
+ return _failed(
811
+ rel,
812
+ f"storage_state was rejected ({resolved.reason or 'unreadable'})",
813
+ request=request,
814
+ )
815
+ try:
816
+ raw_session = resolved.path.read_text(encoding="utf-8")
817
+ session_obj = _load_storage_state_object(raw_session)
818
+ except (OSError, UnicodeError):
819
+ return _failed(
820
+ rel,
821
+ "storage_state is not readable JSON",
822
+ request=request,
823
+ )
824
+ except ValueError as exc:
825
+ return _failed(rel, str(exc), request=request)
826
+ storage_state_path = str(resolved.path)
578
827
  abs_written = str(out_path)
579
- # Refuse every case variant of the manifest execution-record SSOT.
580
- if out_path.name.casefold() == "manifest.jsonl":
581
- return _failed(
582
- rel, "provider never writes manifest.jsonl", abs_written, request=request
583
- )
584
828
  # G6 write boundary: refuse to overwrite an existing artifact unless the
585
829
  # caller explicitly opts in via overwrite=true. Checked before any
586
830
  # Playwright launch so a misconfigured re-run cannot clobber prior evidence.
831
+ will_probe = cap_type == "screenshot" and (
832
+ browser_adapter is None
833
+ or callable(getattr(browser_adapter, "capture_and_probe", None))
834
+ )
835
+ probe_rel = probe_sidecar_rel(rel) if will_probe else ""
836
+ probe_path = None
837
+ if probe_rel:
838
+ try:
839
+ probe_path = _resolve_artifact_path(probe_rel)
840
+ except ValueError as exc:
841
+ return _failed(
842
+ rel,
843
+ f"probe sidecar path rejected: {exc}",
844
+ request=request,
845
+ )
587
846
  if out_path.exists() and not overwrite:
588
847
  return _failed(
589
848
  rel,
@@ -591,6 +850,13 @@ def execute_capture_plan(
591
850
  abs_written,
592
851
  request=request,
593
852
  )
853
+ if probe_path is not None and probe_path.exists() and not overwrite:
854
+ return _failed(
855
+ rel,
856
+ f"artifact already exists: {probe_path} (pass overwrite=true to replace)",
857
+ abs_written,
858
+ request=request,
859
+ )
594
860
 
595
861
  viewport = request["viewport"]
596
862
  freeze = request["freeze"]
@@ -604,18 +870,30 @@ def execute_capture_plan(
604
870
  abs_written,
605
871
  request=request,
606
872
  )
873
+ capture_kwargs: dict[str, Any] = {
874
+ "url": url.strip(),
875
+ "capture_type": cap_type,
876
+ "actions": actions,
877
+ "out_path": out_path,
878
+ "viewport": viewport,
879
+ "freeze": freeze,
880
+ }
881
+ if storage_state_path:
882
+ capture_kwargs["storage_state"] = storage_state_path
883
+ probe_payload: dict[str, Any] | None = None
607
884
  try:
608
- observed = browser_adapter.capture(
609
- url=url.strip(),
610
- capture_type=cap_type,
611
- actions=actions,
612
- out_path=out_path,
613
- viewport=viewport,
614
- freeze=freeze,
615
- )
885
+ probe_fn = getattr(browser_adapter, "capture_and_probe", None)
886
+ if cap_type == "screenshot" and callable(probe_fn):
887
+ probe_payload = probe_fn(**capture_kwargs)
888
+ observed = str(probe_payload.get("observed_state") or "unknown")
889
+ else:
890
+ observed = browser_adapter.capture(**capture_kwargs)
616
891
  except Exception as exc: # noqa: BLE001 — surface as capture failure
617
- _log(f"capture failed: {exc}")
618
- return _failed(rel, str(exc), abs_written, request=request)
892
+ safe = _safe_capture_failure(
893
+ exc, operation="capture", session=session_obj
894
+ )
895
+ _log(safe)
896
+ return _failed(rel, safe, abs_written, request=request)
619
897
 
620
898
  if not out_path.is_file():
621
899
  return _failed(
@@ -625,7 +903,48 @@ def execute_capture_plan(
625
903
  request=request,
626
904
  )
627
905
 
628
- return _captured(rel, observed, abs_written, request)
906
+ wrote_probe = ""
907
+ if probe_payload is not None:
908
+ if not probe_rel:
909
+ return _failed(
910
+ rel,
911
+ "probe sidecar path rejected: missing sidecar path",
912
+ abs_written,
913
+ request=request,
914
+ )
915
+ # OSError (disk full / permission / path-is-a-directory) must reach the
916
+ # orchestrator via the same {result:"failed"} channel as the main
917
+ # capture — escaping to MCP isError loses the structured echo (A3-002).
918
+ # ValueError stays for containment/path-shape sidecar rejection and
919
+ # keeps its full message (debug info, not raw exception text). We do
920
+ # NOT write a clean sidecar; the failure carries written_path (the
921
+ # non-secret absolute artifact path).
922
+ try:
923
+ wrote_probe = _write_probe_sidecar(probe_rel, probe_payload)
924
+ except ValueError as exc:
925
+ return _failed(
926
+ rel,
927
+ f"probe sidecar path rejected: {exc}",
928
+ abs_written,
929
+ request=request,
930
+ )
931
+ except OSError as exc:
932
+ return _failed(
933
+ rel,
934
+ f"probe sidecar write failed: {type(exc).__name__}",
935
+ abs_written,
936
+ request=request,
937
+ )
938
+ if not wrote_probe:
939
+ return _failed(
940
+ rel,
941
+ "probe sidecar was not written",
942
+ abs_written,
943
+ request=request,
944
+ )
945
+ return _captured(
946
+ rel, observed, abs_written, request, probe_artifact=wrote_probe
947
+ )
629
948
 
630
949
 
631
950
  # --------------------------------------------------------------------------- #