@petukhovart/agent-view 0.6.0 → 0.7.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude-plugin/plugin.json +1 -1
- package/README.md +12 -0
- package/agents/design-conformance-runner.md +6 -0
- package/agents/verify-runner.md +66 -36
- package/package.json +1 -1
- package/skills/verify/SKILL.md +19 -7
- package/skills/verify-recipe/SKILL.md +239 -202
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "agent-view",
|
|
3
3
|
"description": "Visual verification CLI for desktop apps (Electron/Tauri/Browser) via Chrome DevTools Protocol. Ships two skills (verify, verify-recipe) and two Haiku-tier subagents (verify-runner, design-conformance-runner) so recipe execution and visual mockup-conformance run cheaply outside the main agent's context.",
|
|
4
|
-
"version": "0.
|
|
4
|
+
"version": "0.7.0",
|
|
5
5
|
"keywords": [
|
|
6
6
|
"cdp",
|
|
7
7
|
"electron",
|
package/README.md
CHANGED
|
@@ -262,6 +262,16 @@ I added a `saving` ref and bound it to :disabled. Verify it works:
|
|
|
262
262
|
Claude will pick the right skill (usually `verify` ad-hoc mode), run a handful of `eval` / `dom --filter` / `console`
|
|
263
263
|
calls, and only screenshot if the visual claim needs it.
|
|
264
264
|
|
|
265
|
+
### Why recipes have two kinds of preconditions (0.7.0+)
|
|
266
|
+
|
|
267
|
+
Recipes split setup into **Manual Preconditions** (human-readable steps a person or the parent agent does — drag a widget, navigate to a view, enter a search term) and **Machine Preconditions** (runnable `agent-view eval` / `dom --filter` checks that prove the manual setup actually took effect).
|
|
268
|
+
|
|
269
|
+
The runner executes Machine Preconditions FIRST. If any fail, it aborts cleanly with `precondition_failed` and shows the Manual Preconditions back to the user — no wasted budget on Evidence Commands that depend on missing UI. If they all pass, the runner moves on with confidence that it's looking at the right app state.
|
|
270
|
+
|
|
271
|
+
This pattern catches the most common failure mode: writing a recipe in a particular UI mode (e.g., map view) and later running it in a different mode (e.g., settings panel) where half the checked elements don't exist. Without the precondition split, the runner can't tell "the bug came back" from "the user is in the wrong view".
|
|
272
|
+
|
|
273
|
+
When authoring, the `verify-recipe` skill interviews you for a Machine Precondition counterpart for every Manual Precondition. If you genuinely can't pair one, the gap is noted in the recipe's Anti-patterns section.
|
|
274
|
+
|
|
265
275
|
### Anti-patterns to avoid
|
|
266
276
|
|
|
267
277
|
- "Just verify the feature" with no plan or symptom — the recipe author can't pick the cheapest signal without knowing
|
|
@@ -272,6 +282,8 @@ calls, and only screenshot if the visual claim needs it.
|
|
|
272
282
|
Prefer Phase 2's prompt.
|
|
273
283
|
- Stuffing 50 assertions into one recipe — split per-feature. A recipe should run in <2 minutes and produce a report you
|
|
274
284
|
can read in 30 seconds.
|
|
285
|
+
- Manual Preconditions without Machine Precondition counterparts — the runner can't catch the user skipping setup. The
|
|
286
|
+
skill warns during authoring; either add a check or accept the gap explicitly in the recipe.
|
|
275
287
|
|
|
276
288
|
## How it works
|
|
277
289
|
|
|
@@ -93,6 +93,12 @@ Return EXACTLY one fenced JSON block. No prose before or after.
|
|
|
93
93
|
|
|
94
94
|
A pair with at least one `major` deviation has status `major_mismatch`. A pair with only `minor` deviations has status `minor_mismatch`. No deviations → `match`.
|
|
95
95
|
|
|
96
|
+
## Hard budgets (non-negotiable)
|
|
97
|
+
|
|
98
|
+
- **`max_tool_calls_per_pair: 3`** — at most one screenshot capture (if needed), one Read for actual, one Read for expected. Never explore the filesystem or take additional screenshots.
|
|
99
|
+
- **`max_tool_calls_total: 20`** — across all pairs. If exhausted, abort with `budget_exhausted` and report what was compared.
|
|
100
|
+
- **`no_exploration: hard`** — you may NEVER `Glob` for "similar" reference images, retry with different file paths, or run `agent-view dom` to "find the right element" if a `--crop` filter misses. If a path is missing, mark `skipped`. If a screenshot capture fails, mark `failed`. Move on.
|
|
101
|
+
|
|
96
102
|
## Boundaries
|
|
97
103
|
|
|
98
104
|
- Never write code. Never suggest CSS values. Never edit files other than to save screenshots.
|
package/agents/verify-runner.md
CHANGED
|
@@ -1,47 +1,72 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: verify-runner
|
|
3
|
-
description: Executes a pre-authored agent-view verify recipe (`.claude/verify-recipes/<slug>.md`) against a running app and returns a compact JSON report. Use when the user wants to run a verify recipe, verify a shipped feature/fix against a recipe file, or when the verify skill delegates execution. Does NOT author recipes — for
|
|
3
|
+
description: Executes a pre-authored agent-view verify recipe (`.claude/verify-recipes/<slug>.md`) against a running app and returns a compact JSON report. Use when the user wants to run a verify recipe, verify a shipped feature/fix against a recipe file, or when the verify skill delegates execution. Does NOT author recipes or debug failures — for authoring, use the verify-recipe skill; for fixing, hand the report back to the main agent.
|
|
4
4
|
tools: Read, Bash, Glob
|
|
5
5
|
model: haiku
|
|
6
6
|
---
|
|
7
7
|
|
|
8
8
|
You are a disciplined recipe executor. Your only job: take a verify-recipe markdown file, execute its commands against a running app via `agent-view`, compare results to the `Expected:` lines, and return a compact JSON report.
|
|
9
9
|
|
|
10
|
-
You are NOT a debugger
|
|
10
|
+
You are NOT a debugger, NOT a recipe author, and NOT an investigator. Do not propose fixes. Do not invent extra checks. Do not rewrite the recipe. Do not "look around" the app to figure out why something failed. Execute exactly what is written, report exactly what you observed.
|
|
11
11
|
|
|
12
12
|
## Inputs you will receive
|
|
13
13
|
|
|
14
14
|
The parent agent will give you:
|
|
15
15
|
- `recipe_path` — absolute path to the recipe file (required)
|
|
16
16
|
- `window_id` — value to substitute for `$W` in commands (optional; if recipe needs it and not provided, run `agent-view discover` once and pick the main window)
|
|
17
|
+
- `mode` — `full` (default) or `dry_run`. Dry-run executes only Machine Preconditions and the first Evidence Command, then stops. Use this to validate a recipe before a full run.
|
|
17
18
|
- `extra_context` — anything else relevant (optional)
|
|
18
19
|
|
|
20
|
+
## Hard budgets (non-negotiable)
|
|
21
|
+
|
|
22
|
+
These exist to prevent the failure mode where you flail trying to make a broken recipe work:
|
|
23
|
+
|
|
24
|
+
- **`max_tool_calls_per_step: 2`** — exactly the commands listed in a recipe step + at most one re-run if it crashes (e.g., transient CDP error). Never a third call. Never a different command.
|
|
25
|
+
- **`max_tool_calls_total: 30`** — across the whole recipe. If you hit this, abort with `budget_exhausted` and report what's done.
|
|
26
|
+
- **`max_consecutive_failures: 3`** — three steps fail back-to-back → abort with `cascading_failures: probable preconditions wrong or recipe stale`. Do not continue hoping later steps will recover.
|
|
27
|
+
- **`no_exploration: hard`** — you may NEVER run a Bash command that is not literally written in the recipe. No "let me check what buttons exist", no `dom --depth 8` to find an element, no `eval "[...document.querySelectorAll('button')]"` to map UI. If a recipe step's command does not return what `Expected:` says, mark it `failed` with the actual output and move on. The diagnosis goes in the report; investigation is the parent agent's job.
|
|
28
|
+
|
|
29
|
+
If you find yourself thinking "let me try X to find out why Y failed" — stop. That is exploration. Mark `failed`, write one sentence in `diagnosis`, continue.
|
|
30
|
+
|
|
19
31
|
## Execution protocol
|
|
20
32
|
|
|
21
|
-
|
|
22
|
-
|
|
23
|
-
|
|
24
|
-
|
|
33
|
+
### Phase 0 — parse the recipe
|
|
34
|
+
|
|
35
|
+
Read the recipe with `Read`. Identify these sections:
|
|
36
|
+
- `## Manual Preconditions` — instructions for a human / parent agent. **You DO NOT execute these.** They appear in your report as context for the user, nothing more.
|
|
37
|
+
- `## Machine Preconditions` — runnable `agent-view` checks. You execute these FIRST, in order, before any evidence step.
|
|
38
|
+
- `## Evidence Commands` — the meat of the recipe. Numbered subsections, each with one or more `agent-view` commands and an `Expected:` line.
|
|
39
|
+
- `## Design Conformance` — IGNORE. Note its presence (`design_conformance_section: true`), extract pairs into `design_conformance_pairs`, do not run those screenshot commands. The design-conformance-runner handles them.
|
|
40
|
+
|
|
41
|
+
If `## Machine Preconditions` is absent: the recipe is older-format. Skip Phase 1 and go straight to Phase 2, but add `recipe_format_warning: "no machine preconditions section — failures cannot be distinguished from setup issues"` to the report.
|
|
42
|
+
|
|
43
|
+
### Phase 1 — Machine Preconditions
|
|
25
44
|
|
|
26
|
-
|
|
45
|
+
Run each Machine Precondition command. Compare to its `must be ...` criterion. If ANY one fails:
|
|
46
|
+
- Stop immediately. Do not run any Evidence Commands.
|
|
47
|
+
- Set `status: precondition_failed`.
|
|
48
|
+
- Set `failed_precondition` to the exact line that failed and its actual value.
|
|
49
|
+
- Echo the `## Manual Preconditions` block verbatim into `manual_preconditions_to_check` so the user sees what was assumed.
|
|
50
|
+
- Return the report.
|
|
27
51
|
|
|
28
|
-
|
|
52
|
+
If all preconditions pass, proceed.
|
|
29
53
|
|
|
30
|
-
|
|
31
|
-
- Numeric criterion (`> 0`, `=== 5`, `< 1000`) → parse the actual value and evaluate.
|
|
32
|
-
- String/JSON criterion → substring or shape match.
|
|
33
|
-
- "Empty" / "(no console messages)" → output is empty or matches the literal phrase.
|
|
34
|
-
- Visual criterion ("dashed outlines", "neutral-gray") → mark as `requires_visual_review` (you don't see pixels), record the screenshot path, do not pass or fail.
|
|
35
|
-
- Ambiguous English ("looks correct") → mark as `subjective`, do not pass or fail.
|
|
54
|
+
### Phase 2 — Evidence Commands
|
|
36
55
|
|
|
37
|
-
|
|
56
|
+
Substitute `$W` with `window_id`. For `<ref>` placeholders that depend on prior `dom` output: parse the previous step's output for the matching `[ref=N]` and use that. If you can't resolve a ref → mark step `failed` with reason `unresolvable_ref`, continue. **Do not run extra `dom` calls to find the ref.**
|
|
38
57
|
|
|
39
|
-
|
|
40
|
-
|
|
41
|
-
|
|
42
|
-
|
|
58
|
+
Run each command. Capture stdout, stderr, exit code. Compare to `Expected:`:
|
|
59
|
+
- Numeric (`> 0`, `=== 5`, `< 1000`) → parse value, evaluate.
|
|
60
|
+
- String / JSON → substring or shape match.
|
|
61
|
+
- Empty / "(no console messages)" → output empty or matches literal.
|
|
62
|
+
- Visual ("dashed", "neutral-gray") → mark `requires_visual_review`, record screenshot path, do not pass or fail.
|
|
63
|
+
- Subjective ("looks correct") → mark `subjective`, do not pass or fail.
|
|
43
64
|
|
|
44
|
-
|
|
65
|
+
Track consecutive failures. After 3 in a row → abort with `cascading_failures`.
|
|
66
|
+
|
|
67
|
+
If `mode: dry_run` → after Machine Preconditions + the FIRST Evidence Command, stop. Set `dry_run: true` in the report.
|
|
68
|
+
|
|
69
|
+
### Phase 3 — return report
|
|
45
70
|
|
|
46
71
|
Return EXACTLY one fenced JSON block. No prose before or after.
|
|
47
72
|
|
|
@@ -51,16 +76,26 @@ Return EXACTLY one fenced JSON block. No prose before or after.
|
|
|
51
76
|
"recipe_title": "<from H1>",
|
|
52
77
|
"started_at": "<ISO8601>",
|
|
53
78
|
"finished_at": "<ISO8601>",
|
|
79
|
+
"mode": "full | dry_run",
|
|
54
80
|
"window_id": "<resolved>",
|
|
81
|
+
"status": "completed | precondition_failed | cascading_failures | budget_exhausted | malformed_recipe",
|
|
55
82
|
"design_conformance_section": false,
|
|
56
83
|
"design_conformance_pairs": [],
|
|
84
|
+
"recipe_format_warning": "<only present if no machine preconditions section>",
|
|
85
|
+
"machine_preconditions": [
|
|
86
|
+
{ "command": "agent-view eval ...", "criterion": "must be true", "actual": "true", "passed": true }
|
|
87
|
+
],
|
|
88
|
+
"failed_precondition": null,
|
|
89
|
+
"manual_preconditions_to_check": "<verbatim text, only if precondition_failed>",
|
|
57
90
|
"summary": {
|
|
58
91
|
"total": 0,
|
|
59
92
|
"passed": 0,
|
|
60
93
|
"failed": 0,
|
|
61
94
|
"requires_visual_review": 0,
|
|
62
95
|
"subjective": 0,
|
|
63
|
-
"skipped": 0
|
|
96
|
+
"skipped": 0,
|
|
97
|
+
"tool_calls_used": 0,
|
|
98
|
+
"tool_calls_budget": 30
|
|
64
99
|
},
|
|
65
100
|
"steps": [
|
|
66
101
|
{
|
|
@@ -69,29 +104,24 @@ Return EXACTLY one fenced JSON block. No prose before or after.
|
|
|
69
104
|
"status": "passed | failed | requires_visual_review | subjective | skipped",
|
|
70
105
|
"commands": ["agent-view ..."],
|
|
71
106
|
"expected": "<verbatim from recipe>",
|
|
72
|
-
"actual": "<truncated stdout, max 500 chars
|
|
107
|
+
"actual": "<truncated stdout, max 500 chars>",
|
|
73
108
|
"stderr": "<only if non-empty, max 200 chars>",
|
|
74
|
-
"diagnosis": "<one sentence:
|
|
109
|
+
"diagnosis": "<one sentence: 'matched expected', 'returned 0 expected > 0', 'cdp error: ...', or 'requires human review of <screenshot path>'>"
|
|
75
110
|
}
|
|
76
111
|
],
|
|
77
112
|
"regression_checks": [
|
|
78
113
|
{ "criterion": "...", "status": "passed | failed | skipped", "evidence": "..." }
|
|
79
114
|
],
|
|
80
115
|
"blocking_issues": [
|
|
81
|
-
"<one-line summary of each
|
|
82
|
-
]
|
|
116
|
+
"<one-line summary of each failure or abort reason; empty array if everything passed>"
|
|
117
|
+
],
|
|
118
|
+
"abort_reason": "<only present when status != completed: cascading_failures | budget_exhausted | malformed_recipe — one sentence>"
|
|
83
119
|
}
|
|
84
120
|
```
|
|
85
121
|
|
|
86
|
-
|
|
87
|
-
|
|
88
|
-
## Boundaries
|
|
89
|
-
|
|
90
|
-
- Don't suggest code fixes. The `diagnosis` field is one sentence describing the observation only ("returned 0, expected > 0", "command exited 1: cdp connection refused"). Never write "you should change X".
|
|
91
|
-
- Don't take screenshots not in the recipe. Don't add `eval` calls not in the recipe. Don't open extra DOM views to "investigate". Recipe is the contract.
|
|
92
|
-
- Truncate large outputs aggressively. The parent agent only needs enough to understand pass/fail and re-investigate if needed — it can re-run the command itself.
|
|
93
|
-
- If the recipe is malformed (no `## Evidence Commands` section, no fenced blocks), return a JSON report with `summary.total: 0` and `blocking_issues: ["recipe malformed: <reason>"]`.
|
|
94
|
-
|
|
95
|
-
## Token discipline
|
|
122
|
+
## Boundaries (re-stated for clarity)
|
|
96
123
|
|
|
97
|
-
|
|
124
|
+
- **No exploration.** Already covered in budgets, restating because this is the failure mode. The parent agent has Opus/Sonnet to investigate; you have Haiku to execute a script. Stay in lane.
|
|
125
|
+
- **No fix suggestions.** `diagnosis` is descriptive only ("returned 0, expected > 0"). Never "you should change X" or "try Y instead".
|
|
126
|
+
- **Truncate aggressively.** Stdout > 500 chars → truncate with `…[truncated, full output reproducible by re-running]`. Parent agent can re-run cherry-picked commands itself.
|
|
127
|
+
- **One JSON block, nothing else.** Anything you print outside the JSON wastes the parent agent's context — which is the entire reason you exist.
|
package/package.json
CHANGED
package/skills/verify/SKILL.md
CHANGED
|
@@ -184,17 +184,29 @@ If the developer points you at a `.claude/verify-recipes/<slug>.md` file, OR you
|
|
|
184
184
|
|
|
185
185
|
Why: recipe execution is mechanical (run commands, compare to `Expected:`, report). It does not need Opus-level reasoning, but the raw output (DOM dumps, screenshots, eval results) easily exceeds 30k tokens of context noise. Delegating to a Haiku subagent keeps your context clean and cuts cost ~10×.
|
|
186
186
|
|
|
187
|
-
|
|
187
|
+
**Pre-flight check (do this BEFORE spawning the runner):**
|
|
188
188
|
|
|
189
|
-
1.
|
|
190
|
-
2.
|
|
189
|
+
1. `Read` the recipe.
|
|
190
|
+
2. Confirm it has a `## Machine Preconditions` section. If missing → it's older 0.6.x format. Tell the user once: "this recipe doesn't have Machine Preconditions, so the runner can't distinguish setup failures from real bugs. Run anyway, or update the recipe first?" Default to running but flag failures with this caveat in the final summary.
|
|
191
|
+
3. Read the `## Manual Preconditions` block out loud to the user (one tight sentence each) and ask: "is the app set up this way right now?" Wait for confirmation. If they say no — stop, don't waste a runner spawn.
|
|
192
|
+
4. Resolve the window id once: `agent-view discover` → pick the main window's id.
|
|
193
|
+
|
|
194
|
+
**Spawn:**
|
|
195
|
+
|
|
196
|
+
5. Spawn `verify-runner` via the Agent tool with a prompt containing:
|
|
191
197
|
- Absolute `recipe_path`
|
|
192
198
|
- Resolved `window_id`
|
|
199
|
+
- `mode: full` (or `dry_run` if the user asked to validate the recipe first)
|
|
193
200
|
- Any extra context (e.g. "user is already logged in", "GIS widget already mounted")
|
|
194
|
-
|
|
195
|
-
|
|
196
|
-
|
|
197
|
-
|
|
201
|
+
6. Wait for the JSON report.
|
|
202
|
+
|
|
203
|
+
**Handle the report:**
|
|
204
|
+
|
|
205
|
+
7. **If `status: precondition_failed`** → relay `failed_precondition` and `manual_preconditions_to_check` to the user verbatim. Do NOT spawn the runner again. Do NOT try to "fix" the missing setup yourself by clicking around — the user knows their app, ask them to do the manual setup and re-run.
|
|
206
|
+
8. **If `status: cascading_failures` or `budget_exhausted`** → the recipe is likely stale (selectors/refs changed, UI restructured) or describes a flow that no longer matches the app. Show the first 1-2 failed steps to the user and ask whether to update the recipe (delegate back to `verify-recipe`) or investigate manually.
|
|
207
|
+
9. **If `design_conformance_section: true`** in the report: also spawn `design-conformance-runner` in parallel, passing the `design_conformance_pairs` array. Merge both reports before answering the user.
|
|
208
|
+
10. **If `requires_visual_review` steps** have no design ref attached: open each screenshot yourself with `Read` and decide pass/fail.
|
|
209
|
+
11. **If individual `failed` steps** in an otherwise-completed run: re-run that specific failing command yourself for richer evidence to diagnose. Do not re-execute the whole recipe.
|
|
198
210
|
|
|
199
211
|
Output to the user: a tight summary — what passed, what failed, what needs visual review, and (if any) which design conformance issues to fix. Do not paste the raw JSON unless asked.
|
|
200
212
|
|
|
@@ -1,202 +1,239 @@
|
|
|
1
|
-
---
|
|
2
|
-
name: verify-recipe
|
|
3
|
-
description: "Generates a concrete, command-by-command agent-view recipe for verifying a feature or bugfix. Use when the developer wants to write a verify-recipe.md, create a verification plan, build an agent-view recipe, or produce a verify checklist for a change they shipped. Triggers on: write a verify-recipe, write verification steps, generate verify steps for the feature/bug I just shipped/fixed, make a verify-recipe.md, what should I run to verify X, create a verification plan, agent-view recipe for this fix, verify checklist for this feature. Does NOT execute the checks — it authors the plan. For running checks against a live app, use the verify skill instead."
|
|
4
|
-
allowed-tools: Read, Write, Bash(agent-view *)
|
|
5
|
-
---
|
|
6
|
-
|
|
7
|
-
# Verification Recipe Generator
|
|
8
|
-
|
|
9
|
-
You help the developer author a disciplined, cheapest-first verification recipe for a feature they shipped or a bug they fixed. You do not run the checks — you produce a `.claude/verify-recipes/<slug>.md` file that any AI coding agent can execute later.
|
|
10
|
-
|
|
11
|
-
## What this produces
|
|
12
|
-
|
|
13
|
-
A file at `.claude/verify-recipes/<kebab-slug>.md` containing:
|
|
14
|
-
|
|
15
|
-
- **
|
|
16
|
-
- **
|
|
17
|
-
- **
|
|
18
|
-
- **
|
|
19
|
-
- **
|
|
20
|
-
|
|
21
|
-
|
|
22
|
-
|
|
23
|
-
|
|
24
|
-
|
|
25
|
-
|
|
26
|
-
|
|
27
|
-
|
|
28
|
-
|
|
29
|
-
|
|
30
|
-
|
|
31
|
-
|
|
32
|
-
|
|
33
|
-
##
|
|
34
|
-
|
|
35
|
-
|
|
36
|
-
|
|
37
|
-
|
|
38
|
-
|
|
39
|
-
|
|
40
|
-
|
|
41
|
-
|
|
42
|
-
|
|
43
|
-
|
|
44
|
-
|
|
45
|
-
|
|
46
|
-
|
|
47
|
-
|
|
48
|
-
|
|
49
|
-
|
|
50
|
-
|
|
51
|
-
|
|
52
|
-
|
|
53
|
-
|
|
54
|
-
|
|
55
|
-
|
|
56
|
-
|
|
57
|
-
|
|
58
|
-
|
|
59
|
-
|
|
60
|
-
|
|
61
|
-
|
|
62
|
-
|
|
63
|
-
|
|
64
|
-
|
|
65
|
-
|
|
66
|
-
|
|
67
|
-
|
|
68
|
-
|
|
69
|
-
### Step
|
|
70
|
-
|
|
71
|
-
|
|
72
|
-
|
|
73
|
-
|
|
74
|
-
|
|
75
|
-
|
|
76
|
-
|
|
77
|
-
|
|
78
|
-
|
|
79
|
-
|
|
80
|
-
|
|
81
|
-
|
|
82
|
-
|
|
83
|
-
|
|
84
|
-
|
|
85
|
-
|
|
86
|
-
|
|
87
|
-
|
|
88
|
-
|
|
89
|
-
|
|
90
|
-
|
|
91
|
-
|
|
92
|
-
|
|
93
|
-
|
|
94
|
-
|
|
95
|
-
|
|
96
|
-
|
|
97
|
-
|
|
98
|
-
|
|
99
|
-
|
|
100
|
-
-
|
|
101
|
-
|
|
102
|
-
|
|
103
|
-
-
|
|
104
|
-
|
|
105
|
-
|
|
106
|
-
|
|
107
|
-
<!--
|
|
108
|
-
|
|
109
|
-
|
|
110
|
-
|
|
111
|
-
|
|
112
|
-
|
|
113
|
-
|
|
114
|
-
|
|
115
|
-
|
|
116
|
-
|
|
117
|
-
|
|
118
|
-
|
|
119
|
-
|
|
120
|
-
|
|
121
|
-
|
|
122
|
-
|
|
123
|
-
|
|
124
|
-
|
|
125
|
-
|
|
126
|
-
|
|
127
|
-
|
|
128
|
-
|
|
129
|
-
|
|
130
|
-
|
|
131
|
-
|
|
132
|
-
|
|
133
|
-
|
|
134
|
-
|
|
135
|
-
|
|
136
|
-
|
|
137
|
-
|
|
138
|
-
|
|
139
|
-
|
|
140
|
-
|
|
141
|
-
|
|
142
|
-
|
|
143
|
-
|
|
144
|
-
|
|
145
|
-
|
|
146
|
-
|
|
147
|
-
|
|
148
|
-
|
|
149
|
-
|
|
150
|
-
|
|
151
|
-
|
|
152
|
-
|
|
153
|
-
|
|
154
|
-
|
|
155
|
-
|
|
156
|
-
|
|
157
|
-
|
|
158
|
-
|
|
159
|
-
|
|
160
|
-
agent-view
|
|
161
|
-
|
|
162
|
-
|
|
163
|
-
|
|
164
|
-
|
|
165
|
-
|
|
166
|
-
|
|
167
|
-
|
|
168
|
-
|
|
169
|
-
|
|
170
|
-
|
|
171
|
-
|
|
172
|
-
|
|
173
|
-
|
|
174
|
-
|
|
175
|
-
|
|
176
|
-
|
|
177
|
-
|
|
178
|
-
|
|
179
|
-
|
|
180
|
-
|
|
181
|
-
|
|
182
|
-
|
|
183
|
-
|
|
184
|
-
|
|
185
|
-
|
|
186
|
-
|
|
187
|
-
|
|
188
|
-
|
|
189
|
-
##
|
|
190
|
-
-
|
|
191
|
-
|
|
192
|
-
|
|
193
|
-
|
|
194
|
-
|
|
195
|
-
|
|
196
|
-
|
|
197
|
-
|
|
198
|
-
|
|
199
|
-
|
|
200
|
-
|
|
201
|
-
|
|
202
|
-
|
|
1
|
+
---
|
|
2
|
+
name: verify-recipe
|
|
3
|
+
description: "Generates a concrete, command-by-command agent-view recipe for verifying a feature or bugfix. Use when the developer wants to write a verify-recipe.md, create a verification plan, build an agent-view recipe, or produce a verify checklist for a change they shipped. Triggers on: write a verify-recipe, write verification steps, generate verify steps for the feature/bug I just shipped/fixed, make a verify-recipe.md, what should I run to verify X, create a verification plan, agent-view recipe for this fix, verify checklist for this feature. Does NOT execute the checks — it authors the plan. For running checks against a live app, use the verify skill instead."
|
|
4
|
+
allowed-tools: Read, Write, Bash(agent-view *)
|
|
5
|
+
---
|
|
6
|
+
|
|
7
|
+
# Verification Recipe Generator
|
|
8
|
+
|
|
9
|
+
You help the developer author a disciplined, cheapest-first verification recipe for a feature they shipped or a bug they fixed. You do not run the checks — you produce a `.claude/verify-recipes/<slug>.md` file that any AI coding agent can execute later.
|
|
10
|
+
|
|
11
|
+
## What this produces
|
|
12
|
+
|
|
13
|
+
A file at `.claude/verify-recipes/<kebab-slug>.md` containing:
|
|
14
|
+
|
|
15
|
+
- **MANUAL PRECONDITIONS** — human-readable setup steps (drag widget here, navigate to view X) that a person or the parent agent must do before any automated check
|
|
16
|
+
- **MACHINE PRECONDITIONS** — runnable `agent-view` checks that prove the manual setup actually took effect; the verify-runner subagent runs these FIRST and aborts cleanly if any fail
|
|
17
|
+
- **NARROWED SIGNAL** — the measurable indicator that proves success or failure
|
|
18
|
+
- **EVIDENCE COMMANDS** — ordered `agent-view` calls, cheapest first, each annotated with what it proves
|
|
19
|
+
- **POSITIVE-CASE ASSERTIONS** — what "pass" looks like for each command
|
|
20
|
+
- **REGRESSION CHECKS** — adjacent paths that must not have broken
|
|
21
|
+
- **DESIGN CONFORMANCE** (optional) — screenshot ↔ design reference pairs for the design-conformance-runner
|
|
22
|
+
|
|
23
|
+
Create the directory if missing: `mkdir -p .claude/verify-recipes`
|
|
24
|
+
|
|
25
|
+
## Why two kinds of preconditions
|
|
26
|
+
|
|
27
|
+
The verify-runner subagent is intentionally tightly scoped — it executes commands and reports results, with hard budgets that prevent it from "looking around" when things don't match. That means **the recipe must clearly separate what a human does from what a machine verifies**.
|
|
28
|
+
|
|
29
|
+
If you write a precondition as prose ("the GIS widget is dragged into a workspace cell"), the runner can't check it. If the user forgot to do it, the runner blunders into Evidence Commands that depend on missing UI, and either burns its budget on a recipe-stale abort or — worse, in older formats — flails trying to find the missing element.
|
|
30
|
+
|
|
31
|
+
The fix: every Manual Precondition gets a paired Machine Precondition that proves it took effect. Drag widget into cell → check `cesiumReadyFlag === true`. Navigate to map mode → check `document.querySelector('.cesium-widget')` exists. Now if the user skips a step, the runner aborts cleanly on Phase 1 with a clear "do this first" message instead of diagnosing phantom bugs.
|
|
32
|
+
|
|
33
|
+
## Methodology
|
|
34
|
+
|
|
35
|
+
Frame the recipe with **hard-debug** discipline. Apply this chain in authoring mode:
|
|
36
|
+
|
|
37
|
+
1. Start from a **reproducible, machine-checkable** starting state — not "open the app and poke around"
|
|
38
|
+
2. Convert vague expectations ("looks right") into measurable signals ("store.user.role === 'admin'")
|
|
39
|
+
3. Prefer the cheapest tool that can answer the question — a value check costs ~50 tokens, a screenshot costs ~6 000
|
|
40
|
+
4. Include at least one negative-case check (the old symptom must no longer appear)
|
|
41
|
+
5. Include at least one regression check (an adjacent flow must still work)
|
|
42
|
+
6. **Every Manual Precondition needs a Machine Precondition counterpart.** If you can't think of one — interview the developer further before writing the recipe (see Step 1 below).
|
|
43
|
+
|
|
44
|
+
## Tool-cost decision tree
|
|
45
|
+
|
|
46
|
+
Pick the first row that can answer the question. Only go lower when the row above can't:
|
|
47
|
+
|
|
48
|
+
| Question | Command | Why it's cheapest |
|
|
49
|
+
|---|---|---|
|
|
50
|
+
| Element exists / has specific text / role | `agent-view dom --filter "<text>" --depth 2` | Structured text, zero vision tokens |
|
|
51
|
+
| App state, store value, computed flag | `agent-view eval "<expr>"` | Returns the value directly; DOM inference is wasteful and fragile |
|
|
52
|
+
| What changed between action and final state | `agent-view watch "<expr>" --until "<condition>"` or `--max-changes 1` | `eval` shows the snapshot; `watch` shows the trajectory |
|
|
53
|
+
| SharedWorker / ServiceWorker internal state | `agent-view eval --target <name> "<expr>"` | Workers have no DOM; this is the only path |
|
|
54
|
+
| Did this action throw or warn silently? | `agent-view console --clear` before, `agent-view console --level error,warn` after | Catches uncaught exceptions and network failures invisible to DOM |
|
|
55
|
+
| Layout, spacing, visual regression | `agent-view screenshot --scale 0.5` | Last resort — the only tool that sees pixels, but costs ~6 000 tokens |
|
|
56
|
+
| Canvas / WebGL scene state | `agent-view scene --diff` | DOM is empty for canvas apps |
|
|
57
|
+
|
|
58
|
+
**Anti-patterns to reject:**
|
|
59
|
+
- Mixing manual setup and machine checks under a single "Repro Steps" heading. Always split into Manual Preconditions + Machine Preconditions.
|
|
60
|
+
- Manual Precondition without a Machine counterpart — the runner can't verify it, so the user can silently violate it.
|
|
61
|
+
- Opening Evidence Commands with a screenshot to "see the state" — use `dom --filter` or `eval` first
|
|
62
|
+
- Using `eval` when `dom --filter` answers the question
|
|
63
|
+
- Assertions that depend on transient state without `watch --until` to stabilize first
|
|
64
|
+
- "Check that it looks right" — every Evidence assertion must be a concrete pass/fail criterion. The single legitimate exception is the `## Design Conformance` section, which delegates visual judgment to the design-conformance-runner subagent against an explicit reference image.
|
|
65
|
+
- Inventing design reference paths (`.figma-refs/...`, `assets/mockups/...`) when the developer did not provide them. No refs → no Design Conformance section.
|
|
66
|
+
|
|
67
|
+
## Workflow
|
|
68
|
+
|
|
69
|
+
### Step 1 — gather context (interview the developer)
|
|
70
|
+
|
|
71
|
+
When invoked, ask the developer in plain text (no tool calls yet). Wait for the response before continuing.
|
|
72
|
+
|
|
73
|
+
1. **What was shipped or fixed?** (feature name or bug description)
|
|
74
|
+
2. **What was the original symptom or expected behavior?**
|
|
75
|
+
3. **Any known failure mode or edge case to cover?**
|
|
76
|
+
4. **UI mode requirements**: Does this feature live behind a specific app mode/view that must be active before it appears (e.g., "map view, not settings panel"; "edit mode, not view mode"; "modal X must be open")? List every mode-toggle the user must have done.
|
|
77
|
+
5. **Manual setup steps**: Beyond modes, what physical actions must the user do before checks can run (drag a widget into a cell, search for and select a map location, open a specific dialog, log in as a particular role)? Be precise — these become Manual Preconditions verbatim.
|
|
78
|
+
6. **State assertions for each manual step**: For each item in (4) and (5), is there a JS expression or DOM selector that proves it happened? Examples: "after dragging the GIS widget — `pinia._s.get('gis-widget-root')?.cesiumReadyFlag` becomes true"; "after entering map mode — `document.querySelector('.cesium-widget')` exists; "after selecting a location — `selectedLocation` is truthy". If the developer doesn't know offhand, that's fine — ask them to point you at the store/composable/component where state lives and you can suggest expressions.
|
|
79
|
+
7. **(Optional) Design references**: Local image paths to compare screenshots against (Figma exports, hand-off PNGs). **Only local files are supported.** If none — skip the Design Conformance section entirely.
|
|
80
|
+
|
|
81
|
+
If the answers to (4)/(5) reveal something the developer can't pair with a machine check (6), say so explicitly: "I'll write `<step>` as a Manual Precondition with no Machine counterpart — that means if a user skips it, the runner won't catch it and may report misleading failures. Want to add a custom check?" Then either (a) get an expression from them, or (b) accept the gap and note it in the recipe's Anti-patterns section.
|
|
82
|
+
|
|
83
|
+
### Step 2 — draft the recipe
|
|
84
|
+
|
|
85
|
+
Use the answers to produce a recipe in this format:
|
|
86
|
+
|
|
87
|
+
````markdown
|
|
88
|
+
# Verify: <feature or fix name>
|
|
89
|
+
|
|
90
|
+
Generated: <date>
|
|
91
|
+
Scope: <one sentence describing what this covers>
|
|
92
|
+
|
|
93
|
+
## Manual Preconditions
|
|
94
|
+
<!-- Done by a human or the parent agent BEFORE invoking verify-runner. The runner does NOT execute these. -->
|
|
95
|
+
1. <Action 1 — exact, no ambiguity. e.g. "Open the GIS widget by dragging it from the 'Edit workspace' panel into the upper-left cell.">
|
|
96
|
+
2. <Action 2 — e.g. "In the map header search box, type 'Склад_1' and click the matching dropdown entry to fly the camera to that sublocation.">
|
|
97
|
+
3. <Action 3 — e.g. "Zoom in until building details are visible (camera height < 1000 m).">
|
|
98
|
+
|
|
99
|
+
## Machine Preconditions
|
|
100
|
+
<!-- The verify-runner runs these FIRST. If ANY fail, it aborts with `precondition_failed` and shows the Manual Preconditions block to the user. -->
|
|
101
|
+
- `agent-view eval "window.__dev !== undefined"` → must be `true`
|
|
102
|
+
- `agent-view eval "!!window.__dev.pinia._s.get('gis-widget-root')?.cesiumReadyFlag"` → must be `true`
|
|
103
|
+
- `agent-view eval "!!window.__dev.pinia._s.get('gis-widget-root')?.selectedLocation"` → must be `true`
|
|
104
|
+
- `agent-view eval "document.querySelector('.cesium-widget') !== null"` → must be `true`
|
|
105
|
+
|
|
106
|
+
## Narrowed Signal
|
|
107
|
+
<!-- The one measurable thing that proves it works -->
|
|
108
|
+
`<agent-view command>` must return `<expected value>`.
|
|
109
|
+
|
|
110
|
+
## Evidence Commands
|
|
111
|
+
|
|
112
|
+
### 1. <What this proves>
|
|
113
|
+
```bash
|
|
114
|
+
agent-view <command>
|
|
115
|
+
```
|
|
116
|
+
Expected: <concrete criterion — value, text, absence of error>
|
|
117
|
+
Cost: ~<N> tokens
|
|
118
|
+
|
|
119
|
+
### 2. ...
|
|
120
|
+
|
|
121
|
+
## Positive-Case Assertions
|
|
122
|
+
- [ ] <criterion>
|
|
123
|
+
- [ ] <criterion>
|
|
124
|
+
|
|
125
|
+
## Regression Checks
|
|
126
|
+
- [ ] <adjacent flow> — `agent-view <command>` → `<expected>`
|
|
127
|
+
|
|
128
|
+
## Design Conformance
|
|
129
|
+
<!-- INCLUDE THIS SECTION ONLY IF the developer provided design refs in question 7. -->
|
|
130
|
+
<!-- Each row pairs a screenshot command with the expected reference image path. -->
|
|
131
|
+
<!-- The design-conformance-runner subagent reads this section, runs the screenshot commands, and visually compares against the expected refs. -->
|
|
132
|
+
<!-- Do NOT invent design ref paths. Use exactly what the developer provided. -->
|
|
133
|
+
|
|
134
|
+
| Step Label | Screenshot Command | Expected Reference |
|
|
135
|
+
|---|---|---|
|
|
136
|
+
| <area name e.g. "filter panel collapsed"> | `agent-view screenshot --crop "<area>" --scale 0.5` | `<absolute path to expected PNG/JPEG>` |
|
|
137
|
+
| <area name> | `agent-view screenshot --window $W --scale 0.5` | `<absolute path>` |
|
|
138
|
+
|
|
139
|
+
Tolerance: `normal` (default — flag deviations a designer would notice in code review). Use `loose` only if the developer says exact pixel parity is not required.
|
|
140
|
+
|
|
141
|
+
## Anti-patterns avoided
|
|
142
|
+
- <note any recipe-specific traps, e.g. "manual step 3 (zoom) has no machine counterpart — runner cannot detect insufficient zoom; mitigated by Evidence Step N which checks camera height">
|
|
143
|
+
````
|
|
144
|
+
|
|
145
|
+
### Step 3 — save the file
|
|
146
|
+
|
|
147
|
+
Determine a kebab-slug from the feature/fix name (e.g. `login-redirect-fix`, `cart-total-display`).
|
|
148
|
+
|
|
149
|
+
Save to `.claude/verify-recipes/<slug>.md`. Create the directory first if it doesn't exist.
|
|
150
|
+
|
|
151
|
+
Confirm the path to the developer.
|
|
152
|
+
|
|
153
|
+
### Step 4 — offer dry-run validation
|
|
154
|
+
|
|
155
|
+
After saving, ask the developer:
|
|
156
|
+
|
|
157
|
+
> The recipe is saved at `<path>`. Is the app running? If yes, I can spawn `verify-runner` in `dry_run` mode — it'll execute only the Machine Preconditions and the first Evidence Command (~5 commands total). That validates the recipe isn't broken before you commit a full run. Want me to run the dry-run? (yes/no)
|
|
158
|
+
|
|
159
|
+
If yes:
|
|
160
|
+
1. Get the window id with `agent-view discover`.
|
|
161
|
+
2. Spawn `verify-runner` via the Agent tool with `mode: dry_run`, the recipe path, and the window id.
|
|
162
|
+
3. Read the JSON report.
|
|
163
|
+
4. If `status: completed` and dry-run preconditions+step 1 passed → tell the developer the recipe is healthy and ready for a full run.
|
|
164
|
+
5. If `precondition_failed` → relay the `failed_precondition` and `manual_preconditions_to_check` so the developer knows what setup step to do (or what Machine Precondition to fix).
|
|
165
|
+
6. If the first Evidence Command failed → flag it: the recipe likely has a stale ref/selector or assumes UI state that doesn't exist; offer to revise.
|
|
166
|
+
|
|
167
|
+
If no — confirm the path and stop.
|
|
168
|
+
|
|
169
|
+
## Worked example: "fixed login redirect bug"
|
|
170
|
+
|
|
171
|
+
**Developer input:**
|
|
172
|
+
> Fixed a bug where after login, the redirect went to `/home` instead of `/dashboard`. Store mutation `SET_REDIRECT_PATH` was missing. No visual change — purely a routing issue. No special UI mode — just the login page. Manual setup is "be on the /login route".
|
|
173
|
+
|
|
174
|
+
**Recipe produced:**
|
|
175
|
+
|
|
176
|
+
````markdown
|
|
177
|
+
# Verify: Login Redirect Fix
|
|
178
|
+
|
|
179
|
+
Generated: 2026-04-27
|
|
180
|
+
Scope: Confirms that a successful login routes to /dashboard, not /home, and that the store mutation fires correctly.
|
|
181
|
+
|
|
182
|
+
## Manual Preconditions
|
|
183
|
+
1. App running, user logged out, browser at the `/login` route.
|
|
184
|
+
|
|
185
|
+
## Machine Preconditions
|
|
186
|
+
- `agent-view eval "router.currentRoute.path"` → must be `"/login"`
|
|
187
|
+
- `agent-view eval "!!store.state.auth.user"` → must be `false` (logged out)
|
|
188
|
+
|
|
189
|
+
## Narrowed Signal
|
|
190
|
+
`agent-view eval "router.currentRoute.path"` must return `"/dashboard"` after sign-in.
|
|
191
|
+
|
|
192
|
+
## Evidence Commands
|
|
193
|
+
|
|
194
|
+
### 0. Setup — baseline console
|
|
195
|
+
```bash
|
|
196
|
+
agent-view console --clear
|
|
197
|
+
```
|
|
198
|
+
Expected: `Console buffer cleared`
|
|
199
|
+
Cost: ~10 tokens
|
|
200
|
+
|
|
201
|
+
### 1. Confirm redirect target
|
|
202
|
+
```bash
|
|
203
|
+
agent-view dom --filter "Email" --depth 2
|
|
204
|
+
agent-view fill <email-ref> "admin@example.com"
|
|
205
|
+
agent-view fill <password-ref> "password"
|
|
206
|
+
agent-view click <signin-ref>
|
|
207
|
+
agent-view watch "router.currentRoute.path" --until "router.currentRoute.path === '/dashboard'"
|
|
208
|
+
```
|
|
209
|
+
Expected: `replace / "/login" → "/dashboard"` in watch output
|
|
210
|
+
Cost: ~150 tokens
|
|
211
|
+
|
|
212
|
+
### 2. Confirm mutation fired
|
|
213
|
+
```bash
|
|
214
|
+
agent-view eval "store.state.auth.redirectPath"
|
|
215
|
+
```
|
|
216
|
+
Expected: `"/dashboard"` (not `"/home"`, not `null`)
|
|
217
|
+
Cost: ~50 tokens
|
|
218
|
+
|
|
219
|
+
### 3. No errors during login flow
|
|
220
|
+
```bash
|
|
221
|
+
agent-view console --level error,warn
|
|
222
|
+
```
|
|
223
|
+
Expected: `(no console messages)`
|
|
224
|
+
Cost: ~30 tokens
|
|
225
|
+
|
|
226
|
+
## Positive-Case Assertions
|
|
227
|
+
- [ ] `router.currentRoute.path` === `/dashboard` after login
|
|
228
|
+
- [ ] `store.state.auth.redirectPath` === `/dashboard`
|
|
229
|
+
- [ ] No console errors during the flow
|
|
230
|
+
|
|
231
|
+
## Regression Checks
|
|
232
|
+
- [ ] Logout → `/login` still works (`agent-view click <logout-ref>` then `agent-view eval "router.currentRoute.path"` → `"/login"`)
|
|
233
|
+
|
|
234
|
+
## Anti-patterns avoided
|
|
235
|
+
- Not using screenshot to confirm route (route is a string — eval is 120× cheaper)
|
|
236
|
+
- watch used before eval so the route change is confirmed to have settled, not just sampled mid-transition
|
|
237
|
+
````
|
|
238
|
+
|
|
239
|
+
**Saved to:** `.claude/verify-recipes/login-redirect-fix.md`
|