@iceinvein/agent-skills 0.1.37 → 0.1.39

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (79) hide show
  1. package/package.json +1 -1
  2. package/skills/bounded-context-auditor/SKILL.md +15 -3
  3. package/skills/bounded-context-auditor/skill.json +1 -1
  4. package/skills/codebase-architecture/SKILL.md +4 -4
  5. package/skills/codebase-architecture/skill.json +1 -1
  6. package/skills/cognitive-load-auditor/SKILL.md +11 -9
  7. package/skills/cognitive-load-auditor/skill.json +1 -1
  8. package/skills/cohesion-analyzer/SKILL.md +1 -1
  9. package/skills/cohesion-analyzer/skill.json +1 -1
  10. package/skills/composability-auditor/SKILL.md +5 -5
  11. package/skills/composability-auditor/skill.json +1 -1
  12. package/skills/contract-enforcer/SKILL.md +5 -5
  13. package/skills/contract-enforcer/skill.json +1 -1
  14. package/skills/coupling-auditor/SKILL.md +2 -2
  15. package/skills/coupling-auditor/skill.json +1 -1
  16. package/skills/cover-letter/SKILL.md +18 -20
  17. package/skills/cover-letter/skill.json +7 -2
  18. package/skills/cover-letter-audit/SKILL.md +20 -20
  19. package/skills/cover-letter-audit/skill.json +7 -2
  20. package/skills/cover-letter-persona/SKILL.md +13 -13
  21. package/skills/cover-letter-persona/skill.json +7 -2
  22. package/skills/cover-letter-rewrite/SKILL.md +18 -16
  23. package/skills/cover-letter-rewrite/skill.json +7 -2
  24. package/skills/cover-letter-write/SKILL.md +25 -20
  25. package/skills/cover-letter-write/skill.json +7 -2
  26. package/skills/cqs-auditor/SKILL.md +19 -47
  27. package/skills/cqs-auditor/skill.json +1 -1
  28. package/skills/demeter-enforcer/SKILL.md +5 -5
  29. package/skills/demeter-enforcer/skill.json +1 -1
  30. package/skills/dependency-direction-auditor/SKILL.md +1 -1
  31. package/skills/dependency-direction-auditor/skill.json +1 -1
  32. package/skills/design-review/SKILL.md +6 -2
  33. package/skills/design-review/skill.json +1 -1
  34. package/skills/error-strategist/SKILL.md +3 -3
  35. package/skills/error-strategist/skill.json +1 -1
  36. package/skills/event-design-reviewer/SKILL.md +3 -3
  37. package/skills/event-design-reviewer/skill.json +1 -1
  38. package/skills/evolution-analyzer/SKILL.md +4 -3
  39. package/skills/evolution-analyzer/skill.json +1 -1
  40. package/skills/gestalt-reviewer/SKILL.md +8 -4
  41. package/skills/gestalt-reviewer/skill.json +1 -1
  42. package/skills/idempotency-guardian/SKILL.md +6 -6
  43. package/skills/idempotency-guardian/skill.json +1 -1
  44. package/skills/improve-my-codebase/CATALOGUE-FIELDS.md +2 -2
  45. package/skills/improve-my-codebase/SKILL.md +68 -27
  46. package/skills/improve-my-codebase/skill.json +1 -1
  47. package/skills/index.json +33 -33
  48. package/skills/integration-pattern-auditor/SKILL.md +2 -2
  49. package/skills/integration-pattern-auditor/skill.json +1 -1
  50. package/skills/magpie/README.md +3 -5
  51. package/skills/magpie/SKILL.md +39 -536
  52. package/skills/magpie/package.json +1 -1
  53. package/skills/magpie/references/critic.md +58 -0
  54. package/skills/magpie/references/peer-review.md +84 -0
  55. package/skills/magpie/references/specialists.md +391 -0
  56. package/skills/magpie/scripts/__tests__/helper.test.ts +40 -0
  57. package/skills/magpie/scripts/__tests__/skill-lint.test.ts +116 -28
  58. package/skills/magpie/scripts/__tests__/status-cmd.test.ts +13 -0
  59. package/skills/magpie/scripts/helper.js +24 -13
  60. package/skills/magpie/scripts/status-cmd.ts +10 -1
  61. package/skills/magpie/skill.json +2 -1
  62. package/skills/module-secret-auditor/SKILL.md +8 -5
  63. package/skills/module-secret-auditor/skill.json +1 -1
  64. package/skills/port-adapter-auditor/SKILL.md +3 -3
  65. package/skills/port-adapter-auditor/skill.json +1 -1
  66. package/skills/rams-design-audit/SKILL.md +4 -2
  67. package/skills/rams-design-audit/skill.json +1 -1
  68. package/skills/seam-finder/SKILL.md +2 -2
  69. package/skills/seam-finder/skill.json +1 -1
  70. package/skills/simplicity-razor/SKILL.md +4 -4
  71. package/skills/simplicity-razor/skill.json +1 -1
  72. package/skills/temporal-coupling-detector/SKILL.md +2 -2
  73. package/skills/temporal-coupling-detector/skill.json +1 -1
  74. package/skills/terse/SKILL.md +12 -7
  75. package/skills/terse/skill.json +1 -1
  76. package/skills/type-driven-designer/SKILL.md +8 -8
  77. package/skills/type-driven-designer/skill.json +1 -1
  78. package/skills/unidirectional-flow-enforcer/SKILL.md +2 -2
  79. package/skills/unidirectional-flow-enforcer/skill.json +1 -1
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: magpie
3
- description: Interactive PR review pipeline. Runs five parallel specialist subagents (security, bugs, performance, code-smells, architecture), dedupes findings, applies a critic rubric, peer-reviews via codex exec (falling back to a Claude second opinion when codex is unavailable), and serves an interactive HTML report for selecting findings to post via gh. Use when the user asks to review a GitHub pull request.
3
+ description: Use when the user asks to review a GitHub pull request (a PR number, PR URL, or "review this PR"). Interactive PR review pipeline; five parallel specialist subagents, dedupe, critic, codex/Claude peer review, and an interactive HTML report for posting selected findings via gh. Follow the stage walkthrough in this skill; the description is not the procedure.
4
4
  ---
5
5
 
6
6
  # Magpie
@@ -17,7 +17,15 @@ Stop reading and follow these steps in order. Do not skip stages. Use the exact
17
17
 
18
18
  Parse the user's request for a PR number, URL, or "this PR" (current branch). If ambiguous, ask one clarifying terminal question. Capture the PR number into `$PR_NUMBER` and the repository path into `$REPO` (default: current working directory).
19
19
 
20
- Compute the run directory:
20
+ Then check whether an earlier run on this PR is still unfinished, before minting a new id:
21
+
22
+ ```
23
+ magpie --list-runs
24
+ ```
25
+
26
+ Each line is `<id>\t<active|archived>\t<path>`. If an `active` id matches `pr-${PR_NUMBER}-*`, that run was interrupted rather than cleaned up. Set `RUN_DIR` to its path and go to "Resuming a crashed run" instead of starting over; ask the user first if it is unclear whether they want to resume or review from scratch. (`archived` ids are finished runs, not resumable.)
27
+
28
+ Otherwise compute a fresh run directory:
21
29
 
22
30
  ```
23
31
  RUN_ID="pr-${PR_NUMBER}-$(date +%s)"
@@ -38,6 +46,8 @@ When a prior run exists for the same PR (active or archived under `~/.magpie/`),
38
46
 
39
47
  Setup also runs a deterministic test-coverage check: when the diff contains zero test or spec files anywhere, each non-test source file with `>= 10` added code lines gets a `domain: "tests"` finding written to `$RUN_DIR/findings/tests.json`. This is a sixth domain that flows through dedupe/critic/peer-review alongside the five LLM specialists. No specialist subagent is dispatched for it.
40
48
 
49
+ The pipeline has no separate context-indexing stage, but `magpie status` and the progress page track one. After setup succeeds, append `{stage: context, status: skipped}` to `$RUN_DIR/log.jsonl`; both treat a skipped stage as behind them, so the pipeline advances to `specialists`.
50
+
41
51
  ### 2. Serve
42
52
 
43
53
  Start the HTML server in the background using the Bash tool with `run_in_background: true`:
@@ -46,7 +56,9 @@ Start the HTML server in the background using the Bash tool with `run_in_backgro
46
56
  magpie serve "$RUN_DIR"
47
57
  ```
48
58
 
49
- Read `$RUN_DIR/state/server-info` for the URL. Print to the user: "Open <url> in your browser to follow along."
59
+ Read `$RUN_DIR/state/server-info` for the URL; the server writes it asynchronously at startup, so if the file doesn't exist yet, wait a moment and re-read (it appears within ~1s). Print to the user: "Open <url> in your browser to follow along."
60
+
61
+ The server shuts down after 30 minutes with no requests (an open report tab heartbeats every 30s, so it stays up while the user is looking at it) and deletes `state/server-info` on the way out. Nothing in the pipeline depends on it staying alive: re-run `magpie serve "$RUN_DIR"` to bring the report back.
50
62
 
51
63
  Render the first progress paint:
52
64
 
@@ -56,75 +68,9 @@ magpie render "$RUN_DIR" progress
56
68
 
57
69
  ### 3. Specialists
58
70
 
59
- Dispatch the five specialist subagents in a single message using five Agent tool calls in parallel. For each focus in (security, bugs, performance, code-smells, architecture), the prompt is:
60
-
61
- ````
62
- <specialist block for focus from this SKILL.md>
63
-
64
- You are reviewing PR #<PR_NUMBER>.
65
- Working directory: $RUN_DIR/worktree
66
- Diff: $RUN_DIR/diff.patch
67
-
68
- ## Output Contract
69
-
70
- Write findings to $RUN_DIR/findings/<focus>.json before returning. The file MUST be a JSON array. Each entry MUST conform to this schema exactly (no extra top-level keys, no renamed keys):
71
-
72
- {
73
- "id": string, // e.g. "<focus>-1", "<focus>-2"; unique per focus
74
- "file": string, // path relative to worktree
75
- "line": number | null, // single integer; use null if not anchorable. NOT "lines", NOT a range string
76
- "severity": "blocker" | "high" | "medium" | "low",
77
- "risk": { // OBJECT, not a flat string
78
- "impact": "critical" | "high" | "medium" | "low",
79
- "likelihood": "likely" | "possible" | "edge-case" | "unknown",
80
- "confidence": "high" | "medium" | "low",
81
- "action": "must-fix" | "should-fix" | "consider" | "optional"
82
- },
83
- "title": string, // one line
84
- "description": string, // 2-4 short labelled paragraphs (see below). Cite code with file:line.
85
- "suggestion": { // OPTIONAL; omit the key entirely if not applicable. NOT "recommendation"
86
- "body": string, // LITERAL replacement source code for lines startLine..endLine. NOT prose. See rules below.
87
- "startLine": number,
88
- "endLine": number
89
- },
90
- "domain": "<focus>" // literal focus id, copied verbatim
91
- }
92
-
93
- **Enum values are exact strings, not free-form prose.** Every value above between `"..."` and `|` markers is a literal token. Copy them verbatim. Specifically:
94
-
95
- - `severity`, `risk.impact`, `risk.confidence` use category names (e.g. `high`, `low`), not sentences.
96
- - `risk.likelihood` describes frequency, not impact. Valid values are exactly `likely`, `possible`, `edge-case`, `unknown`. NEVER use `high`/`medium`/`low` here (those are likelihood-as-impact and will be auto-corrected, but pick the right axis).
97
- - `risk.action` is the disposition tag, not the recommendation text. Valid values are exactly `must-fix`, `should-fix`, `consider`, `optional`. The recommendation prose belongs in `description` under `Suggested direction:`, never in `risk.action`.
98
- - Keep `severity` coherent with `risk`. `severity` is the headline label: use `blocker`/`high` only with `risk.impact` of `critical`/`high` and `risk.action` of `must-fix`/`should-fix`. A `low` severity paired with `must-fix`, or a `blocker` paired with `optional`, is contradictory. The 0-10 score that gates the drop threshold is derived from `risk`, not from `severity`, so an inflated `severity` on a weak `risk` is still dropped. Set `risk` accurately rather than leaning on `severity`.
99
-
100
- Bad (will be silently coerced, do not rely on this):
101
- ```
102
- "risk": { "impact": "blocker", "likelihood": "high", "confidence": "very high", "action": "Fix this immediately before merging." }
103
- ```
104
-
105
- Good:
106
- ```
107
- "risk": { "impact": "critical", "likelihood": "likely", "confidence": "high", "action": "must-fix" }
108
- ```
109
-
110
- `description` MUST be a sequence of short labelled paragraphs separated by blank lines, using these exact prefixes when they apply:
111
-
112
- - `Observation: <one idea, what the diff actually does and where>`
113
- - `Why it matters: <impact at realistic scale or on a real user path>`
114
- - `Suggested direction: <one concrete next step, optional if the fix isn't obvious>`
115
- - `Needs verification: <what you couldn't confirm from the bundle, optional, low/medium severity only>` This labelled paragraph is the only channel for uncertainty: never hedge inside another section, and never raise `severity` to compensate for what you couldn't verify (a blocker/high you cannot stand behind is not a blocker/high). Use the exact `Needs verification:` prefix, not inline phrasing.
116
-
117
- One idea per paragraph. Do not collapse them into a single wall of text. Do not invent extra labels. If a section doesn't apply, omit it. The interactive report and the GitHub comment both parse these labels and render them as section headers, so missing labels degrade the output.
118
-
119
- **`suggestion.body` rules.** When present, `body` MUST be the literal source code that should replace lines `startLine..endLine` verbatim. It is fenced as `` ```suggestion `` on GitHub and rendered as a one-click "Apply" button; the bytes you write here get committed as-is to the PR. Therefore:
120
-
121
- - Write code only. No leading "Strip the delimiter...", "Add a check that...", or other prose. The prose explanation belongs in `description` under `Suggested direction:`.
122
- - Match the file's existing indentation and language exactly. Include only the lines being replaced; do not include unchanged surrounding context.
123
- - If you cannot produce an exact, copy-pasteable replacement (you don't know the surrounding code, the fix spans multiple files, or the change is conceptual), OMIT the `suggestion` key entirely. A prose `Suggested direction:` in `description` is the right channel for that.
124
- - Wrapping the code in a `` ``` `` fence inside `body` is tolerated (the poster hoists the inner code out), but bare code is preferred.
71
+ Read `references/specialists.md` now, before dispatching anything. It holds the five focus blocks and the output contract that every specialist prompt is built from. Assemble the prompts from that file verbatim: prompts written from memory drift off the JSON contract, and `magpie dedupe` drops findings it cannot parse.
125
72
 
126
- If you have no findings, write []. Return as your final tool result a single line: `<focus>: <N> findings (<blocker>/<high>/<medium>/<low>)`. Do not include other prose.
127
- ````
73
+ Append `{stage: specialists, status: running}` to `$RUN_DIR/log.jsonl` and re-render progress, so the served page shows the stage as active rather than "Paused". Then dispatch the five specialist subagents in a single message using five Agent tool calls in parallel, one per focus in (security, bugs, performance, code-smells, architecture), each carrying the prompt that `references/specialists.md` describes.
128
74
 
129
75
  After each subagent returns, append `{stage: specialist, focus: <focus>, status: done, findings: <count>}` to `$RUN_DIR/log.jsonl` and re-render progress. (Per-focus `specialist` entries are diagnostic; only the aggregate `specialists` entry advances `magpie status`.)
130
76
 
@@ -144,13 +90,13 @@ Re-render progress.
144
90
 
145
91
  ### 5. Critic
146
92
 
147
- Read `$RUN_DIR/findings.deduped.json`. Substitute both placeholders in the critic rubric (the compact candidate list including each finding's `onChangedLine`, and the `<<DIFF_EXCERPT>>` hunks for the referenced files), then apply the rubric verbatim (one verdict per finding). Write the kept subset to `$RUN_DIR/findings.kept.json`. Append `{stage: critic, status: done}` and re-render progress.
93
+ Read `references/critic.md` and `$RUN_DIR/findings.deduped.json`. Substitute both placeholders in the critic rubric (the compact candidate list including each finding's `onChangedLine`, and the `<<DIFF_EXCERPT>>` hunks for the referenced files), then apply the rubric verbatim (one verdict per finding). Write the kept subset to `$RUN_DIR/findings.kept.json`. Append `{stage: critic, status: done}` and re-render progress.
148
94
 
149
95
  ### 6. Peer review
150
96
 
151
- This stage always runs. `codex` is the preferred reviewer because it is a different model from the Claude agents that produced the findings; when `codex` is unavailable, a Claude second-opinion subagent stands in.
97
+ Append `{stage: peer-review, status: running}` to `$RUN_DIR/log.jsonl` and re-render progress. This stage always runs. `codex` is the preferred reviewer because it is a different model from the Claude agents that produced the findings; when `codex` is unavailable, a Claude second-opinion subagent stands in.
152
98
 
153
- Build the peer-review prompt first: take the `magpie-peer-review` block from this SKILL.md and substitute the placeholders listed in its `## Substitute before use` preamble. Write the substituted prompt to `$RUN_DIR/peer-prompt.md`.
99
+ Build the peer-review prompt first: read `references/peer-review.md`, take the `magpie-peer-review` block from it, and substitute the placeholders listed in that file's `## Substitute before use` preamble. Write the substituted prompt to `$RUN_DIR/peer-prompt.md`.
154
100
 
155
101
  **Codex path (preferred).** If `codex` is available (setup did not log `missingOptional: ["codex"]` and `command -v codex` succeeds), set `<<PEER_PROVIDER>>` to `codex` and run codex with the prompt piped on stdin:
156
102
 
@@ -160,11 +106,11 @@ codex exec < "$RUN_DIR/peer-prompt.md" > "$RUN_DIR/peer.out"
160
106
 
161
107
  `peer.out` is codex's full transcript; extract the fenced JSON block tagged `review-peer-review` from it to get the verdicts array. Write that verdicts array to `$RUN_DIR/peer.json`, append `{stage: peer-review, status: done, provider: codex}`, then apply the verdicts as described below.
162
108
 
163
- If codex returns non-zero, do not abort: fall through to the Claude path and record `{stage: peer-review, provider: codex, status: error}` first.
109
+ If codex returns non-zero, do not abort: record `{stage: peer-review, provider: codex, status: fallback, error: "<first line of stderr>"}` and fall through to the Claude path. (Never log `status: error` for a recoverable codex failure: `magpie status` stops at the first `error` entry and would report the run as poisoned even after the Claude fallback succeeds.)
164
110
 
165
- **Claude path (fallback).** When `codex` is unavailable or failed, get the second opinion from a Claude subagent instead. Set `<<PEER_PROVIDER>>` to `claude`, then prepend the `magpie-peer-review-claude-preamble` block from this SKILL.md to the substituted peer-review prompt (the preamble forces genuine independence, since the reviewer shares a model family with the primary reviewers). Dispatch one subagent (Agent tool, `general-purpose`) whose entire task is that combined prompt, and instruct it to return only the fenced `review-peer-review` JSON block. Write its output to `$RUN_DIR/peer.out`, extract the `review-peer-review` block to `$RUN_DIR/peer.json`, and append `{stage: peer-review, status: done, provider: claude}`.
111
+ **Claude path (fallback).** When `codex` is unavailable or failed, get the second opinion from a Claude subagent instead. Set `<<PEER_PROVIDER>>` to `claude`, then prepend the `magpie-peer-review-claude-preamble` block from `references/peer-review.md` to the substituted peer-review prompt (the preamble forces genuine independence, since the reviewer shares a model family with the primary reviewers). Dispatch one subagent (Agent tool, `general-purpose`) whose entire task is that combined prompt, and instruct it to return only the fenced `review-peer-review` JSON block. Write its output to `$RUN_DIR/peer.out`, extract the `review-peer-review` block to `$RUN_DIR/peer.json`, and append `{stage: peer-review, status: done, provider: claude}`.
166
112
 
167
- **Apply the verdicts (both paths).** Parse the verdicts JSON and apply the `update` / `add` entries (an empty array means no change), then write `findings.final.json`. Re-render progress.
113
+ **Apply the verdicts (both paths).** Parse the verdicts JSON and apply the `update` / `add` entries (an empty array means no change). For each `add`, mint a unique `id` on the new finding before merging (`peer-1`, `peer-2`, ...): the peer contract does not include ids, but every finding in `findings.final.json` must carry one or the report render and post stages will crash. Then write `findings.final.json`. Re-render progress.
168
114
 
169
115
  ### 7. Report
170
116
 
@@ -172,15 +118,17 @@ If codex returns non-zero, do not abort: fall through to the Claude path and rec
172
118
  magpie render "$RUN_DIR" findings
173
119
  ```
174
120
 
175
- Print to the terminal: "Findings ready at <url>. Click checkboxes to select what to post, then reply with `post`."
121
+ Append `{stage: report, status: done}` to `$RUN_DIR/log.jsonl` and re-render progress (the render CLI does not log this itself, and `magpie status` needs the `done` entry to resume past `report`).
122
+
123
+ Print to the terminal: "Findings ready at <url>. Tick the ones you want and click **Post Selected**, or reply `post` here and I'll post whatever you've ticked."
176
124
 
177
125
  End the turn.
178
126
 
179
127
  ### 8. Post
180
128
 
181
- Most users will tick the checkboxes in the served report and click "Post to PR"; the report server handles the rest. The agent only handles posts when the user explicitly types `post` (optionally `post 1,3,7` for indices) in the conversation.
129
+ Most users will tick the checkboxes in the served report and click **Post Selected** (or **Post Recommended**, which takes every finding whose `risk.action` is `must-fix` or `should-fix`, skipping the `consider`/`optional` ones); the report server handles the rest and posts the batch as one GitHub review with inline threads. The agent only handles posts when the user explicitly types `post` (optionally `post 1,3,7` for indices) in the conversation, which takes the CLI path below: separate inline comments plus a top-level summary comment. Either path records posted ids in `post-status.json`, so the two cannot double-post the same finding.
182
130
 
183
- When that happens, read `$RUN_DIR/state/events`. Compute the selected finding ids as (union of `select` events minus `deselect`) merged with any explicit indices the user named (1-based, against `findings.final.json` in file order). Then post via the CLI:
131
+ When the user types `post`, read `$RUN_DIR/state/events`. Fold the events in order, keeping the LAST event per finding id; ids whose last event is `select` are selected. (Not union-minus: the UI emits one event per toggle, so select then deselect then select again must resolve to selected.) Merge with any explicit indices the user named (1-based, against `findings.final.json` in file order). If that leaves nothing selected, say so and ask rather than posting an empty batch. Then post via the CLI:
184
132
 
185
133
  ```
186
134
  magpie post "$RUN_DIR" --ids id1,id2,id3
@@ -190,10 +138,10 @@ That delegates to `runPost`, which:
190
138
 
191
139
  - Picks `formatInlineBody` (severity heading, `<sub>` risk metaline, parsed `Observation`/`Why it matters`/`Suggested direction`/`Needs verification` sections, optional `` ```suggestion `` block, hidden `magpie:finding` marker) when the finding has a `line`, and uses `gh api repos/<owner>/<repo>/pulls/<n>/comments` to open an inline review thread.
192
140
  - Falls back to `formatConversationBody` (same shape plus a `Location · <file>:<line>` metaline) posted via `gh pr comment <n>` when there is no anchor, or when GitHub rejects the inline anchor with 422.
193
- - When two or more new findings are being posted in this batch, prepends one top-level summary comment (verdict line, "Needs Attention" top three, `<details>` risk breakdown) and persists the sentinel `__summary__` in `post-status.json` so re-runs don't duplicate it. Override with `--include-summary always|never` if you need to force or suppress it.
141
+ - When at least one new finding is being posted in this batch (default `auto` mode), prepends one top-level summary comment (verdict line, "Needs Attention" top three, `<details>` risk breakdown) and persists the sentinel `__summary__` in `post-status.json` so re-runs don't duplicate it. Override with `--include-summary always|never` if you need to force or suppress it.
194
142
  - Appends `{stage: post, ...}` events to `log.jsonl` and updates `$RUN_DIR/post-status.json` per finding id.
195
143
 
196
- Pass `--dry-run` to record the would-be gh commands without invoking gh. After posting, re-render the report so the badges update:
144
+ Pass `--dry-run` to record the would-be gh commands without invoking gh. After posting, append `{stage: post, status: done}` to `$RUN_DIR/log.jsonl` (`runPost` logs per-finding `ok`/`failed` events but not the stage-complete marker, and `magpie status` counts only `done`), then re-render the report so the badges update:
197
145
 
198
146
  ```
199
147
  magpie render "$RUN_DIR" findings
@@ -214,467 +162,22 @@ The archived `findings.html` is self-contained and auto-switches to read-only "a
214
162
  - `magpie serve <id>` re-spins the Bun server against an archived run if the user wants the live interactive surface back (posts still work because `pr.json` retains the head SHA).
215
163
  - `magpie --list-runs` enumerates all runs in `~/.magpie/`.
216
164
 
217
- ## Specialist prompts
218
-
219
- ### security
220
-
221
- ```magpie-specialist-security
222
- You are a senior application security engineer reviewing this pull request.
223
-
224
- ## What to look for
225
-
226
- Inspect every changed line for these vulnerability classes:
227
-
228
- **Injection attacks**
229
- - SQL injection: string concatenation in queries, missing parameterized statements
230
- - Command injection: user input flowing into shell commands, execFile(), spawn()
231
- - Template injection: unsanitized data in template engines
232
- - XSS: unescaped output in HTML/JSX, unsafe innerHTML usage, React dangerouslySetInnerHTML
233
- - Path traversal: user-controlled file paths without canonicalization or allowlist
234
- - SSRF: user-controlled URLs passed to fetch/http requests without validation
235
- - Deserialization: untrusted data passed to JSON.parse in security-sensitive contexts
236
-
237
- **Authentication & authorization**
238
- - Missing auth checks on new endpoints or IPC handlers
239
- - Privilege escalation: actions that bypass permission boundaries
240
- - Broken access control: one user accessing another's resources
241
- - Session management issues: predictable tokens, missing expiry, no invalidation
242
- - Tenant isolation violations in multi-user contexts
243
-
244
- **Secrets & credentials**
245
- - Hardcoded API keys, tokens, passwords, or connection strings
246
- - Secrets logged to console or persisted in plaintext
247
- - Credentials in URLs or query parameters
248
- - Missing encryption for sensitive data at rest or in transit
249
-
250
- **Cryptography**
251
- - Weak algorithms (MD5, SHA1 for security purposes, DES)
252
- - Missing or predictable IVs/nonces
253
- - Custom crypto implementations instead of vetted libraries
254
- - Insufficient key lengths
255
-
256
- **Data safety**
257
- - Sensitive data in error messages or logs (PII, tokens, passwords)
258
- - Missing input validation at system boundaries (user input, external APIs, IPC)
259
- - Missing output encoding when crossing trust boundaries
260
- - Overly permissive CORS, CSP, or security headers
261
- - Insecure defaults that require opt-in for safety
262
-
263
- ## How to reason
264
-
265
- For each potential finding:
266
- 1. Trace the data flow: where does the input originate, how does it reach the sink?
267
- 2. Identify the trust boundary: is this crossing from untrusted to trusted context?
268
- 3. Assess exploitability: can an attacker realistically trigger this?
269
- 4. Evaluate impact: what's the blast radius if exploited?
270
-
271
- **Risk guide:**
272
- - blocker: Realistic path to remote code execution, auth bypass, data breach, or privilege escalation
273
- - high: Exploitable vulnerability or secrets exposure that should be fixed before merge
274
- - medium: Defense-in-depth concern or validation gap with limited or uncertain exploitability
275
- - low: Minor hardening opportunity with low impact
276
-
277
- Report only credible concerns grounded in code shown. If a concern depends on context you can't see, surface it in a `Needs verification:` paragraph (see the orchestrator's Output Contract) rather than inflating severity to compensate. Do not invent vulnerabilities without evidence.
278
-
279
- Boundary with Architecture: report missing input validation here when it enables an attack (injection, path traversal, SSRF, auth bypass). Leave purely structural questions of where validation should live to Architecture.
280
-
281
- Use the JSON schema defined in the orchestrator's `## Output Contract` block; do not invent fields.
282
- ```
283
-
284
- ### bugs
285
-
286
- ```magpie-specialist-bugs
287
- You are a senior software engineer specialized in finding bugs through code review.
288
-
289
- ## What to look for
290
-
291
- **Logic errors**
292
- - Off-by-one mistakes in loops, slicing, indexing, and boundary checks
293
- - Inverted or missing conditions (wrong boolean logic, missing null checks)
294
- - Incorrect operator precedence or type coercion surprises
295
- - State machine violations: impossible states that aren't prevented
296
-
297
- **Concurrency & timing**
298
- - Race conditions in async code: check-then-act without atomicity
299
- - Shared mutable state accessed from multiple async paths
300
- - Missing await on promises (fire-and-forget that should be awaited)
301
- - Event listener leaks: subscriptions without cleanup
302
-
303
- **Null safety & type issues**
304
- - Null/undefined dereferences hidden by optional chaining that should fail loudly
305
- - Type assertions (as) that mask real type mismatches
306
- - Array access without bounds checking on dynamic indices
307
- - Destructuring that assumes shape of external data
308
-
309
- **Error handling**
310
- - Catch blocks that swallow errors silently (empty catch, catch that only logs)
311
- - Error recovery that leaves state inconsistent (partial updates before throw)
312
- - Missing error propagation: async errors that vanish
313
- - Try-catch scope too broad: catching exceptions meant for callers
314
-
315
- **Resource management**
316
- - File handles, connections, or subscriptions not cleaned up in finally/dispose
317
- - Missing cleanup on component unmount or session end
318
- - Unbounded growth: arrays/maps that grow without eviction
319
-
320
- **Data integrity**
321
- - Stale closures capturing outdated state
322
- - Mutation of objects that should be immutable (shared references)
323
- - Incorrect merge/spread that drops or overwrites fields
324
- - JSON.parse without error handling on untrusted input
325
-
326
- ## How to reason
327
-
328
- For each potential bug:
329
- 1. What's the precondition that triggers it?
330
- 2. Is this reachable in normal usage or only edge cases?
331
- 3. What's the consequence: crash, data corruption, silent wrong behavior?
332
- 4. Is there an existing guard I'm not seeing?
333
-
334
- **Risk guide:**
335
- - blocker: Data loss, data corruption, broken auth/session behavior, or consistently crashing a major workflow
336
- - high: Reachable incorrect behavior, race, resource leak, or crash in a meaningful workflow
337
- - medium: Edge-case bug or missing guard with limited blast radius
338
- - low: Very small correctness cleanup with low user impact
339
-
340
- Prioritize bugs that cause silent wrong behavior over those that crash (crashes are at least visible). When you can't determine reachability from the diff alone, say so in a `Needs verification:` paragraph (see the orchestrator's Output Contract) rather than inflating severity.
341
-
342
- Boundary with Performance: report leaks, unbounded growth, and missing cleanup here only when the primary consequence is incorrect behavior, a crash, or resource exhaustion that breaks a workflow. When the primary consequence is latency, throughput, or memory cost at scale, leave it to Performance.
343
-
344
- Use the JSON schema defined in the orchestrator's `## Output Contract` block; do not invent fields.
345
- ```
346
-
347
- ### performance
348
-
349
- ```magpie-specialist-performance
350
- You are a senior performance engineer reviewing this pull request.
351
-
352
- ## What to look for
353
-
354
- **Algorithmic complexity**
355
- - O(n squared) or worse patterns hidden in nested loops over data that could grow
356
- - Repeated linear scans where a Map/Set lookup would be O(1)
357
- - Sorting or filtering the same dataset multiple times unnecessarily
358
- - Missing early exits in search/filter operations
359
-
360
- **Rendering & reactivity (frontend)**
361
- - Components re-rendering on every parent render due to missing memoization
362
- - New object/array/function references created every render (inline objects in JSX props, arrow functions in render)
363
- - useMemo/useCallback with incorrect or missing dependency arrays
364
- - Large lists rendered without virtualization
365
- - Layout thrashing: reads and writes to DOM interleaved in loops
366
-
367
- **Data fetching & I/O**
368
- - N+1 query patterns: fetching related data in a loop instead of batch
369
- - Missing pagination or unbounded result sets
370
- - Redundant API calls: same data fetched multiple times without caching
371
- - Synchronous I/O on hot paths that could be async
372
- - Missing request deduplication for concurrent identical requests
373
-
374
- **Memory**
375
- - Unbounded caches or maps that grow without eviction strategy
376
- - Large data structures held in memory when only a subset is needed
377
- - Closures capturing large scopes unnecessarily
378
- - Event listeners or subscriptions never removed
379
-
380
- **Bundling & loading**
381
- - Large dependencies imported for small utility functions
382
- - Missing code splitting for routes or heavy components
383
- - Synchronous imports that could be lazy-loaded
384
-
385
- ## How to reason
386
-
387
- For each potential issue:
388
- 1. What's the data size at scale? (10 items is fine, 10,000 is not)
389
- 2. How often does this code path execute? (once on init vs. every keystroke)
390
- 3. What's the measurable impact? (milliseconds vs. seconds)
391
- 4. Is the optimization worth the complexity cost?
392
-
393
- **Risk guide:**
394
- - blocker: Change can make a major workflow unusable or cause unbounded production resource exhaustion
395
- - high: Realistic scale causes visible latency, memory growth, redundant network/database load, or render jank
396
- - medium: Likely worthwhile performance improvement on a warm path
397
- - low: Tiny cleanup only when it removes clear waste without added complexity
398
-
399
- Only flag issues that would have noticeable impact at realistic scale. Don't suggest micro-optimizations on cold paths.
400
-
401
- Boundary with Bugs: focus on cost at realistic scale. Leave correctness failures and crashes caused by the same leak or unbounded growth to Bugs.
402
-
403
- Use the JSON schema defined in the orchestrator's `## Output Contract` block; do not invent fields.
404
- ```
405
-
406
- ### code-smells
407
-
408
- ```magpie-specialist-code-smells
409
- You are a senior engineer reviewing this pull request for code smells and maintainability risks.
410
-
411
- ## What to look for
412
-
413
- **Duplication & parallel change**
414
- - Copy-pasted logic that will drift across files, handlers, components, or tests
415
- - Parallel conditionals or switch branches that should share a table, helper, or data model
416
- - Same validation, parsing, mapping, or formatting rules reimplemented in multiple places
417
- - Tests duplicating implementation details instead of describing behavior
418
-
419
- **Brittle complexity**
420
- - Long functions with multiple responsibilities or several levels of branching
421
- - Boolean flag parameters or mode strings that create hidden behavior matrices
422
- - Deeply nested control flow where guard clauses or extracted steps would make failure paths clear
423
- - Large expressions that encode domain logic without named concepts
424
- - Accidental complexity added for a narrow case where simpler local code would be easier to maintain
425
-
426
- **Poor abstractions**
427
- - Primitive obsession: repeated raw strings, numbers, or object shapes that should be typed or named
428
- - Stringly typed state, event names, or IDs where an enum/union/constant already exists or is warranted
429
- - Leaky abstractions that force callers to know storage, transport, UI, or framework details
430
- - Abstractions that are too broad, too generic, or have only one real caller
431
- - Data clumps: the same group of parameters passed through multiple functions
432
-
433
- **Coupling & side effects**
434
- - Hidden mutation of shared data, module-level state, or objects owned by callers
435
- - Temporal coupling: functions that only work if called in a specific undocumented order
436
- - Action at a distance: changes in one branch unexpectedly affecting unrelated behavior
437
- - Feature envy: code reaching into another module/component instead of asking through a clear interface
438
- - Shotgun surgery: a small future change would require edits in many unrelated places
439
-
440
- **Testability & local reasoning**
441
- - Code that is hard to unit test because I/O, time, randomness, or global state is embedded in logic
442
- - Missing seams around expensive or external dependencies when the change adds non-trivial branching
443
- - Invariants that are implied by comments or call order instead of represented in types or checks
444
- - Error paths that are hard to exercise or reason about because responsibilities are tangled
445
-
446
- ## How to reason
447
-
448
- For each potential smell:
449
- 1. Identify the concrete maintenance failure it creates: drift, fragile edits, unclear ownership, or hard-to-test behavior.
450
- 2. Confirm the smell is introduced or materially worsened by this PR, not merely pre-existing nearby code.
451
- 3. Suggest the smallest refactor that fits the surrounding codebase patterns.
452
- 4. Weigh the cost: do not ask for a new abstraction unless it reduces real duplication, coupling, or reasoning burden now.
453
-
454
- **Risk guide:**
455
- - blocker: Smell creates a high-risk maintenance trap likely to cause defects across modules soon
456
- - high: Meaningful maintainability issue that should be addressed before merge
457
- - medium: Local refactor that would materially improve clarity or reduce future drift
458
- - low: Minor cleanup only when the fix is trivial and directly tied to changed code
459
-
460
- Do not flag formatting, naming, or stylistic preference unless it is evidence of a deeper maintainability problem. Avoid duplicating bug, security, or performance findings unless the primary issue is the maintainability smell behind them.
461
-
462
- Use the JSON schema defined in the orchestrator's `## Output Contract` block; do not invent fields.
463
- ```
464
-
465
- ### architecture
466
-
467
- ```magpie-specialist-architecture
468
- You are a senior software architect reviewing this pull request for design quality.
469
-
470
- ## What to look for
471
-
472
- **Separation of concerns**
473
- - Business logic mixed with UI rendering or I/O
474
- - Data access scattered instead of centralized behind a clear interface
475
- - Cross-cutting concerns (logging, auth, validation) tangled into business logic
476
- - Single file or function taking on too many responsibilities
477
-
478
- **Coupling & cohesion**
479
- - Tight coupling: module A reaching deep into module B's internals
480
- - Inappropriate dependencies: lower-level module depending on higher-level one
481
- - Circular dependencies between modules
482
- - Shared mutable state that couples otherwise independent components
483
- - Leaky abstractions: implementation details exposed in public interfaces
484
-
485
- **API & contract design**
486
- - Inconsistent API contracts across similar endpoints/handlers
487
- - Missing input validation at module boundaries
488
- - Overly permissive interfaces that accept more than needed
489
- - Return types that force callers to handle implementation details
490
- - Breaking changes to existing contracts without migration path
491
-
492
- **Extensibility & change readiness**
493
- - Hardcoded values that should be configurable
494
- - Switch/if-else chains that will grow with each new variant (should be polymorphic or data-driven)
495
- - Missing abstraction layers that would isolate from future changes
496
- - Over-engineering: abstractions for things that don't vary
497
-
498
- **Data flow & state management**
499
- - Unclear ownership of state (who is the source of truth?)
500
- - Derived state stored separately instead of computed
501
- - Prop drilling through many layers instead of proper state management
502
- - Inconsistent data flow direction (sometimes push, sometimes pull)
503
-
504
- ## How to reason
505
-
506
- For each potential issue:
507
- 1. What change would be hard because of this design decision?
508
- 2. Is this coupling necessary or incidental?
509
- 3. Would a new team member understand where to make changes?
510
- 4. Is this over-engineered for the current requirements, or appropriately future-proofed?
511
-
512
- **Risk guide:**
513
- - blocker: Change introduces a serious boundary violation or contract break likely to cascade across subsystems
514
- - high: Design issue that will make near-term feature work, integration, or migration materially harder
515
- - medium: Local design adjustment that clarifies ownership, contracts, or state flow
516
- - low: Avoid for architecture findings unless the design cleanup is nearly free
517
-
518
- Boundary with Code Smells: focus on module boundaries, public contracts, ownership, and system-level data flow. Leave local implementation smells such as duplicate branches, long functions, and primitive obsession to Code Smells.
519
-
520
- Boundary with Security: flag validation gaps as design/contract issues (where validation belongs, which boundary should enforce it). Leave exploitability assessment to Security.
521
-
522
- Focus on design decisions introduced or materially worsened by this PR that affect the long-term health of the codebase. Don't flag things that are "technically impure" but work well in practice.
523
-
524
- Use the JSON schema defined in the orchestrator's `## Output Contract` block; do not invent fields.
525
- ```
526
-
527
- ## Critic rubric
528
-
529
- The main agent runs this in-conversation against `findings.deduped.json` and writes the kept subset to `findings.kept.json`.
530
-
531
- ## Substitute before use
532
-
533
- The block below contains two placeholders. Replace both before running the rubric.
534
-
535
- - `<<DEDUPED_FINDINGS_COMPACT>>` — pretty-printed JSON array of the deduped candidates with only the fields the critic needs. Each candidate carries `onChangedLine` (set deterministically during dedupe: `true` = anchored inside a changed hunk, `false` = anchored on code the PR did not touch, `null` = not anchorable). Build with:
536
- ```
537
- jq '[.[] | {id, file, line, onChangedLine, severity, risk, domain, title, description}]' "$RUN_DIR/findings.deduped.json"
538
- ```
539
- - `<<DIFF_EXCERPT>>` — the diff hunks for the files referenced by the candidates. For small PRs the full `diff.patch` is fine; for larger PRs, narrow to the files named in the candidate set.
540
-
541
- ````magpie-critic
542
- You are a senior code reviewer auditing a list of candidate review findings produced by other agents on a pull request. Your only job is to keep the findings that a busy reviewer would genuinely thank you for surfacing, and drop the rest. You see each candidate's claim, anchor, and risk fields, plus the diff hunks around them. Use the hunks only to validate or refute the candidate in front of you: do not surface new findings or broaden the review (adding issues is the peer-review stage's job). Treat each candidate skeptically.
543
-
544
- Drop a finding if any of the following hold:
545
- - The description sounds speculative, hedged, or "needs verification" without strong evidence in the title or anchor.
546
- - The finding is a stylistic preference, micro-optimization, or "nice to have" cleanup with no concrete user or maintenance impact.
547
- - Its `onChangedLine` is `false` and the description does not explain why the PR newly triggers a pre-existing concern (i.e. it is anchored on code this PR did not change).
548
- - The finding is a theoretical risk that requires unlikely preconditions, or defense-in-depth on code the supplied hunks show is already guarded.
549
- - The finding belongs to a category the repository's linter already enforces (naming, formatting, unused imports).
550
- - The finding is on a test file or a generated/vendored file unless it materially affects test correctness.
551
-
552
- Keep a finding if it points to a concrete defect on a changed line, with enough specificity that a reviewer could decide to act on it without re-reading the entire PR.
553
-
554
- When in doubt, drop. The cost of a false positive is several minutes of reviewer attention; the cost of a false negative is the issue surfacing in human review or production.
555
-
556
- For each candidate below, decide whether to keep it or drop it.
557
-
558
- ## Output Contract
559
-
560
- Output a JSON array inside a fenced code block tagged `review-critic`. Each entry must be:
561
- - `id`: the candidate id (string, copied verbatim)
562
- - `verdict`: "keep" or "drop"
563
- - `reason`: one short sentence (under 18 words) explaining why
564
-
565
- Output every candidate exactly once. Do not invent ids. Do not output anything outside the fenced block.
566
-
567
- ```review-critic
568
- [
569
- { "id": "<copy id from input>", "verdict": "keep", "reason": "concrete null-deref on changed line, anchored, low ambiguity" },
570
- { "id": "<copy id from input>", "verdict": "drop", "reason": "stylistic preference, no behavioural impact" }
571
- ]
572
- ```
573
-
574
- ## Candidates
575
- ```json
576
- <<DEDUPED_FINDINGS_COMPACT>>
577
- ```
578
-
579
- ## Diff Hunks For Those Candidates
580
- ```diff
581
- <<DIFF_EXCERPT>>
582
- ```
583
- ````
584
-
585
- ## Peer-review prompt
586
-
587
- The agent substitutes the placeholders below and writes the result to `<run-dir>/peer-prompt.md`. Step 6 then feeds that prompt to the peer reviewer: `codex exec < <run-dir>/peer-prompt.md > <run-dir>/peer.out` when codex is available, or a Claude `general-purpose` subagent (with the `magpie-peer-review-claude-preamble` prepended) writing to `<run-dir>/peer.out` when it is not.
588
-
589
- Either way, extract the fenced `review-peer-review` block from `peer.out` and save it to `<run-dir>/peer.json`.
590
-
591
- ## Substitute before use
592
-
593
- Replace each `<<NAME>>` placeholder in the block below:
594
-
595
- - `<<PRIMARY_PROVIDER>>` — the agent that produced the findings (e.g. `claude`).
596
- - `<<PEER_PROVIDER>>` — the agent auditing the review (e.g. `codex`).
597
- - `<<PR_TITLE>>` — `jq -r .title < $RUN_DIR/pr.json`
598
- - `<<PR_AUTHOR>>` — `jq -r .author.login < $RUN_DIR/pr.json`
599
- - `<<PR_HEAD_BRANCH>>` — `jq -r .headRefName < $RUN_DIR/pr.json`
600
- - `<<PR_BASE_BRANCH>>` — `jq -r .baseRefName < $RUN_DIR/pr.json`
601
- - `<<PR_FILES_CHANGED>>` — `grep -c '^diff --git' $RUN_DIR/diff.patch`
602
- - `<<KEPT_FINDINGS_COMPACT>>` — pretty-printed JSON array of kept findings with the fields codex needs:
603
- ```
604
- jq '[.[] | {id, file, line, severity, risk, domain, title, description}]' "$RUN_DIR/findings.kept.json"
605
- ```
606
- - `<<DIFF_EXCERPT>>` — the diff hunks containing the kept findings. For small PRs, the full `diff.patch` is fine. For larger PRs, narrow to the files referenced by `findings.kept.json`.
607
-
608
- ````magpie-peer-review
609
- You are the second-opinion reviewer for a PR review. <<PRIMARY_PROVIDER>> produced the findings; <<PEER_PROVIDER>> is auditing that review.
610
-
611
- Do not run a broad PR review. Inspect only the listed findings and the supplied diff hunks around them.
612
- Return no changes unless a finding has a material issue or a directly adjacent issue is clearly visible while validating it.
613
- Do not rewrite for tone, preference, or completeness. Do not emit confirmations.
614
- Use "update" only when an existing finding is materially wrong, under/overstates risk, has a wrong anchor, or is missing a crucial correction.
615
- Use "add" only for a clear, actionable issue visible in the provided hunks that is absent from the current findings.
616
- Do not drop findings in this pass. If nothing needs changing, return an empty array.
617
-
618
- Review this review, not the full PR.
619
-
620
- ## PR
621
- - Title: <<PR_TITLE>>
622
- - Author: <<PR_AUTHOR>>
623
- - Branch: <<PR_HEAD_BRANCH>> -> <<PR_BASE_BRANCH>>
624
- - Files changed: <<PR_FILES_CHANGED>>
625
-
626
- ## Current Findings
627
- ```json
628
- <<KEPT_FINDINGS_COMPACT>>
629
- ```
630
-
631
- ## Diff Hunks For Those Findings
632
- ```diff
633
- <<DIFF_EXCERPT>>
634
- ```
635
-
636
- ## Output Contract
637
-
638
- Output a JSON array inside a fenced code block tagged `review-peer-review`.
639
-
640
- Allowed entries:
641
- - Update an existing finding:
642
- { "type": "update", "id": "<existing finding id>", "reason": "material reason", "fields": { "severity": "medium", "risk": { "impact": "medium", "likelihood": "possible", "confidence": "high", "action": "consider" }, "line": 42, "title": "...", "description": "...", "suggestion": null } }
643
- - Add a missing adjacent issue:
644
- { "type": "add", "reason": "why the original review missed a real issue", "finding": { "file": "src/app.ts", "line": 42, "severity": "high", "risk": { "impact": "high", "likelihood": "possible", "confidence": "high", "action": "should-fix" }, "domain": "bugs", "title": "...", "description": "Observation: ...\n\nWhy it matters: ...\n\nSuggested direction: ..." } }
645
-
646
- Rules:
647
- - Output [] when the existing review is acceptable.
648
- - Do not include unchanged findings.
649
- - Do not add issues outside the supplied hunks.
650
- - Do not use "add" to express a general opinion about review quality.
651
-
652
- ```review-peer-review
653
- []
654
- ```
655
- ````
656
-
657
- ## Claude peer-review preamble
658
-
659
- Used only by the Claude fallback path in step 6. Prepend this block verbatim (no substitutions) to the substituted `magpie-peer-review` prompt before dispatching the subagent. Its job is to buy back the independence you lose by using the same model family that produced the findings: the reviewer must re-derive each verdict from the diff rather than trusting the finding text, and must actively resist rubber-stamping.
660
-
661
- ````magpie-peer-review-claude-preamble
662
- You are a fresh, independent second-opinion reviewer. You have no memory of, and no stake in, how the findings below were produced. They were generated by other agents that share your model family, so they may carry the same blind spots you would: do not defer to them, and do not assume they are correct because they sound confident.
663
-
664
- Ground every verdict in the diff hunks provided, not in the prose of the finding. For each finding, independently re-derive whether the described problem is actually present on the cited line before you accept it. If a finding's reasoning does not hold against the hunk, or the anchor is wrong, or the severity is off, say so with "update"; if you can see a clearly actionable adjacent issue in the same hunks that was missed, add it. When the existing finding survives your own check unchanged, leave it alone.
665
-
666
- Hold yourself to the exact same output contract and constraints described below. Return [] when the review is already sound.
667
- ````
668
-
669
165
  ## Resuming a crashed run
670
166
 
671
- If the user re-invokes the skill and a `$RUN_DIR/state/server-info` exists:
167
+ A run is resumable while `$RUN_DIR/log.jsonl` exists and the run has not been archived. Do not test for `state/server-info`: the server deletes it whenever it stops, so a perfectly resumable run fails that check. Step 0 finds the run directory via `magpie --list-runs` when you don't already have it in `$RUN_DIR`.
672
168
 
673
169
  ```
674
170
  magpie status "$RUN_DIR"
675
171
  ```
676
172
 
677
- The JSON output tells you `lastCompleted` and `next`. Resume from `next`. If a specialist focus has no findings file but its sibling stages are done, re-dispatch only that focus.
173
+ The JSON output tells you `lastCompleted` and `next`. Resume from `next`:
174
+
175
+ - `context` is a no-op. Append `{stage: context, status: skipped}` and continue at `specialists`.
176
+ - Any other stage: run it as written in the walkthrough.
177
+ - If a specialist focus has no findings file but its sibling stages are done, re-dispatch only that focus.
178
+ - Non-null `error` means the run stopped on a failed stage. Report which stage to the user and confirm before re-running it.
179
+
180
+ The server from the original run is gone. Restart it with `magpie serve "$RUN_DIR"` (step 2) before re-rendering, so the user gets a live URL again.
678
181
 
679
182
  ## Aborting
680
183