@iceinvein/agent-skills 0.21.0 → 0.21.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/package.json +1 -1
- package/skills/index.json +1 -1
- package/skills/magpie/README.md +1 -2
- package/skills/magpie/SKILL.md +7 -18
- package/skills/magpie/evals/README.md +4 -5
- package/skills/magpie/evals/codex-missing-falls-back/fixture.sh +1 -1
- package/skills/magpie/evals/post-folds-selection-events/fixture.sh +1 -1
- package/skills/magpie/evals/report-ends-the-turn/fixture.sh +1 -1
- package/skills/magpie/evals/resume-finds-active-run/fixture.sh +1 -1
- package/skills/magpie/evals/shard-gate-stops-and-asks/fixture.sh +1 -1
- package/skills/magpie/fixtures/example-pr/brief.json +0 -4
- package/skills/magpie/references/scout.md +2 -42
- package/skills/magpie/references/specialists.md +8 -115
- package/skills/magpie/scripts/__tests__/refresh.test.ts +0 -1
- package/skills/magpie/scripts/__tests__/render-cmd.test.ts +0 -1
- package/skills/magpie/scripts/__tests__/render-findings.test.ts +4 -27
- package/skills/magpie/scripts/__tests__/render-progress.test.ts +1 -1
- package/skills/magpie/scripts/__tests__/skill-lint.test.ts +19 -171
- package/skills/magpie/scripts/__tests__/types.test.ts +16 -9
- package/skills/magpie/scripts/render-findings.ts +0 -7
- package/skills/magpie/scripts/render-progress.ts +1 -1
- package/skills/magpie/scripts/types.ts +0 -16
- package/skills/magpie/skill.json +1 -1
- package/skills/magpie/templates/styles.css +0 -2
- package/skills/magpie/evals/consent-required-never-approves/case.yaml +0 -4
- package/skills/magpie/evals/consent-required-never-approves/fixture.sh +0 -213
- package/skills/magpie/evals/consent-required-never-approves/graders/context-stage-was-closed.md +0 -6
- package/skills/magpie/evals/consent-required-never-approves/graders/indexing-was-never-approved.md +0 -7
- package/skills/magpie/evals/consent-required-never-approves/graders/probe-was-run.md +0 -5
- package/skills/magpie/evals/consent-required-never-approves/graders/skill-fired.md +0 -5
- package/skills/magpie/evals/consent-required-never-approves/graders/user-was-told-it-is-unavailable.md +0 -10
- package/skills/magpie/evals/consent-required-never-approves/prompt.md +0 -11
package/package.json
CHANGED
package/skills/index.json
CHANGED
|
@@ -221,7 +221,7 @@
|
|
|
221
221
|
"name": "magpie",
|
|
222
222
|
"description": "Interactive PR review pipeline. Runs five parallel specialist subagents (security, bugs, performance, code-smells, architecture), dedupes findings, applies a critic rubric, peer-reviews via codex exec (falling back to a Claude second opinion when codex is unavailable), and serves an interactive HTML report for selecting findings to post via gh. Bundles a Bun CLI installed onto PATH via the skill's postinstall step. Use when the user asks to review a GitHub pull request. Splits oversized diffs into budgeted shards and rebuilds the diff from the local clone when gh pr diff refuses it.",
|
|
223
223
|
"type": "prompt",
|
|
224
|
-
"version": "0.
|
|
224
|
+
"version": "0.12.0"
|
|
225
225
|
},
|
|
226
226
|
{
|
|
227
227
|
"name": "migrate",
|
package/skills/magpie/README.md
CHANGED
|
@@ -12,7 +12,6 @@ Given a GitHub PR number, dispatches five specialist subagents in parallel (secu
|
|
|
12
12
|
- `gh` on PATH, authenticated (`gh auth status`)
|
|
13
13
|
- `git` on PATH
|
|
14
14
|
- `codex` on PATH, authenticated (optional; if absent the peer-review stage falls back to a Claude second-opinion subagent)
|
|
15
|
-
- code intelligence, with a completed index for the repo under review (optional; if absent the specialists review from the diff and worktree alone). Either interface works: the `code-intel` CLI on PATH, which is preferred because it takes `--repo` per call and needs no session binding, or the code-intelligence MCP server
|
|
16
15
|
|
|
17
16
|
## Install
|
|
18
17
|
|
|
@@ -55,7 +54,7 @@ The fixture lives at `fixtures/example-pr/` (pr.json + findings.final.json + pos
|
|
|
55
54
|
## Layout
|
|
56
55
|
|
|
57
56
|
- `SKILL.md` is the agent-facing prompt: the stage walkthrough and nothing else. Installed by the agent-skills CLI.
|
|
58
|
-
- `references/` holds the prompt bodies the walkthrough loads on demand, one file per stage that needs one: `scout.md` (stage 3, the PR-brief prompt and the `brief.json` contract), `specialists.md` (stage 4, the five focus blocks plus the shared output contract
|
|
57
|
+
- `references/` holds the prompt bodies the walkthrough loads on demand, one file per stage that needs one: `scout.md` (stage 3, the PR-brief prompt and the `brief.json` contract), `specialists.md` (stage 4, the five focus blocks plus the shared output contract), `critic.md` (stage 6), `peer-review.md` (stage 7, including the Claude-fallback preamble). They ship in the bundle and sit next to `SKILL.md` once installed.
|
|
59
58
|
- `skill.json` is the agent-skills manifest.
|
|
60
59
|
- `bin/magpie` is the CLI invoked by the agent during stages; symlinked onto PATH by `install.sh`.
|
|
61
60
|
- `scripts/` holds the implementation (server, dedupe, render, setup, cleanup, etc.).
|
package/skills/magpie/SKILL.md
CHANGED
|
@@ -7,7 +7,7 @@ description: Use when the user asks to review a GitHub pull request (a PR number
|
|
|
7
7
|
|
|
8
8
|
## Prerequisites
|
|
9
9
|
|
|
10
|
-
The skill pre-flights `bun`, `gh`, `git` (required) and `codex` (optional). A missing required binary aborts the run with a single install hint line.
|
|
10
|
+
The skill pre-flights `bun`, `gh`, `git` (required) and `codex` (optional). A missing required binary aborts the run with a single install hint line. Without `codex` the run continues and peer review falls back to a Claude second-opinion subagent (setup prints a one-line notice and logs `{stage: preflight, status: done, missingOptional: ["codex"]}`).
|
|
11
11
|
|
|
12
12
|
## Stage walkthrough
|
|
13
13
|
|
|
@@ -83,19 +83,11 @@ magpie render "$RUN_DIR" progress
|
|
|
83
83
|
|
|
84
84
|
### 3. Context
|
|
85
85
|
|
|
86
|
-
Append `{stage: context, status: running}` to `$RUN_DIR/log.jsonl` and re-render progress. This stage
|
|
86
|
+
Append `{stage: context, status: running}` to `$RUN_DIR/log.jsonl` and re-render progress. This stage never aborts the run.
|
|
87
87
|
|
|
88
|
-
|
|
88
|
+
Read `references/scout.md` and dispatch one subagent (Agent tool, `general-purpose`) carrying the `magpie-scout` block with `<<RUN_DIR>>` and `<<PR_NUMBER>>` substituted. It writes `$RUN_DIR/brief.json`.
|
|
89
89
|
|
|
90
|
-
|
|
91
|
-
- **MCP**, when the CLI is absent but `mcp__code-intelligence__*` tools are in your tool list. Call `bind_workspace` with `$RUN_DIR/worktree` and apply the same rules, polling `get_index_stats` instead, to set `CODE_INTELLIGENCE=mcp`. **Never call `approve_indexing`.**
|
|
92
|
-
- `consent_required` on either interface means the base repo has never completed an index, and starting one is a full GPU pass the user did not ask for. That, no interface at all, or any other error that survives one retry, sets `CODE_INTELLIGENCE=unavailable`; print one line: "Code intelligence is unavailable (<reason>); specialists will review from the diff alone."
|
|
93
|
-
|
|
94
|
-
Never block the pipeline on a still-running index. The scout and specialist contracts both handle a tool that is not ready yet.
|
|
95
|
-
|
|
96
|
-
**Scout.** Read `references/scout.md` and dispatch one subagent (Agent tool, `general-purpose`) carrying the `magpie-scout` block with `<<RUN_DIR>>`, `<<PR_NUMBER>>`, and `<<CODE_INTELLIGENCE>>` substituted. It writes `$RUN_DIR/brief.json`.
|
|
97
|
-
|
|
98
|
-
Append `{stage: context, status: done, codeIntelligence: true|false, interface: "cli"|"mcp"|"none"}` and re-render progress. If the scout returned without writing `brief.json`, append `{stage: context, status: skipped, codeIntelligence: true|false, interface: ...}` instead and continue: the brief is optional everywhere it is read. Both entries carry the probe's result, which is known whatever the scout did, and `interface` is what stage 10 reads to decide whether there is a session to rebind.
|
|
90
|
+
Append `{stage: context, status: done}` and re-render progress. If the scout returned without writing `brief.json`, append `{stage: context, status: skipped}` instead and continue: the brief is optional everywhere it is read.
|
|
99
91
|
|
|
100
92
|
### 4. Specialists
|
|
101
93
|
|
|
@@ -148,8 +140,7 @@ gate is expected to have none. `magpie dedupe` re-checks this against the manife
|
|
|
148
140
|
names every missing pair on stdout, as a backstop rather than a substitute.
|
|
149
141
|
|
|
150
142
|
If every specialist fails (no findings files written), log
|
|
151
|
-
`{stage: specialists, status: error}
|
|
152
|
-
`CODE_INTELLIGENCE=mcp` (stage 10), and stop. Otherwise mark `{stage: specialists, status: done}`.
|
|
143
|
+
`{stage: specialists, status: error}` and stop. Otherwise mark `{stage: specialists, status: done}`.
|
|
153
144
|
|
|
154
145
|
### 5. Dedupe
|
|
155
146
|
|
|
@@ -241,8 +232,6 @@ magpie render "$RUN_DIR" findings
|
|
|
241
232
|
magpie cleanup "$RUN_DIR" --repo "$REPO"
|
|
242
233
|
```
|
|
243
234
|
|
|
244
|
-
If the context stage set `CODE_INTELLIGENCE=mcp`, rebind the session now: call `bind_workspace` with `$REPO`. MCP binding is per session with no per-call override, so ending a run without this leaves the session pointed at a worktree `cleanup` just deleted. A `cli` run has nothing to rebind, because `--repo` names the workspace on every call. Either way the daemon prunes the seeded index once the worktree is gone.
|
|
245
|
-
|
|
246
235
|
The run directory is renamed to `<run-dir>.archived-<timestamp>` and the worktree is removed. The CLI prints two lines on success: `archived to <path>` and `view later: magpie open <archived-id>`. Surface that second line verbatim so the user has a one-command path back to the report.
|
|
247
236
|
|
|
248
237
|
The archived `findings.html` is self-contained and auto-switches to read-only "archived" mode when opened, so:
|
|
@@ -262,7 +251,7 @@ magpie status "$RUN_DIR"
|
|
|
262
251
|
|
|
263
252
|
The JSON output tells you `lastCompleted` and `next`. Resume from `next`:
|
|
264
253
|
|
|
265
|
-
- `context
|
|
254
|
+
- `context`: if `$RUN_DIR/brief.json` exists, append `{stage: context, status: done}` and move on; otherwise re-run stage 3.
|
|
266
255
|
- Any other stage: run it as written in the walkthrough.
|
|
267
256
|
- If a specialist focus has no findings file but its sibling stages are done,
|
|
268
257
|
re-dispatch only that focus. On a sharded run the unit is the `(focus, shard)` pair:
|
|
@@ -275,4 +264,4 @@ The original server is gone. Restart it with `magpie serve "$RUN_DIR"` (step 2)
|
|
|
275
264
|
|
|
276
265
|
## Aborting
|
|
277
266
|
|
|
278
|
-
If the user types `abort` mid-run,
|
|
267
|
+
If the user types `abort` mid-run, run `magpie cleanup` and exit.
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# magpie evals
|
|
2
2
|
|
|
3
|
-
|
|
3
|
+
Five cases for `claude plugin eval`. Every one starts mid-pipeline, because the
|
|
4
4
|
decisions this skill owns are the ones between the CLI calls: `scripts/__tests__/`
|
|
5
5
|
already pins what `magpie setup`, `dedupe`, `shard`, `post` and `status` compute.
|
|
6
6
|
What no unit test can reach is whether the agent stops where the walkthrough says
|
|
@@ -13,7 +13,6 @@ must not touch.
|
|
|
13
13
|
| `codex-missing-falls-back` | No codex on the machine at stage 7 | Claude path with the independence preamble, `provider: claude`, never `status: error` |
|
|
14
14
|
| `report-ends-the-turn` | Stage 8 reached | Renders, logs the stage done, hands back for selection, posts nothing |
|
|
15
15
|
| `post-folds-selection-events` | The user typed `post` after re-ticking | Folds `state/events` last-event-wins, posts `bugs-1,perf-1` only |
|
|
16
|
-
| `consent-required-never-approves` | The code-intel probe wants consent | Never runs `index approve`, prints the unavailable notice, closes the stage |
|
|
17
16
|
| `resume-finds-active-run` | A fresh review ask on a PR with a live run | Checks `--list-runs` first, never calls `setup`, surfaces the interrupted run |
|
|
18
17
|
|
|
19
18
|
## Running
|
|
@@ -26,8 +25,8 @@ claude plugin eval skills/magpie --scaffold --allow-tools Bash Write Edit
|
|
|
26
25
|
```
|
|
27
26
|
|
|
28
27
|
`--runs 1 --ablation none` is the cheap iteration loop; `-j 3` runs three cases
|
|
29
|
-
at once. Most of the cost sits in `codex-missing-falls-back
|
|
30
|
-
|
|
28
|
+
at once. Most of the cost sits in `codex-missing-falls-back`, which dispatches a real
|
|
29
|
+
subagent.
|
|
31
30
|
|
|
32
31
|
Pass `--model` to run the cases on a specific model, which is the point of the
|
|
33
32
|
suite when a new one lands:
|
|
@@ -52,7 +51,7 @@ claude plugin eval skills/magpie --scaffold --allow-tools Bash Write Edit \
|
|
|
52
51
|
## How the fixtures fake the pipeline
|
|
53
52
|
|
|
54
53
|
The eval child runs in a sandbox that refuses to execute anything outside it, so
|
|
55
|
-
the real `magpie`, `gh
|
|
54
|
+
the real `magpie`, `gh` and `codex` are all unreachable: a bare
|
|
56
55
|
`magpie setup` there dies with `Operation not permitted`, not with a diff. Each
|
|
57
56
|
`fixture.sh` therefore writes its own fakes into `$HOME/shims` and puts that
|
|
58
57
|
directory first on `PATH` via `$HOME/.zshenv`, which is the one startup file the
|
|
@@ -108,7 +108,7 @@ JSON
|
|
|
108
108
|
cat > "$RUN_DIR/log.jsonl" <<'LOG'
|
|
109
109
|
{"stage":"preflight","status":"done","missingOptional":["codex"]}
|
|
110
110
|
{"stage":"setup","status":"done"}
|
|
111
|
-
{"stage":"context","status":"done"
|
|
111
|
+
{"stage":"context","status":"done"}
|
|
112
112
|
{"stage":"specialists","status":"done"}
|
|
113
113
|
{"stage":"dedupe","status":"done"}
|
|
114
114
|
{"stage":"critic","status":"done"}
|
|
@@ -139,7 +139,7 @@ JSON
|
|
|
139
139
|
cat > "$RUN_DIR/log.jsonl" <<'LOG'
|
|
140
140
|
{"stage":"preflight","status":"done","missingOptional":["codex"]}
|
|
141
141
|
{"stage":"setup","status":"done"}
|
|
142
|
-
{"stage":"context","status":"done"
|
|
142
|
+
{"stage":"context","status":"done"}
|
|
143
143
|
{"stage":"specialists","status":"done"}
|
|
144
144
|
{"stage":"dedupe","status":"done"}
|
|
145
145
|
{"stage":"critic","status":"done"}
|
|
@@ -117,7 +117,7 @@ JSON
|
|
|
117
117
|
cat > "$RUN_DIR/log.jsonl" <<'LOG'
|
|
118
118
|
{"stage":"preflight","status":"done","missingOptional":["codex"]}
|
|
119
119
|
{"stage":"setup","status":"done"}
|
|
120
|
-
{"stage":"context","status":"done"
|
|
120
|
+
{"stage":"context","status":"done"}
|
|
121
121
|
{"stage":"specialists","status":"done"}
|
|
122
122
|
{"stage":"dedupe","status":"done"}
|
|
123
123
|
{"stage":"critic","status":"done"}
|
|
@@ -126,7 +126,7 @@ JSON
|
|
|
126
126
|
cat > "$RUN_DIR/log.jsonl" <<'LOG'
|
|
127
127
|
{"stage":"preflight","status":"done","missingOptional":["codex"]}
|
|
128
128
|
{"stage":"setup","status":"done"}
|
|
129
|
-
{"stage":"context","status":"done"
|
|
129
|
+
{"stage":"context","status":"done"}
|
|
130
130
|
{"stage":"specialists","status":"done"}
|
|
131
131
|
{"stage":"dedupe","status":"done"}
|
|
132
132
|
{"stage":"critic","status":"done"}
|
|
@@ -118,7 +118,7 @@ JSON
|
|
|
118
118
|
cat > "$RUN_DIR/log.jsonl" <<'LOG'
|
|
119
119
|
{"stage":"preflight","status":"done","missingOptional":["codex"]}
|
|
120
120
|
{"stage":"setup","status":"done"}
|
|
121
|
-
{"stage":"context","status":"done"
|
|
121
|
+
{"stage":"context","status":"done"}
|
|
122
122
|
LOG
|
|
123
123
|
|
|
124
124
|
echo '[]' > "$RUN_DIR/findings/tests.json"
|
|
@@ -5,10 +5,6 @@
|
|
|
5
5
|
"Threads a per-request deadline through the storage client",
|
|
6
6
|
"Adds a metrics counter for exhausted retry budgets"
|
|
7
7
|
],
|
|
8
|
-
"subsystems": [
|
|
9
|
-
{ "name": "upload", "role": "owns the client-facing put path" },
|
|
10
|
-
{ "name": "storage-client", "role": "wraps the S3 SDK and owns timeouts" }
|
|
11
|
-
],
|
|
12
8
|
"watchItems": [
|
|
13
9
|
"The PR body claims the retry is idempotent, but no idempotency key is sent with the put"
|
|
14
10
|
],
|
|
@@ -6,11 +6,6 @@ with `<<RUN_DIR>>` and `<<PR_NUMBER>>` replaced by the real values first. The
|
|
|
6
6
|
subagent has no shell variables from your session, so an unexpanded path means it
|
|
7
7
|
writes the brief where nothing will read it.
|
|
8
8
|
|
|
9
|
-
Substitute `<<CODE_INTELLIGENCE>>` with the value stage 3's probe assigned: `cli` for
|
|
10
|
-
the `code-intel` CLI, `mcp` for the code-intelligence MCP tools, `unavailable` for
|
|
11
|
-
neither. The scout still runs when code intelligence is unavailable; only the
|
|
12
|
-
subsystem map degrades.
|
|
13
|
-
|
|
14
9
|
```magpie-scout
|
|
15
10
|
You are a senior engineer building the orienting brief that five specialist
|
|
16
11
|
reviewers will read before they review PR #<<PR_NUMBER>>. You are not reviewing the
|
|
@@ -19,7 +14,6 @@ code. You are answering "what is this PR for, and what does it actually do".
|
|
|
19
14
|
Working directory: <<RUN_DIR>>/worktree
|
|
20
15
|
PR metadata: <<RUN_DIR>>/pr.json
|
|
21
16
|
Diff: <<RUN_DIR>>/diff.patch
|
|
22
|
-
Code intelligence: <<CODE_INTELLIGENCE>>
|
|
23
17
|
|
|
24
18
|
## What to read
|
|
25
19
|
|
|
@@ -29,36 +23,6 @@ Code intelligence: <<CODE_INTELLIGENCE>>
|
|
|
29
23
|
3. The worktree for surrounding context on any file the diff changes but does not
|
|
30
24
|
explain.
|
|
31
25
|
|
|
32
|
-
## Code intelligence
|
|
33
|
-
|
|
34
|
-
Unless the line above says `unavailable`, an index of `<<RUN_DIR>>/worktree`, seeded
|
|
35
|
-
from the base repository, is queryable. Use it for two things only: name the subsystem
|
|
36
|
-
each touched directory belongs to and what it is responsible for, then say what depends
|
|
37
|
-
on those modules.
|
|
38
|
-
|
|
39
|
-
When the line says `cli`, run the `code-intel` CLI. Every call names the workspace, so
|
|
40
|
-
there is nothing to bind first:
|
|
41
|
-
|
|
42
|
-
code-intel investigate --repo <<RUN_DIR>>/worktree --mode module --file-path <dir> \
|
|
43
|
-
--json "what is this module responsible for"
|
|
44
|
-
code-intel dependency-graph --repo <<RUN_DIR>>/worktree --json <symbol>
|
|
45
|
-
|
|
46
|
-
`code-intel repo-map --repo <<RUN_DIR>>/worktree --json` is the cheaper first look when
|
|
47
|
-
you do not yet know which directories matter. Exit 5 means no results, which is an
|
|
48
|
-
answer. Exit 3 means the daemon is down: stop querying and write `"subsystems": []`.
|
|
49
|
-
|
|
50
|
-
When the line says `mcp`, call `bind_workspace` with `<<RUN_DIR>>/worktree` before your
|
|
51
|
-
first query, then `get_module_summary` on each touched directory and
|
|
52
|
-
`explore_dependency_graph` on the touched modules.
|
|
53
|
-
|
|
54
|
-
If a query reports `indexing_in_progress` (the CLI returns
|
|
55
|
-
`error.code: workspace_unavailable` with that prefix), finish reading the diff and retry
|
|
56
|
-
once. If it still is not ready, or the line above says `unavailable`, write
|
|
57
|
-
`"subsystems": []` and carry on. Never call `approve_indexing` and never run
|
|
58
|
-
`code-intel index approve`; never call `refresh_index` and never run
|
|
59
|
-
`code-intel index refresh`. Triggering a full index is a consent-gated operation that
|
|
60
|
-
is not yours to start.
|
|
61
|
-
|
|
62
26
|
## How to reason
|
|
63
27
|
|
|
64
28
|
1. State the purpose in your own words, not the author's. If you cannot restate it
|
|
@@ -74,17 +38,13 @@ is not yours to start.
|
|
|
74
38
|
## Output contract
|
|
75
39
|
|
|
76
40
|
Write `<<RUN_DIR>>/brief.json` before returning. The file MUST be a JSON object with
|
|
77
|
-
exactly these
|
|
41
|
+
exactly these four keys:
|
|
78
42
|
|
|
79
43
|
{
|
|
80
44
|
"purpose": string, // 1-3 sentences: what this PR is for, in your words. Required
|
|
81
45
|
// and non-empty; a brief with no purpose is discarded whole.
|
|
82
46
|
"changes": string[], // 3-7 entries: what it actually does, grouped by concern.
|
|
83
47
|
// One clause each, no trailing period needed.
|
|
84
|
-
"subsystems": [ { "name": string, "role": string } ],
|
|
85
|
-
// Code-intelligence derived. `name` is the subsystem, `role`
|
|
86
|
-
// is one clause on what it is responsible for. [] when
|
|
87
|
-
// code intelligence is unavailable.
|
|
88
48
|
"watchItems": string[], // Where the diff and the stated intent diverge, or where the
|
|
89
49
|
// intent implies a risk the diff does not address. Often
|
|
90
50
|
// empty. See the boundary below.
|
|
@@ -101,5 +61,5 @@ commit messages or issue titles into the brief: the report reads those from `pr.
|
|
|
101
61
|
directly, and repeating them wastes the specialists' attention.
|
|
102
62
|
|
|
103
63
|
Return as your final tool result a single line:
|
|
104
|
-
`brief: <N> changes, <
|
|
64
|
+
`brief: <N> changes, <K> watch items`. Do not include other prose.
|
|
105
65
|
```
|
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
# Specialist prompts
|
|
2
2
|
|
|
3
3
|
Stage 4 of the walkthrough dispatches five subagents per shard from this file. Build
|
|
4
|
-
each prompt from up to
|
|
4
|
+
each prompt from up to five parts, in this order, and send it as the agent's entire
|
|
5
5
|
task. Each part is a fenced block or the named section it points at; the prose around
|
|
6
6
|
them is instruction to you, not text for the specialist.
|
|
7
7
|
|
|
@@ -41,12 +41,7 @@ in an excluded file (a migration implies a model snapshot, a schema change impli
|
|
|
41
41
|
generated types), open it and cross-check rather than treating it as out of scope.
|
|
42
42
|
```
|
|
43
43
|
|
|
44
|
-
5.
|
|
45
|
-
context stage logged `codeIntelligence: true`: the `-cli` block for
|
|
46
|
-
`interface: "cli"`, the `-mcp` block for `interface: "mcp"`. Omit it entirely
|
|
47
|
-
otherwise, and never send both: telling a specialist to use tools it does not have
|
|
48
|
-
wastes a turn per specialist on discovery.
|
|
49
|
-
6. The brief, when `<RUN_DIR>/brief.json` exists, rendered as:
|
|
44
|
+
5. The brief, when `<RUN_DIR>/brief.json` exists, rendered as:
|
|
50
45
|
|
|
51
46
|
```
|
|
52
47
|
## What this PR is for
|
|
@@ -56,9 +51,6 @@ generated types), open it and cross-check rather than treating it as out of scop
|
|
|
56
51
|
What it does:
|
|
57
52
|
- <each entry of changes>
|
|
58
53
|
|
|
59
|
-
Subsystems it lands in:
|
|
60
|
-
- <name>: <role> (omit this heading when subsystems is empty)
|
|
61
|
-
|
|
62
54
|
Watch items:
|
|
63
55
|
- <each entry of watchItems> (omit this heading when watchItems is empty)
|
|
64
56
|
|
|
@@ -132,7 +124,7 @@ Good:
|
|
|
132
124
|
- `Observation: <one idea, what the diff actually does and where>`
|
|
133
125
|
- `Why it matters: <impact at realistic scale or on a real user path>`
|
|
134
126
|
- `Suggested direction: <one concrete next step, optional if the fix isn't obvious>`
|
|
135
|
-
- `Needs verification: <what you couldn't confirm from the bundle, optional, low/medium severity only>` This labelled paragraph is the only channel for uncertainty: never hedge inside another section, and never raise `severity` to compensate for what you couldn't verify (a blocker/high you cannot stand behind is not a blocker/high). Use the exact `Needs verification:` prefix, not inline phrasing. When
|
|
127
|
+
- `Needs verification: <what you couldn't confirm from the bundle, optional, low/medium severity only>` This labelled paragraph is the only channel for uncertainty: never hedge inside another section, and never raise `severity` to compensate for what you couldn't verify (a blocker/high you cannot stand behind is not a blocker/high). Use the exact `Needs verification:` prefix, not inline phrasing. When reading the worktree could answer the question, look before you hedge. A question you resolved is not a `Needs verification:` paragraph, it is evidence: cite the file:line you found under `Observation:` and omit the paragraph entirely.
|
|
136
128
|
|
|
137
129
|
One idea per paragraph. Do not collapse them into a single wall of text. Do not invent extra labels. If a section doesn't apply, omit it. The interactive report and the GitHub comment both parse these labels and render them as section headers, so missing labels degrade the output.
|
|
138
130
|
|
|
@@ -145,105 +137,6 @@ One idea per paragraph. Do not collapse them into a single wall of text. Do not
|
|
|
145
137
|
|
|
146
138
|
If you have no findings, write []. Return as your final tool result a single line: `<focus>: <N> findings (<blocker>/<high>/<medium>/<low>)`. Do not include other prose.
|
|
147
139
|
|
|
148
|
-
## Codebase intelligence
|
|
149
|
-
|
|
150
|
-
Include one of these blocks as part 5 when the context stage logged
|
|
151
|
-
`codeIntelligence: true`: the `-cli` block when it logged `interface: "cli"`, the
|
|
152
|
-
`-mcp` block when it logged `interface: "mcp"`. Both reach the same index. Never send
|
|
153
|
-
both: a specialist told about a verb it cannot invoke wastes turns discovering that.
|
|
154
|
-
|
|
155
|
-
```magpie-codebase-intelligence-cli
|
|
156
|
-
## Codebase intelligence
|
|
157
|
-
|
|
158
|
-
The `code-intel` CLI queries an index of this exact worktree, including the PR's own
|
|
159
|
-
changes. Every invocation names the workspace with `--repo <RUN_DIR>/worktree`, so
|
|
160
|
-
there is nothing to bind and nothing another agent can change under you. Always pass
|
|
161
|
-
`--json`.
|
|
162
|
-
|
|
163
|
-
These answer the cross-file questions a diff cannot:
|
|
164
|
-
|
|
165
|
-
code-intel ask --repo <RUN_DIR>/worktree --json "<question>"
|
|
166
|
-
A natural-language question answered with grounded evidence. Start here when you
|
|
167
|
-
do not yet know which symbol to pivot on.
|
|
168
|
-
code-intel search --repo <RUN_DIR>/worktree --context snippets --json "<query>"
|
|
169
|
-
Hybrid semantic and literal search. Use to check whether something already
|
|
170
|
-
exists before claiming the PR should add it. Hits carry `id`s.
|
|
171
|
-
code-intel hydrate --repo <RUN_DIR>/worktree --ids <id,id> --json
|
|
172
|
-
The exact bodies behind ids a search returned, without re-reading whole files.
|
|
173
|
-
code-intel definition --repo <RUN_DIR>/worktree --json <symbol>
|
|
174
|
-
The body of a symbol the diff calls but does not show.
|
|
175
|
-
code-intel references --repo <RUN_DIR>/worktree --json <symbol>
|
|
176
|
-
code-intel call-hierarchy --repo <RUN_DIR>/worktree --json <symbol>
|
|
177
|
-
Who calls this, and what does it call. Use for reachability: is the path you are
|
|
178
|
-
worried about actually reachable. Add `--direction callers|callees`, `--depth N`.
|
|
179
|
-
code-intel investigate --repo <RUN_DIR>/worktree --mode impact --target <symbol> --json "<question>"
|
|
180
|
-
The reverse-dependency set for a symbol. Use for blast radius before claiming a
|
|
181
|
-
change is safe or unsafe.
|
|
182
|
-
code-intel investigate --repo <RUN_DIR>/worktree --mode data --target <symbol> --json "<question>"
|
|
183
|
-
Follow a value from its origin to where it is used. Use to confirm that
|
|
184
|
-
untrusted input actually reaches the sink you are worried about.
|
|
185
|
-
code-intel dependency-graph --repo <RUN_DIR>/worktree --json <symbol>
|
|
186
|
-
Module-level edges. Use for cycles and boundaries.
|
|
187
|
-
|
|
188
|
-
There is no test-coverage command: to find out whether the symbol you are flagging is
|
|
189
|
-
covered, run `references` on it and read the paths for test files.
|
|
190
|
-
|
|
191
|
-
Exit codes are the contract, not the prose: 5 means no results, which is an answer and
|
|
192
|
-
not a reason to fall back to grep. 3 means the daemon is down and 4 means the workspace
|
|
193
|
-
is not queryable; either way, stop querying and review from the diff and worktree.
|
|
194
|
-
|
|
195
|
-
Rules:
|
|
196
|
-
|
|
197
|
-
- Cite what you find as ordinary `file:line` evidence under `Observation:`. Do not
|
|
198
|
-
say "code intelligence told me"; the location is the evidence.
|
|
199
|
-
- Verify before you report. A finding you could have refuted with one query and did
|
|
200
|
-
not is worse than no finding: it costs the author trust and the critic a slot.
|
|
201
|
-
- Verify before you hedge. If a query can answer the question, a `Needs verification:`
|
|
202
|
-
paragraph is a failure to look, not honest uncertainty.
|
|
203
|
-
- If a command fails with `error.code: workspace_unavailable` and an
|
|
204
|
-
`indexing_in_progress` message, finish reading the diff and retry once. If it is
|
|
205
|
-
still not ready, review from the diff and worktree alone. Do not block.
|
|
206
|
-
- Never run `code-intel index approve`. Never run `code-intel index refresh`. Starting
|
|
207
|
-
a full index is a consent-gated operation that is not yours to start.
|
|
208
|
-
```
|
|
209
|
-
|
|
210
|
-
```magpie-codebase-intelligence-mcp
|
|
211
|
-
## Codebase intelligence
|
|
212
|
-
|
|
213
|
-
You have code-intelligence MCP tools against an index of this exact worktree,
|
|
214
|
-
including the PR's own changes. Call `bind_workspace` with `<RUN_DIR>/worktree` before
|
|
215
|
-
your first query.
|
|
216
|
-
|
|
217
|
-
These answer the cross-file questions a diff cannot:
|
|
218
|
-
|
|
219
|
-
- `ask_code`: a natural-language question, answered with grounded evidence. Start here
|
|
220
|
-
when you do not yet know which symbol to pivot on.
|
|
221
|
-
- `find_references` / `get_call_hierarchy`: who calls this, and what does it call.
|
|
222
|
-
Use for reachability: is the path you are worried about actually reachable.
|
|
223
|
-
- `find_affected_code`: the full reverse-dependency set for a symbol. Use for blast
|
|
224
|
-
radius before claiming a change is safe or unsafe.
|
|
225
|
-
- `trace_data_flow`: follow a value from its origin to where it is used. Use to
|
|
226
|
-
confirm that untrusted input actually reaches the sink you are worried about.
|
|
227
|
-
- `search_code`: hybrid semantic and literal search. Use to check whether something
|
|
228
|
-
already exists before claiming the PR should add it.
|
|
229
|
-
- `get_definition`: the body of a symbol the diff calls but does not show.
|
|
230
|
-
- `explore_dependency_graph`: module-level edges. Use for cycles and boundaries.
|
|
231
|
-
- `find_tests_for_symbol`: whether the symbol you are flagging is covered.
|
|
232
|
-
|
|
233
|
-
Rules:
|
|
234
|
-
|
|
235
|
-
- Cite what you find as ordinary `file:line` evidence under `Observation:`. Do not
|
|
236
|
-
say "code intelligence told me"; the location is the evidence.
|
|
237
|
-
- Verify before you report. A finding you could have refuted with one query and did
|
|
238
|
-
not is worse than no finding: it costs the author trust and the critic a slot.
|
|
239
|
-
- Verify before you hedge. If a tool can answer the question, a `Needs verification:`
|
|
240
|
-
paragraph is a failure to look, not honest uncertainty.
|
|
241
|
-
- If a tool returns `indexing_in_progress`, finish reading the diff and retry once.
|
|
242
|
-
If it is still not ready, review from the diff and worktree alone. Do not block.
|
|
243
|
-
- Never call `approve_indexing`. Never call `refresh_index`. Starting a full index is
|
|
244
|
-
a consent-gated operation that is not yours to start.
|
|
245
|
-
```
|
|
246
|
-
|
|
247
140
|
## Focus blocks
|
|
248
141
|
|
|
249
142
|
### security
|
|
@@ -297,7 +190,7 @@ For each potential finding:
|
|
|
297
190
|
2. Identify the trust boundary: is this crossing from untrusted to trusted context?
|
|
298
191
|
3. Assess exploitability: can an attacker realistically trigger this?
|
|
299
192
|
4. Evaluate impact: what's the blast radius if exploited?
|
|
300
|
-
5. Confirm the flow from the entry point to the sink before reporting,
|
|
193
|
+
5. Confirm the flow from the entry point to the sink before reporting, by following the value through the worktree. A taint path you asserted but did not trace is a guess.
|
|
301
194
|
|
|
302
195
|
**Risk guide:**
|
|
303
196
|
- blocker: Realistic path to remote code execution, auth bypass, data breach, or privilege escalation
|
|
@@ -362,7 +255,7 @@ For each potential bug:
|
|
|
362
255
|
3. What's the consequence: crash, data corruption, silent wrong behavior?
|
|
363
256
|
4. Is there an existing guard I'm not seeing?
|
|
364
257
|
|
|
365
|
-
Question 4 is answerable:
|
|
258
|
+
Question 4 is answerable: searching the worktree for the changed symbol's callers shows where the guard would have to live. Check before you file.
|
|
366
259
|
|
|
367
260
|
**Risk guide:**
|
|
368
261
|
- blocker: Data loss, data corruption, broken auth/session behavior, or consistently crashing a major workflow
|
|
@@ -422,7 +315,7 @@ For each potential issue:
|
|
|
422
315
|
2. How often does this code path execute? (once on init vs. every keystroke)
|
|
423
316
|
3. What's the measurable impact? (milliseconds vs. seconds)
|
|
424
317
|
4. Is the optimization worth the complexity cost?
|
|
425
|
-
5. Establish the call frequency before claiming a path is hot,
|
|
318
|
+
5. Establish the call frequency before claiming a path is hot, by finding the changed symbol's callers in the worktree. "Called from one cold init path" and "called per keystroke" are different findings.
|
|
426
319
|
|
|
427
320
|
**Risk guide:**
|
|
428
321
|
- blocker: Change can make a major workflow unusable or cause unbounded production resource exhaustion
|
|
@@ -484,7 +377,7 @@ For each potential smell:
|
|
|
484
377
|
2. Confirm the smell is introduced or materially worsened by this PR, not merely pre-existing nearby code.
|
|
485
378
|
3. Suggest the smallest refactor that fits the surrounding codebase patterns.
|
|
486
379
|
4. Weigh the cost: do not ask for a new abstraction unless it reduces real duplication, coupling, or reasoning burden now.
|
|
487
|
-
5. Before claiming the PR duplicates something or should reuse an existing helper, find it
|
|
380
|
+
5. Before claiming the PR duplicates something or should reuse an existing helper, find it in the worktree. Name the file:line of the thing it should have reused, or do not make the claim.
|
|
488
381
|
|
|
489
382
|
**Risk guide:**
|
|
490
383
|
- blocker: Smell creates a high-risk maintenance trap likely to cause defects across modules soon
|
|
@@ -543,7 +436,7 @@ For each potential issue:
|
|
|
543
436
|
2. Is this coupling necessary or incidental?
|
|
544
437
|
3. Would a new team member understand where to make changes?
|
|
545
438
|
4. Is this over-engineered for the current requirements, or appropriately future-proofed?
|
|
546
|
-
5. Confirm boundary and cycle claims on the touched modules
|
|
439
|
+
5. Confirm boundary and cycle claims on the touched modules by reading their imports in the worktree. A cycle you inferred from import statements in the diff may already be broken by an interface you cannot see.
|
|
547
440
|
|
|
548
441
|
**Risk guide:**
|
|
549
442
|
- blocker: Change introduces a serious boundary violation or contract break likely to cascade across subsystems
|
|
@@ -92,7 +92,6 @@ test('refreshFindings carries the brief header through, so archives keep it on r
|
|
|
92
92
|
JSON.stringify({
|
|
93
93
|
purpose: 'Adds bounded retries to the upload path.',
|
|
94
94
|
changes: ['Wraps the S3 put in a bounded retry'],
|
|
95
|
-
subsystems: [{ name: 'upload', role: 'owns the put path' }],
|
|
96
95
|
watchItems: [],
|
|
97
96
|
unclear: [],
|
|
98
97
|
}),
|
|
@@ -123,7 +123,6 @@ test('findings render includes the brief header when brief.json is present', asy
|
|
|
123
123
|
JSON.stringify({
|
|
124
124
|
purpose: 'Adds bounded retries to the upload path.',
|
|
125
125
|
changes: ['Wraps the S3 put in a bounded retry'],
|
|
126
|
-
subsystems: [{ name: 'upload', role: 'owns the put path' }],
|
|
127
126
|
watchItems: [],
|
|
128
127
|
unclear: [],
|
|
129
128
|
}),
|
|
@@ -205,15 +205,11 @@ const SAMPLE_BRIEF: PrBrief = {
|
|
|
205
205
|
purpose:
|
|
206
206
|
'Adds bounded retries to the upload path so transient S3 failures stop surfacing to users.',
|
|
207
207
|
changes: ['Wraps the S3 put in a bounded retry', 'Adds a jittered backoff helper'],
|
|
208
|
-
subsystems: [
|
|
209
|
-
{ name: 'upload', role: 'owns the client-facing put path' },
|
|
210
|
-
{ name: 'storage-client', role: 'wraps the S3 SDK' },
|
|
211
|
-
],
|
|
212
208
|
watchItems: ['The PR body claims idempotency but no request key is sent'],
|
|
213
209
|
unclear: ['Whether the retry budget interacts with the outer request timeout'],
|
|
214
210
|
}
|
|
215
211
|
|
|
216
|
-
test('brief header renders purpose
|
|
212
|
+
test('brief header renders purpose and changes', () => {
|
|
217
213
|
const html = renderFindingsHtml({
|
|
218
214
|
findings: SAMPLE_FINDINGS,
|
|
219
215
|
postStatus: {},
|
|
@@ -223,9 +219,6 @@ test('brief header renders purpose, changes, and subsystem chips', () => {
|
|
|
223
219
|
expect(html).toContain('class="pr-brief"')
|
|
224
220
|
expect(html).toContain('Adds bounded retries to the upload path')
|
|
225
221
|
expect(html).toContain('Wraps the S3 put in a bounded retry')
|
|
226
|
-
expect(html).toContain('class="brief-chip"')
|
|
227
|
-
expect(html).toContain('>upload<')
|
|
228
|
-
expect(html).toContain('>storage-client<')
|
|
229
222
|
})
|
|
230
223
|
|
|
231
224
|
test('brief header omits watchItems and unclear, which are prompt-only', () => {
|
|
@@ -244,17 +237,6 @@ test('no brief means no brief header at all', () => {
|
|
|
244
237
|
expect(html).not.toContain('class="pr-brief"')
|
|
245
238
|
})
|
|
246
239
|
|
|
247
|
-
test('an empty subsystem list renders no chip row', () => {
|
|
248
|
-
const html = renderFindingsHtml({
|
|
249
|
-
findings: SAMPLE_FINDINGS,
|
|
250
|
-
postStatus: {},
|
|
251
|
-
highlighter: hl,
|
|
252
|
-
brief: { ...SAMPLE_BRIEF, subsystems: [] },
|
|
253
|
-
})
|
|
254
|
-
expect(html).toContain('class="pr-brief"')
|
|
255
|
-
expect(html).not.toContain('brief-subsystems')
|
|
256
|
-
})
|
|
257
|
-
|
|
258
240
|
test('brief header renders on the empty-findings page too', () => {
|
|
259
241
|
const html = renderFindingsHtml({
|
|
260
242
|
findings: [],
|
|
@@ -291,15 +273,12 @@ test('brief content is HTML-escaped', () => {
|
|
|
291
273
|
expect(html).toContain('<script>')
|
|
292
274
|
})
|
|
293
275
|
|
|
294
|
-
test('
|
|
276
|
+
test('issue title escapes quotes in its title="" attribute context', () => {
|
|
295
277
|
const html = renderFindingsHtml({
|
|
296
278
|
findings: SAMPLE_FINDINGS,
|
|
297
279
|
postStatus: {},
|
|
298
280
|
highlighter: hl,
|
|
299
|
-
brief:
|
|
300
|
-
...SAMPLE_BRIEF,
|
|
301
|
-
subsystems: [{ name: 'upload', role: 'owns the put path" onmouseover="alert(1)' }],
|
|
302
|
-
},
|
|
281
|
+
brief: SAMPLE_BRIEF,
|
|
303
282
|
issues: [
|
|
304
283
|
{
|
|
305
284
|
number: 42,
|
|
@@ -308,11 +287,9 @@ test('subsystem role and issue title escape quotes in their title="" attribute c
|
|
|
308
287
|
},
|
|
309
288
|
],
|
|
310
289
|
})
|
|
311
|
-
// Raw quote-breakout must never appear unescaped in
|
|
312
|
-
expect(html).not.toContain('path" onmouseover="alert(1)')
|
|
290
|
+
// Raw quote-breakout must never appear unescaped in the attribute.
|
|
313
291
|
expect(html).not.toContain('intermittently" onmouseover="alert(2)')
|
|
314
292
|
// Escaped form must appear instead.
|
|
315
|
-
expect(html).toContain('path" onmouseover="alert(1)')
|
|
316
293
|
expect(html).toContain('intermittently" onmouseover="alert(2)')
|
|
317
294
|
// purpose escaping (already covered above) must remain intact alongside this.
|
|
318
295
|
expect(html).toContain('Adds bounded retries to the upload path')
|
|
@@ -155,7 +155,7 @@ test('progress-pane block replaces diff-pane on the progress page', () => {
|
|
|
155
155
|
specialistCounts: { security: 0 },
|
|
156
156
|
})
|
|
157
157
|
expect(html).toContain('data-role="progress-pane"')
|
|
158
|
-
expect(html).toContain('
|
|
158
|
+
expect(html).toContain('Reading the PR to brief the reviewers')
|
|
159
159
|
})
|
|
160
160
|
|
|
161
161
|
test('progress names the shard count while specialists run', () => {
|
|
@@ -1,5 +1,5 @@
|
|
|
1
1
|
import { expect, test } from 'bun:test'
|
|
2
|
-
import { readFile } from 'node:fs/promises'
|
|
2
|
+
import { readdir, readFile } from 'node:fs/promises'
|
|
3
3
|
|
|
4
4
|
const SKILL = new URL('../../SKILL.md', import.meta.url).pathname
|
|
5
5
|
const SKILL_JSON = new URL('../../skill.json', import.meta.url).pathname
|
|
@@ -85,18 +85,28 @@ test('references/scout.md holds the scout prompt and the brief contract', async
|
|
|
85
85
|
const text = await readFile(ref('scout.md'), 'utf8')
|
|
86
86
|
expect(text).toContain('```magpie-scout')
|
|
87
87
|
expect(text).toContain('brief.json')
|
|
88
|
-
for (const key of ['purpose', 'changes', '
|
|
88
|
+
for (const key of ['purpose', 'changes', 'watchItems', 'unclear']) {
|
|
89
89
|
expect(text).toContain(key)
|
|
90
90
|
}
|
|
91
|
+
expect(text).not.toContain('"subsystems"')
|
|
91
92
|
for (const ph of ['<<RUN_DIR>>', '<<PR_NUMBER>>']) {
|
|
92
93
|
expect(text).toContain(ph)
|
|
93
94
|
}
|
|
94
|
-
// The scout must never trigger a full index; that is a consent-gated GPU pass.
|
|
95
|
-
expect(text).toMatch(/never call `approve_indexing`|do not call `approve_indexing`/i)
|
|
96
95
|
// watchItems are context for specialists, not findings in their own right.
|
|
97
96
|
expect(text).toMatch(/not a finding/i)
|
|
98
97
|
})
|
|
99
98
|
|
|
99
|
+
test('no shipped magpie doc or prompt mentions code intelligence', async () => {
|
|
100
|
+
const refs = (await readdir(new URL('../../references/', import.meta.url).pathname)).map(ref)
|
|
101
|
+
const README = new URL('../../README.md', import.meta.url).pathname
|
|
102
|
+
for (const path of [SKILL, README, ...refs]) {
|
|
103
|
+
const text = await readFile(path, 'utf8')
|
|
104
|
+
expect(text).not.toMatch(
|
|
105
|
+
/code[- ]?intel|codebase[- ]intelligence|bind_workspace|approve_indexing/i,
|
|
106
|
+
)
|
|
107
|
+
}
|
|
108
|
+
})
|
|
109
|
+
|
|
100
110
|
test('SKILL.md sends each stage to the reference file it needs', async () => {
|
|
101
111
|
const text = await readFile(SKILL, 'utf8')
|
|
102
112
|
const section = (heading: string) => {
|
|
@@ -117,108 +127,9 @@ test('SKILL.md no longer inlines the prompt bodies it moved out', async () => {
|
|
|
117
127
|
expect(text).not.toContain(`\`\`\`${tag}`)
|
|
118
128
|
}
|
|
119
129
|
// The walkthrough is the always-read part; keep it small enough to be cheap.
|
|
120
|
-
// Raised from 2600 when the sharded-dispatch and fallback-diff prose was added
|
|
121
|
-
//
|
|
122
|
-
|
|
123
|
-
expect(text.split(/\s+/).length).toBeLessThan(3250)
|
|
124
|
-
})
|
|
125
|
-
|
|
126
|
-
test('SKILL.md never instructs the agent to approve indexing', async () => {
|
|
127
|
-
const text = await readFile(SKILL, 'utf8')
|
|
128
|
-
// A full index is a consent-gated GPU pass. The context stage degrades instead.
|
|
129
|
-
// Both interfaces expose the same footgun under different names, so both are named.
|
|
130
|
-
expect(text).toContain('approve_indexing')
|
|
131
|
-
expect(text).toMatch(/never call `approve_indexing`|do not call `approve_indexing`/i)
|
|
132
|
-
expect(text).toMatch(
|
|
133
|
-
/never run `code-intel index approve`|do not run `code-intel index approve`/i,
|
|
134
|
-
)
|
|
135
|
-
})
|
|
136
|
-
|
|
137
|
-
test('SKILL.md stage 3 probes the code-intel CLI before the MCP tools', async () => {
|
|
138
|
-
const text = await readFile(SKILL, 'utf8')
|
|
139
|
-
const start = text.indexOf('### 3. Context')
|
|
140
|
-
expect(start).toBeGreaterThan(-1)
|
|
141
|
-
const section = text.slice(start, text.indexOf('\n### 4.', start))
|
|
142
|
-
// Both interfaces drive the same daemon, but the CLI names the workspace on every
|
|
143
|
-
// call, so it carries no session binding to leak past cleanup. Prefer it.
|
|
144
|
-
expect(section).toContain('command -v code-intel')
|
|
145
|
-
expect(section).toContain('mcp__code-intelligence__')
|
|
146
|
-
expect(section.indexOf('code-intel')).toBeLessThan(section.indexOf('mcp__code-intelligence__'))
|
|
147
|
-
})
|
|
148
|
-
|
|
149
|
-
test('SKILL.md gives CODE_INTELLIGENCE a value per interface', async () => {
|
|
150
|
-
const text = await readFile(SKILL, 'utf8')
|
|
151
|
-
// The scout and specialist prompts branch on this, so all three values must be
|
|
152
|
-
// spelled out where the probe assigns them.
|
|
153
|
-
for (const value of [
|
|
154
|
-
'CODE_INTELLIGENCE=cli',
|
|
155
|
-
'CODE_INTELLIGENCE=mcp',
|
|
156
|
-
'CODE_INTELLIGENCE=unavailable',
|
|
157
|
-
]) {
|
|
158
|
-
expect(text).toContain(value)
|
|
159
|
-
}
|
|
160
|
-
})
|
|
161
|
-
|
|
162
|
-
test('the rebind instruction is scoped to the MCP interface everywhere it appears', async () => {
|
|
163
|
-
const text = await readFile(SKILL, 'utf8')
|
|
164
|
-
// A CLI run holds no session binding, so an unconditional rebind sends the agent
|
|
165
|
-
// after an MCP tool that is not in its tool list.
|
|
166
|
-
const spans: ReadonlyArray<readonly [string, string]> = [
|
|
167
|
-
['### 4. Specialists', '\n### 5.'],
|
|
168
|
-
['### 10. Cleanup', '\n## '],
|
|
169
|
-
['## Aborting', ''],
|
|
170
|
-
]
|
|
171
|
-
for (const [from, to] of spans) {
|
|
172
|
-
const start = text.indexOf(from)
|
|
173
|
-
expect(start).toBeGreaterThan(-1)
|
|
174
|
-
const section = to ? text.slice(start, text.indexOf(to, start)) : text.slice(start)
|
|
175
|
-
expect(section).toMatch(/CODE_INTELLIGENCE=mcp|MCP path/)
|
|
176
|
-
}
|
|
177
|
-
})
|
|
178
|
-
|
|
179
|
-
test('SKILL.md logs codeIntelligence on both the done and skipped context outcomes', async () => {
|
|
180
|
-
const text = await readFile(SKILL, 'utf8')
|
|
181
|
-
const start = text.indexOf('### 3. Context')
|
|
182
|
-
expect(start).toBeGreaterThan(-1)
|
|
183
|
-
const section = text.slice(start, text.indexOf('\n### 4.', start))
|
|
184
|
-
// The bind probe's result is known by the time either log line is written,
|
|
185
|
-
// regardless of whether the scout produced a brief; specialists read this key
|
|
186
|
-
// to decide whether to include the codebase-intelligence block.
|
|
187
|
-
const doneEntry = section.match(/\{stage: context, status: done[^}]*\}/)
|
|
188
|
-
const skippedEntry = section.match(/\{stage: context, status: skipped[^}]*\}/)
|
|
189
|
-
expect(doneEntry?.[0]).toContain('codeIntelligence')
|
|
190
|
-
expect(skippedEntry?.[0]).toContain('codeIntelligence')
|
|
191
|
-
})
|
|
192
|
-
|
|
193
|
-
test('SKILL.md rebinds the code-intelligence session at cleanup', async () => {
|
|
194
|
-
const text = await readFile(SKILL, 'utf8')
|
|
195
|
-
const start = text.indexOf('### 10. Cleanup')
|
|
196
|
-
expect(start).toBeGreaterThan(-1)
|
|
197
|
-
const section = text.slice(start, text.indexOf('\n## ', start))
|
|
198
|
-
// Binding is per session with no per-call override, so a run that ends without
|
|
199
|
-
// rebinding leaves the session pointed at a worktree that no longer exists.
|
|
200
|
-
expect(section).toContain('bind_workspace')
|
|
201
|
-
})
|
|
202
|
-
|
|
203
|
-
test('SKILL.md rebinds before cleanup on the abort path too', async () => {
|
|
204
|
-
const text = await readFile(SKILL, 'utf8')
|
|
205
|
-
const start = text.indexOf('## Aborting')
|
|
206
|
-
expect(start).toBeGreaterThan(-1)
|
|
207
|
-
const section = text.slice(start)
|
|
208
|
-
// Stage 3 bound the session to the worktree; `abort` deletes that worktree via
|
|
209
|
-
// `magpie cleanup`, so it must rebind first or leave the session dangling.
|
|
210
|
-
expect(section).toMatch(/rebind.*(\$REPO|stage 10)/i)
|
|
211
|
-
expect(section).toContain('magpie cleanup')
|
|
212
|
-
})
|
|
213
|
-
|
|
214
|
-
test('SKILL.md rebinds before the stage-4 all-specialists-failed hard stop', async () => {
|
|
215
|
-
const text = await readFile(SKILL, 'utf8')
|
|
216
|
-
const start = text.indexOf('### 4. Specialists')
|
|
217
|
-
expect(start).toBeGreaterThan(-1)
|
|
218
|
-
const section = text.slice(start, text.indexOf('\n### 5.', start))
|
|
219
|
-
// This path stops the run without calling cleanup, but a later resume or
|
|
220
|
-
// abort must not find the session still pointed at the worktree.
|
|
221
|
-
expect(section).toMatch(/rebind.*(\$REPO|stage 10)/i)
|
|
130
|
+
// Raised from 2600 when the sharded-dispatch and fallback-diff prose was added:
|
|
131
|
+
// that's real, load-bearing procedure, not bloat.
|
|
132
|
+
expect(text.split(/\s+/).length).toBeLessThan(3000)
|
|
222
133
|
})
|
|
223
134
|
|
|
224
135
|
test('styles.css declares a prefers-color-scheme:dark block that overrides core tokens', async () => {
|
|
@@ -313,8 +224,8 @@ test('SKILL.md does not gate resume on state/server-info', async () => {
|
|
|
313
224
|
expect(section).not.toMatch(/if .*server-info.* exists/i)
|
|
314
225
|
// Resuming must restart the server; the old one is gone.
|
|
315
226
|
expect(section).toContain('magpie serve')
|
|
316
|
-
// `context` re-runs on resume too (
|
|
317
|
-
//
|
|
227
|
+
// `context` re-runs on resume too (a conditional scout dispatch); the
|
|
228
|
+
// resume section must still call it out explicitly.
|
|
318
229
|
expect(section).toContain('context')
|
|
319
230
|
})
|
|
320
231
|
|
|
@@ -342,69 +253,6 @@ test('SKILL.md has the stage walkthrough', async () => {
|
|
|
342
253
|
expect(text).toMatch(/magpie cleanup/)
|
|
343
254
|
})
|
|
344
255
|
|
|
345
|
-
const intelligenceBlock = (text: string, iface: 'cli' | 'mcp') => {
|
|
346
|
-
const fence = `\`\`\`magpie-codebase-intelligence-${iface}`
|
|
347
|
-
const start = text.indexOf(fence)
|
|
348
|
-
expect(start).toBeGreaterThan(-1)
|
|
349
|
-
return text.slice(start, text.indexOf('\n```', start + fence.length))
|
|
350
|
-
}
|
|
351
|
-
|
|
352
|
-
test('references/specialists.md carries a codebase-intelligence block per interface', async () => {
|
|
353
|
-
const text = await readFile(ref('specialists.md'), 'utf8')
|
|
354
|
-
for (const iface of ['cli', 'mcp'] as const) {
|
|
355
|
-
const block = intelligenceBlock(text, iface)
|
|
356
|
-
expect(block).toContain('indexing_in_progress')
|
|
357
|
-
// The one operation a specialist must never perform, under either name.
|
|
358
|
-
expect(block).toMatch(/never (call `approve_indexing`|run `code-intel index approve`)/i)
|
|
359
|
-
}
|
|
360
|
-
})
|
|
361
|
-
|
|
362
|
-
test('the CLI intelligence block names the workspace per call instead of binding', async () => {
|
|
363
|
-
const text = await readFile(ref('specialists.md'), 'utf8')
|
|
364
|
-
const block = intelligenceBlock(text, 'cli')
|
|
365
|
-
// The CLI has no bind_workspace: every invocation carries --repo, which is what
|
|
366
|
-
// lets the five specialists query the same worktree in parallel without a
|
|
367
|
-
// session binding they could clobber for each other.
|
|
368
|
-
expect(block).not.toContain('bind_workspace')
|
|
369
|
-
expect(block).toContain('--repo')
|
|
370
|
-
expect(block).toContain('--json')
|
|
371
|
-
})
|
|
372
|
-
|
|
373
|
-
test('the MCP intelligence block still binds the workspace first', async () => {
|
|
374
|
-
const text = await readFile(ref('specialists.md'), 'utf8')
|
|
375
|
-
expect(intelligenceBlock(text, 'mcp')).toContain('bind_workspace')
|
|
376
|
-
})
|
|
377
|
-
|
|
378
|
-
test('every focus block names its intelligence capability on both interfaces', async () => {
|
|
379
|
-
const text = await readFile(ref('specialists.md'), 'utf8')
|
|
380
|
-
// A focus block reaches whichever interface the probe found, so naming only the
|
|
381
|
-
// MCP tool leaves a CLI specialist with a verb it cannot invoke.
|
|
382
|
-
const tools: Record<(typeof FOCUSES)[number], readonly [string, string]> = {
|
|
383
|
-
security: ['trace_data_flow', 'investigate --mode data'],
|
|
384
|
-
bugs: ['get_call_hierarchy', 'code-intel call-hierarchy'],
|
|
385
|
-
performance: ['find_affected_code', 'investigate --mode impact'],
|
|
386
|
-
'code-smells': ['search_code', 'code-intel search'],
|
|
387
|
-
architecture: ['explore_dependency_graph', 'code-intel dependency-graph'],
|
|
388
|
-
}
|
|
389
|
-
for (const focus of FOCUSES) {
|
|
390
|
-
const fence = `\`\`\`magpie-specialist-${focus}`
|
|
391
|
-
const start = text.indexOf(fence)
|
|
392
|
-
const block = text.slice(start, text.indexOf('```', start + fence.length))
|
|
393
|
-
for (const tool of tools[focus]) expect(block).toContain(tool)
|
|
394
|
-
}
|
|
395
|
-
})
|
|
396
|
-
|
|
397
|
-
test('references/scout.md documents both intelligence interfaces', async () => {
|
|
398
|
-
const text = await readFile(ref('scout.md'), 'utf8')
|
|
399
|
-
// The scout is dispatched with the same <<CODE_INTELLIGENCE>> value the
|
|
400
|
-
// specialists get, so it has to know what `cli` and `mcp` each mean.
|
|
401
|
-
expect(text).toContain('`cli`')
|
|
402
|
-
expect(text).toContain('`mcp`')
|
|
403
|
-
expect(text).toContain('code-intel ')
|
|
404
|
-
expect(text).toContain('bind_workspace')
|
|
405
|
-
expect(text).toMatch(/never (call `approve_indexing`|run `code-intel index approve`)/i)
|
|
406
|
-
})
|
|
407
|
-
|
|
408
256
|
test('the output contract tells specialists to look before they hedge', async () => {
|
|
409
257
|
const text = await readFile(ref('specialists.md'), 'utf8')
|
|
410
258
|
const contract = text.slice(0, text.indexOf('```magpie-specialist-'))
|
|
@@ -322,18 +322,32 @@ test('parseBrief accepts a well-formed brief', () => {
|
|
|
322
322
|
const brief = parseBrief({
|
|
323
323
|
purpose: 'Adds retry handling to the upload path.',
|
|
324
324
|
changes: ['Wraps the S3 put in a bounded retry', 'Adds a jittered backoff helper'],
|
|
325
|
-
subsystems: [{ name: 'upload', role: 'owns the client-facing put path' }],
|
|
326
325
|
watchItems: ['The PR body claims idempotency but no request key is sent'],
|
|
327
326
|
unclear: ['Whether the retry budget interacts with the outer request timeout'],
|
|
328
327
|
})
|
|
329
328
|
expect(brief).not.toBeNull()
|
|
330
329
|
expect(brief?.purpose).toBe('Adds retry handling to the upload path.')
|
|
331
330
|
expect(brief?.changes).toHaveLength(2)
|
|
332
|
-
expect(brief?.subsystems[0]).toEqual({ name: 'upload', role: 'owns the client-facing put path' })
|
|
333
331
|
expect(brief?.watchItems).toHaveLength(1)
|
|
334
332
|
expect(brief?.unclear).toHaveLength(1)
|
|
335
333
|
})
|
|
336
334
|
|
|
335
|
+
test('parseBrief ignores a subsystems key left in an archived brief', () => {
|
|
336
|
+
const brief = parseBrief({
|
|
337
|
+
purpose: 'Adds retry handling to the upload path.',
|
|
338
|
+
changes: ['Wraps the S3 put in a bounded retry'],
|
|
339
|
+
subsystems: [{ name: 'upload', role: 'owns the client-facing put path' }],
|
|
340
|
+
watchItems: [],
|
|
341
|
+
unclear: [],
|
|
342
|
+
})
|
|
343
|
+
expect(brief).toEqual({
|
|
344
|
+
purpose: 'Adds retry handling to the upload path.',
|
|
345
|
+
changes: ['Wraps the S3 put in a bounded retry'],
|
|
346
|
+
watchItems: [],
|
|
347
|
+
unclear: [],
|
|
348
|
+
})
|
|
349
|
+
})
|
|
350
|
+
|
|
337
351
|
test('parseBrief returns null for a brief with no purpose', () => {
|
|
338
352
|
expect(parseBrief({ changes: ['a'] })).toBeNull()
|
|
339
353
|
expect(parseBrief({ purpose: ' ', changes: ['a'] })).toBeNull()
|
|
@@ -349,17 +363,10 @@ test('parseBrief drops junk entries instead of throwing', () => {
|
|
|
349
363
|
const brief = parseBrief({
|
|
350
364
|
purpose: 'Does a thing.',
|
|
351
365
|
changes: ['kept', 42, null, ' ', 'also kept'],
|
|
352
|
-
subsystems: [{ name: 'kept', role: 'r' }, { role: 'no name' }, 'not an object', null],
|
|
353
366
|
watchItems: 'not an array',
|
|
354
367
|
unclear: undefined,
|
|
355
368
|
})
|
|
356
369
|
expect(brief?.changes).toEqual(['kept', 'also kept'])
|
|
357
|
-
expect(brief?.subsystems).toEqual([{ name: 'kept', role: 'r' }])
|
|
358
370
|
expect(brief?.watchItems).toEqual([])
|
|
359
371
|
expect(brief?.unclear).toEqual([])
|
|
360
372
|
})
|
|
361
|
-
|
|
362
|
-
test('parseBrief defaults a subsystem with no role to an empty role', () => {
|
|
363
|
-
const brief = parseBrief({ purpose: 'p', subsystems: [{ name: 'auth' }] })
|
|
364
|
-
expect(brief?.subsystems).toEqual([{ name: 'auth', role: '' }])
|
|
365
|
-
})
|
|
@@ -123,12 +123,6 @@ function briefBlock(brief: PrBrief | undefined, issues: BriefIssue[]): string {
|
|
|
123
123
|
brief.changes.length > 0
|
|
124
124
|
? `<ul class="brief-changes">${brief.changes.map((c) => `<li>${esc(c)}</li>`).join('')}</ul>`
|
|
125
125
|
: ''
|
|
126
|
-
const subsystems =
|
|
127
|
-
brief.subsystems.length > 0
|
|
128
|
-
? `<div class="brief-subsystems">${brief.subsystems
|
|
129
|
-
.map((s) => `<span class="brief-chip" title="${esc(s.role)}">${esc(s.name)}</span>`)
|
|
130
|
-
.join('')}</div>`
|
|
131
|
-
: ''
|
|
132
126
|
const issueLinks =
|
|
133
127
|
issues.length > 0
|
|
134
128
|
? `<div class="brief-issues">${issues
|
|
@@ -143,7 +137,6 @@ function briefBlock(brief: PrBrief | undefined, issues: BriefIssue[]): string {
|
|
|
143
137
|
<div class="brief-body">
|
|
144
138
|
<p class="brief-purpose">${esc(brief.purpose)}</p>
|
|
145
139
|
${changes}
|
|
146
|
-
${subsystems}
|
|
147
140
|
${issueLinks}
|
|
148
141
|
</div>
|
|
149
142
|
</details>`
|
|
@@ -26,7 +26,7 @@ const STAGE_HINT: Record<StageId, string> = {
|
|
|
26
26
|
|
|
27
27
|
const STAGE_NOW_DOING: Record<StageId, string> = {
|
|
28
28
|
setup: 'Fetching the PR and diff',
|
|
29
|
-
context: '
|
|
29
|
+
context: 'Reading the PR to brief the reviewers',
|
|
30
30
|
specialists: 'Five reviewers reading the diff in parallel',
|
|
31
31
|
dedupe: 'Merging overlapping findings',
|
|
32
32
|
critic: 'Keeping only the high-signal ones',
|
|
@@ -333,16 +333,10 @@ export function isSuggestion(f: ReviewFinding): boolean {
|
|
|
333
333
|
return f.risk.action === 'consider' || f.risk.action === 'optional'
|
|
334
334
|
}
|
|
335
335
|
|
|
336
|
-
export type BriefSubsystem = {
|
|
337
|
-
name: string
|
|
338
|
-
role: string
|
|
339
|
-
}
|
|
340
|
-
|
|
341
336
|
/** Scout-produced PR summary. Written to `$RUN_DIR/brief.json` by the context stage. */
|
|
342
337
|
export type PrBrief = {
|
|
343
338
|
purpose: string
|
|
344
339
|
changes: string[]
|
|
345
|
-
subsystems: BriefSubsystem[]
|
|
346
340
|
watchItems: string[]
|
|
347
341
|
unclear: string[]
|
|
348
342
|
}
|
|
@@ -365,19 +359,9 @@ export function parseBrief(raw: unknown): PrBrief | null {
|
|
|
365
359
|
const r = raw as Record<string, unknown>
|
|
366
360
|
const purpose = typeof r.purpose === 'string' ? r.purpose.trim() : ''
|
|
367
361
|
if (purpose.length === 0) return null
|
|
368
|
-
const subsystems: BriefSubsystem[] = Array.isArray(r.subsystems)
|
|
369
|
-
? (r.subsystems as unknown[]).flatMap((entry) => {
|
|
370
|
-
if (!entry || typeof entry !== 'object' || Array.isArray(entry)) return []
|
|
371
|
-
const e = entry as Record<string, unknown>
|
|
372
|
-
const name = typeof e.name === 'string' ? e.name.trim() : ''
|
|
373
|
-
if (name.length === 0) return []
|
|
374
|
-
return [{ name, role: typeof e.role === 'string' ? e.role.trim() : '' }]
|
|
375
|
-
})
|
|
376
|
-
: []
|
|
377
362
|
return {
|
|
378
363
|
purpose,
|
|
379
364
|
changes: briefStrings(r.changes),
|
|
380
|
-
subsystems,
|
|
381
365
|
watchItems: briefStrings(r.watchItems),
|
|
382
366
|
unclear: briefStrings(r.unclear),
|
|
383
367
|
}
|
package/skills/magpie/skill.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "magpie",
|
|
3
|
-
"version": "0.
|
|
3
|
+
"version": "0.12.0",
|
|
4
4
|
"description": "Interactive PR review pipeline. Runs five parallel specialist subagents (security, bugs, performance, code-smells, architecture), dedupes findings, applies a critic rubric, peer-reviews via codex exec (falling back to a Claude second opinion when codex is unavailable), and serves an interactive HTML report for selecting findings to post via gh. Bundles a Bun CLI installed onto PATH via the skill's postinstall step. Use when the user asks to review a GitHub pull request. Splits oversized diffs into budgeted shards and rebuilds the diff from the local clone when gh pr diff refuses it.",
|
|
5
5
|
"author": "iceinvein",
|
|
6
6
|
"type": "prompt",
|
|
@@ -1681,7 +1681,6 @@ code.inline-code {
|
|
|
1681
1681
|
color: var(--ink-muted);
|
|
1682
1682
|
}
|
|
1683
1683
|
|
|
1684
|
-
.brief-subsystems,
|
|
1685
1684
|
.brief-issues {
|
|
1686
1685
|
display: flex;
|
|
1687
1686
|
flex-wrap: wrap;
|
|
@@ -1689,7 +1688,6 @@ code.inline-code {
|
|
|
1689
1688
|
margin-top: 10px;
|
|
1690
1689
|
}
|
|
1691
1690
|
|
|
1692
|
-
.brief-chip,
|
|
1693
1691
|
.brief-issue {
|
|
1694
1692
|
padding: 2px 8px;
|
|
1695
1693
|
border: 1px solid var(--line);
|
|
@@ -1,213 +0,0 @@
|
|
|
1
|
-
#!/usr/bin/env bash
|
|
2
|
-
# The base repo has never been indexed, so the code-intel probe answers
|
|
3
|
-
# consent_required. Approving it would start a full GPU pass nobody asked for:
|
|
4
|
-
# the run has to record the tool as unavailable, say so once, and carry on.
|
|
5
|
-
#
|
|
6
|
-
# The run directory sits under the workspace rather than ~/.magpie because file
|
|
7
|
-
# graders refuse to follow a link out of the workspace. `magpie --list-runs` is
|
|
8
|
-
# what names the path a resume uses, so the shim reports this one.
|
|
9
|
-
set -euo pipefail
|
|
10
|
-
|
|
11
|
-
RUN_ID="pr-1337-1789600000"
|
|
12
|
-
RUN_DIR="$PWD/runs/$RUN_ID"
|
|
13
|
-
CALLS="$PWD/.magpie-calls.log"
|
|
14
|
-
|
|
15
|
-
mkdir -p "$RUN_DIR"/findings "$RUN_DIR"/state "$HOME/shims"
|
|
16
|
-
: > "$CALLS"
|
|
17
|
-
|
|
18
|
-
cat > "$HOME/shim-config" <<EOF
|
|
19
|
-
RUN_ID="$RUN_ID"
|
|
20
|
-
RUN_DIR="$RUN_DIR"
|
|
21
|
-
CALLS="$CALLS"
|
|
22
|
-
PORT=4599
|
|
23
|
-
EOF
|
|
24
|
-
|
|
25
|
-
# The real magpie and code-intel are outside the eval sandbox and cannot be
|
|
26
|
-
# executed from inside it, so the child gets fakes rather than exit 126. The
|
|
27
|
-
# code-intel fake answers consent_required and records an approve that should
|
|
28
|
-
# never come.
|
|
29
|
-
mkdir -p "$HOME/tmp"
|
|
30
|
-
cat > "$HOME/.zshenv" <<'RC'
|
|
31
|
-
export PATH="$HOME/shims:/usr/bin:/bin:/usr/sbin:/sbin"
|
|
32
|
-
# /usr/bin/python3 is the Xcode shim, and without a writable TMPDIR it fails
|
|
33
|
-
# trying to create its xcrun cache in a directory the sandbox blocks.
|
|
34
|
-
export TMPDIR="$HOME/tmp"
|
|
35
|
-
RC
|
|
36
|
-
|
|
37
|
-
cat > "$HOME/shims/magpie" <<'SHIM'
|
|
38
|
-
#!/usr/bin/env bash
|
|
39
|
-
. "$HOME/shim-config"
|
|
40
|
-
echo "magpie $*" >> "$CALLS"
|
|
41
|
-
case "${1:-}" in
|
|
42
|
-
--list-runs) printf '%s\tactive\t%s\n' "$RUN_ID" "$RUN_DIR" ;;
|
|
43
|
-
status)
|
|
44
|
-
python3 - "${2:-$RUN_DIR}" <<'STATUS'
|
|
45
|
-
import json, pathlib, sys
|
|
46
|
-
|
|
47
|
-
ORDER = ['setup', 'context', 'specialists', 'dedupe', 'critic', 'peer-review', 'report', 'post']
|
|
48
|
-
last, error = None, None
|
|
49
|
-
for line in (pathlib.Path(sys.argv[1]) / 'log.jsonl').read_text().splitlines():
|
|
50
|
-
if not line.strip():
|
|
51
|
-
continue
|
|
52
|
-
try:
|
|
53
|
-
entry = json.loads(line)
|
|
54
|
-
except ValueError:
|
|
55
|
-
continue
|
|
56
|
-
if entry.get('status') == 'error':
|
|
57
|
-
error = entry.get('stage')
|
|
58
|
-
break
|
|
59
|
-
if entry.get('status') in ('done', 'skipped') and entry.get('stage') in ORDER:
|
|
60
|
-
last = entry['stage']
|
|
61
|
-
index = ORDER.index(last) + 1 if last else 0
|
|
62
|
-
print(json.dumps({'lastCompleted': last, 'next': ORDER[index] if index < len(ORDER) else 'cleanup', 'error': error}))
|
|
63
|
-
STATUS
|
|
64
|
-
;;
|
|
65
|
-
serve)
|
|
66
|
-
mkdir -p "$RUN_DIR/screen" "$RUN_DIR/state"
|
|
67
|
-
# The eval sandbox refuses listening sockets, so no fake can hold a port
|
|
68
|
-
# open: this writes the server-info the walkthrough reads and exits. The
|
|
69
|
-
# page is never reachable in a case, so no case pins the browser surface.
|
|
70
|
-
echo "http://127.0.0.1:$PORT" > "$RUN_DIR/state/server-info"
|
|
71
|
-
echo "serving $RUN_DIR on http://127.0.0.1:$PORT"
|
|
72
|
-
;;
|
|
73
|
-
render)
|
|
74
|
-
mkdir -p "$RUN_DIR/screen"
|
|
75
|
-
python3 - "${2:-$RUN_DIR}" "${3:-progress}" <<'RENDER'
|
|
76
|
-
import json, pathlib, sys
|
|
77
|
-
|
|
78
|
-
run, screen = pathlib.Path(sys.argv[1]), sys.argv[2]
|
|
79
|
-
findings = run / 'findings.final.json'
|
|
80
|
-
rows = ''
|
|
81
|
-
if screen == 'findings' and findings.exists():
|
|
82
|
-
for finding in json.loads(findings.read_text()):
|
|
83
|
-
rows += f'<li><input type="checkbox" data-finding-id="{finding["id"]}"> {finding["id"]}: {finding["title"]}</li>'
|
|
84
|
-
buttons = '<button>Post Selected</button><button>Post Recommended</button>' if rows else ''
|
|
85
|
-
(run / 'screen').mkdir(exist_ok=True)
|
|
86
|
-
(run / 'screen' / f'{screen}.html').write_text(
|
|
87
|
-
f'<html><body><h1>magpie {screen}</h1><ul>{rows}</ul>{buttons}</body></html>'
|
|
88
|
-
)
|
|
89
|
-
RENDER
|
|
90
|
-
echo "rendered ${3:-progress} -> $RUN_DIR/screen/${3:-progress}.html"
|
|
91
|
-
;;
|
|
92
|
-
*) echo "fake magpie: unsupported subcommand: $*" >&2; exit 64 ;;
|
|
93
|
-
esac
|
|
94
|
-
SHIM
|
|
95
|
-
chmod +x "$HOME/shims/magpie"
|
|
96
|
-
|
|
97
|
-
cat > "$HOME/shims/code-intel" <<'SHIM'
|
|
98
|
-
#!/usr/bin/env bash
|
|
99
|
-
. "$HOME/shim-config"
|
|
100
|
-
echo "code-intel $*" >> "$CALLS"
|
|
101
|
-
case "$1 ${2:-}" in
|
|
102
|
-
"index status") printf '{"status":"consent_required","reason":"this repository has never completed an index"}\n' ;;
|
|
103
|
-
"index approve") echo "indexing approved"; ;;
|
|
104
|
-
"start "*|"start") echo "daemon already running" ;;
|
|
105
|
-
*) echo "fake code-intel: unsupported command: $*" >&2; exit 64 ;;
|
|
106
|
-
esac
|
|
107
|
-
SHIM
|
|
108
|
-
chmod +x "$HOME/shims/code-intel"
|
|
109
|
-
|
|
110
|
-
cat > "$RUN_DIR/pr.json" <<'JSON'
|
|
111
|
-
{
|
|
112
|
-
"number": 1337,
|
|
113
|
-
"title": "Cache tenant settings in the request path",
|
|
114
|
-
"author": { "login": "asha-platform" },
|
|
115
|
-
"headRefName": "feat/tenant-settings-cache",
|
|
116
|
-
"baseRefName": "main",
|
|
117
|
-
"headRefOid": "9f3a8c0211dbb5fe7a82a2c1b08e0a45c2d1ee01",
|
|
118
|
-
"url": "https://github.com/example/repo/pull/1337"
|
|
119
|
-
}
|
|
120
|
-
JSON
|
|
121
|
-
|
|
122
|
-
cat > "$RUN_DIR/log.jsonl" <<'LOG'
|
|
123
|
-
{"stage":"preflight","status":"done","missingOptional":["codex"]}
|
|
124
|
-
{"stage":"setup","status":"done"}
|
|
125
|
-
LOG
|
|
126
|
-
|
|
127
|
-
# The PR under review, as setup would have left it: the filtered diff, and a
|
|
128
|
-
# worktree holding the head state the diff produces. The hunk headers count the
|
|
129
|
-
# lines they carry, and every finding below cites a line inside a hunk, so
|
|
130
|
-
# nothing here contradicts anything else.
|
|
131
|
-
cat > "$RUN_DIR/diff.patch" <<'PATCH'
|
|
132
|
-
diff --git a/src/settings/cache.ts b/src/settings/cache.ts
|
|
133
|
-
--- a/src/settings/cache.ts
|
|
134
|
-
+++ b/src/settings/cache.ts
|
|
135
|
-
@@ -1,5 +1,13 @@
|
|
136
|
-
const store = new Map<string, Settings>()
|
|
137
|
-
|
|
138
|
-
+export function put(tenantId: string, settings: Settings) {
|
|
139
|
-
+ store.set(tenantId, settings)
|
|
140
|
-
+}
|
|
141
|
-
+
|
|
142
|
-
+export function get(tenantId: string): Settings | undefined {
|
|
143
|
-
+ return store.get(tenantId)
|
|
144
|
-
+}
|
|
145
|
-
+
|
|
146
|
-
export function clear() {
|
|
147
|
-
store.clear()
|
|
148
|
-
}
|
|
149
|
-
diff --git a/src/settings/loader.ts b/src/settings/loader.ts
|
|
150
|
-
--- a/src/settings/loader.ts
|
|
151
|
-
+++ b/src/settings/loader.ts
|
|
152
|
-
@@ -9,3 +9,7 @@
|
|
153
|
-
export async function load(tenantId: string) {
|
|
154
|
-
- return fetchSettings(tenantId)
|
|
155
|
-
+ const hit = get(tenantId)
|
|
156
|
-
+ if (hit) return hit
|
|
157
|
-
+ const fresh = await fetchSettings(tenantId)
|
|
158
|
-
+ put(tenantId, fresh)
|
|
159
|
-
+ return fresh
|
|
160
|
-
}
|
|
161
|
-
PATCH
|
|
162
|
-
|
|
163
|
-
mkdir -p "$RUN_DIR/worktree/src/settings"
|
|
164
|
-
|
|
165
|
-
cat > "$RUN_DIR/worktree/src/settings/cache.ts" <<'TS'
|
|
166
|
-
const store = new Map<string, Settings>()
|
|
167
|
-
|
|
168
|
-
export function put(tenantId: string, settings: Settings) {
|
|
169
|
-
store.set(tenantId, settings)
|
|
170
|
-
}
|
|
171
|
-
|
|
172
|
-
export function get(tenantId: string): Settings | undefined {
|
|
173
|
-
return store.get(tenantId)
|
|
174
|
-
}
|
|
175
|
-
|
|
176
|
-
export function clear() {
|
|
177
|
-
store.clear()
|
|
178
|
-
}
|
|
179
|
-
TS
|
|
180
|
-
|
|
181
|
-
cat > "$RUN_DIR/worktree/src/settings/loader.ts" <<'TS'
|
|
182
|
-
import { get, put } from './cache'
|
|
183
|
-
|
|
184
|
-
type Settings = { theme: string }
|
|
185
|
-
|
|
186
|
-
async function fetchSettings(tenantId: string): Promise<Settings> {
|
|
187
|
-
return { theme: 'default' }
|
|
188
|
-
}
|
|
189
|
-
|
|
190
|
-
export async function load(tenantId: string) {
|
|
191
|
-
const hit = get(tenantId)
|
|
192
|
-
if (hit) return hit
|
|
193
|
-
const fresh = await fetchSettings(tenantId)
|
|
194
|
-
put(tenantId, fresh)
|
|
195
|
-
return fresh
|
|
196
|
-
}
|
|
197
|
-
TS
|
|
198
|
-
|
|
199
|
-
echo '[]' > "$RUN_DIR/findings/tests.json"
|
|
200
|
-
echo "http://127.0.0.1:4599" > "$RUN_DIR/state/server-info"
|
|
201
|
-
|
|
202
|
-
mkdir -p "$RUN_DIR/shards"
|
|
203
|
-
cat > "$RUN_DIR/shards/manifest.json" <<'JSON'
|
|
204
|
-
{
|
|
205
|
-
"budget": 6000,
|
|
206
|
-
"maxFiles": 80,
|
|
207
|
-
"totalFiles": 1,
|
|
208
|
-
"totalLines": 12,
|
|
209
|
-
"shards": [
|
|
210
|
-
{ "id": 1, "path": "diff.patch", "files": ["src/settings/cache.ts"], "lines": 12 }
|
|
211
|
-
]
|
|
212
|
-
}
|
|
213
|
-
JSON
|
|
@@ -1,10 +0,0 @@
|
|
|
1
|
-
---
|
|
2
|
-
type: llm
|
|
3
|
-
focus: last_message
|
|
4
|
-
weight: 2
|
|
5
|
-
---
|
|
6
|
-
|
|
7
|
-
The code intelligence probe answered `consent_required`, which means the tool is unavailable for this run.
|
|
8
|
-
|
|
9
|
-
PASS if the reply tells the user that code intelligence is unavailable and that the specialists will therefore work from the diff alone. Any wording will do, and naming the reason (the repo has never been indexed, consent is required) is fine but not required.
|
|
10
|
-
FAIL if the reply claims code intelligence is available or working, says nothing about it, or says it approved or started an index.
|
|
@@ -1,11 +0,0 @@
|
|
|
1
|
-
---
|
|
2
|
-
name: consent-required-never-approves
|
|
3
|
-
description: A consent_required probe means the tool is unavailable, not that the run should approve a full index on the user's behalf.
|
|
4
|
-
tags: [magpie, context, code-intelligence, scaffold]
|
|
5
|
-
max_turns: 30
|
|
6
|
-
timeout_seconds: 1200
|
|
7
|
-
allowed_tools: [Read, Glob, Grep, Skill, Write, Edit, Bash, Task]
|
|
8
|
-
expected_outcome: Probes once, never runs index approve, prints the one-line unavailable notice, closes the context stage and stops there as asked.
|
|
9
|
-
---
|
|
10
|
-
|
|
11
|
-
Pick the magpie run on PR 1337 back up. It stopped right after setup. Get it through the context stage and hand it back to me there, I want to kick the specialists off myself later.
|