@fyeeme/pi-review 2.0.0 → 2.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +78 -132
- package/agents/cleaner-altitude.md +18 -0
- package/agents/cleaner-efficiency.md +19 -0
- package/agents/cleaner-reuse.md +16 -0
- package/agents/cleaner-simplification.md +16 -0
- package/agents/finder-conventions.md +23 -0
- package/agents/finder-cross-file.md +21 -0
- package/agents/finder-diff-scan.md +23 -0
- package/agents/finder-language-pitfall.md +21 -0
- package/agents/finder-removed-behavior.md +21 -0
- package/agents/finder-wrapper-proxy.md +23 -0
- package/agents/gap-hunter.md +24 -0
- package/agents/verifier.md +33 -0
- package/index.ts +34 -29
- package/package.json +17 -15
- package/prompts/review.parallel.md +23 -0
- package/prompts/review.single.md +16 -0
- package/prompts/simplify.parallel.md +45 -0
- package/prompts/simplify.single.md +23 -0
- package/skills/code-review/SKILL.md +292 -97
- package/skills/{simplify → code-simplify}/SKILL.md +93 -36
- package/src/config.ts +109 -0
- package/src/diff.ts +306 -0
- package/src/dispatch.ts +301 -0
- package/src/loop.ts +266 -0
- package/src/strategy.ts +76 -0
- package/src/tools/review_report.ts +36 -20
- package/src/commands/code-review.ts +0 -100
- package/src/commands/code-simplify.ts +0 -100
- package/src/tools/subagent.ts +0 -367
|
@@ -1,16 +1,58 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: code-review
|
|
3
|
-
description: "Review the current diff for correctness bugs and reuse/simplification/efficiency cleanups at the given effort level. Fresh reverse of CC `/code-review`
|
|
3
|
+
description: "Review the current diff, or a PR number/branch/path target, for correctness bugs and reuse/simplification/efficiency cleanups at the given effort level (low/medium/high: single-pass in-session review — medium precision, high recall; xhigh/max: subagent fan-out deep sweep). Fresh reverse of CC `/review` (its own name there is `code-review`), re-verified against CLI v2.1.261 (2026-09-05; originally reversed from v2.1.223). Effort semantics: medium = precision, high+ = recall. Pass --fix to apply, --loop to cycle fix→re-review until no P0/P1 findings remain, --comment to post findings (GitHub inline / GitLab MR note), --share to publish a review page."
|
|
4
4
|
---
|
|
5
5
|
|
|
6
6
|
<!--
|
|
7
|
-
Origin: Claude Code built-in skill `/
|
|
7
|
+
Origin: Claude Code built-in skill `/review` (CLI v2.1.223), freshly
|
|
8
8
|
reverse-engineered 2026-08-06 from bin/claude.exe strings. This file is
|
|
9
9
|
sourced DIRECTLY from the 2.1.223 binary — NOT carried forward from the
|
|
10
10
|
earlier v2.1.220 reconstruction. Every section below was located in the
|
|
11
11
|
extracted strings (cc_strings_223.txt) and verified.
|
|
12
12
|
|
|
13
|
-
|
|
13
|
+
── v2.1 redesign: effort split (single-pass default) ──
|
|
14
|
+
- low/medium/high → SINGLE-PASS FLOW in the main session (no subagents):
|
|
15
|
+
rubric-ported flag criteria, P0–P3 priorities, in-session self-verify.
|
|
16
|
+
Rationale: the medium+ fan-out pipeline (8–10 finder subprocesses +
|
|
17
|
+
grouped verifiers) cost tens of minutes per run and returned zero
|
|
18
|
+
findings when spawned subprocesses failed to boot — unacceptable ROI
|
|
19
|
+
for the default path.
|
|
20
|
+
- xhigh/max keep the fan-out pipeline unchanged (opt-in deep sweep).
|
|
21
|
+
- NEW --loop: extension-driven fix→re-review rounds (≤ maxTurns.loop,
|
|
22
|
+
default 3) until no P0/P1 findings remain; blocking decisions read
|
|
23
|
+
the structured review_report JSON (never markdown scraping).
|
|
24
|
+
- review_report findings gained an optional `priority` (P0–P3).
|
|
25
|
+
|
|
26
|
+
── RE-VERIFIED against CLI v2.1.261 (bin/claude.exe raw bytes, 2026-09-05) ──
|
|
27
|
+
- The 2.1.217-era background Workflow (phases Scope/Find/Verify/Sweep/
|
|
28
|
+
Synthesize) is GONE — the phase prompts now live inline in the skill
|
|
29
|
+
and dispatch via the Agent tool; per the bundled changelog, high/
|
|
30
|
+
xhigh/max now run inside a background agent (a CC host capability Pi
|
|
31
|
+
has no counterpart for — Pi runs inline in the session).
|
|
32
|
+
- 2.1.261 ships flag-gated effort variants (e.g. an inline, dedup-only
|
|
33
|
+
xhigh WITHOUT verify, alongside the fan-out + 1-vote-verify shape).
|
|
34
|
+
This skill keeps the uniform linearization: medium+ = fan-out +
|
|
35
|
+
grouped verify; sweep at xhigh/max.
|
|
36
|
+
- Angle bodies A–E and Reuse/Simplification/Efficiency/Conventions are
|
|
37
|
+
byte-identical to what we carried; **Altitude** gained CC's
|
|
38
|
+
root-cause phrasing + "name that change" (synced below).
|
|
39
|
+
- NEW sweep focus list ("what the first pass tends to miss") — synced
|
|
40
|
+
into Phase 3 and agents/gap-hunter.md.
|
|
41
|
+
- --comment gained GitLab: ONE general MR note via `glab mr note`
|
|
42
|
+
(glab has no single verb for line-anchored comments); GitHub inline
|
|
43
|
+
falls back to `gh api repos/{owner}/{repo}/pulls/{pr}/comments`,
|
|
44
|
+
suggestion block only when it fully fixes the issue — synced below.
|
|
45
|
+
- Output contract unchanged: {level, findings} with file/line/summary/
|
|
46
|
+
short_summary(≤60)/failure_scenario/category/verdict; outcome 三档
|
|
47
|
+
fixed / skipped / no_change_needed on re-report — all still match.
|
|
48
|
+
NEW: CC forbids creating/publishing an artifact of the review ("the
|
|
49
|
+
tool call is the report"); Pi keeps --share as the explicit opt-in
|
|
50
|
+
and forbids UNSOLICITED artifacts instead.
|
|
51
|
+
- `ultra` (deep multi-agent cloud review) still exists upstream; still
|
|
52
|
+
omitted here (requires claude.ai cloud access, which Pi lacks).
|
|
53
|
+
Sticky last-effort (codeReviewLastEffort) unchanged.
|
|
54
|
+
|
|
55
|
+
What CC 2.1.227 contained (historical basis, verified 2026-08-11):
|
|
14
56
|
- Effort quad tuple {correctnessAngles, perAngle, maxFindings, sweep}:
|
|
15
57
|
medium {3,6,8,false} / high {3,6,10,false} / xhigh {5,8,15,true} / max
|
|
16
58
|
same structure as xhigh. medium = precision; high+ = recall
|
|
@@ -32,8 +74,10 @@ description: "Review the current diff for correctness bugs and reuse/simplificat
|
|
|
32
74
|
- Fixed-later obligation (CC Q8m): later fixes in the session must
|
|
33
75
|
re-report findings with updated outcome.
|
|
34
76
|
|
|
35
|
-
Invocation: /code-review [low|medium|high|xhigh|max] [--fix] [--comment] [--share] [<target>]
|
|
77
|
+
Invocation: /code-review [low|medium|high|xhigh|max] [--fix] [--loop] [--comment] [--share] [<target>]
|
|
36
78
|
target = Class#method | file path | PR number | branch name
|
|
79
|
+
--loop = extension-driven fix→re-review rounds (single-pass levels only,
|
|
80
|
+
≤ maxTurns.loop, default 3) until no P0/P1 findings remain
|
|
37
81
|
With no level given, the /code-review HANDLER reuses the last level you
|
|
38
82
|
typed (CC 2.1.223 codeReviewLastEffort); the skill always receives a
|
|
39
83
|
concrete level.
|
|
@@ -53,20 +97,30 @@ description: "Review the current diff for correctness bugs and reuse/simplificat
|
|
|
53
97
|
Markdown as text.
|
|
54
98
|
2. Fan-out — CC uses the Agent tool; Pi uses the `subagent` tool
|
|
55
99
|
(mode: parallel), or runs angles sequentially if unavailable.
|
|
100
|
+
v2.1: fan-out is the XHIGH/MAX path only — low/medium/high
|
|
101
|
+
run as a single pass in the main session (no subagents).
|
|
56
102
|
3. Verify — CC uses the Agent tool; Pi uses `subagent` for the
|
|
57
103
|
independent verify agent (fallback: self-check).
|
|
58
|
-
4. Workflow — CC
|
|
59
|
-
Scope/Find/Verify/Sweep/Synthesize)
|
|
60
|
-
|
|
61
|
-
|
|
104
|
+
4. Workflow — CC 2.1.217 routed high/xhigh/max to a background Workflow
|
|
105
|
+
(phases Scope/Find/Verify/Sweep/Synthesize); 2.1.261 removed
|
|
106
|
+
it and runs the phases inline via its Agent tool, with high+
|
|
107
|
+
inside a background agent. Pi has neither host capability, so
|
|
108
|
+
this skill runs INLINE and linearizes those phases into the
|
|
109
|
+
flow below; the Phase 0.5 scope block is our absorption of
|
|
110
|
+
the old workflow's Scope phase (kept: it still turns N
|
|
111
|
+
repeated subagent discoveries into one).
|
|
62
112
|
5. ultra — dropped (cloud-only).
|
|
63
|
-
6. --share — CC uses the Artifact tool; Pi uses lavish-axi.
|
|
64
|
-
|
|
65
|
-
|
|
66
|
-
|
|
67
|
-
|
|
68
|
-
|
|
69
|
-
|
|
113
|
+
6. --share — CC uses the Artifact tool; Pi uses lavish-axi. CC 2.1.261
|
|
114
|
+
forbids UNSOLICITED review artifacts; --share stays the
|
|
115
|
+
explicit opt-in.
|
|
116
|
+
7. --comment— CC uses mcp__github_inline_comment (fallback `gh api`) and
|
|
117
|
+
posts GitLab MRs as one general note via `glab mr note`; Pi
|
|
118
|
+
mirrors both fallbacks (see the --comment section).
|
|
119
|
+
|
|
120
|
+
Prerequisite: the `subagent` tool (@fyeeme/pi-subagents; parallel mode) for
|
|
121
|
+
xhigh/max only (finder/verifier/gap-hunt fan-out). lavish-axi
|
|
122
|
+
for --share. low/medium/high run standalone in this session
|
|
123
|
+
(no subagents).
|
|
70
124
|
-->
|
|
71
125
|
|
|
72
126
|
You are reviewing the current diff for correctness bugs and reuse /
|
|
@@ -75,22 +129,30 @@ altitude, and conventions findings when the output cap forces a cut.
|
|
|
75
129
|
|
|
76
130
|
## Effort levels
|
|
77
131
|
|
|
78
|
-
| Level |
|
|
132
|
+
| Level | Path | Intent | Verify | Cap |
|
|
79
133
|
|-------|--------|--------|-----------|------------|
|
|
80
|
-
| low (default) | quick scan | no |
|
|
81
|
-
| medium | **precision** — surface only findings a maintainer would act on |
|
|
82
|
-
| high | **recall** — catch every real bug a careful reviewer would; **err on the side of surfacing** |
|
|
83
|
-
| xhigh | recall + **gap-hunt** |
|
|
84
|
-
| max |
|
|
134
|
+
| low (default) | SINGLE-PASS | quick scan | no | `min(files_changed, 4)` |
|
|
135
|
+
| medium | SINGLE-PASS | **precision** — surface only findings a maintainer would act on | self-verify (in-session) | 8 |
|
|
136
|
+
| high | SINGLE-PASS | **recall** — catch every real bug a careful reviewer would; **err on the side of surfacing** | self-verify (in-session) | 10 |
|
|
137
|
+
| xhigh | FAN-OUT (below) | recall + **gap-hunt** | independent verifier agents (grouped) | `{5, 8, 15, true}` |
|
|
138
|
+
| max | FAN-OUT(同 xhigh) | 同 xhigh | 同 xhigh | 同 xhigh |
|
|
139
|
+
|
|
140
|
+
**low/medium/high never dispatch subagents** — one pass in this session:
|
|
141
|
+
read the diff (Turn 1), surface candidates against the rubric (Turn 2),
|
|
142
|
+
self-verify them (Turn 3, medium/high only), report (Turn 4). This is the
|
|
143
|
+
default path: the 8–10 finder + grouped-verifier pipeline cost tens of
|
|
144
|
+
minutes per run and twice produced zero findings when spawned subprocesses
|
|
145
|
+
failed to boot — unacceptable ROI for a daily-driver review.
|
|
85
146
|
|
|
86
147
|
**max 与 xhigh 结构相同**:fan-out / verify / sweep 完全一致,差别仅在模型 reasoning effort(CC v2.1.226 注释实证:`max → same structure as xhigh (the API reasoning effort differs, not the fan-out)`)。若运行时不支持调节 reasoning effort,max 在结构上退化为 xhigh——不要因档名而期待更多 fan-out。
|
|
87
148
|
|
|
88
|
-
The quad tuple parameterizes the
|
|
149
|
+
The quad tuple parameterizes the XHIGH/MAX fan-out only (CC inline semantics,
|
|
150
|
+
verified 2.1.227):
|
|
89
151
|
|
|
90
|
-
- `correctnessAngles` —
|
|
91
|
-
- `perAngle` — candidate cap per finder (
|
|
92
|
-
- `maxFindings` — the report cap after verify (
|
|
93
|
-
- `sweep` — whether Phase 3 gap-hunt runs (
|
|
152
|
+
- `correctnessAngles` — 5 at xhigh/max (angles A–E all run).
|
|
153
|
+
- `perAngle` — candidate cap per finder (8).
|
|
154
|
+
- `maxFindings` — the report cap after verify (15).
|
|
155
|
+
- `sweep` — whether Phase 3 gap-hunt runs (≤ 8 new candidates).
|
|
94
156
|
|
|
95
157
|
Each finder surfaces up to `perAngle` candidate findings with `file`, `line`, a
|
|
96
158
|
one-line `summary`, a ≤60-char `short_summary`, and a concrete
|
|
@@ -141,62 +203,160 @@ actions based on it>
|
|
|
141
203
|
```
|
|
142
204
|
|
|
143
205
|
Embed this block verbatim at the top of **every** finder / verifier / gap-hunt
|
|
144
|
-
subagent prompt. Subagents do not re-discover the diff or
|
|
145
|
-
target argument travels as a scope constraint only, never as an
|
|
146
|
-
a subagent.
|
|
206
|
+
subagent prompt (XHIGH/MAX FLOW). Subagents do not re-discover the diff or
|
|
207
|
+
CLAUDE.md; the target argument travels as a scope constraint only, never as an
|
|
208
|
+
instruction to a subagent. In the SINGLE-PASS FLOW, keep the assembled block
|
|
209
|
+
as your own working notes — conventions come from it, not from re-discovery.
|
|
147
210
|
|
|
148
211
|
---
|
|
149
212
|
|
|
150
|
-
#
|
|
213
|
+
# SINGLE-PASS FLOW (default: low / medium / high — no subagents)
|
|
214
|
+
|
|
215
|
+
You review the diff yourself, in this session. Do NOT dispatch finder or
|
|
216
|
+
verifier agents at these levels, even if the `subagent` tool is available.
|
|
151
217
|
|
|
152
|
-
`low
|
|
218
|
+
- `low` — 1 diff pass, no self-verify, cap `min(files_changed, 4)`.
|
|
219
|
+
- `medium` — 1 pass + self-verify, cap 8, **precision**.
|
|
220
|
+
- `high` — 1 pass + self-verify, cap 10, **recall**.
|
|
153
221
|
|
|
154
222
|
## Turn 1 — read
|
|
155
223
|
|
|
156
224
|
One tool call: read the unified diff (`git diff @{upstream}...HEAD; git diff HEAD`
|
|
157
225
|
to cover both committed and uncommitted changes, or `git diff main...HEAD` / the
|
|
158
|
-
target passed as an argument).
|
|
159
|
-
`__tests__/`, `*_test.*`, `*.test.*`, `fixtures/`, `testdata/`) —
|
|
160
|
-
changes are not reviewed at
|
|
161
|
-
|
|
162
|
-
|
|
163
|
-
|
|
164
|
-
|
|
165
|
-
|
|
166
|
-
|
|
167
|
-
|
|
168
|
-
|
|
169
|
-
|
|
170
|
-
|
|
171
|
-
|
|
172
|
-
|
|
173
|
-
|
|
174
|
-
|
|
175
|
-
|
|
176
|
-
|
|
177
|
-
|
|
178
|
-
|
|
226
|
+
target passed as an argument). At low, skip test/fixture hunks (`test/`,
|
|
227
|
+
`spec/`, `__tests__/`, `*_test.*`, `*.test.*`, `fixtures/`, `testdata/`) —
|
|
228
|
+
test-file changes are not reviewed at that level; medium/high include them.
|
|
229
|
+
Then read the enclosing function for each nontrivial hunk; the applicable
|
|
230
|
+
CLAUDE.md conventions are already pinned in the Phase 0.5 scope block.
|
|
231
|
+
|
|
232
|
+
## Turn 2 — candidates (the rubric)
|
|
233
|
+
|
|
234
|
+
Work the finder angles inline — their definitions live in the XHIGH/MAX FLOW
|
|
235
|
+
below and are shared with the subagent definitions:
|
|
236
|
+
|
|
237
|
+
- **low** — Angle A over the hunks only: runtime-correctness bugs visible
|
|
238
|
+
from the hunk alone (inverted/wrong condition, off-by-one, null/undefined
|
|
239
|
+
deref where adjacent lines show the value can be absent, removed guard,
|
|
240
|
+
falsy-zero check, missing `await`, wrong-variable copy-paste, error
|
|
241
|
+
swallowed in a catch that should propagate), plus new code duplicating an
|
|
242
|
+
existing helper visible in the diff context, plus dead code the diff leaves
|
|
243
|
+
behind. Do **not** flag style, naming, perf, missing tests, or anything
|
|
244
|
+
outside the hunk. If you have fewer than the cap, do one more pass focused
|
|
245
|
+
on the largest changed file and on any **removed** code blocks. Output
|
|
246
|
+
exactly `(none)` only if the diff is trivially correct after that pass.
|
|
247
|
+
- **medium** — Angles A, B, C, then a quick Reuse / Simplification /
|
|
248
|
+
Efficiency pass over the changed code.
|
|
249
|
+
- **high** — the full angle set: A–E, then Reuse / Simplification /
|
|
250
|
+
Efficiency / Altitude / Conventions.
|
|
251
|
+
|
|
252
|
+
Flag issues that (rubric ported from the reference /review implementation):
|
|
253
|
+
|
|
254
|
+
1. Meaningfully impact the accuracy, performance, security, or
|
|
255
|
+
maintainability of the code.
|
|
256
|
+
2. Are discrete and actionable (not general issues or multiple combined
|
|
257
|
+
issues).
|
|
258
|
+
3. Don't demand rigor inconsistent with the rest of the codebase.
|
|
259
|
+
4. Were introduced in the changes being reviewed (not pre-existing bugs).
|
|
260
|
+
5. The author would likely fix if made aware of them.
|
|
261
|
+
6. Don't rely on unstated assumptions about the codebase or the author's
|
|
262
|
+
intent.
|
|
263
|
+
|
|
264
|
+
Every candidate carries `file`, `line`, `category`, a one-line `summary`, a
|
|
265
|
+
≤60-char `short_summary`, a concrete `failure_scenario`, and a **priority**
|
|
266
|
+
(`--loop` treats P0/P1 as blocking):
|
|
267
|
+
|
|
268
|
+
- **P0** — data loss, security hole, crash on a main path, broken build.
|
|
269
|
+
- **P1** — real bug on a plausible path; broken invariant with visible
|
|
270
|
+
effect.
|
|
271
|
+
- **P2** — worthwhile cleanup (duplication, wasted work, wrong altitude) or
|
|
272
|
+
an uncertain-trigger correctness issue.
|
|
273
|
+
- **P3** — nice-to-have.
|
|
274
|
+
|
|
275
|
+
Correctness outranks cleanup when the cap forces a cut.
|
|
276
|
+
|
|
277
|
+
## Turn 3 — self-verify (medium / high; low skips)
|
|
278
|
+
|
|
279
|
+
Re-read every candidate against the code once, in this session:
|
|
280
|
+
|
|
281
|
+
- Drop anything whose `failure_scenario` you cannot make concrete.
|
|
282
|
+
- Set the verdict: **`CONFIRMED`** — you can name the inputs/state that
|
|
283
|
+
trigger it and the wrong output or crash (quote the line); **`PLAUSIBLE`**
|
|
284
|
+
— the mechanism is real but the trigger is uncertain (timing, env,
|
|
285
|
+
config); state what would confirm it.
|
|
286
|
+
- **`PLAUSIBLE` by default** — do not drop a candidate for being
|
|
287
|
+
"speculative" or "depends on runtime state" when the state is realistic:
|
|
288
|
+
concurrency races, nil/undefined on a rare-but-reachable path (error
|
|
289
|
+
handler, cold cache, missing optional field), falsy-zero treated as
|
|
290
|
+
missing, off-by-one on a boundary the code does not exclude, retry storms
|
|
291
|
+
/ partial failures, regex/allowlist that lost an anchor.
|
|
292
|
+
- At medium (precision), additionally drop what a maintainer would not act
|
|
293
|
+
on. At high (recall), keep every surviving candidate — a missed bug ships.
|
|
294
|
+
|
|
295
|
+
## Turn 4 — report
|
|
296
|
+
|
|
297
|
+
Report via the `review_report` tool exactly as the Output section below
|
|
298
|
+
specifies, with `fanned_out: false` (honesty: this was a single-pass
|
|
299
|
+
self-review). At low the candidates ARE the findings (unverified — leave
|
|
300
|
+
`verdict` unset so the reader can discount them); if the `review_report` tool
|
|
301
|
+
is unavailable, print the findings as text (one line per finding:
|
|
302
|
+
`path/to/file.ext:123 — 问题与失败后果`), `(none)` when empty.
|
|
303
|
+
|
|
304
|
+
## Loop fixing (--loop)
|
|
305
|
+
|
|
306
|
+
When the trigger message says loop fixing is armed, the extension takes over
|
|
307
|
+
after your report: it reads the newest `review_report` JSON under
|
|
308
|
+
`.pi/review/`, and while P0/P1 findings remain it sends a fix prompt (apply
|
|
309
|
+
them per the --fix section's rules), waits, then asks you to re-run this
|
|
310
|
+
single-pass flow. Treat each re-review as a fresh pass with a fresh
|
|
311
|
+
`report_id` and an honest fresh findings list — do not rubber-stamp the
|
|
312
|
+
previous run.
|
|
179
313
|
|
|
180
314
|
---
|
|
181
315
|
|
|
182
|
-
#
|
|
183
|
-
|
|
184
|
-
|
|
185
|
-
|
|
186
|
-
|
|
187
|
-
|
|
188
|
-
|
|
189
|
-
|
|
190
|
-
|
|
191
|
-
|
|
192
|
-
|
|
193
|
-
|
|
194
|
-
|
|
195
|
-
|
|
196
|
-
|
|
197
|
-
|
|
198
|
-
|
|
199
|
-
|
|
316
|
+
# XHIGH/MAX FLOW (deep sweep: fan-out + verify)
|
|
317
|
+
|
|
318
|
+
Reached only at effort xhigh/max — low/medium/high use the SINGLE-PASS FLOW
|
|
319
|
+
above. Launch finder agents through the `subagent` tool in a single batch
|
|
320
|
+
(mode: parallel) so they run concurrently; if it is unavailable, do not fake
|
|
321
|
+
the fan-out — work the angles yourself in sequence in this same context, or
|
|
322
|
+
report that the subagent capability is unavailable.
|
|
323
|
+
|
|
324
|
+
**Checking `subagent` availability** — wherever this skill says "if the
|
|
325
|
+
`subagent` tool is available", decide from THIS session's tool list, never by
|
|
326
|
+
probing: `subagent` is a pi extension tool registered alongside
|
|
327
|
+
read/bash/edit, not an MCP server tool, so the `mcp` gateway's tool search
|
|
328
|
+
answers "No tools matching subagent" even when the tool is registered and
|
|
329
|
+
callable. If it is in your toolset, use it without further verification; if it
|
|
330
|
+
is genuinely absent, take the sequential fallback above.
|
|
331
|
+
|
|
332
|
+
**Finder turn budget(Pi adaptation — the same runaway-exploration guard the
|
|
333
|
+
Phase 3 gap-hunt already carries)** — a finder that exhausts its turn cap
|
|
334
|
+
mid-read returns NOTHING and silently loses its whole angle (observed on a
|
|
335
|
+
168-file diff: 7/10 finders burned their full turn budget with zero output,
|
|
336
|
+
and the coverage hole cascaded into two extra compensation waves). Constrain
|
|
337
|
+
every finder batch:
|
|
338
|
+
|
|
339
|
+
1. **Set `maxTurns: 20` on the `subagent` call** — the slowest finder pins
|
|
340
|
+
the wave's wall time; 20 turns covers the highest-risk hunks of any
|
|
341
|
+
single angle, and a capped finder still owes partial output (next item).
|
|
342
|
+
(20 is the built-in default — if the trigger message states a different
|
|
343
|
+
finder budget, use that instead.)
|
|
344
|
+
2. **Declare the budget inside each finder prompt** — e.g. "You have ~15
|
|
345
|
+
tool calls. Spend them on the highest-risk hunks first; when half are
|
|
346
|
+
spent, stop opening new files."
|
|
347
|
+
3. **Final-message contract** — the finder's LAST assistant message must be
|
|
348
|
+
its JSON candidate array (an empty `[]` is a valid answer). Partial
|
|
349
|
+
output beats none: candidates that never reach text never reach verify.
|
|
350
|
+
4. **A finder that hits max-turns with no JSON is a FAILED finder**, not an
|
|
351
|
+
empty angle: re-dispatch that single angle on a narrower file slice
|
|
352
|
+
before Phase 2 (or fold it into the xhigh/max gap-hunt), and note the
|
|
353
|
+
re-dispatch in the report.
|
|
354
|
+
|
|
355
|
+
**Finder allocation** (CC inline, verified 2.1.227): xhigh/max run all five
|
|
356
|
+
correctness angles — **10 finders**: A, B, C, D, E + one finder each for
|
|
357
|
+
Reuse, Simplification, Efficiency + one Altitude + one Conventions. The quad
|
|
358
|
+
tuple's angles are taken **in order A→E** (`slice(0, N)` — do not hand-pick
|
|
359
|
+
angles; that makes runs unreproducible).
|
|
200
360
|
|
|
201
361
|
Each cleanup angle (Reuse / Simplification / Efficiency) gets its own finder;
|
|
202
362
|
Altitude and Conventions are independent finders. Never silently drop an
|
|
@@ -272,10 +432,11 @@ alternative.
|
|
|
272
432
|
|
|
273
433
|
### Altitude
|
|
274
434
|
|
|
275
|
-
Check that each change
|
|
276
|
-
bandaid. Special cases layered on shared
|
|
277
|
-
isn't deep enough — prefer
|
|
278
|
-
special cases
|
|
435
|
+
Check that each change fixes the root cause at the right depth rather than
|
|
436
|
+
patching a symptom with a fragile bandaid. Special cases layered on shared
|
|
437
|
+
infrastructure are a sign the fix isn't deep enough — prefer the simpler,
|
|
438
|
+
more general change to the underlying mechanism over adding special cases,
|
|
439
|
+
and name that change.
|
|
279
440
|
### Conventions (CLAUDE.md)
|
|
280
441
|
Find the CLAUDE.md files that govern the changed code: the user-level
|
|
281
442
|
~/.claude/CLAUDE.md, the repo-root CLAUDE.md, plus any CLAUDE.md or
|
|
@@ -303,7 +464,9 @@ Then verify each candidate **grouped by location**. If the `subagent` tool is
|
|
|
303
464
|
available: group the deduplicated candidates by `(file, line)`; dispatch ONE
|
|
304
465
|
independent verify agent per group (mode: parallel, one prompt per group),
|
|
305
466
|
giving it the scope block, the diff, the relevant file(s), and the full
|
|
306
|
-
candidate list for that location with each candidate's index.
|
|
467
|
+
candidate list for that location with each candidate's index. Set
|
|
468
|
+
`maxTurns: 15` on each verifier call (15 is the built-in default — a
|
|
469
|
+
different verifier budget stated in the trigger message wins). The verifier
|
|
307
470
|
returns a verdict per candidate:
|
|
308
471
|
|
|
309
472
|
```
|
|
@@ -340,11 +503,11 @@ optional field), falsy-zero treated as missing, off-by-one on a boundary the
|
|
|
340
503
|
code does not exclude, retry storms / partial failures, regex/allowlist that
|
|
341
504
|
lost an anchor. These are PLAUSIBLE.
|
|
342
505
|
|
|
343
|
-
**Recall bias
|
|
344
|
-
|
|
345
|
-
|
|
346
|
-
|
|
347
|
-
|
|
506
|
+
**Recall bias** — a single non-REFUTED verdict keeps the candidate: do NOT
|
|
507
|
+
drop it on uncertainty ("speculative", "depends on runtime state"). This
|
|
508
|
+
flow is the recall contract of xhigh/max — a missed bug ships, so err on the
|
|
509
|
+
side of surfacing hardest here. (Medium's precision filter lives in the
|
|
510
|
+
single-pass self-verify; it never reaches this flow.)
|
|
348
511
|
|
|
349
512
|
**REFUTED** only when constructible from the code: factually wrong (quote the
|
|
350
513
|
actual line); provably impossible (type/constant/invariant — show it); already
|
|
@@ -354,8 +517,13 @@ handled in this diff (cite the guard); or pure style with no observable effect.
|
|
|
354
517
|
|
|
355
518
|
At **xhigh and max**, after Phase 2 dedup, dispatch ONE fresh finder agent (the
|
|
356
519
|
`subagent` tool) that has never seen the candidates and hunts only for gaps not
|
|
357
|
-
already listed — **at most 8 new candidates**.
|
|
358
|
-
|
|
520
|
+
already listed — **at most 8 new candidates**. Focus the hunt on what the first
|
|
521
|
+
pass tends to miss (CC 2.1.261 sweep list): moved/extracted code that dropped a
|
|
522
|
+
guard or anchor; second-tier footguns (dataclass default evaluated once,
|
|
523
|
+
`hash()` non-determinism, lock-scope shrink, predicate methods with side
|
|
524
|
+
effects); setup/teardown asymmetry in tests; config defaults flipped. If
|
|
525
|
+
nothing new turns up, return an empty sweep — do not pad. Feed anything it
|
|
526
|
+
finds back through Phase 2 verify before keeping it.
|
|
359
527
|
|
|
360
528
|
Constrain it so exploration can't run away (Pi adaptation — CC's workflow bounds
|
|
361
529
|
this differently):
|
|
@@ -365,7 +533,9 @@ this differently):
|
|
|
365
533
|
results: the diff, the enclosing functions, the deduplicated finding list,
|
|
366
534
|
and any search results. The gap-hunt agent **analyzes**, it does not
|
|
367
535
|
**discover**.
|
|
368
|
-
2. **Set `maxTurns: 15`** on the `subagent` call — caps it at 15 assistant turns
|
|
536
|
+
2. **Set `maxTurns: 15`** on the `subagent` call — caps it at 15 assistant turns
|
|
537
|
+
(built-in default — a different gap-hunt budget stated in the trigger
|
|
538
|
+
message wins).
|
|
369
539
|
3. **Declare a tool-call budget in the prompt** — e.g. "You have ONLY 3 tool
|
|
370
540
|
calls to read files. Read them now, then analyze from this message's
|
|
371
541
|
context."
|
|
@@ -374,8 +544,6 @@ Feed anything it finds back through Phase 2 verify before keeping it. If the
|
|
|
374
544
|
`subagent` tool is unavailable, take one self-sweep instead and note the
|
|
375
545
|
gap-hunt was self-run (lacks the independent fresh-eyes benefit).
|
|
376
546
|
|
|
377
|
-
At **high and below**, skip Phase 3.
|
|
378
|
-
|
|
379
547
|
## Output
|
|
380
548
|
|
|
381
549
|
Report the findings via the `review_report` tool (this extension's counterpart
|
|
@@ -384,12 +552,18 @@ to CC's `ReportFindings`) — call it **once** with
|
|
|
384
552
|
ranked most-severe first (empty array if nothing survived verification). The
|
|
385
553
|
tool renders the Chinese Markdown report (table + details) back to the
|
|
386
554
|
conversation AND writes a machine-readable JSON to `<cwd>/.pi/review/` for CI /
|
|
387
|
-
`--fix` / `--comment`. Do **not** also hand-write the Markdown table.
|
|
555
|
+
`--fix` / `--comment`. Do **not** also hand-write the Markdown table. Also do
|
|
556
|
+
**not** spontaneously produce a review page/artifact when `--share` was not
|
|
557
|
+
passed — the `review_report` call IS the report (CC 2.1.261: "do not create
|
|
558
|
+
or publish an artifact of the review — the tool call is the report");
|
|
559
|
+
`--share` is the only explicit exception.
|
|
388
560
|
|
|
389
561
|
Each finding in the array carries: `file`, `line` (optional), `category`
|
|
390
562
|
(`correctness` / `reuse` / `simplification` / `efficiency` / `altitude` /
|
|
391
563
|
`conventions`, or a more specific slug like `test-coverage`), `verdict`
|
|
392
|
-
(`CONFIRMED` / `PLAUSIBLE`), `
|
|
564
|
+
(`CONFIRMED` / `PLAUSIBLE`), `priority` (`P0`–`P3`; single-pass levels
|
|
565
|
+
always set it — `--loop` treats P0/P1 as blocking; xhigh/max may omit it),
|
|
566
|
+
`short_summary` (≤60 字符、纯声明——去掉理由与
|
|
393
567
|
后果,汇总表概述列优先使用它;示例:`"off-by-one in loop bound"`),
|
|
394
568
|
`summary` (一行中文,含理由与后果,详情块使用), `failure_scenario`
|
|
395
569
|
(concrete input/state → wrong output/crash; for cleanup findings, the
|
|
@@ -416,7 +590,7 @@ the files by id.
|
|
|
416
590
|
`outcome` 作为标识符保留英文 token。
|
|
417
591
|
|
|
418
592
|
**`fanned_out` 诚实** — 准确设置:仅当多智能体 fan-out 真的跑起来(subagent
|
|
419
|
-
finder + verify agent)才为 `true`;low
|
|
593
|
+
finder + verify agent,xhigh/max)才为 `true`;low/medium/high 单遍或任何自审降级为 `false`。该
|
|
420
594
|
字段会出现在报告表头,让读者不被误导(替代旧的 Single-pass honesty 小节)。
|
|
421
595
|
|
|
422
596
|
**降级** — 若 `review_report` 工具未注册(这份 SKILL.md 跑在 pi-review 扩展之外),
|
|
@@ -426,36 +600,57 @@ finder + verify agent)才为 `true`;low effort 或任何单遍/自审降级
|
|
|
426
600
|
|
|
427
601
|
## Applying fixes (--fix)
|
|
428
602
|
|
|
429
|
-
The `--fix` flag was passed
|
|
603
|
+
The `--fix` flag was passed (the extension-driven `--loop` sends the same
|
|
604
|
+
fix prompts between re-review passes — follow them identically). After
|
|
605
|
+
producing the findings list, apply the
|
|
430
606
|
findings to the working tree instead of stopping at the report: fix each one
|
|
431
607
|
directly — correctness bugs and reuse/simplification/efficiency cleanups alike.
|
|
432
608
|
Skip any finding whose fix would change intended behavior, require changes well
|
|
433
609
|
outside the reviewed diff, or that you judge to be a false positive — note the
|
|
434
|
-
skip rather than arguing with it.
|
|
610
|
+
skip rather than arguing with it. If a verification command was detected (see
|
|
611
|
+
the trigger message's verification line), run it BEFORE re-reporting: a finding
|
|
612
|
+
whose fix breaks verification is reverted and re-reported as `skipped`
|
|
613
|
+
(verification is opportunistic — with no detected command, re-report directly
|
|
614
|
+
and say verification was not run). Then call `review_report` once more to
|
|
435
615
|
re-report (same `report_id`), setting `outcome` on each finding (`fixed` =
|
|
436
616
|
applied and verified / `skipped` = real but not applied, incl. reverted /
|
|
437
617
|
`no_change_needed` = not applicable or already handled). This structured
|
|
438
618
|
re-report replaces the hand-written summary and makes the fix result
|
|
439
|
-
machine-consumable.
|
|
619
|
+
machine-consumable. Make that call immediately after the fixes land, before
|
|
620
|
+
any prose summary (CC 2.1.261: the host UI's per-finding status updates only
|
|
621
|
+
from it).
|
|
440
622
|
If `review_report` is unavailable, fall back to a brief text summary of what was
|
|
441
623
|
fixed and what was skipped.
|
|
442
624
|
|
|
443
625
|
## Posting comments (--comment)
|
|
444
626
|
|
|
445
|
-
The `--comment` flag was passed.
|
|
446
|
-
|
|
447
|
-
|
|
448
|
-
|
|
627
|
+
The `--comment` flag was passed. After producing the findings list:
|
|
628
|
+
|
|
629
|
+
- **GitHub PR target** — post each finding as an inline PR comment on the
|
|
630
|
+
corresponding `file`/`line`, one call per finding; include a suggestion
|
|
631
|
+
block only when it fully fixes the issue (CC 2.1.261 rule). If no
|
|
632
|
+
inline-comment tool is available on Pi, fall back to `gh api
|
|
633
|
+
repos/{owner}/{repo}/pulls/{pr}/comments`; if `gh` is unavailable too,
|
|
634
|
+
print the findings as text and note that inline posting was unavailable.
|
|
635
|
+
- **GitLab MR target** — post the findings as ONE general MR note via
|
|
636
|
+
`glab mr note -m "<body>"` from inside the project's checkout — every
|
|
637
|
+
finding with its file:line, the issue, and the suggested fix (CC 2.1.261;
|
|
638
|
+
glab has no single verb for line-anchored comments, so post the general
|
|
639
|
+
note unless the user explicitly asks for inline threads — those need
|
|
640
|
+
`glab api projects/:id/merge_requests/:iid/discussions`). If `glab` is
|
|
641
|
+
unavailable, print the findings instead.
|
|
642
|
+
- **Not a PR/MR target** — print the findings to the terminal and note that
|
|
643
|
+
`--comment` was ignored.
|
|
449
644
|
|
|
450
645
|
## Publishing a shareable review (--share)
|
|
451
646
|
|
|
452
647
|
The `--share` flag was passed. After producing the findings list, also publish
|
|
453
648
|
them as an artifact so they can be shared and iterated on outside the terminal.
|
|
454
649
|
|
|
455
|
-
1. Write a self-contained HTML review page to `.lavish/
|
|
650
|
+
1. Write a self-contained HTML review page to `.lavish/review-<n>.html`
|
|
456
651
|
(create `.lavish/` in the repo root if missing). The page must render with no
|
|
457
652
|
server and carry every finding plus its context.
|
|
458
|
-
2. Open it with `lavish-axi .lavish/
|
|
653
|
+
2. Open it with `lavish-axi .lavish/review-<n>.html` so the reader can
|
|
459
654
|
review, annotate, and send feedback back through the poll.
|
|
460
655
|
|
|
461
656
|
Page structure (follow lavish design guidance — clear visual hierarchy, no
|