@fyeeme/pi-review 2.0.0 → 2.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1,16 +1,58 @@
1
1
  ---
2
2
  name: code-review
3
- description: "Review the current diff for correctness bugs and reuse/simplification/efficiency cleanups at the given effort level. Fresh reverse of CC `/code-review` (CLI v2.1.223). Effort semantics: medium = precision, high+ = recall. Pass --fix to apply, --comment to post inline PR comments, --share to publish a review page."
3
+ description: "Review the current diff, or a PR number/branch/path target, for correctness bugs and reuse/simplification/efficiency cleanups at the given effort level (low/medium/high: single-pass in-session review — medium precision, high recall; xhigh/max: subagent fan-out deep sweep). Fresh reverse of CC `/review` (its own name there is `code-review`), re-verified against CLI v2.1.261 (2026-09-05; originally reversed from v2.1.223). Effort semantics: medium = precision, high+ = recall. Pass --fix to apply, --loop to cycle fix→re-review until no P0/P1 findings remain, --comment to post findings (GitHub inline / GitLab MR note), --share to publish a review page."
4
4
  ---
5
5
 
6
6
  <!--
7
- Origin: Claude Code built-in skill `/code-review` (CLI v2.1.223), freshly
7
+ Origin: Claude Code built-in skill `/review` (CLI v2.1.223), freshly
8
8
  reverse-engineered 2026-08-06 from bin/claude.exe strings. This file is
9
9
  sourced DIRECTLY from the 2.1.223 binary — NOT carried forward from the
10
10
  earlier v2.1.220 reconstruction. Every section below was located in the
11
11
  extracted strings (cc_strings_223.txt) and verified.
12
12
 
13
- What CC 2.1.227 actually contains (verified against the binary):
13
+ ── v2.1 redesign: effort split (single-pass default) ──
14
+ - low/medium/high → SINGLE-PASS FLOW in the main session (no subagents):
15
+ rubric-ported flag criteria, P0–P3 priorities, in-session self-verify.
16
+ Rationale: the medium+ fan-out pipeline (8–10 finder subprocesses +
17
+ grouped verifiers) cost tens of minutes per run and returned zero
18
+ findings when spawned subprocesses failed to boot — unacceptable ROI
19
+ for the default path.
20
+ - xhigh/max keep the fan-out pipeline unchanged (opt-in deep sweep).
21
+ - NEW --loop: extension-driven fix→re-review rounds (≤ maxTurns.loop,
22
+ default 3) until no P0/P1 findings remain; blocking decisions read
23
+ the structured review_report JSON (never markdown scraping).
24
+ - review_report findings gained an optional `priority` (P0–P3).
25
+
26
+ ── RE-VERIFIED against CLI v2.1.261 (bin/claude.exe raw bytes, 2026-09-05) ──
27
+ - The 2.1.217-era background Workflow (phases Scope/Find/Verify/Sweep/
28
+ Synthesize) is GONE — the phase prompts now live inline in the skill
29
+ and dispatch via the Agent tool; per the bundled changelog, high/
30
+ xhigh/max now run inside a background agent (a CC host capability Pi
31
+ has no counterpart for — Pi runs inline in the session).
32
+ - 2.1.261 ships flag-gated effort variants (e.g. an inline, dedup-only
33
+ xhigh WITHOUT verify, alongside the fan-out + 1-vote-verify shape).
34
+ This skill keeps the uniform linearization: medium+ = fan-out +
35
+ grouped verify; sweep at xhigh/max.
36
+ - Angle bodies A–E and Reuse/Simplification/Efficiency/Conventions are
37
+ byte-identical to what we carried; **Altitude** gained CC's
38
+ root-cause phrasing + "name that change" (synced below).
39
+ - NEW sweep focus list ("what the first pass tends to miss") — synced
40
+ into Phase 3 and agents/gap-hunter.md.
41
+ - --comment gained GitLab: ONE general MR note via `glab mr note`
42
+ (glab has no single verb for line-anchored comments); GitHub inline
43
+ falls back to `gh api repos/{owner}/{repo}/pulls/{pr}/comments`,
44
+ suggestion block only when it fully fixes the issue — synced below.
45
+ - Output contract unchanged: {level, findings} with file/line/summary/
46
+ short_summary(≤60)/failure_scenario/category/verdict; outcome 三档
47
+ fixed / skipped / no_change_needed on re-report — all still match.
48
+ NEW: CC forbids creating/publishing an artifact of the review ("the
49
+ tool call is the report"); Pi keeps --share as the explicit opt-in
50
+ and forbids UNSOLICITED artifacts instead.
51
+ - `ultra` (deep multi-agent cloud review) still exists upstream; still
52
+ omitted here (requires claude.ai cloud access, which Pi lacks).
53
+ Sticky last-effort (codeReviewLastEffort) unchanged.
54
+
55
+ What CC 2.1.227 contained (historical basis, verified 2026-08-11):
14
56
  - Effort quad tuple {correctnessAngles, perAngle, maxFindings, sweep}:
15
57
  medium {3,6,8,false} / high {3,6,10,false} / xhigh {5,8,15,true} / max
16
58
  same structure as xhigh. medium = precision; high+ = recall
@@ -32,8 +74,10 @@ description: "Review the current diff for correctness bugs and reuse/simplificat
32
74
  - Fixed-later obligation (CC Q8m): later fixes in the session must
33
75
  re-report findings with updated outcome.
34
76
 
35
- Invocation: /code-review [low|medium|high|xhigh|max] [--fix] [--comment] [--share] [<target>]
77
+ Invocation: /code-review [low|medium|high|xhigh|max] [--fix] [--loop] [--comment] [--share] [<target>]
36
78
  target = Class#method | file path | PR number | branch name
79
+ --loop = extension-driven fix→re-review rounds (single-pass levels only,
80
+ ≤ maxTurns.loop, default 3) until no P0/P1 findings remain
37
81
  With no level given, the /code-review HANDLER reuses the last level you
38
82
  typed (CC 2.1.223 codeReviewLastEffort); the skill always receives a
39
83
  concrete level.
@@ -53,20 +97,30 @@ description: "Review the current diff for correctness bugs and reuse/simplificat
53
97
  Markdown as text.
54
98
  2. Fan-out — CC uses the Agent tool; Pi uses the `subagent` tool
55
99
  (mode: parallel), or runs angles sequentially if unavailable.
100
+ v2.1: fan-out is the XHIGH/MAX path only — low/medium/high
101
+ run as a single pass in the main session (no subagents).
56
102
  3. Verify — CC uses the Agent tool; Pi uses `subagent` for the
57
103
  independent verify agent (fallback: self-check).
58
- 4. Workflow — CC routes high/xhigh/max to a background Workflow (phases
59
- Scope/Find/Verify/Sweep/Synthesize) when workflows are
60
- enabled; Pi has no such tool, so this skill runs INLINE and
61
- linearizes those phases into the flow below.
104
+ 4. Workflow — CC 2.1.217 routed high/xhigh/max to a background Workflow
105
+ (phases Scope/Find/Verify/Sweep/Synthesize); 2.1.261 removed
106
+ it and runs the phases inline via its Agent tool, with high+
107
+ inside a background agent. Pi has neither host capability, so
108
+ this skill runs INLINE and linearizes those phases into the
109
+ flow below; the Phase 0.5 scope block is our absorption of
110
+ the old workflow's Scope phase (kept: it still turns N
111
+ repeated subagent discoveries into one).
62
112
  5. ultra — dropped (cloud-only).
63
- 6. --share — CC uses the Artifact tool; Pi uses lavish-axi.
64
- 7. --comment— CC uses mcp__github_inline_comment; Pi falls back to gh api
65
- or printing.
66
-
67
- Prerequisite: the `subagent` tool (pi-review extension; mode: parallel) for
68
- medium and above, and for the xhigh/max gap-hunter. lavish-axi
69
- for --share. low runs standalone (no subagents).
113
+ 6. --share — CC uses the Artifact tool; Pi uses lavish-axi. CC 2.1.261
114
+ forbids UNSOLICITED review artifacts; --share stays the
115
+ explicit opt-in.
116
+ 7. --comment— CC uses mcp__github_inline_comment (fallback `gh api`) and
117
+ posts GitLab MRs as one general note via `glab mr note`; Pi
118
+ mirrors both fallbacks (see the --comment section).
119
+
120
+ Prerequisite: the `subagent` tool (@fyeeme/pi-subagents; parallel mode) for
121
+ xhigh/max only (finder/verifier/gap-hunt fan-out). lavish-axi
122
+ for --share. low/medium/high run standalone in this session
123
+ (no subagents).
70
124
  -->
71
125
 
72
126
  You are reviewing the current diff for correctness bugs and reuse /
@@ -75,22 +129,30 @@ altitude, and conventions findings when the output cap forces a cut.
75
129
 
76
130
  ## Effort levels
77
131
 
78
- | Level | Intent | Verify | Subagents | 四元组 `{correctnessAngles, perAngle, maxFindings, sweep}` |
132
+ | Level | Path | Intent | Verify | Cap |
79
133
  |-------|--------|--------|-----------|------------|
80
- | low (default) | quick scan | no | no | 上限 `min(files_changed, 4)` |
81
- | medium | **precision** — surface only findings a maintainer would act on | independent verifier (grouped) | 8 finders | `{3, 6, 8, false}` |
82
- | high | **recall** — catch every real bug a careful reviewer would; **err on the side of surfacing** | recall-biased verifier (grouped) | 8 finders | `{3, 6, 10, false}` |
83
- | xhigh | recall + **gap-hunt** | recall-biased verifier (grouped) | 10 finders + 1 gap | `{5, 8, 15, true}` |
84
- | max | 同 xhigh | 同 xhigh | 同 xhigh | 同 xhigh |
134
+ | low (default) | SINGLE-PASS | quick scan | no | `min(files_changed, 4)` |
135
+ | medium | SINGLE-PASS | **precision** — surface only findings a maintainer would act on | self-verify (in-session) | 8 |
136
+ | high | SINGLE-PASS | **recall** — catch every real bug a careful reviewer would; **err on the side of surfacing** | self-verify (in-session) | 10 |
137
+ | xhigh | FAN-OUT (below) | recall + **gap-hunt** | independent verifier agents (grouped) | `{5, 8, 15, true}` |
138
+ | max | FAN-OUT(同 xhigh) | 同 xhigh | 同 xhigh | 同 xhigh |
139
+
140
+ **low/medium/high never dispatch subagents** — one pass in this session:
141
+ read the diff (Turn 1), surface candidates against the rubric (Turn 2),
142
+ self-verify them (Turn 3, medium/high only), report (Turn 4). This is the
143
+ default path: the 8–10 finder + grouped-verifier pipeline cost tens of
144
+ minutes per run and twice produced zero findings when spawned subprocesses
145
+ failed to boot — unacceptable ROI for a daily-driver review.
85
146
 
86
147
  **max 与 xhigh 结构相同**:fan-out / verify / sweep 完全一致,差别仅在模型 reasoning effort(CC v2.1.226 注释实证:`max → same structure as xhigh (the API reasoning effort differs, not the fan-out)`)。若运行时不支持调节 reasoning effort,max 在结构上退化为 xhigh——不要因档名而期待更多 fan-out。
87
148
 
88
- The quad tuple parameterizes the whole pipeline (CC inline semantics, verified 2.1.227):
149
+ The quad tuple parameterizes the XHIGH/MAX fan-out only (CC inline semantics,
150
+ verified 2.1.227):
89
151
 
90
- - `correctnessAngles` — how many correctness angles A–E run, taken **in order** (medium/high: A/B/C; xhigh/max: A–E).
91
- - `perAngle` — candidate cap per finder (6 at medium/high, 8 at xhigh/max).
92
- - `maxFindings` — the report cap after verify (8 / 10 / 15).
93
- - `sweep` — whether Phase 3 gap-hunt runs (xhigh/max only, ≤ 8 new candidates).
152
+ - `correctnessAngles` — 5 at xhigh/max (angles A–E all run).
153
+ - `perAngle` — candidate cap per finder (8).
154
+ - `maxFindings` — the report cap after verify (15).
155
+ - `sweep` — whether Phase 3 gap-hunt runs (≤ 8 new candidates).
94
156
 
95
157
  Each finder surfaces up to `perAngle` candidate findings with `file`, `line`, a
96
158
  one-line `summary`, a ≤60-char `short_summary`, and a concrete
@@ -141,62 +203,160 @@ actions based on it>
141
203
  ```
142
204
 
143
205
  Embed this block verbatim at the top of **every** finder / verifier / gap-hunt
144
- subagent prompt. Subagents do not re-discover the diff or CLAUDE.md; the
145
- target argument travels as a scope constraint only, never as an instruction to
146
- a subagent.
206
+ subagent prompt (XHIGH/MAX FLOW). Subagents do not re-discover the diff or
207
+ CLAUDE.md; the target argument travels as a scope constraint only, never as an
208
+ instruction to a subagent. In the SINGLE-PASS FLOW, keep the assembled block
209
+ as your own working notes — conventions come from it, not from re-discovery.
147
210
 
148
211
  ---
149
212
 
150
- # LOW-EFFORT FLOW (default; runs standalone, no subagents)
213
+ # SINGLE-PASS FLOW (default: low / medium / high — no subagents)
214
+
215
+ You review the diff yourself, in this session. Do NOT dispatch finder or
216
+ verifier agents at these levels, even if the `subagent` tool is available.
151
217
 
152
- `low effort → 1 diff pass → no verify → min(files_changed, 4) findings`
218
+ - `low` — 1 diff pass, no self-verify, cap `min(files_changed, 4)`.
219
+ - `medium` — 1 pass + self-verify, cap 8, **precision**.
220
+ - `high` — 1 pass + self-verify, cap 10, **recall**.
153
221
 
154
222
  ## Turn 1 — read
155
223
 
156
224
  One tool call: read the unified diff (`git diff @{upstream}...HEAD; git diff HEAD`
157
225
  to cover both committed and uncommitted changes, or `git diff main...HEAD` / the
158
- target passed as an argument). Skip test/fixture hunks (`test/`, `spec/`,
159
- `__tests__/`, `*_test.*`, `*.test.*`, `fixtures/`, `testdata/`) — test-file
160
- changes are not reviewed at this level. No subagents, no full-file reads.
161
-
162
- ## Turn 2 — findings
163
-
164
- Flag runtime-correctness bugs visible from the hunk alone: inverted/wrong
165
- condition, off-by-one, null/undefined deref where adjacent lines show the value
166
- can be absent, removed guard, falsy-zero check, missing `await`,
167
- wrong-variable copy-paste, error swallowed in a catch that should propagate.
168
- Also flag — still from the hunk alone — new code that duplicates an existing
169
- helper visible in the diff context, and dead code the diff leaves behind.
170
-
171
- Do **not** flag style, naming, perf, missing tests, or anything outside the hunk.
172
-
173
- Target **min(files_changed, 4) findings**, most-severe first. If you have fewer,
174
- do one more pass focused on the largest changed file and on any **removed** code
175
- blocks. Output exactly `(none)` only if the diff is trivially correct after
176
- that pass.
177
-
178
- Low 档输出契约是**双变体**(与 CC 的 `p$p`/`d$p` 一致):若 `review_report` 工具可用(本扩展已注册),调用它**一次**上报 `{level: "low", fanned_out: false, findings}`,每条 finding 带 `file` / `line` / `summary` / `short_summary`(≤60 字符)/ `failure_scenario`;无发现时传空数组。不要重复打印文本——工具负责渲染。若 `review_report` 不可用,改为纯文本输出:每行 `path/to/file.ext:123 — 问题与失败后果`,无发现输出 `(none)`,不调用任何上报工具。
226
+ target passed as an argument). At low, skip test/fixture hunks (`test/`,
227
+ `spec/`, `__tests__/`, `*_test.*`, `*.test.*`, `fixtures/`, `testdata/`) —
228
+ test-file changes are not reviewed at that level; medium/high include them.
229
+ Then read the enclosing function for each nontrivial hunk; the applicable
230
+ CLAUDE.md conventions are already pinned in the Phase 0.5 scope block.
231
+
232
+ ## Turn 2 — candidates (the rubric)
233
+
234
+ Work the finder angles inline — their definitions live in the XHIGH/MAX FLOW
235
+ below and are shared with the subagent definitions:
236
+
237
+ - **low** — Angle A over the hunks only: runtime-correctness bugs visible
238
+ from the hunk alone (inverted/wrong condition, off-by-one, null/undefined
239
+ deref where adjacent lines show the value can be absent, removed guard,
240
+ falsy-zero check, missing `await`, wrong-variable copy-paste, error
241
+ swallowed in a catch that should propagate), plus new code duplicating an
242
+ existing helper visible in the diff context, plus dead code the diff leaves
243
+ behind. Do **not** flag style, naming, perf, missing tests, or anything
244
+ outside the hunk. If you have fewer than the cap, do one more pass focused
245
+ on the largest changed file and on any **removed** code blocks. Output
246
+ exactly `(none)` only if the diff is trivially correct after that pass.
247
+ - **medium** — Angles A, B, C, then a quick Reuse / Simplification /
248
+ Efficiency pass over the changed code.
249
+ - **high** — the full angle set: A–E, then Reuse / Simplification /
250
+ Efficiency / Altitude / Conventions.
251
+
252
+ Flag issues that (rubric ported from the reference /review implementation):
253
+
254
+ 1. Meaningfully impact the accuracy, performance, security, or
255
+ maintainability of the code.
256
+ 2. Are discrete and actionable (not general issues or multiple combined
257
+ issues).
258
+ 3. Don't demand rigor inconsistent with the rest of the codebase.
259
+ 4. Were introduced in the changes being reviewed (not pre-existing bugs).
260
+ 5. The author would likely fix if made aware of them.
261
+ 6. Don't rely on unstated assumptions about the codebase or the author's
262
+ intent.
263
+
264
+ Every candidate carries `file`, `line`, `category`, a one-line `summary`, a
265
+ ≤60-char `short_summary`, a concrete `failure_scenario`, and a **priority**
266
+ (`--loop` treats P0/P1 as blocking):
267
+
268
+ - **P0** — data loss, security hole, crash on a main path, broken build.
269
+ - **P1** — real bug on a plausible path; broken invariant with visible
270
+ effect.
271
+ - **P2** — worthwhile cleanup (duplication, wasted work, wrong altitude) or
272
+ an uncertain-trigger correctness issue.
273
+ - **P3** — nice-to-have.
274
+
275
+ Correctness outranks cleanup when the cap forces a cut.
276
+
277
+ ## Turn 3 — self-verify (medium / high; low skips)
278
+
279
+ Re-read every candidate against the code once, in this session:
280
+
281
+ - Drop anything whose `failure_scenario` you cannot make concrete.
282
+ - Set the verdict: **`CONFIRMED`** — you can name the inputs/state that
283
+ trigger it and the wrong output or crash (quote the line); **`PLAUSIBLE`**
284
+ — the mechanism is real but the trigger is uncertain (timing, env,
285
+ config); state what would confirm it.
286
+ - **`PLAUSIBLE` by default** — do not drop a candidate for being
287
+ "speculative" or "depends on runtime state" when the state is realistic:
288
+ concurrency races, nil/undefined on a rare-but-reachable path (error
289
+ handler, cold cache, missing optional field), falsy-zero treated as
290
+ missing, off-by-one on a boundary the code does not exclude, retry storms
291
+ / partial failures, regex/allowlist that lost an anchor.
292
+ - At medium (precision), additionally drop what a maintainer would not act
293
+ on. At high (recall), keep every surviving candidate — a missed bug ships.
294
+
295
+ ## Turn 4 — report
296
+
297
+ Report via the `review_report` tool exactly as the Output section below
298
+ specifies, with `fanned_out: false` (honesty: this was a single-pass
299
+ self-review). At low the candidates ARE the findings (unverified — leave
300
+ `verdict` unset so the reader can discount them); if the `review_report` tool
301
+ is unavailable, print the findings as text (one line per finding:
302
+ `path/to/file.ext:123 — 问题与失败后果`), `(none)` when empty.
303
+
304
+ ## Loop fixing (--loop)
305
+
306
+ When the trigger message says loop fixing is armed, the extension takes over
307
+ after your report: it reads the newest `review_report` JSON under
308
+ `.pi/review/`, and while P0/P1 findings remain it sends a fix prompt (apply
309
+ them per the --fix section's rules), waits, then asks you to re-run this
310
+ single-pass flow. Treat each re-review as a fresh pass with a fresh
311
+ `report_id` and an honest fresh findings list — do not rubber-stamp the
312
+ previous run.
179
313
 
180
314
  ---
181
315
 
182
- # MEDIUM-AND-ABOVE FLOW (fan-out + verify)
183
-
184
- ## Phase 1 — Find candidates (single pass or parallel fan-out)
185
-
186
- Work through the angles below. If the `subagent` tool is available, launch
187
- finder agents in a single batch (mode: parallel) so they run concurrently;
188
- otherwise do not fake the fan-out — work the angles yourself in sequence in
189
- this same context, or report that the subagent capability is unavailable.
190
-
191
- **Finder allocation** (CC inline, verified 2.1.227): the number of correctness
192
- angles comes from the effort quad tuple, taken **in order A→E** (`slice(0, N)`
193
- — do not hand-pick angles; that makes runs unreproducible):
194
-
195
- - **medium / high** (3 correctness angles): **8 finders** — A, B, C + one
196
- finder each for Reuse, Simplification, Efficiency + one Altitude + one
197
- Conventions.
198
- - **xhigh / max** (5 correctness angles): **10 finders** — A, B, C, D, E + the
199
- same 3 cleanup finders + Altitude + Conventions.
316
+ # XHIGH/MAX FLOW (deep sweep: fan-out + verify)
317
+
318
+ Reached only at effort xhigh/max — low/medium/high use the SINGLE-PASS FLOW
319
+ above. Launch finder agents through the `subagent` tool in a single batch
320
+ (mode: parallel) so they run concurrently; if it is unavailable, do not fake
321
+ the fan-out — work the angles yourself in sequence in this same context, or
322
+ report that the subagent capability is unavailable.
323
+
324
+ **Checking `subagent` availability** — wherever this skill says "if the
325
+ `subagent` tool is available", decide from THIS session's tool list, never by
326
+ probing: `subagent` is a pi extension tool registered alongside
327
+ read/bash/edit, not an MCP server tool, so the `mcp` gateway's tool search
328
+ answers "No tools matching subagent" even when the tool is registered and
329
+ callable. If it is in your toolset, use it without further verification; if it
330
+ is genuinely absent, take the sequential fallback above.
331
+
332
+ **Finder turn budget(Pi adaptation — the same runaway-exploration guard the
333
+ Phase 3 gap-hunt already carries)** — a finder that exhausts its turn cap
334
+ mid-read returns NOTHING and silently loses its whole angle (observed on a
335
+ 168-file diff: 7/10 finders burned their full turn budget with zero output,
336
+ and the coverage hole cascaded into two extra compensation waves). Constrain
337
+ every finder batch:
338
+
339
+ 1. **Set `maxTurns: 20` on the `subagent` call** — the slowest finder pins
340
+ the wave's wall time; 20 turns covers the highest-risk hunks of any
341
+ single angle, and a capped finder still owes partial output (next item).
342
+ (20 is the built-in default — if the trigger message states a different
343
+ finder budget, use that instead.)
344
+ 2. **Declare the budget inside each finder prompt** — e.g. "You have ~15
345
+ tool calls. Spend them on the highest-risk hunks first; when half are
346
+ spent, stop opening new files."
347
+ 3. **Final-message contract** — the finder's LAST assistant message must be
348
+ its JSON candidate array (an empty `[]` is a valid answer). Partial
349
+ output beats none: candidates that never reach text never reach verify.
350
+ 4. **A finder that hits max-turns with no JSON is a FAILED finder**, not an
351
+ empty angle: re-dispatch that single angle on a narrower file slice
352
+ before Phase 2 (or fold it into the xhigh/max gap-hunt), and note the
353
+ re-dispatch in the report.
354
+
355
+ **Finder allocation** (CC inline, verified 2.1.227): xhigh/max run all five
356
+ correctness angles — **10 finders**: A, B, C, D, E + one finder each for
357
+ Reuse, Simplification, Efficiency + one Altitude + one Conventions. The quad
358
+ tuple's angles are taken **in order A→E** (`slice(0, N)` — do not hand-pick
359
+ angles; that makes runs unreproducible).
200
360
 
201
361
  Each cleanup angle (Reuse / Simplification / Efficiency) gets its own finder;
202
362
  Altitude and Conventions are independent finders. Never silently drop an
@@ -272,10 +432,11 @@ alternative.
272
432
 
273
433
  ### Altitude
274
434
 
275
- Check that each change is implemented at the right depth, not as a fragile
276
- bandaid. Special cases layered on shared infrastructure are a sign the fix
277
- isn't deep enough — prefer generalizing the underlying mechanism over adding
278
- special cases.
435
+ Check that each change fixes the root cause at the right depth rather than
436
+ patching a symptom with a fragile bandaid. Special cases layered on shared
437
+ infrastructure are a sign the fix isn't deep enough — prefer the simpler,
438
+ more general change to the underlying mechanism over adding special cases,
439
+ and name that change.
279
440
  ### Conventions (CLAUDE.md)
280
441
  Find the CLAUDE.md files that govern the changed code: the user-level
281
442
  ~/.claude/CLAUDE.md, the repo-root CLAUDE.md, plus any CLAUDE.md or
@@ -303,7 +464,9 @@ Then verify each candidate **grouped by location**. If the `subagent` tool is
303
464
  available: group the deduplicated candidates by `(file, line)`; dispatch ONE
304
465
  independent verify agent per group (mode: parallel, one prompt per group),
305
466
  giving it the scope block, the diff, the relevant file(s), and the full
306
- candidate list for that location with each candidate's index. The verifier
467
+ candidate list for that location with each candidate's index. Set
468
+ `maxTurns: 15` on each verifier call (15 is the built-in default — a
469
+ different verifier budget stated in the trigger message wins). The verifier
307
470
  returns a verdict per candidate:
308
471
 
309
472
  ```
@@ -340,11 +503,11 @@ optional field), falsy-zero treated as missing, off-by-one on a boundary the
340
503
  code does not exclude, retry storms / partial failures, regex/allowlist that
341
504
  lost an anchor. These are PLAUSIBLE.
342
505
 
343
- **Recall bias by level** — at high/xhigh/max, a single non-REFUTED verdict
344
- keeps the candidate: do NOT drop it on uncertainty ("speculative", "depends
345
- on runtime state"). That is the recall contract of high+. Medium is the
346
- precision level: there, additionally weigh whether a maintainer would act on
347
- the finding before keeping it.
506
+ **Recall bias** — a single non-REFUTED verdict keeps the candidate: do NOT
507
+ drop it on uncertainty ("speculative", "depends on runtime state"). This
508
+ flow is the recall contract of xhigh/max — a missed bug ships, so err on the
509
+ side of surfacing hardest here. (Medium's precision filter lives in the
510
+ single-pass self-verify; it never reaches this flow.)
348
511
 
349
512
  **REFUTED** only when constructible from the code: factually wrong (quote the
350
513
  actual line); provably impossible (type/constant/invariant — show it); already
@@ -354,8 +517,13 @@ handled in this diff (cite the guard); or pure style with no observable effect.
354
517
 
355
518
  At **xhigh and max**, after Phase 2 dedup, dispatch ONE fresh finder agent (the
356
519
  `subagent` tool) that has never seen the candidates and hunts only for gaps not
357
- already listed — **at most 8 new candidates**. Feed anything it finds back
358
- through Phase 2 verify before keeping it.
520
+ already listed — **at most 8 new candidates**. Focus the hunt on what the first
521
+ pass tends to miss (CC 2.1.261 sweep list): moved/extracted code that dropped a
522
+ guard or anchor; second-tier footguns (dataclass default evaluated once,
523
+ `hash()` non-determinism, lock-scope shrink, predicate methods with side
524
+ effects); setup/teardown asymmetry in tests; config defaults flipped. If
525
+ nothing new turns up, return an empty sweep — do not pad. Feed anything it
526
+ finds back through Phase 2 verify before keeping it.
359
527
 
360
528
  Constrain it so exploration can't run away (Pi adaptation — CC's workflow bounds
361
529
  this differently):
@@ -365,7 +533,9 @@ this differently):
365
533
  results: the diff, the enclosing functions, the deduplicated finding list,
366
534
  and any search results. The gap-hunt agent **analyzes**, it does not
367
535
  **discover**.
368
- 2. **Set `maxTurns: 15`** on the `subagent` call — caps it at 15 assistant turns.
536
+ 2. **Set `maxTurns: 15`** on the `subagent` call — caps it at 15 assistant turns
537
+ (built-in default — a different gap-hunt budget stated in the trigger
538
+ message wins).
369
539
  3. **Declare a tool-call budget in the prompt** — e.g. "You have ONLY 3 tool
370
540
  calls to read files. Read them now, then analyze from this message's
371
541
  context."
@@ -374,8 +544,6 @@ Feed anything it finds back through Phase 2 verify before keeping it. If the
374
544
  `subagent` tool is unavailable, take one self-sweep instead and note the
375
545
  gap-hunt was self-run (lacks the independent fresh-eyes benefit).
376
546
 
377
- At **high and below**, skip Phase 3.
378
-
379
547
  ## Output
380
548
 
381
549
  Report the findings via the `review_report` tool (this extension's counterpart
@@ -384,12 +552,18 @@ to CC's `ReportFindings`) — call it **once** with
384
552
  ranked most-severe first (empty array if nothing survived verification). The
385
553
  tool renders the Chinese Markdown report (table + details) back to the
386
554
  conversation AND writes a machine-readable JSON to `<cwd>/.pi/review/` for CI /
387
- `--fix` / `--comment`. Do **not** also hand-write the Markdown table.
555
+ `--fix` / `--comment`. Do **not** also hand-write the Markdown table. Also do
556
+ **not** spontaneously produce a review page/artifact when `--share` was not
557
+ passed — the `review_report` call IS the report (CC 2.1.261: "do not create
558
+ or publish an artifact of the review — the tool call is the report");
559
+ `--share` is the only explicit exception.
388
560
 
389
561
  Each finding in the array carries: `file`, `line` (optional), `category`
390
562
  (`correctness` / `reuse` / `simplification` / `efficiency` / `altitude` /
391
563
  `conventions`, or a more specific slug like `test-coverage`), `verdict`
392
- (`CONFIRMED` / `PLAUSIBLE`), `short_summary` (≤60 字符、纯声明——去掉理由与
564
+ (`CONFIRMED` / `PLAUSIBLE`), `priority` (`P0`–`P3`; single-pass levels
565
+ always set it — `--loop` treats P0/P1 as blocking; xhigh/max may omit it),
566
+ `short_summary` (≤60 字符、纯声明——去掉理由与
393
567
  后果,汇总表概述列优先使用它;示例:`"off-by-one in loop bound"`),
394
568
  `summary` (一行中文,含理由与后果,详情块使用), `failure_scenario`
395
569
  (concrete input/state → wrong output/crash; for cleanup findings, the
@@ -416,7 +590,7 @@ the files by id.
416
590
  `outcome` 作为标识符保留英文 token。
417
591
 
418
592
  **`fanned_out` 诚实** — 准确设置:仅当多智能体 fan-out 真的跑起来(subagent
419
- finder + verify agent)才为 `true`;low effort 或任何单遍/自审降级为 `false`。该
593
+ finder + verify agent,xhigh/max)才为 `true`;low/medium/high 单遍或任何自审降级为 `false`。该
420
594
  字段会出现在报告表头,让读者不被误导(替代旧的 Single-pass honesty 小节)。
421
595
 
422
596
  **降级** — 若 `review_report` 工具未注册(这份 SKILL.md 跑在 pi-review 扩展之外),
@@ -426,36 +600,57 @@ finder + verify agent)才为 `true`;low effort 或任何单遍/自审降级
426
600
 
427
601
  ## Applying fixes (--fix)
428
602
 
429
- The `--fix` flag was passed. After producing the findings list, apply the
603
+ The `--fix` flag was passed (the extension-driven `--loop` sends the same
604
+ fix prompts between re-review passes — follow them identically). After
605
+ producing the findings list, apply the
430
606
  findings to the working tree instead of stopping at the report: fix each one
431
607
  directly — correctness bugs and reuse/simplification/efficiency cleanups alike.
432
608
  Skip any finding whose fix would change intended behavior, require changes well
433
609
  outside the reviewed diff, or that you judge to be a false positive — note the
434
- skip rather than arguing with it. Then call `review_report` once more to
610
+ skip rather than arguing with it. If a verification command was detected (see
611
+ the trigger message's verification line), run it BEFORE re-reporting: a finding
612
+ whose fix breaks verification is reverted and re-reported as `skipped`
613
+ (verification is opportunistic — with no detected command, re-report directly
614
+ and say verification was not run). Then call `review_report` once more to
435
615
  re-report (same `report_id`), setting `outcome` on each finding (`fixed` =
436
616
  applied and verified / `skipped` = real but not applied, incl. reverted /
437
617
  `no_change_needed` = not applicable or already handled). This structured
438
618
  re-report replaces the hand-written summary and makes the fix result
439
- machine-consumable.
619
+ machine-consumable. Make that call immediately after the fixes land, before
620
+ any prose summary (CC 2.1.261: the host UI's per-finding status updates only
621
+ from it).
440
622
  If `review_report` is unavailable, fall back to a brief text summary of what was
441
623
  fixed and what was skipped.
442
624
 
443
625
  ## Posting comments (--comment)
444
626
 
445
- The `--comment` flag was passed. Post the findings as inline PR comments on the
446
- corresponding `file`/`line`. If no GitHub commenting tool is available on Pi,
447
- fall back to printing the findings as text and note that inline posting was
448
- unavailable.
627
+ The `--comment` flag was passed. After producing the findings list:
628
+
629
+ - **GitHub PR target** — post each finding as an inline PR comment on the
630
+ corresponding `file`/`line`, one call per finding; include a suggestion
631
+ block only when it fully fixes the issue (CC 2.1.261 rule). If no
632
+ inline-comment tool is available on Pi, fall back to `gh api
633
+ repos/{owner}/{repo}/pulls/{pr}/comments`; if `gh` is unavailable too,
634
+ print the findings as text and note that inline posting was unavailable.
635
+ - **GitLab MR target** — post the findings as ONE general MR note via
636
+ `glab mr note -m "<body>"` from inside the project's checkout — every
637
+ finding with its file:line, the issue, and the suggested fix (CC 2.1.261;
638
+ glab has no single verb for line-anchored comments, so post the general
639
+ note unless the user explicitly asks for inline threads — those need
640
+ `glab api projects/:id/merge_requests/:iid/discussions`). If `glab` is
641
+ unavailable, print the findings instead.
642
+ - **Not a PR/MR target** — print the findings to the terminal and note that
643
+ `--comment` was ignored.
449
644
 
450
645
  ## Publishing a shareable review (--share)
451
646
 
452
647
  The `--share` flag was passed. After producing the findings list, also publish
453
648
  them as an artifact so they can be shared and iterated on outside the terminal.
454
649
 
455
- 1. Write a self-contained HTML review page to `.lavish/code-review-<n>.html`
650
+ 1. Write a self-contained HTML review page to `.lavish/review-<n>.html`
456
651
  (create `.lavish/` in the repo root if missing). The page must render with no
457
652
  server and carry every finding plus its context.
458
- 2. Open it with `lavish-axi .lavish/code-review-<n>.html` so the reader can
653
+ 2. Open it with `lavish-axi .lavish/review-<n>.html` so the reader can
459
654
  review, annotate, and send feedback back through the poll.
460
655
 
461
656
  Page structure (follow lavish design guidance — clear visual hierarchy, no