@fyeeme/pi-review 1.0.0 → 1.0.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -29,6 +29,9 @@ From the package dir:
29
29
  npm install
30
30
  ```
31
31
 
32
+ This resolves [`@fyeeme/pi-subagent-core`](https://www.npmjs.com/package/@fyeeme/pi-subagent-core)
33
+ (`^0.3.0`, from the npm registry — no sibling-repo layout requirement).
34
+
32
35
  Then point pi at it (e.g. via your extensions config), or symlink into your pi
33
36
  extensions directory.
34
37
 
@@ -63,6 +66,17 @@ This is the deterministic mode selection a pure-prompt skill cannot reproduce
63
66
  `ctx.getContextUsage()`). The decision is announced in the trigger message so
64
67
  it is observable.
65
68
 
69
+ **Apply → verify → revert safety net** (harden-code-simplify): after Phase 2
70
+ applies the cleanups, the handler also injects a verification command detected
71
+ from `package.json` scripts (`check` → `test` → `lint` → `typecheck`). The
72
+ skill snapshots the touched files, applies the fixes, runs that command, and
73
+ on failure auto-reverts per-file (a clean apply runs verify exactly once; only
74
+ a failure escalates to one verify per touched file). The result is reported as
75
+ structured outcomes via `review_report` (`level: "simplify"`), not a free-text
76
+ summary. If no verification command is detectable, fixes are kept but the
77
+ report states no verification was run (verification is opportunistic, never
78
+ blocking).
79
+
66
80
  ### `subagent` tool
67
81
 
68
82
  An LLM-callable tool that spawns one or more real pi subprocesses:
@@ -70,9 +84,22 @@ An LLM-callable tool that spawns one or more real pi subprocesses:
70
84
  | mode | behavior |
71
85
  |---|---|
72
86
  | `single` | run `prompts[0]` once (e.g. an independent verify agent) |
73
- | `parallel` | run all prompts concurrently, capped at 8 (e.g. one finder per angle) |
87
+ | `parallel` | run all prompts concurrently, capped at the ceiling (e.g. one finder per angle) |
74
88
  | `chain` | run sequentially; each later prompt receives prior output |
75
89
 
90
+ **Fan-out guards** (harden-code-simplify, shared with `/code-review`):
91
+
92
+ - **Recursion cap (whitelist-by-default)** — a spawned sub-agent does not
93
+ receive the `subagent` tool in its default toolset, so it cannot recurse. A
94
+ caller opts in by listing `subagent` in the child's `tools` whitelist; set
95
+ `PI_SUBAGENT_MAX_SPAWN_DEPTH` to allow multi-level fan-out up to a hard cap.
96
+ - **Default turn budget** — fan-out agents get a finite default `maxTurns` (50,
97
+ aligned to CC's `FORKED_AGENT_DEFAULT_MAX_TURNS` in 2.1.227)
98
+ when the caller omits it; an explicit `0` is honored.
99
+ - **Configurable concurrency** — `PI_MAX_CONCURRENT_SUBAGENTS` (default 20,
100
+ aligned to CC's `CLAUDE_CODE_MAX_CONCURRENT_SUBAGENTS ?? 20` in 2.1.227;
101
+ invalid values fall back to the default).
102
+
76
103
  Each sub-agent is a full `pi --mode json -p --no-session` run. Progress streams
77
104
  to the TUI via `onUpdate` as each agent completes. ESC aborts the whole batch
78
105
  (SIGTERM → 5s → SIGKILL per subprocess). Errors are thrown (not returned) so
@@ -85,8 +112,6 @@ pi-review/
85
112
  ├── index.ts factory: registerTool(subagent) + 2 commands
86
113
  ├── skills/ bundled SKILL.md files (code-review, simplify)
87
114
  ├── src/
88
- │ ├── agent/dispatch.ts spawnAgent + mapWithConcurrencyLimit (self-contained copy
89
- │ │ from pi-dynamic-workflows; no external dep beyond node + pi-ai)
90
115
  │ ├── skills.ts bundledSkillPath — resolve this extension's own skills/ dir
91
116
  │ ├── tools/subagent.ts defineTool("subagent") — generic capability layer
92
117
  │ └── commands/
@@ -96,15 +121,18 @@ pi-review/
96
121
  ```
97
122
 
98
123
  The layout is deliberately layered: `src/tools/` is the **generic capability
99
- layer** (subagent tool + dispatch), `src/commands/` is the **entry layer** (one
124
+ layer** (subagent tool), `src/commands/` is the **entry layer** (one
100
125
  file per skill). If a third or fourth skill needs the subagent tool, `src/tools/`
101
126
  can be split into its own `pi-subagent` extension with zero refactor — the code
102
127
  is already separated.
103
128
 
104
- `src/agent/dispatch.ts` is a self-contained copy of the spawn pattern from
105
- `examples/extensions/subagent` and `pi-dynamic-workflows/src/agent/dispatch.ts`
106
- (~150 lines). When pi promotes `spawnAgent` to a public `pi-coding-agent`
107
- export, this file should be deleted in favor of that import.
129
+ The dispatch primitive (`spawnAgent`, `mapWithConcurrencyLimit`,
130
+ `createSpawnRegistry`, `abortAgent`, `getPiInvocation` + types) lives in
131
+ [`pi-subagent-core`](../pi-subagent-core) (npm `@fyeeme/pi-subagent-core`), a
132
+ shared library extracted from the duplicated copies that used to live here and
133
+ in `pi-dynamic-workflows`. When pi promotes `spawnAgent` to a public
134
+ `pi-coding-agent` export, `pi-subagent-core` should be deleted in favor of that
135
+ import.
108
136
 
109
137
  ## Relation to the skills
110
138
 
@@ -118,6 +146,12 @@ sub-agents are spawned and *which mode* is chosen.
118
146
 
119
147
  ## Status
120
148
 
121
- MVP. Phase 2 (not yet built): a `review_verify` tool encapsulating 3-vote
122
- adversarial verify, and a `review_report` tool enforcing the output schema /
123
- `--share` lavish artifact.
149
+ `review_report` is built — schema aligned to CC `ReportFindings` (2.1.227
150
+ empirical): 3-state `outcome` (`fixed`/`skipped`/`no_change_needed`), 2-value
151
+ `verdict` (`CONFIRMED`/`PLAUSIBLE`), `short_summary` (≤60, table overview),
152
+ `report_id` for fixed-later re-reports; renders the Chinese Markdown report
153
+ and writes JSON to `<cwd>/.pi/review/` for CI / `--fix` / `--comment`.
154
+
155
+ Remaining Phase 2 item (not yet built): a `review_verify` tool encapsulating
156
+ 3-vote adversarial verify. `--share` already routes through lavish-axi (see
157
+ the code-review skill).
package/index.ts CHANGED
@@ -4,6 +4,9 @@
4
4
  * Registers:
5
5
  * - the `subagent` tool — general-purpose parallel/sequential sub-agent fan-out
6
6
  * via real pi subprocesses. Shared capability used by both skills below;
7
+ * - the `review_report` tool — structured findings sink for the code-review
8
+ * skill (Pi's counterpart to CC's ReportFindings): renders the Markdown
9
+ * report + writes JSON to <cwd>/.pi/review/ for CI;
7
10
  * - the `/code-review` command — effort-level review via the code-review skill;
8
11
  * - the `/code-simplify` command — cleanup via the simplify skill; the handler
9
12
  * decides parallel vs single-pass from ctx.getContextUsage(), mirroring CC's
@@ -13,16 +16,25 @@
13
16
  * provides the entry commands + the fan-out capability they need.
14
17
  *
15
18
  * Layout (layered so the tool layer can be split into its own extension later):
16
- * src/tools/subagent.ts — generic capability (subagent tool + dispatch)
17
- * src/commands/*.ts — per-skill entry commands
19
+ * src/tools/subagent.ts — generic capability (subagent tool; dispatch from pi-subagent-core)
20
+ * src/tools/review_report.ts — structured findings sink (review_report tool; CC ReportFindings counterpart)
21
+ * src/commands/*.ts — per-skill entry commands
18
22
  */
19
23
  import type { ExtensionAPI } from "@earendil-works/pi-coding-agent";
24
+ import { isFanoutToolAllowed } from "@fyeeme/pi-subagent-core";
20
25
  import { registerCodeReview } from "./src/commands/code-review.ts";
21
26
  import { registerSimplify } from "./src/commands/code-simplify.ts";
22
27
  import { subagentTool } from "./src/tools/subagent.ts";
28
+ import { reviewReportTool } from "./src/tools/review_report.ts";
23
29
 
24
30
  export default function (pi: ExtensionAPI): void {
25
- pi.registerTool(subagentTool);
31
+ // The fan-out tool registers only when recursion is allowed for THIS
32
+ // process (top-level, or a child the spawner explicitly opted in AND that is
33
+ // below the max-depth cap). A default child — spawned without the fan-out
34
+ // tool in its whitelist — loads without it, so it physically cannot recurse.
35
+ // This is the whitelist-by-default recursion guard (harden-code-simplify).
36
+ if (isFanoutToolAllowed()) pi.registerTool(subagentTool);
37
+ pi.registerTool(reviewReportTool);
26
38
  registerCodeReview(pi);
27
39
  registerSimplify(pi);
28
40
  }
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@fyeeme/pi-review",
3
- "version": "1.0.0",
3
+ "version": "1.0.2",
4
4
  "description": "Review & cleanup extension for pi. Registers /code-review and /code-simplify commands plus a general-purpose `subagent` tool that spawns parallel pi subprocesses — providing the real fan-out capability the code-review and simplify skills (bundled under `skills/`) need for their multi-agent flows. The /code-simplify handler uses ctx.getContextUsage() to decide parallel vs single-pass mode deterministically.",
5
5
  "type": "module",
6
6
  "license": "MIT",
@@ -34,16 +34,21 @@
34
34
  "test": "vitest --run",
35
35
  "typecheck": "tsc"
36
36
  },
37
+ "dependencies": {
38
+ "@fyeeme/pi-subagent-core": "^0.3.3"
39
+ },
37
40
  "peerDependencies": {
38
- "@earendil-works/pi-ai": ">=0.77.0",
39
- "@earendil-works/pi-coding-agent": ">=0.77.0",
41
+ "@earendil-works/pi-ai": ">=0.84.1",
42
+ "@earendil-works/pi-coding-agent": ">=0.84.1",
43
+ "@earendil-works/pi-tui": ">=0.84.1",
40
44
  "jiti": ">=2.0.0",
41
45
  "typebox": ">=1.0.0",
42
46
  "typescript": ">=5.0.0"
43
47
  },
44
48
  "devDependencies": {
45
- "@earendil-works/pi-ai": "0.77.0",
46
- "@earendil-works/pi-coding-agent": "0.77.0",
49
+ "@earendil-works/pi-ai": "0.84.1",
50
+ "@earendil-works/pi-coding-agent": "0.84.1",
51
+ "@earendil-works/pi-tui": "0.84.1",
47
52
  "@types/node": "22.19.19",
48
53
  "jiti": "2.7.0",
49
54
  "typebox": "1.1.38",
@@ -10,24 +10,27 @@ description: "Review the current diff for correctness bugs and reuse/simplificat
10
10
  earlier v2.1.220 reconstruction. Every section below was located in the
11
11
  extracted strings (cc_strings_223.txt) and verified.
12
12
 
13
- What CC 2.1.223 actually contains (verified against the binary):
14
- - Effort: medium = precision; high = recall ("err on the side of
15
- surfacing"); xhigh/max add a gap-hunt. Each finder surfaces ≤6 candidates.
13
+ What CC 2.1.227 actually contains (verified against the binary):
14
+ - Effort quad tuple {correctnessAngles, perAngle, maxFindings, sweep}:
15
+ medium {3,6,8,false} / high {3,6,10,false} / xhigh {5,8,15,true} / max
16
+ same structure as xhigh. medium = precision; high+ = recall
17
+ ("err on the side of surfacing"); xhigh/max add a gap-hunt (≤8 new
18
+ candidates). Correctness angles are taken in order A→E (`slice(0, N)`).
19
+ - Inline finder allocation: medium/high = 8 finders (A/B/C + 3 cleanup +
20
+ altitude + conventions); xhigh/max = 10 finders (A–E + same). Each
21
+ cleanup angle gets its own finder.
16
22
  - Low effort: 1 diff pass, no verify, target min(files_changed, 4) findings.
17
- - Finder allocation (workflow Find-phase, linearized inline for Pi): one
18
- finder per correctness angle (A–E) + Conventions, plus one combined cleanup
19
- finder, pooled before verify. (CC's native *inline* medium path splits
20
- differently — 8 finders: 3 correctness + 3 cleanup + altitude + conventions.)
23
+ - Verify via an independent agent, grouped by (file, line) (absorbed from
24
+ the workflow GROUP_VERDICT_SCHEMA): CONFIRMED / PLAUSIBLE / REFUTED,
25
+ "PLAUSIBLE by default". Keep CONFIRMED + PLAUSIBLE, drop REFUTED.
26
+ - ReportFindings schema: verdict CONFIRMED|PLAUSIBLE, outcome
27
+ fixed|skipped|no_change_needed, finding carries short_summary (≤60).
21
28
  - Angles A–E + Reuse/Simplification/Efficiency/Altitude + Conventions,
22
29
  verbatim (same source variables the /simplify skill reuses).
23
- - Verify via an independent agent: CONFIRMED / PLAUSIBLE / REFUTED,
24
- "PLAUSIBLE by default". Keep CONFIRMED + PLAUSIBLE, drop REFUTED.
25
30
  - Gap-hunt (xhigh/max): one fresh finder hunting only for gaps not
26
31
  already listed (CC's Sweep phase: "Fresh finder hunting only for gaps").
27
- - Output: Markdown findings table + per-finding details block, printed
28
- as text (no ReportFindings tool on Pi); carries file:line/category/
29
- verdict/summary/failure_scenario per finding.
30
- - --share publishes an Artifact; --fix applies findings to the working tree.
32
+ - Fixed-later obligation (CC Q8m): later fixes in the session must
33
+ re-report findings with updated outcome.
31
34
 
32
35
  Invocation: /code-review [low|medium|high|xhigh|max] [--fix] [--comment] [--share] [<target>]
33
36
  target = Class#method | file path | PR number | branch name
@@ -41,9 +44,13 @@ description: "Review the current diff for correctness bugs and reuse/simplificat
41
44
  Pi ADAPTATIONS (differ from the CC runtime)
42
45
  ════════════════════════════════════════════════════════════════════════
43
46
  1. Output — CC calls a ReportFindings tool with {level, findings}; Pi
44
- PRINTS a Markdown findings table + details block as text
45
- (no such tool on Pi). [was JSON array; switched for
46
- readability]
47
+ uses this extension's `review_report` tool (the Pi counterpart
48
+ to ReportFindings, verdict/outcome enums aligned to CC
49
+ v2.1.227): it renders the Chinese Markdown report (table +
50
+ details) back to the conversation AND writes a
51
+ machine-readable JSON to <cwd>/.pi/review/ for CI / --fix /
52
+ --comment. If the tool is absent, fall back to printing the
53
+ Markdown as text.
47
54
  2. Fan-out — CC uses the Agent tool; Pi uses the `subagent` tool
48
55
  (mode: parallel), or runs angles sequentially if unavailable.
49
56
  3. Verify — CC uses the Agent tool; Pi uses `subagent` for the
@@ -68,15 +75,26 @@ altitude, and conventions findings when the output cap forces a cut.
68
75
 
69
76
  ## Effort levels
70
77
 
71
- | Level | Intent | Verify | Subagents | Output cap |
78
+ | Level | Intent | Verify | Subagents | 四元组 `{correctnessAngles, perAngle, maxFindings, sweep}` |
72
79
  |-------|--------|--------|-----------|------------|
73
- | low (default) | quick scan | no | no | min(files_changed, 4) |
74
- | medium | **precision** — surface only findings a maintainer would act on | independent agent | fan-out (1 finder/angle) | ≤ 8 |
75
- | high | **recall** — catch every real bug a careful reviewer would; **err on the side of surfacing** | independent agent | more angles | ≤ 10 |
76
- | xhigh → max | recall + **gap-hunt** | independent agent | above + 1 fresh gap finder | larger, may include uncertain |
80
+ | low (default) | quick scan | no | no | 上限 `min(files_changed, 4)` |
81
+ | medium | **precision** — surface only findings a maintainer would act on | independent verifier (grouped) | 8 finders | `{3, 6, 8, false}` |
82
+ | high | **recall** — catch every real bug a careful reviewer would; **err on the side of surfacing** | recall-biased verifier (grouped) | 8 finders | `{3, 6, 10, false}` |
83
+ | xhigh | recall + **gap-hunt** | recall-biased verifier (grouped) | 10 finders + 1 gap | `{5, 8, 15, true}` |
84
+ | max | 同 xhigh | 同 xhigh | 同 xhigh | 同 xhigh |
85
+
86
+ **max 与 xhigh 结构相同**:fan-out / verify / sweep 完全一致,差别仅在模型 reasoning effort(CC v2.1.226 注释实证:`max → same structure as xhigh (the API reasoning effort differs, not the fan-out)`)。若运行时不支持调节 reasoning effort,max 在结构上退化为 xhigh——不要因档名而期待更多 fan-out。
87
+
88
+ The quad tuple parameterizes the whole pipeline (CC inline semantics, verified 2.1.227):
89
+
90
+ - `correctnessAngles` — how many correctness angles A–E run, taken **in order** (medium/high: A/B/C; xhigh/max: A–E).
91
+ - `perAngle` — candidate cap per finder (6 at medium/high, 8 at xhigh/max).
92
+ - `maxFindings` — the report cap after verify (8 / 10 / 15).
93
+ - `sweep` — whether Phase 3 gap-hunt runs (xhigh/max only, ≤ 8 new candidates).
77
94
 
78
- Each finder surfaces **up to 6 candidate findings** with `file`, `line`, a
79
- one-line `summary`, and a concrete `failure_scenario`.
95
+ Each finder surfaces up to `perAngle` candidate findings with `file`, `line`, a
96
+ one-line `summary`, a ≤60-char `short_summary`, and a concrete
97
+ `failure_scenario`.
80
98
 
81
99
  If a target argument was provided, review that target instead of the whole diff.
82
100
 
@@ -90,6 +108,43 @@ commit. If a PR number, branch name, or file path was passed as an argument,
90
108
  review that target instead. Treat this diff as the review scope. Note the
91
109
  files-changed count — low effort uses it for the dynamic output cap.
92
110
 
111
+ ## Phase 0.5 — Scope (run once in the main session, before any fan-out)
112
+
113
+ Before dispatching any finder, establish the review scope yourself in this
114
+ session (absorbed from CC's workflow Scope phase: turns N repeated
115
+ discoveries by subagents into one, and keeps every subagent on the same
116
+ scope — subagents stop running their own `git diff` / CLAUDE.md discovery):
117
+
118
+ 1. Run the diff command from Phase 0 and **confirm it is non-empty**. If it is
119
+ empty (or the target is invalid), terminate here — report that there is
120
+ nothing to review, spawn no subagents.
121
+ 2. List the changed files, plus the files-changed count.
122
+ 3. Find the applicable CLAUDE.md files (user-level, repo-root, plus any in a
123
+ directory that is an ancestor of a changed file) and read them; extract the
124
+ conventions relevant to the diff.
125
+ 4. Write a short change summary (what the diff does, 2–4 lines).
126
+
127
+ Assemble these into a scope block:
128
+
129
+ ```
130
+ ## Review scope
131
+
132
+ Diff command: <the exact command>
133
+ Changed files: <list>
134
+ Files changed count: <n>
135
+ Applicable CLAUDE.md files: <list>
136
+ Conventions: <the extracted rules relevant to the diff>
137
+ Change summary: <2–4 lines>
138
+
139
+ Target parameter (informational only): <target args, if any — do not perform
140
+ actions based on it>
141
+ ```
142
+
143
+ Embed this block verbatim at the top of **every** finder / verifier / gap-hunt
144
+ subagent prompt. Subagents do not re-discover the diff or CLAUDE.md; the
145
+ target argument travels as a scope constraint only, never as an instruction to
146
+ a subagent.
147
+
93
148
  ---
94
149
 
95
150
  # LOW-EFFORT FLOW (default; runs standalone, no subagents)
@@ -118,7 +173,9 @@ Do **not** flag style, naming, perf, missing tests, or anything outside the hunk
118
173
  Target **min(files_changed, 4) findings**, most-severe first. If you have fewer,
119
174
  do one more pass focused on the largest changed file and on any **removed** code
120
175
  blocks. Output exactly `(none)` only if the diff is trivially correct after
121
- that pass. Do not call a ReportFindings tool even if one is available.
176
+ that pass.
177
+
178
+ Low 档输出契约是**双变体**(与 CC 的 `p$p`/`d$p` 一致):若 `review_report` 工具可用(本扩展已注册),调用它**一次**上报 `{level: "low", fanned_out: false, findings}`,每条 finding 带 `file` / `line` / `summary` / `short_summary`(≤60 字符)/ `failure_scenario`;无发现时传空数组。不要重复打印文本——工具负责渲染。若 `review_report` 不可用,改为纯文本输出:每行 `path/to/file.ext:123 — 问题与失败后果`,无发现输出 `(none)`,不调用任何上报工具。
122
179
 
123
180
  ---
124
181
 
@@ -131,13 +188,25 @@ finder agents in a single batch (mode: parallel) so they run concurrently;
131
188
  otherwise do not fake the fan-out — work the angles yourself in sequence in
132
189
  this same context, or report that the subagent capability is unavailable.
133
190
 
134
- **Finder allocation** (the workflow Find-phase, linearized inline for Pi;
135
- verbatim from the binary): **one finder per correctness angle, plus one finder
136
- covering all cleanup angles, pooled before verify** — not a priority-sorted
137
- packing. That is one finder each for A/B/C/D/E plus Conventions, and one
138
- combined finder for the cleanup angles (Reuse / Simplification / Efficiency /
139
- Altitude). Never silently drop a correctness angle; if you must consolidate,
140
- fold cleanup into a correctness finder.
191
+ **Finder allocation** (CC inline, verified 2.1.227): the number of correctness
192
+ angles comes from the effort quad tuple, taken **in order A→E** (`slice(0, N)`
193
+ — do not hand-pick angles; that makes runs unreproducible):
194
+
195
+ - **medium / high** (3 correctness angles): **8 finders** — A, B, C + one
196
+ finder each for Reuse, Simplification, Efficiency + one Altitude + one
197
+ Conventions.
198
+ - **xhigh / max** (5 correctness angles): **10 finders** — A, B, C, D, E + the
199
+ same 3 cleanup finders + Altitude + Conventions.
200
+
201
+ Each cleanup angle (Reuse / Simplification / Efficiency) gets its own finder;
202
+ Altitude and Conventions are independent finders. Never silently drop an
203
+ angle — if you must consolidate (subagent unavailable), fold the cleanup
204
+ angles into a correctness finder, but say so in the report.
205
+
206
+ **Suppression 禁令(xhigh/max)** — different finders may surface different
207
+ candidates for the same line with different reasons. At xhigh/max all of them
208
+ are recorded and pass through verify independently: do NOT let one angle's
209
+ conclusions suppress another's — record both.
141
210
 
142
211
  The correctness angles hunt for bugs; the cleanup angles hunt for cleanup in
143
212
  the changed code. Cleanup, altitude, and conventions candidates use the same
@@ -180,29 +249,33 @@ through a registry/session/global — e.g. a caching provider holding a
180
249
  wrapper forwards all the methods the callers actually use.
181
250
 
182
251
  ### Reuse
183
- Flag new code that re-implements something the codebase already has — Grep
184
- shared/utility modules and files adjacent to the change, and name the existing
185
- helper to call instead.
252
+
253
+ Flag new code that re-implements something the codebase
254
+ already has — Grep shared/utility modules and files adjacent to the change,
255
+ and name the existing helper to call instead.
186
256
 
187
257
  ### Simplification
258
+
188
259
  Flag unnecessary complexity the diff adds: redundant or derivable state,
189
- copy-paste with slight variation, deep nesting, dead code left behind. Name the
190
- simpler form that does the same job.
260
+ copy-paste with slight variation, deep nesting, dead code left behind. Name
261
+ the simpler form that does the same job.
191
262
 
192
263
  ### Efficiency
264
+
193
265
  Flag wasted work the diff introduces: redundant computation or repeated I/O,
194
- independent operations run sequentially, blocking work added to startup or hot
195
- paths. Also flag long-lived objects built from closures or captured environments
196
- — they keep the entire enclosing scope alive for the object's lifetime (a memory
197
- leak when that scope holds large values); prefer a class/struct that copies only
198
- the fields it needs. Name the cheaper alternative.
266
+ independent operations run sequentially, blocking work added to startup or
267
+ hot paths. Also flag long-lived objects built from closures or captured
268
+ environments — they keep the entire enclosing scope alive for the object's
269
+ lifetime (a memory leak when that scope holds large values); prefer a
270
+ class/struct that copies only the fields it needs. Name the cheaper
271
+ alternative.
199
272
 
200
273
  ### Altitude
201
- Check that each change is implemented at the right depth, not as a fragile
202
- bandaid. Special cases layered on shared infrastructure are a sign the fix isn't
203
- deep enough — prefer generalizing the underlying mechanism over adding special
204
- cases.
205
274
 
275
+ Check that each change is implemented at the right depth, not as a fragile
276
+ bandaid. Special cases layered on shared infrastructure are a sign the fix
277
+ isn't deep enough — prefer generalizing the underlying mechanism over adding
278
+ special cases.
206
279
  ### Conventions (CLAUDE.md)
207
280
  Find the CLAUDE.md files that govern the changed code: the user-level
208
281
  ~/.claude/CLAUDE.md, the repo-root CLAUDE.md, plus any CLAUDE.md or
@@ -222,15 +295,36 @@ are the dominant cause of misses.
222
295
 
223
296
  ## Phase 2 — Dedup and verify
224
297
 
225
- Dedup near-duplicates (same defect, same location, same reason → keep one).
226
-
227
- Then verify each candidate. If the `subagent` tool is available, dispatch an
228
- independent verify agent (one per candidate, or a small batch): give it the
229
- diff, the relevant file(s), and the candidate; it returns exactly one of
230
- **CONFIRMED / PLAUSIBLE / REFUTED**. An independent agent counters the
231
- confirmation bias of self-review. If `subagent` is unavailable, fall back to
232
- re-checking each candidate yourself (self-check). Keep **CONFIRMED and
233
- PLAUSIBLE**, drop REFUTED. Give each surviving finding a verdict:
298
+ Dedup near-duplicates (same defect, same location, same reason → keep one;
299
+ different reasons for the same line are NOT duplicates — at xhigh/max both
300
+ are kept per the suppression ban).
301
+
302
+ Then verify each candidate **grouped by location**. If the `subagent` tool is
303
+ available: group the deduplicated candidates by `(file, line)`; dispatch ONE
304
+ independent verify agent per group (mode: parallel, one prompt per group),
305
+ giving it the scope block, the diff, the relevant file(s), and the full
306
+ candidate list for that location with each candidate's index. The verifier
307
+ returns a verdict per candidate:
308
+
309
+ ```
310
+ [{ "index": <candidate index>, "verdict": "CONFIRMED" | "PLAUSIBLE" | "REFUTED", "evidence": "<quote/argument>" }, ...]
311
+ ```
312
+
313
+ Grouping is by location, NOT dedup — each candidate is judged independently;
314
+ same-location candidates may describe different defects. A candidate the
315
+ verifier omitted (interrupted or skipped an index) is **dropped** — never
316
+ invent a PLAUSIBLE for it. One verifier failure drops its whole group (the
317
+ same trade-off CC's workflow makes); if you are not confident in a group's
318
+ verifier, fall back to one verifier per candidate for that group.
319
+
320
+ This group-by-location verify is a deliberate absorption of CC's workflow
321
+ optimization into the inline path (CC inline dispatches one verifier per
322
+ candidate): a location with 3 candidates costs 1 verifier instead of 3 —
323
+ ~40% fewer verifier agents is the expectation, not a guarantee.
324
+
325
+ If `subagent` is unavailable, fall back to re-checking each candidate yourself
326
+ (self-check). Keep **CONFIRMED and PLAUSIBLE**, drop REFUTED. Give each
327
+ surviving finding a verdict:
234
328
 
235
329
  - **CONFIRMED** — can name the inputs/state that trigger it and the wrong
236
330
  output or crash. Quote the line.
@@ -246,6 +340,12 @@ optional field), falsy-zero treated as missing, off-by-one on a boundary the
246
340
  code does not exclude, retry storms / partial failures, regex/allowlist that
247
341
  lost an anchor. These are PLAUSIBLE.
248
342
 
343
+ **Recall bias by level** — at high/xhigh/max, a single non-REFUTED verdict
344
+ keeps the candidate: do NOT drop it on uncertainty ("speculative", "depends
345
+ on runtime state"). That is the recall contract of high+. Medium is the
346
+ precision level: there, additionally weigh whether a maintainer would act on
347
+ the finding before keeping it.
348
+
249
349
  **REFUTED** only when constructible from the code: factually wrong (quote the
250
350
  actual line); provably impossible (type/constant/invariant — show it); already
251
351
  handled in this diff (cite the guard); or pure style with no observable effect.
@@ -254,7 +354,8 @@ handled in this diff (cite the guard); or pure style with no observable effect.
254
354
 
255
355
  At **xhigh and max**, after Phase 2 dedup, dispatch ONE fresh finder agent (the
256
356
  `subagent` tool) that has never seen the candidates and hunts only for gaps not
257
- already listed.
357
+ already listed — **at most 8 new candidates**. Feed anything it finds back
358
+ through Phase 2 verify before keeping it.
258
359
 
259
360
  Constrain it so exploration can't run away (Pi adaptation — CC's workflow bounds
260
361
  this differently):
@@ -277,48 +378,49 @@ At **high and below**, skip Phase 3.
277
378
 
278
379
  ## Output
279
380
 
280
- Print the findings as a **Markdown table + a details block** — readable in a
281
- terminal and in rendered Markdown (no JSON, no ReportFindings tool on Pi). Cap =
282
- low's min(files_changed, 4); 8 at medium; 10 at high; larger at xhigh → max.
283
-
284
- **全部用中文输出**:表头、概述、场景一律用中文;`Verdict`、`Category` 作为标识符保留
285
- 英文 token(CONFIRMED / PLAUSIBLE;correctness / reuse …)。
286
-
287
- **1. 表头行** — 单行写明:力度、diff 命令/范围、改动文件数、命中条数、是否真的多智能体
288
- 并发(见下文 Single-pass honesty)。示例:
289
-
290
- > `max` · `git diff HEAD` · 29 个文件 · 3 条发现 · 多智能体(验证 + 查漏)
291
-
292
- **2. 发现汇总表** — 按严重程度从高到低,每条一行:
293
-
294
- | # | 判定 | 类别 | 位置 | 概述 |
295
- |---|------|------|------|------|
296
- | 1 | CONFIRMED | correctness | path/file.ext:123 | 一句话说明这个 bug |
297
- | 2 | PLAUSIBLE | reuse | path/file.ext:45 | … |
298
-
299
- - `判定` — `CONFIRMED` / `PLAUSIBLE` / 留空(未做验证)。
300
- - `类别` — 产生该发现的角度,短横线小写 slug(`correctness`、`simplification`、
301
- `efficiency`、`reuse`、`altitude`、`conventions`,或更具体的如 `test-coverage`)。
302
- - `位置` — `文件:行号`。
303
- - `概述` — 一句话说明(≤ 约 80 字,同时作为紧凑标签)。
304
-
305
- **3. 详情块** — 与表格同序;场景放不进单元格,在这里展开:
306
-
307
- **1. path/file.ext:123 — 类别** *(判定)*
308
- 概述:<一句话>
309
- 场景:<具体的输入/状态 → 错误输出/崩溃;若是清理类发现,写明具体代价——重复了什么、
310
- 浪费了什么、哪里更难维护,或违反了哪条规则>
311
-
312
- If more than `{cap}` survive, keep the `{cap}` most severe (correctness outranks
313
- cleanup/altitude/conventions when cutting). If nothing survives, print the header
314
- line with count 0 and skip the table and details — don't emit an empty table.
315
-
316
- ### Single-pass honesty
317
-
318
- If this review did not actually fan out — low effort, or medium+ where the
319
- `subagent` tool was unavailable so the angles ran sequentially in one context —
320
- state clearly in the header line that this was a single-pass review done without the
321
- multi-agent fan-out, so whoever reads it isn't misled about what actually ran.
381
+ Report the findings via the `review_report` tool (this extension's counterpart
382
+ to CC's `ReportFindings`) — call it **once** with
383
+ `{ level, target, files_changed, fanned_out, report_id, findings }`, findings
384
+ ranked most-severe first (empty array if nothing survived verification). The
385
+ tool renders the Chinese Markdown report (table + details) back to the
386
+ conversation AND writes a machine-readable JSON to `<cwd>/.pi/review/` for CI /
387
+ `--fix` / `--comment`. Do **not** also hand-write the Markdown table.
388
+
389
+ Each finding in the array carries: `file`, `line` (optional), `category`
390
+ (`correctness` / `reuse` / `simplification` / `efficiency` / `altitude` /
391
+ `conventions`, or a more specific slug like `test-coverage`), `verdict`
392
+ (`CONFIRMED` / `PLAUSIBLE`), `short_summary` (≤60 字符、纯声明——去掉理由与
393
+ 后果,汇总表概述列优先使用它;示例:`"off-by-one in loop bound"`),
394
+ `summary` (一行中文,含理由与后果,详情块使用), `failure_scenario`
395
+ (concrete input/state → wrong output/crash; for cleanup findings, the
396
+ concrete cost — Chinese). When re-reporting after applying `--fix`, set
397
+ `outcome` on each finding (`fixed` / `skipped` / `no_change_needed` — CC
398
+ `ReportFindings` 三档,2.1.227 实证).
399
+
400
+ Cap = `maxFindings` from the effort table: `min(files_changed, 4)` at low; 8
401
+ at medium; 10 at high; 15 at xhigh/max. If more than the cap survive, send
402
+ the cap most severe (correctness outranks cleanup/altitude/conventions when
403
+ cutting; CONFIRMED outranks PLAUSIBLE). If nothing survives, send an empty
404
+ `findings` array — the tool prints a zero-count header.
405
+
406
+ **Fixed-later 义务** (CC `Q8m`): if, after this report, any later work in this
407
+ session fixes one of the reported findings (a user-requested fix, or a fix
408
+ that comes along with other changes), you MUST call `review_report` again
409
+ with the same `report_id`, the same findings, and updated `outcome` values —
410
+ before writing any text summary. The re-report updates states only; it does
411
+ not repeat the findings text. Generate the `report_id` (e.g. `review-<ts>`)
412
+ on the first report and reuse it on every re-report so consumers can merge
413
+ the files by id.
414
+
415
+ **全部用中文**:`summary` 与 `failure_scenario` 一律中文;`verdict`、`category`、
416
+ `outcome` 作为标识符保留英文 token。
417
+
418
+ **`fanned_out` 诚实** — 准确设置:仅当多智能体 fan-out 真的跑起来(subagent
419
+ finder + verify agent)才为 `true`;low effort 或任何单遍/自审降级为 `false`。该
420
+ 字段会出现在报告表头,让读者不被误导(替代旧的 Single-pass honesty 小节)。
421
+
422
+ **降级** — 若 `review_report` 工具未注册(这份 SKILL.md 跑在 pi-review 扩展之外),
423
+ 退回到直接打印 Markdown 表格 + 详情块文本;不要报错。
322
424
 
323
425
  ---
324
426
 
@@ -329,8 +431,14 @@ findings to the working tree instead of stopping at the report: fix each one
329
431
  directly — correctness bugs and reuse/simplification/efficiency cleanups alike.
330
432
  Skip any finding whose fix would change intended behavior, require changes well
331
433
  outside the reviewed diff, or that you judge to be a false positive — note the
332
- skip rather than arguing with it. Finish with a brief summary of what was fixed
333
- and what was skipped.
434
+ skip rather than arguing with it. Then call `review_report` once more to
435
+ re-report (same `report_id`), setting `outcome` on each finding (`fixed` =
436
+ applied and verified / `skipped` = real but not applied, incl. reverted /
437
+ `no_change_needed` = not applicable or already handled). This structured
438
+ re-report replaces the hand-written summary and makes the fix result
439
+ machine-consumable.
440
+ If `review_report` is unavailable, fall back to a brief text summary of what was
441
+ fixed and what was skipped.
334
442
 
335
443
  ## Posting comments (--comment)
336
444