@fyeeme/pi-review 1.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/LICENSE ADDED
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 fyeeme
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
package/README.md ADDED
@@ -0,0 +1,123 @@
1
+ # pi-review
2
+
3
+ Review & cleanup extension for [pi](https://github.com/earendil-works/pi).
4
+
5
+ Registers two commands (`/code-review` and `/code-simplify`) and a general-purpose
6
+ `subagent` tool that spawns parallel pi subprocesses — providing the **real
7
+ fan-out capability** that the `code-review` and `simplify` skills
8
+ (bundled in this package under `skills/`) need for their multi-agent flows.
9
+
10
+ ## Why
11
+
12
+ Both skills instruct the agent to fan out sub-agents (finders / verify /
13
+ cleanup angles), but `pi-subagents` does not exist in pi — so the agent
14
+ silently degraded to a sequential self-sweep. This extension ships the actual
15
+ fan-out primitive: an LLM-callable `subagent` tool that spawns real
16
+ `pi --mode json` subprocesses.
17
+
18
+ This is the "tool + prompt" architecture: the **tool** provides deterministic
19
+ dispatch (how many agents, parallelism, abort), the **skill** provides the
20
+ review/cleanup semantics. CC's own `/code-review` and `/code-simplify` work the same
21
+ way — one general Agent tool, prompt decides how to use it.
22
+
23
+ ## Install
24
+
25
+ This package peers on `@earendil-works/pi-coding-agent` / `pi-ai` + `typebox`.
26
+ From the package dir:
27
+
28
+ ```sh
29
+ npm install
30
+ ```
31
+
32
+ Then point pi at it (e.g. via your extensions config), or symlink into your pi
33
+ extensions directory.
34
+
35
+ ## What it registers
36
+
37
+ ### `/code-review` command
38
+
39
+ ```
40
+ /code-review [low|medium|high|xhigh|max] [--fix] [--comment] [--share] [<target>]
41
+ ```
42
+
43
+ Parses args, then asks the agent to load the bundled `skills/code-review/SKILL.md`
44
+ and follow it — using the `subagent` tool for any fan-out / verify / gap-hunt.
45
+
46
+ ### `/code-simplify` command
47
+
48
+ ```
49
+ /code-simplify [<target>]
50
+ ```
51
+
52
+ Cleanup (reuse / simplification / efficiency / altitude) via the `simplify`
53
+ skill. **The handler decides parallel vs single-pass mode deterministically**
54
+ from `ctx.getContextUsage()` (real token count) + whether the `subagent` tool is
55
+ registered — mirroring CC's `Jvo` guard:
56
+
57
+ - context < 80% full AND `subagent` tool available → **parallel** (4 cleanup
58
+ agents via `subagent` mode: parallel)
59
+ - otherwise → **single-pass** (inline 4 angles)
60
+
61
+ This is the deterministic mode selection a pure-prompt skill cannot reproduce
62
+ (the skill has no access to context-token count; only extension code can call
63
+ `ctx.getContextUsage()`). The decision is announced in the trigger message so
64
+ it is observable.
65
+
66
+ ### `subagent` tool
67
+
68
+ An LLM-callable tool that spawns one or more real pi subprocesses:
69
+
70
+ | mode | behavior |
71
+ |---|---|
72
+ | `single` | run `prompts[0]` once (e.g. an independent verify agent) |
73
+ | `parallel` | run all prompts concurrently, capped at 8 (e.g. one finder per angle) |
74
+ | `chain` | run sequentially; each later prompt receives prior output |
75
+
76
+ Each sub-agent is a full `pi --mode json -p --no-session` run. Progress streams
77
+ to the TUI via `onUpdate` as each agent completes. ESC aborts the whole batch
78
+ (SIGTERM → 5s → SIGKILL per subprocess). Errors are thrown (not returned) so
79
+ the agent loop marks the result `isError`.
80
+
81
+ ## Architecture (layered)
82
+
83
+ ```
84
+ pi-review/
85
+ ├── index.ts factory: registerTool(subagent) + 2 commands
86
+ ├── skills/ bundled SKILL.md files (code-review, simplify)
87
+ ├── src/
88
+ │ ├── agent/dispatch.ts spawnAgent + mapWithConcurrencyLimit (self-contained copy
89
+ │ │ from pi-dynamic-workflows; no external dep beyond node + pi-ai)
90
+ │ ├── skills.ts bundledSkillPath — resolve this extension's own skills/ dir
91
+ │ ├── tools/subagent.ts defineTool("subagent") — generic capability layer
92
+ │ └── commands/
93
+ │ ├── code-review.ts /code-review handler + sticky last-used effort (CC 2.1.223)
94
+ │ └── code-simplify.ts /code-simplify handler + decideSimplifyMode (Jvo guard)
95
+ └── test/ commands unit tests
96
+ ```
97
+
98
+ The layout is deliberately layered: `src/tools/` is the **generic capability
99
+ layer** (subagent tool + dispatch), `src/commands/` is the **entry layer** (one
100
+ file per skill). If a third or fourth skill needs the subagent tool, `src/tools/`
101
+ can be split into its own `pi-subagent` extension with zero refactor — the code
102
+ is already separated.
103
+
104
+ `src/agent/dispatch.ts` is a self-contained copy of the spawn pattern from
105
+ `examples/extensions/subagent` and `pi-dynamic-workflows/src/agent/dispatch.ts`
106
+ (~150 lines). When pi promotes `spawnAgent` to a public `pi-coding-agent`
107
+ export, this file should be deleted in favor of that import.
108
+
109
+ ## Relation to the skills
110
+
111
+ | layer | home | role |
112
+ |---|---|---|
113
+ | review/cleanup semantics (angles, verdicts, mode bodies) | `skills/code-review/` + `skills/simplify/` (bundled in this package) | what to look for |
114
+ | fan-out dispatch + mode decision | this extension (`subagent` tool + command handlers) | how to run sub-agents / which mode |
115
+
116
+ Edit a skill to change *what* it hunts; edit this extension to change *how*
117
+ sub-agents are spawned and *which mode* is chosen.
118
+
119
+ ## Status
120
+
121
+ MVP. Phase 2 (not yet built): a `review_verify` tool encapsulating 3-vote
122
+ adversarial verify, and a `review_report` tool enforcing the output schema /
123
+ `--share` lavish artifact.
package/index.ts ADDED
@@ -0,0 +1,28 @@
1
+ /**
2
+ * pi-review — extension entry.
3
+ *
4
+ * Registers:
5
+ * - the `subagent` tool — general-purpose parallel/sequential sub-agent fan-out
6
+ * via real pi subprocesses. Shared capability used by both skills below;
7
+ * - the `/code-review` command — effort-level review via the code-review skill;
8
+ * - the `/code-simplify` command — cleanup via the simplify skill; the handler
9
+ * decides parallel vs single-pass from ctx.getContextUsage(), mirroring CC's
10
+ * Jvo guard (a deterministic decision a pure-prompt skill cannot reproduce).
11
+ *
12
+ * Both skills ship bundled in this package under `skills/` — this extension
13
+ * provides the entry commands + the fan-out capability they need.
14
+ *
15
+ * Layout (layered so the tool layer can be split into its own extension later):
16
+ * src/tools/subagent.ts — generic capability (subagent tool + dispatch)
17
+ * src/commands/*.ts — per-skill entry commands
18
+ */
19
+ import type { ExtensionAPI } from "@earendil-works/pi-coding-agent";
20
+ import { registerCodeReview } from "./src/commands/code-review.ts";
21
+ import { registerSimplify } from "./src/commands/code-simplify.ts";
22
+ import { subagentTool } from "./src/tools/subagent.ts";
23
+
24
+ export default function (pi: ExtensionAPI): void {
25
+ pi.registerTool(subagentTool);
26
+ registerCodeReview(pi);
27
+ registerSimplify(pi);
28
+ }
package/package.json ADDED
@@ -0,0 +1,52 @@
1
+ {
2
+ "name": "@fyeeme/pi-review",
3
+ "version": "1.0.0",
4
+ "description": "Review & cleanup extension for pi. Registers /code-review and /code-simplify commands plus a general-purpose `subagent` tool that spawns parallel pi subprocesses — providing the real fan-out capability the code-review and simplify skills (bundled under `skills/`) need for their multi-agent flows. The /code-simplify handler uses ctx.getContextUsage() to decide parallel vs single-pass mode deterministically.",
5
+ "type": "module",
6
+ "license": "MIT",
7
+ "author": "fyeeme",
8
+ "engines": {
9
+ "node": ">=18"
10
+ },
11
+ "keywords": [
12
+ "pi-package",
13
+ "pi",
14
+ "code-review",
15
+ "simplify",
16
+ "cleanup",
17
+ "subagent",
18
+ "fan-out",
19
+ "review"
20
+ ],
21
+ "files": [
22
+ "*.ts",
23
+ "src/**/*.ts",
24
+ "skills/**/*.md",
25
+ "README.md",
26
+ "LICENSE"
27
+ ],
28
+ "pi": {
29
+ "extensions": [
30
+ "./index.ts"
31
+ ]
32
+ },
33
+ "scripts": {
34
+ "test": "vitest --run",
35
+ "typecheck": "tsc"
36
+ },
37
+ "peerDependencies": {
38
+ "@earendil-works/pi-ai": ">=0.77.0",
39
+ "@earendil-works/pi-coding-agent": ">=0.77.0",
40
+ "jiti": ">=2.0.0",
41
+ "typebox": ">=1.0.0",
42
+ "typescript": ">=5.0.0"
43
+ },
44
+ "devDependencies": {
45
+ "@earendil-works/pi-ai": "0.77.0",
46
+ "@earendil-works/pi-coding-agent": "0.77.0",
47
+ "@types/node": "22.19.19",
48
+ "jiti": "2.7.0",
49
+ "typebox": "1.1.38",
50
+ "typescript": "5.9.3"
51
+ }
52
+ }
@@ -0,0 +1,369 @@
1
+ ---
2
+ name: code-review
3
+ description: "Review the current diff for correctness bugs and reuse/simplification/efficiency cleanups at the given effort level. Fresh reverse of CC `/code-review` (CLI v2.1.223). Effort semantics: medium = precision, high+ = recall. Pass --fix to apply, --comment to post inline PR comments, --share to publish a review page."
4
+ ---
5
+
6
+ <!--
7
+ Origin: Claude Code built-in skill `/code-review` (CLI v2.1.223), freshly
8
+ reverse-engineered 2026-08-06 from bin/claude.exe strings. This file is
9
+ sourced DIRECTLY from the 2.1.223 binary — NOT carried forward from the
10
+ earlier v2.1.220 reconstruction. Every section below was located in the
11
+ extracted strings (cc_strings_223.txt) and verified.
12
+
13
+ What CC 2.1.223 actually contains (verified against the binary):
14
+ - Effort: medium = precision; high = recall ("err on the side of
15
+ surfacing"); xhigh/max add a gap-hunt. Each finder surfaces ≤6 candidates.
16
+ - Low effort: 1 diff pass, no verify, target min(files_changed, 4) findings.
17
+ - Finder allocation (workflow Find-phase, linearized inline for Pi): one
18
+ finder per correctness angle (A–E) + Conventions, plus one combined cleanup
19
+ finder, pooled before verify. (CC's native *inline* medium path splits
20
+ differently — 8 finders: 3 correctness + 3 cleanup + altitude + conventions.)
21
+ - Angles A–E + Reuse/Simplification/Efficiency/Altitude + Conventions,
22
+ verbatim (same source variables the /simplify skill reuses).
23
+ - Verify via an independent agent: CONFIRMED / PLAUSIBLE / REFUTED,
24
+ "PLAUSIBLE by default". Keep CONFIRMED + PLAUSIBLE, drop REFUTED.
25
+ - Gap-hunt (xhigh/max): one fresh finder hunting only for gaps not
26
+ already listed (CC's Sweep phase: "Fresh finder hunting only for gaps").
27
+ - Output: Markdown findings table + per-finding details block, printed
28
+ as text (no ReportFindings tool on Pi); carries file:line/category/
29
+ verdict/summary/failure_scenario per finding.
30
+ - --share publishes an Artifact; --fix applies findings to the working tree.
31
+
32
+ Invocation: /code-review [low|medium|high|xhigh|max] [--fix] [--comment] [--share] [<target>]
33
+ target = Class#method | file path | PR number | branch name
34
+ With no level given, the /code-review HANDLER reuses the last level you
35
+ typed (CC 2.1.223 codeReviewLastEffort); the skill always receives a
36
+ concrete level.
37
+ (CC also supports `ultra` — deep multi-agent review in the cloud.
38
+ OMITTED: requires claude.ai cloud access, which Pi does not provide.)
39
+
40
+ ════════════════════════════════════════════════════════════════════════
41
+ Pi ADAPTATIONS (differ from the CC runtime)
42
+ ════════════════════════════════════════════════════════════════════════
43
+ 1. Output — CC calls a ReportFindings tool with {level, findings}; Pi
44
+ PRINTS a Markdown findings table + details block as text
45
+ (no such tool on Pi). [was JSON array; switched for
46
+ readability]
47
+ 2. Fan-out — CC uses the Agent tool; Pi uses the `subagent` tool
48
+ (mode: parallel), or runs angles sequentially if unavailable.
49
+ 3. Verify — CC uses the Agent tool; Pi uses `subagent` for the
50
+ independent verify agent (fallback: self-check).
51
+ 4. Workflow — CC routes high/xhigh/max to a background Workflow (phases
52
+ Scope/Find/Verify/Sweep/Synthesize) when workflows are
53
+ enabled; Pi has no such tool, so this skill runs INLINE and
54
+ linearizes those phases into the flow below.
55
+ 5. ultra — dropped (cloud-only).
56
+ 6. --share — CC uses the Artifact tool; Pi uses lavish-axi.
57
+ 7. --comment— CC uses mcp__github_inline_comment; Pi falls back to gh api
58
+ or printing.
59
+
60
+ Prerequisite: the `subagent` tool (pi-review extension; mode: parallel) for
61
+ medium and above, and for the xhigh/max gap-hunter. lavish-axi
62
+ for --share. low runs standalone (no subagents).
63
+ -->
64
+
65
+ You are reviewing the current diff for correctness bugs and reuse /
66
+ simplification / efficiency cleanups. Correctness bugs always outrank cleanup,
67
+ altitude, and conventions findings when the output cap forces a cut.
68
+
69
+ ## Effort levels
70
+
71
+ | Level | Intent | Verify | Subagents | Output cap |
72
+ |-------|--------|--------|-----------|------------|
73
+ | low (default) | quick scan | no | no | min(files_changed, 4) |
74
+ | medium | **precision** — surface only findings a maintainer would act on | independent agent | fan-out (1 finder/angle) | ≤ 8 |
75
+ | high | **recall** — catch every real bug a careful reviewer would; **err on the side of surfacing** | independent agent | more angles | ≤ 10 |
76
+ | xhigh → max | recall + **gap-hunt** | independent agent | above + 1 fresh gap finder | larger, may include uncertain |
77
+
78
+ Each finder surfaces **up to 6 candidate findings** with `file`, `line`, a
79
+ one-line `summary`, and a concrete `failure_scenario`.
80
+
81
+ If a target argument was provided, review that target instead of the whole diff.
82
+
83
+ ## Phase 0 — Gather the diff
84
+
85
+ Run `git diff @{upstream}...HEAD` (or `git diff main...HEAD` / `git diff HEAD~1`
86
+ if there's no upstream) to get the unified diff under review. If there are
87
+ uncommitted changes, or the range diff is empty, also run `git diff HEAD` and
88
+ include the working-tree changes in scope — the review often runs before the
89
+ commit. If a PR number, branch name, or file path was passed as an argument,
90
+ review that target instead. Treat this diff as the review scope. Note the
91
+ files-changed count — low effort uses it for the dynamic output cap.
92
+
93
+ ---
94
+
95
+ # LOW-EFFORT FLOW (default; runs standalone, no subagents)
96
+
97
+ `low effort → 1 diff pass → no verify → min(files_changed, 4) findings`
98
+
99
+ ## Turn 1 — read
100
+
101
+ One tool call: read the unified diff (`git diff @{upstream}...HEAD; git diff HEAD`
102
+ to cover both committed and uncommitted changes, or `git diff main...HEAD` / the
103
+ target passed as an argument). Skip test/fixture hunks (`test/`, `spec/`,
104
+ `__tests__/`, `*_test.*`, `*.test.*`, `fixtures/`, `testdata/`) — test-file
105
+ changes are not reviewed at this level. No subagents, no full-file reads.
106
+
107
+ ## Turn 2 — findings
108
+
109
+ Flag runtime-correctness bugs visible from the hunk alone: inverted/wrong
110
+ condition, off-by-one, null/undefined deref where adjacent lines show the value
111
+ can be absent, removed guard, falsy-zero check, missing `await`,
112
+ wrong-variable copy-paste, error swallowed in a catch that should propagate.
113
+ Also flag — still from the hunk alone — new code that duplicates an existing
114
+ helper visible in the diff context, and dead code the diff leaves behind.
115
+
116
+ Do **not** flag style, naming, perf, missing tests, or anything outside the hunk.
117
+
118
+ Target **min(files_changed, 4) findings**, most-severe first. If you have fewer,
119
+ do one more pass focused on the largest changed file and on any **removed** code
120
+ blocks. Output exactly `(none)` only if the diff is trivially correct after
121
+ that pass. Do not call a ReportFindings tool even if one is available.
122
+
123
+ ---
124
+
125
+ # MEDIUM-AND-ABOVE FLOW (fan-out + verify)
126
+
127
+ ## Phase 1 — Find candidates (single pass or parallel fan-out)
128
+
129
+ Work through the angles below. If the `subagent` tool is available, launch
130
+ finder agents in a single batch (mode: parallel) so they run concurrently;
131
+ otherwise do not fake the fan-out — work the angles yourself in sequence in
132
+ this same context, or report that the subagent capability is unavailable.
133
+
134
+ **Finder allocation** (the workflow Find-phase, linearized inline for Pi;
135
+ verbatim from the binary): **one finder per correctness angle, plus one finder
136
+ covering all cleanup angles, pooled before verify** — not a priority-sorted
137
+ packing. That is one finder each for A/B/C/D/E plus Conventions, and one
138
+ combined finder for the cleanup angles (Reuse / Simplification / Efficiency /
139
+ Altitude). Never silently drop a correctness angle; if you must consolidate,
140
+ fold cleanup into a correctness finder.
141
+
142
+ The correctness angles hunt for bugs; the cleanup angles hunt for cleanup in
143
+ the changed code. Cleanup, altitude, and conventions candidates use the same
144
+ `file`/`line`/`summary` shape; in `failure_scenario`, state the concrete cost
145
+ (what is duplicated, wasted, harder to maintain, or which CLAUDE.md rule is
146
+ broken) instead of a crash.
147
+
148
+ ### Angle A — line-by-line diff scan
149
+ Read every hunk in the diff, line by line. Then Read the enclosing function for
150
+ each hunk — bugs in unchanged lines of a touched function are in scope (the PR
151
+ re-exposes or fails to fix them). For every line ask: what input, state, timing,
152
+ or platform makes this line wrong? Look for inverted/wrong conditions,
153
+ off-by-one, null/undefined deref, missing `await`, falsy-zero checks,
154
+ wrong-variable copy-paste, error swallowed in catch, unescaped regex metachars.
155
+
156
+ ### Angle B — removed-behavior auditor
157
+ For every line the diff DELETES or replaces, name the invariant or behavior it
158
+ enforced, then search the new code for where that invariant is re-established.
159
+ If you can't find it, that's a candidate: a removed guard, a dropped error path,
160
+ a narrowed validation, a deleted test that was covering a real case.
161
+
162
+ ### Angle C — cross-file tracer
163
+ For each function the diff changes, find its callers (Grep for the symbol) and
164
+ check whether the change breaks any call site: a new precondition, a changed
165
+ return shape, a new exception, a timing/ordering dependency. Also check callees:
166
+ does a parallel change in the same PR make a call unsafe?
167
+
168
+ ### Angle D — language-pitfall specialist
169
+ Scan for the classic pitfalls of the diff's language/framework — for example:
170
+ JS falsy-zero, `==` coercion, closure-captured loop var; Python mutable default
171
+ args, late-binding closures; Go nil-map write, range-var capture; SQL injection;
172
+ timezone/DST drift; float equality. Flag any instance the diff introduces.
173
+
174
+ ### Angle E — wrapper/proxy correctness
175
+ When the PR adds or modifies a type that wraps another (cache, proxy, decorator,
176
+ adapter): check that every method routes to the wrapped instance and not back
177
+ through a registry/session/global — e.g. a caching provider holding a
178
+ `delegate` field that resolves IDs via `session.get(...)` instead of
179
+ `delegate.get(...)` will re-enter the cache or recurse. Also check that the
180
+ wrapper forwards all the methods the callers actually use.
181
+
182
+ ### Reuse
183
+ Flag new code that re-implements something the codebase already has — Grep
184
+ shared/utility modules and files adjacent to the change, and name the existing
185
+ helper to call instead.
186
+
187
+ ### Simplification
188
+ Flag unnecessary complexity the diff adds: redundant or derivable state,
189
+ copy-paste with slight variation, deep nesting, dead code left behind. Name the
190
+ simpler form that does the same job.
191
+
192
+ ### Efficiency
193
+ Flag wasted work the diff introduces: redundant computation or repeated I/O,
194
+ independent operations run sequentially, blocking work added to startup or hot
195
+ paths. Also flag long-lived objects built from closures or captured environments
196
+ — they keep the entire enclosing scope alive for the object's lifetime (a memory
197
+ leak when that scope holds large values); prefer a class/struct that copies only
198
+ the fields it needs. Name the cheaper alternative.
199
+
200
+ ### Altitude
201
+ Check that each change is implemented at the right depth, not as a fragile
202
+ bandaid. Special cases layered on shared infrastructure are a sign the fix isn't
203
+ deep enough — prefer generalizing the underlying mechanism over adding special
204
+ cases.
205
+
206
+ ### Conventions (CLAUDE.md)
207
+ Find the CLAUDE.md files that govern the changed code: the user-level
208
+ ~/.claude/CLAUDE.md, the repo-root CLAUDE.md, plus any CLAUDE.md or
209
+ CLAUDE.local.md in a directory that is an ancestor of a changed file (a
210
+ directory's CLAUDE.md only applies to files at or below it). Read each one that
211
+ exists, then check the diff for clear violations of the rules they state.
212
+
213
+ Only flag a violation when you can quote the exact rule and the exact line that
214
+ breaks it — no style preferences, no vague "spirit of the doc" inferences. In
215
+ the finding, name the CLAUDE.md path and quote the rule so the report can cite
216
+ it. If no CLAUDE.md applies, return nothing for this angle.
217
+
218
+ ### Pass every candidate through
219
+ Pass every candidate with a nameable failure scenario through to verify —
220
+ finders that silently drop half-believed candidates bypass the verify step and
221
+ are the dominant cause of misses.
222
+
223
+ ## Phase 2 — Dedup and verify
224
+
225
+ Dedup near-duplicates (same defect, same location, same reason → keep one).
226
+
227
+ Then verify each candidate. If the `subagent` tool is available, dispatch an
228
+ independent verify agent (one per candidate, or a small batch): give it the
229
+ diff, the relevant file(s), and the candidate; it returns exactly one of
230
+ **CONFIRMED / PLAUSIBLE / REFUTED**. An independent agent counters the
231
+ confirmation bias of self-review. If `subagent` is unavailable, fall back to
232
+ re-checking each candidate yourself (self-check). Keep **CONFIRMED and
233
+ PLAUSIBLE**, drop REFUTED. Give each surviving finding a verdict:
234
+
235
+ - **CONFIRMED** — can name the inputs/state that trigger it and the wrong
236
+ output or crash. Quote the line.
237
+ - **PLAUSIBLE** — mechanism is real, trigger is uncertain (timing, env,
238
+ config). State what would confirm it.
239
+ - **REFUTED** — factually wrong (code doesn't say that) or guarded elsewhere.
240
+ Quote the line that proves it.
241
+
242
+ **PLAUSIBLE by default** — do not refute a candidate for being "speculative" or
243
+ "depends on runtime state" when the state is realistic: concurrency races,
244
+ nil/undefined on a rare-but-reachable path (error handler, cold cache, missing
245
+ optional field), falsy-zero treated as missing, off-by-one on a boundary the
246
+ code does not exclude, retry storms / partial failures, regex/allowlist that
247
+ lost an anchor. These are PLAUSIBLE.
248
+
249
+ **REFUTED** only when constructible from the code: factually wrong (quote the
250
+ actual line); provably impossible (type/constant/invariant — show it); already
251
+ handled in this diff (cite the guard); or pure style with no observable effect.
252
+
253
+ ## Phase 3 — Gap-hunt (xhigh / max only)
254
+
255
+ At **xhigh and max**, after Phase 2 dedup, dispatch ONE fresh finder agent (the
256
+ `subagent` tool) that has never seen the candidates and hunts only for gaps not
257
+ already listed.
258
+
259
+ Constrain it so exploration can't run away (Pi adaptation — CC's workflow bounds
260
+ this differently):
261
+
262
+ 1. **Pre-embed all context in the prompt** — do not tell the agent to search for
263
+ callers/dependencies itself. Do those searches here first and embed the
264
+ results: the diff, the enclosing functions, the deduplicated finding list,
265
+ and any search results. The gap-hunt agent **analyzes**, it does not
266
+ **discover**.
267
+ 2. **Set `maxTurns: 15`** on the `subagent` call — caps it at 15 assistant turns.
268
+ 3. **Declare a tool-call budget in the prompt** — e.g. "You have ONLY 3 tool
269
+ calls to read files. Read them now, then analyze from this message's
270
+ context."
271
+
272
+ Feed anything it finds back through Phase 2 verify before keeping it. If the
273
+ `subagent` tool is unavailable, take one self-sweep instead and note the
274
+ gap-hunt was self-run (lacks the independent fresh-eyes benefit).
275
+
276
+ At **high and below**, skip Phase 3.
277
+
278
+ ## Output
279
+
280
+ Print the findings as a **Markdown table + a details block** — readable in a
281
+ terminal and in rendered Markdown (no JSON, no ReportFindings tool on Pi). Cap =
282
+ low's min(files_changed, 4); 8 at medium; 10 at high; larger at xhigh → max.
283
+
284
+ **全部用中文输出**:表头、概述、场景一律用中文;`Verdict`、`Category` 作为标识符保留
285
+ 英文 token(CONFIRMED / PLAUSIBLE;correctness / reuse …)。
286
+
287
+ **1. 表头行** — 单行写明:力度、diff 命令/范围、改动文件数、命中条数、是否真的多智能体
288
+ 并发(见下文 Single-pass honesty)。示例:
289
+
290
+ > `max` · `git diff HEAD` · 29 个文件 · 3 条发现 · 多智能体(验证 + 查漏)
291
+
292
+ **2. 发现汇总表** — 按严重程度从高到低,每条一行:
293
+
294
+ | # | 判定 | 类别 | 位置 | 概述 |
295
+ |---|------|------|------|------|
296
+ | 1 | CONFIRMED | correctness | path/file.ext:123 | 一句话说明这个 bug |
297
+ | 2 | PLAUSIBLE | reuse | path/file.ext:45 | … |
298
+
299
+ - `判定` — `CONFIRMED` / `PLAUSIBLE` / 留空(未做验证)。
300
+ - `类别` — 产生该发现的角度,短横线小写 slug(`correctness`、`simplification`、
301
+ `efficiency`、`reuse`、`altitude`、`conventions`,或更具体的如 `test-coverage`)。
302
+ - `位置` — `文件:行号`。
303
+ - `概述` — 一句话说明(≤ 约 80 字,同时作为紧凑标签)。
304
+
305
+ **3. 详情块** — 与表格同序;场景放不进单元格,在这里展开:
306
+
307
+ **1. path/file.ext:123 — 类别** *(判定)*
308
+ 概述:<一句话>
309
+ 场景:<具体的输入/状态 → 错误输出/崩溃;若是清理类发现,写明具体代价——重复了什么、
310
+ 浪费了什么、哪里更难维护,或违反了哪条规则>
311
+
312
+ If more than `{cap}` survive, keep the `{cap}` most severe (correctness outranks
313
+ cleanup/altitude/conventions when cutting). If nothing survives, print the header
314
+ line with count 0 and skip the table and details — don't emit an empty table.
315
+
316
+ ### Single-pass honesty
317
+
318
+ If this review did not actually fan out — low effort, or medium+ where the
319
+ `subagent` tool was unavailable so the angles ran sequentially in one context —
320
+ state clearly in the header line that this was a single-pass review done without the
321
+ multi-agent fan-out, so whoever reads it isn't misled about what actually ran.
322
+
323
+ ---
324
+
325
+ ## Applying fixes (--fix)
326
+
327
+ The `--fix` flag was passed. After producing the findings list, apply the
328
+ findings to the working tree instead of stopping at the report: fix each one
329
+ directly — correctness bugs and reuse/simplification/efficiency cleanups alike.
330
+ Skip any finding whose fix would change intended behavior, require changes well
331
+ outside the reviewed diff, or that you judge to be a false positive — note the
332
+ skip rather than arguing with it. Finish with a brief summary of what was fixed
333
+ and what was skipped.
334
+
335
+ ## Posting comments (--comment)
336
+
337
+ The `--comment` flag was passed. Post the findings as inline PR comments on the
338
+ corresponding `file`/`line`. If no GitHub commenting tool is available on Pi,
339
+ fall back to printing the findings as text and note that inline posting was
340
+ unavailable.
341
+
342
+ ## Publishing a shareable review (--share)
343
+
344
+ The `--share` flag was passed. After producing the findings list, also publish
345
+ them as an artifact so they can be shared and iterated on outside the terminal.
346
+
347
+ 1. Write a self-contained HTML review page to `.lavish/code-review-<n>.html`
348
+ (create `.lavish/` in the repo root if missing). The page must render with no
349
+ server and carry every finding plus its context.
350
+ 2. Open it with `lavish-axi .lavish/code-review-<n>.html` so the reader can
351
+ review, annotate, and send feedback back through the poll.
352
+
353
+ Page structure (follow lavish design guidance — clear visual hierarchy, no
354
+ horizontal overflow at any nesting level, monospace for code/paths, color-code
355
+ verdicts):
356
+
357
+ - **Header**: effort level, target, the diff command that was run, files-changed
358
+ count, and whether the review actually fanned out (single-pass honesty).
359
+ - **Findings table**: one row per finding — `file:line`, `category`, `verdict`
360
+ (CONFIRMED = red, PLAUSIBLE = amber, unset = grey), one-line `summary`, and
361
+ the full `failure_scenario`.
362
+ - **Verdict legend**: a short note on what CONFIRMED vs PLAUSIBLE mean, so a
363
+ non-author reader can discount the uncertain ones.
364
+ - **Context pins**: the changed-files list, the applicable CLAUDE.md files, and
365
+ the conventions that were checked.
366
+
367
+ Skip the artifact if the review was invoked only to feed another tool (e.g.
368
+ `--fix`, where the caller applies its own changes) — note the skip in the
369
+ summary so the absence of a page is not mistaken for a failure.