@fyeeme/pi-review 2.0.0 → 2.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -1,157 +1,103 @@
1
- # pi-review
1
+ # @fyeeme/pi-review (v2)
2
2
 
3
- Review & cleanup extension for [pi](https://github.com/earendil-works/pi).
3
+ **2.0.0 — major release** (from 1.1.1), part of the 2.0 extensions family wave:
4
4
 
5
- Registers two commands (`/code-review` and `/code-simplify`) and a general-purpose
6
- `subagent` tool that spawns parallel pi subprocesses — providing the **real
7
- fan-out capability** that the `code-review` and `simplify` skills
8
- (bundled in this package under `skills/`) need for their multi-agent flows.
5
+ - **Sandwich architecture** — skills carry the methodology, `prompts/` carries the orchestration strategy as data (parallel-when guards in frontmatter, phases in the body), `agents/` carries the 12 review roles (finder A–E, cleaner-reuse/simplification/efficiency, altitude, conventions, verifier, gap-hunter), and a thin plugin entry composes the stack.
9
6
 
10
- ## Why
7
+ - **Batteries-included fan-out** — the entry composes [`@fyeeme/pi-subagents`](https://www.npmjs.com/package/@fyeeme/pi-subagents) 2.1.1 from its npm dependency: the `subagent` tool, real `pi` subprocess spawning, and the live agent UI (widget / FleetView / `/agents`) work out of the box.
11
8
 
12
- Both skills instruct the agent to fan out sub-agents (finders / verify /
13
- cleanup angles), but `pi-subagents` does not exist in pi — so the agent
14
- silently degraded to a sequential self-sweep. This extension ships the actual
15
- fan-out primitive: an LLM-callable `subagent` tool that spawns real
16
- `pi --mode json` subprocesses.
9
+ - **`review_report` structured findings sink** — Chinese Markdown rendered back to the conversation plus machine-readable JSON under `<cwd>/.pi/review/` for CI, `--fix` re-reports, and `--comment`.
17
10
 
18
- This is the "tool + prompt" architecture: the **tool** provides deterministic
19
- dispatch (how many agents, parallelism, abort), the **skill** provides the
20
- review/cleanup semantics. CC's own `/code-review` and `/code-simplify` work the same
21
- way — one general Agent tool, prompt decides how to use it.
11
+ - **Effort split (v2.1)** — `/code-review [low|medium|high|xhigh|max]`: low/medium/high review the diff in ONE pass in the main session (no subagents: rubric-ported flag criteria, P0–P3 priorities, in-session self-verify at medium/high); xhigh/max keep the opt-in deep sweep — quad tuples `{correctnessAngles, perAngle, maxFindings, sweep}`, grouped-by-location independent verification, and the gap-hunt.
22
12
 
23
- ## Install
13
+ - **`/code-simplify` dual-mode** — the dispatcher measures context usage and diff size against the declared strategy, then renders either the PARALLEL template (4 cleaner agents via `subagent`) or the SINGLE-PASS one.
24
14
 
25
- This package peers on `@earendil-works/pi-coding-agent` / `pi-ai` + `typebox`.
26
- From the package dir:
15
+ - Commands: `/code-review` and `/code-simplify` — restored v1 names (2.0.0 briefly renamed them `/review` / `/simplify`).
27
16
 
28
- ```sh
29
- npm install
17
+ Review & cleanup assets for [pi](https://github.com/earendil-works/pi-mono), in the sandwich shape (skills + prompts + agents on top of a thin plugin entry):
18
+
19
+ ```
20
+ skills/ methodology (code-review, code-simplify) — registered natively via the `pi` manifest
21
+ prompts/ orchestration strategy as data — parallel-when guards in frontmatter,
22
+ CC-parity phase structure in the body; rendered by the generic dispatcher
23
+ agents/ the review roles as subagent definitions (finder-*, cleaner-*, verifier,
24
+ gap-hunter) invoked via the `subagent` tool of @fyeeme/pi-subagents
25
+ index.ts plugin entry: the review_report structured findings sink + the
26
+ /code-review and /code-simplify dispatcher commands
27
+ src/ dispatch.ts (variable gathering, guard evaluation, rendering),
28
+ diff.ts (deterministic diff ladder — unchanged v1 semantics),
29
+ strategy.ts (guard evaluator), tools/review_report.ts
30
30
  ```
31
31
 
32
- This resolves [`@fyeeme/pi-subagent-core`](https://www.npmjs.com/package/@fyeeme/pi-subagent-core)
33
- (`^0.3.0`, from the npm registry — no sibling-repo layout requirement).
32
+ **Batteries included**: installing this package is enough. Its extension
33
+ factory composes [`@fyeeme/pi-subagents`](../pi-subagents) (the `subagent`
34
+ tool, spawning, and the live agent UI — widget / FleetView / `/agents`) from
35
+ the version-pinned dependency copy, so every fan-out the skills orchestrate
36
+ works out of the box. A standalone pi-subagents install is optional
37
+ (general-purpose scout/planner/reviewer/worker agents) and coexists —
38
+ composition is idempotent.
39
+
40
+ ## Commands
34
41
 
35
- Then point pi at it (e.g. via your extensions config), or symlink into your pi
36
- extensions directory.
42
+ - `/code-review [low|medium|high|xhigh|max] [--fix] [--loop] [--comment] [--share] [<pr#>|<branch>|<path>]` — effort-level code review via the code-review skill. low/medium/high run as a single pass in this session (fast path, default); xhigh/max fan out finder/verifier/gap-hunt agents through `subagent`. Effort is sticky: an explicit level is remembered; the next bare `/code-review` reuses it. `--loop` (single-pass levels only) drives extension-orchestrated fix→re-review rounds (≤ `maxTurns.loop`, default 3) until no P0/P1 findings remain.
43
+ - `/code-simplify [<target>]` — cleanup of the changed code (reuse/simplification/efficiency/altitude). The dispatcher resolves the diff (upstream merge-base → HEAD worktree → staged → unstaged; submodule-aware), evaluates the strategy declared in `prompts/simplify.parallel.md` frontmatter (context usage < 80%, diff < 400k chars, fan-out available), and renders either the PARALLEL template (Phase 0 visible diff read → `subagent` parallel dispatch of the 4 cleaner agents with `maxTurns: 15` → Phase 2 apply/verify/report) or the SINGLE-PASS template (angles worked inline).
37
44
 
38
- ## What it registers
45
+ Reports land via the `review_report` tool: Chinese Markdown back to the conversation plus machine-readable JSON under `<cwd>/.pi/review/`.
39
46
 
40
- ### `/code-review` command
47
+ ## Strategy is data
41
48
 
42
49
  ```
43
- /code-review [low|medium|high|xhigh|max] [--fix] [--comment] [--share] [<target>]
50
+ # prompts/simplify.parallel.md
51
+ ---
52
+ parallel-when:
53
+ context-below: 0.8
54
+ diff-chars-below: 400000
55
+ ---
44
56
  ```
45
57
 
46
- Parses args, then asks the agent to load the bundled `skills/code-review/SKILL.md`
47
- and follow it — using the `subagent` tool for any fan-out / verify / gap-hunt.
58
+ Edit the file, the strategy changes. The dispatcher only executes what the templates declare (the unmeasurable-context and recursion-guard fallbacks stay as code invariants). See `test/dispatch.test.ts` for the anchored semantics.
48
59
 
49
- ### `/code-simplify` command
60
+ ## Configuration
50
61
 
62
+ Turn budgets are configurable via a JSON file, following the same two-layer
63
+ pattern as pi-subagents' `pi-subagent.json` (project overrides global):
64
+
65
+ - Global: `<agentDir>/pi-review.json`
66
+ - Project: `<cwd>/.pi/pi-review.json`
67
+
68
+ ```jsonc
69
+ // <any layer>/pi-review.json — all keys optional
70
+ {
71
+ "maxTurns": {
72
+ "subagent": 20, // each /code-review finder-batch subagent call
73
+ "verifier": 15, // each /code-review Phase 2 verifier call
74
+ "gapHunt": 15, // the /code-review Phase 3 gap-hunter (xhigh/max)
75
+ "simplify": 15, // each /code-simplify PARALLEL cleaner agent
76
+ "loop": 3 // --loop fix→re-review round cap (single-pass levels)
77
+ }
78
+ }
51
79
  ```
52
- /code-simplify [<target>]
53
- ```
54
80
 
55
- Cleanup (reuse / simplification / efficiency / altitude) via the `simplify`
56
- skill. **The handler decides parallel vs single-pass mode deterministically**
57
- from `ctx.getContextUsage()` (real token count) + whether the `subagent` tool is
58
- registered — mirroring CC's `Jvo` guard:
59
-
60
- - context < 80% full AND `subagent` tool available → **parallel** (4 cleanup
61
- agents via `subagent` mode: parallel)
62
- - otherwise → **single-pass** (inline 4 angles)
63
-
64
- This is the deterministic mode selection a pure-prompt skill cannot reproduce
65
- (the skill has no access to context-token count; only extension code can call
66
- `ctx.getContextUsage()`). The decision is announced in the trigger message so
67
- it is observable.
68
-
69
- **Apply → verify → revert safety net** (harden-code-simplify): after Phase 2
70
- applies the cleanups, the handler also injects a verification command detected
71
- from `package.json` scripts (`check` → `test` → `lint` → `typecheck`). The
72
- skill snapshots the touched files, applies the fixes, runs that command, and
73
- on failure auto-reverts per-file (a clean apply runs verify exactly once; only
74
- a failure escalates to one verify per touched file). The result is reported as
75
- structured outcomes via `review_report` (`level: "simplify"`), not a free-text
76
- summary. If no verification command is detectable, fixes are kept but the
77
- report states no verification was run (verification is opportunistic, never
78
- blocking).
79
-
80
- ### `subagent` tool
81
-
82
- An LLM-callable tool that spawns one or more real pi subprocesses:
83
-
84
- | mode | behavior |
85
- |---|---|
86
- | `single` | run `prompts[0]` once (e.g. an independent verify agent) |
87
- | `parallel` | run all prompts concurrently, capped at the ceiling (e.g. one finder per angle) |
88
- | `chain` | run sequentially; each later prompt receives prior output |
89
-
90
- **Fan-out guards** (harden-code-simplify, shared with `/code-review`):
91
-
92
- - **Recursion cap (whitelist-by-default)** — a spawned sub-agent does not
93
- receive the `subagent` tool in its default toolset, so it cannot recurse. A
94
- caller opts in by listing `subagent` in the child's `tools` whitelist; set
95
- `PI_SUBAGENT_MAX_SPAWN_DEPTH` to allow multi-level fan-out up to a hard cap.
96
- - **Default turn budget** — fan-out agents get a finite default `maxTurns` (50,
97
- aligned to CC's `FORKED_AGENT_DEFAULT_MAX_TURNS` in 2.1.227)
98
- when the caller omits it; an explicit `0` is honored.
99
- - **Configurable concurrency** — `PI_MAX_CONCURRENT_SUBAGENTS` (default 20,
100
- aligned to CC's `CLAUDE_CODE_MAX_CONCURRENT_SUBAGENTS ?? 20` in 2.1.227;
101
- invalid values fall back to the default).
102
-
103
- Each sub-agent is a full `pi --mode json -p --no-session` run. Progress streams
104
- to the TUI via `onUpdate` as each agent completes. ESC aborts the whole batch
105
- (SIGTERM → 5s → SIGKILL per subprocess). Errors are thrown (not returned) so
106
- the agent loop marks the result `isError`.
107
-
108
- ## Architecture (layered)
81
+ Values must be positive integers; anything else (or an absent file) falls back
82
+ to the built-in defaults — `20` / `15` / `15` / `15` / `3`, the numbers the bundled
83
+ prompts and skills were written with — so with no configuration the rendered
84
+ instructions are byte-identical to the pre-config behavior. Files are read at
85
+ command time: an edit takes effect on the next `/code-review` or `/code-simplify`
86
+ without a restart. When a budget is configured, the trigger message states it
87
+ and the skills defer to it over their built-in defaults.
88
+
89
+ ## Requirements
90
+
91
+ None beyond this package. `@fyeeme/pi-subagents` 2.1.1 is a regular npm dependency (exact-pinned) whose extension factory this entry composes (tool + UI). The four cleaner agents and the finder/verifier/gap-hunter definitions ship with this package, registered via `addAgentDir` at extension load.
92
+
93
+ ## Development
109
94
 
110
95
  ```
111
- pi-review/
112
- ├── index.ts factory: registerTool(subagent) + 2 commands
113
- ├── skills/ bundled SKILL.md files (code-review, simplify)
114
- ├── src/
115
- │ ├── skills.ts bundledSkillPath — resolve this extension's own skills/ dir
116
- │ ├── tools/subagent.ts defineTool("subagent") — generic capability layer
117
- │ └── commands/
118
- │ ├── code-review.ts /code-review handler + sticky last-used effort (CC 2.1.223)
119
- │ └── code-simplify.ts /code-simplify handler + decideSimplifyMode (Jvo guard)
120
- └── test/ commands unit tests
96
+ npm install --ignore-scripts # @fyeeme/pi-subagents resolves from the npm registry
97
+ npm test # vitest
98
+ npm run typecheck
121
99
  ```
122
100
 
123
- The layout is deliberately layered: `src/tools/` is the **generic capability
124
- layer** (subagent tool), `src/commands/` is the **entry layer** (one
125
- file per skill). If a third or fourth skill needs the subagent tool, `src/tools/`
126
- can be split into its own `pi-subagent` extension with zero refactor — the code
127
- is already separated.
128
-
129
- The dispatch primitive (`spawnAgent`, `mapWithConcurrencyLimit`,
130
- `createSpawnRegistry`, `abortAgent`, `getPiInvocation` + types) lives in
131
- [`pi-subagent-core`](../pi-subagent-core) (npm `@fyeeme/pi-subagent-core`), a
132
- shared library extracted from the duplicated copies that used to live here and
133
- in `pi-dynamic-workflows`. When pi promotes `spawnAgent` to a public
134
- `pi-coding-agent` export, `pi-subagent-core` should be deleted in favor of that
135
- import.
136
-
137
- ## Relation to the skills
138
-
139
- | layer | home | role |
140
- |---|---|---|
141
- | review/cleanup semantics (angles, verdicts, mode bodies) | `skills/code-review/` + `skills/simplify/` (bundled in this package) | what to look for |
142
- | fan-out dispatch + mode decision | this extension (`subagent` tool + command handlers) | how to run sub-agents / which mode |
143
-
144
- Edit a skill to change *what* it hunts; edit this extension to change *how*
145
- sub-agents are spawned and *which mode* is chosen.
146
-
147
- ## Status
148
-
149
- `review_report` is built — schema aligned to CC `ReportFindings` (2.1.227
150
- empirical): 3-state `outcome` (`fixed`/`skipped`/`no_change_needed`), 2-value
151
- `verdict` (`CONFIRMED`/`PLAUSIBLE`), `short_summary` (≤60, table overview),
152
- `report_id` for fixed-later re-reports; renders the Chinese Markdown report
153
- and writes JSON to `<cwd>/.pi/review/` for CI / `--fix` / `--comment`.
154
-
155
- Remaining Phase 2 item (not yet built): a `review_verify` tool encapsulating
156
- 3-vote adversarial verify. `--share` already routes through lavish-axi (see
157
- the code-review skill).
101
+ To test local pi-subagents changes alongside this package, temporarily point the
102
+ dependency back at the sibling checkout (`file:../pi-subagents`) and reinstall;
103
+ restore the pinned registry version before publishing.
@@ -0,0 +1,18 @@
1
+ ---
2
+ name: cleaner-altitude
3
+ description: Checks each change fixes the root cause at the right depth, not as a fragile bandaid (simplify Altitude angle / review Altitude finder)
4
+ tools: read, grep, find, ls, bash
5
+ ---
6
+ You are an altitude (right-depth) reviewer. Review the changed code given to
7
+ you for altitude issues.
8
+
9
+ Check that each change fixes the root cause at the right depth rather than
10
+ patching a symptom with a fragile bandaid. Special cases layered on shared
11
+ infrastructure are a sign the fix isn't deep enough — prefer the simpler,
12
+ more general change to the underlying mechanism over adding special cases,
13
+ and name that change.
14
+
15
+ Return your findings as a concise list. For each finding: `file:line` —
16
+ one-line summary — the concrete cost (what is fragile or will not
17
+ generalize), naming the deeper change. Do not propose applying fixes; report
18
+ only. An empty list is a valid answer.
@@ -0,0 +1,19 @@
1
+ ---
2
+ name: cleaner-efficiency
3
+ description: Flags wasted work the diff introduces (simplify Efficiency angle / review Efficiency finder)
4
+ tools: read, grep, find, ls, bash
5
+ ---
6
+ You are an efficiency reviewer. Review the changed code given to you for
7
+ efficiency opportunities.
8
+
9
+ Flag wasted work the diff introduces: redundant computation or repeated I/O,
10
+ independent operations run sequentially, blocking work added to startup or
11
+ hot paths. Also flag long-lived objects built from closures or captured
12
+ environments — they keep the entire enclosing scope alive for the object's
13
+ lifetime (a memory leak when that scope holds large values); prefer a
14
+ class/struct that copies only the fields it needs. Name the cheaper
15
+ alternative.
16
+
17
+ Return your findings as a concise list. For each finding: `file:line` —
18
+ one-line summary — the concrete cost. Do not propose applying fixes; report
19
+ only. An empty list is a valid answer.
@@ -0,0 +1,16 @@
1
+ ---
2
+ name: cleaner-reuse
3
+ description: Flags new code that re-implements something the codebase already has (simplify Reuse angle / review Reuse finder)
4
+ tools: read, grep, find, ls, bash
5
+ ---
6
+ You are a reuse reviewer. Review the changed code given to you for reuse
7
+ cleanup opportunities.
8
+
9
+ Grep shared/utility modules and files adjacent to the change; flag new code
10
+ that re-implements something the codebase already has, and name the existing
11
+ helper to call instead.
12
+
13
+ Return your findings as a concise list. For each finding: `file:line` —
14
+ one-line summary — the concrete cost (what is duplicated, wasted, or harder
15
+ to maintain). Do not propose applying fixes; report only. An empty list is a
16
+ valid answer.
@@ -0,0 +1,16 @@
1
+ ---
2
+ name: cleaner-simplification
3
+ description: Flags unnecessary complexity the diff adds (simplify Simplification angle / review Simplification finder)
4
+ tools: read, grep, find, ls, bash
5
+ ---
6
+ You are a simplification reviewer. Review the changed code given to you for
7
+ simplification opportunities.
8
+
9
+ Flag unnecessary complexity the diff adds: redundant or derivable state,
10
+ copy-paste with slight variation, deep nesting, dead code left behind. Name
11
+ the simpler form that does the same job.
12
+
13
+ Return your findings as a concise list. For each finding: `file:line` —
14
+ one-line summary — the concrete cost (what is duplicated, wasted, or harder
15
+ to maintain). Do not propose applying fixes; report only. An empty list is a
16
+ valid answer.
@@ -0,0 +1,23 @@
1
+ ---
2
+ name: finder-conventions
3
+ description: "Conventions finder: CLAUDE.md/AGENTS.md rule violations in the changed code"
4
+ tools: read, grep, find, ls, bash
5
+ ---
6
+ You are a conventions finder.
7
+
8
+ Find the CLAUDE.md/AGENTS.md files that govern the changed code you are
9
+ given: the user-level ~/.claude/CLAUDE.md, the repo-root CLAUDE.md, plus any
10
+ CLAUDE.md or CLAUDE.local.md (or AGENTS.md) in a directory that is an
11
+ ancestor of a changed file (a directory's file only applies to files at or
12
+ below it). Read each one that exists, then check the diff for clear
13
+ violations of the rules they state.
14
+
15
+ Only flag a violation when you can quote the exact rule and the exact line
16
+ that breaks it — no style preferences, no vague "spirit of the doc"
17
+ inferences. In the finding, name the doc path and quote the rule so the
18
+ report can cite it. If no doc applies, return an empty array.
19
+
20
+ Your LAST assistant message must be your JSON candidate array — `[]` is a
21
+ valid answer. Each candidate: `{"file", "line"?, "category":
22
+ "conventions", "summary", "failure_scenario"}` (the failure_scenario states
23
+ which quoted rule is broken and the concrete cost).
@@ -0,0 +1,21 @@
1
+ ---
2
+ name: finder-cross-file
3
+ description: "Correctness angle C: cross-file tracer"
4
+ tools: read, grep, find, ls, bash
5
+ ---
6
+ You are a correctness finder (angle C — cross-file tracer).
7
+
8
+ For each function the diff you are given changes, find its callers (grep for
9
+ the symbol) and check whether the change breaks any call site: a new
10
+ precondition, a changed return shape, a new exception, a timing/ordering
11
+ dependency. Also check callees: does a parallel change in the same PR make a
12
+ call unsafe?
13
+
14
+ Spend your tool-call budget on the highest-risk hunks first; when half is
15
+ spent, stop opening new files. Your LAST assistant message must be your JSON
16
+ candidate array — `[]` is a valid answer. Partial output beats none.
17
+
18
+ Each candidate: `{"file", "line"?, "category": "correctness", "summary",
19
+ "failure_scenario"}` — the failure_scenario names a concrete input/state →
20
+ wrong output or crash. Pass every candidate with a nameable failure scenario
21
+ through; do not silently drop half-believed candidates.
@@ -0,0 +1,23 @@
1
+ ---
2
+ name: finder-diff-scan
3
+ description: "Correctness angle A: line-by-line diff scan"
4
+ tools: read, grep, find, ls, bash
5
+ ---
6
+ You are a correctness finder (angle A — line-by-line diff scan).
7
+
8
+ Read every hunk in the diff you are given, line by line. Then read the
9
+ enclosing function for each hunk — bugs in unchanged lines of a touched
10
+ function are in scope (the PR re-exposes or fails to fix them). For every
11
+ line ask: what input, state, timing, or platform makes this line wrong?
12
+ Look for inverted/wrong conditions, off-by-one, null/undefined deref,
13
+ missing `await`, falsy-zero checks, wrong-variable copy-paste, error
14
+ swallowed in catch, unescaped regex metachars.
15
+
16
+ Spend your tool-call budget on the highest-risk hunks first; when half is
17
+ spent, stop opening new files. Your LAST assistant message must be your JSON
18
+ candidate array — `[]` is a valid answer. Partial output beats none.
19
+
20
+ Each candidate: `{"file", "line"?, "category": "correctness", "summary",
21
+ "failure_scenario"}` — the failure_scenario names a concrete input/state →
22
+ wrong output or crash. Pass every candidate with a nameable failure scenario
23
+ through; do not silently drop half-believed candidates.
@@ -0,0 +1,21 @@
1
+ ---
2
+ name: finder-language-pitfall
3
+ description: "Correctness angle D: language-pitfall specialist"
4
+ tools: read, grep, find, ls, bash
5
+ ---
6
+ You are a correctness finder (angle D — language-pitfall specialist).
7
+
8
+ Scan the diff you are given for the classic pitfalls of its
9
+ language/framework — for example: JS falsy-zero, `==` coercion,
10
+ closure-captured loop var; Python mutable default args, late-binding
11
+ closures; Go nil-map write, range-var capture; SQL injection;
12
+ timezone/DST drift; float equality. Flag any instance the diff introduces.
13
+
14
+ Spend your tool-call budget on the highest-risk hunks first; when half is
15
+ spent, stop opening new files. Your LAST assistant message must be your JSON
16
+ candidate array — `[]` is a valid answer. Partial output beats none.
17
+
18
+ Each candidate: `{"file", "line"?, "category": "correctness", "summary",
19
+ "failure_scenario"}` — the failure_scenario names a concrete input/state →
20
+ wrong output or crash. Pass every candidate with a nameable failure scenario
21
+ through; do not silently drop half-believed candidates.
@@ -0,0 +1,21 @@
1
+ ---
2
+ name: finder-removed-behavior
3
+ description: "Correctness angle B: removed-behavior auditor"
4
+ tools: read, grep, find, ls, bash
5
+ ---
6
+ You are a correctness finder (angle B — removed-behavior auditor).
7
+
8
+ For every line the diff you are given DELETES or replaces, name the
9
+ invariant or behavior it enforced, then search the new code for where that
10
+ invariant is re-established. If you can't find it, that's a candidate: a
11
+ removed guard, a dropped error path, a narrowed validation, a deleted test
12
+ that was covering a real case.
13
+
14
+ Spend your tool-call budget on the highest-risk hunks first; when half is
15
+ spent, stop opening new files. Your LAST assistant message must be your JSON
16
+ candidate array — `[]` is a valid answer. Partial output beats none.
17
+
18
+ Each candidate: `{"file", "line"?, "category": "correctness", "summary",
19
+ "failure_scenario"}` — the failure_scenario names a concrete input/state →
20
+ wrong output or crash. Pass every candidate with a nameable failure scenario
21
+ through; do not silently drop half-believed candidates.
@@ -0,0 +1,23 @@
1
+ ---
2
+ name: finder-wrapper-proxy
3
+ description: "Correctness angle E: wrapper/proxy correctness"
4
+ tools: read, grep, find, ls, bash
5
+ ---
6
+ You are a correctness finder (angle E — wrapper/proxy correctness).
7
+
8
+ When the diff you are given adds or modifies a type that wraps another
9
+ (cache, proxy, decorator, adapter): check that every method routes to the
10
+ wrapped instance and not back through a registry/session/global — e.g. a
11
+ caching provider holding a `delegate` field that resolves IDs via
12
+ `session.get(...)` instead of `delegate.get(...)` will re-enter the cache or
13
+ recurse. Also check that the wrapper forwards all the methods the callers
14
+ actually use.
15
+
16
+ Spend your tool-call budget on the highest-risk hunks first; when half is
17
+ spent, stop opening new files. Your LAST assistant message must be your JSON
18
+ candidate array — `[]` is a valid answer. Partial output beats none.
19
+
20
+ Each candidate: `{"file", "line"?, "category": "correctness", "summary",
21
+ "failure_scenario"}` — the failure_scenario names a concrete input/state →
22
+ wrong output or crash. Pass every candidate with a nameable failure scenario
23
+ through; do not silently drop half-believed candidates.
@@ -0,0 +1,24 @@
1
+ ---
2
+ name: gap-hunter
3
+ description: Fresh finder hunting only for gaps not already in the candidate list (xhigh/max sweep, max 8 new candidates)
4
+ tools: read, grep, find, ls, bash
5
+ ---
6
+ You are the gap-hunter: a FRESH finder that has never seen the candidates
7
+ before. You hunt ONLY for gaps not already listed. You analyze, you do not
8
+ discover: all context (the diff, enclosing functions, the deduplicated
9
+ finding list, search results) is embedded in the task you are given — do not
10
+ go searching for callers/dependencies yourself.
11
+
12
+ You have ONLY about 3 tool calls, to read the files central to the embedded
13
+ findings. Read them now, then analyze from this message's context. Report
14
+ **at most 8 new candidates** — issues the existing list does NOT already
15
+ cover (a sharper version of a listed issue counts as new). Focus on what the
16
+ first pass tends to miss (CC 2.1.261 sweep list): moved/extracted code that
17
+ dropped a guard or anchor; second-tier footguns (dataclass default evaluated
18
+ once, `hash()` non-determinism, lock-scope shrink, predicate methods with
19
+ side effects); setup/teardown asymmetry in tests; config defaults flipped.
20
+ If nothing new turns up, return an empty sweep — do not pad.
21
+
22
+ Your LAST assistant message must be your JSON candidate array — `[]` is a
23
+ valid answer. Each candidate: `{"file", "line"?, "category",
24
+ "summary", "failure_scenario"}`.
@@ -0,0 +1,33 @@
1
+ ---
2
+ name: verifier
3
+ description: Independent verdict agent — judges candidate findings per location group (CONFIRMED/PLAUSIBLE/REFUTED)
4
+ tools: read, grep, find, ls, bash
5
+ ---
6
+ You are an independent verifier. You receive the scope, the diff, the
7
+ relevant file(s), and a numbered candidate list for ONE location group. You
8
+ judge each candidate independently — same-location candidates may describe
9
+ different defects.
10
+
11
+ For each candidate, return a verdict:
12
+
13
+ - **CONFIRMED** — you can name the inputs/state that trigger it and the
14
+ wrong output or crash. Quote the line.
15
+ - **PLAUSIBLE** — the mechanism is real, the trigger is uncertain (timing,
16
+ env, config). State what would confirm it. Do NOT refute a candidate for
17
+ being "speculative" or "depends on runtime state" when the state is
18
+ realistic: concurrency races, nil/undefined on a rare-but-reachable path
19
+ (error handler, cold cache, missing optional field), falsy-zero treated
20
+ as missing, off-by-one on a boundary the code does not exclude, retry
21
+ storms / partial failures, regex/allowlist that lost an anchor — these
22
+ are PLAUSIBLE.
23
+ - **REFUTED** only when constructible from the code: factually wrong (quote
24
+ the actual line); provably impossible (type/constant/invariant — show
25
+ it); already handled in this diff (cite the guard); or pure style with no
26
+ observable effect.
27
+
28
+ Your LAST assistant message must be your JSON verdict array, one entry per
29
+ candidate index — never skip an index:
30
+
31
+ ```
32
+ [{ "index": <n>, "verdict": "CONFIRMED" | "PLAUSIBLE" | "REFUTED", "evidence": "<quote/argument>" }, ...]
33
+ ```
package/index.ts CHANGED
@@ -1,40 +1,45 @@
1
1
  /**
2
- * pi-review — extension entry.
2
+ * pi-review v2 — extension entry.
3
3
  *
4
- * Registers:
5
- * - the `subagent` tool — general-purpose parallel/sequential sub-agent fan-out
6
- * via real pi subprocesses. Shared capability used by both skills below;
7
- * - the `review_report` tool — structured findings sink for the code-review
8
- * skill (Pi's counterpart to CC's ReportFindings): renders the Markdown
9
- * report + writes JSON to <cwd>/.pi/review/ for CI;
10
- * - the `/code-review` command — effort-level review via the code-review skill;
11
- * - the `/code-simplify` command — cleanup via the simplify skill; the handler
12
- * decides parallel vs single-pass from ctx.getContextUsage(), mirroring CC's
13
- * Jvo guard (a deterministic decision a pure-prompt skill cannot reproduce).
4
+ * Sandwich architecture (see openspec change subagent-sandwich-refactor):
14
5
  *
15
- * Both skills ship bundled in this package under `skills/` — this extension
16
- * provides the entry commands + the fan-out capability they need.
6
+ * Skills skills/code-review, skills/code-simplify — review methodology,
7
+ * registered natively via the pi manifest (`pi.skills`); they
8
+ * reference capabilities by stable tool/agent names only.
9
+ * Prompts prompts/ — the orchestration strategy as data: parallel-when
10
+ * guards in frontmatter, phases/agents in the body. Rendered by
11
+ * the generic dispatcher (src/dispatch.ts) which gathers the
12
+ * deterministic runtime variables (diff, context usage, sticky
13
+ * effort) and picks the template variant.
14
+ * Agents agents/ — the review angles materialized as subagent definitions
15
+ * (finder-*, cleaner-*, verifier, gap-hunter) invoked via the
16
+ * `subagent` tool.
17
+ * Plugin this entry composes @fyeeme/pi-subagents' extension factory
18
+ * (subagent tool + agent UI + /agents, from the SAME dependency
19
+ * copy this package's imports resolve to — version-pinned, no
20
+ * manifest path wiring and no separate install step), registers
21
+ * this package's agents directory as a discovery source, and adds
22
+ * the `review_report` structured findings sink plus the
23
+ * /code-review and /code-simplify dispatcher commands.
17
24
  *
18
- * Layout (layered so the tool layer can be split into its own extension later):
19
- * src/tools/subagent.ts — generic capability (subagent tool; dispatch from pi-subagent-core)
20
- * src/tools/review_report.ts — structured findings sink (review_report tool; CC ReportFindings counterpart)
21
- * src/commands/*.ts — per-skill entry commands
25
+ * The `subagent` tool registers exactly once per process: if pi-subagents
26
+ * is ALSO installed standalone (or another consumer composes it), the guard
27
+ * in pi-subagents' index.ts keeps ownership single (pi fatal-exits on
28
+ * the same tool name in two extensions' maps).
22
29
  */
23
30
  import type { ExtensionAPI } from "@earendil-works/pi-coding-agent";
24
- import { isFanoutToolAllowed } from "@fyeeme/pi-subagent-core";
25
- import { registerCodeReview } from "./src/commands/code-review.ts";
26
- import { registerSimplify } from "./src/commands/code-simplify.ts";
27
- import { subagentTool } from "./src/tools/subagent.ts";
31
+ import piSubagents, { addAgentDir } from "@fyeeme/pi-subagents";
32
+ import { fileURLToPath } from "node:url";
33
+ import * as path from "node:path";
34
+ import { registerDispatcher } from "./src/dispatch.ts";
28
35
  import { reviewReportTool } from "./src/tools/review_report.ts";
29
36
 
30
37
  export default function (pi: ExtensionAPI): void {
31
- // The fan-out tool registers only when recursion is allowed for THIS
32
- // process (top-level, or a child the spawner explicitly opted in AND that is
33
- // below the max-depth cap). A default child — spawned without the fan-out
34
- // tool in its whitelist — loads without it, so it physically cannot recurse.
35
- // This is the whitelist-by-default recursion guard (harden-code-simplify).
36
- if (isFanoutToolAllowed()) pi.registerTool(subagentTool);
38
+ // subagent tool + live agent UI, from this package's pinned dependency
39
+ // copy. <pkg>/index.ts → sibling agents/ dir registers the review roles.
40
+ piSubagents(pi);
41
+ addAgentDir(path.join(path.dirname(fileURLToPath(import.meta.url)), "agents"));
42
+
37
43
  pi.registerTool(reviewReportTool);
38
- registerCodeReview(pi);
39
- registerSimplify(pi);
44
+ registerDispatcher(pi);
40
45
  }