@whamp/pi-pstack 0.9.0 → 0.9.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -22,25 +22,25 @@ pi install ~/projects/pi-extensions/packages/pi-pstack
22
22
  This package is ported from the Cursor pstack plugin by Lauren Tan
23
23
  (`LICENSE`).
24
24
 
25
- Requires [`pi-subagents`](https://www.npmjs.com/package/pi-subagents) for the `poteto-agent`, `comment-sicko`, and workflow fan-outs (`how`, `why`, `arena`, `swarm`, `interrogate`, `reflect`).
25
+ Requires [`pi-subagents`](https://www.npmjs.com/package/pi-subagents) for the `poteto-agent`, `comment-sicko`, and workflow fan-outs (`how`, `why`, `arena`, `swarm`, `interrogate`, `reflect`, `code-review`).
26
26
 
27
27
  ## Get started
28
28
 
29
29
  1. Run `/setup-pstack` once to pick which models each role uses (optional; every role inherits the parent session model otherwise).
30
30
  2. Use `/poteto-mode` for sticky Poteto Mode. It stays on until `/poteto-mode off`. `/skill:poteto-mode` also enables it.
31
- 3. Run `/pstack off` to hide even the four Discoverable skills (`how`, `why`, `unslop`, `typescript-best-practices`) from the Skill catalog.
31
+ 3. Run `/pstack off` to hide the Pi-only `code-review` coordinator from the Skill catalog.
32
32
  Off persists in `~/.pi/agent/pstack/models.json`.
33
33
  `/skill:<name>` keeps working.
34
- `/pstack on` restores those four, not all 47.
34
+ `/pstack on` restores `code-review`, not all 48.
35
35
 
36
36
  That is it.
37
37
  The other skills are Hidden; the mode skill uses them as needed.
38
38
 
39
39
  ## What you get
40
40
 
41
- - **47 skills**, including:
41
+ - **48 skills**, including:
42
42
  - `poteto-mode`: the main entry point. Reads your request, matches one of 23 playbooks (bug fix, perf, feature, refactoring, investigation, shipping, orchestrate, autopilot, and more), copies its steps in verbatim, and routes to the other skills as steps fire. Orchestrate refills one shared worker-and-verifier window as each child settles.
43
- - Workflow skills: `how`, `why`, `recall`, `blast-radius`, `architect`, `arena`, `swarm`, `interrogate`, `reflect`, `teach`, `tdd`, `no-comments`, `unslop`, `deslop`, `bro`, `figure-it-out`, `show-me-your-work`, `create-verification-skill`, `maintain-verification-skill`, `automate-me`, `technical-writing`, `typescript-best-practices`.
43
+ - Workflow skills: `code-review`, `how`, `why`, `recall`, `blast-radius`, `architect`, `arena`, `swarm`, `interrogate`, `reflect`, `teach`, `tdd`, `no-comments`, `unslop`, `deslop`, `bro`, `figure-it-out`, `show-me-your-work`, `create-verification-skill`, `maintain-verification-skill`, `automate-me`, `technical-writing`, `typescript-best-practices`.
44
44
  - 23 principle skills (`principle-laziness-protocol`, `principle-model-the-domain`, `principle-prove-it-works`, ...), one rule each, indexed inline by `poteto-mode`.
45
45
  - **`ask_user_question`**: one structured preference question with 2-6 listed options. The user can pick those or type a different answer.
46
46
  - **2 subagents** (loaded by pi-subagents):
@@ -54,15 +54,36 @@ Per-role model choices live in `~/.pi/agent/pstack/models.json`. Run `/setup-pst
54
54
 
55
55
  ## Differences from the Cursor plugin
56
56
 
57
- - Hidden skills set `disable-model-invocation: true`, so they stay out of the Skill catalog.
58
- `/skill:name` still loads the Skill body.
59
- The four Discoverable skills are `how`, `why`, `unslop`, and `typescript-best-practices`.
57
+ - The importer preserves upstream invocation settings. Change upstream behavior only for a necessary Pi adaptation or an explicitly approved exception.
58
+ `how`, `why`, `unslop`, and `typescript-best-practices` retain upstream's `disable-model-invocation: true`.
59
+ Their `agents/openai.yaml` files also set `policy.allow_implicit_invocation: false` for Codex.
60
+ `/skill:name` still loads the Skill body, and Poteto Mode keeps its explicit skill routes.
61
+ The Pi-only `code-review` coordinator remains model-visible.
60
62
  - Slash commands are `/skill:<name>` instead of `/name`.
61
63
  - Subagent delegation uses pi-subagents. Launch one child with `subagent({ action: "execute", input: { agent, task } })`. Set `input.async: true` for background work. Run parallel or dependent children in one `workflowScript` with stable keys. This package does not ship the `subagent` tool.
62
64
  - Session transcripts live under `~/.pi/agent/sessions/--<slug>--/` instead of Cursor `agent-transcripts/`. The active file is `$PI_SESSION_FILE`. `<slug>` is the absolute cwd with the leading slash dropped and each `/` turned into `-`. Stay inside that workspace directory. Do not glob sibling slugs.
63
65
  - The benny automation pack is not ported; it depends on Cursor automations. Model roles live in `~/.pi/agent/pstack/models.json`, written by `/setup-pstack` and read on demand through `model-routing`.
64
66
  - `make-bot-ui` is not ported. It is Cursor Grok Bot / routine webhook UI.
65
67
 
68
+ ## Code review coordinator
69
+
70
+ `code-review` uses Audit for ordinary review requests. It uses Challenge only for explicit adversarial or design interrogation. Requests for both run both routes. PR-status requests stay with the existing Babysit playbook. Audit findings follow `skills/code-review/references/code-review-audit.md`; Challenge reuses the existing `interrogate` skill. The coordinator adds no model role or review registry.
71
+
72
+ The coordinator uses these routes both inside and outside sticky Poteto Mode when Pstack skills are enabled:
73
+
74
+ | Request | Route |
75
+ | --- | --- |
76
+ | "Review this PR" or "review since X" | Audit. Resolve the PR's base and head, or ask for a missing base. |
77
+ | "Review against the issue" | Audit. Pin the base and the issue requirements. |
78
+ | "Challenge the design" | Challenge on the pinned design contents. |
79
+ | "Open a PR" | Opening a PR playbook. Keep its existing Challenge requirement and do not add Audit. |
80
+ | "Check on PR X" | Babysit, not a code review. |
81
+ | "Audit and challenge this change" | Audit + challenge, with independent results. |
82
+
83
+ Audit launches one Standards and one Spec child for each caller-selected model. Explicitly absent specs skip Spec. Challenge keeps one child per configured Interrogate reviewer. With both requested, the counts add; one route does not erase the other.
84
+
85
+ An older `code-review` skill may also appear from a global installation under `~/.agents/skills/code-review`. This change does not alter that installation. While both copies exist, load this package's `skills/code-review/SKILL.md` by its absolute path. A name collision can resolve `/skill:code-review` to the older copy. After the Pstack source is integrated and published, replace the global entry through its installer and verify each affected consumer. Do not remove a shared installation before those consumers have the replacement. Do not run both copies as separate reviewers.
86
+
66
87
  ## Related port
67
88
 
68
89
  [backnotprop/pstack](https://github.com/backnotprop/pstack) is Lauren Tan's standalone mirror of the same Cursor plugin (`npx skills add backnotprop/pstack`). Its `main` branch keeps Cursor wording and adds a [Harness](https://github.com/backnotprop/pstack/blob/main/skills/poteto-mode/SKILL.md#harness) table so one skill body can run in Claude Code, Codex, Pi, and others. This package is the Pi-native port: it rewrites those seams (`/skill:`, `models.json`, pi-subagents) instead of asking the agent to translate. The Pi session path in that Harness table is what this package now writes into skills. Do not install the mirror into Pi if you want this extension.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@whamp/pi-pstack",
3
- "version": "0.9.0",
3
+ "version": "0.9.1",
4
4
  "description": "pstack for Pi: rigorous agent workflows you can parallelize with confidence - poteto-mode playbooks, engineering principles, multi-model review panels, and subagents.",
5
5
  "type": "module",
6
6
  "license": "MIT",
@@ -46,7 +46,7 @@ Default: proceed directly to implementation with the synthesized design. No huma
46
46
 
47
47
  Opt in to a checkpoint when the invoker explicitly asks: "/skill:architect with checkpoint," "stop and show me before implementing," or similar. Then surface the synthesized design and pause for sign-off.
48
48
 
49
- The synthesis can ship as its own commit either way, as the "scaffold first" mode of the **foundational-thinking** principle skill. Planned and scoped breakage during fill-in is fine, per the **outcome-oriented-execution** principle skill. For adversarial pressure on the design before implementing, run the **interrogate** skill on the synthesized sketch.
49
+ The synthesis can ship as its own commit either way, as the "scaffold first" mode of the **foundational-thinking** principle skill. Planned and scoped breakage during fill-in is fine, per the **outcome-oriented-execution** principle skill. For adversarial pressure on the design before implementing, read `../code-review/SKILL.md` relative to this skill directory. Use Challenge on the synthesized sketch.
50
50
 
51
51
  If the human pushes back on the shape (in a checkpoint or after the fact), treat that as Phase A evidence. Re-ground and re-run Phase B before writing more code.
52
52
 
@@ -38,7 +38,7 @@ If a candidate fails to produce output, pass the completed N-1 results to the ju
38
38
 
39
39
  ## Phase C: Cross-judge
40
40
 
41
- After the Phase B workflow completes, choose one model from the `arena judge pool` in `~/.pi/agent/pstack/models.json` when present. Otherwise use inherit-parent. Prefer a different model family from the parent's. Launch the judge with `subagent({ action: "execute", input: { agent: "worker", task, model, async: true } })`. Its task says to inspect only, read the rubric and candidates by path label, score each criterion, and recommend a base with rationale. Read the completed candidate artifacts while the judge runs. The judge never runs while candidates are writing.
41
+ After the Phase B workflow completes, choose one model from the `arena judge pool` in `~/.pi/agent/pstack/models.json` when present. Otherwise use inherit-parent. Prefer a different model family from the parent's. Launch the judge with `subagent({ action: "execute", input: { agent: "reviewer", task, model, async: true } })`. Its task says to inspect only, read the rubric and candidates by path label, score each criterion, and recommend a base with rationale. Read the completed candidate artifacts while the judge runs. The judge never runs while candidates are writing.
42
42
 
43
43
  ## Phase D: Pick a base
44
44
 
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 Matt Pocock
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
@@ -0,0 +1,42 @@
1
+ ---
2
+ name: code-review
3
+ description: "Review a PR, diff, branch, or changes since a fixed point. Audit is the default. Use Challenge only for an explicit adversarial or design review request. Route PR-status requests to Babysit."
4
+ ---
5
+
6
+ # Code review
7
+
8
+ Use this coordinator for code reviews. Keep Audit and Challenge separate. The [Audit procedure](references/code-review-audit.md) defines evidence preparation, reviewer coverage, and the report.
9
+
10
+ ## Route the request
11
+
12
+ - Use Audit for an ordinary or bare request to review a PR, diff, branch, or changes since a point.
13
+ - Use Challenge only when the user explicitly asks for adversarial review or design interrogation. The feature, bug-fix, architect, and PR-opening playbooks can also require Challenge.
14
+ - Run Audit and Challenge when the user explicitly asks for both. Keep their reviewer contexts independent. Audit children do not receive Challenge results. Challenge children do not receive Audit findings.
15
+ - Route PR-status requests, including "check on PR X," to the [Babysit playbook](../poteto-mode/playbooks/babysit.md), not Audit.
16
+ - Opening a PR alone does not start Babysit. The PR-opening playbook still requires Challenge.
17
+
18
+ Ask for a missing Audit base instead of guessing. A named PR supplies immutable base and head commits. Challenge can use pinned design contents without a Git base.
19
+
20
+ ## Freeze evidence before launching reviewers
21
+
22
+ The parent owns preparation and synthesis. Before launching any reviewers, freeze the review intent, exact scope, pinned artifact, relevant spec and standards, completed tool evidence, and limitations. Give each child that same evidence packet. Children inspect only. They do not run Git or shell commands, edit files, write files, or use MCP or extension tools.
23
+
24
+ For Audit, if the spec source is missing and the user has not said that no spec exists, ask for the source before launching Audit reviewers. Challenge needs a clear intent and pinned artifact, not an originating Audit spec. If the user explicitly says no spec is available, record Spec as `SKIPPED (no spec available)`. Do not call it a pass.
25
+
26
+ ## Run Audit
27
+
28
+ Use the shared ordered model lineup from the caller's `model-routing`, spending, and family policy. Each selected model gets one fresh Standards reviewer and one fresh Spec reviewer. These are separate children. Do not combine the axes or choose separate model lineups for them.
29
+
30
+ Use the read-only `reviewer` agent with fresh context through `pi-subagents`. Never substitute a write-capable worker for a reviewer. Follow the [Audit procedure](references/code-review-audit.md). It blocks missing required sources, model families, axes, or reviewer results. Never turn a failed, missing, or partial check into a pass.
31
+
32
+ ## Run Challenge
33
+
34
+ Load the existing [Interrogate skill](../interrogate/SKILL.md). Skip its steps 1 and 2 for a coordinator-delegated Challenge. Start at step 3 with the frozen artifact, scope, and intent verbatim. Do not run Git, rediscover the artifact, or derive intent from the implementation. Keep its reviewer rubric and synthesis. Do not create another Challenge rubric or reviewer roster. Direct `/skill:interrogate` use remains compatible.
35
+
36
+ Do not trigger Challenge based on risk. A completed Audit never satisfies a required Challenge. Reuse prior coverage only when the pinned artifact, intent, rubric, required axes and families, and parent disposition all match.
37
+
38
+ ## Report the review
39
+
40
+ Record the requested and actual model selectors. State any substitution, covered families, missing coverage, child failures, findings, unresolved questions, and the parent's disposition for each finding. Keep Standards, Spec, and Challenge results separate. Do not create a separate review registry or apply fixes automatically.
41
+
42
+ This coordinator adapts Matt Pocock's MIT-licensed `code-review` skill. See [the retained MIT notice](LICENSE).
@@ -0,0 +1,91 @@
1
+ # Audit
2
+
3
+ Audit answers two different questions. Standards asks whether the change follows the repository's documented standards. Spec asks whether the change meets its originating request. Keep the answers separate.
4
+
5
+ ## Prepare the review packet
6
+
7
+ ### Pin the artifact
8
+
9
+ 1. Resolve the supplied fixed base and pinned head to commit SHAs. For a named PR, use its base and head SHAs.
10
+ 2. Preserve three-dot scope. Run `git diff <fixed-base>...<pinned-head>` and resolve and record the merge base it uses. Record commits with `git log <fixed-base>..<pinned-head>`. Do not replace this with direct endpoint subtraction using two dots.
11
+ 3. Ask for the base if the request omits it. Do not guess `main`, `HEAD`, or a merge base.
12
+ 4. For a file or design review, pin the file contents by commit or record an immutable content digest. A design review without Git history can still use Challenge.
13
+ 5. Record the exact resolved commands, commit list, changed-file list, and any scope exclusions.
14
+ 6. Stop before launching children if a ref cannot resolve or the complete requested artifact is empty. An empty committed diff can still have requested WIP to review. Do not reject that WIP merely because no committed changes exist.
15
+
16
+ Review a named PR's committed change by default. Add local work in progress only when requested. When WIP is requested, include staged and unstaged tracked changes and every requested untracked file. A regular `git diff` omits staged and untracked content. Keep the committed change and local WIP evidence distinct so the packet does not duplicate or hide changes.
17
+
18
+ ### Find the spec
19
+
20
+ Trace the request to its source. Check issue references in commits, user-provided paths, matching documents in `docs/`, `specs/`, or `.scratch/`, and the relevant issue tracker. Record the source URL or path, stable requirement reference, and the exact text that governs the change.
21
+
22
+ If no source is found and the user has not said that no spec exists, ask for the spec before reviewers start. If the user explicitly says there is no spec, mark Spec `SKIPPED (no spec available)`. Do not report Spec as passed or infer requirements from the implementation.
23
+
24
+ ### Find the standards
25
+
26
+ Inspect the repository's maintained standards and applicable language guidance. Record each source's URL, path, revision, and relevant rule IDs. The implementation is not its own standard.
27
+
28
+ If a required standard source is missing, block the Standards axis and ask for the source. For a language with no adopted language standard, state `no language standard adopted for <language>; language-rule audit skipped.` Do not invent language rules. Keep Standards open for documented cross-language rules and the labelled Fowler heuristics below. Spec remains a separate axis.
29
+
30
+ For every standards finding, cite a stable rule ID, a source URL, the changed file and line, why the code breaks the rule, and a concrete fix. Use the source's published rule ID. If it has none, label the rule with a stable path and heading anchor. Do not invent an upstream rule number.
31
+
32
+ ### Complete tool checks
33
+
34
+ Before reviewer children start, the parent runs every applicable check required by the standards and records its exact command and result. Inspect tool configuration. Confirm the check covers every changed and new file. Record completed checks with their diagnostics. Record unavailable, failed, incomplete, and inapplicable checks separately. A tool error or incomplete coverage is a gap, not a pass.
35
+
36
+ The parent gathers all spec and standards sources, exact tool output, and limitations into one frozen packet. Reviewer children receive that packet and read-only file pointers. They do not run Git or shell commands, edit or write files, or use MCP or extension tools.
37
+
38
+ ## Select and run Audit reviewers
39
+
40
+ Choose one ordered Audit model lineup for both axes. Follow the caller's `model-routing`, spending, and family policy, plus any explicit user selection. Do not select separate model lists for Standards and Spec. If no private routing policy exists, inherit the parent model and report the actual resolved model and family.
41
+
42
+ Record requested selectors separately from actual selectors. Record every substitution. Do not silently claim a substituted model covers a required family. Missing required families, failed launches, timeouts, and missing outputs remain unresolved coverage.
43
+
44
+ Discover available agents with `subagent({ action: "list", input: { capabilities: true } })`. Launch the Audit children together in one `workflowScript` using `runs.all`. Each item uses `agent: "reviewer"`, `context: "fresh"`, a stable key, the prepared packet, and its resolved model selector. Count all children in the run's spawn budget. The parent synthesizes the returned reports.
45
+
46
+ For each selected model, launch two fresh reviewer children:
47
+
48
+ 1. A Standards child that receives the exact frozen diff, relevant standards, completed diagnostics, and limitations. It reviews Standards only.
49
+ 2. A Spec child that receives the exact frozen diff, originating spec, and relevant context. It reviews Spec only.
50
+
51
+ The children use the same ordered model lineup, but their contexts stay separate. Do not combine both axes in one child. Do not copy one model's answer to fill another model's missing coverage. If the user explicitly says that no spec exists, launch no Spec children and report the axis as skipped.
52
+
53
+ Each child inspects only the prepared artifact. Ask it to report supported findings, file and line, source evidence, and any gaps. A child's zero findings do not turn another missing or failed result into a pass.
54
+
55
+ ## Review Standards findings
56
+
57
+ Start with documented rules. The repository's rule overrides general guidance. Skip a judgment heuristic when the repository endorses the pattern or a tool already enforces it. Label each heuristic finding as a possible smell, never a hard rule violation.
58
+
59
+ Include the twelve Fowler heuristics below in every Standards packet. They are judgment prompts, not hard rules. They remain subordinate to documented standards and tool results.
60
+
61
+ 1. **Mysterious Name.** A function, variable, or type has a name that does not reveal what it does or holds. Rename it. If no honest name fits, clarify the design.
62
+ 2. **Duplicated Code.** The same logic shape appears more than once in the change. Extract the shared shape and call it from both sites.
63
+ 3. **Feature Envy.** A method reads another object's data more than its own. Move the method to the data it uses.
64
+ 4. **Data Clumps.** The same fields or parameters travel together repeatedly. Group them in one type and pass that type.
65
+ 5. **Primitive Obsession.** A primitive or string stands in for a domain concept that needs its own type. Give the concept a small type.
66
+ 6. **Repeated Switches.** The same conditional on the same type recurs in the change. Replace it with polymorphism or one shared map.
67
+ 7. **Shotgun Surgery.** One logical change forces scattered edits across many files. Gather the related behavior in one module.
68
+ 8. **Divergent Change.** One file or module changes for several unrelated reasons. Split it so each module changes for one reason.
69
+ 9. **Speculative Generality.** An abstraction, parameter, or hook serves a need the spec does not have. Delete it and inline the code until a real need appears.
70
+ 10. **Message Chains.** A long chain such as `a.b().c().d()` exposes navigation the caller should not depend on. Hide the walk behind a method on the first object.
71
+ 11. **Middle Man.** A class or function mostly delegates to another. Remove it and call the target directly.
72
+ 12. **Refused Bequest.** A subclass or implementer ignores or overrides most of what it inherits. Replace inheritance with composition.
73
+
74
+ ## Review Spec findings
75
+
76
+ Compare the implementation with each relevant requirement in the frozen spec. Cite the exact requirement and the changed file and line. Report missing or partial requirements, unrequested behavior, and implementations that appear wrong. Do not use code comments or the implementation itself as the spec.
77
+
78
+ ## Synthesize and report coverage
79
+
80
+ The parent checks citations against the pinned packet and keeps disagreement visible. Report a finding total and the worst issue for Standards and Spec separately. Do not rank an issue in one axis against an issue in the other. Report each axis as findings, no findings, blocked, or skipped. `No findings` means the required review completed and found none. `Blocked` means required evidence or execution is missing. `Skipped` means the caller explicitly made the axis inapplicable.
81
+
82
+ Keep a review record in the response, not a new registry. Include:
83
+
84
+ - The intent, exact scope, pinned base and head or content digest, and spec provenance.
85
+ - The standards sources, diagnostics, and completed, blocked, skipped, or incomplete checks.
86
+ - The requested and actual model selectors, substitutions, required and covered families, and child failures.
87
+ - The finding total and worst issue within each axis, with no cross-axis ranking.
88
+ - Separate Standards and Spec results, findings, unresolved coverage, and the parent's disposition for each finding.
89
+ - Spec `SKIPPED (no spec available)` when the user explicitly declares that no spec exists.
90
+
91
+ Do not merge Standards and Spec into one verdict. Do not mark incomplete coverage as passed. Do not apply fixes during the review.
@@ -1,6 +1,7 @@
1
1
  ---
2
2
  name: how
3
3
  description: "Use for \"how does X work\", code walkthroughs before changing something, and placement / ownership / layering questions (\"where should this live\", \"which package owns this\", \"is this the right layer\"). Explains subsystem architecture, runtime flow, onboarding mental models. Use why for motivation."
4
+ disable-model-invocation: true
4
5
  ---
5
6
 
6
7
  # How
@@ -20,10 +21,10 @@ When in doubt, take the simple path.
20
21
 
21
22
  ## Step 2a. Explore (complex questions only)
22
23
 
23
- Decompose the question into 2 to 4 exploration angles, each a distinct slice of the subsystem. Launch the explorers and dependent explainer with one `subagent({ action: "execute", input: { async: true, maxSubagentSpawnsPerRun: N + 1, workflowScript } })` call. In `workflowScript`, await `runs.all([{ key: "explore-<angle>", agent: "worker", task, model }])`, then return `runs.run("explain", { agent: "worker", task, model })` with the explorer outputs.
24
+ Decompose the question into 2 to 4 exploration angles, each a distinct slice of the subsystem. Launch the explorers and dependent explainer with one `subagent({ action: "execute", input: { async: true, maxSubagentSpawnsPerRun: N + 1, workflowScript } })` call. In `workflowScript`, await `runs.all([{ key: "explore-<angle>", agent: "reviewer", task, model }])`, then return `runs.run("explain", { agent: "reviewer", task, model })` with the explorer outputs.
24
25
 
25
26
  Each explorer uses:
26
- - agent: "worker"
27
+ - agent: "reviewer"
27
28
  - `model`: `how explorers` (default inherit-parent)
28
29
  - `task`: the prompt in `references/explorer-prompt.md` with its angle filled in and an instruction to inspect only
29
30
 
@@ -31,8 +32,8 @@ Then go to Step 3.
31
32
 
32
33
  ## Step 2b. Direct Explain (simple questions)
33
34
 
34
- Launch one standalone child with `subagent({ action: "execute", input: { agent: "worker", task, model, async: false } })` using:
35
- - agent: "worker"
35
+ Launch one standalone child with `subagent({ action: "execute", input: { agent: "reviewer", task, model, async: false } })` using:
36
+ - agent: "reviewer"
36
37
  - `model`: `how explainer` (default inherit-parent)
37
38
  - `task`: `references/explainer-prompt.md` without the explorer-findings section and with an instruction to inspect only
38
39
 
@@ -41,7 +42,7 @@ Go to Step 4.
41
42
  ## Step 3. Synthesize (complex questions only)
42
43
 
43
44
  The same workflow launches `explain` after every explorer settles using:
44
- - agent: "worker"
45
+ - agent: "reviewer"
45
46
  - `model`: `how synthesizer` (default inherit-parent)
46
47
  - `task`: `references/explainer-prompt.md` with every explorer result filled in and an instruction to inspect only
47
48
 
@@ -0,0 +1,2 @@
1
+ policy:
2
+ allow_implicit_invocation: false
@@ -33,7 +33,7 @@ Write one clear paragraph. If you're unsure about the intent, ask the user befor
33
33
 
34
34
  ## Step 3, Spawn Reviewers
35
35
 
36
- Launch all reviewers with one `subagent({ action: "execute", input: { async: true, maxSubagentSpawnsPerRun: N, workflowScript } })` call. In `workflowScript`, use `return await runs.all([{ key: "reviewer-a", agent: "worker", task, model }])` with one stable-keyed item per reviewer. Use the `interrogate reviewers` list from `~/.pi/agent/pstack/models.json` when present, one reviewer per entry, extending or shrinking the Reviewer A/B/C labels below to the configured entry count. Otherwise use the table defaults.
36
+ Launch all reviewers with one `subagent({ action: "execute", input: { async: true, maxSubagentSpawnsPerRun: N, workflowScript } })` call. In `workflowScript`, use `return await runs.all([{ key: "reviewer-a", agent: "reviewer", task, model }])` with one stable-keyed item per reviewer. Use the `interrogate reviewers` list from `~/.pi/agent/pstack/models.json` when present, one reviewer per entry, extending or shrinking the Reviewer A/B/C labels below to the configured entry count. Otherwise use the table defaults.
37
37
 
38
38
  | Subagent | Default model |
39
39
  |----------|---------------|
@@ -42,7 +42,7 @@ Launch all reviewers with one `subagent({ action: "execute", input: { async: tru
42
42
  | Reviewer C | inherit-parent |
43
43
 
44
44
  For each reviewer:
45
- - agent: "worker"
45
+ - agent: "reviewer"
46
46
  - `model`: the configured `interrogate reviewers` entry, or the table default with no configured line
47
47
  - `task`: instruct the reviewer to inspect only and not modify files
48
48
 
@@ -11,6 +11,16 @@ disable-model-invocation: true
11
11
  `/skill:poteto-mode` also enables it.
12
12
  Before selecting a delegated model, use `model-routing` to read the configured roles on demand.
13
13
 
14
+ ## Code review routing
15
+
16
+ Read `../code-review/SKILL.md` relative to this skill directory for every route below. Use that file, not the globally registered skill name.
17
+
18
+ - Ordinary requests to review a PR, diff, branch, or changes since a point use `code-review` Audit. A bare `review` also uses Audit.
19
+ - Ask for a missing Audit base. Do not guess. A named PR supplies immutable base and head commits.
20
+ - Use Challenge only for an explicit adversarial or design-interrogation request. Run Audit and Challenge when the user asks for both.
21
+ - Challenge can review pinned design contents without a Git base.
22
+ - PR-status requests such as `check on PR X` use the Babysit playbook.
23
+
14
24
  ## Non-negotiables
15
25
 
16
26
  The Principles section below grounds every trigger. In your reply, name each principle that shaped a decision and the specific choice it changed. Cite only principles whose leaf SKILL.md you read this session.
@@ -22,7 +32,7 @@ Remaining triggers:
22
32
  - Any code → name the data shape first, and choose its organizing structure per **principle-model-the-domain**.
23
33
  - Code crossing a function boundary → the **architect** skill, parallel design exploration before implementing.
24
34
  - Parallel fan-out → the **swarm** skill for coverage matrices, races, gauntlets, and exploration partitions. Use **arena** for design or code bakeoffs with base selection and grafting.
25
- - Contested design → the **interrogate** skill (multi-model adversarial) before shipping.
35
+ - Contested design → read `../code-review/SKILL.md` relative to this skill directory and use Challenge before shipping.
26
36
  - Nontrivial multi-step → write the throughput checkpoint (Feature step 3).
27
37
  - Any prose surface → the **unslop** skill. Your reply is a prose surface. Write it per **Writing the reply**. Agent-facing prose also follows `playbooks/authoring-a-skill.md` and `/skill:unslop`.
28
38
  - Docs, RFCs, readmes, PR descriptions, or commit messages → the **technical-writing** skill (`/skill:technical-writing`).
@@ -93,9 +103,9 @@ Read the leaf skill in full for any principle you apply. Each entry names when i
93
103
 
94
104
  **Defaults for every child launch.** Set `input.async: true` for background work. Pass file pointers instead of inlining context. Select an explicit model per role when `/setup-pstack` configures one. Multiple children or dependent stages use one `subagent({ action: "execute", input: { workflowScript, ... } })` call. Inside the script, use `await runs.all([{ key: "stable-key", ... }])` for fan-out and `return runs.run("stable-key", { ... })` for a direct or final child. Count every later synthesis or review child in `input.maxSubagentSpawnsPerRun` when the workflow sets that limit.
95
105
 
96
- A child does not inherit ambient MCP or extension tools. Keep MCP lookup in the parent for `why`, `reflect`, and `interrogate` unless the selected custom agent lists the tool and loads its provider through `extensions` or `subagentOnlyExtensions`. Do not invent per-call tools.
106
+ A child does not inherit ambient MCP or extension tools. Keep MCP lookup in the parent for `why`, `reflect`, `interrogate`, and `code-review` unless the selected custom agent lists the tool and loads its provider through `extensions` or `subagentOnlyExtensions`. Do not invent per-call tools.
97
107
 
98
- Defaults inherit-parent. Ordinary judgment uses `judgment`. User-facing writing uses `prose`. Escalated difficult work uses `hardest tasks`. Implementation playbooks use `feature implementation`, `refactoring implementation`, `bug-fix`, `perf-issue`, and `hillclimb`. Role lines choose only the model. They never grant tools, authority, or isolation. Code delegates tier by difficulty. The hardest changes (cross-cutting design, gnarly concurrency, subtle algorithms) go to `hardest tasks` when configured, else the parent model, whether the task needs judgment on vague intent or is a precisely specified sequence of steps to execute to the letter. Trivial mechanical edits go to your fast code model. Configured roles resolved through `model-routing` override these defaults and the model choices in the routed skills (`how`, `why`, `arena`, `swarm`, `architect`, `interrogate`, `reflect`). A role with no line keeps its default, and a role line of `inherit-parent` or `auto` runs that role on the parent chat model. Omit `model` in that case.
108
+ Defaults inherit-parent. Ordinary judgment uses `judgment`. User-facing writing uses `prose`. Escalated difficult work uses `hardest tasks`. Implementation playbooks use `feature implementation`, `refactoring implementation`, `bug-fix`, `perf-issue`, and `hillclimb`. Role lines choose only the model. They never grant tools, authority, or isolation. Code delegates tier by difficulty. The hardest changes (cross-cutting design, gnarly concurrency, subtle algorithms) go to `hardest tasks` when configured, else the parent model, whether the task needs judgment on vague intent or is a precisely specified sequence of steps to execute to the letter. Trivial mechanical edits go to your fast code model. Configured roles resolved through `model-routing` override these defaults and the model choices in the routed skills (`how`, `why`, `arena`, `swarm`, `architect`, `interrogate`, `reflect`). A role with no line keeps its default, and a role line of `inherit-parent` or `auto` runs that role on the parent chat model. Omit `model` in that case. The `code-review` coordinator uses the caller's `model-routing`, spending, and family policy for Audit and adds no model role. Challenge mode reuses the existing `interrogate reviewers` role.
99
109
 
100
110
  You own every subagent's work. Review the diff and write your own summary, don't pass through what it said. Interrupt-chained resumes silently drop directives, so fire a fresh subagent with consolidated scope rather than trusting a "done" summary. A second opinion is the same prompt against a different model. Agreement is high-signal.
101
111
 
@@ -5,7 +5,7 @@
5
5
  Be scientific. Every shipped line traces to runtime evidence. Belt-and-suspenders that "might help" is a hypothesis, not a fix. It does not ship. When evidence refutes a hypothesis, revert what it motivated. The smallest change the evidence justifies ships, nothing more.
6
6
 
7
7
  1. Reproduce it yourself on the matching surface via the control skill (Non-negotiables), even when a debug or instrumentation protocol says to ask the user to reproduce. Ask the user only with a stated, specific reason the control surface cannot reach the target, and only after driving it as far as it goes. If it won't reproduce directly, synthesize the trigger, tighten conditions, or instrument until it fires.
8
- 2. Binary-search the cause. Form the candidate hypotheses, then rule them out until one survives. Seed them with `how` over the affected subsystem and the **why** skill for regression history. Each pass, take the split that cuts the most remaining problem space, get runtime evidence, eliminate. When program state is unclear, add instrumentation or logging and read it as the code runs. Don't guess. Drive a long or stubborn hunt with a recurring wake. Confirm the surviving *mechanism* with runtime evidence before the step-3 architect/interrogate fan-out.
8
+ 2. Binary-search the cause. Form the candidate hypotheses, then rule them out until one survives. Seed them with `how` over the affected subsystem and the **why** skill for regression history. Each pass, take the split that cuts the most remaining problem space, get runtime evidence, eliminate. When program state is unclear, add instrumentation or logging and read it as the code runs. Don't guess. Drive a long or stubborn hunt with a recurring wake. Confirm the surviving *mechanism* with runtime evidence before the step-3 `architect` and Pstack Challenge fan-out. For Challenge, read `../../code-review/SKILL.md` relative to this playbook's directory.
9
9
  3. Plan the fix. If it crosses a function boundary, `architect` first. Delegate implementation to a subagent using the `bug-fix` role (default inherit-parent) with a specific scope.
10
10
  4. Verify on the same surface. The original repro now passes. "Inconclusive" or wrong-surface is not a pass. Flag it. Unit tests show branch behavior, not bug absence.
11
11
  5. Stage the commits so the failing repro lands before the fix in git history. See the **tdd** skill for the failing-test-first cadence when the bug has a cheap local test path. Skip it when the test would be expensive, integration-heavy, or unclear.
@@ -13,7 +13,7 @@
13
13
  5. Verify on the matching surface. "Inconclusive" or wrong-surface is not a pass. Flag it.
14
14
  6. Rebase into small, ordered commits. Stack follow-ups.
15
15
  Use the **sequence-verifiable-units** principle skill, building, verifying, and committing each small unit before the next.
16
- 7. If the design is contested, `interrogate` before shipping.
16
+ 7. If the design is contested, read `../../code-review/SKILL.md` relative to this playbook's directory. Use Challenge before shipping.
17
17
  8. Run **Opening a PR**.
18
18
 
19
19
  Code-coupled work (one feature, one migration) goes to a single owner with the checkpoint inline. That owner fans out internally after the blocking phase. Parent-level fan-out is for slices that produce independent artifacts (audits, cross-subsystem investigations, competing experiments). Rewrite the checkpoint at phase boundaries. Spawn a fresh owner rather than chaining interrupts.
@@ -8,6 +8,7 @@ Core discipline: one change, one measurement, keep or revert. Never stack untest
8
8
  2. Build the measurement harness, prove its sensitivity, then freeze it (the **build-the-lever** principle skill). Run contrasting realistic workloads and confirm the target case reproduces the symptom while easier cases separate as expected. If the harness cannot distinguish them, revise the workload or metric. Once frozen, one repeatable command emits the metric, sampled enough to clear the noise (median of N, not a single run). Record the baseline metric and a green run of the regression gate (the tests that must keep passing) before any change.
9
9
  3. Open the decision log via the **show-me-your-work** skill. A `decision.tsv`, one row per attempt: id, hypothesis, change, before, after, delta, tests, verdict (kept or reverted), note. Read it before each attempt. Keep it out of the tree (gitignored).
10
10
  4. Ground each hypothesis in the architecture model from step 1, so it names a specific mechanism ("defer X off the boot path because it blocks first paint"), not "try memoizing something".
11
+ For performance within one executable program, read the installed `~/.agents/skills/perform-like-jeff-and-sanjay/SKILL.md` before the first attempt and use its gated causal hypothesis method for each attempt. Reuse existing measurements. Distributed-systems performance and ML-hardware tuning stay with domain-specific methods.
11
12
  5. Loop, one hypothesis per iteration:
12
13
  - Hand the change to a subagent using the `hillclimb` role (default inherit-parent) with a tight scope. Supervise and review the diff rather than typing it (the **guard-the-context-window** principle skill). When several independent hypotheses are live, fan them to parallel subagents, each in its own worktree (the **separate-before-serializing-shared-state** principle skill).
13
14
  - Measure before and after with the frozen harness, and run the regression gate.
@@ -150,7 +150,7 @@ Each live lane runs at the PR head. Drive through the project's verification ski
150
150
 
151
151
  ## Appendix D. Links and reading list
152
152
 
153
- <Docs to read before editing. Which PRs get `pstack/skills/how/SKILL.md` and `pstack/skills/interrogate/SKILL.md`. The trail per `pstack/skills/show-me-your-work/SKILL.md`.>
153
+ <Docs to read before editing. Which PRs get `pstack/skills/how/SKILL.md` and `../../code-review/SKILL.md` in Challenge mode. Resolve the coordinator path relative to this playbook's directory. The trail per `pstack/skills/show-me-your-work/SKILL.md`.>
154
154
  ````
155
155
 
156
156
  **Reply:** the plan path, the PR ids with their dependencies and the review-gated set, what the prototypes proved and what stays unproven, and the check script's output.
@@ -30,4 +30,4 @@ After these sections, attach videos or screenshots when they prove a claim. Do n
30
30
 
31
31
  **Babysit.** Opening a PR does not start a babysit. Post the URL and keep building. Finish the phase or stack first. Run a separate babysit pass only when the user asks for one after the whole stack exists. A babysit for each new PR stalls the build and spends checks on commits that later waves restart. Push back when feedback drifts from intent.
32
32
 
33
- A subagent that opens a PR runs `interrogate`, `/skill:deslop`, and `/skill:no-comments`, and posts the URL. Then it returns to the parent without babysitting, unless it is an Autopilot-full or Autopilot-stack owner. That owner's brief assigns the babysit loop and is the ask `playbooks/babysit.md` waits for. The owner starts the loop after its code-ready report and reports merge-ready or STACK-READY as its playbook says. The rules here and in `playbooks/babysit.md` that hold babysitting until a whole stack is built do not apply to that owner.
33
+ A subagent that opens a PR first reads `../../code-review/SKILL.md` relative to this playbook's directory. It runs that coordinator in Challenge mode, `/skill:deslop`, and `/skill:no-comments`, and posts the URL. Then it returns to the parent without babysitting, unless it is an Autopilot-full or Autopilot-stack owner. That owner's brief assigns the babysit loop and is the ask `playbooks/babysit.md` waits for. The owner starts the loop after its code-ready report and reports merge-ready or STACK-READY as its playbook says. The rules here and in `playbooks/babysit.md` that hold babysitting until a whole stack is built do not apply to that owner.
@@ -13,6 +13,7 @@
13
13
  - **Redundancy.** The wait hangs on one slow instance or attempt. Duplicate the work (replicas, hedged requests, speculative execution) and take the fastest result. The trace has to show the wait dominates and the system has headroom.
14
14
  - **Lazy evaluation.** Cost lands on results that are never used or not needed yet (eager init on the boot path, rendering offscreen items). Defer the work until first use.
15
15
  - **Scheduling.** The work must happen, but not during the interactive moment. Move it to where nobody is waiting: idle callbacks, a background warmup after boot, precompute before the user arrives, cleanup after the frame commits. The win is perceived latency, so measure the interactive path, not total work done.
16
+ For a bottleneck within one executable program, read the installed `~/.agents/skills/perform-like-jeff-and-sanjay/SKILL.md` and use its Diagnose route before selecting the fix. Reuse existing measurements. Distributed-systems performance and ML-hardware tuning stay with domain-specific methods.
16
17
  3. Plan the fix from the trace. If it crosses a function boundary, `architect` first. Delegate implementation to a subagent using the `perf-issue` role (default inherit-parent). Review the diff. Capture a post-fix trace.
17
18
  Apply the **sequence-verifiable-units** principle skill, verifying each attempt before trying the next.
18
19
  4. Parse and compare the artifacts (JSON to sqlite, diff). "Inconclusive" or wrong-surface is not a pass. Flag it.
@@ -38,11 +38,11 @@ Each child names a role in `~/.pi/agent/pstack/models.json`. Use that role's sel
38
38
  | Tooling | `reflect tooling reviewer` (default inherit-parent) | `references/tooling-reviewer.md` |
39
39
  | Divergent | `reflect divergent reviewer` (default inherit-parent) | `references/divergent-reviewer.md` |
40
40
 
41
- Each reviewer item uses `agent: "worker"`, its configured model, and a task that says to inspect only. Pass each template verbatim, substituting the transcript path or the bounded digest where marked. Reviewers return findings through their workflow results.
41
+ Each reviewer item uses `agent: "reviewer"`, its configured model, and a task that says to inspect only. Pass each template verbatim, substituting the transcript path or the bounded digest where marked. Reviewers return findings through their workflow results.
42
42
 
43
43
  ### 3. Synthesize
44
44
 
45
- The workflow's `synthesize-reviews` child uses `agent: "worker"`. It runs using `reflect synthesizer` (default inherit-parent). Use `references/synthesizer.md` verbatim, with each reviewer's full output inlined where marked. It returns a structured Accepted / Rejected / Backlog list. After the workflow completes, the parent spot-verifies citations with its own MCP and extension tools.
45
+ The workflow's `synthesize-reviews` child uses `agent: "reviewer"`. It runs using `reflect synthesizer` (default inherit-parent). Use `references/synthesizer.md` verbatim, with each reviewer's full output inlined where marked. It returns a structured Accepted / Rejected / Backlog list. After the workflow completes, the parent spot-verifies citations with its own MCP and extension tools.
46
46
 
47
47
  ### 4. Structural enforcement check
48
48
 
@@ -1,6 +1,7 @@
1
1
  ---
2
2
  name: typescript-best-practices
3
3
  description: TypeScript best practices. Use when reading or editing any .ts or .tsx file.
4
+ disable-model-invocation: true
4
5
  ---
5
6
 
6
7
  # TypeScript best practices
@@ -0,0 +1,2 @@
1
+ policy:
2
+ allow_implicit_invocation: false
@@ -1,6 +1,7 @@
1
1
  ---
2
2
  name: unslop
3
3
  description: Cut AI tells from any writing. Must always apply.
4
+ disable-model-invocation: true
4
5
  ---
5
6
 
6
7
  # Unslop
@@ -0,0 +1,2 @@
1
+ policy:
2
+ allow_implicit_invocation: false
@@ -1,6 +1,7 @@
1
1
  ---
2
2
  name: why
3
3
  description: "Use for 'why does X work this way', 'why we picked Y', design rationale, regressions, postmortems, or data-backed thresholds. Discovers available MCPs and queries each evidence category (source control, issue tracker, long-form docs, real-time chat, infrastructure observability, error tracking, product analytics warehouse) in parallel, then returns a cited read on decisions and tradeoffs. Use how for runtime behavior."
4
+ disable-model-invocation: true
4
5
  ---
5
6
 
6
7
  # Why
@@ -74,10 +75,10 @@ Source control is always available through git and `gh`. For the other six, clas
74
75
 
75
76
  Aim for a complete **coverage map**, not a minimal one. Document the null, don't skip the search. The parent queries each available MCP and builds one bounded evidence packet per category before launching children. A child does not inherit ambient MCP or extension tools. Use a custom agent for a child-side lookup only when that agent explicitly lists the tool and loads its provider through `extensions` or `subagentOnlyExtensions`.
76
77
 
77
- Launch all matching investigators and the dependent synthesizer with one `subagent({ action: "execute", input: { async: true, maxSubagentSpawnsPerRun: N + 1, workflowScript } })` call. In `workflowScript`, await the investigators with `runs.all([{ key: "investigate-<category>", agent: "worker", task, model }])`, then return `runs.run("synthesize-why", { agent: "worker", task, model })` with their outputs. `N` is the number of evidence categories launched. Don't ask one agent to cover multiple categories.
78
+ Launch all matching investigators and the dependent synthesizer with one `subagent({ action: "execute", input: { async: true, maxSubagentSpawnsPerRun: N + 1, workflowScript } })` call. In `workflowScript`, await the investigators with `runs.all([{ key: "investigate-<category>", agent: "reviewer", task, model }])`, then return `runs.run("synthesize-why", { agent: "reviewer", task, model })` with their outputs. `N` is the number of evidence categories launched. Don't ask one agent to cover multiple categories.
78
79
 
79
80
  Each investigator uses:
80
- - agent: "worker"
81
+ - agent: "reviewer"
81
82
  - `model`: `why investigators` (default inherit-parent)
82
83
  - `task`: instruct the investigator to inspect only
83
84
 
@@ -121,7 +122,7 @@ If your scope assessment suggests a single-commit trivial target where the PR de
121
122
  ## Step 4. Synthesize
122
123
 
123
124
  The same workflow launches `synthesize-why` after every investigator settles. It uses:
124
- - agent: "worker"
125
+ - agent: "reviewer"
125
126
  - `model`: `why synthesizer` (default inherit-parent)
126
127
 
127
128
  The synthesizer gets:
@@ -0,0 +1,2 @@
1
+ policy:
2
+ allow_implicit_invocation: false