@whamp/pi-pstack 0.7.0 → 0.9.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (55) hide show
  1. package/README.md +38 -11
  2. package/extensions/pstack/index.ts +7 -4
  3. package/extensions/pstack/pstack-role-prompt.ts +3 -29
  4. package/package.json +1 -1
  5. package/skills/architect/SKILL.md +1 -1
  6. package/skills/arena/SKILL.md +2 -2
  7. package/skills/automate-me/SKILL.md +2 -2
  8. package/skills/blast-radius/SKILL.md +3 -3
  9. package/skills/code-review/LICENSE +21 -0
  10. package/skills/code-review/SKILL.md +42 -0
  11. package/skills/code-review/references/code-review-audit.md +91 -0
  12. package/skills/figure-it-out/SKILL.md +3 -3
  13. package/skills/how/SKILL.md +8 -5
  14. package/skills/how/agents/openai.yaml +2 -0
  15. package/skills/how/references/explorer-prompt.md +1 -1
  16. package/skills/interrogate/SKILL.md +2 -3
  17. package/skills/interrogate/references/code-quality-review.md +1 -1
  18. package/skills/interrogate/references/reviewer-prompt.md +1 -3
  19. package/skills/interrogate/references/rubric.md +1 -1
  20. package/skills/poteto-mode/SKILL.md +17 -7
  21. package/skills/poteto-mode/playbooks/autopilot-full.md +6 -6
  22. package/skills/poteto-mode/playbooks/autopilot-stack.md +7 -7
  23. package/skills/poteto-mode/playbooks/babysit.md +1 -1
  24. package/skills/poteto-mode/playbooks/bug-fix.md +3 -5
  25. package/skills/poteto-mode/playbooks/eval.md +1 -1
  26. package/skills/poteto-mode/playbooks/feature.md +3 -3
  27. package/skills/poteto-mode/playbooks/hillclimb.md +1 -0
  28. package/skills/poteto-mode/playbooks/multi-phase-plan.md +8 -8
  29. package/skills/poteto-mode/playbooks/opening-a-pr.md +1 -1
  30. package/skills/poteto-mode/playbooks/pause-safely.md +1 -1
  31. package/skills/poteto-mode/playbooks/perf-issue.md +1 -0
  32. package/skills/poteto-mode/playbooks/refactoring.md +2 -2
  33. package/skills/poteto-mode/playbooks/session-pickup.md +1 -1
  34. package/skills/poteto-mode/playbooks/shipping.md +2 -2
  35. package/skills/principle-guard-the-context-window/SKILL.md +0 -1
  36. package/skills/principle-never-block-on-the-human/SKILL.md +0 -2
  37. package/skills/principle-outcome-oriented-execution/SKILL.md +0 -1
  38. package/skills/principle-prove-it-works/SKILL.md +0 -11
  39. package/skills/principle-sequence-verifiable-units/SKILL.md +0 -5
  40. package/skills/recall/SKILL.md +1 -1
  41. package/skills/reflect/SKILL.md +7 -7
  42. package/skills/reflect/references/divergent-reviewer.md +1 -1
  43. package/skills/reflect/references/judgment-reviewer.md +1 -1
  44. package/skills/reflect/references/tooling-reviewer.md +1 -1
  45. package/skills/show-me-your-work/SKILL.md +7 -7
  46. package/skills/show-me-your-work/scripts/log.sh +4 -2
  47. package/skills/swarm/SKILL.md +4 -4
  48. package/skills/tdd/SKILL.md +1 -3
  49. package/skills/technical-writing/SKILL.md +0 -13
  50. package/skills/typescript-best-practices/SKILL.md +1 -0
  51. package/skills/typescript-best-practices/agents/openai.yaml +2 -0
  52. package/skills/unslop/SKILL.md +1 -1
  53. package/skills/unslop/agents/openai.yaml +2 -0
  54. package/skills/why/SKILL.md +6 -3
  55. package/skills/why/agents/openai.yaml +2 -0
package/README.md CHANGED
@@ -22,25 +22,25 @@ pi install ~/projects/pi-extensions/packages/pi-pstack
22
22
  This package is ported from the Cursor pstack plugin by Lauren Tan
23
23
  (`LICENSE`).
24
24
 
25
- Requires [`pi-subagents`](https://www.npmjs.com/package/pi-subagents) for the `poteto-agent`, `comment-sicko`, and workflow fan-outs (`how`, `why`, `arena`, `swarm`, `interrogate`, `reflect`).
25
+ Requires [`pi-subagents`](https://www.npmjs.com/package/pi-subagents) for the `poteto-agent`, `comment-sicko`, and workflow fan-outs (`how`, `why`, `arena`, `swarm`, `interrogate`, `reflect`, `code-review`).
26
26
 
27
27
  ## Get started
28
28
 
29
29
  1. Run `/setup-pstack` once to pick which models each role uses (optional; every role inherits the parent session model otherwise).
30
30
  2. Use `/poteto-mode` for sticky Poteto Mode. It stays on until `/poteto-mode off`. `/skill:poteto-mode` also enables it.
31
- 3. Run `/pstack off` to hide even the four Discoverable skills (`how`, `why`, `unslop`, `typescript-best-practices`) from the Skill catalog.
31
+ 3. Run `/pstack off` to hide the Pi-only `code-review` coordinator from the Skill catalog.
32
32
  Off persists in `~/.pi/agent/pstack/models.json`.
33
33
  `/skill:<name>` keeps working.
34
- `/pstack on` restores those four, not all 47.
34
+ `/pstack on` restores `code-review`, not all 48.
35
35
 
36
36
  That is it.
37
37
  The other skills are Hidden; the mode skill uses them as needed.
38
38
 
39
39
  ## What you get
40
40
 
41
- - **47 skills**, including:
41
+ - **48 skills**, including:
42
42
  - `poteto-mode`: the main entry point. Reads your request, matches one of 23 playbooks (bug fix, perf, feature, refactoring, investigation, shipping, orchestrate, autopilot, and more), copies its steps in verbatim, and routes to the other skills as steps fire. Orchestrate refills one shared worker-and-verifier window as each child settles.
43
- - Workflow skills: `how`, `why`, `recall`, `blast-radius`, `architect`, `arena`, `swarm`, `interrogate`, `reflect`, `teach`, `tdd`, `no-comments`, `unslop`, `deslop`, `bro`, `figure-it-out`, `show-me-your-work`, `create-verification-skill`, `maintain-verification-skill`, `automate-me`, `technical-writing`, `typescript-best-practices`.
43
+ - Workflow skills: `code-review`, `how`, `why`, `recall`, `blast-radius`, `architect`, `arena`, `swarm`, `interrogate`, `reflect`, `teach`, `tdd`, `no-comments`, `unslop`, `deslop`, `bro`, `figure-it-out`, `show-me-your-work`, `create-verification-skill`, `maintain-verification-skill`, `automate-me`, `technical-writing`, `typescript-best-practices`.
44
44
  - 23 principle skills (`principle-laziness-protocol`, `principle-model-the-domain`, `principle-prove-it-works`, ...), one rule each, indexed inline by `poteto-mode`.
45
45
  - **`ask_user_question`**: one structured preference question with 2-6 listed options. The user can pick those or type a different answer.
46
46
  - **2 subagents** (loaded by pi-subagents):
@@ -50,19 +50,46 @@ The other skills are Hidden; the mode skill uses them as needed.
50
50
 
51
51
  ## Model roles
52
52
 
53
- Per-role model choices live in `~/.pi/agent/pstack/models.json`. Run `/setup-pstack` to write it. The 22 role names and cardinalities are in `skills/setup-pstack/references/MODEL-ROLES.md`. The extension injects the role table only when a role has a real model slug. Default inherit-all injects nothing. `inherit-parent` or `auto` runs on the parent session model.
53
+ Per-role model choices live in `~/.pi/agent/pstack/models.json`. Run `/setup-pstack` to write it. The 22 role names and cardinalities are in `skills/setup-pstack/references/MODEL-ROLES.md`. The extension does not inject model roles into the system prompt. Before delegation, use `model-routing` to read the role configuration and select a model under the caller's routing policy. Install that skill separately from `Whamp/skills`; it is not bundled here. Neither Poteto Mode nor the `/pstack` skills toggle controls role lookup. `inherit-parent` or `auto` runs on the parent session model.
54
54
 
55
55
  ## Differences from the Cursor plugin
56
56
 
57
- - Hidden skills set `disable-model-invocation: true`, so they stay out of the Skill catalog.
58
- `/skill:name` still loads the Skill body.
59
- The four Discoverable skills are `how`, `why`, `unslop`, and `typescript-best-practices`.
57
+ - The importer preserves upstream invocation settings. Change upstream behavior only for a necessary Pi adaptation or an explicitly approved exception.
58
+ `how`, `why`, `unslop`, and `typescript-best-practices` retain upstream's `disable-model-invocation: true`.
59
+ Their `agents/openai.yaml` files also set `policy.allow_implicit_invocation: false` for Codex.
60
+ `/skill:name` still loads the Skill body, and Poteto Mode keeps its explicit skill routes.
61
+ The Pi-only `code-review` coordinator remains model-visible.
60
62
  - Slash commands are `/skill:<name>` instead of `/name`.
61
63
  - Subagent delegation uses pi-subagents. Launch one child with `subagent({ action: "execute", input: { agent, task } })`. Set `input.async: true` for background work. Run parallel or dependent children in one `workflowScript` with stable keys. This package does not ship the `subagent` tool.
62
- - Session transcripts live under `~/.pi/agent/sessions/` instead of `~/.cursor/projects/`. The active file is `$PI_SESSION_FILE`. Files are grouped by cwd slug (`--<cwd>--`, absolute cwd with `/` replaced by `-`).
63
- - The benny automation pack is not ported; it depends on Cursor automations. Model roles live in `~/.pi/agent/pstack/models.json`, written by `/setup-pstack` and injected only when a role has a real model slug.
64
+ - Session transcripts live under `~/.pi/agent/sessions/--<slug>--/` instead of Cursor `agent-transcripts/`. The active file is `$PI_SESSION_FILE`. `<slug>` is the absolute cwd with the leading slash dropped and each `/` turned into `-`. Stay inside that workspace directory. Do not glob sibling slugs.
65
+ - The benny automation pack is not ported; it depends on Cursor automations. Model roles live in `~/.pi/agent/pstack/models.json`, written by `/setup-pstack` and read on demand through `model-routing`.
64
66
  - `make-bot-ui` is not ported. It is Cursor Grok Bot / routine webhook UI.
65
67
 
68
+ ## Code review coordinator
69
+
70
+ `code-review` uses Audit for ordinary review requests. It uses Challenge only for explicit adversarial or design interrogation. Requests for both run both routes. PR-status requests stay with the existing Babysit playbook. Audit findings follow `skills/code-review/references/code-review-audit.md`; Challenge reuses the existing `interrogate` skill. The coordinator adds no model role or review registry.
71
+
72
+ The coordinator uses these routes both inside and outside sticky Poteto Mode when Pstack skills are enabled:
73
+
74
+ | Request | Route |
75
+ | --- | --- |
76
+ | "Review this PR" or "review since X" | Audit. Resolve the PR's base and head, or ask for a missing base. |
77
+ | "Review against the issue" | Audit. Pin the base and the issue requirements. |
78
+ | "Challenge the design" | Challenge on the pinned design contents. |
79
+ | "Open a PR" | Opening a PR playbook. Keep its existing Challenge requirement and do not add Audit. |
80
+ | "Check on PR X" | Babysit, not a code review. |
81
+ | "Audit and challenge this change" | Audit + challenge, with independent results. |
82
+
83
+ Audit launches one Standards and one Spec child for each caller-selected model. Explicitly absent specs skip Spec. Challenge keeps one child per configured Interrogate reviewer. With both requested, the counts add; one route does not erase the other.
84
+
85
+ An older `code-review` skill may also appear from a global installation under `~/.agents/skills/code-review`. This change does not alter that installation. While both copies exist, load this package's `skills/code-review/SKILL.md` by its absolute path. A name collision can resolve `/skill:code-review` to the older copy. After the Pstack source is integrated and published, replace the global entry through its installer and verify each affected consumer. Do not remove a shared installation before those consumers have the replacement. Do not run both copies as separate reviewers.
86
+
87
+ ## Related port
88
+
89
+ [backnotprop/pstack](https://github.com/backnotprop/pstack) is Lauren Tan's standalone mirror of the same Cursor plugin (`npx skills add backnotprop/pstack`). Its `main` branch keeps Cursor wording and adds a [Harness](https://github.com/backnotprop/pstack/blob/main/skills/poteto-mode/SKILL.md#harness) table so one skill body can run in Claude Code, Codex, Pi, and others. This package is the Pi-native port: it rewrites those seams (`/skill:`, `models.json`, pi-subagents) instead of asking the agent to translate. The Pi session path in that Harness table is what this package now writes into skills. Do not install the mirror into Pi if you want this extension.
90
+
91
+ See [MIRROR.md](https://github.com/backnotprop/pstack/blob/main/MIRROR.md) for the mirror's two-branch sync. This package uses `scripts/reground-from-cursor.mjs` instead.
92
+
66
93
  ## License
67
94
 
68
95
  MIT
@@ -31,11 +31,10 @@ import {
31
31
  import { registerAskUserQuestion } from "./ask-user-question.ts";
32
32
  import { stripSkillsByLocationPrefix } from "./skill-strip.ts";
33
33
 
34
- export { systemPromptInjection };
35
-
36
34
  const SKILLS_DIR = join(dirname(fileURLToPath(import.meta.url)), "..", "..", "skills");
37
35
 
38
36
  const POTETO_SKILL = "/skill:poteto-mode";
37
+ const INLINE_POTETO_MODE_RE = /(?<![a-z0-9._%+-])\$poteto-mode(?![a-z0-9_-]|\.[a-z0-9])/;
39
38
 
40
39
  type ModeEntry = {
41
40
  type?: string;
@@ -188,7 +187,11 @@ export default function pstackExtension(pi: ExtensionAPI): void {
188
187
  });
189
188
 
190
189
  pi.on("input", async (event, ctx) => {
191
- if (/^\/skill:poteto-mode(?:\s|$)/.test(event.text)) {
190
+ if (
191
+ event.source !== "extension" &&
192
+ (/^\/skill:poteto-mode(?:\s|$)/.test(event.text) ||
193
+ INLINE_POTETO_MODE_RE.test(event.text))
194
+ ) {
192
195
  persistMode(true, ctx);
193
196
  }
194
197
  return { action: "continue" as const };
@@ -199,7 +202,7 @@ export default function pstackExtension(pi: ExtensionAPI): void {
199
202
  const base = loaded.config.skillsEnabled
200
203
  ? event.systemPrompt
201
204
  : stripSkillsByLocationPrefix(event.systemPrompt, SKILLS_DIR).prompt;
202
- const extra = systemPromptInjection(loaded.config, potetoMode);
205
+ const extra = systemPromptInjection(potetoMode);
203
206
  return {
204
207
  systemPrompt: extra ? `${base}\n\n${extra}` : base,
205
208
  };
@@ -1,33 +1,7 @@
1
- import {
2
- PSTACK_INHERIT_PARENT,
3
- PSTACK_ROLE_NAMES,
4
- PSTACK_ROLES,
5
- type PstackRoleConfig,
6
- } from "./pstack-roles.ts";
7
-
8
1
  const POTETO_PROMPT =
9
2
  "New task? Playbook match or rigor needed -> apply /poteto-mode. Casual turn or user opts out -> don't.";
10
3
 
11
- const ADVISORY =
12
- "Pstack model roles. These are advisory model selections; tools and authority are separate.";
13
-
14
- /** Render configured pstack role lines with cardinality. Inherit-all is empty. */
15
- export function formatPstackRoleTable(config: PstackRoleConfig): string {
16
- const lines: string[] = [];
17
- for (const role of PSTACK_ROLE_NAMES) {
18
- const value = config.roles[role];
19
- if (value === undefined || value === PSTACK_INHERIT_PARENT) continue;
20
- lines.push(`${role} [${PSTACK_ROLES[role].cardinality}]: ${JSON.stringify(value)}`);
21
- }
22
- if (lines.length === 0) return "";
23
- return [ADVISORY, ...lines].join("\n");
24
- }
25
-
26
- /** Assemble the extra system prompt from role table and optional Poteto Mode. */
27
- export function systemPromptInjection(config: PstackRoleConfig, potetoMode: boolean): string {
28
- const parts: string[] = [];
29
- const table = formatPstackRoleTable(config);
30
- if (table) parts.push(table);
31
- if (potetoMode) parts.push(POTETO_PROMPT);
32
- return parts.join("\n\n");
4
+ /** Inject only the Poteto Mode reminder; model-routing resolves roles on demand. */
5
+ export function systemPromptInjection(potetoMode: boolean): string {
6
+ return potetoMode ? POTETO_PROMPT : "";
33
7
  }
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@whamp/pi-pstack",
3
- "version": "0.7.0",
3
+ "version": "0.9.1",
4
4
  "description": "pstack for Pi: rigorous agent workflows you can parallelize with confidence - poteto-mode playbooks, engineering principles, multi-model review panels, and subagents.",
5
5
  "type": "module",
6
6
  "license": "MIT",
@@ -46,7 +46,7 @@ Default: proceed directly to implementation with the synthesized design. No huma
46
46
 
47
47
  Opt in to a checkpoint when the invoker explicitly asks: "/skill:architect with checkpoint," "stop and show me before implementing," or similar. Then surface the synthesized design and pause for sign-off.
48
48
 
49
- The synthesis can ship as its own commit either way, as the "scaffold first" mode of the **foundational-thinking** principle skill. Planned and scoped breakage during fill-in is fine, per the **outcome-oriented-execution** principle skill. For adversarial pressure on the design before implementing, run the **interrogate** skill on the synthesized sketch.
49
+ The synthesis can ship as its own commit either way, as the "scaffold first" mode of the **foundational-thinking** principle skill. Planned and scoped breakage during fill-in is fine, per the **outcome-oriented-execution** principle skill. For adversarial pressure on the design before implementing, read `../code-review/SKILL.md` relative to this skill directory. Use Challenge on the synthesized sketch.
50
50
 
51
51
  If the human pushes back on the shape (in a checkpoint or after the fact), treat that as Phase A evidence. Re-ground and re-run Phase B before writing more code.
52
52
 
@@ -25,7 +25,7 @@ The N candidates will receive the same prompt, so the prompt is the contract.
25
25
 
26
26
  1. State the artifact each candidate is producing.
27
27
  2. Derive the rubric. State what success looks like for *this* task, then turn it into 3-6 concrete gradeable criteria. The rubric is the picker's tool in Phase D. Candidates only see the task.
28
- 3. Pick the runners. Use `arena runners` from `~/.pi/agent/pstack/models.json` when present. Otherwise default to one each on inherit-parent. Spawn more when the arena covers multiple design directions. Same model N times when the work is generation-bound rather than judgment-sensitive.
28
+ 3. Pick the runners. Use `arena runners` from `~/.pi/agent/pstack/models.json` when present. Otherwise default to one each on inherit-parent. An `auto` or `inherit-parent` entry means the parent model, so omit `model` for it. If an explicit selector is unavailable, inspect `subagent({ action: "models", input: {} })`, pick the closest available model (prefer the highest-reasoning tier of the same family), and relaunch. Never treat `inherit-parent` or `auto` as broken selectors. Spawn more when the arena covers multiple design directions. Same model N times when the work is generation-bound rather than judgment-sensitive.
29
29
  4. Assign output paths. Each candidate writes to its own location (a git worktree where possible, otherwise `/tmp/arena-<slug>/candidate-<n>/`), per the **separate-before-serializing-shared-state** principle skill.
30
30
 
31
31
  ## Phase B: Fan out
@@ -38,7 +38,7 @@ If a candidate fails to produce output, pass the completed N-1 results to the ju
38
38
 
39
39
  ## Phase C: Cross-judge
40
40
 
41
- After the Phase B workflow completes, choose one model from the `arena judge pool` in `~/.pi/agent/pstack/models.json` when present. Otherwise use inherit-parent. Prefer a different model family from the parent's. Launch the judge with `subagent({ action: "execute", input: { agent: "worker", task, model, async: true } })`. Its task says to inspect only, read the rubric and candidates by path label, score each criterion, and recommend a base with rationale. Read the completed candidate artifacts while the judge runs. The judge never runs while candidates are writing.
41
+ After the Phase B workflow completes, choose one model from the `arena judge pool` in `~/.pi/agent/pstack/models.json` when present. Otherwise use inherit-parent. Prefer a different model family from the parent's. Launch the judge with `subagent({ action: "execute", input: { agent: "reviewer", task, model, async: true } })`. Its task says to inspect only, read the rubric and candidates by path label, score each criterion, and recommend a base with rationale. Read the completed candidate artifacts while the judge runs. The judge never runs while candidates are writing.
42
42
 
43
43
  ## Phase D: Pick a base
44
44
 
@@ -26,7 +26,7 @@ Update mode changes the rest of the flow:
26
26
 
27
27
  ### 1. Mine their history
28
28
 
29
- Locate the active workspace's transcripts before fanning out. The system prompt names the workspace's `agent-transcripts/` directory. Use only that path. Don't glob across `~/.pi/agent/sessions/`. That crosses workspace boundaries and reads private chats from unrelated projects.
29
+ Locate the active workspace's transcripts before fanning out. Use `~/.pi/agent/sessions/--<slug>--/`, where `<slug>` is the workspace path with the leading slash dropped and each "/" turned into "-". Stay inside that directory. Do not glob sibling slugs under `~/.pi/agent/sessions/`. That crosses workspace boundaries and reads private chats from unrelated projects.
30
30
 
31
31
  Survey recent agent conversations within that scope for recurring patterns. Run multiple parallel subagents across slices of history (e.g. last 2-4 weeks, split into 3 slices so each has enough material). Each slice mining subagent reads transcripts from the workspace-scoped path the parent provides, looks for the signals below, and returns a short structured list of patterns it saw with evidence pointers. Default signals worth hunting:
32
32
 
@@ -41,7 +41,7 @@ Cross-check across slices before elevating a signal. Patterns seen in 2+ slices
41
41
 
42
42
  ### 2. Ask the user directly
43
43
 
44
- Mining misses intent that hasn't come up yet. Use the `ask_user_question` tool. Offer 4-6 options; the user can pick those or type a different answer.
44
+ Mining misses intent that hasn't come up yet. Use the `ask_user_question` tool (structured multi-choice) rather than asking the user to type from scratch.
45
45
 
46
46
  Shape: one or two questions with 4-6 options each, `allow_multiple: true` for category questions. Start broad ("Which areas matter most?"), then follow up on selected areas with specific options. After the structured rounds, one free-form chat question catches anything the options missed.
47
47
 
@@ -26,7 +26,7 @@ For each fact the change's safety depends on, get it as far down this list as is
26
26
  4. You ran it. A script or test that calls the real code and fails loud if you're wrong.
27
27
  5. You reproduced it in the running app.
28
28
 
29
- Any safety fact you can't get to step 4, say so. Don't write it up as settled. Step 4 is usually one small script that imports the same library the app ships and calls the exact function you're worried about.
29
+ Step 4 is usually one small script that imports the same library the app ships and calls the exact function you're worried about.
30
30
 
31
31
  ## Steps
32
32
 
@@ -34,14 +34,14 @@ Any safety fact you can't get to step 4, say so. Don't write it up as settled. S
34
34
  2. Find the one fact it's safe because of. Most changes that look risky are safe because of a single fact, like "this call only drops already-dead cache entries and does nothing else". Find that fact. If it holds, most risky cases are cleared at once. Spend your time here, not on a long list of maybes.
35
35
  3. Look where grep stops. Read the source of the library you call, and check its pinned version and any local patch. Work out when things run: microtasks, unmount and teardown, Solid versus React. Follow what a symbol search misses: the JSON an API returns, a DB column, a wire format, another language reading the same bytes, a feature flag, code three hops downstream.
36
36
  4. Be honest about each risk. Give it a real chance of happening and a real cost if it does. Keep the risks you confirmed. List the ones you checked and cleared separately. Same rules as `why`. Cite a real `file:line`, a search that finds nothing is still an answer, and never make up a caller or an API.
37
- 5. Prove the one fact. Write a script or test that runs the real code, run it, and paste what happened. If you can't prove it cheaply, mark it unproven. Don't overstate.
37
+ 5. Prove the one fact. Write a script or test that runs the real code, run it, and paste what happened.
38
38
  6. For a big or wide change, run it as an `arena`. Ask several models the same question and merge the answers. Different models catch different real bugs.
39
39
 
40
40
  ## What to hand back
41
41
 
42
42
  - **What it does.** What changed, including the part that isn't obvious.
43
43
  - **The one fact it's safe because of.** State it, say which step you got it to, and show the proof. If you couldn't prove it, write unproven.
44
- - **Risks.** Only the real ones. Each names how it breaks, the `file:line`, how likely and how bad, and how to check. Paste the proof for the ones that matter.
44
+ - **Risks.** Each names how it breaks, the `file:line`, how likely and how bad, and how to check. Paste the proof for the ones that matter.
45
45
  - **Cleared.** What you checked and why it's fine.
46
46
  - **Before you merge.** The cheapest test or repro that catches the real bug, including the script you wrote.
47
47
 
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 Matt Pocock
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
@@ -0,0 +1,42 @@
1
+ ---
2
+ name: code-review
3
+ description: "Review a PR, diff, branch, or changes since a fixed point. Audit is the default. Use Challenge only for an explicit adversarial or design review request. Route PR-status requests to Babysit."
4
+ ---
5
+
6
+ # Code review
7
+
8
+ Use this coordinator for code reviews. Keep Audit and Challenge separate. The [Audit procedure](references/code-review-audit.md) defines evidence preparation, reviewer coverage, and the report.
9
+
10
+ ## Route the request
11
+
12
+ - Use Audit for an ordinary or bare request to review a PR, diff, branch, or changes since a point.
13
+ - Use Challenge only when the user explicitly asks for adversarial review or design interrogation. The feature, bug-fix, architect, and PR-opening playbooks can also require Challenge.
14
+ - Run Audit and Challenge when the user explicitly asks for both. Keep their reviewer contexts independent. Audit children do not receive Challenge results. Challenge children do not receive Audit findings.
15
+ - Route PR-status requests, including "check on PR X," to the [Babysit playbook](../poteto-mode/playbooks/babysit.md), not Audit.
16
+ - Opening a PR alone does not start Babysit. The PR-opening playbook still requires Challenge.
17
+
18
+ Ask for a missing Audit base instead of guessing. A named PR supplies immutable base and head commits. Challenge can use pinned design contents without a Git base.
19
+
20
+ ## Freeze evidence before launching reviewers
21
+
22
+ The parent owns preparation and synthesis. Before launching any reviewers, freeze the review intent, exact scope, pinned artifact, relevant spec and standards, completed tool evidence, and limitations. Give each child that same evidence packet. Children inspect only. They do not run Git or shell commands, edit files, write files, or use MCP or extension tools.
23
+
24
+ For Audit, if the spec source is missing and the user has not said that no spec exists, ask for the source before launching Audit reviewers. Challenge needs a clear intent and pinned artifact, not an originating Audit spec. If the user explicitly says no spec is available, record Spec as `SKIPPED (no spec available)`. Do not call it a pass.
25
+
26
+ ## Run Audit
27
+
28
+ Use the shared ordered model lineup from the caller's `model-routing`, spending, and family policy. Each selected model gets one fresh Standards reviewer and one fresh Spec reviewer. These are separate children. Do not combine the axes or choose separate model lineups for them.
29
+
30
+ Use the read-only `reviewer` agent with fresh context through `pi-subagents`. Never substitute a write-capable worker for a reviewer. Follow the [Audit procedure](references/code-review-audit.md). It blocks missing required sources, model families, axes, or reviewer results. Never turn a failed, missing, or partial check into a pass.
31
+
32
+ ## Run Challenge
33
+
34
+ Load the existing [Interrogate skill](../interrogate/SKILL.md). Skip its steps 1 and 2 for a coordinator-delegated Challenge. Start at step 3 with the frozen artifact, scope, and intent verbatim. Do not run Git, rediscover the artifact, or derive intent from the implementation. Keep its reviewer rubric and synthesis. Do not create another Challenge rubric or reviewer roster. Direct `/skill:interrogate` use remains compatible.
35
+
36
+ Do not trigger Challenge based on risk. A completed Audit never satisfies a required Challenge. Reuse prior coverage only when the pinned artifact, intent, rubric, required axes and families, and parent disposition all match.
37
+
38
+ ## Report the review
39
+
40
+ Record the requested and actual model selectors. State any substitution, covered families, missing coverage, child failures, findings, unresolved questions, and the parent's disposition for each finding. Keep Standards, Spec, and Challenge results separate. Do not create a separate review registry or apply fixes automatically.
41
+
42
+ This coordinator adapts Matt Pocock's MIT-licensed `code-review` skill. See [the retained MIT notice](LICENSE).
@@ -0,0 +1,91 @@
1
+ # Audit
2
+
3
+ Audit answers two different questions. Standards asks whether the change follows the repository's documented standards. Spec asks whether the change meets its originating request. Keep the answers separate.
4
+
5
+ ## Prepare the review packet
6
+
7
+ ### Pin the artifact
8
+
9
+ 1. Resolve the supplied fixed base and pinned head to commit SHAs. For a named PR, use its base and head SHAs.
10
+ 2. Preserve three-dot scope. Run `git diff <fixed-base>...<pinned-head>` and resolve and record the merge base it uses. Record commits with `git log <fixed-base>..<pinned-head>`. Do not replace this with direct endpoint subtraction using two dots.
11
+ 3. Ask for the base if the request omits it. Do not guess `main`, `HEAD`, or a merge base.
12
+ 4. For a file or design review, pin the file contents by commit or record an immutable content digest. A design review without Git history can still use Challenge.
13
+ 5. Record the exact resolved commands, commit list, changed-file list, and any scope exclusions.
14
+ 6. Stop before launching children if a ref cannot resolve or the complete requested artifact is empty. An empty committed diff can still have requested WIP to review. Do not reject that WIP merely because no committed changes exist.
15
+
16
+ Review a named PR's committed change by default. Add local work in progress only when requested. When WIP is requested, include staged and unstaged tracked changes and every requested untracked file. A regular `git diff` omits staged and untracked content. Keep the committed change and local WIP evidence distinct so the packet does not duplicate or hide changes.
17
+
18
+ ### Find the spec
19
+
20
+ Trace the request to its source. Check issue references in commits, user-provided paths, matching documents in `docs/`, `specs/`, or `.scratch/`, and the relevant issue tracker. Record the source URL or path, stable requirement reference, and the exact text that governs the change.
21
+
22
+ If no source is found and the user has not said that no spec exists, ask for the spec before reviewers start. If the user explicitly says there is no spec, mark Spec `SKIPPED (no spec available)`. Do not report Spec as passed or infer requirements from the implementation.
23
+
24
+ ### Find the standards
25
+
26
+ Inspect the repository's maintained standards and applicable language guidance. Record each source's URL, path, revision, and relevant rule IDs. The implementation is not its own standard.
27
+
28
+ If a required standard source is missing, block the Standards axis and ask for the source. For a language with no adopted language standard, state `no language standard adopted for <language>; language-rule audit skipped.` Do not invent language rules. Keep Standards open for documented cross-language rules and the labelled Fowler heuristics below. Spec remains a separate axis.
29
+
30
+ For every standards finding, cite a stable rule ID, a source URL, the changed file and line, why the code breaks the rule, and a concrete fix. Use the source's published rule ID. If it has none, label the rule with a stable path and heading anchor. Do not invent an upstream rule number.
31
+
32
+ ### Complete tool checks
33
+
34
+ Before reviewer children start, the parent runs every applicable check required by the standards and records its exact command and result. Inspect tool configuration. Confirm the check covers every changed and new file. Record completed checks with their diagnostics. Record unavailable, failed, incomplete, and inapplicable checks separately. A tool error or incomplete coverage is a gap, not a pass.
35
+
36
+ The parent gathers all spec and standards sources, exact tool output, and limitations into one frozen packet. Reviewer children receive that packet and read-only file pointers. They do not run Git or shell commands, edit or write files, or use MCP or extension tools.
37
+
38
+ ## Select and run Audit reviewers
39
+
40
+ Choose one ordered Audit model lineup for both axes. Follow the caller's `model-routing`, spending, and family policy, plus any explicit user selection. Do not select separate model lists for Standards and Spec. If no private routing policy exists, inherit the parent model and report the actual resolved model and family.
41
+
42
+ Record requested selectors separately from actual selectors. Record every substitution. Do not silently claim a substituted model covers a required family. Missing required families, failed launches, timeouts, and missing outputs remain unresolved coverage.
43
+
44
+ Discover available agents with `subagent({ action: "list", input: { capabilities: true } })`. Launch the Audit children together in one `workflowScript` using `runs.all`. Each item uses `agent: "reviewer"`, `context: "fresh"`, a stable key, the prepared packet, and its resolved model selector. Count all children in the run's spawn budget. The parent synthesizes the returned reports.
45
+
46
+ For each selected model, launch two fresh reviewer children:
47
+
48
+ 1. A Standards child that receives the exact frozen diff, relevant standards, completed diagnostics, and limitations. It reviews Standards only.
49
+ 2. A Spec child that receives the exact frozen diff, originating spec, and relevant context. It reviews Spec only.
50
+
51
+ The children use the same ordered model lineup, but their contexts stay separate. Do not combine both axes in one child. Do not copy one model's answer to fill another model's missing coverage. If the user explicitly says that no spec exists, launch no Spec children and report the axis as skipped.
52
+
53
+ Each child inspects only the prepared artifact. Ask it to report supported findings, file and line, source evidence, and any gaps. A child's zero findings do not turn another missing or failed result into a pass.
54
+
55
+ ## Review Standards findings
56
+
57
+ Start with documented rules. The repository's rule overrides general guidance. Skip a judgment heuristic when the repository endorses the pattern or a tool already enforces it. Label each heuristic finding as a possible smell, never a hard rule violation.
58
+
59
+ Include the twelve Fowler heuristics below in every Standards packet. They are judgment prompts, not hard rules. They remain subordinate to documented standards and tool results.
60
+
61
+ 1. **Mysterious Name.** A function, variable, or type has a name that does not reveal what it does or holds. Rename it. If no honest name fits, clarify the design.
62
+ 2. **Duplicated Code.** The same logic shape appears more than once in the change. Extract the shared shape and call it from both sites.
63
+ 3. **Feature Envy.** A method reads another object's data more than its own. Move the method to the data it uses.
64
+ 4. **Data Clumps.** The same fields or parameters travel together repeatedly. Group them in one type and pass that type.
65
+ 5. **Primitive Obsession.** A primitive or string stands in for a domain concept that needs its own type. Give the concept a small type.
66
+ 6. **Repeated Switches.** The same conditional on the same type recurs in the change. Replace it with polymorphism or one shared map.
67
+ 7. **Shotgun Surgery.** One logical change forces scattered edits across many files. Gather the related behavior in one module.
68
+ 8. **Divergent Change.** One file or module changes for several unrelated reasons. Split it so each module changes for one reason.
69
+ 9. **Speculative Generality.** An abstraction, parameter, or hook serves a need the spec does not have. Delete it and inline the code until a real need appears.
70
+ 10. **Message Chains.** A long chain such as `a.b().c().d()` exposes navigation the caller should not depend on. Hide the walk behind a method on the first object.
71
+ 11. **Middle Man.** A class or function mostly delegates to another. Remove it and call the target directly.
72
+ 12. **Refused Bequest.** A subclass or implementer ignores or overrides most of what it inherits. Replace inheritance with composition.
73
+
74
+ ## Review Spec findings
75
+
76
+ Compare the implementation with each relevant requirement in the frozen spec. Cite the exact requirement and the changed file and line. Report missing or partial requirements, unrequested behavior, and implementations that appear wrong. Do not use code comments or the implementation itself as the spec.
77
+
78
+ ## Synthesize and report coverage
79
+
80
+ The parent checks citations against the pinned packet and keeps disagreement visible. Report a finding total and the worst issue for Standards and Spec separately. Do not rank an issue in one axis against an issue in the other. Report each axis as findings, no findings, blocked, or skipped. `No findings` means the required review completed and found none. `Blocked` means required evidence or execution is missing. `Skipped` means the caller explicitly made the axis inapplicable.
81
+
82
+ Keep a review record in the response, not a new registry. Include:
83
+
84
+ - The intent, exact scope, pinned base and head or content digest, and spec provenance.
85
+ - The standards sources, diagnostics, and completed, blocked, skipped, or incomplete checks.
86
+ - The requested and actual model selectors, substitutions, required and covered families, and child failures.
87
+ - The finding total and worst issue within each axis, with no cross-axis ranking.
88
+ - Separate Standards and Spec results, findings, unresolved coverage, and the parent's disposition for each finding.
89
+ - Spec `SKIPPED (no spec available)` when the user explicitly declares that no spec exists.
90
+
91
+ Do not merge Standards and Spec into one verdict. Do not mark incomplete coverage as passed. Do not apply fixes during the review.
@@ -6,7 +6,7 @@ disable-model-invocation: true
6
6
 
7
7
  # Figure it out
8
8
 
9
- When the task matches no playbook, design one. The deliverable before any code is the workflow itself: a sequence of phases that scales rigor to the task, runs the scientific method, and leaves a decision trail a human can audit after stepping away. Bias toward more rigor. The cost of building the wrong thing dwarfs the cost of being careful.
9
+ When the task matches no playbook, design one. The deliverable before any code is the workflow itself: a sequence of phases that scales rigor to the task, runs the scientific method, and leaves a decision trail a human can audit after stepping away.
10
10
 
11
11
  ## Start
12
12
 
@@ -39,12 +39,12 @@ Each unit is an experiment. State the hypothesis, make the smallest change, meas
39
39
  Apply the **sequence-verifiable-units** principle skill, verifying each unit before starting the next instead of batching checks at the end.
40
40
 
41
41
  - Verify by inspecting the artifact, never a self-report. When something passes too easily, suspect the observation method before the system.
42
- - Pair delegated work with a judge and audit the delegates' artifacts yourself before trusting them. If a worker games the gate, reset and harden the contract. If the gate itself is wrong, fix the gate in its own change rather than routing around it.
42
+ - Pair delegated work with a judge. If a worker games the gate, reset and harden the contract. If the gate itself is wrong, fix the gate in its own change rather than routing around it.
43
43
  - A verdict is VERIFIED, NOT VERIFIED, or INCONCLUSIVE. Inconclusive is not a pass. Don't hide a negative.
44
44
 
45
45
  ## Phase D: Keep the audit trail
46
46
 
47
- Log the run via the **show-me-your-work** skill, one canonical TSV with a row per decision and per unit, evidence as links. figure-it-out's work is usually ambitious enough to commit the trail so the reviewer can read it in the PR. Commit it when confidence has to be shown. Prefer evidence produced by committed scripts. The trail plus the diff is what lets the human come back and trust the work.
47
+ Log the run via the **show-me-your-work** skill. figure-it-out's work is usually ambitious enough to commit the trail so the reviewer can read it in the PR. The trail plus the diff is what lets the human come back and trust the work.
48
48
 
49
49
  ## Phase E: Verify and hand back
50
50
 
@@ -1,12 +1,15 @@
1
1
  ---
2
2
  name: how
3
3
  description: "Use for \"how does X work\", code walkthroughs before changing something, and placement / ownership / layering questions (\"where should this live\", \"which package owns this\", \"is this the right layer\"). Explains subsystem architecture, runtime flow, onboarding mental models. Use why for motivation."
4
+ disable-model-invocation: true
4
5
  ---
5
6
 
6
7
  # How
7
8
 
8
9
  Explore the codebase to answer "how does X work?" questions. Produce architectural explanations at the level of a senior engineer onboarding onto a subsystem, enough to build a working mental model, not so much that it reads like annotated source code.
9
10
 
11
+ Each child names a role in `~/.pi/agent/pstack/models.json`. Use that role's selector. Omit `model` when the value is `inherit-parent` or `auto`. If an explicit selector is unavailable, inspect `subagent({ action: "models", input: {} })`, pick the closest available model (prefer the highest-reasoning tier of the same family), and relaunch. Never treat `inherit-parent` or `auto` as broken selectors.
12
+
10
13
  ## Step 1. Assess Complexity
11
14
 
12
15
  If the scope is ambiguous, state your interpretation and explore. The user can redirect.
@@ -18,10 +21,10 @@ When in doubt, take the simple path.
18
21
 
19
22
  ## Step 2a. Explore (complex questions only)
20
23
 
21
- Decompose the question into 2 to 4 exploration angles, each a distinct slice of the subsystem. Launch the explorers and dependent explainer with one `subagent({ action: "execute", input: { async: true, maxSubagentSpawnsPerRun: N + 1, workflowScript } })` call. In `workflowScript`, await `runs.all([{ key: "explore-<angle>", agent: "worker", task, model }])`, then return `runs.run("explain", { agent: "worker", task, model })` with the explorer outputs.
24
+ Decompose the question into 2 to 4 exploration angles, each a distinct slice of the subsystem. Launch the explorers and dependent explainer with one `subagent({ action: "execute", input: { async: true, maxSubagentSpawnsPerRun: N + 1, workflowScript } })` call. In `workflowScript`, await `runs.all([{ key: "explore-<angle>", agent: "reviewer", task, model }])`, then return `runs.run("explain", { agent: "reviewer", task, model })` with the explorer outputs.
22
25
 
23
26
  Each explorer uses:
24
- - agent: "worker"
27
+ - agent: "reviewer"
25
28
  - `model`: `how explorers` (default inherit-parent)
26
29
  - `task`: the prompt in `references/explorer-prompt.md` with its angle filled in and an instruction to inspect only
27
30
 
@@ -29,8 +32,8 @@ Then go to Step 3.
29
32
 
30
33
  ## Step 2b. Direct Explain (simple questions)
31
34
 
32
- Launch one standalone child with `subagent({ action: "execute", input: { agent: "worker", task, model, async: false } })` using:
33
- - agent: "worker"
35
+ Launch one standalone child with `subagent({ action: "execute", input: { agent: "reviewer", task, model, async: false } })` using:
36
+ - agent: "reviewer"
34
37
  - `model`: `how explainer` (default inherit-parent)
35
38
  - `task`: `references/explainer-prompt.md` without the explorer-findings section and with an instruction to inspect only
36
39
 
@@ -39,7 +42,7 @@ Go to Step 4.
39
42
  ## Step 3. Synthesize (complex questions only)
40
43
 
41
44
  The same workflow launches `explain` after every explorer settles using:
42
- - agent: "worker"
45
+ - agent: "reviewer"
43
46
  - `model`: `how synthesizer` (default inherit-parent)
44
47
  - `task`: `references/explainer-prompt.md` with every explorer result filled in and an instruction to inspect only
45
48
 
@@ -0,0 +1,2 @@
1
+ policy:
2
+ allow_implicit_invocation: false
@@ -4,7 +4,7 @@ Build each explorer subagent's prompt from this template. Fill in the placeholde
4
4
 
5
5
  ---
6
6
 
7
- You are exploring a codebase to understand how something works. Gather facts: trace code paths, read implementations, map components. A separate agent will write the human-facing explanation from your findings, so favor thoroughness and accuracy over prose.
7
+ You are exploring a codebase to understand how something works. Gather facts. Trace code paths, read implementations, map components. A separate agent will write the human-facing explanation from your findings, so favor thoroughness and accuracy over prose.
8
8
 
9
9
  Other explorers are investigating different slices of the same subsystem in parallel. Don't try to cover everything. Focus on your assigned angle and go deep.
10
10
 
@@ -33,17 +33,16 @@ Write one clear paragraph. If you're unsure about the intent, ask the user befor
33
33
 
34
34
  ## Step 3, Spawn Reviewers
35
35
 
36
- Launch all reviewers with one `subagent({ action: "execute", input: { async: true, maxSubagentSpawnsPerRun: N, workflowScript } })` call. In `workflowScript`, use `return await runs.all([{ key: "reviewer-a", agent: "worker", task, model }])` with one stable-keyed item per reviewer. Use the `interrogate reviewers` list from `~/.pi/agent/pstack/models.json` when present, one reviewer per entry, extending or shrinking the Reviewer A/B/C/D labels below to the configured entry count. Otherwise use the table defaults.
36
+ Launch all reviewers with one `subagent({ action: "execute", input: { async: true, maxSubagentSpawnsPerRun: N, workflowScript } })` call. In `workflowScript`, use `return await runs.all([{ key: "reviewer-a", agent: "reviewer", task, model }])` with one stable-keyed item per reviewer. Use the `interrogate reviewers` list from `~/.pi/agent/pstack/models.json` when present, one reviewer per entry, extending or shrinking the Reviewer A/B/C labels below to the configured entry count. Otherwise use the table defaults.
37
37
 
38
38
  | Subagent | Default model |
39
39
  |----------|---------------|
40
40
  | Reviewer A | inherit-parent |
41
41
  | Reviewer B | inherit-parent |
42
42
  | Reviewer C | inherit-parent |
43
- | Reviewer D | inherit-parent |
44
43
 
45
44
  For each reviewer:
46
- - agent: "worker"
45
+ - agent: "reviewer"
47
46
  - `model`: the configured `interrogate reviewers` entry, or the table default with no configured line
48
47
  - `task`: instruct the reviewer to inspect only and not modify files
49
48
 
@@ -36,7 +36,7 @@ Each dimension is stated once. Apply the ones that are relevant.
36
36
 
37
37
  ## Output Expectations
38
38
 
39
- Prioritize structural code-quality regressions and missed simplifications first, then spaghetti and branching complexity, then boundary, type, and file-size concerns, then smaller modularity and legibility issues. Do not flood the review with low-value nits when larger structural issues exist. Prefer a few high-conviction comments over a long list of cosmetic notes.
39
+ Prioritize structural code-quality regressions and missed simplifications first, then spaghetti and branching complexity, then boundary, type, and file-size concerns, then smaller modularity and legibility issues.
40
40
 
41
41
  ## Approval Bar
42
42
 
@@ -35,7 +35,7 @@ For each finding, provide:
35
35
  1. **Severity**: `critical` | `warning` | `nit`
36
36
  - `critical`: Would cause bugs, data loss, security issues, or fundamentally broken behavior
37
37
  - `warning`: Design concern, maintainability risk, or correctness issue that isn't immediately broken but will cause pain
38
- - `nit`: Style, naming, minor improvement. Only include nits if they're genuinely useful, not to pad your review.
38
+ - `nit`: Style, naming, minor improvement.
39
39
  2. **Finding**: What the problem is, in concrete terms. Reference specific lines/functions.
40
40
  3. **Evidence**: Why you believe this is a problem. Show your reasoning. Don't just assert.
41
41
  4. **Suggestion** (optional): What you'd do instead, if you have a concrete alternative. Skip this if you don't have a clear fix.
@@ -50,8 +50,6 @@ For each finding, provide:
50
50
  ## What to Avoid
51
51
 
52
52
  - Restating what the code does without identifying a problem
53
- - Suggesting rewrites for working code because you'd prefer a different style
54
- - Raising hypothetical issues ("what if someone passes null here") without evidence that the code path is reachable
55
53
  - Praising the code. You're an adversary, not a cheerleader. If you find nothing wrong, say "no findings" and stop.
56
54
 
57
55
  ## Output
@@ -69,7 +69,7 @@ Simpler is better unless simpler is wrong. Three lines of duplication beat a pre
69
69
 
70
70
  ## Security
71
71
 
72
- Only flag security issues you can actually trace through the code. "This could be an injection vector" without showing the input path is not useful.
72
+ For each security finding, trace the input path through the code and show it.
73
73
 
74
74
  - User input flowing to dangerous sinks (SQL, shell, eval, innerHTML) without sanitization
75
75
  - Authentication/authorization gaps in new endpoints