@whamp/pi-pstack 0.7.0 → 0.9.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (55) hide show
  1. package/README.md +38 -11
  2. package/extensions/pstack/index.ts +7 -4
  3. package/extensions/pstack/pstack-role-prompt.ts +3 -29
  4. package/package.json +1 -1
  5. package/skills/architect/SKILL.md +1 -1
  6. package/skills/arena/SKILL.md +2 -2
  7. package/skills/automate-me/SKILL.md +2 -2
  8. package/skills/blast-radius/SKILL.md +3 -3
  9. package/skills/code-review/LICENSE +21 -0
  10. package/skills/code-review/SKILL.md +42 -0
  11. package/skills/code-review/references/code-review-audit.md +91 -0
  12. package/skills/figure-it-out/SKILL.md +3 -3
  13. package/skills/how/SKILL.md +8 -5
  14. package/skills/how/agents/openai.yaml +2 -0
  15. package/skills/how/references/explorer-prompt.md +1 -1
  16. package/skills/interrogate/SKILL.md +2 -3
  17. package/skills/interrogate/references/code-quality-review.md +1 -1
  18. package/skills/interrogate/references/reviewer-prompt.md +1 -3
  19. package/skills/interrogate/references/rubric.md +1 -1
  20. package/skills/poteto-mode/SKILL.md +17 -7
  21. package/skills/poteto-mode/playbooks/autopilot-full.md +6 -6
  22. package/skills/poteto-mode/playbooks/autopilot-stack.md +7 -7
  23. package/skills/poteto-mode/playbooks/babysit.md +1 -1
  24. package/skills/poteto-mode/playbooks/bug-fix.md +3 -5
  25. package/skills/poteto-mode/playbooks/eval.md +1 -1
  26. package/skills/poteto-mode/playbooks/feature.md +3 -3
  27. package/skills/poteto-mode/playbooks/hillclimb.md +1 -0
  28. package/skills/poteto-mode/playbooks/multi-phase-plan.md +8 -8
  29. package/skills/poteto-mode/playbooks/opening-a-pr.md +1 -1
  30. package/skills/poteto-mode/playbooks/pause-safely.md +1 -1
  31. package/skills/poteto-mode/playbooks/perf-issue.md +1 -0
  32. package/skills/poteto-mode/playbooks/refactoring.md +2 -2
  33. package/skills/poteto-mode/playbooks/session-pickup.md +1 -1
  34. package/skills/poteto-mode/playbooks/shipping.md +2 -2
  35. package/skills/principle-guard-the-context-window/SKILL.md +0 -1
  36. package/skills/principle-never-block-on-the-human/SKILL.md +0 -2
  37. package/skills/principle-outcome-oriented-execution/SKILL.md +0 -1
  38. package/skills/principle-prove-it-works/SKILL.md +0 -11
  39. package/skills/principle-sequence-verifiable-units/SKILL.md +0 -5
  40. package/skills/recall/SKILL.md +1 -1
  41. package/skills/reflect/SKILL.md +7 -7
  42. package/skills/reflect/references/divergent-reviewer.md +1 -1
  43. package/skills/reflect/references/judgment-reviewer.md +1 -1
  44. package/skills/reflect/references/tooling-reviewer.md +1 -1
  45. package/skills/show-me-your-work/SKILL.md +7 -7
  46. package/skills/show-me-your-work/scripts/log.sh +4 -2
  47. package/skills/swarm/SKILL.md +4 -4
  48. package/skills/tdd/SKILL.md +1 -3
  49. package/skills/technical-writing/SKILL.md +0 -13
  50. package/skills/typescript-best-practices/SKILL.md +1 -0
  51. package/skills/typescript-best-practices/agents/openai.yaml +2 -0
  52. package/skills/unslop/SKILL.md +1 -1
  53. package/skills/unslop/agents/openai.yaml +2 -0
  54. package/skills/why/SKILL.md +6 -3
  55. package/skills/why/agents/openai.yaml +2 -0
@@ -21,7 +21,7 @@ Copy `references/decision-log-template.tsv` (the header row) to start a clean lo
21
21
  - **evidence.** A link or path that proves it: commit SHA, PR number, `file:line`, or an artifact, trace, or screenshot path. Never a paragraph.
22
22
  - **result.** The outcome or predicate state: `tests green`, `reverted`, `pixel-diff 0`, `INCONCLUSIVE`, `open`.
23
23
 
24
- An example, plain-spoken so a reviewer reads it at a glance. This is illustration only. Don't copy these rows into a real log.
24
+ An example, plain-spoken so a reviewer reads it at a glance.
25
25
 
26
26
  ```
27
27
  ts phase decision why evidence result
@@ -39,6 +39,8 @@ Use the helper `scripts/log.sh <logfile> <phase> <decision> <why> <evidence> <re
39
39
 
40
40
  Log decision points and checkpoints, not every action: a fork chosen, a unit completed with its verification result, a pivot or revert with its trigger, a blocker surfaced, a gate fixed. For loop runs, one row per iteration. Skip the trivial and self-evident.
41
41
 
42
+ A run is one agent conversation, including its later turns and any summary of it. A pickup, a replacement agent, or a new chat starts a new run. When a run adds to a log that already has rows, its first row has phase `start`, and so does its first row after another run's `start` row. So a run that comes back to a log in a later turn first reads the log's last rows to see whether another run wrote since. A `start` row names the `ts` range of the rows before it that this run did not write, and its evidence names this run, such as its agent id. Use phase `start` for nothing else.
43
+
42
44
  ## Where it lives
43
45
 
44
46
  By default the log is a working artifact, not committed. Keep it at `decisions.tsv` in the work dir, or `.audit/<task-slug>.tsv` when several efforts run at once, and leave it out of git.
@@ -47,20 +49,18 @@ Commit it only when the work is ambitious enough that a reviewer needs the trail
47
49
 
48
50
  ## Rules
49
51
 
50
- - One row is one decision or checkpoint.
51
52
  - Append-only. A wrong call gets a new row that supersedes it. Never edit or delete history.
52
53
  - Prefer evidence produced by committed scripts over hand-made one-offs (the **encode-lessons-in-structure** principle skill).
53
54
 
54
55
  ## Audit the log against the transcript
55
56
 
56
- At the end of the run, before handing back, check the log told the truth. Read this run's transcript under the active workspace's `agent-transcripts/` directory (the system prompt names the path). Don't glob across `~/.pi/agent/sessions/`. That reads unrelated private chats. Walk the log against what actually happened:
57
+ At the end of the run, before handing back, check the log told the truth. Read this run's transcript. Prefer `$PI_SESSION_FILE`. Otherwise use `~/.pi/agent/sessions/--<slug>--/` (`<slug>` is the workspace path with the leading slash dropped and each "/" turned into "-"). Stay inside that directory. Do not glob sibling slugs under `~/.pi/agent/sessions/`. That reads unrelated private chats. Walk this run's rows against what actually happened. Each stretch of them begins at one of this run's `start` rows, or at the first row if this run created the log, and ends at the next `start` row of another run:
57
58
 
58
- - Every row maps to a real action. Cut invented or aspirational entries.
59
- - Each row's evidence resolves and shows what the row claims.
59
+ - Check that every row maps to a real decision or action.
60
+ - Check that each row's evidence resolves and shows what the row claims.
60
61
  - A fork, pivot, or abandoned approach that shaped the work but isn't logged is a gap. Add it.
61
- - Drop padding.
62
62
 
63
- Fix the log, not the story. If the work diverged from what a row claims, the row is wrong.
63
+ Correct the log, not the story. The audit never edits or removes a row, even an invented one. When a row records neither a real decision nor a real action, or its claim or evidence is wrong, add a row that supersedes it with what actually happened and a pointer that resolves. This audit does not check rows outside this run's stretches. If this run's own work shows one of them is wrong, supersede it like any wrong call.
64
64
 
65
65
  ## Cross-model review of the trail
66
66
 
@@ -16,8 +16,10 @@ if [ -n "$logdir" ] && [ "$logdir" != "." ] && [ ! -d "$logdir" ]; then
16
16
  mkdir -p "$logdir"
17
17
  fi
18
18
 
19
- if [ ! -f "$logfile" ]; then
20
- printf 'ts\tphase\tdecision\twhy\tevidence\tresult\n' > "$logfile"
19
+ # Use `>>` here, never `>`. A network mount can fail this test for a log
20
+ # that exists. Then the cost is one stray header line, not the rows.
21
+ if [ ! -s "$logfile" ]; then
22
+ printf 'ts\tphase\tdecision\twhy\tevidence\tresult\n' >> "$logfile"
21
23
  fi
22
24
 
23
25
  ts="$(date -u +%Y-%m-%dT%H:%M:%SZ)"
@@ -22,8 +22,8 @@ Open a todolist with one entry per phase before launching anything.
22
22
  1. State the done predicate and the artifact or report the swarm must return.
23
23
  2. Choose the shape. Partition into slices, race N workers on identical briefs, or mix both. For a race or mixed shape, declare `first pass`, `rank all`, or `best-of` before spawning.
24
24
  3. Set N from the user or derive it from the shape. N is total workers, not the cloud concurrency limit.
25
- 4. Pick the worker model from `swarm workers` in `~/.pi/agent/pstack/models.json` when present. Otherwise use inherit-parent. For a model race, name each arm's model up front.
26
- 5. Give each worker its own writable output when it writes.
25
+ 4. Pick the worker model from `swarm workers` in `~/.pi/agent/pstack/models.json` when present. Otherwise use inherit-parent. For `auto` or `inherit-parent`, omit `model`. If an explicit selector is unavailable, inspect `subagent({ action: "models", input: {} })`, pick the closest available model (prefer the highest-reasoning tier of the same family), and relaunch. Never treat `inherit-parent` or `auto` as broken selectors. For a model race, name each arm's model up front.
26
+ 5. Give each worker its own writable output when it writes. When workers verify or measure commits, each brief names the exact SHAs. A measurement brief also names the method (sample count, what one sample is, order). The worker records both in its result.
27
27
 
28
28
  ## Phase B: Fan out
29
29
 
@@ -31,13 +31,13 @@ Launch all N workers with one `subagent({ action: "execute", input: { async: tru
31
31
 
32
32
  Set `input.cwd` to an existing checkout. To create a managed checkout from a Git ref, set `input.worktree: true` and `input.baseRef`.
33
33
 
34
- Every brief stands alone. Include the goal, scope, exact slice or race arm, how to verify, and what to report. Reports use `PASS`, `ISSUES`, or `BLOCKED` with evidence.
34
+ Every brief stands alone. Include the goal, scope, exact slice or race arm, how to verify, and what to report. Reports use `PASS`, `ISSUES`, or `BLOCKED` with evidence. A worker that can prove a defect reports `ISSUES` and lists every issue it can prove, not only the first.
35
35
 
36
36
  If a worker drops out, proceed with N-1 and note it.
37
37
 
38
38
  ## Phase C: Aggregate
39
39
 
40
- Read the terminal results. For coverage, every required slice needs a result. For a race, apply the selection rule declared up front. Use first pass, rank all, or best-of. Do not paste raw worker dumps.
40
+ Read the terminal results. Drop a result that does not record the SHAs and method its brief names, and rerun that worker once. After a second miss, record a gap. A gap does not count as a pass. For coverage, every required slice needs a result. For a race, apply the selection rule declared up front. Use first pass, rank all, or best-of. Do not paste raw worker dumps.
41
41
 
42
42
  Keep a compact result table, one-line evidenced issues, and explicit gaps or dropouts.
43
43
 
@@ -18,11 +18,10 @@ Do not force a test when it would be impractical. If the available test would re
18
18
  4. **Run the new test before fixing.** Confirm it fails for the intended reason. If it passes or fails for an unrelated reason, correct the test or reproduction before editing the implementation.
19
19
  5. **Fix the bug.** Make the smallest production change that satisfies the intended behavior while preserving nearby contracts.
20
20
  6. **Rerun the regression test.** Confirm the test now passes.
21
- 7. **Run nearby validation.** Run relevant adjacent tests, type checks, lint, or scenario checks when the change has broader risk.
22
21
 
23
22
  ## If a Failing Test Is Impractical
24
23
 
25
- Do not silently skip the regression step. Before fixing, explicitly explain why a failing test is impossible or not worth the cost, then choose the closest executable regression check available. Examples include a targeted script, manual reproduction command, browser automation, snapshot comparison, log assertion, or focused integration check.
24
+ Use the closest executable regression check instead: a targeted script, manual reproduction command, browser automation, snapshot comparison, log assertion, or focused integration check.
26
25
 
27
26
  Prefer no new test over a bad test. A bad test is one that mostly tests mocks, encodes current implementation details, depends on timing or unrelated global state, needs expensive infrastructure for a small fix, or would be deleted immediately after proving the fix.
28
27
 
@@ -31,7 +30,6 @@ Prefer no new test over a bad test. A bad test is one that mostly tests mocks, e
31
30
  - Do not change tests merely to match a wrong implementation.
32
31
  - Do not weaken existing assertions unless the expected behavior has genuinely changed and the reason is clear.
33
32
  - Keep the regression test focused on the bug. Avoid broad fixture churn or unrelated coverage expansion.
34
- - Do not add tests when the practical signal is weak. Use manual or scripted verification and say why.
35
33
  - If the bug is flaky, make the test deterministic where possible and document the signal being locked down.
36
34
  - If the bug exposes a broader class of failures, first land the focused regression path, then consider additional sibling coverage.
37
35
 
@@ -112,16 +112,3 @@ Before:
112
112
  After:
113
113
 
114
114
  > `budget.mjs` reads the committed budget from `budget.json` and counts the files that import protos. If the count exceeds the budget, CI fails. Run `budget.mjs --write` only to lower the budget.
115
-
116
- ## Review checklist
117
-
118
- Apply to any prose this skill covers. Item 1 applies only to document sets:
119
-
120
- 1. Is each file one Diátaxis mode, with links where modes meet?
121
- 2. Is every instruction written as a command, with its condition in front?
122
- 3. Does any sentence carry two instructions or two thoughts? Split it.
123
- 4. Can any word be cut without losing meaning? Cut it.
124
- 5. Is "only" next to the word it changes? Does every "it" point at one thing? Does every clause keep its verb?
125
- 6. Does each thing have exactly one name across the docs?
126
- 7. Would a developer say these words out loud? Replace invented metaphors and fancy synonyms with the plain word or the real symbol name.
127
- 8. Are all symbols, paths, and counts real at this commit, with the commands that regenerate the counts?
@@ -1,6 +1,7 @@
1
1
  ---
2
2
  name: typescript-best-practices
3
3
  description: TypeScript best practices. Use when reading or editing any .ts or .tsx file.
4
+ disable-model-invocation: true
4
5
  ---
5
6
 
6
7
  # TypeScript best practices
@@ -0,0 +1,2 @@
1
+ policy:
2
+ allow_implicit_invocation: false
@@ -1,6 +1,7 @@
1
1
  ---
2
2
  name: unslop
3
3
  description: Cut AI tells from any writing. Must always apply.
4
+ disable-model-invocation: true
4
5
  ---
5
6
 
6
7
  # Unslop
@@ -11,7 +12,6 @@ Edit text to remove AI patterns.
11
12
 
12
13
  1. Scan for the patterns below.
13
14
  2. Rewrite. Preserve meaning, match intended tone.
14
- 3. Self-audit: "What makes this obviously AI generated?" Fix remaining tells.
15
15
 
16
16
  ## Patterns to detect and fix
17
17
 
@@ -0,0 +1,2 @@
1
+ policy:
2
+ allow_implicit_invocation: false
@@ -1,6 +1,7 @@
1
1
  ---
2
2
  name: why
3
3
  description: "Use for 'why does X work this way', 'why we picked Y', design rationale, regressions, postmortems, or data-backed thresholds. Discovers available MCPs and queries each evidence category (source control, issue tracker, long-form docs, real-time chat, infrastructure observability, error tracking, product analytics warehouse) in parallel, then returns a cited read on decisions and tradeoffs. Use how for runtime behavior."
4
+ disable-model-invocation: true
4
5
  ---
5
6
 
6
7
  # Why
@@ -9,6 +10,8 @@ Investigate the motivation and intent behind code.
9
10
 
10
11
  Companion to the `how` skill. `how` answers what the code does and how it works. `why` answers what forces led to its shape.
11
12
 
13
+ Each child names a role in `~/.pi/agent/pstack/models.json`. Use that role's selector. Omit `model` when the value is `inherit-parent` or `auto`. If an explicit selector is unavailable, inspect `subagent({ action: "models", input: {} })`, pick the closest available model (prefer the highest-reasoning tier of the same family), and relaunch. Never treat `inherit-parent` or `auto` as broken selectors.
14
+
12
15
  ## Operating Posture
13
16
 
14
17
  Operate as a **careful, cautious, and precise investigator**. Be honest about what you know vs what you're inferring. Read `references/epistemics.md` for the full confidence framework and phrasing guide. The synthesizer must follow it.
@@ -72,10 +75,10 @@ Source control is always available through git and `gh`. For the other six, clas
72
75
 
73
76
  Aim for a complete **coverage map**, not a minimal one. Document the null, don't skip the search. The parent queries each available MCP and builds one bounded evidence packet per category before launching children. A child does not inherit ambient MCP or extension tools. Use a custom agent for a child-side lookup only when that agent explicitly lists the tool and loads its provider through `extensions` or `subagentOnlyExtensions`.
74
77
 
75
- Launch all matching investigators and the dependent synthesizer with one `subagent({ action: "execute", input: { async: true, maxSubagentSpawnsPerRun: N + 1, workflowScript } })` call. In `workflowScript`, await the investigators with `runs.all([{ key: "investigate-<category>", agent: "worker", task, model }])`, then return `runs.run("synthesize-why", { agent: "worker", task, model })` with their outputs. `N` is the number of evidence categories launched. Don't ask one agent to cover multiple categories.
78
+ Launch all matching investigators and the dependent synthesizer with one `subagent({ action: "execute", input: { async: true, maxSubagentSpawnsPerRun: N + 1, workflowScript } })` call. In `workflowScript`, await the investigators with `runs.all([{ key: "investigate-<category>", agent: "reviewer", task, model }])`, then return `runs.run("synthesize-why", { agent: "reviewer", task, model })` with their outputs. `N` is the number of evidence categories launched. Don't ask one agent to cover multiple categories.
76
79
 
77
80
  Each investigator uses:
78
- - agent: "worker"
81
+ - agent: "reviewer"
79
82
  - `model`: `why investigators` (default inherit-parent)
80
83
  - `task`: instruct the investigator to inspect only
81
84
 
@@ -119,7 +122,7 @@ If your scope assessment suggests a single-commit trivial target where the PR de
119
122
  ## Step 4. Synthesize
120
123
 
121
124
  The same workflow launches `synthesize-why` after every investigator settles. It uses:
122
- - agent: "worker"
125
+ - agent: "reviewer"
123
126
  - `model`: `why synthesizer` (default inherit-parent)
124
127
 
125
128
  The synthesizer gets:
@@ -0,0 +1,2 @@
1
+ policy:
2
+ allow_implicit_invocation: false