@whamp/pi-pstack 0.7.0 → 0.9.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +38 -11
- package/extensions/pstack/index.ts +7 -4
- package/extensions/pstack/pstack-role-prompt.ts +3 -29
- package/package.json +1 -1
- package/skills/architect/SKILL.md +1 -1
- package/skills/arena/SKILL.md +2 -2
- package/skills/automate-me/SKILL.md +2 -2
- package/skills/blast-radius/SKILL.md +3 -3
- package/skills/code-review/LICENSE +21 -0
- package/skills/code-review/SKILL.md +42 -0
- package/skills/code-review/references/code-review-audit.md +91 -0
- package/skills/figure-it-out/SKILL.md +3 -3
- package/skills/how/SKILL.md +8 -5
- package/skills/how/agents/openai.yaml +2 -0
- package/skills/how/references/explorer-prompt.md +1 -1
- package/skills/interrogate/SKILL.md +2 -3
- package/skills/interrogate/references/code-quality-review.md +1 -1
- package/skills/interrogate/references/reviewer-prompt.md +1 -3
- package/skills/interrogate/references/rubric.md +1 -1
- package/skills/poteto-mode/SKILL.md +17 -7
- package/skills/poteto-mode/playbooks/autopilot-full.md +6 -6
- package/skills/poteto-mode/playbooks/autopilot-stack.md +7 -7
- package/skills/poteto-mode/playbooks/babysit.md +1 -1
- package/skills/poteto-mode/playbooks/bug-fix.md +3 -5
- package/skills/poteto-mode/playbooks/eval.md +1 -1
- package/skills/poteto-mode/playbooks/feature.md +3 -3
- package/skills/poteto-mode/playbooks/hillclimb.md +1 -0
- package/skills/poteto-mode/playbooks/multi-phase-plan.md +8 -8
- package/skills/poteto-mode/playbooks/opening-a-pr.md +1 -1
- package/skills/poteto-mode/playbooks/pause-safely.md +1 -1
- package/skills/poteto-mode/playbooks/perf-issue.md +1 -0
- package/skills/poteto-mode/playbooks/refactoring.md +2 -2
- package/skills/poteto-mode/playbooks/session-pickup.md +1 -1
- package/skills/poteto-mode/playbooks/shipping.md +2 -2
- package/skills/principle-guard-the-context-window/SKILL.md +0 -1
- package/skills/principle-never-block-on-the-human/SKILL.md +0 -2
- package/skills/principle-outcome-oriented-execution/SKILL.md +0 -1
- package/skills/principle-prove-it-works/SKILL.md +0 -11
- package/skills/principle-sequence-verifiable-units/SKILL.md +0 -5
- package/skills/recall/SKILL.md +1 -1
- package/skills/reflect/SKILL.md +7 -7
- package/skills/reflect/references/divergent-reviewer.md +1 -1
- package/skills/reflect/references/judgment-reviewer.md +1 -1
- package/skills/reflect/references/tooling-reviewer.md +1 -1
- package/skills/show-me-your-work/SKILL.md +7 -7
- package/skills/show-me-your-work/scripts/log.sh +4 -2
- package/skills/swarm/SKILL.md +4 -4
- package/skills/tdd/SKILL.md +1 -3
- package/skills/technical-writing/SKILL.md +0 -13
- package/skills/typescript-best-practices/SKILL.md +1 -0
- package/skills/typescript-best-practices/agents/openai.yaml +2 -0
- package/skills/unslop/SKILL.md +1 -1
- package/skills/unslop/agents/openai.yaml +2 -0
- package/skills/why/SKILL.md +6 -3
- package/skills/why/agents/openai.yaml +2 -0
|
@@ -21,7 +21,7 @@ Copy `references/decision-log-template.tsv` (the header row) to start a clean lo
|
|
|
21
21
|
- **evidence.** A link or path that proves it: commit SHA, PR number, `file:line`, or an artifact, trace, or screenshot path. Never a paragraph.
|
|
22
22
|
- **result.** The outcome or predicate state: `tests green`, `reverted`, `pixel-diff 0`, `INCONCLUSIVE`, `open`.
|
|
23
23
|
|
|
24
|
-
An example, plain-spoken so a reviewer reads it at a glance.
|
|
24
|
+
An example, plain-spoken so a reviewer reads it at a glance.
|
|
25
25
|
|
|
26
26
|
```
|
|
27
27
|
ts phase decision why evidence result
|
|
@@ -39,6 +39,8 @@ Use the helper `scripts/log.sh <logfile> <phase> <decision> <why> <evidence> <re
|
|
|
39
39
|
|
|
40
40
|
Log decision points and checkpoints, not every action: a fork chosen, a unit completed with its verification result, a pivot or revert with its trigger, a blocker surfaced, a gate fixed. For loop runs, one row per iteration. Skip the trivial and self-evident.
|
|
41
41
|
|
|
42
|
+
A run is one agent conversation, including its later turns and any summary of it. A pickup, a replacement agent, or a new chat starts a new run. When a run adds to a log that already has rows, its first row has phase `start`, and so does its first row after another run's `start` row. So a run that comes back to a log in a later turn first reads the log's last rows to see whether another run wrote since. A `start` row names the `ts` range of the rows before it that this run did not write, and its evidence names this run, such as its agent id. Use phase `start` for nothing else.
|
|
43
|
+
|
|
42
44
|
## Where it lives
|
|
43
45
|
|
|
44
46
|
By default the log is a working artifact, not committed. Keep it at `decisions.tsv` in the work dir, or `.audit/<task-slug>.tsv` when several efforts run at once, and leave it out of git.
|
|
@@ -47,20 +49,18 @@ Commit it only when the work is ambitious enough that a reviewer needs the trail
|
|
|
47
49
|
|
|
48
50
|
## Rules
|
|
49
51
|
|
|
50
|
-
- One row is one decision or checkpoint.
|
|
51
52
|
- Append-only. A wrong call gets a new row that supersedes it. Never edit or delete history.
|
|
52
53
|
- Prefer evidence produced by committed scripts over hand-made one-offs (the **encode-lessons-in-structure** principle skill).
|
|
53
54
|
|
|
54
55
|
## Audit the log against the transcript
|
|
55
56
|
|
|
56
|
-
At the end of the run, before handing back, check the log told the truth. Read this run's transcript
|
|
57
|
+
At the end of the run, before handing back, check the log told the truth. Read this run's transcript. Prefer `$PI_SESSION_FILE`. Otherwise use `~/.pi/agent/sessions/--<slug>--/` (`<slug>` is the workspace path with the leading slash dropped and each "/" turned into "-"). Stay inside that directory. Do not glob sibling slugs under `~/.pi/agent/sessions/`. That reads unrelated private chats. Walk this run's rows against what actually happened. Each stretch of them begins at one of this run's `start` rows, or at the first row if this run created the log, and ends at the next `start` row of another run:
|
|
57
58
|
|
|
58
|
-
-
|
|
59
|
-
-
|
|
59
|
+
- Check that every row maps to a real decision or action.
|
|
60
|
+
- Check that each row's evidence resolves and shows what the row claims.
|
|
60
61
|
- A fork, pivot, or abandoned approach that shaped the work but isn't logged is a gap. Add it.
|
|
61
|
-
- Drop padding.
|
|
62
62
|
|
|
63
|
-
|
|
63
|
+
Correct the log, not the story. The audit never edits or removes a row, even an invented one. When a row records neither a real decision nor a real action, or its claim or evidence is wrong, add a row that supersedes it with what actually happened and a pointer that resolves. This audit does not check rows outside this run's stretches. If this run's own work shows one of them is wrong, supersede it like any wrong call.
|
|
64
64
|
|
|
65
65
|
## Cross-model review of the trail
|
|
66
66
|
|
|
@@ -16,8 +16,10 @@ if [ -n "$logdir" ] && [ "$logdir" != "." ] && [ ! -d "$logdir" ]; then
|
|
|
16
16
|
mkdir -p "$logdir"
|
|
17
17
|
fi
|
|
18
18
|
|
|
19
|
-
|
|
20
|
-
|
|
19
|
+
# Use `>>` here, never `>`. A network mount can fail this test for a log
|
|
20
|
+
# that exists. Then the cost is one stray header line, not the rows.
|
|
21
|
+
if [ ! -s "$logfile" ]; then
|
|
22
|
+
printf 'ts\tphase\tdecision\twhy\tevidence\tresult\n' >> "$logfile"
|
|
21
23
|
fi
|
|
22
24
|
|
|
23
25
|
ts="$(date -u +%Y-%m-%dT%H:%M:%SZ)"
|
package/skills/swarm/SKILL.md
CHANGED
|
@@ -22,8 +22,8 @@ Open a todolist with one entry per phase before launching anything.
|
|
|
22
22
|
1. State the done predicate and the artifact or report the swarm must return.
|
|
23
23
|
2. Choose the shape. Partition into slices, race N workers on identical briefs, or mix both. For a race or mixed shape, declare `first pass`, `rank all`, or `best-of` before spawning.
|
|
24
24
|
3. Set N from the user or derive it from the shape. N is total workers, not the cloud concurrency limit.
|
|
25
|
-
4. Pick the worker model from `swarm workers` in `~/.pi/agent/pstack/models.json` when present. Otherwise use inherit-parent. For a model race, name each arm's model up front.
|
|
26
|
-
5. Give each worker its own writable output when it writes.
|
|
25
|
+
4. Pick the worker model from `swarm workers` in `~/.pi/agent/pstack/models.json` when present. Otherwise use inherit-parent. For `auto` or `inherit-parent`, omit `model`. If an explicit selector is unavailable, inspect `subagent({ action: "models", input: {} })`, pick the closest available model (prefer the highest-reasoning tier of the same family), and relaunch. Never treat `inherit-parent` or `auto` as broken selectors. For a model race, name each arm's model up front.
|
|
26
|
+
5. Give each worker its own writable output when it writes. When workers verify or measure commits, each brief names the exact SHAs. A measurement brief also names the method (sample count, what one sample is, order). The worker records both in its result.
|
|
27
27
|
|
|
28
28
|
## Phase B: Fan out
|
|
29
29
|
|
|
@@ -31,13 +31,13 @@ Launch all N workers with one `subagent({ action: "execute", input: { async: tru
|
|
|
31
31
|
|
|
32
32
|
Set `input.cwd` to an existing checkout. To create a managed checkout from a Git ref, set `input.worktree: true` and `input.baseRef`.
|
|
33
33
|
|
|
34
|
-
Every brief stands alone. Include the goal, scope, exact slice or race arm, how to verify, and what to report. Reports use `PASS`, `ISSUES`, or `BLOCKED` with evidence.
|
|
34
|
+
Every brief stands alone. Include the goal, scope, exact slice or race arm, how to verify, and what to report. Reports use `PASS`, `ISSUES`, or `BLOCKED` with evidence. A worker that can prove a defect reports `ISSUES` and lists every issue it can prove, not only the first.
|
|
35
35
|
|
|
36
36
|
If a worker drops out, proceed with N-1 and note it.
|
|
37
37
|
|
|
38
38
|
## Phase C: Aggregate
|
|
39
39
|
|
|
40
|
-
Read the terminal results. For coverage, every required slice needs a result. For a race, apply the selection rule declared up front. Use first pass, rank all, or best-of. Do not paste raw worker dumps.
|
|
40
|
+
Read the terminal results. Drop a result that does not record the SHAs and method its brief names, and rerun that worker once. After a second miss, record a gap. A gap does not count as a pass. For coverage, every required slice needs a result. For a race, apply the selection rule declared up front. Use first pass, rank all, or best-of. Do not paste raw worker dumps.
|
|
41
41
|
|
|
42
42
|
Keep a compact result table, one-line evidenced issues, and explicit gaps or dropouts.
|
|
43
43
|
|
package/skills/tdd/SKILL.md
CHANGED
|
@@ -18,11 +18,10 @@ Do not force a test when it would be impractical. If the available test would re
|
|
|
18
18
|
4. **Run the new test before fixing.** Confirm it fails for the intended reason. If it passes or fails for an unrelated reason, correct the test or reproduction before editing the implementation.
|
|
19
19
|
5. **Fix the bug.** Make the smallest production change that satisfies the intended behavior while preserving nearby contracts.
|
|
20
20
|
6. **Rerun the regression test.** Confirm the test now passes.
|
|
21
|
-
7. **Run nearby validation.** Run relevant adjacent tests, type checks, lint, or scenario checks when the change has broader risk.
|
|
22
21
|
|
|
23
22
|
## If a Failing Test Is Impractical
|
|
24
23
|
|
|
25
|
-
|
|
24
|
+
Use the closest executable regression check instead: a targeted script, manual reproduction command, browser automation, snapshot comparison, log assertion, or focused integration check.
|
|
26
25
|
|
|
27
26
|
Prefer no new test over a bad test. A bad test is one that mostly tests mocks, encodes current implementation details, depends on timing or unrelated global state, needs expensive infrastructure for a small fix, or would be deleted immediately after proving the fix.
|
|
28
27
|
|
|
@@ -31,7 +30,6 @@ Prefer no new test over a bad test. A bad test is one that mostly tests mocks, e
|
|
|
31
30
|
- Do not change tests merely to match a wrong implementation.
|
|
32
31
|
- Do not weaken existing assertions unless the expected behavior has genuinely changed and the reason is clear.
|
|
33
32
|
- Keep the regression test focused on the bug. Avoid broad fixture churn or unrelated coverage expansion.
|
|
34
|
-
- Do not add tests when the practical signal is weak. Use manual or scripted verification and say why.
|
|
35
33
|
- If the bug is flaky, make the test deterministic where possible and document the signal being locked down.
|
|
36
34
|
- If the bug exposes a broader class of failures, first land the focused regression path, then consider additional sibling coverage.
|
|
37
35
|
|
|
@@ -112,16 +112,3 @@ Before:
|
|
|
112
112
|
After:
|
|
113
113
|
|
|
114
114
|
> `budget.mjs` reads the committed budget from `budget.json` and counts the files that import protos. If the count exceeds the budget, CI fails. Run `budget.mjs --write` only to lower the budget.
|
|
115
|
-
|
|
116
|
-
## Review checklist
|
|
117
|
-
|
|
118
|
-
Apply to any prose this skill covers. Item 1 applies only to document sets:
|
|
119
|
-
|
|
120
|
-
1. Is each file one Diátaxis mode, with links where modes meet?
|
|
121
|
-
2. Is every instruction written as a command, with its condition in front?
|
|
122
|
-
3. Does any sentence carry two instructions or two thoughts? Split it.
|
|
123
|
-
4. Can any word be cut without losing meaning? Cut it.
|
|
124
|
-
5. Is "only" next to the word it changes? Does every "it" point at one thing? Does every clause keep its verb?
|
|
125
|
-
6. Does each thing have exactly one name across the docs?
|
|
126
|
-
7. Would a developer say these words out loud? Replace invented metaphors and fancy synonyms with the plain word or the real symbol name.
|
|
127
|
-
8. Are all symbols, paths, and counts real at this commit, with the commands that regenerate the counts?
|
package/skills/unslop/SKILL.md
CHANGED
|
@@ -1,6 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: unslop
|
|
3
3
|
description: Cut AI tells from any writing. Must always apply.
|
|
4
|
+
disable-model-invocation: true
|
|
4
5
|
---
|
|
5
6
|
|
|
6
7
|
# Unslop
|
|
@@ -11,7 +12,6 @@ Edit text to remove AI patterns.
|
|
|
11
12
|
|
|
12
13
|
1. Scan for the patterns below.
|
|
13
14
|
2. Rewrite. Preserve meaning, match intended tone.
|
|
14
|
-
3. Self-audit: "What makes this obviously AI generated?" Fix remaining tells.
|
|
15
15
|
|
|
16
16
|
## Patterns to detect and fix
|
|
17
17
|
|
package/skills/why/SKILL.md
CHANGED
|
@@ -1,6 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: why
|
|
3
3
|
description: "Use for 'why does X work this way', 'why we picked Y', design rationale, regressions, postmortems, or data-backed thresholds. Discovers available MCPs and queries each evidence category (source control, issue tracker, long-form docs, real-time chat, infrastructure observability, error tracking, product analytics warehouse) in parallel, then returns a cited read on decisions and tradeoffs. Use how for runtime behavior."
|
|
4
|
+
disable-model-invocation: true
|
|
4
5
|
---
|
|
5
6
|
|
|
6
7
|
# Why
|
|
@@ -9,6 +10,8 @@ Investigate the motivation and intent behind code.
|
|
|
9
10
|
|
|
10
11
|
Companion to the `how` skill. `how` answers what the code does and how it works. `why` answers what forces led to its shape.
|
|
11
12
|
|
|
13
|
+
Each child names a role in `~/.pi/agent/pstack/models.json`. Use that role's selector. Omit `model` when the value is `inherit-parent` or `auto`. If an explicit selector is unavailable, inspect `subagent({ action: "models", input: {} })`, pick the closest available model (prefer the highest-reasoning tier of the same family), and relaunch. Never treat `inherit-parent` or `auto` as broken selectors.
|
|
14
|
+
|
|
12
15
|
## Operating Posture
|
|
13
16
|
|
|
14
17
|
Operate as a **careful, cautious, and precise investigator**. Be honest about what you know vs what you're inferring. Read `references/epistemics.md` for the full confidence framework and phrasing guide. The synthesizer must follow it.
|
|
@@ -72,10 +75,10 @@ Source control is always available through git and `gh`. For the other six, clas
|
|
|
72
75
|
|
|
73
76
|
Aim for a complete **coverage map**, not a minimal one. Document the null, don't skip the search. The parent queries each available MCP and builds one bounded evidence packet per category before launching children. A child does not inherit ambient MCP or extension tools. Use a custom agent for a child-side lookup only when that agent explicitly lists the tool and loads its provider through `extensions` or `subagentOnlyExtensions`.
|
|
74
77
|
|
|
75
|
-
Launch all matching investigators and the dependent synthesizer with one `subagent({ action: "execute", input: { async: true, maxSubagentSpawnsPerRun: N + 1, workflowScript } })` call. In `workflowScript`, await the investigators with `runs.all([{ key: "investigate-<category>", agent: "
|
|
78
|
+
Launch all matching investigators and the dependent synthesizer with one `subagent({ action: "execute", input: { async: true, maxSubagentSpawnsPerRun: N + 1, workflowScript } })` call. In `workflowScript`, await the investigators with `runs.all([{ key: "investigate-<category>", agent: "reviewer", task, model }])`, then return `runs.run("synthesize-why", { agent: "reviewer", task, model })` with their outputs. `N` is the number of evidence categories launched. Don't ask one agent to cover multiple categories.
|
|
76
79
|
|
|
77
80
|
Each investigator uses:
|
|
78
|
-
- agent: "
|
|
81
|
+
- agent: "reviewer"
|
|
79
82
|
- `model`: `why investigators` (default inherit-parent)
|
|
80
83
|
- `task`: instruct the investigator to inspect only
|
|
81
84
|
|
|
@@ -119,7 +122,7 @@ If your scope assessment suggests a single-commit trivial target where the PR de
|
|
|
119
122
|
## Step 4. Synthesize
|
|
120
123
|
|
|
121
124
|
The same workflow launches `synthesize-why` after every investigator settles. It uses:
|
|
122
|
-
- agent: "
|
|
125
|
+
- agent: "reviewer"
|
|
123
126
|
- `model`: `why synthesizer` (default inherit-parent)
|
|
124
127
|
|
|
125
128
|
The synthesizer gets:
|