@jiroamato/pstack 0.0.0-stage → 0.15.15
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/LICENSE +21 -0
- package/README.md +78 -2
- package/bin/pstack.js +95 -0
- package/lib/install.js +103 -0
- package/lib/prompt.js +77 -0
- package/lib/targets.js +43 -0
- package/package.json +38 -5
- package/pstack/.claude-plugin/plugin.json +26 -0
- package/pstack/.codex-plugin/plugin.json +36 -0
- package/pstack/LICENSE +21 -0
- package/pstack/LICENSE-cursor-team-kit +21 -0
- package/pstack/NOTICE +8 -0
- package/pstack/README.md +300 -0
- package/pstack/agents/comment-sicko.md +34 -0
- package/pstack/agents/poteto-agent.md +10 -0
- package/pstack/automations/benny/FOR_AGENTS.md +92 -0
- package/pstack/automations/benny/README.md +28 -0
- package/pstack/automations/benny/skills/reproduce-and-fix-issues/SKILL.md +313 -0
- package/pstack/automations/benny/skills/reproduce-and-fix-issues/references/control-adapter.md +169 -0
- package/pstack/automations/benny/skills/reproduce-and-fix-issues/references/feature-map.example.md +205 -0
- package/pstack/automations/benny/skills/reproduce-and-fix-issues/references/verify-existing-fix.md +93 -0
- package/pstack/automations/benny/skills/setup-benny/SKILL.md +271 -0
- package/pstack/automations/benny/skills/triage-issue-reports/SKILL.md +240 -0
- package/pstack/automations/benny/skills/triage-issue-reports/references/routing.example.md +61 -0
- package/pstack/automations/benny/templates/configuration.example.yaml +84 -0
- package/pstack/automations/benny/templates/reproduce-automation-prompt.md +33 -0
- package/pstack/automations/benny/templates/triage-automation-prompt.md +39 -0
- package/pstack/codex/agents/comment-sicko.toml +36 -0
- package/pstack/codex/agents/poteto-agent.toml +11 -0
- package/pstack/docs/guide/01-setup.md +80 -0
- package/pstack/docs/guide/02-poteto-mode.md +131 -0
- package/pstack/docs/guide/03-understand.md +79 -0
- package/pstack/docs/guide/04-design.md +133 -0
- package/pstack/docs/guide/05-build-and-clean.md +83 -0
- package/pstack/docs/guide/06-verify-and-ship.md +130 -0
- package/pstack/docs/guide/07-overnight.md +120 -0
- package/pstack/docs/guide/08-principles.md +72 -0
- package/pstack/docs/guide/09-make-it-yours.md +100 -0
- package/pstack/docs/guide/10-recipes-and-pitfalls.md +156 -0
- package/pstack/docs/guide/README.md +38 -0
- package/pstack/skills/architect/SKILL.md +85 -0
- package/pstack/skills/architect/agents/openai.yaml +2 -0
- package/pstack/skills/architect/references/design-red-flags.md +57 -0
- package/pstack/skills/architect/references/rationale-template.md +35 -0
- package/pstack/skills/architect/references/runner-prompt.md +20 -0
- package/pstack/skills/arena/SKILL.md +75 -0
- package/pstack/skills/arena/agents/openai.yaml +2 -0
- package/pstack/skills/automate-me/SKILL.md +104 -0
- package/pstack/skills/automate-me/agents/openai.yaml +2 -0
- package/pstack/skills/benchmark-checklist/SKILL.md +39 -0
- package/pstack/skills/benchmark-checklist/agents/openai.yaml +2 -0
- package/pstack/skills/blast-radius/SKILL.md +52 -0
- package/pstack/skills/blast-radius/agents/openai.yaml +2 -0
- package/pstack/skills/bro/SKILL.md +7 -0
- package/pstack/skills/bro/agents/openai.yaml +2 -0
- package/pstack/skills/control-cli/SKILL.md +55 -0
- package/pstack/skills/control-cli/agents/openai.yaml +2 -0
- package/pstack/skills/control-ui/SKILL.md +72 -0
- package/pstack/skills/control-ui/agents/openai.yaml +2 -0
- package/pstack/skills/correct/SKILL.md +34 -0
- package/pstack/skills/correct/agents/openai.yaml +2 -0
- package/pstack/skills/create-verification-skill/SKILL.md +47 -0
- package/pstack/skills/create-verification-skill/agents/openai.yaml +2 -0
- package/pstack/skills/create-verification-skill/references/feature-map-example/README.md +47 -0
- package/pstack/skills/create-verification-skill/references/feature-map-example/create-note.md +39 -0
- package/pstack/skills/create-verification-skill/references/feature-map-example/search.md +45 -0
- package/pstack/skills/deslop/SKILL.md +30 -0
- package/pstack/skills/deslop/agents/openai.yaml +2 -0
- package/pstack/skills/figure-it-out/SKILL.md +55 -0
- package/pstack/skills/figure-it-out/agents/openai.yaml +2 -0
- package/pstack/skills/how/SKILL.md +58 -0
- package/pstack/skills/how/agents/openai.yaml +2 -0
- package/pstack/skills/how/references/explainer-prompt.md +55 -0
- package/pstack/skills/how/references/explorer-prompt.md +52 -0
- package/pstack/skills/interrogate/SKILL.md +111 -0
- package/pstack/skills/interrogate/agents/openai.yaml +2 -0
- package/pstack/skills/interrogate/references/code-quality-review.md +47 -0
- package/pstack/skills/interrogate/references/lead-judgment.md +58 -0
- package/pstack/skills/interrogate/references/reviewer-prompt.md +70 -0
- package/pstack/skills/interrogate/references/rubric.md +77 -0
- package/pstack/skills/maintain-verification-skill/SKILL.md +41 -0
- package/pstack/skills/maintain-verification-skill/agents/openai.yaml +2 -0
- package/pstack/skills/make-bot-ui/SKILL.md +289 -0
- package/pstack/skills/make-bot-ui/agents/openai.yaml +2 -0
- package/pstack/skills/no-comments/SKILL.md +24 -0
- package/pstack/skills/no-comments/agents/openai.yaml +2 -0
- package/pstack/skills/poteto-help/SKILL.md +156 -0
- package/pstack/skills/poteto-help/agents/openai.yaml +2 -0
- package/pstack/skills/poteto-help/references/prompting.md +51 -0
- package/pstack/skills/poteto-help/references/recipes.md +47 -0
- package/pstack/skills/poteto-mode/SKILL.md +143 -0
- package/pstack/skills/poteto-mode/agents/openai.yaml +2 -0
- package/pstack/skills/poteto-mode/playbooks/authoring-a-skill.md +12 -0
- package/pstack/skills/poteto-mode/playbooks/autonomous-run.md +13 -0
- package/pstack/skills/poteto-mode/playbooks/autopilot-full.md +13 -0
- package/pstack/skills/poteto-mode/playbooks/autopilot-stack.md +16 -0
- package/pstack/skills/poteto-mode/playbooks/babysit.md +29 -0
- package/pstack/skills/poteto-mode/playbooks/bug-fix.md +15 -0
- package/pstack/skills/poteto-mode/playbooks/eval.md +25 -0
- package/pstack/skills/poteto-mode/playbooks/feature.md +21 -0
- package/pstack/skills/poteto-mode/playbooks/hillclimb.md +21 -0
- package/pstack/skills/poteto-mode/playbooks/investigation.md +14 -0
- package/pstack/skills/poteto-mode/playbooks/multi-phase-plan.md +155 -0
- package/pstack/skills/poteto-mode/playbooks/opening-a-pr.md +38 -0
- package/pstack/skills/poteto-mode/playbooks/orchestrate.md +114 -0
- package/pstack/skills/poteto-mode/playbooks/pause-safely.md +10 -0
- package/pstack/skills/poteto-mode/playbooks/perf-issue.md +25 -0
- package/pstack/skills/poteto-mode/playbooks/prototype.md +14 -0
- package/pstack/skills/poteto-mode/playbooks/refactoring.md +16 -0
- package/pstack/skills/poteto-mode/playbooks/runtime-forensics.md +11 -0
- package/pstack/skills/poteto-mode/playbooks/session-pickup.md +11 -0
- package/pstack/skills/poteto-mode/playbooks/shipping.md +17 -0
- package/pstack/skills/poteto-mode/playbooks/trace-forensics.md +14 -0
- package/pstack/skills/poteto-mode/playbooks/visual-parity.md +11 -0
- package/pstack/skills/poteto-mode/playbooks/worktree-cleanup.md +14 -0
- package/pstack/skills/poteto-mode/references/bugbot-triage.md +142 -0
- package/pstack/skills/poteto-mode/scripts/bootstrap.ts +62 -0
- package/pstack/skills/poteto-mode/scripts/bun.lock +67 -0
- package/pstack/skills/poteto-mode/scripts/check-plan.mjs +185 -0
- package/pstack/skills/poteto-mode/scripts/orch/orch.test.ts +634 -0
- package/pstack/skills/poteto-mode/scripts/orch/orch.ts +578 -0
- package/pstack/skills/poteto-mode/scripts/orch/store.ts +1607 -0
- package/pstack/skills/poteto-mode/scripts/package.json +16 -0
- package/pstack/skills/poteto-mode/scripts/watch-pr/cli.test.ts +224 -0
- package/pstack/skills/poteto-mode/scripts/watch-pr/cli.ts +223 -0
- package/pstack/skills/poteto-mode/scripts/watch-pr/fakes.test-helper.ts +118 -0
- package/pstack/skills/poteto-mode/scripts/watch-pr/github.test.ts +306 -0
- package/pstack/skills/poteto-mode/scripts/watch-pr/github.ts +699 -0
- package/pstack/skills/poteto-mode/scripts/watch-pr/policy.test.ts +420 -0
- package/pstack/skills/poteto-mode/scripts/watch-pr/policy.ts +832 -0
- package/pstack/skills/poteto-mode/scripts/watch-pr/render.ts +169 -0
- package/pstack/skills/poteto-mode/scripts/watch-pr/tsconfig.json +13 -0
- package/pstack/skills/poteto-mode/scripts/watch-pr/types.compile.ts +93 -0
- package/pstack/skills/poteto-mode/scripts/watch-pr/types.ts +401 -0
- package/pstack/skills/poteto-mode/scripts/watch-pr/watch-pr +6 -0
- package/pstack/skills/poteto-mode/scripts/worktree-audit.sh +92 -0
- package/pstack/skills/principle-attack-the-premise/SKILL.md +23 -0
- package/pstack/skills/principle-attack-the-premise/agents/openai.yaml +2 -0
- package/pstack/skills/principle-boundary-discipline/SKILL.md +34 -0
- package/pstack/skills/principle-boundary-discipline/agents/openai.yaml +2 -0
- package/pstack/skills/principle-build-the-lever/SKILL.md +23 -0
- package/pstack/skills/principle-build-the-lever/agents/openai.yaml +2 -0
- package/pstack/skills/principle-encode-lessons-in-structure/SKILL.md +31 -0
- package/pstack/skills/principle-encode-lessons-in-structure/agents/openai.yaml +2 -0
- package/pstack/skills/principle-exhaust-the-design-space/SKILL.md +21 -0
- package/pstack/skills/principle-exhaust-the-design-space/agents/openai.yaml +2 -0
- package/pstack/skills/principle-experience-first/SKILL.md +19 -0
- package/pstack/skills/principle-experience-first/agents/openai.yaml +2 -0
- package/pstack/skills/principle-explain-the-number/SKILL.md +23 -0
- package/pstack/skills/principle-explain-the-number/agents/openai.yaml +2 -0
- package/pstack/skills/principle-fix-root-causes/SKILL.md +23 -0
- package/pstack/skills/principle-fix-root-causes/agents/openai.yaml +2 -0
- package/pstack/skills/principle-foundational-thinking/SKILL.md +21 -0
- package/pstack/skills/principle-foundational-thinking/agents/openai.yaml +2 -0
- package/pstack/skills/principle-guard-the-context-window/SKILL.md +16 -0
- package/pstack/skills/principle-guard-the-context-window/agents/openai.yaml +2 -0
- package/pstack/skills/principle-laziness-protocol/SKILL.md +18 -0
- package/pstack/skills/principle-laziness-protocol/agents/openai.yaml +2 -0
- package/pstack/skills/principle-make-operations-idempotent/SKILL.md +24 -0
- package/pstack/skills/principle-make-operations-idempotent/agents/openai.yaml +2 -0
- package/pstack/skills/principle-migrate-callers-then-delete-legacy-apis/SKILL.md +22 -0
- package/pstack/skills/principle-migrate-callers-then-delete-legacy-apis/agents/openai.yaml +2 -0
- package/pstack/skills/principle-minimize-reader-load/SKILL.md +23 -0
- package/pstack/skills/principle-minimize-reader-load/agents/openai.yaml +2 -0
- package/pstack/skills/principle-model-the-domain/SKILL.md +26 -0
- package/pstack/skills/principle-model-the-domain/agents/openai.yaml +2 -0
- package/pstack/skills/principle-never-block-on-the-human/SKILL.md +20 -0
- package/pstack/skills/principle-never-block-on-the-human/agents/openai.yaml +2 -0
- package/pstack/skills/principle-outcome-oriented-execution/SKILL.md +21 -0
- package/pstack/skills/principle-outcome-oriented-execution/agents/openai.yaml +2 -0
- package/pstack/skills/principle-prove-it-works/SKILL.md +22 -0
- package/pstack/skills/principle-prove-it-works/agents/openai.yaml +2 -0
- package/pstack/skills/principle-redesign-from-first-principles/SKILL.md +16 -0
- package/pstack/skills/principle-redesign-from-first-principles/agents/openai.yaml +2 -0
- package/pstack/skills/principle-separate-before-serializing-shared-state/SKILL.md +16 -0
- package/pstack/skills/principle-separate-before-serializing-shared-state/agents/openai.yaml +2 -0
- package/pstack/skills/principle-sequence-verifiable-units/SKILL.md +17 -0
- package/pstack/skills/principle-sequence-verifiable-units/agents/openai.yaml +2 -0
- package/pstack/skills/principle-subtract-before-you-add/SKILL.md +21 -0
- package/pstack/skills/principle-subtract-before-you-add/agents/openai.yaml +2 -0
- package/pstack/skills/principle-test-behavior-not-implementation/SKILL.md +25 -0
- package/pstack/skills/principle-test-behavior-not-implementation/agents/openai.yaml +2 -0
- package/pstack/skills/principle-type-system-discipline/SKILL.md +31 -0
- package/pstack/skills/principle-type-system-discipline/agents/openai.yaml +2 -0
- package/pstack/skills/pstack-harness/SKILL.md +67 -0
- package/pstack/skills/recall/SKILL.md +35 -0
- package/pstack/skills/recall/agents/openai.yaml +2 -0
- package/pstack/skills/reflect/SKILL.md +76 -0
- package/pstack/skills/reflect/agents/openai.yaml +2 -0
- package/pstack/skills/reflect/references/divergent-reviewer.md +43 -0
- package/pstack/skills/reflect/references/judgment-reviewer.md +42 -0
- package/pstack/skills/reflect/references/synthesizer.md +56 -0
- package/pstack/skills/reflect/references/tooling-reviewer.md +55 -0
- package/pstack/skills/setup-pstack/SKILL.md +110 -0
- package/pstack/skills/show-me-your-work/SKILL.md +82 -0
- package/pstack/skills/show-me-your-work/agents/openai.yaml +2 -0
- package/pstack/skills/show-me-your-work/references/decision-log-template.tsv +1 -0
- package/pstack/skills/show-me-your-work/scripts/log.sh +42 -0
- package/pstack/skills/swarm/SKILL.md +48 -0
- package/pstack/skills/swarm/agents/openai.yaml +2 -0
- package/pstack/skills/tdd/SKILL.md +44 -0
- package/pstack/skills/tdd/agents/openai.yaml +2 -0
- package/pstack/skills/teach/SKILL.md +21 -0
- package/pstack/skills/teach/agents/openai.yaml +2 -0
- package/pstack/skills/technical-writing/SKILL.md +106 -0
- package/pstack/skills/technical-writing/agents/openai.yaml +2 -0
- package/pstack/skills/typescript-best-practices/SKILL.md +31 -0
- package/pstack/skills/typescript-best-practices/agents/openai.yaml +2 -0
- package/pstack/skills/typescript-best-practices/references/patterns.md +324 -0
- package/pstack/skills/unslop/SKILL.md +67 -0
- package/pstack/skills/unslop/agents/openai.yaml +2 -0
- package/pstack/skills/why/SKILL.md +158 -0
- package/pstack/skills/why/agents/openai.yaml +2 -0
- package/pstack/skills/why/references/epistemics.md +144 -0
- package/pstack/skills/why/references/investigator-prompt.md +103 -0
- package/pstack/skills/why/references/source-playbook.md +17 -0
- package/pstack/skills/why/references/sources/code-archaeology.md +88 -0
- package/pstack/skills/why/references/sources/databricks.md +70 -0
- package/pstack/skills/why/references/sources/datadog.md +99 -0
- package/pstack/skills/why/references/sources/incident-postmortem.md +15 -0
- package/pstack/skills/why/references/sources/linear.md +48 -0
- package/pstack/skills/why/references/sources/notion.md +55 -0
- package/pstack/skills/why/references/sources/sentry.md +100 -0
- package/pstack/skills/why/references/sources/slack.md +54 -0
- package/pstack/skills/why/references/synthesizer-prompt.md +135 -0
|
@@ -0,0 +1,156 @@
|
|
|
1
|
+
# Recipes and pitfalls
|
|
2
|
+
|
|
3
|
+
Prompts worth copying, then the mistakes everyone makes once. Swap in your own paths and finish conditions. The recipes are deliberately informal. That's how they get typed in practice, and the skills read intent fine.
|
|
4
|
+
|
|
5
|
+

|
|
6
|
+
|
|
7
|
+
## Understand an unfamiliar subsystem
|
|
8
|
+
|
|
9
|
+
```text
|
|
10
|
+
use /how first to understand how this initialization works. then use /why to figure out why it broke recently.
|
|
11
|
+
```
|
|
12
|
+
|
|
13
|
+
Mechanics first, history second. Each skill's report tells you which sources it searched, so you know what the answer is grounded in.
|
|
14
|
+
|
|
15
|
+
## Restate a noisy report before touching code
|
|
16
|
+
|
|
17
|
+
```text
|
|
18
|
+
/poteto-mode read this thread. restate the underlying issue in your own words, in plain english. don't change any code yet.
|
|
19
|
+
```
|
|
20
|
+
|
|
21
|
+
A misreading shows up in the restatement, where it costs one message to correct. Keep your own theory to yourself until the agent has stated its own.
|
|
22
|
+
|
|
23
|
+
## Prototype before you pick
|
|
24
|
+
|
|
25
|
+
```text
|
|
26
|
+
/poteto-mode prototype a few options for the settings layout. put them behind a switcher and send me screenshots of each.
|
|
27
|
+
```
|
|
28
|
+
|
|
29
|
+
You pick from things that run, not from descriptions. The agent answers its own layout and timing questions along the way.
|
|
30
|
+
|
|
31
|
+
## Turn a settled design into a plan
|
|
32
|
+
|
|
33
|
+
```text
|
|
34
|
+
/poteto-mode turn this design into a plan. small verifiable PRs, each with its own proof.
|
|
35
|
+
```
|
|
36
|
+
|
|
37
|
+
Ask only after the design settles. The plan is the deliverable, and it names the playbook that will execute it.
|
|
38
|
+
|
|
39
|
+
## Get a second opinion on a design
|
|
40
|
+
|
|
41
|
+
```text
|
|
42
|
+
ask /arena for a second opinion on this thread and our approach
|
|
43
|
+
```
|
|
44
|
+
|
|
45
|
+
Your current design becomes one candidate among several, and the synthesis tells you whether the panel found something better or confirmed what you had. Cheap insurance before a costly commitment.
|
|
46
|
+
|
|
47
|
+
## Check independent slices in parallel
|
|
48
|
+
|
|
49
|
+
```text
|
|
50
|
+
/swarm check every package under packages/ against its check.sh. one worker per package. one report.
|
|
51
|
+
```
|
|
52
|
+
|
|
53
|
+
Each worker owns one package. The parent waits for every slice and returns one `PASS`, `ISSUES`, or `BLOCKED` report instead of raw worker dumps.
|
|
54
|
+
|
|
55
|
+
## Review a branch skeptically
|
|
56
|
+
|
|
57
|
+
```text
|
|
58
|
+
/interrogate the whole branch, but skeptically. don't change anything yet. no nitpicks unless it's an actual bug or regression in behavior.
|
|
59
|
+
```
|
|
60
|
+
|
|
61
|
+
The qualifiers do real work. "don't change anything yet" keeps it read-only, and the nitpick rule pre-filters the noise so `Act on` findings are worth your time.
|
|
62
|
+
|
|
63
|
+
## Fix a bug through a failing test
|
|
64
|
+
|
|
65
|
+
```text
|
|
66
|
+
/poteto-mode repro the duplicate write first. if there's a cheap test path, /tdd it. then fix and rerun.
|
|
67
|
+
```
|
|
68
|
+
|
|
69
|
+
"if there's a cheap test path" matters. Forcing a test through brittle mocks proves less than running the real command, and the playbook is allowed to say so.
|
|
70
|
+
|
|
71
|
+
## Repro and fix a report with proof
|
|
72
|
+
|
|
73
|
+
```text
|
|
74
|
+
/poteto-mode repro this with /verify-<app>. if it repros on main, fix it and show me a video as proof.
|
|
75
|
+
```
|
|
76
|
+
|
|
77
|
+
"if it repros on main" lets the run stop early when the bug is already gone. The video lets you check the fix before you read the diff.
|
|
78
|
+
|
|
79
|
+
## Vet a number before you post it
|
|
80
|
+
|
|
81
|
+
```text
|
|
82
|
+
/benchmark-checklist vet this 40% speedup before it goes in the pr description
|
|
83
|
+
```
|
|
84
|
+
|
|
85
|
+
You get faster, slower, no measurable difference, or inconclusive, with the run count, the range, and what limits the number.
|
|
86
|
+
|
|
87
|
+
## Stop correcting the same mistake
|
|
88
|
+
|
|
89
|
+
```text
|
|
90
|
+
/correct agents keep adding new config flags without registering them in the schema
|
|
91
|
+
```
|
|
92
|
+
|
|
93
|
+
The fix lands in the repo as architecture, a type, a lint, or a test, so the next agent can't make the mistake.
|
|
94
|
+
|
|
95
|
+
## Ask how without starting the work
|
|
96
|
+
|
|
97
|
+
```text
|
|
98
|
+
/poteto-help how do i get poteto-mode to stay on every turn?
|
|
99
|
+
```
|
|
100
|
+
|
|
101
|
+
You get an answer, a prompt to send, and a link to the source. Nothing runs until you send that prompt.
|
|
102
|
+
|
|
103
|
+
## Keep a run honest while you're away
|
|
104
|
+
|
|
105
|
+
```text
|
|
106
|
+
im going to bed, keep going autonomously until every fixture passes. do not stop. keep a decision log i can audit in the morning.
|
|
107
|
+
```
|
|
108
|
+
|
|
109
|
+
The full contract is on the [overnight page](./07-overnight.md). The short form works once the task and finish condition are already in the conversation.
|
|
110
|
+
|
|
111
|
+
## Redirect a drifting run
|
|
112
|
+
|
|
113
|
+
Steering prompts are one line:
|
|
114
|
+
|
|
115
|
+
```text
|
|
116
|
+
i said the goal is to repro. i did not ask for a fix yet.
|
|
117
|
+
```
|
|
118
|
+
|
|
119
|
+
```text
|
|
120
|
+
apply prove it works. show me the real output, not the build log.
|
|
121
|
+
```
|
|
122
|
+
|
|
123
|
+
```text
|
|
124
|
+
/unslop that, no emdashes
|
|
125
|
+
```
|
|
126
|
+
|
|
127
|
+
You rarely need more words. You need the right name, and [the principles page](./08-principles.md) is the vocabulary.
|
|
128
|
+
|
|
129
|
+
## Get the reply in plain words
|
|
130
|
+
|
|
131
|
+
```text
|
|
132
|
+
/bro
|
|
133
|
+
```
|
|
134
|
+
|
|
135
|
+
That's the whole prompt. [`/bro`](../../skills/bro/SKILL.md) restates the last message like one human talking to another, no jargon, shorter. Use it when a reply is technically thorough and you still don't know what it said.
|
|
136
|
+
|
|
137
|
+
## The pitfalls
|
|
138
|
+
|
|
139
|
+
- **Enumerating skills in the prompt.** "use /how then /architect then /arena" reorders steps the playbook already sequences. State the goal and constraints. Name a skill only to override a default.
|
|
140
|
+
- **A vague finish condition.** "make it better" gives the loop nothing to check. Give a command or artifact that can pass or fail.
|
|
141
|
+
- **Leading with your theory of the cause.** The agent searches wherever you pointed. Ask it to restate the problem first, then share your hunch.
|
|
142
|
+
- **Taking the first design.** One attempt locks in the first shape the model thought of. Ask for prototypes or `/architect` and pick from evidence.
|
|
143
|
+
- **Polishing an abstract plan.** Adversarial review of a plan with no code behind it invents risks that will never happen. Settle the open questions with prototypes, then review what got built.
|
|
144
|
+
- **Parallel agents in one worktree.** They overwrite each other and the diff becomes archaeology. Give each its own worktree (harness **isolation** row), or say "own worktree per attempt".
|
|
145
|
+
- **Looping before you trust the loop.** A loop that can't verify its own work only makes unchecked work faster. Get the verification skill working first.
|
|
146
|
+
- **Trusting an unvetted number.** A warm cache or a skipped code path can fake a speedup. Run `/benchmark-checklist` before the number goes anywhere.
|
|
147
|
+
- **Correcting the same mistake by hand.** A correction in chat helps one run. `/correct` fixes the repo so no later run repeats it.
|
|
148
|
+
- **Using `/arena` for coverage.** `/arena` repeats one design or code brief, then picks a base and grafts the best parts. `/swarm` partitions slices or declared race arms and aggregates one report.
|
|
149
|
+
- **Accepting every review comment.** Bots and humans both file real catches and noise in one list. `/interrogate` sorts findings into act-on and dismissed buckets with reasons, and you can override either way.
|
|
150
|
+
- **Treating `auto` as a model slug.** `auto` and `inherit-parent` mean "omit the model field so the subagent inherits the parent chat model." [Setup](./01-setup.md) covers the roles.
|
|
151
|
+
- **Reporting success off a green build.** A build proves it compiles. Ask for the real command, flow, stored value, or profile, and expect the evidence in the reply.
|
|
152
|
+
- **Writing a `SKILL.md` freehand.** Route it through the [Authoring or modifying a skill playbook](../../skills/poteto-mode/playbooks/authoring-a-skill.md) so validation and review happen.
|
|
153
|
+
|
|
154
|
+
That's the guide. If you skipped ahead, go back to [setup](./01-setup.md) and run one real task. The habits stick from use, not from reading.
|
|
155
|
+
|
|
156
|
+
Back to the [guide index](./README.md).
|
|
@@ -0,0 +1,38 @@
|
|
|
1
|
+
# The pstack guide
|
|
2
|
+
|
|
3
|
+
pstack works best when you stop micromanaging the agent. You describe what you want and how you'll know it's done. `/poteto-mode` picks the playbook, runs the other skills as the steps need them, and shows you the evidence. This guide teaches that habit with realistic prompts.
|
|
4
|
+
|
|
5
|
+
Here's what you'll learn:
|
|
6
|
+
|
|
7
|
+
1. [Set up pstack](./01-setup.md). Install the plugin and pick your models.
|
|
8
|
+
2. [Route work through `/poteto-mode`](./02-poteto-mode.md). Give it a goal and watch it pick a playbook.
|
|
9
|
+
3. [Understand the code](./03-understand.md). A read-only investigation, then `/how`, `/why`, `/teach`, and `/recall` before you edit anything.
|
|
10
|
+
4. [Design the change](./04-design.md). `/architect`, `/arena`, `/swarm`, `/interrogate`, prototypes, and plans before code locks in a shape.
|
|
11
|
+
5. [Build and clean the change](./05-build-and-clean.md). The build playbooks, `/tdd`, `/unslop`, and `/no-comments`.
|
|
12
|
+
6. [Verify and ship](./06-verify-and-ship.md). Prove behavior on the real app, vet numbers with `/benchmark-checklist`, then open a focused PR and drive it to merged.
|
|
13
|
+
7. [Run work while you sleep](./07-overnight.md). Trust before loops, an overnight contract, a decision log you can audit, and coordinator sessions and automations that scale past one agent.
|
|
14
|
+
8. [Steer with principle names](./08-principles.md). The 24 names that redirect an agent mid-task.
|
|
15
|
+
9. [Make it yours](./09-make-it-yours.md). Your own mode, `/correct` for repeated mistakes, and how to test a skill change.
|
|
16
|
+
10. [Recipes and pitfalls](./10-recipes-and-pitfalls.md). Prompts to copy and mistakes to skip.
|
|
17
|
+
|
|
18
|
+
Read the pages in order the first time. After that, each page stands alone.
|
|
19
|
+
|
|
20
|
+
When you're stuck, or can't tell which skill fits, type [`/poteto-help`](../../skills/poteto-help/SKILL.md) with your question:
|
|
21
|
+
|
|
22
|
+
```text
|
|
23
|
+
/poteto-help which skill should i use to review this branch?
|
|
24
|
+
```
|
|
25
|
+
|
|
26
|
+
It answers, hands you a prompt to send, and links the skill or guide page the answer came from. It doesn't start the work, because a pstack run spends real tokens, so you send the prompt when you're ready. It runs only when you type it.
|
|
27
|
+
|
|
28
|
+
## If you only remember one thing
|
|
29
|
+
|
|
30
|
+
Give the agent a goal and a way to check it, in your own words:
|
|
31
|
+
|
|
32
|
+
```text
|
|
33
|
+
/poteto-mode the export writes duplicate rows when a retry lands mid-run. repro first, then fix and verify.
|
|
34
|
+
```
|
|
35
|
+
|
|
36
|
+
You don't need to name a playbook or list skills. "repro first" and a checkable outcome are all the routing signal `/poteto-mode` needs. It matches the Bug fix playbook, copies the steps into a todo list, and calls the right skills as each step fires.
|
|
37
|
+
|
|
38
|
+
Next: [Set up pstack](./01-setup.md).
|
|
@@ -0,0 +1,85 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: architect
|
|
3
|
+
description: "Sketch types, signatures, and module structure before code, then stay in the loop while implementation fills in. Use for /architect, 'architect this', 'design this', or non-trivial work where jumping to code would lock in the wrong shape."
|
|
4
|
+
disable-model-invocation: true
|
|
5
|
+
---
|
|
6
|
+
|
|
7
|
+
# Architect
|
|
8
|
+
|
|
9
|
+
Design before implementing. Sketch types, function signatures, class shapes, and module boundaries with `not implemented` bodies and pseudocode. Synthesize across multiple model perspectives, then fill in code against the chosen sketch. If implementation proves the sketch wrong, throw it out and redesign.
|
|
10
|
+
|
|
11
|
+
## Start
|
|
12
|
+
|
|
13
|
+
Open a todolist with one entry per phase before starting.
|
|
14
|
+
|
|
15
|
+
1. Ground
|
|
16
|
+
2. Sketch
|
|
17
|
+
3. Agree
|
|
18
|
+
4. Implement
|
|
19
|
+
5. Scrap
|
|
20
|
+
|
|
21
|
+
## Phase A: Ground the problem
|
|
22
|
+
|
|
23
|
+
Build a real mental model of every system the new code touches. Run the **how** skill over the relevant subsystems.
|
|
24
|
+
|
|
25
|
+
Naming a file isn't grounding. Produce the traced model `how` prescribes. If the design redefines ownership or layering, also run the **why** skill on the existing shape so the rationale becomes a constraint, not a guess.
|
|
26
|
+
|
|
27
|
+
Skip Phase A only when the work is genuinely greenfield with no surrounding system to integrate.
|
|
28
|
+
|
|
29
|
+
## Phase B: Sketch
|
|
30
|
+
|
|
31
|
+
Run the **arena** skill with the design-sketch task and the Phase A grounding artifacts. Pass `references/runner-prompt.md` as each runner's prompt. Each candidate produces a design package shaped per `references/rationale-template.md`.
|
|
32
|
+
|
|
33
|
+
Take the runners from the `architect runners` line in the pstack models file (`~/.pstack/models.md`, harness **models file** row), in place of the `arena runners` line. If the file or that line is missing, use the judgment tier default and the code tier default (harness **tiers** row). Alias and rejected entries follow the runner rules in the **arena** skill's Phase A.
|
|
34
|
+
|
|
35
|
+
Design it twice. Require at least two structurally distinct candidates before synthesis, even when the first looks sufficient. This is the **exhaust-the-design-space** principle skill made concrete. Whole-shape alternatives, not point fixes inside one shape.
|
|
36
|
+
|
|
37
|
+
Screen every candidate against [`references/design-red-flags.md`](references/design-red-flags.md) before synthesis. Assume the next contributor is an agent that sees only the files it opened, copies the nearest example, and takes the shortest path that compiles. Prefer the design where a change that looks right from one file is right for the whole repo.
|
|
38
|
+
|
|
39
|
+
Compare viable candidates on interface depth. Prefer the design that hides more complexity behind a smaller, simpler public surface. A rich interface can keep call chains short by concentrating capability instead of scattering it across layers.
|
|
40
|
+
|
|
41
|
+
Arena returns one synthesized design package. The synthesis decision populates the rationale's "Synthesis decision" section.
|
|
42
|
+
|
|
43
|
+
## Phase C: Agree (opt-in)
|
|
44
|
+
|
|
45
|
+
Default: proceed directly to implementation with the synthesized design. No human checkpoint.
|
|
46
|
+
|
|
47
|
+
Opt in to a checkpoint when the invoker explicitly asks: "/architect with checkpoint," "stop and show me before implementing," or similar. Then surface the synthesized design and pause for sign-off.
|
|
48
|
+
|
|
49
|
+
The synthesis can ship as its own commit either way, as the "scaffold first" mode of the **foundational-thinking** principle skill. Planned and scoped breakage during fill-in is fine, per the **outcome-oriented-execution** principle skill. For adversarial pressure on the design before implementing, run the **interrogate** skill on the synthesized sketch.
|
|
50
|
+
|
|
51
|
+
If the human pushes back on the shape (in a checkpoint or after the fact), treat that as Phase A evidence. Re-ground and re-run Phase B before writing more code.
|
|
52
|
+
|
|
53
|
+
## Phase D: Implement against the sketch
|
|
54
|
+
|
|
55
|
+
Replace `not implemented` bodies with code, pseudocode with logic. The synthesized sketch is the contract.
|
|
56
|
+
|
|
57
|
+
For code you keep, load [`deslop`](../deslop/SKILL.md) before each commit and the final handoff. For web, IDE, or Electron behavior, load [`control-ui`](../control-ui/SKILL.md) and verify the user flow; for CLI/TUI behavior, load [`control-cli`](../control-cli/SKILL.md). Read [`pstack-harness`](../pstack-harness/SKILL.md) for invocation and use a project `verify-<app>` skill for app-specific steps. Put these calls in implementation delegates' briefs too.
|
|
58
|
+
|
|
59
|
+
Deviations from the sketch are signal worth surfacing, not friction to absorb silently. If a function needs a parameter the sketch didn't anticipate, ask whether the sketch was wrong, the requirement was missed, or the implementation is overreaching.
|
|
60
|
+
|
|
61
|
+
## Phase E: Scrap when the architecture is wrong
|
|
62
|
+
|
|
63
|
+
If implementation keeps producing friction the sketch can't absorb, throw the sketch out. Don't bolt fixes onto a wrong design, per the **redesign-from-first-principles** and **fix-root-causes** principle skills.
|
|
64
|
+
|
|
65
|
+
The signal is a *pattern*, not single instances. Tells:
|
|
66
|
+
|
|
67
|
+
- The same shape of workaround appearing repeatedly across unrelated code.
|
|
68
|
+
- Multiple unrelated edge cases that all need special-case branches.
|
|
69
|
+
- Types that need escape hatches (`any`, casts, optional fields always set in practice) to compile.
|
|
70
|
+
- The "we need a lock" reflex when the sketch said the state wasn't shared.
|
|
71
|
+
- Callers having to know the abstraction's internal rules to use it.
|
|
72
|
+
- Two or more independent Phase D deviations of the same shape across the implementation.
|
|
73
|
+
|
|
74
|
+
Use judgment. A few edge cases don't condemn an architecture. Some problems are legitimately complex. Complexity in the data is not complexity in the design.
|
|
75
|
+
|
|
76
|
+
When you scrap:
|
|
77
|
+
|
|
78
|
+
1. Re-run the **how** skill over what's been built.
|
|
79
|
+
2. Redesign as if the new constraints had been day-one assumptions, per redesign-from-first-principles.
|
|
80
|
+
3. Subtract before adding, per the **subtract-before-you-add** principle skill. The new sketch should be smaller than the old one before it grows.
|
|
81
|
+
4. Return to Phase B and re-run arena.
|
|
82
|
+
|
|
83
|
+
## Outputs
|
|
84
|
+
|
|
85
|
+
The caller's usage is written first and the type sketch derived from it. One file with new types and signatures for small changes. Module map plus type definitions for larger work. The rationale ships alongside, shaped per `references/rationale-template.md`, including the usage sketch and the synthesis decision.
|
|
@@ -0,0 +1,57 @@
|
|
|
1
|
+
# Design red flags
|
|
2
|
+
|
|
3
|
+
Screen every candidate before synthesis. A red flag is a reason to revise or reject the shape.
|
|
4
|
+
|
|
5
|
+
## Shallow module
|
|
6
|
+
|
|
7
|
+
A shallow module exposes a large interface while hiding little complexity. Judge depth by the capability and policy hidden behind the public surface relative to the size of that surface. Prefer a simple interface backed by substantial behavior.
|
|
8
|
+
|
|
9
|
+
Do not confuse a deep module with a deep call chain. A deep call chain scatters understanding across layers. A deep module concentrates capability behind one interface.
|
|
10
|
+
|
|
11
|
+
Look for these signs:
|
|
12
|
+
|
|
13
|
+
- Callers coordinate several methods to complete one operation.
|
|
14
|
+
- Public options expose internal stages or implementation choices.
|
|
15
|
+
- Learning the interface does not save the caller from learning the implementation.
|
|
16
|
+
|
|
17
|
+
## Information leakage
|
|
18
|
+
|
|
19
|
+
Information leakage makes multiple modules depend on the same internal decision. A representation, policy, or protocol detail appears in more than one place, so changing it requires coordinated edits.
|
|
20
|
+
|
|
21
|
+
Public re-exports of transport or wire types are leakage. Parse external data into domain types behind the interface. Keep storage schemas, framework objects, and protocol details private.
|
|
22
|
+
|
|
23
|
+
## Temporal decomposition
|
|
24
|
+
|
|
25
|
+
Temporal decomposition organizes modules by execution order instead of the knowledge they own. Separate load, validate, transform, and save stages often repeat one representation and its invariants across several boundaries.
|
|
26
|
+
|
|
27
|
+
Group code around domain knowledge and ownership. Methods that run at different times can still belong to one module when they protect the same decisions.
|
|
28
|
+
|
|
29
|
+
## Pass-through method
|
|
30
|
+
|
|
31
|
+
A pass-through method forwards the same arguments to another method with the same shape. It adds a layer without hiding complexity.
|
|
32
|
+
|
|
33
|
+
Remove it or move responsibility to the module that can complete the operation. Keep a forwarding boundary only when it adds policy, adaptation, or a distinct abstraction.
|
|
34
|
+
|
|
35
|
+
## Split ownership
|
|
36
|
+
|
|
37
|
+
More than one module writes the same state or keeps its own copy of it. An agent that edits one writer can't see the others, so their rules diverge.
|
|
38
|
+
|
|
39
|
+
Give each piece of state one owner. Other modules read it or ask the owner to change it.
|
|
40
|
+
|
|
41
|
+
## Two ways to do one task
|
|
42
|
+
|
|
43
|
+
The design supports more than one way to do the same task. An agent copies whichever way it finds first, so every way keeps gaining callers.
|
|
44
|
+
|
|
45
|
+
Keep one way. Move callers off the others and delete them in the same change.
|
|
46
|
+
|
|
47
|
+
## Importable internals
|
|
48
|
+
|
|
49
|
+
A caller can import a module's internals. An agent takes the shortest path that compiles, so it imports them directly and they become part of the interface.
|
|
50
|
+
|
|
51
|
+
Make internals unreachable from outside the module, so an import from outside fails the build.
|
|
52
|
+
|
|
53
|
+
## Hand-synced list
|
|
54
|
+
|
|
55
|
+
Two or more places list the same items, and adding an item means editing every list. An agent that sees one list updates only that one.
|
|
56
|
+
|
|
57
|
+
Keep one list and derive the others from it. If a list can't be derived, make the build fail when the lists disagree.
|
|
@@ -0,0 +1,35 @@
|
|
|
1
|
+
# Rationale template
|
|
2
|
+
|
|
3
|
+
The prose that ships alongside the type sketch. One page. Sentence-case headings, no boilerplate. Replace the italic notes with actual content.
|
|
4
|
+
|
|
5
|
+
## Problem
|
|
6
|
+
|
|
7
|
+
*One paragraph. What we're trying to do, and what about the existing system or constraints makes the shape non-obvious. If [Phase A](../SKILL.md#phase-a-ground-the-problem) surfaced constraints the design must honor (existing types to interop with, callers we can't break, invariants that crossed our boundary), name them here so the reader sees the same constraints you saw.*
|
|
8
|
+
|
|
9
|
+
## Usage (caller's view)
|
|
10
|
+
|
|
11
|
+
*Write this first, before the type sketch. Show the README or quickstart the consumer reads, plus two or three realistic call sites in their own code. What they import, what they call, what comes back. The type sketch in [Shape](#shape) is derived from this. The two must agree. When they diverge, reconcile the sketch to the usage, not the reverse. The caller's experience is the spec. The types serve it.*
|
|
12
|
+
|
|
13
|
+
## Shape
|
|
14
|
+
|
|
15
|
+
*The recommended architecture. Data structures first. Then how data flows through the signatures. Name the load-bearing decisions. State which invariants are encoded in types, where validation lives, and what the system deliberately does not do. Judge interface depth explicitly. State what complexity the public surface hides, what remains exposed to callers, and why the interface is no larger than needed. Cite the principle behind each decision (e.g., `per boundary-discipline`). Don't restate it.*
|
|
16
|
+
|
|
17
|
+
## Synthesis decision
|
|
18
|
+
|
|
19
|
+
*Filled in by [arena](../../arena/SKILL.md). Records which candidate became the base and why, what was adapted from each of the others, and what was rejected and why.*
|
|
20
|
+
|
|
21
|
+
## Tradeoffs accepted
|
|
22
|
+
|
|
23
|
+
*One bullet per tradeoff the chosen shape makes. Form: "we accept X in exchange for Y." Name anything a future reader might mistake for an oversight, including things that look like premature optimization or premature simplification.*
|
|
24
|
+
|
|
25
|
+
## Alternatives considered
|
|
26
|
+
|
|
27
|
+
*Required. Name at least one concrete alternative shape, with one line on why it lost. Judge each alternative on interface depth, not implementation simplicity alone. Name the complexity it exposes to callers and the complexity it hides. Two or three alternatives belong here when the design space had real contenders. One is fine when the constraints forced the answer, with the conclusion phrased as "this was the only viable shape because..." Avoid listing flavors of the same shape. This section covers design alternatives the chosen shape considered and rejected, not other runner candidates.*
|
|
28
|
+
|
|
29
|
+
## Open questions and risks
|
|
30
|
+
|
|
31
|
+
*Things you noticed during the sketch that the human needs to weigh in on, and risks worth flagging before implementation starts. Phrase as questions, not assertions, so the human's answer is the resolution rather than a comment.*
|
|
32
|
+
|
|
33
|
+
## Next implementation step
|
|
34
|
+
|
|
35
|
+
*The first thing to build against the sketch. One sentence. What you'd start writing immediately after synthesis (or after Phase D sign-off, if a checkpoint was opted into).*
|
|
@@ -0,0 +1,20 @@
|
|
|
1
|
+
# Architect runner prompt
|
|
2
|
+
|
|
3
|
+
The orchestrator passes this file through to every parallel candidate runner during Phase B and fills in the variable inputs around it: the task, the Phase A grounding artifacts, the isolated working directory, and the path to write outputs. The working directory is a git worktree when available, otherwise a per-runner subdirectory under the sketch dir. What matters is independence between candidates.
|
|
4
|
+
|
|
5
|
+
You are producing one candidate design in architect's parallel exploration. Read the **architect** skill in full first. That's the workflow you're inside. Output a candidate design package: type sketch, function signatures, module map, and prose rationale shaped per [`rationale-template.md`](rationale-template.md).
|
|
6
|
+
|
|
7
|
+
Apply the following discipline. The orchestrator compares candidates on these axes to pick a base.
|
|
8
|
+
|
|
9
|
+
- Caller's usage first. Write the README-style usage and two or three real call sites before the types, then derive the type sketch from them. The usage is the spec. The two must agree, so reconcile the sketch to the usage, not the reverse.
|
|
10
|
+
- Data structures first. Get the core types right and the code becomes obvious. Trace each dominant access pattern through the proposed structure. If the answer is "we'll add a map / index / cache later," the structure is wrong.
|
|
11
|
+
- Interface depth. Compare the capability hidden behind the public surface relative to the size of that surface. Prefer a simple interface that pulls complexity into the callee, even when the implementation becomes less simple. Do not put transport or wire types on the public API. Parse into domain types behind the interface.
|
|
12
|
+
- Shared state: if two actors might both write, ask "what happens?" If the answer isn't "nothing," default to per-actor state with a merge at the read boundary, per the **separate-before-serializing-shared-state** principle skill.
|
|
13
|
+
- Make boundaries visible. `not implemented` errors for bodies, `// TODO` pseudocode for tricky logic, doc comments stating intent and invariants. A reader should trace data from input to output by reading types and signatures alone.
|
|
14
|
+
- Encode invariants in types: hard-to-misuse types > runtime checks > prose comments, per the **encode-lessons-in-structure** principle skill.
|
|
15
|
+
- Validate at boundaries, trust types inside, per the **boundary-discipline** principle skill. Business logic as pure functions. The shell stays thin.
|
|
16
|
+
- Single source of truth per invariant. Derive instead of sync.
|
|
17
|
+
- Idempotent state transitions where applicable, per the **make-operations-idempotent** principle skill. Ask what happens if the operation runs twice or crashes halfway.
|
|
18
|
+
- Short call chains. If tracing the flow needs more than three files, flatten the hierarchy, per the **laziness-protocol** and **minimize-reader-load** principle skills.
|
|
19
|
+
|
|
20
|
+
You are one of the parallel runners, each on a different model. Produce the best design your model can make. Don't hedge against the others. Differences between candidates are the signal used to pick a base and graft. Converging on a safe-looking middle defeats the exploration.
|
|
@@ -0,0 +1,75 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: arena
|
|
3
|
+
description: "Spawn N parallel candidates at the same task, pick a base, graft the strongest parts of the losers into it. Use for /arena, 'arena this', 'throw it in the arena', or when one attempt at a non-trivial artifact would lock in the wrong shape."
|
|
4
|
+
disable-model-invocation: true
|
|
5
|
+
---
|
|
6
|
+
|
|
7
|
+
# Arena
|
|
8
|
+
|
|
9
|
+
Fan out N parallel attempts at the same task. Read every candidate end to end. Pick the strongest as the base. Graft the best ideas from the others into it. Verify the synthesized result.
|
|
10
|
+
|
|
11
|
+
## Start
|
|
12
|
+
|
|
13
|
+
Open a todolist with one entry per phase before launching anything.
|
|
14
|
+
|
|
15
|
+
1. Frame
|
|
16
|
+
2. Fan out
|
|
17
|
+
3. Cross-judge
|
|
18
|
+
4. Pick
|
|
19
|
+
5. Graft
|
|
20
|
+
6. Verify
|
|
21
|
+
|
|
22
|
+
## Phase A: Frame
|
|
23
|
+
|
|
24
|
+
The N candidates will receive the same prompt, so the prompt is the contract.
|
|
25
|
+
|
|
26
|
+
1. State the artifact each candidate is producing.
|
|
27
|
+
2. Derive the rubric. State what success looks like for *this* task, then turn it into 3-6 concrete gradeable criteria. The rubric is the picker's tool in Phase D. Candidates only see the task.
|
|
28
|
+
3. Pick the runners. Read the **pstack-harness** skill, then use the `arena runners` line in `~/.pstack/models.md` (harness **models file** row). If the file or that line is missing, default to one each on the judgment tier default and the code tier default (harness **tiers** row). An `auto` or `inherit-parent` entry in this line or the cross-judge line means the parent model, so omit `model` for it. If the spawn rejects a configured entry, run that seat on its family's default and say so. Families follow the harness **families** row. With no family match, use the judgment tier default (harness **tiers** row). If it rejects a default, use the closest valid slug of the same family from its error message. Spawn more when the arena covers multiple design directions. Same model N times when the work is generation-bound rather than judgment-sensitive.
|
|
29
|
+
4. Assign output paths. Each candidate writes to its own location (a git worktree where possible, otherwise `/tmp/arena-<slug>/candidate-<n>/`), per the **separate-before-serializing-shared-state** principle skill.
|
|
30
|
+
|
|
31
|
+
## Phase B: Fan out
|
|
32
|
+
|
|
33
|
+
Spawn all N subagents in one message with `run_in_background: true`, each with the task, the path to the shared grounding, its own output path, and instructions to produce both the artifact and a short rationale.
|
|
34
|
+
|
|
35
|
+
Each rationale names the alternatives the candidate considered and what it rejected.
|
|
36
|
+
|
|
37
|
+
For runnable UI candidates, each brief tells the runner to load [`control-ui`](../control-ui/SKILL.md); for CLI/TUI candidates, [`control-cli`](../control-cli/SKILL.md). Include the project verification skill when one exists. Read the harness **verification harness** row; design-only candidates need no live app run.
|
|
38
|
+
|
|
39
|
+
If a candidate fails to produce output, proceed with N-1 and note the dropout in the synthesis record.
|
|
40
|
+
|
|
41
|
+
## Phase C: Cross-judge
|
|
42
|
+
|
|
43
|
+
After all Phase B candidates complete, choose one model from the `arena cross-judge pool` line in `~/.pstack/models.md`. If the file or that line is missing, choose from the judgment tier default and the code tier default. Prefer a different model family from the parent's (harness **families** row). Spawn one readonly judge subagent on that model. It sees the rubric and the candidates by path label, scores each criterion, and recommends a base with rationale. It runs in parallel with the parent's reading in Phase D, not with the candidates themselves. Don't spawn the judge while candidates are still writing.
|
|
44
|
+
|
|
45
|
+
## Phase D: Pick a base
|
|
46
|
+
|
|
47
|
+
Read every candidate end to end before picking.
|
|
48
|
+
|
|
49
|
+
Score each candidate against the rubric criterion by criterion, not on holistic feel. Compare against the cross-judge. Agreement on the base confirms the pick. Disagreement means one of you is biased or the rubric was ambiguous. Read both rationales before deciding.
|
|
50
|
+
|
|
51
|
+
Pick the base on which candidate a future maintainer can extend most easily without breaking invariants. Prefer the cleaner boundary or smaller API when two feel tied, per the Laziness Protocol.
|
|
52
|
+
|
|
53
|
+
Record the pick and the reason in a short synthesis note alongside the base artifact, including the cross-judge's verdict.
|
|
54
|
+
|
|
55
|
+
## Phase E: Graft
|
|
56
|
+
|
|
57
|
+
Walk each losing candidate once more and identify what is worth porting into the base. The signal is usually one or two things per candidate, not most of it.
|
|
58
|
+
|
|
59
|
+
Fold each graft in by hand, per the **redesign-from-first-principles** principle skill. Don't paste mechanically. The result has to remain coherent under one mental model.
|
|
60
|
+
|
|
61
|
+
Record what was grafted, from which candidate, and what was rejected and why.
|
|
62
|
+
|
|
63
|
+
When N candidates converge on the same shape, that is a strong agreement signal. Note the convergence in the record and ship the consensus shape. No graft is needed. When N candidates wildly diverge, Phase A was under-specified. Reframe and re-run rather than averaging the divergence.
|
|
64
|
+
|
|
65
|
+
## Phase F: Verify
|
|
66
|
+
|
|
67
|
+
The synthesized artifact has to hold up under the same scrutiny as any other output, per the **prove-it-works** principle skill.
|
|
68
|
+
|
|
69
|
+
For production code, load [`deslop`](../deslop/SKILL.md) before committing or handing back the synthesized diff, then rerun affected checks. Verify a UI result through `control-ui`, or a CLI/TUI result through `control-cli`, per the harness **verification harness** row. Throwaway prototypes keep their prototype workflow.
|
|
70
|
+
|
|
71
|
+
If verification surfaces a problem the arena did not catch, either Phase A was wrong (re-frame and re-run) or one candidate caught it and you missed the graft (go back to Phase E). Don't paper over.
|
|
72
|
+
|
|
73
|
+
## Outputs
|
|
74
|
+
|
|
75
|
+
One synthesized artifact. One short synthesis note alongside, naming the base, the grafts (with source candidate), the rejections, the dropouts if any, and the verification result.
|
|
@@ -0,0 +1,104 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: automate-me
|
|
3
|
+
description: "Use for \"automate me\", \"create/update/refresh my -mode skill\", \"turn/capture my preferences or working style into a skill\", or wanting agents to follow how the user works. Drafts or revises a personal -mode skill via create-skill + unslop, optionally pulling fresh evidence from recent transcripts."
|
|
4
|
+
disable-model-invocation: true
|
|
5
|
+
---
|
|
6
|
+
|
|
7
|
+
# Automate me
|
|
8
|
+
|
|
9
|
+
A guided flow for turning the user's working conventions into a skill agents will follow. The output is one `-mode` skill tailored to them (e.g. `jay-mode`, `priya-mode`).
|
|
10
|
+
|
|
11
|
+
This skill orchestrates three others: an inline mining pass (see step 1), `skill-creator` (authoring, harness **skill-creator** row), and the **unslop** skill (prose discipline). It sequences them. It doesn't replace them.
|
|
12
|
+
|
|
13
|
+
## Flow
|
|
14
|
+
|
|
15
|
+
### 0. Check for an existing skill
|
|
16
|
+
|
|
17
|
+
Look recursively for `<project skills>/**/*-mode/SKILL.md` and `<user skills>/*-mode/SKILL.md` matching the user's handle, where the directories come from the harness **project skills** and **user skills** rows. Mode skills can live in a personal category directory (`<project skills>/<handle>/`), not only at the top level. If one exists, confirm intent with the harness **ask** tool (unless they already said "update my skill" or similar):
|
|
18
|
+
|
|
19
|
+
- Update the existing skill (default for repeat runs)
|
|
20
|
+
- Start fresh (rare, ask why before doing it)
|
|
21
|
+
|
|
22
|
+
Update mode changes the rest of the flow:
|
|
23
|
+
- Step 1 mines only history since the skill was last edited (`git log -1 --format=%cI <path>`).
|
|
24
|
+
- Step 2 asks what's changed or missing, not what to capture from zero.
|
|
25
|
+
- Step 4 edits the existing file in place. Preserve sections the user hasn't contradicted. Revise ones with new evidence. Add new sections only for genuinely new rules.
|
|
26
|
+
|
|
27
|
+
### 1. Mine their history
|
|
28
|
+
|
|
29
|
+
Locate the active workspace's transcripts before fanning out. Resolve the active workspace's transcript directory per the harness **transcripts** row and stay inside it. Other workspaces' transcripts are private chats from unrelated projects.
|
|
30
|
+
|
|
31
|
+
Survey recent agent conversations within that scope for recurring patterns. Run multiple parallel subagents across slices of history (e.g. last 2-4 weeks, split into 3 slices so each has enough material). Each slice mining subagent reads transcripts from the workspace-scoped path the parent provides, looks for the signals below, and returns a short structured list of patterns it saw with evidence pointers. Default signals worth hunting:
|
|
32
|
+
|
|
33
|
+
- Response preferences (length, tone, format, "dumb it down" corrections)
|
|
34
|
+
- Delegation habits (subagents, models, specialized workflows, parallelism)
|
|
35
|
+
- Verification posture (what "done" means, unit tests vs live repro, reviewers)
|
|
36
|
+
- Code and prose discipline (style, principles cited, lint/format tools)
|
|
37
|
+
- Process conventions (worktrees, commits, PRs, review/merge tooling)
|
|
38
|
+
- Meta preferences (fixing skills mid-task, proposing new ones)
|
|
39
|
+
|
|
40
|
+
Cross-check across slices before elevating a signal. Patterns seen in 2+ slices are high-confidence. Lone signals are weak and usually get dropped.
|
|
41
|
+
|
|
42
|
+
### 2. Ask the user directly
|
|
43
|
+
|
|
44
|
+
Mining misses intent that hasn't come up yet. Use the harness **ask** tool (structured multi-choice) rather than asking the user to type from scratch.
|
|
45
|
+
|
|
46
|
+
Shape: one or two questions with 4-6 options each, `allow_multiple: true` for category questions. Start broad ("Which areas matter most?"), then follow up on selected areas with specific options. After the structured rounds, one free-form chat question catches anything the options missed.
|
|
47
|
+
|
|
48
|
+
Don't dump 20 questions.
|
|
49
|
+
|
|
50
|
+
### 3. Cluster findings
|
|
51
|
+
|
|
52
|
+
Group the combined signals into sections. Common ones (use only what applies):
|
|
53
|
+
|
|
54
|
+
- **Response style**: length, tone, format.
|
|
55
|
+
- **Autonomy**: how much to do without asking, MCP tool use.
|
|
56
|
+
- **Understand first**: which skills to reach for when scoping or investigating a change.
|
|
57
|
+
- **Subagents**: default, parallelism, model-to-task, specialized workflows.
|
|
58
|
+
- **Prose / code discipline**: principles, lint tools, style guides.
|
|
59
|
+
- **Review and verify**: repro posture, verification skills, live-testing tools.
|
|
60
|
+
- **Process**: git worktrees, commits, PRs, review/merge tooling.
|
|
61
|
+
- **Skills**: skill-authoring habits, fix-the-skill-first, proposing new skills.
|
|
62
|
+
|
|
63
|
+
The **poteto-mode** skill shows the shape. Read it for granularity. Don't copy its content. The user's rules are not the same as poteto-mode's.
|
|
64
|
+
|
|
65
|
+
### 4. Draft the skill
|
|
66
|
+
|
|
67
|
+
Use `skill-creator` (harness **skill-creator** row) to author the skill. Placement:
|
|
68
|
+
|
|
69
|
+
- Path: preserve an existing mode skill's category. For a new mode, use `<project skills>/<handle>/<handle>-mode/SKILL.md` when the repo has an established personal category for that handle. Otherwise default to `<project skills>/<handle>-mode/SKILL.md` in the project (or `<user skills>/<handle>-mode/` if the user prefers a personal skill). The directories come from the harness **project skills** and **user skills** rows.
|
|
70
|
+
- Handle: the user's first name or chosen identifier.
|
|
71
|
+
- Frontmatter `description`: trigger on their name + `/<handle>-mode` + "work in their style", not on generic keywords like "write code" or "review PR".
|
|
72
|
+
- Frontmatter formatting: follow `skill-creator`'s YAML rules. Keep `description` as one YAML scalar. Quote it or use `description: >-` with indented continuation lines when punctuation or wrapping requires it.
|
|
73
|
+
- Frontmatter `disable-model-invocation: true` by default. Opt out only if the user explicitly wants their mode to apply on every turn.
|
|
74
|
+
|
|
75
|
+
### 5. Iterate on prose
|
|
76
|
+
|
|
77
|
+
Apply the **unslop** skill and `skill-creator`'s writing guidelines to every line.
|
|
78
|
+
|
|
79
|
+
Show the draft to the user and take feedback. Expect multiple iterations. Cut ruthlessly. A mode skill is not a manual.
|
|
80
|
+
|
|
81
|
+
### 6. Land it
|
|
82
|
+
|
|
83
|
+
Work in a worktree off main. Commit and open a PR. Don't push to main directly.
|
|
84
|
+
|
|
85
|
+
## Guardrails
|
|
86
|
+
|
|
87
|
+
- **Don't overfit to one conversation.** A preference stated once and contradicted another time is noise. Require multiple instances before codifying it.
|
|
88
|
+
- **Don't be clever.** Restating other skills' contents, inventing metaphors, or writing "poetic" prose for an agent reader is cost without benefit. Keep it operational.
|
|
89
|
+
- **Reference, don't inline.** Other skills the user relies on should appear as path references, not pasted excerpts. Same for any principle docs they maintain elsewhere.
|
|
90
|
+
- **Keep sections minimal.** Only add a section if the user has a specific, non-default rule there. "Communicate clearly" is not a section. "Short paragraphs. Tables when comparing options. Bullets only when items are genuinely parallel." is.
|
|
91
|
+
- **Name conventions generic.** Use "the user" or "the human" in imperatives, not the author's first name.
|
|
92
|
+
- **Don't force symmetry.** If a user has no process rules worth writing down, skip the Process section entirely.
|
|
93
|
+
|
|
94
|
+
## Evaluation
|
|
95
|
+
|
|
96
|
+
A `-mode` skill is subjective output. A `skill-creator`-style test/iterate benchmark loop isn't useful here. Vibe-check with the user: does it read like them? Did it miss anything? Then ship.
|
|
97
|
+
|
|
98
|
+
Run a description-optimization loop only if the skill's trigger accuracy turns out to be a problem in practice.
|
|
99
|
+
|
|
100
|
+
## When not to use
|
|
101
|
+
|
|
102
|
+
- User wants a task-specific skill (not working conventions): `skill-creator` alone, no mining required.
|
|
103
|
+
- User wants to capture one narrow workflow (e.g. "how I write commit messages"). That's a regular skill, not a mode skill.
|
|
104
|
+
|
|
@@ -0,0 +1,39 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: benchmark-checklist
|
|
3
|
+
description: "Vet a perf measurement (limiter, tuning, limits, errors, repeatability, relevance, and whether the work happened) before you report or act on it. Use when you run a benchmark or report a speedup or regression you measured."
|
|
4
|
+
disable-model-invocation: true
|
|
5
|
+
---
|
|
6
|
+
|
|
7
|
+
# Benchmark checklist
|
|
8
|
+
|
|
9
|
+
Use this when you produce a performance number: a PR's before and after, a regression claim, a hillclimb harness, or a library or config choice. [Explain the Number](../principle-explain-the-number/SKILL.md) says why. Answer each question below with evidence from a run, not from a guess about the code.
|
|
10
|
+
|
|
11
|
+
For a quick ballpark the user asked for, one run is enough. Still check questions 4 and 7, and say that it is one run. Skip the rest unless that run looks wrong. A choice between options is never a ballpark.
|
|
12
|
+
|
|
13
|
+
## Before you run anything
|
|
14
|
+
|
|
15
|
+
- Write down the claim you expect to make, in the words you would ship ("export is 30% faster at p50 on the 60k-row dataset"). The questions test that sentence.
|
|
16
|
+
- Read the measurement script. Note what it times, what it counts, and what it ignores.
|
|
17
|
+
- Check the load average with `uptime` and the core count with `nproc`. If the machine is busy, find out what is running. If you cannot stop it, interleave the sides so both see the same noise, and say so in the report.
|
|
18
|
+
|
|
19
|
+
## The questions
|
|
20
|
+
|
|
21
|
+
1. **Why not double?** Name the limiter. Profile in a run you do not report, because profilers and tracers slow the work down. Use CPU per process (`top`, `pidstat`), a profiler for the runtime (`node --cpu-prof`, `py-spy`, `perf`), I/O wait, and syscall counts (`strace -c` on Linux). Then map the hot spot to source. Watch the load generator too. If it saturates first, you measured the load generator. If a change did not move the number, the limiter explains why, so find it before you call the change useless.
|
|
22
|
+
2. **Was it tuned?** Run every side the way production runs it: release builds, production flags and env, batching and transaction settings, connection pools, caches as warm or cold as production sees them, and the same versions and data. If one side runs on defaults, you compared configurations, not implementations. A limiter that is a setting, such as a commit per row, a debug build, or a missing index, means that side is untuned. Tune it and measure again before you pick a winner. If you cannot tune it, do not pick a winner from that run. Narrowing the claim to the code as it ships today does not fix this when the user is choosing what to adopt, because they adopt the option, not today's settings.
|
|
23
|
+
3. **Did it break limits?** Do the arithmetic. Compare bytes per second with disk and network bandwidth, and operations per second times the cost per operation with the cores you have. Compare the time saved with the time the changed piece took. Removing a piece that takes 10% of the run can make the run at most about 11% faster. A result past a limit means the run measured something other than the work, such as a cache, a no-op, or a bug.
|
|
24
|
+
4. **Did it error?** Count failures and non-success responses, and check that the outputs are correct, not just present. Errors behave differently from successes. Rejections are often fast, and timeouts and retries are slow. If the script does not count errors, add the count.
|
|
25
|
+
5. **Does it reproduce?** Run each side at least 5 times, and alternate the sides (A, B, A, B, and so on) so that warmup, lazy initialization, caches, and drift do not favor one side. Report the median and the range. Treat a gap smaller than the run-to-run variation as no measurable difference. When the call is close, use a rank-sum test or the harness's own statistics.
|
|
26
|
+
6. **Does it matter?** Next to any micro result, measure the end-to-end path a user waits on, with realistic data sizes and concurrency. Report the micro result as a share of the whole. A helper that takes 1% of a request can make the request at most 1% faster, however fast the helper gets.
|
|
27
|
+
7. **Did it even happen?** Confirm the work ran inside the timed region. The request reached the server, the rows were written, the bytes were read, and the code used the result. Lazy code (generators nobody iterates, promises nobody awaits, results the JIT can discard) and timeouts all produce numbers for work that never happened.
|
|
28
|
+
|
|
29
|
+
## Report
|
|
30
|
+
|
|
31
|
+
- Lead with the verdict: faster, slower, no measurable difference, or inconclusive.
|
|
32
|
+
- Give the number with its unit, the run count, the range, and the limiter. For example, "p50 41 ms → 33 ms, median of 7 runs per side, range 32 to 35 ms after, bound by JSON parsing on one core."
|
|
33
|
+
- Call the verdict inconclusive when you claim a difference but cannot name the limiter, when a side ran untuned, or when you could not check questions 4 and 7. Name the gap.
|
|
34
|
+
- Keep a PR body to one primary number, per the **Opening a PR** playbook. Put the runs, the range, and the limiter evidence in a linked artifact or a notes file.
|
|
35
|
+
|
|
36
|
+
## How this fits the other perf material
|
|
37
|
+
|
|
38
|
+
- The **Perf issue** playbook finds and fixes slowness, and the performance mantras in its step 2 generate the fixes. This skill vets its baseline before the playbook plans from it, and every number after that.
|
|
39
|
+
- The **Hillclimb** playbook loops on one metric. This skill vets its harness before the harness is frozen. The frozen harness then prints error and work counts, so each keep-or-revert checks questions 4 and 7 for free.
|