@jiroamato/pstack 0.0.0-stage → 0.16.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (229) hide show
  1. package/LICENSE +21 -0
  2. package/README.md +79 -2
  3. package/bin/pstack.js +95 -0
  4. package/lib/install.js +121 -0
  5. package/lib/prompt.js +77 -0
  6. package/lib/targets.js +43 -0
  7. package/package.json +38 -5
  8. package/pstack/.claude-plugin/plugin.json +26 -0
  9. package/pstack/.codex-plugin/plugin.json +36 -0
  10. package/pstack/LICENSE +21 -0
  11. package/pstack/LICENSE-cursor-team-kit +21 -0
  12. package/pstack/NOTICE +8 -0
  13. package/pstack/README.md +307 -0
  14. package/pstack/agents/comment-sicko.md +34 -0
  15. package/pstack/agents/poteto-agent.md +10 -0
  16. package/pstack/automations/benny/FOR_AGENTS.md +92 -0
  17. package/pstack/automations/benny/README.md +28 -0
  18. package/pstack/automations/benny/skills/reproduce-and-fix-issues/SKILL.md +313 -0
  19. package/pstack/automations/benny/skills/reproduce-and-fix-issues/references/control-adapter.md +169 -0
  20. package/pstack/automations/benny/skills/reproduce-and-fix-issues/references/feature-map.example.md +205 -0
  21. package/pstack/automations/benny/skills/reproduce-and-fix-issues/references/verify-existing-fix.md +93 -0
  22. package/pstack/automations/benny/skills/setup-benny/SKILL.md +271 -0
  23. package/pstack/automations/benny/skills/triage-issue-reports/SKILL.md +240 -0
  24. package/pstack/automations/benny/skills/triage-issue-reports/references/routing.example.md +61 -0
  25. package/pstack/automations/benny/templates/configuration.example.yaml +84 -0
  26. package/pstack/automations/benny/templates/reproduce-automation-prompt.md +33 -0
  27. package/pstack/automations/benny/templates/triage-automation-prompt.md +39 -0
  28. package/pstack/codex/agents/comment-sicko.toml +36 -0
  29. package/pstack/codex/agents/poteto-agent.toml +11 -0
  30. package/pstack/docs/guide/01-setup.md +80 -0
  31. package/pstack/docs/guide/02-poteto-mode.md +131 -0
  32. package/pstack/docs/guide/03-understand.md +79 -0
  33. package/pstack/docs/guide/04-design.md +133 -0
  34. package/pstack/docs/guide/05-build-and-clean.md +83 -0
  35. package/pstack/docs/guide/06-verify-and-ship.md +130 -0
  36. package/pstack/docs/guide/07-overnight.md +120 -0
  37. package/pstack/docs/guide/08-principles.md +72 -0
  38. package/pstack/docs/guide/09-make-it-yours.md +100 -0
  39. package/pstack/docs/guide/10-recipes-and-pitfalls.md +156 -0
  40. package/pstack/docs/guide/README.md +38 -0
  41. package/pstack/skills/architect/SKILL.md +85 -0
  42. package/pstack/skills/architect/agents/openai.yaml +2 -0
  43. package/pstack/skills/architect/references/design-red-flags.md +57 -0
  44. package/pstack/skills/architect/references/rationale-template.md +35 -0
  45. package/pstack/skills/architect/references/runner-prompt.md +20 -0
  46. package/pstack/skills/arena/SKILL.md +75 -0
  47. package/pstack/skills/arena/agents/openai.yaml +2 -0
  48. package/pstack/skills/automate-me/SKILL.md +104 -0
  49. package/pstack/skills/automate-me/agents/openai.yaml +2 -0
  50. package/pstack/skills/benchmark-checklist/SKILL.md +39 -0
  51. package/pstack/skills/benchmark-checklist/agents/openai.yaml +2 -0
  52. package/pstack/skills/blast-radius/SKILL.md +52 -0
  53. package/pstack/skills/blast-radius/agents/openai.yaml +2 -0
  54. package/pstack/skills/bro/SKILL.md +7 -0
  55. package/pstack/skills/bro/agents/openai.yaml +2 -0
  56. package/pstack/skills/control-cli/SKILL.md +55 -0
  57. package/pstack/skills/control-cli/agents/openai.yaml +2 -0
  58. package/pstack/skills/control-ui/SKILL.md +72 -0
  59. package/pstack/skills/control-ui/agents/openai.yaml +2 -0
  60. package/pstack/skills/correct/SKILL.md +34 -0
  61. package/pstack/skills/correct/agents/openai.yaml +2 -0
  62. package/pstack/skills/create-verification-skill/SKILL.md +47 -0
  63. package/pstack/skills/create-verification-skill/agents/openai.yaml +2 -0
  64. package/pstack/skills/create-verification-skill/references/feature-map-example/README.md +47 -0
  65. package/pstack/skills/create-verification-skill/references/feature-map-example/create-note.md +39 -0
  66. package/pstack/skills/create-verification-skill/references/feature-map-example/search.md +45 -0
  67. package/pstack/skills/deslop/SKILL.md +30 -0
  68. package/pstack/skills/deslop/agents/openai.yaml +2 -0
  69. package/pstack/skills/figure-it-out/SKILL.md +55 -0
  70. package/pstack/skills/figure-it-out/agents/openai.yaml +2 -0
  71. package/pstack/skills/how/SKILL.md +58 -0
  72. package/pstack/skills/how/agents/openai.yaml +2 -0
  73. package/pstack/skills/how/references/explainer-prompt.md +55 -0
  74. package/pstack/skills/how/references/explorer-prompt.md +52 -0
  75. package/pstack/skills/interrogate/SKILL.md +111 -0
  76. package/pstack/skills/interrogate/agents/openai.yaml +2 -0
  77. package/pstack/skills/interrogate/references/code-quality-review.md +47 -0
  78. package/pstack/skills/interrogate/references/lead-judgment.md +58 -0
  79. package/pstack/skills/interrogate/references/reviewer-prompt.md +70 -0
  80. package/pstack/skills/interrogate/references/rubric.md +77 -0
  81. package/pstack/skills/kiss/SKILL.md +90 -0
  82. package/pstack/skills/kiss/agents/openai.yaml +2 -0
  83. package/pstack/skills/kiss/references/assess.md +110 -0
  84. package/pstack/skills/kiss/references/principles.md +138 -0
  85. package/pstack/skills/maintain-verification-skill/SKILL.md +41 -0
  86. package/pstack/skills/maintain-verification-skill/agents/openai.yaml +2 -0
  87. package/pstack/skills/make-bot-ui/SKILL.md +289 -0
  88. package/pstack/skills/make-bot-ui/agents/openai.yaml +2 -0
  89. package/pstack/skills/no-comments/SKILL.md +24 -0
  90. package/pstack/skills/no-comments/agents/openai.yaml +2 -0
  91. package/pstack/skills/poteto-help/SKILL.md +156 -0
  92. package/pstack/skills/poteto-help/agents/openai.yaml +2 -0
  93. package/pstack/skills/poteto-help/references/prompting.md +51 -0
  94. package/pstack/skills/poteto-help/references/recipes.md +47 -0
  95. package/pstack/skills/poteto-mode/SKILL.md +143 -0
  96. package/pstack/skills/poteto-mode/agents/openai.yaml +2 -0
  97. package/pstack/skills/poteto-mode/playbooks/authoring-a-skill.md +12 -0
  98. package/pstack/skills/poteto-mode/playbooks/autonomous-run.md +13 -0
  99. package/pstack/skills/poteto-mode/playbooks/autopilot-full.md +13 -0
  100. package/pstack/skills/poteto-mode/playbooks/autopilot-stack.md +16 -0
  101. package/pstack/skills/poteto-mode/playbooks/babysit.md +29 -0
  102. package/pstack/skills/poteto-mode/playbooks/bug-fix.md +15 -0
  103. package/pstack/skills/poteto-mode/playbooks/eval.md +25 -0
  104. package/pstack/skills/poteto-mode/playbooks/feature.md +21 -0
  105. package/pstack/skills/poteto-mode/playbooks/hillclimb.md +21 -0
  106. package/pstack/skills/poteto-mode/playbooks/investigation.md +14 -0
  107. package/pstack/skills/poteto-mode/playbooks/multi-phase-plan.md +155 -0
  108. package/pstack/skills/poteto-mode/playbooks/opening-a-pr.md +38 -0
  109. package/pstack/skills/poteto-mode/playbooks/orchestrate.md +114 -0
  110. package/pstack/skills/poteto-mode/playbooks/pause-safely.md +10 -0
  111. package/pstack/skills/poteto-mode/playbooks/perf-issue.md +25 -0
  112. package/pstack/skills/poteto-mode/playbooks/prototype.md +14 -0
  113. package/pstack/skills/poteto-mode/playbooks/refactoring.md +16 -0
  114. package/pstack/skills/poteto-mode/playbooks/runtime-forensics.md +11 -0
  115. package/pstack/skills/poteto-mode/playbooks/session-pickup.md +11 -0
  116. package/pstack/skills/poteto-mode/playbooks/shipping.md +17 -0
  117. package/pstack/skills/poteto-mode/playbooks/trace-forensics.md +14 -0
  118. package/pstack/skills/poteto-mode/playbooks/visual-parity.md +11 -0
  119. package/pstack/skills/poteto-mode/playbooks/worktree-cleanup.md +14 -0
  120. package/pstack/skills/poteto-mode/references/bugbot-triage.md +142 -0
  121. package/pstack/skills/poteto-mode/scripts/bootstrap.ts +62 -0
  122. package/pstack/skills/poteto-mode/scripts/bun.lock +67 -0
  123. package/pstack/skills/poteto-mode/scripts/check-plan.mjs +185 -0
  124. package/pstack/skills/poteto-mode/scripts/orch/orch.test.ts +634 -0
  125. package/pstack/skills/poteto-mode/scripts/orch/orch.ts +578 -0
  126. package/pstack/skills/poteto-mode/scripts/orch/store.ts +1607 -0
  127. package/pstack/skills/poteto-mode/scripts/package.json +16 -0
  128. package/pstack/skills/poteto-mode/scripts/watch-pr/cli.test.ts +224 -0
  129. package/pstack/skills/poteto-mode/scripts/watch-pr/cli.ts +223 -0
  130. package/pstack/skills/poteto-mode/scripts/watch-pr/fakes.test-helper.ts +118 -0
  131. package/pstack/skills/poteto-mode/scripts/watch-pr/github.test.ts +306 -0
  132. package/pstack/skills/poteto-mode/scripts/watch-pr/github.ts +699 -0
  133. package/pstack/skills/poteto-mode/scripts/watch-pr/policy.test.ts +420 -0
  134. package/pstack/skills/poteto-mode/scripts/watch-pr/policy.ts +832 -0
  135. package/pstack/skills/poteto-mode/scripts/watch-pr/render.ts +169 -0
  136. package/pstack/skills/poteto-mode/scripts/watch-pr/tsconfig.json +13 -0
  137. package/pstack/skills/poteto-mode/scripts/watch-pr/types.compile.ts +93 -0
  138. package/pstack/skills/poteto-mode/scripts/watch-pr/types.ts +401 -0
  139. package/pstack/skills/poteto-mode/scripts/watch-pr/watch-pr +6 -0
  140. package/pstack/skills/poteto-mode/scripts/worktree-audit.sh +92 -0
  141. package/pstack/skills/principle-attack-the-premise/SKILL.md +23 -0
  142. package/pstack/skills/principle-attack-the-premise/agents/openai.yaml +2 -0
  143. package/pstack/skills/principle-boundary-discipline/SKILL.md +34 -0
  144. package/pstack/skills/principle-boundary-discipline/agents/openai.yaml +2 -0
  145. package/pstack/skills/principle-build-the-lever/SKILL.md +23 -0
  146. package/pstack/skills/principle-build-the-lever/agents/openai.yaml +2 -0
  147. package/pstack/skills/principle-encode-lessons-in-structure/SKILL.md +31 -0
  148. package/pstack/skills/principle-encode-lessons-in-structure/agents/openai.yaml +2 -0
  149. package/pstack/skills/principle-exhaust-the-design-space/SKILL.md +21 -0
  150. package/pstack/skills/principle-exhaust-the-design-space/agents/openai.yaml +2 -0
  151. package/pstack/skills/principle-experience-first/SKILL.md +19 -0
  152. package/pstack/skills/principle-experience-first/agents/openai.yaml +2 -0
  153. package/pstack/skills/principle-explain-the-number/SKILL.md +23 -0
  154. package/pstack/skills/principle-explain-the-number/agents/openai.yaml +2 -0
  155. package/pstack/skills/principle-fix-root-causes/SKILL.md +23 -0
  156. package/pstack/skills/principle-fix-root-causes/agents/openai.yaml +2 -0
  157. package/pstack/skills/principle-foundational-thinking/SKILL.md +21 -0
  158. package/pstack/skills/principle-foundational-thinking/agents/openai.yaml +2 -0
  159. package/pstack/skills/principle-guard-the-context-window/SKILL.md +16 -0
  160. package/pstack/skills/principle-guard-the-context-window/agents/openai.yaml +2 -0
  161. package/pstack/skills/principle-laziness-protocol/SKILL.md +18 -0
  162. package/pstack/skills/principle-laziness-protocol/agents/openai.yaml +2 -0
  163. package/pstack/skills/principle-make-operations-idempotent/SKILL.md +24 -0
  164. package/pstack/skills/principle-make-operations-idempotent/agents/openai.yaml +2 -0
  165. package/pstack/skills/principle-migrate-callers-then-delete-legacy-apis/SKILL.md +22 -0
  166. package/pstack/skills/principle-migrate-callers-then-delete-legacy-apis/agents/openai.yaml +2 -0
  167. package/pstack/skills/principle-minimize-reader-load/SKILL.md +23 -0
  168. package/pstack/skills/principle-minimize-reader-load/agents/openai.yaml +2 -0
  169. package/pstack/skills/principle-model-the-domain/SKILL.md +26 -0
  170. package/pstack/skills/principle-model-the-domain/agents/openai.yaml +2 -0
  171. package/pstack/skills/principle-never-block-on-the-human/SKILL.md +20 -0
  172. package/pstack/skills/principle-never-block-on-the-human/agents/openai.yaml +2 -0
  173. package/pstack/skills/principle-outcome-oriented-execution/SKILL.md +21 -0
  174. package/pstack/skills/principle-outcome-oriented-execution/agents/openai.yaml +2 -0
  175. package/pstack/skills/principle-prove-it-works/SKILL.md +22 -0
  176. package/pstack/skills/principle-prove-it-works/agents/openai.yaml +2 -0
  177. package/pstack/skills/principle-redesign-from-first-principles/SKILL.md +16 -0
  178. package/pstack/skills/principle-redesign-from-first-principles/agents/openai.yaml +2 -0
  179. package/pstack/skills/principle-separate-before-serializing-shared-state/SKILL.md +16 -0
  180. package/pstack/skills/principle-separate-before-serializing-shared-state/agents/openai.yaml +2 -0
  181. package/pstack/skills/principle-sequence-verifiable-units/SKILL.md +17 -0
  182. package/pstack/skills/principle-sequence-verifiable-units/agents/openai.yaml +2 -0
  183. package/pstack/skills/principle-subtract-before-you-add/SKILL.md +21 -0
  184. package/pstack/skills/principle-subtract-before-you-add/agents/openai.yaml +2 -0
  185. package/pstack/skills/principle-test-behavior-not-implementation/SKILL.md +25 -0
  186. package/pstack/skills/principle-test-behavior-not-implementation/agents/openai.yaml +2 -0
  187. package/pstack/skills/principle-type-system-discipline/SKILL.md +31 -0
  188. package/pstack/skills/principle-type-system-discipline/agents/openai.yaml +2 -0
  189. package/pstack/skills/pstack-harness/SKILL.md +67 -0
  190. package/pstack/skills/recall/SKILL.md +35 -0
  191. package/pstack/skills/recall/agents/openai.yaml +2 -0
  192. package/pstack/skills/reflect/SKILL.md +76 -0
  193. package/pstack/skills/reflect/agents/openai.yaml +2 -0
  194. package/pstack/skills/reflect/references/divergent-reviewer.md +43 -0
  195. package/pstack/skills/reflect/references/judgment-reviewer.md +42 -0
  196. package/pstack/skills/reflect/references/synthesizer.md +56 -0
  197. package/pstack/skills/reflect/references/tooling-reviewer.md +55 -0
  198. package/pstack/skills/setup-pstack/SKILL.md +110 -0
  199. package/pstack/skills/show-me-your-work/SKILL.md +82 -0
  200. package/pstack/skills/show-me-your-work/agents/openai.yaml +2 -0
  201. package/pstack/skills/show-me-your-work/references/decision-log-template.tsv +1 -0
  202. package/pstack/skills/show-me-your-work/scripts/log.sh +42 -0
  203. package/pstack/skills/swarm/SKILL.md +48 -0
  204. package/pstack/skills/swarm/agents/openai.yaml +2 -0
  205. package/pstack/skills/tdd/SKILL.md +44 -0
  206. package/pstack/skills/tdd/agents/openai.yaml +2 -0
  207. package/pstack/skills/teach/SKILL.md +21 -0
  208. package/pstack/skills/teach/agents/openai.yaml +2 -0
  209. package/pstack/skills/technical-writing/SKILL.md +106 -0
  210. package/pstack/skills/technical-writing/agents/openai.yaml +2 -0
  211. package/pstack/skills/typescript-best-practices/SKILL.md +31 -0
  212. package/pstack/skills/typescript-best-practices/agents/openai.yaml +2 -0
  213. package/pstack/skills/typescript-best-practices/references/patterns.md +324 -0
  214. package/pstack/skills/unslop/SKILL.md +67 -0
  215. package/pstack/skills/unslop/agents/openai.yaml +2 -0
  216. package/pstack/skills/why/SKILL.md +158 -0
  217. package/pstack/skills/why/agents/openai.yaml +2 -0
  218. package/pstack/skills/why/references/epistemics.md +144 -0
  219. package/pstack/skills/why/references/investigator-prompt.md +103 -0
  220. package/pstack/skills/why/references/source-playbook.md +17 -0
  221. package/pstack/skills/why/references/sources/code-archaeology.md +88 -0
  222. package/pstack/skills/why/references/sources/databricks.md +70 -0
  223. package/pstack/skills/why/references/sources/datadog.md +99 -0
  224. package/pstack/skills/why/references/sources/incident-postmortem.md +15 -0
  225. package/pstack/skills/why/references/sources/linear.md +48 -0
  226. package/pstack/skills/why/references/sources/notion.md +55 -0
  227. package/pstack/skills/why/references/sources/sentry.md +100 -0
  228. package/pstack/skills/why/references/sources/slack.md +54 -0
  229. package/pstack/skills/why/references/synthesizer-prompt.md +135 -0
@@ -0,0 +1,83 @@
1
+ # Build the change and clean the diff
2
+
3
+ The build playbooks share one discipline. Say what you observed, let the playbook demand the evidence. This page shows what to put in the prompt for each common build task, then the cleanup habit that keeps diffs reviewable.
4
+
5
+ ## Prompt each build playbook with what you know
6
+
7
+ A bug prompt states the symptom and asks for a reproduction first:
8
+
9
+ ```text
10
+ /poteto-mode this command emits two records after a retry. repro first, then fix and verify.
11
+ ```
12
+
13
+ A feature prompt states the behavior and what must not change:
14
+
15
+ ```text
16
+ /poteto-mode add a --json flag. text output stays byte-identical. verify both forms.
17
+ ```
18
+
19
+ A refactoring prompt pins behavior before structure moves:
20
+
21
+ ```text
22
+ /poteto-mode move parsing into one module, zero behavior change. record the current output first and prove it's unchanged after.
23
+ ```
24
+
25
+ A perf prompt states the measurement, not a vibe:
26
+
27
+ ```text
28
+ /poteto-mode startup takes 1.8s on this fixture. trace it, fix the measured cause, show me before and after.
29
+ ```
30
+
31
+ Each of these routes to its playbook ([Bug fix](../../skills/poteto-mode/playbooks/bug-fix.md), [Feature](../../skills/poteto-mode/playbooks/feature.md), [Refactoring](../../skills/poteto-mode/playbooks/refactoring.md), [Perf issue](../../skills/poteto-mode/playbooks/perf-issue.md)), and the playbook supplies the steps you didn't type: reproduce before fixing, name the data shape before implementing, pin behavior before restructuring, profile before optimizing.
32
+
33
+ For sustained improvement of one number, there's the [Hillclimb playbook](../../skills/poteto-mode/playbooks/hillclimb.md). Give it the metric, a target, and a floor on attempts, and it loops one hypothesis at a time with a frozen measurement harness. It keeps wins and reverts everything else.
34
+
35
+ Both perf playbooks run [`/benchmark-checklist`](../../skills/benchmark-checklist/SKILL.md) on their numbers. Perf issue vets its baseline and every number after it, and Hillclimb vets its harness before freezing it. [Verify and ship](./06-verify-and-ship.md#vet-a-measured-number-with-benchmark-checklist) shows when to type it yourself.
36
+
37
+ Sometimes you want the cause before any fix. For a live symptom, such as a leak, an idle CPU spin, or a visual glitch, the [Runtime forensics playbook](../../skills/poteto-mode/playbooks/runtime-forensics.md) instruments the running process. For a profile you already captured, the [Trace forensics playbook](../../skills/poteto-mode/playbooks/trace-forensics.md) reads the artifact and maps the hot frame to source. Both return a diagnosis, not a fix:
38
+
39
+ ```text
40
+ /poteto-mode here's a cpuprofile from the slow startup. tell me where the time goes and which source lines own it. no fix yet.
41
+ ```
42
+
43
+ ## Write the failing test first with `/tdd`
44
+
45
+ When a bug has a cheap local test path, the whole prompt can be two words:
46
+
47
+ ```text
48
+ /tdd implement
49
+ ```
50
+
51
+ In context, that's enough. [`/tdd`](../../skills/tdd/SKILL.md) writes the smallest test that fails for the intended reason, then the fix, then reruns the test. If a test would need broad harness setup or brittle mocks, the skill says so and uses the closest executable check instead. Don't force a test where a real command is stronger evidence.
52
+
53
+ ## Load the TypeScript rules by name
54
+
55
+ [`typescript-best-practices`](../../skills/typescript-best-practices/SKILL.md) turns the type-system principles into concrete rules: discriminated unions, `unknown` at boundaries, exhaustive variants, schema-derived types. It doesn't load on its own, so type `/typescript-best-practices` when a task touches `.ts` or `.tsx` files.
56
+
57
+ ## Clean before you commit
58
+
59
+ The [Opening a PR playbook](../../skills/poteto-mode/playbooks/opening-a-pr.md) runs bundled [`/deslop`](../../skills/deslop/SKILL.md) on the diff before each commit and applies [`/unslop`](../../skills/unslop/SKILL.md) to the PR description and commit bodies. The [`pstack-harness`](../../skills/pstack-harness/SKILL.md) **deslop** row loads it as `/pstack:deslop` on Claude Code or `$deslop` on Codex. It removes unnecessary comments, redundant guards, type escape hatches, and needless abstractions while preserving behavior and local style.
60
+
61
+ For prose, `/unslop` takes a target and any extra rules you have:
62
+
63
+ ```text
64
+ /unslop the readme changes, no emdashes
65
+ ```
66
+
67
+ You'll develop your own shorthand. The skill reads intent fine from terse prompts like `unslop that, tighten it`.
68
+
69
+ ## Strip the comments with `/no-comments`
70
+
71
+ Comments need their own pass, and not from the agent that wrote them. An author defends its comments the way you'd defend yours. So before review, hand them to fresh eyes:
72
+
73
+ ```text
74
+ /no-comments the diff
75
+ ```
76
+
77
+ [`/no-comments`](../../skills/no-comments/SKILL.md) spawns [Comment Sicko](../../agents/comment-sicko.md), a read-only reviewer with a short keep list: license headers, doc comments on a public API, links that explain what code can't, behavior forced by an external dependency you can't reshape. Everything else goes. A surprise in your own code gets no such pass. The comment comes back as a refactor flag, and `/no-comments` fixes the flags it accepts at the root cause. When a comment claims a constraint, "do not remove", the skill offers to encode the claim as a type, test, or lint. Either way, the comment comes out.
78
+
79
+ The division of labor is worth keeping straight. the deslop pass cleans slop out of the code, `/unslop` cleans it out of prose, and `/no-comments` hands the comments to a reviewer who didn't write them.
80
+
81
+ **Pitfall:** cleanup is not optional polish. A diff with narrating comments and defensive dead weight reads as unfinished to reviewers, and the extra code is where the next bug hides. If the diff feels padded, say `deslop it` before you commit, not after review calls it out.
82
+
83
+ Next: [Verify and ship](./06-verify-and-ship.md).
@@ -0,0 +1,130 @@
1
+ # Verify the result and open a PR
2
+
3
+ "It compiles" is not evidence. The [Prove It Works principle](../../skills/principle-prove-it-works/SKILL.md) makes the agent check the real artifact before it reports success, and your job is to make "the real artifact" checkable. This page covers stating a finish condition, vetting a measured number, generating a verification skill for your app, opening the PR, and driving it to merged.
4
+
5
+ Verification is the slowest step in most agent work, because it's the step that usually waits on a human. Make the agent able to do it, and you stop being the bottleneck. Skip it, and running more agents only gets you more unchecked work to review.
6
+
7
+ ![A prototype plane flies a real test course while she times it with a stopwatch and robots film and checklist the run; the terminal reads verify: pass, evidence: captured.](./images/verification.jpg)
8
+
9
+ ## State the finish condition up front
10
+
11
+ Put what done means in the first prompt, in whatever words fit:
12
+
13
+ ```text
14
+ /poteto-mode add json output to this command. text output stays byte-identical, the json parses, both run against the sample project. show me the evidence.
15
+ ```
16
+
17
+ Now the agent has three checks it can run, not a mood to satisfy. When the reply comes back, it should carry the exact commands and outputs. If a check couldn't run, a good reply says "inconclusive", and you should treat a confident reply without evidence as a red flag.
18
+
19
+ Match the check to the change:
20
+
21
+ - A CLI change runs the real command.
22
+ - A UI change walks the changed flow in the running app. When it must match a reference pixel for pixel, the [Visual parity playbook](../../skills/poteto-mode/playbooks/visual-parity.md) diffs screenshots against a frozen baseline instead of judging by eye.
23
+ - A parser or migration replays a saved input.
24
+ - A perf change compares before and after profiles.
25
+ - A storage change reads back the written value.
26
+
27
+ Ask for the proof as an artifact you can inspect yourself: the failing test and then the passing one, a before-and-after video, the trace, the screenshot. If the fix already merged, ask for the same check again on main. An artifact beats a plausible explanation, because you can challenge it without replaying the whole run.
28
+
29
+ For a small diff you don't fully trust, [`/blast-radius`](../../skills/blast-radius/SKILL.md) finds what it could break elsewhere. It picks the one fact the change is safe because of and proves it by running code instead of writing an essay about it.
30
+
31
+ ## Vet a measured number with `/benchmark-checklist`
32
+
33
+ A before-and-after number is the easiest evidence to get wrong by accident. A warm cache, a debug build on one side, or work that never ran inside the timed region can each produce a convincing speedup. Before you report or act on a number, type:
34
+
35
+ ```text
36
+ /benchmark-checklist vet the export speedup before it goes in the pr
37
+ ```
38
+
39
+ [`/benchmark-checklist`](../../skills/benchmark-checklist/SKILL.md) asks seven questions and wants evidence from a run for each:
40
+
41
+ 1. What limits the number, and why isn't it double?
42
+ 2. Did every side run tuned the way production runs?
43
+ 3. Does the result break a physical limit, like disk bandwidth or core count?
44
+ 4. Did anything error or return wrong output?
45
+ 5. Does it reproduce over alternating runs, with a median and a range?
46
+ 6. Does it matter end to end, on the path a user waits on?
47
+ 7. Did the work actually happen inside the timed region?
48
+
49
+ The verdict comes back as faster, slower, no measurable difference, or inconclusive, with the run count, range, and limiter. It says inconclusive when it can't name the limiter or a side ran untuned. `/poteto-mode` already runs the checklist inside the Perf issue and Hillclimb playbooks, so you type it yourself when you measured something outside them, or when someone else's number looks too good. It's the working form of the [Explain the Number principle](../../skills/principle-explain-the-number/SKILL.md).
50
+
51
+ ## Create a project verification skill
52
+
53
+ The UI bullet above hides a real requirement. The agent needs a scripted way to drive your app. Bundled [`/control-ui`](../../skills/control-ui/SKILL.md) and [`/control-cli`](../../skills/control-cli/SKILL.md) provide generic browser and terminal workflows using tools available in the session. `poteto-mode` uses them when no project verification skill exists. For repeatable app-specific commands, selectors, and a feature map, run:
54
+
55
+ ```text
56
+ /create-verification-skill
57
+ ```
58
+
59
+ [`/create-verification-skill`](../../skills/create-verification-skill/SKILL.md) interviews the repository, not you. It works out what a user touches, how the app launches locally, what can drive it (an existing harness first, otherwise browser and CDP, a PTY, or plain HTTP), what evidence proves behavior, and whether two instances can run side by side. It asks you only what the code can't answer.
60
+
61
+ It writes `verify-<app>/` into the project skills directory (`.claude/skills/` on Claude Code, `.agents/skills/` on Codex), agent-facing instructions with exact Launch, Doctor, Drive, Evidence, and Cleanup sections, plus a feature map under `features/` that indexes what the app does and what result proves each feature works. The skill ships a [worked feature-map example](../../skills/create-verification-skill/references/feature-map-example/) with a README index and one file per feature using the four required H2s. Before handing it over, the generator proves the skill once end to end: launch, doctor check, drive one feature, capture evidence, clean up. If that proof fails, don't use the output.
62
+
63
+ From then on, "verify it in the app" is a step any agent can execute, in this repo, with no setup conversation. Name it in the prompt when you want the proof in a specific form:
64
+
65
+ ```text
66
+ /poteto-mode build the bulk-archive action. use /verify-<app> to verify your changes and show me a video and screenshots as proof.
67
+ ```
68
+
69
+ ```text
70
+ /poteto-mode repro this with /verify-<app>. if it repros on main, fix it and show me a video as proof.
71
+ ```
72
+
73
+ Once the verify skill works, a [`/swarm`](../../skills/swarm/SKILL.md) can split a full pass by feature-map entry and aggregate the results. A swarm of verifiers also confirms a perf win over a big enough sample, or fuzzes the app for regressions before a PR ships.
74
+
75
+ Treat the verification skill as infrastructure, not a one-off. Commit it, so every person and every agent on the team drives the app the same way. Then [build the lever](../../skills/principle-build-the-lever/SKILL.md). When agents keep writing throwaway scripts to click through the app, ask for a small control CLI that the skill calls instead. Agents spend fewer tokens, and every run becomes repeatable. A CLI that agents use well has these traits:
76
+
77
+ - A few composable commands, each doing real work, rather than many thin ones.
78
+ - A `--dry-run` option on anything destructive.
79
+ - Subcommands that reveal features gradually instead of all at once.
80
+ - Error messages that say what to do instead.
81
+ - Rich `--help` text.
82
+ - Machine-readable output, such as JSON.
83
+
84
+ While you're there, make the dev setup repeatable too: seeded data, test users, and one command that brings the environment up the same way every time.
85
+
86
+ ## Keep the verification skill honest
87
+
88
+ Apps change and feature maps rot. Run this at least once a day, ideally from a scheduled automation so nobody has to remember:
89
+
90
+ ```text
91
+ /maintain-verification-skill
92
+ ```
93
+
94
+ [`/maintain-verification-skill`](../../skills/maintain-verification-skill/SKILL.md) audits the generated skill: one read-only source reader per feature in parallel, then one live pass that drives every mapped feature. It ends in exactly one of three outcomes. `clean` means full coverage and nothing to ship. `changed` means one PR of proven corrections, confined to the verification skill's own directory. `blocked` names the blocker. It never edits product code. If the live pass catches a product regression, it reports the regression instead of papering over it in docs.
95
+
96
+ ## Open the PR
97
+
98
+ ```text
99
+ /poteto-mode open the pr. small ordered commits, evidence in the description.
100
+ ```
101
+
102
+ The [Opening a PR playbook](../../skills/poteto-mode/playbooks/opening-a-pr.md) works from a worktree, rebases the work into small ordered commits, cleans the diff, unslops the prose, and returns the PR link. Five narrow PRs beat one fat one, and stacked follow-ups beat a growing branch.
103
+
104
+ ## Drive the PR to merge-ready with Babysit
105
+
106
+ An open PR starts collecting blockers immediately. Checks fail, reviewers comment, trunk moves. Hand that churn to the [Babysit playbook](../../skills/poteto-mode/playbooks/babysit.md):
107
+
108
+ ```text
109
+ /poteto-mode babysit this pr. get it green.
110
+ ```
111
+
112
+ Babysit watches the PR with a bundled watcher and takes blockers in order: conflicts, then review threads, then CI. Every known fix batches into one push, so the checks restart once instead of after every fix. The comment triage is skeptical, because humans and bots file real catches and noise in the same list. A real finding gets a fix, and noise gets dismissed with the disproof posted on the thread. When all you want is status, ask smaller and Babysit answers without starting the loop:
113
+
114
+ ```text
115
+ /poteto-mode check on pr 123. anything outstanding?
116
+ ```
117
+
118
+ Babysit stops at merge-ready. It never merges, even with everything green, because merging is a different decision.
119
+
120
+ ## Land the stack with Shipping
121
+
122
+ Green is not the same as safe. When you're ready to land, say so:
123
+
124
+ ```text
125
+ /poteto-mode land the stack.
126
+ ```
127
+
128
+ The [Shipping playbook](../../skills/poteto-mode/playbooks/shipping.md) verifies each PR independently before it arms anything. One fresh agent per PR proves the behavior live, and the agent that judges a change is never the one that wrote it. Then Shipping lands only the contiguous verified run from the bottom, one PR at a time through GitHub by default or Origin when its CLI is available, and reports the first PR that breaks the chain. A verified PR sitting above an unverified one waits, because merging it would pull the gap in underneath.
129
+
130
+ Next: [Run work while you sleep](./07-overnight.md).
@@ -0,0 +1,120 @@
1
+ # Run work while you sleep
2
+
3
+ This is the payoff for everything before it. An agent you can trust to verify its own work is an agent you can leave alone with a hard task. What makes that safe isn't hope. It's a checkable finish condition, an isolated worktree, and a decision log you audit in the morning.
4
+
5
+ ![She waves goodnight from the door while robots keep the factory running, one updating a DECISION LOG wall board under a BUILD LOOP ACTIVE sign.](./images/overnight.jpg)
6
+
7
+ ## Earn the trust before the loop
8
+
9
+ A loop you don't trust just produces unchecked work faster, and the mess compounds with every iteration. Before you leave one running, check that it has earned it:
10
+
11
+ - You've done the task once by hand, or watched an agent do it, so you know what good looks like.
12
+ - The agent has the tools and signals you'd use yourself: the verification skill, the profiler, the logs.
13
+ - Every stage proves its work and can stop the line when the work misses the bar.
14
+ - You've read a few transcripts and turned the repeated failures into tools, skills, or checks.
15
+
16
+ Make the loop autonomous only after all four hold. Until then, run it while you watch.
17
+
18
+ ## The overnight contract
19
+
20
+ A good handoff has the goal, the finish condition, permissions, and an escape hatch. It doesn't need to be long:
21
+
22
+ ```text
23
+ /poteto-mode im going to bed. migrate every caller to the new parser in a fresh worktree off <base>.
24
+ done means zero old callers, all parser fixtures pass, old api deleted.
25
+ keep a decision log. don't ask me before committing.
26
+ loop until done. if you're truly stuck after a few hours, stop and write up why.
27
+ ```
28
+
29
+ Walk through what each line buys you:
30
+
31
+ - "im going to bed" is a session override. The agent stops asking and keeps going.
32
+ - "done means..." turns the goal into checks every iteration can run.
33
+ - "fresh worktree off `<base>`" keeps the run from colliding with anything else you have open.
34
+ - "don't ask me before committing" pre-answers the permission the agent would otherwise block on.
35
+ - "loop until done" asks for a wake mechanism. On Claude Code that is the built-in `/loop`. On Codex it is a polling loop or a monitor agent (the [`pstack-harness`](../../skills/pstack-harness/SKILL.md) **loop** row). The [Autonomous run playbook](../../skills/poteto-mode/playbooks/autonomous-run.md) uses it to re-check the finish condition on events or a heartbeat.
36
+ - The escape hatch lets it stop at a genuine dead end and write up why, which beats eight hours of creative goal reinterpretation.
37
+
38
+ Because you'll review this work after stepping away, `/poteto-mode` routes it through [`/figure-it-out`](../../skills/figure-it-out/SKILL.md), which designs the run's phases before any code and wires in the decision log.
39
+
40
+ To stop a run on purpose, tell the agent to pause, or that you're about to go offline or restart the harness. The [Pause safely playbook](../../skills/poteto-mode/playbooks/pause-safely.md) finishes or backs out of the current step, commits a work-in-progress checkpoint, and writes a resume note. A fresh chat picks the work up from that note through the Session pickup playbook. Saying "keep going" never triggers a pause.
41
+
42
+ ## What the loop does all night
43
+
44
+ ```mermaid
45
+ flowchart TD
46
+ A[Check the finish condition] --> B[Make the smallest justified change]
47
+ B --> C[Verify against the real artifact]
48
+ C --> D{Progress?}
49
+ D -->|Yes| E[Commit]
50
+ D -->|No| F[Discard]
51
+ E --> G[Log one decision row]
52
+ F --> G
53
+ G --> A
54
+ ```
55
+
56
+ One change, one check, one log row, every iteration. Changes that didn't help get discarded, not left to ride. A plateau means pivot, not stop, and the finish condition never quietly relaxes to declare victory.
57
+
58
+ ## The morning audit
59
+
60
+ [`/show-me-your-work`](../../skills/show-me-your-work/SKILL.md) is what makes the run reviewable. Each row records the time, phase, decision, reason, an evidence pointer, and the result, in a TSV at `decisions.tsv` (or `.audit/<task-slug>.tsv` when several runs share a directory). It stays local by default. Commit it when the work is ambitious enough that a reviewer needs the trail to trust the result.
61
+
62
+ When you're back, ask for the run in review form:
63
+
64
+ ```text
65
+ /show-me-your-work catch me up on what you did last night
66
+ ```
67
+
68
+ Before the skill hands back its summary, it spawns a reviewer on a different model family to read the trail and the transcript, and the reply ends with an Attention section listing what deserves your scrutiny. Read that section first, then the log rows it points at. You're auditing decisions, not re-reading the whole night.
69
+
70
+ ## When the night holds a queue, not a task
71
+
72
+ The contract above drives one task to one finish condition. Some nights hold more, a queue of independent changes or a whole program. Three playbooks scale the same trust up.
73
+
74
+ [Autopilot-full](../../skills/poteto-mode/playbooks/autopilot-full.md) runs a queue of independent PRs to merged. Each PR gets one owner agent that carries it from build through merge, and no owner merges on its own verdict. A swarm of fresh verifiers starts a round at the owner's code-ready head and again at every later push that changes the patch. Only a clean verdict on the patch that merges authorizes the merge:
75
+
76
+ ```text
77
+ /poteto-mode full autopilot on this queue. each item is independent. i want them merged by morning.
78
+ ```
79
+
80
+ [Autopilot-stack](../../skills/poteto-mode/playbooks/autopilot-stack.md) runs the same owner loop but ships nothing. You wake up to one linear base-branch stack with a verifier's verdict on every link, and you review and land it yourself. Pick it over Autopilot-full when the changes are coupled, or when you want your own eyes on the work before anything merges:
81
+
82
+ ```text
83
+ /poteto-mode autopilot these five changes but stack them, don't ship. i'll land the stack in the morning.
84
+ ```
85
+
86
+ [Orchestrate](../../skills/poteto-mode/playbooks/orchestrate.md) is for a program that outlives any single agent: multi-day, many stacked PRs, fleets of subagents under one standing coordinator chat. The coordinator authors briefs, collects what its subagents finish, keeps the lowest unmerged PR green, and never writes code itself. It's deliberately heavy machinery. If one agent could finish the work in a session, the playbook itself routes you back to the overnight contract above:
87
+
88
+ ```text
89
+ /poteto-mode orchestrate the store migration. own it until every package is converted and merged. i'll check in twice a day.
90
+ ```
91
+
92
+ ## Run many projects in parallel
93
+
94
+ A long-lived coordinator session gives one agent a persistent thread. The coordinator doesn't write code. It directs subagents, each in its own worktree, so the work keeps moving while you are away. That's the shape the Orchestrate playbook expects. Start your prompts to the coordinator with `/poteto-mode`, and the subagents it spawns follow the playbooks.
95
+
96
+ A few habits help:
97
+
98
+ - Give each body of work its own coordinator session, such as a feature, a migration, a perf push, or a tech-debt cleanup. Several can run side by side.
99
+ - Point the coordinator at related transcripts and PRs, finished ones included. They become context for every agent it spawns.
100
+ - Give each PR a verification swarm before it merges, and let Autopilot-stack or Autopilot-full carry the queue.
101
+ - Ask the coordinator for a plan backed by data, and have it answer open questions with prototypes before it asks you.
102
+
103
+ One prompt can carry a whole Project, from research through execution:
104
+
105
+ ```text
106
+ /poteto-mode refactor this repo so its architecture is more agent friendly. use /correct and /architect on past commits and review comments to find the mistakes agents make most here. use /recall for context from past chats. answer open questions with prototypes instead of asking me. come back with a plan backed by real data. once i approve it, run it with autopilot-stack or autopilot-full, and ask me which.
107
+ ```
108
+
109
+ ## Let loops start themselves
110
+
111
+ Every loop above still waits for you to start it. A scheduled or event-driven automation removes that step. Software maintenance splits into stages that suit this well: triage a report, reproduce it, fix it, verify the fix. Two rules keep such a line trustworthy:
112
+
113
+ - Every stage can stop the line. Triage can decide the report is expected behavior, repro can fail to reproduce it, and the fixer can judge the change too risky. Each of those outcomes is useful, because it keeps bad work from reaching the next stage, where it costs more to undo.
114
+ - Every stage hands over evidence. Repro attaches screenshots and video of the broken state, and the fix attaches before-and-after proof. A human can then check that the agent fixed the right thing before reading a line of code.
115
+
116
+ pstack ships this as a dormant [automation pack](../../automations/benny/README.md) for Slack issue reports. One automation triages each report. The other reproduces confirmed bugs and may prepare a small draft fix. Point an agent at its [`FOR_AGENTS.md`](../../automations/benny/FOR_AGENTS.md) and name the target repository to set it up.
117
+
118
+ **Pitfall:** a duration is not a finish condition. "work on this for 4 hours" gives the agent nothing to check, and you'll wake up to four hours of motion instead of a result. Give the loop a predicate that can pass or fail.
119
+
120
+ Next: [Steer with principle names](./08-principles.md).
@@ -0,0 +1,72 @@
1
+ # Steer with principle names
2
+
3
+ pstack ships 24 principles as individual skills. `/poteto-mode` reads their index at the start of every multi-step task, applies the ones the task triggers, and names each applied principle in its reply along with the decision it changed.
4
+
5
+ You don't invoke principles. You use their names to steer. Each name points at a complete rule the agent has already read, so one phrase redirects the work more precisely than a paragraph of instructions.
6
+
7
+ ## Steering in practice
8
+
9
+ Say the agent is about to bolt a new adapter onto three existing ones:
10
+
11
+ ```text
12
+ use subtract before you add. delete the obsolete adapters first, then design what's left.
13
+ ```
14
+
15
+ Say it claims success because the build passed:
16
+
17
+ ```text
18
+ apply prove it works. run the real import flow and show me the written records.
19
+ ```
20
+
21
+ Say two parallel attempts are about to write to the same branch:
22
+
23
+ ```text
24
+ separate before serializing shared state. give each attempt its own worktree, no locks.
25
+ ```
26
+
27
+ Each phrase lands because the rule behind it is specific. The agent still has to say, in its reply, which decision the rule changed. A principle citation with no decision behind it is the tell that it name-dropped instead of applying.
28
+
29
+ ## The 24, briefly
30
+
31
+ The core principles decide how much to build and when to rethink the design:
32
+
33
+ - [Laziness Protocol](../../skills/principle-laziness-protocol/SKILL.md) prefers deletion and the smallest change that solves the problem.
34
+ - [Foundational Thinking](../../skills/principle-foundational-thinking/SKILL.md) chooses the core data structures before writing logic.
35
+ - [Redesign from First Principles](../../skills/principle-redesign-from-first-principles/SKILL.md) integrates a new requirement as if it had been there from day one.
36
+ - [Attack the Premise](../../skills/principle-attack-the-premise/SKILL.md) questions the premise that two or more failed fixes shared, after a census of which actors hold the imbalance.
37
+ - [Subtract Before You Add](../../skills/principle-subtract-before-you-add/SKILL.md) removes dead weight before building on top of it.
38
+ - [Minimize Reader Load](../../skills/principle-minimize-reader-load/SKILL.md) collapses layers and hidden state a reader must hold in their head.
39
+ - [Outcome-Oriented Execution](../../skills/principle-outcome-oriented-execution/SKILL.md) converges rewrites on the target design instead of preserving throwaway compatibility states.
40
+ - [Experience First](../../skills/principle-experience-first/SKILL.md) chooses the user's result over implementation convenience.
41
+ - [Exhaust the Design Space](../../skills/principle-exhaust-the-design-space/SKILL.md) builds two or three competing prototypes when there's no precedent.
42
+ - [Build the Lever](../../skills/principle-build-the-lever/SKILL.md) builds the script that does or proves the work, so a reviewer can rerun it. When an agent keeps doing the same thing by hand, have it write the tool or skill it wishes it had. If a script can do a step the same way every time, use the script, and save agents for the judgment calls.
43
+
44
+ The architecture principles decide where state, validation, and compatibility live:
45
+
46
+ - [Model the Domain](../../skills/principle-model-the-domain/SKILL.md) encodes repeated rules in one structure, not scattered conditionals.
47
+ - [Boundary Discipline](../../skills/principle-boundary-discipline/SKILL.md) validates at the boundary and trusts internal types.
48
+ - [Type System Discipline](../../skills/principle-type-system-discipline/SKILL.md) makes illegal states unrepresentable.
49
+ - [Make Operations Idempotent](../../skills/principle-make-operations-idempotent/SKILL.md) converges retries on the same end state.
50
+ - [Migrate Callers Then Delete Legacy APIs](../../skills/principle-migrate-callers-then-delete-legacy-apis/SKILL.md) migrates and deletes in one wave.
51
+ - [Separate Before Serializing Shared State](../../skills/principle-separate-before-serializing-shared-state/SKILL.md) removes the sharing before adding coordination.
52
+
53
+ The verification principles define what counts as proof:
54
+
55
+ - [Prove It Works](../../skills/principle-prove-it-works/SKILL.md) verifies the real artifact, not a proxy.
56
+ - [Fix Root Causes](../../skills/principle-fix-root-causes/SKILL.md) reproduces and traces to the cause before changing code.
57
+ - [Sequence Work into Verifiable Units](../../skills/principle-sequence-verifiable-units/SKILL.md) ends each small unit in a check before starting the next.
58
+ - [Test Behavior, Not Implementation](../../skills/principle-test-behavior-not-implementation/SKILL.md) calls the code the way its users do and asserts a literal expected value, and deletes a test that would still pass if every imported function returned `undefined`.
59
+ - [Explain the Number](../../skills/principle-explain-the-number/SKILL.md) names what limits a measured number and rules out that it measured something else, before anyone trusts or reports it. [`/benchmark-checklist`](../../skills/benchmark-checklist/SKILL.md) turns it into seven questions you answer from real runs.
60
+
61
+ The delegation principles keep parallel work sane:
62
+
63
+ - [Guard the Context Window](../../skills/principle-guard-the-context-window/SKILL.md) routes bulk reading to subagents and keeps findings in the main chat.
64
+ - [Never Block on the Human](../../skills/principle-never-block-on-the-human/SKILL.md) proceeds on reversible work and presents the result.
65
+
66
+ And one meta principle:
67
+
68
+ - [Encode Lessons in Structure](../../skills/principle-encode-lessons-in-structure/SKILL.md) turns advice you've repeated twice into a lint, check, or script. [`/correct`](../../skills/correct/SKILL.md) applies it to a whole repo, as shown in [Make it yours](./09-make-it-yours.md#fix-the-environment-with-correct).
69
+
70
+ Don't memorize the list. Skim it now, then come back when you catch the agent doing something a name here would have prevented. That's how the vocabulary sticks.
71
+
72
+ Next: [Make it yours](./09-make-it-yours.md).
@@ -0,0 +1,100 @@
1
+ # Make it yours
2
+
3
+ poteto-mode is one person's style. The machinery underneath, playbooks, routing, model roles, works just as well wearing yours. This page covers generating a personal mode, capturing lessons from a session, fixing the repo so agents stop repeating mistakes, authoring a focused skill, and testing a skill change before you trust it.
4
+
5
+ Start smaller than you think. You don't need many skills on day one, or even this whole plugin. Prompt plainly, watch where agents fail, and add a skill or a check when the same failure shows up twice.
6
+
7
+ ## Generate your own mode with `/automate-me`
8
+
9
+ ```text
10
+ /automate-me
11
+ ```
12
+
13
+ You don't describe your style, because [`/automate-me`](../../skills/automate-me/SKILL.md) reads it out of your history. It mines your recent transcripts in the active workspace for repeated preferences, in how you like replies, delegation, verification, code, prose, and process, then asks you which patterns are really you. It drafts `<your-name>-mode/SKILL.md` in your project skills directory (`.claude/skills/` on Claude Code, `.agents/skills/` on Codex) through the `skill-creator` flow (harness **skill-creator** row), runs the draft through [`/unslop`](../../skills/unslop/SKILL.md), and opens a PR from a worktree so you review it like any other change.
14
+
15
+ Run it again whenever your habits drift:
16
+
17
+ ```text
18
+ /automate-me update my mode skill with everything since its last edit
19
+ ```
20
+
21
+ Update mode mines only the history since the skill last changed. It keeps rules you haven't contradicted, revises the ones with new evidence, and adds sections only for genuinely new patterns.
22
+
23
+ ## Capture a session's lessons with `/reflect`
24
+
25
+ Right after a task that taught you something, run:
26
+
27
+ ```text
28
+ /reflect that took way too long. capture what we learned so the next run doesn't repeat it.
29
+ ```
30
+
31
+ [`/reflect`](../../skills/reflect/SKILL.md) sends the transcript to three parallel reviewers, then a synthesizer sorts the proposals into `Accepted`, `Rejected`, and `Backlog` and waits for your approval before any skill changes. Approve a proposal only if it would change a future decision. One weird session is an anecdote, not a rule.
32
+
33
+ ## Fix the environment with `/correct`
34
+
35
+ When you correct agents for the same mistake again and again, the fix belongs in the repo, not in your next prompt. Rank the options by how well they hold:
36
+
37
+ 1. Make the mistake impossible with architecture or a better data structure.
38
+ 2. Block it with types, or with a lint or CI check whose error names the fix.
39
+ 3. Catch it with a test.
40
+ 4. Write it down as a doc or agent rule. Nothing fails when an agent skips a rule, so this comes last.
41
+
42
+ Human review isn't on the list. A reviewer who must catch the same mistake on every PR is the problem this fixes. [`/correct`](../../skills/correct/SKILL.md) does the work:
43
+
44
+ ```text
45
+ /correct agents keep calling the database client directly instead of going through the repository layer
46
+ ```
47
+
48
+ It reads recent commits, reverts, review comments, and comments that explain workarounds, then groups the mistakes into classes. A class counts once it has happened twice. It fixes the most frequent classes one commit each, at the highest level that works, and proves each new check fails on a real past mistake. It also keeps a table in the agent instruction file that pairs each rule with what enforces it, so a rule that nothing enforces shows up as a repeat. The reply lists each class with its evidence, the level chosen, and why a higher level didn't work.
49
+
50
+ Run it with no argument and it finds the classes from history on its own. `/reflect` and `/correct` split the work. `/reflect` improves skills from one session. `/correct` changes the repo so a mistake class can't come back. Pair it with `/architect` when the fix is a new boundary. [Run many projects in parallel](./07-overnight.md#run-many-projects-in-parallel) has a prompt that does both across a whole repo.
51
+
52
+ ## Author a focused skill
53
+
54
+ When you already know the workflow you want to capture:
55
+
56
+ ```text
57
+ /poteto-mode write a skill for verifying database migrations in this repo
58
+ ```
59
+
60
+ Writing a skill matches the [Authoring or modifying a skill playbook](../../skills/poteto-mode/playbooks/authoring-a-skill.md), which routes through `skill-creator`, validates the frontmatter and links, and ships the result through the Opening a PR playbook. Agent-facing prose has a higher bar than human prose, because an unhelpful sentence becomes an instruction some future agent follows. Let the playbook hold that bar rather than writing a `SKILL.md` freehand.
61
+
62
+ One special case has its own generator. A skill that must drive your app and prove behavior is a verification skill, so use [`/create-verification-skill`](../../skills/create-verification-skill/SKILL.md) and [`/maintain-verification-skill`](../../skills/maintain-verification-skill/SKILL.md) instead. [Verify and ship](./06-verify-and-ship.md#create-a-project-verification-skill) covers both.
63
+
64
+ ## Write docs to a standard with `/technical-writing`
65
+
66
+ Skills aren't the only prose you ship. For docs, RFCs, readmes, PR descriptions, and commit messages:
67
+
68
+ ```text
69
+ /technical-writing review the readme changes
70
+ ```
71
+
72
+ [`/technical-writing`](../../skills/technical-writing/SKILL.md) applies a layered standard with one goal, prose a tired engineer understands on the first read. It picks the document's mode first (tutorial, how-to, reference, or explanation), then works sentence by sentence: who does what, one thought per sentence, nothing readable two ways. Use it to review what you or an agent just wrote, or name it up front when you ask for a doc.
73
+
74
+ ## Test a skill change blind
75
+
76
+ A skill edit affects every future session, so test it like the experiment it is. The same goes for adopting someone else's skill. Check that it makes the agent better on your work before you keep it.
77
+
78
+ ```text
79
+ /poteto-mode run the eval playbook on this skill change. same task for both variants, candidates stay blind.
80
+ ```
81
+
82
+ When a skill keeps missing and you know what it should do, change and test it in one task:
83
+
84
+ ```text
85
+ /poteto-mode update the review skill so it flags missing migrations, and eval the change.
86
+ ```
87
+
88
+ Asking for the eval up front keeps the edit honest. A fix written from one bad session tends to overfit that session, and over many edits the skill drifts. The eval catches the drift before it ships.
89
+
90
+ The [Eval playbook](../../skills/poteto-mode/playbooks/eval.md) is built around one failure mode, the observer effect. An agent that knows it's being evaluated behaves differently. So candidate agents get an organic-looking task in sanitized directories, never the words "eval" or "candidate", and never each other's existence. One judge scores all outputs under neutral labels, and chain-following gets graded from which files each candidate actually read, not from what it claims.
91
+
92
+ Read every output yourself before accepting the verdict. If you disagree with the judge, suspect the rubric before you suspect your judgment.
93
+
94
+ ## Build a bot UI with `/make-bot-ui`
95
+
96
+ One situational skill is for people who want to drive an agent from a page instead of a prompt. [`/make-bot-ui`](../../skills/make-bot-ui/SKILL.md) builds a small page whose buttons wake a bot. For example, you could swipe through a review queue and have each swipe ask the bot to act on that item. It knows three bots: a Claude Code session over a channel on this machine, Claude Code with the reply sent to your phone through the official Telegram channel plugin, and a ChatGPT Dot that watches a Slack channel. A server on your machine holds every token, so no secret reaches the browser or the chat. The skill also covers exposing the page on Tailscale behind a gate token. Channels are a Claude Code research preview, and Codex has none, so Codex users take the Dot route.
97
+
98
+ **Pitfall:** don't edit a skill mid-task because it's misbehaving. Fix it in its own PR and keep the task moving. A skill edit that ships tangled into feature work is invisible to review and impossible to evaluate.
99
+
100
+ Next: [Recipes and pitfalls](./10-recipes-and-pitfalls.md).