@jiroamato/pstack 0.0.0-stage → 0.16.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (229) hide show
  1. package/LICENSE +21 -0
  2. package/README.md +79 -2
  3. package/bin/pstack.js +95 -0
  4. package/lib/install.js +121 -0
  5. package/lib/prompt.js +77 -0
  6. package/lib/targets.js +43 -0
  7. package/package.json +38 -5
  8. package/pstack/.claude-plugin/plugin.json +26 -0
  9. package/pstack/.codex-plugin/plugin.json +36 -0
  10. package/pstack/LICENSE +21 -0
  11. package/pstack/LICENSE-cursor-team-kit +21 -0
  12. package/pstack/NOTICE +8 -0
  13. package/pstack/README.md +307 -0
  14. package/pstack/agents/comment-sicko.md +34 -0
  15. package/pstack/agents/poteto-agent.md +10 -0
  16. package/pstack/automations/benny/FOR_AGENTS.md +92 -0
  17. package/pstack/automations/benny/README.md +28 -0
  18. package/pstack/automations/benny/skills/reproduce-and-fix-issues/SKILL.md +313 -0
  19. package/pstack/automations/benny/skills/reproduce-and-fix-issues/references/control-adapter.md +169 -0
  20. package/pstack/automations/benny/skills/reproduce-and-fix-issues/references/feature-map.example.md +205 -0
  21. package/pstack/automations/benny/skills/reproduce-and-fix-issues/references/verify-existing-fix.md +93 -0
  22. package/pstack/automations/benny/skills/setup-benny/SKILL.md +271 -0
  23. package/pstack/automations/benny/skills/triage-issue-reports/SKILL.md +240 -0
  24. package/pstack/automations/benny/skills/triage-issue-reports/references/routing.example.md +61 -0
  25. package/pstack/automations/benny/templates/configuration.example.yaml +84 -0
  26. package/pstack/automations/benny/templates/reproduce-automation-prompt.md +33 -0
  27. package/pstack/automations/benny/templates/triage-automation-prompt.md +39 -0
  28. package/pstack/codex/agents/comment-sicko.toml +36 -0
  29. package/pstack/codex/agents/poteto-agent.toml +11 -0
  30. package/pstack/docs/guide/01-setup.md +80 -0
  31. package/pstack/docs/guide/02-poteto-mode.md +131 -0
  32. package/pstack/docs/guide/03-understand.md +79 -0
  33. package/pstack/docs/guide/04-design.md +133 -0
  34. package/pstack/docs/guide/05-build-and-clean.md +83 -0
  35. package/pstack/docs/guide/06-verify-and-ship.md +130 -0
  36. package/pstack/docs/guide/07-overnight.md +120 -0
  37. package/pstack/docs/guide/08-principles.md +72 -0
  38. package/pstack/docs/guide/09-make-it-yours.md +100 -0
  39. package/pstack/docs/guide/10-recipes-and-pitfalls.md +156 -0
  40. package/pstack/docs/guide/README.md +38 -0
  41. package/pstack/skills/architect/SKILL.md +85 -0
  42. package/pstack/skills/architect/agents/openai.yaml +2 -0
  43. package/pstack/skills/architect/references/design-red-flags.md +57 -0
  44. package/pstack/skills/architect/references/rationale-template.md +35 -0
  45. package/pstack/skills/architect/references/runner-prompt.md +20 -0
  46. package/pstack/skills/arena/SKILL.md +75 -0
  47. package/pstack/skills/arena/agents/openai.yaml +2 -0
  48. package/pstack/skills/automate-me/SKILL.md +104 -0
  49. package/pstack/skills/automate-me/agents/openai.yaml +2 -0
  50. package/pstack/skills/benchmark-checklist/SKILL.md +39 -0
  51. package/pstack/skills/benchmark-checklist/agents/openai.yaml +2 -0
  52. package/pstack/skills/blast-radius/SKILL.md +52 -0
  53. package/pstack/skills/blast-radius/agents/openai.yaml +2 -0
  54. package/pstack/skills/bro/SKILL.md +7 -0
  55. package/pstack/skills/bro/agents/openai.yaml +2 -0
  56. package/pstack/skills/control-cli/SKILL.md +55 -0
  57. package/pstack/skills/control-cli/agents/openai.yaml +2 -0
  58. package/pstack/skills/control-ui/SKILL.md +72 -0
  59. package/pstack/skills/control-ui/agents/openai.yaml +2 -0
  60. package/pstack/skills/correct/SKILL.md +34 -0
  61. package/pstack/skills/correct/agents/openai.yaml +2 -0
  62. package/pstack/skills/create-verification-skill/SKILL.md +47 -0
  63. package/pstack/skills/create-verification-skill/agents/openai.yaml +2 -0
  64. package/pstack/skills/create-verification-skill/references/feature-map-example/README.md +47 -0
  65. package/pstack/skills/create-verification-skill/references/feature-map-example/create-note.md +39 -0
  66. package/pstack/skills/create-verification-skill/references/feature-map-example/search.md +45 -0
  67. package/pstack/skills/deslop/SKILL.md +30 -0
  68. package/pstack/skills/deslop/agents/openai.yaml +2 -0
  69. package/pstack/skills/figure-it-out/SKILL.md +55 -0
  70. package/pstack/skills/figure-it-out/agents/openai.yaml +2 -0
  71. package/pstack/skills/how/SKILL.md +58 -0
  72. package/pstack/skills/how/agents/openai.yaml +2 -0
  73. package/pstack/skills/how/references/explainer-prompt.md +55 -0
  74. package/pstack/skills/how/references/explorer-prompt.md +52 -0
  75. package/pstack/skills/interrogate/SKILL.md +111 -0
  76. package/pstack/skills/interrogate/agents/openai.yaml +2 -0
  77. package/pstack/skills/interrogate/references/code-quality-review.md +47 -0
  78. package/pstack/skills/interrogate/references/lead-judgment.md +58 -0
  79. package/pstack/skills/interrogate/references/reviewer-prompt.md +70 -0
  80. package/pstack/skills/interrogate/references/rubric.md +77 -0
  81. package/pstack/skills/kiss/SKILL.md +90 -0
  82. package/pstack/skills/kiss/agents/openai.yaml +2 -0
  83. package/pstack/skills/kiss/references/assess.md +110 -0
  84. package/pstack/skills/kiss/references/principles.md +138 -0
  85. package/pstack/skills/maintain-verification-skill/SKILL.md +41 -0
  86. package/pstack/skills/maintain-verification-skill/agents/openai.yaml +2 -0
  87. package/pstack/skills/make-bot-ui/SKILL.md +289 -0
  88. package/pstack/skills/make-bot-ui/agents/openai.yaml +2 -0
  89. package/pstack/skills/no-comments/SKILL.md +24 -0
  90. package/pstack/skills/no-comments/agents/openai.yaml +2 -0
  91. package/pstack/skills/poteto-help/SKILL.md +156 -0
  92. package/pstack/skills/poteto-help/agents/openai.yaml +2 -0
  93. package/pstack/skills/poteto-help/references/prompting.md +51 -0
  94. package/pstack/skills/poteto-help/references/recipes.md +47 -0
  95. package/pstack/skills/poteto-mode/SKILL.md +143 -0
  96. package/pstack/skills/poteto-mode/agents/openai.yaml +2 -0
  97. package/pstack/skills/poteto-mode/playbooks/authoring-a-skill.md +12 -0
  98. package/pstack/skills/poteto-mode/playbooks/autonomous-run.md +13 -0
  99. package/pstack/skills/poteto-mode/playbooks/autopilot-full.md +13 -0
  100. package/pstack/skills/poteto-mode/playbooks/autopilot-stack.md +16 -0
  101. package/pstack/skills/poteto-mode/playbooks/babysit.md +29 -0
  102. package/pstack/skills/poteto-mode/playbooks/bug-fix.md +15 -0
  103. package/pstack/skills/poteto-mode/playbooks/eval.md +25 -0
  104. package/pstack/skills/poteto-mode/playbooks/feature.md +21 -0
  105. package/pstack/skills/poteto-mode/playbooks/hillclimb.md +21 -0
  106. package/pstack/skills/poteto-mode/playbooks/investigation.md +14 -0
  107. package/pstack/skills/poteto-mode/playbooks/multi-phase-plan.md +155 -0
  108. package/pstack/skills/poteto-mode/playbooks/opening-a-pr.md +38 -0
  109. package/pstack/skills/poteto-mode/playbooks/orchestrate.md +114 -0
  110. package/pstack/skills/poteto-mode/playbooks/pause-safely.md +10 -0
  111. package/pstack/skills/poteto-mode/playbooks/perf-issue.md +25 -0
  112. package/pstack/skills/poteto-mode/playbooks/prototype.md +14 -0
  113. package/pstack/skills/poteto-mode/playbooks/refactoring.md +16 -0
  114. package/pstack/skills/poteto-mode/playbooks/runtime-forensics.md +11 -0
  115. package/pstack/skills/poteto-mode/playbooks/session-pickup.md +11 -0
  116. package/pstack/skills/poteto-mode/playbooks/shipping.md +17 -0
  117. package/pstack/skills/poteto-mode/playbooks/trace-forensics.md +14 -0
  118. package/pstack/skills/poteto-mode/playbooks/visual-parity.md +11 -0
  119. package/pstack/skills/poteto-mode/playbooks/worktree-cleanup.md +14 -0
  120. package/pstack/skills/poteto-mode/references/bugbot-triage.md +142 -0
  121. package/pstack/skills/poteto-mode/scripts/bootstrap.ts +62 -0
  122. package/pstack/skills/poteto-mode/scripts/bun.lock +67 -0
  123. package/pstack/skills/poteto-mode/scripts/check-plan.mjs +185 -0
  124. package/pstack/skills/poteto-mode/scripts/orch/orch.test.ts +634 -0
  125. package/pstack/skills/poteto-mode/scripts/orch/orch.ts +578 -0
  126. package/pstack/skills/poteto-mode/scripts/orch/store.ts +1607 -0
  127. package/pstack/skills/poteto-mode/scripts/package.json +16 -0
  128. package/pstack/skills/poteto-mode/scripts/watch-pr/cli.test.ts +224 -0
  129. package/pstack/skills/poteto-mode/scripts/watch-pr/cli.ts +223 -0
  130. package/pstack/skills/poteto-mode/scripts/watch-pr/fakes.test-helper.ts +118 -0
  131. package/pstack/skills/poteto-mode/scripts/watch-pr/github.test.ts +306 -0
  132. package/pstack/skills/poteto-mode/scripts/watch-pr/github.ts +699 -0
  133. package/pstack/skills/poteto-mode/scripts/watch-pr/policy.test.ts +420 -0
  134. package/pstack/skills/poteto-mode/scripts/watch-pr/policy.ts +832 -0
  135. package/pstack/skills/poteto-mode/scripts/watch-pr/render.ts +169 -0
  136. package/pstack/skills/poteto-mode/scripts/watch-pr/tsconfig.json +13 -0
  137. package/pstack/skills/poteto-mode/scripts/watch-pr/types.compile.ts +93 -0
  138. package/pstack/skills/poteto-mode/scripts/watch-pr/types.ts +401 -0
  139. package/pstack/skills/poteto-mode/scripts/watch-pr/watch-pr +6 -0
  140. package/pstack/skills/poteto-mode/scripts/worktree-audit.sh +92 -0
  141. package/pstack/skills/principle-attack-the-premise/SKILL.md +23 -0
  142. package/pstack/skills/principle-attack-the-premise/agents/openai.yaml +2 -0
  143. package/pstack/skills/principle-boundary-discipline/SKILL.md +34 -0
  144. package/pstack/skills/principle-boundary-discipline/agents/openai.yaml +2 -0
  145. package/pstack/skills/principle-build-the-lever/SKILL.md +23 -0
  146. package/pstack/skills/principle-build-the-lever/agents/openai.yaml +2 -0
  147. package/pstack/skills/principle-encode-lessons-in-structure/SKILL.md +31 -0
  148. package/pstack/skills/principle-encode-lessons-in-structure/agents/openai.yaml +2 -0
  149. package/pstack/skills/principle-exhaust-the-design-space/SKILL.md +21 -0
  150. package/pstack/skills/principle-exhaust-the-design-space/agents/openai.yaml +2 -0
  151. package/pstack/skills/principle-experience-first/SKILL.md +19 -0
  152. package/pstack/skills/principle-experience-first/agents/openai.yaml +2 -0
  153. package/pstack/skills/principle-explain-the-number/SKILL.md +23 -0
  154. package/pstack/skills/principle-explain-the-number/agents/openai.yaml +2 -0
  155. package/pstack/skills/principle-fix-root-causes/SKILL.md +23 -0
  156. package/pstack/skills/principle-fix-root-causes/agents/openai.yaml +2 -0
  157. package/pstack/skills/principle-foundational-thinking/SKILL.md +21 -0
  158. package/pstack/skills/principle-foundational-thinking/agents/openai.yaml +2 -0
  159. package/pstack/skills/principle-guard-the-context-window/SKILL.md +16 -0
  160. package/pstack/skills/principle-guard-the-context-window/agents/openai.yaml +2 -0
  161. package/pstack/skills/principle-laziness-protocol/SKILL.md +18 -0
  162. package/pstack/skills/principle-laziness-protocol/agents/openai.yaml +2 -0
  163. package/pstack/skills/principle-make-operations-idempotent/SKILL.md +24 -0
  164. package/pstack/skills/principle-make-operations-idempotent/agents/openai.yaml +2 -0
  165. package/pstack/skills/principle-migrate-callers-then-delete-legacy-apis/SKILL.md +22 -0
  166. package/pstack/skills/principle-migrate-callers-then-delete-legacy-apis/agents/openai.yaml +2 -0
  167. package/pstack/skills/principle-minimize-reader-load/SKILL.md +23 -0
  168. package/pstack/skills/principle-minimize-reader-load/agents/openai.yaml +2 -0
  169. package/pstack/skills/principle-model-the-domain/SKILL.md +26 -0
  170. package/pstack/skills/principle-model-the-domain/agents/openai.yaml +2 -0
  171. package/pstack/skills/principle-never-block-on-the-human/SKILL.md +20 -0
  172. package/pstack/skills/principle-never-block-on-the-human/agents/openai.yaml +2 -0
  173. package/pstack/skills/principle-outcome-oriented-execution/SKILL.md +21 -0
  174. package/pstack/skills/principle-outcome-oriented-execution/agents/openai.yaml +2 -0
  175. package/pstack/skills/principle-prove-it-works/SKILL.md +22 -0
  176. package/pstack/skills/principle-prove-it-works/agents/openai.yaml +2 -0
  177. package/pstack/skills/principle-redesign-from-first-principles/SKILL.md +16 -0
  178. package/pstack/skills/principle-redesign-from-first-principles/agents/openai.yaml +2 -0
  179. package/pstack/skills/principle-separate-before-serializing-shared-state/SKILL.md +16 -0
  180. package/pstack/skills/principle-separate-before-serializing-shared-state/agents/openai.yaml +2 -0
  181. package/pstack/skills/principle-sequence-verifiable-units/SKILL.md +17 -0
  182. package/pstack/skills/principle-sequence-verifiable-units/agents/openai.yaml +2 -0
  183. package/pstack/skills/principle-subtract-before-you-add/SKILL.md +21 -0
  184. package/pstack/skills/principle-subtract-before-you-add/agents/openai.yaml +2 -0
  185. package/pstack/skills/principle-test-behavior-not-implementation/SKILL.md +25 -0
  186. package/pstack/skills/principle-test-behavior-not-implementation/agents/openai.yaml +2 -0
  187. package/pstack/skills/principle-type-system-discipline/SKILL.md +31 -0
  188. package/pstack/skills/principle-type-system-discipline/agents/openai.yaml +2 -0
  189. package/pstack/skills/pstack-harness/SKILL.md +67 -0
  190. package/pstack/skills/recall/SKILL.md +35 -0
  191. package/pstack/skills/recall/agents/openai.yaml +2 -0
  192. package/pstack/skills/reflect/SKILL.md +76 -0
  193. package/pstack/skills/reflect/agents/openai.yaml +2 -0
  194. package/pstack/skills/reflect/references/divergent-reviewer.md +43 -0
  195. package/pstack/skills/reflect/references/judgment-reviewer.md +42 -0
  196. package/pstack/skills/reflect/references/synthesizer.md +56 -0
  197. package/pstack/skills/reflect/references/tooling-reviewer.md +55 -0
  198. package/pstack/skills/setup-pstack/SKILL.md +110 -0
  199. package/pstack/skills/show-me-your-work/SKILL.md +82 -0
  200. package/pstack/skills/show-me-your-work/agents/openai.yaml +2 -0
  201. package/pstack/skills/show-me-your-work/references/decision-log-template.tsv +1 -0
  202. package/pstack/skills/show-me-your-work/scripts/log.sh +42 -0
  203. package/pstack/skills/swarm/SKILL.md +48 -0
  204. package/pstack/skills/swarm/agents/openai.yaml +2 -0
  205. package/pstack/skills/tdd/SKILL.md +44 -0
  206. package/pstack/skills/tdd/agents/openai.yaml +2 -0
  207. package/pstack/skills/teach/SKILL.md +21 -0
  208. package/pstack/skills/teach/agents/openai.yaml +2 -0
  209. package/pstack/skills/technical-writing/SKILL.md +106 -0
  210. package/pstack/skills/technical-writing/agents/openai.yaml +2 -0
  211. package/pstack/skills/typescript-best-practices/SKILL.md +31 -0
  212. package/pstack/skills/typescript-best-practices/agents/openai.yaml +2 -0
  213. package/pstack/skills/typescript-best-practices/references/patterns.md +324 -0
  214. package/pstack/skills/unslop/SKILL.md +67 -0
  215. package/pstack/skills/unslop/agents/openai.yaml +2 -0
  216. package/pstack/skills/why/SKILL.md +158 -0
  217. package/pstack/skills/why/agents/openai.yaml +2 -0
  218. package/pstack/skills/why/references/epistemics.md +144 -0
  219. package/pstack/skills/why/references/investigator-prompt.md +103 -0
  220. package/pstack/skills/why/references/source-playbook.md +17 -0
  221. package/pstack/skills/why/references/sources/code-archaeology.md +88 -0
  222. package/pstack/skills/why/references/sources/databricks.md +70 -0
  223. package/pstack/skills/why/references/sources/datadog.md +99 -0
  224. package/pstack/skills/why/references/sources/incident-postmortem.md +15 -0
  225. package/pstack/skills/why/references/sources/linear.md +48 -0
  226. package/pstack/skills/why/references/sources/notion.md +55 -0
  227. package/pstack/skills/why/references/sources/sentry.md +100 -0
  228. package/pstack/skills/why/references/sources/slack.md +54 -0
  229. package/pstack/skills/why/references/synthesizer-prompt.md +135 -0
@@ -0,0 +1,21 @@
1
+ ### Feature
2
+
3
+ **You own the design. Plan, review, verify.** Delegate implementation. Stay in the lead.
4
+
5
+ 1. `how` over the affected subsystem.
6
+ 2. `architect` for parallel design exploration.
7
+ 3. Write the throughput checkpoint as four todo items. A dimension that genuinely does not apply (single file, no fan-out) keeps its item with `n/a: <reason>` rather than being dropped:
8
+ - **Blocking first steps.** Gates run before fan-out.
9
+ - **Independent workstreams.** Disjoint files, services, or layers parallelize. Shared writes serialize.
10
+ - **Shared mutable state.** Default to splitting the target (the **separate-before-serializing-shared-state** principle skill). Serialize only for real invariants.
11
+ - **Smallest safe decomposition.** If one worker is best, name why.
12
+ 4. Delegate code-writing to a subagent using your configured feature model (default the code tier) with a specific scope (file paths, named data shape and its organizing structure per **principle-model-the-domain**, a state machine over scattered booleans, a table/registry over branching, a typed model over repeated shape assumptions, chosen before the delegate writes logic, and success criteria). When the implementation admits multiple valid shapes (error handling, abstraction layer, test structure), delegate via the **arena** skill instead so the runners surface the alternatives and the cross-judge guards the pick. Mandatory: no skip-with-reason escape, and Laziness Protocol does not override it (the gain is review separation, not lines saved). A subagent forbidden to spawn satisfies this by owning the diff directly with the same review separation. No "standing by" reply that waits on a nested agent. Comments per **Comments**. Surgical edits, re-ground against the source for upstream-derived files. Port shared-primitive improvements to all consumers and verify each. Commit liberally.
13
+ 5. Load `control-ui` for web, IDE, or Electron, or `control-cli` for CLI/TUI, and verify on that surface (harness **verification harness** row). "Inconclusive" or wrong-surface is not a pass. Flag it.
14
+ 6. Rebase into small, ordered commits. Stack follow-ups.
15
+ Use the **sequence-verifiable-units** principle skill, building, verifying, and committing each small unit before the next.
16
+ 7. If the design is contested, `interrogate` before shipping.
17
+ 8. Run **Opening a PR**.
18
+
19
+ Code-coupled work (one feature, one migration) goes to a single owner with the checkpoint inline. That owner fans out internally after the blocking phase. Parent-level fan-out is for slices that produce independent artifacts (audits, cross-subsystem investigations, competing experiments). Rewrite the checkpoint at phase boundaries. Spawn a fresh owner rather than chaining interrupts.
20
+
21
+ **Reply:** what you built, what you chose and why, the throughput checkpoint, open decisions. Tables for design alternatives.
@@ -0,0 +1,21 @@
1
+ ### Hillclimb
2
+
3
+ **You own the metric and the experiment's integrity. Supervise and review. Delegate the attempts.** For sustained, iterative improvement of one measurable thing against a target. A one-off fix is Bug fix or Perf issue. This is the loop.
4
+
5
+ Core discipline: one change, one measurement, keep or revert. Never stack untested changes, and never claim a win from code inspection (the **prove-it-works** principle skill).
6
+
7
+ 1. Ground the workload and architecture before choosing the metric. Run the **how** skill over the target, name the realistic workload dimensions that can move the result (data size, history, state, concurrency), and select a case that reproduces the user's complaint. If no case reproduces it, fix the repro instead of hillclimbing. Then fix one metric, the direction that counts as better, and a checkable stop predicate that pairs a target with a floor on attempts so a lucky early win can't end the run (the example "at least 50% better than baseline and at least 10 iterations" is this shape). Use the user's numbers when given, otherwise agree them.
8
+ 2. For web, IDE, or Electron measurements, load `control-ui`; for CLI/TUI, load `control-cli` (harness **verification harness** row). Build the measurement harness, prove its sensitivity, then freeze it (the **build-the-lever** principle skill). Run contrasting realistic workloads and confirm the target case reproduces the symptom while easier cases separate as expected. If the harness cannot distinguish them, revise the workload or metric. Vet the harness with the **benchmark-checklist** skill before you freeze it, and make it print its error count and a count of the work done. Once frozen, one repeatable command emits the metric, sampled enough to clear the noise (median of N, not a single run). Record the baseline metric and a green run of the regression gate (the tests that must keep passing) before any change.
9
+ 3. Open the decision log via the **show-me-your-work** skill. A `decision.tsv`, one row per attempt: id, hypothesis, change, before, after, delta, tests, verdict (kept or reverted), note. Read it before each attempt. Keep it out of the tree (gitignored).
10
+ 4. Ground each hypothesis in the architecture model from step 1, so it names a specific mechanism ("defer X off the boot path because it blocks first paint"), not "try memoizing something". For a perf metric, order hypotheses by the performance mantras in step 2 of the Perf issue playbook (`playbooks/perf-issue.md`). Borrow only their order, not that step's stop rule.
11
+ 5. Loop, one hypothesis per iteration:
12
+ - Hand the change to a subagent using your configured hillclimb model (default the code tier) with a tight scope. Supervise and review the diff rather than typing it (the **guard-the-context-window** principle skill). When several independent hypotheses are live, fan them to parallel subagents, each in its own worktree (the **separate-before-serializing-shared-state** principle skill).
13
+ - Measure before and after with the frozen harness, and run the regression gate.
14
+ - Accept only when the metric moves past noise and the gate stays green. Otherwise revert the change in full. A tweak that "might help" is not kept.
15
+ - Load bundled `deslop` per the harness **deslop** row before each accepted-fix commit. Rerun affected measurements and the regression gate if cleanup changed code. One commit per accepted fix, staging only the files you changed (`git add <files>`, never `-A`). Log the row either way, kept or reverted.
16
+ Each iteration ends in a check before the next begins (the **sequence-verifiable-units** principle skill). If the run is unattended, borrow only the wake mechanism from the Autonomous run playbook (`playbooks/autonomous-run.md`), not its stop rule.
17
+ 6. Push past the first plateau. On a stall, several rejects in a row, pivot category, combine near-misses, re-read the source, or try something more radical before concluding the hill is climbed. Correctness and simplicity outrank the number. Revert a win that breaks behavior, and keep a simplification that holds the number (the **laziness-protocol** principle skill).
18
+ 7. Stop when the predicate is met, or when the remaining ideas are marginal and not worth their cost. Don't relax the predicate to meet it, and don't quit while cheap untried hypotheses remain. If you are stuck, surface it instead of spinning.
19
+ 8. Run **Opening a PR** with the accepted commits stacked in the order they landed.
20
+
21
+ **Reply:** the metric and target, baseline to final with the percent delta, iterations run (kept vs reverted), each accepted fix on one line, the `decision.tsv` path, and the best idea you would try next if pushed further.
@@ -0,0 +1,14 @@
1
+ ### Investigation
2
+
3
+ **You own the answer. Plan, route, write.**
4
+
5
+ Investigation requests are read-only. They produce a cited explanation or a recommendation, not a code change.
6
+
7
+ 1. Route through the **how** skill. For motivation questions, also route through the **why** skill.
8
+ 2. Throughput checkpoint stays one line: `throughput checkpoint: n/a, read-only investigation`.
9
+ 3. Produce the `how`-shaped output (Overview / Key Concepts / How It Works / Where Things Live / Gotchas), or a recommendation with a tradeoffs table if the request is a decision between alternatives.
10
+ 4. Apply the **unslop** skill to the reply.
11
+
12
+ No PR, no babysit, no `architect` unless the investigation precedes a code change. If it does, hand back to the user and re-route to Bug fix or Feature.
13
+
14
+ **Reply:** the investigation output. For "are we sure?" answers, include your real judgment with reasons. Push back if the premise is wrong (see Autonomy).
@@ -0,0 +1,155 @@
1
+ ### Multi-phase or multi-PR plan
2
+
3
+ **You own the plan, not the code. The plan is a checklist an owner runs box by box and the operator audits from the evidence.** The plan is the deliverable. Do not implement.
4
+
5
+ 1. When the change is one or two files with an obvious approach, skip the plan. Say so and stop.
6
+ 2. Settle open questions by prototype before you write. Run `playbooks/prototype.md` for each. Keep the branch, the SHA, and the screenshots for Appendix A. Ask the operator only about a product or preference call that no run can settle. Give options (the **never-block-on-the-human** principle skill).
7
+ 3. Explore in subagents with `subagent_type: "poteto-agent"` (harness **poteto-agent** row) and an explicit model per the Subagents section (the **guard-the-context-window** principle skill). Each returns file pointers, conventions, test commands, and entry points. No inlined dumps.
8
+ 4. Copy the skeleton below into the plan file and fill every placeholder. Unless the operator names a path, write the file under the agent store's `docs/`. Keep every heading and every sub-block in the order shown. One section per PR. One PR is one change with its own evidence (the **sequence-verifiable-units** principle skill). Name the execution playbook in **How to read this**. Pick between `playbooks/autopilot-full.md` and `playbooks/autopilot-stack.md` per the rule at the end of `playbooks/autopilot-stack.md`. A standing program takes `playbooks/orchestrate.md`.
9
+ 5. Write under `/technical-writing` in full, then `/unslop`. The body is one Diátaxis mode, how-to. Appendices hold explanation and reference. Each heading states the task or the finding. No long dashes. No mid-sentence colons.
10
+ 6. Run `node pstack/skills/poteto-mode/scripts/check-plan.mjs <plan.md>` and fix every line it prints (the **encode-lessons-in-structure** principle skill).
11
+ 7. Hand back. Post the plan path and the script's output, then stop. Execution starts on the operator's explicit go, under the execution playbook the plan names.
12
+
13
+ **Verification.** Tests alone are not sufficient verification. A PR is verified only when its unit, live, and perf boxes are all checked (the **prove-it-works** principle skill). That sentence is the verification rule. Every verification block opens with it. The live block is mandatory. Ten lanes at the PR head drive the real surface through `control-ui` for web, IDE, or Electron, or `control-cli` for CLI/TUI (harness **verification harness** row), per the **swarm** skill, on the `swarm workers` model (default the code tier). Each lane is one box with a concrete scenario, the screenshot it saves, and its pass predicate. One lane is the **Regression lane against trunk.** It runs the same load-bearing scenario on trunk and head. If trunk does not have the feature, the lane records that fact and gates the behavior the diff adds plus the end state the user waits for instead of inventing a trunk result. The perf gate is dual-sided. Trunk and head must both produce the named metric. If trunk lacks the feature, also isolate the work the diff adds and set an absolute budget for that work plus the end-to-end state the user waits for. Do not claim a ratio between unlike scenarios. The perf block names the metric, the interleaved probe, the trunk baseline measured first, and the rule with the number that fails. A PR that changes an interaction is review-gated. The operator reviews it in chat with screenshots and a video before merge. A PR that changes no interaction writes `**Review gate.** None. <PR id> is not review-gated.` and no boxes under it.
14
+
15
+ **Verification driver.** Load bundled `control-ui` for browser, Electron, and web UIs, or `control-cli` for CLI/TUI (harness **verification harness** row). Use the project's `verify-<app>` skill for app-specific commands and feature coverage. Native mobile uses whatever simulator-driving skill the repo has. A PR that touches two surfaces gets lanes on both. A surface with no available verification driver is a risk in Appendix C, and its live block still names how each lane drives it.
16
+
17
+ ````markdown
18
+ # <Program> plan
19
+
20
+ <Under ten lines. What changes, for whom, the rule the program enforces, and the PR ids in order.>
21
+
22
+ ## How to read this
23
+
24
+ One box is one unit of work. Every box names the evidence that checks it. A nested box is a sub-step of the box above it. Check a box only when its evidence exists, a file, a log line, a screenshot, a test run, or a SHA. The body is a how-to. The appendices explain and record.
25
+
26
+ The program runs `pstack/skills/poteto-mode/playbooks/<execution playbook>.md`. <Who merges, and which PR ids are the operator's items that stop at merge-ready.>
27
+
28
+ Tests alone are not sufficient verification. A PR is verified only when its unit, live, and perf boxes are all checked.
29
+
30
+ ## Program checklist
31
+
32
+ ### Arm the program
33
+
34
+ - [ ] State the protocol and this plan to the operator, then stop. Start execution only on the operator's explicit go.
35
+ - [ ] Read these from trunk at program start. Re-read them at every tick.
36
+ - [ ] `git show origin/main:pstack/skills/poteto-mode/playbooks/<execution playbook>.md`
37
+ - [ ] `git show origin/main:pstack/skills/swarm/SKILL.md`
38
+ - [ ] Load bundled `control-ui` for web/IDE/Electron or `control-cli` for CLI/TUI through the harness **skill** row. If a project verification skill exists, also read `git show origin/main:<project verification skill path>`.
39
+ - [ ] `git show origin/main:pstack/skills/poteto-mode/playbooks/opening-a-pr.md`
40
+ - [ ] `git show origin/main:pstack/skills/<each other leaf skill the program uses>`
41
+ - [ ] On the operator's go, arm the hourly audit tick (harness **loop** row, `/loop 1h` on Claude Code) with the tick prompt below. Never leave the cadence to memory.
42
+ - [ ] Use this tick prompt, verbatim. "Re-read the execution playbook from trunk. Audit the operation against it and fix drift in this tick. Probe every active lane and judge progress by side effects only. Stand down a stuck lane and dispatch its replacement now. Then post a short status message to the operator in chat only when the audit found a tracked change that no earlier status message reported, such as a PR opened, a code-ready head, a round launched or closed, a verdict, a merge, a stuck agent and the action taken, a blocker added or cleared, or a decision only the operator can make. Name every such change and nothing else. Do not repeat a table, the merged list, or an unchanged blocker. If the audit found none, end the turn with no reply text. Either way, log this tick's row in your decision trail. The row names the items reported, or none."
43
+ - [ ] On the operator's hold or stand-down, send every owner a zero-writes order at once.
44
+
45
+ ### Spawn owners
46
+
47
+ - [ ] Spawn one owner per PR with the full lifecycle the execution playbook names.
48
+ - [ ] Follow this dependency graph. Start dependent work only after its parent merges, or base it on the parent branch when the execution playbook stacks.
49
+ - [ ] <PR id> and <PR id> are independent and first. Both branch from `main`.
50
+ - [ ] <PR id> after <PR id>.
51
+ - [ ] Hold the file boundaries. <PR id or class> touches only `<glob>`.
52
+ - [ ] Hold the review gate. <PR ids> change an interaction. They wait for the operator's review in chat with screenshots and a video before merge.
53
+
54
+ ### PR mechanics, for every PR
55
+
56
+ - [ ] Resolve the forge once. Default to `gh`; if `command -v origin` succeeds and Origin can resolve the repository, use `origin pr` for every PR operation. Record any fallback to `gh`. Never require `gt`.
57
+ - [ ] Open the PR ready, never draft, per **Opening a PR**. Use the run's built-in PR tool when it has one, else `origin pr create --status open --base <base-branch>` or `gh pr create --base <base-branch>` according to the resolved forge. A stack child targets its parent branch.
58
+ - [ ] Run the repo's lint and typecheck once before the PR-facing push. Push with hooks on.
59
+ - [ ] Deslop the diff (harness **deslop** row) before each commit and run `/no-comments` before review.
60
+ - [ ] Triage every Bugbot and security-reviewer comment per `../references/bugbot-triage.md`.
61
+ - [ ] Rebase onto current trunk before the code-ready report and babysit. Keep that merge base in fix rounds. Rebase again only at merge prep, on a `git merge-tree` conflict with trunk, or on a CI failure that comes from a change on trunk.
62
+
63
+ ### Verdict and merge, for every PR
64
+
65
+ - [ ] At the code-ready head SHA and at each later push that changes the patch, run the swarm per `pstack/skills/swarm/SKILL.md`. One gates lane. The ten live lanes from the PR's **Verify, live** block. The perf lane from its **Verify, perf** block. Two or more audit lanes, each with its own focus, that read the diff and the receipts and distrust the PR body. The root audits the receipts in the merge-ready report before the verdict.
66
+ - [ ] Clean only when every lane is `PASS`. Findings go back to the owner, including a defect that a lane filed as a note. A new head gets a fresh swarm and a fresh verdict, except for results that stay valid under the patch-id rule in `playbooks/shipping.md`.
67
+ - [ ] <The merge or append rule from the execution playbook, with the patch-id rule from `playbooks/shipping.md`.>
68
+
69
+ ### Boot recipe, for every live lane
70
+
71
+ Each live lane runs in its own isolated checkout at the PR head (harness **isolation** row). Drive through the project's verification harness (harness **verification harness** row).
72
+
73
+ - [ ] `git fetch origin <head-branch> && git checkout <head SHA>`.
74
+ - [ ] <Start the backend and the surface. Wait for ready.>
75
+ - [ ] <Deliver input through control-ui for web/IDE/Electron or control-cli for CLI/TUI, using the project verification recipe when present. Name the read-only diagnostics.>
76
+ - [ ] Save every screenshot to `/tmp/swarm-<pr-id>/worker-<n>/<slug>.png` and return the paths with the report.
77
+
78
+ ## <Task as a verb phrase> (<PR id>)
79
+
80
+ **Depends on.** <PR id, or None.>
81
+
82
+ **Files.**
83
+
84
+ - [ ] Edit `<path>`.
85
+ - [ ] Create `<path>`.
86
+ - [ ] Delete `<path>`.
87
+
88
+ **Build.**
89
+
90
+ - [ ] <One change. Name the symbol and the file.>
91
+
92
+ **You see.**
93
+
94
+ - [ ] <One observable result, with the exact log line or screen state.>
95
+
96
+ **Verify, unit.** Tests alone are not sufficient verification. A PR is verified only when its unit, live, and perf boxes are all checked.
97
+
98
+ - [ ] <Test file and the case it gains.> Run `<command>`.
99
+
100
+ **Verify, live.** Tests alone are not sufficient verification. A PR is verified only when its unit, live, and perf boxes are all checked. Ten lanes on `<swarm workers model>` at the PR head, per the boot recipe.
101
+
102
+ - [ ] Lane 1. Regression lane against trunk. Run <the same load-bearing scenario> at trunk and head. If trunk lacks the feature, record that and gate <the behavior the diff adds plus the end state the user waits for>. Save `<slug>.png`. Pass when <predicate>.
103
+ - [ ] Lane 2. <Scenario.> Save `<slug>.png`. Pass when <predicate>.
104
+ - [ ] Lane 3. <Scenario.> Save `<slug>.png`. Pass when <predicate>.
105
+ - [ ] Lane 4. <Scenario.> Save `<slug>.png`. Pass when <predicate>.
106
+ - [ ] Lane 5. <Scenario.> Save `<slug>.png`. Pass when <predicate>.
107
+ - [ ] Lane 6. <Scenario.> Save `<slug>.png`. Pass when <predicate>.
108
+ - [ ] Lane 7. <Scenario.> Save `<slug>.png`. Pass when <predicate>.
109
+ - [ ] Lane 8. <Scenario.> Save `<slug>.png`. Pass when <predicate>.
110
+ - [ ] Lane 9. <Scenario.> Save `<slug>.png`. Pass when <predicate>.
111
+ - [ ] Lane 10. <Scenario.> Save `<slug>.png`. Pass when <predicate>.
112
+
113
+ **Verify, perf.** Tests alone are not sufficient verification. A PR is verified only when its unit, live, and perf boxes are all checked.
114
+
115
+ - [ ] Metric. <What is measured at both trunk and head. If trunk lacks the feature, also name the diff-added work and the end-to-end state the user waits for.>
116
+ - [ ] Probe. <The command or procedure, run at trunk and at the head, interleaved. Both sides must produce the metric.>
117
+ - [ ] Baseline. Record the trunk <value> first.
118
+ - [ ] Rule. <Head against trunk, with the number that fails. If the scenarios differ, add absolute budgets for the diff-added work and the user-visible end state instead of an invalid ratio.>
119
+
120
+ **Review gate.** The operator reviews before merge.
121
+
122
+ - [ ] Copy lane <n> screenshots into `<media path>/<pr-id>-review-<slug>.png`.
123
+ - [ ] Record a 30 to 60 second video of the change on a lane VM. Save it as `<media path>/<pr-id>-review.mp4`.
124
+ - [ ] Post the screenshots and the video in chat. Stop at merge-ready. Wait for the operator's click.
125
+
126
+ **Merge.**
127
+
128
+ - [ ] Root's clean verdict at the exact head SHA.
129
+ - [ ] Review-bot triage done.
130
+ - [ ] Rebased onto current trunk after the verdict, patch-id unchanged.
131
+ - [ ] <The owner squash-merges its own PR, or the root appends it to the base-branch stack and the operator lands it bottom-up.>
132
+
133
+ ## Close the program
134
+
135
+ - [ ] Every box above is checked with its evidence.
136
+ - [ ] Reply to the operator with the report the execution playbook names.
137
+
138
+ ## Appendix A. Prototype evidence
139
+
140
+ <Each open question a prototype answered, with the branch, the SHA, and the artifact links. Each question that stays unproven.>
141
+
142
+ ## Appendix B. Alternatives rejected
143
+
144
+ <Each approach weighed and why it lost.>
145
+
146
+ ## Appendix C. Risks
147
+
148
+ <Each risk with the PR it lands in and what the owner watches.>
149
+
150
+ ## Appendix D. Links and reading list
151
+
152
+ <Docs to read before editing. Which PRs get `pstack/skills/how/SKILL.md` and `pstack/skills/interrogate/SKILL.md`. The trail per `pstack/skills/show-me-your-work/SKILL.md`.>
153
+ ````
154
+
155
+ **Reply:** the plan path, the PR ids with their dependencies and the review-gated set, what the prototypes proved and what stays unproven, and the check script's output.
@@ -0,0 +1,38 @@
1
+ ### Opening a PR
2
+
3
+ Invoked at the end of every other playbook.
4
+
5
+ **Worktree.** Work from a git worktree off main. Subagents inherit it. Multiple spawns on the same branch each get their own worktree, or `git fetch && git reset --hard origin/<branch>` between them. Dirty branch with unrelated work: patch out, fresh worktree, apply. Snarled worktree: reset from main, redo minimally.
6
+
7
+ **Commits.** Commit liberally. Rebase into small, ordered commits before opening PRs. Each commit is a future PR: landable, ordered to tell the story. Amend when the fix belongs in a just-made commit. New commit when separable.
8
+
9
+ **UI proof.** For web, IDE, or Electron changes, load `control-ui` per the harness **verification harness** row and verify the user flow before marking it proven in the PR. For CLI/TUI changes, load `control-cli`. Use any project `verify-<app>` skill for app-specific commands. If that proof already exists for this exact final diff, cite it; cleanup that changes behavior needs fresh proof.
10
+
11
+ **PRs.** Deslop the diff before commit per the harness **deslop** row. Run `/no-comments` before review. Write every PR title, PR description, and commit body with `/technical-writing`, then apply `/unslop`. Apply every technical-writing layer except Diátaxis. Use one word for each action, keep articles, and avoid `-ing` when a plain verb works.
12
+
13
+ **Titles.** Use Conventional Commits in the form `type(scope): subject`. Use `feat`, `fix`, `docs`, `refactor`, `test`, `chore`, or `perf` as the type. Use the changed area, such as `pstack` or `poteto-mode`, as the scope. Keep the subject short and imperative. Name a real symbol when one carries the change. For example, `fix(pstack): retarget opening-a-pr babysit trigger`. Do not add a trailing period.
14
+
15
+ **Descriptions.** The PR body is a briefing, not the lab notebook. A reviewer who has the diff should learn why the change exists, what it leaves out, what it could break, and how you proved it works, in under a minute. Write short, simple sentences with few identifiers. Do not write walls of text. The squash commit body is the PR body. If the body would make the squash commit longer than about 40 lines, cut the body.
16
+
17
+ Put each section under a `##` heading, not a bold lead-in, so the sections stand apart. Use these sections in order. Drop a section when it has nothing to say.
18
+
19
+ - `## Why` gives the problem and the approach in one to three short sentences. Do not list SHAs or rebase genealogy. Do not add a "based on main" preamble.
20
+ - `## What changed` has one to three short bullets. Name a real symbol or path only when it carries the change. Name both sides of a rename or retarget.
21
+ - `## Scope` always names what the PR covers and what it deliberately leaves out, for example a related follow-up or a known gap. Use one to three short items. Do not list symbols or paths, and do not write a file-by-file essay.
22
+ - `## Tradeoffs` names only rejected alternatives that a reviewer would otherwise ask about. Skip this section when there was no real choice.
23
+ - `## Blast Radius` gives one or two sentences on who or what the change touches and why that is safe or risky. If main is red, state the cost of leaving it red.
24
+ - `## Verification` has one to three bullets. Each bullet names a real run path and its outcome. For a performance change, report one primary number with its unit in `before → after` form. Link the arena or swarm directory for the remaining evidence. Do not include sample-size methodology, swarm recitals, or metric tables.
25
+
26
+ After these sections, attach videos or screenshots when they prove a claim. Do not paste full SHAs, swarm or arena lane recitals, lever-correction essays, file-by-file checklists, or "CLEAN" verdicts. Put these details in a linked artifact. A commit body does not restate its subject.
27
+
28
+ **Forge.** Resolve the forge before the first PR operation and keep that choice for create, edit, view, watch, and merge. GitHub CLI (`gh`) is the default. If `command -v origin` succeeds and Origin can resolve the repository, prefer `origin pr ...`. If Origin is absent or cannot resolve the repository, stay on `gh` and record the fallback. Do not require Graphite (`gt`).
29
+
30
+ **Built-in PR tool.** When the run provides a built-in PR tool, create, edit, retarget, and mark ready through it, never through a forge CLI. Its own instructions say how. A PR made with the CLI misses what the tool tracks, such as a description later runs can edit. Use the resolved forge for everything the tool does not cover, and for every PR operation when the run has no such tool.
31
+
32
+ **Size and stacks.** Prefer five narrow PRs to one large PR. A stack is a base-branch chain. The root PR targets trunk. Each child branch rebases onto its parent's exact tip and its PR targets the parent branch. Without a built-in PR tool, create a child with `origin pr create --status open --base <parent-branch>` or `gh pr create --base <parent-branch>` according to the resolved forge, and retarget an existing child with `origin pr edit <pr> --base <parent-branch>` or `gh pr edit <pr> --base <parent-branch>`. Branch from trunk only for independent work. Rebase on trunk before substantial stack work.
33
+
34
+ **Readiness.** Open every PR ready, never as a draft. A built-in PR tool can default to draft, so set `draft: false` on every creation call through it. With Origin, pass `--status open`. With `gh`, omit `--draft`. If a PR still opens as a draft, mark it ready through the PR tool, or run `origin pr ready <number>` or `gh pr ready <number>` according to the resolved forge. Run `origin pr view <number>` or `gh pr view <number>` before you refer to PR status.
35
+
36
+ **Babysit.** Opening a PR does not start a babysit. Post the URL and keep building. Finish the phase or stack first. Run a separate babysit pass only when the user asks for one after the whole stack exists. A babysit for each new PR stalls the build and spends checks on commits that later waves restart. Push back when feedback drifts from intent.
37
+
38
+ A subagent that opens a PR runs `interrogate`, the harness **deslop** row, and `/no-comments`, and posts the URL. Then it returns to the parent without babysitting, unless it is an Autopilot-full or Autopilot-stack owner. That owner's brief assigns the babysit loop and is the ask `playbooks/babysit.md` waits for. The owner starts the loop after its code-ready report and reports merge-ready or STACK-READY as its playbook says. The rules here and in `playbooks/babysit.md` that hold babysitting until a whole stack is built do not apply to that owner.
@@ -0,0 +1,114 @@
1
+ ### Orchestrate
2
+
3
+ **You own the program, never the code. Author briefs, drain the queue, keep the frontier green, decide.** For a whole project handed to one standing coordinator chat: multi-day, many stacked PRs, dozens to hundreds of subagents, the human checking in twice a day instead of every five minutes. One task driven to a predicate is Autonomous run. One ambitious run needing a bespoke workflow is figure-it-out. Route here when the work outlives any single agent. Work one agent could finish inside the session's budget is not a program.
4
+
5
+ Ceremony must scale with the program. On cheap near-identical units, collapse it as each section directs.
6
+
7
+ Three rules carry the rest.
8
+
9
+ - Completions are queue events, not interrupts.
10
+ - Every spawn and every resume carries the standing orders verbatim.
11
+ - The brief is the product. A vague brief fails quietly, because a worker cannot ask you a question.
12
+
13
+ #### Roles and placement
14
+
15
+ - **Coordinator (this chat).** Local. Frames, authors briefs, drains the inbox, owns the human report, makes judgment calls. It never authors or edits code. Conflicted merges, restacks, and code changes are always tasks. Mechanically landing a verified unit (fast-forward or clean cherry-pick of a worker's commit, then push) is bookkeeping the coordinator may do itself on repos where local git is cheap. Queueing finished work behind an idle stacker is how a deadline harvests nothing. The loop is agentic end to end. Agents are spawned, resumed, and drained only through the spawn tool (harness **spawn** row). State reads and writes go through `scripts/orch/orch.ts` at drain points, one command in and one line out. The CLI never spawns, waits, or wakes anything.
16
+ - **Sub-coordinator.** Always local, durable, one per track, and only when the program exceeds what one coordinator's drains can manage. A track the coordinator can drain itself needs no middle layer. Each nested layer re-pays a full orientation preamble, and a blocking sub-coordinator hides its children while the parent idles. Owns its track's units and boards, authors its workers' briefs, spawns its own workers and verifiers (nesting depth is a harness limit; Codex caps it with `agents.max_depth`). Rolls up aggregates at wave boundaries. Never forwards raw child reports. Cap in-flight children at what one drain can process, roughly ten, as a rolling window. Never as blocking batches, which cost the slowest child of every batch.
17
+ - **Worker / verifier.** Always isolated per the harness **isolation** row unless the task needs this machine: runtime verification through the project's verification harness (harness **verification harness** row). Reading local transcripts (harness **transcripts** row). Simulators and local IDE state. Auth that exists only here. An isolated worker cannot read the coordinator's local store, so its brief inlines what it needs or points at repo paths. Prefer fewer, broader workers. One writer per worktree or branch (principle-separate-before-serializing-shared-state). Run a unit's verifier on a different model family from its worker.
18
+
19
+ Depth stays at coordinator, track, worker. Author the track decomposition per project (build, landing, and verification are common cuts, not a required shape). Hard-coded swarm trees were tried and parked as too rigid.
20
+
21
+ #### Store layout
22
+
23
+ Create `orchestrate/<project-slug>/` in the current agent's store (path in the system prompt). Every file has exactly one writer. Owners publish facts, readers aggregate at read time. Use `bun scripts/orch/orch.ts` for bookkeeping, written below as `orch`, while its canonical plain TSV and JSON stay readable without the CLI.
24
+
25
+ - `preferences.md` is the standing-orders register: numbered lines, one constraint each (model policy, stack shape and count, verification bar, forbidden paths, escalation policy). Paste it verbatim into every spawn and every resume. Directives decay across resumes, and each dropped one costs a human turn. When you catch yourself restating an instruction, append the line before you act (principle-encode-lessons-in-structure).
26
+ - `overview.md` is the durable PR and issue DB. Append. Never rewrite wholesale per event.
27
+ - `units.tsv` has one row per unit: id, track, state, branch, PR, head SHA, brief path. Update rows in place.
28
+ - `frontier.json` is the computed merge frontier, per Stack safety.
29
+ - `ledger.tsv` is the verification ledger, per Verification.
30
+ - `inbox/` holds completion pointers. `gates.md` parks human gates (question, options, default on no answer).
31
+ - `decisions.tsv` is the trail via the show-me-your-work skill.
32
+ - `status.md` is derived from `units.tsv` and `ledger.tsv` at each drain, never hand-maintained. Regenerate it from the tables instead of narrating events into it.
33
+
34
+ #### The brief
35
+
36
+ Your prompts to agents are your only product, and a sloppy brief compounds into slop across the whole tree. Every spawn carries all of it. A field you cannot fill is a unit you have not scoped yet.
37
+
38
+ Code-writing briefs require loading bundled `deslop` before each commit and final code handoff (harness **deslop** row). UI verification briefs require loading `control-ui`; CLI/TUI briefs require `control-cli` (harness **verification harness** row). Include any project verification skill for app-specific steps. A skill call loads and follows its `SKILL.md`, even when the worker cannot invoke a slash command.
39
+
40
+ ```
41
+ GOAL one sentence, the outcome, executable by a stranger with no chat access
42
+ SCOPE paths this unit may write; paths it may not; its exclusive worktree or branch
43
+ CONTEXT pointers to files and PRs; upstream reports pasted in full when this unit
44
+ depends on them, because workers cannot see siblings
45
+ ACCEPTANCE checkable criteria, one per line
46
+ VERIFY control-ui for web/IDE/Electron or control-cli for CLI/TUI; project verify
47
+ skill path and exact commands when present, plus known gotchas
48
+ TIMEBOX rough cap on runtime; on expiry, return partial findings and stop rather than run on
49
+ FORBIDDEN no gt, no rebase, no force-push, no fixes outside scope, plus unit-specific bans
50
+ REPORT status, branch, head SHA, PRs, verdict, what you actually ran, deviations,
51
+ suggested follow-ups
52
+ STANDING <preferences.md pasted verbatim>
53
+ ```
54
+
55
+ Size the brief to the unit. A one-command unit gets the template collapsed to a paragraph that still names goal, scope, the verify command, and the report shape. A 4KB scaffold around a two-line edit costs more to write and obey than the edit. Local spawns may reference the standing-orders file by store path. Verbatim paste is for cloud spawns and every resume.
56
+
57
+ A sub-coordinator brief adds its track boundary and unit list, its spawn budget with the cloud default and the local exception list, the drain protocol, and the rollup format (per child: name, status, PR, head SHA, verdict, one line, plus track status and frontier delta).
58
+
59
+ A dependency is a context relay, not just ordering. Undeclared upstream context makes the worker guess. Missing fields are a refuse-to-spawn condition. Audit one sampled worker brief per sub-coordinator per wave, concurrently with the wave it samples, never as a gate in front of it. A failing brief stops that track and fixes the sub-coordinator's instructions, not just the worker, because brief quality decays late in a run. Never resume-chain a brief. Respawn fresh with consolidated scope.
60
+
61
+ #### Steps
62
+
63
+ 1. **Frame.** State the done predicate as something countable ("all 126 units merged, each ledger-verified `unit-test-verified` or better"). Quantify scope: units, rough effort, expected stacks, and the wall-clock budget. If one agent could finish inside that budget, stop here and run Autonomous run instead. Collapsing must not depend on another document being present. It means do the work directly in this session, plain workers where they help, verification inline, landing as you go, and none of the store, register, or pilot machinery below. Schedule landing against the budget. By roughly 70% of it, stop spawning and land what is verified. Name the tracks per project. A contested decomposition or one-way door goes through the arena skill before the pilot. Present the framing once. Reversible prep proceeds without waiting.
64
+ 2. **Install the runtime.** Run `orch init`. Open the trail via the show-me-your-work skill, write the standing orders before any spawn, and seed `frontier.json` from existing PRs with `orch frontier set --repo <repo-dir>`.
65
+ 3. **Pilot.** Push one unit through the whole path: brief, worker, verification, stack entry, ledger row, merge. The pilot exists to falsify the brief template, the verify recipe, and the unit size while that costs one agent instead of fifty. Fix the contract from pilot evidence before any fan-out. Scale the pilot to the unit. On programs of near-identical cheap units, the first unit is the pilot, run as a normal unit with its verify command inline, and fan-out starts the moment it lands. The dedicated pilot pipeline (separate verifier agent, audit gate) is for expensive or novel unit shapes, not for clone-units where a serialized pilot has nothing to falsify.
66
+ 4. **Scale.** Spawn a rolling window of workers up to the in-flight cap, refilling as children finish. Blocking batches pay the slowest child of every batch. Spawn track sub-coordinators only past the one-drain threshold in Roles. Recompute ready work after each drain. Relay upstream reports into downstream briefs. Keep sibling communication upward only. The sampled brief audit runs alongside the wave it samples and stops the next refill on failure, not the current one.
67
+ 5. **Drain.** Run the queue discipline below at every drain point.
68
+ 6. **Land.** Landing is continuous, never a terminal phase. Integration starts with the first verified unit and runs alongside the remaining waves. On heavy repos the stacker is a standing role from wave one, integrating as units verify. On repos where local git is cheap, the coordinator lands verified units itself per Roles. Keep the frontier green before upper-stack work. Stack safety governs. Advance `frontier.json` only on merge or reported new head SHAs.
69
+ 7. **Close.** Drain the final inbox, reconcile every spawned agent to a terminal row (done, abandoned, zombie-reconciled), confirm the predicate on the real artifact, confirm every landed PR has a verdict for its current head SHA, audit the trail per show-me-your-work including its cross-model review, encode recurring corrections into `preferences.md` or the brief template. Leave the store intact. It is the postmortem.
70
+
71
+ #### Queue and drain
72
+
73
+ - On a completion notification, run `orch inbox push <agent> <unit> <status> [--report PATH]` and return to what you were doing. Never deep-review inline. A completion that needs review becomes a verifier unit. Never review a diff inside a drain.
74
+ - Drain in batches at four points: the end of a critical section, a track rollup, a frontier watcher wake (arm it via the loop skill, with a long heartbeat fallback), and before a human report. Begin each batch with `orch inbox drain`. Arrivals during a drain wait for the next one.
75
+ - Critical sections you finish first: authoring a brief, a stack operation, a conflict decision, writing a gate, updating ledger or frontier.
76
+ - Each drain classifies every pointer (landed, needs-verify, failed, zombie, noise), writes the resulting rows through `orch unit add`, `orch unit set`, and `orch ledger record`, runs `orch status`, then spawns the next wave in one message.
77
+ - Account for every spawned child at its track's rollup: arrived, respawned, or its scope explicitly absorbed. Silently redoing a missing child's work hides both the wasted spend and the coverage gap its result existed to close.
78
+ - A drain turn ends with the three lines from `orch status`: counts against the states, what changed, gates open. Detail lives in `status.md`. The full reply contract applies at checkpoints and close.
79
+
80
+ #### Stack safety
81
+
82
+ - The frontier is a computed object, never narrative. Recompute `frontier.json` from `gt` after every merge and stack mutation because GitHub base refs drift mid-restack while gt tracking is authoritative: ordered PR list, branch names, head SHAs, a generation number, the lowest unmerged PR. Resolve it where gt knows the stack, normally the stacker's clone. A checkout whose gt metadata never saw the submits reports no PRs and the command errors rather than guessing.
83
+ - Exactly one stacker per stack may run `gt`, serialized within its stack. Record the holder in the standing orders. Restacks run in cloud. A local restack at this scale takes the laptop down.
84
+ - Workers never rebase and never run `gt`. Babysitters follow `playbooks/babysit.md`, one per stack, scoped to one immutable frontier generation. They report conflicts to the stacker rather than restacking.
85
+ - PR closes and retargets go through the stacker only. Closing a base PR orphans every chain above it. Merges and stack surgery are units with briefs like any other.
86
+ - One retro watcher follows merged PRs for reverts, post-merge CI breaks, and orphaned follow-ups.
87
+
88
+ #### Verification
89
+
90
+ Scale verification to the unit. When VERIFY is a single cheap command, the worker runs it and reports the output, and the coordinator spot-checks receipts. A dedicated verifier agent (on a different model family than the worker) is for units whose verification is expensive, judgment-laden, or high-blast-radius. A verifier agent whose entire product would be rerunning one command is ceremony, not verification.
91
+
92
+ Write ledger rows with `orch ledger record`. Check the current PR and head SHA with `orch ledger check`. `ledger.tsv`, one row per verdict, keyed by PR number plus head SHA: `live-ui-verified | unit-test-verified | type-check-only | verifier-blocked | verifier-failed`. CI green is an input to a verdict, not a verdict. Behavioral work needs better than `type-check-only`. `verifier-blocked` is not a pass. Respawn when the environment heals. `verifier-failed` gets a fix unit, not a re-verify. A worker may self-report. A verifier overrides it on the same key. A new head SHA voids the row, so re-verify after restack. The ledger answers "was this verified", not memory and not the transcript.
93
+
94
+ A unit is not done until its output is externalized the moment it lands, never batched to the end of the run. A worker pushes its branch, a verifier writes its ledger row, receipts land in the store. Work that exists only on one VM when that VM dies was never done.
95
+
96
+ #### Liveness and failure
97
+
98
+ - Never resume an agent to check on it. A resume restarts an idle agent. Probe read-only: the ledger, `units.tsv`, `gh`, pushed branches, the agent's status in the harness's agent list. Transcript mtime is not liveness.
99
+ - A silent death gets a synthetic postmortem row in the inbox (unit, failure mode, last evidence, options). Replan on evidence as it arrives. Never wait for full quiescence.
100
+ - Retry by mode: cap-hit or oom, respawn with smaller scope. Network-drop, retry as-is. Tool-error, retry on a different model. Unknown, retry once. Two retries, then abandon the unit and replan around it.
101
+ - A zombie that returns hours late reconciles against the current frontier and ledger before anything is accepted. Salvage unique findings through a fresh unit, never a blind merge.
102
+ - When continued spawning would produce garbage tree-wide (bad upstream output, broken acceptance, dead infra), write a stop line at the top of the standing orders, let in-flight work finish, fix the cause, clear it.
103
+ - Bound your own infra retries the same way you bound a child's. After a few consecutive tool aborts, stop retrying. Write a terminal handoff to durable state (what is done, where it lives, the exact command to resume) and end the run.
104
+ - After a harness restart: local agents are dead, cloud work is not. Re-read the standing orders and `units.tsv`, recompute the frontier, reattach cloud work by PR and branch rather than agent id, respawn one sub-coordinator per track from its stored brief plus current state, drain, resume. The dead session's store lock clears itself on the next write. `orch` replaces a lock whose holder pid is gone.
105
+
106
+ #### Escalation
107
+
108
+ Reaches the human, batched into the status page rather than per item: irreversible actions (force-push to shared branches, deploys, deletions, closing someone else's PR), genuine product or preference calls no experiment settles, a standing order that contradicts observed reality, a program-level dead end that survived a replan. Park each as a `gates.md` entry before asking, and route work around it.
109
+
110
+ Never reaches the human: frontier nudges, restack mechanics, retries, CI flake triage, review-thread triage, format fixes, scope the brief already forbids (refuse and continue), and "should I keep going". When in doubt, act and log.
111
+
112
+ Mid-run discoveries fix only what blocks the frontier. Everything else parks in follow-ups. At this fan-out a small scope leak multiplies into PRs nobody asked for.
113
+
114
+ **Reply:** at checkpoints and close: the predicate and the count against it from `units.tsv` and `ledger.tsv`, tracks and what each landed, the frontier (PR list plus SHAs), verdicts summary, what was abandoned and why, gates awaiting the human (the only asks), the store path, and the trail path. Numbers from the tables, not narrative. Include PR links.
@@ -0,0 +1,10 @@
1
+ ### Pause safely
2
+
3
+ **You own a clean stop. Leave a checkpoint a cold-start agent can resume from.** This is explicit only. On "keep going", "going to bed, keep going", or "don't stop", do not pause.
4
+
5
+ 1. Stop at a safe boundary. Finish the current atomic step or back out of it. Start nothing new, and cancel any nested subagents.
6
+ 2. Take no irreversible action to pause. No PR and no push unless you already had one out.
7
+ 3. Make the work durable. Commit uncommitted edits as one clear `wip:` commit on the current branch so nothing is lost. If the tree is broken, say so in the commit body in one line.
8
+ 4. Write the resume note off-context. Capture intent, what you were doing, progress and what's verified, current state, next steps, key files, and gotchas. For the compaction trigger write it to a file like `/tmp/<slug>-resume.md`. If a show-me-your-work trail exists, point at it instead of duplicating it.
9
+
10
+ **Reply:** where you are in the loop, what's on disk versus still in your head (paths, no diff dumps), the commits you made and whether the tree is clean, and the first action on resume. This is a pause, not a final report.
@@ -0,0 +1,25 @@
1
+ ### Perf issue
2
+
3
+ **You own the measurement story. Plan, review, verify the numbers.** Tie every fix to a measurement, don't read source instead of measuring.
4
+
5
+ 1. Load `control-ui` for web, IDE, or Electron, or `control-cli` for CLI/TUI, and capture a baseline trace (harness **verification harness** row). Vet the baseline, and each later number, with the **benchmark-checklist** skill.
6
+ 2. `how` to ground hypotheses. Don't claim a perf ceiling without running it first.
7
+ Try the performance mantras in order, cheapest first:
8
+ 1. Don't do it. Stop work whose result nothing uses rather than cheapening it.
9
+ 2. Do it, but don't do it again.
10
+ 3. Do it less.
11
+ 4. Do it later.
12
+ 5. Do it when they're not looking.
13
+ 6. Do it concurrently.
14
+ 7. Do it cheaper.
15
+
16
+ When an earlier mantra meets the target, stop.
17
+ 3. Plan the fix from the trace. If it crosses a function boundary, `architect` first. Delegate implementation to a subagent using your configured perf-issue model (default the code tier). Review the diff. Capture a post-fix trace.
18
+ Apply the **sequence-verifiable-units** principle skill, verifying each attempt before trying the next.
19
+ 4. Parse and compare the artifacts (JSON to sqlite, diff). "Inconclusive" or wrong-surface is not a pass. Flag it.
20
+ 5. Cite the measurement in the PR.
21
+ 6. Run **Opening a PR**.
22
+
23
+ For sustained improvement against a metric rather than a one-off fix, use the Hillclimb playbook (`playbooks/hillclimb.md`).
24
+
25
+ **Reply:** baseline number, post-fix number, delta, artifact path.
@@ -0,0 +1,14 @@
1
+ ### Prototype
2
+
3
+ **You own the design decision, not the code. The prototype is a throwaway instrument. The real build follows Feature.**
4
+
5
+ The one playbook where the Laziness Protocol's "smallest change" and the verification bar invert. Speed over polish, code quality does not matter, no planning. The rigor is in picking the right design cheaply. Propose variations the user didn't ask for, throw an approach away and try another.
6
+
7
+ 1. Scope the decision the prototype exists to make: which layout, which interaction, which density, or for an empirical fork which behavior, timing, or approach. No decision means no prototype. Route to Feature.
8
+ 2. Gather references when the design space is open. Search for prior art, summarize a moodboard of themes, palettes, and layouts, let the user pick directions before building. Skip when the direction is set.
9
+ 3. Build throwaway in an isolated scratch dir, separate from production source. For a visual decision, vanilla HTML/CSS/JS or the lightest stack that renders the idea, CDN deps, a dev server with hot reload. For a behavioral or timing decision, the smallest script that exercises the question. No production framework, no tests, no abstractions.
10
+ 4. When comparing alternatives, build them behind one switcher (buttons or a keypress), each variant labeled. This is the **exhaust-the-design-space** principle skill made cheap.
11
+ 5. Verify on the matching surface. For a visual decision, load `control-ui`, screenshot each variant, and drive the interaction (harness **verification harness** row). For a CLI/TUI decision, load `control-cli`. For a behavioral or timing decision, observe the thing you are deciding by logging the timing, printing the output, or watching the render. The observation is the test here, not an assertion.
12
+ 6. Present alternatives, tradeoffs, and a recommendation. The output is the decision plus the throwaway artifact, not shippable code. Hand the chosen direction to **Feature** (or `architect` for the shape) for the real build.
13
+
14
+ **Reply:** the variants explored, the evidence (screenshots for a visual decision, the observed output or timing for a behavioral one), tradeoffs, your recommendation, and the scratch path. Say plainly that the prototype is throwaway.
@@ -0,0 +1,16 @@
1
+ ### Refactoring
2
+
3
+ **You own the contract. The structure changes. The behavior does not.** Distinct from Feature, which adds behavior, and Bug fix, which corrects it.
4
+
5
+ If the cleanup reveals a missing feature or a real bug, split it out and ship the structural change first against the pinned contract. A redesign is allowed, but name it and route to Feature. Large or cross-cutting structural work belongs to the **figure-it-out** skill. This playbook is the focused-to-medium change.
6
+
7
+ 1. Pin the behavior contract first. Run the **how** skill over the affected subsystem to learn the contract, then write a characterization test, snapshot, or equivalence harness that captures current behavior before any structure moves. If the area has no coverage, write the pin before touching structure. Type check and lint are not a pin.
8
+ 2. Name the structure the code is missing per **principle-model-the-domain**. Boring code stays when the shape is already clear and local. The reshape must delete branches or invalid states, not add indirection.
9
+ 3. Name the target shape. State what the module layout, types, and call graph should be if built today (**principle-foundational-thinking**, **principle-redesign-from-first-principles**). If the target crosses a function boundary, run the **architect** skill for parallel design exploration of the shape before the move.
10
+ 4. Subtract before you add. Delete dead code, collapse one-caller wrappers, drop redundant validators, and remove orphan references before introducing the new shape (**principle-subtract-before-you-add**). The smallest change that reaches the target shape ships (**principle-laziness-protocol**). A speculative cleanup that "might help" gets reverted.
11
+ 5. Move in small behavior-preserving steps, each keeping the pin green. For API reshapes, migrate every caller and delete the old API in the same wave (**principle-migrate-callers-then-delete-legacy-apis**). No compatibility shims, no parallel old-and-new paths. Spot-check every rename against the actual files. Renames silently miss usages in strings, prose, and back-references. Delegate the mechanical edits to a subagent using your configured refactoring model (default the code tier) with a specific scope (file paths, the names being moved, the behavior to hold).
12
+ 6. Prove behavior is unchanged on the real artifact, not "it compiles" (**principle-prove-it-works**). For larger reshapes, run an equivalence check: a script that diffs old-vs-new outputs, a recorded baseline replayed against the new code, or a smoke run through `control-ui` for web, IDE, or Electron, or `control-cli` for CLI/TUI (harness **verification harness** row).
13
+ 7. Confirm the change is worth keeping. The success measure is reduced reader load (**principle-minimize-reader-load**). If the diff does not lower reader load somewhere, revert it.
14
+ 8. Rebase into small ordered commits. A subtraction commit, then the reshape, then any follow-on cleanup. Shape them with the **sequence-verifiable-units** principle skill, so each behavior-preserving slice stays green before the next. Run **Opening a PR**.
15
+
16
+ **Reply:** the structure that changed, the pin you held it against, the equivalence proof, the reader-load delta, what shipped and what got reverted. No new behavior.
@@ -0,0 +1,11 @@
1
+ ### Runtime forensics
2
+
3
+ **You own the diagnosis. Instrument the live process, don't theorize from source.** The deliverable is a cited diagnosis, not a fix.
4
+
5
+ 1. Load `control-ui` for web, IDE, or Electron, or `control-cli` for CLI/TUI, and capture the live signal (harness **verification harness** row): a CPU profile for a spinning process, a heap snapshot for a leak, a CDP trace for a visual glitch. A real artifact, not a guess.
6
+ 2. Reduce the artifact to the smoking gun: the function on the hot path, the retainer chain from the leaked object to a GC root, the loop firing without input. Parse large artifacts in a subagent (the **guard-the-context-window** principle skill), keep the reduced finding in the main thread.
7
+ 3. Prove the mechanism before believing it. Inject instrumentation via CDP eval on the running process, or hotfix the live code without reloading, to confirm the hypothesis cheaply.
8
+ 4. Map the finding back to source: file, symbol, the line that allocates or schedules.
9
+ 5. Throughput checkpoint stays one line: `throughput checkpoint: n/a, read-only forensics`.
10
+
11
+ **Reply:** the signal captured, the reduced finding, how you proved the mechanism, the source location, artifact paths. No fix unless asked. Hand back to Bug fix or Perf once the cause is known.
@@ -0,0 +1,11 @@
1
+ ### Session pickup
2
+
3
+ **You own the resume point. Read the prior trail, don't redo it.**
4
+
5
+ 1. Locate the prior trail. A local transcript in the active workspace's transcript directory (harness **transcripts** row, never another workspace's), a remote session URL, or a pushed branch. Read the metadata overview and last messages first, then scan back for the decision points. Parse a long transcript in a subagent and keep the reduced timeline in the main thread (the **principle-guard-the-context-window** skill).
6
+ 2. Reconstruct operational state. The branch and worktree, what already landed (`git log`, `git diff` against the base), the open todos, the decisions made. The prior trail is authoritative input. Resist the bias to re-derive it.
7
+ 3. Diff done vs pending. Compare what shipped against what was planned, name the resume point, do not re-run the prior repro or redo completed work. A "let me verify from scratch" pass means you're treating the trail as untrustworthy when it's authoritative.
8
+ 4. Route the remaining work to the matching playbook and pick the verdict: continue the execution, ship a finished recommendation, ratify or override a prior conclusion, or postmortem a failed run. The pickup playbook ends here. The routed playbook owns the rest.
9
+ 5. Verify the inherited claims against the original goal on the real artifact (the **principle-prove-it-works** skill). A passing prior self-report is not the proof.
10
+
11
+ **Reply:** where the prior agent stopped, what you inherited vs redid (ideally nothing redone), the resume point, and the outcome.
@@ -0,0 +1,17 @@
1
+ ### Shipping
2
+
3
+ **You own what lands. Verify each PR independently, land only the verified run from the root, then keep your hands off the queue.**
4
+
5
+ This is the half after `playbooks/babysit.md`.
6
+
7
+ 1. **Resolve the forge, then verify every PR independently.** GitHub CLI (`gh`) is the default. If `command -v origin` succeeds and Origin can resolve the repository, use `origin pr ...` for PR view, watch, edit, and merge operations. Otherwise stay on `gh` and record the fallback. Never require Graphite (`gt`). One subagent per PR, not batched, each isolated per the harness **isolation** row, each exercising the real surface through `control-ui` for web, IDE, or Electron, or `control-cli` for CLI/TUI, with any project verification recipe (harness **verification harness** row) against parent versus head. Each returns `PASS`, `PASS+NOTES` or `FAIL` and posts that verdict on its own PR. Safe means a verdict from an agent that did not write the code. CI green is not a verdict, and an approving bot review is not a verdict.
8
+ 2. **Land only the contiguous verified run rooted at the bottom.** Walk up from the lowest unmerged PR and stop at the first one without a passing verdict, where both `PASS` and `PASS+NOTES` pass. A verified PR sitting above an unverified one is not landable. Report the ceiling as a PR number and say what breaks the chain.
9
+ 3. **Re-check that each verdict still describes the patch.** Record the verdict head SHA, base SHA, and stable `git patch-id` of that PR's base-to-head diff. A rebase or base retarget rewrites SHAs and can silently invalidate a verdict without touching a check. Before landing a PR, compare the recorded patch-id with its current base-to-head patch-id. When the two patches differ only in tests, docs, or lint config, build what each lane ran. Build it twice at the verdict SHA and once at the current head. A difference is noise if the two builds at the verdict SHA also show it, or if it is an embedded commit SHA. Judge each difference, not each file, and report each kind of noise with its files. If only noise differs, that lane's result stays valid, and checks and a review of the change run fresh. Do not reuse a lane result from a dev server or from anything else with no build output. Rerun that lane. Re-verify anything else when the patch changed. When it did not, keep the code verdict but re-run mergeability and CI at the current head. Never use matching commit messages or a green check from an older SHA as a substitute.
10
+ 4. **Prepare only the bottom PR.** Fetch current trunk. Rebase the lowest verified branch onto the exact trunk tip when needed, push it, and retarget only that PR to trunk with `origin pr edit <pr> --base <trunk>` or `gh pr edit <pr> --base <trunk>`. Re-run step 3 after the push. Do not retarget, arm, or merge descendants yet.
11
+ 5. **Land one PR at a time.** If the bottom PR is mergeable now, squash it with `origin pr merge <pr> --squash` or `gh pr merge <pr> --squash`. If requirements are still running and the user asked for merge-when-ready, arm only that PR with `origin pr merge <pr> --squash --auto` or `gh pr merge <pr> --squash --auto`. Origin's `--auto` is Origin merge-when-ready. GitHub's `--auto` is GitHub auto-merge. Wait for that PR to merge before preparing the next one.
12
+ 6. **Do not read GitHub `autoMergeRequest` as stack readiness.** At most it says GitHub auto-merge was requested for one GitHub PR. It does not prove Origin merge-when-ready is armed, that a descendant is queued, that a patch verdict is current, or that the contiguous stack is safe. Confirm the active forge's state for the current bottom PR, and say that the state is unknown if the active forge cannot report it.
13
+ 7. **Recompute after every merge.** Fetch trunk, confirm the merged SHA is present, drop the merged PR from the frozen bottom-to-top list, and inspect the new bottom PR's base, head, checks, and patch-id. A host may retarget a child automatically, but do not assume it did. Repeat steps 3 through 6 for that one PR. Independent work stays outside this chain and ships on its own.
14
+ 8. **Watch the current frontier until it merges or fails. Do not mutate the queue around it.** With Origin, use `origin pr view <pr> --checks --comments` and `origin pr checks <pr> --watch`, then re-read the PR until it reports merged or blocked. With GitHub, use `scripts/watch-pr/watch-pr --queued-stack --stack-prs <bottom>` only as an event wake and poll `gh pr view <pr> --json state,mergedAt,mergeStateStatus,statusCheckRollup,autoMergeRequest` after each wake, ignoring `READY` until `mergedAt` is non-null or `state` is `MERGED`. Only then run step 7. Hard-fail only when `state` is `CLOSED` with no `mergedAt`, a required check concludes `FAILURE` or `CANCELLED` and blocks merge after auto-merge is no longer pending, or `mergeStateStatus` is `UNSTABLE` or `DIRTY` with no auto-merge pending. `BLOCKED` while checks are pending or auto-merge is armed is not failure. Do not use Babysit's queued `WAITING`/`merge-queue` stop condition here. Hold the watch under `/loop` in dynamic mode. Report each merge and the new ceiling. If the queue stalls, diagnose before mutating.
15
+ 9. **Stop at the ceiling.** When the verified run is merged, report what landed, what the next unverified PR is, and what verifying it would take. Extending the run is a new pass through step 1.
16
+
17
+ **Reply:** the verified run and its ceiling, each PR's verdict and who produced it, what you armed and how you confirmed it, what landed, and what the next gap needs.
@@ -0,0 +1,14 @@
1
+ ### Trace forensics
2
+
3
+ **You own the diagnosis from the artifact. Load it, shape it, narrow to the cause, attribute to source.**
4
+
5
+ Distinct from **Runtime forensics**, which instruments the live process. Here the capture already exists. The artifact is a fixed dataset, read it, don't re-run it. Keep tooling generic so the playbook stays portable: a DevTools or trace parser for cpuprofile and `.json.gz`, a text editor for a spindump, your heap tooling for a heapsnapshot.
6
+
7
+ 1. Identify the format and load it with the right tool. Parse large artifacts in a subagent (the **principle-guard-the-context-window** skill) and keep the reduced finding in the main thread.
8
+ 2. Transform the raw artifact into a form you can query. Dump the trace or heap snapshot into sqlite, one row per sample, frame, or node. Reach the queryable shape before you read.
9
+ 3. Narrow to the cause. Query for the frames that hold the most time and walk the call tree to the hot path. For a leak, follow the retainer chain from the leaked object to a GC root. For a spindump, find the thread stuck on-CPU or blocked and its wait reason.
10
+ 4. Attribute to source. Map the hot frame to file, symbol, and line via the artifact's own symbols. A frame with no source mapping is not yet a diagnosis. Resolve the symbols, or say plainly the artifact does not carry them.
11
+ 5. Confirm against a paired capture when you have one. Diff a before and after artifact. Without one, mark the finding as the strongest hypothesis the artifact supports, not a confirmed cause.
12
+ 6. Hand back a cited diagnosis, no fix unless asked. Route to Bug fix or Perf issue once the cause is known. Throughput checkpoint stays one line: `throughput checkpoint: n/a, read-only forensics`.
13
+
14
+ **Reply:** the artifact and format, the reduced finding, the source location, the artifact paths, and whether a paired capture confirmed it.
@@ -0,0 +1,11 @@
1
+ ### Visual parity
2
+
3
+ **You own pixel-exact equivalence. The baseline is the spec. You do not touch it.** Equivalence is verified by image diff, not by eye.
4
+
5
+ 1. Load `control-ui` per the harness **verification harness** row. Establish the baseline first, before any migration: a visual regression harness that screenshots the current component across its states, plus the target when matching two implementations. No baseline, no parity claim. A blocking prerequisite, not a follow-up.
6
+ 2. Anti-shortcut clauses, stated and held: no harness modifications, no baseline tampering, no component restructuring to make a diff pass. If the baseline looks wrong, stop and ask, don't edit it.
7
+ 3. Migrate one component at a time. Parallelize across worktrees, one owner per component (the **separate-before-serializing-shared-state** principle skill). Shared primitives migrate first as a blocking phase.
8
+ 4. Verify each component against its baseline via image diff on the matching surface through `control-ui`. A nonzero diff is a fail. Investigate the pixel delta. Loop per component (harness **loop** row) until the diff is zero.
9
+ 5. Run **Opening a PR** per component or per safe batch.
10
+
11
+ **Reply:** components migrated, the diff result for each, the baseline harness location, what's left.
@@ -0,0 +1,14 @@
1
+ ### Worktree and simulator cleanup
2
+
3
+ **You own the disk and the safety gate.** Prune merged or abandoned git worktrees and stale iOS simulators to reclaim space. Deletion is irreversible, so every step guards against deleting something in use or holding uncommitted work.
4
+
5
+ 1. Snapshot and audit. Record `df -h /`, then run `scripts/worktree-audit.sh` (principle-build-the-lever). It reads paths from `git worktree list`, never hand-typed, since a hand-typed `myrepo-worktrees/x` misses one that lives at `.claude/worktrees/x` or `~/.codex/worktrees/<hash>/myrepo` (principle-encode-lessons-in-structure). It classifies each worktree by size, age, merge state, uncommitted work, PR state, and the newest chat that touched it, then suggests a bucket. The transcript scan is slow, so background it.
6
+ 2. The bucket is advice, not permission. The pinned and active chats are the real artifact (principle-prove-it-works). Get that set from the user or sidebar and cross-check every candidate. The lever has marked `safe` a worktree the user had pinned, so the pinned set wins.
7
+ 3. Verify usage before deleting. For every `verify-recent-chat` row, or anything you doubt, fan subagents out to read the transcripts and report whether the chat is pinned or ongoing and which worktrees it touches (principle-guard-the-context-window, transcripts are bulk). A pinned chat spawns arena and repro trees into sibling worktrees via background subagents, and those are in use even when their names never hit the sidebar.
8
+ 4. Pause on irreversible loss. `wip:N` is N tracked uncommitted edits. Show the diff and get a decision first, since removing a clean worktree is recoverable from its branch but uncommitted work is gone. `scratch:N` is untracked throwaway, safe to drop, but name the files. Per Autonomy, clean and merged and not-in-use proceeds. `wip` and in-use pause.
9
+ 5. Prune the confirmed set. Per path, `git worktree remove --force <path>`. If the dir survives on ignored build artifacts, `rm -rf` it, then `git worktree prune`. Branch refs survive, so no commits are lost. Confirm with `df -h /` and re-list.
10
+ 6. Simulators and other reclaimers. Simulators are usually the next-biggest win. `xcrun simctl --set testing delete all` (XCTestDevices clones), `xcrun simctl delete unavailable`, and `xcrun simctl runtime list` then `runtime delete <id>` for old runtimes. More when needed: Xcode `DerivedData` and `iOS DeviceSupport`, `~/Library/Application Support/Cursor` or `~/Library/Application Support/Claude` (`state.vscdb.backup`, and `snapshots/roots/<root>` where a `<root>` named for a folder you opened as a workspace balloons), package caches (pnpm, uv, brew, yarn). Clear only caches the user has not said to keep.
11
+
12
+ This is the one playbook that deletes user state with no code review to catch a slip, so the gates above are the review.
13
+
14
+ **Reply:** `df -h /` before and after with space reclaimed, the worktrees pruned, and a one-line reason for each held back (in-use by which chat, or uncommitted work).