@jiroamato/pstack 0.0.0-stage → 0.15.15

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (225) hide show
  1. package/LICENSE +21 -0
  2. package/README.md +78 -2
  3. package/bin/pstack.js +95 -0
  4. package/lib/install.js +103 -0
  5. package/lib/prompt.js +77 -0
  6. package/lib/targets.js +43 -0
  7. package/package.json +38 -5
  8. package/pstack/.claude-plugin/plugin.json +26 -0
  9. package/pstack/.codex-plugin/plugin.json +36 -0
  10. package/pstack/LICENSE +21 -0
  11. package/pstack/LICENSE-cursor-team-kit +21 -0
  12. package/pstack/NOTICE +8 -0
  13. package/pstack/README.md +300 -0
  14. package/pstack/agents/comment-sicko.md +34 -0
  15. package/pstack/agents/poteto-agent.md +10 -0
  16. package/pstack/automations/benny/FOR_AGENTS.md +92 -0
  17. package/pstack/automations/benny/README.md +28 -0
  18. package/pstack/automations/benny/skills/reproduce-and-fix-issues/SKILL.md +313 -0
  19. package/pstack/automations/benny/skills/reproduce-and-fix-issues/references/control-adapter.md +169 -0
  20. package/pstack/automations/benny/skills/reproduce-and-fix-issues/references/feature-map.example.md +205 -0
  21. package/pstack/automations/benny/skills/reproduce-and-fix-issues/references/verify-existing-fix.md +93 -0
  22. package/pstack/automations/benny/skills/setup-benny/SKILL.md +271 -0
  23. package/pstack/automations/benny/skills/triage-issue-reports/SKILL.md +240 -0
  24. package/pstack/automations/benny/skills/triage-issue-reports/references/routing.example.md +61 -0
  25. package/pstack/automations/benny/templates/configuration.example.yaml +84 -0
  26. package/pstack/automations/benny/templates/reproduce-automation-prompt.md +33 -0
  27. package/pstack/automations/benny/templates/triage-automation-prompt.md +39 -0
  28. package/pstack/codex/agents/comment-sicko.toml +36 -0
  29. package/pstack/codex/agents/poteto-agent.toml +11 -0
  30. package/pstack/docs/guide/01-setup.md +80 -0
  31. package/pstack/docs/guide/02-poteto-mode.md +131 -0
  32. package/pstack/docs/guide/03-understand.md +79 -0
  33. package/pstack/docs/guide/04-design.md +133 -0
  34. package/pstack/docs/guide/05-build-and-clean.md +83 -0
  35. package/pstack/docs/guide/06-verify-and-ship.md +130 -0
  36. package/pstack/docs/guide/07-overnight.md +120 -0
  37. package/pstack/docs/guide/08-principles.md +72 -0
  38. package/pstack/docs/guide/09-make-it-yours.md +100 -0
  39. package/pstack/docs/guide/10-recipes-and-pitfalls.md +156 -0
  40. package/pstack/docs/guide/README.md +38 -0
  41. package/pstack/skills/architect/SKILL.md +85 -0
  42. package/pstack/skills/architect/agents/openai.yaml +2 -0
  43. package/pstack/skills/architect/references/design-red-flags.md +57 -0
  44. package/pstack/skills/architect/references/rationale-template.md +35 -0
  45. package/pstack/skills/architect/references/runner-prompt.md +20 -0
  46. package/pstack/skills/arena/SKILL.md +75 -0
  47. package/pstack/skills/arena/agents/openai.yaml +2 -0
  48. package/pstack/skills/automate-me/SKILL.md +104 -0
  49. package/pstack/skills/automate-me/agents/openai.yaml +2 -0
  50. package/pstack/skills/benchmark-checklist/SKILL.md +39 -0
  51. package/pstack/skills/benchmark-checklist/agents/openai.yaml +2 -0
  52. package/pstack/skills/blast-radius/SKILL.md +52 -0
  53. package/pstack/skills/blast-radius/agents/openai.yaml +2 -0
  54. package/pstack/skills/bro/SKILL.md +7 -0
  55. package/pstack/skills/bro/agents/openai.yaml +2 -0
  56. package/pstack/skills/control-cli/SKILL.md +55 -0
  57. package/pstack/skills/control-cli/agents/openai.yaml +2 -0
  58. package/pstack/skills/control-ui/SKILL.md +72 -0
  59. package/pstack/skills/control-ui/agents/openai.yaml +2 -0
  60. package/pstack/skills/correct/SKILL.md +34 -0
  61. package/pstack/skills/correct/agents/openai.yaml +2 -0
  62. package/pstack/skills/create-verification-skill/SKILL.md +47 -0
  63. package/pstack/skills/create-verification-skill/agents/openai.yaml +2 -0
  64. package/pstack/skills/create-verification-skill/references/feature-map-example/README.md +47 -0
  65. package/pstack/skills/create-verification-skill/references/feature-map-example/create-note.md +39 -0
  66. package/pstack/skills/create-verification-skill/references/feature-map-example/search.md +45 -0
  67. package/pstack/skills/deslop/SKILL.md +30 -0
  68. package/pstack/skills/deslop/agents/openai.yaml +2 -0
  69. package/pstack/skills/figure-it-out/SKILL.md +55 -0
  70. package/pstack/skills/figure-it-out/agents/openai.yaml +2 -0
  71. package/pstack/skills/how/SKILL.md +58 -0
  72. package/pstack/skills/how/agents/openai.yaml +2 -0
  73. package/pstack/skills/how/references/explainer-prompt.md +55 -0
  74. package/pstack/skills/how/references/explorer-prompt.md +52 -0
  75. package/pstack/skills/interrogate/SKILL.md +111 -0
  76. package/pstack/skills/interrogate/agents/openai.yaml +2 -0
  77. package/pstack/skills/interrogate/references/code-quality-review.md +47 -0
  78. package/pstack/skills/interrogate/references/lead-judgment.md +58 -0
  79. package/pstack/skills/interrogate/references/reviewer-prompt.md +70 -0
  80. package/pstack/skills/interrogate/references/rubric.md +77 -0
  81. package/pstack/skills/maintain-verification-skill/SKILL.md +41 -0
  82. package/pstack/skills/maintain-verification-skill/agents/openai.yaml +2 -0
  83. package/pstack/skills/make-bot-ui/SKILL.md +289 -0
  84. package/pstack/skills/make-bot-ui/agents/openai.yaml +2 -0
  85. package/pstack/skills/no-comments/SKILL.md +24 -0
  86. package/pstack/skills/no-comments/agents/openai.yaml +2 -0
  87. package/pstack/skills/poteto-help/SKILL.md +156 -0
  88. package/pstack/skills/poteto-help/agents/openai.yaml +2 -0
  89. package/pstack/skills/poteto-help/references/prompting.md +51 -0
  90. package/pstack/skills/poteto-help/references/recipes.md +47 -0
  91. package/pstack/skills/poteto-mode/SKILL.md +143 -0
  92. package/pstack/skills/poteto-mode/agents/openai.yaml +2 -0
  93. package/pstack/skills/poteto-mode/playbooks/authoring-a-skill.md +12 -0
  94. package/pstack/skills/poteto-mode/playbooks/autonomous-run.md +13 -0
  95. package/pstack/skills/poteto-mode/playbooks/autopilot-full.md +13 -0
  96. package/pstack/skills/poteto-mode/playbooks/autopilot-stack.md +16 -0
  97. package/pstack/skills/poteto-mode/playbooks/babysit.md +29 -0
  98. package/pstack/skills/poteto-mode/playbooks/bug-fix.md +15 -0
  99. package/pstack/skills/poteto-mode/playbooks/eval.md +25 -0
  100. package/pstack/skills/poteto-mode/playbooks/feature.md +21 -0
  101. package/pstack/skills/poteto-mode/playbooks/hillclimb.md +21 -0
  102. package/pstack/skills/poteto-mode/playbooks/investigation.md +14 -0
  103. package/pstack/skills/poteto-mode/playbooks/multi-phase-plan.md +155 -0
  104. package/pstack/skills/poteto-mode/playbooks/opening-a-pr.md +38 -0
  105. package/pstack/skills/poteto-mode/playbooks/orchestrate.md +114 -0
  106. package/pstack/skills/poteto-mode/playbooks/pause-safely.md +10 -0
  107. package/pstack/skills/poteto-mode/playbooks/perf-issue.md +25 -0
  108. package/pstack/skills/poteto-mode/playbooks/prototype.md +14 -0
  109. package/pstack/skills/poteto-mode/playbooks/refactoring.md +16 -0
  110. package/pstack/skills/poteto-mode/playbooks/runtime-forensics.md +11 -0
  111. package/pstack/skills/poteto-mode/playbooks/session-pickup.md +11 -0
  112. package/pstack/skills/poteto-mode/playbooks/shipping.md +17 -0
  113. package/pstack/skills/poteto-mode/playbooks/trace-forensics.md +14 -0
  114. package/pstack/skills/poteto-mode/playbooks/visual-parity.md +11 -0
  115. package/pstack/skills/poteto-mode/playbooks/worktree-cleanup.md +14 -0
  116. package/pstack/skills/poteto-mode/references/bugbot-triage.md +142 -0
  117. package/pstack/skills/poteto-mode/scripts/bootstrap.ts +62 -0
  118. package/pstack/skills/poteto-mode/scripts/bun.lock +67 -0
  119. package/pstack/skills/poteto-mode/scripts/check-plan.mjs +185 -0
  120. package/pstack/skills/poteto-mode/scripts/orch/orch.test.ts +634 -0
  121. package/pstack/skills/poteto-mode/scripts/orch/orch.ts +578 -0
  122. package/pstack/skills/poteto-mode/scripts/orch/store.ts +1607 -0
  123. package/pstack/skills/poteto-mode/scripts/package.json +16 -0
  124. package/pstack/skills/poteto-mode/scripts/watch-pr/cli.test.ts +224 -0
  125. package/pstack/skills/poteto-mode/scripts/watch-pr/cli.ts +223 -0
  126. package/pstack/skills/poteto-mode/scripts/watch-pr/fakes.test-helper.ts +118 -0
  127. package/pstack/skills/poteto-mode/scripts/watch-pr/github.test.ts +306 -0
  128. package/pstack/skills/poteto-mode/scripts/watch-pr/github.ts +699 -0
  129. package/pstack/skills/poteto-mode/scripts/watch-pr/policy.test.ts +420 -0
  130. package/pstack/skills/poteto-mode/scripts/watch-pr/policy.ts +832 -0
  131. package/pstack/skills/poteto-mode/scripts/watch-pr/render.ts +169 -0
  132. package/pstack/skills/poteto-mode/scripts/watch-pr/tsconfig.json +13 -0
  133. package/pstack/skills/poteto-mode/scripts/watch-pr/types.compile.ts +93 -0
  134. package/pstack/skills/poteto-mode/scripts/watch-pr/types.ts +401 -0
  135. package/pstack/skills/poteto-mode/scripts/watch-pr/watch-pr +6 -0
  136. package/pstack/skills/poteto-mode/scripts/worktree-audit.sh +92 -0
  137. package/pstack/skills/principle-attack-the-premise/SKILL.md +23 -0
  138. package/pstack/skills/principle-attack-the-premise/agents/openai.yaml +2 -0
  139. package/pstack/skills/principle-boundary-discipline/SKILL.md +34 -0
  140. package/pstack/skills/principle-boundary-discipline/agents/openai.yaml +2 -0
  141. package/pstack/skills/principle-build-the-lever/SKILL.md +23 -0
  142. package/pstack/skills/principle-build-the-lever/agents/openai.yaml +2 -0
  143. package/pstack/skills/principle-encode-lessons-in-structure/SKILL.md +31 -0
  144. package/pstack/skills/principle-encode-lessons-in-structure/agents/openai.yaml +2 -0
  145. package/pstack/skills/principle-exhaust-the-design-space/SKILL.md +21 -0
  146. package/pstack/skills/principle-exhaust-the-design-space/agents/openai.yaml +2 -0
  147. package/pstack/skills/principle-experience-first/SKILL.md +19 -0
  148. package/pstack/skills/principle-experience-first/agents/openai.yaml +2 -0
  149. package/pstack/skills/principle-explain-the-number/SKILL.md +23 -0
  150. package/pstack/skills/principle-explain-the-number/agents/openai.yaml +2 -0
  151. package/pstack/skills/principle-fix-root-causes/SKILL.md +23 -0
  152. package/pstack/skills/principle-fix-root-causes/agents/openai.yaml +2 -0
  153. package/pstack/skills/principle-foundational-thinking/SKILL.md +21 -0
  154. package/pstack/skills/principle-foundational-thinking/agents/openai.yaml +2 -0
  155. package/pstack/skills/principle-guard-the-context-window/SKILL.md +16 -0
  156. package/pstack/skills/principle-guard-the-context-window/agents/openai.yaml +2 -0
  157. package/pstack/skills/principle-laziness-protocol/SKILL.md +18 -0
  158. package/pstack/skills/principle-laziness-protocol/agents/openai.yaml +2 -0
  159. package/pstack/skills/principle-make-operations-idempotent/SKILL.md +24 -0
  160. package/pstack/skills/principle-make-operations-idempotent/agents/openai.yaml +2 -0
  161. package/pstack/skills/principle-migrate-callers-then-delete-legacy-apis/SKILL.md +22 -0
  162. package/pstack/skills/principle-migrate-callers-then-delete-legacy-apis/agents/openai.yaml +2 -0
  163. package/pstack/skills/principle-minimize-reader-load/SKILL.md +23 -0
  164. package/pstack/skills/principle-minimize-reader-load/agents/openai.yaml +2 -0
  165. package/pstack/skills/principle-model-the-domain/SKILL.md +26 -0
  166. package/pstack/skills/principle-model-the-domain/agents/openai.yaml +2 -0
  167. package/pstack/skills/principle-never-block-on-the-human/SKILL.md +20 -0
  168. package/pstack/skills/principle-never-block-on-the-human/agents/openai.yaml +2 -0
  169. package/pstack/skills/principle-outcome-oriented-execution/SKILL.md +21 -0
  170. package/pstack/skills/principle-outcome-oriented-execution/agents/openai.yaml +2 -0
  171. package/pstack/skills/principle-prove-it-works/SKILL.md +22 -0
  172. package/pstack/skills/principle-prove-it-works/agents/openai.yaml +2 -0
  173. package/pstack/skills/principle-redesign-from-first-principles/SKILL.md +16 -0
  174. package/pstack/skills/principle-redesign-from-first-principles/agents/openai.yaml +2 -0
  175. package/pstack/skills/principle-separate-before-serializing-shared-state/SKILL.md +16 -0
  176. package/pstack/skills/principle-separate-before-serializing-shared-state/agents/openai.yaml +2 -0
  177. package/pstack/skills/principle-sequence-verifiable-units/SKILL.md +17 -0
  178. package/pstack/skills/principle-sequence-verifiable-units/agents/openai.yaml +2 -0
  179. package/pstack/skills/principle-subtract-before-you-add/SKILL.md +21 -0
  180. package/pstack/skills/principle-subtract-before-you-add/agents/openai.yaml +2 -0
  181. package/pstack/skills/principle-test-behavior-not-implementation/SKILL.md +25 -0
  182. package/pstack/skills/principle-test-behavior-not-implementation/agents/openai.yaml +2 -0
  183. package/pstack/skills/principle-type-system-discipline/SKILL.md +31 -0
  184. package/pstack/skills/principle-type-system-discipline/agents/openai.yaml +2 -0
  185. package/pstack/skills/pstack-harness/SKILL.md +67 -0
  186. package/pstack/skills/recall/SKILL.md +35 -0
  187. package/pstack/skills/recall/agents/openai.yaml +2 -0
  188. package/pstack/skills/reflect/SKILL.md +76 -0
  189. package/pstack/skills/reflect/agents/openai.yaml +2 -0
  190. package/pstack/skills/reflect/references/divergent-reviewer.md +43 -0
  191. package/pstack/skills/reflect/references/judgment-reviewer.md +42 -0
  192. package/pstack/skills/reflect/references/synthesizer.md +56 -0
  193. package/pstack/skills/reflect/references/tooling-reviewer.md +55 -0
  194. package/pstack/skills/setup-pstack/SKILL.md +110 -0
  195. package/pstack/skills/show-me-your-work/SKILL.md +82 -0
  196. package/pstack/skills/show-me-your-work/agents/openai.yaml +2 -0
  197. package/pstack/skills/show-me-your-work/references/decision-log-template.tsv +1 -0
  198. package/pstack/skills/show-me-your-work/scripts/log.sh +42 -0
  199. package/pstack/skills/swarm/SKILL.md +48 -0
  200. package/pstack/skills/swarm/agents/openai.yaml +2 -0
  201. package/pstack/skills/tdd/SKILL.md +44 -0
  202. package/pstack/skills/tdd/agents/openai.yaml +2 -0
  203. package/pstack/skills/teach/SKILL.md +21 -0
  204. package/pstack/skills/teach/agents/openai.yaml +2 -0
  205. package/pstack/skills/technical-writing/SKILL.md +106 -0
  206. package/pstack/skills/technical-writing/agents/openai.yaml +2 -0
  207. package/pstack/skills/typescript-best-practices/SKILL.md +31 -0
  208. package/pstack/skills/typescript-best-practices/agents/openai.yaml +2 -0
  209. package/pstack/skills/typescript-best-practices/references/patterns.md +324 -0
  210. package/pstack/skills/unslop/SKILL.md +67 -0
  211. package/pstack/skills/unslop/agents/openai.yaml +2 -0
  212. package/pstack/skills/why/SKILL.md +158 -0
  213. package/pstack/skills/why/agents/openai.yaml +2 -0
  214. package/pstack/skills/why/references/epistemics.md +144 -0
  215. package/pstack/skills/why/references/investigator-prompt.md +103 -0
  216. package/pstack/skills/why/references/source-playbook.md +17 -0
  217. package/pstack/skills/why/references/sources/code-archaeology.md +88 -0
  218. package/pstack/skills/why/references/sources/databricks.md +70 -0
  219. package/pstack/skills/why/references/sources/datadog.md +99 -0
  220. package/pstack/skills/why/references/sources/incident-postmortem.md +15 -0
  221. package/pstack/skills/why/references/sources/linear.md +48 -0
  222. package/pstack/skills/why/references/sources/notion.md +55 -0
  223. package/pstack/skills/why/references/sources/sentry.md +100 -0
  224. package/pstack/skills/why/references/sources/slack.md +54 -0
  225. package/pstack/skills/why/references/synthesizer-prompt.md +135 -0
@@ -0,0 +1,313 @@
1
+ ---
2
+ name: reproduce-and-fix-issues
3
+ description: Reproduce triaged Slack bugs through a configured app-control adapter, verify existing fixes, and open a bounded draft pull request only after before-and-after proof. Use only from the configured Benny repro automation.
4
+ disable-model-invocation: true
5
+ ---
6
+
7
+ # Reproduce and fix issues
8
+
9
+ Wait for a trusted triage marker in the source thread. Reproduce the exact symptom through the target app's real UI. Verify an existing fix when one exists. Attempt a bounded fix only after a confirmed repro.
10
+
11
+ Load the external Benny configuration supplied by the automation. If the config, required actions, control adapter, or completed feature map is missing, fail closed.
12
+
13
+ ## Hard safety rules
14
+
15
+ - Freeze the source channel and root thread coordinates before doing any work.
16
+ - Never post a root message in the source channel.
17
+ - Preflight the source parent before every source-thread post.
18
+ - The coordinator is the only Slack poster.
19
+ - Delegated analysis workers are read-only and return findings or media notes.
20
+ - A fix-phase code worker may edit only when its environment provably excludes Slack credentials and every Slack write action. Otherwise the coordinator edits.
21
+ - Every child prompt must explicitly forbid `SendSlackMessage`, `PostToSlack`, `chat.postMessage`, and all other Slack writes.
22
+ - Never give a child a Slack token, posting instructions, source coordinates for posting, or permission to report externally.
23
+ - If a child needs Slack write access to run, do not launch it.
24
+ - Utility bots are evidence sources. They do not own the fix unless a person explicitly delegated the fix to them.
25
+ - The exact discriminating symptom must appear twice through real UI interaction.
26
+ - State inspection may confirm an observation. It must not inject or force the symptom.
27
+ - No confirmed repro means no authored fix.
28
+ - Existing pull requests or commits switch the run to verify mode. Do not author over them.
29
+ - Use `github.com` pull request links.
30
+ - Keep captures, recordings, logs, and tokens out of source control.
31
+ - Use pstack's `principle-guard-the-context-window` for delegated analysis.
32
+ - Apply pstack's `principle-sequence-verifiable-units`, `principle-fix-root-causes`, and `principle-prove-it-works` through repro, fix, and verification.
33
+
34
+ ## 1. Freeze source coordinates
35
+
36
+ Before making a work list or delegating:
37
+
38
+ 1. Require the trigger channel to equal the configured source channel.
39
+ 2. Set `SOURCE_THREAD_TS` to `trigger.thread_ts` when present. Otherwise use `trigger.ts`.
40
+ 3. Require a nonempty `SOURCE_THREAD_TS`.
41
+ 4. Store `SOURCE_CHANNEL_ID` and `SOURCE_THREAD_TS` as immutable values.
42
+ 5. Read the source thread and verify its root has those exact coordinates.
43
+ 6. Fetch the source permalink.
44
+
45
+ Never replace these values with a reply timestamp, operations timestamp, or status-message timestamp.
46
+
47
+ Before every source-channel post:
48
+
49
+ 1. Read the thread by the immutable coordinates.
50
+ 2. Confirm the parent exists, is not deleted, and still belongs to the source channel.
51
+ 3. Send only with `channel=SOURCE_CHANNEL_ID` and `thread_ts=SOURCE_THREAD_TS`.
52
+ 4. Read the thread again and verify the new message is a reply.
53
+
54
+ If any check fails, post nothing. Never retry at the root or in a fallback channel.
55
+
56
+ ## 2. Wait for the triage contract
57
+
58
+ Watch the source thread for the configured verdict budget. Stay silent while waiting.
59
+
60
+ Accept a verdict only when:
61
+
62
+ - Its author matches `slack.triage_identity_user_id`.
63
+ - It is a reply under `SOURCE_THREAD_TS`.
64
+ - It contains exactly one configured marker.
65
+
66
+ Public marker forms:
67
+
68
+ ```text
69
+ [benny:bug]
70
+ [benny:bug] tracker=https://tracker.example/issue/123
71
+ [benny:performance]
72
+ [benny:performance] tracker=https://tracker.example/issue/123
73
+ [benny:other]
74
+ ```
75
+
76
+ Proceed only for `bug` or `performance`. Capture the optional tracker URL. Stop silently for `other`, a missing verdict, an untrusted author, conflicting markers, or a timeout.
77
+
78
+ This marker replaces private bot identities and free-form verdict matching.
79
+
80
+ ## 3. Apply ownership and fix-artifact gates
81
+
82
+ Re-read the thread immediately before starting work.
83
+
84
+ ### Someone is explicitly fixing it
85
+
86
+ Stop when a person clearly claims the fix, gives a concrete implementation plan, or asks another agent to implement, patch, fix, or open a pull request.
87
+
88
+ Do not treat these as fix ownership:
89
+
90
+ - A bot summarizes evidence.
91
+ - A tool looks up logs or tickets.
92
+ - Someone asks a bot to diagnose, explain, inspect, or reproduce.
93
+ - A bot posts a cause hypothesis without agreeing to implement it.
94
+
95
+ Judge the requested action, not the presence of a bot.
96
+
97
+ ### A fix artifact already exists
98
+
99
+ If an open pull request or merged commit plausibly fixes this report, switch to `references/verify-existing-fix.md`.
100
+
101
+ An artifact may come from the thread, tracker issue, repository history, or pull request search. A claim without a commit or pull request is not a fix artifact.
102
+
103
+ If a person owns the work but has not produced an artifact, stop. Do not race them.
104
+
105
+ ## 4. Open an optional operations thread
106
+
107
+ If `slack.operations_channel_id` is configured, the coordinator may create one root status message there. This is the only allowed root post in the repro workflow.
108
+
109
+ Store its coordinates as `OPERATIONS_CHANNEL_ID` and `OPERATIONS_THREAD_TS`. Never confuse them with the source coordinates.
110
+
111
+ Use the configured plain Unicode status strings. Keep status text short:
112
+
113
+ - Reproducing
114
+ - Could not reproduce
115
+ - Blocked
116
+ - Reproduced
117
+ - Verifying existing fix
118
+ - Attempting bounded fix
119
+ - Draft pull request opened
120
+ - Fix did not land
121
+
122
+ Prefer the Slack actions your harness provides (a Slack MCP server or channel). Use `BENNY_SLACK_BOT_TOKEN` only when the user configured it for a narrow missing capability such as editing this one status message. Never expose the token to a worker.
123
+
124
+ If no operations channel is configured, keep detailed status in the automation run output. Do not substitute a source-channel root message.
125
+
126
+ ## 5. Load and check the control adapter
127
+
128
+ Read `references/control-adapter.md` and the completed map at `control.feature_map_path`, then invoke the skill named by `control.skill_name`.
129
+
130
+ For web, IDE, or Electron, load bundled pstack `control-ui` via the `pstack-harness` **skill** row before driving the app. The configured adapter supplies the app-specific launch steps, selectors, and feature map; it is not a separate generic `control` skill. Other UI platforms need their configured platform driver.
131
+
132
+ Find the feature-map section that matches the reported user path. Read it before driving the app. If no section covers the feature, mark the run blocked instead of inventing a path or selector.
133
+
134
+ Require all seven capabilities:
135
+
136
+ 1. Bring up the configured target app and test environment.
137
+ 2. Navigate the mapped feature and exercise its documented states.
138
+ 3. Drive the real UI with clicks, typing, keys, scrolling, drag, resize, or navigation.
139
+ 4. Inspect state without mutating it.
140
+ 5. Capture screenshots.
141
+ 6. Start and stop a screen recording.
142
+ 7. Clean up processes, sessions, profiles, and temporary data.
143
+
144
+ If the adapter is absent or any required capability is missing, mark the operations status as blocked and stop. Do not pretend a screenshot, unit test, state mutation, or source reading is a UI repro.
145
+
146
+ ## 6. Study the report
147
+
148
+ Read the full source thread and tracker issue when present.
149
+
150
+ Collect:
151
+
152
+ - Exact action path
153
+ - Expected behavior
154
+ - Observed behavior
155
+ - Discriminating state where they diverge
156
+ - Frequency
157
+ - Version, environment, and platform
158
+ - Attachments and error signatures
159
+ - Candidate code area
160
+
161
+ Inspect screenshots and video. Use read-only parallel workers for code history, test ideas, blast-radius mapping, and media review when useful. Each worker gets a narrow question and the Slack-write prohibition.
162
+
163
+ Use pstack's `how` skill to trace the action through the repository. Use `why` for regression history and defensive code. Form competing cause hypotheses and identify evidence that would separate them.
164
+
165
+ ## 7. Reproduce
166
+
167
+ Bring up the target app through the control adapter.
168
+
169
+ Confirm the correct app, workspace, account, data set, and feature state before acting. Use stable app markers. Do not rely on window order or a familiar title alone.
170
+
171
+ Drive the reported path through real UI actions.
172
+
173
+ Before calling it reproduced:
174
+
175
+ 1. Name the correct final state.
176
+ 2. Name the broken final state.
177
+ 3. Reach the point where they diverge.
178
+ 4. Observe the broken state.
179
+ 5. Reset enough state to make the second attempt independent.
180
+ 6. Repeat the same path and observe the same broken state again.
181
+ 7. Cross-check a real state value when possible.
182
+
183
+ An expected dialog, loading state, or setup step is not the bug. Capture the final state that distinguishes correct from broken behavior.
184
+
185
+ Use the configured repro budget. If the symptom does not reproduce within it, report a clean `Could not reproduce` outcome. If the environment cannot provide a required capability, report `Blocked` and state what was missing.
186
+
187
+ ## 8. Capture and review evidence
188
+
189
+ For a successful repro:
190
+
191
+ - Record the full path through the symptom.
192
+ - Capture a screenshot of the broken final state.
193
+ - Save a short note with the exact steps and observed state.
194
+ - Keep artifacts in the configured temporary artifact directory.
195
+
196
+ Have a read-only media reviewer answer one question: does the evidence visibly show the discriminating broken state?
197
+
198
+ If the answer is no or uncertain, the repro is not confirmed. Capture better evidence or use `Could not reproduce`.
199
+
200
+ Post detailed evidence only in the operations thread when configured. Keep the source update concise.
201
+
202
+ ## 9. Report the repro outcome
203
+
204
+ Update the operations status first.
205
+
206
+ For `Could not reproduce` or `Blocked`, post nothing in the source thread. The operations thread or run output carries the result.
207
+
208
+ For a confirmed repro, run the source preflight and post at most one unprompted source reply:
209
+
210
+ - Say the issue reproduced.
211
+ - Link the operations evidence thread when one exists.
212
+ - Include at most three short findings.
213
+ - Link the tracker issue when one exists.
214
+ - Do not ping an owner by default.
215
+
216
+ Attach evidence only when the configured Slack action keeps it inside the same source thread and the organization's retention policy allows it.
217
+
218
+ Wait for the configured rejection window. If a person shows that the setup or interpretation was wrong, correct the repro once. Do not start the fix phase until the window closes without a valid rejection.
219
+
220
+ ## 10. Verify an existing fix
221
+
222
+ When a fix artifact exists, follow `references/verify-existing-fix.md`.
223
+
224
+ Verification must show the symptom on the baseline and its absence on the patched build. Both paths use the real UI twice.
225
+
226
+ Do not edit the existing fix, add a competing patch, or open a replacement pull request.
227
+
228
+ ## 11. Qualify a bounded fix
229
+
230
+ Attempt a fix only when all of these hold:
231
+
232
+ - The outcome is a plain confirmed repro.
233
+ - Media review confirmed the broken final state.
234
+ - No existing fix artifact appeared.
235
+ - No person claimed the fix during the rejection window.
236
+ - Runtime evidence identifies the root cause.
237
+ - The likely change fits the configured fix budget and repository scope.
238
+ - The control adapter can run both baseline and patched builds.
239
+
240
+ If any condition fails, keep the repro report and stop without a pull request.
241
+
242
+ When the gate passes, update operations status to `Attempting bounded fix`.
243
+
244
+ ## 12. Root-cause and implement
245
+
246
+ The coordinator owns every Slack post, the final diff review, commits, and the pull request.
247
+
248
+ Read-only workers may:
249
+
250
+ - Trace code and history
251
+ - Propose tests
252
+ - Map blast radius
253
+ - Review a diff
254
+ - Review media
255
+
256
+ They do not edit, run external writes, post status, or own the fix.
257
+
258
+ A tightly scoped code edit may be delegated during this phase only when tool isolation removes Slack credentials and every Slack write action from that worker. Its prompt must still carry the explicit Slack-write ban. The coordinator reviews the edit and runs or verifies the required tests. If tool isolation is uncertain, keep the edit in the coordinator.
259
+
260
+ Confirm the mechanism with runtime evidence. Eliminate competing hypotheses before editing.
261
+
262
+ Fix the root cause with the smallest justified change.
263
+
264
+ - Invoke pstack's `tdd` skill when there is a cheap local test target, and write the failing test before the fix.
265
+ - State why TDD was skipped when the path is expensive, unclear, or integration-heavy.
266
+ - Keep unrelated cleanup out.
267
+ - Stop if the change grows beyond the configured effort or risk budget.
268
+
269
+ ## 13. Prove the fix
270
+
271
+ Keep the original baseline evidence.
272
+
273
+ On the patched build:
274
+
275
+ 1. Run the same real UI path.
276
+ 2. Repeat it twice.
277
+ 3. Show that the broken state is gone.
278
+ 4. Show the expected state in its place.
279
+ 5. Capture an after recording and screenshot.
280
+ 6. Cross-check the same real state value used for the baseline.
281
+
282
+ A compile, unit test, code review, or plausible diff is not after evidence.
283
+
284
+ Run focused tests, then smoke the blast radius around the changed behavior. Cover nearby states, inputs, permissions, platforms, and failure paths that the change could affect. Stop without a pull request if a regression remains.
285
+
286
+ ## 14. Open a draft pull request
287
+
288
+ Only after before-and-after proof:
289
+
290
+ - Review the final diff for unrelated changes and secrets.
291
+ - Load bundled pstack `deslop` per the `pstack-harness` **deslop** row before any fix commit. If cleanup changes code, rerun affected tests and the mapped UI repro through the configured adapter, with `control-ui` for web/IDE/Electron, before opening the pull request.
292
+ - Run the repository's required checks.
293
+ - Create small ordered commits when the repository workflow allows it.
294
+ - Open a draft pull request. Never merge or deploy from this workflow.
295
+ - Link the configured tracker issue using the tracker's supported pull request syntax.
296
+ - Use the configured public URL form, normally `https://github.com/{owner}/{repo}/pull/{number}`.
297
+ - Include the repro steps, root cause, test result, before and after evidence, and blast-radius checks.
298
+ - Run the pull request text and all Slack updates through pstack's `unslop` skill.
299
+
300
+ If pull request creation fails, do not claim success. Keep the commit or branch state in the run output and mark operations status `Fix did not land`.
301
+
302
+ On success, mark operations status `Draft pull request opened` and post one concise reply in the operations thread with the linked pull request. Do not create a second source-channel root or unprompted source reply.
303
+
304
+ ## 15. Follow-ups and cleanup
305
+
306
+ Watch the configured operations thread for one follow-up window.
307
+
308
+ - Answer a direct question from evidence already gathered.
309
+ - Apply one concrete correction and rerun the repro once when it invalidates the setup.
310
+ - Stay out of human coordination and side chatter.
311
+ - Stop when asked.
312
+
313
+ Always call the control adapter's cleanup capability. Keep artifacts only as long as the configured retention policy allows.
@@ -0,0 +1,169 @@
1
+ # Control-adapter contract
2
+
3
+ Benny does not know how to start or drive every app. Configure an app-specific verification skill or adapter that implements this contract. For web, IDE, or Electron driving, that skill loads bundled pstack `control-ui` and supplies the app-specific launch commands, selectors, and feature map. There is no separate generic `control` skill. The generic `control-ui` skill alone does not supply all of Benny's app-specific capabilities.
4
+
5
+ Set its skill name in `control.skill_name`.
6
+
7
+ Set the completed user-facing feature map path in `control.feature_map_path`. Copy and fill [`feature-map.example.md`](./feature-map.example.md) outside `.pstack/automations/benny/` instead of editing the copied example.
8
+
9
+ If the skill, feature map, or a required capability is absent, ambiguous, or incomplete, repro and fix work must fail closed.
10
+
11
+ ## Required capabilities
12
+
13
+ ### Bring up
14
+
15
+ Start the requested app revision in the requested test environment.
16
+
17
+ Input:
18
+
19
+ - Repository and revision
20
+ - Build or start mode
21
+ - Workspace, account, fixture, and feature-state requirements
22
+ - Artifact directory
23
+ - Completed feature-map path
24
+
25
+ Return:
26
+
27
+ - Session identifier
28
+ - How the adapter confirmed the correct app and environment
29
+ - Stable app markers
30
+ - Running process or target details needed by later calls
31
+ - Any missing capability
32
+
33
+ The adapter must distinguish the target app from a similar window, shell, or production instance.
34
+
35
+ ### Drive UI
36
+
37
+ Perform real user actions:
38
+
39
+ - Click
40
+ - Type
41
+ - Press keys
42
+ - Scroll
43
+ - Drag
44
+ - Resize
45
+ - Navigate through app controls
46
+
47
+ Prefer roles, labels, and stable selectors. Use coordinates only after a fresh screenshot.
48
+
49
+ Return each action and the observed state change.
50
+
51
+ Do not set internal state, call hidden app methods, write directly to storage, or inject DOM changes to create the symptom.
52
+
53
+ ### Drive mapped features and states
54
+
55
+ Read the relevant feature-map section before driving the app.
56
+
57
+ The adapter must expose ways to:
58
+
59
+ - Navigate every mapped feature through the user-visible path.
60
+ - Invoke the adapter action names listed for that feature.
61
+ - Interact with default, hover, focus-visible, active, disabled, loading, empty, error, selected, open, expanded, and feature-specific states when they apply.
62
+ - Arrange a state through safe fixture data, permissions, flags, service responses, or supported test controls.
63
+ - Reset the feature for a second independent repro attempt.
64
+ - Capture the screenshot, video, and read-only cross-check named by the feature map.
65
+
66
+ Use roles, accessible names, ARIA relationships, stable component markers, and purpose-named data attributes. Never use generated CSS or StyleX classes, dynamic hashes, child indexes, or brittle DOM position.
67
+
68
+ Arranging a precondition is not permission to inject the reported symptom. The repro itself must still come from real user interaction.
69
+
70
+ ### Inspect state
71
+
72
+ Read state to confirm what the UI shows.
73
+
74
+ Examples:
75
+
76
+ - Accessibility tree
77
+ - DOM or view hierarchy
78
+ - Process state
79
+ - Local logs
80
+ - Network request status
81
+ - App-exposed debug state
82
+
83
+ Inspection is read-only. If a query changes state, it belongs in `drive UI` and must represent a real user action.
84
+
85
+ ### Screenshot
86
+
87
+ Capture the current app state to a requested path.
88
+
89
+ Return:
90
+
91
+ - File path
92
+ - Capture time
93
+ - App marker or window title
94
+ - Short description of what should be visible
95
+
96
+ The screenshot must show enough app chrome to prove that the correct app is under test.
97
+
98
+ ### Recording
99
+
100
+ Start and stop a screen recording around the full repro path.
101
+
102
+ Return:
103
+
104
+ - File path
105
+ - Start and stop times
106
+ - Captured window or region
107
+ - Whether audio or sensitive overlays were omitted
108
+
109
+ The recording must show the discriminating final state, not only setup or a loading screen.
110
+
111
+ ### Cleanup
112
+
113
+ Stop processes and sessions created by the adapter.
114
+
115
+ Remove disposable:
116
+
117
+ - Browser or app profiles
118
+ - Temporary workspaces
119
+ - Test accounts or fixtures when the adapter created them
120
+ - Debug ports and tunnels
121
+ - Captures past their retention window
122
+
123
+ Return what was stopped, removed, retained, or left for a person.
124
+
125
+ Cleanup must not delete user work.
126
+
127
+ ## Adapter behavior
128
+
129
+ The adapter must:
130
+
131
+ - Report capabilities before the repro starts.
132
+ - Report which feature-map sections it can drive and which are blocked.
133
+ - Use the same environment inputs for baseline and patched builds.
134
+ - Surface startup failures as failures.
135
+ - Bound retries.
136
+ - Keep secrets out of logs and artifacts.
137
+ - Keep captures outside the repository.
138
+ - Support a fresh or reset state between the two repro attempts.
139
+ - Avoid production changes unless the user explicitly configured a safe test action.
140
+
141
+ ## Environment translation
142
+
143
+ Before declaring an environment block, restate the defect without platform-specific nouns and ask whether the same behavior can be tested safely in the available environment.
144
+
145
+ Examples:
146
+
147
+ - A named browser may mean any external browser.
148
+ - A named key may mean the configured shortcut.
149
+ - A named remote host may mean a delayed or disconnected remote target.
150
+
151
+ Use a translated attempt only when it tests the same underlying behavior. Label it as translated evidence. Do not call it an exact repro when the missing environment is part of the defect.
152
+
153
+ Hardware prompts, operating-system permission dialogs, device-only APIs, and unavailable account states may be real blocks.
154
+
155
+ ## Setup check
156
+
157
+ Before enabling the repro automation, run one harmless adapter check:
158
+
159
+ 1. Bring up the app.
160
+ 2. Confirm the stable app marker.
161
+ 3. Load one completed feature-map section.
162
+ 4. Navigate to that feature through its user path.
163
+ 5. Exercise one disposable state through mapped adapter actions.
164
+ 6. Inspect the resulting state.
165
+ 7. Capture a screenshot.
166
+ 8. Record a short clip.
167
+ 9. Clean up.
168
+
169
+ Enable repro work only when all nine steps succeed and no source-channel Slack post is involved.
@@ -0,0 +1,205 @@
1
+ # Feature-map example
2
+
3
+ Map every user-facing feature Benny may reproduce. Read the relevant section before driving the app. Keep this map at the user point of view. Discover internals and current code paths at runtime instead of freezing them here.
4
+
5
+ Copy this file outside `.pstack/automations/benny/`, for example to `.pstack/benny/feature-map.md`, and set `control.feature_map_path` to the copy. Pack refreshes must not overwrite it.
6
+
7
+ ## Per-feature template
8
+
9
+ ### `<feature name>`
10
+
11
+ `<one-line user-visible purpose>`
12
+
13
+ #### How a user gets there
14
+
15
+ - Click path: `<screen> -> <menu, tab, or panel> -> <control>`
16
+ - Keyboard shortcut: `<shortcut or none>`
17
+
18
+ #### How the control adapter drives it
19
+
20
+ - `<adapter action>` with `<inputs>` should `<visible result>`.
21
+ - Reset: `<how the adapter returns to a fresh state>`.
22
+
23
+ #### Stable selectors
24
+
25
+ - `<role and accessible name>`
26
+ - `<ARIA relationship>`
27
+ - `<data-component or purpose-named data attribute>`
28
+
29
+ Never use generated CSS or StyleX classes, dynamic hashes, child indexes, or brittle DOM position.
30
+
31
+ #### States to exercise
32
+
33
+ - Default, hover, focus-visible, active, disabled
34
+ - Loading, empty, error
35
+ - Selected, open, expanded
36
+ - `<relevant feature-specific variants>`
37
+
38
+ Mark states that do not apply.
39
+
40
+ #### Preconditions and setup
41
+
42
+ - Auth: `<account state>`
43
+ - Data: `<fixture>`
44
+ - Permissions: `<role>`
45
+ - Flags: `<flag or none>`
46
+ - Services: `<required availability>`
47
+
48
+ #### Evidence and cross-check
49
+
50
+ - Screenshot: `<app identity, feature, and discriminating state>`
51
+ - Video: `<entry path, interaction, and final state>`
52
+ - Cross-check: `<read-only state or value that confirms the UI>`
53
+
54
+ #### Gotchas
55
+
56
+ - `<known dead end or wrong surface>`
57
+ - `<safe environment translation>`
58
+
59
+ ## Fictional example
60
+
61
+ These features belong to a fictional task app. They are examples, not required Benny features.
62
+
63
+ ### Sign in
64
+
65
+ Lets a user enter the task app.
66
+
67
+ #### How a user gets there
68
+
69
+ - Open the app and choose `Sign in`. No shortcut.
70
+
71
+ #### How the control adapter drives it
72
+
73
+ - `open_app`, `click Sign in`, `fill credentials`, and `click Continue` should open the item list.
74
+ - Reset by signing out and clearing the disposable session.
75
+
76
+ #### Stable selectors
77
+
78
+ - Button `Sign in`, textboxes `Email` and `Password`, `data-component="sign-in-form"`
79
+
80
+ #### States to exercise
81
+
82
+ - Default, focus-visible, submitting, disabled, loading, error
83
+
84
+ #### Preconditions and setup
85
+
86
+ - Disposable account and available authentication service
87
+
88
+ #### Evidence and cross-check
89
+
90
+ - Record landing page through item list. Check read-only session state.
91
+
92
+ #### Gotchas
93
+
94
+ - A marketing page is the wrong surface. A missing auth service is a block.
95
+
96
+ ### Item list and detail
97
+
98
+ Lets a user browse items and open one.
99
+
100
+ #### How a user gets there
101
+
102
+ - Open the `Items` tab, then choose a row.
103
+
104
+ #### How the control adapter drives it
105
+
106
+ - `select_tab Items` and `click <fixture item>` should open its detail.
107
+ - Reset by closing the detail and clearing selection.
108
+
109
+ #### Stable selectors
110
+
111
+ - Tab and list named `Items`, fixture-named row, `data-component="item-detail"`
112
+
113
+ #### States to exercise
114
+
115
+ - Loading, empty, error, selected, open, expanded
116
+
117
+ #### Preconditions and setup
118
+
119
+ - Named fixture items, read permission, available item service
120
+
121
+ #### Evidence and cross-check
122
+
123
+ - Show selection and matching detail title. Check selected-item ID.
124
+
125
+ #### Gotchas
126
+
127
+ - Search results may look similar but use a different path.
128
+
129
+ ### Item editor
130
+
131
+ Lets a user create or edit an item.
132
+
133
+ #### How a user gets there
134
+
135
+ - Choose `Edit` from detail or `New item` from the list.
136
+
137
+ #### How the control adapter drives it
138
+
139
+ - `click Edit`, `fill <field>`, and `click Save` should update detail.
140
+ - Reset by restoring the fixture.
141
+
142
+ #### Stable selectors
143
+
144
+ - Buttons `Edit`, `New item`, `Save`, form `Item editor`, label-linked fields
145
+
146
+ #### States to exercise
147
+
148
+ - Default, focus-visible, dirty, validating, disabled, saving, error, success
149
+
150
+ #### Preconditions and setup
151
+
152
+ - Editable fixture, write permission, available save service
153
+
154
+ #### Evidence and cross-check
155
+
156
+ - Show field change through updated detail. Check the stored item value read-only.
157
+
158
+ #### Gotchas
159
+
160
+ - Do not inject form state. A read-only detail field is not the editor.
161
+
162
+ ### Settings
163
+
164
+ Lets a user change personal preferences.
165
+
166
+ #### How a user gets there
167
+
168
+ - Open the profile menu, then choose `Settings`.
169
+
170
+ #### How the control adapter drives it
171
+
172
+ - `open_menu Profile`, `click Settings`, and `toggle <preference>` should update the control.
173
+ - Reset by restoring the starting preference.
174
+
175
+ #### Stable selectors
176
+
177
+ - Button `Profile`, menu item `Settings`, region `Settings`, purpose-named preference attribute
178
+
179
+ #### States to exercise
180
+
181
+ - Closed, open, selected, focus-visible, disabled, loading, error
182
+
183
+ #### Preconditions and setup
184
+
185
+ - Signed-in test account, known preferences, available preference service
186
+
187
+ #### Evidence and cross-check
188
+
189
+ - Show the menu path and final control state. Check the preference value read-only.
190
+
191
+ #### Gotchas
192
+
193
+ - Operating-system settings are a different surface.
194
+
195
+ ## Completeness checklist
196
+
197
+ - Every reproducible user-facing feature has a section.
198
+ - Every section names a user path, adapter actions, and reset.
199
+ - Selectors use roles, names, ARIA, stable component markers, or purpose-named attributes.
200
+ - No selector uses generated classes or DOM position.
201
+ - Relevant interaction, loading, empty, error, selected, and expanded states are covered.
202
+ - Auth, fixtures, permissions, flags, and services are explicit.
203
+ - Screenshot, video, and underlying cross-check requirements are explicit.
204
+ - Wrong surfaces, dead ends, and safe environment translations are listed.
205
+ - Implementation details remain runtime discoveries.