@jiroamato/pstack 0.0.0-stage → 0.15.15
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/LICENSE +21 -0
- package/README.md +78 -2
- package/bin/pstack.js +95 -0
- package/lib/install.js +103 -0
- package/lib/prompt.js +77 -0
- package/lib/targets.js +43 -0
- package/package.json +38 -5
- package/pstack/.claude-plugin/plugin.json +26 -0
- package/pstack/.codex-plugin/plugin.json +36 -0
- package/pstack/LICENSE +21 -0
- package/pstack/LICENSE-cursor-team-kit +21 -0
- package/pstack/NOTICE +8 -0
- package/pstack/README.md +300 -0
- package/pstack/agents/comment-sicko.md +34 -0
- package/pstack/agents/poteto-agent.md +10 -0
- package/pstack/automations/benny/FOR_AGENTS.md +92 -0
- package/pstack/automations/benny/README.md +28 -0
- package/pstack/automations/benny/skills/reproduce-and-fix-issues/SKILL.md +313 -0
- package/pstack/automations/benny/skills/reproduce-and-fix-issues/references/control-adapter.md +169 -0
- package/pstack/automations/benny/skills/reproduce-and-fix-issues/references/feature-map.example.md +205 -0
- package/pstack/automations/benny/skills/reproduce-and-fix-issues/references/verify-existing-fix.md +93 -0
- package/pstack/automations/benny/skills/setup-benny/SKILL.md +271 -0
- package/pstack/automations/benny/skills/triage-issue-reports/SKILL.md +240 -0
- package/pstack/automations/benny/skills/triage-issue-reports/references/routing.example.md +61 -0
- package/pstack/automations/benny/templates/configuration.example.yaml +84 -0
- package/pstack/automations/benny/templates/reproduce-automation-prompt.md +33 -0
- package/pstack/automations/benny/templates/triage-automation-prompt.md +39 -0
- package/pstack/codex/agents/comment-sicko.toml +36 -0
- package/pstack/codex/agents/poteto-agent.toml +11 -0
- package/pstack/docs/guide/01-setup.md +80 -0
- package/pstack/docs/guide/02-poteto-mode.md +131 -0
- package/pstack/docs/guide/03-understand.md +79 -0
- package/pstack/docs/guide/04-design.md +133 -0
- package/pstack/docs/guide/05-build-and-clean.md +83 -0
- package/pstack/docs/guide/06-verify-and-ship.md +130 -0
- package/pstack/docs/guide/07-overnight.md +120 -0
- package/pstack/docs/guide/08-principles.md +72 -0
- package/pstack/docs/guide/09-make-it-yours.md +100 -0
- package/pstack/docs/guide/10-recipes-and-pitfalls.md +156 -0
- package/pstack/docs/guide/README.md +38 -0
- package/pstack/skills/architect/SKILL.md +85 -0
- package/pstack/skills/architect/agents/openai.yaml +2 -0
- package/pstack/skills/architect/references/design-red-flags.md +57 -0
- package/pstack/skills/architect/references/rationale-template.md +35 -0
- package/pstack/skills/architect/references/runner-prompt.md +20 -0
- package/pstack/skills/arena/SKILL.md +75 -0
- package/pstack/skills/arena/agents/openai.yaml +2 -0
- package/pstack/skills/automate-me/SKILL.md +104 -0
- package/pstack/skills/automate-me/agents/openai.yaml +2 -0
- package/pstack/skills/benchmark-checklist/SKILL.md +39 -0
- package/pstack/skills/benchmark-checklist/agents/openai.yaml +2 -0
- package/pstack/skills/blast-radius/SKILL.md +52 -0
- package/pstack/skills/blast-radius/agents/openai.yaml +2 -0
- package/pstack/skills/bro/SKILL.md +7 -0
- package/pstack/skills/bro/agents/openai.yaml +2 -0
- package/pstack/skills/control-cli/SKILL.md +55 -0
- package/pstack/skills/control-cli/agents/openai.yaml +2 -0
- package/pstack/skills/control-ui/SKILL.md +72 -0
- package/pstack/skills/control-ui/agents/openai.yaml +2 -0
- package/pstack/skills/correct/SKILL.md +34 -0
- package/pstack/skills/correct/agents/openai.yaml +2 -0
- package/pstack/skills/create-verification-skill/SKILL.md +47 -0
- package/pstack/skills/create-verification-skill/agents/openai.yaml +2 -0
- package/pstack/skills/create-verification-skill/references/feature-map-example/README.md +47 -0
- package/pstack/skills/create-verification-skill/references/feature-map-example/create-note.md +39 -0
- package/pstack/skills/create-verification-skill/references/feature-map-example/search.md +45 -0
- package/pstack/skills/deslop/SKILL.md +30 -0
- package/pstack/skills/deslop/agents/openai.yaml +2 -0
- package/pstack/skills/figure-it-out/SKILL.md +55 -0
- package/pstack/skills/figure-it-out/agents/openai.yaml +2 -0
- package/pstack/skills/how/SKILL.md +58 -0
- package/pstack/skills/how/agents/openai.yaml +2 -0
- package/pstack/skills/how/references/explainer-prompt.md +55 -0
- package/pstack/skills/how/references/explorer-prompt.md +52 -0
- package/pstack/skills/interrogate/SKILL.md +111 -0
- package/pstack/skills/interrogate/agents/openai.yaml +2 -0
- package/pstack/skills/interrogate/references/code-quality-review.md +47 -0
- package/pstack/skills/interrogate/references/lead-judgment.md +58 -0
- package/pstack/skills/interrogate/references/reviewer-prompt.md +70 -0
- package/pstack/skills/interrogate/references/rubric.md +77 -0
- package/pstack/skills/maintain-verification-skill/SKILL.md +41 -0
- package/pstack/skills/maintain-verification-skill/agents/openai.yaml +2 -0
- package/pstack/skills/make-bot-ui/SKILL.md +289 -0
- package/pstack/skills/make-bot-ui/agents/openai.yaml +2 -0
- package/pstack/skills/no-comments/SKILL.md +24 -0
- package/pstack/skills/no-comments/agents/openai.yaml +2 -0
- package/pstack/skills/poteto-help/SKILL.md +156 -0
- package/pstack/skills/poteto-help/agents/openai.yaml +2 -0
- package/pstack/skills/poteto-help/references/prompting.md +51 -0
- package/pstack/skills/poteto-help/references/recipes.md +47 -0
- package/pstack/skills/poteto-mode/SKILL.md +143 -0
- package/pstack/skills/poteto-mode/agents/openai.yaml +2 -0
- package/pstack/skills/poteto-mode/playbooks/authoring-a-skill.md +12 -0
- package/pstack/skills/poteto-mode/playbooks/autonomous-run.md +13 -0
- package/pstack/skills/poteto-mode/playbooks/autopilot-full.md +13 -0
- package/pstack/skills/poteto-mode/playbooks/autopilot-stack.md +16 -0
- package/pstack/skills/poteto-mode/playbooks/babysit.md +29 -0
- package/pstack/skills/poteto-mode/playbooks/bug-fix.md +15 -0
- package/pstack/skills/poteto-mode/playbooks/eval.md +25 -0
- package/pstack/skills/poteto-mode/playbooks/feature.md +21 -0
- package/pstack/skills/poteto-mode/playbooks/hillclimb.md +21 -0
- package/pstack/skills/poteto-mode/playbooks/investigation.md +14 -0
- package/pstack/skills/poteto-mode/playbooks/multi-phase-plan.md +155 -0
- package/pstack/skills/poteto-mode/playbooks/opening-a-pr.md +38 -0
- package/pstack/skills/poteto-mode/playbooks/orchestrate.md +114 -0
- package/pstack/skills/poteto-mode/playbooks/pause-safely.md +10 -0
- package/pstack/skills/poteto-mode/playbooks/perf-issue.md +25 -0
- package/pstack/skills/poteto-mode/playbooks/prototype.md +14 -0
- package/pstack/skills/poteto-mode/playbooks/refactoring.md +16 -0
- package/pstack/skills/poteto-mode/playbooks/runtime-forensics.md +11 -0
- package/pstack/skills/poteto-mode/playbooks/session-pickup.md +11 -0
- package/pstack/skills/poteto-mode/playbooks/shipping.md +17 -0
- package/pstack/skills/poteto-mode/playbooks/trace-forensics.md +14 -0
- package/pstack/skills/poteto-mode/playbooks/visual-parity.md +11 -0
- package/pstack/skills/poteto-mode/playbooks/worktree-cleanup.md +14 -0
- package/pstack/skills/poteto-mode/references/bugbot-triage.md +142 -0
- package/pstack/skills/poteto-mode/scripts/bootstrap.ts +62 -0
- package/pstack/skills/poteto-mode/scripts/bun.lock +67 -0
- package/pstack/skills/poteto-mode/scripts/check-plan.mjs +185 -0
- package/pstack/skills/poteto-mode/scripts/orch/orch.test.ts +634 -0
- package/pstack/skills/poteto-mode/scripts/orch/orch.ts +578 -0
- package/pstack/skills/poteto-mode/scripts/orch/store.ts +1607 -0
- package/pstack/skills/poteto-mode/scripts/package.json +16 -0
- package/pstack/skills/poteto-mode/scripts/watch-pr/cli.test.ts +224 -0
- package/pstack/skills/poteto-mode/scripts/watch-pr/cli.ts +223 -0
- package/pstack/skills/poteto-mode/scripts/watch-pr/fakes.test-helper.ts +118 -0
- package/pstack/skills/poteto-mode/scripts/watch-pr/github.test.ts +306 -0
- package/pstack/skills/poteto-mode/scripts/watch-pr/github.ts +699 -0
- package/pstack/skills/poteto-mode/scripts/watch-pr/policy.test.ts +420 -0
- package/pstack/skills/poteto-mode/scripts/watch-pr/policy.ts +832 -0
- package/pstack/skills/poteto-mode/scripts/watch-pr/render.ts +169 -0
- package/pstack/skills/poteto-mode/scripts/watch-pr/tsconfig.json +13 -0
- package/pstack/skills/poteto-mode/scripts/watch-pr/types.compile.ts +93 -0
- package/pstack/skills/poteto-mode/scripts/watch-pr/types.ts +401 -0
- package/pstack/skills/poteto-mode/scripts/watch-pr/watch-pr +6 -0
- package/pstack/skills/poteto-mode/scripts/worktree-audit.sh +92 -0
- package/pstack/skills/principle-attack-the-premise/SKILL.md +23 -0
- package/pstack/skills/principle-attack-the-premise/agents/openai.yaml +2 -0
- package/pstack/skills/principle-boundary-discipline/SKILL.md +34 -0
- package/pstack/skills/principle-boundary-discipline/agents/openai.yaml +2 -0
- package/pstack/skills/principle-build-the-lever/SKILL.md +23 -0
- package/pstack/skills/principle-build-the-lever/agents/openai.yaml +2 -0
- package/pstack/skills/principle-encode-lessons-in-structure/SKILL.md +31 -0
- package/pstack/skills/principle-encode-lessons-in-structure/agents/openai.yaml +2 -0
- package/pstack/skills/principle-exhaust-the-design-space/SKILL.md +21 -0
- package/pstack/skills/principle-exhaust-the-design-space/agents/openai.yaml +2 -0
- package/pstack/skills/principle-experience-first/SKILL.md +19 -0
- package/pstack/skills/principle-experience-first/agents/openai.yaml +2 -0
- package/pstack/skills/principle-explain-the-number/SKILL.md +23 -0
- package/pstack/skills/principle-explain-the-number/agents/openai.yaml +2 -0
- package/pstack/skills/principle-fix-root-causes/SKILL.md +23 -0
- package/pstack/skills/principle-fix-root-causes/agents/openai.yaml +2 -0
- package/pstack/skills/principle-foundational-thinking/SKILL.md +21 -0
- package/pstack/skills/principle-foundational-thinking/agents/openai.yaml +2 -0
- package/pstack/skills/principle-guard-the-context-window/SKILL.md +16 -0
- package/pstack/skills/principle-guard-the-context-window/agents/openai.yaml +2 -0
- package/pstack/skills/principle-laziness-protocol/SKILL.md +18 -0
- package/pstack/skills/principle-laziness-protocol/agents/openai.yaml +2 -0
- package/pstack/skills/principle-make-operations-idempotent/SKILL.md +24 -0
- package/pstack/skills/principle-make-operations-idempotent/agents/openai.yaml +2 -0
- package/pstack/skills/principle-migrate-callers-then-delete-legacy-apis/SKILL.md +22 -0
- package/pstack/skills/principle-migrate-callers-then-delete-legacy-apis/agents/openai.yaml +2 -0
- package/pstack/skills/principle-minimize-reader-load/SKILL.md +23 -0
- package/pstack/skills/principle-minimize-reader-load/agents/openai.yaml +2 -0
- package/pstack/skills/principle-model-the-domain/SKILL.md +26 -0
- package/pstack/skills/principle-model-the-domain/agents/openai.yaml +2 -0
- package/pstack/skills/principle-never-block-on-the-human/SKILL.md +20 -0
- package/pstack/skills/principle-never-block-on-the-human/agents/openai.yaml +2 -0
- package/pstack/skills/principle-outcome-oriented-execution/SKILL.md +21 -0
- package/pstack/skills/principle-outcome-oriented-execution/agents/openai.yaml +2 -0
- package/pstack/skills/principle-prove-it-works/SKILL.md +22 -0
- package/pstack/skills/principle-prove-it-works/agents/openai.yaml +2 -0
- package/pstack/skills/principle-redesign-from-first-principles/SKILL.md +16 -0
- package/pstack/skills/principle-redesign-from-first-principles/agents/openai.yaml +2 -0
- package/pstack/skills/principle-separate-before-serializing-shared-state/SKILL.md +16 -0
- package/pstack/skills/principle-separate-before-serializing-shared-state/agents/openai.yaml +2 -0
- package/pstack/skills/principle-sequence-verifiable-units/SKILL.md +17 -0
- package/pstack/skills/principle-sequence-verifiable-units/agents/openai.yaml +2 -0
- package/pstack/skills/principle-subtract-before-you-add/SKILL.md +21 -0
- package/pstack/skills/principle-subtract-before-you-add/agents/openai.yaml +2 -0
- package/pstack/skills/principle-test-behavior-not-implementation/SKILL.md +25 -0
- package/pstack/skills/principle-test-behavior-not-implementation/agents/openai.yaml +2 -0
- package/pstack/skills/principle-type-system-discipline/SKILL.md +31 -0
- package/pstack/skills/principle-type-system-discipline/agents/openai.yaml +2 -0
- package/pstack/skills/pstack-harness/SKILL.md +67 -0
- package/pstack/skills/recall/SKILL.md +35 -0
- package/pstack/skills/recall/agents/openai.yaml +2 -0
- package/pstack/skills/reflect/SKILL.md +76 -0
- package/pstack/skills/reflect/agents/openai.yaml +2 -0
- package/pstack/skills/reflect/references/divergent-reviewer.md +43 -0
- package/pstack/skills/reflect/references/judgment-reviewer.md +42 -0
- package/pstack/skills/reflect/references/synthesizer.md +56 -0
- package/pstack/skills/reflect/references/tooling-reviewer.md +55 -0
- package/pstack/skills/setup-pstack/SKILL.md +110 -0
- package/pstack/skills/show-me-your-work/SKILL.md +82 -0
- package/pstack/skills/show-me-your-work/agents/openai.yaml +2 -0
- package/pstack/skills/show-me-your-work/references/decision-log-template.tsv +1 -0
- package/pstack/skills/show-me-your-work/scripts/log.sh +42 -0
- package/pstack/skills/swarm/SKILL.md +48 -0
- package/pstack/skills/swarm/agents/openai.yaml +2 -0
- package/pstack/skills/tdd/SKILL.md +44 -0
- package/pstack/skills/tdd/agents/openai.yaml +2 -0
- package/pstack/skills/teach/SKILL.md +21 -0
- package/pstack/skills/teach/agents/openai.yaml +2 -0
- package/pstack/skills/technical-writing/SKILL.md +106 -0
- package/pstack/skills/technical-writing/agents/openai.yaml +2 -0
- package/pstack/skills/typescript-best-practices/SKILL.md +31 -0
- package/pstack/skills/typescript-best-practices/agents/openai.yaml +2 -0
- package/pstack/skills/typescript-best-practices/references/patterns.md +324 -0
- package/pstack/skills/unslop/SKILL.md +67 -0
- package/pstack/skills/unslop/agents/openai.yaml +2 -0
- package/pstack/skills/why/SKILL.md +158 -0
- package/pstack/skills/why/agents/openai.yaml +2 -0
- package/pstack/skills/why/references/epistemics.md +144 -0
- package/pstack/skills/why/references/investigator-prompt.md +103 -0
- package/pstack/skills/why/references/source-playbook.md +17 -0
- package/pstack/skills/why/references/sources/code-archaeology.md +88 -0
- package/pstack/skills/why/references/sources/databricks.md +70 -0
- package/pstack/skills/why/references/sources/datadog.md +99 -0
- package/pstack/skills/why/references/sources/incident-postmortem.md +15 -0
- package/pstack/skills/why/references/sources/linear.md +48 -0
- package/pstack/skills/why/references/sources/notion.md +55 -0
- package/pstack/skills/why/references/sources/sentry.md +100 -0
- package/pstack/skills/why/references/sources/slack.md +54 -0
- package/pstack/skills/why/references/synthesizer-prompt.md +135 -0
|
@@ -0,0 +1,52 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: blast-radius
|
|
3
|
+
description: "Find what a change could break somewhere else before it ships, beyond the diff, and prove the one fact it's safe because of by running real code instead of writing it up. Use for 'blast radius of X', 'what could this break', or reviewing a small diff you don't trust."
|
|
4
|
+
disable-model-invocation: true
|
|
5
|
+
---
|
|
6
|
+
|
|
7
|
+
# Blast radius
|
|
8
|
+
|
|
9
|
+
Find what a change breaks somewhere else, before it ships. Use for "blast radius of X", "what could this break", or reviewing a small diff you don't trust yet.
|
|
10
|
+
|
|
11
|
+
Companion to `how` and `why`. `how` tells you what the code does. `why` tells you why it's shaped that way. Blast radius tells you what it breaks somewhere else.
|
|
12
|
+
|
|
13
|
+
Listing the callers is not the job. The agent can grep those in a second. The job is the breakage grep won't show you.
|
|
14
|
+
|
|
15
|
+
## Don't trust your own writeup
|
|
16
|
+
|
|
17
|
+
A blast-radius writeup that sounds right is worthless. It reads as convincing whether or not it's true. So don't hand back the writeup. Find the one or two facts the whole thing depends on and prove them by running code.
|
|
18
|
+
|
|
19
|
+
### How sure are you
|
|
20
|
+
|
|
21
|
+
For each fact the change's safety depends on, get it as far down this list as is cheap, and say where it stopped.
|
|
22
|
+
|
|
23
|
+
1. You said so. Worthless on its own.
|
|
24
|
+
2. You pointed at the line. A real `file:line`, or the library's own source.
|
|
25
|
+
3. You showed the bad case can't happen. You walked the failure step by step and it doesn't reach.
|
|
26
|
+
4. You ran it. A script or test that calls the real code and fails loud if you're wrong.
|
|
27
|
+
5. You reproduced it in the running app.
|
|
28
|
+
|
|
29
|
+
Step 4 is usually one small script that imports the same library the app ships and calls the exact function you're worried about.
|
|
30
|
+
|
|
31
|
+
## Steps
|
|
32
|
+
|
|
33
|
+
If the proof needs a running web, IDE, or Electron user flow, load [`control-ui`](../control-ui/SKILL.md); for CLI/TUI, load [`control-cli`](../control-cli/SKILL.md). Read [`pstack-harness`](../pstack-harness/SKILL.md) for invocation and include any project verification skill. Keep this risk investigation read-only with respect to the reviewed diff.
|
|
34
|
+
|
|
35
|
+
1. Read the change. The diff, the symbols it adds, changes, and deletes, and what it now does differently, including the part the diff doesn't spell out. Use `why` step 2 to pull the PR and commits.
|
|
36
|
+
2. Find the one fact it's safe because of. Most changes that look risky are safe because of a single fact, like "this call only drops already-dead cache entries and does nothing else". Find that fact. If it holds, most risky cases are cleared at once. Spend your time here, not on a long list of maybes.
|
|
37
|
+
3. Look where grep stops. Read the source of the library you call, and check its pinned version and any local patch. Work out when things run: microtasks, unmount and teardown, Solid versus React. Follow what a symbol search misses: the JSON an API returns, a DB column, a wire format, another language reading the same bytes, a feature flag, code three hops downstream.
|
|
38
|
+
4. Be honest about each risk. Give it a real chance of happening and a real cost if it does. Keep the risks you confirmed. List the ones you checked and cleared separately. Same rules as `why`. Cite a real `file:line`, a search that finds nothing is still an answer, and never make up a caller or an API.
|
|
39
|
+
5. Prove the one fact. Write a script or test that runs the real code, run it, and paste what happened.
|
|
40
|
+
6. For a big or wide change, run it as an `arena`. Ask more than one model the same question and merge the answers. Different models catch different real bugs.
|
|
41
|
+
|
|
42
|
+
## What to hand back
|
|
43
|
+
|
|
44
|
+
- **What it does.** What changed, including the part that isn't obvious.
|
|
45
|
+
- **The one fact it's safe because of.** State it, say which step you got it to, and show the proof. If you couldn't prove it, write unproven.
|
|
46
|
+
- **Risks.** Each names how it breaks, the `file:line`, how likely and how bad, and how to check. Paste the proof for the ones that matter.
|
|
47
|
+
- **Cleared.** What you checked and why it's fine.
|
|
48
|
+
- **Before you merge.** The cheapest test or repro that catches the real bug, including the script you wrote.
|
|
49
|
+
|
|
50
|
+
Write it through `unslop`, cite real code, and strip anything private before it goes anywhere public.
|
|
51
|
+
|
|
52
|
+
**Reply:** the writeup above, with the one safety fact either proven or marked unproven.
|
|
@@ -0,0 +1,7 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: bro
|
|
3
|
+
description: Restate the last message in plain human language, with no jargon.
|
|
4
|
+
disable-model-invocation: true
|
|
5
|
+
---
|
|
6
|
+
|
|
7
|
+
Restate your last message. Stop using jargon and speak coherently. State it more simply and concisely, like one human talking to another.
|
|
@@ -0,0 +1,55 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: control-cli
|
|
3
|
+
description: Drive, inspect, and profile a local CLI or TUI with a repeatable terminal harness. Use for CLI UX checks, startup regressions, memory leaks, hangs, prompt flows, keyboard behavior, or terminal demos.
|
|
4
|
+
disable-model-invocation: true
|
|
5
|
+
---
|
|
6
|
+
|
|
7
|
+
# Control CLI
|
|
8
|
+
|
|
9
|
+
Exercise the real command with repeatable input and capture evidence. Reuse the project's `verify-<app>` skill or checked-in test/demo harness first. Otherwise use terminal tools already available in the session. This skill supplies the workflow; it does not install a terminal driver.
|
|
10
|
+
|
|
11
|
+
## Choose the driver
|
|
12
|
+
|
|
13
|
+
- For a non-interactive command, run it through Claude Code's `Bash` or Codex's shell tool and capture stdout, stderr, exit status, and resulting files.
|
|
14
|
+
- For prompts or a TUI, use a real PTY. On Codex, use `exec_command` with `tty: true` and `write_stdin` when those tools are available and support the host. Poll the returned session ID and send input to that same session.
|
|
15
|
+
- On Claude Code, prefer a repo-native PTY helper, Expect, or `tmux` through `Bash`. A background shell with piped stdin is not proof of terminal behavior.
|
|
16
|
+
- On Windows, use an available ConPTY driver such as a repo's `node-pty` or `pywinpty` harness, or the session's native PTY tool. Unix `pty`, Expect, and `tmux` recipes require a Unix host or WSL with the app running there too.
|
|
17
|
+
- If no suitable driver exists, identify the missing capability and report what remains unverified. Do not substitute piped input for a TTY-sensitive flow or silently add a project dependency.
|
|
18
|
+
|
|
19
|
+
## Harness loop
|
|
20
|
+
|
|
21
|
+
1. Identify the command, terminal size, and smallest reproducible workspace. Discover package scripts, e2e tests, expect scripts, PTY helpers, and demo recorders.
|
|
22
|
+
2. Launch an isolated session with the repo's required environment and disposable input/data. Record its session ID or PID, command, cwd, and terminal dimensions.
|
|
23
|
+
3. Capture the screen before interacting. Keep terminal escape sequences when layout matters; use the existing terminal renderer for a visual capture if available.
|
|
24
|
+
4. Send one action at a time: text, Enter, arrows, Escape, Ctrl-C, or resize. Use the driver's supported key encoding.
|
|
25
|
+
5. Wait for a concrete prompt or screen pattern with a deadline before the next action. A successful input write does not prove the command handled it.
|
|
26
|
+
6. Assert the user-visible result and side effects, including files, exit codes, or API responses. Save the transcript and any profile artifacts. Capture before/after runs for a bug fix.
|
|
27
|
+
7. Close sessions and processes this run created, including failed attempts. Preserve evidence outside disposable state and confirm it survives cleanup.
|
|
28
|
+
|
|
29
|
+
## tmux recipe
|
|
30
|
+
|
|
31
|
+
On a Unix host with `tmux`, adapt this to the current command and prompt:
|
|
32
|
+
|
|
33
|
+
```bash
|
|
34
|
+
SESSION="cli-harness-$(date +%s)-$$"
|
|
35
|
+
tmux new-session -d -s "$SESSION" -- <command-under-test>
|
|
36
|
+
trap 'tmux kill-session -t "$SESSION" 2>/dev/null || true' EXIT
|
|
37
|
+
tmux capture-pane -pt "$SESSION"
|
|
38
|
+
# Wait for the app's ready prompt with a bounded poll before sending input.
|
|
39
|
+
tmux send-keys -t "$SESSION" "help" Enter
|
|
40
|
+
# Wait for the expected result, then capture and save the transcript.
|
|
41
|
+
tmux capture-pane -pt "$SESSION"
|
|
42
|
+
```
|
|
43
|
+
|
|
44
|
+
Use a temporary directory from the host's temp API or environment, rather than assuming `/tmp`. Keep a one-off harness temporary unless the task calls for a reusable test. Do not hard-code paths from another repository.
|
|
45
|
+
|
|
46
|
+
## Profiling and demos
|
|
47
|
+
|
|
48
|
+
- Startup regression: compare baseline and treatment with the same machine, environment, command, and readiness condition.
|
|
49
|
+
- Slow operation: collect a CPU profile around the operation and compare top self-time functions.
|
|
50
|
+
- Memory leak: repeat the operation between heap snapshots; force GC if the runtime supports it and record whether you did.
|
|
51
|
+
- Hang: capture the screen, active handles/resources, and a stack or CPU sample before interrupting.
|
|
52
|
+
- Node or Bun inspector: use the runtime's supported inspector flag with a loopback address and an ephemeral port. Read the inspector URL from startup output and use available DevTools-compatible tooling.
|
|
53
|
+
- Terminal demo: use an existing recorder or an asciinema-compatible tool when the user asks. A recording supplements the assertions above.
|
|
54
|
+
|
|
55
|
+
Use deterministic waits rather than fixed sleeps. Keep credentials out of transcripts. Report the command, actions, observed result, evidence paths, and any unverified behavior.
|
|
@@ -0,0 +1,72 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: control-ui
|
|
3
|
+
description: Drive and inspect a local web, IDE, or Electron UI with browser or CDP tooling. Use for UI verification, screenshots, accessibility snapshots, performance profiles, visual diffs, or reproducing UI bugs.
|
|
4
|
+
disable-model-invocation: true
|
|
5
|
+
---
|
|
6
|
+
|
|
7
|
+
# Control UI
|
|
8
|
+
|
|
9
|
+
Verify the running UI with evidence. Reuse the project's `verify-<app>` skill or checked-in Playwright, Cypress, browser, or Electron harness first. Otherwise use browser tools already available in the session. This skill supplies the workflow; it does not install a browser driver.
|
|
10
|
+
|
|
11
|
+
## Setup
|
|
12
|
+
|
|
13
|
+
1. Discover the app's documented dev command, local URL, required environment, and existing tests. Start an isolated instance and wait for a concrete readiness condition.
|
|
14
|
+
2. On Claude Code, use available browser MCP tools or the repo's browser harness through `Bash`. On Codex, use the session's browser tools or an installed browser harness through the shell. Read the selected tool's instructions before driving it; availability differs by environment.
|
|
15
|
+
3. For a web app, connect to its local URL. For Electron/Chromium, use the existing Electron harness or a loopback remote debugging port when supported.
|
|
16
|
+
4. Select the correct page by URL and a stable app marker, not tab order alone. If none matches, list available page titles and URLs rather than guessing.
|
|
17
|
+
5. Prefer accessibility roles, labels, and stable `data-*` selectors over coordinates. If no suitable tool exists, report the missing capability and what remains unverified. Do not silently add a project dependency.
|
|
18
|
+
|
|
19
|
+
## Interaction loop
|
|
20
|
+
|
|
21
|
+
1. Capture a snapshot or screenshot before acting.
|
|
22
|
+
2. Choose a target from the latest page structure.
|
|
23
|
+
3. Perform one structural action: click, type, keypress, drag, scroll, navigate, or resize.
|
|
24
|
+
4. Wait for the expected state with a bounded assertion, then capture fresh evidence. Do not reuse stale element references after navigation or structural changes.
|
|
25
|
+
5. Verify the visible result and the relevant side effect, such as a saved file, persisted record, or response. A screenshot alone does not prove a workflow succeeded.
|
|
26
|
+
6. Save before/after artifacts for bug fixes or comparisons. Keep evidence outside disposable app state.
|
|
27
|
+
7. Clean up dev servers, test profiles, and browser instances this run created, including failed attempts. Confirm the evidence still exists. Leave pre-existing user sessions running.
|
|
28
|
+
|
|
29
|
+
Use coordinates only when semantic targeting is unavailable and a fresh screenshot identifies the target. For visual parity, compare against the reference at the same viewport, scale, theme, and font state.
|
|
30
|
+
|
|
31
|
+
## Playwright recipe
|
|
32
|
+
|
|
33
|
+
If the repo already has Playwright, adapt the URL, action, and assertion to the actual user flow. Run the probe where the existing package resolves. This example creates its own browser and evidence directory:
|
|
34
|
+
|
|
35
|
+
```javascript
|
|
36
|
+
import { chromium } from "playwright";
|
|
37
|
+
import { mkdtemp } from "node:fs/promises";
|
|
38
|
+
import { tmpdir } from "node:os";
|
|
39
|
+
import { join } from "node:path";
|
|
40
|
+
|
|
41
|
+
const evidence = await mkdtemp(join(tmpdir(), "ui-harness-"));
|
|
42
|
+
const browser = await chromium.launch();
|
|
43
|
+
try {
|
|
44
|
+
const page = await browser.newPage({ viewport: { width: 1280, height: 800 } });
|
|
45
|
+
await page.goto("http://127.0.0.1:<port>");
|
|
46
|
+
await page.screenshot({ path: join(evidence, "before.png"), fullPage: true });
|
|
47
|
+
await page.getByRole("button", { name: /submit/i }).click();
|
|
48
|
+
await page.getByRole("status").filter({ hasText: /saved/i }).waitFor();
|
|
49
|
+
await page.screenshot({ path: join(evidence, "after.png"), fullPage: true });
|
|
50
|
+
console.log(evidence);
|
|
51
|
+
} finally {
|
|
52
|
+
await browser.close();
|
|
53
|
+
}
|
|
54
|
+
```
|
|
55
|
+
|
|
56
|
+
Do not add Playwright solely for this probe when existing external browser tools can drive the app. Adapt selectors, ports, and paths from the current repo.
|
|
57
|
+
|
|
58
|
+
## Electron and CDP
|
|
59
|
+
|
|
60
|
+
Use the project's Electron automation when available. For a Chromium debug port, Playwright's `chromium.connectOverCDP("http://127.0.0.1:<debug-port>")` can expose pages through `browser.contexts().flatMap(context => context.pages())`. Select the page with a positive app-root marker; use a negative marker to exclude another surface if needed. Inventory titles and URLs when no page matches.
|
|
61
|
+
|
|
62
|
+
Track ownership when attaching. Disconnect the automation connection using the driver's supported mechanism; close the app or browser only if this run launched it. Follow the environment's browser-control restrictions when deciding whether direct CDP is available.
|
|
63
|
+
|
|
64
|
+
Use raw CDP only when higher-level APIs are insufficient:
|
|
65
|
+
|
|
66
|
+
- Performance: CPU profiles, traces, paint flashing, FPS, and layout shift inspection.
|
|
67
|
+
- Memory: heap snapshots and forced GC for leak investigations.
|
|
68
|
+
- Network: request blocking, throttling, cache controls, and request/response logs.
|
|
69
|
+
- Rendering: viewport, color scheme, reduced motion, and accessibility checks.
|
|
70
|
+
- Debugging: console streaming, exceptions, and DOM snapshots.
|
|
71
|
+
|
|
72
|
+
Keep test data disposable and credentials out of artifacts. Report the route, actions, observed result, evidence paths, and any unverified behavior.
|
|
@@ -0,0 +1,34 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: correct
|
|
3
|
+
description: "Find the mistakes agents keep repeating in this repo and make each one impossible. Try architecture first, then types, then a lint whose error names the fix, then a test, and write docs last. Prove each check fails on a real past mistake. Repeat this each time the operator corrects you. Use for /correct."
|
|
4
|
+
disable-model-invocation: true
|
|
5
|
+
---
|
|
6
|
+
|
|
7
|
+
# Correct
|
|
8
|
+
|
|
9
|
+
The operator keeps correcting agents in this repo for the same mistakes. Change the repo so the next agent can't make them.
|
|
10
|
+
|
|
11
|
+
Assume every contributor is an agent that sees only the files it opened, copies the nearest example, and takes the shortest path that compiles. Design the repo so a change that looks right from one file is right for the whole repo.
|
|
12
|
+
|
|
13
|
+
## Find the mistake classes
|
|
14
|
+
|
|
15
|
+
First, read recent commits, reverts, review comments, agent instruction files, and comments that explain workarounds. Group the mistakes into classes. A class counts once it has happened twice.
|
|
16
|
+
|
|
17
|
+
## Fix each class at the highest level that works
|
|
18
|
+
|
|
19
|
+
1. **Eliminate it with architecture.** Give each piece of state one owner and each task one supported way. Hide internals so the wrong import fails. Replace hand-synced lists with one source of truth. Delete old ways and dead code an agent would copy.
|
|
20
|
+
2. **Enforce it with types so the bad state can't be written.** If bad code still compiles, add a lint or CI check whose error names the file, type, or function to use instead. If the pattern is already common, fail only when a change adds more.
|
|
21
|
+
3. **Test the behavior.** Fix or delete any test that would still pass if every function it calls returned nothing.
|
|
22
|
+
4. **Write docs or agent rules last, only for judgment calls.** Nothing fails when an agent skips them.
|
|
23
|
+
|
|
24
|
+
## Fix and prove
|
|
25
|
+
|
|
26
|
+
Then fix the most frequent classes now, one commit each. Prove each new check fails on a real past mistake. Run the same command locally and in CI. Exceptions go on the offending line with a reason, an expiry date, and a human's approval.
|
|
27
|
+
|
|
28
|
+
Load [`deslop`](../deslop/SKILL.md) before each fix commit and rerun affected checks after cleanup. If a correction affects web, IDE, or Electron behavior, load [`control-ui`](../control-ui/SKILL.md) for before/after proof; for CLI/TUI, load [`control-cli`](../control-cli/SKILL.md). Read [`pstack-harness`](../pstack-harness/SKILL.md) for invocation and use any project verification skill for app-specific steps.
|
|
29
|
+
|
|
30
|
+
## Keep the rule table
|
|
31
|
+
|
|
32
|
+
Last, keep a table in the agent instruction file that pairs each rule with what enforces it. When the operator corrects you, fix the mistake and add the rule. If the rule was already there and nothing enforces it, that's a repeat, so fix it at the highest level in the same change. Drop a rule once its mistake can't happen.
|
|
33
|
+
|
|
34
|
+
**Reply:** each class with its evidence, the level you picked, and why a higher level didn't work.
|
|
@@ -0,0 +1,47 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: create-verification-skill
|
|
3
|
+
description: "Generate a project-local verification skill that drives your app the way a user does, for any language, framework, or platform. Use for /create-verification-skill, \"make a verification skill for this repo\", or when a project has no scripted way to prove UI/CLI/service behavior."
|
|
4
|
+
disable-model-invocation: true
|
|
5
|
+
---
|
|
6
|
+
|
|
7
|
+
# Create a verification skill
|
|
8
|
+
|
|
9
|
+
Every serious project needs a scripted way to drive the real app and prove behavior: launch it, exercise a feature the way a user would, and capture evidence. This skill generates that as a project-local skill (`<project skills>/verify-<app>/`, where `<project skills>` is `.claude/skills` on Claude Code and `.agents/skills` on Codex, per the harness **project skills** row) tailored to the repo. You write the generator's output for the next agent, not for a human: it will be read cold, mid-task, by an agent that has never seen the app.
|
|
10
|
+
|
|
11
|
+
## 1. Interview the repo, not the user
|
|
12
|
+
|
|
13
|
+
Answer these from the codebase and only ask the user what you cannot observe:
|
|
14
|
+
|
|
15
|
+
- **Surface:** what does a user actually touch? A web UI, a CLI/TUI, a desktop app, an API, a mobile app, a library? A repo can have several; pick the primary one and note the rest.
|
|
16
|
+
- **Run:** how does the app start locally? Prefer the repo's own documented dev command (package scripts, Makefile, README quickstart). Note ports, env vars, seed data, auth.
|
|
17
|
+
- **Drive:** how can an agent interact with it programmatically? Load bundled [`control-ui`](../control-ui/SKILL.md) for web, IDE, and Electron or [`control-cli`](../control-cli/SKILL.md) for CLI/TUI. Adapt its workflow using existing Playwright/Cypress specs, expect scripts, PTY helpers, or a debug port before choosing another available driver. Use plain HTTP for services.
|
|
18
|
+
- **Observe:** what evidence can be captured? Screenshots, terminal transcripts, response bodies, logs, exit codes, DB state.
|
|
19
|
+
- **Isolate:** can two instances run side by side (ports, data dirs, profiles)? If not, say so in the generated skill: refusing to double-drive a shared instance beats corrupting the user's session.
|
|
20
|
+
|
|
21
|
+
If the checkout doesn't build or start as-is, fix that first (or report it precisely) before generating; a skill written against a broken base teaches wrong steps. When an irrelevant missing asset blocks startup (a static dir the API never serves, a sample config), the generated skill may create it, clearly marked as verification scaffolding, and remove it in cleanup.
|
|
22
|
+
|
|
23
|
+
## 2. Generate the skill
|
|
24
|
+
|
|
25
|
+
Write `<project skills>/verify-<app>/SKILL.md` with YAML frontmatter (`name: verify-<app>` and a `description` that names the app, the surface, and when to reach for it — without frontmatter the skill never registers) and these sections, each grounded in what the interview actually found (no placeholders left):
|
|
26
|
+
|
|
27
|
+
- **Launch:** the exact command that starts the app for verification, and how to tell it's ready (a log line, a port answering, a prompt). Include teardown. For a short-lived CLI or TUI there is no server to keep alive: launch means build the binary (or install deps) once, then start each drive in its own isolated PTY or tmux session.
|
|
28
|
+
- **Doctor:** one read-only check that answers "is this instance worth driving?" — process up, right version/build, port owned by us, auth valid. An agent runs this first whenever anything looks off.
|
|
29
|
+
- **Drive:** the harness recipe with real selectors/commands from this repo, not examples. Prefer stable handles (ARIA labels, data attributes, prompt strings, route paths) over coordinates and tab order.
|
|
30
|
+
For web, IDE, or Electron, the generated skill explicitly loads bundled `control-ui`; for CLI/TUI, `control-cli`, through the harness **skill** row. Then supply this app's exact recipe and feature map. This is an app-specific verification skill, not a separate generic `control` skill.
|
|
31
|
+
- **Evidence:** what to capture for a proof and where it goes. State the proof standards: exercise the real user path, not internal setters or test-only endpoints; capture the action and the resulting state, not just the final screen; verify side effects (files written, rows inserted, messages sent) alongside what's visible; mocks only where a production boundary already isolates the external system. When the safe path is a dry-run or test mode, verify what it actually skips by observing (files, network, git refs) rather than trusting its name: some dry-runs still touch the network or open a browser.
|
|
32
|
+
- **Cleanup:** how to tear down instances the run created. Never kill by process name; kill what you started. Cleanup removes instances and scratch state, never the evidence: proof artifacts survive the teardown, in a location the skill names.
|
|
33
|
+
- **Helpers:** any script the skill ships is executable and its invocation is shown in the skill body. A helper the reader has to reverse-engineer is not a helper.
|
|
34
|
+
|
|
35
|
+
## 3. Seed the feature map
|
|
36
|
+
|
|
37
|
+
Create `<project skills>/verify-<app>/features/README.md` plus one file per user-facing feature you can identify (aim for the top 3-5 to start, from routes, commands, menus, or docs). Follow the shape in [`references/feature-map-example/`](references/feature-map-example/), with a README index and one file per feature. Each file answers, from the user's point of view: what the feature is, how to reach it, how to drive it with the harness, and what observable end state proves it works. The four H2s are `Sub-features`, `How to get to it (user POV)`, `Driving it with <harness>`, and `Gotchas`. The map is the repo's maintained verification source; a proof that drives one convenient entry point is incomplete when the map lists others.
|
|
38
|
+
|
|
39
|
+
## 4. Prove the generated skill before handing it over
|
|
40
|
+
|
|
41
|
+
If generation adds executable helpers, load [`deslop`](../deslop/SKILL.md) on that code before committing or the final proof run.
|
|
42
|
+
|
|
43
|
+
Run its own instructions end to end once: launch, doctor, drive ONE mapped feature (one is enough; the map exists so later runs can cover the rest), capture evidence, clean up. After cleanup, confirm the evidence still exists at the named location — a cleanup that eats the proof fails this step. Fix what fails, and run the generated cleanup after every failed iteration too, so broken attempts don't strand processes and ports. A generated skill that was never executed is a draft, not a deliverable.
|
|
44
|
+
|
|
45
|
+
## 5. Offer the maintenance loop
|
|
46
|
+
|
|
47
|
+
Point the user at `/maintain-verification-skill` for keeping the map honest as the app changes. Suggest a cadence only if they ask.
|
|
@@ -0,0 +1,47 @@
|
|
|
1
|
+
# Notes verification map
|
|
2
|
+
|
|
3
|
+
This directory is the maintained source for verifying the user-facing behavior of Notes. Read the index before driving the app, then use the matching feature file as the recipe.
|
|
4
|
+
|
|
5
|
+
## Baseline preconditions
|
|
6
|
+
|
|
7
|
+
- Launch Notes at `http://127.0.0.1:4173` with a disposable data directory.
|
|
8
|
+
- Set `NOTES_DATA_DIR=/tmp/notes-verify-$RUN_ID` so concurrent runs do not share state.
|
|
9
|
+
- Seed notes titled `Quarterly plan` and `Grocery list`.
|
|
10
|
+
- Put `control-notes` and the `notes` CLI on `PATH`.
|
|
11
|
+
- Run `control-notes doctor` and require the expected URL, data directory, and build revision.
|
|
12
|
+
- Never drive an instance that was not started by this verification run.
|
|
13
|
+
|
|
14
|
+
## Driving conventions
|
|
15
|
+
|
|
16
|
+
- Start every recipe from the baseline state unless its preconditions say otherwise.
|
|
17
|
+
- Prefer ARIA roles and accessible names over CSS selectors or DOM position.
|
|
18
|
+
- Treat every command as literal. Keep quoted names and flags unchanged.
|
|
19
|
+
- Run browser actions through `control-notes browser`.
|
|
20
|
+
- Run terminal actions through `control-notes cli -- <command>`.
|
|
21
|
+
- Restore seeded data after a mutation. Do not remove proof artifacts during cleanup.
|
|
22
|
+
|
|
23
|
+
## Proof and skip reporting
|
|
24
|
+
|
|
25
|
+
- Capture the user action and the resulting state, not only the final screen.
|
|
26
|
+
- UI proof includes an ARIA snapshot and a screenshot with the app identity visible.
|
|
27
|
+
- CLI proof includes the command, stdout, stderr, and exit code.
|
|
28
|
+
- Mutation proof includes a read-only second view of the stored value.
|
|
29
|
+
- Record the feature ID and entry point used with every artifact.
|
|
30
|
+
- Report an unreachable path with the attempted command and the unmet precondition.
|
|
31
|
+
- Do not report a skipped entry point as verified through a different path.
|
|
32
|
+
|
|
33
|
+
## Feature entry contract
|
|
34
|
+
|
|
35
|
+
Each feature file starts with an H1 title and one paragraph describing the user-visible behavior. It then uses exactly four H2 sections in this order.
|
|
36
|
+
|
|
37
|
+
1. `Sub-features` lists short IDs with one line for each behavior.
|
|
38
|
+
2. `How to get to it (user POV)` lists every user entry point.
|
|
39
|
+
3. `Driving it with <harness>` starts with `Preconditions:` and uses labeled bullets that pair each user action with an exact command and observable result.
|
|
40
|
+
4. `Gotchas` lists traps that can waste or invalidate a verification run.
|
|
41
|
+
|
|
42
|
+
Keep implementation details out of the map. Name only user paths, stable handles, required state, commands, and observable proof.
|
|
43
|
+
|
|
44
|
+
## Features
|
|
45
|
+
|
|
46
|
+
- [Create a note](./create-note.md) covers browser and CLI creation, cancellation, persistence, and cleanup.
|
|
47
|
+
- [Search notes](./search.md) covers toolbar, keyboard, and CLI search with matching, empty, and clear states.
|
|
@@ -0,0 +1,39 @@
|
|
|
1
|
+
# Create a note
|
|
2
|
+
|
|
3
|
+
Create note lets a user save a titled note from the browser or CLI, cancel an unfinished draft, and confirm the saved note from a second user-facing view.
|
|
4
|
+
|
|
5
|
+
## Sub-features
|
|
6
|
+
|
|
7
|
+
- `create-open` opens a blank editor from each browser entry point.
|
|
8
|
+
- `create-save` persists a title and body.
|
|
9
|
+
- `create-cancel` discards an unfinished browser draft.
|
|
10
|
+
- `create-cli` creates the same note shape from the terminal.
|
|
11
|
+
|
|
12
|
+
## How to get to it (user POV)
|
|
13
|
+
|
|
14
|
+
- Choose the `New note` button in the browser toolbar.
|
|
15
|
+
- Press `n` in the browser while focus is outside an editable field.
|
|
16
|
+
- Run `notes create --title <title> --body <body>` in a terminal.
|
|
17
|
+
|
|
18
|
+
## Driving it with control-notes
|
|
19
|
+
|
|
20
|
+
Preconditions:
|
|
21
|
+
|
|
22
|
+
- Notes is healthy at `http://127.0.0.1:4173`.
|
|
23
|
+
- No note is titled `Release checklist`.
|
|
24
|
+
- `control-notes doctor` reports the expected URL and disposable data directory.
|
|
25
|
+
|
|
26
|
+
- **Open editor.** Choose `New note`. Run `control-notes browser click --role button --name "New note"`. A form named `Note editor` appears with focus in the `Title` textbox.
|
|
27
|
+
- **Enter content.** Type the title and body. Run `control-notes browser fill --role textbox --name "Title" --value "Release checklist"` and `control-notes browser fill --role textbox --name "Body" --value "Tag and publish"`. The `Save note` button becomes enabled.
|
|
28
|
+
- **Save note.** Choose `Save note`. Run `control-notes browser click --role button --name "Save note"`. A status named `Note saved` appears and the heading reads `Release checklist`.
|
|
29
|
+
- **Confirm persistence.** Return to the note list and reopen the note. Run `control-notes browser click --role link --name "All notes"` and `control-notes browser click --role link --name "Release checklist"`. The editor shows both saved values.
|
|
30
|
+
- **Cancel draft.** Open a new note, enter `Discard me`, and choose `Cancel`. Run `control-notes browser click --role button --name "New note"`, `control-notes browser fill --role textbox --name "Title" --value "Discard me"`, and `control-notes browser click --role button --name "Cancel"`. The note list returns and has no `Discard me` link.
|
|
31
|
+
- **CLI entry.** Create a second note. Run `control-notes cli -- notes create --title "CLI note" --body "Created from terminal" --format json`. Exit code `0` and stdout contain the new note ID and title.
|
|
32
|
+
- **Proof.** Reopen both saved notes from `All notes`. Run `control-notes browser snapshot --aria --path artifacts/create-note/list.aria.txt` and `control-notes browser screenshot --path artifacts/create-note/list.png`. The artifacts show `Release checklist` and `CLI note`.
|
|
33
|
+
|
|
34
|
+
## Gotchas
|
|
35
|
+
|
|
36
|
+
- Pressing `n` while a textbox has focus types the character instead of opening a new editor.
|
|
37
|
+
- Titles are trimmed on save. Assert the rendered title, not the draft input value.
|
|
38
|
+
- A save status alone is insufficient proof. Reopen the note from the list.
|
|
39
|
+
- Remove `Release checklist` and `CLI note` during fixture cleanup, but retain their proof artifacts.
|
|
@@ -0,0 +1,45 @@
|
|
|
1
|
+
# Search notes
|
|
2
|
+
|
|
3
|
+
Search lets a user find notes by title or body text, inspect a matching note, and distinguish no matches from an unavailable search.
|
|
4
|
+
|
|
5
|
+
## Sub-features
|
|
6
|
+
|
|
7
|
+
- `search-open` opens search from each supported browser entry point.
|
|
8
|
+
- `search-match` returns title and body matches without changing note data.
|
|
9
|
+
- `search-open-result` opens a result in the note editor.
|
|
10
|
+
- `search-empty` shows a complete empty state for a query with no matches.
|
|
11
|
+
- `search-clear` removes the query and restores the recent-notes view.
|
|
12
|
+
- `search-cli` returns the same matching notes from the terminal.
|
|
13
|
+
|
|
14
|
+
## How to get to it (user POV)
|
|
15
|
+
|
|
16
|
+
- Choose the `Search` button in the browser toolbar.
|
|
17
|
+
- Press `/` in the browser while focus is outside an editable field.
|
|
18
|
+
- Run `notes search <query>` in a terminal.
|
|
19
|
+
|
|
20
|
+
## Driving it with control-notes
|
|
21
|
+
|
|
22
|
+
Preconditions:
|
|
23
|
+
|
|
24
|
+
- Notes is healthy at `http://127.0.0.1:4173`.
|
|
25
|
+
- The disposable data directory contains `Quarterly plan` with body text `Draft budget`.
|
|
26
|
+
- `control-notes doctor` reports the expected URL and data directory.
|
|
27
|
+
|
|
28
|
+
- **Toolbar entry.** Choose the `Search` button. Run `control-notes browser click --role button --name "Search"`. A dialog named `Search notes` appears with focus in its searchbox.
|
|
29
|
+
- **Keyboard entry.** Close the dialog, focus the page, and press `/`. Run `control-notes browser press --key "/"`. The same dialog appears and the page does not insert a slash.
|
|
30
|
+
- **Title match.** Type `quarterly`. Run `control-notes browser fill --role searchbox --name "Search notes" --value "quarterly"`. The `Search results` list contains `Quarterly plan` and does not contain `Grocery list`.
|
|
31
|
+
- **Body match.** Replace the query with `budget`. Run `control-notes browser fill --role searchbox --name "Search notes" --value "budget"`. The result `Quarterly plan` remains visible with a body-match excerpt.
|
|
32
|
+
- **Open result.** Choose `Quarterly plan`. Run `control-notes browser click --role link --name "Quarterly plan"`. The dialog closes and the editor heading reads `Quarterly plan`.
|
|
33
|
+
- **Empty state.** Reopen search and enter `volcano`. Run `control-notes browser fill --role searchbox --name "Search notes" --value "volcano"`. A status named `No matching notes` appears after search completes.
|
|
34
|
+
- **Clear query.** Choose `Clear search`. Run `control-notes browser click --role button --name "Clear search"`. The searchbox is empty and the `Recent notes` region replaces the result list.
|
|
35
|
+
- **CLI match.** Search from the terminal. Run `control-notes cli -- notes search "quarterly" --format json`. Exit code `0` and stdout contain one object whose title is `Quarterly plan`.
|
|
36
|
+
- **CLI miss.** Search for an absent value. Run `control-notes cli -- notes search "volcano" --format json`. Exit code `0` and stdout are `[]`.
|
|
37
|
+
- **Proof.** Capture the populated result state. Run `control-notes browser snapshot --aria --path artifacts/search/results.aria.txt` and `control-notes browser screenshot --path artifacts/search/results.png`. Both artifacts identify Notes, the query, and `Quarterly plan`.
|
|
38
|
+
|
|
39
|
+
## Gotchas
|
|
40
|
+
|
|
41
|
+
- Pressing `/` while the editor or searchbox has focus inserts text instead of opening search.
|
|
42
|
+
- Results update after a short debounce. Wait for the results list or empty status, not a fixed sleep.
|
|
43
|
+
- Archived notes are excluded unless the user enables `Include archived`.
|
|
44
|
+
- The CLI defaults to human-readable output. Use `--format json` for stable assertions.
|
|
45
|
+
- Opening a result changes browser state. Reopen search before proving another query.
|
|
@@ -0,0 +1,30 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: deslop
|
|
3
|
+
description: Remove AI-generated code slop from a branch or working diff before committing. Use for /deslop, code cleanup, unnecessary comments, redundant guards, type escape hatches, or patterns that do not fit the surrounding code.
|
|
4
|
+
disable-model-invocation: true
|
|
5
|
+
---
|
|
6
|
+
|
|
7
|
+
# Remove AI code slop
|
|
8
|
+
|
|
9
|
+
Check the diff and remove AI-generated slop introduced by the current work. This is code cleanup; use `unslop` for prose and `no-comments` for the independent comment review.
|
|
10
|
+
|
|
11
|
+
## Choose the diff
|
|
12
|
+
|
|
13
|
+
Use the base the user or PR names. Otherwise resolve the repository's default branch from its documentation or `refs/remotes/origin/HEAD`; do not assume it is called `main`. Compare from the merge base (`git diff <base>...HEAD`) and inspect staged and unstaged changes (`git diff --cached`, `git diff`). Read new untracked files that belong to the task too. If the base is unavailable locally, inspect the working diff and report the limited scope.
|
|
14
|
+
|
|
15
|
+
Read the surrounding code before editing. Clean the current task's changes; leave unrelated user edits alone.
|
|
16
|
+
|
|
17
|
+
## Focus areas
|
|
18
|
+
|
|
19
|
+
- Extra comments that are unnecessary or inconsistent with local style.
|
|
20
|
+
- Defensive checks or try/catch blocks that are abnormal for trusted code paths.
|
|
21
|
+
- Casts to `any` used only to bypass type issues.
|
|
22
|
+
- Deeply nested code that should be simplified with early returns.
|
|
23
|
+
- Dead compatibility paths, needless abstractions, and placeholder names.
|
|
24
|
+
- Other patterns inconsistent with the file and surrounding codebase.
|
|
25
|
+
|
|
26
|
+
## Finish
|
|
27
|
+
|
|
28
|
+
Keep behavior unchanged and prefer minimal edits over broad rewrites. Preserve validation at external boundaries and comments that explain a real constraint. If cleanup exposes a bug, handle it within the user's task or report it separately.
|
|
29
|
+
|
|
30
|
+
Inspect the resulting diff and run the checks appropriate to the code you changed. In a pstack commit workflow, follow with `no-comments` per the harness **deslop** row. Keep the final summary to 1-3 sentences, including any checks that could not run.
|
|
@@ -0,0 +1,55 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: figure-it-out
|
|
3
|
+
description: "Design an auditable playbook when no narrower one fits: a large migration, an ambitious multi-part change, or work a human reviews after stepping away. Scales rigor to the task, runs a hypothesis loop, and logs decisions via show-me-your-work. Use for /figure-it-out, 'figure it out', a large migration, or when no narrower playbook applies."
|
|
4
|
+
disable-model-invocation: true
|
|
5
|
+
---
|
|
6
|
+
|
|
7
|
+
# Figure it out
|
|
8
|
+
|
|
9
|
+
When the task matches no playbook, design one. The deliverable before any code is the workflow itself: a sequence of phases that scales rigor to the task, runs the scientific method, and leaves a decision trail a human can audit after stepping away.
|
|
10
|
+
|
|
11
|
+
## Start
|
|
12
|
+
|
|
13
|
+
Open a todolist whose first item is to read the Principles section of the **poteto-mode** skill. Then add the phases below as todos.
|
|
14
|
+
|
|
15
|
+
## Phase A: Frame
|
|
16
|
+
|
|
17
|
+
Ground first, then commit. Don't start the run until you can state:
|
|
18
|
+
|
|
19
|
+
- The definition of done as a falsifiable predicate (the **prove-it-works** principle skill).
|
|
20
|
+
- Scope, quantified: rough units and effort, plus the blockers grounding surfaced.
|
|
21
|
+
- The rigor level, biased high. One-way doors and high blast radius get more. Reversible low-stakes steps get less. Rigor is gates and artifacts, not "try harder".
|
|
22
|
+
|
|
23
|
+
Present the framing and tradeoffs before committing to a long run. Reversible work proceeds (the **never-block-on-the-human** principle skill), but a multi-hour run earns one checkpoint.
|
|
24
|
+
|
|
25
|
+
## Phase B: Design the workflow
|
|
26
|
+
|
|
27
|
+
Decompose into atomic, independently-landable units. Sequence riskiest-unknown-first. Scaffold and verification come before features (the **foundational-thinking** principle skill).
|
|
28
|
+
|
|
29
|
+
- Build the verification harness before the work, with the baseline captured from the pre-change state, so the check reads as "old value vs new value".
|
|
30
|
+
- For web, IDE, or Electron units, load [`control-ui`](../control-ui/SKILL.md); for CLI/TUI units, [`control-cli`](../control-cli/SKILL.md). Read [`pstack-harness`](../pstack-harness/SKILL.md) and apply its **verification harness** row with any project verification skill. Put the driver path and calls in each delegated unit's brief.
|
|
31
|
+
- For one-way-door design decisions, run the **architect** skill (it runs **arena**). Skip it for mechanical work whose shape is already concrete. A second arena over a settled design is over-engineering (the **laziness-protocol** principle skill).
|
|
32
|
+
- Decide what fans out. Parallelize only across seams, and give each worker its own worktree or branch (the **separate-before-serializing-shared-state** principle skill). Don't over-fan.
|
|
33
|
+
- Write the designed phase list down. That list is what the human reviews.
|
|
34
|
+
|
|
35
|
+
Then execute the design. Add its steps to the todolist as concrete items, after the Phase C entry and before Phase D. Run each under the Phase C loop discipline, and weave the Phase D log through them, a row as each step lands, rather than saving the whole trail for the end.
|
|
36
|
+
|
|
37
|
+
## Phase C: Run the loop
|
|
38
|
+
|
|
39
|
+
Each unit is an experiment. State the hypothesis, make the smallest change, measure against the predicate on the real artifact, keep it if it advanced, revert it if it didn't.
|
|
40
|
+
Load [`deslop`](../deslop/SKILL.md) before committing or handing back each kept code unit, then rerun the checks its cleanup affects. Include this in code-writing workers' briefs.
|
|
41
|
+
Apply the **sequence-verifiable-units** principle skill, verifying each unit before starting the next instead of batching checks at the end.
|
|
42
|
+
|
|
43
|
+
- Verify by inspecting the artifact, never a self-report. When something passes too easily, suspect the observation method before the system.
|
|
44
|
+
- Pair delegated work with a judge. If a worker games the gate, reset and harden the contract. If the gate itself is wrong, fix the gate in its own change rather than routing around it.
|
|
45
|
+
- A verdict is VERIFIED, NOT VERIFIED, or INCONCLUSIVE. Inconclusive is not a pass. Don't hide a negative.
|
|
46
|
+
|
|
47
|
+
## Phase D: Keep the audit trail
|
|
48
|
+
|
|
49
|
+
Log the run via the **show-me-your-work** skill. figure-it-out's work is usually ambitious enough to commit the trail so the reviewer can read it in the PR. The trail plus the diff is what lets the human come back and trust the work.
|
|
50
|
+
|
|
51
|
+
## Phase E: Verify and hand back
|
|
52
|
+
|
|
53
|
+
Check the whole against the Phase A predicate on the real product, not just the harness. Encode any recurring correction as a gate, a lint rule, a check, or a script (the **encode-lessons-in-structure** principle skill).
|
|
54
|
+
|
|
55
|
+
**Reply:** the playbook you designed, the rigor level and why, the decision-trail path, what's verified against the predicate, and what's still open.
|
|
@@ -0,0 +1,58 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: how
|
|
3
|
+
description: "Use for \"how does X work\", code walkthroughs before changing something, and placement / ownership / layering questions (\"where should this live\", \"which package owns this\", \"is this the right layer\"). Explains subsystem architecture, runtime flow, onboarding mental models. Use why for motivation."
|
|
4
|
+
disable-model-invocation: true
|
|
5
|
+
---
|
|
6
|
+
|
|
7
|
+
# How
|
|
8
|
+
|
|
9
|
+
Explore the codebase to answer "how does X work?" questions. Produce architectural explanations at the level of a senior engineer onboarding onto a subsystem, enough to build a working mental model, not so much that it reads like annotated source code.
|
|
10
|
+
|
|
11
|
+
Read the **pstack-harness** skill first. Each spawn below names a role line in the pstack models file (`~/.pstack/models.md`, harness **models file** row) and a default (harness **tiers** row). Set `model` to that line's value, or to the default if the file or the line is missing. Leave `model` unset when the value is `auto` or `inherit-parent`. If the spawn rejects a model, use the default and say so. If it rejects the default, use the closest valid model of the same family from its error message.
|
|
12
|
+
|
|
13
|
+
## Step 1. Assess Complexity
|
|
14
|
+
|
|
15
|
+
If the scope is ambiguous, state your interpretation and explore. The user can redirect.
|
|
16
|
+
|
|
17
|
+
- **Simple** (a single module, a small utility, a narrow question such as "how does function X work"): no explorers. One explainer explores and explains in a single pass. Go to Step 2b.
|
|
18
|
+
- **Complex** (a subsystem spanning multiple files or services, a cross-cutting feature, a full architectural overview): spawn parallel explorers first, then hand off to the explainer. Go to Step 2a.
|
|
19
|
+
|
|
20
|
+
When in doubt, take the simple path.
|
|
21
|
+
|
|
22
|
+
## Step 2a. Explore (complex questions only)
|
|
23
|
+
|
|
24
|
+
Decompose the question into 2 to 4 exploration angles, each a distinct slice of the subsystem. Spawn all explorers in a single message:
|
|
25
|
+
|
|
26
|
+
- `subagent_type`: `general-purpose` (harness **default agent** row)
|
|
27
|
+
- `model`: the `how explorer` line, default the code tier
|
|
28
|
+
- `readonly`: `true` (harness **read-only** row)
|
|
29
|
+
|
|
30
|
+
Each explorer gets the prompt in `references/explorer-prompt.md` with its angle filled in. Then go to Step 3.
|
|
31
|
+
|
|
32
|
+
## Step 2b. Direct Explain (simple questions)
|
|
33
|
+
|
|
34
|
+
Spawn one subagent that explores and explains in one pass:
|
|
35
|
+
|
|
36
|
+
- `subagent_type`: `general-purpose` (harness **default agent** row)
|
|
37
|
+
- `model`: the `how explainer` line, default the judgment tier
|
|
38
|
+
- `readonly`: `true` (harness **read-only** row)
|
|
39
|
+
|
|
40
|
+
Build its prompt from `references/explainer-prompt.md` without the explorer-findings section. Go to Step 4.
|
|
41
|
+
|
|
42
|
+
## Step 3. Synthesize (complex questions only)
|
|
43
|
+
|
|
44
|
+
Once all explorers have returned, spawn one subagent to synthesize their findings into one explanation:
|
|
45
|
+
|
|
46
|
+
- `subagent_type`: `general-purpose` (harness **default agent** row)
|
|
47
|
+
- `model`: the `how explainer` line, default the judgment tier
|
|
48
|
+
- `readonly`: `true` (harness **read-only** row)
|
|
49
|
+
|
|
50
|
+
Build its prompt from `references/explainer-prompt.md` with every explorer's findings filled in.
|
|
51
|
+
|
|
52
|
+
## Step 4. Present
|
|
53
|
+
|
|
54
|
+
Present the explainer's output to the user. Light edits for clarity or context from the conversation are fine. Do not substantially rewrite it.
|
|
55
|
+
|
|
56
|
+
## Output Format
|
|
57
|
+
|
|
58
|
+
The explanation uses the sections defined in `references/explainer-prompt.md`, dropping any that do not apply: Overview, Key Concepts, How It Works, Where Things Live, Gotchas.
|