leos-agent 6.3.0 → 10.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (84) hide show
  1. package/LICENSE +21 -0
  2. package/README.md +547 -24
  3. package/commands/handoff.md +11 -0
  4. package/commands/handon.md +10 -0
  5. package/commands/leo-doctor.md +22 -0
  6. package/commands/leo-install.md +9 -0
  7. package/commands/review-pr.md +9 -0
  8. package/commands-claude/watch-review.md +9 -0
  9. package/index.js +12 -0
  10. package/package.json +30 -18
  11. package/payload/codex-agents/leo-executor.toml +36 -0
  12. package/payload/codex-agents/leo-runner.toml +28 -0
  13. package/rules/preferences.md +97 -0
  14. package/scripts/check.py +244 -0
  15. package/scripts/ghreview.py +24 -6
  16. package/scripts/handoff.py +183 -0
  17. package/scripts/leo-install.py +509 -0
  18. package/scripts/measure_context.py +113 -0
  19. package/scripts/publish-npm.py +138 -0
  20. package/scripts/resolve_attach_target.py +45 -13
  21. package/scripts/watch_review.py +169 -0
  22. package/skills/doctor/SKILL.md +73 -96
  23. package/skills/doctor/agents/openai.yaml +5 -0
  24. package/skills/handoff/SKILL.md +99 -0
  25. package/skills/handoff/agents/openai.yaml +5 -0
  26. package/skills/handon/SKILL.md +61 -0
  27. package/skills/install/SKILL.md +79 -0
  28. package/skills/install/agents/openai.yaml +5 -0
  29. package/skills/review-pr/SKILL.md +59 -308
  30. package/skills/review-pr/reference/lenses.md +67 -0
  31. package/skills/review-pr/reference/procedure.md +348 -0
  32. package/skills-claude/attach-pr/SKILL.md +178 -0
  33. package/skills-claude/watch-review/SKILL.md +91 -0
  34. package/adapters/cursor/agents/executor.md +0 -17
  35. package/adapters/cursor/agents/expert.md +0 -70
  36. package/adapters/cursor/agents/explore.md +0 -16
  37. package/adapters/cursor/agents/implementer.md +0 -18
  38. package/adapters/cursor/agents/investigator.md +0 -18
  39. package/adapters/cursor/agents/planner.md +0 -28
  40. package/adapters/cursor/agents/reviewer.md +0 -34
  41. package/adapters/opencode/agents.json +0 -66
  42. package/adapters/opencode/plugin.js +0 -288
  43. package/config/models.json +0 -408
  44. package/hooks/bash-guard.py +0 -541
  45. package/hooks/cursor-guard.py +0 -84
  46. package/hooks/hooks-cursor.json +0 -11
  47. package/hooks/hooks.json +0 -20
  48. package/hooks/session-start.py +0 -148
  49. package/roles/executor.md +0 -15
  50. package/roles/expert.md +0 -67
  51. package/roles/explore.md +0 -13
  52. package/roles/implementer.md +0 -16
  53. package/roles/investigator.md +0 -15
  54. package/roles/planner.md +0 -25
  55. package/roles/reviewer.md +0 -31
  56. package/scripts/doctor.py +0 -284
  57. package/scripts/memory.py +0 -705
  58. package/scripts/render_adapters.py +0 -473
  59. package/scripts/setup.py +0 -161
  60. package/settings.json +0 -7
  61. package/skills/.gitkeep +0 -0
  62. package/skills/brainstorming/SKILL.md +0 -109
  63. package/skills/debugging/SKILL.md +0 -98
  64. package/skills/delegation/SKILL.md +0 -141
  65. package/skills/executing-plans/SKILL.md +0 -116
  66. package/skills/finishing-a-branch/SKILL.md +0 -123
  67. package/skills/freshness/SKILL.md +0 -118
  68. package/skills/memory/SKILL.md +0 -144
  69. package/skills/resolve-ticket/SKILL.md +0 -269
  70. package/skills/setup/SKILL.md +0 -85
  71. package/skills/test-first/SKILL.md +0 -90
  72. package/skills/using-leo/SKILL.md +0 -96
  73. package/skills/using-leo/references/claude-mapping.md +0 -32
  74. package/skills/using-leo/references/codex-mapping.md +0 -34
  75. package/skills/using-leo/references/cursor-mapping.md +0 -34
  76. package/skills/using-leo/references/hermes-mapping.md +0 -36
  77. package/skills/using-leo/references/opencode-mapping.md +0 -36
  78. package/skills/verification/SKILL.md +0 -109
  79. package/skills/visual-verification/SKILL.md +0 -114
  80. package/skills/watch-review/SKILL.md +0 -125
  81. package/skills/worktrees/SKILL.md +0 -129
  82. package/skills/writing-plans/SKILL.md +0 -96
  83. package/skills/writing-skills/SKILL.md +0 -134
  84. package/workflows/cost-tiered-fix.js +0 -259
@@ -1,109 +0,0 @@
1
- ---
2
- name: brainstorming
3
- description: >
4
- Design gate before non-trivial code — proportional to blast radius and
5
- reversibility, the deliberate opposite of an unconditional gate. Contained,
6
- easily reversible changes clear with one sentence of rationale; changes with
7
- wide blast radius, hard to reverse, or that introduce new surface need
8
- genuine, viable alternatives with trade-offs weighed before any code gets
9
- written. Produces the chosen approach and its trade-offs, sized to the gate,
10
- handed off to leo:writing-plans.
11
- when_to_use: >
12
- Before starting non-trivial code: a new feature, a new integration surface,
13
- a schema or data-model change, anything that's expensive or awkward to
14
- undo. NOT for a contained, easily reversible tweak (that just needs one
15
- sentence of rationale, not this skill's full procedure), NOT for pure
16
- investigation (use investigator), and NOT for writing the plan itself
17
- (leo:writing-plans) — brainstorming stops at a chosen approach, it never
18
- slides into implementation.
19
- ---
20
-
21
- # brainstorming
22
-
23
- Core rule: the depth of the design gate is proportional to blast radius and
24
- reversibility, not to how the task felt when it landed. A one-line change to
25
- a private helper does not need three alternatives; a new public API or a
26
- schema migration does.
27
-
28
- ## Size the gate first
29
-
30
- Before generating anything, classify the change:
31
-
32
- - **Contained + easily reversible** (a local refactor, an internal helper, a
33
- flag you can flip back) → one sentence of rationale is enough. Say what
34
- you're doing and why, then move to leo:writing-plans or straight to
35
- implementation per the routing table.
36
- - **Wide blast radius, hard to reverse, or new surface** (public API, schema
37
- or data-model change, cross-service contract, anything users or other
38
- systems will come to depend on) → full gate: genuine alternatives with
39
- trade-offs, written down, before any code.
40
-
41
- When unsure which bucket, treat it as the wider one — the cost of one extra
42
- paragraph is nothing next to the cost of an unreversible wrong turn.
43
-
44
- ## Alternatives must be viable
45
-
46
- Every alternative in a full gate has to be something a reasonable engineer
47
- could actually ship and defend, not a strawman stood up to make the first
48
- idea look good by comparison. If you can't articulate a real reason someone
49
- would pick alternative B, it isn't an alternative — go find one that's
50
- actually competing for the job, or drop down to the one-sentence gate because
51
- there's really only one sane approach.
52
-
53
- Test: could you argue for this option in front of Leo without a "but
54
- obviously we won't do this" tone? If not, it's a strawman — cut it.
55
-
56
- ## Generation method
57
-
58
- To surface genuinely different options, vary along a different axis each
59
- time rather than producing three cosmetic variants of the same idea:
60
-
61
- 1. **Data model vs. control flow vs. boundary/interface** — would this
62
- problem look different if you moved the complexity into the data shape,
63
- into how execution flows, or into where the interface/boundary sits?
64
- 2. **Prior art in the repo** — grep for how this repo already solved a
65
- similar problem (via explore, not inline digging) and steal that pattern
66
- before inventing a new one. Consistency with existing structure is a real
67
- trade-off, not a tie-breaker of last resort.
68
- 3. **The 10x-simpler version** — what would this look like with one-tenth
69
- the code/config/moving parts? Even when you don't ship it, it's usually
70
- the sharpest lens on what the "proper" version is paying for.
71
-
72
- ## Output
73
-
74
- - Contained/reversible: one sentence of rationale, folded into the plan or
75
- the commit itself.
76
- - Full gate: the chosen approach plus the trade-offs record — a paragraph for
77
- a medium decision, a short doc for a genuinely high-stakes one. Sized to
78
- the gate, not padded to look thorough.
79
-
80
- Either way, the output is a decision, not code. Hand it to
81
- leo:writing-plans for the actual plan; brainstorming never slides into
82
- implementation itself.
83
-
84
- ## Self-talk to catch
85
-
86
- - "I'll just list two options so it looks considered" — if you can't defend
87
- both, that's a strawman, not a gate.
88
- - "This is a big change but I already know the answer" — blast radius and
89
- reversibility decide the gate size, not your confidence.
90
- - "I'll sketch the plan while I'm at it" — that's leo:writing-plans' job;
91
- stop at the chosen approach.
92
- - "One sentence feels thin for something this exciting" — excitement isn't
93
- blast radius; if it's contained and reversible, one sentence is correct.
94
-
95
- ## Escalation
96
-
97
- Planning-tier work runs at Opus per the routing table (plan mode in an Opus
98
- session, or the `planner` subagent otherwise). Escalate per the standard
99
- ladder: two failed passes at reaching a defensible set of alternatives step
100
- up a tier; a genuine deadlock between two Opus-tier framings goes to
101
- `expert`, announced in one line, never silently.
102
-
103
- ## Works with
104
-
105
- - leo:writing-plans — takes the chosen approach and turns it into an
106
- executable plan.
107
- - investigator — for questions that need evidence before a design question
108
- can even be framed.
109
- - explore — cheap prior-art search feeding the generation method above.
@@ -1,98 +0,0 @@
1
- ---
2
- name: debugging
3
- description: >
4
- Root-cause-before-fix loop for bugs, failing tests, crashes, and surprising
5
- behavior. Five named phases — Reproduce, Localize, Hypothesize, Prove, Fix —
6
- each with an exit criterion, so a fix never lands before the cause is
7
- pinned to file:line. Diagnosis is read-only judge work (investigator); the
8
- fix happens separately, at the routed tier.
9
- when_to_use: >
10
- Any bug report, failing test, crash, stack trace, or "why does X happen"
11
- before proposing a fix — used by the investigator agent and by the main
12
- loop ahead of any edit that touches broken behavior. NOT for planned
13
- feature work with no defect (that's planner), NOT for judging someone
14
- else's diff (that's reviewer), and NOT a substitute for leo:verification
15
- after the fix lands — this skill ends at Fix, verification is separate.
16
- ---
17
-
18
- # debugging
19
-
20
- Core rule: no fix before the cause is REPRODUCED and LOCATED at file:line. A
21
- symptom going away is not proof — it's a coincidence until the loop below
22
- says otherwise.
23
-
24
- ## When it fires
25
-
26
- Bug reports, failing tests, crashes, stack traces, flaky behavior, "this
27
- should work but doesn't." Route the diagnosis itself through `investigator`
28
- (Opus, read-only) per the model-routing table — this skill is its loop.
29
- Doesn't fire for greenfield feature work (no defect exists yet) or for
30
- diffing someone else's change (that's `reviewer`).
31
-
32
- ## The five phases
33
-
34
- Named exactly, run in order, each with an exit criterion. Do not skip a
35
- phase because the bug "looks obvious" — obvious bugs are exactly the ones
36
- where a wrong guess ships fastest.
37
-
38
- | Phase | Exit criterion |
39
- |---|---|
40
- | **Reproduce** | The failure fires on command — a test, a script, a repro sequence — not "worked once." No stable repro yet is itself a finding: report it, don't guess past it. |
41
- | **Localize** | The failure is traced to a specific **file:line**, not a subsystem or a vibe ("something in auth"). Read the actual code path the repro exercises; don't infer from names or docs. |
42
- | **Hypothesize** | One sentence: "X happens because file:line does Y instead of Z." One hypothesis at a time — write it down before touching anything. |
43
- | **Prove** | The smallest evidence that the hypothesis IS the cause, not just correlated with it. Where the surface is testable, that's a failing test written per `leo:test-first` — red on the bug, and its assertion names the file:line from Localize. Where nothing is testable (infra, timing, external system), the next-smallest evidence: a log line, a debugger break, a minimal repro script. |
44
- | **Fix** | The change that makes Prove's evidence pass. Happens at the routed tier (`executor` for mechanical, `implementer` for real changes) — never by the same pass that diagnosed it. |
45
-
46
- Reproduce and Localize can compress into one step for a trivial case (a
47
- crash with a one-frame stack trace pointing straight at the bug) — but
48
- Hypothesize and Prove never collapse into Fix. If you catch yourself editing
49
- code before you've written the hypothesis sentence, stop and back up.
50
-
51
- ## One hypothesis, one change
52
-
53
- Test one hypothesis at a time. If Fix doesn't clear Prove's evidence, the
54
- hypothesis was wrong — REVERT the change before forming the next one. Never
55
- stack a second speculative edit on top of a first that didn't pan out; you
56
- lose the ability to tell which change did what, and the diff stops being
57
- reviewable. Revert, re-enter Hypothesize with what the failed attempt taught
58
- you, and go again.
59
-
60
- ## Stuck: the ladder
61
-
62
- After two failures on the same cause (two hypotheses tried and reverted, still
63
- no Prove), step up one tier rather than retrying at the same one — investigator
64
- haiku-assist steps to full investigator, investigator itself steps to a
65
- second, more evidence-fed pass, capped at Opus. A genuine deadlock, or two
66
- Opus verdicts on the same cause that disagree → `expert`, announced in one
67
- line ("escalating to expert: <question>") before it's invoked, never silent.
68
- Don't loop a third time at the same tier hoping the next guess lands — that's
69
- the same failure mode as skipping Prove, just slower.
70
-
71
- ## Diagnosis and fix stay separate
72
-
73
- The phase that reaches the verdict (Reproduce through Prove) is read-only
74
- judge work — no edits, no reverts-of-other-people's-code, just evidence and a
75
- file:line. Whoever ran that pass hands the hypothesis and its proof to the
76
- executing tier for Fix. This mirrors why `reviewer` never patches what it
77
- finds: the same pass that wants to be right about the cause is a bad judge of
78
- whether it actually is. After Fix lands, `leo:verification` (or a plain
79
- `reviewer` pass on the diff) is the separate check that the fix is real and
80
- didn't just make Prove's specific probe go quiet.
81
-
82
- ## Self-talk to catch
83
-
84
- - "It's obviously the timeout" — obvious is not file:line; go Localize it.
85
- - "Passing now, good enough" — passing isn't Prove; did you write the
86
- failing-first check, or did the symptom just stop reproducing?
87
- - "I'll patch this and see if it helps" — that's skipping Hypothesize; name
88
- the mechanism before touching code.
89
- - "One more tweak on top, I'm close" — that's the stacked-edit trap; revert
90
- first.
91
- - "Third guess this tier, one more won't hurt" — it's the two-failures
92
- trigger; escalate instead.
93
-
94
- ## Works with
95
-
96
- `leo:test-first` for writing Prove's failing test. `leo:verification` for
97
- the post-Fix check. `investigator` runs this loop; `reviewer` judges the
98
- resulting diff once Fix is applied.
@@ -1,141 +0,0 @@
1
- ---
2
- name: delegation
3
- description: >
4
- Operational mechanics for dispatching subagents — a single spawn or a large
5
- fan-out — the companion to the policy's "Delegate the labor" section.
6
- Covers brief construction, model/effort pinning, the four-state return
7
- contract, and ledger-backed progress tracking for long multi-agent runs.
8
- when_to_use: >
9
- Any time work is routed to a subagent (explore, investigator, executor,
10
- implementer, reviewer, expert) rather than done inline — single dispatch or
11
- fan-out. NOT for deciding *which* tier a task belongs in (that's the
12
- routing table in the injected leo:using-leo policy); this skill covers what
13
- to do once the tier is already chosen.
14
- ---
15
-
16
- # delegation
17
-
18
- Core rule: a subagent gets one shot at the brief and no session history. If
19
- the brief doesn't stand alone, the dispatch is already broken.
20
-
21
- ## Writing the brief
22
-
23
- Every dispatch is self-contained: goal, constraints, exact file paths, the
24
- checks to run, and what the return must contain. Write it as if for a
25
- stranger who will never see this conversation — because that's what a
26
- subagent is. A brief missing a file path or a check produces a report that
27
- looks done and isn't.
28
-
29
- Bad: "fix the flaky auth test." Good: "`tests/auth/session_test.py::test_expiry`
30
- fails intermittently (repro: run it 20x, ~1 in 8 fails). Fix the race, keep
31
- the test's intent unchanged, don't touch other tests in the file. Run
32
- `pytest tests/auth/session_test.py -x` 20 times clean before reporting done.
33
- Return: files touched, the race you found, the check output." The second
34
- version needs no follow-up question; the first invites three.
35
-
36
- ## Pin model and effort
37
-
38
- Every dispatch pins **model AND effort** from the routing table — opus for
39
- judges (reviewer, investigator), sonnet for execution (implementer, executor
40
- on normal work), haiku for mechanical work (executor on boilerplate). expert
41
- never appears in a fan-out — one at a time, never fanned. An unpinned call
42
- silently inherits the session's tier: in an opus session that means every
43
- executor spawn quietly runs at opus, and a ten-item fan-out burns
44
- opus-fan-out money for haiku-shaped work. Pin both fields on every spawn, not
45
- just the ones that "obviously" need it.
46
-
47
- ## The four-state return contract
48
-
49
- A subagent's report must resolve to exactly one of four states. Don't accept
50
- a report that hedges across two of them.
51
-
52
- | State | Means | Your response |
53
- |---|---|---|
54
- | `done` | Work finished, matches the brief | Verify against artifacts — see leo:verification — never take the self-report at face value |
55
- | `concerns` | Finished, but flags something worth a second look | Read the concerns before accepting; they're often the real finding |
56
- | `needs-context` | Blocked on missing information you can supply | Send the missing piece to the same agent (`SendMessage` on Claude Code — elsewhere see the *Follow-up to a live agent* row of your mapping, and where none is established, cold re-dispatch with the context restated is the whole mechanism) so it keeps the context it already built. Either way **once** — a second needs-context on the same gap means the brief itself is broken, escalate the tier |
57
- | `blocked` | Blocked on something you can't hand over inline | Resolve the blocker, or escalate per the ladder — never a silent same-tier retry |
58
-
59
- `needs-context` and `blocked` look similar; the test is whether the missing
60
- piece is something *you* hold (needs-context — a file path, a decision, a
61
- credential) or something neither of you can supply without more work
62
- (blocked — a failing external service, a genuinely ambiguous requirement).
63
-
64
- Each role's own prompt carries the state line it must emit, so the contract
65
- is enforced at both ends. Two roles are deliberately narrowed: `reviewer`
66
- emits only `done` / `needs-context` (severity already lives in
67
- `blocking`/`non-blocking`, and the diff's own verdict in
68
- `approved`/`needs-changes`), and `expert` never emits `blocked` — it is the
69
- ceiling, so there is nothing left to escalate to. `status` is a separate axis
70
- from `confidence`: status routes your next move, confidence rates the work.
71
-
72
- ## Long multi-agent runs: the ledger
73
-
74
- A run spanning many dispatches survives context compaction only if progress
75
- is persisted outside the conversation. Use
76
- `${CLAUDE_PLUGIN_ROOT}/scripts/state.py` (get / merge / path — flock-guarded,
77
- atomic writes, keyed per repo) as the ledger, not ad hoc notes in the
78
- transcript. Each entry: item id, status (one of the four states above, plus
79
- `pending` / `in-progress`), artifact path (branch name, file, or diff). On
80
- resume, read the ledger first — anything already `done` or `concerns` is not
81
- re-dispatched; anything `blocked` is reported, not silently retried.
82
-
83
- A ledger entry is small — `{"items": {"<id>": {"status": "done", "artifact":
84
- "branch:fix/eng-123-slug"}}}` merged via `python3 "${CLAUDE_PLUGIN_ROOT}/scripts/state.py"
85
- merge <skill-name> <owner/repo> '<patch>'` — but it's the only thing standing between a
86
- compaction mid-run and forty items silently re-dispatched from item 1.
87
- Update it after every dispatch resolves, not in a batch at the end: a crash
88
- between "agent finished" and "ledger written" is exactly the gap this
89
- exists to close.
90
-
91
- `${CLAUDE_PLUGIN_ROOT}` above is the Claude Code spelling of the plugin root,
92
- and it is substituted into this text only there. On another harness, read the
93
- plugin-root form from that harness's appendix in the injected policy (Codex
94
- uses a real `$PLUGIN_ROOT` env var, Cursor `$CURSOR_PLUGIN_ROOT`) — the path
95
- after the root is identical everywhere.
96
-
97
- For a batch of independent, well-scoped fixes, don't hand-roll this loop —
98
- the reusable workflow at `${CLAUDE_PLUGIN_ROOT}/workflows/cost-tiered-fix.js`
99
- (Workflow tool, `scriptPath`) already implements plan → tiered execute →
100
- opus verify with escalation built in, including its own progress tracking.
101
- Reach for it before writing a bespoke fan-out loop; write the ledger
102
- approach above only when the run doesn't fit that workflow's shape (e.g. one
103
- dispatch at a time inside a larger interactive flow, not a clean batch).
104
- That workflow needs Claude Code's Workflow tool; on every other harness the
105
- ledger above is the whole mechanism, so use it directly rather than looking
106
- for a runner that isn't there.
107
-
108
- ## Parallel dispatch: own your files
109
-
110
- Fan-out is safe only when each spawn writes to **disjoint** files — no two
111
- concurrent dispatches touching the same path. If the work can't be split
112
- into disjoint file sets (one coherent change that happens to span many
113
- files, like a single ticket fix), don't fan out — either run it sequentially
114
- in one dispatch, or give each spawn its own isolated tree via leo:worktrees
115
- so parallel edits can't collide even when the file sets overlap.
116
-
117
- ## Self-talk to catch
118
-
119
- - "I'll skip pinning effort, model is enough" — no; an unpinned effort on an
120
- opus judge still runs at opus prices, at auto effort, which is not what
121
- the routing table costed out.
122
- - "The brief is short, they'll infer the rest" — a subagent infers nothing;
123
- it has this brief and nothing else.
124
- - "It said needs-context, I'll just re-ask the same way" — re-dispatching
125
- with the identical brief reproduces the identical gap; either add the
126
- missing piece or step up a tier. And prefer messaging the same agent over
127
- a fresh spawn: a cold re-dispatch pays again for the context it already
128
- built and can rediscover the same gap from a different angle.
129
- - "Two spawns editing the same file will probably be fine, they touch
130
- different functions" — same file is not disjoint; sequence them or
131
- isolate with a worktree.
132
- - "This ten-item fan-out is basically cost-tiered-fix, I'll just write the
133
- loop myself" — the workflow already handles escalation and orphan
134
- tracking; reinventing it inline drops that for no reason.
135
-
136
- ## Works with
137
-
138
- - leo:verification — how a `done` report gets checked against real
139
- artifacts, not trusted as stated.
140
- - leo:worktrees — file isolation for parallel dispatches that can't be made
141
- disjoint by scope alone.
@@ -1,116 +0,0 @@
1
- ---
2
- name: executing-plans
3
- description: >
4
- Checkpoint discipline for carrying out a written plan — batch execution
5
- with a check at every batch boundary, plan-intent-wins-on-architecture /
6
- reality-wins-on-mechanics arbitration, and one fix-then-re-review cycle
7
- before stopping to report. Used by the implementer agent, or the main loop
8
- when it executes a plan directly.
9
- when_to_use: >
10
- A written plan (from planner, an issue, or Leo's own outline) is about to
11
- be turned into code. NOT for open-ended implementation with no plan
12
- (normal execute-then-review flow) and NOT for the review step itself
13
- (the reviewer agent judges the diff; this skill only carries out the plan).
14
- ---
15
-
16
- # executing-plans
17
-
18
- Core rule: a plan is executed in checkpointed batches, never as one long
19
- uninterrupted run. Each checkpoint is a place execution is allowed to stop
20
- without having made things worse.
21
-
22
- ## Before edit one
23
-
24
- Sanity-check the plan against the tree it's about to touch:
25
-
26
- - Base ref matches what the plan assumed — `git rev-parse HEAD` against the
27
- base the plan was written against. Drifted → say so before touching
28
- anything; the plan may already be stale.
29
- - Files/symbols the plan names actually exist at the paths/shapes it
30
- describes. A plan step that references a function that moved or a file
31
- that's gone is a stop-and-report, not a guess-and-proceed.
32
-
33
- This is cheap — a few Read/Grep calls — and skipping it is how a plan
34
- written against yesterday's tree silently corrupts today's.
35
-
36
- ## Execute in batches
37
-
38
- Break the plan into batches along its own natural seams (usually: one
39
- plan-step or one cohesive file group per batch). At each batch boundary:
40
-
41
- 1. Finish the batch's edits.
42
- 2. Run the narrowest relevant checks for what that batch touched — the
43
- touched test file, a targeted typecheck, not the full suite every time.
44
- 3. Green → advance to the next batch. Red → stop the batch right there; fix
45
- it or report it. Never carry a red check into the next batch hoping it
46
- resolves itself — a checkpoint exists precisely to catch this before the
47
- failure compounds across three more batches of edits built on top of it.
48
-
49
- This is the same shape as leo:delegation's tiering: cheap, frequent checks
50
- bound the blast radius so the expensive step (review) isn't debugging a
51
- pile of unrelated regressions.
52
-
53
- ## Plan intent wins on architecture; reality wins on mechanical detail
54
-
55
- Two different kinds of mismatch between plan and tree call for two different
56
- responses:
57
-
58
- - **Mechanical drift** (a renamed variable, a moved file, a slightly
59
- different function signature than the plan assumed) — reality wins. Adapt
60
- the mechanics silently and keep going; that's normal execution, not a
61
- deviation worth flagging.
62
- - **Architectural disagreement** (the plan's approach doesn't fit the actual
63
- structure, a step contradicts how the system actually works, following it
64
- as written would build on a wrong premise) — the plan's intent still wins
65
- over improvising a fix, but only the plan's author can resolve a real
66
- conflict. Stop and report the disagreement; never silently redesign around
67
- it. Silent redesign is worse than executing a flawed plan, because it
68
- hides the disagreement instead of surfacing it.
69
-
70
- When genuinely unsure which kind of mismatch it is, treat it as
71
- architectural and stop — reporting an unnecessary pause costs a message;
72
- silently redesigning costs trust.
73
-
74
- ## Behavior changes still default to test-first, done still means verification
75
-
76
- A plan step that changes behavior doesn't get a pass on process because it's
77
- already written down. Default to leo:test-first for those steps, and treat
78
- "the plan is implemented" and "the plan is done" as different states — done
79
- still means the change clears leo:verification, not just that every step got
80
- executed.
81
-
82
- ## One fix-then-re-review cycle
83
-
84
- Once all batches are in, this hands off to the standard review gate — spawn
85
- `reviewer` on the actual diff against the recorded base ref,
86
- with the plan text as the original request. If it comes back with blocking
87
- findings: fix at the executing tier, then re-review only the fix. That's
88
- **one fix-then-re-review cycle**, full stop. A second block on the same
89
- findings means stop the loop and report to Leo with options, expert
90
- arbitration (the `expert` agent) among them — never a third pass, never quietly
91
- loosening what counts as blocking to escape the loop.
92
-
93
- ## Delegation and workspace boundaries
94
-
95
- Executing a written plan is `implementer`'s job per leo:delegation — the
96
- main loop only executes inline when it's already the implementer context or
97
- the touch is genuinely trivial. If the plan spans a branch of nontrivial
98
- size, it runs on a dedicated branch per leo:worktrees, and finishing it
99
- follows leo:finishing-a-branch rather than improvising a merge/cleanup
100
- sequence at the end.
101
-
102
- ## Self-talk to catch
103
-
104
- - "The plan says step 4, I'll just push through to step 7 before checking
105
- anything" — that's skipping checkpoints, not saving time; a break at step
106
- 5 now costs one batch's rework instead of three.
107
- - "This isn't quite what the plan says but it's obviously what they meant" —
108
- if it's mechanical, fine; if it's architectural, that's the silent
109
- redesign this skill exists to block. Report it instead.
110
- - "The re-review still isn't clean but it's close enough" — close enough on
111
- a second block is the definition of stop-and-report, not a third fix.
112
-
113
- ## Works with
114
-
115
- leo:test-first, leo:verification, leo:delegation, leo:worktrees,
116
- leo:finishing-a-branch — plus the `reviewer` and `expert` agents.
@@ -1,123 +0,0 @@
1
- ---
2
- name: finishing-a-branch
3
- description: >
4
- End-of-branch state machine: what happens once implementation on a
5
- branch/worktree is complete. Gates on a clean review verdict, then offers
6
- a closed set of next steps — merge / PR / keep / discard — routes the
7
- chosen path through the right ordering (land the work before removing the
8
- worktree, remove the worktree before deleting the branch), and leaves the
9
- repo clean.
10
- when_to_use: >
11
- A branch or worktree has reached "implementation done" and Leo needs to
12
- decide what happens to it. Fires after execute-then-review completes, or
13
- when Leo says finish/wrap up/close out/clean up this branch. NOT for
14
- starting or managing a worktree mid-task (that's leo:worktrees) and NOT a
15
- substitute for the review cycle itself (that's execute-then-review) — this
16
- skill starts only once a review verdict already exists.
17
- ---
18
-
19
- # finishing-a-branch
20
-
21
- Core rule: a branch doesn't get disposed of by momentum. It reaches one of
22
- four terminal states, each chosen explicitly, and destructive ones require
23
- saying out loud what gets lost.
24
-
25
- ## Precondition: review verdict, not vibes
26
-
27
- Do not enter this skill's decision step without a clean **review verdict**
28
- on the final diff. "Implementation looks done" is not a review verdict.
29
-
30
- - If review hasn't run yet, or the last verdict was `needs-changes`: stop
31
- here, go run/finish the review cycle (see execute-then-review), come back.
32
- - If review is `approved`: proceed.
33
- - Never offer merge/PR on unreviewed or still-blocked work. "It's a small
34
- change" or "I already read through it" does not substitute for the
35
- reviewer's verdict — those are exactly the rationalizations this gate
36
- exists to block.
37
-
38
- ## The option set is closed
39
-
40
- Once the gate passes, present exactly these four options — never an
41
- open-ended "what would you like to do next?":
42
-
43
- - **merge** — into the target branch, locally or via `gh pr merge`
44
- - **PR** — open a pull request and stop (no local merge)
45
- - **keep** — leave the branch/worktree exactly as-is, decide later
46
- - **discard** — delete the branch and its worktree, work is gone
47
-
48
- State the branch name, commit count ahead of the target, and the review
49
- verdict when you present the set. Leo picks one; do not infer a choice from
50
- silence, from a prior unrelated "yes," or from tone.
51
-
52
- ## Ordering (prevents self-referential failures)
53
-
54
- Regardless of which path Leo picks, sequence matters — doing this out of
55
- order breaks the tools that need the worktree or branch to still exist:
56
-
57
- 1. **cd out of the worktree first.** A shell sitting inside the worktree
58
- directory blocks its own removal.
59
- 2. **Merge (or push for a PR) BEFORE removing the worktree.** Land or
60
- publish the commits while the worktree still exists to operate from.
61
- 3. **Remove the worktree BEFORE deleting the branch.** Deleting the branch
62
- out from under a live worktree leaves the worktree metadata dangling and
63
- git in an inconsistent state.
64
- 4. Mechanics of steps 1–3 (which git worktree commands, how to prune) are
65
- owned by `leo:worktrees` — call into it rather than hand-rolling worktree
66
- surgery here. This skill decides *what* happens and *in what order*;
67
- `leo:worktrees` executes *how*.
68
-
69
- Per option:
70
-
71
- | Option | Sequence |
72
- |---|---|
73
- | merge | merge locally or `gh pr merge` → remove worktree (`leo:worktrees`) → delete local branch |
74
- | PR | push branch → open PR → **stop** (worktree and branch stay; nothing is unmerged yet) |
75
- | keep | do nothing destructive; leave worktree and branch as-is |
76
- | discard | typed confirmation (below) → remove worktree (`leo:worktrees`) → force-delete branch |
77
-
78
- ## Destructive paths require a typed confirmation
79
-
80
- `discard`, and any force-delete of a branch with unmerged commits, requires
81
- Leo to type back a confirmation that **names exactly what will be lost** —
82
- not a plain "yes" or "go ahead". Prompt with the specific string, e.g.:
83
-
84
- > Type `discard` to delete branch `feature/foo`, 4 commits, no PR — this
85
- > cannot be undone.
86
-
87
- - An implied or inferred yes never triggers deletion — silence, "sounds
88
- good," or approval of some *other* step in the conversation does not
89
- count.
90
- - If Leo's typed text doesn't match what was asked for, ask again; don't
91
- guess at intent.
92
- - `keep` never needs this — it's non-destructive by construction.
93
- - If the branch is already merged, force-delete is not "destructive" in the
94
- data-loss sense (git still warns) — a plain confirmation is enough since
95
- nothing unmerged is at risk; use the typed-confirmation form when in doubt.
96
-
97
- ## Leave the repo clean
98
-
99
- After any path except `keep`:
100
-
101
- - Prune worktree metadata (`leo:worktrees` handles this as part of removal
102
- — don't leave a stale entry in `git worktree list`).
103
- - Confirm `git status` is clean from the directory you're now in.
104
- - Note the outcome (merged / PR opened + link / kept / discarded) in the
105
- done report per `leo:verification` — the report's job is to make the
106
- terminal state legible later, not just at the moment it happened.
107
-
108
- ## Self-talk to catch
109
-
110
- - "The diff was tiny, I basically reviewed it while writing it" — that's
111
- not a review verdict; go get one.
112
- - "Leo said 'sounds good' earlier, close it out" — sounds-good is not a
113
- typed confirmation naming what's lost.
114
- - "I'll just clean up the worktree now and merge after" — wrong order,
115
- breaks the merge step; land first.
116
- - "Discard is obviously right here, I'll skip the prompt to save a round
117
- trip" — the option set is closed and explicit for a reason; present it.
118
-
119
- ## Works with
120
-
121
- - `leo:worktrees` — owns worktree creation/removal mechanics.
122
- - `leo:verification` — owns the shape of the done report this skill feeds
123
- its outcome line into.