leos-agent 6.1.1 → 7.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (61) hide show
  1. package/LICENSE +21 -0
  2. package/README.md +18 -11
  3. package/adapters/cursor/agents/executor.md +2 -2
  4. package/adapters/cursor/agents/implementer.md +2 -2
  5. package/adapters/cursor/agents/review-lens.md +22 -0
  6. package/adapters/cursor/agents/reviewer.md +3 -2
  7. package/adapters/opencode/agents.json +44 -5
  8. package/adapters/opencode/plugin.js +435 -47
  9. package/config/MCP_PINS.md +17 -0
  10. package/config/models.json +647 -33
  11. package/hooks/bash-guard.py +51 -9
  12. package/hooks/session-start.py +27 -0
  13. package/package.json +19 -8
  14. package/roles/executor.md +2 -2
  15. package/roles/implementer.md +2 -2
  16. package/roles/review-lens.md +20 -0
  17. package/roles/reviewer.md +3 -2
  18. package/scripts/doctor.py +520 -0
  19. package/scripts/ghreview.py +558 -0
  20. package/scripts/jsonc_bridge.cjs +23 -0
  21. package/scripts/memory.py +744 -0
  22. package/scripts/render_adapters.py +253 -121
  23. package/scripts/resolve_attach_target.py +389 -0
  24. package/scripts/setup.py +1753 -0
  25. package/skills/brainstorming/SKILL.md +3 -1
  26. package/skills/debugging/SKILL.md +4 -2
  27. package/skills/delegation/SKILL.md +10 -8
  28. package/skills/doctor/SKILL.md +124 -0
  29. package/skills/executing-plans/SKILL.md +2 -1
  30. package/skills/finishing-a-branch/SKILL.md +4 -2
  31. package/skills/freshness/SKILL.md +131 -0
  32. package/skills/memory/SKILL.md +154 -0
  33. package/skills/resolve-ticket/SKILL.md +275 -0
  34. package/skills/review-pr/SKILL.md +327 -0
  35. package/skills/setup/SKILL.md +199 -0
  36. package/skills/setup/agents/openai.yaml +5 -0
  37. package/skills/test-first/SKILL.md +3 -1
  38. package/skills/using-leo/SKILL.md +18 -6
  39. package/skills/using-leo/references/claude-mapping.md +23 -1
  40. package/skills/using-leo/references/codex-mapping.md +18 -9
  41. package/skills/using-leo/references/cursor-mapping.md +19 -6
  42. package/skills/using-leo/references/hermes-mapping.md +18 -7
  43. package/skills/using-leo/references/opencode-mapping.md +18 -9
  44. package/skills/verification/SKILL.md +9 -1
  45. package/skills/visual-verification/SKILL.md +115 -0
  46. package/skills/watch-review/SKILL.md +128 -0
  47. package/skills/watch-review/agents/openai.yaml +5 -0
  48. package/skills/worktrees/SKILL.md +3 -1
  49. package/skills/writing-plans/SKILL.md +2 -1
  50. package/skills/writing-skills/SKILL.md +141 -0
  51. package/vendor/jsonc-parser-3.3.1/LICENSE.md +21 -0
  52. package/vendor/jsonc-parser-3.3.1/README.md +26 -0
  53. package/vendor/jsonc-parser-3.3.1/lib/umd/impl/edit.js +201 -0
  54. package/vendor/jsonc-parser-3.3.1/lib/umd/impl/format.js +275 -0
  55. package/vendor/jsonc-parser-3.3.1/lib/umd/impl/parser.js +682 -0
  56. package/vendor/jsonc-parser-3.3.1/lib/umd/impl/scanner.js +456 -0
  57. package/vendor/jsonc-parser-3.3.1/lib/umd/impl/string-intern.js +42 -0
  58. package/vendor/jsonc-parser-3.3.1/lib/umd/main.d.ts +351 -0
  59. package/vendor/jsonc-parser-3.3.1/lib/umd/main.js +194 -0
  60. package/vendor/jsonc-parser-3.3.1/package.json +37 -0
  61. package/workflows/cost-tiered-fix.js +32 -4
@@ -7,7 +7,9 @@ description: >
7
7
  wide blast radius, hard to reverse, or that introduce new surface need
8
8
  genuine, viable alternatives with trade-offs weighed before any code gets
9
9
  written. Produces the chosen approach and its trade-offs, sized to the gate,
10
- handed off to leo:writing-plans.
10
+ handed off to leo:writing-plans. Use when choosing an approach before
11
+ non-trivial code. Do not use for contained reversible tweaks, investigation,
12
+ or writing the plan itself.
11
13
  when_to_use: >
12
14
  Before starting non-trivial code: a new feature, a new integration surface,
13
15
  a schema or data-model change, anything that's expensive or awkward to
@@ -5,7 +5,9 @@ description: >
5
5
  behavior. Five named phases — Reproduce, Localize, Hypothesize, Prove, Fix —
6
6
  each with an exit criterion, so a fix never lands before the cause is
7
7
  pinned to file:line. Diagnosis is read-only judge work (investigator); the
8
- fix happens separately, at the routed tier.
8
+ fix happens separately, at the routed tier. Use when a bug, failure, crash,
9
+ or surprising behavior needs diagnosis. Do not use for planned features,
10
+ reviewing a diff, or post-fix completion verification.
9
11
  when_to_use: >
10
12
  Any bug report, failing test, crash, stack trace, or "why does X happen"
11
13
  before proposing a fix — used by the investigator agent and by the main
@@ -61,7 +63,7 @@ you, and go again.
61
63
 
62
64
  After two failures on the same cause (two hypotheses tried and reverted, still
63
65
  no Prove), step up one tier rather than retrying at the same one — investigator
64
- haiku-assist steps to full investigator, investigator itself steps to a
66
+ explore findings feed a full investigator, and investigator itself steps to a
65
67
  second, more evidence-fed pass, capped at Opus. A genuine deadlock, or two
66
68
  Opus verdicts on the same cause that disagree → `expert`, announced in one
67
69
  line ("escalating to expert: <question>") before it's invoked, never silent.
@@ -4,7 +4,8 @@ description: >
4
4
  Operational mechanics for dispatching subagents — a single spawn or a large
5
5
  fan-out — the companion to the policy's "Delegate the labor" section.
6
6
  Covers brief construction, model/effort pinning, the four-state return
7
- contract, and ledger-backed progress tracking for long multi-agent runs.
7
+ contract, and ledger-backed progress tracking for long multi-agent runs. Use
8
+ when dispatching any subagent or fan-out. Do not use to choose a task's tier.
8
9
  when_to_use: >
9
10
  Any time work is routed to a subagent (explore, investigator, executor,
10
11
  implementer, reviewer, expert) rather than done inline — single dispatch or
@@ -36,8 +37,8 @@ version needs no follow-up question; the first invites three.
36
37
  ## Pin model and effort
37
38
 
38
39
  Every dispatch pins **model AND effort** from the routing table — opus for
39
- judges (reviewer, investigator), sonnet for execution (implementer, executor
40
- on normal work), haiku for mechanical work (executor on boilerplate). expert
40
+ judges (reviewer, investigator), sonnet for normal implementation
41
+ (implementer), haiku for mechanical work (executor). expert
41
42
  never appears in a fan-out — one at a time, never fanned. An unpinned call
42
43
  silently inherits the session's tier: in an opus session that means every
43
44
  executor spawn quietly runs at opus, and a ten-item fan-out burns
@@ -53,7 +54,7 @@ a report that hedges across two of them.
53
54
  |---|---|---|
54
55
  | `done` | Work finished, matches the brief | Verify against artifacts — see leo:verification — never take the self-report at face value |
55
56
  | `concerns` | Finished, but flags something worth a second look | Read the concerns before accepting; they're often the real finding |
56
- | `needs-context` | Blocked on missing information you can supply | Send the missing piece to the same agent (SendMessage) so it keeps the context it already built; cold re-dispatch only if that agent is gone. Either way **once** — a second needs-context on the same gap means the brief itself is broken, escalate the tier |
57
+ | `needs-context` | Blocked on missing information you can supply | Send the missing piece to the same agent (`SendMessage` on Claude Code, `followup_task` on Codex — elsewhere see the *Follow-up to a live agent* row of your mapping, and where none is established, cold re-dispatch with the context restated is the whole mechanism) so it keeps the context it already built. Either way **once** — a second needs-context on the same gap means the brief itself is broken, escalate the tier |
57
58
  | `blocked` | Blocked on something you can't hand over inline | Resolve the blocker, or escalate per the ladder — never a silent same-tier retry |
58
59
 
59
60
  `needs-context` and `blocked` look similar; the test is whether the missing
@@ -89,10 +90,11 @@ between "agent finished" and "ledger written" is exactly the gap this
89
90
  exists to close.
90
91
 
91
92
  `${CLAUDE_PLUGIN_ROOT}` above is the Claude Code spelling of the plugin root,
92
- and it is substituted into this text only there. On another harness, read the
93
- plugin-root form from that harness's appendix in the injected policy (Codex
94
- uses a real `$PLUGIN_ROOT` env var, Cursor `$CURSOR_PLUGIN_ROOT`) — the path
95
- after the root is identical everywhere.
93
+ and it is substituted into this text only there. Codex exposes `$PLUGIN_ROOT`
94
+ and Cursor `$CURSOR_PLUGIN_ROOT`. Hermes and OpenCode expose no root variable;
95
+ their injected policy substitutes an absolute payload path into the `state.py`
96
+ and `memory.py` commands, which is the discoverable source to reuse. Do not
97
+ invent an environment variable where the harness exposes none.
96
98
 
97
99
  For a batch of independent, well-scoped fixes, don't hand-roll this loop —
98
100
  the reusable workflow at `${CLAUDE_PLUGIN_ROOT}/workflows/cost-tiered-fix.js`
@@ -0,0 +1,124 @@
1
+ ---
2
+ name: doctor
3
+ description: >
4
+ Self-check for Leo's own wiring. Reports which harness this is, what each
5
+ tier name resolves to here, whether the bootstrap is installed, where
6
+ machine-local state and the memory store live, and which skills shipped
7
+ versus which this session can actually invoke. Disk facts come from a
8
+ helper script; the context facts only the running session can answer, and
9
+ a disagreement between the two columns is the diagnosis. Use when Leo asks
10
+ about Leo's loading, routing, or skill wiring. Do not use for project health
11
+ checks, project-code debugging, or unprompted inspection.
12
+ when_to_use: >
13
+ Leo asks whether the policy loaded, why routing or a skill is misbehaving,
14
+ or invokes doctor by name after installing, updating, or switching harness.
15
+ Also the first move when a leo skill cannot be found. NOT a general
16
+ environment or project health check, NOT for debugging the project's own
17
+ code (that is leo:debugging), and never run unprompted — it reports on the
18
+ agent, not on the work.
19
+ ---
20
+
21
+ # doctor
22
+
23
+ Doctor answers two questions that look like one: what shipped to disk, and what
24
+ reached this session. A skill the harness never registered is indistinguishable
25
+ from a skill that does not exist, right up until the moment you invoke it.
26
+
27
+ ## Run the script
28
+
29
+ ```sh
30
+ python3 "${CLAUDE_PLUGIN_ROOT}/scripts/doctor.py" --harness <name>
31
+ ```
32
+
33
+ Pass `--harness` with the harness you are on — the mapping appendix in your
34
+ context names it in its own heading (`# Hermes mapping` → `hermes`). Detection
35
+ without it relies on a plugin-root variable that Hermes and OpenCode do not
36
+ export, so on those two the script reports `unknown` rather than guessing.
37
+ `unknown` on a harness whose mapping you can plainly read is a missing
38
+ argument, not a fault.
39
+
40
+ `${CLAUDE_PLUGIN_ROOT}` is the Claude Code spelling. Codex exports
41
+ `$PLUGIN_ROOT` and Cursor `$CURSOR_PLUGIN_ROOT`. On Hermes and OpenCode no
42
+ plugin-root variable exists at all — the injected policy instead substitutes an
43
+ absolute payload path into its `state.py` and `memory.py` commands. Read that
44
+ command path from the policy as the discoverable source. Being unable to locate
45
+ the payload at all is itself the first finding: the harness is not looking where
46
+ the plugin was installed.
47
+
48
+ Add `--json` when you want the same facts as data.
49
+
50
+ Doctor validates the bootstrap that actually belongs to the named harness:
51
+ the session hook and manifest for Claude, Codex, and Cursor;
52
+ `config.instructions` in OpenCode's plugin; and Hermes registration plus its
53
+ first-tool-result fallback. It also reports the running Python version against
54
+ the supported 3.9+ floor. Codex hook trust is not provable from disk: review
55
+ the plugin in `/hooks` and confirm it is trusted before treating the on-disk
56
+ hook as active.
57
+
58
+ ## Then answer the three it cannot
59
+
60
+ A script can prove the hook is installed and that the policy renders. It cannot
61
+ prove the policy arrived. Only you can see your own context.
62
+
63
+ 1. **Did the policy load?** Look for the policy wrapper in your context, and
64
+ check that the mapping following it names *this* harness. A policy present
65
+ but carrying another harness's mapping is worse than none, because routing
66
+ then points at models that do not exist here.
67
+ 2. **Which skills are actually invocable?** Compare your own skill list against
68
+ the script's shipped roster. Mind the naming rule: most harnesses namespace
69
+ them as `leo:<name>`, while OpenCode has no namespace and requires a
70
+ skill's frontmatter name to match its directory, so the plugin registers a
71
+ generated shadow copy with every skill renamed `leo-<name>`. A skill that
72
+ looks missing on OpenCode may simply be listed as `leo-<name>` rather than
73
+ `leo:<name>`.
74
+ 3. **Is memory present and delivered?** The script reports whether the store
75
+ exists and whether each native surface received its generated copy. Whether
76
+ those facts are in front of you right now is something only you can confirm.
77
+ Report the two separately; they disagree more often than expected. Hermes's
78
+ projection is opt-in, so doctor reports it explicitly as disabled rather
79
+ than silently omitting it.
80
+
81
+ ## Reading the report
82
+
83
+ Every row carries its source — `env`, `disk`, `config`, or `context` — so a
84
+ reader can tell a fact from an inference. Close with one verdict from exactly
85
+ three: **healthy**, **degraded**, or **not loaded**. Never free prose. `not
86
+ loaded` outranks everything else: if the policy did not arrive, nothing else in
87
+ the report describes how this session will actually behave.
88
+
89
+ **Most breadcrumb logs are history, not a verdict.** Some older logs carry no
90
+ timestamps, and the test suite drives failure paths deliberately, so entries
91
+ can accumulate on a development machine. Quote the newest line if useful, but
92
+ never conclude "the hook failed this session" from history alone. The one
93
+ capability exception is `opencode-skills.log`: its presence means namespaced
94
+ OpenCode skill registration degraded and doctor reports that state until the
95
+ breadcrumb is cleared after the underlying problem is understood.
96
+
97
+ ## Failure modes
98
+
99
+ | Symptom | Likely cause | Fix |
100
+ |---|---|---|
101
+ | Policy absent, bootstrap installed | the hook fired and failed open | read the newest breadcrumb, then confirm it describes this session before believing it |
102
+ | Policy present, mapping names another harness | detection resolved wrong, usually a stray plugin-root variable exported in an unrelated shell | unset it, restart the session |
103
+ | Harness reported as `unknown` | no `--harness`, and this harness exports no plugin-root variable | degraded until re-run with `--harness <name>` read off your mapping heading; a still-unknown explicit run is invalid wiring |
104
+ | Codex hook is on disk but policy is absent | the new or changed hook may not be trusted | open `/hooks`, review the hook, and explicitly trust it |
105
+ | OpenCode reports `opencode-skills.log` | the namespaced shadow tree failed and no bare-name fallback was registered | inspect the newest breadcrumb, fix the path/permission failure, and restart OpenCode |
106
+ | Shipped roster exceeds what you can invoke | the harness cached an older payload, or the skills directory is not registered | update the plugin; on OpenCode check `opencode debug skill` for each skill's `location` |
107
+ | Skills listed as `leo-<name>` instead of `leo:<name>` | OpenCode, working as designed | invoke them as `leo-<name>`; not a fault |
108
+ | Tier names resolve to models this harness cannot run | mapping and harness disagree | same as row 2 |
109
+ | Machine-local state not writable | the path override points somewhere unwritable | fix or unset it |
110
+ | A skill is genuinely absent from disk | it was never added | see leo:writing-skills |
111
+
112
+ ## Doctor never repairs
113
+
114
+ It reports, and it names the fix. It does not reinstall, rewrite configuration,
115
+ or delete state — which is what keeps it safe to run at any tier and at any
116
+ moment.
117
+
118
+ ## Works with
119
+
120
+ - leo:writing-skills — for a skill that turned out to be missing because nobody
121
+ wrote it yet.
122
+ - leo:memory — doctor reports whether the store exists and reached each surface.
123
+ - leo:verification — this report is a claim like any other: the script ran this
124
+ turn and its output was read.
@@ -5,7 +5,8 @@ description: >
5
5
  with a check at every batch boundary, plan-intent-wins-on-architecture /
6
6
  reality-wins-on-mechanics arbitration, and one fix-then-re-review cycle
7
7
  before stopping to report. Used by the implementer agent, or the main loop
8
- when it executes a plan directly.
8
+ when it executes a plan directly. Use when a written plan is about to become
9
+ code. Do not use for open-ended work without a plan or for reviewing a diff.
9
10
  when_to_use: >
10
11
  A written plan (from planner, an issue, or Leo's own outline) is about to
11
12
  be turned into code. NOT for open-ended implementation with no plan
@@ -6,7 +6,9 @@ description: >
6
6
  a closed set of next steps — merge / PR / keep / discard — routes the
7
7
  chosen path through the right ordering (land the work before removing the
8
8
  worktree, remove the worktree before deleting the branch), and leaves the
9
- repo clean.
9
+ repo clean. Use when reviewed implementation on a branch/worktree needs a
10
+ terminal disposition. Do not use to manage a worktree mid-task or replace
11
+ the review cycle.
10
12
  when_to_use: >
11
13
  A branch or worktree has reached "implementation done" and Leo needs to
12
14
  decide what happens to it. Fires after execute-then-review completes, or
@@ -71,7 +73,7 @@ Per option:
71
73
  | Option | Sequence |
72
74
  |---|---|
73
75
  | merge | merge locally or `gh pr merge` → remove worktree (`leo:worktrees`) → delete local branch |
74
- | PR | push branch → open PR → **stop** (worktree and branch stay; nothing is unmerged yet) |
76
+ | PR | push branch → open PR → **stop** (worktree and branch stay; nothing is merged yet.) |
75
77
  | keep | do nothing destructive; leave worktree and branch as-is |
76
78
  | discard | typed confirmation (below) → remove worktree (`leo:worktrees`) → force-delete branch |
77
79
 
@@ -0,0 +1,131 @@
1
+ ---
2
+ name: freshness
3
+ description: >
4
+ Currency gate for code written against anything outside this repository.
5
+ Before a library call, CLI flag, endpoint field, or vendor number is
6
+ committed to, its shape is confirmed against a source that reflects the
7
+ version this project actually runs — the installed package, the lockfile
8
+ pin, or documentation fetched this turn. Each check is recorded by symbol
9
+ and source in the report. Use when writing, reviewing, or asserting a
10
+ third-party surface. Do not use for first-party code, pinned standard
11
+ libraries, or as a substitute for verification.
12
+ when_to_use: >
13
+ About to write, review, or assert the shape of a third-party surface — an
14
+ import path, an argument list, a config key, an HTTP field, an auth
15
+ scheme, a model id, a price, a deprecation claim. NOT for first-party code
16
+ in this workspace (read it instead), NOT for the standard library of a
17
+ pinned runtime, and NOT a substitute for running anything — leo:verification
18
+ still governs the completion claim built on top of it.
19
+ ---
20
+
21
+ # freshness
22
+
23
+ A third-party surface you have not read this session is a guess, however
24
+ familiar it feels. Recall of a library is a snapshot of some arbitrary past
25
+ version; it is not a snapshot of the one pinned in this lockfile. The cost is
26
+ a call that reads perfectly and does not exist.
27
+
28
+ Recall is not a source. The package installed on disk is.
29
+
30
+ ## When it fires
31
+
32
+ A closed list of five.
33
+
34
+ 1. **A symbol you did not read this session** — a function, method, class,
35
+ decorator, flag, or config key belonging to something not defined in this
36
+ working tree.
37
+ 2. **A version-sensitive call shape** — argument order, keyword names, return
38
+ type, or import path for a dependency whose installed version you have not
39
+ confirmed.
40
+ 3. **A service contract** — endpoint path, request or response field, auth
41
+ scheme, pagination rule, error code.
42
+ 4. **A vendor-schedule fact** — a model id, context window, price, rate limit,
43
+ or regional availability. These move on someone else's calendar.
44
+ 5. **A deprecation or removal claim** — "that was dropped in v3" is an
45
+ assertion about a moving target and needs the same check as a signature.
46
+
47
+ Outside these five, write the code.
48
+
49
+ ## What counts as a source
50
+
51
+ Two different questions — which to reach for, and which one wins.
52
+
53
+ **Lookup order.** Cheapest first; stop at the first that answers.
54
+
55
+ 1. A documentation tool the harness exposes for that vendor (Context7 and
56
+ the like) — one call, cheap.
57
+ 2. Official documentation fetched this turn — cheap.
58
+ 3. The lockfile pin plus that version's changelog — a narrow read.
59
+ 4. The installed package read on disk — `node_modules`, `site-packages`,
60
+ `vendor` — expensive; grep for the specific symbol, never read whole
61
+ files.
62
+
63
+ Rungs 1 and 2 answer for whichever version they happen to describe, which is
64
+ not always yours. Note the version each one reports and compare it to the pin;
65
+ a cheap answer that cannot say which version it describes has not answered.
66
+
67
+ **Authority.** When two sources disagree, the installed package wins — it
68
+ is the version that will execute. A cheap source that contradicts it is
69
+ wrong.
70
+
71
+ Not sources: your recollection; an older file in this repo calling the same API,
72
+ which may be the stale thing you are about to copy; a blog post; a search
73
+ snippet you did not open.
74
+
75
+ ## When it doesn't — Exemptions
76
+
77
+ A closed, named list. Outside it the default holds — no free pass by analogy.
78
+
79
+ 1. **First-party code** — defined in this repo or a sibling package in the same
80
+ workspace. Read it; a fetch would answer a question the tree already answers.
81
+ 2. **Standard library at a pinned runtime** — those shapes do not move between
82
+ two runs of the same interpreter.
83
+ 3. **Already checked this session** — one check per symbol. Cite the earlier
84
+ check rather than repeating it.
85
+ 4. **Covered by a red-to-green run against the real dependency** — a
86
+ leo:test-first cycle that exercises the actual library is this check, and its
87
+ transition is the record. Do not manufacture weaker evidence beside it.
88
+ 5. **No fetch capability in this session** — offline, or no docs tool reachable.
89
+ Then the claim is reported as unchecked and this exemption is named.
90
+
91
+ A skip must name its exemption in the report — "skipped freshness: first-party,
92
+ read src/auth/session.ts". An unnamed skip is an unchecked claim.
93
+
94
+ ## Recording the check
95
+
96
+ One line per check, in the done report:
97
+
98
+ ```
99
+ checked <symbol> against <source> (<version>)
100
+ ```
101
+
102
+ The version in parentheses is what makes it auditable — a reviewer compares it
103
+ to the lockfile without rerunning anything.
104
+
105
+ ## Self-talk to catch
106
+
107
+ - "I've used this library for years" — across how many major versions, and
108
+ which one is pinned here?
109
+ - "The docs will only confirm what I know" — then it costs nothing, and the
110
+ case where they do not is the entire reason for the step.
111
+ - "Another file here calls it this way" — that file may be what you are about
112
+ to propagate.
113
+ - "The typechecker will catch it" — a typechecker reads installed stubs, the
114
+ authority source. Say so and cite it, rather than skipping and hoping.
115
+ - "It's one argument" — argument names are exactly what moves between majors.
116
+
117
+ ## Reviewable finding
118
+
119
+ An unchecked third-party surface with no named exemption is a finding:
120
+ blocking when the call sits on the path the task was about, non-blocking
121
+ otherwise.
122
+
123
+ ## Works with
124
+
125
+ - leo:verification — that gate proves the code you wrote runs; this one governs
126
+ whether the API you wrote it against exists. A green test against a mocked
127
+ dependency satisfies that skill and not this one.
128
+ - leo:test-first — exemption 4; a red-to-green run against the real dependency
129
+ has already done this work.
130
+ - leo:debugging — when Localize follows a path into a dependency, this says
131
+ which copy of it to read.
@@ -0,0 +1,154 @@
1
+ ---
2
+ name: memory
3
+ description: >
4
+ Durable cross-harness facts, one per file, in a store that outlives the
5
+ session and every plugin update. Covers what earns a place in the store,
6
+ how a fact is written and revised, how to read one before acting on it,
7
+ and when to throw one away. The store is canonical; each harness's own
8
+ memory surface receives a generated copy of the global facts. Use when a
9
+ durable preference, repo rule, decision, or machine quirk surfaces. Do not
10
+ use for current-task state or facts already recorded in the repository.
11
+ when_to_use: >
12
+ A fact surfaces that will still be true next month — a stated preference,
13
+ a repo rule the code does not spell out, a settled decision, a machine
14
+ quirk that cost you a detour. Also when a remembered fact turns out wrong
15
+ and has to be revised or dropped. NOT for anything scoped to the current
16
+ task (branch names, what is failing right now — that is machine-local
17
+ JSON state), and NOT for material the repository already records.
18
+ ---
19
+
20
+ # memory
21
+
22
+ One fact per file, written the moment it is learned. A fact you intend to
23
+ record at the end of the session is a fact you will lose, because the end of
24
+ the session is exactly where context runs out.
25
+
26
+ The store is the only place you write. Each harness's native memory file
27
+ receives a generated copy of the global facts, so a preference learned on one
28
+ harness is in front of you on the next one. Those copies are derived — editing
29
+ one changes nothing and is overwritten on the next write.
30
+
31
+ ## What earns a place
32
+
33
+ All three must hold. Miss one and it is not a memory.
34
+
35
+ 1. **It is durable.** Still true a month from now. Not the branch you are on,
36
+ not the test that is failing, not where you are in the current task.
37
+ 2. **It is not cheaply re-derivable.** You could not recover it from one grep
38
+ or one file read in the repo you are already sitting in.
39
+ 3. **It fits one of the five types below.** There is no sixth type, and that
40
+ closed set is the whole gate.
41
+
42
+ ## The five types
43
+
44
+ 1. **preference** — Leo said how he wants something done, and it outlives this
45
+ task. *"Squash-merge, never a merge commit."*
46
+ 2. **convention** — a rule of this repo the code does not state, usually
47
+ learned the hard way. *"The adapters directory is generated; hand edits are
48
+ swept on the next render."*
49
+ 3. **environment** — a machine or tooling fact that cost a detour to establish.
50
+ Never a credential.
51
+ 4. **decision** — a settled choice and its one-line reason, where reopening it
52
+ would cost a conversation.
53
+ 5. **person** — who owns or decides what, and how to reach them about it.
54
+
55
+ ## When it doesn't — Exemptions
56
+
57
+ A closed, named list. Outside it the default holds — no free pass by analogy.
58
+
59
+ 1. **Task state** — anything true only until this task ends. Branch names, PR
60
+ numbers, what you are about to do next. That belongs in machine-local JSON
61
+ via `${CLAUDE_PLUGIN_ROOT}/scripts/state.py`, not here. (`${CLAUDE_PLUGIN_ROOT}`
62
+ is the Claude Code spelling of the plugin root and is not substituted into
63
+ this text; leo:delegation's ledger section gives the per-harness forms.)
64
+ 2. **Re-readable facts** — anything one search away in the working tree. The
65
+ repository is not something to memorize.
66
+ 3. **Your own conclusions** — an analysis, a diagnosis, a plan. A memory
67
+ records what Leo or the world asserted, not your reasoning about it.
68
+ 4. **Restatements of policy** — anything already in leo:using-leo or another
69
+ leo skill. Two copies of one rule drift apart, and the copy wins by being
70
+ nearer to hand.
71
+ 5. **Secrets** — tokens, keys, passwords, private URLs. Never, under any type:
72
+ the store is plain text on disk.
73
+ 6. **One-off corrections** — Leo redirecting you inside this task. Only a
74
+ correction he generalizes becomes a preference.
75
+
76
+ ## Rate discipline
77
+
78
+ Automatic capture without a brake becomes a log, and nobody trusts a log.
79
+
80
+ - At most **three** unprompted writes in a session. Reaching for a fourth means
81
+ you are recording activity, not learning facts — consolidate instead.
82
+ - Announce every write in one line: `remembered: <title> (preference)`. A store
83
+ that grows invisibly is a store Leo cannot audit.
84
+ - Check the scope before writing. A fact that restates one already there is a
85
+ revision of that file, never a second file beside it.
86
+
87
+ ## Procedure
88
+
89
+ Write, with the body on standard input:
90
+
91
+ ```sh
92
+ python3 "${CLAUDE_PLUGIN_ROOT}/scripts/memory.py" write global preference "Squash merge"
93
+ ```
94
+
95
+ Repo-scoped facts take an explicit key — the working directory is never
96
+ guessed, because a worktree would attribute the fact to the wrong project:
97
+
98
+ ```sh
99
+ python3 "${CLAUDE_PLUGIN_ROOT}/scripts/memory.py" write repo convention "Generated adapters" --repo owner/name
100
+ ```
101
+
102
+ Writing the same title again revises that file in place and keeps its original
103
+ creation date only when both title and type match exactly. A slug collision with
104
+ a different title or type receives the next `-N` suffix; a corrupt occupied
105
+ slot is never overwritten. `list` shows what is stored; `read <ref>` returns
106
+ one fact whole.
107
+
108
+ Leo-owned memory directories use mode `0700`; fact files and the generated
109
+ `index.json` and `MEMORY.md` use `0600`. A newly generated projection is also
110
+ private, while an existing user-owned projection target retains the mode the
111
+ user chose.
112
+
113
+ ## Read path
114
+
115
+ 1. The index arrives in context on every harness whose mapping says so. Each
116
+ line is a pointer, not the fact — the one-line hook is lossy by design.
117
+ 2. Read the file before you rely on it.
118
+ 3. **What you can see beats what you remember.** When a stored fact disagrees
119
+ with the repository in front of you, the repository is right. Use the
120
+ observation, then revise the memory. Never act on a fact you just watched
121
+ fail.
122
+
123
+ ## Forget path
124
+
125
+ Three triggers, and no others: Leo says it is wrong or has changed; you
126
+ observed it to be false; or its subject no longer exists.
127
+
128
+ ```sh
129
+ python3 "${CLAUDE_PLUGIN_ROOT}/scripts/memory.py" forget global/squash-merge
130
+ ```
131
+
132
+ Forgetting moves the file aside rather than destroying it, so a wrong call is
133
+ recoverable. A superseded fact is a revision, not a forget followed by a write.
134
+ Suspicion that something looks stale is not grounds to drop it — that needs an
135
+ assertion or an observation.
136
+
137
+ ## Self-talk to catch
138
+
139
+ - "I'll write this down once the task settles" — the task ending is what takes
140
+ the fact with it.
141
+ - "This is worth keeping, roughly" — name its type, or it does not go in.
142
+ - "The memory says the flag is called that" — the memory says what was true
143
+ when someone wrote it; check the flag.
144
+ - "Leo corrected me, that's a preference" — inside one task it is a
145
+ correction; only a generalization is a preference.
146
+
147
+ ## Works with
148
+
149
+ - leo:using-leo — draws the line this skill sits on: per-task JSON state on one
150
+ side, durable facts on the other.
151
+ - leo:doctor — reports whether the store exists and whether each harness
152
+ actually received its copy.
153
+ - leo:verification — a stored fact is not evidence. Claims still need a fresh
154
+ command run this turn.