leos-agent 6.1.1 → 7.0.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/LICENSE +21 -0
- package/README.md +18 -11
- package/adapters/cursor/agents/executor.md +2 -2
- package/adapters/cursor/agents/implementer.md +2 -2
- package/adapters/cursor/agents/review-lens.md +22 -0
- package/adapters/cursor/agents/reviewer.md +3 -2
- package/adapters/opencode/agents.json +44 -5
- package/adapters/opencode/plugin.js +435 -47
- package/config/MCP_PINS.md +17 -0
- package/config/models.json +647 -33
- package/hooks/bash-guard.py +51 -9
- package/hooks/session-start.py +27 -0
- package/package.json +19 -8
- package/roles/executor.md +2 -2
- package/roles/implementer.md +2 -2
- package/roles/review-lens.md +20 -0
- package/roles/reviewer.md +3 -2
- package/scripts/doctor.py +520 -0
- package/scripts/ghreview.py +558 -0
- package/scripts/jsonc_bridge.cjs +23 -0
- package/scripts/memory.py +744 -0
- package/scripts/render_adapters.py +253 -121
- package/scripts/resolve_attach_target.py +389 -0
- package/scripts/setup.py +1753 -0
- package/skills/brainstorming/SKILL.md +3 -1
- package/skills/debugging/SKILL.md +4 -2
- package/skills/delegation/SKILL.md +10 -8
- package/skills/doctor/SKILL.md +124 -0
- package/skills/executing-plans/SKILL.md +2 -1
- package/skills/finishing-a-branch/SKILL.md +4 -2
- package/skills/freshness/SKILL.md +131 -0
- package/skills/memory/SKILL.md +154 -0
- package/skills/resolve-ticket/SKILL.md +275 -0
- package/skills/review-pr/SKILL.md +327 -0
- package/skills/setup/SKILL.md +199 -0
- package/skills/setup/agents/openai.yaml +5 -0
- package/skills/test-first/SKILL.md +3 -1
- package/skills/using-leo/SKILL.md +18 -6
- package/skills/using-leo/references/claude-mapping.md +23 -1
- package/skills/using-leo/references/codex-mapping.md +18 -9
- package/skills/using-leo/references/cursor-mapping.md +19 -6
- package/skills/using-leo/references/hermes-mapping.md +18 -7
- package/skills/using-leo/references/opencode-mapping.md +18 -9
- package/skills/verification/SKILL.md +9 -1
- package/skills/visual-verification/SKILL.md +115 -0
- package/skills/watch-review/SKILL.md +128 -0
- package/skills/watch-review/agents/openai.yaml +5 -0
- package/skills/worktrees/SKILL.md +3 -1
- package/skills/writing-plans/SKILL.md +2 -1
- package/skills/writing-skills/SKILL.md +141 -0
- package/vendor/jsonc-parser-3.3.1/LICENSE.md +21 -0
- package/vendor/jsonc-parser-3.3.1/README.md +26 -0
- package/vendor/jsonc-parser-3.3.1/lib/umd/impl/edit.js +201 -0
- package/vendor/jsonc-parser-3.3.1/lib/umd/impl/format.js +275 -0
- package/vendor/jsonc-parser-3.3.1/lib/umd/impl/parser.js +682 -0
- package/vendor/jsonc-parser-3.3.1/lib/umd/impl/scanner.js +456 -0
- package/vendor/jsonc-parser-3.3.1/lib/umd/impl/string-intern.js +42 -0
- package/vendor/jsonc-parser-3.3.1/lib/umd/main.d.ts +351 -0
- package/vendor/jsonc-parser-3.3.1/lib/umd/main.js +194 -0
- package/vendor/jsonc-parser-3.3.1/package.json +37 -0
- package/workflows/cost-tiered-fix.js +32 -4
|
@@ -7,7 +7,9 @@ description: >
|
|
|
7
7
|
wide blast radius, hard to reverse, or that introduce new surface need
|
|
8
8
|
genuine, viable alternatives with trade-offs weighed before any code gets
|
|
9
9
|
written. Produces the chosen approach and its trade-offs, sized to the gate,
|
|
10
|
-
handed off to leo:writing-plans.
|
|
10
|
+
handed off to leo:writing-plans. Use when choosing an approach before
|
|
11
|
+
non-trivial code. Do not use for contained reversible tweaks, investigation,
|
|
12
|
+
or writing the plan itself.
|
|
11
13
|
when_to_use: >
|
|
12
14
|
Before starting non-trivial code: a new feature, a new integration surface,
|
|
13
15
|
a schema or data-model change, anything that's expensive or awkward to
|
|
@@ -5,7 +5,9 @@ description: >
|
|
|
5
5
|
behavior. Five named phases — Reproduce, Localize, Hypothesize, Prove, Fix —
|
|
6
6
|
each with an exit criterion, so a fix never lands before the cause is
|
|
7
7
|
pinned to file:line. Diagnosis is read-only judge work (investigator); the
|
|
8
|
-
fix happens separately, at the routed tier.
|
|
8
|
+
fix happens separately, at the routed tier. Use when a bug, failure, crash,
|
|
9
|
+
or surprising behavior needs diagnosis. Do not use for planned features,
|
|
10
|
+
reviewing a diff, or post-fix completion verification.
|
|
9
11
|
when_to_use: >
|
|
10
12
|
Any bug report, failing test, crash, stack trace, or "why does X happen"
|
|
11
13
|
before proposing a fix — used by the investigator agent and by the main
|
|
@@ -61,7 +63,7 @@ you, and go again.
|
|
|
61
63
|
|
|
62
64
|
After two failures on the same cause (two hypotheses tried and reverted, still
|
|
63
65
|
no Prove), step up one tier rather than retrying at the same one — investigator
|
|
64
|
-
|
|
66
|
+
explore findings feed a full investigator, and investigator itself steps to a
|
|
65
67
|
second, more evidence-fed pass, capped at Opus. A genuine deadlock, or two
|
|
66
68
|
Opus verdicts on the same cause that disagree → `expert`, announced in one
|
|
67
69
|
line ("escalating to expert: <question>") before it's invoked, never silent.
|
|
@@ -4,7 +4,8 @@ description: >
|
|
|
4
4
|
Operational mechanics for dispatching subagents — a single spawn or a large
|
|
5
5
|
fan-out — the companion to the policy's "Delegate the labor" section.
|
|
6
6
|
Covers brief construction, model/effort pinning, the four-state return
|
|
7
|
-
contract, and ledger-backed progress tracking for long multi-agent runs.
|
|
7
|
+
contract, and ledger-backed progress tracking for long multi-agent runs. Use
|
|
8
|
+
when dispatching any subagent or fan-out. Do not use to choose a task's tier.
|
|
8
9
|
when_to_use: >
|
|
9
10
|
Any time work is routed to a subagent (explore, investigator, executor,
|
|
10
11
|
implementer, reviewer, expert) rather than done inline — single dispatch or
|
|
@@ -36,8 +37,8 @@ version needs no follow-up question; the first invites three.
|
|
|
36
37
|
## Pin model and effort
|
|
37
38
|
|
|
38
39
|
Every dispatch pins **model AND effort** from the routing table — opus for
|
|
39
|
-
judges (reviewer, investigator), sonnet for
|
|
40
|
-
|
|
40
|
+
judges (reviewer, investigator), sonnet for normal implementation
|
|
41
|
+
(implementer), haiku for mechanical work (executor). expert
|
|
41
42
|
never appears in a fan-out — one at a time, never fanned. An unpinned call
|
|
42
43
|
silently inherits the session's tier: in an opus session that means every
|
|
43
44
|
executor spawn quietly runs at opus, and a ten-item fan-out burns
|
|
@@ -53,7 +54,7 @@ a report that hedges across two of them.
|
|
|
53
54
|
|---|---|---|
|
|
54
55
|
| `done` | Work finished, matches the brief | Verify against artifacts — see leo:verification — never take the self-report at face value |
|
|
55
56
|
| `concerns` | Finished, but flags something worth a second look | Read the concerns before accepting; they're often the real finding |
|
|
56
|
-
| `needs-context` | Blocked on missing information you can supply | Send the missing piece to the same agent (SendMessage
|
|
57
|
+
| `needs-context` | Blocked on missing information you can supply | Send the missing piece to the same agent (`SendMessage` on Claude Code, `followup_task` on Codex — elsewhere see the *Follow-up to a live agent* row of your mapping, and where none is established, cold re-dispatch with the context restated is the whole mechanism) so it keeps the context it already built. Either way **once** — a second needs-context on the same gap means the brief itself is broken, escalate the tier |
|
|
57
58
|
| `blocked` | Blocked on something you can't hand over inline | Resolve the blocker, or escalate per the ladder — never a silent same-tier retry |
|
|
58
59
|
|
|
59
60
|
`needs-context` and `blocked` look similar; the test is whether the missing
|
|
@@ -89,10 +90,11 @@ between "agent finished" and "ledger written" is exactly the gap this
|
|
|
89
90
|
exists to close.
|
|
90
91
|
|
|
91
92
|
`${CLAUDE_PLUGIN_ROOT}` above is the Claude Code spelling of the plugin root,
|
|
92
|
-
and it is substituted into this text only there.
|
|
93
|
-
|
|
94
|
-
|
|
95
|
-
|
|
93
|
+
and it is substituted into this text only there. Codex exposes `$PLUGIN_ROOT`
|
|
94
|
+
and Cursor `$CURSOR_PLUGIN_ROOT`. Hermes and OpenCode expose no root variable;
|
|
95
|
+
their injected policy substitutes an absolute payload path into the `state.py`
|
|
96
|
+
and `memory.py` commands, which is the discoverable source to reuse. Do not
|
|
97
|
+
invent an environment variable where the harness exposes none.
|
|
96
98
|
|
|
97
99
|
For a batch of independent, well-scoped fixes, don't hand-roll this loop —
|
|
98
100
|
the reusable workflow at `${CLAUDE_PLUGIN_ROOT}/workflows/cost-tiered-fix.js`
|
|
@@ -0,0 +1,124 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: doctor
|
|
3
|
+
description: >
|
|
4
|
+
Self-check for Leo's own wiring. Reports which harness this is, what each
|
|
5
|
+
tier name resolves to here, whether the bootstrap is installed, where
|
|
6
|
+
machine-local state and the memory store live, and which skills shipped
|
|
7
|
+
versus which this session can actually invoke. Disk facts come from a
|
|
8
|
+
helper script; the context facts only the running session can answer, and
|
|
9
|
+
a disagreement between the two columns is the diagnosis. Use when Leo asks
|
|
10
|
+
about Leo's loading, routing, or skill wiring. Do not use for project health
|
|
11
|
+
checks, project-code debugging, or unprompted inspection.
|
|
12
|
+
when_to_use: >
|
|
13
|
+
Leo asks whether the policy loaded, why routing or a skill is misbehaving,
|
|
14
|
+
or invokes doctor by name after installing, updating, or switching harness.
|
|
15
|
+
Also the first move when a leo skill cannot be found. NOT a general
|
|
16
|
+
environment or project health check, NOT for debugging the project's own
|
|
17
|
+
code (that is leo:debugging), and never run unprompted — it reports on the
|
|
18
|
+
agent, not on the work.
|
|
19
|
+
---
|
|
20
|
+
|
|
21
|
+
# doctor
|
|
22
|
+
|
|
23
|
+
Doctor answers two questions that look like one: what shipped to disk, and what
|
|
24
|
+
reached this session. A skill the harness never registered is indistinguishable
|
|
25
|
+
from a skill that does not exist, right up until the moment you invoke it.
|
|
26
|
+
|
|
27
|
+
## Run the script
|
|
28
|
+
|
|
29
|
+
```sh
|
|
30
|
+
python3 "${CLAUDE_PLUGIN_ROOT}/scripts/doctor.py" --harness <name>
|
|
31
|
+
```
|
|
32
|
+
|
|
33
|
+
Pass `--harness` with the harness you are on — the mapping appendix in your
|
|
34
|
+
context names it in its own heading (`# Hermes mapping` → `hermes`). Detection
|
|
35
|
+
without it relies on a plugin-root variable that Hermes and OpenCode do not
|
|
36
|
+
export, so on those two the script reports `unknown` rather than guessing.
|
|
37
|
+
`unknown` on a harness whose mapping you can plainly read is a missing
|
|
38
|
+
argument, not a fault.
|
|
39
|
+
|
|
40
|
+
`${CLAUDE_PLUGIN_ROOT}` is the Claude Code spelling. Codex exports
|
|
41
|
+
`$PLUGIN_ROOT` and Cursor `$CURSOR_PLUGIN_ROOT`. On Hermes and OpenCode no
|
|
42
|
+
plugin-root variable exists at all — the injected policy instead substitutes an
|
|
43
|
+
absolute payload path into its `state.py` and `memory.py` commands. Read that
|
|
44
|
+
command path from the policy as the discoverable source. Being unable to locate
|
|
45
|
+
the payload at all is itself the first finding: the harness is not looking where
|
|
46
|
+
the plugin was installed.
|
|
47
|
+
|
|
48
|
+
Add `--json` when you want the same facts as data.
|
|
49
|
+
|
|
50
|
+
Doctor validates the bootstrap that actually belongs to the named harness:
|
|
51
|
+
the session hook and manifest for Claude, Codex, and Cursor;
|
|
52
|
+
`config.instructions` in OpenCode's plugin; and Hermes registration plus its
|
|
53
|
+
first-tool-result fallback. It also reports the running Python version against
|
|
54
|
+
the supported 3.9+ floor. Codex hook trust is not provable from disk: review
|
|
55
|
+
the plugin in `/hooks` and confirm it is trusted before treating the on-disk
|
|
56
|
+
hook as active.
|
|
57
|
+
|
|
58
|
+
## Then answer the three it cannot
|
|
59
|
+
|
|
60
|
+
A script can prove the hook is installed and that the policy renders. It cannot
|
|
61
|
+
prove the policy arrived. Only you can see your own context.
|
|
62
|
+
|
|
63
|
+
1. **Did the policy load?** Look for the policy wrapper in your context, and
|
|
64
|
+
check that the mapping following it names *this* harness. A policy present
|
|
65
|
+
but carrying another harness's mapping is worse than none, because routing
|
|
66
|
+
then points at models that do not exist here.
|
|
67
|
+
2. **Which skills are actually invocable?** Compare your own skill list against
|
|
68
|
+
the script's shipped roster. Mind the naming rule: most harnesses namespace
|
|
69
|
+
them as `leo:<name>`, while OpenCode has no namespace and requires a
|
|
70
|
+
skill's frontmatter name to match its directory, so the plugin registers a
|
|
71
|
+
generated shadow copy with every skill renamed `leo-<name>`. A skill that
|
|
72
|
+
looks missing on OpenCode may simply be listed as `leo-<name>` rather than
|
|
73
|
+
`leo:<name>`.
|
|
74
|
+
3. **Is memory present and delivered?** The script reports whether the store
|
|
75
|
+
exists and whether each native surface received its generated copy. Whether
|
|
76
|
+
those facts are in front of you right now is something only you can confirm.
|
|
77
|
+
Report the two separately; they disagree more often than expected. Hermes's
|
|
78
|
+
projection is opt-in, so doctor reports it explicitly as disabled rather
|
|
79
|
+
than silently omitting it.
|
|
80
|
+
|
|
81
|
+
## Reading the report
|
|
82
|
+
|
|
83
|
+
Every row carries its source — `env`, `disk`, `config`, or `context` — so a
|
|
84
|
+
reader can tell a fact from an inference. Close with one verdict from exactly
|
|
85
|
+
three: **healthy**, **degraded**, or **not loaded**. Never free prose. `not
|
|
86
|
+
loaded` outranks everything else: if the policy did not arrive, nothing else in
|
|
87
|
+
the report describes how this session will actually behave.
|
|
88
|
+
|
|
89
|
+
**Most breadcrumb logs are history, not a verdict.** Some older logs carry no
|
|
90
|
+
timestamps, and the test suite drives failure paths deliberately, so entries
|
|
91
|
+
can accumulate on a development machine. Quote the newest line if useful, but
|
|
92
|
+
never conclude "the hook failed this session" from history alone. The one
|
|
93
|
+
capability exception is `opencode-skills.log`: its presence means namespaced
|
|
94
|
+
OpenCode skill registration degraded and doctor reports that state until the
|
|
95
|
+
breadcrumb is cleared after the underlying problem is understood.
|
|
96
|
+
|
|
97
|
+
## Failure modes
|
|
98
|
+
|
|
99
|
+
| Symptom | Likely cause | Fix |
|
|
100
|
+
|---|---|---|
|
|
101
|
+
| Policy absent, bootstrap installed | the hook fired and failed open | read the newest breadcrumb, then confirm it describes this session before believing it |
|
|
102
|
+
| Policy present, mapping names another harness | detection resolved wrong, usually a stray plugin-root variable exported in an unrelated shell | unset it, restart the session |
|
|
103
|
+
| Harness reported as `unknown` | no `--harness`, and this harness exports no plugin-root variable | degraded until re-run with `--harness <name>` read off your mapping heading; a still-unknown explicit run is invalid wiring |
|
|
104
|
+
| Codex hook is on disk but policy is absent | the new or changed hook may not be trusted | open `/hooks`, review the hook, and explicitly trust it |
|
|
105
|
+
| OpenCode reports `opencode-skills.log` | the namespaced shadow tree failed and no bare-name fallback was registered | inspect the newest breadcrumb, fix the path/permission failure, and restart OpenCode |
|
|
106
|
+
| Shipped roster exceeds what you can invoke | the harness cached an older payload, or the skills directory is not registered | update the plugin; on OpenCode check `opencode debug skill` for each skill's `location` |
|
|
107
|
+
| Skills listed as `leo-<name>` instead of `leo:<name>` | OpenCode, working as designed | invoke them as `leo-<name>`; not a fault |
|
|
108
|
+
| Tier names resolve to models this harness cannot run | mapping and harness disagree | same as row 2 |
|
|
109
|
+
| Machine-local state not writable | the path override points somewhere unwritable | fix or unset it |
|
|
110
|
+
| A skill is genuinely absent from disk | it was never added | see leo:writing-skills |
|
|
111
|
+
|
|
112
|
+
## Doctor never repairs
|
|
113
|
+
|
|
114
|
+
It reports, and it names the fix. It does not reinstall, rewrite configuration,
|
|
115
|
+
or delete state — which is what keeps it safe to run at any tier and at any
|
|
116
|
+
moment.
|
|
117
|
+
|
|
118
|
+
## Works with
|
|
119
|
+
|
|
120
|
+
- leo:writing-skills — for a skill that turned out to be missing because nobody
|
|
121
|
+
wrote it yet.
|
|
122
|
+
- leo:memory — doctor reports whether the store exists and reached each surface.
|
|
123
|
+
- leo:verification — this report is a claim like any other: the script ran this
|
|
124
|
+
turn and its output was read.
|
|
@@ -5,7 +5,8 @@ description: >
|
|
|
5
5
|
with a check at every batch boundary, plan-intent-wins-on-architecture /
|
|
6
6
|
reality-wins-on-mechanics arbitration, and one fix-then-re-review cycle
|
|
7
7
|
before stopping to report. Used by the implementer agent, or the main loop
|
|
8
|
-
when it executes a plan directly.
|
|
8
|
+
when it executes a plan directly. Use when a written plan is about to become
|
|
9
|
+
code. Do not use for open-ended work without a plan or for reviewing a diff.
|
|
9
10
|
when_to_use: >
|
|
10
11
|
A written plan (from planner, an issue, or Leo's own outline) is about to
|
|
11
12
|
be turned into code. NOT for open-ended implementation with no plan
|
|
@@ -6,7 +6,9 @@ description: >
|
|
|
6
6
|
a closed set of next steps — merge / PR / keep / discard — routes the
|
|
7
7
|
chosen path through the right ordering (land the work before removing the
|
|
8
8
|
worktree, remove the worktree before deleting the branch), and leaves the
|
|
9
|
-
repo clean.
|
|
9
|
+
repo clean. Use when reviewed implementation on a branch/worktree needs a
|
|
10
|
+
terminal disposition. Do not use to manage a worktree mid-task or replace
|
|
11
|
+
the review cycle.
|
|
10
12
|
when_to_use: >
|
|
11
13
|
A branch or worktree has reached "implementation done" and Leo needs to
|
|
12
14
|
decide what happens to it. Fires after execute-then-review completes, or
|
|
@@ -71,7 +73,7 @@ Per option:
|
|
|
71
73
|
| Option | Sequence |
|
|
72
74
|
|---|---|
|
|
73
75
|
| merge | merge locally or `gh pr merge` → remove worktree (`leo:worktrees`) → delete local branch |
|
|
74
|
-
| PR | push branch → open PR → **stop** (worktree and branch stay; nothing is
|
|
76
|
+
| PR | push branch → open PR → **stop** (worktree and branch stay; nothing is merged yet.) |
|
|
75
77
|
| keep | do nothing destructive; leave worktree and branch as-is |
|
|
76
78
|
| discard | typed confirmation (below) → remove worktree (`leo:worktrees`) → force-delete branch |
|
|
77
79
|
|
|
@@ -0,0 +1,131 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: freshness
|
|
3
|
+
description: >
|
|
4
|
+
Currency gate for code written against anything outside this repository.
|
|
5
|
+
Before a library call, CLI flag, endpoint field, or vendor number is
|
|
6
|
+
committed to, its shape is confirmed against a source that reflects the
|
|
7
|
+
version this project actually runs — the installed package, the lockfile
|
|
8
|
+
pin, or documentation fetched this turn. Each check is recorded by symbol
|
|
9
|
+
and source in the report. Use when writing, reviewing, or asserting a
|
|
10
|
+
third-party surface. Do not use for first-party code, pinned standard
|
|
11
|
+
libraries, or as a substitute for verification.
|
|
12
|
+
when_to_use: >
|
|
13
|
+
About to write, review, or assert the shape of a third-party surface — an
|
|
14
|
+
import path, an argument list, a config key, an HTTP field, an auth
|
|
15
|
+
scheme, a model id, a price, a deprecation claim. NOT for first-party code
|
|
16
|
+
in this workspace (read it instead), NOT for the standard library of a
|
|
17
|
+
pinned runtime, and NOT a substitute for running anything — leo:verification
|
|
18
|
+
still governs the completion claim built on top of it.
|
|
19
|
+
---
|
|
20
|
+
|
|
21
|
+
# freshness
|
|
22
|
+
|
|
23
|
+
A third-party surface you have not read this session is a guess, however
|
|
24
|
+
familiar it feels. Recall of a library is a snapshot of some arbitrary past
|
|
25
|
+
version; it is not a snapshot of the one pinned in this lockfile. The cost is
|
|
26
|
+
a call that reads perfectly and does not exist.
|
|
27
|
+
|
|
28
|
+
Recall is not a source. The package installed on disk is.
|
|
29
|
+
|
|
30
|
+
## When it fires
|
|
31
|
+
|
|
32
|
+
A closed list of five.
|
|
33
|
+
|
|
34
|
+
1. **A symbol you did not read this session** — a function, method, class,
|
|
35
|
+
decorator, flag, or config key belonging to something not defined in this
|
|
36
|
+
working tree.
|
|
37
|
+
2. **A version-sensitive call shape** — argument order, keyword names, return
|
|
38
|
+
type, or import path for a dependency whose installed version you have not
|
|
39
|
+
confirmed.
|
|
40
|
+
3. **A service contract** — endpoint path, request or response field, auth
|
|
41
|
+
scheme, pagination rule, error code.
|
|
42
|
+
4. **A vendor-schedule fact** — a model id, context window, price, rate limit,
|
|
43
|
+
or regional availability. These move on someone else's calendar.
|
|
44
|
+
5. **A deprecation or removal claim** — "that was dropped in v3" is an
|
|
45
|
+
assertion about a moving target and needs the same check as a signature.
|
|
46
|
+
|
|
47
|
+
Outside these five, write the code.
|
|
48
|
+
|
|
49
|
+
## What counts as a source
|
|
50
|
+
|
|
51
|
+
Two different questions — which to reach for, and which one wins.
|
|
52
|
+
|
|
53
|
+
**Lookup order.** Cheapest first; stop at the first that answers.
|
|
54
|
+
|
|
55
|
+
1. A documentation tool the harness exposes for that vendor (Context7 and
|
|
56
|
+
the like) — one call, cheap.
|
|
57
|
+
2. Official documentation fetched this turn — cheap.
|
|
58
|
+
3. The lockfile pin plus that version's changelog — a narrow read.
|
|
59
|
+
4. The installed package read on disk — `node_modules`, `site-packages`,
|
|
60
|
+
`vendor` — expensive; grep for the specific symbol, never read whole
|
|
61
|
+
files.
|
|
62
|
+
|
|
63
|
+
Rungs 1 and 2 answer for whichever version they happen to describe, which is
|
|
64
|
+
not always yours. Note the version each one reports and compare it to the pin;
|
|
65
|
+
a cheap answer that cannot say which version it describes has not answered.
|
|
66
|
+
|
|
67
|
+
**Authority.** When two sources disagree, the installed package wins — it
|
|
68
|
+
is the version that will execute. A cheap source that contradicts it is
|
|
69
|
+
wrong.
|
|
70
|
+
|
|
71
|
+
Not sources: your recollection; an older file in this repo calling the same API,
|
|
72
|
+
which may be the stale thing you are about to copy; a blog post; a search
|
|
73
|
+
snippet you did not open.
|
|
74
|
+
|
|
75
|
+
## When it doesn't — Exemptions
|
|
76
|
+
|
|
77
|
+
A closed, named list. Outside it the default holds — no free pass by analogy.
|
|
78
|
+
|
|
79
|
+
1. **First-party code** — defined in this repo or a sibling package in the same
|
|
80
|
+
workspace. Read it; a fetch would answer a question the tree already answers.
|
|
81
|
+
2. **Standard library at a pinned runtime** — those shapes do not move between
|
|
82
|
+
two runs of the same interpreter.
|
|
83
|
+
3. **Already checked this session** — one check per symbol. Cite the earlier
|
|
84
|
+
check rather than repeating it.
|
|
85
|
+
4. **Covered by a red-to-green run against the real dependency** — a
|
|
86
|
+
leo:test-first cycle that exercises the actual library is this check, and its
|
|
87
|
+
transition is the record. Do not manufacture weaker evidence beside it.
|
|
88
|
+
5. **No fetch capability in this session** — offline, or no docs tool reachable.
|
|
89
|
+
Then the claim is reported as unchecked and this exemption is named.
|
|
90
|
+
|
|
91
|
+
A skip must name its exemption in the report — "skipped freshness: first-party,
|
|
92
|
+
read src/auth/session.ts". An unnamed skip is an unchecked claim.
|
|
93
|
+
|
|
94
|
+
## Recording the check
|
|
95
|
+
|
|
96
|
+
One line per check, in the done report:
|
|
97
|
+
|
|
98
|
+
```
|
|
99
|
+
checked <symbol> against <source> (<version>)
|
|
100
|
+
```
|
|
101
|
+
|
|
102
|
+
The version in parentheses is what makes it auditable — a reviewer compares it
|
|
103
|
+
to the lockfile without rerunning anything.
|
|
104
|
+
|
|
105
|
+
## Self-talk to catch
|
|
106
|
+
|
|
107
|
+
- "I've used this library for years" — across how many major versions, and
|
|
108
|
+
which one is pinned here?
|
|
109
|
+
- "The docs will only confirm what I know" — then it costs nothing, and the
|
|
110
|
+
case where they do not is the entire reason for the step.
|
|
111
|
+
- "Another file here calls it this way" — that file may be what you are about
|
|
112
|
+
to propagate.
|
|
113
|
+
- "The typechecker will catch it" — a typechecker reads installed stubs, the
|
|
114
|
+
authority source. Say so and cite it, rather than skipping and hoping.
|
|
115
|
+
- "It's one argument" — argument names are exactly what moves between majors.
|
|
116
|
+
|
|
117
|
+
## Reviewable finding
|
|
118
|
+
|
|
119
|
+
An unchecked third-party surface with no named exemption is a finding:
|
|
120
|
+
blocking when the call sits on the path the task was about, non-blocking
|
|
121
|
+
otherwise.
|
|
122
|
+
|
|
123
|
+
## Works with
|
|
124
|
+
|
|
125
|
+
- leo:verification — that gate proves the code you wrote runs; this one governs
|
|
126
|
+
whether the API you wrote it against exists. A green test against a mocked
|
|
127
|
+
dependency satisfies that skill and not this one.
|
|
128
|
+
- leo:test-first — exemption 4; a red-to-green run against the real dependency
|
|
129
|
+
has already done this work.
|
|
130
|
+
- leo:debugging — when Localize follows a path into a dependency, this says
|
|
131
|
+
which copy of it to read.
|
|
@@ -0,0 +1,154 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: memory
|
|
3
|
+
description: >
|
|
4
|
+
Durable cross-harness facts, one per file, in a store that outlives the
|
|
5
|
+
session and every plugin update. Covers what earns a place in the store,
|
|
6
|
+
how a fact is written and revised, how to read one before acting on it,
|
|
7
|
+
and when to throw one away. The store is canonical; each harness's own
|
|
8
|
+
memory surface receives a generated copy of the global facts. Use when a
|
|
9
|
+
durable preference, repo rule, decision, or machine quirk surfaces. Do not
|
|
10
|
+
use for current-task state or facts already recorded in the repository.
|
|
11
|
+
when_to_use: >
|
|
12
|
+
A fact surfaces that will still be true next month — a stated preference,
|
|
13
|
+
a repo rule the code does not spell out, a settled decision, a machine
|
|
14
|
+
quirk that cost you a detour. Also when a remembered fact turns out wrong
|
|
15
|
+
and has to be revised or dropped. NOT for anything scoped to the current
|
|
16
|
+
task (branch names, what is failing right now — that is machine-local
|
|
17
|
+
JSON state), and NOT for material the repository already records.
|
|
18
|
+
---
|
|
19
|
+
|
|
20
|
+
# memory
|
|
21
|
+
|
|
22
|
+
One fact per file, written the moment it is learned. A fact you intend to
|
|
23
|
+
record at the end of the session is a fact you will lose, because the end of
|
|
24
|
+
the session is exactly where context runs out.
|
|
25
|
+
|
|
26
|
+
The store is the only place you write. Each harness's native memory file
|
|
27
|
+
receives a generated copy of the global facts, so a preference learned on one
|
|
28
|
+
harness is in front of you on the next one. Those copies are derived — editing
|
|
29
|
+
one changes nothing and is overwritten on the next write.
|
|
30
|
+
|
|
31
|
+
## What earns a place
|
|
32
|
+
|
|
33
|
+
All three must hold. Miss one and it is not a memory.
|
|
34
|
+
|
|
35
|
+
1. **It is durable.** Still true a month from now. Not the branch you are on,
|
|
36
|
+
not the test that is failing, not where you are in the current task.
|
|
37
|
+
2. **It is not cheaply re-derivable.** You could not recover it from one grep
|
|
38
|
+
or one file read in the repo you are already sitting in.
|
|
39
|
+
3. **It fits one of the five types below.** There is no sixth type, and that
|
|
40
|
+
closed set is the whole gate.
|
|
41
|
+
|
|
42
|
+
## The five types
|
|
43
|
+
|
|
44
|
+
1. **preference** — Leo said how he wants something done, and it outlives this
|
|
45
|
+
task. *"Squash-merge, never a merge commit."*
|
|
46
|
+
2. **convention** — a rule of this repo the code does not state, usually
|
|
47
|
+
learned the hard way. *"The adapters directory is generated; hand edits are
|
|
48
|
+
swept on the next render."*
|
|
49
|
+
3. **environment** — a machine or tooling fact that cost a detour to establish.
|
|
50
|
+
Never a credential.
|
|
51
|
+
4. **decision** — a settled choice and its one-line reason, where reopening it
|
|
52
|
+
would cost a conversation.
|
|
53
|
+
5. **person** — who owns or decides what, and how to reach them about it.
|
|
54
|
+
|
|
55
|
+
## When it doesn't — Exemptions
|
|
56
|
+
|
|
57
|
+
A closed, named list. Outside it the default holds — no free pass by analogy.
|
|
58
|
+
|
|
59
|
+
1. **Task state** — anything true only until this task ends. Branch names, PR
|
|
60
|
+
numbers, what you are about to do next. That belongs in machine-local JSON
|
|
61
|
+
via `${CLAUDE_PLUGIN_ROOT}/scripts/state.py`, not here. (`${CLAUDE_PLUGIN_ROOT}`
|
|
62
|
+
is the Claude Code spelling of the plugin root and is not substituted into
|
|
63
|
+
this text; leo:delegation's ledger section gives the per-harness forms.)
|
|
64
|
+
2. **Re-readable facts** — anything one search away in the working tree. The
|
|
65
|
+
repository is not something to memorize.
|
|
66
|
+
3. **Your own conclusions** — an analysis, a diagnosis, a plan. A memory
|
|
67
|
+
records what Leo or the world asserted, not your reasoning about it.
|
|
68
|
+
4. **Restatements of policy** — anything already in leo:using-leo or another
|
|
69
|
+
leo skill. Two copies of one rule drift apart, and the copy wins by being
|
|
70
|
+
nearer to hand.
|
|
71
|
+
5. **Secrets** — tokens, keys, passwords, private URLs. Never, under any type:
|
|
72
|
+
the store is plain text on disk.
|
|
73
|
+
6. **One-off corrections** — Leo redirecting you inside this task. Only a
|
|
74
|
+
correction he generalizes becomes a preference.
|
|
75
|
+
|
|
76
|
+
## Rate discipline
|
|
77
|
+
|
|
78
|
+
Automatic capture without a brake becomes a log, and nobody trusts a log.
|
|
79
|
+
|
|
80
|
+
- At most **three** unprompted writes in a session. Reaching for a fourth means
|
|
81
|
+
you are recording activity, not learning facts — consolidate instead.
|
|
82
|
+
- Announce every write in one line: `remembered: <title> (preference)`. A store
|
|
83
|
+
that grows invisibly is a store Leo cannot audit.
|
|
84
|
+
- Check the scope before writing. A fact that restates one already there is a
|
|
85
|
+
revision of that file, never a second file beside it.
|
|
86
|
+
|
|
87
|
+
## Procedure
|
|
88
|
+
|
|
89
|
+
Write, with the body on standard input:
|
|
90
|
+
|
|
91
|
+
```sh
|
|
92
|
+
python3 "${CLAUDE_PLUGIN_ROOT}/scripts/memory.py" write global preference "Squash merge"
|
|
93
|
+
```
|
|
94
|
+
|
|
95
|
+
Repo-scoped facts take an explicit key — the working directory is never
|
|
96
|
+
guessed, because a worktree would attribute the fact to the wrong project:
|
|
97
|
+
|
|
98
|
+
```sh
|
|
99
|
+
python3 "${CLAUDE_PLUGIN_ROOT}/scripts/memory.py" write repo convention "Generated adapters" --repo owner/name
|
|
100
|
+
```
|
|
101
|
+
|
|
102
|
+
Writing the same title again revises that file in place and keeps its original
|
|
103
|
+
creation date only when both title and type match exactly. A slug collision with
|
|
104
|
+
a different title or type receives the next `-N` suffix; a corrupt occupied
|
|
105
|
+
slot is never overwritten. `list` shows what is stored; `read <ref>` returns
|
|
106
|
+
one fact whole.
|
|
107
|
+
|
|
108
|
+
Leo-owned memory directories use mode `0700`; fact files and the generated
|
|
109
|
+
`index.json` and `MEMORY.md` use `0600`. A newly generated projection is also
|
|
110
|
+
private, while an existing user-owned projection target retains the mode the
|
|
111
|
+
user chose.
|
|
112
|
+
|
|
113
|
+
## Read path
|
|
114
|
+
|
|
115
|
+
1. The index arrives in context on every harness whose mapping says so. Each
|
|
116
|
+
line is a pointer, not the fact — the one-line hook is lossy by design.
|
|
117
|
+
2. Read the file before you rely on it.
|
|
118
|
+
3. **What you can see beats what you remember.** When a stored fact disagrees
|
|
119
|
+
with the repository in front of you, the repository is right. Use the
|
|
120
|
+
observation, then revise the memory. Never act on a fact you just watched
|
|
121
|
+
fail.
|
|
122
|
+
|
|
123
|
+
## Forget path
|
|
124
|
+
|
|
125
|
+
Three triggers, and no others: Leo says it is wrong or has changed; you
|
|
126
|
+
observed it to be false; or its subject no longer exists.
|
|
127
|
+
|
|
128
|
+
```sh
|
|
129
|
+
python3 "${CLAUDE_PLUGIN_ROOT}/scripts/memory.py" forget global/squash-merge
|
|
130
|
+
```
|
|
131
|
+
|
|
132
|
+
Forgetting moves the file aside rather than destroying it, so a wrong call is
|
|
133
|
+
recoverable. A superseded fact is a revision, not a forget followed by a write.
|
|
134
|
+
Suspicion that something looks stale is not grounds to drop it — that needs an
|
|
135
|
+
assertion or an observation.
|
|
136
|
+
|
|
137
|
+
## Self-talk to catch
|
|
138
|
+
|
|
139
|
+
- "I'll write this down once the task settles" — the task ending is what takes
|
|
140
|
+
the fact with it.
|
|
141
|
+
- "This is worth keeping, roughly" — name its type, or it does not go in.
|
|
142
|
+
- "The memory says the flag is called that" — the memory says what was true
|
|
143
|
+
when someone wrote it; check the flag.
|
|
144
|
+
- "Leo corrected me, that's a preference" — inside one task it is a
|
|
145
|
+
correction; only a generalization is a preference.
|
|
146
|
+
|
|
147
|
+
## Works with
|
|
148
|
+
|
|
149
|
+
- leo:using-leo — draws the line this skill sits on: per-task JSON state on one
|
|
150
|
+
side, durable facts on the other.
|
|
151
|
+
- leo:doctor — reports whether the store exists and whether each harness
|
|
152
|
+
actually received its copy.
|
|
153
|
+
- leo:verification — a stored fact is not evidence. Claims still need a fresh
|
|
154
|
+
command run this turn.
|