leos-agent 6.3.0 → 10.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/LICENSE +21 -0
- package/README.md +547 -24
- package/commands/handoff.md +11 -0
- package/commands/handon.md +10 -0
- package/commands/leo-doctor.md +22 -0
- package/commands/leo-install.md +9 -0
- package/commands/review-pr.md +9 -0
- package/commands-claude/watch-review.md +9 -0
- package/index.js +12 -0
- package/package.json +30 -18
- package/payload/codex-agents/leo-executor.toml +36 -0
- package/payload/codex-agents/leo-runner.toml +28 -0
- package/rules/preferences.md +97 -0
- package/scripts/check.py +244 -0
- package/scripts/ghreview.py +24 -6
- package/scripts/handoff.py +183 -0
- package/scripts/leo-install.py +509 -0
- package/scripts/measure_context.py +113 -0
- package/scripts/publish-npm.py +138 -0
- package/scripts/resolve_attach_target.py +45 -13
- package/scripts/watch_review.py +169 -0
- package/skills/doctor/SKILL.md +73 -96
- package/skills/doctor/agents/openai.yaml +5 -0
- package/skills/handoff/SKILL.md +99 -0
- package/skills/handoff/agents/openai.yaml +5 -0
- package/skills/handon/SKILL.md +61 -0
- package/skills/install/SKILL.md +79 -0
- package/skills/install/agents/openai.yaml +5 -0
- package/skills/review-pr/SKILL.md +59 -308
- package/skills/review-pr/reference/lenses.md +67 -0
- package/skills/review-pr/reference/procedure.md +348 -0
- package/skills-claude/attach-pr/SKILL.md +178 -0
- package/skills-claude/watch-review/SKILL.md +91 -0
- package/adapters/cursor/agents/executor.md +0 -17
- package/adapters/cursor/agents/expert.md +0 -70
- package/adapters/cursor/agents/explore.md +0 -16
- package/adapters/cursor/agents/implementer.md +0 -18
- package/adapters/cursor/agents/investigator.md +0 -18
- package/adapters/cursor/agents/planner.md +0 -28
- package/adapters/cursor/agents/reviewer.md +0 -34
- package/adapters/opencode/agents.json +0 -66
- package/adapters/opencode/plugin.js +0 -288
- package/config/models.json +0 -408
- package/hooks/bash-guard.py +0 -541
- package/hooks/cursor-guard.py +0 -84
- package/hooks/hooks-cursor.json +0 -11
- package/hooks/hooks.json +0 -20
- package/hooks/session-start.py +0 -148
- package/roles/executor.md +0 -15
- package/roles/expert.md +0 -67
- package/roles/explore.md +0 -13
- package/roles/implementer.md +0 -16
- package/roles/investigator.md +0 -15
- package/roles/planner.md +0 -25
- package/roles/reviewer.md +0 -31
- package/scripts/doctor.py +0 -284
- package/scripts/memory.py +0 -705
- package/scripts/render_adapters.py +0 -473
- package/scripts/setup.py +0 -161
- package/settings.json +0 -7
- package/skills/.gitkeep +0 -0
- package/skills/brainstorming/SKILL.md +0 -109
- package/skills/debugging/SKILL.md +0 -98
- package/skills/delegation/SKILL.md +0 -141
- package/skills/executing-plans/SKILL.md +0 -116
- package/skills/finishing-a-branch/SKILL.md +0 -123
- package/skills/freshness/SKILL.md +0 -118
- package/skills/memory/SKILL.md +0 -144
- package/skills/resolve-ticket/SKILL.md +0 -269
- package/skills/setup/SKILL.md +0 -85
- package/skills/test-first/SKILL.md +0 -90
- package/skills/using-leo/SKILL.md +0 -96
- package/skills/using-leo/references/claude-mapping.md +0 -32
- package/skills/using-leo/references/codex-mapping.md +0 -34
- package/skills/using-leo/references/cursor-mapping.md +0 -34
- package/skills/using-leo/references/hermes-mapping.md +0 -36
- package/skills/using-leo/references/opencode-mapping.md +0 -36
- package/skills/verification/SKILL.md +0 -109
- package/skills/visual-verification/SKILL.md +0 -114
- package/skills/watch-review/SKILL.md +0 -125
- package/skills/worktrees/SKILL.md +0 -129
- package/skills/writing-plans/SKILL.md +0 -96
- package/skills/writing-skills/SKILL.md +0 -134
- package/workflows/cost-tiered-fix.js +0 -259
|
@@ -1,90 +0,0 @@
|
|
|
1
|
-
---
|
|
2
|
-
name: test-first
|
|
3
|
-
description: >
|
|
4
|
-
Failing-test-first as the default for runtime-behavior changes. Before
|
|
5
|
-
writing the change, write a test that fails for the intended reason, watch
|
|
6
|
-
it fail, then make it pass with the change — the red-to-green transition is
|
|
7
|
-
the evidence a real guard exists, not a rubber stamp added after the fact.
|
|
8
|
-
Applies to implementer and executor during implementation; reviewer checks
|
|
9
|
-
the resulting coverage as a rubric line.
|
|
10
|
-
when_to_use: >
|
|
11
|
-
Any implementation task that changes runtime behavior — fix, feature,
|
|
12
|
-
refactor with observable effect — routed through implementer or executor.
|
|
13
|
-
NOT for spikes/throwaway exploration that gets deleted, NOT for
|
|
14
|
-
docs/comments/config/dependency-bump edits, and NOT for pure UI copy or
|
|
15
|
-
styling tweaks — see Exemptions below for the full closed list.
|
|
16
|
-
---
|
|
17
|
-
|
|
18
|
-
# test-first
|
|
19
|
-
|
|
20
|
-
**Core rule**: before writing the change, write a failing test — and watch
|
|
21
|
-
it fail — for the reason the change is supposed to fix. Then make it pass.
|
|
22
|
-
A test that passes on its first run proves nothing about what it guards; it
|
|
23
|
-
could be checking the wrong thing, hitting a no-op path, or asserting
|
|
24
|
-
something already true.
|
|
25
|
-
|
|
26
|
-
## When it fires
|
|
27
|
-
|
|
28
|
-
Any task that changes runtime behavior: a bug fix, a new code path, a
|
|
29
|
-
refactor that alters observable output. If the diff can make a program do
|
|
30
|
-
something different, this skill applies before the diff is written.
|
|
31
|
-
|
|
32
|
-
## When it doesn't — Exemptions
|
|
33
|
-
|
|
34
|
-
A closed, named list. Outside it, the default holds — no free pass by
|
|
35
|
-
analogy, no "this one's basically like a spike."
|
|
36
|
-
|
|
37
|
-
1. **Spike** — throwaway exploration that gets deleted, never merged. If it
|
|
38
|
-
survives into the diff, it was not a spike; go back and cover it.
|
|
39
|
-
2. **Docs / comments / config / dependency bumps** — no runtime behavior
|
|
40
|
-
changes, nothing to guard with a test.
|
|
41
|
-
3. **Pure UI copy or styling tweaks** — text or CSS changes with no logic
|
|
42
|
-
branch behind them.
|
|
43
|
-
|
|
44
|
-
A skip must name its exemption in the report — "skipped test-first: spike,
|
|
45
|
-
deleted before merge" or "skipped test-first: config only." An unnamed skip
|
|
46
|
-
is not a skip; treat it as coverage missing.
|
|
47
|
-
|
|
48
|
-
## Procedure
|
|
49
|
-
|
|
50
|
-
1. Write the test first, targeting the exact failure the change is meant to
|
|
51
|
-
fix (the bug's symptom, or the new behavior's absence).
|
|
52
|
-
2. Run it. Watch it fail — and confirm it fails for the intended reason, not
|
|
53
|
-
a typo, import error, or wrong assertion. A red test that fails for the
|
|
54
|
-
wrong reason is as useless as one that never went red.
|
|
55
|
-
3. Make the change.
|
|
56
|
-
4. Run the test again. Green confirms the change closed the gap the red run
|
|
57
|
-
opened — this red-to-green transition is the evidence, and it's the same
|
|
58
|
-
evidence leo:verification asks for when confirming a change actually
|
|
59
|
-
works end-to-end: don't produce it twice in different words, point to it.
|
|
60
|
-
5. Report which exemption applied, or report the red-then-green pair (what
|
|
61
|
-
failed, what changed, what passed).
|
|
62
|
-
|
|
63
|
-
## Self-talk to catch
|
|
64
|
-
|
|
65
|
-
- "I'll add the test after, same effect" — it isn't. A test written against
|
|
66
|
-
passing code never proves it can fail; you've verified the assertion
|
|
67
|
-
compiles, not that it guards anything.
|
|
68
|
-
- "This is basically a spike" — if it's in the diff you're about to submit,
|
|
69
|
-
it isn't a spike. Spikes get deleted, not merged.
|
|
70
|
-
- "It's small, not worth a test" — size isn't in the exemption list.
|
|
71
|
-
Behavior change is the trigger, not line count.
|
|
72
|
-
- "I ran it and it passed, close enough" — passing without ever having seen
|
|
73
|
-
it fail is not evidence. Go back and force the fail first.
|
|
74
|
-
|
|
75
|
-
## Reviewable finding
|
|
76
|
-
|
|
77
|
-
Changed runtime behavior with no test that would fail without the change is
|
|
78
|
-
a reviewable finding — blocking when the behavior is load-bearing (the
|
|
79
|
-
user-facing or system-critical path the task was actually about), otherwise
|
|
80
|
-
non-blocking. The reviewer checks for the red-to-green evidence, not for
|
|
81
|
-
test existence alone: a test that was never watched failing doesn't clear
|
|
82
|
-
the bar even if one exists in the diff.
|
|
83
|
-
|
|
84
|
-
## Works with
|
|
85
|
-
|
|
86
|
-
- leo:verification — shares the red-to-green transition as evidence of a
|
|
87
|
-
real fix; don't duplicate the check, cite it.
|
|
88
|
-
- reviewer — enforces the coverage rubric line above on the actual diff.
|
|
89
|
-
- implementer, executor — the tiers that own writing the failing test and
|
|
90
|
-
then the fix.
|
|
@@ -1,96 +0,0 @@
|
|
|
1
|
-
---
|
|
2
|
-
name: using-leo
|
|
3
|
-
description: >
|
|
4
|
-
Leo's global operating policy: cost-tiered model routing, the
|
|
5
|
-
execute-then-review gate, delegation rules, orchestration triggers,
|
|
6
|
-
machine-local state, and the index of leo:* process skills. Injected
|
|
7
|
-
into every session by the harness bootstrap (with a per-harness mapping
|
|
8
|
-
appended) — it is context, not a skill to run.
|
|
9
|
-
---
|
|
10
|
-
|
|
11
|
-
# Leo's global agent directives
|
|
12
|
-
|
|
13
|
-
These apply in every session on every machine and every harness. Canonical copy: `skills/using-leo/SKILL.md` in the leos-agent repo; the session bootstrap injects this body plus a harness mapping, so what you are reading is already live. Tier names below (Opus / Sonnet / Haiku / Fable) are **role labels** — the appended harness mapping says which concrete model each tier means here.
|
|
14
|
-
|
|
15
|
-
## Model routing
|
|
16
|
-
|
|
17
|
-
Tier every task by the kind of work, not per session. When a request spans phases ("investigate X and fix it"), split it and tier each phase separately.
|
|
18
|
-
|
|
19
|
-
| Work type | Typical verbs | Tier | Do it via |
|
|
20
|
-
|---|---|---|---|
|
|
21
|
-
| Investigation | investigate, diagnose, debug, root-cause, "why does…" | Opus | the `investigator` role |
|
|
22
|
-
| Planning / design | plan, design, architect, decide | Opus | the `planner` role (or the harness's native plan flow at the Opus tier) |
|
|
23
|
-
| Implementation | implement, fix, build, refactor, execute | Sonnet | main loop if the session runs at the Sonnet tier, else the `implementer` role |
|
|
24
|
-
| Mechanical | rename, codemod, apply known pattern, boilerplate, format | Haiku | the `executor` role |
|
|
25
|
-
| Review / verification | review, verify, audit, judge | Opus | the `reviewer` role on the real diff |
|
|
26
|
-
| Hardest problems / arbitration | "use expert", "deep thinking", "deep investigate", Fable by name | Fable | the `expert` role |
|
|
27
|
-
|
|
28
|
-
Code location and structure-mapping that precedes any tiered work above goes to `explore` (Haiku tier, read-only) — cheap scouting that feeds the roles in the table; it returns file:line locations, never verdicts.
|
|
29
|
-
|
|
30
|
-
**Escalate, don't struggle**: if a cheap-tier task turns out ambiguous or fails twice, step up one tier rather than retrying at the same tier. When the right tier is unclear, default up — **capped at Opus**. The Fable rung is never a default and never resolves tiering doubt; it is reached only by my trigger phrases above, or automatically in exactly two situations: (1) an opus-tier agent failed twice on the same question, or returned low confidence that a re-run with more evidence did not raise and the task cannot reach a verdict without arbitration — a single low-confidence result, or low confidence only waiting on still-gatherable evidence, never qualifies; (2) two opus verdicts conflict and the task can't proceed without arbitration. Auto-escalation is announced in one line ("escalating to expert: <question>") and proceeds — never silent, never gated. On a harness with no Fable tier (see the mapping), escalation caps at Opus: stop and report to Leo instead, offering to continue at the Opus tier or hand off to a harness that has the expert rung.
|
|
31
|
-
|
|
32
|
-
## Execute means execute-then-review
|
|
33
|
-
|
|
34
|
-
Every implementation request — "fix", "implement", "execute the plan", anything that changes code — implicitly includes a review phase, whether or not review was mentioned. Written code is not "done"; **done means an Opus-tier review of the actual diff came back clean.**
|
|
35
|
-
|
|
36
|
-
1. Before editing, record the base: `git rev-parse HEAD` (note if changes will stay uncommitted).
|
|
37
|
-
2. Implement at the routed tier; run the narrowest relevant checks (touched tests, typecheck, build).
|
|
38
|
-
3. Have the `reviewer` role judge the actual diff, passing the base ref (or "uncommitted working tree") and the original request/plan text. Never self-review instead. Review runs at the Opus tier by default. Downscale to a Sonnet-tier review ONLY for a clearly-trivial diff — ALL of: ≤ 2 files, ≤ ~60 changed lines, mechanical/boilerplate class (rename, format, comment, constant/string tweak, dependency-version bump, test-data edit), and no risky-path match (auth, payments/billing, crypto/secrets, DB migration or schema, CI/CD config, access control). If any condition fails or you are unsure, keep the full Opus-tier review — the default bucket is today's behavior. Never skip review because the change "is small". Only exemptions (no review at all): docs/comment-only diffs, and edits Leo dictated verbatim — and "verbatim" means I gave you the literal text or the literal command, so claiming this exemption requires quoting what I said back in the done report. A paraphrase, an interpretation, or "this is what he meant" is not dictation and gets the normal review.
|
|
39
|
-
4. Blocking findings: fix at the executing tier, re-review the fix only. ONE cycle — if the second review still blocks, stop and report the findings to Leo instead of looping, offering `expert` arbitration as one of the options (where the Fable rung exists).
|
|
40
|
-
5. Report done as three lines: what changed / checks run / review verdict.
|
|
41
|
-
|
|
42
|
-
## Delegate the labor
|
|
43
|
-
|
|
44
|
-
The main loop orchestrates; delegated roles do the volume. In an expensive-tier session, inline bulk work burns the expensive tier — delegate down:
|
|
45
|
-
|
|
46
|
-
- Locating code, mapping structure → `explore` (Haiku tier), in parallel when questions are independent.
|
|
47
|
-
- Diagnosis needing a verdict → `investigator` (Opus tier) — ONE per question, fed by cheap exploration; distinct questions may run in parallel, but never fan the same question across multiple Opus-tier agents.
|
|
48
|
-
- Mechanical edits → `executor` (Haiku tier), fanned across independent items.
|
|
49
|
-
- Executing a written plan → `implementer` (Sonnet tier).
|
|
50
|
-
- Judging a diff → `reviewer` (Opus tier).
|
|
51
|
-
- Hardest verdicts and deadlocks → `expert` (Fable tier) — one at a time, never fanned out, never implements; hand it the outcome wanted, the raw artifact paths, and the full failure history (it reads sources itself — don't pre-digest for a stronger model).
|
|
52
|
-
|
|
53
|
-
Scale to complexity: simple lookup = 1 agent; comparing a few areas = 2–4 in parallel; large parallel workloads = orchestration triggers below. A fan-out costs roughly an order of magnitude more than a single chat — reserve it for genuinely parallel, high-value work. In an expensive-tier session this is a hard rule, not a heuristic: implementation and mechanical edits MUST go to `implementer`/`executor`, and code searches to `explore`; editing or grepping inline is the exception, reserved for a trivial single-file touch (< ~10 lines) where writing the spec would cost more than the change. More than ~3 inline file edits or ~5 inline searches in an expensive-tier session means the work should have been delegated. Dispatch mechanics — brief structure, the return-status contract, durable progress — live in leo:delegation.
|
|
54
|
-
|
|
55
|
-
## Agent teams
|
|
56
|
-
|
|
57
|
-
Where the harness offers persistent teammates rather than one-shot subagents, the topology changes but nothing above relaxes. Every rule in this policy binds a teammate exactly as it binds a dispatched role: execute-then-review still gates "done", one investigator per question still holds, `expert` is still never fanned out, and every teammate is still tier-pinned — an unpinned teammate inherits the session tier and quietly bills a judge's rate for an executor's work. Verdicts route through the main loop: peer messages coordinate, they never approve. Reach for a team only when a role must *accompany* the work — a reviewer watching an implementer's long run, an investigator unblocking it live. Batch-shaped work stays with fan-out or the workflow tool, which is cheaper and already has ledger-backed progress.
|
|
58
|
-
|
|
59
|
-
## Orchestration triggers
|
|
60
|
-
|
|
61
|
-
These phrases are my standing opt-in to multi-agent orchestration: **"fan this out"**, **"workflow this"**, **"grind on this"**, **"do this properly"**.
|
|
62
|
-
|
|
63
|
-
For a non-trivial task where I haven't used a trigger phrase, propose orchestration in one line (rough shape: agent count + model mix) and proceed single-agent unless I take the offer. Never launch a large fan-out silently. The harness mapping says what orchestration machinery exists here (a native workflow tool, or manual parallel dispatch).
|
|
64
|
-
|
|
65
|
-
## Machine-local state
|
|
66
|
-
|
|
67
|
-
Any skill or agent that needs to persist information writes JSON to `$LEOS_AGENT_LOCAL_PATH/<skill-or-agent-name>.json` — `LEOS_AGENT_LOCAL_PATH` is an optional override, unset it defaults to `~/.leos-agent-local` (in bash: `${LEOS_AGENT_LOCAL_PATH:-$HOME/.leos-agent-local}`). Top-level keys are `owner/repo` (or the absolute project path when there's no GitHub repo): **data always stays separate per repo/project**. Read and write through `python3 "${CLAUDE_PLUGIN_ROOT}/scripts/state.py"` (`get` / `merge` / `path`) instead of hand-rolling read-modify-write — the code ships with the plugin, the data stays under `${LEOS_AGENT_LOCAL_PATH:-$HOME/.leos-agent-local}/`, gitignored, per-machine, never synced, and survives plugin updates. Examples: `review-watcher.json` (PR numbers already auto-reviewed), `resolve-ticket.json` (ticket-prefix → tracker mappings).
|
|
68
|
-
|
|
69
|
-
Durable facts are a different thing and do not belong in those JSON files: a preference, a repo rule the code never states, a settled decision goes to the memory store at `$LEOS_AGENT_LOCAL_PATH/memory/`, one fact per file, through `python3 "${CLAUDE_PLUGIN_ROOT}/scripts/memory.py"` and leo:memory. That store is canonical and each harness's own memory surface receives a generated copy of the global facts, so a preference learned on one harness is in front of me on the next. Its index is appended below when the store is not empty.
|
|
70
|
-
|
|
71
|
-
## Cost discipline
|
|
72
|
-
|
|
73
|
-
Spend expensive tokens on planning, verification, and synthesis (low volume, high leverage); spend cheap tokens on execution volume. When dispatching delegated work, pin the tier per task — the `executor` role runs at the Haiku tier for mechanical and boilerplate work and at the Sonnet tier at low effort for ordinary implementation, judges/verifiers at the Opus tier. The Fable tier is the most expensive per call and cheap as a policy only because it fires rarely and only on verdicts — batch fan-outs never auto-use it (that is exactly where a Fable jump silently multiplies cost).
|
|
74
|
-
|
|
75
|
-
## Skill index
|
|
76
|
-
|
|
77
|
-
Reach for the matching skill at the decision point — each one encodes the mechanics these directives already assume, sized to the work, not extra ritual.
|
|
78
|
-
|
|
79
|
-
| At this point | Consult |
|
|
80
|
-
|---|---|
|
|
81
|
-
| A bug or failing test, before any fix | leo:debugging |
|
|
82
|
-
| An approach not yet settled, before non-trivial code | leo:brainstorming |
|
|
83
|
-
| Turning a chosen approach into a plan | leo:writing-plans |
|
|
84
|
-
| Carrying out a written plan | leo:executing-plans |
|
|
85
|
-
| Adding or changing runtime behavior | leo:test-first |
|
|
86
|
-
| Coding against a third-party API | leo:freshness |
|
|
87
|
-
| Before claiming anything done / fixed / passing | leo:verification |
|
|
88
|
-
| A UI-visible change, before done | leo:visual-verification |
|
|
89
|
-
| Dispatching subagents or a fan-out | leo:delegation |
|
|
90
|
-
| Isolating branch work | leo:worktrees |
|
|
91
|
-
| Landing or cleaning up a finished branch | leo:finishing-a-branch |
|
|
92
|
-
| A durable fact surfaces, or one turns out wrong | leo:memory |
|
|
93
|
-
| Policy or harness wiring in doubt | leo:doctor |
|
|
94
|
-
| Authoring or editing a skill | leo:writing-skills |
|
|
95
|
-
|
|
96
|
-
Four operational skills — `leo:resolve-ticket`, `leo:review-pr`, `leo:watch-review`, `leo:setup` — are invoked by name rather than reached from the table above. One more, `leo:attach-pr`, is Claude Code only and is not registered on any other harness (the harness mapping appended below says so explicitly, and names what else differs here).
|
|
@@ -1,32 +0,0 @@
|
|
|
1
|
-
<!-- Generated by scripts/render_adapters.py; do not edit. -->
|
|
2
|
-
# Claude Code mapping
|
|
3
|
-
|
|
4
|
-
| Tier | Model | Effort |
|
|
5
|
-
|---|---|---|
|
|
6
|
-
| Fable | `fable` | max |
|
|
7
|
-
| Opus | `opus` | native default |
|
|
8
|
-
| Sonnet | `sonnet` | native default |
|
|
9
|
-
| Haiku | `haiku` | native default |
|
|
10
|
-
|
|
11
|
-
## Capabilities here
|
|
12
|
-
|
|
13
|
-
| Capability | Here |
|
|
14
|
-
|---|---|
|
|
15
|
-
| Policy injection | `SessionStart` hook, on every startup / resume / clear / compact |
|
|
16
|
-
| Subagent spawn | spawn the named native agent; its generated frontmatter pins the model |
|
|
17
|
-
| Per-spawn model | yes — the agent's own frontmatter |
|
|
18
|
-
| Read-only roles | harness-enforced — the tool allowlist omits Write and Edit |
|
|
19
|
-
| Worktrees | `EnterWorktree` / `ExitWorktree`, session-tracked and auto-cleaned; pair every Enter with an Exit |
|
|
20
|
-
| Workflow runner | the Workflow tool runs `workflows/cost-tiered-fix.js` by `scriptPath` |
|
|
21
|
-
| Follow-up to a live agent | `SendMessage` to the same agent, which keeps the context it already built |
|
|
22
|
-
| Skill names | `leo:<name>` |
|
|
23
|
-
|
|
24
|
-
Visual evidence here: the Browser pane (start or attach a preview, then take a screenshot), an attached Chrome, or the iOS Simulator control tool; some arrive only after a tool search, so an empty tool list is not proof of absence. When no rung answers, leo:visual-verification requires the unverified-change warning in place of a done report.
|
|
25
|
-
|
|
26
|
-
Memory projection here writes to the per-user `CLAUDE.md` in the Claude config directory. Only global-scope facts are projected — every per-user surface loads in every repository, so repo facts would leak across projects; they reach the model through the session context block instead. Leo's block is marker-delimited; the rest of the file is untouched.
|
|
27
|
-
|
|
28
|
-
## Leo skills only available here
|
|
29
|
-
|
|
30
|
-
- `leo:attach-pr` — its entire product is a side effect in Claude Code Desktop's PR-card detector, which no other harness has — the same commands would run here, succeed, and produce nothing observable.
|
|
31
|
-
|
|
32
|
-
Every other skill in the policy's Skill index is registered on every harness and behaves the same. These are not, so a procedure that leans on one of them does not transfer.
|
|
@@ -1,34 +0,0 @@
|
|
|
1
|
-
<!-- Generated by scripts/render_adapters.py; do not edit. -->
|
|
2
|
-
# Codex mapping
|
|
3
|
-
|
|
4
|
-
| Tier | Model | Effort |
|
|
5
|
-
|---|---|---|
|
|
6
|
-
| Fable | `gpt-5.6-sol` | max |
|
|
7
|
-
| Opus | `gpt-5.6-sol` | high |
|
|
8
|
-
| Sonnet | `gpt-5.6-terra` | medium |
|
|
9
|
-
| Haiku | `gpt-5.6-luna` | low |
|
|
10
|
-
|
|
11
|
-
## Capabilities here
|
|
12
|
-
|
|
13
|
-
| Capability | Here |
|
|
14
|
-
|---|---|
|
|
15
|
-
| Policy injection | `SessionStart` hook, on every startup / resume / clear / compact |
|
|
16
|
-
| Subagent spawn | generic subagent with `roles/<role>.md` pasted in |
|
|
17
|
-
| Per-spawn model | yes — pass `model` and `reasoning_effort` explicitly; a user or `AGENTS.md` override still wins |
|
|
18
|
-
| Read-only roles | prompt only — a convention, never a guarantee; never route work here that depends on it |
|
|
19
|
-
| Worktrees | no native tool — raw `git worktree` at `.claude/worktrees/<name>` |
|
|
20
|
-
| Workflow runner | none — fan out by hand and keep the ledger in `<plugin-root>/scripts/state.py` |
|
|
21
|
-
| Follow-up to a live agent | none established — re-dispatch cold with the context restated |
|
|
22
|
-
| Skill names | `leo:<name>` |
|
|
23
|
-
|
|
24
|
-
Visual evidence here: the bundled browser plugin, else computer-use, else Playwright driven from the shell. When no rung answers, leo:visual-verification requires the unverified-change warning in place of a done report.
|
|
25
|
-
|
|
26
|
-
Memory projection here writes to the per-user `AGENTS.md` in the Codex home directory. Only global-scope facts are projected — every per-user surface loads in every repository, so repo facts would leak across projects; they reach the model through the session context block instead. Leo's block is marker-delimited; the rest of the file is untouched.
|
|
27
|
-
|
|
28
|
-
Tier collapse here: Fable≡Opus (`gpt-5.6-sol`) — routing between collapsed rungs buys role, not power. Fable is not a real rung: `expert` cannot break a deadlock a collapsed Opus already lost, so cap escalation at Opus and report.
|
|
29
|
-
|
|
30
|
-
## Leo skills not available here
|
|
31
|
-
|
|
32
|
-
- `leo:attach-pr` — its entire product is a side effect in Claude Code Desktop's PR-card detector, which no other harness has — the same commands would run here, succeed, and produce nothing observable.
|
|
33
|
-
|
|
34
|
-
Every other skill in the policy's Skill index is registered here and behaves the same, and so are the operational skills — `leo:review-pr`, `leo:resolve-ticket` and `leo:watch-review` all run on this harness. Where they name a capability the table above says is missing, take the fallback each one documents.
|
|
@@ -1,34 +0,0 @@
|
|
|
1
|
-
<!-- Generated by scripts/render_adapters.py; do not edit. -->
|
|
2
|
-
# Cursor mapping
|
|
3
|
-
|
|
4
|
-
| Tier | Model | Effort |
|
|
5
|
-
|---|---|---|
|
|
6
|
-
| Fable | `GPT-5.6 Sol` | native default |
|
|
7
|
-
| Opus | `Grok 4.5` | native default |
|
|
8
|
-
| Sonnet | `Grok 4.5` | native default |
|
|
9
|
-
| Haiku | `Composer 2.5` | native default |
|
|
10
|
-
|
|
11
|
-
## Capabilities here
|
|
12
|
-
|
|
13
|
-
| Capability | Here |
|
|
14
|
-
|---|---|
|
|
15
|
-
| Policy injection | `sessionStart` hook, every session |
|
|
16
|
-
| Subagent spawn | plugin agent from the generated Cursor agents directory |
|
|
17
|
-
| Per-spawn model | no — agents are `model: inherit`; select the tier's model in the UI before a homogeneous batch |
|
|
18
|
-
| Read-only roles | harness-enforced — generated `readonly: true` |
|
|
19
|
-
| Worktrees | no native tool — raw `git worktree` at `.claude/worktrees/<name>` |
|
|
20
|
-
| Workflow runner | none — fan out by hand and keep the ledger in `<plugin-root>/scripts/state.py` |
|
|
21
|
-
| Follow-up to a live agent | none established — re-dispatch cold with the context restated |
|
|
22
|
-
| Skill names | `leo:<name>` |
|
|
23
|
-
|
|
24
|
-
Visual evidence here: Browser Preview against a running dev server, else a Playwright server if one is registered. When no rung answers, leo:visual-verification requires the unverified-change warning in place of a done report.
|
|
25
|
-
|
|
26
|
-
Memory projection here writes to a generated rules file in the per-user Cursor rules directory. Only global-scope facts are projected — every per-user surface loads in every repository, so repo facts would leak across projects; they reach the model through the session context block instead. Leo's block is marker-delimited; the rest of the file is untouched.
|
|
27
|
-
|
|
28
|
-
Tier collapse here: Opus≡Sonnet (`Grok 4.5`) — routing between collapsed rungs buys role, not power.
|
|
29
|
-
|
|
30
|
-
## Leo skills not available here
|
|
31
|
-
|
|
32
|
-
- `leo:attach-pr` — its entire product is a side effect in Claude Code Desktop's PR-card detector, which no other harness has — the same commands would run here, succeed, and produce nothing observable.
|
|
33
|
-
|
|
34
|
-
Every other skill in the policy's Skill index is registered here and behaves the same, and so are the operational skills — `leo:review-pr`, `leo:resolve-ticket` and `leo:watch-review` all run on this harness. Where they name a capability the table above says is missing, take the fallback each one documents.
|
|
@@ -1,36 +0,0 @@
|
|
|
1
|
-
<!-- Generated by scripts/render_adapters.py; do not edit. -->
|
|
2
|
-
# Hermes mapping
|
|
3
|
-
|
|
4
|
-
Provider: `openrouter`
|
|
5
|
-
|
|
6
|
-
| Tier | Model | Effort |
|
|
7
|
-
|---|---|---|
|
|
8
|
-
| Fable | `moonshotai/kimi-k3` | native default |
|
|
9
|
-
| Opus | `moonshotai/kimi-k3` | native default |
|
|
10
|
-
| Sonnet | `z-ai/glm-5.2` | native default |
|
|
11
|
-
| Haiku | `z-ai/glm-5.2` | native default |
|
|
12
|
-
|
|
13
|
-
## Capabilities here
|
|
14
|
-
|
|
15
|
-
| Capability | Here |
|
|
16
|
-
|---|---|
|
|
17
|
-
| Policy injection | rides the session's first tool result — so a session that runs no tool gets none; read `leo:using-leo` if the policy is not already in your context |
|
|
18
|
-
| Subagent spawn | native `delegate_task`, canonical role prompt pasted in |
|
|
19
|
-
| Per-spawn model | no — one `delegation.model` for every child, so batch homogeneous Kimi or GLM work and switch it between batches |
|
|
20
|
-
| Read-only roles | prompt only — a convention, never a guarantee; never route work here that depends on it |
|
|
21
|
-
| Worktrees | no native tool — raw `git worktree` at `.claude/worktrees/<name>` |
|
|
22
|
-
| Workflow runner | none — fan out by hand and keep the ledger in `<plugin-root>/scripts/state.py` |
|
|
23
|
-
| Follow-up to a live agent | none established — re-dispatch cold with the context restated |
|
|
24
|
-
| Skill names | `leo:<name>` |
|
|
25
|
-
|
|
26
|
-
Visual evidence here: no built-in renderer; Playwright driven from the shell is the only rung, and only when the project already depends on it. When no rung answers, leo:visual-verification requires the unverified-change warning in place of a done report.
|
|
27
|
-
|
|
28
|
-
Memory projection here writes to `SOUL.md` in the Hermes home, and only once `leo:setup enable hermes-memory` turns it on — that file is the agent's own identity prompt, so it is never written to unasked and never created. Only global-scope facts are projected — every per-user surface loads in every repository, so repo facts would leak across projects; they reach the model through the session context block instead. Leo's block is marker-delimited; the rest of the file is untouched.
|
|
29
|
-
|
|
30
|
-
Tier collapse here: Fable≡Opus (`moonshotai/kimi-k3`), Sonnet≡Haiku (`z-ai/glm-5.2`) — routing between collapsed rungs buys role, not power. Fable is not a real rung: `expert` cannot break a deadlock a collapsed Opus already lost, so cap escalation at Opus and report.
|
|
31
|
-
|
|
32
|
-
## Leo skills not available here
|
|
33
|
-
|
|
34
|
-
- `leo:attach-pr` — its entire product is a side effect in Claude Code Desktop's PR-card detector, which no other harness has — the same commands would run here, succeed, and produce nothing observable.
|
|
35
|
-
|
|
36
|
-
Every other skill in the policy's Skill index is registered here and behaves the same, and so are the operational skills — `leo:review-pr`, `leo:resolve-ticket` and `leo:watch-review` all run on this harness. Where they name a capability the table above says is missing, take the fallback each one documents.
|
|
@@ -1,36 +0,0 @@
|
|
|
1
|
-
<!-- Generated by scripts/render_adapters.py; do not edit. -->
|
|
2
|
-
# OpenCode mapping
|
|
3
|
-
|
|
4
|
-
Provider: `openrouter`
|
|
5
|
-
|
|
6
|
-
| Tier | Model | Effort |
|
|
7
|
-
|---|---|---|
|
|
8
|
-
| Fable | `moonshotai/kimi-k3` | native default |
|
|
9
|
-
| Opus | `moonshotai/kimi-k3` | native default |
|
|
10
|
-
| Sonnet | `z-ai/glm-5.2` | native default |
|
|
11
|
-
| Haiku | `z-ai/glm-5.2` | native default |
|
|
12
|
-
|
|
13
|
-
## Capabilities here
|
|
14
|
-
|
|
15
|
-
| Capability | Here |
|
|
16
|
-
|---|---|
|
|
17
|
-
| Policy injection | `config.instructions`, with a system-prompt transform as backstop |
|
|
18
|
-
| Subagent spawn | registered agent from `agents.json`, spawned via the task tool |
|
|
19
|
-
| Per-spawn model | no — each agent always runs its registered model, so `reviewer` never downscales on a trivial diff |
|
|
20
|
-
| Read-only roles | harness-enforced — generated `permission.edit: deny`, refused by OpenCode itself |
|
|
21
|
-
| Worktrees | no native tool — raw `git worktree` at `.claude/worktrees/<name>` |
|
|
22
|
-
| Workflow runner | no runner — `cost-tiered-fix.js` ships in the package but nothing here executes it; fan out by hand and keep the ledger in `<plugin-root>/scripts/state.py` |
|
|
23
|
-
| Follow-up to a live agent | none established — re-dispatch cold with the context restated |
|
|
24
|
-
| Skill names | bare `<name>` — read every `leo:<x>` above as `<x>` |
|
|
25
|
-
|
|
26
|
-
Visual evidence here: no built-in renderer; a registered Playwright server or the Playwright CLI. When no rung answers, leo:visual-verification requires the unverified-change warning in place of a done report.
|
|
27
|
-
|
|
28
|
-
Memory projection here writes to the per-user `AGENTS.md` in the OpenCode config directory. Only global-scope facts are projected — every per-user surface loads in every repository, so repo facts would leak across projects; they reach the model through the session context block instead. Leo's block is marker-delimited; the rest of the file is untouched.
|
|
29
|
-
|
|
30
|
-
Tier collapse here: Fable≡Opus (`moonshotai/kimi-k3`), Sonnet≡Haiku (`z-ai/glm-5.2`) — routing between collapsed rungs buys role, not power. Fable is not a real rung: `expert` cannot break a deadlock a collapsed Opus already lost, so cap escalation at Opus and report.
|
|
31
|
-
|
|
32
|
-
## Leo skills not available here
|
|
33
|
-
|
|
34
|
-
- `leo:attach-pr` — its entire product is a side effect in Claude Code Desktop's PR-card detector, which no other harness has — the same commands would run here, succeed, and produce nothing observable.
|
|
35
|
-
|
|
36
|
-
Every other skill in the policy's Skill index is registered here and behaves the same, and so are the operational skills — `leo:review-pr`, `leo:resolve-ticket` and `leo:watch-review` all run on this harness. Where they name a capability the table above says is missing, take the fallback each one documents.
|
|
@@ -1,109 +0,0 @@
|
|
|
1
|
-
---
|
|
2
|
-
name: verification
|
|
3
|
-
description: >
|
|
4
|
-
Fresh-evidence gate before claiming done, fixed, or passing. Applies to
|
|
5
|
-
the main loop, implementer, executor, and anyone reporting completion:
|
|
6
|
-
a completion claim needs a proving command run in the current turn, whose
|
|
7
|
-
output was actually read — never a prior run, a "should pass now," or a
|
|
8
|
-
subagent's self-report relayed as fact.
|
|
9
|
-
when_to_use: >
|
|
10
|
-
Before writing any completion claim — "tests pass," "build is green,"
|
|
11
|
-
"bug fixed," "agent finished the task." Also applies when relaying a
|
|
12
|
-
subagent's own "done"/"success" report up the chain. NOT for routine
|
|
13
|
-
in-progress status updates that make no completion claim, and NOT a
|
|
14
|
-
substitute for the review phase itself — this gates the evidence behind
|
|
15
|
-
it, execute-then-review still owns the verdict.
|
|
16
|
-
---
|
|
17
|
-
|
|
18
|
-
# verification
|
|
19
|
-
|
|
20
|
-
Core rule: no completion claim without a proving command run fresh, this
|
|
21
|
-
turn, with output actually read. A claim resting on memory, a prior run, or
|
|
22
|
-
someone else's word is not verification — it is a guess wearing the shape
|
|
23
|
-
of one.
|
|
24
|
-
|
|
25
|
-
## When it fires
|
|
26
|
-
|
|
27
|
-
Any sentence of the form "X passes," "X is fixed," "X works now," "agent Y
|
|
28
|
-
finished." That sentence is a claim. A claim needs proof, and proof has an
|
|
29
|
-
expiration: the moment code changes again, prior proof is stale.
|
|
30
|
-
|
|
31
|
-
Does not fire for: status updates that don't assert completion ("still
|
|
32
|
-
running," "found the bug, fixing now"), or work that genuinely has no
|
|
33
|
-
runtime surface (docs/comment-only diffs — see the execute-then-review
|
|
34
|
-
exemptions).
|
|
35
|
-
|
|
36
|
-
## The discipline
|
|
37
|
-
|
|
38
|
-
1. **Name the command that would falsify the claim.** Not "does it look
|
|
39
|
-
right" — the specific test/build/repro that fails if the claim is false.
|
|
40
|
-
If no such command exists, the claim isn't verifiable yet; say so instead
|
|
41
|
-
of asserting it.
|
|
42
|
-
2. **Run it fresh, this turn.** A green run from three edits ago proves
|
|
43
|
-
nothing about the code as it stands now.
|
|
44
|
-
3. **Read the exit status and failure counts.** Not the last line of
|
|
45
|
-
scrollback, not a summary a subagent wrote — the actual output.
|
|
46
|
-
4. **Only then claim, and state the evidence** — the command and what it
|
|
47
|
-
returned, not just "verified."
|
|
48
|
-
|
|
49
|
-
Skipping step 1 is how "should be fine" sneaks in. Skipping step 2 is how a
|
|
50
|
-
stale green run gets relayed as current. Skipping step 3 is how a nonzero
|
|
51
|
-
exit gets read as success because the output scrolled by fast.
|
|
52
|
-
|
|
53
|
-
## Claim → proof
|
|
54
|
-
|
|
55
|
-
| Claim | Falsifying command |
|
|
56
|
-
|---|---|
|
|
57
|
-
| Tests pass | the actual test command, run now, exit code + failure count read |
|
|
58
|
-
| Build is green | the build command, run now, read for errors/warnings |
|
|
59
|
-
| Bug is fixed | the reproducer that showed the bug — now green, run fresh |
|
|
60
|
-
| Agent reports done | its diff and output, inspected directly — its "success" is a claim, not evidence |
|
|
61
|
-
| This library call is correct | the installed package or current docs, read this session — see leo:freshness |
|
|
62
|
-
| A UI change looks right | a render produced after the edit, looked at — see leo:visual-verification |
|
|
63
|
-
|
|
64
|
-
## Subagent reports are claims, not evidence
|
|
65
|
-
|
|
66
|
-
A subagent saying "success," "all tests pass," or "implemented as
|
|
67
|
-
requested" is exactly as unverified as your own untested assertion would be
|
|
68
|
-
— it is one more claim to check against artifacts. Inspect the diff it
|
|
69
|
-
produced. Run the check it says it ran. If it reports a test command, that
|
|
70
|
-
command's output belongs in your evidence, not its summary of the output.
|
|
71
|
-
Relaying a subagent's self-report upward without this check just moves the
|
|
72
|
-
gap in provenance one level up the chain.
|
|
73
|
-
|
|
74
|
-
## Done is the three-line report
|
|
75
|
-
|
|
76
|
-
Writing code is not done. Done is the execute-then-review report:
|
|
77
|
-
|
|
78
|
-
- **what changed** — the diff, in one line
|
|
79
|
-
- **checks run** — the fresh commands from this gate, with results
|
|
80
|
-
- **review verdict** — clean per the execute-then-review policy, not
|
|
81
|
-
self-assessed
|
|
82
|
-
|
|
83
|
-
Verification is the evidence behind lines two and three. A report with line
|
|
84
|
-
one but not two and three is a status update, not a completion claim — label
|
|
85
|
-
it as such.
|
|
86
|
-
|
|
87
|
-
## Self-talk to catch
|
|
88
|
-
|
|
89
|
-
- "It should pass now, I fixed the obvious thing" — run it.
|
|
90
|
-
- "The tests were green before this edit" — before this edit is not now.
|
|
91
|
-
- "The subagent said it's done" — done according to whom, checked how.
|
|
92
|
-
- "I read the code and it looks correct" — reading is not running.
|
|
93
|
-
- "Re-running is wasteful, nothing changed" — if nothing changed, the prior
|
|
94
|
-
run is fine to cite as fresh evidence; if anything did, it isn't.
|
|
95
|
-
|
|
96
|
-
## Works with
|
|
97
|
-
|
|
98
|
-
- The execute-then-review policy (injected leo:using-leo) — that gate is the
|
|
99
|
-
outer loop this skill feeds into; the reviewer subagent judges the diff,
|
|
100
|
-
this skill governs the evidence claimed leading up to that judgment.
|
|
101
|
-
- reviewer — its verdict is itself a claim to relay accurately, not to
|
|
102
|
-
soften or summarize away.
|
|
103
|
-
- End-to-end exercise — when the falsifying command is "does the real flow
|
|
104
|
-
work," drive the actual app or flow, not just the test suite.
|
|
105
|
-
- leo:freshness — this gate proves the code you wrote runs; that one governs
|
|
106
|
-
whether the third-party API you wrote it against actually exists. A green
|
|
107
|
-
test against a mocked dependency clears this skill and not that one.
|
|
108
|
-
- leo:visual-verification — for a change someone sees, the falsifying artifact
|
|
109
|
-
is a render, not an exit status. That skill owns what the render must show.
|
|
@@ -1,114 +0,0 @@
|
|
|
1
|
-
---
|
|
2
|
-
name: visual-verification
|
|
3
|
-
description: >
|
|
4
|
-
Render-evidence gate for changes a person can see. A UI-visible edit is not
|
|
5
|
-
reported done until a render produced after the edit has been looked at.
|
|
6
|
-
Detection walks a ranked ladder of whatever browser, preview, or simulator
|
|
7
|
-
tooling this harness exposes; when nothing on the ladder answers, the change
|
|
8
|
-
is reported with an explicit unverified warning instead of a completion
|
|
9
|
-
claim.
|
|
10
|
-
when_to_use: >
|
|
11
|
-
A change whose result someone would notice by looking — layout, styling,
|
|
12
|
-
on-screen text, a new view or route, a chart, a generated image or rendered
|
|
13
|
-
document. Fires just before the completion claim, beside leo:verification.
|
|
14
|
-
NOT for logic with no rendered surface, NOT for a component behind a flag
|
|
15
|
-
that is off, and NOT a replacement for tests — a render shows one state, a
|
|
16
|
-
test covers the branch.
|
|
17
|
-
---
|
|
18
|
-
|
|
19
|
-
# visual-verification
|
|
20
|
-
|
|
21
|
-
A pixel claim needs a pixel. Reporting that a visible change works, having
|
|
22
|
-
never rendered it, is an assertion about something nobody looked at — and the
|
|
23
|
-
whole suite can be green while the element sits clipped, transparent, or
|
|
24
|
-
underneath its own container.
|
|
25
|
-
|
|
26
|
-
This is the complement of the test gate. leo:test-first exempts pure copy and
|
|
27
|
-
styling tweaks precisely because a test would say nothing useful about them;
|
|
28
|
-
this skill is what catches them instead. What one gate waves through, the other
|
|
29
|
-
holds.
|
|
30
|
-
|
|
31
|
-
## When it fires
|
|
32
|
-
|
|
33
|
-
Rendered layout or styling; on-screen text; a new or changed view, route, or
|
|
34
|
-
component; a chart, canvas, or generated image; a rendered document artifact;
|
|
35
|
-
a state someone reaches by clicking.
|
|
36
|
-
|
|
37
|
-
It does not fire for data-layer changes, logging, build configuration, or a
|
|
38
|
-
component behind a disabled flag.
|
|
39
|
-
|
|
40
|
-
## The detection ladder
|
|
41
|
-
|
|
42
|
-
Walk in order, stop at the first rung that answers. A rung that exists but
|
|
43
|
-
errors or returns nothing counts as absent for this purpose.
|
|
44
|
-
|
|
45
|
-
1. **A harness-native preview or browser pane** — starts or attaches to the
|
|
46
|
-
project's own dev server and screenshots the running app. Highest fidelity:
|
|
47
|
-
it renders the real build.
|
|
48
|
-
2. **A harness-native attached browser** — drives an already-running browser at
|
|
49
|
-
a URL. Right when the app is deployed or served outside this session.
|
|
50
|
-
3. **A platform simulator** — for native UI no browser can show.
|
|
51
|
-
4. **A scriptable driver through the shell** — Playwright or Puppeteer, or an
|
|
52
|
-
existing end-to-end test that captures a screenshot. Check the lockfile
|
|
53
|
-
before concluding the project does not have one.
|
|
54
|
-
5. **A rendering assertion the project already owns** — a snapshot or visual
|
|
55
|
-
regression suite. Weaker than a render you looked at, but it is evidence
|
|
56
|
-
produced after the edit. Name which one you used.
|
|
57
|
-
|
|
58
|
-
**The ladder is about capability, not brand names.** A harness that renames its
|
|
59
|
-
browser tool still has rung 1. Where a harness defers part of its tool
|
|
60
|
-
inventory until it is searched, an empty tool list is not evidence of absence —
|
|
61
|
-
search first, then conclude. That distinction is the most likely way this gate
|
|
62
|
-
degrades to a warning when something was in fact available.
|
|
63
|
-
|
|
64
|
-
The mapping appended to the session policy names which rungs exist here.
|
|
65
|
-
|
|
66
|
-
## What counts as evidence
|
|
67
|
-
|
|
68
|
-
The render is produced **after** the edit, this turn, and is actually looked
|
|
69
|
-
at. Name what you checked in it — the element, where it sits, what state it is
|
|
70
|
-
in. A screenshot captured is not a screenshot read.
|
|
71
|
-
|
|
72
|
-
## When nothing answers
|
|
73
|
-
|
|
74
|
-
There is no exemption list here. A UI-visible change either carries a render or
|
|
75
|
-
carries this block. Emit it **instead of** the word done:
|
|
76
|
-
|
|
77
|
-
```
|
|
78
|
-
UNVERIFIED UI CHANGE — no render tool answered on this harness.
|
|
79
|
-
|
|
80
|
-
Changed: <the visible change, one line>
|
|
81
|
-
Expected: <what should look different, and where>
|
|
82
|
-
Probed: <the rungs tried, by name, in order>
|
|
83
|
-
Verify by: <the one concrete thing Leo can do — a URL, a command, a screen>
|
|
84
|
-
```
|
|
85
|
-
|
|
86
|
-
The completion line then reads "implemented, unverified" — never "done". A
|
|
87
|
-
`Probed:` line that names nothing means the ladder was skipped, not that it came
|
|
88
|
-
up empty. Never suppress the block because the change looks obviously correct
|
|
89
|
-
or the diff was one line of CSS.
|
|
90
|
-
|
|
91
|
-
## Self-talk to catch
|
|
92
|
-
|
|
93
|
-
- "It's one line of CSS" — one line of CSS is what collapses a flex container.
|
|
94
|
-
- "The component tests pass" — tests assert a tree; an element can be present
|
|
95
|
-
and invisible.
|
|
96
|
-
- "There's no browser tool here" — did you probe, or read a tool list that
|
|
97
|
-
hides half its inventory until asked?
|
|
98
|
-
- "I'll mention it wasn't verified in passing" — in passing is how it gets read
|
|
99
|
-
as done. Use the block.
|
|
100
|
-
- "I rendered it earlier" — then you have a picture of the previous version.
|
|
101
|
-
|
|
102
|
-
## Reviewable finding
|
|
103
|
-
|
|
104
|
-
A UI-visible diff reported done with neither render evidence nor the warning
|
|
105
|
-
block is a blocking finding.
|
|
106
|
-
|
|
107
|
-
## Works with
|
|
108
|
-
|
|
109
|
-
- leo:verification — the same rule about evidence being fresh, applied to a
|
|
110
|
-
render rather than an exit status. That gate owns the completion claim; this
|
|
111
|
-
one owns what a visible claim needs behind it.
|
|
112
|
-
- leo:test-first — its copy-and-styling exemption is this skill's inbox. A
|
|
113
|
-
snapshot test is rung 5 and satisfies both, once.
|
|
114
|
-
- reviewer — the warning block is an artifact to judge, not prose to skim past.
|