@lorekit/cli 1.69.0 → 1.70.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/package.json +1 -1
- package/skill/lorekit-groom/rules/grooming-pass.md +34 -0
- package/skill/lorekit-setup/SKILL.md +141 -104
- package/skill/lorekit-setup/rules/cold-start-seeding.md +102 -0
- package/skill/lorekit-setup/rules/loop-health.md +117 -0
- package/skill/lorekit-setup/rules/proving-improvement.md +111 -0
- package/skill/lorekit-setup/rules/self-improvement-loops.md +52 -4
- package/skill/lorekit-setup/rules/team-and-portfolio.md +107 -0
- package/skill/lorekit-setup/templates/README.md +26 -0
- package/skill/lorekit-setup/templates/ci-job.md +90 -0
- package/skill/lorekit-setup/templates/code-changing-agent.md +89 -0
- package/skill/lorekit-setup/templates/multi-step-orchestrator.md +76 -0
- package/skill/lorekit-setup/templates/reviewer-reconcile-host.md +86 -0
- package/src/commands/lint.mjs +4 -2
- package/src/shared/lessons-view.mjs +71 -0
|
@@ -0,0 +1,111 @@
|
|
|
1
|
+
# Proving a loop improved something
|
|
2
|
+
|
|
3
|
+
A loop that nobody can see helping is indistinguishable from no loop. Wiring is not
|
|
4
|
+
the deliverable — **a real failure that stops recurring** is. This file is how you show
|
|
5
|
+
that, and — as important — how you avoid fooling yourself with numbers that cannot prove
|
|
6
|
+
it.
|
|
7
|
+
|
|
8
|
+
## Contents
|
|
9
|
+
|
|
10
|
+
- [The required proof: the immunity re-challenge](#the-required-proof-the-immunity-re-challenge)
|
|
11
|
+
- [Why the raw counters cannot prove causation](#why-the-raw-counters-cannot-prove-causation)
|
|
12
|
+
- [The optional counter view (advisory only)](#the-optional-counter-view-advisory-only)
|
|
13
|
+
- [Declare the metric at wiring time](#declare-the-metric-at-wiring-time)
|
|
14
|
+
- [The north star: a withheld-memory control arm](#the-north-star-a-withheld-memory-control-arm)
|
|
15
|
+
- [Guards](#guards)
|
|
16
|
+
|
|
17
|
+
---
|
|
18
|
+
|
|
19
|
+
## The required proof: the immunity re-challenge
|
|
20
|
+
|
|
21
|
+
There is exactly one proof this skill *requires*, because it is the one a doc can make
|
|
22
|
+
a person actually run and the one that survives the confounds below: **after you promote
|
|
23
|
+
a lesson to a rule, its failure signature should stop recurring.**
|
|
24
|
+
|
|
25
|
+
Borrowed from immune memory: promoting a lesson registers its original failure
|
|
26
|
+
signature as a "pathogen"; a later write bearing that same signature is a re-infection.
|
|
27
|
+
|
|
28
|
+
1. **At promotion, record the signature.** The lesson's **Applies when** line *is* the
|
|
29
|
+
signature — a file glob, a task type, a tool name, an error shape. Note the source
|
|
30
|
+
lesson's key and its `seen_count` at the moment of promotion.
|
|
31
|
+
2. **Watch for recurrence.** The direct measurement is the source lesson's `seen_count`
|
|
32
|
+
going **flat** — the loop stops writing that lesson because the failure stopped
|
|
33
|
+
happening. Read the column (`memory.read` / `lorekit show`), never a number in the
|
|
34
|
+
body.
|
|
35
|
+
3. **Escalate a breakthrough.** Any new write matching the promoted signature is a
|
|
36
|
+
vaccine-breakthrough case: the rule did not hold. That is a signal to fix the rule
|
|
37
|
+
(too narrow, wrong shape), not to shrug.
|
|
38
|
+
|
|
39
|
+
A promoted rule's value is the **silence** of its failure signature afterward. You
|
|
40
|
+
measure the absence. This needs no new product feature — `seen_count` and the
|
|
41
|
+
**Applies when** line already exist.
|
|
42
|
+
|
|
43
|
+
## Why the raw counters cannot prove causation
|
|
44
|
+
|
|
45
|
+
It is tempting to prove a loop "works" by pointing at rising counters. Do not lead with
|
|
46
|
+
this — LoreKit's own analysis records why the absolute counters are confounded:
|
|
47
|
+
|
|
48
|
+
- **`read_count` is ~99.8% bulk ride-alongs** — it tracks scope breadth, not whether a
|
|
49
|
+
lesson mattered.
|
|
50
|
+
- **`seen_count` is `1` for ~88% of rows** — most lessons never recur, so a raw count is
|
|
51
|
+
mostly noise.
|
|
52
|
+
- **Pull-through (`opened_count / read_count`) is the only trustworthy value signal** —
|
|
53
|
+
it cancels the ride-along confound by dividing it out.
|
|
54
|
+
|
|
55
|
+
And even pull-through is *correlation*: a lesson can pull through because it was
|
|
56
|
+
genuinely useful, or because the task was easy and any nudge looked right. Counter
|
|
57
|
+
movement alone can never separate "the lesson helped" from "the run would have
|
|
58
|
+
succeeded anyway." Only a run that deliberately did **not** get the memory can — see
|
|
59
|
+
[the north star](#the-north-star-a-withheld-memory-control-arm).
|
|
60
|
+
|
|
61
|
+
## The optional counter view (advisory only)
|
|
62
|
+
|
|
63
|
+
Once the immunity check is in place, the counters are a useful *secondary* read — never
|
|
64
|
+
the headline:
|
|
65
|
+
|
|
66
|
+
- **Pull-through per bucket** — `opened_count / read_count` across `loop::<host>-lessons`.
|
|
67
|
+
A bucket read often and opened rarely is dead weight to prune, not proof of value.
|
|
68
|
+
- **`/insights`** — the dashboard's Lore Utility grid classifies each lesson
|
|
69
|
+
(load-bearing / specialist / noise-tax / dormant) from pull-through against an
|
|
70
|
+
evidence floor. Consult it to find lessons to promote (load-bearing) or retire
|
|
71
|
+
(noise-tax), not to claim the loop improved a host.
|
|
72
|
+
- **`cited_count`** — when a host cites the lessons it applied (`cited` on the write),
|
|
73
|
+
a rising `cited_count` on a lesson is *evidence* it is being used. It is evidence,
|
|
74
|
+
never a denominator: a `0` means "nobody said so", not "unused".
|
|
75
|
+
|
|
76
|
+
State the confound wherever you report these: raw counters describe supply, not lift.
|
|
77
|
+
|
|
78
|
+
## Declare the metric at wiring time
|
|
79
|
+
|
|
80
|
+
Do this at step 6 of the quickstart, not months later when you have forgotten what
|
|
81
|
+
"better" meant:
|
|
82
|
+
|
|
83
|
+
- **Required:** name the failure signature (the **Applies when** line) whose recurrence
|
|
84
|
+
you will watch after promotion.
|
|
85
|
+
- **Optional:** name the one bucket pull-through you will glance at, and the interval.
|
|
86
|
+
- Write it into the host's own docs next to the loop's read/write steps, so a future
|
|
87
|
+
maintainer inherits the definition of "working" instead of re-inventing it.
|
|
88
|
+
|
|
89
|
+
## The north star: a withheld-memory control arm
|
|
90
|
+
|
|
91
|
+
The only *rigorous* attribution is a control arm: deterministically bucket a small
|
|
92
|
+
fraction of runs into a **lessons-suppressed** arm (reuse LoreKit's feature-flags
|
|
93
|
+
FNV-1a bucketing on a stable run id), then compare outcome rates — retry count,
|
|
94
|
+
first-try pass, wall-clock — between treated and control. That yields lift with a
|
|
95
|
+
confidence interval instead of a ratio.
|
|
96
|
+
|
|
97
|
+
This is the aspiration, **not a requirement**, for two honest reasons: it needs product
|
|
98
|
+
infrastructure that does not exist yet (a per-loop lift harness), and deliberately
|
|
99
|
+
withholding memory has a real UX/ethics cost, so it must be opt-in. Treat it as the
|
|
100
|
+
direction of travel; ship the immunity re-challenge today.
|
|
101
|
+
|
|
102
|
+
## Guards
|
|
103
|
+
|
|
104
|
+
- **Do not measure a loop you seeded.** Cold-start seeding erases the clean baseline the
|
|
105
|
+
immunity check reads — see [cold-start-seeding.md](./cold-start-seeding.md). Pick, per
|
|
106
|
+
loop: run-one value (seed it) OR provable lift (leave it cold).
|
|
107
|
+
- **A flat `seen_count` after promotion is proof only if the loop is still firing.**
|
|
108
|
+
Confirm the loop is alive first ([loop-health.md](./loop-health.md#is-the-loop-firing));
|
|
109
|
+
a dead loop also has a flat count, for the wrong reason.
|
|
110
|
+
- **Never gate anything on a counter.** These are proof for a human, not a machine gate.
|
|
111
|
+
The only mechanical gate is the compiled invariant ([compiled-invariants.md](./compiled-invariants.md)).
|
|
@@ -15,6 +15,7 @@ The design has two tiers connected by a recurrence gate. Both run on LoreKit.
|
|
|
15
15
|
|
|
16
16
|
## Contents
|
|
17
17
|
|
|
18
|
+
- [Find where a loop pays off](#find-where-a-loop-pays-off)
|
|
18
19
|
- [When to add a loop (and when not to)](#when-to-add-a-loop-and-when-not-to)
|
|
19
20
|
- [The two tiers](#the-two-tiers)
|
|
20
21
|
- [Conventions](#conventions)
|
|
@@ -30,6 +31,41 @@ The design has two tiers connected by a recurrence gate. Both run on LoreKit.
|
|
|
30
31
|
|
|
31
32
|
---
|
|
32
33
|
|
|
34
|
+
## Find where a loop pays off
|
|
35
|
+
|
|
36
|
+
Do not ask a newcomer to *guess* where memory helps — the best place for a loop is a
|
|
37
|
+
statistical property of the codebase's own history, and the data already knows it. Answer
|
|
38
|
+
"where?" before "how?".
|
|
39
|
+
|
|
40
|
+
**Start from recurring pain, not from a host you happen to be looking at.** A loop earns
|
|
41
|
+
its keep only where the same class of failure happens more than once. Two ways to surface
|
|
42
|
+
that:
|
|
43
|
+
|
|
44
|
+
- **From the store, if loops already exist:** `lorekit dedupe` clusters near-duplicate
|
|
45
|
+
lessons, and `lorekit invariants candidates` ranks clusters by summed `seen_count` ×
|
|
46
|
+
distinct scopes — literally "how often, in how many places, did this recur." A high-
|
|
47
|
+
ranked cluster with no dedicated host is a loop waiting to be wired.
|
|
48
|
+
- **From history, on a cold repo:** the same pattern fixed the same way 3+ times, a
|
|
49
|
+
revert-then-refix chain, a review note left across many PRs, or an existing CI guard —
|
|
50
|
+
each is a lesson someone already learned. `git log`, `git log --grep=revert`, and the
|
|
51
|
+
`obligations-map` are the seams. These double as seed sources —
|
|
52
|
+
[cold-start-seeding.md](./cold-start-seeding.md).
|
|
53
|
+
|
|
54
|
+
**Then match the host to an archetype** and copy its recipe card rather than wiring from
|
|
55
|
+
scratch:
|
|
56
|
+
|
|
57
|
+
| The host… | Card |
|
|
58
|
+
| --------- | ---- |
|
|
59
|
+
| edits code, fails in recurring ways | [code-changing-agent](../templates/code-changing-agent.md) |
|
|
60
|
+
| posts durable outputs at a target it revisits | [reviewer-reconcile-host](../templates/reviewer-reconcile-host.md) |
|
|
61
|
+
| is a pipeline that fails in classifiable ways | [multi-step-orchestrator](../templates/multi-step-orchestrator.md) |
|
|
62
|
+
| is a deterministic CI job needing last-run state | [ci-job](../templates/ci-job.md) |
|
|
63
|
+
|
|
64
|
+
The abstract test in the next section is the fallback when no recurrence data exists yet —
|
|
65
|
+
apply it, but prefer the evidence above when you have it.
|
|
66
|
+
|
|
67
|
+
---
|
|
68
|
+
|
|
33
69
|
## When to add a loop (and when not to)
|
|
34
70
|
|
|
35
71
|
Add a loop when **all** of these hold:
|
|
@@ -220,13 +256,17 @@ readable start to finish.
|
|
|
220
256
|
|
|
221
257
|
## Read step (start of every run)
|
|
222
258
|
|
|
223
|
-
Read narrow-to-broad, filtered by the bucket tag, and merge
|
|
259
|
+
Read narrow-to-broad, filtered by the bucket tag, and merge — but **capped**. A loop
|
|
260
|
+
injects at most a small N (`limit: 5` per scope by default) into a run. This is a wiring
|
|
261
|
+
precondition, not a tuning knob: many uncapped loops on the same scopes tax every
|
|
262
|
+
session's context until agents learn to ignore injected lore entirely. Raise N only with
|
|
263
|
+
a reason.
|
|
224
264
|
|
|
225
265
|
```text
|
|
226
|
-
memory.list { scope: "repo::{owner}/{repo}", tags: ["loop::<host>-lessons"], limit:
|
|
227
|
-
memory.list { scope: "global", tags: ["loop::<host>-lessons"], limit:
|
|
266
|
+
memory.list { scope: "repo::{owner}/{repo}", tags: ["loop::<host>-lessons"], limit: 5 } # skips silently if memory.* not connected
|
|
267
|
+
memory.list { scope: "global", tags: ["loop::<host>-lessons"], limit: 5 }
|
|
228
268
|
# when the run names a subsystem / error, add a search:
|
|
229
|
-
memory.search { q: "<keywords>", scopes: ["repo::{owner}/*", "global"], limit:
|
|
269
|
+
memory.search { q: "<keywords>", scopes: ["repo::{owner}/*", "global"], limit: 5 }
|
|
230
270
|
```
|
|
231
271
|
|
|
232
272
|
Then:
|
|
@@ -557,6 +597,14 @@ To add a loop to a host called `<host>`:
|
|
|
557
597
|
read** at its plan/apply seam (match `hotspot::<path>` /
|
|
558
598
|
`knowledge::<symbol>@<path>` to the files it will change). If it **verifies**
|
|
559
599
|
a structural fact, wire the **write** behind that section's contract.
|
|
600
|
+
- [ ] Set the **injection cap** (`limit: 5` per scope by default) on the read step —
|
|
601
|
+
a wiring precondition, not a tuning knob.
|
|
602
|
+
- [ ] Decide **seeding**: value on run one ([cold-start-seeding.md](./cold-start-seeding.md),
|
|
603
|
+
a handful, curated) **or** provable lift (leave it cold) — never both.
|
|
604
|
+
- [ ] Declare the **proof** — the immunity re-challenge whose failure signature you will
|
|
605
|
+
watch after promotion ([proving-improvement.md](./proving-improvement.md)).
|
|
606
|
+
- [ ] Add the **health checks** — a canary, the write-time matchability check, and the
|
|
607
|
+
quarantine-not-delete rollback ([loop-health.md](./loop-health.md)).
|
|
560
608
|
|
|
561
609
|
---
|
|
562
610
|
|
|
@@ -0,0 +1,107 @@
|
|
|
1
|
+
# Team & portfolio — from one loop to a practice
|
|
2
|
+
|
|
3
|
+
Everything else in this skill is written for one person wiring one host. But the store
|
|
4
|
+
is shared, and that changes the economics: a lesson one person's loop writes can rescue
|
|
5
|
+
another person's host, the same pain relearned by three people is a louder signal than
|
|
6
|
+
any single run, and a new hire can inherit years of accumulated workflow intelligence on
|
|
7
|
+
day one. This file is how a team turns scattered loops into a compounding practice.
|
|
8
|
+
|
|
9
|
+
Some of this is guidance a person acts on today; some names a dashboard/product surface
|
|
10
|
+
that would make it turnkey (flagged **[product]**).
|
|
11
|
+
|
|
12
|
+
## Contents
|
|
13
|
+
|
|
14
|
+
- [The rediscovery signal: where a team loop pays off](#the-rediscovery-signal-where-a-team-loop-pays-off)
|
|
15
|
+
- [The maturity model](#the-maturity-model)
|
|
16
|
+
- [Promotion as review](#promotion-as-review)
|
|
17
|
+
- [Onboarding: inherit the lore](#onboarding-inherit-the-lore)
|
|
18
|
+
- [Cross-host portability](#cross-host-portability)
|
|
19
|
+
- [Governance is mostly self-executing](#governance-is-mostly-self-executing)
|
|
20
|
+
|
|
21
|
+
---
|
|
22
|
+
|
|
23
|
+
## The rediscovery signal: where a team loop pays off
|
|
24
|
+
|
|
25
|
+
At single-host scale, `seen_count` is a promotion trigger. At team scale it is a
|
|
26
|
+
**rediscovery counter**: the same lesson climbing in count across **distinct authors and
|
|
27
|
+
scopes** is proof a pain is *shared*, not personal — and shared pain is exactly where a
|
|
28
|
+
team-wide loop or rule pays off most.
|
|
29
|
+
|
|
30
|
+
Find it with the tools that already exist:
|
|
31
|
+
|
|
32
|
+
- `lorekit dedupe` — near-duplicate lessons written independently by different people are
|
|
33
|
+
one shared lesson wearing several names; the cluster's summed `seen_count` is the team
|
|
34
|
+
signal.
|
|
35
|
+
- `lorekit invariants candidates` — ranks clusters by summed `seen_count` × distinct
|
|
36
|
+
scopes, which is precisely "how many people, in how many places, hit this."
|
|
37
|
+
- **[product]** A dashboard "rediscovery radar" that surfaces cross-author clusters with
|
|
38
|
+
a one-click "wire a loop here" would make this turnkey.
|
|
39
|
+
|
|
40
|
+
## The maturity model
|
|
41
|
+
|
|
42
|
+
The FAST → SLOW → compiled-invariant rungs describe one loop's lifecycle. Aggregated
|
|
43
|
+
across every host, they become a **portfolio view** that shows where the team's loops are
|
|
44
|
+
thin:
|
|
45
|
+
|
|
46
|
+
| Level | State | Evidence |
|
|
47
|
+
| ----- | ----- | -------- |
|
|
48
|
+
| **L0** | No loop | The host has no `loop::<host>-lessons` bucket |
|
|
49
|
+
| **L1** | Fast advisories only | Lessons accrue and are read; nothing promoted yet |
|
|
50
|
+
| **L2** | Slow promotions landing | Recurring lessons hardened into host rules (`status::promoted`) |
|
|
51
|
+
| **L3** | Compiled invariant | A mechanically-checked `obligations-map` entry, `gating` |
|
|
52
|
+
|
|
53
|
+
Plot each workflow's level against how many teammates actually run it. A high-traffic
|
|
54
|
+
workflow stuck at L0/L1 is the team's biggest missed compounding. **[product]** A
|
|
55
|
+
maturity board on `/insights` would render this from the tags and counters that already
|
|
56
|
+
exist.
|
|
57
|
+
|
|
58
|
+
## Promotion as review
|
|
59
|
+
|
|
60
|
+
Promoting a `repo::` lesson to a team rule is a governance moment, and it maps cleanly
|
|
61
|
+
onto a process the team already trusts — code review:
|
|
62
|
+
|
|
63
|
+
1. A promotion-eligible lesson (`seen_count >= 3` or `status::structural`) enters a
|
|
64
|
+
"pending team rule" state.
|
|
65
|
+
2. A **second engineer** reviews it, exactly as they would a PR. The gate the loop
|
|
66
|
+
computes automatically becomes the human checkpoint.
|
|
67
|
+
3. On approval, the promotion is a normal reviewed edit to the repo's rules/docs; add a
|
|
68
|
+
`status::promoted` tag so it stops re-suggesting and stands as an audit trail.
|
|
69
|
+
4. On rejection, write a lesson about **why it was not generalized** — that reasoning is
|
|
70
|
+
itself reusable.
|
|
71
|
+
|
|
72
|
+
## Onboarding: inherit the lore
|
|
73
|
+
|
|
74
|
+
Onboarding content does not need to be written — it is the exact bucket an agent already
|
|
75
|
+
reads at SessionStart, re-aimed at a human. Generate a new hire's "what this team learned
|
|
76
|
+
the hard way" brief from:
|
|
77
|
+
|
|
78
|
+
- the highest-`seen_count` and `status::promoted` lessons per repo (rules first),
|
|
79
|
+
- the `codebase-knowledge` hotspots and invariants (where the bodies are buried),
|
|
80
|
+
- ordered by rung: promoted rules, then advisories.
|
|
81
|
+
|
|
82
|
+
Because it is derived from the live store, it is never stale and needs no maintenance.
|
|
83
|
+
|
|
84
|
+
## Cross-host portability
|
|
85
|
+
|
|
86
|
+
Eight people on Claude Code, Cursor, and Codex normally fragment knowledge into per-tool
|
|
87
|
+
silos. The shared store lets host-diversity be an *asset* instead:
|
|
88
|
+
|
|
89
|
+
- **Host-tag every lesson** (the `host` field). A lesson read by a *different* host than
|
|
90
|
+
wrote it is portable; one only ever read within its authoring host may be host-local.
|
|
91
|
+
- **A lesson proven across two hosts is stronger evidence** it is a real team rule than
|
|
92
|
+
one seen within a single tool — weight it higher at promotion.
|
|
93
|
+
- At promotion to a team rule, **strip host-specific phrasing** ("in Claude Code, …") so
|
|
94
|
+
the rule reads for any host.
|
|
95
|
+
|
|
96
|
+
## Governance is mostly self-executing
|
|
97
|
+
|
|
98
|
+
The failure mode of "team governance" is that it assumes a human process watching loops
|
|
99
|
+
decay — which solo devs and small teams skip, so entrenchment guards rot. Lean on the
|
|
100
|
+
mechanisms that need no staffing:
|
|
101
|
+
|
|
102
|
+
- **TTL expiry** drops stale lessons automatically — decay is built in, not a chore.
|
|
103
|
+
- **The recurrence gate** means a single bad run can never rewrite a host.
|
|
104
|
+
- **The injection cap** bounds read cost no matter how the bucket grows.
|
|
105
|
+
|
|
106
|
+
The only step that genuinely needs a human is **promotion review** above. Keep the
|
|
107
|
+
staffed surface that small; let the defaults do the rest.
|
|
@@ -0,0 +1,26 @@
|
|
|
1
|
+
# Loop recipe cards
|
|
2
|
+
|
|
3
|
+
Copy-paste starters for wiring a self-improvement loop into a host. Pick the card
|
|
4
|
+
that matches your host, paste its read/write steps at the seams it names, and fill in
|
|
5
|
+
`<host>` / `<owner>/<repo>`. Each card already bakes in the five non-negotiables from
|
|
6
|
+
the skill's [Non-negotiables](../SKILL.md#non-negotiables-do-not-optimize-these-away):
|
|
7
|
+
a per-host **injection cap**, a **TTL**, the **advisory-only** contract, the **body
|
|
8
|
+
contract** (pure markdown), and a **fire-once check** so you prove the loop closes
|
|
9
|
+
before you trust it.
|
|
10
|
+
|
|
11
|
+
| Card | Use it for | Shape |
|
|
12
|
+
| ---- | ---------- | ----- |
|
|
13
|
+
| [code-changing-agent.md](./code-changing-agent.md) | A host that edits code and can fail in recurring ways (an autonomous workflow, a fix/implement agent) | Lessons loop + the shared `codebase-knowledge` read/write |
|
|
14
|
+
| [reviewer-reconcile-host.md](./reviewer-reconcile-host.md) | A host that posts durable outputs at a shared target it revisits (a PR reviewer, a triager, a linter that files tickets) | Lessons loop + the reconcile-on-re-run Signal bucket |
|
|
15
|
+
| [multi-step-orchestrator.md](./multi-step-orchestrator.md) | A pipeline that fails in classifiable ways (wrong triage, a missed step, a false-green gate) | The plain two-tier lessons loop |
|
|
16
|
+
| [ci-job.md](./ci-job.md) | A deterministic job that needs last-run state (flaky set, last deployed SHA, a baseline) | JSON state records, not lessons |
|
|
17
|
+
|
|
18
|
+
Every card is a *starting point*, not a spec. The authoritative rules are one level up:
|
|
19
|
+
[self-improvement-loops.md](../rules/self-improvement-loops.md) (lessons),
|
|
20
|
+
[ci-state-records.md](../rules/ci-state-records.md) (state records),
|
|
21
|
+
[proving-improvement.md](../rules/proving-improvement.md) (proof),
|
|
22
|
+
[loop-health.md](../rules/loop-health.md) (maintenance).
|
|
23
|
+
|
|
24
|
+
**`N` — the injection cap — is `5` in every card by default.** Raise it only with a
|
|
25
|
+
reason; a loop that reads 30 lessons into every run is why agents learn to ignore
|
|
26
|
+
injected lore. See [Non-negotiables #2](../SKILL.md#non-negotiables-do-not-optimize-these-away).
|
|
@@ -0,0 +1,90 @@
|
|
|
1
|
+
# Recipe card — CI job (state record, not a lesson)
|
|
2
|
+
|
|
3
|
+
For a **deterministic job** — a GitHub Actions workflow, a cron script, a release
|
|
4
|
+
pipeline — that needs to know *what happened last time*: which tests flaked, the last
|
|
5
|
+
benchmark number, the last deployed SHA. This is **not** a lessons loop: nothing here
|
|
6
|
+
is authored or interpreted by a model, so there is no `seen_count`, no promotion, no
|
|
7
|
+
entrenchment guard. A different, smaller set of guards applies. Full rule:
|
|
8
|
+
[ci-state-records.md](../rules/ci-state-records.md).
|
|
9
|
+
|
|
10
|
+
Fill in: `<job>` (e.g. `test`), `<slug>` (e.g. `flaky-tests`), `<owner>/<repo>`.
|
|
11
|
+
|
|
12
|
+
> Use `actions/cache` instead if **only** the next CI run will ever read this. LoreKit
|
|
13
|
+
> earns its place only when the fact is *also* useful to an agent or a human.
|
|
14
|
+
|
|
15
|
+
## Bucket
|
|
16
|
+
|
|
17
|
+
- Tag `ci::<job>-state` (deliberately NOT `loop::…`). Key `ci-state::<slug>`, one slug
|
|
18
|
+
per fact, **overwritten in place** — never one key per run.
|
|
19
|
+
- Taxonomy: pass `--kind bus --host ci` explicitly, or the record leaks into every
|
|
20
|
+
session's SessionStart digest as a raw JSON blob.
|
|
21
|
+
|
|
22
|
+
## Record shape (versioned JSON)
|
|
23
|
+
|
|
24
|
+
```json
|
|
25
|
+
{
|
|
26
|
+
"v": 1,
|
|
27
|
+
"updated_by_run": "https://github.com/<owner>/<repo>/actions/runs/123",
|
|
28
|
+
"commit": "0f4a1c9…",
|
|
29
|
+
"data": { "flaky": ["src/queue.test.ts::retries on 429"], "consecutive_green": 3 }
|
|
30
|
+
}
|
|
31
|
+
```
|
|
32
|
+
|
|
33
|
+
## Read step — early in the job
|
|
34
|
+
|
|
35
|
+
```bash
|
|
36
|
+
set -euo pipefail
|
|
37
|
+
STATE_JSON='{}'
|
|
38
|
+
if lorekit show --scope "repo::<owner>/<repo>" --key 'ci-state::<slug>' --remote --json > state.json 2>&1; then
|
|
39
|
+
cat state.json # always echo — never `|| true`
|
|
40
|
+
STATE_JSON=$(jq -r '.remote.record.value // "{}"' state.json)
|
|
41
|
+
else
|
|
42
|
+
cat state.json
|
|
43
|
+
echo "No prior state (first run, or LoreKit unreachable) — continuing with defaults."
|
|
44
|
+
fi
|
|
45
|
+
VERSION=$(jq -r '.v // 0' <<<"$STATE_JSON")
|
|
46
|
+
[ "$VERSION" = "1" ] || { echo "State schema v${VERSION} != v1 — rebuilding."; STATE_JSON='{}'; }
|
|
47
|
+
```
|
|
48
|
+
|
|
49
|
+
`lorekit show` exits 1 on a miss — that is the first run, branch on it; do not `|| true`
|
|
50
|
+
it away (that swallows a real auth/network failure too).
|
|
51
|
+
|
|
52
|
+
## Write step — end of job (`if: always()` if state should survive a failing run)
|
|
53
|
+
|
|
54
|
+
```bash
|
|
55
|
+
set -euo pipefail
|
|
56
|
+
jq -nc \
|
|
57
|
+
--arg run "${GITHUB_SERVER_URL}/${GITHUB_REPOSITORY}/actions/runs/${GITHUB_RUN_ID}" \
|
|
58
|
+
--arg sha "${GITHUB_SHA}" \
|
|
59
|
+
--argjson data "$NEW_DATA" \
|
|
60
|
+
'{v: 1, updated_by_run: $run, commit: $sha, data: $data}' \
|
|
61
|
+
| lorekit write \
|
|
62
|
+
--scope "repo::<owner>/<repo>" --key 'ci-state::<slug>' \
|
|
63
|
+
--tags 'ci::<job>-state' --kind bus --host ci \
|
|
64
|
+
--ttl-days 7 --remote --json \
|
|
65
|
+
| tee write.log \
|
|
66
|
+
|| echo "LoreKit write failed (exit $?) — not failing the build; see output above."
|
|
67
|
+
```
|
|
68
|
+
|
|
69
|
+
`--ttl-days` is **not optional** and should be short (7 for daily jobs) — it refreshes
|
|
70
|
+
every write, so it expires only when the *job* goes silent. See
|
|
71
|
+
[TTL is a liveness guard](../rules/ci-state-records.md#ttl-is-a-liveness-guard).
|
|
72
|
+
|
|
73
|
+
## Tokens
|
|
74
|
+
|
|
75
|
+
Write with a **write-only `lk_wo_*`** token (a leaked CI token then cannot exfiltrate
|
|
76
|
+
the team's lore); read with `lk_ro_*`; use `lk_rw_*` only where one job genuinely needs
|
|
77
|
+
both.
|
|
78
|
+
|
|
79
|
+
## Guards (do not skip)
|
|
80
|
+
|
|
81
|
+
Bounded cardinality (one key per fact) · no secrets, ever (build the payload from an
|
|
82
|
+
allow-list, never `env | jq -R`) · explicit short TTL · never on the critical path
|
|
83
|
+
(a store outage logs and continues) · last-write-wins (serialise with a `concurrency`
|
|
84
|
+
group) · version every record. Full list:
|
|
85
|
+
[Guards](../rules/ci-state-records.md#guards-do-not-skip-these).
|
|
86
|
+
|
|
87
|
+
## Verify
|
|
88
|
+
|
|
89
|
+
Run the job twice; confirm `lorekit list --kind bus --host ci` still shows **one** row
|
|
90
|
+
per fact (key count must not grow with run count).
|
|
@@ -0,0 +1,89 @@
|
|
|
1
|
+
# Recipe card — code-changing agent
|
|
2
|
+
|
|
3
|
+
For a host that **edits code** and fails in recurring ways: an autonomous workflow, a
|
|
4
|
+
fix-bug agent, an implement-suggestion worker. It wants two reads — its own lessons
|
|
5
|
+
*and* the shared `codebase-knowledge` for the files it is about to touch — and two
|
|
6
|
+
writes — a lesson on failure, and a verified structural fact back to
|
|
7
|
+
`codebase-knowledge`.
|
|
8
|
+
|
|
9
|
+
Fill in: `<host>` (e.g. `fix-bug`), `<owner>/<repo>`. `N = 5`.
|
|
10
|
+
|
|
11
|
+
## Bucket
|
|
12
|
+
|
|
13
|
+
- Lessons: tag `loop::<host>-lessons`, key `<host>-lessons::<slug>`.
|
|
14
|
+
- Shared: tag `codebase-knowledge` (read by every code-touching loop; see
|
|
15
|
+
[the shared layer](../rules/self-improvement-loops.md#shared-codebase-knowledge-the-standard-cross-loop-layer)).
|
|
16
|
+
|
|
17
|
+
## Read step — at the start of the run AND at the plan/apply seam
|
|
18
|
+
|
|
19
|
+
```text
|
|
20
|
+
# 1. Own lessons, narrow-to-broad, capped at N per scope:
|
|
21
|
+
memory.list { scope: "repo::<owner>/<repo>", tags: ["loop::<host>-lessons"], limit: 5 }
|
|
22
|
+
memory.list { scope: "global", tags: ["loop::<host>-lessons"], limit: 5 }
|
|
23
|
+
|
|
24
|
+
# 2. When the run names a subsystem/error, one targeted search:
|
|
25
|
+
memory.search { q: "<keywords>", scopes: ["repo::<owner>/*", "global"], limit: 5 }
|
|
26
|
+
|
|
27
|
+
# 3. At the plan/apply seam — once you have the concrete file/symbol list:
|
|
28
|
+
memory.list { scope: "repo::<owner>/<repo>", tags: ["codebase-knowledge"], limit: 100 }
|
|
29
|
+
# keep ONLY hotspot::<path> / knowledge::<symbol>@<path> whose <path>/<symbol>
|
|
30
|
+
# this run will touch. Apply as PLANNING INPUT: raise coverage on a hotspot,
|
|
31
|
+
# design around a known invariant. Advisory; re-verify against the code.
|
|
32
|
+
```
|
|
33
|
+
|
|
34
|
+
Match each lesson's **Applies when** line against the current run; apply matching
|
|
35
|
+
**Do this instead** lines as *considerations*, not commands. On a `repo::` vs `global`
|
|
36
|
+
collision, `repo::` wins. An absent `codebase-knowledge` record is never evidence of
|
|
37
|
+
safety.
|
|
38
|
+
|
|
39
|
+
## Write step — on failure / at end of run
|
|
40
|
+
|
|
41
|
+
```text
|
|
42
|
+
# Dedup first so a recurrence UPDATEs in place:
|
|
43
|
+
memory.search { q: "<key words of the lesson>", scopes: ["repo::<owner>/<repo>", "global"], limit: 10 }
|
|
44
|
+
|
|
45
|
+
memory.write {
|
|
46
|
+
scope: "<global | repo::<owner>/<repo>>", # repo:: if it names a repo path/term, else global
|
|
47
|
+
key: "<host>-lessons::<slug>",
|
|
48
|
+
value: "<markdown lesson body — no hidden blocks>",
|
|
49
|
+
tags: ["loop::<host>-lessons", "source::<trigger>"],
|
|
50
|
+
trigger: "<stuck-loop | command-failure | gotcha | near-miss | assumption-wrong | paid-off>",
|
|
51
|
+
ttl_days: 90
|
|
52
|
+
}
|
|
53
|
+
```
|
|
54
|
+
|
|
55
|
+
If the host **verified** a structural fact this run (an invariant a symbol holds, a
|
|
56
|
+
file that produced a defect), write it back to `codebase-knowledge` under the
|
|
57
|
+
multi-writer contract — `verified_at_sha`, `source_agent`, merge-not-clobber, raise
|
|
58
|
+
care never suppress:
|
|
59
|
+
[write side](../rules/self-improvement-loops.md#write-side--how-the-layer-fills-and-why-many-writers-stay-safe).
|
|
60
|
+
|
|
61
|
+
## Lesson body — copy this shape
|
|
62
|
+
|
|
63
|
+
```markdown
|
|
64
|
+
# <one-line takeaway — what to do, not what it is about>
|
|
65
|
+
|
|
66
|
+
**Applies when:** <concrete signal — file glob, task type, tool name, error shape>
|
|
67
|
+
|
|
68
|
+
**What happened:** <the concrete observable>
|
|
69
|
+
**Why:** <root cause, or "unknown">
|
|
70
|
+
**Do this instead:** <prescriptive, testable instruction>
|
|
71
|
+
**Promotion target:** <the host rule this would harden if promoted, or "none">
|
|
72
|
+
```
|
|
73
|
+
|
|
74
|
+
## Fire-once check (do this before trusting the loop)
|
|
75
|
+
|
|
76
|
+
1. Trigger the failure this loop is meant to catch.
|
|
77
|
+
2. Confirm the lesson landed: `memory.list { scope: "<expected>", tags: ["loop::<host>-lessons"], limit: 10 }`.
|
|
78
|
+
3. Re-run over the same situation; confirm it surfaces in the read step and biases the run.
|
|
79
|
+
|
|
80
|
+
## Cold start (optional, capped)
|
|
81
|
+
|
|
82
|
+
Pre-seed a *handful* of real lessons from git history / prior review threads so run one
|
|
83
|
+
is not empty — [cold-start-seeding.md](../rules/cold-start-seeding.md). Do **not** seed
|
|
84
|
+
a loop whose lift you intend to measure; it erases the baseline.
|
|
85
|
+
|
|
86
|
+
## Prove it
|
|
87
|
+
|
|
88
|
+
Declare the immunity re-challenge at wiring time: after you promote a lesson to a rule,
|
|
89
|
+
its failure signature should stop recurring — [proving-improvement.md](../rules/proving-improvement.md).
|
|
@@ -0,0 +1,76 @@
|
|
|
1
|
+
# Recipe card — multi-step orchestrator
|
|
2
|
+
|
|
3
|
+
For a **pipeline that fails in classifiable ways** but does not post durable outputs at
|
|
4
|
+
a shared target and does not (necessarily) edit code: a triage → plan → execute
|
|
5
|
+
workflow, a release orchestrator, a batch processor. The plain two-tier lessons loop
|
|
6
|
+
fits it exactly. This is the card to start from when no other card matches.
|
|
7
|
+
|
|
8
|
+
Fill in: `<host>` (e.g. `deploy`), `<owner>/<repo>`. `N = 5`.
|
|
9
|
+
|
|
10
|
+
## Bucket
|
|
11
|
+
|
|
12
|
+
Tag `loop::<host>-lessons`, key `<host>-lessons::<slug>`. One bucket per host.
|
|
13
|
+
|
|
14
|
+
## Read step — start of every run
|
|
15
|
+
|
|
16
|
+
```text
|
|
17
|
+
memory.list { scope: "repo::<owner>/<repo>", tags: ["loop::<host>-lessons"], limit: 5 }
|
|
18
|
+
memory.list { scope: "global", tags: ["loop::<host>-lessons"], limit: 5 }
|
|
19
|
+
# when the run names a stage/error, one targeted search:
|
|
20
|
+
memory.search { q: "<keywords>", scopes: ["repo::<owner>/*", "global"], limit: 5 }
|
|
21
|
+
```
|
|
22
|
+
|
|
23
|
+
Match each **Applies when** against the current run; apply matching **Do this instead**
|
|
24
|
+
lines as considerations. `repo::` beats `global` on a collision.
|
|
25
|
+
|
|
26
|
+
## Write step — at the host's EXISTING failure points
|
|
27
|
+
|
|
28
|
+
Do not add a new reflection stage — hook the points the host already detects (a stuck
|
|
29
|
+
loop, a repeated failure, a gate that should have caught something, a near-miss, a
|
|
30
|
+
guess that paid off). Not on smooth successes.
|
|
31
|
+
|
|
32
|
+
```text
|
|
33
|
+
memory.search { q: "<key words of the lesson>", scopes: ["repo::<owner>/<repo>", "global"], limit: 10 }
|
|
34
|
+
|
|
35
|
+
memory.write {
|
|
36
|
+
scope: "<global | repo::<owner>/<repo>>",
|
|
37
|
+
key: "<host>-lessons::<slug>",
|
|
38
|
+
value: "<markdown lesson body — no hidden blocks>",
|
|
39
|
+
tags: ["loop::<host>-lessons", "source::<trigger>"], # + "status::structural" when it is
|
|
40
|
+
trigger: "<stuck-loop | command-failure | gotcha | near-miss | assumption-wrong | paid-off>",
|
|
41
|
+
ttl_days: 90
|
|
42
|
+
}
|
|
43
|
+
```
|
|
44
|
+
|
|
45
|
+
Same `scope` + `key` UPDATEs in place — the store increments `seen_count` for you and
|
|
46
|
+
re-passing `ttl_days` refreshes the expiry. Never hand-write a count into the body.
|
|
47
|
+
|
|
48
|
+
## Lesson body — copy this shape
|
|
49
|
+
|
|
50
|
+
```markdown
|
|
51
|
+
# <one-line takeaway — what to do, not what it is about>
|
|
52
|
+
|
|
53
|
+
**Applies when:** <concrete signal — stage name, task type, tool name, error shape>
|
|
54
|
+
|
|
55
|
+
**What happened:** <the concrete observable>
|
|
56
|
+
**Why:** <root cause, or "unknown">
|
|
57
|
+
**Do this instead:** <prescriptive, testable instruction>
|
|
58
|
+
**Promotion target:** <the host step this would harden if promoted, or "none">
|
|
59
|
+
```
|
|
60
|
+
|
|
61
|
+
## Promotion
|
|
62
|
+
|
|
63
|
+
When a lesson hits `seen_count >= 3` (read the column, not the body) or carries
|
|
64
|
+
`status::structural`, surface a one-line promotion suggestion — never act silently.
|
|
65
|
+
See [Promotion](../rules/self-improvement-loops.md#promotion-fast--slow).
|
|
66
|
+
|
|
67
|
+
## Fire-once check
|
|
68
|
+
|
|
69
|
+
1. Force the failure this loop targets.
|
|
70
|
+
2. `memory.list { scope: "<expected>", tags: ["loop::<host>-lessons"], limit: 10 }` — confirm the lesson.
|
|
71
|
+
3. Re-run; confirm it surfaces and biases the run.
|
|
72
|
+
|
|
73
|
+
## Cold start & proof
|
|
74
|
+
|
|
75
|
+
Optional capped seeding: [cold-start-seeding.md](../rules/cold-start-seeding.md).
|
|
76
|
+
Required proof — the immunity re-challenge: [proving-improvement.md](../rules/proving-improvement.md).
|