@lorekit/cli 1.68.0 → 1.70.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +29 -8
- package/package.json +1 -1
- package/skill/lorekit-groom/rules/grooming-pass.md +34 -0
- package/skill/lorekit-memory/SKILL.md +20 -2
- package/skill/lorekit-memory/references/scope-resolution.md +23 -1
- package/skill/lorekit-setup/SKILL.md +153 -92
- package/skill/lorekit-setup/rules/ci-state-records.md +1 -1
- package/skill/lorekit-setup/rules/cold-start-seeding.md +102 -0
- package/skill/lorekit-setup/rules/compiled-invariants.md +9 -8
- package/skill/lorekit-setup/rules/loop-health.md +117 -0
- package/skill/lorekit-setup/rules/proving-improvement.md +111 -0
- package/skill/lorekit-setup/rules/self-improvement-loops.md +196 -41
- package/skill/lorekit-setup/rules/team-and-portfolio.md +107 -0
- package/skill/lorekit-setup/templates/README.md +26 -0
- package/skill/lorekit-setup/templates/ci-job.md +90 -0
- package/skill/lorekit-setup/templates/code-changing-agent.md +89 -0
- package/skill/lorekit-setup/templates/multi-step-orchestrator.md +76 -0
- package/skill/lorekit-setup/templates/reviewer-reconcile-host.md +86 -0
- package/src/commands/invariants.mjs +17 -10
- package/src/commands/lint.mjs +4 -2
- package/src/commands/show.mjs +244 -3
- package/src/shared/candidates-pure.mjs +94 -5
- package/src/shared/lessons-view.mjs +71 -0
- package/src/shared/mirror-pairs.mjs +9 -0
- package/src/shared/scope-precedence.mjs +82 -0
- package/src/store/local.mjs +95 -1
- package/src/store/remote.mjs +60 -3
- package/src/surfaces.generated.mjs +14 -17
|
@@ -0,0 +1,117 @@
|
|
|
1
|
+
# Loop health — keeping it firing, clean, and honest
|
|
2
|
+
|
|
3
|
+
A wired loop is not a working loop. This file is the maintenance layer: how to tell the
|
|
4
|
+
loop is firing at all, how to catch a lesson that will never match the read that should
|
|
5
|
+
surface it, how to roll back a bad lesson without losing the evidence, how to stop the
|
|
6
|
+
bucket bloating, and how to keep every body honest. Most of it is a check you run once
|
|
7
|
+
at wiring time; the rest the `lorekit-groom` skill automates.
|
|
8
|
+
|
|
9
|
+
## Contents
|
|
10
|
+
|
|
11
|
+
- [Is the loop firing?](#is-the-loop-firing)
|
|
12
|
+
- [Does the write match the read? (matchability check)](#does-the-write-match-the-read-matchability-check)
|
|
13
|
+
- [Rolling back a bad lesson: quarantine, not delete](#rolling-back-a-bad-lesson-quarantine-not-delete)
|
|
14
|
+
- [Keep the bucket from bloating](#keep-the-bucket-from-bloating)
|
|
15
|
+
- [The body contract is now enforced](#the-body-contract-is-now-enforced)
|
|
16
|
+
|
|
17
|
+
---
|
|
18
|
+
|
|
19
|
+
## Is the loop firing?
|
|
20
|
+
|
|
21
|
+
The most common silent failure is a loop that never runs — a read step behind a
|
|
22
|
+
condition that is never true, or a `memory.*` connection that quietly dropped. Silence
|
|
23
|
+
looks identical to "no failures happened."
|
|
24
|
+
|
|
25
|
+
Turn the absence into a presence with a **canary**: at wiring time, plant one canary
|
|
26
|
+
lesson per bucket with a unique key and a `ttl_days` shorter than the loop's cadence.
|
|
27
|
+
|
|
28
|
+
```text
|
|
29
|
+
memory.write {
|
|
30
|
+
scope: "repo::<owner>/<repo>",
|
|
31
|
+
key: "<host>-lessons::canary",
|
|
32
|
+
value: "# Canary — if the read step surfaced this, the loop's read path works.",
|
|
33
|
+
tags: ["loop::<host>-lessons", "status::canary"],
|
|
34
|
+
trigger: "manual",
|
|
35
|
+
ttl_days: <shorter than the loop's run cadence>
|
|
36
|
+
}
|
|
37
|
+
```
|
|
38
|
+
|
|
39
|
+
If the loop is alive, its own read/write touches the canary and the store's usage
|
|
40
|
+
telemetry records it; if the canary quietly expires without ever being touched, the loop
|
|
41
|
+
is not firing — a diagnosable signal where silence was not. Check recency with
|
|
42
|
+
`lorekit stats` / `lorekit list --tags loop::<host>-lessons`. Exclude `status::canary`
|
|
43
|
+
from the run's actual considerations.
|
|
44
|
+
|
|
45
|
+
## Does the write match the read? (matchability check)
|
|
46
|
+
|
|
47
|
+
A lesson that the read step will never surface is dead on arrival — and you only find
|
|
48
|
+
out much later, when the same failure recurs and the loop "did not help." The only
|
|
49
|
+
moment you can catch it is at **write** time, when you still know what the lesson was
|
|
50
|
+
supposed to catch.
|
|
51
|
+
|
|
52
|
+
Right after writing a failure lesson, re-query the store with the **same terms the read
|
|
53
|
+
step uses** and confirm the just-written lesson comes back:
|
|
54
|
+
|
|
55
|
+
```text
|
|
56
|
+
memory.search { q: "<the failure's key terms>", scopes: ["repo::<owner>/<repo>", "global"], limit: 10 }
|
|
57
|
+
# Is the lesson you just wrote in the results? If not, its tags / Applies-when / scope
|
|
58
|
+
# do not match how the read step looks for it — fix them NOW, while you have the context.
|
|
59
|
+
```
|
|
60
|
+
|
|
61
|
+
A miss means a mismatch between how the lesson was written and how it will be sought —
|
|
62
|
+
usually a too-specific **Applies when**, a wrong scope, or a missing tag. Fixing it at
|
|
63
|
+
write time is cheap; discovering it at the next failure is not.
|
|
64
|
+
|
|
65
|
+
## Rolling back a bad lesson: quarantine, not delete
|
|
66
|
+
|
|
67
|
+
A loop can store a wrong conclusion. Deleting it destroys the evidence of *what the
|
|
68
|
+
agent believed* when it got stuck — which is exactly the forensic value you want right
|
|
69
|
+
after the harm. Prefer a reversible demotion:
|
|
70
|
+
|
|
71
|
+
1. **Quarantine** the suspect lesson: add a `status::quarantine` tag and shorten its
|
|
72
|
+
`ttl_days`. The start-of-run read filters `status::quarantine` OUT, so it stops
|
|
73
|
+
biasing runs, but it survives long enough for a human to inspect it.
|
|
74
|
+
2. **Inspect**, decide: rewrite it (a real lesson, badly phrased) or let it expire (a
|
|
75
|
+
genuine mistake).
|
|
76
|
+
3. **Protect** a lesson a human has vetted with `memory.protect`, so grooming and
|
|
77
|
+
auto-quarantine never touch it.
|
|
78
|
+
|
|
79
|
+
Delete only a lesson that is both wrong *and* worthless to inspect. Quarantine is the
|
|
80
|
+
default; deletion is the exception.
|
|
81
|
+
|
|
82
|
+
## Keep the bucket from bloating
|
|
83
|
+
|
|
84
|
+
Every lesson in a bucket competes for the run's read budget. Two mechanisms keep it
|
|
85
|
+
bounded:
|
|
86
|
+
|
|
87
|
+
- **The injection cap (a wiring precondition).** The loop reads at most N lessons into a
|
|
88
|
+
run (the recipe cards default N = 5). This bounds read cost *regardless* of how big the
|
|
89
|
+
bucket grows — it is the backstop.
|
|
90
|
+
- **Grooming.** Run the `lorekit-groom` skill periodically to merge near-duplicates,
|
|
91
|
+
expire the stale, and retire the never-opened (low pull-through). Grooming is where a
|
|
92
|
+
bucket's total size is actually reduced; the cap only bounds what a single run reads.
|
|
93
|
+
|
|
94
|
+
A bucket whose read is always dominated by low-value lessons is a grooming problem, not
|
|
95
|
+
a reason to raise the cap.
|
|
96
|
+
|
|
97
|
+
## The body contract is now enforced
|
|
98
|
+
|
|
99
|
+
Every lesson body is markdown for humans — no HTML comment, no front-matter, no JSON
|
|
100
|
+
blob, no `key=value` header ([the full contract](./self-improvement-loops.md#never-put-machine-metadata-in-the-body)).
|
|
101
|
+
This is no longer honor-system:
|
|
102
|
+
|
|
103
|
+
- **`lorekit lint` flags a violation.** A body containing `<!-- meta` / any `<!--`
|
|
104
|
+
block, leading front-matter (`---`), or a `key=value` header before the `#` title is
|
|
105
|
+
reported by the `hidden-metadata` lint rule. Run `lorekit lint` after a batch of
|
|
106
|
+
writes.
|
|
107
|
+
- **The `lorekit-groom` skill cleans legacy offenders.** The store still contains records
|
|
108
|
+
written under the old `<!-- meta: seen_count=… status=… trigger-context=… -->`
|
|
109
|
+
convention (which this skill once prescribed and now forbids); grooming detects them,
|
|
110
|
+
folds any recoverable value into the proper column/tag/field, and strips the block. See
|
|
111
|
+
the `lorekit-groom` skill.
|
|
112
|
+
- **The matchability check above** is the natural moment to also eyeball the body: if you
|
|
113
|
+
find yourself reaching for a compact hidden header, stop — every fact in it has a
|
|
114
|
+
first-class field.
|
|
115
|
+
|
|
116
|
+
Never write a hidden block. It renders to nothing for a human and its baked-in values
|
|
117
|
+
silently disagree with the store's own columns.
|
|
@@ -0,0 +1,111 @@
|
|
|
1
|
+
# Proving a loop improved something
|
|
2
|
+
|
|
3
|
+
A loop that nobody can see helping is indistinguishable from no loop. Wiring is not
|
|
4
|
+
the deliverable — **a real failure that stops recurring** is. This file is how you show
|
|
5
|
+
that, and — as important — how you avoid fooling yourself with numbers that cannot prove
|
|
6
|
+
it.
|
|
7
|
+
|
|
8
|
+
## Contents
|
|
9
|
+
|
|
10
|
+
- [The required proof: the immunity re-challenge](#the-required-proof-the-immunity-re-challenge)
|
|
11
|
+
- [Why the raw counters cannot prove causation](#why-the-raw-counters-cannot-prove-causation)
|
|
12
|
+
- [The optional counter view (advisory only)](#the-optional-counter-view-advisory-only)
|
|
13
|
+
- [Declare the metric at wiring time](#declare-the-metric-at-wiring-time)
|
|
14
|
+
- [The north star: a withheld-memory control arm](#the-north-star-a-withheld-memory-control-arm)
|
|
15
|
+
- [Guards](#guards)
|
|
16
|
+
|
|
17
|
+
---
|
|
18
|
+
|
|
19
|
+
## The required proof: the immunity re-challenge
|
|
20
|
+
|
|
21
|
+
There is exactly one proof this skill *requires*, because it is the one a doc can make
|
|
22
|
+
a person actually run and the one that survives the confounds below: **after you promote
|
|
23
|
+
a lesson to a rule, its failure signature should stop recurring.**
|
|
24
|
+
|
|
25
|
+
Borrowed from immune memory: promoting a lesson registers its original failure
|
|
26
|
+
signature as a "pathogen"; a later write bearing that same signature is a re-infection.
|
|
27
|
+
|
|
28
|
+
1. **At promotion, record the signature.** The lesson's **Applies when** line *is* the
|
|
29
|
+
signature — a file glob, a task type, a tool name, an error shape. Note the source
|
|
30
|
+
lesson's key and its `seen_count` at the moment of promotion.
|
|
31
|
+
2. **Watch for recurrence.** The direct measurement is the source lesson's `seen_count`
|
|
32
|
+
going **flat** — the loop stops writing that lesson because the failure stopped
|
|
33
|
+
happening. Read the column (`memory.read` / `lorekit show`), never a number in the
|
|
34
|
+
body.
|
|
35
|
+
3. **Escalate a breakthrough.** Any new write matching the promoted signature is a
|
|
36
|
+
vaccine-breakthrough case: the rule did not hold. That is a signal to fix the rule
|
|
37
|
+
(too narrow, wrong shape), not to shrug.
|
|
38
|
+
|
|
39
|
+
A promoted rule's value is the **silence** of its failure signature afterward. You
|
|
40
|
+
measure the absence. This needs no new product feature — `seen_count` and the
|
|
41
|
+
**Applies when** line already exist.
|
|
42
|
+
|
|
43
|
+
## Why the raw counters cannot prove causation
|
|
44
|
+
|
|
45
|
+
It is tempting to prove a loop "works" by pointing at rising counters. Do not lead with
|
|
46
|
+
this — LoreKit's own analysis records why the absolute counters are confounded:
|
|
47
|
+
|
|
48
|
+
- **`read_count` is ~99.8% bulk ride-alongs** — it tracks scope breadth, not whether a
|
|
49
|
+
lesson mattered.
|
|
50
|
+
- **`seen_count` is `1` for ~88% of rows** — most lessons never recur, so a raw count is
|
|
51
|
+
mostly noise.
|
|
52
|
+
- **Pull-through (`opened_count / read_count`) is the only trustworthy value signal** —
|
|
53
|
+
it cancels the ride-along confound by dividing it out.
|
|
54
|
+
|
|
55
|
+
And even pull-through is *correlation*: a lesson can pull through because it was
|
|
56
|
+
genuinely useful, or because the task was easy and any nudge looked right. Counter
|
|
57
|
+
movement alone can never separate "the lesson helped" from "the run would have
|
|
58
|
+
succeeded anyway." Only a run that deliberately did **not** get the memory can — see
|
|
59
|
+
[the north star](#the-north-star-a-withheld-memory-control-arm).
|
|
60
|
+
|
|
61
|
+
## The optional counter view (advisory only)
|
|
62
|
+
|
|
63
|
+
Once the immunity check is in place, the counters are a useful *secondary* read — never
|
|
64
|
+
the headline:
|
|
65
|
+
|
|
66
|
+
- **Pull-through per bucket** — `opened_count / read_count` across `loop::<host>-lessons`.
|
|
67
|
+
A bucket read often and opened rarely is dead weight to prune, not proof of value.
|
|
68
|
+
- **`/insights`** — the dashboard's Lore Utility grid classifies each lesson
|
|
69
|
+
(load-bearing / specialist / noise-tax / dormant) from pull-through against an
|
|
70
|
+
evidence floor. Consult it to find lessons to promote (load-bearing) or retire
|
|
71
|
+
(noise-tax), not to claim the loop improved a host.
|
|
72
|
+
- **`cited_count`** — when a host cites the lessons it applied (`cited` on the write),
|
|
73
|
+
a rising `cited_count` on a lesson is *evidence* it is being used. It is evidence,
|
|
74
|
+
never a denominator: a `0` means "nobody said so", not "unused".
|
|
75
|
+
|
|
76
|
+
State the confound wherever you report these: raw counters describe supply, not lift.
|
|
77
|
+
|
|
78
|
+
## Declare the metric at wiring time
|
|
79
|
+
|
|
80
|
+
Do this at step 6 of the quickstart, not months later when you have forgotten what
|
|
81
|
+
"better" meant:
|
|
82
|
+
|
|
83
|
+
- **Required:** name the failure signature (the **Applies when** line) whose recurrence
|
|
84
|
+
you will watch after promotion.
|
|
85
|
+
- **Optional:** name the one bucket pull-through you will glance at, and the interval.
|
|
86
|
+
- Write it into the host's own docs next to the loop's read/write steps, so a future
|
|
87
|
+
maintainer inherits the definition of "working" instead of re-inventing it.
|
|
88
|
+
|
|
89
|
+
## The north star: a withheld-memory control arm
|
|
90
|
+
|
|
91
|
+
The only *rigorous* attribution is a control arm: deterministically bucket a small
|
|
92
|
+
fraction of runs into a **lessons-suppressed** arm (reuse LoreKit's feature-flags
|
|
93
|
+
FNV-1a bucketing on a stable run id), then compare outcome rates — retry count,
|
|
94
|
+
first-try pass, wall-clock — between treated and control. That yields lift with a
|
|
95
|
+
confidence interval instead of a ratio.
|
|
96
|
+
|
|
97
|
+
This is the aspiration, **not a requirement**, for two honest reasons: it needs product
|
|
98
|
+
infrastructure that does not exist yet (a per-loop lift harness), and deliberately
|
|
99
|
+
withholding memory has a real UX/ethics cost, so it must be opt-in. Treat it as the
|
|
100
|
+
direction of travel; ship the immunity re-challenge today.
|
|
101
|
+
|
|
102
|
+
## Guards
|
|
103
|
+
|
|
104
|
+
- **Do not measure a loop you seeded.** Cold-start seeding erases the clean baseline the
|
|
105
|
+
immunity check reads — see [cold-start-seeding.md](./cold-start-seeding.md). Pick, per
|
|
106
|
+
loop: run-one value (seed it) OR provable lift (leave it cold).
|
|
107
|
+
- **A flat `seen_count` after promotion is proof only if the loop is still firing.**
|
|
108
|
+
Confirm the loop is alive first ([loop-health.md](./loop-health.md#is-the-loop-firing));
|
|
109
|
+
a dead loop also has a flat count, for the wrong reason.
|
|
110
|
+
- **Never gate anything on a counter.** These are proof for a human, not a machine gate.
|
|
111
|
+
The only mechanical gate is the compiled invariant ([compiled-invariants.md](./compiled-invariants.md)).
|
|
@@ -15,6 +15,7 @@ The design has two tiers connected by a recurrence gate. Both run on LoreKit.
|
|
|
15
15
|
|
|
16
16
|
## Contents
|
|
17
17
|
|
|
18
|
+
- [Find where a loop pays off](#find-where-a-loop-pays-off)
|
|
18
19
|
- [When to add a loop (and when not to)](#when-to-add-a-loop-and-when-not-to)
|
|
19
20
|
- [The two tiers](#the-two-tiers)
|
|
20
21
|
- [Conventions](#conventions)
|
|
@@ -30,6 +31,41 @@ The design has two tiers connected by a recurrence gate. Both run on LoreKit.
|
|
|
30
31
|
|
|
31
32
|
---
|
|
32
33
|
|
|
34
|
+
## Find where a loop pays off
|
|
35
|
+
|
|
36
|
+
Do not ask a newcomer to *guess* where memory helps — the best place for a loop is a
|
|
37
|
+
statistical property of the codebase's own history, and the data already knows it. Answer
|
|
38
|
+
"where?" before "how?".
|
|
39
|
+
|
|
40
|
+
**Start from recurring pain, not from a host you happen to be looking at.** A loop earns
|
|
41
|
+
its keep only where the same class of failure happens more than once. Two ways to surface
|
|
42
|
+
that:
|
|
43
|
+
|
|
44
|
+
- **From the store, if loops already exist:** `lorekit dedupe` clusters near-duplicate
|
|
45
|
+
lessons, and `lorekit invariants candidates` ranks clusters by summed `seen_count` ×
|
|
46
|
+
distinct scopes — literally "how often, in how many places, did this recur." A high-
|
|
47
|
+
ranked cluster with no dedicated host is a loop waiting to be wired.
|
|
48
|
+
- **From history, on a cold repo:** the same pattern fixed the same way 3+ times, a
|
|
49
|
+
revert-then-refix chain, a review note left across many PRs, or an existing CI guard —
|
|
50
|
+
each is a lesson someone already learned. `git log`, `git log --grep=revert`, and the
|
|
51
|
+
`obligations-map` are the seams. These double as seed sources —
|
|
52
|
+
[cold-start-seeding.md](./cold-start-seeding.md).
|
|
53
|
+
|
|
54
|
+
**Then match the host to an archetype** and copy its recipe card rather than wiring from
|
|
55
|
+
scratch:
|
|
56
|
+
|
|
57
|
+
| The host… | Card |
|
|
58
|
+
| --------- | ---- |
|
|
59
|
+
| edits code, fails in recurring ways | [code-changing-agent](../templates/code-changing-agent.md) |
|
|
60
|
+
| posts durable outputs at a target it revisits | [reviewer-reconcile-host](../templates/reviewer-reconcile-host.md) |
|
|
61
|
+
| is a pipeline that fails in classifiable ways | [multi-step-orchestrator](../templates/multi-step-orchestrator.md) |
|
|
62
|
+
| is a deterministic CI job needing last-run state | [ci-job](../templates/ci-job.md) |
|
|
63
|
+
|
|
64
|
+
The abstract test in the next section is the fallback when no recurrence data exists yet —
|
|
65
|
+
apply it, but prefer the evidence above when you have it.
|
|
66
|
+
|
|
67
|
+
---
|
|
68
|
+
|
|
33
69
|
## When to add a loop (and when not to)
|
|
34
70
|
|
|
35
71
|
Add a loop when **all** of these hold:
|
|
@@ -71,6 +107,11 @@ across runs** — recurrence is the cheap external signal that the lesson is rea
|
|
|
71
107
|
not a one-off. This is the episodic → procedural promotion path, with a human
|
|
72
108
|
gate on the slow tier so a single bad run can never rewrite the host.
|
|
73
109
|
|
|
110
|
+
Both tiers are **read by people as well as agents** — a teammate browsing the
|
|
111
|
+
dashboard, a reviewer asking why a rule exists. That is why a lesson body is
|
|
112
|
+
plain markdown with no hidden payload; see
|
|
113
|
+
[the lesson record](#the-lesson-record).
|
|
114
|
+
|
|
74
115
|
The fast tier is optional: if LoreKit's `memory.*` tools are not connected, the
|
|
75
116
|
loop is a silent no-op (log one line, continue). The slow tier is just editing
|
|
76
117
|
the host and is unaffected.
|
|
@@ -101,44 +142,139 @@ Reserve `branch::` for throwaway notes; a loop normally writes `global` or
|
|
|
101
142
|
### The lesson record
|
|
102
143
|
|
|
103
144
|
A loop lesson is **procedural** ("how to do better next time"), not a fact about
|
|
104
|
-
the user.
|
|
105
|
-
|
|
145
|
+
the user. The `value` is **markdown a human reads** — in the dashboard, in a
|
|
146
|
+
SessionStart injection, in another agent's context — and nothing else. Write it
|
|
147
|
+
to this shape:
|
|
106
148
|
|
|
107
149
|
```markdown
|
|
108
|
-
|
|
150
|
+
# <one-line takeaway — what to do, not what the lesson is about>
|
|
109
151
|
|
|
110
|
-
|
|
152
|
+
**Applies when:** <concrete signal — file glob, task type, tool name, error shape>
|
|
111
153
|
|
|
112
|
-
**What
|
|
154
|
+
**What happened:** <the concrete observable from the run>
|
|
113
155
|
**Why:** <root cause, if known; "unknown" is allowed>
|
|
114
|
-
**
|
|
156
|
+
**Do this instead:** <prescriptive, actionable, testable instruction>
|
|
115
157
|
**Promotion target:** <the host rule/step this would harden if promoted, or "none">
|
|
116
158
|
```
|
|
117
159
|
|
|
118
|
-
`
|
|
119
|
-
shapes) — never "when it feels relevant" — so the read step can match it
|
|
120
|
-
|
|
121
|
-
|
|
160
|
+
`Applies when` must be **concrete** (globs, task types, tool names, error
|
|
161
|
+
shapes) — never "when it feels relevant" — so the read step can match it against
|
|
162
|
+
the current run. It is a visible line, not hidden metadata: the reader deciding
|
|
163
|
+
whether a lesson is theirs needs it first, which is why it sits directly under
|
|
164
|
+
the title.
|
|
165
|
+
|
|
166
|
+
#### Never put machine metadata in the body
|
|
167
|
+
|
|
168
|
+
**A lesson body carries no hidden or machine-only payload — no HTML comment, no
|
|
169
|
+
front-matter block, no JSON blob, no `key=value` header.** Every byte of `value`
|
|
170
|
+
must render as prose a person can read. This is not a style preference; a body
|
|
171
|
+
field is the wrong home for each of these on the merits:
|
|
172
|
+
|
|
173
|
+
| Fact | Where it belongs | Why not the body |
|
|
174
|
+
| ---- | ---------------- | ---------------- |
|
|
175
|
+
| **Recurrence count** | `seen_count`, a store column | `memory_write` sets `seen_count = memories.seen_count + 1` on every overwrite. A count written into prose is a snapshot the writer guessed at and nothing ever updates — stale the first time the lesson recurs, and readers of the real counter never see it |
|
|
176
|
+
| **Expiry** | `ttl_days` on the write (`clear_ttl` to make permanent) | Expiry is enforced against the column by `memory.purge_expired` and the read filters. A date in prose expires nothing |
|
|
177
|
+
| **Status** (`structural`, `promoted`) | a `status::<value>` tag | Tags are first-class and filterable — `memory.list { tags: ["status::structural"] }` finds them. Prose is not queryable |
|
|
178
|
+
| **Owning host / bucket kind** | the `host` and `kind` write fields | Both are first-class, inferred from the `loop::<host>-lessons` tag when omitted, and drive the Explorer's own facets |
|
|
179
|
+
| **Provenance** (repo, branch, commit, PR) | `origin_repo` / `origin_branch` / `origin_commit` / `origin_pr` | First-class, and the dashboard renders them as links |
|
|
180
|
+
| **Trigger** (`stuck-loop`, `command-failure`, …) | the `trigger` write field | Already a facet; restating it in prose adds a line and no information |
|
|
181
|
+
|
|
182
|
+
An HTML comment is the worst of both worlds: markdown renders it to *nothing*, so
|
|
183
|
+
a human sees a lesson that starts mid-thought, while the fields inside it silently
|
|
184
|
+
disagree with the store's own columns. If a fact has a first-class home, put it
|
|
185
|
+
there and leave it out of the prose entirely.
|
|
186
|
+
|
|
187
|
+
This rule governs **lesson** bodies. A CI **state record** is a different shape
|
|
188
|
+
on purpose — its whole `value` is a JSON object, authoritative and parsed rather
|
|
189
|
+
than read, with no prose wrapped around it. See
|
|
190
|
+
[ci-state-records.md](./ci-state-records.md).
|
|
191
|
+
|
|
192
|
+
#### Writing rules
|
|
193
|
+
|
|
194
|
+
Enforce these on every write, autonomous or not:
|
|
195
|
+
|
|
196
|
+
- **One lesson per record.** Two takeaways are two keys.
|
|
197
|
+
- **Lead with the takeaway.** The `#` title is the instruction, not the topic —
|
|
198
|
+
"Pass `--node-modules-dir=none` to `deno check`", not "Notes on deno check".
|
|
199
|
+
- **Bold label, then one short paragraph.** No nesting past one list level, no
|
|
200
|
+
sub-headings; the whole record is read at a glance or not at all.
|
|
201
|
+
- **Fence every command or snippet**, with a language tag.
|
|
202
|
+
- **Budget ~1,500 characters.** A lesson longer than a screen is a document, and
|
|
203
|
+
a loop that injects documents crowds out the run it was meant to help. Cut the
|
|
204
|
+
narrative, keep the instruction.
|
|
205
|
+
- **No run residue** — no transcript excerpts, reasoning traces, session or
|
|
206
|
+
correlation IDs, timestamps, or "in this run I…" framing. A lesson is written
|
|
207
|
+
for the *next* run, which has none of that context.
|
|
208
|
+
- **No secrets or PII**, per the privacy pre-flight in the write step.
|
|
209
|
+
|
|
210
|
+
#### Worked example
|
|
211
|
+
|
|
212
|
+
❌ **Don't** — a hidden header, a title that names a topic, and a run narrative:
|
|
213
|
+
|
|
214
|
+
```markdown
|
|
215
|
+
<!-- meta: seen_count=1 status=active expires=2026-12-11 trigger-context="bash: line 5: dash0link: command not found" -->
|
|
216
|
+
# Notes on the Slack tool
|
|
217
|
+
**What failed:** In this run I called `slackSendMessage --args='{ … `dash0link` … }'`
|
|
218
|
+
and got `bash: line 5: dash0link: command not found`, then retried twice with
|
|
219
|
+
different escaping before it worked. Session 4b1c2efe, 2026-09-11.
|
|
220
|
+
```
|
|
221
|
+
|
|
222
|
+
Everything before the title renders to nothing for a human; `seen_count` and
|
|
223
|
+
`expires` contradict the store the moment the lesson recurs; the title says what
|
|
224
|
+
the lesson is *about* rather than what to do; and the run residue is dead weight
|
|
225
|
+
in every future context window.
|
|
226
|
+
|
|
227
|
+
✅ **Do** — the same lesson, all metadata in its own field:
|
|
228
|
+
|
|
229
|
+
````markdown
|
|
230
|
+
# Send Slack messages via a heredoc, not inline `--args='…'`
|
|
231
|
+
|
|
232
|
+
**Applies when:** shelling out to `slackSendMessage` with text interpolated from
|
|
233
|
+
an external source (ticket titles, PR titles) that may contain backticks or apostrophes.
|
|
234
|
+
|
|
235
|
+
**What happened:** Inline `--args='{ … }'` failed with `bash: dash0link: command not found` —
|
|
236
|
+
the shell expanded backticks in the JSON, and an apostrophe closed the quoted argument early.
|
|
237
|
+
**Why:** Single quotes do not protect the string once the outer shell re-processes it.
|
|
238
|
+
**Do this instead:** Pass the JSON through a quoted heredoc, which needs no escaping:
|
|
239
|
+
|
|
240
|
+
```bash
|
|
241
|
+
tools invoke slack.slackSendMessage --args="$(cat <<'ENDJSON'
|
|
242
|
+
{ "text": "…" }
|
|
243
|
+
ENDJSON
|
|
244
|
+
)"
|
|
245
|
+
```
|
|
246
|
+
|
|
247
|
+
**Promotion target:** the automation prompt's message-building step.
|
|
248
|
+
````
|
|
249
|
+
|
|
250
|
+
Written with `tags: ["loop::<host>-lessons", "source::command-failure"]`,
|
|
251
|
+
`trigger: "command-failure"`, and `ttl_days: 90` — so the recurrence count, the
|
|
252
|
+
expiry, the owning host and the trigger are all queryable, and the body is
|
|
253
|
+
readable start to finish.
|
|
122
254
|
|
|
123
255
|
---
|
|
124
256
|
|
|
125
257
|
## Read step (start of every run)
|
|
126
258
|
|
|
127
|
-
Read narrow-to-broad, filtered by the bucket tag, and merge
|
|
259
|
+
Read narrow-to-broad, filtered by the bucket tag, and merge — but **capped**. A loop
|
|
260
|
+
injects at most a small N (`limit: 5` per scope by default) into a run. This is a wiring
|
|
261
|
+
precondition, not a tuning knob: many uncapped loops on the same scopes tax every
|
|
262
|
+
session's context until agents learn to ignore injected lore entirely. Raise N only with
|
|
263
|
+
a reason.
|
|
128
264
|
|
|
129
265
|
```text
|
|
130
|
-
memory.list { scope: "repo::{owner}/{repo}", tags: ["loop::<host>-lessons"], limit:
|
|
131
|
-
memory.list { scope: "global", tags: ["loop::<host>-lessons"], limit:
|
|
266
|
+
memory.list { scope: "repo::{owner}/{repo}", tags: ["loop::<host>-lessons"], limit: 5 } # skips silently if memory.* not connected
|
|
267
|
+
memory.list { scope: "global", tags: ["loop::<host>-lessons"], limit: 5 }
|
|
132
268
|
# when the run names a subsystem / error, add a search:
|
|
133
|
-
memory.search { q: "<keywords>", scopes: ["repo::{owner}/*", "global"], limit:
|
|
269
|
+
memory.search { q: "<keywords>", scopes: ["repo::{owner}/*", "global"], limit: 5 }
|
|
134
270
|
```
|
|
135
271
|
|
|
136
272
|
Then:
|
|
137
273
|
|
|
138
|
-
1. Match each lesson's
|
|
139
|
-
matches.
|
|
140
|
-
|
|
141
|
-
2. Apply each matching
|
|
274
|
+
1. Match each lesson's **Applies when** line against the current run. Consider
|
|
275
|
+
only matches. Expired lessons do not come back — the store's own TTL drops
|
|
276
|
+
them, so the read never has to filter on a date in the prose.
|
|
277
|
+
2. Apply each matching **Do this instead** line as a **consideration**, not a
|
|
142
278
|
command — it biases the run unless it conflicts with the user's stated intent
|
|
143
279
|
or a task-specific constraint. On conflict, the user's intent wins; surface it.
|
|
144
280
|
3. On a `repo::` vs `global` collision, the `repo::` lesson wins (closer scope).
|
|
@@ -161,23 +297,31 @@ caught something, a near-miss, a guess that paid off. Not on smooth successes.
|
|
|
161
297
|
memory.search { q: "<key words of the lesson>", scopes: ["repo::{owner}/{repo}", "global"], limit: 10 }
|
|
162
298
|
```
|
|
163
299
|
|
|
164
|
-
3. **Write** to the classified scope
|
|
300
|
+
3. **Write** to the classified scope, putting every fact in its own field and
|
|
301
|
+
nothing but prose in `value`:
|
|
165
302
|
|
|
166
303
|
```text
|
|
167
304
|
memory.write {
|
|
168
|
-
scope:
|
|
169
|
-
key:
|
|
170
|
-
value:
|
|
171
|
-
tags:
|
|
172
|
-
trigger:
|
|
305
|
+
scope: "<global | repo::{owner}/{repo}>",
|
|
306
|
+
key: "<host>-lessons::<slug>",
|
|
307
|
+
value: "<the markdown lesson body above — no hidden blocks>",
|
|
308
|
+
tags: ["loop::<host>-lessons", "source::<trigger>"], # + "status::structural" when it is
|
|
309
|
+
trigger: "<stuck-loop | command-failure | gotcha | near-miss | assumption-wrong | paid-off | manual>",
|
|
310
|
+
ttl_days: 90
|
|
173
311
|
}
|
|
174
312
|
```
|
|
175
313
|
|
|
176
|
-
|
|
177
|
-
|
|
178
|
-
|
|
179
|
-
|
|
180
|
-
|
|
314
|
+
`host` and `kind` are inferred from the `loop::<host>-lessons` tag; pass them
|
|
315
|
+
explicitly only when the tag does not carry them. Add the `origin_*` fields
|
|
316
|
+
when the run knows them.
|
|
317
|
+
|
|
318
|
+
Same `scope` + `key` overwrites in place. **A recurrence resolves to an UPDATE:
|
|
319
|
+
the store increments `seen_count` by 1 for you, and re-passing `ttl_days`
|
|
320
|
+
refreshes the expiry** — that is what makes recurrence countable and drives
|
|
321
|
+
promotion. Never hand-write a count into the body to track this; the column is
|
|
322
|
+
the only copy that stays true. If a lesson you applied at the start of the run
|
|
323
|
+
worked (the failure did not recur), still write the UPDATE: successful
|
|
324
|
+
application is recurrence evidence.
|
|
181
325
|
|
|
182
326
|
The privacy pre-flight is never skipped, autonomous or not: a candidate lesson
|
|
183
327
|
containing a secret, token, credential, or PII is **dropped, not written**. The
|
|
@@ -381,8 +525,9 @@ hotspots during review), and every code-changing host — `aw`, `implement-sugge
|
|
|
381
525
|
|
|
382
526
|
After a read or write, a lesson is **promotion-eligible** when either:
|
|
383
527
|
|
|
384
|
-
- `seen_count >= 3` — the same failure recurred across at least three runs
|
|
385
|
-
|
|
528
|
+
- `seen_count >= 3` — the same failure recurred across at least three runs (read
|
|
529
|
+
the store's column, never a number written into the body), or
|
|
530
|
+
- it carries the `status::structural` tag because it reflects a design gap, not a
|
|
386
531
|
one-off.
|
|
387
532
|
|
|
388
533
|
For an eligible lesson, **surface a one-line suggestion — never act silently**:
|
|
@@ -394,8 +539,8 @@ The promotion target follows the lesson's scope: a `global` lesson hardens the
|
|
|
394
539
|
**host's own source** (every user of the host benefits); a `repo::` lesson
|
|
395
540
|
hardens the **repo's own rules / docs** (every teammate in that repo benefits).
|
|
396
541
|
Promotion is a normal, human-reviewed edit — LoreKit does not apply it. After a
|
|
397
|
-
successful promotion, write an UPDATE
|
|
398
|
-
stops re-suggesting and stands as an audit trail of why the rule exists.
|
|
542
|
+
successful promotion, write an UPDATE adding the `status::promoted` tag so the
|
|
543
|
+
lesson stops re-suggesting and stands as an audit trail of why the rule exists.
|
|
399
544
|
|
|
400
545
|
A recurring lesson can be promoted a second time, past the prose rule above,
|
|
401
546
|
into a **compiled invariant** — a declarative, mechanically-checked assertion
|
|
@@ -420,9 +565,10 @@ guards are what make the loop safe:
|
|
|
420
565
|
run; it can never silently disable a gate, skip a step, or change a limit. The
|
|
421
566
|
only path from a lesson to changed behavior is the human-reviewed slow tier.
|
|
422
567
|
2. **Recurrence gates promotion, not a single run** (`seen_count >= 3`, or an
|
|
423
|
-
explicit `status
|
|
424
|
-
3. **Every lesson expires**
|
|
425
|
-
|
|
568
|
+
explicit `status::structural` tag on the lesson).
|
|
569
|
+
3. **Every lesson expires** — `ttl_days: 90` on the write, refreshed on each
|
|
570
|
+
recurrence, so stale beliefs decay instead of entrenching. The store enforces
|
|
571
|
+
this; an expiry stated only in prose is decoration.
|
|
426
572
|
4. **Contradiction is surfaced, not silently overwritten** — the dedup search
|
|
427
573
|
finds the prior lesson; a genuine reversal is a reviewed decision.
|
|
428
574
|
5. **The privacy pre-flight is never bypassed** — secrets / PII are dropped, not
|
|
@@ -436,13 +582,14 @@ To add a loop to a host called `<host>`:
|
|
|
436
582
|
|
|
437
583
|
- [ ] Pick the bucket: tag `loop::<host>-lessons`, key `<host>-lessons::<slug>`.
|
|
438
584
|
- [ ] Add the **read step** at the start of the host's run (narrow-to-broad
|
|
439
|
-
`memory.list` filtered by the tag; apply matches as considerations
|
|
440
|
-
expired).
|
|
585
|
+
`memory.list` filtered by the tag; apply matches as considerations).
|
|
441
586
|
- [ ] Add the **write step** at the host's existing failure / end-of-run points
|
|
442
|
-
(classify scope, `memory.search` to dedup, `memory.write`).
|
|
443
|
-
reflection stage — hook the points the host already detects.
|
|
587
|
+
(classify scope, `memory.search` to dedup, `memory.write` with `ttl_days`).
|
|
588
|
+
No new reflection stage — hook the points the host already detects.
|
|
589
|
+
- [ ] State the **body contract** in the host's own write step: markdown to the
|
|
590
|
+
shape above, no hidden blocks, every store-backed fact in its own field.
|
|
444
591
|
- [ ] Add the **promotion suggestion** when a read/written lesson hits
|
|
445
|
-
`seen_count >= 3` or `status
|
|
592
|
+
`seen_count >= 3` or carries `status::structural`.
|
|
446
593
|
- [ ] State the **entrenchment guards** so a future maintainer does not "optimize
|
|
447
594
|
them away".
|
|
448
595
|
- [ ] Confirm the loop **degrades silently** when `memory.*` is not connected.
|
|
@@ -450,6 +597,14 @@ To add a loop to a host called `<host>`:
|
|
|
450
597
|
read** at its plan/apply seam (match `hotspot::<path>` /
|
|
451
598
|
`knowledge::<symbol>@<path>` to the files it will change). If it **verifies**
|
|
452
599
|
a structural fact, wire the **write** behind that section's contract.
|
|
600
|
+
- [ ] Set the **injection cap** (`limit: 5` per scope by default) on the read step —
|
|
601
|
+
a wiring precondition, not a tuning knob.
|
|
602
|
+
- [ ] Decide **seeding**: value on run one ([cold-start-seeding.md](./cold-start-seeding.md),
|
|
603
|
+
a handful, curated) **or** provable lift (leave it cold) — never both.
|
|
604
|
+
- [ ] Declare the **proof** — the immunity re-challenge whose failure signature you will
|
|
605
|
+
watch after promotion ([proving-improvement.md](./proving-improvement.md)).
|
|
606
|
+
- [ ] Add the **health checks** — a canary, the write-time matchability check, and the
|
|
607
|
+
quarantine-not-delete rollback ([loop-health.md](./loop-health.md)).
|
|
453
608
|
|
|
454
609
|
---
|
|
455
610
|
|