akm-cli 0.9.26 → 0.9.27-alpha.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (39) hide show
  1. package/CHANGELOG.md +237 -0
  2. package/LICENSE +3 -4
  3. package/dist/assets/prompts/consolidate-system.md +2 -2
  4. package/dist/assets/prompts/distill-lesson-system.md +29 -7
  5. package/dist/assets/prompts/extract-session.md +2 -2
  6. package/dist/commands/improve/consolidate/coverage.js +71 -17
  7. package/dist/commands/improve/consolidate/pair-pass.js +12 -9
  8. package/dist/commands/improve/consolidate.js +77 -5
  9. package/dist/commands/improve/distill-guards.js +9 -10
  10. package/dist/commands/improve/distill.js +70 -22
  11. package/dist/commands/improve/extract-prompt.js +61 -39
  12. package/dist/commands/improve/extract.js +2 -1
  13. package/dist/commands/improve/preparation.js +14 -1
  14. package/dist/commands/improve/reflect.js +36 -4
  15. package/dist/commands/improve/retrieval-gate.js +1 -1
  16. package/dist/commands/improve/session-asset.js +3 -2
  17. package/dist/commands/improve/stage.js +35 -53
  18. package/dist/commands/proposal/drain.js +85 -21
  19. package/dist/commands/proposal/proposal-types.js +1 -1
  20. package/dist/core/config/schema/improve-processes.js +3 -3
  21. package/dist/core/paths.js +0 -9
  22. package/dist/indexer/indexer.js +35 -19
  23. package/dist/integrations/harnesses/codex/agent-builder.js +23 -15
  24. package/dist/llm/client.js +44 -14
  25. package/dist/llm/memory-infer.js +1 -1
  26. package/dist/scripts/akm-migrate-node.js +25 -20
  27. package/dist/scripts/akm-migrate.js +25 -20
  28. package/dist/storage/repositories/improve-ledger-repository.js +4 -1
  29. package/dist/storage/repositories/index-entry-schema.js +20 -4
  30. package/dist/storage/repositories/index-fts-repository.js +44 -3
  31. package/dist/storage/repositories/index-schema.js +14 -8
  32. package/docs/README.md +1 -2
  33. package/docs/integration/bundling-akm.md +1 -1
  34. package/docs/migration/v0.8-to-v0.9.md +3 -1
  35. package/docs/reference/README.md +1 -1
  36. package/docs/reference/cli.md +8 -4
  37. package/docs/reference/configuration.md +5 -2
  38. package/docs/reference/data-and-telemetry.md +1 -1
  39. package/package.json +1 -1
package/CHANGELOG.md CHANGED
@@ -6,6 +6,243 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/).
6
6
 
7
7
  ## [Unreleased]
8
8
 
9
+ ## [0.9.27-alpha.2] - 2026-10-07
10
+
11
+ ### Removed
12
+
13
+ - **The `scripts/akm-eval` toolkit has moved out of this repository.** The
14
+ read-only measurement toolkit (the case runner and its suites, the twin
15
+ experiment, the real-query verdict for the proactive lane, the state
16
+ analyzers and the curate benchmark) is retired; every live eval is in
17
+ [itlackey/akm-eval](https://github.com/itlackey/akm-eval). Its code is kept
18
+ there, to read and not to run, in `retired/akm-scripts-akm-eval/`, copied from
19
+ commit `f57a7fd44b37`. It imports akm's `src/` by relative path, so it runs
20
+ only in a checkout at that commit. Removed here with it: `scripts/akm-eval/`,
21
+ its tests (`tests/integration/akm-eval/`, `tests/akm-eval-*.test.ts`,
22
+ `tests/curate-metrics.test.ts`) and fixtures (`tests/fixtures/akm-eval/`, and
23
+ the `curate-golden` stash, which only the curate benchmark read), the
24
+ `akm-eval determinism` CI job, and `getMeasurementVerdictsDir`, whose only
25
+ caller was the verdict runner. akm no longer names
26
+ `$STATE/improve/measurement/verdicts/<stash>/`; a file already there is inert.
27
+ `docs/maintainers/eval.md` is now a pointer to the new home.
28
+
29
+ ### Fixed
30
+
31
+ - **The drain's judge no longer sees a note with a code block as truncated.**
32
+ The judgment prompt fenced the proposed content (and the live asset, sibling
33
+ proposals and neighbour excerpts) in three backticks, so a note holding its
34
+ own code block closed the fence early and read as cut off; real rejections said
35
+ "ends in an empty code block" or "truncated". Each block now uses a fence longer
36
+ than any backtick run inside it. The judge's reason is also kept on accepts,
37
+ staged accepts and defers (as the gate decision's `judgeReason`, until now
38
+ rejections only), and a judge reply that is not a verdict is stamped
39
+ `judgment-parse-failure`, and a runner failure `judgment-error`, instead of
40
+ looking like a defer.
41
+
42
+ - **Distill writes a lesson only when its memory holds one, says only what the
43
+ memory says, and its judge rejects what a reviewer would.** 2 of the 22 distill
44
+ proposals since 0.9.26 began were accepted, and 17 of the 19 queued on
45
+ 2026-10-05 were bad (they restated their memory, filed a dated status as a
46
+ lesson, claimed what the memory does not say, or repeated an asset the library
47
+ holds). Four causes, found in the code and the rejected proposals, and fixed:
48
+ (1) the prompt and schema forced a lesson from every memory, and 18 of the 19
49
+ were records of what was done; the writer now says why a memory holds a
50
+ lesson or none (`reason`, then `decision: lesson|none`, or the word `NONE`),
51
+ defined as a cause and what to do about it, or a rule with its reason, and
52
+ writes only what the memory and its feedback state, in the scope they have; a
53
+ `NONE` is a `skipped` distill (`skipReason: nothing_reusable`, the writer's
54
+ reason in the message) with no proposal and no judge call, and the loop keeps
55
+ its ledger row `unchanged`. (2) The judge asked for "information not already
56
+ present in the source", so an invented claim scored as novel and a faithful
57
+ lesson of a lesson-worthy memory as a restatement, and it passed anything that
58
+ "goes beyond the source" because it "may draw on feedback you are not shown".
59
+ The rubric is now reusable (a rule with its reason, not a record of what was
60
+ done), non-redundancy and grounding (every cause, step, number and limit is
61
+ in the source or its feedback), and the judge is shown the feedback the writer
62
+ saw. (3) A mean hid a decisive score (4 and 1 average 2.5, a review), and a
63
+ reviewer read everything the judge did not reject; any criterion at 2 or
64
+ below, grounding included, is now `quality_rejected`, and the reason names it
65
+ (`grounding 2/5: …`). The "borderline grounding" routing is gone. (4) Neither
66
+ the writer nor the judge could see a knowledge note or a skill that already
67
+ states the rule (the judge saw the 3 lexically nearest lessons, none of them
68
+ related); both now see the lessons, knowledge notes and skills nearest the
69
+ memory, which is the existing `processes.distill.cls` context turned on by
70
+ default (`enabled: false` turns it off). Judge scores are keyed `reusable`
71
+ where they were `novelty`. Measured on the local qwen3.8-27b with akm-eval's
72
+ `evals/distill` (30 fictional memories, 5 runs each side): good lessons 7/14 on
73
+ average (5 to 9) against 3.7/14 (3 to 5), lessons queued for memories that
74
+ deserve none 0.4/16 against 3.3/16. On 37 real memories with their feedback
75
+ (34 reviewed bad, 3 good; 3 runs against 2): a lesson was queued for 11% of the
76
+ bad ones against 44%, and for 6 of 9 good ones against 4 of 6; of the memories
77
+ that pass 0.9.26's skip of bare positive feedback, 20% of the bad against 50%.
78
+ No new settings.
79
+
80
+ - **Reflect no longer plans an asset whose negative feedback is already acted
81
+ on.** A negative `akm feedback` that came with an exact fix (`--replace` and
82
+ `--with`, `--outdated` or `--superseded-by`) makes a `feedback` proposal, and
83
+ once that proposal is accepted the feedback has done its work. Reflect still
84
+ took the ref as having fresh negative feedback, and on 2026-10-07 44 of its 50
85
+ refs were of that kind: the judge refused or the model changed nothing for
86
+ most of them. A negative event with a fix is now left out of the reflect
87
+ cursor when an accepted `feedback` proposal for the ref was created at or
88
+ after it. A negative with no fix, one given after the proposal, and one whose
89
+ proposal is still pending or was rejected plan a reflect as before.
90
+
91
+ - **Consolidate stops re-offering memories a reviewer already turned down, and
92
+ the nightly judge sees what a promotion may duplicate.** About 53 promotions a
93
+ night reached review at ~5% precision, 64-70% of them a memory body already
94
+ proposed or rejected. Four causes, four changes: a memory whose body equals
95
+ that of a consolidate promotion rejected on or after 2026-09-29 is held until
96
+ its body changes, under any name (earlier rejections, the bulk audits of
97
+ 2026-08, do not count); a memory the model judged and left alone is held by its
98
+ body hash instead of a 7-day clock, so an unchanged memory is no longer judged
99
+ every week (a row recorded without a hash keeps the 7 days); the coverage gate
100
+ skips a memory when 30% of its text, not 50%, is in a neighbouring knowledge
101
+ doc, which catches paraphrases; and the drain's judgment tier, which judged a
102
+ promotion seeing only the proposal and never `knowledge/`, is now shown the 5
103
+ nearest knowledge notes (ref, description, excerpt) and told to reject a
104
+ promotion they already cover. No new settings.
105
+
106
+ - **A confident `subsumed` or `supersedes` retirement resolves unattended, as a
107
+ `duplicate` already did.** The pair pass staged a retire proposal for the
108
+ triage drain only when the judge's label was `duplicate`; every other retirement
109
+ waited for a person. It now stages any of the three retire labels when the
110
+ second look (what does the retired note hold that the kept one lacks?) comes
111
+ back empty and there is no continuity risk; the retired side's claim list is
112
+ already empty for any proposal, and a `duplicate` must still have an empty list
113
+ on the kept side too. The staged gate reason is the judge's label, and the drain records it. Replay over 360
114
+ judged pairs: 336 safe (0.93): `duplicate` 0.98, `subsumed` 0.92, `supersedes`
115
+ 0.875. In production, unstaged `subsumed` retirements were accepted 45 of 54
116
+ times by hand, and in the latest run 41 of 47 pair proposals would have resolved
117
+ without a person. `docs/architecture/internals/improve-workflow.md` said triage
118
+ never auto-accepts a retire proposal, which stopped being true in 0.9.26; it,
119
+ and the matching lines in `improvement.md`, now describe the staging rule.
120
+
121
+ ## [0.9.27-alpha.1] - 2026-10-06
122
+
123
+ ### Fixed
124
+
125
+ - **A codex dispatch with an output schema no longer leaves a temp folder
126
+ behind.** Every build of the codex command for a request with a schema made a
127
+ new `akm-codex-schema-*` folder in the OS temp dir for `--output-schema` and
128
+ nothing ever removed it (the builder has no post-run hook, and the file is read
129
+ after it returns), so a machine whose `/tmp` is tmpfs held a folder in RAM per
130
+ dispatch until reboot. The schema is now written once to akm's cache dir, in a
131
+ file named by its hash: concurrent units dispatching the same schema share it,
132
+ a rewrite is an atomic rename of identical bytes, and the only residue is one
133
+ small file per distinct schema.
134
+ - **An index that has been updated ranks like a fresh index of the same files,
135
+ and `akm index --full` no longer doubles the full-text totals.** `entries_fts`
136
+ is contentless, and FTS5 cannot take a deleted row out of a contentless
137
+ table's BM25 totals (its row count, and the token counts the average document
138
+ length comes from). Every replaced or removed row left them one row too high:
139
+ 40 notes read 40, then 41 after one edit, then 81 after `--full`, and a delete
140
+ never lowered them, so scores drifted away from what a fresh index gives.
141
+ SQLite has no command that recomputes them (`delete` and `rebuild` are refused
142
+ on a contentless table), so a delete that removes a row now stamps
143
+ `index_meta.ftsTotalsStale` in its own transaction, whichever process made it,
144
+ and the next `akm index` rebuilds the table from `entries` before it finishes:
145
+ about a second at 25,000 entries, and only when rows have left the table. A
146
+ row the write-path index replaced after an accepted proposal is settled the
147
+ same way, and an index that has already drifted is corrected by the first run
148
+ that replaces or removes a row.
149
+ - **`akm improve` runs against an API that rejects `chat_template_kwargs`,
150
+ OpenAI's among them.** Improve's reflect, consolidate and judge calls always
151
+ ask for thinking off, and the client sends that as
152
+ `chat_template_kwargs.enable_thinking` and a top-level `enable_thinking`. A
153
+ strict API answers 400 `Unknown parameter: 'chat_template_kwargs'`, the retry
154
+ without the response schema sent both fields again, and `akm improve judge`
155
+ reported `judge timeout/error — routed to review`. No engine setting could
156
+ stop it: the call sites override the engine's `enableThinking`, and
157
+ `extraParams` can only add fields. A 4xx that names either field is now
158
+ answered by one retry without both (which may in turn fall back without the
159
+ schema), and akm stops sending them to that endpoint and model for the rest
160
+ of the process, as it already does for `response_format`.
161
+ - **A lesson the model wrote without a `when_to_use` is no longer thrown away
162
+ silently by `akm proposal extract`.** The extract schema left `when_to_use`
163
+ optional while the parser dropped a lesson without one (or with one under 15
164
+ characters) and said nothing, so a model that followed the schema could
165
+ write a sound lesson that akm discarded and the session reported no
166
+ candidates: on akm-eval's extract eval qwen3.8-27b kept 3 of 11 expected
167
+ insights against 9 for gpt-oss-120b. Every property of the schema is now
168
+ required, `when_to_use` and `rationale_if_empty` included, with an empty
169
+ string for none (a memory or knowledge candidate needs no trigger, a
170
+ non-empty answer no rationale), which is also what a strict structured-output
171
+ provider needs; the prompt's output contract says the same. Any candidate the
172
+ contract still refuses, for this reason or another, is named in its session's
173
+ `warnings` as `<type>:<name> dropped: <reason>`.
174
+ - **Distill's response schemas are valid for a strict structured-output
175
+ provider.** The client sends a response schema `strict: true`, and OpenAI
176
+ refuses one whose objects leave a property out of `required`: `400 Invalid
177
+ schema for response_format 'akm_response': ... Missing 'tags'`. The lesson
178
+ schema left out `tags`, and the knowledge schema `tags` and `sources`, so
179
+ the first distill request on such a provider always failed. After a 4xx the
180
+ client retries once without the schema, but a gateway that answers the same
181
+ rejection with a 502 is not retried, and every distill call through it
182
+ failed; the workaround, `supportsJsonSchema: false`, loses the guidance that
183
+ keeps a model from leaving out `when_to_use`. Every property is now
184
+ required, and an empty array stands for none (distill already dropped an
185
+ empty `tags` or `sources`).
186
+ - **`akm improve` no longer reflects on an asset whose file a proposal would not
187
+ write, so accepting a reflect proposal no longer adds a second file.** A
188
+ skill's `references/a.md` is indexed as `knowledge/skills/<name>/references/a`,
189
+ but a proposal writes the path derived from that ref,
190
+ `knowledge/skills/<name>/references/a.md`. Nothing is there, so the proposal
191
+ was a `create`, and accepting it wrote a copy beside the skill's own file;
192
+ later proposals then revised the copy while the skill's file drifted. Reflect
193
+ now refuses before it calls the model, naming the file and the path a
194
+ proposal would write, and the loop records it as a skip (`unsupported_type`,
195
+ `file_outside_layout` in the `reflect_completed` event). This is the rule
196
+ 0.9.26 added to `akm feedback --replace`. An asset that another bundle owns is
197
+ still refused by `createProposal` (#1000).
198
+ - **An index that 0.9.1 wrote no longer crashes akm.** Every `akm index` on it
199
+ exited 70 with `null is not an object (evaluating 'doc.xrefs')`, `akm migrate
200
+ apply` and `akm index --full` did not help, and `akm search` failed with
201
+ `null is not an object (evaluating 'item.entry.quality')`; the only way out
202
+ was moving `index.db` aside. That layout (20) keeps the transitional
203
+ `entry_key`, `dir_path`, `stash_dir`, `entry_json` and `entry_type` columns
204
+ that layout 21 removed, each NOT NULL, beside the current columns, and leaves
205
+ `document_json` NULL on every row. The table had every column akm checks for,
206
+ so it was taken for a current one: the links migration read the NULL
207
+ documents, and no insert could ever have succeeded (`NOT NULL constraint
208
+ failed: entries.entry_key`). akm now treats a table that still has a retired
209
+ column as older than layout 21, as the compat notes already said: the
210
+ writable opener recreates its entries-keyed tables (the LLM enrichment cache
211
+ is kept) and the next `akm index` re-walks every source, and a read rebuilds
212
+ inline. An index that a failed open already half-migrated (layout stamp still
213
+ 20, `search_text` dropped, `asset_links` created) recovers the same way.
214
+ - **Two files that claim one ref no longer trade places in the index, and `akm
215
+ index` says so.** A skill's `references/a.md` and a note at
216
+ `knowledge/skills/x/references/a.md` are both the ref
217
+ `knowledge/skills/x/references/a`, and the index holds one row for it. The
218
+ first file a run persisted held it, so a full build followed the filesystem's
219
+ listing order (one bundle indexed on tmpfs and on ext4 held different files),
220
+ and the first incremental run after a full build handed the row to the other
221
+ file, because it drains only the directory that lost: with no file touched,
222
+ the row, its search entry and the text its vector is embedded from changed,
223
+ and the vector was dropped and recomputed. When a smaller-path file was added
224
+ later and then deleted, the ref also left the index until `--full`, although
225
+ its other file was still on disk. The file with the smaller path (code-point
226
+ order, as `akm show`'s refusal lists them) now holds the ref however the
227
+ directories are drained and the walk is ordered, a directory that gives a ref
228
+ up is drained again so the ref passes back when its holder goes, and each
229
+ pair is reported in the `warnings` of `akm index`, naming the file indexed
230
+ and the one skipped.
231
+ - **Consolidate's plan schema and the session summary schema are valid for a
232
+ strict structured-output provider, and a new response schema can no longer
233
+ skip the rule.** The client sends a response schema `strict: true`, and
234
+ OpenAI refuses one whose objects leave a property out of `required`. The
235
+ consolidate plan left out `description` and `confidence`, and the session
236
+ summary `tags`, so the first request of every plan and every summary to such
237
+ a provider was refused. The client then retries without the schema and
238
+ remembers that per connection (endpoint and model), not per schema, so one
239
+ invalid schema also switched the response schema off for every valid one
240
+ that followed on that connection. Every property of both is now required: an
241
+ empty `description` keeps the memory's own, a null `confidence` records none
242
+ and an empty `tags` array is no tags, which is how all three were already
243
+ read. A contract test runs every schema akm sends through the rule, so a
244
+ property added later without being required fails CI.
245
+
9
246
  ## [0.9.26] - 2026-10-05
10
247
 
11
248
  The stable release of the 0.9.26 line: 0.9.26-alpha.1 and alpha.2, and the
package/LICENSE CHANGED
@@ -225,10 +225,9 @@ statute, judicial order, or regulation then You must: (a) comply with
225
225
  the terms of this License to the maximum extent possible; and (b)
226
226
  describe the limitations and the code they affect. Such description must
227
227
  be placed in a text file included with all distributions of the Covered
228
- Software under the name "LEGAL", with additions for new restrictions
229
- placed at the end of the file. Except to the extent prohibited by
230
- statute or regulation, such description must be sufficiently detailed
231
- for a recipient of ordinary skill to be able to understand it.
228
+ Software under this License. Except to the extent prohibited by statute
229
+ or regulation, such description must be sufficiently detailed for a
230
+ recipient of ordinary skill to be able to understand it.
232
231
 
233
232
  5. Termination
234
233
  --------------
@@ -7,10 +7,10 @@ Rules:
7
7
  Return ONLY JSON (no prose, no code fences):
8
8
  {
9
9
  "operations": [
10
- { "op": "promote", "ref": "memories/<name>", "knowledgeRef": "knowledge/<suggested-slug>", "reason": "<brief reason>", "description": "<one sentence describing the new knowledge asset>", "confidence": 0.92 }
10
+ { "op": "promote", "ref": "memories/<name>", "knowledgeRef": "knowledge/<suggested-slug>", "reason": "<brief reason>", "description": "<one sentence describing the new knowledge asset, or an empty string to keep the memory's own>", "confidence": 0.92 }
11
11
  ]
12
12
  }
13
13
 
14
- For every operation, emit a `confidence` field in [0, 1] expressing your certainty that the operation is correct and safe. Use 0.95+ only when evidence is unambiguous. Omit the field rather than guessing if you are uncertain.
14
+ For every operation, emit a `confidence` field in [0, 1] expressing your certainty that the operation is correct and safe. Use 0.95+ only when evidence is unambiguous. Use `null` rather than guessing if you are uncertain.
15
15
 
16
16
  When the merged content includes an `updated` frontmatter field, the value MUST be a real ISO date string (e.g. `updated: 2026-05-20`). NEVER emit `updated: today`, `updated: {today}`, `updated: {today: null}`, `updated: now`, or any other literal placeholder/template-variable. If you do not have a real source-of-truth date, OMIT the `updated` field entirely — the post-processor will not invent one for you.
@@ -1,7 +1,29 @@
1
1
  You are the akm `distill` distiller.
2
- Given an asset and recent feedback events about it, produce a single
3
- concise *lesson* an agent should remember next time it works on this
4
- asset's domain.
2
+ You are given a memory and the feedback recorded about it. Decide whether it
3
+ holds a lesson and, if it does, write the lesson.
4
+
5
+ A memory holds a lesson when it states a cause and what to do about it: a
6
+ failure or surprise with its cause and the fix that worked, or a rule with the
7
+ reason it holds. It holds a lesson even when it is short and names one project,
8
+ tool or incident, if the cause and the fix would help someone in a similar
9
+ situation.
10
+
11
+ A memory holds NO lesson when all it states is what was done, shipped, released
12
+ or decided, what is pending or planned, how a system is set up now, or the steps
13
+ of a procedure, with no failure and cause behind it, or when its feedback says
14
+ only that it is out of date or superseded. ANSWER NONE then: the single word and
15
+ nothing else. A reply bound to a JSON schema answers NONE by
16
+ setting `decision` to `none` and leaving the other fields empty. Answer NONE too
17
+ when a related asset listed below the memory already states the rule the memory
18
+ would give.
19
+
20
+ When the memory holds a lesson, write it from what the memory and its feedback
21
+ say, and nothing more:
22
+ - Add no cause, step, rule, number, check or safeguard that neither states.
23
+ - Keep the scope the memory has. A fix verified in one place is a fix for that
24
+ place, and what was not checked stays unchecked. One case is not "always" or
25
+ "never".
26
+ - Be shorter than the memory.
5
27
 
6
28
  YOUR RESPONSE MUST START EXACTLY WITH `---` ON THE VERY FIRST LINE.
7
29
  DO NOT output any prose, explanation, or code fences before or after.
@@ -12,7 +34,7 @@ description: <one complete sentence (ending with `.`) summarising what the lesso
12
34
  when_to_use: <one complete sentence describing the concrete trigger condition>
13
35
  ---
14
36
 
15
- <lesson body — plain markdown, 1–3 short paragraphs of practical guidance>
37
+ <lesson body — plain markdown, as short as the memory allows>
16
38
 
17
39
  ## description field (MANDATORY)
18
40
  - A single complete sentence in present tense, 20–400 chars, NO markdown.
@@ -21,7 +43,7 @@ when_to_use: <one complete sentence describing the concrete trigger condition>
21
43
  - DO NOT copy a section heading ("Key takeaways", "For example", "Key pitfalls").
22
44
  - DO NOT begin with a numbered list marker, code fence, or markdown heading.
23
45
 
24
- GOOD: "Always validate ref existence before promoting a memory to knowledge; missing refs surface as silent 404s during accept."
46
+ GOOD: "Pin the container image tag, because the `latest` tag moved under the nightly job and its output changed with no code change."
25
47
  BAD: "Key pitfalls"
26
48
  BAD: "When working with the akm CLI"
27
49
  BAD: "For example, you might..."
@@ -32,5 +54,5 @@ RULES:
32
54
  - `description` and `when_to_use` MUST differ from each other.
33
55
  - The lesson body MUST be non-empty markdown prose. Do NOT restate `description:` or `when_to_use:` inside the body (no `**description:** ...` or `**when_to_use:** ...` lines — the frontmatter is the only place those keys belong).
34
56
  - Do NOT emit a second `---` fence after the opening frontmatter — there are exactly two `---` lines in the output, both belonging to the single frontmatter block at the top.
35
- - Do NOT reproduce the source asset verbatim — distil what a caller needs to know.
36
- - Output ONLY the lesson file. No preamble, no code fences, no trailing prose.
57
+ - Do NOT reproduce the source asset verbatim.
58
+ - Output ONLY the lesson file. No preamble, no code fences, no trailing prose.
@@ -48,13 +48,13 @@ Respond with EXACTLY one JSON object matching this shape:
48
48
  "type": "memory" | "lesson" | "knowledge",
49
49
  "name": "<kebab-case name, e.g. jwt-token; optionally under one kebab-case scope, e.g. auth/jwt-token>",
50
50
  "description": "<one sentence 20-400 chars>",
51
- "when_to_use": "<one sentence 15-400 chars; REQUIRED only when type=lesson>",
51
+ "when_to_use": "<one sentence 15-400 chars for a lesson; an empty string for a memory or knowledge candidate>",
52
52
  "body": "<markdown body, 200-3000 chars typical>",
53
53
  "confidence": <number 0.0-1.0>,
54
54
  "evidence": "<one-line pointer to the moment in the session>"
55
55
  }
56
56
  ],
57
- "rationale_if_empty": "<one sentence; REQUIRED when candidates is empty>"
57
+ "rationale_if_empty": "<one sentence when candidates is empty; an empty string otherwise>"
58
58
  }
59
59
  ```
60
60
 
@@ -6,7 +6,7 @@
6
6
  * `knowledge/` proposal, ask whether `knowledge/` already says it.
7
7
  *
8
8
  * The rule: a memory is covered when at least {@link COVERAGE_MIN_CONTAINMENT}
9
- * (half) of its distinct {@link COVERAGE_SHINGLE_WORDS}-word shingles appear in
9
+ * (30%) of its distinct {@link COVERAGE_SHINGLE_WORDS}-word shingles appear in
10
10
  * one of the knowledge docs nearest to it. It is a containment of the MEMORY in
11
11
  * the doc, not a similarity: a long guide that quotes the memory covers it; a
12
12
  * memory that quotes a short doc and adds claims of its own does not.
@@ -25,9 +25,13 @@
25
25
  * this gate achieves: it reads only the {@link PAIR_NEIGHBOR_FETCH_K} nearest
26
26
  * knowledge docs (below), and a covering doc that ranks lower goes unseen. The
27
27
  * recall over that candidate set is unmeasured. The rest of the rejected
28
- * proposals (paraphrases, partial overlaps) still reach review: the next cut
29
- * measured, 0.2 (169 of 224 rejected), would also have skipped 2 of the 103
30
- * accepted ones, and no cosine cut was measured at all. A wrong skip is a
28
+ * proposals (paraphrases, partial overlaps) still reach review. 0.5 let
29
+ * paraphrases through: of the 54 promotions minted on 2026-10-07, 18 of the 51
30
+ * later rejected held 30% or more of their text in a neighbouring doc. The cut is
31
+ * now 0.3, between the two measured points: 0.2 (169 of 224 rejected) also
32
+ * skipped 2 of the 103 accepted ones. At 0.3, none of the 5 promotions graded
33
+ * good in the 2026-10-05 review sample would be skipped (their best doc holds
34
+ * at most 1% of them) while 5 of its 15 bad ones would. A wrong skip is a
31
35
  * promotion nobody gets to review.
32
36
  *
33
37
  * Candidates are the memory's {@link PAIR_NEIGHBOR_FETCH_K} nearest knowledge
@@ -49,7 +53,7 @@ import { PAIR_NEIGHBOR_FETCH_K } from "./pair-pass.js";
49
53
  /** Words per shingle. */
50
54
  export const COVERAGE_SHINGLE_WORDS = 5;
51
55
  /** Share of a memory's distinct shingles one knowledge doc must hold for the memory to count as covered. */
52
- export const COVERAGE_MIN_CONTAINMENT = 0.5;
56
+ export const COVERAGE_MIN_CONTAINMENT = 0.3;
53
57
  const WORD = /[\p{L}\p{N}]+/gu;
54
58
  /** The distinct lower-cased word n-grams of `text`; empty when it has fewer than {@link COVERAGE_SHINGLE_WORDS} words. */
55
59
  export function wordShingles(text) {
@@ -72,19 +76,15 @@ export function shingleContainment(memory, doc) {
72
76
  return shared / memory.size;
73
77
  }
74
78
  /**
75
- * The best-covering knowledge doc among the {@link PAIR_NEIGHBOR_FETCH_K}
76
- * knowledge docs in `bundleId` nearest to the memory. `filePath` is the
77
- * memory's indexed file; a memory the index does not know has no stored vector
78
- * and so no candidates.
79
+ * The {@link PAIR_NEIGHBOR_FETCH_K} knowledge docs in `bundleId` nearest to the
80
+ * memory at `filePath`, nearest first. A memory the index does not know has no
81
+ * stored vector and so no neighbours.
79
82
  */
80
- export function findCoveringKnowledge(db, bundleId, filePath, body) {
81
- const shingles = wordShingles(body);
82
- if (shingles.size === 0)
83
- return undefined;
83
+ function knowledgeNeighbours(db, bundleId, filePath) {
84
84
  const entryId = getEntryIdByFilePath(db, filePath);
85
85
  if (entryId === undefined)
86
- return undefined;
87
- let best;
86
+ return [];
87
+ const out = [];
88
88
  for (const hit of getNeighborsByEntryId(db, entryId, PAIR_NEIGHBOR_FETCH_K, { type: "knowledge", bundleId })) {
89
89
  const neighbour = getEntryById(db, hit.id);
90
90
  if (!neighbour)
@@ -96,13 +96,67 @@ export function findCoveringKnowledge(db, bundleId, filePath, body) {
96
96
  catch {
97
97
  continue; // the index outlived the file
98
98
  }
99
- const containment = shingleContainment(shingles, stripFrontmatterBody(raw));
99
+ out.push({
100
+ ref: neighbour.conceptId,
101
+ description: neighbour.entry.description ?? "",
102
+ body: stripFrontmatterBody(raw),
103
+ });
104
+ }
105
+ return out;
106
+ }
107
+ /**
108
+ * The best-covering knowledge doc among the {@link PAIR_NEIGHBOR_FETCH_K}
109
+ * knowledge docs in `bundleId` nearest to the memory. `filePath` is the
110
+ * memory's indexed file; a memory the index does not know has no stored
111
+ * vector and so no candidates.
112
+ */
113
+ export function findCoveringKnowledge(db, bundleId, filePath, body) {
114
+ const shingles = wordShingles(body);
115
+ if (shingles.size === 0)
116
+ return undefined;
117
+ let best;
118
+ for (const neighbour of knowledgeNeighbours(db, bundleId, filePath)) {
119
+ const containment = shingleContainment(shingles, neighbour.body);
100
120
  if (containment >= COVERAGE_MIN_CONTAINMENT && (best === undefined || containment > best.containment)) {
101
- best = { ref: neighbour.conceptId, containment };
121
+ best = { ref: neighbour.ref, containment };
102
122
  }
103
123
  }
104
124
  return best;
105
125
  }
126
+ /** Knowledge docs a reviewer is shown for a promotion, nearest first. */
127
+ export const NEIGHBOUR_NOTE_COUNT = 5;
128
+ const NEIGHBOUR_EXCERPT_CHARS = 300;
129
+ /**
130
+ * The knowledge notes nearest to the memory at `memoryPath`, for the drain's
131
+ * judge to compare a promotion against: the nearest {@link NEIGHBOUR_NOTE_COUNT}
132
+ * of the same candidates the coverage gate reads. The memory's bundle is the one
133
+ * the index recorded for it. Empty when the index has no vector for the memory
134
+ * or cannot be opened; never throws.
135
+ */
136
+ export function nearestKnowledgeNotes(memoryPath) {
137
+ let db;
138
+ try {
139
+ db = openExistingDatabase();
140
+ const entryId = getEntryIdByFilePath(db, memoryPath);
141
+ const bundleId = entryId === undefined ? undefined : getEntryById(db, entryId)?.bundleId;
142
+ if (bundleId === undefined)
143
+ return [];
144
+ return knowledgeNeighbours(db, bundleId, memoryPath)
145
+ .slice(0, NEIGHBOUR_NOTE_COUNT)
146
+ .map((n) => ({
147
+ ref: n.ref,
148
+ description: n.description,
149
+ excerpt: n.body.length > NEIGHBOUR_EXCERPT_CHARS ? `${n.body.slice(0, NEIGHBOUR_EXCERPT_CHARS)}...` : n.body,
150
+ }));
151
+ }
152
+ catch {
153
+ return [];
154
+ }
155
+ finally {
156
+ if (db)
157
+ closeDatabase(db);
158
+ }
159
+ }
106
160
  /**
107
161
  * The gate for one run, holding its own read handle on `index.db` for the
108
162
  * promotions that run emits; `undefined` when there is no bundle or no index
@@ -92,7 +92,7 @@ export const PAIR_JUDGE_JSON_SCHEMA = {
92
92
  reason: { type: "string", maxLength: 400 },
93
93
  },
94
94
  };
95
- const PAIR_CHECK_JSON_SCHEMA = {
95
+ export const PAIR_CHECK_JSON_SCHEMA = {
96
96
  type: "object",
97
97
  required: ["missing"],
98
98
  additionalProperties: false,
@@ -503,7 +503,7 @@ function checkSection(label, side) {
503
503
  ].join("\n");
504
504
  }
505
505
  /**
506
- * The second look a duplicate gets before it may retire unattended: one call
506
+ * The second look a retirement gets before it may retire unattended: one call
507
507
  * that asks only what the retired note holds that the kept one lacks. True
508
508
  * only on a clean, empty answer (it caught 2 of 4 duplicates the judge got
509
509
  * wrong, and held back none of 109 right ones).
@@ -658,17 +658,20 @@ async function judgeOne(ctx, candidate) {
658
658
  }, ctx.opts.proposalsCtx);
659
659
  ctx.retired.push(proposal.id);
660
660
  ctx.perInitiatorProposed.add(candidate.initiator.ref);
661
- // A duplicate with nothing unique on either side, confirmed by a second
662
- // look, is the one class that retires unattended (109 of 111 safe on the
663
- // owner's reviewed pairs, 2026-10-04): the triage drain accepts it under
664
- // its usual applyMode. Every other retirement waits for a person.
665
- if (verdict.relation === "duplicate" &&
666
- verdict.onlyInA.length + verdict.onlyInB.length === 0 &&
661
+ // A retirement the second look confirms loses nothing retires unattended:
662
+ // the triage drain accepts it under its usual applyMode. The retired side
663
+ // holds no claim of its own for any label (`decideRetirement` mints nothing
664
+ // else); a duplicate must also leave the kept side with none, while a
665
+ // subsumed or superseding successor holds more by definition. Replay
666
+ // precision 336/360 (duplicate 0.98, subsumed 0.92, supersedes 0.875,
667
+ // 2026-10-07); a duplicate alone was 109 of 111 safe on the owner's
668
+ // reviewed pairs (2026-10-04). Anything else waits for a person.
669
+ if ((verdict.relation !== "duplicate" || verdict.onlyInA.length + verdict.onlyInB.length === 0) &&
667
670
  !continuityRisk &&
668
671
  (await confirmNothingLost(ctx, retired, successor))) {
669
672
  recordGateDecision(ctx.stashDir, proposal.id, {
670
673
  outcome: "staged",
671
- reason: "duplicate",
674
+ reason: verdict.relation,
672
675
  gate: PAIR_PASS_GATE,
673
676
  contentHash: proposalContentHash(proposal),
674
677
  }, ctx.opts.proposalsCtx);