akm-cli 0.9.27-alpha.1 → 0.9.27-alpha.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -6,6 +6,118 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/).
6
6
 
7
7
  ## [Unreleased]
8
8
 
9
+ ## [0.9.27-alpha.2] - 2026-10-07
10
+
11
+ ### Removed
12
+
13
+ - **The `scripts/akm-eval` toolkit has moved out of this repository.** The
14
+ read-only measurement toolkit (the case runner and its suites, the twin
15
+ experiment, the real-query verdict for the proactive lane, the state
16
+ analyzers and the curate benchmark) is retired; every live eval is in
17
+ [itlackey/akm-eval](https://github.com/itlackey/akm-eval). Its code is kept
18
+ there, to read and not to run, in `retired/akm-scripts-akm-eval/`, copied from
19
+ commit `f57a7fd44b37`. It imports akm's `src/` by relative path, so it runs
20
+ only in a checkout at that commit. Removed here with it: `scripts/akm-eval/`,
21
+ its tests (`tests/integration/akm-eval/`, `tests/akm-eval-*.test.ts`,
22
+ `tests/curate-metrics.test.ts`) and fixtures (`tests/fixtures/akm-eval/`, and
23
+ the `curate-golden` stash, which only the curate benchmark read), the
24
+ `akm-eval determinism` CI job, and `getMeasurementVerdictsDir`, whose only
25
+ caller was the verdict runner. akm no longer names
26
+ `$STATE/improve/measurement/verdicts/<stash>/`; a file already there is inert.
27
+ `docs/maintainers/eval.md` is now a pointer to the new home.
28
+
29
+ ### Fixed
30
+
31
+ - **The drain's judge no longer sees a note with a code block as truncated.**
32
+ The judgment prompt fenced the proposed content (and the live asset, sibling
33
+ proposals and neighbour excerpts) in three backticks, so a note holding its
34
+ own code block closed the fence early and read as cut off; real rejections said
35
+ "ends in an empty code block" or "truncated". Each block now uses a fence longer
36
+ than any backtick run inside it. The judge's reason is also kept on accepts,
37
+ staged accepts and defers (as the gate decision's `judgeReason`, until now
38
+ rejections only), and a judge reply that is not a verdict is stamped
39
+ `judgment-parse-failure`, and a runner failure `judgment-error`, instead of
40
+ looking like a defer.
41
+
42
+ - **Distill writes a lesson only when its memory holds one, says only what the
43
+ memory says, and its judge rejects what a reviewer would.** 2 of the 22 distill
44
+ proposals since 0.9.26 began were accepted, and 17 of the 19 queued on
45
+ 2026-10-05 were bad (they restated their memory, filed a dated status as a
46
+ lesson, claimed what the memory does not say, or repeated an asset the library
47
+ holds). Four causes, found in the code and the rejected proposals, and fixed:
48
+ (1) the prompt and schema forced a lesson from every memory, and 18 of the 19
49
+ were records of what was done; the writer now says why a memory holds a
50
+ lesson or none (`reason`, then `decision: lesson|none`, or the word `NONE`),
51
+ defined as a cause and what to do about it, or a rule with its reason, and
52
+ writes only what the memory and its feedback state, in the scope they have; a
53
+ `NONE` is a `skipped` distill (`skipReason: nothing_reusable`, the writer's
54
+ reason in the message) with no proposal and no judge call, and the loop keeps
55
+ its ledger row `unchanged`. (2) The judge asked for "information not already
56
+ present in the source", so an invented claim scored as novel and a faithful
57
+ lesson of a lesson-worthy memory as a restatement, and it passed anything that
58
+ "goes beyond the source" because it "may draw on feedback you are not shown".
59
+ The rubric is now reusable (a rule with its reason, not a record of what was
60
+ done), non-redundancy and grounding (every cause, step, number and limit is
61
+ in the source or its feedback), and the judge is shown the feedback the writer
62
+ saw. (3) A mean hid a decisive score (4 and 1 average 2.5, a review), and a
63
+ reviewer read everything the judge did not reject; any criterion at 2 or
64
+ below, grounding included, is now `quality_rejected`, and the reason names it
65
+ (`grounding 2/5: …`). The "borderline grounding" routing is gone. (4) Neither
66
+ the writer nor the judge could see a knowledge note or a skill that already
67
+ states the rule (the judge saw the 3 lexically nearest lessons, none of them
68
+ related); both now see the lessons, knowledge notes and skills nearest the
69
+ memory, which is the existing `processes.distill.cls` context turned on by
70
+ default (`enabled: false` turns it off). Judge scores are keyed `reusable`
71
+ where they were `novelty`. Measured on the local qwen3.8-27b with akm-eval's
72
+ `evals/distill` (30 fictional memories, 5 runs each side): good lessons 7/14 on
73
+ average (5 to 9) against 3.7/14 (3 to 5), lessons queued for memories that
74
+ deserve none 0.4/16 against 3.3/16. On 37 real memories with their feedback
75
+ (34 reviewed bad, 3 good; 3 runs against 2): a lesson was queued for 11% of the
76
+ bad ones against 44%, and for 6 of 9 good ones against 4 of 6; of the memories
77
+ that pass 0.9.26's skip of bare positive feedback, 20% of the bad against 50%.
78
+ No new settings.
79
+
80
+ - **Reflect no longer plans an asset whose negative feedback is already acted
81
+ on.** A negative `akm feedback` that came with an exact fix (`--replace` and
82
+ `--with`, `--outdated` or `--superseded-by`) makes a `feedback` proposal, and
83
+ once that proposal is accepted the feedback has done its work. Reflect still
84
+ took the ref as having fresh negative feedback, and on 2026-10-07 44 of its 50
85
+ refs were of that kind: the judge refused or the model changed nothing for
86
+ most of them. A negative event with a fix is now left out of the reflect
87
+ cursor when an accepted `feedback` proposal for the ref was created at or
88
+ after it. A negative with no fix, one given after the proposal, and one whose
89
+ proposal is still pending or was rejected plan a reflect as before.
90
+
91
+ - **Consolidate stops re-offering memories a reviewer already turned down, and
92
+ the nightly judge sees what a promotion may duplicate.** About 53 promotions a
93
+ night reached review at ~5% precision, 64-70% of them a memory body already
94
+ proposed or rejected. Four causes, four changes: a memory whose body equals
95
+ that of a consolidate promotion rejected on or after 2026-09-29 is held until
96
+ its body changes, under any name (earlier rejections, the bulk audits of
97
+ 2026-08, do not count); a memory the model judged and left alone is held by its
98
+ body hash instead of a 7-day clock, so an unchanged memory is no longer judged
99
+ every week (a row recorded without a hash keeps the 7 days); the coverage gate
100
+ skips a memory when 30% of its text, not 50%, is in a neighbouring knowledge
101
+ doc, which catches paraphrases; and the drain's judgment tier, which judged a
102
+ promotion seeing only the proposal and never `knowledge/`, is now shown the 5
103
+ nearest knowledge notes (ref, description, excerpt) and told to reject a
104
+ promotion they already cover. No new settings.
105
+
106
+ - **A confident `subsumed` or `supersedes` retirement resolves unattended, as a
107
+ `duplicate` already did.** The pair pass staged a retire proposal for the
108
+ triage drain only when the judge's label was `duplicate`; every other retirement
109
+ waited for a person. It now stages any of the three retire labels when the
110
+ second look (what does the retired note hold that the kept one lacks?) comes
111
+ back empty and there is no continuity risk; the retired side's claim list is
112
+ already empty for any proposal, and a `duplicate` must still have an empty list
113
+ on the kept side too. The staged gate reason is the judge's label, and the drain records it. Replay over 360
114
+ judged pairs: 336 safe (0.93): `duplicate` 0.98, `subsumed` 0.92, `supersedes`
115
+ 0.875. In production, unstaged `subsumed` retirements were accepted 45 of 54
116
+ times by hand, and in the latest run 41 of 47 pair proposals would have resolved
117
+ without a person. `docs/architecture/internals/improve-workflow.md` said triage
118
+ never auto-accepts a retire proposal, which stopped being true in 0.9.26; it,
119
+ and the matching lines in `improvement.md`, now describe the staging rule.
120
+
9
121
  ## [0.9.27-alpha.1] - 2026-10-06
10
122
 
11
123
  ### Fixed
package/LICENSE CHANGED
@@ -225,10 +225,9 @@ statute, judicial order, or regulation then You must: (a) comply with
225
225
  the terms of this License to the maximum extent possible; and (b)
226
226
  describe the limitations and the code they affect. Such description must
227
227
  be placed in a text file included with all distributions of the Covered
228
- Software under the name "LEGAL", with additions for new restrictions
229
- placed at the end of the file. Except to the extent prohibited by
230
- statute or regulation, such description must be sufficiently detailed
231
- for a recipient of ordinary skill to be able to understand it.
228
+ Software under this License. Except to the extent prohibited by statute
229
+ or regulation, such description must be sufficiently detailed for a
230
+ recipient of ordinary skill to be able to understand it.
232
231
 
233
232
  5. Termination
234
233
  --------------
@@ -1,7 +1,29 @@
1
1
  You are the akm `distill` distiller.
2
- Given an asset and recent feedback events about it, produce a single
3
- concise *lesson* an agent should remember next time it works on this
4
- asset's domain.
2
+ You are given a memory and the feedback recorded about it. Decide whether it
3
+ holds a lesson and, if it does, write the lesson.
4
+
5
+ A memory holds a lesson when it states a cause and what to do about it: a
6
+ failure or surprise with its cause and the fix that worked, or a rule with the
7
+ reason it holds. It holds a lesson even when it is short and names one project,
8
+ tool or incident, if the cause and the fix would help someone in a similar
9
+ situation.
10
+
11
+ A memory holds NO lesson when all it states is what was done, shipped, released
12
+ or decided, what is pending or planned, how a system is set up now, or the steps
13
+ of a procedure, with no failure and cause behind it, or when its feedback says
14
+ only that it is out of date or superseded. ANSWER NONE then: the single word and
15
+ nothing else. A reply bound to a JSON schema answers NONE by
16
+ setting `decision` to `none` and leaving the other fields empty. Answer NONE too
17
+ when a related asset listed below the memory already states the rule the memory
18
+ would give.
19
+
20
+ When the memory holds a lesson, write it from what the memory and its feedback
21
+ say, and nothing more:
22
+ - Add no cause, step, rule, number, check or safeguard that neither states.
23
+ - Keep the scope the memory has. A fix verified in one place is a fix for that
24
+ place, and what was not checked stays unchecked. One case is not "always" or
25
+ "never".
26
+ - Be shorter than the memory.
5
27
 
6
28
  YOUR RESPONSE MUST START EXACTLY WITH `---` ON THE VERY FIRST LINE.
7
29
  DO NOT output any prose, explanation, or code fences before or after.
@@ -12,7 +34,7 @@ description: <one complete sentence (ending with `.`) summarising what the lesso
12
34
  when_to_use: <one complete sentence describing the concrete trigger condition>
13
35
  ---
14
36
 
15
- <lesson body — plain markdown, 1–3 short paragraphs of practical guidance>
37
+ <lesson body — plain markdown, as short as the memory allows>
16
38
 
17
39
  ## description field (MANDATORY)
18
40
  - A single complete sentence in present tense, 20–400 chars, NO markdown.
@@ -21,7 +43,7 @@ when_to_use: <one complete sentence describing the concrete trigger condition>
21
43
  - DO NOT copy a section heading ("Key takeaways", "For example", "Key pitfalls").
22
44
  - DO NOT begin with a numbered list marker, code fence, or markdown heading.
23
45
 
24
- GOOD: "Always validate ref existence before promoting a memory to knowledge; missing refs surface as silent 404s during accept."
46
+ GOOD: "Pin the container image tag, because the `latest` tag moved under the nightly job and its output changed with no code change."
25
47
  BAD: "Key pitfalls"
26
48
  BAD: "When working with the akm CLI"
27
49
  BAD: "For example, you might..."
@@ -32,5 +54,5 @@ RULES:
32
54
  - `description` and `when_to_use` MUST differ from each other.
33
55
  - The lesson body MUST be non-empty markdown prose. Do NOT restate `description:` or `when_to_use:` inside the body (no `**description:** ...` or `**when_to_use:** ...` lines — the frontmatter is the only place those keys belong).
34
56
  - Do NOT emit a second `---` fence after the opening frontmatter — there are exactly two `---` lines in the output, both belonging to the single frontmatter block at the top.
35
- - Do NOT reproduce the source asset verbatim — distil what a caller needs to know.
36
- - Output ONLY the lesson file. No preamble, no code fences, no trailing prose.
57
+ - Do NOT reproduce the source asset verbatim.
58
+ - Output ONLY the lesson file. No preamble, no code fences, no trailing prose.
@@ -6,7 +6,7 @@
6
6
  * `knowledge/` proposal, ask whether `knowledge/` already says it.
7
7
  *
8
8
  * The rule: a memory is covered when at least {@link COVERAGE_MIN_CONTAINMENT}
9
- * (half) of its distinct {@link COVERAGE_SHINGLE_WORDS}-word shingles appear in
9
+ * (30%) of its distinct {@link COVERAGE_SHINGLE_WORDS}-word shingles appear in
10
10
  * one of the knowledge docs nearest to it. It is a containment of the MEMORY in
11
11
  * the doc, not a similarity: a long guide that quotes the memory covers it; a
12
12
  * memory that quotes a short doc and adds claims of its own does not.
@@ -25,9 +25,13 @@
25
25
  * this gate achieves: it reads only the {@link PAIR_NEIGHBOR_FETCH_K} nearest
26
26
  * knowledge docs (below), and a covering doc that ranks lower goes unseen. The
27
27
  * recall over that candidate set is unmeasured. The rest of the rejected
28
- * proposals (paraphrases, partial overlaps) still reach review: the next cut
29
- * measured, 0.2 (169 of 224 rejected), would also have skipped 2 of the 103
30
- * accepted ones, and no cosine cut was measured at all. A wrong skip is a
28
+ * proposals (paraphrases, partial overlaps) still reach review. 0.5 let
29
+ * paraphrases through: of the 54 promotions minted on 2026-10-07, 18 of the 51
30
+ * later rejected held 30% or more of their text in a neighbouring doc. The cut is
31
+ * now 0.3, between the two measured points: 0.2 (169 of 224 rejected) also
32
+ * skipped 2 of the 103 accepted ones. At 0.3, none of the 5 promotions graded
33
+ * good in the 2026-10-05 review sample would be skipped (their best doc holds
34
+ * at most 1% of them) while 5 of its 15 bad ones would. A wrong skip is a
31
35
  * promotion nobody gets to review.
32
36
  *
33
37
  * Candidates are the memory's {@link PAIR_NEIGHBOR_FETCH_K} nearest knowledge
@@ -49,7 +53,7 @@ import { PAIR_NEIGHBOR_FETCH_K } from "./pair-pass.js";
49
53
  /** Words per shingle. */
50
54
  export const COVERAGE_SHINGLE_WORDS = 5;
51
55
  /** Share of a memory's distinct shingles one knowledge doc must hold for the memory to count as covered. */
52
- export const COVERAGE_MIN_CONTAINMENT = 0.5;
56
+ export const COVERAGE_MIN_CONTAINMENT = 0.3;
53
57
  const WORD = /[\p{L}\p{N}]+/gu;
54
58
  /** The distinct lower-cased word n-grams of `text`; empty when it has fewer than {@link COVERAGE_SHINGLE_WORDS} words. */
55
59
  export function wordShingles(text) {
@@ -72,19 +76,15 @@ export function shingleContainment(memory, doc) {
72
76
  return shared / memory.size;
73
77
  }
74
78
  /**
75
- * The best-covering knowledge doc among the {@link PAIR_NEIGHBOR_FETCH_K}
76
- * knowledge docs in `bundleId` nearest to the memory. `filePath` is the
77
- * memory's indexed file; a memory the index does not know has no stored vector
78
- * and so no candidates.
79
+ * The {@link PAIR_NEIGHBOR_FETCH_K} knowledge docs in `bundleId` nearest to the
80
+ * memory at `filePath`, nearest first. A memory the index does not know has no
81
+ * stored vector and so no neighbours.
79
82
  */
80
- export function findCoveringKnowledge(db, bundleId, filePath, body) {
81
- const shingles = wordShingles(body);
82
- if (shingles.size === 0)
83
- return undefined;
83
+ function knowledgeNeighbours(db, bundleId, filePath) {
84
84
  const entryId = getEntryIdByFilePath(db, filePath);
85
85
  if (entryId === undefined)
86
- return undefined;
87
- let best;
86
+ return [];
87
+ const out = [];
88
88
  for (const hit of getNeighborsByEntryId(db, entryId, PAIR_NEIGHBOR_FETCH_K, { type: "knowledge", bundleId })) {
89
89
  const neighbour = getEntryById(db, hit.id);
90
90
  if (!neighbour)
@@ -96,13 +96,67 @@ export function findCoveringKnowledge(db, bundleId, filePath, body) {
96
96
  catch {
97
97
  continue; // the index outlived the file
98
98
  }
99
- const containment = shingleContainment(shingles, stripFrontmatterBody(raw));
99
+ out.push({
100
+ ref: neighbour.conceptId,
101
+ description: neighbour.entry.description ?? "",
102
+ body: stripFrontmatterBody(raw),
103
+ });
104
+ }
105
+ return out;
106
+ }
107
+ /**
108
+ * The best-covering knowledge doc among the {@link PAIR_NEIGHBOR_FETCH_K}
109
+ * knowledge docs in `bundleId` nearest to the memory. `filePath` is the
110
+ * memory's indexed file; a memory the index does not know has no stored
111
+ * vector and so no candidates.
112
+ */
113
+ export function findCoveringKnowledge(db, bundleId, filePath, body) {
114
+ const shingles = wordShingles(body);
115
+ if (shingles.size === 0)
116
+ return undefined;
117
+ let best;
118
+ for (const neighbour of knowledgeNeighbours(db, bundleId, filePath)) {
119
+ const containment = shingleContainment(shingles, neighbour.body);
100
120
  if (containment >= COVERAGE_MIN_CONTAINMENT && (best === undefined || containment > best.containment)) {
101
- best = { ref: neighbour.conceptId, containment };
121
+ best = { ref: neighbour.ref, containment };
102
122
  }
103
123
  }
104
124
  return best;
105
125
  }
126
+ /** Knowledge docs a reviewer is shown for a promotion, nearest first. */
127
+ export const NEIGHBOUR_NOTE_COUNT = 5;
128
+ const NEIGHBOUR_EXCERPT_CHARS = 300;
129
+ /**
130
+ * The knowledge notes nearest to the memory at `memoryPath`, for the drain's
131
+ * judge to compare a promotion against: the nearest {@link NEIGHBOUR_NOTE_COUNT}
132
+ * of the same candidates the coverage gate reads. The memory's bundle is the one
133
+ * the index recorded for it. Empty when the index has no vector for the memory
134
+ * or cannot be opened; never throws.
135
+ */
136
+ export function nearestKnowledgeNotes(memoryPath) {
137
+ let db;
138
+ try {
139
+ db = openExistingDatabase();
140
+ const entryId = getEntryIdByFilePath(db, memoryPath);
141
+ const bundleId = entryId === undefined ? undefined : getEntryById(db, entryId)?.bundleId;
142
+ if (bundleId === undefined)
143
+ return [];
144
+ return knowledgeNeighbours(db, bundleId, memoryPath)
145
+ .slice(0, NEIGHBOUR_NOTE_COUNT)
146
+ .map((n) => ({
147
+ ref: n.ref,
148
+ description: n.description,
149
+ excerpt: n.body.length > NEIGHBOUR_EXCERPT_CHARS ? `${n.body.slice(0, NEIGHBOUR_EXCERPT_CHARS)}...` : n.body,
150
+ }));
151
+ }
152
+ catch {
153
+ return [];
154
+ }
155
+ finally {
156
+ if (db)
157
+ closeDatabase(db);
158
+ }
159
+ }
106
160
  /**
107
161
  * The gate for one run, holding its own read handle on `index.db` for the
108
162
  * promotions that run emits; `undefined` when there is no bundle or no index
@@ -503,7 +503,7 @@ function checkSection(label, side) {
503
503
  ].join("\n");
504
504
  }
505
505
  /**
506
- * The second look a duplicate gets before it may retire unattended: one call
506
+ * The second look a retirement gets before it may retire unattended: one call
507
507
  * that asks only what the retired note holds that the kept one lacks. True
508
508
  * only on a clean, empty answer (it caught 2 of 4 duplicates the judge got
509
509
  * wrong, and held back none of 109 right ones).
@@ -658,17 +658,20 @@ async function judgeOne(ctx, candidate) {
658
658
  }, ctx.opts.proposalsCtx);
659
659
  ctx.retired.push(proposal.id);
660
660
  ctx.perInitiatorProposed.add(candidate.initiator.ref);
661
- // A duplicate with nothing unique on either side, confirmed by a second
662
- // look, is the one class that retires unattended (109 of 111 safe on the
663
- // owner's reviewed pairs, 2026-10-04): the triage drain accepts it under
664
- // its usual applyMode. Every other retirement waits for a person.
665
- if (verdict.relation === "duplicate" &&
666
- verdict.onlyInA.length + verdict.onlyInB.length === 0 &&
661
+ // A retirement the second look confirms loses nothing retires unattended:
662
+ // the triage drain accepts it under its usual applyMode. The retired side
663
+ // holds no claim of its own for any label (`decideRetirement` mints nothing
664
+ // else); a duplicate must also leave the kept side with none, while a
665
+ // subsumed or superseding successor holds more by definition. Replay
666
+ // precision 336/360 (duplicate 0.98, subsumed 0.92, supersedes 0.875,
667
+ // 2026-10-07); a duplicate alone was 109 of 111 safe on the owner's
668
+ // reviewed pairs (2026-10-04). Anything else waits for a person.
669
+ if ((verdict.relation !== "duplicate" || verdict.onlyInA.length + verdict.onlyInB.length === 0) &&
667
670
  !continuityRisk &&
668
671
  (await confirmNothingLost(ctx, retired, successor))) {
669
672
  recordGateDecision(ctx.stashDir, proposal.id, {
670
673
  outcome: "staged",
671
- reason: "duplicate",
674
+ reason: verdict.relation,
672
675
  gate: PAIR_PASS_GATE,
673
676
  contentHash: proposalContentHash(proposal),
674
677
  }, ctx.opts.proposalsCtx);
@@ -260,6 +260,39 @@ function loadPendingConsolidateProposalHashes(stashDir, proposalsCtx) {
260
260
  }
261
261
  return hashes;
262
262
  }
263
+ /**
264
+ * Rejections decided before this are not a verdict on the memory's text: the
265
+ * 2026-08-02 and 2026-08-18 bulk audits rejected hundreds of promotions
266
+ * wholesale, and holding their bodies would skip good memories for good.
267
+ */
268
+ const REJECTED_BODY_HOLD_FROM = "2026-09-29";
269
+ /**
270
+ * Body hashes of the memories whose promotion was rejected on review. A
271
+ * proposal minted with `promotionSourceHash` names the memory's raw body; an
272
+ * older one is hashed from its own body, which is the memory's unless
273
+ * sanitization changed it.
274
+ */
275
+ function loadRejectedPromotionBodyHashes(stashDir, proposalsCtx) {
276
+ const hashes = new Set();
277
+ try {
278
+ for (const p of listProposalsReadOnly(stashDir, { status: "rejected", includeArchive: true }, proposalsCtx)) {
279
+ if (p.source !== "consolidate")
280
+ continue;
281
+ if ((p.review?.decidedAt ?? p.updatedAt) < REJECTED_BODY_HOLD_FROM)
282
+ continue;
283
+ try {
284
+ hashes.add(p.promotionSourceHash ?? contentHash(proposalContent(p), "body"));
285
+ }
286
+ catch {
287
+ // A malformed payload cannot hold a memory.
288
+ }
289
+ }
290
+ }
291
+ catch {
292
+ // Best-effort: a failed read never blocks judging.
293
+ }
294
+ return hashes;
295
+ }
263
296
  /**
264
297
  * Body hashes of the live knowledge assets, read from disk (the index may lag
265
298
  * a just-written asset), so an accepted promotion is not proposed again.
@@ -469,6 +502,18 @@ export function inspectConsolidationPool(opts, stashDir, warnings, existingKnowl
469
502
  return !isLedgerBlocked(row, nowIso, changedAt);
470
503
  });
471
504
  }
505
+ // A memory whose text a reviewer already rejected as a promotion waits for an edit, whatever it is called.
506
+ const rejectedBodies = loadRejectedPromotionBodyHashes(stashDir, opts.proposalsCtx);
507
+ if (rejectedBodies.size > 0) {
508
+ memories = memories.filter((memory) => {
509
+ try {
510
+ return !rejectedBodies.has(contentHash(fs.readFileSync(memory.filePath, "utf8"), "body"));
511
+ }
512
+ catch {
513
+ return true;
514
+ }
515
+ });
516
+ }
472
517
  const judgedUnchanged = poolSize - memories.length;
473
518
  // Only what retrieval returned or new material improve never processed (#986).
474
519
  const retrievalScope = loadRetrievalScope({ proposalsCtx: opts.proposalsCtx, readOnly }, stashDir);
@@ -768,7 +813,13 @@ async function consolidate(opts, config, stashDir, startMs, stateDb) {
768
813
  recordLedgerAttempt({ proposalsCtx: opts.proposalsCtx }, [...acc.judgedRefs]
769
814
  .filter((ref) => !ctx.promotedSourceRefs.has(ref) &&
770
815
  !acc.skipReasonByRef.get(ref)?.skips.some((skip) => skip.reason === "promote_create_failed"))
771
- .map((ref) => ({ stashDir, ref, source: "consolidate", outcome: "judged_no_action" })));
816
+ .map((ref) => ({
817
+ stashDir,
818
+ ref,
819
+ source: "consolidate",
820
+ outcome: "judged_no_action",
821
+ ...bodyHashOf(ctx.memoryByRef.get(ref)),
822
+ })));
772
823
  return makeConsolidateResult({
773
824
  ...summary(),
774
825
  promoted: ctx.promoted,
@@ -783,6 +834,17 @@ async function consolidate(opts, config, stashDir, startMs, stateDb) {
783
834
  },
784
835
  });
785
836
  }
837
+ /** The memory's current body hash as a ledger input field; empty when it cannot be read (the row then keeps its 7-day window). */
838
+ function bodyHashOf(memory) {
839
+ if (!memory)
840
+ return {};
841
+ try {
842
+ return { contentHash: contentHash(fs.readFileSync(memory.filePath, "utf8"), "body") };
843
+ }
844
+ catch {
845
+ return {};
846
+ }
847
+ }
786
848
  /** The conceptId a ref maps to, or undefined for an invalid ref. */
787
849
  function conceptIdForRef(ref) {
788
850
  try {
@@ -2,25 +2,24 @@
2
2
  // License, v. 2.0. If a copy of the MPL was not distributed with this
3
3
  // file, You can obtain one at https://mozilla.org/MPL/2.0/.
4
4
  /**
5
- * Distill guards: related lessons/knowledge shown to the model so it does not
6
- * overwrite prior generalizations (CLS context), and a cheap check that a
7
- * proposal does not contradict the memories it came from.
5
+ * Distill guards: the related lessons, knowledge notes and skills shown to the
6
+ * writer so it does not repeat or overwrite them (CLS context), and a cheap check
7
+ * that a proposal does not contradict the memories it came from.
8
8
  */
9
9
  export const DEFAULT_CLS_ADJACENT_COUNT = 3;
10
- /** The CLS prompt section (each entry capped at 400 chars); empty when disabled or nothing is related. */
10
+ /** The CLS prompt section (each entry capped at 600 chars); empty when disabled (on unless `enabled: false`) or nothing is related. */
11
11
  export function buildClsContext(adjacentItems, config) {
12
- if (!config.enabled || adjacentItems.length === 0)
12
+ if (config.enabled === false || adjacentItems.length === 0)
13
13
  return "";
14
14
  const lines = [
15
15
  "",
16
- "## Existing adjacent lessons / knowledge (CLS context)",
17
- "The following are semantically related entries already in the stash.",
18
- "Your proposal MUST NOT contradict or silently overwrite these — if you",
19
- "disagree with one, flag it as contradicted (do not ignore it).",
16
+ "## Related assets already in the library",
17
+ "The library already holds these lessons, knowledge notes and skills near this memory. They may be about another subject.",
18
+ "If one of them already states the rule the memory would give, answer NONE. Do not contradict or overwrite them.",
20
19
  "",
21
20
  ];
22
21
  for (const item of adjacentItems)
23
- lines.push(`### ${item.ref}`, item.content.trim().slice(0, 400), "");
22
+ lines.push(`### ${item.ref}`, item.content.trim().slice(0, 600), "");
24
23
  return lines.join("\n");
25
24
  }
26
25
  /**
@@ -78,23 +78,29 @@ export function deriveLessonRef(inputRef) {
78
78
  // properties, so every property is required and "none" is an empty array (#1046).
79
79
  export const DISTILL_LESSON_JSON_SCHEMA = {
80
80
  type: "object",
81
- required: ["description", "when_to_use", "body", "tags"],
81
+ required: ["reason", "decision", "description", "when_to_use", "body", "tags"],
82
82
  additionalProperties: false,
83
83
  properties: {
84
+ reason: {
85
+ type: "string",
86
+ description: "One sentence, written first: the cause and the fix the memory states, or why it states none.",
87
+ },
88
+ decision: {
89
+ type: "string",
90
+ enum: ["lesson", "none"],
91
+ description: "`none` when the memory holds no lesson (it records what was done, a design, or steps already written elsewhere): leave the other fields empty. Otherwise `lesson`.",
92
+ },
84
93
  description: {
85
94
  type: "string",
86
- minLength: 10,
87
- description: "Single complete sentence (80-200 chars) summarising what the lesson teaches. No markdown, no leading 'When'/'If'.",
95
+ description: "Single complete sentence summarising what the lesson teaches. No markdown, no leading 'When'/'If'. Empty for `none`.",
88
96
  },
89
97
  when_to_use: {
90
98
  type: "string",
91
- minLength: 10,
92
- description: "Single complete sentence describing the concrete trigger condition for the lesson.",
99
+ description: "Single complete sentence describing the concrete trigger condition for the lesson. Empty for `none`.",
93
100
  },
94
101
  body: {
95
102
  type: "string",
96
- minLength: 1,
97
- description: "Lesson body — plain markdown, 1-3 short paragraphs of practical guidance.",
103
+ description: "Lesson body: plain markdown, shorter than the memory, stating only what it and its feedback say. Empty for `none`.",
98
104
  },
99
105
  tags: {
100
106
  type: "array",
@@ -154,6 +160,19 @@ export function assembleStructuredDistillMarkdown(payload, kind) {
154
160
  fm.xrefs = sources;
155
161
  return assembleAssetFromString(serializeFrontmatterQuoted(fm), body);
156
162
  }
163
+ /**
164
+ * The writer's answer when it found no lesson: the word NONE, or `decision: "none"` in a reply bound to the schema,
165
+ * with the reason it gave (`""` for the bare word). `null` for any other reply.
166
+ */
167
+ function answeredNone(raw) {
168
+ if (/^[\s"'`*_]*none[\s.!"'`*_]*$/i.test(stripMarkdownFences(raw)))
169
+ return { reason: "" };
170
+ const payload = parseEmbeddedJsonResponse(raw);
171
+ if (payload === null || typeof payload !== "object" || Array.isArray(payload) || payload.decision !== "none") {
172
+ return null;
173
+ }
174
+ return { reason: typeof payload.reason === "string" ? payload.reason.trim() : "" };
175
+ }
157
176
  function validateKnowledgeContent(content, inputRef) {
158
177
  const findings = [];
159
178
  const parsed = parseFrontmatter(content);
@@ -251,7 +270,7 @@ export function buildDistillPrompt(input) {
251
270
  }
252
271
  lines.push(input.proposalKind === "knowledge"
253
272
  ? "Produce the knowledge markdown file now. Start your response with `---` on the first line, followed by a `description:` field whose value is a 1-sentence summary (20–400 chars). Never use placeholder values like `---`, `tbd`, `n/a`, or a single dash. If the source has nothing meaningful to summarize, do NOT produce a proposal — return an empty response instead. The frontmatter block ends with a second `---` line; do not emit any additional `---` fences in the body."
254
- : "Produce the lesson markdown file now. Start your response with `---` on the first line, followed by `description:` and `when_to_use:` fields. Both must be real one-sentence summaries (20–400 chars) — never placeholder values like `---`, `tbd`, or `n/a`. The frontmatter block ends with a second `---` line; do not emit any additional `---` fences in the body.");
273
+ : "Produce the lesson markdown file now. Start your response with `---` on the first line, followed by `description:` and `when_to_use:` fields. Both must be real one-sentence summaries (20–400 chars) — never placeholder values like `---`, `tbd`, or `n/a`. The frontmatter block ends with a second `---` line; do not emit any additional `---` fences in the body. If the memory holds no lesson, answer NONE instead.");
255
274
  return lines.join("\n");
256
275
  }
257
276
  // ── Invocation ───────────────────────────────────────────────────────────────
@@ -343,7 +362,7 @@ export async function akmDistill(options) {
343
362
  asset,
344
363
  vocabulary: loadRefVocabulary(),
345
364
  outcomeWeightEnabled: config.improve?.salience?.outcomeWeightEnabled !== false,
346
- similar: options.fetchSimilarLessonsFn ?? fetchTopSimilarLessons,
365
+ related: options.fetchRelatedFn ?? fetchRelatedAssets,
347
366
  lookup,
348
367
  };
349
368
  const feedbackEvents = readDistillFeedback(run);
@@ -381,6 +400,8 @@ async function distill(run, targetKind, kind, outputRef, feedbackEvents) {
381
400
  system,
382
401
  prompt,
383
402
  gate: { config: run.config, enabled: true },
403
+ // NONE is an answer: a parser that rejected it would ask for a lesson again.
404
+ parse: (raw) => (answeredNone(raw) ? raw : parseEmbeddedJsonResponse(raw)),
384
405
  // The injected test transport never sees the schema.
385
406
  request: {
386
407
  ...(run.options.chat === undefined
@@ -413,6 +434,10 @@ async function distill(run, targetKind, kind, outputRef, feedbackEvents) {
413
434
  ...exclusionMeta(run, true),
414
435
  };
415
436
  }
437
+ const none = answeredNone(call.raw);
438
+ if (none) {
439
+ return skipDistill(run, outputRef, kind, "nothing_reusable", `The writer found no lesson in ${run.inputRef}${none.reason ? `: ${none.reason}` : "."}`);
440
+ }
416
441
  const assembled = assembleDistilledContent(run, call.raw, kind, outputRef);
417
442
  if ("rejection" in assembled)
418
443
  return assembled.rejection;
@@ -422,8 +447,20 @@ async function distill(run, targetKind, kind, outputRef, feedbackEvents) {
422
447
  content: assembled.content,
423
448
  source: run.asset.content,
424
449
  descriptionSwapped: assembled.descriptionSwapped,
450
+ feedback: feedbackLines(feedback),
425
451
  });
426
452
  }
453
+ /** The feedback that says something, one line each, for the judge. A bare signal says nothing the writer could use. */
454
+ function feedbackLines(feedback) {
455
+ const lines = [];
456
+ for (const event of feedback) {
457
+ const meta = event.metadata ?? {};
458
+ const detail = (typeof meta.reason === "string" ? meta.reason : "") || (typeof meta.note === "string" ? meta.note : "");
459
+ if (detail.trim())
460
+ lines.push(`- [${typeof meta.signal === "string" ? meta.signal : event.eventType}] ${detail.trim()}`);
461
+ }
462
+ return lines;
463
+ }
427
464
  /** Whether a file already holds the lesson `ref` in the stash the proposal would be filed in. */
428
465
  function lessonExists(run, ref) {
429
466
  const { type, name } = parseRefInput(ref);
@@ -489,11 +526,12 @@ async function judgeAndQueue(run, out) {
489
526
  let confidence;
490
527
  let judged;
491
528
  if (qualityGateEnabled(run)) {
492
- const similarLessons = await run.similar(content.slice(0, 500), 3);
529
+ const related = await run.related(content.slice(0, 500), RELATED_COUNT);
493
530
  // The judge reads what the generator read: the source body, without its frontmatter (buildDistillPrompt).
494
531
  const source = out.source ? parseFrontmatter(out.source).content.trim() : "";
495
532
  const verdict = await runLessonQualityJudge(run.config, content, source, run.options.chat, {
496
- ...(similarLessons.length > 0 ? { similarLessons } : {}),
533
+ ...(related.length > 0 ? { related } : {}),
534
+ ...(out.feedback && out.feedback.length > 0 ? { feedback: out.feedback } : {}),
497
535
  ...((run.judgeRunner ?? run.runner) ? { llmRunner: run.judgeRunner ?? run.runner } : {}),
498
536
  ...(run.options.signal ? { signal: run.options.signal } : {}),
499
537
  onNotices: run.notices.add,
@@ -873,13 +911,13 @@ function readDistillFeedback(run) {
873
911
  /** System + user prompt: rejected-proposal context, optional CLS neighbours, stash standards. */
874
912
  async function buildDistillMessages(run, feedback, kind, outputRef) {
875
913
  const rejectedProposals = rejectedProposalContext(run.stash, run.inputRef, run.options.ctx, run.options.eventsCtx);
876
- // CLS interleaving (default off): show related lessons so the model does not overwrite them.
914
+ // CLS interleaving (default on): show the related lessons, knowledge notes and skills, so the writer neither repeats nor overwrites them.
877
915
  const cls = getImproveProcessConfig("distill", run.profile)?.cls ?? {};
878
916
  let clsContext = "";
879
- if (cls.enabled) {
917
+ if (cls.enabled !== false) {
880
918
  try {
881
919
  const query = run.asset.content ? run.asset.content.slice(0, 500) : run.inputRef;
882
- clsContext = buildClsContext(await run.similar(query, cls.adjacentCount ?? DEFAULT_CLS_ADJACENT_COUNT), cls);
920
+ clsContext = buildClsContext(await run.related(query, cls.adjacentCount ?? DEFAULT_CLS_ADJACENT_COUNT), cls);
883
921
  }
884
922
  catch {
885
923
  // CLS context is supplemental.
@@ -908,12 +946,18 @@ async function defaultLookup(ref, stashDir) {
908
946
  honorOrigin: false,
909
947
  });
910
948
  }
911
- /** Top-N existing lessons similar to `query` (empty when search is unavailable). */
912
- async function fetchTopSimilarLessons(query, n) {
949
+ /** What the library already holds on a subject: lessons and knowledge notes say it, a skill is how to do it. */
950
+ const RELATED_TYPES = ["lesson", "knowledge", "skill"];
951
+ const RELATED_COUNT = 3;
952
+ /** The top-N lessons, knowledge notes and skills related to `query`, best first (empty when search is unavailable). */
953
+ async function fetchRelatedAssets(query, n) {
913
954
  try {
914
- const result = await akmSearch({ query, type: "lesson", limit: n, skipLogging: true, eventSource: "improve" });
915
- return (result?.hits ?? [])
955
+ // One search per type: memories outnumber the rest and would fill an untyped list.
956
+ const results = await Promise.all(RELATED_TYPES.map((type) => akmSearch({ query, type, limit: n, skipLogging: true, eventSource: "improve" })));
957
+ return results
958
+ .flatMap((result) => result?.hits ?? [])
916
959
  .filter((h) => "path" in h && typeof h.path === "string")
960
+ .sort((a, b) => (b.score ?? 0) - (a.score ?? 0))
917
961
  .slice(0, n)
918
962
  .map((h) => {
919
963
  let content = "";
@@ -601,6 +601,16 @@ export function buildSnapshotManifest(args) {
601
601
  const latestNegativeTs = new Map();
602
602
  const feedback = new Map(candidates.map((r) => [r.ref, { hasSignal: false, positive: 0, negative: 0 }]));
603
603
  if (candidates.length > 0) {
604
+ // When each ref's accepted feedback proposals were created: a fix event is acted on once one exists at or after it.
605
+ const fixedAt = new Map();
606
+ if (stashDir) {
607
+ withRunState(eventsCtx, args.readOnly !== true, (db) => {
608
+ for (const p of listStateProposals(db, { stashDir, status: "accepted" })) {
609
+ if (p.source === "feedback")
610
+ fixedAt.set(p.ref, [...(fixedAt.get(p.ref) ?? []), p.createdAt]);
611
+ }
612
+ });
613
+ }
604
614
  for (const e of readEvents({ type: "feedback" }, eventsCtx).events) {
605
615
  const ref = e.ref ? refByKey.get(e.ref) : undefined;
606
616
  const entry = ref ? feedback.get(ref) : undefined;
@@ -612,8 +622,11 @@ export function buildSnapshotManifest(args) {
612
622
  entry.hasSignal = true;
613
623
  if (ts > (latestFeedbackTs.get(ref) ?? ""))
614
624
  latestFeedbackTs.set(ref, ts);
615
- if (signal === "negative" && ts > (latestNegativeTs.get(ref) ?? ""))
625
+ const fixApplied = e.metadata?.fix !== undefined &&
626
+ (fixedAt.get(e.ref ?? "") ?? []).some((createdAt) => createdAt >= ts);
627
+ if (signal === "negative" && !fixApplied && ts > (latestNegativeTs.get(ref) ?? "")) {
616
628
  latestNegativeTs.set(ref, ts);
629
+ }
617
630
  }
618
631
  if (signal === "positive")
619
632
  entry.positive++;
@@ -216,28 +216,31 @@ export function resolveQualityGateJudge(config, profile, processName, onNotices)
216
216
  onNotices?.(resolved.notices);
217
217
  return resolved.runner;
218
218
  }
219
- /** Lesson judge prompt; similar existing lessons let it mark near-duplicates down. */
220
- export function buildJudgePrompt(lessonContent, sourceContent, similarLessons) {
219
+ /** Lesson judge prompt: what the writer was given (the source and its feedback), the assets nearest the new lesson, the lesson. */
220
+ export function buildJudgePrompt(lessonContent, sourceContent, related, feedback) {
221
221
  const lines = [
222
- "You are evaluating a proposed lesson asset for an akm knowledge base.",
222
+ "You are evaluating a lesson an agent wrote from a memory and the feedback about it, for an akm knowledge base.",
223
223
  "",
224
224
  "Score this lesson on each criterion from 1 (poor) to 5 (excellent):",
225
- "1. NOVELTY: Does the lesson add information not already present in the source asset?",
226
- "2. NON-REDUNDANCY: Is this lesson meaningfully different from what the source already says?",
227
- "3. GROUNDING: Is the lesson about what the source asset is about? Score 1-2 only if it is about a different subject than the source; 3 if it is on the source's subject but goes beyond or corrects what the source says (it may draw on feedback you are not shown); 4-5 if the source supports it. A lesson may generalize the source's point.",
225
+ "1. REUSABLE: Does the lesson state a rule an agent can use on another occasion, with the reason it holds? Score 1-2 when it only records what was done, shipped, decided, found or is pending, on a date or for one build, machine or project, or how a system is set up now, however it is phrased. Score 4-5 for a rule with its reason.",
226
+ "2. NON-REDUNDANCY: Is the lesson new next to the existing assets shown below? Score 1-2 only when one of them already states the same rule. Assets on other subjects change nothing: score 4-5 when none is shown or none is on the same subject.",
227
+ "3. GROUNDING: Is every statement in the lesson stated by the source or its feedback, in any words? Check each cause, step, number, rule and limit in the lesson against them. Score 4-5 when each is stated. Score 3 when one stretches what the source says. Score 1-2 when any is in neither, when the lesson drops a limit the source states (one place checked, not confirmed, a guess) and says more than it, or when it is about another subject than the source.",
228
228
  "",
229
- "Source asset content:",
229
+ "Source memory:",
230
230
  "```",
231
231
  // The window distill generates from (buildDistillPrompt): grounding can reject, so the judge reads all of it.
232
232
  sourceContent.slice(0, 3000),
233
233
  "```",
234
234
  ];
235
- if (similarLessons && similarLessons.length > 0) {
236
- lines.push("", "Existing similar lessons (top-3 by similarity). Rate NOVELTY and NON-REDUNDANCY lower if the proposed lesson is substantially similar to any of these:");
237
- for (const sl of similarLessons)
238
- lines.push(`\nExisting lesson ref: ${sl.ref}`, "```", sl.content.slice(0, 500), "```");
235
+ if (feedback && feedback.length > 0) {
236
+ lines.push("", "Feedback recorded about the memory (the writer saw it too):", "```", feedback.join("\n").slice(0, 1500), "```");
237
+ }
238
+ if (related && related.length > 0) {
239
+ lines.push("", "Existing assets nearest the new lesson (they may be on another subject):");
240
+ for (const asset of related)
241
+ lines.push(`\nExisting asset ref: ${asset.ref}`, "```", asset.content.slice(0, 600), "```");
239
242
  }
240
- lines.push("", "Proposed lesson content:", "```", lessonContent.slice(0, 1000), "```", "", 'Return ONLY valid JSON, no prose: {"scores": {"novelty": <1-5 integer>, "nonRedundancy": <1-5 integer>, "grounding": <1-5 integer>}, "reason": "<one sentence>"}');
243
+ lines.push("", "Proposed lesson:", "```", lessonContent.slice(0, 2000), "```", "", 'Return ONLY valid JSON, no prose: {"scores": {"reusable": <1-5 integer>, "nonRedundancy": <1-5 integer>, "grounding": <1-5 integer>}, "reason": "<one sentence naming the weakest criterion>"}');
241
244
  return lines.join("\n");
242
245
  }
243
246
  function boundedDocument(content, maxChars = 6000) {
@@ -317,23 +320,19 @@ export function buildReflectJudgePrompt(candidateContent, sourceContent, feedbac
317
320
  }
318
321
  /**
319
322
  * `grounding` is scored with the other lesson criteria but left out of their
320
- * mean: a lesson about a different subject than its source reads as novel and
321
- * non-redundant, so they would pass it (or, in the review band, mint it as
322
- * a pending proposal). The rubric reserves 1-2 for a different subject. A score
323
- * of {@link UNGROUNDED_MAX_SCORE} or less is a rejection whatever the mean says
324
- * (#999). A higher score up to {@link BORDERLINE_GROUNDING_MAX_SCORE} is only
325
- * borderline: a lesson on its source's subject that advises beyond it has scored
326
- * 2, and a score can move a point between runs (see `runQualityJudge`), so it
327
- * goes to a person unless the mean alone already rejects it. A lesson that goes
328
- * beyond or corrects its source is on its subject: distill folds feedback into
329
- * the lesson, and the judge is never shown it. A contradiction of the source is
330
- * the optional fidelity check's to send to a human (`judgeAndQueue` in
331
- * distill.ts), so the rubric must not pre-empt it.
323
+ * mean. A lesson criterion scored {@link LESSON_REJECT_MAX_SCORE} or less is a
324
+ * rejection whatever the mean says: the mean would hide it (4 and 1 average
325
+ * 2.5, a review), and a reviewer was reading every lesson that was not rejected,
326
+ * 17 of 19 of them bad on 2026-10-05. The rubric reserves 1-2 for a lesson that
327
+ * records what was done instead of a rule, repeats an asset the library holds,
328
+ * or states what neither its source nor its feedback does. The judge is shown
329
+ * the feedback the writer saw, so a statement it supports is not an invention. A
330
+ * contradiction of the source is the optional fidelity check's to send to a
331
+ * human (`judgeAndQueue` in distill.ts).
332
332
  */
333
333
  const GROUNDING_CRITERION = "grounding";
334
- const UNGROUNDED_MAX_SCORE = 1;
335
- const BORDERLINE_GROUNDING_MAX_SCORE = 2;
336
- const LESSON_JUDGE_CRITERIA = ["novelty", "nonRedundancy", GROUNDING_CRITERION];
334
+ const LESSON_REJECT_MAX_SCORE = 2;
335
+ const LESSON_JUDGE_CRITERIA = ["reusable", "nonRedundancy", GROUNDING_CRITERION];
337
336
  const REFLECT_JUDGE_CRITERIA = ["need", "preservation", "quality"];
338
337
  /**
339
338
  * Read a judge response: the per-criterion shape (averaged here, `grounding`
@@ -390,9 +389,8 @@ export function judgeResponseSchema(keys) {
390
389
  * The quality judge. Fails closed: no runner, an unparseable verdict or a
391
390
  * provider failure never passes content. Bands: every criterion in the mean
392
391
  * >= 4 passes, otherwise a mean >= 2.5 is review and a lower one reject; a
393
- * `grounding` score of {@link UNGROUNDED_MAX_SCORE} or less rejects whatever
394
- * the mean is, and one of {@link BORDERLINE_GROUNDING_MAX_SCORE} routes a lesson
395
- * that would pass to review (a mean that rejects stays a rejection).
392
+ * lesson criterion (`grounding` included) of {@link LESSON_REJECT_MAX_SCORE}
393
+ * or less rejects whatever the mean is.
396
394
  * Temperature is set to 0, which reduces run-to-run variation but does not
397
395
  * remove it: on some servers (llama.cpp batching, for one) the same request can
398
396
  * score a point apart, so the routing rules are chosen with that margin in mind.
@@ -430,34 +428,18 @@ async function runQualityJudge(feature, config, prompt, keys, chat, options) {
430
428
  if (!parsed)
431
429
  return { pass: false, score: -1, reason: "judge parse failed — routed to review", reviewNeeded: true };
432
430
  const { score, lowest, reason, criteria } = parsed;
433
- const grounding = criteria?.[GROUNDING_CRITERION];
434
- if (criteria && grounding !== undefined && grounding <= UNGROUNDED_MAX_SCORE) {
435
- return {
436
- pass: false,
437
- score,
438
- reason: `Off-subject for its source (grounding ${grounding}/5): ${reason}`,
439
- criteria,
440
- };
431
+ // A lesson criterion at 2 or below is a defect the mean would hide (4 and 1 average 2.5, a review): it rejects.
432
+ if (criteria && criteria[GROUNDING_CRITERION] !== undefined) {
433
+ const [weakest, low] = Object.entries(criteria).sort((x, y) => x[1] - y[1])[0];
434
+ if (low <= LESSON_REJECT_MAX_SCORE)
435
+ return { pass: false, score, reason: `${weakest} ${low}/5: ${reason}`, criteria };
441
436
  }
442
437
  const verdict = lowest >= 4 ? { pass: true } : score >= 2.5 ? { pass: false, reviewNeeded: true } : { pass: false };
443
- // Borderline grounding is a person's call even when the lesson would pass; a mean that rejects stays rejected.
444
- if (criteria &&
445
- grounding !== undefined &&
446
- grounding <= BORDERLINE_GROUNDING_MAX_SCORE &&
447
- (verdict.pass || verdict.reviewNeeded)) {
448
- return {
449
- pass: false,
450
- reviewNeeded: true,
451
- score,
452
- reason: `Borderline on grounding (${grounding}/5), routed to review: ${reason}`,
453
- criteria,
454
- };
455
- }
456
438
  return { ...verdict, score, reason, ...(criteria ? { criteria } : {}) };
457
439
  }
458
440
  /** Judge a proposed lesson (or knowledge promotion) against its source. */
459
441
  export function runLessonQualityJudge(config, lessonContent, sourceContent, chat, options = {}) {
460
- const prompt = buildJudgePrompt(lessonContent, sourceContent, options.similarLessons);
442
+ const prompt = buildJudgePrompt(lessonContent, sourceContent, options.related, options.feedback);
461
443
  return runQualityJudge("lesson_quality_gate", config, prompt, LESSON_JUDGE_CRITERIA, chat, options);
462
444
  }
463
445
  /** Judge an in-place reflect revision without new-lesson novelty criteria. */
@@ -28,6 +28,7 @@ import { info, warn } from "../../core/warn.js";
28
28
  import { DEFAULT_LLM_TIMEOUT_MS } from "../../integrations/agent/config.js";
29
29
  import { buildExecution, resolveExecution } from "../../integrations/agent/execution.js";
30
30
  import { assertRunnerCredentials, runExecution, } from "../../integrations/agent/runner-dispatch.js";
31
+ import { nearestKnowledgeNotes } from "../improve/consolidate/coverage.js";
31
32
  import { errMessage, noticeSet } from "../improve/stage.js";
32
33
  import { akmProposalAccept, akmProposalReject } from "./proposal.js";
33
34
  import { isRetireProposal, PAIR_PASS_GATE, STALE_TARGET_GATE_REASON } from "./proposal-types.js";
@@ -57,8 +58,13 @@ function categorizeDrainFailure(message, fallback) {
57
58
  * rejection, so instead of failing identically every run it is auto-rejected
58
59
  * once; the ledger records `failed`, keeping the ref re-proposable.
59
60
  */
60
- async function acceptProposal(opts, proposal, id, reason, promoteFn, rejectFn) {
61
- const gateDecision = { outcome: "auto-accepted", reason, gate: DRAIN_GATE };
61
+ async function acceptProposal(opts, proposal, id, reason, promoteFn, rejectFn, judgeReason) {
62
+ const gateDecision = {
63
+ outcome: "auto-accepted",
64
+ reason,
65
+ gate: DRAIN_GATE,
66
+ ...(judgeReason ? { judgeReason } : {}),
67
+ };
62
68
  try {
63
69
  if (!opts.dryRun) {
64
70
  await promoteFn({
@@ -102,7 +108,7 @@ async function acceptProposal(opts, proposal, id, reason, promoteFn, rejectFn) {
102
108
  }
103
109
  }
104
110
  /** Reject one proposal (nothing in a dry run); the error message on failure. */
105
- async function rejectProposal(opts, id, reason, gateReason, rejectFn) {
111
+ async function rejectProposal(opts, id, reason, gateReason, rejectFn, judgeReason) {
106
112
  if (opts.dryRun)
107
113
  return undefined;
108
114
  try {
@@ -110,7 +116,12 @@ async function rejectProposal(opts, id, reason, gateReason, rejectFn) {
110
116
  stashDir: opts.stashDir,
111
117
  id,
112
118
  reason,
113
- gateDecision: { outcome: "auto-rejected", reason: gateReason, gate: DRAIN_GATE },
119
+ gateDecision: {
120
+ outcome: "auto-rejected",
121
+ reason: gateReason,
122
+ gate: DRAIN_GATE,
123
+ ...(judgeReason ? { judgeReason } : {}),
124
+ },
114
125
  });
115
126
  return undefined;
116
127
  }
@@ -118,7 +129,20 @@ async function rejectProposal(opts, id, reason, gateReason, rejectFn) {
118
129
  return errMessage(err);
119
130
  }
120
131
  }
121
- /** The judgment prompt: the proposal, the live asset it would overwrite, and same-ref siblings. */
132
+ /**
133
+ * `text` as a fenced block whose fence is longer than any backtick run inside
134
+ * it (the CommonMark rule), so a note holding its own code block is not cut
135
+ * short at the first inner fence.
136
+ */
137
+ export function fencedBlock(text) {
138
+ let longest = 2;
139
+ for (const run of text.match(/`+/g) ?? [])
140
+ if (run.length > longest)
141
+ longest = run.length;
142
+ const fence = "`".repeat(longest + 1);
143
+ return [fence, text, fence];
144
+ }
145
+ /** The judgment prompt: the proposal, the live asset it would overwrite, same-ref siblings, and for a promotion the nearest knowledge notes. */
122
146
  export function buildJudgmentPrompt(proposal, reason, ctx) {
123
147
  const sections = [
124
148
  "You are adjudicating a pending knowledge-base proposal no quality judge has",
@@ -129,12 +153,10 @@ export function buildJudgmentPrompt(proposal, reason, ctx) {
129
153
  `Left for judgment because: ${reason === "needs-judgment" ? "no quality judge has passed this content yet" : reason}`,
130
154
  "",
131
155
  "## Proposed content",
132
- "```",
133
- proposalContent(proposal),
134
- "```",
156
+ ...fencedBlock(proposalContent(proposal)),
135
157
  ];
136
158
  if (ctx.liveAsset !== undefined) {
137
- sections.push("", "## Current live asset (would be overwritten on accept)", "```", ctx.liveAsset, "```");
159
+ sections.push("", "## Current live asset (would be overwritten on accept)", ...fencedBlock(ctx.liveAsset));
138
160
  }
139
161
  else {
140
162
  sections.push("", "## Current live asset", "(none — this proposal would create a new asset)");
@@ -142,8 +164,15 @@ export function buildJudgmentPrompt(proposal, reason, ctx) {
142
164
  if (ctx.siblings.length > 0) {
143
165
  sections.push("", "## Other pending proposals for the same ref (dedup context)");
144
166
  for (const sib of ctx.siblings) {
145
- sections.push("", `### Sibling ${sib.id} (source: ${sib.source})`, "```", proposalContent(sib), "```");
167
+ sections.push("", `### Sibling ${sib.id} (source: ${sib.source})`, ...fencedBlock(proposalContent(sib)));
168
+ }
169
+ }
170
+ if (ctx.neighbours && ctx.neighbours.length > 0) {
171
+ sections.push("", "## Existing knowledge notes nearest to this promotion's source memory");
172
+ for (const note of ctx.neighbours) {
173
+ sections.push("", `### ${note.ref}`, note.description, ...fencedBlock(note.excerpt));
146
174
  }
175
+ sections.push("", "Reject the promotion if these notes already cover what it says, even in other words.");
147
176
  }
148
177
  sections.push("", "## Your task", 'Return ONLY a JSON object: {"decision": "accept" | "reject" | "defer", "reason": "<short reason>"}.', "- accept: the proposed content is a correct, valuable update worth committing.", "- reject: the proposal is wrong, a duplicate, or contradicts the live asset.", "- defer: you cannot decide from the provided context (leave it pending).", "Output the JSON object and nothing else.");
149
178
  return sections.join("\n");
@@ -205,7 +234,7 @@ async function dispatchJudgment(runner, prompt, seams) {
205
234
  * (under `applyMode` and the remaining accept budget) or the reject. A defer, an
206
235
  * unparseable verdict or a runner error leaves the item undecided.
207
236
  */
208
- async function runJudgmentTier(opts, result, pending, acceptBudget, promoteFn, rejectFn, seams) {
237
+ async function runJudgmentTier(opts, result, pending, acceptBudget, promoteFn, rejectFn, seams, deferNotes) {
209
238
  const byId = new Map(pending.map((p) => [p.id, p]));
210
239
  const notices = noticeSet();
211
240
  const stillDeferred = [];
@@ -216,21 +245,33 @@ async function runJudgmentTier(opts, result, pending, acceptBudget, promoteFn, r
216
245
  stillDeferred.push(item);
217
246
  continue;
218
247
  }
248
+ const liveAsset = readLiveAssetContent(opts.stashDir, proposal.ref);
219
249
  const prompt = buildJudgmentPrompt(proposal, item.reason, {
220
- liveAsset: readLiveAssetContent(opts.stashDir, proposal.ref),
250
+ liveAsset,
221
251
  siblings: pending.filter((p) => p.ref === proposal.ref && p.id !== proposal.id),
252
+ // A create the model otherwise judges blind: the model never sees knowledge/.
253
+ ...(liveAsset === undefined ? { neighbours: promotionNeighbours(opts.stashDir, proposal) } : {}),
222
254
  });
223
255
  const dispatch = await dispatchJudgment(opts.judgment, prompt, seams);
224
256
  notices.add(dispatch.notices);
225
257
  if (dispatch.error)
226
258
  warn(`[triage] judgment dispatch failed for ${item.id}: ${dispatch.error}`);
227
259
  const verdict = dispatch.error ? null : dispatch.verdict;
228
- if (!verdict || verdict.decision === "defer") {
260
+ if (!verdict) {
261
+ deferNotes.set(item.id, { reason: dispatch.error ? "judgment-error" : "judgment-parse-failure" });
262
+ stillDeferred.push(item);
263
+ continue;
264
+ }
265
+ if (verdict.decision === "defer") {
266
+ deferNotes.set(item.id, {
267
+ reason: "judgment-deferred",
268
+ ...(verdict.reason ? { judgeReason: verdict.reason } : {}),
269
+ });
229
270
  stillDeferred.push(item);
230
271
  continue;
231
272
  }
232
273
  if (verdict.decision === "reject") {
233
- const failure = await rejectProposal(opts, item.id, verdict.reason || "judgment: reject", "judgment-reject", rejectFn);
274
+ const failure = await rejectProposal(opts, item.id, verdict.reason || "judgment: reject", "judgment-reject", rejectFn, verdict.reason);
234
275
  if (failure === undefined) {
235
276
  result.rejected.push(item.id);
236
277
  }
@@ -252,6 +293,7 @@ async function runJudgmentTier(opts, result, pending, acceptBudget, promoteFn, r
252
293
  reason: "judgment-accept",
253
294
  contentHash: proposalContentHash(proposal),
254
295
  gate: DRAIN_GATE,
296
+ ...(verdict.reason ? { judgeReason: verdict.reason } : {}),
255
297
  });
256
298
  result.staged.push(item.id);
257
299
  }
@@ -265,7 +307,7 @@ async function runJudgmentTier(opts, result, pending, acceptBudget, promoteFn, r
265
307
  result.skippedByCap.push(item.id);
266
308
  continue;
267
309
  }
268
- const outcome = await acceptProposal(opts, proposal, item.id, "judgment-accept", promoteFn, rejectFn);
310
+ const outcome = await acceptProposal(opts, proposal, item.id, "judgment-accept", promoteFn, rejectFn, verdict.reason);
269
311
  if (outcome === "promoted") {
270
312
  result.promoted.push(item.id);
271
313
  acceptBudget -= 1;
@@ -286,6 +328,21 @@ async function runJudgmentTier(opts, result, pending, acceptBudget, promoteFn, r
286
328
  result.notices = notices.list();
287
329
  result.deferred = stillDeferred;
288
330
  }
331
+ /** The knowledge notes nearest to a consolidate promotion's source memory; none for any other proposal. */
332
+ function promotionNeighbours(stashDir, proposal) {
333
+ if (proposal.source !== "consolidate" || proposal.promotionSource === undefined)
334
+ return [];
335
+ try {
336
+ const parsed = parseRefInput(proposal.promotionSource);
337
+ const typeDir = stashDirFor(parsed.type);
338
+ if (!typeDir)
339
+ return [];
340
+ return nearestKnowledgeNotes(assetPathForName(parsed.type, path.join(stashDir, typeDir), parsed.name));
341
+ }
342
+ catch {
343
+ return [];
344
+ }
345
+ }
289
346
  /** The live asset a proposal would overwrite, if any. */
290
347
  function readLiveAssetContent(stashDir, ref) {
291
348
  try {
@@ -314,8 +371,8 @@ export async function drainProposals(opts, promoteFn = akmProposalAccept, reject
314
371
  const empties = [];
315
372
  for (const proposal of pending) {
316
373
  // A consolidate pair-pass `retire` proposal is auto-accepted only when the
317
- // pair judge staged it as a duplicate with nothing unique on either side
318
- // (spec §25.9, equivalent content); every other one waits for a direct
374
+ // pair pass staged it: nothing unique on either side, confirmed by a
375
+ // second look, no continuity risk; every other one waits for a direct
319
376
  // `akm proposal accept` (spec §25.6). Checked before isEmptyDiff, which
320
377
  // has nothing meaningful to read on a delete-primary change.
321
378
  if (isRetireProposal(proposal)) {
@@ -324,7 +381,7 @@ export async function drainProposals(opts, promoteFn = akmProposalAccept, reject
324
381
  staged.gate === PAIR_PASS_GATE &&
325
382
  staged.contentHash === proposalContentHash(proposal) &&
326
383
  !proposal.retirement?.continuityRisk) {
327
- accepts.push({ id: proposal.id, reason: "duplicate" });
384
+ accepts.push({ id: proposal.id, reason: staged.reason });
328
385
  }
329
386
  continue;
330
387
  }
@@ -394,15 +451,22 @@ export async function drainProposals(opts, promoteFn = akmProposalAccept, reject
394
451
  }
395
452
  }
396
453
  }
454
+ const deferNotes = new Map();
397
455
  if (opts.judgment && result.deferred.length > 0) {
398
- await runJudgmentTier({ ...opts, judgment: opts.judgment }, result, pending, cap - promotedHere, promoteFn, rejectFn, judgmentSeams);
456
+ await runJudgmentTier({ ...opts, judgment: opts.judgment }, result, pending, cap - promotedHere, promoteFn, rejectFn, judgmentSeams, deferNotes);
399
457
  }
400
458
  // #577: whatever stays undecided is left for review (`review_needed` in the ledger).
401
459
  if (!opts.dryRun) {
402
- const reviewReason = opts.judgment ? "judgment-deferred" : "no-judge-configured";
403
460
  for (const item of result.deferred) {
461
+ const note = deferNotes.get(item.id);
462
+ const reviewReason = note?.reason ?? (opts.judgment ? "judgment-deferred" : "no-judge-configured");
404
463
  try {
405
- recordGateDecision(opts.stashDir, item.id, { outcome: "deferred", reason: reviewReason, gate: DRAIN_GATE });
464
+ recordGateDecision(opts.stashDir, item.id, {
465
+ outcome: "deferred",
466
+ reason: reviewReason,
467
+ gate: DRAIN_GATE,
468
+ ...(note?.judgeReason ? { judgeReason: note.judgeReason } : {}),
469
+ });
406
470
  }
407
471
  catch (err) {
408
472
  warn(`[triage] failed to record gate decision for ${item.id}: ${errMessage(err)}`);
@@ -53,7 +53,7 @@ export function isRetireProposal(proposal) {
53
53
  }
54
54
  /** A promote refused because the target changed after mint (STALE, R20) — not a merit judgement. */
55
55
  export const STALE_TARGET_GATE_REASON = "stale-target";
56
- /** The gate on a retire proposal the triage drain may accept unattended: a pair-judged duplicate. */
56
+ /** The gate on a retire proposal the triage drain may accept unattended: a staged pair-judged retirement. */
57
57
  export const PAIR_PASS_GATE = "consolidate-pair";
58
58
  export const EXPIRED_GATE_REASON = "expired";
59
59
  export const ASSET_MISSING_GATE_REASON = "asset-missing";
@@ -78,9 +78,9 @@ const qualityGateField = z
78
78
  .optional();
79
79
  /**
80
80
  * WS-3b: CLS (Complementary Learning System) interleaving (step 9).
81
- * distill/memoryInference prompts include embedding-retrieved existing adjacent
82
- * lessons/knowledge to prevent catastrophic interference with prior generalizations.
83
- * Default OFF. Only meaningful on `distill` and `memoryInference` processes.
81
+ * The distill prompt includes the lessons, knowledge notes and skills the library already holds near the
82
+ * memory, so the writer answers NONE for a rule one of them states and does not overwrite a prior
83
+ * generalization. Default ON; `enabled: false` turns it off. Only meaningful on the `distill` process.
84
84
  */
85
85
  const clsField = z
86
86
  .object({
@@ -332,15 +332,6 @@ export function getStashStateKey(stashDir) {
332
332
  function stashScopedDir(base, stashDir) {
333
333
  return path.join(base, getStashStateKey(stashDir));
334
334
  }
335
- /**
336
- * `$STATE/improve/measurement/verdicts/<stash>/` — `akm-eval-proactive-verdict`
337
- * reports. Moved out of `$STASH/.akm/measurement/verdicts/` (itlackey/akm#890);
338
- * the pilot treatment file at `$STASH/.akm/measurement/` is manually-authored
339
- * measurement input and stays put.
340
- */
341
- export function getMeasurementVerdictsDir(stashDir) {
342
- return stashScopedDir(path.join(getStateDir(), "improve", "measurement", "verdicts"), stashDir);
343
- }
344
335
  /**
345
336
  * `$CACHE/index/unresolved-sources/<stash>/` — synthetic placeholder path for
346
337
  * a configured source whose content root did not resolve this run. Never
@@ -50,11 +50,14 @@ export const CONSOLIDATE_LEDGER_SOURCE = "consolidate";
50
50
  * promotion is a verdict on that memory's text; asking the model about the
51
51
  * same text again can only reproduce the proposal (accepted used to be
52
52
  * eligible at once, rejected after 7 days), so the memory waits for an edit.
53
+ * So does a memory the model judged and left alone (`judged_no_action`):
54
+ * the same text would be judged weekly with the same answer.
53
55
  * The clock stays for a row with no recorded hash — one decided before the
54
56
  * hash was recorded — see {@link nextEligibleAt} and {@link isContentDrivenRow}.
55
57
  */
56
58
  export function isContentDrivenDecision(source, outcome) {
57
- return source === CONSOLIDATE_LEDGER_SOURCE && (outcome === "accepted" || outcome === "rejected");
59
+ return (source === CONSOLIDATE_LEDGER_SOURCE &&
60
+ (outcome === "accepted" || outcome === "rejected" || outcome === "judged_no_action"));
58
61
  }
59
62
  /**
60
63
  * Whether this row is held by its content hash: a decided consolidate
package/docs/README.md CHANGED
@@ -49,7 +49,6 @@ Working on akm itself, not just using it.
49
49
 
50
50
  - [Maintainer Docs](https://github.com/itlackey/akm/blob/main/docs/maintainers/README.md) -- Start here: local development, measuring improvement, and the curate contract
51
51
  - [Local Development](https://github.com/itlackey/akm/blob/main/docs/maintainers/local-development.md) -- Dogfooding akm while editing its own source
52
- - [akm-eval](https://github.com/itlackey/akm/blob/main/docs/maintainers/eval.md) -- Standalone toolkit for measuring whether `akm improve` is working
53
52
  - [Curate Workmap](https://github.com/itlackey/akm/blob/main/docs/maintainers/curate-workmap.md) -- The current `akm curate` contract and the highest-value next fixes
54
53
 
55
54
  ## Look up details
@@ -100,7 +99,7 @@ Source articles for the dev.to publishing pipeline (historical record). See
100
99
  - [itlackey/akm-registry](https://github.com/itlackey/akm-registry) -- the official registry index that powers built-in discovery
101
100
  - [itlackey/akm-plugins](https://github.com/itlackey/akm-plugins) -- optional integrations for tools like OpenCode
102
101
  - [itlackey/akm-bench](https://github.com/itlackey/akm-bench) -- the standalone benchmark harness for measuring agent performance with akm
103
- - [itlackey/akm-eval](https://github.com/itlackey/akm-eval) -- the eval framework and tools for akm asset quality (distinct from the in-repo [`scripts/akm-eval/` toolkit](https://github.com/itlackey/akm/blob/main/docs/maintainers/eval.md))
102
+ - [itlackey/akm-eval](https://github.com/itlackey/akm-eval) -- the eval framework and tools for akm asset quality
104
103
 
105
104
  ---
106
105
 
@@ -266,7 +266,7 @@ rather than rely on `$HOME`-derived defaults (names verified against
266
266
  | `AKM_CONFIG_DIR` | `config.json`'s directory. |
267
267
  | `AKM_DATA_DIR` | Durable, non-regenerable data: **`index.db` and `state.db` live here.** This is the directory a migration snapshot's safety copy sits beside. |
268
268
  | `AKM_CACHE_DIR` | Regenerable cache: registry downloads, config backups, task logs. Safe to discard between image builds (not between boots of the same running install). |
269
- | `AKM_STATE_DIR` | **Not** where `state.db` lives, despite the name — this is the XDG "state" directory. Holds scheduled-task invocation context, companion-plugin hook state (Claude Code / OpenCode hook logs), and, per stash, `akm improve`'s machine-local writers (`improve/measurement/verdicts/`) and whole-run lock (`locks/`) — see [Storage locations](https://github.com/itlackey/akm/blob/main/docs/architecture/internals/storage-locations.md). Set it anyway if you schedule akm tasks inside the image, so that context is captured consistently rather than falling back to `$HOME/.local/state/akm`. |
269
+ | `AKM_STATE_DIR` | **Not** where `state.db` lives, despite the name — this is the XDG "state" directory. Holds scheduled-task invocation context, companion-plugin hook state (Claude Code / OpenCode hook logs), and, per stash, `akm improve`'s whole-run lock (`locks/`) — see [Storage locations](https://github.com/itlackey/akm/blob/main/docs/architecture/internals/storage-locations.md). Set it anyway if you schedule akm tasks inside the image, so that context is captured consistently rather than falling back to `$HOME/.local/state/akm`. |
270
270
 
271
271
  Set all five to paths that persist across container restarts (a mounted
272
272
  volume), or `akm migrate apply` will see an empty `state.db` on every boot
@@ -144,7 +144,9 @@ step is idempotent — a second run reports nothing pending. 0.9.17-alpha.4
144
144
  removed that step: `akm migrate` no longer relocates these files, and one left
145
145
  at an old path is inert (nothing reads it). Two of the five writers no longer
146
146
  exist either — the improve ledger replaced `distill-rejected/`, and the
147
- write-only `eval-cases/` path was removed.
147
+ write-only `eval-cases/` path was removed. A third, `measurement/verdicts/`, went
148
+ with the `scripts/akm-eval` toolkit that wrote it, which has since left this
149
+ repository ([akm-eval](https://github.com/itlackey/akm/blob/main/docs/maintainers/eval.md)).
148
150
  `$STASH/.akm/memory-cleanup/` did not move; it is the one confirmed exception
149
151
  to the rule (see Storage locations, above).
150
152
 
@@ -17,4 +17,4 @@ Authoritative reference documentation for the akm CLI and its data.
17
17
  - [Website Sources](https://github.com/itlackey/akm/blob/main/docs/reference/website-sources.md) -- The pluggable fetcher API behind `akm import <url>` and other URL-based knowledge reads
18
18
  - [Data & Telemetry](data-and-telemetry.md) -- Exactly what akm reads and writes on your machine (no remote telemetry)
19
19
 
20
- See also: [akm-eval](https://github.com/itlackey/akm/blob/main/docs/maintainers/eval.md) -- the standalone toolkit for measuring whether `akm improve` is working (maintainer docs), and the repo-root [Roadmap](https://github.com/itlackey/akm/blob/main/ROADMAP.md) -- high-level focus for upcoming releases.
20
+ See also: [akm-eval](https://github.com/itlackey/akm-eval) -- the evals and benchmarks for measuring whether `akm improve` is working, and the repo-root [Roadmap](https://github.com/itlackey/akm/blob/main/ROADMAP.md) -- high-level focus for upcoming releases.
@@ -2492,16 +2492,19 @@ day; an asset a stage looked at and left unchanged is revisited after 7 days,
2492
2492
  or as soon as new feedback (or, for consolidation, an edit) arrives.
2493
2493
 
2494
2494
  Consolidation's promotion of a memory into `knowledge/` is the exception to the
2495
- 7-day rule: once a promotion is accepted or rejected, its memory is not offered
2496
- to the model again until its body changes, however long that takes. The ledger
2495
+ 7-day rule: once a promotion is accepted or rejected, or the model judged the
2496
+ memory and proposed nothing, the memory is not offered to the model again until
2497
+ its body changes, however long that takes. The ledger
2497
2498
  records the body hash the promotion was decided against and compares it with
2498
2499
  the memory's current body (frontmatter edits do not count), the same
2499
2500
  content-driven rule the consolidate pair pass uses. A promotion decided by an
2500
- older release, which recorded no hash, keeps the old windows. Consolidation
2501
+ older release, which recorded no hash, keeps the old windows. A memory whose
2502
+ body equals that of a consolidate promotion rejected on or after 2026-09-29 is
2503
+ held the same way, under whatever name it has. Consolidation
2501
2504
  also does not promote a memory that `knowledge/` already covers: before it
2502
2505
  queues a promotion it compares the memory with the 20 `knowledge/` docs in its
2503
2506
  bundle nearest to it by stored vector, and skips the memory when one of them
2504
- holds at least half of its distinct 5-word shingles (skip reason
2507
+ holds at least 30% of its distinct 5-word shingles (skip reason
2505
2508
  `dedup_covered_by_knowledge` in the result's `consolidation.skipReasons`). A
2506
2509
  covering doc that ranks lower than the 20th nearest goes unseen. With no stored
2507
2510
  vector (semantic search off, or the memory not indexed yet) that check does
@@ -195,7 +195,7 @@ the set of types the code actually emits at HEAD (verified against every
195
195
  | `reflect_completed` | Reflect phase produced a proposal | `ref` |
196
196
  | `improve_reflect_outcome` | Per-asset reflect result | `ref`, `ok`, `durationMs`, `reason` |
197
197
  | `propose_invoked` | `akm proposal new` | `ref` |
198
- | `distill_invoked` | Distill phase inside the `akm improve`/`akm proposal new` pipeline. **`akm distill` is not a CLI command** — there is no standalone verb by that name | `ref`, outcome (`queued`, `skipped` with a `skipReason` such as `lesson_exists` or `conflict_noop`, `llm_failed`, `validation_failed`, `quality_rejected`, `review_needed`) |
198
+ | `distill_invoked` | Distill phase inside the `akm improve`/`akm proposal new` pipeline. **`akm distill` is not a CLI command** — there is no standalone verb by that name | `ref`, outcome (`queued`, `skipped` with a `skipReason` such as `lesson_exists`, `nothing_reusable` or `conflict_noop`, `llm_failed`, `validation_failed`, `quality_rejected`, `review_needed`) |
199
199
  | `extract_invoked` | `akm proposal extract --type <harness>` / `--auto`, or improve-stage session extraction | `outcome`, `sessionId`, `harness` |
200
200
  | `extract_triaged` | The pre-LLM extract triage gate evaluated at least one session | `evaluated`, `passed`, `triagedOut`, `sourceRun` (aggregated) |
201
201
  | `schema_repair_invoked` | The schema-repair pass inside `akm improve` (`runSchemaRepairPass`) attempts to patch missing frontmatter on an asset that failed schema validation. **There is no `akm lint --repair` flag** — `lint` has `--fix`/`--auto-fix`, unrelated to this event | `ref`, outcome |
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "akm-cli",
3
- "version": "0.9.27-alpha.1",
3
+ "version": "0.9.27-alpha.2",
4
4
  "type": "module",
5
5
  "description": "akm (Agent Knowledge Manager) — a portable, local-first capability library for AI agents. Discover, load, share, and improve reusable skills, scripts, workflows, and knowledge across any shell-capable coding agent, including Claude Code, OpenCode, and Cursor.",
6
6
  "keywords": [