akm-cli 0.9.27-alpha.1 → 0.9.27-alpha.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +112 -0
- package/LICENSE +3 -4
- package/dist/assets/prompts/distill-lesson-system.md +29 -7
- package/dist/commands/improve/consolidate/coverage.js +71 -17
- package/dist/commands/improve/consolidate/pair-pass.js +11 -8
- package/dist/commands/improve/consolidate.js +63 -1
- package/dist/commands/improve/distill-guards.js +9 -10
- package/dist/commands/improve/distill.js +62 -18
- package/dist/commands/improve/preparation.js +14 -1
- package/dist/commands/improve/stage.js +34 -52
- package/dist/commands/proposal/drain.js +85 -21
- package/dist/commands/proposal/proposal-types.js +1 -1
- package/dist/core/config/schema/improve-processes.js +3 -3
- package/dist/core/paths.js +0 -9
- package/dist/storage/repositories/improve-ledger-repository.js +4 -1
- package/docs/README.md +1 -2
- package/docs/integration/bundling-akm.md +1 -1
- package/docs/migration/v0.8-to-v0.9.md +3 -1
- package/docs/reference/README.md +1 -1
- package/docs/reference/cli.md +7 -4
- package/docs/reference/data-and-telemetry.md +1 -1
- package/package.json +1 -1
package/CHANGELOG.md
CHANGED
|
@@ -6,6 +6,118 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/).
|
|
|
6
6
|
|
|
7
7
|
## [Unreleased]
|
|
8
8
|
|
|
9
|
+
## [0.9.27-alpha.2] - 2026-10-07
|
|
10
|
+
|
|
11
|
+
### Removed
|
|
12
|
+
|
|
13
|
+
- **The `scripts/akm-eval` toolkit has moved out of this repository.** The
|
|
14
|
+
read-only measurement toolkit (the case runner and its suites, the twin
|
|
15
|
+
experiment, the real-query verdict for the proactive lane, the state
|
|
16
|
+
analyzers and the curate benchmark) is retired; every live eval is in
|
|
17
|
+
[itlackey/akm-eval](https://github.com/itlackey/akm-eval). Its code is kept
|
|
18
|
+
there, to read and not to run, in `retired/akm-scripts-akm-eval/`, copied from
|
|
19
|
+
commit `f57a7fd44b37`. It imports akm's `src/` by relative path, so it runs
|
|
20
|
+
only in a checkout at that commit. Removed here with it: `scripts/akm-eval/`,
|
|
21
|
+
its tests (`tests/integration/akm-eval/`, `tests/akm-eval-*.test.ts`,
|
|
22
|
+
`tests/curate-metrics.test.ts`) and fixtures (`tests/fixtures/akm-eval/`, and
|
|
23
|
+
the `curate-golden` stash, which only the curate benchmark read), the
|
|
24
|
+
`akm-eval determinism` CI job, and `getMeasurementVerdictsDir`, whose only
|
|
25
|
+
caller was the verdict runner. akm no longer names
|
|
26
|
+
`$STATE/improve/measurement/verdicts/<stash>/`; a file already there is inert.
|
|
27
|
+
`docs/maintainers/eval.md` is now a pointer to the new home.
|
|
28
|
+
|
|
29
|
+
### Fixed
|
|
30
|
+
|
|
31
|
+
- **The drain's judge no longer sees a note with a code block as truncated.**
|
|
32
|
+
The judgment prompt fenced the proposed content (and the live asset, sibling
|
|
33
|
+
proposals and neighbour excerpts) in three backticks, so a note holding its
|
|
34
|
+
own code block closed the fence early and read as cut off; real rejections said
|
|
35
|
+
"ends in an empty code block" or "truncated". Each block now uses a fence longer
|
|
36
|
+
than any backtick run inside it. The judge's reason is also kept on accepts,
|
|
37
|
+
staged accepts and defers (as the gate decision's `judgeReason`, until now
|
|
38
|
+
rejections only), and a judge reply that is not a verdict is stamped
|
|
39
|
+
`judgment-parse-failure`, and a runner failure `judgment-error`, instead of
|
|
40
|
+
looking like a defer.
|
|
41
|
+
|
|
42
|
+
- **Distill writes a lesson only when its memory holds one, says only what the
|
|
43
|
+
memory says, and its judge rejects what a reviewer would.** 2 of the 22 distill
|
|
44
|
+
proposals since 0.9.26 began were accepted, and 17 of the 19 queued on
|
|
45
|
+
2026-10-05 were bad (they restated their memory, filed a dated status as a
|
|
46
|
+
lesson, claimed what the memory does not say, or repeated an asset the library
|
|
47
|
+
holds). Four causes, found in the code and the rejected proposals, and fixed:
|
|
48
|
+
(1) the prompt and schema forced a lesson from every memory, and 18 of the 19
|
|
49
|
+
were records of what was done; the writer now says why a memory holds a
|
|
50
|
+
lesson or none (`reason`, then `decision: lesson|none`, or the word `NONE`),
|
|
51
|
+
defined as a cause and what to do about it, or a rule with its reason, and
|
|
52
|
+
writes only what the memory and its feedback state, in the scope they have; a
|
|
53
|
+
`NONE` is a `skipped` distill (`skipReason: nothing_reusable`, the writer's
|
|
54
|
+
reason in the message) with no proposal and no judge call, and the loop keeps
|
|
55
|
+
its ledger row `unchanged`. (2) The judge asked for "information not already
|
|
56
|
+
present in the source", so an invented claim scored as novel and a faithful
|
|
57
|
+
lesson of a lesson-worthy memory as a restatement, and it passed anything that
|
|
58
|
+
"goes beyond the source" because it "may draw on feedback you are not shown".
|
|
59
|
+
The rubric is now reusable (a rule with its reason, not a record of what was
|
|
60
|
+
done), non-redundancy and grounding (every cause, step, number and limit is
|
|
61
|
+
in the source or its feedback), and the judge is shown the feedback the writer
|
|
62
|
+
saw. (3) A mean hid a decisive score (4 and 1 average 2.5, a review), and a
|
|
63
|
+
reviewer read everything the judge did not reject; any criterion at 2 or
|
|
64
|
+
below, grounding included, is now `quality_rejected`, and the reason names it
|
|
65
|
+
(`grounding 2/5: …`). The "borderline grounding" routing is gone. (4) Neither
|
|
66
|
+
the writer nor the judge could see a knowledge note or a skill that already
|
|
67
|
+
states the rule (the judge saw the 3 lexically nearest lessons, none of them
|
|
68
|
+
related); both now see the lessons, knowledge notes and skills nearest the
|
|
69
|
+
memory, which is the existing `processes.distill.cls` context turned on by
|
|
70
|
+
default (`enabled: false` turns it off). Judge scores are keyed `reusable`
|
|
71
|
+
where they were `novelty`. Measured on the local qwen3.8-27b with akm-eval's
|
|
72
|
+
`evals/distill` (30 fictional memories, 5 runs each side): good lessons 7/14 on
|
|
73
|
+
average (5 to 9) against 3.7/14 (3 to 5), lessons queued for memories that
|
|
74
|
+
deserve none 0.4/16 against 3.3/16. On 37 real memories with their feedback
|
|
75
|
+
(34 reviewed bad, 3 good; 3 runs against 2): a lesson was queued for 11% of the
|
|
76
|
+
bad ones against 44%, and for 6 of 9 good ones against 4 of 6; of the memories
|
|
77
|
+
that pass 0.9.26's skip of bare positive feedback, 20% of the bad against 50%.
|
|
78
|
+
No new settings.
|
|
79
|
+
|
|
80
|
+
- **Reflect no longer plans an asset whose negative feedback is already acted
|
|
81
|
+
on.** A negative `akm feedback` that came with an exact fix (`--replace` and
|
|
82
|
+
`--with`, `--outdated` or `--superseded-by`) makes a `feedback` proposal, and
|
|
83
|
+
once that proposal is accepted the feedback has done its work. Reflect still
|
|
84
|
+
took the ref as having fresh negative feedback, and on 2026-10-07 44 of its 50
|
|
85
|
+
refs were of that kind: the judge refused or the model changed nothing for
|
|
86
|
+
most of them. A negative event with a fix is now left out of the reflect
|
|
87
|
+
cursor when an accepted `feedback` proposal for the ref was created at or
|
|
88
|
+
after it. A negative with no fix, one given after the proposal, and one whose
|
|
89
|
+
proposal is still pending or was rejected plan a reflect as before.
|
|
90
|
+
|
|
91
|
+
- **Consolidate stops re-offering memories a reviewer already turned down, and
|
|
92
|
+
the nightly judge sees what a promotion may duplicate.** About 53 promotions a
|
|
93
|
+
night reached review at ~5% precision, 64-70% of them a memory body already
|
|
94
|
+
proposed or rejected. Four causes, four changes: a memory whose body equals
|
|
95
|
+
that of a consolidate promotion rejected on or after 2026-09-29 is held until
|
|
96
|
+
its body changes, under any name (earlier rejections, the bulk audits of
|
|
97
|
+
2026-08, do not count); a memory the model judged and left alone is held by its
|
|
98
|
+
body hash instead of a 7-day clock, so an unchanged memory is no longer judged
|
|
99
|
+
every week (a row recorded without a hash keeps the 7 days); the coverage gate
|
|
100
|
+
skips a memory when 30% of its text, not 50%, is in a neighbouring knowledge
|
|
101
|
+
doc, which catches paraphrases; and the drain's judgment tier, which judged a
|
|
102
|
+
promotion seeing only the proposal and never `knowledge/`, is now shown the 5
|
|
103
|
+
nearest knowledge notes (ref, description, excerpt) and told to reject a
|
|
104
|
+
promotion they already cover. No new settings.
|
|
105
|
+
|
|
106
|
+
- **A confident `subsumed` or `supersedes` retirement resolves unattended, as a
|
|
107
|
+
`duplicate` already did.** The pair pass staged a retire proposal for the
|
|
108
|
+
triage drain only when the judge's label was `duplicate`; every other retirement
|
|
109
|
+
waited for a person. It now stages any of the three retire labels when the
|
|
110
|
+
second look (what does the retired note hold that the kept one lacks?) comes
|
|
111
|
+
back empty and there is no continuity risk; the retired side's claim list is
|
|
112
|
+
already empty for any proposal, and a `duplicate` must still have an empty list
|
|
113
|
+
on the kept side too. The staged gate reason is the judge's label, and the drain records it. Replay over 360
|
|
114
|
+
judged pairs: 336 safe (0.93): `duplicate` 0.98, `subsumed` 0.92, `supersedes`
|
|
115
|
+
0.875. In production, unstaged `subsumed` retirements were accepted 45 of 54
|
|
116
|
+
times by hand, and in the latest run 41 of 47 pair proposals would have resolved
|
|
117
|
+
without a person. `docs/architecture/internals/improve-workflow.md` said triage
|
|
118
|
+
never auto-accepts a retire proposal, which stopped being true in 0.9.26; it,
|
|
119
|
+
and the matching lines in `improvement.md`, now describe the staging rule.
|
|
120
|
+
|
|
9
121
|
## [0.9.27-alpha.1] - 2026-10-06
|
|
10
122
|
|
|
11
123
|
### Fixed
|
package/LICENSE
CHANGED
|
@@ -225,10 +225,9 @@ statute, judicial order, or regulation then You must: (a) comply with
|
|
|
225
225
|
the terms of this License to the maximum extent possible; and (b)
|
|
226
226
|
describe the limitations and the code they affect. Such description must
|
|
227
227
|
be placed in a text file included with all distributions of the Covered
|
|
228
|
-
Software under
|
|
229
|
-
|
|
230
|
-
|
|
231
|
-
for a recipient of ordinary skill to be able to understand it.
|
|
228
|
+
Software under this License. Except to the extent prohibited by statute
|
|
229
|
+
or regulation, such description must be sufficiently detailed for a
|
|
230
|
+
recipient of ordinary skill to be able to understand it.
|
|
232
231
|
|
|
233
232
|
5. Termination
|
|
234
233
|
--------------
|
|
@@ -1,7 +1,29 @@
|
|
|
1
1
|
You are the akm `distill` distiller.
|
|
2
|
-
|
|
3
|
-
|
|
4
|
-
|
|
2
|
+
You are given a memory and the feedback recorded about it. Decide whether it
|
|
3
|
+
holds a lesson and, if it does, write the lesson.
|
|
4
|
+
|
|
5
|
+
A memory holds a lesson when it states a cause and what to do about it: a
|
|
6
|
+
failure or surprise with its cause and the fix that worked, or a rule with the
|
|
7
|
+
reason it holds. It holds a lesson even when it is short and names one project,
|
|
8
|
+
tool or incident, if the cause and the fix would help someone in a similar
|
|
9
|
+
situation.
|
|
10
|
+
|
|
11
|
+
A memory holds NO lesson when all it states is what was done, shipped, released
|
|
12
|
+
or decided, what is pending or planned, how a system is set up now, or the steps
|
|
13
|
+
of a procedure, with no failure and cause behind it, or when its feedback says
|
|
14
|
+
only that it is out of date or superseded. ANSWER NONE then: the single word and
|
|
15
|
+
nothing else. A reply bound to a JSON schema answers NONE by
|
|
16
|
+
setting `decision` to `none` and leaving the other fields empty. Answer NONE too
|
|
17
|
+
when a related asset listed below the memory already states the rule the memory
|
|
18
|
+
would give.
|
|
19
|
+
|
|
20
|
+
When the memory holds a lesson, write it from what the memory and its feedback
|
|
21
|
+
say, and nothing more:
|
|
22
|
+
- Add no cause, step, rule, number, check or safeguard that neither states.
|
|
23
|
+
- Keep the scope the memory has. A fix verified in one place is a fix for that
|
|
24
|
+
place, and what was not checked stays unchecked. One case is not "always" or
|
|
25
|
+
"never".
|
|
26
|
+
- Be shorter than the memory.
|
|
5
27
|
|
|
6
28
|
YOUR RESPONSE MUST START EXACTLY WITH `---` ON THE VERY FIRST LINE.
|
|
7
29
|
DO NOT output any prose, explanation, or code fences before or after.
|
|
@@ -12,7 +34,7 @@ description: <one complete sentence (ending with `.`) summarising what the lesso
|
|
|
12
34
|
when_to_use: <one complete sentence describing the concrete trigger condition>
|
|
13
35
|
---
|
|
14
36
|
|
|
15
|
-
<lesson body — plain markdown,
|
|
37
|
+
<lesson body — plain markdown, as short as the memory allows>
|
|
16
38
|
|
|
17
39
|
## description field (MANDATORY)
|
|
18
40
|
- A single complete sentence in present tense, 20–400 chars, NO markdown.
|
|
@@ -21,7 +43,7 @@ when_to_use: <one complete sentence describing the concrete trigger condition>
|
|
|
21
43
|
- DO NOT copy a section heading ("Key takeaways", "For example", "Key pitfalls").
|
|
22
44
|
- DO NOT begin with a numbered list marker, code fence, or markdown heading.
|
|
23
45
|
|
|
24
|
-
GOOD: "
|
|
46
|
+
GOOD: "Pin the container image tag, because the `latest` tag moved under the nightly job and its output changed with no code change."
|
|
25
47
|
BAD: "Key pitfalls"
|
|
26
48
|
BAD: "When working with the akm CLI"
|
|
27
49
|
BAD: "For example, you might..."
|
|
@@ -32,5 +54,5 @@ RULES:
|
|
|
32
54
|
- `description` and `when_to_use` MUST differ from each other.
|
|
33
55
|
- The lesson body MUST be non-empty markdown prose. Do NOT restate `description:` or `when_to_use:` inside the body (no `**description:** ...` or `**when_to_use:** ...` lines — the frontmatter is the only place those keys belong).
|
|
34
56
|
- Do NOT emit a second `---` fence after the opening frontmatter — there are exactly two `---` lines in the output, both belonging to the single frontmatter block at the top.
|
|
35
|
-
- Do NOT reproduce the source asset verbatim
|
|
36
|
-
- Output ONLY the lesson file. No preamble, no code fences, no trailing prose.
|
|
57
|
+
- Do NOT reproduce the source asset verbatim.
|
|
58
|
+
- Output ONLY the lesson file. No preamble, no code fences, no trailing prose.
|
|
@@ -6,7 +6,7 @@
|
|
|
6
6
|
* `knowledge/` proposal, ask whether `knowledge/` already says it.
|
|
7
7
|
*
|
|
8
8
|
* The rule: a memory is covered when at least {@link COVERAGE_MIN_CONTAINMENT}
|
|
9
|
-
* (
|
|
9
|
+
* (30%) of its distinct {@link COVERAGE_SHINGLE_WORDS}-word shingles appear in
|
|
10
10
|
* one of the knowledge docs nearest to it. It is a containment of the MEMORY in
|
|
11
11
|
* the doc, not a similarity: a long guide that quotes the memory covers it; a
|
|
12
12
|
* memory that quotes a short doc and adds claims of its own does not.
|
|
@@ -25,9 +25,13 @@
|
|
|
25
25
|
* this gate achieves: it reads only the {@link PAIR_NEIGHBOR_FETCH_K} nearest
|
|
26
26
|
* knowledge docs (below), and a covering doc that ranks lower goes unseen. The
|
|
27
27
|
* recall over that candidate set is unmeasured. The rest of the rejected
|
|
28
|
-
* proposals (paraphrases, partial overlaps) still reach review
|
|
29
|
-
*
|
|
30
|
-
*
|
|
28
|
+
* proposals (paraphrases, partial overlaps) still reach review. 0.5 let
|
|
29
|
+
* paraphrases through: of the 54 promotions minted on 2026-10-07, 18 of the 51
|
|
30
|
+
* later rejected held 30% or more of their text in a neighbouring doc. The cut is
|
|
31
|
+
* now 0.3, between the two measured points: 0.2 (169 of 224 rejected) also
|
|
32
|
+
* skipped 2 of the 103 accepted ones. At 0.3, none of the 5 promotions graded
|
|
33
|
+
* good in the 2026-10-05 review sample would be skipped (their best doc holds
|
|
34
|
+
* at most 1% of them) while 5 of its 15 bad ones would. A wrong skip is a
|
|
31
35
|
* promotion nobody gets to review.
|
|
32
36
|
*
|
|
33
37
|
* Candidates are the memory's {@link PAIR_NEIGHBOR_FETCH_K} nearest knowledge
|
|
@@ -49,7 +53,7 @@ import { PAIR_NEIGHBOR_FETCH_K } from "./pair-pass.js";
|
|
|
49
53
|
/** Words per shingle. */
|
|
50
54
|
export const COVERAGE_SHINGLE_WORDS = 5;
|
|
51
55
|
/** Share of a memory's distinct shingles one knowledge doc must hold for the memory to count as covered. */
|
|
52
|
-
export const COVERAGE_MIN_CONTAINMENT = 0.
|
|
56
|
+
export const COVERAGE_MIN_CONTAINMENT = 0.3;
|
|
53
57
|
const WORD = /[\p{L}\p{N}]+/gu;
|
|
54
58
|
/** The distinct lower-cased word n-grams of `text`; empty when it has fewer than {@link COVERAGE_SHINGLE_WORDS} words. */
|
|
55
59
|
export function wordShingles(text) {
|
|
@@ -72,19 +76,15 @@ export function shingleContainment(memory, doc) {
|
|
|
72
76
|
return shared / memory.size;
|
|
73
77
|
}
|
|
74
78
|
/**
|
|
75
|
-
* The
|
|
76
|
-
*
|
|
77
|
-
*
|
|
78
|
-
* and so no candidates.
|
|
79
|
+
* The {@link PAIR_NEIGHBOR_FETCH_K} knowledge docs in `bundleId` nearest to the
|
|
80
|
+
* memory at `filePath`, nearest first. A memory the index does not know has no
|
|
81
|
+
* stored vector and so no neighbours.
|
|
79
82
|
*/
|
|
80
|
-
|
|
81
|
-
const shingles = wordShingles(body);
|
|
82
|
-
if (shingles.size === 0)
|
|
83
|
-
return undefined;
|
|
83
|
+
function knowledgeNeighbours(db, bundleId, filePath) {
|
|
84
84
|
const entryId = getEntryIdByFilePath(db, filePath);
|
|
85
85
|
if (entryId === undefined)
|
|
86
|
-
return
|
|
87
|
-
|
|
86
|
+
return [];
|
|
87
|
+
const out = [];
|
|
88
88
|
for (const hit of getNeighborsByEntryId(db, entryId, PAIR_NEIGHBOR_FETCH_K, { type: "knowledge", bundleId })) {
|
|
89
89
|
const neighbour = getEntryById(db, hit.id);
|
|
90
90
|
if (!neighbour)
|
|
@@ -96,13 +96,67 @@ export function findCoveringKnowledge(db, bundleId, filePath, body) {
|
|
|
96
96
|
catch {
|
|
97
97
|
continue; // the index outlived the file
|
|
98
98
|
}
|
|
99
|
-
|
|
99
|
+
out.push({
|
|
100
|
+
ref: neighbour.conceptId,
|
|
101
|
+
description: neighbour.entry.description ?? "",
|
|
102
|
+
body: stripFrontmatterBody(raw),
|
|
103
|
+
});
|
|
104
|
+
}
|
|
105
|
+
return out;
|
|
106
|
+
}
|
|
107
|
+
/**
|
|
108
|
+
* The best-covering knowledge doc among the {@link PAIR_NEIGHBOR_FETCH_K}
|
|
109
|
+
* knowledge docs in `bundleId` nearest to the memory. `filePath` is the
|
|
110
|
+
* memory's indexed file; a memory the index does not know has no stored
|
|
111
|
+
* vector and so no candidates.
|
|
112
|
+
*/
|
|
113
|
+
export function findCoveringKnowledge(db, bundleId, filePath, body) {
|
|
114
|
+
const shingles = wordShingles(body);
|
|
115
|
+
if (shingles.size === 0)
|
|
116
|
+
return undefined;
|
|
117
|
+
let best;
|
|
118
|
+
for (const neighbour of knowledgeNeighbours(db, bundleId, filePath)) {
|
|
119
|
+
const containment = shingleContainment(shingles, neighbour.body);
|
|
100
120
|
if (containment >= COVERAGE_MIN_CONTAINMENT && (best === undefined || containment > best.containment)) {
|
|
101
|
-
best = { ref: neighbour.
|
|
121
|
+
best = { ref: neighbour.ref, containment };
|
|
102
122
|
}
|
|
103
123
|
}
|
|
104
124
|
return best;
|
|
105
125
|
}
|
|
126
|
+
/** Knowledge docs a reviewer is shown for a promotion, nearest first. */
|
|
127
|
+
export const NEIGHBOUR_NOTE_COUNT = 5;
|
|
128
|
+
const NEIGHBOUR_EXCERPT_CHARS = 300;
|
|
129
|
+
/**
|
|
130
|
+
* The knowledge notes nearest to the memory at `memoryPath`, for the drain's
|
|
131
|
+
* judge to compare a promotion against: the nearest {@link NEIGHBOUR_NOTE_COUNT}
|
|
132
|
+
* of the same candidates the coverage gate reads. The memory's bundle is the one
|
|
133
|
+
* the index recorded for it. Empty when the index has no vector for the memory
|
|
134
|
+
* or cannot be opened; never throws.
|
|
135
|
+
*/
|
|
136
|
+
export function nearestKnowledgeNotes(memoryPath) {
|
|
137
|
+
let db;
|
|
138
|
+
try {
|
|
139
|
+
db = openExistingDatabase();
|
|
140
|
+
const entryId = getEntryIdByFilePath(db, memoryPath);
|
|
141
|
+
const bundleId = entryId === undefined ? undefined : getEntryById(db, entryId)?.bundleId;
|
|
142
|
+
if (bundleId === undefined)
|
|
143
|
+
return [];
|
|
144
|
+
return knowledgeNeighbours(db, bundleId, memoryPath)
|
|
145
|
+
.slice(0, NEIGHBOUR_NOTE_COUNT)
|
|
146
|
+
.map((n) => ({
|
|
147
|
+
ref: n.ref,
|
|
148
|
+
description: n.description,
|
|
149
|
+
excerpt: n.body.length > NEIGHBOUR_EXCERPT_CHARS ? `${n.body.slice(0, NEIGHBOUR_EXCERPT_CHARS)}...` : n.body,
|
|
150
|
+
}));
|
|
151
|
+
}
|
|
152
|
+
catch {
|
|
153
|
+
return [];
|
|
154
|
+
}
|
|
155
|
+
finally {
|
|
156
|
+
if (db)
|
|
157
|
+
closeDatabase(db);
|
|
158
|
+
}
|
|
159
|
+
}
|
|
106
160
|
/**
|
|
107
161
|
* The gate for one run, holding its own read handle on `index.db` for the
|
|
108
162
|
* promotions that run emits; `undefined` when there is no bundle or no index
|
|
@@ -503,7 +503,7 @@ function checkSection(label, side) {
|
|
|
503
503
|
].join("\n");
|
|
504
504
|
}
|
|
505
505
|
/**
|
|
506
|
-
* The second look a
|
|
506
|
+
* The second look a retirement gets before it may retire unattended: one call
|
|
507
507
|
* that asks only what the retired note holds that the kept one lacks. True
|
|
508
508
|
* only on a clean, empty answer (it caught 2 of 4 duplicates the judge got
|
|
509
509
|
* wrong, and held back none of 109 right ones).
|
|
@@ -658,17 +658,20 @@ async function judgeOne(ctx, candidate) {
|
|
|
658
658
|
}, ctx.opts.proposalsCtx);
|
|
659
659
|
ctx.retired.push(proposal.id);
|
|
660
660
|
ctx.perInitiatorProposed.add(candidate.initiator.ref);
|
|
661
|
-
// A
|
|
662
|
-
//
|
|
663
|
-
//
|
|
664
|
-
//
|
|
665
|
-
|
|
666
|
-
|
|
661
|
+
// A retirement the second look confirms loses nothing retires unattended:
|
|
662
|
+
// the triage drain accepts it under its usual applyMode. The retired side
|
|
663
|
+
// holds no claim of its own for any label (`decideRetirement` mints nothing
|
|
664
|
+
// else); a duplicate must also leave the kept side with none, while a
|
|
665
|
+
// subsumed or superseding successor holds more by definition. Replay
|
|
666
|
+
// precision 336/360 (duplicate 0.98, subsumed 0.92, supersedes 0.875,
|
|
667
|
+
// 2026-10-07); a duplicate alone was 109 of 111 safe on the owner's
|
|
668
|
+
// reviewed pairs (2026-10-04). Anything else waits for a person.
|
|
669
|
+
if ((verdict.relation !== "duplicate" || verdict.onlyInA.length + verdict.onlyInB.length === 0) &&
|
|
667
670
|
!continuityRisk &&
|
|
668
671
|
(await confirmNothingLost(ctx, retired, successor))) {
|
|
669
672
|
recordGateDecision(ctx.stashDir, proposal.id, {
|
|
670
673
|
outcome: "staged",
|
|
671
|
-
reason:
|
|
674
|
+
reason: verdict.relation,
|
|
672
675
|
gate: PAIR_PASS_GATE,
|
|
673
676
|
contentHash: proposalContentHash(proposal),
|
|
674
677
|
}, ctx.opts.proposalsCtx);
|
|
@@ -260,6 +260,39 @@ function loadPendingConsolidateProposalHashes(stashDir, proposalsCtx) {
|
|
|
260
260
|
}
|
|
261
261
|
return hashes;
|
|
262
262
|
}
|
|
263
|
+
/**
|
|
264
|
+
* Rejections decided before this are not a verdict on the memory's text: the
|
|
265
|
+
* 2026-08-02 and 2026-08-18 bulk audits rejected hundreds of promotions
|
|
266
|
+
* wholesale, and holding their bodies would skip good memories for good.
|
|
267
|
+
*/
|
|
268
|
+
const REJECTED_BODY_HOLD_FROM = "2026-09-29";
|
|
269
|
+
/**
|
|
270
|
+
* Body hashes of the memories whose promotion was rejected on review. A
|
|
271
|
+
* proposal minted with `promotionSourceHash` names the memory's raw body; an
|
|
272
|
+
* older one is hashed from its own body, which is the memory's unless
|
|
273
|
+
* sanitization changed it.
|
|
274
|
+
*/
|
|
275
|
+
function loadRejectedPromotionBodyHashes(stashDir, proposalsCtx) {
|
|
276
|
+
const hashes = new Set();
|
|
277
|
+
try {
|
|
278
|
+
for (const p of listProposalsReadOnly(stashDir, { status: "rejected", includeArchive: true }, proposalsCtx)) {
|
|
279
|
+
if (p.source !== "consolidate")
|
|
280
|
+
continue;
|
|
281
|
+
if ((p.review?.decidedAt ?? p.updatedAt) < REJECTED_BODY_HOLD_FROM)
|
|
282
|
+
continue;
|
|
283
|
+
try {
|
|
284
|
+
hashes.add(p.promotionSourceHash ?? contentHash(proposalContent(p), "body"));
|
|
285
|
+
}
|
|
286
|
+
catch {
|
|
287
|
+
// A malformed payload cannot hold a memory.
|
|
288
|
+
}
|
|
289
|
+
}
|
|
290
|
+
}
|
|
291
|
+
catch {
|
|
292
|
+
// Best-effort: a failed read never blocks judging.
|
|
293
|
+
}
|
|
294
|
+
return hashes;
|
|
295
|
+
}
|
|
263
296
|
/**
|
|
264
297
|
* Body hashes of the live knowledge assets, read from disk (the index may lag
|
|
265
298
|
* a just-written asset), so an accepted promotion is not proposed again.
|
|
@@ -469,6 +502,18 @@ export function inspectConsolidationPool(opts, stashDir, warnings, existingKnowl
|
|
|
469
502
|
return !isLedgerBlocked(row, nowIso, changedAt);
|
|
470
503
|
});
|
|
471
504
|
}
|
|
505
|
+
// A memory whose text a reviewer already rejected as a promotion waits for an edit, whatever it is called.
|
|
506
|
+
const rejectedBodies = loadRejectedPromotionBodyHashes(stashDir, opts.proposalsCtx);
|
|
507
|
+
if (rejectedBodies.size > 0) {
|
|
508
|
+
memories = memories.filter((memory) => {
|
|
509
|
+
try {
|
|
510
|
+
return !rejectedBodies.has(contentHash(fs.readFileSync(memory.filePath, "utf8"), "body"));
|
|
511
|
+
}
|
|
512
|
+
catch {
|
|
513
|
+
return true;
|
|
514
|
+
}
|
|
515
|
+
});
|
|
516
|
+
}
|
|
472
517
|
const judgedUnchanged = poolSize - memories.length;
|
|
473
518
|
// Only what retrieval returned or new material improve never processed (#986).
|
|
474
519
|
const retrievalScope = loadRetrievalScope({ proposalsCtx: opts.proposalsCtx, readOnly }, stashDir);
|
|
@@ -768,7 +813,13 @@ async function consolidate(opts, config, stashDir, startMs, stateDb) {
|
|
|
768
813
|
recordLedgerAttempt({ proposalsCtx: opts.proposalsCtx }, [...acc.judgedRefs]
|
|
769
814
|
.filter((ref) => !ctx.promotedSourceRefs.has(ref) &&
|
|
770
815
|
!acc.skipReasonByRef.get(ref)?.skips.some((skip) => skip.reason === "promote_create_failed"))
|
|
771
|
-
.map((ref) => ({
|
|
816
|
+
.map((ref) => ({
|
|
817
|
+
stashDir,
|
|
818
|
+
ref,
|
|
819
|
+
source: "consolidate",
|
|
820
|
+
outcome: "judged_no_action",
|
|
821
|
+
...bodyHashOf(ctx.memoryByRef.get(ref)),
|
|
822
|
+
})));
|
|
772
823
|
return makeConsolidateResult({
|
|
773
824
|
...summary(),
|
|
774
825
|
promoted: ctx.promoted,
|
|
@@ -783,6 +834,17 @@ async function consolidate(opts, config, stashDir, startMs, stateDb) {
|
|
|
783
834
|
},
|
|
784
835
|
});
|
|
785
836
|
}
|
|
837
|
+
/** The memory's current body hash as a ledger input field; empty when it cannot be read (the row then keeps its 7-day window). */
|
|
838
|
+
function bodyHashOf(memory) {
|
|
839
|
+
if (!memory)
|
|
840
|
+
return {};
|
|
841
|
+
try {
|
|
842
|
+
return { contentHash: contentHash(fs.readFileSync(memory.filePath, "utf8"), "body") };
|
|
843
|
+
}
|
|
844
|
+
catch {
|
|
845
|
+
return {};
|
|
846
|
+
}
|
|
847
|
+
}
|
|
786
848
|
/** The conceptId a ref maps to, or undefined for an invalid ref. */
|
|
787
849
|
function conceptIdForRef(ref) {
|
|
788
850
|
try {
|
|
@@ -2,25 +2,24 @@
|
|
|
2
2
|
// License, v. 2.0. If a copy of the MPL was not distributed with this
|
|
3
3
|
// file, You can obtain one at https://mozilla.org/MPL/2.0/.
|
|
4
4
|
/**
|
|
5
|
-
* Distill guards: related lessons
|
|
6
|
-
* overwrite
|
|
7
|
-
* proposal does not contradict the memories it came from.
|
|
5
|
+
* Distill guards: the related lessons, knowledge notes and skills shown to the
|
|
6
|
+
* writer so it does not repeat or overwrite them (CLS context), and a cheap check
|
|
7
|
+
* that a proposal does not contradict the memories it came from.
|
|
8
8
|
*/
|
|
9
9
|
export const DEFAULT_CLS_ADJACENT_COUNT = 3;
|
|
10
|
-
/** The CLS prompt section (each entry capped at
|
|
10
|
+
/** The CLS prompt section (each entry capped at 600 chars); empty when disabled (on unless `enabled: false`) or nothing is related. */
|
|
11
11
|
export function buildClsContext(adjacentItems, config) {
|
|
12
|
-
if (
|
|
12
|
+
if (config.enabled === false || adjacentItems.length === 0)
|
|
13
13
|
return "";
|
|
14
14
|
const lines = [
|
|
15
15
|
"",
|
|
16
|
-
"##
|
|
17
|
-
"The
|
|
18
|
-
"
|
|
19
|
-
"disagree with one, flag it as contradicted (do not ignore it).",
|
|
16
|
+
"## Related assets already in the library",
|
|
17
|
+
"The library already holds these lessons, knowledge notes and skills near this memory. They may be about another subject.",
|
|
18
|
+
"If one of them already states the rule the memory would give, answer NONE. Do not contradict or overwrite them.",
|
|
20
19
|
"",
|
|
21
20
|
];
|
|
22
21
|
for (const item of adjacentItems)
|
|
23
|
-
lines.push(`### ${item.ref}`, item.content.trim().slice(0,
|
|
22
|
+
lines.push(`### ${item.ref}`, item.content.trim().slice(0, 600), "");
|
|
24
23
|
return lines.join("\n");
|
|
25
24
|
}
|
|
26
25
|
/**
|
|
@@ -78,23 +78,29 @@ export function deriveLessonRef(inputRef) {
|
|
|
78
78
|
// properties, so every property is required and "none" is an empty array (#1046).
|
|
79
79
|
export const DISTILL_LESSON_JSON_SCHEMA = {
|
|
80
80
|
type: "object",
|
|
81
|
-
required: ["description", "when_to_use", "body", "tags"],
|
|
81
|
+
required: ["reason", "decision", "description", "when_to_use", "body", "tags"],
|
|
82
82
|
additionalProperties: false,
|
|
83
83
|
properties: {
|
|
84
|
+
reason: {
|
|
85
|
+
type: "string",
|
|
86
|
+
description: "One sentence, written first: the cause and the fix the memory states, or why it states none.",
|
|
87
|
+
},
|
|
88
|
+
decision: {
|
|
89
|
+
type: "string",
|
|
90
|
+
enum: ["lesson", "none"],
|
|
91
|
+
description: "`none` when the memory holds no lesson (it records what was done, a design, or steps already written elsewhere): leave the other fields empty. Otherwise `lesson`.",
|
|
92
|
+
},
|
|
84
93
|
description: {
|
|
85
94
|
type: "string",
|
|
86
|
-
|
|
87
|
-
description: "Single complete sentence (80-200 chars) summarising what the lesson teaches. No markdown, no leading 'When'/'If'.",
|
|
95
|
+
description: "Single complete sentence summarising what the lesson teaches. No markdown, no leading 'When'/'If'. Empty for `none`.",
|
|
88
96
|
},
|
|
89
97
|
when_to_use: {
|
|
90
98
|
type: "string",
|
|
91
|
-
|
|
92
|
-
description: "Single complete sentence describing the concrete trigger condition for the lesson.",
|
|
99
|
+
description: "Single complete sentence describing the concrete trigger condition for the lesson. Empty for `none`.",
|
|
93
100
|
},
|
|
94
101
|
body: {
|
|
95
102
|
type: "string",
|
|
96
|
-
|
|
97
|
-
description: "Lesson body — plain markdown, 1-3 short paragraphs of practical guidance.",
|
|
103
|
+
description: "Lesson body: plain markdown, shorter than the memory, stating only what it and its feedback say. Empty for `none`.",
|
|
98
104
|
},
|
|
99
105
|
tags: {
|
|
100
106
|
type: "array",
|
|
@@ -154,6 +160,19 @@ export function assembleStructuredDistillMarkdown(payload, kind) {
|
|
|
154
160
|
fm.xrefs = sources;
|
|
155
161
|
return assembleAssetFromString(serializeFrontmatterQuoted(fm), body);
|
|
156
162
|
}
|
|
163
|
+
/**
|
|
164
|
+
* The writer's answer when it found no lesson: the word NONE, or `decision: "none"` in a reply bound to the schema,
|
|
165
|
+
* with the reason it gave (`""` for the bare word). `null` for any other reply.
|
|
166
|
+
*/
|
|
167
|
+
function answeredNone(raw) {
|
|
168
|
+
if (/^[\s"'`*_]*none[\s.!"'`*_]*$/i.test(stripMarkdownFences(raw)))
|
|
169
|
+
return { reason: "" };
|
|
170
|
+
const payload = parseEmbeddedJsonResponse(raw);
|
|
171
|
+
if (payload === null || typeof payload !== "object" || Array.isArray(payload) || payload.decision !== "none") {
|
|
172
|
+
return null;
|
|
173
|
+
}
|
|
174
|
+
return { reason: typeof payload.reason === "string" ? payload.reason.trim() : "" };
|
|
175
|
+
}
|
|
157
176
|
function validateKnowledgeContent(content, inputRef) {
|
|
158
177
|
const findings = [];
|
|
159
178
|
const parsed = parseFrontmatter(content);
|
|
@@ -251,7 +270,7 @@ export function buildDistillPrompt(input) {
|
|
|
251
270
|
}
|
|
252
271
|
lines.push(input.proposalKind === "knowledge"
|
|
253
272
|
? "Produce the knowledge markdown file now. Start your response with `---` on the first line, followed by a `description:` field whose value is a 1-sentence summary (20–400 chars). Never use placeholder values like `---`, `tbd`, `n/a`, or a single dash. If the source has nothing meaningful to summarize, do NOT produce a proposal — return an empty response instead. The frontmatter block ends with a second `---` line; do not emit any additional `---` fences in the body."
|
|
254
|
-
: "Produce the lesson markdown file now. Start your response with `---` on the first line, followed by `description:` and `when_to_use:` fields. Both must be real one-sentence summaries (20–400 chars) — never placeholder values like `---`, `tbd`, or `n/a`. The frontmatter block ends with a second `---` line; do not emit any additional `---` fences in the body.");
|
|
273
|
+
: "Produce the lesson markdown file now. Start your response with `---` on the first line, followed by `description:` and `when_to_use:` fields. Both must be real one-sentence summaries (20–400 chars) — never placeholder values like `---`, `tbd`, or `n/a`. The frontmatter block ends with a second `---` line; do not emit any additional `---` fences in the body. If the memory holds no lesson, answer NONE instead.");
|
|
255
274
|
return lines.join("\n");
|
|
256
275
|
}
|
|
257
276
|
// ── Invocation ───────────────────────────────────────────────────────────────
|
|
@@ -343,7 +362,7 @@ export async function akmDistill(options) {
|
|
|
343
362
|
asset,
|
|
344
363
|
vocabulary: loadRefVocabulary(),
|
|
345
364
|
outcomeWeightEnabled: config.improve?.salience?.outcomeWeightEnabled !== false,
|
|
346
|
-
|
|
365
|
+
related: options.fetchRelatedFn ?? fetchRelatedAssets,
|
|
347
366
|
lookup,
|
|
348
367
|
};
|
|
349
368
|
const feedbackEvents = readDistillFeedback(run);
|
|
@@ -381,6 +400,8 @@ async function distill(run, targetKind, kind, outputRef, feedbackEvents) {
|
|
|
381
400
|
system,
|
|
382
401
|
prompt,
|
|
383
402
|
gate: { config: run.config, enabled: true },
|
|
403
|
+
// NONE is an answer: a parser that rejected it would ask for a lesson again.
|
|
404
|
+
parse: (raw) => (answeredNone(raw) ? raw : parseEmbeddedJsonResponse(raw)),
|
|
384
405
|
// The injected test transport never sees the schema.
|
|
385
406
|
request: {
|
|
386
407
|
...(run.options.chat === undefined
|
|
@@ -413,6 +434,10 @@ async function distill(run, targetKind, kind, outputRef, feedbackEvents) {
|
|
|
413
434
|
...exclusionMeta(run, true),
|
|
414
435
|
};
|
|
415
436
|
}
|
|
437
|
+
const none = answeredNone(call.raw);
|
|
438
|
+
if (none) {
|
|
439
|
+
return skipDistill(run, outputRef, kind, "nothing_reusable", `The writer found no lesson in ${run.inputRef}${none.reason ? `: ${none.reason}` : "."}`);
|
|
440
|
+
}
|
|
416
441
|
const assembled = assembleDistilledContent(run, call.raw, kind, outputRef);
|
|
417
442
|
if ("rejection" in assembled)
|
|
418
443
|
return assembled.rejection;
|
|
@@ -422,8 +447,20 @@ async function distill(run, targetKind, kind, outputRef, feedbackEvents) {
|
|
|
422
447
|
content: assembled.content,
|
|
423
448
|
source: run.asset.content,
|
|
424
449
|
descriptionSwapped: assembled.descriptionSwapped,
|
|
450
|
+
feedback: feedbackLines(feedback),
|
|
425
451
|
});
|
|
426
452
|
}
|
|
453
|
+
/** The feedback that says something, one line each, for the judge. A bare signal says nothing the writer could use. */
|
|
454
|
+
function feedbackLines(feedback) {
|
|
455
|
+
const lines = [];
|
|
456
|
+
for (const event of feedback) {
|
|
457
|
+
const meta = event.metadata ?? {};
|
|
458
|
+
const detail = (typeof meta.reason === "string" ? meta.reason : "") || (typeof meta.note === "string" ? meta.note : "");
|
|
459
|
+
if (detail.trim())
|
|
460
|
+
lines.push(`- [${typeof meta.signal === "string" ? meta.signal : event.eventType}] ${detail.trim()}`);
|
|
461
|
+
}
|
|
462
|
+
return lines;
|
|
463
|
+
}
|
|
427
464
|
/** Whether a file already holds the lesson `ref` in the stash the proposal would be filed in. */
|
|
428
465
|
function lessonExists(run, ref) {
|
|
429
466
|
const { type, name } = parseRefInput(ref);
|
|
@@ -489,11 +526,12 @@ async function judgeAndQueue(run, out) {
|
|
|
489
526
|
let confidence;
|
|
490
527
|
let judged;
|
|
491
528
|
if (qualityGateEnabled(run)) {
|
|
492
|
-
const
|
|
529
|
+
const related = await run.related(content.slice(0, 500), RELATED_COUNT);
|
|
493
530
|
// The judge reads what the generator read: the source body, without its frontmatter (buildDistillPrompt).
|
|
494
531
|
const source = out.source ? parseFrontmatter(out.source).content.trim() : "";
|
|
495
532
|
const verdict = await runLessonQualityJudge(run.config, content, source, run.options.chat, {
|
|
496
|
-
...(
|
|
533
|
+
...(related.length > 0 ? { related } : {}),
|
|
534
|
+
...(out.feedback && out.feedback.length > 0 ? { feedback: out.feedback } : {}),
|
|
497
535
|
...((run.judgeRunner ?? run.runner) ? { llmRunner: run.judgeRunner ?? run.runner } : {}),
|
|
498
536
|
...(run.options.signal ? { signal: run.options.signal } : {}),
|
|
499
537
|
onNotices: run.notices.add,
|
|
@@ -873,13 +911,13 @@ function readDistillFeedback(run) {
|
|
|
873
911
|
/** System + user prompt: rejected-proposal context, optional CLS neighbours, stash standards. */
|
|
874
912
|
async function buildDistillMessages(run, feedback, kind, outputRef) {
|
|
875
913
|
const rejectedProposals = rejectedProposalContext(run.stash, run.inputRef, run.options.ctx, run.options.eventsCtx);
|
|
876
|
-
// CLS interleaving (default
|
|
914
|
+
// CLS interleaving (default on): show the related lessons, knowledge notes and skills, so the writer neither repeats nor overwrites them.
|
|
877
915
|
const cls = getImproveProcessConfig("distill", run.profile)?.cls ?? {};
|
|
878
916
|
let clsContext = "";
|
|
879
|
-
if (cls.enabled) {
|
|
917
|
+
if (cls.enabled !== false) {
|
|
880
918
|
try {
|
|
881
919
|
const query = run.asset.content ? run.asset.content.slice(0, 500) : run.inputRef;
|
|
882
|
-
clsContext = buildClsContext(await run.
|
|
920
|
+
clsContext = buildClsContext(await run.related(query, cls.adjacentCount ?? DEFAULT_CLS_ADJACENT_COUNT), cls);
|
|
883
921
|
}
|
|
884
922
|
catch {
|
|
885
923
|
// CLS context is supplemental.
|
|
@@ -908,12 +946,18 @@ async function defaultLookup(ref, stashDir) {
|
|
|
908
946
|
honorOrigin: false,
|
|
909
947
|
});
|
|
910
948
|
}
|
|
911
|
-
/**
|
|
912
|
-
|
|
949
|
+
/** What the library already holds on a subject: lessons and knowledge notes say it, a skill is how to do it. */
|
|
950
|
+
const RELATED_TYPES = ["lesson", "knowledge", "skill"];
|
|
951
|
+
const RELATED_COUNT = 3;
|
|
952
|
+
/** The top-N lessons, knowledge notes and skills related to `query`, best first (empty when search is unavailable). */
|
|
953
|
+
async function fetchRelatedAssets(query, n) {
|
|
913
954
|
try {
|
|
914
|
-
|
|
915
|
-
|
|
955
|
+
// One search per type: memories outnumber the rest and would fill an untyped list.
|
|
956
|
+
const results = await Promise.all(RELATED_TYPES.map((type) => akmSearch({ query, type, limit: n, skipLogging: true, eventSource: "improve" })));
|
|
957
|
+
return results
|
|
958
|
+
.flatMap((result) => result?.hits ?? [])
|
|
916
959
|
.filter((h) => "path" in h && typeof h.path === "string")
|
|
960
|
+
.sort((a, b) => (b.score ?? 0) - (a.score ?? 0))
|
|
917
961
|
.slice(0, n)
|
|
918
962
|
.map((h) => {
|
|
919
963
|
let content = "";
|
|
@@ -601,6 +601,16 @@ export function buildSnapshotManifest(args) {
|
|
|
601
601
|
const latestNegativeTs = new Map();
|
|
602
602
|
const feedback = new Map(candidates.map((r) => [r.ref, { hasSignal: false, positive: 0, negative: 0 }]));
|
|
603
603
|
if (candidates.length > 0) {
|
|
604
|
+
// When each ref's accepted feedback proposals were created: a fix event is acted on once one exists at or after it.
|
|
605
|
+
const fixedAt = new Map();
|
|
606
|
+
if (stashDir) {
|
|
607
|
+
withRunState(eventsCtx, args.readOnly !== true, (db) => {
|
|
608
|
+
for (const p of listStateProposals(db, { stashDir, status: "accepted" })) {
|
|
609
|
+
if (p.source === "feedback")
|
|
610
|
+
fixedAt.set(p.ref, [...(fixedAt.get(p.ref) ?? []), p.createdAt]);
|
|
611
|
+
}
|
|
612
|
+
});
|
|
613
|
+
}
|
|
604
614
|
for (const e of readEvents({ type: "feedback" }, eventsCtx).events) {
|
|
605
615
|
const ref = e.ref ? refByKey.get(e.ref) : undefined;
|
|
606
616
|
const entry = ref ? feedback.get(ref) : undefined;
|
|
@@ -612,8 +622,11 @@ export function buildSnapshotManifest(args) {
|
|
|
612
622
|
entry.hasSignal = true;
|
|
613
623
|
if (ts > (latestFeedbackTs.get(ref) ?? ""))
|
|
614
624
|
latestFeedbackTs.set(ref, ts);
|
|
615
|
-
|
|
625
|
+
const fixApplied = e.metadata?.fix !== undefined &&
|
|
626
|
+
(fixedAt.get(e.ref ?? "") ?? []).some((createdAt) => createdAt >= ts);
|
|
627
|
+
if (signal === "negative" && !fixApplied && ts > (latestNegativeTs.get(ref) ?? "")) {
|
|
616
628
|
latestNegativeTs.set(ref, ts);
|
|
629
|
+
}
|
|
617
630
|
}
|
|
618
631
|
if (signal === "positive")
|
|
619
632
|
entry.positive++;
|
|
@@ -216,28 +216,31 @@ export function resolveQualityGateJudge(config, profile, processName, onNotices)
|
|
|
216
216
|
onNotices?.(resolved.notices);
|
|
217
217
|
return resolved.runner;
|
|
218
218
|
}
|
|
219
|
-
/** Lesson judge prompt
|
|
220
|
-
export function buildJudgePrompt(lessonContent, sourceContent,
|
|
219
|
+
/** Lesson judge prompt: what the writer was given (the source and its feedback), the assets nearest the new lesson, the lesson. */
|
|
220
|
+
export function buildJudgePrompt(lessonContent, sourceContent, related, feedback) {
|
|
221
221
|
const lines = [
|
|
222
|
-
"You are evaluating a
|
|
222
|
+
"You are evaluating a lesson an agent wrote from a memory and the feedback about it, for an akm knowledge base.",
|
|
223
223
|
"",
|
|
224
224
|
"Score this lesson on each criterion from 1 (poor) to 5 (excellent):",
|
|
225
|
-
"1.
|
|
226
|
-
"2. NON-REDUNDANCY: Is
|
|
227
|
-
"3. GROUNDING: Is the lesson
|
|
225
|
+
"1. REUSABLE: Does the lesson state a rule an agent can use on another occasion, with the reason it holds? Score 1-2 when it only records what was done, shipped, decided, found or is pending, on a date or for one build, machine or project, or how a system is set up now, however it is phrased. Score 4-5 for a rule with its reason.",
|
|
226
|
+
"2. NON-REDUNDANCY: Is the lesson new next to the existing assets shown below? Score 1-2 only when one of them already states the same rule. Assets on other subjects change nothing: score 4-5 when none is shown or none is on the same subject.",
|
|
227
|
+
"3. GROUNDING: Is every statement in the lesson stated by the source or its feedback, in any words? Check each cause, step, number, rule and limit in the lesson against them. Score 4-5 when each is stated. Score 3 when one stretches what the source says. Score 1-2 when any is in neither, when the lesson drops a limit the source states (one place checked, not confirmed, a guess) and says more than it, or when it is about another subject than the source.",
|
|
228
228
|
"",
|
|
229
|
-
"Source
|
|
229
|
+
"Source memory:",
|
|
230
230
|
"```",
|
|
231
231
|
// The window distill generates from (buildDistillPrompt): grounding can reject, so the judge reads all of it.
|
|
232
232
|
sourceContent.slice(0, 3000),
|
|
233
233
|
"```",
|
|
234
234
|
];
|
|
235
|
-
if (
|
|
236
|
-
lines.push("", "
|
|
237
|
-
|
|
238
|
-
|
|
235
|
+
if (feedback && feedback.length > 0) {
|
|
236
|
+
lines.push("", "Feedback recorded about the memory (the writer saw it too):", "```", feedback.join("\n").slice(0, 1500), "```");
|
|
237
|
+
}
|
|
238
|
+
if (related && related.length > 0) {
|
|
239
|
+
lines.push("", "Existing assets nearest the new lesson (they may be on another subject):");
|
|
240
|
+
for (const asset of related)
|
|
241
|
+
lines.push(`\nExisting asset ref: ${asset.ref}`, "```", asset.content.slice(0, 600), "```");
|
|
239
242
|
}
|
|
240
|
-
lines.push("", "Proposed lesson
|
|
243
|
+
lines.push("", "Proposed lesson:", "```", lessonContent.slice(0, 2000), "```", "", 'Return ONLY valid JSON, no prose: {"scores": {"reusable": <1-5 integer>, "nonRedundancy": <1-5 integer>, "grounding": <1-5 integer>}, "reason": "<one sentence naming the weakest criterion>"}');
|
|
241
244
|
return lines.join("\n");
|
|
242
245
|
}
|
|
243
246
|
function boundedDocument(content, maxChars = 6000) {
|
|
@@ -317,23 +320,19 @@ export function buildReflectJudgePrompt(candidateContent, sourceContent, feedbac
|
|
|
317
320
|
}
|
|
318
321
|
/**
|
|
319
322
|
* `grounding` is scored with the other lesson criteria but left out of their
|
|
320
|
-
* mean
|
|
321
|
-
*
|
|
322
|
-
* a
|
|
323
|
-
* of
|
|
324
|
-
*
|
|
325
|
-
*
|
|
326
|
-
*
|
|
327
|
-
*
|
|
328
|
-
*
|
|
329
|
-
* the lesson, and the judge is never shown it. A contradiction of the source is
|
|
330
|
-
* the optional fidelity check's to send to a human (`judgeAndQueue` in
|
|
331
|
-
* distill.ts), so the rubric must not pre-empt it.
|
|
323
|
+
* mean. A lesson criterion scored {@link LESSON_REJECT_MAX_SCORE} or less is a
|
|
324
|
+
* rejection whatever the mean says: the mean would hide it (4 and 1 average
|
|
325
|
+
* 2.5, a review), and a reviewer was reading every lesson that was not rejected,
|
|
326
|
+
* 17 of 19 of them bad on 2026-10-05. The rubric reserves 1-2 for a lesson that
|
|
327
|
+
* records what was done instead of a rule, repeats an asset the library holds,
|
|
328
|
+
* or states what neither its source nor its feedback does. The judge is shown
|
|
329
|
+
* the feedback the writer saw, so a statement it supports is not an invention. A
|
|
330
|
+
* contradiction of the source is the optional fidelity check's to send to a
|
|
331
|
+
* human (`judgeAndQueue` in distill.ts).
|
|
332
332
|
*/
|
|
333
333
|
const GROUNDING_CRITERION = "grounding";
|
|
334
|
-
const
|
|
335
|
-
const
|
|
336
|
-
const LESSON_JUDGE_CRITERIA = ["novelty", "nonRedundancy", GROUNDING_CRITERION];
|
|
334
|
+
const LESSON_REJECT_MAX_SCORE = 2;
|
|
335
|
+
const LESSON_JUDGE_CRITERIA = ["reusable", "nonRedundancy", GROUNDING_CRITERION];
|
|
337
336
|
const REFLECT_JUDGE_CRITERIA = ["need", "preservation", "quality"];
|
|
338
337
|
/**
|
|
339
338
|
* Read a judge response: the per-criterion shape (averaged here, `grounding`
|
|
@@ -390,9 +389,8 @@ export function judgeResponseSchema(keys) {
|
|
|
390
389
|
* The quality judge. Fails closed: no runner, an unparseable verdict or a
|
|
391
390
|
* provider failure never passes content. Bands: every criterion in the mean
|
|
392
391
|
* >= 4 passes, otherwise a mean >= 2.5 is review and a lower one reject; a
|
|
393
|
-
* `grounding`
|
|
394
|
-
*
|
|
395
|
-
* that would pass to review (a mean that rejects stays a rejection).
|
|
392
|
+
* lesson criterion (`grounding` included) of {@link LESSON_REJECT_MAX_SCORE}
|
|
393
|
+
* or less rejects whatever the mean is.
|
|
396
394
|
* Temperature is set to 0, which reduces run-to-run variation but does not
|
|
397
395
|
* remove it: on some servers (llama.cpp batching, for one) the same request can
|
|
398
396
|
* score a point apart, so the routing rules are chosen with that margin in mind.
|
|
@@ -430,34 +428,18 @@ async function runQualityJudge(feature, config, prompt, keys, chat, options) {
|
|
|
430
428
|
if (!parsed)
|
|
431
429
|
return { pass: false, score: -1, reason: "judge parse failed — routed to review", reviewNeeded: true };
|
|
432
430
|
const { score, lowest, reason, criteria } = parsed;
|
|
433
|
-
|
|
434
|
-
if (criteria &&
|
|
435
|
-
|
|
436
|
-
|
|
437
|
-
score,
|
|
438
|
-
reason: `Off-subject for its source (grounding ${grounding}/5): ${reason}`,
|
|
439
|
-
criteria,
|
|
440
|
-
};
|
|
431
|
+
// A lesson criterion at 2 or below is a defect the mean would hide (4 and 1 average 2.5, a review): it rejects.
|
|
432
|
+
if (criteria && criteria[GROUNDING_CRITERION] !== undefined) {
|
|
433
|
+
const [weakest, low] = Object.entries(criteria).sort((x, y) => x[1] - y[1])[0];
|
|
434
|
+
if (low <= LESSON_REJECT_MAX_SCORE)
|
|
435
|
+
return { pass: false, score, reason: `${weakest} ${low}/5: ${reason}`, criteria };
|
|
441
436
|
}
|
|
442
437
|
const verdict = lowest >= 4 ? { pass: true } : score >= 2.5 ? { pass: false, reviewNeeded: true } : { pass: false };
|
|
443
|
-
// Borderline grounding is a person's call even when the lesson would pass; a mean that rejects stays rejected.
|
|
444
|
-
if (criteria &&
|
|
445
|
-
grounding !== undefined &&
|
|
446
|
-
grounding <= BORDERLINE_GROUNDING_MAX_SCORE &&
|
|
447
|
-
(verdict.pass || verdict.reviewNeeded)) {
|
|
448
|
-
return {
|
|
449
|
-
pass: false,
|
|
450
|
-
reviewNeeded: true,
|
|
451
|
-
score,
|
|
452
|
-
reason: `Borderline on grounding (${grounding}/5), routed to review: ${reason}`,
|
|
453
|
-
criteria,
|
|
454
|
-
};
|
|
455
|
-
}
|
|
456
438
|
return { ...verdict, score, reason, ...(criteria ? { criteria } : {}) };
|
|
457
439
|
}
|
|
458
440
|
/** Judge a proposed lesson (or knowledge promotion) against its source. */
|
|
459
441
|
export function runLessonQualityJudge(config, lessonContent, sourceContent, chat, options = {}) {
|
|
460
|
-
const prompt = buildJudgePrompt(lessonContent, sourceContent, options.
|
|
442
|
+
const prompt = buildJudgePrompt(lessonContent, sourceContent, options.related, options.feedback);
|
|
461
443
|
return runQualityJudge("lesson_quality_gate", config, prompt, LESSON_JUDGE_CRITERIA, chat, options);
|
|
462
444
|
}
|
|
463
445
|
/** Judge an in-place reflect revision without new-lesson novelty criteria. */
|
|
@@ -28,6 +28,7 @@ import { info, warn } from "../../core/warn.js";
|
|
|
28
28
|
import { DEFAULT_LLM_TIMEOUT_MS } from "../../integrations/agent/config.js";
|
|
29
29
|
import { buildExecution, resolveExecution } from "../../integrations/agent/execution.js";
|
|
30
30
|
import { assertRunnerCredentials, runExecution, } from "../../integrations/agent/runner-dispatch.js";
|
|
31
|
+
import { nearestKnowledgeNotes } from "../improve/consolidate/coverage.js";
|
|
31
32
|
import { errMessage, noticeSet } from "../improve/stage.js";
|
|
32
33
|
import { akmProposalAccept, akmProposalReject } from "./proposal.js";
|
|
33
34
|
import { isRetireProposal, PAIR_PASS_GATE, STALE_TARGET_GATE_REASON } from "./proposal-types.js";
|
|
@@ -57,8 +58,13 @@ function categorizeDrainFailure(message, fallback) {
|
|
|
57
58
|
* rejection, so instead of failing identically every run it is auto-rejected
|
|
58
59
|
* once; the ledger records `failed`, keeping the ref re-proposable.
|
|
59
60
|
*/
|
|
60
|
-
async function acceptProposal(opts, proposal, id, reason, promoteFn, rejectFn) {
|
|
61
|
-
const gateDecision = {
|
|
61
|
+
async function acceptProposal(opts, proposal, id, reason, promoteFn, rejectFn, judgeReason) {
|
|
62
|
+
const gateDecision = {
|
|
63
|
+
outcome: "auto-accepted",
|
|
64
|
+
reason,
|
|
65
|
+
gate: DRAIN_GATE,
|
|
66
|
+
...(judgeReason ? { judgeReason } : {}),
|
|
67
|
+
};
|
|
62
68
|
try {
|
|
63
69
|
if (!opts.dryRun) {
|
|
64
70
|
await promoteFn({
|
|
@@ -102,7 +108,7 @@ async function acceptProposal(opts, proposal, id, reason, promoteFn, rejectFn) {
|
|
|
102
108
|
}
|
|
103
109
|
}
|
|
104
110
|
/** Reject one proposal (nothing in a dry run); the error message on failure. */
|
|
105
|
-
async function rejectProposal(opts, id, reason, gateReason, rejectFn) {
|
|
111
|
+
async function rejectProposal(opts, id, reason, gateReason, rejectFn, judgeReason) {
|
|
106
112
|
if (opts.dryRun)
|
|
107
113
|
return undefined;
|
|
108
114
|
try {
|
|
@@ -110,7 +116,12 @@ async function rejectProposal(opts, id, reason, gateReason, rejectFn) {
|
|
|
110
116
|
stashDir: opts.stashDir,
|
|
111
117
|
id,
|
|
112
118
|
reason,
|
|
113
|
-
gateDecision: {
|
|
119
|
+
gateDecision: {
|
|
120
|
+
outcome: "auto-rejected",
|
|
121
|
+
reason: gateReason,
|
|
122
|
+
gate: DRAIN_GATE,
|
|
123
|
+
...(judgeReason ? { judgeReason } : {}),
|
|
124
|
+
},
|
|
114
125
|
});
|
|
115
126
|
return undefined;
|
|
116
127
|
}
|
|
@@ -118,7 +129,20 @@ async function rejectProposal(opts, id, reason, gateReason, rejectFn) {
|
|
|
118
129
|
return errMessage(err);
|
|
119
130
|
}
|
|
120
131
|
}
|
|
121
|
-
/**
|
|
132
|
+
/**
|
|
133
|
+
* `text` as a fenced block whose fence is longer than any backtick run inside
|
|
134
|
+
* it (the CommonMark rule), so a note holding its own code block is not cut
|
|
135
|
+
* short at the first inner fence.
|
|
136
|
+
*/
|
|
137
|
+
export function fencedBlock(text) {
|
|
138
|
+
let longest = 2;
|
|
139
|
+
for (const run of text.match(/`+/g) ?? [])
|
|
140
|
+
if (run.length > longest)
|
|
141
|
+
longest = run.length;
|
|
142
|
+
const fence = "`".repeat(longest + 1);
|
|
143
|
+
return [fence, text, fence];
|
|
144
|
+
}
|
|
145
|
+
/** The judgment prompt: the proposal, the live asset it would overwrite, same-ref siblings, and for a promotion the nearest knowledge notes. */
|
|
122
146
|
export function buildJudgmentPrompt(proposal, reason, ctx) {
|
|
123
147
|
const sections = [
|
|
124
148
|
"You are adjudicating a pending knowledge-base proposal no quality judge has",
|
|
@@ -129,12 +153,10 @@ export function buildJudgmentPrompt(proposal, reason, ctx) {
|
|
|
129
153
|
`Left for judgment because: ${reason === "needs-judgment" ? "no quality judge has passed this content yet" : reason}`,
|
|
130
154
|
"",
|
|
131
155
|
"## Proposed content",
|
|
132
|
-
|
|
133
|
-
proposalContent(proposal),
|
|
134
|
-
"```",
|
|
156
|
+
...fencedBlock(proposalContent(proposal)),
|
|
135
157
|
];
|
|
136
158
|
if (ctx.liveAsset !== undefined) {
|
|
137
|
-
sections.push("", "## Current live asset (would be overwritten on accept)",
|
|
159
|
+
sections.push("", "## Current live asset (would be overwritten on accept)", ...fencedBlock(ctx.liveAsset));
|
|
138
160
|
}
|
|
139
161
|
else {
|
|
140
162
|
sections.push("", "## Current live asset", "(none — this proposal would create a new asset)");
|
|
@@ -142,8 +164,15 @@ export function buildJudgmentPrompt(proposal, reason, ctx) {
|
|
|
142
164
|
if (ctx.siblings.length > 0) {
|
|
143
165
|
sections.push("", "## Other pending proposals for the same ref (dedup context)");
|
|
144
166
|
for (const sib of ctx.siblings) {
|
|
145
|
-
sections.push("", `### Sibling ${sib.id} (source: ${sib.source})`,
|
|
167
|
+
sections.push("", `### Sibling ${sib.id} (source: ${sib.source})`, ...fencedBlock(proposalContent(sib)));
|
|
168
|
+
}
|
|
169
|
+
}
|
|
170
|
+
if (ctx.neighbours && ctx.neighbours.length > 0) {
|
|
171
|
+
sections.push("", "## Existing knowledge notes nearest to this promotion's source memory");
|
|
172
|
+
for (const note of ctx.neighbours) {
|
|
173
|
+
sections.push("", `### ${note.ref}`, note.description, ...fencedBlock(note.excerpt));
|
|
146
174
|
}
|
|
175
|
+
sections.push("", "Reject the promotion if these notes already cover what it says, even in other words.");
|
|
147
176
|
}
|
|
148
177
|
sections.push("", "## Your task", 'Return ONLY a JSON object: {"decision": "accept" | "reject" | "defer", "reason": "<short reason>"}.', "- accept: the proposed content is a correct, valuable update worth committing.", "- reject: the proposal is wrong, a duplicate, or contradicts the live asset.", "- defer: you cannot decide from the provided context (leave it pending).", "Output the JSON object and nothing else.");
|
|
149
178
|
return sections.join("\n");
|
|
@@ -205,7 +234,7 @@ async function dispatchJudgment(runner, prompt, seams) {
|
|
|
205
234
|
* (under `applyMode` and the remaining accept budget) or the reject. A defer, an
|
|
206
235
|
* unparseable verdict or a runner error leaves the item undecided.
|
|
207
236
|
*/
|
|
208
|
-
async function runJudgmentTier(opts, result, pending, acceptBudget, promoteFn, rejectFn, seams) {
|
|
237
|
+
async function runJudgmentTier(opts, result, pending, acceptBudget, promoteFn, rejectFn, seams, deferNotes) {
|
|
209
238
|
const byId = new Map(pending.map((p) => [p.id, p]));
|
|
210
239
|
const notices = noticeSet();
|
|
211
240
|
const stillDeferred = [];
|
|
@@ -216,21 +245,33 @@ async function runJudgmentTier(opts, result, pending, acceptBudget, promoteFn, r
|
|
|
216
245
|
stillDeferred.push(item);
|
|
217
246
|
continue;
|
|
218
247
|
}
|
|
248
|
+
const liveAsset = readLiveAssetContent(opts.stashDir, proposal.ref);
|
|
219
249
|
const prompt = buildJudgmentPrompt(proposal, item.reason, {
|
|
220
|
-
liveAsset
|
|
250
|
+
liveAsset,
|
|
221
251
|
siblings: pending.filter((p) => p.ref === proposal.ref && p.id !== proposal.id),
|
|
252
|
+
// A create the model otherwise judges blind: the model never sees knowledge/.
|
|
253
|
+
...(liveAsset === undefined ? { neighbours: promotionNeighbours(opts.stashDir, proposal) } : {}),
|
|
222
254
|
});
|
|
223
255
|
const dispatch = await dispatchJudgment(opts.judgment, prompt, seams);
|
|
224
256
|
notices.add(dispatch.notices);
|
|
225
257
|
if (dispatch.error)
|
|
226
258
|
warn(`[triage] judgment dispatch failed for ${item.id}: ${dispatch.error}`);
|
|
227
259
|
const verdict = dispatch.error ? null : dispatch.verdict;
|
|
228
|
-
if (!verdict
|
|
260
|
+
if (!verdict) {
|
|
261
|
+
deferNotes.set(item.id, { reason: dispatch.error ? "judgment-error" : "judgment-parse-failure" });
|
|
262
|
+
stillDeferred.push(item);
|
|
263
|
+
continue;
|
|
264
|
+
}
|
|
265
|
+
if (verdict.decision === "defer") {
|
|
266
|
+
deferNotes.set(item.id, {
|
|
267
|
+
reason: "judgment-deferred",
|
|
268
|
+
...(verdict.reason ? { judgeReason: verdict.reason } : {}),
|
|
269
|
+
});
|
|
229
270
|
stillDeferred.push(item);
|
|
230
271
|
continue;
|
|
231
272
|
}
|
|
232
273
|
if (verdict.decision === "reject") {
|
|
233
|
-
const failure = await rejectProposal(opts, item.id, verdict.reason || "judgment: reject", "judgment-reject", rejectFn);
|
|
274
|
+
const failure = await rejectProposal(opts, item.id, verdict.reason || "judgment: reject", "judgment-reject", rejectFn, verdict.reason);
|
|
234
275
|
if (failure === undefined) {
|
|
235
276
|
result.rejected.push(item.id);
|
|
236
277
|
}
|
|
@@ -252,6 +293,7 @@ async function runJudgmentTier(opts, result, pending, acceptBudget, promoteFn, r
|
|
|
252
293
|
reason: "judgment-accept",
|
|
253
294
|
contentHash: proposalContentHash(proposal),
|
|
254
295
|
gate: DRAIN_GATE,
|
|
296
|
+
...(verdict.reason ? { judgeReason: verdict.reason } : {}),
|
|
255
297
|
});
|
|
256
298
|
result.staged.push(item.id);
|
|
257
299
|
}
|
|
@@ -265,7 +307,7 @@ async function runJudgmentTier(opts, result, pending, acceptBudget, promoteFn, r
|
|
|
265
307
|
result.skippedByCap.push(item.id);
|
|
266
308
|
continue;
|
|
267
309
|
}
|
|
268
|
-
const outcome = await acceptProposal(opts, proposal, item.id, "judgment-accept", promoteFn, rejectFn);
|
|
310
|
+
const outcome = await acceptProposal(opts, proposal, item.id, "judgment-accept", promoteFn, rejectFn, verdict.reason);
|
|
269
311
|
if (outcome === "promoted") {
|
|
270
312
|
result.promoted.push(item.id);
|
|
271
313
|
acceptBudget -= 1;
|
|
@@ -286,6 +328,21 @@ async function runJudgmentTier(opts, result, pending, acceptBudget, promoteFn, r
|
|
|
286
328
|
result.notices = notices.list();
|
|
287
329
|
result.deferred = stillDeferred;
|
|
288
330
|
}
|
|
331
|
+
/** The knowledge notes nearest to a consolidate promotion's source memory; none for any other proposal. */
|
|
332
|
+
function promotionNeighbours(stashDir, proposal) {
|
|
333
|
+
if (proposal.source !== "consolidate" || proposal.promotionSource === undefined)
|
|
334
|
+
return [];
|
|
335
|
+
try {
|
|
336
|
+
const parsed = parseRefInput(proposal.promotionSource);
|
|
337
|
+
const typeDir = stashDirFor(parsed.type);
|
|
338
|
+
if (!typeDir)
|
|
339
|
+
return [];
|
|
340
|
+
return nearestKnowledgeNotes(assetPathForName(parsed.type, path.join(stashDir, typeDir), parsed.name));
|
|
341
|
+
}
|
|
342
|
+
catch {
|
|
343
|
+
return [];
|
|
344
|
+
}
|
|
345
|
+
}
|
|
289
346
|
/** The live asset a proposal would overwrite, if any. */
|
|
290
347
|
function readLiveAssetContent(stashDir, ref) {
|
|
291
348
|
try {
|
|
@@ -314,8 +371,8 @@ export async function drainProposals(opts, promoteFn = akmProposalAccept, reject
|
|
|
314
371
|
const empties = [];
|
|
315
372
|
for (const proposal of pending) {
|
|
316
373
|
// A consolidate pair-pass `retire` proposal is auto-accepted only when the
|
|
317
|
-
// pair
|
|
318
|
-
//
|
|
374
|
+
// pair pass staged it: nothing unique on either side, confirmed by a
|
|
375
|
+
// second look, no continuity risk; every other one waits for a direct
|
|
319
376
|
// `akm proposal accept` (spec §25.6). Checked before isEmptyDiff, which
|
|
320
377
|
// has nothing meaningful to read on a delete-primary change.
|
|
321
378
|
if (isRetireProposal(proposal)) {
|
|
@@ -324,7 +381,7 @@ export async function drainProposals(opts, promoteFn = akmProposalAccept, reject
|
|
|
324
381
|
staged.gate === PAIR_PASS_GATE &&
|
|
325
382
|
staged.contentHash === proposalContentHash(proposal) &&
|
|
326
383
|
!proposal.retirement?.continuityRisk) {
|
|
327
|
-
accepts.push({ id: proposal.id, reason:
|
|
384
|
+
accepts.push({ id: proposal.id, reason: staged.reason });
|
|
328
385
|
}
|
|
329
386
|
continue;
|
|
330
387
|
}
|
|
@@ -394,15 +451,22 @@ export async function drainProposals(opts, promoteFn = akmProposalAccept, reject
|
|
|
394
451
|
}
|
|
395
452
|
}
|
|
396
453
|
}
|
|
454
|
+
const deferNotes = new Map();
|
|
397
455
|
if (opts.judgment && result.deferred.length > 0) {
|
|
398
|
-
await runJudgmentTier({ ...opts, judgment: opts.judgment }, result, pending, cap - promotedHere, promoteFn, rejectFn, judgmentSeams);
|
|
456
|
+
await runJudgmentTier({ ...opts, judgment: opts.judgment }, result, pending, cap - promotedHere, promoteFn, rejectFn, judgmentSeams, deferNotes);
|
|
399
457
|
}
|
|
400
458
|
// #577: whatever stays undecided is left for review (`review_needed` in the ledger).
|
|
401
459
|
if (!opts.dryRun) {
|
|
402
|
-
const reviewReason = opts.judgment ? "judgment-deferred" : "no-judge-configured";
|
|
403
460
|
for (const item of result.deferred) {
|
|
461
|
+
const note = deferNotes.get(item.id);
|
|
462
|
+
const reviewReason = note?.reason ?? (opts.judgment ? "judgment-deferred" : "no-judge-configured");
|
|
404
463
|
try {
|
|
405
|
-
recordGateDecision(opts.stashDir, item.id, {
|
|
464
|
+
recordGateDecision(opts.stashDir, item.id, {
|
|
465
|
+
outcome: "deferred",
|
|
466
|
+
reason: reviewReason,
|
|
467
|
+
gate: DRAIN_GATE,
|
|
468
|
+
...(note?.judgeReason ? { judgeReason: note.judgeReason } : {}),
|
|
469
|
+
});
|
|
406
470
|
}
|
|
407
471
|
catch (err) {
|
|
408
472
|
warn(`[triage] failed to record gate decision for ${item.id}: ${errMessage(err)}`);
|
|
@@ -53,7 +53,7 @@ export function isRetireProposal(proposal) {
|
|
|
53
53
|
}
|
|
54
54
|
/** A promote refused because the target changed after mint (STALE, R20) — not a merit judgement. */
|
|
55
55
|
export const STALE_TARGET_GATE_REASON = "stale-target";
|
|
56
|
-
/** The gate on a retire proposal the triage drain may accept unattended: a pair-judged
|
|
56
|
+
/** The gate on a retire proposal the triage drain may accept unattended: a staged pair-judged retirement. */
|
|
57
57
|
export const PAIR_PASS_GATE = "consolidate-pair";
|
|
58
58
|
export const EXPIRED_GATE_REASON = "expired";
|
|
59
59
|
export const ASSET_MISSING_GATE_REASON = "asset-missing";
|
|
@@ -78,9 +78,9 @@ const qualityGateField = z
|
|
|
78
78
|
.optional();
|
|
79
79
|
/**
|
|
80
80
|
* WS-3b: CLS (Complementary Learning System) interleaving (step 9).
|
|
81
|
-
* distill
|
|
82
|
-
*
|
|
83
|
-
* Default
|
|
81
|
+
* The distill prompt includes the lessons, knowledge notes and skills the library already holds near the
|
|
82
|
+
* memory, so the writer answers NONE for a rule one of them states and does not overwrite a prior
|
|
83
|
+
* generalization. Default ON; `enabled: false` turns it off. Only meaningful on the `distill` process.
|
|
84
84
|
*/
|
|
85
85
|
const clsField = z
|
|
86
86
|
.object({
|
package/dist/core/paths.js
CHANGED
|
@@ -332,15 +332,6 @@ export function getStashStateKey(stashDir) {
|
|
|
332
332
|
function stashScopedDir(base, stashDir) {
|
|
333
333
|
return path.join(base, getStashStateKey(stashDir));
|
|
334
334
|
}
|
|
335
|
-
/**
|
|
336
|
-
* `$STATE/improve/measurement/verdicts/<stash>/` — `akm-eval-proactive-verdict`
|
|
337
|
-
* reports. Moved out of `$STASH/.akm/measurement/verdicts/` (itlackey/akm#890);
|
|
338
|
-
* the pilot treatment file at `$STASH/.akm/measurement/` is manually-authored
|
|
339
|
-
* measurement input and stays put.
|
|
340
|
-
*/
|
|
341
|
-
export function getMeasurementVerdictsDir(stashDir) {
|
|
342
|
-
return stashScopedDir(path.join(getStateDir(), "improve", "measurement", "verdicts"), stashDir);
|
|
343
|
-
}
|
|
344
335
|
/**
|
|
345
336
|
* `$CACHE/index/unresolved-sources/<stash>/` — synthetic placeholder path for
|
|
346
337
|
* a configured source whose content root did not resolve this run. Never
|
|
@@ -50,11 +50,14 @@ export const CONSOLIDATE_LEDGER_SOURCE = "consolidate";
|
|
|
50
50
|
* promotion is a verdict on that memory's text; asking the model about the
|
|
51
51
|
* same text again can only reproduce the proposal (accepted used to be
|
|
52
52
|
* eligible at once, rejected after 7 days), so the memory waits for an edit.
|
|
53
|
+
* So does a memory the model judged and left alone (`judged_no_action`):
|
|
54
|
+
* the same text would be judged weekly with the same answer.
|
|
53
55
|
* The clock stays for a row with no recorded hash — one decided before the
|
|
54
56
|
* hash was recorded — see {@link nextEligibleAt} and {@link isContentDrivenRow}.
|
|
55
57
|
*/
|
|
56
58
|
export function isContentDrivenDecision(source, outcome) {
|
|
57
|
-
return source === CONSOLIDATE_LEDGER_SOURCE &&
|
|
59
|
+
return (source === CONSOLIDATE_LEDGER_SOURCE &&
|
|
60
|
+
(outcome === "accepted" || outcome === "rejected" || outcome === "judged_no_action"));
|
|
58
61
|
}
|
|
59
62
|
/**
|
|
60
63
|
* Whether this row is held by its content hash: a decided consolidate
|
package/docs/README.md
CHANGED
|
@@ -49,7 +49,6 @@ Working on akm itself, not just using it.
|
|
|
49
49
|
|
|
50
50
|
- [Maintainer Docs](https://github.com/itlackey/akm/blob/main/docs/maintainers/README.md) -- Start here: local development, measuring improvement, and the curate contract
|
|
51
51
|
- [Local Development](https://github.com/itlackey/akm/blob/main/docs/maintainers/local-development.md) -- Dogfooding akm while editing its own source
|
|
52
|
-
- [akm-eval](https://github.com/itlackey/akm/blob/main/docs/maintainers/eval.md) -- Standalone toolkit for measuring whether `akm improve` is working
|
|
53
52
|
- [Curate Workmap](https://github.com/itlackey/akm/blob/main/docs/maintainers/curate-workmap.md) -- The current `akm curate` contract and the highest-value next fixes
|
|
54
53
|
|
|
55
54
|
## Look up details
|
|
@@ -100,7 +99,7 @@ Source articles for the dev.to publishing pipeline (historical record). See
|
|
|
100
99
|
- [itlackey/akm-registry](https://github.com/itlackey/akm-registry) -- the official registry index that powers built-in discovery
|
|
101
100
|
- [itlackey/akm-plugins](https://github.com/itlackey/akm-plugins) -- optional integrations for tools like OpenCode
|
|
102
101
|
- [itlackey/akm-bench](https://github.com/itlackey/akm-bench) -- the standalone benchmark harness for measuring agent performance with akm
|
|
103
|
-
- [itlackey/akm-eval](https://github.com/itlackey/akm-eval) -- the eval framework and tools for akm asset quality
|
|
102
|
+
- [itlackey/akm-eval](https://github.com/itlackey/akm-eval) -- the eval framework and tools for akm asset quality
|
|
104
103
|
|
|
105
104
|
---
|
|
106
105
|
|
|
@@ -266,7 +266,7 @@ rather than rely on `$HOME`-derived defaults (names verified against
|
|
|
266
266
|
| `AKM_CONFIG_DIR` | `config.json`'s directory. |
|
|
267
267
|
| `AKM_DATA_DIR` | Durable, non-regenerable data: **`index.db` and `state.db` live here.** This is the directory a migration snapshot's safety copy sits beside. |
|
|
268
268
|
| `AKM_CACHE_DIR` | Regenerable cache: registry downloads, config backups, task logs. Safe to discard between image builds (not between boots of the same running install). |
|
|
269
|
-
| `AKM_STATE_DIR` | **Not** where `state.db` lives, despite the name — this is the XDG "state" directory. Holds scheduled-task invocation context, companion-plugin hook state (Claude Code / OpenCode hook logs), and, per stash, `akm improve`'s
|
|
269
|
+
| `AKM_STATE_DIR` | **Not** where `state.db` lives, despite the name — this is the XDG "state" directory. Holds scheduled-task invocation context, companion-plugin hook state (Claude Code / OpenCode hook logs), and, per stash, `akm improve`'s whole-run lock (`locks/`) — see [Storage locations](https://github.com/itlackey/akm/blob/main/docs/architecture/internals/storage-locations.md). Set it anyway if you schedule akm tasks inside the image, so that context is captured consistently rather than falling back to `$HOME/.local/state/akm`. |
|
|
270
270
|
|
|
271
271
|
Set all five to paths that persist across container restarts (a mounted
|
|
272
272
|
volume), or `akm migrate apply` will see an empty `state.db` on every boot
|
|
@@ -144,7 +144,9 @@ step is idempotent — a second run reports nothing pending. 0.9.17-alpha.4
|
|
|
144
144
|
removed that step: `akm migrate` no longer relocates these files, and one left
|
|
145
145
|
at an old path is inert (nothing reads it). Two of the five writers no longer
|
|
146
146
|
exist either — the improve ledger replaced `distill-rejected/`, and the
|
|
147
|
-
write-only `eval-cases/` path was removed.
|
|
147
|
+
write-only `eval-cases/` path was removed. A third, `measurement/verdicts/`, went
|
|
148
|
+
with the `scripts/akm-eval` toolkit that wrote it, which has since left this
|
|
149
|
+
repository ([akm-eval](https://github.com/itlackey/akm/blob/main/docs/maintainers/eval.md)).
|
|
148
150
|
`$STASH/.akm/memory-cleanup/` did not move; it is the one confirmed exception
|
|
149
151
|
to the rule (see Storage locations, above).
|
|
150
152
|
|
package/docs/reference/README.md
CHANGED
|
@@ -17,4 +17,4 @@ Authoritative reference documentation for the akm CLI and its data.
|
|
|
17
17
|
- [Website Sources](https://github.com/itlackey/akm/blob/main/docs/reference/website-sources.md) -- The pluggable fetcher API behind `akm import <url>` and other URL-based knowledge reads
|
|
18
18
|
- [Data & Telemetry](data-and-telemetry.md) -- Exactly what akm reads and writes on your machine (no remote telemetry)
|
|
19
19
|
|
|
20
|
-
See also: [akm-eval](https://github.com/itlackey/akm
|
|
20
|
+
See also: [akm-eval](https://github.com/itlackey/akm-eval) -- the evals and benchmarks for measuring whether `akm improve` is working, and the repo-root [Roadmap](https://github.com/itlackey/akm/blob/main/ROADMAP.md) -- high-level focus for upcoming releases.
|
package/docs/reference/cli.md
CHANGED
|
@@ -2492,16 +2492,19 @@ day; an asset a stage looked at and left unchanged is revisited after 7 days,
|
|
|
2492
2492
|
or as soon as new feedback (or, for consolidation, an edit) arrives.
|
|
2493
2493
|
|
|
2494
2494
|
Consolidation's promotion of a memory into `knowledge/` is the exception to the
|
|
2495
|
-
7-day rule: once a promotion is accepted or rejected,
|
|
2496
|
-
|
|
2495
|
+
7-day rule: once a promotion is accepted or rejected, or the model judged the
|
|
2496
|
+
memory and proposed nothing, the memory is not offered to the model again until
|
|
2497
|
+
its body changes, however long that takes. The ledger
|
|
2497
2498
|
records the body hash the promotion was decided against and compares it with
|
|
2498
2499
|
the memory's current body (frontmatter edits do not count), the same
|
|
2499
2500
|
content-driven rule the consolidate pair pass uses. A promotion decided by an
|
|
2500
|
-
older release, which recorded no hash, keeps the old windows.
|
|
2501
|
+
older release, which recorded no hash, keeps the old windows. A memory whose
|
|
2502
|
+
body equals that of a consolidate promotion rejected on or after 2026-09-29 is
|
|
2503
|
+
held the same way, under whatever name it has. Consolidation
|
|
2501
2504
|
also does not promote a memory that `knowledge/` already covers: before it
|
|
2502
2505
|
queues a promotion it compares the memory with the 20 `knowledge/` docs in its
|
|
2503
2506
|
bundle nearest to it by stored vector, and skips the memory when one of them
|
|
2504
|
-
holds at least
|
|
2507
|
+
holds at least 30% of its distinct 5-word shingles (skip reason
|
|
2505
2508
|
`dedup_covered_by_knowledge` in the result's `consolidation.skipReasons`). A
|
|
2506
2509
|
covering doc that ranks lower than the 20th nearest goes unseen. With no stored
|
|
2507
2510
|
vector (semantic search off, or the memory not indexed yet) that check does
|
|
@@ -195,7 +195,7 @@ the set of types the code actually emits at HEAD (verified against every
|
|
|
195
195
|
| `reflect_completed` | Reflect phase produced a proposal | `ref` |
|
|
196
196
|
| `improve_reflect_outcome` | Per-asset reflect result | `ref`, `ok`, `durationMs`, `reason` |
|
|
197
197
|
| `propose_invoked` | `akm proposal new` | `ref` |
|
|
198
|
-
| `distill_invoked` | Distill phase inside the `akm improve`/`akm proposal new` pipeline. **`akm distill` is not a CLI command** — there is no standalone verb by that name | `ref`, outcome (`queued`, `skipped` with a `skipReason` such as `lesson_exists` or `conflict_noop`, `llm_failed`, `validation_failed`, `quality_rejected`, `review_needed`) |
|
|
198
|
+
| `distill_invoked` | Distill phase inside the `akm improve`/`akm proposal new` pipeline. **`akm distill` is not a CLI command** — there is no standalone verb by that name | `ref`, outcome (`queued`, `skipped` with a `skipReason` such as `lesson_exists`, `nothing_reusable` or `conflict_noop`, `llm_failed`, `validation_failed`, `quality_rejected`, `review_needed`) |
|
|
199
199
|
| `extract_invoked` | `akm proposal extract --type <harness>` / `--auto`, or improve-stage session extraction | `outcome`, `sessionId`, `harness` |
|
|
200
200
|
| `extract_triaged` | The pre-LLM extract triage gate evaluated at least one session | `evaluated`, `passed`, `triagedOut`, `sourceRun` (aggregated) |
|
|
201
201
|
| `schema_repair_invoked` | The schema-repair pass inside `akm improve` (`runSchemaRepairPass`) attempts to patch missing frontmatter on an asset that failed schema validation. **There is no `akm lint --repair` flag** — `lint` has `--fix`/`--auto-fix`, unrelated to this event | `ref`, outcome |
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "akm-cli",
|
|
3
|
-
"version": "0.9.27-alpha.
|
|
3
|
+
"version": "0.9.27-alpha.2",
|
|
4
4
|
"type": "module",
|
|
5
5
|
"description": "akm (Agent Knowledge Manager) — a portable, local-first capability library for AI agents. Discover, load, share, and improve reusable skills, scripts, workflows, and knowledge across any shell-capable coding agent, including Claude Code, OpenCode, and Cursor.",
|
|
6
6
|
"keywords": [
|