akm-cli 0.9.26 → 0.9.27-alpha.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +237 -0
- package/LICENSE +3 -4
- package/dist/assets/prompts/consolidate-system.md +2 -2
- package/dist/assets/prompts/distill-lesson-system.md +29 -7
- package/dist/assets/prompts/extract-session.md +2 -2
- package/dist/commands/improve/consolidate/coverage.js +71 -17
- package/dist/commands/improve/consolidate/pair-pass.js +12 -9
- package/dist/commands/improve/consolidate.js +77 -5
- package/dist/commands/improve/distill-guards.js +9 -10
- package/dist/commands/improve/distill.js +70 -22
- package/dist/commands/improve/extract-prompt.js +61 -39
- package/dist/commands/improve/extract.js +2 -1
- package/dist/commands/improve/preparation.js +14 -1
- package/dist/commands/improve/reflect.js +36 -4
- package/dist/commands/improve/retrieval-gate.js +1 -1
- package/dist/commands/improve/session-asset.js +3 -2
- package/dist/commands/improve/stage.js +35 -53
- package/dist/commands/proposal/drain.js +85 -21
- package/dist/commands/proposal/proposal-types.js +1 -1
- package/dist/core/config/schema/improve-processes.js +3 -3
- package/dist/core/paths.js +0 -9
- package/dist/indexer/indexer.js +35 -19
- package/dist/integrations/harnesses/codex/agent-builder.js +23 -15
- package/dist/llm/client.js +44 -14
- package/dist/llm/memory-infer.js +1 -1
- package/dist/scripts/akm-migrate-node.js +25 -20
- package/dist/scripts/akm-migrate.js +25 -20
- package/dist/storage/repositories/improve-ledger-repository.js +4 -1
- package/dist/storage/repositories/index-entry-schema.js +20 -4
- package/dist/storage/repositories/index-fts-repository.js +44 -3
- package/dist/storage/repositories/index-schema.js +14 -8
- package/docs/README.md +1 -2
- package/docs/integration/bundling-akm.md +1 -1
- package/docs/migration/v0.8-to-v0.9.md +3 -1
- package/docs/reference/README.md +1 -1
- package/docs/reference/cli.md +8 -4
- package/docs/reference/configuration.md +5 -2
- package/docs/reference/data-and-telemetry.md +1 -1
- package/package.json +1 -1
package/CHANGELOG.md
CHANGED
|
@@ -6,6 +6,243 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/).
|
|
|
6
6
|
|
|
7
7
|
## [Unreleased]
|
|
8
8
|
|
|
9
|
+
## [0.9.27-alpha.2] - 2026-10-07
|
|
10
|
+
|
|
11
|
+
### Removed
|
|
12
|
+
|
|
13
|
+
- **The `scripts/akm-eval` toolkit has moved out of this repository.** The
|
|
14
|
+
read-only measurement toolkit (the case runner and its suites, the twin
|
|
15
|
+
experiment, the real-query verdict for the proactive lane, the state
|
|
16
|
+
analyzers and the curate benchmark) is retired; every live eval is in
|
|
17
|
+
[itlackey/akm-eval](https://github.com/itlackey/akm-eval). Its code is kept
|
|
18
|
+
there, to read and not to run, in `retired/akm-scripts-akm-eval/`, copied from
|
|
19
|
+
commit `f57a7fd44b37`. It imports akm's `src/` by relative path, so it runs
|
|
20
|
+
only in a checkout at that commit. Removed here with it: `scripts/akm-eval/`,
|
|
21
|
+
its tests (`tests/integration/akm-eval/`, `tests/akm-eval-*.test.ts`,
|
|
22
|
+
`tests/curate-metrics.test.ts`) and fixtures (`tests/fixtures/akm-eval/`, and
|
|
23
|
+
the `curate-golden` stash, which only the curate benchmark read), the
|
|
24
|
+
`akm-eval determinism` CI job, and `getMeasurementVerdictsDir`, whose only
|
|
25
|
+
caller was the verdict runner. akm no longer names
|
|
26
|
+
`$STATE/improve/measurement/verdicts/<stash>/`; a file already there is inert.
|
|
27
|
+
`docs/maintainers/eval.md` is now a pointer to the new home.
|
|
28
|
+
|
|
29
|
+
### Fixed
|
|
30
|
+
|
|
31
|
+
- **The drain's judge no longer sees a note with a code block as truncated.**
|
|
32
|
+
The judgment prompt fenced the proposed content (and the live asset, sibling
|
|
33
|
+
proposals and neighbour excerpts) in three backticks, so a note holding its
|
|
34
|
+
own code block closed the fence early and read as cut off; real rejections said
|
|
35
|
+
"ends in an empty code block" or "truncated". Each block now uses a fence longer
|
|
36
|
+
than any backtick run inside it. The judge's reason is also kept on accepts,
|
|
37
|
+
staged accepts and defers (as the gate decision's `judgeReason`, until now
|
|
38
|
+
rejections only), and a judge reply that is not a verdict is stamped
|
|
39
|
+
`judgment-parse-failure`, and a runner failure `judgment-error`, instead of
|
|
40
|
+
looking like a defer.
|
|
41
|
+
|
|
42
|
+
- **Distill writes a lesson only when its memory holds one, says only what the
|
|
43
|
+
memory says, and its judge rejects what a reviewer would.** 2 of the 22 distill
|
|
44
|
+
proposals since 0.9.26 began were accepted, and 17 of the 19 queued on
|
|
45
|
+
2026-10-05 were bad (they restated their memory, filed a dated status as a
|
|
46
|
+
lesson, claimed what the memory does not say, or repeated an asset the library
|
|
47
|
+
holds). Four causes, found in the code and the rejected proposals, and fixed:
|
|
48
|
+
(1) the prompt and schema forced a lesson from every memory, and 18 of the 19
|
|
49
|
+
were records of what was done; the writer now says why a memory holds a
|
|
50
|
+
lesson or none (`reason`, then `decision: lesson|none`, or the word `NONE`),
|
|
51
|
+
defined as a cause and what to do about it, or a rule with its reason, and
|
|
52
|
+
writes only what the memory and its feedback state, in the scope they have; a
|
|
53
|
+
`NONE` is a `skipped` distill (`skipReason: nothing_reusable`, the writer's
|
|
54
|
+
reason in the message) with no proposal and no judge call, and the loop keeps
|
|
55
|
+
its ledger row `unchanged`. (2) The judge asked for "information not already
|
|
56
|
+
present in the source", so an invented claim scored as novel and a faithful
|
|
57
|
+
lesson of a lesson-worthy memory as a restatement, and it passed anything that
|
|
58
|
+
"goes beyond the source" because it "may draw on feedback you are not shown".
|
|
59
|
+
The rubric is now reusable (a rule with its reason, not a record of what was
|
|
60
|
+
done), non-redundancy and grounding (every cause, step, number and limit is
|
|
61
|
+
in the source or its feedback), and the judge is shown the feedback the writer
|
|
62
|
+
saw. (3) A mean hid a decisive score (4 and 1 average 2.5, a review), and a
|
|
63
|
+
reviewer read everything the judge did not reject; any criterion at 2 or
|
|
64
|
+
below, grounding included, is now `quality_rejected`, and the reason names it
|
|
65
|
+
(`grounding 2/5: …`). The "borderline grounding" routing is gone. (4) Neither
|
|
66
|
+
the writer nor the judge could see a knowledge note or a skill that already
|
|
67
|
+
states the rule (the judge saw the 3 lexically nearest lessons, none of them
|
|
68
|
+
related); both now see the lessons, knowledge notes and skills nearest the
|
|
69
|
+
memory, which is the existing `processes.distill.cls` context turned on by
|
|
70
|
+
default (`enabled: false` turns it off). Judge scores are keyed `reusable`
|
|
71
|
+
where they were `novelty`. Measured on the local qwen3.8-27b with akm-eval's
|
|
72
|
+
`evals/distill` (30 fictional memories, 5 runs each side): good lessons 7/14 on
|
|
73
|
+
average (5 to 9) against 3.7/14 (3 to 5), lessons queued for memories that
|
|
74
|
+
deserve none 0.4/16 against 3.3/16. On 37 real memories with their feedback
|
|
75
|
+
(34 reviewed bad, 3 good; 3 runs against 2): a lesson was queued for 11% of the
|
|
76
|
+
bad ones against 44%, and for 6 of 9 good ones against 4 of 6; of the memories
|
|
77
|
+
that pass 0.9.26's skip of bare positive feedback, 20% of the bad against 50%.
|
|
78
|
+
No new settings.
|
|
79
|
+
|
|
80
|
+
- **Reflect no longer plans an asset whose negative feedback is already acted
|
|
81
|
+
on.** A negative `akm feedback` that came with an exact fix (`--replace` and
|
|
82
|
+
`--with`, `--outdated` or `--superseded-by`) makes a `feedback` proposal, and
|
|
83
|
+
once that proposal is accepted the feedback has done its work. Reflect still
|
|
84
|
+
took the ref as having fresh negative feedback, and on 2026-10-07 44 of its 50
|
|
85
|
+
refs were of that kind: the judge refused or the model changed nothing for
|
|
86
|
+
most of them. A negative event with a fix is now left out of the reflect
|
|
87
|
+
cursor when an accepted `feedback` proposal for the ref was created at or
|
|
88
|
+
after it. A negative with no fix, one given after the proposal, and one whose
|
|
89
|
+
proposal is still pending or was rejected plan a reflect as before.
|
|
90
|
+
|
|
91
|
+
- **Consolidate stops re-offering memories a reviewer already turned down, and
|
|
92
|
+
the nightly judge sees what a promotion may duplicate.** About 53 promotions a
|
|
93
|
+
night reached review at ~5% precision, 64-70% of them a memory body already
|
|
94
|
+
proposed or rejected. Four causes, four changes: a memory whose body equals
|
|
95
|
+
that of a consolidate promotion rejected on or after 2026-09-29 is held until
|
|
96
|
+
its body changes, under any name (earlier rejections, the bulk audits of
|
|
97
|
+
2026-08, do not count); a memory the model judged and left alone is held by its
|
|
98
|
+
body hash instead of a 7-day clock, so an unchanged memory is no longer judged
|
|
99
|
+
every week (a row recorded without a hash keeps the 7 days); the coverage gate
|
|
100
|
+
skips a memory when 30% of its text, not 50%, is in a neighbouring knowledge
|
|
101
|
+
doc, which catches paraphrases; and the drain's judgment tier, which judged a
|
|
102
|
+
promotion seeing only the proposal and never `knowledge/`, is now shown the 5
|
|
103
|
+
nearest knowledge notes (ref, description, excerpt) and told to reject a
|
|
104
|
+
promotion they already cover. No new settings.
|
|
105
|
+
|
|
106
|
+
- **A confident `subsumed` or `supersedes` retirement resolves unattended, as a
|
|
107
|
+
`duplicate` already did.** The pair pass staged a retire proposal for the
|
|
108
|
+
triage drain only when the judge's label was `duplicate`; every other retirement
|
|
109
|
+
waited for a person. It now stages any of the three retire labels when the
|
|
110
|
+
second look (what does the retired note hold that the kept one lacks?) comes
|
|
111
|
+
back empty and there is no continuity risk; the retired side's claim list is
|
|
112
|
+
already empty for any proposal, and a `duplicate` must still have an empty list
|
|
113
|
+
on the kept side too. The staged gate reason is the judge's label, and the drain records it. Replay over 360
|
|
114
|
+
judged pairs: 336 safe (0.93): `duplicate` 0.98, `subsumed` 0.92, `supersedes`
|
|
115
|
+
0.875. In production, unstaged `subsumed` retirements were accepted 45 of 54
|
|
116
|
+
times by hand, and in the latest run 41 of 47 pair proposals would have resolved
|
|
117
|
+
without a person. `docs/architecture/internals/improve-workflow.md` said triage
|
|
118
|
+
never auto-accepts a retire proposal, which stopped being true in 0.9.26; it,
|
|
119
|
+
and the matching lines in `improvement.md`, now describe the staging rule.
|
|
120
|
+
|
|
121
|
+
## [0.9.27-alpha.1] - 2026-10-06
|
|
122
|
+
|
|
123
|
+
### Fixed
|
|
124
|
+
|
|
125
|
+
- **A codex dispatch with an output schema no longer leaves a temp folder
|
|
126
|
+
behind.** Every build of the codex command for a request with a schema made a
|
|
127
|
+
new `akm-codex-schema-*` folder in the OS temp dir for `--output-schema` and
|
|
128
|
+
nothing ever removed it (the builder has no post-run hook, and the file is read
|
|
129
|
+
after it returns), so a machine whose `/tmp` is tmpfs held a folder in RAM per
|
|
130
|
+
dispatch until reboot. The schema is now written once to akm's cache dir, in a
|
|
131
|
+
file named by its hash: concurrent units dispatching the same schema share it,
|
|
132
|
+
a rewrite is an atomic rename of identical bytes, and the only residue is one
|
|
133
|
+
small file per distinct schema.
|
|
134
|
+
- **An index that has been updated ranks like a fresh index of the same files,
|
|
135
|
+
and `akm index --full` no longer doubles the full-text totals.** `entries_fts`
|
|
136
|
+
is contentless, and FTS5 cannot take a deleted row out of a contentless
|
|
137
|
+
table's BM25 totals (its row count, and the token counts the average document
|
|
138
|
+
length comes from). Every replaced or removed row left them one row too high:
|
|
139
|
+
40 notes read 40, then 41 after one edit, then 81 after `--full`, and a delete
|
|
140
|
+
never lowered them, so scores drifted away from what a fresh index gives.
|
|
141
|
+
SQLite has no command that recomputes them (`delete` and `rebuild` are refused
|
|
142
|
+
on a contentless table), so a delete that removes a row now stamps
|
|
143
|
+
`index_meta.ftsTotalsStale` in its own transaction, whichever process made it,
|
|
144
|
+
and the next `akm index` rebuilds the table from `entries` before it finishes:
|
|
145
|
+
about a second at 25,000 entries, and only when rows have left the table. A
|
|
146
|
+
row the write-path index replaced after an accepted proposal is settled the
|
|
147
|
+
same way, and an index that has already drifted is corrected by the first run
|
|
148
|
+
that replaces or removes a row.
|
|
149
|
+
- **`akm improve` runs against an API that rejects `chat_template_kwargs`,
|
|
150
|
+
OpenAI's among them.** Improve's reflect, consolidate and judge calls always
|
|
151
|
+
ask for thinking off, and the client sends that as
|
|
152
|
+
`chat_template_kwargs.enable_thinking` and a top-level `enable_thinking`. A
|
|
153
|
+
strict API answers 400 `Unknown parameter: 'chat_template_kwargs'`, the retry
|
|
154
|
+
without the response schema sent both fields again, and `akm improve judge`
|
|
155
|
+
reported `judge timeout/error — routed to review`. No engine setting could
|
|
156
|
+
stop it: the call sites override the engine's `enableThinking`, and
|
|
157
|
+
`extraParams` can only add fields. A 4xx that names either field is now
|
|
158
|
+
answered by one retry without both (which may in turn fall back without the
|
|
159
|
+
schema), and akm stops sending them to that endpoint and model for the rest
|
|
160
|
+
of the process, as it already does for `response_format`.
|
|
161
|
+
- **A lesson the model wrote without a `when_to_use` is no longer thrown away
|
|
162
|
+
silently by `akm proposal extract`.** The extract schema left `when_to_use`
|
|
163
|
+
optional while the parser dropped a lesson without one (or with one under 15
|
|
164
|
+
characters) and said nothing, so a model that followed the schema could
|
|
165
|
+
write a sound lesson that akm discarded and the session reported no
|
|
166
|
+
candidates: on akm-eval's extract eval qwen3.8-27b kept 3 of 11 expected
|
|
167
|
+
insights against 9 for gpt-oss-120b. Every property of the schema is now
|
|
168
|
+
required, `when_to_use` and `rationale_if_empty` included, with an empty
|
|
169
|
+
string for none (a memory or knowledge candidate needs no trigger, a
|
|
170
|
+
non-empty answer no rationale), which is also what a strict structured-output
|
|
171
|
+
provider needs; the prompt's output contract says the same. Any candidate the
|
|
172
|
+
contract still refuses, for this reason or another, is named in its session's
|
|
173
|
+
`warnings` as `<type>:<name> dropped: <reason>`.
|
|
174
|
+
- **Distill's response schemas are valid for a strict structured-output
|
|
175
|
+
provider.** The client sends a response schema `strict: true`, and OpenAI
|
|
176
|
+
refuses one whose objects leave a property out of `required`: `400 Invalid
|
|
177
|
+
schema for response_format 'akm_response': ... Missing 'tags'`. The lesson
|
|
178
|
+
schema left out `tags`, and the knowledge schema `tags` and `sources`, so
|
|
179
|
+
the first distill request on such a provider always failed. After a 4xx the
|
|
180
|
+
client retries once without the schema, but a gateway that answers the same
|
|
181
|
+
rejection with a 502 is not retried, and every distill call through it
|
|
182
|
+
failed; the workaround, `supportsJsonSchema: false`, loses the guidance that
|
|
183
|
+
keeps a model from leaving out `when_to_use`. Every property is now
|
|
184
|
+
required, and an empty array stands for none (distill already dropped an
|
|
185
|
+
empty `tags` or `sources`).
|
|
186
|
+
- **`akm improve` no longer reflects on an asset whose file a proposal would not
|
|
187
|
+
write, so accepting a reflect proposal no longer adds a second file.** A
|
|
188
|
+
skill's `references/a.md` is indexed as `knowledge/skills/<name>/references/a`,
|
|
189
|
+
but a proposal writes the path derived from that ref,
|
|
190
|
+
`knowledge/skills/<name>/references/a.md`. Nothing is there, so the proposal
|
|
191
|
+
was a `create`, and accepting it wrote a copy beside the skill's own file;
|
|
192
|
+
later proposals then revised the copy while the skill's file drifted. Reflect
|
|
193
|
+
now refuses before it calls the model, naming the file and the path a
|
|
194
|
+
proposal would write, and the loop records it as a skip (`unsupported_type`,
|
|
195
|
+
`file_outside_layout` in the `reflect_completed` event). This is the rule
|
|
196
|
+
0.9.26 added to `akm feedback --replace`. An asset that another bundle owns is
|
|
197
|
+
still refused by `createProposal` (#1000).
|
|
198
|
+
- **An index that 0.9.1 wrote no longer crashes akm.** Every `akm index` on it
|
|
199
|
+
exited 70 with `null is not an object (evaluating 'doc.xrefs')`, `akm migrate
|
|
200
|
+
apply` and `akm index --full` did not help, and `akm search` failed with
|
|
201
|
+
`null is not an object (evaluating 'item.entry.quality')`; the only way out
|
|
202
|
+
was moving `index.db` aside. That layout (20) keeps the transitional
|
|
203
|
+
`entry_key`, `dir_path`, `stash_dir`, `entry_json` and `entry_type` columns
|
|
204
|
+
that layout 21 removed, each NOT NULL, beside the current columns, and leaves
|
|
205
|
+
`document_json` NULL on every row. The table had every column akm checks for,
|
|
206
|
+
so it was taken for a current one: the links migration read the NULL
|
|
207
|
+
documents, and no insert could ever have succeeded (`NOT NULL constraint
|
|
208
|
+
failed: entries.entry_key`). akm now treats a table that still has a retired
|
|
209
|
+
column as older than layout 21, as the compat notes already said: the
|
|
210
|
+
writable opener recreates its entries-keyed tables (the LLM enrichment cache
|
|
211
|
+
is kept) and the next `akm index` re-walks every source, and a read rebuilds
|
|
212
|
+
inline. An index that a failed open already half-migrated (layout stamp still
|
|
213
|
+
20, `search_text` dropped, `asset_links` created) recovers the same way.
|
|
214
|
+
- **Two files that claim one ref no longer trade places in the index, and `akm
|
|
215
|
+
index` says so.** A skill's `references/a.md` and a note at
|
|
216
|
+
`knowledge/skills/x/references/a.md` are both the ref
|
|
217
|
+
`knowledge/skills/x/references/a`, and the index holds one row for it. The
|
|
218
|
+
first file a run persisted held it, so a full build followed the filesystem's
|
|
219
|
+
listing order (one bundle indexed on tmpfs and on ext4 held different files),
|
|
220
|
+
and the first incremental run after a full build handed the row to the other
|
|
221
|
+
file, because it drains only the directory that lost: with no file touched,
|
|
222
|
+
the row, its search entry and the text its vector is embedded from changed,
|
|
223
|
+
and the vector was dropped and recomputed. When a smaller-path file was added
|
|
224
|
+
later and then deleted, the ref also left the index until `--full`, although
|
|
225
|
+
its other file was still on disk. The file with the smaller path (code-point
|
|
226
|
+
order, as `akm show`'s refusal lists them) now holds the ref however the
|
|
227
|
+
directories are drained and the walk is ordered, a directory that gives a ref
|
|
228
|
+
up is drained again so the ref passes back when its holder goes, and each
|
|
229
|
+
pair is reported in the `warnings` of `akm index`, naming the file indexed
|
|
230
|
+
and the one skipped.
|
|
231
|
+
- **Consolidate's plan schema and the session summary schema are valid for a
|
|
232
|
+
strict structured-output provider, and a new response schema can no longer
|
|
233
|
+
skip the rule.** The client sends a response schema `strict: true`, and
|
|
234
|
+
OpenAI refuses one whose objects leave a property out of `required`. The
|
|
235
|
+
consolidate plan left out `description` and `confidence`, and the session
|
|
236
|
+
summary `tags`, so the first request of every plan and every summary to such
|
|
237
|
+
a provider was refused. The client then retries without the schema and
|
|
238
|
+
remembers that per connection (endpoint and model), not per schema, so one
|
|
239
|
+
invalid schema also switched the response schema off for every valid one
|
|
240
|
+
that followed on that connection. Every property of both is now required: an
|
|
241
|
+
empty `description` keeps the memory's own, a null `confidence` records none
|
|
242
|
+
and an empty `tags` array is no tags, which is how all three were already
|
|
243
|
+
read. A contract test runs every schema akm sends through the rule, so a
|
|
244
|
+
property added later without being required fails CI.
|
|
245
|
+
|
|
9
246
|
## [0.9.26] - 2026-10-05
|
|
10
247
|
|
|
11
248
|
The stable release of the 0.9.26 line: 0.9.26-alpha.1 and alpha.2, and the
|
package/LICENSE
CHANGED
|
@@ -225,10 +225,9 @@ statute, judicial order, or regulation then You must: (a) comply with
|
|
|
225
225
|
the terms of this License to the maximum extent possible; and (b)
|
|
226
226
|
describe the limitations and the code they affect. Such description must
|
|
227
227
|
be placed in a text file included with all distributions of the Covered
|
|
228
|
-
Software under
|
|
229
|
-
|
|
230
|
-
|
|
231
|
-
for a recipient of ordinary skill to be able to understand it.
|
|
228
|
+
Software under this License. Except to the extent prohibited by statute
|
|
229
|
+
or regulation, such description must be sufficiently detailed for a
|
|
230
|
+
recipient of ordinary skill to be able to understand it.
|
|
232
231
|
|
|
233
232
|
5. Termination
|
|
234
233
|
--------------
|
|
@@ -7,10 +7,10 @@ Rules:
|
|
|
7
7
|
Return ONLY JSON (no prose, no code fences):
|
|
8
8
|
{
|
|
9
9
|
"operations": [
|
|
10
|
-
{ "op": "promote", "ref": "memories/<name>", "knowledgeRef": "knowledge/<suggested-slug>", "reason": "<brief reason>", "description": "<one sentence describing the new knowledge asset>", "confidence": 0.92 }
|
|
10
|
+
{ "op": "promote", "ref": "memories/<name>", "knowledgeRef": "knowledge/<suggested-slug>", "reason": "<brief reason>", "description": "<one sentence describing the new knowledge asset, or an empty string to keep the memory's own>", "confidence": 0.92 }
|
|
11
11
|
]
|
|
12
12
|
}
|
|
13
13
|
|
|
14
|
-
For every operation, emit a `confidence` field in [0, 1] expressing your certainty that the operation is correct and safe. Use 0.95+ only when evidence is unambiguous.
|
|
14
|
+
For every operation, emit a `confidence` field in [0, 1] expressing your certainty that the operation is correct and safe. Use 0.95+ only when evidence is unambiguous. Use `null` rather than guessing if you are uncertain.
|
|
15
15
|
|
|
16
16
|
When the merged content includes an `updated` frontmatter field, the value MUST be a real ISO date string (e.g. `updated: 2026-05-20`). NEVER emit `updated: today`, `updated: {today}`, `updated: {today: null}`, `updated: now`, or any other literal placeholder/template-variable. If you do not have a real source-of-truth date, OMIT the `updated` field entirely — the post-processor will not invent one for you.
|
|
@@ -1,7 +1,29 @@
|
|
|
1
1
|
You are the akm `distill` distiller.
|
|
2
|
-
|
|
3
|
-
|
|
4
|
-
|
|
2
|
+
You are given a memory and the feedback recorded about it. Decide whether it
|
|
3
|
+
holds a lesson and, if it does, write the lesson.
|
|
4
|
+
|
|
5
|
+
A memory holds a lesson when it states a cause and what to do about it: a
|
|
6
|
+
failure or surprise with its cause and the fix that worked, or a rule with the
|
|
7
|
+
reason it holds. It holds a lesson even when it is short and names one project,
|
|
8
|
+
tool or incident, if the cause and the fix would help someone in a similar
|
|
9
|
+
situation.
|
|
10
|
+
|
|
11
|
+
A memory holds NO lesson when all it states is what was done, shipped, released
|
|
12
|
+
or decided, what is pending or planned, how a system is set up now, or the steps
|
|
13
|
+
of a procedure, with no failure and cause behind it, or when its feedback says
|
|
14
|
+
only that it is out of date or superseded. ANSWER NONE then: the single word and
|
|
15
|
+
nothing else. A reply bound to a JSON schema answers NONE by
|
|
16
|
+
setting `decision` to `none` and leaving the other fields empty. Answer NONE too
|
|
17
|
+
when a related asset listed below the memory already states the rule the memory
|
|
18
|
+
would give.
|
|
19
|
+
|
|
20
|
+
When the memory holds a lesson, write it from what the memory and its feedback
|
|
21
|
+
say, and nothing more:
|
|
22
|
+
- Add no cause, step, rule, number, check or safeguard that neither states.
|
|
23
|
+
- Keep the scope the memory has. A fix verified in one place is a fix for that
|
|
24
|
+
place, and what was not checked stays unchecked. One case is not "always" or
|
|
25
|
+
"never".
|
|
26
|
+
- Be shorter than the memory.
|
|
5
27
|
|
|
6
28
|
YOUR RESPONSE MUST START EXACTLY WITH `---` ON THE VERY FIRST LINE.
|
|
7
29
|
DO NOT output any prose, explanation, or code fences before or after.
|
|
@@ -12,7 +34,7 @@ description: <one complete sentence (ending with `.`) summarising what the lesso
|
|
|
12
34
|
when_to_use: <one complete sentence describing the concrete trigger condition>
|
|
13
35
|
---
|
|
14
36
|
|
|
15
|
-
<lesson body — plain markdown,
|
|
37
|
+
<lesson body — plain markdown, as short as the memory allows>
|
|
16
38
|
|
|
17
39
|
## description field (MANDATORY)
|
|
18
40
|
- A single complete sentence in present tense, 20–400 chars, NO markdown.
|
|
@@ -21,7 +43,7 @@ when_to_use: <one complete sentence describing the concrete trigger condition>
|
|
|
21
43
|
- DO NOT copy a section heading ("Key takeaways", "For example", "Key pitfalls").
|
|
22
44
|
- DO NOT begin with a numbered list marker, code fence, or markdown heading.
|
|
23
45
|
|
|
24
|
-
GOOD: "
|
|
46
|
+
GOOD: "Pin the container image tag, because the `latest` tag moved under the nightly job and its output changed with no code change."
|
|
25
47
|
BAD: "Key pitfalls"
|
|
26
48
|
BAD: "When working with the akm CLI"
|
|
27
49
|
BAD: "For example, you might..."
|
|
@@ -32,5 +54,5 @@ RULES:
|
|
|
32
54
|
- `description` and `when_to_use` MUST differ from each other.
|
|
33
55
|
- The lesson body MUST be non-empty markdown prose. Do NOT restate `description:` or `when_to_use:` inside the body (no `**description:** ...` or `**when_to_use:** ...` lines — the frontmatter is the only place those keys belong).
|
|
34
56
|
- Do NOT emit a second `---` fence after the opening frontmatter — there are exactly two `---` lines in the output, both belonging to the single frontmatter block at the top.
|
|
35
|
-
- Do NOT reproduce the source asset verbatim
|
|
36
|
-
- Output ONLY the lesson file. No preamble, no code fences, no trailing prose.
|
|
57
|
+
- Do NOT reproduce the source asset verbatim.
|
|
58
|
+
- Output ONLY the lesson file. No preamble, no code fences, no trailing prose.
|
|
@@ -48,13 +48,13 @@ Respond with EXACTLY one JSON object matching this shape:
|
|
|
48
48
|
"type": "memory" | "lesson" | "knowledge",
|
|
49
49
|
"name": "<kebab-case name, e.g. jwt-token; optionally under one kebab-case scope, e.g. auth/jwt-token>",
|
|
50
50
|
"description": "<one sentence 20-400 chars>",
|
|
51
|
-
"when_to_use": "<one sentence 15-400 chars;
|
|
51
|
+
"when_to_use": "<one sentence 15-400 chars for a lesson; an empty string for a memory or knowledge candidate>",
|
|
52
52
|
"body": "<markdown body, 200-3000 chars typical>",
|
|
53
53
|
"confidence": <number 0.0-1.0>,
|
|
54
54
|
"evidence": "<one-line pointer to the moment in the session>"
|
|
55
55
|
}
|
|
56
56
|
],
|
|
57
|
-
"rationale_if_empty": "<one sentence
|
|
57
|
+
"rationale_if_empty": "<one sentence when candidates is empty; an empty string otherwise>"
|
|
58
58
|
}
|
|
59
59
|
```
|
|
60
60
|
|
|
@@ -6,7 +6,7 @@
|
|
|
6
6
|
* `knowledge/` proposal, ask whether `knowledge/` already says it.
|
|
7
7
|
*
|
|
8
8
|
* The rule: a memory is covered when at least {@link COVERAGE_MIN_CONTAINMENT}
|
|
9
|
-
* (
|
|
9
|
+
* (30%) of its distinct {@link COVERAGE_SHINGLE_WORDS}-word shingles appear in
|
|
10
10
|
* one of the knowledge docs nearest to it. It is a containment of the MEMORY in
|
|
11
11
|
* the doc, not a similarity: a long guide that quotes the memory covers it; a
|
|
12
12
|
* memory that quotes a short doc and adds claims of its own does not.
|
|
@@ -25,9 +25,13 @@
|
|
|
25
25
|
* this gate achieves: it reads only the {@link PAIR_NEIGHBOR_FETCH_K} nearest
|
|
26
26
|
* knowledge docs (below), and a covering doc that ranks lower goes unseen. The
|
|
27
27
|
* recall over that candidate set is unmeasured. The rest of the rejected
|
|
28
|
-
* proposals (paraphrases, partial overlaps) still reach review
|
|
29
|
-
*
|
|
30
|
-
*
|
|
28
|
+
* proposals (paraphrases, partial overlaps) still reach review. 0.5 let
|
|
29
|
+
* paraphrases through: of the 54 promotions minted on 2026-10-07, 18 of the 51
|
|
30
|
+
* later rejected held 30% or more of their text in a neighbouring doc. The cut is
|
|
31
|
+
* now 0.3, between the two measured points: 0.2 (169 of 224 rejected) also
|
|
32
|
+
* skipped 2 of the 103 accepted ones. At 0.3, none of the 5 promotions graded
|
|
33
|
+
* good in the 2026-10-05 review sample would be skipped (their best doc holds
|
|
34
|
+
* at most 1% of them) while 5 of its 15 bad ones would. A wrong skip is a
|
|
31
35
|
* promotion nobody gets to review.
|
|
32
36
|
*
|
|
33
37
|
* Candidates are the memory's {@link PAIR_NEIGHBOR_FETCH_K} nearest knowledge
|
|
@@ -49,7 +53,7 @@ import { PAIR_NEIGHBOR_FETCH_K } from "./pair-pass.js";
|
|
|
49
53
|
/** Words per shingle. */
|
|
50
54
|
export const COVERAGE_SHINGLE_WORDS = 5;
|
|
51
55
|
/** Share of a memory's distinct shingles one knowledge doc must hold for the memory to count as covered. */
|
|
52
|
-
export const COVERAGE_MIN_CONTAINMENT = 0.
|
|
56
|
+
export const COVERAGE_MIN_CONTAINMENT = 0.3;
|
|
53
57
|
const WORD = /[\p{L}\p{N}]+/gu;
|
|
54
58
|
/** The distinct lower-cased word n-grams of `text`; empty when it has fewer than {@link COVERAGE_SHINGLE_WORDS} words. */
|
|
55
59
|
export function wordShingles(text) {
|
|
@@ -72,19 +76,15 @@ export function shingleContainment(memory, doc) {
|
|
|
72
76
|
return shared / memory.size;
|
|
73
77
|
}
|
|
74
78
|
/**
|
|
75
|
-
* The
|
|
76
|
-
*
|
|
77
|
-
*
|
|
78
|
-
* and so no candidates.
|
|
79
|
+
* The {@link PAIR_NEIGHBOR_FETCH_K} knowledge docs in `bundleId` nearest to the
|
|
80
|
+
* memory at `filePath`, nearest first. A memory the index does not know has no
|
|
81
|
+
* stored vector and so no neighbours.
|
|
79
82
|
*/
|
|
80
|
-
|
|
81
|
-
const shingles = wordShingles(body);
|
|
82
|
-
if (shingles.size === 0)
|
|
83
|
-
return undefined;
|
|
83
|
+
function knowledgeNeighbours(db, bundleId, filePath) {
|
|
84
84
|
const entryId = getEntryIdByFilePath(db, filePath);
|
|
85
85
|
if (entryId === undefined)
|
|
86
|
-
return
|
|
87
|
-
|
|
86
|
+
return [];
|
|
87
|
+
const out = [];
|
|
88
88
|
for (const hit of getNeighborsByEntryId(db, entryId, PAIR_NEIGHBOR_FETCH_K, { type: "knowledge", bundleId })) {
|
|
89
89
|
const neighbour = getEntryById(db, hit.id);
|
|
90
90
|
if (!neighbour)
|
|
@@ -96,13 +96,67 @@ export function findCoveringKnowledge(db, bundleId, filePath, body) {
|
|
|
96
96
|
catch {
|
|
97
97
|
continue; // the index outlived the file
|
|
98
98
|
}
|
|
99
|
-
|
|
99
|
+
out.push({
|
|
100
|
+
ref: neighbour.conceptId,
|
|
101
|
+
description: neighbour.entry.description ?? "",
|
|
102
|
+
body: stripFrontmatterBody(raw),
|
|
103
|
+
});
|
|
104
|
+
}
|
|
105
|
+
return out;
|
|
106
|
+
}
|
|
107
|
+
/**
|
|
108
|
+
* The best-covering knowledge doc among the {@link PAIR_NEIGHBOR_FETCH_K}
|
|
109
|
+
* knowledge docs in `bundleId` nearest to the memory. `filePath` is the
|
|
110
|
+
* memory's indexed file; a memory the index does not know has no stored
|
|
111
|
+
* vector and so no candidates.
|
|
112
|
+
*/
|
|
113
|
+
export function findCoveringKnowledge(db, bundleId, filePath, body) {
|
|
114
|
+
const shingles = wordShingles(body);
|
|
115
|
+
if (shingles.size === 0)
|
|
116
|
+
return undefined;
|
|
117
|
+
let best;
|
|
118
|
+
for (const neighbour of knowledgeNeighbours(db, bundleId, filePath)) {
|
|
119
|
+
const containment = shingleContainment(shingles, neighbour.body);
|
|
100
120
|
if (containment >= COVERAGE_MIN_CONTAINMENT && (best === undefined || containment > best.containment)) {
|
|
101
|
-
best = { ref: neighbour.
|
|
121
|
+
best = { ref: neighbour.ref, containment };
|
|
102
122
|
}
|
|
103
123
|
}
|
|
104
124
|
return best;
|
|
105
125
|
}
|
|
126
|
+
/** Knowledge docs a reviewer is shown for a promotion, nearest first. */
|
|
127
|
+
export const NEIGHBOUR_NOTE_COUNT = 5;
|
|
128
|
+
const NEIGHBOUR_EXCERPT_CHARS = 300;
|
|
129
|
+
/**
|
|
130
|
+
* The knowledge notes nearest to the memory at `memoryPath`, for the drain's
|
|
131
|
+
* judge to compare a promotion against: the nearest {@link NEIGHBOUR_NOTE_COUNT}
|
|
132
|
+
* of the same candidates the coverage gate reads. The memory's bundle is the one
|
|
133
|
+
* the index recorded for it. Empty when the index has no vector for the memory
|
|
134
|
+
* or cannot be opened; never throws.
|
|
135
|
+
*/
|
|
136
|
+
export function nearestKnowledgeNotes(memoryPath) {
|
|
137
|
+
let db;
|
|
138
|
+
try {
|
|
139
|
+
db = openExistingDatabase();
|
|
140
|
+
const entryId = getEntryIdByFilePath(db, memoryPath);
|
|
141
|
+
const bundleId = entryId === undefined ? undefined : getEntryById(db, entryId)?.bundleId;
|
|
142
|
+
if (bundleId === undefined)
|
|
143
|
+
return [];
|
|
144
|
+
return knowledgeNeighbours(db, bundleId, memoryPath)
|
|
145
|
+
.slice(0, NEIGHBOUR_NOTE_COUNT)
|
|
146
|
+
.map((n) => ({
|
|
147
|
+
ref: n.ref,
|
|
148
|
+
description: n.description,
|
|
149
|
+
excerpt: n.body.length > NEIGHBOUR_EXCERPT_CHARS ? `${n.body.slice(0, NEIGHBOUR_EXCERPT_CHARS)}...` : n.body,
|
|
150
|
+
}));
|
|
151
|
+
}
|
|
152
|
+
catch {
|
|
153
|
+
return [];
|
|
154
|
+
}
|
|
155
|
+
finally {
|
|
156
|
+
if (db)
|
|
157
|
+
closeDatabase(db);
|
|
158
|
+
}
|
|
159
|
+
}
|
|
106
160
|
/**
|
|
107
161
|
* The gate for one run, holding its own read handle on `index.db` for the
|
|
108
162
|
* promotions that run emits; `undefined` when there is no bundle or no index
|
|
@@ -92,7 +92,7 @@ export const PAIR_JUDGE_JSON_SCHEMA = {
|
|
|
92
92
|
reason: { type: "string", maxLength: 400 },
|
|
93
93
|
},
|
|
94
94
|
};
|
|
95
|
-
const PAIR_CHECK_JSON_SCHEMA = {
|
|
95
|
+
export const PAIR_CHECK_JSON_SCHEMA = {
|
|
96
96
|
type: "object",
|
|
97
97
|
required: ["missing"],
|
|
98
98
|
additionalProperties: false,
|
|
@@ -503,7 +503,7 @@ function checkSection(label, side) {
|
|
|
503
503
|
].join("\n");
|
|
504
504
|
}
|
|
505
505
|
/**
|
|
506
|
-
* The second look a
|
|
506
|
+
* The second look a retirement gets before it may retire unattended: one call
|
|
507
507
|
* that asks only what the retired note holds that the kept one lacks. True
|
|
508
508
|
* only on a clean, empty answer (it caught 2 of 4 duplicates the judge got
|
|
509
509
|
* wrong, and held back none of 109 right ones).
|
|
@@ -658,17 +658,20 @@ async function judgeOne(ctx, candidate) {
|
|
|
658
658
|
}, ctx.opts.proposalsCtx);
|
|
659
659
|
ctx.retired.push(proposal.id);
|
|
660
660
|
ctx.perInitiatorProposed.add(candidate.initiator.ref);
|
|
661
|
-
// A
|
|
662
|
-
//
|
|
663
|
-
//
|
|
664
|
-
//
|
|
665
|
-
|
|
666
|
-
|
|
661
|
+
// A retirement the second look confirms loses nothing retires unattended:
|
|
662
|
+
// the triage drain accepts it under its usual applyMode. The retired side
|
|
663
|
+
// holds no claim of its own for any label (`decideRetirement` mints nothing
|
|
664
|
+
// else); a duplicate must also leave the kept side with none, while a
|
|
665
|
+
// subsumed or superseding successor holds more by definition. Replay
|
|
666
|
+
// precision 336/360 (duplicate 0.98, subsumed 0.92, supersedes 0.875,
|
|
667
|
+
// 2026-10-07); a duplicate alone was 109 of 111 safe on the owner's
|
|
668
|
+
// reviewed pairs (2026-10-04). Anything else waits for a person.
|
|
669
|
+
if ((verdict.relation !== "duplicate" || verdict.onlyInA.length + verdict.onlyInB.length === 0) &&
|
|
667
670
|
!continuityRisk &&
|
|
668
671
|
(await confirmNothingLost(ctx, retired, successor))) {
|
|
669
672
|
recordGateDecision(ctx.stashDir, proposal.id, {
|
|
670
673
|
outcome: "staged",
|
|
671
|
-
reason:
|
|
674
|
+
reason: verdict.relation,
|
|
672
675
|
gate: PAIR_PASS_GATE,
|
|
673
676
|
contentHash: proposalContentHash(proposal),
|
|
674
677
|
}, ctx.opts.proposalsCtx);
|