akm-cli 0.9.19-alpha.1 → 0.9.19
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +142 -112
- package/dist/commands/improve/stage.js +31 -8
- package/docs/migration/release-notes/0.9.19.md +14 -4
- package/docs/reference/cli.md +3 -2
- package/package.json +1 -1
package/CHANGELOG.md
CHANGED
|
@@ -6,140 +6,170 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/).
|
|
|
6
6
|
|
|
7
7
|
## [Unreleased]
|
|
8
8
|
|
|
9
|
-
## [0.9.19
|
|
9
|
+
## [0.9.19] - 2026-09-30
|
|
10
10
|
|
|
11
11
|
### Added
|
|
12
12
|
|
|
13
13
|
- **`akm proposal reopen <id...> [--reason <text>]` (#997).** A rejection was
|
|
14
14
|
final: nothing undid it, and a rejected `consolidate-pair` retire proposal
|
|
15
15
|
also kept the pair pass from ever proposing that retirement again while both
|
|
16
|
-
documents were unchanged. Reopen moves rejected proposals back to `pending
|
|
17
|
-
|
|
18
|
-
|
|
19
|
-
|
|
20
|
-
|
|
21
|
-
|
|
22
|
-
|
|
23
|
-
proposal
|
|
24
|
-
|
|
25
|
-
|
|
26
|
-
|
|
27
|
-
|
|
28
|
-
|
|
29
|
-
|
|
30
|
-
|
|
31
|
-
|
|
32
|
-
|
|
33
|
-
|
|
16
|
+
documents were unchanged. Reopen moves rejected proposals back to `pending`
|
|
17
|
+
(a proposal that retention expiry archived is a rejected one too), keeping
|
|
18
|
+
the rejection, and the gate verdict that came with it, in the proposal's new
|
|
19
|
+
`reviewHistory`, which `proposal show` prints. The verdict itself is cleared
|
|
20
|
+
so the drain sees the proposal as undecided, except a `deferred` one (the
|
|
21
|
+
quality gate's hand-off to a person), which stays. It takes a full id or an
|
|
22
|
+
asset ref (an id prefix only matches pending proposals) and is refused
|
|
23
|
+
unless the proposal is `rejected` and `accept` would not refuse it as stale: an update's target
|
|
24
|
+
unchanged, a create's target still absent, a retire proposal's successor
|
|
25
|
+
present and both documents' body hashes as recorded. A retire proposal is
|
|
26
|
+
also refused while another pending retire proposal involves either of its
|
|
27
|
+
documents. Several ids are all-or-nothing. The pair pass follows the status:
|
|
28
|
+
a reopened proposal is no longer a settled pair and, pending, is not minted
|
|
29
|
+
twice. Its `improve_ledger` row is reset (a retire proposal's rejection row
|
|
30
|
+
is dropped, any other goes back to `proposed`), the age that retention expiry
|
|
31
|
+
and `--older-than` (bulk accept/reject, `drain`) see restarts at the reopen,
|
|
32
|
+
so a scheduled sweep does not take a proposal a person just put back, and a
|
|
33
|
+
`proposal_reopened` event is appended. `akm proposal reject`'s confirmation
|
|
34
|
+
prompt no longer says a rejection cannot be undone.
|
|
34
35
|
|
|
35
|
-
###
|
|
36
|
+
### Changed
|
|
36
37
|
|
|
37
|
-
-
|
|
38
|
-
|
|
39
|
-
|
|
40
|
-
|
|
41
|
-
|
|
42
|
-
|
|
43
|
-
|
|
44
|
-
|
|
45
|
-
|
|
46
|
-
|
|
47
|
-
|
|
48
|
-
|
|
49
|
-
|
|
50
|
-
|
|
51
|
-
|
|
52
|
-
|
|
53
|
-
|
|
54
|
-
`quality_rejected` (an `improve_ledger` row and a `distill_invoked` event, no
|
|
55
|
-
proposal) whatever the mean of novelty and non-redundancy is. Such a lesson
|
|
56
|
-
reads as novel and non-redundant, so it used to pass or, in the review band,
|
|
57
|
-
be minted as a pending `review_needed` proposal. The judge also reads the
|
|
58
|
-
same slice of the source the lesson was generated from (its body without
|
|
59
|
-
frontmatter, first 3000 characters) instead of the raw file's first 2000. A
|
|
60
|
-
lesson that contradicts its source still reaches a human through the
|
|
61
|
-
optional fidelity check (`processes.distill.fidelityCheck.enabled`, off by
|
|
62
|
-
default), and every other `review_needed` reason is unchanged. `TODO:`
|
|
63
|
-
lines already in a memory are not removed.
|
|
64
|
-
- **Consolidation stops re-proposing memories that `knowledge/` already
|
|
65
|
-
covers (#998).** The promote pass copied a memory into a new `knowledge/`
|
|
66
|
-
proposal with no notion of what `knowledge/` already held: the model never
|
|
67
|
-
sees it, the mint-time checks only caught the same slug or a byte-identical
|
|
68
|
-
body, and an accepted promotion's memory was eligible again at once (a
|
|
69
|
-
rejected one after 7 days). On one bundle 88% of a run's proposals came from
|
|
70
|
-
memories promoted before, one of them 14 times, and 215 of 224 rejections
|
|
71
|
-
read "covered by an existing knowledge doc". Two changes, no new setting:
|
|
72
|
-
before queuing a promotion, consolidate now compares the memory with the 20
|
|
73
|
-
`knowledge/` docs in its bundle nearest to it by stored vector (the lookup
|
|
74
|
-
the pair pass uses) and skips it, with skip reason
|
|
75
|
-
`dedup_covered_by_knowledge`, when one of them holds at least half of the
|
|
76
|
-
memory's distinct 5-word shingles (measured against every knowledge doc,
|
|
77
|
-
that share was at least 0.5 for 122 of the 224 rejected proposals and for
|
|
78
|
-
none of the 103 accepted ones; a covering doc past the 20 nearest goes
|
|
79
|
-
unseen); and a memory whose promotion was accepted or rejected is offered
|
|
80
|
-
again only when its body changes, the same content-driven rule the pair pass
|
|
81
|
-
uses, instead of at once or after 7 days. A promotion decided by an older
|
|
82
|
-
release recorded no body hash and keeps its old windows. With no stored
|
|
83
|
-
vector (semantic search off) the coverage check does nothing.
|
|
84
|
-
- **`akm improve` no longer files one bundle's assets into another (#1000).**
|
|
85
|
-
Candidate selection admitted assets from every writable bundle, but every
|
|
86
|
-
proposal is filed in the run's write target and reflect reads each asset
|
|
87
|
-
from the bundle that owns it. An asset owned by another bundle therefore
|
|
88
|
-
came back as a `create` fork in the write target, or as an `update` of the
|
|
89
|
-
write target's own copy built from the other bundle's copy. A run now plans
|
|
90
|
-
only the bundle it writes to (`--bundle`, else `defaultWriteTarget`, else
|
|
91
|
-
the working bundle), and a bare ref scope (`akm improve skills/x`) resolves
|
|
92
|
-
inside that bundle. Distill's memory-to-knowledge promotion likewise merges
|
|
93
|
-
only with a doc that already exists in the write target. As a second line of
|
|
94
|
-
defence, `createProposal` refuses a rewrite whose `itemRef` names an asset
|
|
95
|
-
owned by a different configured bundle than the queue's. **Narrowed
|
|
96
|
-
behaviour:** a run no longer picks up assets from your other writable
|
|
97
|
-
bundles (it used to read them and queue the result in its own write target),
|
|
98
|
-
so a scheduled `akm improve` now covers only its write target: add one
|
|
99
|
-
`akm improve --bundle <name>` run per other bundle you want improved.
|
|
38
|
+
- **`akm improve` improves only the bundle it writes to (#1000).** Candidate
|
|
39
|
+
selection admitted assets from every writable bundle, but every proposal is
|
|
40
|
+
filed in the run's write target and reflect reads each asset from the bundle
|
|
41
|
+
that owns it. An asset owned by another bundle therefore came back as a
|
|
42
|
+
`create` fork in the write target, or as an `update` of the write target's
|
|
43
|
+
own copy built from the other bundle's copy. A run now plans only the bundle
|
|
44
|
+
it writes to (`--bundle`, else `defaultWriteTarget`, else the working bundle:
|
|
45
|
+
`AKM_BUNDLE_DIR` when set, otherwise `defaultBundle`), and a bare ref scope
|
|
46
|
+
(`akm improve skills/x`) resolves inside that bundle. Distill's
|
|
47
|
+
memory-to-knowledge promotion likewise merges only with a doc that already
|
|
48
|
+
exists in the write target. As a second line of defence, `createProposal`
|
|
49
|
+
refuses a rewrite whose `itemRef` names an asset owned by a different
|
|
50
|
+
configured bundle than the queue's. **Narrowed behaviour:** a run no longer
|
|
51
|
+
picks up assets from your other writable bundles (it used to read them and
|
|
52
|
+
queue the result in its own write target), so a scheduled `akm improve`, the
|
|
53
|
+
shipped improve task templates included, now covers only its write target:
|
|
54
|
+
add one `akm improve --bundle <name>` run per other bundle you want improved.
|
|
100
55
|
`akm improve <ref>` for an asset that lives only in another bundle now fails
|
|
101
56
|
with a not-found error whose hint names the remedy (`--bundle team`, or
|
|
102
57
|
`akm improve team//skills/x`). Proposals the old behaviour already queued
|
|
103
58
|
stay in the queue; review them with `akm proposal list`. `--bundle`'s help
|
|
104
59
|
text now says it selects the bundle a run improves and writes to.
|
|
105
|
-
-
|
|
106
|
-
|
|
107
|
-
|
|
108
|
-
only
|
|
109
|
-
|
|
60
|
+
- **Consolidation stops re-proposing memories that `knowledge/` already covers
|
|
61
|
+
(#998).** The promote pass copied a memory into a new `knowledge/` proposal
|
|
62
|
+
with no notion of what `knowledge/` already held: the model never sees it,
|
|
63
|
+
the mint-time checks only caught the same slug or an identical body, and an
|
|
64
|
+
accepted promotion's memory was eligible again at once (a rejected one after
|
|
65
|
+
7 days). On one bundle 88% of a run's proposals came from memories promoted
|
|
66
|
+
before, one of them 14 times, and 215 of 224 rejections said the memory
|
|
67
|
+
duplicated, was covered by, or overlapped an existing knowledge doc. Two
|
|
68
|
+
changes, no new setting: before queuing a promotion, consolidate now compares
|
|
69
|
+
the memory with the 20 `knowledge/` docs in its bundle nearest to it by
|
|
70
|
+
stored vector (the lookup the pair pass uses) and skips it, with skip reason
|
|
71
|
+
`dedup_covered_by_knowledge` and a warning naming the covering doc, when one
|
|
72
|
+
of them holds at least half of the memory's distinct 5-word shingles
|
|
73
|
+
(measured against every knowledge doc, that share was at least 0.5 for 122 of
|
|
74
|
+
the 224 rejected proposals and for none of the 103 accepted ones; a covering
|
|
75
|
+
doc past the 20 nearest goes unseen); and a memory whose promotion was
|
|
76
|
+
accepted or rejected is offered again only when its body (not its
|
|
77
|
+
frontmatter) changes, the same content-driven rule the pair pass uses,
|
|
78
|
+
instead of at once or after 7 days. A promotion decided by an older release
|
|
79
|
+
recorded no body hash and keeps its old windows. With no stored vector
|
|
80
|
+
(semantic search off) the coverage check does nothing.
|
|
81
|
+
- **Distill's quality judge also scores grounding: a 1 is vetoed, a 2 goes to
|
|
82
|
+
review (#999).** A lesson about a different subject than its source reads as
|
|
83
|
+
novel and non-redundant, so it used to pass or, in the review band, be minted
|
|
84
|
+
as a pending `review_needed` proposal. The judge, which gates
|
|
85
|
+
memory-to-knowledge promotions the same way, now also scores **grounding**,
|
|
86
|
+
whether the lesson is about what its source is about: 1 or 2 only for a
|
|
87
|
+
different subject, 3 for a lesson on the source's subject that goes beyond or
|
|
88
|
+
corrects it (distill folds feedback into the lesson, and the judge is not
|
|
89
|
+
shown that feedback), 4 or 5 when the source supports it. Grounding is left
|
|
90
|
+
out of the mean of novelty and non-redundancy. A grounding score of 1 is
|
|
91
|
+
`quality_rejected` whatever the mean is (an `improve_ledger` row and a
|
|
92
|
+
`distill_invoked` event, no proposal, reason
|
|
93
|
+
`Off-subject for its source (grounding 1/5): …`). A 2 is borderline and is
|
|
94
|
+
`review_needed`: a pending proposal for a person, with the reason
|
|
95
|
+
`Borderline on grounding (2/5), routed to review: …`, even when the mean
|
|
96
|
+
alone would pass it; a mean that alone rejects the lesson stays
|
|
97
|
+
`quality_rejected`. A 3 to 5 changes nothing. Only a 1 is a veto because the
|
|
98
|
+
judge is not repeatable enough to discard a 2 unseen: on a calibration on a
|
|
99
|
+
local llama.cpp model (34 cases, 5 passes), the only legitimate lesson it
|
|
100
|
+
scored 2 or lower (2 of 120 judgments) was an on-subject lesson that adds
|
|
101
|
+
advice beyond its source, and temperature 0 did not make the scores
|
|
102
|
+
repeatable there: they moved by up to a point between passes. The judge also
|
|
103
|
+
reads the same slice of the source the lesson was generated from (its body
|
|
104
|
+
without frontmatter, first 3000 characters) instead of the raw file's first
|
|
105
|
+
2000. A lesson that contradicts its source still reaches a human through the
|
|
106
|
+
optional fidelity check (`processes.distill.fidelityCheck.enabled`, off by
|
|
107
|
+
default), and every other `review_needed` reason is unchanged.
|
|
108
|
+
|
|
109
|
+
### Fixed
|
|
110
|
+
|
|
110
111
|
- **`akm proposal diff` shows a retire proposal as a retirement (#997).** It
|
|
111
112
|
rendered the retired file as replaced by one blank line (`----`, a lone `+`,
|
|
112
113
|
then every other line as a removal, under an `(update: <ref>)` header) and
|
|
113
114
|
said nothing about the retirement, so one reviewer rejected all 65 of a
|
|
114
115
|
bundle's `consolidate-pair` proposals as "would destroy content". The diff
|
|
115
|
-
now lists only the removed lines, under a
|
|
116
|
-
|
|
117
|
-
|
|
118
|
-
|
|
119
|
-
`
|
|
120
|
-
|
|
121
|
-
`.akm/memory-cleanup/archive/` and
|
|
122
|
-
|
|
123
|
-
|
|
124
|
-
|
|
116
|
+
now lists only the removed lines, under a
|
|
117
|
+
`(retire: <retired> -> <successor>)` header and
|
|
118
|
+
`+++ /dev/null (retired: archived; successor <ref>)`, and its JSON result
|
|
119
|
+
gains `op: "delete"`, a `retirement` block under the keys `proposal show`
|
|
120
|
+
uses (`retiredRef`, `successorRef`, `judgeLabel`, `judgeReason`, `cosine`,
|
|
121
|
+
and `continuityRisk` when the pair was flagged) and a `note` that accepting
|
|
122
|
+
archives the file under `.akm/memory-cleanup/archive/` and
|
|
123
|
+
`akm proposal revert` restores it byte-exactly; the text output prints the
|
|
124
|
+
same verdict and note above the removed lines. The new fields are additive
|
|
125
|
+
and appear on retire proposals only. `akm proposal show --detail full` also
|
|
126
|
+
stops ending a retire proposal with a bare `payload:` heading over nothing.
|
|
127
|
+
Rejections made over the old rendering can be taken back with
|
|
128
|
+
`akm proposal reopen`.
|
|
129
|
+
- **A tool failure recorded with `akm feedback` no longer becomes a lesson
|
|
130
|
+
about the error or a `TODO` placeholder in a memory (#999).** Agents recorded
|
|
131
|
+
`akm show` failing on a memory with a `.derived.md` child (fixed in 0.9.17)
|
|
132
|
+
as negative feedback, and distill and reflect read it as evidence about the
|
|
133
|
+
memory's content. On one bundle, 9 such events on 8 memories produced 4
|
|
134
|
+
distill lessons about "duplicate physical owners" for memories on unrelated
|
|
135
|
+
subjects (one auto-accepted and live), and a reflect proposal, also
|
|
136
|
+
auto-accepted, that added a `TODO: verify physical owner` section to a
|
|
137
|
+
memory, on which a fifth lesson was then built. The shipped hints had told
|
|
138
|
+
agents to record `--negative` "when it fails"; they, and the help for
|
|
139
|
+
`akm feedback --reason`, now say a failed akm command is not feedback on the
|
|
140
|
+
asset (if you pasted `akm help agents` into an AGENTS.md or a system prompt,
|
|
141
|
+
regenerate it). Reflect's feedback caveat no longer offers a `TODO: verify …`
|
|
142
|
+
placeholder: when feedback asks for information the asset lacks, it says only
|
|
143
|
+
to leave the section unchanged. Distill's judge now also checks that a lesson
|
|
144
|
+
is about its source (see Changed). `TODO:` lines already in a memory are not
|
|
145
|
+
removed.
|
|
146
|
+
- **`akm improve --dry-run`/`--plan` previews the bundle a live run improves.**
|
|
147
|
+
With no `--bundle` and no `defaultWriteTarget`, a live run starts from
|
|
148
|
+
`AKM_BUNDLE_DIR` before `defaultBundle`, but a dry run read `defaultBundle`
|
|
149
|
+
only, so the two could plan different bundles. The preview now resolves the
|
|
150
|
+
working bundle the same way.
|
|
125
151
|
- **Errors name the flag the command takes, not the retired `--target`
|
|
126
152
|
(`improve`, `remember`, `clone`, `task`, `proposal --queue`).** A `--bundle`
|
|
127
153
|
that names no configured bundle, names a read-only one, or differs from a
|
|
128
|
-
bundle-qualified ref (`akm improve team//skills/x --bundle stash`) failed
|
|
129
|
-
"--target must reference a source name", "or pass --target to a
|
|
130
|
-
source" or "conflicts with --target". `akm improve`, `remember`,
|
|
131
|
-
every `task` verb reject `--target` (renamed `--bundle` in 0.9),
|
|
132
|
-
`proposal --queue` takes `--queue`, so each message sent the user to a
|
|
133
|
-
that does not work. They now name the flag the command takes, and the
|
|
154
|
+
bundle-qualified ref (`akm improve team//skills/x --bundle stash`) failed
|
|
155
|
+
with "--target must reference a source name", "or pass --target to a
|
|
156
|
+
different source" or "conflicts with --target". `akm improve`, `remember`,
|
|
157
|
+
`clone` and every `task` verb reject `--target` (renamed `--bundle` in 0.9),
|
|
158
|
+
and `proposal --queue` takes `--queue`, so each message sent the user to a
|
|
159
|
+
flag that does not work. They now name the flag the command takes, and the
|
|
134
160
|
remedy in `akm remember --supersedes` says "re-run with --bundle" instead of
|
|
135
|
-
"--target". A bundle that came from a ref inside a task
|
|
136
|
-
is described as such rather than blamed on a flag. The
|
|
137
|
-
take `--target` (`import`, `env`, `secret`,
|
|
138
|
-
|
|
139
|
-
|
|
140
|
-
|
|
141
|
-
|
|
142
|
-
`--
|
|
161
|
+
"--target". A bundle that came from a ref inside a task
|
|
162
|
+
(`ghost//workflows/x`) is described as such rather than blamed on a flag. The
|
|
163
|
+
commands that really take `--target` (`import`, `env`, `secret`, and
|
|
164
|
+
`proposal accept`/`diff`/`revert`) are unchanged.
|
|
165
|
+
- **Stale help and warning text corrected.** The help for
|
|
166
|
+
`akm improve --skip-if-locked` said a lock collision without the flag exits
|
|
167
|
+
78. It exits 75 (`IMPROVE_LOCK_HELD`), as the CLI reference says. The
|
|
168
|
+
`--limit` help says "highest salience first" (it said utility), the
|
|
169
|
+
`task add --command` help example uses `--strategy reflect-distill` (the
|
|
170
|
+
`frequent` strategy no longer exists), and improve's warning about missing
|
|
171
|
+
`show` events now speaks of the retrieval scope (it named a zero-feedback
|
|
172
|
+
fallback).
|
|
143
173
|
|
|
144
174
|
## [0.9.18] - 2026-09-29
|
|
145
175
|
|
|
@@ -235,15 +235,20 @@ export function buildReflectJudgePrompt(candidateContent, sourceContent, feedbac
|
|
|
235
235
|
* `grounding` is scored with the other lesson criteria but left out of their
|
|
236
236
|
* mean: a lesson about a different subject than its source reads as novel and
|
|
237
237
|
* non-redundant, so the mean would pass it (or, in the review band, mint it as
|
|
238
|
-
* a pending proposal).
|
|
239
|
-
*
|
|
240
|
-
*
|
|
241
|
-
*
|
|
242
|
-
*
|
|
243
|
-
*
|
|
238
|
+
* a pending proposal). The rubric reserves 1-2 for a different subject. A score
|
|
239
|
+
* of {@link UNGROUNDED_MAX_SCORE} or less is a rejection whatever the mean says
|
|
240
|
+
* (#999). A higher score up to {@link BORDERLINE_GROUNDING_MAX_SCORE} is only
|
|
241
|
+
* borderline: a lesson on its source's subject that advises beyond it has scored
|
|
242
|
+
* 2, and a score can move a point between runs (see `runQualityJudge`), so it
|
|
243
|
+
* goes to a person unless the mean alone already rejects it. A lesson that goes
|
|
244
|
+
* beyond or corrects its source is on its subject: distill folds feedback into
|
|
245
|
+
* the lesson, and the judge is never shown it. A contradiction of the source is
|
|
246
|
+
* the optional fidelity check's to send to a human (`judgeAndQueue` in
|
|
247
|
+
* distill.ts), so the rubric must not pre-empt it.
|
|
244
248
|
*/
|
|
245
249
|
const GROUNDING_CRITERION = "grounding";
|
|
246
|
-
const UNGROUNDED_MAX_SCORE =
|
|
250
|
+
const UNGROUNDED_MAX_SCORE = 1;
|
|
251
|
+
const BORDERLINE_GROUNDING_MAX_SCORE = 2;
|
|
247
252
|
const LESSON_JUDGE_CRITERIA = ["novelty", "nonRedundancy", GROUNDING_CRITERION];
|
|
248
253
|
const REFLECT_JUDGE_CRITERIA = ["feedbackAlignment", "preservation", "quality"];
|
|
249
254
|
/**
|
|
@@ -296,7 +301,12 @@ function judgeResponseSchema(keys) {
|
|
|
296
301
|
* The quality judge. Fails closed: no runner, an unparseable verdict or a
|
|
297
302
|
* provider failure never passes content. Bands: >= 3.5 pass, 2.5-3.5 review,
|
|
298
303
|
* < 2.5 reject; a `grounding` score of {@link UNGROUNDED_MAX_SCORE} or less
|
|
299
|
-
* rejects whatever the mean is
|
|
304
|
+
* rejects whatever the mean is, and one of {@link BORDERLINE_GROUNDING_MAX_SCORE}
|
|
305
|
+
* routes a lesson the mean would pass to review (a mean that rejects stays a
|
|
306
|
+
* rejection). Temperature is set to 0, which reduces run-to-run variation but
|
|
307
|
+
* does not remove it: on some servers (llama.cpp batching, for one) the same
|
|
308
|
+
* request can score a point apart, so the routing rules are chosen with that
|
|
309
|
+
* margin in mind.
|
|
300
310
|
*/
|
|
301
311
|
async function runQualityJudge(feature, config, prompt, keys, chat, options) {
|
|
302
312
|
const resolved = !options.runnerSelectionFrozen && !options.llmRunner
|
|
@@ -339,6 +349,19 @@ async function runQualityJudge(feature, config, prompt, keys, chat, options) {
|
|
|
339
349
|
};
|
|
340
350
|
}
|
|
341
351
|
const verdict = score >= 3.5 ? { pass: true } : score >= 2.5 ? { pass: false, reviewNeeded: true } : { pass: false };
|
|
352
|
+
// Borderline grounding is a person's call even when the mean would pass; a mean that rejects stays rejected.
|
|
353
|
+
if (criteria &&
|
|
354
|
+
grounding !== undefined &&
|
|
355
|
+
grounding <= BORDERLINE_GROUNDING_MAX_SCORE &&
|
|
356
|
+
(verdict.pass || verdict.reviewNeeded)) {
|
|
357
|
+
return {
|
|
358
|
+
pass: false,
|
|
359
|
+
reviewNeeded: true,
|
|
360
|
+
score,
|
|
361
|
+
reason: `Borderline on grounding (${grounding}/5), routed to review: ${reason}`,
|
|
362
|
+
criteria,
|
|
363
|
+
};
|
|
364
|
+
}
|
|
342
365
|
return { ...verdict, score, reason, ...(criteria ? { criteria } : {}) };
|
|
343
366
|
}
|
|
344
367
|
/** Judge a proposed lesson (or knowledge promotion) against its source. */
|
|
@@ -101,16 +101,26 @@ those old windows. Expect a shorter promotion queue, with the skipped
|
|
|
101
101
|
memories showing up as skip reasons and warnings in the run's result; to have
|
|
102
102
|
a memory considered again, edit its text.
|
|
103
103
|
|
|
104
|
+
One consequence of those old windows: if you deleted the knowledge copies an
|
|
105
|
+
older release promoted, to keep only the memories, those memories have
|
|
106
|
+
neither a hold nor a copy for the coverage check to find, so the first 0.9.19
|
|
107
|
+
run proposes each of them once more, word for word. Reject them (akm proposal
|
|
108
|
+
reject <id> --yes --reason "..."): 0.9.19 records the rejection with the
|
|
109
|
+
memory's body hash, and they stay held until their text changes.
|
|
110
|
+
|
|
104
111
|
Three changes keep a tool failure recorded as feedback from becoming a lesson
|
|
105
112
|
about the error or a TODO placeholder in a memory. Reflect's feedback caveat no
|
|
106
113
|
longer offers a TODO: verify placeholder: when feedback asks for something the
|
|
107
114
|
asset lacks, reflect is told to leave the section unchanged. TODO lines earlier
|
|
108
115
|
runs already put in your assets stay until you remove them; grep -rniE
|
|
109
116
|
"TODO:? *verify" over the bundle finds them. Distill's quality judge now also
|
|
110
|
-
scores whether a lesson is about what its source is about
|
|
111
|
-
off-subject is dropped as quality_rejected (a ledger row and a
|
|
112
|
-
event, no proposal) instead of passing or waiting in the queue
|
|
113
|
-
review_needed
|
|
117
|
+
scores whether a lesson is about what its source is about. A lesson scored 1
|
|
118
|
+
(off-subject) is dropped as quality_rejected (a ledger row and a
|
|
119
|
+
distill_invoked event, no proposal) instead of passing or waiting in the queue
|
|
120
|
+
as review_needed. A lesson scored 2 is borderline: it is queued as
|
|
121
|
+
review_needed for you to decide, even when its novelty and non-redundancy
|
|
122
|
+
alone would pass it, unless those alone would reject it, which stays
|
|
123
|
+
quality_rejected. The judge also reads the same first 3000 characters of the
|
|
114
124
|
source body, without frontmatter, that the generator saw. The shipped agent
|
|
115
125
|
guidance changed with it: akm help agents now says to record feedback about an
|
|
116
126
|
asset's content, that it helped or turned out wrong, stale or unhelpful, and
|
package/docs/reference/cli.md
CHANGED
|
@@ -3045,8 +3045,9 @@ Write the reason about the asset's content. Reflect treats it as an unverified
|
|
|
3045
3045
|
report to investigate, not a fact to insert, and is told to leave the section
|
|
3046
3046
|
unchanged when the reason asks for information the asset lacks. Distill's
|
|
3047
3047
|
quality gate rejects a lesson that is off-subject for the asset it was
|
|
3048
|
-
distilled from
|
|
3049
|
-
says nothing about the asset, so it is
|
|
3048
|
+
distilled from and sends a borderline one to review. A command that failed
|
|
3049
|
+
(`akm show` erroring on the ref, say) says nothing about the asset, so it is
|
|
3050
|
+
not a reason to record against it.
|
|
3050
3051
|
|
|
3051
3052
|
### task
|
|
3052
3053
|
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "akm-cli",
|
|
3
|
-
"version": "0.9.19
|
|
3
|
+
"version": "0.9.19",
|
|
4
4
|
"type": "module",
|
|
5
5
|
"description": "akm (Agent Knowledge Manager) — a portable, local-first capability library for AI agents. Discover, load, share, and improve reusable skills, scripts, workflows, and knowledge across any shell-capable coding agent, including Claude Code, OpenCode, and Cursor.",
|
|
6
6
|
"keywords": [
|