akm-cli 0.9.19-alpha.2 → 0.9.19

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -6,161 +6,170 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/).
6
6
 
7
7
  ## [Unreleased]
8
8
 
9
- ## [0.9.19-alpha.2] - 2026-09-30
10
-
11
- ### Changed
12
-
13
- - **Distill's grounding check vetoes only a score of 1; a 2 goes to a person.**
14
- 0.9.19-alpha.1 made a grounding score of 2 or less `quality_rejected`. A
15
- calibration of the lesson judge on a local llama.cpp model (34 cases, 5
16
- passes, 199 calls) scored all 4 off-subject lessons grounding 1 in nearly
17
- every pass (one scored 2 in 4 of 5 full passes and 1 otherwise), and its only
18
- false vetoes, 2 of 120 legitimate judgments, were one on-subject lesson that
19
- adds advice beyond its source and scored 2. Temperature 0 did not make the
20
- judge repeatable on that server either: scores moved by up to a point between
21
- passes. A grounding score of 1 is therefore still `quality_rejected` (an
22
- `improve_ledger` row and a `distill_invoked` event, no proposal, reason
23
- `Off-subject for its source (grounding 1/5): …`), and a 2 is now
24
- `review_needed`, a pending proposal for a person with the reason
25
- `Borderline on grounding (2/5), routed to review: …`, even when the mean of
26
- novelty and non-redundancy alone would pass it. A mean that alone rejects the
27
- lesson stays `quality_rejected`, and a grounding score of 3 to 5 is
28
- unchanged.
29
-
30
- ## [0.9.19-alpha.1] - 2026-09-30
9
+ ## [0.9.19] - 2026-09-30
31
10
 
32
11
  ### Added
33
12
 
34
13
  - **`akm proposal reopen <id...> [--reason <text>]` (#997).** A rejection was
35
14
  final: nothing undid it, and a rejected `consolidate-pair` retire proposal
36
15
  also kept the pair pass from ever proposing that retirement again while both
37
- documents were unchanged. Reopen moves rejected proposals back to `pending`,
38
- keeping the rejection (and the gate verdict that came with it) in the
39
- proposal's new `reviewHistory`, which `proposal show` prints; the verdict
40
- itself is cleared so the drain sees the proposal as undecided, except a
41
- `deferred` one (the quality gate's hand-off to a person), which stays. It is
42
- refused unless the proposal is `rejected` and `accept` would not refuse it as
43
- stale (an update's target unchanged; a create's target still absent; a retire
44
- proposal's successor present and both documents' body hashes as recorded),
45
- and a retire proposal is refused while another pending retire proposal
46
- involves either of its documents. Several ids are all-or-nothing. The pair pass
47
- follows the status: a reopened proposal is no longer a settled pair and,
48
- pending, is not minted twice. Its `improve_ledger` row is reset (a retire
49
- proposal's rejection row is dropped, any other goes back to `proposed`), the
50
- age that retention expiry and `--older-than` (bulk accept/reject, `drain`)
51
- see restarts at the reopen, so a scheduled sweep does not take a proposal a
52
- person just put back, and a `proposal_reopened` event is appended.
53
- `akm proposal reject`'s confirmation prompt no longer says a rejection
54
- cannot be undone.
16
+ documents were unchanged. Reopen moves rejected proposals back to `pending`
17
+ (a proposal that retention expiry archived is a rejected one too), keeping
18
+ the rejection, and the gate verdict that came with it, in the proposal's new
19
+ `reviewHistory`, which `proposal show` prints. The verdict itself is cleared
20
+ so the drain sees the proposal as undecided, except a `deferred` one (the
21
+ quality gate's hand-off to a person), which stays. It takes a full id or an
22
+ asset ref (an id prefix only matches pending proposals) and is refused
23
+ unless the proposal is `rejected` and `accept` would not refuse it as stale: an update's target
24
+ unchanged, a create's target still absent, a retire proposal's successor
25
+ present and both documents' body hashes as recorded. A retire proposal is
26
+ also refused while another pending retire proposal involves either of its
27
+ documents. Several ids are all-or-nothing. The pair pass follows the status:
28
+ a reopened proposal is no longer a settled pair and, pending, is not minted
29
+ twice. Its `improve_ledger` row is reset (a retire proposal's rejection row
30
+ is dropped, any other goes back to `proposed`), the age that retention expiry
31
+ and `--older-than` (bulk accept/reject, `drain`) see restarts at the reopen,
32
+ so a scheduled sweep does not take a proposal a person just put back, and a
33
+ `proposal_reopened` event is appended. `akm proposal reject`'s confirmation
34
+ prompt no longer says a rejection cannot be undone.
55
35
 
56
- ### Fixed
36
+ ### Changed
57
37
 
58
- - **A tool failure recorded with `akm feedback` no longer becomes a lesson
59
- about the error or a `TODO` placeholder in a memory (#999).** Agents
60
- recorded `akm show` failing on a memory with a `.derived.md` child (fixed in
61
- 0.9.17) as negative feedback, and distill and reflect read it as evidence
62
- about the memory's content. On one bundle, 9 such events on 8 memories
63
- produced 4 distill lessons about "duplicate physical owners" for memories on
64
- unrelated subjects (one auto-accepted and live), and a reflect proposal,
65
- also auto-accepted, that added a `TODO: verify physical owner` section to a
66
- memory, on which a fifth lesson was then built. The shipped hints had told
67
- agents to record `--negative` "when it fails"; they, and the help for
68
- `akm feedback --reason`, now say a failed akm command is not feedback on the
69
- asset. Reflect's feedback caveat no longer offers a `TODO: verify …`
70
- placeholder: when feedback asks for information the asset lacks, it says
71
- only to leave the section unchanged. The distill quality judge now also
72
- scores **grounding**, whether the lesson is about what its source is about
73
- (1–2 only for a different subject; a lesson that corrects its source from
74
- feedback is not off-subject), and a grounding score of 2 or less is
75
- `quality_rejected` (an `improve_ledger` row and a `distill_invoked` event, no
76
- proposal) whatever the mean of novelty and non-redundancy is. Such a lesson
77
- reads as novel and non-redundant, so it used to pass or, in the review band,
78
- be minted as a pending `review_needed` proposal. The judge also reads the
79
- same slice of the source the lesson was generated from (its body without
80
- frontmatter, first 3000 characters) instead of the raw file's first 2000. A
81
- lesson that contradicts its source still reaches a human through the
82
- optional fidelity check (`processes.distill.fidelityCheck.enabled`, off by
83
- default), and every other `review_needed` reason is unchanged. `TODO:`
84
- lines already in a memory are not removed.
85
- - **Consolidation stops re-proposing memories that `knowledge/` already
86
- covers (#998).** The promote pass copied a memory into a new `knowledge/`
87
- proposal with no notion of what `knowledge/` already held: the model never
88
- sees it, the mint-time checks only caught the same slug or a byte-identical
89
- body, and an accepted promotion's memory was eligible again at once (a
90
- rejected one after 7 days). On one bundle 88% of a run's proposals came from
91
- memories promoted before, one of them 14 times, and 215 of 224 rejections
92
- read "covered by an existing knowledge doc". Two changes, no new setting:
93
- before queuing a promotion, consolidate now compares the memory with the 20
94
- `knowledge/` docs in its bundle nearest to it by stored vector (the lookup
95
- the pair pass uses) and skips it, with skip reason
96
- `dedup_covered_by_knowledge`, when one of them holds at least half of the
97
- memory's distinct 5-word shingles (measured against every knowledge doc,
98
- that share was at least 0.5 for 122 of the 224 rejected proposals and for
99
- none of the 103 accepted ones; a covering doc past the 20 nearest goes
100
- unseen); and a memory whose promotion was accepted or rejected is offered
101
- again only when its body changes, the same content-driven rule the pair pass
102
- uses, instead of at once or after 7 days. A promotion decided by an older
103
- release recorded no body hash and keeps its old windows. With no stored
104
- vector (semantic search off) the coverage check does nothing.
105
- - **`akm improve` no longer files one bundle's assets into another (#1000).**
106
- Candidate selection admitted assets from every writable bundle, but every
107
- proposal is filed in the run's write target and reflect reads each asset
108
- from the bundle that owns it. An asset owned by another bundle therefore
109
- came back as a `create` fork in the write target, or as an `update` of the
110
- write target's own copy built from the other bundle's copy. A run now plans
111
- only the bundle it writes to (`--bundle`, else `defaultWriteTarget`, else
112
- the working bundle), and a bare ref scope (`akm improve skills/x`) resolves
113
- inside that bundle. Distill's memory-to-knowledge promotion likewise merges
114
- only with a doc that already exists in the write target. As a second line of
115
- defence, `createProposal` refuses a rewrite whose `itemRef` names an asset
116
- owned by a different configured bundle than the queue's. **Narrowed
117
- behaviour:** a run no longer picks up assets from your other writable
118
- bundles (it used to read them and queue the result in its own write target),
119
- so a scheduled `akm improve` now covers only its write target: add one
120
- `akm improve --bundle <name>` run per other bundle you want improved.
38
+ - **`akm improve` improves only the bundle it writes to (#1000).** Candidate
39
+ selection admitted assets from every writable bundle, but every proposal is
40
+ filed in the run's write target and reflect reads each asset from the bundle
41
+ that owns it. An asset owned by another bundle therefore came back as a
42
+ `create` fork in the write target, or as an `update` of the write target's
43
+ own copy built from the other bundle's copy. A run now plans only the bundle
44
+ it writes to (`--bundle`, else `defaultWriteTarget`, else the working bundle:
45
+ `AKM_BUNDLE_DIR` when set, otherwise `defaultBundle`), and a bare ref scope
46
+ (`akm improve skills/x`) resolves inside that bundle. Distill's
47
+ memory-to-knowledge promotion likewise merges only with a doc that already
48
+ exists in the write target. As a second line of defence, `createProposal`
49
+ refuses a rewrite whose `itemRef` names an asset owned by a different
50
+ configured bundle than the queue's. **Narrowed behaviour:** a run no longer
51
+ picks up assets from your other writable bundles (it used to read them and
52
+ queue the result in its own write target), so a scheduled `akm improve`, the
53
+ shipped improve task templates included, now covers only its write target:
54
+ add one `akm improve --bundle <name>` run per other bundle you want improved.
121
55
  `akm improve <ref>` for an asset that lives only in another bundle now fails
122
56
  with a not-found error whose hint names the remedy (`--bundle team`, or
123
57
  `akm improve team//skills/x`). Proposals the old behaviour already queued
124
58
  stay in the queue; review them with `akm proposal list`. `--bundle`'s help
125
59
  text now says it selects the bundle a run improves and writes to.
126
- - **`akm improve --dry-run`/`--plan` previews the bundle a live run improves.**
127
- With no `--bundle` and no `defaultWriteTarget`, a live run starts from
128
- `AKM_BUNDLE_DIR` before `defaultBundle`, but a dry run read `defaultBundle`
129
- only, so the two could plan different bundles. The preview now resolves the
130
- working bundle the same way.
60
+ - **Consolidation stops re-proposing memories that `knowledge/` already covers
61
+ (#998).** The promote pass copied a memory into a new `knowledge/` proposal
62
+ with no notion of what `knowledge/` already held: the model never sees it,
63
+ the mint-time checks only caught the same slug or an identical body, and an
64
+ accepted promotion's memory was eligible again at once (a rejected one after
65
+ 7 days). On one bundle 88% of a run's proposals came from memories promoted
66
+ before, one of them 14 times, and 215 of 224 rejections said the memory
67
+ duplicated, was covered by, or overlapped an existing knowledge doc. Two
68
+ changes, no new setting: before queuing a promotion, consolidate now compares
69
+ the memory with the 20 `knowledge/` docs in its bundle nearest to it by
70
+ stored vector (the lookup the pair pass uses) and skips it, with skip reason
71
+ `dedup_covered_by_knowledge` and a warning naming the covering doc, when one
72
+ of them holds at least half of the memory's distinct 5-word shingles
73
+ (measured against every knowledge doc, that share was at least 0.5 for 122 of
74
+ the 224 rejected proposals and for none of the 103 accepted ones; a covering
75
+ doc past the 20 nearest goes unseen); and a memory whose promotion was
76
+ accepted or rejected is offered again only when its body (not its
77
+ frontmatter) changes, the same content-driven rule the pair pass uses,
78
+ instead of at once or after 7 days. A promotion decided by an older release
79
+ recorded no body hash and keeps its old windows. With no stored vector
80
+ (semantic search off) the coverage check does nothing.
81
+ - **Distill's quality judge also scores grounding: a 1 is vetoed, a 2 goes to
82
+ review (#999).** A lesson about a different subject than its source reads as
83
+ novel and non-redundant, so it used to pass or, in the review band, be minted
84
+ as a pending `review_needed` proposal. The judge, which gates
85
+ memory-to-knowledge promotions the same way, now also scores **grounding**,
86
+ whether the lesson is about what its source is about: 1 or 2 only for a
87
+ different subject, 3 for a lesson on the source's subject that goes beyond or
88
+ corrects it (distill folds feedback into the lesson, and the judge is not
89
+ shown that feedback), 4 or 5 when the source supports it. Grounding is left
90
+ out of the mean of novelty and non-redundancy. A grounding score of 1 is
91
+ `quality_rejected` whatever the mean is (an `improve_ledger` row and a
92
+ `distill_invoked` event, no proposal, reason
93
+ `Off-subject for its source (grounding 1/5): …`). A 2 is borderline and is
94
+ `review_needed`: a pending proposal for a person, with the reason
95
+ `Borderline on grounding (2/5), routed to review: …`, even when the mean
96
+ alone would pass it; a mean that alone rejects the lesson stays
97
+ `quality_rejected`. A 3 to 5 changes nothing. Only a 1 is a veto because the
98
+ judge is not repeatable enough to discard a 2 unseen: on a calibration on a
99
+ local llama.cpp model (34 cases, 5 passes), the only legitimate lesson it
100
+ scored 2 or lower (2 of 120 judgments) was an on-subject lesson that adds
101
+ advice beyond its source, and temperature 0 did not make the scores
102
+ repeatable there: they moved by up to a point between passes. The judge also
103
+ reads the same slice of the source the lesson was generated from (its body
104
+ without frontmatter, first 3000 characters) instead of the raw file's first
105
+ 2000. A lesson that contradicts its source still reaches a human through the
106
+ optional fidelity check (`processes.distill.fidelityCheck.enabled`, off by
107
+ default), and every other `review_needed` reason is unchanged.
108
+
109
+ ### Fixed
110
+
131
111
  - **`akm proposal diff` shows a retire proposal as a retirement (#997).** It
132
112
  rendered the retired file as replaced by one blank line (`----`, a lone `+`,
133
113
  then every other line as a removal, under an `(update: <ref>)` header) and
134
114
  said nothing about the retirement, so one reviewer rejected all 65 of a
135
115
  bundle's `consolidate-pair` proposals as "would destroy content". The diff
136
- now lists only the removed lines, under a `(retire: <retired> -> <successor>)`
137
- header and `+++ /dev/null (retired: archived; successor <ref>)`, and its JSON
138
- result gains `op: "delete"`, a `retirement` block under the keys `proposal
139
- show` uses (`retiredRef`, `successorRef`, `judgeLabel`, `judgeReason`,
140
- `cosine`, and `continuityRisk` when the pair was flagged, which the text
141
- output prints as well) and a `note` that accepting archives the file under
142
- `.akm/memory-cleanup/archive/` and `akm proposal revert` restores it
143
- byte-exactly. The new fields are additive and appear on retire proposals
144
- only. `akm proposal show --detail full` also stops ending a retire proposal
145
- with a bare `payload:` heading over nothing.
116
+ now lists only the removed lines, under a
117
+ `(retire: <retired> -> <successor>)` header and
118
+ `+++ /dev/null (retired: archived; successor <ref>)`, and its JSON result
119
+ gains `op: "delete"`, a `retirement` block under the keys `proposal show`
120
+ uses (`retiredRef`, `successorRef`, `judgeLabel`, `judgeReason`, `cosine`,
121
+ and `continuityRisk` when the pair was flagged) and a `note` that accepting
122
+ archives the file under `.akm/memory-cleanup/archive/` and
123
+ `akm proposal revert` restores it byte-exactly; the text output prints the
124
+ same verdict and note above the removed lines. The new fields are additive
125
+ and appear on retire proposals only. `akm proposal show --detail full` also
126
+ stops ending a retire proposal with a bare `payload:` heading over nothing.
127
+ Rejections made over the old rendering can be taken back with
128
+ `akm proposal reopen`.
129
+ - **A tool failure recorded with `akm feedback` no longer becomes a lesson
130
+ about the error or a `TODO` placeholder in a memory (#999).** Agents recorded
131
+ `akm show` failing on a memory with a `.derived.md` child (fixed in 0.9.17)
132
+ as negative feedback, and distill and reflect read it as evidence about the
133
+ memory's content. On one bundle, 9 such events on 8 memories produced 4
134
+ distill lessons about "duplicate physical owners" for memories on unrelated
135
+ subjects (one auto-accepted and live), and a reflect proposal, also
136
+ auto-accepted, that added a `TODO: verify physical owner` section to a
137
+ memory, on which a fifth lesson was then built. The shipped hints had told
138
+ agents to record `--negative` "when it fails"; they, and the help for
139
+ `akm feedback --reason`, now say a failed akm command is not feedback on the
140
+ asset (if you pasted `akm help agents` into an AGENTS.md or a system prompt,
141
+ regenerate it). Reflect's feedback caveat no longer offers a `TODO: verify …`
142
+ placeholder: when feedback asks for information the asset lacks, it says only
143
+ to leave the section unchanged. Distill's judge now also checks that a lesson
144
+ is about its source (see Changed). `TODO:` lines already in a memory are not
145
+ removed.
146
+ - **`akm improve --dry-run`/`--plan` previews the bundle a live run improves.**
147
+ With no `--bundle` and no `defaultWriteTarget`, a live run starts from
148
+ `AKM_BUNDLE_DIR` before `defaultBundle`, but a dry run read `defaultBundle`
149
+ only, so the two could plan different bundles. The preview now resolves the
150
+ working bundle the same way.
146
151
  - **Errors name the flag the command takes, not the retired `--target`
147
152
  (`improve`, `remember`, `clone`, `task`, `proposal --queue`).** A `--bundle`
148
153
  that names no configured bundle, names a read-only one, or differs from a
149
- bundle-qualified ref (`akm improve team//skills/x --bundle stash`) failed with
150
- "--target must reference a source name", "or pass --target to a different
151
- source" or "conflicts with --target". `akm improve`, `remember`, `clone` and
152
- every `task` verb reject `--target` (renamed `--bundle` in 0.9), and
153
- `proposal --queue` takes `--queue`, so each message sent the user to a flag
154
- that does not work. They now name the flag the command takes, and the
154
+ bundle-qualified ref (`akm improve team//skills/x --bundle stash`) failed
155
+ with "--target must reference a source name", "or pass --target to a
156
+ different source" or "conflicts with --target". `akm improve`, `remember`,
157
+ `clone` and every `task` verb reject `--target` (renamed `--bundle` in 0.9),
158
+ and `proposal --queue` takes `--queue`, so each message sent the user to a
159
+ flag that does not work. They now name the flag the command takes, and the
155
160
  remedy in `akm remember --supersedes` says "re-run with --bundle" instead of
156
- "--target". A bundle that came from a ref inside a task (`ghost//workflows/x`)
157
- is described as such rather than blamed on a flag. The commands that really
158
- take `--target` (`import`, `env`, `secret`, `proposal accept`) are unchanged.
159
- Separately, the help for `akm improve --skip-if-locked` said a lock collision
160
- without the flag exits 78. It exits 75 (`IMPROVE_LOCK_HELD`), as the CLI
161
- reference says. The `--limit` help says "highest salience first" (it said
162
- utility), and the `task add --command` help example uses
163
- `--strategy reflect-distill` (the `frequent` strategy no longer exists).
161
+ "--target". A bundle that came from a ref inside a task
162
+ (`ghost//workflows/x`) is described as such rather than blamed on a flag. The
163
+ commands that really take `--target` (`import`, `env`, `secret`, and
164
+ `proposal accept`/`diff`/`revert`) are unchanged.
165
+ - **Stale help and warning text corrected.** The help for
166
+ `akm improve --skip-if-locked` said a lock collision without the flag exits
167
+ 78. It exits 75 (`IMPROVE_LOCK_HELD`), as the CLI reference says. The
168
+ `--limit` help says "highest salience first" (it said utility), the
169
+ `task add --command` help example uses `--strategy reflect-distill` (the
170
+ `frequent` strategy no longer exists), and improve's warning about missing
171
+ `show` events now speaks of the retrieval scope (it named a zero-feedback
172
+ fallback).
164
173
 
165
174
  ## [0.9.18] - 2026-09-29
166
175
 
@@ -101,6 +101,13 @@ those old windows. Expect a shorter promotion queue, with the skipped
101
101
  memories showing up as skip reasons and warnings in the run's result; to have
102
102
  a memory considered again, edit its text.
103
103
 
104
+ One consequence of those old windows: if you deleted the knowledge copies an
105
+ older release promoted, to keep only the memories, those memories have
106
+ neither a hold nor a copy for the coverage check to find, so the first 0.9.19
107
+ run proposes each of them once more, word for word. Reject them (akm proposal
108
+ reject <id> --yes --reason "..."): 0.9.19 records the rejection with the
109
+ memory's body hash, and they stay held until their text changes.
110
+
104
111
  Three changes keep a tool failure recorded as feedback from becoming a lesson
105
112
  about the error or a TODO placeholder in a memory. Reflect's feedback caveat no
106
113
  longer offers a TODO: verify placeholder: when feedback asks for something the
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "akm-cli",
3
- "version": "0.9.19-alpha.2",
3
+ "version": "0.9.19",
4
4
  "type": "module",
5
5
  "description": "akm (Agent Knowledge Manager) — a portable, local-first capability library for AI agents. Discover, load, share, and improve reusable skills, scripts, workflows, and knowledge across any shell-capable coding agent, including Claude Code, OpenCode, and Cursor.",
6
6
  "keywords": [