akm-cli 0.9.23 → 0.9.24
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md
CHANGED
|
@@ -4,6 +4,51 @@ All notable changes to this project will be documented in this file.
|
|
|
4
4
|
|
|
5
5
|
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/).
|
|
6
6
|
|
|
7
|
+
## [0.9.24] - 2026-10-02
|
|
8
|
+
|
|
9
|
+
### Changed
|
|
10
|
+
|
|
11
|
+
- **Reflect never auto-accepts a revision that changes the body; it waits for
|
|
12
|
+
review.** On 396 labelled reflect edits, the judge-passed edits that changed
|
|
13
|
+
the body were good 12 times in 37, and those that changed only the
|
|
14
|
+
frontmatter 13 times in 13. A revision whose body differs from the asset's,
|
|
15
|
+
ignoring whitespace, or whose source could not be read, is now created
|
|
16
|
+
pending and deferred for review with the reason `body-edit` (gate
|
|
17
|
+
`reflect`). When the quality judge passed it, the judge's scores and reason
|
|
18
|
+
stay on its gate decision. With `processes.reflect.qualityGate` off, no
|
|
19
|
+
judge ran, so the decision carries none. Before, the triage drain accepted a
|
|
20
|
+
judge-passed body edit, and its judgment tier could accept one made with the
|
|
21
|
+
gate off. These body edits now appear in the review queue
|
|
22
|
+
(`akm proposal list`), and `akm proposal show <id>` gives the reason and any
|
|
23
|
+
judge scores. A judge-passed revision that leaves the body unchanged is
|
|
24
|
+
still staged and accepted by the drain. A frontmatter-only revision made
|
|
25
|
+
with the gate off is still left to the drain, and a revision the judge fails
|
|
26
|
+
is still refused.
|
|
27
|
+
|
|
28
|
+
### Fixed
|
|
29
|
+
|
|
30
|
+
- **Reflect sends a revision to review when its judge fails, instead of
|
|
31
|
+
rejecting it.** A quality judge that timed out, errored or returned a reply
|
|
32
|
+
that could not be parsed gave no verdict, but reflect treated that as a
|
|
33
|
+
rejection. It created no proposal, recorded a `quality_rejected` attempt that
|
|
34
|
+
kept the asset from being reflected again for 14 days, and reported
|
|
35
|
+
`quality gate rejected: score=-1` with a reason saying the revision had been
|
|
36
|
+
routed to review. Reflect now creates the proposal and leaves it pending for
|
|
37
|
+
a person, deferred by the quality gate with the reason `judge-error`, as
|
|
38
|
+
distill does when its judge fails. The triage drain leaves it alone. It has
|
|
39
|
+
not been through the retrieval regression check, which runs only on a
|
|
40
|
+
revision the judge passes. A revision the judge scores too low is still
|
|
41
|
+
refused.
|
|
42
|
+
- **The triage drain leaves every proposal a stage deferred for review to a
|
|
43
|
+
person.** It skipped a deferred proposal only when the quality gate had
|
|
44
|
+
deferred it. Reflect defers with its own gate, `reflect`: a revision whose
|
|
45
|
+
size the size guard flagged, one that echoed the truncation notice, one made
|
|
46
|
+
with no judge configured, and now a body edit. With
|
|
47
|
+
`processes.triage.judgment` enabled, the drain's judgment tier decided those
|
|
48
|
+
proposals, and under `applyMode: promote` it could accept them before anyone
|
|
49
|
+
saw them. The drain now skips every deferral it did not make itself, and
|
|
50
|
+
judges again, as before, the ones it did.
|
|
51
|
+
|
|
7
52
|
## [0.9.23] - 2026-10-01
|
|
8
53
|
|
|
9
54
|
### Added
|
|
@@ -932,8 +932,10 @@ const NOISE_SUBREASONS = {
|
|
|
932
932
|
};
|
|
933
933
|
/**
|
|
934
934
|
* Sanitize, drop a no-op/cosmetic (and optionally low-value) change, judge the
|
|
935
|
-
* exact content that would be persisted, then mint.
|
|
936
|
-
*
|
|
935
|
+
* exact content that would be persisted, then mint. A judge pass is staged only
|
|
936
|
+
* when the body is unchanged: a body edit the judge passes, or one made with the
|
|
937
|
+
* gate off, waits for review. Size-flagged or truncation-leaking content skips
|
|
938
|
+
* the judge and waits for review.
|
|
937
939
|
*/
|
|
938
940
|
async function finalizeReflectProposal(args) {
|
|
939
941
|
const { run, assetContent, result, judge, feedback } = args;
|
|
@@ -980,6 +982,7 @@ async function finalizeReflectProposal(args) {
|
|
|
980
982
|
return reflectFailure(run, result, "quality_rejected", message, false);
|
|
981
983
|
};
|
|
982
984
|
let verdict;
|
|
985
|
+
let judgeFailed = false;
|
|
983
986
|
if (judged) {
|
|
984
987
|
verdict = await runReflectQualityJudge(run.config, payload.content, assetContent ?? "", feedback, options.chat, {
|
|
985
988
|
runnerSelectionFrozen: true,
|
|
@@ -988,7 +991,9 @@ async function finalizeReflectProposal(args) {
|
|
|
988
991
|
...(options.signal ? { signal: options.signal } : {}),
|
|
989
992
|
onNotices: run.notices.add,
|
|
990
993
|
});
|
|
991
|
-
|
|
994
|
+
// A judge that timed out, errored or replied unparseably gave no verdict: a person reviews the revision.
|
|
995
|
+
judgeFailed = verdict.reviewNeeded === true && verdict.score === -1;
|
|
996
|
+
if (!verdict.pass && !judgeFailed) {
|
|
992
997
|
return refuse(verdict.reason, {
|
|
993
998
|
qualityScore: verdict.score,
|
|
994
999
|
qualityReason: verdict.reason,
|
|
@@ -997,7 +1002,7 @@ async function finalizeReflectProposal(args) {
|
|
|
997
1002
|
}
|
|
998
1003
|
}
|
|
999
1004
|
// #722: a rewrite of an existing asset must not grade lower on its own retrieval queries.
|
|
1000
|
-
if (
|
|
1005
|
+
if (verdict?.pass && judge.runner && assetContent !== undefined) {
|
|
1001
1006
|
const retrieval = await runRetrievalRegressionGate({
|
|
1002
1007
|
ref: payload.ref,
|
|
1003
1008
|
before: assetContent,
|
|
@@ -1026,9 +1031,18 @@ async function finalizeReflectProposal(args) {
|
|
|
1026
1031
|
};
|
|
1027
1032
|
const reviewReasons = [
|
|
1028
1033
|
...(judge.skippedNoJudge ? ["no-judge-configured"] : []),
|
|
1034
|
+
...(judgeFailed ? ["judge-error"] : []),
|
|
1029
1035
|
...(sanitized.sizeGuardRatio ? ["reflect-size-ratio"] : []),
|
|
1030
1036
|
...(sanitized.truncationMarkerLeaked ? ["reflect-truncation-leak"] : []),
|
|
1031
1037
|
];
|
|
1038
|
+
// A revision that changes the body is never auto-accepted: on labelled edits, the judge's
|
|
1039
|
+
// passes on body edits were good 12 times in 37, and on frontmatter-only edits 13 in 13.
|
|
1040
|
+
// One that nothing above holds for review (the judge passed it, or the gate is off) waits
|
|
1041
|
+
// for a person, as does a revision with no source to compare.
|
|
1042
|
+
const bodyOf = (content) => splitFrontmatter(content).body.replace(/\s+/g, " ").trim();
|
|
1043
|
+
const bodyEdit = reviewReasons.length === 0 && (assetContent === undefined || bodyOf(assetContent) !== bodyOf(payload.content));
|
|
1044
|
+
if (bodyEdit)
|
|
1045
|
+
reviewReasons.push("body-edit");
|
|
1032
1046
|
const proposal = mintProposal(run.stash, options.ctx, {
|
|
1033
1047
|
ref: payload.ref,
|
|
1034
1048
|
...(options.target ? { target: options.target } : {}),
|
|
@@ -1042,8 +1056,12 @@ async function finalizeReflectProposal(args) {
|
|
|
1042
1056
|
? {
|
|
1043
1057
|
review: {
|
|
1044
1058
|
reason: reviewReasons.join("+"),
|
|
1045
|
-
gate:
|
|
1059
|
+
// The quality gate's hand-off to a person, as distill's: the triage drain leaves it alone.
|
|
1060
|
+
gate: judgeFailed ? "quality-gate" : "reflect",
|
|
1046
1061
|
...(sanitized.sizeGuardRatio ? { measured: Math.round(sanitized.sizeGuardRatio.ratio * 100) } : {}),
|
|
1062
|
+
// The reviewer sees why the judge passed it (with the gate off, nothing judged it).
|
|
1063
|
+
...(bodyEdit && verdict?.criteria ? { scores: verdict.criteria } : {}),
|
|
1064
|
+
...(bodyEdit && verdict ? { judgeReason: verdict.reason } : {}),
|
|
1047
1065
|
},
|
|
1048
1066
|
}
|
|
1049
1067
|
: { judged: verdict });
|
|
@@ -1055,6 +1073,7 @@ async function finalizeReflectProposal(args) {
|
|
|
1055
1073
|
source: "reflect",
|
|
1056
1074
|
engine: run.engineName,
|
|
1057
1075
|
...(judge.skippedNoJudge ? { qualityGateSkippedNoJudge: true } : {}),
|
|
1076
|
+
...(judgeFailed ? { qualityReason: verdict?.reason } : {}),
|
|
1058
1077
|
...(sanitized.sizeGuardRatio
|
|
1059
1078
|
? { sizeGuardRatio: sanitized.sizeGuardRatio.code, sizeGuardRatioValue: sanitized.sizeGuardRatio.ratio }
|
|
1060
1079
|
: {}),
|
|
@@ -13,8 +13,8 @@
|
|
|
13
13
|
* is configured, and whatever stays undecided is left for review
|
|
14
14
|
* (`review_needed` in the improve ledger).
|
|
15
15
|
* `maxAccepts` caps promotions across both tiers; `applyMode: "queue"` never
|
|
16
|
-
* promotes; `excludeIds` keeps this run's fresh proposals out; a proposal
|
|
17
|
-
*
|
|
16
|
+
* promotes; `excludeIds` keeps this run's fresh proposals out; a proposal that
|
|
17
|
+
* a generating stage routed to a person is left for that person.
|
|
18
18
|
*/
|
|
19
19
|
import fs from "node:fs";
|
|
20
20
|
import path from "node:path";
|
|
@@ -313,11 +313,11 @@ export async function drainProposals(opts, promoteFn = akmProposalAccept, reject
|
|
|
313
313
|
if (isRetireProposal(proposal))
|
|
314
314
|
continue;
|
|
315
315
|
const decision = proposal.gateDecision;
|
|
316
|
-
// Another gate's rejection stands
|
|
317
|
-
//
|
|
316
|
+
// Another gate's rejection stands, and another gate's deferral is a
|
|
317
|
+
// generating stage's hand-off to a person: it is left for that person.
|
|
318
318
|
if (decision?.outcome === "auto-rejected" && !decision.gate?.startsWith(DRAIN_GATE))
|
|
319
319
|
continue;
|
|
320
|
-
if (decision?.outcome === "deferred" && decision.gate
|
|
320
|
+
if (decision?.outcome === "deferred" && !decision.gate?.startsWith(DRAIN_GATE))
|
|
321
321
|
continue;
|
|
322
322
|
if (isEmptyDiff(proposal)) {
|
|
323
323
|
empties.push(proposal.id);
|
package/docs/reference/cli.md
CHANGED
|
@@ -3039,9 +3039,11 @@ Drain the standing pending-proposal backlog instead of adjudicating proposals
|
|
|
3039
3039
|
one at a time. One rule decides each proposal: a proposal whose quality judge
|
|
3040
3040
|
passed on its current content is accepted (unless its target changed since it
|
|
3041
3041
|
was minted — that one is auto-rejected as `stale-target`); an empty diff is
|
|
3042
|
-
rejected;
|
|
3043
|
-
|
|
3044
|
-
|
|
3042
|
+
rejected; a proposal that reflect or distill deferred for review is left for a
|
|
3043
|
+
person; everything else goes to the judgment tier when one is enabled, and
|
|
3044
|
+
is otherwise left for review. A reflect revision that changes the body is
|
|
3045
|
+
deferred for review even when its judge passes it. Default mode stages
|
|
3046
|
+
decisions (queue mode); pass `--promote` to actually accept.
|
|
3045
3047
|
|
|
3046
3048
|
```sh
|
|
3047
3049
|
akm proposal drain --dry-run # Preview without writing
|
|
@@ -356,7 +356,9 @@ guidance. When enabled, engine selection is judgment → triage → strategy →
|
|
|
356
356
|
|
|
357
357
|
`processes.reflect.qualityGate` and `processes.distill.qualityGate` control
|
|
358
358
|
each process's LLM-as-judge quality gate. Each is on unless it sets
|
|
359
|
-
`enabled: false`, and each follows only its own switch.
|
|
359
|
+
`enabled: false`, and each follows only its own switch. A reflect revision
|
|
360
|
+
that changes the body is never auto-accepted; when the judge passes it, it
|
|
361
|
+
waits for review. With the gate off, it waits for review too. The judge is the
|
|
360
362
|
process's own LLM engine, or `defaults.llmEngine` when an agent generates.
|
|
361
363
|
`engine`, `model`, `timeoutMs` and `llm` give the gate a judge of its own,
|
|
362
364
|
resolved over the process's settings the way `triage.judgment` resolves over
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "akm-cli",
|
|
3
|
-
"version": "0.9.
|
|
3
|
+
"version": "0.9.24",
|
|
4
4
|
"type": "module",
|
|
5
5
|
"description": "akm (Agent Knowledge Manager) — a portable, local-first capability library for AI agents. Discover, load, share, and improve reusable skills, scripts, workflows, and knowledge across any shell-capable coding agent, including Claude Code, OpenCode, and Cursor.",
|
|
6
6
|
"keywords": [
|