akm-cli 0.9.22 → 0.9.24

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -4,6 +4,89 @@ All notable changes to this project will be documented in this file.
4
4
 
5
5
  The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/).
6
6
 
7
+ ## [0.9.24] - 2026-10-02
8
+
9
+ ### Changed
10
+
11
+ - **Reflect never auto-accepts a revision that changes the body; it waits for
12
+ review.** On 396 labelled reflect edits, the judge-passed edits that changed
13
+ the body were good 12 times in 37, and those that changed only the
14
+ frontmatter 13 times in 13. A revision whose body differs from the asset's,
15
+ ignoring whitespace, or whose source could not be read, is now created
16
+ pending and deferred for review with the reason `body-edit` (gate
17
+ `reflect`). When the quality judge passed it, the judge's scores and reason
18
+ stay on its gate decision. With `processes.reflect.qualityGate` off, no
19
+ judge ran, so the decision carries none. Before, the triage drain accepted a
20
+ judge-passed body edit, and its judgment tier could accept one made with the
21
+ gate off. These body edits now appear in the review queue
22
+ (`akm proposal list`), and `akm proposal show <id>` gives the reason and any
23
+ judge scores. A judge-passed revision that leaves the body unchanged is
24
+ still staged and accepted by the drain. A frontmatter-only revision made
25
+ with the gate off is still left to the drain, and a revision the judge fails
26
+ is still refused.
27
+
28
+ ### Fixed
29
+
30
+ - **Reflect sends a revision to review when its judge fails, instead of
31
+ rejecting it.** A quality judge that timed out, errored or returned a reply
32
+ that could not be parsed gave no verdict, but reflect treated that as a
33
+ rejection. It created no proposal, recorded a `quality_rejected` attempt that
34
+ kept the asset from being reflected again for 14 days, and reported
35
+ `quality gate rejected: score=-1` with a reason saying the revision had been
36
+ routed to review. Reflect now creates the proposal and leaves it pending for
37
+ a person, deferred by the quality gate with the reason `judge-error`, as
38
+ distill does when its judge fails. The triage drain leaves it alone. It has
39
+ not been through the retrieval regression check, which runs only on a
40
+ revision the judge passes. A revision the judge scores too low is still
41
+ refused.
42
+ - **The triage drain leaves every proposal a stage deferred for review to a
43
+ person.** It skipped a deferred proposal only when the quality gate had
44
+ deferred it. Reflect defers with its own gate, `reflect`: a revision whose
45
+ size the size guard flagged, one that echoed the truncation notice, one made
46
+ with no judge configured, and now a body edit. With
47
+ `processes.triage.judgment` enabled, the drain's judgment tier decided those
48
+ proposals, and under `applyMode: promote` it could accept them before anyone
49
+ saw them. The drain now skips every deferral it did not make itself, and
50
+ judges again, as before, the ones it did.
51
+
52
+ ## [0.9.23] - 2026-10-01
53
+
54
+ ### Added
55
+
56
+ - **A quality gate can judge with an engine of its own (#1011).**
57
+ `processes.reflect.qualityGate` and `processes.distill.qualityGate` take
58
+ `engine`, `model`, `timeoutMs` and `llm`. These resolve over the process's own
59
+ settings, the way `processes.triage.judgment` resolves over triage's. Before,
60
+ a gate could not choose its judge. Config validation
61
+ rejects a gate `engine` that is missing or is not an LLM engine. A gate whose
62
+ settings resolve to no LLM engine fails before anything is generated, and
63
+ never falls back to another engine. The judge's LLM usage is recorded under
64
+ its own engine, so the improve report shows it on its own row.
65
+
66
+ ### Changed
67
+
68
+ - **A judge whose engine sets `enableThinking: true` now thinks.** The reflect
69
+ and distill judges always turned thinking off. They still do, unless the
70
+ judge's engine turns it on, so a thinking judge can be configured. To keep a
71
+ thinking engine's judge from thinking, set
72
+ `qualityGate.llm.enableThinking: false`. In the maintainer's replay of
73
+ labelled reflect edits, a thinking judge passed more of the good edits, and
74
+ more of the bad ones too. It used about 20 times the tokens and took a median
75
+ of 30–67 s per judgment, against about 5 s, so it stays opt-in.
76
+
77
+ ### Fixed
78
+
79
+ - **The reflect quality gate follows its own switch.** It ran when either
80
+ `processes.reflect.qualityGate.enabled` was true or distill's gate was on,
81
+ which it is by default. So `reflect.qualityGate.enabled: false` did nothing
82
+ unless distill's gate was off too, and turning distill's gate off also turned
83
+ off reflect's. Each gate now follows only its own `enabled`, which defaults to
84
+ on.
85
+ - **Memory inference uses the configured temperature.** It always sent 0.1,
86
+ ignoring the engine's `temperature` and
87
+ `processes.memoryInference.llm.temperature`. It now sends the configured
88
+ temperature, and 0.1 only when nothing sets one.
89
+
7
90
  ## [0.9.22] - 2026-10-01
8
91
 
9
92
  ### Fixed
@@ -46,7 +46,7 @@ import { buildRefVocabulary, scoreEncodingSalience } from "./encoding-salience.j
46
46
  import { resolveImproveStrategy, resolveProcessEnabled } from "./improve-strategies.js";
47
47
  import { recordLedgerAttempt } from "./ledger.js";
48
48
  import { computeSalience, upsertAssetSalience } from "./salience.js";
49
- import { callStage, mintProposal, noticeSet, rejectedProposalContext, runLessonQualityJudge, stageRunner, } from "./stage.js";
49
+ import { callStage, mintProposal, noticeSet, rejectedProposalContext, resolveQualityGateJudge, runLessonQualityJudge, stageRunner, } from "./stage.js";
50
50
  /**
51
51
  * Input types distill structurally refuses: a lesson is the distilled form
52
52
  * (distilling one would mint `lessons/lesson-…-lesson`), and env/secret bytes
@@ -306,6 +306,7 @@ export async function akmDistill(options) {
306
306
  config,
307
307
  profile,
308
308
  runner: stageRunner(options, config, profile, "distill", notices.add),
309
+ judgeRunner: resolveQualityGateJudge(config, profile, "distill", notices.add),
309
310
  notices,
310
311
  eligMeta,
311
312
  asset,
@@ -322,8 +323,11 @@ async function distill(run, targetKind, kind, outputRef, feedbackEvents) {
322
323
  // A reinforced memory graduates to knowledge without a generation call.
323
324
  const promotion = targetKind === "lesson" ? null : await planPromotion(run, feedbackEvents);
324
325
  if (promotion) {
325
- if (run.runner && (promotion.existing || qualityGateEnabled(run)))
326
+ if (run.runner && promotion.existing)
326
327
  assertRunnerCredentials(run.runner);
328
+ const judge = run.judgeRunner ?? run.runner;
329
+ if (judge && qualityGateEnabled(run))
330
+ assertRunnerCredentials(judge);
327
331
  const promoted = await promoteToKnowledge(run, promotion);
328
332
  stampInputSalience(run);
329
333
  return promoted;
@@ -441,7 +445,7 @@ async function judgeAndQueue(run, out) {
441
445
  const source = out.source ? parseFrontmatter(out.source).content.trim() : "";
442
446
  const verdict = await runLessonQualityJudge(run.config, content, source, run.options.chat, {
443
447
  ...(similarLessons.length > 0 ? { similarLessons } : {}),
444
- ...(run.runner ? { llmRunner: run.runner } : {}),
448
+ ...((run.judgeRunner ?? run.runner) ? { llmRunner: run.judgeRunner ?? run.runner } : {}),
445
449
  ...(run.options.signal ? { signal: run.options.signal } : {}),
446
450
  onNotices: run.notices.add,
447
451
  });
@@ -44,7 +44,7 @@ import { resolveImproveLlmExecution } from "./execution.js";
44
44
  import { recordLedgerAttempt } from "./ledger.js";
45
45
  import { classifyReflectChange, splitFrontmatter } from "./reflect-noise.js";
46
46
  import { loadRetrievalQueries, runRetrievalRegressionGate } from "./retrieval-gate.js";
47
- import { callStage, mintProposal, noticeSet, rejectedProposalContext, runReflectQualityJudge, } from "./stage.js";
47
+ import { callStage, mintProposal, noticeSet, rejectedProposalContext, resolveQualityGateJudge, runReflectQualityJudge, } from "./stage.js";
48
48
  const MAX_FEEDBACK_LINES = 10;
49
49
  const MAX_GLOBAL_FEEDBACK_LINES = 20;
50
50
  function readOnlyEventsContext(ctx) {
@@ -932,8 +932,10 @@ const NOISE_SUBREASONS = {
932
932
  };
933
933
  /**
934
934
  * Sanitize, drop a no-op/cosmetic (and optionally low-value) change, judge the
935
- * exact content that would be persisted, then mint. Size-flagged or
936
- * truncation-leaking content skips the judge and waits for review.
935
+ * exact content that would be persisted, then mint. A judge pass is staged only
936
+ * when the body is unchanged: a body edit the judge passes, or one made with the
937
+ * gate off, waits for review. Size-flagged or truncation-leaking content skips
938
+ * the judge and waits for review.
937
939
  */
938
940
  async function finalizeReflectProposal(args) {
939
941
  const { run, assetContent, result, judge, feedback } = args;
@@ -980,6 +982,7 @@ async function finalizeReflectProposal(args) {
980
982
  return reflectFailure(run, result, "quality_rejected", message, false);
981
983
  };
982
984
  let verdict;
985
+ let judgeFailed = false;
983
986
  if (judged) {
984
987
  verdict = await runReflectQualityJudge(run.config, payload.content, assetContent ?? "", feedback, options.chat, {
985
988
  runnerSelectionFrozen: true,
@@ -988,7 +991,9 @@ async function finalizeReflectProposal(args) {
988
991
  ...(options.signal ? { signal: options.signal } : {}),
989
992
  onNotices: run.notices.add,
990
993
  });
991
- if (!verdict.pass) {
994
+ // A judge that timed out, errored or replied unparseably gave no verdict: a person reviews the revision.
995
+ judgeFailed = verdict.reviewNeeded === true && verdict.score === -1;
996
+ if (!verdict.pass && !judgeFailed) {
992
997
  return refuse(verdict.reason, {
993
998
  qualityScore: verdict.score,
994
999
  qualityReason: verdict.reason,
@@ -997,7 +1002,7 @@ async function finalizeReflectProposal(args) {
997
1002
  }
998
1003
  }
999
1004
  // #722: a rewrite of an existing asset must not grade lower on its own retrieval queries.
1000
- if (judged && judge.runner && assetContent !== undefined) {
1005
+ if (verdict?.pass && judge.runner && assetContent !== undefined) {
1001
1006
  const retrieval = await runRetrievalRegressionGate({
1002
1007
  ref: payload.ref,
1003
1008
  before: assetContent,
@@ -1026,9 +1031,18 @@ async function finalizeReflectProposal(args) {
1026
1031
  };
1027
1032
  const reviewReasons = [
1028
1033
  ...(judge.skippedNoJudge ? ["no-judge-configured"] : []),
1034
+ ...(judgeFailed ? ["judge-error"] : []),
1029
1035
  ...(sanitized.sizeGuardRatio ? ["reflect-size-ratio"] : []),
1030
1036
  ...(sanitized.truncationMarkerLeaked ? ["reflect-truncation-leak"] : []),
1031
1037
  ];
1038
+ // A revision that changes the body is never auto-accepted: on labelled edits, the judge's
1039
+ // passes on body edits were good 12 times in 37, and on frontmatter-only edits 13 in 13.
1040
+ // One that nothing above holds for review (the judge passed it, or the gate is off) waits
1041
+ // for a person, as does a revision with no source to compare.
1042
+ const bodyOf = (content) => splitFrontmatter(content).body.replace(/\s+/g, " ").trim();
1043
+ const bodyEdit = reviewReasons.length === 0 && (assetContent === undefined || bodyOf(assetContent) !== bodyOf(payload.content));
1044
+ if (bodyEdit)
1045
+ reviewReasons.push("body-edit");
1032
1046
  const proposal = mintProposal(run.stash, options.ctx, {
1033
1047
  ref: payload.ref,
1034
1048
  ...(options.target ? { target: options.target } : {}),
@@ -1042,8 +1056,12 @@ async function finalizeReflectProposal(args) {
1042
1056
  ? {
1043
1057
  review: {
1044
1058
  reason: reviewReasons.join("+"),
1045
- gate: "reflect",
1059
+ // The quality gate's hand-off to a person, as distill's: the triage drain leaves it alone.
1060
+ gate: judgeFailed ? "quality-gate" : "reflect",
1046
1061
  ...(sanitized.sizeGuardRatio ? { measured: Math.round(sanitized.sizeGuardRatio.ratio * 100) } : {}),
1062
+ // The reviewer sees why the judge passed it (with the gate off, nothing judged it).
1063
+ ...(bodyEdit && verdict?.criteria ? { scores: verdict.criteria } : {}),
1064
+ ...(bodyEdit && verdict ? { judgeReason: verdict.reason } : {}),
1047
1065
  },
1048
1066
  }
1049
1067
  : { judged: verdict });
@@ -1055,6 +1073,7 @@ async function finalizeReflectProposal(args) {
1055
1073
  source: "reflect",
1056
1074
  engine: run.engineName,
1057
1075
  ...(judge.skippedNoJudge ? { qualityGateSkippedNoJudge: true } : {}),
1076
+ ...(judgeFailed ? { qualityReason: verdict?.reason } : {}),
1058
1077
  ...(sanitized.sizeGuardRatio
1059
1078
  ? { sizeGuardRatio: sanitized.sizeGuardRatio.code, sizeGuardRatioValue: sanitized.sizeGuardRatio.ratio }
1060
1079
  : {}),
@@ -1113,11 +1132,14 @@ export async function akmReflect(options = {}) {
1113
1132
  notices.add(resolutionNotices);
1114
1133
  const run = { options, stash, config, runnerSpec, engineName, notices, emitInvoked, emitFailed };
1115
1134
  // Judge selection is frozen before dispatch so a missing judge credential fails first.
1116
- const judgeWanted = (activeStrategy?.processes?.reflect?.qualityGate?.enabled ?? false) ||
1117
- (activeStrategy?.processes?.distill?.qualityGate?.enabled ?? true);
1135
+ const judgeWanted = activeStrategy?.processes?.reflect?.qualityGate?.enabled ?? true;
1118
1136
  let judgeRunner;
1119
1137
  if (judgeWanted) {
1120
- if (runnerIsLlm(runnerSpec)) {
1138
+ const gateJudge = resolveQualityGateJudge(config, activeStrategy, "reflect", notices.add);
1139
+ if (gateJudge) {
1140
+ judgeRunner = gateJudge;
1141
+ }
1142
+ else if (runnerIsLlm(runnerSpec)) {
1121
1143
  judgeRunner = runnerSpec;
1122
1144
  }
1123
1145
  else {
@@ -7,7 +7,7 @@ import { parseEmbeddedJsonResponse } from "../../core/parse.js";
7
7
  import { warn } from "../../core/warn.js";
8
8
  import { LlmCallError } from "../../llm/client.js";
9
9
  import { callStructured } from "../../llm/structured-call.js";
10
- import { withLlmStage } from "../../llm/usage-telemetry.js";
10
+ import { currentLlmStage, withLlmStage } from "../../llm/usage-telemetry.js";
11
11
  import { isProceduralRejection } from "../proposal/proposal-types.js";
12
12
  import { createProposal, listProposalsReadOnly, proposalContentHash, recordGateDecision, } from "../proposal/repository.js";
13
13
  import { resolveImproveLlmExecution } from "./execution.js";
@@ -55,7 +55,7 @@ export async function callStage(call) {
55
55
  ];
56
56
  let failure;
57
57
  try {
58
- const raw = await callStructured({
58
+ const dispatch = () => callStructured({
59
59
  feature: call.feature,
60
60
  ...(call.gate
61
61
  ? { akmConfig: call.gate.config, ...(call.gate.enabled !== undefined ? { enabled: call.gate.enabled } : {}) }
@@ -74,6 +74,12 @@ export async function callStage(call) {
74
74
  failure ??= { ok: false, reason: event.reason, ...(event.error ? { error: event.error.message } : {}) };
75
75
  },
76
76
  });
77
+ // Usage is credited to the engine that serves the call, which for a gate's own judge
78
+ // (#1011) is not the stage's planned engine.
79
+ const stage = currentLlmStage();
80
+ const raw = await (stage === undefined
81
+ ? dispatch()
82
+ : withLlmStage(stage, dispatch, { engine: call.runner.engine }));
77
83
  return raw === undefined ? (failure ?? { ok: false, reason: "error" }) : { ok: true, raw };
78
84
  }
79
85
  catch (err) {
@@ -152,6 +158,34 @@ export function stageJudgedProposal(stash, proposal, judged, proposalsCtx) {
152
158
  return proposal;
153
159
  }
154
160
  }
161
+ /**
162
+ * The judge a quality gate names for itself (#1011): the gate's `engine`,
163
+ * `model`, `timeoutMs` and `llm` over the process's own settings, as
164
+ * `processes.triage.judgment` resolves over triage. `undefined` when the gate
165
+ * is off or sets none of them, so the caller keeps its own judge. Throws when
166
+ * they resolve to no LLM engine, before anything is generated: a judge must be
167
+ * one, and the gate never falls back to another.
168
+ */
169
+ export function resolveQualityGateJudge(config, profile, processName, onNotices) {
170
+ const process = profile?.processes?.[processName];
171
+ const gate = process?.qualityGate;
172
+ if (!gate || gate.enabled === false)
173
+ return undefined;
174
+ if (!["engine", "model", "timeoutMs", "llm"].some((key) => Object.hasOwn(gate, key)))
175
+ return undefined;
176
+ const resolved = resolveImproveLlmExecution({
177
+ config,
178
+ processName: `${processName}-quality-judge`,
179
+ ...(profile ? { profile } : {}),
180
+ ...(process ? { process } : {}),
181
+ current: gate,
182
+ });
183
+ if (!resolved) {
184
+ throw new ConfigError(`The ${processName} quality gate's judge must be an LLM engine. Set processes.${processName}.qualityGate.engine to one.`, "INVALID_CONFIG_FILE");
185
+ }
186
+ onNotices?.(resolved.notices);
187
+ return resolved.runner;
188
+ }
155
189
  /** Lesson judge prompt; similar existing lessons let it mark near-duplicates down. */
156
190
  export function buildJudgePrompt(lessonContent, sourceContent, similarLessons) {
157
191
  const lines = [
@@ -328,7 +362,8 @@ async function runQualityJudge(feature, config, prompt, keys, chat, options) {
328
362
  system: "Return only valid JSON. No prose.",
329
363
  prompt,
330
364
  request: {
331
- enableThinking: false,
365
+ // Off unless the judge's own engine enables thinking (a slower, separate judge engine, #1011).
366
+ enableThinking: runner.connection.enableThinking === true,
332
367
  temperature: 0,
333
368
  responseSchema: judgeResponseSchema(keys),
334
369
  ...(Object.hasOwn(options, "timeoutMs") ? { timeoutMs: options.timeoutMs } : {}),
@@ -13,8 +13,8 @@
13
13
  * is configured, and whatever stays undecided is left for review
14
14
  * (`review_needed` in the improve ledger).
15
15
  * `maxAccepts` caps promotions across both tiers; `applyMode: "queue"` never
16
- * promotes; `excludeIds` keeps this run's fresh proposals out; a proposal the
17
- * distill quality gate routed to a human is left for that human.
16
+ * promotes; `excludeIds` keeps this run's fresh proposals out; a proposal that
17
+ * a generating stage routed to a person is left for that person.
18
18
  */
19
19
  import fs from "node:fs";
20
20
  import path from "node:path";
@@ -313,11 +313,11 @@ export async function drainProposals(opts, promoteFn = akmProposalAccept, reject
313
313
  if (isRetireProposal(proposal))
314
314
  continue;
315
315
  const decision = proposal.gateDecision;
316
- // Another gate's rejection stands; a human-review deferral from the distill
317
- // quality gate is left for that human.
316
+ // Another gate's rejection stands, and another gate's deferral is a
317
+ // generating stage's hand-off to a person: it is left for that person.
318
318
  if (decision?.outcome === "auto-rejected" && !decision.gate?.startsWith(DRAIN_GATE))
319
319
  continue;
320
- if (decision?.outcome === "deferred" && decision.gate === "quality-gate")
320
+ if (decision?.outcome === "deferred" && !decision.gate?.startsWith(DRAIN_GATE))
321
321
  continue;
322
322
  if (isEmptyDiff(proposal)) {
323
323
  empties.push(proposal.id);
@@ -308,6 +308,16 @@ export const AkmConfigSchema = AkmConfigBaseSchema.superRefine((config, ctx) =>
308
308
  });
309
309
  }
310
310
  }
311
+ const gateEngine = processConfig.qualityGate?.engine;
312
+ if (gateEngine && config.engines?.[gateEngine]?.kind !== "llm") {
313
+ ctx.addIssue({
314
+ code: z.ZodIssueCode.custom,
315
+ path: ["improve", "strategies", strategyName, "processes", processName, "qualityGate", "engine"],
316
+ message: config.engines?.[gateEngine]
317
+ ? "a quality-gate judge must be an LLM engine"
318
+ : "engine does not name a configured engine",
319
+ });
320
+ }
311
321
  }
312
322
  }
313
323
  // #464.a: defaultWriteTarget must name a configured source. 0.9.0 (spec
@@ -59,12 +59,23 @@ const excludeRefPrefixesField = z.array(z.string().min(1)).optional();
59
59
  */
60
60
  const processLimitField = positiveInt.optional();
61
61
  /**
62
- * Distill process: LLM-as-judge lesson quality gate. Default ON (R3);
63
- * fail-open — judge failure/timeout/parse errors pass through. Set
64
- * `enabled: false` on the distill process to opt out. Also read on the
65
- * `reflect` process (proposal-side quality gate; see reflect.ts).
62
+ * The distill and reflect LLM-as-judge quality gates. Default ON; set
63
+ * `enabled: false` on the process to opt out. A judge error, timeout or
64
+ * unparsable reply never passes: distill sends the lesson to review and
65
+ * reflect refuses the edit. `engine`, `model`, `timeoutMs` and `llm` give the
66
+ * gate a judge of its own (#1011), resolved over the process's own settings;
67
+ * with none set, the gate judges with the generator's runner.
66
68
  */
67
- const qualityGateField = z.object({ enabled: z.boolean().optional() }).passthrough().optional();
69
+ const qualityGateField = z
70
+ .object({
71
+ enabled: z.boolean().optional(),
72
+ engine: engineName.optional(),
73
+ model: nonEmptyString.optional(),
74
+ timeoutMs: z.union([positiveInt, z.null()]).optional(),
75
+ llm: LlmInvocationOverridesSchema.passthrough().optional(),
76
+ })
77
+ .passthrough()
78
+ .optional();
68
79
  /**
69
80
  * WS-3b: CLS (Complementary Learning System) interleaving (step 9).
70
81
  * distill/memoryInference prompts include embedding-retrieved existing adjacent
@@ -88,7 +88,8 @@ export async function compressMemoryToDerivedMemory(llmRunner, body, signal, akm
88
88
  { role: "user", content: userPrompt },
89
89
  ],
90
90
  request: {
91
- temperature: 0.1,
91
+ // The engine's or the process's configured temperature; 0.1 only when neither sets one.
92
+ temperature: llmRunner.connection.temperature ?? 0.1,
92
93
  timeoutMs: llmRunner.timeoutMs,
93
94
  signal,
94
95
  responseSchema: DERIVED_MEMORY_JSON_SCHEMA,
@@ -35851,7 +35851,13 @@ var IMPROVE_PROCESS_BASE_FIELDS = {
35851
35851
  var allowedTypesField = exports_external.array(exports_external.string().min(1)).optional();
35852
35852
  var excludeRefPrefixesField = exports_external.array(exports_external.string().min(1)).optional();
35853
35853
  var processLimitField = positiveInt.optional();
35854
- var qualityGateField = exports_external.object({ enabled: exports_external.boolean().optional() }).passthrough().optional();
35854
+ var qualityGateField = exports_external.object({
35855
+ enabled: exports_external.boolean().optional(),
35856
+ engine: engineName.optional(),
35857
+ model: nonEmptyString.optional(),
35858
+ timeoutMs: exports_external.union([positiveInt, exports_external.null()]).optional(),
35859
+ llm: LlmInvocationOverridesSchema.passthrough().optional()
35860
+ }).passthrough().optional();
35855
35861
  var clsField = exports_external.object({
35856
35862
  enabled: exports_external.boolean().optional(),
35857
35863
  adjacentCount: exports_external.number().int().min(1).optional()
@@ -36468,6 +36474,14 @@ var AkmConfigSchema = AkmConfigBaseSchema.superRefine((config, ctx) => {
36468
36474
  });
36469
36475
  }
36470
36476
  }
36477
+ const gateEngine = processConfig.qualityGate?.engine;
36478
+ if (gateEngine && config.engines?.[gateEngine]?.kind !== "llm") {
36479
+ ctx.addIssue({
36480
+ code: exports_external.ZodIssueCode.custom,
36481
+ path: ["improve", "strategies", strategyName, "processes", processName, "qualityGate", "engine"],
36482
+ message: config.engines?.[gateEngine] ? "a quality-gate judge must be an LLM engine" : "engine does not name a configured engine"
36483
+ });
36484
+ }
36471
36485
  }
36472
36486
  }
36473
36487
  if (config.defaultWriteTarget !== undefined) {
@@ -35179,7 +35179,13 @@ var IMPROVE_PROCESS_BASE_FIELDS = {
35179
35179
  var allowedTypesField = exports_external.array(exports_external.string().min(1)).optional();
35180
35180
  var excludeRefPrefixesField = exports_external.array(exports_external.string().min(1)).optional();
35181
35181
  var processLimitField = positiveInt.optional();
35182
- var qualityGateField = exports_external.object({ enabled: exports_external.boolean().optional() }).passthrough().optional();
35182
+ var qualityGateField = exports_external.object({
35183
+ enabled: exports_external.boolean().optional(),
35184
+ engine: engineName.optional(),
35185
+ model: nonEmptyString.optional(),
35186
+ timeoutMs: exports_external.union([positiveInt, exports_external.null()]).optional(),
35187
+ llm: LlmInvocationOverridesSchema.passthrough().optional()
35188
+ }).passthrough().optional();
35183
35189
  var clsField = exports_external.object({
35184
35190
  enabled: exports_external.boolean().optional(),
35185
35191
  adjacentCount: exports_external.number().int().min(1).optional()
@@ -35796,6 +35802,14 @@ var AkmConfigSchema = AkmConfigBaseSchema.superRefine((config, ctx) => {
35796
35802
  });
35797
35803
  }
35798
35804
  }
35805
+ const gateEngine = processConfig.qualityGate?.engine;
35806
+ if (gateEngine && config.engines?.[gateEngine]?.kind !== "llm") {
35807
+ ctx.addIssue({
35808
+ code: exports_external.ZodIssueCode.custom,
35809
+ path: ["improve", "strategies", strategyName, "processes", processName, "qualityGate", "engine"],
35810
+ message: config.engines?.[gateEngine] ? "a quality-gate judge must be an LLM engine" : "engine does not name a configured engine"
35811
+ });
35812
+ }
35799
35813
  }
35800
35814
  }
35801
35815
  if (config.defaultWriteTarget !== undefined) {
@@ -3039,9 +3039,11 @@ Drain the standing pending-proposal backlog instead of adjudicating proposals
3039
3039
  one at a time. One rule decides each proposal: a proposal whose quality judge
3040
3040
  passed on its current content is accepted (unless its target changed since it
3041
3041
  was minted — that one is auto-rejected as `stale-target`); an empty diff is
3042
- rejected; everything else goes to the judgment tier when one is enabled, and
3043
- is otherwise left for review. Default mode stages decisions (queue mode); pass
3044
- `--promote` to actually accept.
3042
+ rejected; a proposal that reflect or distill deferred for review is left for a
3043
+ person; everything else goes to the judgment tier when one is enabled, and
3044
+ is otherwise left for review. A reflect revision that changes the body is
3045
+ deferred for review even when its judge passes it. Default mode stages
3046
+ decisions (queue mode); pass `--promote` to actually accept.
3045
3047
 
3046
3048
  ```sh
3047
3049
  akm proposal drain --dry-run # Preview without writing
@@ -103,10 +103,11 @@ An LLM engine may set `enableThinking: false` to turn thinking off and
103
103
  `reasoningEffort` to a value such as `"none"`, `"low"`, or `"high"`. AKM sends
104
104
  **both** wire forms — `chat_template_kwargs.enable_thinking` and top-level
105
105
  `enable_thinking` — whenever `enableThinking` resolves, from engine config
106
- or a calling process (improve's `consolidate`/`reflect` and the distill
107
- quality gate always request `enableThinking: false` for a machine-readable
108
- payload); `reasoningEffort` is always sent as top-level `reasoning_effort`
109
- when set. Backend support: llama.cpp direct honors both forms
106
+ or a calling process (improve's `consolidate`/`reflect` always request
107
+ `enableThinking: false` for a machine-readable payload, and the reflect and
108
+ distill quality-gate judges do too unless the judge's engine sets
109
+ `enableThinking: true`); `reasoningEffort` is always sent as top-level
110
+ `reasoning_effort` when set. Backend support: llama.cpp direct honors both forms
110
111
  (`reasoning_effort` from build ≥ b10644); vLLM honors
111
112
  `chat_template_kwargs`; Bifrost drops `chat_template_kwargs` and passes
112
113
  `reasoning_effort` through, so also set `reasoningEffort: "none"` behind it; a
@@ -353,6 +354,44 @@ guidance. When enabled, engine selection is judgment → triage → strategy →
353
354
  }
354
355
  ```
355
356
 
357
+ `processes.reflect.qualityGate` and `processes.distill.qualityGate` control
358
+ each process's LLM-as-judge quality gate. Each is on unless it sets
359
+ `enabled: false`, and each follows only its own switch. A reflect revision
360
+ that changes the body is never auto-accepted; when the judge passes it, it
361
+ waits for review. With the gate off, it waits for review too. The judge is the
362
+ process's own LLM engine, or `defaults.llmEngine` when an agent generates.
363
+ `engine`, `model`, `timeoutMs` and `llm` give the gate a judge of its own,
364
+ resolved over the process's settings the way `triage.judgment` resolves over
365
+ triage's. A gate whose settings resolve to no LLM engine fails before anything
366
+ is generated; it never falls back to another engine. The judge runs at
367
+ temperature 0 with thinking off unless its engine sets `enableThinking: true`.
368
+ Thinking is slow: on a 27B llama.cpp server, a thinking judgment took a median
369
+ of 30–67 s and up to about 3 minutes, against about 5 s without.
370
+ Behind a gateway that drops `chat_template_kwargs` (Bifrost), point a thinking
371
+ judge's engine at the server directly.
372
+
373
+ ```jsonc
374
+ {
375
+ "engines": {
376
+ "judge": {
377
+ "kind": "llm",
378
+ "endpoint": "http://127.0.0.1:8080/v1/chat/completions",
379
+ "model": "qwen3-27b",
380
+ "enableThinking": true
381
+ }
382
+ },
383
+ "improve": {
384
+ "strategies": {
385
+ "nightly": {
386
+ "processes": {
387
+ "reflect": { "qualityGate": { "engine": "judge" } }
388
+ }
389
+ }
390
+ }
391
+ }
392
+ }
393
+ ```
394
+
356
395
  No shipped strategy turns improve-stage session extraction on.
357
396
  `proactiveMaintenance` is on only in the `proactive-maintenance` preset; run
358
397
  `akm improve --strategy proactive-maintenance` to use that opt-in preset.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "akm-cli",
3
- "version": "0.9.22",
3
+ "version": "0.9.24",
4
4
  "type": "module",
5
5
  "description": "akm (Agent Knowledge Manager) — a portable, local-first capability library for AI agents. Discover, load, share, and improve reusable skills, scripts, workflows, and knowledge across any shell-capable coding agent, including Claude Code, OpenCode, and Cursor.",
6
6
  "keywords": [
@@ -756,6 +756,57 @@
756
756
  "properties": {
757
757
  "enabled": {
758
758
  "type": "boolean"
759
+ },
760
+ "engine": {
761
+ "type": "string",
762
+ "maxLength": 63,
763
+ "pattern": "^(?!akm-)[a-z][a-z0-9]*(?:-[a-z0-9]+)*$"
764
+ },
765
+ "model": {
766
+ "type": "string",
767
+ "minLength": 1
768
+ },
769
+ "timeoutMs": {
770
+ "anyOf": [
771
+ {
772
+ "type": "integer",
773
+ "exclusiveMinimum": 0
774
+ },
775
+ {
776
+ "type": "null"
777
+ }
778
+ ]
779
+ },
780
+ "llm": {
781
+ "type": "object",
782
+ "properties": {
783
+ "temperature": {
784
+ "type": "number"
785
+ },
786
+ "maxTokens": {
787
+ "type": "integer",
788
+ "exclusiveMinimum": 0
789
+ },
790
+ "supportsJsonSchema": {
791
+ "type": "boolean"
792
+ },
793
+ "extraParams": {
794
+ "type": "object",
795
+ "additionalProperties": {}
796
+ },
797
+ "contextLength": {
798
+ "type": "integer",
799
+ "exclusiveMinimum": 0
800
+ },
801
+ "enableThinking": {
802
+ "type": "boolean"
803
+ },
804
+ "reasoningEffort": {
805
+ "type": "string",
806
+ "minLength": 1
807
+ }
808
+ },
809
+ "additionalProperties": true
759
810
  }
760
811
  },
761
812
  "additionalProperties": true
@@ -845,6 +896,57 @@
845
896
  "properties": {
846
897
  "enabled": {
847
898
  "type": "boolean"
899
+ },
900
+ "engine": {
901
+ "type": "string",
902
+ "maxLength": 63,
903
+ "pattern": "^(?!akm-)[a-z][a-z0-9]*(?:-[a-z0-9]+)*$"
904
+ },
905
+ "model": {
906
+ "type": "string",
907
+ "minLength": 1
908
+ },
909
+ "timeoutMs": {
910
+ "anyOf": [
911
+ {
912
+ "type": "integer",
913
+ "exclusiveMinimum": 0
914
+ },
915
+ {
916
+ "type": "null"
917
+ }
918
+ ]
919
+ },
920
+ "llm": {
921
+ "type": "object",
922
+ "properties": {
923
+ "temperature": {
924
+ "type": "number"
925
+ },
926
+ "maxTokens": {
927
+ "type": "integer",
928
+ "exclusiveMinimum": 0
929
+ },
930
+ "supportsJsonSchema": {
931
+ "type": "boolean"
932
+ },
933
+ "extraParams": {
934
+ "type": "object",
935
+ "additionalProperties": {}
936
+ },
937
+ "contextLength": {
938
+ "type": "integer",
939
+ "exclusiveMinimum": 0
940
+ },
941
+ "enableThinking": {
942
+ "type": "boolean"
943
+ },
944
+ "reasoningEffort": {
945
+ "type": "string",
946
+ "minLength": 1
947
+ }
948
+ },
949
+ "additionalProperties": true
848
950
  }
849
951
  },
850
952
  "additionalProperties": true
@@ -2306,6 +2408,57 @@
2306
2408
  "properties": {
2307
2409
  "enabled": {
2308
2410
  "type": "boolean"
2411
+ },
2412
+ "engine": {
2413
+ "type": "string",
2414
+ "maxLength": 63,
2415
+ "pattern": "^(?!akm-)[a-z][a-z0-9]*(?:-[a-z0-9]+)*$"
2416
+ },
2417
+ "model": {
2418
+ "type": "string",
2419
+ "minLength": 1
2420
+ },
2421
+ "timeoutMs": {
2422
+ "anyOf": [
2423
+ {
2424
+ "type": "integer",
2425
+ "exclusiveMinimum": 0
2426
+ },
2427
+ {
2428
+ "type": "null"
2429
+ }
2430
+ ]
2431
+ },
2432
+ "llm": {
2433
+ "type": "object",
2434
+ "properties": {
2435
+ "temperature": {
2436
+ "type": "number"
2437
+ },
2438
+ "maxTokens": {
2439
+ "type": "integer",
2440
+ "exclusiveMinimum": 0
2441
+ },
2442
+ "supportsJsonSchema": {
2443
+ "type": "boolean"
2444
+ },
2445
+ "extraParams": {
2446
+ "type": "object",
2447
+ "additionalProperties": {}
2448
+ },
2449
+ "contextLength": {
2450
+ "type": "integer",
2451
+ "exclusiveMinimum": 0
2452
+ },
2453
+ "enableThinking": {
2454
+ "type": "boolean"
2455
+ },
2456
+ "reasoningEffort": {
2457
+ "type": "string",
2458
+ "minLength": 1
2459
+ }
2460
+ },
2461
+ "additionalProperties": true
2309
2462
  }
2310
2463
  },
2311
2464
  "additionalProperties": true
@@ -2395,6 +2548,57 @@
2395
2548
  "properties": {
2396
2549
  "enabled": {
2397
2550
  "type": "boolean"
2551
+ },
2552
+ "engine": {
2553
+ "type": "string",
2554
+ "maxLength": 63,
2555
+ "pattern": "^(?!akm-)[a-z][a-z0-9]*(?:-[a-z0-9]+)*$"
2556
+ },
2557
+ "model": {
2558
+ "type": "string",
2559
+ "minLength": 1
2560
+ },
2561
+ "timeoutMs": {
2562
+ "anyOf": [
2563
+ {
2564
+ "type": "integer",
2565
+ "exclusiveMinimum": 0
2566
+ },
2567
+ {
2568
+ "type": "null"
2569
+ }
2570
+ ]
2571
+ },
2572
+ "llm": {
2573
+ "type": "object",
2574
+ "properties": {
2575
+ "temperature": {
2576
+ "type": "number"
2577
+ },
2578
+ "maxTokens": {
2579
+ "type": "integer",
2580
+ "exclusiveMinimum": 0
2581
+ },
2582
+ "supportsJsonSchema": {
2583
+ "type": "boolean"
2584
+ },
2585
+ "extraParams": {
2586
+ "type": "object",
2587
+ "additionalProperties": {}
2588
+ },
2589
+ "contextLength": {
2590
+ "type": "integer",
2591
+ "exclusiveMinimum": 0
2592
+ },
2593
+ "enableThinking": {
2594
+ "type": "boolean"
2595
+ },
2596
+ "reasoningEffort": {
2597
+ "type": "string",
2598
+ "minLength": 1
2599
+ }
2600
+ },
2601
+ "additionalProperties": true
2398
2602
  }
2399
2603
  },
2400
2604
  "additionalProperties": true
@@ -3316,6 +3520,57 @@
3316
3520
  "properties": {
3317
3521
  "enabled": {
3318
3522
  "type": "boolean"
3523
+ },
3524
+ "engine": {
3525
+ "type": "string",
3526
+ "maxLength": 63,
3527
+ "pattern": "^(?!akm-)[a-z][a-z0-9]*(?:-[a-z0-9]+)*$"
3528
+ },
3529
+ "model": {
3530
+ "type": "string",
3531
+ "minLength": 1
3532
+ },
3533
+ "timeoutMs": {
3534
+ "anyOf": [
3535
+ {
3536
+ "type": "integer",
3537
+ "exclusiveMinimum": 0
3538
+ },
3539
+ {
3540
+ "type": "null"
3541
+ }
3542
+ ]
3543
+ },
3544
+ "llm": {
3545
+ "type": "object",
3546
+ "properties": {
3547
+ "temperature": {
3548
+ "type": "number"
3549
+ },
3550
+ "maxTokens": {
3551
+ "type": "integer",
3552
+ "exclusiveMinimum": 0
3553
+ },
3554
+ "supportsJsonSchema": {
3555
+ "type": "boolean"
3556
+ },
3557
+ "extraParams": {
3558
+ "type": "object",
3559
+ "additionalProperties": {}
3560
+ },
3561
+ "contextLength": {
3562
+ "type": "integer",
3563
+ "exclusiveMinimum": 0
3564
+ },
3565
+ "enableThinking": {
3566
+ "type": "boolean"
3567
+ },
3568
+ "reasoningEffort": {
3569
+ "type": "string",
3570
+ "minLength": 1
3571
+ }
3572
+ },
3573
+ "additionalProperties": true
3319
3574
  }
3320
3575
  },
3321
3576
  "additionalProperties": true
@@ -3655,6 +3910,57 @@
3655
3910
  "properties": {
3656
3911
  "enabled": {
3657
3912
  "type": "boolean"
3913
+ },
3914
+ "engine": {
3915
+ "type": "string",
3916
+ "maxLength": 63,
3917
+ "pattern": "^(?!akm-)[a-z][a-z0-9]*(?:-[a-z0-9]+)*$"
3918
+ },
3919
+ "model": {
3920
+ "type": "string",
3921
+ "minLength": 1
3922
+ },
3923
+ "timeoutMs": {
3924
+ "anyOf": [
3925
+ {
3926
+ "type": "integer",
3927
+ "exclusiveMinimum": 0
3928
+ },
3929
+ {
3930
+ "type": "null"
3931
+ }
3932
+ ]
3933
+ },
3934
+ "llm": {
3935
+ "type": "object",
3936
+ "properties": {
3937
+ "temperature": {
3938
+ "type": "number"
3939
+ },
3940
+ "maxTokens": {
3941
+ "type": "integer",
3942
+ "exclusiveMinimum": 0
3943
+ },
3944
+ "supportsJsonSchema": {
3945
+ "type": "boolean"
3946
+ },
3947
+ "extraParams": {
3948
+ "type": "object",
3949
+ "additionalProperties": {}
3950
+ },
3951
+ "contextLength": {
3952
+ "type": "integer",
3953
+ "exclusiveMinimum": 0
3954
+ },
3955
+ "enableThinking": {
3956
+ "type": "boolean"
3957
+ },
3958
+ "reasoningEffort": {
3959
+ "type": "string",
3960
+ "minLength": 1
3961
+ }
3962
+ },
3963
+ "additionalProperties": true
3658
3964
  }
3659
3965
  },
3660
3966
  "additionalProperties": true
@@ -3744,6 +4050,57 @@
3744
4050
  "properties": {
3745
4051
  "enabled": {
3746
4052
  "type": "boolean"
4053
+ },
4054
+ "engine": {
4055
+ "type": "string",
4056
+ "maxLength": 63,
4057
+ "pattern": "^(?!akm-)[a-z][a-z0-9]*(?:-[a-z0-9]+)*$"
4058
+ },
4059
+ "model": {
4060
+ "type": "string",
4061
+ "minLength": 1
4062
+ },
4063
+ "timeoutMs": {
4064
+ "anyOf": [
4065
+ {
4066
+ "type": "integer",
4067
+ "exclusiveMinimum": 0
4068
+ },
4069
+ {
4070
+ "type": "null"
4071
+ }
4072
+ ]
4073
+ },
4074
+ "llm": {
4075
+ "type": "object",
4076
+ "properties": {
4077
+ "temperature": {
4078
+ "type": "number"
4079
+ },
4080
+ "maxTokens": {
4081
+ "type": "integer",
4082
+ "exclusiveMinimum": 0
4083
+ },
4084
+ "supportsJsonSchema": {
4085
+ "type": "boolean"
4086
+ },
4087
+ "extraParams": {
4088
+ "type": "object",
4089
+ "additionalProperties": {}
4090
+ },
4091
+ "contextLength": {
4092
+ "type": "integer",
4093
+ "exclusiveMinimum": 0
4094
+ },
4095
+ "enableThinking": {
4096
+ "type": "boolean"
4097
+ },
4098
+ "reasoningEffort": {
4099
+ "type": "string",
4100
+ "minLength": 1
4101
+ }
4102
+ },
4103
+ "additionalProperties": true
3747
4104
  }
3748
4105
  },
3749
4106
  "additionalProperties": true