akm-cli 0.9.19-alpha.1 → 0.9.19-alpha.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -6,6 +6,27 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/).
6
6
 
7
7
  ## [Unreleased]
8
8
 
9
+ ## [0.9.19-alpha.2] - 2026-09-30
10
+
11
+ ### Changed
12
+
13
+ - **Distill's grounding check vetoes only a score of 1; a 2 goes to a person.**
14
+ 0.9.19-alpha.1 made a grounding score of 2 or less `quality_rejected`. A
15
+ calibration of the lesson judge on a local llama.cpp model (34 cases, 5
16
+ passes, 199 calls) scored all 4 off-subject lessons grounding 1 in nearly
17
+ every pass (one scored 2 in 4 of 5 full passes and 1 otherwise), and its only
18
+ false vetoes, 2 of 120 legitimate judgments, were one on-subject lesson that
19
+ adds advice beyond its source and scored 2. Temperature 0 did not make the
20
+ judge repeatable on that server either: scores moved by up to a point between
21
+ passes. A grounding score of 1 is therefore still `quality_rejected` (an
22
+ `improve_ledger` row and a `distill_invoked` event, no proposal, reason
23
+ `Off-subject for its source (grounding 1/5): …`), and a 2 is now
24
+ `review_needed`, a pending proposal for a person with the reason
25
+ `Borderline on grounding (2/5), routed to review: …`, even when the mean of
26
+ novelty and non-redundancy alone would pass it. A mean that alone rejects the
27
+ lesson stays `quality_rejected`, and a grounding score of 3 to 5 is
28
+ unchanged.
29
+
9
30
  ## [0.9.19-alpha.1] - 2026-09-30
10
31
 
11
32
  ### Added
@@ -235,15 +235,20 @@ export function buildReflectJudgePrompt(candidateContent, sourceContent, feedbac
235
235
  * `grounding` is scored with the other lesson criteria but left out of their
236
236
  * mean: a lesson about a different subject than its source reads as novel and
237
237
  * non-redundant, so the mean would pass it (or, in the review band, mint it as
238
- * a pending proposal). A score of {@link UNGROUNDED_MAX_SCORE} or less is a
239
- * rejection whatever the mean says (#999). Only a different subject scores that
240
- * low. A lesson that goes beyond or corrects its source is on its subject:
241
- * distill folds feedback into the lesson, and the judge is never shown it. A
242
- * contradiction of the source is the optional fidelity check's to send to a
243
- * human (`judgeAndQueue` in distill.ts), so the rubric must not pre-empt it.
238
+ * a pending proposal). The rubric reserves 1-2 for a different subject. A score
239
+ * of {@link UNGROUNDED_MAX_SCORE} or less is a rejection whatever the mean says
240
+ * (#999). A higher score up to {@link BORDERLINE_GROUNDING_MAX_SCORE} is only
241
+ * borderline: a lesson on its source's subject that advises beyond it has scored
242
+ * 2, and a score can move a point between runs (see `runQualityJudge`), so it
243
+ * goes to a person unless the mean alone already rejects it. A lesson that goes
244
+ * beyond or corrects its source is on its subject: distill folds feedback into
245
+ * the lesson, and the judge is never shown it. A contradiction of the source is
246
+ * the optional fidelity check's to send to a human (`judgeAndQueue` in
247
+ * distill.ts), so the rubric must not pre-empt it.
244
248
  */
245
249
  const GROUNDING_CRITERION = "grounding";
246
- const UNGROUNDED_MAX_SCORE = 2;
250
+ const UNGROUNDED_MAX_SCORE = 1;
251
+ const BORDERLINE_GROUNDING_MAX_SCORE = 2;
247
252
  const LESSON_JUDGE_CRITERIA = ["novelty", "nonRedundancy", GROUNDING_CRITERION];
248
253
  const REFLECT_JUDGE_CRITERIA = ["feedbackAlignment", "preservation", "quality"];
249
254
  /**
@@ -296,7 +301,12 @@ function judgeResponseSchema(keys) {
296
301
  * The quality judge. Fails closed: no runner, an unparseable verdict or a
297
302
  * provider failure never passes content. Bands: >= 3.5 pass, 2.5-3.5 review,
298
303
  * < 2.5 reject; a `grounding` score of {@link UNGROUNDED_MAX_SCORE} or less
299
- * rejects whatever the mean is. Temperature is pinned to 0 so verdicts do not flip.
304
+ * rejects whatever the mean is, and one of {@link BORDERLINE_GROUNDING_MAX_SCORE}
305
+ * routes a lesson the mean would pass to review (a mean that rejects stays a
306
+ * rejection). Temperature is set to 0, which reduces run-to-run variation but
307
+ * does not remove it: on some servers (llama.cpp batching, for one) the same
308
+ * request can score a point apart, so the routing rules are chosen with that
309
+ * margin in mind.
300
310
  */
301
311
  async function runQualityJudge(feature, config, prompt, keys, chat, options) {
302
312
  const resolved = !options.runnerSelectionFrozen && !options.llmRunner
@@ -339,6 +349,19 @@ async function runQualityJudge(feature, config, prompt, keys, chat, options) {
339
349
  };
340
350
  }
341
351
  const verdict = score >= 3.5 ? { pass: true } : score >= 2.5 ? { pass: false, reviewNeeded: true } : { pass: false };
352
+ // Borderline grounding is a person's call even when the mean would pass; a mean that rejects stays rejected.
353
+ if (criteria &&
354
+ grounding !== undefined &&
355
+ grounding <= BORDERLINE_GROUNDING_MAX_SCORE &&
356
+ (verdict.pass || verdict.reviewNeeded)) {
357
+ return {
358
+ pass: false,
359
+ reviewNeeded: true,
360
+ score,
361
+ reason: `Borderline on grounding (${grounding}/5), routed to review: ${reason}`,
362
+ criteria,
363
+ };
364
+ }
342
365
  return { ...verdict, score, reason, ...(criteria ? { criteria } : {}) };
343
366
  }
344
367
  /** Judge a proposed lesson (or knowledge promotion) against its source. */
@@ -107,10 +107,13 @@ longer offers a TODO: verify placeholder: when feedback asks for something the
107
107
  asset lacks, reflect is told to leave the section unchanged. TODO lines earlier
108
108
  runs already put in your assets stay until you remove them; grep -rniE
109
109
  "TODO:? *verify" over the bundle finds them. Distill's quality judge now also
110
- scores whether a lesson is about what its source is about, and a lesson judged
111
- off-subject is dropped as quality_rejected (a ledger row and a distill_invoked
112
- event, no proposal) instead of passing or waiting in the queue as
113
- review_needed; the judge also reads the same first 3000 characters of the
110
+ scores whether a lesson is about what its source is about. A lesson scored 1
111
+ (off-subject) is dropped as quality_rejected (a ledger row and a
112
+ distill_invoked event, no proposal) instead of passing or waiting in the queue
113
+ as review_needed. A lesson scored 2 is borderline: it is queued as
114
+ review_needed for you to decide, even when its novelty and non-redundancy
115
+ alone would pass it, unless those alone would reject it, which stays
116
+ quality_rejected. The judge also reads the same first 3000 characters of the
114
117
  source body, without frontmatter, that the generator saw. The shipped agent
115
118
  guidance changed with it: akm help agents now says to record feedback about an
116
119
  asset's content, that it helped or turned out wrong, stale or unhelpful, and
@@ -3045,8 +3045,9 @@ Write the reason about the asset's content. Reflect treats it as an unverified
3045
3045
  report to investigate, not a fact to insert, and is told to leave the section
3046
3046
  unchanged when the reason asks for information the asset lacks. Distill's
3047
3047
  quality gate rejects a lesson that is off-subject for the asset it was
3048
- distilled from. A command that failed (`akm show` erroring on the ref, say)
3049
- says nothing about the asset, so it is not a reason to record against it.
3048
+ distilled from and sends a borderline one to review. A command that failed
3049
+ (`akm show` erroring on the ref, say) says nothing about the asset, so it is
3050
+ not a reason to record against it.
3050
3051
 
3051
3052
  ### task
3052
3053
 
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "akm-cli",
3
- "version": "0.9.19-alpha.1",
3
+ "version": "0.9.19-alpha.2",
4
4
  "type": "module",
5
5
  "description": "akm (Agent Knowledge Manager) — a portable, local-first capability library for AI agents. Discover, load, share, and improve reusable skills, scripts, workflows, and knowledge across any shell-capable coding agent, including Claude Code, OpenCode, and Cursor.",
6
6
  "keywords": [