akm-cli 0.9.19-alpha.1 → 0.9.19-alpha.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md
CHANGED
|
@@ -6,6 +6,27 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/).
|
|
|
6
6
|
|
|
7
7
|
## [Unreleased]
|
|
8
8
|
|
|
9
|
+
## [0.9.19-alpha.2] - 2026-09-30
|
|
10
|
+
|
|
11
|
+
### Changed
|
|
12
|
+
|
|
13
|
+
- **Distill's grounding check vetoes only a score of 1; a 2 goes to a person.**
|
|
14
|
+
0.9.19-alpha.1 made a grounding score of 2 or less `quality_rejected`. A
|
|
15
|
+
calibration of the lesson judge on a local llama.cpp model (34 cases, 5
|
|
16
|
+
passes, 199 calls) scored all 4 off-subject lessons grounding 1 in nearly
|
|
17
|
+
every pass (one scored 2 in 4 of 5 full passes and 1 otherwise), and its only
|
|
18
|
+
false vetoes, 2 of 120 legitimate judgments, were one on-subject lesson that
|
|
19
|
+
adds advice beyond its source and scored 2. Temperature 0 did not make the
|
|
20
|
+
judge repeatable on that server either: scores moved by up to a point between
|
|
21
|
+
passes. A grounding score of 1 is therefore still `quality_rejected` (an
|
|
22
|
+
`improve_ledger` row and a `distill_invoked` event, no proposal, reason
|
|
23
|
+
`Off-subject for its source (grounding 1/5): …`), and a 2 is now
|
|
24
|
+
`review_needed`, a pending proposal for a person with the reason
|
|
25
|
+
`Borderline on grounding (2/5), routed to review: …`, even when the mean of
|
|
26
|
+
novelty and non-redundancy alone would pass it. A mean that alone rejects the
|
|
27
|
+
lesson stays `quality_rejected`, and a grounding score of 3 to 5 is
|
|
28
|
+
unchanged.
|
|
29
|
+
|
|
9
30
|
## [0.9.19-alpha.1] - 2026-09-30
|
|
10
31
|
|
|
11
32
|
### Added
|
|
@@ -235,15 +235,20 @@ export function buildReflectJudgePrompt(candidateContent, sourceContent, feedbac
|
|
|
235
235
|
* `grounding` is scored with the other lesson criteria but left out of their
|
|
236
236
|
* mean: a lesson about a different subject than its source reads as novel and
|
|
237
237
|
* non-redundant, so the mean would pass it (or, in the review band, mint it as
|
|
238
|
-
* a pending proposal).
|
|
239
|
-
*
|
|
240
|
-
*
|
|
241
|
-
*
|
|
242
|
-
*
|
|
243
|
-
*
|
|
238
|
+
* a pending proposal). The rubric reserves 1-2 for a different subject. A score
|
|
239
|
+
* of {@link UNGROUNDED_MAX_SCORE} or less is a rejection whatever the mean says
|
|
240
|
+
* (#999). A higher score up to {@link BORDERLINE_GROUNDING_MAX_SCORE} is only
|
|
241
|
+
* borderline: a lesson on its source's subject that advises beyond it has scored
|
|
242
|
+
* 2, and a score can move a point between runs (see `runQualityJudge`), so it
|
|
243
|
+
* goes to a person unless the mean alone already rejects it. A lesson that goes
|
|
244
|
+
* beyond or corrects its source is on its subject: distill folds feedback into
|
|
245
|
+
* the lesson, and the judge is never shown it. A contradiction of the source is
|
|
246
|
+
* the optional fidelity check's to send to a human (`judgeAndQueue` in
|
|
247
|
+
* distill.ts), so the rubric must not pre-empt it.
|
|
244
248
|
*/
|
|
245
249
|
const GROUNDING_CRITERION = "grounding";
|
|
246
|
-
const UNGROUNDED_MAX_SCORE =
|
|
250
|
+
const UNGROUNDED_MAX_SCORE = 1;
|
|
251
|
+
const BORDERLINE_GROUNDING_MAX_SCORE = 2;
|
|
247
252
|
const LESSON_JUDGE_CRITERIA = ["novelty", "nonRedundancy", GROUNDING_CRITERION];
|
|
248
253
|
const REFLECT_JUDGE_CRITERIA = ["feedbackAlignment", "preservation", "quality"];
|
|
249
254
|
/**
|
|
@@ -296,7 +301,12 @@ function judgeResponseSchema(keys) {
|
|
|
296
301
|
* The quality judge. Fails closed: no runner, an unparseable verdict or a
|
|
297
302
|
* provider failure never passes content. Bands: >= 3.5 pass, 2.5-3.5 review,
|
|
298
303
|
* < 2.5 reject; a `grounding` score of {@link UNGROUNDED_MAX_SCORE} or less
|
|
299
|
-
* rejects whatever the mean is
|
|
304
|
+
* rejects whatever the mean is, and one of {@link BORDERLINE_GROUNDING_MAX_SCORE}
|
|
305
|
+
* routes a lesson the mean would pass to review (a mean that rejects stays a
|
|
306
|
+
* rejection). Temperature is set to 0, which reduces run-to-run variation but
|
|
307
|
+
* does not remove it: on some servers (llama.cpp batching, for one) the same
|
|
308
|
+
* request can score a point apart, so the routing rules are chosen with that
|
|
309
|
+
* margin in mind.
|
|
300
310
|
*/
|
|
301
311
|
async function runQualityJudge(feature, config, prompt, keys, chat, options) {
|
|
302
312
|
const resolved = !options.runnerSelectionFrozen && !options.llmRunner
|
|
@@ -339,6 +349,19 @@ async function runQualityJudge(feature, config, prompt, keys, chat, options) {
|
|
|
339
349
|
};
|
|
340
350
|
}
|
|
341
351
|
const verdict = score >= 3.5 ? { pass: true } : score >= 2.5 ? { pass: false, reviewNeeded: true } : { pass: false };
|
|
352
|
+
// Borderline grounding is a person's call even when the mean would pass; a mean that rejects stays rejected.
|
|
353
|
+
if (criteria &&
|
|
354
|
+
grounding !== undefined &&
|
|
355
|
+
grounding <= BORDERLINE_GROUNDING_MAX_SCORE &&
|
|
356
|
+
(verdict.pass || verdict.reviewNeeded)) {
|
|
357
|
+
return {
|
|
358
|
+
pass: false,
|
|
359
|
+
reviewNeeded: true,
|
|
360
|
+
score,
|
|
361
|
+
reason: `Borderline on grounding (${grounding}/5), routed to review: ${reason}`,
|
|
362
|
+
criteria,
|
|
363
|
+
};
|
|
364
|
+
}
|
|
342
365
|
return { ...verdict, score, reason, ...(criteria ? { criteria } : {}) };
|
|
343
366
|
}
|
|
344
367
|
/** Judge a proposed lesson (or knowledge promotion) against its source. */
|
|
@@ -107,10 +107,13 @@ longer offers a TODO: verify placeholder: when feedback asks for something the
|
|
|
107
107
|
asset lacks, reflect is told to leave the section unchanged. TODO lines earlier
|
|
108
108
|
runs already put in your assets stay until you remove them; grep -rniE
|
|
109
109
|
"TODO:? *verify" over the bundle finds them. Distill's quality judge now also
|
|
110
|
-
scores whether a lesson is about what its source is about
|
|
111
|
-
off-subject is dropped as quality_rejected (a ledger row and a
|
|
112
|
-
event, no proposal) instead of passing or waiting in the queue
|
|
113
|
-
review_needed
|
|
110
|
+
scores whether a lesson is about what its source is about. A lesson scored 1
|
|
111
|
+
(off-subject) is dropped as quality_rejected (a ledger row and a
|
|
112
|
+
distill_invoked event, no proposal) instead of passing or waiting in the queue
|
|
113
|
+
as review_needed. A lesson scored 2 is borderline: it is queued as
|
|
114
|
+
review_needed for you to decide, even when its novelty and non-redundancy
|
|
115
|
+
alone would pass it, unless those alone would reject it, which stays
|
|
116
|
+
quality_rejected. The judge also reads the same first 3000 characters of the
|
|
114
117
|
source body, without frontmatter, that the generator saw. The shipped agent
|
|
115
118
|
guidance changed with it: akm help agents now says to record feedback about an
|
|
116
119
|
asset's content, that it helped or turned out wrong, stale or unhelpful, and
|
package/docs/reference/cli.md
CHANGED
|
@@ -3045,8 +3045,9 @@ Write the reason about the asset's content. Reflect treats it as an unverified
|
|
|
3045
3045
|
report to investigate, not a fact to insert, and is told to leave the section
|
|
3046
3046
|
unchanged when the reason asks for information the asset lacks. Distill's
|
|
3047
3047
|
quality gate rejects a lesson that is off-subject for the asset it was
|
|
3048
|
-
distilled from
|
|
3049
|
-
says nothing about the asset, so it is
|
|
3048
|
+
distilled from and sends a borderline one to review. A command that failed
|
|
3049
|
+
(`akm show` erroring on the ref, say) says nothing about the asset, so it is
|
|
3050
|
+
not a reason to record against it.
|
|
3050
3051
|
|
|
3051
3052
|
### task
|
|
3052
3053
|
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "akm-cli",
|
|
3
|
-
"version": "0.9.19-alpha.
|
|
3
|
+
"version": "0.9.19-alpha.2",
|
|
4
4
|
"type": "module",
|
|
5
5
|
"description": "akm (Agent Knowledge Manager) — a portable, local-first capability library for AI agents. Discover, load, share, and improve reusable skills, scripts, workflows, and knowledge across any shell-capable coding agent, including Claude Code, OpenCode, and Cursor.",
|
|
6
6
|
"keywords": [
|