akm-cli 0.9.25-alpha.3 → 0.9.25-alpha.4
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +18 -0
- package/dist/commands/improve/stage.js +15 -5
- package/package.json +1 -1
package/CHANGELOG.md
CHANGED
|
@@ -6,6 +6,24 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/).
|
|
|
6
6
|
|
|
7
7
|
## [Unreleased]
|
|
8
8
|
|
|
9
|
+
## [0.9.25-alpha.4] - 2026-10-04
|
|
10
|
+
|
|
11
|
+
### Changed
|
|
12
|
+
|
|
13
|
+
- **The reflect quality judge's rubric is written for the frontmatter-only
|
|
14
|
+
revisions reflect makes, and says what drove each one.** A revision answering
|
|
15
|
+
negative feedback (any `[negative]` line) is judged only on the fields it
|
|
16
|
+
changes, since it cannot change the body and need not resolve feedback about
|
|
17
|
+
it. A field change is justified when the field is missing or broken, claims
|
|
18
|
+
more than the body covers, or, for a note the feedback calls stale, needs to
|
|
19
|
+
name the version or date the body records. A maintenance revision needs a
|
|
20
|
+
missing or broken field, and rewording a sound one is churn. New values are
|
|
21
|
+
checked as claims against the body. On 113 frontmatter-only proposals
|
|
22
|
+
reviewed by Claude Opus, the production model passed 78 of the 95 good ones
|
|
23
|
+
and 3 of the 18 bad ones, against 53 and 7 under the previous rubric, which
|
|
24
|
+
was tuned on body rewrites. The agent judge's tool paragraph drops its
|
|
25
|
+
feedback-points check, which contradicted that rule.
|
|
26
|
+
|
|
9
27
|
## [0.9.25-alpha.3] - 2026-10-04
|
|
10
28
|
|
|
11
29
|
### Added
|
|
@@ -269,17 +269,27 @@ function buildChangedRegion(sourceContent, candidateContent) {
|
|
|
269
269
|
*/
|
|
270
270
|
function reflectJudgeToolRules(ref) {
|
|
271
271
|
const asset = ref ? `The asset is \`${ref}\`: read it with akm_show, ` : "Read an asset with akm_show ";
|
|
272
|
-
return `Tools: ${asset}or an asset the changed region names, only to verify a fact the revision adds or alters; the text above already shows every change. Do not search, do not read anything else, and do not use a tool to judge structure or wording. One or two reads at most. Before scoring, check
|
|
272
|
+
return `Tools: ${asset}or an asset the changed region names, only to verify a fact the revision adds or alters; the text above already shows every change. Do not search, do not read anything else, and do not use a tool to judge structure or wording. One or two reads at most. Before scoring, check two lists: (1) every statement the revision adds: find each in the asset, or as a fact the feedback states about the subject, and score QUALITY 1-2 if any is in neither; a statement is found only when the asset or the feedback says it, in any words: a new step, cause, consequence or detail that merely seems to follow is not found; feedback says what to fix and is not content, so an added statement about how the asset was used, found or verified is unsupported; (2) every fact, caveat and field of the source: find each in the revision, and score PRESERVATION 1-3 if any is missing. A read that finds nothing wrong raises no score above what these lists support. Then reply with the JSON.`;
|
|
273
273
|
}
|
|
274
|
+
/**
|
|
275
|
+
* The reflect judge's rubric, for the frontmatter-only revisions reflect makes, and the paragraph that says what
|
|
276
|
+
* drove the revision: negative feedback (any `[negative]` line) or maintenance. Tuned on the production model
|
|
277
|
+
* against 113 reviewed proposals (the stash's eval/judge-gate/tuning/reflect/judge-fm, rubric f07).
|
|
278
|
+
*/
|
|
279
|
+
const REFLECT_JUDGE_INTRO = "You are evaluating a proposed revision of an existing akm asset's frontmatter. The revision may change only the `description`, the `when_to_use` and the title (a level-1 heading added when the body has none); the body is unchanged.";
|
|
280
|
+
const REFLECT_JUDGE_NEGATIVE = "This revision answers negative feedback. It cannot change the body, so it need not resolve the feedback: judge only the fields it changes, and never fault it for a field it leaves as it was. Most negative feedback says the asset did not help with a task it was retrieved for: a retrieval miss, not a defect. It justifies changing a field only when the field claims more than the body covers (narrow it to what the body covers), or when the feedback calls the asset stale, outdated, superseded or historical (a new `when_to_use` must then name the version or date the body records). Feedback about a task the asset never claims to cover justifies no change. Repairing a missing or broken field is always needed, whatever the feedback says.";
|
|
281
|
+
const REFLECT_JUDGE_MAINTENANCE = "This revision is maintenance: there is no negative feedback. Only a missing or broken field needs a change; rewording a sound `description` or `when_to_use` is churn, however accurate.";
|
|
274
282
|
/** Judge prompt for an in-place revision. `tools` is set when the judge runs on an agent engine. */
|
|
275
283
|
export function buildReflectJudgePrompt(candidateContent, sourceContent, feedback, tools) {
|
|
276
284
|
return [
|
|
277
|
-
|
|
285
|
+
REFLECT_JUDGE_INTRO,
|
|
286
|
+
"",
|
|
287
|
+
feedback.some((line) => line.startsWith("[negative]")) ? REFLECT_JUDGE_NEGATIVE : REFLECT_JUDGE_MAINTENANCE,
|
|
278
288
|
"",
|
|
279
289
|
"Score this revision on each criterion from 1 (poor) to 5 (excellent):",
|
|
280
|
-
"1. NEED: Does
|
|
281
|
-
"2. PRESERVATION: Does
|
|
282
|
-
"3. QUALITY: Is
|
|
290
|
+
"1. NEED: Does every changed field fix a real problem? Real problems: a missing `description`, `when_to_use` or title; a broken description (a sentence split by a stray period at a line wrap, an escaped or unbalanced quote, a truncated ending, a heading fragment); and, for a revision answering negative feedback, a field that claims more than the body covers, or a stale note's `when_to_use` that does not name the version or date the body records. Replacing a stray period that splits a sentence with a comma, a word or nothing repairs a broken description, however small the change looks. Score 4-5 when every changed field fixes one. Score 1-2 when any changed field rewrites a sound field. Score NEED on the changed fields alone: leaving negative feedback about the body unresolved never lowers it.",
|
|
291
|
+
"2. PRESERVATION: Does the new description keep every fact the old one carried: names, identifiers, numbers, versions, paths, qualifiers and status words such as 'Proposal' or 'draft'? Are all other frontmatter fields unchanged? Score 1-3 when anything is dropped or changed.",
|
|
292
|
+
"3. QUALITY: Is every new value supported by the body, without inventing, over-claiming or misdescribing? Check each new value as a claim against the body; restating the body in other words is supported. Score only values the revision adds or changes: a field it leaves as it was, however stale, is never this revision's fault. Score 1-2 when a new value says something the body does not support or the opposite of what it says, keeps a truncated or garbled fragment, or offers a dated or historical note for current work: when the feedback calls the note stale, outdated, superseded or historical, or the body records the version or date it was true for, a new `when_to_use` that does not name that version or date, or a new value that calls a dated snapshot 'current', scores 1-2.",
|
|
283
293
|
"",
|
|
284
294
|
"Feedback:",
|
|
285
295
|
"```",
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "akm-cli",
|
|
3
|
-
"version": "0.9.25-alpha.
|
|
3
|
+
"version": "0.9.25-alpha.4",
|
|
4
4
|
"type": "module",
|
|
5
5
|
"description": "akm (Agent Knowledge Manager) — a portable, local-first capability library for AI agents. Discover, load, share, and improve reusable skills, scripts, workflows, and knowledge across any shell-capable coding agent, including Claude Code, OpenCode, and Cursor.",
|
|
6
6
|
"keywords": [
|