akm-cli 0.9.25-alpha.3 → 0.9.25-alpha.4

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -6,6 +6,24 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/).
6
6
 
7
7
  ## [Unreleased]
8
8
 
9
+ ## [0.9.25-alpha.4] - 2026-10-04
10
+
11
+ ### Changed
12
+
13
+ - **The reflect quality judge's rubric is written for the frontmatter-only
14
+ revisions reflect makes, and says what drove each one.** A revision answering
15
+ negative feedback (any `[negative]` line) is judged only on the fields it
16
+ changes, since it cannot change the body and need not resolve feedback about
17
+ it. A field change is justified when the field is missing or broken, claims
18
+ more than the body covers, or, for a note the feedback calls stale, needs to
19
+ name the version or date the body records. A maintenance revision needs a
20
+ missing or broken field, and rewording a sound one is churn. New values are
21
+ checked as claims against the body. On 113 frontmatter-only proposals
22
+ reviewed by Claude Opus, the production model passed 78 of the 95 good ones
23
+ and 3 of the 18 bad ones, against 53 and 7 under the previous rubric, which
24
+ was tuned on body rewrites. The agent judge's tool paragraph drops its
25
+ feedback-points check, which contradicted that rule.
26
+
9
27
  ## [0.9.25-alpha.3] - 2026-10-04
10
28
 
11
29
  ### Added
@@ -269,17 +269,27 @@ function buildChangedRegion(sourceContent, candidateContent) {
269
269
  */
270
270
  function reflectJudgeToolRules(ref) {
271
271
  const asset = ref ? `The asset is \`${ref}\`: read it with akm_show, ` : "Read an asset with akm_show ";
272
- return `Tools: ${asset}or an asset the changed region names, only to verify a fact the revision adds or alters; the text above already shows every change. Do not search, do not read anything else, and do not use a tool to judge structure or wording. One or two reads at most. Before scoring, check three lists: (1) every statement the revision adds: find each in the asset, or as a fact the feedback states about the subject, and score QUALITY 1-2 if any is in neither; a statement is found only when the asset or the feedback says it, in any words: a new step, cause, consequence or detail that merely seems to follow is not found; feedback says what to fix and is not content, so an added statement about how the asset was used, found or verified is unsupported; (2) every fact, caveat and field of the source: find each in the revision, and score PRESERVATION 1-3 if any is missing; (3) every point the feedback makes: find the text it is about changed in the revision, and score NEED 2-3 if any is not; a note that restates the feedback does not address it. A read that finds nothing wrong raises no score above what these lists support. Then reply with the JSON.`;
272
+ return `Tools: ${asset}or an asset the changed region names, only to verify a fact the revision adds or alters; the text above already shows every change. Do not search, do not read anything else, and do not use a tool to judge structure or wording. One or two reads at most. Before scoring, check two lists: (1) every statement the revision adds: find each in the asset, or as a fact the feedback states about the subject, and score QUALITY 1-2 if any is in neither; a statement is found only when the asset or the feedback says it, in any words: a new step, cause, consequence or detail that merely seems to follow is not found; feedback says what to fix and is not content, so an added statement about how the asset was used, found or verified is unsupported; (2) every fact, caveat and field of the source: find each in the revision, and score PRESERVATION 1-3 if any is missing. A read that finds nothing wrong raises no score above what these lists support. Then reply with the JSON.`;
273
273
  }
274
+ /**
275
+ * The reflect judge's rubric, for the frontmatter-only revisions reflect makes, and the paragraph that says what
276
+ * drove the revision: negative feedback (any `[negative]` line) or maintenance. Tuned on the production model
277
+ * against 113 reviewed proposals (the stash's eval/judge-gate/tuning/reflect/judge-fm, rubric f07).
278
+ */
279
+ const REFLECT_JUDGE_INTRO = "You are evaluating a proposed revision of an existing akm asset's frontmatter. The revision may change only the `description`, the `when_to_use` and the title (a level-1 heading added when the body has none); the body is unchanged.";
280
+ const REFLECT_JUDGE_NEGATIVE = "This revision answers negative feedback. It cannot change the body, so it need not resolve the feedback: judge only the fields it changes, and never fault it for a field it leaves as it was. Most negative feedback says the asset did not help with a task it was retrieved for: a retrieval miss, not a defect. It justifies changing a field only when the field claims more than the body covers (narrow it to what the body covers), or when the feedback calls the asset stale, outdated, superseded or historical (a new `when_to_use` must then name the version or date the body records). Feedback about a task the asset never claims to cover justifies no change. Repairing a missing or broken field is always needed, whatever the feedback says.";
281
+ const REFLECT_JUDGE_MAINTENANCE = "This revision is maintenance: there is no negative feedback. Only a missing or broken field needs a change; rewording a sound `description` or `when_to_use` is churn, however accurate.";
274
282
  /** Judge prompt for an in-place revision. `tools` is set when the judge runs on an agent engine. */
275
283
  export function buildReflectJudgePrompt(candidateContent, sourceContent, feedback, tools) {
276
284
  return [
277
- "You are evaluating a proposed revision to an existing akm asset.",
285
+ REFLECT_JUDGE_INTRO,
286
+ "",
287
+ feedback.some((line) => line.startsWith("[negative]")) ? REFLECT_JUDGE_NEGATIVE : REFLECT_JUDGE_MAINTENANCE,
278
288
  "",
279
289
  "Score this revision on each criterion from 1 (poor) to 5 (excellent):",
280
- "1. NEED: Does the revision fix a concrete problem in the source? Concrete problems are: something the feedback reports as wrong or missing; a factual error; broken, garbled, truncated or missing text; and a missing or broken title, description or when_to_use field. Compare the source's description with the revision's: a description with a sentence split in its middle by a stray period or line-wrap artifact (as in 'calls. asset writes'), an unbalanced or escaped quote, or a truncated ending is broken, and repairing it is a concrete problem fixed even when the rest of the revision only adds stamps or reformats; adding or removing a trailing period, or rewording a readable description, repairs nothing. A missing title, description or when_to_use is a concrete problem whether or not the feedback mentions it: empty or positive feedback does not mean the source was complete, and the frontmatter is complete only when it has all three. A when_to_use is missing when the frontmatter has none, even if the body has a 'when to use' section; moving or copying that text into the field is the fix. A missing type field or a provenance stamp such as generated or verified is not a concrete problem. Score 4-5 when the revision fixes one, even a small one, whatever else it also reformats. Score 1-2 when the source was already complete and correct and the revision only rewords, restates, reformats, adds a type field or a stamp, or adds headings, an introduction or a table of contents.",
281
- "2. PRESERVATION: Does it keep every concrete fact, identifier, command, path, number, example, caveat and frontmatter field from the source, without truncation? Check the changed region line by line. Score 1-3 when any of them is dropped or weakened, even when the revision also fixes something. Trimming narrative that states no fact, or removing what the feedback asks to remove or rescope, is not a drop.",
282
- "3. QUALITY: Is it coherent and accurate, with no claims, steps or details that the source or the feedback does not support? Check each added or changed statement against the source and the feedback: a statement that follows from either counts as supported, and frontmatter stamps such as type, generated, verified or quality, whatever their values, and a restatement of existing content are not claims. Score 1-2 when the revision adds a claim neither supports, turns a draft or proposal into a decision, strengthens a statement beyond the source (a preference into a requirement, a possibility into a fact, a pending fix into a done one), adds a hedge such as 'may be outdated', or a placeholder such as 'TODO' or 'verify'; a fix elsewhere in the revision does not raise this score.",
290
+ "1. NEED: Does every changed field fix a real problem? Real problems: a missing `description`, `when_to_use` or title; a broken description (a sentence split by a stray period at a line wrap, an escaped or unbalanced quote, a truncated ending, a heading fragment); and, for a revision answering negative feedback, a field that claims more than the body covers, or a stale note's `when_to_use` that does not name the version or date the body records. Replacing a stray period that splits a sentence with a comma, a word or nothing repairs a broken description, however small the change looks. Score 4-5 when every changed field fixes one. Score 1-2 when any changed field rewrites a sound field. Score NEED on the changed fields alone: leaving negative feedback about the body unresolved never lowers it.",
291
+ "2. PRESERVATION: Does the new description keep every fact the old one carried: names, identifiers, numbers, versions, paths, qualifiers and status words such as 'Proposal' or 'draft'? Are all other frontmatter fields unchanged? Score 1-3 when anything is dropped or changed.",
292
+ "3. QUALITY: Is every new value supported by the body, without inventing, over-claiming or misdescribing? Check each new value as a claim against the body; restating the body in other words is supported. Score only values the revision adds or changes: a field it leaves as it was, however stale, is never this revision's fault. Score 1-2 when a new value says something the body does not support or the opposite of what it says, keeps a truncated or garbled fragment, or offers a dated or historical note for current work: when the feedback calls the note stale, outdated, superseded or historical, or the body records the version or date it was true for, a new `when_to_use` that does not name that version or date, or a new value that calls a dated snapshot 'current', scores 1-2.",
283
293
  "",
284
294
  "Feedback:",
285
295
  "```",
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "akm-cli",
3
- "version": "0.9.25-alpha.3",
3
+ "version": "0.9.25-alpha.4",
4
4
  "type": "module",
5
5
  "description": "akm (Agent Knowledge Manager) — a portable, local-first capability library for AI agents. Discover, load, share, and improve reusable skills, scripts, workflows, and knowledge across any shell-capable coding agent, including Claude Code, OpenCode, and Cursor.",
6
6
  "keywords": [