akm-cli 0.9.20 → 0.9.21
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +85 -1
- package/dist/assets/hints/cli-hints-full.md +9 -4
- package/dist/assets/hints/cli-hints-short.md +7 -4
- package/dist/assets/improve-strategies/default.json +2 -2
- package/dist/assets/improve-strategies/proactive-maintenance.json +1 -1
- package/dist/assets/improve-strategies/thorough.json +1 -1
- package/dist/assets/stash-skeleton/README.md +5 -1
- package/dist/commands/feedback-cli.js +14 -6
- package/dist/commands/health/improve-metrics.js +1 -1
- package/dist/commands/improve/distill.js +3 -1
- package/dist/commands/improve/improve-cli.js +1 -1
- package/dist/commands/improve/improve.js +3 -27
- package/dist/commands/improve/preparation.js +35 -29
- package/dist/commands/improve/proactive-maintenance.js +2 -9
- package/dist/commands/improve/reflect.js +3 -2
- package/dist/commands/improve/stage.js +36 -31
- package/dist/commands/proposal/repository.js +14 -11
- package/dist/commands/remember.js +7 -1
- package/dist/core/improve-result.js +4 -1
- package/dist/core/write-source.js +28 -5
- package/dist/scripts/akm-migrate-node.js +4 -4
- package/dist/scripts/akm-migrate.js +4 -4
- package/dist/sources/providers/git-stash.js +16 -9
- package/dist/storage/repositories/improve-runs-repository.js +1 -3
- package/docs/reference/cli.md +41 -15
- package/package.json +1 -1
package/CHANGELOG.md
CHANGED
|
@@ -4,7 +4,91 @@ All notable changes to this project will be documented in this file.
|
|
|
4
4
|
|
|
5
5
|
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/).
|
|
6
6
|
|
|
7
|
-
## [
|
|
7
|
+
## [0.9.21] - 2026-10-01
|
|
8
|
+
|
|
9
|
+
### Changed
|
|
10
|
+
|
|
11
|
+
- **`akm improve` rewrites an asset only from negative feedback.** Reflect's
|
|
12
|
+
signal delta reads negative feedback only, so an asset whose recent feedback
|
|
13
|
+
is positive or a note is no longer planned for a rewrite; any signal planned
|
|
14
|
+
one before. `akm feedback <ref> --negative --reason "<what is wrong and what
|
|
15
|
+
should change>"` flags the asset for review, and the next improve run
|
|
16
|
+
proposes a fix based on the reason. `--positive` records that the asset
|
|
17
|
+
helped (it raises its ranking) and never triggers a rewrite. An explicit ref
|
|
18
|
+
(`akm improve skills/x`) still plans one, and distill still reads any recent
|
|
19
|
+
signal on a memory.
|
|
20
|
+
- **The fallback lanes score assets and no longer plan them.** High salience,
|
|
21
|
+
and proactive maintenance in a strategy that enables it (the shipped
|
|
22
|
+
`proactive-maintenance` strategy), still pick assets and score them
|
|
23
|
+
(salience and outcome), but a pick no longer enters the loop, so nothing is
|
|
24
|
+
reflected or distilled for it and improve no longer rewrites assets on a
|
|
25
|
+
proactive cadence. `eligibilitySource` is now `signal-delta` or `scope`;
|
|
26
|
+
`proactive` and `high-salience` stay valid on the rows an older release wrote.
|
|
27
|
+
`--require-feedback-signal` still turns the lanes off.
|
|
28
|
+
- **`distill.requirePlannedRefs` is `false` in the `default` and `thorough`
|
|
29
|
+
strategies** (the other presets inherit it from `default`). With reflect
|
|
30
|
+
planned only from negative feedback, `true` would have skipped distill on
|
|
31
|
+
every run where no ref has negative feedback; distill now keeps working
|
|
32
|
+
through memories with fresh feedback on each run. A strategy that sets it
|
|
33
|
+
`true` still skips distill when no ref is planned for reflect.
|
|
34
|
+
- **`akm feedback`'s help and errors, both CLI hints, the stash README and the
|
|
35
|
+
docs say what each signal does:** `--negative --reason` flags the asset for
|
|
36
|
+
review and improve proposes a fix from the reason, so the reason should say
|
|
37
|
+
what is wrong and what should change; `--positive` raises the ranking and
|
|
38
|
+
does not trigger a rewrite. They also say that improve no longer rewrites on
|
|
39
|
+
a proactive cadence or from positive signals. The `default` and
|
|
40
|
+
`proactive-maintenance` strategy descriptions match.
|
|
41
|
+
- **The reflect judge asks whether a rewrite is needed, and both judges pass
|
|
42
|
+
only when every criterion scores 4 or more.** Judging on the mean (3.5 or
|
|
43
|
+
more) let rewrites that only reworded a correct asset through. The reflect
|
|
44
|
+
judge's criteria are now **need** (does the revision fix a concrete problem
|
|
45
|
+
in the source: something the feedback reports, a factual error, or broken,
|
|
46
|
+
garbled, truncated or missing text, frontmatter fields included; rewording,
|
|
47
|
+
restating, reformatting or adding headings scores 1 or 2), **preservation**
|
|
48
|
+
and **quality**, replacing feedback alignment, preservation and quality, and
|
|
49
|
+
the "overlap with the source is expected" line is gone. Its JSON keys are
|
|
50
|
+
`need`, `preservation` and `quality`. The lesson judge passes only when
|
|
51
|
+
novelty and non-redundancy both score 4 or more. A verdict that does not
|
|
52
|
+
pass keeps its routing (a mean of 2.5 or more is a review, below that a
|
|
53
|
+
rejection) and the grounding rules are unchanged, grounding staying outside
|
|
54
|
+
the mean and the pass rule.
|
|
55
|
+
- **A quality-judge pass keeps its evidence.** The `staged` decision the
|
|
56
|
+
quality gate stamps on a proposal now carries the per-criterion `scores` and
|
|
57
|
+
the judge's `judgeReason` (`gateDecision.scores`, `gateDecision.judgeReason`
|
|
58
|
+
in `akm proposal show --format json`), and they stay on the proposal when the
|
|
59
|
+
drain accepts it, so a later audit can read why a rewrite passed.
|
|
60
|
+
- **Each accepted proposal is committed as it happens when its bundle is a git
|
|
61
|
+
repository.** Manual `akm proposal accept`, `akm proposal drain`, the improve
|
|
62
|
+
triage pre-pass and judgment accepts all commit exactly the paths the accept
|
|
63
|
+
wrote or removed, a retirement's archived copy and tombstone and the source
|
|
64
|
+
memory a consolidate promotion retires included, as one local commit
|
|
65
|
+
`akm accept: <generator> <proposal-id-8> <ref>`. Before, only a `kind: "git"`
|
|
66
|
+
bundle committed at accept; an accept into a filesystem bundle with a `.git`
|
|
67
|
+
directory (the working bundle `akm init` creates is one) stayed uncommitted
|
|
68
|
+
until a sync, and a standalone `akm proposal accept` never committed it. A
|
|
69
|
+
commit that fails warns and the accept stands. A non-git bundle is
|
|
70
|
+
unchanged.
|
|
71
|
+
- **The end-of-run sync and `akm sync` commit even when they cannot push.** A
|
|
72
|
+
branch with no upstream, or behind or diverged from it, used to fail before
|
|
73
|
+
committing, leaving the run's changes in the working tree. They are now
|
|
74
|
+
committed and only the push is skipped, with `not pushed: ...` as the result's
|
|
75
|
+
`reason`. A branch ahead of its upstream (an accept commits locally) is pushed
|
|
76
|
+
along with the sync commit, where it used to be refused.
|
|
77
|
+
|
|
78
|
+
### Fixed
|
|
79
|
+
|
|
80
|
+
- **A dotted token no longer splits a synthesized description.** `akm
|
|
81
|
+
remember` ended a sentence at every `.`, so `192.168.0.203` became `192. 168.
|
|
82
|
+
0. 203` and `0.9.12` became `0. 9. 12`. A `.`, `!` or `?` now ends a sentence
|
|
83
|
+
only when whitespace or the end of the text follows it.
|
|
84
|
+
- **The nightly commit message's `{accepted}` and the run metrics count the
|
|
85
|
+
proposals triage promoted.** `{accepted}`, `autoAcceptedCount` in
|
|
86
|
+
`improve_runs.metrics_json` and `improve.autoAccept.promoted` in `akm health`
|
|
87
|
+
read `gateAutoAcceptedCount`, which nothing has set since the confidence gate
|
|
88
|
+
was deleted, so they were always 0. They read the triage pre-pass's promoted
|
|
89
|
+
count, and the dead field is gone from the result type. A result row an older
|
|
90
|
+
akm wrote with the field still decodes in `akm health`: the decoder accepts
|
|
91
|
+
the key and ignores its value.
|
|
8
92
|
|
|
9
93
|
## [0.9.20] - 2026-09-30
|
|
10
94
|
|
|
@@ -94,15 +94,20 @@ akm import ./doc.md --target my-other-bundle # Route import to a named writab
|
|
|
94
94
|
akm workflow create ship-release # Create a workflow asset in the bundle
|
|
95
95
|
akm lint --type workflows # Parse and compile every .md/.yml workflow source; list every error
|
|
96
96
|
akm workflow run workflows/ship-release # Start or resume and execute the workflow
|
|
97
|
-
akm feedback skills/code-review --positive # Record that an asset helped
|
|
98
|
-
akm feedback agents/reviewer --negative --reason "wrong framework" #
|
|
97
|
+
akm feedback skills/code-review --positive # Record that an asset helped (ranks it higher; no rewrite)
|
|
98
|
+
akm feedback agents/reviewer --negative --reason "wrong framework" # Flag it for review: improve proposes a fix from the reason
|
|
99
99
|
akm feedback memories/deployment-notes --positive # Works for memories too
|
|
100
100
|
akm feedback env/prod --positive # Records env feedback without surfacing values
|
|
101
101
|
```
|
|
102
102
|
|
|
103
103
|
Use `akm feedback` whenever an asset's content materially helps, or proves wrong,
|
|
104
|
-
stale or unhelpful, so future search ranking can learn from actual usage.
|
|
105
|
-
|
|
104
|
+
stale or unhelpful, so future search ranking can learn from actual usage.
|
|
105
|
+
`akm feedback <ref> --negative --reason "<what is wrong and what should change>"`
|
|
106
|
+
flags the asset for review: the next improve run proposes a fix based on your
|
|
107
|
+
reason, so be specific. `--positive` records that an asset helped (it raises its
|
|
108
|
+
ranking) and does not trigger a rewrite; improve no longer rewrites assets from
|
|
109
|
+
positive signals or on a proactive cadence. An akm command that fails says
|
|
110
|
+
nothing about the asset; don't record it as feedback.
|
|
106
111
|
|
|
107
112
|
## LLM Wiki bundles
|
|
108
113
|
|
|
@@ -8,7 +8,7 @@ For any task, follow this loop:
|
|
|
8
8
|
1. `akm curate "<task>"` — find the best matching asset
|
|
9
9
|
2. `akm show <ref>` — read the schema (field names and structure)
|
|
10
10
|
3. Edit the workspace file using schema field names + task-specific values from your README
|
|
11
|
-
4. `akm feedback <ref> --positive` — record that the asset helped
|
|
11
|
+
4. `akm feedback <ref> --positive` — record that the asset helped (it raises its ranking and does not trigger a rewrite); when its content was wrong, stale or unhelpful, `akm feedback <ref> --negative --reason "<what is wrong and what should change>"` flags it for review: the next improve run proposes a fix based on your reason, so be specific. A failed akm command (e.g. `akm show` erroring) is not feedback on the asset — don't record it.
|
|
12
12
|
|
|
13
13
|
For workflow tasks:
|
|
14
14
|
1. `akm show workflows/<name>` — inspect the procedure before executing it
|
|
@@ -39,7 +39,8 @@ akm import ./doc.md --target my-bundle # Route import to a named writabl
|
|
|
39
39
|
akm proposal diff skills/akm-dream # Diff proposal by ref, UUID, or 8-char prefix
|
|
40
40
|
akm proposal accept 7c115132 # Accept by UUID prefix
|
|
41
41
|
akm proposal reject skills/my-skill --reason "..." # Reject by ref
|
|
42
|
-
akm feedback <ref> --positive
|
|
42
|
+
akm feedback <ref> --positive # Record that an asset helped (ranks it higher; no rewrite)
|
|
43
|
+
akm feedback <ref> --negative --reason "..." # Flag it for review: the next improve run proposes a fix from your reason
|
|
43
44
|
akm bundle add <ref> # Add a source (npm, GitHub, git, local dir)
|
|
44
45
|
akm clone <ref> # Copy an asset to the working bundle (optional --dest arg to clone to specific location)
|
|
45
46
|
akm sync # Commit (and push if writable remote) changes in the primary bundle (--no-push to commit only)
|
|
@@ -65,8 +66,10 @@ akm search "<query>" --from registry # Search all registries (registry
|
|
|
65
66
|
|
|
66
67
|
When an asset's content meaningfully helps, or proves wrong, stale or unhelpful,
|
|
67
68
|
record that with `akm feedback` so future search ranking can learn from real
|
|
68
|
-
usage.
|
|
69
|
-
|
|
69
|
+
usage. Only negative feedback with a specific reason gets the asset reviewed and
|
|
70
|
+
fixed: improve no longer rewrites assets from positive signals or on a
|
|
71
|
+
proactive cadence. An akm command that fails says nothing about the asset;
|
|
72
|
+
don't record it as feedback.
|
|
70
73
|
|
|
71
74
|
## Error Shapes and Exit Codes
|
|
72
75
|
|
|
@@ -1,12 +1,12 @@
|
|
|
1
1
|
{
|
|
2
|
-
"description": "Standard improve pass — reflect, distill, consolidation (promotion plus the reviewed pair-pass retire/supersede proposals), and validation. Memory inference is listed below but only runs when experimental.improveAutonomy is set; improve-stage extract and proactive maintenance off.",
|
|
2
|
+
"description": "Standard improve pass — reflect (rewrites only from negative feedback), distill, consolidation (promotion plus the reviewed pair-pass retire/supersede proposals), and validation. Memory inference is listed below but only runs when experimental.improveAutonomy is set; improve-stage extract and proactive maintenance off.",
|
|
3
3
|
"processes": {
|
|
4
4
|
"reflect": {
|
|
5
5
|
"enabled": true,
|
|
6
6
|
"limit": 25,
|
|
7
7
|
"allowedTypes": ["agent", "command", "knowledge", "lesson", "memory", "skill", "workflow"]
|
|
8
8
|
},
|
|
9
|
-
"distill": { "enabled": true, "allowedTypes": ["memory"], "requirePlannedRefs":
|
|
9
|
+
"distill": { "enabled": true, "allowedTypes": ["memory"], "requirePlannedRefs": false },
|
|
10
10
|
"consolidate": { "enabled": true, "allowedTypes": ["memory"] },
|
|
11
11
|
"memoryInference": { "enabled": true },
|
|
12
12
|
"extract": { "enabled": false, "triage": { "enabled": true, "minScore": 2 } },
|
|
@@ -1,5 +1,5 @@
|
|
|
1
1
|
{
|
|
2
|
-
"description": "Opt-in proactive-maintenance pass — reflect, distill, proposal triage (promote, high budget), and the proactive-maintenance lane (maxPerRun 100); consolidate/memoryInference/extract off. Sync disabled: an interrupted run would otherwise leave an uncommitted backlog.",
|
|
2
|
+
"description": "Opt-in proactive-maintenance pass — reflect (on negative feedback), distill, proposal triage (promote, high budget), and the proactive-maintenance lane (maxPerRun 100), which picks due assets for scoring but plans no rewrite; consolidate/memoryInference/extract off. Sync disabled: an interrupted run would otherwise leave an uncommitted backlog.",
|
|
3
3
|
"processes": {
|
|
4
4
|
"reflect": {
|
|
5
5
|
"enabled": true,
|
|
@@ -77,9 +77,13 @@ akm search "<query>" --type skill
|
|
|
77
77
|
### Recording feedback and new knowledge
|
|
78
78
|
|
|
79
79
|
```sh
|
|
80
|
-
# Mark an asset as helpful (
|
|
80
|
+
# Mark an asset as helpful (raises its ranking; does not trigger a rewrite)
|
|
81
81
|
akm feedback <ref> --positive
|
|
82
82
|
|
|
83
|
+
# Flag an asset for review: the next improve run proposes a fix based on your
|
|
84
|
+
# reason, so be specific about what is wrong and what should change
|
|
85
|
+
akm feedback <ref> --negative --reason "<what is wrong and what should change>"
|
|
86
|
+
|
|
83
87
|
# Capture a durable lesson or memory from the current session
|
|
84
88
|
akm remember "<fact or lesson>"
|
|
85
89
|
```
|
|
@@ -190,6 +190,10 @@ export const feedbackCommand = defineJsonCommand({
|
|
|
190
190
|
meta: {
|
|
191
191
|
name: "feedback",
|
|
192
192
|
description: "Record positive or negative feedback for any indexed bundle asset.\n\n" +
|
|
193
|
+
'`akm feedback <ref> --negative --reason "<what is wrong and what should change>"` flags\n' +
|
|
194
|
+
"the asset for review: the next improve run proposes a fix based on your reason, so be\n" +
|
|
195
|
+
"specific. `--positive` records that an asset helped (it raises its ranking) and does not\n" +
|
|
196
|
+
"trigger a rewrite.\n\n" +
|
|
193
197
|
"Both signals adjust the asset's usefulness score right away, in the same\n" +
|
|
194
198
|
"process: positive feedback raises it, negative lowers it, and recent\n" +
|
|
195
199
|
"feedback counts for more than old feedback. No reindex is needed — the new\n" +
|
|
@@ -200,15 +204,19 @@ export const feedbackCommand = defineJsonCommand({
|
|
|
200
204
|
// and throw a structured UsageError below so exit code is 2 (USAGE) rather
|
|
201
205
|
// than citty's default 0 (help banner).
|
|
202
206
|
ref: { type: "positional", description: "Asset ref ([bundle//]conceptId, e.g. lessons/deploy)", required: false },
|
|
203
|
-
positive: {
|
|
207
|
+
positive: {
|
|
208
|
+
type: "boolean",
|
|
209
|
+
description: "Record that the asset helped (raises its ranking immediately; does not trigger a rewrite)",
|
|
210
|
+
default: false,
|
|
211
|
+
},
|
|
204
212
|
negative: {
|
|
205
213
|
type: "boolean",
|
|
206
|
-
description: "
|
|
214
|
+
description: "Flag the asset for review: the next improve run proposes a fix from --reason (also lowers its ranking immediately, no reindex needed).",
|
|
207
215
|
default: false,
|
|
208
216
|
},
|
|
209
217
|
reason: {
|
|
210
218
|
type: "string",
|
|
211
|
-
description: "What
|
|
219
|
+
description: "What is wrong with the asset's content and what should change; the next improve run proposes a fix from it, so be specific (required for negative feedback by default). Not for akm command errors.",
|
|
212
220
|
},
|
|
213
221
|
"failure-mode": {
|
|
214
222
|
type: "string",
|
|
@@ -260,12 +268,12 @@ export const feedbackCommand = defineJsonCommand({
|
|
|
260
268
|
const cfg = loadConfig();
|
|
261
269
|
const requireReason = cfg.feedback?.requireReason ?? true; // Default: true (F-3 / #384)
|
|
262
270
|
if (requireReason) {
|
|
263
|
-
throw new UsageError("Negative feedback requires --reason
|
|
271
|
+
throw new UsageError("Negative feedback requires --reason: the next improve run proposes a fix from it, so say what is wrong and what should change. " +
|
|
264
272
|
"Use --failure-mode for a curated taxonomy or --reason for free text. " +
|
|
265
|
-
"Set feedback.requireReason: false in akm.json to downgrade to a warning.", "MISSING_REQUIRED_ARGUMENT", `Hint: akm feedback ${ref} --negative --reason "
|
|
273
|
+
"Set feedback.requireReason: false in akm.json to downgrade to a warning.", "MISSING_REQUIRED_ARGUMENT", `Hint: akm feedback ${ref} --negative --reason "<what is wrong and what should change>" [--failure-mode incorrect|outdated|dangerous|incomplete|redundant]`);
|
|
266
274
|
}
|
|
267
275
|
else {
|
|
268
|
-
warn("Warning: negative feedback without --reason
|
|
276
|
+
warn("Warning: negative feedback without --reason gives the next improve run nothing to base a fix on.");
|
|
269
277
|
}
|
|
270
278
|
}
|
|
271
279
|
const rawTags = parseAllFlagValues("--tag");
|
|
@@ -190,7 +190,7 @@ function projectRunMetrics(result) {
|
|
|
190
190
|
(metrics.actions.distill.skippedByReason[reason] ?? 0) + toFiniteNumber(count);
|
|
191
191
|
}
|
|
192
192
|
}
|
|
193
|
-
metrics.autoAccept.promoted += toFiniteNumber(result.
|
|
193
|
+
metrics.autoAccept.promoted += toFiniteNumber(result.triage?.promoted);
|
|
194
194
|
metrics.autoAccept.validationFailed += toFiniteNumber(result.gateAutoAcceptFailedCount);
|
|
195
195
|
const memorySummary = result.memorySummary;
|
|
196
196
|
if (memorySummary) {
|
|
@@ -434,6 +434,7 @@ function qualityGateEnabled(run) {
|
|
|
434
434
|
async function judgeAndQueue(run, out) {
|
|
435
435
|
let content = out.content;
|
|
436
436
|
let confidence;
|
|
437
|
+
let judged;
|
|
437
438
|
if (qualityGateEnabled(run)) {
|
|
438
439
|
const similarLessons = await run.similar(content.slice(0, 500), 3);
|
|
439
440
|
// The judge reads what the generator read: the source body, without its frontmatter (buildDistillPrompt).
|
|
@@ -452,6 +453,7 @@ async function judgeAndQueue(run, out) {
|
|
|
452
453
|
}
|
|
453
454
|
if (verdict.score > 0)
|
|
454
455
|
confidence = verdict.score / 5;
|
|
456
|
+
judged = verdict;
|
|
455
457
|
}
|
|
456
458
|
let frontmatter;
|
|
457
459
|
if (out.promotion) {
|
|
@@ -489,7 +491,7 @@ async function judgeAndQueue(run, out) {
|
|
|
489
491
|
...(run.options.eligibilitySource ? { eligibilitySource: run.options.eligibilitySource } : {}),
|
|
490
492
|
// The ledger keys the attempt by the input, not the output.
|
|
491
493
|
attemptedRefs: [run.ledgerRef],
|
|
492
|
-
}, { judged
|
|
494
|
+
}, { judged });
|
|
493
495
|
persistOutputEncodingSalience(run, out.ref, content);
|
|
494
496
|
const swapped = out.descriptionSwapped ? { descriptionSwapped: out.descriptionSwapped } : {};
|
|
495
497
|
emitDistill(run, {
|
|
@@ -220,7 +220,7 @@ export const improveCommand = defineCommand({
|
|
|
220
220
|
},
|
|
221
221
|
"require-feedback-signal": {
|
|
222
222
|
type: "boolean",
|
|
223
|
-
description: "
|
|
223
|
+
description: "Turn the proactive/high-salience fallback lanes off (they only select and score assets; a rewrite needs negative feedback)",
|
|
224
224
|
default: false,
|
|
225
225
|
},
|
|
226
226
|
"json-to-stdout": {
|
|
@@ -9,7 +9,7 @@
|
|
|
9
9
|
import fs from "node:fs";
|
|
10
10
|
import path from "node:path";
|
|
11
11
|
import { parseRefInput } from "../../core/asset/resolve-ref.js";
|
|
12
|
-
import { bundlesToSourceEntries, loadConfig
|
|
12
|
+
import { bundlesToSourceEntries, loadConfig } from "../../core/config/config.js";
|
|
13
13
|
import { ConfigError, rethrowIfTestIsolationError, UsageError } from "../../core/errors.js";
|
|
14
14
|
import { appendEvent, readEvents } from "../../core/events.js";
|
|
15
15
|
import { classifyImproveAction, foldDistillSkipped } from "../../core/improve-types.js";
|
|
@@ -39,13 +39,11 @@ import { akmDistill } from "./distill.js";
|
|
|
39
39
|
import { collectEligibleRefs, collectEligibleRefsReadOnly, memoryCleanupParentRef, resolveImproveScope, shouldAnalyzeMemoryCleanup, } from "./eligibility.js";
|
|
40
40
|
import { eligibleRefCount, projectResolvedProcessRouting, resolveImprovePlan, resolveImproveStrategy, } from "./improve-strategies.js";
|
|
41
41
|
import { buildImproveUsageReport } from "./improve-usage-report.js";
|
|
42
|
-
import { lastAttemptByRef, loadLedgerSnapshot } from "./ledger.js";
|
|
43
42
|
import { improveLockPath, releaseImproveLock, tryAcquireImproveLock } from "./locks.js";
|
|
44
43
|
import { runImproveLoopStage, runImprovePostLoopStage } from "./loop-stages.js";
|
|
45
44
|
import { analyzeMemoryCleanup, purgeGracedArchive, RETIRE_GRACE_DAYS, } from "./memory/memory-improve.js";
|
|
46
45
|
import { buildImproveExecutionPlan } from "./planner.js";
|
|
47
46
|
import { CONSOLIDATION_CONFIG_KEYS, pickDefined, recordImproveSkip, runImprovePreparationStage } from "./preparation.js";
|
|
48
|
-
import { DEFAULT_DUE_DAYS, filterProactiveDue } from "./proactive-maintenance.js";
|
|
49
47
|
import { akmReflect } from "./reflect.js";
|
|
50
48
|
import { errMessage, noticeSet } from "./stage.js";
|
|
51
49
|
export { runImproveMaintenancePasses } from "./loop-stages.js";
|
|
@@ -57,7 +55,7 @@ export function renderSyncCommitMessage(template, result, nowMs) {
|
|
|
57
55
|
time: iso.slice(11, 19),
|
|
58
56
|
scope: result.scope.value ?? result.scope.mode,
|
|
59
57
|
refs: String(result.plannedRefs.length),
|
|
60
|
-
accepted: String(result.
|
|
58
|
+
accepted: String(result.triage?.promoted ?? 0),
|
|
61
59
|
triage_promoted: String(result.triage?.promoted ?? 0),
|
|
62
60
|
triage_rejected: String(result.triage?.rejected ?? 0),
|
|
63
61
|
runId: result.runId ?? "",
|
|
@@ -758,25 +756,6 @@ function makeCommitStashBatch(deps) {
|
|
|
758
756
|
}
|
|
759
757
|
};
|
|
760
758
|
}
|
|
761
|
-
/**
|
|
762
|
-
* Re-read the improve ledger under the lock and drop proactive refs another run
|
|
763
|
-
* attempted after this one planned.
|
|
764
|
-
*/
|
|
765
|
-
export function refilterProactiveLoopRefs(loopRefs, improveProfile, ledgerAccess) {
|
|
766
|
-
const proactiveLoopRefs = loopRefs.filter((r) => r.eligibilitySource === "proactive");
|
|
767
|
-
if (proactiveLoopRefs.length === 0 || !ledgerAccess.stashDir)
|
|
768
|
-
return loopRefs;
|
|
769
|
-
const ledger = loadLedgerSnapshot({ eventsCtx: ledgerAccess.eventsCtx }, ledgerAccess.stashDir, [
|
|
770
|
-
"reflect",
|
|
771
|
-
"distill",
|
|
772
|
-
]);
|
|
773
|
-
const stillDue = new Set(filterProactiveDue(proactiveLoopRefs, lastAttemptByRef(ledger, "reflect", proactiveLoopRefs), lastAttemptByRef(ledger, "distill", proactiveLoopRefs), improveProfile.processes?.proactiveMaintenance?.dueDays ?? DEFAULT_DUE_DAYS, Date.now()).map((r) => r.ref));
|
|
774
|
-
const dropped = proactiveLoopRefs.filter((r) => !stillDue.has(r.ref));
|
|
775
|
-
if (dropped.length === 0)
|
|
776
|
-
return loopRefs;
|
|
777
|
-
info(`[improve] post-lock cooldown re-filter: dropped ${dropped.length} proactive ref(s) claimed by concurrent run (${dropped.map((r) => r.ref).join(", ")})`);
|
|
778
|
-
return loopRefs.filter((r) => r.eligibilitySource !== "proactive" || stillDue.has(r.ref));
|
|
779
|
-
}
|
|
780
759
|
/**
|
|
781
760
|
* The audit events for refs and lanes this run will not touch, then
|
|
782
761
|
* preparation → loop → post-loop. No post-loop work starts past the budget; the
|
|
@@ -824,10 +803,7 @@ async function runImproveStageSequence(run, collected, preEnsureCleanupWarnings,
|
|
|
824
803
|
options,
|
|
825
804
|
reflectFn: run.reflectFn,
|
|
826
805
|
distillFn: run.distillFn,
|
|
827
|
-
loopRefs:
|
|
828
|
-
stashDir: primaryStashDir ?? options.stashDir,
|
|
829
|
-
eventsCtx,
|
|
830
|
-
}),
|
|
806
|
+
loopRefs: preparation.loopRefs,
|
|
831
807
|
actions: preparation.actions,
|
|
832
808
|
signalBearingSet: preparation.signalBearingSet,
|
|
833
809
|
distillCooledRefs: preparation.distillCooledRefs,
|
|
@@ -8,10 +8,11 @@
|
|
|
8
8
|
*
|
|
9
9
|
* Candidate selection reads the improve ledger plus one set of signals: a ref
|
|
10
10
|
* is eligible for a source when feedback newer than its last attempt landed and
|
|
11
|
-
* no ledger window holds it
|
|
12
|
-
* by the fallback lanes (proactive
|
|
13
|
-
*
|
|
14
|
-
*
|
|
11
|
+
* no ledger window holds it, and reflect reads only negative feedback. Refs
|
|
12
|
+
* without recent feedback can still be picked by the fallback lanes (proactive
|
|
13
|
+
* maintenance, high salience), which only score them; the survivors are ranked
|
|
14
|
+
* by salience, checked on disk and capped. A plan-only run evaluates the same
|
|
15
|
+
* selectors against read snapshots and writes nothing.
|
|
15
16
|
*/
|
|
16
17
|
import fs from "node:fs";
|
|
17
18
|
import path from "node:path";
|
|
@@ -549,6 +550,7 @@ export function buildSnapshotManifest(args) {
|
|
|
549
550
|
const candidates = args.postCleanupRefs.filter((r) => !args.validationFailureRefs.has(r.ref));
|
|
550
551
|
const refByKey = new Map(candidates.map((r) => [keyOf(r), r.ref]));
|
|
551
552
|
const latestFeedbackTs = new Map();
|
|
553
|
+
const latestNegativeTs = new Map();
|
|
552
554
|
const feedback = new Map(candidates.map((r) => [r.ref, { hasSignal: false, positive: 0, negative: 0 }]));
|
|
553
555
|
if (candidates.length > 0) {
|
|
554
556
|
for (const e of readEvents({ type: "feedback" }, eventsCtx).events) {
|
|
@@ -557,12 +559,14 @@ export function buildSnapshotManifest(args) {
|
|
|
557
559
|
if (!ref || !entry)
|
|
558
560
|
continue;
|
|
559
561
|
const ts = e.ts ?? "";
|
|
562
|
+
const signal = e.metadata?.signal;
|
|
560
563
|
if (ts >= feedbackSinceCutoff && isSignalEvent(e.metadata)) {
|
|
561
564
|
entry.hasSignal = true;
|
|
562
565
|
if (ts > (latestFeedbackTs.get(ref) ?? ""))
|
|
563
566
|
latestFeedbackTs.set(ref, ts);
|
|
567
|
+
if (signal === "negative" && ts > (latestNegativeTs.get(ref) ?? ""))
|
|
568
|
+
latestNegativeTs.set(ref, ts);
|
|
564
569
|
}
|
|
565
|
-
const signal = e.metadata?.signal;
|
|
566
570
|
if (signal === "positive")
|
|
567
571
|
entry.positive++;
|
|
568
572
|
else if (signal === "negative")
|
|
@@ -576,6 +580,7 @@ export function buildSnapshotManifest(args) {
|
|
|
576
580
|
feedbackSinceCutoff,
|
|
577
581
|
nowIso: new Date().toISOString(),
|
|
578
582
|
latestFeedbackTs,
|
|
583
|
+
latestNegativeTs,
|
|
579
584
|
ledger,
|
|
580
585
|
lastReflectAttemptAt: lastAttemptByRef(ledger, "reflect", candidates),
|
|
581
586
|
lastDistillAttemptAt: lastAttemptByRef(ledger, "distill", candidates),
|
|
@@ -584,19 +589,21 @@ export function buildSnapshotManifest(args) {
|
|
|
584
589
|
}
|
|
585
590
|
/**
|
|
586
591
|
* Partition the post-cleanup refs against the ledger:
|
|
587
|
-
* - eligibleRefs: reflect's signal delta passes
|
|
588
|
-
*
|
|
592
|
+
* - eligibleRefs: reflect's signal delta passes, which only fresh negative
|
|
593
|
+
* feedback can do (distill may still be cooled);
|
|
594
|
+
* - distillOnlyRefs: only distill's passes (any signal, a positive or a note
|
|
595
|
+
* included), on a distill candidate;
|
|
589
596
|
* - noFeedbackPool: no recent feedback and no reflect window, left to the
|
|
590
|
-
* fallback lanes;
|
|
597
|
+
* fallback lanes, which only score them;
|
|
591
598
|
* - fullySkippedCount: feedback on record but nothing new, or a live window.
|
|
592
599
|
* An explicit `--scope <ref>` bypasses every gate.
|
|
593
600
|
*/
|
|
594
601
|
export function partitionBySignalDelta(args) {
|
|
595
602
|
const { postCleanupRefs, validationFailureRefs } = args;
|
|
596
|
-
const { latestFeedbackTs, ledger, nowIso } = args.snapshot;
|
|
597
|
-
// Newer feedback lifts a revisit window, never a rejection.
|
|
603
|
+
const { latestFeedbackTs, latestNegativeTs, ledger, nowIso } = args.snapshot;
|
|
604
|
+
// Newer feedback lifts a revisit window, never a rejection. Reflect reads negative feedback only.
|
|
598
605
|
const deltaPasses = (candidate, source) => {
|
|
599
|
-
const feedbackAt = latestFeedbackTs.get(candidate.ref);
|
|
606
|
+
const feedbackAt = (source === "reflect" ? latestNegativeTs : latestFeedbackTs).get(candidate.ref);
|
|
600
607
|
if (!feedbackAt)
|
|
601
608
|
return false;
|
|
602
609
|
const row = ledgerRowFor(ledger, source, candidate.ref, candidate.itemRef);
|
|
@@ -639,8 +646,8 @@ export function partitionBySignalDelta(args) {
|
|
|
639
646
|
}
|
|
640
647
|
/**
|
|
641
648
|
* Pick the loop's refs: signal delta, the fallback lanes (unless
|
|
642
|
-
* `--require-feedback-signal
|
|
643
|
-
* no-op-dampened ranking, the disk check and the limit.
|
|
649
|
+
* `--require-feedback-signal`; they score, never plan), lane attribution,
|
|
650
|
+
* salience, the no-op-dampened ranking, the disk check and the limit.
|
|
644
651
|
*/
|
|
645
652
|
async function selectLoopCandidates(args, postCleanupRefs, validationFailureRefs, actions, persist) {
|
|
646
653
|
const { scope, options, primaryStashDir, eventsCtx, improveProfile } = args;
|
|
@@ -678,16 +685,13 @@ async function selectLoopCandidates(args, postCleanupRefs, validationFailureRefs
|
|
|
678
685
|
const highSalienceRefs = allowFallbacks
|
|
679
686
|
? selectHighSalienceLane(options, improveProfile, eventsCtx, noFeedbackCandidates.filter((r) => !proactive.proactiveRefs.some((p) => p.ref === r.ref)), snapshot.lastReflectAttemptAt, persist)
|
|
680
687
|
: [];
|
|
681
|
-
// An explicit ref scope always acts on its ref; otherwise
|
|
688
|
+
// An explicit ref scope always acts on its ref; otherwise only feedback plans the loop. The fallback lanes'
|
|
689
|
+
// picks are scored with it, never planned: a rewrite needs negative feedback.
|
|
682
690
|
const signalAndRetrievalRefs = dedupeRefs([...signalFiltered, ...proactive.proactiveRefs, ...highSalienceRefs]);
|
|
683
|
-
const mergedRefs = scope.mode === "ref" ? processableRefs :
|
|
684
|
-
|
|
685
|
-
//
|
|
691
|
+
const mergedRefs = scope.mode === "ref" ? processableRefs : signalFiltered;
|
|
692
|
+
const scoredRefs = scope.mode === "ref" ? processableRefs : signalAndRetrievalRefs;
|
|
693
|
+
// Lane attribution: signal-delta, and an explicit ref scope over it.
|
|
686
694
|
const sourceByRef = new Map();
|
|
687
|
-
for (const r of highSalienceRefs)
|
|
688
|
-
sourceByRef.set(r.ref, "high-salience");
|
|
689
|
-
for (const r of proactive.proactiveRefs)
|
|
690
|
-
sourceByRef.set(r.ref, "proactive");
|
|
691
695
|
for (const r of signalFiltered)
|
|
692
696
|
sourceByRef.set(r.ref, "signal-delta");
|
|
693
697
|
if (scope.mode === "ref")
|
|
@@ -695,7 +699,7 @@ async function selectLoopCandidates(args, postCleanupRefs, validationFailureRefs
|
|
|
695
699
|
sourceByRef.set(r.ref, "scope");
|
|
696
700
|
for (const r of mergedRefs)
|
|
697
701
|
r.eligibilitySource = sourceByRef.get(r.ref) ?? "unknown";
|
|
698
|
-
const salienceMap = scoreSalience(args,
|
|
702
|
+
const salienceMap = scoreSalience(args, scoredRefs, snapshot.feedback, retrieval.retrievalCounts, persist);
|
|
699
703
|
// Rank by salience; a ref skipped as a no-op repeatedly sorts lower (its stored rank is untouched).
|
|
700
704
|
const noOps = new Map();
|
|
701
705
|
withRunState(eventsCtx, persist, (db) => {
|
|
@@ -724,8 +728,7 @@ async function selectLoopCandidates(args, postCleanupRefs, validationFailureRefs
|
|
|
724
728
|
const deferred = actionableRefs.length - selection.loopRefs.length;
|
|
725
729
|
info(`[improve] ${actionableRefs.length} actionable; ${selection.loopRefs.length} will be processed` +
|
|
726
730
|
(options.limit && deferred > 0 ? ` (--limit ${options.limit} applied; ${deferred} deferred)` : ""));
|
|
727
|
-
// Skip observability
|
|
728
|
-
// survivors, so a rescued ref is never also reported skipped.
|
|
731
|
+
// Skip observability: every candidate that did not survive into the loop pool is reported skipped.
|
|
729
732
|
const survivors = new Set(sorted.map((c) => c.ref));
|
|
730
733
|
const skipped = fallbackEligible.filter((c) => !survivors.has(c.ref));
|
|
731
734
|
const retrievalSkipped = skipped.filter((c) => outOfScope.has(c.ref));
|
|
@@ -770,7 +773,7 @@ async function selectLoopCandidates(args, postCleanupRefs, validationFailureRefs
|
|
|
770
773
|
{
|
|
771
774
|
name: "signal",
|
|
772
775
|
removed: signalSkipped.length,
|
|
773
|
-
reason: "no fresh signal since the last attempt
|
|
776
|
+
reason: "no fresh negative feedback (or, for distill, any signal) since the last attempt, or a ledger window holds it",
|
|
774
777
|
},
|
|
775
778
|
{ name: "disk", removed: missing.length, reason: "backing asset is absent on disk" },
|
|
776
779
|
{ name: "limit", removed: selection.limitRemoved, reason: "deferred by the effective run limit" },
|
|
@@ -803,9 +806,11 @@ function fetchRetrievalSignals(options, signalFiltered, noFeedbackCandidates, ev
|
|
|
803
806
|
return out;
|
|
804
807
|
}
|
|
805
808
|
/**
|
|
806
|
-
* Proactive maintenance (default off, whole-stash/type runs):
|
|
807
|
-
* assets
|
|
808
|
-
*
|
|
809
|
+
* Proactive maintenance (default off, whole-stash/type runs): pick stable
|
|
810
|
+
* assets due for a revisit. The picks are scored and never planned: improve
|
|
811
|
+
* does not rewrite on a proactive cadence. The due gate doubles as the
|
|
812
|
+
* rotation cooldown: a freshly reflected asset waits `dueDays` before it is
|
|
813
|
+
* picked again.
|
|
809
814
|
*/
|
|
810
815
|
function selectProactiveMaintenanceLane(args, candidates, snapshot, retrieval, persist) {
|
|
811
816
|
if (args.scope.mode === "ref" || !args.resolvedPlan.processes.proactiveMaintenance.enabled) {
|
|
@@ -856,7 +861,8 @@ function selectProactiveMaintenanceLane(args, candidates, snapshot, retrieval, p
|
|
|
856
861
|
/**
|
|
857
862
|
* High salience: zero-feedback refs whose content-derived encoding score (not
|
|
858
863
|
* a per-type stub) reaches `salienceThreshold` and that were never reflected,
|
|
859
|
-
* top-N by score, capped at 10% of the effective limit.
|
|
864
|
+
* top-N by score, capped at 10% of the effective limit. The picks are scored,
|
|
865
|
+
* never planned.
|
|
860
866
|
*/
|
|
861
867
|
function selectHighSalienceLane(options, improveProfile, eventsCtx, candidates, lastReflectAttemptAt, persist) {
|
|
862
868
|
const threshold = (options.config ?? loadConfig()).improve?.salience?.salienceThreshold ?? 0.75;
|
|
@@ -11,8 +11,8 @@ export const DEFAULT_MAX_PER_RUN = 25;
|
|
|
11
11
|
/** Size floor for the cost term, so tiny files don't divide by ~0. */
|
|
12
12
|
const SIZE_FLOOR_BYTES = 200;
|
|
13
13
|
/**
|
|
14
|
-
* The due gate,
|
|
15
|
-
*
|
|
14
|
+
* The due gate: never touched, or last touched more than `dueDays` ago. It
|
|
15
|
+
* doubles as the rotation cooldown.
|
|
16
16
|
*/
|
|
17
17
|
function staleness(ref, lastReflectTs, lastDistillTs, dueDays, now) {
|
|
18
18
|
const lastTouchMs = Math.max(0, Date.parse(lastReflectTs.get(ref) ?? "") || 0, Date.parse(lastDistillTs.get(ref) ?? "") || 0);
|
|
@@ -64,10 +64,3 @@ export function selectProactiveMaintenanceRefs(params) {
|
|
|
64
64
|
scored,
|
|
65
65
|
};
|
|
66
66
|
}
|
|
67
|
-
/**
|
|
68
|
-
* Re-apply the due gate under the run lock with fresh timestamps, dropping refs
|
|
69
|
-
* another run attempted after this one planned.
|
|
70
|
-
*/
|
|
71
|
-
export function filterProactiveDue(selected, lastReflectTs, lastDistillTs, dueDays, now) {
|
|
72
|
-
return selected.filter((c) => staleness(c.ref, lastReflectTs, lastDistillTs, dueDays, now).due);
|
|
73
|
-
}
|
|
@@ -979,8 +979,9 @@ async function finalizeReflectProposal(args) {
|
|
|
979
979
|
}, options.eventsCtx);
|
|
980
980
|
return reflectFailure(run, result, "quality_rejected", message, false);
|
|
981
981
|
};
|
|
982
|
+
let verdict;
|
|
982
983
|
if (judged) {
|
|
983
|
-
|
|
984
|
+
verdict = await runReflectQualityJudge(run.config, payload.content, assetContent ?? "", feedback, options.chat, {
|
|
984
985
|
runnerSelectionFrozen: true,
|
|
985
986
|
...(judge.runner ? { llmRunner: judge.runner } : {}),
|
|
986
987
|
...(Object.hasOwn(options, "timeoutMs") ? { timeoutMs: options.timeoutMs } : {}),
|
|
@@ -1045,7 +1046,7 @@ async function finalizeReflectProposal(args) {
|
|
|
1045
1046
|
...(sanitized.sizeGuardRatio ? { measured: Math.round(sanitized.sizeGuardRatio.ratio * 100) } : {}),
|
|
1046
1047
|
},
|
|
1047
1048
|
}
|
|
1048
|
-
: { judged });
|
|
1049
|
+
: { judged: verdict });
|
|
1049
1050
|
appendEvent({
|
|
1050
1051
|
eventType: "reflect_completed",
|
|
1051
1052
|
ref: proposal.ref,
|
|
@@ -120,29 +120,31 @@ export function rejectedProposalContext(stash, ref, ctx, eventsCtx) {
|
|
|
120
120
|
}
|
|
121
121
|
// ── Mint ─────────────────────────────────────────────────────────────────────
|
|
122
122
|
/**
|
|
123
|
-
* Create a stage's proposal. `judged` stamps a `staged`
|
|
124
|
-
* judged content's hash
|
|
125
|
-
*
|
|
126
|
-
* improve ledger).
|
|
123
|
+
* Create a stage's proposal. `judged` (the passing verdict) stamps a `staged`
|
|
124
|
+
* gate decision with the judged content's hash and the judge's scores and reason
|
|
125
|
+
* (the triage drain accepts it while the content still matches); `review`
|
|
126
|
+
* leaves it `deferred` for a human (`review_needed` in the improve ledger).
|
|
127
127
|
*/
|
|
128
128
|
export function mintProposal(stash, proposalsCtx, input, verdict = {}) {
|
|
129
129
|
const proposal = createProposal(stash, input, proposalsCtx);
|
|
130
130
|
if (verdict.review) {
|
|
131
131
|
return recordGateDecision(stash, proposal.id, { outcome: "deferred", ...verdict.review }, proposalsCtx) ?? proposal;
|
|
132
132
|
}
|
|
133
|
-
return verdict.judged ? stageJudgedProposal(stash, proposal, proposalsCtx) : proposal;
|
|
133
|
+
return verdict.judged ? stageJudgedProposal(stash, proposal, verdict.judged, proposalsCtx) : proposal;
|
|
134
134
|
}
|
|
135
135
|
/**
|
|
136
136
|
* Stamp a proposal the quality judge passed. Best-effort: a failed stamp only
|
|
137
137
|
* means the triage drain judges it again.
|
|
138
138
|
*/
|
|
139
|
-
export function stageJudgedProposal(stash, proposal, proposalsCtx) {
|
|
139
|
+
export function stageJudgedProposal(stash, proposal, judged, proposalsCtx) {
|
|
140
140
|
try {
|
|
141
141
|
return (recordGateDecision(stash, proposal.id, {
|
|
142
142
|
outcome: "staged",
|
|
143
143
|
reason: "quality-judge",
|
|
144
144
|
gate: "quality-gate",
|
|
145
145
|
contentHash: proposalContentHash(proposal),
|
|
146
|
+
...(judged?.criteria ? { scores: judged.criteria } : {}),
|
|
147
|
+
...(judged ? { judgeReason: judged.reason } : {}),
|
|
146
148
|
}, proposalsCtx) ?? proposal);
|
|
147
149
|
}
|
|
148
150
|
catch (error) {
|
|
@@ -196,17 +198,15 @@ function buildChangedRegion(sourceContent, candidateContent) {
|
|
|
196
198
|
const added = candidate.slice(prefix, candidate.length - suffix).join("\n");
|
|
197
199
|
return boundedDocument(`Removed or replaced:\n${removed || "(none)"}\n\nAdded or replacement:\n${added || "(none)"}`);
|
|
198
200
|
}
|
|
199
|
-
/** Judge prompt for an in-place revision
|
|
201
|
+
/** Judge prompt for an in-place revision. */
|
|
200
202
|
export function buildReflectJudgePrompt(candidateContent, sourceContent, feedback) {
|
|
201
203
|
return [
|
|
202
204
|
"You are evaluating a proposed revision to an existing akm asset.",
|
|
203
205
|
"",
|
|
204
206
|
"Score this revision on each criterion from 1 (poor) to 5 (excellent):",
|
|
205
|
-
"1.
|
|
206
|
-
"2. PRESERVATION: Does it
|
|
207
|
-
"3. QUALITY: Is
|
|
208
|
-
"",
|
|
209
|
-
"Overlap with the source is expected and must not lower the score by itself; this is an in-place revision, not a new lesson.",
|
|
207
|
+
"1. NEED: Does the revision fix a concrete problem in the source? Concrete problems are: something the feedback reports as wrong or missing; a factual error; or broken, garbled, truncated or missing text, including frontmatter fields such as description or when_to_use. Score 4-5 when it fixes one, even a small one. Score 1-2 when the source was already correct and the revision only rewords, restates, reformats, or adds headings, an introduction or a table of contents.",
|
|
208
|
+
"2. PRESERVATION: Does it keep every concrete fact, identifier, command, path, number and example from the source, without truncation?",
|
|
209
|
+
"3. QUALITY: Is it coherent and accurate, with no claims, steps or details that the source or the feedback does not support?",
|
|
210
210
|
"",
|
|
211
211
|
"Feedback:",
|
|
212
212
|
"```",
|
|
@@ -228,13 +228,13 @@ export function buildReflectJudgePrompt(candidateContent, sourceContent, feedbac
|
|
|
228
228
|
buildChangedRegion(sourceContent, candidateContent),
|
|
229
229
|
"```",
|
|
230
230
|
"",
|
|
231
|
-
'Return ONLY valid JSON, no prose: {"scores": {"
|
|
231
|
+
'Return ONLY valid JSON, no prose: {"scores": {"need": <1-5 integer>, "preservation": <1-5 integer>, "quality": <1-5 integer>}, "reason": "<one sentence>"}',
|
|
232
232
|
].join("\n");
|
|
233
233
|
}
|
|
234
234
|
/**
|
|
235
235
|
* `grounding` is scored with the other lesson criteria but left out of their
|
|
236
236
|
* mean: a lesson about a different subject than its source reads as novel and
|
|
237
|
-
* non-redundant, so
|
|
237
|
+
* non-redundant, so they would pass it (or, in the review band, mint it as
|
|
238
238
|
* a pending proposal). The rubric reserves 1-2 for a different subject. A score
|
|
239
239
|
* of {@link UNGROUNDED_MAX_SCORE} or less is a rejection whatever the mean says
|
|
240
240
|
* (#999). A higher score up to {@link BORDERLINE_GROUNDING_MAX_SCORE} is only
|
|
@@ -250,12 +250,12 @@ const GROUNDING_CRITERION = "grounding";
|
|
|
250
250
|
const UNGROUNDED_MAX_SCORE = 1;
|
|
251
251
|
const BORDERLINE_GROUNDING_MAX_SCORE = 2;
|
|
252
252
|
const LESSON_JUDGE_CRITERIA = ["novelty", "nonRedundancy", GROUNDING_CRITERION];
|
|
253
|
-
const REFLECT_JUDGE_CRITERIA = ["
|
|
253
|
+
const REFLECT_JUDGE_CRITERIA = ["need", "preservation", "quality"];
|
|
254
254
|
/**
|
|
255
255
|
* Read a judge response: the per-criterion shape (averaged here, `grounding`
|
|
256
|
-
* aside) or the older `{"score"}`
|
|
257
|
-
* any missing or out-of-range
|
|
258
|
-
* ignored.
|
|
256
|
+
* aside; `lowest` is the lowest score in that mean) or the older `{"score"}`
|
|
257
|
+
* shape. Only the expected criteria are read; any missing or out-of-range
|
|
258
|
+
* (1..5) value is a parse failure, extra keys are ignored.
|
|
259
259
|
*/
|
|
260
260
|
function parseJudgeResponse(raw, keys) {
|
|
261
261
|
const parsed = parseEmbeddedJsonResponse(raw);
|
|
@@ -277,9 +277,14 @@ function parseJudgeResponse(raw, keys) {
|
|
|
277
277
|
const averaged = Object.entries(criteria)
|
|
278
278
|
.filter(([key]) => key !== GROUNDING_CRITERION)
|
|
279
279
|
.map(([, value]) => value);
|
|
280
|
-
return {
|
|
280
|
+
return {
|
|
281
|
+
score: averaged.reduce((a, b) => a + b, 0) / averaged.length,
|
|
282
|
+
lowest: Math.min(...averaged),
|
|
283
|
+
reason,
|
|
284
|
+
criteria,
|
|
285
|
+
};
|
|
281
286
|
}
|
|
282
|
-
return inRange(parsed.score) ? { score: parsed.score, reason } : undefined;
|
|
287
|
+
return inRange(parsed.score) ? { score: parsed.score, lowest: parsed.score, reason } : undefined;
|
|
283
288
|
}
|
|
284
289
|
function judgeResponseSchema(keys) {
|
|
285
290
|
return {
|
|
@@ -299,14 +304,14 @@ function judgeResponseSchema(keys) {
|
|
|
299
304
|
}
|
|
300
305
|
/**
|
|
301
306
|
* The quality judge. Fails closed: no runner, an unparseable verdict or a
|
|
302
|
-
* provider failure never passes content. Bands:
|
|
303
|
-
*
|
|
304
|
-
*
|
|
305
|
-
*
|
|
306
|
-
*
|
|
307
|
-
*
|
|
308
|
-
*
|
|
309
|
-
* margin in mind.
|
|
307
|
+
* provider failure never passes content. Bands: every criterion in the mean
|
|
308
|
+
* >= 4 passes, otherwise a mean >= 2.5 is review and a lower one reject; a
|
|
309
|
+
* `grounding` score of {@link UNGROUNDED_MAX_SCORE} or less rejects whatever
|
|
310
|
+
* the mean is, and one of {@link BORDERLINE_GROUNDING_MAX_SCORE} routes a lesson
|
|
311
|
+
* that would pass to review (a mean that rejects stays a rejection).
|
|
312
|
+
* Temperature is set to 0, which reduces run-to-run variation but does not
|
|
313
|
+
* remove it: on some servers (llama.cpp batching, for one) the same request can
|
|
314
|
+
* score a point apart, so the routing rules are chosen with that margin in mind.
|
|
310
315
|
*/
|
|
311
316
|
async function runQualityJudge(feature, config, prompt, keys, chat, options) {
|
|
312
317
|
const resolved = !options.runnerSelectionFrozen && !options.llmRunner
|
|
@@ -338,7 +343,7 @@ async function runQualityJudge(feature, config, prompt, keys, chat, options) {
|
|
|
338
343
|
const parsed = parseJudgeResponse(outcome.raw, keys);
|
|
339
344
|
if (!parsed)
|
|
340
345
|
return { pass: false, score: -1, reason: "judge parse failed — routed to review", reviewNeeded: true };
|
|
341
|
-
const { score, reason, criteria } = parsed;
|
|
346
|
+
const { score, lowest, reason, criteria } = parsed;
|
|
342
347
|
const grounding = criteria?.[GROUNDING_CRITERION];
|
|
343
348
|
if (criteria && grounding !== undefined && grounding <= UNGROUNDED_MAX_SCORE) {
|
|
344
349
|
return {
|
|
@@ -348,8 +353,8 @@ async function runQualityJudge(feature, config, prompt, keys, chat, options) {
|
|
|
348
353
|
criteria,
|
|
349
354
|
};
|
|
350
355
|
}
|
|
351
|
-
const verdict =
|
|
352
|
-
// Borderline grounding is a person's call even when the
|
|
356
|
+
const verdict = lowest >= 4 ? { pass: true } : score >= 2.5 ? { pass: false, reviewNeeded: true } : { pass: false };
|
|
357
|
+
// Borderline grounding is a person's call even when the lesson would pass; a mean that rejects stays rejected.
|
|
353
358
|
if (criteria &&
|
|
354
359
|
grounding !== undefined &&
|
|
355
360
|
grounding <= BORDERLINE_GROUNDING_MAX_SCORE &&
|
|
@@ -28,7 +28,7 @@ import { canonicalBundleIdForTarget, resolveBundleWriteTarget } from "../../core
|
|
|
28
28
|
import { getStateDbPath, withImmediateTransaction, withStateDb } from "../../core/state-db.js";
|
|
29
29
|
import { warn } from "../../core/warn.js";
|
|
30
30
|
import { recordWrittenPath } from "../../core/write-provenance.js";
|
|
31
|
-
import { assertAkmAssetWrite, commitWriteTargetBoundary, prepareWriteTargetForMutation, resolveWriteTarget, } from "../../core/write-source.js";
|
|
31
|
+
import { assertAkmAssetWrite, commitAcceptedPaths, commitWriteTargetBoundary, prepareWriteTargetForMutation, resolveWriteTarget, } from "../../core/write-source.js";
|
|
32
32
|
import { withAssetMutationLease } from "../../indexer/index-writer-lock.js";
|
|
33
33
|
import { indexWrittenAssets } from "../../indexer/index-written-assets.js";
|
|
34
34
|
import { deriveInstallations } from "../../indexer/installations.js";
|
|
@@ -708,6 +708,9 @@ function persistProposalDecision(stashDir, proposal, decision, ctx) {
|
|
|
708
708
|
...(decision.gateDecision
|
|
709
709
|
? {
|
|
710
710
|
gateDecision: {
|
|
711
|
+
// The drain's verdict replaces the quality judge's stamp; the judge's evidence stays on the row.
|
|
712
|
+
...(current.gateDecision?.scores ? { scores: current.gateDecision.scores } : {}),
|
|
713
|
+
...(current.gateDecision?.judgeReason ? { judgeReason: current.gateDecision.judgeReason } : {}),
|
|
711
714
|
...decision.gateDecision,
|
|
712
715
|
decidedAt: decision.gateDecision.decidedAt ?? decision.decidedAt,
|
|
713
716
|
},
|
|
@@ -1040,6 +1043,10 @@ function requireAcceptedTarget(proposal) {
|
|
|
1040
1043
|
}
|
|
1041
1044
|
return proposal.acceptedTarget;
|
|
1042
1045
|
}
|
|
1046
|
+
/** An accept's commit subject: the generator, the proposal id's first 8 characters and the asset ref. */
|
|
1047
|
+
function acceptCommitMessage(proposal) {
|
|
1048
|
+
return `akm accept: ${proposal.source} ${proposal.id.slice(0, 8)} ${proposal.ref}`;
|
|
1049
|
+
}
|
|
1043
1050
|
/**
|
|
1044
1051
|
* O1 (alpha.9): an accepted consolidate PROMOTION retires its source memory
|
|
1045
1052
|
* (and its `.derived` twin), so promotion no longer leaves a memory/
|
|
@@ -1059,7 +1066,7 @@ function requireAcceptedTarget(proposal) {
|
|
|
1059
1066
|
* existed carries no hash at all, so it is treated the same way: never
|
|
1060
1067
|
* archived, not verified against a hash that was never recorded.
|
|
1061
1068
|
*/
|
|
1062
|
-
function retirePromotionSource(mutationTarget, accepted) {
|
|
1069
|
+
function retirePromotionSource(mutationTarget, accepted, paths) {
|
|
1063
1070
|
if (!accepted.promotionSource)
|
|
1064
1071
|
return;
|
|
1065
1072
|
try {
|
|
@@ -1086,17 +1093,12 @@ function retirePromotionSource(mutationTarget, accepted) {
|
|
|
1086
1093
|
successorRefs: [accepted.ref],
|
|
1087
1094
|
};
|
|
1088
1095
|
const record = archiveCleanupCandidate(mutationTarget.source.path, candidate, sourcePath);
|
|
1089
|
-
|
|
1090
|
-
sourcePath,
|
|
1091
|
-
path.join(mutationTarget.source.path, record.archivedPath),
|
|
1092
|
-
path.join(mutationTarget.source.path, record.auditPath),
|
|
1093
|
-
];
|
|
1096
|
+
paths.push(sourcePath, path.join(mutationTarget.source.path, record.archivedPath), path.join(mutationTarget.source.path, record.auditPath));
|
|
1094
1097
|
const twin = derivedTwinPath(sourcePath, sourceRef.type);
|
|
1095
1098
|
if (twin) {
|
|
1096
1099
|
const twinRecord = archiveCleanupCandidate(mutationTarget.source.path, candidate, twin);
|
|
1097
1100
|
paths.push(twin, path.join(mutationTarget.source.path, twinRecord.archivedPath), path.join(mutationTarget.source.path, twinRecord.auditPath));
|
|
1098
1101
|
}
|
|
1099
|
-
commitWriteTargetBoundary(mutationTarget, `Retire promoted source ${accepted.promotionSource}`, { paths });
|
|
1100
1102
|
}
|
|
1101
1103
|
catch (error) {
|
|
1102
1104
|
warn(`[proposal] O1: failed to retire promotion source ${accepted.promotionSource} for ${accepted.id}: ${error instanceof Error ? error.message : String(error)}`);
|
|
@@ -1133,7 +1135,6 @@ async function promoteProposalWithLease(stashDir, config, id, options, ctx) {
|
|
|
1133
1135
|
const decidedAt = nowIso(ctx);
|
|
1134
1136
|
const content = preflight.stampedContent.endsWith("\n") ? preflight.stampedContent : `${preflight.stampedContent}\n`;
|
|
1135
1137
|
writeProposalAssetFile(assetPath, content);
|
|
1136
|
-
commitWriteTargetBoundary(mutationTarget, `Update ${proposalForMutation.ref}`, { paths: [assetPath] });
|
|
1137
1138
|
const accepted = persistProposalDecision(stashDir, proposalForMutation, {
|
|
1138
1139
|
operation: "accept",
|
|
1139
1140
|
target: mutationTarget,
|
|
@@ -1146,8 +1147,10 @@ async function promoteProposalWithLease(stashDir, config, id, options, ctx) {
|
|
|
1146
1147
|
decidedAt,
|
|
1147
1148
|
}, ctx);
|
|
1148
1149
|
await indexWrittenProposalAsset(mutationTarget, assetPath);
|
|
1150
|
+
const paths = [assetPath];
|
|
1149
1151
|
if (accepted.status === "accepted" && accepted.source === "consolidate")
|
|
1150
|
-
retirePromotionSource(mutationTarget, accepted);
|
|
1152
|
+
retirePromotionSource(mutationTarget, accepted, paths);
|
|
1153
|
+
commitAcceptedPaths(mutationTarget, acceptCommitMessage(accepted), paths);
|
|
1151
1154
|
return { proposal: accepted, assetPath, ref: accepted.ref };
|
|
1152
1155
|
}
|
|
1153
1156
|
/**
|
|
@@ -1406,7 +1409,7 @@ async function retireProposalWithLease(stashDir, config, proposal, options, ctx)
|
|
|
1406
1409
|
paths.push(twinPath, path.join(mutationTarget.source.path, twinRecord.archivedPath), path.join(mutationTarget.source.path, twinRecord.auditPath));
|
|
1407
1410
|
}
|
|
1408
1411
|
if (paths.length > 0)
|
|
1409
|
-
|
|
1412
|
+
commitAcceptedPaths(mutationTarget, acceptCommitMessage(proposal), paths);
|
|
1410
1413
|
if (archiveDirs.length === 0) {
|
|
1411
1414
|
// Recorded intent, but neither file is at its original location NOR
|
|
1412
1415
|
// archived under this proposal's id: something else removed the target
|
|
@@ -101,7 +101,9 @@ export function readMemoryContent(contentArg) {
|
|
|
101
101
|
* Split `text` into sentence-shaped chunks on `.`/`!`/`?`, swallowing any
|
|
102
102
|
* immediately-trailing closing quotes/brackets/repeated terminators into the
|
|
103
103
|
* same sentence (so `Alice said, "hi there."` ends the sentence at the
|
|
104
|
-
* closing quote, not the period).
|
|
104
|
+
* closing quote, not the period). A terminator ends a sentence only when
|
|
105
|
+
* whitespace or the end of the text follows, so `192.168.0.203`, `0.9.12`,
|
|
106
|
+
* `example.com` and `notes.md` stay whole.
|
|
105
107
|
*
|
|
106
108
|
* Ported from akm-eval's memory backend (`splitIntoSentences` /
|
|
107
109
|
* `firstSentencesCapped` in akm-eval/src/memory/backends/akm.ts), which
|
|
@@ -119,6 +121,10 @@ function splitIntoSentences(text) {
|
|
|
119
121
|
let end = i + 1;
|
|
120
122
|
while (end < text.length && /["'”’)\]!?.]/.test(text.charAt(end)))
|
|
121
123
|
end += 1;
|
|
124
|
+
if (end < text.length && !/\s/.test(text.charAt(end))) {
|
|
125
|
+
i = end;
|
|
126
|
+
continue;
|
|
127
|
+
}
|
|
122
128
|
sentences.push(text.slice(start, end));
|
|
123
129
|
while (end < text.length && /\s/.test(text.charAt(end)))
|
|
124
130
|
end += 1;
|
|
@@ -49,6 +49,10 @@ const COMMON_FIELDS = [
|
|
|
49
49
|
"reflectCooldownActions",
|
|
50
50
|
"reflectSkippedActions",
|
|
51
51
|
"reflectGuardRejectedActions",
|
|
52
|
+
// No longer written (nothing has set it since the confidence gate was
|
|
53
|
+
// deleted), but kept in the allow-list so `decodeImproveResult` still reads
|
|
54
|
+
// the improve_runs rows an older release wrote with it, its value ignored —
|
|
55
|
+
// AGENTS.md "Reading persisted data".
|
|
52
56
|
"gateAutoAcceptedCount",
|
|
53
57
|
"gateAutoAcceptFailedCount",
|
|
54
58
|
"triage",
|
|
@@ -494,7 +498,6 @@ function validateCommon(value) {
|
|
|
494
498
|
"reflectCooldownActions",
|
|
495
499
|
"reflectSkippedActions",
|
|
496
500
|
"reflectGuardRejectedActions",
|
|
497
|
-
"gateAutoAcceptedCount",
|
|
498
501
|
"gateAutoAcceptFailedCount",
|
|
499
502
|
]) {
|
|
500
503
|
if (value[field] !== undefined && typeof value[field] !== "number")
|
|
@@ -20,7 +20,9 @@
|
|
|
20
20
|
* Nothing commits per asset (issue #507). Callers write and delete through
|
|
21
21
|
* {@link writeAssetToSource} / {@link deleteAssetFromSource}, then fire
|
|
22
22
|
* {@link commitWriteTargetBoundary} once — or wrap a custom mutation in
|
|
23
|
-
* {@link withWriteTargetMutation}, which does both under the asset lease.
|
|
23
|
+
* {@link withWriteTargetMutation}, which does both under the asset lease. An
|
|
24
|
+
* accepted proposal commits its own paths ({@link commitAcceptedPaths}), on a
|
|
25
|
+
* filesystem stash that is a git repository too.
|
|
24
26
|
*/
|
|
25
27
|
import fs from "node:fs";
|
|
26
28
|
import path from "node:path";
|
|
@@ -394,11 +396,12 @@ export function withWriteTargetMutation(target, paths, options, mutate) {
|
|
|
394
396
|
* Commit a git target's recorded paths plus `options.paths` (absolute or
|
|
395
397
|
* repository-relative) as one commit, and push it with `--force-with-lease`
|
|
396
398
|
* when the target is writable, has an upstream, and `push !== false`. A no-op
|
|
397
|
-
* for filesystem targets. Ignored paths stay local:
|
|
398
|
-
* commit with a warning rather than failing a write
|
|
399
|
+
* for filesystem targets unless `filesystem` is set. Ignored paths stay local:
|
|
400
|
+
* they are dropped from the commit with a warning rather than failing a write
|
|
401
|
+
* that already landed.
|
|
399
402
|
*/
|
|
400
403
|
export function commitWriteTargetBoundary(target, message, options) {
|
|
401
|
-
if (target.source.kind !== "git")
|
|
404
|
+
if (target.source.kind !== "git" && !options?.filesystem)
|
|
402
405
|
return;
|
|
403
406
|
const repoDir = repoDirFor(target.source);
|
|
404
407
|
const recorded = pendingGitPaths.get(repoDir) ?? new Set();
|
|
@@ -417,11 +420,13 @@ export function commitWriteTargetBoundary(target, message, options) {
|
|
|
417
420
|
if (committable.length === 0)
|
|
418
421
|
return;
|
|
419
422
|
try {
|
|
420
|
-
saveGitStash(undefined, message, resolveWritable(target.config), {
|
|
423
|
+
const saved = saveGitStash(undefined, message, resolveWritable(target.config), {
|
|
421
424
|
repoDir,
|
|
422
425
|
paths: committable,
|
|
423
426
|
...(options?.push === undefined ? {} : { push: options.push }),
|
|
424
427
|
});
|
|
428
|
+
if (saved.reason)
|
|
429
|
+
warn(`warning: "${target.source.name}" ${saved.reason}`);
|
|
425
430
|
}
|
|
426
431
|
catch (error) {
|
|
427
432
|
if (error instanceof GitStashPushError) {
|
|
@@ -432,3 +437,21 @@ export function commitWriteTargetBoundary(target, message, options) {
|
|
|
432
437
|
throw error;
|
|
433
438
|
}
|
|
434
439
|
}
|
|
440
|
+
/**
|
|
441
|
+
* Commit exactly the paths an accepted proposal wrote or removed, when the
|
|
442
|
+
* target is a git repository: a filesystem stash with a `.git` too, which the
|
|
443
|
+
* boundary commit skips, locally (the end-of-run sync pushes). A failed commit
|
|
444
|
+
* only warns, since the accept already landed.
|
|
445
|
+
*/
|
|
446
|
+
export function commitAcceptedPaths(target, message, paths) {
|
|
447
|
+
try {
|
|
448
|
+
commitWriteTargetBoundary(target, message, {
|
|
449
|
+
paths,
|
|
450
|
+
filesystem: true,
|
|
451
|
+
...(target.source.kind === "git" ? {} : { push: false }),
|
|
452
|
+
});
|
|
453
|
+
}
|
|
454
|
+
catch (error) {
|
|
455
|
+
warn(`warning: could not commit the accept (${message}): ${error instanceof Error ? error.message : String(error)}`);
|
|
456
|
+
}
|
|
457
|
+
}
|
|
@@ -31339,14 +31339,14 @@ var consolidate_default = {
|
|
|
31339
31339
|
};
|
|
31340
31340
|
// src/assets/improve-strategies/default.json
|
|
31341
31341
|
var default_default = {
|
|
31342
|
-
description: "Standard improve pass — reflect, distill, consolidation (promotion plus the reviewed pair-pass retire/supersede proposals), and validation. Memory inference is listed below but only runs when experimental.improveAutonomy is set; improve-stage extract and proactive maintenance off.",
|
|
31342
|
+
description: "Standard improve pass — reflect (rewrites only from negative feedback), distill, consolidation (promotion plus the reviewed pair-pass retire/supersede proposals), and validation. Memory inference is listed below but only runs when experimental.improveAutonomy is set; improve-stage extract and proactive maintenance off.",
|
|
31343
31343
|
processes: {
|
|
31344
31344
|
reflect: {
|
|
31345
31345
|
enabled: true,
|
|
31346
31346
|
limit: 25,
|
|
31347
31347
|
allowedTypes: ["agent", "command", "knowledge", "lesson", "memory", "skill", "workflow"]
|
|
31348
31348
|
},
|
|
31349
|
-
distill: { enabled: true, allowedTypes: ["memory"], requirePlannedRefs:
|
|
31349
|
+
distill: { enabled: true, allowedTypes: ["memory"], requirePlannedRefs: false },
|
|
31350
31350
|
consolidate: { enabled: true, allowedTypes: ["memory"] },
|
|
31351
31351
|
memoryInference: { enabled: true },
|
|
31352
31352
|
extract: { enabled: false, triage: { enabled: true, minScore: 2 } },
|
|
@@ -31358,7 +31358,7 @@ var default_default = {
|
|
|
31358
31358
|
};
|
|
31359
31359
|
// src/assets/improve-strategies/proactive-maintenance.json
|
|
31360
31360
|
var proactive_maintenance_default = {
|
|
31361
|
-
description: "Opt-in proactive-maintenance pass — reflect, distill, proposal triage (promote, high budget), and the proactive-maintenance lane (maxPerRun 100); consolidate/memoryInference/extract off. Sync disabled: an interrupted run would otherwise leave an uncommitted backlog.",
|
|
31361
|
+
description: "Opt-in proactive-maintenance pass — reflect (on negative feedback), distill, proposal triage (promote, high budget), and the proactive-maintenance lane (maxPerRun 100), which picks due assets for scoring but plans no rewrite; consolidate/memoryInference/extract off. Sync disabled: an interrupted run would otherwise leave an uncommitted backlog.",
|
|
31362
31362
|
processes: {
|
|
31363
31363
|
reflect: {
|
|
31364
31364
|
enabled: true,
|
|
@@ -31441,7 +31441,7 @@ var thorough_default = {
|
|
|
31441
31441
|
distill: {
|
|
31442
31442
|
enabled: true,
|
|
31443
31443
|
allowedTypes: ["memory"],
|
|
31444
|
-
requirePlannedRefs:
|
|
31444
|
+
requirePlannedRefs: false
|
|
31445
31445
|
},
|
|
31446
31446
|
consolidate: {
|
|
31447
31447
|
enabled: true,
|
|
@@ -30667,14 +30667,14 @@ var consolidate_default = {
|
|
|
30667
30667
|
};
|
|
30668
30668
|
// src/assets/improve-strategies/default.json
|
|
30669
30669
|
var default_default = {
|
|
30670
|
-
description: "Standard improve pass \u2014 reflect, distill, consolidation (promotion plus the reviewed pair-pass retire/supersede proposals), and validation. Memory inference is listed below but only runs when experimental.improveAutonomy is set; improve-stage extract and proactive maintenance off.",
|
|
30670
|
+
description: "Standard improve pass \u2014 reflect (rewrites only from negative feedback), distill, consolidation (promotion plus the reviewed pair-pass retire/supersede proposals), and validation. Memory inference is listed below but only runs when experimental.improveAutonomy is set; improve-stage extract and proactive maintenance off.",
|
|
30671
30671
|
processes: {
|
|
30672
30672
|
reflect: {
|
|
30673
30673
|
enabled: true,
|
|
30674
30674
|
limit: 25,
|
|
30675
30675
|
allowedTypes: ["agent", "command", "knowledge", "lesson", "memory", "skill", "workflow"]
|
|
30676
30676
|
},
|
|
30677
|
-
distill: { enabled: true, allowedTypes: ["memory"], requirePlannedRefs:
|
|
30677
|
+
distill: { enabled: true, allowedTypes: ["memory"], requirePlannedRefs: false },
|
|
30678
30678
|
consolidate: { enabled: true, allowedTypes: ["memory"] },
|
|
30679
30679
|
memoryInference: { enabled: true },
|
|
30680
30680
|
extract: { enabled: false, triage: { enabled: true, minScore: 2 } },
|
|
@@ -30686,7 +30686,7 @@ var default_default = {
|
|
|
30686
30686
|
};
|
|
30687
30687
|
// src/assets/improve-strategies/proactive-maintenance.json
|
|
30688
30688
|
var proactive_maintenance_default = {
|
|
30689
|
-
description: "Opt-in proactive-maintenance pass \u2014 reflect, distill, proposal triage (promote, high budget), and the proactive-maintenance lane (maxPerRun 100); consolidate/memoryInference/extract off. Sync disabled: an interrupted run would otherwise leave an uncommitted backlog.",
|
|
30689
|
+
description: "Opt-in proactive-maintenance pass \u2014 reflect (on negative feedback), distill, proposal triage (promote, high budget), and the proactive-maintenance lane (maxPerRun 100), which picks due assets for scoring but plans no rewrite; consolidate/memoryInference/extract off. Sync disabled: an interrupted run would otherwise leave an uncommitted backlog.",
|
|
30690
30690
|
processes: {
|
|
30691
30691
|
reflect: {
|
|
30692
30692
|
enabled: true,
|
|
@@ -30769,7 +30769,7 @@ var thorough_default = {
|
|
|
30769
30769
|
distill: {
|
|
30770
30770
|
enabled: true,
|
|
30771
30771
|
allowedTypes: ["memory"],
|
|
30772
|
-
requirePlannedRefs:
|
|
30772
|
+
requirePlannedRefs: false
|
|
30773
30773
|
},
|
|
30774
30774
|
consolidate: {
|
|
30775
30775
|
enabled: true,
|
|
@@ -167,7 +167,9 @@ export function resolveWritableOverride(config) {
|
|
|
167
167
|
* - Not a git repo → skipped (no-op)
|
|
168
168
|
* - Git repo, no remote → commit only
|
|
169
169
|
* - Git repo, has remote, but stash is not writable → commit only
|
|
170
|
-
* - Git repo, has remote, stash is writable → commit + push
|
|
170
|
+
* - Git repo, has remote, stash is writable → commit + push, unless the
|
|
171
|
+
* branch has no upstream or is behind or diverged from it: then commit
|
|
172
|
+
* only, and `reason` says why
|
|
171
173
|
*
|
|
172
174
|
* When `name` is omitted the primary stash directory is used.
|
|
173
175
|
* When `message` is omitted a timestamp is used.
|
|
@@ -288,7 +290,9 @@ export function saveGitStash(name, message, writableOverride, options) {
|
|
|
288
290
|
throw new Error(`git remote failed: ${remoteResult.stderr?.trim() || "unknown error"}`);
|
|
289
291
|
}
|
|
290
292
|
const hasRemote = remoteResult.stdout.trim().length > 0;
|
|
291
|
-
|
|
293
|
+
// A branch that cannot be pushed still gets its commit: only the push is skipped, and the result says why.
|
|
294
|
+
const upstream = hasRemote && writable && allowPush ? readActualUpstream(repoDir, baseHead) : undefined;
|
|
295
|
+
const pushTarget = typeof upstream === "object" ? upstream : undefined;
|
|
292
296
|
const exactCommit = createExactPathCommit(repoDir, {
|
|
293
297
|
baseHead,
|
|
294
298
|
commitMessage,
|
|
@@ -304,6 +308,7 @@ export function saveGitStash(name, message, writableOverride, options) {
|
|
|
304
308
|
committed: true,
|
|
305
309
|
pushed: false,
|
|
306
310
|
skipped: false,
|
|
311
|
+
...(typeof upstream === "string" ? { reason: `not pushed: ${upstream}` } : {}),
|
|
307
312
|
output: `commit ${exactCommit}`,
|
|
308
313
|
commit: exactCommit,
|
|
309
314
|
};
|
|
@@ -387,20 +392,22 @@ function readBranchRef(repoDir) {
|
|
|
387
392
|
}
|
|
388
393
|
return result.stdout.trim();
|
|
389
394
|
}
|
|
395
|
+
/**
|
|
396
|
+
* Where a commit on top of `baseHead` can be pushed: the branch's upstream, as
|
|
397
|
+
* long as it is an ancestor of `baseHead` (ahead is fine, the push fast-forwards
|
|
398
|
+
* it). Otherwise why not: no upstream, or behind or diverged from it.
|
|
399
|
+
*/
|
|
390
400
|
function readActualUpstream(repoDir, baseHead) {
|
|
391
|
-
if (!baseHead)
|
|
392
|
-
throw new UsageError(`Writable Git target at ${repoDir} has no commit to publish.`);
|
|
393
401
|
const branchRef = readBranchRef(repoDir);
|
|
394
402
|
const branch = branchRef.replace(/^refs\/heads\//, "");
|
|
395
403
|
const remote = runGit(["-C", repoDir, "config", "--get", `branch.${branch}.remote`]);
|
|
396
404
|
const merge = runGit(["-C", repoDir, "config", "--get", `branch.${branch}.merge`]);
|
|
397
405
|
const upstream = runGit(["-C", repoDir, "rev-parse", "--verify", "@{u}"]);
|
|
398
|
-
if (remote.status !== 0 || merge.status !== 0 || upstream.status !== 0)
|
|
399
|
-
|
|
400
|
-
}
|
|
406
|
+
if (remote.status !== 0 || merge.status !== 0 || upstream.status !== 0)
|
|
407
|
+
return "no upstream branch is configured";
|
|
401
408
|
const upstreamHead = upstream.stdout.trim();
|
|
402
|
-
if (!
|
|
403
|
-
|
|
409
|
+
if (!baseHead || runGit(["-C", repoDir, "merge-base", "--is-ancestor", upstreamHead, baseHead]).status !== 0) {
|
|
410
|
+
return "the branch is behind its upstream or has diverged from it";
|
|
404
411
|
}
|
|
405
412
|
return { remote: remote.stdout.trim(), mergeRef: merge.stdout.trim(), upstreamHead };
|
|
406
413
|
}
|
|
@@ -16,7 +16,6 @@ export function computeImproveRunMetrics(result) {
|
|
|
16
16
|
let acceptedCount = 0;
|
|
17
17
|
let rejectedCount = 0;
|
|
18
18
|
let skippedCount = 0;
|
|
19
|
-
let autoAcceptedCount = 0;
|
|
20
19
|
let errorCount = 0;
|
|
21
20
|
for (const action of actions) {
|
|
22
21
|
// Bucketing delegated to the shared classifyImproveAction so this aggregate
|
|
@@ -42,8 +41,7 @@ export function computeImproveRunMetrics(result) {
|
|
|
42
41
|
break;
|
|
43
42
|
}
|
|
44
43
|
}
|
|
45
|
-
|
|
46
|
-
autoAcceptedCount += result.gateAutoAcceptedCount ?? 0;
|
|
44
|
+
const autoAcceptedCount = result.triage?.promoted ?? 0;
|
|
47
45
|
// C1 (13-bus-factor): distill-skipped rows are folded into the bounded
|
|
48
46
|
// `distillSkipped` aggregate and no longer live in `actions`. Add the
|
|
49
47
|
// aggregate total to the skipped + total-actions counters so metrics_json
|
package/docs/reference/cli.md
CHANGED
|
@@ -1395,6 +1395,8 @@ value also looks like a format token.
|
|
|
1395
1395
|
| Git repo, no remote | Stage and commit only |
|
|
1396
1396
|
| Git repo, has remote, not writable | Stage and commit only |
|
|
1397
1397
|
| Git repo, has remote, `writable: true` | Stage, commit, and push |
|
|
1398
|
+
| Writable with a remote, but no upstream branch, or behind or diverged from it | Stage and commit; the push is skipped and `reason` says `not pushed: ...` |
|
|
1399
|
+
| Writable with a remote and ahead of its upstream | Stage, commit, and push, unpushed commits included |
|
|
1398
1400
|
| Any writable repo with `--no-push` | Stage and commit only (push suppressed) |
|
|
1399
1401
|
|
|
1400
1402
|
**Primary bundle writable config:**
|
|
@@ -1592,9 +1594,12 @@ preserves it byte-for-byte.
|
|
|
1592
1594
|
|
|
1593
1595
|
### feedback
|
|
1594
1596
|
|
|
1595
|
-
Record positive or negative feedback for any indexed bundle asset.
|
|
1596
|
-
|
|
1597
|
-
|
|
1597
|
+
Record positive or negative feedback for any indexed bundle asset.
|
|
1598
|
+
`akm feedback <ref> --negative --reason "<what is wrong and what should change>"`
|
|
1599
|
+
flags the asset for review: the next improve run proposes a fix based on your
|
|
1600
|
+
reason, so be specific. `--positive` records that an asset helped (it raises
|
|
1601
|
+
its ranking) and does not trigger a rewrite. Both signals update the asset's
|
|
1602
|
+
utility score right away, so highly-rated assets rank higher in search results.
|
|
1598
1603
|
|
|
1599
1604
|
```sh
|
|
1600
1605
|
akm feedback scripts/deploy.sh --positive
|
|
@@ -1608,9 +1613,9 @@ akm feedback skills/code-review --negative --reason "flaky" --tag slice:train --
|
|
|
1608
1613
|
|
|
1609
1614
|
| Flag | Description |
|
|
1610
1615
|
| --- | --- |
|
|
1611
|
-
| `--positive` | Record
|
|
1612
|
-
| `--negative` |
|
|
1613
|
-
| `--reason` | What
|
|
1616
|
+
| `--positive` | Record that an asset helped: it raises its ranking and does not trigger a rewrite |
|
|
1617
|
+
| `--negative` | Flag the asset for review: the next improve run proposes a fix based on `--reason`, so be specific |
|
|
1618
|
+
| `--reason` | What is wrong with the asset's content and what should change; not for `akm` command errors. Attached to the feedback event and read by the next improve run's fix proposal (required for negative feedback by default) |
|
|
1614
1619
|
| `--failure-mode` | Structured failure-mode taxonomy for negative feedback: `incorrect`, `outdated`, `dangerous`, `incomplete`, `redundant`. Stored alongside `--reason` in event metadata for the distill pipeline. |
|
|
1615
1620
|
| `--tag` | Tag to attach to the feedback (repeatable, e.g. `--tag slice:train --tag team:platform`) |
|
|
1616
1621
|
| `--applied-to <ref>` | Credit a `lessons/<name>` lesson that helped resolve this task. When combined with `--positive`, appends this feedback ref to the target lesson's `lessonStrength[]` frontmatter array (dedup, idempotent). A non-lesson target, or a missing `--positive`, produces a warning rather than silently doing nothing. |
|
|
@@ -1618,6 +1623,11 @@ akm feedback skills/code-review --negative --reason "flaky" --tag slice:train --
|
|
|
1618
1623
|
Specify exactly one of `--positive` or `--negative`. The ref must already be
|
|
1619
1624
|
present in the current local index.
|
|
1620
1625
|
|
|
1626
|
+
Only negative feedback with a specific reason gets an asset reviewed and fixed:
|
|
1627
|
+
`akm improve` plans a rewrite (reflect) only for an asset with fresh negative
|
|
1628
|
+
feedback, or for an explicit ref. A positive or note-only signal never plans
|
|
1629
|
+
one, and improve no longer rewrites assets on a proactive cadence.
|
|
1630
|
+
|
|
1621
1631
|
The `--applied-to` flag records the lesson-strength signal: each credit is
|
|
1622
1632
|
kept in the lesson's `lessonStrength[]` frontmatter. Search ranking does not
|
|
1623
1633
|
use it.
|
|
@@ -2432,7 +2442,7 @@ akm improve report --since 7d # ...aggregated over every real run start
|
|
|
2432
2442
|
| `--bundle` | Select the bundle the run improves and writes to (default: `defaultWriteTarget`, else the working bundle); only that bundle's assets are planned. When the ref scope is bundle-qualified, it must name the same bundle |
|
|
2433
2443
|
| `--limit <n>` | Cap the refs the run processes, highest salience first (refs routed to distill only come last). Overrides the strategy's `processes.reflect.limit` and `limit` |
|
|
2434
2444
|
| `--timeout-ms <ms>` | Wall-clock budget for the run (default: `7200000` = 2 hours) |
|
|
2435
|
-
| `--require-feedback-signal` |
|
|
2445
|
+
| `--require-feedback-signal` | Turn the fallback lanes (high salience, proactive maintenance) off for the run: they only select and score assets, and a rewrite needs negative feedback |
|
|
2436
2446
|
| `--strategy <name>` | Override the active improve strategy (a built-in or entry under `improve.strategies`) |
|
|
2437
2447
|
| `--json-to-stdout` | Also emit the full persisted JSON result on stdout for a live run. Without this flag, stdout stays empty. Dry-runs always emit their result and are never persisted. |
|
|
2438
2448
|
| `--skip-if-locked` | If another improve run already holds the lock, skip gracefully (exit 0) instead of failing with "already running" (exit 75, `TransientError`, code `IMPROVE_LOCK_HELD` — field follow-up to #948: two legitimate `improve` invocations colliding on this lock is ordinary, retryable contention, not a broken config file). Use for high-frequency scheduled runs so they don't pile up failures while a longer run is in progress. |
|
|
@@ -2481,7 +2491,8 @@ vector (semantic search off, or the memory not indexed yet) that check does
|
|
|
2481
2491
|
nothing and the exact slug and whole-body checks still apply.
|
|
2482
2492
|
|
|
2483
2493
|
No built-in strategy turns the improve-stage extract process on, and only
|
|
2484
|
-
`proactive-maintenance` turns proactive maintenance on
|
|
2494
|
+
`proactive-maintenance` turns proactive maintenance on, which selects and
|
|
2495
|
+
scores due assets but plans no rewrite. Use that strategy or
|
|
2485
2496
|
set the selected strategy's process `enabled: true` to opt in. The stage toggle does not disable a direct
|
|
2486
2497
|
`akm proposal extract --type <harness>` or `akm proposal extract --auto`
|
|
2487
2498
|
invocation.
|
|
@@ -2502,16 +2513,22 @@ the drain engine. Reflect still emits a `confidence` score (0..1) in its JSON
|
|
|
2502
2513
|
response schema; it is recorded on the proposal for triage and ranking, but no
|
|
2503
2514
|
threshold auto-accepts anything.
|
|
2504
2515
|
|
|
2505
|
-
Selection
|
|
2506
|
-
days
|
|
2507
|
-
|
|
2516
|
+
Selection plans a reflect (a proposed rewrite) only for refs with negative
|
|
2517
|
+
feedback in the last 30 days newer than the stage's last ledger attempt, or for
|
|
2518
|
+
an explicit ref scope. A positive or note-only signal never plans one, so
|
|
2519
|
+
improve does not rewrite an asset from a positive signal. Distill keeps its own
|
|
2520
|
+
trigger: a memory with feedback of any kind (a signal or a note) in that window,
|
|
2521
|
+
newer than distill's last attempt. Two fallback lanes pick refs with no such
|
|
2522
|
+
feedback: high salience (content-scored refs at or above
|
|
2508
2523
|
`improve.salience.salienceThreshold`, default `0.75`, that were never reflected,
|
|
2509
2524
|
capped at 10% of the limit, at least one ref) and, in a strategy that enables
|
|
2510
|
-
`proactiveMaintenance`, refs due for a revisit.
|
|
2525
|
+
`proactiveMaintenance`, refs due for a revisit. They only select and score refs
|
|
2526
|
+
(salience and outcome) and plan nothing, so improve does not rewrite on a
|
|
2527
|
+
proactive cadence; they pick only refs in the
|
|
2511
2528
|
[retrieval scope](https://github.com/itlackey/akm/blob/main/docs/architecture/improvement.md#retrieval-scope): returned by
|
|
2512
2529
|
`search`, `curate` or `show`, or named by feedback, in the last 90 days, or new
|
|
2513
|
-
material no improve stage has processed. The
|
|
2514
|
-
cut to the limit; an explicit ref scope bypasses every gate. Use
|
|
2530
|
+
material no improve stage has processed. The planned refs are ranked by salience
|
|
2531
|
+
and cut to the limit; an explicit ref scope bypasses every gate. Use
|
|
2515
2532
|
`--require-feedback-signal` to turn the fallback lanes off for the run.
|
|
2516
2533
|
|
|
2517
2534
|
When the active strategy enables a process (or the triage judgment engine)
|
|
@@ -2818,6 +2835,14 @@ Bulk-accept all pending proposals from one generator with `--generator <name>`
|
|
|
2818
2835
|
(e.g. `reflect`, `distill`) and no positional id. Bulk accept requires
|
|
2819
2836
|
`-y`/`--yes` in non-interactive shells.
|
|
2820
2837
|
|
|
2838
|
+
When the destination bundle is a git repository (a `.git` directory, whatever
|
|
2839
|
+
its source kind), each accept is committed as it happens, with exactly the paths
|
|
2840
|
+
it wrote or removed and the subject `akm accept: <generator> <proposal-id-8>
|
|
2841
|
+
<ref>`. A retirement's archived copy and tombstone, and the source memory a
|
|
2842
|
+
consolidate promotion retires, are part of the same commit. The commit is local:
|
|
2843
|
+
`akm sync`, or the end-of-run sync of `akm improve`, pushes it. A commit that
|
|
2844
|
+
fails warns and the accept stands.
|
|
2845
|
+
|
|
2821
2846
|
#### proposal reject
|
|
2822
2847
|
|
|
2823
2848
|
Reject a proposal and archive the reason. Accepts a full UUID, an 8-character
|
|
@@ -3039,7 +3064,8 @@ akm proposal drain --strategy default --promote -y # Read the triage block from
|
|
|
3039
3064
|
|
|
3040
3065
|
`akm feedback` accepts an optional `--reason <text>` flag whose value is
|
|
3041
3066
|
forwarded into feedback metadata and consumed by improve/distill proposal
|
|
3042
|
-
prompts. Negative feedback requires a reason by default
|
|
3067
|
+
prompts. Negative feedback requires a reason by default: the next improve run
|
|
3068
|
+
proposes a fix from it, so say what is wrong and what should change.
|
|
3043
3069
|
|
|
3044
3070
|
Write the reason about the asset's content. Reflect treats it as an unverified
|
|
3045
3071
|
report to investigate, not a fact to insert, and is told to leave the section
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "akm-cli",
|
|
3
|
-
"version": "0.9.
|
|
3
|
+
"version": "0.9.21",
|
|
4
4
|
"type": "module",
|
|
5
5
|
"description": "akm (Agent Knowledge Manager) — a portable, local-first capability library for AI agents. Discover, load, share, and improve reusable skills, scripts, workflows, and knowledge across any shell-capable coding agent, including Claude Code, OpenCode, and Cursor.",
|
|
6
6
|
"keywords": [
|