akm-cli 0.9.17-alpha.6 → 0.9.17-alpha.7

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (38) hide show
  1. package/CHANGELOG.md +146 -0
  2. package/STABILITY.md +2 -2
  3. package/dist/akm +7 -7
  4. package/dist/assets/prompts/retrieval-relevance-judge.md +6 -0
  5. package/dist/commands/improve/consolidate.js +11 -0
  6. package/dist/commands/improve/improve-cli.js +27 -7
  7. package/dist/commands/improve/ledger.js +7 -3
  8. package/dist/commands/improve/preparation.js +40 -10
  9. package/dist/commands/improve/reflect.js +46 -22
  10. package/dist/commands/improve/retrieval-gate.js +127 -0
  11. package/dist/commands/improve/retrieval-scope.js +77 -0
  12. package/dist/commands/read/curate.js +1 -17
  13. package/dist/commands/tasks/tasks-cli.js +10 -12
  14. package/dist/commands/tasks/tasks.js +57 -56
  15. package/dist/commands/tasks/validate.js +27 -46
  16. package/dist/core/adapter/adapters/akm-task-adapter.js +29 -8
  17. package/dist/core/improve-result.js +4 -1
  18. package/dist/core/non-task-input.js +20 -0
  19. package/dist/core/paths.js +0 -4
  20. package/dist/indexer/indexer.js +1 -3
  21. package/dist/indexer/usage/usage-events.js +34 -0
  22. package/dist/scripts/akm-migrate-node.js +5822 -5833
  23. package/dist/scripts/akm-migrate.js +6301 -6312
  24. package/dist/storage/repositories/proposals-repository.js +4 -0
  25. package/dist/tasks/backends/cron.js +80 -43
  26. package/dist/tasks/backends/launchd.js +28 -15
  27. package/dist/tasks/backends/schtasks.js +25 -10
  28. package/dist/tasks/run/load-task.js +1 -1
  29. package/dist/tasks/scheduler-binding.js +4 -2
  30. package/dist/tasks/scheduler-invocation.js +127 -235
  31. package/dist/tasks/scheduler-sync.js +13 -8
  32. package/dist/tasks/source/parse-task-source.js +22 -126
  33. package/dist/tasks/source/task-to-v3.js +1 -55
  34. package/dist/tasks/source/task-to-v4.js +1 -13
  35. package/docs/migration/v0.9.1-to-v0.9.2.md +7 -3
  36. package/docs/reference/cli.md +4 -3
  37. package/docs/reference/tasks.md +58 -40
  38. package/package.json +1 -1
package/CHANGELOG.md CHANGED
@@ -6,6 +6,152 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/).
6
6
 
7
7
  ## [Unreleased]
8
8
 
9
+ ## [0.9.17-alpha.7] - 2026-09-28
10
+
11
+ A scheduled task is now just a command and a schedule. Each native row
12
+ carries its own `AKM_BUNDLE_DIR` instead of pointing at a descriptor file,
13
+ and the runtime reads only v4 task files; older files convert once with
14
+ `akm migrate apply`. The first `akm task sync` after upgrading rewrites each
15
+ row once, keeping every task and schedule. `akm improve` reworks only assets
16
+ that retrieval returned or that are new, and reflect refuses a rewrite that
17
+ grades worse on the asset's own searches. `--require-engines` no longer
18
+ skips a run because the LLM endpoint is busy.
19
+
20
+ ### Changed
21
+
22
+ - **akm reads only task source v4.** A `version: 2` or `version: 3` task
23
+ file, or a `version: 4` file that still carries 0.9.15's retired
24
+ `schedule[].enabled`, now fails on its own with a message naming
25
+ `akm migrate apply`, which converts it once, under a backup (`akm upgrade`
26
+ runs it after an install). Until now every read converted such a file in
27
+ memory. `akm task sync` reports each one as a failure, leaves its installed
28
+ row as it is, and keeps reconciling every other task; `akm task run`,
29
+ `akm lint` and `akm task validate` report it the same way, and
30
+ `akm task validate`'s `converts` outcome is gone (such a file is
31
+ `blocked`). `akm migrate apply` now also converts the tasks of the stash
32
+ `AKM_BUNDLE_DIR` selects when no configured bundle names it, since the
33
+ runtime reads those too, and a root whose top-level task files are all
34
+ v2/v3 is still detected as an `akm-task` bundle, so they are found and
35
+ converted. A host whose task files are all v4 (`akm migrate status`
36
+ reports `current`) sees no difference.
37
+ (`src/tasks/source/parse-task-source.ts`, `src/commands/tasks/validate.ts`,
38
+ `scripts/akm-migrate/task-migrate.ts`,
39
+ `src/core/adapter/adapters/akm-task-adapter.ts`)
40
+ - **A scheduled row is its command plus its schedule, and carries its own
41
+ context.** Rows no longer name a `--scheduler-context` descriptor file;
42
+ they set what it held themselves. Every row sets `AKM_BUNDLE_DIR` to the
43
+ working stash of the shell that ran `akm task sync`, plus any
44
+ `AKM_CONFIG_DIR`, `AKM_DATA_DIR`, `AKM_CACHE_DIR` or `AKM_STATE_DIR` that
45
+ shell set explicitly: a `VAR=value` prefix in the crontab, an
46
+ `EnvironmentVariables` entry in a launchd plist, a `$env:VAR='value';`
47
+ assignment ahead of the command in Task Scheduler (its action has no
48
+ environment of its own). Scheduled runs see the same environment as
49
+ before, and sync still tells installations sharing a crontab apart by
50
+ that path (#846).
51
+ **What hosts see:** the first `akm task sync` after upgrading rewrites
52
+ every akm row once. `akm task sync --dry-run` lists each one as an update,
53
+ never an add or a remove; each keeps its launcher and its schedule, and
54
+ the `--scheduler-context <file>` argument becomes an inline
55
+ `AKM_BUNDLE_DIR=<working stash>`. On a host whose working stash is
56
+ `/home/u/akm` a row changes from
57
+
58
+ ```text
59
+ 30 8 * * * /home/u/.bun/bin/bun /home/u/.bun/lib/node_modules/akm-cli/dist/akm --scheduler-context /home/u/.local/share/akm/tasks/context/e898….json task run capture --bundle akm --scheduled > /home/u/.cache/akm/tasks/logs/capture.log 2>&1
60
+ ```
61
+
62
+ to
63
+
64
+ ```text
65
+ 30 8 * * * AKM_BUNDLE_DIR=/home/u/akm /home/u/.bun/bin/bun /home/u/.bun/lib/node_modules/akm-cli/dist/akm task run capture --bundle akm --scheduled > /home/u/.cache/akm/tasks/logs/capture.log 2>&1
66
+ ```
67
+
68
+ Rows written by 0.9.0 through 0.9.17-alpha.6 keep firing until that sync:
69
+ the CLI still accepts `--scheduler-context <file>` and applies the file's
70
+ environment (PATH included, for a 0.9.16 row). The files under
71
+ `$DATA/tasks/context/` are no longer written, and the uid, mode, symlink
72
+ and content-hash checks made on every scheduled run are gone. A sync
73
+ leaves some rows as they are (one whose task file failed to load, one a
74
+ `--bundle` sync did not cover), and those still name their file: once
75
+ `akm task doctor` lists no binding with a `contextPath`, nothing reads
76
+ them and they can be deleted. `akm task prune` now
77
+ finds rows whose `AKM_BUNDLE_DIR` names a directory that is gone, and older
78
+ rows whose descriptor cannot be read; `akm task doctor` lists `contextPath`
79
+ only for an older row. (`src/tasks/scheduler-invocation.ts`,
80
+ `src/tasks/backends/cron.ts`, `src/tasks/backends/launchd.ts`,
81
+ `src/tasks/backends/schtasks.ts`, `src/tasks/scheduler-sync.ts`,
82
+ `src/commands/tasks/tasks.ts`)
83
+ - **`akm improve` reworks only what gets read (#986).** An asset with fresh
84
+ feedback, or one you name (`akm improve skills/x`), is handled as before.
85
+ Every other pick must now be in the retrieval scope. That covers the
86
+ proactive-maintenance, high-salience and forgetting-safety lanes, and the
87
+ memories consolidation judges. An asset is in scope if a user `search`,
88
+ `curate` or `show` returned it, or user `feedback` named it, in the last 90
89
+ days, which is the usage log's retention. A hit on a `.derived` memory counts
90
+ for its parent. New material that no improve stage has processed yet is also
91
+ in scope. There is no new config key.
92
+
93
+ Measured with `akm improve --dry-run` on a copy of the maintainer's bundle
94
+ (19,870 assets), against 0.9.17-alpha.6:
95
+ - The fallback lanes pick from 6,450 assets instead of 15,686, and 9,236
96
+ refs are left out. No lane setting can reach the unread tail any more. In
97
+ July the proactive lane rewrote 3,069 assets, and 3,059 of them had not
98
+ been retrieved since the usage log began on 1 July.
99
+ - Consolidation judges 59 memories instead of 69.
100
+ - The high-salience lane no longer admits distill outputs nobody has read (2
101
+ today).
102
+ - Under the scheduled caps, today's nightly work is unchanged. The default
103
+ strategy selects the same 50 feedback-driven refs, and weekly proactive
104
+ maintenance selects the same 25, because salience ranking already puts
105
+ retrieved assets first.
106
+
107
+ `akm improve --dry-run` and the run result report the left-out refs as a new
108
+ `retrieval` gate. Health reports them under the skip reason `not_retrieved`.
109
+ Improve results stored by earlier releases, which have no such gate, still
110
+ decode. (`src/commands/improve/retrieval-scope.ts`,
111
+ `src/commands/improve/preparation.ts`, `src/commands/improve/consolidate.ts`)
112
+
113
+ - **Reflect refuses a rewrite that makes an asset worse for its own searches
114
+ (#722).** Before reflect proposes a rewrite of an existing asset, it grades
115
+ the old and the new content on up to five of the queries that actually
116
+ retrieved the asset (user `search` and `curate`). It uses the retrieval
117
+ eval's relevance prompt, which agrees with human grades at kappa 0.83. When
118
+ the new content grades lower on average, the rewrite is refused the same way
119
+ a quality-judge rejection is: `quality_rejected`, with the 14-day reflect
120
+ window. An asset without retrieval queries is not graded.
121
+
122
+ This was measured before it was built. Of 60 accepted rewrites since July,
123
+ judged this way, 14 graded lower (23%, 95% CI 14–35%) and 12 graded higher.
124
+ The gate was built because the lower bound cleared the 10% threshold set
125
+ before any judging. It costs two judge calls per query on the engine that
126
+ already runs the quality judge. On the maintainer's 2026-09-28 nightly run,
127
+ whose 30 rewrites had 102 usable queries, that is 204 calls, about 6 more
128
+ minutes on a 73-minute run. (`src/commands/improve/retrieval-gate.ts`,
129
+ `src/commands/improve/reflect.ts`)
130
+
131
+ ### Fixed
132
+
133
+ - **A cron row too long for one line is seen by `akm task sync` again.** A
134
+ command over 1,000 bytes runs from a wrapper script, and sync could not
135
+ read which task such a row ran: every sync, and every `--dry-run`, showed
136
+ it as an add and wrote it again. Sync now reads the script, so the row is
137
+ unchanged or an update like any other. (`src/tasks/backends/cron.ts`)
138
+ - **A `$` or a backslash in a scheduled row's value is kept.** launchd and
139
+ Task Scheduler rows passed their values through a string replacement that
140
+ read `$'`, `$&` and `$$` as patterns, so a path such as a Windows admin
141
+ share (`\\nas\share\akm$`) came out corrupted; reading a crontab row
142
+ back dropped a backslash inside a single-quoted value.
143
+ (`src/tasks/backends/launchd.ts`, `src/tasks/backends/schtasks.ts`,
144
+ `src/tasks/backends/cron.ts`)
145
+ - **`akm improve --require-engines` no longer skips a run because the LLM
146
+ endpoint is busy.** Its reachability probe, one short completion, gave up
147
+ after 3 seconds, so a local server busy with another job looked
148
+ unreachable and the whole scheduled run failed (all four scheduled runs on
149
+ 2026-09-27). The probe now waits up to the engine's own `timeoutMs`, at
150
+ most two minutes, so a busy server can answer while a hung one still fails
151
+ fast. The error's hint now says to check the endpoint rather than to run
152
+ `akm setup`.
153
+ (`src/commands/improve/improve-cli.ts`)
154
+
9
155
  ## [0.9.17-alpha.6] - 2026-09-27
10
156
 
11
157
  Graph extraction stops losing and wasting work. A timed-out extraction is
package/STABILITY.md CHANGED
@@ -281,8 +281,8 @@ CHANGELOG with a migration note.
281
281
  print unredacted. `akm task validate <path>` (new in 0.9.11) is the same
282
282
  kind of zero-write introspection as `explain`, but takes a bare filesystem
283
283
  path rather than a bundle-qualified ref — it reports whether that ONE file
284
- would parse cleanly (`valid`), auto-convert from task v2/v3 (`converts`),
285
- need a human decision the deterministic migrator can't make (`blocked`),
284
+ would parse cleanly (`valid`), need `akm migrate apply` first (`blocked`:
285
+ task v2/v3, or a retired `schedule[].enabled`),
286
286
  fail schema validation (`invalid`), or isn't a task source at all
287
287
  (`not-a-task`) — exactly the diagnostic `akm task sync` would produce for
288
288
  it, before the file is ever wired into a bundle or the scheduler.
package/dist/akm CHANGED
@@ -6,13 +6,13 @@
6
6
  import { spawn, spawnSync } from "node:child_process";
7
7
  import { fileURLToPath } from "node:url";
8
8
 
9
- // `--scheduler-context <descriptor>` passes through untouched: the CLI loads
10
- // and validates it before anything else (`consumeSchedulerContextArg` in
11
- // src/cli.ts). The scheduler supplies PATH itself (the crontab's `# akm:env`
12
- // line, a launchd plist's EnvironmentVariables), so runtime selection below
13
- // needs nothing from the descriptor. This launcher used to re-validate the
14
- // descriptor with its own copy of the schema, and that copy rejected every
15
- // descriptor the CLI writes since 0.9.17.
9
+ // A scheduled row sets its environment itself (a cron `VAR=value` prefix, a
10
+ // launchd plist's EnvironmentVariables), and the scheduler supplies PATH, so
11
+ // runtime selection below needs nothing from it. A row written by 0.9.0 –
12
+ // 0.9.17-alpha.6 passes `--scheduler-context <descriptor>` instead; it passes
13
+ // through untouched, and the CLI applies it (`consumeSchedulerContextArg` in
14
+ // src/cli.ts). This launcher used to re-validate the descriptor with its own
15
+ // copy of the schema, and that copy rejected every descriptor 0.9.17 wrote.
16
16
 
17
17
  process.env.AKM_LAUNCHER_NODE = process.execPath;
18
18
  process.env.AKM_LAUNCHER_PATH = fileURLToPath(import.meta.url);
@@ -0,0 +1,6 @@
1
+ You are grading search-and-retrieval results for an AI coding agent's knowledge base (the akm tool). For the given query and ONE candidate asset, grade how useful loading this asset would be, on this scale:
2
+ 3 = exactly the asset an agent should load for this query or task; it directly answers or performs the request.
3
+ 2 = relevant and clearly useful, though not the single best asset for the query.
4
+ 1 = same general topic as the query, but would not actually help complete this specific query or task.
5
+ 0 = unrelated to the query.
6
+ Reply with ONLY a JSON object: {"grade": <integer 0-3>, "reason": "<=25 words"}.
@@ -26,6 +26,7 @@ import { warn, warnVerbose } from "../../core/warn.js";
26
26
  import { resolveWriteTarget } from "../../core/write-source.js";
27
27
  import { deriveInstallations } from "../../indexer/installations.js";
28
28
  import { resolveSourceEntries } from "../../indexer/search/search-source.js";
29
+ import { USAGE_EVENT_RETENTION_DAYS } from "../../indexer/usage/usage-events.js";
29
30
  import { assertRunnerCredentials } from "../../integrations/agent/runner-dispatch.js";
30
31
  import { cosineSimilarity, embedBatch, resolveEmbeddingModelId } from "../../llm/embedder.js";
31
32
  import { getBodyEmbeddings, upsertBodyEmbeddings } from "../../storage/repositories/embeddings-repository.js";
@@ -39,6 +40,7 @@ import { sanitizeMergedContent } from "./consolidate/sanitize.js";
39
40
  import { contentHash } from "./content-hash.js";
40
41
  import { resolveImproveStrategy, resolveProcessEnabled } from "./improve-strategies.js";
41
42
  import { isLedgerBlocked, ledgerKey, loadLedgerSnapshot, recordLedgerAttempt } from "./ledger.js";
43
+ import { isInRetrievalScope, loadRetrievalScope } from "./retrieval-scope.js";
42
44
  import { callStage, mintProposal, noticeSet, stageRunner } from "./stage.js";
43
45
  /** A plan op worth acting on. Retired advisory ops (merge/delete/contradict) are dropped, never thrown on. */
44
46
  export function isValidOp(op) {
@@ -428,6 +430,11 @@ export function inspectConsolidationPool(opts, stashDir, warnings, existingKnowl
428
430
  });
429
431
  }
430
432
  const judgedUnchanged = poolSize - memories.length;
433
+ // Only what retrieval returned or new material improve never processed (#986).
434
+ const retrievalScope = loadRetrievalScope({ proposalsCtx: opts.proposalsCtx, readOnly }, stashDir);
435
+ const beforeScope = memories.length;
436
+ memories = memories.filter((memory) => isInRetrievalScope(retrievalScope, conceptIdFromTypeName("memory", memory.name), memory.filePath));
437
+ const outsideRetrievalScope = beforeScope - memories.length;
431
438
  if (opts.incrementalSince && memories.length > 0) {
432
439
  memories = narrowToIncrementalCandidates(memories, opts.incrementalSince, warnings, opts.neighborsPerChanged, readOnly);
433
440
  }
@@ -465,6 +472,7 @@ export function inspectConsolidationPool(opts, stashDir, warnings, existingKnowl
465
472
  memories,
466
473
  prefilteredAlreadyPromoted,
467
474
  judgedUnchanged,
475
+ outsideRetrievalScope,
468
476
  };
469
477
  }
470
478
  const ABORT_MIN_CHUNKS = 4;
@@ -641,6 +649,9 @@ async function consolidate(opts, config, stashDir, startMs, stateDb) {
641
649
  if (pool.judgedUnchanged > 0) {
642
650
  warnings.push(`Consolidation: skipped ${pool.judgedUnchanged} ${plural(pool.judgedUnchanged)} judged within the revisit window and unchanged since.`);
643
651
  }
652
+ if (pool.outsideRetrievalScope > 0) {
653
+ warnings.push(`Consolidation: skipped ${pool.outsideRetrievalScope} ${plural(pool.outsideRetrievalScope)} already judged that retrieval has not returned in the last ${USAGE_EVENT_RETENTION_DAYS} days.`);
654
+ }
644
655
  if (prefilteredAlreadyPromoted > 0) {
645
656
  warnings.push(`Consolidation: pre-filtered ${prefilteredAlreadyPromoted} ${plural(prefilteredAlreadyPromoted)} whose body already exists verbatim in knowledge/ before chunking.`);
646
657
  }
@@ -14,6 +14,7 @@ import { getCacheDir } from "../../core/paths.js";
14
14
  import { redactSensitiveText } from "../../core/redaction.js";
15
15
  import { clearLogFile, setLogFile, warn } from "../../core/warn.js";
16
16
  import { resolveWriteTarget } from "../../core/write-source.js";
17
+ import { DEFAULT_LLM_TIMEOUT_MS } from "../../integrations/agent/config.js";
17
18
  import { collectEngineCredentialValues } from "../../integrations/agent/engine-resolution.js";
18
19
  import { probeLlmReachable } from "../../llm/client.js";
19
20
  import { getOutputMode } from "../../output/context.js";
@@ -77,25 +78,44 @@ function collectRequiredEngineTargets(plan) {
77
78
  const targets = [];
78
79
  for (const [processName, process] of Object.entries(plan.processes)) {
79
80
  if (process.runner) {
80
- targets.push({ process: processName, engine: process.runner.engine, connection: process.runner.connection });
81
+ targets.push({
82
+ process: processName,
83
+ engine: process.runner.engine,
84
+ connection: probeConnection(process.runner),
85
+ });
81
86
  }
82
87
  }
83
88
  if (plan.triageJudgment?.kind === "llm") {
84
89
  targets.push({
85
90
  process: "triage.judgment",
86
91
  engine: plan.triageJudgment.engine,
87
- connection: plan.triageJudgment.connection,
92
+ connection: probeConnection(plan.triageJudgment),
88
93
  });
89
94
  }
90
95
  return targets;
91
96
  }
97
+ /** The resolved engine keeps its request timeout beside the connection (the runtime merges it in); the probe needs it on the connection. */
98
+ function probeConnection(runner) {
99
+ return runner.timeoutMs !== undefined ? { ...runner.connection, timeoutMs: runner.timeoutMs } : runner.connection;
100
+ }
101
+ /**
102
+ * The bound on `--require-engines`' probe: the connection's own request
103
+ * timeout, at most two minutes. A local server busy with another job queues
104
+ * the probe behind that job, and a fixed 3s bound failed every scheduled
105
+ * improve run on 2026-09-27 against a reachable endpoint; the cap still ends a
106
+ * hung endpoint (#957) long before a run's own multi-minute calls would.
107
+ */
108
+ export function requiredEngineProbeTimeoutMs(connection) {
109
+ return Math.min(connection.timeoutMs ?? DEFAULT_LLM_TIMEOUT_MS, REQUIRED_ENGINE_PROBE_MAX_MS);
110
+ }
111
+ const REQUIRED_ENGINE_PROBE_MAX_MS = 120_000;
92
112
  /**
93
113
  * `--require-engines`, live: probe each connection's real completion path
94
- * (a gateway can list a model whose completion route is dead, #980) with a 3s
95
- * bound, once per endpoint + model. Returns each target's latency for the run
96
- * result (R17); an unreachable one fails the run.
114
+ * (a gateway can list a model whose completion route is dead, #980), once per
115
+ * endpoint + model, within {@link requiredEngineProbeTimeoutMs}. Returns each
116
+ * target's latency for the run result (R17); an unreachable one fails the run.
97
117
  */
98
- export async function assertRequiredEnginesReachable(plan, probeReachable = (connection) => probeLlmReachable(connection, 3_000)) {
118
+ export async function assertRequiredEnginesReachable(plan, probeReachable = (connection) => probeLlmReachable(connection, requiredEngineProbeTimeoutMs(connection))) {
99
119
  const targets = collectRequiredEngineTargets(plan);
100
120
  if (targets.length === 0)
101
121
  return [];
@@ -117,7 +137,7 @@ export async function assertRequiredEnginesReachable(plan, probeReachable = (con
117
137
  const unreachable = probed.filter((item) => !item.reach.reachable);
118
138
  if (unreachable.length > 0) {
119
139
  const lines = unreachable.map((item) => ` - ${item.process} (engine "${item.engine}", ${item.connection.endpoint}): ${item.reach.error ?? "did not respond"}`);
120
- throw new ConfigError(`--require-engines: ${unreachable.length} improve process${unreachable.length === 1 ? "" : "es"} cannot run because ${unreachable.length === 1 ? "its" : "their"} engine completion path is not reachable:\n${lines.join("\n")}`, "LLM_NOT_CONFIGURED");
140
+ throw new ConfigError(`--require-engines: ${unreachable.length} improve process${unreachable.length === 1 ? "" : "es"} cannot run because ${unreachable.length === 1 ? "its" : "their"} engine completion path is not reachable:\n${lines.join("\n")}`, "LLM_NOT_CONFIGURED", "Check that each listed endpoint is up and serves its model. The probe is one short completion, bounded by the engine's timeoutMs (at most two minutes).");
121
141
  }
122
142
  return probed.map((item) => ({
123
143
  process: item.process,
@@ -82,13 +82,17 @@ export function recordLedgerAttempt(access, inputs) {
82
82
  export function ledgerKey(source, ref) {
83
83
  return `${source}\0${ref}`;
84
84
  }
85
+ /** Read state.db through `fn` without creating it: `undefined` when there is none yet. */
86
+ export function readLedgerDb(access, fn) {
87
+ if (!access?.eventsCtx?.db && !fs.existsSync(ledgerDbPath(access) ?? getStateDbPath()))
88
+ return undefined;
89
+ return withLedgerDb(access, fn);
90
+ }
85
91
  /** Every ledger row for `sources` in one query; no state.db yet means nothing was attempted. */
86
92
  export function loadLedgerSnapshot(access, stashDir, sources) {
87
93
  const out = new Map();
88
- if (!access?.eventsCtx?.db && !fs.existsSync(ledgerDbPath(access) ?? getStateDbPath()))
89
- return out;
90
94
  try {
91
- withLedgerDb(access, (db) => {
95
+ readLedgerDb(access, (db) => {
92
96
  for (const row of listImproveLedgerRows(db, stashDir, sources))
93
97
  out.set(ledgerKey(row.source, row.ref), row);
94
98
  });
@@ -23,7 +23,7 @@ import { ConfigError, rethrowIfTestIsolationError } from "../../core/errors.js";
23
23
  import { appendEvent, readEvents } from "../../core/events.js";
24
24
  import { withStateDb } from "../../core/state-db.js";
25
25
  import { info, warn } from "../../core/warn.js";
26
- import { countUsageEventsByType } from "../../indexer/usage/usage-events.js";
26
+ import { countUsageEventsByType, USAGE_EVENT_RETENTION_DAYS } from "../../indexer/usage/usage-events.js";
27
27
  import { getAvailableHarnesses } from "../../integrations/session-logs/index.js";
28
28
  import { getZeroResultSearches } from "../../storage/repositories/index-entries-repository.js";
29
29
  import { getRetrievalCounts } from "../../storage/repositories/index-utility-repository.js";
@@ -41,6 +41,7 @@ import { applyMemoryCleanup } from "./memory/memory-improve.js";
41
41
  import { getAllAssetOutcomes, getAssetOutcome, getOutcomeScoresByRef, OUTCOME_SCORE_MAX, outcomeScoreToSalience, projectAssetOutcome, updateAssetOutcome, } from "./outcome-loop.js";
42
42
  import { projectMemoryCleanup, selectEffectiveImproveRefs } from "./planner.js";
43
43
  import { DEFAULT_DUE_DAYS, DEFAULT_MAX_PER_RUN, selectProactiveMaintenanceRefs } from "./proactive-maintenance.js";
44
+ import { isInRetrievalScope, loadRetrievalScope } from "./retrieval-scope.js";
44
45
  import { buildRankChangeReport, computeSalience, getAllRankScores, getAssetSalience, getLastUseMsByRef, isContentEncodingRow, SALIENCE_NO_OP_DAMPEN_FACTOR, SALIENCE_NO_OP_DAMPEN_THRESHOLD, upsertAssetSalience, } from "./salience.js";
45
46
  import { attributeStage, errMessage } from "./stage.js";
46
47
  /** The candidate's durable state key (salience, outcome, ledger). */
@@ -663,12 +664,23 @@ async function selectLoopCandidates(args, postCleanupRefs, validationFailureRefs
663
664
  const processableRefs = [...partition.eligibleRefs, ...partition.distillOnlyRefs];
664
665
  const signalFiltered = processableRefs.filter((c) => snapshot.feedback.get(c.ref)?.hasSignal === true);
665
666
  const signalBearingSet = new Set(signalFiltered.map((r) => r.ref));
666
- const noFeedbackCandidates = dedupeRefs([
667
+ // The fallback lanes (proactive, high salience, forgetting safety) have no
668
+ // usage evidence of their own: they pick only what retrieval returned or new
669
+ // material improve never processed (#986). Evaluated once per candidate.
670
+ const fallbackEligible = postCleanupRefs.filter((c) => !validationFailureRefs.has(c.ref));
671
+ const allowFallbacks = options.requireFeedbackSignal !== true;
672
+ const retrievalScope = scope.mode === "ref" || !allowFallbacks
673
+ ? undefined
674
+ : loadRetrievalScope({ eventsCtx, ...(persist ? {} : { readOnly: true }) }, primaryStashDir ?? options.stashDir);
675
+ const unscoped = new Set(fallbackEligible.filter((c) => !isInRetrievalScope(retrievalScope, c.ref, c.filePath)).map((c) => c.ref));
676
+ const noFeedbackPool = dedupeRefs([
667
677
  ...processableRefs.filter((r) => !signalBearingSet.has(r.ref)),
668
678
  ...partition.noFeedbackPool,
669
679
  ]);
680
+ const noFeedbackCandidates = noFeedbackPool.filter((r) => !unscoped.has(r.ref));
681
+ // Only a ref no fallback lane may pick anymore is charged to the retrieval gate.
682
+ const outOfScope = new Set(noFeedbackPool.filter((r) => unscoped.has(r.ref)).map((r) => r.ref));
670
683
  const retrieval = fetchRetrievalSignals(options, signalFiltered, noFeedbackCandidates, eventsCtx, persist);
671
- const allowFallbacks = options.requireFeedbackSignal !== true;
672
684
  const proactive = allowFallbacks
673
685
  ? selectProactiveMaintenanceLane(args, noFeedbackCandidates, snapshot, retrieval, persist)
674
686
  : { proactiveRefs: [] };
@@ -692,10 +704,10 @@ async function selectLoopCandidates(args, postCleanupRefs, validationFailureRefs
692
704
  sourceByRef.set(r.ref, "scope");
693
705
  for (const r of mergedRefs)
694
706
  r.eligibilitySource = sourceByRef.get(r.ref) ?? "unknown";
695
- // Forgetting safety may only reuse this plan's own surviving objects, and
696
- // never a ref whose reflect window is still open.
697
- const fallbackEligible = postCleanupRefs.filter((c) => !validationFailureRefs.has(c.ref));
698
- const forgettingEligible = fallbackEligible.filter((c) => !isLedgerBlocked(ledgerRowFor(snapshot.ledger, "reflect", c.ref, c.itemRef), snapshot.nowIso));
707
+ // Forgetting safety may only reuse this plan's own surviving objects inside
708
+ // the retrieval scope, and never a ref whose reflect window is still open.
709
+ const forgettingEligible = fallbackEligible.filter((c) => !unscoped.has(c.ref) &&
710
+ !isLedgerBlocked(ledgerRowFor(snapshot.ledger, "reflect", c.ref, c.itemRef), snapshot.nowIso));
699
711
  const scored = scoreSalience(args, mergedRefs, snapshot.feedback, retrieval.retrievalCounts, persist);
700
712
  mergedRefs = applyForgettingSafety({
701
713
  pendingForgettingRefs: scored.pendingForgettingRefs,
@@ -739,28 +751,46 @@ async function selectLoopCandidates(args, postCleanupRefs, validationFailureRefs
739
751
  // Skip observability waits until every fallback lane has finalized the
740
752
  // survivors, so a rescued ref is never also reported skipped.
741
753
  const survivors = new Set(sorted.map((c) => c.ref));
742
- const signalSkipped = fallbackEligible.filter((c) => !survivors.has(c.ref));
754
+ const skipped = fallbackEligible.filter((c) => !survivors.has(c.ref));
755
+ const retrievalSkipped = skipped.filter((c) => outOfScope.has(c.ref));
756
+ const signalSkipped = skipped.filter((c) => !outOfScope.has(c.ref));
743
757
  for (const ref of partition.distillCooledRefs) {
744
758
  actions.push({ ref, mode: "distill-skipped", result: { ok: true, reason: "distill signal-delta" } });
745
759
  if (persist)
746
760
  recordImproveSkip(eventsCtx, ref, { reason: "distill_no_new_signal" });
747
761
  }
748
- for (const candidate of signalSkipped) {
762
+ for (const candidate of skipped) {
749
763
  actions.push({
750
764
  ref: candidate.ref,
751
765
  mode: "distill-skipped",
752
- result: { ok: true, reason: "no new signal since last proposal" },
766
+ result: {
767
+ ok: true,
768
+ reason: outOfScope.has(candidate.ref)
769
+ ? "not retrieved inside the usage window"
770
+ : "no new signal since last proposal",
771
+ },
753
772
  });
754
773
  }
755
774
  if (persist && signalSkipped.length > 0) {
756
775
  recordImproveSkip(eventsCtx, undefined, { reason: "no_new_signal", count: signalSkipped.length });
757
776
  }
777
+ if (persist && retrievalSkipped.length > 0) {
778
+ recordImproveSkip(eventsCtx, undefined, { reason: "not_retrieved", count: retrievalSkipped.length });
779
+ }
758
780
  const blocked = signalSkipped.length + partition.distillOnlyRefs.length;
759
781
  if (blocked > 0) {
760
782
  info(`[improve] ${blocked} of ${partition.preCooldownCount} indexed refs blocked by reflect signal-delta ` +
761
783
  `(${signalSkipped.length} fully skipped, ${partition.distillOnlyRefs.length} routed to distill-only)`);
762
784
  }
785
+ if (retrievalSkipped.length > 0) {
786
+ info(`[improve] ${retrievalSkipped.length} refs left out: not retrieved in the last ${USAGE_EVENT_RETENTION_DAYS} days, and not new material`);
787
+ }
763
788
  const gates = [
789
+ {
790
+ name: "retrieval",
791
+ removed: retrievalSkipped.length,
792
+ reason: `no feedback, and neither returned by search, curate or show in the last ${USAGE_EVENT_RETENTION_DAYS} days nor new material improve never processed`,
793
+ },
764
794
  {
765
795
  name: "signal",
766
796
  removed: signalSkipped.length,
@@ -43,6 +43,7 @@ import { findAssetFilePath } from "./eligibility.js";
43
43
  import { resolveImproveLlmExecution } from "./execution.js";
44
44
  import { recordLedgerAttempt } from "./ledger.js";
45
45
  import { classifyReflectChange, splitFrontmatter } from "./reflect-noise.js";
46
+ import { loadRetrievalQueries, runRetrievalRegressionGate } from "./retrieval-gate.js";
46
47
  import { callStage, mintProposal, noticeSet, rejectedProposalContext, runReflectQualityJudge, } from "./stage.js";
47
48
  const MAX_FEEDBACK_LINES = 10;
48
49
  const MAX_GLOBAL_FEEDBACK_LINES = 20;
@@ -960,6 +961,24 @@ async function finalizeReflectProposal(args) {
960
961
  }
961
962
  const flagged = Boolean(sanitized.sizeGuardRatio || sanitized.truncationMarkerLeaked);
962
963
  const judged = judge.enabled && !flagged;
964
+ /** A judge refused the revision: record it for the ledger's rejection window and stop. */
965
+ const refuse = (detail, metadata, message) => {
966
+ if (options.ref) {
967
+ recordLedgerAttempt({ proposalsCtx: options.ctx, eventsCtx: options.eventsCtx }, {
968
+ stashDir: run.stash,
969
+ ref: options.itemRef ?? options.ref,
970
+ source: "reflect",
971
+ outcome: "quality_rejected",
972
+ detail,
973
+ });
974
+ }
975
+ appendEvent({
976
+ eventType: "reflect_completed",
977
+ ref: payload.ref,
978
+ metadata: { source: "reflect", qualityRejected: true, ...metadata, ...telemetry },
979
+ }, options.eventsCtx);
980
+ return reflectFailure(run, result, "quality_rejected", message, false);
981
+ };
963
982
  if (judged) {
964
983
  const verdict = await runReflectQualityJudge(run.config, payload.content, assetContent ?? "", feedback, options.chat, {
965
984
  runnerSelectionFrozen: true,
@@ -969,28 +988,33 @@ async function finalizeReflectProposal(args) {
969
988
  onNotices: run.notices.add,
970
989
  });
971
990
  if (!verdict.pass) {
972
- if (options.ref) {
973
- recordLedgerAttempt({ proposalsCtx: options.ctx, eventsCtx: options.eventsCtx }, {
974
- stashDir: run.stash,
975
- ref: options.itemRef ?? options.ref,
976
- source: "reflect",
977
- outcome: "quality_rejected",
978
- detail: verdict.reason,
979
- });
980
- }
981
- appendEvent({
982
- eventType: "reflect_completed",
983
- ref: payload.ref,
984
- metadata: {
985
- source: "reflect",
986
- qualityRejected: true,
987
- qualityScore: verdict.score,
988
- qualityReason: verdict.reason,
989
- ...(verdict.criteria ? { qualityCriteria: verdict.criteria } : {}),
990
- ...telemetry,
991
- },
992
- }, options.eventsCtx);
993
- return reflectFailure(run, result, "quality_rejected", `Reflect proposal quality gate rejected: score=${verdict.score}, reason="${verdict.reason}"`, false);
991
+ return refuse(verdict.reason, {
992
+ qualityScore: verdict.score,
993
+ qualityReason: verdict.reason,
994
+ ...(verdict.criteria ? { qualityCriteria: verdict.criteria } : {}),
995
+ }, `Reflect proposal quality gate rejected: score=${verdict.score}, reason="${verdict.reason}"`);
996
+ }
997
+ }
998
+ // #722: a rewrite of an existing asset must not grade lower on its own retrieval queries.
999
+ if (judged && judge.runner && assetContent !== undefined) {
1000
+ const retrieval = await runRetrievalRegressionGate({
1001
+ ref: payload.ref,
1002
+ before: assetContent,
1003
+ after: payload.content,
1004
+ queries: loadRetrievalQueries({ proposalsCtx: options.ctx, eventsCtx: options.eventsCtx }, payload.ref),
1005
+ runner: judge.runner,
1006
+ ...(options.chat ? { chat: options.chat } : {}),
1007
+ ...(Object.hasOwn(options, "timeoutMs") ? { timeoutMs: options.timeoutMs } : {}),
1008
+ ...(options.signal ? { signal: options.signal } : {}),
1009
+ onNotices: run.notices.add,
1010
+ });
1011
+ if (!retrieval.pass) {
1012
+ return refuse(retrieval.reason, {
1013
+ retrievalRegression: true,
1014
+ retrievalQueries: retrieval.queries,
1015
+ ...(retrieval.oldMean !== undefined ? { retrievalGradeBefore: retrieval.oldMean } : {}),
1016
+ ...(retrieval.newMean !== undefined ? { retrievalGradeAfter: retrieval.newMean } : {}),
1017
+ }, `Reflect proposal refused: ${retrieval.reason}`);
994
1018
  }
995
1019
  }
996
1020
  // A lesson reflect wrote is marked so a later reflect on the same skill does