akm-cli 0.9.15-beta.3 → 0.9.15-beta.5

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -8,6 +8,36 @@ now, retry shortly." If a script or scheduler wrapper special-cases exit 2 to
8
8
  detect a held lease, switch it to exit 75, or read the JSON envelope's `code`
9
9
  field instead.
10
10
 
11
+ A concurrent `akm index` now also exits 75 instead of exit 78 or exit 70. A
12
+ 2026-09-10 field report found a second `akm index` (no `--skip-if-locked`)
13
+ colliding on the short internal barrier that registers the opt-in rebuild
14
+ lock could fail with `{"code":"INVALID_CONFIG_FILE"}` at exit 78 — a
15
+ config-error exit that told a supervisor to stop retrying ordinary
16
+ contention between two legitimate runs. That barrier now retries briefly
17
+ before giving up, and a busy barrier is reclassified as `TransientError`
18
+ code `MAINTENANCE_BARRIER_BUSY` (exit 75) instead. Contention on index.db
19
+ itself was already reclassified from a raw `database is locked` error (exit
20
+ 70) to `TransientError` code `INDEX_DB_CONTENDED` (exit 75). Either way, the
21
+ fix for a scheduled or opportunistic index run is the same:
22
+ `akm index --skip-if-locked` steps aside (exit 0) instead of contending at
23
+ all.
24
+
25
+ Two more contention paths collided on the same "config error instead of
26
+ transient" anti-pattern (#948 field follow-up, 2026-09-10). `akm improve`'s
27
+ whole-run lock (only one `improve` runs at a time) used to fail a losing
28
+ contender with `{"code":"INVALID_CONFIG_FILE"}` at exit 78 when no
29
+ `--skip-if-locked` was passed — a config-error exit for ordinary contention
30
+ between two legitimate `improve` invocations. It is now `TransientError`
31
+ code `IMPROVE_LOCK_HELD` at exit 75, naming the current holder's pid and
32
+ start time; `--skip-if-locked` is unchanged (still exit 0). Separately, the
33
+ index rebuild's embedding-verification step (`[index:verify] Semantic search
34
+ verification failed: ...`) could still surface a raw, unclassified
35
+ `database is locked` message even though the acquisition path was already
36
+ fixed — that message is now built through the same `INDEX_DB_CONTENDED`
37
+ reclassification as the rest of the index path. This step does not throw, so
38
+ it does not change `akm index`'s own exit code; only the message text
39
+ changed.
40
+
11
41
  The six shipped scheduled `improve` task templates now run with
12
42
  `--require-engines`, which aborts (exit 78) before any index work when a
13
43
  process's engine or credential cannot be resolved in the task's own
@@ -17,6 +47,23 @@ already-materialized task files are not rewritten. To get the same protection
17
47
  on an existing scheduled task, add `--require-engines` to its `run:` command
18
48
  yourself, then run `akm task sync`.
19
49
 
50
+ `--require-engines` now also runs a bounded reachability probe against each
51
+ distinct engine endpoint (the same probe `akm health` already uses), not
52
+ only a config/credential check. A field re-test found the flag let a run
53
+ through to a fully dead endpoint, which then sat silent for minutes making
54
+ no progress and no exit — the flag's exit-78 abort now catches that case up
55
+ front, naming the unreachable engine and endpoint, before any index work.
56
+ If a scheduled `--require-engines` run starts failing at exit 78 after
57
+ upgrading, check that the engine's endpoint actually answers — this is the
58
+ flag doing its documented job on a condition it previously missed, not a
59
+ new failure mode. Separately, `--timeout-ms` and an engine's own configured
60
+ timeout already aborted an in-flight request correctly, and SIGTERM/SIGINT
61
+ already ended a run within its documented grace period — both confirmed,
62
+ not changed, by this investigation. A live run now also prints one
63
+ default-level line if it waits more than a few seconds on its first engine
64
+ response, so a slow-but-alive run and a dead one are never indistinguishable
65
+ from silence alone.
66
+
20
67
  `akm health --no-probe` now also skips the `cli-version` update check (a GitHub
21
68
  release lookup), alongside the engine-reachability checks it already skipped.
22
69
  An air-gapped or offline host's existing `--no-probe` habit now suppresses both
@@ -114,6 +161,40 @@ instead of ever failing a whole batch over one oversized document.
114
161
  set `contextLength` specifically to control request batching (not your
115
162
  Ollama server's context window), set `embedding.maxTokens` instead.
116
163
 
164
+ **Which token knob fixed the original 8k-context overflow.** A 0.9.15-beta
165
+ field report described documents estimated under the request budget that
166
+ still tokenized to 8.5k-12.4k real tokens against an 8192-token endpoint,
167
+ because the 4-chars-per-token estimator undercounts dense technical text.
168
+ `maxInputTokens`, `maxTokens`, and `contextLength` are easy to confuse, and
169
+ only one of them makes that overflow structurally unreachable:
170
+
171
+ - `embedding.maxInputTokens` (default `512`) is the per-DOCUMENT cap. It
172
+ truncates a document's embedded text to its head before the document is
173
+ ever counted toward a request, so no single document can contribute more
174
+ than 512 estimated tokens. This is the fix: it makes the original
175
+ single-document overflow structurally unreachable, independent of the
176
+ other two knobs.
177
+ - `embedding.maxTokens` (default `6000`) is the per-REQUEST budget — how
178
+ many already-capped documents fit in one HTTP request — plus a same-run
179
+ adaptive shrink on the first context-size rejection. It reduces how often
180
+ a request lands near an endpoint's real limit, but a request-level budget
181
+ alone cannot stop one oversized document from overflowing a request.
182
+ - `embedding.contextLength` sets Ollama's `num_ctx` only. It no longer feeds
183
+ the request token budget the way it used to (the two fields used to share
184
+ this one value), and it has no effect at all against a non-Ollama
185
+ endpoint.
186
+
187
+ The field's exact 0.9.15-beta config — `contextLength: 8192` and
188
+ `maxTokens: 8000` — produces no 400s on 0.9.15. `maxTokens` now defaults
189
+ lower anyway (6000), but that is not why the overflow stopped: every
190
+ document is truncated to `maxInputTokens` (512 tokens) before it is counted
191
+ toward any request, so the 8.5k-12.4k-token documents that used to overflow
192
+ an 8192-token endpoint can no longer reach the request budget in the first
193
+ place. Set `embedding.maxInputTokens` higher only if you need documents
194
+ longer than ~2000 characters embedded in full — for a corpus with such
195
+ documents, size `embedding.maxTokens` to still fit the worst case, or the
196
+ overflow risk returns.
197
+
117
198
  `akm index --full` and an index-generation bump no longer re-embed
118
199
  unchanged content: vectors about to be discarded are salvaged and handed
119
200
  back to unchanged entries at the start of the next embedding pass instead
@@ -285,6 +285,13 @@ long enough to exhaust the driver's retry window, the run now fails with
285
285
  exit 75 (`TransientError`, code `INDEX_DB_CONTENDED`) instead of the raw
286
286
  driver error at exit 70 — the same retry-shortly contract as
287
287
  `STATE_DB_CONTENDED`, so a scheduler can branch on it instead of alerting.
288
+ The rebuild lock itself is registered through a brief internal barrier
289
+ (`getMaintenanceBarrierPath()`) shared with every other akm lock/lease; two
290
+ `akm index` runs launched close enough together to collide on that
291
+ registration step retry briefly and then, if it is still busy, also exit 75
292
+ (code `MAINTENANCE_BARRIER_BUSY`) rather than the config-error exit 78 a
293
+ 2026-09-10 field report found — a busy registration barrier is ordinary
294
+ contention between two legitimate runs, never a broken config file.
288
295
  `--skip-if-locked` changes that only for the invocation that passes it: if
289
296
  the lock is already held by a live process, it skips gracefully (exit 0,
290
297
  `{ ok: true, skipped: { reason: "lock-held", pid, launcherPid, startedAt } }`
@@ -2336,6 +2343,7 @@ akm improve --require-engines # for scheduled runs: abort (exit 78) ins
2336
2343
  akm improve --no-sync # skip the end-of-run git commit entirely (default: on for git-backed bundles)
2337
2344
  akm improve --sync --no-push # commit only, skip the push after it
2338
2345
  akm improve --plan --strategy thorough # preview thorough's resolved engine/model routing; nothing is dispatched
2346
+ akm improve lessons/my-lesson --show-prompt --format text # print the composed reflect prompt for one asset, unwrapped; no lock/index/engine call
2339
2347
  akm improve report # LLM usage/routing report for the most recent real run
2340
2348
  akm improve report --run <id> # ...for one specific improve_runs id
2341
2349
  akm improve report --since 7d # ...aggregated over every real run started in the last 7 days
@@ -2354,8 +2362,9 @@ akm improve report --since 7d # ...aggregated over every real run start
2354
2362
  | `--require-feedback-signal` | Only process assets with recent feedback signals |
2355
2363
  | `--strategy <name>` | Override the active improve strategy (a built-in or entry under `improve.strategies`) |
2356
2364
  | `--json-to-stdout` | Also emit the full persisted JSON result on stdout for a live run. Without this flag, stdout stays empty. Dry-runs always emit their result and are never persisted. |
2357
- | `--skip-if-locked` | If another improve run already holds the lock, skip gracefully (exit 0) instead of failing with "already running" (exit 78). Use for high-frequency scheduled runs so they don't pile up failures while a longer run is in progress. |
2358
- | `--require-engines` | Abort (exit 78, before any indexing, lock, or log side effect) if the active strategy would enable a process whose engine or credential cannot be resolved in this process's environment. Without this flag, improve degrades gracefully: it skips the affected processes and reports them in the result's `skippedProcesses`. Recommended alongside `--skip-if-locked` for scheduled runs, since the operator's own shell can pass config validation while a scheduler's stripped-down environment (see #953) cannot. |
2365
+ | `--skip-if-locked` | If another improve run already holds the lock, skip gracefully (exit 0) instead of failing with "already running" (exit 75, `TransientError`, code `IMPROVE_LOCK_HELD` — field follow-up to #948: two legitimate `improve` invocations colliding on this lock is ordinary, retryable contention, not a broken config file). Use for high-frequency scheduled runs so they don't pile up failures while a longer run is in progress. |
2366
+ | `--require-engines` | Abort (exit 78, before any indexing, lock, or log side effect) if the active strategy would enable a process whose engine or credential cannot be resolved in this process's environment, OR whose endpoint fails a bounded reachability probe — the same probe `akm health`'s `default-llm-engine`/`configured-engines` checks run, once per distinct endpoint. Without this flag, improve degrades gracefully: it skips the affected processes and reports them in the result's `skippedProcesses`. Recommended alongside `--skip-if-locked` for scheduled runs, since the operator's own shell can pass config validation while a scheduler's stripped-down environment (see #953) cannot. |
2367
+ | `--show-prompt` | Print the composed reflect prompt (#952) for one asset and exit — before any lock, index write, or engine dispatch. Requires a fully-qualified asset ref as the scope (`akm improve lessons/my-lesson --show-prompt`); rejected with a type or whole-bundle scope. The default output format is JSON, which carries the prompt as a `prompt` field (escaped into one line) alongside the resolved `engine`/`engineKind`; pass `--format text` to print the prompt itself, unwrapped and readable by eye. |
2359
2368
  | `--sync` / `--no-sync` | Commit (and optionally push) the git-backed primary bundle when the run finishes. Default: on for git-backed bundles (per profile config). |
2360
2369
  | `--push` / `--no-push` | Push after the end-of-run sync commit when writable with a remote configured. `--no-push` commits only, skipping the push. Default: per profile config (`true`). `sync.push` stays outside the autonomy gate — this is a per-run opt-out, not a default change. |
2361
2370
 
@@ -2411,6 +2420,18 @@ on an unavailable credential either — even a strategy left with every process
2411
2420
  disabled this way still returns its plan, with the affected processes in
2412
2421
  `skippedProcesses`.
2413
2422
 
2423
+ `--timeout-ms` is a run-wide wall-clock budget: when it expires, the run
2424
+ cooperatively aborts any in-flight engine request (the same `AbortSignal`
2425
+ every LLM call already honors) instead of waiting out the engine's own,
2426
+ much longer, per-call timeout — the run then finishes and reports normally
2427
+ rather than hanging past its budget. SIGTERM/SIGINT/SIGHUP end a live run
2428
+ the same way, within a short bounded grace period, and the process exits
2429
+ with a stable per-signal code (`143`/`130`/`129`) rather than needing a
2430
+ `kill -9`. If a live run has waited more than a few seconds without any
2431
+ engine response at all, one default-level line ("Still waiting for the
2432
+ first engine response...") is printed so a scheduled run's log is never
2433
+ silently empty while an engine is slow or dead.
2434
+
2414
2435
  For dry runs, `plannedRefs` is the effective post-limit work set, not every
2415
2436
  ref in the requested scope. The `plan` object preserves both views: raw scope
2416
2437
  size and per-gate removals, configured and effective caps, final ranked refs
@@ -2452,6 +2473,17 @@ no model or per-process notices. Neither `--dry-run` nor `--plan` probes
2452
2473
  engine reachability over the network — pair with `akm health --probe` (or the
2453
2474
  default probe-on behavior) to check whether a named engine actually answers.
2454
2475
 
2476
+ `--show-prompt` (#952) is the cheapest way to exercise reflect alone: it
2477
+ builds the exact prompt reflect would send for one asset — the same source
2478
+ resolution, runner selection, feedback/schema-hint/related-lesson/rejected-
2479
+ proposal gathering `akm improve`'s live reflect step uses — and prints it
2480
+ without acquiring a dispatch lease, so it never calls an engine. Add
2481
+ `--format text` (the default JSON/yaml envelope escapes the prompt into one
2482
+ line, which defeats a by-eye read) to confirm by eye that recent feedback is
2483
+ framed as an unverified report to investigate (never a fact to insert
2484
+ verbatim) and that the response contract tells the model never to emit the
2485
+ truncation marker or any content from outside the shown asset.
2486
+
2455
2487
  When reinforced facts need promotion, `knowledge` is the higher-authority
2456
2488
  destination than `memory`. The deterministic search ranking also prefers
2457
2489
  `knowledge` over `memory` hits, including inferred `.derived` memories, when
@@ -411,6 +411,40 @@ apply when unset), for a remote endpoint (`src/llm/embedders/remote.ts`):
411
411
  | `embedding.timeoutMs` | `120000` (120s) | Per-request wall timeout — see below. |
412
412
  | `embedding.concurrency` | `1` loopback / `2` remote | In-flight request window — see below. |
413
413
 
414
+ **Which knob fixed the field's 8k-context overflow, worked examples.** A
415
+ 0.9.15-beta field report described documents estimated under the request
416
+ budget that still tokenized to 8.5k-12.4k real tokens against an
417
+ 8192-token endpoint, because the 4-chars-per-token estimator undercounts
418
+ dense technical text. Three knobs changed shape between beta and this
419
+ release; only one of them makes that overflow structurally unreachable:
420
+
421
+ - `embedding.maxInputTokens: 512` — per-DOCUMENT cap, applied before
422
+ batching. Example: a 6,000-character API reference page is truncated to
423
+ its first ~2,000 characters (512 estimated tokens) before it is ever
424
+ counted toward a request. This is the fix for the original overflow: no
425
+ single document can contribute more than 512 estimated tokens to a
426
+ request, no matter how `maxTokens` or `contextLength` are set.
427
+ - `embedding.maxTokens: 6000` — per-REQUEST budget: how many already-capped
428
+ documents' estimated tokens fit in one HTTP request. Example: with the
429
+ default 512-token document cap, a request packs about 11 documents before
430
+ this budget is reached and the request is sent; if the run's first
431
+ request is still rejected for exceeding the endpoint's real context
432
+ window, akm shrinks this budget to three quarters of its value (floored
433
+ at twice `maxInputTokens`) for every later request in the same run. A
434
+ request-level budget alone cannot stop one oversized document from
435
+ overflowing a request — only the per-document cap above does that.
436
+ - `embedding.contextLength: 8192` — Ollama's `num_ctx` only, forwarded
437
+ verbatim on a native `/api/embed` request. It has no effect on request or
438
+ document sizing, and no effect at all against a non-Ollama endpoint — see
439
+ below for why that used not to be true.
440
+
441
+ A field config of `contextLength: 8192` + `maxTokens: 8000` (the exact
442
+ 0.9.15-beta values from the original report) produces no 400s on 0.9.15:
443
+ `maxInputTokens` (512, new this release) caps every document before it is
444
+ counted, so the original 8.5k-12.4k-token documents that overflowed the
445
+ 8192-token endpoint can never reach the request budget in the first place —
446
+ independent of whatever `maxTokens` or `contextLength` are set to.
447
+
414
448
  `embedding.timeoutMs` (positive integer, default `120000` — 120s) is the
415
449
  budget for a request at the FULL token budget (`embedding.maxTokens`); a
416
450
  local model server on a large, token-budget-bounded batch legitimately takes
@@ -480,6 +514,23 @@ taking about the same wall time as a single one against a healthy endpoint.
480
514
  | --- | --- |
481
515
  | `search.graphBoost.*` | Entity-graph relevance boost: `directBoostPerEntity`/`directBoostCap` (directly related entities), `hopBoostPerEntity`/`hopBoostCap` (multi-hop, capped at `maxHops` ≤ 3), `confidenceMode` (`blend`, the only supported value), `confidenceWeight` (0–1, default `0.2`) |
482
516
 
517
+ ### Curate rerank (#951)
518
+
519
+ An optional cross-encoder rerank pass over `akm curate`'s already-selected
520
+ candidates, via a standalone `/rerank`-style HTTP endpoint (NOT one of the
521
+ `engines.*` `"llm"`/`"agent"` kinds). Disabled by default; a misconfigured
522
+ endpoint, network failure, timeout, or malformed response falls back to
523
+ curate's own ranking unchanged.
524
+
525
+ | Key | Purpose |
526
+ | --- | --- |
527
+ | `search.curateRerank.enabled` | Turn the rerank pass on (default `false`) |
528
+ | `search.curateRerank.endpoint` | Full URL of the reranker's rerank endpoint |
529
+ | `search.curateRerank.model` | Model name sent to the endpoint (optional) |
530
+ | `search.curateRerank.apiKey` | `$VAR`/`secret://<name>` credential reference (optional) |
531
+ | `search.curateRerank.timeoutMs` | Request timeout (default `10000`) |
532
+ | `search.curateRerank.topN` | How many of curate's ranked candidates to send (default `8`, max `50`) |
533
+
483
534
  ## Feedback
484
535
 
485
536
  `feedback` shapes the `akm feedback` taxonomy:
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "akm-cli",
3
- "version": "0.9.15-beta.3",
3
+ "version": "0.9.15-beta.5",
4
4
  "type": "module",
5
5
  "description": "akm (Agent Knowledge Manager) — a portable, local-first capability library for AI agents. Discover, load, share, and improve reusable skills, scripts, workflows, and knowledge across any shell-capable coding agent, including Claude Code, OpenCode, and Cursor.",
6
6
  "keywords": [
@@ -596,6 +596,34 @@
596
596
  }
597
597
  },
598
598
  "additionalProperties": true
599
+ },
600
+ "curateRerank": {
601
+ "type": "object",
602
+ "properties": {
603
+ "enabled": {
604
+ "type": "boolean"
605
+ },
606
+ "endpoint": {
607
+ "type": "string"
608
+ },
609
+ "model": {
610
+ "type": "string",
611
+ "minLength": 1
612
+ },
613
+ "apiKey": {
614
+ "type": "string"
615
+ },
616
+ "timeoutMs": {
617
+ "type": "integer",
618
+ "exclusiveMinimum": 0
619
+ },
620
+ "topN": {
621
+ "type": "integer",
622
+ "exclusiveMinimum": 0,
623
+ "maximum": 50
624
+ }
625
+ },
626
+ "additionalProperties": true
599
627
  }
600
628
  },
601
629
  "additionalProperties": true
@@ -2296,6 +2324,34 @@
2296
2324
  }
2297
2325
  },
2298
2326
  "additionalProperties": true
2327
+ },
2328
+ "curateRerank": {
2329
+ "type": "object",
2330
+ "properties": {
2331
+ "enabled": {
2332
+ "type": "boolean"
2333
+ },
2334
+ "endpoint": {
2335
+ "type": "string"
2336
+ },
2337
+ "model": {
2338
+ "type": "string",
2339
+ "minLength": 1
2340
+ },
2341
+ "apiKey": {
2342
+ "type": "string"
2343
+ },
2344
+ "timeoutMs": {
2345
+ "type": "integer",
2346
+ "exclusiveMinimum": 0
2347
+ },
2348
+ "topN": {
2349
+ "type": "integer",
2350
+ "exclusiveMinimum": 0,
2351
+ "maximum": 50
2352
+ }
2353
+ },
2354
+ "additionalProperties": true
2299
2355
  }
2300
2356
  },
2301
2357
  "additionalProperties": true