akm-cli 0.9.15-beta.1 → 0.9.15-beta.3

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -396,21 +396,74 @@ unless a remote `embedding` config is provided.
396
396
  `akm improve`'s memory-inference/consolidate passes when they call an
397
397
  embedding model: `provider`, `endpoint`, `model`, `apiKey` (symbolic
398
398
  reference, same rules as engine `apiKey`), `dimension`, `localModel`,
399
- `maxTokens`, `batchSize`, `chunkSize`, `contextLength`, and
400
- `ollamaOptions.num_ctx`.
401
-
402
- `akm index` keeps a small, fixed number of `/v1/embeddings` requests in
403
- flight at once (a remote endpoint only; the local transformer path is
404
- unaffected): `1` for a loopback endpoint (`localhost`, `127.0.0.0/8`, etc. —
405
- a local model server serves one inference at a time, and parallel requests
406
- thrash it) and `2` for a remote one. This width is not configurable. The
407
- actual throughput knob is request SIZE, not request count: `embedding.batchSize`
408
- (a document-count cap, default 100) together with `embedding.maxTokens` /
409
- `embedding.contextLength` (an estimated token budget per request, default
410
- 8000) control how many documents land in one request — a batch of 16-32
411
- documents takes about the same wall time as a single one against a healthy
412
- endpoint, so growing the batch is where most of the win is, not adding more
413
- concurrent requests.
399
+ `maxInputTokens`, `maxTokens`, `batchSize`, `contextLength`, `timeoutMs`,
400
+ `concurrency`, and `ollamaOptions.num_ctx`.
401
+
402
+ The knobs that bound request/document size and rate, all optional (defaults
403
+ apply when unset), for a remote endpoint (`src/llm/embedders/remote.ts`):
404
+
405
+ | Key | Default | Bounds |
406
+ | --- | --- | --- |
407
+ | `embedding.maxInputTokens` | `512` | Per-DOCUMENT cap, applied before batching (#956). A document's embedded text is truncated to its head (unicode-safe) at this many estimated tokens instead of ever being skipped for size alone — a document is skipped only when its truncated head is empty. |
408
+ | `embedding.maxTokens` | `6000` (`DEFAULT_TOKEN_BUDGET`) | Per-REQUEST token budget: how many (already-capped) documents' estimated tokens fit in one HTTP request. With the 512-token default document cap, a request carries about 11 documents by default. Lowered from 8000 to 6000 (#954): the 4-chars-per-token estimator undercounts dense technical text by 7-55%, so 8000 regularly overshot an 8192-token endpoint's real context window. |
409
+ | `embedding.batchSize` | `100` | Per-REQUEST document-COUNT safety cap, independent of the token budget — guards against many tiny documents packing an oversized request. |
410
+ | `embedding.contextLength` | unset | Ollama's `num_ctx` ONLY, forwarded verbatim as `options.num_ctx` on the native `/api/embed` request. Does **not** feed the request token budget above (#956) — the two used to share this one field, so setting it for the server's context window silently changed request batching too. |
411
+ | `embedding.timeoutMs` | `120000` (120s) | Per-request wall timeout — see below. |
412
+ | `embedding.concurrency` | `1` loopback / `2` remote | In-flight request window — see below. |
413
+
414
+ `embedding.timeoutMs` (positive integer, default `120000` — 120s) is the
415
+ budget for a request at the FULL token budget (`embedding.maxTokens`); a
416
+ local model server on a large, token-budget-bounded batch legitimately takes
417
+ longer than the prior fixed 30s cut off. A smaller request gets a
418
+ proportionally smaller timeout —
419
+ `clamp(timeoutMs × requestTokens / tokenBudget, 30000, timeoutMs)` — so a
420
+ dead endpoint is still detected in seconds on the common case of small
421
+ documents. Set `embedding.timeoutMs` lower to fail fast against a
422
+ known-fast endpoint, or higher for a slow local server on large batches.
423
+
424
+ `embedding.maxTokens` (or its default) is also a run-scoped adaptive
425
+ starting point, not a hard ceiling (#954): on the FIRST rejection of an
426
+ `akm index` run for exceeding the endpoint's context window, akm shrinks
427
+ the request budget to three quarters of its current value — floored at
428
+ twice `embedding.maxInputTokens` — for every request not yet sent, and
429
+ prints one line naming the new value. This never changes the rejected
430
+ request's own split-and-retry (below), never shrinks a second time in the
431
+ same run, and never grows the budget back up. Users who set
432
+ `embedding.maxTokens` explicitly are unaffected by the LOWERED DEFAULT
433
+ above but still benefit from this same-run recovery if their own value
434
+ turns out to be too high for the endpoint.
435
+
436
+ A request TIMEOUT (not a rejection for exceeding the context window) never
437
+ drops its batch immediately: field confirmation showed that once akm
438
+ abandons a timed-out request, the endpoint (e.g. llama-server) keeps
439
+ computing it anyway, so dropping it right away just grows the provider's
440
+ queue while every following batch dies the same way. Instead akm backs off
441
+ (5s, doubling, capped at 60s) and retries the same request once; a second
442
+ timeout splits it in half and retries each half the same way, down to
443
+ individual documents, and a single document that still times out is finally
444
+ skipped (logged at the default `warn` level). After 3 consecutive failures
445
+ at single-document size (timeout or network error), or 3 consecutive
446
+ network errors at any size, the embedding phase stops dispatching further
447
+ requests and reports failure — batches already committed are kept; rerun
448
+ `akm index` once the endpoint is healthy.
449
+
450
+ `akm index` keeps a small number of `/v1/embeddings` requests in flight at
451
+ once (a remote endpoint only; the local transformer path is unaffected):
452
+ `1` for a loopback endpoint (`localhost`, `127.0.0.0/8`, etc. — a local
453
+ model server serves one inference at a time, and parallel requests thrash
454
+ it) and `2` for a remote one, unless `embedding.concurrency` (positive
455
+ integer, 1-16) overrides it. This default holds for the overwhelming
456
+ majority of setups; set the override only for an endpoint that genuinely
457
+ serves parallel requests — a local server started with a multi-slot flag
458
+ (llama.cpp's `--parallel N`, vLLM) — not to "speed up" an ordinary
459
+ single-slot model server, which the default already protects from
460
+ reload-thrash. Request SIZE remains the first throughput lever regardless:
461
+ `embedding.batchSize` (a document-count cap, default 100) together with
462
+ `embedding.maxTokens` (an estimated token budget per request, default 6000
463
+ — NOT `embedding.contextLength`, see the table above) control how many
464
+ documents land in one request — with the default 512-token
465
+ `embedding.maxInputTokens` document cap, that is about 11 documents,
466
+ taking about the same wall time as a single one against a healthy endpoint.
414
467
 
415
468
  ## Search tuning
416
469
 
@@ -674,6 +727,17 @@ embedding calls have since 0.9.13 (#917); resolution order for a single
674
727
  `apiKeyFile`, then `secret://<name>` — though in practice a config sets only
675
728
  one of the three per engine.
676
729
 
730
+ `embedding.apiKey` accepts the same three forms and resolves `secret://` the
731
+ same way, on every path that sends an embedding request: `akm index`
732
+ (including its `bundle update` post-commit embedding pass and the targeted
733
+ re-embed a write command like `akm remember` triggers), `akm improve`'s
734
+ consolidate pass (memory dedup and similarity clustering), and the
735
+ fingerprint-rename canary `akm index` runs when the embedding config
736
+ changes. All of them build the
737
+ provider request through the same `RemoteEmbedder`/`resolveSecret` boundary,
738
+ so a `secret://` reference resolves identically regardless of which command
739
+ triggered the request (#953).
740
+
677
741
  Use `AKM_SQLITE_JOURNAL_MODE=DELETE` or `TRUNCATE` when WAL is unavailable,
678
742
  such as on some NFS/SMB mounts. With the default `WAL` setting, AKM detects a
679
743
  network filesystem for the data directory and falls back to `DELETE`.
@@ -685,3 +749,7 @@ network filesystem for the data directory and falls back to `DELETE`.
685
749
  configuration using `engines`, `defaults.engine`, `defaults.llmEngine`, and
686
750
  `improve.strategies`; AKM deliberately does not infer or rename ambiguous
687
751
  profile identities.
752
+
753
+ `embedding.chunkSize` was never read by anything under `src/` (#954), so a
754
+ config that still sets it is simply ignored — it still loads, unvalidated
755
+ and without warning.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "akm-cli",
3
- "version": "0.9.15-beta.1",
3
+ "version": "0.9.15-beta.3",
4
4
  "type": "module",
5
5
  "description": "akm (Agent Knowledge Manager) — a portable, local-first capability library for AI agents. Discover, load, share, and improve reusable skills, scripts, workflows, and knowledge across any shell-capable coding agent, including Claude Code, OpenCode, and Cursor.",
6
6
  "keywords": [
@@ -211,15 +211,15 @@
211
211
  "type": "string",
212
212
  "minLength": 1
213
213
  },
214
- "maxTokens": {
214
+ "maxInputTokens": {
215
215
  "type": "integer",
216
216
  "exclusiveMinimum": 0
217
217
  },
218
- "batchSize": {
218
+ "maxTokens": {
219
219
  "type": "integer",
220
220
  "exclusiveMinimum": 0
221
221
  },
222
- "chunkSize": {
222
+ "batchSize": {
223
223
  "type": "integer",
224
224
  "exclusiveMinimum": 0
225
225
  },
@@ -236,6 +236,15 @@
236
236
  }
237
237
  },
238
238
  "additionalProperties": true
239
+ },
240
+ "timeoutMs": {
241
+ "type": "integer",
242
+ "exclusiveMinimum": 0
243
+ },
244
+ "concurrency": {
245
+ "type": "integer",
246
+ "exclusiveMinimum": 0,
247
+ "maximum": 16
239
248
  }
240
249
  },
241
250
  "additionalProperties": true
@@ -1902,15 +1911,15 @@
1902
1911
  "type": "string",
1903
1912
  "minLength": 1
1904
1913
  },
1905
- "maxTokens": {
1914
+ "maxInputTokens": {
1906
1915
  "type": "integer",
1907
1916
  "exclusiveMinimum": 0
1908
1917
  },
1909
- "batchSize": {
1918
+ "maxTokens": {
1910
1919
  "type": "integer",
1911
1920
  "exclusiveMinimum": 0
1912
1921
  },
1913
- "chunkSize": {
1922
+ "batchSize": {
1914
1923
  "type": "integer",
1915
1924
  "exclusiveMinimum": 0
1916
1925
  },
@@ -1927,6 +1936,15 @@
1927
1936
  }
1928
1937
  },
1929
1938
  "additionalProperties": true
1939
+ },
1940
+ "timeoutMs": {
1941
+ "type": "integer",
1942
+ "exclusiveMinimum": 0
1943
+ },
1944
+ "concurrency": {
1945
+ "type": "integer",
1946
+ "exclusiveMinimum": 0,
1947
+ "maximum": 16
1930
1948
  }
1931
1949
  },
1932
1950
  "additionalProperties": true
@@ -3471,15 +3489,15 @@
3471
3489
  "type": "string",
3472
3490
  "minLength": 1
3473
3491
  },
3474
- "maxTokens": {
3492
+ "maxInputTokens": {
3475
3493
  "type": "integer",
3476
3494
  "exclusiveMinimum": 0
3477
3495
  },
3478
- "batchSize": {
3496
+ "maxTokens": {
3479
3497
  "type": "integer",
3480
3498
  "exclusiveMinimum": 0
3481
3499
  },
3482
- "chunkSize": {
3500
+ "batchSize": {
3483
3501
  "type": "integer",
3484
3502
  "exclusiveMinimum": 0
3485
3503
  },
@@ -3496,6 +3514,15 @@
3496
3514
  }
3497
3515
  },
3498
3516
  "additionalProperties": true
3517
+ },
3518
+ "timeoutMs": {
3519
+ "type": "integer",
3520
+ "exclusiveMinimum": 0
3521
+ },
3522
+ "concurrency": {
3523
+ "type": "integer",
3524
+ "exclusiveMinimum": 0,
3525
+ "maximum": 16
3499
3526
  }
3500
3527
  },
3501
3528
  "additionalProperties": true