@aria-framework/ai 0.24.1 → 0.26.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -29,11 +29,14 @@ const ai = createAiClient({
29
29
  timeoutMs: 60000, maxTokens: 1024, contextTokens: 8192
30
30
  }),
31
31
  budget: { async assertWithinBudget(cfg, meta) {}, async record(cfg, result, meta) {} }, // optional
32
- logger: console, // optional
33
- maxRetryAfterMs: 10000 // optional: the longest Retry-After a call will wait out (see 429 below)
32
+ logger: console // optional
34
33
  });
35
34
 
36
- const r = await ai.complete({ system, messages, maxTokens: 400, schema /* optional */ });
35
+ const r = await ai.complete({
36
+ system, messages, maxTokens: 400,
37
+ schema, // optional: constrained JSON
38
+ // reasoning: 'full', // optional, lmx only: let the model think fully (see Reasoning below)
39
+ });
37
40
  // r.text, r.json (when schema), r.model, r.usage.total, r.ms
38
41
  ```
39
42
 
@@ -55,16 +58,16 @@ to write yourself:
55
58
 
56
59
  | Contract rule | What the adapter does |
57
60
  |---|---|
58
- | Poll `/status` about every 2 s | One poller per `(instance, statusUrl)`, shared by every engine on that stack, with `unref()` called on it |
61
+ | Poll `/status` about every 2 s | One poller per stack (keyed by `instance`), shared by every engine on it, `unref()`'d. Each read is bounded (`statusTimeoutMs`, default 5 s), concurrent reads share one fetch, and `pollMs` is floored at 500 ms |
59
62
  | Route only to `healthy`; `draining` gets nothing new | Selects **for** `healthy`, so a state lmx adds later can never receive work by accident |
60
63
  | Engine URLs come from the document | Resolved on every call, and never stored |
61
64
  | A status outage is not an inference outage | Keeps routing on the last good document for 3 minutes (`staleMs`), logging loudly |
62
65
  | Pin the certificate; never disable verification | Uses an undici `Agent({ connect: { ca } })` per stack. Never `NODE_EXTRA_CA_CERTS`, never `rejectUnauthorized:false` |
63
- | Choose the reasoning flag from the model | `qwen` → `chat_template_kwargs.enable_thinking=false`, `gpt-oss` → `reasoning_effort:'low'`, anything else gets neither (with a warning) |
66
+ | Choose the reasoning flag from the model | By default thinking is kept to a minimum: `qwen` → `chat_template_kwargs.enable_thinking=false`, `gpt-oss` → `reasoning_effort:'low'`, anything else gets neither (with a warning). `reasoning: 'full'` on the call (since 0.26.0) sends `enable_thinking=true` / `reasoning_effort:'high'` instead. The flag always overrides a raw `extra`; only the named option changes it. See [Reasoning](#reasoning-reasoning-full) |
64
67
  | Size requests against `maxInputTokens` (per slot) | `contextTokens` is taken from the engine, not from config |
65
68
  | Identity is `(instance, id)`, falling back to `name` | `lmx.engineId` is matched first; `rekeyPlan()` migrates rows when ids are minted or engines renamed |
66
- | A `429` is about the key, not the engine | Waits `Retry-After` once (up to `maxRetryAfterMs`), then throws with `lmxSkip: 'lmx_throttled'` |
67
- | The document's `instance` must match | A document for another instance is refused, because engine names collide across stacks |
69
+ | A `429` is about the key, not the engine | Fails fast (since 0.25.0): throws `rate_limit` with `lmxSkip: 'lmx_throttled'` and `retryAfterMs`. The client never waits; whether to wait or move on is the caller's decision |
70
+ | The document's `instance` must match | A document for another instance is refused, because engine names collide across stacks. `lmxVerify` then returns `engines: null`, so nothing downstream (`rekeyPlan`) can act on another stack's list |
68
71
 
69
72
  ### Ask the operator for
70
73
 
@@ -94,7 +97,8 @@ The adapter takes the status URL exactly as given; it does not append `/status`.
94
97
  ca: '-----BEGIN CERTIFICATE-----…', // PEM to pin; null = system trust store
95
98
  engine: 'advanced', // engine name: always stored, used for display and logs
96
99
  engineId: '7c9e6679-…', // engine id when the stack publishes one; null before mint-ids
97
- pollMs: 2000, staleMs: 180000 // optional
100
+ pollMs: 2000, staleMs: 180000, // optional; pollMs is floored at 500, null/undefined = default
101
+ statusTimeoutMs: 5000 // optional; one status read's deadline
98
102
  }
99
103
  }
100
104
  ```
@@ -142,6 +146,14 @@ const { findEngine } = require('@aria-framework/ai');
142
146
  const engine = findEngine(doc.engines, { id: row.lmx_engine_id || null, name: row.lmx_engine });
143
147
  ```
144
148
 
149
+ ### Rotating credentials, and removing a stack
150
+
151
+ The poller takes its token, certificate and intervals from the config you pass on **every** call
152
+ (`discoveryFor` reconfigures the live poller since 0.25.0), so a rotated token or a replaced
153
+ certificate takes effect on the next call — no restart. A changed `statusUrl` replaces the poller.
154
+ When you delete or disable a supervisor, call
155
+ `require('@aria-framework/ai/providers/lmx').forgetDiscovery(instance)` so its poller stops.
156
+
145
157
  ### Errors and routing
146
158
 
147
159
  When an lmx call cannot run, the thrown `AiError` carries `err.lmxSkip`. None of these mean the
@@ -153,18 +165,65 @@ instead:
153
165
  | `lmx_not_healthy` | draining, restarting or starting; the supervisor is doing planned work |
154
166
  | `lmx_stale` | no fresh status document, so we cannot see the stack |
155
167
  | `lmx_unknown_engine` | the engine is not in the document (removed, or hidden from this gateway key) |
156
- | `lmx_throttled` | a `429` on the key; the one `Retry-After` wait has already been spent |
168
+ | `lmx_throttled` | a `429` on the gateway key. Not retried by the client; `err.retryAfterMs` says how long the gateway asked for. Every engine on the stack shares the key, so the next endpoint on the same stack will likely be throttled too |
157
169
 
158
170
  Failover across several engines for one job is the app's job; lmx provides none. List
159
171
  endpoints in preference order and take the first one that serves.
160
172
 
173
+ ### Reasoning (`reasoning: 'full'`)
174
+
175
+ By default the adapter makes the model think as little as its family allows, so a quick answer
176
+ stays quick. Pass `reasoning: 'full'` on a `complete()` call to let the model think fully:
177
+
178
+ ```js
179
+ const r = await ai.complete({
180
+ system, messages, schema,
181
+ reasoning: 'full',
182
+ maxTokens: 4096, // at least 4096; see the trap below
183
+ }); // and give the endpoint timeoutMs: 120000 or more
184
+ ```
185
+
186
+ **When to use it.** Use it for judgement tasks where the answer depends on weighing evidence,
187
+ such as alert triage. Do not use it for quick summaries, rewrites or extraction. They do not get
188
+ better, only slower.
189
+
190
+ **What it costs.** It is much slower, roughly 30–45 s per answer on lab hardware instead of a few
191
+ seconds, and it spends many more tokens, because the thinking is generated and billed like any
192
+ other output. Budget for both.
193
+
194
+ **The trap.** Thinking comes out of `maxTokens`. If the limit is too small, the model spends all
195
+ of it thinking and returns **an empty answer, with no error from the server** (`finish_reason:
196
+ "length"`). The adapter turns that, and a reply cut off inside an unclosed `<think>` block, into
197
+ a `bad_response` AiError saying it ran out of room (and, when the server says so, that the
198
+ tokens went on internal reasoning), but the work is lost either way. Use **`maxTokens` ≥ 4096** and a
199
+ **`timeoutMs` ≥ 120 s** (120000) on the endpoint, or the deadline cuts the answer off first.
200
+
201
+ **What is sent, per model family** (read from the model the engine is running on each call):
202
+
203
+ | Family | Default (no `reasoning`) | `reasoning: 'full'` |
204
+ |---|---|---|
205
+ | qwen | `{"chat_template_kwargs":{"enable_thinking":false}}` | `{"chat_template_kwargs":{"enable_thinking":true}}` |
206
+ | gpt-oss | `{"reasoning_effort":"low"}` | `{"reasoning_effort":"high"}` |
207
+ | anything else | nothing (warning logged) | nothing (warning logged that `'full'` could not be honoured) |
208
+
209
+ Rules:
210
+
211
+ - `'full'` is the only value. Leaving the option out (or `null`) means the default. Any other value,
212
+ for example `'high'`, `true` or `''`, is refused with `AiError` kind `refused` before any
213
+ request is made, whichever provider the endpoint uses.
214
+ - The flag is merged **after** `extra`, so a raw `extra: { chat_template_kwargs: … }` or
215
+ `extra: { reasoning_effort: … }` cannot change it. Only `reasoning` can.
216
+ - Only `lmx` honours it. On `openai-compatible`, `lmstudio` or `anthropic` the client logs a
217
+ warning, removes the option, and nothing reaches that provider's request body.
218
+ - `REASONING_MODES` (package root) lists the accepted values for an app that validates its own
219
+ settings.
220
+
161
221
  ### Embeddings
162
222
 
163
223
  Call `client.embed(cfg, texts, { signal })` and pass your own resolved embedding config; it never
164
- falls back to `resolveConfig()`. It applies the same retry as `complete()`: a `429` with
165
- `Retry-After` is waited out once, and a fast dropped connection is retried once. Calling the
166
- provider adapter's `embed()` directly skips both, and a throttled batch then fails outright.
167
- Embeddings are not metered against the token budget.
224
+ falls back to `resolveConfig()`. It applies the same policy as `complete()`: a fast dropped
225
+ connection is retried once, a `429` fails fast with `retryAfterMs` on the error, and a config with
226
+ `enabled: false` is refused. Embeddings are not metered against the token budget.
168
227
 
169
228
  On lmx, `embed()` puts the engine's `dimensions` on the config as `embeddingDimensions`. Record the width
170
229
  you built an index with and compare it on startup. A model swap changes the width, and an index