@aria-framework/ai 0.24.1 → 0.26.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +72 -13
- package/index.js +333 -316
- package/lmxStatus.js +6 -2
- package/lmxVerify.js +5 -0
- package/package.json +2 -2
- package/providers/lmx.js +99 -27
- package/providers/lmxDiscovery.js +71 -10
- package/providers/openai-compatible.js +19 -3
package/README.md
CHANGED
|
@@ -29,11 +29,14 @@ const ai = createAiClient({
|
|
|
29
29
|
timeoutMs: 60000, maxTokens: 1024, contextTokens: 8192
|
|
30
30
|
}),
|
|
31
31
|
budget: { async assertWithinBudget(cfg, meta) {}, async record(cfg, result, meta) {} }, // optional
|
|
32
|
-
logger: console
|
|
33
|
-
maxRetryAfterMs: 10000 // optional: the longest Retry-After a call will wait out (see 429 below)
|
|
32
|
+
logger: console // optional
|
|
34
33
|
});
|
|
35
34
|
|
|
36
|
-
const r = await ai.complete({
|
|
35
|
+
const r = await ai.complete({
|
|
36
|
+
system, messages, maxTokens: 400,
|
|
37
|
+
schema, // optional: constrained JSON
|
|
38
|
+
// reasoning: 'full', // optional, lmx only: let the model think fully (see Reasoning below)
|
|
39
|
+
});
|
|
37
40
|
// r.text, r.json (when schema), r.model, r.usage.total, r.ms
|
|
38
41
|
```
|
|
39
42
|
|
|
@@ -55,16 +58,16 @@ to write yourself:
|
|
|
55
58
|
|
|
56
59
|
| Contract rule | What the adapter does |
|
|
57
60
|
|---|---|
|
|
58
|
-
| Poll `/status` about every 2 s | One poller per `
|
|
61
|
+
| Poll `/status` about every 2 s | One poller per stack (keyed by `instance`), shared by every engine on it, `unref()`'d. Each read is bounded (`statusTimeoutMs`, default 5 s), concurrent reads share one fetch, and `pollMs` is floored at 500 ms |
|
|
59
62
|
| Route only to `healthy`; `draining` gets nothing new | Selects **for** `healthy`, so a state lmx adds later can never receive work by accident |
|
|
60
63
|
| Engine URLs come from the document | Resolved on every call, and never stored |
|
|
61
64
|
| A status outage is not an inference outage | Keeps routing on the last good document for 3 minutes (`staleMs`), logging loudly |
|
|
62
65
|
| Pin the certificate; never disable verification | Uses an undici `Agent({ connect: { ca } })` per stack. Never `NODE_EXTRA_CA_CERTS`, never `rejectUnauthorized:false` |
|
|
63
|
-
| Choose the reasoning flag from the model | `qwen` → `chat_template_kwargs.enable_thinking=false`, `gpt-oss` → `reasoning_effort:'low'`, anything else gets neither (with a warning) |
|
|
66
|
+
| Choose the reasoning flag from the model | By default thinking is kept to a minimum: `qwen` → `chat_template_kwargs.enable_thinking=false`, `gpt-oss` → `reasoning_effort:'low'`, anything else gets neither (with a warning). `reasoning: 'full'` on the call (since 0.26.0) sends `enable_thinking=true` / `reasoning_effort:'high'` instead. The flag always overrides a raw `extra`; only the named option changes it. See [Reasoning](#reasoning-reasoning-full) |
|
|
64
67
|
| Size requests against `maxInputTokens` (per slot) | `contextTokens` is taken from the engine, not from config |
|
|
65
68
|
| Identity is `(instance, id)`, falling back to `name` | `lmx.engineId` is matched first; `rekeyPlan()` migrates rows when ids are minted or engines renamed |
|
|
66
|
-
| A `429` is about the key, not the engine |
|
|
67
|
-
| The document's `instance` must match | A document for another instance is refused, because engine names collide across stacks |
|
|
69
|
+
| A `429` is about the key, not the engine | Fails fast (since 0.25.0): throws `rate_limit` with `lmxSkip: 'lmx_throttled'` and `retryAfterMs`. The client never waits; whether to wait or move on is the caller's decision |
|
|
70
|
+
| The document's `instance` must match | A document for another instance is refused, because engine names collide across stacks. `lmxVerify` then returns `engines: null`, so nothing downstream (`rekeyPlan`) can act on another stack's list |
|
|
68
71
|
|
|
69
72
|
### Ask the operator for
|
|
70
73
|
|
|
@@ -94,7 +97,8 @@ The adapter takes the status URL exactly as given; it does not append `/status`.
|
|
|
94
97
|
ca: '-----BEGIN CERTIFICATE-----…', // PEM to pin; null = system trust store
|
|
95
98
|
engine: 'advanced', // engine name: always stored, used for display and logs
|
|
96
99
|
engineId: '7c9e6679-…', // engine id when the stack publishes one; null before mint-ids
|
|
97
|
-
pollMs: 2000, staleMs: 180000
|
|
100
|
+
pollMs: 2000, staleMs: 180000, // optional; pollMs is floored at 500, null/undefined = default
|
|
101
|
+
statusTimeoutMs: 5000 // optional; one status read's deadline
|
|
98
102
|
}
|
|
99
103
|
}
|
|
100
104
|
```
|
|
@@ -142,6 +146,14 @@ const { findEngine } = require('@aria-framework/ai');
|
|
|
142
146
|
const engine = findEngine(doc.engines, { id: row.lmx_engine_id || null, name: row.lmx_engine });
|
|
143
147
|
```
|
|
144
148
|
|
|
149
|
+
### Rotating credentials, and removing a stack
|
|
150
|
+
|
|
151
|
+
The poller takes its token, certificate and intervals from the config you pass on **every** call
|
|
152
|
+
(`discoveryFor` reconfigures the live poller since 0.25.0), so a rotated token or a replaced
|
|
153
|
+
certificate takes effect on the next call — no restart. A changed `statusUrl` replaces the poller.
|
|
154
|
+
When you delete or disable a supervisor, call
|
|
155
|
+
`require('@aria-framework/ai/providers/lmx').forgetDiscovery(instance)` so its poller stops.
|
|
156
|
+
|
|
145
157
|
### Errors and routing
|
|
146
158
|
|
|
147
159
|
When an lmx call cannot run, the thrown `AiError` carries `err.lmxSkip`. None of these mean the
|
|
@@ -153,18 +165,65 @@ instead:
|
|
|
153
165
|
| `lmx_not_healthy` | draining, restarting or starting; the supervisor is doing planned work |
|
|
154
166
|
| `lmx_stale` | no fresh status document, so we cannot see the stack |
|
|
155
167
|
| `lmx_unknown_engine` | the engine is not in the document (removed, or hidden from this gateway key) |
|
|
156
|
-
| `lmx_throttled` | a `429` on the key
|
|
168
|
+
| `lmx_throttled` | a `429` on the gateway key. Not retried by the client; `err.retryAfterMs` says how long the gateway asked for. Every engine on the stack shares the key, so the next endpoint on the same stack will likely be throttled too |
|
|
157
169
|
|
|
158
170
|
Failover across several engines for one job is the app's job; lmx provides none. List
|
|
159
171
|
endpoints in preference order and take the first one that serves.
|
|
160
172
|
|
|
173
|
+
### Reasoning (`reasoning: 'full'`)
|
|
174
|
+
|
|
175
|
+
By default the adapter makes the model think as little as its family allows, so a quick answer
|
|
176
|
+
stays quick. Pass `reasoning: 'full'` on a `complete()` call to let the model think fully:
|
|
177
|
+
|
|
178
|
+
```js
|
|
179
|
+
const r = await ai.complete({
|
|
180
|
+
system, messages, schema,
|
|
181
|
+
reasoning: 'full',
|
|
182
|
+
maxTokens: 4096, // at least 4096; see the trap below
|
|
183
|
+
}); // and give the endpoint timeoutMs: 120000 or more
|
|
184
|
+
```
|
|
185
|
+
|
|
186
|
+
**When to use it.** Use it for judgement tasks where the answer depends on weighing evidence,
|
|
187
|
+
such as alert triage. Do not use it for quick summaries, rewrites or extraction. They do not get
|
|
188
|
+
better, only slower.
|
|
189
|
+
|
|
190
|
+
**What it costs.** It is much slower, roughly 30–45 s per answer on lab hardware instead of a few
|
|
191
|
+
seconds, and it spends many more tokens, because the thinking is generated and billed like any
|
|
192
|
+
other output. Budget for both.
|
|
193
|
+
|
|
194
|
+
**The trap.** Thinking comes out of `maxTokens`. If the limit is too small, the model spends all
|
|
195
|
+
of it thinking and returns **an empty answer, with no error from the server** (`finish_reason:
|
|
196
|
+
"length"`). The adapter turns that, and a reply cut off inside an unclosed `<think>` block, into
|
|
197
|
+
a `bad_response` AiError saying it ran out of room (and, when the server says so, that the
|
|
198
|
+
tokens went on internal reasoning), but the work is lost either way. Use **`maxTokens` ≥ 4096** and a
|
|
199
|
+
**`timeoutMs` ≥ 120 s** (120000) on the endpoint, or the deadline cuts the answer off first.
|
|
200
|
+
|
|
201
|
+
**What is sent, per model family** (read from the model the engine is running on each call):
|
|
202
|
+
|
|
203
|
+
| Family | Default (no `reasoning`) | `reasoning: 'full'` |
|
|
204
|
+
|---|---|---|
|
|
205
|
+
| qwen | `{"chat_template_kwargs":{"enable_thinking":false}}` | `{"chat_template_kwargs":{"enable_thinking":true}}` |
|
|
206
|
+
| gpt-oss | `{"reasoning_effort":"low"}` | `{"reasoning_effort":"high"}` |
|
|
207
|
+
| anything else | nothing (warning logged) | nothing (warning logged that `'full'` could not be honoured) |
|
|
208
|
+
|
|
209
|
+
Rules:
|
|
210
|
+
|
|
211
|
+
- `'full'` is the only value. Leaving the option out (or `null`) means the default. Any other value,
|
|
212
|
+
for example `'high'`, `true` or `''`, is refused with `AiError` kind `refused` before any
|
|
213
|
+
request is made, whichever provider the endpoint uses.
|
|
214
|
+
- The flag is merged **after** `extra`, so a raw `extra: { chat_template_kwargs: … }` or
|
|
215
|
+
`extra: { reasoning_effort: … }` cannot change it. Only `reasoning` can.
|
|
216
|
+
- Only `lmx` honours it. On `openai-compatible`, `lmstudio` or `anthropic` the client logs a
|
|
217
|
+
warning, removes the option, and nothing reaches that provider's request body.
|
|
218
|
+
- `REASONING_MODES` (package root) lists the accepted values for an app that validates its own
|
|
219
|
+
settings.
|
|
220
|
+
|
|
161
221
|
### Embeddings
|
|
162
222
|
|
|
163
223
|
Call `client.embed(cfg, texts, { signal })` and pass your own resolved embedding config; it never
|
|
164
|
-
falls back to `resolveConfig()`. It applies the same
|
|
165
|
-
|
|
166
|
-
|
|
167
|
-
Embeddings are not metered against the token budget.
|
|
224
|
+
falls back to `resolveConfig()`. It applies the same policy as `complete()`: a fast dropped
|
|
225
|
+
connection is retried once, a `429` fails fast with `retryAfterMs` on the error, and a config with
|
|
226
|
+
`enabled: false` is refused. Embeddings are not metered against the token budget.
|
|
168
227
|
|
|
169
228
|
On lmx, `embed()` puts the engine's `dimensions` on the config as `embeddingDimensions`. Record the width
|
|
170
229
|
you built an index with and compare it on startup. A model swap changes the width, and an index
|