libfx 0.0.11 → 0.0.13

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -39,25 +39,50 @@ await agent.close();
39
39
 
40
40
  `apiKey` is required. `model` is optional and defaults to fx's built-in model.
41
41
  Agent configuration uses named options; `env` is reserved for
42
- `createFxTerminal()`.
43
-
44
- `effort` sets the reasoning effort for models that advertise effort levels. It
45
- uses the same vocabulary as the fx CLI's `--effort` flag: `"default"` leaves
46
- the choice to the model, and named levels such as `"low"`, `"medium"`,
47
- `"high"`, or `"xhigh"` request a specific level. A named level is validated
48
- against the selected model's advertised levels at creation; an unsupported
49
- level rejects with an error (code `LIBFX_UNSUPPORTED_EFFORT`) naming the
50
- supported set. When `effort` is omitted or `"default"`, the model default
51
- applies.
52
-
53
- `fast` enables the fast lane for models that advertise one, matching the fx
54
- CLI's `--fast` flag. Enabling it is validated against the selected model at
55
- creation; a model without a fast path rejects with an error (code
56
- `LIBFX_UNSUPPORTED_FAST`). When `fast` is omitted or `false`, the model
57
- default applies.
42
+ `createFxTerminal()`. The canonical model configuration groups the model ID
43
+ and model-specific options:
44
+
45
+ ```js
46
+ const agent = await createFxAgent({
47
+ apiKey,
48
+ model: { id: "anthropic/claude-opus-5.5-fast", effort: "low", fast: true },
49
+ });
50
+ ```
51
+
52
+ A string `model` remains supported as shorthand. Top-level `effort` and
53
+ `fast` are deprecated but remain supported with a string model or no model;
54
+ they cannot be mixed with a model object. New code should use the model object.
55
+
56
+ `model.effort` sets the reasoning effort for models that advertise effort
57
+ levels. It uses the same vocabulary as the fx CLI's `--effort` flag:
58
+ `"default"` leaves the choice to the model; named levels such as `"low"`,
59
+ `"medium"`, `"high"`, or `"xhigh"` request a specific level. A named level is
60
+ validated at creation; an unsupported level rejects with an Error carrying
61
+ `code: "LIBFX_MODEL_UNSUPPORTED_EFFORT"`, `model`, and
62
+ `capability: "effort"`. Its message names the supported set when available.
63
+ Omitting effort or using `"default"` leaves the model default in place.
64
+
65
+ `model.fast` enables the fast lane for models that advertise one, matching the
66
+ fx CLI's `--fast` flag. A model without a fast path rejects at creation with
67
+ `code: "LIBFX_MODEL_UNSUPPORTED_FAST"`, `model`, and `capability: "fast"`.
68
+ Omitting fast or setting it to `false` leaves the model default in place.
69
+
70
+ `model.ultrafast` requests Ultra mode. It is off by default and maps to
71
+ `openai.serviceTier: "ultrafast"` through the Vercel AI Gateway for models
72
+ whose metadata advertises Ultra eligibility. It uses the higher-cost service tier.
73
+ Set it to `true` only after the host has selected an eligible model; an
74
+ unsupported request rejects with `code: "LIBFX_MODEL_UNSUPPORTED_ULTRAFAST"`,
75
+ `model`, and `capability: "ultrafast"`. Set it to `false` to explicitly disable
76
+ an inherited request. Ultra and Fast are mutually exclusive, so `ultrafast:
77
+ true` disables Fast for the agent. Gateway metadata currently marks Astra
78
+ eligible; libfx does not select Ultra automatically.
79
+
80
+ The new codes replace `LIBFX_UNSUPPORTED_EFFORT` and `LIBFX_UNSUPPORTED_FAST`
81
+ for both nested and legacy top-level settings. Callers that check the old codes
82
+ must update their error handling.
58
83
 
59
84
  The host selects the model. Agent creation does not fetch the Gateway model
60
- catalog unless `effort` requests a named level or `fast` is enabled. Prompting
85
+ catalog unless effort requests a named level, fast is enabled, or ultrafast is enabled. Prompting
61
86
  can resolve model capabilities and context capacity through the supplied
62
87
  `fetch`; fx caches that metadata for the agent.
63
88
 
@@ -96,9 +121,11 @@ const result = await turn.result;
96
121
  ```
97
122
 
98
123
  A turn has one event consumer. Breaking out of its iterator cancels the turn;
99
- `turn.cancel()` and `agent.close()` also release blocked output. Transport or
100
- message-decoding failures reject the result instead of returning success with
101
- missing text.
124
+ `turn.cancel()` and `agent.close()` also release blocked output. Embedded agents
125
+ have no implicit model-step cap, so hosts should cancel turns that exceed their
126
+ own budgets. The CLI's `max_agent_steps` setting does not apply to
127
+ `createFxAgent()`. Transport or message-decoding failures reject the result
128
+ instead of returning success with missing text.
102
129
 
103
130
  Native transport buffers at most 8 MiB of output bytes. Unread SDK events apply
104
131
  backpressure at 1 MiB of encoded messages or 256 events. One message can exceed
@@ -106,26 +133,137 @@ that threshold when the queue is empty; an individual encoded ACP message is
106
133
  limited to 64 MiB on both backends. These are transport bounds, not a total
107
134
  answer-size limit or a bound on retained conversation history.
108
135
 
109
- Image blocks use ACP's content shape and carry canonical base64 (no line
110
- wrapping) of a PNG, JPEG, GIF, or WebP payload:
136
+ Image blocks accept a `Blob` or `File` with a non-empty `type`, raw bytes
137
+ (`Uint8Array`, `Buffer`, `ArrayBuffer`, or another typed array) with an
138
+ explicit `mimeType`, or canonical base64 (no line wrapping) with an explicit
139
+ `mimeType`. The payload must be PNG, JPEG, GIF, or WebP. For a Blob, the MIME
140
+ type is inferred from `Blob.type`; an empty type or a conflicting explicit
141
+ `mimeType` rejects the input:
111
142
 
112
143
  ```js
113
144
  const turn = agent.prompt([
114
145
  { type: "text", text: "What does this screenshot show?" },
115
- { type: "image", data: base64Png, mimeType: "image/png" },
146
+ { type: "image", data: file, sourceRef: "uploads:screenshot-1" }, // File or Blob, with file.type
147
+ // Or: { type: "image", data: pngBytes, mimeType: "image/png" }
148
+ // Or: { type: "image", data: base64Png, mimeType: "image/png" }
116
149
  ]);
117
150
  ```
118
151
 
119
- A prompt may contain up to 8 images, each with up to 5 MiB of base64 data,
120
- with at most 8 MiB of image data per prompt; the SDK rejects larger input with
121
- typed `RangeError`s before any request. The kernel then validates the decoded
122
- bytes against the declared `mimeType` and fails the turn with
123
- `Invalid image prompt block` on a mismatch. Images are routed only to models
124
- that advertise image input; for any other model the turn fails with
152
+ An optional `sourceRef` identifies a host-owned original. It must be a non-empty
153
+ UTF-8 string of at most 512 bytes, without ASCII control characters (0–31 or
154
+ DEL). With a reference, you may omit `data` entirely:
155
+
156
+ ```js
157
+ agent.prompt([
158
+ { type: "image", mimeType: "image/png", sourceRef: "uploads:screenshot-1" },
159
+ ]);
160
+ ```
161
+
162
+ Image bytes reach the agent core beside the ACP prompt message rather than
163
+ inside it. Base64 input is decoded before transfer; prompt image bytes are
164
+ encoded as base64 only for the model request. A prompt may contain up to
165
+ 8 images, including reference-only blocks. Each image may contain up to
166
+ 3.75 MiB (3,932,160 bytes) of raw data, with at most 6 MiB of raw image data
167
+ per prompt. Once encoded for the model request, those limits are 5 MiB per
168
+ image and 8 MiB per prompt. The prompt's text, encoded images, and metadata
169
+ must fit within the same 8 MiB frame budget.
170
+
171
+ Without `resizeImage`, the SDK checks Blob size before reading it and the actual
172
+ byte count after reading it. If referenced image data would exceed a per-image,
173
+ aggregate image, or combined frame limit, the SDK sends its MIME type and
174
+ reference without image data. Referenced Blobs known to exceed those limits
175
+ remain unread and host-owned. A reference does not increase the limits;
176
+ unreferenced oversized input rejects with typed `RangeError`s.
177
+
178
+ Base64 that is not canonical throws a `TypeError` from `prompt()`. Raw bytes
179
+ are copied before `prompt()` returns, so the caller can reuse its buffer. Blob
180
+ reads are asynchronous: `prompt()` returns a turn, and read failures reject
181
+ `turn.result`. Cancelling or closing during a Blob read settles the turn without
182
+ sending its prompt. Without `resizeImage`, size errors for unreferenced base64
183
+ and raw-byte input throw synchronously from `prompt()`.
184
+
185
+ To downscale or convert prompt images before they are sent, pass the optional
186
+ `resizeImage` hook when creating the agent. It receives `{ bytes, mimeType }`,
187
+ where `bytes` is a `Uint8Array`, and returns `{ bytes, mimeType }` directly or
188
+ as a promise. Returned `bytes` may be any typed array or `ArrayBuffer`.
189
+ For a Node.js host that already uses `sharp`:
190
+
191
+ ```js
192
+ import sharp from "sharp";
193
+
194
+ const agent = await createFxAgent({
195
+ apiKey,
196
+ model,
197
+ async resizeImage({ bytes }) {
198
+ const png = await sharp(bytes).resize({ width: 1568, withoutEnlargement: true }).png().toBuffer();
199
+ return { bytes: png, mimeType: "image/png" };
200
+ },
201
+ });
202
+ ```
203
+
204
+ In a browser, return the `ArrayBuffer` from `Blob.arrayBuffer()`, including a
205
+ Blob produced by `OffscreenCanvas.convertToBlob()`. The returned bytes are
206
+ copied, so the hook may reuse its buffer. With `resizeImage`, image size limits
207
+ apply to its output rather than its input; image count and frame metadata
208
+ limits still apply. Image prompts with data are prepared asynchronously like
209
+ a Blob prompt, and a failure inside the hook rejects `turn.result`. This
210
+ opt-in preprocessing reads supplied Blob data even when its original size
211
+ exceeds the limits. Reference-only blocks have no bytes and do not call the
212
+ hook or fetch the original. The hook is host-provided, not a built-in image
213
+ converter or a libfx runtime dependency.
214
+
215
+ The kernel sniffs the final bytes and compares them with the claimed MIME type
216
+ for every input form that carries data; a mismatch fails the turn with
217
+ `Invalid image prompt block`. Eligible images reach the model unchanged unless
218
+ your hook changes them. Images are routed only to models that advertise image
219
+ input; for any other model the turn fails with
125
220
  `Image prompts are unavailable for the selected model` and no image bytes
126
- leave the process. Prompt images are retained in checkpoints within the
127
- existing 4 MiB checkpoint bound, so a restored agent can refer to earlier
128
- images on either backend.
221
+ leave the process. Images outside the request's pixel or encoded-size limits,
222
+ and reference-only originals, are withheld with model-visible recovery feedback
223
+ that includes their `sourceRef` when supplied.
224
+
225
+ A source reference is metadata, not an access grant or an automatic fetch.
226
+ Your host owns the original's lifetime, reference resolution, and authorization.
227
+ Expose preparation or retrieval through your existing tools' `execute()`
228
+ callbacks if the agent needs a smaller copy. libfx has no built-in image resizer,
229
+ shell, global converter, or source store, and a reference grants no native
230
+ filesystem authority. If a tool is absent or fails, the original remains
231
+ withheld rather than being sent anyway.
232
+
233
+ For example, a tool can delegate to your app's authorized image store:
234
+
235
+ ```js
236
+ const prepareImage = {
237
+ name: "prepare_image",
238
+ description: "Return a new smaller copy of a host-owned image. Use the request limit for maxSide.",
239
+ inputSchema: {
240
+ type: "object",
241
+ properties: { sourceRef: { type: "string" }, maxSide: { type: "integer" } },
242
+ required: ["sourceRef", "maxSide"],
243
+ },
244
+ async execute({ sourceRef, maxSide }, { signal }) {
245
+ const copy = await imageStore.prepareCopy(sourceRef, { maxSide, signal });
246
+ return {
247
+ type: "libfx.tool-result",
248
+ text: "Prepared a new copy; original unchanged.",
249
+ images: [{ type: "image", data: copy.base64, mimeType: copy.mimeType, sourceRef }],
250
+ };
251
+ },
252
+ };
253
+ // Supply prepareImage in createFxAgent({ tools: [prepareImage], ... }).
254
+ ```
255
+
256
+ `imageStore` is application code, not a libfx API. It must authorize references,
257
+ validate the requested dimensions, and stop conversion when `signal` aborts.
258
+ The returned copy is checked again by the kernel before model submission.
259
+ Typed host-tool images still use base64 `data`; `resizeImage` prepares prompt
260
+ images only.
261
+
262
+ Version 2 checkpoints retain prompt images as raw image blobs and preserve
263
+ source refs within the existing 4 MiB checkpoint bound on both backends. A
264
+ checkpoint does not store a source file or host-owned original for a
265
+ reference-only block. On restoration, resupply your tools and restore the
266
+ sources those refs identify.
129
267
 
130
268
  Only one top-level prompt may run at a time. While it runs,
131
269
  `await turn.steer(text)` appends guidance at the next safe model boundary
@@ -133,8 +271,10 @@ without discarding the in-flight response or completed tool work. Steering also
133
271
  accepts an array of text blocks; image and resource steering blocks are rejected.
134
272
  Each message is limited to 64 KiB, with at most 64 queued messages and 1 MiB of
135
273
  queued steering text. Accepted steering appears as a `user_message` event before
136
- the model's continued output. Calling `steer()` after the turn settles rejects
137
- with `no prompt is running`.
274
+ the model's continued output. For a Blob prompt or one prepared by
275
+ `resizeImage`, steering during that preparation waits for the prompt to be
276
+ sent; cancelling before then rejects the pending steering. Calling `steer()` after the turn settles rejects with
277
+ `no prompt is running`.
138
278
 
139
279
  ```js
140
280
  const turn = agent.prompt("Build the feature.");
@@ -148,8 +288,11 @@ for await (const event of turn) {
148
288
  Cancelling a steered turn drops any guidance that has not reached a safe
149
289
  boundary and releases its queue. Applied guidance is part of the same history
150
290
  turn, so an idle `checkpoint()` includes the full steered conversation.
151
- `checkpoint()` returns opaque, bounded, versioned bytes. Restore them only when
152
- creating a fresh agent:
291
+ `checkpoint()` returns opaque, bounded, versioned bytes. Concurrent calls run
292
+ one at a time, and a call still waiting for an earlier one fails the same way a
293
+ direct call would if a prompt starts or the agent closes first. A newer libfx
294
+ restores checkpoints from older versions, but an older libfx cannot restore one
295
+ written by a newer version. Restore them only when creating a fresh agent:
153
296
 
154
297
  ```js
155
298
  const restored = await createFxAgent({ apiKey, model, checkpoint });
@@ -160,10 +303,10 @@ or a history change. The next prompt can run normally.
160
303
 
161
304
  The checkpoint contains conversation history and usage only. The host owns
162
305
  durable storage and must resupply models, credentials, instructions, tools,
163
- MCP clients, and skill records. Reasoning effort and fast mode are
306
+ MCP clients, and skill records. Reasoning effort, Fast mode, and Ultra mode are
164
307
  agent-creation options and are not stored in a checkpoint: recreate the agent
165
- with new `effort` or `fast` values to change them, the same path as switching
166
- models.
308
+ with new `model.effort`, `model.fast`, or `model.ultrafast` values to change
309
+ them, the same path as switching models.
167
310
 
168
311
  ## Models
169
312
 
@@ -233,6 +376,29 @@ A host tool may use any name, including the kernel's builtin names such as
233
376
  `write_file` and `edit_file`: the kernel routes by the registered executor, so
234
377
  a host-defined `write_file` calls the host's `execute()` rather than the
235
378
  builtin file mutation.
379
+
380
+ Ordinary objects returned by tools are JSON text. To return rich image content,
381
+ use the typed result:
382
+
383
+ ```js
384
+ return {
385
+ type: "libfx.tool-result",
386
+ text: "Original is available through the host image tools.",
387
+ images: [
388
+ { type: "image", mimeType: "image/png", sourceRef: "uploads:screenshot-1" },
389
+ // A prepared copy may also include data: base64Png.
390
+ ],
391
+ };
392
+ ```
393
+
394
+ Tool images require `mimeType` and accept base64 `data`, or a `sourceRef` with no
395
+ `data`. They do not accept raw bytes or Blobs and do not run `resizeImage`.
396
+ References use the same validation, ownership, recovery feedback, and checkpoint
397
+ rules as prompt images. A result may contain up to 8 images, each with up to
398
+ 5 MiB of base64-encoded data. Its text, encoded images, and metadata must fit
399
+ within the 8 MiB result/frame bound. Referenced data is omitted if those bounds
400
+ would overflow; an unreferenced oversized result remains a tool error.
401
+
236
402
  Instructions are limited to 64 KiB of UTF-8 text, including text assembled by
237
403
  the MCP and skills adapters. They are the complete host-owned system context:
238
404
  libfx adds no hidden base prompt, and omitting `instructions` sends no system
package/fx-core.wasm CHANGED
Binary file