libfx 0.0.12 → 0.0.13

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -66,12 +66,23 @@ Omitting effort or using `"default"` leaves the model default in place.
66
66
  fx CLI's `--fast` flag. A model without a fast path rejects at creation with
67
67
  `code: "LIBFX_MODEL_UNSUPPORTED_FAST"`, `model`, and `capability: "fast"`.
68
68
  Omitting fast or setting it to `false` leaves the model default in place.
69
+
70
+ `model.ultrafast` requests Ultra mode. It is off by default and maps to
71
+ `openai.serviceTier: "ultrafast"` through the Vercel AI Gateway for models
72
+ whose metadata advertises Ultra eligibility. It uses the higher-cost service tier.
73
+ Set it to `true` only after the host has selected an eligible model; an
74
+ unsupported request rejects with `code: "LIBFX_MODEL_UNSUPPORTED_ULTRAFAST"`,
75
+ `model`, and `capability: "ultrafast"`. Set it to `false` to explicitly disable
76
+ an inherited request. Ultra and Fast are mutually exclusive, so `ultrafast:
77
+ true` disables Fast for the agent. Gateway metadata currently marks Astra
78
+ eligible; libfx does not select Ultra automatically.
79
+
69
80
  The new codes replace `LIBFX_UNSUPPORTED_EFFORT` and `LIBFX_UNSUPPORTED_FAST`
70
81
  for both nested and legacy top-level settings. Callers that check the old codes
71
82
  must update their error handling.
72
83
 
73
84
  The host selects the model. Agent creation does not fetch the Gateway model
74
- catalog unless effort requests a named level or fast is enabled. Prompting
85
+ catalog unless effort requests a named level, fast is enabled, or ultrafast is enabled. Prompting
75
86
  can resolve model capabilities and context capacity through the supplied
76
87
  `fetch`; fx caches that metadata for the agent.
77
88
 
@@ -122,36 +133,137 @@ that threshold when the queue is empty; an individual encoded ACP message is
122
133
  limited to 64 MiB on both backends. These are transport bounds, not a total
123
134
  answer-size limit or a bound on retained conversation history.
124
135
 
125
- Image blocks accept a `Blob` or `File` with a non-empty `type`, or the
126
- existing canonical base64 (no line wrapping) and explicit `mimeType` of a PNG,
127
- JPEG, GIF, or WebP payload:
136
+ Image blocks accept a `Blob` or `File` with a non-empty `type`, raw bytes
137
+ (`Uint8Array`, `Buffer`, `ArrayBuffer`, or another typed array) with an
138
+ explicit `mimeType`, or canonical base64 (no line wrapping) with an explicit
139
+ `mimeType`. The payload must be PNG, JPEG, GIF, or WebP. For a Blob, the MIME
140
+ type is inferred from `Blob.type`; an empty type or a conflicting explicit
141
+ `mimeType` rejects the input:
128
142
 
129
143
  ```js
130
144
  const turn = agent.prompt([
131
145
  { type: "text", text: "What does this screenshot show?" },
132
- { type: "image", data: file }, // File or Blob, with file.type
146
+ { type: "image", data: file, sourceRef: "uploads:screenshot-1" }, // File or Blob, with file.type
147
+ // Or: { type: "image", data: pngBytes, mimeType: "image/png" }
133
148
  // Or: { type: "image", data: base64Png, mimeType: "image/png" }
134
149
  ]);
135
150
  ```
136
151
 
137
- A prompt may contain up to 8 images, each with up to 5 MiB of base64 data,
138
- with at most 8 MiB of image data per prompt. The SDK checks Blob size before
139
- reading it, encodes it for the same ACP wire format, and rejects larger input
140
- with typed `RangeError`s. The total frame size is checked before reading a
141
- Blob, and the actual byte count is checked before encoding it. Blob reads are
142
- asynchronous: `prompt()` returns a turn, and read failures reject
143
- `turn.result`. Cancelling or closing while a Blob is being read settles the
144
- turn without sending its prompt. For base64 input, size errors still throw
145
- synchronously from `prompt()`.
146
-
147
- The kernel sniffs decoded bytes and compares them with the claimed MIME type
148
- for both input forms; a mismatch fails the turn with
149
- `Invalid image prompt block`. Images are routed only to models that advertise
150
- image input; for any other model the turn fails with
152
+ An optional `sourceRef` identifies a host-owned original. It must be a non-empty
153
+ UTF-8 string of at most 512 bytes, without ASCII control characters (0–31 or
154
+ DEL). With a reference, you may omit `data` entirely:
155
+
156
+ ```js
157
+ agent.prompt([
158
+ { type: "image", mimeType: "image/png", sourceRef: "uploads:screenshot-1" },
159
+ ]);
160
+ ```
161
+
162
+ Image bytes reach the agent core beside the ACP prompt message rather than
163
+ inside it. Base64 input is decoded before transfer; prompt image bytes are
164
+ encoded as base64 only for the model request. A prompt may contain up to
165
+ 8 images, including reference-only blocks. Each image may contain up to
166
+ 3.75 MiB (3,932,160 bytes) of raw data, with at most 6 MiB of raw image data
167
+ per prompt. Once encoded for the model request, those limits are 5 MiB per
168
+ image and 8 MiB per prompt. The prompt's text, encoded images, and metadata
169
+ must fit within the same 8 MiB frame budget.
170
+
171
+ Without `resizeImage`, the SDK checks Blob size before reading it and the actual
172
+ byte count after reading it. If referenced image data would exceed a per-image,
173
+ aggregate image, or combined frame limit, the SDK sends its MIME type and
174
+ reference without image data. Referenced Blobs known to exceed those limits
175
+ remain unread and host-owned. A reference does not increase the limits;
176
+ unreferenced oversized input rejects with typed `RangeError`s.
177
+
178
+ Base64 that is not canonical throws a `TypeError` from `prompt()`. Raw bytes
179
+ are copied before `prompt()` returns, so the caller can reuse its buffer. Blob
180
+ reads are asynchronous: `prompt()` returns a turn, and read failures reject
181
+ `turn.result`. Cancelling or closing during a Blob read settles the turn without
182
+ sending its prompt. Without `resizeImage`, size errors for unreferenced base64
183
+ and raw-byte input throw synchronously from `prompt()`.
184
+
185
+ To downscale or convert prompt images before they are sent, pass the optional
186
+ `resizeImage` hook when creating the agent. It receives `{ bytes, mimeType }`,
187
+ where `bytes` is a `Uint8Array`, and returns `{ bytes, mimeType }` directly or
188
+ as a promise. Returned `bytes` may be any typed array or `ArrayBuffer`.
189
+ For a Node.js host that already uses `sharp`:
190
+
191
+ ```js
192
+ import sharp from "sharp";
193
+
194
+ const agent = await createFxAgent({
195
+ apiKey,
196
+ model,
197
+ async resizeImage({ bytes }) {
198
+ const png = await sharp(bytes).resize({ width: 1568, withoutEnlargement: true }).png().toBuffer();
199
+ return { bytes: png, mimeType: "image/png" };
200
+ },
201
+ });
202
+ ```
203
+
204
+ In a browser, return the `ArrayBuffer` from `Blob.arrayBuffer()`, including a
205
+ Blob produced by `OffscreenCanvas.convertToBlob()`. The returned bytes are
206
+ copied, so the hook may reuse its buffer. With `resizeImage`, image size limits
207
+ apply to its output rather than its input; image count and frame metadata
208
+ limits still apply. Image prompts with data are prepared asynchronously like
209
+ a Blob prompt, and a failure inside the hook rejects `turn.result`. This
210
+ opt-in preprocessing reads supplied Blob data even when its original size
211
+ exceeds the limits. Reference-only blocks have no bytes and do not call the
212
+ hook or fetch the original. The hook is host-provided, not a built-in image
213
+ converter or a libfx runtime dependency.
214
+
215
+ The kernel sniffs the final bytes and compares them with the claimed MIME type
216
+ for every input form that carries data; a mismatch fails the turn with
217
+ `Invalid image prompt block`. Eligible images reach the model unchanged unless
218
+ your hook changes them. Images are routed only to models that advertise image
219
+ input; for any other model the turn fails with
151
220
  `Image prompts are unavailable for the selected model` and no image bytes
152
- leave the process. Prompt images are retained in checkpoints within the
153
- existing 4 MiB checkpoint bound, so a restored agent can refer to earlier
154
- images on either backend.
221
+ leave the process. Images outside the request's pixel or encoded-size limits,
222
+ and reference-only originals, are withheld with model-visible recovery feedback
223
+ that includes their `sourceRef` when supplied.
224
+
225
+ A source reference is metadata, not an access grant or an automatic fetch.
226
+ Your host owns the original's lifetime, reference resolution, and authorization.
227
+ Expose preparation or retrieval through your existing tools' `execute()`
228
+ callbacks if the agent needs a smaller copy. libfx has no built-in image resizer,
229
+ shell, global converter, or source store, and a reference grants no native
230
+ filesystem authority. If a tool is absent or fails, the original remains
231
+ withheld rather than being sent anyway.
232
+
233
+ For example, a tool can delegate to your app's authorized image store:
234
+
235
+ ```js
236
+ const prepareImage = {
237
+ name: "prepare_image",
238
+ description: "Return a new smaller copy of a host-owned image. Use the request limit for maxSide.",
239
+ inputSchema: {
240
+ type: "object",
241
+ properties: { sourceRef: { type: "string" }, maxSide: { type: "integer" } },
242
+ required: ["sourceRef", "maxSide"],
243
+ },
244
+ async execute({ sourceRef, maxSide }, { signal }) {
245
+ const copy = await imageStore.prepareCopy(sourceRef, { maxSide, signal });
246
+ return {
247
+ type: "libfx.tool-result",
248
+ text: "Prepared a new copy; original unchanged.",
249
+ images: [{ type: "image", data: copy.base64, mimeType: copy.mimeType, sourceRef }],
250
+ };
251
+ },
252
+ };
253
+ // Supply prepareImage in createFxAgent({ tools: [prepareImage], ... }).
254
+ ```
255
+
256
+ `imageStore` is application code, not a libfx API. It must authorize references,
257
+ validate the requested dimensions, and stop conversion when `signal` aborts.
258
+ The returned copy is checked again by the kernel before model submission.
259
+ Typed host-tool images still use base64 `data`; `resizeImage` prepares prompt
260
+ images only.
261
+
262
+ Version 2 checkpoints retain prompt images as raw image blobs and preserve
263
+ source refs within the existing 4 MiB checkpoint bound on both backends. A
264
+ checkpoint does not store a source file or host-owned original for a
265
+ reference-only block. On restoration, resupply your tools and restore the
266
+ sources those refs identify.
155
267
 
156
268
  Only one top-level prompt may run at a time. While it runs,
157
269
  `await turn.steer(text)` appends guidance at the next safe model boundary
@@ -159,9 +271,9 @@ without discarding the in-flight response or completed tool work. Steering also
159
271
  accepts an array of text blocks; image and resource steering blocks are rejected.
160
272
  Each message is limited to 64 KiB, with at most 64 queued messages and 1 MiB of
161
273
  queued steering text. Accepted steering appears as a `user_message` event before
162
- the model's continued output. For a Blob prompt, steering during the read
163
- waits for the prompt to be sent; cancelling before then rejects the pending
164
- steering. Calling `steer()` after the turn settles rejects with
274
+ the model's continued output. For a Blob prompt or one prepared by
275
+ `resizeImage`, steering during that preparation waits for the prompt to be
276
+ sent; cancelling before then rejects the pending steering. Calling `steer()` after the turn settles rejects with
165
277
  `no prompt is running`.
166
278
 
167
279
  ```js
@@ -176,8 +288,11 @@ for await (const event of turn) {
176
288
  Cancelling a steered turn drops any guidance that has not reached a safe
177
289
  boundary and releases its queue. Applied guidance is part of the same history
178
290
  turn, so an idle `checkpoint()` includes the full steered conversation.
179
- `checkpoint()` returns opaque, bounded, versioned bytes. Restore them only when
180
- creating a fresh agent:
291
+ `checkpoint()` returns opaque, bounded, versioned bytes. Concurrent calls run
292
+ one at a time, and a call still waiting for an earlier one fails the same way a
293
+ direct call would if a prompt starts or the agent closes first. A newer libfx
294
+ restores checkpoints from older versions, but an older libfx cannot restore one
295
+ written by a newer version. Restore them only when creating a fresh agent:
181
296
 
182
297
  ```js
183
298
  const restored = await createFxAgent({ apiKey, model, checkpoint });
@@ -188,10 +303,10 @@ or a history change. The next prompt can run normally.
188
303
 
189
304
  The checkpoint contains conversation history and usage only. The host owns
190
305
  durable storage and must resupply models, credentials, instructions, tools,
191
- MCP clients, and skill records. Reasoning effort and fast mode are
306
+ MCP clients, and skill records. Reasoning effort, Fast mode, and Ultra mode are
192
307
  agent-creation options and are not stored in a checkpoint: recreate the agent
193
- with new `model.effort` or `model.fast` values to change them, the same path as
194
- switching models.
308
+ with new `model.effort`, `model.fast`, or `model.ultrafast` values to change
309
+ them, the same path as switching models.
195
310
 
196
311
  ## Models
197
312
 
@@ -261,6 +376,29 @@ A host tool may use any name, including the kernel's builtin names such as
261
376
  `write_file` and `edit_file`: the kernel routes by the registered executor, so
262
377
  a host-defined `write_file` calls the host's `execute()` rather than the
263
378
  builtin file mutation.
379
+
380
+ Ordinary objects returned by tools are JSON text. To return rich image content,
381
+ use the typed result:
382
+
383
+ ```js
384
+ return {
385
+ type: "libfx.tool-result",
386
+ text: "Original is available through the host image tools.",
387
+ images: [
388
+ { type: "image", mimeType: "image/png", sourceRef: "uploads:screenshot-1" },
389
+ // A prepared copy may also include data: base64Png.
390
+ ],
391
+ };
392
+ ```
393
+
394
+ Tool images require `mimeType` and accept base64 `data`, or a `sourceRef` with no
395
+ `data`. They do not accept raw bytes or Blobs and do not run `resizeImage`.
396
+ References use the same validation, ownership, recovery feedback, and checkpoint
397
+ rules as prompt images. A result may contain up to 8 images, each with up to
398
+ 5 MiB of base64-encoded data. Its text, encoded images, and metadata must fit
399
+ within the 8 MiB result/frame bound. Referenced data is omitted if those bounds
400
+ would overflow; an unreferenced oversized result remains a tool error.
401
+
264
402
  Instructions are limited to 64 KiB of UTF-8 text, including text assembled by
265
403
  the MCP and skills adapters. They are the complete host-owned system context:
266
404
  libfx adds no hidden base prompt, and omitting `instructions` sends no system
package/fx-core.wasm CHANGED
Binary file