@shardflux/sdk 0.6.2 → 0.8.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -3,7 +3,115 @@
3
3
  Every API the README shows is available from the version named here. Below 1.0, a minor release may break
4
4
  compatibility; breaking changes are marked **Breaking**.
5
5
 
6
- ## 0.6.2 (not yet published; npm `latest` is 0.6.1)
6
+ ## 0.8.0 (2026-09-28)
7
+
8
+ Types only; nothing changes at run time and the API is unchanged.
9
+
10
+ ### Provider tool exports type-check without casts
11
+
12
+ - `toAnthropicTools(tools)` is assignable to `Anthropic.Tool[]` (`@anthropic-ai/sdk`) and `toOpenAITools(tools, { api:
13
+ 'responses' })` to `OpenAI.Responses.FunctionTool[]` (`openai`) under TypeScript's `strict` checks, with no `as`
14
+ casts. `toOpenAITools(tools)` (and `{ api: 'chat' }`) is assignable to `OpenAI.Chat.ChatCompletionTool[]`.
15
+ - `toOpenAITools` has one return type per format (overloads): an `api` known only at run time still returns either
16
+ array, as before.
17
+ - `JsonSchema` is a type alias instead of an interface, so a schema is assignable to the providers' open schema types
18
+ (`{ [key: string]: unknown }`). An object literal typed as `JsonSchema` still rejects a misspelt keyword.
19
+ - `executeToolCall(tools, call)` takes `input` as `unknown`, as the Anthropic SDK's `ToolUseBlock` types it, so a
20
+ `tool_use` block is passed as it is (no `block.input as Record<string, unknown>`); `execute` validates it as before.
21
+ - New exported types for the three formats: `AnthropicToolDefinition`, `OpenAIChatToolDefinition`,
22
+ `OpenAIResponsesToolDefinition`.
23
+ - Checked at compile time against `@anthropic-ai/sdk` 0.128.0 and `openai` 7.23.0 (devDependencies only; the SDK still
24
+ has no runtime dependencies).
25
+
26
+ ## 0.7.0 (2026-09-28)
27
+
28
+ Needs an API with the template editor (contracts §24); every new field is additive and older fields are unchanged.
29
+
30
+ ### Template editor: build a template from template.yaml
31
+
32
+ - `templates.buildFromFile(path, { templateSlug, autoPublish?, description?, displayName?, acknowledgedScanFindings?,
33
+ organizationId?, wait?, onProgress?, root?, parseYaml? })` (Node only): reads template.yaml (or a `.json` file with
34
+ the same document), uploads each local `from` path of `build.files` (folders packed as a reproducible tar, files as
35
+ they are; each distinct path once; bytes the organization already has are not sent), creates the recipe v2 build and
36
+ with `wait` follows it. Returns `{ build, recipe, uploads }`. Progress events: `pack`, `upload`, `build`.
37
+ `templates.buildFromRecipe(doc, { baseDir, ... })` does the same for a document in memory.
38
+ - YAML is read with the optional peer dependency `yaml` (`npm install yaml`) or `parseYaml`; JSON needs nothing. The
39
+ SDK keeps no runtime dependencies, and the Node-only code is loaded with dynamic imports, so browser bundles are
40
+ unaffected.
41
+ - The folder tar is byte-identical to the Python SDK's (`shardflux` 0.3.0) for the same folder: members sorted by UTF-8
42
+ path, mtime 0, uid/gid 0 without names, permission bits kept, symlinks as symlinks, pax records for long or
43
+ non-ASCII names. Absolute or escaping symlinks, devices, FIFOs and sockets are refused before any request, as the
44
+ build host would.
45
+ - `TemplateFileError` (a file that cannot be read, parsed or packed; nothing was sent) and `TemplateUploadError` (the
46
+ storage refused a PUT: `status`, `code` such as `BadDigest`; never the presigned URL). Exported: `packDirectory`,
47
+ `readTemplateFile`, `parseTemplateText`, the tar writer (`tarHeader`, `tarPadding`, `tarEnd`).
48
+
49
+ ### Template editor: the API surface (contracts §24.6)
50
+
51
+ - `templates.uploads.request({ sha256, size, kind })`, `templates.uploads.put(bytes | Blob | stream, { kind, sha256?,
52
+ size? })` (PUT with exactly the presigned headers, skipped when the organization has the bytes, confirmed after) and
53
+ `templates.uploads.putPath(path)` (Node).
54
+ - `templates.builds.create()` takes recipe v2 (`TemplateRecipeV2`), `description` and `acknowledgedScanFindings`.
55
+ `TemplateBuild.denied_hosts`. **Breaking (types only):** `TemplateBuild.recipe` is `{dockerfile}` or the stored
56
+ recipe v2 (`TemplateBuildRecipeV2`); narrow with `'dockerfile' in build.recipe` before reading `dockerfile`.
57
+ `waitForBuild()` takes `onChange`.
58
+ - `templates.versions.recipe(slug, version)`: the recipe and settings a version was built from, ready to build again.
59
+ - `templates.versionTestInstances.create(slug, version, { inputs?, key?, caps?, ... })`: a session workspace on a
60
+ registered version, published or not.
61
+ - `templates.languages(base)`: the languages and versions a base offers `build.languages` (`included` when the base
62
+ already has one).
63
+ - `templates.packages.search(ecosystem, query, { base?, limit? })` and `templates.packages.get(ecosystem, name)`.
64
+ - `workspaces.open({ inputs })`, `workspace.inputs()` / `workspaces.inputs(id)`, `workspace.startup` (start commands
65
+ and services: pending, running, ready or failed with the step, exit code and output tail).
66
+ - Drafts: `create({ displayName, inputs })`, `openTestInstance({ inputs })`, `publish({ settings })`;
67
+ `saveAsTemplate({ settings })`.
68
+ - Views: `TemplateVersion.settings`, `TemplateSummary.category` / `TemplateDetail.category`,
69
+ `WorkspaceEgressPolicy.template_egress` and `effective_policy`. Types: `TemplateSettings`, `TemplateSettingsInput`,
70
+ `TemplateInput`, `TemplateStartCommand`, `TemplateService`, `TemplateEgressDefault`, `TemplateUpload*`,
71
+ `TemplateVersionRecipe`, `TemplateLanguages`, `TemplatePackage*`, `WorkspaceStartup`, `WorkspaceInputs`.
72
+ - Error reasons: `invalid_recipe`, `base_not_layered`, `language_unavailable`, `language_conflict`, `upload_required`,
73
+ `upload_missing`, `upload_digest_mismatch`, `upload_too_large`, `invalid_settings`, `services_unsupported`,
74
+ `input_required`, `input_unknown`, `input_invalid`, `egress_widening`, `package_index_unavailable` and the operation
75
+ error `startup_failed` (`KnownErrorReason`).
76
+
77
+ ### Tool-call capture
78
+
79
+ `workspace.captureToolCalls(options)` saves the tool calls of your own agent harness (web search, SQL, HTTP APIs, MCP
80
+ servers) into the workspace: each call's input and full output become files under `/home/user/tool-calls/<run>/`
81
+ with a JSON-lines index, so the agent can compute on them with code and snapshots and forks keep them. The format is
82
+ docs/decisions/0006-tool-call-capture.md (shared with the Python SDK 0.3.0).
83
+
84
+ - Invisible: a wrapped tool returns the same value, the same promise object and the same thrown error; a synchronous
85
+ tool stays synchronous; capture never throws into the harness (failures go to `onError` and `capture.stats`).
86
+ Serialization runs in `setImmediate`, writes in the background (4 in parallel, index lines batched every 50 ms).
87
+ - Explicit capture: `capture.record(call)`, `capture.run(toolCall, fn)` (Anthropic `tool_use`, Responses
88
+ `function_call`, Chat `tool_calls`, `{ name, input }`), `capture.wrap(name, fn)` (async generators pass through and
89
+ are stored as `.jsonl`), `capture.tools(toolsObjOrArray)`, and `captureTool(name, fn)` with `capture.activate(fn)`
90
+ for tools defined at import time.
91
+ - Adapters (structural types, no dependency): `capture.aiSdk.tools()` / `.callbacks()` (Vercel AI SDK 7, streaming
92
+ tools included), `capture.mastra.hooks()` / `.tools()`, `capture.anthropic.tools()` (tool runner),
93
+ `capture.openaiAgents.tools()` / `.execute()` / `.attach(runner | agent)`, `capture.claude.hooks()` (Claude Agent
94
+ SDK, with a PreToolUse barrier for `mcp__shardflux__` tools), `capture.langchain.handler()`, `capture.mcp.instrument()`.
95
+ - Selection: explicit capture always records; hook-level adapters apply `include` / `exclude`. Call ids are
96
+ deduplicated (last 10 000), so a wrapper and a hook can be combined.
97
+ - Output encoding: JSON, text, HTML, bytes by magic number (PNG, JPEG, GIF, WebP, PDF, ZIP, gzip, Parquet), MCP results
98
+ and content blocks as a directory of parts; limits `maxOutputBytes` (32 MiB, text cut as `.part`),
99
+ `maxInlineInputBytes` (64 KiB), `maxPendingBytes` (128 MiB of files, index lines and inline inputs, then
100
+ `dropped: "queue_full"` with a small line), `maxPendingCalls` (10 000; past it a call is not recorded and `onError`
101
+ says so); `transform` for redaction (fails closed); streams are never read; out-of-range timestamps are recorded as
102
+ `started_at: null`. A wrapped async tool's promise is marked handled (no `unhandledRejection` if nobody awaits it).
103
+ - Read-your-writes: calls through the same client (exec and files through `workspace.cell()`, `workspaceTools`,
104
+ snapshot, fork, suspend, saveAsTemplate, close; also `cloud.workspaces.*(id)`) wait for capture writes recorded
105
+ before them, bounded by `settleTimeoutMs` (30 s). Lifecycle timing shows the wait as a new `capture_flush` phase.
106
+ `delete` and `reset` drop pending writes. Writes wake a suspended workspace (`wake: null` opts out).
107
+ - `capture.flush()` / `capture.close()` never reject (`{ complete, written, failed, dropped }`); `capture.promptHint()`
108
+ returns a paragraph for your system prompt.
109
+ - `executeToolCall()` passes the call's `id` / `call_id` to `execute` as `toolCallId`, and `WorkspaceTool.execute`'s
110
+ options accept it.
111
+ - **Breaking (types only):** `LifecyclePhase` gains `capture_flush`; an exhaustive `switch` over phases without a
112
+ `default` case needs the new member.
113
+
114
+ ## 0.6.2 (2026-09-28)
7
115
 
8
116
  ### Starts that wait for capacity end
9
117
 
package/README.md CHANGED
@@ -4,16 +4,18 @@ TypeScript SDK for [Shardflux](https://shardflux.dev): cloud computers for AI ag
4
4
 
5
5
  Open a persistent workspace by key, run commands and move files in it, suspend it when idle, resume
6
6
  it later with its disk and memory intact, and fork it. Hand your agent framework-neutral workspace
7
- tools (exec, files, processes, PTY, git, browser) that plug into any model provider.
7
+ tools (exec, files, processes, PTY, git, browser) that plug into any model provider, and save your harness's own
8
+ tool calls into the workspace ([tool-call capture](#tool-call-capture-070)).
8
9
 
9
10
  > **Early access.** Shardflux is in early access. The API is versioned (`/v1`), but this SDK is
10
11
  > below 1.0: a minor release may contain breaking changes (see [Compatibility](#compatibility)).
11
12
 
12
- > **Versions.** This README describes 0.6.2. Anything marked **(0.6.0+)** is not in 0.5.0;
13
- > [CHANGELOG.md](./CHANGELOG.md) lists what each version added. Check yours with
13
+ > **Versions.** This README describes 0.8.0. Anything marked **(0.8.0+)** is not in 0.7.x, **(0.7.0+)** not in 0.6.x
14
+ > and **(0.6.0+)** not in 0.5.0; [CHANGELOG.md](./CHANGELOG.md) lists what each version added. Check yours with
14
15
  > `npm ls @shardflux/sdk` or the exported `SDK_VERSION`.
15
16
 
16
- - ESM only, no runtime dependencies, Node.js 24 or later.
17
+ - ESM only, no runtime dependencies, Node.js 24 or later. Reading a YAML template file uses the optional peer
18
+ dependency `yaml` (`npm install yaml`); JSON template files need nothing.
17
19
  - Typed from the published OpenAPI documents.
18
20
  - Retries, idempotency keys, operation polling and tool-token refresh are handled for you.
19
21
 
@@ -300,6 +302,80 @@ await draft.discard(); // deletes the draft and en
300
302
  `states()`, `statesAll()` and `testInstances({ includeEnded })` list the draft's states and test instances. Only
301
303
  owners, admins and API keys with a tool permission may change drafts; others get 403 `template_dev_mode_role`.
302
304
 
305
+ ## Build a template from template.yaml (0.7.0+)
306
+
307
+ A template is a recipe v2: a base, what the build adds (languages, apt/pip/npm packages, files, named build steps) and
308
+ the settings a workspace gets when it opens (environment, open-time inputs, start commands, services, defaults).
309
+ `template.yaml` is that document in YAML. A file entry may name a local `from` path (relative to the file): a folder is
310
+ copied as a tar, a file as it is.
311
+
312
+ ```yaml
313
+ # acme/template.yaml
314
+ base: ubuntu-24.04@1
315
+ build:
316
+ languages: [{ id: python }, { id: node, version: "22" }]
317
+ packages:
318
+ apt: [jq]
319
+ pip: { packages: [pandas==2.3.2], requirements: [/home/user/app/requirements.txt] }
320
+ files:
321
+ - { from: ./app, to: /home/user/app, owner: user } # a folder: uploaded as a tar
322
+ - { from: ./config/settings.toml, to: /home/user/.config/acme/settings.toml, owner: user, mode: "0600" }
323
+ steps:
324
+ - { name: install, run: npm ci, user: user, cwd: /home/user/app }
325
+ settings:
326
+ env: { APP_ENV: development }
327
+ inputs:
328
+ PROJECT_NAME: { kind: text, required: true }
329
+ OPENAI_API_KEY: { kind: secret } # a stored secret of that name, bound at open
330
+ start: [{ name: seed, when: create, run: python seed.py, user: user, cwd: /home/user/app }]
331
+ services:
332
+ web: { run: npm start, user: user, cwd: /home/user/app, ready: { port: 3000 } }
333
+ ```
334
+
335
+ ```ts
336
+ const { build, uploads } = await cloud.templates.buildFromFile('acme/template.yaml', {
337
+ templateSlug: 'acme-dev',
338
+ autoPublish: false, // register it unpublished, test it, publish it later
339
+ wait: true, // until registered (or failed); default: return the queued build
340
+ onProgress: (e) => console.log(e.type, e.type === 'build' ? e.build.state : e.from),
341
+ });
342
+ console.log(build.state, build.template_version, build.provenance.recipe_sha256, uploads);
343
+ ```
344
+
345
+ `buildFromFile` is Node only (it reads the disk; browsers never load that code). Folders are packed as a
346
+ reproducible tar (sorted, mtime 0, no owner names; symlinks must stay inside the folder), the same bytes as the Python
347
+ SDK packs, so the same inputs give the same `recipe_sha256`. Uploads the organization already has are not sent
348
+ again. It throws `TemplateFileError` before any request for a file it cannot read or pack, and `TemplateUploadError`
349
+ when the storage refuses the bytes. A recipe already in memory: `templates.buildFromRecipe(doc, { templateSlug,
350
+ baseDir })`. Pass `root` to refuse local paths outside a directory, and `parseYaml` to use another YAML parser.
351
+
352
+ The pieces on their own:
353
+
354
+ ```ts
355
+ const up = await cloud.templates.uploads.put(bytes, { kind: 'file' }); // Uint8Array, Blob or a stream
356
+ const dir = await cloud.templates.uploads.putPath('./app'); // Node: a folder as the tar
357
+ await cloud.templates.builds.create(orgId, { templateSlug: 'acme-dev', recipe: { schema: 'shardflux.template-recipe.v2',
358
+ base: 'ubuntu-24.04@1', build: { files: [{ upload: up.ref, kind: 'file', to: '/etc/acme.conf' }] }, settings: {} } });
359
+ const exported = await cloud.templates.versions.recipe('acme-dev', 3); // the recipe to build again, and settings
360
+ const langs = await cloud.templates.languages('ubuntu-24.04@1'); // what build.languages offers on a base
361
+ const pkgs = await cloud.templates.packages.search('apt', 'ffmpeg', { base: 'ubuntu-24.04@1' });
362
+ ```
363
+
364
+ Test a version before publishing it, with its inputs, then open workspaces with inputs:
365
+
366
+ ```ts
367
+ const test = await cloud.templates.versionTestInstances.create('acme-dev', 4, { inputs: { PROJECT_NAME: 'demo' } });
368
+ console.log(test.startup); // start commands and services: pending | running | ready | failed
369
+ await test.close();
370
+ const ws = await cloud.workspaces.open({ key: 'customer-42/main', template: 'acme-dev', inputs: { PROJECT_NAME: 'acme' } });
371
+ console.log(await ws.inputs(), ws.startup);
372
+ ```
373
+
374
+ A failed start command or service fails the open (`OperationFailedError`, code `startup_failed`, retryable); the
375
+ workspace keeps running so you can inspect it, and `workspace.startup` names the step, its exit code and output tail.
376
+ The next open runs the failed step again. Versions report their `settings`, platform templates their `category`
377
+ (`os` or `stack`), builds the `denied_hosts` their build network refused (add them to `build.network.extra_hosts`).
378
+
303
379
  ## Agent tools
304
380
 
305
381
  `workspaceTools(workspace)` returns tools with a name, a description, a JSON Schema for the
@@ -307,18 +383,135 @@ parameters and an `execute` function. Export them for your model provider and di
307
383
  calls:
308
384
 
309
385
  ```ts
386
+ import Anthropic from '@anthropic-ai/sdk';
310
387
  import { executeToolCall, toAnthropicTools, toOpenAITools, workspaceTools } from '@shardflux/sdk';
311
388
 
312
389
  const tools = workspaceTools(workspace);
313
- const anthropicTools = toAnthropicTools(tools); // or toOpenAITools(tools)
390
+ const anthropicTools: Anthropic.Tool[] = toAnthropicTools(tools);
391
+ const chatTools = toOpenAITools(tools); // OpenAI Chat Completions
392
+ const responsesTools = toOpenAITools(tools, { api: 'responses' }); // OpenAI Responses API
314
393
 
315
- // For each tool call the model makes:
316
- const output = await executeToolCall(tools, { name: call.name, input: call.input });
394
+ // For each tool call the model makes (an Anthropic tool_use block, or an OpenAI function_call item, as it is):
395
+ const output = await executeToolCall(tools, block);
317
396
  ```
318
397
 
398
+ **(0.8.0+)** The exports type-check as the provider SDKs' own types under `strict`, with no casts:
399
+ `Anthropic.Tool[]`, `OpenAI.Chat.ChatCompletionTool[]` and, with `{ api: 'responses' }`,
400
+ `OpenAI.Responses.FunctionTool[]`. `executeToolCall` takes a `tool_use` block's `unknown` input as it is and validates
401
+ it.
402
+
319
403
  Your agent loop and model calls stay in your application; the workspace is the computer the tools
320
404
  act on.
321
405
 
406
+ ## Tool-call capture (0.7.0+)
407
+
408
+ Your harness's tools (web search, SQL, HTTP APIs, MCP servers) run in your application, so their results reach the
409
+ model but not the workspace. `workspace.captureToolCalls()` saves every call's input and full output as files in the
410
+ workspace, where the agent can process them with `jq` or Python, and where snapshots and forks keep them.
411
+
412
+ ```ts
413
+ const capture = workspace.captureToolCalls();
414
+
415
+ // A hand-rolled loop (Anthropic tool_use, Responses function_call, Chat tool_calls item, or { name, input }):
416
+ for (const block of message.content) {
417
+ if (block.type !== 'tool_use') continue;
418
+ const output = await capture.run(block, () => myTools[block.name](block.input));
419
+ results.push({ type: 'tool_result', tool_use_id: block.id, content: JSON.stringify(output) });
420
+ }
421
+ ```
422
+
423
+ Capture is invisible to the harness. A wrapped tool returns the same value, the same promise object and the same
424
+ thrown error; a synchronous tool stays synchronous. Nothing capture does throws into your code: write failures,
425
+ drops and a throwing `transform` go to `onError` and `capture.stats`. Serialization runs after the call returns and
426
+ the writes run in the background.
427
+
428
+ One line per framework:
429
+
430
+ | Harness | Integration |
431
+ |---|---|
432
+ | Hand-rolled loop | `capture.run(call, () => ...)`, `capture.wrap('name', fn)`, or `capture.record({ tool, input, output, callId })` |
433
+ | Shardflux tools | `executeToolCall(capture.tools(workspaceTools(workspace)), call)` (the call's `id` is recorded) |
434
+ | Vercel AI SDK 7 | `generateText({ tools: capture.aiSdk.tools(tools) })`, or `...capture.aiSdk.callbacks()` (`onToolExecutionEnd`; `{ legacy: true }` for `experimental_onToolCallFinish`) |
435
+ | Mastra | `new Agent({ tools: capture.mastra.tools({ weather }), hooks: capture.mastra.hooks() })` |
436
+ | Anthropic tool runner | `client.beta.messages.toolRunner({ tools: capture.anthropic.tools([weatherTool]) })` |
437
+ | OpenAI Agents JS | `const detach = capture.openaiAgents.attach(runner)`, or `tools: capture.openaiAgents.tools([...])` |
438
+ | Claude Agent SDK | `query({ prompt, options: { hooks: capture.claude.hooks(myHooks) } })` |
439
+ | LangChain.js / LangGraph.js | `agent.invoke(input, { callbacks: [capture.langchain.handler()] })` |
440
+ | MCP client | `const release = capture.mcp.instrument(client, { server: 'github' })` |
441
+
442
+ ```ts
443
+ // Vercel AI SDK: the wrapper keeps toolCallId, abortSignal and every other tool property. A streaming tool
444
+ // (async generator execute) passes through unchanged; its final value is recorded.
445
+ const result = await generateText({ model, tools: capture.aiSdk.tools({ weather, search }), prompt });
446
+
447
+ // Claude Agent SDK: your own hooks are kept. PostToolUse / PostToolUseFailure record every tool, and a
448
+ // PreToolUse hook matching ^mcp__shardflux__ waits for pending writes, so the Shardflux MCP server sees them.
449
+ for await (const m of query({ prompt, options: { hooks: capture.claude.hooks(), mcpServers } })) handle(m);
450
+
451
+ // OpenAI Agents JS: agent_tool_end reports a tool's error as the string the model got (status ok). For the exact
452
+ // error, wrap execute: tool({ name: 'get_weather', parameters, execute: capture.openaiAgents.execute('get_weather', fn) }).
453
+ const detach = capture.openaiAgents.attach(runner);
454
+
455
+ // Tools defined at import time in a multi-tenant server: late-bound to the capture active for the request.
456
+ export const lookup = captureTool('lookup', async (id: string) => db.find(id));
457
+ await capture.activate(() => handleRequest(req)); // lookup() records into this capture; elsewhere it passes through
458
+ ```
459
+
460
+ **Selection.** Explicit capture (`record`, `run`, `wrap`, `tools`, `captureTool`) always records. Hook-level
461
+ adapters (`callbacks()`, `hooks()`, `attach()`, `handler()`, `instrument()`) see every tool, Shardflux's own
462
+ included, filtered by `include` / `exclude` (names, a RegExp, or `(tool, source) => boolean`). A capture remembers the
463
+ last 10 000 call ids and records each once, so a wrapper and a hook can be combined.
464
+
465
+ **In the workspace** (`dir` defaults to `/home/user/tool-calls`; point it at a subdirectory, e.g.
466
+ `/home/user/tool-calls/conv-123`, to group runs):
467
+
468
+ ```
469
+ <dir>/README.md layout and jq recipes, for the agent
470
+ <dir>/<run>/index.jsonl one JSON line per call: seq, call_id, tool, status, input, output_path, ...
471
+ <dir>/<run>/000007-web_search.json the output (.json, .txt, .html, .png, .pdf, ... from its content)
472
+ <dir>/<run>/000010-github.search/ an MCP result or content blocks: part-1.txt, part-2.png, result.json
473
+ <dir>/<run>/000011-sql.input.json an input over 64 KiB
474
+ ```
475
+
476
+ `<run>` is `capture.runId` (start time plus a random suffix; `capture.runDir` is the full path). Read the index with
477
+ `jq -cR 'fromjson? // empty' <dir>/*/index.jsonl`, which skips a line torn by a failed append. `capture.promptHint()`
478
+ returns a paragraph that tells the agent where its tool calls are; add it to your system prompt if you want (it is
479
+ never injected).
480
+
481
+ **Read-your-writes.** Calls through the same client first wait for capture writes recorded before them (bounded by
482
+ `settleTimeoutMs`, 30 s; they never fail because of capture): `exec` and files calls through `workspace.cell()`,
483
+ `workspaceTools`, `snapshot`, `fork`, `suspend`, `saveAsTemplate` and `close` (on the workspace handle and on
484
+ `cloud.workspaces.*(id)`). A lifecycle call's timing shows the wait as a `capture_flush` phase. `delete` and `reset`
485
+ drop pending writes. A write to a suspended workspace wakes it (`wake: null` opts out).
486
+
487
+ **Serverless.** Writes finish in the background, so let them finish before the function is frozen:
488
+ `waitUntil(capture.flush())` on Vercel (or Next.js `after(() => capture.flush())`), `await capture.flush()` before
489
+ returning on AWS Lambda. `flush()` and `close()` never reject; they return `{ complete, written, failed, dropped }`.
490
+
491
+ **Redaction.** Nothing is redacted by default (the input came from the model). `transform(event)` gets a copy of each
492
+ call and returns it (changed or not) or `null` to drop it; if it throws, the call is dropped, never written unredacted.
493
+
494
+ ```ts
495
+ const capture = workspace.captureToolCalls({
496
+ exclude: ['exec'],
497
+ transform: (e) => (e.tool === 'crm_lookup' ? { ...e, output: redact(e.output) } : e),
498
+ onError: (err) => log.warn(err.kind, err.message),
499
+ });
500
+ ```
501
+
502
+ **Limits.** An output over `maxOutputBytes` (32 MiB) is cut when it is text or JSON (`.part`, `truncated: true`) and
503
+ not stored when it is binary (`dropped: "too_large"`). What pending calls hold in memory (files, index lines, inline
504
+ inputs) is bounded by `maxPendingBytes` (128 MiB): over it the output and a large input are dropped and the call keeps
505
+ a small index line with `dropped: "queue_full"`; the call never waits. At most `maxPendingCalls` (10 000) calls are
506
+ pending: past that, while the workspace takes no writes, a call is not recorded at all (`record()` returns null,
507
+ `onError` gets `queue_full`). A write is retried for `retryWindowMs` (120 s) and then dropped (`write_failed`).
508
+ Observing an async tool marks its promise handled: if your code never awaits a wrapped tool's promise and it rejects,
509
+ Node reports no `unhandledRejection` for it (the error is still in the index; Node has no way to observe a rejection
510
+ without handling it). About one write per call: more than roughly 100 calls per second per
511
+ workspace reaches the pending limit. `Response`, `ReadableStream`, Node streams and `Blob` results are never read
512
+ (`meta.note: "stream_not_captured"`). On template v1 the files API writes as root, so the agent can read the captured
513
+ files but not change them. The format is specified in `docs/decisions/0006-tool-call-capture.md`.
514
+
322
515
  ## Secrets
323
516
 
324
517
  Store credentials once and give them to a workspace's processes as environment variables. Values
@@ -366,7 +559,7 @@ Treat unknown error codes and reasons as generic errors: show `message`, and use
366
559
 
367
560
  ## More of the API
368
561
 
369
- The `Shardflux` object also has `templates` (including custom template builds), `volumes`
562
+ The `Shardflux` object also has `templates` (including custom template builds and the template editor), `volumes`
370
563
  (shared persistent storage attached to workspaces), `secrets` (see above), `egress` (outbound allowlists),
371
564
  `usage`, `billing`, `me()`, `entitlements(orgId)` and `request(method, path)` for any `/v1`
372
565
  route. The package exports the OpenAPI-generated types as well (`paths`, `components`,
@@ -376,7 +569,7 @@ route. The package exports the OpenAPI-generated types as well (`paths`, `compon
376
569
 
377
570
  - The SDK follows the API's `/v1` contract. New fields, enum values and error codes can appear in
378
571
  any release; ignore unknown fields.
379
- - While below 1.0, a breaking change bumps the minor version (0.5 to 0.6).
572
+ - While below 1.0, a breaking change bumps the minor version (0.7 to 0.8).
380
573
  - `SDK_VERSION` is exported; requests send `User-Agent: shardflux-sdk-ts/<version>`.
381
574
  - Examples in this README, in `examples/` and on shardflux.dev name the version they need. The examples on the
382
575
  website and in the console are checked against the version published on npm before they ship.
@@ -0,0 +1,178 @@
1
+ /**
2
+ * Framework adapters for tool-call capture. Structural types only: no framework is a dependency of the SDK, and each
3
+ * adapter's output is checked against the real framework types in test/capture-types.ts.
4
+ *
5
+ * Explicit wrappers (`tools()` of each framework, `capture.tools()`) always record. Hook-level adapters (AI SDK
6
+ * callbacks, Mastra hooks, OpenAI Agents `attach`, Claude Agent SDK hooks, LangChain callbacks, MCP instrumentation)
7
+ * see every tool and apply `include` / `exclude`. A capture remembers the last 10 000 call ids, so a wrapper and a hook
8
+ * on the same call record it once.
9
+ */
10
+ import type { CallRef, CaptureCall, CaptureFlushResult, CaptureSource, CaptureStatus } from './capture.js';
11
+ /** How the capture observes one call (internal). */
12
+ export interface ObserveSpec {
13
+ tool: string;
14
+ source: CaptureSource;
15
+ input: unknown;
16
+ callId: string | null;
17
+ meta?: Record<string, unknown>;
18
+ /** An async iterable result: record its last item (AI SDK streaming tools) instead of every item. */
19
+ final?: boolean;
20
+ /** Filtered by include/exclude (hook-level adapters). */
21
+ hookLevel?: boolean;
22
+ /** Maps a successful result to what is recorded (LangChain ToolMessage, MCP isError, Mastra ValidationError). */
23
+ post?: (output: unknown) => {
24
+ output: unknown;
25
+ error?: unknown;
26
+ status?: CaptureStatus;
27
+ };
28
+ }
29
+ /** What the adapters need from a capture (internal). */
30
+ export interface CaptureCore {
31
+ observe<T>(spec: ObserveSpec, fn: (...args: never[]) => T, thisArg: unknown, args: unknown[]): T;
32
+ hook(call: CaptureCall): CallRef | null;
33
+ settle(): Promise<unknown> | undefined;
34
+ flush(timeoutMs?: number): Promise<CaptureFlushResult>;
35
+ accepts(tool: string, source: CaptureSource): boolean;
36
+ }
37
+ /** A copy of `target` with its prototype and own properties, and `key` replaced by `value` (the original is untouched). */
38
+ export declare function copyWith<T extends object>(target: T, key: string, value: unknown): T;
39
+ /** A LangChain ToolMessage: recorded as its content (plus artifact), status error when the message says so. */
40
+ export declare function isToolMessage(v: unknown): v is {
41
+ content: unknown;
42
+ artifact?: unknown;
43
+ status?: string;
44
+ tool_call_id: string;
45
+ name?: string;
46
+ };
47
+ /** The part of an AI SDK 7 `onToolExecutionEnd` event capture reads (also the v6 `experimental_onToolCallFinish` shape). */
48
+ export interface AiSdkToolEndEvent {
49
+ toolCall?: {
50
+ toolCallId?: string;
51
+ toolName?: string;
52
+ input?: unknown;
53
+ } | undefined;
54
+ toolExecutionMs?: number | undefined;
55
+ toolOutput?: {
56
+ type?: string;
57
+ output?: unknown;
58
+ error?: unknown;
59
+ } | undefined;
60
+ success?: boolean | undefined;
61
+ output?: unknown;
62
+ error?: unknown;
63
+ durationMs?: number | undefined;
64
+ }
65
+ export interface AiSdkAdapter {
66
+ /** `generateText({ tools: capture.aiSdk.tools(tools) })`: wraps each `execute` (keeps toolCallId, abortSignal and every other property). */
67
+ tools<T>(tools: T): T;
68
+ /** `generateText({ ...capture.aiSdk.callbacks() })`: `onToolExecutionEnd` (hook-level). */
69
+ callbacks(opts?: {
70
+ legacy?: false;
71
+ }): {
72
+ onToolExecutionEnd: (event: AiSdkToolEndEvent) => void;
73
+ };
74
+ /** `{ legacy: true }`: the deprecated alias `experimental_onToolCallFinish`. */
75
+ callbacks(opts: {
76
+ legacy: true;
77
+ }): {
78
+ experimental_onToolCallFinish: (event: AiSdkToolEndEvent) => void;
79
+ };
80
+ }
81
+ /** Mastra `afterToolCall` context (ToolAfterHookContext). */
82
+ export interface MastraAfterToolCallContext {
83
+ toolName?: string;
84
+ input?: unknown;
85
+ context?: unknown;
86
+ metadata?: unknown;
87
+ output?: unknown;
88
+ error?: unknown;
89
+ }
90
+ export interface MastraHooksLike {
91
+ beforeToolCall?: (...args: any[]) => any;
92
+ afterToolCall?: (ctx: any) => void | Promise<void>;
93
+ }
94
+ export interface MastraAdapter {
95
+ /** `new Agent({ hooks: capture.mastra.hooks(myHooks) })`: records in `afterToolCall`, then calls yours. */
96
+ hooks<H extends MastraHooksLike = Record<never, never>>(existing?: H): H & {
97
+ afterToolCall: (ctx: MastraAfterToolCallContext) => Promise<void>;
98
+ };
99
+ /** `new Agent({ tools: capture.mastra.tools({ weather }) })`: wraps each `execute(inputData, context)`. */
100
+ tools<T>(tools: T): T;
101
+ }
102
+ export interface AnthropicAdapter {
103
+ /** `client.beta.messages.toolRunner({ tools: capture.anthropic.tools([...]) })`: wraps `run(input, context)`. */
104
+ tools<T>(tools: T): T;
105
+ }
106
+ export interface OpenAIAgentsAdapter {
107
+ /** `new Agent({ tools: capture.openaiAgents.tools([...]) })`: wraps each FunctionTool's `invoke`. */
108
+ tools<T>(tools: T): T;
109
+ /**
110
+ * `tool({ execute: capture.openaiAgents.execute('name', fn) })`: records around your execute, with the exact error
111
+ * (a FunctionTool's own error function turns a thrown error into a string result before `invoke` returns).
112
+ */
113
+ execute<F extends (...args: any[]) => any>(name: string, fn: F): F;
114
+ /** Subscribes `agent_tool_end` on a Runner or an Agent (hook-level); returns the function that unsubscribes. */
115
+ attach(target: {
116
+ on(event: 'agent_tool_end', listener: (...args: any[]) => void): unknown;
117
+ off(event: 'agent_tool_end', listener: (...args: any[]) => void): unknown;
118
+ }): () => void;
119
+ }
120
+ export type ClaudeHookCallback = (input: any, toolUseID: string | undefined, options: {
121
+ signal: AbortSignal;
122
+ }) => Promise<{
123
+ continue?: boolean;
124
+ }>;
125
+ export interface ClaudeHookMatcher {
126
+ matcher?: string;
127
+ hooks: ClaudeHookCallback[];
128
+ timeout?: number;
129
+ }
130
+ export interface ClaudeHookMatcherLike {
131
+ matcher?: string;
132
+ hooks: Array<(input: any, toolUseID: string | undefined, options: {
133
+ signal: AbortSignal;
134
+ }) => Promise<unknown>>;
135
+ timeout?: number;
136
+ }
137
+ export type ClaudeHooksLike = Partial<Record<string, ClaudeHookMatcherLike[]>>;
138
+ export type ClaudeCaptureHooks = {
139
+ PreToolUse: ClaudeHookMatcher[];
140
+ PostToolUse: ClaudeHookMatcher[];
141
+ PostToolUseFailure: ClaudeHookMatcher[];
142
+ };
143
+ export interface ClaudeAdapter {
144
+ /**
145
+ * `query({ prompt, options: { hooks: capture.claude.hooks(myHooks) } })`: PostToolUse and PostToolUseFailure record
146
+ * (hook-level); a PreToolUse hook matching `^mcp__shardflux__` waits for pending writes (bounded) so the Shardflux MCP
147
+ * server sees them. Your hooks are kept and run as before.
148
+ */
149
+ hooks<H extends ClaudeHooksLike = Record<never, never>>(existing?: H): Omit<H, keyof ClaudeCaptureHooks> & ClaudeCaptureHooks;
150
+ }
151
+ export interface LangChainToolHandler {
152
+ handleToolStart(tool: unknown, input: string, runId: string, parentRunId?: string, tags?: string[], metadata?: Record<string, unknown>, runName?: string, toolCallId?: string): void;
153
+ handleToolEnd(output: unknown, runId: string, parentRunId?: string, tags?: string[]): void;
154
+ handleToolError(err: unknown, runId: string, parentRunId?: string, tags?: string[]): void;
155
+ }
156
+ export interface LangChainAdapter {
157
+ /** `{ callbacks: [capture.langchain.handler()] }` (a plain methods object: LangChain wraps it with fromMethods). */
158
+ handler(): LangChainToolHandler;
159
+ }
160
+ export interface McpAdapter {
161
+ /** Patches this client instance's `callTool` (hook-level). Returns the function that restores it. */
162
+ instrument(client: {
163
+ callTool: (...args: any[]) => any;
164
+ }, opts?: {
165
+ server?: string;
166
+ }): () => void;
167
+ }
168
+ export interface CaptureAdapters {
169
+ wrapAny<T>(tools: T): T;
170
+ aiSdk: AiSdkAdapter;
171
+ mastra: MastraAdapter;
172
+ anthropic: AnthropicAdapter;
173
+ openaiAgents: OpenAIAgentsAdapter;
174
+ claude: ClaudeAdapter;
175
+ langchain: LangChainAdapter;
176
+ mcp: McpAdapter;
177
+ }
178
+ export declare function createAdapters(core: CaptureCore): CaptureAdapters;