@gea-ai/agent-sdk 0.1.260910-alpha.0 → 0.1.260910-alpha.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (51) hide show
  1. package/README.md +3 -1253
  2. package/dist/agent-channel-author-runtime.d.ts +79 -0
  3. package/dist/agent-channel-author-runtime.d.ts.map +1 -0
  4. package/dist/agent-channel-author-runtime.js +511 -0
  5. package/dist/agent-channel-author-runtime.js.map +1 -0
  6. package/dist/agent-channel-receiver.d.ts +46 -0
  7. package/dist/agent-channel-receiver.d.ts.map +1 -0
  8. package/dist/agent-channel-receiver.js +356 -0
  9. package/dist/agent-channel-receiver.js.map +1 -0
  10. package/dist/agent-channel-runtime.d.ts +6 -0
  11. package/dist/agent-channel-runtime.d.ts.map +1 -0
  12. package/dist/agent-channel-runtime.js +623 -0
  13. package/dist/agent-channel-runtime.js.map +1 -0
  14. package/dist/agent-channel-worker.d.ts +22 -0
  15. package/dist/agent-channel-worker.d.ts.map +1 -0
  16. package/dist/agent-channel-worker.js +335 -0
  17. package/dist/agent-channel-worker.js.map +1 -0
  18. package/dist/agent-session.d.ts.map +1 -1
  19. package/dist/agent-session.js +177 -4
  20. package/dist/agent-session.js.map +1 -1
  21. package/dist/agent-worker.d.ts +27 -1
  22. package/dist/agent-worker.d.ts.map +1 -1
  23. package/dist/agent-worker.js +867 -27
  24. package/dist/agent-worker.js.map +1 -1
  25. package/dist/channels/feishu-codec.d.ts +18 -0
  26. package/dist/channels/feishu-codec.d.ts.map +1 -0
  27. package/dist/channels/feishu-codec.js +109 -0
  28. package/dist/channels/feishu-codec.js.map +1 -0
  29. package/dist/channels/feishu-inbound.d.ts +35 -0
  30. package/dist/channels/feishu-inbound.d.ts.map +1 -0
  31. package/dist/channels/feishu-inbound.js +136 -0
  32. package/dist/channels/feishu-inbound.js.map +1 -0
  33. package/dist/channels/feishu-websocket.d.ts +5 -0
  34. package/dist/channels/feishu-websocket.d.ts.map +1 -0
  35. package/dist/channels/feishu-websocket.js +424 -0
  36. package/dist/channels/feishu-websocket.js.map +1 -0
  37. package/dist/channels/feishu.d.ts +32 -0
  38. package/dist/channels/feishu.d.ts.map +1 -0
  39. package/dist/channels/feishu.js +233 -0
  40. package/dist/channels/feishu.js.map +1 -0
  41. package/dist/channels/index.d.ts +180 -0
  42. package/dist/channels/index.d.ts.map +1 -0
  43. package/dist/channels/index.js +93 -0
  44. package/dist/channels/index.js.map +1 -0
  45. package/dist/index.d.ts +18 -0
  46. package/dist/index.d.ts.map +1 -1
  47. package/dist/index.js +38 -8
  48. package/dist/index.js.map +1 -1
  49. package/package.json +16 -3
  50. package/src/channels/README.md +298 -0
  51. package/src/channels/feishu.ts +295 -0
package/README.md CHANGED
@@ -1,1259 +1,9 @@
1
1
  # `@gea-ai/agent-sdk`
2
2
 
3
- `@gea-ai/agent-sdk` is the code-first authoring SDK for GEA Agent
4
- applications. It defines main Agents, package-local executable Tools, Skill
5
- references, Agent-owned Durable Objects, and secret-free Connector definitions
6
- and requirements.
7
-
8
- The SDK retains executable handlers for Worker runtime use.
9
- `createAgentPackageSnapshot()` produces deterministic JSON-compatible metadata
10
- for validation, Studio display, and exact-version execution. Snapshots never
11
- contain handler functions, credentials, Connector Connections, tenant identity,
12
- provider configuration, or environment values.
13
-
14
- ## Agent Authoring Shapes
15
-
16
- A single-main application starts with root `agent.ts` and `AGENTS.md`:
17
-
18
- ```ts
19
- import { defineAgent } from "@gea-ai/agent-sdk";
20
-
21
- export default defineAgent({
22
- model: "gea-model-1",
23
- name: "research-agent",
24
- slug: "research-agent",
25
- });
26
- ```
27
-
28
- An Agent may instead declare capability thresholds for the SDK's first-party
29
- Creative Reasoning router:
30
-
31
- ```ts
32
- export default defineAgent({
33
- model: "auto",
34
- modelRequirements: {
35
- agentic: 0.8,
36
- copywriting: 0.6,
37
- multimodal: 0.4,
38
- speed: 0.7,
39
- },
40
- name: "research-agent",
41
- slug: "research-agent",
42
- });
43
- ```
44
-
45
- Each requirement is an optional minimum score from `0` to `1`; omitted
46
- dimensions do not constrain selection. `modelRequirements` is required for
47
- `model: "auto"` and rejected for a concrete model. The SDK filters its curated
48
- Creative Reasoning capability table and selects the eligible model with the
49
- lowest configured cost rank. It writes only the selected concrete CR gateway
50
- model ID into the package snapshot, so runtime model selection does not change
51
- when a later SDK updates the table.
52
-
53
- This table is maintained in the SDK; it does not query CR model discovery,
54
- verify API-key access, or read the deployment's model catalog or prices.
55
- Capability scores and cost ranks are provisional routing parameters rather
56
- than official platform measurements or actual catalog prices. The following
57
- candidate IDs and upstream mappings were supplied by the platform owner on
58
- 2026-09-05; authenticated access still needs verification for each deployment.
59
- The mapping is developer information and is not a product display label.
60
-
61
- | CR gateway model ID | Upstream mapping | Cost rank | Agentic | Copywriting | Multimodal | Speed |
62
- | ------------------------ | ----------------- | --------- | ------- | ----------- | ---------- | ----- |
63
- | `deepseek-v4-flash` | DeepSeek V4 Flash | 1 | 0.5 | 0.5 | 0 | 1 |
64
- | `crr-q-flash-20260826` | Qwen3.8 Flash | 2 | 0.55 | 0.6 | 0 | 1 |
65
- | `crr-o-mini-20260710` | gpt-5.6-luna | 3 | 0.55 | 0.45 | 0.55 | 1 |
66
- | `deepseek-v4-pro` | DeepSeek V4 Pro | 4 | 0.8 | 0.75 | 0 | 0.65 |
67
- | `crr-q-pro-20260804` | Qwen3.8 Max | 5 | 0.85 | 0.8 | 0 | 0.6 |
68
- | `crr-o-20260710` | gpt-5.6-terra | 6 | 0.8 | 0.7 | 0.75 | 0.7 |
69
- | `creative-reasoning-1.5` | Claude Sonnet 5 | 7 | 0.8 | 0.9 | 0.7 | 0.65 |
70
- | `crr-o-pro-20260710` | gpt-5.6-sol | 8 | 0.95 | 0.82 | 0.85 | 0.45 |
71
- | `crr-a-pro-20260724` | Claude Opus 5 | 9 | 1 | 1 | 0.85 | 0.35 |
72
-
73
- Qwen and DeepSeek currently have a zero multimodal routing score until their
74
- gateway support is verified; any positive multimodal requirement excludes
75
- them. Existing GPT and Claude routing thresholds are retained. With this
76
- candidate set, `multimodal: 0.9` or `{ multimodal: 0.8, speed: 0.9 }` has no
77
- eligible model and fails during Agent definition. Removing an auto candidate
78
- does not invalidate an explicit model declaration or an existing snapshot.
79
-
80
- For hosted execution, register these nine IDs in
81
- `LLM_MODEL_CATALOG_JSON.models`, using `creative-reasoning/<gateway-model-id>`
82
- targets and actual target-keyed prices. Keep the existing `gea-fast` and
83
- `gea-pro` aliases and set `selectableModels: ["gea-fast", "gea-pro"]` so normal
84
- product selectors show only those aliases. Studio developers can inspect the
85
- immutable version's model but cannot override it in Playground.
86
-
87
- `defineAgent({ model: "auto", ... })` provides contextual field completion for
88
- all four requirements. Standalone input annotations require a model type:
89
- `DefineAgentInput<"auto">` requires the capability fields, while
90
- `DefineAgentInput<"crr-a-pro-20260724">` forbids them. There is no implicit
91
- `string` model parameter, because it would also accept `"auto"` without its
92
- requirements. Direct `defineAgent(...)` calls infer the model type.
93
-
94
- A concrete first-party model may use any unqualified public ID, such as
95
- `crr-a-pro-20260724`, `creative-reasoning-1.5`, or
96
- `creative-reasoning-1.5-flash`. An explicit Creative Reasoning provider ID such
97
- as `creative-reasoning/kimi-k3` is also supported. The snapshot preserves the
98
- declared concrete ID.
99
-
100
- Tools, Connectors, Hooks, referenced Skills, and local Skills are discovered
101
- from conventional sibling directories. The required `slug` is the Agent's
102
- stable identity inside this Worker application. It becomes the manifest and
103
- runtime `agentKey`; `name` remains display metadata.
104
-
105
- A root Agent may stay in place when the application adds more main Agents under
106
- `agents/<slug>/agent.ts`. This preserves the original Agent's identity and
107
- root `benchmarks/`. A directory Agent's declared slug must match its directory
108
- name. Applications without a root Agent may use `agents/default/` as the
109
- default source alias; that Agent still declares a real non-reserved slug. If
110
- there is no root or `agents/default/`, the lexically first directory Agent owns
111
- the default route. `default`, `run`, and `sessions` are reserved protocol
112
- segments and cannot be Agent slugs.
113
-
114
- ```text
115
- agents/default/agent.ts
116
- agents/default/AGENTS.md
117
- agents/support/agent.ts
118
- agents/support/AGENTS.md
119
- ```
120
-
121
- Add root `worker.ts` only when the same Worker also serves UI, business APIs,
122
- assets, or application Durable Objects. The CLI generates Agent assembly and
123
- prepends it to the application handler. Developers do not list Tools, Skills,
124
- or Connectors again in `worker.ts`.
125
-
126
- Generated code uses `createAgentApplicationFetch()`, which owns:
127
-
128
- ```text
129
- /gea/agents/run
130
- /gea/agents/<agentKey>/run
131
- /gea/agents[/<agentKey>]/sessions/<sessionId>/v1/<operation>
132
- ```
133
-
134
- Hosted Agent HTTP handlers require a Project API Key by default. This applies
135
- to the Agent routes, while ordinary routes in `worker.ts` remain public.
136
-
137
- ```ts
138
- // Inside defineAgent(...):
139
- http: {
140
- auth: false;
141
- } // Explicit public access to this Agent's HTTP API.
142
- ```
143
-
144
- An application can replace the default with its own Session authenticator:
145
-
146
- ```ts
147
- http: {
148
- auth: async (request, { environment }) => {
149
- const session = await getSession(request, environment); // Your verified session.
150
- return session
151
- ? { principalType: "user", principalId: session.userId }
152
- : null;
153
- },
154
- }
155
- ```
156
-
157
- The callback returns an external user, a `Response`, or `null` (401). It may
158
- also return the result of `projectKeyAuth()` to explicitly compose key auth.
159
- There is no implicit fallback after a custom authenticator rejects. Project
160
- Keys stay on the backend; they never represent the creator's GEA user.
161
- `authenticateProjectKey(request, environment.PROJECT)` also works in ordinary
162
- Worker routes with a declared Project binding and a Project attachment.
163
-
164
- Application pages and browser Sessions require the operator's isolated
165
- `WORKER_APP_BASE_URL`. GEA-domain and path ingress preserve API access but strip
166
- platform cookies and force CSP sandbox plus `nosniff`; Worker pages cannot run
167
- scripts there. Application-domain pages retain their own CSP. Local `gea agent dev` uses its
168
- existing trusted development identity. Raw AgentSession and Connector auth
169
- operations remain private regardless of `http.auth`. Rebuild and republish old
170
- Agent packages with the matching SDK/CLI to enable the new hosted handler.
171
-
172
- The public run request is intentionally small:
173
-
174
- ```json
175
- { "chatId": "optional-continuation-id", "message": "Hello" }
176
- ```
177
-
178
- If `chatId` is absent, the hosted managed API creates a Chat and Run and returns
179
- their IDs in response headers. Local execution creates its local Session identity.
180
- Callers cannot supply model-loop messages, instructions, provider options, or
181
- platform-owned Run/message IDs. Each main Agent's exact `AGENTS.md` text is
182
- compiled into the Worker and is never accepted from the run request.
183
-
184
- `composeFetch()` delegates unclaimed requests to the optional application
185
- Worker. Worker Runtime remains route-agnostic.
186
-
187
- Every main Agent declares one concrete public GEA model ID or the SDK-owned
188
- `auto` selector above. GEA resolves the concrete snapshot model through the
189
- active Model Gateway catalog; source and artifacts contain no provider
190
- credential or target configuration.
191
-
192
- ## Context strategies
193
-
194
- Agents use the existing tool-output offload and rolling summary by default.
195
- Configure that recipe with `defaultContext`, replace it with a custom strategy,
196
- or set `context: false` to disable automatic summaries, file writes and tool-result
197
- replacement. Disabling compression preserves already committed summaries and
198
- continues normal message persistence.
199
-
200
- ```ts
201
- import { defineAgent } from "@gea-ai/agent-sdk";
202
- import { defaultContext } from "@gea-ai/agent-sdk/context";
203
-
204
- export default defineAgent({
205
- name: "assistant",
206
- slug: "assistant",
207
- model: "gea-model-1",
208
- context: defaultContext({
209
- toolOutput: { offloadAtCharacters: 4_000, replaceAtCharacters: 8_000 },
210
- preserveRecentGroups: 5,
211
- summarization: {
212
- triggerAtTotalTokens: 64_000,
213
- // Optional: model, instructions, maxOutputTokens.
214
- },
215
- }),
216
- });
217
- ```
218
-
219
- `toolOutput: false` and `summarization: false` independently disable the two
220
- parts of the default recipe. Thresholds use model-visible characters and the
221
- last provider-reported `usage.totalTokens`; the SDK does not estimate tokens.
222
- Recent groups keep tool calls and their results together. Without Computer
223
- storage, the default recipe skips offload but can still summarize.
224
-
225
- The optional Rust/WASM engine uses the same context strategies and host tools:
226
-
227
- ```ts
228
- import { agentCore } from "@gea-ai/agent-sdk/agent-core";
229
-
230
- export default defineAgent({
231
- name: "assistant",
232
- slug: "assistant",
233
- model: "gea-model-1",
234
- engine: agentCore({
235
- maxToolConcurrency: 4,
236
- maxSteps: 20,
237
- maxRetries: 2,
238
- timeout: {
239
- totalMs: 120_000,
240
- stepMs: 60_000,
241
- chunkMs: 15_000,
242
- toolMs: 30_000,
243
- },
244
- onError: ({ error }) => console.error(error.code, error.details),
245
- guardrail: ({ call }) =>
246
- call.name === "deleteAsset" ? "user-approval" : "not-applicable",
247
- }),
248
- });
249
- ```
250
-
251
- `gea agent` bundles Core's WASM bytes with the Agent. The factory is an authoring
252
- entry for that bundler; installed projects do not need Rust. The default engine
253
- remains AI SDK. Core supports local TS tools, pre-execution guardrail and approval,
254
- context strategies and the existing managed Session/Artifact lifecycle. An approval
255
- pause skips custom `onEnd` until continuation completes. Guardrail is checked again
256
- before dispatch and may deny or require approval; it cannot override a Tool denial.
257
- Core initiates preparation, step-end and run-end callbacks into TS, and waits for
258
- context settlement before returning its outcome. A failed final context commit
259
- cannot report success; cancellation still waits for cleanup.
260
- Declared Tools keep their typed `tool-<name>` UI parts and dynamic Tools keep
261
- `dynamic-tool`, so existing Tool renderers and code Judges can use either engine.
262
- Provider-hosted/custom and client-executed tools are outside the current Core subset.
263
- The Worker Runtime must support WebAssembly JSPI. This optional engine supports
264
- a subset of AI SDK model and output features.
265
- Rust owns error classification, model request retries and timeout decisions. The
266
- TS `onError` callback observes terminal failures once; recoverable Tool errors,
267
- retry attempts and caller cancellation do not invoke it. `maxRetries` defaults to
268
- 2 and `timeout` also accepts a number for the total active execution budget.
269
- Established streams and Tools are never replayed automatically. `maxSteps` stops
270
- after a complete step. Cleanup remains awaited, so a timeout does not forcibly
271
- terminate uncooperative host code or reverse an effect.
272
-
273
- Dynamic step settings are supplied by TS callbacks and applied by Rust:
274
-
275
- ```ts
276
- engine: agentCore({
277
- models: ["gea-model-2"],
278
- prepareStep: ({ stepNumber }) =>
279
- stepNumber === 0
280
- ? { activeTools: ["lookup"], toolChoice: { type: "tool", toolName: "lookup" } }
281
- : { model: "gea-model-2", activeTools: [], generation: { max_output_tokens: 2048 } },
282
- stopWhen: ({ steps }) =>
283
- steps.reduce((sum, step) => sum + (step.usage?.totalTokens ?? 0), 0) >= 10_000,
284
- }),
285
- ```
286
-
287
- Declare additional model IDs in `models` so the CLI can prepare their local routes.
288
- `stepNumber` is zero-based. `prepareStep` receives current messages, available tools
289
- and completed step summaries; it can also set `toolOrder`, `bodyOptions` or `stop`.
290
- Model/tool/generation overrides apply only to that step. `generation` uses Core's
291
- neutral snake_case keys; `bodyOptions` uses the selected protocol's raw option keys.
292
- Omitted/undefined values leave provider defaults intact. `stopWhen` runs after a
293
- completed tool step is committed; a final text response ends naturally. The
294
- example's token threshold is checked between steps and may be exceeded by the
295
- last completed call. Approval preserves selected tools/model and prior step
296
- usage, then evaluates the predicate after continuation without another model call.
297
- Both callbacks receive `signal`. Existing `context` strategies remain responsible
298
- for replacing the durable message projection.
299
-
300
- Final structured output accepts a convertible TS schema:
301
-
302
- ```ts
303
- engine: agentCore({
304
- output: z.object({ answer: z.number().int(), explanation: z.string() }),
305
- }),
306
- ```
307
-
308
- Zod 4 and libraries implementing Standard JSON Schema v1 are supported. TS converts
309
- only the schema; Rust sends the protocol request, parses final JSON and validates
310
- it. Consume the successful result from `message.metadata.output`. Intermediate
311
- text is still a preview. Tool steps and approvals work normally, and output errors
312
- retain model-call usage. TS-only refinements/transforms and lossy conversions are
313
- unsupported; there is no partial-object stream or automatic repair.
314
-
315
- This engine does not yet provide the complete AI SDK Agent/`streamText` contract.
316
- The Worker engine uses built-in Rust HTTP/SSE protocols for the main model.
317
- The host supplies raw fetch, credentials and scoped effects. Default compression
318
- and custom context strategies remain TS hooks; `runtimeContext.model()` resolves
319
- an ordinary AI SDK model lazily, and strategies can call `streamText` directly.
320
- Their model features are independent of Core protocol coverage. The optional
321
- Core-backed AI SDK model facade is an outward adapter. Hosted development requires
322
- the matching raw proxy server for the main Core model.
323
-
324
- A custom strategy is an object with optional `prepareStep`, `onStepEnd` and
325
- `onEnd` callbacks, using the AI SDK 7 names. These callbacks replace the default
326
- recipe. `prepareStep` receives durable `messages`, `contextSize`, `stepNumber`
327
- and `runtimeContext`; returning `{ messages }` replaces the active projection
328
- for subsequent steps and Runs. Returning nothing keeps it. `onStepEnd` also
329
- receives that step's `usage` and `finishReason`; `onEnd` receives the main loop's
330
- `totalUsage` and `finishReason`.
331
-
332
- ```ts
333
- const context = {
334
- async prepareStep({ messages, runtimeContext }) {
335
- const memory = await loadCommittedMemory(runtimeContext);
336
- return { messages: applyMemoryToCoveredPrefix(messages, memory) };
337
- },
338
- async onStepEnd({ messages, runtimeContext }) {
339
- runtimeContext.waitUntil(precomputeMemory(messages, runtimeContext));
340
- },
341
- };
342
- ```
343
-
344
- The functions in this short sketch are application-owned.
345
-
346
- `runtimeContext` provides trusted `identity`, `auth`, declared `env`, execution
347
- `environment`, typed `durableObjects`, `signal`, and the main public `modelId`.
348
- Call `model(modelId?)` to resolve a model through the existing Host, then use AI
349
- SDK `generateText` or `streamText` with that model and the cancellation signal.
350
- Register each auxiliary call's `{ model, usage, label?, status? }` through
351
- `recordUsage`. Status defaults to `completed`; use `failed` or `aborted` with
352
- `usage: null` when the provider did not return usage. Register a successful call
353
- before validating its text or committing memory, since those operations can fail
354
- after the model has consumed tokens. The SDK assigns the call id and tracks its
355
- Session write automatically.
356
- `writeToolOutput` is a writer returning `{ artifactUri }`, or `null` if storage
357
- is unavailable. Its input is `{ content, sha256, toolCallId, toolName }`.
358
-
359
- `runtimeContext.updateMessages(messages)` persists a replacement from an awaited
360
- foreground callback, including the final step or `onEnd`. Background work stores
361
- candidate memory in a DO; a later foreground callback applies it. The Worker
362
- waits for registered work before publishing its final success and usage, including
363
- when Computer is disabled. `metadata.modelCalls` records each main-model step and
364
- auxiliary call with a stable call id, Run id, model, source, optional label, status,
365
- completion time and usage. `metadata.totalUsage` is their known usage sum;
366
- `metadata.contextUsage` retains the auxiliary view. Hosted settlement prices each
367
- call using its own catalog target and writes a separate idempotent usage event.
368
- A completed model call stays completed even if the containing Run later fails.
369
- Unhandled callback or background errors fail the Run; completed model messages
370
- and registered usage survive that failure. Pass `signal` into external work so cancellation can finish promptly.
371
-
372
- AgentSession owns the active projection and separate transcript. Custom
373
- observations, reflection and coverage state belong in the application's chat DO.
374
- Commit memory and coverage together before removing the covered message prefix.
375
- Existing summary metadata is included when a custom callback reads messages, and
376
- is no longer separately injected after that callback replaces the projection.
377
- `stepNumber` is local to a Run, not a durable chat cursor.
378
-
379
- Context callbacks execute in the Agent bundle; snapshots contain only serializable
380
- descriptors and default recipe options. Ordinary imported modules are sufficient.
381
- Upgrade the SDK and rebuild/publish existing Agents to use the API. `waitUntil`
382
- covers the current invocation; it does not schedule future Runs or replay a model
383
- call after a crash.
384
-
385
- ## CLI Workflow
386
-
387
- Local validation, development, and packing need no hosted Agent record:
3
+ The TypeScript SDK for building GEA Agent applications.
388
4
 
389
5
  ```bash
390
- gea agent validate --json '{"cwd":"."}'
391
- gea agent dev --json '{"cwd":"."}'
392
- gea agent eval --json '{"cwd":".","judgeModel":"gea-model-1"}'
393
- gea agent pack --json '{"cwd":"."}'
6
+ npm install @gea-ai/agent-sdk
394
7
  ```
395
8
 
396
- `gea agent dev` treats every unqualified model ID, plus explicit
397
- `creative-reasoning/<model>` IDs, as first-party Creative Reasoning models. It
398
- calls Creative Reasoning directly when `CREATIVE_REASONING_API_KEY` is present;
399
- when the key is absent it falls back to the hosted web model proxy and the
400
- current `gea login` Workspace. An explicit `{"modelSource":"hosted"}` keeps
401
- hosted routing. `{"modelSource":"local"}` forces local routing and therefore
402
- requires the Creative Reasoning key for first-party models; other qualified
403
- local providers continue to use their configured local catalog.
404
- `gea agent eval` applies the same decision to both Agent models and the active
405
- `judgeModel`, so a direct evaluation catalog contains every model the run uses.
406
-
407
- Native Anthropic calls, including `crr-a-*` and `creative-reasoning-1.5`, apply
408
- the SDK's shared rolling prompt-cache policy before provider serialization.
409
- System prompts, conversation prefixes, and eligible Tool definitions retain
410
- cache markers across model steps and turns. Markers belong only to outbound
411
- requests and never enter the stored AgentSession context. Hosted calls retain
412
- their existing proxy-owned cache policy.
413
-
414
- `gea agent eval` discovers Cases and inherited Judges from the Benchmark next
415
- to each main Agent's `tools/`: root `benchmarks/` for the root Agent, including
416
- after additional main Agents are added, and `agents/<slug>/benchmarks/` for
417
- directory Agents. It runs each Case through that main Agent's stable slug route
418
- and writes results under `.gea/evals`. Use `{"agentKey":"support"}` to select
419
- one main Agent. A Case's optional
420
- `expected` value enables the built-in Autoeval Judge. Author semantic rubrics
421
- as `*.judge.md`; author deterministic trajectory checks as `*.judge.ts` with
422
- the `./evals` SDK export. Autoeval may query bounded run evidence across
423
- multiple model steps and finishes by calling a structured result Tool; it does
424
- not require the model's text response to be JSON:
425
-
426
- ```ts
427
- import { defineJudge } from "@gea-ai/agent-sdk/evals";
428
-
429
- export default defineJudge(({ messages }) => {
430
- const usedSearch = messages.some((message) =>
431
- message.parts.some((part) => part.type === "tool-search"),
432
- );
433
- return {
434
- reason: usedSearch ? "Search was used." : "Search was not used.",
435
- score: usedSearch ? 1 : 0,
436
- };
437
- });
438
- ```
439
-
440
- `gea agent dev` and `gea agent eval` generate `.gea/bindings.d.ts` from the
441
- owning main Agent's `tools/` and snapshot-known runtime Tool names. This
442
- registers the canonical root SDK types:
443
-
444
- ```ts
445
- import agent from "./agent";
446
- import type {
447
- AgentMessage,
448
- AgentMessageFor,
449
- InferAgentMessage,
450
- } from "@gea-ai/agent-sdk";
451
-
452
- type DefaultMessage = AgentMessage;
453
- type SupportMessage = AgentMessageFor<"support">;
454
- type ThisAgentMessage = InferAgentMessage<typeof agent>;
455
- ```
456
-
457
- All three are ordinary AI SDK `UIMessage` types. Static custom `tool-*` parts
458
- preserve Tool names, Zod input types, and return types. Snapshot-known Computer
459
- and Connector Tools preserve their concrete `tool-*` names with generic input
460
- and output. MCP Tool parts use the qualified template
461
- `tool-<alias>__${string}` because their catalog is runtime-discovered; truly
462
- unqualified runtime Tools remain `dynamic-tool` parts. SDK-first Agent Workers
463
- currently emit no custom `data-*` parts, so Legacy chat data parts are not
464
- included. Code Judges consume the same types rather than defining an
465
- Eval-specific message model.
466
- `InferAgentMessage<typeof agent>` resolves to `never` when the generated
467
- registry is absent or its Agent slug does not match, rather than silently
468
- falling back to an unrelated Tool catalog.
469
-
470
- Studio also runs `defineApiConnector`, `defineMcpConnector`, and deployment-enabled
471
- built-in `geaConnect` providers through the managed Host. All use existing
472
- Project/Agent Connections, isolated by Preview/Production and application
473
- principal. No database migration or personal credential inheritance is involved.
474
- API keys default to `connection: { principalType: "agent" }`; OAuth and
475
- `geaConnect` default to `"user"`. All three helpers accept an explicit
476
- `connection: { principalType: "agent" | "user" }`. No-auth API/MCP endpoints
477
- need no Connection. Custom executable `defineConnector` ownership stays unchanged.
478
-
479
- Configure Agent-owned credentials in Project Connections or an Agent override.
480
- User-owned authorization links bind the application's supplied principal; missing
481
- user identity returns `principal_required`. OAuth client variables come from the
482
- selected environment. Register `${APP_BASE_URL}/api/agent-connectors/oauth/callback`
483
- for hosted OAuth, including the deployment Google OAuth client used by Google Drive.
484
- API keys are submitted on a public, state-bound setup page; OAuth uses PKCE;
485
- Feishu uses the host's isolated managed CLI. Credentials and refresh state remain
486
- encrypted in the original Connection scope and never enter Worker code.
487
-
488
- Rebuild old `geaConnect` snapshots to declare Studio ownership. Supported native
489
- providers are Google Drive, Feishu, Social Search, and WeChat Official Account,
490
- subject to deployment availability. Other platform resource references should
491
- use `defineApiConnector` or `defineMcpConnector` with an explicit endpoint.
492
- CLI upload preparation rejects old personal-authorization declarations and GEA
493
- user identity injection with the Connector alias, key, runtime, and reason.
494
- The server additionally checks native provider availability before publication.
495
-
496
- Configure a package-local API or MCP Connector without uploading its
497
- credential:
498
-
499
- ```bash
500
- gea agent connect exa-market-intelligence
501
- ```
502
-
503
- An MCP Connector declares only its stable server and authentication boundary:
504
-
505
- ```ts
506
- import { defineMcpConnector, noAuth } from "@gea-ai/agent-sdk";
507
-
508
- export default defineMcpConnector({
509
- auth: noAuth(),
510
- key: "product-docs",
511
- name: "Product Docs",
512
- serverUrl: "https://mcp.example.com/mcp",
513
- }).require();
514
- ```
515
-
516
- Its current Tools and schemas are discovered at the start of every run and
517
- exposed as `product-docs__<tool>`. They are not persisted in the Agent Package.
518
-
519
- API-key input is hidden. OAuth uses Authorization Code plus S256 PKCE through
520
- `http://127.0.0.1:8788/oauth/callback`. Credentials stay under `~/.gea`,
521
- scoped to canonical project path and Connector definition, and are reread for
522
- each local Tool call.
523
-
524
- Platform Connectors declared with `geaConnect(...)` are called through the
525
- Workspace selected by `gea login` and `gea workspace use`; their platform-owned
526
- credentials are never copied into the local project. This same local Connector
527
- Host behavior is used by both `gea agent dev` and `gea agent eval`.
528
-
529
- After `gea workspace use`, publish the complete Worker application:
530
-
531
- ```bash
532
- gea agent push --json '{"cwd":".","slug":"product-worker"}'
533
- ```
534
-
535
- The command uploads one generic Worker deployment and registers every main
536
- Agent in Studio by Worker plus Agent key. The new deployment becomes Preview.
537
- It never creates or updates Legacy `app_resource` or resource-package rows.
538
- Preview and Production promotion are Worker-wide.
539
-
540
- Use a Studio Agent ID to start its selected environment:
541
-
542
- ```bash
543
- gea agent run --json '{"agentId":"<studio-agent-id>","prompt":"Inspect the project"}'
544
- ```
545
-
546
- The server pins the exact Studio Agent version, Agent key, Worker ID, and Worker
547
- deployment to the Run. Recovery never follows a later environment target or
548
- falls back to Legacy.
549
-
550
- Custom Tools may omit `execute` to delegate execution to the calling application.
551
- These Tools are recorded as `execution: "client"` in the package snapshot and
552
- pause with an `input-available` Tool part. The application returns the result
553
- through AI SDK `addToolOutput`. Studio Tool approvals and client outputs use the
554
- existing `/run` endpoint with the server's assistant message ID.
555
- This capability requires updated SDK, Web/ORPC and Agent deployments.
556
-
557
- Package Tools may declare a static or input-dependent approval policy. A
558
- `user-approval` decision pauses before execution and can be resumed from any
559
- supported Chat client. For example:
560
-
561
- ```ts
562
- import { defineTool } from "@gea-ai/agent-sdk";
563
- import { z } from "zod";
564
-
565
- export default defineTool({
566
- approval: ({ amount }) => (amount > 100 ? "user-approval" : "not-applicable"),
567
- description: "Transfer credits after any required user approval.",
568
- execute: async ({ amount }) => ({ transferred: amount }),
569
- input: z.object({ amount: z.number().positive() }),
570
- name: "transfer_credits",
571
- });
572
- ```
573
-
574
- ```text
575
- coding-agent/
576
- agent.ts
577
- AGENTS.md
578
- connectors/github.ts
579
- hooks/before-turn.ts
580
- durable-objects/agent-counter.ts
581
- skills/project-conventions/SKILL.md
582
- tools/agent.ts
583
- tools/search-github.ts
584
- tools/chat-list.ts
585
- tools/chat-read.ts
586
- tools/chat-send.ts
587
- tools/computer.ts
588
- tools/web-search.ts
589
- ```
590
-
591
- ```ts
592
- // agent.ts
593
- import { DurableObject } from "cloudflare:workers";
594
- import { defineAgent } from "@gea-ai/agent-sdk";
595
- import { AgentCounter } from "./durable-objects/agent-counter";
596
-
597
- export { AgentCounter };
598
-
599
- export default defineAgent({
600
- computer: {
601
- enabled: true,
602
- filesystem: { durability: "durable", scope: "user" },
603
- },
604
- durableObjects: { COUNTER: AgentCounter },
605
- model: "gea-model-1",
606
- name: "coding-agent",
607
- slug: "coding-agent",
608
- });
609
- ```
610
-
611
- ```ts
612
- // hooks/before-turn.ts
613
- import { defineHook } from "@gea-ai/agent-sdk";
614
-
615
- export default defineHook({
616
- event: "beforeTurn",
617
- run: ({ identity, now }) => ({
618
- prepend: [
619
- `Current time: ${now.toISOString()}`,
620
- `Act only for workspace ${identity.workspace.id}.`,
621
- ].join("\n"),
622
- }),
623
- });
624
- ```
625
-
626
- Durable Object implementations declared by `defineAgent()` must be named
627
- exports from `agent.ts` in a headless project, or from `worker.ts` when an
628
- explicit application Worker exists. Keep the same exported class name,
629
- binding, and object ID to retain state across releases. Install
630
- `@cloudflare/workers-types` for Cloudflare-compatible authoring types.
631
-
632
- Public methods on the class are native typed RPC methods. The generated
633
- `.gea/bindings.d.ts` connects the binding to the exported class, so separate
634
- Tool files get method, input, and result autocomplete without internal URLs or
635
- a second contract definition:
636
-
637
- ```ts
638
- // durable-objects/agent-counter.ts
639
- import { DurableObject } from "cloudflare:workers";
640
-
641
- export class AgentCounter extends DurableObject {
642
- async increment({ by }: { by: number }) {
643
- const value = ((await this.ctx.storage.get<number>("value")) ?? 0) + by;
644
- await this.ctx.storage.put("value", value);
645
- return { value };
646
- }
647
- }
648
- ```
649
-
650
- ```ts
651
- // tools/increment-counter.ts
652
- import { defineTool } from "@gea-ai/agent-sdk";
653
- import { z } from "zod";
654
-
655
- export default defineTool({
656
- description: "Increment the workspace counter.",
657
- execute: async ({ by }, context) => {
658
- const counter = context.durableObjects.COUNTER.getByName(
659
- `workspace:${context.identity.workspace.id}`,
660
- );
661
- return await counter.increment({ by });
662
- },
663
- input: z.object({ by: z.number().int() }),
664
- name: "increment_counter",
665
- });
666
- ```
667
-
668
- RPC arguments and results use structured clone. `fetch`, alarms, and WebSocket
669
- lifecycle handlers remain lifecycle methods rather than model-facing RPC.
670
- Browser code should still call an authenticated Worker API and must not receive
671
- a Durable Object namespace or choose an identity-derived object key.
672
-
673
- `AGENTS.md` is required and automatically becomes the only instruction
674
- entrypoint. Every discovered `tools/**/*.ts`, `connectors/**/*.ts`,
675
- `hooks/**/*.ts`, and direct `skills/*.ts` file must default-export its SDK
676
- definition. Local Skills are discovered from `skills/<name>/SKILL.md`. Tests,
677
- specs, and `_`-prefixed files are ignored.
678
-
679
- SDK-first Agents do not inherit legacy runtime reminders. Use the optional
680
- `beforeTurn` hook for dynamic hidden context derived from the current message,
681
- stable identity, declared environment, or current time. The runtime prepends
682
- the returned text to the active user message, stores that transformed message
683
- in AgentSession model context before calling the model, and reuses it for Tool
684
- steps, approval continuation, and same-Run recovery. It does not change the
685
- public Chat message. Keep stable instructions in `AGENTS.md` so prompt caching
686
- can reuse the system prefix.
687
-
688
- When the exact Package Version contains Skills, the runtime adds one stable
689
- `loadSkill` Tool, including when Computer is disabled. Call
690
- `loadSkill({ skill: "report" })` to read `SKILL.md`, or
691
- `loadSkill({ skill: "report", path: "references/format.md" })` to read another
692
- UTF-8 text file relative to that Skill's root. Absolute paths and traversal are
693
- rejected. Each read returns `{ skill, path, content }`, with a relative `path`,
694
- and is limited to 512 KiB. When the existing Computer initialization has prepared
695
- the mount, the result also includes `executionDirectory`, the selected Skill's
696
- read-only root in that Computer. Binary files are not a text-reading surface.
697
-
698
- Both entry instructions and supporting files come directly from the exact
699
- immutable Worker bundle; reading them never opens Computer. The Tool description
700
- lists all available Skill slugs and descriptions and explains the optional
701
- `executionDirectory`. The directory comes from the initialized session's actual
702
- roots; the description contains no physical path assumptions. Script execution
703
- uses the Agent's provided execution Tools. Skill loading does not grant execution,
704
- install interpreters, or add system instructions.
705
-
706
- The CLI uses the same SDK Tool builder as the Worker to emit
707
- `definition/skill-tools/<agentKey>.json` in each expanded dev artifact and Agent
708
- archive. Inspect `.gea/agent-dev/definition/skill-tools/<agentKey>.json` after
709
- `gea agent dev` to see the final `description` and `inputSchema`. The preview
710
- contains the complete static Skill catalog, not file contents or an additional
711
- runtime configuration. Use matching updated CLI and Agent SDK versions to build
712
- these artifacts; existing published Workers keep their bundled behavior.
713
-
714
- `createAgentSkillTool({ files, snapshot, mountedRoot? })` accepts only bundled
715
- files, the snapshot, and an optional already prepared mount root. It never creates
716
- a Computer or mounts files. Provider session initialization prepares Skills before
717
- the model runs and passes this root to the Tool builder. Legacy Computer
718
- configuration retains lazy mounting and omits `executionDirectory` when the Tool
719
- is constructed before mounting. Existing Computer capability validation still
720
- applies. `readFile` remains available for Computer filesystem access, while
721
- `loadSkill` is the portable Skill text reader.
722
-
723
- The generated Agent Worker constructs the complete model-facing Tool set inside
724
- V8 from the bundled package snapshot and Tool runtime. Web initializes only
725
- runtime capabilities, Connector state, and dynamic Connector input schemas;
726
- the run body contains no serialized AI SDK Tools. Privileged Connector,
727
- conversation, and search execution remains behind the private Tools binding.
728
-
729
- ```ts
730
- // tools/web-search.ts
731
- import { webSearch } from "@gea-ai/agent-sdk";
732
-
733
- export default webSearch();
734
- ```
735
-
736
- Conversation capabilities are selected the same way and remain independent:
737
-
738
- ```ts
739
- // tools/chat-list.ts
740
- import { chatList } from "@gea-ai/agent-sdk";
741
-
742
- export default chatList();
743
- ```
744
-
745
- Use `agentTool()` to let the model create an isolated child Agent conversation.
746
- The model-visible Tool is named `agent`; omitting its optional `agentSlug`
747
- starts the current Agent in a fresh context. The returned `chatId` is then
748
- handled by the ordinary Conversation Tools: `chatList()` discovers child
749
- conversations, `chatRead()` reads their bounded transcript, and `chatSend()`
750
- continues an idle conversation. There is no separate `subagent_*` Tool family.
751
-
752
- ```ts
753
- // tools/agent.ts
754
- import { agentTool } from "@gea-ai/agent-sdk";
755
-
756
- export default agentTool();
757
- ```
758
-
759
- Select `chatRead()` and `chatSend()` in their corresponding files only when the
760
- Agent should be able to inspect or continue conversations.
761
-
762
- One Tool module may group reusable selections without changing the serialized
763
- Agent Package contract:
764
-
765
- ```ts
766
- // tools/research.ts
767
- import { chatRead, defineToolSet, webSearch } from "@gea-ai/agent-sdk";
768
-
769
- export default defineToolSet([chatRead(), webSearch()]);
770
- ```
771
-
772
- Tool Sets may contain other Tool Sets. Their entries are flattened from left to
773
- right, and the CLI preserves that order after deterministic file discovery.
774
-
775
- Computer setup has two independent parts. First enable the runtime in
776
- `agent.ts`:
777
-
778
- ```ts
779
- // agent.ts
780
- import { defineAgent } from "@gea-ai/agent-sdk";
781
-
782
- export default defineAgent({
783
- computer: {
784
- enabled: true,
785
- filesystem: { durability: "durable", scope: "chat" },
786
- },
787
- model: "gea-model-1",
788
- name: "research-agent",
789
- slug: "research-agent",
790
- });
791
- ```
792
-
793
- New Agent code can choose its Computer provider directly:
794
-
795
- ```ts
796
- import {
797
- defaultComputer,
798
- localComputer,
799
- remoteComputer,
800
- secret,
801
- } from "@gea-ai/agent-sdk/computer";
802
- import { justbash } from "@gea-ai/agent-sdk/computer/just-bash";
803
- import { e2b } from "@gea-ai/computer-e2b";
804
-
805
- const computer = defaultComputer({
806
- filesystem: {
807
- scope: ({ principal }) => {
808
- if (!principal) throw new Error("A business principal is required");
809
- return `${principal.principalType}:${principal.principalId}`;
810
- },
811
- files: { "README.md": "Workspace initialized by this Agent." },
812
- },
813
- async onSession({ use, ctx }) {
814
- const session = await use();
815
- await session.writeTextFile({
816
- path: `${session.roots.tmp}/principal.json`,
817
- content: JSON.stringify(ctx.principal),
818
- });
819
- },
820
- });
821
-
822
- // Other choices for defineAgent({ computer: ... }):
823
- const virtual = justbash({ filesystem: { durability: "ephemeral" } });
824
- const external = e2b({
825
- apiKey: secret("E2B_API_KEY"),
826
- filesystem: { durability: "ephemeral" },
827
- });
828
- const remote = remoteComputer({
829
- url: "https://computer.example.com",
830
- token: secret("COMPUTER_TOKEN"),
831
- });
832
- const local = localComputer({
833
- mode: "sandbox",
834
- filesystem: { source: "local-project" },
835
- });
836
- ```
837
-
838
- Each factory returns one Computer with a matching filesystem. The code selects
839
- the provider and any remote endpoint. The existing COMPUTER Native binding connects
840
- directly to E2B or the developer service; just-bash is bundled into the Worker.
841
- Provider secret references are resolved only for Native Runtime. Do not separately declare those keys in the Agent's
842
- `environment` unless the Agent itself intentionally needs them.
843
-
844
- Computer file I/O follows Eve / AI SDK: `readFile/writeFile` stream bytes,
845
- `readBinaryFile/writeBinaryFile` handle Uint8Array, and text variants decode or
846
- encode strings. Options contain `abortSignal`. Reads return `null` for absent
847
- files; permission/I/O failures reject. Buffered reads are limited to 16 MiB;
848
- use streams for larger files. `stat`, `mkdir`, `rename`, `remove`, and `listFiles`
849
- operate on that same filesystem. Relative paths use `computer.roots.workspace`.
850
- Remote services use the Remote Computer HTTP v1 protocol for files and sessions.
851
-
852
- `computer.commands.run()` accepts a shell `command`, literal `args`, `env`,
853
- `cwd`, `timeoutMs` and initial text `stdin`. Simple foreground calls return
854
- `{ stdout, stderr, exitCode }`. Providers advertising `processes` also support
855
- background handles, incremental output and interactive stdin:
856
-
857
- ```ts
858
- const process = await computer.commands.run({
859
- command: "python",
860
- args: ["-u", "script.py"],
861
- env: { MODE: "test" },
862
- timeoutMs: 30_000,
863
- stdin: true,
864
- background: true,
865
- });
866
- const finished = process.wait({ onStdout: (text) => console.log(text) });
867
- await process.sendStdin("input\n", { end: true });
868
- const result = await finished;
869
- // To interrupt a running command: await process.kill().
870
- ```
871
-
872
- `wait()` starts output consumption and closes the handle when finished; repeated
873
- calls share its result and its first callbacks. Use `background: true` to supply
874
- interactive stdin, and `end: true` to close it. The command/owner signal or timeout
875
- also interrupts the process. Commands are limited to 120 seconds and collected
876
- output to 2 MiB (a provider may impose lower limits); stdin writes are at most
877
- 64 KiB. Callbacks receive decoded stdout/stderr chunks, not a terminal/PTY.
878
- E2B and local allow eight retained handles; wait for finished handles to release
879
- capacity. Hosted default uses the existing daemon's configured process limits.
880
-
881
- Hosted default, E2B and authorized local execution implement `processes`.
882
- just-bash and default dev support foreground options only and explicitly reject
883
- background/interactive/streaming requests. A remote service must implement the
884
- optional process-control protocol before advertising the `processes` capability.
885
- Every process belongs to the current invocation: Worker stop terminates unfinished
886
- processes before releasing the session or settling its filesystem. Background
887
- processes cannot outlive a completed Agent run. Normal foreground completion also
888
- cleans up redirected descendants. Local foreground/background commands share UTF-8
889
- decoding and process cleanup; the foreground API still uses one request.
890
- Execution user, PTY, listening
891
- ports, pause/resume and snapshots are not added by this API.
892
-
893
- `defaultComputer()` retains durable/chat defaults and supports a synchronous
894
- scope callback with `chatId`, `principal`, and `metadata`. Business code supplies
895
- `principal`; Computer never reads a GEA `userId`. `files` is a map of normalized
896
- relative workspace paths to strings or `Uint8Array` (2 MiB per file, 4 MiB total,
897
- 1000 files). It seeds only a new allocation. Upgrades, reconnections and deleted
898
- seed files do not reseed an existing durable workspace; legacy allocations are
899
- preserved. Ephemeral providers seed each new physical session.
900
-
901
- `onSession({ use, ctx })` runs after seeds and bundled Skills are ready, before
902
- model Tools. `use()` supplies that same session. It runs once per physical
903
- session, including replacement after confirmed expiry, and is skipped for live
904
- reconnects. Make the callback idempotent and check nonzero command exit codes;
905
- failed initialization prevents Tool execution. No unknown command is replayed.
906
-
907
- Computer lifecycle uses `create()` / `stop()`. The Worker owns it: create and
908
- initialize before tools, then stop after the model, tools, context writes and
909
- response stream have settled. Initialization/setup failures and cancellation also
910
- release the session, with a separate cleanup signal. Stop preserves durable files
911
- and the local project; ephemeral files disappear. Repeated stop is safe and must
912
- not create a new runtime or affect a newer generation. Services need their own
913
- expiry policy for Worker crashes or unreachable networks. The updated SDK requires
914
- a lifecycle-capable Worker Runtime; Native retains open/close aliases for older
915
- immutable Agent bundles. Lifecycle methods are not exposed as model Tools.
916
-
917
- `justbash()` provides virtual shell commands in Worker memory, with a 32 MiB public-write
918
- limit. It and E2B currently create ephemeral sessions per invocation; use Artifacts
919
- for retained files. No independent GEA Computer Host or additional Pod is required.
920
- Developer-owned services implement the Remote Computer HTTP v1 protocol
921
- and connect through `remoteComputer`. Developers choose their image, installed
922
- CLI tools, environment, and filesystem persistence. The URL's `filesystemId`
923
- identifies authorized storage; sessionId/invocationId/generation identify execution
924
- state. There is no Computer SDK, prescribed service framework, or base image in
925
- this release. Declared capabilities are checked against the actual session.
926
-
927
- `localComputer` requires local execution and matching explicit CLI authorization:
928
- `gea agent dev --json '{"localComputer":"sandbox","sandboxdPath":"/absolute/path/to/sandboxd"}'`.
929
- It uses the existing project, accepts no seed files, and never deletes the project
930
- on stop. Tools use returned roots and preserve literal shell commands.
931
- Trusted mode runs as the OS user and cannot enforce read-only Skills. macOS
932
- sandbox mode has native tests; Linux requires host-supplied bwrap configuration,
933
- and Windows local mode is currently unsupported. GUI operations are not exposed.
934
- Hosted Agents reject local mode. The factories do not add model-visible Tools;
935
- select them using the Tool files below.
936
-
937
- This provides the isolated Computer and makes `context.computer` available to
938
- custom Tools. `scope` declares whether the durable filesystem belongs to one
939
- chat, the current user, or a custom partition; it does not select a physical storage provider.
940
- It does not expose any built-in Computer Tool to the model.
941
-
942
- Runtime sessions and their read-only Skill mounts can be reclaimed while the
943
- durable filesystem remains. With the matching updated Worker Runtime, the SDK
944
- reconstructs the bundled Skill mount and retries a confirmed pre-execution
945
- `MOUNT_MISMATCH` once. It does not automatically replay commands after network
946
- or persistence failures. Publishing the SDK update requires updating hosted
947
- Worker Runtime first; existing immutable Agent bundles retain their old behavior.
948
-
949
- Then explicitly select the model-visible Computer Tools through the same Tool
950
- file convention. Each SDK factory fixes the Tool ID, input schema, and runtime
951
- implementation; the caller may override only its description:
952
-
953
- ```ts
954
- // tools/computer.ts
955
- import {
956
- applyPatch,
957
- bash,
958
- defineToolSet,
959
- editFile,
960
- findFiles,
961
- listDirectory,
962
- readFile,
963
- searchText,
964
- writeFile,
965
- } from "@gea-ai/agent-sdk";
966
-
967
- export default defineToolSet([
968
- readFile(),
969
- listDirectory(),
970
- findFiles(),
971
- searchText(),
972
- writeFile({
973
- description:
974
- "Write new or small complete files. Build long reports in bounded editFile steps.",
975
- }),
976
- editFile(),
977
- applyPatch(),
978
- bash(),
979
- ]);
980
- ```
981
-
982
- This array is the complete model-visible Computer Tool set and order. Only
983
- listed Computer Tools are visible: enabling Computer alone does not add
984
- `bash`, `readFile`, `writeFile`, or any other built-in Tool. Declaring no
985
- Computer Tools leaves the model with none, while custom Tools may still use
986
- `context.computer`.
987
-
988
- | Factory | Fixed Tool ID | Purpose |
989
- | ----------------- | --------------- | ------------------------------------------------------------- |
990
- | `readFile()` | `readFile` | Read bounded text from files, the Agent package, or artifacts |
991
- | `listDirectory()` | `listDirectory` | List one bounded directory |
992
- | `findFiles()` | `findFiles` | Find files with a glob |
993
- | `searchText()` | `searchText` | Search text across files |
994
- | `writeFile()` | `writeFile` | Create a file or completely write a small file |
995
- | `editFile()` | `editFile` | Prefer for exact, unique replacements in an existing file |
996
- | `applyPatch()` | `applyPatch` | Apply a Codex patch to create, update, move, or delete files |
997
- | `bash()` | `bash` | Run a Bash command from `/workspace` |
998
-
999
- The CLI discovers `tools/**/*.ts` deterministically and preserves Tool Set
1000
- array order. The default `writeFile()` description tells the model to create a
1001
- small skeleton for a long report and fill it with bounded `editFile()` calls
1002
- instead of sending one large full-report Tool input.
1003
-
1004
- Prefer `editFile()` for ordinary edits to an existing file. Read the file
1005
- first and provide exact, unique `oldText`/`newText` replacements; replacements
1006
- are applied in order, so later edits see earlier edits in the same call.
1007
-
1008
- `applyPatch()` accepts **only Codex patch syntax**. It supports Add File,
1009
- Delete File, Update File with optional Move to, multiple chunks, `@@` context
1010
- anchors, and `*** End of File`. It uses a pinned upstream TypeScript port;
1011
- see `THIRD_PARTY_NOTICES.md` included in this package. Update matching follows
1012
- Codex's context and whitespace rules. Prefer enough surrounding context or a
1013
- named `@@` anchor when the same text appears more than once.
1014
-
1015
- ```text
1016
- *** Begin Patch
1017
- *** Update File: notes.md
1018
- @@ ## Next steps
1019
- -status: pending
1020
- +status: done
1021
- *** Add File: notes/summary.md
1022
- +Work completed.
1023
- *** End Patch
1024
- ```
1025
-
1026
- Paths are literal workspace-relative paths or absolute paths under
1027
- `/workspace`; `a/` and `b/` are actual directory names. Unified diffs and
1028
- Environment ID routing are not supported by this Computer Tool. All paths
1029
- are validated before mutation, and file operations run in order, stopping at
1030
- the first failure. A multi-file patch is not a transaction: `changedPaths`
1031
- reports completed writes/deletes even on failure, `results` lists completed
1032
- operations (a move uses its destination path), and `conflicts` contains the
1033
- failure message. Inspect these outputs and reread changed files before retrying.
1034
-
1035
- This replaces the previous unified-diff input contract. Existing immutable
1036
- Agent bundles retain their embedded SDK; rebuild and publish each Agent with
1037
- the updated SDK to adopt the new behavior. Update custom tool descriptions
1038
- and any callers that still generate unified diffs at the same time.
1039
-
1040
- The package tree above is the minimal starting point. Add Connector-dependent
1041
- custom Tools and referenced Skills only when the Agent needs those capabilities.
1042
-
1043
- `gea agent dev` also generates `.gea/bindings.d.ts`. A custom Tool receives
1044
- configured environment values, capability-bound stable identity, and declared
1045
- Durable Object namespaces without receiving model, Connector, Computer, or
1046
- Durable Object service credentials. A configured `tsconfig.json` should include
1047
- `.gea/bindings.d.ts`; inferred TypeScript projects discover it automatically.
1048
-
1049
- ```ts
1050
- const counter = context.durableObjects.COUNTER.getByName(
1051
- context.identity.chat.id,
1052
- );
1053
- const response = await counter.fetch("https://counter.internal/increment", {
1054
- method: "POST",
1055
- });
1056
- ```
1057
-
1058
- ## Call a Studio Agent with AI SDK
1059
-
1060
- Use `StudioAgentChatTransport` from `@gea-ai/agent-sdk/studio-client` with
1061
- `useChat` from `@ai-sdk/react` in the browser, and `StudioAgentClient` from
1062
- `@gea-ai/agent-sdk/studio-server` on your backend. The browser authenticates to
1063
- your application; only your backend sends the Project key to GEA. The transport
1064
- reads response identity headers and reuses the server Chat on later turns.
1065
-
1066
- ```tsx
1067
- import { useChat } from "@ai-sdk/react";
1068
- import { StudioAgentChatTransport } from "@gea-ai/agent-sdk/studio-client";
1069
- import { useState } from "react";
1070
-
1071
- // Inside your component:
1072
- const [transport] = useState(
1073
- () =>
1074
- new StudioAgentChatTransport({
1075
- api: "/api/agent", // Your authenticated application path, without /run.
1076
- headers: { "x-csrf-token": csrfToken }, // From your application session.
1077
- onRun: ({ chatId, runId, requestId }) => {
1078
- console.log({ chatId, runId, requestId });
1079
- },
1080
- }),
1081
- );
1082
- const { messages, sendMessage, resumeStream } = useChat({ transport });
1083
- ```
1084
-
1085
- The transport sends your application's session cookies by default. It accepts
1086
- only same-origin paths and sends no business metadata. Set `headers` to your
1087
- application's CSRF token or user access token as appropriate.
1088
-
1089
- Create a key in **Project → API Key** and copy the matching URL from the
1090
- Agent detail overview. Production uses `worker--<workerId>.<apex>`; preview
1091
- uses `preview--worker--<workerId>.<apex>`. Both use `/gea/agents/<agentKey>/run`.
1092
- Remove the trailing `/run` for the SDK `api` option, which appends operation
1093
- paths itself. The key must match the URL environment.
1094
-
1095
- On the backend, configure the GEA destination and key from server environment:
1096
-
1097
- ```ts
1098
- import { StudioAgentClient } from "@gea-ai/agent-sdk/studio-server";
1099
- import { env } from "./env";
1100
-
1101
- const agent = new StudioAgentClient({
1102
- api: env.GEA_AGENT_API_URL, // Agent base URL without /run.
1103
- apiKey: env.GEA_PROJECT_API_KEY,
1104
- });
1105
- // In an application route, after authentication and Chat ownership checks:
1106
- const response = await agent.run(
1107
- { chatId, message },
1108
- { signal: request.signal },
1109
- );
1110
- ```
1111
-
1112
- The backend provides new-Chat metadata from trusted application state. Do not
1113
- forward browser Cookie/Authorization headers to GEA or expose an anonymous
1114
- proxy. The server client constructs its own headers, refuses redirects, and
1115
- filters response headers while preserving the stream. Its browser export is
1116
- blocked. See the [application integration guide](https://musegea.com/developers/agent-studio-quick-start#api-invocation)
1117
- for session/CSRF integration, Chat ownership checks and reconnecting. These
1118
- entrypoints are available in `0.1.260906-alpha.0` and later.
1119
-
1120
- Applications select the user through the backend Run request's `principal`.
1121
- Computer scope callbacks receive it directly as `principal`; existing Tool and
1122
- hook contexts expose the same value through `auth.current`. GEA resource identity
1123
- is separate and must not be used to infer the application principal. The SDK
1124
- neither fabricates a GEA user nor requires principal IDs to equal GEA user IDs.
1125
-
1126
- ## Lazy Computer allocation and scope resolvers
1127
-
1128
- SDK-first Web execution passes trusted identity and caller metadata to the Worker;
1129
- it no longer allocates a filesystem before dispatch. An Agent may declare `chat`,
1130
- `user`, or a synchronous function for `computer.filesystem.scope`. The function
1131
- receives `{ chatId, principal, metadata }` and returns a non-empty string of at most
1132
- 256 characters without control characters. The Worker bundle contains the function;
1133
- the immutable snapshot records only `scope: "custom"`.
1134
-
1135
- AgentSession pins the first resolved scope. For custom resolvers, later metadata
1136
- or principal changes never recompute the key; applications must authorize Chat
1137
- access accordingly. Built-in `"user"` scope instead selects the current Run's
1138
- user partition on each allocation and rechecks that a user principal is present. The function receives business metadata
1139
- from Chat creation / first-run input; the application backend must authenticate its
1140
- users and authorize that data. Metadata never grants GEA user permissions.
1141
-
1142
- The native Computer binding requests an allocation from the existing ORPC registry
1143
- on first use. Trusted Run/version context determines ownership. Custom partitions
1144
- include Workspace, Project, Agent and environment; changing API keys preserves
1145
- storage. Runtime coalesces concurrent first operations using an invocation-owned
1146
- cache and keeps allocation IDs, provider details and credentials out of JavaScript.
1147
- Cancellation during allocation prevents Computer dispatch. No Computer use means
1148
- no allocation. Computer opening and Skill mount initialization also remain lazy.
1149
-
1150
- New SDK-first Computer filesystems use shared POSIX for every scope. Locally this
1151
- is the configured filesystem directory; production mounts the common filesystem
1152
- at `AGENT_FILESYSTEM_ROOT` on the Computer hosts. SDK, API and Studio expose scope,
1153
- not provider selection. The Legacy deployment mode does not enable new DO filesystems
1154
- for Studio. Existing allocation identities/providers remain readable without migration.
1155
- User scope requires `principalType: "user"` and uses the current Run's
1156
- `principalId`, including opaque external application identifiers. Studio user
1157
- partitions include Workspace, Project, Agent and environment. Attributes do not
1158
- change that identity. Hosted omission selects the Project service principal and
1159
- cannot reuse a previous Run's user files. The caller supplies `principal`; the
1160
- platform business adapter converts GEA session users before SDK execution.
1161
- Computer does not infer identity from `userId` or `identity.user`.
1162
-
1163
- The local CLI forwards the same selected principal to its allocation callback.
1164
- User files are partitioned by local project, Agent and principal ID; omission
1165
- selects `local-user` and preserves its existing directory. Existing non-Studio
1166
- GEA-user allocation IDs and Chat/custom allocation IDs are retained. Its just-bash
1167
- bridge is a development runtime.
1168
-
1169
- Rollout requires matching Web/ORPC and Worker Runtime, plus Agents rebuilt and
1170
- published with this SDK. No database migration or virtual user is introduced.
1171
-
1172
- ```ts
1173
- computer: {
1174
- enabled: true,
1175
- filesystem: {
1176
- durability: "durable",
1177
- scope: ({ metadata }) => {
1178
- if (typeof metadata.customerId !== "string" || !metadata.customerId)
1179
- throw new Error("customerId is required");
1180
- return `customer:${metadata.customerId}`;
1181
- },
1182
- },
1183
- }
1184
- ```
1185
-
1186
- ## Project Artifacts
1187
-
1188
- Declare `artifacts: {}` to enable `ctx.artifacts` in custom Tools. The default scope
1189
- is the runtime principal; Project and environment are always fixed by the host.
1190
- Enable `artifacts: { tools: true }` for three standard model tools:
1191
-
1192
- | Tool | Behavior |
1193
- | ----------------------------------------------------- | ----------------------------------------------------------------------------------------------- |
1194
- | `listArtifacts({ query?, prefix?, cursor?, limit? })` | Case-insensitive filename substring search and scoped pagination. |
1195
- | `getArtifact({ key })` | Metadata and a temporary download URL, bounded text/JSON preview, and current-turn image reads. |
1196
- | `createUpload({ filename, contentType?, key? })` | A temporary HTTP upload target; optional key selects an output to replace. |
1197
-
1198
- A target includes `key`, `uploadUrl`, `method`, `headers`, `expiresAt`, and `maxBytes`.
1199
- Use every returned header and upload original bytes with an ordinary HTTP client.
1200
- For example, in a Computer with curl and network access:
1201
-
1202
- ```bash
1203
- curl --fail-with-body --silent --show-error -X PUT \
1204
- -H 'Authorization: Bearer <returned authorization>' \
1205
- -H 'Content-Type: application/pdf' \
1206
- --upload-file /workspace/report.pdf '<uploadUrl>'
1207
- ```
1208
-
1209
- The server verifies and publishes in that request; HTTP success returns the
1210
- Artifact. No completion tool is needed. To download, use the `downloadUrl` from
1211
- `getArtifact`, or share that URL directly with the user. If optional text/image
1212
- preview fails, `getArtifact` still returns metadata and the download URL together
1213
- with `previewError`; lookup/authorization failures remain Tool errors. No Computer is allocated
1214
- for Artifact metadata, image reads, or upload authorization. A network-disabled
1215
- Computer cannot use these HTTP URLs; the tools do not bypass its network policy.
1216
-
1217
- Custom Tools use `ctx.artifacts.createUpload`, `get`, `readText`, `list`, `delete`,
1218
- and bounded text `put`. For example:
1219
-
1220
- ```ts
1221
- const target = await ctx.artifacts.createUpload({
1222
- filename: "report.pdf",
1223
- contentType: "application/pdf",
1224
- });
1225
- const files = await ctx.artifacts.list({ query: "report.pdf", limit: 20 });
1226
- ```
1227
-
1228
- `put({key, content, contentType?, title?, filename?, metadata?})` stores UTF-8 text
1229
- up to 2 MiB. `readText(key)` accepts text/JSON up to 2 MiB. `getArtifact` limits its
1230
- model text preview to 8,000 characters and keeps image bytes transient. The SDK's
1231
- old Computer path-copy methods and dedicated transfer protocol have been removed.
1232
-
1233
- A developer can share a scope with trusted code:
1234
-
1235
- ```ts
1236
- artifacts: {
1237
- tools: true,
1238
- scope: ({ metadata }) => typeof metadata.customerId === "string" ? `customer:${metadata.customerId}` : null,
1239
- },
1240
- ```
1241
-
1242
- Authorize customer membership in the application before using business metadata
1243
- as a scope. Null/empty results fail closed. Matching scopes share across Agents
1244
- and Chats. An omitted upload key is `files/<uuid>/<filename>`. Explicit repeated
1245
- keys replace content while preserving Artifact ID and original creation provenance;
1246
- last committed write wins. The reserved `uploads/` namespace is immutable message
1247
- input and cannot be overwritten by Agent uploads. `createUpload` authority lasts
1248
- one hour and can be reused for its same key until expiration. After an uncertain
1249
- HTTP response, inspect the key; failure to receive a reply does not imply rollback.
1250
-
1251
- Downloads use attachment disposition, including HTML/SVG. Temporary links are
1252
- bearer access; get a fresh link from the persistent Artifact key when needed.
1253
- Hosted uploads use bounded disk staging and streaming object-store writes, with a
1254
- 1 GiB default limit (configurable up to 10 GiB); local dev supports 1 GiB and stores
1255
- files under `.gea/artifacts`. Ingress and client limits still apply. The saved copy
1256
- is independent of later Computer edits.
1257
-
1258
- For Agent authoring and local development, see the public
1259
- [Agent development guide](https://musegea.com/developers/agent-development).
9
+ [Developer documentation](https://musegea.com/developers/agent-development)