orchajs 0.4.0 → 0.6.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (55) hide show
  1. package/README.md +122 -2
  2. package/dist/cli-agents-template.d.ts +1 -1
  3. package/dist/cli-agents-template.d.ts.map +1 -1
  4. package/dist/cli-agents-template.js +140 -15
  5. package/dist/cli-agents-template.js.map +1 -1
  6. package/dist/cli.js +2 -0
  7. package/dist/cli.js.map +1 -1
  8. package/dist/compiler/build-project.d.ts.map +1 -1
  9. package/dist/compiler/build-project.js +7 -1
  10. package/dist/compiler/build-project.js.map +1 -1
  11. package/dist/compiler/compile.d.ts +2 -2
  12. package/dist/compiler/compile.d.ts.map +1 -1
  13. package/dist/compiler/compile.js +146 -67
  14. package/dist/compiler/compile.js.map +1 -1
  15. package/dist/index.d.ts +1 -1
  16. package/dist/index.d.ts.map +1 -1
  17. package/dist/providers/anthropic.js +1 -0
  18. package/dist/providers/anthropic.js.map +1 -1
  19. package/dist/providers/bedrock.d.ts +3 -1
  20. package/dist/providers/bedrock.d.ts.map +1 -1
  21. package/dist/providers/bedrock.js +1 -1
  22. package/dist/providers/bedrock.js.map +1 -1
  23. package/dist/providers/deepseek.js +1 -0
  24. package/dist/providers/deepseek.js.map +1 -1
  25. package/dist/providers/google-gemini.d.ts.map +1 -1
  26. package/dist/providers/google-gemini.js +15 -3
  27. package/dist/providers/google-gemini.js.map +1 -1
  28. package/dist/providers/openai.js +6 -0
  29. package/dist/providers/openai.js.map +1 -1
  30. package/dist/providers/types.d.ts +1 -0
  31. package/dist/providers/types.d.ts.map +1 -1
  32. package/dist/runtime/agent.d.ts +15 -2
  33. package/dist/runtime/agent.d.ts.map +1 -1
  34. package/dist/runtime/agent.js +817 -37
  35. package/dist/runtime/agent.js.map +1 -1
  36. package/dist/runtime/context.d.ts +2 -2
  37. package/dist/runtime/context.d.ts.map +1 -1
  38. package/dist/runtime/context.js.map +1 -1
  39. package/dist/runtime/create-orcha.d.ts.map +1 -1
  40. package/dist/runtime/create-orcha.js +22 -1
  41. package/dist/runtime/create-orcha.js.map +1 -1
  42. package/dist/runtime/execution.d.ts +3 -2
  43. package/dist/runtime/execution.d.ts.map +1 -1
  44. package/dist/runtime/execution.js +6 -1
  45. package/dist/runtime/execution.js.map +1 -1
  46. package/dist/runtime/session-events.d.ts.map +1 -1
  47. package/dist/runtime/session-events.js +6 -3
  48. package/dist/runtime/session-events.js.map +1 -1
  49. package/dist/testing.d.ts +2 -1
  50. package/dist/testing.d.ts.map +1 -1
  51. package/dist/testing.js +23 -1
  52. package/dist/testing.js.map +1 -1
  53. package/dist/types.d.ts +65 -10
  54. package/dist/types.d.ts.map +1 -1
  55. package/package.json +7 -2
package/README.md CHANGED
@@ -14,11 +14,30 @@ Read the full concept: [`docs/concept.md`](./docs/concept.md)
14
14
 
15
15
  ---
16
16
 
17
+ ## Create a project
18
+
19
+ Scaffold a complete Node.js application with example agents and a local
20
+ browser playground:
21
+
22
+ ```bash
23
+ npm create orcha@latest my-agent-app
24
+ cd my-agent-app
25
+ npm install
26
+ cp .env.example .env
27
+ npm run dev
28
+ ```
29
+
30
+ The starter demonstrates durable sessions, actions, client-action pauses,
31
+ skills, tests, evaluations, and multimodal input without a frontend framework
32
+ or hosted dependency.
33
+
34
+ ---
35
+
17
36
  ## The convention
18
37
 
19
38
  ```
20
39
  /agentname
21
- index.js | index.json → provider, model, generation config, output schema
40
+ index.json → name, description, model, limits, output schema
22
41
  instructions.md → required base agent instructions
23
42
  /skills
24
43
  /skillname
@@ -239,7 +258,7 @@ npx orcha build
239
258
  `run` and `test` support `--json` for scripts and coding agents.
240
259
  `orcha init` also creates a comprehensive `AGENTS.md` describing Orcha's
241
260
  filesystem contract and APIs. The CLI does not provide a UI; applications can
242
- build one from the runtime session APIs or durable JSONL logs.
261
+ build one from the storage-neutral runtime session APIs.
243
262
 
244
263
  Orcha owns its built-in provider adapters. Messages, tool calls and results,
245
264
  streaming snapshots, usage, and errors are normalized before reaching the
@@ -247,6 +266,88 @@ runtime, so application and action APIs do not change when an agent switches
247
266
  providers. Provider-specific capabilities that cannot be represented safely
248
267
  fail with an explicit Orcha error instead of silently degrading.
249
268
 
269
+ ### Parent agents and subagents
270
+
271
+ Register private subagents on a top-level agent:
272
+
273
+ ```js
274
+ orcha.init({
275
+ providers: {
276
+ openai: process.env.OPENAI_API_KEY,
277
+ },
278
+ agents: {
279
+ coordinator: {
280
+ path: "./coordinator",
281
+ subagents: {
282
+ researcher: "./researcher",
283
+ },
284
+ },
285
+ researcher: "./researcher",
286
+ },
287
+ });
288
+ ```
289
+
290
+ Only top-level registrations are exposed directly, so `orcha.coordinator`
291
+ and `orcha.researcher` exist in this example. Removing the top-level
292
+ `researcher` registration makes it private while preserving delegation.
293
+
294
+ Every agent `index.json` requires a model-facing `name`; `description` is
295
+ optional and helps parent models understand when to delegate. A parent may
296
+ also configure delegation limits:
297
+
298
+ ```json
299
+ {
300
+ "name": "Coordinator",
301
+ "description": "Delegate research and assemble the final answer.",
302
+ "provider": "openai",
303
+ "model": "gpt-5-mini",
304
+ "subagents": {
305
+ "maxPerRun": 3
306
+ }
307
+ }
308
+ ```
309
+
310
+ `maxPerRun` limits new child sessions in one parent run and defaults to `10`.
311
+
312
+ Starting or continuing a child is a synchronous internal tool call. While it
313
+ runs, the parent reports `waiting_for_subagent`. Child text, pause state,
314
+ client-action request, failure, or completion returns to the parent as a tool
315
+ result. Delegated agents cannot start further subagents.
316
+
317
+ Applications inspect a durable child through its owning parent agent:
318
+
319
+ ```js
320
+ const history = await orcha.coordinator.subagentHistory(childSessionId);
321
+ ```
322
+
323
+ The child session ID is included in the parent's projected history and
324
+ `subagent.initiated` event. Access is lineage-checked, so another parent agent
325
+ cannot inspect it. Parent traces record each child transition as
326
+ `subagent.initiated`, `subagent.paused`, `subagent.resumed`,
327
+ `subagent.completed`, or `subagent.failed`.
328
+
329
+ `orcha.coordinator.pause(sessionId)` immediately aborts active provider work
330
+ for the parent and active children. Already streamed text is persisted as an
331
+ incomplete assistant message. `resume()` always starts a new run with the
332
+ complete prior history, including that incomplete message:
333
+
334
+ ```js
335
+ await orcha.coordinator.pause(parentSessionId);
336
+
337
+ // Adds "Continue." as the next user message.
338
+ await orcha.coordinator.resume(parentSessionId).result;
339
+
340
+ // Or provide different instructions.
341
+ await orcha.coordinator.resume(parentSessionId, {
342
+ content: "Continue, but only use confirmed findings.",
343
+ }).result;
344
+ ```
345
+
346
+ A child waiting for a client action returns control to the parent. The parent
347
+ can resolve it through its internal `resume_agent` tool, ask its own client
348
+ for information, inspect the child history, or finish without resuming the
349
+ paused child. Paused children do not block parent completion.
350
+
250
351
  Anthropic, Amazon Bedrock Converse, DeepSeek Chat Completions, OpenAI
251
352
  Responses, the Gemini Developer API through Google GenAI, and Vertex AI are
252
353
  built in:
@@ -300,6 +401,8 @@ ID, or ARN:
300
401
 
301
402
  ```json
302
403
  {
404
+ "name": "Bedrock Agent",
405
+ "description": "Handle requests using the configured Bedrock model.",
303
406
  "provider": "bedrock",
304
407
  "model": "us.anthropic.claude-sonnet-4-6",
305
408
  "outputType": "text"
@@ -352,6 +455,8 @@ Choose the adapter and model in the agent's `index.json`:
352
455
 
353
456
  ```json
354
457
  {
458
+ "name": "Gemini Agent",
459
+ "description": "Handle requests using the Gemini Developer API.",
355
460
  "provider": "googlegenai",
356
461
  "model": "gemini-2.5-flash",
357
462
  "outputType": "text"
@@ -401,6 +506,7 @@ Sessions are managed through the same registered agent:
401
506
  ```js
402
507
  await orcha.exampleAgent.get(sessionId);
403
508
  await orcha.exampleAgent.history(sessionId, { page: 1, pageSize: 50 });
509
+ await orcha.exampleAgent.events(sessionId, { page: 1, pageSize: 100 });
404
510
  await orcha.exampleAgent.list({
405
511
  metadata: { customerId: "cus_123" },
406
512
  page: 1,
@@ -412,6 +518,13 @@ await orcha.exampleAgent.update(sessionId, {
412
518
  });
413
519
  ```
414
520
 
521
+ `history()` returns a user-facing projection. `events()` returns the canonical
522
+ durable event records for observability and debugging without exposing the
523
+ active storage adapter. Applications should never read session files directly.
524
+ When paging while a session is still changing, pass the first response's
525
+ `throughSequence` into later `events()` calls to keep every page on the same
526
+ event boundary.
527
+
415
528
  ## Agent tests
416
529
 
417
530
  Agent tests live beside the agent and use the same compiled instructions,
@@ -507,6 +620,8 @@ export default defineEvaluations({
507
620
  ```json
508
621
  {
509
622
  "name": "response_quality",
623
+ "description": "Grounding of agent responses.",
624
+ "instructions": "Judge the complete response using only confirmed evidence recorded in the session.",
510
625
  "enabled": true,
511
626
  "provider": "openai",
512
627
  "model": "gpt-5-mini",
@@ -527,6 +642,11 @@ Evaluation errors and missed thresholds do not change a successful production
527
642
  agent result, but they do fail an agent test. Evaluator usage is reported
528
643
  separately from the agent's usage.
529
644
 
645
+ `description` is a short human-facing summary. `instructions` contains the
646
+ potentially detailed prompt that directs the evaluator; when omitted,
647
+ `description` is used for backward compatibility. Metric descriptions define
648
+ the individual scoring criteria.
649
+
530
650
  Evaluations do not delay the agent result. Await them only where the caller
531
651
  needs the scores:
532
652
 
@@ -1,2 +1,2 @@
1
- export declare const AGENTS_MD = "# Orcha project guide\n\nThis repository uses OrchaJS, a filesystem-convention framework for durable\nAI agents. Treat the `orcha/` directory as source code. Do not edit generated\nfiles under `.orcha/`.\n\n## Commands\n\n- `orcha init` creates the initial Orcha files without overwriting files.\n- `orcha dev` validates the registry and watches `orcha/**` for changes.\n- `orcha run <agent> --input \"\u2026\"` executes one registered agent.\n- `orcha run <agent> --input-file request.json` accepts structured input.\n- `orcha run <agent> --session <id> --input \"\u2026\"` continues a session.\n- `orcha run <agent> --session <id> --tool-results results.json` submits\n pending client-action results.\n- `orcha test` runs every registered agent test.\n- `orcha test <agent>` or `orcha test <agent>/<case>` narrows the run.\n- `orcha build` creates the production Orcha bundle without calling models.\n- Add `--json` to `run` and `test` for machine-readable output.\n\n`run` and `test` use real providers and require credentials. `dev` and\n`build` are offline. Session logs are JSONL files under\n`.orcha/sessions/<sessionId>.jsonl`.\n\n## Registry\n\n`orcha/index.ts` initializes providers and explicitly registers agents:\n\n```ts\nimport { orcha } from \"orchajs\";\n\norcha.init({\n providers: {\n anthropic: process.env.ANTHROPIC_API_KEY ?? \"\",\n openai: process.env.OPENAI_API_KEY ?? \"\",\n },\n actions: { runtime: \"sandbox\" },\n agents: {\n supportBot: \"./supportBot\",\n },\n});\n```\n\nOnly registered folders are compiled. Agent keys become runtime properties\nsuch as `orcha.supportBot`. Use `actions.runtime: \"sandbox\"` for isolated\nlocal action execution or `\"native\"` when the application intentionally\nallows action modules to execute in its Node.js process.\n\n`orcha.init()` fields:\n\n- `providers` (required): provider configurations keyed by built-in provider\n name.\n- `agents` (required): runtime property names mapped to folders relative to\n `orcha/`. At least one agent is required.\n- `actions` (required only when a registered agent has local actions):\n selects the local execution runtime and its environment/sandbox settings.\n- `storage.strategy` (optional): currently only `\"node-jsonl\"`.\n- `storage.directory` (optional): session directory relative to project\n root; defaults to `.orcha/sessions`.\n- `root` (optional): absolute or working-directory-relative project root;\n defaults to `ORCHA_PROJECT_ROOT` and then `process.cwd()`.\n\nProvider configuration shapes:\n\n```ts\nproviders: {\n anthropic: process.env.ANTHROPIC_API_KEY ?? \"\",\n deepseek: process.env.DEEPSEEK_API_KEY ?? \"\",\n googlegenai: process.env.GOOGLE_API_KEY ?? \"\",\n openai: {\n apiKey: process.env.OPENAI_API_KEY ?? \"\",\n baseUrl: \"https://api.openai.com/v1\", // optional override\n },\n vertexai: {\n project: process.env.GOOGLE_CLOUD_PROJECT ?? \"\",\n location: process.env.GOOGLE_CLOUD_LOCATION ?? \"us-central1\",\n // credentials is optional; omit it to use Google ADC.\n credentials: {\n clientEmail: process.env.GOOGLE_CLIENT_EMAIL ?? \"\",\n privateKey: process.env.GOOGLE_PRIVATE_KEY ?? \"\",\n },\n baseUrl: undefined, // optional override\n },\n bedrock: {\n region: process.env.AWS_REGION ?? \"us-east-1\",\n // credentials is optional; omit it to use the AWS credential chain.\n credentials: {\n accessKeyId: process.env.AWS_ACCESS_KEY_ID ?? \"\",\n secretAccessKey: process.env.AWS_SECRET_ACCESS_KEY ?? \"\",\n sessionToken: process.env.AWS_SESSION_TOKEN,\n },\n baseUrl: undefined, // optional override\n },\n}\n```\n\nAPI-key providers accept either a string shorthand or\n`{ apiKey, baseUrl? }`. Vertex AI requires `project` and `location`;\nexplicit service-account credentials are optional. Bedrock requires `region`;\nexplicit AWS credentials are optional. Never place credentials in\n`index.json`, instructions, tests, session metadata, or committed files.\n\n## Agent folders\n\n```text\norcha/\n index.ts\n supportBot/\n index.json\n instructions.md\n actions/\n skills/\n tests/\n evaluations/\n```\n\n`index.json` selects the model:\n\n```json\n{\n \"provider\": \"anthropic\",\n \"model\": \"claude-sonnet-4-6\",\n \"region\": \"provider_managed\",\n \"maxTokens\": 10240,\n \"outputType\": \"text\"\n}\n```\n\nOptional fields include `reasoningLevel`, `outputType: \"json\"`, and an\n`outputSchema` JSON Schema. Provider-specific reasoning values are forwarded\nwithout translation. Put the agent's stable role, boundaries, and operating\ninstructions in `instructions.md`.\n\nAgent `index.json` fields:\n\n- `provider` (required): `\"anthropic\"`, `\"bedrock\"`, `\"deepseek\"`,\n `\"openai\"`, `\"googlegenai\"`, or `\"vertexai\"`.\n- `model` (required): exact provider model identifier.\n- `region` (optional): provider/model routing hint; defaults in durable\n metadata to `\"provider_managed\"`.\n- `maxTokens` (optional): positive integer. If omitted, the provider adapter\n chooses its default.\n- `reasoningLevel` (optional): non-empty provider-native string. Orcha does\n not translate values between providers.\n- `outputType` (optional): `\"text\"` (default) or `\"json\"`. Image and\n audio are reserved but not implemented.\n- `outputSchema` (required for JSON output): JSON Schema used for provider\n structured output and final validation.\n\n`instructions.md` is required and cannot be empty. At compile time it becomes\nthe base system prompt. Orcha appends the compact available-skill catalog and\nthe full instructions for skills already loaded in this durable session.\n\n## Core execution model\n\nAn **agent** is the compiled definition: model settings, instructions, actions,\nskills, and evaluations. An agent can create many independent sessions.\n\nA **session** is one durable conversation owned by one agent. It has one\n`sessionId`, optional name and metadata, fixed prompt variables, and one\nappend-only JSONL timeline. Completing one response does not close the\nsession\u2014the application can resume it later. A session cannot be transferred\nto another registered agent, but later runs may use a different provider or\nmodel if that same agent's configuration changes.\n\nA **run** is one attempt to advance a session. `run()` creates a session and\nits first run. A conversational `resume(sessionId, { content })` creates the\nnext numbered run in that session. Each run accumulates its own model usage and\nends in exactly one of these states:\n\n- `completed`: the model produced final output.\n- `waiting_for_client_action`: the model requested work that only the\n application can perform. The run is paused, not completed.\n- `failed`: validation, provider, storage, or execution failed. The durable\n events remain available for diagnosis.\n\nAn **execution** is the in-process handle returned by one call to `run()` or\n`resume()`. It exposes a cumulative output stream, latest snapshot, final\nresult promise, and evaluation promise. An execution ends when that invocation\ncompletes, pauses, or fails; the durable session may continue through another\nexecution.\n\nA **model round** is one provider request inside a run. One run may contain\nseveral rounds:\n\n```text\nuser input\n \u2192 model round\n \u2192 tool calls\n \u2192 tool results\n \u2192 another model round\n \u2192 final answer\n```\n\nLocal actions and skill loads are handled automatically inside the same\nexecution. Their results are sent back to the model and the model loop\ncontinues without application involvement.\n\nA **client action** deliberately crosses the application boundary. Orcha can\ndescribe the tool to the model but cannot execute it because the operation\nbelongs to a browser, mobile app, approval system, or other caller-owned\nenvironment. The complete pause/continue flow is:\n\n```text\n1. Application calls agent.run(...) or agent.resume(...content).\n2. Model requests one or more client actions.\n3. Orcha stores client_action.requested and run.paused.\n4. execution.result resolves with:\n {\n status: \"waiting_for_client_action\",\n sessionId,\n clientToolCalls: [{ callId, name, arguments }]\n }\n5. Application executes every requested action.\n6. Application calls agent.resume(sessionId, {\n toolResults: [{ callId, output, isError? }]\n }).\n7. Orcha validates every callId and output, stores the results, and continues\n the same paused run from its prior model context.\n8. The resumed execution either completes, requests more client actions, or\n fails.\n```\n\nEvery pending call must be resolved exactly once in one resume operation.\n`callId` links the submitted result to the model's request; the action name\nmust not be substituted for it. Re-submitting the identical resolved result is\nidempotent and returns the prior completed result. Submitting different data\nfor an already-resolved call fails with `action_result_conflict`.\n\n`clientCapabilities` is supplied per invocation because different callers\nmay support different client actions. Orcha exposes only declared client\nactions to that model round. Local actions are always available when compiled.\n\nOnly one execution may mutate a session at a time. Concurrent calls for the\nsame `sessionId` return `session_busy`; different sessions can run\nindependently.\n\n## Running and resuming\n\n```ts\nconst execution = orcha.supportBot.run({\n content: \"Check subscription sub_123.\",\n name: \"Subscription check\",\n metadata: { accountId: \"acct_123\" },\n clientCapabilities: [\"request_human_approval\"],\n});\n\nfor await (const snapshot of execution.stream) {\n console.log(snapshot);\n}\n\nconst result = await execution.result;\nconst evaluations = await execution.evaluations;\n```\n\n`run()` input fields:\n\n- `content` (required): a non-empty string or array of text/file blocks.\n File blocks contain `type`, `mimeType`, and `fileUri`; unsupported\n provider/content combinations fail explicitly.\n- `name` (optional): trimmed session label from 1 through 200 characters.\n- `metadata` (optional): at most 50 fields with non-empty keys and finite\n string, number, boolean, or null values. Metadata is durable and available\n to local action context; never place secrets in it.\n- `variables` (optional): at most 50 string values whose keys are JavaScript\n identifiers. They replace `{{ variableName }}` placeholders in\n `instructions.md`, are fixed when the session is created, and are reused\n by later resumes. A missing referenced variable fails the run.\n- `clientCapabilities` (optional): action names the current caller can\n execute. Client actions not declared here are withheld from the model.\n\n`run()` always creates a new durable session. Continue one with:\n\n```ts\nconst execution = orcha.supportBot.resume(sessionId, {\n content: \"Continue with the confirmed account.\",\n});\n```\n\nIf a result has `status: \"waiting_for_client_action\"`, execute the requested\nclient actions in the application and submit every result:\n\n```ts\norcha.supportBot.resume(sessionId, {\n toolResults: [\n { callId: \"call_123\", output: { approved: true } }\n ],\n});\n```\n\nNever invent call IDs. Use the IDs returned in `clientToolCalls`.\n\n## Actions\n\nEach action has metadata and, for local actions, executable code:\n\n```text\nactions/\n lookupAccount/\n index.json\n index.js\n```\n\n```json\n{\n \"name\": \"lookup_account\",\n \"description\": \"Look up one account.\",\n \"execution\": \"local\",\n \"parameters\": {\n \"type\": \"object\",\n \"properties\": {\n \"accountId\": { \"type\": \"string\" }\n },\n \"required\": [\"accountId\"],\n \"additionalProperties\": false\n },\n \"outputSchema\": {\n \"type\": \"object\",\n \"properties\": {\n \"status\": { \"type\": \"string\" }\n },\n \"required\": [\"status\"],\n \"additionalProperties\": false\n }\n}\n```\n\n```js\nexport default async function lookupAccount({ accountId }) {\n return { status: \"active\" };\n}\n```\n\nClient actions use `\"execution\": \"client\"` and do not include executable\ncode. Orcha pauses until the caller submits their results. Keep action names,\ndescriptions, schemas, and implementations aligned.\n\nAction `index.json` fields:\n\n- `name` (required): model-facing tool name, 1\u201364 letters, numbers,\n underscores, or hyphens. `load_skill` is reserved.\n- `description` (required): tells the model when and why to call the action.\n- `execution` (required): `\"local\"` executes `index.js`; `\"client\"`\n pauses the run and delegates execution to the application.\n- `parameters` (required): JSON Schema for model-generated arguments.\n- `outputSchema` (optional): JSON Schema validated against local or submitted\n client output before the model receives it.\n- `timeoutMs` (optional): integer from 1 through 120000; defaults to 10000.\n- `permissions.env` (optional): names copied from `orcha.init().actions.env`\n into the action context.\n- `permissions.network` (optional): exact hosts or wildcard subdomains such\n as `\"api.example.com\"` or `\"*.example.com\"` allowed through\n `context.fetch`. Redirects are rejected.\n- `sideEffect` (optional): descriptive metadata for whether the operation\n mutates external state. It does not currently change execution behavior.\n\n`orcha.init().actions.runtime` and an action's `execution` solve different\nproblems:\n\n- `execution: \"client\"`: Orcha never executes code for this action.\n- `execution: \"local\"` + `runtime: \"sandbox\"`: compiled code runs in a\n QuickJS isolate with JSON-only inputs/outputs, default 32 MB memory, default\n 512 KB stack, interruptible timeout, declared environment values, and\n allowlisted network access through the provided context.\n- `execution: \"local\"` + `runtime: \"native\"`: code runs in the host Node.js\n process. It can use host privileges directly. The timeout rejects slow\n asynchronous work but cannot interrupt synchronous blocking code.\n\nGlobal local-action configuration:\n\n```ts\nactions: {\n runtime: \"sandbox\", // required when any registered action is local\n env: {\n BILLING_API_TOKEN: process.env.BILLING_API_TOKEN,\n },\n sandbox: {\n memoryLimitMb: 32,\n stackLimitKb: 512,\n },\n}\n```\n\nThe local action signature is\n`(parameters, context) => output | Promise<output>`. Context contains\n`sessionId`, immutable session `metadata`, a stable `idempotencyKey`,\nallowlisted `env`, guarded `fetch`, and prefixed `log`.\n\n## Skills\n\nSkills are lazy-loaded procedural instructions. Register only intended skills:\n\n```js\n// skills/index.js\nimport { defineSkills } from \"orchajs/skills\";\n\nexport default defineSkills({\n incidentTriage: \"./incidentTriage\",\n});\n```\n\nEach skill folder contains `index.json` metadata and `instructions.md`.\nThe model receives a compact catalog and can call the internal `load_skill`\ntool. Loaded instructions remain active for the durable session. Lifecycle\nevents are `skill.requested`, `skill.loaded`, and `skill.failed`.\n\nSkill `index.json` fields:\n\n- `name` (required): model-facing name, 1\u201364 letters, numbers, underscores,\n or hyphens; unique within the agent.\n- `description` (required): compact catalog description shown before loading.\n- `triggers` (optional): non-empty array of non-empty situations describing\n when the model should load the skill.\n\n`instructions.md` is required and cannot be empty. The key in\n`skills/index.js` is only a registration label; `index.json.name` is the\nname used by the model and durable events. Unregistered folders are ignored.\n\n## Tests\n\nRegister tests in `tests/index.js`:\n\n```js\nimport { defineTests } from \"orchajs/testing\";\n\nexport default defineTests({\n activeAccount: \"./activeAccount\",\n});\n```\n\nEach case's `index.json` defines `input`, mocked responses for every action,\nand `expect`. Tests run the real compiled agent and provider but never execute\nreal actions. The mocked action set must exactly match the compiled action set.\n\n```json\n{\n \"input\": { \"content\": \"Check account acct_123.\" },\n \"actions\": {\n \"lookup_account\": {\n \"responses\": [\n { \"output\": { \"status\": \"active\" } }\n ]\n }\n },\n \"expect\": {\n \"status\": \"completed\",\n \"text\": { \"contains\": [\"active\"] },\n \"actions\": [\n {\n \"name\": \"lookup_account\",\n \"arguments\": { \"equals\": { \"accountId\": \"acct_123\" } }\n }\n ]\n }\n}\n```\n\nTest sessions use the `ses_test_` prefix and end with a `test.completed`\nevent. Prefer semantic output assertions; verify exact identifiers and values\nthrough action-argument assertions.\n\nTest `index.json` fields:\n\n- `description` (optional): human-readable purpose.\n- `input.content` (required): string or multimodal content array.\n- `input.variables` (optional): string map available to the session.\n- `input.metadata` (optional): string, number, boolean, or null values.\n- `actions` (required): exactly one key for every compiled action, including\n local actions. Every `responses` array is consumed in call order.\n- `responses[].output` (required): mocked action result.\n- `responses[].isError` (optional): marks the mocked result as an error.\n- `expect.status` (optional): `\"completed\"` or `\"failed\"`; defaults to\n `\"completed\"`.\n- `expect.output.equals` / `partial` (optional): exact or recursive partial\n comparison against structured output.\n- `expect.text.contains` / `excludes` (optional): case-sensitive semantic\n text checks.\n- `expect.actions` (optional): ordered expected calls. Each may assert\n `arguments.equals` or `arguments.partial`.\n\nThe registration key in `tests/index.js` is the test selector used by\n`orcha test agent/testName`; its value resolves to the case folder.\n\n## Evaluations\n\nEvaluations are asynchronous LLM judges registered in\n`evaluations/index.js` with `defineEvaluations` from\n`orchajs/evaluations`. Each folder's `index.json` defines its provider,\nmodel, metrics, and thresholds:\n\n```json\n{\n \"name\": \"response_quality\",\n \"enabled\": true,\n \"provider\": \"openai\",\n \"model\": \"gpt-5-mini\",\n \"metrics\": [\n {\n \"name\": \"groundedness\",\n \"description\": \"The answer relies on confirmed session evidence.\",\n \"threshold\": 0.8\n }\n ]\n}\n```\n\n`execution.result` does not wait for judges. Await\n`execution.evaluations` when results must finish before process exit.\nEvaluations always finish during `orcha test`; judge errors and missed\nthresholds fail the test. Lifecycle events are `evaluation.requested`,\n`evaluation.completed`, and `evaluation.failed`.\n\nEvaluation `index.json` fields:\n\n- `name` (required): durable model-facing identifier, 1\u201364 letters, numbers,\n underscores, or hyphens; unique within the agent.\n- `description` (optional): overall judging objective.\n- `enabled` (optional): defaults to `true`. Disabled evaluations are\n compiled but do not run.\n- `provider` and `model` (required): independently select the judge. The\n provider must also exist in `orcha.init().providers`.\n- `maxTokens` (optional): positive integer; defaults to 2000 for judges.\n- `reasoningLevel` (optional): non-empty provider-native string forwarded\n without translation.\n- `metrics` (required): non-empty array with unique metric names.\n- `metrics[].name`: 1\u201364 letters, numbers, underscores, or hyphens.\n- `metrics[].description`: exact criterion supplied to the judge.\n- `metrics[].threshold`: inclusive number from 0 to 1. A metric passes when\n the returned score is greater than or equal to this threshold.\n\nThe judge sees a sanitized transcript of user/assistant messages, action\nrequests and outcomes, client-action activity, and loaded skill names. It does\nnot receive internal reasoning blocks, replay metadata, previous evaluation\nresults, or test assertions. It must return exactly one score, reasoning\nstring, and non-empty evidence array for every configured metric.\n\n## Sessions and logs\n\nJSONL is the durable source of truth. Each line is one complete JSON object;\nnever treat the file as one JSON array. Events are append-only and ordered by\n`sequence`.\n\nAll events use this envelope:\n\n```ts\ntype SessionEvent<T> = {\n sequence: number; // starts at 1 and increases across the whole session\n type: SessionEventType;\n timestamp: string; // ISO-8601 UTC timestamp\n run?: number; // present for run-scoped events\n data: T;\n};\n```\n\nSession-scoped events omit `run`. Optional properties whose values are\n`undefined` are omitted from serialized JSON.\n\nShared stored structures:\n\n```ts\ntype Usage = {\n inputTokens: number;\n outputTokens: number;\n reasoningTokens: number | null;\n cacheReadTokens: number;\n cacheWriteTokens: number;\n};\n\ntype ErrorData = {\n code: string;\n message: string;\n retryable?: boolean;\n};\n\ntype UserContent =\n | { type: \"text\"; text: string }\n | {\n type: \"image\" | \"video\" | \"audio\" | \"url\";\n mimeType: string;\n fileUri: string;\n };\n\ntype AssistantContent =\n | { type: \"text\"; text: string }\n | {\n type: \"reasoning\";\n text: string;\n replay?: { providerId?: string; opaqueData?: string };\n }\n | {\n type: \"tool_call\";\n callId: string;\n name: string;\n arguments: Record<string, unknown>;\n replay?: { providerId?: string; opaqueData?: string };\n };\n\ntype ToolResult = {\n callId: string;\n output: unknown;\n isError?: boolean;\n};\n```\n\nExact event payloads:\n\n```ts\ntype SessionCreated = SessionEvent<{\n schemaVersion: 1;\n sessionId: string;\n agent: string;\n status: \"active\";\n name?: string;\n metadata: Record<string, string | number | boolean | null>;\n variables: Record<string, string>;\n}>; // type \"session.created\", no run\n\ntype SessionUpdated = SessionEvent<{\n name?: string;\n metadata?: Record<string, string | number | boolean | null>;\n}>; // type \"session.updated\", no run\n\ntype RunStarted = SessionEvent<{\n status: \"running\";\n agent: string;\n provider: string;\n model: string;\n region: string;\n reasoningLevel?: string;\n outputType: \"text\" | \"json\";\n clientCapabilities: string[];\n}>; // type \"run.started\"\n\ntype UserMessageCreated = SessionEvent<{\n role: \"user\";\n content: UserContent[];\n}>; // type \"message.created\"\n\ntype AssistantMessageCreated = SessionEvent<{\n status: \"completed\" | \"incomplete\";\n provider: string;\n model: string;\n responseId?: string;\n stopReason?: \"end_turn\" | \"tool_call\" | \"max_tokens\" |\n \"content_filter\" | \"unknown\";\n role: \"assistant\";\n content: AssistantContent[];\n parsedOutput?: unknown; // final JSON output only\n usage: Usage;\n durationMs: number;\n}>; // type \"message.created\"\n\ntype ToolMessageCreated = SessionEvent<{\n role: \"tool\";\n content: ToolResult[];\n}>; // type \"message.created\"\n\ntype ActionRequested = SessionEvent<{\n callId: string;\n name: string;\n arguments: Record<string, unknown>;\n sourceHash?: string;\n idempotencyKey: string; // sessionId:callId\n}>; // type \"action.requested\"\n\ntype ActionCompleted = SessionEvent<{\n callId: string;\n name: string;\n output: unknown;\n sourceHash?: string;\n durationMs: number;\n}>; // type \"action.completed\"\n\ntype ActionFailed = SessionEvent<{\n callId: string;\n name: string;\n sourceHash?: string;\n durationMs: number;\n error: {\n code: \"action_execution_failed\";\n message: string;\n };\n}>; // type \"action.failed\"\n\ntype ClientActionRequested = SessionEvent<{\n status: \"waiting\";\n calls: Array<{\n callId: string;\n name: string;\n arguments: Record<string, unknown>;\n }>;\n localResults: ToolResult[];\n toolCallOrder: string[];\n}>; // type \"client_action.requested\"\n\ntype ClientActionResolved = SessionEvent<{\n status: \"completed\";\n results: ToolResult[];\n}>; // type \"client_action.resolved\"\n\ntype SkillRequested = SessionEvent<{\n callId: string;\n name: unknown;\n}>; // type \"skill.requested\"\n\ntype SkillLoaded = SessionEvent<{\n callId: string;\n name: string;\n alreadyLoaded: boolean;\n}>; // type \"skill.loaded\", no run\n\ntype SkillFailed = SessionEvent<{\n callId: string;\n name: unknown;\n error: {\n code: \"skill_not_found\";\n message: string;\n };\n}>; // type \"skill.failed\"\n\ntype EvaluationRequested = SessionEvent<{\n name: string;\n provider: string;\n model: string;\n evaluatedThroughSequence: number;\n}>; // type \"evaluation.requested\"\n\ntype EvaluationMetric = {\n name: string;\n score: number;\n threshold: number;\n passed: boolean;\n reasoning: string;\n evidence: string[];\n};\n\ntype EvaluationCompleted = SessionEvent<{\n name: string;\n status: \"passed\" | \"failed\";\n metrics: EvaluationMetric[];\n usage?: Usage;\n durationMs: number;\n provider: string;\n model: string;\n evaluatedThroughSequence: number;\n}>; // type \"evaluation.completed\"\n\ntype EvaluationFailed = SessionEvent<{\n name: string;\n status: \"error\";\n metrics: [];\n durationMs: number;\n error: { message: string };\n provider: string;\n model: string;\n evaluatedThroughSequence: number;\n}>; // type \"evaluation.failed\"\n\ntype RunPaused = SessionEvent<{\n status: \"waiting_for_client_action\";\n clientToolCalls: Array<{\n callId: string;\n name: string;\n arguments: Record<string, unknown>;\n }>;\n usage: Usage;\n}>; // type \"run.paused\"\n\ntype RunCompleted = SessionEvent<{\n status: \"completed\";\n durationMs: number;\n usage: Usage;\n}>; // type \"run.completed\"\n\ntype RunFailed = SessionEvent<{\n status: \"failed\";\n durationMs: number;\n usage?: Usage;\n error: ErrorData;\n}>; // type \"run.failed\"\n\ntype TestCompleted = SessionEvent<{\n suiteId: string;\n agent: string;\n test: string;\n status: \"passed\" | \"failed\";\n durationMs: number;\n assertions: Array<{\n path: string;\n passed: boolean;\n message: string;\n expected?: unknown;\n actual?: unknown;\n }>;\n usage?: Usage;\n evaluations?: Array<{\n name: string;\n status: \"passed\" | \"failed\" | \"error\";\n metrics: EvaluationMetric[];\n usage?: Usage;\n durationMs: number;\n error?: { message: string };\n }>;\n error?: ErrorData;\n}>; // type \"test.completed\", no run\n```\n\nTypical event order:\n\n```text\nsession.created\nrun.started\nmessage.created (user)\nmessage.created (assistant, possibly with tool_call)\naction.requested \u2192 action.completed|action.failed # local action\nmessage.created (tool)\n...additional model/action rounds...\nmessage.created (assistant final)\nrun.completed\nevaluation.requested\nevaluation.completed|evaluation.failed\n```\n\nFor client actions, `client_action.requested` and `run.paused` replace the\nimmediate tool message. A later `resume(...toolResults)` appends\n`client_action.resolved`, the tool message, and continues the same run\nnumber. A conversational `resume(...content)` starts a new run number.\n\n## How Orcha works behind the scenes\n\n### Compilation\n\n1. `orcha/index.ts` calls `orcha.init()` with explicit agent paths.\n2. The compiler reads each registered agent's `index.json` and\n `instructions.md`.\n3. Every directory under `actions/` is compiled. Skills, tests, and\n evaluations are included only through their local `index.js` registry.\n4. Local action source is bundled and SHA-256 hashed. Production bundles keep\n only provider adapters required by agents and enabled evaluations.\n5. Invalid paths, duplicate model-facing names, missing files, unsupported\n configuration values, and missing schema objects fail before execution.\n Concrete action arguments and outputs are validated against their schemas\n when the action is used.\n\n`orcha dev` repeats validation when files change. `orcha build` performs\noffline production compilation. Neither command invokes a provider.\n\n### Run lifecycle\n\n1. `run()` creates a `ses_<uuid>`, acquires the per-session execution lock,\n appends `session.created`, then starts run 1.\n2. `resume()` reads and validates the existing session. Message continuation\n starts a new run; submitted client results continue the paused run.\n3. The provider receives the base instructions, available-skill catalog,\n loaded skill instructions, normalized conversation messages, action\n schemas, model settings, and current client capabilities.\n4. Provider-specific responses are normalized into text, reasoning, and tool\n call blocks. Opaque replay metadata is stored only when a provider needs it\n to replay its own prior block correctly.\n5. Tool calls are checked against compiled actions and declared client\n capabilities. Arguments and outputs are validated against JSON Schema.\n6. `load_skill` updates durable session instructions. Local actions execute\n through the configured runtime. Client actions pause safely. Tool results\n are reordered to match the model's original call order.\n7. The model loop continues until final output, failure, a client pause, or\n the maximum of 10 action rounds.\n8. Usage is normalized and aggregated across every model call in the run.\n9. After `run.completed`, enabled evaluations start in the background.\n `execution.result` is already available; `execution.evaluations` waits\n for judge completion and durable persistence.\n\nThe per-session lock prevents two model executions from mutating one session\nat once. Background evaluation writes queue behind active runs so they cannot\ncause `resume()` to fail spuriously or reuse sequence numbers.\n\n### Replay and context\n\nOrcha does not send raw JSONL back to the model. It projects durable events\ninto provider-neutral conversation messages. Completed conversational runs\nbecome user, assistant, and tool messages; lifecycle bookkeeping such as\ndurations, test assertions, and evaluation events is excluded from model\ncontext. Reasoning text and provider replay metadata are retained where needed\nfor faithful continuation but are omitted from evaluation transcripts.\n\n### Reading sessions\n\n- `agent.get(sessionId)` projects the latest status, pending client actions,\n last output, metadata, and aggregate usage.\n- `agent.history(sessionId, { page, pageSize })` returns a safe user-facing\n timeline rather than raw provider bookkeeping.\n- `agent.list({ page, pageSize, status, metadata })` lists projected session\n snapshots.\n- Read JSONL directly when building observability, audit, or debugging tools\n that require the exact append-only event stream described above.\n\n## Change rules\n\n- Register every new agent, skill, test, and evaluation explicitly.\n- Keep runtime behavior provider-neutral.\n- Do not call real actions from tests.\n- Do not commit `.orcha/`; it contains generated output and session data.\n- Run `orcha dev` after filesystem changes and `orcha test` when behavior\n changes.\n- Do not weaken assertions to match incorrect behavior. Remove an assertion\n only when it is stricter than the documented agent contract.\n";
1
+ export declare const AGENTS_MD = "# Orcha project guide\n\nThis repository uses OrchaJS, a filesystem-convention framework for durable\nAI agents. Treat the `orcha/` directory as source code. Do not edit generated\nfiles under `.orcha/`.\n\n## Commands\n\n- `orcha init` creates the initial Orcha files without overwriting files.\n- `orcha dev` validates the registry and watches `orcha/**` for changes.\n- `orcha run <agent> --input \"\u2026\"` executes one registered agent.\n- `orcha run <agent> --input-file request.json` accepts structured input.\n- `orcha run <agent> --session <id> --input \"\u2026\"` continues a session.\n- `orcha run <agent> --session <id> --tool-results results.json` submits\n pending client-action results.\n- `orcha test` runs every registered agent test.\n- `orcha test <agent>` or `orcha test <agent>/<case>` narrows the run.\n- `orcha build` creates the production Orcha bundle without calling models.\n- Add `--json` to `run` and `test` for machine-readable output.\n\n`run` and `test` use real providers and require credentials. `dev` and\n`build` are offline. Session logs are JSONL files under\n`.orcha/sessions/<sessionId>.jsonl`.\n\n## Registry\n\n`orcha/index.ts` initializes providers and explicitly registers agents:\n\n```ts\nimport { orcha } from \"orchajs\";\n\norcha.init({\n providers: {\n anthropic: process.env.ANTHROPIC_API_KEY ?? \"\",\n openai: process.env.OPENAI_API_KEY ?? \"\",\n },\n actions: { runtime: \"sandbox\" },\n agents: {\n supportBot: \"./supportBot\",\n },\n});\n```\n\nOnly registered folders are compiled. Agent keys become runtime properties\nsuch as `orcha.supportBot`. Use `actions.runtime: \"sandbox\"` for isolated\nlocal action execution or `\"native\"` when the application intentionally\nallows action modules to execute in its Node.js process.\n\n`orcha.init()` fields:\n\n- `providers` (required): provider configurations keyed by built-in provider\n name.\n- `agents` (required): runtime property names mapped to folders relative to\n `orcha/`. At least one agent is required.\n- `actions` (required only when a registered agent has local actions):\n selects the local execution runtime and its environment/sandbox settings.\n- `storage.strategy` (optional): currently only `\"node-jsonl\"`.\n- `storage.directory` (optional): session directory relative to project\n root; defaults to `.orcha/sessions`.\n- `root` (optional): absolute or working-directory-relative project root;\n defaults to `ORCHA_PROJECT_ROOT` and then `process.cwd()`.\n\nProvider configuration shapes:\n\n```ts\nproviders: {\n anthropic: process.env.ANTHROPIC_API_KEY ?? \"\",\n deepseek: process.env.DEEPSEEK_API_KEY ?? \"\",\n googlegenai: process.env.GOOGLE_API_KEY ?? \"\",\n openai: {\n apiKey: process.env.OPENAI_API_KEY ?? \"\",\n baseUrl: \"https://api.openai.com/v1\", // optional override\n },\n vertexai: {\n project: process.env.GOOGLE_CLOUD_PROJECT ?? \"\",\n location: process.env.GOOGLE_CLOUD_LOCATION ?? \"us-central1\",\n // credentials is optional; omit it to use Google ADC.\n credentials: {\n clientEmail: process.env.GOOGLE_CLIENT_EMAIL ?? \"\",\n privateKey: process.env.GOOGLE_PRIVATE_KEY ?? \"\",\n },\n baseUrl: undefined, // optional override\n },\n bedrock: {\n region: process.env.AWS_REGION ?? \"us-east-1\",\n // credentials is optional; omit it to use the AWS credential chain.\n credentials: {\n accessKeyId: process.env.AWS_ACCESS_KEY_ID ?? \"\",\n secretAccessKey: process.env.AWS_SECRET_ACCESS_KEY ?? \"\",\n sessionToken: process.env.AWS_SESSION_TOKEN,\n },\n baseUrl: undefined, // optional override\n },\n}\n```\n\nAPI-key providers accept either a string shorthand or\n`{ apiKey, baseUrl? }`. Vertex AI requires `project` and `location`;\nexplicit service-account credentials are optional. Bedrock requires `region`;\nexplicit AWS credentials are optional. Never place credentials in\n`index.json`, instructions, tests, session metadata, or committed files.\n\n## Agent folders\n\n```text\norcha/\n index.ts\n supportBot/\n index.json\n instructions.md\n actions/\n skills/\n tests/\n evaluations/\n```\n\n`index.json` selects the model:\n\n```json\n{\n \"name\": \"Support Agent\",\n \"description\": \"Resolve customer support questions using confirmed account data.\",\n \"provider\": \"anthropic\",\n \"model\": \"claude-sonnet-4-6\",\n \"region\": \"provider_managed\",\n \"maxTokens\": 10240,\n \"outputType\": \"text\"\n}\n```\n\nOptional fields include `reasoningLevel`, `outputType: \"json\"`, and an\n`outputSchema` JSON Schema. Provider-specific reasoning values are forwarded\nwithout translation. Put the agent's stable role, boundaries, and operating\ninstructions in `instructions.md`.\n\nAgent `index.json` fields:\n\n- `name` (required): concise human-readable agent name.\n- `description` (optional): what the agent does and when it should be used.\n- `provider` (required): `\"anthropic\"`, `\"bedrock\"`, `\"deepseek\"`,\n `\"openai\"`, `\"googlegenai\"`, or `\"vertexai\"`.\n- `model` (required): exact provider model identifier.\n- `region` (optional): provider/model routing hint; defaults in durable\n metadata to `\"provider_managed\"`.\n- `maxTokens` (optional): positive integer. If omitted, the provider adapter\n chooses its default.\n- `reasoningLevel` (optional): non-empty provider-native string. Orcha does\n not translate values between providers.\n- `outputType` (optional): `\"text\"` (default) or `\"json\"`. Image and\n audio are reserved but not implemented.\n- `outputSchema` (required for JSON output): JSON Schema used for provider\n structured output and final validation.\n\n`instructions.md` is required and cannot be empty. At compile time it becomes\nthe base system prompt. Orcha appends the compact available-skill catalog and\nthe full instructions for skills already loaded in this durable session.\n\n## Subagents\n\nRegister private subagents alongside a parent in `orcha/index.ts`:\n\n```js\nagents: {\n coordinator: {\n path: \"./coordinator\",\n subagents: {\n researcher: \"./researcher\",\n },\n },\n researcher: \"./researcher\",\n}\n```\n\nOnly top-level keys become `orcha.<agentName>`. In this example the\nresearcher is both directly accessible and available to the coordinator.\nRemove its top-level entry to make it private.\n\nA parent may configure delegation limits in its `index.json`:\n\n```json\n{\n \"subagents\": {\n \"maxPerRun\": 3\n }\n}\n```\n\n`maxPerRun` limits newly created child sessions in one parent run and\ndefaults to 10.\n\nThe parent receives a compact catalog containing each subagent's registered\nname and optional description. Internal `run_agent`, `resume_agent`,\n`inspect_agent` tools let it start, continue, and inspect only child sessions\ninitiated by its current session. A delegated agent cannot delegate again.\n\nSubagent calls are synchronous. The parent becomes\n`waiting_for_subagent` while the child runs, then receives the child's text,\npause state, client-action request, failure, or completion as a normal tool\nresult. The parent and child keep separate linked JSONL sessions.\n\n## Core execution model\n\nAn **agent** is the compiled definition: model settings, instructions, actions,\nskills, and evaluations. An agent can create many independent sessions.\n\nA **session** is one durable conversation owned by one agent. It has one\n`sessionId`, optional name and metadata, fixed prompt variables, and one\nappend-only JSONL timeline. Completing one response does not close the\nsession\u2014the application can resume it later. A session cannot be transferred\nto another registered agent, but later runs may use a different provider or\nmodel if that same agent's configuration changes.\n\nA **run** is one attempt to advance a session. `run()` creates a session and\nits first run. A conversational `resume(sessionId, { content })` creates the\nnext numbered run in that session. Each run accumulates its own model usage and\nends in exactly one of these states:\n\n- `completed`: the model produced final output.\n- `waiting_for_subagent`: the parent is waiting for a synchronous child\n response and continues automatically when it arrives.\n- `waiting_for_client_action`: the model requested work that only the\n application can perform. The run is paused, not completed.\n- `paused`: active provider work was aborted, streamed output was preserved,\n and a later `resume()` starts a new run.\n- `failed`: validation, provider, storage, or execution failed. The durable\n events remain available for diagnosis.\n\nAn **execution** is the in-process handle returned by one call to `run()` or\n`resume()`. It exposes a cumulative output stream, latest snapshot, final\nresult promise, and evaluation promise. An execution ends when that invocation\ncompletes, pauses, or fails; the durable session may continue through another\nexecution.\n\nA **model round** is one provider request inside a run. One run may contain\nseveral rounds:\n\n```text\nuser input\n \u2192 model round\n \u2192 tool calls\n \u2192 tool results\n \u2192 another model round\n \u2192 final answer\n```\n\nLocal actions and skill loads are handled automatically inside the same\nexecution. Their results are sent back to the model and the model loop\ncontinues without application involvement.\n\nA **client action** deliberately crosses the application boundary. Orcha can\ndescribe the tool to the model but cannot execute it because the operation\nbelongs to a browser, mobile app, approval system, or other caller-owned\nenvironment. The complete pause/continue flow is:\n\n```text\n1. Application calls agent.run(...) or agent.resume(...content).\n2. Model requests one or more client actions.\n3. Orcha stores client_action.requested and run.paused.\n4. execution.result resolves with:\n {\n status: \"waiting_for_client_action\",\n sessionId,\n clientToolCalls: [{ callId, name, arguments }]\n }\n5. Application executes every requested action.\n6. Application calls agent.resume(sessionId, {\n toolResults: [{ callId, output, isError? }]\n }).\n7. Orcha validates every callId and output, stores the results, and continues\n the same paused run from its prior model context.\n8. The resumed execution either completes, requests more client actions, or\n fails.\n```\n\nEvery pending call must be resolved exactly once in one resume operation.\n`callId` links the submitted result to the model's request; the action name\nmust not be substituted for it. Re-submitting the identical resolved result is\nidempotent and returns the prior completed result. Submitting different data\nfor an already-resolved call fails with `action_result_conflict`.\n\n`clientCapabilities` is supplied per invocation because different callers\nmay support different client actions. Orcha exposes only declared client\nactions to that model round. Local actions are always available when compiled.\n\nOnly one execution may mutate a session at a time. Concurrent calls for the\nsame `sessionId` return `session_busy`; different sessions can run\nindependently.\n\n## Running and resuming\n\n```ts\nconst execution = orcha.supportBot.run({\n content: \"Check subscription sub_123.\",\n name: \"Subscription check\",\n metadata: { accountId: \"acct_123\" },\n clientCapabilities: [\"request_human_approval\"],\n});\n\nfor await (const snapshot of execution.stream) {\n console.log(snapshot);\n}\n\nconst result = await execution.result;\nconst evaluations = await execution.evaluations;\n```\n\n`run()` input fields:\n\n- `content` (required): a non-empty string or array of text/file blocks.\n File blocks contain `type`, `mimeType`, and `fileUri`; unsupported\n provider/content combinations fail explicitly.\n- `name` (optional): trimmed session label from 1 through 200 characters.\n- `metadata` (optional): at most 50 fields with non-empty keys and finite\n string, number, boolean, or null values. Metadata is durable and available\n to local action context; never place secrets in it.\n- `variables` (optional): at most 50 string values whose keys are JavaScript\n identifiers. They replace `{{ variableName }}` placeholders in\n `instructions.md`, are fixed when the session is created, and are reused\n by later resumes. A missing referenced variable fails the run.\n- `clientCapabilities` (optional): action names the current caller can\n execute. Client actions not declared here are withheld from the model.\n\n`run()` always creates a new durable session. Continue one with:\n\n```ts\nconst execution = orcha.supportBot.resume(sessionId, {\n content: \"Continue with the confirmed account.\",\n});\n```\n\nIf a result has `status: \"waiting_for_client_action\"`, execute the requested\nclient actions in the application and submit every result:\n\n```ts\norcha.supportBot.resume(sessionId, {\n toolResults: [\n { callId: \"call_123\", output: { approved: true } }\n ],\n});\n```\n\nNever invent call IDs. Use the IDs returned in `clientToolCalls`.\n\n## Actions\n\nEach action has metadata and, for local actions, executable code:\n\n```text\nactions/\n lookupAccount/\n index.json\n index.js\n```\n\n```json\n{\n \"name\": \"lookup_account\",\n \"description\": \"Look up one account.\",\n \"execution\": \"local\",\n \"parameters\": {\n \"type\": \"object\",\n \"properties\": {\n \"accountId\": { \"type\": \"string\" }\n },\n \"required\": [\"accountId\"],\n \"additionalProperties\": false\n },\n \"outputSchema\": {\n \"type\": \"object\",\n \"properties\": {\n \"status\": { \"type\": \"string\" }\n },\n \"required\": [\"status\"],\n \"additionalProperties\": false\n }\n}\n```\n\n```js\nexport default async function lookupAccount({ accountId }) {\n return { status: \"active\" };\n}\n```\n\nClient actions use `\"execution\": \"client\"` and do not include executable\ncode. Orcha pauses until the caller submits their results. Keep action names,\ndescriptions, schemas, and implementations aligned.\n\nAction `index.json` fields:\n\n- `name` (required): model-facing tool name, 1\u201364 letters, numbers,\n underscores, or hyphens. `load_skill` is reserved.\n- `description` (required): tells the model when and why to call the action.\n- `execution` (required): `\"local\"` executes `index.js`; `\"client\"`\n pauses the run and delegates execution to the application.\n- `parameters` (required): JSON Schema for model-generated arguments.\n- `outputSchema` (optional): JSON Schema validated against local or submitted\n client output before the model receives it.\n- `timeoutMs` (optional): integer from 1 through 120000; defaults to 10000.\n- `permissions.env` (optional): names copied from `orcha.init().actions.env`\n into the action context.\n- `permissions.network` (optional): exact hosts or wildcard subdomains such\n as `\"api.example.com\"` or `\"*.example.com\"` allowed through\n `context.fetch`. Redirects are rejected.\n- `sideEffect` (optional): descriptive metadata for whether the operation\n mutates external state. It does not currently change execution behavior.\n\n`orcha.init().actions.runtime` and an action's `execution` solve different\nproblems:\n\n- `execution: \"client\"`: Orcha never executes code for this action.\n- `execution: \"local\"` + `runtime: \"sandbox\"`: compiled code runs in a\n QuickJS isolate with JSON-only inputs/outputs, default 32 MB memory, default\n 512 KB stack, interruptible timeout, declared environment values, and\n allowlisted network access through the provided context.\n- `execution: \"local\"` + `runtime: \"native\"`: code runs in the host Node.js\n process. It can use host privileges directly. The timeout rejects slow\n asynchronous work but cannot interrupt synchronous blocking code.\n\nGlobal local-action configuration:\n\n```ts\nactions: {\n runtime: \"sandbox\", // required when any registered action is local\n env: {\n BILLING_API_TOKEN: process.env.BILLING_API_TOKEN,\n },\n sandbox: {\n memoryLimitMb: 32,\n stackLimitKb: 512,\n },\n}\n```\n\nThe local action signature is\n`(parameters, context) => output | Promise<output>`. Context contains\n`sessionId`, immutable session `metadata`, a stable `idempotencyKey`,\nallowlisted `env`, guarded `fetch`, and prefixed `log`.\n\n## Skills\n\nSkills are lazy-loaded procedural instructions. Register only intended skills:\n\n```js\n// skills/index.js\nimport { defineSkills } from \"orchajs/skills\";\n\nexport default defineSkills({\n incidentTriage: \"./incidentTriage\",\n});\n```\n\nEach skill folder contains `index.json` metadata and `instructions.md`.\nThe model receives a compact catalog and can call the internal `load_skill`\ntool. Loaded instructions remain active for the durable session. Lifecycle\nevents are `skill.requested`, `skill.loaded`, and `skill.failed`.\n\nSkill `index.json` fields:\n\n- `name` (required): model-facing name, 1\u201364 letters, numbers, underscores,\n or hyphens; unique within the agent.\n- `description` (required): compact catalog description shown before loading.\n- `triggers` (optional): non-empty array of non-empty situations describing\n when the model should load the skill.\n\n`instructions.md` is required and cannot be empty. The key in\n`skills/index.js` is only a registration label; `index.json.name` is the\nname used by the model and durable events. Unregistered folders are ignored.\n\n## Tests\n\nRegister tests in `tests/index.js`:\n\n```js\nimport { defineTests } from \"orchajs/testing\";\n\nexport default defineTests({\n activeAccount: \"./activeAccount\",\n});\n```\n\nEach case's `index.json` defines `input`, mocked responses for every action,\nand `expect`. Tests run the real compiled agent and provider but never execute\nreal actions. The mocked action set must exactly match the compiled action set.\n\n```json\n{\n \"input\": { \"content\": \"Check account acct_123.\" },\n \"actions\": {\n \"lookup_account\": {\n \"responses\": [\n { \"output\": { \"status\": \"active\" } }\n ]\n }\n },\n \"expect\": {\n \"status\": \"completed\",\n \"text\": { \"contains\": [\"active\"] },\n \"actions\": [\n {\n \"name\": \"lookup_account\",\n \"arguments\": { \"equals\": { \"accountId\": \"acct_123\" } }\n }\n ]\n }\n}\n```\n\nTest sessions use the `ses_test_` prefix and end with a `test.completed`\nevent. Prefer semantic output assertions; verify exact identifiers and values\nthrough action-argument assertions.\n\nTest `index.json` fields:\n\n- `description` (optional): human-readable purpose.\n- `input.content` (required): string or multimodal content array.\n- `input.variables` (optional): string map available to the session.\n- `input.metadata` (optional): string, number, boolean, or null values.\n- `actions` (required): exactly one key for every compiled action, including\n local actions. Every `responses` array is consumed in call order.\n- `responses[].output` (required): mocked action result.\n- `responses[].isError` (optional): marks the mocked result as an error.\n- `expect.status` (optional): `\"completed\"` or `\"failed\"`; defaults to\n `\"completed\"`.\n- `expect.output.equals` / `partial` (optional): exact or recursive partial\n comparison against structured output.\n- `expect.text.contains` / `excludes` (optional): case-sensitive semantic\n text checks.\n- `expect.actions` (optional): ordered expected calls. Each may assert\n `arguments.equals` or `arguments.partial`.\n\nThe registration key in `tests/index.js` is the test selector used by\n`orcha test agent/testName`; its value resolves to the case folder.\n\n## Evaluations\n\nEvaluations are asynchronous LLM judges registered in\n`evaluations/index.js` with `defineEvaluations` from\n`orchajs/evaluations`. Each folder's `index.json` defines its provider,\nmodel, metrics, and thresholds:\n\n```json\n{\n \"name\": \"response_quality\",\n \"description\": \"Grounding of agent responses.\",\n \"instructions\": \"Judge the complete response using only confirmed evidence recorded in the session.\",\n \"enabled\": true,\n \"provider\": \"openai\",\n \"model\": \"gpt-5-mini\",\n \"metrics\": [\n {\n \"name\": \"groundedness\",\n \"description\": \"The answer relies on confirmed session evidence.\",\n \"threshold\": 0.8\n }\n ]\n}\n```\n\n`execution.result` does not wait for judges. Await\n`execution.evaluations` when results must finish before process exit.\nEvaluations always finish during `orcha test`; judge errors and missed\nthresholds fail the test. Lifecycle events are `evaluation.requested`,\n`evaluation.completed`, and `evaluation.failed`.\n\nEvaluation `index.json` fields:\n\n- `name` (required): durable model-facing identifier, 1\u201364 letters, numbers,\n underscores, or hyphens; unique within the agent.\n- `description` (optional): short human-facing summary of the evaluator.\n- `instructions` (optional): detailed prompt supplied to the judge. When\n omitted, `description` is used for backward compatibility.\n- `enabled` (optional): defaults to `true`. Disabled evaluations are\n compiled but do not run.\n- `provider` and `model` (required): independently select the judge. The\n provider must also exist in `orcha.init().providers`.\n- `maxTokens` (optional): positive integer; defaults to 2000 for judges.\n- `reasoningLevel` (optional): non-empty provider-native string forwarded\n without translation.\n- `metrics` (required): non-empty array with unique metric names.\n- `metrics[].name`: 1\u201364 letters, numbers, underscores, or hyphens.\n- `metrics[].description`: exact criterion supplied to the judge.\n- `metrics[].threshold`: inclusive number from 0 to 1. A metric passes when\n the returned score is greater than or equal to this threshold.\n\nThe judge sees a sanitized transcript of user/assistant messages, action\nrequests and outcomes, client-action activity, and loaded skill names. It does\nnot receive internal reasoning blocks, replay metadata, previous evaluation\nresults, or test assertions. It must return exactly one score, reasoning\nstring, and non-empty evidence array for every configured metric.\n\n## Sessions and logs\n\nJSONL is the durable source of truth. Each line is one complete JSON object;\nnever treat the file as one JSON array. Events are append-only and ordered by\n`sequence`.\n\nAll events use this envelope:\n\n```ts\ntype SessionEvent<T> = {\n sequence: number; // starts at 1 and increases across the whole session\n type: SessionEventType;\n timestamp: string; // ISO-8601 UTC timestamp\n run?: number; // present for run-scoped events\n data: T;\n};\n```\n\nSession-scoped events omit `run`. Optional properties whose values are\n`undefined` are omitted from serialized JSON.\n\nShared stored structures:\n\n```ts\ntype Usage = {\n inputTokens: number;\n outputTokens: number;\n reasoningTokens: number | null;\n cacheReadTokens: number;\n cacheWriteTokens: number;\n};\n\ntype ErrorData = {\n code: string;\n message: string;\n retryable?: boolean;\n};\n\ntype UserContent =\n | { type: \"text\"; text: string }\n | {\n type: \"image\" | \"video\" | \"audio\" | \"url\";\n mimeType: string;\n fileUri: string;\n };\n\ntype AssistantContent =\n | { type: \"text\"; text: string }\n | {\n type: \"reasoning\";\n text: string;\n replay?: { providerId?: string; opaqueData?: string };\n }\n | {\n type: \"tool_call\";\n callId: string;\n name: string;\n arguments: Record<string, unknown>;\n replay?: { providerId?: string; opaqueData?: string };\n };\n\ntype ToolResult = {\n callId: string;\n output: unknown;\n isError?: boolean;\n};\n```\n\nExact event payloads:\n\n```ts\ntype SessionCreated = SessionEvent<{\n schemaVersion: 1;\n sessionId: string;\n agent: string;\n status: \"active\";\n name?: string;\n metadata: Record<string, string | number | boolean | null>;\n variables: Record<string, string>;\n lineage?: {\n origin: \"delegated\";\n parentAgent: string;\n parentSessionId: string;\n parentCallId: string;\n };\n}>; // type \"session.created\", no run\n\ntype SessionUpdated = SessionEvent<{\n name?: string;\n metadata?: Record<string, string | number | boolean | null>;\n}>; // type \"session.updated\", no run\n\ntype RunStarted = SessionEvent<{\n status: \"running\";\n agent: string;\n provider: string;\n model: string;\n region: string;\n reasoningLevel?: string;\n outputType: \"text\" | \"json\";\n clientCapabilities: string[];\n}>; // type \"run.started\"\n\ntype UserMessageCreated = SessionEvent<{\n role: \"user\";\n content: UserContent[];\n}>; // type \"message.created\"\n\ntype AssistantMessageCreated = SessionEvent<{\n status: \"completed\" | \"incomplete\";\n provider: string;\n model: string;\n responseId?: string;\n stopReason?: \"end_turn\" | \"tool_call\" | \"max_tokens\" |\n \"content_filter\" | \"unknown\";\n role: \"assistant\";\n content: AssistantContent[];\n parsedOutput?: unknown; // final JSON output only\n usage: Usage;\n durationMs: number;\n}>; // type \"message.created\"\n\ntype ToolMessageCreated = SessionEvent<{\n role: \"tool\";\n content: ToolResult[];\n}>; // type \"message.created\"\n\ntype ActionRequested = SessionEvent<{\n callId: string;\n name: string;\n arguments: Record<string, unknown>;\n sourceHash?: string;\n idempotencyKey: string; // sessionId:callId\n}>; // type \"action.requested\"\n\ntype ActionCompleted = SessionEvent<{\n callId: string;\n name: string;\n output: unknown;\n sourceHash?: string;\n durationMs: number;\n}>; // type \"action.completed\"\n\ntype ActionFailed = SessionEvent<{\n callId: string;\n name: string;\n sourceHash?: string;\n durationMs: number;\n error: {\n code: \"action_execution_failed\";\n message: string;\n };\n}>; // type \"action.failed\"\n\ntype ClientActionRequested = SessionEvent<{\n status: \"waiting\";\n calls: Array<{\n callId: string;\n name: string;\n arguments: Record<string, unknown>;\n }>;\n localResults: ToolResult[];\n toolCallOrder: string[];\n}>; // type \"client_action.requested\"\n\ntype ClientActionResolved = SessionEvent<{\n status: \"completed\";\n results: ToolResult[];\n}>; // type \"client_action.resolved\"\n\ntype SkillRequested = SessionEvent<{\n callId: string;\n name: unknown;\n}>; // type \"skill.requested\"\n\ntype SkillLoaded = SessionEvent<{\n callId: string;\n name: string;\n alreadyLoaded: boolean;\n}>; // type \"skill.loaded\", no run\n\ntype SkillFailed = SessionEvent<{\n callId: string;\n name: unknown;\n error: {\n code: \"skill_not_found\";\n message: string;\n };\n}>; // type \"skill.failed\"\n\ntype EvaluationRequested = SessionEvent<{\n name: string;\n provider: string;\n model: string;\n evaluatedThroughSequence: number;\n}>; // type \"evaluation.requested\"\n\ntype EvaluationMetric = {\n name: string;\n score: number;\n threshold: number;\n passed: boolean;\n reasoning: string;\n evidence: string[];\n};\n\ntype EvaluationCompleted = SessionEvent<{\n name: string;\n status: \"passed\" | \"failed\";\n metrics: EvaluationMetric[];\n usage?: Usage;\n durationMs: number;\n provider: string;\n model: string;\n evaluatedThroughSequence: number;\n}>; // type \"evaluation.completed\"\n\ntype EvaluationFailed = SessionEvent<{\n name: string;\n status: \"error\";\n metrics: [];\n durationMs: number;\n error: { message: string };\n provider: string;\n model: string;\n evaluatedThroughSequence: number;\n}>; // type \"evaluation.failed\"\n\ntype SubagentInitiated = SessionEvent<{\n callId: string;\n agent: string;\n sessionId: string;\n status: \"running\";\n}>; // type \"subagent.initiated\"\n\ntype SubagentResumed = SessionEvent<{\n callId: string;\n agent: string;\n sessionId: string;\n status: \"running\";\n}>; // type \"subagent.resumed\"\n\ntype SubagentCompletedOrPausedOrFailed = SessionEvent<{\n callId: string;\n agent: string;\n sessionId: string;\n status:\n | \"completed\"\n | \"waiting_for_client_action\"\n | \"paused\"\n | \"failed\";\n output?: unknown;\n usage?: Usage;\n clientToolCalls?: Array<{\n callId: string;\n name: string;\n arguments: Record<string, unknown>;\n }>;\n error?: ErrorData;\n}>; // type \"subagent.completed\" | \"subagent.paused\" | \"subagent.failed\"\n\ntype RunPaused = SessionEvent<{\n status: \"waiting_for_client_action\" | \"paused\";\n clientToolCalls?: Array<{\n callId: string;\n name: string;\n arguments: Record<string, unknown>;\n }>;\n usage?: Usage;\n output?: unknown;\n durationMs?: number;\n}>; // type \"run.paused\"\n\ntype RunCompleted = SessionEvent<{\n status: \"completed\";\n durationMs: number;\n usage: Usage;\n}>; // type \"run.completed\"\n\ntype RunFailed = SessionEvent<{\n status: \"failed\";\n durationMs: number;\n usage?: Usage;\n error: ErrorData;\n}>; // type \"run.failed\"\n\ntype SessionPaused = SessionEvent<{\n status: \"paused\";\n}>; // type \"session.paused\", no run\n\ntype SessionResumed = SessionEvent<{\n status: \"active\";\n}>; // type \"session.resumed\", no run\n\ntype TestCompleted = SessionEvent<{\n suiteId: string;\n agent: string;\n test: string;\n status: \"passed\" | \"failed\";\n durationMs: number;\n assertions: Array<{\n path: string;\n passed: boolean;\n message: string;\n expected?: unknown;\n actual?: unknown;\n }>;\n usage?: Usage;\n evaluations?: Array<{\n name: string;\n status: \"passed\" | \"failed\" | \"error\";\n metrics: EvaluationMetric[];\n usage?: Usage;\n durationMs: number;\n error?: { message: string };\n }>;\n error?: ErrorData;\n}>; // type \"test.completed\", no run\n```\n\nTypical event order:\n\n```text\nsession.created\nrun.started\nmessage.created (user)\nmessage.created (assistant, possibly with tool_call)\naction.requested \u2192 action.completed|action.failed # local action\nsubagent.initiated\n...child session advances independently...\nsubagent.paused\nsubagent.resumed\n...child session advances independently...\nsubagent.completed|subagent.paused|subagent.failed\nmessage.created (tool)\n...additional model/action rounds...\nmessage.created (assistant final)\nrun.completed\nevaluation.requested\nevaluation.completed|evaluation.failed\n```\n\nFor client actions, `client_action.requested` and `run.paused` replace the\nimmediate tool message. A later `resume(...toolResults)` appends\n`client_action.resolved`, the tool message, and continues the same run\nnumber. A conversational `resume(...content)` starts a new run number.\n\n## How Orcha works behind the scenes\n\n### Compilation\n\n1. `orcha/index.ts` calls `orcha.init()` with explicit agent paths.\n2. The compiler reads each registered agent's `index.json` and\n `instructions.md`.\n3. Every directory under `actions/` is compiled. Skills, tests, and\n evaluations are included only through their local `index.js` registry.\n4. Local action source is bundled and SHA-256 hashed. Production bundles keep\n only provider adapters required by agents and enabled evaluations.\n5. Invalid paths, duplicate model-facing names, missing files, unsupported\n configuration values, and missing schema objects fail before execution.\n Concrete action arguments and outputs are validated against their schemas\n when the action is used.\n\n`orcha dev` repeats validation when files change. `orcha build` performs\noffline production compilation. Neither command invokes a provider.\n\n### Run lifecycle\n\n1. `run()` creates a `ses_<uuid>`, acquires the per-session execution lock,\n appends `session.created`, then starts run 1.\n2. `resume()` reads and validates the existing session. Message continuation\n starts a new run; submitted client results continue the paused run.\n3. The provider receives the base instructions, available-skill catalog,\n loaded skill instructions, normalized conversation messages, action\n schemas, model settings, and current client capabilities.\n4. Provider-specific responses are normalized into text, reasoning, and tool\n call blocks. Opaque replay metadata is stored only when a provider needs it\n to replay its own prior block correctly.\n5. Tool calls are checked against compiled actions and declared client\n capabilities. Arguments and outputs are validated against JSON Schema.\n6. `load_skill` updates durable session instructions. Local actions execute\n through the configured runtime. Client actions pause safely. Tool results\n are reordered to match the model's original call order.\n7. Internal agent tools start or continue linked child sessions. The parent\n reports `waiting_for_subagent` until each synchronous child call returns.\n8. The model loop continues until final output, failure, a pause, or the\n maximum of 10 action rounds.\n9. Usage is normalized and aggregated across every model call in the run.\n10. After `run.completed`, enabled evaluations start in the background.\n `execution.result` is already available; `execution.evaluations` waits\n for judge completion and durable persistence.\n\nThe per-session lock prevents two model executions from mutating one session\nat once. Background evaluation writes queue behind active runs so they cannot\ncause `resume()` to fail spuriously or reuse sequence numbers.\n\n### Replay and context\n\nOrcha does not send raw JSONL back to the model. It projects durable events\ninto provider-neutral conversation messages. Completed and explicitly paused\nconversational runs become user, assistant, and tool messages, so a new run\nsees incomplete assistant text preserved by `pause()`. Lifecycle bookkeeping\nsuch as durations, test assertions, and evaluation events is excluded from\nmodel context. Reasoning text and provider replay metadata are retained where\nneeded for faithful continuation but are omitted from evaluation transcripts.\n\n### Reading sessions\n\n- `agent.get(sessionId)` projects the latest status, pending client actions,\n last output, metadata, and aggregate usage.\n- `agent.history(sessionId, { page, pageSize })` returns a safe user-facing\n timeline rather than raw provider bookkeeping.\n- `agent.events(sessionId, { page, pageSize })` returns the canonical durable\n event records for observability, audit, and debugging tools. For subsequent\n pages of a changing session, pass the first response's `throughSequence`\n back in the options to keep pagination on a stable event boundary.\n- `agent.list({ page, pageSize, status, metadata })` lists projected session\n snapshots.\n- `agent.subagentHistory(childSessionId, { page, pageSize })` returns the\n projected history of a child owned by this parent agent. Lineage checks\n prevent access through unrelated agents.\n- `agent.pause(sessionId)` aborts active provider work for the session and\n active children, preserving streamed text as an incomplete assistant\n message. `agent.resume(sessionId)` starts a new run with `\"Continue.\"`;\n pass content to give different instructions. Pending client actions still\n require their exact tool results.\n- Never read storage files directly. Use these methods so applications remain\n compatible with JSONL, SQLite, IndexedDB, remote, and future storage adapters.\n\n## Change rules\n\n- Register every new agent, skill, test, and evaluation explicitly.\n- Keep runtime behavior provider-neutral.\n- Do not call real actions from tests.\n- Do not commit `.orcha/`; it contains generated output and session data.\n- Run `orcha dev` after filesystem changes and `orcha test` when behavior\n changes.\n- Do not weaken assertions to match incorrect behavior. Remove an assertion\n only when it is stricter than the documented agent contract.\n";
2
2
  //# sourceMappingURL=cli-agents-template.d.ts.map
@@ -1 +1 @@
1
- {"version":3,"file":"cli-agents-template.d.ts","sourceRoot":"","sources":["../src/cli-agents-template.ts"],"names":[],"mappings":"AAAA,eAAO,MAAM,SAAS,4s9BAy4BrB,CAAC"}
1
+ {"version":3,"file":"cli-agents-template.d.ts","sourceRoot":"","sources":["../src/cli-agents-template.ts"],"names":[],"mappings":"AAAA,eAAO,MAAM,SAAS,0qmCAsgCrB,CAAC"}
@@ -118,6 +118,8 @@ orcha/
118
118
 
119
119
  \`\`\`json
120
120
  {
121
+ "name": "Support Agent",
122
+ "description": "Resolve customer support questions using confirmed account data.",
121
123
  "provider": "anthropic",
122
124
  "model": "claude-sonnet-4-6",
123
125
  "region": "provider_managed",
@@ -133,6 +135,8 @@ instructions in \`instructions.md\`.
133
135
 
134
136
  Agent \`index.json\` fields:
135
137
 
138
+ - \`name\` (required): concise human-readable agent name.
139
+ - \`description\` (optional): what the agent does and when it should be used.
136
140
  - \`provider\` (required): \`"anthropic"\`, \`"bedrock"\`, \`"deepseek"\`,
137
141
  \`"openai"\`, \`"googlegenai"\`, or \`"vertexai"\`.
138
142
  - \`model\` (required): exact provider model identifier.
@@ -151,6 +155,49 @@ Agent \`index.json\` fields:
151
155
  the base system prompt. Orcha appends the compact available-skill catalog and
152
156
  the full instructions for skills already loaded in this durable session.
153
157
 
158
+ ## Subagents
159
+
160
+ Register private subagents alongside a parent in \`orcha/index.ts\`:
161
+
162
+ \`\`\`js
163
+ agents: {
164
+ coordinator: {
165
+ path: "./coordinator",
166
+ subagents: {
167
+ researcher: "./researcher",
168
+ },
169
+ },
170
+ researcher: "./researcher",
171
+ }
172
+ \`\`\`
173
+
174
+ Only top-level keys become \`orcha.<agentName>\`. In this example the
175
+ researcher is both directly accessible and available to the coordinator.
176
+ Remove its top-level entry to make it private.
177
+
178
+ A parent may configure delegation limits in its \`index.json\`:
179
+
180
+ \`\`\`json
181
+ {
182
+ "subagents": {
183
+ "maxPerRun": 3
184
+ }
185
+ }
186
+ \`\`\`
187
+
188
+ \`maxPerRun\` limits newly created child sessions in one parent run and
189
+ defaults to 10.
190
+
191
+ The parent receives a compact catalog containing each subagent's registered
192
+ name and optional description. Internal \`run_agent\`, \`resume_agent\`,
193
+ \`inspect_agent\` tools let it start, continue, and inspect only child sessions
194
+ initiated by its current session. A delegated agent cannot delegate again.
195
+
196
+ Subagent calls are synchronous. The parent becomes
197
+ \`waiting_for_subagent\` while the child runs, then receives the child's text,
198
+ pause state, client-action request, failure, or completion as a normal tool
199
+ result. The parent and child keep separate linked JSONL sessions.
200
+
154
201
  ## Core execution model
155
202
 
156
203
  An **agent** is the compiled definition: model settings, instructions, actions,
@@ -169,8 +216,12 @@ next numbered run in that session. Each run accumulates its own model usage and
169
216
  ends in exactly one of these states:
170
217
 
171
218
  - \`completed\`: the model produced final output.
219
+ - \`waiting_for_subagent\`: the parent is waiting for a synchronous child
220
+ response and continues automatically when it arrives.
172
221
  - \`waiting_for_client_action\`: the model requested work that only the
173
222
  application can perform. The run is paused, not completed.
223
+ - \`paused\`: active provider work was aborted, streamed output was preserved,
224
+ and a later \`resume()\` starts a new run.
174
225
  - \`failed\`: validation, provider, storage, or execution failed. The durable
175
226
  events remain available for diagnosis.
176
227
 
@@ -491,6 +542,8 @@ model, metrics, and thresholds:
491
542
  \`\`\`json
492
543
  {
493
544
  "name": "response_quality",
545
+ "description": "Grounding of agent responses.",
546
+ "instructions": "Judge the complete response using only confirmed evidence recorded in the session.",
494
547
  "enabled": true,
495
548
  "provider": "openai",
496
549
  "model": "gpt-5-mini",
@@ -514,7 +567,9 @@ Evaluation \`index.json\` fields:
514
567
 
515
568
  - \`name\` (required): durable model-facing identifier, 1–64 letters, numbers,
516
569
  underscores, or hyphens; unique within the agent.
517
- - \`description\` (optional): overall judging objective.
570
+ - \`description\` (optional): short human-facing summary of the evaluator.
571
+ - \`instructions\` (optional): detailed prompt supplied to the judge. When
572
+ omitted, \`description\` is used for backward compatibility.
518
573
  - \`enabled\` (optional): defaults to \`true\`. Disabled evaluations are
519
574
  compiled but do not run.
520
575
  - \`provider\` and \`model\` (required): independently select the judge. The
@@ -613,6 +668,12 @@ type SessionCreated = SessionEvent<{
613
668
  name?: string;
614
669
  metadata: Record<string, string | number | boolean | null>;
615
670
  variables: Record<string, string>;
671
+ lineage?: {
672
+ origin: "delegated";
673
+ parentAgent: string;
674
+ parentSessionId: string;
675
+ parentCallId: string;
676
+ };
616
677
  }>; // type "session.created", no run
617
678
 
618
679
  type SessionUpdated = SessionEvent<{
@@ -756,14 +817,49 @@ type EvaluationFailed = SessionEvent<{
756
817
  evaluatedThroughSequence: number;
757
818
  }>; // type "evaluation.failed"
758
819
 
820
+ type SubagentInitiated = SessionEvent<{
821
+ callId: string;
822
+ agent: string;
823
+ sessionId: string;
824
+ status: "running";
825
+ }>; // type "subagent.initiated"
826
+
827
+ type SubagentResumed = SessionEvent<{
828
+ callId: string;
829
+ agent: string;
830
+ sessionId: string;
831
+ status: "running";
832
+ }>; // type "subagent.resumed"
833
+
834
+ type SubagentCompletedOrPausedOrFailed = SessionEvent<{
835
+ callId: string;
836
+ agent: string;
837
+ sessionId: string;
838
+ status:
839
+ | "completed"
840
+ | "waiting_for_client_action"
841
+ | "paused"
842
+ | "failed";
843
+ output?: unknown;
844
+ usage?: Usage;
845
+ clientToolCalls?: Array<{
846
+ callId: string;
847
+ name: string;
848
+ arguments: Record<string, unknown>;
849
+ }>;
850
+ error?: ErrorData;
851
+ }>; // type "subagent.completed" | "subagent.paused" | "subagent.failed"
852
+
759
853
  type RunPaused = SessionEvent<{
760
- status: "waiting_for_client_action";
761
- clientToolCalls: Array<{
854
+ status: "waiting_for_client_action" | "paused";
855
+ clientToolCalls?: Array<{
762
856
  callId: string;
763
857
  name: string;
764
858
  arguments: Record<string, unknown>;
765
859
  }>;
766
- usage: Usage;
860
+ usage?: Usage;
861
+ output?: unknown;
862
+ durationMs?: number;
767
863
  }>; // type "run.paused"
768
864
 
769
865
  type RunCompleted = SessionEvent<{
@@ -779,6 +875,14 @@ type RunFailed = SessionEvent<{
779
875
  error: ErrorData;
780
876
  }>; // type "run.failed"
781
877
 
878
+ type SessionPaused = SessionEvent<{
879
+ status: "paused";
880
+ }>; // type "session.paused", no run
881
+
882
+ type SessionResumed = SessionEvent<{
883
+ status: "active";
884
+ }>; // type "session.resumed", no run
885
+
782
886
  type TestCompleted = SessionEvent<{
783
887
  suiteId: string;
784
888
  agent: string;
@@ -813,6 +917,12 @@ run.started
813
917
  message.created (user)
814
918
  message.created (assistant, possibly with tool_call)
815
919
  action.requested → action.completed|action.failed # local action
920
+ subagent.initiated
921
+ ...child session advances independently...
922
+ subagent.paused
923
+ subagent.resumed
924
+ ...child session advances independently...
925
+ subagent.completed|subagent.paused|subagent.failed
816
926
  message.created (tool)
817
927
  ...additional model/action rounds...
818
928
  message.created (assistant final)
@@ -862,10 +972,12 @@ offline production compilation. Neither command invokes a provider.
862
972
  6. \`load_skill\` updates durable session instructions. Local actions execute
863
973
  through the configured runtime. Client actions pause safely. Tool results
864
974
  are reordered to match the model's original call order.
865
- 7. The model loop continues until final output, failure, a client pause, or
866
- the maximum of 10 action rounds.
867
- 8. Usage is normalized and aggregated across every model call in the run.
868
- 9. After \`run.completed\`, enabled evaluations start in the background.
975
+ 7. Internal agent tools start or continue linked child sessions. The parent
976
+ reports \`waiting_for_subagent\` until each synchronous child call returns.
977
+ 8. The model loop continues until final output, failure, a pause, or the
978
+ maximum of 10 action rounds.
979
+ 9. Usage is normalized and aggregated across every model call in the run.
980
+ 10. After \`run.completed\`, enabled evaluations start in the background.
869
981
  \`execution.result\` is already available; \`execution.evaluations\` waits
870
982
  for judge completion and durable persistence.
871
983
 
@@ -876,11 +988,12 @@ cause \`resume()\` to fail spuriously or reuse sequence numbers.
876
988
  ### Replay and context
877
989
 
878
990
  Orcha does not send raw JSONL back to the model. It projects durable events
879
- into provider-neutral conversation messages. Completed conversational runs
880
- become user, assistant, and tool messages; lifecycle bookkeeping such as
881
- durations, test assertions, and evaluation events is excluded from model
882
- context. Reasoning text and provider replay metadata are retained where needed
883
- for faithful continuation but are omitted from evaluation transcripts.
991
+ into provider-neutral conversation messages. Completed and explicitly paused
992
+ conversational runs become user, assistant, and tool messages, so a new run
993
+ sees incomplete assistant text preserved by \`pause()\`. Lifecycle bookkeeping
994
+ such as durations, test assertions, and evaluation events is excluded from
995
+ model context. Reasoning text and provider replay metadata are retained where
996
+ needed for faithful continuation but are omitted from evaluation transcripts.
884
997
 
885
998
  ### Reading sessions
886
999
 
@@ -888,10 +1001,22 @@ for faithful continuation but are omitted from evaluation transcripts.
888
1001
  last output, metadata, and aggregate usage.
889
1002
  - \`agent.history(sessionId, { page, pageSize })\` returns a safe user-facing
890
1003
  timeline rather than raw provider bookkeeping.
1004
+ - \`agent.events(sessionId, { page, pageSize })\` returns the canonical durable
1005
+ event records for observability, audit, and debugging tools. For subsequent
1006
+ pages of a changing session, pass the first response's \`throughSequence\`
1007
+ back in the options to keep pagination on a stable event boundary.
891
1008
  - \`agent.list({ page, pageSize, status, metadata })\` lists projected session
892
1009
  snapshots.
893
- - Read JSONL directly when building observability, audit, or debugging tools
894
- that require the exact append-only event stream described above.
1010
+ - \`agent.subagentHistory(childSessionId, { page, pageSize })\` returns the
1011
+ projected history of a child owned by this parent agent. Lineage checks
1012
+ prevent access through unrelated agents.
1013
+ - \`agent.pause(sessionId)\` aborts active provider work for the session and
1014
+ active children, preserving streamed text as an incomplete assistant
1015
+ message. \`agent.resume(sessionId)\` starts a new run with \`"Continue."\`;
1016
+ pass content to give different instructions. Pending client actions still
1017
+ require their exact tool results.
1018
+ - Never read storage files directly. Use these methods so applications remain
1019
+ compatible with JSONL, SQLite, IndexedDB, remote, and future storage adapters.
895
1020
 
896
1021
  ## Change rules
897
1022
 
@@ -1 +1 @@
1
- {"version":3,"file":"cli-agents-template.js","sourceRoot":"","sources":["../src/cli-agents-template.ts"],"names":[],"mappings":"AAAA,MAAM,CAAC,MAAM,SAAS,GAAG;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;CAy4BxB,CAAC"}
1
+ {"version":3,"file":"cli-agents-template.js","sourceRoot":"","sources":["../src/cli-agents-template.ts"],"names":[],"mappings":"AAAA,MAAM,CAAC,MAAM,SAAS,GAAG;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;CAsgCxB,CAAC"}
package/dist/cli.js CHANGED
@@ -53,6 +53,8 @@ async function initializeProject(projectRoot) {
53
53
  [
54
54
  "orcha/exampleAgent/index.json",
55
55
  `${JSON.stringify({
56
+ name: "Example Agent",
57
+ description: "Answer general questions clearly and concisely.",
56
58
  provider: "anthropic",
57
59
  model: "claude-sonnet-4-6",
58
60
  region: "provider_managed",