orchajs 0.3.0 → 0.5.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,915 @@
1
+ export const AGENTS_MD = `# Orcha project guide
2
+
3
+ This repository uses OrchaJS, a filesystem-convention framework for durable
4
+ AI agents. Treat the \`orcha/\` directory as source code. Do not edit generated
5
+ files under \`.orcha/\`.
6
+
7
+ ## Commands
8
+
9
+ - \`orcha init\` creates the initial Orcha files without overwriting files.
10
+ - \`orcha dev\` validates the registry and watches \`orcha/**\` for changes.
11
+ - \`orcha run <agent> --input "…"\` executes one registered agent.
12
+ - \`orcha run <agent> --input-file request.json\` accepts structured input.
13
+ - \`orcha run <agent> --session <id> --input "…"\` continues a session.
14
+ - \`orcha run <agent> --session <id> --tool-results results.json\` submits
15
+ pending client-action results.
16
+ - \`orcha test\` runs every registered agent test.
17
+ - \`orcha test <agent>\` or \`orcha test <agent>/<case>\` narrows the run.
18
+ - \`orcha build\` creates the production Orcha bundle without calling models.
19
+ - Add \`--json\` to \`run\` and \`test\` for machine-readable output.
20
+
21
+ \`run\` and \`test\` use real providers and require credentials. \`dev\` and
22
+ \`build\` are offline. Session logs are JSONL files under
23
+ \`.orcha/sessions/<sessionId>.jsonl\`.
24
+
25
+ ## Registry
26
+
27
+ \`orcha/index.ts\` initializes providers and explicitly registers agents:
28
+
29
+ \`\`\`ts
30
+ import { orcha } from "orchajs";
31
+
32
+ orcha.init({
33
+ providers: {
34
+ anthropic: process.env.ANTHROPIC_API_KEY ?? "",
35
+ openai: process.env.OPENAI_API_KEY ?? "",
36
+ },
37
+ actions: { runtime: "sandbox" },
38
+ agents: {
39
+ supportBot: "./supportBot",
40
+ },
41
+ });
42
+ \`\`\`
43
+
44
+ Only registered folders are compiled. Agent keys become runtime properties
45
+ such as \`orcha.supportBot\`. Use \`actions.runtime: "sandbox"\` for isolated
46
+ local action execution or \`"native"\` when the application intentionally
47
+ allows action modules to execute in its Node.js process.
48
+
49
+ \`orcha.init()\` fields:
50
+
51
+ - \`providers\` (required): provider configurations keyed by built-in provider
52
+ name.
53
+ - \`agents\` (required): runtime property names mapped to folders relative to
54
+ \`orcha/\`. At least one agent is required.
55
+ - \`actions\` (required only when a registered agent has local actions):
56
+ selects the local execution runtime and its environment/sandbox settings.
57
+ - \`storage.strategy\` (optional): currently only \`"node-jsonl"\`.
58
+ - \`storage.directory\` (optional): session directory relative to project
59
+ root; defaults to \`.orcha/sessions\`.
60
+ - \`root\` (optional): absolute or working-directory-relative project root;
61
+ defaults to \`ORCHA_PROJECT_ROOT\` and then \`process.cwd()\`.
62
+
63
+ Provider configuration shapes:
64
+
65
+ \`\`\`ts
66
+ providers: {
67
+ anthropic: process.env.ANTHROPIC_API_KEY ?? "",
68
+ deepseek: process.env.DEEPSEEK_API_KEY ?? "",
69
+ googlegenai: process.env.GOOGLE_API_KEY ?? "",
70
+ openai: {
71
+ apiKey: process.env.OPENAI_API_KEY ?? "",
72
+ baseUrl: "https://api.openai.com/v1", // optional override
73
+ },
74
+ vertexai: {
75
+ project: process.env.GOOGLE_CLOUD_PROJECT ?? "",
76
+ location: process.env.GOOGLE_CLOUD_LOCATION ?? "us-central1",
77
+ // credentials is optional; omit it to use Google ADC.
78
+ credentials: {
79
+ clientEmail: process.env.GOOGLE_CLIENT_EMAIL ?? "",
80
+ privateKey: process.env.GOOGLE_PRIVATE_KEY ?? "",
81
+ },
82
+ baseUrl: undefined, // optional override
83
+ },
84
+ bedrock: {
85
+ region: process.env.AWS_REGION ?? "us-east-1",
86
+ // credentials is optional; omit it to use the AWS credential chain.
87
+ credentials: {
88
+ accessKeyId: process.env.AWS_ACCESS_KEY_ID ?? "",
89
+ secretAccessKey: process.env.AWS_SECRET_ACCESS_KEY ?? "",
90
+ sessionToken: process.env.AWS_SESSION_TOKEN,
91
+ },
92
+ baseUrl: undefined, // optional override
93
+ },
94
+ }
95
+ \`\`\`
96
+
97
+ API-key providers accept either a string shorthand or
98
+ \`{ apiKey, baseUrl? }\`. Vertex AI requires \`project\` and \`location\`;
99
+ explicit service-account credentials are optional. Bedrock requires \`region\`;
100
+ explicit AWS credentials are optional. Never place credentials in
101
+ \`index.json\`, instructions, tests, session metadata, or committed files.
102
+
103
+ ## Agent folders
104
+
105
+ \`\`\`text
106
+ orcha/
107
+ index.ts
108
+ supportBot/
109
+ index.json
110
+ instructions.md
111
+ actions/
112
+ skills/
113
+ tests/
114
+ evaluations/
115
+ \`\`\`
116
+
117
+ \`index.json\` selects the model:
118
+
119
+ \`\`\`json
120
+ {
121
+ "provider": "anthropic",
122
+ "model": "claude-sonnet-4-6",
123
+ "region": "provider_managed",
124
+ "maxTokens": 10240,
125
+ "outputType": "text"
126
+ }
127
+ \`\`\`
128
+
129
+ Optional fields include \`reasoningLevel\`, \`outputType: "json"\`, and an
130
+ \`outputSchema\` JSON Schema. Provider-specific reasoning values are forwarded
131
+ without translation. Put the agent's stable role, boundaries, and operating
132
+ instructions in \`instructions.md\`.
133
+
134
+ Agent \`index.json\` fields:
135
+
136
+ - \`provider\` (required): \`"anthropic"\`, \`"bedrock"\`, \`"deepseek"\`,
137
+ \`"openai"\`, \`"googlegenai"\`, or \`"vertexai"\`.
138
+ - \`model\` (required): exact provider model identifier.
139
+ - \`region\` (optional): provider/model routing hint; defaults in durable
140
+ metadata to \`"provider_managed"\`.
141
+ - \`maxTokens\` (optional): positive integer. If omitted, the provider adapter
142
+ chooses its default.
143
+ - \`reasoningLevel\` (optional): non-empty provider-native string. Orcha does
144
+ not translate values between providers.
145
+ - \`outputType\` (optional): \`"text"\` (default) or \`"json"\`. Image and
146
+ audio are reserved but not implemented.
147
+ - \`outputSchema\` (required for JSON output): JSON Schema used for provider
148
+ structured output and final validation.
149
+
150
+ \`instructions.md\` is required and cannot be empty. At compile time it becomes
151
+ the base system prompt. Orcha appends the compact available-skill catalog and
152
+ the full instructions for skills already loaded in this durable session.
153
+
154
+ ## Core execution model
155
+
156
+ An **agent** is the compiled definition: model settings, instructions, actions,
157
+ skills, and evaluations. An agent can create many independent sessions.
158
+
159
+ A **session** is one durable conversation owned by one agent. It has one
160
+ \`sessionId\`, optional name and metadata, fixed prompt variables, and one
161
+ append-only JSONL timeline. Completing one response does not close the
162
+ session—the application can resume it later. A session cannot be transferred
163
+ to another registered agent, but later runs may use a different provider or
164
+ model if that same agent's configuration changes.
165
+
166
+ A **run** is one attempt to advance a session. \`run()\` creates a session and
167
+ its first run. A conversational \`resume(sessionId, { content })\` creates the
168
+ next numbered run in that session. Each run accumulates its own model usage and
169
+ ends in exactly one of these states:
170
+
171
+ - \`completed\`: the model produced final output.
172
+ - \`waiting_for_client_action\`: the model requested work that only the
173
+ application can perform. The run is paused, not completed.
174
+ - \`failed\`: validation, provider, storage, or execution failed. The durable
175
+ events remain available for diagnosis.
176
+
177
+ An **execution** is the in-process handle returned by one call to \`run()\` or
178
+ \`resume()\`. It exposes a cumulative output stream, latest snapshot, final
179
+ result promise, and evaluation promise. An execution ends when that invocation
180
+ completes, pauses, or fails; the durable session may continue through another
181
+ execution.
182
+
183
+ A **model round** is one provider request inside a run. One run may contain
184
+ several rounds:
185
+
186
+ \`\`\`text
187
+ user input
188
+ → model round
189
+ → tool calls
190
+ → tool results
191
+ → another model round
192
+ → final answer
193
+ \`\`\`
194
+
195
+ Local actions and skill loads are handled automatically inside the same
196
+ execution. Their results are sent back to the model and the model loop
197
+ continues without application involvement.
198
+
199
+ A **client action** deliberately crosses the application boundary. Orcha can
200
+ describe the tool to the model but cannot execute it because the operation
201
+ belongs to a browser, mobile app, approval system, or other caller-owned
202
+ environment. The complete pause/continue flow is:
203
+
204
+ \`\`\`text
205
+ 1. Application calls agent.run(...) or agent.resume(...content).
206
+ 2. Model requests one or more client actions.
207
+ 3. Orcha stores client_action.requested and run.paused.
208
+ 4. execution.result resolves with:
209
+ {
210
+ status: "waiting_for_client_action",
211
+ sessionId,
212
+ clientToolCalls: [{ callId, name, arguments }]
213
+ }
214
+ 5. Application executes every requested action.
215
+ 6. Application calls agent.resume(sessionId, {
216
+ toolResults: [{ callId, output, isError? }]
217
+ }).
218
+ 7. Orcha validates every callId and output, stores the results, and continues
219
+ the same paused run from its prior model context.
220
+ 8. The resumed execution either completes, requests more client actions, or
221
+ fails.
222
+ \`\`\`
223
+
224
+ Every pending call must be resolved exactly once in one resume operation.
225
+ \`callId\` links the submitted result to the model's request; the action name
226
+ must not be substituted for it. Re-submitting the identical resolved result is
227
+ idempotent and returns the prior completed result. Submitting different data
228
+ for an already-resolved call fails with \`action_result_conflict\`.
229
+
230
+ \`clientCapabilities\` is supplied per invocation because different callers
231
+ may support different client actions. Orcha exposes only declared client
232
+ actions to that model round. Local actions are always available when compiled.
233
+
234
+ Only one execution may mutate a session at a time. Concurrent calls for the
235
+ same \`sessionId\` return \`session_busy\`; different sessions can run
236
+ independently.
237
+
238
+ ## Running and resuming
239
+
240
+ \`\`\`ts
241
+ const execution = orcha.supportBot.run({
242
+ content: "Check subscription sub_123.",
243
+ name: "Subscription check",
244
+ metadata: { accountId: "acct_123" },
245
+ clientCapabilities: ["request_human_approval"],
246
+ });
247
+
248
+ for await (const snapshot of execution.stream) {
249
+ console.log(snapshot);
250
+ }
251
+
252
+ const result = await execution.result;
253
+ const evaluations = await execution.evaluations;
254
+ \`\`\`
255
+
256
+ \`run()\` input fields:
257
+
258
+ - \`content\` (required): a non-empty string or array of text/file blocks.
259
+ File blocks contain \`type\`, \`mimeType\`, and \`fileUri\`; unsupported
260
+ provider/content combinations fail explicitly.
261
+ - \`name\` (optional): trimmed session label from 1 through 200 characters.
262
+ - \`metadata\` (optional): at most 50 fields with non-empty keys and finite
263
+ string, number, boolean, or null values. Metadata is durable and available
264
+ to local action context; never place secrets in it.
265
+ - \`variables\` (optional): at most 50 string values whose keys are JavaScript
266
+ identifiers. They replace \`{{ variableName }}\` placeholders in
267
+ \`instructions.md\`, are fixed when the session is created, and are reused
268
+ by later resumes. A missing referenced variable fails the run.
269
+ - \`clientCapabilities\` (optional): action names the current caller can
270
+ execute. Client actions not declared here are withheld from the model.
271
+
272
+ \`run()\` always creates a new durable session. Continue one with:
273
+
274
+ \`\`\`ts
275
+ const execution = orcha.supportBot.resume(sessionId, {
276
+ content: "Continue with the confirmed account.",
277
+ });
278
+ \`\`\`
279
+
280
+ If a result has \`status: "waiting_for_client_action"\`, execute the requested
281
+ client actions in the application and submit every result:
282
+
283
+ \`\`\`ts
284
+ orcha.supportBot.resume(sessionId, {
285
+ toolResults: [
286
+ { callId: "call_123", output: { approved: true } }
287
+ ],
288
+ });
289
+ \`\`\`
290
+
291
+ Never invent call IDs. Use the IDs returned in \`clientToolCalls\`.
292
+
293
+ ## Actions
294
+
295
+ Each action has metadata and, for local actions, executable code:
296
+
297
+ \`\`\`text
298
+ actions/
299
+ lookupAccount/
300
+ index.json
301
+ index.js
302
+ \`\`\`
303
+
304
+ \`\`\`json
305
+ {
306
+ "name": "lookup_account",
307
+ "description": "Look up one account.",
308
+ "execution": "local",
309
+ "parameters": {
310
+ "type": "object",
311
+ "properties": {
312
+ "accountId": { "type": "string" }
313
+ },
314
+ "required": ["accountId"],
315
+ "additionalProperties": false
316
+ },
317
+ "outputSchema": {
318
+ "type": "object",
319
+ "properties": {
320
+ "status": { "type": "string" }
321
+ },
322
+ "required": ["status"],
323
+ "additionalProperties": false
324
+ }
325
+ }
326
+ \`\`\`
327
+
328
+ \`\`\`js
329
+ export default async function lookupAccount({ accountId }) {
330
+ return { status: "active" };
331
+ }
332
+ \`\`\`
333
+
334
+ Client actions use \`"execution": "client"\` and do not include executable
335
+ code. Orcha pauses until the caller submits their results. Keep action names,
336
+ descriptions, schemas, and implementations aligned.
337
+
338
+ Action \`index.json\` fields:
339
+
340
+ - \`name\` (required): model-facing tool name, 1–64 letters, numbers,
341
+ underscores, or hyphens. \`load_skill\` is reserved.
342
+ - \`description\` (required): tells the model when and why to call the action.
343
+ - \`execution\` (required): \`"local"\` executes \`index.js\`; \`"client"\`
344
+ pauses the run and delegates execution to the application.
345
+ - \`parameters\` (required): JSON Schema for model-generated arguments.
346
+ - \`outputSchema\` (optional): JSON Schema validated against local or submitted
347
+ client output before the model receives it.
348
+ - \`timeoutMs\` (optional): integer from 1 through 120000; defaults to 10000.
349
+ - \`permissions.env\` (optional): names copied from \`orcha.init().actions.env\`
350
+ into the action context.
351
+ - \`permissions.network\` (optional): exact hosts or wildcard subdomains such
352
+ as \`"api.example.com"\` or \`"*.example.com"\` allowed through
353
+ \`context.fetch\`. Redirects are rejected.
354
+ - \`sideEffect\` (optional): descriptive metadata for whether the operation
355
+ mutates external state. It does not currently change execution behavior.
356
+
357
+ \`orcha.init().actions.runtime\` and an action's \`execution\` solve different
358
+ problems:
359
+
360
+ - \`execution: "client"\`: Orcha never executes code for this action.
361
+ - \`execution: "local"\` + \`runtime: "sandbox"\`: compiled code runs in a
362
+ QuickJS isolate with JSON-only inputs/outputs, default 32 MB memory, default
363
+ 512 KB stack, interruptible timeout, declared environment values, and
364
+ allowlisted network access through the provided context.
365
+ - \`execution: "local"\` + \`runtime: "native"\`: code runs in the host Node.js
366
+ process. It can use host privileges directly. The timeout rejects slow
367
+ asynchronous work but cannot interrupt synchronous blocking code.
368
+
369
+ Global local-action configuration:
370
+
371
+ \`\`\`ts
372
+ actions: {
373
+ runtime: "sandbox", // required when any registered action is local
374
+ env: {
375
+ BILLING_API_TOKEN: process.env.BILLING_API_TOKEN,
376
+ },
377
+ sandbox: {
378
+ memoryLimitMb: 32,
379
+ stackLimitKb: 512,
380
+ },
381
+ }
382
+ \`\`\`
383
+
384
+ The local action signature is
385
+ \`(parameters, context) => output | Promise<output>\`. Context contains
386
+ \`sessionId\`, immutable session \`metadata\`, a stable \`idempotencyKey\`,
387
+ allowlisted \`env\`, guarded \`fetch\`, and prefixed \`log\`.
388
+
389
+ ## Skills
390
+
391
+ Skills are lazy-loaded procedural instructions. Register only intended skills:
392
+
393
+ \`\`\`js
394
+ // skills/index.js
395
+ import { defineSkills } from "orchajs/skills";
396
+
397
+ export default defineSkills({
398
+ incidentTriage: "./incidentTriage",
399
+ });
400
+ \`\`\`
401
+
402
+ Each skill folder contains \`index.json\` metadata and \`instructions.md\`.
403
+ The model receives a compact catalog and can call the internal \`load_skill\`
404
+ tool. Loaded instructions remain active for the durable session. Lifecycle
405
+ events are \`skill.requested\`, \`skill.loaded\`, and \`skill.failed\`.
406
+
407
+ Skill \`index.json\` fields:
408
+
409
+ - \`name\` (required): model-facing name, 1–64 letters, numbers, underscores,
410
+ or hyphens; unique within the agent.
411
+ - \`description\` (required): compact catalog description shown before loading.
412
+ - \`triggers\` (optional): non-empty array of non-empty situations describing
413
+ when the model should load the skill.
414
+
415
+ \`instructions.md\` is required and cannot be empty. The key in
416
+ \`skills/index.js\` is only a registration label; \`index.json.name\` is the
417
+ name used by the model and durable events. Unregistered folders are ignored.
418
+
419
+ ## Tests
420
+
421
+ Register tests in \`tests/index.js\`:
422
+
423
+ \`\`\`js
424
+ import { defineTests } from "orchajs/testing";
425
+
426
+ export default defineTests({
427
+ activeAccount: "./activeAccount",
428
+ });
429
+ \`\`\`
430
+
431
+ Each case's \`index.json\` defines \`input\`, mocked responses for every action,
432
+ and \`expect\`. Tests run the real compiled agent and provider but never execute
433
+ real actions. The mocked action set must exactly match the compiled action set.
434
+
435
+ \`\`\`json
436
+ {
437
+ "input": { "content": "Check account acct_123." },
438
+ "actions": {
439
+ "lookup_account": {
440
+ "responses": [
441
+ { "output": { "status": "active" } }
442
+ ]
443
+ }
444
+ },
445
+ "expect": {
446
+ "status": "completed",
447
+ "text": { "contains": ["active"] },
448
+ "actions": [
449
+ {
450
+ "name": "lookup_account",
451
+ "arguments": { "equals": { "accountId": "acct_123" } }
452
+ }
453
+ ]
454
+ }
455
+ }
456
+ \`\`\`
457
+
458
+ Test sessions use the \`ses_test_\` prefix and end with a \`test.completed\`
459
+ event. Prefer semantic output assertions; verify exact identifiers and values
460
+ through action-argument assertions.
461
+
462
+ Test \`index.json\` fields:
463
+
464
+ - \`description\` (optional): human-readable purpose.
465
+ - \`input.content\` (required): string or multimodal content array.
466
+ - \`input.variables\` (optional): string map available to the session.
467
+ - \`input.metadata\` (optional): string, number, boolean, or null values.
468
+ - \`actions\` (required): exactly one key for every compiled action, including
469
+ local actions. Every \`responses\` array is consumed in call order.
470
+ - \`responses[].output\` (required): mocked action result.
471
+ - \`responses[].isError\` (optional): marks the mocked result as an error.
472
+ - \`expect.status\` (optional): \`"completed"\` or \`"failed"\`; defaults to
473
+ \`"completed"\`.
474
+ - \`expect.output.equals\` / \`partial\` (optional): exact or recursive partial
475
+ comparison against structured output.
476
+ - \`expect.text.contains\` / \`excludes\` (optional): case-sensitive semantic
477
+ text checks.
478
+ - \`expect.actions\` (optional): ordered expected calls. Each may assert
479
+ \`arguments.equals\` or \`arguments.partial\`.
480
+
481
+ The registration key in \`tests/index.js\` is the test selector used by
482
+ \`orcha test agent/testName\`; its value resolves to the case folder.
483
+
484
+ ## Evaluations
485
+
486
+ Evaluations are asynchronous LLM judges registered in
487
+ \`evaluations/index.js\` with \`defineEvaluations\` from
488
+ \`orchajs/evaluations\`. Each folder's \`index.json\` defines its provider,
489
+ model, metrics, and thresholds:
490
+
491
+ \`\`\`json
492
+ {
493
+ "name": "response_quality",
494
+ "description": "Grounding of agent responses.",
495
+ "instructions": "Judge the complete response using only confirmed evidence recorded in the session.",
496
+ "enabled": true,
497
+ "provider": "openai",
498
+ "model": "gpt-5-mini",
499
+ "metrics": [
500
+ {
501
+ "name": "groundedness",
502
+ "description": "The answer relies on confirmed session evidence.",
503
+ "threshold": 0.8
504
+ }
505
+ ]
506
+ }
507
+ \`\`\`
508
+
509
+ \`execution.result\` does not wait for judges. Await
510
+ \`execution.evaluations\` when results must finish before process exit.
511
+ Evaluations always finish during \`orcha test\`; judge errors and missed
512
+ thresholds fail the test. Lifecycle events are \`evaluation.requested\`,
513
+ \`evaluation.completed\`, and \`evaluation.failed\`.
514
+
515
+ Evaluation \`index.json\` fields:
516
+
517
+ - \`name\` (required): durable model-facing identifier, 1–64 letters, numbers,
518
+ underscores, or hyphens; unique within the agent.
519
+ - \`description\` (optional): short human-facing summary of the evaluator.
520
+ - \`instructions\` (optional): detailed prompt supplied to the judge. When
521
+ omitted, \`description\` is used for backward compatibility.
522
+ - \`enabled\` (optional): defaults to \`true\`. Disabled evaluations are
523
+ compiled but do not run.
524
+ - \`provider\` and \`model\` (required): independently select the judge. The
525
+ provider must also exist in \`orcha.init().providers\`.
526
+ - \`maxTokens\` (optional): positive integer; defaults to 2000 for judges.
527
+ - \`reasoningLevel\` (optional): non-empty provider-native string forwarded
528
+ without translation.
529
+ - \`metrics\` (required): non-empty array with unique metric names.
530
+ - \`metrics[].name\`: 1–64 letters, numbers, underscores, or hyphens.
531
+ - \`metrics[].description\`: exact criterion supplied to the judge.
532
+ - \`metrics[].threshold\`: inclusive number from 0 to 1. A metric passes when
533
+ the returned score is greater than or equal to this threshold.
534
+
535
+ The judge sees a sanitized transcript of user/assistant messages, action
536
+ requests and outcomes, client-action activity, and loaded skill names. It does
537
+ not receive internal reasoning blocks, replay metadata, previous evaluation
538
+ results, or test assertions. It must return exactly one score, reasoning
539
+ string, and non-empty evidence array for every configured metric.
540
+
541
+ ## Sessions and logs
542
+
543
+ JSONL is the durable source of truth. Each line is one complete JSON object;
544
+ never treat the file as one JSON array. Events are append-only and ordered by
545
+ \`sequence\`.
546
+
547
+ All events use this envelope:
548
+
549
+ \`\`\`ts
550
+ type SessionEvent<T> = {
551
+ sequence: number; // starts at 1 and increases across the whole session
552
+ type: SessionEventType;
553
+ timestamp: string; // ISO-8601 UTC timestamp
554
+ run?: number; // present for run-scoped events
555
+ data: T;
556
+ };
557
+ \`\`\`
558
+
559
+ Session-scoped events omit \`run\`. Optional properties whose values are
560
+ \`undefined\` are omitted from serialized JSON.
561
+
562
+ Shared stored structures:
563
+
564
+ \`\`\`ts
565
+ type Usage = {
566
+ inputTokens: number;
567
+ outputTokens: number;
568
+ reasoningTokens: number | null;
569
+ cacheReadTokens: number;
570
+ cacheWriteTokens: number;
571
+ };
572
+
573
+ type ErrorData = {
574
+ code: string;
575
+ message: string;
576
+ retryable?: boolean;
577
+ };
578
+
579
+ type UserContent =
580
+ | { type: "text"; text: string }
581
+ | {
582
+ type: "image" | "video" | "audio" | "url";
583
+ mimeType: string;
584
+ fileUri: string;
585
+ };
586
+
587
+ type AssistantContent =
588
+ | { type: "text"; text: string }
589
+ | {
590
+ type: "reasoning";
591
+ text: string;
592
+ replay?: { providerId?: string; opaqueData?: string };
593
+ }
594
+ | {
595
+ type: "tool_call";
596
+ callId: string;
597
+ name: string;
598
+ arguments: Record<string, unknown>;
599
+ replay?: { providerId?: string; opaqueData?: string };
600
+ };
601
+
602
+ type ToolResult = {
603
+ callId: string;
604
+ output: unknown;
605
+ isError?: boolean;
606
+ };
607
+ \`\`\`
608
+
609
+ Exact event payloads:
610
+
611
+ \`\`\`ts
612
+ type SessionCreated = SessionEvent<{
613
+ schemaVersion: 1;
614
+ sessionId: string;
615
+ agent: string;
616
+ status: "active";
617
+ name?: string;
618
+ metadata: Record<string, string | number | boolean | null>;
619
+ variables: Record<string, string>;
620
+ }>; // type "session.created", no run
621
+
622
+ type SessionUpdated = SessionEvent<{
623
+ name?: string;
624
+ metadata?: Record<string, string | number | boolean | null>;
625
+ }>; // type "session.updated", no run
626
+
627
+ type RunStarted = SessionEvent<{
628
+ status: "running";
629
+ agent: string;
630
+ provider: string;
631
+ model: string;
632
+ region: string;
633
+ reasoningLevel?: string;
634
+ outputType: "text" | "json";
635
+ clientCapabilities: string[];
636
+ }>; // type "run.started"
637
+
638
+ type UserMessageCreated = SessionEvent<{
639
+ role: "user";
640
+ content: UserContent[];
641
+ }>; // type "message.created"
642
+
643
+ type AssistantMessageCreated = SessionEvent<{
644
+ status: "completed" | "incomplete";
645
+ provider: string;
646
+ model: string;
647
+ responseId?: string;
648
+ stopReason?: "end_turn" | "tool_call" | "max_tokens" |
649
+ "content_filter" | "unknown";
650
+ role: "assistant";
651
+ content: AssistantContent[];
652
+ parsedOutput?: unknown; // final JSON output only
653
+ usage: Usage;
654
+ durationMs: number;
655
+ }>; // type "message.created"
656
+
657
+ type ToolMessageCreated = SessionEvent<{
658
+ role: "tool";
659
+ content: ToolResult[];
660
+ }>; // type "message.created"
661
+
662
+ type ActionRequested = SessionEvent<{
663
+ callId: string;
664
+ name: string;
665
+ arguments: Record<string, unknown>;
666
+ sourceHash?: string;
667
+ idempotencyKey: string; // sessionId:callId
668
+ }>; // type "action.requested"
669
+
670
+ type ActionCompleted = SessionEvent<{
671
+ callId: string;
672
+ name: string;
673
+ output: unknown;
674
+ sourceHash?: string;
675
+ durationMs: number;
676
+ }>; // type "action.completed"
677
+
678
+ type ActionFailed = SessionEvent<{
679
+ callId: string;
680
+ name: string;
681
+ sourceHash?: string;
682
+ durationMs: number;
683
+ error: {
684
+ code: "action_execution_failed";
685
+ message: string;
686
+ };
687
+ }>; // type "action.failed"
688
+
689
+ type ClientActionRequested = SessionEvent<{
690
+ status: "waiting";
691
+ calls: Array<{
692
+ callId: string;
693
+ name: string;
694
+ arguments: Record<string, unknown>;
695
+ }>;
696
+ localResults: ToolResult[];
697
+ toolCallOrder: string[];
698
+ }>; // type "client_action.requested"
699
+
700
+ type ClientActionResolved = SessionEvent<{
701
+ status: "completed";
702
+ results: ToolResult[];
703
+ }>; // type "client_action.resolved"
704
+
705
+ type SkillRequested = SessionEvent<{
706
+ callId: string;
707
+ name: unknown;
708
+ }>; // type "skill.requested"
709
+
710
+ type SkillLoaded = SessionEvent<{
711
+ callId: string;
712
+ name: string;
713
+ alreadyLoaded: boolean;
714
+ }>; // type "skill.loaded", no run
715
+
716
+ type SkillFailed = SessionEvent<{
717
+ callId: string;
718
+ name: unknown;
719
+ error: {
720
+ code: "skill_not_found";
721
+ message: string;
722
+ };
723
+ }>; // type "skill.failed"
724
+
725
+ type EvaluationRequested = SessionEvent<{
726
+ name: string;
727
+ provider: string;
728
+ model: string;
729
+ evaluatedThroughSequence: number;
730
+ }>; // type "evaluation.requested"
731
+
732
+ type EvaluationMetric = {
733
+ name: string;
734
+ score: number;
735
+ threshold: number;
736
+ passed: boolean;
737
+ reasoning: string;
738
+ evidence: string[];
739
+ };
740
+
741
+ type EvaluationCompleted = SessionEvent<{
742
+ name: string;
743
+ status: "passed" | "failed";
744
+ metrics: EvaluationMetric[];
745
+ usage?: Usage;
746
+ durationMs: number;
747
+ provider: string;
748
+ model: string;
749
+ evaluatedThroughSequence: number;
750
+ }>; // type "evaluation.completed"
751
+
752
+ type EvaluationFailed = SessionEvent<{
753
+ name: string;
754
+ status: "error";
755
+ metrics: [];
756
+ durationMs: number;
757
+ error: { message: string };
758
+ provider: string;
759
+ model: string;
760
+ evaluatedThroughSequence: number;
761
+ }>; // type "evaluation.failed"
762
+
763
+ type RunPaused = SessionEvent<{
764
+ status: "waiting_for_client_action";
765
+ clientToolCalls: Array<{
766
+ callId: string;
767
+ name: string;
768
+ arguments: Record<string, unknown>;
769
+ }>;
770
+ usage: Usage;
771
+ }>; // type "run.paused"
772
+
773
+ type RunCompleted = SessionEvent<{
774
+ status: "completed";
775
+ durationMs: number;
776
+ usage: Usage;
777
+ }>; // type "run.completed"
778
+
779
+ type RunFailed = SessionEvent<{
780
+ status: "failed";
781
+ durationMs: number;
782
+ usage?: Usage;
783
+ error: ErrorData;
784
+ }>; // type "run.failed"
785
+
786
+ type TestCompleted = SessionEvent<{
787
+ suiteId: string;
788
+ agent: string;
789
+ test: string;
790
+ status: "passed" | "failed";
791
+ durationMs: number;
792
+ assertions: Array<{
793
+ path: string;
794
+ passed: boolean;
795
+ message: string;
796
+ expected?: unknown;
797
+ actual?: unknown;
798
+ }>;
799
+ usage?: Usage;
800
+ evaluations?: Array<{
801
+ name: string;
802
+ status: "passed" | "failed" | "error";
803
+ metrics: EvaluationMetric[];
804
+ usage?: Usage;
805
+ durationMs: number;
806
+ error?: { message: string };
807
+ }>;
808
+ error?: ErrorData;
809
+ }>; // type "test.completed", no run
810
+ \`\`\`
811
+
812
+ Typical event order:
813
+
814
+ \`\`\`text
815
+ session.created
816
+ run.started
817
+ message.created (user)
818
+ message.created (assistant, possibly with tool_call)
819
+ action.requested → action.completed|action.failed # local action
820
+ message.created (tool)
821
+ ...additional model/action rounds...
822
+ message.created (assistant final)
823
+ run.completed
824
+ evaluation.requested
825
+ evaluation.completed|evaluation.failed
826
+ \`\`\`
827
+
828
+ For client actions, \`client_action.requested\` and \`run.paused\` replace the
829
+ immediate tool message. A later \`resume(...toolResults)\` appends
830
+ \`client_action.resolved\`, the tool message, and continues the same run
831
+ number. A conversational \`resume(...content)\` starts a new run number.
832
+
833
+ ## How Orcha works behind the scenes
834
+
835
+ ### Compilation
836
+
837
+ 1. \`orcha/index.ts\` calls \`orcha.init()\` with explicit agent paths.
838
+ 2. The compiler reads each registered agent's \`index.json\` and
839
+ \`instructions.md\`.
840
+ 3. Every directory under \`actions/\` is compiled. Skills, tests, and
841
+ evaluations are included only through their local \`index.js\` registry.
842
+ 4. Local action source is bundled and SHA-256 hashed. Production bundles keep
843
+ only provider adapters required by agents and enabled evaluations.
844
+ 5. Invalid paths, duplicate model-facing names, missing files, unsupported
845
+ configuration values, and missing schema objects fail before execution.
846
+ Concrete action arguments and outputs are validated against their schemas
847
+ when the action is used.
848
+
849
+ \`orcha dev\` repeats validation when files change. \`orcha build\` performs
850
+ offline production compilation. Neither command invokes a provider.
851
+
852
+ ### Run lifecycle
853
+
854
+ 1. \`run()\` creates a \`ses_<uuid>\`, acquires the per-session execution lock,
855
+ appends \`session.created\`, then starts run 1.
856
+ 2. \`resume()\` reads and validates the existing session. Message continuation
857
+ starts a new run; submitted client results continue the paused run.
858
+ 3. The provider receives the base instructions, available-skill catalog,
859
+ loaded skill instructions, normalized conversation messages, action
860
+ schemas, model settings, and current client capabilities.
861
+ 4. Provider-specific responses are normalized into text, reasoning, and tool
862
+ call blocks. Opaque replay metadata is stored only when a provider needs it
863
+ to replay its own prior block correctly.
864
+ 5. Tool calls are checked against compiled actions and declared client
865
+ capabilities. Arguments and outputs are validated against JSON Schema.
866
+ 6. \`load_skill\` updates durable session instructions. Local actions execute
867
+ through the configured runtime. Client actions pause safely. Tool results
868
+ are reordered to match the model's original call order.
869
+ 7. The model loop continues until final output, failure, a client pause, or
870
+ the maximum of 10 action rounds.
871
+ 8. Usage is normalized and aggregated across every model call in the run.
872
+ 9. After \`run.completed\`, enabled evaluations start in the background.
873
+ \`execution.result\` is already available; \`execution.evaluations\` waits
874
+ for judge completion and durable persistence.
875
+
876
+ The per-session lock prevents two model executions from mutating one session
877
+ at once. Background evaluation writes queue behind active runs so they cannot
878
+ cause \`resume()\` to fail spuriously or reuse sequence numbers.
879
+
880
+ ### Replay and context
881
+
882
+ Orcha does not send raw JSONL back to the model. It projects durable events
883
+ into provider-neutral conversation messages. Completed conversational runs
884
+ become user, assistant, and tool messages; lifecycle bookkeeping such as
885
+ durations, test assertions, and evaluation events is excluded from model
886
+ context. Reasoning text and provider replay metadata are retained where needed
887
+ for faithful continuation but are omitted from evaluation transcripts.
888
+
889
+ ### Reading sessions
890
+
891
+ - \`agent.get(sessionId)\` projects the latest status, pending client actions,
892
+ last output, metadata, and aggregate usage.
893
+ - \`agent.history(sessionId, { page, pageSize })\` returns a safe user-facing
894
+ timeline rather than raw provider bookkeeping.
895
+ - \`agent.events(sessionId, { page, pageSize })\` returns the canonical durable
896
+ event records for observability, audit, and debugging tools. For subsequent
897
+ pages of a changing session, pass the first response's \`throughSequence\`
898
+ back in the options to keep pagination on a stable event boundary.
899
+ - \`agent.list({ page, pageSize, status, metadata })\` lists projected session
900
+ snapshots.
901
+ - Never read storage files directly. Use these methods so applications remain
902
+ compatible with JSONL, SQLite, IndexedDB, remote, and future storage adapters.
903
+
904
+ ## Change rules
905
+
906
+ - Register every new agent, skill, test, and evaluation explicitly.
907
+ - Keep runtime behavior provider-neutral.
908
+ - Do not call real actions from tests.
909
+ - Do not commit \`.orcha/\`; it contains generated output and session data.
910
+ - Run \`orcha dev\` after filesystem changes and \`orcha test\` when behavior
911
+ changes.
912
+ - Do not weaken assertions to match incorrect behavior. Remove an assertion
913
+ only when it is stricter than the documented agent contract.
914
+ `;
915
+ //# sourceMappingURL=cli-agents-template.js.map