orchajs 0.3.0 → 0.4.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +29 -1
- package/dist/cli-agents-template.d.ts +2 -0
- package/dist/cli-agents-template.d.ts.map +1 -0
- package/dist/cli-agents-template.js +907 -0
- package/dist/cli-agents-template.js.map +1 -0
- package/dist/cli.js +335 -16
- package/dist/cli.js.map +1 -1
- package/dist/runtime/agent.d.ts.map +1 -1
- package/dist/runtime/agent.js +10 -2
- package/dist/runtime/agent.js.map +1 -1
- package/dist/testing.js +1 -0
- package/dist/testing.js.map +1 -1
- package/dist/types.d.ts +1 -0
- package/dist/types.d.ts.map +1 -1
- package/package.json +1 -1
package/README.md
CHANGED
|
@@ -213,6 +213,34 @@ streaming are supported.
|
|
|
213
213
|
`run()` always starts a new session; `resume()` continues one with either a new
|
|
214
214
|
message or pending client tool results.
|
|
215
215
|
|
|
216
|
+
The CLI provides the same focused workflows without introducing a separate
|
|
217
|
+
application runtime:
|
|
218
|
+
|
|
219
|
+
```bash
|
|
220
|
+
# Offline validation that watches orcha/**
|
|
221
|
+
npx orcha dev
|
|
222
|
+
|
|
223
|
+
# Execute a registered agent. Input can also come from a JSON file or stdin.
|
|
224
|
+
npx orcha run exampleAgent --input "Help me understand this invoice"
|
|
225
|
+
|
|
226
|
+
# Continue an existing session or resolve pending client actions.
|
|
227
|
+
npx orcha run exampleAgent --session ses_123 --input "Summarize it"
|
|
228
|
+
npx orcha run exampleAgent --session ses_123 --tool-results results.json
|
|
229
|
+
|
|
230
|
+
# Run all tests, one agent, or one case.
|
|
231
|
+
npx orcha test
|
|
232
|
+
npx orcha test exampleAgent
|
|
233
|
+
npx orcha test exampleAgent/exampleCase
|
|
234
|
+
|
|
235
|
+
# Offline production compilation.
|
|
236
|
+
npx orcha build
|
|
237
|
+
```
|
|
238
|
+
|
|
239
|
+
`run` and `test` support `--json` for scripts and coding agents.
|
|
240
|
+
`orcha init` also creates a comprehensive `AGENTS.md` describing Orcha's
|
|
241
|
+
filesystem contract and APIs. The CLI does not provide a UI; applications can
|
|
242
|
+
build one from the runtime session APIs or durable JSONL logs.
|
|
243
|
+
|
|
216
244
|
Orcha owns its built-in provider adapters. Messages, tool calls and results,
|
|
217
245
|
streaming snapshots, usage, and errors are normalized before reaching the
|
|
218
246
|
runtime, so application and action APIs do not change when an agent switches
|
|
@@ -531,7 +559,7 @@ application continues to use its normal `npm run build` command.
|
|
|
531
559
|
- [x] Registered LLM-judged `/evaluations`
|
|
532
560
|
- [ ] Agent `/guardrails`
|
|
533
561
|
- [x] Production compiler and `orcha build`
|
|
534
|
-
- [
|
|
562
|
+
- [x] Lightweight CLI for init, validation, execution, tests, and builds
|
|
535
563
|
|
|
536
564
|
See [`ROADMAP.md`](./ROADMAP.md) for what's coming after — guardrails,
|
|
537
565
|
OpenTelemetry tracing, MCP interop adapters, and remote session storage remain
|
|
@@ -0,0 +1,2 @@
|
|
|
1
|
+
export declare const AGENTS_MD = "# Orcha project guide\n\nThis repository uses OrchaJS, a filesystem-convention framework for durable\nAI agents. Treat the `orcha/` directory as source code. Do not edit generated\nfiles under `.orcha/`.\n\n## Commands\n\n- `orcha init` creates the initial Orcha files without overwriting files.\n- `orcha dev` validates the registry and watches `orcha/**` for changes.\n- `orcha run <agent> --input \"\u2026\"` executes one registered agent.\n- `orcha run <agent> --input-file request.json` accepts structured input.\n- `orcha run <agent> --session <id> --input \"\u2026\"` continues a session.\n- `orcha run <agent> --session <id> --tool-results results.json` submits\n pending client-action results.\n- `orcha test` runs every registered agent test.\n- `orcha test <agent>` or `orcha test <agent>/<case>` narrows the run.\n- `orcha build` creates the production Orcha bundle without calling models.\n- Add `--json` to `run` and `test` for machine-readable output.\n\n`run` and `test` use real providers and require credentials. `dev` and\n`build` are offline. Session logs are JSONL files under\n`.orcha/sessions/<sessionId>.jsonl`.\n\n## Registry\n\n`orcha/index.ts` initializes providers and explicitly registers agents:\n\n```ts\nimport { orcha } from \"orchajs\";\n\norcha.init({\n providers: {\n anthropic: process.env.ANTHROPIC_API_KEY ?? \"\",\n openai: process.env.OPENAI_API_KEY ?? \"\",\n },\n actions: { runtime: \"sandbox\" },\n agents: {\n supportBot: \"./supportBot\",\n },\n});\n```\n\nOnly registered folders are compiled. Agent keys become runtime properties\nsuch as `orcha.supportBot`. Use `actions.runtime: \"sandbox\"` for isolated\nlocal action execution or `\"native\"` when the application intentionally\nallows action modules to execute in its Node.js process.\n\n`orcha.init()` fields:\n\n- `providers` (required): provider configurations keyed by built-in provider\n name.\n- `agents` (required): runtime property names mapped to folders relative to\n `orcha/`. At least one agent is required.\n- `actions` (required only when a registered agent has local actions):\n selects the local execution runtime and its environment/sandbox settings.\n- `storage.strategy` (optional): currently only `\"node-jsonl\"`.\n- `storage.directory` (optional): session directory relative to project\n root; defaults to `.orcha/sessions`.\n- `root` (optional): absolute or working-directory-relative project root;\n defaults to `ORCHA_PROJECT_ROOT` and then `process.cwd()`.\n\nProvider configuration shapes:\n\n```ts\nproviders: {\n anthropic: process.env.ANTHROPIC_API_KEY ?? \"\",\n deepseek: process.env.DEEPSEEK_API_KEY ?? \"\",\n googlegenai: process.env.GOOGLE_API_KEY ?? \"\",\n openai: {\n apiKey: process.env.OPENAI_API_KEY ?? \"\",\n baseUrl: \"https://api.openai.com/v1\", // optional override\n },\n vertexai: {\n project: process.env.GOOGLE_CLOUD_PROJECT ?? \"\",\n location: process.env.GOOGLE_CLOUD_LOCATION ?? \"us-central1\",\n // credentials is optional; omit it to use Google ADC.\n credentials: {\n clientEmail: process.env.GOOGLE_CLIENT_EMAIL ?? \"\",\n privateKey: process.env.GOOGLE_PRIVATE_KEY ?? \"\",\n },\n baseUrl: undefined, // optional override\n },\n bedrock: {\n region: process.env.AWS_REGION ?? \"us-east-1\",\n // credentials is optional; omit it to use the AWS credential chain.\n credentials: {\n accessKeyId: process.env.AWS_ACCESS_KEY_ID ?? \"\",\n secretAccessKey: process.env.AWS_SECRET_ACCESS_KEY ?? \"\",\n sessionToken: process.env.AWS_SESSION_TOKEN,\n },\n baseUrl: undefined, // optional override\n },\n}\n```\n\nAPI-key providers accept either a string shorthand or\n`{ apiKey, baseUrl? }`. Vertex AI requires `project` and `location`;\nexplicit service-account credentials are optional. Bedrock requires `region`;\nexplicit AWS credentials are optional. Never place credentials in\n`index.json`, instructions, tests, session metadata, or committed files.\n\n## Agent folders\n\n```text\norcha/\n index.ts\n supportBot/\n index.json\n instructions.md\n actions/\n skills/\n tests/\n evaluations/\n```\n\n`index.json` selects the model:\n\n```json\n{\n \"provider\": \"anthropic\",\n \"model\": \"claude-sonnet-4-6\",\n \"region\": \"provider_managed\",\n \"maxTokens\": 10240,\n \"outputType\": \"text\"\n}\n```\n\nOptional fields include `reasoningLevel`, `outputType: \"json\"`, and an\n`outputSchema` JSON Schema. Provider-specific reasoning values are forwarded\nwithout translation. Put the agent's stable role, boundaries, and operating\ninstructions in `instructions.md`.\n\nAgent `index.json` fields:\n\n- `provider` (required): `\"anthropic\"`, `\"bedrock\"`, `\"deepseek\"`,\n `\"openai\"`, `\"googlegenai\"`, or `\"vertexai\"`.\n- `model` (required): exact provider model identifier.\n- `region` (optional): provider/model routing hint; defaults in durable\n metadata to `\"provider_managed\"`.\n- `maxTokens` (optional): positive integer. If omitted, the provider adapter\n chooses its default.\n- `reasoningLevel` (optional): non-empty provider-native string. Orcha does\n not translate values between providers.\n- `outputType` (optional): `\"text\"` (default) or `\"json\"`. Image and\n audio are reserved but not implemented.\n- `outputSchema` (required for JSON output): JSON Schema used for provider\n structured output and final validation.\n\n`instructions.md` is required and cannot be empty. At compile time it becomes\nthe base system prompt. Orcha appends the compact available-skill catalog and\nthe full instructions for skills already loaded in this durable session.\n\n## Core execution model\n\nAn **agent** is the compiled definition: model settings, instructions, actions,\nskills, and evaluations. An agent can create many independent sessions.\n\nA **session** is one durable conversation owned by one agent. It has one\n`sessionId`, optional name and metadata, fixed prompt variables, and one\nappend-only JSONL timeline. Completing one response does not close the\nsession\u2014the application can resume it later. A session cannot be transferred\nto another registered agent, but later runs may use a different provider or\nmodel if that same agent's configuration changes.\n\nA **run** is one attempt to advance a session. `run()` creates a session and\nits first run. A conversational `resume(sessionId, { content })` creates the\nnext numbered run in that session. Each run accumulates its own model usage and\nends in exactly one of these states:\n\n- `completed`: the model produced final output.\n- `waiting_for_client_action`: the model requested work that only the\n application can perform. The run is paused, not completed.\n- `failed`: validation, provider, storage, or execution failed. The durable\n events remain available for diagnosis.\n\nAn **execution** is the in-process handle returned by one call to `run()` or\n`resume()`. It exposes a cumulative output stream, latest snapshot, final\nresult promise, and evaluation promise. An execution ends when that invocation\ncompletes, pauses, or fails; the durable session may continue through another\nexecution.\n\nA **model round** is one provider request inside a run. One run may contain\nseveral rounds:\n\n```text\nuser input\n \u2192 model round\n \u2192 tool calls\n \u2192 tool results\n \u2192 another model round\n \u2192 final answer\n```\n\nLocal actions and skill loads are handled automatically inside the same\nexecution. Their results are sent back to the model and the model loop\ncontinues without application involvement.\n\nA **client action** deliberately crosses the application boundary. Orcha can\ndescribe the tool to the model but cannot execute it because the operation\nbelongs to a browser, mobile app, approval system, or other caller-owned\nenvironment. The complete pause/continue flow is:\n\n```text\n1. Application calls agent.run(...) or agent.resume(...content).\n2. Model requests one or more client actions.\n3. Orcha stores client_action.requested and run.paused.\n4. execution.result resolves with:\n {\n status: \"waiting_for_client_action\",\n sessionId,\n clientToolCalls: [{ callId, name, arguments }]\n }\n5. Application executes every requested action.\n6. Application calls agent.resume(sessionId, {\n toolResults: [{ callId, output, isError? }]\n }).\n7. Orcha validates every callId and output, stores the results, and continues\n the same paused run from its prior model context.\n8. The resumed execution either completes, requests more client actions, or\n fails.\n```\n\nEvery pending call must be resolved exactly once in one resume operation.\n`callId` links the submitted result to the model's request; the action name\nmust not be substituted for it. Re-submitting the identical resolved result is\nidempotent and returns the prior completed result. Submitting different data\nfor an already-resolved call fails with `action_result_conflict`.\n\n`clientCapabilities` is supplied per invocation because different callers\nmay support different client actions. Orcha exposes only declared client\nactions to that model round. Local actions are always available when compiled.\n\nOnly one execution may mutate a session at a time. Concurrent calls for the\nsame `sessionId` return `session_busy`; different sessions can run\nindependently.\n\n## Running and resuming\n\n```ts\nconst execution = orcha.supportBot.run({\n content: \"Check subscription sub_123.\",\n name: \"Subscription check\",\n metadata: { accountId: \"acct_123\" },\n clientCapabilities: [\"request_human_approval\"],\n});\n\nfor await (const snapshot of execution.stream) {\n console.log(snapshot);\n}\n\nconst result = await execution.result;\nconst evaluations = await execution.evaluations;\n```\n\n`run()` input fields:\n\n- `content` (required): a non-empty string or array of text/file blocks.\n File blocks contain `type`, `mimeType`, and `fileUri`; unsupported\n provider/content combinations fail explicitly.\n- `name` (optional): trimmed session label from 1 through 200 characters.\n- `metadata` (optional): at most 50 fields with non-empty keys and finite\n string, number, boolean, or null values. Metadata is durable and available\n to local action context; never place secrets in it.\n- `variables` (optional): at most 50 string values whose keys are JavaScript\n identifiers. They replace `{{ variableName }}` placeholders in\n `instructions.md`, are fixed when the session is created, and are reused\n by later resumes. A missing referenced variable fails the run.\n- `clientCapabilities` (optional): action names the current caller can\n execute. Client actions not declared here are withheld from the model.\n\n`run()` always creates a new durable session. Continue one with:\n\n```ts\nconst execution = orcha.supportBot.resume(sessionId, {\n content: \"Continue with the confirmed account.\",\n});\n```\n\nIf a result has `status: \"waiting_for_client_action\"`, execute the requested\nclient actions in the application and submit every result:\n\n```ts\norcha.supportBot.resume(sessionId, {\n toolResults: [\n { callId: \"call_123\", output: { approved: true } }\n ],\n});\n```\n\nNever invent call IDs. Use the IDs returned in `clientToolCalls`.\n\n## Actions\n\nEach action has metadata and, for local actions, executable code:\n\n```text\nactions/\n lookupAccount/\n index.json\n index.js\n```\n\n```json\n{\n \"name\": \"lookup_account\",\n \"description\": \"Look up one account.\",\n \"execution\": \"local\",\n \"parameters\": {\n \"type\": \"object\",\n \"properties\": {\n \"accountId\": { \"type\": \"string\" }\n },\n \"required\": [\"accountId\"],\n \"additionalProperties\": false\n },\n \"outputSchema\": {\n \"type\": \"object\",\n \"properties\": {\n \"status\": { \"type\": \"string\" }\n },\n \"required\": [\"status\"],\n \"additionalProperties\": false\n }\n}\n```\n\n```js\nexport default async function lookupAccount({ accountId }) {\n return { status: \"active\" };\n}\n```\n\nClient actions use `\"execution\": \"client\"` and do not include executable\ncode. Orcha pauses until the caller submits their results. Keep action names,\ndescriptions, schemas, and implementations aligned.\n\nAction `index.json` fields:\n\n- `name` (required): model-facing tool name, 1\u201364 letters, numbers,\n underscores, or hyphens. `load_skill` is reserved.\n- `description` (required): tells the model when and why to call the action.\n- `execution` (required): `\"local\"` executes `index.js`; `\"client\"`\n pauses the run and delegates execution to the application.\n- `parameters` (required): JSON Schema for model-generated arguments.\n- `outputSchema` (optional): JSON Schema validated against local or submitted\n client output before the model receives it.\n- `timeoutMs` (optional): integer from 1 through 120000; defaults to 10000.\n- `permissions.env` (optional): names copied from `orcha.init().actions.env`\n into the action context.\n- `permissions.network` (optional): exact hosts or wildcard subdomains such\n as `\"api.example.com\"` or `\"*.example.com\"` allowed through\n `context.fetch`. Redirects are rejected.\n- `sideEffect` (optional): descriptive metadata for whether the operation\n mutates external state. It does not currently change execution behavior.\n\n`orcha.init().actions.runtime` and an action's `execution` solve different\nproblems:\n\n- `execution: \"client\"`: Orcha never executes code for this action.\n- `execution: \"local\"` + `runtime: \"sandbox\"`: compiled code runs in a\n QuickJS isolate with JSON-only inputs/outputs, default 32 MB memory, default\n 512 KB stack, interruptible timeout, declared environment values, and\n allowlisted network access through the provided context.\n- `execution: \"local\"` + `runtime: \"native\"`: code runs in the host Node.js\n process. It can use host privileges directly. The timeout rejects slow\n asynchronous work but cannot interrupt synchronous blocking code.\n\nGlobal local-action configuration:\n\n```ts\nactions: {\n runtime: \"sandbox\", // required when any registered action is local\n env: {\n BILLING_API_TOKEN: process.env.BILLING_API_TOKEN,\n },\n sandbox: {\n memoryLimitMb: 32,\n stackLimitKb: 512,\n },\n}\n```\n\nThe local action signature is\n`(parameters, context) => output | Promise<output>`. Context contains\n`sessionId`, immutable session `metadata`, a stable `idempotencyKey`,\nallowlisted `env`, guarded `fetch`, and prefixed `log`.\n\n## Skills\n\nSkills are lazy-loaded procedural instructions. Register only intended skills:\n\n```js\n// skills/index.js\nimport { defineSkills } from \"orchajs/skills\";\n\nexport default defineSkills({\n incidentTriage: \"./incidentTriage\",\n});\n```\n\nEach skill folder contains `index.json` metadata and `instructions.md`.\nThe model receives a compact catalog and can call the internal `load_skill`\ntool. Loaded instructions remain active for the durable session. Lifecycle\nevents are `skill.requested`, `skill.loaded`, and `skill.failed`.\n\nSkill `index.json` fields:\n\n- `name` (required): model-facing name, 1\u201364 letters, numbers, underscores,\n or hyphens; unique within the agent.\n- `description` (required): compact catalog description shown before loading.\n- `triggers` (optional): non-empty array of non-empty situations describing\n when the model should load the skill.\n\n`instructions.md` is required and cannot be empty. The key in\n`skills/index.js` is only a registration label; `index.json.name` is the\nname used by the model and durable events. Unregistered folders are ignored.\n\n## Tests\n\nRegister tests in `tests/index.js`:\n\n```js\nimport { defineTests } from \"orchajs/testing\";\n\nexport default defineTests({\n activeAccount: \"./activeAccount\",\n});\n```\n\nEach case's `index.json` defines `input`, mocked responses for every action,\nand `expect`. Tests run the real compiled agent and provider but never execute\nreal actions. The mocked action set must exactly match the compiled action set.\n\n```json\n{\n \"input\": { \"content\": \"Check account acct_123.\" },\n \"actions\": {\n \"lookup_account\": {\n \"responses\": [\n { \"output\": { \"status\": \"active\" } }\n ]\n }\n },\n \"expect\": {\n \"status\": \"completed\",\n \"text\": { \"contains\": [\"active\"] },\n \"actions\": [\n {\n \"name\": \"lookup_account\",\n \"arguments\": { \"equals\": { \"accountId\": \"acct_123\" } }\n }\n ]\n }\n}\n```\n\nTest sessions use the `ses_test_` prefix and end with a `test.completed`\nevent. Prefer semantic output assertions; verify exact identifiers and values\nthrough action-argument assertions.\n\nTest `index.json` fields:\n\n- `description` (optional): human-readable purpose.\n- `input.content` (required): string or multimodal content array.\n- `input.variables` (optional): string map available to the session.\n- `input.metadata` (optional): string, number, boolean, or null values.\n- `actions` (required): exactly one key for every compiled action, including\n local actions. Every `responses` array is consumed in call order.\n- `responses[].output` (required): mocked action result.\n- `responses[].isError` (optional): marks the mocked result as an error.\n- `expect.status` (optional): `\"completed\"` or `\"failed\"`; defaults to\n `\"completed\"`.\n- `expect.output.equals` / `partial` (optional): exact or recursive partial\n comparison against structured output.\n- `expect.text.contains` / `excludes` (optional): case-sensitive semantic\n text checks.\n- `expect.actions` (optional): ordered expected calls. Each may assert\n `arguments.equals` or `arguments.partial`.\n\nThe registration key in `tests/index.js` is the test selector used by\n`orcha test agent/testName`; its value resolves to the case folder.\n\n## Evaluations\n\nEvaluations are asynchronous LLM judges registered in\n`evaluations/index.js` with `defineEvaluations` from\n`orchajs/evaluations`. Each folder's `index.json` defines its provider,\nmodel, metrics, and thresholds:\n\n```json\n{\n \"name\": \"response_quality\",\n \"enabled\": true,\n \"provider\": \"openai\",\n \"model\": \"gpt-5-mini\",\n \"metrics\": [\n {\n \"name\": \"groundedness\",\n \"description\": \"The answer relies on confirmed session evidence.\",\n \"threshold\": 0.8\n }\n ]\n}\n```\n\n`execution.result` does not wait for judges. Await\n`execution.evaluations` when results must finish before process exit.\nEvaluations always finish during `orcha test`; judge errors and missed\nthresholds fail the test. Lifecycle events are `evaluation.requested`,\n`evaluation.completed`, and `evaluation.failed`.\n\nEvaluation `index.json` fields:\n\n- `name` (required): durable model-facing identifier, 1\u201364 letters, numbers,\n underscores, or hyphens; unique within the agent.\n- `description` (optional): overall judging objective.\n- `enabled` (optional): defaults to `true`. Disabled evaluations are\n compiled but do not run.\n- `provider` and `model` (required): independently select the judge. The\n provider must also exist in `orcha.init().providers`.\n- `maxTokens` (optional): positive integer; defaults to 2000 for judges.\n- `reasoningLevel` (optional): non-empty provider-native string forwarded\n without translation.\n- `metrics` (required): non-empty array with unique metric names.\n- `metrics[].name`: 1\u201364 letters, numbers, underscores, or hyphens.\n- `metrics[].description`: exact criterion supplied to the judge.\n- `metrics[].threshold`: inclusive number from 0 to 1. A metric passes when\n the returned score is greater than or equal to this threshold.\n\nThe judge sees a sanitized transcript of user/assistant messages, action\nrequests and outcomes, client-action activity, and loaded skill names. It does\nnot receive internal reasoning blocks, replay metadata, previous evaluation\nresults, or test assertions. It must return exactly one score, reasoning\nstring, and non-empty evidence array for every configured metric.\n\n## Sessions and logs\n\nJSONL is the durable source of truth. Each line is one complete JSON object;\nnever treat the file as one JSON array. Events are append-only and ordered by\n`sequence`.\n\nAll events use this envelope:\n\n```ts\ntype SessionEvent<T> = {\n sequence: number; // starts at 1 and increases across the whole session\n type: SessionEventType;\n timestamp: string; // ISO-8601 UTC timestamp\n run?: number; // present for run-scoped events\n data: T;\n};\n```\n\nSession-scoped events omit `run`. Optional properties whose values are\n`undefined` are omitted from serialized JSON.\n\nShared stored structures:\n\n```ts\ntype Usage = {\n inputTokens: number;\n outputTokens: number;\n reasoningTokens: number | null;\n cacheReadTokens: number;\n cacheWriteTokens: number;\n};\n\ntype ErrorData = {\n code: string;\n message: string;\n retryable?: boolean;\n};\n\ntype UserContent =\n | { type: \"text\"; text: string }\n | {\n type: \"image\" | \"video\" | \"audio\" | \"url\";\n mimeType: string;\n fileUri: string;\n };\n\ntype AssistantContent =\n | { type: \"text\"; text: string }\n | {\n type: \"reasoning\";\n text: string;\n replay?: { providerId?: string; opaqueData?: string };\n }\n | {\n type: \"tool_call\";\n callId: string;\n name: string;\n arguments: Record<string, unknown>;\n replay?: { providerId?: string; opaqueData?: string };\n };\n\ntype ToolResult = {\n callId: string;\n output: unknown;\n isError?: boolean;\n};\n```\n\nExact event payloads:\n\n```ts\ntype SessionCreated = SessionEvent<{\n schemaVersion: 1;\n sessionId: string;\n agent: string;\n status: \"active\";\n name?: string;\n metadata: Record<string, string | number | boolean | null>;\n variables: Record<string, string>;\n}>; // type \"session.created\", no run\n\ntype SessionUpdated = SessionEvent<{\n name?: string;\n metadata?: Record<string, string | number | boolean | null>;\n}>; // type \"session.updated\", no run\n\ntype RunStarted = SessionEvent<{\n status: \"running\";\n agent: string;\n provider: string;\n model: string;\n region: string;\n reasoningLevel?: string;\n outputType: \"text\" | \"json\";\n clientCapabilities: string[];\n}>; // type \"run.started\"\n\ntype UserMessageCreated = SessionEvent<{\n role: \"user\";\n content: UserContent[];\n}>; // type \"message.created\"\n\ntype AssistantMessageCreated = SessionEvent<{\n status: \"completed\" | \"incomplete\";\n provider: string;\n model: string;\n responseId?: string;\n stopReason?: \"end_turn\" | \"tool_call\" | \"max_tokens\" |\n \"content_filter\" | \"unknown\";\n role: \"assistant\";\n content: AssistantContent[];\n parsedOutput?: unknown; // final JSON output only\n usage: Usage;\n durationMs: number;\n}>; // type \"message.created\"\n\ntype ToolMessageCreated = SessionEvent<{\n role: \"tool\";\n content: ToolResult[];\n}>; // type \"message.created\"\n\ntype ActionRequested = SessionEvent<{\n callId: string;\n name: string;\n arguments: Record<string, unknown>;\n sourceHash?: string;\n idempotencyKey: string; // sessionId:callId\n}>; // type \"action.requested\"\n\ntype ActionCompleted = SessionEvent<{\n callId: string;\n name: string;\n output: unknown;\n sourceHash?: string;\n durationMs: number;\n}>; // type \"action.completed\"\n\ntype ActionFailed = SessionEvent<{\n callId: string;\n name: string;\n sourceHash?: string;\n durationMs: number;\n error: {\n code: \"action_execution_failed\";\n message: string;\n };\n}>; // type \"action.failed\"\n\ntype ClientActionRequested = SessionEvent<{\n status: \"waiting\";\n calls: Array<{\n callId: string;\n name: string;\n arguments: Record<string, unknown>;\n }>;\n localResults: ToolResult[];\n toolCallOrder: string[];\n}>; // type \"client_action.requested\"\n\ntype ClientActionResolved = SessionEvent<{\n status: \"completed\";\n results: ToolResult[];\n}>; // type \"client_action.resolved\"\n\ntype SkillRequested = SessionEvent<{\n callId: string;\n name: unknown;\n}>; // type \"skill.requested\"\n\ntype SkillLoaded = SessionEvent<{\n callId: string;\n name: string;\n alreadyLoaded: boolean;\n}>; // type \"skill.loaded\", no run\n\ntype SkillFailed = SessionEvent<{\n callId: string;\n name: unknown;\n error: {\n code: \"skill_not_found\";\n message: string;\n };\n}>; // type \"skill.failed\"\n\ntype EvaluationRequested = SessionEvent<{\n name: string;\n provider: string;\n model: string;\n evaluatedThroughSequence: number;\n}>; // type \"evaluation.requested\"\n\ntype EvaluationMetric = {\n name: string;\n score: number;\n threshold: number;\n passed: boolean;\n reasoning: string;\n evidence: string[];\n};\n\ntype EvaluationCompleted = SessionEvent<{\n name: string;\n status: \"passed\" | \"failed\";\n metrics: EvaluationMetric[];\n usage?: Usage;\n durationMs: number;\n provider: string;\n model: string;\n evaluatedThroughSequence: number;\n}>; // type \"evaluation.completed\"\n\ntype EvaluationFailed = SessionEvent<{\n name: string;\n status: \"error\";\n metrics: [];\n durationMs: number;\n error: { message: string };\n provider: string;\n model: string;\n evaluatedThroughSequence: number;\n}>; // type \"evaluation.failed\"\n\ntype RunPaused = SessionEvent<{\n status: \"waiting_for_client_action\";\n clientToolCalls: Array<{\n callId: string;\n name: string;\n arguments: Record<string, unknown>;\n }>;\n usage: Usage;\n}>; // type \"run.paused\"\n\ntype RunCompleted = SessionEvent<{\n status: \"completed\";\n durationMs: number;\n usage: Usage;\n}>; // type \"run.completed\"\n\ntype RunFailed = SessionEvent<{\n status: \"failed\";\n durationMs: number;\n usage?: Usage;\n error: ErrorData;\n}>; // type \"run.failed\"\n\ntype TestCompleted = SessionEvent<{\n suiteId: string;\n agent: string;\n test: string;\n status: \"passed\" | \"failed\";\n durationMs: number;\n assertions: Array<{\n path: string;\n passed: boolean;\n message: string;\n expected?: unknown;\n actual?: unknown;\n }>;\n usage?: Usage;\n evaluations?: Array<{\n name: string;\n status: \"passed\" | \"failed\" | \"error\";\n metrics: EvaluationMetric[];\n usage?: Usage;\n durationMs: number;\n error?: { message: string };\n }>;\n error?: ErrorData;\n}>; // type \"test.completed\", no run\n```\n\nTypical event order:\n\n```text\nsession.created\nrun.started\nmessage.created (user)\nmessage.created (assistant, possibly with tool_call)\naction.requested \u2192 action.completed|action.failed # local action\nmessage.created (tool)\n...additional model/action rounds...\nmessage.created (assistant final)\nrun.completed\nevaluation.requested\nevaluation.completed|evaluation.failed\n```\n\nFor client actions, `client_action.requested` and `run.paused` replace the\nimmediate tool message. A later `resume(...toolResults)` appends\n`client_action.resolved`, the tool message, and continues the same run\nnumber. A conversational `resume(...content)` starts a new run number.\n\n## How Orcha works behind the scenes\n\n### Compilation\n\n1. `orcha/index.ts` calls `orcha.init()` with explicit agent paths.\n2. The compiler reads each registered agent's `index.json` and\n `instructions.md`.\n3. Every directory under `actions/` is compiled. Skills, tests, and\n evaluations are included only through their local `index.js` registry.\n4. Local action source is bundled and SHA-256 hashed. Production bundles keep\n only provider adapters required by agents and enabled evaluations.\n5. Invalid paths, duplicate model-facing names, missing files, unsupported\n configuration values, and missing schema objects fail before execution.\n Concrete action arguments and outputs are validated against their schemas\n when the action is used.\n\n`orcha dev` repeats validation when files change. `orcha build` performs\noffline production compilation. Neither command invokes a provider.\n\n### Run lifecycle\n\n1. `run()` creates a `ses_<uuid>`, acquires the per-session execution lock,\n appends `session.created`, then starts run 1.\n2. `resume()` reads and validates the existing session. Message continuation\n starts a new run; submitted client results continue the paused run.\n3. The provider receives the base instructions, available-skill catalog,\n loaded skill instructions, normalized conversation messages, action\n schemas, model settings, and current client capabilities.\n4. Provider-specific responses are normalized into text, reasoning, and tool\n call blocks. Opaque replay metadata is stored only when a provider needs it\n to replay its own prior block correctly.\n5. Tool calls are checked against compiled actions and declared client\n capabilities. Arguments and outputs are validated against JSON Schema.\n6. `load_skill` updates durable session instructions. Local actions execute\n through the configured runtime. Client actions pause safely. Tool results\n are reordered to match the model's original call order.\n7. The model loop continues until final output, failure, a client pause, or\n the maximum of 10 action rounds.\n8. Usage is normalized and aggregated across every model call in the run.\n9. After `run.completed`, enabled evaluations start in the background.\n `execution.result` is already available; `execution.evaluations` waits\n for judge completion and durable persistence.\n\nThe per-session lock prevents two model executions from mutating one session\nat once. Background evaluation writes queue behind active runs so they cannot\ncause `resume()` to fail spuriously or reuse sequence numbers.\n\n### Replay and context\n\nOrcha does not send raw JSONL back to the model. It projects durable events\ninto provider-neutral conversation messages. Completed conversational runs\nbecome user, assistant, and tool messages; lifecycle bookkeeping such as\ndurations, test assertions, and evaluation events is excluded from model\ncontext. Reasoning text and provider replay metadata are retained where needed\nfor faithful continuation but are omitted from evaluation transcripts.\n\n### Reading sessions\n\n- `agent.get(sessionId)` projects the latest status, pending client actions,\n last output, metadata, and aggregate usage.\n- `agent.history(sessionId, { page, pageSize })` returns a safe user-facing\n timeline rather than raw provider bookkeeping.\n- `agent.list({ page, pageSize, status, metadata })` lists projected session\n snapshots.\n- Read JSONL directly when building observability, audit, or debugging tools\n that require the exact append-only event stream described above.\n\n## Change rules\n\n- Register every new agent, skill, test, and evaluation explicitly.\n- Keep runtime behavior provider-neutral.\n- Do not call real actions from tests.\n- Do not commit `.orcha/`; it contains generated output and session data.\n- Run `orcha dev` after filesystem changes and `orcha test` when behavior\n changes.\n- Do not weaken assertions to match incorrect behavior. Remove an assertion\n only when it is stricter than the documented agent contract.\n";
|
|
2
|
+
//# sourceMappingURL=cli-agents-template.d.ts.map
|
|
@@ -0,0 +1 @@
|
|
|
1
|
+
{"version":3,"file":"cli-agents-template.d.ts","sourceRoot":"","sources":["../src/cli-agents-template.ts"],"names":[],"mappings":"AAAA,eAAO,MAAM,SAAS,4s9BAy4BrB,CAAC"}
|