@enderfga/claw-orchestrator 7.5.3 → 7.5.5
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +21 -22
- package/configs/engines/README.md +7 -6
- package/dist/bin/cli.js +1 -1
- package/dist/bin/cli.js.map +1 -1
- package/dist/src/acp-server.d.ts +1 -1
- package/dist/src/acp-server.js +7 -5
- package/dist/src/acp-server.js.map +1 -1
- package/dist/src/autoloop/notify.d.ts +5 -7
- package/dist/src/autoloop/notify.js +21 -20
- package/dist/src/autoloop/notify.js.map +1 -1
- package/dist/src/embedded-server.js +8 -5
- package/dist/src/embedded-server.js.map +1 -1
- package/dist/src/fanout.d.ts +6 -0
- package/dist/src/fanout.js +1 -0
- package/dist/src/fanout.js.map +1 -1
- package/dist/src/index.js +21 -13
- package/dist/src/index.js.map +1 -1
- package/dist/src/kernel/engine.d.ts +39 -1
- package/dist/src/kernel/engine.js +120 -10
- package/dist/src/kernel/engine.js.map +1 -1
- package/dist/src/kernel/nodes/fanout.js +1 -0
- package/dist/src/kernel/nodes/fanout.js.map +1 -1
- package/dist/src/kernel/types.d.ts +9 -0
- package/dist/src/kernel/types.js.map +1 -1
- package/dist/src/openai-compat.d.ts +2 -2
- package/dist/src/openai-compat.js +5 -2
- package/dist/src/openai-compat.js.map +1 -1
- package/dist/src/session-manager.d.ts +7 -2
- package/dist/src/session-manager.js +21 -7
- package/dist/src/session-manager.js.map +1 -1
- package/dist/src/types.d.ts +2 -0
- package/openclaw.plugin.json +1 -1
- package/package.json +2 -2
- package/skills/SKILL.md +31 -32
- package/skills/references/acp.md +19 -36
- package/skills/references/autoloop.md +158 -180
- package/skills/references/claude-cli-tracking.md +27 -27
- package/skills/references/cli.md +62 -79
- package/skills/references/council.md +40 -63
- package/skills/references/dashboard.md +42 -55
- package/skills/references/getting-started.md +20 -14
- package/skills/references/inbox.md +6 -4
- package/skills/references/mcp.md +29 -24
- package/skills/references/multi-engine.md +105 -153
- package/skills/references/observability.md +42 -32
- package/skills/references/openai-compat.md +169 -303
- package/skills/references/sessions.md +20 -29
- package/skills/references/tools.md +67 -78
- package/skills/references/ultra.md +17 -16
- package/skills/references/ultraapp.md +59 -64
- package/skills/references/verification.md +29 -52
- package/skills/references/workflow.md +49 -107
- package/skills/ultraapp/SKILL.md +9 -10
|
@@ -17,7 +17,7 @@ SessionManager
|
|
|
17
17
|
├── engine: 'grok' → PersistentGrokSession
|
|
18
18
|
│ └── Wraps: grok -p --output-format json (xAI Grok Build, per-message spawning)
|
|
19
19
|
├── engine: 'cursor' → PersistentCursorSession (legacy)
|
|
20
|
-
│ └── Wraps: agent -p --
|
|
20
|
+
│ └── Wraps: cursor-agent -p --trust --output-format stream-json (per-message spawning)
|
|
21
21
|
├── engine: 'opencode' → PersistentOpencodeSession
|
|
22
22
|
│ └── Wraps: opencode run --format json (per-message spawning)
|
|
23
23
|
└── engine: 'custom' → PersistentCustomSession
|
|
@@ -28,33 +28,28 @@ SessionManager
|
|
|
28
28
|
|
|
29
29
|
### Claude Code (`engine: 'claude'`)
|
|
30
30
|
|
|
31
|
-
Default engine. Long-running subprocess with streaming JSON I/O. Tested with Claude Code CLI **2.1.
|
|
31
|
+
Default engine. Long-running subprocess with streaming JSON I/O. Tested with Claude Code CLI **2.1.280**.
|
|
32
32
|
|
|
33
33
|
- Persistent multi-turn conversations
|
|
34
34
|
- Real-time streaming (text, tool_use, tool_result, system events)
|
|
35
35
|
- Session resume via `--resume`
|
|
36
36
|
- Full cost tracking from API usage data
|
|
37
|
-
- Cross-session peer messaging (`crossSessionInbound`):
|
|
37
|
+
- Cross-session peer messaging (`crossSessionInbound`): this session's policy for messages from
|
|
38
|
+
other Claude Code sessions on the same machine — `accept` delivers them, `hold` waits for approval
|
|
39
|
+
in that session's terminal, `refuse` rejects them. It is a settings key, passed through
|
|
40
|
+
`--settings`. Set it explicitly on orchestrated sessions: when it is unset, the CLI holds messages
|
|
41
|
+
whenever the two sides' permission modes differ, so a message can wait for approval in a terminal
|
|
42
|
+
nobody is watching. A user-level `~/.claude/settings.json` value may take precedence over the
|
|
43
|
+
per-session one.
|
|
38
44
|
- Hook lifecycle events (`includeHookEvents`), subagent output forwarding (`forwardSubagentText`), permission delegation (`permissionPromptTool`), prompt cache optimization (`bare` + `excludeDynamicSystemPromptSections` + `enablePromptCaching1H`), debug control, `--from-pr` resume, and MCP channel subscriptions
|
|
39
|
-
- `--permission-prompts none` is passed whenever no `permissionPromptTool` is configured
|
|
40
|
-
|
|
41
|
-
|
|
42
|
-
|
|
43
|
-
|
|
44
|
-
|
|
45
|
-
|
|
46
|
-
|
|
47
|
-
`sandboxMode: 'read-only'`, which maps to plan mode — measured against 2.1.251, plan mode alone
|
|
48
|
-
refused a direct write, a shell write and a delegated subagent write, so this is defence in depth
|
|
49
|
-
rather than a fix, and switching it on implicitly would silently drop the caller's CLAUDE.md
|
|
50
|
-
- Fork subagent (`forkSubagent`), tool search (`enableToolSearch`), OpenTelemetry logging toggles (`otelLogUserPrompts`, `otelLogRawApiBodies`), `xhigh` effort tier (Opus 4.7), and `stats.pluginErrors` capture — see [CLI 2.1.121 options in SKILL.md](../SKILL.md) and [tools.md](./tools.md)
|
|
51
|
-
|
|
52
|
-
> **Behavior changes from upstream Claude CLI 2.1.121** (worth knowing if you set permission rules):
|
|
53
|
-
>
|
|
54
|
-
> - `--agent` / `--print` now enforce agent frontmatter `permissionMode`, `tools`, `disallowedTools` (was advisory). Affects `council` agent personas.
|
|
55
|
-
> - `Bash(find:*)` permission rule no longer auto-approves `find -exec` or `find -delete`. Add explicit rules if you depend on these.
|
|
56
|
-
> - `--dangerously-skip-permissions` also skips prompts for `.claude/skills/` directory. Treat with care.
|
|
57
|
-
> - Distributed tracing context (`TRACEPARENT` / `TRACESTATE`) is automatically forwarded to the child process — set them in the parent before starting the session.
|
|
45
|
+
- `--permission-prompts none` is passed whenever no `permissionPromptTool` is configured. The
|
|
46
|
+
session has no TTY and no prompt tool, so a tool call the permission mode does not already decide
|
|
47
|
+
is denied (the model sees the denial and can adapt) instead of waiting until the turn timeout.
|
|
48
|
+
With a prompt tool configured, the CLI's default (`host`) is kept so the tool is asked
|
|
49
|
+
- `restricted` → `--restricted`: removes the command- and code-running tools and `WebFetch` from
|
|
50
|
+
the session, and ignores user/project/local settings files (including CLAUDE.md). It is separate
|
|
51
|
+
from `sandboxMode: 'read-only'`, which maps to plan mode
|
|
52
|
+
- Fork subagent (`forkSubagent`), tool search (`enableToolSearch`), OpenTelemetry logging toggles (`otelLogUserPrompts`, `otelLogRawApiBodies`), the `xhigh` effort tier, and `stats.pluginErrors` capture — see [SKILL.md](../SKILL.md#claude-engine-options) and [tools.md](./tools.md)
|
|
58
53
|
|
|
59
54
|
```typescript
|
|
60
55
|
await manager.startSession({
|
|
@@ -67,26 +62,23 @@ await manager.startSession({
|
|
|
67
62
|
|
|
68
63
|
### OpenAI Codex (`engine: 'codex'`)
|
|
69
64
|
|
|
70
|
-
Wraps the `codex exec` subcommand. Each `send()` spawns a new process. Tested with `codex` CLI **0.
|
|
65
|
+
Wraps the `codex exec` subcommand. Each `send()` spawns a new process. Tested with `codex` CLI **0.156.1**.
|
|
71
66
|
|
|
72
|
-
- Non-interactive execution via `codex exec --sandbox workspace-write --json`
|
|
73
|
-
- Real `usage` from the `turn.completed` JSON event (input, output, cached, reasoning tokens). **These are cumulative over the thread, not per turn
|
|
74
|
-
- `contextPercent` is
|
|
67
|
+
- Non-interactive execution via `codex exec --sandbox workspace-write --skip-git-repo-check --json`
|
|
68
|
+
- Real `usage` from the `turn.completed` JSON event (input, output, cached, reasoning tokens). **These are cumulative over the thread, not per turn**, so they replace the session totals rather than being added to them; subtracting consecutive values gives one turn's prompt
|
|
69
|
+
- `contextPercent` is the per-turn prompt measured against **codex's own context limit** (`model_context_window`, read from the thread's rollout file), not the model's published window in the registry. Resuming a thread also seeds the token baseline from the rollout, so the first send does not count the whole thread history as one prompt. This is best-effort: an unreadable or `--ephemeral` thread falls back to the registry window
|
|
75
70
|
- `item.completed` parsing distinguishes `reasoning` / `todo_list` (logged, not counted) from real tool items (`command_execution`, `file_change`, `mcp_tool_call`, `web_search`, which increment `toolCalls`; a non-zero `command_execution.exit_code` increments `toolErrors`)
|
|
76
|
-
- Reasoning effort: the engine-agnostic `effort` maps to `-c model_reasoning_effort=<level>`
|
|
77
|
-
- `
|
|
78
|
-
- `
|
|
79
|
-
- `--
|
|
80
|
-
|
|
81
|
-
|
|
82
|
-
here that checks work — acceptance contracts, evidence diffs, the baseline change set — reads the
|
|
83
|
-
session's `cwd`, so passing the flag would have them verify an untouched tree and report on it.
|
|
84
|
-
Council's own per-agent git worktrees cover the isolation use case with paths the orchestrator owns
|
|
71
|
+
- Reasoning effort: the engine-agnostic `effort` maps to `-c model_reasoning_effort=<level>` for `low|medium|high|xhigh|max|ultra`; `auto` and `ultracode` pass nothing. `-c` values are not validated at spawn, so an unknown level fails at the API rather than at the command line
|
|
72
|
+
- `appendSystemPrompt`: codex has no system-prompt flag, so it is placed at the top of the first message of each new conversation (a resumed thread already carries it)
|
|
73
|
+
- `jsonSchema` → `--output-schema <file>` (written to a temp file, accepted by `exec` and `exec resume`)
|
|
74
|
+
- `noSessionPersistence` → `--ephemeral` (accepted by `exec` and `exec resume`); `ignoreUserConfig` → `--ignore-user-config`, which stops `$CODEX_HOME/config.toml` from choosing the model for an orchestrated run (auth still resolves from `CODEX_HOME`); `addDir` → `--add-dir` on the first turn only, since `exec resume` rejects it and the resumed thread keeps the roots it opened with
|
|
75
|
+
- `codexProfile` → `--profile <name>` (named config profile from `~/.codex/config.toml`), first turn only
|
|
76
|
+
- `--worktree` is not passed. With it, codex writes the turn's edits to `~/.codex/worktrees/<hash>/<repo>` instead of the session's `cwd`, while acceptance contracts, evidence diffs and the baseline change set all read the session's `cwd`. Council's per-agent git worktrees cover the isolation use case
|
|
85
77
|
- Per-session continuity: the `thread_id` from the first turn's `thread.started` event is captured and reused via `codex exec resume <id>` for subsequent sends, so the model sees prior turns
|
|
86
|
-
- `sandboxMode` maps to `--sandbox <mode>` on the first turn.
|
|
78
|
+
- `sandboxMode` maps to `--sandbox <mode>` on the first turn. A resumed thread does not keep it, and `codex exec resume` rejects `--sandbox`, so the policy is restated as `-c sandbox_mode="<mode>"` on every resumed turn
|
|
87
79
|
- One-shot execution per message (no persistent subprocess between sends)
|
|
88
80
|
- Captures the real Codex thread ID and persists it, so later sends and process-level session resume use `codex exec resume <thread_id>`
|
|
89
|
-
- Working directory passed via `-C`
|
|
81
|
+
- Working directory passed via `-C` on the first turn
|
|
90
82
|
- Default model: `gpt-5.5`
|
|
91
83
|
- Requires `codex` CLI >= 0.119 (for `exec resume`): `npm install -g @openai/codex`
|
|
92
84
|
- **Does not support `/goal`** — for that, use `engine: 'codex-app'` below
|
|
@@ -108,13 +100,14 @@ Wraps `codex app-server --listen stdio:// --enable goals` as a long-running JSON
|
|
|
108
100
|
- Long-running subprocess; one `codex app-server` per session
|
|
109
101
|
- JSON-RPC 2.0 over stdio with v2 protocol method names (`initialize`, `thread/start`, `turn/start`, ...)
|
|
110
102
|
- Real-time streaming via `item/agentMessage/delta` notifications
|
|
111
|
-
- Cumulative token tracking from `thread/tokenUsage/updated` notifications. The same notification's `last` breakdown and `modelContextWindow` drive `contextPercent`, so it reports live occupancy against the window the server
|
|
103
|
+
- Cumulative token tracking from `thread/tokenUsage/updated` notifications. The same notification's `last` breakdown and `modelContextWindow` drive `contextPercent`, so it reports live occupancy against the window the server enforces rather than a running total over the model's published window
|
|
104
|
+
- `appendSystemPrompt` is placed at the top of the first message of a new thread (app-server has no system-prompt flag); a resumed thread already carries it
|
|
112
105
|
- Goal lifecycle observation via `thread/goal/updated` and `thread/goal/cleared` notifications
|
|
113
106
|
- Goal control via the `codex_goal_*` tools (which internally send the `/goal` slash command as user text — see [tools.md](./tools.md#codex-13))
|
|
114
107
|
- v2 RPC tools (Codex 0.137): `codex_interrupt` (`turn/interrupt`), `codex_steer` (`turn/steer`), `codex_fork` (`thread/fork`), `codex_rollback` (`thread/rollback`), `codex_models` (`model/list`), `codex_thread_list` (`thread/list`). A `turn/completed` with `status: 'failed'` rejects the turn and increments `toolErrors`.
|
|
115
108
|
- Thread resume: starting with `resumeSessionId` loads the existing thread via `thread/resume` instead of `thread/start`.
|
|
116
109
|
|
|
117
|
-
> **Feature
|
|
110
|
+
> **Feature flag.** `goals` is an experimental Codex feature. The session always passes `--enable goals`; some goal commands may still fail or be ignored by the server.
|
|
118
111
|
|
|
119
112
|
```typescript
|
|
120
113
|
await manager.startSession({
|
|
@@ -133,7 +126,7 @@ await manager.startSession({
|
|
|
133
126
|
|
|
134
127
|
Wraps Google's **Antigravity CLI** (`agy`) — the successor to Gemini CLI (consumer
|
|
135
128
|
Gemini CLI tiers stopped serving 2026-06-18). Each `send()` spawns a new process
|
|
136
|
-
in print mode.
|
|
129
|
+
in print mode. Tested with `agy` **1.2.8**.
|
|
137
130
|
|
|
138
131
|
- One-shot execution per message (no persistent subprocess)
|
|
139
132
|
- **Structured output and real usage** — `--output-format stream-json` emits an
|
|
@@ -145,21 +138,18 @@ in print mode. Adapter behavior is covered through `agy` **1.2.2**.
|
|
|
145
138
|
scrape remains as a fallback for turns that die before emitting `init`. Seed it
|
|
146
139
|
externally via `resumeSessionId` (bare UUID only); read it back from
|
|
147
140
|
`getStats().agyConversationId`.
|
|
148
|
-
- **Empty responses fail
|
|
149
|
-
|
|
150
|
-
|
|
151
|
-
|
|
152
|
-
|
|
153
|
-
|
|
154
|
-
|
|
155
|
-
|
|
156
|
-
|
|
157
|
-
denial can accompany `status: SUCCESS` and a non-empty reply; the refused tool
|
|
158
|
-
names are emitted as `permission_denials`, which SessionManager exposes as
|
|
159
|
-
`SendResult.permissionDenials` without discarding the reply.
|
|
141
|
+
- **Empty responses fail**: an exit-0 result with a missing or blank response
|
|
142
|
+
is treated as a failed turn, not a successful empty reply, and is not retried.
|
|
143
|
+
A conversation id already received from `init` is kept for the next send. When
|
|
144
|
+
the failure follows a tool confirmation refused in plan mode, the error says so
|
|
145
|
+
(a fixed message; native log content is not exposed).
|
|
146
|
+
- **Refused tools**: when agy refuses tools but still replies, the refused tool
|
|
147
|
+
names are exposed as `SendResult.permissionDenials` alongside the reply.
|
|
148
|
+
- `appendSystemPrompt`: agy has no system-prompt flag, so it is placed at the top
|
|
149
|
+
of the first message of each new conversation.
|
|
160
150
|
- **Reasoning effort**: session `effort` and per-turn `session_send` overrides map
|
|
161
151
|
to `--effort`. agy accepts `low`, `medium`, and `high`; everything above that
|
|
162
|
-
(`xhigh`, `max`, `ultra`) clamps to `high`. agy
|
|
152
|
+
(`xhigh`, `max`, `ultra`) clamps to `high`. agy requires an effort with unsuffixed base
|
|
163
153
|
slugs such as `gemini-3.7-flash`, so `auto` resolves those to `high`; a model
|
|
164
154
|
already ending in `-low`, `-medium`, or `-high` keeps that qualified effort.
|
|
165
155
|
Per-turn overrides also work with qualified slugs: the adapter removes a
|
|
@@ -173,14 +163,15 @@ in print mode. Adapter behavior is covered through `agy` **1.2.2**.
|
|
|
173
163
|
did not ask for; run `agy models` to see the tiers a slug actually exposes.
|
|
174
164
|
|
|
175
165
|
- Permission modes: `bypassPermissions` → `--dangerously-skip-permissions`,
|
|
176
|
-
`default` → `--sandbox` (terminal-restricted), and
|
|
166
|
+
`default` and `manual` → `--sandbox` (terminal-restricted), and
|
|
177
167
|
`sandboxMode: 'read-only'` → `--mode plan` (takes precedence). Other modes
|
|
178
168
|
run agy's own approval flow, which can block in headless print mode. A caller
|
|
179
169
|
must explicitly choose `bypassPermissions` for a write-enabled session; it is
|
|
180
170
|
not a recovery mechanism. In particular, an Autoloop Planner stays on
|
|
181
171
|
`--mode plan` when its preserved conversation is resumed.
|
|
182
|
-
-
|
|
183
|
-
|
|
172
|
+
- The engine always passes `--print-timeout` (the send timeout plus 5s), so the
|
|
173
|
+
wrapper's timer decides when a turn ends; without it a stuck headless agy turn
|
|
174
|
+
can run indefinitely
|
|
184
175
|
- Do not rely on an unknown `--model` falling back: current agy versions can
|
|
185
176
|
report `status: ERROR` with no usable response. The adapter rejects result
|
|
186
177
|
errors, non-success statuses, and empty responses. `agy-flash` and the engine
|
|
@@ -216,49 +207,35 @@ await manager.startSession({
|
|
|
216
207
|
### Grok Build (`engine: 'grok'`)
|
|
217
208
|
|
|
218
209
|
Wraps xAI's **Grok Build** CLI. Each `send()` spawns `grok -p <msg> --output-format json`, which
|
|
219
|
-
prints a single JSON object and exits.
|
|
220
|
-
|
|
221
|
-
- **Cost comes from the engine, not from
|
|
222
|
-
`total_cost_usd`, and the wrapper writes it straight into the session's spend
|
|
223
|
-
|
|
224
|
-
|
|
225
|
-
`grok-4.6` is still registered, for its context window and an indicative breakdown; its two price
|
|
226
|
-
tiers ($2/$0.50/$6 under a 200K prompt, $4/$1/$12 at or above, charged across the whole request)
|
|
227
|
-
therefore never have to be modelled here.
|
|
210
|
+
prints a single JSON object and exits. Tested with `grok` **1.0.41**.
|
|
211
|
+
|
|
212
|
+
- **Cost comes from the engine, not from the price table.** The result object carries
|
|
213
|
+
`total_cost_usd`, and the wrapper writes it straight into the session's spend, so the run ledger
|
|
214
|
+
and the `maxBudgetUsd` gate both read what xAI charged. Other engines multiply tokens by a rate in
|
|
215
|
+
`models.ts`. `grok-4.6` is still registered, for its context window and an indicative breakdown.
|
|
228
216
|
- **Real conversation continuity**: the `sessionId` from turn 1 is replayed as `--resume <id>`.
|
|
229
|
-
`--continue` is
|
|
230
|
-
|
|
217
|
+
`--continue` is not used — it means "the most recent session for this cwd", which collides
|
|
218
|
+
between concurrent sessions.
|
|
231
219
|
- Real token counts from `usage` (`input_tokens`, `output_tokens`, `cache_read_input_tokens`).
|
|
232
|
-
These are **per-turn**,
|
|
233
|
-
because the same-looking field on codex is a running total.
|
|
220
|
+
These are **per-turn**, unlike codex, where the same field is cumulative.
|
|
234
221
|
- Permission modes pass straight through: grok's `--permission-mode` takes the same vocabulary we
|
|
235
222
|
use. The one exception is our `manual`, which grok spells `default`.
|
|
236
|
-
- Reasoning effort maps to `--effort`; grok
|
|
237
|
-
|
|
238
|
-
- Session options that reach grok
|
|
223
|
+
- Reasoning effort maps to `--effort`; grok accepts `low|medium|high|xhigh`, so `max` and `ultra`
|
|
224
|
+
clamp to `xhigh`.
|
|
225
|
+
- Session options that reach grok: `appendSystemPrompt` → `--rules` (appends, unlike
|
|
239
226
|
`systemPrompt` → `--system-prompt-override`, which replaces), `allowedTools` → `--tools`,
|
|
240
227
|
`disallowedTools` → `--disallowed-tools`, `jsonSchema` → `--json-schema` (inline, and it implies the
|
|
241
228
|
JSON output format already asked for), `agent` → `--agent`, `agents` → `--agents`,
|
|
242
229
|
`dangerouslySkipPermissions` → `--always-approve`, `customSessionId` → `--session-id`, `forkSession`
|
|
243
230
|
→ `--fork-session`. **grok validates neither tool list**: a name that does not exist is ignored
|
|
244
231
|
rather than rejected, so a typo in a denylist leaves the tool enabled. Prefer an allowlist.
|
|
245
|
-
-
|
|
246
|
-
|
|
247
|
-
|
|
248
|
-
|
|
249
|
-
|
|
250
|
-
|
|
251
|
-
|
|
252
|
-
denying the write tools without denying `task` left the delegation path open. grok also ships
|
|
253
|
-
`--no-subagents`, the obvious next probe, deliberately not wired: the run that would have
|
|
254
|
-
confirmed it hit the account's free-tier limit, and a probe that fails for lack of quota writes no
|
|
255
|
-
file either. A read-only grok session throws rather than running writable under a read-only label.
|
|
256
|
-
- **On a spent free tier, `grok -p` can hang silently instead of erroring.** Earlier in the same
|
|
257
|
-
session it printed a usage-limit message to stderr and exited 1; later invocations produced nothing
|
|
258
|
-
on either stream and never exited — in any directory, with or without `--no-leader`. The session's
|
|
259
|
-
turn timeout is what ends such a turn, so a caller on that tier pays the full timeout before seeing
|
|
260
|
-
a failure. Nothing in this wrapper can distinguish that hang from a slow turn; check `grok -p` by
|
|
261
|
-
hand when a grok session times out with no output.
|
|
232
|
+
- **`sandboxMode: 'read-only'` is refused.** A read-only `--tools` allowlist plus
|
|
233
|
+
`--permission-mode plan` does not stop a delegated subagent from writing, because the subagent
|
|
234
|
+
does not inherit the parent's tool restriction. A read-only grok session therefore throws at start instead of running
|
|
235
|
+
writable.
|
|
236
|
+
- **On an exhausted free tier, `grok -p` may hang with no output instead of exiting with an error.**
|
|
237
|
+
The session's turn timeout is then the only thing that ends the turn. If a grok turn times out with
|
|
238
|
+
no output, run `grok -p` by hand to check your quota.
|
|
262
239
|
- Binary: `grok` (set `GROK_BIN` to override). Not `agent`: xAI's installer claims that name too,
|
|
263
240
|
and so did Cursor's.
|
|
264
241
|
- Requires Grok Build: see `x.ai/cli`.
|
|
@@ -278,20 +255,21 @@ await manager.startSession({
|
|
|
278
255
|
> The `cursor` engine still exists and still works — existing callers are not broken — but it is
|
|
279
256
|
> no longer a documented option, is not version-tracked, and gets no new work.
|
|
280
257
|
>
|
|
281
|
-
>
|
|
282
|
-
>
|
|
283
|
-
>
|
|
284
|
-
> — its `system` init event says `"model": "Auto"` — so a router that spans Claude, GPT and Grok
|
|
285
|
-
> leaves every cost row attributed to a hardcoded proxy rate. And xAI's Grok installer now claims
|
|
286
|
-
> the bare `agent` name, so the binary that name resolves to depends on install order.
|
|
258
|
+
> Cursor itself is still maintained. It left the tracked set because it does not report which
|
|
259
|
+
> model ran (its `system` init event says `"model": "Auto"`), so costs cannot be attributed to a
|
|
260
|
+
> model, and because xAI's Grok installer also claims the bare `agent` binary name.
|
|
287
261
|
|
|
288
|
-
Wraps the Cursor Agent CLI with
|
|
262
|
+
Wraps the Cursor Agent CLI with `-p --output-format stream-json`. Write-enabled sessions use `--force`. Each `send()` spawns a new process.
|
|
289
263
|
|
|
290
264
|
- Conversation continuity: the chat id from the first turn's `system` event is captured and passed back as `--resume <chatId>` on later sends, so the model sees prior turns. `--continue` is deliberately not used: it resumes "the latest chat", which collides between concurrent sessions.
|
|
291
265
|
- One-shot execution per message (no persistent subprocess)
|
|
292
266
|
- Working directory via `--workspace` flag
|
|
293
267
|
- Real token counts from stream-json `result` events (camelCase: `inputTokens`, `outputTokens`, `cacheReadTokens`)
|
|
294
|
-
- `--force` enables auto-approval of file changes. `sandboxMode: 'read-only'` does **not** use `--force
|
|
268
|
+
- `--force` enables auto-approval of file changes. `sandboxMode: 'read-only'` does **not** use `--force`:
|
|
269
|
+
- read-only is enforced by a `.cursor/cli.json` deny config (`Write`/`Edit`/`Shell` denied), written into an isolated temp dir used as the process cwd, with `--workspace` pointing at the real project (the repo tree is never modified)
|
|
270
|
+
- `--mode plan` is passed as well, but it only steers the model; the deny config is the boundary
|
|
271
|
+
- `--sandbox` is not added, since it does not restrict in-workspace writes and overrides the mode
|
|
272
|
+
- read/grep/search remain available
|
|
295
273
|
- `--trust` auto-trusts the workspace without prompting
|
|
296
274
|
- Cursor uses its own model routing (e.g., `sonnet-4`, `gpt-5`, `auto`)
|
|
297
275
|
- Requires Cursor Agent CLI: `curl https://cursor.com/install -fsSL | bash`
|
|
@@ -314,27 +292,22 @@ await manager.startSession({
|
|
|
314
292
|
bundled in `configs/engines/`. The preset form exists so a third-party CLI is
|
|
315
293
|
described once and shipped, rather than retyped by every caller.
|
|
316
294
|
|
|
317
|
-
|
|
318
|
-
|
|
295
|
+
Engines fall into three tiers, decided by how they are verified. Bundled presets
|
|
296
|
+
are always `community`:
|
|
319
297
|
|
|
320
|
-
| Tier | Maintainer | What
|
|
298
|
+
| Tier | Maintainer | What is verified |
|
|
321
299
|
| ------------- | ----------- | --------------------------------------------------------------------------------------------------- |
|
|
322
|
-
| **core** | Maintainers | Wrapped
|
|
300
|
+
| **core** | Maintainers | Wrapped in code, exercised live weekly, pinned to a tested version |
|
|
323
301
|
| **community** | Contributor | The schema is validated and the preset ships. Whether it runs is the maintainer's dated attestation |
|
|
324
|
-
| **legacy** | — | Still wired, no longer tracked
|
|
302
|
+
| **legacy** | — | Still wired, no longer tracked |
|
|
325
303
|
|
|
326
|
-
|
|
327
|
-
|
|
328
|
-
|
|
329
|
-
permanently empty, and an empty tier looks the same as never having done the work.
|
|
330
|
-
So the bar is an attributable, dated, falsifiable claim: `provenance.maintainer`,
|
|
331
|
-
`provenance.verifiedAgainst` (the engine version), `provenance.verifiedOn`, and a
|
|
332
|
-
link to the captured smoke transcript.
|
|
304
|
+
A community preset must name its maintainer and the engine version and date it
|
|
305
|
+
was verified against (`provenance.maintainer`, `provenance.verifiedAgainst`,
|
|
306
|
+
`provenance.verifiedOn`), with a link to the captured smoke-test transcript.
|
|
333
307
|
|
|
334
308
|
A preset never carries protocol translation. A CLI speaking its own wire format
|
|
335
309
|
needs an adapter binary in its author's package, with the preset pointing `bin`
|
|
336
|
-
at it —
|
|
337
|
-
something this project can ship without owning a protocol it cannot test.
|
|
310
|
+
at it — as `@enderfga/dsh-clawo` does.
|
|
338
311
|
|
|
339
312
|
Over HTTP — which includes the `clawo` CLI — a custom engine may be given **only**
|
|
340
313
|
as a preset id. An inline config names a binary and its arguments and is refused
|
|
@@ -355,11 +328,17 @@ Wraps the [sst/opencode](https://github.com/sst/opencode) CLI with `run --format
|
|
|
355
328
|
- NDJSON event stream with envelope `{ type, timestamp, sessionID, ... }`
|
|
356
329
|
- Event types: `text`, `reasoning`, `tool_use`, `step_start`, `step_finish`, `error`
|
|
357
330
|
- `text` and `tool_use` are **cumulative snapshots** keyed by `part.id` / `part.callID`; the wrapper diffs them to produce streaming deltas for `onText` callbacks and counts each tool invocation once
|
|
358
|
-
- Real token counts from `step_finish.part.tokens.{input,output,cache.{read,write}}`. **`input` is the uncached remainder only** — opencode's own `total` is `input + output + cache.read + cache.write` — so the cached part is billed on top of it, not carved out of it, and `contextPercent` is measured against the whole input side. On a resumed turn the remainder is
|
|
331
|
+
- Real token counts from `step_finish.part.tokens.{input,output,cache.{read,write}}`. **`input` is the uncached remainder only** — opencode's own `total` is `input + output + cache.read + cache.write` — so the cached part is billed on top of it, not carved out of it, and `contextPercent` is measured against the whole input side. On a resumed turn the uncached remainder is usually a small fraction of the cached part
|
|
359
332
|
- Reasoning effort maps to `--variant`, opencode's provider-specific effort knob. opencode does not validate the value: a level the provider does not offer runs the turn at its default rather than failing
|
|
333
|
+
- `appendSystemPrompt`: opencode has no system-prompt flag, so it is placed at the top of the first message of each new conversation
|
|
360
334
|
- The wrapper closes the subprocess's stdin immediately after spawn (opencode otherwise reads stdin and blocks on EOF, hanging the call)
|
|
361
335
|
- Provider-agnostic: opencode's `--model` expects `provider/model` form (e.g. `anthropic/claude-sonnet-4`). The wrapper passes `--model` through only when the value contains a `/`; otherwise opencode's own default applies
|
|
362
|
-
- `sandboxMode: 'read-only'`
|
|
336
|
+
- `sandboxMode: 'read-only'` runs a generated `clawo-readonly` agent (`--agent clawo-readonly` plus an `OPENCODE_CONFIG_CONTENT` env var defining it):
|
|
337
|
+
- it denies `edit`, `bash`, `external_directory`, `webfetch` and `task` at the permission level, and also removes those tools via the agent's `tools` map
|
|
338
|
+
- denying `task` matters: otherwise the agent can hand a write to a subagent that runs under the default writable agent
|
|
339
|
+
- OpenCode's built-in `plan` agent is not used, because it permits both `bash` and `edit`
|
|
340
|
+
- if the `clawo-readonly` agent fails to load, the turn is refused rather than run with write access
|
|
341
|
+
- to test a change to this config, use adversarial prompts that include asking the agent to delegate; `opencode agent list` shows compiled rules that look the same for a safe and an unsafe agent
|
|
363
342
|
- Requires opencode installed: `brew install sst/tap/opencode` or `npm install -g opencode-ai`. Auth via `opencode auth login` **or** any provider env var (`ANTHROPIC_API_KEY`, `OPENAI_API_KEY`, `GEMINI_API_KEY`, etc.) — opencode picks up either path
|
|
364
343
|
- Binary: `opencode` (set `OPENCODE_BIN` env var to override)
|
|
365
344
|
|
|
@@ -422,12 +401,14 @@ is not the answer on every engine. See "Stats & Monitoring" in `sessions.md`.
|
|
|
422
401
|
|
|
423
402
|
Team tools (`team_list`, `team_send`) operate on the same virtual-team layer for **every** engine: the "team" is the set of all active sessions managed by SessionManager.
|
|
424
403
|
|
|
425
|
-
| Engine
|
|
426
|
-
|
|
|
427
|
-
| Claude
|
|
428
|
-
| Codex
|
|
429
|
-
| Antigravity
|
|
430
|
-
|
|
|
404
|
+
| Engine | `team_list` | `team_send` |
|
|
405
|
+
| ------------------- | ------------------------------------------ | ------------------------------ |
|
|
406
|
+
| Claude | Lists other active SessionManager sessions | Routes via cross-session inbox |
|
|
407
|
+
| Codex / `codex-app` | Lists other active SessionManager sessions | Routes via cross-session inbox |
|
|
408
|
+
| Antigravity | Lists other active SessionManager sessions | Routes via cross-session inbox |
|
|
409
|
+
| Grok | Lists other active SessionManager sessions | Routes via cross-session inbox |
|
|
410
|
+
| OpenCode | Lists other active SessionManager sessions | Routes via cross-session inbox |
|
|
411
|
+
| Custom | Lists other active SessionManager sessions | Routes via cross-session inbox |
|
|
431
412
|
|
|
432
413
|
Messages are delivered via the inbox system — idle sessions receive immediately, busy sessions queue for later delivery.
|
|
433
414
|
|
|
@@ -591,37 +572,7 @@ await manager.startSession({
|
|
|
591
572
|
|
|
592
573
|
### Example: Google Antigravity CLI (`agy`)
|
|
593
574
|
|
|
594
|
-
|
|
595
|
-
> instead, which adds conversation resume and timeout coherence the recipe below
|
|
596
|
-
> lacks. This recipe remains as a reference for driving older agy builds or
|
|
597
|
-
> forks with a diverged flag surface:
|
|
598
|
-
|
|
599
|
-
```typescript
|
|
600
|
-
await manager.startSession({
|
|
601
|
-
name: 'antigravity-task',
|
|
602
|
-
engine: 'custom',
|
|
603
|
-
cwd: '/project',
|
|
604
|
-
dangerouslySkipPermissions: true,
|
|
605
|
-
customEngine: {
|
|
606
|
-
name: 'antigravity',
|
|
607
|
-
bin: 'agy', // install: curl -fsSL https://antigravity.google/cli/install.sh | bash
|
|
608
|
-
binEnv: 'AGY_BIN',
|
|
609
|
-
persistent: false,
|
|
610
|
-
args: {
|
|
611
|
-
print: '-p', // single-prompt headless mode
|
|
612
|
-
skipPermissions: '--dangerously-skip-permissions',
|
|
613
|
-
workspace: '--add-dir',
|
|
614
|
-
// NOTE: agy 1.0.2 has NO --output-format flag — output is plain text only.
|
|
615
|
-
// Omitting outputFormat makes the wrapper parse plain text and *estimate*
|
|
616
|
-
// tokens (no real usage / tool-call events). Watch for a JSON output mode.
|
|
617
|
-
},
|
|
618
|
-
},
|
|
619
|
-
});
|
|
620
|
-
```
|
|
621
|
-
|
|
622
|
-
Caveats with `agy` 1.0.2: (1) no structured/stream-json output → token counts are
|
|
623
|
-
estimated, not real; (2) requires a one-time `agy` Google OAuth login; (3) resume
|
|
624
|
-
by conversation ID isn't wired (no JSON stream to capture the ID from).
|
|
575
|
+
Use the built-in [`engine: 'agy'`](#google-antigravity-engine-agy) instead of a custom config.
|
|
625
576
|
|
|
626
577
|
### Custom Engine in Council
|
|
627
578
|
|
|
@@ -649,9 +600,10 @@ manager.councilStart('Build feature X', {
|
|
|
649
600
|
To add a built-in engine (for CLIs that need custom protocol handling beyond what `CustomEngineConfig` supports):
|
|
650
601
|
|
|
651
602
|
1. Create `src/persistent-<engine>-session.ts` implementing `ISession`
|
|
652
|
-
2. Add the engine name to `EngineType` in `src/types.ts`
|
|
653
|
-
3. Add a case to `
|
|
654
|
-
4. Add
|
|
603
|
+
2. Add the engine name to `ENGINE_TYPES` (from which `EngineType` is derived) and its binary to `ENGINE_BINARY_NAMES` in `src/types.ts`
|
|
604
|
+
3. Add a case to `engineHasNativeConversation()` in `src/types.ts` and, if the engine resumes a conversation by id, to `nativeThreadIsLive()` in `src/openai-compat.ts`
|
|
605
|
+
4. Add a case to `SessionManager._createSession()`
|
|
606
|
+
5. Add model pricing to `MODELS[]` in `src/models.ts`
|
|
655
607
|
|
|
656
608
|
The `ISession` interface is deliberately minimal — each engine handles its own subprocess bootstrapping, I/O protocol, and cleanup internally.
|
|
657
609
|
|
|
@@ -4,17 +4,12 @@ Two related surfaces: a durable record of every turn this runtime executes, and
|
|
|
4
4
|
spend cap that is enforced by the runtime rather than by whichever CLI happens to
|
|
5
5
|
support a budget flag.
|
|
6
6
|
|
|
7
|
-
##
|
|
7
|
+
## Purpose
|
|
8
8
|
|
|
9
|
-
`getStats()` / `getCost()` describe a **live** session
|
|
10
|
-
|
|
11
|
-
|
|
12
|
-
|
|
13
|
-
|
|
14
|
-
The same gap made `maxBudgetUsd` a promise the runtime did not keep: it was only
|
|
15
|
-
ever translated into Claude Code's `--max-budget-usd` flag, so a council of Codex
|
|
16
|
-
agents ran with no cap at all. The cap is now applied in `SessionManager`, which
|
|
17
|
-
every engine passes through.
|
|
9
|
+
`getStats()` / `getCost()` describe a **live** session and are held in memory.
|
|
10
|
+
The run ledger is the durable record of every turn across restarts: what ran, on
|
|
11
|
+
which engine, for how much. `maxBudgetUsd` is enforced in `SessionManager`, which
|
|
12
|
+
every engine passes through, not only through Claude Code's `--max-budget-usd`.
|
|
18
13
|
|
|
19
14
|
## The ledger
|
|
20
15
|
|
|
@@ -34,7 +29,7 @@ every engine passes through.
|
|
|
34
29
|
| ----------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
35
30
|
| `ts` | ISO timestamp of turn completion |
|
|
36
31
|
| `session` | SessionManager session name |
|
|
37
|
-
| `engine` | `claude` / `codex` / `codex-app` / `grok` / `opencode` / `agy` / `custom`
|
|
32
|
+
| `engine` | `claude` / `codex` / `codex-app` / `grok` / `opencode` / `agy` / `cursor` (legacy) / `custom` |
|
|
38
33
|
| `model` | Configured model, or the engine's own reported model when none was set (Claude Code: the model named in its `init` event). Absent when neither is known |
|
|
39
34
|
| `cwd` | Working directory the turn ran in |
|
|
40
35
|
| `turn` | 1-based index of the send within the session, counted by the process that recorded it |
|
|
@@ -45,7 +40,7 @@ every engine passes through.
|
|
|
45
40
|
| `toolCalls` / `toolErrors` | Per-turn deltas |
|
|
46
41
|
| `ok` | `false` for a turn that threw, or that the session's own `turnsSucceeded` counter did not count (see `sessions.md`). Falls back to "nothing was thrown" when the counter cannot be read |
|
|
47
42
|
| `error` | Failure text, truncated to 500 chars. Absent when the turn resolved but the engine did not count it as succeeded (an interrupted or non-SUCCESS turn), so a failed row does not always carry one |
|
|
48
|
-
| `parent` | council
|
|
43
|
+
| `parent` | council / fanout / autoloop / workflow run id, when the turn belongs to one |
|
|
49
44
|
|
|
50
45
|
Deltas rather than totals means summing a query window gives that window's spend
|
|
51
46
|
without double-counting.
|
|
@@ -74,7 +69,8 @@ curl "http://127.0.0.1:18796/runs?since=24h&limit=200" -H "Authorization: Bearer
|
|
|
74
69
|
```
|
|
75
70
|
|
|
76
71
|
Returns `{ ok, rows, summary }`, where `summary` carries `rows`, `costUsd`,
|
|
77
|
-
`tokensIn`, `tokensOut`, `estimatedRows`
|
|
72
|
+
`tokensIn`, `tokensOut`, `estimatedRows`, `verifiedRows`, `refutedRows`,
|
|
73
|
+
`unverifiedRows` and a per-engine breakdown `byEngine`. The dashboard
|
|
78
74
|
header shows the 24-hour figure from the same endpoint.
|
|
79
75
|
|
|
80
76
|
Programmatically: `manager.getRunLedger({ since, session, engine, parent, limit })`.
|
|
@@ -106,8 +102,8 @@ Notes:
|
|
|
106
102
|
|
|
107
103
|
### Zeroing a model's pricing
|
|
108
104
|
|
|
109
|
-
`pricingOverrides`
|
|
110
|
-
|
|
105
|
+
`pricingOverrides` sets a model's per-token rate, so the cost figures for a subscription
|
|
106
|
+
seat read 0 instead of an API-rate equivalent. Zeroing a rate also **disables `maxBudgetUsd`
|
|
111
107
|
for that model**: the cap compares the session's accrued `getCost().totalUsd` against it,
|
|
112
108
|
and that number is pricing-derived, so at a rate of 0 it never reaches any cap. There is
|
|
113
109
|
no separate token or turn ceiling behind it.
|
|
@@ -138,16 +134,30 @@ Where the engine reports usage, those counts are the engine's own. Where it does
|
|
|
138
134
|
not, the wrapper falls back to `estimateTokens()` (characters ÷ 4) and the row is
|
|
139
135
|
flagged `tokensEstimated: true`; the CLI marks those costs with a trailing `~`.
|
|
140
136
|
|
|
141
|
-
| Engine | Token counts
|
|
142
|
-
| ----------------- |
|
|
143
|
-
| `claude` | Engine-reported
|
|
144
|
-
| `codex` | Engine-reported
|
|
145
|
-
| `codex-app` | Engine-reported
|
|
146
|
-
| `grok` | Engine-reported — and so is the **cost**: this engine reports `total_cost_usd`, which the wrapper passes through instead of pricing tokens from the registry, so registry drift cannot affect a grok row
|
|
147
|
-
| `cursor` (legacy) | Engine-reported when the stream carries `usage`, else estimated
|
|
148
|
-
| `opencode` | Engine-reported when the run JSON carries `tokens`, else estimated
|
|
149
|
-
| `agy` | Engine-reported when the result event carries usage, else estimated
|
|
150
|
-
| `custom` | Depends on the CLI; estimated when it emits no usage
|
|
137
|
+
| Engine | Token counts |
|
|
138
|
+
| ----------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
139
|
+
| `claude` | Engine-reported, including the **cost** (see below) |
|
|
140
|
+
| `codex` | Engine-reported |
|
|
141
|
+
| `codex-app` | Engine-reported |
|
|
142
|
+
| `grok` | Engine-reported — and so is the **cost**: this engine reports `total_cost_usd`, which the wrapper passes through instead of pricing tokens from the registry, so registry drift cannot affect a grok row |
|
|
143
|
+
| `cursor` (legacy) | Engine-reported when the stream carries `usage`, else estimated |
|
|
144
|
+
| `opencode` | Engine-reported when the run JSON carries `tokens`, else estimated |
|
|
145
|
+
| `agy` | Engine-reported when the result event carries usage, else estimated |
|
|
146
|
+
| `custom` | Depends on the CLI; estimated when it emits no usage |
|
|
147
|
+
|
|
148
|
+
How a `claude` row's cost is taken:
|
|
149
|
+
|
|
150
|
+
- The `result` event's `total_cost_usd` is used as-is, so pricing-table drift
|
|
151
|
+
cannot affect a Claude row.
|
|
152
|
+
- That figure is the session's running total, so a turn's cost is the difference
|
|
153
|
+
from the previous report.
|
|
154
|
+
- A process started with `--resume` (a model switch, a session recovered after a
|
|
155
|
+
restart) reports a total that includes the resumed history. Its first report is
|
|
156
|
+
taken as a baseline, and that one turn is priced from the registry instead, so
|
|
157
|
+
the history is not billed twice (`maxBudgetUsd` gates against this number).
|
|
158
|
+
- Proxy sessions (`baseUrl` set) use the registry estimate instead: the CLI
|
|
159
|
+
believes it is running `opus` while another provider serves the tokens, so its
|
|
160
|
+
own figure would be Opus list price.
|
|
151
161
|
|
|
152
162
|
### What "input tokens" means is not the same on every engine
|
|
153
163
|
|
|
@@ -183,14 +193,13 @@ Cost figures are also only as good as the pricing table: a model missing from
|
|
|
183
193
|
ChatGPT Pro) bill nothing per token while the ledger still reports the API-rate
|
|
184
194
|
equivalent. Read `costUsd` as "what this would cost at API rates".
|
|
185
195
|
|
|
186
|
-
## `ok` vs `verified`
|
|
196
|
+
## `ok` vs `verified`
|
|
187
197
|
|
|
188
|
-
A row
|
|
189
|
-
this section exists to prevent.
|
|
198
|
+
A row carries two different judgements.
|
|
190
199
|
|
|
191
200
|
- **`ok`** — the engine's own terminal verdict for that turn. Codex fails a turn
|
|
192
|
-
that emits `turn.failed` while exiting 0
|
|
193
|
-
|
|
201
|
+
that emits `turn.failed` while exiting 0. It is a careful signal, but it is the
|
|
202
|
+
engine talking about itself.
|
|
194
203
|
- **`verified`** — an acceptance contract ran against the work and every required
|
|
195
204
|
check passed. That is the runtime's own measurement.
|
|
196
205
|
|
|
@@ -210,7 +219,8 @@ either would make the ledger useless for the thing it is for.
|
|
|
210
219
|
`verified` is **not written at turn time**, deliberately. The turns that produce
|
|
211
220
|
the work all finish before the verifier that judges it, so stamping a verdict on
|
|
212
221
|
them as they are written would be inventing one. It is joined in at read time
|
|
213
|
-
from the run record via the row's `parent`, by `annotateVerdicts()`.
|
|
222
|
+
from the run record via the row's `parent`, by `annotateVerdicts()`. The exception
|
|
223
|
+
is a standalone `verify_run`, which writes its verdict directly.
|
|
214
224
|
|
|
215
225
|
Two consequences worth knowing:
|
|
216
226
|
|
|
@@ -219,7 +229,7 @@ Two consequences worth knowing:
|
|
|
219
229
|
- Filtering on `--verified` happens _after_ the join. Pushing the filter into the
|
|
220
230
|
ledger read would match on a field no row carries yet and return nothing.
|
|
221
231
|
|
|
222
|
-
###
|
|
232
|
+
### Verification row fields
|
|
223
233
|
|
|
224
234
|
| Field | Source |
|
|
225
235
|
| -------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|