@enderfga/claw-orchestrator 6.0.4 → 6.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -30,22 +30,22 @@ every engine passes through.
30
30
 
31
31
  ### Row schema
32
32
 
33
- | Field | Meaning |
34
- | ----------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
35
- | `ts` | ISO timestamp of turn completion |
36
- | `session` | SessionManager session name |
37
- | `engine` | `claude` / `codex` / `codex-app` / `grok` / `opencode` / `agy` / `custom` |
38
- | `model` | Configured model, or the engine's own reported model when none was set |
39
- | `cwd` | Working directory the turn ran in |
40
- | `turn` | 1-based turn index within the session |
41
- | `tokensIn` / `tokensOut` / `cachedTokens` | **Per-turn deltas**, not session totals |
33
+ | Field | Meaning |
34
+ | ----------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
35
+ | `ts` | ISO timestamp of turn completion |
36
+ | `session` | SessionManager session name |
37
+ | `engine` | `claude` / `codex` / `codex-app` / `grok` / `opencode` / `agy` / `custom` |
38
+ | `model` | Configured model, or the engine's own reported model when none was set |
39
+ | `cwd` | Working directory the turn ran in |
40
+ | `turn` | 1-based turn index within the session |
41
+ | `tokensIn` / `tokensOut` / `cachedTokens` | **Per-turn deltas**, not session totals |
42
42
  | `costUsd` | Per-turn delta in USD — token count × the registry rate. On a flat-rate plan (a subscription seat) nobody was billed this: set `pricingOverrides` in the plugin config to zero the model out (`{"gpt-5.5": {"input": 0, "output": 0, "cached": 0}}`). Read [Zeroing a model's pricing](#zeroing-a-models-pricing) first — it disables `maxBudgetUsd` for that model, except on `grok`, which prices itself |
43
- | `tokensEstimated` | `true` when the counts came from `estimateTokens()` (see below) |
44
- | `durationMs` | Wall-clock for the turn |
45
- | `toolCalls` / `toolErrors` | Per-turn deltas |
46
- | `ok` | `false` for a turn that threw, or that the session's own `turnsSucceeded` counter did not count (see `sessions.md`). Falls back to "nothing was thrown" when the counter cannot be read |
47
- | `error` | Failure text, truncated to 500 chars. Absent when the turn resolved but the engine did not count it as succeeded (an interrupted or non-SUCCESS turn), so a failed row does not always carry one |
48
- | `parent` | council id / fanout id / autoloop run id, when the turn belongs to one |
43
+ | `tokensEstimated` | `true` when the counts came from `estimateTokens()` (see below) |
44
+ | `durationMs` | Wall-clock for the turn |
45
+ | `toolCalls` / `toolErrors` | Per-turn deltas |
46
+ | `ok` | `false` for a turn that threw, or that the session's own `turnsSucceeded` counter did not count (see `sessions.md`). Falls back to "nothing was thrown" when the counter cannot be read |
47
+ | `error` | Failure text, truncated to 500 chars. Absent when the turn resolved but the engine did not count it as succeeded (an interrupted or non-SUCCESS turn), so a failed row does not always carry one |
48
+ | `parent` | council id / fanout id / autoloop run id, when the turn belongs to one |
49
49
 
50
50
  Deltas rather than totals means summing a query window gives that window's spend
51
51
  without double-counting.
@@ -138,16 +138,41 @@ Where the engine reports usage, those counts are the engine's own. Where it does
138
138
  not, the wrapper falls back to `estimateTokens()` (characters ÷ 4) and the row is
139
139
  flagged `tokensEstimated: true`; the CLI marks those costs with a trailing `~`.
140
140
 
141
- | Engine | Token counts |
142
- | ----------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
143
- | `claude` | Engine-reported |
144
- | `codex` | Engine-reported |
145
- | `codex-app` | Engine-reported |
146
- | `grok` | Engine-reported — and so is the **cost**: this engine reports `total_cost_usd`, which the wrapper passes through instead of pricing tokens from the registry, so registry drift cannot affect a grok row |
147
- | `cursor` (legacy) | Engine-reported when the stream carries `usage`, else estimated |
148
- | `opencode` | Engine-reported when the run JSON carries `tokens`, else estimated |
149
- | `agy` | Engine-reported when the result event carries usage, else estimated |
150
- | `custom` | Depends on the CLI; estimated when it emits no usage |
141
+ | Engine | Token counts |
142
+ | ----------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
143
+ | `claude` | Engine-reported — and so is the **cost**: the `result` event's `total_cost_usd` is taken as-is, so registry drift cannot affect a Claude row. It is the session running total rather than the turn's, so spend advances by the difference between turns. Proxy sessions (`baseUrl` set) are the exception: the CLI is told it is running `opus` while another provider serves the tokens, so its figure is Opus list price for someone else's model and the registry estimate is used instead |
144
+ | `codex` | Engine-reported |
145
+ | `codex-app` | Engine-reported |
146
+ | `grok` | Engine-reported — and so is the **cost**: this engine reports `total_cost_usd`, which the wrapper passes through instead of pricing tokens from the registry, so registry drift cannot affect a grok row |
147
+ | `cursor` (legacy) | Engine-reported when the stream carries `usage`, else estimated |
148
+ | `opencode` | Engine-reported when the run JSON carries `tokens`, else estimated |
149
+ | `agy` | Engine-reported when the result event carries usage, else estimated |
150
+ | `custom` | Depends on the CLI; estimated when it emits no usage |
151
+
152
+ ### What "input tokens" means is not the same on every engine
153
+
154
+ Two engines can report `input` and `cached` and mean different things by them,
155
+ and the difference decides whether the cost math may subtract one from the other:
156
+
157
+ | Engine | Its own arithmetic for one turn | `input` contains cached reads |
158
+ | ---------- | ---------------------------------------------------- | ----------------------------- |
159
+ | `codex` | `total 19704 = input 19699 + output 5` | yes |
160
+ | `grok` | `total 30034 = input 19393 + output 17 + read 10624` | no |
161
+ | `opencode` | `total 26315 = input 58 + output 17 + read 26240` | no |
162
+ | `claude` | `input_tokens 2` against `cache_read 47371` | no |
163
+
164
+ On an engine that excludes them, cached reads and cache writes are billed **on
165
+ top of** `input`, and the prompt the turn actually carried is the sum of all
166
+ three — which is what `contextPercent` measures. Reading `input` alone reports a
167
+ nearly-full context as empty on any resumed conversation, because the history
168
+ arrives as cached reads.
169
+
170
+ Anthropic bills cache _writes_ above the input rate, and the premium depends on
171
+ the TTL (1.25x for the 5-minute cache, 2x for the 1-hour one); the Claude
172
+ estimate prices both tiers from the split the engine reports. Elsewhere cache
173
+ writes are priced at the plain input rate, which is a floor rather than an exact
174
+ figure — one more reason to prefer the engine's own `total_cost_usd` where it
175
+ offers one.
151
176
 
152
177
  So on an estimating engine the cap is best-effort. It will stop a runaway session;
153
178
  it is not an accounting guarantee, and it is not a substitute for the spend limits
@@ -33,7 +33,8 @@ Start a persistent coding session with full CLI flag support.
33
33
  | `mcpConfig` | string \| string[] | MCP server config file(s) |
34
34
  | `settings` | string | Settings.json path or inline JSON |
35
35
  | `ultracode` | boolean | Claude only. Enable "ultracode" / dynamic workflows — Claude plans a JS orchestration script per substantive task and fans out to subagents. Injected as the `ultracode:true` settings key (merged into `settings`), **not** a `--effort` value (the CLI rejects `--effort ultracode`). |
36
- | `noSessionPersistence` | boolean | Do not save session to disk — both the engine's own transcript and this orchestrator's resume registry, so a later start under the same name does not reattach |
36
+ | `noSessionPersistence` | boolean | Do not save session to disk — both the engine's own transcript and this orchestrator's resume registry, so a later start under the same name does not reattach |
37
+ | `ignoreUserConfig` | boolean | Codex only. Run without loading `$CODEX_HOME/config.toml`, so an orchestrated run is decided by what the caller passed rather than by the machine's own Codex config — notably a `model = …` line in that file, which otherwise picks the model while the ledger records this engine's default. Auth still resolves from `CODEX_HOME`. |
37
38
  | `betas` | string \| string[] | Custom beta headers |
38
39
  | `enableAgentTeams` | boolean | Enable experimental agent teams |
39
40
  | `enableAutoMode` | boolean | Enable auto permission mode |
@@ -54,7 +55,7 @@ Start a persistent coding session with full CLI flag support.
54
55
  | `otelLogUserPrompts` | boolean | OpenTelemetry: log user prompts (sets `OTEL_LOG_USER_PROMPTS=1`) |
55
56
  | `otelLogRawApiBodies` | boolean | OpenTelemetry: log raw API request/response bodies (sets `OTEL_LOG_RAW_API_BODIES=1`); debug only |
56
57
  | `bedrockServiceTier` | `'default'` \| `'flex'` \| `'priority'` | AWS Bedrock service tier (sets `ANTHROPIC_BEDROCK_SERVICE_TIER`); only effective when routing through Bedrock |
57
- | `effort` | `'low'` \| `'medium'` \| `'high'` \| `'xhigh'` \| `'max'` \| `'auto'` | Reasoning effort level. `xhigh` is Opus 4.7-only (between `high` and `max`); triggers `ultrathink` prefix on user messages, same as `high` and `max`. |
58
+ | `effort` | `'low'` \| `'medium'` \| `'high'` \| `'xhigh'` \| `'max'` \| `'ultra'` \| `'auto'` | Reasoning effort level. Engines do not share one ladder: Claude Code takes up to `max`, Codex adds `ultra`, Grok stops at `xhigh`, Antigravity at `high`, and OpenCode forwards the level to its provider unvalidated. A level above an engine's ceiling clamps rather than being dropped. `high` and above trigger the `ultrathink` prefix on Claude user messages. |
58
59
 
59
60
  ### `session_send`
60
61