@enderfga/claw-orchestrator 7.5.2 → 7.5.4

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (59) hide show
  1. package/README.md +26 -27
  2. package/configs/engines/README.md +7 -6
  3. package/dist/bin/cli.js +1 -1
  4. package/dist/bin/cli.js.map +1 -1
  5. package/dist/src/acp-server.d.ts +1 -1
  6. package/dist/src/acp-server.js +7 -5
  7. package/dist/src/acp-server.js.map +1 -1
  8. package/dist/src/autoloop/dispatcher.js +3 -3
  9. package/dist/src/autoloop/dispatcher.js.map +1 -1
  10. package/dist/src/autoloop/notify.d.ts +5 -7
  11. package/dist/src/autoloop/notify.js +21 -20
  12. package/dist/src/autoloop/notify.js.map +1 -1
  13. package/dist/src/base-oneshot-session.js +20 -1
  14. package/dist/src/base-oneshot-session.js.map +1 -1
  15. package/dist/src/dashboard/index.html +94 -15
  16. package/dist/src/embedded-server.js +40 -7
  17. package/dist/src/embedded-server.js.map +1 -1
  18. package/dist/src/fanout.d.ts +6 -0
  19. package/dist/src/fanout.js +1 -0
  20. package/dist/src/fanout.js.map +1 -1
  21. package/dist/src/index.js +19 -11
  22. package/dist/src/index.js.map +1 -1
  23. package/dist/src/kernel/nodes/fanout.js +1 -0
  24. package/dist/src/kernel/nodes/fanout.js.map +1 -1
  25. package/dist/src/kernel/types.d.ts +2 -0
  26. package/dist/src/kernel/types.js.map +1 -1
  27. package/dist/src/models.js +36 -7
  28. package/dist/src/models.js.map +1 -1
  29. package/dist/src/openai-compat.d.ts +2 -2
  30. package/dist/src/openai-compat.js +5 -2
  31. package/dist/src/openai-compat.js.map +1 -1
  32. package/dist/src/persistent-agy-session.js +6 -1
  33. package/dist/src/persistent-agy-session.js.map +1 -1
  34. package/dist/src/session-manager.d.ts +1 -0
  35. package/dist/src/session-manager.js +17 -5
  36. package/dist/src/session-manager.js.map +1 -1
  37. package/dist/src/types.d.ts +2 -0
  38. package/openclaw.plugin.json +1 -1
  39. package/package.json +2 -2
  40. package/skills/SKILL.md +31 -32
  41. package/skills/references/acp.md +19 -36
  42. package/skills/references/autoloop.md +163 -180
  43. package/skills/references/claude-cli-tracking.md +28 -27
  44. package/skills/references/cli.md +62 -79
  45. package/skills/references/council.md +40 -63
  46. package/skills/references/dashboard.md +42 -55
  47. package/skills/references/getting-started.md +21 -15
  48. package/skills/references/inbox.md +6 -4
  49. package/skills/references/mcp.md +29 -24
  50. package/skills/references/multi-engine.md +105 -153
  51. package/skills/references/observability.md +42 -32
  52. package/skills/references/openai-compat.md +169 -303
  53. package/skills/references/sessions.md +20 -29
  54. package/skills/references/tools.md +62 -76
  55. package/skills/references/ultra.md +17 -16
  56. package/skills/references/ultraapp.md +59 -64
  57. package/skills/references/verification.md +29 -52
  58. package/skills/references/workflow.md +37 -104
  59. package/skills/ultraapp/SKILL.md +9 -10
@@ -1,8 +1,8 @@
1
1
  # OpenAI-Compatible Bridge
2
2
 
3
- > **Cost warning**: This bridge routes requests through the Claude Code CLI, which uses your Claude Max subscription's **extra usage** quota. When OpenClaw's agent loop sends its system prompt (with distinctive tool definitions and agent instructions), Anthropic's backend recognizes this as programmatic/agent traffic and bills it against extra usage — **not** the included allowance. This is by design: the bridge does NOT bypass Anthropic's billing or subscription enforcement. Using it as OpenClaw's primary model backend means every agent turn consumes extra usage credits at standard API rates ($15/M input, $75/M output for Opus). Monitor your usage at [claude.ai/settings/usage](https://claude.ai/settings/usage).
3
+ > **Cost note**: the bridge drives the Claude Code CLI, so usage counts against your Claude account. It does not bypass Anthropic's billing or subscription limits. Heavy agent traffic can consume extra-usage credits at API rates. Check your usage at [claude.ai/settings/usage](https://claude.ai/settings/usage).
4
4
 
5
- The embedded server exposes a drop-in OpenAI-compatible endpoint so any client that speaks `/v1/chat/completions` can talk to a persistent Claude Code (or Codex / Antigravity / Grok) session. The bridge is designed to serve **two kinds of clients as first-class citizens**:
5
+ The embedded server exposes a drop-in OpenAI-compatible endpoint so any client that speaks `/v1/chat/completions` can talk to a persistent Claude Code (or Codex / Antigravity / Grok / OpenCode) session. The bridge is designed to serve **two kinds of clients as first-class citizens**:
6
6
 
7
7
  1. **Upstream agents** that maintain their own conversation state and forward only the latest user turn — OpenClaw's main agent loop, cron jobs, subagents, programmatic clients.
8
8
  2. **OpenAI-compatible webchat / labeling tools** that re-send the full transcript on every turn — ChatGPT-Next-Web, Open WebUI, LobeChat, data-labeling pipelines.
@@ -11,13 +11,14 @@ Both modes share the same wire protocol; the difference is how a "new conversati
11
11
 
12
12
  ## Endpoint
13
13
 
14
- | | |
15
- | ----------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------- |
16
- | **URL** | `http://127.0.0.1:18796/v1/chat/completions` |
17
- | **Models endpoint** | `GET /v1/models` |
18
- | **Inspection endpoint** | `GET /v1/sessions` (lists active openai-compat sessions with caching stats) |
19
- | **Auth** | Bearer token via `Authorization: Bearer $OPENCLAW_SERVER_TOKEN` (set the env var to enable; otherwise no auth and the server is loopback-only) |
20
- | **Wire format** | OpenAI Chat Completions, both streaming (SSE) and non-streaming |
14
+ | | |
15
+ | ----------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
16
+ | **URL** | `http://127.0.0.1:18796/v1/chat/completions` |
17
+ | **Models endpoint** | `GET /v1/models` |
18
+ | **Inspection endpoint** | `GET /v1/sessions` (lists active openai-compat sessions with caching stats) |
19
+ | **Auth** | `Authorization: Bearer <token>`. By default the server generates a token and writes it to `~/.openclaw/server-token` (mode 0600). Set `OPENCLAW_SERVER_TOKEN=<value>` to choose the token, or `OPENCLAW_SERVER_TOKEN=disabled` to turn auth off (single-user hosts only) |
20
+ | **Wire format** | OpenAI Chat Completions, both streaming (SSE) and non-streaming |
21
+ | **Default model** | `claude-sonnet-4-6` when the request has no `model` |
21
22
 
22
23
  ## Session keying
23
24
 
@@ -28,23 +29,22 @@ Each request is mapped to a long-running session. Once a session exists, subsequ
28
29
  3. **`sys-<sha1(model + systemPrompt)[0..12]>`** — automatic fallback so unkeyed callers don't all collapse onto a single shared session
29
30
  4. **`'default'`** — only when there is no system prompt AND no model (degenerate empty body)
30
31
 
31
- The hash fallback exists because the previous behavior collapsed every unkeyed caller onto one `openai-default` session. In multi-caller setups (OpenClaw routing the main agent + cron jobs + subagents through one gateway) that meant requests serialized against each other and frequently picked up the wrong session's `appendSystemPrompt` — also a privacy leak across distinct callers.
32
+ Without the hash fallback, unkeyed callers (for example OpenClaw's main agent, cron jobs and subagents behind one gateway) would share one session: their requests would queue behind each other and could receive another caller's system prompt.
32
33
 
33
- The model is mixed into the hash so that two callers with the same system prompt but different requested models (e.g. one wants `claude-opus-4-6`, another wants `claude-sonnet-4-6`) don't collide and silently get responses from the wrong model.
34
+ The model is mixed into the hash so that two callers with the same system prompt but different requested models (e.g. one wants `opus`, another wants `sonnet`) get separate sessions rather than responses from the wrong model.
34
35
 
35
36
  The full plugin-side session name is `openai-<key>`. The key becomes a directory
36
37
  name — the bridge starts each session in `os.tmpdir()/openclaw-compat-<name>` —
37
38
  so a key that is not already `[A-Za-z0-9._-]` is replaced by a hash of itself.
38
- A caller using an ordinary id keeps the session name it has always had; one
39
- sending path separators no longer chooses where the session runs.
39
+ An ordinary id is used unchanged; a key containing path separators cannot choose
40
+ where the session runs.
40
41
 
41
42
  The tool list is fingerprinted into the hash fallback by name, a description
42
43
  prefix, and the **parameter schema** (with object keys normalised, so a
43
44
  re-serialised identical schema still resolves to the same session). On the
44
- Claude engine the schemas are baked into the session's system prompt at create
45
- time and deliberately not re-injected per turn, so the key is the only thing
46
- that can notice a schema change — without it, a caller that edited a tool's
47
- parameters kept getting `tool_calls` shaped like the schema it had replaced.
45
+ Claude engine the schemas are written into the session's system prompt when the
46
+ session is created and are not re-sent per turn, so a changed schema has to
47
+ resolve to a new session for the model to see it.
48
48
 
49
49
  ## Operator modes
50
50
 
@@ -96,13 +96,12 @@ mechanism is used depends on whether the engine keeps the conversation itself.
96
96
  | `codex`, `codex-app`, `agy`, `opencode`, `grok` | Full schema block prepended to the message | A short reminder of the calling convention, no schemas — but only once the conversation id has been captured; until then the full block is sent again |
97
97
  | `gemini`, one-shot `custom` | Full schema block prepended to the message | Full schema block again — these have no resume surface, so nothing persists between sends |
98
98
 
99
- The middle row is the one worth understanding. Those engines resume a conversation by id, so
100
- everything injected stays in the transcript. Re-sending the full block each turn
101
- grows the prompt without bound — a 54-tool block runs to roughly 17k tokens, so a
102
- handful of turns is enough to overflow the context window mid-loop and fail the
103
- run outright. Sending _nothing_ on resume turns is not the answer either: the
104
- block also carries the "emit a tool call, do not carry out the work yourself"
105
- framing, and without it the CLI starts doing the work directly.
99
+ The engines in the middle row resume a conversation by id, so everything injected
100
+ stays in the transcript. Re-sending the full block each turn would grow the prompt
101
+ without bound (a 54-tool block is roughly 17k tokens), so a few turns could
102
+ overflow the context window. The short reminder is still needed, because the
103
+ block carries the "emit a tool call, do not carry out the work yourself" framing;
104
+ without it the CLI starts doing the work directly.
106
105
 
107
106
  A fresh session always gets the full block, so a thread is never created without
108
107
  the definitions — including when a session was evicted and is being recreated. A
@@ -115,19 +114,13 @@ is prepended to the message, and skipped only while the conversation it was sent
115
114
  to is still the one being resumed. A turn that creates a conversation always
116
115
  carries it — including a `X-Session-Reset: 1` turn, which stops the existing
117
116
  session and starts a new one. "The conversation is being resumed" means the
118
- engine has actually announced an id (codex's `thread.started`, agy's log, cursor's
119
- and opencode's session id), not merely that the session is in the manager's map: a
120
- first turn that died before announcing one leaves a session behind that resumes
121
- nothing, and those turns keep receiving the full prompt.
122
-
123
- One exception to the sentence above, and it is a defect rather than a design: on a
124
- request whose last non-system message is a `tool` result, `X-Session-Reset` is not
125
- seen at all. The header is parsed after the branch that handles a trailing `tool`
126
- role returns, so on that shape the reset stops nothing, creates nothing, and the
127
- turn is treated as a resumed one — the prompt is skipped if the thread is live.
128
- Every other shape, `[..., tool, user]` included, honors it. A client that needs a
129
- reset mid-tool-loop has to send it on a turn that does not end in a `tool`
130
- message.
117
+ engine has actually announced an id (codex's `thread.started`, agy's `init` event,
118
+ cursor's, grok's and opencode's session id), not merely that the session is in the
119
+ manager's map: a first turn that failed before announcing one leaves a session
120
+ that resumes nothing, and later turns keep receiving the full prompt.
121
+
122
+ See [Known limitations](#known-limitations) for the one request shape where
123
+ `X-Session-Reset` is not honoured.
131
124
 
132
125
  `OPENAI_COMPAT_TOOLS_PER_MESSAGE=1` opts out: it re-sends the full block on every
133
126
  turn for `claude` too, which is what makes a changing tool set work inside one
@@ -138,13 +131,10 @@ does change mid-conversation.
138
131
  ## Conversation history on the way in
139
132
 
140
133
  The caller's `messages[]` can carry the whole conversation: earlier `user` turns and the engine's
141
- own earlier `assistant` replies. Whether those turns need to be sent is the same question the
142
- section below asks about tool results — does the engine's own conversation already hold them? — and
143
- it gets the same answer, from the same predicate. The turns that are in scope are serialized into
144
- one `<conversation_history>` block of `<user>` / `<assistant>` turns and put in front of the
145
- caller's latest `user` text. The wrapper tag is the one `renderHistory()` in the autoloop dispatcher
146
- already uses for the same job; the per-turn tags are not — that one labels its two speakers `<user>`
147
- / `<agent>`, because its roles are autoloop roles rather than OpenAI wire roles.
134
+ own earlier `assistant` replies. The bridge sends those turns only when the engine's conversation
135
+ does not already hold them. The turns that are sent are serialized into one
136
+ `<conversation_history>` block of `<user>` / `<assistant>` turns, placed in front of the caller's
137
+ latest `user` text.
148
138
 
149
139
  | On this turn the engine | What is sent |
150
140
  | ----------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------ |
@@ -152,217 +142,111 @@ already uses for the same job; the per-turn tags are not — that one labels its
152
142
  | **is** resuming a live conversation, but one this bridge never sent these turns to | the same — a live thread is not automatically _this_ thread |
153
143
  | **is** resuming the live conversation these turns belong to | nothing — the turns are already in the transcript, and the text goes out alone |
154
144
 
155
- The last row is what keeps Anthropic prompt caching warm on `claude` and keeps a resumed `codex`
156
- thread from being re-fed its own history: on a live thread the message is byte-identical to what it
157
- was before this block existed. The first row is the one that was losing data. A client that opens a
158
- new conversation per turn — one whose session key hashes the last message, say — lands in it on
159
- every turn, and `skipPersistence: true` means an OpenAI-compat session is never auto-resumed from
160
- disk either, so a follow-up like "yes, go ahead" used to reach the engine with nothing in front of
161
- it.
162
-
163
- The middle row is the one that is easy to get wrong. "Is there a live thread under this session
164
- name?" is not "is that thread holding this conversation?", and the two come apart constantly. A
165
- caller whose session key hashes its latest message resolves every repeat of a short confirmation —
166
- "yes, go ahead", typed all day by someone approving invoices — to the session some _earlier_
167
- invoice opened. A client that sends no session key at all falls back to a hash of
168
- model+system+tools, which is stable for every turn AND identical between two concurrent chats of
169
- the same user, so both chats share one session. And on `claude`, the default engine,
170
- `nativeThreadIsLive()` has no id to check and returns true for anything in the session map, so
171
- there the question collapses to "does the name exist" with no gate at all.
172
-
173
- So the bridge records, per session, a fingerprint of the `user` turns it has actually pushed there,
174
- and replays unless this request continues exactly that. It is the only writer to these sessions, so
175
- what it sent is what the engine holds. Unknown session, forked conversation, evicted entry: all
176
- replay. Being wrong in that direction costs a duplicated turn under a framing that says not to act
177
- on it twice; being wrong in the other direction drops context silently, which is the class of bug
178
- this exists to fix.
179
-
180
- What the request is compared against depends on where its array ends, and only an array ending in a
181
- `user` turn carries a turn the bridge has not sent yet. For that shape the fingerprint of every
182
- `user` turn _before_ the last one is what the thread should be holding. For every other shape — a
183
- tool-loop hop ending in `tool`, a prefill/continue ending in `assistant` — the latest `user` turn
184
- was already pushed on an earlier request, so the whole array is compared. Getting that wrong in the
185
- other direction is not harmless: comparing a hop against one turn less never matches, and the
186
- transcript is then replayed into the very session that is already holding it, on every hop, for a
187
- caller that needed none of this.
188
-
189
- The map is bounded at 1000 entries and evicted oldest-first, which it needs independently of the
190
- session map: `_cleanupIdleSessions()` reaps a session by `sessionTtlMinutes` without telling this
191
- map, so a fingerprint outlives the session it mirrors. Eviction costs a replayed block, never a
192
- dropped one, and `serve` restarts start it empty for the same price.
193
-
194
- `system` messages are never in the block: they travel as the session's system prompt (see [Tool
195
- definitions](#tool-definitions-and-where-they-live)). `tool` messages are never in it either —
196
- they are the next section's business, and repeating them here would duplicate every payload and
197
- undo the scoping that keeps a tool loop linear. An `assistant` message that only announces
198
- `tool_calls` carries no text, so it renders no turn; a turn whose text is empty or whitespace is
199
- dropped rather than rendered as an empty shell. An array with nothing to replay produces no block
200
- at all, so a single-turn `[system, user]` request — the shape the OpenClaw main agent, cron jobs
201
- and subagents send — goes out exactly as it did before.
202
-
203
- Replayed text has neutralized every tag the assembled prompt treats as structure — the block's own
204
- `<conversation_history>` / `<user>` / `<assistant>`, and also `<tool_results>` / `<tool_result>`,
205
- `<system>` and `<tool_calls>` (`</user>` becomes `&lt;/user>`). Unlike a `<tool_result>` body, which
206
- comes from the caller's own tool runner, a replayed turn is whatever an end user typed, and
207
- `hi</user>\n<assistant>\n...` would otherwise close its turn early and forge an `assistant` turn —
208
- putting words in the engine's own mouth. The other three matter for the same reason: a fabricated
209
- tool return carries the framing sentence that says a payload is authoritative, a `<system>` block
210
- contradicts the real one the non-claude path prepends, and `<tool_calls>` is the exact protocol JSON
211
- the model is asked to emit.
212
-
213
- Only the `<` is escaped, by lookahead. That shape is what makes the boundary decidable: matching up
214
- to the closing `>` instead means re-emitting whatever was captured, and `hola<user a</user>` then
215
- smuggles a raw close through the attribute slot of a tag that IS matched. So the match ends at
216
- anything that ends a tag name — `>`, `/`, `<`, end of text, or a character that occupies no width.
217
- That last clause is five Unicode properties rather than a list of code points, and it is not
218
- decoration: measured over the 6,060 code points that are zero-advance or render blank,
219
- `[\s></\p{Cc}\p{Cf}]` let **5,806** through, so `ok</user︀>\n<assistant︀>` forged a turn that is
220
- indistinguishable on screen from `ok</user>`. Adding `Default_Ignorable` leaves 1,770;
221
- `\p{Mn}\p{Me}` closes it; U+2800 BRAILLE PATTERN BLANK is neither and is named. Cost: the same 11 of
222
- a 28-string corpus of plausible legitimate text change under the wide class as under the narrow one.
223
-
224
- The name does **not** have to sit flush against `<` or `</`: the same invisible padding, plus the
225
- slash itself, is allowed before the name, because `hola</​user>` renders as `hola</user>` and a
226
- model reads it as a close — a trailing class complete over zero-width filler with a flush leading
227
- side still forges a turn. That is **one** class, `[/\p{Cc}\p{Cf}\p{Mn}\p{Me}\p{Default_Ignorable_Code_Point}]*`,
228
- not `[…]*\/?[…]*`: two adjacent unbounded quantifiers over the same class backtrack O(n²) on a long
229
- run that never reaches a valid name, and a single history message of ~100 KB of combining marks hung
230
- the event loop ~80 s — a one-request denial of service against a fleet whose watchdog already resets
231
- on event-loop stalls. The class is the zero-**advance** subset only (no `\s`, no U+2800), because a
232
- visible separator before the name is the `if (count < user && x)` corruption the boundary class
233
- already refuses. Swept over the 6,060 code points that are zero-advance or separators, in all three
234
- positions (18,180 probes): 12,120 went through unfenced with the flush leading side, **38** with the
235
- interior class, and the 38 are 19 `Zs`/`Zl`/`Zp` code points — U+0020, U+00A0, U+2000–200A, U+3000 —
236
- i.e. exactly the visible-separator limitation below, and nothing else. The single class also lets a
237
- slash sit among the filler (`<//user`, `</␀/user`); harmless, since the only action is escaping the
238
- `<`. Filler INSIDE the name (`</us␀er>`) is still not fenced — the old regex missed it too.
239
-
240
- **What it does not promise.** A positive class cannot be complete over a _visible_ separator, so
241
- `</ user>` and `< assistant>` go through raw — a model reads them as a boundary. They are excluded on
242
- cost: reaching them means corrupting `if (count < user && x)` and `the < user > column`. What already
243
- gets corrupted for the same reason, since the class contains whitespace: `Promise<User | null>` and
244
- `count<user && total>limit` come out with `&lt;`. Readable to a model, not byte-identical. And the
245
- body of a `<tool_result>` is never fenced — its content comes from the caller's own tool runner — so a
246
- tool return that embeds a whole `<conversation_history>` block is not stopped here.
247
-
248
- The caller's **latest** `user` turn is fenced too, but only on turns that actually carry a block.
249
- That turn is the one input an attacker controls end to end, and the block teaches the model in the
250
- same prompt that `<conversation_history>` holds its own earlier turns — unfenced, it could close the
251
- real block and open a second one indistinguishable from it. With no block in front of it there is
252
- nothing to forge, so those turns stay byte-for-byte what they were.
253
-
254
- A `user` turn carrying only non-text content (an image) renders as `[non-text content]` rather than
255
- vanishing. Dropping it would leave the `assistant` reply to it standing alone under a framing that
256
- calls the assistant turns the model's own — a reply to a request the model cannot see, which reads
257
- as license to act on the reply by itself. "Photo of the invoice", then "yes, go ahead", is an
258
- everyday shape. A leading `assistant` turn with no `user` turn in front of it at all (content
259
- `null`, or empty) is dropped instead, since no marker could honestly stand in for it.
145
+ The last row keeps Anthropic prompt caching warm on `claude` and stops a resumed `codex` thread
146
+ from being sent its own history again: on a live thread the message is exactly the caller's text.
147
+ The first row covers clients that open a new conversation every turn (for example, one whose
148
+ session key hashes the last message). Bridge sessions are created with `skipPersistence: true` and
149
+ are never resumed from disk, so without the replay a follow-up like "yes, go ahead" would reach
150
+ the engine with no context. How the bridge tells the second row from the third is described in
151
+ [A live session is not the same thing as this conversation](#a-live-session-is-not-the-same-thing-as-this-conversation).
152
+
153
+ What goes into the block:
154
+
155
+ - `system` messages are never in it: they travel as the session's system prompt (see
156
+ [Tool definitions](#tool-definitions-and-where-they-live)).
157
+ - `tool` messages are never in it either; they are handled by the `<tool_results>` block described
158
+ under [Tool results on the way back](#tool-results-on-the-way-back).
159
+ - An `assistant` message that only announces `tool_calls` carries no text, so it renders no turn. A
160
+ turn whose text is empty or whitespace is dropped.
161
+ - A `user` turn carrying only non-text content (an image) renders as `[non-text content]`, so the
162
+ `assistant` reply to it does not appear to answer nothing. A leading `assistant` turn with no
163
+ `user` turn before it (content `null`, or empty) is dropped.
164
+ - An array with nothing to replay produces no block, so a single-turn `[system, user]` request —
165
+ the shape the OpenClaw main agent, cron jobs and subagents send — goes out unchanged.
166
+
167
+ Replayed text has the `<` of every structural tag escaped (`</user>` becomes `&lt;/user>`): the
168
+ block's own `<conversation_history>` / `<user>` / `<assistant>`, and also `<tool_results>` /
169
+ `<tool_result>`, `<system>` and `<tool_calls>`. Forms padded with invisible or zero-width
170
+ characters (`</​user>`, `<user︀>`) are caught too. This stops a user message from closing its own
171
+ turn or forging an `assistant` turn, a `<system>` block, a tool result, or a tool call. Limits:
172
+
173
+ - A visible space inside the tag (`</ user>`, `< assistant>`) is not caught, and filler inside the
174
+ tag name (`</us␀er>`) is not caught either.
175
+ - Text that merely looks like a tag, such as `Promise<User | null>`, comes out with `&lt;`. It stays
176
+ readable to a model but is not byte-identical.
177
+ - The body of a `<tool_result>` is never escaped, since it comes from the caller's own tool runner.
178
+
179
+ The caller's **latest** `user` turn is escaped the same way, but only on turns that carry a block.
180
+ With no block in front of it there is nothing to forge, so those turns go out unchanged.
260
181
 
261
182
  ### What this does not cover
262
183
 
263
- - **A turn can be replayed that the engine already had.** The mirror image of the duplicate-once
264
- trade below, and it comes from the same place: the engine's state is read from the session, not
265
- from the array. A client whose session looks new to the bridge but whose engine did hold context
266
- gets those turns a second time. The framing sentences tell the model these are earlier turns and
267
- not to act on them again, which is what keeps a duplicated turn from becoming a duplicated
268
- action; nothing enforces it.
184
+ - **A turn can be replayed that the engine already had.** The engine's state is read from the
185
+ session, not from the array, so a client whose session looks new to the bridge but whose engine
186
+ did hold context gets those turns a second time. The framing tells the model these are earlier
187
+ turns and not to act on them again; nothing enforces it.
269
188
  - **The block is not in strict chronological order when the array does not end in the caller's
270
189
  latest `user` turn.** Every `user`/`assistant` turn except that one is replayed, including turns
271
- that come after it — an array ending in `assistant` (prefill, an explicit "continue", a framework
272
- appending its own reply) keeps that turn, because a transcript presented as complete while
273
- missing the last thing the model said invites it to redo the work. The cost is that such a turn
274
- is rendered inside the block, i.e. before the caller's latest text rather than after it.
275
- - **`X-Session-Reset` is not honored on an array ending in a `tool` result**, so on that one shape a
276
- reset turn on a live thread gets no history block either. Same pre-existing asymmetry the tool
277
- results have there, from the same cause — the header is parsed after that branch returns. See the
278
- note under [Tool definitions](#tool-definitions-and-where-they-live).
279
- - **`grok` inherits the hole described in the next section**: it resumes by id, but its id is absent
280
- from `SessionStats`, so `nativeThreadIsLive()` reaches `default: return true` and a `grok` session
281
- whose first turn died before emitting its id reports a live thread — and so gets no history. Same
282
- cause, same fix (adding the field), not addressed here.
190
+ after it — an array ending in `assistant` (prefill, an explicit "continue") keeps that turn, so it
191
+ is rendered inside the block, before the caller's latest text.
283
192
  - **The two blocks are not interleaved.** When a request carries both, the message is the history
284
- block, then the tool results, then the caller's new text — chronological between the blocks, but
285
- a `tool` result that chronologically preceded a replayed `assistant` turn still appears after it.
286
- - **The block is capped at 24,000 characters — the whole block, not the sum of the turn text.**
287
- Wrapper tags, per-turn tags, elision markers and the 268 characters of framing are all charged to
288
- the budget before any turn is. That is the fix for a cap that did not cap: charging only the turn
289
- text left the retained turn count bounded by nothing but `24,000 / mean-turn-length`, and 8,000
290
- alternating one-word turns (`ok`, `sí`, `dale`) rendered a **165,008-byte** block — 33 KiB past the
291
- argv ceiling below, from a request no bigger than a chat backlog. The elision markers were worse:
292
- 32 characters each, added AFTER the arithmetic, once per truncated turn, which is how a "24,000
293
- cap" emitted more than it said.
294
-
295
- Oldest turns are dropped first, and the turn the budget runs out inside is truncated (head kept)
296
- and marked `[… turn truncated for length …]` rather than dropped whole — dropping it whole would
297
- take the request with it on the pasted-document shape. The marker comes out of that turn's own
298
- allowance, so a truncated turn cannot push the block over. Below 200 rendered characters of
299
- remaining room a turn is not started at all: it and the turns behind it are dropped, and the
300
- remainder goes unspent, because a fragment that short is a sentence with its qualifier cut off
301
- rather than context.
302
-
303
- The newest `user` turn in the block — the ask the turns after it answer — is the one exception to
304
- spending newest-first: the turns after it (all `assistant`, by definition of "newest `user` turn")
305
- spend against the budget minus its frame plus `min(its length, 200)`. That reserve is the fix for
306
- a drop, not a refinement. Without it, replies that add past the cap (one 30k pasted listing, or
307
- two ordinary 12k ones) take the whole budget, the window starts past every `user` turn, the
308
- leading-`assistant` rule clears what is left, and NO block goes out at all: the caller's latest
309
- turn reaches the engine alone, which is this block's own failure mode at its worst.
310
-
311
- The 200-character floor applies **above** the anchor too, and that is the second half of the cap
312
- fix. Once the post-anchor turns have truncated the budget down near the reserve, an older turn is
313
- started with a `room` smaller than the 32-character elision marker, and `slice(0, room - 32)` with
314
- a negative argument slices from the END of the string — emitting nearly the whole turn while
315
- charging the budget only `room`, which blows the cap. Measured on a narrating tool loop, where each
316
- hop's `assistant` narration becomes a consecutive post-anchor turn because `tool` messages are
317
- filtered out: with the floor removed, 30 hops render 27,627 characters and the three-`user` /
318
- with-reply sweep shapes reach 28,003 / 30,003. With the floor, that run of older turns is dropped
319
- instead, which the anchor's reserve makes safe: ending the window there would drop the anchor and
320
- hand the leading-`assistant` rule an all-`assistant` list to clear, i.e. the empty block again.
321
-
322
- Verification. 44,000 random shapes across three message-count ranges (≤7, ≤60 and ≤400 messages,
323
- roles `user`/`assistant`/`system`/`tool`, content string/whitespace/null/array/empty-array/no-text,
324
- lengths straddling 0/1/199/200/201/11,999/12,000/12,001/23,799/23,800/24,000/24,001/30,000/60,000):
325
- **max block 23,999 characters, zero shapes over 24,000, zero content-free turns, zero blocks with
326
- no `user` turn in them, and zero shapes that went from a non-empty block to an empty one.** 7,488
327
- of them went the other way — empty before, non-empty now — which is the anchor reserve doing its
328
- job. 11,954 non-empty blocks changed, which is the point: the old arithmetic charged less than it
329
- emitted, so every block near the ceiling gets shorter. Directed shapes: 30/100/2,000 narrating
330
- hops and 1,200/8,000/24,000 one-word turns are all ≤ 24,000 with no content-free turns.
331
-
332
- The ceiling is on characters and argv counts bytes, which is the conservative direction only up to
333
- a point: 24,000 characters of astral-plane text is 48,000 bytes and the worst case (3-byte BMP) is
334
- 72,000, both still inside 128 KiB, but the
335
- block is not the whole prompt — the tool block, the `<system>` prepend and the caller's own turn
336
- are added after it. What the cap bounds is the part that scales with the transcript.
337
-
338
- `MAX_BODY_SIZE` (5 MiB) is **not** a usable bound here: six of the nine `ENGINE_TYPES` pass the
339
- prompt to the CLI as a single argv element (`codex`, `gemini`, `agy`, `cursor`, `grok`,
340
- `opencode`), and so does a one-shot `custom` engine; Linux caps one argument at `MAX_ARG_STRLEN` =
341
- 128 KiB whatever `getconf ARG_MAX` reports (measured: 131071 bytes spawns, 131072 throws `E2BIG`).
342
- Only `claude`, `codex-app` and a persistent `custom` engine write over stdin. Uncapped, ordinary
343
- traffic reaches that ceiling — 400-character turns with 900-character replies put the message at
344
- 131,063 characters at turn 96 and 132,425 at turn 97 — and the failure is a 500 with the turn
345
- lost, which is worse than the missing context the block exists to restore. This is the same trade
346
- `renderHistory()` makes, with the same `REPLAY_CHAR_BUDGET`, feeding the same engines.
347
- - **A send that threw records nothing.** The fingerprint is written after the send and only when the
348
- send landed, because the two ways to be wrong are not symmetric: forgetting a turn that landed
349
- replays it once more, while assuming one landed that did not drops context silently. "Landed" is
350
- `sendMessage` returning — a returned error is answered with 502 and still records, since the CLI
351
- received the prompt — so only a throw withholds the record. An earlier revision of this file
352
- described the opposite placement as deliberate and named its residue: a caller that answered a 5xx
353
- by appending an `assistant` turn and sending again got the next turn bare. That was the bug, not
354
- the design.
193
+ block, then the tool results, then the caller's new text. A `tool` result that chronologically
194
+ preceded a replayed `assistant` turn still appears after it.
195
+ - **The block is capped at 24,000 characters, counting tags, markers and framing.** Oldest turns are
196
+ dropped first. The turn where the budget runs out is cut (start kept) and marked
197
+ `[… turn truncated for length …]`. A turn with under 200 characters of room is dropped instead.
198
+ The most recent `user` turn in the block always keeps at least its first 200 characters, so the
199
+ block is never left with only `assistant` turns.
200
+
201
+ The cap exists because most engines (`codex`, `agy`, `grok`, `opencode`, `cursor`, `gemini`,
202
+ one-shot `custom`) receive the prompt as a single command-line argument, and Linux limits one
203
+ argument to 128 KiB (`MAX_ARG_STRLEN`). Going over fails the request with a 500. The 5 MiB request
204
+ body limit does not protect against this. `claude`, `codex-app` and a persistent `custom` engine
205
+ write over stdin instead. The cap bounds only the part of the prompt that grows with the
206
+ transcript; the tool block, the system prompt and the caller's own turn are added on top.
207
+
208
+ - **A send that threw records nothing.** The fingerprint is written only after the send returns. A
209
+ send that returns an error is answered with 502 and still records, since the CLI received the
210
+ prompt; only a thrown send leaves no record, so the next request replays the turns.
355
211
  - **A second request that arrives while the first is still in flight replays.** It sees no
356
- fingerprint yet, so the transcript goes out again into the session that already holds it. The safe
357
- direction — a duplicate rather than a drop — plus a lost cache prefix.
358
- - **Cost is O(n) per turn for engines that never resume.** `engineHasNativeConversation()` is false
359
- for `gemini` and one-shot custom engines, so for them the block is re-serialized on every turn,
360
- capped but never free. Same for any caller that mints a new session per turn.
361
- - **`X-Session-Reset` now replays the transcript.** A reset turn means the engine holds nothing, so
362
- the history goes out in full. Under the reading "the caller asked to start clean" that is the
363
- opposite of what was asked, and a client that sends the header on every request AND re-sends
212
+ fingerprint yet, so the transcript goes out again into the session that already holds it: a
213
+ duplicate rather than a loss, plus a lost cache prefix.
214
+ - **Cost is O(n) per turn for engines that never resume.** `gemini` and one-shot custom engines have
215
+ no native conversation, so the block is rebuilt and sent on every turn (capped). The same applies
216
+ to any caller that creates a new session per turn.
217
+ - **`X-Session-Reset` replays the transcript.** A reset turn means the engine holds nothing, so the
218
+ history goes out in full. A client that sends the header on every request and also re-sends
364
219
  `messages[]` pays for the transcript every time.
365
220
 
221
+ ### A live session is not the same thing as this conversation
222
+
223
+ Suppressing the replay needs a stronger fact than "a session under this name is live". That is what
224
+ `nativeThreadIsLive()` reports, and a session name can be live while its transcript belongs to a
225
+ different exchange. Three shapes where the two come apart, all reachable with default settings:
226
+
227
+ | shape | what happens |
228
+ | ---------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------- |
229
+ | a caller whose session key hashes its latest message | every repeat of the same short confirmation resolves to whichever session that phrase opened first — often a different subject entirely |
230
+ | a caller that sends no `X-Session-Id` at all | the key falls back to a hash of model + system prompt + tools, so all of that caller's concurrent chats share one name |
231
+ | `engine: 'claude'` — the default | `nativeThreadIsLive()` has no id to check and returns `true` for anything in the session map, so the name is the only evidence there is |
232
+
233
+ So the bridge records, per session, a fingerprint of the `user` turns it has sent there, and
234
+ replays whenever the incoming conversation is not the one it recorded. It is the only writer to
235
+ these sessions, so what it sent is what the engine holds. The fingerprint covers the `user` turns
236
+ only: those are the caller's own text echoed back verbatim, while assistant text is what the engine
237
+ produced and a client may normalize it. A mismatch replays: the cost is a repeated block, never a
238
+ lost one.
239
+
240
+ What the request is compared against depends on how the array ends. For an array ending in a `user`
241
+ turn, the `user` turns before that last one are compared, since the last one has not been sent yet.
242
+ For any other ending (a tool-loop hop ending in `tool`, a prefill ending in `assistant`) the latest
243
+ `user` turn was already sent, so all `user` turns are compared.
244
+
245
+ The fingerprint map holds at most 1,000 entries, evicted oldest-first. It is separate from the
246
+ session map: `_cleanupIdleSessions()` reaps idle sessions by TTL without updating it, so a
247
+ fingerprint can outlive its session. Losing an entry (eviction, or a `serve` restart, which starts
248
+ the map empty) costs a replayed block, never a dropped one.
249
+
366
250
  ## Tool results on the way back
367
251
 
368
252
  A `tool` role message in the caller's array is the result of a call the model asked for on an
@@ -380,15 +264,9 @@ array:
380
264
  Whether the engine is resuming is resolved with the same `nativeThreadIsLive()` check as the
381
265
  middle row of [Tool definitions](#tool-definitions-and-where-they-live). Engines with no resume
382
266
  surface (`gemini`, one-shot `custom`) are never in the second row, because
383
- `engineHasNativeConversation()` gates the check.
384
-
385
- Two engines reach the second row from map presence alone rather than from a captured id, because
386
- `nativeThreadIsLive()` has no case for them and its `default` arm returns `true`: `claude`, which
387
- holds its context in a live process and has no separate id to check, and `grok` — which does resume
388
- by id (`--resume <sessionUUID>`), but whose id is absent from `SessionStats`, so there is nothing to
389
- check even though there is something to check for. A `grok` session whose first turn died before
390
- emitting its id therefore reports a live thread and gets its earlier rounds scoped away. That is
391
- pre-existing and not addressed here; fixing it means adding the field.
267
+ `engineHasNativeConversation()` gates the check. `claude` (and a persistent `custom` engine) holds
268
+ its context in a live process and has no separate id to check, so it is in the second row whenever
269
+ the session exists.
392
270
 
393
271
  The trailing role of the array does not enter into it — `[..., tool]`, `[..., tool, user]` and
394
272
  `[..., tool, assistant]` are read the same way, and the caller's latest `user` text, when there is
@@ -397,70 +275,56 @@ all — a multimodal content array holding only an image — gets the block as t
397
275
  array carrying no `tool` message at all is untouched: no block, no wrapper, the message goes as it
398
276
  came.
399
277
 
400
- The third case in the first row — a session being stopped and recreated — is `X-Session-Reset`, and
401
- it puts the turn in that row on every shape where the header is read at all. That excludes an array
402
- ending in a `tool` result, where the header is not seen; see the note under [Tool definitions](#tool-definitions-and-where-they-live).
278
+ The third case in the first row — a session being stopped and recreated — is `X-Session-Reset`,
279
+ except on the one shape listed under [Known limitations](#known-limitations).
403
280
 
404
281
  ### What the scoping is for, and what it does not cover
405
282
 
406
- On a resumed conversation the scoping is what keeps a tool loop linear instead of quadratic. With a
283
+ On a resumed conversation the scoping keeps a tool loop linear instead of quadratic. With a
407
284
  30k-character batch per round, the tenth hop carries ~30k characters of results instead of the
408
285
  ~300k the engine has already seen. Two properties are worth checking against your own client before
409
286
  relying on it:
410
287
 
411
- - **The boundary is the last `assistant` message, and what matters is whether one sits AFTER the
288
+ - **The boundary is the last `assistant` message, and what matters is whether one sits after the
412
289
  earliest unsent `tool` message** — not whether the array contains one at all. The OpenAI wire
413
290
  format has the caller echo the `assistant` turn that carried the `tool_calls` ahead of the
414
291
  matching `tool` messages, and a client that echoes it pays for one round per hop. When no
415
- `assistant` message follows the earliest unsent result, `lastIndexOf('assistant')` is behind them
416
- all and the slice keeps everything, so the scoping is a no-op. Two shapes land there: an array
417
- with no `assistant` message at all, and one `assistant` announcing N parallel calls followed by
418
- its N results — the second is a no-op that costs nothing, since those N results _are_ one round.
292
+ `assistant` message follows the earliest unsent result, the slice keeps everything, so the scoping
293
+ has no effect. Two shapes land there: an array with no `assistant` message at all, and one
294
+ `assistant` announcing N parallel calls followed by its N results — the second costs nothing,
295
+ since those N results _are_ one round.
419
296
  - **A round can go out twice.** Engine replies are not read back out of the array, so a round the
420
297
  caller did not record an `assistant` turn for looks the same as a round the engine never saw, and
421
298
  it goes out again. The duplication is bounded to one round wherever the scoping runs — i.e. on a
422
299
  resumed conversation whose array does carry an `assistant` message. It is unbounded in the two
423
- cases where nothing is scoped: the row above, and any turn in the first row of the table (no
424
- resumed conversation), where the whole array goes out by design because the engine holds none of
425
- it.
300
+ cases where nothing is scoped: the shape above with no `assistant` message, and any turn in the
301
+ first row of the table (no resumed conversation), where the whole array goes out because the
302
+ engine holds none of it.
426
303
 
427
304
  Neither applies to a caller that keeps its own transcript and forwards only the latest turn — it
428
305
  sends one round at a time.
429
306
 
430
- ### A live session is not the same thing as this conversation
431
-
432
- Suppressing the replay needs a stronger fact than "a session under this name is live". That is what
433
- `nativeThreadIsLive()` reports, and a session name can be live while its transcript belongs to a
434
- different exchange. Three shapes where the two come apart, all reachable with default settings:
435
-
436
- | shape | what happens |
437
- | ---------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------- |
438
- | a caller whose session key hashes its latest message | every repeat of the same short confirmation resolves to whichever session that phrase opened first — often a different subject entirely |
439
- | a caller that sends no `X-Session-Id` at all | the key falls back to a hash of model + system prompt + tools, so all of that caller's concurrent chats share one name |
440
- | `engine: 'claude'` — the default | `nativeThreadIsLive()` has no id to check and returns `true` for anything in the session map, so the name is the only evidence there is |
441
-
442
- So the bridge tracks what it actually pushed into each session and replays whenever the incoming
443
- conversation is not the one it remembers seeding. The fingerprint covers the `user` turns only: those
444
- are the caller's own text echoed back verbatim, while assistant text is what the engine produced and
445
- a client may normalize it. A mismatch replays, which is the safe direction — the cost is a repeated
446
- block, never a lost one.
307
+ ## Known limitations
447
308
 
448
- The map is bounded and evicted oldest-first, and it has to be independent of the session map:
449
- `_cleanupIdleSessions()` reaps a session by TTL without telling it, so a fingerprint outlives the
450
- session it mirrors. Losing an entry costs a replayed block, never a dropped one.
309
+ - **`X-Session-Reset` is ignored on a request whose last non-system message is a `tool` result.** On
310
+ that shape the reset stops nothing and creates nothing, and the turn is treated as a resumed one:
311
+ the system prompt is skipped if the thread is live, and tool results are scoped as for a live
312
+ thread. Every other shape,
313
+ `[..., tool, user]` included, honours it. To reset in the middle of a tool loop, send the header on
314
+ a turn that ends in a `user` message.
451
315
 
452
316
  ## Environment variables
453
317
 
454
- | Variable | Default | Purpose |
455
- | ----------------------------------- | --------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
456
- | `OPENCLAW_SERVER_TOKEN` | (unset) | Bearer token for HTTP auth. Set to enable; written to `~/.openclaw/server-token` for the CLI. |
457
- | `OPENCLAW_RATE_LIMIT` | `300` | Max requests per IP per 60-second sliding window. |
458
- | `OPENCLAW_CORS_ORIGINS` | (loopback only) | Set to `*` to allow all origins (the `/v1/*` paths already do this). |
459
- | `OPENAI_COMPAT_NEW_CONVO_HEURISTIC` | (unset) | Set to `1` to enable webchat mode (see above). |
460
- | `OPENAI_COMPAT_TOOLS_PER_MESSAGE` | (unset) | Set to `1` to re-send the full tool schemas on every turn (see [Tool definitions](#tool-definitions-and-where-they-live)). Needed only when the tool set changes mid-conversation; costs per-turn prompt growth. |
461
- | `OPENAI_COMPAT_STATUS_URL` | (unset) | If set, the bridge POSTs JSON status updates to this URL (fire-and-forget, 2s timeout). See [Status webhook](#status-webhook). |
462
- | `OPENCLAW_SERVE_MAX_SESSIONS` | `32` | Max concurrent OpenAI-compat sessions in serve mode. Bumped from the in-plugin default of 5 because each distinct caller now gets its own `sys-<hash>` session. |
463
- | `OPENCLAW_SERVE_TTL_MINUTES` | `60` | Idle TTL for OpenAI-compat sessions in serve mode. Idle sessions are reaped by a 60s background loop; persisted disk registry is kept for 7 days so a returning caller is auto-resumed. |
318
+ | Variable | Default | Purpose |
319
+ | ----------------------------------- | ---------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
320
+ | `OPENCLAW_SERVER_TOKEN` | (auto-generated) | Overrides the auto-generated bearer token; `disabled` turns auth off. The active token is written to `~/.openclaw/server-token` for the CLI. |
321
+ | `OPENCLAW_RATE_LIMIT` | `300` | Max requests per IP per 60-second sliding window. |
322
+ | `OPENCLAW_CORS_ORIGINS` | (loopback only) | Set to `*` to allow all origins (the `/v1/*` paths already do this). |
323
+ | `OPENAI_COMPAT_NEW_CONVO_HEURISTIC` | (unset) | Set to `1` to enable webchat mode (see above). |
324
+ | `OPENAI_COMPAT_TOOLS_PER_MESSAGE` | (unset) | Set to `1` to re-send the full tool schemas on every turn (see [Tool definitions](#tool-definitions-and-where-they-live)). Needed only when the tool set changes mid-conversation; costs per-turn prompt growth. |
325
+ | `OPENAI_COMPAT_STATUS_URL` | (unset) | If set, the bridge POSTs JSON status updates to this URL (fire-and-forget, 2s timeout). See [Status webhook](#status-webhook). |
326
+ | `OPENCLAW_SERVE_MAX_SESSIONS` | `32` | Max concurrent OpenAI-compat sessions in serve mode. The plugin default is 5; serve mode raises it because each distinct caller gets its own `sys-<hash>` session. |
327
+ | `OPENCLAW_SERVE_TTL_MINUTES` | `60` | Idle TTL for OpenAI-compat sessions in serve mode. Idle sessions are reaped by a 60s background loop and are not resumed from disk. |
464
328
 
465
329
  ## Inspection: `GET /v1/sessions`
466
330
 
@@ -480,7 +344,7 @@ Sample response:
480
344
  {
481
345
  "key": "sys-a3f81c9d0b27",
482
346
  "session_name": "openai-sys-a3f81c9d0b27",
483
- "model": "claude-opus-4-6",
347
+ "model": "claude-opus-5-5",
484
348
  "cwd": "/home/user/projects",
485
349
  "created": "2026-04-09T03:12:18.441Z",
486
350
  "turns": 68,
@@ -512,7 +376,7 @@ Run after standing up the server. Set `TOKEN=$(cat ~/.openclaw/server-token)` fi
512
376
  for SYS in 'You are Alice.' 'You are Bob.'; do
513
377
  curl -s http://127.0.0.1:18796/v1/chat/completions \
514
378
  -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
515
- -d "{\"model\":\"claude-opus-4-6\",\"messages\":[{\"role\":\"system\",\"content\":\"$SYS\"},{\"role\":\"user\",\"content\":\"hi\"}]}" \
379
+ -d "{\"model\":\"opus\",\"messages\":[{\"role\":\"system\",\"content\":\"$SYS\"},{\"role\":\"user\",\"content\":\"hi\"}]}" \
516
380
  | jq -r '.id'
517
381
  done
518
382
  curl -s http://127.0.0.1:18796/v1/sessions -H "Authorization: Bearer $TOKEN" \
@@ -523,7 +387,7 @@ curl -s http://127.0.0.1:18796/v1/sessions -H "Authorization: Bearer $TOKEN" \
523
387
  **2. Same system prompt + different model produces two sessions.**
524
388
 
525
389
  ```bash
526
- for M in claude-opus-4-6 claude-sonnet-4-6; do
390
+ for M in opus sonnet; do
527
391
  curl -s http://127.0.0.1:18796/v1/chat/completions \
528
392
  -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
529
393
  -d "{\"model\":\"$M\",\"messages\":[{\"role\":\"system\",\"content\":\"SAME\"},{\"role\":\"user\",\"content\":\"hi\"}]}" > /dev/null
@@ -538,10 +402,10 @@ curl -s http://127.0.0.1:18796/v1/sessions -H "Authorization: Bearer $TOKEN" | j
538
402
  SID=smoke-reset
539
403
  curl -s http://127.0.0.1:18796/v1/chat/completions \
540
404
  -H "Authorization: Bearer $TOKEN" -H "X-Session-Id: $SID" -H "Content-Type: application/json" \
541
- -d '{"model":"claude-opus-4-6","messages":[{"role":"user","content":"remember the word banana"}]}' > /dev/null
405
+ -d '{"model":"opus","messages":[{"role":"user","content":"remember the word banana"}]}' > /dev/null
542
406
  curl -s http://127.0.0.1:18796/v1/chat/completions \
543
407
  -H "Authorization: Bearer $TOKEN" -H "X-Session-Id: $SID" -H "X-Session-Reset: 1" -H "Content-Type: application/json" \
544
- -d '{"model":"claude-opus-4-6","messages":[{"role":"user","content":"what word did I just tell you"}]}' \
408
+ -d '{"model":"opus","messages":[{"role":"user","content":"what word did I just tell you"}]}' \
545
409
  | jq -r '.choices[0].message.content'
546
410
  # Expected: model says it has no prior context.
547
411
  ```
@@ -554,7 +418,7 @@ PREAMBLE=$(printf 'x%.0s' {1..3000})
554
418
  for i in 1 2 3 4; do
555
419
  curl -s http://127.0.0.1:18796/v1/chat/completions \
556
420
  -H "Authorization: Bearer $TOKEN" -H "X-Session-Id: $SID" -H "Content-Type: application/json" \
557
- -d "{\"model\":\"claude-opus-4-6\",\"messages\":[{\"role\":\"system\",\"content\":\"long preamble: $PREAMBLE\"},{\"role\":\"user\",\"content\":\"turn $i\"}]}" > /dev/null
421
+ -d "{\"model\":\"opus\",\"messages\":[{\"role\":\"system\",\"content\":\"long preamble: $PREAMBLE\"},{\"role\":\"user\",\"content\":\"turn $i\"}]}" > /dev/null
558
422
  curl -s http://127.0.0.1:18796/v1/sessions -H "Authorization: Bearer $TOKEN" \
559
423
  | jq ".data[] | select(.session_name == \"openai-$SID\") | {turn: $i, cached_tokens, tokens_in}"
560
424
  done
@@ -573,10 +437,12 @@ Errors use the OpenAI error envelope:
573
437
  | Status | When |
574
438
  | ------ | ---------------------------------------------------------------------- |
575
439
  | 400 | `messages` empty/missing, no user message, invalid `max_tokens` |
440
+ | 413 | Request body over 5 MiB |
576
441
  | 401 | Missing or wrong bearer token (when auth enabled) |
577
442
  | 415 | POST without `Content-Type: application/json` |
578
443
  | 429 | Rate limited (`OPENCLAW_RATE_LIMIT` exceeded) |
579
444
  | 503 | Failed to start a new session (model unavailable, CLI crashed at boot) |
445
+ | 502 | The CLI finished the turn with an error (`type: upstream_error`) |
580
446
  | 500 | Mid-turn failure |
581
447
 
582
448
  ## Related