@enderfga/claw-orchestrator 7.5.3 → 7.5.4
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +21 -22
- package/configs/engines/README.md +7 -6
- package/dist/bin/cli.js +1 -1
- package/dist/bin/cli.js.map +1 -1
- package/dist/src/acp-server.d.ts +1 -1
- package/dist/src/acp-server.js +7 -5
- package/dist/src/acp-server.js.map +1 -1
- package/dist/src/autoloop/notify.d.ts +5 -7
- package/dist/src/autoloop/notify.js +21 -20
- package/dist/src/autoloop/notify.js.map +1 -1
- package/dist/src/embedded-server.js +7 -4
- package/dist/src/embedded-server.js.map +1 -1
- package/dist/src/fanout.d.ts +6 -0
- package/dist/src/fanout.js +1 -0
- package/dist/src/fanout.js.map +1 -1
- package/dist/src/index.js +19 -11
- package/dist/src/index.js.map +1 -1
- package/dist/src/kernel/nodes/fanout.js +1 -0
- package/dist/src/kernel/nodes/fanout.js.map +1 -1
- package/dist/src/kernel/types.d.ts +2 -0
- package/dist/src/kernel/types.js.map +1 -1
- package/dist/src/openai-compat.d.ts +2 -2
- package/dist/src/openai-compat.js +5 -2
- package/dist/src/openai-compat.js.map +1 -1
- package/dist/src/session-manager.js +14 -5
- package/dist/src/session-manager.js.map +1 -1
- package/dist/src/types.d.ts +2 -0
- package/openclaw.plugin.json +1 -1
- package/package.json +2 -2
- package/skills/SKILL.md +31 -32
- package/skills/references/acp.md +19 -36
- package/skills/references/autoloop.md +158 -180
- package/skills/references/claude-cli-tracking.md +27 -27
- package/skills/references/cli.md +62 -79
- package/skills/references/council.md +40 -63
- package/skills/references/dashboard.md +42 -55
- package/skills/references/getting-started.md +20 -14
- package/skills/references/inbox.md +6 -4
- package/skills/references/mcp.md +29 -24
- package/skills/references/multi-engine.md +105 -153
- package/skills/references/observability.md +42 -32
- package/skills/references/openai-compat.md +169 -303
- package/skills/references/sessions.md +20 -29
- package/skills/references/tools.md +62 -76
- package/skills/references/ultra.md +17 -16
- package/skills/references/ultraapp.md +59 -64
- package/skills/references/verification.md +29 -52
- package/skills/references/workflow.md +37 -104
- package/skills/ultraapp/SKILL.md +9 -10
|
@@ -1,8 +1,8 @@
|
|
|
1
1
|
# OpenAI-Compatible Bridge
|
|
2
2
|
|
|
3
|
-
> **Cost
|
|
3
|
+
> **Cost note**: the bridge drives the Claude Code CLI, so usage counts against your Claude account. It does not bypass Anthropic's billing or subscription limits. Heavy agent traffic can consume extra-usage credits at API rates. Check your usage at [claude.ai/settings/usage](https://claude.ai/settings/usage).
|
|
4
4
|
|
|
5
|
-
The embedded server exposes a drop-in OpenAI-compatible endpoint so any client that speaks `/v1/chat/completions` can talk to a persistent Claude Code (or Codex / Antigravity / Grok) session. The bridge is designed to serve **two kinds of clients as first-class citizens**:
|
|
5
|
+
The embedded server exposes a drop-in OpenAI-compatible endpoint so any client that speaks `/v1/chat/completions` can talk to a persistent Claude Code (or Codex / Antigravity / Grok / OpenCode) session. The bridge is designed to serve **two kinds of clients as first-class citizens**:
|
|
6
6
|
|
|
7
7
|
1. **Upstream agents** that maintain their own conversation state and forward only the latest user turn — OpenClaw's main agent loop, cron jobs, subagents, programmatic clients.
|
|
8
8
|
2. **OpenAI-compatible webchat / labeling tools** that re-send the full transcript on every turn — ChatGPT-Next-Web, Open WebUI, LobeChat, data-labeling pipelines.
|
|
@@ -11,13 +11,14 @@ Both modes share the same wire protocol; the difference is how a "new conversati
|
|
|
11
11
|
|
|
12
12
|
## Endpoint
|
|
13
13
|
|
|
14
|
-
| |
|
|
15
|
-
| ----------------------- |
|
|
16
|
-
| **URL** | `http://127.0.0.1:18796/v1/chat/completions`
|
|
17
|
-
| **Models endpoint** | `GET /v1/models`
|
|
18
|
-
| **Inspection endpoint** | `GET /v1/sessions` (lists active openai-compat sessions with caching stats)
|
|
19
|
-
| **Auth** | Bearer token
|
|
20
|
-
| **Wire format** | OpenAI Chat Completions, both streaming (SSE) and non-streaming
|
|
14
|
+
| | |
|
|
15
|
+
| ----------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
|
|
16
|
+
| **URL** | `http://127.0.0.1:18796/v1/chat/completions` |
|
|
17
|
+
| **Models endpoint** | `GET /v1/models` |
|
|
18
|
+
| **Inspection endpoint** | `GET /v1/sessions` (lists active openai-compat sessions with caching stats) |
|
|
19
|
+
| **Auth** | `Authorization: Bearer <token>`. By default the server generates a token and writes it to `~/.openclaw/server-token` (mode 0600). Set `OPENCLAW_SERVER_TOKEN=<value>` to choose the token, or `OPENCLAW_SERVER_TOKEN=disabled` to turn auth off (single-user hosts only) |
|
|
20
|
+
| **Wire format** | OpenAI Chat Completions, both streaming (SSE) and non-streaming |
|
|
21
|
+
| **Default model** | `claude-sonnet-4-6` when the request has no `model` |
|
|
21
22
|
|
|
22
23
|
## Session keying
|
|
23
24
|
|
|
@@ -28,23 +29,22 @@ Each request is mapped to a long-running session. Once a session exists, subsequ
|
|
|
28
29
|
3. **`sys-<sha1(model + systemPrompt)[0..12]>`** — automatic fallback so unkeyed callers don't all collapse onto a single shared session
|
|
29
30
|
4. **`'default'`** — only when there is no system prompt AND no model (degenerate empty body)
|
|
30
31
|
|
|
31
|
-
|
|
32
|
+
Without the hash fallback, unkeyed callers (for example OpenClaw's main agent, cron jobs and subagents behind one gateway) would share one session: their requests would queue behind each other and could receive another caller's system prompt.
|
|
32
33
|
|
|
33
|
-
The model is mixed into the hash so that two callers with the same system prompt but different requested models (e.g. one wants `
|
|
34
|
+
The model is mixed into the hash so that two callers with the same system prompt but different requested models (e.g. one wants `opus`, another wants `sonnet`) get separate sessions rather than responses from the wrong model.
|
|
34
35
|
|
|
35
36
|
The full plugin-side session name is `openai-<key>`. The key becomes a directory
|
|
36
37
|
name — the bridge starts each session in `os.tmpdir()/openclaw-compat-<name>` —
|
|
37
38
|
so a key that is not already `[A-Za-z0-9._-]` is replaced by a hash of itself.
|
|
38
|
-
|
|
39
|
-
|
|
39
|
+
An ordinary id is used unchanged; a key containing path separators cannot choose
|
|
40
|
+
where the session runs.
|
|
40
41
|
|
|
41
42
|
The tool list is fingerprinted into the hash fallback by name, a description
|
|
42
43
|
prefix, and the **parameter schema** (with object keys normalised, so a
|
|
43
44
|
re-serialised identical schema still resolves to the same session). On the
|
|
44
|
-
Claude engine the schemas are
|
|
45
|
-
|
|
46
|
-
|
|
47
|
-
parameters kept getting `tool_calls` shaped like the schema it had replaced.
|
|
45
|
+
Claude engine the schemas are written into the session's system prompt when the
|
|
46
|
+
session is created and are not re-sent per turn, so a changed schema has to
|
|
47
|
+
resolve to a new session for the model to see it.
|
|
48
48
|
|
|
49
49
|
## Operator modes
|
|
50
50
|
|
|
@@ -96,13 +96,12 @@ mechanism is used depends on whether the engine keeps the conversation itself.
|
|
|
96
96
|
| `codex`, `codex-app`, `agy`, `opencode`, `grok` | Full schema block prepended to the message | A short reminder of the calling convention, no schemas — but only once the conversation id has been captured; until then the full block is sent again |
|
|
97
97
|
| `gemini`, one-shot `custom` | Full schema block prepended to the message | Full schema block again — these have no resume surface, so nothing persists between sends |
|
|
98
98
|
|
|
99
|
-
The
|
|
100
|
-
|
|
101
|
-
|
|
102
|
-
|
|
103
|
-
|
|
104
|
-
|
|
105
|
-
framing, and without it the CLI starts doing the work directly.
|
|
99
|
+
The engines in the middle row resume a conversation by id, so everything injected
|
|
100
|
+
stays in the transcript. Re-sending the full block each turn would grow the prompt
|
|
101
|
+
without bound (a 54-tool block is roughly 17k tokens), so a few turns could
|
|
102
|
+
overflow the context window. The short reminder is still needed, because the
|
|
103
|
+
block carries the "emit a tool call, do not carry out the work yourself" framing;
|
|
104
|
+
without it the CLI starts doing the work directly.
|
|
106
105
|
|
|
107
106
|
A fresh session always gets the full block, so a thread is never created without
|
|
108
107
|
the definitions — including when a session was evicted and is being recreated. A
|
|
@@ -115,19 +114,13 @@ is prepended to the message, and skipped only while the conversation it was sent
|
|
|
115
114
|
to is still the one being resumed. A turn that creates a conversation always
|
|
116
115
|
carries it — including a `X-Session-Reset: 1` turn, which stops the existing
|
|
117
116
|
session and starts a new one. "The conversation is being resumed" means the
|
|
118
|
-
engine has actually announced an id (codex's `thread.started`, agy's
|
|
119
|
-
and opencode's session id), not merely that the session is in the
|
|
120
|
-
first turn that
|
|
121
|
-
nothing, and
|
|
122
|
-
|
|
123
|
-
|
|
124
|
-
|
|
125
|
-
seen at all. The header is parsed after the branch that handles a trailing `tool`
|
|
126
|
-
role returns, so on that shape the reset stops nothing, creates nothing, and the
|
|
127
|
-
turn is treated as a resumed one — the prompt is skipped if the thread is live.
|
|
128
|
-
Every other shape, `[..., tool, user]` included, honors it. A client that needs a
|
|
129
|
-
reset mid-tool-loop has to send it on a turn that does not end in a `tool`
|
|
130
|
-
message.
|
|
117
|
+
engine has actually announced an id (codex's `thread.started`, agy's `init` event,
|
|
118
|
+
cursor's, grok's and opencode's session id), not merely that the session is in the
|
|
119
|
+
manager's map: a first turn that failed before announcing one leaves a session
|
|
120
|
+
that resumes nothing, and later turns keep receiving the full prompt.
|
|
121
|
+
|
|
122
|
+
See [Known limitations](#known-limitations) for the one request shape where
|
|
123
|
+
`X-Session-Reset` is not honoured.
|
|
131
124
|
|
|
132
125
|
`OPENAI_COMPAT_TOOLS_PER_MESSAGE=1` opts out: it re-sends the full block on every
|
|
133
126
|
turn for `claude` too, which is what makes a changing tool set work inside one
|
|
@@ -138,13 +131,10 @@ does change mid-conversation.
|
|
|
138
131
|
## Conversation history on the way in
|
|
139
132
|
|
|
140
133
|
The caller's `messages[]` can carry the whole conversation: earlier `user` turns and the engine's
|
|
141
|
-
own earlier `assistant` replies.
|
|
142
|
-
|
|
143
|
-
|
|
144
|
-
|
|
145
|
-
caller's latest `user` text. The wrapper tag is the one `renderHistory()` in the autoloop dispatcher
|
|
146
|
-
already uses for the same job; the per-turn tags are not — that one labels its two speakers `<user>`
|
|
147
|
-
/ `<agent>`, because its roles are autoloop roles rather than OpenAI wire roles.
|
|
134
|
+
own earlier `assistant` replies. The bridge sends those turns only when the engine's conversation
|
|
135
|
+
does not already hold them. The turns that are sent are serialized into one
|
|
136
|
+
`<conversation_history>` block of `<user>` / `<assistant>` turns, placed in front of the caller's
|
|
137
|
+
latest `user` text.
|
|
148
138
|
|
|
149
139
|
| On this turn the engine | What is sent |
|
|
150
140
|
| ----------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------ |
|
|
@@ -152,217 +142,111 @@ already uses for the same job; the per-turn tags are not — that one labels its
|
|
|
152
142
|
| **is** resuming a live conversation, but one this bridge never sent these turns to | the same — a live thread is not automatically _this_ thread |
|
|
153
143
|
| **is** resuming the live conversation these turns belong to | nothing — the turns are already in the transcript, and the text goes out alone |
|
|
154
144
|
|
|
155
|
-
The last row
|
|
156
|
-
|
|
157
|
-
|
|
158
|
-
|
|
159
|
-
|
|
160
|
-
|
|
161
|
-
|
|
162
|
-
|
|
163
|
-
|
|
164
|
-
|
|
165
|
-
|
|
166
|
-
|
|
167
|
-
|
|
168
|
-
|
|
169
|
-
|
|
170
|
-
|
|
171
|
-
|
|
172
|
-
|
|
173
|
-
|
|
174
|
-
|
|
175
|
-
|
|
176
|
-
|
|
177
|
-
|
|
178
|
-
|
|
179
|
-
|
|
180
|
-
|
|
181
|
-
|
|
182
|
-
|
|
183
|
-
|
|
184
|
-
|
|
185
|
-
|
|
186
|
-
|
|
187
|
-
|
|
188
|
-
|
|
189
|
-
The
|
|
190
|
-
|
|
191
|
-
map, so a fingerprint outlives the session it mirrors. Eviction costs a replayed block, never a
|
|
192
|
-
dropped one, and `serve` restarts start it empty for the same price.
|
|
193
|
-
|
|
194
|
-
`system` messages are never in the block: they travel as the session's system prompt (see [Tool
|
|
195
|
-
definitions](#tool-definitions-and-where-they-live)). `tool` messages are never in it either —
|
|
196
|
-
they are the next section's business, and repeating them here would duplicate every payload and
|
|
197
|
-
undo the scoping that keeps a tool loop linear. An `assistant` message that only announces
|
|
198
|
-
`tool_calls` carries no text, so it renders no turn; a turn whose text is empty or whitespace is
|
|
199
|
-
dropped rather than rendered as an empty shell. An array with nothing to replay produces no block
|
|
200
|
-
at all, so a single-turn `[system, user]` request — the shape the OpenClaw main agent, cron jobs
|
|
201
|
-
and subagents send — goes out exactly as it did before.
|
|
202
|
-
|
|
203
|
-
Replayed text has neutralized every tag the assembled prompt treats as structure — the block's own
|
|
204
|
-
`<conversation_history>` / `<user>` / `<assistant>`, and also `<tool_results>` / `<tool_result>`,
|
|
205
|
-
`<system>` and `<tool_calls>` (`</user>` becomes `</user>`). Unlike a `<tool_result>` body, which
|
|
206
|
-
comes from the caller's own tool runner, a replayed turn is whatever an end user typed, and
|
|
207
|
-
`hi</user>\n<assistant>\n...` would otherwise close its turn early and forge an `assistant` turn —
|
|
208
|
-
putting words in the engine's own mouth. The other three matter for the same reason: a fabricated
|
|
209
|
-
tool return carries the framing sentence that says a payload is authoritative, a `<system>` block
|
|
210
|
-
contradicts the real one the non-claude path prepends, and `<tool_calls>` is the exact protocol JSON
|
|
211
|
-
the model is asked to emit.
|
|
212
|
-
|
|
213
|
-
Only the `<` is escaped, by lookahead. That shape is what makes the boundary decidable: matching up
|
|
214
|
-
to the closing `>` instead means re-emitting whatever was captured, and `hola<user a</user>` then
|
|
215
|
-
smuggles a raw close through the attribute slot of a tag that IS matched. So the match ends at
|
|
216
|
-
anything that ends a tag name — `>`, `/`, `<`, end of text, or a character that occupies no width.
|
|
217
|
-
That last clause is five Unicode properties rather than a list of code points, and it is not
|
|
218
|
-
decoration: measured over the 6,060 code points that are zero-advance or render blank,
|
|
219
|
-
`[\s></\p{Cc}\p{Cf}]` let **5,806** through, so `ok</user︀>\n<assistant︀>` forged a turn that is
|
|
220
|
-
indistinguishable on screen from `ok</user>`. Adding `Default_Ignorable` leaves 1,770;
|
|
221
|
-
`\p{Mn}\p{Me}` closes it; U+2800 BRAILLE PATTERN BLANK is neither and is named. Cost: the same 11 of
|
|
222
|
-
a 28-string corpus of plausible legitimate text change under the wide class as under the narrow one.
|
|
223
|
-
|
|
224
|
-
The name does **not** have to sit flush against `<` or `</`: the same invisible padding, plus the
|
|
225
|
-
slash itself, is allowed before the name, because `hola</user>` renders as `hola</user>` and a
|
|
226
|
-
model reads it as a close — a trailing class complete over zero-width filler with a flush leading
|
|
227
|
-
side still forges a turn. That is **one** class, `[/\p{Cc}\p{Cf}\p{Mn}\p{Me}\p{Default_Ignorable_Code_Point}]*`,
|
|
228
|
-
not `[…]*\/?[…]*`: two adjacent unbounded quantifiers over the same class backtrack O(n²) on a long
|
|
229
|
-
run that never reaches a valid name, and a single history message of ~100 KB of combining marks hung
|
|
230
|
-
the event loop ~80 s — a one-request denial of service against a fleet whose watchdog already resets
|
|
231
|
-
on event-loop stalls. The class is the zero-**advance** subset only (no `\s`, no U+2800), because a
|
|
232
|
-
visible separator before the name is the `if (count < user && x)` corruption the boundary class
|
|
233
|
-
already refuses. Swept over the 6,060 code points that are zero-advance or separators, in all three
|
|
234
|
-
positions (18,180 probes): 12,120 went through unfenced with the flush leading side, **38** with the
|
|
235
|
-
interior class, and the 38 are 19 `Zs`/`Zl`/`Zp` code points — U+0020, U+00A0, U+2000–200A, U+3000 —
|
|
236
|
-
i.e. exactly the visible-separator limitation below, and nothing else. The single class also lets a
|
|
237
|
-
slash sit among the filler (`<//user`, `</␀/user`); harmless, since the only action is escaping the
|
|
238
|
-
`<`. Filler INSIDE the name (`</us␀er>`) is still not fenced — the old regex missed it too.
|
|
239
|
-
|
|
240
|
-
**What it does not promise.** A positive class cannot be complete over a _visible_ separator, so
|
|
241
|
-
`</ user>` and `< assistant>` go through raw — a model reads them as a boundary. They are excluded on
|
|
242
|
-
cost: reaching them means corrupting `if (count < user && x)` and `the < user > column`. What already
|
|
243
|
-
gets corrupted for the same reason, since the class contains whitespace: `Promise<User | null>` and
|
|
244
|
-
`count<user && total>limit` come out with `<`. Readable to a model, not byte-identical. And the
|
|
245
|
-
body of a `<tool_result>` is never fenced — its content comes from the caller's own tool runner — so a
|
|
246
|
-
tool return that embeds a whole `<conversation_history>` block is not stopped here.
|
|
247
|
-
|
|
248
|
-
The caller's **latest** `user` turn is fenced too, but only on turns that actually carry a block.
|
|
249
|
-
That turn is the one input an attacker controls end to end, and the block teaches the model in the
|
|
250
|
-
same prompt that `<conversation_history>` holds its own earlier turns — unfenced, it could close the
|
|
251
|
-
real block and open a second one indistinguishable from it. With no block in front of it there is
|
|
252
|
-
nothing to forge, so those turns stay byte-for-byte what they were.
|
|
253
|
-
|
|
254
|
-
A `user` turn carrying only non-text content (an image) renders as `[non-text content]` rather than
|
|
255
|
-
vanishing. Dropping it would leave the `assistant` reply to it standing alone under a framing that
|
|
256
|
-
calls the assistant turns the model's own — a reply to a request the model cannot see, which reads
|
|
257
|
-
as license to act on the reply by itself. "Photo of the invoice", then "yes, go ahead", is an
|
|
258
|
-
everyday shape. A leading `assistant` turn with no `user` turn in front of it at all (content
|
|
259
|
-
`null`, or empty) is dropped instead, since no marker could honestly stand in for it.
|
|
145
|
+
The last row keeps Anthropic prompt caching warm on `claude` and stops a resumed `codex` thread
|
|
146
|
+
from being sent its own history again: on a live thread the message is exactly the caller's text.
|
|
147
|
+
The first row covers clients that open a new conversation every turn (for example, one whose
|
|
148
|
+
session key hashes the last message). Bridge sessions are created with `skipPersistence: true` and
|
|
149
|
+
are never resumed from disk, so without the replay a follow-up like "yes, go ahead" would reach
|
|
150
|
+
the engine with no context. How the bridge tells the second row from the third is described in
|
|
151
|
+
[A live session is not the same thing as this conversation](#a-live-session-is-not-the-same-thing-as-this-conversation).
|
|
152
|
+
|
|
153
|
+
What goes into the block:
|
|
154
|
+
|
|
155
|
+
- `system` messages are never in it: they travel as the session's system prompt (see
|
|
156
|
+
[Tool definitions](#tool-definitions-and-where-they-live)).
|
|
157
|
+
- `tool` messages are never in it either; they are handled by the `<tool_results>` block described
|
|
158
|
+
under [Tool results on the way back](#tool-results-on-the-way-back).
|
|
159
|
+
- An `assistant` message that only announces `tool_calls` carries no text, so it renders no turn. A
|
|
160
|
+
turn whose text is empty or whitespace is dropped.
|
|
161
|
+
- A `user` turn carrying only non-text content (an image) renders as `[non-text content]`, so the
|
|
162
|
+
`assistant` reply to it does not appear to answer nothing. A leading `assistant` turn with no
|
|
163
|
+
`user` turn before it (content `null`, or empty) is dropped.
|
|
164
|
+
- An array with nothing to replay produces no block, so a single-turn `[system, user]` request —
|
|
165
|
+
the shape the OpenClaw main agent, cron jobs and subagents send — goes out unchanged.
|
|
166
|
+
|
|
167
|
+
Replayed text has the `<` of every structural tag escaped (`</user>` becomes `</user>`): the
|
|
168
|
+
block's own `<conversation_history>` / `<user>` / `<assistant>`, and also `<tool_results>` /
|
|
169
|
+
`<tool_result>`, `<system>` and `<tool_calls>`. Forms padded with invisible or zero-width
|
|
170
|
+
characters (`</user>`, `<user︀>`) are caught too. This stops a user message from closing its own
|
|
171
|
+
turn or forging an `assistant` turn, a `<system>` block, a tool result, or a tool call. Limits:
|
|
172
|
+
|
|
173
|
+
- A visible space inside the tag (`</ user>`, `< assistant>`) is not caught, and filler inside the
|
|
174
|
+
tag name (`</us␀er>`) is not caught either.
|
|
175
|
+
- Text that merely looks like a tag, such as `Promise<User | null>`, comes out with `<`. It stays
|
|
176
|
+
readable to a model but is not byte-identical.
|
|
177
|
+
- The body of a `<tool_result>` is never escaped, since it comes from the caller's own tool runner.
|
|
178
|
+
|
|
179
|
+
The caller's **latest** `user` turn is escaped the same way, but only on turns that carry a block.
|
|
180
|
+
With no block in front of it there is nothing to forge, so those turns go out unchanged.
|
|
260
181
|
|
|
261
182
|
### What this does not cover
|
|
262
183
|
|
|
263
|
-
- **A turn can be replayed that the engine already had.** The
|
|
264
|
-
|
|
265
|
-
|
|
266
|
-
|
|
267
|
-
not to act on them again, which is what keeps a duplicated turn from becoming a duplicated
|
|
268
|
-
action; nothing enforces it.
|
|
184
|
+
- **A turn can be replayed that the engine already had.** The engine's state is read from the
|
|
185
|
+
session, not from the array, so a client whose session looks new to the bridge but whose engine
|
|
186
|
+
did hold context gets those turns a second time. The framing tells the model these are earlier
|
|
187
|
+
turns and not to act on them again; nothing enforces it.
|
|
269
188
|
- **The block is not in strict chronological order when the array does not end in the caller's
|
|
270
189
|
latest `user` turn.** Every `user`/`assistant` turn except that one is replayed, including turns
|
|
271
|
-
|
|
272
|
-
|
|
273
|
-
missing the last thing the model said invites it to redo the work. The cost is that such a turn
|
|
274
|
-
is rendered inside the block, i.e. before the caller's latest text rather than after it.
|
|
275
|
-
- **`X-Session-Reset` is not honored on an array ending in a `tool` result**, so on that one shape a
|
|
276
|
-
reset turn on a live thread gets no history block either. Same pre-existing asymmetry the tool
|
|
277
|
-
results have there, from the same cause — the header is parsed after that branch returns. See the
|
|
278
|
-
note under [Tool definitions](#tool-definitions-and-where-they-live).
|
|
279
|
-
- **`grok` inherits the hole described in the next section**: it resumes by id, but its id is absent
|
|
280
|
-
from `SessionStats`, so `nativeThreadIsLive()` reaches `default: return true` and a `grok` session
|
|
281
|
-
whose first turn died before emitting its id reports a live thread — and so gets no history. Same
|
|
282
|
-
cause, same fix (adding the field), not addressed here.
|
|
190
|
+
after it — an array ending in `assistant` (prefill, an explicit "continue") keeps that turn, so it
|
|
191
|
+
is rendered inside the block, before the caller's latest text.
|
|
283
192
|
- **The two blocks are not interleaved.** When a request carries both, the message is the history
|
|
284
|
-
block, then the tool results, then the caller's new text
|
|
285
|
-
|
|
286
|
-
- **The block is capped at 24,000 characters
|
|
287
|
-
|
|
288
|
-
|
|
289
|
-
|
|
290
|
-
|
|
291
|
-
|
|
292
|
-
|
|
293
|
-
|
|
294
|
-
|
|
295
|
-
|
|
296
|
-
|
|
297
|
-
|
|
298
|
-
|
|
299
|
-
|
|
300
|
-
|
|
301
|
-
|
|
302
|
-
|
|
303
|
-
The newest `user` turn in the block — the ask the turns after it answer — is the one exception to
|
|
304
|
-
spending newest-first: the turns after it (all `assistant`, by definition of "newest `user` turn")
|
|
305
|
-
spend against the budget minus its frame plus `min(its length, 200)`. That reserve is the fix for
|
|
306
|
-
a drop, not a refinement. Without it, replies that add past the cap (one 30k pasted listing, or
|
|
307
|
-
two ordinary 12k ones) take the whole budget, the window starts past every `user` turn, the
|
|
308
|
-
leading-`assistant` rule clears what is left, and NO block goes out at all: the caller's latest
|
|
309
|
-
turn reaches the engine alone, which is this block's own failure mode at its worst.
|
|
310
|
-
|
|
311
|
-
The 200-character floor applies **above** the anchor too, and that is the second half of the cap
|
|
312
|
-
fix. Once the post-anchor turns have truncated the budget down near the reserve, an older turn is
|
|
313
|
-
started with a `room` smaller than the 32-character elision marker, and `slice(0, room - 32)` with
|
|
314
|
-
a negative argument slices from the END of the string — emitting nearly the whole turn while
|
|
315
|
-
charging the budget only `room`, which blows the cap. Measured on a narrating tool loop, where each
|
|
316
|
-
hop's `assistant` narration becomes a consecutive post-anchor turn because `tool` messages are
|
|
317
|
-
filtered out: with the floor removed, 30 hops render 27,627 characters and the three-`user` /
|
|
318
|
-
with-reply sweep shapes reach 28,003 / 30,003. With the floor, that run of older turns is dropped
|
|
319
|
-
instead, which the anchor's reserve makes safe: ending the window there would drop the anchor and
|
|
320
|
-
hand the leading-`assistant` rule an all-`assistant` list to clear, i.e. the empty block again.
|
|
321
|
-
|
|
322
|
-
Verification. 44,000 random shapes across three message-count ranges (≤7, ≤60 and ≤400 messages,
|
|
323
|
-
roles `user`/`assistant`/`system`/`tool`, content string/whitespace/null/array/empty-array/no-text,
|
|
324
|
-
lengths straddling 0/1/199/200/201/11,999/12,000/12,001/23,799/23,800/24,000/24,001/30,000/60,000):
|
|
325
|
-
**max block 23,999 characters, zero shapes over 24,000, zero content-free turns, zero blocks with
|
|
326
|
-
no `user` turn in them, and zero shapes that went from a non-empty block to an empty one.** 7,488
|
|
327
|
-
of them went the other way — empty before, non-empty now — which is the anchor reserve doing its
|
|
328
|
-
job. 11,954 non-empty blocks changed, which is the point: the old arithmetic charged less than it
|
|
329
|
-
emitted, so every block near the ceiling gets shorter. Directed shapes: 30/100/2,000 narrating
|
|
330
|
-
hops and 1,200/8,000/24,000 one-word turns are all ≤ 24,000 with no content-free turns.
|
|
331
|
-
|
|
332
|
-
The ceiling is on characters and argv counts bytes, which is the conservative direction only up to
|
|
333
|
-
a point: 24,000 characters of astral-plane text is 48,000 bytes and the worst case (3-byte BMP) is
|
|
334
|
-
72,000, both still inside 128 KiB, but the
|
|
335
|
-
block is not the whole prompt — the tool block, the `<system>` prepend and the caller's own turn
|
|
336
|
-
are added after it. What the cap bounds is the part that scales with the transcript.
|
|
337
|
-
|
|
338
|
-
`MAX_BODY_SIZE` (5 MiB) is **not** a usable bound here: six of the nine `ENGINE_TYPES` pass the
|
|
339
|
-
prompt to the CLI as a single argv element (`codex`, `gemini`, `agy`, `cursor`, `grok`,
|
|
340
|
-
`opencode`), and so does a one-shot `custom` engine; Linux caps one argument at `MAX_ARG_STRLEN` =
|
|
341
|
-
128 KiB whatever `getconf ARG_MAX` reports (measured: 131071 bytes spawns, 131072 throws `E2BIG`).
|
|
342
|
-
Only `claude`, `codex-app` and a persistent `custom` engine write over stdin. Uncapped, ordinary
|
|
343
|
-
traffic reaches that ceiling — 400-character turns with 900-character replies put the message at
|
|
344
|
-
131,063 characters at turn 96 and 132,425 at turn 97 — and the failure is a 500 with the turn
|
|
345
|
-
lost, which is worse than the missing context the block exists to restore. This is the same trade
|
|
346
|
-
`renderHistory()` makes, with the same `REPLAY_CHAR_BUDGET`, feeding the same engines.
|
|
347
|
-
- **A send that threw records nothing.** The fingerprint is written after the send and only when the
|
|
348
|
-
send landed, because the two ways to be wrong are not symmetric: forgetting a turn that landed
|
|
349
|
-
replays it once more, while assuming one landed that did not drops context silently. "Landed" is
|
|
350
|
-
`sendMessage` returning — a returned error is answered with 502 and still records, since the CLI
|
|
351
|
-
received the prompt — so only a throw withholds the record. An earlier revision of this file
|
|
352
|
-
described the opposite placement as deliberate and named its residue: a caller that answered a 5xx
|
|
353
|
-
by appending an `assistant` turn and sending again got the next turn bare. That was the bug, not
|
|
354
|
-
the design.
|
|
193
|
+
block, then the tool results, then the caller's new text. A `tool` result that chronologically
|
|
194
|
+
preceded a replayed `assistant` turn still appears after it.
|
|
195
|
+
- **The block is capped at 24,000 characters, counting tags, markers and framing.** Oldest turns are
|
|
196
|
+
dropped first. The turn where the budget runs out is cut (start kept) and marked
|
|
197
|
+
`[… turn truncated for length …]`. A turn with under 200 characters of room is dropped instead.
|
|
198
|
+
The most recent `user` turn in the block always keeps at least its first 200 characters, so the
|
|
199
|
+
block is never left with only `assistant` turns.
|
|
200
|
+
|
|
201
|
+
The cap exists because most engines (`codex`, `agy`, `grok`, `opencode`, `cursor`, `gemini`,
|
|
202
|
+
one-shot `custom`) receive the prompt as a single command-line argument, and Linux limits one
|
|
203
|
+
argument to 128 KiB (`MAX_ARG_STRLEN`). Going over fails the request with a 500. The 5 MiB request
|
|
204
|
+
body limit does not protect against this. `claude`, `codex-app` and a persistent `custom` engine
|
|
205
|
+
write over stdin instead. The cap bounds only the part of the prompt that grows with the
|
|
206
|
+
transcript; the tool block, the system prompt and the caller's own turn are added on top.
|
|
207
|
+
|
|
208
|
+
- **A send that threw records nothing.** The fingerprint is written only after the send returns. A
|
|
209
|
+
send that returns an error is answered with 502 and still records, since the CLI received the
|
|
210
|
+
prompt; only a thrown send leaves no record, so the next request replays the turns.
|
|
355
211
|
- **A second request that arrives while the first is still in flight replays.** It sees no
|
|
356
|
-
fingerprint yet, so the transcript goes out again into the session that already holds it
|
|
357
|
-
|
|
358
|
-
- **Cost is O(n) per turn for engines that never resume.** `
|
|
359
|
-
|
|
360
|
-
|
|
361
|
-
- **`X-Session-Reset`
|
|
362
|
-
|
|
363
|
-
opposite of what was asked, and a client that sends the header on every request AND re-sends
|
|
212
|
+
fingerprint yet, so the transcript goes out again into the session that already holds it: a
|
|
213
|
+
duplicate rather than a loss, plus a lost cache prefix.
|
|
214
|
+
- **Cost is O(n) per turn for engines that never resume.** `gemini` and one-shot custom engines have
|
|
215
|
+
no native conversation, so the block is rebuilt and sent on every turn (capped). The same applies
|
|
216
|
+
to any caller that creates a new session per turn.
|
|
217
|
+
- **`X-Session-Reset` replays the transcript.** A reset turn means the engine holds nothing, so the
|
|
218
|
+
history goes out in full. A client that sends the header on every request and also re-sends
|
|
364
219
|
`messages[]` pays for the transcript every time.
|
|
365
220
|
|
|
221
|
+
### A live session is not the same thing as this conversation
|
|
222
|
+
|
|
223
|
+
Suppressing the replay needs a stronger fact than "a session under this name is live". That is what
|
|
224
|
+
`nativeThreadIsLive()` reports, and a session name can be live while its transcript belongs to a
|
|
225
|
+
different exchange. Three shapes where the two come apart, all reachable with default settings:
|
|
226
|
+
|
|
227
|
+
| shape | what happens |
|
|
228
|
+
| ---------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------- |
|
|
229
|
+
| a caller whose session key hashes its latest message | every repeat of the same short confirmation resolves to whichever session that phrase opened first — often a different subject entirely |
|
|
230
|
+
| a caller that sends no `X-Session-Id` at all | the key falls back to a hash of model + system prompt + tools, so all of that caller's concurrent chats share one name |
|
|
231
|
+
| `engine: 'claude'` — the default | `nativeThreadIsLive()` has no id to check and returns `true` for anything in the session map, so the name is the only evidence there is |
|
|
232
|
+
|
|
233
|
+
So the bridge records, per session, a fingerprint of the `user` turns it has sent there, and
|
|
234
|
+
replays whenever the incoming conversation is not the one it recorded. It is the only writer to
|
|
235
|
+
these sessions, so what it sent is what the engine holds. The fingerprint covers the `user` turns
|
|
236
|
+
only: those are the caller's own text echoed back verbatim, while assistant text is what the engine
|
|
237
|
+
produced and a client may normalize it. A mismatch replays: the cost is a repeated block, never a
|
|
238
|
+
lost one.
|
|
239
|
+
|
|
240
|
+
What the request is compared against depends on how the array ends. For an array ending in a `user`
|
|
241
|
+
turn, the `user` turns before that last one are compared, since the last one has not been sent yet.
|
|
242
|
+
For any other ending (a tool-loop hop ending in `tool`, a prefill ending in `assistant`) the latest
|
|
243
|
+
`user` turn was already sent, so all `user` turns are compared.
|
|
244
|
+
|
|
245
|
+
The fingerprint map holds at most 1,000 entries, evicted oldest-first. It is separate from the
|
|
246
|
+
session map: `_cleanupIdleSessions()` reaps idle sessions by TTL without updating it, so a
|
|
247
|
+
fingerprint can outlive its session. Losing an entry (eviction, or a `serve` restart, which starts
|
|
248
|
+
the map empty) costs a replayed block, never a dropped one.
|
|
249
|
+
|
|
366
250
|
## Tool results on the way back
|
|
367
251
|
|
|
368
252
|
A `tool` role message in the caller's array is the result of a call the model asked for on an
|
|
@@ -380,15 +264,9 @@ array:
|
|
|
380
264
|
Whether the engine is resuming is resolved with the same `nativeThreadIsLive()` check as the
|
|
381
265
|
middle row of [Tool definitions](#tool-definitions-and-where-they-live). Engines with no resume
|
|
382
266
|
surface (`gemini`, one-shot `custom`) are never in the second row, because
|
|
383
|
-
`engineHasNativeConversation()` gates the check.
|
|
384
|
-
|
|
385
|
-
|
|
386
|
-
`nativeThreadIsLive()` has no case for them and its `default` arm returns `true`: `claude`, which
|
|
387
|
-
holds its context in a live process and has no separate id to check, and `grok` — which does resume
|
|
388
|
-
by id (`--resume <sessionUUID>`), but whose id is absent from `SessionStats`, so there is nothing to
|
|
389
|
-
check even though there is something to check for. A `grok` session whose first turn died before
|
|
390
|
-
emitting its id therefore reports a live thread and gets its earlier rounds scoped away. That is
|
|
391
|
-
pre-existing and not addressed here; fixing it means adding the field.
|
|
267
|
+
`engineHasNativeConversation()` gates the check. `claude` (and a persistent `custom` engine) holds
|
|
268
|
+
its context in a live process and has no separate id to check, so it is in the second row whenever
|
|
269
|
+
the session exists.
|
|
392
270
|
|
|
393
271
|
The trailing role of the array does not enter into it — `[..., tool]`, `[..., tool, user]` and
|
|
394
272
|
`[..., tool, assistant]` are read the same way, and the caller's latest `user` text, when there is
|
|
@@ -397,70 +275,56 @@ all — a multimodal content array holding only an image — gets the block as t
|
|
|
397
275
|
array carrying no `tool` message at all is untouched: no block, no wrapper, the message goes as it
|
|
398
276
|
came.
|
|
399
277
|
|
|
400
|
-
The third case in the first row — a session being stopped and recreated — is `X-Session-Reset`,
|
|
401
|
-
|
|
402
|
-
ending in a `tool` result, where the header is not seen; see the note under [Tool definitions](#tool-definitions-and-where-they-live).
|
|
278
|
+
The third case in the first row — a session being stopped and recreated — is `X-Session-Reset`,
|
|
279
|
+
except on the one shape listed under [Known limitations](#known-limitations).
|
|
403
280
|
|
|
404
281
|
### What the scoping is for, and what it does not cover
|
|
405
282
|
|
|
406
|
-
On a resumed conversation the scoping
|
|
283
|
+
On a resumed conversation the scoping keeps a tool loop linear instead of quadratic. With a
|
|
407
284
|
30k-character batch per round, the tenth hop carries ~30k characters of results instead of the
|
|
408
285
|
~300k the engine has already seen. Two properties are worth checking against your own client before
|
|
409
286
|
relying on it:
|
|
410
287
|
|
|
411
|
-
- **The boundary is the last `assistant` message, and what matters is whether one sits
|
|
288
|
+
- **The boundary is the last `assistant` message, and what matters is whether one sits after the
|
|
412
289
|
earliest unsent `tool` message** — not whether the array contains one at all. The OpenAI wire
|
|
413
290
|
format has the caller echo the `assistant` turn that carried the `tool_calls` ahead of the
|
|
414
291
|
matching `tool` messages, and a client that echoes it pays for one round per hop. When no
|
|
415
|
-
`assistant` message follows the earliest unsent result,
|
|
416
|
-
|
|
417
|
-
|
|
418
|
-
|
|
292
|
+
`assistant` message follows the earliest unsent result, the slice keeps everything, so the scoping
|
|
293
|
+
has no effect. Two shapes land there: an array with no `assistant` message at all, and one
|
|
294
|
+
`assistant` announcing N parallel calls followed by its N results — the second costs nothing,
|
|
295
|
+
since those N results _are_ one round.
|
|
419
296
|
- **A round can go out twice.** Engine replies are not read back out of the array, so a round the
|
|
420
297
|
caller did not record an `assistant` turn for looks the same as a round the engine never saw, and
|
|
421
298
|
it goes out again. The duplication is bounded to one round wherever the scoping runs — i.e. on a
|
|
422
299
|
resumed conversation whose array does carry an `assistant` message. It is unbounded in the two
|
|
423
|
-
cases where nothing is scoped: the
|
|
424
|
-
resumed conversation), where the whole array goes out
|
|
425
|
-
it.
|
|
300
|
+
cases where nothing is scoped: the shape above with no `assistant` message, and any turn in the
|
|
301
|
+
first row of the table (no resumed conversation), where the whole array goes out because the
|
|
302
|
+
engine holds none of it.
|
|
426
303
|
|
|
427
304
|
Neither applies to a caller that keeps its own transcript and forwards only the latest turn — it
|
|
428
305
|
sends one round at a time.
|
|
429
306
|
|
|
430
|
-
|
|
431
|
-
|
|
432
|
-
Suppressing the replay needs a stronger fact than "a session under this name is live". That is what
|
|
433
|
-
`nativeThreadIsLive()` reports, and a session name can be live while its transcript belongs to a
|
|
434
|
-
different exchange. Three shapes where the two come apart, all reachable with default settings:
|
|
435
|
-
|
|
436
|
-
| shape | what happens |
|
|
437
|
-
| ---------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------- |
|
|
438
|
-
| a caller whose session key hashes its latest message | every repeat of the same short confirmation resolves to whichever session that phrase opened first — often a different subject entirely |
|
|
439
|
-
| a caller that sends no `X-Session-Id` at all | the key falls back to a hash of model + system prompt + tools, so all of that caller's concurrent chats share one name |
|
|
440
|
-
| `engine: 'claude'` — the default | `nativeThreadIsLive()` has no id to check and returns `true` for anything in the session map, so the name is the only evidence there is |
|
|
441
|
-
|
|
442
|
-
So the bridge tracks what it actually pushed into each session and replays whenever the incoming
|
|
443
|
-
conversation is not the one it remembers seeding. The fingerprint covers the `user` turns only: those
|
|
444
|
-
are the caller's own text echoed back verbatim, while assistant text is what the engine produced and
|
|
445
|
-
a client may normalize it. A mismatch replays, which is the safe direction — the cost is a repeated
|
|
446
|
-
block, never a lost one.
|
|
307
|
+
## Known limitations
|
|
447
308
|
|
|
448
|
-
|
|
449
|
-
|
|
450
|
-
|
|
309
|
+
- **`X-Session-Reset` is ignored on a request whose last non-system message is a `tool` result.** On
|
|
310
|
+
that shape the reset stops nothing and creates nothing, and the turn is treated as a resumed one:
|
|
311
|
+
the system prompt is skipped if the thread is live, and tool results are scoped as for a live
|
|
312
|
+
thread. Every other shape,
|
|
313
|
+
`[..., tool, user]` included, honours it. To reset in the middle of a tool loop, send the header on
|
|
314
|
+
a turn that ends in a `user` message.
|
|
451
315
|
|
|
452
316
|
## Environment variables
|
|
453
317
|
|
|
454
|
-
| Variable | Default
|
|
455
|
-
| ----------------------------------- |
|
|
456
|
-
| `OPENCLAW_SERVER_TOKEN` | (
|
|
457
|
-
| `OPENCLAW_RATE_LIMIT` | `300`
|
|
458
|
-
| `OPENCLAW_CORS_ORIGINS` | (loopback only)
|
|
459
|
-
| `OPENAI_COMPAT_NEW_CONVO_HEURISTIC` | (unset)
|
|
460
|
-
| `OPENAI_COMPAT_TOOLS_PER_MESSAGE` | (unset)
|
|
461
|
-
| `OPENAI_COMPAT_STATUS_URL` | (unset)
|
|
462
|
-
| `OPENCLAW_SERVE_MAX_SESSIONS` | `32`
|
|
463
|
-
| `OPENCLAW_SERVE_TTL_MINUTES` | `60`
|
|
318
|
+
| Variable | Default | Purpose |
|
|
319
|
+
| ----------------------------------- | ---------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
320
|
+
| `OPENCLAW_SERVER_TOKEN` | (auto-generated) | Overrides the auto-generated bearer token; `disabled` turns auth off. The active token is written to `~/.openclaw/server-token` for the CLI. |
|
|
321
|
+
| `OPENCLAW_RATE_LIMIT` | `300` | Max requests per IP per 60-second sliding window. |
|
|
322
|
+
| `OPENCLAW_CORS_ORIGINS` | (loopback only) | Set to `*` to allow all origins (the `/v1/*` paths already do this). |
|
|
323
|
+
| `OPENAI_COMPAT_NEW_CONVO_HEURISTIC` | (unset) | Set to `1` to enable webchat mode (see above). |
|
|
324
|
+
| `OPENAI_COMPAT_TOOLS_PER_MESSAGE` | (unset) | Set to `1` to re-send the full tool schemas on every turn (see [Tool definitions](#tool-definitions-and-where-they-live)). Needed only when the tool set changes mid-conversation; costs per-turn prompt growth. |
|
|
325
|
+
| `OPENAI_COMPAT_STATUS_URL` | (unset) | If set, the bridge POSTs JSON status updates to this URL (fire-and-forget, 2s timeout). See [Status webhook](#status-webhook). |
|
|
326
|
+
| `OPENCLAW_SERVE_MAX_SESSIONS` | `32` | Max concurrent OpenAI-compat sessions in serve mode. The plugin default is 5; serve mode raises it because each distinct caller gets its own `sys-<hash>` session. |
|
|
327
|
+
| `OPENCLAW_SERVE_TTL_MINUTES` | `60` | Idle TTL for OpenAI-compat sessions in serve mode. Idle sessions are reaped by a 60s background loop and are not resumed from disk. |
|
|
464
328
|
|
|
465
329
|
## Inspection: `GET /v1/sessions`
|
|
466
330
|
|
|
@@ -480,7 +344,7 @@ Sample response:
|
|
|
480
344
|
{
|
|
481
345
|
"key": "sys-a3f81c9d0b27",
|
|
482
346
|
"session_name": "openai-sys-a3f81c9d0b27",
|
|
483
|
-
"model": "claude-opus-
|
|
347
|
+
"model": "claude-opus-5-5",
|
|
484
348
|
"cwd": "/home/user/projects",
|
|
485
349
|
"created": "2026-04-09T03:12:18.441Z",
|
|
486
350
|
"turns": 68,
|
|
@@ -512,7 +376,7 @@ Run after standing up the server. Set `TOKEN=$(cat ~/.openclaw/server-token)` fi
|
|
|
512
376
|
for SYS in 'You are Alice.' 'You are Bob.'; do
|
|
513
377
|
curl -s http://127.0.0.1:18796/v1/chat/completions \
|
|
514
378
|
-H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
|
|
515
|
-
-d "{\"model\":\"
|
|
379
|
+
-d "{\"model\":\"opus\",\"messages\":[{\"role\":\"system\",\"content\":\"$SYS\"},{\"role\":\"user\",\"content\":\"hi\"}]}" \
|
|
516
380
|
| jq -r '.id'
|
|
517
381
|
done
|
|
518
382
|
curl -s http://127.0.0.1:18796/v1/sessions -H "Authorization: Bearer $TOKEN" \
|
|
@@ -523,7 +387,7 @@ curl -s http://127.0.0.1:18796/v1/sessions -H "Authorization: Bearer $TOKEN" \
|
|
|
523
387
|
**2. Same system prompt + different model produces two sessions.**
|
|
524
388
|
|
|
525
389
|
```bash
|
|
526
|
-
for M in
|
|
390
|
+
for M in opus sonnet; do
|
|
527
391
|
curl -s http://127.0.0.1:18796/v1/chat/completions \
|
|
528
392
|
-H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
|
|
529
393
|
-d "{\"model\":\"$M\",\"messages\":[{\"role\":\"system\",\"content\":\"SAME\"},{\"role\":\"user\",\"content\":\"hi\"}]}" > /dev/null
|
|
@@ -538,10 +402,10 @@ curl -s http://127.0.0.1:18796/v1/sessions -H "Authorization: Bearer $TOKEN" | j
|
|
|
538
402
|
SID=smoke-reset
|
|
539
403
|
curl -s http://127.0.0.1:18796/v1/chat/completions \
|
|
540
404
|
-H "Authorization: Bearer $TOKEN" -H "X-Session-Id: $SID" -H "Content-Type: application/json" \
|
|
541
|
-
-d '{"model":"
|
|
405
|
+
-d '{"model":"opus","messages":[{"role":"user","content":"remember the word banana"}]}' > /dev/null
|
|
542
406
|
curl -s http://127.0.0.1:18796/v1/chat/completions \
|
|
543
407
|
-H "Authorization: Bearer $TOKEN" -H "X-Session-Id: $SID" -H "X-Session-Reset: 1" -H "Content-Type: application/json" \
|
|
544
|
-
-d '{"model":"
|
|
408
|
+
-d '{"model":"opus","messages":[{"role":"user","content":"what word did I just tell you"}]}' \
|
|
545
409
|
| jq -r '.choices[0].message.content'
|
|
546
410
|
# Expected: model says it has no prior context.
|
|
547
411
|
```
|
|
@@ -554,7 +418,7 @@ PREAMBLE=$(printf 'x%.0s' {1..3000})
|
|
|
554
418
|
for i in 1 2 3 4; do
|
|
555
419
|
curl -s http://127.0.0.1:18796/v1/chat/completions \
|
|
556
420
|
-H "Authorization: Bearer $TOKEN" -H "X-Session-Id: $SID" -H "Content-Type: application/json" \
|
|
557
|
-
-d "{\"model\":\"
|
|
421
|
+
-d "{\"model\":\"opus\",\"messages\":[{\"role\":\"system\",\"content\":\"long preamble: $PREAMBLE\"},{\"role\":\"user\",\"content\":\"turn $i\"}]}" > /dev/null
|
|
558
422
|
curl -s http://127.0.0.1:18796/v1/sessions -H "Authorization: Bearer $TOKEN" \
|
|
559
423
|
| jq ".data[] | select(.session_name == \"openai-$SID\") | {turn: $i, cached_tokens, tokens_in}"
|
|
560
424
|
done
|
|
@@ -573,10 +437,12 @@ Errors use the OpenAI error envelope:
|
|
|
573
437
|
| Status | When |
|
|
574
438
|
| ------ | ---------------------------------------------------------------------- |
|
|
575
439
|
| 400 | `messages` empty/missing, no user message, invalid `max_tokens` |
|
|
440
|
+
| 413 | Request body over 5 MiB |
|
|
576
441
|
| 401 | Missing or wrong bearer token (when auth enabled) |
|
|
577
442
|
| 415 | POST without `Content-Type: application/json` |
|
|
578
443
|
| 429 | Rate limited (`OPENCLAW_RATE_LIMIT` exceeded) |
|
|
579
444
|
| 503 | Failed to start a new session (model unavailable, CLI crashed at boot) |
|
|
445
|
+
| 502 | The CLI finished the turn with an error (`type: upstream_error`) |
|
|
580
446
|
| 500 | Mid-turn failure |
|
|
581
447
|
|
|
582
448
|
## Related
|