glm-acp-agent 1.1.3 → 1.2.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (37) hide show
  1. package/README.md +118 -51
  2. package/dist/index.js +2 -2
  3. package/dist/llm/glm-client.d.ts +59 -1
  4. package/dist/llm/glm-client.d.ts.map +1 -1
  5. package/dist/llm/glm-client.js +135 -15
  6. package/dist/llm/glm-client.js.map +1 -1
  7. package/dist/protocol/agent.d.ts +16 -3
  8. package/dist/protocol/agent.d.ts.map +1 -1
  9. package/dist/protocol/agent.js +422 -49
  10. package/dist/protocol/agent.js.map +1 -1
  11. package/dist/protocol/image-preprocessor.js +1 -1
  12. package/dist/protocol/image-preprocessor.js.map +1 -1
  13. package/dist/protocol/session-store.d.ts +14 -2
  14. package/dist/protocol/session-store.d.ts.map +1 -1
  15. package/dist/protocol/session-store.js +27 -6
  16. package/dist/protocol/session-store.js.map +1 -1
  17. package/dist/protocol/system-prompt.d.ts.map +1 -1
  18. package/dist/protocol/system-prompt.js +1 -3
  19. package/dist/protocol/system-prompt.js.map +1 -1
  20. package/dist/tests/agent.test.js +395 -4
  21. package/dist/tests/agent.test.js.map +1 -1
  22. package/dist/tests/executor.test.js +69 -2
  23. package/dist/tests/executor.test.js.map +1 -1
  24. package/dist/tests/glm-client.test.js +152 -2
  25. package/dist/tests/glm-client.test.js.map +1 -1
  26. package/dist/tests/integration.test.js +50 -0
  27. package/dist/tests/integration.test.js.map +1 -1
  28. package/dist/tools/executor.d.ts +10 -1
  29. package/dist/tools/executor.d.ts.map +1 -1
  30. package/dist/tools/executor.js +77 -53
  31. package/dist/tools/executor.js.map +1 -1
  32. package/dist/tools/session-mcp-client.js +3 -1
  33. package/dist/tools/session-mcp-client.js.map +1 -1
  34. package/dist/tools/vision-mcp-client.d.ts.map +1 -1
  35. package/dist/tools/vision-mcp-client.js +6 -2
  36. package/dist/tools/vision-mcp-client.js.map +1 -1
  37. package/package.json +8 -4
package/README.md CHANGED
@@ -1,6 +1,6 @@
1
1
  # glm-acp-agent
2
2
 
3
- An [Agent Client Protocol (ACP)](https://agentclientprotocol.com) agent written in TypeScript that uses the **Z.AI / Zhipu AI GLM** model family (GLM-5.1, GLM-4.7, GLM-4.6, …) as its reasoning core.
3
+ An [Agent Client Protocol (ACP)](https://agentclientprotocol.com) agent written in TypeScript that uses the **Z.AI / Zhipu AI GLM** model family (GLM-5.2, GLM-5.1, GLM-4.7, …) as its reasoning core.
4
4
 
5
5
  The agent connects to any ACP-compatible IDE or client over **stdio**, streams responses back in real time, and can call a rich set of tools to interact with the user's file system, terminal, and the web.
6
6
 
@@ -28,12 +28,13 @@ Built-in web tools use Coding Plan-compatible MCP endpoints, not the general `/a
28
28
  - **Streaming** – assistant text and reasoning tokens are forwarded as incremental ACP chunks
29
29
  - **Tool calling** – agentic loop with up to 20 turns of GLM function calling
30
30
  - **Thinking mode** – GLM's `reasoning_content` tokens are surfaced as `agent_thought_chunk` blocks so the client can show the model's chain of thought
31
+ - **Session permission modes** – supports `default`, `accept_edits`, and `bypass_permissions` via `session/set_mode`. Clients like DevFlow can use this to toggle between prompting for every edit, auto-approving edits while prompting for commands, or bypassing permissions entirely.
31
32
  - **Per-session model switching** – `session/set_model` lets clients change the active GLM model mid-conversation; `session/new` returns the curated `availableModels` list
32
- - **Image input via Coding Plan Vision MCP** – `promptCapabilities.image` is advertised; pasted ACP image blocks are routed through Z.AI Vision MCP (`@z_ai/mcp-server`) and the resulting analysis is fed into the main coding model. Direct chat-image (e.g. `glm-4v-plus`) calls are intentionally not used.
33
+ - **Image input via Coding Plan-native vision or Vision MCP** – `promptCapabilities.image` is advertised; `glm-5v-turbo` sessions send supported image parts directly to the model, while non-native coding models route pasted ACP image blocks through Z.AI Vision MCP (`@z_ai/mcp-server`). Direct chat-image-only models (e.g. `glm-4v-plus`) are intentionally not used.
33
34
  - **Session persistence** – conversations are written to `~/.local/state/glm-acp-agent/sessions/` and can be reloaded via `session/load`, branched via `session/fork`, or resumed without replay via `session/resume`
34
35
  - **Six built-in tools** (see below)
35
36
  - **Self-sufficient local tools** – file reads/writes, directory listings, and shell commands run in the agent process, so they do not depend on ACP client `fs` or `terminal` capabilities
36
- - **Permissioned writes / commands** – every `write_file` and `run_command` call asks the user via `session/request_permission` before doing anything
37
+ - **Configurable permissions** – `write_file` and `run_command` behavior depends on the active session mode (prompts by default)
37
38
  - **Protocol-correct stop reasons** – maps model and runtime conditions to ACP `end_turn`, `max_tokens`, `max_turn_requests`, `refusal`, and `cancelled`
38
39
  - **Protocol-correct tool statuses** – `pending` → `in_progress` → `completed` / `failed`
39
40
  - **Token usage reporting** – aggregated usage is returned on the `session/prompt` response
@@ -65,15 +66,27 @@ The agent process needs network access to `api.z.ai` for chat completions and We
65
66
 
66
67
  ## Available Tools
67
68
 
68
- | Tool | Runs on | Requires client capability | Description |
69
- |------|---------|----------------------------|-------------|
70
- | `read_file` | Agent process | – | Read the text content of a file |
71
- | `write_file` | Agent process | ACP permission prompt | Write or overwrite a text file (asks for permission) |
72
- | `list_files` | Agent process | – | List a directory using Node filesystem APIs |
73
- | `run_command` | Agent process | ACP permission prompt | Run an arbitrary shell command via `sh -c` in the session cwd (asks for permission) |
74
- | `web_search` | Agent (Z.AI Coding Plan MCP) | – | Search the web — returns titles, URLs, and summaries |
75
- | `web_reader` | Agent (Z.AI Coding Plan MCP) | – | Fetch and parse a web page (markdown or plain text) |
76
- | `image_analysis` | Agent (Z.AI Vision MCP, stdio) | – | Analyze a local image path or remote URL using `@z_ai/mcp-server` |
69
+ | Tool | Runs on | Permission behavior | Description |
70
+ |------|---------|---------------------|-------------|
71
+ | `read_file` | Agent process | Always silent | Read the text content of a file |
72
+ | `write_file` | Agent process | Mode-dependent | Write or overwrite a text file. Silent in `accept_edits` and `bypass_permissions`. |
73
+ | `list_files` | Agent process | Always silent | List a directory using Node filesystem APIs |
74
+ | `run_command` | Agent process | Mode-dependent | Run an arbitrary shell command. Silent only in `bypass_permissions`. |
75
+ | `web_search` | Agent (Z.AI Coding Plan MCP) | Always silent | Search the web — returns titles, URLs, and summaries |
76
+ | `web_reader` | Agent (Z.AI Coding Plan MCP) | Always silent | Fetch and parse a web page (markdown or plain text) |
77
+ | `image_analysis` | Agent (Z.AI Vision MCP, stdio) | Always silent | Analyze a local image path or remote URL using `@z_ai/mcp-server` |
78
+
79
+ ### Session Modes
80
+
81
+ Clients can use `session/set_mode` to drive the permission policy:
82
+
83
+ | Mode ID | Name | `write_file` | `run_command` |
84
+ |---|---|---|---|
85
+ | `default` | Ask for permission | **Prompt** | **Prompt** |
86
+ | `accept_edits` | Auto-approve edits | Silent | **Prompt** |
87
+ | `bypass_permissions` | Bypass all permissions | Silent | Silent |
88
+
89
+ Reads, listings, and MCP tool calls are always silent across all modes.
77
90
 
78
91
  ---
79
92
 
@@ -87,6 +100,27 @@ The agent process needs network access to `api.z.ai` for chat completions and We
87
100
 
88
101
  ## Installation
89
102
 
103
+ ### Quick Start
104
+
105
+ Install the published package globally to get the `glm-acp-agent` command on your `PATH`:
106
+
107
+ ```bash
108
+ npm install -g glm-acp-agent@latest
109
+ ```
110
+
111
+ Then either export your Z.AI API key and run it directly:
112
+
113
+ ```bash
114
+ export Z_AI_API_KEY=your_key_here
115
+ glm-acp-agent
116
+ ```
117
+
118
+ …or run the interactive setup once to persist the key to disk (see [One-time setup](#one-time-setup)) and point any ACP-compatible client at the `glm-acp-agent` command.
119
+
120
+ ### From Source (Development)
121
+
122
+ Clone the repository if you want to hack on the agent or pin a specific commit:
123
+
90
124
  ```bash
91
125
  git clone https://github.com/stefandevo/glm-acp-agent.git
92
126
  cd glm-acp-agent
@@ -94,6 +128,8 @@ npm install
94
128
  npm run build
95
129
  ```
96
130
 
131
+ The build output lands in `dist/`; the rest of this README uses `node dist/index.js` whenever it refers to the source-build entry point.
132
+
97
133
  ---
98
134
 
99
135
  ## Configuration
@@ -103,7 +139,7 @@ The agent reads its configuration from environment variables, plus an optional c
103
139
  | Variable | Required | Default | Description |
104
140
  |----------|----------|---------|-------------|
105
141
  | `Z_AI_API_KEY` | One of env / `--setup` | — | API key for the Z.AI / Zhipu AI service. If unset, the credentials file is consulted. |
106
- | `ACP_GLM_MODEL` | No | `glm-5.1` | Default GLM model for new sessions |
142
+ | `ACP_GLM_MODEL` | No | `glm-5.2` | Default GLM model for new sessions |
107
143
  | `ACP_GLM_AVAILABLE_MODELS` | No | built-in list | Comma-separated list of model ids advertised in `session/set_model` |
108
144
  | `ACP_GLM_BASE_URL` | No | `https://api.z.ai/api/coding/paas/v4` | Override the API base URL |
109
145
  | `ACP_GLM_MAX_TOKENS` | No | `8192` | Cap on `max_tokens` for each completion |
@@ -117,13 +153,15 @@ The agent reads its configuration from environment variables, plus an optional c
117
153
  If you'd rather not pass `Z_AI_API_KEY` through your ACP client's environment block, run the interactive setup once and the agent will read the key from disk on subsequent launches:
118
154
 
119
155
  ```bash
120
- # From the project directory (after npm run build)
121
- node dist/index.js --setup
122
-
123
- # Or, if you installed globally
124
156
  glm-acp-agent --setup
125
157
  ```
126
158
 
159
+ If you're working from a source clone instead of the published package, use the build output directly:
160
+
161
+ ```bash
162
+ node dist/index.js --setup
163
+ ```
164
+
127
165
  The key is written to `$XDG_CONFIG_HOME/glm-acp-agent/credentials.json` (default: `~/.config/glm-acp-agent/credentials.json`) with `0600` permissions. The `Z_AI_API_KEY` environment variable, when set, always wins over the file.
128
166
 
129
167
  ### Supported models
@@ -132,20 +170,37 @@ The agent advertises only the models on the current Z.AI Coding Plan allowlist:
132
170
 
133
171
  | Model | Notes |
134
172
  |-------|-------|
135
- | `glm-5.1` | **Default.** Long-horizon coding model; thinking mode auto-enabled |
173
+ | `glm-5.2` | **Default.** Newest 1M-context coding model; thinking mode auto-enabled |
174
+ | `glm-5.1` | Long-horizon coding model; thinking mode auto-enabled |
136
175
  | `glm-5-turbo` | Faster Coding Plan reasoning model |
176
+ | `glm-5v-turbo` | Multimodal Coding Plan model; native image understanding for jpg / jpeg / png inputs |
137
177
  | `glm-4.7` | 200K-context reasoning model |
138
178
  | `glm-4.5-air` | Lightweight, lower-latency model |
139
179
 
140
- `ACP_GLM_AVAILABLE_MODELS` still lets you advertise custom IDs, but custom IDs sit outside the supported Coding Plan list — the Coding Plan endpoint will reject any model code Z.AI hasn't whitelisted (business code `1211`).
180
+ `ACP_GLM_AVAILABLE_MODELS` still lets you advertise custom IDs, but custom IDs sit outside the supported Coding Plan list — the Coding Plan endpoint will reject any model code Z.AI hasn't whitelisted (business code `1211`). If you override the model list, include `glm-5v-turbo` yourself to keep it visible in the picker.
181
+
182
+ Vision-only chat models (`glm-4v-plus` etc.) are **not** advertised. The multimodal coding model `glm-5v-turbo` is advertised because it is on the Coding Plan and receives supported ACP image blocks directly as native `image_url` content parts. Other advertised models keep using the [Vision MCP](#vision-mcp) path for image analysis.
141
183
 
142
- Vision (`glm-4v-plus` etc.) is **not** advertised as a chat model: vision on the Coding Plan flows through the [Vision MCP](#vision-mcp) section below instead.
184
+ When the model name matches `glm-4.5`, `glm-4.6`, `glm-4.7`, or the `glm-5` family, the agent enables Z.AI's `thinking: { type: "enabled" }` extension and forwards reasoning tokens to the client as `agent_thought_chunk` blocks. This includes `glm-5v-turbo`. Override with `ACP_GLM_THINKING=false` if you want plain completions only.
143
185
 
144
- When the model name matches `glm-4.5`, `glm-4.6`, `glm-4.7`, or `glm-5.x`, the agent enables Z.AI's `thinking: { type: "enabled" }` extension and forwards reasoning tokens to the client as `agent_thought_chunk` blocks. Override with `ACP_GLM_THINKING=false` if you want plain completions only.
186
+ #### Thought level (reasoning effort)
187
+
188
+ The agent advertises a `thought_level` [SessionConfigOption](https://agentclientprotocol.com) so clients that support config options (e.g. a "Thinking" selector) can control reasoning effort per session. The available levels depend on the active model:
189
+
190
+ | Model | Levels | Mapping to the Z.AI request |
191
+ |-------|--------|-----------------------------|
192
+ | `glm-5.2` | `Off` / `High` / `Max` | `Off` → `thinking: { type: "disabled" }`; `High`/`Max` → `thinking: { type: "enabled" }` + `reasoning_effort: "high"` / `"max"` |
193
+ | other thinking-capable models | `Off` / `On` | `Off` → `thinking: { type: "disabled" }`; `On` → `thinking: { type: "enabled" }` |
194
+
195
+ `reasoning_effort` is a GLM-5.2 extra — other models never receive it. New sessions default to the model's own default effort (`Max` on GLM-5.2, which is also Z.AI's default when the field is omitted), so out-of-the-box behaviour is unchanged. Switching models re-clamps the level (e.g. a `High` selection reverts to `On` when you move off GLM-5.2) and pushes a `config_option_update`. The `ACP_GLM_THINKING` env override still wins: `false` forces thinking off regardless of the selected level. Clients that don't support config options simply ignore the advertised option and get the default behaviour.
196
+
197
+ `ACP_GLM_PROMPT_IMAGES=false` still hides the image-attachment capability at session startup. With that flag set, users can pick `glm-5v-turbo` for text work but clients should not offer image attachments.
145
198
 
146
199
  ### Vision MCP
147
200
 
148
- Pasted ACP image blocks are not sent to the chat-completions endpoint. Instead, the agent boots `@z_ai/mcp-server` over stdio (via `npx -y @z_ai/mcp-server@latest`) and calls its `image_analysis` tool. The text result is spliced into the user message as `<image_analysis index="N">…</image_analysis>` so the regular Coding Plan model can reason about it.
201
+ For `glm-5v-turbo`, pasted ACP image blocks with `image/jpeg`, `image/jpg`, or `image/png` are sent directly to chat completions as `image_url` content parts. HTTPS image URLs are forwarded as URLs; inline base64 data is sent as a `data:<mime>;base64,...` URI. Unsupported image MIME types are rejected client-side with an inline `<image_unsupported_format>` annotation so the prompt can continue without a provider 4xx.
202
+
203
+ For non-native models, pasted ACP image blocks are not sent to the chat-completions endpoint. Instead, the agent boots `@z_ai/mcp-server` over stdio (via `npx -y @z_ai/mcp-server@latest`) and calls its `image_analysis` tool. The text result is spliced into the user message as `<image_analysis index="N">…</image_analysis>` so the regular Coding Plan model can reason about it.
149
204
 
150
205
  Prerequisites:
151
206
 
@@ -160,23 +215,26 @@ The model can also call `image_analysis` explicitly with `{ image_source: "/path
160
215
 
161
216
  ### Standalone (stdio)
162
217
 
218
+ If you installed the published package globally, just run the CLI:
219
+
163
220
  ```bash
164
221
  export Z_AI_API_KEY=your_key_here
165
- node dist/index.js
222
+ glm-acp-agent
166
223
  ```
167
224
 
168
- The agent speaks the ACP newline-delimited JSON protocol over stdin/stdout. You can connect any ACP-compatible client to it.
169
-
170
- ### As a global CLI
225
+ If you built from source, invoke the entry point directly:
171
226
 
172
227
  ```bash
173
- npm install -g .
174
228
  export Z_AI_API_KEY=your_key_here
175
- glm-acp-agent
229
+ node dist/index.js
176
230
  ```
177
231
 
232
+ The agent speaks the ACP newline-delimited JSON protocol over stdin/stdout. You can connect any ACP-compatible client to it.
233
+
178
234
  ### Development mode (watch)
179
235
 
236
+ When working from a source clone, run `tsc` in watch mode so `dist/` rebuilds on every save:
237
+
180
238
  ```bash
181
239
  export Z_AI_API_KEY=your_key_here
182
240
  npm run dev # tsc --watch
@@ -196,20 +254,16 @@ npm run dev # tsc --watch
196
254
  - Node.js 20 or later on your `PATH` (`node --version`)
197
255
  - A Z.AI API key — create one at <https://z.ai/manage-apikey/apikey-list>
198
256
 
199
- #### 2. Build the agent
257
+ #### 2. Install the agent
200
258
 
201
- ```bash
202
- git clone https://github.com/stefandevo/glm-acp-agent.git
203
- cd glm-acp-agent
204
- npm install
205
- npm run build
206
- ```
259
+ Follow either path from the [Installation](#installation) section above:
207
260
 
208
- Take note of the absolute path to the freshly built entry point — Zed needs it (no `~`, no `$HOME` shortcuts):
261
+ - **Quick Start** — `npm install -g glm-acp-agent@latest`. Zed can then spawn the agent with `"command": "glm-acp-agent"`.
262
+ - **From source** — clone, `npm install`, `npm run build`. Zed needs the absolute path to the built entry point (no `~`, no `$HOME` shortcuts):
209
263
 
210
- ```bash
211
- echo "$(pwd)/dist/index.js"
212
- ```
264
+ ```bash
265
+ echo "$(pwd)/dist/index.js"
266
+ ```
213
267
 
214
268
  #### 3. Wire it into Zed
215
269
 
@@ -217,6 +271,19 @@ Open (or create) `~/.config/zed/settings.json` and add an `agent_servers` entry.
217
271
 
218
272
  **Option A — inline `env` block (simplest):**
219
273
 
274
+ ```json
275
+ {
276
+ "agent_servers": {
277
+ "glm": {
278
+ "command": "glm-acp-agent",
279
+ "env": { "Z_AI_API_KEY": "sk-…" }
280
+ }
281
+ }
282
+ }
283
+ ```
284
+
285
+ If you installed from source instead of the published package, replace `command` / `args` with the absolute path to the build output:
286
+
220
287
  ```json
221
288
  {
222
289
  "agent_servers": {
@@ -231,26 +298,21 @@ Open (or create) `~/.config/zed/settings.json` and add an `agent_servers` entry.
231
298
 
232
299
  **Option B — credentials file (no key in your editor settings):**
233
300
 
234
- Run the interactive setup once. From the project directory:
235
-
236
- ```bash
237
- node dist/index.js --setup
238
- ```
239
-
240
- …or, if you `npm install -g .`'d the package:
301
+ Run the interactive setup once:
241
302
 
242
303
  ```bash
243
304
  glm-acp-agent --setup
244
305
  ```
245
306
 
307
+ If you're working from a source clone, use the build output directly: `node dist/index.js --setup`.
308
+
246
309
  The key is written to `~/.config/glm-acp-agent/credentials.json` with `0600` permissions. Then drop the `env` block from the Zed entry — the agent will read the file on launch:
247
310
 
248
311
  ```json
249
312
  {
250
313
  "agent_servers": {
251
314
  "glm": {
252
- "command": "node",
253
- "args": ["/absolute/path/to/glm-acp-agent/dist/index.js"]
315
+ "command": "glm-acp-agent"
254
316
  }
255
317
  }
256
318
  }
@@ -264,7 +326,7 @@ If `Z_AI_API_KEY` is set in the environment **and** a credentials file exists, t
264
326
  2. Open the **agent panel** (use the command palette: `agent panel: toggle focus`).
265
327
  3. In the agent picker, select **glm** — Zed labels external agents by their `agent_servers` key.
266
328
  4. Start a new thread and send a small prompt that exercises a tool, e.g. `Read package.json and tell me the project name.`
267
- 5. You should see streaming text, a `read_file` tool call awaiting permission, and (with a thinking-capable model like `glm-5.1`) reasoning surfaced as a separate thought block.
329
+ 5. You should see streaming text, a `read_file` tool call awaiting permission, and (with a thinking-capable model like `glm-5.2`) reasoning surfaced as a separate thought block.
268
330
 
269
331
  #### 5. Iterating on the agent
270
332
 
@@ -288,7 +350,12 @@ For a tighter inner loop, run `npm run dev` in a terminal so `dist/` rebuilds on
288
350
 
289
351
  ### Neovim / VS Code / JetBrains / any ACP client
290
352
 
291
- Any client that supports configuring an ACP agent via a `command` + `args` invocation works the same way: point it at `node /absolute/path/to/glm-acp-agent/dist/index.js` and supply `Z_AI_API_KEY` in the environment.
353
+ Any client that supports configuring an ACP agent via a `command` + `args` invocation works the same way:
354
+
355
+ - If you installed via `npm install -g glm-acp-agent@latest`, set `command` to `glm-acp-agent` (no args required).
356
+ - If you built from source, set `command` to `node` and `args` to `["/absolute/path/to/glm-acp-agent/dist/index.js"]`.
357
+
358
+ Supply `Z_AI_API_KEY` in the environment, or run `glm-acp-agent --setup` once so the agent reads the key from disk on launch.
292
359
 
293
360
  ### Authentication
294
361
 
@@ -353,7 +420,7 @@ The test suite covers:
353
420
 
354
421
  ## Troubleshooting
355
422
 
356
- - **`No API key found.`** — either set `Z_AI_API_KEY` in the environment, or run `node dist/index.js --setup` (or `glm-acp-agent --setup` if installed globally) once to store the key on disk.
423
+ - **`No API key found.`** — either set `Z_AI_API_KEY` in the environment, or run `glm-acp-agent --setup` (or `node dist/index.js --setup` from a source clone) once to store the key on disk.
357
424
  - **`HTTP 401: Invalid API key`** — your key is wrong or expired; rotate it on <https://z.ai/manage-apikey/apikey-list>.
358
425
  - **Writes or commands never get to run.** — make sure your ACP client supports `session/request_permission` and that you approve the prompt for the specific tool call.
359
426
 
package/dist/index.js CHANGED
@@ -9,7 +9,7 @@
9
9
  * Environment variables:
10
10
  * Z_AI_API_KEY - API key for the Z.AI / Zhipu AI service. If unset,
11
11
  * falls back to the credentials file written by --setup.
12
- * ACP_GLM_MODEL - (optional) Override the default model (default: glm-5.1)
12
+ * ACP_GLM_MODEL - (optional) Override the default model (default: glm-5.2)
13
13
  */
14
14
  import { startConnection } from "./protocol/connection.js";
15
15
  import { runSetup } from "./setup.js";
@@ -34,7 +34,7 @@ else if (args.includes("--help") || args.includes("-h")) {
34
34
  "",
35
35
  "Environment variables:",
36
36
  " Z_AI_API_KEY API key (overrides the stored credentials)",
37
- " ACP_GLM_MODEL Default model id (e.g. glm-5.1)",
37
+ " ACP_GLM_MODEL Default model id (e.g. glm-5.2)",
38
38
  " ACP_GLM_AVAILABLE_MODELS Comma-separated list of advertised models",
39
39
  " ACP_GLM_BASE_URL Override the Z.AI API base URL",
40
40
  " ACP_GLM_MAX_TOKENS Per-call max output tokens (default 8192)",
@@ -1,6 +1,32 @@
1
1
  import type { ChatCompletionMessageParam } from "openai/resources/index.js";
2
2
  import type { ModelInfo, Usage } from "@agentclientprotocol/sdk";
3
3
  import { type ToolDefinition } from "../tools/definitions.js";
4
+ /**
5
+ * Reasoning effort levels exposed to ACP clients via the `thought_level`
6
+ * SessionConfigOption. These map onto Z.AI's `thinking` / `reasoning_effort`
7
+ * request parameters (see {@link buildThinkingParams}).
8
+ *
9
+ * GLM-5.2 supports three levels: `none`, `high`, `max`. Other thinking-capable
10
+ * models (GLM-5.1, 5-turbo, 4.7, …) only distinguish thinking on vs. off, so
11
+ * they use `none` and `on`.
12
+ */
13
+ export type ThoughtLevel = "none" | "on" | "high" | "max";
14
+ /** Type guard: whether an arbitrary string is a known ThoughtLevel. */
15
+ export declare function isThoughtLevel(value: string): value is ThoughtLevel;
16
+ /**
17
+ * Resolve which thought-level options a model supports.
18
+ *
19
+ * `reasoning_effort` is a GLM-5.2 exclusive per the Z.AI docs — other models
20
+ * accept the field but it has no effect, so we only expose high/max for 5.2.
21
+ */
22
+ export declare function getThoughtLevels(model: string): ThoughtLevel[];
23
+ /**
24
+ * Resolve a stored ThoughtLevel to one that's valid for the given model.
25
+ * Used when switching models or restoring a persisted session: if the old
26
+ * level isn't in the new model's option list, fall back to the model's
27
+ * default (max for 5.2, on for everything else — the last entry in the list).
28
+ */
29
+ export declare function resolveThoughtLevel(model: string, level: ThoughtLevel): ThoughtLevel;
4
30
  /**
5
31
  * A single message in the GLM conversation history.
6
32
  */
@@ -32,9 +58,22 @@ export interface StreamChatOptions {
32
58
  model: string;
33
59
  /** Tool schemas available in this specific session. */
34
60
  tools?: ToolDefinition[];
61
+ /** Reasoning effort for this call, or undefined to use the model defaults. */
62
+ reasoningEffort?: ThoughtLevel;
35
63
  }
36
64
  /** Default GLM model when neither client nor user has chosen one. */
37
- export declare const DEFAULT_MODEL = "glm-5.1";
65
+ export declare const DEFAULT_MODEL = "glm-5.2";
66
+ /**
67
+ * Z.AI / Zhipu AI error code returned when the total prompt length (messages +
68
+ * tools) exceeds the model's context window.
69
+ */
70
+ export declare const ERR_CONTEXT_OVERFLOW = 1261;
71
+ export declare function isVisionNativeModel(modelId: string): boolean;
72
+ /**
73
+ * Resolve the context window size for a given model ID. Falls back to a safe
74
+ * default (128K) for uncatalogued models.
75
+ */
76
+ export declare function getContextWindow(modelId: string): number;
38
77
  /**
39
78
  * Resolve the list of advertised models, allowing the user to override the
40
79
  * built-in list via `ACP_GLM_AVAILABLE_MODELS`.
@@ -63,4 +102,23 @@ export declare class GlmClient {
63
102
  */
64
103
  streamChat(messages: GlmMessage[], signal?: AbortSignal, options?: StreamChatOptions): AsyncGenerator<GlmStreamChunk>;
65
104
  }
105
+ /**
106
+ * Build the Z.AI `thinking` / `reasoning_effort` extra-body params for a call.
107
+ *
108
+ * Z.AI exposes two parameters:
109
+ * - `thinking` — an on/off gate: `{"type":"enabled"}` or `{"type":"disabled"}`
110
+ * - `reasoning_effort` — GLM-5.2 only; controls thinking depth (`high`/`max`,
111
+ * defaulting to `max` when omitted).
112
+ *
113
+ * {@link ThoughtLevel} values map as follows:
114
+ * - `none` → thinking disabled
115
+ * - `on` → thinking enabled (no reasoning_effort)
116
+ * - `high` → thinking enabled + reasoning_effort=high (5.2 only)
117
+ * - `max` → thinking enabled + reasoning_effort=max (5.2 only)
118
+ *
119
+ * The `ACP_GLM_THINKING` env override still wins: `false` forces thinking off,
120
+ * `true` forces it on (ignoring a `none` level). When `effort` is unset the
121
+ * model defaults are used, preserving the pre-thought-level behaviour.
122
+ */
123
+ export declare function buildThinkingParams(model: string, effort?: ThoughtLevel): Record<string, unknown>;
66
124
  //# sourceMappingURL=glm-client.d.ts.map
@@ -1 +1 @@
1
- {"version":3,"file":"glm-client.d.ts","sourceRoot":"","sources":["../../src/llm/glm-client.ts"],"names":[],"mappings":"AACA,OAAO,KAAK,EAAE,0BAA0B,EAAsB,MAAM,2BAA2B,CAAC;AAChG,OAAO,KAAK,EAAE,SAAS,EAAE,KAAK,EAAE,MAAM,0BAA0B,CAAC;AACjE,OAAO,EAAoB,KAAK,cAAc,EAAE,MAAM,yBAAyB,CAAC;AAIhF;;GAEG;AACH,MAAM,MAAM,UAAU,GAAG,0BAA0B,CAAC;AAEpD;;GAEG;AACH,MAAM,WAAW,cAAc;IAC7B,iCAAiC;IACjC,IAAI,CAAC,EAAE,MAAM,CAAC;IACd,0CAA0C;IAC1C,QAAQ,CAAC,EAAE,MAAM,CAAC;IAClB,6DAA6D;IAC7D,QAAQ,CAAC,EAAE;QACT,EAAE,EAAE,MAAM,CAAC;QACX,IAAI,EAAE,MAAM,CAAC;QACb,SAAS,EAAE,MAAM,CAAC;KACnB,CAAC;IACF,oDAAoD;IACpD,KAAK,CAAC,EAAE,KAAK,CAAC;IACd,kCAAkC;IAClC,IAAI,CAAC,EAAE,OAAO,CAAC;IACf,4BAA4B;IAC5B,UAAU,CAAC,EAAE,MAAM,CAAC;CACrB;AAED,qDAAqD;AACrD,MAAM,WAAW,iBAAiB;IAChC,iDAAiD;IACjD,KAAK,EAAE,MAAM,CAAC;IACd,uDAAuD;IACvD,KAAK,CAAC,EAAE,cAAc,EAAE,CAAC;CAC1B;AAKD,qEAAqE;AACrE,eAAO,MAAM,aAAa,YAAY,CAAC;AAiCvC;;;GAGG;AACH,wBAAgB,kBAAkB,IAAI,SAAS,EAAE,CAYhD;AAED,uEAAuE;AACvE,wBAAgB,eAAe,IAAI,MAAM,CAExC;AAED;;;;;;;GAOG;AACH,qBAAa,SAAS;IACpB,OAAO,CAAC,MAAM,CAAS;IACvB,OAAO,CAAC,SAAS,CAAS;;IAgB1B;;;;;;OAMG;IACI,UAAU,CACf,QAAQ,EAAE,UAAU,EAAE,EACtB,MAAM,CAAC,EAAE,WAAW,EACpB,OAAO,CAAC,EAAE,iBAAiB,GAC1B,cAAc,CAAC,cAAc,CAAC;CA6IlC"}
1
+ {"version":3,"file":"glm-client.d.ts","sourceRoot":"","sources":["../../src/llm/glm-client.ts"],"names":[],"mappings":"AACA,OAAO,KAAK,EAAE,0BAA0B,EAAsB,MAAM,2BAA2B,CAAC;AAChG,OAAO,KAAK,EAAE,SAAS,EAAE,KAAK,EAAE,MAAM,0BAA0B,CAAC;AACjE,OAAO,EAAoB,KAAK,cAAc,EAAE,MAAM,yBAAyB,CAAC;AAIhF;;;;;;;;GAQG;AACH,MAAM,MAAM,YAAY,GAAG,MAAM,GAAG,IAAI,GAAG,MAAM,GAAG,KAAK,CAAC;AAgB1D,uEAAuE;AACvE,wBAAgB,cAAc,CAAC,KAAK,EAAE,MAAM,GAAG,KAAK,IAAI,YAAY,CAEnE;AAED;;;;;GAKG;AACH,wBAAgB,gBAAgB,CAAC,KAAK,EAAE,MAAM,GAAG,YAAY,EAAE,CAE9D;AAED;;;;;GAKG;AACH,wBAAgB,mBAAmB,CAAC,KAAK,EAAE,MAAM,EAAE,KAAK,EAAE,YAAY,GAAG,YAAY,CAGpF;AAED;;GAEG;AACH,MAAM,MAAM,UAAU,GAAG,0BAA0B,CAAC;AAEpD;;GAEG;AACH,MAAM,WAAW,cAAc;IAC7B,iCAAiC;IACjC,IAAI,CAAC,EAAE,MAAM,CAAC;IACd,0CAA0C;IAC1C,QAAQ,CAAC,EAAE,MAAM,CAAC;IAClB,6DAA6D;IAC7D,QAAQ,CAAC,EAAE;QACT,EAAE,EAAE,MAAM,CAAC;QACX,IAAI,EAAE,MAAM,CAAC;QACb,SAAS,EAAE,MAAM,CAAC;KACnB,CAAC;IACF,oDAAoD;IACpD,KAAK,CAAC,EAAE,KAAK,CAAC;IACd,kCAAkC;IAClC,IAAI,CAAC,EAAE,OAAO,CAAC;IACf,4BAA4B;IAC5B,UAAU,CAAC,EAAE,MAAM,CAAC;CACrB;AAED,qDAAqD;AACrD,MAAM,WAAW,iBAAiB;IAChC,iDAAiD;IACjD,KAAK,EAAE,MAAM,CAAC;IACd,uDAAuD;IACvD,KAAK,CAAC,EAAE,cAAc,EAAE,CAAC;IACzB,8EAA8E;IAC9E,eAAe,CAAC,EAAE,YAAY,CAAC;CAChC;AAKD,qEAAqE;AACrE,eAAO,MAAM,aAAa,YAAY,CAAC;AA2CvC;;;GAGG;AACH,eAAO,MAAM,oBAAoB,OAAO,CAAC;AAiBzC,wBAAgB,mBAAmB,CAAC,OAAO,EAAE,MAAM,GAAG,OAAO,CAE5D;AAED;;;GAGG;AACH,wBAAgB,gBAAgB,CAAC,OAAO,EAAE,MAAM,GAAG,MAAM,CAExD;AAED;;;GAGG;AACH,wBAAgB,kBAAkB,IAAI,SAAS,EAAE,CAYhD;AAED,uEAAuE;AACvE,wBAAgB,eAAe,IAAI,MAAM,CAExC;AAED;;;;;;;GAOG;AACH,qBAAa,SAAS;IACpB,OAAO,CAAC,MAAM,CAAS;IACvB,OAAO,CAAC,SAAS,CAAS;;IAgB1B;;;;;;OAMG;IACI,UAAU,CACf,QAAQ,EAAE,UAAU,EAAE,EACtB,MAAM,CAAC,EAAE,WAAW,EACpB,OAAO,CAAC,EAAE,iBAAiB,GAC1B,cAAc,CAAC,cAAc,CAAC;CA0IlC;AA+BD;;;;;;;;;;;;;;;;;GAiBG;AACH,wBAAgB,mBAAmB,CACjC,KAAK,EAAE,MAAM,EACb,MAAM,CAAC,EAAE,YAAY,GACpB,MAAM,CAAC,MAAM,EAAE,OAAO,CAAC,CAiCzB"}
@@ -2,10 +2,43 @@ import OpenAI from "openai";
2
2
  import { TOOL_DEFINITIONS } from "../tools/definitions.js";
3
3
  import { resolveApiKey } from "./credentials.js";
4
4
  import { debug, error } from "./logger.js";
5
+ /** Levels shown when the selected model is GLM-5.2. */
6
+ const LEVELS_52 = ["none", "high", "max"];
7
+ /** Levels shown for every other thinking-capable model. */
8
+ const LEVELS_DEFAULT = ["none", "on"];
9
+ /** Whether a model id is in the GLM-5.2 family (case-insensitive). */
10
+ function isGlm52(model) {
11
+ return model.toLowerCase().startsWith("glm-5.2");
12
+ }
13
+ /** The complete set of thought levels the agent understands. */
14
+ const ALL_THOUGHT_LEVELS = ["none", "on", "high", "max"];
15
+ /** Type guard: whether an arbitrary string is a known ThoughtLevel. */
16
+ export function isThoughtLevel(value) {
17
+ return ALL_THOUGHT_LEVELS.includes(value);
18
+ }
19
+ /**
20
+ * Resolve which thought-level options a model supports.
21
+ *
22
+ * `reasoning_effort` is a GLM-5.2 exclusive per the Z.AI docs — other models
23
+ * accept the field but it has no effect, so we only expose high/max for 5.2.
24
+ */
25
+ export function getThoughtLevels(model) {
26
+ return isGlm52(model) ? LEVELS_52 : LEVELS_DEFAULT;
27
+ }
28
+ /**
29
+ * Resolve a stored ThoughtLevel to one that's valid for the given model.
30
+ * Used when switching models or restoring a persisted session: if the old
31
+ * level isn't in the new model's option list, fall back to the model's
32
+ * default (max for 5.2, on for everything else — the last entry in the list).
33
+ */
34
+ export function resolveThoughtLevel(model, level) {
35
+ const valid = getThoughtLevels(model);
36
+ return valid.includes(level) ? level : valid[valid.length - 1];
37
+ }
5
38
  /** Default base URL for the Z.AI / Zhipu OpenAI-compatible API (Coding endpoint). */
6
39
  const DEFAULT_BASE_URL = "https://api.z.ai/api/coding/paas/v4";
7
40
  /** Default GLM model when neither client nor user has chosen one. */
8
- export const DEFAULT_MODEL = "glm-5.1";
41
+ export const DEFAULT_MODEL = "glm-5.2";
9
42
  /**
10
43
  * Curated list of GLM models the agent advertises to ACP clients via
11
44
  * `SessionModelState.availableModels`. Users can override this list via the
@@ -15,16 +48,26 @@ export const DEFAULT_MODEL = "glm-5.1";
15
48
  * model pickers.
16
49
  */
17
50
  const BUILTIN_AVAILABLE_MODELS = [
51
+ {
52
+ modelId: "glm-5.2",
53
+ name: "GLM-5.2",
54
+ description: "Newest 1M-context coding model with thinking mode",
55
+ },
18
56
  {
19
57
  modelId: "glm-5.1",
20
58
  name: "GLM-5.1",
21
- description: "Latest GLM reasoning model with thinking mode",
59
+ description: "Long-horizon coding model with thinking mode",
22
60
  },
23
61
  {
24
62
  modelId: "glm-5-turbo",
25
63
  name: "GLM-5 Turbo",
26
64
  description: "Faster Coding Plan reasoning model",
27
65
  },
66
+ {
67
+ modelId: "glm-5v-turbo",
68
+ name: "GLM-5V Turbo",
69
+ description: "Multimodal Coding Plan model with native vision",
70
+ },
28
71
  {
29
72
  modelId: "glm-4.7",
30
73
  name: "GLM-4.7",
@@ -36,6 +79,34 @@ const BUILTIN_AVAILABLE_MODELS = [
36
79
  description: "Lightweight, lower-latency model",
37
80
  },
38
81
  ];
82
+ /**
83
+ * Z.AI / Zhipu AI error code returned when the total prompt length (messages +
84
+ * tools) exceeds the model's context window.
85
+ */
86
+ export const ERR_CONTEXT_OVERFLOW = 1261;
87
+ /**
88
+ * Context window sizes (in tokens) for GLM series models.
89
+ */
90
+ const MODEL_METADATA = {
91
+ "glm-5.2": { contextWindow: 1_000_000 },
92
+ "glm-5.1": { contextWindow: 128_000 },
93
+ "glm-5-turbo": { contextWindow: 128_000 },
94
+ "glm-5v-turbo": { contextWindow: 200_000 },
95
+ "glm-4.7": { contextWindow: 200_000 },
96
+ "glm-4.5-air": { contextWindow: 128_000 },
97
+ };
98
+ /** Models that accept image content parts directly through chat completions. */
99
+ const VISION_NATIVE_MODELS = new Set(["glm-5v-turbo"]);
100
+ export function isVisionNativeModel(modelId) {
101
+ return VISION_NATIVE_MODELS.has(modelId.toLowerCase());
102
+ }
103
+ /**
104
+ * Resolve the context window size for a given model ID. Falls back to a safe
105
+ * default (128K) for uncatalogued models.
106
+ */
107
+ export function getContextWindow(modelId) {
108
+ return MODEL_METADATA[modelId]?.contextWindow ?? 128_000;
109
+ }
39
110
  /**
40
111
  * Resolve the list of advertised models, allowing the user to override the
41
112
  * built-in list via `ACP_GLM_AVAILABLE_MODELS`.
@@ -88,7 +159,7 @@ export class GlmClient {
88
159
  */
89
160
  async *streamChat(messages, signal, options) {
90
161
  const model = options?.model ?? getDefaultModel();
91
- const thinkingEnabled = parseThinking(model);
162
+ const thinkingParams = buildThinkingParams(model, options?.reasoningEffort);
92
163
  const tools = (options?.tools ?? TOOL_DEFINITIONS).map((t) => ({
93
164
  type: "function",
94
165
  function: {
@@ -97,13 +168,10 @@ export class GlmClient {
97
168
  parameters: t.function.parameters,
98
169
  },
99
170
  }));
100
- debug(`streamChat: model=${model} baseURL=${this.client.baseURL} messages=${messages.length} tools=${tools.length} thinking=${thinkingEnabled}`);
171
+ debug(`streamChat: model=${model} baseURL=${this.client.baseURL} messages=${messages.length} tools=${tools.length} thinking=${JSON.stringify(thinkingParams)}`);
101
172
  // The OpenAI SDK forwards unknown extra body fields verbatim, so we use
102
- // that to pass GLM-specific fields like `thinking`.
103
- const extraBody = {};
104
- if (thinkingEnabled) {
105
- extraBody["thinking"] = { type: "enabled" };
106
- }
173
+ // that to pass GLM-specific fields like `thinking` and `reasoning_effort`.
174
+ const extraBody = { ...thinkingParams };
107
175
  let stream;
108
176
  try {
109
177
  stream = await this.client.chat.completions.create({
@@ -201,17 +269,69 @@ function parseIntEnv(name, fallback) {
201
269
  return Number.isFinite(n) && n > 0 ? n : fallback;
202
270
  }
203
271
  /**
204
- * Decide whether to enable GLM "thinking" mode for a given model name.
205
- *
206
- * The user can force on/off via `ACP_GLM_THINKING=true|false`. Otherwise,
207
- * we enable thinking for any model whose name suggests it supports it
272
+ * Decide whether a model supports GLM "thinking" mode based on its name
208
273
  * (the GLM-4.5 / GLM-4.6 / GLM-4.7 / GLM-5.x families).
209
274
  */
210
- function parseThinking(model) {
275
+ function modelSupportsThinking(model) {
276
+ return /^glm-(?:4\.[567]|5)/i.test(model);
277
+ }
278
+ /**
279
+ * Parse the `ACP_GLM_THINKING` env override into an explicit boolean, or
280
+ * `undefined` when the variable is unset. `true`/`1` force thinking on;
281
+ * anything else forces it off.
282
+ */
283
+ function thinkingOverride() {
211
284
  const override = process.env["ACP_GLM_THINKING"];
212
285
  if (override !== undefined) {
213
286
  return override.toLowerCase() === "true" || override === "1";
214
287
  }
215
- return /^glm-(?:4\.[567]|5)/i.test(model);
288
+ return undefined;
289
+ }
290
+ /**
291
+ * Build the Z.AI `thinking` / `reasoning_effort` extra-body params for a call.
292
+ *
293
+ * Z.AI exposes two parameters:
294
+ * - `thinking` — an on/off gate: `{"type":"enabled"}` or `{"type":"disabled"}`
295
+ * - `reasoning_effort` — GLM-5.2 only; controls thinking depth (`high`/`max`,
296
+ * defaulting to `max` when omitted).
297
+ *
298
+ * {@link ThoughtLevel} values map as follows:
299
+ * - `none` → thinking disabled
300
+ * - `on` → thinking enabled (no reasoning_effort)
301
+ * - `high` → thinking enabled + reasoning_effort=high (5.2 only)
302
+ * - `max` → thinking enabled + reasoning_effort=max (5.2 only)
303
+ *
304
+ * The `ACP_GLM_THINKING` env override still wins: `false` forces thinking off,
305
+ * `true` forces it on (ignoring a `none` level). When `effort` is unset the
306
+ * model defaults are used, preserving the pre-thought-level behaviour.
307
+ */
308
+ export function buildThinkingParams(model, effort) {
309
+ const params = {};
310
+ const override = thinkingOverride();
311
+ const modelCanThink = modelSupportsThinking(model);
312
+ // Explicit env override to disable wins over everything.
313
+ if (override === false) {
314
+ if (modelCanThink)
315
+ params["thinking"] = { type: "disabled" };
316
+ return params;
317
+ }
318
+ const canThink = override === true || modelCanThink;
319
+ // A `none` level disables thinking, unless the env override forces it on.
320
+ if (effort === "none" && override !== true) {
321
+ if (canThink)
322
+ params["thinking"] = { type: "disabled" };
323
+ return params;
324
+ }
325
+ if (!canThink)
326
+ return params;
327
+ params["thinking"] = { type: "enabled" };
328
+ // reasoning_effort is only meaningful for GLM-5.2 per the Z.AI docs, and Z.AI
329
+ // only accepts the "high"/"max" values there. Other levels ("on", and "none"
330
+ // when thinking is force-enabled via the env override) carry no valid
331
+ // reasoning_effort, so we omit the field entirely.
332
+ if ((effort === "high" || effort === "max") && isGlm52(model)) {
333
+ params["reasoning_effort"] = effort;
334
+ }
335
+ return params;
216
336
  }
217
337
  //# sourceMappingURL=glm-client.js.map