glm-acp-agent 1.7.0 → 1.8.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +34 -1
- package/dist/protocol/agent.d.ts +16 -0
- package/dist/protocol/agent.d.ts.map +1 -1
- package/dist/protocol/agent.js +215 -39
- package/dist/protocol/agent.js.map +1 -1
- package/dist/protocol/session-store.d.ts +12 -1
- package/dist/protocol/session-store.d.ts.map +1 -1
- package/dist/protocol/session-store.js +10 -4
- package/dist/protocol/session-store.js.map +1 -1
- package/dist/protocol/slash-commands.d.ts +44 -0
- package/dist/protocol/slash-commands.d.ts.map +1 -0
- package/dist/protocol/slash-commands.js +251 -0
- package/dist/protocol/slash-commands.js.map +1 -0
- package/dist/tests/agent.test.js +644 -6
- package/dist/tests/agent.test.js.map +1 -1
- package/dist/tests/integration.test.js +39 -1
- package/dist/tests/integration.test.js.map +1 -1
- package/dist/tests/slash-commands.test.d.ts +2 -0
- package/dist/tests/slash-commands.test.d.ts.map +1 -0
- package/dist/tests/slash-commands.test.js +214 -0
- package/dist/tests/slash-commands.test.js.map +1 -0
- package/package.json +1 -1
package/README.md
CHANGED
|
@@ -24,12 +24,13 @@ Built-in web tools use Coding Plan-compatible MCP endpoints, not the general `/a
|
|
|
24
24
|
|
|
25
25
|
## Features
|
|
26
26
|
|
|
27
|
-
- **Full ACP compliance** – implements `initialize`, `authenticate`, `session/new`, `session/set_mode`, `session/prompt`, `session/cancel`, `session/close`, `session/list`, `session/load`, `session/fork`, `session/resume`, and `session/set_model`
|
|
27
|
+
- **Full ACP compliance** – implements `initialize`, `authenticate`, `session/new`, `session/set_mode`, `session/prompt`, `session/cancel`, `session/close`, `session/list`, `session/load`, `session/fork`, `session/resume`, and `session/set_model`, and pushes `session_info_update`, `config_option_update`, `current_mode_update`, and `available_commands_update` notifications
|
|
28
28
|
- **Streaming** – assistant text and reasoning tokens are forwarded as incremental ACP chunks
|
|
29
29
|
- **Tool calling** – agentic loop with a configurable cap of GLM function-calling turns (default 20; see `ACP_GLM_MAX_TURNS` / `--max-turns`)
|
|
30
30
|
- **Thinking mode** – GLM's `reasoning_content` tokens are surfaced as `agent_thought_chunk` blocks so the client can show the model's chain of thought
|
|
31
31
|
- **Session permission modes** – supports `default`, `accept_edits`, and `bypass_permissions` via `session/set_mode`. Clients like DevFlow can use this to toggle between prompting for every edit, auto-approving edits while prompting for commands, or bypassing permissions entirely.
|
|
32
32
|
- **Per-session model switching** – `session/set_model` lets clients change the active GLM model mid-conversation; `session/new` returns the curated `availableModels` list
|
|
33
|
+
- **Slash commands** – commands and skills found under the session's `.claude/` directory are advertised to the client with `available_commands_update`, so `/` autocomplete is populated (see [Slash commands](#slash-commands))
|
|
33
34
|
- **Image input via Coding Plan-native vision or Vision MCP** – `promptCapabilities.image` is advertised; `glm-5.3-flash` sends supported pasted ACP image blocks directly as native `image_url` content parts, while the other advertised coding models (including default `glm-5.3`) route them through Z.AI Vision MCP (`@z_ai/mcp-server`). `glm-5v-turbo` keeps the same native-vision path when re-added via `ACP_GLM_AVAILABLE_MODELS` — it is no longer on the Coding Plan allowlist. Direct chat-image-only models (e.g. `glm-4v-plus`) are intentionally not used.
|
|
34
35
|
- **Session persistence** – conversations are written to `~/.local/state/glm-acp-agent/sessions/` and can be reloaded via `session/load`, branched via `session/fork`, or resumed without replay via `session/resume`
|
|
35
36
|
- **Seven built-in tools** (see below)
|
|
@@ -90,6 +91,36 @@ Reads, listings, and MCP tool calls are always silent across all modes.
|
|
|
90
91
|
|
|
91
92
|
---
|
|
92
93
|
|
|
94
|
+
## Slash commands
|
|
95
|
+
|
|
96
|
+
After `session/new`, `session/load`, `session/fork`, and `session/resume`, the agent
|
|
97
|
+
sends an ACP [`available_commands_update`](https://agentclientprotocol.com/protocol/v2/slash-commands)
|
|
98
|
+
notification. Clients such as Zed and Paseo use it to populate their `/` autocomplete.
|
|
99
|
+
Each notification is a full snapshot that replaces the previous list.
|
|
100
|
+
|
|
101
|
+
The snapshot is built from the same on-disk layout Claude Code uses, read from the
|
|
102
|
+
session's working directory and from `~/.claude` for user-level definitions:
|
|
103
|
+
|
|
104
|
+
| Path | Command name |
|
|
105
|
+
|---|---|
|
|
106
|
+
| `<cwd>/.claude/commands/commit.md` | `/commit` |
|
|
107
|
+
| `<cwd>/.claude/commands/review/pr.md` | `/review:pr` |
|
|
108
|
+
| `<cwd>/.claude/skills/audit/SKILL.md` | `/audit` |
|
|
109
|
+
|
|
110
|
+
Project definitions shadow user-level ones with the same name. The `description`
|
|
111
|
+
and `argument-hint` keys of a file's YAML frontmatter become the menu entry's
|
|
112
|
+
description and input hint; without a `description`, the first heading (then the
|
|
113
|
+
first line) of the body is used instead.
|
|
114
|
+
|
|
115
|
+
Advertised names are sent **without** a leading slash — the client prepends it for
|
|
116
|
+
display and invokes the command by sending `/name …` as ordinary prompt text. When
|
|
117
|
+
that text names an advertised command, the agent injects the definition's body into
|
|
118
|
+
the user message, substituting `$ARGUMENTS` where the file asks for it and otherwise
|
|
119
|
+
appending the typed arguments. Unknown `/foo` is left untouched and reaches the model
|
|
120
|
+
as prose, so typing a slash by accident never fails the turn.
|
|
121
|
+
|
|
122
|
+
---
|
|
123
|
+
|
|
93
124
|
## Prerequisites
|
|
94
125
|
|
|
95
126
|
- **Node.js** 20 or later (native `fetch` and Web Streams required)
|
|
@@ -386,6 +417,7 @@ src/
|
|
|
386
417
|
├── protocol/
|
|
387
418
|
│ ├── connection.ts # Sets up the ACP stdio connection
|
|
388
419
|
│ ├── agent.ts # GlmAcpAgent – ACP protocol implementation
|
|
420
|
+
│ ├── slash-commands.ts # Discovers .claude commands/skills; expands `/name`
|
|
389
421
|
│ └── session-store.ts # On-disk persistence for load/fork/resume
|
|
390
422
|
├── tools/
|
|
391
423
|
│ ├── definitions.ts # Tool JSON schemas (function-calling format)
|
|
@@ -395,6 +427,7 @@ src/
|
|
|
395
427
|
├── credentials.test.ts # Credential resolution and --setup persistence
|
|
396
428
|
├── executor.test.ts # Tests for ToolExecutor
|
|
397
429
|
├── glm-client.test.ts # Tests for streaming / tool-call assembly
|
|
430
|
+
├── slash-commands.test.ts # Command discovery and `/name` expansion
|
|
398
431
|
└── integration.test.ts # End-to-end tests over the real ACP ndjson transport
|
|
399
432
|
```
|
|
400
433
|
|
package/dist/protocol/agent.d.ts
CHANGED
|
@@ -58,6 +58,22 @@ export declare class GlmAcpAgent implements Agent {
|
|
|
58
58
|
initialize(params: InitializeRequest): Promise<InitializeResponse>;
|
|
59
59
|
authenticate(params: AuthenticateRequest): Promise<AuthenticateResponse>;
|
|
60
60
|
newSession(params: NewSessionRequest): Promise<NewSessionResponse>;
|
|
61
|
+
/**
|
|
62
|
+
* Queue an `available_commands_update` snapshot for a session.
|
|
63
|
+
*
|
|
64
|
+
* The notification is the only channel ACP gives us for slash-command
|
|
65
|
+
* autocomplete — it is not part of any method's response — so every session
|
|
66
|
+
* entry point (create / load / fork / resume) has to send one.
|
|
67
|
+
*
|
|
68
|
+
* It is deliberately *deferred* rather than awaited inline: a client learns a
|
|
69
|
+
* session's id from the `session/new` / `session/fork` response, so a
|
|
70
|
+
* notification written ahead of that response arrives for a session the client
|
|
71
|
+
* has never heard of, and clients drop those. Sending on the next macrotask
|
|
72
|
+
* puts it behind the response the caller is about to return (and, on load,
|
|
73
|
+
* behind the replayed transcript), which is also when a client is ready to
|
|
74
|
+
* paint the menu. Each send replaces the previous list wholesale.
|
|
75
|
+
*/
|
|
76
|
+
private scheduleAvailableCommands;
|
|
61
77
|
unstable_setSessionModel(params: SetSessionModelRequest): Promise<SetSessionModelResponse>;
|
|
62
78
|
/**
|
|
63
79
|
* Switch a session to a new model id — shared by `session/set_model` and the
|
|
@@ -1 +1 @@
|
|
|
1
|
-
{"version":3,"file":"agent.d.ts","sourceRoot":"","sources":["../../src/protocol/agent.ts"],"names":[],"mappings":"AAIA,OAAO,KAAK,EACV,KAAK,EACL,mBAAmB,EAEnB,iBAAiB,EACjB,kBAAkB,EAClB,iBAAiB,EACjB,kBAAkB,EAClB,aAAa,EACb,cAAc,EACd,kBAAkB,EAClB,mBAAmB,EACnB,oBAAoB,EACpB,qBAAqB,EACrB,sBAAsB,EACtB,mBAAmB,EACnB,mBAAmB,EACnB,oBAAoB,EACpB,kBAAkB,EAClB,mBAAmB,EACnB,kBAAkB,EAClB,mBAAmB,EACnB,oBAAoB,EACpB,qBAAqB,EACrB,sBAAsB,EACtB,uBAAuB,EACvB,6BAA6B,EAC7B,8BAA8B,
|
|
1
|
+
{"version":3,"file":"agent.d.ts","sourceRoot":"","sources":["../../src/protocol/agent.ts"],"names":[],"mappings":"AAIA,OAAO,KAAK,EACV,KAAK,EACL,mBAAmB,EAEnB,iBAAiB,EACjB,kBAAkB,EAClB,iBAAiB,EACjB,kBAAkB,EAClB,aAAa,EACb,cAAc,EACd,kBAAkB,EAClB,mBAAmB,EACnB,oBAAoB,EACpB,qBAAqB,EACrB,sBAAsB,EACtB,mBAAmB,EACnB,mBAAmB,EACnB,oBAAoB,EACpB,kBAAkB,EAClB,mBAAmB,EACnB,kBAAkB,EAClB,mBAAmB,EACnB,oBAAoB,EACpB,qBAAqB,EACrB,sBAAsB,EACtB,uBAAuB,EACvB,6BAA6B,EAC7B,8BAA8B,EAM/B,MAAM,0BAA0B,CAAC;AAElC,OAAO,EAUL,KAAK,UAAU,EACf,KAAK,cAAc,EACnB,KAAK,iBAAiB,EAEvB,MAAM,sBAAsB,CAAC;AAI9B,OAAO,EAAE,YAAY,EAAyB,MAAM,oBAAoB,CAAC;AASzE,OAAO,EAAwB,KAAK,eAAe,EAAE,MAAM,+BAA+B,CAAC;AAa3F;;;GAGG;AACH,MAAM,MAAM,aAAa,GAAG,SAAS,GAAG,cAAc,GAAG,oBAAoB,CAAC;AAgG9E;;GAEG;AACH,MAAM,WAAW,kBAAkB;IACjC,+CAA+C;IAC/C,GAAG,CAAC,EAAE;QACJ,UAAU,EAAE,CACV,QAAQ,EAAE,UAAU,EAAE,EACtB,MAAM,CAAC,EAAE,WAAW,EACpB,OAAO,CAAC,EAAE,iBAAiB,KACxB,aAAa,CAAC,cAAc,CAAC,CAAC;KACpC,CAAC;IACF;;;OAGG;IACH,QAAQ,CAAC,EAAE,MAAM,CAAC;IAClB;;;;;OAKG;IACH,YAAY,CAAC,EAAE,YAAY,GAAG,IAAI,CAAC;IACnC;;;;;OAKG;IACH,YAAY,CAAC,EAAE,eAAe,GAAG,IAAI,CAAC;CACvC;AAED;;;;;;GAMG;AACH,eAAO,MAAM,iBAAiB,KAAK,CAAC;AAsBpC,qBAAa,WAAY,YAAW,KAAK;IAUrC,OAAO,CAAC,UAAU;IATpB,OAAO,CAAC,QAAQ,CAAwC;IACxD,OAAO,CAAC,IAAI,CAAgD;IAC5D,OAAO,CAAC,QAAQ,CAAS;IACzB,OAAO,CAAC,kBAAkB,CAAmC;IAC7D,OAAO,CAAC,YAAY,CAAsB;IAC1C,OAAO,CAAC,aAAa,CAAyB;IAC9C,OAAO,CAAC,oBAAoB,CAAU;gBAG5B,UAAU,EAAE,mBAAmB,EACvC,OAAO,GAAE,kBAAuB;IAelC,OAAO,KAAK,GAAG,GAKd;IAED,OAAO,KAAK,YAAY,GAQvB;IAMK,UAAU,CAAC,MAAM,EAAE,iBAAiB,GAAG,OAAO,CAAC,kBAAkB,CAAC;IAsElE,YAAY,CAChB,MAAM,EAAE,mBAAmB,GAC1B,OAAO,CAAC,oBAAoB,CAAC;IAS1B,UAAU,CAAC,MAAM,EAAE,iBAAiB,GAAG,OAAO,CAAC,kBAAkB,CAAC;IA+CxE;;;;;;;;;;;;;;OAcG;IACH,OAAO,CAAC,yBAAyB;IAc3B,wBAAwB,CAC5B,MAAM,EAAE,sBAAsB,GAC7B,OAAO,CAAC,uBAAuB,CAAC;IASnC;;;;;;;OAOG;YACW,iBAAiB;IA0C/B;;;;;;;;;;;;;;;;OAgBG;IACH,OAAO,CAAC,kBAAkB;IA+CpB,sBAAsB,CAC1B,MAAM,EAAE,6BAA6B,GACpC,OAAO,CAAC,8BAA8B,CAAC;IAsE1C;;;;;;;;;OASG;IACH,OAAO,CAAC,WAAW;IAOnB,kFAAkF;IAClF,OAAO,CAAC,UAAU;IAUZ,cAAc,CAClB,MAAM,EAAE,qBAAqB,GAC5B,OAAO,CAAC,sBAAsB,CAAC;IA0C5B,MAAM,CAAC,MAAM,EAAE,aAAa,GAAG,OAAO,CAAC,cAAc,CAAC;IAuJtD,MAAM,CAAC,MAAM,EAAE,kBAAkB,GAAG,OAAO,CAAC,IAAI,CAAC;IAKjD,YAAY,CAAC,MAAM,EAAE,mBAAmB,GAAG,OAAO,CAAC,IAAI,CAAC;IAaxD,YAAY,CAAC,MAAM,EAAE,mBAAmB,GAAG,OAAO,CAAC,oBAAoB,CAAC;IAqDxE,WAAW,CAAC,MAAM,EAAE,kBAAkB,GAAG,OAAO,CAAC,mBAAmB,CAAC;IA+CrE,oBAAoB,CACxB,MAAM,EAAE,kBAAkB,GACzB,OAAO,CAAC,mBAAmB,CAAC;IAiDzB,aAAa,CACjB,MAAM,EAAE,oBAAoB,GAC3B,OAAO,CAAC,qBAAqB,CAAC;IA0CjC,OAAO,CAAC,QAAQ;IAiBhB,OAAO,CAAC,cAAc;IAYtB,OAAO,CAAC,gBAAgB;YAWV,cAAc;IAkD5B;;;;OAIG;YACW,aAAa;IAwK3B,OAAO,CAAC,aAAa;IAiBrB,iFAAiF;IACjF,OAAO,CAAC,wBAAwB;CAWjC"}
|
package/dist/protocol/agent.js
CHANGED
|
@@ -8,6 +8,7 @@ import { TOOL_DEFINITIONS } from "../tools/definitions.js";
|
|
|
8
8
|
import { connectSessionMcpServers } from "../tools/session-mcp-client.js";
|
|
9
9
|
import { SessionStore } from "./session-store.js";
|
|
10
10
|
import { buildSystemPrompt } from "./system-prompt.js";
|
|
11
|
+
import { discoverSlashCommands, parseSlashCommand, renderSlashCommand, } from "./slash-commands.js";
|
|
11
12
|
import { preprocessImageBlocks, buildPromptBlockDiagnosticLines } from "./image-preprocessor.js";
|
|
12
13
|
import { StdioVisionMcpClient } from "../tools/vision-mcp-client.js";
|
|
13
14
|
import { resolveApiKey } from "../llm/credentials.js";
|
|
@@ -235,7 +236,10 @@ export class GlmAcpAgent {
|
|
|
235
236
|
mcpTools,
|
|
236
237
|
mode: "default",
|
|
237
238
|
thoughtLevel,
|
|
239
|
+
commands: discoverSlashCommands(params.cwd),
|
|
240
|
+
displayText: new WeakMap(),
|
|
238
241
|
});
|
|
242
|
+
this.scheduleAvailableCommands(sessionId);
|
|
239
243
|
return {
|
|
240
244
|
sessionId,
|
|
241
245
|
models: this.modelsState(model),
|
|
@@ -243,6 +247,35 @@ export class GlmAcpAgent {
|
|
|
243
247
|
configOptions: this.configOptionsState(model, thoughtLevel, "default"),
|
|
244
248
|
};
|
|
245
249
|
}
|
|
250
|
+
/**
|
|
251
|
+
* Queue an `available_commands_update` snapshot for a session.
|
|
252
|
+
*
|
|
253
|
+
* The notification is the only channel ACP gives us for slash-command
|
|
254
|
+
* autocomplete — it is not part of any method's response — so every session
|
|
255
|
+
* entry point (create / load / fork / resume) has to send one.
|
|
256
|
+
*
|
|
257
|
+
* It is deliberately *deferred* rather than awaited inline: a client learns a
|
|
258
|
+
* session's id from the `session/new` / `session/fork` response, so a
|
|
259
|
+
* notification written ahead of that response arrives for a session the client
|
|
260
|
+
* has never heard of, and clients drop those. Sending on the next macrotask
|
|
261
|
+
* puts it behind the response the caller is about to return (and, on load,
|
|
262
|
+
* behind the replayed transcript), which is also when a client is ready to
|
|
263
|
+
* paint the menu. Each send replaces the previous list wholesale.
|
|
264
|
+
*/
|
|
265
|
+
scheduleAvailableCommands(sessionId) {
|
|
266
|
+
setTimeout(() => {
|
|
267
|
+
const session = this.sessions.get(sessionId);
|
|
268
|
+
if (!session)
|
|
269
|
+
return;
|
|
270
|
+
void safeSessionUpdate(this.connection, {
|
|
271
|
+
sessionId,
|
|
272
|
+
update: {
|
|
273
|
+
sessionUpdate: "available_commands_update",
|
|
274
|
+
availableCommands: availableCommandsState(session.commands),
|
|
275
|
+
},
|
|
276
|
+
});
|
|
277
|
+
}, 0);
|
|
278
|
+
}
|
|
246
279
|
async unstable_setSessionModel(params) {
|
|
247
280
|
const session = this.sessions.get(params.sessionId);
|
|
248
281
|
if (!session) {
|
|
@@ -500,14 +533,28 @@ export class GlmAcpAgent {
|
|
|
500
533
|
debug(line);
|
|
501
534
|
}
|
|
502
535
|
}
|
|
536
|
+
// Clients invoke an advertised command by sending `/name …` as ordinary
|
|
537
|
+
// prompt text, so expand it here into the instructions its definition
|
|
538
|
+
// holds. Unknown `/foo` is left alone and reaches the model as prose.
|
|
539
|
+
const promptBlocks = expandPromptCommand(params.prompt, session.commands);
|
|
503
540
|
const visionNative = isVisionNativeModel(session.model);
|
|
504
541
|
const preprocessed = visionNative
|
|
505
|
-
? { blocks:
|
|
506
|
-
: await preprocessImageBlocks(
|
|
507
|
-
const
|
|
542
|
+
? { blocks: promptBlocks, cleanups: [] }
|
|
543
|
+
: await preprocessImageBlocks(promptBlocks, this.visionClient, abortController.signal);
|
|
544
|
+
const userContent = visionNative
|
|
508
545
|
? renderVisionNativePromptBlocks(preprocessed.blocks)
|
|
509
|
-
: renderPromptBlocks(preprocessed.blocks);
|
|
510
|
-
|
|
546
|
+
: renderPromptBlocks(preprocessed.blocks).content;
|
|
547
|
+
const userMessage = { role: "user", content: userContent };
|
|
548
|
+
session.messages.push(userMessage);
|
|
549
|
+
// What the user typed, rendered from the blocks as they arrived — before
|
|
550
|
+
// command expansion and image analysis rewrote them for the model. Kept
|
|
551
|
+
// separately so `session/load` replays the conversation the user had (and
|
|
552
|
+
// the title names it), not the one the model saw. Persisted only when the
|
|
553
|
+
// two actually diverge.
|
|
554
|
+
const displayText = renderPromptBlocks(params.prompt).plainText;
|
|
555
|
+
if (displayText !== stringifyUserMessage(userContent)) {
|
|
556
|
+
session.displayText.set(userMessage, displayText);
|
|
557
|
+
}
|
|
511
558
|
// Echo back the client-supplied messageId on every response (success,
|
|
512
559
|
// cancelled, or error) so the client can correlate the turn.
|
|
513
560
|
const userMessageId = params.messageId ?? undefined;
|
|
@@ -523,7 +570,10 @@ export class GlmAcpAgent {
|
|
|
523
570
|
// updated timestamp so clients can show fresh metadata.
|
|
524
571
|
const titleUpdate = session.title === null
|
|
525
572
|
? (() => {
|
|
526
|
-
const derived =
|
|
573
|
+
const derived = displayText
|
|
574
|
+
.slice(0, 80)
|
|
575
|
+
.replace(/\s+/g, " ")
|
|
576
|
+
.trim();
|
|
527
577
|
session.title = derived.length > 0 ? derived : "New conversation";
|
|
528
578
|
return { title: session.title };
|
|
529
579
|
})()
|
|
@@ -659,12 +709,15 @@ export class GlmAcpAgent {
|
|
|
659
709
|
mcpTools,
|
|
660
710
|
mode: persisted.mode,
|
|
661
711
|
thoughtLevel: resolveThoughtLevel(persisted.model, persisted.thoughtLevel ?? "max"),
|
|
712
|
+
commands: discoverSlashCommands(params.cwd),
|
|
713
|
+
displayText: deserializeDisplayText(persisted.messages, persisted.displayText),
|
|
662
714
|
};
|
|
663
715
|
this.sessions.set(params.sessionId, restored);
|
|
664
716
|
// Replay user/assistant text turns so the client can rehydrate its UI.
|
|
665
717
|
// Tool / system messages are skipped — they're internal and the client
|
|
666
718
|
// doesn't render them on its own.
|
|
667
|
-
await this.replayMessages(params.sessionId, persisted.messages);
|
|
719
|
+
await this.replayMessages(params.sessionId, persisted.messages, restored.displayText);
|
|
720
|
+
this.scheduleAvailableCommands(params.sessionId);
|
|
668
721
|
return {
|
|
669
722
|
models: this.modelsState(persisted.model),
|
|
670
723
|
modes: this.modesState(persisted.mode),
|
|
@@ -680,10 +733,11 @@ export class GlmAcpAgent {
|
|
|
680
733
|
const toolDefinitions = this.availableToolDefinitions(mcpTools);
|
|
681
734
|
const newSessionId = randomUUID();
|
|
682
735
|
const forkedTitle = persisted.title === null ? null : `${persisted.title} (fork)`;
|
|
736
|
+
const forkedMessages = structuredClone(persisted.messages);
|
|
683
737
|
const forked = {
|
|
684
738
|
cwd: params.cwd,
|
|
685
739
|
// Deep-clone messages so the fork doesn't share state with the parent.
|
|
686
|
-
messages:
|
|
740
|
+
messages: forkedMessages,
|
|
687
741
|
abortController: null,
|
|
688
742
|
promptPromise: null,
|
|
689
743
|
title: forkedTitle,
|
|
@@ -693,9 +747,16 @@ export class GlmAcpAgent {
|
|
|
693
747
|
mcpTools,
|
|
694
748
|
mode: persisted.mode,
|
|
695
749
|
thoughtLevel: resolveThoughtLevel(persisted.model, persisted.thoughtLevel ?? "max"),
|
|
750
|
+
commands: discoverSlashCommands(params.cwd),
|
|
751
|
+
// Re-key onto the cloned messages: the parent's map is keyed by the
|
|
752
|
+
// originals, which the fork no longer holds.
|
|
753
|
+
displayText: deserializeDisplayText(forkedMessages, persisted.displayText),
|
|
696
754
|
};
|
|
697
755
|
this.sessions.set(newSessionId, forked);
|
|
698
756
|
this.persistSession(newSessionId, forked);
|
|
757
|
+
// Notify on the *created* session id — the parent thread's command list is
|
|
758
|
+
// unchanged and a notify there would repaint the wrong menu.
|
|
759
|
+
this.scheduleAvailableCommands(newSessionId);
|
|
699
760
|
return {
|
|
700
761
|
sessionId: newSessionId,
|
|
701
762
|
models: this.modelsState(forked.model),
|
|
@@ -720,8 +781,11 @@ export class GlmAcpAgent {
|
|
|
720
781
|
mcpTools,
|
|
721
782
|
mode: persisted.mode,
|
|
722
783
|
thoughtLevel: resolveThoughtLevel(persisted.model, persisted.thoughtLevel ?? "max"),
|
|
784
|
+
commands: discoverSlashCommands(params.cwd),
|
|
785
|
+
displayText: deserializeDisplayText(persisted.messages, persisted.displayText),
|
|
723
786
|
};
|
|
724
787
|
this.sessions.set(params.sessionId, restored);
|
|
788
|
+
this.scheduleAvailableCommands(params.sessionId);
|
|
725
789
|
// Resume does NOT replay history — the client keeps its own UI state and
|
|
726
790
|
// just wants the agent to pick up where it left off.
|
|
727
791
|
return {
|
|
@@ -734,6 +798,7 @@ export class GlmAcpAgent {
|
|
|
734
798
|
// Persistence helpers
|
|
735
799
|
// ---------------------------------------------------------------------------
|
|
736
800
|
snapshot(sessionId, session) {
|
|
801
|
+
const displayText = serializeDisplayText(session.messages, session.displayText);
|
|
737
802
|
return {
|
|
738
803
|
sessionId,
|
|
739
804
|
cwd: session.cwd,
|
|
@@ -743,6 +808,9 @@ export class GlmAcpAgent {
|
|
|
743
808
|
model: session.model,
|
|
744
809
|
mode: session.mode,
|
|
745
810
|
thoughtLevel: session.thoughtLevel,
|
|
811
|
+
// Absent for the common case where nothing diverges, so an ordinary
|
|
812
|
+
// session gains no on-disk weight.
|
|
813
|
+
...(displayText ? { displayText } : {}),
|
|
746
814
|
};
|
|
747
815
|
}
|
|
748
816
|
persistSession(sessionId, session) {
|
|
@@ -766,10 +834,10 @@ export class GlmAcpAgent {
|
|
|
766
834
|
}
|
|
767
835
|
return persisted;
|
|
768
836
|
}
|
|
769
|
-
async replayMessages(sessionId, messages) {
|
|
837
|
+
async replayMessages(sessionId, messages, displayText) {
|
|
770
838
|
for (const msg of messages) {
|
|
771
839
|
if (msg.role === "user") {
|
|
772
|
-
const text = stringifyUserMessage(msg.content);
|
|
840
|
+
const text = displayText.get(msg) ?? stringifyUserMessage(msg.content);
|
|
773
841
|
if (text.length === 0)
|
|
774
842
|
continue;
|
|
775
843
|
await this.connection.sessionUpdate({
|
|
@@ -884,7 +952,9 @@ export class GlmAcpAgent {
|
|
|
884
952
|
if (isOverflow && overflowRetryCount < 1) {
|
|
885
953
|
debug(`promptLoop: context overflow (1261) detected, performing emergency compaction`);
|
|
886
954
|
const window = getContextWindow(session.model);
|
|
887
|
-
session.messages = compactMessages(session.messages, Math.floor(window * 0.7)
|
|
955
|
+
session.messages = compactMessages(session.messages, Math.floor(window * 0.7), {
|
|
956
|
+
force: true,
|
|
957
|
+
});
|
|
888
958
|
overflowRetryCount++;
|
|
889
959
|
retryTurn = true;
|
|
890
960
|
}
|
|
@@ -1005,7 +1075,8 @@ function renderPromptBlocks(blocks) {
|
|
|
1005
1075
|
break;
|
|
1006
1076
|
}
|
|
1007
1077
|
case "image": {
|
|
1008
|
-
//
|
|
1078
|
+
// Only the display-text pass sees these; the model-facing pass has had
|
|
1079
|
+
// its images preprocessed into `<image_analysis>` annotations upstream.
|
|
1009
1080
|
textParts.push(`[image: ${block.mimeType}]`);
|
|
1010
1081
|
break;
|
|
1011
1082
|
}
|
|
@@ -1021,13 +1092,11 @@ function renderPromptBlocks(blocks) {
|
|
|
1021
1092
|
}
|
|
1022
1093
|
function renderVisionNativePromptBlocks(blocks) {
|
|
1023
1094
|
const content = [];
|
|
1024
|
-
const plainTextParts = [];
|
|
1025
1095
|
let imageIndex = 0;
|
|
1026
1096
|
const pushText = (text) => {
|
|
1027
1097
|
if (text.length === 0)
|
|
1028
1098
|
return;
|
|
1029
1099
|
content.push({ type: "text", text });
|
|
1030
|
-
plainTextParts.push(text);
|
|
1031
1100
|
};
|
|
1032
1101
|
for (const block of blocks) {
|
|
1033
1102
|
switch (block.type) {
|
|
@@ -1052,7 +1121,6 @@ function renderVisionNativePromptBlocks(blocks) {
|
|
|
1052
1121
|
const imageUrl = toVisionNativeImageUrl(block);
|
|
1053
1122
|
if (imageUrl) {
|
|
1054
1123
|
content.push({ type: "image_url", image_url: { url: imageUrl } });
|
|
1055
|
-
plainTextParts.push(`[image: ${block.mimeType}]`);
|
|
1056
1124
|
}
|
|
1057
1125
|
else {
|
|
1058
1126
|
pushText(`<image_unsupported_format index="${imageIndex}" mime="${escapeAttribute(block.mimeType)}">Only image/jpeg, image/jpg, image/png inputs can be sent to the native vision model. Attach a supported HTTPS image URL or supported base64 image data.</image_unsupported_format>`);
|
|
@@ -1066,7 +1134,7 @@ function renderVisionNativePromptBlocks(blocks) {
|
|
|
1066
1134
|
pushText(`[unknown block type ${block.type}]`);
|
|
1067
1135
|
}
|
|
1068
1136
|
}
|
|
1069
|
-
return
|
|
1137
|
+
return content;
|
|
1070
1138
|
}
|
|
1071
1139
|
function toVisionNativeImageUrl(block) {
|
|
1072
1140
|
if (!isSupportedVisionMime(block.mimeType))
|
|
@@ -1105,6 +1173,36 @@ function escapeAttribute(value) {
|
|
|
1105
1173
|
.replaceAll("<", "<")
|
|
1106
1174
|
.replaceAll(">", ">");
|
|
1107
1175
|
}
|
|
1176
|
+
/**
|
|
1177
|
+
* Project the identity-keyed display map onto indices into `messages`, ready
|
|
1178
|
+
* to persist. Returns `undefined` when nothing diverges so the field can be
|
|
1179
|
+
* left off the record entirely.
|
|
1180
|
+
*/
|
|
1181
|
+
function serializeDisplayText(messages, displayText) {
|
|
1182
|
+
let out;
|
|
1183
|
+
for (const [index, message] of messages.entries()) {
|
|
1184
|
+
const text = displayText.get(message);
|
|
1185
|
+
if (text === undefined)
|
|
1186
|
+
continue;
|
|
1187
|
+
out ??= {};
|
|
1188
|
+
out[String(index)] = text;
|
|
1189
|
+
}
|
|
1190
|
+
return out;
|
|
1191
|
+
}
|
|
1192
|
+
/** Rebuild the identity-keyed display map from a persisted index map. */
|
|
1193
|
+
function deserializeDisplayText(messages, persisted) {
|
|
1194
|
+
const out = new WeakMap();
|
|
1195
|
+
if (!persisted)
|
|
1196
|
+
return out;
|
|
1197
|
+
for (const [key, text] of Object.entries(persisted)) {
|
|
1198
|
+
const message = messages[Number(key)];
|
|
1199
|
+
// A hand-edited or truncated record can point past the end of the array;
|
|
1200
|
+
// dropping the entry just falls back to replaying the stored content.
|
|
1201
|
+
if (message)
|
|
1202
|
+
out.set(message, text);
|
|
1203
|
+
}
|
|
1204
|
+
return out;
|
|
1205
|
+
}
|
|
1108
1206
|
/** Flatten the `content` of a user message into a plain string for replay. */
|
|
1109
1207
|
function stringifyUserMessage(content) {
|
|
1110
1208
|
if (typeof content === "string")
|
|
@@ -1142,6 +1240,45 @@ function availableModelsWith(currentModelId) {
|
|
|
1142
1240
|
? available
|
|
1143
1241
|
: [...available, { modelId: currentModelId, name: currentModelId }];
|
|
1144
1242
|
}
|
|
1243
|
+
/**
|
|
1244
|
+
* Map discovered commands onto the ACP wire shape. Names are sent *without* a
|
|
1245
|
+
* leading slash — the client prepends it for display and sends `/name …` back
|
|
1246
|
+
* as prompt text. `input` is omitted for commands that declare no
|
|
1247
|
+
* `argument-hint`, which is how a client learns the command takes no extra text.
|
|
1248
|
+
*/
|
|
1249
|
+
function availableCommandsState(commands) {
|
|
1250
|
+
return commands.map((command) => ({
|
|
1251
|
+
name: command.name,
|
|
1252
|
+
description: command.description,
|
|
1253
|
+
...(command.argumentHint !== undefined
|
|
1254
|
+
? { input: { hint: command.argumentHint } }
|
|
1255
|
+
: {}),
|
|
1256
|
+
}));
|
|
1257
|
+
}
|
|
1258
|
+
/**
|
|
1259
|
+
* Expand a leading `/name` in the prompt's first written text into the
|
|
1260
|
+
* command's body.
|
|
1261
|
+
*
|
|
1262
|
+
* The target is the first text block that actually has content, not block 0:
|
|
1263
|
+
* a client may place an image or resource block ahead of what the user typed,
|
|
1264
|
+
* and the command would otherwise reach the model as a literal `/name`.
|
|
1265
|
+
*
|
|
1266
|
+
* Returns the blocks unchanged when that text does not start with an
|
|
1267
|
+
* advertised command. The caller keeps the original blocks around to derive
|
|
1268
|
+
* the replay/title text, so nothing here needs to report the typed form back.
|
|
1269
|
+
*/
|
|
1270
|
+
function expandPromptCommand(blocks, commands) {
|
|
1271
|
+
const index = blocks.findIndex((block) => block.type === "text" && block.text.trim().length > 0);
|
|
1272
|
+
const typed = blocks[index];
|
|
1273
|
+
if (typed?.type !== "text")
|
|
1274
|
+
return blocks;
|
|
1275
|
+
const parsed = parseSlashCommand(typed.text, commands);
|
|
1276
|
+
if (!parsed)
|
|
1277
|
+
return blocks;
|
|
1278
|
+
const expanded = [...blocks];
|
|
1279
|
+
expanded[index] = { ...typed, text: renderSlashCommand(parsed) };
|
|
1280
|
+
return expanded;
|
|
1281
|
+
}
|
|
1145
1282
|
/**
|
|
1146
1283
|
* Read an `AGENTS.md` (preferred) or `CLAUDE.md` from the session's cwd, returning
|
|
1147
1284
|
* its contents capped to {@link PROJECT_CONTEXT_CAP_CHARS} characters. Read errors
|
|
@@ -1176,13 +1313,26 @@ async function safeSessionUpdate(connection, params) {
|
|
|
1176
1313
|
// best-effort
|
|
1177
1314
|
}
|
|
1178
1315
|
}
|
|
1316
|
+
/**
|
|
1317
|
+
* Token cost charged to a single `image_url` content part.
|
|
1318
|
+
*
|
|
1319
|
+
* Vision-native GLM models tile an image into 14px patches and merge them 2x2,
|
|
1320
|
+
* so a full-size input (capped at 1120x1120) costs (1120/14)^2 / 4 = 1600
|
|
1321
|
+
* tokens. The wire form — an HTTPS URL or a base64 data URL — says nothing
|
|
1322
|
+
* about the decoded dimensions, so charge every image that ceiling:
|
|
1323
|
+
* over-counting only makes compaction fire early, while under-counting is what
|
|
1324
|
+
* lets an image-heavy session sail past the window and get rejected.
|
|
1325
|
+
*/
|
|
1326
|
+
const IMAGE_PART_TOKENS = 1600;
|
|
1179
1327
|
/**
|
|
1180
1328
|
* Heuristically estimate the number of tokens in a list of messages.
|
|
1181
1329
|
* Uses a simple 4-character-per-token rule, which is a safe baseline for
|
|
1182
|
-
* English and code
|
|
1330
|
+
* English and code, plus a flat per-image charge for native `image_url` parts
|
|
1331
|
+
* (see {@link IMAGE_PART_TOKENS}).
|
|
1183
1332
|
*/
|
|
1184
1333
|
function estimateTokens(messages) {
|
|
1185
1334
|
let chars = 0;
|
|
1335
|
+
let tokens = 0;
|
|
1186
1336
|
for (const m of messages) {
|
|
1187
1337
|
if (typeof m.content === "string") {
|
|
1188
1338
|
chars += m.content.length;
|
|
@@ -1192,6 +1342,9 @@ function estimateTokens(messages) {
|
|
|
1192
1342
|
if ("text" in part && typeof part.text === "string") {
|
|
1193
1343
|
chars += part.text.length;
|
|
1194
1344
|
}
|
|
1345
|
+
else if (part.type === "image_url") {
|
|
1346
|
+
tokens += IMAGE_PART_TOKENS;
|
|
1347
|
+
}
|
|
1195
1348
|
}
|
|
1196
1349
|
}
|
|
1197
1350
|
if (m.role === "assistant" && m.tool_calls) {
|
|
@@ -1203,7 +1356,7 @@ function estimateTokens(messages) {
|
|
|
1203
1356
|
}
|
|
1204
1357
|
}
|
|
1205
1358
|
}
|
|
1206
|
-
return Math.ceil(chars / 4);
|
|
1359
|
+
return tokens + Math.ceil(chars / 4);
|
|
1207
1360
|
}
|
|
1208
1361
|
/**
|
|
1209
1362
|
* Prune message history to stay within a target token limit.
|
|
@@ -1215,8 +1368,15 @@ function estimateTokens(messages) {
|
|
|
1215
1368
|
* with a user message followed by assistant and tool responses.
|
|
1216
1369
|
* 3. Evict the largest remaining interaction groups until the total estimate
|
|
1217
1370
|
* is below `targetTokens`.
|
|
1371
|
+
*
|
|
1372
|
+
* `force` is for the emergency path after the provider itself rejected the
|
|
1373
|
+
* history. There {@link estimateTokens} has been proven wrong, so it can set
|
|
1374
|
+
* neither the stopping point nor what is off limits: the target is halved, and
|
|
1375
|
+
* the preserved tail becomes evictable from its oldest end. Only the final turn
|
|
1376
|
+
* is sacred — it carries the live user message, and a request without it is not
|
|
1377
|
+
* a retry of anything.
|
|
1218
1378
|
*/
|
|
1219
|
-
function compactMessages(messages, targetTokens, preserveTurns = 10) {
|
|
1379
|
+
function compactMessages(messages, targetTokens, { preserveTurns = 10, force = false } = {}) {
|
|
1220
1380
|
if (messages.length <= 1)
|
|
1221
1381
|
return messages;
|
|
1222
1382
|
const systemPrompt = messages[0];
|
|
@@ -1234,40 +1394,56 @@ function compactMessages(messages, targetTokens, preserveTurns = 10) {
|
|
|
1234
1394
|
if (currentTurn.length > 0) {
|
|
1235
1395
|
turns.push(currentTurn);
|
|
1236
1396
|
}
|
|
1237
|
-
//
|
|
1238
|
-
|
|
1397
|
+
// The final turn holds the live user message, so it is never evictable —
|
|
1398
|
+
// with nothing else to drop there is no compaction to do.
|
|
1399
|
+
if (turns.length < 2)
|
|
1400
|
+
return messages;
|
|
1401
|
+
if (!force && turns.length <= preserveTurns)
|
|
1239
1402
|
return messages;
|
|
1240
1403
|
let currentEstimate = estimateTokens(messages);
|
|
1241
|
-
if (currentEstimate <= targetTokens)
|
|
1404
|
+
if (!force && currentEstimate <= targetTokens)
|
|
1242
1405
|
return messages;
|
|
1243
|
-
|
|
1244
|
-
//
|
|
1245
|
-
|
|
1246
|
-
|
|
1247
|
-
|
|
1248
|
-
|
|
1249
|
-
|
|
1250
|
-
|
|
1251
|
-
|
|
1252
|
-
})
|
|
1253
|
-
|
|
1254
|
-
|
|
1406
|
+
// An estimate already above target names a real reduction to aim for, forced
|
|
1407
|
+
// or not. One that sits *below* target while the provider is rejecting the
|
|
1408
|
+
// payload has been disproven, and stopping on it would spend the single retry
|
|
1409
|
+
// on another oversized request — so halve it instead. Wrong by an unknown
|
|
1410
|
+
// factor still shrinks geometrically, and only that case pays the extra loss.
|
|
1411
|
+
const estimateDisproven = force && currentEstimate <= targetTokens;
|
|
1412
|
+
const effectiveTarget = estimateDisproven
|
|
1413
|
+
? Math.floor(currentEstimate / 2)
|
|
1414
|
+
: targetTokens;
|
|
1415
|
+
debug(`compactMessages: currentEstimate=${currentEstimate} target=${effectiveTarget} turns=${turns.length} force=${force}`);
|
|
1416
|
+
const sized = turns.map((turn, index) => ({ index, tokens: estimateTokens(turn) }));
|
|
1417
|
+
const protectedFrom = turns.length - preserveTurns;
|
|
1255
1418
|
const evictedIndices = new Set();
|
|
1419
|
+
// Largest first, among the turns outside the preserved tail.
|
|
1420
|
+
const candidateTurns = sized
|
|
1421
|
+
.filter((c) => c.index < protectedFrom)
|
|
1422
|
+
.sort((a, b) => b.tokens - a.tokens);
|
|
1256
1423
|
for (const c of candidateTurns) {
|
|
1257
|
-
if (currentEstimate <=
|
|
1424
|
+
if (currentEstimate <= effectiveTarget)
|
|
1258
1425
|
break;
|
|
1259
1426
|
evictedIndices.add(c.index);
|
|
1260
1427
|
currentEstimate -= c.tokens;
|
|
1261
1428
|
}
|
|
1429
|
+
// Still over, and forced? Then the preserved tail is itself the problem — a
|
|
1430
|
+
// run of image-heavy prompts can exceed the window on its own, leaving the
|
|
1431
|
+
// candidates above unable to reach the target however many are dropped. Eat
|
|
1432
|
+
// into the tail from its oldest end so the freshest context survives.
|
|
1433
|
+
if (force) {
|
|
1434
|
+
for (let i = Math.max(protectedFrom, 0); i < turns.length - 1; i++) {
|
|
1435
|
+
if (currentEstimate <= effectiveTarget)
|
|
1436
|
+
break;
|
|
1437
|
+
evictedIndices.add(i);
|
|
1438
|
+
currentEstimate -= sized[i].tokens;
|
|
1439
|
+
}
|
|
1440
|
+
}
|
|
1262
1441
|
const compacted = [systemPrompt];
|
|
1263
|
-
for (let i = 0; i < turns.length
|
|
1442
|
+
for (let i = 0; i < turns.length; i++) {
|
|
1264
1443
|
if (!evictedIndices.has(i)) {
|
|
1265
1444
|
compacted.push(...turns[i]);
|
|
1266
1445
|
}
|
|
1267
1446
|
}
|
|
1268
|
-
for (const turn of tail) {
|
|
1269
|
-
compacted.push(...turn);
|
|
1270
|
-
}
|
|
1271
1447
|
debug(`compactMessages: done, newEstimate=${estimateTokens(compacted)} messageCount=${compacted.length}`);
|
|
1272
1448
|
return compacted;
|
|
1273
1449
|
}
|