pi-makora-provider 1.1.0 → 1.3.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +15 -13
- package/index.ts +16 -29
- package/models.json +13 -13
- package/package.json +11 -11
- package/patch.json +28 -66
package/README.md
CHANGED
|
@@ -4,7 +4,7 @@
|
|
|
4
4
|
|
|
5
5
|
**Open-weight models through [Makora](https://inference.makora.com)**
|
|
6
6
|
|
|
7
|
-
_DeepSeek V4, Kimi K2.
|
|
7
|
+
_DeepSeek V4, Kimi K2.7 Code, GLM 5.2, Qwen 3.6 — with client-side tool call repair and preserved-thinking flags for [pi](https://github.com/earendil-works/pi-coding-agent)._
|
|
8
8
|
|
|
9
9
|
[](https://github.com/earendil-works/pi-coding-agent)
|
|
10
10
|
[](./LICENSE)
|
|
@@ -19,14 +19,14 @@ _DeepSeek V4, Kimi K2.6, GLM 5.1, Qwen 3.6 — with client-side tool call repair
|
|
|
19
19
|
| Model | ID | Reasoning | Notes |
|
|
20
20
|
|-------|----|-----------|-------|
|
|
21
21
|
| DeepSeek V4 Flash | `deepseek-ai/DeepSeek-V4-Flash` | Yes | `include_reasoning` + `chat_template_kwargs.thinking` via `before_provider_request` payload rewrite; returns `reasoning` field |
|
|
22
|
-
| DeepSeek V4 Pro | `deepseek-ai/DeepSeek-V4-Pro` | Yes | `chat_template_kwargs.thinking` via `before_provider_request` payload rewrite; returns `
|
|
23
|
-
| GLM 5.2 FP8 | `zai-org/GLM-5.2-FP8` | Yes | `enable_thinking` via `qwen-chat-template`; thinking levels
|
|
24
|
-
|
|
|
22
|
+
| DeepSeek V4 Pro | `deepseek-ai/DeepSeek-V4-Pro` | Yes | `chat_template_kwargs.thinking` via `before_provider_request` payload rewrite; returns `reasoning` field |
|
|
23
|
+
| GLM 5.2 FP8 | `zai-org/GLM-5.2-FP8` | Yes | `enable_thinking` via `qwen-chat-template`; effort via `reasoning_effort` (only `high`/`max` distinct, per vLLM GLM-5.2 recipe); thinking levels aligned with neuralwatt GLM 5.2; returns `reasoning` field |
|
|
24
|
+
| GLM 5.2 NVFP4 | `zai-org/GLM-5.2-NVFP4` | Yes | `enable_thinking` via `qwen-chat-template`; effort via `reasoning_effort` (only `high`/`max` distinct, per vLLM GLM-5.2 recipe); returns `reasoning` field |
|
|
25
|
+
| Kimi K2.7 Code | `moonshotai/Kimi-K2.7-Code` | Yes | Reasoning on by default (thinking-only model); `preserve_thinking` via `chatTemplateKwargs` for multi-turn continuity; returns `reasoning` field; client-side tool call parsing (vLLM streaming parser bypass) |
|
|
25
26
|
| Llama 3.3 70B FP8 | `amd/Llama-3.3-70B-Instruct-FP8-KV` | No | |
|
|
26
27
|
| Llama 3.3 70B Instruct | `meta-llama/Llama-3.3-70B-Instruct` | No | |
|
|
27
|
-
|
|
|
28
|
-
| Qwen 3.6
|
|
29
|
-
| Qwen 3.6 35B A3B NVFP4 | `unsloth/Qwen3.6-35B-A3B-NVFP4` | Yes | `enable_thinking` via `qwen-chat-template`; client-side tool call parsing (vLLM streaming parser bypass) |
|
|
28
|
+
| Qwen 3.6 27B NVFP4 | `unsloth/Qwen3.6-27B-NVFP4` | Yes | `enable_thinking` via `qwen-chat-template`; `preserve_thinking` via `chatTemplateKwargs` for multi-turn continuity; returns `reasoning` field; client-side tool call parsing (vLLM streaming parser bypass) |
|
|
29
|
+
| Qwen 3.6 35B A3B NVFP4 | `unsloth/Qwen3.6-35B-A3B-NVFP4` | Yes | `enable_thinking` via `qwen-chat-template`; `preserve_thinking` via `chatTemplateKwargs` for multi-turn continuity; returns `reasoning` field; client-side tool call parsing (vLLM streaming parser bypass) |
|
|
30
30
|
<!-- MODELS_TABLE_END -->
|
|
31
31
|
|
|
32
32
|
## Installation
|
|
@@ -163,15 +163,17 @@ These issues are common to all vLLM-hosted providers and affect Makora models:
|
|
|
163
163
|
|
|
164
164
|
- **GLM tool calls — follow-up `tool_calls` crash workaround**: the ZAI chat template calls `.items()` on any assistant message that carries a `tool_calls` field (on the `tool_calls` list or on its JSON-string `arguments` field), raising a leaked Python `AttributeError` that vLLM returns as HTTP 400 (`'list object' has no attribute 'items'` / `'str object' has no attribute 'items'`). The `before_provider_request` hook (`stripGlmToolCalls`) strips `tool_calls` from assistant messages and converts them to GLM-native `<tool_call>` XML text in `content` — which the model natively understands in conversation history — before the follow-up request is sent. The hook is scoped to Makora's own GLM models via an exact model-id match (`isMakoraGlmVllmModel`): `before_provider_request` is global in pi (it fires for every loaded provider's requests and the event carries no `provider` field), so gating by model id is what keeps it from touching GLM models served by other providers. If upstream fixes the `.items()` crash, this transform is a harmless no-op (the XML text is valid GLM input either way).
|
|
165
165
|
|
|
166
|
-
- **Kimi K2.
|
|
166
|
+
- **Kimi K2.7 Code + Qwen 3.6 tool calling**: vLLM's server-side streaming tool-call handling is broken or missing for these models — Makora's vLLM is missing both `--enable-auto-tool-choice` and `--tool-call-parser` for them, so they emit tool calls as raw text tokens rather than structured `tool_call` deltas.
|
|
167
167
|
|
|
168
|
-
- **Kimi K2.
|
|
168
|
+
- **Kimi K2.7 Code**: Uses `<|tool_call_begin|>...<|tool_call_end|>` tokens.
|
|
169
169
|
- **Qwen 3.6**: Uses hermes-style `<function=...>` XML, sometimes with `█` delimiters.
|
|
170
170
|
|
|
171
|
-
- **GLM 5.
|
|
171
|
+
- **GLM 5.2 CoT leak**: On some vLLM builds, disabling reasoning may still leak chain-of-thought into `content` terminated by a ``` marker. See [vllm-project/vllm#31319](https://github.com/vllm-project/vllm/issues/31319).
|
|
172
172
|
|
|
173
173
|
- **DeepSeek V4 reasoning**: The official DeepSeek API uses `thinking: { type: "enabled" }` which Makora's vLLM silently ignores. The `before_provider_request` hook rewrites the payload to use vLLM-native params instead:
|
|
174
|
-
- **DS V4 Pro**: `chat_template_kwargs: { thinking: true }`. Returns `reasoning_content
|
|
174
|
+
- **DS V4 Pro**: `chat_template_kwargs: { thinking: true }`. Returns `reasoning` (vLLM reasoning parser; the official DeepSeek API returns `reasoning_content`, but Makora's vLLM build returns `reasoning`).
|
|
175
175
|
- **DS V4 Flash**: `include_reasoning: true` + `chat_template_kwargs: { thinking: true }`. `include_reasoning` alone returns `reasoning: null` on this vLLM build — both params are required. Returns `reasoning`.
|
|
176
|
-
- **
|
|
177
|
-
- **
|
|
176
|
+
- **Preserved thinking (multi-turn reasoning continuity)**: every reasoning model on Makora returns reasoning in the `reasoning` field (not `reasoning_content`). Pi sends prior assistant reasoning back on follow-up turns; Makora's vLLM chat templates accept it under either `reasoning` or `reasoning_content`. For models whose chat template defines an explicit preserved-thinking flag, the flag is sent via `compat.chatTemplateKwargs`:
|
|
177
|
+
- **Kimi K2.7 Code**: `preserve_thinking: true` (the model forces thinking-only mode with preserve_thinking on, per the vLLM recipe).
|
|
178
|
+
- **Qwen 3.6 27B / 35B**: `preserve_thinking: true` (Qwen 3.6 ships with preserve_thinking).
|
|
179
|
+
- **GLM 5.2**: no extra flag — the vLLM GLM-5.2 recipe uses only `enable_thinking` + `reasoning_effort`; preserved thinking works without a dedicated flag. Effort honors only `high` and `max` (lower pi levels resolve to the default `max`).
|
package/index.ts
CHANGED
|
@@ -17,7 +17,7 @@
|
|
|
17
17
|
* pi sends thinking: { type } via the "deepseek" thinkingFormat, but vLLM
|
|
18
18
|
* ignores that — the before_provider_request hook rewrites the payload to
|
|
19
19
|
* use chat_template_kwargs: { thinking: true } instead.
|
|
20
|
-
* Returns
|
|
20
|
+
* Returns reasoning field (vLLM reasoning parser; not reasoning_content).
|
|
21
21
|
* - DeepSeek V4 Flash: reasoning via include_reasoning +
|
|
22
22
|
* chat_template_kwargs.thinking on vLLM.
|
|
23
23
|
* The before_provider_request hook rewrites the payload to replace
|
|
@@ -25,26 +25,18 @@
|
|
|
25
25
|
* chat_template_kwargs: { thinking: true }.
|
|
26
26
|
* include_reasoning alone returns reasoning: null on this vLLM build.
|
|
27
27
|
* Returns reasoning field.
|
|
28
|
-
* - GLM 5.
|
|
29
|
-
*
|
|
30
|
-
*
|
|
31
|
-
*
|
|
32
|
-
*
|
|
33
|
-
*
|
|
34
|
-
*
|
|
35
|
-
*
|
|
36
|
-
*
|
|
37
|
-
* - GLM 5.2 FP8: reasoning via chat_template_kwargs.enable_thinking.
|
|
38
|
-
* Thinking levels (minimal/low/medium/high/max) are copied from the
|
|
39
|
-
* neuralwatt provider's GLM 5.2 configuration and mapped through pi's
|
|
40
|
-
* qwen-chat-template thinkingFormat.
|
|
41
|
-
* - GPT-OSS 120B: reasoning always on; returns `reasoning` field.
|
|
42
|
-
* - Kimi K2.6 NVFP4 / Kimi K2.7 Code: reasoning always on by default;
|
|
43
|
-
* returns `reasoning` field. Can be toggled via enable_thinking.
|
|
28
|
+
* - GLM 5.2 FP8 / NVFP4: reasoning via chat_template_kwargs.enable_thinking.
|
|
29
|
+
* Effort uses vLLM's reasoning_effort field (only `high` and `max` are
|
|
30
|
+
* distinct levels per the vLLM GLM-5.2 recipe; lower pi levels resolve to
|
|
31
|
+
* the default). Thinking levels aligned with the neuralwatt provider's
|
|
32
|
+
* GLM 5.2 configuration and mapped through pi's qwen-chat-template
|
|
33
|
+
* thinkingFormat. Returns `reasoning` field.
|
|
34
|
+
* - Kimi K2.7 Code: reasoning always on (thinking-only model);
|
|
35
|
+
* chatTemplateKwargs.preserve_thinking forces multi-turn reasoning
|
|
36
|
+
* continuity. Returns `reasoning` field. Can be toggled via enable_thinking.
|
|
44
37
|
* - Qwen 3.6 models: reasoning via chat_template_kwargs.enable_thinking;
|
|
45
|
-
*
|
|
46
|
-
*
|
|
47
|
-
* returns reasoning_content field.
|
|
38
|
+
* chatTemplateKwargs.preserve_thinking for multi-turn continuity.
|
|
39
|
+
* Returns `reasoning` field.
|
|
48
40
|
* - Llama 3.3 70B: not a reasoning model.
|
|
49
41
|
*
|
|
50
42
|
* Developer role is NOT supported by any of the chat templates on Makora's
|
|
@@ -109,6 +101,10 @@ interface JsonModel {
|
|
|
109
101
|
requiresToolResultName?: boolean;
|
|
110
102
|
requiresAssistantAfterToolResult?: boolean;
|
|
111
103
|
cacheControlFormat?: "anthropic";
|
|
104
|
+
/** Extra keys merged into vLLM `chat_template_kwargs` on every request.
|
|
105
|
+
* Used for preserved-thinking flags like `preserve_thinking` / `clear_thinking`
|
|
106
|
+
* that some chat templates require for multi-turn reasoning continuity. */
|
|
107
|
+
chatTemplateKwargs?: Record<string, unknown>;
|
|
112
108
|
};
|
|
113
109
|
}
|
|
114
110
|
|
|
@@ -218,10 +214,8 @@ const BASE_URL = "https://inference.makora.com/v1";
|
|
|
218
214
|
|
|
219
215
|
const DS_PRO_ID = "deepseek-ai/DeepSeek-V4-Pro";
|
|
220
216
|
const DS_FLASH_ID = "deepseek-ai/DeepSeek-V4-Flash";
|
|
221
|
-
const MINIMAX_M3_ID = "MiniMaxAI/MiniMax-M3-MXFP8";
|
|
222
217
|
|
|
223
218
|
const DS_VLLM_MODELS = new Set([DS_PRO_ID, DS_FLASH_ID]);
|
|
224
|
-
const ENABLE_THINKING_VLLM_MODELS = new Set([MINIMAX_M3_ID]);
|
|
225
219
|
|
|
226
220
|
/**
|
|
227
221
|
* Makora's GLM models, built from the same models list this extension
|
|
@@ -266,8 +260,6 @@ export function isMakoraGlmVllmModel(model: string): boolean {
|
|
|
266
260
|
* - DS V4 Flash: `include_reasoning: true` + `chat_template_kwargs: { thinking: true }`
|
|
267
261
|
* + `reasoning_effort`. `include_reasoning` alone returns `reasoning: null`
|
|
268
262
|
* on this vLLM build — both params are required.
|
|
269
|
-
* - MiniMax M3: `chat_template_kwargs: { enable_thinking: true }` +
|
|
270
|
-
* `reasoning_effort`. Returns `reasoning_content` field.
|
|
271
263
|
*
|
|
272
264
|
* This hook rewrites the payload accordingly.
|
|
273
265
|
*/
|
|
@@ -293,11 +285,6 @@ function rewriteVllmPayload(payload: Record<string, unknown>): Record<string, un
|
|
|
293
285
|
const ctq = (p.chat_template_kwargs as Record<string, unknown>) ?? {};
|
|
294
286
|
p.chat_template_kwargs = { ...ctq, thinking: true };
|
|
295
287
|
}
|
|
296
|
-
} else if (ENABLE_THINKING_VLLM_MODELS.has(model)) {
|
|
297
|
-
// Models using chat_template_kwargs.enable_thinking (e.g. MiniMax M3)
|
|
298
|
-
delete p.thinking;
|
|
299
|
-
const ctq = (p.chat_template_kwargs as Record<string, unknown>) ?? {};
|
|
300
|
-
p.chat_template_kwargs = { ...ctq, enable_thinking: true };
|
|
301
288
|
}
|
|
302
289
|
|
|
303
290
|
return p;
|
package/models.json
CHANGED
|
@@ -63,8 +63,8 @@
|
|
|
63
63
|
}
|
|
64
64
|
},
|
|
65
65
|
{
|
|
66
|
-
"id": "
|
|
67
|
-
"name": "
|
|
66
|
+
"id": "moonshotai/Kimi-K2.7-Code",
|
|
67
|
+
"name": "Kimi K2.7 Code",
|
|
68
68
|
"reasoning": false,
|
|
69
69
|
"input": [
|
|
70
70
|
"text"
|
|
@@ -75,7 +75,7 @@
|
|
|
75
75
|
"cacheRead": 0,
|
|
76
76
|
"cacheWrite": 0
|
|
77
77
|
},
|
|
78
|
-
"contextWindow":
|
|
78
|
+
"contextWindow": 262144,
|
|
79
79
|
"maxTokens": 0,
|
|
80
80
|
"compat": {
|
|
81
81
|
"supportsDeveloperRole": false,
|
|
@@ -84,8 +84,8 @@
|
|
|
84
84
|
}
|
|
85
85
|
},
|
|
86
86
|
{
|
|
87
|
-
"id": "
|
|
88
|
-
"name": "
|
|
87
|
+
"id": "unsloth/Qwen3.6-27B-NVFP4",
|
|
88
|
+
"name": "Qwen 3.6 27B NVFP4",
|
|
89
89
|
"reasoning": false,
|
|
90
90
|
"input": [
|
|
91
91
|
"text"
|
|
@@ -105,8 +105,8 @@
|
|
|
105
105
|
}
|
|
106
106
|
},
|
|
107
107
|
{
|
|
108
|
-
"id": "unsloth/Qwen3.6-
|
|
109
|
-
"name": "Qwen 3.6
|
|
108
|
+
"id": "unsloth/Qwen3.6-35B-A3B-NVFP4",
|
|
109
|
+
"name": "Qwen 3.6 35B A3B NVFP4",
|
|
110
110
|
"reasoning": false,
|
|
111
111
|
"input": [
|
|
112
112
|
"text"
|
|
@@ -126,8 +126,8 @@
|
|
|
126
126
|
}
|
|
127
127
|
},
|
|
128
128
|
{
|
|
129
|
-
"id": "
|
|
130
|
-
"name": "
|
|
129
|
+
"id": "zai-org/GLM-5.2-FP8",
|
|
130
|
+
"name": "GLM 5.2 FP8",
|
|
131
131
|
"reasoning": false,
|
|
132
132
|
"input": [
|
|
133
133
|
"text"
|
|
@@ -138,7 +138,7 @@
|
|
|
138
138
|
"cacheRead": 0,
|
|
139
139
|
"cacheWrite": 0
|
|
140
140
|
},
|
|
141
|
-
"contextWindow":
|
|
141
|
+
"contextWindow": 1048576,
|
|
142
142
|
"maxTokens": 0,
|
|
143
143
|
"compat": {
|
|
144
144
|
"supportsDeveloperRole": false,
|
|
@@ -147,8 +147,8 @@
|
|
|
147
147
|
}
|
|
148
148
|
},
|
|
149
149
|
{
|
|
150
|
-
"id": "zai-org/GLM-5.2-
|
|
151
|
-
"name": "GLM 5.2
|
|
150
|
+
"id": "zai-org/GLM-5.2-NVFP4",
|
|
151
|
+
"name": "GLM 5.2 NVFP4",
|
|
152
152
|
"reasoning": false,
|
|
153
153
|
"input": [
|
|
154
154
|
"text"
|
|
@@ -159,7 +159,7 @@
|
|
|
159
159
|
"cacheRead": 0,
|
|
160
160
|
"cacheWrite": 0
|
|
161
161
|
},
|
|
162
|
-
"contextWindow":
|
|
162
|
+
"contextWindow": 1048576,
|
|
163
163
|
"maxTokens": 0,
|
|
164
164
|
"compat": {
|
|
165
165
|
"supportsDeveloperRole": false,
|
package/package.json
CHANGED
|
@@ -1,17 +1,9 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "pi-makora-provider",
|
|
3
|
-
"version": "1.
|
|
4
|
-
"description": "Makora provider extension for pi - Access DeepSeek V4, GLM 5.
|
|
3
|
+
"version": "1.3.0",
|
|
4
|
+
"description": "Makora provider extension for pi - Access DeepSeek V4, GLM 5.2, Kimi K2.7 Code, Llama 3.3, Qwen 3.6, and more through the Makora inference API",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"main": "index.ts",
|
|
7
|
-
"scripts": {
|
|
8
|
-
"clean": "echo 'nothing to clean'",
|
|
9
|
-
"build": "echo 'nothing to build'",
|
|
10
|
-
"check": "echo 'nothing to check'",
|
|
11
|
-
"test": "vitest run",
|
|
12
|
-
"test:watch": "vitest",
|
|
13
|
-
"update-models": "node scripts/update-models.js"
|
|
14
|
-
},
|
|
15
7
|
"keywords": [
|
|
16
8
|
"pi",
|
|
17
9
|
"extension",
|
|
@@ -42,5 +34,13 @@
|
|
|
42
34
|
},
|
|
43
35
|
"devDependencies": {
|
|
44
36
|
"vitest": "^4.1.9"
|
|
37
|
+
},
|
|
38
|
+
"scripts": {
|
|
39
|
+
"clean": "echo 'nothing to clean'",
|
|
40
|
+
"build": "echo 'nothing to build'",
|
|
41
|
+
"check": "echo 'nothing to check'",
|
|
42
|
+
"test": "vitest run",
|
|
43
|
+
"test:watch": "vitest",
|
|
44
|
+
"update-models": "node scripts/update-models.js"
|
|
45
45
|
}
|
|
46
|
-
}
|
|
46
|
+
}
|
package/patch.json
CHANGED
|
@@ -16,7 +16,7 @@
|
|
|
16
16
|
},
|
|
17
17
|
"deepseek-ai/DeepSeek-V4-Pro": {
|
|
18
18
|
"reasoning": true,
|
|
19
|
-
"notes": "`chat_template_kwargs.thinking` via `before_provider_request` payload rewrite; returns `
|
|
19
|
+
"notes": "`chat_template_kwargs.thinking` via `before_provider_request` payload rewrite; returns `reasoning` field",
|
|
20
20
|
"thinkingLevelMap": {
|
|
21
21
|
"minimal": null,
|
|
22
22
|
"low": null,
|
|
@@ -30,76 +30,34 @@
|
|
|
30
30
|
"requiresReasoningContentOnAssistantMessages": true
|
|
31
31
|
}
|
|
32
32
|
},
|
|
33
|
-
"nvidia/Kimi-K2.6-NVFP4": {
|
|
34
|
-
"reasoning": true,
|
|
35
|
-
"input": [
|
|
36
|
-
"text",
|
|
37
|
-
"image"
|
|
38
|
-
],
|
|
39
|
-
"notes": "Reasoning on by default; client-side tool call parsing (vLLM streaming parser bypass)",
|
|
40
|
-
"thinkingLevelMap": {
|
|
41
|
-
"minimal": "low",
|
|
42
|
-
"xhigh": "high"
|
|
43
|
-
},
|
|
44
|
-
"compat": {
|
|
45
|
-
"thinkingFormat": "qwen-chat-template",
|
|
46
|
-
"supportsReasoningEffort": true
|
|
47
|
-
}
|
|
48
|
-
},
|
|
49
|
-
"openai/gpt-oss-120b": {
|
|
50
|
-
"reasoning": true,
|
|
51
|
-
"notes": "Reasoning always on",
|
|
52
|
-
"thinkingLevelMap": {
|
|
53
|
-
"minimal": "low",
|
|
54
|
-
"xhigh": "high"
|
|
55
|
-
},
|
|
56
|
-
"compat": {
|
|
57
|
-
"thinkingFormat": "qwen-chat-template",
|
|
58
|
-
"supportsReasoningEffort": true
|
|
59
|
-
}
|
|
60
|
-
},
|
|
61
33
|
"unsloth/Qwen3.6-27B-NVFP4": {
|
|
62
34
|
"reasoning": true,
|
|
63
|
-
"notes": "`enable_thinking` via `qwen-chat-template`; client-side tool call parsing (vLLM streaming parser bypass)",
|
|
35
|
+
"notes": "`enable_thinking` via `qwen-chat-template`; `preserve_thinking` via `chatTemplateKwargs` for multi-turn continuity; returns `reasoning` field; client-side tool call parsing (vLLM streaming parser bypass)",
|
|
64
36
|
"thinkingLevelMap": {
|
|
65
37
|
"minimal": "low",
|
|
66
38
|
"xhigh": "high"
|
|
67
39
|
},
|
|
68
40
|
"compat": {
|
|
69
41
|
"thinkingFormat": "qwen-chat-template",
|
|
70
|
-
"supportsReasoningEffort": true
|
|
42
|
+
"supportsReasoningEffort": true,
|
|
43
|
+
"chatTemplateKwargs": {
|
|
44
|
+
"preserve_thinking": true
|
|
45
|
+
}
|
|
71
46
|
}
|
|
72
47
|
},
|
|
73
48
|
"unsloth/Qwen3.6-35B-A3B-NVFP4": {
|
|
74
49
|
"reasoning": true,
|
|
75
|
-
"notes": "`enable_thinking` via `qwen-chat-template`; client-side tool call parsing (vLLM streaming parser bypass)",
|
|
50
|
+
"notes": "`enable_thinking` via `qwen-chat-template`; `preserve_thinking` via `chatTemplateKwargs` for multi-turn continuity; returns `reasoning` field; client-side tool call parsing (vLLM streaming parser bypass)",
|
|
76
51
|
"thinkingLevelMap": {
|
|
77
52
|
"minimal": "low",
|
|
78
53
|
"xhigh": "high"
|
|
79
54
|
},
|
|
80
55
|
"compat": {
|
|
81
56
|
"thinkingFormat": "qwen-chat-template",
|
|
82
|
-
"supportsReasoningEffort": true
|
|
83
|
-
}
|
|
84
|
-
},
|
|
85
|
-
"MiniMaxAI/MiniMax-M3-MXFP8": {
|
|
86
|
-
"reasoning": true,
|
|
87
|
-
"input": [
|
|
88
|
-
"text",
|
|
89
|
-
"image"
|
|
90
|
-
],
|
|
91
|
-
"notes": "Reasoning via `chat_template_kwargs.enable_thinking`; returns `reasoning_content` field",
|
|
92
|
-
"thinkingLevelMap": {
|
|
93
|
-
"minimal": null,
|
|
94
|
-
"low": null,
|
|
95
|
-
"medium": null,
|
|
96
|
-
"high": "high",
|
|
97
|
-
"xhigh": "max"
|
|
98
|
-
},
|
|
99
|
-
"compat": {
|
|
100
|
-
"thinkingFormat": "deepseek",
|
|
101
57
|
"supportsReasoningEffort": true,
|
|
102
|
-
"
|
|
58
|
+
"chatTemplateKwargs": {
|
|
59
|
+
"preserve_thinking": true
|
|
60
|
+
}
|
|
103
61
|
}
|
|
104
62
|
},
|
|
105
63
|
"moonshotai/Kimi-K2.7-Code": {
|
|
@@ -108,37 +66,41 @@
|
|
|
108
66
|
"text",
|
|
109
67
|
"image"
|
|
110
68
|
],
|
|
111
|
-
"notes": "Reasoning on by default; client-side tool call parsing (vLLM streaming parser bypass)",
|
|
69
|
+
"notes": "Reasoning on by default (thinking-only model); `preserve_thinking` via `chatTemplateKwargs` for multi-turn continuity; returns `reasoning` field; client-side tool call parsing (vLLM streaming parser bypass)",
|
|
112
70
|
"thinkingLevelMap": {
|
|
113
71
|
"minimal": "low",
|
|
114
72
|
"xhigh": "high"
|
|
115
73
|
},
|
|
116
74
|
"compat": {
|
|
117
75
|
"thinkingFormat": "qwen-chat-template",
|
|
118
|
-
"supportsReasoningEffort": true
|
|
76
|
+
"supportsReasoningEffort": true,
|
|
77
|
+
"chatTemplateKwargs": {
|
|
78
|
+
"preserve_thinking": true
|
|
79
|
+
}
|
|
119
80
|
}
|
|
120
81
|
},
|
|
121
|
-
"zai-org/GLM-5.
|
|
122
|
-
"contextWindow": 200000,
|
|
82
|
+
"zai-org/GLM-5.2-FP8": {
|
|
123
83
|
"reasoning": true,
|
|
124
|
-
"notes": "`enable_thinking` via `qwen-chat-template`;
|
|
84
|
+
"notes": "`enable_thinking` via `qwen-chat-template`; effort via `reasoning_effort` (only `high`/`max` distinct, per vLLM GLM-5.2 recipe); thinking levels aligned with neuralwatt GLM 5.2; returns `reasoning` field",
|
|
125
85
|
"thinkingLevelMap": {
|
|
126
|
-
"minimal":
|
|
127
|
-
"
|
|
86
|
+
"minimal": null,
|
|
87
|
+
"low": null,
|
|
88
|
+
"medium": null,
|
|
89
|
+
"high": "high",
|
|
90
|
+
"xhigh": "max"
|
|
128
91
|
},
|
|
129
92
|
"compat": {
|
|
130
93
|
"thinkingFormat": "qwen-chat-template",
|
|
131
|
-
"supportsReasoningEffort": true
|
|
132
|
-
"zaiToolStream": true
|
|
94
|
+
"supportsReasoningEffort": true
|
|
133
95
|
}
|
|
134
96
|
},
|
|
135
|
-
"zai-org/GLM-5.2-
|
|
97
|
+
"zai-org/GLM-5.2-NVFP4": {
|
|
136
98
|
"reasoning": true,
|
|
137
|
-
"notes": "`enable_thinking` via `qwen-chat-template`;
|
|
99
|
+
"notes": "`enable_thinking` via `qwen-chat-template`; effort via `reasoning_effort` (only `high`/`max` distinct, per vLLM GLM-5.2 recipe); returns `reasoning` field",
|
|
138
100
|
"thinkingLevelMap": {
|
|
139
|
-
"minimal":
|
|
140
|
-
"low":
|
|
141
|
-
"medium":
|
|
101
|
+
"minimal": null,
|
|
102
|
+
"low": null,
|
|
103
|
+
"medium": null,
|
|
142
104
|
"high": "high",
|
|
143
105
|
"xhigh": "max"
|
|
144
106
|
},
|