pi-makora-provider 1.0.0 → 1.2.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +36 -19
- package/index.ts +134 -10
- package/models.json +10 -52
- package/package.json +14 -9
- package/patch.json +46 -7
package/README.md
CHANGED
|
@@ -4,7 +4,7 @@
|
|
|
4
4
|
|
|
5
5
|
**Open-weight models through [Makora](https://inference.makora.com)**
|
|
6
6
|
|
|
7
|
-
_DeepSeek V4, Kimi K2.
|
|
7
|
+
_DeepSeek V4, Kimi K2.7 Code, GLM 5.2, Qwen 3.6 — with client-side tool call repair and preserved-thinking flags for [pi](https://github.com/earendil-works/pi-coding-agent)._
|
|
8
8
|
|
|
9
9
|
[](https://github.com/earendil-works/pi-coding-agent)
|
|
10
10
|
[](./LICENSE)
|
|
@@ -19,21 +19,36 @@ _DeepSeek V4, Kimi K2.6, GLM 5.1, Qwen 3.6 — with client-side tool call repair
|
|
|
19
19
|
| Model | ID | Reasoning | Notes |
|
|
20
20
|
|-------|----|-----------|-------|
|
|
21
21
|
| DeepSeek V4 Flash | `deepseek-ai/DeepSeek-V4-Flash` | Yes | `include_reasoning` + `chat_template_kwargs.thinking` via `before_provider_request` payload rewrite; returns `reasoning` field |
|
|
22
|
-
| DeepSeek V4 Pro | `deepseek-ai/DeepSeek-V4-Pro` | Yes | `chat_template_kwargs.thinking` via `before_provider_request` payload rewrite; returns `
|
|
23
|
-
| GLM 5.
|
|
24
|
-
|
|
|
25
|
-
| Kimi K2.
|
|
26
|
-
| Kimi K2.7 Code | `moonshotai/Kimi-K2.7-Code` | Yes | Reasoning on by default; client-side tool call parsing (vLLM streaming parser bypass) |
|
|
22
|
+
| DeepSeek V4 Pro | `deepseek-ai/DeepSeek-V4-Pro` | Yes | `chat_template_kwargs.thinking` via `before_provider_request` payload rewrite; returns `reasoning` field |
|
|
23
|
+
| GLM 5.2 FP8 | `zai-org/GLM-5.2-FP8` | Yes | `enable_thinking` via `qwen-chat-template`; effort via `reasoning_effort` (only `high`/`max` distinct, per vLLM GLM-5.2 recipe); thinking levels aligned with neuralwatt GLM 5.2; returns `reasoning` field |
|
|
24
|
+
| GLM 5.2 NVFP4 | `zai-org/GLM-5.2-NVFP4` | Yes | `enable_thinking` via `qwen-chat-template`; effort via `reasoning_effort` (only `high`/`max` distinct, per vLLM GLM-5.2 recipe); returns `reasoning` field |
|
|
25
|
+
| Kimi K2.7 Code | `moonshotai/Kimi-K2.7-Code` | Yes | Reasoning on by default (thinking-only model); `preserve_thinking` via `chatTemplateKwargs` for multi-turn continuity; returns `reasoning` field; client-side tool call parsing (vLLM streaming parser bypass) |
|
|
27
26
|
| Llama 3.3 70B FP8 | `amd/Llama-3.3-70B-Instruct-FP8-KV` | No | |
|
|
28
27
|
| Llama 3.3 70B Instruct | `meta-llama/Llama-3.3-70B-Instruct` | No | |
|
|
29
|
-
|
|
|
30
|
-
| Qwen 3.6
|
|
31
|
-
| Qwen 3.6 35B A3B NVFP4 | `unsloth/Qwen3.6-35B-A3B-NVFP4` | Yes | `enable_thinking` via `qwen-chat-template`; client-side tool call parsing (vLLM streaming parser bypass) |
|
|
28
|
+
| Qwen 3.6 27B NVFP4 | `unsloth/Qwen3.6-27B-NVFP4` | Yes | `enable_thinking` via `qwen-chat-template`; `preserve_thinking` via `chatTemplateKwargs` for multi-turn continuity; returns `reasoning` field; client-side tool call parsing (vLLM streaming parser bypass) |
|
|
29
|
+
| Qwen 3.6 35B A3B NVFP4 | `unsloth/Qwen3.6-35B-A3B-NVFP4` | Yes | `enable_thinking` via `qwen-chat-template`; `preserve_thinking` via `chatTemplateKwargs` for multi-turn continuity; returns `reasoning` field; client-side tool call parsing (vLLM streaming parser bypass) |
|
|
32
30
|
<!-- MODELS_TABLE_END -->
|
|
33
31
|
|
|
34
32
|
## Installation
|
|
35
33
|
|
|
36
|
-
### Option 1:
|
|
34
|
+
### Option 1: Install from npm (Recommended)
|
|
35
|
+
|
|
36
|
+
```bash
|
|
37
|
+
pi install npm:pi-makora-provider
|
|
38
|
+
```
|
|
39
|
+
|
|
40
|
+
Then set your API key and run pi:
|
|
41
|
+
```bash
|
|
42
|
+
# Recommended: add to auth.json
|
|
43
|
+
# See Authentication section below
|
|
44
|
+
|
|
45
|
+
# Or set as environment variable
|
|
46
|
+
export MAKORA_OPTIMIZE_TOKEN=your-api-key-here
|
|
47
|
+
|
|
48
|
+
pi
|
|
49
|
+
```
|
|
50
|
+
|
|
51
|
+
### Option 2: Using `pi install` from GitHub
|
|
37
52
|
|
|
38
53
|
Install directly from GitHub:
|
|
39
54
|
|
|
@@ -52,7 +67,7 @@ export MAKORA_OPTIMIZE_TOKEN=your-api-key-here
|
|
|
52
67
|
pi
|
|
53
68
|
```
|
|
54
69
|
|
|
55
|
-
### Option
|
|
70
|
+
### Option 3: Manual Clone
|
|
56
71
|
|
|
57
72
|
1. Clone this repository:
|
|
58
73
|
```bash
|
|
@@ -146,17 +161,19 @@ Do **not** edit `models.json` directly — it is auto-generated from the API. To
|
|
|
146
161
|
|
|
147
162
|
These issues are common to all vLLM-hosted providers and affect Makora models:
|
|
148
163
|
|
|
149
|
-
- **GLM
|
|
164
|
+
- **GLM tool calls — follow-up `tool_calls` crash workaround**: the ZAI chat template calls `.items()` on any assistant message that carries a `tool_calls` field (on the `tool_calls` list or on its JSON-string `arguments` field), raising a leaked Python `AttributeError` that vLLM returns as HTTP 400 (`'list object' has no attribute 'items'` / `'str object' has no attribute 'items'`). The `before_provider_request` hook (`stripGlmToolCalls`) strips `tool_calls` from assistant messages and converts them to GLM-native `<tool_call>` XML text in `content` — which the model natively understands in conversation history — before the follow-up request is sent. The hook is scoped to Makora's own GLM models via an exact model-id match (`isMakoraGlmVllmModel`): `before_provider_request` is global in pi (it fires for every loaded provider's requests and the event carries no `provider` field), so gating by model id is what keeps it from touching GLM models served by other providers. If upstream fixes the `.items()` crash, this transform is a harmless no-op (the XML text is valid GLM input either way).
|
|
150
165
|
|
|
151
|
-
|
|
166
|
+
- **Kimi K2.7 Code + Qwen 3.6 tool calling**: vLLM's server-side streaming tool-call handling is broken or missing for these models — Makora's vLLM is missing both `--enable-auto-tool-choice` and `--tool-call-parser` for them, so they emit tool calls as raw text tokens rather than structured `tool_call` deltas.
|
|
152
167
|
|
|
153
|
-
- **Kimi K2.
|
|
154
|
-
- **Qwen 3.6**: Uses hermes-style `<function=...>` XML, sometimes with `█` delimiters.
|
|
168
|
+
- **Kimi K2.7 Code**: Uses `<|tool_call_begin|>...<|tool_call_end|>` tokens.
|
|
169
|
+
- **Qwen 3.6**: Uses hermes-style `<function=...>` XML, sometimes with `█` delimiters.
|
|
155
170
|
|
|
156
|
-
- **GLM 5.
|
|
171
|
+
- **GLM 5.2 CoT leak**: On some vLLM builds, disabling reasoning may still leak chain-of-thought into `content` terminated by a ``` marker. See [vllm-project/vllm#31319](https://github.com/vllm-project/vllm/issues/31319).
|
|
157
172
|
|
|
158
173
|
- **DeepSeek V4 reasoning**: The official DeepSeek API uses `thinking: { type: "enabled" }` which Makora's vLLM silently ignores. The `before_provider_request` hook rewrites the payload to use vLLM-native params instead:
|
|
159
|
-
- **DS V4 Pro**: `chat_template_kwargs: { thinking: true }`. Returns `reasoning_content
|
|
174
|
+
- **DS V4 Pro**: `chat_template_kwargs: { thinking: true }`. Returns `reasoning` (vLLM reasoning parser; the official DeepSeek API returns `reasoning_content`, but Makora's vLLM build returns `reasoning`).
|
|
160
175
|
- **DS V4 Flash**: `include_reasoning: true` + `chat_template_kwargs: { thinking: true }`. `include_reasoning` alone returns `reasoning: null` on this vLLM build — both params are required. Returns `reasoning`.
|
|
161
|
-
- **
|
|
162
|
-
- **
|
|
176
|
+
- **Preserved thinking (multi-turn reasoning continuity)**: every reasoning model on Makora returns reasoning in the `reasoning` field (not `reasoning_content`). Pi sends prior assistant reasoning back on follow-up turns; Makora's vLLM chat templates accept it under either `reasoning` or `reasoning_content`. For models whose chat template defines an explicit preserved-thinking flag, the flag is sent via `compat.chatTemplateKwargs`:
|
|
177
|
+
- **Kimi K2.7 Code**: `preserve_thinking: true` (the model forces thinking-only mode with preserve_thinking on, per the vLLM recipe).
|
|
178
|
+
- **Qwen 3.6 27B / 35B**: `preserve_thinking: true` (Qwen 3.6 ships with preserve_thinking).
|
|
179
|
+
- **GLM 5.2**: no extra flag — the vLLM GLM-5.2 recipe uses only `enable_thinking` + `reasoning_effort`; preserved thinking works without a dedicated flag. Effort honors only `high` and `max` (lower pi levels resolve to the default `max`).
|
package/index.ts
CHANGED
|
@@ -17,7 +17,7 @@
|
|
|
17
17
|
* pi sends thinking: { type } via the "deepseek" thinkingFormat, but vLLM
|
|
18
18
|
* ignores that — the before_provider_request hook rewrites the payload to
|
|
19
19
|
* use chat_template_kwargs: { thinking: true } instead.
|
|
20
|
-
* Returns
|
|
20
|
+
* Returns reasoning field (vLLM reasoning parser; not reasoning_content).
|
|
21
21
|
* - DeepSeek V4 Flash: reasoning via include_reasoning +
|
|
22
22
|
* chat_template_kwargs.thinking on vLLM.
|
|
23
23
|
* The before_provider_request hook rewrites the payload to replace
|
|
@@ -34,11 +34,19 @@
|
|
|
34
34
|
* delta. Setting zaiToolStream: true sends tool_stream: true in the
|
|
35
35
|
* request, which forces vLLM to use the explicit tool streaming path
|
|
36
36
|
* that correctly emits tool call chunks.
|
|
37
|
+
* - GLM 5.2 FP8 / NVFP4: reasoning via chat_template_kwargs.enable_thinking.
|
|
38
|
+
* Effort uses vLLM's reasoning_effort field (only `high` and `max` are
|
|
39
|
+
* distinct levels per the vLLM GLM-5.2 recipe; lower pi levels resolve to
|
|
40
|
+
* the default). Thinking levels aligned with the neuralwatt provider's
|
|
41
|
+
* GLM 5.2 configuration and mapped through pi's qwen-chat-template
|
|
42
|
+
* thinkingFormat. Returns `reasoning` field.
|
|
37
43
|
* - GPT-OSS 120B: reasoning always on; returns `reasoning` field.
|
|
38
|
-
* - Kimi K2.
|
|
39
|
-
*
|
|
44
|
+
* - Kimi K2.7 Code: reasoning always on (thinking-only model);
|
|
45
|
+
* chatTemplateKwargs.preserve_thinking forces multi-turn reasoning
|
|
46
|
+
* continuity. Returns `reasoning` field. Can be toggled via enable_thinking.
|
|
40
47
|
* - Qwen 3.6 models: reasoning via chat_template_kwargs.enable_thinking;
|
|
41
|
-
*
|
|
48
|
+
* chatTemplateKwargs.preserve_thinking for multi-turn continuity.
|
|
49
|
+
* Returns `reasoning` field.
|
|
42
50
|
* - MiniMax M3 MXFP8: reasoning via chat_template_kwargs.enable_thinking;
|
|
43
51
|
* returns reasoning_content field.
|
|
44
52
|
* - Llama 3.3 70B: not a reasoning model.
|
|
@@ -105,6 +113,10 @@ interface JsonModel {
|
|
|
105
113
|
requiresToolResultName?: boolean;
|
|
106
114
|
requiresAssistantAfterToolResult?: boolean;
|
|
107
115
|
cacheControlFormat?: "anthropic";
|
|
116
|
+
/** Extra keys merged into vLLM `chat_template_kwargs` on every request.
|
|
117
|
+
* Used for preserved-thinking flags like `preserve_thinking` / `clear_thinking`
|
|
118
|
+
* that some chat templates require for multi-turn reasoning continuity. */
|
|
119
|
+
chatTemplateKwargs?: Record<string, unknown>;
|
|
108
120
|
};
|
|
109
121
|
}
|
|
110
122
|
|
|
@@ -129,6 +141,11 @@ interface PatchEntry {
|
|
|
129
141
|
|
|
130
142
|
type PatchMap = Record<string, PatchEntry>;
|
|
131
143
|
|
|
144
|
+
/** Type guard: non-null, non-array object. */
|
|
145
|
+
export function isObject(value: unknown): value is Record<string, unknown> {
|
|
146
|
+
return value !== null && typeof value === "object" && !Array.isArray(value);
|
|
147
|
+
}
|
|
148
|
+
|
|
132
149
|
// Patch Application
|
|
133
150
|
|
|
134
151
|
function applyPatch(model: JsonModel, patch: PatchEntry): JsonModel {
|
|
@@ -214,6 +231,38 @@ const MINIMAX_M3_ID = "MiniMaxAI/MiniMax-M3-MXFP8";
|
|
|
214
231
|
const DS_VLLM_MODELS = new Set([DS_PRO_ID, DS_FLASH_ID]);
|
|
215
232
|
const ENABLE_THINKING_VLLM_MODELS = new Set([MINIMAX_M3_ID]);
|
|
216
233
|
|
|
234
|
+
/**
|
|
235
|
+
* Makora's GLM models, built from the same models list this extension
|
|
236
|
+
* registers. `before_provider_request` is GLOBAL in pi — it fires for every
|
|
237
|
+
* loaded provider's requests, not just makora's — and its event carries only
|
|
238
|
+
* `payload` (no `provider` field to gate on). So the GLM tool_calls strip MUST
|
|
239
|
+
* be scoped to makora's own GLM model ids. A loose `/glm/i` regex would also
|
|
240
|
+
* rewrite tool_calls for every sibling provider that serves any GLM model
|
|
241
|
+
* (baseten, io, lilac, neuralwatt, tensorix, fireworks, crofai, hypercharm,
|
|
242
|
+
* wafer, parasail, umans, opencode, ...), silently breaking their tool calling
|
|
243
|
+
* and clobbering the input their own before_provider_request handlers expect.
|
|
244
|
+
*
|
|
245
|
+
* Derived (not hardcoded) so it auto-tracks models.json refreshes — models.json
|
|
246
|
+
* is auto-generated by scripts/update-models.js. All makora models run on vLLM,
|
|
247
|
+
* so every GLM id here shares the ZAI chat-template `.items()` crash that
|
|
248
|
+
* stripGlmToolCalls works around.
|
|
249
|
+
*/
|
|
250
|
+
const allMakoraModels = buildModels(
|
|
251
|
+
modelsData as JsonModel[],
|
|
252
|
+
customModelsData as JsonModel[],
|
|
253
|
+
patchData as PatchMap,
|
|
254
|
+
);
|
|
255
|
+
const GLM_VLLM_MODELS = new Set(
|
|
256
|
+
allMakoraModels.filter((m) => /glm/i.test(m.id)).map((m) => m.id),
|
|
257
|
+
);
|
|
258
|
+
|
|
259
|
+
/** Whether `model` is one of makora's GLM models — the only ids the
|
|
260
|
+
* before_provider_request hook should strip tool_calls for. Exported so the
|
|
261
|
+
* scoping can be unit-tested in isolation. */
|
|
262
|
+
export function isMakoraGlmVllmModel(model: string): boolean {
|
|
263
|
+
return GLM_VLLM_MODELS.has(model);
|
|
264
|
+
}
|
|
265
|
+
|
|
217
266
|
/**
|
|
218
267
|
* Intercept the request payload for models that need vLLM-specific thinking
|
|
219
268
|
* param rewrites.
|
|
@@ -262,12 +311,82 @@ function rewriteVllmPayload(payload: Record<string, unknown>): Record<string, un
|
|
|
262
311
|
return p;
|
|
263
312
|
}
|
|
264
313
|
|
|
265
|
-
|
|
266
|
-
|
|
267
|
-
|
|
268
|
-
|
|
314
|
+
/**
|
|
315
|
+
* GLM models on Makora's vLLM crash with a leaked Python AttributeError
|
|
316
|
+
* (`'list object' has no attribute 'items'` or `'str object' has no attribute 'items'`)
|
|
317
|
+
* when any assistant message in the request contains a `tool_calls` field.
|
|
318
|
+
* The ZAI/vLLM chat template calls `.items()` on the tool_calls list (or on the
|
|
319
|
+
* JSON-string `arguments` field), which raises AttributeError and leaks into the
|
|
320
|
+
* HTTP 400 response body.
|
|
321
|
+
*
|
|
322
|
+
* Fix: for makora's GLM models only (see GLM_VLLM_MODELS / isMakoraGlmVllmModel),
|
|
323
|
+
* strip `tool_calls` from assistant messages in the before_provider_request hook
|
|
324
|
+
* and convert them back to GLM's native
|
|
325
|
+
* `<tool_call>` XML text in `content`. The model natively understands this
|
|
326
|
+
* format in conversation history, and the `role: "tool"` result messages that
|
|
327
|
+
* follow are rendered fine by the chat template's tool-observation branch.
|
|
328
|
+
*
|
|
329
|
+
* If upstream fixes both the streaming parser and the 500/400 crash, this
|
|
330
|
+
* transform becomes a harmless no-op (the XML text is still valid GLM input).
|
|
331
|
+
*/
|
|
269
332
|
|
|
270
|
-
|
|
333
|
+
export function toolCallToGlmXml(tc: Record<string, unknown>): string {
|
|
334
|
+
const fn = (isObject(tc.function) ? tc.function : {}) as Record<string, unknown>;
|
|
335
|
+
const name = typeof fn.name === "string" ? fn.name : "";
|
|
336
|
+
const argsStr = typeof fn.arguments === "string" ? fn.arguments : "{}";
|
|
337
|
+
let args: Record<string, unknown>;
|
|
338
|
+
try {
|
|
339
|
+
args = JSON.parse(argsStr) as Record<string, unknown>;
|
|
340
|
+
} catch {
|
|
341
|
+
args = {};
|
|
342
|
+
}
|
|
343
|
+
if (!isObject(args)) args = {};
|
|
344
|
+
const argLines = Object.entries(args).map(
|
|
345
|
+
([key, value]) =>
|
|
346
|
+
`<arg_key>${key}</arg_key>\n<arg_value>${
|
|
347
|
+
typeof value === "string" ? value : JSON.stringify(value)
|
|
348
|
+
}</arg_value>`,
|
|
349
|
+
);
|
|
350
|
+
return `<tool_call>${name}\n${argLines.join("\n")}\n</tool_call>`;
|
|
351
|
+
}
|
|
352
|
+
|
|
353
|
+
export function stripGlmToolCalls(payload: Record<string, unknown>): Record<string, unknown> {
|
|
354
|
+
const messages = payload.messages;
|
|
355
|
+
if (!Array.isArray(messages)) return payload;
|
|
356
|
+
|
|
357
|
+
let modified = false;
|
|
358
|
+
const newMessages = messages.map((msg) => {
|
|
359
|
+
if (!isObject(msg)) return msg;
|
|
360
|
+
if (msg.role !== "assistant") return msg;
|
|
361
|
+
|
|
362
|
+
const toolCalls = msg.tool_calls;
|
|
363
|
+
if (!Array.isArray(toolCalls) || toolCalls.length === 0) return msg;
|
|
364
|
+
|
|
365
|
+
const xmlBlocks = toolCalls
|
|
366
|
+
.filter((tc): tc is Record<string, unknown> => isObject(tc))
|
|
367
|
+
.map(toolCallToGlmXml);
|
|
368
|
+
if (xmlBlocks.length === 0) return msg;
|
|
369
|
+
|
|
370
|
+
const toolCallText = xmlBlocks.join("\n");
|
|
371
|
+
// pi's openai-completions provider always serializes assistant content as a
|
|
372
|
+
// plain string, but guard against null/array content for robustness.
|
|
373
|
+
const existingContent = typeof msg.content === "string" ? msg.content : "";
|
|
374
|
+
const newContent = existingContent ? `${existingContent}\n${toolCallText}` : toolCallText;
|
|
375
|
+
|
|
376
|
+
modified = true;
|
|
377
|
+
const rest: Record<string, unknown> = {};
|
|
378
|
+
for (const [k, v] of Object.entries(msg)) {
|
|
379
|
+
if (k !== "tool_calls") rest[k] = v;
|
|
380
|
+
}
|
|
381
|
+
return { ...rest, content: newContent };
|
|
382
|
+
});
|
|
383
|
+
|
|
384
|
+
if (!modified) return payload;
|
|
385
|
+
return { ...payload, messages: newMessages };
|
|
386
|
+
}
|
|
387
|
+
|
|
388
|
+
export default function (pi: ExtensionAPI) {
|
|
389
|
+
const models = allMakoraModels;
|
|
271
390
|
|
|
272
391
|
// apiKey resolution order: auth.json ("makora" key) → MAKORA_OPTIMIZE_TOKEN env var
|
|
273
392
|
pi.registerProvider(PROVIDER_ID, {
|
|
@@ -281,7 +400,12 @@ export default function (pi: ExtensionAPI) {
|
|
|
281
400
|
pi.on("before_provider_request", (event) => {
|
|
282
401
|
const payload = event.payload as Record<string, unknown> | undefined;
|
|
283
402
|
if (!payload || typeof payload.model !== "string") return;
|
|
284
|
-
|
|
403
|
+
|
|
404
|
+
let result = rewriteVllmPayload(payload);
|
|
405
|
+
if (isMakoraGlmVllmModel(payload.model)) {
|
|
406
|
+
result = stripGlmToolCalls(result);
|
|
407
|
+
}
|
|
408
|
+
return result;
|
|
285
409
|
});
|
|
286
410
|
}
|
|
287
411
|
|
package/models.json
CHANGED
|
@@ -62,27 +62,6 @@
|
|
|
62
62
|
"maxTokensField": "max_completion_tokens"
|
|
63
63
|
}
|
|
64
64
|
},
|
|
65
|
-
{
|
|
66
|
-
"id": "MiniMaxAI/MiniMax-M3-MXFP8",
|
|
67
|
-
"name": "MiniMax M3 MXFP8",
|
|
68
|
-
"reasoning": false,
|
|
69
|
-
"input": [
|
|
70
|
-
"text"
|
|
71
|
-
],
|
|
72
|
-
"cost": {
|
|
73
|
-
"input": 0,
|
|
74
|
-
"output": 0,
|
|
75
|
-
"cacheRead": 0,
|
|
76
|
-
"cacheWrite": 0
|
|
77
|
-
},
|
|
78
|
-
"contextWindow": 1048576,
|
|
79
|
-
"maxTokens": 0,
|
|
80
|
-
"compat": {
|
|
81
|
-
"supportsDeveloperRole": false,
|
|
82
|
-
"supportsStore": false,
|
|
83
|
-
"maxTokensField": "max_completion_tokens"
|
|
84
|
-
}
|
|
85
|
-
},
|
|
86
65
|
{
|
|
87
66
|
"id": "moonshotai/Kimi-K2.7-Code",
|
|
88
67
|
"name": "Kimi K2.7 Code",
|
|
@@ -105,8 +84,8 @@
|
|
|
105
84
|
}
|
|
106
85
|
},
|
|
107
86
|
{
|
|
108
|
-
"id": "
|
|
109
|
-
"name": "
|
|
87
|
+
"id": "unsloth/Qwen3.6-27B-NVFP4",
|
|
88
|
+
"name": "Qwen 3.6 27B NVFP4",
|
|
110
89
|
"reasoning": false,
|
|
111
90
|
"input": [
|
|
112
91
|
"text"
|
|
@@ -126,29 +105,8 @@
|
|
|
126
105
|
}
|
|
127
106
|
},
|
|
128
107
|
{
|
|
129
|
-
"id": "
|
|
130
|
-
"name": "
|
|
131
|
-
"reasoning": false,
|
|
132
|
-
"input": [
|
|
133
|
-
"text"
|
|
134
|
-
],
|
|
135
|
-
"cost": {
|
|
136
|
-
"input": 0,
|
|
137
|
-
"output": 0,
|
|
138
|
-
"cacheRead": 0,
|
|
139
|
-
"cacheWrite": 0
|
|
140
|
-
},
|
|
141
|
-
"contextWindow": 131072,
|
|
142
|
-
"maxTokens": 0,
|
|
143
|
-
"compat": {
|
|
144
|
-
"supportsDeveloperRole": false,
|
|
145
|
-
"supportsStore": false,
|
|
146
|
-
"maxTokensField": "max_completion_tokens"
|
|
147
|
-
}
|
|
148
|
-
},
|
|
149
|
-
{
|
|
150
|
-
"id": "unsloth/Qwen3.6-27B-NVFP4",
|
|
151
|
-
"name": "Qwen 3.6 27B NVFP4",
|
|
108
|
+
"id": "unsloth/Qwen3.6-35B-A3B-NVFP4",
|
|
109
|
+
"name": "Qwen 3.6 35B A3B NVFP4",
|
|
152
110
|
"reasoning": false,
|
|
153
111
|
"input": [
|
|
154
112
|
"text"
|
|
@@ -168,8 +126,8 @@
|
|
|
168
126
|
}
|
|
169
127
|
},
|
|
170
128
|
{
|
|
171
|
-
"id": "
|
|
172
|
-
"name": "
|
|
129
|
+
"id": "zai-org/GLM-5.2-FP8",
|
|
130
|
+
"name": "GLM 5.2 FP8",
|
|
173
131
|
"reasoning": false,
|
|
174
132
|
"input": [
|
|
175
133
|
"text"
|
|
@@ -180,7 +138,7 @@
|
|
|
180
138
|
"cacheRead": 0,
|
|
181
139
|
"cacheWrite": 0
|
|
182
140
|
},
|
|
183
|
-
"contextWindow":
|
|
141
|
+
"contextWindow": 1048576,
|
|
184
142
|
"maxTokens": 0,
|
|
185
143
|
"compat": {
|
|
186
144
|
"supportsDeveloperRole": false,
|
|
@@ -189,8 +147,8 @@
|
|
|
189
147
|
}
|
|
190
148
|
},
|
|
191
149
|
{
|
|
192
|
-
"id": "zai-org/GLM-5.
|
|
193
|
-
"name": "GLM 5.
|
|
150
|
+
"id": "zai-org/GLM-5.2-NVFP4",
|
|
151
|
+
"name": "GLM 5.2 NVFP4",
|
|
194
152
|
"reasoning": false,
|
|
195
153
|
"input": [
|
|
196
154
|
"text"
|
|
@@ -201,7 +159,7 @@
|
|
|
201
159
|
"cacheRead": 0,
|
|
202
160
|
"cacheWrite": 0
|
|
203
161
|
},
|
|
204
|
-
"contextWindow":
|
|
162
|
+
"contextWindow": 1048576,
|
|
205
163
|
"maxTokens": 0,
|
|
206
164
|
"compat": {
|
|
207
165
|
"supportsDeveloperRole": false,
|
package/package.json
CHANGED
|
@@ -1,15 +1,9 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "pi-makora-provider",
|
|
3
|
-
"version": "1.
|
|
4
|
-
"description": "Makora provider extension for pi - Access DeepSeek V4, GLM 5.
|
|
3
|
+
"version": "1.2.0",
|
|
4
|
+
"description": "Makora provider extension for pi - Access DeepSeek V4, GLM 5.2, Kimi K2.7 Code, Llama 3.3, Qwen 3.6, and more through the Makora inference API",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"main": "index.ts",
|
|
7
|
-
"scripts": {
|
|
8
|
-
"clean": "echo 'nothing to clean'",
|
|
9
|
-
"build": "echo 'nothing to build'",
|
|
10
|
-
"check": "echo 'nothing to check'",
|
|
11
|
-
"update-models": "node scripts/update-models.js"
|
|
12
|
-
},
|
|
13
7
|
"keywords": [
|
|
14
8
|
"pi",
|
|
15
9
|
"extension",
|
|
@@ -37,5 +31,16 @@
|
|
|
37
31
|
"extensions": [
|
|
38
32
|
"./index.ts"
|
|
39
33
|
]
|
|
34
|
+
},
|
|
35
|
+
"devDependencies": {
|
|
36
|
+
"vitest": "^4.1.9"
|
|
37
|
+
},
|
|
38
|
+
"scripts": {
|
|
39
|
+
"clean": "echo 'nothing to clean'",
|
|
40
|
+
"build": "echo 'nothing to build'",
|
|
41
|
+
"check": "echo 'nothing to check'",
|
|
42
|
+
"test": "vitest run",
|
|
43
|
+
"test:watch": "vitest",
|
|
44
|
+
"update-models": "node scripts/update-models.js"
|
|
40
45
|
}
|
|
41
|
-
}
|
|
46
|
+
}
|
package/patch.json
CHANGED
|
@@ -16,7 +16,7 @@
|
|
|
16
16
|
},
|
|
17
17
|
"deepseek-ai/DeepSeek-V4-Pro": {
|
|
18
18
|
"reasoning": true,
|
|
19
|
-
"notes": "`chat_template_kwargs.thinking` via `before_provider_request` payload rewrite; returns `
|
|
19
|
+
"notes": "`chat_template_kwargs.thinking` via `before_provider_request` payload rewrite; returns `reasoning` field",
|
|
20
20
|
"thinkingLevelMap": {
|
|
21
21
|
"minimal": null,
|
|
22
22
|
"low": null,
|
|
@@ -60,26 +60,32 @@
|
|
|
60
60
|
},
|
|
61
61
|
"unsloth/Qwen3.6-27B-NVFP4": {
|
|
62
62
|
"reasoning": true,
|
|
63
|
-
"notes": "`enable_thinking` via `qwen-chat-template`; client-side tool call parsing (vLLM streaming parser bypass)",
|
|
63
|
+
"notes": "`enable_thinking` via `qwen-chat-template`; `preserve_thinking` via `chatTemplateKwargs` for multi-turn continuity; returns `reasoning` field; client-side tool call parsing (vLLM streaming parser bypass)",
|
|
64
64
|
"thinkingLevelMap": {
|
|
65
65
|
"minimal": "low",
|
|
66
66
|
"xhigh": "high"
|
|
67
67
|
},
|
|
68
68
|
"compat": {
|
|
69
69
|
"thinkingFormat": "qwen-chat-template",
|
|
70
|
-
"supportsReasoningEffort": true
|
|
70
|
+
"supportsReasoningEffort": true,
|
|
71
|
+
"chatTemplateKwargs": {
|
|
72
|
+
"preserve_thinking": true
|
|
73
|
+
}
|
|
71
74
|
}
|
|
72
75
|
},
|
|
73
76
|
"unsloth/Qwen3.6-35B-A3B-NVFP4": {
|
|
74
77
|
"reasoning": true,
|
|
75
|
-
"notes": "`enable_thinking` via `qwen-chat-template`; client-side tool call parsing (vLLM streaming parser bypass)",
|
|
78
|
+
"notes": "`enable_thinking` via `qwen-chat-template`; `preserve_thinking` via `chatTemplateKwargs` for multi-turn continuity; returns `reasoning` field; client-side tool call parsing (vLLM streaming parser bypass)",
|
|
76
79
|
"thinkingLevelMap": {
|
|
77
80
|
"minimal": "low",
|
|
78
81
|
"xhigh": "high"
|
|
79
82
|
},
|
|
80
83
|
"compat": {
|
|
81
84
|
"thinkingFormat": "qwen-chat-template",
|
|
82
|
-
"supportsReasoningEffort": true
|
|
85
|
+
"supportsReasoningEffort": true,
|
|
86
|
+
"chatTemplateKwargs": {
|
|
87
|
+
"preserve_thinking": true
|
|
88
|
+
}
|
|
83
89
|
}
|
|
84
90
|
},
|
|
85
91
|
"MiniMaxAI/MiniMax-M3-MXFP8": {
|
|
@@ -108,14 +114,17 @@
|
|
|
108
114
|
"text",
|
|
109
115
|
"image"
|
|
110
116
|
],
|
|
111
|
-
"notes": "Reasoning on by default; client-side tool call parsing (vLLM streaming parser bypass)",
|
|
117
|
+
"notes": "Reasoning on by default (thinking-only model); `preserve_thinking` via `chatTemplateKwargs` for multi-turn continuity; returns `reasoning` field; client-side tool call parsing (vLLM streaming parser bypass)",
|
|
112
118
|
"thinkingLevelMap": {
|
|
113
119
|
"minimal": "low",
|
|
114
120
|
"xhigh": "high"
|
|
115
121
|
},
|
|
116
122
|
"compat": {
|
|
117
123
|
"thinkingFormat": "qwen-chat-template",
|
|
118
|
-
"supportsReasoningEffort": true
|
|
124
|
+
"supportsReasoningEffort": true,
|
|
125
|
+
"chatTemplateKwargs": {
|
|
126
|
+
"preserve_thinking": true
|
|
127
|
+
}
|
|
119
128
|
}
|
|
120
129
|
},
|
|
121
130
|
"zai-org/GLM-5.1-FP8": {
|
|
@@ -131,5 +140,35 @@
|
|
|
131
140
|
"supportsReasoningEffort": true,
|
|
132
141
|
"zaiToolStream": true
|
|
133
142
|
}
|
|
143
|
+
},
|
|
144
|
+
"zai-org/GLM-5.2-FP8": {
|
|
145
|
+
"reasoning": true,
|
|
146
|
+
"notes": "`enable_thinking` via `qwen-chat-template`; effort via `reasoning_effort` (only `high`/`max` distinct, per vLLM GLM-5.2 recipe); thinking levels aligned with neuralwatt GLM 5.2; returns `reasoning` field",
|
|
147
|
+
"thinkingLevelMap": {
|
|
148
|
+
"minimal": null,
|
|
149
|
+
"low": null,
|
|
150
|
+
"medium": null,
|
|
151
|
+
"high": "high",
|
|
152
|
+
"xhigh": "max"
|
|
153
|
+
},
|
|
154
|
+
"compat": {
|
|
155
|
+
"thinkingFormat": "qwen-chat-template",
|
|
156
|
+
"supportsReasoningEffort": true
|
|
157
|
+
}
|
|
158
|
+
},
|
|
159
|
+
"zai-org/GLM-5.2-NVFP4": {
|
|
160
|
+
"reasoning": true,
|
|
161
|
+
"notes": "`enable_thinking` via `qwen-chat-template`; effort via `reasoning_effort` (only `high`/`max` distinct, per vLLM GLM-5.2 recipe); returns `reasoning` field",
|
|
162
|
+
"thinkingLevelMap": {
|
|
163
|
+
"minimal": null,
|
|
164
|
+
"low": null,
|
|
165
|
+
"medium": null,
|
|
166
|
+
"high": "high",
|
|
167
|
+
"xhigh": "max"
|
|
168
|
+
},
|
|
169
|
+
"compat": {
|
|
170
|
+
"thinkingFormat": "qwen-chat-template",
|
|
171
|
+
"supportsReasoningEffort": true
|
|
172
|
+
}
|
|
134
173
|
}
|
|
135
174
|
}
|