pi-makora-provider 1.0.0 → 1.2.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -4,7 +4,7 @@
4
4
 
5
5
  **Open-weight models through [Makora](https://inference.makora.com)**
6
6
 
7
- _DeepSeek V4, Kimi K2.6, GLM 5.1, Qwen 3.6 — with client-side tool call repair for [pi](https://github.com/earendil-works/pi-coding-agent)._
7
+ _DeepSeek V4, Kimi K2.7 Code, GLM 5.2, Qwen 3.6 — with client-side tool call repair and preserved-thinking flags for [pi](https://github.com/earendil-works/pi-coding-agent)._
8
8
 
9
9
  [![pi extension](https://img.shields.io/badge/pi-extension-blueviolet)](https://github.com/earendil-works/pi-coding-agent)
10
10
  [![license](https://img.shields.io/badge/license-MIT-blue)](./LICENSE)
@@ -19,21 +19,36 @@ _DeepSeek V4, Kimi K2.6, GLM 5.1, Qwen 3.6 — with client-side tool call repair
19
19
  | Model | ID | Reasoning | Notes |
20
20
  |-------|----|-----------|-------|
21
21
  | DeepSeek V4 Flash | `deepseek-ai/DeepSeek-V4-Flash` | Yes | `include_reasoning` + `chat_template_kwargs.thinking` via `before_provider_request` payload rewrite; returns `reasoning` field |
22
- | DeepSeek V4 Pro | `deepseek-ai/DeepSeek-V4-Pro` | Yes | `chat_template_kwargs.thinking` via `before_provider_request` payload rewrite; returns `reasoning_content` field |
23
- | GLM 5.1 FP8 | `zai-org/GLM-5.1-FP8` | Yes | `enable_thinking` via `qwen-chat-template`; returns `reasoning_content` field; client-side tool call parsing (vLLM streaming parser bypass) |
24
- | GPT-OSS 120B | `openai/gpt-oss-120b` | Yes | Reasoning always on |
25
- | Kimi K2.6 NVFP4 | `nvidia/Kimi-K2.6-NVFP4` | Yes | Reasoning on by default; client-side tool call parsing (vLLM streaming parser bypass) |
26
- | Kimi K2.7 Code | `moonshotai/Kimi-K2.7-Code` | Yes | Reasoning on by default; client-side tool call parsing (vLLM streaming parser bypass) |
22
+ | DeepSeek V4 Pro | `deepseek-ai/DeepSeek-V4-Pro` | Yes | `chat_template_kwargs.thinking` via `before_provider_request` payload rewrite; returns `reasoning` field |
23
+ | GLM 5.2 FP8 | `zai-org/GLM-5.2-FP8` | Yes | `enable_thinking` via `qwen-chat-template`; effort via `reasoning_effort` (only `high`/`max` distinct, per vLLM GLM-5.2 recipe); thinking levels aligned with neuralwatt GLM 5.2; returns `reasoning` field |
24
+ | GLM 5.2 NVFP4 | `zai-org/GLM-5.2-NVFP4` | Yes | `enable_thinking` via `qwen-chat-template`; effort via `reasoning_effort` (only `high`/`max` distinct, per vLLM GLM-5.2 recipe); returns `reasoning` field |
25
+ | Kimi K2.7 Code | `moonshotai/Kimi-K2.7-Code` | Yes | Reasoning on by default (thinking-only model); `preserve_thinking` via `chatTemplateKwargs` for multi-turn continuity; returns `reasoning` field; client-side tool call parsing (vLLM streaming parser bypass) |
27
26
  | Llama 3.3 70B FP8 | `amd/Llama-3.3-70B-Instruct-FP8-KV` | No | |
28
27
  | Llama 3.3 70B Instruct | `meta-llama/Llama-3.3-70B-Instruct` | No | |
29
- | MiniMax M3 MXFP8 | `MiniMaxAI/MiniMax-M3-MXFP8` | Yes | Reasoning via `chat_template_kwargs.enable_thinking`; returns `reasoning_content` field |
30
- | Qwen 3.6 27B NVFP4 | `unsloth/Qwen3.6-27B-NVFP4` | Yes | `enable_thinking` via `qwen-chat-template`; client-side tool call parsing (vLLM streaming parser bypass) |
31
- | Qwen 3.6 35B A3B NVFP4 | `unsloth/Qwen3.6-35B-A3B-NVFP4` | Yes | `enable_thinking` via `qwen-chat-template`; client-side tool call parsing (vLLM streaming parser bypass) |
28
+ | Qwen 3.6 27B NVFP4 | `unsloth/Qwen3.6-27B-NVFP4` | Yes | `enable_thinking` via `qwen-chat-template`; `preserve_thinking` via `chatTemplateKwargs` for multi-turn continuity; returns `reasoning` field; client-side tool call parsing (vLLM streaming parser bypass) |
29
+ | Qwen 3.6 35B A3B NVFP4 | `unsloth/Qwen3.6-35B-A3B-NVFP4` | Yes | `enable_thinking` via `qwen-chat-template`; `preserve_thinking` via `chatTemplateKwargs` for multi-turn continuity; returns `reasoning` field; client-side tool call parsing (vLLM streaming parser bypass) |
32
30
  <!-- MODELS_TABLE_END -->
33
31
 
34
32
  ## Installation
35
33
 
36
- ### Option 1: Using `pi install` (Recommended)
34
+ ### Option 1: Install from npm (Recommended)
35
+
36
+ ```bash
37
+ pi install npm:pi-makora-provider
38
+ ```
39
+
40
+ Then set your API key and run pi:
41
+ ```bash
42
+ # Recommended: add to auth.json
43
+ # See Authentication section below
44
+
45
+ # Or set as environment variable
46
+ export MAKORA_OPTIMIZE_TOKEN=your-api-key-here
47
+
48
+ pi
49
+ ```
50
+
51
+ ### Option 2: Using `pi install` from GitHub
37
52
 
38
53
  Install directly from GitHub:
39
54
 
@@ -52,7 +67,7 @@ export MAKORA_OPTIMIZE_TOKEN=your-api-key-here
52
67
  pi
53
68
  ```
54
69
 
55
- ### Option 2: Manual Clone
70
+ ### Option 3: Manual Clone
56
71
 
57
72
  1. Clone this repository:
58
73
  ```bash
@@ -146,17 +161,19 @@ Do **not** edit `models.json` directly — it is auto-generated from the API. To
146
161
 
147
162
  These issues are common to all vLLM-hosted providers and affect Makora models:
148
163
 
149
- - **GLM 5.1 tool calling**: vLLM's streaming tool call handling is broken for GLM — the model outputs Zhipu's native `<tool_call>` XML format as raw text. The `message_end` hook parses this into `toolCall` blocks so pi can execute the tools. A `context` hook then strips `tool_calls` from assistant messages before follow-up requests, converting them back to `<tool_call>` text to avoid a ZAI/vLLM server crash (500: `'str object' has no attribute 'items'`) that occurs when any assistant message contains a `tool_calls` field. If upstream fixes both the streaming parser and the 500 crash, the `message_end` hook gracefully skips (existing valid `toolCall` blocks are preserved), and the `context` hook's text-stripping is harmless (GLM natively understands `<tool_call>` format in conversation history).
164
+ - **GLM tool calls — follow-up `tool_calls` crash workaround**: the ZAI chat template calls `.items()` on any assistant message that carries a `tool_calls` field (on the `tool_calls` list or on its JSON-string `arguments` field), raising a leaked Python `AttributeError` that vLLM returns as HTTP 400 (`'list object' has no attribute 'items'` / `'str object' has no attribute 'items'`). The `before_provider_request` hook (`stripGlmToolCalls`) strips `tool_calls` from assistant messages and converts them to GLM-native `<tool_call>` XML text in `content` — which the model natively understands in conversation history — before the follow-up request is sent. The hook is scoped to Makora's own GLM models via an exact model-id match (`isMakoraGlmVllmModel`): `before_provider_request` is global in pi (it fires for every loaded provider's requests and the event carries no `provider` field), so gating by model id is what keeps it from touching GLM models served by other providers. If upstream fixes the `.items()` crash, this transform is a harmless no-op (the XML text is valid GLM input either way).
150
165
 
151
- - **Kimi K2.6 + Qwen 3.6 tool calling**: vLLM's streaming tool call handling is broken or missing for these models. The `before_provider_request` hook sets `tool_choice: "none"` and `skip_special_tokens: false` so the model's tool call tokens pass through as plain text. The `message_end` hook then re-parses into `toolCall` blocks:
166
+ - **Kimi K2.7 Code + Qwen 3.6 tool calling**: vLLM's server-side streaming tool-call handling is broken or missing for these models — Makora's vLLM is missing both `--enable-auto-tool-choice` and `--tool-call-parser` for them, so they emit tool calls as raw text tokens rather than structured `tool_call` deltas.
152
167
 
153
- - **Kimi K2.6**: Uses `<|tool_call_begin|>...<|tool_call_end|>` tokens. Makora's vLLM is missing both `--enable-auto-tool-choice` and `--tool-call-parser` for this model.
154
- - **Qwen 3.6**: Uses hermes-style `<function=...>` XML, sometimes with `█` delimiters. Same vLLM flag limitation as Kimi.
168
+ - **Kimi K2.7 Code**: Uses `<|tool_call_begin|>...<|tool_call_end|>` tokens.
169
+ - **Qwen 3.6**: Uses hermes-style `<function=...>` XML, sometimes with `█` delimiters.
155
170
 
156
- - **GLM 5.1 CoT leak**: On some vLLM builds, disabling reasoning may still leak chain-of-thought into `content` terminated by a ``` marker. See [vllm-project/vllm#31319](https://github.com/vllm-project/vllm/issues/31319).
171
+ - **GLM 5.2 CoT leak**: On some vLLM builds, disabling reasoning may still leak chain-of-thought into `content` terminated by a ``` marker. See [vllm-project/vllm#31319](https://github.com/vllm-project/vllm/issues/31319).
157
172
 
158
173
  - **DeepSeek V4 reasoning**: The official DeepSeek API uses `thinking: { type: "enabled" }` which Makora's vLLM silently ignores. The `before_provider_request` hook rewrites the payload to use vLLM-native params instead:
159
- - **DS V4 Pro**: `chat_template_kwargs: { thinking: true }`. Returns `reasoning_content`.
174
+ - **DS V4 Pro**: `chat_template_kwargs: { thinking: true }`. Returns `reasoning` (vLLM reasoning parser; the official DeepSeek API returns `reasoning_content`, but Makora's vLLM build returns `reasoning`).
160
175
  - **DS V4 Flash**: `include_reasoning: true` + `chat_template_kwargs: { thinking: true }`. `include_reasoning` alone returns `reasoning: null` on this vLLM build — both params are required. Returns `reasoning`.
161
- - **GLM 5.1 reasoning**: Returns `reasoning_content` (not `reasoning`). pi's OpenAI completions handler checks `reasoning_content` first, so this is handled correctly.
162
- - **MiniMax M3 reasoning**: Uses `chat_template_kwargs.enable_thinking` to toggle thinking (not `chat_template_kwargs.thinking` like DeepSeek). The `before_provider_request` hook rewrites the DeepSeek API-style `thinking` param into vLLM-native `chat_template_kwargs: { enable_thinking: true }`. Returns `reasoning_content` field.
176
+ - **Preserved thinking (multi-turn reasoning continuity)**: every reasoning model on Makora returns reasoning in the `reasoning` field (not `reasoning_content`). Pi sends prior assistant reasoning back on follow-up turns; Makora's vLLM chat templates accept it under either `reasoning` or `reasoning_content`. For models whose chat template defines an explicit preserved-thinking flag, the flag is sent via `compat.chatTemplateKwargs`:
177
+ - **Kimi K2.7 Code**: `preserve_thinking: true` (the model forces thinking-only mode with preserve_thinking on, per the vLLM recipe).
178
+ - **Qwen 3.6 27B / 35B**: `preserve_thinking: true` (Qwen 3.6 ships with preserve_thinking).
179
+ - **GLM 5.2**: no extra flag — the vLLM GLM-5.2 recipe uses only `enable_thinking` + `reasoning_effort`; preserved thinking works without a dedicated flag. Effort honors only `high` and `max` (lower pi levels resolve to the default `max`).
package/index.ts CHANGED
@@ -17,7 +17,7 @@
17
17
  * pi sends thinking: { type } via the "deepseek" thinkingFormat, but vLLM
18
18
  * ignores that — the before_provider_request hook rewrites the payload to
19
19
  * use chat_template_kwargs: { thinking: true } instead.
20
- * Returns reasoning_content field.
20
+ * Returns reasoning field (vLLM reasoning parser; not reasoning_content).
21
21
  * - DeepSeek V4 Flash: reasoning via include_reasoning +
22
22
  * chat_template_kwargs.thinking on vLLM.
23
23
  * The before_provider_request hook rewrites the payload to replace
@@ -34,11 +34,19 @@
34
34
  * delta. Setting zaiToolStream: true sends tool_stream: true in the
35
35
  * request, which forces vLLM to use the explicit tool streaming path
36
36
  * that correctly emits tool call chunks.
37
+ * - GLM 5.2 FP8 / NVFP4: reasoning via chat_template_kwargs.enable_thinking.
38
+ * Effort uses vLLM's reasoning_effort field (only `high` and `max` are
39
+ * distinct levels per the vLLM GLM-5.2 recipe; lower pi levels resolve to
40
+ * the default). Thinking levels aligned with the neuralwatt provider's
41
+ * GLM 5.2 configuration and mapped through pi's qwen-chat-template
42
+ * thinkingFormat. Returns `reasoning` field.
37
43
  * - GPT-OSS 120B: reasoning always on; returns `reasoning` field.
38
- * - Kimi K2.6 NVFP4 / Kimi K2.7 Code: reasoning always on by default;
39
- * returns `reasoning` field. Can be toggled via enable_thinking.
44
+ * - Kimi K2.7 Code: reasoning always on (thinking-only model);
45
+ * chatTemplateKwargs.preserve_thinking forces multi-turn reasoning
46
+ * continuity. Returns `reasoning` field. Can be toggled via enable_thinking.
40
47
  * - Qwen 3.6 models: reasoning via chat_template_kwargs.enable_thinking;
41
- * returns `reasoning` field.
48
+ * chatTemplateKwargs.preserve_thinking for multi-turn continuity.
49
+ * Returns `reasoning` field.
42
50
  * - MiniMax M3 MXFP8: reasoning via chat_template_kwargs.enable_thinking;
43
51
  * returns reasoning_content field.
44
52
  * - Llama 3.3 70B: not a reasoning model.
@@ -105,6 +113,10 @@ interface JsonModel {
105
113
  requiresToolResultName?: boolean;
106
114
  requiresAssistantAfterToolResult?: boolean;
107
115
  cacheControlFormat?: "anthropic";
116
+ /** Extra keys merged into vLLM `chat_template_kwargs` on every request.
117
+ * Used for preserved-thinking flags like `preserve_thinking` / `clear_thinking`
118
+ * that some chat templates require for multi-turn reasoning continuity. */
119
+ chatTemplateKwargs?: Record<string, unknown>;
108
120
  };
109
121
  }
110
122
 
@@ -129,6 +141,11 @@ interface PatchEntry {
129
141
 
130
142
  type PatchMap = Record<string, PatchEntry>;
131
143
 
144
+ /** Type guard: non-null, non-array object. */
145
+ export function isObject(value: unknown): value is Record<string, unknown> {
146
+ return value !== null && typeof value === "object" && !Array.isArray(value);
147
+ }
148
+
132
149
  // Patch Application
133
150
 
134
151
  function applyPatch(model: JsonModel, patch: PatchEntry): JsonModel {
@@ -214,6 +231,38 @@ const MINIMAX_M3_ID = "MiniMaxAI/MiniMax-M3-MXFP8";
214
231
  const DS_VLLM_MODELS = new Set([DS_PRO_ID, DS_FLASH_ID]);
215
232
  const ENABLE_THINKING_VLLM_MODELS = new Set([MINIMAX_M3_ID]);
216
233
 
234
+ /**
235
+ * Makora's GLM models, built from the same models list this extension
236
+ * registers. `before_provider_request` is GLOBAL in pi — it fires for every
237
+ * loaded provider's requests, not just makora's — and its event carries only
238
+ * `payload` (no `provider` field to gate on). So the GLM tool_calls strip MUST
239
+ * be scoped to makora's own GLM model ids. A loose `/glm/i` regex would also
240
+ * rewrite tool_calls for every sibling provider that serves any GLM model
241
+ * (baseten, io, lilac, neuralwatt, tensorix, fireworks, crofai, hypercharm,
242
+ * wafer, parasail, umans, opencode, ...), silently breaking their tool calling
243
+ * and clobbering the input their own before_provider_request handlers expect.
244
+ *
245
+ * Derived (not hardcoded) so it auto-tracks models.json refreshes — models.json
246
+ * is auto-generated by scripts/update-models.js. All makora models run on vLLM,
247
+ * so every GLM id here shares the ZAI chat-template `.items()` crash that
248
+ * stripGlmToolCalls works around.
249
+ */
250
+ const allMakoraModels = buildModels(
251
+ modelsData as JsonModel[],
252
+ customModelsData as JsonModel[],
253
+ patchData as PatchMap,
254
+ );
255
+ const GLM_VLLM_MODELS = new Set(
256
+ allMakoraModels.filter((m) => /glm/i.test(m.id)).map((m) => m.id),
257
+ );
258
+
259
+ /** Whether `model` is one of makora's GLM models — the only ids the
260
+ * before_provider_request hook should strip tool_calls for. Exported so the
261
+ * scoping can be unit-tested in isolation. */
262
+ export function isMakoraGlmVllmModel(model: string): boolean {
263
+ return GLM_VLLM_MODELS.has(model);
264
+ }
265
+
217
266
  /**
218
267
  * Intercept the request payload for models that need vLLM-specific thinking
219
268
  * param rewrites.
@@ -262,12 +311,82 @@ function rewriteVllmPayload(payload: Record<string, unknown>): Record<string, un
262
311
  return p;
263
312
  }
264
313
 
265
- export default function (pi: ExtensionAPI) {
266
- const embeddedModels = modelsData as JsonModel[];
267
- const customModels = customModelsData as JsonModel[];
268
- const patches = patchData as PatchMap;
314
+ /**
315
+ * GLM models on Makora's vLLM crash with a leaked Python AttributeError
316
+ * (`'list object' has no attribute 'items'` or `'str object' has no attribute 'items'`)
317
+ * when any assistant message in the request contains a `tool_calls` field.
318
+ * The ZAI/vLLM chat template calls `.items()` on the tool_calls list (or on the
319
+ * JSON-string `arguments` field), which raises AttributeError and leaks into the
320
+ * HTTP 400 response body.
321
+ *
322
+ * Fix: for makora's GLM models only (see GLM_VLLM_MODELS / isMakoraGlmVllmModel),
323
+ * strip `tool_calls` from assistant messages in the before_provider_request hook
324
+ * and convert them back to GLM's native
325
+ * `<tool_call>` XML text in `content`. The model natively understands this
326
+ * format in conversation history, and the `role: "tool"` result messages that
327
+ * follow are rendered fine by the chat template's tool-observation branch.
328
+ *
329
+ * If upstream fixes both the streaming parser and the 500/400 crash, this
330
+ * transform becomes a harmless no-op (the XML text is still valid GLM input).
331
+ */
269
332
 
270
- const models = buildModels(embeddedModels, customModels, patches);
333
+ export function toolCallToGlmXml(tc: Record<string, unknown>): string {
334
+ const fn = (isObject(tc.function) ? tc.function : {}) as Record<string, unknown>;
335
+ const name = typeof fn.name === "string" ? fn.name : "";
336
+ const argsStr = typeof fn.arguments === "string" ? fn.arguments : "{}";
337
+ let args: Record<string, unknown>;
338
+ try {
339
+ args = JSON.parse(argsStr) as Record<string, unknown>;
340
+ } catch {
341
+ args = {};
342
+ }
343
+ if (!isObject(args)) args = {};
344
+ const argLines = Object.entries(args).map(
345
+ ([key, value]) =>
346
+ `<arg_key>${key}</arg_key>\n<arg_value>${
347
+ typeof value === "string" ? value : JSON.stringify(value)
348
+ }</arg_value>`,
349
+ );
350
+ return `<tool_call>${name}\n${argLines.join("\n")}\n</tool_call>`;
351
+ }
352
+
353
+ export function stripGlmToolCalls(payload: Record<string, unknown>): Record<string, unknown> {
354
+ const messages = payload.messages;
355
+ if (!Array.isArray(messages)) return payload;
356
+
357
+ let modified = false;
358
+ const newMessages = messages.map((msg) => {
359
+ if (!isObject(msg)) return msg;
360
+ if (msg.role !== "assistant") return msg;
361
+
362
+ const toolCalls = msg.tool_calls;
363
+ if (!Array.isArray(toolCalls) || toolCalls.length === 0) return msg;
364
+
365
+ const xmlBlocks = toolCalls
366
+ .filter((tc): tc is Record<string, unknown> => isObject(tc))
367
+ .map(toolCallToGlmXml);
368
+ if (xmlBlocks.length === 0) return msg;
369
+
370
+ const toolCallText = xmlBlocks.join("\n");
371
+ // pi's openai-completions provider always serializes assistant content as a
372
+ // plain string, but guard against null/array content for robustness.
373
+ const existingContent = typeof msg.content === "string" ? msg.content : "";
374
+ const newContent = existingContent ? `${existingContent}\n${toolCallText}` : toolCallText;
375
+
376
+ modified = true;
377
+ const rest: Record<string, unknown> = {};
378
+ for (const [k, v] of Object.entries(msg)) {
379
+ if (k !== "tool_calls") rest[k] = v;
380
+ }
381
+ return { ...rest, content: newContent };
382
+ });
383
+
384
+ if (!modified) return payload;
385
+ return { ...payload, messages: newMessages };
386
+ }
387
+
388
+ export default function (pi: ExtensionAPI) {
389
+ const models = allMakoraModels;
271
390
 
272
391
  // apiKey resolution order: auth.json ("makora" key) → MAKORA_OPTIMIZE_TOKEN env var
273
392
  pi.registerProvider(PROVIDER_ID, {
@@ -281,7 +400,12 @@ export default function (pi: ExtensionAPI) {
281
400
  pi.on("before_provider_request", (event) => {
282
401
  const payload = event.payload as Record<string, unknown> | undefined;
283
402
  if (!payload || typeof payload.model !== "string") return;
284
- return rewriteVllmPayload(payload);
403
+
404
+ let result = rewriteVllmPayload(payload);
405
+ if (isMakoraGlmVllmModel(payload.model)) {
406
+ result = stripGlmToolCalls(result);
407
+ }
408
+ return result;
285
409
  });
286
410
  }
287
411
 
package/models.json CHANGED
@@ -62,27 +62,6 @@
62
62
  "maxTokensField": "max_completion_tokens"
63
63
  }
64
64
  },
65
- {
66
- "id": "MiniMaxAI/MiniMax-M3-MXFP8",
67
- "name": "MiniMax M3 MXFP8",
68
- "reasoning": false,
69
- "input": [
70
- "text"
71
- ],
72
- "cost": {
73
- "input": 0,
74
- "output": 0,
75
- "cacheRead": 0,
76
- "cacheWrite": 0
77
- },
78
- "contextWindow": 1048576,
79
- "maxTokens": 0,
80
- "compat": {
81
- "supportsDeveloperRole": false,
82
- "supportsStore": false,
83
- "maxTokensField": "max_completion_tokens"
84
- }
85
- },
86
65
  {
87
66
  "id": "moonshotai/Kimi-K2.7-Code",
88
67
  "name": "Kimi K2.7 Code",
@@ -105,8 +84,8 @@
105
84
  }
106
85
  },
107
86
  {
108
- "id": "nvidia/Kimi-K2.6-NVFP4",
109
- "name": "Kimi K2.6 NVFP4",
87
+ "id": "unsloth/Qwen3.6-27B-NVFP4",
88
+ "name": "Qwen 3.6 27B NVFP4",
110
89
  "reasoning": false,
111
90
  "input": [
112
91
  "text"
@@ -126,29 +105,8 @@
126
105
  }
127
106
  },
128
107
  {
129
- "id": "openai/gpt-oss-120b",
130
- "name": "GPT-OSS 120B",
131
- "reasoning": false,
132
- "input": [
133
- "text"
134
- ],
135
- "cost": {
136
- "input": 0,
137
- "output": 0,
138
- "cacheRead": 0,
139
- "cacheWrite": 0
140
- },
141
- "contextWindow": 131072,
142
- "maxTokens": 0,
143
- "compat": {
144
- "supportsDeveloperRole": false,
145
- "supportsStore": false,
146
- "maxTokensField": "max_completion_tokens"
147
- }
148
- },
149
- {
150
- "id": "unsloth/Qwen3.6-27B-NVFP4",
151
- "name": "Qwen 3.6 27B NVFP4",
108
+ "id": "unsloth/Qwen3.6-35B-A3B-NVFP4",
109
+ "name": "Qwen 3.6 35B A3B NVFP4",
152
110
  "reasoning": false,
153
111
  "input": [
154
112
  "text"
@@ -168,8 +126,8 @@
168
126
  }
169
127
  },
170
128
  {
171
- "id": "unsloth/Qwen3.6-35B-A3B-NVFP4",
172
- "name": "Qwen 3.6 35B A3B NVFP4",
129
+ "id": "zai-org/GLM-5.2-FP8",
130
+ "name": "GLM 5.2 FP8",
173
131
  "reasoning": false,
174
132
  "input": [
175
133
  "text"
@@ -180,7 +138,7 @@
180
138
  "cacheRead": 0,
181
139
  "cacheWrite": 0
182
140
  },
183
- "contextWindow": 262144,
141
+ "contextWindow": 1048576,
184
142
  "maxTokens": 0,
185
143
  "compat": {
186
144
  "supportsDeveloperRole": false,
@@ -189,8 +147,8 @@
189
147
  }
190
148
  },
191
149
  {
192
- "id": "zai-org/GLM-5.1-FP8",
193
- "name": "GLM 5.1 FP8",
150
+ "id": "zai-org/GLM-5.2-NVFP4",
151
+ "name": "GLM 5.2 NVFP4",
194
152
  "reasoning": false,
195
153
  "input": [
196
154
  "text"
@@ -201,7 +159,7 @@
201
159
  "cacheRead": 0,
202
160
  "cacheWrite": 0
203
161
  },
204
- "contextWindow": 202752,
162
+ "contextWindow": 1048576,
205
163
  "maxTokens": 0,
206
164
  "compat": {
207
165
  "supportsDeveloperRole": false,
package/package.json CHANGED
@@ -1,15 +1,9 @@
1
1
  {
2
2
  "name": "pi-makora-provider",
3
- "version": "1.0.0",
4
- "description": "Makora provider extension for pi - Access DeepSeek V4, GLM 5.1, Kimi K2.6, Llama 3.3, Qwen 3.6, and more through the Makora inference API",
3
+ "version": "1.2.0",
4
+ "description": "Makora provider extension for pi - Access DeepSeek V4, GLM 5.2, Kimi K2.7 Code, Llama 3.3, Qwen 3.6, and more through the Makora inference API",
5
5
  "type": "module",
6
6
  "main": "index.ts",
7
- "scripts": {
8
- "clean": "echo 'nothing to clean'",
9
- "build": "echo 'nothing to build'",
10
- "check": "echo 'nothing to check'",
11
- "update-models": "node scripts/update-models.js"
12
- },
13
7
  "keywords": [
14
8
  "pi",
15
9
  "extension",
@@ -37,5 +31,16 @@
37
31
  "extensions": [
38
32
  "./index.ts"
39
33
  ]
34
+ },
35
+ "devDependencies": {
36
+ "vitest": "^4.1.9"
37
+ },
38
+ "scripts": {
39
+ "clean": "echo 'nothing to clean'",
40
+ "build": "echo 'nothing to build'",
41
+ "check": "echo 'nothing to check'",
42
+ "test": "vitest run",
43
+ "test:watch": "vitest",
44
+ "update-models": "node scripts/update-models.js"
40
45
  }
41
- }
46
+ }
package/patch.json CHANGED
@@ -16,7 +16,7 @@
16
16
  },
17
17
  "deepseek-ai/DeepSeek-V4-Pro": {
18
18
  "reasoning": true,
19
- "notes": "`chat_template_kwargs.thinking` via `before_provider_request` payload rewrite; returns `reasoning_content` field",
19
+ "notes": "`chat_template_kwargs.thinking` via `before_provider_request` payload rewrite; returns `reasoning` field",
20
20
  "thinkingLevelMap": {
21
21
  "minimal": null,
22
22
  "low": null,
@@ -60,26 +60,32 @@
60
60
  },
61
61
  "unsloth/Qwen3.6-27B-NVFP4": {
62
62
  "reasoning": true,
63
- "notes": "`enable_thinking` via `qwen-chat-template`; client-side tool call parsing (vLLM streaming parser bypass)",
63
+ "notes": "`enable_thinking` via `qwen-chat-template`; `preserve_thinking` via `chatTemplateKwargs` for multi-turn continuity; returns `reasoning` field; client-side tool call parsing (vLLM streaming parser bypass)",
64
64
  "thinkingLevelMap": {
65
65
  "minimal": "low",
66
66
  "xhigh": "high"
67
67
  },
68
68
  "compat": {
69
69
  "thinkingFormat": "qwen-chat-template",
70
- "supportsReasoningEffort": true
70
+ "supportsReasoningEffort": true,
71
+ "chatTemplateKwargs": {
72
+ "preserve_thinking": true
73
+ }
71
74
  }
72
75
  },
73
76
  "unsloth/Qwen3.6-35B-A3B-NVFP4": {
74
77
  "reasoning": true,
75
- "notes": "`enable_thinking` via `qwen-chat-template`; client-side tool call parsing (vLLM streaming parser bypass)",
78
+ "notes": "`enable_thinking` via `qwen-chat-template`; `preserve_thinking` via `chatTemplateKwargs` for multi-turn continuity; returns `reasoning` field; client-side tool call parsing (vLLM streaming parser bypass)",
76
79
  "thinkingLevelMap": {
77
80
  "minimal": "low",
78
81
  "xhigh": "high"
79
82
  },
80
83
  "compat": {
81
84
  "thinkingFormat": "qwen-chat-template",
82
- "supportsReasoningEffort": true
85
+ "supportsReasoningEffort": true,
86
+ "chatTemplateKwargs": {
87
+ "preserve_thinking": true
88
+ }
83
89
  }
84
90
  },
85
91
  "MiniMaxAI/MiniMax-M3-MXFP8": {
@@ -108,14 +114,17 @@
108
114
  "text",
109
115
  "image"
110
116
  ],
111
- "notes": "Reasoning on by default; client-side tool call parsing (vLLM streaming parser bypass)",
117
+ "notes": "Reasoning on by default (thinking-only model); `preserve_thinking` via `chatTemplateKwargs` for multi-turn continuity; returns `reasoning` field; client-side tool call parsing (vLLM streaming parser bypass)",
112
118
  "thinkingLevelMap": {
113
119
  "minimal": "low",
114
120
  "xhigh": "high"
115
121
  },
116
122
  "compat": {
117
123
  "thinkingFormat": "qwen-chat-template",
118
- "supportsReasoningEffort": true
124
+ "supportsReasoningEffort": true,
125
+ "chatTemplateKwargs": {
126
+ "preserve_thinking": true
127
+ }
119
128
  }
120
129
  },
121
130
  "zai-org/GLM-5.1-FP8": {
@@ -131,5 +140,35 @@
131
140
  "supportsReasoningEffort": true,
132
141
  "zaiToolStream": true
133
142
  }
143
+ },
144
+ "zai-org/GLM-5.2-FP8": {
145
+ "reasoning": true,
146
+ "notes": "`enable_thinking` via `qwen-chat-template`; effort via `reasoning_effort` (only `high`/`max` distinct, per vLLM GLM-5.2 recipe); thinking levels aligned with neuralwatt GLM 5.2; returns `reasoning` field",
147
+ "thinkingLevelMap": {
148
+ "minimal": null,
149
+ "low": null,
150
+ "medium": null,
151
+ "high": "high",
152
+ "xhigh": "max"
153
+ },
154
+ "compat": {
155
+ "thinkingFormat": "qwen-chat-template",
156
+ "supportsReasoningEffort": true
157
+ }
158
+ },
159
+ "zai-org/GLM-5.2-NVFP4": {
160
+ "reasoning": true,
161
+ "notes": "`enable_thinking` via `qwen-chat-template`; effort via `reasoning_effort` (only `high`/`max` distinct, per vLLM GLM-5.2 recipe); returns `reasoning` field",
162
+ "thinkingLevelMap": {
163
+ "minimal": null,
164
+ "low": null,
165
+ "medium": null,
166
+ "high": "high",
167
+ "xhigh": "max"
168
+ },
169
+ "compat": {
170
+ "thinkingFormat": "qwen-chat-template",
171
+ "supportsReasoningEffort": true
172
+ }
134
173
  }
135
174
  }