pi-makora-provider 1.3.0 → 1.5.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -4,7 +4,7 @@
4
4
 
5
5
  **Open-weight models through [Makora](https://inference.makora.com)**
6
6
 
7
- _DeepSeek V4, Kimi K2.7 Code, GLM 5.2, Qwen 3.6 — with client-side tool call repair and preserved-thinking flags for [pi](https://github.com/earendil-works/pi-coding-agent)._
7
+ _DeepSeek V4, Kimi K2.7 Code, GLM 5.2, Qwen 3.6 for [pi](https://github.com/earendil-works/pi-coding-agent)._
8
8
 
9
9
  [![pi extension](https://img.shields.io/badge/pi-extension-blueviolet)](https://github.com/earendil-works/pi-coding-agent)
10
10
  [![license](https://img.shields.io/badge/license-MIT-blue)](./LICENSE)
@@ -18,15 +18,16 @@ _DeepSeek V4, Kimi K2.7 Code, GLM 5.2, Qwen 3.6 — with client-side tool call r
18
18
  <!-- MODELS_TABLE_START -->
19
19
  | Model | ID | Reasoning | Notes |
20
20
  |-------|----|-----------|-------|
21
- | DeepSeek V4 Flash | `deepseek-ai/DeepSeek-V4-Flash` | Yes | `include_reasoning` + `chat_template_kwargs.thinking` via `before_provider_request` payload rewrite; returns `reasoning` field |
22
- | DeepSeek V4 Pro | `deepseek-ai/DeepSeek-V4-Pro` | Yes | `chat_template_kwargs.thinking` via `before_provider_request` payload rewrite; returns `reasoning` field |
21
+ | DeepSeek V4 Flash | `deepseek-ai/DeepSeek-V4-Flash` | Yes | returns `reasoning` field |
22
+ | DeepSeek V4 Pro | `deepseek-ai/DeepSeek-V4-Pro` | Yes | returns `reasoning` field |
23
+ | Gemma 4 26B A4B | `google/gemma-4-26B-A4B` | No | |
23
24
  | GLM 5.2 FP8 | `zai-org/GLM-5.2-FP8` | Yes | `enable_thinking` via `qwen-chat-template`; effort via `reasoning_effort` (only `high`/`max` distinct, per vLLM GLM-5.2 recipe); thinking levels aligned with neuralwatt GLM 5.2; returns `reasoning` field |
24
25
  | GLM 5.2 NVFP4 | `zai-org/GLM-5.2-NVFP4` | Yes | `enable_thinking` via `qwen-chat-template`; effort via `reasoning_effort` (only `high`/`max` distinct, per vLLM GLM-5.2 recipe); returns `reasoning` field |
25
- | Kimi K2.7 Code | `moonshotai/Kimi-K2.7-Code` | Yes | Reasoning on by default (thinking-only model); `preserve_thinking` via `chatTemplateKwargs` for multi-turn continuity; returns `reasoning` field; client-side tool call parsing (vLLM streaming parser bypass) |
26
+ | Kimi K2.7 Code | `moonshotai/Kimi-K2.7-Code` | Yes | Reasoning on by default (thinking-only model); `preserve_thinking` via `chatTemplateKwargs` for multi-turn continuity; returns `reasoning` field |
26
27
  | Llama 3.3 70B FP8 | `amd/Llama-3.3-70B-Instruct-FP8-KV` | No | |
27
28
  | Llama 3.3 70B Instruct | `meta-llama/Llama-3.3-70B-Instruct` | No | |
28
- | Qwen 3.6 27B NVFP4 | `unsloth/Qwen3.6-27B-NVFP4` | Yes | `enable_thinking` via `qwen-chat-template`; `preserve_thinking` via `chatTemplateKwargs` for multi-turn continuity; returns `reasoning` field; client-side tool call parsing (vLLM streaming parser bypass) |
29
- | Qwen 3.6 35B A3B NVFP4 | `unsloth/Qwen3.6-35B-A3B-NVFP4` | Yes | `enable_thinking` via `qwen-chat-template`; `preserve_thinking` via `chatTemplateKwargs` for multi-turn continuity; returns `reasoning` field; client-side tool call parsing (vLLM streaming parser bypass) |
29
+ | Qwen 3.6 27B NVFP4 | `unsloth/Qwen3.6-27B-NVFP4` | Yes | `enable_thinking` via `qwen-chat-template`; `preserve_thinking` via `chatTemplateKwargs` for multi-turn continuity; returns `reasoning` field |
30
+ | Qwen 3.6 35B A3B NVFP4 | `unsloth/Qwen3.6-35B-A3B-NVFP4` | Yes | `enable_thinking` via `qwen-chat-template`; `preserve_thinking` via `chatTemplateKwargs` for multi-turn continuity; returns `reasoning` field |
30
31
  <!-- MODELS_TABLE_END -->
31
32
 
32
33
  ## Installation
@@ -157,23 +158,32 @@ Do **not** edit `models.json` directly — it is auto-generated from the API. To
157
158
  - All models are hosted on vLLM
158
159
  - The `developer` role is not supported (prompts are silently dropped); `supportsDeveloperRole` is set to `false` for all models
159
160
 
160
- ## vLLM Caveats
161
-
162
- These issues are common to all vLLM-hosted providers and affect Makora models:
163
-
164
- - **GLM tool calls — follow-up `tool_calls` crash workaround**: the ZAI chat template calls `.items()` on any assistant message that carries a `tool_calls` field (on the `tool_calls` list or on its JSON-string `arguments` field), raising a leaked Python `AttributeError` that vLLM returns as HTTP 400 (`'list object' has no attribute 'items'` / `'str object' has no attribute 'items'`). The `before_provider_request` hook (`stripGlmToolCalls`) strips `tool_calls` from assistant messages and converts them to GLM-native `<tool_call>` XML text in `content` — which the model natively understands in conversation history — before the follow-up request is sent. The hook is scoped to Makora's own GLM models via an exact model-id match (`isMakoraGlmVllmModel`): `before_provider_request` is global in pi (it fires for every loaded provider's requests and the event carries no `provider` field), so gating by model id is what keeps it from touching GLM models served by other providers. If upstream fixes the `.items()` crash, this transform is a harmless no-op (the XML text is valid GLM input either way).
165
-
166
- - **Kimi K2.7 Code + Qwen 3.6 tool calling**: vLLM's server-side streaming tool-call handling is broken or missing for these models — Makora's vLLM is missing both `--enable-auto-tool-choice` and `--tool-call-parser` for them, so they emit tool calls as raw text tokens rather than structured `tool_call` deltas.
167
-
168
- - **Kimi K2.7 Code**: Uses `<|tool_call_begin|>...<|tool_call_end|>` tokens.
169
- - **Qwen 3.6**: Uses hermes-style `<function=...>` XML, sometimes with `█` delimiters.
170
-
171
- - **GLM 5.2 CoT leak**: On some vLLM builds, disabling reasoning may still leak chain-of-thought into `content` terminated by a ``` marker. See [vllm-project/vllm#31319](https://github.com/vllm-project/vllm/issues/31319).
172
-
173
- - **DeepSeek V4 reasoning**: The official DeepSeek API uses `thinking: { type: "enabled" }` which Makora's vLLM silently ignores. The `before_provider_request` hook rewrites the payload to use vLLM-native params instead:
174
- - **DS V4 Pro**: `chat_template_kwargs: { thinking: true }`. Returns `reasoning` (vLLM reasoning parser; the official DeepSeek API returns `reasoning_content`, but Makora's vLLM build returns `reasoning`).
175
- - **DS V4 Flash**: `include_reasoning: true` + `chat_template_kwargs: { thinking: true }`. `include_reasoning` alone returns `reasoning: null` on this vLLM build — both params are required. Returns `reasoning`.
176
- - **Preserved thinking (multi-turn reasoning continuity)**: every reasoning model on Makora returns reasoning in the `reasoning` field (not `reasoning_content`). Pi sends prior assistant reasoning back on follow-up turns; Makora's vLLM chat templates accept it under either `reasoning` or `reasoning_content`. For models whose chat template defines an explicit preserved-thinking flag, the flag is sent via `compat.chatTemplateKwargs`:
177
- - **Kimi K2.7 Code**: `preserve_thinking: true` (the model forces thinking-only mode with preserve_thinking on, per the vLLM recipe).
178
- - **Qwen 3.6 27B / 35B**: `preserve_thinking: true` (Qwen 3.6 ships with preserve_thinking).
179
- - **GLM 5.2**: no extra flag — the vLLM GLM-5.2 recipe uses only `enable_thinking` + `reasoning_effort`; preserved thinking works without a dedicated flag. Effort honors only `high` and `max` (lower pi levels resolve to the default `max`).
161
+ ## Death-Loop Guard
162
+
163
+ GLM 5.2 (NVFP4 / FP8) occasionally degenerates into an unbroken `!` repetition
164
+ loop (`!!!!...`) that eats the whole response. This extension ships a guard
165
+ that watches the streamed assistant output (both the visible answer and the
166
+ reasoning trace) and, when it detects a long run of `!` characters, **aborts
167
+ the runaway generation and resumes the agentic loop invisibly** — no new user
168
+ message is injected, using the
169
+ [pi-invisible-continue](https://github.com/monotykamary/pi-invisible-continue)
170
+ pattern (`agent.prompt([])`).
171
+
172
+ On abort, pi finalizes the in-flight assistant message **with its accumulated
173
+ `!!!` content** into the transcript, so the guard drops that partial message
174
+ before resuming — otherwise the model would just re-read its own `!!!` and loop
175
+ again. The recovery is bounded per prompt (default 3 attempts) to avoid
176
+ abort/continue thrash.
177
+
178
+ The guard is scoped to the Makora GLM 5.2 family by default and is tunable via
179
+ constants at the top of [`death-loop-guard.ts`](./death-loop-guard.ts):
180
+
181
+ | Constant | Default | Meaning |
182
+ |---|---|---|
183
+ | `GUARDED_MODEL_IDS` | `zai-org/GLM-5.2-NVFP4`, `zai-org/GLM-5.2-FP8` | Which model IDs to guard. Add `'*'` to guard every Makora model. |
184
+ | `BANG_THRESHOLD` | `40` | Consecutive `!` characters that trip the guard. 40 is far above anything normal prose/code produces. |
185
+ | `MAX_RECOVERS_PER_RUN` | `3` | Max invisible recoveries per user prompt before the guard stops intervening. |
186
+
187
+ It is only active for the `makora` provider, so it never interferes when you
188
+ switch to another provider. If you also run `pi-invisible-continue`, the two
189
+ coexist — both chain the `Agent.prototype.subscribe` patch.
@@ -0,0 +1,238 @@
1
+ /**
2
+ * Death-loop guard for Makora reasoning models.
3
+ *
4
+ * Some Makora models (notably GLM 5.2 NVFP4 / FP8) occasionally fall into a
5
+ * degenerate repetition loop, emitting an unbroken run of '!' characters
6
+ * (e.g. "!!!!...") that consumes the whole response. This guard watches the
7
+ * streamed assistant output (both the visible answer and the reasoning
8
+ * trace); when that run is detected it aborts the runaway generation, drops
9
+ * the partial (toxic) assistant message from the transcript, and resumes the
10
+ * agentic loop invisibly via agent.prompt([]) — the same pattern
11
+ * pi-invisible-continue uses, so no new user message pollutes the context.
12
+ *
13
+ * Why trim the aborted message: on abort, pi finalizes the in-flight
14
+ * assistant message WITH its accumulated '!!!' content and stopReason
15
+ * "aborted" into the transcript. Resuming from that context would re-feed
16
+ * the toxic text to the model and likely re-trigger the loop. Dropping the
17
+ * last (aborted) assistant message leaves the context ending at the prior
18
+ * user/toolResult message — a clean continuation point.
19
+ *
20
+ * Module resolution: @earendil-works/pi-agent-core is a devDependency only
21
+ * (types + test resolution). At runtime pi's extension loader aliases that
22
+ * specifier to its bundled copy, so the Agent class patched below is the
23
+ * SAME class AgentSession uses. A static import is required — jiti's alias
24
+ * applies to static imports (which it rewrites to its own resolver) but not
25
+ * to native dynamic import() calls.
26
+ */
27
+
28
+ import type { ExtensionAPI } from "@earendil-works/pi-coding-agent";
29
+ import { Agent } from "@earendil-works/pi-agent-core";
30
+
31
+ const PROVIDER_ID = "makora";
32
+
33
+ /** Makora model IDs to guard. Add ids to widen coverage, or include "*"
34
+ * to guard every Makora model. Defaults to the GLM 5.2 family (the known
35
+ * offender). Kept exported so tests and downstream forks can introspect. */
36
+ export const GUARDED_MODEL_IDS = new Set<string>([
37
+ "zai-org/GLM-5.2-NVFP4",
38
+ "zai-org/GLM-5.2-FP8",
39
+ ]);
40
+
41
+ /** Trip after this many consecutive '!' characters in the streamed answer.
42
+ * 40 is far above anything normal prose or code produces. */
43
+ export const BANG_THRESHOLD = 40;
44
+
45
+ /** Max invisible recoveries per user prompt, to bound abort/continue thrash. */
46
+ export const MAX_RECOVERS_PER_RUN = 3;
47
+
48
+ export interface GuardedMessage {
49
+ role: string;
50
+ stopReason?: string;
51
+ content?: unknown;
52
+ }
53
+
54
+ export interface GuardedAgent {
55
+ abort(): void;
56
+ waitForIdle(): Promise<void>;
57
+ prompt(input: unknown[] | string): Promise<void>;
58
+ state: { messages: GuardedMessage[] };
59
+ }
60
+
61
+ let _agent: GuardedAgent | null = null;
62
+
63
+ /** Mutex held only across the abort+trim critical section — NOT across the
64
+ * resumed prompt([]) run, so the resumed stream stays monitored. */
65
+ let _recovering = false;
66
+
67
+ /** Trailing '!' run length in the current text block. */
68
+ let _trailingBangs = 0;
69
+
70
+ /** Latch: already tripped for the current assistant message. */
71
+ let _tripped = false;
72
+
73
+ /** Recoveries performed in the current user-initiated agent run. */
74
+ let _recoversThisRun = 0;
75
+
76
+ export function isGuardedModel(
77
+ model: { provider?: string; id?: string } | undefined | null,
78
+ ): boolean {
79
+ if (!model) return false;
80
+ if (model.provider !== PROVIDER_ID) return false;
81
+ if (GUARDED_MODEL_IDS.has("*")) return true;
82
+ return model.id != null && GUARDED_MODEL_IDS.has(model.id);
83
+ }
84
+
85
+ /** Update the trailing-'!' run length given a new text delta. O(len(delta)).
86
+ * Trailing run depends only on the delta's suffix: if the delta contains any
87
+ * non-'!' char, the prior run is cut off at that char; if the delta is all
88
+ * '!', it extends the prior run. */
89
+ export function nextTrailingBangs(prev: number, delta: string): number {
90
+ const len = delta.length;
91
+ if (len === 0) return prev;
92
+ let i = len - 1;
93
+ while (i >= 0 && delta.charCodeAt(i) === 0x21) i--; // '!' === 0x21
94
+ const trailingInDelta = len - 1 - i;
95
+ return i >= 0 ? trailingInDelta : prev + len;
96
+ }
97
+
98
+ export function extractText(content: unknown): string {
99
+ if (typeof content === "string") return content;
100
+ if (Array.isArray(content)) {
101
+ let out = "";
102
+ for (const block of content) {
103
+ if (
104
+ block &&
105
+ typeof block === "object" &&
106
+ (block as { type?: string }).type === "text" &&
107
+ typeof (block as { text?: unknown }).text === "string"
108
+ ) {
109
+ out += (block as { text: string }).text;
110
+ }
111
+ }
112
+ return out;
113
+ }
114
+ return "";
115
+ }
116
+
117
+ /** Trailing '!' run length of a finalized message's text content. */
118
+ export function messageTrailingBangs(msg: GuardedMessage): number {
119
+ const text = extractText(msg.content);
120
+ let i = text.length - 1;
121
+ while (i >= 0 && text.charCodeAt(i) === 0x21) i--;
122
+ return text.length - 1 - i;
123
+ }
124
+
125
+ async function recover(agent: GuardedAgent): Promise<void> {
126
+ _recovering = true;
127
+ try {
128
+ agent.abort();
129
+ await agent.waitForIdle();
130
+ const msgs = agent.state.messages;
131
+ const last = msgs[msgs.length - 1];
132
+ if (
133
+ last &&
134
+ last.role === "assistant" &&
135
+ (last.stopReason === "aborted" || messageTrailingBangs(last) >= BANG_THRESHOLD)
136
+ ) {
137
+ // Drop the toxic partial so it isn't re-fed to the model on resume.
138
+ // The state setter copies the array, leaving a clean transcript that
139
+ // ends at the prior user/toolResult message.
140
+ agent.state.messages = msgs.slice(0, -1);
141
+ }
142
+ } finally {
143
+ // Release before resuming so the resumed stream stays monitored.
144
+ _recovering = false;
145
+ }
146
+ try {
147
+ // Invisible continue: fresh agent loop, no new message injected.
148
+ await agent.prompt([]);
149
+ } catch {
150
+ // "Agent is already processing" or other transient error — best effort.
151
+ }
152
+ }
153
+
154
+ export function registerDeathLoopGuard(pi: ExtensionAPI): void {
155
+ // Capture the live Agent instance by chaining Agent.prototype.subscribe.
156
+ // subscribe() fires when AgentSession attaches — on every fresh session
157
+ // and every resume — so _agent always points at the active Agent. Chain
158
+ // any prior patch (e.g. pi-invisible-continue) so both extensions coexist.
159
+ const proto = Agent.prototype as unknown as {
160
+ subscribe: (this: GuardedAgent, ...args: unknown[]) => unknown;
161
+ };
162
+ const origSubscribe = proto.subscribe;
163
+ proto.subscribe = function (this: GuardedAgent, ...args: unknown[]) {
164
+ _agent = this;
165
+ return origSubscribe.apply(this, args);
166
+ };
167
+
168
+ pi.on("session_start", () => {
169
+ _recovering = false;
170
+ _trailingBangs = 0;
171
+ _tripped = false;
172
+ _recoversThisRun = 0;
173
+ });
174
+
175
+ // before_agent_start fires only for user prompts (the AgentSession path),
176
+ // not for the recovery's direct agent.prompt([]), so the recovery cap is
177
+ // bounded per user prompt instead of reset on every recovery continuation.
178
+ pi.on("before_agent_start", () => {
179
+ _recovering = false;
180
+ _trailingBangs = 0;
181
+ _tripped = false;
182
+ _recoversThisRun = 0;
183
+ });
184
+
185
+ pi.on("message_start", (event) => {
186
+ if (event.message?.role === "assistant") {
187
+ _trailingBangs = 0;
188
+ _tripped = false;
189
+ }
190
+ });
191
+
192
+ pi.on("message_update", (event, ctx) => {
193
+ if (_recovering || _tripped) return;
194
+ const ame = event.assistantMessageEvent;
195
+ if (ame.type === "text_start" || ame.type === "thinking_start") {
196
+ // New content block (answer or reasoning) — trailing run starts fresh.
197
+ _trailingBangs = 0;
198
+ return;
199
+ }
200
+ // Watch both the visible answer and the reasoning trace: GLM 5.2 is a
201
+ // reasoning model, and the loop can surface in either.
202
+ if (ame.type !== "text_delta" && ame.type !== "thinking_delta") return;
203
+ if (!isGuardedModel(ctx.model)) return;
204
+
205
+ _trailingBangs = nextTrailingBangs(_trailingBangs, ame.delta);
206
+ if (_trailingBangs < BANG_THRESHOLD) return;
207
+
208
+ if (_recoversThisRun >= MAX_RECOVERS_PER_RUN) {
209
+ _tripped = true;
210
+ ctx.ui.notify(
211
+ "Makora death-loop guard: runaway '!' output detected, but the " +
212
+ "recovery limit for this prompt was reached — stopping " +
213
+ "intervention. Try /continue or rephrase.",
214
+ "warning",
215
+ );
216
+ return;
217
+ }
218
+ _tripped = true;
219
+ _recoversThisRun++;
220
+ const agent = _agent;
221
+ if (!agent) {
222
+ ctx.ui.notify(
223
+ "Makora death-loop guard: runaway '!' output detected but the Agent " +
224
+ "instance was not captured; cannot recover automatically.",
225
+ "warning",
226
+ );
227
+ return;
228
+ }
229
+ ctx.ui.notify(
230
+ `Makora death-loop guard: aborting runaway '!' output and resuming ` +
231
+ `(${_recoversThisRun}/${MAX_RECOVERS_PER_RUN}).`,
232
+ "warning",
233
+ );
234
+ // Detach: the handler must return so the run can unwind to idle before
235
+ // recover() awaits waitForIdle() and calls prompt([]).
236
+ void recover(agent);
237
+ });
238
+ }
package/index.ts CHANGED
@@ -13,32 +13,19 @@
13
13
  * Model resolution strategy: static models.json merged with custom-models.json
14
14
  *
15
15
  * Reasoning notes:
16
- * - DeepSeek V4 Pro: reasoning via chat_template_kwargs.thinking on vLLM.
17
- * pi sends thinking: { type } via the "deepseek" thinkingFormat, but vLLM
18
- * ignores that — the before_provider_request hook rewrites the payload to
19
- * use chat_template_kwargs: { thinking: true } instead.
20
- * Returns reasoning field (vLLM reasoning parser; not reasoning_content).
21
- * - DeepSeek V4 Flash: reasoning via include_reasoning +
22
- * chat_template_kwargs.thinking on vLLM.
23
- * The before_provider_request hook rewrites the payload to replace
24
- * thinking: { type } with include_reasoning: true +
25
- * chat_template_kwargs: { thinking: true }.
26
- * include_reasoning alone returns reasoning: null on this vLLM build.
27
- * Returns reasoning field.
28
- * - GLM 5.2 FP8 / NVFP4: reasoning via chat_template_kwargs.enable_thinking.
29
- * Effort uses vLLM's reasoning_effort field (only `high` and `max` are
30
- * distinct levels per the vLLM GLM-5.2 recipe; lower pi levels resolve to
31
- * the default). Thinking levels aligned with the neuralwatt provider's
32
- * GLM 5.2 configuration and mapped through pi's qwen-chat-template
33
- * thinkingFormat. Returns `reasoning` field.
34
- * - Kimi K2.7 Code: reasoning always on (thinking-only model);
35
- * chatTemplateKwargs.preserve_thinking forces multi-turn reasoning
36
- * continuity. Returns `reasoning` field. Can be toggled via enable_thinking.
37
- * - Qwen 3.6 models: reasoning via chat_template_kwargs.enable_thinking;
38
- * chatTemplateKwargs.preserve_thinking for multi-turn continuity.
39
- * Returns `reasoning` field.
16
+ * - DeepSeek V4 Pro: returns `reasoning` field.
17
+ * - DeepSeek V4 Flash: returns `reasoning` field.
18
+ * - GLM 5.2 FP8 / NVFP4: returns `reasoning` field.
19
+ * - Kimi K2.7 Code: returns `reasoning` field.
20
+ * - Qwen 3.6 models: returns `reasoning` field.
40
21
  * - Llama 3.3 70B: not a reasoning model.
41
22
  *
23
+ * A death-loop guard (see ./death-loop-guard.ts) is registered alongside the
24
+ * provider. It watches the assistant text stream on the GLM 5.2 family and,
25
+ * if the model falls into an unbroken '!' repetition loop, aborts the runaway
26
+ * generation and resumes the agentic loop invisibly via agent.prompt([]) (the
27
+ * pi-invisible-continue pattern) — no new user message is injected.
28
+ *
42
29
  * Developer role is NOT supported by any of the chat templates on Makora's
43
30
  * vLLM deployment (prompts with role: "developer" are silently dropped).
44
31
  * supportsDeveloperRole is set to false for all models.
@@ -61,6 +48,7 @@ import type { ExtensionAPI } from "@earendil-works/pi-coding-agent";
61
48
  import modelsData from "./models.json" with { type: "json" };
62
49
  import customModelsData from "./custom-models.json" with { type: "json" };
63
50
  import patchData from "./patch.json" with { type: "json" };
51
+ import { registerDeathLoopGuard } from "./death-loop-guard.js";
64
52
 
65
53
  // Types
66
54
 
@@ -129,11 +117,6 @@ interface PatchEntry {
129
117
 
130
118
  type PatchMap = Record<string, PatchEntry>;
131
119
 
132
- /** Type guard: non-null, non-array object. */
133
- export function isObject(value: unknown): value is Record<string, unknown> {
134
- return value !== null && typeof value === "object" && !Array.isArray(value);
135
- }
136
-
137
120
  // Patch Application
138
121
 
139
122
  function applyPatch(model: JsonModel, patch: PatchEntry): JsonModel {
@@ -212,157 +195,11 @@ function buildModels(
212
195
  const PROVIDER_ID = "makora";
213
196
  const BASE_URL = "https://inference.makora.com/v1";
214
197
 
215
- const DS_PRO_ID = "deepseek-ai/DeepSeek-V4-Pro";
216
- const DS_FLASH_ID = "deepseek-ai/DeepSeek-V4-Flash";
217
-
218
- const DS_VLLM_MODELS = new Set([DS_PRO_ID, DS_FLASH_ID]);
219
-
220
- /**
221
- * Makora's GLM models, built from the same models list this extension
222
- * registers. `before_provider_request` is GLOBAL in pi — it fires for every
223
- * loaded provider's requests, not just makora's — and its event carries only
224
- * `payload` (no `provider` field to gate on). So the GLM tool_calls strip MUST
225
- * be scoped to makora's own GLM model ids. A loose `/glm/i` regex would also
226
- * rewrite tool_calls for every sibling provider that serves any GLM model
227
- * (baseten, io, lilac, neuralwatt, tensorix, fireworks, crofai, hypercharm,
228
- * wafer, parasail, umans, opencode, ...), silently breaking their tool calling
229
- * and clobbering the input their own before_provider_request handlers expect.
230
- *
231
- * Derived (not hardcoded) so it auto-tracks models.json refreshes — models.json
232
- * is auto-generated by scripts/update-models.js. All makora models run on vLLM,
233
- * so every GLM id here shares the ZAI chat-template `.items()` crash that
234
- * stripGlmToolCalls works around.
235
- */
236
198
  const allMakoraModels = buildModels(
237
199
  modelsData as JsonModel[],
238
200
  customModelsData as JsonModel[],
239
201
  patchData as PatchMap,
240
202
  );
241
- const GLM_VLLM_MODELS = new Set(
242
- allMakoraModels.filter((m) => /glm/i.test(m.id)).map((m) => m.id),
243
- );
244
-
245
- /** Whether `model` is one of makora's GLM models — the only ids the
246
- * before_provider_request hook should strip tool_calls for. Exported so the
247
- * scoping can be unit-tested in isolation. */
248
- export function isMakoraGlmVllmModel(model: string): boolean {
249
- return GLM_VLLM_MODELS.has(model);
250
- }
251
-
252
- /**
253
- * Intercept the request payload for models that need vLLM-specific thinking
254
- * param rewrites.
255
- *
256
- * pi's "deepseek" thinkingFormat sends `thinking: { type: "enabled" }` which
257
- * is the official DeepSeek API format — but Makora's vLLM deployment ignores
258
- * it. vLLM requires different params depending on the model:
259
- * - DS V4 Pro: `chat_template_kwargs: { thinking: true }` + `reasoning_effort`
260
- * - DS V4 Flash: `include_reasoning: true` + `chat_template_kwargs: { thinking: true }`
261
- * + `reasoning_effort`. `include_reasoning` alone returns `reasoning: null`
262
- * on this vLLM build — both params are required.
263
- *
264
- * This hook rewrites the payload accordingly.
265
- */
266
- function rewriteVllmPayload(payload: Record<string, unknown>): Record<string, unknown> {
267
- const model = payload.model as string | undefined;
268
- if (!model) return payload;
269
-
270
- const p = { ...payload };
271
-
272
- if (DS_VLLM_MODELS.has(model)) {
273
- // Remove the DeepSeek API-style `thinking` param that vLLM ignores
274
- delete p.thinking;
275
-
276
- if (model === DS_PRO_ID) {
277
- // DS Pro: chat_template_kwargs.thinking + reasoning_effort
278
- const ctq = (p.chat_template_kwargs as Record<string, unknown>) ?? {};
279
- p.chat_template_kwargs = { ...ctq, thinking: true };
280
- } else if (model === DS_FLASH_ID) {
281
- // DS Flash: include_reasoning + chat_template_kwargs.thinking + reasoning_effort
282
- // vLLM requires *both* include_reasoning and chat_template_kwargs.thinking:
283
- // include_reasoning alone returns reasoning: null.
284
- p.include_reasoning = true;
285
- const ctq = (p.chat_template_kwargs as Record<string, unknown>) ?? {};
286
- p.chat_template_kwargs = { ...ctq, thinking: true };
287
- }
288
- }
289
-
290
- return p;
291
- }
292
-
293
- /**
294
- * GLM models on Makora's vLLM crash with a leaked Python AttributeError
295
- * (`'list object' has no attribute 'items'` or `'str object' has no attribute 'items'`)
296
- * when any assistant message in the request contains a `tool_calls` field.
297
- * The ZAI/vLLM chat template calls `.items()` on the tool_calls list (or on the
298
- * JSON-string `arguments` field), which raises AttributeError and leaks into the
299
- * HTTP 400 response body.
300
- *
301
- * Fix: for makora's GLM models only (see GLM_VLLM_MODELS / isMakoraGlmVllmModel),
302
- * strip `tool_calls` from assistant messages in the before_provider_request hook
303
- * and convert them back to GLM's native
304
- * `<tool_call>` XML text in `content`. The model natively understands this
305
- * format in conversation history, and the `role: "tool"` result messages that
306
- * follow are rendered fine by the chat template's tool-observation branch.
307
- *
308
- * If upstream fixes both the streaming parser and the 500/400 crash, this
309
- * transform becomes a harmless no-op (the XML text is still valid GLM input).
310
- */
311
-
312
- export function toolCallToGlmXml(tc: Record<string, unknown>): string {
313
- const fn = (isObject(tc.function) ? tc.function : {}) as Record<string, unknown>;
314
- const name = typeof fn.name === "string" ? fn.name : "";
315
- const argsStr = typeof fn.arguments === "string" ? fn.arguments : "{}";
316
- let args: Record<string, unknown>;
317
- try {
318
- args = JSON.parse(argsStr) as Record<string, unknown>;
319
- } catch {
320
- args = {};
321
- }
322
- if (!isObject(args)) args = {};
323
- const argLines = Object.entries(args).map(
324
- ([key, value]) =>
325
- `<arg_key>${key}</arg_key>\n<arg_value>${
326
- typeof value === "string" ? value : JSON.stringify(value)
327
- }</arg_value>`,
328
- );
329
- return `<tool_call>${name}\n${argLines.join("\n")}\n</tool_call>`;
330
- }
331
-
332
- export function stripGlmToolCalls(payload: Record<string, unknown>): Record<string, unknown> {
333
- const messages = payload.messages;
334
- if (!Array.isArray(messages)) return payload;
335
-
336
- let modified = false;
337
- const newMessages = messages.map((msg) => {
338
- if (!isObject(msg)) return msg;
339
- if (msg.role !== "assistant") return msg;
340
-
341
- const toolCalls = msg.tool_calls;
342
- if (!Array.isArray(toolCalls) || toolCalls.length === 0) return msg;
343
-
344
- const xmlBlocks = toolCalls
345
- .filter((tc): tc is Record<string, unknown> => isObject(tc))
346
- .map(toolCallToGlmXml);
347
- if (xmlBlocks.length === 0) return msg;
348
-
349
- const toolCallText = xmlBlocks.join("\n");
350
- // pi's openai-completions provider always serializes assistant content as a
351
- // plain string, but guard against null/array content for robustness.
352
- const existingContent = typeof msg.content === "string" ? msg.content : "";
353
- const newContent = existingContent ? `${existingContent}\n${toolCallText}` : toolCallText;
354
-
355
- modified = true;
356
- const rest: Record<string, unknown> = {};
357
- for (const [k, v] of Object.entries(msg)) {
358
- if (k !== "tool_calls") rest[k] = v;
359
- }
360
- return { ...rest, content: newContent };
361
- });
362
-
363
- if (!modified) return payload;
364
- return { ...payload, messages: newMessages };
365
- }
366
203
 
367
204
  export default function (pi: ExtensionAPI) {
368
205
  const models = allMakoraModels;
@@ -376,15 +213,7 @@ export default function (pi: ExtensionAPI) {
376
213
  models,
377
214
  });
378
215
 
379
- pi.on("before_provider_request", (event) => {
380
- const payload = event.payload as Record<string, unknown> | undefined;
381
- if (!payload || typeof payload.model !== "string") return;
382
-
383
- let result = rewriteVllmPayload(payload);
384
- if (isMakoraGlmVllmModel(payload.model)) {
385
- result = stripGlmToolCalls(result);
386
- }
387
- return result;
388
- });
216
+ // Abort runaway '!' repetition loops on the GLM 5.2 family and resume the
217
+ // agentic loop invisibly (no new user message). See ./death-loop-guard.ts.
218
+ registerDeathLoopGuard(pi);
389
219
  }
390
-
package/models.json CHANGED
@@ -41,6 +41,27 @@
41
41
  "maxTokensField": "max_completion_tokens"
42
42
  }
43
43
  },
44
+ {
45
+ "id": "google/gemma-4-26B-A4B",
46
+ "name": "Gemma 4 26B A4B",
47
+ "reasoning": false,
48
+ "input": [
49
+ "text"
50
+ ],
51
+ "cost": {
52
+ "input": 0,
53
+ "output": 0,
54
+ "cacheRead": 0,
55
+ "cacheWrite": 0
56
+ },
57
+ "contextWindow": 262144,
58
+ "maxTokens": 0,
59
+ "compat": {
60
+ "supportsDeveloperRole": false,
61
+ "supportsStore": false,
62
+ "maxTokensField": "max_completion_tokens"
63
+ }
64
+ },
44
65
  {
45
66
  "id": "meta-llama/Llama-3.3-70B-Instruct",
46
67
  "name": "Llama 3.3 70B Instruct",
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "pi-makora-provider",
3
- "version": "1.3.0",
3
+ "version": "1.5.0",
4
4
  "description": "Makora provider extension for pi - Access DeepSeek V4, GLM 5.2, Kimi K2.7 Code, Llama 3.3, Qwen 3.6, and more through the Makora inference API",
5
5
  "type": "module",
6
6
  "main": "index.ts",
@@ -21,6 +21,7 @@
21
21
  "license": "MIT",
22
22
  "files": [
23
23
  "index.ts",
24
+ "death-loop-guard.ts",
24
25
  "models.json",
25
26
  "custom-models.json",
26
27
  "patch.json",
@@ -33,6 +34,8 @@
33
34
  ]
34
35
  },
35
36
  "devDependencies": {
37
+ "@earendil-works/pi-agent-core": "^0.80.2",
38
+ "@earendil-works/pi-coding-agent": "^0.80.2",
36
39
  "vitest": "^4.1.9"
37
40
  },
38
41
  "scripts": {
package/patch.json CHANGED
@@ -1,7 +1,7 @@
1
1
  {
2
2
  "deepseek-ai/DeepSeek-V4-Flash": {
3
3
  "reasoning": true,
4
- "notes": "`include_reasoning` + `chat_template_kwargs.thinking` via `before_provider_request` payload rewrite; returns `reasoning` field",
4
+ "notes": "returns `reasoning` field",
5
5
  "thinkingLevelMap": {
6
6
  "minimal": null,
7
7
  "low": null,
@@ -16,7 +16,7 @@
16
16
  },
17
17
  "deepseek-ai/DeepSeek-V4-Pro": {
18
18
  "reasoning": true,
19
- "notes": "`chat_template_kwargs.thinking` via `before_provider_request` payload rewrite; returns `reasoning` field",
19
+ "notes": "returns `reasoning` field",
20
20
  "thinkingLevelMap": {
21
21
  "minimal": null,
22
22
  "low": null,
@@ -32,7 +32,7 @@
32
32
  },
33
33
  "unsloth/Qwen3.6-27B-NVFP4": {
34
34
  "reasoning": true,
35
- "notes": "`enable_thinking` via `qwen-chat-template`; `preserve_thinking` via `chatTemplateKwargs` for multi-turn continuity; returns `reasoning` field; client-side tool call parsing (vLLM streaming parser bypass)",
35
+ "notes": "`enable_thinking` via `qwen-chat-template`; `preserve_thinking` via `chatTemplateKwargs` for multi-turn continuity; returns `reasoning` field",
36
36
  "thinkingLevelMap": {
37
37
  "minimal": "low",
38
38
  "xhigh": "high"
@@ -47,7 +47,7 @@
47
47
  },
48
48
  "unsloth/Qwen3.6-35B-A3B-NVFP4": {
49
49
  "reasoning": true,
50
- "notes": "`enable_thinking` via `qwen-chat-template`; `preserve_thinking` via `chatTemplateKwargs` for multi-turn continuity; returns `reasoning` field; client-side tool call parsing (vLLM streaming parser bypass)",
50
+ "notes": "`enable_thinking` via `qwen-chat-template`; `preserve_thinking` via `chatTemplateKwargs` for multi-turn continuity; returns `reasoning` field",
51
51
  "thinkingLevelMap": {
52
52
  "minimal": "low",
53
53
  "xhigh": "high"
@@ -66,7 +66,7 @@
66
66
  "text",
67
67
  "image"
68
68
  ],
69
- "notes": "Reasoning on by default (thinking-only model); `preserve_thinking` via `chatTemplateKwargs` for multi-turn continuity; returns `reasoning` field; client-side tool call parsing (vLLM streaming parser bypass)",
69
+ "notes": "Reasoning on by default (thinking-only model); `preserve_thinking` via `chatTemplateKwargs` for multi-turn continuity; returns `reasoning` field",
70
70
  "thinkingLevelMap": {
71
71
  "minimal": "low",
72
72
  "xhigh": "high"