pi-makora-provider 1.3.0 → 1.5.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +36 -26
- package/death-loop-guard.ts +238 -0
- package/index.ts +15 -186
- package/models.json +21 -0
- package/package.json +4 -1
- package/patch.json +5 -5
package/README.md
CHANGED
|
@@ -4,7 +4,7 @@
|
|
|
4
4
|
|
|
5
5
|
**Open-weight models through [Makora](https://inference.makora.com)**
|
|
6
6
|
|
|
7
|
-
_DeepSeek V4, Kimi K2.7 Code, GLM 5.2, Qwen 3.6
|
|
7
|
+
_DeepSeek V4, Kimi K2.7 Code, GLM 5.2, Qwen 3.6 for [pi](https://github.com/earendil-works/pi-coding-agent)._
|
|
8
8
|
|
|
9
9
|
[](https://github.com/earendil-works/pi-coding-agent)
|
|
10
10
|
[](./LICENSE)
|
|
@@ -18,15 +18,16 @@ _DeepSeek V4, Kimi K2.7 Code, GLM 5.2, Qwen 3.6 — with client-side tool call r
|
|
|
18
18
|
<!-- MODELS_TABLE_START -->
|
|
19
19
|
| Model | ID | Reasoning | Notes |
|
|
20
20
|
|-------|----|-----------|-------|
|
|
21
|
-
| DeepSeek V4 Flash | `deepseek-ai/DeepSeek-V4-Flash` | Yes |
|
|
22
|
-
| DeepSeek V4 Pro | `deepseek-ai/DeepSeek-V4-Pro` | Yes |
|
|
21
|
+
| DeepSeek V4 Flash | `deepseek-ai/DeepSeek-V4-Flash` | Yes | returns `reasoning` field |
|
|
22
|
+
| DeepSeek V4 Pro | `deepseek-ai/DeepSeek-V4-Pro` | Yes | returns `reasoning` field |
|
|
23
|
+
| Gemma 4 26B A4B | `google/gemma-4-26B-A4B` | No | |
|
|
23
24
|
| GLM 5.2 FP8 | `zai-org/GLM-5.2-FP8` | Yes | `enable_thinking` via `qwen-chat-template`; effort via `reasoning_effort` (only `high`/`max` distinct, per vLLM GLM-5.2 recipe); thinking levels aligned with neuralwatt GLM 5.2; returns `reasoning` field |
|
|
24
25
|
| GLM 5.2 NVFP4 | `zai-org/GLM-5.2-NVFP4` | Yes | `enable_thinking` via `qwen-chat-template`; effort via `reasoning_effort` (only `high`/`max` distinct, per vLLM GLM-5.2 recipe); returns `reasoning` field |
|
|
25
|
-
| Kimi K2.7 Code | `moonshotai/Kimi-K2.7-Code` | Yes | Reasoning on by default (thinking-only model); `preserve_thinking` via `chatTemplateKwargs` for multi-turn continuity; returns `reasoning` field
|
|
26
|
+
| Kimi K2.7 Code | `moonshotai/Kimi-K2.7-Code` | Yes | Reasoning on by default (thinking-only model); `preserve_thinking` via `chatTemplateKwargs` for multi-turn continuity; returns `reasoning` field |
|
|
26
27
|
| Llama 3.3 70B FP8 | `amd/Llama-3.3-70B-Instruct-FP8-KV` | No | |
|
|
27
28
|
| Llama 3.3 70B Instruct | `meta-llama/Llama-3.3-70B-Instruct` | No | |
|
|
28
|
-
| Qwen 3.6 27B NVFP4 | `unsloth/Qwen3.6-27B-NVFP4` | Yes | `enable_thinking` via `qwen-chat-template`; `preserve_thinking` via `chatTemplateKwargs` for multi-turn continuity; returns `reasoning` field
|
|
29
|
-
| Qwen 3.6 35B A3B NVFP4 | `unsloth/Qwen3.6-35B-A3B-NVFP4` | Yes | `enable_thinking` via `qwen-chat-template`; `preserve_thinking` via `chatTemplateKwargs` for multi-turn continuity; returns `reasoning` field
|
|
29
|
+
| Qwen 3.6 27B NVFP4 | `unsloth/Qwen3.6-27B-NVFP4` | Yes | `enable_thinking` via `qwen-chat-template`; `preserve_thinking` via `chatTemplateKwargs` for multi-turn continuity; returns `reasoning` field |
|
|
30
|
+
| Qwen 3.6 35B A3B NVFP4 | `unsloth/Qwen3.6-35B-A3B-NVFP4` | Yes | `enable_thinking` via `qwen-chat-template`; `preserve_thinking` via `chatTemplateKwargs` for multi-turn continuity; returns `reasoning` field |
|
|
30
31
|
<!-- MODELS_TABLE_END -->
|
|
31
32
|
|
|
32
33
|
## Installation
|
|
@@ -157,23 +158,32 @@ Do **not** edit `models.json` directly — it is auto-generated from the API. To
|
|
|
157
158
|
- All models are hosted on vLLM
|
|
158
159
|
- The `developer` role is not supported (prompts are silently dropped); `supportsDeveloperRole` is set to `false` for all models
|
|
159
160
|
|
|
160
|
-
##
|
|
161
|
-
|
|
162
|
-
|
|
163
|
-
|
|
164
|
-
|
|
165
|
-
|
|
166
|
-
|
|
167
|
-
|
|
168
|
-
|
|
169
|
-
|
|
170
|
-
|
|
171
|
-
|
|
172
|
-
|
|
173
|
-
|
|
174
|
-
|
|
175
|
-
|
|
176
|
-
|
|
177
|
-
|
|
178
|
-
|
|
179
|
-
|
|
161
|
+
## Death-Loop Guard
|
|
162
|
+
|
|
163
|
+
GLM 5.2 (NVFP4 / FP8) occasionally degenerates into an unbroken `!` repetition
|
|
164
|
+
loop (`!!!!...`) that eats the whole response. This extension ships a guard
|
|
165
|
+
that watches the streamed assistant output (both the visible answer and the
|
|
166
|
+
reasoning trace) and, when it detects a long run of `!` characters, **aborts
|
|
167
|
+
the runaway generation and resumes the agentic loop invisibly** — no new user
|
|
168
|
+
message is injected, using the
|
|
169
|
+
[pi-invisible-continue](https://github.com/monotykamary/pi-invisible-continue)
|
|
170
|
+
pattern (`agent.prompt([])`).
|
|
171
|
+
|
|
172
|
+
On abort, pi finalizes the in-flight assistant message **with its accumulated
|
|
173
|
+
`!!!` content** into the transcript, so the guard drops that partial message
|
|
174
|
+
before resuming — otherwise the model would just re-read its own `!!!` and loop
|
|
175
|
+
again. The recovery is bounded per prompt (default 3 attempts) to avoid
|
|
176
|
+
abort/continue thrash.
|
|
177
|
+
|
|
178
|
+
The guard is scoped to the Makora GLM 5.2 family by default and is tunable via
|
|
179
|
+
constants at the top of [`death-loop-guard.ts`](./death-loop-guard.ts):
|
|
180
|
+
|
|
181
|
+
| Constant | Default | Meaning |
|
|
182
|
+
|---|---|---|
|
|
183
|
+
| `GUARDED_MODEL_IDS` | `zai-org/GLM-5.2-NVFP4`, `zai-org/GLM-5.2-FP8` | Which model IDs to guard. Add `'*'` to guard every Makora model. |
|
|
184
|
+
| `BANG_THRESHOLD` | `40` | Consecutive `!` characters that trip the guard. 40 is far above anything normal prose/code produces. |
|
|
185
|
+
| `MAX_RECOVERS_PER_RUN` | `3` | Max invisible recoveries per user prompt before the guard stops intervening. |
|
|
186
|
+
|
|
187
|
+
It is only active for the `makora` provider, so it never interferes when you
|
|
188
|
+
switch to another provider. If you also run `pi-invisible-continue`, the two
|
|
189
|
+
coexist — both chain the `Agent.prototype.subscribe` patch.
|
|
@@ -0,0 +1,238 @@
|
|
|
1
|
+
/**
|
|
2
|
+
* Death-loop guard for Makora reasoning models.
|
|
3
|
+
*
|
|
4
|
+
* Some Makora models (notably GLM 5.2 NVFP4 / FP8) occasionally fall into a
|
|
5
|
+
* degenerate repetition loop, emitting an unbroken run of '!' characters
|
|
6
|
+
* (e.g. "!!!!...") that consumes the whole response. This guard watches the
|
|
7
|
+
* streamed assistant output (both the visible answer and the reasoning
|
|
8
|
+
* trace); when that run is detected it aborts the runaway generation, drops
|
|
9
|
+
* the partial (toxic) assistant message from the transcript, and resumes the
|
|
10
|
+
* agentic loop invisibly via agent.prompt([]) — the same pattern
|
|
11
|
+
* pi-invisible-continue uses, so no new user message pollutes the context.
|
|
12
|
+
*
|
|
13
|
+
* Why trim the aborted message: on abort, pi finalizes the in-flight
|
|
14
|
+
* assistant message WITH its accumulated '!!!' content and stopReason
|
|
15
|
+
* "aborted" into the transcript. Resuming from that context would re-feed
|
|
16
|
+
* the toxic text to the model and likely re-trigger the loop. Dropping the
|
|
17
|
+
* last (aborted) assistant message leaves the context ending at the prior
|
|
18
|
+
* user/toolResult message — a clean continuation point.
|
|
19
|
+
*
|
|
20
|
+
* Module resolution: @earendil-works/pi-agent-core is a devDependency only
|
|
21
|
+
* (types + test resolution). At runtime pi's extension loader aliases that
|
|
22
|
+
* specifier to its bundled copy, so the Agent class patched below is the
|
|
23
|
+
* SAME class AgentSession uses. A static import is required — jiti's alias
|
|
24
|
+
* applies to static imports (which it rewrites to its own resolver) but not
|
|
25
|
+
* to native dynamic import() calls.
|
|
26
|
+
*/
|
|
27
|
+
|
|
28
|
+
import type { ExtensionAPI } from "@earendil-works/pi-coding-agent";
|
|
29
|
+
import { Agent } from "@earendil-works/pi-agent-core";
|
|
30
|
+
|
|
31
|
+
const PROVIDER_ID = "makora";
|
|
32
|
+
|
|
33
|
+
/** Makora model IDs to guard. Add ids to widen coverage, or include "*"
|
|
34
|
+
* to guard every Makora model. Defaults to the GLM 5.2 family (the known
|
|
35
|
+
* offender). Kept exported so tests and downstream forks can introspect. */
|
|
36
|
+
export const GUARDED_MODEL_IDS = new Set<string>([
|
|
37
|
+
"zai-org/GLM-5.2-NVFP4",
|
|
38
|
+
"zai-org/GLM-5.2-FP8",
|
|
39
|
+
]);
|
|
40
|
+
|
|
41
|
+
/** Trip after this many consecutive '!' characters in the streamed answer.
|
|
42
|
+
* 40 is far above anything normal prose or code produces. */
|
|
43
|
+
export const BANG_THRESHOLD = 40;
|
|
44
|
+
|
|
45
|
+
/** Max invisible recoveries per user prompt, to bound abort/continue thrash. */
|
|
46
|
+
export const MAX_RECOVERS_PER_RUN = 3;
|
|
47
|
+
|
|
48
|
+
export interface GuardedMessage {
|
|
49
|
+
role: string;
|
|
50
|
+
stopReason?: string;
|
|
51
|
+
content?: unknown;
|
|
52
|
+
}
|
|
53
|
+
|
|
54
|
+
export interface GuardedAgent {
|
|
55
|
+
abort(): void;
|
|
56
|
+
waitForIdle(): Promise<void>;
|
|
57
|
+
prompt(input: unknown[] | string): Promise<void>;
|
|
58
|
+
state: { messages: GuardedMessage[] };
|
|
59
|
+
}
|
|
60
|
+
|
|
61
|
+
let _agent: GuardedAgent | null = null;
|
|
62
|
+
|
|
63
|
+
/** Mutex held only across the abort+trim critical section — NOT across the
|
|
64
|
+
* resumed prompt([]) run, so the resumed stream stays monitored. */
|
|
65
|
+
let _recovering = false;
|
|
66
|
+
|
|
67
|
+
/** Trailing '!' run length in the current text block. */
|
|
68
|
+
let _trailingBangs = 0;
|
|
69
|
+
|
|
70
|
+
/** Latch: already tripped for the current assistant message. */
|
|
71
|
+
let _tripped = false;
|
|
72
|
+
|
|
73
|
+
/** Recoveries performed in the current user-initiated agent run. */
|
|
74
|
+
let _recoversThisRun = 0;
|
|
75
|
+
|
|
76
|
+
export function isGuardedModel(
|
|
77
|
+
model: { provider?: string; id?: string } | undefined | null,
|
|
78
|
+
): boolean {
|
|
79
|
+
if (!model) return false;
|
|
80
|
+
if (model.provider !== PROVIDER_ID) return false;
|
|
81
|
+
if (GUARDED_MODEL_IDS.has("*")) return true;
|
|
82
|
+
return model.id != null && GUARDED_MODEL_IDS.has(model.id);
|
|
83
|
+
}
|
|
84
|
+
|
|
85
|
+
/** Update the trailing-'!' run length given a new text delta. O(len(delta)).
|
|
86
|
+
* Trailing run depends only on the delta's suffix: if the delta contains any
|
|
87
|
+
* non-'!' char, the prior run is cut off at that char; if the delta is all
|
|
88
|
+
* '!', it extends the prior run. */
|
|
89
|
+
export function nextTrailingBangs(prev: number, delta: string): number {
|
|
90
|
+
const len = delta.length;
|
|
91
|
+
if (len === 0) return prev;
|
|
92
|
+
let i = len - 1;
|
|
93
|
+
while (i >= 0 && delta.charCodeAt(i) === 0x21) i--; // '!' === 0x21
|
|
94
|
+
const trailingInDelta = len - 1 - i;
|
|
95
|
+
return i >= 0 ? trailingInDelta : prev + len;
|
|
96
|
+
}
|
|
97
|
+
|
|
98
|
+
export function extractText(content: unknown): string {
|
|
99
|
+
if (typeof content === "string") return content;
|
|
100
|
+
if (Array.isArray(content)) {
|
|
101
|
+
let out = "";
|
|
102
|
+
for (const block of content) {
|
|
103
|
+
if (
|
|
104
|
+
block &&
|
|
105
|
+
typeof block === "object" &&
|
|
106
|
+
(block as { type?: string }).type === "text" &&
|
|
107
|
+
typeof (block as { text?: unknown }).text === "string"
|
|
108
|
+
) {
|
|
109
|
+
out += (block as { text: string }).text;
|
|
110
|
+
}
|
|
111
|
+
}
|
|
112
|
+
return out;
|
|
113
|
+
}
|
|
114
|
+
return "";
|
|
115
|
+
}
|
|
116
|
+
|
|
117
|
+
/** Trailing '!' run length of a finalized message's text content. */
|
|
118
|
+
export function messageTrailingBangs(msg: GuardedMessage): number {
|
|
119
|
+
const text = extractText(msg.content);
|
|
120
|
+
let i = text.length - 1;
|
|
121
|
+
while (i >= 0 && text.charCodeAt(i) === 0x21) i--;
|
|
122
|
+
return text.length - 1 - i;
|
|
123
|
+
}
|
|
124
|
+
|
|
125
|
+
async function recover(agent: GuardedAgent): Promise<void> {
|
|
126
|
+
_recovering = true;
|
|
127
|
+
try {
|
|
128
|
+
agent.abort();
|
|
129
|
+
await agent.waitForIdle();
|
|
130
|
+
const msgs = agent.state.messages;
|
|
131
|
+
const last = msgs[msgs.length - 1];
|
|
132
|
+
if (
|
|
133
|
+
last &&
|
|
134
|
+
last.role === "assistant" &&
|
|
135
|
+
(last.stopReason === "aborted" || messageTrailingBangs(last) >= BANG_THRESHOLD)
|
|
136
|
+
) {
|
|
137
|
+
// Drop the toxic partial so it isn't re-fed to the model on resume.
|
|
138
|
+
// The state setter copies the array, leaving a clean transcript that
|
|
139
|
+
// ends at the prior user/toolResult message.
|
|
140
|
+
agent.state.messages = msgs.slice(0, -1);
|
|
141
|
+
}
|
|
142
|
+
} finally {
|
|
143
|
+
// Release before resuming so the resumed stream stays monitored.
|
|
144
|
+
_recovering = false;
|
|
145
|
+
}
|
|
146
|
+
try {
|
|
147
|
+
// Invisible continue: fresh agent loop, no new message injected.
|
|
148
|
+
await agent.prompt([]);
|
|
149
|
+
} catch {
|
|
150
|
+
// "Agent is already processing" or other transient error — best effort.
|
|
151
|
+
}
|
|
152
|
+
}
|
|
153
|
+
|
|
154
|
+
export function registerDeathLoopGuard(pi: ExtensionAPI): void {
|
|
155
|
+
// Capture the live Agent instance by chaining Agent.prototype.subscribe.
|
|
156
|
+
// subscribe() fires when AgentSession attaches — on every fresh session
|
|
157
|
+
// and every resume — so _agent always points at the active Agent. Chain
|
|
158
|
+
// any prior patch (e.g. pi-invisible-continue) so both extensions coexist.
|
|
159
|
+
const proto = Agent.prototype as unknown as {
|
|
160
|
+
subscribe: (this: GuardedAgent, ...args: unknown[]) => unknown;
|
|
161
|
+
};
|
|
162
|
+
const origSubscribe = proto.subscribe;
|
|
163
|
+
proto.subscribe = function (this: GuardedAgent, ...args: unknown[]) {
|
|
164
|
+
_agent = this;
|
|
165
|
+
return origSubscribe.apply(this, args);
|
|
166
|
+
};
|
|
167
|
+
|
|
168
|
+
pi.on("session_start", () => {
|
|
169
|
+
_recovering = false;
|
|
170
|
+
_trailingBangs = 0;
|
|
171
|
+
_tripped = false;
|
|
172
|
+
_recoversThisRun = 0;
|
|
173
|
+
});
|
|
174
|
+
|
|
175
|
+
// before_agent_start fires only for user prompts (the AgentSession path),
|
|
176
|
+
// not for the recovery's direct agent.prompt([]), so the recovery cap is
|
|
177
|
+
// bounded per user prompt instead of reset on every recovery continuation.
|
|
178
|
+
pi.on("before_agent_start", () => {
|
|
179
|
+
_recovering = false;
|
|
180
|
+
_trailingBangs = 0;
|
|
181
|
+
_tripped = false;
|
|
182
|
+
_recoversThisRun = 0;
|
|
183
|
+
});
|
|
184
|
+
|
|
185
|
+
pi.on("message_start", (event) => {
|
|
186
|
+
if (event.message?.role === "assistant") {
|
|
187
|
+
_trailingBangs = 0;
|
|
188
|
+
_tripped = false;
|
|
189
|
+
}
|
|
190
|
+
});
|
|
191
|
+
|
|
192
|
+
pi.on("message_update", (event, ctx) => {
|
|
193
|
+
if (_recovering || _tripped) return;
|
|
194
|
+
const ame = event.assistantMessageEvent;
|
|
195
|
+
if (ame.type === "text_start" || ame.type === "thinking_start") {
|
|
196
|
+
// New content block (answer or reasoning) — trailing run starts fresh.
|
|
197
|
+
_trailingBangs = 0;
|
|
198
|
+
return;
|
|
199
|
+
}
|
|
200
|
+
// Watch both the visible answer and the reasoning trace: GLM 5.2 is a
|
|
201
|
+
// reasoning model, and the loop can surface in either.
|
|
202
|
+
if (ame.type !== "text_delta" && ame.type !== "thinking_delta") return;
|
|
203
|
+
if (!isGuardedModel(ctx.model)) return;
|
|
204
|
+
|
|
205
|
+
_trailingBangs = nextTrailingBangs(_trailingBangs, ame.delta);
|
|
206
|
+
if (_trailingBangs < BANG_THRESHOLD) return;
|
|
207
|
+
|
|
208
|
+
if (_recoversThisRun >= MAX_RECOVERS_PER_RUN) {
|
|
209
|
+
_tripped = true;
|
|
210
|
+
ctx.ui.notify(
|
|
211
|
+
"Makora death-loop guard: runaway '!' output detected, but the " +
|
|
212
|
+
"recovery limit for this prompt was reached — stopping " +
|
|
213
|
+
"intervention. Try /continue or rephrase.",
|
|
214
|
+
"warning",
|
|
215
|
+
);
|
|
216
|
+
return;
|
|
217
|
+
}
|
|
218
|
+
_tripped = true;
|
|
219
|
+
_recoversThisRun++;
|
|
220
|
+
const agent = _agent;
|
|
221
|
+
if (!agent) {
|
|
222
|
+
ctx.ui.notify(
|
|
223
|
+
"Makora death-loop guard: runaway '!' output detected but the Agent " +
|
|
224
|
+
"instance was not captured; cannot recover automatically.",
|
|
225
|
+
"warning",
|
|
226
|
+
);
|
|
227
|
+
return;
|
|
228
|
+
}
|
|
229
|
+
ctx.ui.notify(
|
|
230
|
+
`Makora death-loop guard: aborting runaway '!' output and resuming ` +
|
|
231
|
+
`(${_recoversThisRun}/${MAX_RECOVERS_PER_RUN}).`,
|
|
232
|
+
"warning",
|
|
233
|
+
);
|
|
234
|
+
// Detach: the handler must return so the run can unwind to idle before
|
|
235
|
+
// recover() awaits waitForIdle() and calls prompt([]).
|
|
236
|
+
void recover(agent);
|
|
237
|
+
});
|
|
238
|
+
}
|
package/index.ts
CHANGED
|
@@ -13,32 +13,19 @@
|
|
|
13
13
|
* Model resolution strategy: static models.json merged with custom-models.json
|
|
14
14
|
*
|
|
15
15
|
* Reasoning notes:
|
|
16
|
-
* - DeepSeek V4 Pro: reasoning
|
|
17
|
-
*
|
|
18
|
-
*
|
|
19
|
-
*
|
|
20
|
-
*
|
|
21
|
-
* - DeepSeek V4 Flash: reasoning via include_reasoning +
|
|
22
|
-
* chat_template_kwargs.thinking on vLLM.
|
|
23
|
-
* The before_provider_request hook rewrites the payload to replace
|
|
24
|
-
* thinking: { type } with include_reasoning: true +
|
|
25
|
-
* chat_template_kwargs: { thinking: true }.
|
|
26
|
-
* include_reasoning alone returns reasoning: null on this vLLM build.
|
|
27
|
-
* Returns reasoning field.
|
|
28
|
-
* - GLM 5.2 FP8 / NVFP4: reasoning via chat_template_kwargs.enable_thinking.
|
|
29
|
-
* Effort uses vLLM's reasoning_effort field (only `high` and `max` are
|
|
30
|
-
* distinct levels per the vLLM GLM-5.2 recipe; lower pi levels resolve to
|
|
31
|
-
* the default). Thinking levels aligned with the neuralwatt provider's
|
|
32
|
-
* GLM 5.2 configuration and mapped through pi's qwen-chat-template
|
|
33
|
-
* thinkingFormat. Returns `reasoning` field.
|
|
34
|
-
* - Kimi K2.7 Code: reasoning always on (thinking-only model);
|
|
35
|
-
* chatTemplateKwargs.preserve_thinking forces multi-turn reasoning
|
|
36
|
-
* continuity. Returns `reasoning` field. Can be toggled via enable_thinking.
|
|
37
|
-
* - Qwen 3.6 models: reasoning via chat_template_kwargs.enable_thinking;
|
|
38
|
-
* chatTemplateKwargs.preserve_thinking for multi-turn continuity.
|
|
39
|
-
* Returns `reasoning` field.
|
|
16
|
+
* - DeepSeek V4 Pro: returns `reasoning` field.
|
|
17
|
+
* - DeepSeek V4 Flash: returns `reasoning` field.
|
|
18
|
+
* - GLM 5.2 FP8 / NVFP4: returns `reasoning` field.
|
|
19
|
+
* - Kimi K2.7 Code: returns `reasoning` field.
|
|
20
|
+
* - Qwen 3.6 models: returns `reasoning` field.
|
|
40
21
|
* - Llama 3.3 70B: not a reasoning model.
|
|
41
22
|
*
|
|
23
|
+
* A death-loop guard (see ./death-loop-guard.ts) is registered alongside the
|
|
24
|
+
* provider. It watches the assistant text stream on the GLM 5.2 family and,
|
|
25
|
+
* if the model falls into an unbroken '!' repetition loop, aborts the runaway
|
|
26
|
+
* generation and resumes the agentic loop invisibly via agent.prompt([]) (the
|
|
27
|
+
* pi-invisible-continue pattern) — no new user message is injected.
|
|
28
|
+
*
|
|
42
29
|
* Developer role is NOT supported by any of the chat templates on Makora's
|
|
43
30
|
* vLLM deployment (prompts with role: "developer" are silently dropped).
|
|
44
31
|
* supportsDeveloperRole is set to false for all models.
|
|
@@ -61,6 +48,7 @@ import type { ExtensionAPI } from "@earendil-works/pi-coding-agent";
|
|
|
61
48
|
import modelsData from "./models.json" with { type: "json" };
|
|
62
49
|
import customModelsData from "./custom-models.json" with { type: "json" };
|
|
63
50
|
import patchData from "./patch.json" with { type: "json" };
|
|
51
|
+
import { registerDeathLoopGuard } from "./death-loop-guard.js";
|
|
64
52
|
|
|
65
53
|
// Types
|
|
66
54
|
|
|
@@ -129,11 +117,6 @@ interface PatchEntry {
|
|
|
129
117
|
|
|
130
118
|
type PatchMap = Record<string, PatchEntry>;
|
|
131
119
|
|
|
132
|
-
/** Type guard: non-null, non-array object. */
|
|
133
|
-
export function isObject(value: unknown): value is Record<string, unknown> {
|
|
134
|
-
return value !== null && typeof value === "object" && !Array.isArray(value);
|
|
135
|
-
}
|
|
136
|
-
|
|
137
120
|
// Patch Application
|
|
138
121
|
|
|
139
122
|
function applyPatch(model: JsonModel, patch: PatchEntry): JsonModel {
|
|
@@ -212,157 +195,11 @@ function buildModels(
|
|
|
212
195
|
const PROVIDER_ID = "makora";
|
|
213
196
|
const BASE_URL = "https://inference.makora.com/v1";
|
|
214
197
|
|
|
215
|
-
const DS_PRO_ID = "deepseek-ai/DeepSeek-V4-Pro";
|
|
216
|
-
const DS_FLASH_ID = "deepseek-ai/DeepSeek-V4-Flash";
|
|
217
|
-
|
|
218
|
-
const DS_VLLM_MODELS = new Set([DS_PRO_ID, DS_FLASH_ID]);
|
|
219
|
-
|
|
220
|
-
/**
|
|
221
|
-
* Makora's GLM models, built from the same models list this extension
|
|
222
|
-
* registers. `before_provider_request` is GLOBAL in pi — it fires for every
|
|
223
|
-
* loaded provider's requests, not just makora's — and its event carries only
|
|
224
|
-
* `payload` (no `provider` field to gate on). So the GLM tool_calls strip MUST
|
|
225
|
-
* be scoped to makora's own GLM model ids. A loose `/glm/i` regex would also
|
|
226
|
-
* rewrite tool_calls for every sibling provider that serves any GLM model
|
|
227
|
-
* (baseten, io, lilac, neuralwatt, tensorix, fireworks, crofai, hypercharm,
|
|
228
|
-
* wafer, parasail, umans, opencode, ...), silently breaking their tool calling
|
|
229
|
-
* and clobbering the input their own before_provider_request handlers expect.
|
|
230
|
-
*
|
|
231
|
-
* Derived (not hardcoded) so it auto-tracks models.json refreshes — models.json
|
|
232
|
-
* is auto-generated by scripts/update-models.js. All makora models run on vLLM,
|
|
233
|
-
* so every GLM id here shares the ZAI chat-template `.items()` crash that
|
|
234
|
-
* stripGlmToolCalls works around.
|
|
235
|
-
*/
|
|
236
198
|
const allMakoraModels = buildModels(
|
|
237
199
|
modelsData as JsonModel[],
|
|
238
200
|
customModelsData as JsonModel[],
|
|
239
201
|
patchData as PatchMap,
|
|
240
202
|
);
|
|
241
|
-
const GLM_VLLM_MODELS = new Set(
|
|
242
|
-
allMakoraModels.filter((m) => /glm/i.test(m.id)).map((m) => m.id),
|
|
243
|
-
);
|
|
244
|
-
|
|
245
|
-
/** Whether `model` is one of makora's GLM models — the only ids the
|
|
246
|
-
* before_provider_request hook should strip tool_calls for. Exported so the
|
|
247
|
-
* scoping can be unit-tested in isolation. */
|
|
248
|
-
export function isMakoraGlmVllmModel(model: string): boolean {
|
|
249
|
-
return GLM_VLLM_MODELS.has(model);
|
|
250
|
-
}
|
|
251
|
-
|
|
252
|
-
/**
|
|
253
|
-
* Intercept the request payload for models that need vLLM-specific thinking
|
|
254
|
-
* param rewrites.
|
|
255
|
-
*
|
|
256
|
-
* pi's "deepseek" thinkingFormat sends `thinking: { type: "enabled" }` which
|
|
257
|
-
* is the official DeepSeek API format — but Makora's vLLM deployment ignores
|
|
258
|
-
* it. vLLM requires different params depending on the model:
|
|
259
|
-
* - DS V4 Pro: `chat_template_kwargs: { thinking: true }` + `reasoning_effort`
|
|
260
|
-
* - DS V4 Flash: `include_reasoning: true` + `chat_template_kwargs: { thinking: true }`
|
|
261
|
-
* + `reasoning_effort`. `include_reasoning` alone returns `reasoning: null`
|
|
262
|
-
* on this vLLM build — both params are required.
|
|
263
|
-
*
|
|
264
|
-
* This hook rewrites the payload accordingly.
|
|
265
|
-
*/
|
|
266
|
-
function rewriteVllmPayload(payload: Record<string, unknown>): Record<string, unknown> {
|
|
267
|
-
const model = payload.model as string | undefined;
|
|
268
|
-
if (!model) return payload;
|
|
269
|
-
|
|
270
|
-
const p = { ...payload };
|
|
271
|
-
|
|
272
|
-
if (DS_VLLM_MODELS.has(model)) {
|
|
273
|
-
// Remove the DeepSeek API-style `thinking` param that vLLM ignores
|
|
274
|
-
delete p.thinking;
|
|
275
|
-
|
|
276
|
-
if (model === DS_PRO_ID) {
|
|
277
|
-
// DS Pro: chat_template_kwargs.thinking + reasoning_effort
|
|
278
|
-
const ctq = (p.chat_template_kwargs as Record<string, unknown>) ?? {};
|
|
279
|
-
p.chat_template_kwargs = { ...ctq, thinking: true };
|
|
280
|
-
} else if (model === DS_FLASH_ID) {
|
|
281
|
-
// DS Flash: include_reasoning + chat_template_kwargs.thinking + reasoning_effort
|
|
282
|
-
// vLLM requires *both* include_reasoning and chat_template_kwargs.thinking:
|
|
283
|
-
// include_reasoning alone returns reasoning: null.
|
|
284
|
-
p.include_reasoning = true;
|
|
285
|
-
const ctq = (p.chat_template_kwargs as Record<string, unknown>) ?? {};
|
|
286
|
-
p.chat_template_kwargs = { ...ctq, thinking: true };
|
|
287
|
-
}
|
|
288
|
-
}
|
|
289
|
-
|
|
290
|
-
return p;
|
|
291
|
-
}
|
|
292
|
-
|
|
293
|
-
/**
|
|
294
|
-
* GLM models on Makora's vLLM crash with a leaked Python AttributeError
|
|
295
|
-
* (`'list object' has no attribute 'items'` or `'str object' has no attribute 'items'`)
|
|
296
|
-
* when any assistant message in the request contains a `tool_calls` field.
|
|
297
|
-
* The ZAI/vLLM chat template calls `.items()` on the tool_calls list (or on the
|
|
298
|
-
* JSON-string `arguments` field), which raises AttributeError and leaks into the
|
|
299
|
-
* HTTP 400 response body.
|
|
300
|
-
*
|
|
301
|
-
* Fix: for makora's GLM models only (see GLM_VLLM_MODELS / isMakoraGlmVllmModel),
|
|
302
|
-
* strip `tool_calls` from assistant messages in the before_provider_request hook
|
|
303
|
-
* and convert them back to GLM's native
|
|
304
|
-
* `<tool_call>` XML text in `content`. The model natively understands this
|
|
305
|
-
* format in conversation history, and the `role: "tool"` result messages that
|
|
306
|
-
* follow are rendered fine by the chat template's tool-observation branch.
|
|
307
|
-
*
|
|
308
|
-
* If upstream fixes both the streaming parser and the 500/400 crash, this
|
|
309
|
-
* transform becomes a harmless no-op (the XML text is still valid GLM input).
|
|
310
|
-
*/
|
|
311
|
-
|
|
312
|
-
export function toolCallToGlmXml(tc: Record<string, unknown>): string {
|
|
313
|
-
const fn = (isObject(tc.function) ? tc.function : {}) as Record<string, unknown>;
|
|
314
|
-
const name = typeof fn.name === "string" ? fn.name : "";
|
|
315
|
-
const argsStr = typeof fn.arguments === "string" ? fn.arguments : "{}";
|
|
316
|
-
let args: Record<string, unknown>;
|
|
317
|
-
try {
|
|
318
|
-
args = JSON.parse(argsStr) as Record<string, unknown>;
|
|
319
|
-
} catch {
|
|
320
|
-
args = {};
|
|
321
|
-
}
|
|
322
|
-
if (!isObject(args)) args = {};
|
|
323
|
-
const argLines = Object.entries(args).map(
|
|
324
|
-
([key, value]) =>
|
|
325
|
-
`<arg_key>${key}</arg_key>\n<arg_value>${
|
|
326
|
-
typeof value === "string" ? value : JSON.stringify(value)
|
|
327
|
-
}</arg_value>`,
|
|
328
|
-
);
|
|
329
|
-
return `<tool_call>${name}\n${argLines.join("\n")}\n</tool_call>`;
|
|
330
|
-
}
|
|
331
|
-
|
|
332
|
-
export function stripGlmToolCalls(payload: Record<string, unknown>): Record<string, unknown> {
|
|
333
|
-
const messages = payload.messages;
|
|
334
|
-
if (!Array.isArray(messages)) return payload;
|
|
335
|
-
|
|
336
|
-
let modified = false;
|
|
337
|
-
const newMessages = messages.map((msg) => {
|
|
338
|
-
if (!isObject(msg)) return msg;
|
|
339
|
-
if (msg.role !== "assistant") return msg;
|
|
340
|
-
|
|
341
|
-
const toolCalls = msg.tool_calls;
|
|
342
|
-
if (!Array.isArray(toolCalls) || toolCalls.length === 0) return msg;
|
|
343
|
-
|
|
344
|
-
const xmlBlocks = toolCalls
|
|
345
|
-
.filter((tc): tc is Record<string, unknown> => isObject(tc))
|
|
346
|
-
.map(toolCallToGlmXml);
|
|
347
|
-
if (xmlBlocks.length === 0) return msg;
|
|
348
|
-
|
|
349
|
-
const toolCallText = xmlBlocks.join("\n");
|
|
350
|
-
// pi's openai-completions provider always serializes assistant content as a
|
|
351
|
-
// plain string, but guard against null/array content for robustness.
|
|
352
|
-
const existingContent = typeof msg.content === "string" ? msg.content : "";
|
|
353
|
-
const newContent = existingContent ? `${existingContent}\n${toolCallText}` : toolCallText;
|
|
354
|
-
|
|
355
|
-
modified = true;
|
|
356
|
-
const rest: Record<string, unknown> = {};
|
|
357
|
-
for (const [k, v] of Object.entries(msg)) {
|
|
358
|
-
if (k !== "tool_calls") rest[k] = v;
|
|
359
|
-
}
|
|
360
|
-
return { ...rest, content: newContent };
|
|
361
|
-
});
|
|
362
|
-
|
|
363
|
-
if (!modified) return payload;
|
|
364
|
-
return { ...payload, messages: newMessages };
|
|
365
|
-
}
|
|
366
203
|
|
|
367
204
|
export default function (pi: ExtensionAPI) {
|
|
368
205
|
const models = allMakoraModels;
|
|
@@ -376,15 +213,7 @@ export default function (pi: ExtensionAPI) {
|
|
|
376
213
|
models,
|
|
377
214
|
});
|
|
378
215
|
|
|
379
|
-
|
|
380
|
-
|
|
381
|
-
|
|
382
|
-
|
|
383
|
-
let result = rewriteVllmPayload(payload);
|
|
384
|
-
if (isMakoraGlmVllmModel(payload.model)) {
|
|
385
|
-
result = stripGlmToolCalls(result);
|
|
386
|
-
}
|
|
387
|
-
return result;
|
|
388
|
-
});
|
|
216
|
+
// Abort runaway '!' repetition loops on the GLM 5.2 family and resume the
|
|
217
|
+
// agentic loop invisibly (no new user message). See ./death-loop-guard.ts.
|
|
218
|
+
registerDeathLoopGuard(pi);
|
|
389
219
|
}
|
|
390
|
-
|
package/models.json
CHANGED
|
@@ -41,6 +41,27 @@
|
|
|
41
41
|
"maxTokensField": "max_completion_tokens"
|
|
42
42
|
}
|
|
43
43
|
},
|
|
44
|
+
{
|
|
45
|
+
"id": "google/gemma-4-26B-A4B",
|
|
46
|
+
"name": "Gemma 4 26B A4B",
|
|
47
|
+
"reasoning": false,
|
|
48
|
+
"input": [
|
|
49
|
+
"text"
|
|
50
|
+
],
|
|
51
|
+
"cost": {
|
|
52
|
+
"input": 0,
|
|
53
|
+
"output": 0,
|
|
54
|
+
"cacheRead": 0,
|
|
55
|
+
"cacheWrite": 0
|
|
56
|
+
},
|
|
57
|
+
"contextWindow": 262144,
|
|
58
|
+
"maxTokens": 0,
|
|
59
|
+
"compat": {
|
|
60
|
+
"supportsDeveloperRole": false,
|
|
61
|
+
"supportsStore": false,
|
|
62
|
+
"maxTokensField": "max_completion_tokens"
|
|
63
|
+
}
|
|
64
|
+
},
|
|
44
65
|
{
|
|
45
66
|
"id": "meta-llama/Llama-3.3-70B-Instruct",
|
|
46
67
|
"name": "Llama 3.3 70B Instruct",
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "pi-makora-provider",
|
|
3
|
-
"version": "1.
|
|
3
|
+
"version": "1.5.0",
|
|
4
4
|
"description": "Makora provider extension for pi - Access DeepSeek V4, GLM 5.2, Kimi K2.7 Code, Llama 3.3, Qwen 3.6, and more through the Makora inference API",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"main": "index.ts",
|
|
@@ -21,6 +21,7 @@
|
|
|
21
21
|
"license": "MIT",
|
|
22
22
|
"files": [
|
|
23
23
|
"index.ts",
|
|
24
|
+
"death-loop-guard.ts",
|
|
24
25
|
"models.json",
|
|
25
26
|
"custom-models.json",
|
|
26
27
|
"patch.json",
|
|
@@ -33,6 +34,8 @@
|
|
|
33
34
|
]
|
|
34
35
|
},
|
|
35
36
|
"devDependencies": {
|
|
37
|
+
"@earendil-works/pi-agent-core": "^0.80.2",
|
|
38
|
+
"@earendil-works/pi-coding-agent": "^0.80.2",
|
|
36
39
|
"vitest": "^4.1.9"
|
|
37
40
|
},
|
|
38
41
|
"scripts": {
|
package/patch.json
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
{
|
|
2
2
|
"deepseek-ai/DeepSeek-V4-Flash": {
|
|
3
3
|
"reasoning": true,
|
|
4
|
-
"notes": "
|
|
4
|
+
"notes": "returns `reasoning` field",
|
|
5
5
|
"thinkingLevelMap": {
|
|
6
6
|
"minimal": null,
|
|
7
7
|
"low": null,
|
|
@@ -16,7 +16,7 @@
|
|
|
16
16
|
},
|
|
17
17
|
"deepseek-ai/DeepSeek-V4-Pro": {
|
|
18
18
|
"reasoning": true,
|
|
19
|
-
"notes": "
|
|
19
|
+
"notes": "returns `reasoning` field",
|
|
20
20
|
"thinkingLevelMap": {
|
|
21
21
|
"minimal": null,
|
|
22
22
|
"low": null,
|
|
@@ -32,7 +32,7 @@
|
|
|
32
32
|
},
|
|
33
33
|
"unsloth/Qwen3.6-27B-NVFP4": {
|
|
34
34
|
"reasoning": true,
|
|
35
|
-
"notes": "`enable_thinking` via `qwen-chat-template`; `preserve_thinking` via `chatTemplateKwargs` for multi-turn continuity; returns `reasoning` field
|
|
35
|
+
"notes": "`enable_thinking` via `qwen-chat-template`; `preserve_thinking` via `chatTemplateKwargs` for multi-turn continuity; returns `reasoning` field",
|
|
36
36
|
"thinkingLevelMap": {
|
|
37
37
|
"minimal": "low",
|
|
38
38
|
"xhigh": "high"
|
|
@@ -47,7 +47,7 @@
|
|
|
47
47
|
},
|
|
48
48
|
"unsloth/Qwen3.6-35B-A3B-NVFP4": {
|
|
49
49
|
"reasoning": true,
|
|
50
|
-
"notes": "`enable_thinking` via `qwen-chat-template`; `preserve_thinking` via `chatTemplateKwargs` for multi-turn continuity; returns `reasoning` field
|
|
50
|
+
"notes": "`enable_thinking` via `qwen-chat-template`; `preserve_thinking` via `chatTemplateKwargs` for multi-turn continuity; returns `reasoning` field",
|
|
51
51
|
"thinkingLevelMap": {
|
|
52
52
|
"minimal": "low",
|
|
53
53
|
"xhigh": "high"
|
|
@@ -66,7 +66,7 @@
|
|
|
66
66
|
"text",
|
|
67
67
|
"image"
|
|
68
68
|
],
|
|
69
|
-
"notes": "Reasoning on by default (thinking-only model); `preserve_thinking` via `chatTemplateKwargs` for multi-turn continuity; returns `reasoning` field
|
|
69
|
+
"notes": "Reasoning on by default (thinking-only model); `preserve_thinking` via `chatTemplateKwargs` for multi-turn continuity; returns `reasoning` field",
|
|
70
70
|
"thinkingLevelMap": {
|
|
71
71
|
"minimal": "low",
|
|
72
72
|
"xhigh": "high"
|