pi-lilac-provider 1.4.1 → 1.6.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -128,6 +128,23 @@ for M3's adaptive "model decides" mode. (The selector/footer show pi's level
128
128
  names — `minimal`/`high` — not the `thinking_mode` values; pi has no per-model
129
129
  level-relabel hook.)
130
130
 
131
+ **Preserved thinking (full-history reasoning).** By default these templates
132
+ trim older assistant reasoning between turns (each vendor's default), which
133
+ degrades multi-turn recall. Three models opt into full-history preservation via
134
+ a template flag sent alongside the reasoning key:
135
+
136
+ | Model | Flag | Effect |
137
+ |-------|------|--------|
138
+ | Kimi K2.6 | `preserve_thinking: true` | keeps every assistant turn's reasoning (default: only the last) |
139
+ | GLM 5.1 | `clear_thinking: false` | keeps reasoning for all turns (default: clears before the last user message) |
140
+ | GLM 5.2 | `clear_thinking: false` | keeps reasoning for all turns (default: clears before the last user message) |
141
+
142
+ Kimi K2.6 and GLM 5.2 are E2E-verified on the sibling neuralwatt provider via a
143
+ 3-turn, two-20-digit-number recall test (Kimi 0/6 → 6/6, GLM 5.2 1/4 → 4/4);
144
+ GLM 5.1 uses the same `clear_thinking` mechanism (confirmed in its HuggingFace
145
+ chat template). Gemma 4 and MiniMax M2.7/M3 expose no family-wide preserve flag,
146
+ so their older assistant reasoning is trimmed per the template default.
147
+
131
148
  In pi, reasoning models automatically use the appropriate thinking format. Use
132
149
  Shift+Tab to control thinking level.
133
150
 
@@ -173,7 +190,7 @@ Add to your pi configuration for automatic loading:
173
190
 
174
191
  Lilac's API is OpenAI-compatible with these specifics:
175
192
 
176
- - **`thinkingFormat: "chat-template"`** — All reasoning models. Lilac's vLLM backend toggles reasoning via `chat_template_kwargs`, but the honored key differs per model family. Per-model `chatTemplateKwargs` in `patch.json` send the right key(s): `thinking`+`enable_thinking` (bool) for Kimi K2.6, GLM 5.1, Gemma 4, and MiniMax M2.7; `enable_thinking` + `reasoning_effort` for GLM 5.2; `thinking_mode` (adaptive|enabled|disabled) for MiniMax M3.
193
+ - **`thinkingFormat: "chat-template"`** — All reasoning models. Lilac's vLLM backend toggles reasoning via `chat_template_kwargs`, but the honored key differs per model family. Per-model `chatTemplateKwargs` in `patch.json` send the right key(s): `thinking`+`enable_thinking` (bool) for Kimi K2.6, GLM 5.1, Gemma 4, and MiniMax M2.7; `enable_thinking` + `reasoning_effort` for GLM 5.2; `thinking_mode` (adaptive|enabled|disabled) for MiniMax M3. Kimi K2.6, GLM 5.1, and GLM 5.2 additionally send a preservation flag (`preserve_thinking: true` / `clear_thinking: false`) to retain full reasoning history across turns — see [Preserved thinking](#thinking-mode) above. Override these per-model via [Model Overrides](#model-overrides).
177
194
  - **`maxTokensField: "max_completion_tokens"`** — All models. Lilac supports `max_completion_tokens` (preferred for reasoning models as it includes reasoning tokens).
178
195
  - **`supportsDeveloperRole: true`** — All models. Lilac's vLLM backend maps the developer role to system.
179
196
  - **`supportsStore: false`** — All models. Lilac doesn't support the `store` parameter.
@@ -193,6 +210,27 @@ The `patch.json` file contains overrides that are applied on top of `models.json
193
210
  - Adding compat settings that the API doesn't provide
194
211
  - Overriding pricing when official rates change
195
212
 
213
+ ### Model Overrides
214
+
215
+ `modelOverrides` lets you override compat flags and other model properties per model id, **on top of** `patch.json` + `custom-models.json`, without editing the extension. Keyed by model id; `compat` (including nested `chatTemplateKwargs`), `thinkingLevelMap`, and `cost` are deep-merged **recursively** (toggle one flag without redeclaring the rest), scalars and arrays are replaced. Applied at session start, so edits take effect on the next `pi` session.
216
+
217
+ Create `~/.pi/agent/extensions/lilac.json` (auto-populated with defaults on first run):
218
+
219
+ ```jsonc
220
+ {
221
+ "modelOverrides": {
222
+ // Disable full-history reasoning for kimi-k2.6 (e.g. to save tokens):
223
+ "moonshotai/kimi-k2.6": { "compat": { "chatTemplateKwargs": { "preserve_thinking": false } } },
224
+ // Toggle the GLM 5.2 clear_thinking flag without redeclaring the rest of compat:
225
+ "zai-org/glm-5.2": { "compat": { "chatTemplateKwargs": { "clear_thinking": true } } },
226
+ // Override a single thinking level without redeclaring the whole map:
227
+ "zai-org/glm-5.1": { "thinkingLevelMap": { "high": "max" } }
228
+ }
229
+ }
230
+ ```
231
+
232
+ The full set of overridable fields matches the model schema (`compat`, `thinkingLevelMap`, `cost`, `contextWindow`, `maxTokens`, `reasoning`, `input`). See [Compat Settings](#compat-settings) for the catalog of compat flags and what `chatTemplateKwargs` values mean per family. An invalid JSON file is left untouched (defaults are used) so a typo isn't silently wiped — fix the file and restart pi.
233
+
196
234
  ## Updating Models
197
235
 
198
236
  Run the update script to fetch the latest models from Lilac's API:
package/index.ts CHANGED
@@ -178,6 +178,120 @@ interface PatchEntry {
178
178
 
179
179
  type PatchData = Record<string, PatchEntry>;
180
180
 
181
+ // User Configuration: ~/.pi/agent/extensions/lilac.json lets a user override model
182
+ // properties per id ON TOP of patch.json + custom-models.json (so they win).
183
+ // Recursively deep-merges `compat` (incl. nested `chatTemplateKwargs`),
184
+ // `thinkingLevelMap`, and `cost` (toggle one flag without redeclaring the rest),
185
+ // replaces scalars and arrays. Lets a user toggle chat_template_kwargs (e.g.
186
+ // preserve_thinking / clear_thinking) or a single thinking level without editing
187
+ // the extension. See README "Model Overrides".
188
+ interface ModelOverride {
189
+ thinkingLevelMap?: ThinkingLevelMap;
190
+ compat?: Record<string, unknown>;
191
+ }
192
+
193
+ interface LilacConfig {
194
+ modelOverrides?: Record<string, ModelOverride>;
195
+ }
196
+
197
+ const CONFIG_PATH = path.join(os.homedir(), ".pi", "agent", "extensions", "lilac.json");
198
+ const DEFAULT_CONFIG: LilacConfig = { modelOverrides: {} };
199
+
200
+ // Validate user-supplied modelOverrides from the config file. Non-object ids and
201
+ // non-object overrides are dropped silently so a malformed file doesn't crash
202
+ // model registration.
203
+ function parseModelOverrides(raw: unknown): Record<string, ModelOverride> | undefined {
204
+ if (!raw || typeof raw !== "object" || Array.isArray(raw)) return undefined;
205
+ const result: Record<string, ModelOverride> = {};
206
+ for (const [id, override] of Object.entries(raw as Record<string, unknown>)) {
207
+ if (typeof id !== "string" || !override || typeof override !== "object" || Array.isArray(override)) continue;
208
+ const o = override as Record<string, unknown>;
209
+ const parsed: ModelOverride = {};
210
+ if (o.thinkingLevelMap && typeof o.thinkingLevelMap === "object") {
211
+ const m: Record<string, string | null> = {};
212
+ for (const [k, v] of Object.entries(o.thinkingLevelMap as Record<string, unknown>)) {
213
+ if (v === null || typeof v === "string") m[k] = v;
214
+ }
215
+ if (Object.keys(m).length > 0) parsed.thinkingLevelMap = m as ThinkingLevelMap;
216
+ }
217
+ if (o.compat && typeof o.compat === "object") parsed.compat = o.compat as Record<string, unknown>;
218
+ if (Object.keys(parsed).length > 0) result[id] = parsed;
219
+ }
220
+ return Object.keys(result).length > 0 ? result : undefined;
221
+ }
222
+
223
+ // Reads ~/.pi/agent/extensions/lilac.json. Missing file → populate with defaults
224
+ // so the user can discover it, then return defaults. An existing-but-invalid file
225
+ // is left untouched (defaults returned) so a user's typo isn't silently wiped —
226
+ // they fix the file and restart pi. Loaded lazily on first use (not at import) so
227
+ // importing the module has no filesystem side effects and unit tests can import
228
+ // the pure helpers safely.
229
+ function loadConfig(): LilacConfig {
230
+ let rawText: string;
231
+ try {
232
+ rawText = fs.readFileSync(CONFIG_PATH, "utf8");
233
+ } catch {
234
+ try {
235
+ fs.mkdirSync(path.dirname(CONFIG_PATH), { recursive: true });
236
+ fs.writeFileSync(CONFIG_PATH, JSON.stringify(DEFAULT_CONFIG, null, 2) + "\n");
237
+ } catch {
238
+ // Write failure is non-fatal — defaults still work in memory
239
+ }
240
+ return { ...DEFAULT_CONFIG };
241
+ }
242
+ try {
243
+ const raw = JSON.parse(rawText);
244
+ return { modelOverrides: parseModelOverrides(raw.modelOverrides) };
245
+ } catch {
246
+ // File exists but is invalid JSON — return defaults WITHOUT overwriting.
247
+ return { ...DEFAULT_CONFIG };
248
+ }
249
+ }
250
+
251
+ // Recursively deep-merge `override` into `base`. Plain objects are merged
252
+ // key-by-key (so a user can toggle a single chatTemplateKwargs flag without
253
+ // redeclaring the rest, and a single compat flag without redeclaring
254
+ // chatTemplateKwargs); arrays and non-plain-object values replace the base value
255
+ // (so an overridden { $var } schema object, or an `input` array, replaces wholesale).
256
+ function isPlainObject(v: unknown): v is Record<string, unknown> {
257
+ return v !== null && typeof v === "object" && !Array.isArray(v);
258
+ }
259
+
260
+ function deepMerge<T>(base: T, override: unknown): T {
261
+ if (!isPlainObject(override) || !isPlainObject(base)) return override as T;
262
+ const result: Record<string, unknown> = { ...(base as Record<string, unknown>) };
263
+ for (const [k, v] of Object.entries(override)) {
264
+ result[k] = isPlainObject(v) && isPlainObject(result[k]) ? deepMerge(result[k], v) : v;
265
+ }
266
+ return result as T;
267
+ }
268
+
269
+ // Apply a user-supplied modelOverride (from lilac.json) on top of a built model.
270
+ // Recursively deep-merges compat (incl. nested chatTemplateKwargs) /
271
+ // thinkingLevelMap / cost so a user can toggle a single flag (e.g.
272
+ // chatTemplateKwargs.preserve_thinking) without redeclaring the rest; replaces
273
+ // scalars and arrays. No reasoning-cleanup (unlike applyPatch) — the override is
274
+ // authoritative.
275
+ function applyModelOverride(model: JsonModel, override: ModelOverride): JsonModel {
276
+ const result = { ...model };
277
+ for (const [key, value] of Object.entries(override)) {
278
+ (result as any)[key] = isPlainObject(value) && isPlainObject((result as any)[key])
279
+ ? deepMerge((result as any)[key], value)
280
+ : value;
281
+ }
282
+ return result;
283
+ }
284
+
285
+ let config: LilacConfig | undefined;
286
+ function getConfig(): LilacConfig {
287
+ if (!config) config = loadConfig();
288
+ return config;
289
+ }
290
+
291
+ function activeOverrides(): Record<string, ModelOverride> {
292
+ return getConfig().modelOverrides ?? {};
293
+ }
294
+
181
295
  // ─── Patch Application ────────────────────────────────────────────────────────
182
296
 
183
297
  function applyPatch(model: JsonModel, patch: PatchEntry): JsonModel {
@@ -217,8 +331,13 @@ function applyPatch(model: JsonModel, patch: PatchEntry): JsonModel {
217
331
  return result;
218
332
  }
219
333
 
220
- /** Full pipeline: base models → patch → custom → result */
221
- function buildModels(base: JsonModel[], custom: JsonModel[], patch: PatchData): JsonModel[] {
334
+ /** Full pipeline: base models → patch → custom → user modelOverrides → result */
335
+ function buildModels(
336
+ base: JsonModel[],
337
+ custom: JsonModel[],
338
+ patch: PatchData,
339
+ overrides: Record<string, ModelOverride> = {},
340
+ ): JsonModel[] {
222
341
  const modelMap = new Map<string, JsonModel>();
223
342
 
224
343
  for (const model of base) {
@@ -246,6 +365,18 @@ function buildModels(base: JsonModel[], custom: JsonModel[], patch: PatchData):
246
365
  }
247
366
  }
248
367
 
368
+ // User-supplied modelOverrides (from ~/.pi/agent/extensions/lilac.json) applied
369
+ // LAST so they win over patch.json + custom-models.json. Recursively deep-merges
370
+ // compat (incl. chatTemplateKwargs) / thinkingLevelMap / cost so a user can
371
+ // toggle a single flag (e.g. chatTemplateKwargs.preserve_thinking) without
372
+ // redeclaring the rest.
373
+ for (const [id, override] of Object.entries(overrides)) {
374
+ const existing = modelMap.get(id);
375
+ if (existing) {
376
+ modelMap.set(id, applyModelOverride(existing, override));
377
+ }
378
+ }
379
+
249
380
  return Array.from(modelMap.values());
250
381
  }
251
382
 
@@ -617,14 +748,14 @@ export default function (pi: ExtensionAPI) {
617
748
  // the in-flight model's cost without compounding an already-applied discount.
618
749
  function getListModels(): JsonModel[] {
619
750
  if (!listModelsCache) {
620
- listModelsCache = buildModels(loadStaleModels(embeddedModels), customModels, patches);
751
+ listModelsCache = buildModels(loadStaleModels(embeddedModels), customModels, patches, activeOverrides());
621
752
  }
622
753
  return listModelsCache;
623
754
  }
624
755
 
625
756
  const staleBase = loadStaleModels(embeddedModels);
626
757
  latestDiscounts = loadCachedDiscounts();
627
- const staleModels = applyDiscounts(buildModels(staleBase, customModels, patches), latestDiscounts);
758
+ const staleModels = applyDiscounts(buildModels(staleBase, customModels, patches, activeOverrides()), latestDiscounts);
628
759
 
629
760
  pi.registerProvider("lilac", {
630
761
  baseUrl: BASE_URL,
@@ -731,14 +862,14 @@ export default function (pi: ExtensionAPI) {
731
862
  baseUrl: BASE_URL,
732
863
  apiKey: "$LILAC_API_KEY",
733
864
  api: "openai-completions",
734
- models: applyDiscounts(buildModels(merged, customModels, patches), latestDiscounts),
865
+ models: applyDiscounts(buildModels(merged, customModels, patches, activeOverrides()), latestDiscounts),
735
866
  });
736
867
  } else if (discounts) {
737
868
  pi.registerProvider("lilac", {
738
869
  baseUrl: BASE_URL,
739
870
  apiKey: "$LILAC_API_KEY",
740
871
  api: "openai-completions",
741
- models: applyDiscounts(buildModels(staleBase, customModels, patches), latestDiscounts),
872
+ models: applyDiscounts(buildModels(staleBase, customModels, patches, activeOverrides()), latestDiscounts),
742
873
  });
743
874
  }
744
875
 
@@ -887,5 +1018,5 @@ export default function (pi: ExtensionAPI) {
887
1018
  });
888
1019
  }
889
1020
 
890
- export { fetchStatusDiscounts, applyDiscounts, applyDiscountInPlace, loadCachedDiscounts, cacheDiscounts };
891
- export type { JsonDiscount, JsonModel, PatchEntry, PatchData };
1021
+ export { fetchStatusDiscounts, applyDiscounts, applyDiscountInPlace, loadCachedDiscounts, cacheDiscounts, buildModels, applyModelOverride, parseModelOverrides, loadConfig, getConfig };
1022
+ export type { JsonDiscount, JsonModel, PatchEntry, PatchData, ModelOverride, LilacConfig };
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "pi-lilac-provider",
3
- "version": "1.4.1",
3
+ "version": "1.6.0",
4
4
  "description": "Lilac provider extension for pi - Access Kimi K2.6, GLM 5.1, and Gemma 4 models through Lilac's OpenAI-compatible API on idle GPUs",
5
5
  "type": "module",
6
6
  "main": "index.ts",
@@ -34,7 +34,10 @@
34
34
  "clean": "echo 'nothing to clean'",
35
35
  "build": "echo 'nothing to build'",
36
36
  "check": "echo 'nothing to check'",
37
- "test": "node scripts/test-discounts.ts",
37
+ "test": "node scripts/test-discounts.ts && node scripts/test-preserved-thinking.ts && node scripts/test-model-overrides.ts",
38
+ "test:discounts": "node scripts/test-discounts.ts",
39
+ "test:thinking": "node scripts/test-preserved-thinking.ts",
40
+ "test:overrides": "node scripts/test-model-overrides.ts",
38
41
  "update-models": "node scripts/update-models.js"
39
42
  }
40
43
  }
package/patch.json CHANGED
@@ -10,7 +10,8 @@
10
10
  "thinkingFormat": "chat-template",
11
11
  "chatTemplateKwargs": {
12
12
  "thinking": { "$var": "thinking.enabled" },
13
- "enable_thinking": { "$var": "thinking.enabled" }
13
+ "enable_thinking": { "$var": "thinking.enabled" },
14
+ "preserve_thinking": true
14
15
  },
15
16
  "maxTokensField": "max_completion_tokens",
16
17
  "supportsDeveloperRole": false,
@@ -29,7 +30,8 @@
29
30
  "thinkingFormat": "chat-template",
30
31
  "chatTemplateKwargs": {
31
32
  "thinking": { "$var": "thinking.enabled" },
32
- "enable_thinking": { "$var": "thinking.enabled" }
33
+ "enable_thinking": { "$var": "thinking.enabled" },
34
+ "clear_thinking": false
33
35
  },
34
36
  "maxTokensField": "max_completion_tokens",
35
37
  "supportsDeveloperRole": false,
@@ -42,7 +44,8 @@
42
44
  "thinkingFormat": "chat-template",
43
45
  "chatTemplateKwargs": {
44
46
  "enable_thinking": { "$var": "thinking.enabled" },
45
- "reasoning_effort": { "$var": "thinking.effort", "omitWhenOff": true }
47
+ "reasoning_effort": { "$var": "thinking.effort", "omitWhenOff": true },
48
+ "clear_thinking": false
46
49
  }
47
50
  },
48
51
  "thinkingLevelMap": {
@@ -17,7 +17,7 @@
17
17
  * recomputed from list price so re-applied discounts never compound.
18
18
  * 11. before_provider_request mutates the bound (in-flight) model object so the
19
19
  * current turn's cost calc sees the discount in real time.
20
- * 12. session_start schedules a 10-minute background /status poll to cover idle
20
+ * 12. session_start schedules a 5-minute background /status poll to cover idle
21
21
  * sessions; session_shutdown clears it.
22
22
  */
23
23
 
@@ -700,12 +700,12 @@ assert(
700
700
  "deferred session_start paints the NEW lilac model's discount (glm), not the stale kimi capture",
701
701
  );
702
702
 
703
- // ─── Test 17: session_start schedules a 10-min idle poll; shutdown clears it ─
703
+ // ─── Test 17: session_start schedules a 5-min idle poll; shutdown clears it ─
704
704
 
705
705
  console.log("\n--- Test 17: session_start schedules idle poll; session_shutdown clears it ---");
706
706
 
707
- // Lilac refreshes discounts ~every 10 minutes. session_start must schedule a
708
- // background /status poll at that cadence to cover idle sessions (turn fetches
707
+ // Lilac refreshes discounts ~every 10 minutes, so session_start polls /status
708
+ // every 5 min (half the refresh window) to cover idle sessions (turn fetches
709
709
  // only run when the user sends a message), and session_shutdown must clear it so
710
710
  // it neither leaks nor keeps the process alive. Wrap the global timer APIs to
711
711
  // capture the scheduled delay + handle, then confirm shutdown clears it.
@@ -729,7 +729,7 @@ globalThis.clearInterval = ((handle: ReturnType<typeof setInterval>) => {
729
729
 
730
730
  try {
731
731
  // Benign fetch mock so the session_start fire-and-forget /models + /status
732
- // fetch doesn't hit the network. (The poll itself never fires — 10 min — so
732
+ // fetch doesn't hit the network. (The poll itself never fires — 5 min — so
733
733
  // only the startup fetch needs mocking here.)
734
734
  globalThis.fetch = mockFetch({
735
735
  "/models": { body: { data: [] } },
@@ -759,8 +759,8 @@ try {
759
759
  }
760
760
 
761
761
  assert(
762
- scheduledDelay === 10 * 60 * 1000,
763
- "session_start schedules a 10-minute (600000ms) /status poll for idle sessions",
762
+ scheduledDelay === 5 * 60 * 1000,
763
+ "session_start schedules a 5-minute (300000ms) /status poll for idle sessions",
764
764
  );
765
765
  assert(scheduledHandle !== null, "poll interval handle was captured");
766
766
  const capturedHandle = scheduledHandle;
@@ -0,0 +1,262 @@
1
+ #!/usr/bin/env node
2
+ /**
3
+ * Test for user model overrides (~/.pi/agent/extensions/lilac.json).
4
+ *
5
+ * Verifies, against the REAL exported helpers from index.ts (not a re-implementation):
6
+ * - applyModelOverride: override wins over the base value; deep-merge compat /
7
+ * cost / thinkingLevelMap preserves non-overridden fields; scalars replaced;
8
+ * input model not mutated.
9
+ * - parseModelOverrides: drops invalid entries (non-object override, non-string
10
+ * thinkingLevelMap values), keeps valid ones.
11
+ * - loadConfig: auto-populates a scaffold on a missing file; parses an existing
12
+ * file's modelOverrides; returns defaults on invalid JSON WITHOUT overwriting
13
+ * the user's file; returns undefined overrides for a file lacking the key.
14
+ * - buildModels end-to-end: a user override on a real model id wins over
15
+ * patch.json (e.g. disable preserve_thinking on kimi-k2.6) while the rest of
16
+ * compat survives; with no overrides the built models are unchanged; an
17
+ * override for an unknown id adds no models.
18
+ *
19
+ * Config FS is isolated to a temp HOME so nothing touches the real ~/.pi.
20
+ */
21
+
22
+ import fs from "fs";
23
+ import os from "os";
24
+ import path from "path";
25
+
26
+ const root = path.resolve(import.meta.dirname, "..");
27
+ const modelsData = JSON.parse(fs.readFileSync(path.join(root, "models.json"), "utf8"));
28
+ const customModelsData = JSON.parse(fs.readFileSync(path.join(root, "custom-models.json"), "utf8"));
29
+ const patchData = JSON.parse(fs.readFileSync(path.join(root, "patch.json"), "utf8"));
30
+
31
+ // Isolate config + cache to a temp HOME so loadConfig never touches the real ~/.pi.
32
+ // Must be set before importing index.ts, which computes CONFIG_PATH at module scope.
33
+ const tmpHome = `/tmp/pi-lilac-override-test-${Date.now()}`;
34
+ fs.mkdirSync(tmpHome, { recursive: true });
35
+ process.env.HOME = tmpHome;
36
+
37
+ const {
38
+ buildModels,
39
+ applyModelOverride,
40
+ parseModelOverrides,
41
+ loadConfig,
42
+ } = await import("../index.ts");
43
+
44
+ let passed = 0;
45
+ let failed = 0;
46
+ function assert(condition: boolean, message: string) {
47
+ if (condition) {
48
+ console.log(` ✓ ${message}`);
49
+ passed++;
50
+ } else {
51
+ console.error(` ✗ ${message}`);
52
+ failed++;
53
+ }
54
+ }
55
+ function eq<T>(actual: T, expected: T, message: string) {
56
+ const ok = JSON.stringify(actual) === JSON.stringify(expected);
57
+ if (ok) {
58
+ console.log(` ✓ ${message}`);
59
+ passed++;
60
+ } else {
61
+ console.error(` ✗ ${message}\n expected: ${JSON.stringify(expected)}\n actual: ${JSON.stringify(actual)}`);
62
+ failed++;
63
+ }
64
+ }
65
+
66
+ const KIMI = "moonshotai/kimi-k2.6";
67
+ const GLM52 = "zai-org/glm-5.2";
68
+ const GLM51 = "zai-org/glm-5.1";
69
+
70
+ // ─── applyModelOverride ────────────────────────────────────────────────────────
71
+
72
+ console.log("\n--- applyModelOverride ---");
73
+
74
+ {
75
+ const base = {
76
+ id: KIMI,
77
+ reasoning: true,
78
+ compat: {
79
+ thinkingFormat: "chat-template",
80
+ supportsDeveloperRole: false,
81
+ chatTemplateKwargs: { thinking: { $var: "thinking.enabled" }, preserve_thinking: true },
82
+ },
83
+ thinkingLevelMap: { low: "low", high: "high" },
84
+ } as any;
85
+
86
+ const out = applyModelOverride(base, { compat: { chatTemplateKwargs: { preserve_thinking: false } } } as any);
87
+ assert(out.compat.chatTemplateKwargs.preserve_thinking === false, "override wins over base for a compat flag it sets");
88
+ assert(out.compat.thinkingFormat === "chat-template", "deep-merge compat preserves non-overridden thinkingFormat");
89
+ assert(out.compat.supportsDeveloperRole === false, "deep-merge compat preserves non-overridden supportsDeveloperRole");
90
+ assert((out.compat.chatTemplateKwargs as any).thinking?.$var === "thinking.enabled", "deep-merge chatTemplateKwargs preserves non-overridden thinking key");
91
+ assert((base.compat.chatTemplateKwargs as any).preserve_thinking === true, "does not mutate the input model");
92
+ }
93
+
94
+ {
95
+ const base = { id: GLM52, thinkingLevelMap: { low: "low", high: "high" } } as any;
96
+ const out = applyModelOverride(base, { thinkingLevelMap: { high: "max" } } as any);
97
+ eq(out.thinkingLevelMap, { low: "low", high: "max" }, "deep-merge thinkingLevelMap overrides a single level, keeps the rest");
98
+ }
99
+
100
+ {
101
+ const base = { id: KIMI, reasoning: true, contextWindow: 131072 } as any;
102
+ const out = applyModelOverride(base, { reasoning: false, contextWindow: 65536 } as any);
103
+ assert(out.reasoning === false, "scalar override replaces reasoning");
104
+ assert(out.contextWindow === 65536, "scalar override replaces contextWindow");
105
+ }
106
+
107
+ {
108
+ const base = { id: GLM52, cost: { input: 1, output: 2, cacheRead: 0.2, cacheWrite: 0 } } as any;
109
+ // cost is a plain object -> recursively deep-merged (not on the ModelOverride
110
+ // type, so exercise it via a loose cast).
111
+ const out = applyModelOverride(base, { cost: { input: 5 } } as any);
112
+ eq(out.cost, { input: 5, output: 2, cacheRead: 0.2, cacheWrite: 0 }, "deep-merge cost overrides one field, keeps the rest");
113
+ }
114
+
115
+ {
116
+ // Overriding a { $var } schema object wholesale REPLACES it (no deep-merge into $var)
117
+ const base = { id: KIMI, compat: { chatTemplateKwargs: { thinking: { $var: "thinking.enabled" }, preserve_thinking: true } } } as any;
118
+ const out = applyModelOverride(base, { compat: { chatTemplateKwargs: { thinking: { $var: "thinking.effort" } } } } as any);
119
+ eq(out.compat.chatTemplateKwargs.thinking, { $var: "thinking.effort" }, "overriding a { $var } object replaces it (no merge into $var)");
120
+ assert(out.compat.chatTemplateKwargs.preserve_thinking === true, "sibling chatTemplateKwargs key survives a $var replacement");
121
+ }
122
+
123
+ {
124
+ // Overriding an array (input) replaces it wholesale (no index-wise merge)
125
+ const base = { id: KIMI, input: ["text", "image"] } as any;
126
+ const out = applyModelOverride(base, { input: ["text"] } as any);
127
+ eq(out.input, ["text"], "array override replaces wholesale (no index-wise merge)");
128
+ }
129
+
130
+ // ─── parseModelOverrides ───────────────────────────────────────────────────────
131
+
132
+ console.log("\n--- parseModelOverrides ---");
133
+
134
+ eq(parseModelOverrides(undefined), undefined, "undefined input -> undefined");
135
+ eq(parseModelOverrides("nope"), undefined, "non-object input -> undefined");
136
+ eq(parseModelOverrides({}), undefined, "empty object -> undefined (no valid entries)");
137
+ eq(parseModelOverrides({ badId: "not-an-object" } as any), undefined, "non-object override value dropped -> undefined");
138
+ eq(parseModelOverrides({ id: 123 } as any), undefined, "non-object override (number) dropped -> undefined");
139
+ {
140
+ const r = parseModelOverrides({
141
+ [KIMI]: { compat: { chatTemplateKwargs: { preserve_thinking: false } } },
142
+ bad: 123,
143
+ alsobad: "string",
144
+ });
145
+ assert(r !== undefined && Object.keys(r!).length === 1 && r![KIMI] !== undefined, "keeps valid override, drops invalid ids");
146
+ assert((r![KIMI] as any).compat.chatTemplateKwargs.preserve_thinking === false, "valid override compat preserved");
147
+ }
148
+ {
149
+ // thinkingLevelMap: only string/null values kept; non-string values dropped
150
+ const r = parseModelOverrides({ [GLM52]: { thinkingLevelMap: { high: "max", bad: 42, off: null } } });
151
+ eq(r, { [GLM52]: { thinkingLevelMap: { high: "max", off: null } } }, "thinkingLevelMap keeps string/null values, drops others");
152
+ }
153
+ {
154
+ // an override whose fields all parse to nothing is dropped
155
+ const r = parseModelOverrides({ [KIMI]: { thinkingLevelMap: { bad: 42 } } });
156
+ eq(r, undefined, "override with no usable fields is dropped -> undefined");
157
+ }
158
+
159
+ // ─── loadConfig ────────────────────────────────────────────────────────────────
160
+
161
+ console.log("\n--- loadConfig ---");
162
+
163
+ const cfgPath = path.join(os.homedir(), ".pi", "agent", "extensions", "lilac.json");
164
+
165
+ {
166
+ // Fresh tmpHome: no config file -> loadConfig auto-populates the scaffold and returns defaults
167
+ assert(!fs.existsSync(cfgPath), "scaffold not present before first loadConfig");
168
+ const cfg = loadConfig();
169
+ eq(cfg, { modelOverrides: {} }, "missing file -> defaults (empty modelOverrides)");
170
+ assert(fs.existsSync(cfgPath), "loadConfig auto-populates the scaffold file on missing file");
171
+ eq(JSON.parse(fs.readFileSync(cfgPath, "utf8")), { modelOverrides: {} }, "scaffold file contains the default shape");
172
+ }
173
+
174
+ {
175
+ // Existing file with valid modelOverrides -> parsed
176
+ fs.writeFileSync(cfgPath, JSON.stringify({
177
+ modelOverrides: { [KIMI]: { compat: { chatTemplateKwargs: { preserve_thinking: false } } } },
178
+ }));
179
+ const cfg = loadConfig();
180
+ assert((cfg.modelOverrides as any)?.[KIMI]?.compat?.chatTemplateKwargs?.preserve_thinking === false, "existing file's modelOverrides parsed");
181
+ }
182
+
183
+ {
184
+ // Existing file WITHOUT a modelOverrides key -> undefined overrides (not an error)
185
+ fs.writeFileSync(cfgPath, JSON.stringify({ unrelatedKey: true }));
186
+ const cfg = loadConfig();
187
+ assert(cfg.modelOverrides === undefined, "file without modelOverrides key -> undefined overrides");
188
+ assert(JSON.parse(fs.readFileSync(cfgPath, "utf8")).unrelatedKey === true, "file without modelOverrides key is not rewritten");
189
+ }
190
+
191
+ {
192
+ // Existing file with invalid JSON -> defaults returned, file left UNTOUCHED (typo not wiped)
193
+ fs.writeFileSync(cfgPath, "not json {{{");
194
+ const cfg = loadConfig();
195
+ eq(cfg, { modelOverrides: {} }, "invalid JSON -> defaults");
196
+ assert(fs.readFileSync(cfgPath, "utf8") === "not json {{{", "invalid file is not overwritten (typo preserved)");
197
+ }
198
+
199
+ // ─── buildModels end-to-end (real models.json + patch.json) ────────────────────
200
+
201
+ console.log("\n--- buildModels end-to-end ---");
202
+
203
+ function find(models: any[], id: string): any {
204
+ const m = models.find((x) => x.id === id);
205
+ if (!m) throw new Error(`model ${id} not built`);
206
+ return m;
207
+ }
208
+
209
+ {
210
+ // No overrides -> identical to today: kimi preserve_thinking: true survives patch
211
+ const models = buildModels(modelsData, customModelsData, patchData, {});
212
+ const kimi = find(models, KIMI);
213
+ assert(kimi.compat.chatTemplateKwargs.preserve_thinking === true, "no overrides -> kimi preserve_thinking stays true (patch wins)");
214
+ assert(kimi.compat.thinkingFormat === "chat-template", "no overrides -> kimi thinkingFormat intact");
215
+ const glm52 = find(models, GLM52);
216
+ assert(glm52.compat.chatTemplateKwargs.clear_thinking === false, "no overrides -> glm-5.2 clear_thinking stays false (patch wins)");
217
+ }
218
+
219
+ {
220
+ // User override wins over patch.json: disable preserve_thinking on kimi
221
+ const overrides = { [KIMI]: { compat: { chatTemplateKwargs: { preserve_thinking: false } } } } as any;
222
+ const models = buildModels(modelsData, customModelsData, patchData, overrides);
223
+ const kimi = find(models, KIMI);
224
+ assert(kimi.compat.chatTemplateKwargs.preserve_thinking === false, "override wins over patch: kimi preserve_thinking -> false");
225
+ // deep-merge: the $var thinking keys + thinkingFormat survive
226
+ assert((kimi.compat.chatTemplateKwargs as any).thinking?.$var === "thinking.enabled", "override deep-merges: thinking $var key survives");
227
+ assert((kimi.compat.chatTemplateKwargs as any).enable_thinking?.$var === "thinking.enabled", "override deep-merges: enable_thinking $var key survives");
228
+ assert(kimi.compat.thinkingFormat === "chat-template", "override deep-merges: thinkingFormat survives");
229
+ assert(kimi.compat.supportsDeveloperRole === false, "override deep-merges: supportsDeveloperRole survives");
230
+ }
231
+
232
+ {
233
+ // Override a single thinking level on glm-5.2 without redeclaring the map;
234
+ // the patch-applied clear_thinking flag survives a thinkingLevelMap-only override
235
+ const overrides = { [GLM52]: { thinkingLevelMap: { high: "max" } } } as any;
236
+ const models = buildModels(modelsData, customModelsData, patchData, overrides);
237
+ const glm = find(models, GLM52);
238
+ assert((glm.thinkingLevelMap as any)?.high === "max", "override thinkingLevelMap.high wins over patch");
239
+ assert(glm.compat.chatTemplateKwargs.clear_thinking === false, "non-overridden clear_thinking survives a thinkingLevelMap-only override");
240
+ }
241
+
242
+ {
243
+ // Override on glm-5.1 toggles clear_thinking (patch sets false); other compat survives
244
+ const overrides = { [GLM51]: { compat: { chatTemplateKwargs: { clear_thinking: true } } } } as any;
245
+ const models = buildModels(modelsData, customModelsData, patchData, overrides);
246
+ const glm = find(models, GLM51);
247
+ assert(glm.compat.chatTemplateKwargs.clear_thinking === true, "override wins over patch: glm-5.1 clear_thinking -> true");
248
+ assert((glm.compat.chatTemplateKwargs as any).thinking?.$var === "thinking.enabled", "override deep-merges: glm-5.1 thinking $var key survives");
249
+ assert(glm.compat.zaiToolStream === true, "override deep-merges: glm-5.1 zaiToolStream survives");
250
+ }
251
+
252
+ {
253
+ // Override for an unknown id is a no-op (adds no models)
254
+ const before = buildModels(modelsData, customModelsData, patchData, {});
255
+ const after = buildModels(modelsData, customModelsData, patchData, { "no/such-model": { reasoning: false } } as any);
256
+ eq(after.length, before.length, "override for an unknown id adds no models");
257
+ }
258
+
259
+ // ─── Summary ───────────────────────────────────────────────────────────────────
260
+
261
+ console.log(`\n${failed === 0 ? "ALL PASS" : `${failed} FAILED`}`);
262
+ process.exit(failed === 0 ? 0 : 1);
@@ -0,0 +1,237 @@
1
+ #!/usr/bin/env node
2
+ /**
3
+ * Wire-level test for preserved-thinking (full-history reasoning) flags.
4
+ *
5
+ * Verifies, against the REAL pi-ai streamSimple + a stubbed fetch, that the
6
+ * `chat_template_kwargs` each Lilac reasoning model puts on the wire include the
7
+ * preservation flags configured in patch.json — and that those static flags
8
+ * coexist with the { $var }-resolved thinking keys at every thinking level.
9
+ *
10
+ * Background: reasoning models trim older assistant reasoning across turns by
11
+ * default (each vendor's template default). Two template-level flags opt into
12
+ * full-history preservation (confirmed in each model's HuggingFace chat
13
+ * template; Kimi K2.6 and GLM 5.2 additionally E2E-verified on the sibling
14
+ * neuralwatt provider's 3-turn / two-20-digit-number recall test):
15
+ *
16
+ * Kimi K2.6 → preserve_thinking: true (template: "if preserve_thinking, keep -1 ... retain reasoning")
17
+ * GLM 5.1 → clear_thinking: false (same clear_thinking mechanism as GLM 5.2)
18
+ * GLM 5.2 → clear_thinking: false (template: keep reasoning when "clear_thinking is defined and not clear_thinking")
19
+ *
20
+ * Lilac uses pi-ai's `chat-template` thinkingFormat, which calls
21
+ * buildChatTemplateKwargs → resolveChatTemplateKwargValue. Static primitive
22
+ * kwargs pass through verbatim; only { $var } objects are resolved against the
23
+ * turn's thinking state. So the preserve flags ride onto the wire as plain
24
+ * booleans next to the { $var } thinking/enable_thinking keys — no onPayload
25
+ * hook needed (unlike neuralwatt, which drives the openai reasoning_effort path
26
+ * and must inject via onPayload because the two paths are mutually exclusive).
27
+ *
28
+ * Gemma 4 and MiniMax M2.7/M3 expose NO family-wide preserve flag (their HF
29
+ * templates read only enable_thinking / thinking_mode + reasoning_content), so
30
+ * no flag is added for them here — asserted as a regression guard.
31
+ *
32
+ * Run: node scripts/test-preserved-thinking.ts
33
+ */
34
+ import fs from "fs";
35
+ import path from "path";
36
+ import { pathToFileURL } from "url";
37
+
38
+ const HERE = path.dirname(new URL(import.meta.url).pathname);
39
+
40
+ // ─── Faithful replica of index.ts applyPatch / buildModels ────────────────────
41
+ // Mirrors the provider's model pipeline so the test exercises the same objects
42
+ // the runtime registers. (Kept inline so the test has no TS-import dependency
43
+ // on index.ts, which pulls in pi-coding-agent types.)
44
+
45
+ function applyPatch(model, patch) {
46
+ const result = { ...model };
47
+ if (patch.name !== undefined) result.name = patch.name;
48
+ if (patch.reasoning !== undefined) result.reasoning = patch.reasoning;
49
+ if (patch.input !== undefined) result.input = patch.input;
50
+ if (patch.contextWindow !== undefined) result.contextWindow = patch.contextWindow;
51
+ if (patch.maxTokens !== undefined) result.maxTokens = patch.maxTokens;
52
+ if (patch.cost) {
53
+ result.cost = {
54
+ input: patch.cost.input ?? result.cost.input,
55
+ output: patch.cost.output ?? result.cost.output,
56
+ cacheRead: patch.cost.cacheRead ?? result.cost.cacheRead,
57
+ cacheWrite: patch.cost.cacheWrite ?? result.cost.cacheWrite,
58
+ };
59
+ }
60
+ if (patch.compat) result.compat = { ...(result.compat || {}), ...patch.compat };
61
+ if (patch.thinkingLevelMap !== undefined) result.thinkingLevelMap = patch.thinkingLevelMap;
62
+ if (!result.reasoning && result.compat?.thinkingFormat) delete result.compat.thinkingFormat;
63
+ if (!result.reasoning && result.thinkingLevelMap) delete result.thinkingLevelMap;
64
+ if (result.compat && Object.keys(result.compat).length === 0) delete result.compat;
65
+ return result;
66
+ }
67
+
68
+ function buildModels(base, custom, patch) {
69
+ const map = new Map();
70
+ for (const m of base) map.set(m.id, m);
71
+ for (const [id, p] of Object.entries(patch)) {
72
+ const ex = map.get(id);
73
+ if (ex) map.set(id, applyPatch(ex, p));
74
+ }
75
+ for (const m of custom) {
76
+ const ex = map.get(m.id);
77
+ const p = patch[m.id];
78
+ if (ex && p) map.set(m.id, applyPatch(m, p));
79
+ else if (ex) map.set(m.id, m);
80
+ else if (p) map.set(m.id, applyPatch(m, p));
81
+ else map.set(m.id, m);
82
+ }
83
+ return Array.from(map.values());
84
+ }
85
+
86
+ // ─── Load REAL pi-ai streamSimple from the global pi install ──────────────────
87
+ // Imported by absolute path so its relative deps (openai, ../models.js, ...) and
88
+ // the `openai` SDK resolve from the global pi node_modules tree.
89
+ const PI_AI_API = path.join(
90
+ os_home_pi_ai(),
91
+ "dist/api/openai-completions.js",
92
+ );
93
+ function os_home_pi_ai() {
94
+ // Resolve the pi-ai package shipped inside the globally-installed pi agent.
95
+ const candidates = [
96
+ "/Users/monotykamary/.npm-global/lib/node_modules/@earendil-works/pi-coding-agent/node_modules/@earendil-works/pi-ai",
97
+ ];
98
+ for (const c of candidates) if (fs.existsSync(c)) return c;
99
+ throw new Error("Could not locate pi-ai package in the global pi install.");
100
+ }
101
+
102
+ const { streamSimple } = await import(pathToFileURL(PI_AI_API).href);
103
+
104
+ // ─── Build the exact model objects the provider registers ────────────────────
105
+ const embedded = JSON.parse(fs.readFileSync(path.join(HERE, "..", "models.json"), "utf8"));
106
+ const custom = JSON.parse(fs.readFileSync(path.join(HERE, "..", "custom-models.json"), "utf8"));
107
+ const patch = JSON.parse(fs.readFileSync(path.join(HERE, "..", "patch.json"), "utf8"));
108
+ const models = new Map(buildModels(embedded, custom, patch).map((m) => [m.id, m]));
109
+
110
+ // ─── Stub fetch to capture the request body (fires before response parsing) ──
111
+ const captured = [];
112
+ const originalFetch = globalThis.fetch;
113
+ globalThis.fetch = async (url, init) => {
114
+ captured.push({ url: String(url), body: init?.body ?? null });
115
+ // Minimal SSE response that ends immediately — enough for the OpenAI SDK to
116
+ // construct the stream; we only need the request body, captured above.
117
+ return new Response(new ReadableStream({ start(c) { c.close(); } }), {
118
+ headers: { "content-type": "text/event-stream" },
119
+ });
120
+ };
121
+
122
+ async function wire(modelId, reasoning) {
123
+ const model = models.get(modelId);
124
+ if (!model) throw new Error(`unknown model: ${modelId}`);
125
+ captured.length = 0;
126
+ const ctx = { messages: [{ role: "user", content: [{ type: "text", text: "hi" }] }] };
127
+ const s = streamSimple(
128
+ { ...model, provider: "lilac", api: "openai-completions", baseUrl: "https://api.getlilac.com/v1" },
129
+ ctx,
130
+ { apiKey: "sk-test", reasoning },
131
+ );
132
+ // The fetch fires inside stream()'s async IIFE; poll for it.
133
+ const deadline = Date.now() + 2000;
134
+ while (captured.length === 0 && Date.now() < deadline) await new Promise((r) => setTimeout(r, 20));
135
+ try { s.end?.(); } catch {}
136
+ const hit = captured.find((c) => c.url.includes("/chat/completions"));
137
+ if (!hit) throw new Error(`no /chat/completions request captured for ${modelId} (${reasoning})`);
138
+ return JSON.parse(hit.body);
139
+ }
140
+
141
+ // ─── Assertions ───────────────────────────────────────────────────────────────
142
+ let failures = 0;
143
+ function eq(actual, expected, msg) {
144
+ const a = JSON.stringify(actual);
145
+ const e = JSON.stringify(expected);
146
+ const ok = a === e;
147
+ console.log(`${ok ? "✓" : "✗"} ${msg}`);
148
+ if (!ok) {
149
+ failures++;
150
+ console.log(` expected ${e}`);
151
+ console.log(` actual ${a}`);
152
+ }
153
+ }
154
+ function truthy(actual, msg) {
155
+ const ok = !!actual;
156
+ console.log(`${ok ? "✓" : "✗"} ${msg}`);
157
+ if (!ok) {
158
+ failures++;
159
+ console.log(` expected truthy, got ${JSON.stringify(actual)}`);
160
+ }
161
+ }
162
+ function falsy(actual, msg) {
163
+ const ok = !actual;
164
+ console.log(`${ok ? "✓" : "✗"} ${msg}`);
165
+ if (!ok) {
166
+ failures++;
167
+ console.log(` expected falsy/undefined, got ${JSON.stringify(actual)}`);
168
+ }
169
+ }
170
+
171
+ console.log("\n=== patch.json data ===");
172
+ const kimiPatch = patch["moonshotai/kimi-k2.6"]?.compat?.chatTemplateKwargs;
173
+ eq(kimiPatch?.preserve_thinking, true, "kimi-k2.6 patch sets preserve_thinking: true");
174
+ const glm51Patch = patch["zai-org/glm-5.1"]?.compat?.chatTemplateKwargs;
175
+ eq(glm51Patch?.clear_thinking, false, "glm-5.1 patch sets clear_thinking: false");
176
+ const glm52Patch = patch["zai-org/glm-5.2"]?.compat?.chatTemplateKwargs;
177
+ eq(glm52Patch?.clear_thinking, false, "glm-5.2 patch sets clear_thinking: false");
178
+
179
+ console.log("\n=== Kimi K2.6 on the wire (real pi-ai) ===");
180
+ {
181
+ const high = await wire("moonshotai/kimi-k2.6", "high");
182
+ eq(high.chat_template_kwargs, { thinking: true, enable_thinking: true, preserve_thinking: true },
183
+ "kimi @ high → thinking+enable_thinking true AND preserve_thinking true");
184
+ const off = await wire("moonshotai/kimi-k2.6", "off");
185
+ eq(off.chat_template_kwargs, { thinking: false, enable_thinking: false, preserve_thinking: true },
186
+ "kimi @ off → thinking false but preserve_thinking still true (level-independent)");
187
+ const minimal = await wire("moonshotai/kimi-k2.6", "minimal");
188
+ eq(minimal.chat_template_kwargs, { thinking: true, enable_thinking: true, preserve_thinking: true },
189
+ "kimi @ minimal → preserve_thinking present at every level");
190
+ }
191
+
192
+ console.log("\n=== GLM 5.2 on the wire (real pi-ai) ===");
193
+ {
194
+ const high = await wire("zai-org/glm-5.2", "high");
195
+ eq(high.chat_template_kwargs, { enable_thinking: true, reasoning_effort: "high", clear_thinking: false },
196
+ "glm-5.2 @ high → enable_thinking+reasoning_effort AND clear_thinking false");
197
+ const xhigh = await wire("zai-org/glm-5.2", "xhigh");
198
+ eq(xhigh.chat_template_kwargs, { enable_thinking: true, reasoning_effort: "max", clear_thinking: false },
199
+ "glm-5.2 @ xhigh → reasoning_effort max, clear_thinking false");
200
+ const off = await wire("zai-org/glm-5.2", "off");
201
+ eq(off.chat_template_kwargs, { enable_thinking: false, clear_thinking: false },
202
+ "glm-5.2 @ off → reasoning_effort omitted (omitWhenOff), clear_thinking false persists");
203
+ }
204
+
205
+ console.log("\n=== GLM 5.1 on the wire (real pi-ai) ===");
206
+ {
207
+ const high = await wire("zai-org/glm-5.1", "high");
208
+ eq(high.chat_template_kwargs, { thinking: true, enable_thinking: true, clear_thinking: false },
209
+ "glm-5.1 @ high → thinking+enable_thinking true AND clear_thinking false");
210
+ const off = await wire("zai-org/glm-5.1", "off");
211
+ eq(off.chat_template_kwargs, { thinking: false, enable_thinking: false, clear_thinking: false },
212
+ "glm-5.1 @ off → thinking false but clear_thinking false persists");
213
+ }
214
+
215
+ console.log("\n=== Gemma 4 / MiniMax (no family-wide preserve flag — regression guard) ===");
216
+ {
217
+ const gemma = await wire("google/gemma-4-31b-it", "high");
218
+ eq(gemma.chat_template_kwargs, { thinking: true, enable_thinking: true },
219
+ "gemma-4 @ high → only thinking/enable_thinking (no preserve/clear flag)");
220
+ falsy(gemma.chat_template_kwargs?.preserve_thinking, "gemma-4 has no preserve_thinking");
221
+ falsy(gemma.chat_template_kwargs?.clear_thinking, "gemma-4 has no clear_thinking");
222
+
223
+ const m3 = await wire("minimaxai/minimax-m3", "high");
224
+ eq(m3.chat_template_kwargs, { thinking_mode: "enabled" },
225
+ "minimax-m3 @ high → only thinking_mode (no preserve/clear flag)");
226
+ falsy(m3.chat_template_kwargs?.preserve_thinking, "minimax-m3 has no preserve_thinking");
227
+
228
+ const m27 = await wire("minimaxai/minimax-m2.7", "high");
229
+ eq(m27.chat_template_kwargs, { thinking: true, enable_thinking: true },
230
+ "minimax-m2.7 @ high → only thinking/enable_thinking (no preserve/clear flag)");
231
+ falsy(m27.chat_template_kwargs?.clear_thinking, "minimax-m2.7 has no clear_thinking");
232
+ }
233
+
234
+ globalThis.fetch = originalFetch;
235
+
236
+ console.log(`\n${failures === 0 ? "ALL PASS" : `${failures} FAILURE(S)`}`);
237
+ if (failures > 0) process.exit(1);