pi-lilac-provider 1.4.1 → 1.6.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +39 -1
- package/index.ts +139 -8
- package/package.json +5 -2
- package/patch.json +6 -3
- package/scripts/test-discounts.ts +7 -7
- package/scripts/test-model-overrides.ts +262 -0
- package/scripts/test-preserved-thinking.ts +237 -0
package/README.md
CHANGED
|
@@ -128,6 +128,23 @@ for M3's adaptive "model decides" mode. (The selector/footer show pi's level
|
|
|
128
128
|
names — `minimal`/`high` — not the `thinking_mode` values; pi has no per-model
|
|
129
129
|
level-relabel hook.)
|
|
130
130
|
|
|
131
|
+
**Preserved thinking (full-history reasoning).** By default these templates
|
|
132
|
+
trim older assistant reasoning between turns (each vendor's default), which
|
|
133
|
+
degrades multi-turn recall. Three models opt into full-history preservation via
|
|
134
|
+
a template flag sent alongside the reasoning key:
|
|
135
|
+
|
|
136
|
+
| Model | Flag | Effect |
|
|
137
|
+
|-------|------|--------|
|
|
138
|
+
| Kimi K2.6 | `preserve_thinking: true` | keeps every assistant turn's reasoning (default: only the last) |
|
|
139
|
+
| GLM 5.1 | `clear_thinking: false` | keeps reasoning for all turns (default: clears before the last user message) |
|
|
140
|
+
| GLM 5.2 | `clear_thinking: false` | keeps reasoning for all turns (default: clears before the last user message) |
|
|
141
|
+
|
|
142
|
+
Kimi K2.6 and GLM 5.2 are E2E-verified on the sibling neuralwatt provider via a
|
|
143
|
+
3-turn, two-20-digit-number recall test (Kimi 0/6 → 6/6, GLM 5.2 1/4 → 4/4);
|
|
144
|
+
GLM 5.1 uses the same `clear_thinking` mechanism (confirmed in its HuggingFace
|
|
145
|
+
chat template). Gemma 4 and MiniMax M2.7/M3 expose no family-wide preserve flag,
|
|
146
|
+
so their older assistant reasoning is trimmed per the template default.
|
|
147
|
+
|
|
131
148
|
In pi, reasoning models automatically use the appropriate thinking format. Use
|
|
132
149
|
Shift+Tab to control thinking level.
|
|
133
150
|
|
|
@@ -173,7 +190,7 @@ Add to your pi configuration for automatic loading:
|
|
|
173
190
|
|
|
174
191
|
Lilac's API is OpenAI-compatible with these specifics:
|
|
175
192
|
|
|
176
|
-
- **`thinkingFormat: "chat-template"`** — All reasoning models. Lilac's vLLM backend toggles reasoning via `chat_template_kwargs`, but the honored key differs per model family. Per-model `chatTemplateKwargs` in `patch.json` send the right key(s): `thinking`+`enable_thinking` (bool) for Kimi K2.6, GLM 5.1, Gemma 4, and MiniMax M2.7; `enable_thinking` + `reasoning_effort` for GLM 5.2; `thinking_mode` (adaptive|enabled|disabled) for MiniMax M3.
|
|
193
|
+
- **`thinkingFormat: "chat-template"`** — All reasoning models. Lilac's vLLM backend toggles reasoning via `chat_template_kwargs`, but the honored key differs per model family. Per-model `chatTemplateKwargs` in `patch.json` send the right key(s): `thinking`+`enable_thinking` (bool) for Kimi K2.6, GLM 5.1, Gemma 4, and MiniMax M2.7; `enable_thinking` + `reasoning_effort` for GLM 5.2; `thinking_mode` (adaptive|enabled|disabled) for MiniMax M3. Kimi K2.6, GLM 5.1, and GLM 5.2 additionally send a preservation flag (`preserve_thinking: true` / `clear_thinking: false`) to retain full reasoning history across turns — see [Preserved thinking](#thinking-mode) above. Override these per-model via [Model Overrides](#model-overrides).
|
|
177
194
|
- **`maxTokensField: "max_completion_tokens"`** — All models. Lilac supports `max_completion_tokens` (preferred for reasoning models as it includes reasoning tokens).
|
|
178
195
|
- **`supportsDeveloperRole: true`** — All models. Lilac's vLLM backend maps the developer role to system.
|
|
179
196
|
- **`supportsStore: false`** — All models. Lilac doesn't support the `store` parameter.
|
|
@@ -193,6 +210,27 @@ The `patch.json` file contains overrides that are applied on top of `models.json
|
|
|
193
210
|
- Adding compat settings that the API doesn't provide
|
|
194
211
|
- Overriding pricing when official rates change
|
|
195
212
|
|
|
213
|
+
### Model Overrides
|
|
214
|
+
|
|
215
|
+
`modelOverrides` lets you override compat flags and other model properties per model id, **on top of** `patch.json` + `custom-models.json`, without editing the extension. Keyed by model id; `compat` (including nested `chatTemplateKwargs`), `thinkingLevelMap`, and `cost` are deep-merged **recursively** (toggle one flag without redeclaring the rest), scalars and arrays are replaced. Applied at session start, so edits take effect on the next `pi` session.
|
|
216
|
+
|
|
217
|
+
Create `~/.pi/agent/extensions/lilac.json` (auto-populated with defaults on first run):
|
|
218
|
+
|
|
219
|
+
```jsonc
|
|
220
|
+
{
|
|
221
|
+
"modelOverrides": {
|
|
222
|
+
// Disable full-history reasoning for kimi-k2.6 (e.g. to save tokens):
|
|
223
|
+
"moonshotai/kimi-k2.6": { "compat": { "chatTemplateKwargs": { "preserve_thinking": false } } },
|
|
224
|
+
// Toggle the GLM 5.2 clear_thinking flag without redeclaring the rest of compat:
|
|
225
|
+
"zai-org/glm-5.2": { "compat": { "chatTemplateKwargs": { "clear_thinking": true } } },
|
|
226
|
+
// Override a single thinking level without redeclaring the whole map:
|
|
227
|
+
"zai-org/glm-5.1": { "thinkingLevelMap": { "high": "max" } }
|
|
228
|
+
}
|
|
229
|
+
}
|
|
230
|
+
```
|
|
231
|
+
|
|
232
|
+
The full set of overridable fields matches the model schema (`compat`, `thinkingLevelMap`, `cost`, `contextWindow`, `maxTokens`, `reasoning`, `input`). See [Compat Settings](#compat-settings) for the catalog of compat flags and what `chatTemplateKwargs` values mean per family. An invalid JSON file is left untouched (defaults are used) so a typo isn't silently wiped — fix the file and restart pi.
|
|
233
|
+
|
|
196
234
|
## Updating Models
|
|
197
235
|
|
|
198
236
|
Run the update script to fetch the latest models from Lilac's API:
|
package/index.ts
CHANGED
|
@@ -178,6 +178,120 @@ interface PatchEntry {
|
|
|
178
178
|
|
|
179
179
|
type PatchData = Record<string, PatchEntry>;
|
|
180
180
|
|
|
181
|
+
// User Configuration: ~/.pi/agent/extensions/lilac.json lets a user override model
|
|
182
|
+
// properties per id ON TOP of patch.json + custom-models.json (so they win).
|
|
183
|
+
// Recursively deep-merges `compat` (incl. nested `chatTemplateKwargs`),
|
|
184
|
+
// `thinkingLevelMap`, and `cost` (toggle one flag without redeclaring the rest),
|
|
185
|
+
// replaces scalars and arrays. Lets a user toggle chat_template_kwargs (e.g.
|
|
186
|
+
// preserve_thinking / clear_thinking) or a single thinking level without editing
|
|
187
|
+
// the extension. See README "Model Overrides".
|
|
188
|
+
interface ModelOverride {
|
|
189
|
+
thinkingLevelMap?: ThinkingLevelMap;
|
|
190
|
+
compat?: Record<string, unknown>;
|
|
191
|
+
}
|
|
192
|
+
|
|
193
|
+
interface LilacConfig {
|
|
194
|
+
modelOverrides?: Record<string, ModelOverride>;
|
|
195
|
+
}
|
|
196
|
+
|
|
197
|
+
const CONFIG_PATH = path.join(os.homedir(), ".pi", "agent", "extensions", "lilac.json");
|
|
198
|
+
const DEFAULT_CONFIG: LilacConfig = { modelOverrides: {} };
|
|
199
|
+
|
|
200
|
+
// Validate user-supplied modelOverrides from the config file. Non-object ids and
|
|
201
|
+
// non-object overrides are dropped silently so a malformed file doesn't crash
|
|
202
|
+
// model registration.
|
|
203
|
+
function parseModelOverrides(raw: unknown): Record<string, ModelOverride> | undefined {
|
|
204
|
+
if (!raw || typeof raw !== "object" || Array.isArray(raw)) return undefined;
|
|
205
|
+
const result: Record<string, ModelOverride> = {};
|
|
206
|
+
for (const [id, override] of Object.entries(raw as Record<string, unknown>)) {
|
|
207
|
+
if (typeof id !== "string" || !override || typeof override !== "object" || Array.isArray(override)) continue;
|
|
208
|
+
const o = override as Record<string, unknown>;
|
|
209
|
+
const parsed: ModelOverride = {};
|
|
210
|
+
if (o.thinkingLevelMap && typeof o.thinkingLevelMap === "object") {
|
|
211
|
+
const m: Record<string, string | null> = {};
|
|
212
|
+
for (const [k, v] of Object.entries(o.thinkingLevelMap as Record<string, unknown>)) {
|
|
213
|
+
if (v === null || typeof v === "string") m[k] = v;
|
|
214
|
+
}
|
|
215
|
+
if (Object.keys(m).length > 0) parsed.thinkingLevelMap = m as ThinkingLevelMap;
|
|
216
|
+
}
|
|
217
|
+
if (o.compat && typeof o.compat === "object") parsed.compat = o.compat as Record<string, unknown>;
|
|
218
|
+
if (Object.keys(parsed).length > 0) result[id] = parsed;
|
|
219
|
+
}
|
|
220
|
+
return Object.keys(result).length > 0 ? result : undefined;
|
|
221
|
+
}
|
|
222
|
+
|
|
223
|
+
// Reads ~/.pi/agent/extensions/lilac.json. Missing file → populate with defaults
|
|
224
|
+
// so the user can discover it, then return defaults. An existing-but-invalid file
|
|
225
|
+
// is left untouched (defaults returned) so a user's typo isn't silently wiped —
|
|
226
|
+
// they fix the file and restart pi. Loaded lazily on first use (not at import) so
|
|
227
|
+
// importing the module has no filesystem side effects and unit tests can import
|
|
228
|
+
// the pure helpers safely.
|
|
229
|
+
function loadConfig(): LilacConfig {
|
|
230
|
+
let rawText: string;
|
|
231
|
+
try {
|
|
232
|
+
rawText = fs.readFileSync(CONFIG_PATH, "utf8");
|
|
233
|
+
} catch {
|
|
234
|
+
try {
|
|
235
|
+
fs.mkdirSync(path.dirname(CONFIG_PATH), { recursive: true });
|
|
236
|
+
fs.writeFileSync(CONFIG_PATH, JSON.stringify(DEFAULT_CONFIG, null, 2) + "\n");
|
|
237
|
+
} catch {
|
|
238
|
+
// Write failure is non-fatal — defaults still work in memory
|
|
239
|
+
}
|
|
240
|
+
return { ...DEFAULT_CONFIG };
|
|
241
|
+
}
|
|
242
|
+
try {
|
|
243
|
+
const raw = JSON.parse(rawText);
|
|
244
|
+
return { modelOverrides: parseModelOverrides(raw.modelOverrides) };
|
|
245
|
+
} catch {
|
|
246
|
+
// File exists but is invalid JSON — return defaults WITHOUT overwriting.
|
|
247
|
+
return { ...DEFAULT_CONFIG };
|
|
248
|
+
}
|
|
249
|
+
}
|
|
250
|
+
|
|
251
|
+
// Recursively deep-merge `override` into `base`. Plain objects are merged
|
|
252
|
+
// key-by-key (so a user can toggle a single chatTemplateKwargs flag without
|
|
253
|
+
// redeclaring the rest, and a single compat flag without redeclaring
|
|
254
|
+
// chatTemplateKwargs); arrays and non-plain-object values replace the base value
|
|
255
|
+
// (so an overridden { $var } schema object, or an `input` array, replaces wholesale).
|
|
256
|
+
function isPlainObject(v: unknown): v is Record<string, unknown> {
|
|
257
|
+
return v !== null && typeof v === "object" && !Array.isArray(v);
|
|
258
|
+
}
|
|
259
|
+
|
|
260
|
+
function deepMerge<T>(base: T, override: unknown): T {
|
|
261
|
+
if (!isPlainObject(override) || !isPlainObject(base)) return override as T;
|
|
262
|
+
const result: Record<string, unknown> = { ...(base as Record<string, unknown>) };
|
|
263
|
+
for (const [k, v] of Object.entries(override)) {
|
|
264
|
+
result[k] = isPlainObject(v) && isPlainObject(result[k]) ? deepMerge(result[k], v) : v;
|
|
265
|
+
}
|
|
266
|
+
return result as T;
|
|
267
|
+
}
|
|
268
|
+
|
|
269
|
+
// Apply a user-supplied modelOverride (from lilac.json) on top of a built model.
|
|
270
|
+
// Recursively deep-merges compat (incl. nested chatTemplateKwargs) /
|
|
271
|
+
// thinkingLevelMap / cost so a user can toggle a single flag (e.g.
|
|
272
|
+
// chatTemplateKwargs.preserve_thinking) without redeclaring the rest; replaces
|
|
273
|
+
// scalars and arrays. No reasoning-cleanup (unlike applyPatch) — the override is
|
|
274
|
+
// authoritative.
|
|
275
|
+
function applyModelOverride(model: JsonModel, override: ModelOverride): JsonModel {
|
|
276
|
+
const result = { ...model };
|
|
277
|
+
for (const [key, value] of Object.entries(override)) {
|
|
278
|
+
(result as any)[key] = isPlainObject(value) && isPlainObject((result as any)[key])
|
|
279
|
+
? deepMerge((result as any)[key], value)
|
|
280
|
+
: value;
|
|
281
|
+
}
|
|
282
|
+
return result;
|
|
283
|
+
}
|
|
284
|
+
|
|
285
|
+
let config: LilacConfig | undefined;
|
|
286
|
+
function getConfig(): LilacConfig {
|
|
287
|
+
if (!config) config = loadConfig();
|
|
288
|
+
return config;
|
|
289
|
+
}
|
|
290
|
+
|
|
291
|
+
function activeOverrides(): Record<string, ModelOverride> {
|
|
292
|
+
return getConfig().modelOverrides ?? {};
|
|
293
|
+
}
|
|
294
|
+
|
|
181
295
|
// ─── Patch Application ────────────────────────────────────────────────────────
|
|
182
296
|
|
|
183
297
|
function applyPatch(model: JsonModel, patch: PatchEntry): JsonModel {
|
|
@@ -217,8 +331,13 @@ function applyPatch(model: JsonModel, patch: PatchEntry): JsonModel {
|
|
|
217
331
|
return result;
|
|
218
332
|
}
|
|
219
333
|
|
|
220
|
-
/** Full pipeline: base models → patch → custom → result */
|
|
221
|
-
function buildModels(
|
|
334
|
+
/** Full pipeline: base models → patch → custom → user modelOverrides → result */
|
|
335
|
+
function buildModels(
|
|
336
|
+
base: JsonModel[],
|
|
337
|
+
custom: JsonModel[],
|
|
338
|
+
patch: PatchData,
|
|
339
|
+
overrides: Record<string, ModelOverride> = {},
|
|
340
|
+
): JsonModel[] {
|
|
222
341
|
const modelMap = new Map<string, JsonModel>();
|
|
223
342
|
|
|
224
343
|
for (const model of base) {
|
|
@@ -246,6 +365,18 @@ function buildModels(base: JsonModel[], custom: JsonModel[], patch: PatchData):
|
|
|
246
365
|
}
|
|
247
366
|
}
|
|
248
367
|
|
|
368
|
+
// User-supplied modelOverrides (from ~/.pi/agent/extensions/lilac.json) applied
|
|
369
|
+
// LAST so they win over patch.json + custom-models.json. Recursively deep-merges
|
|
370
|
+
// compat (incl. chatTemplateKwargs) / thinkingLevelMap / cost so a user can
|
|
371
|
+
// toggle a single flag (e.g. chatTemplateKwargs.preserve_thinking) without
|
|
372
|
+
// redeclaring the rest.
|
|
373
|
+
for (const [id, override] of Object.entries(overrides)) {
|
|
374
|
+
const existing = modelMap.get(id);
|
|
375
|
+
if (existing) {
|
|
376
|
+
modelMap.set(id, applyModelOverride(existing, override));
|
|
377
|
+
}
|
|
378
|
+
}
|
|
379
|
+
|
|
249
380
|
return Array.from(modelMap.values());
|
|
250
381
|
}
|
|
251
382
|
|
|
@@ -617,14 +748,14 @@ export default function (pi: ExtensionAPI) {
|
|
|
617
748
|
// the in-flight model's cost without compounding an already-applied discount.
|
|
618
749
|
function getListModels(): JsonModel[] {
|
|
619
750
|
if (!listModelsCache) {
|
|
620
|
-
listModelsCache = buildModels(loadStaleModels(embeddedModels), customModels, patches);
|
|
751
|
+
listModelsCache = buildModels(loadStaleModels(embeddedModels), customModels, patches, activeOverrides());
|
|
621
752
|
}
|
|
622
753
|
return listModelsCache;
|
|
623
754
|
}
|
|
624
755
|
|
|
625
756
|
const staleBase = loadStaleModels(embeddedModels);
|
|
626
757
|
latestDiscounts = loadCachedDiscounts();
|
|
627
|
-
const staleModels = applyDiscounts(buildModels(staleBase, customModels, patches), latestDiscounts);
|
|
758
|
+
const staleModels = applyDiscounts(buildModels(staleBase, customModels, patches, activeOverrides()), latestDiscounts);
|
|
628
759
|
|
|
629
760
|
pi.registerProvider("lilac", {
|
|
630
761
|
baseUrl: BASE_URL,
|
|
@@ -731,14 +862,14 @@ export default function (pi: ExtensionAPI) {
|
|
|
731
862
|
baseUrl: BASE_URL,
|
|
732
863
|
apiKey: "$LILAC_API_KEY",
|
|
733
864
|
api: "openai-completions",
|
|
734
|
-
models: applyDiscounts(buildModels(merged, customModels, patches), latestDiscounts),
|
|
865
|
+
models: applyDiscounts(buildModels(merged, customModels, patches, activeOverrides()), latestDiscounts),
|
|
735
866
|
});
|
|
736
867
|
} else if (discounts) {
|
|
737
868
|
pi.registerProvider("lilac", {
|
|
738
869
|
baseUrl: BASE_URL,
|
|
739
870
|
apiKey: "$LILAC_API_KEY",
|
|
740
871
|
api: "openai-completions",
|
|
741
|
-
models: applyDiscounts(buildModels(staleBase, customModels, patches), latestDiscounts),
|
|
872
|
+
models: applyDiscounts(buildModels(staleBase, customModels, patches, activeOverrides()), latestDiscounts),
|
|
742
873
|
});
|
|
743
874
|
}
|
|
744
875
|
|
|
@@ -887,5 +1018,5 @@ export default function (pi: ExtensionAPI) {
|
|
|
887
1018
|
});
|
|
888
1019
|
}
|
|
889
1020
|
|
|
890
|
-
export { fetchStatusDiscounts, applyDiscounts, applyDiscountInPlace, loadCachedDiscounts, cacheDiscounts };
|
|
891
|
-
export type { JsonDiscount, JsonModel, PatchEntry, PatchData };
|
|
1021
|
+
export { fetchStatusDiscounts, applyDiscounts, applyDiscountInPlace, loadCachedDiscounts, cacheDiscounts, buildModels, applyModelOverride, parseModelOverrides, loadConfig, getConfig };
|
|
1022
|
+
export type { JsonDiscount, JsonModel, PatchEntry, PatchData, ModelOverride, LilacConfig };
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "pi-lilac-provider",
|
|
3
|
-
"version": "1.
|
|
3
|
+
"version": "1.6.0",
|
|
4
4
|
"description": "Lilac provider extension for pi - Access Kimi K2.6, GLM 5.1, and Gemma 4 models through Lilac's OpenAI-compatible API on idle GPUs",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"main": "index.ts",
|
|
@@ -34,7 +34,10 @@
|
|
|
34
34
|
"clean": "echo 'nothing to clean'",
|
|
35
35
|
"build": "echo 'nothing to build'",
|
|
36
36
|
"check": "echo 'nothing to check'",
|
|
37
|
-
"test": "node scripts/test-discounts.ts",
|
|
37
|
+
"test": "node scripts/test-discounts.ts && node scripts/test-preserved-thinking.ts && node scripts/test-model-overrides.ts",
|
|
38
|
+
"test:discounts": "node scripts/test-discounts.ts",
|
|
39
|
+
"test:thinking": "node scripts/test-preserved-thinking.ts",
|
|
40
|
+
"test:overrides": "node scripts/test-model-overrides.ts",
|
|
38
41
|
"update-models": "node scripts/update-models.js"
|
|
39
42
|
}
|
|
40
43
|
}
|
package/patch.json
CHANGED
|
@@ -10,7 +10,8 @@
|
|
|
10
10
|
"thinkingFormat": "chat-template",
|
|
11
11
|
"chatTemplateKwargs": {
|
|
12
12
|
"thinking": { "$var": "thinking.enabled" },
|
|
13
|
-
"enable_thinking": { "$var": "thinking.enabled" }
|
|
13
|
+
"enable_thinking": { "$var": "thinking.enabled" },
|
|
14
|
+
"preserve_thinking": true
|
|
14
15
|
},
|
|
15
16
|
"maxTokensField": "max_completion_tokens",
|
|
16
17
|
"supportsDeveloperRole": false,
|
|
@@ -29,7 +30,8 @@
|
|
|
29
30
|
"thinkingFormat": "chat-template",
|
|
30
31
|
"chatTemplateKwargs": {
|
|
31
32
|
"thinking": { "$var": "thinking.enabled" },
|
|
32
|
-
"enable_thinking": { "$var": "thinking.enabled" }
|
|
33
|
+
"enable_thinking": { "$var": "thinking.enabled" },
|
|
34
|
+
"clear_thinking": false
|
|
33
35
|
},
|
|
34
36
|
"maxTokensField": "max_completion_tokens",
|
|
35
37
|
"supportsDeveloperRole": false,
|
|
@@ -42,7 +44,8 @@
|
|
|
42
44
|
"thinkingFormat": "chat-template",
|
|
43
45
|
"chatTemplateKwargs": {
|
|
44
46
|
"enable_thinking": { "$var": "thinking.enabled" },
|
|
45
|
-
"reasoning_effort": { "$var": "thinking.effort", "omitWhenOff": true }
|
|
47
|
+
"reasoning_effort": { "$var": "thinking.effort", "omitWhenOff": true },
|
|
48
|
+
"clear_thinking": false
|
|
46
49
|
}
|
|
47
50
|
},
|
|
48
51
|
"thinkingLevelMap": {
|
|
@@ -17,7 +17,7 @@
|
|
|
17
17
|
* recomputed from list price so re-applied discounts never compound.
|
|
18
18
|
* 11. before_provider_request mutates the bound (in-flight) model object so the
|
|
19
19
|
* current turn's cost calc sees the discount in real time.
|
|
20
|
-
* 12. session_start schedules a
|
|
20
|
+
* 12. session_start schedules a 5-minute background /status poll to cover idle
|
|
21
21
|
* sessions; session_shutdown clears it.
|
|
22
22
|
*/
|
|
23
23
|
|
|
@@ -700,12 +700,12 @@ assert(
|
|
|
700
700
|
"deferred session_start paints the NEW lilac model's discount (glm), not the stale kimi capture",
|
|
701
701
|
);
|
|
702
702
|
|
|
703
|
-
// ─── Test 17: session_start schedules a
|
|
703
|
+
// ─── Test 17: session_start schedules a 5-min idle poll; shutdown clears it ─
|
|
704
704
|
|
|
705
705
|
console.log("\n--- Test 17: session_start schedules idle poll; session_shutdown clears it ---");
|
|
706
706
|
|
|
707
|
-
// Lilac refreshes discounts ~every 10 minutes
|
|
708
|
-
//
|
|
707
|
+
// Lilac refreshes discounts ~every 10 minutes, so session_start polls /status
|
|
708
|
+
// every 5 min (half the refresh window) to cover idle sessions (turn fetches
|
|
709
709
|
// only run when the user sends a message), and session_shutdown must clear it so
|
|
710
710
|
// it neither leaks nor keeps the process alive. Wrap the global timer APIs to
|
|
711
711
|
// capture the scheduled delay + handle, then confirm shutdown clears it.
|
|
@@ -729,7 +729,7 @@ globalThis.clearInterval = ((handle: ReturnType<typeof setInterval>) => {
|
|
|
729
729
|
|
|
730
730
|
try {
|
|
731
731
|
// Benign fetch mock so the session_start fire-and-forget /models + /status
|
|
732
|
-
// fetch doesn't hit the network. (The poll itself never fires —
|
|
732
|
+
// fetch doesn't hit the network. (The poll itself never fires — 5 min — so
|
|
733
733
|
// only the startup fetch needs mocking here.)
|
|
734
734
|
globalThis.fetch = mockFetch({
|
|
735
735
|
"/models": { body: { data: [] } },
|
|
@@ -759,8 +759,8 @@ try {
|
|
|
759
759
|
}
|
|
760
760
|
|
|
761
761
|
assert(
|
|
762
|
-
scheduledDelay ===
|
|
763
|
-
"session_start schedules a
|
|
762
|
+
scheduledDelay === 5 * 60 * 1000,
|
|
763
|
+
"session_start schedules a 5-minute (300000ms) /status poll for idle sessions",
|
|
764
764
|
);
|
|
765
765
|
assert(scheduledHandle !== null, "poll interval handle was captured");
|
|
766
766
|
const capturedHandle = scheduledHandle;
|
|
@@ -0,0 +1,262 @@
|
|
|
1
|
+
#!/usr/bin/env node
|
|
2
|
+
/**
|
|
3
|
+
* Test for user model overrides (~/.pi/agent/extensions/lilac.json).
|
|
4
|
+
*
|
|
5
|
+
* Verifies, against the REAL exported helpers from index.ts (not a re-implementation):
|
|
6
|
+
* - applyModelOverride: override wins over the base value; deep-merge compat /
|
|
7
|
+
* cost / thinkingLevelMap preserves non-overridden fields; scalars replaced;
|
|
8
|
+
* input model not mutated.
|
|
9
|
+
* - parseModelOverrides: drops invalid entries (non-object override, non-string
|
|
10
|
+
* thinkingLevelMap values), keeps valid ones.
|
|
11
|
+
* - loadConfig: auto-populates a scaffold on a missing file; parses an existing
|
|
12
|
+
* file's modelOverrides; returns defaults on invalid JSON WITHOUT overwriting
|
|
13
|
+
* the user's file; returns undefined overrides for a file lacking the key.
|
|
14
|
+
* - buildModels end-to-end: a user override on a real model id wins over
|
|
15
|
+
* patch.json (e.g. disable preserve_thinking on kimi-k2.6) while the rest of
|
|
16
|
+
* compat survives; with no overrides the built models are unchanged; an
|
|
17
|
+
* override for an unknown id adds no models.
|
|
18
|
+
*
|
|
19
|
+
* Config FS is isolated to a temp HOME so nothing touches the real ~/.pi.
|
|
20
|
+
*/
|
|
21
|
+
|
|
22
|
+
import fs from "fs";
|
|
23
|
+
import os from "os";
|
|
24
|
+
import path from "path";
|
|
25
|
+
|
|
26
|
+
const root = path.resolve(import.meta.dirname, "..");
|
|
27
|
+
const modelsData = JSON.parse(fs.readFileSync(path.join(root, "models.json"), "utf8"));
|
|
28
|
+
const customModelsData = JSON.parse(fs.readFileSync(path.join(root, "custom-models.json"), "utf8"));
|
|
29
|
+
const patchData = JSON.parse(fs.readFileSync(path.join(root, "patch.json"), "utf8"));
|
|
30
|
+
|
|
31
|
+
// Isolate config + cache to a temp HOME so loadConfig never touches the real ~/.pi.
|
|
32
|
+
// Must be set before importing index.ts, which computes CONFIG_PATH at module scope.
|
|
33
|
+
const tmpHome = `/tmp/pi-lilac-override-test-${Date.now()}`;
|
|
34
|
+
fs.mkdirSync(tmpHome, { recursive: true });
|
|
35
|
+
process.env.HOME = tmpHome;
|
|
36
|
+
|
|
37
|
+
const {
|
|
38
|
+
buildModels,
|
|
39
|
+
applyModelOverride,
|
|
40
|
+
parseModelOverrides,
|
|
41
|
+
loadConfig,
|
|
42
|
+
} = await import("../index.ts");
|
|
43
|
+
|
|
44
|
+
let passed = 0;
|
|
45
|
+
let failed = 0;
|
|
46
|
+
function assert(condition: boolean, message: string) {
|
|
47
|
+
if (condition) {
|
|
48
|
+
console.log(` ✓ ${message}`);
|
|
49
|
+
passed++;
|
|
50
|
+
} else {
|
|
51
|
+
console.error(` ✗ ${message}`);
|
|
52
|
+
failed++;
|
|
53
|
+
}
|
|
54
|
+
}
|
|
55
|
+
function eq<T>(actual: T, expected: T, message: string) {
|
|
56
|
+
const ok = JSON.stringify(actual) === JSON.stringify(expected);
|
|
57
|
+
if (ok) {
|
|
58
|
+
console.log(` ✓ ${message}`);
|
|
59
|
+
passed++;
|
|
60
|
+
} else {
|
|
61
|
+
console.error(` ✗ ${message}\n expected: ${JSON.stringify(expected)}\n actual: ${JSON.stringify(actual)}`);
|
|
62
|
+
failed++;
|
|
63
|
+
}
|
|
64
|
+
}
|
|
65
|
+
|
|
66
|
+
const KIMI = "moonshotai/kimi-k2.6";
|
|
67
|
+
const GLM52 = "zai-org/glm-5.2";
|
|
68
|
+
const GLM51 = "zai-org/glm-5.1";
|
|
69
|
+
|
|
70
|
+
// ─── applyModelOverride ────────────────────────────────────────────────────────
|
|
71
|
+
|
|
72
|
+
console.log("\n--- applyModelOverride ---");
|
|
73
|
+
|
|
74
|
+
{
|
|
75
|
+
const base = {
|
|
76
|
+
id: KIMI,
|
|
77
|
+
reasoning: true,
|
|
78
|
+
compat: {
|
|
79
|
+
thinkingFormat: "chat-template",
|
|
80
|
+
supportsDeveloperRole: false,
|
|
81
|
+
chatTemplateKwargs: { thinking: { $var: "thinking.enabled" }, preserve_thinking: true },
|
|
82
|
+
},
|
|
83
|
+
thinkingLevelMap: { low: "low", high: "high" },
|
|
84
|
+
} as any;
|
|
85
|
+
|
|
86
|
+
const out = applyModelOverride(base, { compat: { chatTemplateKwargs: { preserve_thinking: false } } } as any);
|
|
87
|
+
assert(out.compat.chatTemplateKwargs.preserve_thinking === false, "override wins over base for a compat flag it sets");
|
|
88
|
+
assert(out.compat.thinkingFormat === "chat-template", "deep-merge compat preserves non-overridden thinkingFormat");
|
|
89
|
+
assert(out.compat.supportsDeveloperRole === false, "deep-merge compat preserves non-overridden supportsDeveloperRole");
|
|
90
|
+
assert((out.compat.chatTemplateKwargs as any).thinking?.$var === "thinking.enabled", "deep-merge chatTemplateKwargs preserves non-overridden thinking key");
|
|
91
|
+
assert((base.compat.chatTemplateKwargs as any).preserve_thinking === true, "does not mutate the input model");
|
|
92
|
+
}
|
|
93
|
+
|
|
94
|
+
{
|
|
95
|
+
const base = { id: GLM52, thinkingLevelMap: { low: "low", high: "high" } } as any;
|
|
96
|
+
const out = applyModelOverride(base, { thinkingLevelMap: { high: "max" } } as any);
|
|
97
|
+
eq(out.thinkingLevelMap, { low: "low", high: "max" }, "deep-merge thinkingLevelMap overrides a single level, keeps the rest");
|
|
98
|
+
}
|
|
99
|
+
|
|
100
|
+
{
|
|
101
|
+
const base = { id: KIMI, reasoning: true, contextWindow: 131072 } as any;
|
|
102
|
+
const out = applyModelOverride(base, { reasoning: false, contextWindow: 65536 } as any);
|
|
103
|
+
assert(out.reasoning === false, "scalar override replaces reasoning");
|
|
104
|
+
assert(out.contextWindow === 65536, "scalar override replaces contextWindow");
|
|
105
|
+
}
|
|
106
|
+
|
|
107
|
+
{
|
|
108
|
+
const base = { id: GLM52, cost: { input: 1, output: 2, cacheRead: 0.2, cacheWrite: 0 } } as any;
|
|
109
|
+
// cost is a plain object -> recursively deep-merged (not on the ModelOverride
|
|
110
|
+
// type, so exercise it via a loose cast).
|
|
111
|
+
const out = applyModelOverride(base, { cost: { input: 5 } } as any);
|
|
112
|
+
eq(out.cost, { input: 5, output: 2, cacheRead: 0.2, cacheWrite: 0 }, "deep-merge cost overrides one field, keeps the rest");
|
|
113
|
+
}
|
|
114
|
+
|
|
115
|
+
{
|
|
116
|
+
// Overriding a { $var } schema object wholesale REPLACES it (no deep-merge into $var)
|
|
117
|
+
const base = { id: KIMI, compat: { chatTemplateKwargs: { thinking: { $var: "thinking.enabled" }, preserve_thinking: true } } } as any;
|
|
118
|
+
const out = applyModelOverride(base, { compat: { chatTemplateKwargs: { thinking: { $var: "thinking.effort" } } } } as any);
|
|
119
|
+
eq(out.compat.chatTemplateKwargs.thinking, { $var: "thinking.effort" }, "overriding a { $var } object replaces it (no merge into $var)");
|
|
120
|
+
assert(out.compat.chatTemplateKwargs.preserve_thinking === true, "sibling chatTemplateKwargs key survives a $var replacement");
|
|
121
|
+
}
|
|
122
|
+
|
|
123
|
+
{
|
|
124
|
+
// Overriding an array (input) replaces it wholesale (no index-wise merge)
|
|
125
|
+
const base = { id: KIMI, input: ["text", "image"] } as any;
|
|
126
|
+
const out = applyModelOverride(base, { input: ["text"] } as any);
|
|
127
|
+
eq(out.input, ["text"], "array override replaces wholesale (no index-wise merge)");
|
|
128
|
+
}
|
|
129
|
+
|
|
130
|
+
// ─── parseModelOverrides ───────────────────────────────────────────────────────
|
|
131
|
+
|
|
132
|
+
console.log("\n--- parseModelOverrides ---");
|
|
133
|
+
|
|
134
|
+
eq(parseModelOverrides(undefined), undefined, "undefined input -> undefined");
|
|
135
|
+
eq(parseModelOverrides("nope"), undefined, "non-object input -> undefined");
|
|
136
|
+
eq(parseModelOverrides({}), undefined, "empty object -> undefined (no valid entries)");
|
|
137
|
+
eq(parseModelOverrides({ badId: "not-an-object" } as any), undefined, "non-object override value dropped -> undefined");
|
|
138
|
+
eq(parseModelOverrides({ id: 123 } as any), undefined, "non-object override (number) dropped -> undefined");
|
|
139
|
+
{
|
|
140
|
+
const r = parseModelOverrides({
|
|
141
|
+
[KIMI]: { compat: { chatTemplateKwargs: { preserve_thinking: false } } },
|
|
142
|
+
bad: 123,
|
|
143
|
+
alsobad: "string",
|
|
144
|
+
});
|
|
145
|
+
assert(r !== undefined && Object.keys(r!).length === 1 && r![KIMI] !== undefined, "keeps valid override, drops invalid ids");
|
|
146
|
+
assert((r![KIMI] as any).compat.chatTemplateKwargs.preserve_thinking === false, "valid override compat preserved");
|
|
147
|
+
}
|
|
148
|
+
{
|
|
149
|
+
// thinkingLevelMap: only string/null values kept; non-string values dropped
|
|
150
|
+
const r = parseModelOverrides({ [GLM52]: { thinkingLevelMap: { high: "max", bad: 42, off: null } } });
|
|
151
|
+
eq(r, { [GLM52]: { thinkingLevelMap: { high: "max", off: null } } }, "thinkingLevelMap keeps string/null values, drops others");
|
|
152
|
+
}
|
|
153
|
+
{
|
|
154
|
+
// an override whose fields all parse to nothing is dropped
|
|
155
|
+
const r = parseModelOverrides({ [KIMI]: { thinkingLevelMap: { bad: 42 } } });
|
|
156
|
+
eq(r, undefined, "override with no usable fields is dropped -> undefined");
|
|
157
|
+
}
|
|
158
|
+
|
|
159
|
+
// ─── loadConfig ────────────────────────────────────────────────────────────────
|
|
160
|
+
|
|
161
|
+
console.log("\n--- loadConfig ---");
|
|
162
|
+
|
|
163
|
+
const cfgPath = path.join(os.homedir(), ".pi", "agent", "extensions", "lilac.json");
|
|
164
|
+
|
|
165
|
+
{
|
|
166
|
+
// Fresh tmpHome: no config file -> loadConfig auto-populates the scaffold and returns defaults
|
|
167
|
+
assert(!fs.existsSync(cfgPath), "scaffold not present before first loadConfig");
|
|
168
|
+
const cfg = loadConfig();
|
|
169
|
+
eq(cfg, { modelOverrides: {} }, "missing file -> defaults (empty modelOverrides)");
|
|
170
|
+
assert(fs.existsSync(cfgPath), "loadConfig auto-populates the scaffold file on missing file");
|
|
171
|
+
eq(JSON.parse(fs.readFileSync(cfgPath, "utf8")), { modelOverrides: {} }, "scaffold file contains the default shape");
|
|
172
|
+
}
|
|
173
|
+
|
|
174
|
+
{
|
|
175
|
+
// Existing file with valid modelOverrides -> parsed
|
|
176
|
+
fs.writeFileSync(cfgPath, JSON.stringify({
|
|
177
|
+
modelOverrides: { [KIMI]: { compat: { chatTemplateKwargs: { preserve_thinking: false } } } },
|
|
178
|
+
}));
|
|
179
|
+
const cfg = loadConfig();
|
|
180
|
+
assert((cfg.modelOverrides as any)?.[KIMI]?.compat?.chatTemplateKwargs?.preserve_thinking === false, "existing file's modelOverrides parsed");
|
|
181
|
+
}
|
|
182
|
+
|
|
183
|
+
{
|
|
184
|
+
// Existing file WITHOUT a modelOverrides key -> undefined overrides (not an error)
|
|
185
|
+
fs.writeFileSync(cfgPath, JSON.stringify({ unrelatedKey: true }));
|
|
186
|
+
const cfg = loadConfig();
|
|
187
|
+
assert(cfg.modelOverrides === undefined, "file without modelOverrides key -> undefined overrides");
|
|
188
|
+
assert(JSON.parse(fs.readFileSync(cfgPath, "utf8")).unrelatedKey === true, "file without modelOverrides key is not rewritten");
|
|
189
|
+
}
|
|
190
|
+
|
|
191
|
+
{
|
|
192
|
+
// Existing file with invalid JSON -> defaults returned, file left UNTOUCHED (typo not wiped)
|
|
193
|
+
fs.writeFileSync(cfgPath, "not json {{{");
|
|
194
|
+
const cfg = loadConfig();
|
|
195
|
+
eq(cfg, { modelOverrides: {} }, "invalid JSON -> defaults");
|
|
196
|
+
assert(fs.readFileSync(cfgPath, "utf8") === "not json {{{", "invalid file is not overwritten (typo preserved)");
|
|
197
|
+
}
|
|
198
|
+
|
|
199
|
+
// ─── buildModels end-to-end (real models.json + patch.json) ────────────────────
|
|
200
|
+
|
|
201
|
+
console.log("\n--- buildModels end-to-end ---");
|
|
202
|
+
|
|
203
|
+
function find(models: any[], id: string): any {
|
|
204
|
+
const m = models.find((x) => x.id === id);
|
|
205
|
+
if (!m) throw new Error(`model ${id} not built`);
|
|
206
|
+
return m;
|
|
207
|
+
}
|
|
208
|
+
|
|
209
|
+
{
|
|
210
|
+
// No overrides -> identical to today: kimi preserve_thinking: true survives patch
|
|
211
|
+
const models = buildModels(modelsData, customModelsData, patchData, {});
|
|
212
|
+
const kimi = find(models, KIMI);
|
|
213
|
+
assert(kimi.compat.chatTemplateKwargs.preserve_thinking === true, "no overrides -> kimi preserve_thinking stays true (patch wins)");
|
|
214
|
+
assert(kimi.compat.thinkingFormat === "chat-template", "no overrides -> kimi thinkingFormat intact");
|
|
215
|
+
const glm52 = find(models, GLM52);
|
|
216
|
+
assert(glm52.compat.chatTemplateKwargs.clear_thinking === false, "no overrides -> glm-5.2 clear_thinking stays false (patch wins)");
|
|
217
|
+
}
|
|
218
|
+
|
|
219
|
+
{
|
|
220
|
+
// User override wins over patch.json: disable preserve_thinking on kimi
|
|
221
|
+
const overrides = { [KIMI]: { compat: { chatTemplateKwargs: { preserve_thinking: false } } } } as any;
|
|
222
|
+
const models = buildModels(modelsData, customModelsData, patchData, overrides);
|
|
223
|
+
const kimi = find(models, KIMI);
|
|
224
|
+
assert(kimi.compat.chatTemplateKwargs.preserve_thinking === false, "override wins over patch: kimi preserve_thinking -> false");
|
|
225
|
+
// deep-merge: the $var thinking keys + thinkingFormat survive
|
|
226
|
+
assert((kimi.compat.chatTemplateKwargs as any).thinking?.$var === "thinking.enabled", "override deep-merges: thinking $var key survives");
|
|
227
|
+
assert((kimi.compat.chatTemplateKwargs as any).enable_thinking?.$var === "thinking.enabled", "override deep-merges: enable_thinking $var key survives");
|
|
228
|
+
assert(kimi.compat.thinkingFormat === "chat-template", "override deep-merges: thinkingFormat survives");
|
|
229
|
+
assert(kimi.compat.supportsDeveloperRole === false, "override deep-merges: supportsDeveloperRole survives");
|
|
230
|
+
}
|
|
231
|
+
|
|
232
|
+
{
|
|
233
|
+
// Override a single thinking level on glm-5.2 without redeclaring the map;
|
|
234
|
+
// the patch-applied clear_thinking flag survives a thinkingLevelMap-only override
|
|
235
|
+
const overrides = { [GLM52]: { thinkingLevelMap: { high: "max" } } } as any;
|
|
236
|
+
const models = buildModels(modelsData, customModelsData, patchData, overrides);
|
|
237
|
+
const glm = find(models, GLM52);
|
|
238
|
+
assert((glm.thinkingLevelMap as any)?.high === "max", "override thinkingLevelMap.high wins over patch");
|
|
239
|
+
assert(glm.compat.chatTemplateKwargs.clear_thinking === false, "non-overridden clear_thinking survives a thinkingLevelMap-only override");
|
|
240
|
+
}
|
|
241
|
+
|
|
242
|
+
{
|
|
243
|
+
// Override on glm-5.1 toggles clear_thinking (patch sets false); other compat survives
|
|
244
|
+
const overrides = { [GLM51]: { compat: { chatTemplateKwargs: { clear_thinking: true } } } } as any;
|
|
245
|
+
const models = buildModels(modelsData, customModelsData, patchData, overrides);
|
|
246
|
+
const glm = find(models, GLM51);
|
|
247
|
+
assert(glm.compat.chatTemplateKwargs.clear_thinking === true, "override wins over patch: glm-5.1 clear_thinking -> true");
|
|
248
|
+
assert((glm.compat.chatTemplateKwargs as any).thinking?.$var === "thinking.enabled", "override deep-merges: glm-5.1 thinking $var key survives");
|
|
249
|
+
assert(glm.compat.zaiToolStream === true, "override deep-merges: glm-5.1 zaiToolStream survives");
|
|
250
|
+
}
|
|
251
|
+
|
|
252
|
+
{
|
|
253
|
+
// Override for an unknown id is a no-op (adds no models)
|
|
254
|
+
const before = buildModels(modelsData, customModelsData, patchData, {});
|
|
255
|
+
const after = buildModels(modelsData, customModelsData, patchData, { "no/such-model": { reasoning: false } } as any);
|
|
256
|
+
eq(after.length, before.length, "override for an unknown id adds no models");
|
|
257
|
+
}
|
|
258
|
+
|
|
259
|
+
// ─── Summary ───────────────────────────────────────────────────────────────────
|
|
260
|
+
|
|
261
|
+
console.log(`\n${failed === 0 ? "ALL PASS" : `${failed} FAILED`}`);
|
|
262
|
+
process.exit(failed === 0 ? 0 : 1);
|
|
@@ -0,0 +1,237 @@
|
|
|
1
|
+
#!/usr/bin/env node
|
|
2
|
+
/**
|
|
3
|
+
* Wire-level test for preserved-thinking (full-history reasoning) flags.
|
|
4
|
+
*
|
|
5
|
+
* Verifies, against the REAL pi-ai streamSimple + a stubbed fetch, that the
|
|
6
|
+
* `chat_template_kwargs` each Lilac reasoning model puts on the wire include the
|
|
7
|
+
* preservation flags configured in patch.json — and that those static flags
|
|
8
|
+
* coexist with the { $var }-resolved thinking keys at every thinking level.
|
|
9
|
+
*
|
|
10
|
+
* Background: reasoning models trim older assistant reasoning across turns by
|
|
11
|
+
* default (each vendor's template default). Two template-level flags opt into
|
|
12
|
+
* full-history preservation (confirmed in each model's HuggingFace chat
|
|
13
|
+
* template; Kimi K2.6 and GLM 5.2 additionally E2E-verified on the sibling
|
|
14
|
+
* neuralwatt provider's 3-turn / two-20-digit-number recall test):
|
|
15
|
+
*
|
|
16
|
+
* Kimi K2.6 → preserve_thinking: true (template: "if preserve_thinking, keep -1 ... retain reasoning")
|
|
17
|
+
* GLM 5.1 → clear_thinking: false (same clear_thinking mechanism as GLM 5.2)
|
|
18
|
+
* GLM 5.2 → clear_thinking: false (template: keep reasoning when "clear_thinking is defined and not clear_thinking")
|
|
19
|
+
*
|
|
20
|
+
* Lilac uses pi-ai's `chat-template` thinkingFormat, which calls
|
|
21
|
+
* buildChatTemplateKwargs → resolveChatTemplateKwargValue. Static primitive
|
|
22
|
+
* kwargs pass through verbatim; only { $var } objects are resolved against the
|
|
23
|
+
* turn's thinking state. So the preserve flags ride onto the wire as plain
|
|
24
|
+
* booleans next to the { $var } thinking/enable_thinking keys — no onPayload
|
|
25
|
+
* hook needed (unlike neuralwatt, which drives the openai reasoning_effort path
|
|
26
|
+
* and must inject via onPayload because the two paths are mutually exclusive).
|
|
27
|
+
*
|
|
28
|
+
* Gemma 4 and MiniMax M2.7/M3 expose NO family-wide preserve flag (their HF
|
|
29
|
+
* templates read only enable_thinking / thinking_mode + reasoning_content), so
|
|
30
|
+
* no flag is added for them here — asserted as a regression guard.
|
|
31
|
+
*
|
|
32
|
+
* Run: node scripts/test-preserved-thinking.ts
|
|
33
|
+
*/
|
|
34
|
+
import fs from "fs";
|
|
35
|
+
import path from "path";
|
|
36
|
+
import { pathToFileURL } from "url";
|
|
37
|
+
|
|
38
|
+
const HERE = path.dirname(new URL(import.meta.url).pathname);
|
|
39
|
+
|
|
40
|
+
// ─── Faithful replica of index.ts applyPatch / buildModels ────────────────────
|
|
41
|
+
// Mirrors the provider's model pipeline so the test exercises the same objects
|
|
42
|
+
// the runtime registers. (Kept inline so the test has no TS-import dependency
|
|
43
|
+
// on index.ts, which pulls in pi-coding-agent types.)
|
|
44
|
+
|
|
45
|
+
function applyPatch(model, patch) {
|
|
46
|
+
const result = { ...model };
|
|
47
|
+
if (patch.name !== undefined) result.name = patch.name;
|
|
48
|
+
if (patch.reasoning !== undefined) result.reasoning = patch.reasoning;
|
|
49
|
+
if (patch.input !== undefined) result.input = patch.input;
|
|
50
|
+
if (patch.contextWindow !== undefined) result.contextWindow = patch.contextWindow;
|
|
51
|
+
if (patch.maxTokens !== undefined) result.maxTokens = patch.maxTokens;
|
|
52
|
+
if (patch.cost) {
|
|
53
|
+
result.cost = {
|
|
54
|
+
input: patch.cost.input ?? result.cost.input,
|
|
55
|
+
output: patch.cost.output ?? result.cost.output,
|
|
56
|
+
cacheRead: patch.cost.cacheRead ?? result.cost.cacheRead,
|
|
57
|
+
cacheWrite: patch.cost.cacheWrite ?? result.cost.cacheWrite,
|
|
58
|
+
};
|
|
59
|
+
}
|
|
60
|
+
if (patch.compat) result.compat = { ...(result.compat || {}), ...patch.compat };
|
|
61
|
+
if (patch.thinkingLevelMap !== undefined) result.thinkingLevelMap = patch.thinkingLevelMap;
|
|
62
|
+
if (!result.reasoning && result.compat?.thinkingFormat) delete result.compat.thinkingFormat;
|
|
63
|
+
if (!result.reasoning && result.thinkingLevelMap) delete result.thinkingLevelMap;
|
|
64
|
+
if (result.compat && Object.keys(result.compat).length === 0) delete result.compat;
|
|
65
|
+
return result;
|
|
66
|
+
}
|
|
67
|
+
|
|
68
|
+
function buildModels(base, custom, patch) {
|
|
69
|
+
const map = new Map();
|
|
70
|
+
for (const m of base) map.set(m.id, m);
|
|
71
|
+
for (const [id, p] of Object.entries(patch)) {
|
|
72
|
+
const ex = map.get(id);
|
|
73
|
+
if (ex) map.set(id, applyPatch(ex, p));
|
|
74
|
+
}
|
|
75
|
+
for (const m of custom) {
|
|
76
|
+
const ex = map.get(m.id);
|
|
77
|
+
const p = patch[m.id];
|
|
78
|
+
if (ex && p) map.set(m.id, applyPatch(m, p));
|
|
79
|
+
else if (ex) map.set(m.id, m);
|
|
80
|
+
else if (p) map.set(m.id, applyPatch(m, p));
|
|
81
|
+
else map.set(m.id, m);
|
|
82
|
+
}
|
|
83
|
+
return Array.from(map.values());
|
|
84
|
+
}
|
|
85
|
+
|
|
86
|
+
// ─── Load REAL pi-ai streamSimple from the global pi install ──────────────────
|
|
87
|
+
// Imported by absolute path so its relative deps (openai, ../models.js, ...) and
|
|
88
|
+
// the `openai` SDK resolve from the global pi node_modules tree.
|
|
89
|
+
const PI_AI_API = path.join(
|
|
90
|
+
os_home_pi_ai(),
|
|
91
|
+
"dist/api/openai-completions.js",
|
|
92
|
+
);
|
|
93
|
+
function os_home_pi_ai() {
|
|
94
|
+
// Resolve the pi-ai package shipped inside the globally-installed pi agent.
|
|
95
|
+
const candidates = [
|
|
96
|
+
"/Users/monotykamary/.npm-global/lib/node_modules/@earendil-works/pi-coding-agent/node_modules/@earendil-works/pi-ai",
|
|
97
|
+
];
|
|
98
|
+
for (const c of candidates) if (fs.existsSync(c)) return c;
|
|
99
|
+
throw new Error("Could not locate pi-ai package in the global pi install.");
|
|
100
|
+
}
|
|
101
|
+
|
|
102
|
+
const { streamSimple } = await import(pathToFileURL(PI_AI_API).href);
|
|
103
|
+
|
|
104
|
+
// ─── Build the exact model objects the provider registers ────────────────────
|
|
105
|
+
const embedded = JSON.parse(fs.readFileSync(path.join(HERE, "..", "models.json"), "utf8"));
|
|
106
|
+
const custom = JSON.parse(fs.readFileSync(path.join(HERE, "..", "custom-models.json"), "utf8"));
|
|
107
|
+
const patch = JSON.parse(fs.readFileSync(path.join(HERE, "..", "patch.json"), "utf8"));
|
|
108
|
+
const models = new Map(buildModels(embedded, custom, patch).map((m) => [m.id, m]));
|
|
109
|
+
|
|
110
|
+
// ─── Stub fetch to capture the request body (fires before response parsing) ──
|
|
111
|
+
const captured = [];
|
|
112
|
+
const originalFetch = globalThis.fetch;
|
|
113
|
+
globalThis.fetch = async (url, init) => {
|
|
114
|
+
captured.push({ url: String(url), body: init?.body ?? null });
|
|
115
|
+
// Minimal SSE response that ends immediately — enough for the OpenAI SDK to
|
|
116
|
+
// construct the stream; we only need the request body, captured above.
|
|
117
|
+
return new Response(new ReadableStream({ start(c) { c.close(); } }), {
|
|
118
|
+
headers: { "content-type": "text/event-stream" },
|
|
119
|
+
});
|
|
120
|
+
};
|
|
121
|
+
|
|
122
|
+
async function wire(modelId, reasoning) {
|
|
123
|
+
const model = models.get(modelId);
|
|
124
|
+
if (!model) throw new Error(`unknown model: ${modelId}`);
|
|
125
|
+
captured.length = 0;
|
|
126
|
+
const ctx = { messages: [{ role: "user", content: [{ type: "text", text: "hi" }] }] };
|
|
127
|
+
const s = streamSimple(
|
|
128
|
+
{ ...model, provider: "lilac", api: "openai-completions", baseUrl: "https://api.getlilac.com/v1" },
|
|
129
|
+
ctx,
|
|
130
|
+
{ apiKey: "sk-test", reasoning },
|
|
131
|
+
);
|
|
132
|
+
// The fetch fires inside stream()'s async IIFE; poll for it.
|
|
133
|
+
const deadline = Date.now() + 2000;
|
|
134
|
+
while (captured.length === 0 && Date.now() < deadline) await new Promise((r) => setTimeout(r, 20));
|
|
135
|
+
try { s.end?.(); } catch {}
|
|
136
|
+
const hit = captured.find((c) => c.url.includes("/chat/completions"));
|
|
137
|
+
if (!hit) throw new Error(`no /chat/completions request captured for ${modelId} (${reasoning})`);
|
|
138
|
+
return JSON.parse(hit.body);
|
|
139
|
+
}
|
|
140
|
+
|
|
141
|
+
// ─── Assertions ───────────────────────────────────────────────────────────────
|
|
142
|
+
let failures = 0;
|
|
143
|
+
function eq(actual, expected, msg) {
|
|
144
|
+
const a = JSON.stringify(actual);
|
|
145
|
+
const e = JSON.stringify(expected);
|
|
146
|
+
const ok = a === e;
|
|
147
|
+
console.log(`${ok ? "✓" : "✗"} ${msg}`);
|
|
148
|
+
if (!ok) {
|
|
149
|
+
failures++;
|
|
150
|
+
console.log(` expected ${e}`);
|
|
151
|
+
console.log(` actual ${a}`);
|
|
152
|
+
}
|
|
153
|
+
}
|
|
154
|
+
function truthy(actual, msg) {
|
|
155
|
+
const ok = !!actual;
|
|
156
|
+
console.log(`${ok ? "✓" : "✗"} ${msg}`);
|
|
157
|
+
if (!ok) {
|
|
158
|
+
failures++;
|
|
159
|
+
console.log(` expected truthy, got ${JSON.stringify(actual)}`);
|
|
160
|
+
}
|
|
161
|
+
}
|
|
162
|
+
function falsy(actual, msg) {
|
|
163
|
+
const ok = !actual;
|
|
164
|
+
console.log(`${ok ? "✓" : "✗"} ${msg}`);
|
|
165
|
+
if (!ok) {
|
|
166
|
+
failures++;
|
|
167
|
+
console.log(` expected falsy/undefined, got ${JSON.stringify(actual)}`);
|
|
168
|
+
}
|
|
169
|
+
}
|
|
170
|
+
|
|
171
|
+
console.log("\n=== patch.json data ===");
|
|
172
|
+
const kimiPatch = patch["moonshotai/kimi-k2.6"]?.compat?.chatTemplateKwargs;
|
|
173
|
+
eq(kimiPatch?.preserve_thinking, true, "kimi-k2.6 patch sets preserve_thinking: true");
|
|
174
|
+
const glm51Patch = patch["zai-org/glm-5.1"]?.compat?.chatTemplateKwargs;
|
|
175
|
+
eq(glm51Patch?.clear_thinking, false, "glm-5.1 patch sets clear_thinking: false");
|
|
176
|
+
const glm52Patch = patch["zai-org/glm-5.2"]?.compat?.chatTemplateKwargs;
|
|
177
|
+
eq(glm52Patch?.clear_thinking, false, "glm-5.2 patch sets clear_thinking: false");
|
|
178
|
+
|
|
179
|
+
console.log("\n=== Kimi K2.6 on the wire (real pi-ai) ===");
|
|
180
|
+
{
|
|
181
|
+
const high = await wire("moonshotai/kimi-k2.6", "high");
|
|
182
|
+
eq(high.chat_template_kwargs, { thinking: true, enable_thinking: true, preserve_thinking: true },
|
|
183
|
+
"kimi @ high → thinking+enable_thinking true AND preserve_thinking true");
|
|
184
|
+
const off = await wire("moonshotai/kimi-k2.6", "off");
|
|
185
|
+
eq(off.chat_template_kwargs, { thinking: false, enable_thinking: false, preserve_thinking: true },
|
|
186
|
+
"kimi @ off → thinking false but preserve_thinking still true (level-independent)");
|
|
187
|
+
const minimal = await wire("moonshotai/kimi-k2.6", "minimal");
|
|
188
|
+
eq(minimal.chat_template_kwargs, { thinking: true, enable_thinking: true, preserve_thinking: true },
|
|
189
|
+
"kimi @ minimal → preserve_thinking present at every level");
|
|
190
|
+
}
|
|
191
|
+
|
|
192
|
+
console.log("\n=== GLM 5.2 on the wire (real pi-ai) ===");
|
|
193
|
+
{
|
|
194
|
+
const high = await wire("zai-org/glm-5.2", "high");
|
|
195
|
+
eq(high.chat_template_kwargs, { enable_thinking: true, reasoning_effort: "high", clear_thinking: false },
|
|
196
|
+
"glm-5.2 @ high → enable_thinking+reasoning_effort AND clear_thinking false");
|
|
197
|
+
const xhigh = await wire("zai-org/glm-5.2", "xhigh");
|
|
198
|
+
eq(xhigh.chat_template_kwargs, { enable_thinking: true, reasoning_effort: "max", clear_thinking: false },
|
|
199
|
+
"glm-5.2 @ xhigh → reasoning_effort max, clear_thinking false");
|
|
200
|
+
const off = await wire("zai-org/glm-5.2", "off");
|
|
201
|
+
eq(off.chat_template_kwargs, { enable_thinking: false, clear_thinking: false },
|
|
202
|
+
"glm-5.2 @ off → reasoning_effort omitted (omitWhenOff), clear_thinking false persists");
|
|
203
|
+
}
|
|
204
|
+
|
|
205
|
+
console.log("\n=== GLM 5.1 on the wire (real pi-ai) ===");
|
|
206
|
+
{
|
|
207
|
+
const high = await wire("zai-org/glm-5.1", "high");
|
|
208
|
+
eq(high.chat_template_kwargs, { thinking: true, enable_thinking: true, clear_thinking: false },
|
|
209
|
+
"glm-5.1 @ high → thinking+enable_thinking true AND clear_thinking false");
|
|
210
|
+
const off = await wire("zai-org/glm-5.1", "off");
|
|
211
|
+
eq(off.chat_template_kwargs, { thinking: false, enable_thinking: false, clear_thinking: false },
|
|
212
|
+
"glm-5.1 @ off → thinking false but clear_thinking false persists");
|
|
213
|
+
}
|
|
214
|
+
|
|
215
|
+
console.log("\n=== Gemma 4 / MiniMax (no family-wide preserve flag — regression guard) ===");
|
|
216
|
+
{
|
|
217
|
+
const gemma = await wire("google/gemma-4-31b-it", "high");
|
|
218
|
+
eq(gemma.chat_template_kwargs, { thinking: true, enable_thinking: true },
|
|
219
|
+
"gemma-4 @ high → only thinking/enable_thinking (no preserve/clear flag)");
|
|
220
|
+
falsy(gemma.chat_template_kwargs?.preserve_thinking, "gemma-4 has no preserve_thinking");
|
|
221
|
+
falsy(gemma.chat_template_kwargs?.clear_thinking, "gemma-4 has no clear_thinking");
|
|
222
|
+
|
|
223
|
+
const m3 = await wire("minimaxai/minimax-m3", "high");
|
|
224
|
+
eq(m3.chat_template_kwargs, { thinking_mode: "enabled" },
|
|
225
|
+
"minimax-m3 @ high → only thinking_mode (no preserve/clear flag)");
|
|
226
|
+
falsy(m3.chat_template_kwargs?.preserve_thinking, "minimax-m3 has no preserve_thinking");
|
|
227
|
+
|
|
228
|
+
const m27 = await wire("minimaxai/minimax-m2.7", "high");
|
|
229
|
+
eq(m27.chat_template_kwargs, { thinking: true, enable_thinking: true },
|
|
230
|
+
"minimax-m2.7 @ high → only thinking/enable_thinking (no preserve/clear flag)");
|
|
231
|
+
falsy(m27.chat_template_kwargs?.clear_thinking, "minimax-m2.7 has no clear_thinking");
|
|
232
|
+
}
|
|
233
|
+
|
|
234
|
+
globalThis.fetch = originalFetch;
|
|
235
|
+
|
|
236
|
+
console.log(`\n${failures === 0 ? "ALL PASS" : `${failures} FAILURE(S)`}`);
|
|
237
|
+
if (failures > 0) process.exit(1);
|