pi-lilac-provider 1.7.2 → 1.8.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +17 -11
- package/index.ts +175 -65
- package/models.json +0 -45
- package/package.json +1 -1
- package/patch.json +1 -1
- package/scripts/test-flex.ts +33 -78
- package/scripts/test-model-overrides.ts +0 -11
- package/scripts/test-preserved-thinking.ts +4 -19
package/README.md
CHANGED
|
@@ -23,7 +23,7 @@ Access Kimi K2.6, GLM 5.1, MiniMax M2.7, and Gemma 4 models through Lilac's Open
|
|
|
23
23
|
- **Reasoning Models** — Chain-of-thought via `chat_template_kwargs` (all models)
|
|
24
24
|
- **Vision Support** — Image input on Kimi K2.6 and Gemma 4
|
|
25
25
|
- **Context Caching** — Cache read pricing on Kimi K2.6 and GLM 5.1
|
|
26
|
-
- **Flex (Discount Gating)** — Only let the LLM respond when the active model's discount meets a threshold you set (`/lilac-
|
|
26
|
+
- **Flex (Discount Gating)** — Only let the LLM respond when the active model's discount meets a threshold you set (`/lilac-settings`)
|
|
27
27
|
- **Idle GPU Scheduling** — Lilac leverages idle GPU capacity for cost-efficient inference
|
|
28
28
|
|
|
29
29
|
## Installation
|
|
@@ -74,10 +74,8 @@ pi
|
|
|
74
74
|
| Model | Context | Vision | Reasoning | Input $/M | Cache Read $/M | Output $/M |
|
|
75
75
|
|-------|---------|--------|-----------|-----------|-----------------|------------|
|
|
76
76
|
| Gemma 4 | 262K | ✅ | ✅ | $0.11 | — | $0.35 |
|
|
77
|
-
| GLM 5.1 | 203K | ❌ | ✅ | $0.90 | $0.27 | $3.00 |
|
|
78
77
|
| GLM 5.2 | 524K | ❌ | ✅ | $0.90 | $0.27 | $3.00 |
|
|
79
78
|
| Kimi K2.6 | 262K | ✅ | ✅ | $0.70 | $0.20 | $3.50 |
|
|
80
|
-
| MiniMax M2.7 | 205K | ❌ | ✅ | $0.30 | $0.06 | $1.20 |
|
|
81
79
|
| MiniMax M3 | 1.0M | ✅ | ✅ | $0.28 | $0.05 | $1.10 |
|
|
82
80
|
|
|
83
81
|
*Costs are per million tokens. Prices subject to change — check [getlilac.com](https://getlilac.com/) for current pricing.*
|
|
@@ -121,7 +119,7 @@ model's chat template honors differs per family. The provider uses pi's
|
|
|
121
119
|
Kimi K2.6, GLM 5.1, Gemma 4, and MiniMax M2.7 use the forward-compatible form
|
|
122
120
|
that sends **both** `thinking` and `enable_thinking`, so whichever key the
|
|
123
121
|
template honors is set. GLM 5.2 additionally maps pi's thinking levels to
|
|
124
|
-
`reasoning_effort` (`high` = lower-latency, `
|
|
122
|
+
`reasoning_effort` (`high` = lower-latency, `max` = deepest). MiniMax M3 uses
|
|
125
123
|
the `thinking_mode` enum, exposed as three pi thinking levels: `off` →
|
|
126
124
|
`disabled` (never think), `minimal` → `adaptive` (the model decides), `high` →
|
|
127
125
|
`enabled` (always think). Pi starts at `off` (`disabled`); cycle to `minimal`
|
|
@@ -220,7 +218,7 @@ Create `~/.pi/agent/extensions/lilac.json` (auto-populated with defaults on firs
|
|
|
220
218
|
```jsonc
|
|
221
219
|
{
|
|
222
220
|
// Only respond when the active model's discount is >= this percent. null = off.
|
|
223
|
-
// See "Flex (Discount Gating)" below. Set
|
|
221
|
+
// See "Flex (Discount Gating)" below. Set via /lilac-settings.
|
|
224
222
|
"flexThreshold": null,
|
|
225
223
|
"modelOverrides": {
|
|
226
224
|
// Disable full-history reasoning for kimi-k2.6 (e.g. to save tokens):
|
|
@@ -239,16 +237,13 @@ The full set of overridable fields matches the model schema (`compat`, `thinking
|
|
|
239
237
|
|
|
240
238
|
Lilac's per-model discount fluctuates with idle-GPU supply. **Flex** lets you set a discount threshold so pi **only sends a prompt to the LLM when the active model's current discount is at or above it** — e.g. "only respond when the discount is ≥ 75%". Below the threshold, the prompt is blocked (dropped with a warning) until the next discount poll brings the discount back up. This is a spend-control feature: you only spend when supply is cheap.
|
|
241
239
|
|
|
242
|
-
Set it
|
|
240
|
+
Set it via the **Flex threshold** row in [`/lilac-settings`](#settings-ui) — cycle `off` / `50` / `75`:
|
|
243
241
|
|
|
244
242
|
```
|
|
245
|
-
/lilac-
|
|
246
|
-
/lilac-flex 75 # set threshold directly (only respond at ≥75% discount)
|
|
247
|
-
/lilac-flex 50% # trailing % accepted on the command line
|
|
248
|
-
/lilac-flex off # disable flex (allow all discounts)
|
|
243
|
+
/lilac-settings # open the settings panel → Flex threshold row
|
|
249
244
|
```
|
|
250
245
|
|
|
251
|
-
The threshold persists in `~/.pi/agent/extensions/lilac.json` as `flexThreshold` (a number `0`–`100`, or `null` for off) alongside `modelOverrides`, so it survives restarts. `/lilac-
|
|
246
|
+
The threshold persists in `~/.pi/agent/extensions/lilac.json` as `flexThreshold` (a number `0`–`100`, or `null` for off) alongside `modelOverrides`, so it survives restarts. `/lilac-settings` updates it live — no restart needed. For a custom value (e.g. 60), edit `flexThreshold` in that file directly.
|
|
252
247
|
|
|
253
248
|
Behavior notes:
|
|
254
249
|
|
|
@@ -258,6 +253,17 @@ Behavior notes:
|
|
|
258
253
|
- **Freshness:** when a prompt is blocked, the extension triggers an immediate `/status` refresh (throttled to once per ~5s) so you're not stuck on a stale low value from the 5-minute idle poll. The next submission sees the fresh discount. The footer status reflects the gate: `… · flex ≥75% ok` or `… · flex ≥75% blocked`.
|
|
259
254
|
- **Discount lock-in:** per Lilac, a discount is locked in when a request starts. Flex gates on the best-known discount at submit time, which is what gets locked in for that turn.
|
|
260
255
|
|
|
256
|
+
### Settings UI
|
|
257
|
+
|
|
258
|
+
`/lilac-settings` opens an interactive settings panel (mirrors pi core's `/settings` — bordered `SettingsList`, Esc to go back) to configure Lilac without editing JSON by hand:
|
|
259
|
+
|
|
260
|
+
- **Flex threshold** (`off` / `50` / `75`) — the spend-control gate from [Flex (Discount Gating)](#flex-discount-gating); custom values via `lilac.json`.
|
|
261
|
+
- **Preserved thinking** (nested submenu, one row per model) — toggles `clear_thinking` (GLM-5.2 / GLM-5.1) / `preserve_thinking` (Kimi K2.6) in `modelOverrides` between **Preserve Thinking** (keep full reasoning history across turns; the default, `clear_thinking: false` / `preserve_thinking: true`) and **Clear Thinking** (let the template drop older reasoning; saves tokens, but can degrade multi-turn recall / cause overthinking).
|
|
262
|
+
|
|
263
|
+
Changes write to `~/.pi/agent/extensions/lilac.json`, refresh the in-memory config, and re-register the provider, so they take effect immediately — no restart needed.
|
|
264
|
+
|
|
265
|
+
When you switch to — or start pi on — a Lilac model that carries a preserved-thinking flag (e.g. GLM-5.2 / GLM-5.1, Kimi K2.6), an info notification reports the state and how to change it, e.g. `Preserved thinking ON for glm-5.2 (clear_thinking: false) — suited for coding, but not for prose. Open /lilac-settings to change.` (OFF reads `... reasoning trimmed each turn (lighter; better for prose) ...`). It's an ordinary info notification (not a warning), so it doesn't paint bright yellow.
|
|
266
|
+
|
|
261
267
|
## Updating Models
|
|
262
268
|
|
|
263
269
|
Run the update script to fetch the latest models from Lilac's API:
|
package/index.ts
CHANGED
|
@@ -21,7 +21,7 @@
|
|
|
21
21
|
* MiniMax M3). The forward-compatible form sends BOTH `thinking` and
|
|
22
22
|
* `enable_thinking` so whichever key the template honors is set:
|
|
23
23
|
* { chat_template_kwargs: { thinking: <bool>, enable_thinking: <bool> } }
|
|
24
|
-
* GLM 5.2 adds `reasoning_effort` (high
|
|
24
|
+
* GLM 5.2 adds `reasoning_effort` (high for lower-latency, max for deep) via a
|
|
25
25
|
* thinkingLevelMap. MiniMax M3 maps to the `thinking_mode` enum as three pi
|
|
26
26
|
* thinking levels — off→disabled, minimal→adaptive (model decides), high→enabled
|
|
27
27
|
* — so adaptive is selectable via pi's Shift+Tab cycle (off→minimal→high). Pi
|
|
@@ -99,7 +99,7 @@ interface JsonDiscount {
|
|
|
99
99
|
creditMultiplier: number;
|
|
100
100
|
}
|
|
101
101
|
|
|
102
|
-
// Maps pi's thinking levels (off, minimal, low, medium, high, xhigh) to the
|
|
102
|
+
// Maps pi's thinking levels (off, minimal, low, medium, high, xhigh, max) to the
|
|
103
103
|
// provider-specific effort string sent on the wire. A `null` value marks a
|
|
104
104
|
// level as unsupported — clampThinkingLevel skips it when resolving the
|
|
105
105
|
// user's selection. Mirrors pi-ai's ThinkingLevelMap shape.
|
|
@@ -110,6 +110,7 @@ type ThinkingLevelMap = {
|
|
|
110
110
|
medium?: string | null;
|
|
111
111
|
high?: string | null;
|
|
112
112
|
xhigh?: string | null;
|
|
113
|
+
max?: string | null;
|
|
113
114
|
};
|
|
114
115
|
|
|
115
116
|
// A chat_template_kwargs value, mirroring pi-ai's ChatTemplateKwargSchema. Scalar
|
|
@@ -193,7 +194,7 @@ interface LilacConfig {
|
|
|
193
194
|
modelOverrides?: Record<string, ModelOverride>;
|
|
194
195
|
// Flex discount threshold: only allow interactive prompts to reach the LLM
|
|
195
196
|
// when the active lilac model's discountPercent is >= this. null/undefined =
|
|
196
|
-
// off (allow all). Set via /lilac-
|
|
197
|
+
// off (allow all). Set via /lilac-settings; persisted in lilac.json.
|
|
197
198
|
flexThreshold?: number | null;
|
|
198
199
|
}
|
|
199
200
|
|
|
@@ -312,12 +313,12 @@ function getConfig(): LilacConfig {
|
|
|
312
313
|
}
|
|
313
314
|
|
|
314
315
|
// Read-modify-write the config file and refresh the in-memory cache. Used by
|
|
315
|
-
// /lilac-
|
|
316
|
+
// /lilac-settings so a threshold change takes effect immediately (the input gate and
|
|
316
317
|
// footer read getConfig()) without a restart, and without clobbering the user's
|
|
317
318
|
// modelOverrides. Reads via loadConfig (which validates), so modelOverrides the
|
|
318
319
|
// user hand-edited since startup survive the spread. The file is normalized to a
|
|
319
320
|
// discoverable shape (modelOverrides: {}, flexThreshold: null always present) so
|
|
320
|
-
// /lilac-
|
|
321
|
+
// /lilac-settings never strips the modelOverrides scaffold from the file.
|
|
321
322
|
function updateConfig(mutator: (cfg: LilacConfig) => LilacConfig): LilacConfig {
|
|
322
323
|
const next = mutator(loadConfig());
|
|
323
324
|
const toWrite: LilacConfig = {
|
|
@@ -760,20 +761,9 @@ function syncStatus(ctx: any): void {
|
|
|
760
761
|
}
|
|
761
762
|
}
|
|
762
763
|
|
|
763
|
-
// Parse a /lilac-flex argument or custom-threshold input: "off"/"none"/"" → null
|
|
764
|
-
// (disabled), a number in [0,100] (optionally with a trailing %) → that number,
|
|
765
|
-
// anything else → undefined (invalid).
|
|
766
|
-
function parseFlexArg(raw: string): number | null | undefined {
|
|
767
|
-
const s = raw.trim().toLowerCase().replace(/%$/, "");
|
|
768
|
-
if (s === "" || s === "off" || s === "none") return null;
|
|
769
|
-
const n = Number(s);
|
|
770
|
-
if (!Number.isFinite(n) || n < 0 || n > 100) return undefined;
|
|
771
|
-
return Math.round(n);
|
|
772
|
-
}
|
|
773
|
-
|
|
774
764
|
// Persist a flex threshold via updateConfig and report the resulting state against
|
|
775
765
|
// the live model's current discount. The footer is re-painted (syncStatus) so the
|
|
776
|
-
// flex indicator appears immediately. Used by the /lilac-flex
|
|
766
|
+
// flex indicator appears immediately. Used by the /lilac-settings flex row.
|
|
777
767
|
function applyFlexThreshold(value: number | null, ctx: any): void {
|
|
778
768
|
updateConfig((cfg) => ({ ...cfg, flexThreshold: value }));
|
|
779
769
|
const threshold = getConfig().flexThreshold ?? null;
|
|
@@ -784,16 +774,16 @@ function applyFlexThreshold(value: number | null, ctx: any): void {
|
|
|
784
774
|
let msg: string;
|
|
785
775
|
let level: "info" | "warning";
|
|
786
776
|
if (threshold == null) {
|
|
787
|
-
msg = "
|
|
777
|
+
msg = "Flex: off — all discounts allowed";
|
|
788
778
|
level = "info";
|
|
789
779
|
} else if (pct == null) {
|
|
790
|
-
msg = `
|
|
780
|
+
msg = `Flex: ≥ ${threshold}% — no discount data yet; will block until the next poll`;
|
|
791
781
|
level = "warning";
|
|
792
782
|
} else if (pct >= threshold) {
|
|
793
|
-
msg = `
|
|
783
|
+
msg = `Flex: ≥ ${threshold}% — current ${pct}% discount allowed`;
|
|
794
784
|
level = "info";
|
|
795
785
|
} else {
|
|
796
|
-
msg = `
|
|
786
|
+
msg = `Flex: ≥ ${threshold}% — current ${pct}% discount blocked until it improves`;
|
|
797
787
|
level = "warning";
|
|
798
788
|
}
|
|
799
789
|
try {
|
|
@@ -862,6 +852,31 @@ export default function (pi: ExtensionAPI) {
|
|
|
862
852
|
const customModels = customModelsData as JsonModel[];
|
|
863
853
|
const patches = patchData as PatchData;
|
|
864
854
|
|
|
855
|
+
// Deferred model_select notify timer — see the model_select handler. Cleared on
|
|
856
|
+
// rapid re-switch and on session_shutdown so only the latest switch notifies.
|
|
857
|
+
let modelSelectNotifyTimer: ReturnType<typeof setTimeout> | null = null;
|
|
858
|
+
const MODEL_SELECT_NOTIFY_DELAY_MS = 250;
|
|
859
|
+
|
|
860
|
+
// Notify preserved-thinking state for a preserve-flag model. Computed from the
|
|
861
|
+
// build pipeline (config as source of truth, not event.model.compat), deferred
|
|
862
|
+
// so pi core's (and other extensions') notifications land first, and cancelled
|
|
863
|
+
// on re-switch/shutdown so only the latest shows. Always level "info" (not a
|
|
864
|
+
// warning) — the text conveys the coding/prose tradeoff.
|
|
865
|
+
function notifyPreservedThinkingFor(model: any, ctx: any): void {
|
|
866
|
+
if (!model || model.provider !== PROVIDER_ID) return;
|
|
867
|
+
const entry = collectPreserveState().find((e: any) => e.id === model.id);
|
|
868
|
+
if (!entry) return;
|
|
869
|
+
const flagValue = entry.flag === "clear_thinking" ? !entry.preserved : entry.preserved;
|
|
870
|
+
const msg = entry.preserved
|
|
871
|
+
? `Preserved thinking ON for ${entry.name} (${entry.flag}: ${flagValue}) — suited for coding, but not for prose. Open /lilac-settings to change.`
|
|
872
|
+
: `Preserved thinking OFF for ${entry.name} (${entry.flag}: ${flagValue}) — reasoning trimmed each turn (lighter; better for prose). Open /lilac-settings to change.`;
|
|
873
|
+
if (modelSelectNotifyTimer) clearTimeout(modelSelectNotifyTimer);
|
|
874
|
+
modelSelectNotifyTimer = setTimeout(() => {
|
|
875
|
+
modelSelectNotifyTimer = null;
|
|
876
|
+
try { ctx.ui.notify(msg, "info"); } catch { /* notify is a no-op without a UI runner */ }
|
|
877
|
+
}, MODEL_SELECT_NOTIFY_DELAY_MS);
|
|
878
|
+
}
|
|
879
|
+
|
|
865
880
|
// List-price models (patch applied, pre-discount), cached at module scope and
|
|
866
881
|
// rebuilt only when the base set changes (see cacheModels). Used to recompute
|
|
867
882
|
// the in-flight model's cost without compounding an already-applied discount.
|
|
@@ -981,6 +996,9 @@ export default function (pi: ExtensionAPI) {
|
|
|
981
996
|
// syncStatus reads the LIVE ctx.model so a switch away from lilac before the
|
|
982
997
|
// background fetch resolves never leaves a stale discount painted.
|
|
983
998
|
syncStatus(ctx);
|
|
999
|
+
// Show the preserved-thinking notification on first load / resume if the
|
|
1000
|
+
// active model carries a preserve flag (model_select may not fire on startup).
|
|
1001
|
+
notifyPreservedThinkingFor(ctx.model, ctx);
|
|
984
1002
|
|
|
985
1003
|
// Fire-and-forget: resolve API key, then fetch live data in background.
|
|
986
1004
|
// Provider and status are hot-swapped when results arrive.
|
|
@@ -1110,11 +1128,12 @@ export default function (pi: ExtensionAPI) {
|
|
|
1110
1128
|
syncStatus(ctx);
|
|
1111
1129
|
});
|
|
1112
1130
|
|
|
1113
|
-
pi.on("model_select", async (
|
|
1131
|
+
pi.on("model_select", async (event, ctx) => {
|
|
1114
1132
|
// ctx.model is the live session model (pi sets state.model before emitting
|
|
1115
1133
|
// this event), so syncStatus paints/clears consistently with every other
|
|
1116
1134
|
// handler — one source of truth for the footer.
|
|
1117
1135
|
syncStatus(ctx);
|
|
1136
|
+
notifyPreservedThinkingFor(event.model ?? ctx.model, ctx);
|
|
1118
1137
|
});
|
|
1119
1138
|
|
|
1120
1139
|
pi.on("session_tree", async (_event, ctx) => {
|
|
@@ -1153,7 +1172,7 @@ export default function (pi: ExtensionAPI) {
|
|
|
1153
1172
|
};
|
|
1154
1173
|
});
|
|
1155
1174
|
|
|
1156
|
-
//
|
|
1175
|
+
// Flex: gate interactive prompts on the active lilac model's discount.
|
|
1157
1176
|
// Only interactive prompts are gated — extension-injected messages are skipped
|
|
1158
1177
|
// (programmatic; would loop) and rpc/print are skipped (a silent block would be
|
|
1159
1178
|
// a confusing failure in automation). A missing discount entry or no data yet
|
|
@@ -1177,62 +1196,153 @@ export default function (pi: ExtensionAPI) {
|
|
|
1177
1196
|
? `${discountPercent}% discount`
|
|
1178
1197
|
: (latestDiscounts ? "no discount on this model (list price)" : "no discount data yet");
|
|
1179
1198
|
ctx.ui.notify(
|
|
1180
|
-
`
|
|
1199
|
+
`Flex: ${desc} < ${threshold}% threshold — blocked until the next poll. Re-submit once the discount improves, or open /lilac-settings to adjust.`,
|
|
1181
1200
|
"warning",
|
|
1182
1201
|
);
|
|
1183
1202
|
triggerDiscountRefresh(ctx);
|
|
1184
1203
|
return { action: "handled" };
|
|
1185
1204
|
});
|
|
1186
1205
|
|
|
1187
|
-
|
|
1188
|
-
|
|
1189
|
-
|
|
1190
|
-
|
|
1191
|
-
|
|
1192
|
-
|
|
1206
|
+
// ─── /lilac-settings: settings UI (mirrors pi core /settings) ──────────────
|
|
1207
|
+
// Opens a SettingsList (lazy-imported from pi-tui) via ctx.ui.custom(). Toggles
|
|
1208
|
+
// write to ~/.pi/agent/extensions/lilac.json (modelOverrides for preserved
|
|
1209
|
+
// thinking; flexThreshold via the shared applyFlexThreshold helper), refresh
|
|
1210
|
+
// the in-memory config, bust the list-price model cache, and re-register the
|
|
1211
|
+
// provider so the change takes effect immediately.
|
|
1212
|
+
function collectPreserveState(): Array<{ id: string; name: string; flag: "clear_thinking" | "preserve_thinking"; preserved: boolean }> {
|
|
1213
|
+
const resolved = buildModels(loadStaleModels(embeddedModels), customModels, patches, activeOverrides());
|
|
1214
|
+
const out: Array<{ id: string; name: string; flag: "clear_thinking" | "preserve_thinking"; preserved: boolean }> = [];
|
|
1215
|
+
for (const m of resolved) {
|
|
1216
|
+
const kwargs = (m as any).compat?.chatTemplateKwargs;
|
|
1217
|
+
if (!kwargs || typeof kwargs !== "object") continue;
|
|
1218
|
+
if (typeof kwargs.clear_thinking === "boolean") {
|
|
1219
|
+
out.push({ id: m.id, name: (m as any).name || m.id, flag: "clear_thinking", preserved: kwargs.clear_thinking === false });
|
|
1220
|
+
} else if (typeof kwargs.preserve_thinking === "boolean") {
|
|
1221
|
+
out.push({ id: m.id, name: (m as any).name || m.id, flag: "preserve_thinking", preserved: kwargs.preserve_thinking === true });
|
|
1193
1222
|
}
|
|
1223
|
+
}
|
|
1224
|
+
return out;
|
|
1225
|
+
}
|
|
1194
1226
|
|
|
1195
|
-
|
|
1196
|
-
|
|
1197
|
-
|
|
1198
|
-
|
|
1199
|
-
|
|
1200
|
-
return;
|
|
1201
|
-
}
|
|
1202
|
-
applyFlexThreshold(value, ctx);
|
|
1227
|
+
pi.registerCommand("lilac-settings", {
|
|
1228
|
+
description: "Configure Lilac: flex discount threshold + preserved thinking per model",
|
|
1229
|
+
async handler(_args, ctx) {
|
|
1230
|
+
if (ctx.mode !== "tui") {
|
|
1231
|
+
ctx.ui.notify("/lilac-settings requires TUI mode.", "error");
|
|
1203
1232
|
return;
|
|
1204
1233
|
}
|
|
1205
|
-
|
|
1206
|
-
const
|
|
1207
|
-
|
|
1208
|
-
|
|
1209
|
-
|
|
1210
|
-
|
|
1211
|
-
"
|
|
1212
|
-
|
|
1213
|
-
|
|
1214
|
-
|
|
1215
|
-
|
|
1216
|
-
|
|
1217
|
-
|
|
1218
|
-
|
|
1219
|
-
|
|
1220
|
-
|
|
1221
|
-
|
|
1222
|
-
|
|
1223
|
-
|
|
1224
|
-
|
|
1225
|
-
|
|
1226
|
-
|
|
1227
|
-
|
|
1228
|
-
|
|
1229
|
-
|
|
1230
|
-
|
|
1234
|
+
const { SettingsList, Container } = await import("@earendil-works/pi-tui");
|
|
1235
|
+
const { getSettingsListTheme, DynamicBorder } = await import("@earendil-works/pi-coding-agent");
|
|
1236
|
+
|
|
1237
|
+
const currentFlex = getConfig().flexThreshold ?? null;
|
|
1238
|
+
|
|
1239
|
+
await ctx.ui.custom((_tui, theme, _kb, done) => {
|
|
1240
|
+
const border = () => new DynamicBorder((s: string) => theme.fg("border", s));
|
|
1241
|
+
// SettingsList left-aligns the value column after the widest label (capped
|
|
1242
|
+
// at 30 cols). A label wider than 30 shifts that row's value out of
|
|
1243
|
+
// alignment, so cap model-name labels.
|
|
1244
|
+
const truncateLabel = (s: string) => (s.length > 30 ? s.slice(0, 27) + "..." : s);
|
|
1245
|
+
|
|
1246
|
+
const items: any[] = [
|
|
1247
|
+
{
|
|
1248
|
+
id: "flex",
|
|
1249
|
+
label: "Flex threshold",
|
|
1250
|
+
description: "Only respond when the active model's discount is at/above this. 'off' allows all. (Custom values: edit ~/.pi/agent/extensions/lilac.json.)",
|
|
1251
|
+
currentValue: currentFlex == null ? "off" : String(currentFlex),
|
|
1252
|
+
values: ["off", "50", "75"],
|
|
1253
|
+
},
|
|
1254
|
+
{
|
|
1255
|
+
id: "preserved-thinking",
|
|
1256
|
+
label: "Preserved thinking",
|
|
1257
|
+
description: "Per-model Preserve Thinking / Clear Thinking (full-history reasoning). Preserve Thinking keeps all turns' reasoning; Clear Thinking lets the template drop older reasoning (saves tokens, can hurt multi-turn recall / cause overthinking).",
|
|
1258
|
+
currentValue: "configure",
|
|
1259
|
+
submenu: (_currentValue: string, subDone: (v?: string) => void) => {
|
|
1260
|
+
// Re-read state on each open so toggles from a previous visit (which
|
|
1261
|
+
// wrote lilac.json + refreshed config) are reflected — a snapshot
|
|
1262
|
+
// captured at panel-open time would show stale values after a toggle.
|
|
1263
|
+
const fresh = collectPreserveState();
|
|
1264
|
+
const subItems = fresh.map((e) => ({
|
|
1265
|
+
id: `preserve:${e.id}`,
|
|
1266
|
+
label: truncateLabel(e.name),
|
|
1267
|
+
description: `${e.id} — Preserve Thinking keeps full reasoning history across turns; Clear Thinking lets the template drop older reasoning (saves tokens, can hurt multi-turn recall / cause overthinking).`,
|
|
1268
|
+
currentValue: e.preserved ? "Preserve Thinking" : "Clear Thinking",
|
|
1269
|
+
values: ["Preserve Thinking", "Clear Thinking"],
|
|
1270
|
+
}));
|
|
1271
|
+
const subList = new SettingsList(
|
|
1272
|
+
subItems,
|
|
1273
|
+
Math.min(subItems.length + 2, 15),
|
|
1274
|
+
getSettingsListTheme(),
|
|
1275
|
+
(id: string, newValue: string) => {
|
|
1276
|
+
const modelId = id.slice("preserve:".length);
|
|
1277
|
+
const entry = fresh.find((p) => p.id === modelId);
|
|
1278
|
+
if (!entry) return;
|
|
1279
|
+
const preservedOn = newValue === "Preserve Thinking";
|
|
1280
|
+
const flagValue = entry.flag === "clear_thinking" ? !preservedOn : preservedOn;
|
|
1281
|
+
updateConfig((cfg) => {
|
|
1282
|
+
const overrides = cfg.modelOverrides ?? (cfg.modelOverrides = {});
|
|
1283
|
+
const ov = overrides[modelId] ?? (overrides[modelId] = {});
|
|
1284
|
+
const compat = ov.compat ?? (ov.compat = {});
|
|
1285
|
+
const kwargs = compat.chatTemplateKwargs ?? (compat.chatTemplateKwargs = {});
|
|
1286
|
+
kwargs[entry.flag] = flagValue;
|
|
1287
|
+
return cfg;
|
|
1288
|
+
});
|
|
1289
|
+
listModelsCache = null;
|
|
1290
|
+
pi.registerProvider("lilac", {
|
|
1291
|
+
baseUrl: BASE_URL,
|
|
1292
|
+
apiKey: "$LILAC_API_KEY",
|
|
1293
|
+
api: "openai-completions",
|
|
1294
|
+
models: applyDiscounts(getListModels(), latestDiscounts),
|
|
1295
|
+
});
|
|
1296
|
+
syncStatus(ctx);
|
|
1297
|
+
ctx.ui.notify(`Preserved thinking ${preservedOn ? "on" : "off"} for ${entry.name} — takes effect now.`, "info");
|
|
1298
|
+
},
|
|
1299
|
+
() => subDone(),
|
|
1300
|
+
{ enableSearch: true },
|
|
1301
|
+
);
|
|
1302
|
+
// The outer container's borders already frame the panel; return the
|
|
1303
|
+
// list directly so we don't render a second border pair.
|
|
1304
|
+
return subList;
|
|
1305
|
+
},
|
|
1306
|
+
},
|
|
1307
|
+
];
|
|
1308
|
+
|
|
1309
|
+
const container = new Container();
|
|
1310
|
+
container.addChild(border());
|
|
1311
|
+
|
|
1312
|
+
const settingsList = new SettingsList(
|
|
1313
|
+
items,
|
|
1314
|
+
Math.min(items.length + 2, 15),
|
|
1315
|
+
getSettingsListTheme(),
|
|
1316
|
+
(id: string, newValue: string) => {
|
|
1317
|
+
if (id === "flex") {
|
|
1318
|
+
const value = newValue === "off" ? null : Number(newValue);
|
|
1319
|
+
applyFlexThreshold(value, ctx);
|
|
1320
|
+
}
|
|
1321
|
+
},
|
|
1322
|
+
() => done(undefined),
|
|
1323
|
+
{ enableSearch: true },
|
|
1324
|
+
);
|
|
1325
|
+
container.addChild(settingsList);
|
|
1326
|
+
container.addChild(border());
|
|
1327
|
+
|
|
1328
|
+
return {
|
|
1329
|
+
render(width: number) {
|
|
1330
|
+
return container.render(width);
|
|
1331
|
+
},
|
|
1332
|
+
invalidate() {
|
|
1333
|
+
container.invalidate();
|
|
1334
|
+
},
|
|
1335
|
+
handleInput(data: string) {
|
|
1336
|
+
settingsList.handleInput?.(data);
|
|
1337
|
+
},
|
|
1338
|
+
};
|
|
1339
|
+
});
|
|
1231
1340
|
},
|
|
1232
1341
|
});
|
|
1233
1342
|
|
|
1234
1343
|
pi.on("session_shutdown", () => {
|
|
1235
1344
|
revalidateAbort?.abort();
|
|
1345
|
+
if (modelSelectNotifyTimer) { clearTimeout(modelSelectNotifyTimer); modelSelectNotifyTimer = null; }
|
|
1236
1346
|
if (pollInterval) {
|
|
1237
1347
|
clearInterval(pollInterval);
|
|
1238
1348
|
pollInterval = null;
|
|
@@ -1240,5 +1350,5 @@ export default function (pi: ExtensionAPI) {
|
|
|
1240
1350
|
});
|
|
1241
1351
|
}
|
|
1242
1352
|
|
|
1243
|
-
export { fetchStatusDiscounts, applyDiscounts, applyDiscountInPlace, loadCachedDiscounts, cacheDiscounts, buildModels, applyModelOverride, parseModelOverrides, parseFlexThreshold, updateConfig, loadConfig, getConfig };
|
|
1353
|
+
export { fetchStatusDiscounts, applyDiscounts, applyDiscountInPlace, loadCachedDiscounts, cacheDiscounts, buildModels, applyModelOverride, parseModelOverrides, parseFlexThreshold, updateConfig, loadConfig, getConfig, applyFlexThreshold };
|
|
1244
1354
|
export type { JsonDiscount, JsonModel, PatchEntry, PatchData, ModelOverride, LilacConfig };
|
package/models.json
CHANGED
|
@@ -23,29 +23,6 @@
|
|
|
23
23
|
"zaiToolStream": true
|
|
24
24
|
}
|
|
25
25
|
},
|
|
26
|
-
{
|
|
27
|
-
"id": "zai-org/glm-5.1",
|
|
28
|
-
"name": "GLM 5.1",
|
|
29
|
-
"reasoning": true,
|
|
30
|
-
"input": [
|
|
31
|
-
"text"
|
|
32
|
-
],
|
|
33
|
-
"cost": {
|
|
34
|
-
"input": 0.9,
|
|
35
|
-
"output": 3,
|
|
36
|
-
"cacheRead": 0.27,
|
|
37
|
-
"cacheWrite": 0
|
|
38
|
-
},
|
|
39
|
-
"contextWindow": 202752,
|
|
40
|
-
"maxTokens": 131072,
|
|
41
|
-
"compat": {
|
|
42
|
-
"supportsDeveloperRole": false,
|
|
43
|
-
"supportsStore": false,
|
|
44
|
-
"maxTokensField": "max_completion_tokens",
|
|
45
|
-
"thinkingFormat": "qwen-chat-template",
|
|
46
|
-
"zaiToolStream": true
|
|
47
|
-
}
|
|
48
|
-
},
|
|
49
26
|
{
|
|
50
27
|
"id": "zai-org/glm-5.2",
|
|
51
28
|
"name": "GLM 5.2",
|
|
@@ -92,28 +69,6 @@
|
|
|
92
69
|
"zaiToolStream": true
|
|
93
70
|
}
|
|
94
71
|
},
|
|
95
|
-
{
|
|
96
|
-
"id": "minimaxai/minimax-m2.7",
|
|
97
|
-
"name": "MiniMax M2.7",
|
|
98
|
-
"reasoning": true,
|
|
99
|
-
"input": [
|
|
100
|
-
"text"
|
|
101
|
-
],
|
|
102
|
-
"cost": {
|
|
103
|
-
"input": 0.3,
|
|
104
|
-
"output": 1.2,
|
|
105
|
-
"cacheRead": 0.055,
|
|
106
|
-
"cacheWrite": 0
|
|
107
|
-
},
|
|
108
|
-
"contextWindow": 204800,
|
|
109
|
-
"maxTokens": 204800,
|
|
110
|
-
"compat": {
|
|
111
|
-
"supportsDeveloperRole": true,
|
|
112
|
-
"supportsStore": false,
|
|
113
|
-
"maxTokensField": "max_completion_tokens",
|
|
114
|
-
"thinkingFormat": "qwen-chat-template"
|
|
115
|
-
}
|
|
116
|
-
},
|
|
117
72
|
{
|
|
118
73
|
"id": "minimaxai/minimax-m3",
|
|
119
74
|
"name": "MiniMax M3",
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "pi-lilac-provider",
|
|
3
|
-
"version": "1.
|
|
3
|
+
"version": "1.8.1",
|
|
4
4
|
"description": "Lilac provider extension for pi - Access Kimi K2.6, GLM 5.1, and Gemma 4 models through Lilac's OpenAI-compatible API on idle GPUs",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"main": "index.ts",
|
package/patch.json
CHANGED
package/scripts/test-flex.ts
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
#!/usr/bin/env node
|
|
2
2
|
/**
|
|
3
|
-
* Tests for the
|
|
3
|
+
* Tests for the Lilac flex feature: discount-threshold gating.
|
|
4
4
|
*
|
|
5
5
|
* Verifies, against the REAL exported helpers and registered handlers/commands
|
|
6
6
|
* from index.ts (not a re-implementation):
|
|
@@ -14,9 +14,9 @@
|
|
|
14
14
|
* discount < threshold; allows (continue) when >= threshold, when flex is
|
|
15
15
|
* off, for non-lilac models, for non-interactive sources, and for models
|
|
16
16
|
* with no discount entry (treated as 0% -> blocked when flex on).
|
|
17
|
-
* -
|
|
18
|
-
*
|
|
19
|
-
*
|
|
17
|
+
* - applyFlexThreshold (flex configuration, now driven by the /lilac-settings
|
|
18
|
+
* flex row): sets flexThreshold, persists to disk, reports state against the
|
|
19
|
+
* live model's discount, and is visible to the input gate.
|
|
20
20
|
*
|
|
21
21
|
* Config FS + discount cache are isolated to a temp HOME so nothing touches the
|
|
22
22
|
* real ~/.pi.
|
|
@@ -29,7 +29,7 @@ import { getAgentDir, type ExtensionAPI } from "@earendil-works/pi-coding-agent"
|
|
|
29
29
|
// Isolate config + cache to a temp agent dir so loadConfig/cacheDiscounts never touch
|
|
30
30
|
// the real ~/.pi. Must be set before importing index.ts, which computes
|
|
31
31
|
// CONFIG_PATH / CACHE_PATH at module scope.
|
|
32
|
-
const tmpHome = `/tmp/pi-lilac-
|
|
32
|
+
const tmpHome = `/tmp/pi-lilac-test-${Date.now()}`;
|
|
33
33
|
fs.mkdirSync(tmpHome, { recursive: true });
|
|
34
34
|
process.env.HOME = tmpHome;
|
|
35
35
|
process.env.PI_CODING_AGENT_DIR = path.join(tmpHome, ".pi", "agent");
|
|
@@ -41,6 +41,7 @@ const {
|
|
|
41
41
|
getConfig,
|
|
42
42
|
updateConfig,
|
|
43
43
|
cacheDiscounts,
|
|
44
|
+
applyFlexThreshold,
|
|
44
45
|
} = await import("../index.ts");
|
|
45
46
|
|
|
46
47
|
let passed = 0;
|
|
@@ -167,7 +168,6 @@ cacheDiscounts(new Map([
|
|
|
167
168
|
]));
|
|
168
169
|
|
|
169
170
|
const handlers = new Map<string, ((...args: any[]) => any)[]>();
|
|
170
|
-
const commands = new Map<string, { handler: (...args: any[]) => any }>();
|
|
171
171
|
|
|
172
172
|
const mockApi: ExtensionAPI = {
|
|
173
173
|
registerProvider: () => {},
|
|
@@ -175,9 +175,7 @@ const mockApi: ExtensionAPI = {
|
|
|
175
175
|
if (!handlers.has(event)) handlers.set(event, []);
|
|
176
176
|
handlers.get(event)!.push(handler);
|
|
177
177
|
},
|
|
178
|
-
registerCommand: (
|
|
179
|
-
commands.set(name, opts);
|
|
180
|
-
},
|
|
178
|
+
registerCommand: () => {},
|
|
181
179
|
appendEntry: () => {},
|
|
182
180
|
exec: async () => ({ exitCode: 0, stdout: "", stderr: "" }),
|
|
183
181
|
} as any;
|
|
@@ -266,102 +264,59 @@ updateConfig((c) => ({ ...c, flexThreshold: 25 }));
|
|
|
266
264
|
eq(result, { action: "continue" }, "25% >= 25% threshold (inclusive) -> continue");
|
|
267
265
|
}
|
|
268
266
|
|
|
269
|
-
// ─── /lilac-
|
|
270
|
-
|
|
271
|
-
console.log("\n--- /lilac-flex command ---");
|
|
267
|
+
// ─── applyFlexThreshold (flex configuration, driven by the /lilac-settings row) ─
|
|
272
268
|
|
|
273
|
-
|
|
274
|
-
const cmd = commands.get("lilac-flex")!;
|
|
269
|
+
console.log("\n--- applyFlexThreshold ---");
|
|
275
270
|
|
|
276
|
-
function
|
|
271
|
+
function runApply(value: number | null, model: any = { id: KIMI, provider: "lilac" }) {
|
|
277
272
|
const notifications: { msg: string; level: string }[] = [];
|
|
278
273
|
const ctx = {
|
|
279
|
-
|
|
280
|
-
model: opts.model ?? { id: KIMI, provider: "lilac" },
|
|
274
|
+
model,
|
|
281
275
|
ui: {
|
|
282
276
|
notify: (msg: string, level: string) => notifications.push({ msg, level }),
|
|
283
277
|
setStatus: () => {},
|
|
284
278
|
theme: { fg: (_c: string, t: string) => t },
|
|
285
|
-
select: async (_title: string, _labels: string[]) => opts.select,
|
|
286
|
-
input: async (_title: string, _placeholder: string) => opts.input,
|
|
287
279
|
},
|
|
288
280
|
};
|
|
289
|
-
|
|
290
|
-
}
|
|
291
|
-
|
|
292
|
-
// direct numeric arg
|
|
293
|
-
{
|
|
294
|
-
const { flexThreshold } = await runCmd("75");
|
|
295
|
-
assert(flexThreshold === 75, "/lilac-flex 75 -> flexThreshold 75");
|
|
296
|
-
assert(JSON.parse(fs.readFileSync(cfgPath, "utf8")).flexThreshold === 75, "/lilac-flex 75 persisted to disk");
|
|
297
|
-
}
|
|
298
|
-
|
|
299
|
-
// trailing % accepted
|
|
300
|
-
{
|
|
301
|
-
const { flexThreshold } = await runCmd("50%");
|
|
302
|
-
assert(flexThreshold === 50, "/lilac-flex 50% -> flexThreshold 50");
|
|
303
|
-
}
|
|
304
|
-
|
|
305
|
-
// off keyword
|
|
306
|
-
{
|
|
307
|
-
const { flexThreshold } = await runCmd("off");
|
|
308
|
-
assert(flexThreshold === null, "/lilac-flex off -> flexThreshold null");
|
|
309
|
-
}
|
|
310
|
-
|
|
311
|
-
// invalid arg -> no change, usage notify
|
|
312
|
-
{
|
|
313
|
-
const { flexThreshold, notifications } = await runCmd("bogus");
|
|
314
|
-
assert(flexThreshold === null, "/lilac-flex bogus -> no change (still off)");
|
|
315
|
-
assert(notifications.some((n) => n.msg.includes("Usage")), "/lilac-flex bogus -> usage notify");
|
|
316
|
-
}
|
|
317
|
-
|
|
318
|
-
// picker: "≥ 75% discount" preset
|
|
319
|
-
{
|
|
320
|
-
const { flexThreshold } = await runCmd("", { select: "≥ 75% discount" });
|
|
321
|
-
assert(flexThreshold === 75, "picker '≥ 75% discount' -> 75");
|
|
322
|
-
}
|
|
323
|
-
|
|
324
|
-
// picker: "Off" preset
|
|
325
|
-
{
|
|
326
|
-
const { flexThreshold } = await runCmd("", { select: "Off — allow all discounts" });
|
|
327
|
-
assert(flexThreshold === null, "picker 'Off' -> null");
|
|
281
|
+
applyFlexThreshold(value, ctx);
|
|
282
|
+
return { notifications, flexThreshold: getConfig().flexThreshold ?? null };
|
|
328
283
|
}
|
|
329
284
|
|
|
330
|
-
//
|
|
285
|
+
// set 75 (kimi seeded at 25% -> below threshold -> blocked-warning message)
|
|
331
286
|
{
|
|
332
|
-
const { flexThreshold } =
|
|
333
|
-
assert(flexThreshold ===
|
|
287
|
+
const { flexThreshold, notifications } = runApply(75);
|
|
288
|
+
assert(flexThreshold === 75, "applyFlexThreshold(75) -> flexThreshold 75");
|
|
289
|
+
assert(JSON.parse(fs.readFileSync(cfgPath, "utf8")).flexThreshold === 75, "persisted to disk");
|
|
290
|
+
assert(notifications.some((n) => n.msg.includes("75%") && n.level === "warning"), "75 with kimi@25% -> warning, mentions threshold");
|
|
334
291
|
}
|
|
335
292
|
|
|
336
|
-
//
|
|
293
|
+
// off
|
|
337
294
|
{
|
|
338
|
-
const { flexThreshold, notifications } =
|
|
339
|
-
assert(flexThreshold ===
|
|
340
|
-
assert(notifications.some((n) => n.
|
|
295
|
+
const { flexThreshold, notifications } = runApply(null);
|
|
296
|
+
assert(flexThreshold === null, "applyFlexThreshold(null) -> null");
|
|
297
|
+
assert(notifications.some((n) => n.msg.includes("off")), "off -> notify mentions off");
|
|
341
298
|
}
|
|
342
299
|
|
|
343
|
-
//
|
|
300
|
+
// threshold below current discount -> allowed (info)
|
|
344
301
|
{
|
|
345
|
-
const {
|
|
346
|
-
assert(
|
|
347
|
-
assert(notifications.length === 0, "picker cancelled -> no notify");
|
|
302
|
+
const { notifications } = runApply(20);
|
|
303
|
+
assert(notifications.some((n) => n.level === "info" && n.msg.includes("allowed")), "20 with kimi@25% -> info, allowed");
|
|
348
304
|
}
|
|
349
305
|
|
|
350
|
-
//
|
|
306
|
+
// no discount data for model -> warning
|
|
351
307
|
{
|
|
352
|
-
const {
|
|
353
|
-
assert(
|
|
354
|
-
assert(notifications.some((n) => n.level === "error" && n.msg.includes("interactive")), "non-interactive -> error notify");
|
|
308
|
+
const { notifications } = runApply(75, { id: "zai-org/glm-5.1", provider: "lilac" });
|
|
309
|
+
assert(notifications.some((n) => n.level === "warning" && n.msg.includes("no discount data")), "uncached model + threshold -> warning, no discount data");
|
|
355
310
|
}
|
|
356
311
|
|
|
357
|
-
// setting flex via
|
|
312
|
+
// setting flex via applyFlexThreshold is visible to the gate (config cache in sync)
|
|
358
313
|
{
|
|
359
|
-
|
|
314
|
+
runApply(75);
|
|
360
315
|
const { result } = await runInput("interactive", { id: KIMI, provider: "lilac" });
|
|
361
|
-
eq(result, { action: "handled" }, "after
|
|
362
|
-
|
|
316
|
+
eq(result, { action: "handled" }, "after applyFlexThreshold(75), gate blocks kimi@25%");
|
|
317
|
+
runApply(null);
|
|
363
318
|
const { result: afterOff } = await runInput("interactive", { id: KIMI, provider: "lilac" });
|
|
364
|
-
eq(afterOff, { action: "continue" }, "after
|
|
319
|
+
eq(afterOff, { action: "continue" }, "after applyFlexThreshold(null), gate allows kimi@25%");
|
|
365
320
|
}
|
|
366
321
|
|
|
367
322
|
// ─── Summary ───────────────────────────────────────────────────────────────────
|
|
@@ -66,7 +66,6 @@ function eq<T>(actual: T, expected: T, message: string) {
|
|
|
66
66
|
|
|
67
67
|
const KIMI = "moonshotai/kimi-k2.6";
|
|
68
68
|
const GLM52 = "zai-org/glm-5.2";
|
|
69
|
-
const GLM51 = "zai-org/glm-5.1";
|
|
70
69
|
|
|
71
70
|
// ─── applyModelOverride ────────────────────────────────────────────────────────
|
|
72
71
|
|
|
@@ -240,16 +239,6 @@ function find(models: any[], id: string): any {
|
|
|
240
239
|
assert(glm.compat.chatTemplateKwargs.clear_thinking === false, "non-overridden clear_thinking survives a thinkingLevelMap-only override");
|
|
241
240
|
}
|
|
242
241
|
|
|
243
|
-
{
|
|
244
|
-
// Override on glm-5.1 toggles clear_thinking (patch sets false); other compat survives
|
|
245
|
-
const overrides = { [GLM51]: { compat: { chatTemplateKwargs: { clear_thinking: true } } } } as any;
|
|
246
|
-
const models = buildModels(modelsData, customModelsData, patchData, overrides);
|
|
247
|
-
const glm = find(models, GLM51);
|
|
248
|
-
assert(glm.compat.chatTemplateKwargs.clear_thinking === true, "override wins over patch: glm-5.1 clear_thinking -> true");
|
|
249
|
-
assert((glm.compat.chatTemplateKwargs as any).thinking?.$var === "thinking.enabled", "override deep-merges: glm-5.1 thinking $var key survives");
|
|
250
|
-
assert(glm.compat.zaiToolStream === true, "override deep-merges: glm-5.1 zaiToolStream survives");
|
|
251
|
-
}
|
|
252
|
-
|
|
253
242
|
{
|
|
254
243
|
// Override for an unknown id is a no-op (adds no models)
|
|
255
244
|
const before = buildModels(modelsData, customModelsData, patchData, {});
|
|
@@ -194,25 +194,15 @@ console.log("\n=== GLM 5.2 on the wire (real pi-ai) ===");
|
|
|
194
194
|
const high = await wire("zai-org/glm-5.2", "high");
|
|
195
195
|
eq(high.chat_template_kwargs, { enable_thinking: true, reasoning_effort: "high", clear_thinking: false },
|
|
196
196
|
"glm-5.2 @ high → enable_thinking+reasoning_effort AND clear_thinking false");
|
|
197
|
-
const
|
|
198
|
-
eq(
|
|
199
|
-
"glm-5.2 @
|
|
197
|
+
const max = await wire("zai-org/glm-5.2", "max");
|
|
198
|
+
eq(max.chat_template_kwargs, { enable_thinking: true, reasoning_effort: "max", clear_thinking: false },
|
|
199
|
+
"glm-5.2 @ max → reasoning_effort max, clear_thinking false");
|
|
200
200
|
const off = await wire("zai-org/glm-5.2", "off");
|
|
201
201
|
eq(off.chat_template_kwargs, { enable_thinking: false, clear_thinking: false },
|
|
202
202
|
"glm-5.2 @ off → reasoning_effort omitted (omitWhenOff), clear_thinking false persists");
|
|
203
203
|
}
|
|
204
204
|
|
|
205
|
-
console.log("\n===
|
|
206
|
-
{
|
|
207
|
-
const high = await wire("zai-org/glm-5.1", "high");
|
|
208
|
-
eq(high.chat_template_kwargs, { thinking: true, enable_thinking: true, clear_thinking: false },
|
|
209
|
-
"glm-5.1 @ high → thinking+enable_thinking true AND clear_thinking false");
|
|
210
|
-
const off = await wire("zai-org/glm-5.1", "off");
|
|
211
|
-
eq(off.chat_template_kwargs, { thinking: false, enable_thinking: false, clear_thinking: false },
|
|
212
|
-
"glm-5.1 @ off → thinking false but clear_thinking false persists");
|
|
213
|
-
}
|
|
214
|
-
|
|
215
|
-
console.log("\n=== Gemma 4 / MiniMax (no family-wide preserve flag — regression guard) ===");
|
|
205
|
+
console.log("\n=== Gemma / MiniMax (no family-wide preserve flag — regression guard) ===");
|
|
216
206
|
{
|
|
217
207
|
const gemma = await wire("google/gemma-4-31b-it", "high");
|
|
218
208
|
eq(gemma.chat_template_kwargs, { thinking: true, enable_thinking: true },
|
|
@@ -224,11 +214,6 @@ console.log("\n=== Gemma 4 / MiniMax (no family-wide preserve flag — regressio
|
|
|
224
214
|
eq(m3.chat_template_kwargs, { thinking_mode: "enabled" },
|
|
225
215
|
"minimax-m3 @ high → only thinking_mode (no preserve/clear flag)");
|
|
226
216
|
falsy(m3.chat_template_kwargs?.preserve_thinking, "minimax-m3 has no preserve_thinking");
|
|
227
|
-
|
|
228
|
-
const m27 = await wire("minimaxai/minimax-m2.7", "high");
|
|
229
|
-
eq(m27.chat_template_kwargs, { thinking: true, enable_thinking: true },
|
|
230
|
-
"minimax-m2.7 @ high → only thinking/enable_thinking (no preserve/clear flag)");
|
|
231
|
-
falsy(m27.chat_template_kwargs?.clear_thinking, "minimax-m2.7 has no clear_thinking");
|
|
232
217
|
}
|
|
233
218
|
|
|
234
219
|
globalThis.fetch = originalFetch;
|