pi-lilac-provider 1.7.2 → 1.8.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -23,7 +23,7 @@ Access Kimi K2.6, GLM 5.1, MiniMax M2.7, and Gemma 4 models through Lilac's Open
23
23
  - **Reasoning Models** — Chain-of-thought via `chat_template_kwargs` (all models)
24
24
  - **Vision Support** — Image input on Kimi K2.6 and Gemma 4
25
25
  - **Context Caching** — Cache read pricing on Kimi K2.6 and GLM 5.1
26
- - **Flex (Discount Gating)** — Only let the LLM respond when the active model's discount meets a threshold you set (`/lilac-flex`)
26
+ - **Flex (Discount Gating)** — Only let the LLM respond when the active model's discount meets a threshold you set (`/lilac-settings`)
27
27
  - **Idle GPU Scheduling** — Lilac leverages idle GPU capacity for cost-efficient inference
28
28
 
29
29
  ## Installation
@@ -74,10 +74,8 @@ pi
74
74
  | Model | Context | Vision | Reasoning | Input $/M | Cache Read $/M | Output $/M |
75
75
  |-------|---------|--------|-----------|-----------|-----------------|------------|
76
76
  | Gemma 4 | 262K | ✅ | ✅ | $0.11 | — | $0.35 |
77
- | GLM 5.1 | 203K | ❌ | ✅ | $0.90 | $0.27 | $3.00 |
78
77
  | GLM 5.2 | 524K | ❌ | ✅ | $0.90 | $0.27 | $3.00 |
79
78
  | Kimi K2.6 | 262K | ✅ | ✅ | $0.70 | $0.20 | $3.50 |
80
- | MiniMax M2.7 | 205K | ❌ | ✅ | $0.30 | $0.06 | $1.20 |
81
79
  | MiniMax M3 | 1.0M | ✅ | ✅ | $0.28 | $0.05 | $1.10 |
82
80
 
83
81
  *Costs are per million tokens. Prices subject to change — check [getlilac.com](https://getlilac.com/) for current pricing.*
@@ -121,7 +119,7 @@ model's chat template honors differs per family. The provider uses pi's
121
119
  Kimi K2.6, GLM 5.1, Gemma 4, and MiniMax M2.7 use the forward-compatible form
122
120
  that sends **both** `thinking` and `enable_thinking`, so whichever key the
123
121
  template honors is set. GLM 5.2 additionally maps pi's thinking levels to
124
- `reasoning_effort` (`high` = lower-latency, `xhigh` = `max`). MiniMax M3 uses
122
+ `reasoning_effort` (`high` = lower-latency, `max` = deepest). MiniMax M3 uses
125
123
  the `thinking_mode` enum, exposed as three pi thinking levels: `off` →
126
124
  `disabled` (never think), `minimal` → `adaptive` (the model decides), `high` →
127
125
  `enabled` (always think). Pi starts at `off` (`disabled`); cycle to `minimal`
@@ -220,7 +218,7 @@ Create `~/.pi/agent/extensions/lilac.json` (auto-populated with defaults on firs
220
218
  ```jsonc
221
219
  {
222
220
  // Only respond when the active model's discount is >= this percent. null = off.
223
- // See "Flex (Discount Gating)" below. Set interactively with /lilac-flex.
221
+ // See "Flex (Discount Gating)" below. Set via /lilac-settings.
224
222
  "flexThreshold": null,
225
223
  "modelOverrides": {
226
224
  // Disable full-history reasoning for kimi-k2.6 (e.g. to save tokens):
@@ -239,16 +237,13 @@ The full set of overridable fields matches the model schema (`compat`, `thinking
239
237
 
240
238
  Lilac's per-model discount fluctuates with idle-GPU supply. **Flex** lets you set a discount threshold so pi **only sends a prompt to the LLM when the active model's current discount is at or above it** — e.g. "only respond when the discount is ≥ 75%". Below the threshold, the prompt is blocked (dropped with a warning) until the next discount poll brings the discount back up. This is a spend-control feature: you only spend when supply is cheap.
241
239
 
242
- Set it interactively with the `/lilac-flex` command:
240
+ Set it via the **Flex threshold** row in [`/lilac-settings`](#settings-ui) — cycle `off` / `50` / `75`:
243
241
 
244
242
  ```
245
- /lilac-flex # picker: Off / ≥50% / ≥75% / Custom…
246
- /lilac-flex 75 # set threshold directly (only respond at ≥75% discount)
247
- /lilac-flex 50% # trailing % accepted on the command line
248
- /lilac-flex off # disable flex (allow all discounts)
243
+ /lilac-settings # open the settings panel → Flex threshold row
249
244
  ```
250
245
 
251
- The threshold persists in `~/.pi/agent/extensions/lilac.json` as `flexThreshold` (a number `0`–`100`, or `null` for off) alongside `modelOverrides`, so it survives restarts. `/lilac-flex` updates it live — no restart needed.
246
+ The threshold persists in `~/.pi/agent/extensions/lilac.json` as `flexThreshold` (a number `0`–`100`, or `null` for off) alongside `modelOverrides`, so it survives restarts. `/lilac-settings` updates it live — no restart needed. For a custom value (e.g. 60), edit `flexThreshold` in that file directly.
252
247
 
253
248
  Behavior notes:
254
249
 
@@ -258,6 +253,17 @@ Behavior notes:
258
253
  - **Freshness:** when a prompt is blocked, the extension triggers an immediate `/status` refresh (throttled to once per ~5s) so you're not stuck on a stale low value from the 5-minute idle poll. The next submission sees the fresh discount. The footer status reflects the gate: `… · flex ≥75% ok` or `… · flex ≥75% blocked`.
259
254
  - **Discount lock-in:** per Lilac, a discount is locked in when a request starts. Flex gates on the best-known discount at submit time, which is what gets locked in for that turn.
260
255
 
256
+ ### Settings UI
257
+
258
+ `/lilac-settings` opens an interactive settings panel (mirrors pi core's `/settings` — bordered `SettingsList`, Esc to go back) to configure Lilac without editing JSON by hand:
259
+
260
+ - **Flex threshold** (`off` / `50` / `75`) — the spend-control gate from [Flex (Discount Gating)](#flex-discount-gating); custom values via `lilac.json`.
261
+ - **Preserved thinking** (nested submenu, one row per model) — toggles `clear_thinking` (GLM-5.2 / GLM-5.1) / `preserve_thinking` (Kimi K2.6) in `modelOverrides` between **Preserve Thinking** (keep full reasoning history across turns; the default, `clear_thinking: false` / `preserve_thinking: true`) and **Clear Thinking** (let the template drop older reasoning; saves tokens, but can degrade multi-turn recall / cause overthinking).
262
+
263
+ Changes write to `~/.pi/agent/extensions/lilac.json`, refresh the in-memory config, and re-register the provider, so they take effect immediately — no restart needed.
264
+
265
+ When you switch to — or start pi on — a Lilac model that carries a preserved-thinking flag (e.g. GLM-5.2 / GLM-5.1, Kimi K2.6), an info notification reports the state and how to change it, e.g. `Preserved thinking ON for glm-5.2 (clear_thinking: false) — suited for coding, but not for prose. Open /lilac-settings to change.` (OFF reads `... reasoning trimmed each turn (lighter; better for prose) ...`). It's an ordinary info notification (not a warning), so it doesn't paint bright yellow.
266
+
261
267
  ## Updating Models
262
268
 
263
269
  Run the update script to fetch the latest models from Lilac's API:
package/index.ts CHANGED
@@ -21,7 +21,7 @@
21
21
  * MiniMax M3). The forward-compatible form sends BOTH `thinking` and
22
22
  * `enable_thinking` so whichever key the template honors is set:
23
23
  * { chat_template_kwargs: { thinking: <bool>, enable_thinking: <bool> } }
24
- * GLM 5.2 adds `reasoning_effort` (high = lower-latency, xhigh = max) via a
24
+ * GLM 5.2 adds `reasoning_effort` (high for lower-latency, max for deep) via a
25
25
  * thinkingLevelMap. MiniMax M3 maps to the `thinking_mode` enum as three pi
26
26
  * thinking levels — off→disabled, minimal→adaptive (model decides), high→enabled
27
27
  * — so adaptive is selectable via pi's Shift+Tab cycle (off→minimal→high). Pi
@@ -99,7 +99,7 @@ interface JsonDiscount {
99
99
  creditMultiplier: number;
100
100
  }
101
101
 
102
- // Maps pi's thinking levels (off, minimal, low, medium, high, xhigh) to the
102
+ // Maps pi's thinking levels (off, minimal, low, medium, high, xhigh, max) to the
103
103
  // provider-specific effort string sent on the wire. A `null` value marks a
104
104
  // level as unsupported — clampThinkingLevel skips it when resolving the
105
105
  // user's selection. Mirrors pi-ai's ThinkingLevelMap shape.
@@ -110,6 +110,7 @@ type ThinkingLevelMap = {
110
110
  medium?: string | null;
111
111
  high?: string | null;
112
112
  xhigh?: string | null;
113
+ max?: string | null;
113
114
  };
114
115
 
115
116
  // A chat_template_kwargs value, mirroring pi-ai's ChatTemplateKwargSchema. Scalar
@@ -193,7 +194,7 @@ interface LilacConfig {
193
194
  modelOverrides?: Record<string, ModelOverride>;
194
195
  // Flex discount threshold: only allow interactive prompts to reach the LLM
195
196
  // when the active lilac model's discountPercent is >= this. null/undefined =
196
- // off (allow all). Set via /lilac-flex; persisted in lilac.json.
197
+ // off (allow all). Set via /lilac-settings; persisted in lilac.json.
197
198
  flexThreshold?: number | null;
198
199
  }
199
200
 
@@ -312,12 +313,12 @@ function getConfig(): LilacConfig {
312
313
  }
313
314
 
314
315
  // Read-modify-write the config file and refresh the in-memory cache. Used by
315
- // /lilac-flex so a threshold change takes effect immediately (the input gate and
316
+ // /lilac-settings so a threshold change takes effect immediately (the input gate and
316
317
  // footer read getConfig()) without a restart, and without clobbering the user's
317
318
  // modelOverrides. Reads via loadConfig (which validates), so modelOverrides the
318
319
  // user hand-edited since startup survive the spread. The file is normalized to a
319
320
  // discoverable shape (modelOverrides: {}, flexThreshold: null always present) so
320
- // /lilac-flex never strips the modelOverrides scaffold from the file.
321
+ // /lilac-settings never strips the modelOverrides scaffold from the file.
321
322
  function updateConfig(mutator: (cfg: LilacConfig) => LilacConfig): LilacConfig {
322
323
  const next = mutator(loadConfig());
323
324
  const toWrite: LilacConfig = {
@@ -760,20 +761,9 @@ function syncStatus(ctx: any): void {
760
761
  }
761
762
  }
762
763
 
763
- // Parse a /lilac-flex argument or custom-threshold input: "off"/"none"/"" → null
764
- // (disabled), a number in [0,100] (optionally with a trailing %) → that number,
765
- // anything else → undefined (invalid).
766
- function parseFlexArg(raw: string): number | null | undefined {
767
- const s = raw.trim().toLowerCase().replace(/%$/, "");
768
- if (s === "" || s === "off" || s === "none") return null;
769
- const n = Number(s);
770
- if (!Number.isFinite(n) || n < 0 || n > 100) return undefined;
771
- return Math.round(n);
772
- }
773
-
774
764
  // Persist a flex threshold via updateConfig and report the resulting state against
775
765
  // the live model's current discount. The footer is re-painted (syncStatus) so the
776
- // flex indicator appears immediately. Used by the /lilac-flex command.
766
+ // flex indicator appears immediately. Used by the /lilac-settings flex row.
777
767
  function applyFlexThreshold(value: number | null, ctx: any): void {
778
768
  updateConfig((cfg) => ({ ...cfg, flexThreshold: value }));
779
769
  const threshold = getConfig().flexThreshold ?? null;
@@ -784,16 +774,16 @@ function applyFlexThreshold(value: number | null, ctx: any): void {
784
774
  let msg: string;
785
775
  let level: "info" | "warning";
786
776
  if (threshold == null) {
787
- msg = "lilac-flex: off — all discounts allowed";
777
+ msg = "Flex: off — all discounts allowed";
788
778
  level = "info";
789
779
  } else if (pct == null) {
790
- msg = `lilac-flex: ≥ ${threshold}% — no discount data yet; will block until the next poll`;
780
+ msg = `Flex: ≥ ${threshold}% — no discount data yet; will block until the next poll`;
791
781
  level = "warning";
792
782
  } else if (pct >= threshold) {
793
- msg = `lilac-flex: ≥ ${threshold}% — current ${pct}% discount allowed`;
783
+ msg = `Flex: ≥ ${threshold}% — current ${pct}% discount allowed`;
794
784
  level = "info";
795
785
  } else {
796
- msg = `lilac-flex: ≥ ${threshold}% — current ${pct}% discount blocked until it improves`;
786
+ msg = `Flex: ≥ ${threshold}% — current ${pct}% discount blocked until it improves`;
797
787
  level = "warning";
798
788
  }
799
789
  try {
@@ -862,6 +852,31 @@ export default function (pi: ExtensionAPI) {
862
852
  const customModels = customModelsData as JsonModel[];
863
853
  const patches = patchData as PatchData;
864
854
 
855
+ // Deferred model_select notify timer — see the model_select handler. Cleared on
856
+ // rapid re-switch and on session_shutdown so only the latest switch notifies.
857
+ let modelSelectNotifyTimer: ReturnType<typeof setTimeout> | null = null;
858
+ const MODEL_SELECT_NOTIFY_DELAY_MS = 250;
859
+
860
+ // Notify preserved-thinking state for a preserve-flag model. Computed from the
861
+ // build pipeline (config as source of truth, not event.model.compat), deferred
862
+ // so pi core's (and other extensions') notifications land first, and cancelled
863
+ // on re-switch/shutdown so only the latest shows. Always level "info" (not a
864
+ // warning) — the text conveys the coding/prose tradeoff.
865
+ function notifyPreservedThinkingFor(model: any, ctx: any): void {
866
+ if (!model || model.provider !== PROVIDER_ID) return;
867
+ const entry = collectPreserveState().find((e: any) => e.id === model.id);
868
+ if (!entry) return;
869
+ const flagValue = entry.flag === "clear_thinking" ? !entry.preserved : entry.preserved;
870
+ const msg = entry.preserved
871
+ ? `Preserved thinking ON for ${entry.name} (${entry.flag}: ${flagValue}) — suited for coding, but not for prose. Open /lilac-settings to change.`
872
+ : `Preserved thinking OFF for ${entry.name} (${entry.flag}: ${flagValue}) — reasoning trimmed each turn (lighter; better for prose). Open /lilac-settings to change.`;
873
+ if (modelSelectNotifyTimer) clearTimeout(modelSelectNotifyTimer);
874
+ modelSelectNotifyTimer = setTimeout(() => {
875
+ modelSelectNotifyTimer = null;
876
+ try { ctx.ui.notify(msg, "info"); } catch { /* notify is a no-op without a UI runner */ }
877
+ }, MODEL_SELECT_NOTIFY_DELAY_MS);
878
+ }
879
+
865
880
  // List-price models (patch applied, pre-discount), cached at module scope and
866
881
  // rebuilt only when the base set changes (see cacheModels). Used to recompute
867
882
  // the in-flight model's cost without compounding an already-applied discount.
@@ -981,6 +996,9 @@ export default function (pi: ExtensionAPI) {
981
996
  // syncStatus reads the LIVE ctx.model so a switch away from lilac before the
982
997
  // background fetch resolves never leaves a stale discount painted.
983
998
  syncStatus(ctx);
999
+ // Show the preserved-thinking notification on first load / resume if the
1000
+ // active model carries a preserve flag (model_select may not fire on startup).
1001
+ notifyPreservedThinkingFor(ctx.model, ctx);
984
1002
 
985
1003
  // Fire-and-forget: resolve API key, then fetch live data in background.
986
1004
  // Provider and status are hot-swapped when results arrive.
@@ -1110,11 +1128,12 @@ export default function (pi: ExtensionAPI) {
1110
1128
  syncStatus(ctx);
1111
1129
  });
1112
1130
 
1113
- pi.on("model_select", async (_event, ctx) => {
1131
+ pi.on("model_select", async (event, ctx) => {
1114
1132
  // ctx.model is the live session model (pi sets state.model before emitting
1115
1133
  // this event), so syncStatus paints/clears consistently with every other
1116
1134
  // handler — one source of truth for the footer.
1117
1135
  syncStatus(ctx);
1136
+ notifyPreservedThinkingFor(event.model ?? ctx.model, ctx);
1118
1137
  });
1119
1138
 
1120
1139
  pi.on("session_tree", async (_event, ctx) => {
@@ -1153,7 +1172,7 @@ export default function (pi: ExtensionAPI) {
1153
1172
  };
1154
1173
  });
1155
1174
 
1156
- // lilac-flex: gate interactive prompts on the active lilac model's discount.
1175
+ // Flex: gate interactive prompts on the active lilac model's discount.
1157
1176
  // Only interactive prompts are gated — extension-injected messages are skipped
1158
1177
  // (programmatic; would loop) and rpc/print are skipped (a silent block would be
1159
1178
  // a confusing failure in automation). A missing discount entry or no data yet
@@ -1177,62 +1196,153 @@ export default function (pi: ExtensionAPI) {
1177
1196
  ? `${discountPercent}% discount`
1178
1197
  : (latestDiscounts ? "no discount on this model (list price)" : "no discount data yet");
1179
1198
  ctx.ui.notify(
1180
- `lilac-flex: ${desc} < ${threshold}% threshold — blocked until the next poll. Re-submit once the discount improves, or run /lilac-flex to adjust.`,
1199
+ `Flex: ${desc} < ${threshold}% threshold — blocked until the next poll. Re-submit once the discount improves, or open /lilac-settings to adjust.`,
1181
1200
  "warning",
1182
1201
  );
1183
1202
  triggerDiscountRefresh(ctx);
1184
1203
  return { action: "handled" };
1185
1204
  });
1186
1205
 
1187
- pi.registerCommand("lilac-flex", {
1188
- description: "Set lilac flex discount threshold — only respond at/above this discount",
1189
- async handler(args, ctx) {
1190
- if (!ctx.hasUI) {
1191
- ctx.ui.notify("/lilac-flex requires interactive mode.", "error");
1192
- return;
1206
+ // ─── /lilac-settings: settings UI (mirrors pi core /settings) ──────────────
1207
+ // Opens a SettingsList (lazy-imported from pi-tui) via ctx.ui.custom(). Toggles
1208
+ // write to ~/.pi/agent/extensions/lilac.json (modelOverrides for preserved
1209
+ // thinking; flexThreshold via the shared applyFlexThreshold helper), refresh
1210
+ // the in-memory config, bust the list-price model cache, and re-register the
1211
+ // provider so the change takes effect immediately.
1212
+ function collectPreserveState(): Array<{ id: string; name: string; flag: "clear_thinking" | "preserve_thinking"; preserved: boolean }> {
1213
+ const resolved = buildModels(loadStaleModels(embeddedModels), customModels, patches, activeOverrides());
1214
+ const out: Array<{ id: string; name: string; flag: "clear_thinking" | "preserve_thinking"; preserved: boolean }> = [];
1215
+ for (const m of resolved) {
1216
+ const kwargs = (m as any).compat?.chatTemplateKwargs;
1217
+ if (!kwargs || typeof kwargs !== "object") continue;
1218
+ if (typeof kwargs.clear_thinking === "boolean") {
1219
+ out.push({ id: m.id, name: (m as any).name || m.id, flag: "clear_thinking", preserved: kwargs.clear_thinking === false });
1220
+ } else if (typeof kwargs.preserve_thinking === "boolean") {
1221
+ out.push({ id: m.id, name: (m as any).name || m.id, flag: "preserve_thinking", preserved: kwargs.preserve_thinking === true });
1193
1222
  }
1223
+ }
1224
+ return out;
1225
+ }
1194
1226
 
1195
- const arg = (args ?? "").trim();
1196
- if (arg) {
1197
- const value = parseFlexArg(arg);
1198
- if (value === undefined) {
1199
- ctx.ui.notify("Usage: /lilac-flex [off | 0-100]", "warning");
1200
- return;
1201
- }
1202
- applyFlexThreshold(value, ctx);
1227
+ pi.registerCommand("lilac-settings", {
1228
+ description: "Configure Lilac: flex discount threshold + preserved thinking per model",
1229
+ async handler(_args, ctx) {
1230
+ if (ctx.mode !== "tui") {
1231
+ ctx.ui.notify("/lilac-settings requires TUI mode.", "error");
1203
1232
  return;
1204
1233
  }
1205
-
1206
- const current = getConfig().flexThreshold ?? null;
1207
- const labels = [
1208
- "Off — allow all discounts",
1209
- "≥ 50% discount",
1210
- "≥ 75% discount",
1211
- "Custom…",
1212
- ];
1213
- const choice = await ctx.ui.select("Lilac flex: only respond at/above this discount", labels);
1214
- if (choice === undefined) return; // cancelled
1215
-
1216
- let value: number | null;
1217
- if (choice === labels[0]) value = null;
1218
- else if (choice === labels[1]) value = 50;
1219
- else if (choice === labels[2]) value = 75;
1220
- else {
1221
- const input = await ctx.ui.input("Threshold (0–100, or 'off'):", String(current ?? ""));
1222
- if (input === undefined) return; // cancelled
1223
- const parsed = parseFlexArg(input);
1224
- if (parsed === undefined) {
1225
- ctx.ui.notify("Invalid threshold. Use a number 0–100 or 'off'.", "warning");
1226
- return;
1227
- }
1228
- value = parsed;
1229
- }
1230
- applyFlexThreshold(value, ctx);
1234
+ const { SettingsList, Container } = await import("@earendil-works/pi-tui");
1235
+ const { getSettingsListTheme, DynamicBorder } = await import("@earendil-works/pi-coding-agent");
1236
+
1237
+ const currentFlex = getConfig().flexThreshold ?? null;
1238
+
1239
+ await ctx.ui.custom((_tui, theme, _kb, done) => {
1240
+ const border = () => new DynamicBorder((s: string) => theme.fg("border", s));
1241
+ // SettingsList left-aligns the value column after the widest label (capped
1242
+ // at 30 cols). A label wider than 30 shifts that row's value out of
1243
+ // alignment, so cap model-name labels.
1244
+ const truncateLabel = (s: string) => (s.length > 30 ? s.slice(0, 27) + "..." : s);
1245
+
1246
+ const items: any[] = [
1247
+ {
1248
+ id: "flex",
1249
+ label: "Flex threshold",
1250
+ description: "Only respond when the active model's discount is at/above this. 'off' allows all. (Custom values: edit ~/.pi/agent/extensions/lilac.json.)",
1251
+ currentValue: currentFlex == null ? "off" : String(currentFlex),
1252
+ values: ["off", "50", "75"],
1253
+ },
1254
+ {
1255
+ id: "preserved-thinking",
1256
+ label: "Preserved thinking",
1257
+ description: "Per-model Preserve Thinking / Clear Thinking (full-history reasoning). Preserve Thinking keeps all turns' reasoning; Clear Thinking lets the template drop older reasoning (saves tokens, can hurt multi-turn recall / cause overthinking).",
1258
+ currentValue: "configure",
1259
+ submenu: (_currentValue: string, subDone: (v?: string) => void) => {
1260
+ // Re-read state on each open so toggles from a previous visit (which
1261
+ // wrote lilac.json + refreshed config) are reflected — a snapshot
1262
+ // captured at panel-open time would show stale values after a toggle.
1263
+ const fresh = collectPreserveState();
1264
+ const subItems = fresh.map((e) => ({
1265
+ id: `preserve:${e.id}`,
1266
+ label: truncateLabel(e.name),
1267
+ description: `${e.id} — Preserve Thinking keeps full reasoning history across turns; Clear Thinking lets the template drop older reasoning (saves tokens, can hurt multi-turn recall / cause overthinking).`,
1268
+ currentValue: e.preserved ? "Preserve Thinking" : "Clear Thinking",
1269
+ values: ["Preserve Thinking", "Clear Thinking"],
1270
+ }));
1271
+ const subList = new SettingsList(
1272
+ subItems,
1273
+ Math.min(subItems.length + 2, 15),
1274
+ getSettingsListTheme(),
1275
+ (id: string, newValue: string) => {
1276
+ const modelId = id.slice("preserve:".length);
1277
+ const entry = fresh.find((p) => p.id === modelId);
1278
+ if (!entry) return;
1279
+ const preservedOn = newValue === "Preserve Thinking";
1280
+ const flagValue = entry.flag === "clear_thinking" ? !preservedOn : preservedOn;
1281
+ updateConfig((cfg) => {
1282
+ const overrides = cfg.modelOverrides ?? (cfg.modelOverrides = {});
1283
+ const ov = overrides[modelId] ?? (overrides[modelId] = {});
1284
+ const compat = ov.compat ?? (ov.compat = {});
1285
+ const kwargs = compat.chatTemplateKwargs ?? (compat.chatTemplateKwargs = {});
1286
+ kwargs[entry.flag] = flagValue;
1287
+ return cfg;
1288
+ });
1289
+ listModelsCache = null;
1290
+ pi.registerProvider("lilac", {
1291
+ baseUrl: BASE_URL,
1292
+ apiKey: "$LILAC_API_KEY",
1293
+ api: "openai-completions",
1294
+ models: applyDiscounts(getListModels(), latestDiscounts),
1295
+ });
1296
+ syncStatus(ctx);
1297
+ ctx.ui.notify(`Preserved thinking ${preservedOn ? "on" : "off"} for ${entry.name} — takes effect now.`, "info");
1298
+ },
1299
+ () => subDone(),
1300
+ { enableSearch: true },
1301
+ );
1302
+ // The outer container's borders already frame the panel; return the
1303
+ // list directly so we don't render a second border pair.
1304
+ return subList;
1305
+ },
1306
+ },
1307
+ ];
1308
+
1309
+ const container = new Container();
1310
+ container.addChild(border());
1311
+
1312
+ const settingsList = new SettingsList(
1313
+ items,
1314
+ Math.min(items.length + 2, 15),
1315
+ getSettingsListTheme(),
1316
+ (id: string, newValue: string) => {
1317
+ if (id === "flex") {
1318
+ const value = newValue === "off" ? null : Number(newValue);
1319
+ applyFlexThreshold(value, ctx);
1320
+ }
1321
+ },
1322
+ () => done(undefined),
1323
+ { enableSearch: true },
1324
+ );
1325
+ container.addChild(settingsList);
1326
+ container.addChild(border());
1327
+
1328
+ return {
1329
+ render(width: number) {
1330
+ return container.render(width);
1331
+ },
1332
+ invalidate() {
1333
+ container.invalidate();
1334
+ },
1335
+ handleInput(data: string) {
1336
+ settingsList.handleInput?.(data);
1337
+ },
1338
+ };
1339
+ });
1231
1340
  },
1232
1341
  });
1233
1342
 
1234
1343
  pi.on("session_shutdown", () => {
1235
1344
  revalidateAbort?.abort();
1345
+ if (modelSelectNotifyTimer) { clearTimeout(modelSelectNotifyTimer); modelSelectNotifyTimer = null; }
1236
1346
  if (pollInterval) {
1237
1347
  clearInterval(pollInterval);
1238
1348
  pollInterval = null;
@@ -1240,5 +1350,5 @@ export default function (pi: ExtensionAPI) {
1240
1350
  });
1241
1351
  }
1242
1352
 
1243
- export { fetchStatusDiscounts, applyDiscounts, applyDiscountInPlace, loadCachedDiscounts, cacheDiscounts, buildModels, applyModelOverride, parseModelOverrides, parseFlexThreshold, updateConfig, loadConfig, getConfig };
1353
+ export { fetchStatusDiscounts, applyDiscounts, applyDiscountInPlace, loadCachedDiscounts, cacheDiscounts, buildModels, applyModelOverride, parseModelOverrides, parseFlexThreshold, updateConfig, loadConfig, getConfig, applyFlexThreshold };
1244
1354
  export type { JsonDiscount, JsonModel, PatchEntry, PatchData, ModelOverride, LilacConfig };
package/models.json CHANGED
@@ -23,29 +23,6 @@
23
23
  "zaiToolStream": true
24
24
  }
25
25
  },
26
- {
27
- "id": "zai-org/glm-5.1",
28
- "name": "GLM 5.1",
29
- "reasoning": true,
30
- "input": [
31
- "text"
32
- ],
33
- "cost": {
34
- "input": 0.9,
35
- "output": 3,
36
- "cacheRead": 0.27,
37
- "cacheWrite": 0
38
- },
39
- "contextWindow": 202752,
40
- "maxTokens": 131072,
41
- "compat": {
42
- "supportsDeveloperRole": false,
43
- "supportsStore": false,
44
- "maxTokensField": "max_completion_tokens",
45
- "thinkingFormat": "qwen-chat-template",
46
- "zaiToolStream": true
47
- }
48
- },
49
26
  {
50
27
  "id": "zai-org/glm-5.2",
51
28
  "name": "GLM 5.2",
@@ -92,28 +69,6 @@
92
69
  "zaiToolStream": true
93
70
  }
94
71
  },
95
- {
96
- "id": "minimaxai/minimax-m2.7",
97
- "name": "MiniMax M2.7",
98
- "reasoning": true,
99
- "input": [
100
- "text"
101
- ],
102
- "cost": {
103
- "input": 0.3,
104
- "output": 1.2,
105
- "cacheRead": 0.055,
106
- "cacheWrite": 0
107
- },
108
- "contextWindow": 204800,
109
- "maxTokens": 204800,
110
- "compat": {
111
- "supportsDeveloperRole": true,
112
- "supportsStore": false,
113
- "maxTokensField": "max_completion_tokens",
114
- "thinkingFormat": "qwen-chat-template"
115
- }
116
- },
117
72
  {
118
73
  "id": "minimaxai/minimax-m3",
119
74
  "name": "MiniMax M3",
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "pi-lilac-provider",
3
- "version": "1.7.2",
3
+ "version": "1.8.1",
4
4
  "description": "Lilac provider extension for pi - Access Kimi K2.6, GLM 5.1, and Gemma 4 models through Lilac's OpenAI-compatible API on idle GPUs",
5
5
  "type": "module",
6
6
  "main": "index.ts",
package/patch.json CHANGED
@@ -53,7 +53,7 @@
53
53
  "low": "high",
54
54
  "medium": "high",
55
55
  "high": "high",
56
- "xhigh": "max"
56
+ "max": "max"
57
57
  }
58
58
  },
59
59
  "google/gemma-4-31b-it": {
@@ -1,6 +1,6 @@
1
1
  #!/usr/bin/env node
2
2
  /**
3
- * Tests for the lilac-flex feature: discount-threshold gating.
3
+ * Tests for the Lilac flex feature: discount-threshold gating.
4
4
  *
5
5
  * Verifies, against the REAL exported helpers and registered handlers/commands
6
6
  * from index.ts (not a re-implementation):
@@ -14,9 +14,9 @@
14
14
  * discount < threshold; allows (continue) when >= threshold, when flex is
15
15
  * off, for non-lilac models, for non-interactive sources, and for models
16
16
  * with no discount entry (treated as 0% -> blocked when flex on).
17
- * - /lilac-flex command (integration via the registered handler): direct arg
18
- * ("/lilac-flex 75", "off", "50%"), picker presets, custom input, cancel,
19
- * invalid arg, and non-interactive guard.
17
+ * - applyFlexThreshold (flex configuration, now driven by the /lilac-settings
18
+ * flex row): sets flexThreshold, persists to disk, reports state against the
19
+ * live model's discount, and is visible to the input gate.
20
20
  *
21
21
  * Config FS + discount cache are isolated to a temp HOME so nothing touches the
22
22
  * real ~/.pi.
@@ -29,7 +29,7 @@ import { getAgentDir, type ExtensionAPI } from "@earendil-works/pi-coding-agent"
29
29
  // Isolate config + cache to a temp agent dir so loadConfig/cacheDiscounts never touch
30
30
  // the real ~/.pi. Must be set before importing index.ts, which computes
31
31
  // CONFIG_PATH / CACHE_PATH at module scope.
32
- const tmpHome = `/tmp/pi-lilac-flex-test-${Date.now()}`;
32
+ const tmpHome = `/tmp/pi-lilac-test-${Date.now()}`;
33
33
  fs.mkdirSync(tmpHome, { recursive: true });
34
34
  process.env.HOME = tmpHome;
35
35
  process.env.PI_CODING_AGENT_DIR = path.join(tmpHome, ".pi", "agent");
@@ -41,6 +41,7 @@ const {
41
41
  getConfig,
42
42
  updateConfig,
43
43
  cacheDiscounts,
44
+ applyFlexThreshold,
44
45
  } = await import("../index.ts");
45
46
 
46
47
  let passed = 0;
@@ -167,7 +168,6 @@ cacheDiscounts(new Map([
167
168
  ]));
168
169
 
169
170
  const handlers = new Map<string, ((...args: any[]) => any)[]>();
170
- const commands = new Map<string, { handler: (...args: any[]) => any }>();
171
171
 
172
172
  const mockApi: ExtensionAPI = {
173
173
  registerProvider: () => {},
@@ -175,9 +175,7 @@ const mockApi: ExtensionAPI = {
175
175
  if (!handlers.has(event)) handlers.set(event, []);
176
176
  handlers.get(event)!.push(handler);
177
177
  },
178
- registerCommand: (name: string, opts: any) => {
179
- commands.set(name, opts);
180
- },
178
+ registerCommand: () => {},
181
179
  appendEntry: () => {},
182
180
  exec: async () => ({ exitCode: 0, stdout: "", stderr: "" }),
183
181
  } as any;
@@ -266,102 +264,59 @@ updateConfig((c) => ({ ...c, flexThreshold: 25 }));
266
264
  eq(result, { action: "continue" }, "25% >= 25% threshold (inclusive) -> continue");
267
265
  }
268
266
 
269
- // ─── /lilac-flex command ──────────────────────────────────────────────────────
270
-
271
- console.log("\n--- /lilac-flex command ---");
267
+ // ─── applyFlexThreshold (flex configuration, driven by the /lilac-settings row) ─
272
268
 
273
- assert(commands.has("lilac-flex"), "/lilac-flex command registered");
274
- const cmd = commands.get("lilac-flex")!;
269
+ console.log("\n--- applyFlexThreshold ---");
275
270
 
276
- function runCmd(args: string, opts: { select?: string; input?: string; hasUI?: boolean; model?: any } = {}) {
271
+ function runApply(value: number | null, model: any = { id: KIMI, provider: "lilac" }) {
277
272
  const notifications: { msg: string; level: string }[] = [];
278
273
  const ctx = {
279
- hasUI: opts.hasUI ?? true,
280
- model: opts.model ?? { id: KIMI, provider: "lilac" },
274
+ model,
281
275
  ui: {
282
276
  notify: (msg: string, level: string) => notifications.push({ msg, level }),
283
277
  setStatus: () => {},
284
278
  theme: { fg: (_c: string, t: string) => t },
285
- select: async (_title: string, _labels: string[]) => opts.select,
286
- input: async (_title: string, _placeholder: string) => opts.input,
287
279
  },
288
280
  };
289
- return cmd.handler(args, ctx).then(() => ({ notifications, flexThreshold: getConfig().flexThreshold ?? null }));
290
- }
291
-
292
- // direct numeric arg
293
- {
294
- const { flexThreshold } = await runCmd("75");
295
- assert(flexThreshold === 75, "/lilac-flex 75 -> flexThreshold 75");
296
- assert(JSON.parse(fs.readFileSync(cfgPath, "utf8")).flexThreshold === 75, "/lilac-flex 75 persisted to disk");
297
- }
298
-
299
- // trailing % accepted
300
- {
301
- const { flexThreshold } = await runCmd("50%");
302
- assert(flexThreshold === 50, "/lilac-flex 50% -> flexThreshold 50");
303
- }
304
-
305
- // off keyword
306
- {
307
- const { flexThreshold } = await runCmd("off");
308
- assert(flexThreshold === null, "/lilac-flex off -> flexThreshold null");
309
- }
310
-
311
- // invalid arg -> no change, usage notify
312
- {
313
- const { flexThreshold, notifications } = await runCmd("bogus");
314
- assert(flexThreshold === null, "/lilac-flex bogus -> no change (still off)");
315
- assert(notifications.some((n) => n.msg.includes("Usage")), "/lilac-flex bogus -> usage notify");
316
- }
317
-
318
- // picker: "≥ 75% discount" preset
319
- {
320
- const { flexThreshold } = await runCmd("", { select: "≥ 75% discount" });
321
- assert(flexThreshold === 75, "picker '≥ 75% discount' -> 75");
322
- }
323
-
324
- // picker: "Off" preset
325
- {
326
- const { flexThreshold } = await runCmd("", { select: "Off — allow all discounts" });
327
- assert(flexThreshold === null, "picker 'Off' -> null");
281
+ applyFlexThreshold(value, ctx);
282
+ return { notifications, flexThreshold: getConfig().flexThreshold ?? null };
328
283
  }
329
284
 
330
- // picker: "Custom…" then input "60"
285
+ // set 75 (kimi seeded at 25% -> below threshold -> blocked-warning message)
331
286
  {
332
- const { flexThreshold } = await runCmd("", { select: "Custom…", input: "60" });
333
- assert(flexThreshold === 60, "picker Custom + input 60 -> 60");
287
+ const { flexThreshold, notifications } = runApply(75);
288
+ assert(flexThreshold === 75, "applyFlexThreshold(75) -> flexThreshold 75");
289
+ assert(JSON.parse(fs.readFileSync(cfgPath, "utf8")).flexThreshold === 75, "persisted to disk");
290
+ assert(notifications.some((n) => n.msg.includes("75%") && n.level === "warning"), "75 with kimi@25% -> warning, mentions threshold");
334
291
  }
335
292
 
336
- // picker: custom input invalid -> no change, warning
293
+ // off
337
294
  {
338
- const { flexThreshold, notifications } = await runCmd("", { select: "Custom…", input: "abc" });
339
- assert(flexThreshold === 60, "invalid custom input -> no change (still 60)");
340
- assert(notifications.some((n) => n.level === "warning" && n.msg.includes("Invalid")), "invalid custom input -> warning notify");
295
+ const { flexThreshold, notifications } = runApply(null);
296
+ assert(flexThreshold === null, "applyFlexThreshold(null) -> null");
297
+ assert(notifications.some((n) => n.msg.includes("off")), "off -> notify mentions off");
341
298
  }
342
299
 
343
- // picker cancelled (select returns undefined) -> no change, no notify
300
+ // threshold below current discount -> allowed (info)
344
301
  {
345
- const { flexThreshold, notifications } = await runCmd("", { select: undefined });
346
- assert(flexThreshold === 60, "picker cancelled -> no change (still 60)");
347
- assert(notifications.length === 0, "picker cancelled -> no notify");
302
+ const { notifications } = runApply(20);
303
+ assert(notifications.some((n) => n.level === "info" && n.msg.includes("allowed")), "20 with kimi@25% -> info, allowed");
348
304
  }
349
305
 
350
- // non-interactive guard -> error notify, no change
306
+ // no discount data for model -> warning
351
307
  {
352
- const { flexThreshold, notifications } = await runCmd("75", { hasUI: false });
353
- assert(flexThreshold === 60, "non-interactive -> no change (still 60)");
354
- assert(notifications.some((n) => n.level === "error" && n.msg.includes("interactive")), "non-interactive -> error notify");
308
+ const { notifications } = runApply(75, { id: "zai-org/glm-5.1", provider: "lilac" });
309
+ assert(notifications.some((n) => n.level === "warning" && n.msg.includes("no discount data")), "uncached model + threshold -> warning, no discount data");
355
310
  }
356
311
 
357
- // setting flex via command is visible to the gate (config cache in sync)
312
+ // setting flex via applyFlexThreshold is visible to the gate (config cache in sync)
358
313
  {
359
- await runCmd("75");
314
+ runApply(75);
360
315
  const { result } = await runInput("interactive", { id: KIMI, provider: "lilac" });
361
- eq(result, { action: "handled" }, "after /lilac-flex 75, gate blocks kimi@25%");
362
- await runCmd("off");
316
+ eq(result, { action: "handled" }, "after applyFlexThreshold(75), gate blocks kimi@25%");
317
+ runApply(null);
363
318
  const { result: afterOff } = await runInput("interactive", { id: KIMI, provider: "lilac" });
364
- eq(afterOff, { action: "continue" }, "after /lilac-flex off, gate allows kimi@25%");
319
+ eq(afterOff, { action: "continue" }, "after applyFlexThreshold(null), gate allows kimi@25%");
365
320
  }
366
321
 
367
322
  // ─── Summary ───────────────────────────────────────────────────────────────────
@@ -66,7 +66,6 @@ function eq<T>(actual: T, expected: T, message: string) {
66
66
 
67
67
  const KIMI = "moonshotai/kimi-k2.6";
68
68
  const GLM52 = "zai-org/glm-5.2";
69
- const GLM51 = "zai-org/glm-5.1";
70
69
 
71
70
  // ─── applyModelOverride ────────────────────────────────────────────────────────
72
71
 
@@ -240,16 +239,6 @@ function find(models: any[], id: string): any {
240
239
  assert(glm.compat.chatTemplateKwargs.clear_thinking === false, "non-overridden clear_thinking survives a thinkingLevelMap-only override");
241
240
  }
242
241
 
243
- {
244
- // Override on glm-5.1 toggles clear_thinking (patch sets false); other compat survives
245
- const overrides = { [GLM51]: { compat: { chatTemplateKwargs: { clear_thinking: true } } } } as any;
246
- const models = buildModels(modelsData, customModelsData, patchData, overrides);
247
- const glm = find(models, GLM51);
248
- assert(glm.compat.chatTemplateKwargs.clear_thinking === true, "override wins over patch: glm-5.1 clear_thinking -> true");
249
- assert((glm.compat.chatTemplateKwargs as any).thinking?.$var === "thinking.enabled", "override deep-merges: glm-5.1 thinking $var key survives");
250
- assert(glm.compat.zaiToolStream === true, "override deep-merges: glm-5.1 zaiToolStream survives");
251
- }
252
-
253
242
  {
254
243
  // Override for an unknown id is a no-op (adds no models)
255
244
  const before = buildModels(modelsData, customModelsData, patchData, {});
@@ -194,25 +194,15 @@ console.log("\n=== GLM 5.2 on the wire (real pi-ai) ===");
194
194
  const high = await wire("zai-org/glm-5.2", "high");
195
195
  eq(high.chat_template_kwargs, { enable_thinking: true, reasoning_effort: "high", clear_thinking: false },
196
196
  "glm-5.2 @ high → enable_thinking+reasoning_effort AND clear_thinking false");
197
- const xhigh = await wire("zai-org/glm-5.2", "xhigh");
198
- eq(xhigh.chat_template_kwargs, { enable_thinking: true, reasoning_effort: "max", clear_thinking: false },
199
- "glm-5.2 @ xhigh → reasoning_effort max, clear_thinking false");
197
+ const max = await wire("zai-org/glm-5.2", "max");
198
+ eq(max.chat_template_kwargs, { enable_thinking: true, reasoning_effort: "max", clear_thinking: false },
199
+ "glm-5.2 @ max → reasoning_effort max, clear_thinking false");
200
200
  const off = await wire("zai-org/glm-5.2", "off");
201
201
  eq(off.chat_template_kwargs, { enable_thinking: false, clear_thinking: false },
202
202
  "glm-5.2 @ off → reasoning_effort omitted (omitWhenOff), clear_thinking false persists");
203
203
  }
204
204
 
205
- console.log("\n=== GLM 5.1 on the wire (real pi-ai) ===");
206
- {
207
- const high = await wire("zai-org/glm-5.1", "high");
208
- eq(high.chat_template_kwargs, { thinking: true, enable_thinking: true, clear_thinking: false },
209
- "glm-5.1 @ high → thinking+enable_thinking true AND clear_thinking false");
210
- const off = await wire("zai-org/glm-5.1", "off");
211
- eq(off.chat_template_kwargs, { thinking: false, enable_thinking: false, clear_thinking: false },
212
- "glm-5.1 @ off → thinking false but clear_thinking false persists");
213
- }
214
-
215
- console.log("\n=== Gemma 4 / MiniMax (no family-wide preserve flag — regression guard) ===");
205
+ console.log("\n=== Gemma / MiniMax (no family-wide preserve flag — regression guard) ===");
216
206
  {
217
207
  const gemma = await wire("google/gemma-4-31b-it", "high");
218
208
  eq(gemma.chat_template_kwargs, { thinking: true, enable_thinking: true },
@@ -224,11 +214,6 @@ console.log("\n=== Gemma 4 / MiniMax (no family-wide preserve flag — regressio
224
214
  eq(m3.chat_template_kwargs, { thinking_mode: "enabled" },
225
215
  "minimax-m3 @ high → only thinking_mode (no preserve/clear flag)");
226
216
  falsy(m3.chat_template_kwargs?.preserve_thinking, "minimax-m3 has no preserve_thinking");
227
-
228
- const m27 = await wire("minimaxai/minimax-m2.7", "high");
229
- eq(m27.chat_template_kwargs, { thinking: true, enable_thinking: true },
230
- "minimax-m2.7 @ high → only thinking/enable_thinking (no preserve/clear flag)");
231
- falsy(m27.chat_template_kwargs?.clear_thinking, "minimax-m2.7 has no clear_thinking");
232
217
  }
233
218
 
234
219
  globalThis.fetch = originalFetch;