pi-magi-theme 0.2.4 → 0.3.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -36,14 +36,18 @@ Clone the repo and point pi at it instead (edits in the repo are live on the nex
36
36
 
37
37
  | Command | What it does |
38
38
  |---------|--------------|
39
- | `/magi <question>` | the council answers a question (recent conversation as context) |
40
- | `/magi review [focus]` | the council reviews your pending changes (`git diff HEAD` plus untracked file names) before you commit |
39
+ | `/magi council <question>` | the council answers a question (recent conversation as context) |
40
+ | `/magi council review [focus]` | the council reviews your pending changes (`git diff HEAD` plus untracked file names) before you commit |
41
41
  | `/magi config` | pick a model for each MAGI |
42
42
  | `/magi mecha` | MECHA SELECT: pick the llama-swap model to activate, each shown as a mecha head lit by its real state |
43
- | `/magi-ui compact` | toggle the compact side panel (basic info and animations only); remembered across sessions |
44
- | `/magi-ui status` | llama-swap report from its last 100 requests: speed, tokens, cache hits, MTP draft acceptance, durations, errors per model |
45
- | `/magi-ui config` | set the electricity price per kWh and the currency (EUR or USD) for the COST row |
46
- | `/magi-ui panel` · `on` · `off` | hide/show the side panel, enable/disable the whole chrome |
43
+ | `/magi compact` | toggle the compact side panel (basic info and animations only); remembered across sessions |
44
+ | `/magi status` | llama-swap report from its last 100 requests: speed, tokens, cache hits, MTP draft acceptance, durations, errors per model |
45
+ | `/magi cost` | set the electricity price per kWh and the currency (EUR or USD) for the COST row |
46
+ | `/magi panel` · `on` · `off` | hide/show the side panel, enable/disable the whole chrome |
47
+ | `/magi hygiene` · `on` · `off` · `step <tokens>` · `<turns> <results>` | show how much context the hygiene pruned; enable/disable it; how many prunable tokens make a pruning step (e.g. `step 40k`, default 15k); how many recent turns keep their thinking and how many tool results stay whole (e.g. `3 5`) |
48
+ | `/magi budget` · `auto` · `off` · `reset` · `message` · `<planning> <acting>` | show the thinking budget of the current model; learn it per model (default); leave it to llama-server; forget what was learned for this model; turn the closing message off/on; or fix it (e.g. `16k 4k`) |
49
+
50
+ Everything lives under `/magi`: type `/magi ` (with the space) to see every option with a short description; keep typing to narrow it down, Tab or Enter to pick one. Anything that is not an option is a question for the council.
47
51
 
48
52
  ## Lore ↔ function
49
53
 
@@ -84,9 +88,39 @@ pi install npm:pi-smart-compact
84
88
 
85
89
  It extracts files, errors, decisions and open loops locally (no LLM calls), then synthesizes and verifies the summary. Point its `summaryModel` at a local model to keep compaction free. When it is installed, the sixth seal suggests `/smart-compact`, the seals name it while they break (`✶ BREAKING THE SEALS · smart-compact · 4s`) and the seventh seal reports who actually produced the summary: `smart-compact`, or `pi native` if it fell back to pi's own compactor.
86
90
 
91
+ ## Local models: context hygiene, thinking budget, loop guard, MAGI.md
92
+
93
+ Local models run out of context on long tasks well before they run out of work. They also tend to think for minutes between two tool calls and to repeat the same command when stuck. MAGI works on all three, automatically (the thinking budget for llama-swap models, the rest for any model):
94
+
95
+ - **Context hygiene: the model's memory stays lean.** *Problem:* every file the agent reads and every long reasoning stays in the conversation, until the model's context is full and the task falls apart. *What MAGI does:* before each request it replaces old reasoning, old tool outputs and old file writes with a one-line note (`<<pruned by MAGI…>>`). Only what the model sees is trimmed: your saved session stays complete. The latest 3 turns keep their reasoning and the latest 5 tool outputs stay whole. On two real sessions it brought 104k and 119k tokens down to ~64k and ~47k.
96
+ - **Thinking budget: no more ten-minute thinks.** *Problem:* a local model can reason for thousands of tokens before a simple step, and at 10 tokens/s that is minutes of waiting. *What MAGI does:* it gives the model a maximum length of thinking on every request: larger right after you write (planning), smaller between tool calls (acting). When the limit is reached the model is stopped mid-thought, says *"Time is up. I will take the smallest safe next step…"* and acts. The limit adapts to each model on its own. The side panel counts these cuts (`HYGIENE -18.2k ✂2`).
97
+ - **Loop guard: no endless retries.** *Problem:* a stuck model runs the same command again and again. *What MAGI does:* the third identical tool call in a row is blocked, with a message asking the model to try something else.
98
+ - **MAGI.md: house rules for the model.** *Problem:* local models repeat the same mistakes: reading whole files, inventing paths, claiming success without checking. *What MAGI does:* it creates `MAGI.md` in your project on the first start (never overwritten) and adds it to the model's instructions on every run: short rules against these mistakes, plus the habit of keeping the task plan in `PLAN.md` and findings in `NOTES.md`, so nothing important is lost when old context is trimmed. Edit it per project; delete it to get the defaults back.
99
+
100
+ Nothing needs setting up in llama-swap. The defaults suit most tasks; two adjustments are worth knowing:
101
+
102
+ - **Long tasks: `/magi hygiene step 40k`.** Each time the hygiene trims, the server has to re-read part of the conversation, and on some models (see *Hybrid models* below) almost all of it, which can take a few minutes. A step of 40k trims less often: on a real session it cut the re-reading from ~6 minutes to ~1.
103
+ - **A model that really needs to think longer: `/magi budget <planning> <acting>`**, e.g. `/magi budget 16k 8k`. Frequent ✂ cuts in the side panel are the sign.
104
+
105
+ ### How it works
106
+
107
+ For the curious, and for tuning.
108
+
109
+ **Hygiene.** Trimming happens in steps, not on every request: the conversation up to a mark is trimmed, and the mark only moves forward once ~15k more tokens (the step) could be trimmed. Between steps the conversation only grows at the end, so llama.cpp reuses what it already processed (its KV cache) and only reads the new messages. Each step changes the conversation from the first newly trimmed message on, and the server re-reads from there: that is why steps are large and rare. Old tool outputs are replaced, not summarized: [simple observation masking matches LLM summarization at half the cost](https://arxiv.org/abs/2508.21433).
110
+
111
+ **Hybrid models.** Some models (e.g. Qwen3.8 Flash Next; dense or MoE does not matter) mix a few classic attention layers with recurrent ones, which squeeze the whole conversation into a fixed-size state instead of keeping each token. The server cannot rewind that state to an arbitrary point: it can only restore a saved copy (a checkpoint) and re-read from there, and the only copy before the trimmed part is usually the end of the system prompt. So on these models each step re-reads almost the whole prompt. With 10k-token thoughts and 4k-token file reads, a 15k step is crossed every 2–3 turns: on a real 70k-token session that was 3 re-reads in 6 requests (~6 min at 185 tokens/s); with a 40k step, 1 re-read (~1 min) for a prompt at most 6k larger. A model is hybrid if its llama-server log shows `restored context checkpoint` lines.
112
+
113
+ **Thinking budget.** MAGI sends the limit with every request (`thinking_budget_tokens`), and llama.cpp applies it. It learns per model and per phase: 1.5 × the 95th percentile of the model's last 30 thinking lengths, rounded up to 1k. A cut is recorded as 1/1.5 of the limit it hit, so cuts never raise the limit: a model that often runs away is held, not chased. Replies that end close under the limit raise it; a model that thinks little lowers it. Limits: planning 4k–32k, acting 2k–4k (past ~4k between two tool calls it is overthinking). Until a model has 10 replies in a phase it uses 16k / 4k.
114
+
115
+ When the limit is reached llama.cpp does not abort the reply: it inserts the closing sentence and the end-of-thinking tag, and the model goes on to act. MAGI sends that sentence with every request too (`reasoning_budget_message`; `/magi budget message` turns it off and on). `MAGI.md` tells the model what to do after a cut (one small verifiable step, the open plan into `PLAN.md`), and the next turn gets a fresh budget.
116
+
117
+ **llama-server versions.** The per-request limit works whenever llama-server was started without `--reasoning-budget`, which is the default. If it was started with one, the server's limit wins: MAGI notices the model thinking well past its own limit and `/magi budget` says so. A llama-server too old for the closing sentence ends the thinking silently, and MAGI detects the cut by its length.
118
+
119
+ **Loop guard details.** A tool call counts as identical when both the tool and its arguments match. A file write or edit that copies a `<<pruned by MAGI…>>` note into a file is blocked too.
120
+
87
121
  ## The council
88
122
 
89
- `/magi <question>` asks three models in parallel, each with its own nature, then shows the votes and a majority verdict:
123
+ `/magi council <question>` asks three models in parallel, each with its own nature, then shows the votes and a majority verdict:
90
124
 
91
125
  | Unit | Nature | Looks at |
92
126
  |------|--------|----------|
@@ -96,18 +130,20 @@ It extracts files, errors, decisions and open loops locally (no LLM calls), then
96
130
 
97
131
  Each nature is a lens, not a specialty, so the council answers any question, not only software ones. Every MAGI first answers the question, then judges it through its lens, naming concrete tools, numbers and scenarios from your question instead of generic advice. Votes: **APPROVE** = go ahead or clear recommendation; **CONDITIONAL** = only if the named conditions hold, or when information is missing (it says what it needs); **REJECT** = a concrete problem, with what to do instead. A MAGI never rejects because a topic is outside its nature. Answers come back in the language of your question.
98
132
 
99
- `/magi <question>` gives the MAGI the recent conversation as context; `/magi review` gives them the pending diff (truncated at 24k characters). Full opinions are added to the chat (not sent to the agent), and the last verdict stays under the MAGI in the side panel.
133
+ `/magi council <question>` gives the MAGI the recent conversation as context; `/magi council review` gives them the pending diff (truncated at 24k characters). Full opinions are added to the chat (not sent to the agent), and the last verdict stays under the MAGI in the side panel.
100
134
 
101
135
  ## Configuration
102
136
 
103
- `~/.pi/agent/magi.json` (written by `/magi config`, `/magi-ui compact` and `/magi-ui config`, editable by hand):
137
+ `~/.pi/agent/magi.json` (written by `/magi config`, `/magi compact` and `/magi cost`, editable by hand):
104
138
 
105
139
  ```json
106
140
  {
107
141
  "MELCHIOR": { "model": "llama-swap/Qwen3.8 27B Q4_K_M - Thinking", "thinking": "low" },
108
142
  "ui": { "compact": false, "kwhPrice": 0.30, "currency": "EUR" },
109
143
  "loads": { "qwen3.8-27b": 41200 },
110
- "totalWh": 1843.2
144
+ "totalWh": 1843.2,
145
+ "hygiene": { "enabled": true, "keepThinkingTurns": 3, "keepToolResults": 5, "stepTokens": 15000, "minPruneChars": 600 },
146
+ "thinkingBudget": { "mode": "auto", "planning": 16384, "acting": 4096, "message": true, "learned": { "qwen3.8-27b": { "acting": [812, 430, 2211] } } }
111
147
  }
112
148
  ```
113
149
 
@@ -115,6 +151,7 @@ Each nature is a lens, not a specialty, so the council answers any question, not
115
151
  - `ui.compact`: start with the compact side panel;
116
152
  - `ui.kwhPrice` and `ui.currency` (`EUR` or `USD`): the COST row multiplies the GPU energy by this price, showing the running total of every session with the current one in brackets;
117
153
  - `totalWh`: written by the theme, GPU energy summed over every session (delete the key to reset the COST total);
154
+ - `hygiene` and `thinkingBudget` are set with `/magi hygiene` and `/magi budget` (`hygiene.minPruneChars` by hand only); `thinkingBudget.message` sends the closing message with every request (default `true`); `thinkingBudget.learned` is written by the theme (recent thinking lengths per model and phase);
118
155
  - `loads`: written by the theme, how long each llama-swap model took to load last time (paces the angel attack; 60s when unknown).
119
156
 
120
157
  ## Release
@@ -129,6 +166,8 @@ No token: npmjs is configured to trust this repository's `publish.yml` (npm trus
129
166
 
130
167
  ## llama-swap
131
168
 
169
+ **Thinking levels.** [pi-llama-swap](https://www.npmjs.com/package/@danielmeneses/pi-llama-swap) registers every model with reasoning off, so `/thinking` only offers `off`. At session start MAGI reads the reasoning levels llama-swap publishes in `/v1/models` (`meta.llamaswap.reasoning.levels`) and re-registers those models with thinking on: `/thinking` then offers exactly those levels, sent as `chat_template_kwargs` (`enable_thinking`, `reasoning_effort`). Aliases (e.g. the Instruct twin of a Thinking model) stay off, and image input follows `architecture.input_modalities`. A new or renamed model works without a `modelOverrides` entry in `models.json`; entries you keep there still apply on top (e.g. `samplingParams`). The starting level is pi's usual one for a model switch: the level saved for that model (`Ctrl+S` in `/thinking`), else `defaultThinkingLevel`.
170
+
132
171
  When the session model uses the `llama-swap` provider, the side panel:
133
172
 
134
173
  - loads nothing at startup: a new session opens MECHA SELECT, a resumed one shows whether its model is already in VRAM;
@@ -17,22 +17,41 @@
17
17
  * - while a model loads an angel attacks the MAGI: red spreads through BALTHASAR, MELCHIOR and CASPAR at the pace of
18
18
  * the model's last load, a corner of CASPAR holds out blinking; once loaded, blue takes the MAGI back from that corner
19
19
  * - llama-swap telemetry: VRAM, GPU load/temp/power, energy used, RAM, server-side tok/s, prompt tok/s, cache hits
20
- * - /magi config → assign a model to each MAGI; /magi-ui compact|status → smaller panel, llama-swap report
20
+ * - /magi config → assign a model to each MAGI; /magi compact|status → smaller panel, llama-swap report
21
21
  *
22
22
  * Fan art: the MAGI and their screen come from Neon Genesis Evangelion, all rights reserved to khara, Inc.
23
23
  * Use with the theme ../../themes/magi.json
24
24
  */
25
25
 
26
26
  import { execFile } from "node:child_process";
27
- import { readFileSync, writeFileSync } from "node:fs";
27
+ import { existsSync, readFileSync, writeFileSync } from "node:fs";
28
28
  import { homedir } from "node:os";
29
29
  import { join } from "node:path";
30
30
  import { promisify } from "node:util";
31
31
  import type { AssistantMessage, Model } from "@earendil-works/pi-ai";
32
32
  import { completeSimple } from "@earendil-works/pi-ai";
33
- import type { ExtensionAPI, ExtensionContext, Theme, ThemeColor } from "@earendil-works/pi-coding-agent";
33
+ import type { ExtensionAPI, ExtensionCommandContext, ExtensionContext, Theme, ThemeColor } from "@earendil-works/pi-coding-agent";
34
34
  import type { Component, OverlayHandle, TUI } from "@earendil-works/pi-tui";
35
35
  import { HStack, matchesKey, sliceByColumn, truncateToWidth, visibleWidth, wrapTextWithAnsi } from "@earendil-works/pi-tui";
36
+ import {
37
+ BUDGET_DEFAULTS,
38
+ BUDGET_MESSAGE,
39
+ BUDGET_WINDOW,
40
+ HYGIENE_DEFAULTS,
41
+ MAGI_MD,
42
+ PRUNED_MARK,
43
+ learnedBudget,
44
+ budgetSample,
45
+ pruneContext,
46
+ requestPhase,
47
+ thinkingTokens,
48
+ budgetVerdict,
49
+ thinkingWasCut,
50
+ type BudgetMode,
51
+ type BudgetPhase,
52
+ type HygieneOptions,
53
+ type HygieneStats,
54
+ } from "./local-models.ts";
36
55
 
37
56
  /* ────────────────────────────────────────────────────────────── art ── */
38
57
 
@@ -317,6 +336,8 @@ const state = {
317
336
  rebornAt: 0,
318
337
  hasSmartCompact: false,
319
338
  lastCouncil: undefined as { verdict: Vote | null; tally: number; question: string } | undefined,
339
+ pruned: 0, // tokens the context hygiene keeps away from the model
340
+ thinkCuts: 0, // replies whose thinking llama.cpp cut at the budget
320
341
  };
321
342
 
322
343
  /** Panel preferences, persisted under "ui" in ~/.pi/agent/magi.json. */
@@ -324,7 +345,7 @@ type Currency = "EUR" | "USD";
324
345
 
325
346
  const ui = {
326
347
  compact: false,
327
- kwhPrice: undefined as number | undefined, // price per kWh, for the COST row (/magi-ui config)
348
+ kwhPrice: undefined as number | undefined, // price per kWh, for the COST row (/magi cost)
328
349
  currency: "EUR" as Currency,
329
350
  };
330
351
 
@@ -421,7 +442,9 @@ interface TokenStats {
421
442
 
422
443
  const tokens: TokenStats = { input: 0, output: 0, cacheRead: 0, cost: 0 };
423
444
 
424
- function addUsage(m: AssistantMessage): void {
445
+ /** Session totals from one assistant reply: tokens, cost, and whether its thinking was cut at the budget. */
446
+ function countAssistant(m: AssistantMessage): void {
447
+ if (thinkingWasCut(m as any)) state.thinkCuts++;
425
448
  tokens.input += m.usage?.input ?? 0;
426
449
  tokens.output += m.usage?.output ?? 0;
427
450
  tokens.cacheRead += m.usage?.cacheRead ?? 0;
@@ -432,8 +455,9 @@ function addUsage(m: AssistantMessage): void {
432
455
  function recountSession(ctx: ExtensionContext): void {
433
456
  Object.assign(tokens, { input: 0, output: 0, cacheRead: 0, cost: 0 });
434
457
  state.lastCouncil = undefined;
458
+ state.thinkCuts = 0;
435
459
  for (const entry of ctx.sessionManager.getBranch()) {
436
- if (entry.type === "message" && entry.message.role === "assistant") addUsage(entry.message as AssistantMessage);
460
+ if (entry.type === "message" && entry.message.role === "assistant") countAssistant(entry.message as AssistantMessage);
437
461
  if (entry.type === "custom" && entry.customType === "magi-verdict") {
438
462
  const d = entry.data as { verdict?: Vote | null; tally?: number; question?: string } | undefined;
439
463
  if (d && Array.isArray((d as any).opinions)) state.lastCouncil = { verdict: d.verdict ?? null, tally: d.tally ?? 0, question: d.question ?? "" };
@@ -781,10 +805,11 @@ let lastPersistAt = 0;
781
805
  /**
782
806
  * Adds this session's new energy to the running total in magi.json, so the COST row survives restarts.
783
807
  * ponytail: writes at most once a minute, and adds a delta so parallel sessions do not overwrite each other.
808
+ * A crash therefore loses up to a minute of energy, which is what a kill loses anyway.
784
809
  */
785
- function persistEnergy(force = false): void {
810
+ function persistEnergy(): void {
786
811
  const delta = swap.energyWh - swap.savedWh;
787
- if (delta <= 0 || (!force && Date.now() - lastPersistAt < ENERGY_SAVE_MS)) return;
812
+ if (delta <= 0 || Date.now() - lastPersistAt < ENERGY_SAVE_MS) return;
788
813
  lastPersistAt = Date.now();
789
814
  const cfg = loadMagiConfig();
790
815
  cfg.totalWh = (cfg.totalWh ?? 0) + delta;
@@ -882,6 +907,56 @@ async function swapAliases(): Promise<Map<string, string>> {
882
907
  return aliases;
883
908
  }
884
909
 
910
+ const PI_THINKING_LEVELS = ["minimal", "low", "medium", "high", "xhigh", "max"];
911
+
912
+ /**
913
+ * pi-llama-swap registers every model with reasoning off, so /thinking only offers "off".
914
+ * llama-swap publishes each model's reasoning levels in /v1/models: re-register the provider with them,
915
+ * sent as chat_template_kwargs (enable_thinking + reasoning_effort). Aliases (the Instruct twins) stay off.
916
+ * models.json modelOverrides still apply on top.
917
+ */
918
+ async function enableSwapReasoning(pi: ExtensionAPI, ctx: ExtensionContext): Promise<void> {
919
+ const config = ctx.modelRegistry.getRegisteredProviderConfig("llama-swap") as any;
920
+ if (!config?.models?.length || !config.baseUrl) return;
921
+ try {
922
+ const headers: Record<string, string> = config.apiKey ? { Authorization: `Bearer ${config.apiKey}` } : {};
923
+ const res = await fetch(`${config.baseUrl.replace(/\/$/, "")}/models`, { headers, signal: AbortSignal.timeout(5000) });
924
+ const { data } = (await res.json()) as { data?: any[] };
925
+ const meta = new Map((data ?? []).map((m) => [m.id, m]));
926
+ let changed = false;
927
+ const models = config.models.map((m: any) => {
928
+ const entry = meta.get(m.id);
929
+ const levels: string[] | undefined = entry?.meta?.llamaswap?.reasoning?.levels;
930
+ const input = entry?.architecture?.input_modalities?.includes("image") ? ["text", "image"] : ["text"];
931
+ if (!levels?.length || entry.meta.llamaswap.type === "alias") return { ...m, input };
932
+ changed = true;
933
+ return {
934
+ ...m,
935
+ input,
936
+ reasoning: true,
937
+ thinkingLevelMap: Object.fromEntries(PI_THINKING_LEVELS.map((l) => [l, levels.includes(l) ? l : null])),
938
+ compat: {
939
+ ...m.compat,
940
+ thinkingFormat: "chat-template",
941
+ chatTemplateKwargs: {
942
+ enable_thinking: { $var: "thinking.enabled" },
943
+ preserve_thinking: true,
944
+ reasoning_effort: { $var: "thinking.effort", omitWhenOff: true },
945
+ },
946
+ },
947
+ };
948
+ });
949
+ // ponytail: a /llama-swap refresh re-registers the plain models until the next session start
950
+ if (!changed) return;
951
+ ctx.modelRegistry.registerProvider("llama-swap", { ...config, models });
952
+ // the session already holds the old model object: swap in the new one
953
+ const fresh = ctx.model?.provider === "llama-swap" ? ctx.modelRegistry.find("llama-swap", ctx.model.id) : undefined;
954
+ if (fresh?.reasoning) await pi.setModel(fresh);
955
+ } catch {
956
+ // server unreachable: models stay as pi-llama-swap registered them
957
+ }
958
+ }
959
+
885
960
  /** Models llama-swap keeps in memory: real id → "ready" | "starting" | … (empty when the server is unreachable). */
886
961
  async function swapRunning(): Promise<Map<string, string>> {
887
962
  try {
@@ -1030,7 +1105,7 @@ async function prewarmPrefix(cwd: string): Promise<void> {
1030
1105
  }
1031
1106
  }
1032
1107
 
1033
- /* ── /magi-ui status: a report built from the last requests llama-swap recorded ── */
1108
+ /* ── /magi status: a report built from the last requests llama-swap recorded ── */
1034
1109
 
1035
1110
  interface ActivityRow {
1036
1111
  timestamp: string;
@@ -1315,6 +1390,11 @@ class MagiPanel implements Component {
1315
1390
  if (!compact && usage?.contextWindow) {
1316
1391
  out.push(this.field("CONTEXT", `${usage.tokens == null ? "?" : fmtTokens(usage.tokens)} / ${fmtTokens(usage.contextWindow)}`, inner, "muted"));
1317
1392
  }
1393
+ // HYGIENE: tokens pruned from what the model sees, ✂ = thinking cut at the budget
1394
+ if (!compact && (state.pruned || state.thinkCuts)) {
1395
+ const cuts = state.thinkCuts ? ` ✂${state.thinkCuts}` : "";
1396
+ out.push(this.field("HYGIENE", `-${fmtTokens(state.pruned)}${cuts}`, inner, state.thinkCuts ? "warning" : "muted"));
1397
+ }
1318
1398
  return out;
1319
1399
  }
1320
1400
 
@@ -1346,14 +1426,12 @@ class MagiPanel implements Component {
1346
1426
  return out;
1347
1427
  }
1348
1428
 
1349
- /** COST: GPU energy × price per kWh set with /magi-ui config — all sessions, with this one in brackets. */
1429
+ /** COST: GPU energy × price per kWh set with /magi cost — all sessions, with this one in brackets. */
1350
1430
  private costRow(inner: number): string {
1351
1431
  if (!swap.gpus.length) return this.field("COST", "—", inner, "muted");
1352
- if (ui.kwhPrice === undefined) return this.field("COST", "→ /magi-ui config", inner, "dim");
1353
- const th = this.theme;
1432
+ if (ui.kwhPrice === undefined) return this.field("COST", "→ /magi cost", inner, "dim");
1354
1433
  const price = (wh: number) => fmtMoney((wh / 1000) * ui.kwhPrice!);
1355
- const label = th.fg("dim", "COST".padEnd(9));
1356
- return this.frameLine(" " + label + th.fg("warning", price(totalWh())) + th.fg("dim", ` (ses ${price(swap.energyWh)})`), inner);
1434
+ return this.field("COST", price(totalWh()) + this.theme.fg("dim", ` (ses ${price(swap.energyWh)})`), inner, "warning");
1357
1435
  }
1358
1436
 
1359
1437
  private swapRows(inner: number, compact: boolean): string[] {
@@ -1505,9 +1583,49 @@ type MagiConfig = Partial<Record<MagiUnit, MagiUnitConfig>> & {
1505
1583
  ui?: { compact?: boolean; kwhPrice?: number; currency?: Currency };
1506
1584
  loads?: Record<string, number>; // real model id → ms its last load took
1507
1585
  totalWh?: number; // GPU energy summed over every session, for the COST row
1586
+ hygiene?: Partial<HygieneOptions> & { enabled?: boolean }; // context pruning for local models, see local-models.ts
1587
+ thinkingBudget?: {
1588
+ mode?: BudgetMode; // auto (learned per model) · fixed · off, set with /magi budget
1589
+ planning?: number; // fixed budgets
1590
+ acting?: number;
1591
+ message?: boolean; // send BUDGET_MESSAGE as reasoning_budget_message (default on), /magi budget message
1592
+ learned?: Record<string, Partial<Record<BudgetPhase, number[]>>>; // written by the theme: recent thinking lengths per model
1593
+ };
1508
1594
  };
1509
1595
 
1510
1596
  const MAGI_CONFIG_PATH = join(homedir(), ".pi", "agent", "magi.json");
1597
+ /** /magi arguments that manage the theme instead of asking the council. */
1598
+ const UI_ARGS = /^(on|off|panel|compact|status|cost)$|^(hygiene|budget)(\s|$)/i;
1599
+ /** /magi arguments offered by autocomplete: the full argument, and what it does. */
1600
+ const MAGI_ARGS: [string, string][] = [
1601
+ ["council", "ask the three MAGI a question (recent conversation as context)"],
1602
+ ["council review", "the council reviews your pending changes before you commit"],
1603
+ ["config", "pick a model for each MAGI"],
1604
+ ["mecha", "MECHA SELECT: pick the llama-swap model to activate"],
1605
+ ["status", "llama-swap report: speed, tokens, cache hits, errors per model"],
1606
+ ["panel", "hide/show the side panel"],
1607
+ ["compact", "toggle the compact side panel"],
1608
+ ["cost", "electricity price and currency for the COST row"],
1609
+ ["on", "enable the MAGI chrome"],
1610
+ ["off", "disable the MAGI chrome"],
1611
+ ["hygiene", "show how much context was pruned"],
1612
+ ["hygiene on", "enable context pruning"],
1613
+ ["hygiene off", "disable context pruning"],
1614
+ ["hygiene step 40k", "prune less often: fewer prompt re-reads on long tasks (default 15k)"],
1615
+ ["hygiene 3 5", "recent turns that keep their thinking, tool results kept whole"],
1616
+ ["budget", "show the thinking budget of the current model"],
1617
+ ["budget auto", "learn the budget per model (default)"],
1618
+ ["budget off", "no budget: llama-server decides"],
1619
+ ["budget reset", "forget what was learned for the current model"],
1620
+ ["budget message", "turn the closing message after a cut off/on"],
1621
+ ["budget 16k 4k", "fixed budget: planning, acting"],
1622
+ ];
1623
+ function argCompletions(table: [string, string][], prefix: string) {
1624
+ const p = prefix.trimStart().toLowerCase();
1625
+ const items = table.filter(([value]) => value.startsWith(p)).map(([value, description]) => ({ value, label: value, description }));
1626
+ return items.length ? items : null;
1627
+ }
1628
+ const LOOP_REPEATS = 3; // the same tool call this many times in a row is blocked
1511
1629
 
1512
1630
  function loadMagiConfig(): MagiConfig {
1513
1631
  try {
@@ -1561,7 +1679,7 @@ function conversationExcerpt(ctx: ExtensionContext, maxChars = 6000): string {
1561
1679
 
1562
1680
  const REVIEW_MAX_CHARS = 24_000;
1563
1681
 
1564
- /** The pending changes for /magi review: tracked changes against HEAD plus the names of untracked files. */
1682
+ /** The pending changes for /magi council review: tracked changes against HEAD plus the names of untracked files. */
1565
1683
  async function pendingChanges(cwd: string): Promise<{ diff: string; untracked: string[] }> {
1566
1684
  const git = (args: string[]) => promisify(execFile)("git", args, { cwd, maxBuffer: 32 * 1024 * 1024 }).then((r) => r.stdout);
1567
1685
  let diff: string;
@@ -1723,7 +1841,7 @@ function buildDeliberationView(
1723
1841
  };
1724
1842
  }
1725
1843
 
1726
- /** A read-only boxed report (used by /magi-ui status); any key closes it. */
1844
+ /** A read-only boxed report (used by /magi status); any key closes it. */
1727
1845
  function buildReportView(theme: Theme, lines: string[], close: () => void) {
1728
1846
  return {
1729
1847
  render(width: number): string[] {
@@ -1984,7 +2102,7 @@ function buildFooter(tui: TUI, theme: Theme, footerData: any) {
1984
2102
  // change would shift it under the window. A fresh line is taken when it has the same
1985
2103
  // width, so nothing moves, otherwise at the end of the loop.
1986
2104
  const line = left + dim(FOOTER_SCROLL_GAP) + right + dim(FOOTER_SCROLL_GAP);
1987
- const period = Math.max(1, visibleWidth(line));
2105
+ const period = visibleWidth(line);
1988
2106
  if (!scrollLine || scrollOff === 0 || period === scrollPeriod) {
1989
2107
  scrollLine = line;
1990
2108
  scrollPeriod = period;
@@ -2021,6 +2139,22 @@ export default function (pi: ExtensionAPI) {
2021
2139
  let panelHandle: OverlayHandle | undefined;
2022
2140
  let panel: MagiPanel | undefined;
2023
2141
  let titleDone = false;
2142
+ let hygiene = { enabled: true, ...HYGIENE_DEFAULTS };
2143
+ let hygieneStats: HygieneStats | undefined;
2144
+ let budgetMode: BudgetMode = "auto";
2145
+ let fixedBudget = { ...BUDGET_DEFAULTS };
2146
+ let budgetMessage = true;
2147
+ let learned: Record<string, Partial<Record<BudgetPhase, number[]>>> = {};
2148
+ let pendingBudget: { model: string; phase: BudgetPhase; tokens: number } | undefined; // the request in flight
2149
+ const budgetIgnored = new Set<string>(); // models whose llama-server thought past the budget: it sets its own
2150
+ const budgetFor = (model: string, phase: BudgetPhase) =>
2151
+ budgetMode === "fixed" ? fixedBudget[phase] : (learnedBudget(learned[model]?.[phase] ?? [], phase) ?? BUDGET_DEFAULTS[phase]);
2152
+ const persistBudget = () => {
2153
+ const cfg = loadMagiConfig();
2154
+ saveMagiConfig({ ...cfg, thinkingBudget: { mode: budgetMode, ...fixedBudget, message: budgetMessage, learned } });
2155
+ };
2156
+ let lastCall = ""; // loop guard: fingerprint of the previous tool call and how often it repeated
2157
+ let repeats = 0;
2024
2158
 
2025
2159
  const repaint = () => {
2026
2160
  panel?.invalidate();
@@ -2181,6 +2315,24 @@ export default function (pi: ExtensionAPI) {
2181
2315
 
2182
2316
  pi.on("session_start", async (event, ctx) => {
2183
2317
  liveCtx = ctx;
2318
+ const magiCfg = loadMagiConfig();
2319
+ hygiene = { enabled: true, ...HYGIENE_DEFAULTS, ...magiCfg.hygiene };
2320
+ const tb = magiCfg.thinkingBudget ?? {};
2321
+ budgetMode = tb.mode ?? "auto";
2322
+ fixedBudget = { planning: tb.planning ?? BUDGET_DEFAULTS.planning, acting: tb.acting ?? BUDGET_DEFAULTS.acting };
2323
+ budgetMessage = tb.message ?? true;
2324
+ learned = tb.learned ?? {};
2325
+ await enableSwapReasoning(pi, ctx);
2326
+ // the local-model rules live next to AGENTS.md, created once so the user can edit them
2327
+ const rules = join(ctx.cwd, "MAGI.md");
2328
+ if (!existsSync(rules)) {
2329
+ try {
2330
+ writeFileSync(rules, MAGI_MD);
2331
+ if (ctx.mode === "tui") ctx.ui.notify("MAGI.md created: rules for local models, appended to the system prompt", "info");
2332
+ } catch {
2333
+ // read-only directory: the built-in rules are used as they are
2334
+ }
2335
+ }
2184
2336
  void refreshUnit(ctx);
2185
2337
  recountSession(ctx);
2186
2338
  state.hasSmartCompact = pi.getCommands().some((c) => c.name.replace(/^\//, "") === "smart-compact");
@@ -2233,6 +2385,12 @@ export default function (pi: ExtensionAPI) {
2233
2385
 
2234
2386
  pi.on("before_provider_request", (event, ctx) => {
2235
2387
  rememberPrefix(ctx.cwd, event.payload);
2388
+ // llama.cpp honours a per-request thinking budget only when llama-server runs without --reasoning-budget
2389
+ if (budgetMode === "off" || !swap.base || !ctx.model) return;
2390
+ const phase = requestPhase(event.payload);
2391
+ pendingBudget = { model: ctx.model.id, phase, tokens: budgetFor(ctx.model.id, phase) };
2392
+ const payload = { ...(event.payload as object), thinking_budget_tokens: pendingBudget.tokens };
2393
+ return budgetMessage ? { ...payload, reasoning_budget_message: `\n\n${BUDGET_MESSAGE}` } : payload;
2236
2394
  });
2237
2395
 
2238
2396
  pi.on("model_select", async (event, ctx) => {
@@ -2242,7 +2400,7 @@ export default function (pi: ExtensionAPI) {
2242
2400
  });
2243
2401
 
2244
2402
  pi.on("session_shutdown", async () => {
2245
- persistEnergy(true);
2403
+ persistEnergy();
2246
2404
  prewarm.abort?.abort();
2247
2405
  clearInterval(metricsTimer);
2248
2406
  metricsTimer = undefined;
@@ -2257,6 +2415,40 @@ export default function (pi: ExtensionAPI) {
2257
2415
 
2258
2416
  pi.on("agent_start", async () => {
2259
2417
  state.runStart = Date.now();
2418
+ lastCall = "";
2419
+ });
2420
+
2421
+ pi.on("before_agent_start", async (event, ctx) => {
2422
+ let rules = MAGI_MD;
2423
+ try {
2424
+ rules = readFileSync(join(ctx.cwd, "MAGI.md"), "utf8");
2425
+ } catch {
2426
+ // no MAGI.md (deleted, or unwritable cwd): the built-in rules
2427
+ }
2428
+ return rules.trim() ? { systemPrompt: `${event.systemPrompt}\n\n${rules.trim()}` } : undefined;
2429
+ });
2430
+
2431
+ // before every request: drop old thinking and tool outputs from what the model sees, never from the session
2432
+ pi.on("context", async (event) => {
2433
+ if (!hygiene.enabled) return;
2434
+ const pruned = pruneContext(event.messages as any, hygiene);
2435
+ hygieneStats = pruned.stats;
2436
+ state.pruned = pruned.stats.prunedTokens;
2437
+ return { messages: pruned.messages as any };
2438
+ });
2439
+
2440
+ // local models loop on the same call, and may copy a pruned placeholder into a file
2441
+ pi.on("tool_call", async (event) => {
2442
+ const args = JSON.stringify(event.input);
2443
+ if ((event.toolName === "write" || event.toolName === "edit") && args.includes(PRUNED_MARK)) {
2444
+ return { block: true, reason: `MAGI: this ${event.toolName} contains a "${PRUNED_MARK}" placeholder, not real content. Read the file and use the actual text.` };
2445
+ }
2446
+ const fingerprint = event.toolName + args;
2447
+ repeats = fingerprint === lastCall ? repeats + 1 : 1;
2448
+ lastCall = fingerprint;
2449
+ if (repeats >= LOOP_REPEATS) {
2450
+ return { block: true, reason: `MAGI: identical ${event.toolName} call ${repeats} times in a row. Repeating it will not change the outcome: change approach, or tell the user what blocks you.` };
2451
+ }
2260
2452
  });
2261
2453
 
2262
2454
  pi.on("turn_start", async (_event, ctx) => {
@@ -2300,7 +2492,19 @@ export default function (pi: ExtensionAPI) {
2300
2492
  pi.on("message_end", async (event) => {
2301
2493
  if (event.message.role !== "assistant") return;
2302
2494
  const m = event.message as AssistantMessage;
2303
- addUsage(m);
2495
+ countAssistant(m);
2496
+ // learn how long this model thinks in this phase; a cut must not raise the budget
2497
+ const thought = thinkingTokens(m as any);
2498
+ if (pendingBudget && thought > 0 && m.stopReason !== "aborted" && m.stopReason !== "error") {
2499
+ const verdict = budgetVerdict(m as any, pendingBudget.tokens);
2500
+ if (verdict === "ignored") budgetIgnored.add(pendingBudget.model);
2501
+ if (verdict === "cut" && !thinkingWasCut(m as any)) state.thinkCuts++; // silent cut: no budget message on the server
2502
+ const samples = ((learned[pendingBudget.model] ??= {})[pendingBudget.phase] ??= []);
2503
+ samples.push(budgetSample(thought, pendingBudget.tokens, verdict === "cut"));
2504
+ samples.splice(0, samples.length - BUDGET_WINDOW);
2505
+ persistBudget();
2506
+ }
2507
+ pendingBudget = undefined;
2304
2508
  if (perf.start) {
2305
2509
  const end = Date.now();
2306
2510
  perf.lastMs = end - perf.start;
@@ -2436,9 +2640,11 @@ export default function (pi: ExtensionAPI) {
2436
2640
  }
2437
2641
 
2438
2642
  pi.registerCommand("magi", {
2439
- description: "Ask the three MAGI (pragmatist, guardian, visionary); /magi review [focus] judges the git diff; /magi config assigns models; /magi mecha picks the model",
2643
+ description: "Ask the three MAGI, or manage them: review [focus] · config · mecha · status · panel · compact · cost · on · off · hygiene · budget (type a space to see them all)",
2644
+ getArgumentCompletions: (prefix) => argCompletions(MAGI_ARGS, prefix),
2440
2645
  handler: async (args, ctx) => {
2441
2646
  const arg = args.trim();
2647
+ if (UI_ARGS.test(arg)) return manageUi(arg, ctx);
2442
2648
  if (arg === "config") return configureMagi(ctx);
2443
2649
  if (arg === "mecha") {
2444
2650
  if (ctx.mode !== "tui" || !swap.base) return ctx.ui.notify("MECHA SELECT needs the TUI and a llama-swap model", "error");
@@ -2446,8 +2652,12 @@ export default function (pi: ExtensionAPI) {
2446
2652
  return pickModel(ctx);
2447
2653
  }
2448
2654
 
2449
- if (arg === "review" || arg.startsWith("review ")) {
2450
- const focus = arg.slice("review".length).trim();
2655
+ if (arg !== "council" && !arg.startsWith("council ")) {
2656
+ return ctx.ui.notify(arg ? `Unknown /magi command "${arg}": to ask the MAGI, /magi council <question>` : "Ask the MAGI with /magi council <question>; type /magi and a space to see every command", arg ? "warning" : "info");
2657
+ }
2658
+ const ask = arg.slice("council".length).trim();
2659
+ if (ask === "review" || ask.startsWith("review ")) {
2660
+ const focus = ask.slice("review".length).trim();
2451
2661
  let changes: { diff: string; untracked: string[] };
2452
2662
  try {
2453
2663
  changes = await pendingChanges(ctx.cwd);
@@ -2470,93 +2680,151 @@ export default function (pi: ExtensionAPI) {
2470
2680
  return runCouncil(ctx, question, "```diff\n" + diff + "\n```" + untracked, "Pending changes (git diff HEAD)");
2471
2681
  }
2472
2682
 
2473
- const question = arg || (await ctx.ui.input("Question for the MAGI:", "should we …?"))?.trim() || "";
2683
+ const question = ask || (await ctx.ui.input("Question for the MAGI:", "should we …?"))?.trim() || "";
2474
2684
  if (!question) return;
2475
2685
  return runCouncil(ctx, question, conversationExcerpt(ctx));
2476
2686
  },
2477
2687
  });
2478
2688
 
2479
- pi.registerCommand("magi-ui", {
2480
- description: "MAGI chrome: enable the theme, or manage it (on|off|panel|compact|status|config)",
2481
- handler: async (args, ctx) => {
2482
- liveCtx = ctx;
2483
- const arg = args.trim().toLowerCase();
2484
-
2485
- if (arg === "panel") {
2486
- panelEnabled = !panelEnabled;
2487
- if (panelEnabled) showPanel(ctx.ui.theme);
2488
- else hidePanel();
2489
- ctx.ui.notify(`Side panel ${panelEnabled ? "enabled" : "disabled"}`, "info");
2689
+ /** /magi on|off|panel|compact|status|cost|hygiene|budget: the theme, its panel and the local-model settings. */
2690
+ async function manageUi(args: string, ctx: ExtensionCommandContext) {
2691
+ liveCtx = ctx;
2692
+ const arg = args.trim().toLowerCase();
2693
+
2694
+ if (arg === "hygiene" || arg.startsWith("hygiene ")) {
2695
+ const sub = arg.slice("hygiene".length).trim();
2696
+ const keep = /^(\d+)\s+(\d+)$/.exec(sub);
2697
+ const step = /^step\s+(\d+)(k?)$/.exec(sub);
2698
+ if (sub === "on" || sub === "off" || keep || step) {
2699
+ if (keep) Object.assign(hygiene, { enabled: true, keepThinkingTurns: Number(keep[1]), keepToolResults: Number(keep[2]) });
2700
+ else if (step) hygiene.stepTokens = Math.max(1000, Number(step[1]) * (step[2] ? 1000 : 1));
2701
+ else hygiene.enabled = sub === "on";
2702
+ const cfg = loadMagiConfig();
2703
+ const { enabled, keepThinkingTurns, keepToolResults, stepTokens } = hygiene;
2704
+ saveMagiConfig({ ...cfg, hygiene: { ...cfg.hygiene, enabled, keepThinkingTurns, keepToolResults, stepTokens } });
2705
+ } else if (sub) {
2706
+ ctx.ui.notify("Usage: /magi hygiene [on|off|step <tokens>|<thinking turns kept> <tool results kept>], e.g. step 40k, 3 5", "warning");
2490
2707
  return;
2491
2708
  }
2492
- if (arg === "compact") {
2493
- ui.compact = !ui.compact;
2494
- const cfg = loadMagiConfig();
2495
- saveMagiConfig({ ...cfg, ui: { ...cfg.ui, compact: ui.compact } });
2496
- repaint();
2497
- ctx.ui.notify(`Side panel ${ui.compact ? "compact" : "detailed"}`, "info");
2709
+ const s = hygieneStats;
2710
+ ctx.ui.notify(
2711
+ (!hygiene.enabled
2712
+ ? "Context hygiene off: /magi hygiene on"
2713
+ : s
2714
+ ? `Context hygiene: ${fmtTokens(s.prunedTokens)} tokens pruned in the first ${s.watermark}/${s.messages} messages, ${fmtTokens(s.pendingTokens)} waiting for the next step (every ${fmtTokens(hygiene.stepTokens)})`
2715
+ : "Context hygiene on: nothing sent to the model yet") +
2716
+ ` · keeps thinking of the last ${hygiene.keepThinkingTurns} turns, the last ${hygiene.keepToolResults} tool results` +
2717
+ (state.thinkCuts ? ` · thinking cut at the budget ${state.thinkCuts}×` : ""),
2718
+ "info",
2719
+ );
2720
+ return;
2721
+ }
2722
+ if (arg === "budget" || arg.startsWith("budget ")) {
2723
+ const sub = arg.slice("budget".length).trim();
2724
+ const model = ctx.model?.id ?? "";
2725
+ const tokens = (t: string) => Math.round(Number(t.replace(/k$/, "")) * (t.endsWith("k") ? 1024 : 1));
2726
+ const fixed = /^(\d+k?)\s+(\d+k?)$/.exec(sub);
2727
+ if (sub === "auto" || sub === "off") budgetMode = sub;
2728
+ else if (sub === "reset") delete learned[model];
2729
+ else if (sub === "message") budgetMessage = !budgetMessage;
2730
+ else if (fixed) {
2731
+ budgetMode = "fixed";
2732
+ fixedBudget = { planning: tokens(fixed[1]!), acting: tokens(fixed[2]!) };
2733
+ } else if (sub) {
2734
+ ctx.ui.notify("Usage: /magi budget [auto|off|reset|message|<planning> <acting>], e.g. 16k 4k", "warning");
2498
2735
  return;
2499
2736
  }
2500
- if (arg === "config") {
2501
- const currency = await ctx.ui.select(`Currency for COST (current: ${ui.currency})`, ["EUR", "USD"]);
2502
- if (!currency) return;
2503
- const current = ui.kwhPrice !== undefined ? ` (current: ${ui.kwhPrice})` : "";
2504
- const raw = (await ctx.ui.input(`Electricity price per kWh in ${currency}${current}:`, "0.30"))?.trim();
2505
- if (raw === undefined) return;
2506
- const price = raw === "" && ui.kwhPrice !== undefined ? ui.kwhPrice : Number(raw.replace(",", "."));
2507
- if (raw === "" && ui.kwhPrice === undefined) {
2508
- ctx.ui.notify("No price entered: COST unchanged", "warning");
2509
- return;
2510
- }
2511
- if (!Number.isFinite(price) || price < 0) {
2512
- ctx.ui.notify(`Invalid price: "${raw}"`, "error");
2513
- return;
2514
- }
2515
- ui.currency = currency === "USD" ? "USD" : "EUR";
2516
- ui.kwhPrice = price;
2517
- const cfg = loadMagiConfig();
2518
- saveMagiConfig({ ...cfg, ui: { ...cfg.ui, currency: ui.currency, kwhPrice: price } });
2519
- repaint();
2520
- ctx.ui.notify(`COST: ${fmtMoney(price)} per kWh`, "info");
2737
+ if (sub) persistBudget();
2738
+ const phase = (p: BudgetPhase) => {
2739
+ const n = learned[model]?.[p]?.length ?? 0;
2740
+ const how = budgetMode === "fixed" ? "fixed" : learnedBudget(learned[model]?.[p] ?? [], p) ? `learned from ${n}` : `default, learning ${n}/10`;
2741
+ return `${p} ${budgetFor(model, p) / 1024}k (${how})`;
2742
+ };
2743
+ ctx.ui.notify(
2744
+ budgetMode === "off"
2745
+ ? "Thinking budget off: llama-server decides. /magi budget auto"
2746
+ : `Thinking budget ${budgetMode.toUpperCase()} · ${model || "no model"}: ${phase("planning")} · ${phase("acting")}` +
2747
+ ` · closing message ${budgetMessage ? "on" : "off"}` +
2748
+ (swap.base ? "" : " · applies to llama-swap models only") +
2749
+ (budgetIgnored.has(model) ? " · ⚠ llama-server thinks past it: it was started with its own --reasoning-budget (or is too old), which wins" : ""),
2750
+ "info",
2751
+ );
2752
+ return;
2753
+ }
2754
+ if (arg === "panel") {
2755
+ panelEnabled = !panelEnabled;
2756
+ if (panelEnabled) showPanel(ctx.ui.theme);
2757
+ else hidePanel();
2758
+ ctx.ui.notify(`Side panel ${panelEnabled ? "enabled" : "disabled"}`, "info");
2759
+ return;
2760
+ }
2761
+ if (arg === "compact") {
2762
+ ui.compact = !ui.compact;
2763
+ const cfg = loadMagiConfig();
2764
+ saveMagiConfig({ ...cfg, ui: { ...cfg.ui, compact: ui.compact } });
2765
+ repaint();
2766
+ ctx.ui.notify(`Side panel ${ui.compact ? "compact" : "detailed"}`, "info");
2767
+ return;
2768
+ }
2769
+ if (arg === "cost") {
2770
+ const currency = await ctx.ui.select(`Currency for COST (current: ${ui.currency})`, ["EUR", "USD"]);
2771
+ if (!currency) return;
2772
+ const current = ui.kwhPrice !== undefined ? ` (current: ${ui.kwhPrice})` : "";
2773
+ const raw = (await ctx.ui.input(`Electricity price per kWh in ${currency}${current}:`, "0.30"))?.trim();
2774
+ if (raw === undefined) return;
2775
+ const price = raw === "" && ui.kwhPrice !== undefined ? ui.kwhPrice : Number(raw.replace(",", "."));
2776
+ if (raw === "" && ui.kwhPrice === undefined) {
2777
+ ctx.ui.notify("No price entered: COST unchanged", "warning");
2521
2778
  return;
2522
2779
  }
2523
- if (arg === "status") {
2524
- if (!swap.base) {
2525
- ctx.ui.notify("/magi-ui status needs a llama-swap session model", "warning");
2526
- return;
2527
- }
2528
- let report: { data?: ActivityRow[]; total?: number };
2529
- try {
2530
- report = (await (await swapGet(`/api/metrics/activity?limit=${ACTIVITY_REPORT_ROWS}`, 15_000)).json()) as typeof report;
2531
- } catch (err) {
2532
- ctx.ui.notify(`llama-swap status failed: ${err instanceof Error ? err.message : String(err)}`, "error");
2533
- return;
2534
- }
2535
- const rows = report.data ?? [];
2536
- if (!rows.length) {
2537
- ctx.ui.notify("llama-swap has no recorded requests yet", "info");
2538
- return;
2539
- }
2540
- await ctx.ui.custom<void>((_tui, theme, _keys, done) =>
2541
- buildReportView(theme, activityReport(theme, rows, report.total ?? rows.length), () => done(undefined)),
2542
- );
2780
+ if (!Number.isFinite(price) || price < 0) {
2781
+ ctx.ui.notify(`Invalid price: "${raw}"`, "error");
2543
2782
  return;
2544
2783
  }
2545
- if (arg === "off" || arg === "on") {
2546
- chrome = arg === "on";
2547
- applyChrome(ctx);
2548
- ctx.ui.notify(`MAGI chrome ${chrome ? "enabled" : "disabled"}`, "info");
2784
+ ui.currency = currency === "USD" ? "USD" : "EUR";
2785
+ ui.kwhPrice = price;
2786
+ const cfg = loadMagiConfig();
2787
+ saveMagiConfig({ ...cfg, ui: { ...cfg.ui, currency: ui.currency, kwhPrice: price } });
2788
+ repaint();
2789
+ ctx.ui.notify(`COST: ${fmtMoney(price)} per kWh`, "info");
2790
+ return;
2791
+ }
2792
+ if (arg === "status") {
2793
+ if (!swap.base) {
2794
+ ctx.ui.notify("/magi status needs a llama-swap session model", "warning");
2549
2795
  return;
2550
2796
  }
2551
-
2552
- const res = ctx.ui.setTheme("magi");
2553
- if (!res.success) {
2554
- ctx.ui.notify(`Theme magi not found: ${res.error}`, "error");
2797
+ let report: { data?: ActivityRow[]; total?: number };
2798
+ try {
2799
+ report = (await (await swapGet(`/api/metrics/activity?limit=${ACTIVITY_REPORT_ROWS}`, 15_000)).json()) as typeof report;
2800
+ } catch (err) {
2801
+ ctx.ui.notify(`llama-swap status failed: ${err instanceof Error ? err.message : String(err)}`, "error");
2555
2802
  return;
2556
2803
  }
2557
- chrome = true;
2804
+ const rows = report.data ?? [];
2805
+ if (!rows.length) {
2806
+ ctx.ui.notify("llama-swap has no recorded requests yet", "info");
2807
+ return;
2808
+ }
2809
+ await ctx.ui.custom<void>((_tui, theme, _keys, done) =>
2810
+ buildReportView(theme, activityReport(theme, rows, report.total ?? rows.length), () => done(undefined)),
2811
+ );
2812
+ return;
2813
+ }
2814
+ if (arg === "off") {
2815
+ chrome = false;
2558
2816
  applyChrome(ctx);
2559
- ctx.ui.notify("MAGI online — /magi-ui panel|compact|status|off", "info");
2560
- },
2561
- });
2817
+ ctx.ui.notify("MAGI chrome disabled", "info");
2818
+ return;
2819
+ }
2820
+
2821
+ const res = ctx.ui.setTheme("magi");
2822
+ if (!res.success) {
2823
+ ctx.ui.notify(`Theme magi not found: ${res.error}`, "error");
2824
+ return;
2825
+ }
2826
+ chrome = true;
2827
+ applyChrome(ctx);
2828
+ ctx.ui.notify("MAGI online — /magi panel|compact|status|off", "info");
2829
+ }
2562
2830
  }
@@ -0,0 +1,240 @@
1
+ /**
2
+ * Local-model helpers for MAGI: context hygiene, loop guard, and the MAGI.md rules template.
3
+ *
4
+ * Context hygiene rewrites the messages pi sends to the model (never the saved session):
5
+ * - old thinking blocks are dropped: Qwen templates resend every reasoning block of an agent run,
6
+ * one 18k-token think stays in the context until compaction
7
+ * - old tool outputs become a one-line marker (observation masking, Lindenbauer et al. 2025:
8
+ * as good as LLM summarization, at half the cost)
9
+ * - old write/edit payloads become a marker too: the file on disk is the source of truth
10
+ * Pruning advances in steps behind a watermark, so between steps the prompt only grows at the end
11
+ * and llama.cpp keeps reusing its KV cache.
12
+ */
13
+
14
+ export type Msg = { role: string; content?: any; [key: string]: any };
15
+
16
+ export interface HygieneOptions {
17
+ keepThinkingTurns: number; // newest assistant messages that keep their thinking
18
+ keepToolResults: number; // newest tool results kept whole
19
+ stepTokens: number; // the watermark moves each time the prunable total crosses another multiple of this
20
+ minPruneChars: number; // smaller outputs/payloads are left alone
21
+ }
22
+
23
+ export interface HygieneStats {
24
+ messages: number;
25
+ watermark: number; // messages before this index are pruned
26
+ prunedTokens: number;
27
+ pendingTokens: number; // prunable, waiting for the next step
28
+ stepTokens: number;
29
+ }
30
+
31
+ export const HYGIENE_DEFAULTS: HygieneOptions = { keepThinkingTurns: 3, keepToolResults: 5, stepTokens: 15_000, minPruneChars: 600 };
32
+
33
+ // ponytail: chars/token measured on Qwen3.x pi sessions (2.9–3.2); a tokenizer call would be exact but costs a request
34
+ const CHARS_PER_TOKEN = 3;
35
+ const IMAGE_CHARS = 4800;
36
+
37
+ /** Marks every placeholder, so a model that copies one into a real write/edit gets blocked (see loop guard). */
38
+ export const PRUNED_MARK = "<<pruned by MAGI";
39
+
40
+ const lines = (s: string) => s.split("\n").length;
41
+ const clip = (s: string, n = 80) => (s.length > n ? s.slice(0, n - 1) + "…" : s);
42
+
43
+ function callTarget(call: any): string {
44
+ const a = call?.arguments ?? {};
45
+ return clip(String(a.path ?? a.command ?? a.pattern ?? a.query ?? a.url ?? ""));
46
+ }
47
+
48
+ function resultChars(m: Msg): number {
49
+ let n = 0;
50
+ for (const c of m.content ?? []) n += c.type === "image" ? IMAGE_CHARS : (c.text?.length ?? 0);
51
+ return n;
52
+ }
53
+
54
+ function editChars(call: any): number {
55
+ return (call.arguments?.edits ?? []).reduce((n: number, e: any) => n + (e.oldText?.length ?? 0) + (e.newText?.length ?? 0), 0);
56
+ }
57
+
58
+ /** Chars pruning would remove from this assistant message. */
59
+ function assistantSavings(m: Msg, o: HygieneOptions): number {
60
+ let n = 0;
61
+ for (const c of m.content ?? []) {
62
+ if (c.type === "thinking") n += c.thinking?.length ?? 0;
63
+ else if (c.type === "toolCall" && c.name === "write" && (c.arguments?.content?.length ?? 0) >= o.minPruneChars) n += c.arguments.content.length;
64
+ else if (c.type === "toolCall" && c.name === "edit" && editChars(c) >= o.minPruneChars) n += editChars(c);
65
+ }
66
+ return n;
67
+ }
68
+
69
+ function pruneAssistant(m: Msg, o: HygieneOptions): Msg {
70
+ const content = [];
71
+ for (const c of m.content ?? []) {
72
+ if (c.type === "thinking") continue;
73
+ if (c.type === "toolCall" && c.name === "write" && (c.arguments?.content?.length ?? 0) >= o.minPruneChars) {
74
+ const text = c.arguments.content as string;
75
+ content.push({ ...c, arguments: { ...c.arguments, content: `${PRUNED_MARK}: ${lines(text)} lines written, the file on disk is the source of truth>>` } });
76
+ } else if (c.type === "toolCall" && c.name === "edit" && editChars(c) >= o.minPruneChars) {
77
+ const edits = c.arguments.edits.map((e: any) => ({ oldText: `${PRUNED_MARK}>>`, newText: `${PRUNED_MARK}: ${lines(e.newText ?? "")} lines>>` }));
78
+ content.push({ ...c, arguments: { ...c.arguments, edits } });
79
+ } else content.push(c);
80
+ }
81
+ // a message that only held thinking keeps a marker: some templates choke on an empty assistant turn
82
+ if (!content.length) content.push({ type: "text", text: `${PRUNED_MARK}: reasoning>>` });
83
+ return { ...m, content };
84
+ }
85
+
86
+ function pruneResult(m: Msg, call: any): Msg {
87
+ const text = (m.content ?? []).map((c: any) => c.text ?? "").join("\n");
88
+ const what = [m.toolName, callTarget(call)].filter(Boolean).join(" ");
89
+ return { ...m, content: [{ type: "text", text: `${PRUNED_MARK}: output of ${what}, ${lines(text)} lines. Run it again if you need it.>>` }] };
90
+ }
91
+
92
+ /**
93
+ * Prunes everything before the watermark. The watermark sits on the last point where the prunable chars,
94
+ * summed from the first message, crossed a multiple of stepTokens. Only messages before the protected tail
95
+ * count, and the tail only moves forward, so earlier crossings never move: the pruned prefix is stable.
96
+ */
97
+ export function pruneContext(messages: Msg[], o: HygieneOptions = HYGIENE_DEFAULTS): { messages: Msg[]; stats: HygieneStats } {
98
+ const lastIndices = (role: string, keep: number) =>
99
+ messages.flatMap((m, i) => (m.role === role ? [i] : [])).slice(-keep);
100
+ const tail = [...lastIndices("assistant", o.keepThinkingTurns), ...lastIndices("toolResult", o.keepToolResults)];
101
+ const tailStart = tail.length ? Math.min(...tail) : messages.length;
102
+
103
+ const calls = new Map<string, any>();
104
+ for (const m of messages) if (m.role === "assistant") for (const c of m.content ?? []) if (c.type === "toolCall") calls.set(c.id, c);
105
+
106
+ const savings = (m: Msg) =>
107
+ m.role === "assistant" ? assistantSavings(m, o) : m.role === "toolResult" && resultChars(m) >= o.minPruneChars ? resultChars(m) : 0;
108
+
109
+ const step = o.stepTokens * CHARS_PER_TOKEN;
110
+ let sum = 0;
111
+ let watermark = 0;
112
+ let prunedChars = 0;
113
+ for (let i = 0; i < tailStart; i++) {
114
+ sum += savings(messages[i]!);
115
+ if (Math.floor(sum / step) > Math.floor(prunedChars / step)) {
116
+ watermark = i + 1;
117
+ prunedChars = sum;
118
+ }
119
+ }
120
+
121
+ const out = messages.map((m, i) => {
122
+ if (i >= watermark || !savings(m)) return m;
123
+ return m.role === "assistant" ? pruneAssistant(m, o) : pruneResult(m, calls.get(m.toolCallId));
124
+ });
125
+ return {
126
+ messages: out,
127
+ stats: {
128
+ messages: messages.length,
129
+ watermark,
130
+ prunedTokens: Math.round(prunedChars / CHARS_PER_TOKEN),
131
+ pendingTokens: Math.round((sum - prunedChars) / CHARS_PER_TOKEN),
132
+ stepTokens: o.stepTokens,
133
+ },
134
+ };
135
+ }
136
+
137
+ /**
138
+ * Sent as reasoning_budget_message with every request: when the thinking budget runs out llama.cpp forces this text,
139
+ * then the end-of-thinking tag, and the model reads it as its own words, so it says what to do next.
140
+ * A llama-server older than the per-request message closes the thinking silently; cuts are then detected by length (budgetVerdict).
141
+ */
142
+ export const BUDGET_MESSAGE = "Time is up. I will take the smallest safe next step with what I know, and write my open plan into PLAN.md.";
143
+
144
+ export type BudgetPhase = "planning" | "acting"; // right after the user spoke / between tool calls
145
+ export type BudgetMode = "auto" | "fixed" | "off";
146
+
147
+ /** Starting budgets, used until a model has enough samples to learn its own. */
148
+ export const BUDGET_DEFAULTS: Record<BudgetPhase, number> = { planning: 16384, acting: 4096 };
149
+ export const BUDGET_LIMITS: Record<BudgetPhase, [number, number]> = { planning: [4096, 32768], acting: [2048, 4096] }; // acting: past ~4k between tool calls it is overthinking
150
+
151
+ const BUDGET_SAMPLES_MIN = 10;
152
+ export const BUDGET_WINDOW = 30; // thinking lengths kept per model and phase
153
+ const BUDGET_PERCENTILE = 0.95;
154
+ const BUDGET_HEADROOM = 1.5;
155
+
156
+ /** Phase of the next request, from the OpenAI-style payload: a tool result last means the agent is mid-task. */
157
+ export function requestPhase(payload: any): BudgetPhase {
158
+ return payload?.messages?.at(-1)?.role === "tool" ? "acting" : "planning";
159
+ }
160
+
161
+ /**
162
+ * Learned budget: 1.5 × the 95th percentile of recent thinking lengths, clamped and rounded to 1k.
163
+ * Cuts never raise it (see budgetSample): a model that often runs away is held at its budget instead of chased.
164
+ * Replies that finish close under the budget raise it, a model that thinks little pulls it down.
165
+ * Undefined until there are enough samples.
166
+ */
167
+ export function learnedBudget(samples: number[], phase: BudgetPhase): number | undefined {
168
+ if (samples.length < BUDGET_SAMPLES_MIN) return undefined;
169
+ const sorted = [...samples].sort((a, b) => a - b);
170
+ const p = sorted[Math.min(sorted.length - 1, Math.floor(sorted.length * BUDGET_PERCENTILE))]!;
171
+ const [lo, hi] = BUDGET_LIMITS[phase];
172
+ return Math.min(hi, Math.max(lo, Math.ceil((p * BUDGET_HEADROOM) / 1024) * 1024));
173
+ }
174
+
175
+ /** The sample to learn from a reply: a cut counts as budget / headroom, so cuts alone give back the same budget. */
176
+ export function budgetSample(thought: number, budget: number, cut: boolean): number {
177
+ return cut ? budget / BUDGET_HEADROOM : thought;
178
+ }
179
+
180
+ /**
181
+ * What the server did with the budget sent for this reply, from the thinking length (estimated, ±10%):
182
+ * cut near the budget, "ignored" well past it (llama-server started with its own --reasoning-budget, or too old
183
+ * to read thinking_budget_tokens), otherwise within.
184
+ */
185
+ export function budgetVerdict(m: Msg, budget: number): "cut" | "ignored" | "within" {
186
+ if (thinkingWasCut(m)) return "cut";
187
+ const thought = thinkingTokens(m);
188
+ return thought > budget * 1.3 ? "ignored" : thought >= budget * 0.9 ? "cut" : "within";
189
+ }
190
+
191
+ /** Estimated thinking tokens of a reply (same chars/token as the hygiene). */
192
+ export function thinkingTokens(m: Msg): number {
193
+ const chars = (m.content ?? []).reduce((n: number, c: any) => n + (c.type === "thinking" ? (c.thinking?.length ?? 0) : 0), 0);
194
+ return Math.round(chars / CHARS_PER_TOKEN);
195
+ }
196
+
197
+ /** Whether llama.cpp cut this message's thinking at the budget. */
198
+ export function thinkingWasCut(m: Msg): boolean {
199
+ return (m.content ?? []).some((c: any) => c.type === "thinking" && (c.thinking ?? "").trimEnd().endsWith(BUDGET_MESSAGE));
200
+ }
201
+
202
+ /**
203
+ * Written to <project>/MAGI.md when missing, then appended to the system prompt on every run.
204
+ * English on purpose: Qwen-family models reason in English and follow English rules more reliably.
205
+ * Kept short: it is paid for on every request.
206
+ */
207
+ export const MAGI_MD = `# MAGI rules for local models
208
+
209
+ These rules are appended to the system prompt by the MAGI extension. Edit them freely; delete the file to get the defaults back.
210
+
211
+ ## Think less, act more
212
+ - Keep reasoning short: understand the step, decide, act. Do not re-plan what is already decided.
213
+ - Never draft code or file contents in your reasoning. Write them directly with the write/edit tool.
214
+ - If two attempts at the same approach fail, stop and change approach, or ask the user. Do not retry blindly.
215
+ - If your reasoning stopped abruptly mid-thought, or ends with "${BUDGET_MESSAGE.split(". ")[0]}.", your thinking budget ran out: in that turn make no large or irreversible change. Write your open plan into PLAN.md or take one small step you can verify; the next turn gives you a fresh budget.
216
+
217
+ ## Your context is small: spend it carefully
218
+ - Search before reading: use rg/find to locate code, then read only the needed range (offset/limit).
219
+ - Do not read whole large files, and do not re-read a file you just wrote.
220
+ - Trim long command output: pipe through tail, head or rg.
221
+ - Old tool outputs and old reasoning are removed from your context automatically (marked "${PRUNED_MARK}…>>"). Anything you will need later must be written down (see below). Never copy those markers into a file.
222
+
223
+ ## Keep the task state on disk
224
+ - For any task longer than a few steps, keep PLAN.md (a checklist) and NOTES.md (findings, decisions, commands that work).
225
+ - Update them after every completed step. If the context is compacted or a new session starts, continue from PLAN.md.
226
+
227
+ ## Edit safely
228
+ - edit oldText must match the file exactly: copy it from a fresh read, keep it small and unique.
229
+ - If an edit fails, read that region again instead of guessing.
230
+ - For big new files: write a skeleton first, then fill it with edits. Never replace a whole file with a partial version.
231
+
232
+ ## Verify, never assume
233
+ - Never invent paths, functions, APIs or flags: check with ls, rg, --help or the docs first.
234
+ - After a change, run the build, the tests or the program. Claim success only when a tool result proves it.
235
+ - If you cannot verify something, say so explicitly.
236
+
237
+ ## Finish cleanly
238
+ - When done: say briefly what changed, how you verified it, and what is left.
239
+ - Answer in the user's language.
240
+ `;
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "pi-magi-theme",
3
- "version": "0.2.4",
3
+ "version": "0.3.1",
4
4
  "description": "MAGI SYSTEM theme + extension for pi (Evangelion fan art): MAGI control screen panel, three-model /magi council, MECHA SELECT model picker, angel-attack loading, seven-seal context gauge, llama-swap telemetry",
5
5
  "keywords": [
6
6
  "pi-package",