pi-magi-theme 0.2.3 → 0.3.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -40,10 +40,14 @@ Clone the repo and point pi at it instead (edits in the repo are live on the nex
40
40
  | `/magi review [focus]` | the council reviews your pending changes (`git diff HEAD` plus untracked file names) before you commit |
41
41
  | `/magi config` | pick a model for each MAGI |
42
42
  | `/magi mecha` | MECHA SELECT: pick the llama-swap model to activate, each shown as a mecha head lit by its real state |
43
- | `/magi-ui compact` | toggle the compact side panel (basic info and animations only); remembered across sessions |
44
- | `/magi-ui status` | llama-swap report from its last 100 requests: speed, tokens, cache hits, MTP draft acceptance, durations, errors per model |
45
- | `/magi-ui config` | set the electricity price per kWh and the currency (EUR or USD) for the COST row |
46
- | `/magi-ui panel` · `on` · `off` | hide/show the side panel, enable/disable the whole chrome |
43
+ | `/magi compact` | toggle the compact side panel (basic info and animations only); remembered across sessions |
44
+ | `/magi status` | llama-swap report from its last 100 requests: speed, tokens, cache hits, MTP draft acceptance, durations, errors per model |
45
+ | `/magi cost` | set the electricity price per kWh and the currency (EUR or USD) for the COST row |
46
+ | `/magi panel` · `on` · `off` | hide/show the side panel, enable/disable the whole chrome |
47
+ | `/magi hygiene` · `on` · `off` · `step <tokens>` · `<turns> <results>` | show how much context the hygiene pruned; enable/disable it; how many prunable tokens make a pruning step (e.g. `step 40k`, default 15k); how many recent turns keep their thinking and how many tool results stay whole (e.g. `3 5`) |
48
+ | `/magi budget` · `auto` · `off` · `reset` · `message` · `<planning> <acting>` | show the thinking budget of the current model; learn it per model (default); leave it to llama-server; forget what was learned for this model; turn the closing message off/on; or fix it (e.g. `16k 4k`) |
49
+
50
+ Everything lives under `/magi`: type `/magi ` (with the space) to see every option with a short description; keep typing to narrow it down, Tab or Enter to pick one. Anything that is not an option is a question for the council.
47
51
 
48
52
  ## Lore ↔ function
49
53
 
@@ -72,6 +76,8 @@ The window title is `π - Magi - <working directory>`. After a run longer than 3
72
76
 
73
77
  In fullscreen mode the side panel always reaches the bottom of the terminal.
74
78
 
79
+ In a terminal too narrow for it, the footer scrolls as one line instead of being cut off.
80
+
75
81
  ## The seventh seal: smart compaction
76
82
 
77
83
  The theme does not compact anything itself: it shows who does. For better compaction install [pi-smart-compact](https://www.npmjs.com/package/pi-smart-compact):
@@ -82,6 +88,36 @@ pi install npm:pi-smart-compact
82
88
 
83
89
  It extracts files, errors, decisions and open loops locally (no LLM calls), then synthesizes and verifies the summary. Point its `summaryModel` at a local model to keep compaction free. When it is installed, the sixth seal suggests `/smart-compact`, the seals name it while they break (`✶ BREAKING THE SEALS · smart-compact · 4s`) and the seventh seal reports who actually produced the summary: `smart-compact`, or `pi native` if it fell back to pi's own compactor.
84
90
 
91
+ ## Local models: context hygiene, thinking budget, loop guard, MAGI.md
92
+
93
+ Local models run out of context on long tasks well before they run out of work. They also tend to think for minutes between two tool calls and to repeat the same command when stuck. MAGI works on all three, automatically (the thinking budget for llama-swap models, the rest for any model):
94
+
95
+ - **Context hygiene: the model's memory stays lean.** *Problem:* every file the agent reads and every long reasoning stays in the conversation, until the model's context is full and the task falls apart. *What MAGI does:* before each request it replaces old reasoning, old tool outputs and old file writes with a one-line note (`<<pruned by MAGI…>>`). Only what the model sees is trimmed: your saved session stays complete. The latest 3 turns keep their reasoning and the latest 5 tool outputs stay whole. On two real sessions it brought 104k and 119k tokens down to ~64k and ~47k.
96
+ - **Thinking budget: no more ten-minute thinks.** *Problem:* a local model can reason for thousands of tokens before a simple step, and at 10 tokens/s that is minutes of waiting. *What MAGI does:* it gives the model a maximum length of thinking on every request: larger right after you write (planning), smaller between tool calls (acting). When the limit is reached the model is stopped mid-thought, says *"Time is up. I will take the smallest safe next step…"* and acts. The limit adapts to each model on its own. The side panel counts these cuts (`HYGIENE -18.2k ✂2`).
97
+ - **Loop guard: no endless retries.** *Problem:* a stuck model runs the same command again and again. *What MAGI does:* the third identical tool call in a row is blocked, with a message asking the model to try something else.
98
+ - **MAGI.md: house rules for the model.** *Problem:* local models repeat the same mistakes: reading whole files, inventing paths, claiming success without checking. *What MAGI does:* it creates `MAGI.md` in your project on the first start (never overwritten) and adds it to the model's instructions on every run: short rules against these mistakes, plus the habit of keeping the task plan in `PLAN.md` and findings in `NOTES.md`, so nothing important is lost when old context is trimmed. Edit it per project; delete it to get the defaults back.
99
+
100
+ Nothing needs setting up in llama-swap. The defaults suit most tasks; two adjustments are worth knowing:
101
+
102
+ - **Long tasks: `/magi hygiene step 40k`.** Each time the hygiene trims, the server has to re-read part of the conversation, and on some models (see *Hybrid models* below) almost all of it, which can take a few minutes. A step of 40k trims less often: on a real session it cut the re-reading from ~6 minutes to ~1.
103
+ - **A model that really needs to think longer: `/magi budget <planning> <acting>`**, e.g. `/magi budget 16k 8k`. Frequent ✂ cuts in the side panel are the sign.
104
+
105
+ ### How it works
106
+
107
+ For the curious, and for tuning.
108
+
109
+ **Hygiene.** Trimming happens in steps, not on every request: the conversation up to a mark is trimmed, and the mark only moves forward once ~15k more tokens (the step) could be trimmed. Between steps the conversation only grows at the end, so llama.cpp reuses what it already processed (its KV cache) and only reads the new messages. Each step changes the conversation from the first newly trimmed message on, and the server re-reads from there: that is why steps are large and rare. Old tool outputs are replaced, not summarized: [simple observation masking matches LLM summarization at half the cost](https://arxiv.org/abs/2508.21433).
110
+
111
+ **Hybrid models.** Some models (e.g. Qwen3.8 Flash Next; dense or MoE does not matter) mix a few classic attention layers with recurrent ones, which squeeze the whole conversation into a fixed-size state instead of keeping each token. The server cannot rewind that state to an arbitrary point: it can only restore a saved copy (a checkpoint) and re-read from there, and the only copy before the trimmed part is usually the end of the system prompt. So on these models each step re-reads almost the whole prompt. With 10k-token thoughts and 4k-token file reads, a 15k step is crossed every 2–3 turns: on a real 70k-token session that was 3 re-reads in 6 requests (~6 min at 185 tokens/s); with a 40k step, 1 re-read (~1 min) for a prompt at most 6k larger. A model is hybrid if its llama-server log shows `restored context checkpoint` lines.
112
+
113
+ **Thinking budget.** MAGI sends the limit with every request (`thinking_budget_tokens`), and llama.cpp applies it. It learns per model and per phase: 1.5 × the 95th percentile of the model's last 30 thinking lengths, rounded up to 1k. A cut is recorded as 1/1.5 of the limit it hit, so cuts never raise the limit: a model that often runs away is held, not chased. Replies that end close under the limit raise it; a model that thinks little lowers it. Limits: planning 4k–32k, acting 2k–4k (past ~4k between two tool calls it is overthinking). Until a model has 10 replies in a phase it uses 16k / 4k.
114
+
115
+ When the limit is reached llama.cpp does not abort the reply: it inserts the closing sentence and the end-of-thinking tag, and the model goes on to act. MAGI sends that sentence with every request too (`reasoning_budget_message`; `/magi budget message` turns it off and on). `MAGI.md` tells the model what to do after a cut (one small verifiable step, the open plan into `PLAN.md`), and the next turn gets a fresh budget.
116
+
117
+ **llama-server versions.** The per-request limit works whenever llama-server was started without `--reasoning-budget`, which is the default. If it was started with one, the server's limit wins: MAGI notices the model thinking well past its own limit and `/magi budget` says so. A llama-server too old for the closing sentence ends the thinking silently, and MAGI detects the cut by its length.
118
+
119
+ **Loop guard details.** A tool call counts as identical when both the tool and its arguments match. A file write or edit that copies a `<<pruned by MAGI…>>` note into a file is blocked too.
120
+
85
121
  ## The council
86
122
 
87
123
  `/magi <question>` asks three models in parallel, each with its own nature, then shows the votes and a majority verdict:
@@ -98,21 +134,36 @@ Each nature is a lens, not a specialty, so the council answers any question, not
98
134
 
99
135
  ## Configuration
100
136
 
101
- `~/.pi/agent/magi.json` (written by `/magi config`, `/magi-ui compact` and `/magi-ui config`, editable by hand):
137
+ `~/.pi/agent/magi.json` (written by `/magi config`, `/magi compact` and `/magi cost`, editable by hand):
102
138
 
103
139
  ```json
104
140
  {
105
141
  "MELCHIOR": { "model": "llama-swap/Qwen3.8 27B Q4_K_M - Thinking", "thinking": "low" },
106
142
  "ui": { "compact": false, "kwhPrice": 0.30, "currency": "EUR" },
107
- "loads": { "qwen3.8-27b": 41200 }
143
+ "loads": { "qwen3.8-27b": 41200 },
144
+ "totalWh": 1843.2,
145
+ "hygiene": { "enabled": true, "keepThinkingTurns": 3, "keepToolResults": 5, "stepTokens": 15000, "minPruneChars": 600 },
146
+ "thinkingBudget": { "mode": "auto", "planning": 16384, "acting": 4096, "message": true, "learned": { "qwen3.8-27b": { "acting": [812, 430, 2211] } } }
108
147
  }
109
148
  ```
110
149
 
111
150
  - per MAGI: `model` (unset = current session model) and optional `thinking` level;
112
151
  - `ui.compact`: start with the compact side panel;
113
- - `ui.kwhPrice` and `ui.currency` (`EUR` or `USD`): the COST row multiplies the GPU energy used in the session by this price;
152
+ - `ui.kwhPrice` and `ui.currency` (`EUR` or `USD`): the COST row multiplies the GPU energy by this price, showing the running total of every session with the current one in brackets;
153
+ - `totalWh`: written by the theme, GPU energy summed over every session (delete the key to reset the COST total);
154
+ - `hygiene` and `thinkingBudget` are set with `/magi hygiene` and `/magi budget` (`hygiene.minPruneChars` by hand only); `thinkingBudget.message` sends the closing message with every request (default `true`); `thinkingBudget.learned` is written by the theme (recent thinking lengths per model and phase);
114
155
  - `loads`: written by the theme, how long each llama-swap model took to load last time (paces the angel attack; 60s when unknown).
115
156
 
157
+ ## Release
158
+
159
+ `.github/workflows/publish.yml` publishes to npm when a `v*` tag is pushed, and refuses if the tag does not match `package.json`:
160
+
161
+ ```
162
+ npm version patch && git push --follow-tags
163
+ ```
164
+
165
+ No token: npmjs is configured to trust this repository's `publish.yml` (npm trusted publishing, OIDC), which also signs the provenance.
166
+
116
167
  ## llama-swap
117
168
 
118
169
  When the session model uses the `llama-swap` provider, the side panel:
@@ -17,22 +17,41 @@
17
17
  * - while a model loads an angel attacks the MAGI: red spreads through BALTHASAR, MELCHIOR and CASPAR at the pace of
18
18
  * the model's last load, a corner of CASPAR holds out blinking; once loaded, blue takes the MAGI back from that corner
19
19
  * - llama-swap telemetry: VRAM, GPU load/temp/power, energy used, RAM, server-side tok/s, prompt tok/s, cache hits
20
- * - /magi config → assign a model to each MAGI; /magi-ui compact|status → smaller panel, llama-swap report
20
+ * - /magi config → assign a model to each MAGI; /magi compact|status → smaller panel, llama-swap report
21
21
  *
22
22
  * Fan art: the MAGI and their screen come from Neon Genesis Evangelion, all rights reserved to khara, Inc.
23
23
  * Use with the theme ../../themes/magi.json
24
24
  */
25
25
 
26
26
  import { execFile } from "node:child_process";
27
- import { readFileSync, writeFileSync } from "node:fs";
27
+ import { existsSync, readFileSync, writeFileSync } from "node:fs";
28
28
  import { homedir } from "node:os";
29
29
  import { join } from "node:path";
30
30
  import { promisify } from "node:util";
31
31
  import type { AssistantMessage, Model } from "@earendil-works/pi-ai";
32
32
  import { completeSimple } from "@earendil-works/pi-ai";
33
- import type { ExtensionAPI, ExtensionContext, Theme, ThemeColor } from "@earendil-works/pi-coding-agent";
33
+ import type { ExtensionAPI, ExtensionCommandContext, ExtensionContext, Theme, ThemeColor } from "@earendil-works/pi-coding-agent";
34
34
  import type { Component, OverlayHandle, TUI } from "@earendil-works/pi-tui";
35
- import { HStack, matchesKey, truncateToWidth, visibleWidth, wrapTextWithAnsi } from "@earendil-works/pi-tui";
35
+ import { HStack, matchesKey, sliceByColumn, truncateToWidth, visibleWidth, wrapTextWithAnsi } from "@earendil-works/pi-tui";
36
+ import {
37
+ BUDGET_DEFAULTS,
38
+ BUDGET_MESSAGE,
39
+ BUDGET_WINDOW,
40
+ HYGIENE_DEFAULTS,
41
+ MAGI_MD,
42
+ PRUNED_MARK,
43
+ learnedBudget,
44
+ budgetSample,
45
+ pruneContext,
46
+ requestPhase,
47
+ thinkingTokens,
48
+ budgetVerdict,
49
+ thinkingWasCut,
50
+ type BudgetMode,
51
+ type BudgetPhase,
52
+ type HygieneOptions,
53
+ type HygieneStats,
54
+ } from "./local-models.ts";
36
55
 
37
56
  /* ────────────────────────────────────────────────────────────── art ── */
38
57
 
@@ -317,6 +336,8 @@ const state = {
317
336
  rebornAt: 0,
318
337
  hasSmartCompact: false,
319
338
  lastCouncil: undefined as { verdict: Vote | null; tally: number; question: string } | undefined,
339
+ pruned: 0, // tokens the context hygiene keeps away from the model
340
+ thinkCuts: 0, // replies whose thinking llama.cpp cut at the budget
320
341
  };
321
342
 
322
343
  /** Panel preferences, persisted under "ui" in ~/.pi/agent/magi.json. */
@@ -324,7 +345,7 @@ type Currency = "EUR" | "USD";
324
345
 
325
346
  const ui = {
326
347
  compact: false,
327
- kwhPrice: undefined as number | undefined, // price per kWh, for the COST row (/magi-ui config)
348
+ kwhPrice: undefined as number | undefined, // price per kWh, for the COST row (/magi cost)
328
349
  currency: "EUR" as Currency,
329
350
  };
330
351
 
@@ -340,10 +361,15 @@ function setPhase(p: Phase): void {
340
361
  }
341
362
 
342
363
  const ANIM_STEP_MS = 220;
364
+ /** How long one Sephirah stays lit in the footer: its name and meaning are text, they need time to be read. */
365
+ const SEPHIRAH_STEP_MS = 1400;
343
366
  const FAIL_FLASH_MS = 2500;
344
367
  const REBIRTH_MS = 6000;
345
368
  const DONE_TITLE_AFTER_MS = 30_000;
346
369
  const SIXTH_SEAL_PERCENT = (6 / 7) * 100;
370
+ /** Footer marquee: one cell every FOOTER_SCROLL_MS, with this gap between the loop's end and its start. */
371
+ const FOOTER_SCROLL_MS = 260;
372
+ const FOOTER_SCROLL_GAP = " · ";
347
373
 
348
374
  /** Sync ratio: share of tool calls that succeeded (null before the first tool). */
349
375
  function syncPercent(): number | null {
@@ -416,7 +442,9 @@ interface TokenStats {
416
442
 
417
443
  const tokens: TokenStats = { input: 0, output: 0, cacheRead: 0, cost: 0 };
418
444
 
419
- function addUsage(m: AssistantMessage): void {
445
+ /** Session totals from one assistant reply: tokens, cost, and whether its thinking was cut at the budget. */
446
+ function countAssistant(m: AssistantMessage): void {
447
+ if (thinkingWasCut(m as any)) state.thinkCuts++;
420
448
  tokens.input += m.usage?.input ?? 0;
421
449
  tokens.output += m.usage?.output ?? 0;
422
450
  tokens.cacheRead += m.usage?.cacheRead ?? 0;
@@ -427,8 +455,9 @@ function addUsage(m: AssistantMessage): void {
427
455
  function recountSession(ctx: ExtensionContext): void {
428
456
  Object.assign(tokens, { input: 0, output: 0, cacheRead: 0, cost: 0 });
429
457
  state.lastCouncil = undefined;
458
+ state.thinkCuts = 0;
430
459
  for (const entry of ctx.sessionManager.getBranch()) {
431
- if (entry.type === "message" && entry.message.role === "assistant") addUsage(entry.message as AssistantMessage);
460
+ if (entry.type === "message" && entry.message.role === "assistant") countAssistant(entry.message as AssistantMessage);
432
461
  if (entry.type === "custom" && entry.customType === "magi-verdict") {
433
462
  const d = entry.data as { verdict?: Vote | null; tally?: number; question?: string } | undefined;
434
463
  if (d && Array.isArray((d as any).opinions)) state.lastCouncil = { verdict: d.verdict ?? null, tally: d.tally ?? 0, question: d.question ?? "" };
@@ -710,6 +739,8 @@ const swap = {
710
739
  cacheTokens: 0,
711
740
  inputTokens: 0,
712
741
  energyWh: 0, // GPU energy since the session started
742
+ diskWh: 0, // GPU energy of all sessions, as last read from/written to magi.json
743
+ savedWh: 0, // the part of energyWh already added to diskWh
713
744
  lastSampleAt: 0,
714
745
  lastWatts: 0,
715
746
  };
@@ -760,6 +791,35 @@ function sampleEnergy(now = Date.now()): void {
760
791
  }
761
792
  swap.lastSampleAt = now;
762
793
  swap.lastWatts = watts;
794
+ persistEnergy();
795
+ }
796
+
797
+ /** GPU energy of every session so far: what is on disk plus this session's not-yet-written part. */
798
+ function totalWh(): number {
799
+ return swap.diskWh + swap.energyWh - swap.savedWh;
800
+ }
801
+
802
+ const ENERGY_SAVE_MS = 60_000;
803
+ let lastPersistAt = 0;
804
+
805
+ /**
806
+ * Adds this session's new energy to the running total in magi.json, so the COST row survives restarts.
807
+ * ponytail: writes at most once a minute, and adds a delta so parallel sessions do not overwrite each other.
808
+ * A crash therefore loses up to a minute of energy, which is what a kill loses anyway.
809
+ */
810
+ function persistEnergy(): void {
811
+ const delta = swap.energyWh - swap.savedWh;
812
+ if (delta <= 0 || Date.now() - lastPersistAt < ENERGY_SAVE_MS) return;
813
+ lastPersistAt = Date.now();
814
+ const cfg = loadMagiConfig();
815
+ cfg.totalWh = (cfg.totalWh ?? 0) + delta;
816
+ try {
817
+ saveMagiConfig(cfg);
818
+ } catch {
819
+ return; // disk unavailable: keep the delta and retry on the next sample
820
+ }
821
+ swap.savedWh = swap.energyWh;
822
+ swap.diskWh = cfg.totalWh;
763
823
  }
764
824
 
765
825
  async function refreshSwapMetrics(): Promise<void> {
@@ -995,7 +1055,7 @@ async function prewarmPrefix(cwd: string): Promise<void> {
995
1055
  }
996
1056
  }
997
1057
 
998
- /* ── /magi-ui status: a report built from the last requests llama-swap recorded ── */
1058
+ /* ── /magi status: a report built from the last requests llama-swap recorded ── */
999
1059
 
1000
1060
  interface ActivityRow {
1001
1061
  timestamp: string;
@@ -1280,6 +1340,11 @@ class MagiPanel implements Component {
1280
1340
  if (!compact && usage?.contextWindow) {
1281
1341
  out.push(this.field("CONTEXT", `${usage.tokens == null ? "?" : fmtTokens(usage.tokens)} / ${fmtTokens(usage.contextWindow)}`, inner, "muted"));
1282
1342
  }
1343
+ // HYGIENE: tokens pruned from what the model sees, ✂ = thinking cut at the budget
1344
+ if (!compact && (state.pruned || state.thinkCuts)) {
1345
+ const cuts = state.thinkCuts ? ` ✂${state.thinkCuts}` : "";
1346
+ out.push(this.field("HYGIENE", `-${fmtTokens(state.pruned)}${cuts}`, inner, state.thinkCuts ? "warning" : "muted"));
1347
+ }
1283
1348
  return out;
1284
1349
  }
1285
1350
 
@@ -1311,11 +1376,12 @@ class MagiPanel implements Component {
1311
1376
  return out;
1312
1377
  }
1313
1378
 
1314
- /** COST: GPU energy used in the session × price per kWh set with /magi-ui config. */
1379
+ /** COST: GPU energy × price per kWh set with /magi cost — all sessions, with this one in brackets. */
1315
1380
  private costRow(inner: number): string {
1316
1381
  if (!swap.gpus.length) return this.field("COST", "—", inner, "muted");
1317
- if (ui.kwhPrice === undefined) return this.field("COST", "→ /magi-ui config", inner, "dim");
1318
- return this.field("COST", fmtMoney((swap.energyWh / 1000) * ui.kwhPrice), inner, "warning");
1382
+ if (ui.kwhPrice === undefined) return this.field("COST", "→ /magi cost", inner, "dim");
1383
+ const price = (wh: number) => fmtMoney((wh / 1000) * ui.kwhPrice!);
1384
+ return this.field("COST", price(totalWh()) + this.theme.fg("dim", ` (ses ${price(swap.energyWh)})`), inner, "warning");
1319
1385
  }
1320
1386
 
1321
1387
  private swapRows(inner: number, compact: boolean): string[] {
@@ -1357,7 +1423,7 @@ class MagiPanel implements Component {
1357
1423
  );
1358
1424
  }
1359
1425
  if (swap.ramTotal) out.push(this.field("RAM", `${gib(swap.ramUsed).toFixed(1)} / ${Math.round(gib(swap.ramTotal))}G`, inner, "muted"));
1360
- if (swap.gpus.length) out.push(this.field("ENERGY", `${(swap.energyWh / 1000).toFixed(3)} kWh`, inner, "muted"));
1426
+ if (swap.gpus.length) out.push(this.field("ENERGY", `${(totalWh() / 1000).toFixed(3)} kWh`, inner, "muted"));
1361
1427
  out.push(this.costRow(inner));
1362
1428
  return out;
1363
1429
  }
@@ -1466,9 +1532,49 @@ interface MagiUnitConfig {
1466
1532
  type MagiConfig = Partial<Record<MagiUnit, MagiUnitConfig>> & {
1467
1533
  ui?: { compact?: boolean; kwhPrice?: number; currency?: Currency };
1468
1534
  loads?: Record<string, number>; // real model id → ms its last load took
1535
+ totalWh?: number; // GPU energy summed over every session, for the COST row
1536
+ hygiene?: Partial<HygieneOptions> & { enabled?: boolean }; // context pruning for local models, see local-models.ts
1537
+ thinkingBudget?: {
1538
+ mode?: BudgetMode; // auto (learned per model) · fixed · off, set with /magi budget
1539
+ planning?: number; // fixed budgets
1540
+ acting?: number;
1541
+ message?: boolean; // send BUDGET_MESSAGE as reasoning_budget_message (default on), /magi budget message
1542
+ learned?: Record<string, Partial<Record<BudgetPhase, number[]>>>; // written by the theme: recent thinking lengths per model
1543
+ };
1469
1544
  };
1470
1545
 
1471
1546
  const MAGI_CONFIG_PATH = join(homedir(), ".pi", "agent", "magi.json");
1547
+ /** /magi arguments that manage the theme instead of asking the council. */
1548
+ const UI_ARGS = /^(on|off|panel|compact|status|cost)$|^(hygiene|budget)(\s|$)/i; // anything else is a question
1549
+ /** /magi arguments offered by autocomplete: the full argument, and what it does. */
1550
+ const MAGI_ARGS: [string, string][] = [
1551
+ ["review", "the council reviews your pending changes before you commit"],
1552
+ ["config", "pick a model for each MAGI"],
1553
+ ["mecha", "MECHA SELECT: pick the llama-swap model to activate"],
1554
+ ["status", "llama-swap report: speed, tokens, cache hits, errors per model"],
1555
+ ["panel", "hide/show the side panel"],
1556
+ ["compact", "toggle the compact side panel"],
1557
+ ["cost", "electricity price and currency for the COST row"],
1558
+ ["on", "enable the MAGI chrome"],
1559
+ ["off", "disable the MAGI chrome"],
1560
+ ["hygiene", "show how much context was pruned"],
1561
+ ["hygiene on", "enable context pruning"],
1562
+ ["hygiene off", "disable context pruning"],
1563
+ ["hygiene step 40k", "prune less often: fewer prompt re-reads on long tasks (default 15k)"],
1564
+ ["hygiene 3 5", "recent turns that keep their thinking, tool results kept whole"],
1565
+ ["budget", "show the thinking budget of the current model"],
1566
+ ["budget auto", "learn the budget per model (default)"],
1567
+ ["budget off", "no budget: llama-server decides"],
1568
+ ["budget reset", "forget what was learned for the current model"],
1569
+ ["budget message", "turn the closing message after a cut off/on"],
1570
+ ["budget 16k 4k", "fixed budget: planning, acting"],
1571
+ ];
1572
+ function argCompletions(table: [string, string][], prefix: string) {
1573
+ const p = prefix.trimStart().toLowerCase();
1574
+ const items = table.filter(([value]) => value.startsWith(p)).map(([value, description]) => ({ value, label: value, description }));
1575
+ return items.length ? items : null;
1576
+ }
1577
+ const LOOP_REPEATS = 3; // the same tool call this many times in a row is blocked
1472
1578
 
1473
1579
  function loadMagiConfig(): MagiConfig {
1474
1580
  try {
@@ -1684,7 +1790,7 @@ function buildDeliberationView(
1684
1790
  };
1685
1791
  }
1686
1792
 
1687
- /** A read-only boxed report (used by /magi-ui status); any key closes it. */
1793
+ /** A read-only boxed report (used by /magi status); any key closes it. */
1688
1794
  function buildReportView(theme: Theme, lines: string[], close: () => void) {
1689
1795
  return {
1690
1796
  render(width: number): string[] {
@@ -1807,6 +1913,7 @@ async function configureMagi(ctx: ExtensionContext): Promise<void> {
1807
1913
  function footerLeft(th: Theme, now = Date.now()): string {
1808
1914
  const dim = (s: string) => th.fg("dim", s);
1809
1915
  const step = Math.floor(now / ANIM_STEP_MS);
1916
+ const slowStep = Math.floor(now / SEPHIRAH_STEP_MS);
1810
1917
  const lights = (fn: (i: number) => NodeLight) => SEPHIROT.map((_, i) => fn(i));
1811
1918
 
1812
1919
  if (state.compacting) {
@@ -1860,7 +1967,7 @@ function footerLeft(th: Theme, now = Date.now()): string {
1860
1967
  );
1861
1968
  }
1862
1969
  if (state.phase === "thinking") {
1863
- const cur = Math.floor(step / 2) % 3; // Keter, Chokmah, Binah
1970
+ const cur = slowStep % 3; // Keter, Chokmah, Binah
1864
1971
  const s = SEPHIROT[cur]!;
1865
1972
  return (
1866
1973
  th.fg("accent", "◆ ") +
@@ -1870,7 +1977,7 @@ function footerLeft(th: Theme, now = Date.now()): string {
1870
1977
  );
1871
1978
  }
1872
1979
  if (state.phase === "responding") {
1873
- const cur = 5 + (step % 5); // Tiferet → Malkuth
1980
+ const cur = 5 + (slowStep % 5); // Tiferet → Malkuth
1874
1981
  const s = SEPHIROT[cur]!;
1875
1982
  const tps = liveTps();
1876
1983
  return (
@@ -1908,9 +2015,21 @@ function footerLeft(th: Theme, now = Date.now()): string {
1908
2015
 
1909
2016
  /** Footer: the animation on the left, other extensions' statuses on the right. */
1910
2017
  function buildFooter(tui: TUI, theme: Theme, footerData: any) {
2018
+ let scrolling = false;
2019
+ let scrollOff = 0;
2020
+ let scrollPeriod = 1;
2021
+ let scrollLine = "";
1911
2022
  const timer = setInterval(() => {
1912
2023
  if (animating()) tui.requestRender();
1913
2024
  }, ANIM_STEP_MS);
2025
+ // The marquee keeps its own cadence: one cell per tick, never derived from the clock, so it
2026
+ // does not stutter when the line is re-rendered at some other pace (animation, streamed tokens)
2027
+ // nor jump when the left side changes length (sephirah name, tok/s) and with it the loop period.
2028
+ const scrollTimer = setInterval(() => {
2029
+ if (!scrolling) return;
2030
+ scrollOff = (scrollOff + 1) % scrollPeriod;
2031
+ tui.requestRender();
2032
+ }, FOOTER_SCROLL_MS);
1914
2033
 
1915
2034
  const unsub = footerData?.onBranchChange?.(() => tui.requestRender());
1916
2035
 
@@ -1919,14 +2038,30 @@ function buildFooter(tui: TUI, theme: Theme, footerData: any) {
1919
2038
  const dim = (s: string) => theme.fg("dim", s);
1920
2039
  const statuses = footerData?.getExtensionStatuses?.();
1921
2040
  const right = statuses ? [...statuses.values()].filter(Boolean).join(dim(" │ ")) : "";
1922
- const room = Math.max(12, width - visibleWidth(right) - 2);
1923
- const l = truncateToWidth(footerLeft(theme), room);
1924
- const gap = Math.max(1, width - visibleWidth(l) - visibleWidth(right));
1925
- return [truncateToWidth(l + " ".repeat(gap) + right, width)];
2041
+ const left = footerLeft(theme);
2042
+ const gap = width - visibleWidth(left) - visibleWidth(right);
2043
+ scrolling = gap < 1;
2044
+ if (!scrolling) {
2045
+ scrollLine = "";
2046
+ scrollOff = 0;
2047
+ return [left + " ".repeat(gap) + right];
2048
+ }
2049
+ // Too narrow to fit: scroll the whole line instead of cutting it off. The line is frozen
2050
+ // while it scrolls — live text changes width (seconds, tok/s, sephirah names) and every
2051
+ // change would shift it under the window. A fresh line is taken when it has the same
2052
+ // width, so nothing moves, otherwise at the end of the loop.
2053
+ const line = left + dim(FOOTER_SCROLL_GAP) + right + dim(FOOTER_SCROLL_GAP);
2054
+ const period = visibleWidth(line);
2055
+ if (!scrollLine || scrollOff === 0 || period === scrollPeriod) {
2056
+ scrollLine = line;
2057
+ scrollPeriod = period;
2058
+ }
2059
+ return [sliceByColumn(scrollLine + scrollLine, scrollOff % scrollPeriod, width, true)];
1926
2060
  },
1927
2061
  invalidate() {},
1928
2062
  dispose() {
1929
2063
  clearInterval(timer);
2064
+ clearInterval(scrollTimer);
1930
2065
  unsub?.();
1931
2066
  },
1932
2067
  };
@@ -1953,6 +2088,22 @@ export default function (pi: ExtensionAPI) {
1953
2088
  let panelHandle: OverlayHandle | undefined;
1954
2089
  let panel: MagiPanel | undefined;
1955
2090
  let titleDone = false;
2091
+ let hygiene = { enabled: true, ...HYGIENE_DEFAULTS };
2092
+ let hygieneStats: HygieneStats | undefined;
2093
+ let budgetMode: BudgetMode = "auto";
2094
+ let fixedBudget = { ...BUDGET_DEFAULTS };
2095
+ let budgetMessage = true;
2096
+ let learned: Record<string, Partial<Record<BudgetPhase, number[]>>> = {};
2097
+ let pendingBudget: { model: string; phase: BudgetPhase; tokens: number } | undefined; // the request in flight
2098
+ const budgetIgnored = new Set<string>(); // models whose llama-server thought past the budget: it sets its own
2099
+ const budgetFor = (model: string, phase: BudgetPhase) =>
2100
+ budgetMode === "fixed" ? fixedBudget[phase] : (learnedBudget(learned[model]?.[phase] ?? [], phase) ?? BUDGET_DEFAULTS[phase]);
2101
+ const persistBudget = () => {
2102
+ const cfg = loadMagiConfig();
2103
+ saveMagiConfig({ ...cfg, thinkingBudget: { mode: budgetMode, ...fixedBudget, message: budgetMessage, learned } });
2104
+ };
2105
+ let lastCall = ""; // loop guard: fingerprint of the previous tool call and how often it repeated
2106
+ let repeats = 0;
1956
2107
 
1957
2108
  const repaint = () => {
1958
2109
  panel?.invalidate();
@@ -2113,6 +2264,23 @@ export default function (pi: ExtensionAPI) {
2113
2264
 
2114
2265
  pi.on("session_start", async (event, ctx) => {
2115
2266
  liveCtx = ctx;
2267
+ const magiCfg = loadMagiConfig();
2268
+ hygiene = { enabled: true, ...HYGIENE_DEFAULTS, ...magiCfg.hygiene };
2269
+ const tb = magiCfg.thinkingBudget ?? {};
2270
+ budgetMode = tb.mode ?? "auto";
2271
+ fixedBudget = { planning: tb.planning ?? BUDGET_DEFAULTS.planning, acting: tb.acting ?? BUDGET_DEFAULTS.acting };
2272
+ budgetMessage = tb.message ?? true;
2273
+ learned = tb.learned ?? {};
2274
+ // the local-model rules live next to AGENTS.md, created once so the user can edit them
2275
+ const rules = join(ctx.cwd, "MAGI.md");
2276
+ if (!existsSync(rules)) {
2277
+ try {
2278
+ writeFileSync(rules, MAGI_MD);
2279
+ if (ctx.mode === "tui") ctx.ui.notify("MAGI.md created: rules for local models, appended to the system prompt", "info");
2280
+ } catch {
2281
+ // read-only directory: the built-in rules are used as they are
2282
+ }
2283
+ }
2116
2284
  void refreshUnit(ctx);
2117
2285
  recountSession(ctx);
2118
2286
  state.hasSmartCompact = pi.getCommands().some((c) => c.name.replace(/^\//, "") === "smart-compact");
@@ -2120,6 +2288,7 @@ export default function (pi: ExtensionAPI) {
2120
2288
  ui.compact = cfg.ui?.compact ?? false;
2121
2289
  ui.kwhPrice = cfg.ui?.kwhPrice;
2122
2290
  ui.currency = cfg.ui?.currency === "USD" ? "USD" : "EUR";
2291
+ swap.diskWh = cfg.totalWh ?? 0;
2123
2292
  if (ctx.mode !== "tui") return;
2124
2293
  applyChrome(ctx);
2125
2294
  // nothing is loaded at startup: a new session picks its MECHA unit, a resumed one shows whether its model is in VRAM
@@ -2164,6 +2333,12 @@ export default function (pi: ExtensionAPI) {
2164
2333
 
2165
2334
  pi.on("before_provider_request", (event, ctx) => {
2166
2335
  rememberPrefix(ctx.cwd, event.payload);
2336
+ // llama.cpp honours a per-request thinking budget only when llama-server runs without --reasoning-budget
2337
+ if (budgetMode === "off" || !swap.base || !ctx.model) return;
2338
+ const phase = requestPhase(event.payload);
2339
+ pendingBudget = { model: ctx.model.id, phase, tokens: budgetFor(ctx.model.id, phase) };
2340
+ const payload = { ...(event.payload as object), thinking_budget_tokens: pendingBudget.tokens };
2341
+ return budgetMessage ? { ...payload, reasoning_budget_message: `\n\n${BUDGET_MESSAGE}` } : payload;
2167
2342
  });
2168
2343
 
2169
2344
  pi.on("model_select", async (event, ctx) => {
@@ -2173,6 +2348,7 @@ export default function (pi: ExtensionAPI) {
2173
2348
  });
2174
2349
 
2175
2350
  pi.on("session_shutdown", async () => {
2351
+ persistEnergy();
2176
2352
  prewarm.abort?.abort();
2177
2353
  clearInterval(metricsTimer);
2178
2354
  metricsTimer = undefined;
@@ -2187,6 +2363,40 @@ export default function (pi: ExtensionAPI) {
2187
2363
 
2188
2364
  pi.on("agent_start", async () => {
2189
2365
  state.runStart = Date.now();
2366
+ lastCall = "";
2367
+ });
2368
+
2369
+ pi.on("before_agent_start", async (event, ctx) => {
2370
+ let rules = MAGI_MD;
2371
+ try {
2372
+ rules = readFileSync(join(ctx.cwd, "MAGI.md"), "utf8");
2373
+ } catch {
2374
+ // no MAGI.md (deleted, or unwritable cwd): the built-in rules
2375
+ }
2376
+ return rules.trim() ? { systemPrompt: `${event.systemPrompt}\n\n${rules.trim()}` } : undefined;
2377
+ });
2378
+
2379
+ // before every request: drop old thinking and tool outputs from what the model sees, never from the session
2380
+ pi.on("context", async (event) => {
2381
+ if (!hygiene.enabled) return;
2382
+ const pruned = pruneContext(event.messages as any, hygiene);
2383
+ hygieneStats = pruned.stats;
2384
+ state.pruned = pruned.stats.prunedTokens;
2385
+ return { messages: pruned.messages as any };
2386
+ });
2387
+
2388
+ // local models loop on the same call, and may copy a pruned placeholder into a file
2389
+ pi.on("tool_call", async (event) => {
2390
+ const args = JSON.stringify(event.input);
2391
+ if ((event.toolName === "write" || event.toolName === "edit") && args.includes(PRUNED_MARK)) {
2392
+ return { block: true, reason: `MAGI: this ${event.toolName} contains a "${PRUNED_MARK}" placeholder, not real content. Read the file and use the actual text.` };
2393
+ }
2394
+ const fingerprint = event.toolName + args;
2395
+ repeats = fingerprint === lastCall ? repeats + 1 : 1;
2396
+ lastCall = fingerprint;
2397
+ if (repeats >= LOOP_REPEATS) {
2398
+ return { block: true, reason: `MAGI: identical ${event.toolName} call ${repeats} times in a row. Repeating it will not change the outcome: change approach, or tell the user what blocks you.` };
2399
+ }
2190
2400
  });
2191
2401
 
2192
2402
  pi.on("turn_start", async (_event, ctx) => {
@@ -2230,7 +2440,19 @@ export default function (pi: ExtensionAPI) {
2230
2440
  pi.on("message_end", async (event) => {
2231
2441
  if (event.message.role !== "assistant") return;
2232
2442
  const m = event.message as AssistantMessage;
2233
- addUsage(m);
2443
+ countAssistant(m);
2444
+ // learn how long this model thinks in this phase; a cut must not raise the budget
2445
+ const thought = thinkingTokens(m as any);
2446
+ if (pendingBudget && thought > 0 && m.stopReason !== "aborted" && m.stopReason !== "error") {
2447
+ const verdict = budgetVerdict(m as any, pendingBudget.tokens);
2448
+ if (verdict === "ignored") budgetIgnored.add(pendingBudget.model);
2449
+ if (verdict === "cut" && !thinkingWasCut(m as any)) state.thinkCuts++; // silent cut: no budget message on the server
2450
+ const samples = ((learned[pendingBudget.model] ??= {})[pendingBudget.phase] ??= []);
2451
+ samples.push(budgetSample(thought, pendingBudget.tokens, verdict === "cut"));
2452
+ samples.splice(0, samples.length - BUDGET_WINDOW);
2453
+ persistBudget();
2454
+ }
2455
+ pendingBudget = undefined;
2234
2456
  if (perf.start) {
2235
2457
  const end = Date.now();
2236
2458
  perf.lastMs = end - perf.start;
@@ -2366,9 +2588,11 @@ export default function (pi: ExtensionAPI) {
2366
2588
  }
2367
2589
 
2368
2590
  pi.registerCommand("magi", {
2369
- description: "Ask the three MAGI (pragmatist, guardian, visionary); /magi review [focus] judges the git diff; /magi config assigns models; /magi mecha picks the model",
2591
+ description: "Ask the three MAGI, or manage them: review [focus] · config · mecha · status · panel · compact · cost · on · off · hygiene · budget (type a space to see them all)",
2592
+ getArgumentCompletions: (prefix) => argCompletions(MAGI_ARGS, prefix),
2370
2593
  handler: async (args, ctx) => {
2371
2594
  const arg = args.trim();
2595
+ if (UI_ARGS.test(arg)) return manageUi(arg, ctx);
2372
2596
  if (arg === "config") return configureMagi(ctx);
2373
2597
  if (arg === "mecha") {
2374
2598
  if (ctx.mode !== "tui" || !swap.base) return ctx.ui.notify("MECHA SELECT needs the TUI and a llama-swap model", "error");
@@ -2406,87 +2630,145 @@ export default function (pi: ExtensionAPI) {
2406
2630
  },
2407
2631
  });
2408
2632
 
2409
- pi.registerCommand("magi-ui", {
2410
- description: "MAGI chrome: enable the theme, or manage it (on|off|panel|compact|status|config)",
2411
- handler: async (args, ctx) => {
2412
- liveCtx = ctx;
2413
- const arg = args.trim().toLowerCase();
2414
-
2415
- if (arg === "panel") {
2416
- panelEnabled = !panelEnabled;
2417
- if (panelEnabled) showPanel(ctx.ui.theme);
2418
- else hidePanel();
2419
- ctx.ui.notify(`Side panel ${panelEnabled ? "enabled" : "disabled"}`, "info");
2633
+ /** /magi on|off|panel|compact|status|cost|hygiene|budget: the theme, its panel and the local-model settings. */
2634
+ async function manageUi(args: string, ctx: ExtensionCommandContext) {
2635
+ liveCtx = ctx;
2636
+ const arg = args.trim().toLowerCase();
2637
+
2638
+ if (arg === "hygiene" || arg.startsWith("hygiene ")) {
2639
+ const sub = arg.slice("hygiene".length).trim();
2640
+ const keep = /^(\d+)\s+(\d+)$/.exec(sub);
2641
+ const step = /^step\s+(\d+)(k?)$/.exec(sub);
2642
+ if (sub === "on" || sub === "off" || keep || step) {
2643
+ if (keep) Object.assign(hygiene, { enabled: true, keepThinkingTurns: Number(keep[1]), keepToolResults: Number(keep[2]) });
2644
+ else if (step) hygiene.stepTokens = Math.max(1000, Number(step[1]) * (step[2] ? 1000 : 1));
2645
+ else hygiene.enabled = sub === "on";
2646
+ const cfg = loadMagiConfig();
2647
+ const { enabled, keepThinkingTurns, keepToolResults, stepTokens } = hygiene;
2648
+ saveMagiConfig({ ...cfg, hygiene: { ...cfg.hygiene, enabled, keepThinkingTurns, keepToolResults, stepTokens } });
2649
+ } else if (sub) {
2650
+ ctx.ui.notify("Usage: /magi hygiene [on|off|step <tokens>|<thinking turns kept> <tool results kept>], e.g. step 40k, 3 5", "warning");
2420
2651
  return;
2421
2652
  }
2422
- if (arg === "compact") {
2423
- ui.compact = !ui.compact;
2424
- const cfg = loadMagiConfig();
2425
- saveMagiConfig({ ...cfg, ui: { ...cfg.ui, compact: ui.compact } });
2426
- repaint();
2427
- ctx.ui.notify(`Side panel ${ui.compact ? "compact" : "detailed"}`, "info");
2653
+ const s = hygieneStats;
2654
+ ctx.ui.notify(
2655
+ (!hygiene.enabled
2656
+ ? "Context hygiene off: /magi hygiene on"
2657
+ : s
2658
+ ? `Context hygiene: ${fmtTokens(s.prunedTokens)} tokens pruned in the first ${s.watermark}/${s.messages} messages, ${fmtTokens(s.pendingTokens)} waiting for the next step (every ${fmtTokens(hygiene.stepTokens)})`
2659
+ : "Context hygiene on: nothing sent to the model yet") +
2660
+ ` · keeps thinking of the last ${hygiene.keepThinkingTurns} turns, the last ${hygiene.keepToolResults} tool results` +
2661
+ (state.thinkCuts ? ` · thinking cut at the budget ${state.thinkCuts}×` : ""),
2662
+ "info",
2663
+ );
2664
+ return;
2665
+ }
2666
+ if (arg === "budget" || arg.startsWith("budget ")) {
2667
+ const sub = arg.slice("budget".length).trim();
2668
+ const model = ctx.model?.id ?? "";
2669
+ const tokens = (t: string) => Math.round(Number(t.replace(/k$/, "")) * (t.endsWith("k") ? 1024 : 1));
2670
+ const fixed = /^(\d+k?)\s+(\d+k?)$/.exec(sub);
2671
+ if (sub === "auto" || sub === "off") budgetMode = sub;
2672
+ else if (sub === "reset") delete learned[model];
2673
+ else if (sub === "message") budgetMessage = !budgetMessage;
2674
+ else if (fixed) {
2675
+ budgetMode = "fixed";
2676
+ fixedBudget = { planning: tokens(fixed[1]!), acting: tokens(fixed[2]!) };
2677
+ } else if (sub) {
2678
+ ctx.ui.notify("Usage: /magi budget [auto|off|reset|message|<planning> <acting>], e.g. 16k 4k", "warning");
2428
2679
  return;
2429
2680
  }
2430
- if (arg === "config") {
2431
- const currency = await ctx.ui.select(`Currency for COST (current: ${ui.currency})`, ["EUR", "USD"]);
2432
- if (!currency) return;
2433
- const current = ui.kwhPrice !== undefined ? ` (current: ${ui.kwhPrice})` : "";
2434
- const raw = (await ctx.ui.input(`Electricity price per kWh in ${currency}${current}:`, "0.30"))?.trim();
2435
- if (raw === undefined) return;
2436
- const price = raw === "" && ui.kwhPrice !== undefined ? ui.kwhPrice : Number(raw.replace(",", "."));
2437
- if (raw === "" && ui.kwhPrice === undefined) {
2438
- ctx.ui.notify("No price entered: COST unchanged", "warning");
2439
- return;
2440
- }
2441
- if (!Number.isFinite(price) || price < 0) {
2442
- ctx.ui.notify(`Invalid price: "${raw}"`, "error");
2443
- return;
2444
- }
2445
- ui.currency = currency === "USD" ? "USD" : "EUR";
2446
- ui.kwhPrice = price;
2447
- const cfg = loadMagiConfig();
2448
- saveMagiConfig({ ...cfg, ui: { ...cfg.ui, currency: ui.currency, kwhPrice: price } });
2449
- repaint();
2450
- ctx.ui.notify(`COST: ${fmtMoney(price)} per kWh`, "info");
2681
+ if (sub) persistBudget();
2682
+ const phase = (p: BudgetPhase) => {
2683
+ const n = learned[model]?.[p]?.length ?? 0;
2684
+ const how = budgetMode === "fixed" ? "fixed" : learnedBudget(learned[model]?.[p] ?? [], p) ? `learned from ${n}` : `default, learning ${n}/10`;
2685
+ return `${p} ${budgetFor(model, p) / 1024}k (${how})`;
2686
+ };
2687
+ ctx.ui.notify(
2688
+ budgetMode === "off"
2689
+ ? "Thinking budget off: llama-server decides. /magi budget auto"
2690
+ : `Thinking budget ${budgetMode.toUpperCase()} · ${model || "no model"}: ${phase("planning")} · ${phase("acting")}` +
2691
+ ` · closing message ${budgetMessage ? "on" : "off"}` +
2692
+ (swap.base ? "" : " · applies to llama-swap models only") +
2693
+ (budgetIgnored.has(model) ? " · ⚠ llama-server thinks past it: it was started with its own --reasoning-budget (or is too old), which wins" : ""),
2694
+ "info",
2695
+ );
2696
+ return;
2697
+ }
2698
+ if (arg === "panel") {
2699
+ panelEnabled = !panelEnabled;
2700
+ if (panelEnabled) showPanel(ctx.ui.theme);
2701
+ else hidePanel();
2702
+ ctx.ui.notify(`Side panel ${panelEnabled ? "enabled" : "disabled"}`, "info");
2703
+ return;
2704
+ }
2705
+ if (arg === "compact") {
2706
+ ui.compact = !ui.compact;
2707
+ const cfg = loadMagiConfig();
2708
+ saveMagiConfig({ ...cfg, ui: { ...cfg.ui, compact: ui.compact } });
2709
+ repaint();
2710
+ ctx.ui.notify(`Side panel ${ui.compact ? "compact" : "detailed"}`, "info");
2711
+ return;
2712
+ }
2713
+ if (arg === "cost") {
2714
+ const currency = await ctx.ui.select(`Currency for COST (current: ${ui.currency})`, ["EUR", "USD"]);
2715
+ if (!currency) return;
2716
+ const current = ui.kwhPrice !== undefined ? ` (current: ${ui.kwhPrice})` : "";
2717
+ const raw = (await ctx.ui.input(`Electricity price per kWh in ${currency}${current}:`, "0.30"))?.trim();
2718
+ if (raw === undefined) return;
2719
+ const price = raw === "" && ui.kwhPrice !== undefined ? ui.kwhPrice : Number(raw.replace(",", "."));
2720
+ if (raw === "" && ui.kwhPrice === undefined) {
2721
+ ctx.ui.notify("No price entered: COST unchanged", "warning");
2451
2722
  return;
2452
2723
  }
2453
- if (arg === "status") {
2454
- if (!swap.base) {
2455
- ctx.ui.notify("/magi-ui status needs a llama-swap session model", "warning");
2456
- return;
2457
- }
2458
- let report: { data?: ActivityRow[]; total?: number };
2459
- try {
2460
- report = (await (await swapGet(`/api/metrics/activity?limit=${ACTIVITY_REPORT_ROWS}`, 15_000)).json()) as typeof report;
2461
- } catch (err) {
2462
- ctx.ui.notify(`llama-swap status failed: ${err instanceof Error ? err.message : String(err)}`, "error");
2463
- return;
2464
- }
2465
- const rows = report.data ?? [];
2466
- if (!rows.length) {
2467
- ctx.ui.notify("llama-swap has no recorded requests yet", "info");
2468
- return;
2469
- }
2470
- await ctx.ui.custom<void>((_tui, theme, _keys, done) =>
2471
- buildReportView(theme, activityReport(theme, rows, report.total ?? rows.length), () => done(undefined)),
2472
- );
2724
+ if (!Number.isFinite(price) || price < 0) {
2725
+ ctx.ui.notify(`Invalid price: "${raw}"`, "error");
2726
+ return;
2727
+ }
2728
+ ui.currency = currency === "USD" ? "USD" : "EUR";
2729
+ ui.kwhPrice = price;
2730
+ const cfg = loadMagiConfig();
2731
+ saveMagiConfig({ ...cfg, ui: { ...cfg.ui, currency: ui.currency, kwhPrice: price } });
2732
+ repaint();
2733
+ ctx.ui.notify(`COST: ${fmtMoney(price)} per kWh`, "info");
2734
+ return;
2735
+ }
2736
+ if (arg === "status") {
2737
+ if (!swap.base) {
2738
+ ctx.ui.notify("/magi status needs a llama-swap session model", "warning");
2473
2739
  return;
2474
2740
  }
2475
- if (arg === "off" || arg === "on") {
2476
- chrome = arg === "on";
2477
- applyChrome(ctx);
2478
- ctx.ui.notify(`MAGI chrome ${chrome ? "enabled" : "disabled"}`, "info");
2741
+ let report: { data?: ActivityRow[]; total?: number };
2742
+ try {
2743
+ report = (await (await swapGet(`/api/metrics/activity?limit=${ACTIVITY_REPORT_ROWS}`, 15_000)).json()) as typeof report;
2744
+ } catch (err) {
2745
+ ctx.ui.notify(`llama-swap status failed: ${err instanceof Error ? err.message : String(err)}`, "error");
2479
2746
  return;
2480
2747
  }
2481
-
2482
- const res = ctx.ui.setTheme("magi");
2483
- if (!res.success) {
2484
- ctx.ui.notify(`Theme magi not found: ${res.error}`, "error");
2748
+ const rows = report.data ?? [];
2749
+ if (!rows.length) {
2750
+ ctx.ui.notify("llama-swap has no recorded requests yet", "info");
2485
2751
  return;
2486
2752
  }
2487
- chrome = true;
2753
+ await ctx.ui.custom<void>((_tui, theme, _keys, done) =>
2754
+ buildReportView(theme, activityReport(theme, rows, report.total ?? rows.length), () => done(undefined)),
2755
+ );
2756
+ return;
2757
+ }
2758
+ if (arg === "off") {
2759
+ chrome = false;
2488
2760
  applyChrome(ctx);
2489
- ctx.ui.notify("MAGI online — /magi-ui panel|compact|status|off", "info");
2490
- },
2491
- });
2761
+ ctx.ui.notify("MAGI chrome disabled", "info");
2762
+ return;
2763
+ }
2764
+
2765
+ const res = ctx.ui.setTheme("magi");
2766
+ if (!res.success) {
2767
+ ctx.ui.notify(`Theme magi not found: ${res.error}`, "error");
2768
+ return;
2769
+ }
2770
+ chrome = true;
2771
+ applyChrome(ctx);
2772
+ ctx.ui.notify("MAGI online — /magi panel|compact|status|off", "info");
2773
+ }
2492
2774
  }
@@ -0,0 +1,240 @@
1
+ /**
2
+ * Local-model helpers for MAGI: context hygiene, loop guard, and the MAGI.md rules template.
3
+ *
4
+ * Context hygiene rewrites the messages pi sends to the model (never the saved session):
5
+ * - old thinking blocks are dropped: Qwen templates resend every reasoning block of an agent run,
6
+ * one 18k-token think stays in the context until compaction
7
+ * - old tool outputs become a one-line marker (observation masking, Lindenbauer et al. 2025:
8
+ * as good as LLM summarization, at half the cost)
9
+ * - old write/edit payloads become a marker too: the file on disk is the source of truth
10
+ * Pruning advances in steps behind a watermark, so between steps the prompt only grows at the end
11
+ * and llama.cpp keeps reusing its KV cache.
12
+ */
13
+
14
+ export type Msg = { role: string; content?: any; [key: string]: any };
15
+
16
+ export interface HygieneOptions {
17
+ keepThinkingTurns: number; // newest assistant messages that keep their thinking
18
+ keepToolResults: number; // newest tool results kept whole
19
+ stepTokens: number; // the watermark moves each time the prunable total crosses another multiple of this
20
+ minPruneChars: number; // smaller outputs/payloads are left alone
21
+ }
22
+
23
+ export interface HygieneStats {
24
+ messages: number;
25
+ watermark: number; // messages before this index are pruned
26
+ prunedTokens: number;
27
+ pendingTokens: number; // prunable, waiting for the next step
28
+ stepTokens: number;
29
+ }
30
+
31
+ export const HYGIENE_DEFAULTS: HygieneOptions = { keepThinkingTurns: 3, keepToolResults: 5, stepTokens: 15_000, minPruneChars: 600 };
32
+
33
+ // ponytail: chars/token measured on Qwen3.x pi sessions (2.9–3.2); a tokenizer call would be exact but costs a request
34
+ const CHARS_PER_TOKEN = 3;
35
+ const IMAGE_CHARS = 4800;
36
+
37
+ /** Marks every placeholder, so a model that copies one into a real write/edit gets blocked (see loop guard). */
38
+ export const PRUNED_MARK = "<<pruned by MAGI";
39
+
40
+ const lines = (s: string) => s.split("\n").length;
41
+ const clip = (s: string, n = 80) => (s.length > n ? s.slice(0, n - 1) + "…" : s);
42
+
43
+ function callTarget(call: any): string {
44
+ const a = call?.arguments ?? {};
45
+ return clip(String(a.path ?? a.command ?? a.pattern ?? a.query ?? a.url ?? ""));
46
+ }
47
+
48
+ function resultChars(m: Msg): number {
49
+ let n = 0;
50
+ for (const c of m.content ?? []) n += c.type === "image" ? IMAGE_CHARS : (c.text?.length ?? 0);
51
+ return n;
52
+ }
53
+
54
+ function editChars(call: any): number {
55
+ return (call.arguments?.edits ?? []).reduce((n: number, e: any) => n + (e.oldText?.length ?? 0) + (e.newText?.length ?? 0), 0);
56
+ }
57
+
58
+ /** Chars pruning would remove from this assistant message. */
59
+ function assistantSavings(m: Msg, o: HygieneOptions): number {
60
+ let n = 0;
61
+ for (const c of m.content ?? []) {
62
+ if (c.type === "thinking") n += c.thinking?.length ?? 0;
63
+ else if (c.type === "toolCall" && c.name === "write" && (c.arguments?.content?.length ?? 0) >= o.minPruneChars) n += c.arguments.content.length;
64
+ else if (c.type === "toolCall" && c.name === "edit" && editChars(c) >= o.minPruneChars) n += editChars(c);
65
+ }
66
+ return n;
67
+ }
68
+
69
+ function pruneAssistant(m: Msg, o: HygieneOptions): Msg {
70
+ const content = [];
71
+ for (const c of m.content ?? []) {
72
+ if (c.type === "thinking") continue;
73
+ if (c.type === "toolCall" && c.name === "write" && (c.arguments?.content?.length ?? 0) >= o.minPruneChars) {
74
+ const text = c.arguments.content as string;
75
+ content.push({ ...c, arguments: { ...c.arguments, content: `${PRUNED_MARK}: ${lines(text)} lines written, the file on disk is the source of truth>>` } });
76
+ } else if (c.type === "toolCall" && c.name === "edit" && editChars(c) >= o.minPruneChars) {
77
+ const edits = c.arguments.edits.map((e: any) => ({ oldText: `${PRUNED_MARK}>>`, newText: `${PRUNED_MARK}: ${lines(e.newText ?? "")} lines>>` }));
78
+ content.push({ ...c, arguments: { ...c.arguments, edits } });
79
+ } else content.push(c);
80
+ }
81
+ // a message that only held thinking keeps a marker: some templates choke on an empty assistant turn
82
+ if (!content.length) content.push({ type: "text", text: `${PRUNED_MARK}: reasoning>>` });
83
+ return { ...m, content };
84
+ }
85
+
86
+ function pruneResult(m: Msg, call: any): Msg {
87
+ const text = (m.content ?? []).map((c: any) => c.text ?? "").join("\n");
88
+ const what = [m.toolName, callTarget(call)].filter(Boolean).join(" ");
89
+ return { ...m, content: [{ type: "text", text: `${PRUNED_MARK}: output of ${what}, ${lines(text)} lines. Run it again if you need it.>>` }] };
90
+ }
91
+
92
+ /**
93
+ * Prunes everything before the watermark. The watermark sits on the last point where the prunable chars,
94
+ * summed from the first message, crossed a multiple of stepTokens. Only messages before the protected tail
95
+ * count, and the tail only moves forward, so earlier crossings never move: the pruned prefix is stable.
96
+ */
97
+ export function pruneContext(messages: Msg[], o: HygieneOptions = HYGIENE_DEFAULTS): { messages: Msg[]; stats: HygieneStats } {
98
+ const lastIndices = (role: string, keep: number) =>
99
+ messages.flatMap((m, i) => (m.role === role ? [i] : [])).slice(-keep);
100
+ const tail = [...lastIndices("assistant", o.keepThinkingTurns), ...lastIndices("toolResult", o.keepToolResults)];
101
+ const tailStart = tail.length ? Math.min(...tail) : messages.length;
102
+
103
+ const calls = new Map<string, any>();
104
+ for (const m of messages) if (m.role === "assistant") for (const c of m.content ?? []) if (c.type === "toolCall") calls.set(c.id, c);
105
+
106
+ const savings = (m: Msg) =>
107
+ m.role === "assistant" ? assistantSavings(m, o) : m.role === "toolResult" && resultChars(m) >= o.minPruneChars ? resultChars(m) : 0;
108
+
109
+ const step = o.stepTokens * CHARS_PER_TOKEN;
110
+ let sum = 0;
111
+ let watermark = 0;
112
+ let prunedChars = 0;
113
+ for (let i = 0; i < tailStart; i++) {
114
+ sum += savings(messages[i]!);
115
+ if (Math.floor(sum / step) > Math.floor(prunedChars / step)) {
116
+ watermark = i + 1;
117
+ prunedChars = sum;
118
+ }
119
+ }
120
+
121
+ const out = messages.map((m, i) => {
122
+ if (i >= watermark || !savings(m)) return m;
123
+ return m.role === "assistant" ? pruneAssistant(m, o) : pruneResult(m, calls.get(m.toolCallId));
124
+ });
125
+ return {
126
+ messages: out,
127
+ stats: {
128
+ messages: messages.length,
129
+ watermark,
130
+ prunedTokens: Math.round(prunedChars / CHARS_PER_TOKEN),
131
+ pendingTokens: Math.round((sum - prunedChars) / CHARS_PER_TOKEN),
132
+ stepTokens: o.stepTokens,
133
+ },
134
+ };
135
+ }
136
+
137
+ /**
138
+ * Sent as reasoning_budget_message with every request: when the thinking budget runs out llama.cpp forces this text,
139
+ * then the end-of-thinking tag, and the model reads it as its own words, so it says what to do next.
140
+ * A llama-server older than the per-request message closes the thinking silently; cuts are then detected by length (budgetVerdict).
141
+ */
142
+ export const BUDGET_MESSAGE = "Time is up. I will take the smallest safe next step with what I know, and write my open plan into PLAN.md.";
143
+
144
+ export type BudgetPhase = "planning" | "acting"; // right after the user spoke / between tool calls
145
+ export type BudgetMode = "auto" | "fixed" | "off";
146
+
147
+ /** Starting budgets, used until a model has enough samples to learn its own. */
148
+ export const BUDGET_DEFAULTS: Record<BudgetPhase, number> = { planning: 16384, acting: 4096 };
149
+ export const BUDGET_LIMITS: Record<BudgetPhase, [number, number]> = { planning: [4096, 32768], acting: [2048, 4096] }; // acting: past ~4k between tool calls it is overthinking
150
+
151
+ const BUDGET_SAMPLES_MIN = 10;
152
+ export const BUDGET_WINDOW = 30; // thinking lengths kept per model and phase
153
+ const BUDGET_PERCENTILE = 0.95;
154
+ const BUDGET_HEADROOM = 1.5;
155
+
156
+ /** Phase of the next request, from the OpenAI-style payload: a tool result last means the agent is mid-task. */
157
+ export function requestPhase(payload: any): BudgetPhase {
158
+ return payload?.messages?.at(-1)?.role === "tool" ? "acting" : "planning";
159
+ }
160
+
161
+ /**
162
+ * Learned budget: 1.5 × the 95th percentile of recent thinking lengths, clamped and rounded to 1k.
163
+ * Cuts never raise it (see budgetSample): a model that often runs away is held at its budget instead of chased.
164
+ * Replies that finish close under the budget raise it, a model that thinks little pulls it down.
165
+ * Undefined until there are enough samples.
166
+ */
167
+ export function learnedBudget(samples: number[], phase: BudgetPhase): number | undefined {
168
+ if (samples.length < BUDGET_SAMPLES_MIN) return undefined;
169
+ const sorted = [...samples].sort((a, b) => a - b);
170
+ const p = sorted[Math.min(sorted.length - 1, Math.floor(sorted.length * BUDGET_PERCENTILE))]!;
171
+ const [lo, hi] = BUDGET_LIMITS[phase];
172
+ return Math.min(hi, Math.max(lo, Math.ceil((p * BUDGET_HEADROOM) / 1024) * 1024));
173
+ }
174
+
175
+ /** The sample to learn from a reply: a cut counts as budget / headroom, so cuts alone give back the same budget. */
176
+ export function budgetSample(thought: number, budget: number, cut: boolean): number {
177
+ return cut ? budget / BUDGET_HEADROOM : thought;
178
+ }
179
+
180
+ /**
181
+ * What the server did with the budget sent for this reply, from the thinking length (estimated, ±10%):
182
+ * cut near the budget, "ignored" well past it (llama-server started with its own --reasoning-budget, or too old
183
+ * to read thinking_budget_tokens), otherwise within.
184
+ */
185
+ export function budgetVerdict(m: Msg, budget: number): "cut" | "ignored" | "within" {
186
+ if (thinkingWasCut(m)) return "cut";
187
+ const thought = thinkingTokens(m);
188
+ return thought > budget * 1.3 ? "ignored" : thought >= budget * 0.9 ? "cut" : "within";
189
+ }
190
+
191
+ /** Estimated thinking tokens of a reply (same chars/token as the hygiene). */
192
+ export function thinkingTokens(m: Msg): number {
193
+ const chars = (m.content ?? []).reduce((n: number, c: any) => n + (c.type === "thinking" ? (c.thinking?.length ?? 0) : 0), 0);
194
+ return Math.round(chars / CHARS_PER_TOKEN);
195
+ }
196
+
197
+ /** Whether llama.cpp cut this message's thinking at the budget. */
198
+ export function thinkingWasCut(m: Msg): boolean {
199
+ return (m.content ?? []).some((c: any) => c.type === "thinking" && (c.thinking ?? "").trimEnd().endsWith(BUDGET_MESSAGE));
200
+ }
201
+
202
+ /**
203
+ * Written to <project>/MAGI.md when missing, then appended to the system prompt on every run.
204
+ * English on purpose: Qwen-family models reason in English and follow English rules more reliably.
205
+ * Kept short: it is paid for on every request.
206
+ */
207
+ export const MAGI_MD = `# MAGI rules for local models
208
+
209
+ These rules are appended to the system prompt by the MAGI extension. Edit them freely; delete the file to get the defaults back.
210
+
211
+ ## Think less, act more
212
+ - Keep reasoning short: understand the step, decide, act. Do not re-plan what is already decided.
213
+ - Never draft code or file contents in your reasoning. Write them directly with the write/edit tool.
214
+ - If two attempts at the same approach fail, stop and change approach, or ask the user. Do not retry blindly.
215
+ - If your reasoning stopped abruptly mid-thought, or ends with "${BUDGET_MESSAGE.split(". ")[0]}.", your thinking budget ran out: in that turn make no large or irreversible change. Write your open plan into PLAN.md or take one small step you can verify; the next turn gives you a fresh budget.
216
+
217
+ ## Your context is small: spend it carefully
218
+ - Search before reading: use rg/find to locate code, then read only the needed range (offset/limit).
219
+ - Do not read whole large files, and do not re-read a file you just wrote.
220
+ - Trim long command output: pipe through tail, head or rg.
221
+ - Old tool outputs and old reasoning are removed from your context automatically (marked "${PRUNED_MARK}…>>"). Anything you will need later must be written down (see below). Never copy those markers into a file.
222
+
223
+ ## Keep the task state on disk
224
+ - For any task longer than a few steps, keep PLAN.md (a checklist) and NOTES.md (findings, decisions, commands that work).
225
+ - Update them after every completed step. If the context is compacted or a new session starts, continue from PLAN.md.
226
+
227
+ ## Edit safely
228
+ - edit oldText must match the file exactly: copy it from a fresh read, keep it small and unique.
229
+ - If an edit fails, read that region again instead of guessing.
230
+ - For big new files: write a skeleton first, then fill it with edits. Never replace a whole file with a partial version.
231
+
232
+ ## Verify, never assume
233
+ - Never invent paths, functions, APIs or flags: check with ls, rg, --help or the docs first.
234
+ - After a change, run the build, the tests or the program. Claim success only when a tool result proves it.
235
+ - If you cannot verify something, say so explicitly.
236
+
237
+ ## Finish cleanly
238
+ - When done: say briefly what changed, how you verified it, and what is left.
239
+ - Answer in the user's language.
240
+ `;
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "pi-magi-theme",
3
- "version": "0.2.3",
3
+ "version": "0.3.0",
4
4
  "description": "MAGI SYSTEM theme + extension for pi (Evangelion fan art): MAGI control screen panel, three-model /magi council, MECHA SELECT model picker, angel-attack loading, seven-seal context gauge, llama-swap telemetry",
5
5
  "keywords": [
6
6
  "pi-package",