pi-magi-theme 0.2.3 → 0.3.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +58 -7
- package/extensions/magi/index.ts +373 -91
- package/extensions/magi/local-models.ts +240 -0
- package/package.json +1 -1
package/README.md
CHANGED
|
@@ -40,10 +40,14 @@ Clone the repo and point pi at it instead (edits in the repo are live on the nex
|
|
|
40
40
|
| `/magi review [focus]` | the council reviews your pending changes (`git diff HEAD` plus untracked file names) before you commit |
|
|
41
41
|
| `/magi config` | pick a model for each MAGI |
|
|
42
42
|
| `/magi mecha` | MECHA SELECT: pick the llama-swap model to activate, each shown as a mecha head lit by its real state |
|
|
43
|
-
| `/magi
|
|
44
|
-
| `/magi
|
|
45
|
-
| `/magi
|
|
46
|
-
| `/magi
|
|
43
|
+
| `/magi compact` | toggle the compact side panel (basic info and animations only); remembered across sessions |
|
|
44
|
+
| `/magi status` | llama-swap report from its last 100 requests: speed, tokens, cache hits, MTP draft acceptance, durations, errors per model |
|
|
45
|
+
| `/magi cost` | set the electricity price per kWh and the currency (EUR or USD) for the COST row |
|
|
46
|
+
| `/magi panel` · `on` · `off` | hide/show the side panel, enable/disable the whole chrome |
|
|
47
|
+
| `/magi hygiene` · `on` · `off` · `step <tokens>` · `<turns> <results>` | show how much context the hygiene pruned; enable/disable it; how many prunable tokens make a pruning step (e.g. `step 40k`, default 15k); how many recent turns keep their thinking and how many tool results stay whole (e.g. `3 5`) |
|
|
48
|
+
| `/magi budget` · `auto` · `off` · `reset` · `message` · `<planning> <acting>` | show the thinking budget of the current model; learn it per model (default); leave it to llama-server; forget what was learned for this model; turn the closing message off/on; or fix it (e.g. `16k 4k`) |
|
|
49
|
+
|
|
50
|
+
Everything lives under `/magi`: type `/magi ` (with the space) to see every option with a short description; keep typing to narrow it down, Tab or Enter to pick one. Anything that is not an option is a question for the council.
|
|
47
51
|
|
|
48
52
|
## Lore ↔ function
|
|
49
53
|
|
|
@@ -72,6 +76,8 @@ The window title is `π - Magi - <working directory>`. After a run longer than 3
|
|
|
72
76
|
|
|
73
77
|
In fullscreen mode the side panel always reaches the bottom of the terminal.
|
|
74
78
|
|
|
79
|
+
In a terminal too narrow for it, the footer scrolls as one line instead of being cut off.
|
|
80
|
+
|
|
75
81
|
## The seventh seal: smart compaction
|
|
76
82
|
|
|
77
83
|
The theme does not compact anything itself: it shows who does. For better compaction install [pi-smart-compact](https://www.npmjs.com/package/pi-smart-compact):
|
|
@@ -82,6 +88,36 @@ pi install npm:pi-smart-compact
|
|
|
82
88
|
|
|
83
89
|
It extracts files, errors, decisions and open loops locally (no LLM calls), then synthesizes and verifies the summary. Point its `summaryModel` at a local model to keep compaction free. When it is installed, the sixth seal suggests `/smart-compact`, the seals name it while they break (`✶ BREAKING THE SEALS · smart-compact · 4s`) and the seventh seal reports who actually produced the summary: `smart-compact`, or `pi native` if it fell back to pi's own compactor.
|
|
84
90
|
|
|
91
|
+
## Local models: context hygiene, thinking budget, loop guard, MAGI.md
|
|
92
|
+
|
|
93
|
+
Local models run out of context on long tasks well before they run out of work. They also tend to think for minutes between two tool calls and to repeat the same command when stuck. MAGI works on all three, automatically (the thinking budget for llama-swap models, the rest for any model):
|
|
94
|
+
|
|
95
|
+
- **Context hygiene: the model's memory stays lean.** *Problem:* every file the agent reads and every long reasoning stays in the conversation, until the model's context is full and the task falls apart. *What MAGI does:* before each request it replaces old reasoning, old tool outputs and old file writes with a one-line note (`<<pruned by MAGI…>>`). Only what the model sees is trimmed: your saved session stays complete. The latest 3 turns keep their reasoning and the latest 5 tool outputs stay whole. On two real sessions it brought 104k and 119k tokens down to ~64k and ~47k.
|
|
96
|
+
- **Thinking budget: no more ten-minute thinks.** *Problem:* a local model can reason for thousands of tokens before a simple step, and at 10 tokens/s that is minutes of waiting. *What MAGI does:* it gives the model a maximum length of thinking on every request: larger right after you write (planning), smaller between tool calls (acting). When the limit is reached the model is stopped mid-thought, says *"Time is up. I will take the smallest safe next step…"* and acts. The limit adapts to each model on its own. The side panel counts these cuts (`HYGIENE -18.2k ✂2`).
|
|
97
|
+
- **Loop guard: no endless retries.** *Problem:* a stuck model runs the same command again and again. *What MAGI does:* the third identical tool call in a row is blocked, with a message asking the model to try something else.
|
|
98
|
+
- **MAGI.md: house rules for the model.** *Problem:* local models repeat the same mistakes: reading whole files, inventing paths, claiming success without checking. *What MAGI does:* it creates `MAGI.md` in your project on the first start (never overwritten) and adds it to the model's instructions on every run: short rules against these mistakes, plus the habit of keeping the task plan in `PLAN.md` and findings in `NOTES.md`, so nothing important is lost when old context is trimmed. Edit it per project; delete it to get the defaults back.
|
|
99
|
+
|
|
100
|
+
Nothing needs setting up in llama-swap. The defaults suit most tasks; two adjustments are worth knowing:
|
|
101
|
+
|
|
102
|
+
- **Long tasks: `/magi hygiene step 40k`.** Each time the hygiene trims, the server has to re-read part of the conversation, and on some models (see *Hybrid models* below) almost all of it, which can take a few minutes. A step of 40k trims less often: on a real session it cut the re-reading from ~6 minutes to ~1.
|
|
103
|
+
- **A model that really needs to think longer: `/magi budget <planning> <acting>`**, e.g. `/magi budget 16k 8k`. Frequent ✂ cuts in the side panel are the sign.
|
|
104
|
+
|
|
105
|
+
### How it works
|
|
106
|
+
|
|
107
|
+
For the curious, and for tuning.
|
|
108
|
+
|
|
109
|
+
**Hygiene.** Trimming happens in steps, not on every request: the conversation up to a mark is trimmed, and the mark only moves forward once ~15k more tokens (the step) could be trimmed. Between steps the conversation only grows at the end, so llama.cpp reuses what it already processed (its KV cache) and only reads the new messages. Each step changes the conversation from the first newly trimmed message on, and the server re-reads from there: that is why steps are large and rare. Old tool outputs are replaced, not summarized: [simple observation masking matches LLM summarization at half the cost](https://arxiv.org/abs/2508.21433).
|
|
110
|
+
|
|
111
|
+
**Hybrid models.** Some models (e.g. Qwen3.8 Flash Next; dense or MoE does not matter) mix a few classic attention layers with recurrent ones, which squeeze the whole conversation into a fixed-size state instead of keeping each token. The server cannot rewind that state to an arbitrary point: it can only restore a saved copy (a checkpoint) and re-read from there, and the only copy before the trimmed part is usually the end of the system prompt. So on these models each step re-reads almost the whole prompt. With 10k-token thoughts and 4k-token file reads, a 15k step is crossed every 2–3 turns: on a real 70k-token session that was 3 re-reads in 6 requests (~6 min at 185 tokens/s); with a 40k step, 1 re-read (~1 min) for a prompt at most 6k larger. A model is hybrid if its llama-server log shows `restored context checkpoint` lines.
|
|
112
|
+
|
|
113
|
+
**Thinking budget.** MAGI sends the limit with every request (`thinking_budget_tokens`), and llama.cpp applies it. It learns per model and per phase: 1.5 × the 95th percentile of the model's last 30 thinking lengths, rounded up to 1k. A cut is recorded as 1/1.5 of the limit it hit, so cuts never raise the limit: a model that often runs away is held, not chased. Replies that end close under the limit raise it; a model that thinks little lowers it. Limits: planning 4k–32k, acting 2k–4k (past ~4k between two tool calls it is overthinking). Until a model has 10 replies in a phase it uses 16k / 4k.
|
|
114
|
+
|
|
115
|
+
When the limit is reached llama.cpp does not abort the reply: it inserts the closing sentence and the end-of-thinking tag, and the model goes on to act. MAGI sends that sentence with every request too (`reasoning_budget_message`; `/magi budget message` turns it off and on). `MAGI.md` tells the model what to do after a cut (one small verifiable step, the open plan into `PLAN.md`), and the next turn gets a fresh budget.
|
|
116
|
+
|
|
117
|
+
**llama-server versions.** The per-request limit works whenever llama-server was started without `--reasoning-budget`, which is the default. If it was started with one, the server's limit wins: MAGI notices the model thinking well past its own limit and `/magi budget` says so. A llama-server too old for the closing sentence ends the thinking silently, and MAGI detects the cut by its length.
|
|
118
|
+
|
|
119
|
+
**Loop guard details.** A tool call counts as identical when both the tool and its arguments match. A file write or edit that copies a `<<pruned by MAGI…>>` note into a file is blocked too.
|
|
120
|
+
|
|
85
121
|
## The council
|
|
86
122
|
|
|
87
123
|
`/magi <question>` asks three models in parallel, each with its own nature, then shows the votes and a majority verdict:
|
|
@@ -98,21 +134,36 @@ Each nature is a lens, not a specialty, so the council answers any question, not
|
|
|
98
134
|
|
|
99
135
|
## Configuration
|
|
100
136
|
|
|
101
|
-
`~/.pi/agent/magi.json` (written by `/magi config`, `/magi
|
|
137
|
+
`~/.pi/agent/magi.json` (written by `/magi config`, `/magi compact` and `/magi cost`, editable by hand):
|
|
102
138
|
|
|
103
139
|
```json
|
|
104
140
|
{
|
|
105
141
|
"MELCHIOR": { "model": "llama-swap/Qwen3.8 27B Q4_K_M - Thinking", "thinking": "low" },
|
|
106
142
|
"ui": { "compact": false, "kwhPrice": 0.30, "currency": "EUR" },
|
|
107
|
-
"loads": { "qwen3.8-27b": 41200 }
|
|
143
|
+
"loads": { "qwen3.8-27b": 41200 },
|
|
144
|
+
"totalWh": 1843.2,
|
|
145
|
+
"hygiene": { "enabled": true, "keepThinkingTurns": 3, "keepToolResults": 5, "stepTokens": 15000, "minPruneChars": 600 },
|
|
146
|
+
"thinkingBudget": { "mode": "auto", "planning": 16384, "acting": 4096, "message": true, "learned": { "qwen3.8-27b": { "acting": [812, 430, 2211] } } }
|
|
108
147
|
}
|
|
109
148
|
```
|
|
110
149
|
|
|
111
150
|
- per MAGI: `model` (unset = current session model) and optional `thinking` level;
|
|
112
151
|
- `ui.compact`: start with the compact side panel;
|
|
113
|
-
- `ui.kwhPrice` and `ui.currency` (`EUR` or `USD`): the COST row multiplies the GPU energy
|
|
152
|
+
- `ui.kwhPrice` and `ui.currency` (`EUR` or `USD`): the COST row multiplies the GPU energy by this price, showing the running total of every session with the current one in brackets;
|
|
153
|
+
- `totalWh`: written by the theme, GPU energy summed over every session (delete the key to reset the COST total);
|
|
154
|
+
- `hygiene` and `thinkingBudget` are set with `/magi hygiene` and `/magi budget` (`hygiene.minPruneChars` by hand only); `thinkingBudget.message` sends the closing message with every request (default `true`); `thinkingBudget.learned` is written by the theme (recent thinking lengths per model and phase);
|
|
114
155
|
- `loads`: written by the theme, how long each llama-swap model took to load last time (paces the angel attack; 60s when unknown).
|
|
115
156
|
|
|
157
|
+
## Release
|
|
158
|
+
|
|
159
|
+
`.github/workflows/publish.yml` publishes to npm when a `v*` tag is pushed, and refuses if the tag does not match `package.json`:
|
|
160
|
+
|
|
161
|
+
```
|
|
162
|
+
npm version patch && git push --follow-tags
|
|
163
|
+
```
|
|
164
|
+
|
|
165
|
+
No token: npmjs is configured to trust this repository's `publish.yml` (npm trusted publishing, OIDC), which also signs the provenance.
|
|
166
|
+
|
|
116
167
|
## llama-swap
|
|
117
168
|
|
|
118
169
|
When the session model uses the `llama-swap` provider, the side panel:
|
package/extensions/magi/index.ts
CHANGED
|
@@ -17,22 +17,41 @@
|
|
|
17
17
|
* - while a model loads an angel attacks the MAGI: red spreads through BALTHASAR, MELCHIOR and CASPAR at the pace of
|
|
18
18
|
* the model's last load, a corner of CASPAR holds out blinking; once loaded, blue takes the MAGI back from that corner
|
|
19
19
|
* - llama-swap telemetry: VRAM, GPU load/temp/power, energy used, RAM, server-side tok/s, prompt tok/s, cache hits
|
|
20
|
-
* - /magi config → assign a model to each MAGI; /magi
|
|
20
|
+
* - /magi config → assign a model to each MAGI; /magi compact|status → smaller panel, llama-swap report
|
|
21
21
|
*
|
|
22
22
|
* Fan art: the MAGI and their screen come from Neon Genesis Evangelion, all rights reserved to khara, Inc.
|
|
23
23
|
* Use with the theme ../../themes/magi.json
|
|
24
24
|
*/
|
|
25
25
|
|
|
26
26
|
import { execFile } from "node:child_process";
|
|
27
|
-
import { readFileSync, writeFileSync } from "node:fs";
|
|
27
|
+
import { existsSync, readFileSync, writeFileSync } from "node:fs";
|
|
28
28
|
import { homedir } from "node:os";
|
|
29
29
|
import { join } from "node:path";
|
|
30
30
|
import { promisify } from "node:util";
|
|
31
31
|
import type { AssistantMessage, Model } from "@earendil-works/pi-ai";
|
|
32
32
|
import { completeSimple } from "@earendil-works/pi-ai";
|
|
33
|
-
import type { ExtensionAPI, ExtensionContext, Theme, ThemeColor } from "@earendil-works/pi-coding-agent";
|
|
33
|
+
import type { ExtensionAPI, ExtensionCommandContext, ExtensionContext, Theme, ThemeColor } from "@earendil-works/pi-coding-agent";
|
|
34
34
|
import type { Component, OverlayHandle, TUI } from "@earendil-works/pi-tui";
|
|
35
|
-
import { HStack, matchesKey, truncateToWidth, visibleWidth, wrapTextWithAnsi } from "@earendil-works/pi-tui";
|
|
35
|
+
import { HStack, matchesKey, sliceByColumn, truncateToWidth, visibleWidth, wrapTextWithAnsi } from "@earendil-works/pi-tui";
|
|
36
|
+
import {
|
|
37
|
+
BUDGET_DEFAULTS,
|
|
38
|
+
BUDGET_MESSAGE,
|
|
39
|
+
BUDGET_WINDOW,
|
|
40
|
+
HYGIENE_DEFAULTS,
|
|
41
|
+
MAGI_MD,
|
|
42
|
+
PRUNED_MARK,
|
|
43
|
+
learnedBudget,
|
|
44
|
+
budgetSample,
|
|
45
|
+
pruneContext,
|
|
46
|
+
requestPhase,
|
|
47
|
+
thinkingTokens,
|
|
48
|
+
budgetVerdict,
|
|
49
|
+
thinkingWasCut,
|
|
50
|
+
type BudgetMode,
|
|
51
|
+
type BudgetPhase,
|
|
52
|
+
type HygieneOptions,
|
|
53
|
+
type HygieneStats,
|
|
54
|
+
} from "./local-models.ts";
|
|
36
55
|
|
|
37
56
|
/* ────────────────────────────────────────────────────────────── art ── */
|
|
38
57
|
|
|
@@ -317,6 +336,8 @@ const state = {
|
|
|
317
336
|
rebornAt: 0,
|
|
318
337
|
hasSmartCompact: false,
|
|
319
338
|
lastCouncil: undefined as { verdict: Vote | null; tally: number; question: string } | undefined,
|
|
339
|
+
pruned: 0, // tokens the context hygiene keeps away from the model
|
|
340
|
+
thinkCuts: 0, // replies whose thinking llama.cpp cut at the budget
|
|
320
341
|
};
|
|
321
342
|
|
|
322
343
|
/** Panel preferences, persisted under "ui" in ~/.pi/agent/magi.json. */
|
|
@@ -324,7 +345,7 @@ type Currency = "EUR" | "USD";
|
|
|
324
345
|
|
|
325
346
|
const ui = {
|
|
326
347
|
compact: false,
|
|
327
|
-
kwhPrice: undefined as number | undefined, // price per kWh, for the COST row (/magi
|
|
348
|
+
kwhPrice: undefined as number | undefined, // price per kWh, for the COST row (/magi cost)
|
|
328
349
|
currency: "EUR" as Currency,
|
|
329
350
|
};
|
|
330
351
|
|
|
@@ -340,10 +361,15 @@ function setPhase(p: Phase): void {
|
|
|
340
361
|
}
|
|
341
362
|
|
|
342
363
|
const ANIM_STEP_MS = 220;
|
|
364
|
+
/** How long one Sephirah stays lit in the footer: its name and meaning are text, they need time to be read. */
|
|
365
|
+
const SEPHIRAH_STEP_MS = 1400;
|
|
343
366
|
const FAIL_FLASH_MS = 2500;
|
|
344
367
|
const REBIRTH_MS = 6000;
|
|
345
368
|
const DONE_TITLE_AFTER_MS = 30_000;
|
|
346
369
|
const SIXTH_SEAL_PERCENT = (6 / 7) * 100;
|
|
370
|
+
/** Footer marquee: one cell every FOOTER_SCROLL_MS, with this gap between the loop's end and its start. */
|
|
371
|
+
const FOOTER_SCROLL_MS = 260;
|
|
372
|
+
const FOOTER_SCROLL_GAP = " · ";
|
|
347
373
|
|
|
348
374
|
/** Sync ratio: share of tool calls that succeeded (null before the first tool). */
|
|
349
375
|
function syncPercent(): number | null {
|
|
@@ -416,7 +442,9 @@ interface TokenStats {
|
|
|
416
442
|
|
|
417
443
|
const tokens: TokenStats = { input: 0, output: 0, cacheRead: 0, cost: 0 };
|
|
418
444
|
|
|
419
|
-
|
|
445
|
+
/** Session totals from one assistant reply: tokens, cost, and whether its thinking was cut at the budget. */
|
|
446
|
+
function countAssistant(m: AssistantMessage): void {
|
|
447
|
+
if (thinkingWasCut(m as any)) state.thinkCuts++;
|
|
420
448
|
tokens.input += m.usage?.input ?? 0;
|
|
421
449
|
tokens.output += m.usage?.output ?? 0;
|
|
422
450
|
tokens.cacheRead += m.usage?.cacheRead ?? 0;
|
|
@@ -427,8 +455,9 @@ function addUsage(m: AssistantMessage): void {
|
|
|
427
455
|
function recountSession(ctx: ExtensionContext): void {
|
|
428
456
|
Object.assign(tokens, { input: 0, output: 0, cacheRead: 0, cost: 0 });
|
|
429
457
|
state.lastCouncil = undefined;
|
|
458
|
+
state.thinkCuts = 0;
|
|
430
459
|
for (const entry of ctx.sessionManager.getBranch()) {
|
|
431
|
-
if (entry.type === "message" && entry.message.role === "assistant")
|
|
460
|
+
if (entry.type === "message" && entry.message.role === "assistant") countAssistant(entry.message as AssistantMessage);
|
|
432
461
|
if (entry.type === "custom" && entry.customType === "magi-verdict") {
|
|
433
462
|
const d = entry.data as { verdict?: Vote | null; tally?: number; question?: string } | undefined;
|
|
434
463
|
if (d && Array.isArray((d as any).opinions)) state.lastCouncil = { verdict: d.verdict ?? null, tally: d.tally ?? 0, question: d.question ?? "" };
|
|
@@ -710,6 +739,8 @@ const swap = {
|
|
|
710
739
|
cacheTokens: 0,
|
|
711
740
|
inputTokens: 0,
|
|
712
741
|
energyWh: 0, // GPU energy since the session started
|
|
742
|
+
diskWh: 0, // GPU energy of all sessions, as last read from/written to magi.json
|
|
743
|
+
savedWh: 0, // the part of energyWh already added to diskWh
|
|
713
744
|
lastSampleAt: 0,
|
|
714
745
|
lastWatts: 0,
|
|
715
746
|
};
|
|
@@ -760,6 +791,35 @@ function sampleEnergy(now = Date.now()): void {
|
|
|
760
791
|
}
|
|
761
792
|
swap.lastSampleAt = now;
|
|
762
793
|
swap.lastWatts = watts;
|
|
794
|
+
persistEnergy();
|
|
795
|
+
}
|
|
796
|
+
|
|
797
|
+
/** GPU energy of every session so far: what is on disk plus this session's not-yet-written part. */
|
|
798
|
+
function totalWh(): number {
|
|
799
|
+
return swap.diskWh + swap.energyWh - swap.savedWh;
|
|
800
|
+
}
|
|
801
|
+
|
|
802
|
+
const ENERGY_SAVE_MS = 60_000;
|
|
803
|
+
let lastPersistAt = 0;
|
|
804
|
+
|
|
805
|
+
/**
|
|
806
|
+
* Adds this session's new energy to the running total in magi.json, so the COST row survives restarts.
|
|
807
|
+
* ponytail: writes at most once a minute, and adds a delta so parallel sessions do not overwrite each other.
|
|
808
|
+
* A crash therefore loses up to a minute of energy, which is what a kill loses anyway.
|
|
809
|
+
*/
|
|
810
|
+
function persistEnergy(): void {
|
|
811
|
+
const delta = swap.energyWh - swap.savedWh;
|
|
812
|
+
if (delta <= 0 || Date.now() - lastPersistAt < ENERGY_SAVE_MS) return;
|
|
813
|
+
lastPersistAt = Date.now();
|
|
814
|
+
const cfg = loadMagiConfig();
|
|
815
|
+
cfg.totalWh = (cfg.totalWh ?? 0) + delta;
|
|
816
|
+
try {
|
|
817
|
+
saveMagiConfig(cfg);
|
|
818
|
+
} catch {
|
|
819
|
+
return; // disk unavailable: keep the delta and retry on the next sample
|
|
820
|
+
}
|
|
821
|
+
swap.savedWh = swap.energyWh;
|
|
822
|
+
swap.diskWh = cfg.totalWh;
|
|
763
823
|
}
|
|
764
824
|
|
|
765
825
|
async function refreshSwapMetrics(): Promise<void> {
|
|
@@ -995,7 +1055,7 @@ async function prewarmPrefix(cwd: string): Promise<void> {
|
|
|
995
1055
|
}
|
|
996
1056
|
}
|
|
997
1057
|
|
|
998
|
-
/* ── /magi
|
|
1058
|
+
/* ── /magi status: a report built from the last requests llama-swap recorded ── */
|
|
999
1059
|
|
|
1000
1060
|
interface ActivityRow {
|
|
1001
1061
|
timestamp: string;
|
|
@@ -1280,6 +1340,11 @@ class MagiPanel implements Component {
|
|
|
1280
1340
|
if (!compact && usage?.contextWindow) {
|
|
1281
1341
|
out.push(this.field("CONTEXT", `${usage.tokens == null ? "?" : fmtTokens(usage.tokens)} / ${fmtTokens(usage.contextWindow)}`, inner, "muted"));
|
|
1282
1342
|
}
|
|
1343
|
+
// HYGIENE: tokens pruned from what the model sees, ✂ = thinking cut at the budget
|
|
1344
|
+
if (!compact && (state.pruned || state.thinkCuts)) {
|
|
1345
|
+
const cuts = state.thinkCuts ? ` ✂${state.thinkCuts}` : "";
|
|
1346
|
+
out.push(this.field("HYGIENE", `-${fmtTokens(state.pruned)}${cuts}`, inner, state.thinkCuts ? "warning" : "muted"));
|
|
1347
|
+
}
|
|
1283
1348
|
return out;
|
|
1284
1349
|
}
|
|
1285
1350
|
|
|
@@ -1311,11 +1376,12 @@ class MagiPanel implements Component {
|
|
|
1311
1376
|
return out;
|
|
1312
1377
|
}
|
|
1313
1378
|
|
|
1314
|
-
/** COST: GPU energy
|
|
1379
|
+
/** COST: GPU energy × price per kWh set with /magi cost — all sessions, with this one in brackets. */
|
|
1315
1380
|
private costRow(inner: number): string {
|
|
1316
1381
|
if (!swap.gpus.length) return this.field("COST", "—", inner, "muted");
|
|
1317
|
-
if (ui.kwhPrice === undefined) return this.field("COST", "→ /magi
|
|
1318
|
-
|
|
1382
|
+
if (ui.kwhPrice === undefined) return this.field("COST", "→ /magi cost", inner, "dim");
|
|
1383
|
+
const price = (wh: number) => fmtMoney((wh / 1000) * ui.kwhPrice!);
|
|
1384
|
+
return this.field("COST", price(totalWh()) + this.theme.fg("dim", ` (ses ${price(swap.energyWh)})`), inner, "warning");
|
|
1319
1385
|
}
|
|
1320
1386
|
|
|
1321
1387
|
private swapRows(inner: number, compact: boolean): string[] {
|
|
@@ -1357,7 +1423,7 @@ class MagiPanel implements Component {
|
|
|
1357
1423
|
);
|
|
1358
1424
|
}
|
|
1359
1425
|
if (swap.ramTotal) out.push(this.field("RAM", `${gib(swap.ramUsed).toFixed(1)} / ${Math.round(gib(swap.ramTotal))}G`, inner, "muted"));
|
|
1360
|
-
if (swap.gpus.length) out.push(this.field("ENERGY", `${(
|
|
1426
|
+
if (swap.gpus.length) out.push(this.field("ENERGY", `${(totalWh() / 1000).toFixed(3)} kWh`, inner, "muted"));
|
|
1361
1427
|
out.push(this.costRow(inner));
|
|
1362
1428
|
return out;
|
|
1363
1429
|
}
|
|
@@ -1466,9 +1532,49 @@ interface MagiUnitConfig {
|
|
|
1466
1532
|
type MagiConfig = Partial<Record<MagiUnit, MagiUnitConfig>> & {
|
|
1467
1533
|
ui?: { compact?: boolean; kwhPrice?: number; currency?: Currency };
|
|
1468
1534
|
loads?: Record<string, number>; // real model id → ms its last load took
|
|
1535
|
+
totalWh?: number; // GPU energy summed over every session, for the COST row
|
|
1536
|
+
hygiene?: Partial<HygieneOptions> & { enabled?: boolean }; // context pruning for local models, see local-models.ts
|
|
1537
|
+
thinkingBudget?: {
|
|
1538
|
+
mode?: BudgetMode; // auto (learned per model) · fixed · off, set with /magi budget
|
|
1539
|
+
planning?: number; // fixed budgets
|
|
1540
|
+
acting?: number;
|
|
1541
|
+
message?: boolean; // send BUDGET_MESSAGE as reasoning_budget_message (default on), /magi budget message
|
|
1542
|
+
learned?: Record<string, Partial<Record<BudgetPhase, number[]>>>; // written by the theme: recent thinking lengths per model
|
|
1543
|
+
};
|
|
1469
1544
|
};
|
|
1470
1545
|
|
|
1471
1546
|
const MAGI_CONFIG_PATH = join(homedir(), ".pi", "agent", "magi.json");
|
|
1547
|
+
/** /magi arguments that manage the theme instead of asking the council. */
|
|
1548
|
+
const UI_ARGS = /^(on|off|panel|compact|status|cost)$|^(hygiene|budget)(\s|$)/i; // anything else is a question
|
|
1549
|
+
/** /magi arguments offered by autocomplete: the full argument, and what it does. */
|
|
1550
|
+
const MAGI_ARGS: [string, string][] = [
|
|
1551
|
+
["review", "the council reviews your pending changes before you commit"],
|
|
1552
|
+
["config", "pick a model for each MAGI"],
|
|
1553
|
+
["mecha", "MECHA SELECT: pick the llama-swap model to activate"],
|
|
1554
|
+
["status", "llama-swap report: speed, tokens, cache hits, errors per model"],
|
|
1555
|
+
["panel", "hide/show the side panel"],
|
|
1556
|
+
["compact", "toggle the compact side panel"],
|
|
1557
|
+
["cost", "electricity price and currency for the COST row"],
|
|
1558
|
+
["on", "enable the MAGI chrome"],
|
|
1559
|
+
["off", "disable the MAGI chrome"],
|
|
1560
|
+
["hygiene", "show how much context was pruned"],
|
|
1561
|
+
["hygiene on", "enable context pruning"],
|
|
1562
|
+
["hygiene off", "disable context pruning"],
|
|
1563
|
+
["hygiene step 40k", "prune less often: fewer prompt re-reads on long tasks (default 15k)"],
|
|
1564
|
+
["hygiene 3 5", "recent turns that keep their thinking, tool results kept whole"],
|
|
1565
|
+
["budget", "show the thinking budget of the current model"],
|
|
1566
|
+
["budget auto", "learn the budget per model (default)"],
|
|
1567
|
+
["budget off", "no budget: llama-server decides"],
|
|
1568
|
+
["budget reset", "forget what was learned for the current model"],
|
|
1569
|
+
["budget message", "turn the closing message after a cut off/on"],
|
|
1570
|
+
["budget 16k 4k", "fixed budget: planning, acting"],
|
|
1571
|
+
];
|
|
1572
|
+
function argCompletions(table: [string, string][], prefix: string) {
|
|
1573
|
+
const p = prefix.trimStart().toLowerCase();
|
|
1574
|
+
const items = table.filter(([value]) => value.startsWith(p)).map(([value, description]) => ({ value, label: value, description }));
|
|
1575
|
+
return items.length ? items : null;
|
|
1576
|
+
}
|
|
1577
|
+
const LOOP_REPEATS = 3; // the same tool call this many times in a row is blocked
|
|
1472
1578
|
|
|
1473
1579
|
function loadMagiConfig(): MagiConfig {
|
|
1474
1580
|
try {
|
|
@@ -1684,7 +1790,7 @@ function buildDeliberationView(
|
|
|
1684
1790
|
};
|
|
1685
1791
|
}
|
|
1686
1792
|
|
|
1687
|
-
/** A read-only boxed report (used by /magi
|
|
1793
|
+
/** A read-only boxed report (used by /magi status); any key closes it. */
|
|
1688
1794
|
function buildReportView(theme: Theme, lines: string[], close: () => void) {
|
|
1689
1795
|
return {
|
|
1690
1796
|
render(width: number): string[] {
|
|
@@ -1807,6 +1913,7 @@ async function configureMagi(ctx: ExtensionContext): Promise<void> {
|
|
|
1807
1913
|
function footerLeft(th: Theme, now = Date.now()): string {
|
|
1808
1914
|
const dim = (s: string) => th.fg("dim", s);
|
|
1809
1915
|
const step = Math.floor(now / ANIM_STEP_MS);
|
|
1916
|
+
const slowStep = Math.floor(now / SEPHIRAH_STEP_MS);
|
|
1810
1917
|
const lights = (fn: (i: number) => NodeLight) => SEPHIROT.map((_, i) => fn(i));
|
|
1811
1918
|
|
|
1812
1919
|
if (state.compacting) {
|
|
@@ -1860,7 +1967,7 @@ function footerLeft(th: Theme, now = Date.now()): string {
|
|
|
1860
1967
|
);
|
|
1861
1968
|
}
|
|
1862
1969
|
if (state.phase === "thinking") {
|
|
1863
|
-
const cur =
|
|
1970
|
+
const cur = slowStep % 3; // Keter, Chokmah, Binah
|
|
1864
1971
|
const s = SEPHIROT[cur]!;
|
|
1865
1972
|
return (
|
|
1866
1973
|
th.fg("accent", "◆ ") +
|
|
@@ -1870,7 +1977,7 @@ function footerLeft(th: Theme, now = Date.now()): string {
|
|
|
1870
1977
|
);
|
|
1871
1978
|
}
|
|
1872
1979
|
if (state.phase === "responding") {
|
|
1873
|
-
const cur = 5 + (
|
|
1980
|
+
const cur = 5 + (slowStep % 5); // Tiferet → Malkuth
|
|
1874
1981
|
const s = SEPHIROT[cur]!;
|
|
1875
1982
|
const tps = liveTps();
|
|
1876
1983
|
return (
|
|
@@ -1908,9 +2015,21 @@ function footerLeft(th: Theme, now = Date.now()): string {
|
|
|
1908
2015
|
|
|
1909
2016
|
/** Footer: the animation on the left, other extensions' statuses on the right. */
|
|
1910
2017
|
function buildFooter(tui: TUI, theme: Theme, footerData: any) {
|
|
2018
|
+
let scrolling = false;
|
|
2019
|
+
let scrollOff = 0;
|
|
2020
|
+
let scrollPeriod = 1;
|
|
2021
|
+
let scrollLine = "";
|
|
1911
2022
|
const timer = setInterval(() => {
|
|
1912
2023
|
if (animating()) tui.requestRender();
|
|
1913
2024
|
}, ANIM_STEP_MS);
|
|
2025
|
+
// The marquee keeps its own cadence: one cell per tick, never derived from the clock, so it
|
|
2026
|
+
// does not stutter when the line is re-rendered at some other pace (animation, streamed tokens)
|
|
2027
|
+
// nor jump when the left side changes length (sephirah name, tok/s) and with it the loop period.
|
|
2028
|
+
const scrollTimer = setInterval(() => {
|
|
2029
|
+
if (!scrolling) return;
|
|
2030
|
+
scrollOff = (scrollOff + 1) % scrollPeriod;
|
|
2031
|
+
tui.requestRender();
|
|
2032
|
+
}, FOOTER_SCROLL_MS);
|
|
1914
2033
|
|
|
1915
2034
|
const unsub = footerData?.onBranchChange?.(() => tui.requestRender());
|
|
1916
2035
|
|
|
@@ -1919,14 +2038,30 @@ function buildFooter(tui: TUI, theme: Theme, footerData: any) {
|
|
|
1919
2038
|
const dim = (s: string) => theme.fg("dim", s);
|
|
1920
2039
|
const statuses = footerData?.getExtensionStatuses?.();
|
|
1921
2040
|
const right = statuses ? [...statuses.values()].filter(Boolean).join(dim(" │ ")) : "";
|
|
1922
|
-
const
|
|
1923
|
-
const
|
|
1924
|
-
|
|
1925
|
-
|
|
2041
|
+
const left = footerLeft(theme);
|
|
2042
|
+
const gap = width - visibleWidth(left) - visibleWidth(right);
|
|
2043
|
+
scrolling = gap < 1;
|
|
2044
|
+
if (!scrolling) {
|
|
2045
|
+
scrollLine = "";
|
|
2046
|
+
scrollOff = 0;
|
|
2047
|
+
return [left + " ".repeat(gap) + right];
|
|
2048
|
+
}
|
|
2049
|
+
// Too narrow to fit: scroll the whole line instead of cutting it off. The line is frozen
|
|
2050
|
+
// while it scrolls — live text changes width (seconds, tok/s, sephirah names) and every
|
|
2051
|
+
// change would shift it under the window. A fresh line is taken when it has the same
|
|
2052
|
+
// width, so nothing moves, otherwise at the end of the loop.
|
|
2053
|
+
const line = left + dim(FOOTER_SCROLL_GAP) + right + dim(FOOTER_SCROLL_GAP);
|
|
2054
|
+
const period = visibleWidth(line);
|
|
2055
|
+
if (!scrollLine || scrollOff === 0 || period === scrollPeriod) {
|
|
2056
|
+
scrollLine = line;
|
|
2057
|
+
scrollPeriod = period;
|
|
2058
|
+
}
|
|
2059
|
+
return [sliceByColumn(scrollLine + scrollLine, scrollOff % scrollPeriod, width, true)];
|
|
1926
2060
|
},
|
|
1927
2061
|
invalidate() {},
|
|
1928
2062
|
dispose() {
|
|
1929
2063
|
clearInterval(timer);
|
|
2064
|
+
clearInterval(scrollTimer);
|
|
1930
2065
|
unsub?.();
|
|
1931
2066
|
},
|
|
1932
2067
|
};
|
|
@@ -1953,6 +2088,22 @@ export default function (pi: ExtensionAPI) {
|
|
|
1953
2088
|
let panelHandle: OverlayHandle | undefined;
|
|
1954
2089
|
let panel: MagiPanel | undefined;
|
|
1955
2090
|
let titleDone = false;
|
|
2091
|
+
let hygiene = { enabled: true, ...HYGIENE_DEFAULTS };
|
|
2092
|
+
let hygieneStats: HygieneStats | undefined;
|
|
2093
|
+
let budgetMode: BudgetMode = "auto";
|
|
2094
|
+
let fixedBudget = { ...BUDGET_DEFAULTS };
|
|
2095
|
+
let budgetMessage = true;
|
|
2096
|
+
let learned: Record<string, Partial<Record<BudgetPhase, number[]>>> = {};
|
|
2097
|
+
let pendingBudget: { model: string; phase: BudgetPhase; tokens: number } | undefined; // the request in flight
|
|
2098
|
+
const budgetIgnored = new Set<string>(); // models whose llama-server thought past the budget: it sets its own
|
|
2099
|
+
const budgetFor = (model: string, phase: BudgetPhase) =>
|
|
2100
|
+
budgetMode === "fixed" ? fixedBudget[phase] : (learnedBudget(learned[model]?.[phase] ?? [], phase) ?? BUDGET_DEFAULTS[phase]);
|
|
2101
|
+
const persistBudget = () => {
|
|
2102
|
+
const cfg = loadMagiConfig();
|
|
2103
|
+
saveMagiConfig({ ...cfg, thinkingBudget: { mode: budgetMode, ...fixedBudget, message: budgetMessage, learned } });
|
|
2104
|
+
};
|
|
2105
|
+
let lastCall = ""; // loop guard: fingerprint of the previous tool call and how often it repeated
|
|
2106
|
+
let repeats = 0;
|
|
1956
2107
|
|
|
1957
2108
|
const repaint = () => {
|
|
1958
2109
|
panel?.invalidate();
|
|
@@ -2113,6 +2264,23 @@ export default function (pi: ExtensionAPI) {
|
|
|
2113
2264
|
|
|
2114
2265
|
pi.on("session_start", async (event, ctx) => {
|
|
2115
2266
|
liveCtx = ctx;
|
|
2267
|
+
const magiCfg = loadMagiConfig();
|
|
2268
|
+
hygiene = { enabled: true, ...HYGIENE_DEFAULTS, ...magiCfg.hygiene };
|
|
2269
|
+
const tb = magiCfg.thinkingBudget ?? {};
|
|
2270
|
+
budgetMode = tb.mode ?? "auto";
|
|
2271
|
+
fixedBudget = { planning: tb.planning ?? BUDGET_DEFAULTS.planning, acting: tb.acting ?? BUDGET_DEFAULTS.acting };
|
|
2272
|
+
budgetMessage = tb.message ?? true;
|
|
2273
|
+
learned = tb.learned ?? {};
|
|
2274
|
+
// the local-model rules live next to AGENTS.md, created once so the user can edit them
|
|
2275
|
+
const rules = join(ctx.cwd, "MAGI.md");
|
|
2276
|
+
if (!existsSync(rules)) {
|
|
2277
|
+
try {
|
|
2278
|
+
writeFileSync(rules, MAGI_MD);
|
|
2279
|
+
if (ctx.mode === "tui") ctx.ui.notify("MAGI.md created: rules for local models, appended to the system prompt", "info");
|
|
2280
|
+
} catch {
|
|
2281
|
+
// read-only directory: the built-in rules are used as they are
|
|
2282
|
+
}
|
|
2283
|
+
}
|
|
2116
2284
|
void refreshUnit(ctx);
|
|
2117
2285
|
recountSession(ctx);
|
|
2118
2286
|
state.hasSmartCompact = pi.getCommands().some((c) => c.name.replace(/^\//, "") === "smart-compact");
|
|
@@ -2120,6 +2288,7 @@ export default function (pi: ExtensionAPI) {
|
|
|
2120
2288
|
ui.compact = cfg.ui?.compact ?? false;
|
|
2121
2289
|
ui.kwhPrice = cfg.ui?.kwhPrice;
|
|
2122
2290
|
ui.currency = cfg.ui?.currency === "USD" ? "USD" : "EUR";
|
|
2291
|
+
swap.diskWh = cfg.totalWh ?? 0;
|
|
2123
2292
|
if (ctx.mode !== "tui") return;
|
|
2124
2293
|
applyChrome(ctx);
|
|
2125
2294
|
// nothing is loaded at startup: a new session picks its MECHA unit, a resumed one shows whether its model is in VRAM
|
|
@@ -2164,6 +2333,12 @@ export default function (pi: ExtensionAPI) {
|
|
|
2164
2333
|
|
|
2165
2334
|
pi.on("before_provider_request", (event, ctx) => {
|
|
2166
2335
|
rememberPrefix(ctx.cwd, event.payload);
|
|
2336
|
+
// llama.cpp honours a per-request thinking budget only when llama-server runs without --reasoning-budget
|
|
2337
|
+
if (budgetMode === "off" || !swap.base || !ctx.model) return;
|
|
2338
|
+
const phase = requestPhase(event.payload);
|
|
2339
|
+
pendingBudget = { model: ctx.model.id, phase, tokens: budgetFor(ctx.model.id, phase) };
|
|
2340
|
+
const payload = { ...(event.payload as object), thinking_budget_tokens: pendingBudget.tokens };
|
|
2341
|
+
return budgetMessage ? { ...payload, reasoning_budget_message: `\n\n${BUDGET_MESSAGE}` } : payload;
|
|
2167
2342
|
});
|
|
2168
2343
|
|
|
2169
2344
|
pi.on("model_select", async (event, ctx) => {
|
|
@@ -2173,6 +2348,7 @@ export default function (pi: ExtensionAPI) {
|
|
|
2173
2348
|
});
|
|
2174
2349
|
|
|
2175
2350
|
pi.on("session_shutdown", async () => {
|
|
2351
|
+
persistEnergy();
|
|
2176
2352
|
prewarm.abort?.abort();
|
|
2177
2353
|
clearInterval(metricsTimer);
|
|
2178
2354
|
metricsTimer = undefined;
|
|
@@ -2187,6 +2363,40 @@ export default function (pi: ExtensionAPI) {
|
|
|
2187
2363
|
|
|
2188
2364
|
pi.on("agent_start", async () => {
|
|
2189
2365
|
state.runStart = Date.now();
|
|
2366
|
+
lastCall = "";
|
|
2367
|
+
});
|
|
2368
|
+
|
|
2369
|
+
pi.on("before_agent_start", async (event, ctx) => {
|
|
2370
|
+
let rules = MAGI_MD;
|
|
2371
|
+
try {
|
|
2372
|
+
rules = readFileSync(join(ctx.cwd, "MAGI.md"), "utf8");
|
|
2373
|
+
} catch {
|
|
2374
|
+
// no MAGI.md (deleted, or unwritable cwd): the built-in rules
|
|
2375
|
+
}
|
|
2376
|
+
return rules.trim() ? { systemPrompt: `${event.systemPrompt}\n\n${rules.trim()}` } : undefined;
|
|
2377
|
+
});
|
|
2378
|
+
|
|
2379
|
+
// before every request: drop old thinking and tool outputs from what the model sees, never from the session
|
|
2380
|
+
pi.on("context", async (event) => {
|
|
2381
|
+
if (!hygiene.enabled) return;
|
|
2382
|
+
const pruned = pruneContext(event.messages as any, hygiene);
|
|
2383
|
+
hygieneStats = pruned.stats;
|
|
2384
|
+
state.pruned = pruned.stats.prunedTokens;
|
|
2385
|
+
return { messages: pruned.messages as any };
|
|
2386
|
+
});
|
|
2387
|
+
|
|
2388
|
+
// local models loop on the same call, and may copy a pruned placeholder into a file
|
|
2389
|
+
pi.on("tool_call", async (event) => {
|
|
2390
|
+
const args = JSON.stringify(event.input);
|
|
2391
|
+
if ((event.toolName === "write" || event.toolName === "edit") && args.includes(PRUNED_MARK)) {
|
|
2392
|
+
return { block: true, reason: `MAGI: this ${event.toolName} contains a "${PRUNED_MARK}" placeholder, not real content. Read the file and use the actual text.` };
|
|
2393
|
+
}
|
|
2394
|
+
const fingerprint = event.toolName + args;
|
|
2395
|
+
repeats = fingerprint === lastCall ? repeats + 1 : 1;
|
|
2396
|
+
lastCall = fingerprint;
|
|
2397
|
+
if (repeats >= LOOP_REPEATS) {
|
|
2398
|
+
return { block: true, reason: `MAGI: identical ${event.toolName} call ${repeats} times in a row. Repeating it will not change the outcome: change approach, or tell the user what blocks you.` };
|
|
2399
|
+
}
|
|
2190
2400
|
});
|
|
2191
2401
|
|
|
2192
2402
|
pi.on("turn_start", async (_event, ctx) => {
|
|
@@ -2230,7 +2440,19 @@ export default function (pi: ExtensionAPI) {
|
|
|
2230
2440
|
pi.on("message_end", async (event) => {
|
|
2231
2441
|
if (event.message.role !== "assistant") return;
|
|
2232
2442
|
const m = event.message as AssistantMessage;
|
|
2233
|
-
|
|
2443
|
+
countAssistant(m);
|
|
2444
|
+
// learn how long this model thinks in this phase; a cut must not raise the budget
|
|
2445
|
+
const thought = thinkingTokens(m as any);
|
|
2446
|
+
if (pendingBudget && thought > 0 && m.stopReason !== "aborted" && m.stopReason !== "error") {
|
|
2447
|
+
const verdict = budgetVerdict(m as any, pendingBudget.tokens);
|
|
2448
|
+
if (verdict === "ignored") budgetIgnored.add(pendingBudget.model);
|
|
2449
|
+
if (verdict === "cut" && !thinkingWasCut(m as any)) state.thinkCuts++; // silent cut: no budget message on the server
|
|
2450
|
+
const samples = ((learned[pendingBudget.model] ??= {})[pendingBudget.phase] ??= []);
|
|
2451
|
+
samples.push(budgetSample(thought, pendingBudget.tokens, verdict === "cut"));
|
|
2452
|
+
samples.splice(0, samples.length - BUDGET_WINDOW);
|
|
2453
|
+
persistBudget();
|
|
2454
|
+
}
|
|
2455
|
+
pendingBudget = undefined;
|
|
2234
2456
|
if (perf.start) {
|
|
2235
2457
|
const end = Date.now();
|
|
2236
2458
|
perf.lastMs = end - perf.start;
|
|
@@ -2366,9 +2588,11 @@ export default function (pi: ExtensionAPI) {
|
|
|
2366
2588
|
}
|
|
2367
2589
|
|
|
2368
2590
|
pi.registerCommand("magi", {
|
|
2369
|
-
description: "Ask the three MAGI
|
|
2591
|
+
description: "Ask the three MAGI, or manage them: review [focus] · config · mecha · status · panel · compact · cost · on · off · hygiene · budget (type a space to see them all)",
|
|
2592
|
+
getArgumentCompletions: (prefix) => argCompletions(MAGI_ARGS, prefix),
|
|
2370
2593
|
handler: async (args, ctx) => {
|
|
2371
2594
|
const arg = args.trim();
|
|
2595
|
+
if (UI_ARGS.test(arg)) return manageUi(arg, ctx);
|
|
2372
2596
|
if (arg === "config") return configureMagi(ctx);
|
|
2373
2597
|
if (arg === "mecha") {
|
|
2374
2598
|
if (ctx.mode !== "tui" || !swap.base) return ctx.ui.notify("MECHA SELECT needs the TUI and a llama-swap model", "error");
|
|
@@ -2406,87 +2630,145 @@ export default function (pi: ExtensionAPI) {
|
|
|
2406
2630
|
},
|
|
2407
2631
|
});
|
|
2408
2632
|
|
|
2409
|
-
|
|
2410
|
-
|
|
2411
|
-
|
|
2412
|
-
|
|
2413
|
-
|
|
2414
|
-
|
|
2415
|
-
|
|
2416
|
-
|
|
2417
|
-
|
|
2418
|
-
|
|
2419
|
-
|
|
2633
|
+
/** /magi on|off|panel|compact|status|cost|hygiene|budget: the theme, its panel and the local-model settings. */
|
|
2634
|
+
async function manageUi(args: string, ctx: ExtensionCommandContext) {
|
|
2635
|
+
liveCtx = ctx;
|
|
2636
|
+
const arg = args.trim().toLowerCase();
|
|
2637
|
+
|
|
2638
|
+
if (arg === "hygiene" || arg.startsWith("hygiene ")) {
|
|
2639
|
+
const sub = arg.slice("hygiene".length).trim();
|
|
2640
|
+
const keep = /^(\d+)\s+(\d+)$/.exec(sub);
|
|
2641
|
+
const step = /^step\s+(\d+)(k?)$/.exec(sub);
|
|
2642
|
+
if (sub === "on" || sub === "off" || keep || step) {
|
|
2643
|
+
if (keep) Object.assign(hygiene, { enabled: true, keepThinkingTurns: Number(keep[1]), keepToolResults: Number(keep[2]) });
|
|
2644
|
+
else if (step) hygiene.stepTokens = Math.max(1000, Number(step[1]) * (step[2] ? 1000 : 1));
|
|
2645
|
+
else hygiene.enabled = sub === "on";
|
|
2646
|
+
const cfg = loadMagiConfig();
|
|
2647
|
+
const { enabled, keepThinkingTurns, keepToolResults, stepTokens } = hygiene;
|
|
2648
|
+
saveMagiConfig({ ...cfg, hygiene: { ...cfg.hygiene, enabled, keepThinkingTurns, keepToolResults, stepTokens } });
|
|
2649
|
+
} else if (sub) {
|
|
2650
|
+
ctx.ui.notify("Usage: /magi hygiene [on|off|step <tokens>|<thinking turns kept> <tool results kept>], e.g. step 40k, 3 5", "warning");
|
|
2420
2651
|
return;
|
|
2421
2652
|
}
|
|
2422
|
-
|
|
2423
|
-
|
|
2424
|
-
|
|
2425
|
-
|
|
2426
|
-
|
|
2427
|
-
|
|
2653
|
+
const s = hygieneStats;
|
|
2654
|
+
ctx.ui.notify(
|
|
2655
|
+
(!hygiene.enabled
|
|
2656
|
+
? "Context hygiene off: /magi hygiene on"
|
|
2657
|
+
: s
|
|
2658
|
+
? `Context hygiene: ${fmtTokens(s.prunedTokens)} tokens pruned in the first ${s.watermark}/${s.messages} messages, ${fmtTokens(s.pendingTokens)} waiting for the next step (every ${fmtTokens(hygiene.stepTokens)})`
|
|
2659
|
+
: "Context hygiene on: nothing sent to the model yet") +
|
|
2660
|
+
` · keeps thinking of the last ${hygiene.keepThinkingTurns} turns, the last ${hygiene.keepToolResults} tool results` +
|
|
2661
|
+
(state.thinkCuts ? ` · thinking cut at the budget ${state.thinkCuts}×` : ""),
|
|
2662
|
+
"info",
|
|
2663
|
+
);
|
|
2664
|
+
return;
|
|
2665
|
+
}
|
|
2666
|
+
if (arg === "budget" || arg.startsWith("budget ")) {
|
|
2667
|
+
const sub = arg.slice("budget".length).trim();
|
|
2668
|
+
const model = ctx.model?.id ?? "";
|
|
2669
|
+
const tokens = (t: string) => Math.round(Number(t.replace(/k$/, "")) * (t.endsWith("k") ? 1024 : 1));
|
|
2670
|
+
const fixed = /^(\d+k?)\s+(\d+k?)$/.exec(sub);
|
|
2671
|
+
if (sub === "auto" || sub === "off") budgetMode = sub;
|
|
2672
|
+
else if (sub === "reset") delete learned[model];
|
|
2673
|
+
else if (sub === "message") budgetMessage = !budgetMessage;
|
|
2674
|
+
else if (fixed) {
|
|
2675
|
+
budgetMode = "fixed";
|
|
2676
|
+
fixedBudget = { planning: tokens(fixed[1]!), acting: tokens(fixed[2]!) };
|
|
2677
|
+
} else if (sub) {
|
|
2678
|
+
ctx.ui.notify("Usage: /magi budget [auto|off|reset|message|<planning> <acting>], e.g. 16k 4k", "warning");
|
|
2428
2679
|
return;
|
|
2429
2680
|
}
|
|
2430
|
-
if (
|
|
2431
|
-
|
|
2432
|
-
|
|
2433
|
-
const
|
|
2434
|
-
|
|
2435
|
-
|
|
2436
|
-
|
|
2437
|
-
|
|
2438
|
-
|
|
2439
|
-
|
|
2440
|
-
|
|
2441
|
-
|
|
2442
|
-
|
|
2443
|
-
|
|
2444
|
-
|
|
2445
|
-
|
|
2446
|
-
|
|
2447
|
-
|
|
2448
|
-
|
|
2449
|
-
|
|
2450
|
-
|
|
2681
|
+
if (sub) persistBudget();
|
|
2682
|
+
const phase = (p: BudgetPhase) => {
|
|
2683
|
+
const n = learned[model]?.[p]?.length ?? 0;
|
|
2684
|
+
const how = budgetMode === "fixed" ? "fixed" : learnedBudget(learned[model]?.[p] ?? [], p) ? `learned from ${n}` : `default, learning ${n}/10`;
|
|
2685
|
+
return `${p} ${budgetFor(model, p) / 1024}k (${how})`;
|
|
2686
|
+
};
|
|
2687
|
+
ctx.ui.notify(
|
|
2688
|
+
budgetMode === "off"
|
|
2689
|
+
? "Thinking budget off: llama-server decides. /magi budget auto"
|
|
2690
|
+
: `Thinking budget ${budgetMode.toUpperCase()} · ${model || "no model"}: ${phase("planning")} · ${phase("acting")}` +
|
|
2691
|
+
` · closing message ${budgetMessage ? "on" : "off"}` +
|
|
2692
|
+
(swap.base ? "" : " · applies to llama-swap models only") +
|
|
2693
|
+
(budgetIgnored.has(model) ? " · ⚠ llama-server thinks past it: it was started with its own --reasoning-budget (or is too old), which wins" : ""),
|
|
2694
|
+
"info",
|
|
2695
|
+
);
|
|
2696
|
+
return;
|
|
2697
|
+
}
|
|
2698
|
+
if (arg === "panel") {
|
|
2699
|
+
panelEnabled = !panelEnabled;
|
|
2700
|
+
if (panelEnabled) showPanel(ctx.ui.theme);
|
|
2701
|
+
else hidePanel();
|
|
2702
|
+
ctx.ui.notify(`Side panel ${panelEnabled ? "enabled" : "disabled"}`, "info");
|
|
2703
|
+
return;
|
|
2704
|
+
}
|
|
2705
|
+
if (arg === "compact") {
|
|
2706
|
+
ui.compact = !ui.compact;
|
|
2707
|
+
const cfg = loadMagiConfig();
|
|
2708
|
+
saveMagiConfig({ ...cfg, ui: { ...cfg.ui, compact: ui.compact } });
|
|
2709
|
+
repaint();
|
|
2710
|
+
ctx.ui.notify(`Side panel ${ui.compact ? "compact" : "detailed"}`, "info");
|
|
2711
|
+
return;
|
|
2712
|
+
}
|
|
2713
|
+
if (arg === "cost") {
|
|
2714
|
+
const currency = await ctx.ui.select(`Currency for COST (current: ${ui.currency})`, ["EUR", "USD"]);
|
|
2715
|
+
if (!currency) return;
|
|
2716
|
+
const current = ui.kwhPrice !== undefined ? ` (current: ${ui.kwhPrice})` : "";
|
|
2717
|
+
const raw = (await ctx.ui.input(`Electricity price per kWh in ${currency}${current}:`, "0.30"))?.trim();
|
|
2718
|
+
if (raw === undefined) return;
|
|
2719
|
+
const price = raw === "" && ui.kwhPrice !== undefined ? ui.kwhPrice : Number(raw.replace(",", "."));
|
|
2720
|
+
if (raw === "" && ui.kwhPrice === undefined) {
|
|
2721
|
+
ctx.ui.notify("No price entered: COST unchanged", "warning");
|
|
2451
2722
|
return;
|
|
2452
2723
|
}
|
|
2453
|
-
if (
|
|
2454
|
-
|
|
2455
|
-
|
|
2456
|
-
|
|
2457
|
-
|
|
2458
|
-
|
|
2459
|
-
|
|
2460
|
-
|
|
2461
|
-
|
|
2462
|
-
|
|
2463
|
-
|
|
2464
|
-
|
|
2465
|
-
|
|
2466
|
-
|
|
2467
|
-
|
|
2468
|
-
return;
|
|
2469
|
-
}
|
|
2470
|
-
await ctx.ui.custom<void>((_tui, theme, _keys, done) =>
|
|
2471
|
-
buildReportView(theme, activityReport(theme, rows, report.total ?? rows.length), () => done(undefined)),
|
|
2472
|
-
);
|
|
2724
|
+
if (!Number.isFinite(price) || price < 0) {
|
|
2725
|
+
ctx.ui.notify(`Invalid price: "${raw}"`, "error");
|
|
2726
|
+
return;
|
|
2727
|
+
}
|
|
2728
|
+
ui.currency = currency === "USD" ? "USD" : "EUR";
|
|
2729
|
+
ui.kwhPrice = price;
|
|
2730
|
+
const cfg = loadMagiConfig();
|
|
2731
|
+
saveMagiConfig({ ...cfg, ui: { ...cfg.ui, currency: ui.currency, kwhPrice: price } });
|
|
2732
|
+
repaint();
|
|
2733
|
+
ctx.ui.notify(`COST: ${fmtMoney(price)} per kWh`, "info");
|
|
2734
|
+
return;
|
|
2735
|
+
}
|
|
2736
|
+
if (arg === "status") {
|
|
2737
|
+
if (!swap.base) {
|
|
2738
|
+
ctx.ui.notify("/magi status needs a llama-swap session model", "warning");
|
|
2473
2739
|
return;
|
|
2474
2740
|
}
|
|
2475
|
-
|
|
2476
|
-
|
|
2477
|
-
|
|
2478
|
-
|
|
2741
|
+
let report: { data?: ActivityRow[]; total?: number };
|
|
2742
|
+
try {
|
|
2743
|
+
report = (await (await swapGet(`/api/metrics/activity?limit=${ACTIVITY_REPORT_ROWS}`, 15_000)).json()) as typeof report;
|
|
2744
|
+
} catch (err) {
|
|
2745
|
+
ctx.ui.notify(`llama-swap status failed: ${err instanceof Error ? err.message : String(err)}`, "error");
|
|
2479
2746
|
return;
|
|
2480
2747
|
}
|
|
2481
|
-
|
|
2482
|
-
|
|
2483
|
-
|
|
2484
|
-
ctx.ui.notify(`Theme magi not found: ${res.error}`, "error");
|
|
2748
|
+
const rows = report.data ?? [];
|
|
2749
|
+
if (!rows.length) {
|
|
2750
|
+
ctx.ui.notify("llama-swap has no recorded requests yet", "info");
|
|
2485
2751
|
return;
|
|
2486
2752
|
}
|
|
2487
|
-
|
|
2753
|
+
await ctx.ui.custom<void>((_tui, theme, _keys, done) =>
|
|
2754
|
+
buildReportView(theme, activityReport(theme, rows, report.total ?? rows.length), () => done(undefined)),
|
|
2755
|
+
);
|
|
2756
|
+
return;
|
|
2757
|
+
}
|
|
2758
|
+
if (arg === "off") {
|
|
2759
|
+
chrome = false;
|
|
2488
2760
|
applyChrome(ctx);
|
|
2489
|
-
ctx.ui.notify("MAGI
|
|
2490
|
-
|
|
2491
|
-
|
|
2761
|
+
ctx.ui.notify("MAGI chrome disabled", "info");
|
|
2762
|
+
return;
|
|
2763
|
+
}
|
|
2764
|
+
|
|
2765
|
+
const res = ctx.ui.setTheme("magi");
|
|
2766
|
+
if (!res.success) {
|
|
2767
|
+
ctx.ui.notify(`Theme magi not found: ${res.error}`, "error");
|
|
2768
|
+
return;
|
|
2769
|
+
}
|
|
2770
|
+
chrome = true;
|
|
2771
|
+
applyChrome(ctx);
|
|
2772
|
+
ctx.ui.notify("MAGI online — /magi panel|compact|status|off", "info");
|
|
2773
|
+
}
|
|
2492
2774
|
}
|
|
@@ -0,0 +1,240 @@
|
|
|
1
|
+
/**
|
|
2
|
+
* Local-model helpers for MAGI: context hygiene, loop guard, and the MAGI.md rules template.
|
|
3
|
+
*
|
|
4
|
+
* Context hygiene rewrites the messages pi sends to the model (never the saved session):
|
|
5
|
+
* - old thinking blocks are dropped: Qwen templates resend every reasoning block of an agent run,
|
|
6
|
+
* one 18k-token think stays in the context until compaction
|
|
7
|
+
* - old tool outputs become a one-line marker (observation masking, Lindenbauer et al. 2025:
|
|
8
|
+
* as good as LLM summarization, at half the cost)
|
|
9
|
+
* - old write/edit payloads become a marker too: the file on disk is the source of truth
|
|
10
|
+
* Pruning advances in steps behind a watermark, so between steps the prompt only grows at the end
|
|
11
|
+
* and llama.cpp keeps reusing its KV cache.
|
|
12
|
+
*/
|
|
13
|
+
|
|
14
|
+
export type Msg = { role: string; content?: any; [key: string]: any };
|
|
15
|
+
|
|
16
|
+
export interface HygieneOptions {
|
|
17
|
+
keepThinkingTurns: number; // newest assistant messages that keep their thinking
|
|
18
|
+
keepToolResults: number; // newest tool results kept whole
|
|
19
|
+
stepTokens: number; // the watermark moves each time the prunable total crosses another multiple of this
|
|
20
|
+
minPruneChars: number; // smaller outputs/payloads are left alone
|
|
21
|
+
}
|
|
22
|
+
|
|
23
|
+
export interface HygieneStats {
|
|
24
|
+
messages: number;
|
|
25
|
+
watermark: number; // messages before this index are pruned
|
|
26
|
+
prunedTokens: number;
|
|
27
|
+
pendingTokens: number; // prunable, waiting for the next step
|
|
28
|
+
stepTokens: number;
|
|
29
|
+
}
|
|
30
|
+
|
|
31
|
+
export const HYGIENE_DEFAULTS: HygieneOptions = { keepThinkingTurns: 3, keepToolResults: 5, stepTokens: 15_000, minPruneChars: 600 };
|
|
32
|
+
|
|
33
|
+
// ponytail: chars/token measured on Qwen3.x pi sessions (2.9–3.2); a tokenizer call would be exact but costs a request
|
|
34
|
+
const CHARS_PER_TOKEN = 3;
|
|
35
|
+
const IMAGE_CHARS = 4800;
|
|
36
|
+
|
|
37
|
+
/** Marks every placeholder, so a model that copies one into a real write/edit gets blocked (see loop guard). */
|
|
38
|
+
export const PRUNED_MARK = "<<pruned by MAGI";
|
|
39
|
+
|
|
40
|
+
const lines = (s: string) => s.split("\n").length;
|
|
41
|
+
const clip = (s: string, n = 80) => (s.length > n ? s.slice(0, n - 1) + "…" : s);
|
|
42
|
+
|
|
43
|
+
function callTarget(call: any): string {
|
|
44
|
+
const a = call?.arguments ?? {};
|
|
45
|
+
return clip(String(a.path ?? a.command ?? a.pattern ?? a.query ?? a.url ?? ""));
|
|
46
|
+
}
|
|
47
|
+
|
|
48
|
+
function resultChars(m: Msg): number {
|
|
49
|
+
let n = 0;
|
|
50
|
+
for (const c of m.content ?? []) n += c.type === "image" ? IMAGE_CHARS : (c.text?.length ?? 0);
|
|
51
|
+
return n;
|
|
52
|
+
}
|
|
53
|
+
|
|
54
|
+
function editChars(call: any): number {
|
|
55
|
+
return (call.arguments?.edits ?? []).reduce((n: number, e: any) => n + (e.oldText?.length ?? 0) + (e.newText?.length ?? 0), 0);
|
|
56
|
+
}
|
|
57
|
+
|
|
58
|
+
/** Chars pruning would remove from this assistant message. */
|
|
59
|
+
function assistantSavings(m: Msg, o: HygieneOptions): number {
|
|
60
|
+
let n = 0;
|
|
61
|
+
for (const c of m.content ?? []) {
|
|
62
|
+
if (c.type === "thinking") n += c.thinking?.length ?? 0;
|
|
63
|
+
else if (c.type === "toolCall" && c.name === "write" && (c.arguments?.content?.length ?? 0) >= o.minPruneChars) n += c.arguments.content.length;
|
|
64
|
+
else if (c.type === "toolCall" && c.name === "edit" && editChars(c) >= o.minPruneChars) n += editChars(c);
|
|
65
|
+
}
|
|
66
|
+
return n;
|
|
67
|
+
}
|
|
68
|
+
|
|
69
|
+
function pruneAssistant(m: Msg, o: HygieneOptions): Msg {
|
|
70
|
+
const content = [];
|
|
71
|
+
for (const c of m.content ?? []) {
|
|
72
|
+
if (c.type === "thinking") continue;
|
|
73
|
+
if (c.type === "toolCall" && c.name === "write" && (c.arguments?.content?.length ?? 0) >= o.minPruneChars) {
|
|
74
|
+
const text = c.arguments.content as string;
|
|
75
|
+
content.push({ ...c, arguments: { ...c.arguments, content: `${PRUNED_MARK}: ${lines(text)} lines written, the file on disk is the source of truth>>` } });
|
|
76
|
+
} else if (c.type === "toolCall" && c.name === "edit" && editChars(c) >= o.minPruneChars) {
|
|
77
|
+
const edits = c.arguments.edits.map((e: any) => ({ oldText: `${PRUNED_MARK}>>`, newText: `${PRUNED_MARK}: ${lines(e.newText ?? "")} lines>>` }));
|
|
78
|
+
content.push({ ...c, arguments: { ...c.arguments, edits } });
|
|
79
|
+
} else content.push(c);
|
|
80
|
+
}
|
|
81
|
+
// a message that only held thinking keeps a marker: some templates choke on an empty assistant turn
|
|
82
|
+
if (!content.length) content.push({ type: "text", text: `${PRUNED_MARK}: reasoning>>` });
|
|
83
|
+
return { ...m, content };
|
|
84
|
+
}
|
|
85
|
+
|
|
86
|
+
function pruneResult(m: Msg, call: any): Msg {
|
|
87
|
+
const text = (m.content ?? []).map((c: any) => c.text ?? "").join("\n");
|
|
88
|
+
const what = [m.toolName, callTarget(call)].filter(Boolean).join(" ");
|
|
89
|
+
return { ...m, content: [{ type: "text", text: `${PRUNED_MARK}: output of ${what}, ${lines(text)} lines. Run it again if you need it.>>` }] };
|
|
90
|
+
}
|
|
91
|
+
|
|
92
|
+
/**
|
|
93
|
+
* Prunes everything before the watermark. The watermark sits on the last point where the prunable chars,
|
|
94
|
+
* summed from the first message, crossed a multiple of stepTokens. Only messages before the protected tail
|
|
95
|
+
* count, and the tail only moves forward, so earlier crossings never move: the pruned prefix is stable.
|
|
96
|
+
*/
|
|
97
|
+
export function pruneContext(messages: Msg[], o: HygieneOptions = HYGIENE_DEFAULTS): { messages: Msg[]; stats: HygieneStats } {
|
|
98
|
+
const lastIndices = (role: string, keep: number) =>
|
|
99
|
+
messages.flatMap((m, i) => (m.role === role ? [i] : [])).slice(-keep);
|
|
100
|
+
const tail = [...lastIndices("assistant", o.keepThinkingTurns), ...lastIndices("toolResult", o.keepToolResults)];
|
|
101
|
+
const tailStart = tail.length ? Math.min(...tail) : messages.length;
|
|
102
|
+
|
|
103
|
+
const calls = new Map<string, any>();
|
|
104
|
+
for (const m of messages) if (m.role === "assistant") for (const c of m.content ?? []) if (c.type === "toolCall") calls.set(c.id, c);
|
|
105
|
+
|
|
106
|
+
const savings = (m: Msg) =>
|
|
107
|
+
m.role === "assistant" ? assistantSavings(m, o) : m.role === "toolResult" && resultChars(m) >= o.minPruneChars ? resultChars(m) : 0;
|
|
108
|
+
|
|
109
|
+
const step = o.stepTokens * CHARS_PER_TOKEN;
|
|
110
|
+
let sum = 0;
|
|
111
|
+
let watermark = 0;
|
|
112
|
+
let prunedChars = 0;
|
|
113
|
+
for (let i = 0; i < tailStart; i++) {
|
|
114
|
+
sum += savings(messages[i]!);
|
|
115
|
+
if (Math.floor(sum / step) > Math.floor(prunedChars / step)) {
|
|
116
|
+
watermark = i + 1;
|
|
117
|
+
prunedChars = sum;
|
|
118
|
+
}
|
|
119
|
+
}
|
|
120
|
+
|
|
121
|
+
const out = messages.map((m, i) => {
|
|
122
|
+
if (i >= watermark || !savings(m)) return m;
|
|
123
|
+
return m.role === "assistant" ? pruneAssistant(m, o) : pruneResult(m, calls.get(m.toolCallId));
|
|
124
|
+
});
|
|
125
|
+
return {
|
|
126
|
+
messages: out,
|
|
127
|
+
stats: {
|
|
128
|
+
messages: messages.length,
|
|
129
|
+
watermark,
|
|
130
|
+
prunedTokens: Math.round(prunedChars / CHARS_PER_TOKEN),
|
|
131
|
+
pendingTokens: Math.round((sum - prunedChars) / CHARS_PER_TOKEN),
|
|
132
|
+
stepTokens: o.stepTokens,
|
|
133
|
+
},
|
|
134
|
+
};
|
|
135
|
+
}
|
|
136
|
+
|
|
137
|
+
/**
|
|
138
|
+
* Sent as reasoning_budget_message with every request: when the thinking budget runs out llama.cpp forces this text,
|
|
139
|
+
* then the end-of-thinking tag, and the model reads it as its own words, so it says what to do next.
|
|
140
|
+
* A llama-server older than the per-request message closes the thinking silently; cuts are then detected by length (budgetVerdict).
|
|
141
|
+
*/
|
|
142
|
+
export const BUDGET_MESSAGE = "Time is up. I will take the smallest safe next step with what I know, and write my open plan into PLAN.md.";
|
|
143
|
+
|
|
144
|
+
export type BudgetPhase = "planning" | "acting"; // right after the user spoke / between tool calls
|
|
145
|
+
export type BudgetMode = "auto" | "fixed" | "off";
|
|
146
|
+
|
|
147
|
+
/** Starting budgets, used until a model has enough samples to learn its own. */
|
|
148
|
+
export const BUDGET_DEFAULTS: Record<BudgetPhase, number> = { planning: 16384, acting: 4096 };
|
|
149
|
+
export const BUDGET_LIMITS: Record<BudgetPhase, [number, number]> = { planning: [4096, 32768], acting: [2048, 4096] }; // acting: past ~4k between tool calls it is overthinking
|
|
150
|
+
|
|
151
|
+
const BUDGET_SAMPLES_MIN = 10;
|
|
152
|
+
export const BUDGET_WINDOW = 30; // thinking lengths kept per model and phase
|
|
153
|
+
const BUDGET_PERCENTILE = 0.95;
|
|
154
|
+
const BUDGET_HEADROOM = 1.5;
|
|
155
|
+
|
|
156
|
+
/** Phase of the next request, from the OpenAI-style payload: a tool result last means the agent is mid-task. */
|
|
157
|
+
export function requestPhase(payload: any): BudgetPhase {
|
|
158
|
+
return payload?.messages?.at(-1)?.role === "tool" ? "acting" : "planning";
|
|
159
|
+
}
|
|
160
|
+
|
|
161
|
+
/**
|
|
162
|
+
* Learned budget: 1.5 × the 95th percentile of recent thinking lengths, clamped and rounded to 1k.
|
|
163
|
+
* Cuts never raise it (see budgetSample): a model that often runs away is held at its budget instead of chased.
|
|
164
|
+
* Replies that finish close under the budget raise it, a model that thinks little pulls it down.
|
|
165
|
+
* Undefined until there are enough samples.
|
|
166
|
+
*/
|
|
167
|
+
export function learnedBudget(samples: number[], phase: BudgetPhase): number | undefined {
|
|
168
|
+
if (samples.length < BUDGET_SAMPLES_MIN) return undefined;
|
|
169
|
+
const sorted = [...samples].sort((a, b) => a - b);
|
|
170
|
+
const p = sorted[Math.min(sorted.length - 1, Math.floor(sorted.length * BUDGET_PERCENTILE))]!;
|
|
171
|
+
const [lo, hi] = BUDGET_LIMITS[phase];
|
|
172
|
+
return Math.min(hi, Math.max(lo, Math.ceil((p * BUDGET_HEADROOM) / 1024) * 1024));
|
|
173
|
+
}
|
|
174
|
+
|
|
175
|
+
/** The sample to learn from a reply: a cut counts as budget / headroom, so cuts alone give back the same budget. */
|
|
176
|
+
export function budgetSample(thought: number, budget: number, cut: boolean): number {
|
|
177
|
+
return cut ? budget / BUDGET_HEADROOM : thought;
|
|
178
|
+
}
|
|
179
|
+
|
|
180
|
+
/**
|
|
181
|
+
* What the server did with the budget sent for this reply, from the thinking length (estimated, ±10%):
|
|
182
|
+
* cut near the budget, "ignored" well past it (llama-server started with its own --reasoning-budget, or too old
|
|
183
|
+
* to read thinking_budget_tokens), otherwise within.
|
|
184
|
+
*/
|
|
185
|
+
export function budgetVerdict(m: Msg, budget: number): "cut" | "ignored" | "within" {
|
|
186
|
+
if (thinkingWasCut(m)) return "cut";
|
|
187
|
+
const thought = thinkingTokens(m);
|
|
188
|
+
return thought > budget * 1.3 ? "ignored" : thought >= budget * 0.9 ? "cut" : "within";
|
|
189
|
+
}
|
|
190
|
+
|
|
191
|
+
/** Estimated thinking tokens of a reply (same chars/token as the hygiene). */
|
|
192
|
+
export function thinkingTokens(m: Msg): number {
|
|
193
|
+
const chars = (m.content ?? []).reduce((n: number, c: any) => n + (c.type === "thinking" ? (c.thinking?.length ?? 0) : 0), 0);
|
|
194
|
+
return Math.round(chars / CHARS_PER_TOKEN);
|
|
195
|
+
}
|
|
196
|
+
|
|
197
|
+
/** Whether llama.cpp cut this message's thinking at the budget. */
|
|
198
|
+
export function thinkingWasCut(m: Msg): boolean {
|
|
199
|
+
return (m.content ?? []).some((c: any) => c.type === "thinking" && (c.thinking ?? "").trimEnd().endsWith(BUDGET_MESSAGE));
|
|
200
|
+
}
|
|
201
|
+
|
|
202
|
+
/**
|
|
203
|
+
* Written to <project>/MAGI.md when missing, then appended to the system prompt on every run.
|
|
204
|
+
* English on purpose: Qwen-family models reason in English and follow English rules more reliably.
|
|
205
|
+
* Kept short: it is paid for on every request.
|
|
206
|
+
*/
|
|
207
|
+
export const MAGI_MD = `# MAGI rules for local models
|
|
208
|
+
|
|
209
|
+
These rules are appended to the system prompt by the MAGI extension. Edit them freely; delete the file to get the defaults back.
|
|
210
|
+
|
|
211
|
+
## Think less, act more
|
|
212
|
+
- Keep reasoning short: understand the step, decide, act. Do not re-plan what is already decided.
|
|
213
|
+
- Never draft code or file contents in your reasoning. Write them directly with the write/edit tool.
|
|
214
|
+
- If two attempts at the same approach fail, stop and change approach, or ask the user. Do not retry blindly.
|
|
215
|
+
- If your reasoning stopped abruptly mid-thought, or ends with "${BUDGET_MESSAGE.split(". ")[0]}.", your thinking budget ran out: in that turn make no large or irreversible change. Write your open plan into PLAN.md or take one small step you can verify; the next turn gives you a fresh budget.
|
|
216
|
+
|
|
217
|
+
## Your context is small: spend it carefully
|
|
218
|
+
- Search before reading: use rg/find to locate code, then read only the needed range (offset/limit).
|
|
219
|
+
- Do not read whole large files, and do not re-read a file you just wrote.
|
|
220
|
+
- Trim long command output: pipe through tail, head or rg.
|
|
221
|
+
- Old tool outputs and old reasoning are removed from your context automatically (marked "${PRUNED_MARK}…>>"). Anything you will need later must be written down (see below). Never copy those markers into a file.
|
|
222
|
+
|
|
223
|
+
## Keep the task state on disk
|
|
224
|
+
- For any task longer than a few steps, keep PLAN.md (a checklist) and NOTES.md (findings, decisions, commands that work).
|
|
225
|
+
- Update them after every completed step. If the context is compacted or a new session starts, continue from PLAN.md.
|
|
226
|
+
|
|
227
|
+
## Edit safely
|
|
228
|
+
- edit oldText must match the file exactly: copy it from a fresh read, keep it small and unique.
|
|
229
|
+
- If an edit fails, read that region again instead of guessing.
|
|
230
|
+
- For big new files: write a skeleton first, then fill it with edits. Never replace a whole file with a partial version.
|
|
231
|
+
|
|
232
|
+
## Verify, never assume
|
|
233
|
+
- Never invent paths, functions, APIs or flags: check with ls, rg, --help or the docs first.
|
|
234
|
+
- After a change, run the build, the tests or the program. Claim success only when a tool result proves it.
|
|
235
|
+
- If you cannot verify something, say so explicitly.
|
|
236
|
+
|
|
237
|
+
## Finish cleanly
|
|
238
|
+
- When done: say briefly what changed, how you verified it, and what is left.
|
|
239
|
+
- Answer in the user's language.
|
|
240
|
+
`;
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "pi-magi-theme",
|
|
3
|
-
"version": "0.
|
|
3
|
+
"version": "0.3.0",
|
|
4
4
|
"description": "MAGI SYSTEM theme + extension for pi (Evangelion fan art): MAGI control screen panel, three-model /magi council, MECHA SELECT model picker, angel-attack loading, seven-seal context gauge, llama-swap telemetry",
|
|
5
5
|
"keywords": [
|
|
6
6
|
"pi-package",
|