pi-magi-theme 0.2.4 → 0.3.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +49 -10
- package/extensions/magi/index.ts +361 -93
- package/extensions/magi/local-models.ts +240 -0
- package/package.json +1 -1
package/README.md
CHANGED
|
@@ -36,14 +36,18 @@ Clone the repo and point pi at it instead (edits in the repo are live on the nex
|
|
|
36
36
|
|
|
37
37
|
| Command | What it does |
|
|
38
38
|
|---------|--------------|
|
|
39
|
-
| `/magi <question>` | the council answers a question (recent conversation as context) |
|
|
40
|
-
| `/magi review [focus]` | the council reviews your pending changes (`git diff HEAD` plus untracked file names) before you commit |
|
|
39
|
+
| `/magi council <question>` | the council answers a question (recent conversation as context) |
|
|
40
|
+
| `/magi council review [focus]` | the council reviews your pending changes (`git diff HEAD` plus untracked file names) before you commit |
|
|
41
41
|
| `/magi config` | pick a model for each MAGI |
|
|
42
42
|
| `/magi mecha` | MECHA SELECT: pick the llama-swap model to activate, each shown as a mecha head lit by its real state |
|
|
43
|
-
| `/magi
|
|
44
|
-
| `/magi
|
|
45
|
-
| `/magi
|
|
46
|
-
| `/magi
|
|
43
|
+
| `/magi compact` | toggle the compact side panel (basic info and animations only); remembered across sessions |
|
|
44
|
+
| `/magi status` | llama-swap report from its last 100 requests: speed, tokens, cache hits, MTP draft acceptance, durations, errors per model |
|
|
45
|
+
| `/magi cost` | set the electricity price per kWh and the currency (EUR or USD) for the COST row |
|
|
46
|
+
| `/magi panel` · `on` · `off` | hide/show the side panel, enable/disable the whole chrome |
|
|
47
|
+
| `/magi hygiene` · `on` · `off` · `step <tokens>` · `<turns> <results>` | show how much context the hygiene pruned; enable/disable it; how many prunable tokens make a pruning step (e.g. `step 40k`, default 15k); how many recent turns keep their thinking and how many tool results stay whole (e.g. `3 5`) |
|
|
48
|
+
| `/magi budget` · `auto` · `off` · `reset` · `message` · `<planning> <acting>` | show the thinking budget of the current model; learn it per model (default); leave it to llama-server; forget what was learned for this model; turn the closing message off/on; or fix it (e.g. `16k 4k`) |
|
|
49
|
+
|
|
50
|
+
Everything lives under `/magi`: type `/magi ` (with the space) to see every option with a short description; keep typing to narrow it down, Tab or Enter to pick one. Anything that is not an option is a question for the council.
|
|
47
51
|
|
|
48
52
|
## Lore ↔ function
|
|
49
53
|
|
|
@@ -84,9 +88,39 @@ pi install npm:pi-smart-compact
|
|
|
84
88
|
|
|
85
89
|
It extracts files, errors, decisions and open loops locally (no LLM calls), then synthesizes and verifies the summary. Point its `summaryModel` at a local model to keep compaction free. When it is installed, the sixth seal suggests `/smart-compact`, the seals name it while they break (`✶ BREAKING THE SEALS · smart-compact · 4s`) and the seventh seal reports who actually produced the summary: `smart-compact`, or `pi native` if it fell back to pi's own compactor.
|
|
86
90
|
|
|
91
|
+
## Local models: context hygiene, thinking budget, loop guard, MAGI.md
|
|
92
|
+
|
|
93
|
+
Local models run out of context on long tasks well before they run out of work. They also tend to think for minutes between two tool calls and to repeat the same command when stuck. MAGI works on all three, automatically (the thinking budget for llama-swap models, the rest for any model):
|
|
94
|
+
|
|
95
|
+
- **Context hygiene: the model's memory stays lean.** *Problem:* every file the agent reads and every long reasoning stays in the conversation, until the model's context is full and the task falls apart. *What MAGI does:* before each request it replaces old reasoning, old tool outputs and old file writes with a one-line note (`<<pruned by MAGI…>>`). Only what the model sees is trimmed: your saved session stays complete. The latest 3 turns keep their reasoning and the latest 5 tool outputs stay whole. On two real sessions it brought 104k and 119k tokens down to ~64k and ~47k.
|
|
96
|
+
- **Thinking budget: no more ten-minute thinks.** *Problem:* a local model can reason for thousands of tokens before a simple step, and at 10 tokens/s that is minutes of waiting. *What MAGI does:* it gives the model a maximum length of thinking on every request: larger right after you write (planning), smaller between tool calls (acting). When the limit is reached the model is stopped mid-thought, says *"Time is up. I will take the smallest safe next step…"* and acts. The limit adapts to each model on its own. The side panel counts these cuts (`HYGIENE -18.2k ✂2`).
|
|
97
|
+
- **Loop guard: no endless retries.** *Problem:* a stuck model runs the same command again and again. *What MAGI does:* the third identical tool call in a row is blocked, with a message asking the model to try something else.
|
|
98
|
+
- **MAGI.md: house rules for the model.** *Problem:* local models repeat the same mistakes: reading whole files, inventing paths, claiming success without checking. *What MAGI does:* it creates `MAGI.md` in your project on the first start (never overwritten) and adds it to the model's instructions on every run: short rules against these mistakes, plus the habit of keeping the task plan in `PLAN.md` and findings in `NOTES.md`, so nothing important is lost when old context is trimmed. Edit it per project; delete it to get the defaults back.
|
|
99
|
+
|
|
100
|
+
Nothing needs setting up in llama-swap. The defaults suit most tasks; two adjustments are worth knowing:
|
|
101
|
+
|
|
102
|
+
- **Long tasks: `/magi hygiene step 40k`.** Each time the hygiene trims, the server has to re-read part of the conversation, and on some models (see *Hybrid models* below) almost all of it, which can take a few minutes. A step of 40k trims less often: on a real session it cut the re-reading from ~6 minutes to ~1.
|
|
103
|
+
- **A model that really needs to think longer: `/magi budget <planning> <acting>`**, e.g. `/magi budget 16k 8k`. Frequent ✂ cuts in the side panel are the sign.
|
|
104
|
+
|
|
105
|
+
### How it works
|
|
106
|
+
|
|
107
|
+
For the curious, and for tuning.
|
|
108
|
+
|
|
109
|
+
**Hygiene.** Trimming happens in steps, not on every request: the conversation up to a mark is trimmed, and the mark only moves forward once ~15k more tokens (the step) could be trimmed. Between steps the conversation only grows at the end, so llama.cpp reuses what it already processed (its KV cache) and only reads the new messages. Each step changes the conversation from the first newly trimmed message on, and the server re-reads from there: that is why steps are large and rare. Old tool outputs are replaced, not summarized: [simple observation masking matches LLM summarization at half the cost](https://arxiv.org/abs/2508.21433).
|
|
110
|
+
|
|
111
|
+
**Hybrid models.** Some models (e.g. Qwen3.8 Flash Next; dense or MoE does not matter) mix a few classic attention layers with recurrent ones, which squeeze the whole conversation into a fixed-size state instead of keeping each token. The server cannot rewind that state to an arbitrary point: it can only restore a saved copy (a checkpoint) and re-read from there, and the only copy before the trimmed part is usually the end of the system prompt. So on these models each step re-reads almost the whole prompt. With 10k-token thoughts and 4k-token file reads, a 15k step is crossed every 2–3 turns: on a real 70k-token session that was 3 re-reads in 6 requests (~6 min at 185 tokens/s); with a 40k step, 1 re-read (~1 min) for a prompt at most 6k larger. A model is hybrid if its llama-server log shows `restored context checkpoint` lines.
|
|
112
|
+
|
|
113
|
+
**Thinking budget.** MAGI sends the limit with every request (`thinking_budget_tokens`), and llama.cpp applies it. It learns per model and per phase: 1.5 × the 95th percentile of the model's last 30 thinking lengths, rounded up to 1k. A cut is recorded as 1/1.5 of the limit it hit, so cuts never raise the limit: a model that often runs away is held, not chased. Replies that end close under the limit raise it; a model that thinks little lowers it. Limits: planning 4k–32k, acting 2k–4k (past ~4k between two tool calls it is overthinking). Until a model has 10 replies in a phase it uses 16k / 4k.
|
|
114
|
+
|
|
115
|
+
When the limit is reached llama.cpp does not abort the reply: it inserts the closing sentence and the end-of-thinking tag, and the model goes on to act. MAGI sends that sentence with every request too (`reasoning_budget_message`; `/magi budget message` turns it off and on). `MAGI.md` tells the model what to do after a cut (one small verifiable step, the open plan into `PLAN.md`), and the next turn gets a fresh budget.
|
|
116
|
+
|
|
117
|
+
**llama-server versions.** The per-request limit works whenever llama-server was started without `--reasoning-budget`, which is the default. If it was started with one, the server's limit wins: MAGI notices the model thinking well past its own limit and `/magi budget` says so. A llama-server too old for the closing sentence ends the thinking silently, and MAGI detects the cut by its length.
|
|
118
|
+
|
|
119
|
+
**Loop guard details.** A tool call counts as identical when both the tool and its arguments match. A file write or edit that copies a `<<pruned by MAGI…>>` note into a file is blocked too.
|
|
120
|
+
|
|
87
121
|
## The council
|
|
88
122
|
|
|
89
|
-
`/magi <question>` asks three models in parallel, each with its own nature, then shows the votes and a majority verdict:
|
|
123
|
+
`/magi council <question>` asks three models in parallel, each with its own nature, then shows the votes and a majority verdict:
|
|
90
124
|
|
|
91
125
|
| Unit | Nature | Looks at |
|
|
92
126
|
|------|--------|----------|
|
|
@@ -96,18 +130,20 @@ It extracts files, errors, decisions and open loops locally (no LLM calls), then
|
|
|
96
130
|
|
|
97
131
|
Each nature is a lens, not a specialty, so the council answers any question, not only software ones. Every MAGI first answers the question, then judges it through its lens, naming concrete tools, numbers and scenarios from your question instead of generic advice. Votes: **APPROVE** = go ahead or clear recommendation; **CONDITIONAL** = only if the named conditions hold, or when information is missing (it says what it needs); **REJECT** = a concrete problem, with what to do instead. A MAGI never rejects because a topic is outside its nature. Answers come back in the language of your question.
|
|
98
132
|
|
|
99
|
-
`/magi <question>` gives the MAGI the recent conversation as context; `/magi review` gives them the pending diff (truncated at 24k characters). Full opinions are added to the chat (not sent to the agent), and the last verdict stays under the MAGI in the side panel.
|
|
133
|
+
`/magi council <question>` gives the MAGI the recent conversation as context; `/magi council review` gives them the pending diff (truncated at 24k characters). Full opinions are added to the chat (not sent to the agent), and the last verdict stays under the MAGI in the side panel.
|
|
100
134
|
|
|
101
135
|
## Configuration
|
|
102
136
|
|
|
103
|
-
`~/.pi/agent/magi.json` (written by `/magi config`, `/magi
|
|
137
|
+
`~/.pi/agent/magi.json` (written by `/magi config`, `/magi compact` and `/magi cost`, editable by hand):
|
|
104
138
|
|
|
105
139
|
```json
|
|
106
140
|
{
|
|
107
141
|
"MELCHIOR": { "model": "llama-swap/Qwen3.8 27B Q4_K_M - Thinking", "thinking": "low" },
|
|
108
142
|
"ui": { "compact": false, "kwhPrice": 0.30, "currency": "EUR" },
|
|
109
143
|
"loads": { "qwen3.8-27b": 41200 },
|
|
110
|
-
"totalWh": 1843.2
|
|
144
|
+
"totalWh": 1843.2,
|
|
145
|
+
"hygiene": { "enabled": true, "keepThinkingTurns": 3, "keepToolResults": 5, "stepTokens": 15000, "minPruneChars": 600 },
|
|
146
|
+
"thinkingBudget": { "mode": "auto", "planning": 16384, "acting": 4096, "message": true, "learned": { "qwen3.8-27b": { "acting": [812, 430, 2211] } } }
|
|
111
147
|
}
|
|
112
148
|
```
|
|
113
149
|
|
|
@@ -115,6 +151,7 @@ Each nature is a lens, not a specialty, so the council answers any question, not
|
|
|
115
151
|
- `ui.compact`: start with the compact side panel;
|
|
116
152
|
- `ui.kwhPrice` and `ui.currency` (`EUR` or `USD`): the COST row multiplies the GPU energy by this price, showing the running total of every session with the current one in brackets;
|
|
117
153
|
- `totalWh`: written by the theme, GPU energy summed over every session (delete the key to reset the COST total);
|
|
154
|
+
- `hygiene` and `thinkingBudget` are set with `/magi hygiene` and `/magi budget` (`hygiene.minPruneChars` by hand only); `thinkingBudget.message` sends the closing message with every request (default `true`); `thinkingBudget.learned` is written by the theme (recent thinking lengths per model and phase);
|
|
118
155
|
- `loads`: written by the theme, how long each llama-swap model took to load last time (paces the angel attack; 60s when unknown).
|
|
119
156
|
|
|
120
157
|
## Release
|
|
@@ -129,6 +166,8 @@ No token: npmjs is configured to trust this repository's `publish.yml` (npm trus
|
|
|
129
166
|
|
|
130
167
|
## llama-swap
|
|
131
168
|
|
|
169
|
+
**Thinking levels.** [pi-llama-swap](https://www.npmjs.com/package/@danielmeneses/pi-llama-swap) registers every model with reasoning off, so `/thinking` only offers `off`. At session start MAGI reads the reasoning levels llama-swap publishes in `/v1/models` (`meta.llamaswap.reasoning.levels`) and re-registers those models with thinking on: `/thinking` then offers exactly those levels, sent as `chat_template_kwargs` (`enable_thinking`, `reasoning_effort`). Aliases (e.g. the Instruct twin of a Thinking model) stay off, and image input follows `architecture.input_modalities`. A new or renamed model works without a `modelOverrides` entry in `models.json`; entries you keep there still apply on top (e.g. `samplingParams`). The starting level is pi's usual one for a model switch: the level saved for that model (`Ctrl+S` in `/thinking`), else `defaultThinkingLevel`.
|
|
170
|
+
|
|
132
171
|
When the session model uses the `llama-swap` provider, the side panel:
|
|
133
172
|
|
|
134
173
|
- loads nothing at startup: a new session opens MECHA SELECT, a resumed one shows whether its model is already in VRAM;
|
package/extensions/magi/index.ts
CHANGED
|
@@ -17,22 +17,41 @@
|
|
|
17
17
|
* - while a model loads an angel attacks the MAGI: red spreads through BALTHASAR, MELCHIOR and CASPAR at the pace of
|
|
18
18
|
* the model's last load, a corner of CASPAR holds out blinking; once loaded, blue takes the MAGI back from that corner
|
|
19
19
|
* - llama-swap telemetry: VRAM, GPU load/temp/power, energy used, RAM, server-side tok/s, prompt tok/s, cache hits
|
|
20
|
-
* - /magi config → assign a model to each MAGI; /magi
|
|
20
|
+
* - /magi config → assign a model to each MAGI; /magi compact|status → smaller panel, llama-swap report
|
|
21
21
|
*
|
|
22
22
|
* Fan art: the MAGI and their screen come from Neon Genesis Evangelion, all rights reserved to khara, Inc.
|
|
23
23
|
* Use with the theme ../../themes/magi.json
|
|
24
24
|
*/
|
|
25
25
|
|
|
26
26
|
import { execFile } from "node:child_process";
|
|
27
|
-
import { readFileSync, writeFileSync } from "node:fs";
|
|
27
|
+
import { existsSync, readFileSync, writeFileSync } from "node:fs";
|
|
28
28
|
import { homedir } from "node:os";
|
|
29
29
|
import { join } from "node:path";
|
|
30
30
|
import { promisify } from "node:util";
|
|
31
31
|
import type { AssistantMessage, Model } from "@earendil-works/pi-ai";
|
|
32
32
|
import { completeSimple } from "@earendil-works/pi-ai";
|
|
33
|
-
import type { ExtensionAPI, ExtensionContext, Theme, ThemeColor } from "@earendil-works/pi-coding-agent";
|
|
33
|
+
import type { ExtensionAPI, ExtensionCommandContext, ExtensionContext, Theme, ThemeColor } from "@earendil-works/pi-coding-agent";
|
|
34
34
|
import type { Component, OverlayHandle, TUI } from "@earendil-works/pi-tui";
|
|
35
35
|
import { HStack, matchesKey, sliceByColumn, truncateToWidth, visibleWidth, wrapTextWithAnsi } from "@earendil-works/pi-tui";
|
|
36
|
+
import {
|
|
37
|
+
BUDGET_DEFAULTS,
|
|
38
|
+
BUDGET_MESSAGE,
|
|
39
|
+
BUDGET_WINDOW,
|
|
40
|
+
HYGIENE_DEFAULTS,
|
|
41
|
+
MAGI_MD,
|
|
42
|
+
PRUNED_MARK,
|
|
43
|
+
learnedBudget,
|
|
44
|
+
budgetSample,
|
|
45
|
+
pruneContext,
|
|
46
|
+
requestPhase,
|
|
47
|
+
thinkingTokens,
|
|
48
|
+
budgetVerdict,
|
|
49
|
+
thinkingWasCut,
|
|
50
|
+
type BudgetMode,
|
|
51
|
+
type BudgetPhase,
|
|
52
|
+
type HygieneOptions,
|
|
53
|
+
type HygieneStats,
|
|
54
|
+
} from "./local-models.ts";
|
|
36
55
|
|
|
37
56
|
/* ────────────────────────────────────────────────────────────── art ── */
|
|
38
57
|
|
|
@@ -317,6 +336,8 @@ const state = {
|
|
|
317
336
|
rebornAt: 0,
|
|
318
337
|
hasSmartCompact: false,
|
|
319
338
|
lastCouncil: undefined as { verdict: Vote | null; tally: number; question: string } | undefined,
|
|
339
|
+
pruned: 0, // tokens the context hygiene keeps away from the model
|
|
340
|
+
thinkCuts: 0, // replies whose thinking llama.cpp cut at the budget
|
|
320
341
|
};
|
|
321
342
|
|
|
322
343
|
/** Panel preferences, persisted under "ui" in ~/.pi/agent/magi.json. */
|
|
@@ -324,7 +345,7 @@ type Currency = "EUR" | "USD";
|
|
|
324
345
|
|
|
325
346
|
const ui = {
|
|
326
347
|
compact: false,
|
|
327
|
-
kwhPrice: undefined as number | undefined, // price per kWh, for the COST row (/magi
|
|
348
|
+
kwhPrice: undefined as number | undefined, // price per kWh, for the COST row (/magi cost)
|
|
328
349
|
currency: "EUR" as Currency,
|
|
329
350
|
};
|
|
330
351
|
|
|
@@ -421,7 +442,9 @@ interface TokenStats {
|
|
|
421
442
|
|
|
422
443
|
const tokens: TokenStats = { input: 0, output: 0, cacheRead: 0, cost: 0 };
|
|
423
444
|
|
|
424
|
-
|
|
445
|
+
/** Session totals from one assistant reply: tokens, cost, and whether its thinking was cut at the budget. */
|
|
446
|
+
function countAssistant(m: AssistantMessage): void {
|
|
447
|
+
if (thinkingWasCut(m as any)) state.thinkCuts++;
|
|
425
448
|
tokens.input += m.usage?.input ?? 0;
|
|
426
449
|
tokens.output += m.usage?.output ?? 0;
|
|
427
450
|
tokens.cacheRead += m.usage?.cacheRead ?? 0;
|
|
@@ -432,8 +455,9 @@ function addUsage(m: AssistantMessage): void {
|
|
|
432
455
|
function recountSession(ctx: ExtensionContext): void {
|
|
433
456
|
Object.assign(tokens, { input: 0, output: 0, cacheRead: 0, cost: 0 });
|
|
434
457
|
state.lastCouncil = undefined;
|
|
458
|
+
state.thinkCuts = 0;
|
|
435
459
|
for (const entry of ctx.sessionManager.getBranch()) {
|
|
436
|
-
if (entry.type === "message" && entry.message.role === "assistant")
|
|
460
|
+
if (entry.type === "message" && entry.message.role === "assistant") countAssistant(entry.message as AssistantMessage);
|
|
437
461
|
if (entry.type === "custom" && entry.customType === "magi-verdict") {
|
|
438
462
|
const d = entry.data as { verdict?: Vote | null; tally?: number; question?: string } | undefined;
|
|
439
463
|
if (d && Array.isArray((d as any).opinions)) state.lastCouncil = { verdict: d.verdict ?? null, tally: d.tally ?? 0, question: d.question ?? "" };
|
|
@@ -781,10 +805,11 @@ let lastPersistAt = 0;
|
|
|
781
805
|
/**
|
|
782
806
|
* Adds this session's new energy to the running total in magi.json, so the COST row survives restarts.
|
|
783
807
|
* ponytail: writes at most once a minute, and adds a delta so parallel sessions do not overwrite each other.
|
|
808
|
+
* A crash therefore loses up to a minute of energy, which is what a kill loses anyway.
|
|
784
809
|
*/
|
|
785
|
-
function persistEnergy(
|
|
810
|
+
function persistEnergy(): void {
|
|
786
811
|
const delta = swap.energyWh - swap.savedWh;
|
|
787
|
-
if (delta <= 0 ||
|
|
812
|
+
if (delta <= 0 || Date.now() - lastPersistAt < ENERGY_SAVE_MS) return;
|
|
788
813
|
lastPersistAt = Date.now();
|
|
789
814
|
const cfg = loadMagiConfig();
|
|
790
815
|
cfg.totalWh = (cfg.totalWh ?? 0) + delta;
|
|
@@ -882,6 +907,56 @@ async function swapAliases(): Promise<Map<string, string>> {
|
|
|
882
907
|
return aliases;
|
|
883
908
|
}
|
|
884
909
|
|
|
910
|
+
const PI_THINKING_LEVELS = ["minimal", "low", "medium", "high", "xhigh", "max"];
|
|
911
|
+
|
|
912
|
+
/**
|
|
913
|
+
* pi-llama-swap registers every model with reasoning off, so /thinking only offers "off".
|
|
914
|
+
* llama-swap publishes each model's reasoning levels in /v1/models: re-register the provider with them,
|
|
915
|
+
* sent as chat_template_kwargs (enable_thinking + reasoning_effort). Aliases (the Instruct twins) stay off.
|
|
916
|
+
* models.json modelOverrides still apply on top.
|
|
917
|
+
*/
|
|
918
|
+
async function enableSwapReasoning(pi: ExtensionAPI, ctx: ExtensionContext): Promise<void> {
|
|
919
|
+
const config = ctx.modelRegistry.getRegisteredProviderConfig("llama-swap") as any;
|
|
920
|
+
if (!config?.models?.length || !config.baseUrl) return;
|
|
921
|
+
try {
|
|
922
|
+
const headers: Record<string, string> = config.apiKey ? { Authorization: `Bearer ${config.apiKey}` } : {};
|
|
923
|
+
const res = await fetch(`${config.baseUrl.replace(/\/$/, "")}/models`, { headers, signal: AbortSignal.timeout(5000) });
|
|
924
|
+
const { data } = (await res.json()) as { data?: any[] };
|
|
925
|
+
const meta = new Map((data ?? []).map((m) => [m.id, m]));
|
|
926
|
+
let changed = false;
|
|
927
|
+
const models = config.models.map((m: any) => {
|
|
928
|
+
const entry = meta.get(m.id);
|
|
929
|
+
const levels: string[] | undefined = entry?.meta?.llamaswap?.reasoning?.levels;
|
|
930
|
+
const input = entry?.architecture?.input_modalities?.includes("image") ? ["text", "image"] : ["text"];
|
|
931
|
+
if (!levels?.length || entry.meta.llamaswap.type === "alias") return { ...m, input };
|
|
932
|
+
changed = true;
|
|
933
|
+
return {
|
|
934
|
+
...m,
|
|
935
|
+
input,
|
|
936
|
+
reasoning: true,
|
|
937
|
+
thinkingLevelMap: Object.fromEntries(PI_THINKING_LEVELS.map((l) => [l, levels.includes(l) ? l : null])),
|
|
938
|
+
compat: {
|
|
939
|
+
...m.compat,
|
|
940
|
+
thinkingFormat: "chat-template",
|
|
941
|
+
chatTemplateKwargs: {
|
|
942
|
+
enable_thinking: { $var: "thinking.enabled" },
|
|
943
|
+
preserve_thinking: true,
|
|
944
|
+
reasoning_effort: { $var: "thinking.effort", omitWhenOff: true },
|
|
945
|
+
},
|
|
946
|
+
},
|
|
947
|
+
};
|
|
948
|
+
});
|
|
949
|
+
// ponytail: a /llama-swap refresh re-registers the plain models until the next session start
|
|
950
|
+
if (!changed) return;
|
|
951
|
+
ctx.modelRegistry.registerProvider("llama-swap", { ...config, models });
|
|
952
|
+
// the session already holds the old model object: swap in the new one
|
|
953
|
+
const fresh = ctx.model?.provider === "llama-swap" ? ctx.modelRegistry.find("llama-swap", ctx.model.id) : undefined;
|
|
954
|
+
if (fresh?.reasoning) await pi.setModel(fresh);
|
|
955
|
+
} catch {
|
|
956
|
+
// server unreachable: models stay as pi-llama-swap registered them
|
|
957
|
+
}
|
|
958
|
+
}
|
|
959
|
+
|
|
885
960
|
/** Models llama-swap keeps in memory: real id → "ready" | "starting" | … (empty when the server is unreachable). */
|
|
886
961
|
async function swapRunning(): Promise<Map<string, string>> {
|
|
887
962
|
try {
|
|
@@ -1030,7 +1105,7 @@ async function prewarmPrefix(cwd: string): Promise<void> {
|
|
|
1030
1105
|
}
|
|
1031
1106
|
}
|
|
1032
1107
|
|
|
1033
|
-
/* ── /magi
|
|
1108
|
+
/* ── /magi status: a report built from the last requests llama-swap recorded ── */
|
|
1034
1109
|
|
|
1035
1110
|
interface ActivityRow {
|
|
1036
1111
|
timestamp: string;
|
|
@@ -1315,6 +1390,11 @@ class MagiPanel implements Component {
|
|
|
1315
1390
|
if (!compact && usage?.contextWindow) {
|
|
1316
1391
|
out.push(this.field("CONTEXT", `${usage.tokens == null ? "?" : fmtTokens(usage.tokens)} / ${fmtTokens(usage.contextWindow)}`, inner, "muted"));
|
|
1317
1392
|
}
|
|
1393
|
+
// HYGIENE: tokens pruned from what the model sees, ✂ = thinking cut at the budget
|
|
1394
|
+
if (!compact && (state.pruned || state.thinkCuts)) {
|
|
1395
|
+
const cuts = state.thinkCuts ? ` ✂${state.thinkCuts}` : "";
|
|
1396
|
+
out.push(this.field("HYGIENE", `-${fmtTokens(state.pruned)}${cuts}`, inner, state.thinkCuts ? "warning" : "muted"));
|
|
1397
|
+
}
|
|
1318
1398
|
return out;
|
|
1319
1399
|
}
|
|
1320
1400
|
|
|
@@ -1346,14 +1426,12 @@ class MagiPanel implements Component {
|
|
|
1346
1426
|
return out;
|
|
1347
1427
|
}
|
|
1348
1428
|
|
|
1349
|
-
/** COST: GPU energy × price per kWh set with /magi
|
|
1429
|
+
/** COST: GPU energy × price per kWh set with /magi cost — all sessions, with this one in brackets. */
|
|
1350
1430
|
private costRow(inner: number): string {
|
|
1351
1431
|
if (!swap.gpus.length) return this.field("COST", "—", inner, "muted");
|
|
1352
|
-
if (ui.kwhPrice === undefined) return this.field("COST", "→ /magi
|
|
1353
|
-
const th = this.theme;
|
|
1432
|
+
if (ui.kwhPrice === undefined) return this.field("COST", "→ /magi cost", inner, "dim");
|
|
1354
1433
|
const price = (wh: number) => fmtMoney((wh / 1000) * ui.kwhPrice!);
|
|
1355
|
-
|
|
1356
|
-
return this.frameLine(" " + label + th.fg("warning", price(totalWh())) + th.fg("dim", ` (ses ${price(swap.energyWh)})`), inner);
|
|
1434
|
+
return this.field("COST", price(totalWh()) + this.theme.fg("dim", ` (ses ${price(swap.energyWh)})`), inner, "warning");
|
|
1357
1435
|
}
|
|
1358
1436
|
|
|
1359
1437
|
private swapRows(inner: number, compact: boolean): string[] {
|
|
@@ -1505,9 +1583,49 @@ type MagiConfig = Partial<Record<MagiUnit, MagiUnitConfig>> & {
|
|
|
1505
1583
|
ui?: { compact?: boolean; kwhPrice?: number; currency?: Currency };
|
|
1506
1584
|
loads?: Record<string, number>; // real model id → ms its last load took
|
|
1507
1585
|
totalWh?: number; // GPU energy summed over every session, for the COST row
|
|
1586
|
+
hygiene?: Partial<HygieneOptions> & { enabled?: boolean }; // context pruning for local models, see local-models.ts
|
|
1587
|
+
thinkingBudget?: {
|
|
1588
|
+
mode?: BudgetMode; // auto (learned per model) · fixed · off, set with /magi budget
|
|
1589
|
+
planning?: number; // fixed budgets
|
|
1590
|
+
acting?: number;
|
|
1591
|
+
message?: boolean; // send BUDGET_MESSAGE as reasoning_budget_message (default on), /magi budget message
|
|
1592
|
+
learned?: Record<string, Partial<Record<BudgetPhase, number[]>>>; // written by the theme: recent thinking lengths per model
|
|
1593
|
+
};
|
|
1508
1594
|
};
|
|
1509
1595
|
|
|
1510
1596
|
const MAGI_CONFIG_PATH = join(homedir(), ".pi", "agent", "magi.json");
|
|
1597
|
+
/** /magi arguments that manage the theme instead of asking the council. */
|
|
1598
|
+
const UI_ARGS = /^(on|off|panel|compact|status|cost)$|^(hygiene|budget)(\s|$)/i;
|
|
1599
|
+
/** /magi arguments offered by autocomplete: the full argument, and what it does. */
|
|
1600
|
+
const MAGI_ARGS: [string, string][] = [
|
|
1601
|
+
["council", "ask the three MAGI a question (recent conversation as context)"],
|
|
1602
|
+
["council review", "the council reviews your pending changes before you commit"],
|
|
1603
|
+
["config", "pick a model for each MAGI"],
|
|
1604
|
+
["mecha", "MECHA SELECT: pick the llama-swap model to activate"],
|
|
1605
|
+
["status", "llama-swap report: speed, tokens, cache hits, errors per model"],
|
|
1606
|
+
["panel", "hide/show the side panel"],
|
|
1607
|
+
["compact", "toggle the compact side panel"],
|
|
1608
|
+
["cost", "electricity price and currency for the COST row"],
|
|
1609
|
+
["on", "enable the MAGI chrome"],
|
|
1610
|
+
["off", "disable the MAGI chrome"],
|
|
1611
|
+
["hygiene", "show how much context was pruned"],
|
|
1612
|
+
["hygiene on", "enable context pruning"],
|
|
1613
|
+
["hygiene off", "disable context pruning"],
|
|
1614
|
+
["hygiene step 40k", "prune less often: fewer prompt re-reads on long tasks (default 15k)"],
|
|
1615
|
+
["hygiene 3 5", "recent turns that keep their thinking, tool results kept whole"],
|
|
1616
|
+
["budget", "show the thinking budget of the current model"],
|
|
1617
|
+
["budget auto", "learn the budget per model (default)"],
|
|
1618
|
+
["budget off", "no budget: llama-server decides"],
|
|
1619
|
+
["budget reset", "forget what was learned for the current model"],
|
|
1620
|
+
["budget message", "turn the closing message after a cut off/on"],
|
|
1621
|
+
["budget 16k 4k", "fixed budget: planning, acting"],
|
|
1622
|
+
];
|
|
1623
|
+
function argCompletions(table: [string, string][], prefix: string) {
|
|
1624
|
+
const p = prefix.trimStart().toLowerCase();
|
|
1625
|
+
const items = table.filter(([value]) => value.startsWith(p)).map(([value, description]) => ({ value, label: value, description }));
|
|
1626
|
+
return items.length ? items : null;
|
|
1627
|
+
}
|
|
1628
|
+
const LOOP_REPEATS = 3; // the same tool call this many times in a row is blocked
|
|
1511
1629
|
|
|
1512
1630
|
function loadMagiConfig(): MagiConfig {
|
|
1513
1631
|
try {
|
|
@@ -1561,7 +1679,7 @@ function conversationExcerpt(ctx: ExtensionContext, maxChars = 6000): string {
|
|
|
1561
1679
|
|
|
1562
1680
|
const REVIEW_MAX_CHARS = 24_000;
|
|
1563
1681
|
|
|
1564
|
-
/** The pending changes for /magi review: tracked changes against HEAD plus the names of untracked files. */
|
|
1682
|
+
/** The pending changes for /magi council review: tracked changes against HEAD plus the names of untracked files. */
|
|
1565
1683
|
async function pendingChanges(cwd: string): Promise<{ diff: string; untracked: string[] }> {
|
|
1566
1684
|
const git = (args: string[]) => promisify(execFile)("git", args, { cwd, maxBuffer: 32 * 1024 * 1024 }).then((r) => r.stdout);
|
|
1567
1685
|
let diff: string;
|
|
@@ -1723,7 +1841,7 @@ function buildDeliberationView(
|
|
|
1723
1841
|
};
|
|
1724
1842
|
}
|
|
1725
1843
|
|
|
1726
|
-
/** A read-only boxed report (used by /magi
|
|
1844
|
+
/** A read-only boxed report (used by /magi status); any key closes it. */
|
|
1727
1845
|
function buildReportView(theme: Theme, lines: string[], close: () => void) {
|
|
1728
1846
|
return {
|
|
1729
1847
|
render(width: number): string[] {
|
|
@@ -1984,7 +2102,7 @@ function buildFooter(tui: TUI, theme: Theme, footerData: any) {
|
|
|
1984
2102
|
// change would shift it under the window. A fresh line is taken when it has the same
|
|
1985
2103
|
// width, so nothing moves, otherwise at the end of the loop.
|
|
1986
2104
|
const line = left + dim(FOOTER_SCROLL_GAP) + right + dim(FOOTER_SCROLL_GAP);
|
|
1987
|
-
const period =
|
|
2105
|
+
const period = visibleWidth(line);
|
|
1988
2106
|
if (!scrollLine || scrollOff === 0 || period === scrollPeriod) {
|
|
1989
2107
|
scrollLine = line;
|
|
1990
2108
|
scrollPeriod = period;
|
|
@@ -2021,6 +2139,22 @@ export default function (pi: ExtensionAPI) {
|
|
|
2021
2139
|
let panelHandle: OverlayHandle | undefined;
|
|
2022
2140
|
let panel: MagiPanel | undefined;
|
|
2023
2141
|
let titleDone = false;
|
|
2142
|
+
let hygiene = { enabled: true, ...HYGIENE_DEFAULTS };
|
|
2143
|
+
let hygieneStats: HygieneStats | undefined;
|
|
2144
|
+
let budgetMode: BudgetMode = "auto";
|
|
2145
|
+
let fixedBudget = { ...BUDGET_DEFAULTS };
|
|
2146
|
+
let budgetMessage = true;
|
|
2147
|
+
let learned: Record<string, Partial<Record<BudgetPhase, number[]>>> = {};
|
|
2148
|
+
let pendingBudget: { model: string; phase: BudgetPhase; tokens: number } | undefined; // the request in flight
|
|
2149
|
+
const budgetIgnored = new Set<string>(); // models whose llama-server thought past the budget: it sets its own
|
|
2150
|
+
const budgetFor = (model: string, phase: BudgetPhase) =>
|
|
2151
|
+
budgetMode === "fixed" ? fixedBudget[phase] : (learnedBudget(learned[model]?.[phase] ?? [], phase) ?? BUDGET_DEFAULTS[phase]);
|
|
2152
|
+
const persistBudget = () => {
|
|
2153
|
+
const cfg = loadMagiConfig();
|
|
2154
|
+
saveMagiConfig({ ...cfg, thinkingBudget: { mode: budgetMode, ...fixedBudget, message: budgetMessage, learned } });
|
|
2155
|
+
};
|
|
2156
|
+
let lastCall = ""; // loop guard: fingerprint of the previous tool call and how often it repeated
|
|
2157
|
+
let repeats = 0;
|
|
2024
2158
|
|
|
2025
2159
|
const repaint = () => {
|
|
2026
2160
|
panel?.invalidate();
|
|
@@ -2181,6 +2315,24 @@ export default function (pi: ExtensionAPI) {
|
|
|
2181
2315
|
|
|
2182
2316
|
pi.on("session_start", async (event, ctx) => {
|
|
2183
2317
|
liveCtx = ctx;
|
|
2318
|
+
const magiCfg = loadMagiConfig();
|
|
2319
|
+
hygiene = { enabled: true, ...HYGIENE_DEFAULTS, ...magiCfg.hygiene };
|
|
2320
|
+
const tb = magiCfg.thinkingBudget ?? {};
|
|
2321
|
+
budgetMode = tb.mode ?? "auto";
|
|
2322
|
+
fixedBudget = { planning: tb.planning ?? BUDGET_DEFAULTS.planning, acting: tb.acting ?? BUDGET_DEFAULTS.acting };
|
|
2323
|
+
budgetMessage = tb.message ?? true;
|
|
2324
|
+
learned = tb.learned ?? {};
|
|
2325
|
+
await enableSwapReasoning(pi, ctx);
|
|
2326
|
+
// the local-model rules live next to AGENTS.md, created once so the user can edit them
|
|
2327
|
+
const rules = join(ctx.cwd, "MAGI.md");
|
|
2328
|
+
if (!existsSync(rules)) {
|
|
2329
|
+
try {
|
|
2330
|
+
writeFileSync(rules, MAGI_MD);
|
|
2331
|
+
if (ctx.mode === "tui") ctx.ui.notify("MAGI.md created: rules for local models, appended to the system prompt", "info");
|
|
2332
|
+
} catch {
|
|
2333
|
+
// read-only directory: the built-in rules are used as they are
|
|
2334
|
+
}
|
|
2335
|
+
}
|
|
2184
2336
|
void refreshUnit(ctx);
|
|
2185
2337
|
recountSession(ctx);
|
|
2186
2338
|
state.hasSmartCompact = pi.getCommands().some((c) => c.name.replace(/^\//, "") === "smart-compact");
|
|
@@ -2233,6 +2385,12 @@ export default function (pi: ExtensionAPI) {
|
|
|
2233
2385
|
|
|
2234
2386
|
pi.on("before_provider_request", (event, ctx) => {
|
|
2235
2387
|
rememberPrefix(ctx.cwd, event.payload);
|
|
2388
|
+
// llama.cpp honours a per-request thinking budget only when llama-server runs without --reasoning-budget
|
|
2389
|
+
if (budgetMode === "off" || !swap.base || !ctx.model) return;
|
|
2390
|
+
const phase = requestPhase(event.payload);
|
|
2391
|
+
pendingBudget = { model: ctx.model.id, phase, tokens: budgetFor(ctx.model.id, phase) };
|
|
2392
|
+
const payload = { ...(event.payload as object), thinking_budget_tokens: pendingBudget.tokens };
|
|
2393
|
+
return budgetMessage ? { ...payload, reasoning_budget_message: `\n\n${BUDGET_MESSAGE}` } : payload;
|
|
2236
2394
|
});
|
|
2237
2395
|
|
|
2238
2396
|
pi.on("model_select", async (event, ctx) => {
|
|
@@ -2242,7 +2400,7 @@ export default function (pi: ExtensionAPI) {
|
|
|
2242
2400
|
});
|
|
2243
2401
|
|
|
2244
2402
|
pi.on("session_shutdown", async () => {
|
|
2245
|
-
persistEnergy(
|
|
2403
|
+
persistEnergy();
|
|
2246
2404
|
prewarm.abort?.abort();
|
|
2247
2405
|
clearInterval(metricsTimer);
|
|
2248
2406
|
metricsTimer = undefined;
|
|
@@ -2257,6 +2415,40 @@ export default function (pi: ExtensionAPI) {
|
|
|
2257
2415
|
|
|
2258
2416
|
pi.on("agent_start", async () => {
|
|
2259
2417
|
state.runStart = Date.now();
|
|
2418
|
+
lastCall = "";
|
|
2419
|
+
});
|
|
2420
|
+
|
|
2421
|
+
pi.on("before_agent_start", async (event, ctx) => {
|
|
2422
|
+
let rules = MAGI_MD;
|
|
2423
|
+
try {
|
|
2424
|
+
rules = readFileSync(join(ctx.cwd, "MAGI.md"), "utf8");
|
|
2425
|
+
} catch {
|
|
2426
|
+
// no MAGI.md (deleted, or unwritable cwd): the built-in rules
|
|
2427
|
+
}
|
|
2428
|
+
return rules.trim() ? { systemPrompt: `${event.systemPrompt}\n\n${rules.trim()}` } : undefined;
|
|
2429
|
+
});
|
|
2430
|
+
|
|
2431
|
+
// before every request: drop old thinking and tool outputs from what the model sees, never from the session
|
|
2432
|
+
pi.on("context", async (event) => {
|
|
2433
|
+
if (!hygiene.enabled) return;
|
|
2434
|
+
const pruned = pruneContext(event.messages as any, hygiene);
|
|
2435
|
+
hygieneStats = pruned.stats;
|
|
2436
|
+
state.pruned = pruned.stats.prunedTokens;
|
|
2437
|
+
return { messages: pruned.messages as any };
|
|
2438
|
+
});
|
|
2439
|
+
|
|
2440
|
+
// local models loop on the same call, and may copy a pruned placeholder into a file
|
|
2441
|
+
pi.on("tool_call", async (event) => {
|
|
2442
|
+
const args = JSON.stringify(event.input);
|
|
2443
|
+
if ((event.toolName === "write" || event.toolName === "edit") && args.includes(PRUNED_MARK)) {
|
|
2444
|
+
return { block: true, reason: `MAGI: this ${event.toolName} contains a "${PRUNED_MARK}" placeholder, not real content. Read the file and use the actual text.` };
|
|
2445
|
+
}
|
|
2446
|
+
const fingerprint = event.toolName + args;
|
|
2447
|
+
repeats = fingerprint === lastCall ? repeats + 1 : 1;
|
|
2448
|
+
lastCall = fingerprint;
|
|
2449
|
+
if (repeats >= LOOP_REPEATS) {
|
|
2450
|
+
return { block: true, reason: `MAGI: identical ${event.toolName} call ${repeats} times in a row. Repeating it will not change the outcome: change approach, or tell the user what blocks you.` };
|
|
2451
|
+
}
|
|
2260
2452
|
});
|
|
2261
2453
|
|
|
2262
2454
|
pi.on("turn_start", async (_event, ctx) => {
|
|
@@ -2300,7 +2492,19 @@ export default function (pi: ExtensionAPI) {
|
|
|
2300
2492
|
pi.on("message_end", async (event) => {
|
|
2301
2493
|
if (event.message.role !== "assistant") return;
|
|
2302
2494
|
const m = event.message as AssistantMessage;
|
|
2303
|
-
|
|
2495
|
+
countAssistant(m);
|
|
2496
|
+
// learn how long this model thinks in this phase; a cut must not raise the budget
|
|
2497
|
+
const thought = thinkingTokens(m as any);
|
|
2498
|
+
if (pendingBudget && thought > 0 && m.stopReason !== "aborted" && m.stopReason !== "error") {
|
|
2499
|
+
const verdict = budgetVerdict(m as any, pendingBudget.tokens);
|
|
2500
|
+
if (verdict === "ignored") budgetIgnored.add(pendingBudget.model);
|
|
2501
|
+
if (verdict === "cut" && !thinkingWasCut(m as any)) state.thinkCuts++; // silent cut: no budget message on the server
|
|
2502
|
+
const samples = ((learned[pendingBudget.model] ??= {})[pendingBudget.phase] ??= []);
|
|
2503
|
+
samples.push(budgetSample(thought, pendingBudget.tokens, verdict === "cut"));
|
|
2504
|
+
samples.splice(0, samples.length - BUDGET_WINDOW);
|
|
2505
|
+
persistBudget();
|
|
2506
|
+
}
|
|
2507
|
+
pendingBudget = undefined;
|
|
2304
2508
|
if (perf.start) {
|
|
2305
2509
|
const end = Date.now();
|
|
2306
2510
|
perf.lastMs = end - perf.start;
|
|
@@ -2436,9 +2640,11 @@ export default function (pi: ExtensionAPI) {
|
|
|
2436
2640
|
}
|
|
2437
2641
|
|
|
2438
2642
|
pi.registerCommand("magi", {
|
|
2439
|
-
description: "Ask the three MAGI
|
|
2643
|
+
description: "Ask the three MAGI, or manage them: review [focus] · config · mecha · status · panel · compact · cost · on · off · hygiene · budget (type a space to see them all)",
|
|
2644
|
+
getArgumentCompletions: (prefix) => argCompletions(MAGI_ARGS, prefix),
|
|
2440
2645
|
handler: async (args, ctx) => {
|
|
2441
2646
|
const arg = args.trim();
|
|
2647
|
+
if (UI_ARGS.test(arg)) return manageUi(arg, ctx);
|
|
2442
2648
|
if (arg === "config") return configureMagi(ctx);
|
|
2443
2649
|
if (arg === "mecha") {
|
|
2444
2650
|
if (ctx.mode !== "tui" || !swap.base) return ctx.ui.notify("MECHA SELECT needs the TUI and a llama-swap model", "error");
|
|
@@ -2446,8 +2652,12 @@ export default function (pi: ExtensionAPI) {
|
|
|
2446
2652
|
return pickModel(ctx);
|
|
2447
2653
|
}
|
|
2448
2654
|
|
|
2449
|
-
if (arg
|
|
2450
|
-
|
|
2655
|
+
if (arg !== "council" && !arg.startsWith("council ")) {
|
|
2656
|
+
return ctx.ui.notify(arg ? `Unknown /magi command "${arg}": to ask the MAGI, /magi council <question>` : "Ask the MAGI with /magi council <question>; type /magi and a space to see every command", arg ? "warning" : "info");
|
|
2657
|
+
}
|
|
2658
|
+
const ask = arg.slice("council".length).trim();
|
|
2659
|
+
if (ask === "review" || ask.startsWith("review ")) {
|
|
2660
|
+
const focus = ask.slice("review".length).trim();
|
|
2451
2661
|
let changes: { diff: string; untracked: string[] };
|
|
2452
2662
|
try {
|
|
2453
2663
|
changes = await pendingChanges(ctx.cwd);
|
|
@@ -2470,93 +2680,151 @@ export default function (pi: ExtensionAPI) {
|
|
|
2470
2680
|
return runCouncil(ctx, question, "```diff\n" + diff + "\n```" + untracked, "Pending changes (git diff HEAD)");
|
|
2471
2681
|
}
|
|
2472
2682
|
|
|
2473
|
-
const question =
|
|
2683
|
+
const question = ask || (await ctx.ui.input("Question for the MAGI:", "should we …?"))?.trim() || "";
|
|
2474
2684
|
if (!question) return;
|
|
2475
2685
|
return runCouncil(ctx, question, conversationExcerpt(ctx));
|
|
2476
2686
|
},
|
|
2477
2687
|
});
|
|
2478
2688
|
|
|
2479
|
-
|
|
2480
|
-
|
|
2481
|
-
|
|
2482
|
-
|
|
2483
|
-
|
|
2484
|
-
|
|
2485
|
-
|
|
2486
|
-
|
|
2487
|
-
|
|
2488
|
-
|
|
2489
|
-
|
|
2689
|
+
/** /magi on|off|panel|compact|status|cost|hygiene|budget: the theme, its panel and the local-model settings. */
|
|
2690
|
+
async function manageUi(args: string, ctx: ExtensionCommandContext) {
|
|
2691
|
+
liveCtx = ctx;
|
|
2692
|
+
const arg = args.trim().toLowerCase();
|
|
2693
|
+
|
|
2694
|
+
if (arg === "hygiene" || arg.startsWith("hygiene ")) {
|
|
2695
|
+
const sub = arg.slice("hygiene".length).trim();
|
|
2696
|
+
const keep = /^(\d+)\s+(\d+)$/.exec(sub);
|
|
2697
|
+
const step = /^step\s+(\d+)(k?)$/.exec(sub);
|
|
2698
|
+
if (sub === "on" || sub === "off" || keep || step) {
|
|
2699
|
+
if (keep) Object.assign(hygiene, { enabled: true, keepThinkingTurns: Number(keep[1]), keepToolResults: Number(keep[2]) });
|
|
2700
|
+
else if (step) hygiene.stepTokens = Math.max(1000, Number(step[1]) * (step[2] ? 1000 : 1));
|
|
2701
|
+
else hygiene.enabled = sub === "on";
|
|
2702
|
+
const cfg = loadMagiConfig();
|
|
2703
|
+
const { enabled, keepThinkingTurns, keepToolResults, stepTokens } = hygiene;
|
|
2704
|
+
saveMagiConfig({ ...cfg, hygiene: { ...cfg.hygiene, enabled, keepThinkingTurns, keepToolResults, stepTokens } });
|
|
2705
|
+
} else if (sub) {
|
|
2706
|
+
ctx.ui.notify("Usage: /magi hygiene [on|off|step <tokens>|<thinking turns kept> <tool results kept>], e.g. step 40k, 3 5", "warning");
|
|
2490
2707
|
return;
|
|
2491
2708
|
}
|
|
2492
|
-
|
|
2493
|
-
|
|
2494
|
-
|
|
2495
|
-
|
|
2496
|
-
|
|
2497
|
-
|
|
2709
|
+
const s = hygieneStats;
|
|
2710
|
+
ctx.ui.notify(
|
|
2711
|
+
(!hygiene.enabled
|
|
2712
|
+
? "Context hygiene off: /magi hygiene on"
|
|
2713
|
+
: s
|
|
2714
|
+
? `Context hygiene: ${fmtTokens(s.prunedTokens)} tokens pruned in the first ${s.watermark}/${s.messages} messages, ${fmtTokens(s.pendingTokens)} waiting for the next step (every ${fmtTokens(hygiene.stepTokens)})`
|
|
2715
|
+
: "Context hygiene on: nothing sent to the model yet") +
|
|
2716
|
+
` · keeps thinking of the last ${hygiene.keepThinkingTurns} turns, the last ${hygiene.keepToolResults} tool results` +
|
|
2717
|
+
(state.thinkCuts ? ` · thinking cut at the budget ${state.thinkCuts}×` : ""),
|
|
2718
|
+
"info",
|
|
2719
|
+
);
|
|
2720
|
+
return;
|
|
2721
|
+
}
|
|
2722
|
+
if (arg === "budget" || arg.startsWith("budget ")) {
|
|
2723
|
+
const sub = arg.slice("budget".length).trim();
|
|
2724
|
+
const model = ctx.model?.id ?? "";
|
|
2725
|
+
const tokens = (t: string) => Math.round(Number(t.replace(/k$/, "")) * (t.endsWith("k") ? 1024 : 1));
|
|
2726
|
+
const fixed = /^(\d+k?)\s+(\d+k?)$/.exec(sub);
|
|
2727
|
+
if (sub === "auto" || sub === "off") budgetMode = sub;
|
|
2728
|
+
else if (sub === "reset") delete learned[model];
|
|
2729
|
+
else if (sub === "message") budgetMessage = !budgetMessage;
|
|
2730
|
+
else if (fixed) {
|
|
2731
|
+
budgetMode = "fixed";
|
|
2732
|
+
fixedBudget = { planning: tokens(fixed[1]!), acting: tokens(fixed[2]!) };
|
|
2733
|
+
} else if (sub) {
|
|
2734
|
+
ctx.ui.notify("Usage: /magi budget [auto|off|reset|message|<planning> <acting>], e.g. 16k 4k", "warning");
|
|
2498
2735
|
return;
|
|
2499
2736
|
}
|
|
2500
|
-
if (
|
|
2501
|
-
|
|
2502
|
-
|
|
2503
|
-
const
|
|
2504
|
-
|
|
2505
|
-
|
|
2506
|
-
|
|
2507
|
-
|
|
2508
|
-
|
|
2509
|
-
|
|
2510
|
-
|
|
2511
|
-
|
|
2512
|
-
|
|
2513
|
-
|
|
2514
|
-
|
|
2515
|
-
|
|
2516
|
-
|
|
2517
|
-
|
|
2518
|
-
|
|
2519
|
-
|
|
2520
|
-
|
|
2737
|
+
if (sub) persistBudget();
|
|
2738
|
+
const phase = (p: BudgetPhase) => {
|
|
2739
|
+
const n = learned[model]?.[p]?.length ?? 0;
|
|
2740
|
+
const how = budgetMode === "fixed" ? "fixed" : learnedBudget(learned[model]?.[p] ?? [], p) ? `learned from ${n}` : `default, learning ${n}/10`;
|
|
2741
|
+
return `${p} ${budgetFor(model, p) / 1024}k (${how})`;
|
|
2742
|
+
};
|
|
2743
|
+
ctx.ui.notify(
|
|
2744
|
+
budgetMode === "off"
|
|
2745
|
+
? "Thinking budget off: llama-server decides. /magi budget auto"
|
|
2746
|
+
: `Thinking budget ${budgetMode.toUpperCase()} · ${model || "no model"}: ${phase("planning")} · ${phase("acting")}` +
|
|
2747
|
+
` · closing message ${budgetMessage ? "on" : "off"}` +
|
|
2748
|
+
(swap.base ? "" : " · applies to llama-swap models only") +
|
|
2749
|
+
(budgetIgnored.has(model) ? " · ⚠ llama-server thinks past it: it was started with its own --reasoning-budget (or is too old), which wins" : ""),
|
|
2750
|
+
"info",
|
|
2751
|
+
);
|
|
2752
|
+
return;
|
|
2753
|
+
}
|
|
2754
|
+
if (arg === "panel") {
|
|
2755
|
+
panelEnabled = !panelEnabled;
|
|
2756
|
+
if (panelEnabled) showPanel(ctx.ui.theme);
|
|
2757
|
+
else hidePanel();
|
|
2758
|
+
ctx.ui.notify(`Side panel ${panelEnabled ? "enabled" : "disabled"}`, "info");
|
|
2759
|
+
return;
|
|
2760
|
+
}
|
|
2761
|
+
if (arg === "compact") {
|
|
2762
|
+
ui.compact = !ui.compact;
|
|
2763
|
+
const cfg = loadMagiConfig();
|
|
2764
|
+
saveMagiConfig({ ...cfg, ui: { ...cfg.ui, compact: ui.compact } });
|
|
2765
|
+
repaint();
|
|
2766
|
+
ctx.ui.notify(`Side panel ${ui.compact ? "compact" : "detailed"}`, "info");
|
|
2767
|
+
return;
|
|
2768
|
+
}
|
|
2769
|
+
if (arg === "cost") {
|
|
2770
|
+
const currency = await ctx.ui.select(`Currency for COST (current: ${ui.currency})`, ["EUR", "USD"]);
|
|
2771
|
+
if (!currency) return;
|
|
2772
|
+
const current = ui.kwhPrice !== undefined ? ` (current: ${ui.kwhPrice})` : "";
|
|
2773
|
+
const raw = (await ctx.ui.input(`Electricity price per kWh in ${currency}${current}:`, "0.30"))?.trim();
|
|
2774
|
+
if (raw === undefined) return;
|
|
2775
|
+
const price = raw === "" && ui.kwhPrice !== undefined ? ui.kwhPrice : Number(raw.replace(",", "."));
|
|
2776
|
+
if (raw === "" && ui.kwhPrice === undefined) {
|
|
2777
|
+
ctx.ui.notify("No price entered: COST unchanged", "warning");
|
|
2521
2778
|
return;
|
|
2522
2779
|
}
|
|
2523
|
-
if (
|
|
2524
|
-
|
|
2525
|
-
ctx.ui.notify("/magi-ui status needs a llama-swap session model", "warning");
|
|
2526
|
-
return;
|
|
2527
|
-
}
|
|
2528
|
-
let report: { data?: ActivityRow[]; total?: number };
|
|
2529
|
-
try {
|
|
2530
|
-
report = (await (await swapGet(`/api/metrics/activity?limit=${ACTIVITY_REPORT_ROWS}`, 15_000)).json()) as typeof report;
|
|
2531
|
-
} catch (err) {
|
|
2532
|
-
ctx.ui.notify(`llama-swap status failed: ${err instanceof Error ? err.message : String(err)}`, "error");
|
|
2533
|
-
return;
|
|
2534
|
-
}
|
|
2535
|
-
const rows = report.data ?? [];
|
|
2536
|
-
if (!rows.length) {
|
|
2537
|
-
ctx.ui.notify("llama-swap has no recorded requests yet", "info");
|
|
2538
|
-
return;
|
|
2539
|
-
}
|
|
2540
|
-
await ctx.ui.custom<void>((_tui, theme, _keys, done) =>
|
|
2541
|
-
buildReportView(theme, activityReport(theme, rows, report.total ?? rows.length), () => done(undefined)),
|
|
2542
|
-
);
|
|
2780
|
+
if (!Number.isFinite(price) || price < 0) {
|
|
2781
|
+
ctx.ui.notify(`Invalid price: "${raw}"`, "error");
|
|
2543
2782
|
return;
|
|
2544
2783
|
}
|
|
2545
|
-
|
|
2546
|
-
|
|
2547
|
-
|
|
2548
|
-
|
|
2784
|
+
ui.currency = currency === "USD" ? "USD" : "EUR";
|
|
2785
|
+
ui.kwhPrice = price;
|
|
2786
|
+
const cfg = loadMagiConfig();
|
|
2787
|
+
saveMagiConfig({ ...cfg, ui: { ...cfg.ui, currency: ui.currency, kwhPrice: price } });
|
|
2788
|
+
repaint();
|
|
2789
|
+
ctx.ui.notify(`COST: ${fmtMoney(price)} per kWh`, "info");
|
|
2790
|
+
return;
|
|
2791
|
+
}
|
|
2792
|
+
if (arg === "status") {
|
|
2793
|
+
if (!swap.base) {
|
|
2794
|
+
ctx.ui.notify("/magi status needs a llama-swap session model", "warning");
|
|
2549
2795
|
return;
|
|
2550
2796
|
}
|
|
2551
|
-
|
|
2552
|
-
|
|
2553
|
-
|
|
2554
|
-
|
|
2797
|
+
let report: { data?: ActivityRow[]; total?: number };
|
|
2798
|
+
try {
|
|
2799
|
+
report = (await (await swapGet(`/api/metrics/activity?limit=${ACTIVITY_REPORT_ROWS}`, 15_000)).json()) as typeof report;
|
|
2800
|
+
} catch (err) {
|
|
2801
|
+
ctx.ui.notify(`llama-swap status failed: ${err instanceof Error ? err.message : String(err)}`, "error");
|
|
2555
2802
|
return;
|
|
2556
2803
|
}
|
|
2557
|
-
|
|
2804
|
+
const rows = report.data ?? [];
|
|
2805
|
+
if (!rows.length) {
|
|
2806
|
+
ctx.ui.notify("llama-swap has no recorded requests yet", "info");
|
|
2807
|
+
return;
|
|
2808
|
+
}
|
|
2809
|
+
await ctx.ui.custom<void>((_tui, theme, _keys, done) =>
|
|
2810
|
+
buildReportView(theme, activityReport(theme, rows, report.total ?? rows.length), () => done(undefined)),
|
|
2811
|
+
);
|
|
2812
|
+
return;
|
|
2813
|
+
}
|
|
2814
|
+
if (arg === "off") {
|
|
2815
|
+
chrome = false;
|
|
2558
2816
|
applyChrome(ctx);
|
|
2559
|
-
ctx.ui.notify("MAGI
|
|
2560
|
-
|
|
2561
|
-
|
|
2817
|
+
ctx.ui.notify("MAGI chrome disabled", "info");
|
|
2818
|
+
return;
|
|
2819
|
+
}
|
|
2820
|
+
|
|
2821
|
+
const res = ctx.ui.setTheme("magi");
|
|
2822
|
+
if (!res.success) {
|
|
2823
|
+
ctx.ui.notify(`Theme magi not found: ${res.error}`, "error");
|
|
2824
|
+
return;
|
|
2825
|
+
}
|
|
2826
|
+
chrome = true;
|
|
2827
|
+
applyChrome(ctx);
|
|
2828
|
+
ctx.ui.notify("MAGI online — /magi panel|compact|status|off", "info");
|
|
2829
|
+
}
|
|
2562
2830
|
}
|
|
@@ -0,0 +1,240 @@
|
|
|
1
|
+
/**
|
|
2
|
+
* Local-model helpers for MAGI: context hygiene, loop guard, and the MAGI.md rules template.
|
|
3
|
+
*
|
|
4
|
+
* Context hygiene rewrites the messages pi sends to the model (never the saved session):
|
|
5
|
+
* - old thinking blocks are dropped: Qwen templates resend every reasoning block of an agent run,
|
|
6
|
+
* one 18k-token think stays in the context until compaction
|
|
7
|
+
* - old tool outputs become a one-line marker (observation masking, Lindenbauer et al. 2025:
|
|
8
|
+
* as good as LLM summarization, at half the cost)
|
|
9
|
+
* - old write/edit payloads become a marker too: the file on disk is the source of truth
|
|
10
|
+
* Pruning advances in steps behind a watermark, so between steps the prompt only grows at the end
|
|
11
|
+
* and llama.cpp keeps reusing its KV cache.
|
|
12
|
+
*/
|
|
13
|
+
|
|
14
|
+
export type Msg = { role: string; content?: any; [key: string]: any };
|
|
15
|
+
|
|
16
|
+
export interface HygieneOptions {
|
|
17
|
+
keepThinkingTurns: number; // newest assistant messages that keep their thinking
|
|
18
|
+
keepToolResults: number; // newest tool results kept whole
|
|
19
|
+
stepTokens: number; // the watermark moves each time the prunable total crosses another multiple of this
|
|
20
|
+
minPruneChars: number; // smaller outputs/payloads are left alone
|
|
21
|
+
}
|
|
22
|
+
|
|
23
|
+
export interface HygieneStats {
|
|
24
|
+
messages: number;
|
|
25
|
+
watermark: number; // messages before this index are pruned
|
|
26
|
+
prunedTokens: number;
|
|
27
|
+
pendingTokens: number; // prunable, waiting for the next step
|
|
28
|
+
stepTokens: number;
|
|
29
|
+
}
|
|
30
|
+
|
|
31
|
+
export const HYGIENE_DEFAULTS: HygieneOptions = { keepThinkingTurns: 3, keepToolResults: 5, stepTokens: 15_000, minPruneChars: 600 };
|
|
32
|
+
|
|
33
|
+
// ponytail: chars/token measured on Qwen3.x pi sessions (2.9–3.2); a tokenizer call would be exact but costs a request
|
|
34
|
+
const CHARS_PER_TOKEN = 3;
|
|
35
|
+
const IMAGE_CHARS = 4800;
|
|
36
|
+
|
|
37
|
+
/** Marks every placeholder, so a model that copies one into a real write/edit gets blocked (see loop guard). */
|
|
38
|
+
export const PRUNED_MARK = "<<pruned by MAGI";
|
|
39
|
+
|
|
40
|
+
const lines = (s: string) => s.split("\n").length;
|
|
41
|
+
const clip = (s: string, n = 80) => (s.length > n ? s.slice(0, n - 1) + "…" : s);
|
|
42
|
+
|
|
43
|
+
function callTarget(call: any): string {
|
|
44
|
+
const a = call?.arguments ?? {};
|
|
45
|
+
return clip(String(a.path ?? a.command ?? a.pattern ?? a.query ?? a.url ?? ""));
|
|
46
|
+
}
|
|
47
|
+
|
|
48
|
+
function resultChars(m: Msg): number {
|
|
49
|
+
let n = 0;
|
|
50
|
+
for (const c of m.content ?? []) n += c.type === "image" ? IMAGE_CHARS : (c.text?.length ?? 0);
|
|
51
|
+
return n;
|
|
52
|
+
}
|
|
53
|
+
|
|
54
|
+
function editChars(call: any): number {
|
|
55
|
+
return (call.arguments?.edits ?? []).reduce((n: number, e: any) => n + (e.oldText?.length ?? 0) + (e.newText?.length ?? 0), 0);
|
|
56
|
+
}
|
|
57
|
+
|
|
58
|
+
/** Chars pruning would remove from this assistant message. */
|
|
59
|
+
function assistantSavings(m: Msg, o: HygieneOptions): number {
|
|
60
|
+
let n = 0;
|
|
61
|
+
for (const c of m.content ?? []) {
|
|
62
|
+
if (c.type === "thinking") n += c.thinking?.length ?? 0;
|
|
63
|
+
else if (c.type === "toolCall" && c.name === "write" && (c.arguments?.content?.length ?? 0) >= o.minPruneChars) n += c.arguments.content.length;
|
|
64
|
+
else if (c.type === "toolCall" && c.name === "edit" && editChars(c) >= o.minPruneChars) n += editChars(c);
|
|
65
|
+
}
|
|
66
|
+
return n;
|
|
67
|
+
}
|
|
68
|
+
|
|
69
|
+
function pruneAssistant(m: Msg, o: HygieneOptions): Msg {
|
|
70
|
+
const content = [];
|
|
71
|
+
for (const c of m.content ?? []) {
|
|
72
|
+
if (c.type === "thinking") continue;
|
|
73
|
+
if (c.type === "toolCall" && c.name === "write" && (c.arguments?.content?.length ?? 0) >= o.minPruneChars) {
|
|
74
|
+
const text = c.arguments.content as string;
|
|
75
|
+
content.push({ ...c, arguments: { ...c.arguments, content: `${PRUNED_MARK}: ${lines(text)} lines written, the file on disk is the source of truth>>` } });
|
|
76
|
+
} else if (c.type === "toolCall" && c.name === "edit" && editChars(c) >= o.minPruneChars) {
|
|
77
|
+
const edits = c.arguments.edits.map((e: any) => ({ oldText: `${PRUNED_MARK}>>`, newText: `${PRUNED_MARK}: ${lines(e.newText ?? "")} lines>>` }));
|
|
78
|
+
content.push({ ...c, arguments: { ...c.arguments, edits } });
|
|
79
|
+
} else content.push(c);
|
|
80
|
+
}
|
|
81
|
+
// a message that only held thinking keeps a marker: some templates choke on an empty assistant turn
|
|
82
|
+
if (!content.length) content.push({ type: "text", text: `${PRUNED_MARK}: reasoning>>` });
|
|
83
|
+
return { ...m, content };
|
|
84
|
+
}
|
|
85
|
+
|
|
86
|
+
function pruneResult(m: Msg, call: any): Msg {
|
|
87
|
+
const text = (m.content ?? []).map((c: any) => c.text ?? "").join("\n");
|
|
88
|
+
const what = [m.toolName, callTarget(call)].filter(Boolean).join(" ");
|
|
89
|
+
return { ...m, content: [{ type: "text", text: `${PRUNED_MARK}: output of ${what}, ${lines(text)} lines. Run it again if you need it.>>` }] };
|
|
90
|
+
}
|
|
91
|
+
|
|
92
|
+
/**
|
|
93
|
+
* Prunes everything before the watermark. The watermark sits on the last point where the prunable chars,
|
|
94
|
+
* summed from the first message, crossed a multiple of stepTokens. Only messages before the protected tail
|
|
95
|
+
* count, and the tail only moves forward, so earlier crossings never move: the pruned prefix is stable.
|
|
96
|
+
*/
|
|
97
|
+
export function pruneContext(messages: Msg[], o: HygieneOptions = HYGIENE_DEFAULTS): { messages: Msg[]; stats: HygieneStats } {
|
|
98
|
+
const lastIndices = (role: string, keep: number) =>
|
|
99
|
+
messages.flatMap((m, i) => (m.role === role ? [i] : [])).slice(-keep);
|
|
100
|
+
const tail = [...lastIndices("assistant", o.keepThinkingTurns), ...lastIndices("toolResult", o.keepToolResults)];
|
|
101
|
+
const tailStart = tail.length ? Math.min(...tail) : messages.length;
|
|
102
|
+
|
|
103
|
+
const calls = new Map<string, any>();
|
|
104
|
+
for (const m of messages) if (m.role === "assistant") for (const c of m.content ?? []) if (c.type === "toolCall") calls.set(c.id, c);
|
|
105
|
+
|
|
106
|
+
const savings = (m: Msg) =>
|
|
107
|
+
m.role === "assistant" ? assistantSavings(m, o) : m.role === "toolResult" && resultChars(m) >= o.minPruneChars ? resultChars(m) : 0;
|
|
108
|
+
|
|
109
|
+
const step = o.stepTokens * CHARS_PER_TOKEN;
|
|
110
|
+
let sum = 0;
|
|
111
|
+
let watermark = 0;
|
|
112
|
+
let prunedChars = 0;
|
|
113
|
+
for (let i = 0; i < tailStart; i++) {
|
|
114
|
+
sum += savings(messages[i]!);
|
|
115
|
+
if (Math.floor(sum / step) > Math.floor(prunedChars / step)) {
|
|
116
|
+
watermark = i + 1;
|
|
117
|
+
prunedChars = sum;
|
|
118
|
+
}
|
|
119
|
+
}
|
|
120
|
+
|
|
121
|
+
const out = messages.map((m, i) => {
|
|
122
|
+
if (i >= watermark || !savings(m)) return m;
|
|
123
|
+
return m.role === "assistant" ? pruneAssistant(m, o) : pruneResult(m, calls.get(m.toolCallId));
|
|
124
|
+
});
|
|
125
|
+
return {
|
|
126
|
+
messages: out,
|
|
127
|
+
stats: {
|
|
128
|
+
messages: messages.length,
|
|
129
|
+
watermark,
|
|
130
|
+
prunedTokens: Math.round(prunedChars / CHARS_PER_TOKEN),
|
|
131
|
+
pendingTokens: Math.round((sum - prunedChars) / CHARS_PER_TOKEN),
|
|
132
|
+
stepTokens: o.stepTokens,
|
|
133
|
+
},
|
|
134
|
+
};
|
|
135
|
+
}
|
|
136
|
+
|
|
137
|
+
/**
|
|
138
|
+
* Sent as reasoning_budget_message with every request: when the thinking budget runs out llama.cpp forces this text,
|
|
139
|
+
* then the end-of-thinking tag, and the model reads it as its own words, so it says what to do next.
|
|
140
|
+
* A llama-server older than the per-request message closes the thinking silently; cuts are then detected by length (budgetVerdict).
|
|
141
|
+
*/
|
|
142
|
+
export const BUDGET_MESSAGE = "Time is up. I will take the smallest safe next step with what I know, and write my open plan into PLAN.md.";
|
|
143
|
+
|
|
144
|
+
export type BudgetPhase = "planning" | "acting"; // right after the user spoke / between tool calls
|
|
145
|
+
export type BudgetMode = "auto" | "fixed" | "off";
|
|
146
|
+
|
|
147
|
+
/** Starting budgets, used until a model has enough samples to learn its own. */
|
|
148
|
+
export const BUDGET_DEFAULTS: Record<BudgetPhase, number> = { planning: 16384, acting: 4096 };
|
|
149
|
+
export const BUDGET_LIMITS: Record<BudgetPhase, [number, number]> = { planning: [4096, 32768], acting: [2048, 4096] }; // acting: past ~4k between tool calls it is overthinking
|
|
150
|
+
|
|
151
|
+
const BUDGET_SAMPLES_MIN = 10;
|
|
152
|
+
export const BUDGET_WINDOW = 30; // thinking lengths kept per model and phase
|
|
153
|
+
const BUDGET_PERCENTILE = 0.95;
|
|
154
|
+
const BUDGET_HEADROOM = 1.5;
|
|
155
|
+
|
|
156
|
+
/** Phase of the next request, from the OpenAI-style payload: a tool result last means the agent is mid-task. */
|
|
157
|
+
export function requestPhase(payload: any): BudgetPhase {
|
|
158
|
+
return payload?.messages?.at(-1)?.role === "tool" ? "acting" : "planning";
|
|
159
|
+
}
|
|
160
|
+
|
|
161
|
+
/**
|
|
162
|
+
* Learned budget: 1.5 × the 95th percentile of recent thinking lengths, clamped and rounded to 1k.
|
|
163
|
+
* Cuts never raise it (see budgetSample): a model that often runs away is held at its budget instead of chased.
|
|
164
|
+
* Replies that finish close under the budget raise it, a model that thinks little pulls it down.
|
|
165
|
+
* Undefined until there are enough samples.
|
|
166
|
+
*/
|
|
167
|
+
export function learnedBudget(samples: number[], phase: BudgetPhase): number | undefined {
|
|
168
|
+
if (samples.length < BUDGET_SAMPLES_MIN) return undefined;
|
|
169
|
+
const sorted = [...samples].sort((a, b) => a - b);
|
|
170
|
+
const p = sorted[Math.min(sorted.length - 1, Math.floor(sorted.length * BUDGET_PERCENTILE))]!;
|
|
171
|
+
const [lo, hi] = BUDGET_LIMITS[phase];
|
|
172
|
+
return Math.min(hi, Math.max(lo, Math.ceil((p * BUDGET_HEADROOM) / 1024) * 1024));
|
|
173
|
+
}
|
|
174
|
+
|
|
175
|
+
/** The sample to learn from a reply: a cut counts as budget / headroom, so cuts alone give back the same budget. */
|
|
176
|
+
export function budgetSample(thought: number, budget: number, cut: boolean): number {
|
|
177
|
+
return cut ? budget / BUDGET_HEADROOM : thought;
|
|
178
|
+
}
|
|
179
|
+
|
|
180
|
+
/**
|
|
181
|
+
* What the server did with the budget sent for this reply, from the thinking length (estimated, ±10%):
|
|
182
|
+
* cut near the budget, "ignored" well past it (llama-server started with its own --reasoning-budget, or too old
|
|
183
|
+
* to read thinking_budget_tokens), otherwise within.
|
|
184
|
+
*/
|
|
185
|
+
export function budgetVerdict(m: Msg, budget: number): "cut" | "ignored" | "within" {
|
|
186
|
+
if (thinkingWasCut(m)) return "cut";
|
|
187
|
+
const thought = thinkingTokens(m);
|
|
188
|
+
return thought > budget * 1.3 ? "ignored" : thought >= budget * 0.9 ? "cut" : "within";
|
|
189
|
+
}
|
|
190
|
+
|
|
191
|
+
/** Estimated thinking tokens of a reply (same chars/token as the hygiene). */
|
|
192
|
+
export function thinkingTokens(m: Msg): number {
|
|
193
|
+
const chars = (m.content ?? []).reduce((n: number, c: any) => n + (c.type === "thinking" ? (c.thinking?.length ?? 0) : 0), 0);
|
|
194
|
+
return Math.round(chars / CHARS_PER_TOKEN);
|
|
195
|
+
}
|
|
196
|
+
|
|
197
|
+
/** Whether llama.cpp cut this message's thinking at the budget. */
|
|
198
|
+
export function thinkingWasCut(m: Msg): boolean {
|
|
199
|
+
return (m.content ?? []).some((c: any) => c.type === "thinking" && (c.thinking ?? "").trimEnd().endsWith(BUDGET_MESSAGE));
|
|
200
|
+
}
|
|
201
|
+
|
|
202
|
+
/**
|
|
203
|
+
* Written to <project>/MAGI.md when missing, then appended to the system prompt on every run.
|
|
204
|
+
* English on purpose: Qwen-family models reason in English and follow English rules more reliably.
|
|
205
|
+
* Kept short: it is paid for on every request.
|
|
206
|
+
*/
|
|
207
|
+
export const MAGI_MD = `# MAGI rules for local models
|
|
208
|
+
|
|
209
|
+
These rules are appended to the system prompt by the MAGI extension. Edit them freely; delete the file to get the defaults back.
|
|
210
|
+
|
|
211
|
+
## Think less, act more
|
|
212
|
+
- Keep reasoning short: understand the step, decide, act. Do not re-plan what is already decided.
|
|
213
|
+
- Never draft code or file contents in your reasoning. Write them directly with the write/edit tool.
|
|
214
|
+
- If two attempts at the same approach fail, stop and change approach, or ask the user. Do not retry blindly.
|
|
215
|
+
- If your reasoning stopped abruptly mid-thought, or ends with "${BUDGET_MESSAGE.split(". ")[0]}.", your thinking budget ran out: in that turn make no large or irreversible change. Write your open plan into PLAN.md or take one small step you can verify; the next turn gives you a fresh budget.
|
|
216
|
+
|
|
217
|
+
## Your context is small: spend it carefully
|
|
218
|
+
- Search before reading: use rg/find to locate code, then read only the needed range (offset/limit).
|
|
219
|
+
- Do not read whole large files, and do not re-read a file you just wrote.
|
|
220
|
+
- Trim long command output: pipe through tail, head or rg.
|
|
221
|
+
- Old tool outputs and old reasoning are removed from your context automatically (marked "${PRUNED_MARK}…>>"). Anything you will need later must be written down (see below). Never copy those markers into a file.
|
|
222
|
+
|
|
223
|
+
## Keep the task state on disk
|
|
224
|
+
- For any task longer than a few steps, keep PLAN.md (a checklist) and NOTES.md (findings, decisions, commands that work).
|
|
225
|
+
- Update them after every completed step. If the context is compacted or a new session starts, continue from PLAN.md.
|
|
226
|
+
|
|
227
|
+
## Edit safely
|
|
228
|
+
- edit oldText must match the file exactly: copy it from a fresh read, keep it small and unique.
|
|
229
|
+
- If an edit fails, read that region again instead of guessing.
|
|
230
|
+
- For big new files: write a skeleton first, then fill it with edits. Never replace a whole file with a partial version.
|
|
231
|
+
|
|
232
|
+
## Verify, never assume
|
|
233
|
+
- Never invent paths, functions, APIs or flags: check with ls, rg, --help or the docs first.
|
|
234
|
+
- After a change, run the build, the tests or the program. Claim success only when a tool result proves it.
|
|
235
|
+
- If you cannot verify something, say so explicitly.
|
|
236
|
+
|
|
237
|
+
## Finish cleanly
|
|
238
|
+
- When done: say briefly what changed, how you verified it, and what is left.
|
|
239
|
+
- Answer in the user's language.
|
|
240
|
+
`;
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "pi-magi-theme",
|
|
3
|
-
"version": "0.
|
|
3
|
+
"version": "0.3.1",
|
|
4
4
|
"description": "MAGI SYSTEM theme + extension for pi (Evangelion fan art): MAGI control screen panel, three-model /magi council, MECHA SELECT model picker, angel-attack loading, seven-seal context gauge, llama-swap telemetry",
|
|
5
5
|
"keywords": [
|
|
6
6
|
"pi-package",
|