context-doctor 0.3.4 → 0.3.6

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -64,7 +64,7 @@ One run of `npx context-doctor install` writes five things (each config edit mak
64
64
  **In every Claude Code / Cowork session afterward:** all of the above via MCP, plus two more layers:
65
65
 
66
66
  - the **skill** loads whenever context work is relevant, and
67
- - the **hook runs on every single prompt you send**: it measures the session's real size in ~100ms. Under 80k tokens it stays completely silent. Above, it injects a note the model sees with your message — actual token count, cost per message, the single largest recoverable waste — with instructions to work leaner and offer you compaction. It re-fires only after ~40% further growth, and can never break a prompt (any failure exits silently).
67
+ - the **hook runs on every single prompt you send**: lean sessions cost a ~1ms file-size check; once a session is heavy it profiles on growth events and injects a note the model sees with your message — actual token count, cost per message, the single largest recoverable waste — with instructions to work leaner and offer you compaction. It re-fires only after ~40% further growth, can never break a prompt (any failure exits silently), and logs each deep check to a small local ledger that feeds `context-doctor report`.
68
68
 
69
69
  **What it never does:** delete or rewrite your history without asking (pruning is consent-only, and the model writes the replacement summary so nothing is lost silently), send data anywhere (everything runs on your machine), or touch an API key.
70
70
 
@@ -83,6 +83,18 @@ MCP is just one of six delivery mechanisms. It's only required when you want the
83
83
 
84
84
  Practical upshot: a developer who only wants cheaper, faster API calls never touches MCP (proxy + CLI). A Claude Code user gets the hook and skill without MCP either — the MCP server just adds in-chat tools on top. `install` sets up all of it at once precisely so you don't have to think about which mechanism is which.
85
85
 
86
+ ## All commands at a glance
87
+
88
+ | Command | What it does |
89
+ |---|---|
90
+ | `context-doctor install` / `uninstall` | Wire (or remove) everything: MCP for Claude Desktop/Code/Cursor, the Agent Skill, the every-prompt hook |
91
+ | `context-doctor analyze <file>` | Profile a conversation: token breakdown, findings, cost + latency estimates |
92
+ | `context-doctor optimize <file>` | Apply the safe fixes; `--strategy prune-history` for consented lossy compaction |
93
+ | `context-doctor session [file]` | Profile a Claude Code session transcript (defaults to your most recent; `--list` to browse) |
94
+ | `context-doctor report` | Machine-wide impact report: exact proxy savings, hook activity, recoverable waste in recent sessions |
95
+ | `context-doctor proxy` | Always-on local proxy that optimizes every Anthropic/OpenAI API request in flight (`/stats` for cumulative savings) |
96
+ | `context-doctor hook` | The every-prompt Claude Code hook (registered by `install`; you never run this yourself) |
97
+
86
98
  ## What "always-on" means, per surface
87
99
 
88
100
  | Where you run LLMs | Mechanism | Guarantee |
@@ -246,6 +258,34 @@ Everything the optimizer does is inspectable: it prints exactly which messages c
246
258
 
247
259
  `skills/context-doctor/SKILL.md` (installed by `npx context-doctor install`) teaches Claude to practice context hygiene proactively: summarize big tool results after consuming them, never re-paste duplicated content, keep stable content cache-friendly, and offer compaction when a session gets heavy — so sessions get inherently leaner without you asking.
248
260
 
261
+ ## Measuring the impact: `context-doctor report`
262
+
263
+ ```bash
264
+ npx context-doctor report
265
+ ```
266
+
267
+ One report for your whole machine, led by a headline of **tokens context-doctor saved**, built only from measured sources:
268
+
269
+ - **exact** proxy savings (real before/after on every request),
270
+ - **exact** savings from every optimization applied via the CLI or the in-chat tools — split by model family (Claude vs GPT), with dollar estimates,
271
+ - **observed per-session shrinkage**: real context reductions recorded between the hook's deep checks after hygiene warnings — shown per session in the table alongside remaining waste.
272
+
273
+ Honest measurement note: proxy numbers are exact. Session numbers are measured-now. What no tool can report is the counterfactual — tokens Claude *avoided* adding because of the hygiene guidance — since the same session can't be re-run without it. The report says so instead of inventing a number.
274
+
275
+ ## Performance: what context-doctor itself costs
276
+
277
+ A tool that promises speed must be near-free. Measured overhead per touchpoint:
278
+
279
+ | Touchpoint | When it runs | Overhead |
280
+ |---|---|---|
281
+ | Every-prompt hook (Claude Code) | Every prompt | **~80ms** (Node startup; logic ~1ms). Lean sessions exit on a single `stat()` — the transcript is never read. Full profiling (~200ms on a 4MB session) happens only when the transcript has grown ~40% since last checked |
282
+ | MCP server | Spawned once per app session | Tools run only when called; standing instructions cost **~110 tokens per conversation** — deliberately terse |
283
+ | Proxy | Per API request | ~1–3ms of CPU (parse → optimize → re-serialize) against typical model latencies of hundreds of ms; responses stream through chunk-by-chunk, never buffered |
284
+ | Skill | Loads only when relevant | ~1k tokens while active; its always-present description is ~60 tokens |
285
+ | CLI / library | Only when you run it | Not in any hot path |
286
+
287
+ Net effect is strongly negative overhead: the tokens these touchpoints save on every subsequent call dwarf what they cost.
288
+
249
289
  ## Why token counts are "~"
250
290
 
251
291
  Exact counts require each provider's private tokenizer. `context-doctor` uses a calibrated chars-per-token heuristic (denser for code/JSON) that lands within ~10% — plenty accurate for finding what's heavy and measuring savings, and it keeps the tool fully offline with zero configuration.
package/dist/cli.js CHANGED
@@ -18,6 +18,8 @@ import { startProxy } from "./proxy.js";
18
18
  import { runInstall, runUninstall } from "./install.js";
19
19
  import { listSessions, parseSessionFile } from "./session.js";
20
20
  import { runHook } from "./hook.js";
21
+ import { buildImpactReport } from "./impact.js";
22
+ import { recordLedger } from "./ledger.js";
21
23
  const HELP = `context-doctor — profile and optimize LLM context windows
22
24
 
23
25
  Usage:
@@ -32,6 +34,8 @@ Usage:
32
34
  (default: the most recent session; --list to browse)
33
35
  context-doctor hook Claude Code UserPromptSubmit hook (installed
34
36
  automatically by \`install\`; reads hook JSON on stdin)
37
+ context-doctor report Impact report: exact proxy savings, hook activity,
38
+ and remaining recoverable waste in recent sessions
35
39
 
36
40
  Input: a conversation JSON file (OpenAI or Anthropic message format, or a bare
37
41
  message array). Use "-" to read from stdin.
@@ -119,6 +123,10 @@ function main() {
119
123
  void runHook();
120
124
  return;
121
125
  }
126
+ if (args.command === "report") {
127
+ void buildImpactReport(args.port).then((r) => console.log(r));
128
+ return;
129
+ }
122
130
  if (args.command === "session") {
123
131
  if (args.list) {
124
132
  const sessions = listSessions();
@@ -215,6 +223,9 @@ function main() {
215
223
  console.log(output);
216
224
  }
217
225
  const saved = result.tokensBefore - result.tokensAfter;
226
+ if (saved > 0) {
227
+ recordLedger({ ev: "optimize", src: "cli", saved, model: result.conversation?.model });
228
+ }
218
229
  const pct = result.tokensBefore > 0 ? Math.round((saved / result.tokensBefore) * 100) : 0;
219
230
  console.error(`\ncontext-doctor: ${formatTokens(result.tokensBefore)} → ${formatTokens(result.tokensAfter)} tokens ` +
220
231
  `(saved ~${formatTokens(saved)}, ${pct}%) via ${result.applied.length} change(s)` +
package/dist/hook.js CHANGED
@@ -10,9 +10,8 @@
10
10
  * Registered by `context-doctor install` under hooks.UserPromptSubmit in
11
11
  * ~/.claude/settings.json; removed by `context-doctor uninstall`.
12
12
  */
13
- import { existsSync, readFileSync, writeFileSync } from "node:fs";
14
- import { homedir } from "node:os";
15
- import { join } from "node:path";
13
+ import { existsSync, readFileSync, statSync, writeFileSync } from "node:fs";
14
+ import { recordLedger, statePath } from "./ledger.js";
16
15
  import { parseConversation } from "./parse.js";
17
16
  import { profileConversation } from "./profile.js";
18
17
  import { parseSessionFile } from "./session.js";
@@ -22,9 +21,13 @@ import { formatUsd } from "./pricing.js";
22
21
  const WARN_TOKENS = 80_000;
23
22
  /** Re-nudge only after the context grows another 40% — one reminder, not a nag. */
24
23
  const REGROWTH_FACTOR = 1.4;
25
- function statePath() {
26
- return process.env.CONTEXT_DOCTOR_HOOK_STATE ?? join(homedir(), ".claude", ".context-doctor-hook-state.json");
27
- }
24
+ /**
25
+ * Fast-path gate: text tokens are at least ~4 bytes each and the transcript
26
+ * carries JSON overhead on top, so a file smaller than this cannot possibly
27
+ * hold WARN_TOKENS of context. Lean sessions cost one stat() call — the
28
+ * transcript is never even read.
29
+ */
30
+ const MIN_BYTES_FOR_WARN = WARN_TOKENS * 4;
28
31
  async function readStdin() {
29
32
  const chunks = [];
30
33
  for await (const chunk of process.stdin)
@@ -38,13 +41,15 @@ export async function runHook() {
38
41
  const transcriptPath = input.transcript_path;
39
42
  if (!transcriptPath || !existsSync(transcriptPath))
40
43
  return;
41
- const parsed = parseSessionFile(transcriptPath);
42
- if (parsed.messageCount === 0)
43
- return;
44
- const profile = profileConversation(parseConversation(parsed.conversationJson), parsed.model);
45
- if (profile.totalTokens < WARN_TOKENS)
44
+ // Fast path 1: a small transcript cannot exceed the threshold — exit on a
45
+ // single stat() without reading the file. This is the every-prompt cost
46
+ // for lean sessions: ~1ms.
47
+ const sizeBytes = statSync(transcriptPath).size;
48
+ if (sizeBytes < MIN_BYTES_FOR_WARN)
46
49
  return;
47
- // Per-session rate limit so the nudge fires on growth, not on every prompt.
50
+ // Fast path 2: growth gate BEFORE parsing. If the file hasn't grown ~40%
51
+ // since the last full parse, nothing new can trigger — exit without the
52
+ // expensive read. Heavy-but-quiet sessions cost one stat + tiny state read.
48
53
  const sessionId = input.session_id ?? transcriptPath;
49
54
  let state = {};
50
55
  try {
@@ -53,11 +58,24 @@ export async function runHook() {
53
58
  catch {
54
59
  /* first run */
55
60
  }
56
- const lastWarnedAt = state[sessionId] ?? 0;
57
- if (profile.totalTokens < lastWarnedAt * REGROWTH_FACTOR)
61
+ const rawPrev = state[sessionId];
62
+ // Migrate pre-0.3.5 numeric entries ({tokens only}) to the new shape.
63
+ const prev = typeof rawPrev === "number" ? { t: rawPrev, b: 0 } : rawPrev ?? { t: 0, b: 0 };
64
+ if (prev.b > 0 && sizeBytes < prev.b * REGROWTH_FACTOR)
65
+ return;
66
+ // Slow path (growth events only): full parse + profile.
67
+ const parsed = parseSessionFile(transcriptPath);
68
+ if (parsed.messageCount === 0)
58
69
  return;
59
- const entries = Object.entries({ ...state, [sessionId]: profile.totalTokens });
70
+ const profile = profileConversation(parseConversation(parsed.conversationJson), parsed.model);
71
+ // Record this parse so the next prompts take fast path 2.
72
+ const shouldWarn = profile.totalTokens >= WARN_TOKENS && profile.totalTokens >= prev.t * REGROWTH_FACTOR;
73
+ const nextState = { t: shouldWarn ? profile.totalTokens : prev.t, b: sizeBytes };
74
+ const entries = Object.entries({ ...state, [sessionId]: nextState });
60
75
  writeFileSync(statePath(), JSON.stringify(Object.fromEntries(entries.slice(-100))));
76
+ recordLedger({ ev: "check", sid: sessionId.slice(0, 12), tok: profile.totalTokens, warn: shouldWarn });
77
+ if (!shouldWarn)
78
+ return;
61
79
  const lines = [
62
80
  `This session's context is at ~${formatTokens(profile.totalTokens)} tokens` +
63
81
  (profile.usagePct ? ` (${profile.usagePct.toFixed(0)}% of the window)` : "") +
@@ -0,0 +1,11 @@
1
+ /**
2
+ * `context-doctor report` — one impact report for this machine.
3
+ *
4
+ * Honesty rules baked in: proxy numbers are EXACT (real before/after on every
5
+ * request). Session numbers are MEASURED-NOW (current size + what optimization
6
+ * would still recover). The behavioral counterfactual — what Claude avoided
7
+ * wasting because of hygiene guidance — cannot be measured by anyone: the same
8
+ * session cannot be re-run without it. The report says so instead of inventing
9
+ * a number.
10
+ */
11
+ export declare function buildImpactReport(proxyPort?: number): Promise<string>;
package/dist/impact.js ADDED
@@ -0,0 +1,147 @@
1
+ /**
2
+ * `context-doctor report` — one impact report for this machine.
3
+ *
4
+ * Honesty rules baked in: proxy numbers are EXACT (real before/after on every
5
+ * request). Session numbers are MEASURED-NOW (current size + what optimization
6
+ * would still recover). The behavioral counterfactual — what Claude avoided
7
+ * wasting because of hygiene guidance — cannot be measured by anyone: the same
8
+ * session cannot be re-run without it. The report says so instead of inventing
9
+ * a number.
10
+ */
11
+ import { readLedger } from "./ledger.js";
12
+ import { listSessions, parseSessionFile } from "./session.js";
13
+ import { parseConversation } from "./parse.js";
14
+ import { profileConversation } from "./profile.js";
15
+ import { formatTokens } from "./tokens.js";
16
+ import { formatUsd, inputCostUsd, pricingFor } from "./pricing.js";
17
+ /** Sessions larger than this are skipped in the report (keeps it snappy). */
18
+ const MAX_SESSION_BYTES = 30 * 1024 * 1024;
19
+ async function fetchProxyStats(port) {
20
+ try {
21
+ const res = await fetch(`http://127.0.0.1:${port}/stats`, { signal: AbortSignal.timeout(500) });
22
+ if (!res.ok)
23
+ return null;
24
+ return (await res.json());
25
+ }
26
+ catch {
27
+ return null;
28
+ }
29
+ }
30
+ export async function buildImpactReport(proxyPort = 8787) {
31
+ const lines = [];
32
+ lines.push("CONTEXT DOCTOR — impact report");
33
+ lines.push("═".repeat(56));
34
+ const ledger = readLedger();
35
+ const checks = ledger.filter((e) => e.ev === "check" || e.ev === undefined);
36
+ const optimizes = ledger.filter((e) => e.ev === "optimize");
37
+ // Observed per-session reductions: when a session SHRANK between two deep
38
+ // checks (compaction/cleanup after a warning), that drop is measured fact.
39
+ const bySession = new Map();
40
+ for (const c of checks) {
41
+ if (!c.sid || typeof c.tok !== "number")
42
+ continue;
43
+ const s = bySession.get(c.sid) ?? { toks: [], warns: 0 };
44
+ s.toks.push(c.tok);
45
+ if (c.warn)
46
+ s.warns++;
47
+ bySession.set(c.sid, s);
48
+ }
49
+ const reductionBySession = new Map();
50
+ for (const [sid, s] of bySession) {
51
+ let reduction = 0;
52
+ for (let i = 1; i < s.toks.length; i++) {
53
+ if (s.toks[i] < s.toks[i - 1])
54
+ reduction += s.toks[i - 1] - s.toks[i];
55
+ }
56
+ reductionBySession.set(sid, reduction);
57
+ }
58
+ const totalReduction = [...reductionBySession.values()].reduce((a, b) => a + b, 0);
59
+ // Optimize-event savings, split by model family (claude / gpt / other).
60
+ const optimizeSaved = optimizes.reduce((s, e) => s + (e.saved ?? 0), 0);
61
+ const savedByFamily = new Map();
62
+ let optimizeUsd = 0;
63
+ for (const e of optimizes) {
64
+ const family = /claude/i.test(e.model ?? "") ? "claude" : /gpt|^o\d/i.test(e.model ?? "") ? "gpt" : "other";
65
+ savedByFamily.set(family, (savedByFamily.get(family) ?? 0) + (e.saved ?? 0));
66
+ const pricing = pricingFor(e.model);
67
+ if (pricing && e.saved)
68
+ optimizeUsd += inputCostUsd(e.saved, pricing);
69
+ }
70
+ const proxy = await fetchProxyStats(proxyPort);
71
+ const proxySaved = proxy?.tokensSaved ?? 0;
72
+ // -- Headline: what context-doctor has saved ----------------------------------
73
+ const totalSaved = proxySaved + optimizeSaved + totalReduction;
74
+ lines.push("Tokens context-doctor saved (measured)");
75
+ lines.push("─".repeat(56));
76
+ lines.push(`TOTAL: ~${formatTokens(totalSaved)} tokens`);
77
+ if (proxy) {
78
+ lines.push(` · proxy (exact, current run): ${formatTokens(proxySaved)} across ${proxy.optimizedRequests}/${proxy.requests} requests ≈ ${formatUsd(proxy.estUsdSaved)}`);
79
+ }
80
+ else {
81
+ lines.push(` · proxy: not running on :${proxyPort} (its exact savings appear here while it runs)`);
82
+ }
83
+ const familyNote = [...savedByFamily.entries()]
84
+ .filter(([, v]) => v > 0)
85
+ .map(([k, v]) => `${k}: ${formatTokens(v)}`)
86
+ .join(", ");
87
+ lines.push(` · optimizations applied via CLI/chat tools (exact): ${formatTokens(optimizeSaved)} over ${optimizes.length} run(s)` +
88
+ (familyNote ? ` [${familyNote}]` : "") +
89
+ (optimizeUsd > 0 ? ` ≈ ${formatUsd(optimizeUsd)}` : ""));
90
+ lines.push(` · observed session shrinkage after hygiene warnings: ${formatTokens(totalReduction)}`);
91
+ lines.push("");
92
+ // -- Hook activity ------------------------------------------------------------
93
+ lines.push("Hygiene activity (every-prompt hook)");
94
+ lines.push("─".repeat(56));
95
+ if (checks.length > 0) {
96
+ const warnings = checks.filter((e) => e.warn).length;
97
+ lines.push(`${checks.length} deep context checks across ${bySession.size} session(s); ${warnings} warning(s) delivered to the model.`);
98
+ lines.push("(Prompt-level fast checks are not logged — they cost ~1ms and leave no trace by design.)");
99
+ }
100
+ else {
101
+ lines.push("No hook activity recorded yet (ledger appears after the first deep check of a heavy session).");
102
+ }
103
+ lines.push("");
104
+ // -- Measured-now: recent session profiles ------------------------------------
105
+ lines.push("Your recent sessions — waste still recoverable today");
106
+ lines.push("─".repeat(56));
107
+ const sessions = listSessions(8).filter((s) => s.sizeBytes <= MAX_SESSION_BYTES);
108
+ if (sessions.length === 0) {
109
+ lines.push("No Claude Code session transcripts found.");
110
+ }
111
+ else {
112
+ let totalTokens = 0;
113
+ let totalWaste = 0;
114
+ let totalWasteUsd = 0;
115
+ for (const s of sessions) {
116
+ try {
117
+ const parsed = parseSessionFile(s.path);
118
+ if (parsed.messageCount === 0)
119
+ continue;
120
+ const p = profileConversation(parseConversation(parsed.conversationJson), parsed.model);
121
+ totalTokens += p.totalTokens;
122
+ totalWaste += p.totalEstSavings;
123
+ if (p.cost)
124
+ totalWasteUsd += p.cost.savingsPerCallUsd;
125
+ const wastePct = p.totalTokens > 0 ? Math.round((p.totalEstSavings / p.totalTokens) * 100) : 0;
126
+ // Ledger sids are the session UUID's first 12 chars (= filename prefix).
127
+ const sid = (s.path.split("/").pop() ?? "").slice(0, 12);
128
+ const saved = reductionBySession.get(sid) ?? 0;
129
+ const warns = bySession.get(sid)?.warns ?? 0;
130
+ lines.push(` ${(parsed.title ?? s.path.split("/").pop() ?? "session").slice(0, 40).padEnd(42)} ` +
131
+ `${formatTokens(p.totalTokens).padStart(6)} tok · waste ${String(wastePct).padStart(2)}%` +
132
+ (saved > 0 ? ` · saved ${formatTokens(saved)}` : warns > 0 ? ` · ${warns} warning(s)` : ""));
133
+ }
134
+ catch {
135
+ /* unreadable session — skip */
136
+ }
137
+ }
138
+ lines.push("");
139
+ lines.push(`Across these sessions: ~${formatTokens(totalTokens)} tokens held; ~${formatTokens(totalWaste)} still recoverable` +
140
+ (totalWasteUsd > 0 ? ` (≈ ${formatUsd(totalWasteUsd)} of input per message sent in them)` : ""));
141
+ }
142
+ lines.push("");
143
+ lines.push("What no report can show: tokens Claude AVOIDED adding thanks to the standing");
144
+ lines.push("hygiene guidance — the same session cannot be re-run without it. Proxy numbers");
145
+ lines.push("above are exact; session numbers are what optimization would still save now.");
146
+ return lines.join("\n");
147
+ }
@@ -0,0 +1,24 @@
1
+ /**
2
+ * Local activity ledger: one JSONL line per notable event, feeding
3
+ * `context-doctor report`. Best-effort by design — a ledger failure must
4
+ * never break a prompt, a tool call, or an optimize run.
5
+ *
6
+ * Event shapes (all carry ts):
7
+ * check — hook deep-parsed a session {ev?: undefined|"check", sid, tok, warn}
8
+ * (pre-0.3.6 hook entries have no `ev` field; treated as checks)
9
+ * optimize — an optimization was applied {ev: "optimize", src: "cli"|"mcp", saved, model?}
10
+ */
11
+ export interface LedgerEntry {
12
+ ts: number;
13
+ ev?: "check" | "optimize";
14
+ sid?: string;
15
+ tok?: number;
16
+ warn?: boolean;
17
+ src?: "cli" | "mcp";
18
+ saved?: number;
19
+ model?: string;
20
+ }
21
+ export declare function statePath(): string;
22
+ export declare function ledgerPath(): string;
23
+ export declare function recordLedger(entry: Omit<LedgerEntry, "ts">): void;
24
+ export declare function readLedger(): LedgerEntry[];
package/dist/ledger.js ADDED
@@ -0,0 +1,54 @@
1
+ /**
2
+ * Local activity ledger: one JSONL line per notable event, feeding
3
+ * `context-doctor report`. Best-effort by design — a ledger failure must
4
+ * never break a prompt, a tool call, or an optimize run.
5
+ *
6
+ * Event shapes (all carry ts):
7
+ * check — hook deep-parsed a session {ev?: undefined|"check", sid, tok, warn}
8
+ * (pre-0.3.6 hook entries have no `ev` field; treated as checks)
9
+ * optimize — an optimization was applied {ev: "optimize", src: "cli"|"mcp", saved, model?}
10
+ */
11
+ import { appendFileSync, existsSync, readFileSync, statSync, writeFileSync } from "node:fs";
12
+ import { homedir } from "node:os";
13
+ import { dirname, join } from "node:path";
14
+ export function statePath() {
15
+ return process.env.CONTEXT_DOCTOR_HOOK_STATE ?? join(homedir(), ".claude", ".context-doctor-hook-state.json");
16
+ }
17
+ export function ledgerPath() {
18
+ return join(dirname(statePath()), ".context-doctor-ledger.jsonl");
19
+ }
20
+ export function recordLedger(entry) {
21
+ const path = ledgerPath();
22
+ try {
23
+ // Cap growth: past ~256KB keep the most recent 500 entries.
24
+ if (existsSync(path) && statSync(path).size > 256 * 1024) {
25
+ const lines = readFileSync(path, "utf8").trimEnd().split("\n");
26
+ writeFileSync(path, lines.slice(-500).join("\n") + "\n");
27
+ }
28
+ appendFileSync(path, JSON.stringify({ ts: Date.now(), ...entry }) + "\n");
29
+ }
30
+ catch {
31
+ /* best-effort */
32
+ }
33
+ }
34
+ export function readLedger() {
35
+ const path = ledgerPath();
36
+ if (!existsSync(path))
37
+ return [];
38
+ try {
39
+ return readFileSync(path, "utf8")
40
+ .trimEnd()
41
+ .split("\n")
42
+ .flatMap((line) => {
43
+ try {
44
+ return [JSON.parse(line)];
45
+ }
46
+ catch {
47
+ return [];
48
+ }
49
+ });
50
+ }
51
+ catch {
52
+ return [];
53
+ }
54
+ }
package/dist/mcp.js CHANGED
@@ -20,18 +20,17 @@ import { profileConversation } from "./profile.js";
20
20
  import { optimizeConversation } from "./optimize.js";
21
21
  import { renderProfile } from "./report.js";
22
22
  import { formatTokens } from "./tokens.js";
23
+ import { recordLedger } from "./ledger.js";
23
24
  /**
24
25
  * Server instructions are injected by MCP clients (Claude Desktop, Cursor, …)
25
26
  * into the system context of EVERY conversation where this server is enabled.
26
27
  * This is what upgrades plain-chat apps from "tools available on request" to
27
28
  * "standing context-hygiene instructions in every chat" — no hook needed.
28
29
  */
29
- const SERVER_INSTRUCTIONS = `context-doctor keeps this conversation's context window lean, fast, and cheap. Standing instructions for every conversation:
30
- 1. Practice context hygiene continuously: after consuming a large tool result or pasted document, carry forward a one-sentence summary instead of re-quoting it; reference earlier content rather than repeating it; never inline base64.
31
- 2. When the conversation has grown long (roughly 30+ turns) or contains several large pastes, proactively offer to profile it and call profile_context do not wait to be asked.
32
- 3. When the user asks anything about token usage, cost, latency, or "what's eating my context", call profile_context immediately rather than estimating by eye.
33
- 4. When optimize_context returns a pruned-turns digest, YOU write the replacement summary (dense, factual, ≤150 tokens) as instructed in the result.`;
34
- const server = new McpServer({ name: "context-doctor", version: "0.3.4" }, { instructions: SERVER_INSTRUCTIONS });
30
+ // Kept deliberately terse: these ride in EVERY conversation's context, and a
31
+ // context-saving tool must not itself be context overhead (~110 tokens).
32
+ const SERVER_INSTRUCTIONS = `Context hygiene, always: summarize large pastes/tool results instead of carrying them verbatim; reference earlier content, don't re-quote; never inline base64. Past ~30 turns or several large pastes, proactively offer to run profile_context. Any question about tokens, cost, or latency: call profile_context, don't estimate. If optimize_context returns a pruned-turns digest, you write the ≤150-token replacement summary.`;
33
+ const server = new McpServer({ name: "context-doctor", version: "0.3.6" }, { instructions: SERVER_INSTRUCTIONS });
35
34
  const STRATEGY_IDS = ["dedupe", "trim-tool-results", "strip-base64", "prune-history"];
36
35
  server.tool("profile_context", "Profile an LLM conversation or prompt: token breakdown by category, largest messages, and actionable findings about wasted context (duplicates, oversized tool results, base64 blobs, cache-unfriendly ordering). Accepts OpenAI/Anthropic conversation JSON or raw text. Call this immediately whenever the user asks about token usage, context size, LLM cost, or latency — and proactively offer it once a conversation grows long or accumulates large pasted content.", {
37
36
  conversation: z.string().describe("Conversation JSON (OpenAI or Anthropic format, or bare message array) or raw prompt text"),
@@ -53,6 +52,9 @@ server.tool("optimize_context", "Rewrite a conversation to reclaim tokens using
53
52
  maxToolResultTokens: max_tool_result_tokens,
54
53
  });
55
54
  const saved = result.tokensBefore - result.tokensAfter;
55
+ if (saved > 0) {
56
+ recordLedger({ ev: "optimize", src: "mcp", saved, model: result.conversation?.model });
57
+ }
56
58
  const summary = `Saved ~${formatTokens(saved)} tokens (${formatTokens(result.tokensBefore)} → ${formatTokens(result.tokensAfter)}) ` +
57
59
  `via ${result.applied.length} change(s):\n` +
58
60
  result.applied.map((c) => `- [${c.strategy}] message #${c.messageIndex}: ${c.note} (~${formatTokens(c.tokensSaved)})`).join("\n");
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "context-doctor",
3
- "version": "0.3.4",
3
+ "version": "0.3.6",
4
4
  "description": "Profile and optimize LLM context windows. See what's eating your tokens and fix it — works with Claude, GPT, Gemini, and any MCP-capable AI app.",
5
5
  "keywords": [
6
6
  "llm",
@@ -38,7 +38,7 @@
38
38
  "node": ">=18"
39
39
  },
40
40
  "scripts": {
41
- "build": "tsc",
41
+ "build": "tsc && node -e \"const fs=require('fs');['dist/cli.js','dist/mcp.js'].forEach(f=>fs.chmodSync(f,0o755))\"",
42
42
  "prepublishOnly": "npm run build",
43
43
  "dev": "tsc --watch",
44
44
  "test": "npm run build && node --test dist/test/smoke.test.js dist/test/proxy.test.js dist/test/hook.test.js"