claude-token-saver 3.35.0 → 3.36.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.ko.md +605 -0
- package/README.md +493 -352
- package/package.json +2 -2
- package/src/model-rules.js +61 -0
- package/README.en.md +0 -706
package/README.md
CHANGED
|
@@ -1,585 +1,726 @@
|
|
|
1
|
-
|
|
1
|
+
**English** · [한국어](./README.ko.md)
|
|
2
2
|
|
|
3
3
|
[](https://www.npmjs.com/package/claude-token-saver)
|
|
4
4
|
|
|
5
|
-
🌐 **[
|
|
5
|
+
🌐 **[Project page](https://rootstudioyaml.github.io/claude-token-saver/)**
|
|
6
6
|
|
|
7
7
|
# claude-token-saver
|
|
8
8
|
|
|
9
|
-
|
|
9
|
+
**Shows what it saved, on two lines.** It moves the easy work your expensive model keeps repeating onto cheaper ones, and turns documents the model cannot read into Markdown. Both figures are ledger entries rather than estimates, and whichever saved more takes the top line. Zero dependencies, one-line install.
|
|
10
10
|
|
|
11
|
-

|
|
12
12
|
|
|
13
13
|
```bash
|
|
14
14
|
npm i -g claude-token-saver
|
|
15
15
|
```
|
|
16
16
|
|
|
17
|
-
|
|
17
|
+
Four numbers are the whole pitch.
|
|
18
18
|
|
|
19
|
-
-
|
|
20
|
-
-
|
|
21
|
-
-
|
|
19
|
+
- **Beats every single model on public benchmark data**: the shipped tier criteria score 59.1% on 11,696 LLMRouterBench instances against the best single model's 57.9% — at 31% less than gpt-5, 64% less than gemini-2.5-pro ([benchmark](./docs/BENCHMARK.md))
|
|
20
|
+
- **95.8% fewer tokens per document**: a 30MB deck read as Markdown cost 22,610 tokens instead of 540,429 ([evidence](#-doc2md--documents-become-markdown-before-the-model-reads-them))
|
|
21
|
+
- **18.6% lower cost**: measured before/after adopting the Harness principles ([evidence](#real-world-impact--beforeafter-report))
|
|
22
|
+
- **Routing savings are a per-run ledger**: the price difference of each delegated run, not an estimate ([evidence](#-the-savings-figure-is-a-ledger-entry-not-an-estimate))
|
|
22
23
|
|
|
23
|
-
|
|
24
|
+
```text
|
|
25
|
+
accuracy total cost, 11,696 queries
|
|
26
|
+
tier-criteria router 59.1% ← best $268 ▰▰▰▰▰▱▱▱▱▱▱▱▱▱▱
|
|
27
|
+
gpt-5 57.8% $388 ▰▰▰▰▰▰▰▰▱▱▱▱▱▱▱
|
|
28
|
+
gemini-2.5-pro 57.9% $734 ▰▰▰▰▰▰▰▰▰▰▰▰▰▰▰
|
|
29
|
+
```
|
|
30
|
+
|
|
31
|
+
Since v3.35.0 spend is visible too: month-to-date spend shows as `💵 Sep $42`, and on LiteLLM gateways (Bedrock and friends) with no 5h/7d caps, your key budget renders as a `🔑 budget ▰▱ 34% $34/$100` gauge.
|
|
32
|
+
|
|
33
|
+
## Four parts, working together
|
|
34
|
+
|
|
35
|
+
| | What it does | Effect |
|
|
36
|
+
|---|---|---|
|
|
37
|
+
| 🔀 **Routing** | Delegates recurring easy work to cheaper models | Savings recorded per run in a ledger; criteria [benchmarked](./docs/BENCHMARK.md) on public data |
|
|
38
|
+
| 📄 **Document conversion** | Turns pptx/xlsx/pdf/docx/fig into Markdown before the model reads them | **510,000 tokens** saved on one deck ([below](#-doc2md--documents-become-markdown-before-the-model-reads-them)) |
|
|
39
|
+
| 🅷 **Harness** | Blocks the token-burning habits: unevidenced "done", skipped verification (5 principles) | **−18.6% cost** ([measured](#real-world-impact--beforeafter-report)) |
|
|
40
|
+
| ⚙️ **Ratchet** | Freezes each error you hit into a rule | Same mistake stops recurring |
|
|
41
|
+
|
|
42
|
+
One install sets up all four. The measured −18.6% comes from the harness and ratchet; routing and conversion savings sit on top of it.
|
|
43
|
+
|
|
44
|
+
The two savings figures are never added together, because they answer different questions. Routing says "the same work ran on a cheaper model". Conversion says "a file you could not read became readable, without pushing the original through the context window". The statusline gives each its own line and puts the larger one first.
|
|
45
|
+
|
|
46
|
+
## Contents
|
|
47
|
+
|
|
48
|
+
- **Start here**: [Getting started](#getting-started) · [Reading the statusline](#reading-the-statusline) · [Commands](#commands)
|
|
49
|
+
- **Savings**: [The routing ledger](#-the-savings-figure-is-a-ledger-entry-not-an-estimate) · [route-scan](#-route-scan--this-recurring-task-could-run-on-a-cheaper-tier) · [doc2md](#-doc2md--documents-become-markdown-before-the-model-reads-them) · [seed](#-seed-delegation-that-works-from-the-first-session) · [Benchmark](./docs/BENCHMARK.md)
|
|
50
|
+
- **Guardrails**: [Harness](#-harness-mode) · [compact-window](#-compact-window--pin-where-a-1m-session-compacts) · [Korean writing guidance](#-korean-writing-guidance)
|
|
51
|
+
- **Spend & environments**: [Monthly spend · LiteLLM key budget](#litellm-your-key-budget-stands-in-for-the-missing-5h7d-caps-v3350) · [Gateways (Bedrock/Vertex)](#-behind-a-gateway-bedrock--vertex) · [Spike issue codes](#spike-issue-codes) · [Measured impact](#real-world-impact--beforeafter-report)
|
|
52
|
+
|
|
53
|
+
## Getting started
|
|
54
|
+
|
|
55
|
+
**Prerequisite:** Node.js ≥ 18 (`node -v` · macOS `brew install node` · Windows `winget install OpenJS.NodeJS.LTS` · Linux/WSL: [nvm](https://github.com/nvm-sh/nvm) recommended)
|
|
56
|
+
|
|
57
|
+
```bash
|
|
58
|
+
npm uninstall -g claude-cache-monitor # (previous-package users only)
|
|
59
|
+
npm i -g claude-token-saver
|
|
60
|
+
```
|
|
61
|
+
|
|
62
|
+
The statusline appears at the bottom of Claude Code right away. If auto-registration was skipped (`--ignore-scripts`, sudo, sandboxed installs), run `claude-token-saver install`.
|
|
63
|
+
|
|
64
|
+
One install sets up everything: **statusline, Skill, SessionStart hook, the 🅷 Harness (5 principles), and a first route-scan.** The harness and the Korean writing guidance **show what they add and ask before enabling it.** The harness is **appended** to `~/.claude/CLAUDE.md` as a marked block (your existing content is backed up and preserved) and is left alone if one is already there.
|
|
65
|
+
|
|
66
|
+
Outside a terminal — npm `postinstall`, CI, piped stdin — the question is skipped and the old defaults apply. Use `--yes` or `--no-input` to skip it deliberately, `CTS_NO_HARNESS=1 npm i -g claude-token-saver` to skip the harness entirely, and `claude-token-saver harness uninit --global` to undo it.
|
|
24
67
|
|
|
25
|
-
|
|
68
|
+
> ⚠️ Avoid `sudo` global installs — the Skill lands in root's `~/.claude` instead of yours. Use nvm/fnm/Volta or `npm config set prefix ~/.npm-global`.
|
|
69
|
+
|
|
70
|
+
### What the install turns on, and what stays manual
|
|
71
|
+
|
|
72
|
+
Everything that costs nothing until it is needed is on after a plain install. The only manual items are the ones that change Claude Code's own settings or need a human to pick a scope.
|
|
73
|
+
|
|
74
|
+
| Feature | After install | How to turn it off |
|
|
26
75
|
|---|---|---|
|
|
27
|
-
|
|
|
28
|
-
|
|
|
29
|
-
|
|
|
30
|
-
|
|
|
76
|
+
| statusline (diagnostic chips, savings ledger) | on | `claude-token-saver uninstall` |
|
|
77
|
+
| `/claude-token-saver` Skill | on | same |
|
|
78
|
+
| SessionStart hook (route-scan refresh) | on | same |
|
|
79
|
+
| UserPromptSubmit hook (brief injection) | on | same |
|
|
80
|
+
| First route-scan (last 14 days of logs) | runs once during the install | n/a |
|
|
81
|
+
| 🅷 Harness 5 principles (`~/.claude/CLAUDE.md`) | on (shown and confirmed once at a terminal) | `harness uninit --global`, `CTS_NO_HARNESS=1` |
|
|
82
|
+
| doc2md hooks (Read, Edit/Write, prompt) | on | `doc2md off`, `CTS_NO_DOC2MD=1` |
|
|
83
|
+
| doc2md converter (markitdown venv) | offered at a terminal; an unattended install prints the command | install later with `doc2md install-converter` |
|
|
84
|
+
| Korean writing guidance | on when the locale is Korean (asked at a terminal) | `korean off`, `CTS_NO_KOREAN=1` |
|
|
85
|
+
| compact-window warning chip | on | `compact-window off` |
|
|
86
|
+
| Update-available chip | on | `CTS_NO_UPDATE_CHECK=1` |
|
|
87
|
+
| **Pinning compact-window** (`autoCompactWindow` 500k) | **off — run it yourself** | `compact-window set --global` or `--project` |
|
|
88
|
+
| **Model-fitting rules** (`ratchet-model.md` delegations) | **candidates are proposed only** | review and approve with `route-scan rules` |
|
|
89
|
+
| `handoff` (back up work before a cap) | an on-demand command | n/a |
|
|
31
90
|
|
|
32
|
-
|
|
91
|
+
`compact-window set` writes into Claude Code's `settings.json` and a human has to choose global or project scope, so it is never run for you. Model-fitting rules keep an approval step for the same reason: which work belongs on a cheaper tier is your call.
|
|
33
92
|
|
|
34
|
-
## 🔀 절감액은 추정이 아니라 원장 기록입니다
|
|
35
93
|
|
|
36
|
-
|
|
94
|
+
## 🔀 The savings figure is a ledger entry, not an estimate
|
|
95
|
+
|
|
96
|
+
Every delegated run is recorded like this:
|
|
37
97
|
|
|
38
98
|
```
|
|
39
|
-
|
|
99
|
+
before after gap
|
|
40
100
|
claude-opus-5 → haiku-4-5 = $0.57
|
|
41
|
-
(
|
|
42
|
-
|
|
43
|
-
|
|
101
|
+
(the model (what (same token counts,
|
|
102
|
+
handling this actually priced against
|
|
103
|
+
before the rule) ran it) both models)
|
|
44
104
|
```
|
|
45
105
|
|
|
46
106
|
```bash
|
|
47
|
-
$ claude-token-saver route-scan savings #
|
|
107
|
+
$ claude-token-saver route-scan savings # trace every dollar back to its rule
|
|
48
108
|
|
|
49
|
-
🔀
|
|
109
|
+
🔀 Routing saved, lifetime $2.09 (last 7d $1.40 · 30d $2.09)
|
|
50
110
|
|
|
51
|
-
|
|
52
|
-
claude-fable-5 → claude-sonnet-5 — 1
|
|
53
|
-
claude-opus-5 → claude-haiku-4-5 — 1
|
|
111
|
+
By model change:
|
|
112
|
+
claude-fable-5 → claude-sonnet-5 — 1 run, $0.72
|
|
113
|
+
claude-opus-5 → claude-haiku-4-5 — 1 run, $0.57
|
|
54
114
|
|
|
55
|
-
|
|
115
|
+
By run (newest first):
|
|
56
116
|
2026-08-22 $0.51 claude-fable-5 → claude-haiku-4-5
|
|
57
|
-
|
|
117
|
+
rule: T2|paste|-Users-me-projects-my-app
|
|
58
118
|
```
|
|
59
119
|
|
|
60
|
-
|
|
120
|
+
**What is excluded** — an honest number beats a big one:
|
|
61
121
|
|
|
62
|
-
-
|
|
63
|
-
-
|
|
122
|
+
- Delegations no registered rule covers (`Explore`, your own agents, plugin subagents): this tool did not route them.
|
|
123
|
+
- Model ids the pricing table cannot recognize: the run is dropped rather than priced wrong.
|
|
64
124
|
|
|
65
125
|
---
|
|
66
126
|
|
|
67
|
-
## ⚡
|
|
127
|
+
## ⚡ What else the statusline catches
|
|
68
128
|
|
|
69
129
|
| | |
|
|
70
130
|
|---|---|
|
|
71
|
-
| 🚨
|
|
72
|
-
| 🧠
|
|
73
|
-
|
|
|
131
|
+
| 🚨 **No surprise rate limits** | Warns when the 5H/7D window hits 90%; `handoff` backs up your work |
|
|
132
|
+
| 🧠 **Cache waste detection** | Hit rate, TTL, 1M-context detection — spikes diagnosed with issue codes |
|
|
133
|
+
| 💵 **Spend visibility** | Month-to-date spend (`💵 Sep $42`) always on; behind a LiteLLM gateway the key budget gauge (`🔑 budget 34% $34/$100`) stands in for the missing 5h/7d caps ([below](#litellm-your-key-budget-stands-in-for-the-missing-5h7d-caps-v3350)) |
|
|
134
|
+
| 🇰🇷 **Korean writing guidance** | Offered at install time, defaulting to your locale ([below](#-korean-writing-guidance)) |
|
|
74
135
|
|
|
75
|
-
##
|
|
136
|
+
## Not a router — 60 seconds
|
|
76
137
|
|
|
77
|
-
|
|
78
|
-
|
|
79
|
-
|
|
138
|
+
It never intercepts a request in realtime.
|
|
139
|
+
**After a session ends** it reads your local logs, finds the easy patterns your expensive model
|
|
140
|
+
kept handling, and promotes them into rules so a cheaper model takes them **from the next session
|
|
141
|
+
onward**. Rules are scoped global or per-project.
|
|
80
142
|
|
|
81
|
-
###
|
|
143
|
+
### Why realtime model routing can cost more, not less
|
|
82
144
|
|
|
83
|
-
|
|
145
|
+
Never switching models mid-session is the point of this design.
|
|
84
146
|
|
|
85
|
-
|
|
147
|
+
Prompt caches are **kept per model.** Switch to a cheaper model mid-session and it starts from a cold cache, re-reading the whole conversation at full input price. A cache hit costs about a tenth of that, so past roughly 20k tokens of history **one switch can erase everything the cheaper model was going to save.** You moved the work down a tier and the bill went up: the central paradox of realtime routing.
|
|
86
148
|
|
|
87
|
-
|
|
149
|
+
Teams shipping routing products have turned the feature off for exactly this reason: [LLM 라우터를 만든 사람들이 직접 껐습니다 #Shorts](https://www.youtube.com/shorts/SK-GoAABjbg) (Korean).
|
|
88
150
|
|
|
89
|
-
|
|
151
|
+
So this tool never touches the main session's model. It delegates to **subagents only**, which leaves the main session's cache intact and runs the delegated work on a cheap model in its own context. That is why the savings are not cancelled out by cache loss.
|
|
90
152
|
|
|
91
153
|
```bash
|
|
92
154
|
npm i -g claude-token-saver@latest
|
|
93
|
-
claude-token-saver route-scan #
|
|
94
|
-
claude-token-saver route-scan rules #
|
|
95
|
-
claude-token-saver route-scan savings #
|
|
155
|
+
claude-token-saver route-scan # find delegation candidates in your own history (0 LLM calls)
|
|
156
|
+
claude-token-saver route-scan rules # list promoted rules · rm <N> to remove
|
|
157
|
+
claude-token-saver route-scan savings # audit every dollar the routing saved
|
|
96
158
|
```
|
|
97
159
|
|
|
98
|
-
|
|
99
|
-
|
|
160
|
+
Thresholds come from **your own last-14-day distribution (p25/p75)**, not someone else's benchmark.
|
|
161
|
+
Measured rule-health — whether a delegated run actually succeeded — landed in [v3.9.0](#v390-2026-08-01).
|
|
100
162
|
|
|
101
163
|
---
|
|
102
164
|
|
|
103
|
-
##
|
|
104
|
-
|
|
105
|
-
**사전 준비:** Node.js ≥ 18 (`node -v`로 확인 · macOS `brew install node` · Windows `winget install OpenJS.NodeJS.LTS` · Linux/WSL은 [nvm](https://github.com/nvm-sh/nvm) 권장)
|
|
106
|
-
|
|
107
|
-
```bash
|
|
108
|
-
npm uninstall -g claude-cache-monitor # (구 패키지 사용자만)
|
|
109
|
-
npm i -g claude-token-saver
|
|
110
|
-
```
|
|
111
|
-
|
|
112
|
-
설치하면 Claude Code 화면 하단에 statusline이 곧바로 나타납니다. `--ignore-scripts` 옵션이나 sudo 사용 등으로 자동 등록이 되지 않았다면 `claude-token-saver install`을 실행해 직접 등록하십시오.
|
|
113
|
-
|
|
114
|
-
설치 한 번으로 **statusline과 Skill, SessionStart 훅, 🅷 Harness(5원칙), 최초 route-scan이** 모두 준비됩니다. Harness와 한국어 문체 지침은 **무엇이 추가되는지 보여 준 뒤 켤지 물어봅니다.** Harness는 `~/.claude/CLAUDE.md`에 표시가 붙은 블록으로 **추가되며**, 기존에 작성해 둔 내용은 백업한 뒤 그대로 보존합니다. 이미 설정되어 있는 경우에는 아무것도 바꾸지 않습니다.
|
|
115
|
-
|
|
116
|
-
터미널이 아닌 환경(npm의 `postinstall`, CI, 파이프 입력)에서는 질문을 건너뛰고 기존 기본값을 적용합니다. 질문 없이 진행하려면 `--yes`나 `--no-input`, 아예 건너뛰려면 `CTS_NO_HARNESS=1 npm i -g claude-token-saver`를 쓰십시오. 이미 적용한 설정을 되돌리려면 `claude-token-saver harness uninit --global`을 실행하십시오.
|
|
165
|
+
## Reading the statusline
|
|
117
166
|
|
|
118
|
-
|
|
119
|
-
|
|
120
|
-
### 설치하면 켜지는 기능과 직접 켜야 하는 기능
|
|
121
|
-
|
|
122
|
-
설치 시점에 비용이 들지 않는 기능은 전부 자동으로 켜집니다. 수동으로 남겨 둔 항목은 Claude Code 자체의 설정을 바꾸거나, 적용 범위를 사람이 정해 주어야 하는 것들뿐입니다.
|
|
123
|
-
|
|
124
|
-
| 기능 | 설치 직후 상태 | 끄는 방법 |
|
|
125
|
-
|---|---|---|
|
|
126
|
-
| statusline (진단 칩·절감 원장) | 켜짐 | `claude-token-saver uninstall` |
|
|
127
|
-
| `/claude-token-saver` Skill | 켜짐 | 위와 같습니다 |
|
|
128
|
-
| SessionStart 훅 (route-scan 재분석) | 켜짐 | 위와 같습니다 |
|
|
129
|
-
| UserPromptSubmit 훅 (brief 주입) | 켜짐 | 위와 같습니다 |
|
|
130
|
-
| 최초 route-scan (최근 14일 로그 분석) | 설치 중 즉시 1회 실행 | 해당 없습니다 |
|
|
131
|
-
| 🅷 Harness 5원칙 (`~/.claude/CLAUDE.md`) | 켜짐 (터미널에서는 내용을 보여 주고 한 번 묻습니다) | `harness uninit --global` · `CTS_NO_HARNESS=1` |
|
|
132
|
-
| doc2md 훅 (Read·Edit·Write·프롬프트) | 켜짐 | `doc2md off` · `CTS_NO_DOC2MD=1` |
|
|
133
|
-
| doc2md 변환기(markitdown venv) | 터미널에서 설치 여부를 묻고, 비대화형 설치에서는 명령만 안내합니다 | `doc2md install-converter` 로 나중에 설치 |
|
|
134
|
-
| 한국어 문체 지침 | 시스템 로케일이 한국어면 켜짐 (터미널에서는 묻습니다) | `korean off` · `CTS_NO_KOREAN=1` |
|
|
135
|
-
| compact-window 경고 칩 | 켜짐 | `compact-window off` |
|
|
136
|
-
| 업데이트 안내 칩 | 켜짐 | `CTS_NO_UPDATE_CHECK=1` |
|
|
137
|
-
| **compact-window 값 고정** (`autoCompactWindow` 500k) | **꺼짐 — 직접 실행해야 합니다** | `compact-window set --global` 또는 `--project` |
|
|
138
|
-
| **모델 피팅 룰 승인** (`ratchet-model.md` 위임 룰) | **후보만 제안합니다** | `route-scan rules` 로 확인하고 승인·삭제 |
|
|
139
|
-
| `handoff` (한도 임박 시 작업 백업) | 필요할 때 직접 실행하는 명령입니다 | 해당 없습니다 |
|
|
140
|
-
|
|
141
|
-
`compact-window set` 은 Claude Code 의 `settings.json` 에 값을 적고, 전역과 프로젝트 중 어디에 적을지는 사람이 정해야 하므로 자동으로 실행하지 않습니다. 모델 피팅 룰도 같은 이유로 승인 단계를 남겨 두었습니다. 어떤 작업을 더 싼 티어에 맡길지는 사용자의 판단이 필요합니다.
|
|
142
|
-
|
|
143
|
-
## statusline 읽는 법
|
|
144
|
-
|
|
145
|
-
절감 원장에 기록이 쌓이면 statusline이 **두 줄로** 출력됩니다. 첫째 줄에는 라우팅 절감액만 표시하고, 둘째 줄에는 진단 칩을 표시합니다.
|
|
167
|
+
Once the savings ledger has entries it renders as **two rows** — routing savings on row 1, diagnostics on row 2.
|
|
146
168
|
|
|
147
169
|
```
|
|
148
170
|
🔀 Routing saved $2.09 | fable→sonnet 1× $0.72 · opus→haiku 1× $0.57
|
|
149
171
|
⚠ Ctx 500k+ · 🅷 5/5 · 🤖 Opus 5 · 🧠 Cache hit 98.8% · ⏳ Cache expires 59:46 · ✦ current ███▓░░ 62% 🔄 21:33 · 📅 weekly ██▒░░░ 38% 🔄 Tue 19:33 · 📦 Ctx 47% of 1M · 💰 Cache saved $1.0K · last 1d
|
|
150
172
|
```
|
|
151
173
|
|
|
152
|
-
|
|
174
|
+
With an empty ledger (no measured delegation yet) row 1 is not drawn and the layout stays single-line. If your build renders only the first row (some macOS Claude Code versions), pass `--single-line`.
|
|
153
175
|
|
|
154
|
-
|
|
|
176
|
+
| Segment | Meaning |
|
|
155
177
|
|---|---|
|
|
156
|
-
| `🔀`
|
|
157
|
-
|
|
|
158
|
-
|
|
|
159
|
-
|
|
|
160
|
-
|
|
|
161
|
-
|
|
|
162
|
-
|
|
|
163
|
-
|
|
|
164
|
-
|
|
|
165
|
-
|
|
|
166
|
-
|
|
|
167
|
-
|
|
|
168
|
-
|
|
169
|
-
|
|
178
|
+
| `🔀` **row 1** | Lifetime routing savings + the model moves behind them. The breakdown sums exactly to the total, model names keep only the family (`opus→haiku`). Full audit: `route-scan savings` |
|
|
179
|
+
| `📄` **row 2** | Lifetime doc2md conversion savings with a per-format breakdown. Whichever of routing/conversion saved more takes row 1 |
|
|
180
|
+
| `🤖` | Active model |
|
|
181
|
+
| `🅷 5/5` | Harness principle score ([Harness mode](#-harness-mode)) |
|
|
182
|
+
| `🧠` | Cache hit rate (green at 85%+) |
|
|
183
|
+
| `⏳` | Cache TTL countdown — send a message before expiry to keep the cache warm. Ticking while idle requires Claude Code v2.1.97+ (see [If the countdown looks frozen](#how-it-works--environment)) |
|
|
184
|
+
| `✦ current` / `📅 weekly` | 5-hour / 7-day rate-limit window usage + reset time |
|
|
185
|
+
| `📦` | Context usage (e.g. `Ctx 68% of 1M`) — colored by fill. Current models default to 1M with no premium, but token volume itself drives per-turn cost and 5H/7D burn |
|
|
186
|
+
| `💵 Sep $42` | **Estimated spend since 00:00 on the 1st of this month** (local time). Summed per session with that session's model pricing; always shown, even on gateways with no 5h/7d caps (v3.35.0) |
|
|
187
|
+
| `🔑 budget` | **LiteLLM key budget gauge.** When stdin carries no rate_limits, shows the key's `spend` against `max_budget` as `🔑 budget ▰▱ 34% $34/$100` (v3.35.0, [below](#-behind-a-gateway-bedrock--vertex)) |
|
|
188
|
+
| `💰` | Cumulative savings from prompt caching — a **different** number from row 1's `🔀` (model routing) |
|
|
189
|
+
| `v3.24.0` | The version you are running. Gray, at the tail, when it is the latest one |
|
|
190
|
+
| `⬆ v3.24.0 → 3.25.0` | A newer release exists. Actionable, so it moves to the front of the line ([Update notifications](#-update-notifications)) |
|
|
191
|
+
|
|
192
|
+
When something is wrong, a **warning chip leads the line**:
|
|
170
193
|
|
|
171
194
|
```
|
|
172
195
|
🚨 5H █████▓ 94% 🔄 12:36 · 🅷 5/5 · 🤖 Opus 4.8 · 🧠 Cache hit 72.1% · ⚠ Cache miss · 📅 weekly ▓░░░░░ 12% 🔄 Sun 14:26 · 📦 Ctx 200k · last 1d
|
|
173
196
|
```
|
|
174
197
|
|
|
175
|
-
|
|
198
|
+
Chips — `🚨 5H/7D NN%` (cap imminent) · `⚠ Ctx 500k+` (a single request actually exceeded 500k) · `⚠ Cache miss` · `⚠ Input spike` · `⚠ Output heavy` · `⚠ Call surge` · `⚠ Rebuild churn` · `⚠ 5m TTL`. When both windows cross 90% at once, the sooner-resetting one is promoted to 🚨 and the other stays visible as a red segment (v2.16.0+).
|
|
176
199
|
|
|
177
|
-
###
|
|
200
|
+
### When a chip appears
|
|
178
201
|
|
|
179
|
-
|
|
202
|
+
Run the `/claude-token-saver` Skill inside Claude — or just say the chip wording ("5H cap is up", "cache miss") and it auto-activates. The Skill surfaces the **root-cause code + step-by-step fix**. When a cap is imminent, run `claude-token-saver handoff` to back up your work state to markdown and continue in a fresh session.
|
|
180
203
|
|
|
181
|
-
##
|
|
204
|
+
## Commands
|
|
182
205
|
|
|
183
|
-
|
|
206
|
+
Run these in your shell (inside Claude Code, the `/claude-token-saver` Skill is the only entry point):
|
|
184
207
|
|
|
185
|
-
|
|
|
208
|
+
| Command | What it does |
|
|
186
209
|
|---|---|
|
|
187
|
-
| `claude-token-saver` |
|
|
188
|
-
| `claude-token-saver last` |
|
|
189
|
-
| `claude-token-saver history` |
|
|
190
|
-
| `claude-token-saver handoff` |
|
|
191
|
-
| `claude-token-saver mode [keywords...]` |
|
|
192
|
-
| `claude-token-saver harness ...` | 🅷 Harness
|
|
193
|
-
| `claude-token-saver route-scan` |
|
|
194
|
-
| `claude-token-saver route-scan savings` |
|
|
195
|
-
| `claude-token-saver compact-window` | 1M
|
|
196
|
-
| `claude-token-saver korean on\|off\|status` |
|
|
197
|
-
| `claude-token-saver korean lint block\|warn\|off` |
|
|
198
|
-
| `claude-token-saver korean lint scope all\|prose` |
|
|
199
|
-
| `claude-token-saver doc2md on\|off` |
|
|
200
|
-
| `claude-token-saver doc2md
|
|
201
|
-
| `claude-token-saver mode ttl=5m\|1h\|auto` |
|
|
202
|
-
| `claude-token-saver --version` |
|
|
203
|
-
| `claude-token-saver update-check` |
|
|
204
|
-
| `claude-token-saver upgrade` |
|
|
205
|
-
| `claude-token-saver install` | Skill
|
|
206
|
-
| `claude-token-saver uninstall [--purge]` |
|
|
207
|
-
|
|
208
|
-
|
|
209
|
-
|
|
210
|
-
|
|
211
|
-
|
|
212
|
-
|
|
213
|
-
|
|
214
|
-
|
|
215
|
-
|
|
216
|
-
|
|
217
|
-
|
|
218
|
-
|
|
219
|
-
|
|
220
|
-
|
|
221
|
-
|
|
222
|
-
|
|
210
|
+
| `claude-token-saver` | Last-1-day diagnostic report (`--days N` / `--hours N`) |
|
|
211
|
+
| `claude-token-saver last` | Most recent warning + remediation |
|
|
212
|
+
| `claude-token-saver history` | Last 7 days of warning transitions |
|
|
213
|
+
| `claude-token-saver handoff` | Back work up to `HANDOFF-*.md` before a cap blocks you |
|
|
214
|
+
| `claude-token-saver mode [keywords...]` | Output config (`icon`/`text`, `en`/`ko`, `1h`–`30d` window, …) |
|
|
215
|
+
| `claude-token-saver harness ...` | 🅷 Harness management (below) |
|
|
216
|
+
| `claude-token-saver route-scan` | Detect recurring easy work on expensive models → propose haiku-delegation ratchet rules (below) |
|
|
217
|
+
| `claude-token-saver route-scan savings` | The routing-savings ledger — per-model-change rollup + per-run log (the evidence behind the figure) |
|
|
218
|
+
| `claude-token-saver compact-window` | Warn when a 1M-context session has no auto-compact cap → pin 400k with `set` (below) |
|
|
219
|
+
| `claude-token-saver korean on\|off\|status` | Inject Korean writing guidance at session start and install the write-time check (below) |
|
|
220
|
+
| `claude-token-saver korean lint block\|warn\|off` | How the write-time check handles findings |
|
|
221
|
+
| `claude-token-saver korean lint scope all\|prose` | Check every text file, or documents only |
|
|
222
|
+
| `claude-token-saver doc2md on\|off` | Convert attached documents to Markdown before the model reads them (below) |
|
|
223
|
+
| `claude-token-saver doc2md <file>` | Convert one file by hand. Diagnostic: it prints the refusal reason instead of swallowing it |
|
|
224
|
+
| `claude-token-saver mode ttl=5m\|1h\|auto` | Pin the cache TTL bucket. The default `auto` trusts the measured split, then falls back to gateway detection |
|
|
225
|
+
| `claude-token-saver --version` | Print the installed version |
|
|
226
|
+
| `claude-token-saver update-check` | Is a newer version out? (`--refresh` to ask now, `--dismiss` to mute this version's offer) |
|
|
227
|
+
| `claude-token-saver upgrade` | Install the latest release with the package manager that installed this copy (`--print` shows the command only) |
|
|
228
|
+
| `claude-token-saver install` | Manually register Skill + statusline |
|
|
229
|
+
| `claude-token-saver uninstall [--purge]` | Remove the hooks, statusline and skill it registered. Recorded savings are kept unless `--purge` is given |
|
|
230
|
+
|
|
231
|
+
The output language is decided once, at install time: a terminal install proposes the system locale and asks whether to use Korean, while an unattended install records what the locale says. Once recorded it is never asked again, not even on an upgrade. Change it later with `mode ko` / `mode en`, or pin it for a scripted install with `CTS_LANG=ko` / `CTS_LANG=en`. Statusline chips stay symbolic either way.
|
|
232
|
+
|
|
233
|
+
<details>
|
|
234
|
+
<summary>All CLI options</summary>
|
|
235
|
+
|
|
236
|
+
| Flag | Description | Default |
|
|
237
|
+
|------|-------------|---------|
|
|
238
|
+
| `--days, -d` | Analysis period in days | 30 |
|
|
239
|
+
| `--hours` | Analysis window in hours (overrides `--days`) | – |
|
|
240
|
+
| `--format, -f` | `table` / `json` / `csv` | table |
|
|
241
|
+
| `--project, -p` | Filter by project directory | all |
|
|
242
|
+
| `--threshold` | Hit-rate alert threshold (0.0–1.0) | 0.7 |
|
|
243
|
+
| `--statusline` | One-line statusline output | – |
|
|
244
|
+
| `--icon` | Use 🧠 / ⏳ / 💰 / 📦 icons | text |
|
|
245
|
+
| `--verbose` | Longer labels | – |
|
|
246
|
+
| `--no-timer` | Hide TTL countdown | show |
|
|
247
|
+
| `--no-color` | Strip ANSI codes | – |
|
|
248
|
+
| `--segments=…` | Limit statusline segments (e.g. `model,five_hour,seven_day,saved`) | all |
|
|
249
|
+
| `--install-hook` / `--uninstall-hook` | Manage the PostToolUse hook | – |
|
|
250
|
+
</details>
|
|
251
|
+
|
|
252
|
+
## ⬆ Update notifications
|
|
253
|
+
|
|
254
|
+
A statusline cannot open a dialog, and it re-renders every ~300ms, so it can never touch the network while drawing. The notification is therefore split in two:
|
|
255
|
+
|
|
256
|
+
- **The statusline tells you.** Up to date: a quiet gray `v3.24.0` at the tail. Newer release out: `⬆ v3.24.0 → 3.25.0` in yellow, moved to the front. Never red — nothing is broken.
|
|
257
|
+
- **Session start asks you.** On a new session or `/clear`, the SessionStart hook injects one line telling the model a newer version exists and to ask before installing anything. Only after you agree does it run `claude-token-saver upgrade`.
|
|
258
|
+
- **Declining sticks.** `claude-token-saver update-check --dismiss` mutes the offer for that version; the next release asks again. The statusline chip stays — you declined the question, not the fact.
|
|
259
|
+
|
|
260
|
+
The registry lookup runs at most once every 24h in a detached background process and only ever writes a cache file (`update-check.json`) — the same shape npm's `update-notifier` uses. A failed check still stamps its timestamp, so an offline machine backs off instead of retrying on every render. Turn checks off entirely with `CTS_NO_UPDATE_CHECK=1` or `NO_UPDATE_NOTIFIER`.
|
|
261
|
+
|
|
262
|
+
## 🅷 Harness mode
|
|
263
|
+
|
|
264
|
+
Bootstrap five engineering principles (Ratchet · Evidence · PEV · Structured Task · Default Safe Path) into `CLAUDE.md` with one command; the statusline scores it as `🅷 5/5`. When the same error keeps recurring, a `🅷⚠ ratchet?` nudge appears so you can promote it to a rule.
|
|
223
265
|
|
|
224
266
|
```bash
|
|
225
|
-
claude-token-saver harness init #
|
|
226
|
-
claude-token-saver harness init --global # ~/.claude/CLAUDE.md
|
|
227
|
-
claude-token-saver harness check #
|
|
228
|
-
claude-token-saver harness analyze #
|
|
229
|
-
claude-token-saver harness promote <N> --project|--global #
|
|
230
|
-
claude-token-saver harness promote "
|
|
231
|
-
claude-token-saver harness pull #
|
|
232
|
-
claude-token-saver harness list / rm <N> #
|
|
233
|
-
claude-token-saver harness off | on # 🅷
|
|
267
|
+
claude-token-saver harness init # this project
|
|
268
|
+
claude-token-saver harness init --global # ~/.claude/CLAUDE.md — every project
|
|
269
|
+
claude-token-saver harness check # current score (global fallback honored)
|
|
270
|
+
claude-token-saver harness analyze # run the transcript analysis manually (no hook needed); refreshes harness-state.json
|
|
271
|
+
claude-token-saver harness promote <N> --project|--global # warning #N → ratchet rule (scope required)
|
|
272
|
+
claude-token-saver harness promote "<rule text>" --project|--global # register your own hand-written rules the same way
|
|
273
|
+
claude-token-saver harness pull # register the package's curated ratchet rules into your global ratchet (opt-in, dedupes)
|
|
274
|
+
claude-token-saver harness list / rm <N> # view / delete rules (auto .bak)
|
|
275
|
+
claude-token-saver harness off | on # toggle the 🅷 chip
|
|
234
276
|
```
|
|
235
277
|
|
|
236
|
-
- `promote
|
|
237
|
-
- `pull
|
|
238
|
-
- `seed
|
|
239
|
-
- 🅷⚠
|
|
278
|
+
- `promote` **requires** `--project`/`--global` in non-TTY contexts (scripts, LLM calls) — a scope choice is never silently made for the caller.
|
|
279
|
+
- `pull` registers the **author-curated ratchet rules** bundled with the package (`presets/ratchet-rules.json` — only general-purpose rules promoted from real recurring mistakes) into your global ratchet (`~/.claude/ratchet.md`). `install`/`init` never auto-inject anything; `pull` is always opt-in and idempotent. Drop any rule you dislike with `harness rm`.
|
|
280
|
+
- `seed` offers the same presets **one at a time**. Where `pull` registers the whole ratchet set in one go, `seed` covers the model-fitting presets too and asks about each of them in the first session after an install or upgrade ([below](#-seed-delegation-that-works-from-the-first-session)).
|
|
281
|
+
- 🅷⚠ runtime warnings (`ratchet?` `no-evidence` `PEV-skip`) expire after 30 minutes, subdirectory sessions match their project correctly, and PEV-skip counts only mutating tools (Edit/Write/Bash) so read-only research sessions don't trip it (v2.16.0+).
|
|
240
282
|
|
|
241
283
|
<details>
|
|
242
|
-
<summary>⚠️ <code>harness rm</code
|
|
284
|
+
<summary>⚠️ <code>harness rm</code> — checklist before deleting</summary>
|
|
243
285
|
|
|
244
|
-
|
|
286
|
+
The whole point of the ratchet is **one-direction accumulation**. Deleting rules casually means the same mistakes return.
|
|
245
287
|
|
|
246
|
-
-
|
|
247
|
-
-
|
|
248
|
-
-
|
|
288
|
+
- **Rule too broad, blocking valid cases?** → ❌ delete ✅ narrow the condition (e.g. `"no hardcoded values"` → `"no hardcoded values outside tests"`)
|
|
289
|
+
- **Rule too narrow, almost never fires?** → ❌ delete ✅ leave it (zero cost)
|
|
290
|
+
- **Genuinely wrong?** → ✅ delete then
|
|
249
291
|
|
|
250
|
-
|
|
292
|
+
An auto `.bak` is kept, but **the session context that earned the rule its place is not recoverable.**
|
|
251
293
|
</details>
|
|
252
294
|
|
|
253
295
|
|
|
254
|
-
## 📦 compact-window
|
|
296
|
+
## 📦 compact-window — pin where a 1M session compacts
|
|
255
297
|
|
|
256
|
-
Claude Code
|
|
298
|
+
Claude Code compacts when usage approaches `min(autoCompactWindow, model max context)`. On a 1M window, with that value unset, compaction only fires near 800k — and until then every request re-bills the whole context. **1M is too large; the recommendation is a 400k–700k band** — 2–3.5x a 200k session's headroom for the genuinely large pastes, with the runaway tail cut off.
|
|
257
299
|
|
|
258
|
-
|
|
300
|
+
**Anything inside the band is left alone.** 400k is the floor where the saving beats the extra compactions, and long sessions often want more room than that. Only an unset window, or one above 700k, is warned about (a smaller one is a deliberate, more aggressive choice).
|
|
259
301
|
|
|
260
|
-
**200k
|
|
302
|
+
**200k sessions are never warned** — their window is already at or below 200k, so the setting cannot change anything.
|
|
261
303
|
|
|
262
304
|
```bash
|
|
263
|
-
claude-token-saver compact-window #
|
|
264
|
-
claude-token-saver compact-window set --global # ~/.claude/settings.json
|
|
265
|
-
claude-token-saver compact-window set --project # <root>/.claude/settings.json
|
|
266
|
-
claude-token-saver compact-window set --global --value 600k
|
|
267
|
-
claude-token-saver compact-window off | on #
|
|
305
|
+
claude-token-saver compact-window # status (model, window, value, source)
|
|
306
|
+
claude-token-saver compact-window set --global # pin 500k (mid-band) in ~/.claude/settings.json
|
|
307
|
+
claude-token-saver compact-window set --project # pin it in <root>/.claude/settings.json
|
|
308
|
+
claude-token-saver compact-window set --global --value 600k # explicit value (100k–1M)
|
|
309
|
+
claude-token-saver compact-window off | on # toggle the warning
|
|
268
310
|
```
|
|
269
311
|
|
|
270
|
-
- 1M
|
|
271
|
-
-
|
|
272
|
-
-
|
|
273
|
-
-
|
|
312
|
+
- On a 1M model with the value unset or above 700k, the statusline shows `🅷⚠ compact-window?` and the session briefing hands the model the exact registration command.
|
|
313
|
+
- Scope (`--global`/`--project`) is **required** for `set` — a global settings file is never edited on a guess.
|
|
314
|
+
- Every other key in `settings.json` is preserved and a `.bak` is written first. Malformed JSON aborts the write untouched.
|
|
315
|
+
- An exported `CLAUDE_CODE_AUTO_COMPACT_WINDOW` beats settings.json; `set` detects that and says so.
|
|
274
316
|
|
|
275
|
-
## 🔀 route-scan
|
|
317
|
+
## 🔀 route-scan — "this recurring task could run on a cheaper tier"
|
|
276
318
|
|
|
277
|
-
|
|
319
|
+
Finds the easy work your expensive model (opus/fable) keeps redoing in your session logs and proposes **haiku/sonnet delegation rules**. Fully local, zero token cost.
|
|
278
320
|
|
|
279
|
-
- **T2 → haiku
|
|
280
|
-
- **T1 → sonnet
|
|
281
|
-
- **T0
|
|
321
|
+
- **T2 → haiku**: lookups, pasted-screen Q&A, simple runs — zero errors, near-zero mutation
|
|
322
|
+
- **T1 → sonnet**: build pipelines, status checks — few mutations, ≤1 error
|
|
323
|
+
- **T0 stays**: repeated errors, heavy mutation, design/analysis — the session model keeps it
|
|
282
324
|
|
|
283
|
-
|
|
284
|
-
1.
|
|
285
|
-
2.
|
|
286
|
-
3.
|
|
325
|
+
Three design pillars:
|
|
326
|
+
1. Difficulty is judged by **outcome, not text guessing** — tool errors, mutating tool calls, output tokens
|
|
327
|
+
2. Thresholds **auto-calibrate to your own 14-day distribution** — fixed constants drift with workload
|
|
328
|
+
3. Promoted rules live in a tool-owned file (`.claude/ratchet-model.md`) that **refreshes itself every scan**, and a `⚠ rule-health` flag fires when a delegated category's error rate climbs — rules report their own staleness
|
|
287
329
|
|
|
288
330
|
```bash
|
|
289
|
-
claude-token-saver route-scan #
|
|
290
|
-
claude-token-saver harness promote R1 --project #
|
|
291
|
-
claude-token-saver route-scan dismiss 1 #
|
|
292
|
-
claude-token-saver route-scan rules #
|
|
293
|
-
claude-token-saver route-scan savings #
|
|
294
|
-
```
|
|
295
|
-
|
|
296
|
-
`route-scan savings`는 statusline의 `🔀 Routing saved` 한 줄 뒤에 있는 근거를 그대로 보여줍니다. 모델 이동별 합계와 실행별 내역이 함께 나오므로, 금액이 어디서 나왔는지 추적할 수 있습니다.
|
|
297
|
-
|
|
298
|
-
```
|
|
299
|
-
🔀 라우팅 절감 누적 $2.09 (최근 7일 $1.40 · 30일 $2.09)
|
|
300
|
-
|
|
301
|
-
모델 이동별:
|
|
302
|
-
claude-fable-5 → claude-sonnet-5 — 1회, $0.72
|
|
303
|
-
claude-opus-5 → claude-haiku-4-5 — 1회, $0.57
|
|
331
|
+
claude-token-saver route-scan # scan (24h cache) + tiered candidates
|
|
332
|
+
claude-token-saver harness promote R1 --project # promote candidate R1 to a model-fitting rule
|
|
333
|
+
claude-token-saver route-scan dismiss 1 # not interested — won't resurface
|
|
334
|
+
claude-token-saver route-scan rules # list model-fitting rules (rm <N> to remove)
|
|
335
|
+
claude-token-saver route-scan savings # the savings ledger — which rule moved work off which model, onto which
|
|
304
336
|
```
|
|
305
337
|
|
|
306
|
-
|
|
338
|
+
Dig deeper: **tier criteria & research evidence** → [docs/TIER_CRITERIA.md](./docs/TIER_CRITERIA.md) (Korean) · **rule-file mechanics, scan triggers, subagent setup** → [docs/ROUTE_SCAN.md](./docs/ROUTE_SCAN.md) (Korean + English)
|
|
307
339
|
|
|
308
|
-
###
|
|
340
|
+
### Behind a gateway (Bedrock / LiteLLM)
|
|
309
341
|
|
|
310
|
-
|
|
342
|
+
Through a corporate gateway the transcript records an inference-profile ARN where the model id belongs. That string says nothing about `opus` or `haiku`, so older versions read every session as Sonnet — which made **T1 (→sonnet) rules unreachable and zeroed the savings figures**.
|
|
311
343
|
|
|
312
|
-
v3.10.0
|
|
344
|
+
Since v3.10.0 the profile id is mapped back to a role (main, opus, sonnet, haiku) and then to the alias your `ANTHROPIC_DEFAULT_*_MODEL` variables declare. The mapping is learned by joining each parent `Task` call to the subagent run it spawned via `toolUseId`. Below three observations, or when the role votes agree less than 80% of the time, the id stays `unknown` and drops out of the delegation aggregate rather than being guessed at.
|
|
313
345
|
|
|
314
|
-
|
|
346
|
+
For environments the learner cannot reach, write the mapping yourself in `<userDataDir>/profile-map.json`. Account id and region may be wildcarded:
|
|
315
347
|
|
|
316
348
|
```jsonc
|
|
317
349
|
{
|
|
318
350
|
"modelAliases": {
|
|
319
351
|
"arn:aws:bedrock:*:*:application-inference-profile/<PROFILE_ID>": "claude-opus-5",
|
|
320
|
-
"prod-large": "claude-opus-5", //
|
|
352
|
+
"prod-large": "claude-opus-5", // house aliases map the same way
|
|
321
353
|
"team-*": "claude-haiku-4-5"
|
|
322
354
|
}
|
|
323
355
|
}
|
|
324
356
|
```
|
|
325
357
|
|
|
326
|
-
|
|
358
|
+
**Map house aliases that carry no family name** (`prod-large`, `team-fast`) here too. Shapes that keep the family name are recognized as-is — Bedrock (`anthropic.claude-opus-4-5-v1:0`), Vertex (`claude-opus-4-5@20251101`), and the 1M suffix (`claude-sonnet-4-5[1m]`) — but an alias without one cannot be priced. Rather than report a wrong figure, routing-savings **drops those runs from the aggregate** (both sides of the comparison must be recognizable); one line in the table above brings them back.
|
|
327
359
|
|
|
328
|
-
|
|
360
|
+
That file holds internal identifiers in plain text — do not commit it. On a direct-API machine it is never created and behaviour is unchanged.
|
|
329
361
|
|
|
330
|
-
## 🌱 seed:
|
|
362
|
+
## 🌱 seed: delegation that works from the first session
|
|
331
363
|
|
|
332
|
-
|
|
364
|
+
The model-fitting ratchet (`ratchet-model.md`) **starts empty.** A rule exists only after route-scan has seen the same kind of work recur in your own logs and you have approved that candidate. So a fresh install delegates nothing, and keeps delegating nothing for days — precisely the stretch where the savings would matter most.
|
|
333
365
|
|
|
334
|
-
`seed
|
|
366
|
+
`seed` fills that gap from presets bundled with the package.
|
|
335
367
|
|
|
336
|
-
|
|
|
368
|
+
| Presets | What they cover | File |
|
|
337
369
|
|---|---|---|
|
|
338
|
-
|
|
|
339
|
-
|
|
|
370
|
+
| 9 model-fitting | running commands, lookup, status checks, questions about pasted logs, read-and-summarize — each with a T2 (haiku) and a T1 (sonnet) rule | `presets/model-rules.json` |
|
|
371
|
+
| 6 ratchet | general-purpose rules promoted from mistakes that actually recurred | `presets/ratchet-rules.json` |
|
|
340
372
|
|
|
341
|
-
|
|
373
|
+
**How they get registered:** in the first session after an install or upgrade, the SessionStart hook hands the pending presets to the model, which walks the user through them **one at a time**. Each answer runs one of these immediately:
|
|
342
374
|
|
|
343
375
|
```bash
|
|
344
|
-
claude-token-saver seed #
|
|
345
|
-
claude-token-saver seed accept <id> --global|--project #
|
|
346
|
-
claude-token-saver seed accept all --global #
|
|
347
|
-
claude-token-saver seed skip <id> #
|
|
348
|
-
claude-token-saver seed reset #
|
|
376
|
+
claude-token-saver seed # pending presets + recorded answers
|
|
377
|
+
claude-token-saver seed accept <id> --global|--project # register one (scope required)
|
|
378
|
+
claude-token-saver seed accept all --global # when the user says "register them all"
|
|
379
|
+
claude-token-saver seed skip <id> # decline — never offered again
|
|
380
|
+
claude-token-saver seed reset # clear the answers and offer everything again
|
|
349
381
|
```
|
|
350
382
|
|
|
351
|
-
-
|
|
352
|
-
-
|
|
353
|
-
-
|
|
354
|
-
-
|
|
383
|
+
- **Nothing is written without a yes to that specific rule.** A declined rule stays declined across upgrades; a later release only surfaces the presets it actually added.
|
|
384
|
+
- A preset is withheld when you already approved a rule of the same shape (same tier and category).
|
|
385
|
+
- A seeded rule **does not pass someone else's statistics off as yours.** It is recorded as `preset (curated)` until a scan measures real firings and delegations, and then those numbers replace it. If its delegated error rate crosses the threshold it gets the same review flag as any other rule.
|
|
386
|
+
- The scope must be stated as `--global` or `--project`. The hook environment is non-TTY, so the CLI cannot ask — the model confirms with the user and passes the flag.
|
|
355
387
|
|
|
356
|
-
## 🇰🇷
|
|
388
|
+
## 🇰🇷 Korean writing guidance
|
|
357
389
|
|
|
358
|
-
|
|
390
|
+
Injects guidance that corrects how Claude writes Korean (dropped sentence parts, noun-stopped sentences, translationese, em-dash overuse) **once per session.**
|
|
359
391
|
|
|
360
392
|
```bash
|
|
361
|
-
claude-token-saver korean on #
|
|
362
|
-
claude-token-saver korean status #
|
|
363
|
-
claude-token-saver korean show #
|
|
364
|
-
claude-token-saver korean off #
|
|
393
|
+
claude-token-saver korean on # on, for every project
|
|
394
|
+
claude-token-saver korean status # state, cost, provenance
|
|
395
|
+
claude-token-saver korean show # print the guidance itself
|
|
396
|
+
claude-token-saver korean off # off
|
|
365
397
|
```
|
|
366
398
|
|
|
367
|
-
Claude Code
|
|
399
|
+
Claude Code's output styles can do the same thing, but an output style is **a single global slot**: turning it on takes that slot away from anything else and has to be configured per machine. This ships the guidance inside the package and delivers it through the SessionStart hook that is already installed, so it **applies wherever the CLI is installed and leaves the output-style slot free.** It survives `/clear`, because the hook fires again.
|
|
368
400
|
|
|
369
|
-
|
|
401
|
+
Cost is **~1.5k tokens per session, injected once at session start rather than per turn**, and covered by the prompt cache from the second request on. When it is on, a `✍️` chip appears in the statusline.
|
|
370
402
|
|
|
371
|
-
###
|
|
403
|
+
### Before and after
|
|
372
404
|
|
|
373
|
-
|
|
405
|
+
From a real pipeline (yaml-sns-agent), where a Slack bot shells out to `claude -p` to investigate the repository and answers in a thread. The guidance was toggled off and then on, with the same question both times: "explain in three Korean sentences how to narrow down a render failure."
|
|
374
406
|
|
|
375
|
-
|
|
407
|
+
**Before** (guidance off)
|
|
376
408
|
|
|
377
409
|
> 렌더 실패 원인 좁히기 3단계: 먼저 `hn:still --check`로 스틸이 명세보다 최신인지(FRESH/STALE) 확인 — STALE이면 재생성 안 해서 옛 자산으로 렌더 시도한 게 원인. 다음 caption-blocks가 최신 나레이션과 단어열 일치하는지 직접 대조 — TTS 재생성 후 blocks 재생성 빠지면 자막·음성 어긋나 렌더 깨짐.
|
|
378
410
|
|
|
379
|
-
|
|
411
|
+
**After** (guidance on)
|
|
380
412
|
|
|
381
413
|
> 먼저 `hn:still --check` 돌려 스틸이 명세보다 최신인지 확인한다. STALE이면 재생성 안 해서 생긴 문제.
|
|
382
414
|
>
|
|
383
415
|
> 다음 caption-blocks가 captions.json 단어열과 일치하는지 본다. 내레이션 재TTS 후 blocks 재생성 빠지면 옛 자막이 새 음성 위에 뜬다.
|
|
384
416
|
|
|
385
|
-
|
|
417
|
+
Three things change. Clauses chained with em dashes become separate sentences, so one sentence carries one fact. Noun-stopped phrases (확인, 대조, 렌더 깨짐 — "check", "compare", "render breaks") become predicates (확인한다, 본다, 뜬다), which makes it explicit that these are steps to take. And the particles come back where they had been dropped, so subject and object are legible on the first read.
|
|
386
418
|
|
|
387
|
-
|
|
419
|
+
The technical content is identical in both. The guidance touches sentence construction only, not judgement or accuracy: the answer does not change, it just stops needing a second read. In a channel people scroll through, that difference cuts follow-up questions — and the tokens those follow-ups would have cost.
|
|
388
420
|
|
|
389
|
-
###
|
|
421
|
+
### The write-time check (v3.24.0)
|
|
390
422
|
|
|
391
|
-
|
|
423
|
+
Injecting the guidance once at session start turned out to be half the job. The model reads it, then writes dozens of files over the next hours with nothing re-reading the output. Sessions with the guidance active still shipped violations into documents, and it surfaced only when a human read the finished artifact. An August 2026 fix reworded the scope sentence to address this; it recurred, because rewording an instruction does not add a checkpoint.
|
|
392
424
|
|
|
393
|
-
v3.24.0
|
|
425
|
+
From v3.24.0 `korean on` also installs a PostToolUse hook. It opens the file the model just wrote, runs the clauses a machine can decide, and hands any findings back. The file is already saved, so nothing is lost — the model fixes it on the spot.
|
|
394
426
|
|
|
395
427
|
```bash
|
|
396
|
-
claude-token-saver korean lint block #
|
|
397
|
-
claude-token-saver korean lint warn #
|
|
398
|
-
claude-token-saver korean lint off #
|
|
428
|
+
claude-token-saver korean lint block # default: findings are handed back as blocking feedback
|
|
429
|
+
claude-token-saver korean lint warn # print findings, do not block
|
|
430
|
+
claude-token-saver korean lint off # disable the check
|
|
399
431
|
|
|
400
|
-
claude-token-saver korean lint scope all #
|
|
401
|
-
claude-token-saver korean lint scope prose #
|
|
432
|
+
claude-token-saver korean lint scope all # default: every text file the session writes
|
|
433
|
+
claude-token-saver korean lint scope prose # documents only
|
|
402
434
|
|
|
403
|
-
claude-token-saver korean lint docs/*.md #
|
|
435
|
+
claude-token-saver korean lint docs/*.md # check files already on disk
|
|
404
436
|
```
|
|
405
437
|
|
|
406
|
-
|
|
438
|
+
Checked: 15 figurative phrases, translationese markers, separators (`—`·`ㅡ`·`|`), three or more `의` particles in one phrase, and a period after a nominal ending. Clauses that need judgement stay with the guidance text.
|
|
407
439
|
|
|
408
|
-
|
|
440
|
+
The default `all` scope covers code comments, UI strings, subtitles, templates, and build output, not just documents. The vendored guidance exempts comments, but comments are read by people and generated artifacts (PDF, HTML) are assembled from those strings, so exempting them reopens the exact gap that was reported. Only installed dependencies, VCS internals, lockfiles, and binary or image files are skipped; `dist/` and `build/` are checked. `korean lint scope prose` restores the narrow reading.
|
|
409
441
|
|
|
410
|
-
|
|
442
|
+
The scope sentence in the injected guidance is generated from the same setting, so the model is never told one rule while being corrected against another.
|
|
411
443
|
|
|
412
|
-
###
|
|
444
|
+
### The encoding rule that ships with it (v3.23.2)
|
|
413
445
|
|
|
414
|
-
|
|
446
|
+
Alongside the writing guidance, one more line is injected: **non-ASCII strings in tool-call parameters must be written as literal UTF-8, never as `\uXXXX` unicode escapes.**
|
|
415
447
|
|
|
416
|
-
|
|
448
|
+
When the model puts Korean into a Write or Edit parameter as escapes, those escapes are sometimes not decoded into code points at all: the literal text `한` lands in the file. The artifact carries mojibake, and the model keeps editing on top of it without noticing that what it wrote and what the file holds have diverged. Not writing escapes in the first place removes the path entirely, so the rule blocks the input instead of repairing the output.
|
|
417
449
|
|
|
418
|
-
|
|
450
|
+
This line lives in claude-token-saver's own framing paragraph, not in the vendored fluent-korean text. It governs encoding rather than style, and the vendored wording is kept unmodified. For the same reason it carries no exceptions, unlike the style rules that skip code and commit messages. It adds roughly 60 tokens per session.
|
|
419
451
|
|
|
420
|
-
>
|
|
421
|
-
>
|
|
452
|
+
> **Evidence**
|
|
453
|
+
> The same failure is reported against Claude Code: [#12417, unicode handling regression](https://github.com/anthropics/claude-code/issues/12417) and [#26141, Edit silently corrupting unicode](https://github.com/anthropics/claude-code/issues/26141).
|
|
422
454
|
|
|
423
|
-
###
|
|
455
|
+
### Asked at install time
|
|
424
456
|
|
|
425
|
-
|
|
457
|
+
The install **prints what the guidance changes, its per-session cost and its source, then asks.** A Korean system locale (`ko_KR` and friends; on macOS the system setting is checked too) makes the question default to yes; anything else defaults to no, so users who never write Korean are not billed 1.5k tokens a session. The locale is only a default, so an English-locale machine used for Korean work can still turn it on right there.
|
|
426
458
|
|
|
427
|
-
npm
|
|
459
|
+
Installs with nobody attached — npm `postinstall`, CI, piped stdin — skip the question and apply the locale default, because a blocked prompt hangs the install. In that case, if the locale is not Korean the setting is **left undecided rather than recorded**, so a later run at a terminal still gets to ask. Use `--yes` or `--no-input` to force the non-interactive path, or `CTS_NO_KOREAN=1` to skip the feature entirely. **Once you have turned it on or off yourself, that choice sticks — an upgrade never overrides it.**
|
|
428
460
|
|
|
429
|
-
>
|
|
430
|
-
>
|
|
431
|
-
>
|
|
461
|
+
> **Source and license**
|
|
462
|
+
> The guidance text comes from [fluent-korean](https://github.com/snflkd/fluent-korean). Copyright (c) 2026 snflkd, MIT License.
|
|
463
|
+
> The wording is unmodified; only the output-style frontmatter was removed. The full license ships with the package at `presets/korean-style/LICENSE-fluent-korean`.
|
|
432
464
|
|
|
433
|
-
## 📄 doc2md
|
|
465
|
+
## 📄 doc2md — documents become Markdown before the model reads them
|
|
434
466
|
|
|
435
|
-
|
|
467
|
+
`Read` a pptx, xlsx, pdf or docx and the raw bytes go into the context window, where the model cannot read them. This intercepts that `Read`, converts the file once, and hands over the Markdown instead.
|
|
436
468
|
|
|
437
|
-
|
|
469
|
+
**This is opt-in.** Installing the CLI does not turn it on: both commands below are required, and a registered hook with no converter behind it does nothing at all.
|
|
438
470
|
|
|
439
|
-
|
|
471
|
+
Three situations, three different interception points:
|
|
440
472
|
|
|
441
|
-
|
|
442
|
-
|
|
443
|
-
| 상황 | 개입 지점 |
|
|
473
|
+
| Situation | Where it is caught |
|
|
444
474
|
|---|---|
|
|
445
|
-
|
|
|
446
|
-
|
|
|
447
|
-
|
|
|
475
|
+
| A document path typed in the prompt (`@path`, quoted, or relative) | `UserPromptSubmit`: converted, and the conversion's path is handed back as context |
|
|
476
|
+
| A document opened with `Read` mid-task | PDFs are caught by `PreToolUse(Read)`. pptx/xlsx/docx/fig are not: Claude Code refuses them as binary *before* any hook runs, so the session-start note tells the model to run `doc2md <path>` instead |
|
|
477
|
+
| A document attached to the message | **Not catchable.** No hook event receives attachment content. The session-start note has the model ask for a path next time |
|
|
448
478
|
|
|
449
|
-
|
|
479
|
+
That second row is measured, not assumed: a `.pdf` Read fires the hook, and a `.pptx` Read in the same session leaves no hook log entry at all.
|
|
450
480
|
|
|
451
481
|
```bash
|
|
452
|
-
claude-token-saver doc2md on #
|
|
453
|
-
claude-token-saver doc2md #
|
|
454
|
-
claude-token-saver doc2md
|
|
455
|
-
claude-token-saver doc2md install-converter #
|
|
482
|
+
claude-token-saver doc2md on # register the hooks (the converter installs itself)
|
|
483
|
+
claude-token-saver doc2md # check converter + hook registration
|
|
484
|
+
claude-token-saver doc2md report.pptx # convert by hand and see the result
|
|
485
|
+
claude-token-saver doc2md install-converter # only to get the install out of the way early
|
|
456
486
|
```
|
|
457
487
|
|
|
458
|
-
|
|
488
|
+
**The converter installs itself.** Any rollout step a person has to be told about is a step some of them skip, so the converter installs in the background the moment a document first shows up, and converts as soon as it is ready. Measured: about 30s for the first document (15s install plus markitdown's first import), then 3.7s for a new document and 0.1s on a cache hit. The `.fig` parser installs in half a second on the first Figma file.
|
|
489
|
+
|
|
490
|
+
It installs on first use rather than at `install` time: the venv is 47MB, and someone who never opens a document should not pay for it. Set `CTS_DOC2MD_NO_AUTOINSTALL=1` to turn the automatic install off.
|
|
491
|
+
|
|
492
|
+
**Python 3.10+ is required** — markitdown's own floor, and macOS still ships 3.9 as `/usr/bin/python3`. The venv is built on an interpreter chosen by version rather than by PATH order. Built on 3.9, pip resolves markitdown to a 2019 placeholder release (0.0.1a1): the install looks like it worked and every conversion then dies at import. This was found by walking into it. When nothing on the machine is new enough, the message points at `brew install python` instead of at an install command that cannot succeed.
|
|
493
|
+
|
|
494
|
+
The converter goes into a venv this tool owns (`<state dir>/doc2md-venv`): no system interpreter is touched, and uninstalling the CLI takes it along. An existing markitdown on `uv tool` or `PATH` is preferred over building a new one.
|
|
495
|
+
|
|
496
|
+
Conversion is [markitdown](https://github.com/microsoft/markitdown). Slide numbers, heading levels, tables, speaker notes and per-sheet headings all survive, and non-Latin text comes through intact.
|
|
497
|
+
|
|
498
|
+
Several things it deliberately does not do:
|
|
499
|
+
|
|
500
|
+
- **Images are not converted.** markitdown returns nothing for them, and OCR misread resource names in testing (`c5.xlarge` as `c.xlarge`). In a document where those names *are* the content, wrong text is worse than none. The model reads images natively anyway.
|
|
501
|
+
- **A missing converter never fails silently.** The install command is shown once, then the original `Read` proceeds untouched. Repeating the notice on every read would be its own nuisance; saying nothing is how a broken converter hides. Run `doc2md` with no arguments to see the converter and hook registration together.
|
|
502
|
+
- **Conversions never land in your project.** They go under the tool's own state directory with mode `0700`, so there is nothing to add to `.gitignore`. Filenames matching payroll/contract/secret patterns are skipped entirely.
|
|
503
|
+
- **Zip bombs are refused.** pptx/xlsx/docx are zip containers: the declared sizes are checked first, and since those are written by whoever built the file, the real decompressed bytes are counted against a ceiling too.
|
|
504
|
+
- **Spreadsheets are capped by rows, not bytes.** Conversion time tracks row count (measured: a 6.3MB PDF in 0.9s, a 5.8MB workbook in 47.75s). Past 50,000 rows only the head is converted, and **the truncation and the true row count are both stated** in what the model is told.
|
|
505
|
+
|
|
506
|
+
### What a conversion saves
|
|
507
|
+
|
|
508
|
+
Every conversion is stamped with a provenance header: which original, when, how many tokens. Savings show up on the statusline's own `📄 Doc2md saved` line.
|
|
509
|
+
|
|
510
|
+
The baseline is what you would have done without a converter, and that differs by format. Both were measured on 2026-09-06.
|
|
511
|
+
|
|
512
|
+
**PDF is priced against attaching it.** The same one-line prompt was sent through `claude --print --input-format stream-json` with and without the file as a document block. The control turn cost 42,204 tokens, twice, to the token.
|
|
513
|
+
|
|
514
|
+
| Attached file | Size | Extra tokens | Per page |
|
|
515
|
+
|---|---|---|---|
|
|
516
|
+
| Résumé PDF | 7 pages | +20,537 | 2,934 |
|
|
517
|
+
| Résumé PDF | 5 pages | +12,709 | 2,542 |
|
|
459
518
|
|
|
460
|
-
|
|
519
|
+
An attached PDF is read whole, but every page costs 2,500–2,900 tokens against 5,531 for the conversion. The coefficient used is 2,500 per page — below both measurements, so the figure understates rather than flatters.
|
|
461
520
|
|
|
462
|
-
|
|
521
|
+
**pptx/xlsx/docx are priced against unpacking the container.** These never reach the model as attachments at all: the same probe on a docx added 78 tokens and the model replied that it had no file, and `Read` refuses the format outright. What you actually do without a converter is unzip the archive and read its XML, where tags and style attributes outweigh the words.
|
|
463
522
|
|
|
464
|
-
|
|
523
|
+
| Original | Body XML | Conversion | Ratio |
|
|
524
|
+
|---|---|---|---|
|
|
525
|
+
| Deck, pptx (31.8MB) | ~540,429 tokens | ~22,610 tokens | 23.8× |
|
|
526
|
+
| Résumé, docx (189KB) | ~79,621 tokens | ~1,684 tokens | 47.3× |
|
|
465
527
|
|
|
466
|
-
|
|
528
|
+
This baseline is measured per file from the real XML size, not applied as a per-format ratio. `.xls` is not a zip container and has no markup to measure, so it claims nothing.
|
|
467
529
|
|
|
468
|
-
|
|
530
|
+
### Figma `.fig` converts too
|
|
469
531
|
|
|
470
|
-
|
|
471
|
-
- **변환기가 없으면 조용히 실패하지 않습니다.** 설치 명령을 한 번 안내한 뒤 원본 `Read` 를 그대로 통과시킵니다. 매번 알리면 그것대로 방해가 되고, 아무 말도 하지 않으면 고장을 숨기게 됩니다. `doc2md` 를 인자 없이 실행하면 변환기와 훅 등록 상태를 한 번에 확인할 수 있습니다.
|
|
472
|
-
- **변환본은 프로젝트 안에 남기지 않습니다.** 도구의 상태 디렉터리 아래 권한 `0700` 으로 저장하므로 `.gitignore` 에 무엇을 추가할 필요가 없습니다. 파일명이 급여·계약·개인정보 같은 패턴에 걸리면 아예 변환하지 않습니다.
|
|
473
|
-
- **압축 폭탄은 막습니다.** pptx·xlsx·docx 는 zip 컨테이너입니다. 선언된 크기를 먼저 걸러 내고, 선언은 조작될 수 있으므로 실제 해제 바이트도 상한과 대조합니다.
|
|
474
|
-
- **엑셀은 행 수로 자릅니다.** 변환 시간은 파일 크기가 아니라 행 수를 따릅니다(실측: PDF 6.3MB 0.9초, 엑셀 5.8MB 47.75초). 5만 행을 넘으면 앞부분만 변환하고, **잘랐다는 사실과 전체 행 수를 안내에 함께 적습니다.**
|
|
532
|
+
Planning documents are moving from PowerPoint to Figma, so the same hook catches `.fig`. A `.fig` is a zip, but the `canvas.fig` inside it is Figma's private binary (kiwi format), which markitdown cannot open — so this one format is converted in Node with [openfig-core](https://github.com/OpenFig-org/openfig-core) (MIT). `doc2md install-converter` places it beside markitdown in the tool's state directory; the package itself still ships zero dependencies.
|
|
475
533
|
|
|
476
|
-
|
|
534
|
+
The result is an outline: pages and frames become headings, text nodes become body lines, and shapes are counted rather than listed — in a planning document the words are the content, and two hundred `Rectangle 173` lines would drown them. A file with no text at all is refused rather than dressed up as an empty document.
|
|
477
535
|
|
|
478
|
-
|
|
536
|
+
Verified against real files: a community Bootstrap UI kit (8.1MB, 4,155 nodes, 1,312 of them text) and a 52MB Tailwind kit, each converting in under a second. Both `.fig` vintages parse — the current zip container and the older bare fig-kiwi stream.
|
|
479
537
|
|
|
480
|
-
|
|
481
|
-
- `.fig` 절감이 가장 큽니다. `Read` 가 이진 파일을 거부하지 않고 그대로 읽어 회당 약 44,000 토큰을 태우기 때문입니다.
|
|
482
|
-
- 암호 문서와 DRM 문서는 오류가 아니라 상태로 판별해 안내합니다. 원본 `Read` 를 막지 않으므로 작업이 중단되지 않습니다.
|
|
538
|
+
**`.fig` saves the most of any format.** Unlike the Office containers, `Read` does not refuse a `.fig`: the extension means nothing to it, so it pulls the binary in as text and the context window fills with tokenised noise. Measured against the same 42,760-token control:
|
|
483
539
|
|
|
484
|
-
|
|
540
|
+
| File | Size | Extra tokens for a Read | Conversion |
|
|
541
|
+
|---|---|---|---|
|
|
542
|
+
| plan.fig | 26KB | +44,195 | 100 tokens |
|
|
543
|
+
| bootstrap-kit.fig | 8.1MB | +43,994 | 18,397 tokens |
|
|
485
544
|
|
|
486
|
-
|
|
545
|
+
Two files three hundred times apart in size cost the same, because Read truncates long before the file ends — you pay for a whole document and receive a fraction of one. The baseline is therefore a flat 44,000 tokens. For comparison, the same probe on a pptx cost +317 tokens and on a docx +185: a refusal message, and nothing else.
|
|
487
546
|
|
|
488
|
-
|
|
547
|
+
#### Why the baseline does not scale with file size
|
|
489
548
|
|
|
490
|
-
|
|
491
|
-
- 5분 버킷에서는 카운트다운 색이 비율이 아니라 절대 시간을 따릅니다. 5분의 30%는 90초여서, 초록이 주는 여유가 실제와 어긋났습니다.
|
|
492
|
-
- `⚠ 5m TTL` 경고가 이 환경에도 도달합니다. 다만 조언 문구는 다릅니다. 구독 플랜을 바꿔도 해소되지 않는 환경이므로 플랜 전환을 권하지 않습니다.
|
|
493
|
-
- `Extra cost if 5m-only` 는 1시간 쓰기가 있는 경우에만 묻습니다. 이미 5분 전용인 환경에서는 질문 자체가 성립하지 않아 `+$0` 이 잘못 읽혔습니다.
|
|
494
|
-
- 위임 건이 모델 ID 해석 실패로 버려졌으면 statusline 에 `🔀 N unresolved` 로 알립니다. 이전에는 "위임한 적 없음"과 화면상 구별되지 않았습니다.
|
|
495
|
-
- 환경변수를 `foundation-model` ARN 으로 지정한 경우에도 모델을 해석합니다. 이름을 담고 있지 않은 `application-inference-profile` ID 는 그대로 거부합니다. 값을 추측해 넣으면 원장에 틀린 금액이 들어가기 때문입니다.
|
|
549
|
+
A baseline has to be what would actually have been spent without the converter. Intuition says a bigger file burns more, but the `Read` tool has a cap (2,000 lines by default, plus a per-line character limit), and a binary file hits it almost immediately: even the 26KB file was already truncated, which is why two files 300× apart came out 201 tokens apart. Had the 8.1MB file gone in whole it would have been millions of tokens — money nobody could have spent, since it does not fit in a 200k context window. Claiming to have saved unspendable money is flattery, not measurement.
|
|
496
550
|
|
|
497
|
-
|
|
551
|
+
The same principle runs through every baseline here:
|
|
498
552
|
|
|
499
|
-
|
|
553
|
+
- **`.fig`, flat 44,000** — set below both measurements (44,195 and 43,994). A model could burn size-proportional tokens by re-Reading at successive offsets, but one Read is what a sane agent does once the bytes turn out to be binary noise, so one Read is the honest counterfactual.
|
|
554
|
+
- **PDF, 2,500 per page** — below both measured values (2,542 and 2,934).
|
|
555
|
+
- **Office formats, the file's actual XML size** — the one case where proportional is right, because a person really does end up reading that XML; it is measured per file rather than applied as a ratio.
|
|
500
556
|
|
|
501
|
-
|
|
557
|
+
The common rule: wherever an estimate and a measurement diverge, the lower number wins. A figure the user can trust is worth more than one that flatters the tool.
|
|
502
558
|
|
|
503
|
-
|
|
504
|
-
- 조회는 LiteLLM 의 `GET /key/info` 와 `GET /user/info` 로 하고, 호출 키 자신의 정보만 받습니다. 예산 출처는 실무에서 가장 많이 쓰는 **팀 멤버십 예산**(team_memberships 의 spend·max_budget)을 먼저 보고, 없으면 키 자체의 max_budget, 그다음 internal user 예산 순으로 고릅니다. 렌더는 캐시 파일만 읽으며, 갱신은 5분에 한 번 분리된 백그라운드 프로세스가 수행합니다 (update-check 와 같은 구조라 statusline 이 네트워크를 기다리지 않습니다).
|
|
505
|
-
- `max_budget` 이 없는 무제한 키는 게이지를 만들지 않습니다. 이 경우에도 `💵` 월 지출 세그먼트는 세션 로그 기반이라 그대로 표시됩니다.
|
|
506
|
-
- 상태 확인: `claude-token-saver litellm-budget` (캐시 출력) · `litellm-budget --refresh` (즉시 조회).
|
|
559
|
+
### Editing a document: copy, then script
|
|
507
560
|
|
|
508
|
-
|
|
561
|
+
Conversion is one-way — editing the cached `.md` changes nothing in the source. The hook refuses `Edit`/`Write` on both the cache and the original binary, and points at the right path instead: copy the original, edit the copy with a script, re-convert the copy to verify.
|
|
509
562
|
|
|
510
|
-
|
|
563
|
+
`install-converter` puts the editing libraries (python-pptx, python-docx, openpyxl) in the same venv, so a structural request like "swap the chart on slide 23 for a line chart" is a short script the agent writes on the spot. `.fig` edits go through openfig-core, which encodes as well as parses.
|
|
511
564
|
|
|
512
|
-
|
|
565
|
+
All four formats were exercised end to end on 2026-09-06: 10 docx run replacements plus three consecutive re-saves, a pptx bar-to-line chart swap with an added data point, xlsx value edits and a new row, and a fig text edit with re-encode and re-parse. In every case the original was byte-identical afterwards and the re-converted copy showed the change. One caveat: removing a chart shape from a pptx leaves the old chart XML part orphaned — PowerPoint ignores it, but delete the part and its rels for a clean file. Charts and images never appear in a conversion, so visual edits must be confirmed in the application itself.
|
|
566
|
+
|
|
567
|
+
### DRM-wrapped documents
|
|
568
|
+
|
|
569
|
+
Encryption and DRM are different problems with different answers. Enterprise DRM (Fasoo, MarkAny, SoftCamp and the like) does not password a document — it wraps the whole file, and only processes the vendor's agent has whitelisted ever see plaintext. Python is not one of them, so what sits on disk is ciphertext behind a vendor header, and **no password will open it.**
|
|
570
|
+
|
|
571
|
+
The first bytes decide which story to tell: a zip header means a truncated download, an OLE container means a password, and neither means the file is not that format at all.
|
|
572
|
+
|
|
573
|
+
```
|
|
574
|
+
✗ bad-archive: File is not a zip file → download it again
|
|
575
|
+
✗ encrypted: password-protected Office file → ask for an unlocked copy
|
|
576
|
+
✗ drm-protected: DRM-wrapped file (FASOO) → ask for a copy released from DRM
|
|
577
|
+
```
|
|
578
|
+
|
|
579
|
+
Vendor names are matched only to say which client to go to; the classification stands without recognising the vendor. PDFs are judged the same way through their public DRM security-handler names (FOPN_foweb, EBX_HANDLER, Adobe.APS).
|
|
580
|
+
|
|
581
|
+
### Locked documents, and Windows
|
|
582
|
+
|
|
583
|
+
**A password-protected document is a state, not an error.** Office encrypts by wrapping the package in an OLE compound file rather than a zip, so opening one as a zip used to report "not a zip file" — which reads as a broken download and sends the user after the wrong problem. It is now identified before conversion:
|
|
584
|
+
|
|
585
|
+
```
|
|
586
|
+
✗ encrypted: password-protected Office file (OLE-wrapped)
|
|
587
|
+
✗ encrypted: password-protected PDF
|
|
588
|
+
```
|
|
589
|
+
|
|
590
|
+
The model is told to ask for an unlocked copy. This tool never asks for or stores a password, and never blocks the original `Read`, so work continues either way. A PDF that merely restricts printing still opens and still converts — checked against a false positive — and a legacy `.xls`, which is an OLE file by design, is not mistaken for an encrypted one.
|
|
591
|
+
|
|
592
|
+
**Windows is supported.** For teams with Windows machines:
|
|
593
|
+
|
|
594
|
+
- The Python search uses the `py -3` launcher. `python3` is rarely on PATH there, and a bare `python` may be the Store alias stub that opens a web page instead of running anything. Venv interpreters are looked for at `Scripts\python.exe`.
|
|
595
|
+
- The `.fig` parser installs through `npm.cmd` via the shell, and the package spec dropped its caret (`openfig-core@0.4.x`): in cmd.exe `^` is the escape character and never reaches npm.
|
|
596
|
+
- The background install and every child process set `windowsHide`, so no console window appears in the middle of someone's prompt.
|
|
597
|
+
|
|
598
|
+
`claude-token-saver doc2md --clean` empties the conversion cache; `doc2md off` removes the hook. Removal filters for this tool's own entry, so anything else you registered under `PreToolUse` stays.
|
|
599
|
+
|
|
600
|
+
## 🌐 Behind a gateway (Bedrock / Vertex)
|
|
601
|
+
|
|
602
|
+
A gateway reports the cache-creation total but never the 5m/1h split. That left the tool unable to tell "nothing cached yet" from "this provider does not say", and the fallback assumed an hour — for a window that is really five minutes on Bedrock, overstating it twelvefold.
|
|
603
|
+
|
|
604
|
+
Since v3.26.0 the gateway is detected from the model ids in the transcript, which fixes:
|
|
605
|
+
|
|
606
|
+
- The countdown falls back to 5 minutes, labelled `5m?`. Three grades of certainty get three labels: measured (`5m`), inferred (`5m?`), unknown (`?`).
|
|
607
|
+
- In a 5-minute bucket the countdown colour follows absolute time rather than a percentage. 30% of five minutes is 90 seconds, and green there promised comfort that was not there.
|
|
608
|
+
- The `⚠ 5m TTL` warning finally reaches these users — with different advice, since no subscription plan changes a gateway's TTL.
|
|
609
|
+
- `Extra cost if 5m-only` is only asked of sessions that have 1h writes to lose. Elsewhere the arithmetically honest `+$0` read as an endorsement of the bucket you are already stuck in.
|
|
610
|
+
- Delegated runs dropped for an unpriceable model id show as `🔀 N unresolved` instead of nothing, which used to be indistinguishable from never having delegated.
|
|
611
|
+
- Environment variables set to a `foundation-model` ARN now resolve. An opaque `application-inference-profile` id still does not: guessing at it is how wrong prices enter the ledger.
|
|
612
|
+
|
|
613
|
+
If the detection is wrong, pin it with `claude-token-saver mode ttl=5m` (or `ttl=1h`). An explicit value outranks the measurement.
|
|
614
|
+
|
|
615
|
+
### LiteLLM: your key budget stands in for the missing 5h/7d caps (v3.35.0)
|
|
616
|
+
|
|
617
|
+
Behind a LiteLLM proxy (Bedrock and friends), Claude Code's stdin never carries `rate_limits`, so the `✦ current` / `📅 weekly` gauges simply do not exist. LiteLLM does track per-key budgets, so the statusline draws a budget gauge in their place.
|
|
618
|
+
|
|
619
|
+
- Detection: `ANTHROPIC_BASE_URL` points somewhere other than the official endpoint, and `ANTHROPIC_AUTH_TOKEN` (or `ANTHROPIC_API_KEY`) is set.
|
|
620
|
+
- The proxy is asked via `GET /key/info` and `GET /user/info` — only the calling key's own data. Renders read a cache file; a detached background process refreshes it every 5 minutes (same shape as the update check), so the statusline never waits on the network.
|
|
621
|
+
- Budget source priority follows real-world usage: the **team-membership budget** (`team_memberships[].spend` + its linked budget table row) first, then the key's own `max_budget`, then the internal-user budget. Verified against a Dockerized LiteLLM, including memberships whose budget diverges from the team max into a separate budget-table row.
|
|
622
|
+
- Unlimited keys (no `max_budget`) get no gauge. The `💵` monthly-spend segment still shows, since it comes from session logs.
|
|
623
|
+
- Inspect with `claude-token-saver litellm-budget` (cached) or `litellm-budget --refresh` (query now).
|
|
624
|
+
|
|
625
|
+
One related non-bug: if your session model is already sonnet, a sonnet-delegation (T1) rule can never save anything, because there is no price gap to capture. That is correct, but `route-scan rules` displayed it identically to "no delegations yet", so it now says outright that the rule does not apply at the current default model.
|
|
626
|
+
|
|
627
|
+
## Spike issue codes
|
|
628
|
+
|
|
629
|
+
| Code | Meaning |
|
|
513
630
|
|---|---|
|
|
514
|
-
| `LARGE_INPUT_PER_REQUEST` |
|
|
515
|
-
| `LOW_HIT_RATE` |
|
|
516
|
-
| `BUCKET_5M_DOMINANT` |
|
|
517
|
-
| `HIGH_OUTPUT_RATIO` |
|
|
518
|
-
| `HIGH_REQUEST_COUNT` |
|
|
519
|
-
| `FREQUENT_CACHE_REBUILD` |
|
|
631
|
+
| `LARGE_INPUT_PER_REQUEST` | single request > 200k input tokens — per-turn re-billing and cap burn spike |
|
|
632
|
+
| `LOW_HIT_RATE` | cache hit rate < 50% |
|
|
633
|
+
| `BUCKET_5M_DOMINANT` | > 70% of cache writes hit the 5m bucket |
|
|
634
|
+
| `HIGH_OUTPUT_RATIO` | output/input > 0.15 (output is 5× input price) |
|
|
635
|
+
| `HIGH_REQUEST_COUNT` | session made 3×+ your median (tool loop?) |
|
|
636
|
+
| `FREQUENT_CACHE_REBUILD` | `cache_creation` > `cache_read` |
|
|
520
637
|
|
|
521
|
-
|
|
638
|
+
Remediation commands are OS-aware (`~/.zshrc` for macOS/Linux/WSL, `setx` for Windows).
|
|
522
639
|
|
|
523
|
-
##
|
|
640
|
+
## Real-world impact — before/after report
|
|
524
641
|
|
|
525
|
-

|
|
526
643
|
|
|
527
|
-
harness 5/5 + ratchet
|
|
644
|
+
harness 5/5 + ratchet applied to the author's own Claude Code work, normalized **per user message** (cutoff 2026-05-02, Opus 4.7 pricing):
|
|
528
645
|
|
|
529
|
-
|
|
|
646
|
+
| metric | before (7d / 739 msgs) | after (2d / 157 msgs) | Δ |
|
|
530
647
|
|---|---:|---:|---:|
|
|
531
|
-
|
|
|
532
|
-
|
|
|
533
|
-
|
|
|
534
|
-
|
|
|
648
|
+
| cost / user message | $2.345 | $1.910 | **−18.6%** |
|
|
649
|
+
| output tokens / message | 7,391 | 6,052 | −18.1% |
|
|
650
|
+
| assistant turns / message | 9.73 | 8.83 | −9.2% |
|
|
651
|
+
| tool calls / message | 5.72 | 5.25 | −8.2% |
|
|
535
652
|
|
|
536
|
-
|
|
653
|
+
Same request resolved in fewer round-trips → first-try success rate up — the effect of PEV + Structured Task forcing one-shot delivery.
|
|
537
654
|
|
|
538
655
|
<details>
|
|
539
|
-
<summary
|
|
656
|
+
<summary>Measurement notes — why cache hit rate isn't included · sample caveats</summary>
|
|
540
657
|
|
|
541
|
-
-
|
|
542
|
-
-
|
|
543
|
-
- ⚠️
|
|
658
|
+
- The author is on the Max plan (1-hour cache TTL) with hit rate already converged near ~98%, so little headroom there. **Pro-plan users (5-minute TTL)** likely see hit rate itself rise with the handoff-before-expiry workflow.
|
|
659
|
+
- Handoff-before-expiry: watch the TTL countdown, run `claude-token-saver handoff` just before expiry to dump work state into a markdown brief, start a fresh cache cycle. Same flow handles the 1M warning and cap chips.
|
|
660
|
+
- ⚠️ POST window is only 2 days (157 msgs); statistical confidence is low, and week-to-week topic mix differs, so the tool effect isn't cleanly isolated.
|
|
544
661
|
</details>
|
|
545
662
|
|
|
546
|
-
##
|
|
663
|
+
## Pricing (Jul 2026)
|
|
547
664
|
|
|
548
|
-
|
|
665
|
+
Per million tokens (USD), as used by the cost estimator:
|
|
549
666
|
|
|
550
|
-
|
|
667
|
+
| Tier | Models | Input | 5m Write | 1h Write | Read | Output |
|
|
668
|
+
|---|---|---|---|---|---|---|
|
|
669
|
+
| `claude-fable-5` | Fable 5 / Mythos 5 | $10 | $12.50 | $20 | $1 | $50 |
|
|
670
|
+
| `claude-opus-new` | Opus 4.5 / 4.6 / 4.7 / 4.8 | $5 | $6.25 | $10 | $0.50 | $25 |
|
|
671
|
+
| `claude-opus-legacy` | Opus 4 / 4.1 / 3 | $15 | $18.75 | $30 | $1.50 | $75 |
|
|
672
|
+
| `claude-sonnet` | Sonnet 3.7 / 4 / 4.5 / 4.6 / 5 | $3 | $3.75 | $6 | $0.30 | $15 |
|
|
673
|
+
| `claude-haiku-4-5` | Haiku 4.5 | $1 | $1.25 | $2 | $0.10 | $5 |
|
|
674
|
+
|
|
675
|
+
Source: [Anthropic pricing docs](https://platform.claude.com/docs/en/about-claude/pricing). Sonnet 5 has an introductory $2/$10 rate through 2026-08-31; the estimator uses the standard sticker. Versions ≤ 2.16.x priced Fable 5 at the Sonnet tier (~3× under-estimate) — upgrade to 2.17.0+.
|
|
676
|
+
|
|
677
|
+
### Cache TTL by plan
|
|
678
|
+
|
|
679
|
+
| Plan | TTL | Controlled by |
|
|
680
|
+
|---|---|---|
|
|
681
|
+
| Max ($100–200/mo) | **1h auto** | `tengu_prompt_cache_1h_config` flag |
|
|
682
|
+
| Pro ($20/mo) | **5m fixed** | not configurable |
|
|
683
|
+
| API key | 5m default (1h via beta header) | `cache_control.ttl` |
|
|
684
|
+
|
|
685
|
+
## How it works · Environment
|
|
686
|
+
|
|
687
|
+
Claude Code logs every API call to `~/.claude/projects/<dir>/<session>.jsonl`. This tool dedupes streaming chunks by `requestId` and aggregates `cache_read_input_tokens` / `cache_creation.ephemeral_5m/1h_input_tokens` by day and session.
|
|
688
|
+
|
|
689
|
+
Node.js ≥ 18 · macOS / Linux / Windows / WSL · **zero dependencies**.
|
|
551
690
|
|
|
552
691
|
<details>
|
|
553
|
-
<summary
|
|
692
|
+
<summary>Known quirks · Migration · Background</summary>
|
|
554
693
|
|
|
555
|
-
**IntelliJ Claude Code plugin
|
|
694
|
+
**IntelliJ Claude Code plugin** — the statusline widget fuses frames at the character level when emoji are present (`59:548` artifacts). v2.8.5+ detects `TERMINAL_EMULATOR=JetBrains-JediTerm` and falls back to text mode automatically.
|
|
556
695
|
|
|
557
|
-
**
|
|
696
|
+
**If the countdown looks frozen:** ticking while idle requires Claude Code to re-run the statusline command on a timer, controlled by `statusLine.refreshInterval` (seconds, Claude Code v2.1.97+) in `~/.claude/settings.json`. Without it the line only redraws when the conversation updates. If behavior differs per terminal, check three things: ① that machine's Claude Code is ≥ 2.1.97; ② no project `.claude/settings.json` / `settings.local.json` overrides `statusLine` without a refreshInterval; ③ the statusline wrapper actually finds `claude-token-saver` on PATH instead of falling back to a multi-second `npx` run on every render (typical when nvm is not loaded in non-login shells). Re-running `claude-token-saver install` restores refreshInterval=5.
|
|
558
697
|
|
|
559
|
-
**claude-cache-monitor
|
|
698
|
+
**Migration from claude-cache-monitor:**
|
|
560
699
|
```bash
|
|
561
700
|
npm uninstall -g claude-cache-monitor && npm i -g claude-token-saver
|
|
562
701
|
```
|
|
563
|
-
`~/.claude/settings.json
|
|
702
|
+
Also update `statusLine.command` in `~/.claude/settings.json` to `claude-token-saver …`.
|
|
703
|
+
|
|
704
|
+
**Background:** [GitHub Issue #46829](https://github.com/anthropics/claude-code/issues/46829) (cache TTL regression) · [HN discussion](https://news.ycombinator.com/item?id=47736476)
|
|
564
705
|
</details>
|
|
565
706
|
|
|
566
|
-
##
|
|
707
|
+
## Release notes
|
|
567
708
|
|
|
568
|
-
|
|
709
|
+
The full history moved to [CHANGELOG.md](./CHANGELOG.md) (Korean; version headings and command names are language-neutral). Recent changes:
|
|
569
710
|
|
|
570
|
-
- **v3.35.0**:
|
|
571
|
-
- **v3.34.0**: seed
|
|
711
|
+
- **v3.35.0**: A `💵 Sep $42` segment now shows estimated spend since 00:00 on the 1st of the current month, always on — including gateway setups with no 5h/7d caps. LiteLLM gateway users get a `🔑 budget ▰▱ 34% $34/$100` gauge built from the key's budget (`GET /key/info` + `GET /user/info`, team-membership budget first, then key, then internal user — verified against a Dockerized LiteLLM).
|
|
712
|
+
- **v3.34.0**: seed presets offered one at a time, output-language choice at install, context warning raised to 500k.
|
|
572
713
|
|
|
573
|
-
##
|
|
714
|
+
## License
|
|
574
715
|
|
|
575
716
|
MIT
|
|
576
717
|
|
|
577
718
|
---
|
|
578
719
|
|
|
579
|
-
##
|
|
720
|
+
## Who makes this
|
|
580
721
|
|
|
581
722
|
[](https://www.youtube.com/@DeepPulseKR)
|
|
582
723
|
[](https://www.youtube.com/@DeepPulseEN)
|
|
583
724
|
[](https://rootstudioyaml.github.io/)
|
|
584
725
|
|
|
585
|
-
|
|
726
|
+
Built and used at **DeepPulse**, a channel about AI developer tooling. The [launch Short (60s)](https://www.youtube.com/shorts/RaD8qMsPTnA) covers where this came from and how it is used.
|