claude-token-saver 3.35.1 → 3.37.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.ko.md +620 -0
- package/README.md +489 -345
- package/package.json +2 -2
- package/presets/korean-style/supplement.md +93 -0
- package/src/korean-lint.cjs +9 -0
- package/src/korean-style.js +16 -1
- package/src/model-rules.js +61 -0
- package/README.en.md +0 -718
package/README.md
CHANGED
|
@@ -1,597 +1,741 @@
|
|
|
1
|
-
|
|
1
|
+
**English** · [한국어](./README.ko.md)
|
|
2
2
|
|
|
3
3
|
[](https://www.npmjs.com/package/claude-token-saver)
|
|
4
4
|
|
|
5
|
-
🌐 **[
|
|
5
|
+
🌐 **[Project page](https://rootstudioyaml.github.io/claude-token-saver/)**
|
|
6
6
|
|
|
7
7
|
# claude-token-saver
|
|
8
8
|
|
|
9
|
-
|
|
9
|
+
**Shows what it saved, on two lines.** It moves the easy work your expensive model keeps repeating onto cheaper ones, and turns documents the model cannot read into Markdown. Both figures are ledger entries rather than estimates, and whichever saved more takes the top line. Zero dependencies, one-line install.
|
|
10
10
|
|
|
11
|
-

|
|
12
12
|
|
|
13
13
|
```bash
|
|
14
14
|
npm i -g claude-token-saver
|
|
15
15
|
```
|
|
16
16
|
|
|
17
|
-
|
|
17
|
+
Four numbers are the whole pitch.
|
|
18
18
|
|
|
19
|
-
-
|
|
20
|
-
-
|
|
21
|
-
-
|
|
19
|
+
- **Beats every single model on public benchmark data**: the shipped tier criteria score 59.1% on 11,696 LLMRouterBench instances against the best single model's 57.9% — at 31% less than gpt-5, 64% less than gemini-2.5-pro ([benchmark](./docs/BENCHMARK.md))
|
|
20
|
+
- **95.8% fewer tokens per document**: a 30MB deck read as Markdown cost 22,610 tokens instead of 540,429 ([evidence](#-doc2md--documents-become-markdown-before-the-model-reads-them))
|
|
21
|
+
- **18.6% lower cost**: measured before/after adopting the Harness principles ([evidence](#real-world-impact--beforeafter-report))
|
|
22
|
+
- **Routing savings are a per-run ledger**: the price difference of each delegated run, not an estimate ([evidence](#-the-savings-figure-is-a-ledger-entry-not-an-estimate))
|
|
22
23
|
|
|
23
|
-
|
|
24
|
+
```text
|
|
25
|
+
accuracy total cost, 11,696 queries
|
|
26
|
+
tier-criteria router 59.1% ← best $268 ▰▰▰▰▰▱▱▱▱▱▱▱▱▱▱
|
|
27
|
+
gpt-5 57.8% $388 ▰▰▰▰▰▰▰▰▱▱▱▱▱▱▱
|
|
28
|
+
gemini-2.5-pro 57.9% $734 ▰▰▰▰▰▰▰▰▰▰▰▰▰▰▰
|
|
29
|
+
```
|
|
30
|
+
|
|
31
|
+
Since v3.35.0 spend is visible too: month-to-date spend shows as `💵 Sep $42`, and on LiteLLM gateways (Bedrock and friends) with no 5h/7d caps, your key budget renders as a `🔑 budget ▰▱ 34% $34/$100` gauge.
|
|
24
32
|
|
|
25
|
-
##
|
|
33
|
+
## Four parts, working together
|
|
26
34
|
|
|
27
|
-
| |
|
|
35
|
+
| | What it does | Effect |
|
|
28
36
|
|---|---|---|
|
|
29
|
-
| 🔀
|
|
30
|
-
| 📄
|
|
31
|
-
| 🅷 **Harness** |
|
|
32
|
-
| ⚙️ **Ratchet** |
|
|
37
|
+
| 🔀 **Routing** | Delegates recurring easy work to cheaper models | Savings recorded per run in a ledger; criteria [benchmarked](./docs/BENCHMARK.md) on public data |
|
|
38
|
+
| 📄 **Document conversion** | Turns pptx/xlsx/pdf/docx/fig into Markdown before the model reads them | **510,000 tokens** saved on one deck ([below](#-doc2md--documents-become-markdown-before-the-model-reads-them)) |
|
|
39
|
+
| 🅷 **Harness** | Blocks the token-burning habits: unevidenced "done", skipped verification (5 principles) | **−18.6% cost** ([measured](#real-world-impact--beforeafter-report)) |
|
|
40
|
+
| ⚙️ **Ratchet** | Freezes each error you hit into a rule | Same mistake stops recurring |
|
|
41
|
+
|
|
42
|
+
One install sets up all four. The measured −18.6% comes from the harness and ratchet; routing and conversion savings sit on top of it.
|
|
33
43
|
|
|
34
|
-
|
|
44
|
+
The two savings figures are never added together, because they answer different questions. Routing says "the same work ran on a cheaper model". Conversion says "a file you could not read became readable, without pushing the original through the context window". The statusline gives each its own line and puts the larger one first.
|
|
35
45
|
|
|
36
|
-
##
|
|
46
|
+
## Contents
|
|
37
47
|
|
|
38
|
-
-
|
|
39
|
-
-
|
|
40
|
-
-
|
|
41
|
-
-
|
|
48
|
+
- **Start here**: [Getting started](#getting-started) · [Reading the statusline](#reading-the-statusline) · [Commands](#commands)
|
|
49
|
+
- **Savings**: [The routing ledger](#-the-savings-figure-is-a-ledger-entry-not-an-estimate) · [route-scan](#-route-scan--this-recurring-task-could-run-on-a-cheaper-tier) · [doc2md](#-doc2md--documents-become-markdown-before-the-model-reads-them) · [seed](#-seed-delegation-that-works-from-the-first-session) · [Benchmark](./docs/BENCHMARK.md)
|
|
50
|
+
- **Guardrails**: [Harness](#-harness-mode) · [compact-window](#-compact-window--pin-where-a-1m-session-compacts) · [Korean writing guidance](#-korean-writing-guidance)
|
|
51
|
+
- **Spend & environments**: [Monthly spend · LiteLLM key budget](#litellm-your-key-budget-stands-in-for-the-missing-5h7d-caps-v3350) · [Gateways (Bedrock/Vertex)](#-behind-a-gateway-bedrock--vertex) · [Spike issue codes](#spike-issue-codes) · [Measured impact](#real-world-impact--beforeafter-report)
|
|
42
52
|
|
|
43
|
-
##
|
|
53
|
+
## Getting started
|
|
44
54
|
|
|
45
|
-
|
|
55
|
+
**Prerequisite:** Node.js ≥ 18 (`node -v` · macOS `brew install node` · Windows `winget install OpenJS.NodeJS.LTS` · Linux/WSL: [nvm](https://github.com/nvm-sh/nvm) recommended)
|
|
46
56
|
|
|
47
57
|
```bash
|
|
48
|
-
npm uninstall -g claude-cache-monitor # (
|
|
58
|
+
npm uninstall -g claude-cache-monitor # (previous-package users only)
|
|
49
59
|
npm i -g claude-token-saver
|
|
50
60
|
```
|
|
51
61
|
|
|
52
|
-
|
|
62
|
+
The statusline appears at the bottom of Claude Code right away. If auto-registration was skipped (`--ignore-scripts`, sudo, sandboxed installs), run `claude-token-saver install`.
|
|
53
63
|
|
|
54
|
-
|
|
64
|
+
One install sets up everything: **statusline, Skill, SessionStart hook, the 🅷 Harness (5 principles), and a first route-scan.** The harness and the Korean writing guidance **show what they add and ask before enabling it.** The harness is **appended** to `~/.claude/CLAUDE.md` as a marked block (your existing content is backed up and preserved) and is left alone if one is already there.
|
|
55
65
|
|
|
56
|
-
|
|
66
|
+
Outside a terminal — npm `postinstall`, CI, piped stdin — the question is skipped and the old defaults apply. Use `--yes` or `--no-input` to skip it deliberately, `CTS_NO_HARNESS=1 npm i -g claude-token-saver` to skip the harness entirely, and `claude-token-saver harness uninit --global` to undo it.
|
|
57
67
|
|
|
58
|
-
> ⚠️ sudo
|
|
68
|
+
> ⚠️ Avoid `sudo` global installs — the Skill lands in root's `~/.claude` instead of yours. Use nvm/fnm/Volta or `npm config set prefix ~/.npm-global`.
|
|
59
69
|
|
|
60
|
-
###
|
|
70
|
+
### What the install turns on, and what stays manual
|
|
61
71
|
|
|
62
|
-
|
|
72
|
+
Everything that costs nothing until it is needed is on after a plain install. The only manual items are the ones that change Claude Code's own settings or need a human to pick a scope.
|
|
63
73
|
|
|
64
|
-
|
|
|
74
|
+
| Feature | After install | How to turn it off |
|
|
65
75
|
|---|---|---|
|
|
66
|
-
| statusline (
|
|
67
|
-
| `/claude-token-saver` Skill |
|
|
68
|
-
| SessionStart
|
|
69
|
-
| UserPromptSubmit
|
|
70
|
-
|
|
|
71
|
-
| 🅷 Harness 5
|
|
72
|
-
| doc2md
|
|
73
|
-
| doc2md
|
|
74
|
-
|
|
|
75
|
-
| compact-window
|
|
76
|
-
|
|
|
77
|
-
| **compact-window
|
|
78
|
-
|
|
|
79
|
-
| `handoff` (
|
|
76
|
+
| statusline (diagnostic chips, savings ledger) | on | `claude-token-saver uninstall` |
|
|
77
|
+
| `/claude-token-saver` Skill | on | same |
|
|
78
|
+
| SessionStart hook (route-scan refresh) | on | same |
|
|
79
|
+
| UserPromptSubmit hook (brief injection) | on | same |
|
|
80
|
+
| First route-scan (last 14 days of logs) | runs once during the install | n/a |
|
|
81
|
+
| 🅷 Harness 5 principles (`~/.claude/CLAUDE.md`) | on (shown and confirmed once at a terminal) | `harness uninit --global`, `CTS_NO_HARNESS=1` |
|
|
82
|
+
| doc2md hooks (Read, Edit/Write, prompt) | on | `doc2md off`, `CTS_NO_DOC2MD=1` |
|
|
83
|
+
| doc2md converter (markitdown venv) | offered at a terminal; an unattended install prints the command | install later with `doc2md install-converter` |
|
|
84
|
+
| Korean writing guidance | on when the locale is Korean (asked at a terminal) | `korean off`, `CTS_NO_KOREAN=1` |
|
|
85
|
+
| compact-window warning chip | on | `compact-window off` |
|
|
86
|
+
| Update-available chip | on | `CTS_NO_UPDATE_CHECK=1` |
|
|
87
|
+
| **Pinning compact-window** (`autoCompactWindow` 500k) | **off — run it yourself** | `compact-window set --global` or `--project` |
|
|
88
|
+
| **Model-fitting rules** (`ratchet-model.md` delegations) | **candidates are proposed only** | review and approve with `route-scan rules` |
|
|
89
|
+
| `handoff` (back up work before a cap) | an on-demand command | n/a |
|
|
80
90
|
|
|
81
|
-
`compact-window set`
|
|
91
|
+
`compact-window set` writes into Claude Code's `settings.json` and a human has to choose global or project scope, so it is never run for you. Model-fitting rules keep an approval step for the same reason: which work belongs on a cheaper tier is your call.
|
|
82
92
|
|
|
83
93
|
|
|
84
|
-
## 🔀
|
|
94
|
+
## 🔀 The savings figure is a ledger entry, not an estimate
|
|
85
95
|
|
|
86
|
-
|
|
96
|
+
Every delegated run is recorded like this:
|
|
87
97
|
|
|
88
98
|
```
|
|
89
|
-
|
|
99
|
+
before after gap
|
|
90
100
|
claude-opus-5 → haiku-4-5 = $0.57
|
|
91
|
-
(
|
|
92
|
-
|
|
93
|
-
|
|
101
|
+
(the model (what (same token counts,
|
|
102
|
+
handling this actually priced against
|
|
103
|
+
before the rule) ran it) both models)
|
|
94
104
|
```
|
|
95
105
|
|
|
96
106
|
```bash
|
|
97
|
-
$ claude-token-saver route-scan savings #
|
|
107
|
+
$ claude-token-saver route-scan savings # trace every dollar back to its rule
|
|
98
108
|
|
|
99
|
-
🔀
|
|
109
|
+
🔀 Routing saved, lifetime $2.09 (last 7d $1.40 · 30d $2.09)
|
|
100
110
|
|
|
101
|
-
|
|
102
|
-
claude-fable-5 → claude-sonnet-5 — 1
|
|
103
|
-
claude-opus-5 → claude-haiku-4-5 — 1
|
|
111
|
+
By model change:
|
|
112
|
+
claude-fable-5 → claude-sonnet-5 — 1 run, $0.72
|
|
113
|
+
claude-opus-5 → claude-haiku-4-5 — 1 run, $0.57
|
|
104
114
|
|
|
105
|
-
|
|
115
|
+
By run (newest first):
|
|
106
116
|
2026-08-22 $0.51 claude-fable-5 → claude-haiku-4-5
|
|
107
|
-
|
|
117
|
+
rule: T2|paste|-Users-me-projects-my-app
|
|
108
118
|
```
|
|
109
119
|
|
|
110
|
-
|
|
120
|
+
**What is excluded** — an honest number beats a big one:
|
|
111
121
|
|
|
112
|
-
-
|
|
113
|
-
-
|
|
122
|
+
- Delegations no registered rule covers (`Explore`, your own agents, plugin subagents): this tool did not route them.
|
|
123
|
+
- Model ids the pricing table cannot recognize: the run is dropped rather than priced wrong.
|
|
114
124
|
|
|
115
125
|
---
|
|
116
126
|
|
|
117
|
-
## ⚡
|
|
127
|
+
## ⚡ What else the statusline catches
|
|
118
128
|
|
|
119
129
|
| | |
|
|
120
130
|
|---|---|
|
|
121
|
-
| 🚨
|
|
122
|
-
| 🧠
|
|
123
|
-
| 💵
|
|
124
|
-
| 🇰🇷
|
|
131
|
+
| 🚨 **No surprise rate limits** | Warns when the 5H/7D window hits 90%; `handoff` backs up your work |
|
|
132
|
+
| 🧠 **Cache waste detection** | Hit rate, TTL, 1M-context detection — spikes diagnosed with issue codes |
|
|
133
|
+
| 💵 **Spend visibility** | Month-to-date spend (`💵 Sep $42`) always on; behind a LiteLLM gateway the key budget gauge (`🔑 budget 34% $34/$100`) stands in for the missing 5h/7d caps ([below](#litellm-your-key-budget-stands-in-for-the-missing-5h7d-caps-v3350)) |
|
|
134
|
+
| 🇰🇷 **Korean writing guidance** | Offered at install time, defaulting to your locale ([below](#-korean-writing-guidance)) |
|
|
125
135
|
|
|
126
|
-
##
|
|
136
|
+
## Not a router — 60 seconds
|
|
127
137
|
|
|
128
|
-
|
|
129
|
-
|
|
130
|
-
|
|
138
|
+
It never intercepts a request in realtime.
|
|
139
|
+
**After a session ends** it reads your local logs, finds the easy patterns your expensive model
|
|
140
|
+
kept handling, and promotes them into rules so a cheaper model takes them **from the next session
|
|
141
|
+
onward**. Rules are scoped global or per-project.
|
|
131
142
|
|
|
132
|
-
###
|
|
143
|
+
### Why realtime model routing can cost more, not less
|
|
133
144
|
|
|
134
|
-
|
|
145
|
+
Never switching models mid-session is the point of this design.
|
|
135
146
|
|
|
136
|
-
|
|
147
|
+
Prompt caches are **kept per model.** Switch to a cheaper model mid-session and it starts from a cold cache, re-reading the whole conversation at full input price. A cache hit costs about a tenth of that, so past roughly 20k tokens of history **one switch can erase everything the cheaper model was going to save.** You moved the work down a tier and the bill went up: the central paradox of realtime routing.
|
|
137
148
|
|
|
138
|
-
|
|
149
|
+
Teams shipping routing products have turned the feature off for exactly this reason: [LLM 라우터를 만든 사람들이 직접 껐습니다 #Shorts](https://www.youtube.com/shorts/SK-GoAABjbg) (Korean).
|
|
139
150
|
|
|
140
|
-
|
|
151
|
+
So this tool never touches the main session's model. It delegates to **subagents only**, which leaves the main session's cache intact and runs the delegated work on a cheap model in its own context. That is why the savings are not cancelled out by cache loss.
|
|
141
152
|
|
|
142
153
|
```bash
|
|
143
154
|
npm i -g claude-token-saver@latest
|
|
144
|
-
claude-token-saver route-scan #
|
|
145
|
-
claude-token-saver route-scan rules #
|
|
146
|
-
claude-token-saver route-scan savings #
|
|
155
|
+
claude-token-saver route-scan # find delegation candidates in your own history (0 LLM calls)
|
|
156
|
+
claude-token-saver route-scan rules # list promoted rules · rm <N> to remove
|
|
157
|
+
claude-token-saver route-scan savings # audit every dollar the routing saved
|
|
147
158
|
```
|
|
148
159
|
|
|
149
|
-
|
|
150
|
-
|
|
160
|
+
Thresholds come from **your own last-14-day distribution (p25/p75)**, not someone else's benchmark.
|
|
161
|
+
Measured rule-health — whether a delegated run actually succeeded — landed in [v3.9.0](#v390-2026-08-01).
|
|
151
162
|
|
|
152
163
|
---
|
|
153
164
|
|
|
154
|
-
##
|
|
165
|
+
## Reading the statusline
|
|
155
166
|
|
|
156
|
-
|
|
167
|
+
Once the savings ledger has entries it renders as **two rows** — routing savings on row 1, diagnostics on row 2.
|
|
157
168
|
|
|
158
169
|
```
|
|
159
170
|
🔀 Routing saved $2.09 | fable→sonnet 1× $0.72 · opus→haiku 1× $0.57
|
|
160
171
|
⚠ Ctx 500k+ · 🅷 5/5 · 🤖 Opus 5 · 🧠 Cache hit 98.8% · ⏳ Cache expires 59:46 · ✦ current ███▓░░ 62% 🔄 21:33 · 📅 weekly ██▒░░░ 38% 🔄 Tue 19:33 · 📦 Ctx 47% of 1M · 💰 Cache saved $1.0K · last 1d
|
|
161
172
|
```
|
|
162
173
|
|
|
163
|
-
|
|
174
|
+
With an empty ledger (no measured delegation yet) row 1 is not drawn and the layout stays single-line. If your build renders only the first row (some macOS Claude Code versions), pass `--single-line`.
|
|
164
175
|
|
|
165
|
-
|
|
|
176
|
+
| Segment | Meaning |
|
|
166
177
|
|---|---|
|
|
167
|
-
| `🔀`
|
|
168
|
-
| `📄`
|
|
169
|
-
| `🤖` |
|
|
170
|
-
| `🅷 5/5` |
|
|
171
|
-
| `🧠` |
|
|
172
|
-
| `⏳` |
|
|
173
|
-
| `✦ current` / `📅 weekly` | 5
|
|
174
|
-
| `📦` |
|
|
175
|
-
| `💵 Sep $42` |
|
|
176
|
-
| `🔑 budget` | **LiteLLM
|
|
177
|
-
| `💰` |
|
|
178
|
-
| `v3.24.0` |
|
|
179
|
-
| `⬆ v3.24.0 → 3.25.0` |
|
|
180
|
-
|
|
181
|
-
|
|
178
|
+
| `🔀` **row 1** | Lifetime routing savings + the model moves behind them. The breakdown sums exactly to the total, model names keep only the family (`opus→haiku`). Full audit: `route-scan savings` |
|
|
179
|
+
| `📄` **row 2** | Lifetime doc2md conversion savings with a per-format breakdown. Whichever of routing/conversion saved more takes row 1 |
|
|
180
|
+
| `🤖` | Active model |
|
|
181
|
+
| `🅷 5/5` | Harness principle score ([Harness mode](#-harness-mode)) |
|
|
182
|
+
| `🧠` | Cache hit rate (green at 85%+) |
|
|
183
|
+
| `⏳` | Cache TTL countdown — send a message before expiry to keep the cache warm. Ticking while idle requires Claude Code v2.1.97+ (see [If the countdown looks frozen](#how-it-works--environment)) |
|
|
184
|
+
| `✦ current` / `📅 weekly` | 5-hour / 7-day rate-limit window usage + reset time |
|
|
185
|
+
| `📦` | Context usage (e.g. `Ctx 68% of 1M`) — colored by fill. Current models default to 1M with no premium, but token volume itself drives per-turn cost and 5H/7D burn |
|
|
186
|
+
| `💵 Sep $42` | **Estimated spend since 00:00 on the 1st of this month** (local time). Summed per session with that session's model pricing; always shown, even on gateways with no 5h/7d caps (v3.35.0) |
|
|
187
|
+
| `🔑 budget` | **LiteLLM key budget gauge.** When stdin carries no rate_limits, shows the key's `spend` against `max_budget` as `🔑 budget ▰▱ 34% $34/$100` (v3.35.0, [below](#-behind-a-gateway-bedrock--vertex)) |
|
|
188
|
+
| `💰` | Cumulative savings from prompt caching — a **different** number from row 1's `🔀` (model routing) |
|
|
189
|
+
| `v3.24.0` | The version you are running. Gray, at the tail, when it is the latest one |
|
|
190
|
+
| `⬆ v3.24.0 → 3.25.0` | A newer release exists. Actionable, so it moves to the front of the line ([Update notifications](#-update-notifications)) |
|
|
191
|
+
|
|
192
|
+
When something is wrong, a **warning chip leads the line**:
|
|
182
193
|
|
|
183
194
|
```
|
|
184
195
|
🚨 5H █████▓ 94% 🔄 12:36 · 🅷 5/5 · 🤖 Opus 4.8 · 🧠 Cache hit 72.1% · ⚠ Cache miss · 📅 weekly ▓░░░░░ 12% 🔄 Sun 14:26 · 📦 Ctx 200k · last 1d
|
|
185
196
|
```
|
|
186
197
|
|
|
187
|
-
|
|
198
|
+
Chips — `🚨 5H/7D NN%` (cap imminent) · `⚠ Ctx 500k+` (a single request actually exceeded 500k) · `⚠ Cache miss` · `⚠ Input spike` · `⚠ Output heavy` · `⚠ Call surge` · `⚠ Rebuild churn` · `⚠ 5m TTL`. When both windows cross 90% at once, the sooner-resetting one is promoted to 🚨 and the other stays visible as a red segment (v2.16.0+).
|
|
188
199
|
|
|
189
|
-
###
|
|
200
|
+
### When a chip appears
|
|
190
201
|
|
|
191
|
-
|
|
202
|
+
Run the `/claude-token-saver` Skill inside Claude — or just say the chip wording ("5H cap is up", "cache miss") and it auto-activates. The Skill surfaces the **root-cause code + step-by-step fix**. When a cap is imminent, run `claude-token-saver handoff` to back up your work state to markdown and continue in a fresh session.
|
|
192
203
|
|
|
193
|
-
##
|
|
204
|
+
## Commands
|
|
194
205
|
|
|
195
|
-
|
|
206
|
+
Run these in your shell (inside Claude Code, the `/claude-token-saver` Skill is the only entry point):
|
|
196
207
|
|
|
197
|
-
|
|
|
208
|
+
| Command | What it does |
|
|
198
209
|
|---|---|
|
|
199
|
-
| `claude-token-saver` |
|
|
200
|
-
| `claude-token-saver last` |
|
|
201
|
-
| `claude-token-saver history` |
|
|
202
|
-
| `claude-token-saver handoff` |
|
|
203
|
-
| `claude-token-saver mode [keywords...]` |
|
|
204
|
-
| `claude-token-saver harness ...` | 🅷 Harness
|
|
205
|
-
| `claude-token-saver route-scan` |
|
|
206
|
-
| `claude-token-saver route-scan savings` |
|
|
207
|
-
| `claude-token-saver compact-window` | 1M
|
|
208
|
-
| `claude-token-saver korean on\|off\|status` |
|
|
209
|
-
| `claude-token-saver korean lint block\|warn\|off` |
|
|
210
|
-
| `claude-token-saver korean lint scope all\|prose` |
|
|
211
|
-
| `claude-token-saver doc2md on\|off` |
|
|
212
|
-
| `claude-token-saver doc2md
|
|
213
|
-
| `claude-token-saver mode ttl=5m\|1h\|auto` |
|
|
214
|
-
| `claude-token-saver --version` |
|
|
215
|
-
| `claude-token-saver update-check` |
|
|
216
|
-
| `claude-token-saver upgrade` |
|
|
217
|
-
| `claude-token-saver install` | Skill
|
|
218
|
-
| `claude-token-saver uninstall [--purge]` |
|
|
219
|
-
|
|
220
|
-
|
|
221
|
-
|
|
222
|
-
|
|
223
|
-
|
|
224
|
-
|
|
225
|
-
|
|
226
|
-
|
|
227
|
-
|
|
228
|
-
|
|
229
|
-
|
|
230
|
-
|
|
231
|
-
|
|
232
|
-
|
|
233
|
-
|
|
234
|
-
|
|
210
|
+
| `claude-token-saver` | Last-1-day diagnostic report (`--days N` / `--hours N`) |
|
|
211
|
+
| `claude-token-saver last` | Most recent warning + remediation |
|
|
212
|
+
| `claude-token-saver history` | Last 7 days of warning transitions |
|
|
213
|
+
| `claude-token-saver handoff` | Back work up to `HANDOFF-*.md` before a cap blocks you |
|
|
214
|
+
| `claude-token-saver mode [keywords...]` | Output config (`icon`/`text`, `en`/`ko`, `1h`–`30d` window, …) |
|
|
215
|
+
| `claude-token-saver harness ...` | 🅷 Harness management (below) |
|
|
216
|
+
| `claude-token-saver route-scan` | Detect recurring easy work on expensive models → propose haiku-delegation ratchet rules (below) |
|
|
217
|
+
| `claude-token-saver route-scan savings` | The routing-savings ledger — per-model-change rollup + per-run log (the evidence behind the figure) |
|
|
218
|
+
| `claude-token-saver compact-window` | Warn when a 1M-context session has no auto-compact cap → pin 400k with `set` (below) |
|
|
219
|
+
| `claude-token-saver korean on\|off\|status` | Inject Korean writing guidance at session start and install the write-time check (below) |
|
|
220
|
+
| `claude-token-saver korean lint block\|warn\|off` | How the write-time check handles findings |
|
|
221
|
+
| `claude-token-saver korean lint scope all\|prose` | Check every text file, or documents only |
|
|
222
|
+
| `claude-token-saver doc2md on\|off` | Convert attached documents to Markdown before the model reads them (below) |
|
|
223
|
+
| `claude-token-saver doc2md <file>` | Convert one file by hand. Diagnostic: it prints the refusal reason instead of swallowing it |
|
|
224
|
+
| `claude-token-saver mode ttl=5m\|1h\|auto` | Pin the cache TTL bucket. The default `auto` trusts the measured split, then falls back to gateway detection |
|
|
225
|
+
| `claude-token-saver --version` | Print the installed version |
|
|
226
|
+
| `claude-token-saver update-check` | Is a newer version out? (`--refresh` to ask now, `--dismiss` to mute this version's offer) |
|
|
227
|
+
| `claude-token-saver upgrade` | Install the latest release with the package manager that installed this copy (`--print` shows the command only) |
|
|
228
|
+
| `claude-token-saver install` | Manually register Skill + statusline |
|
|
229
|
+
| `claude-token-saver uninstall [--purge]` | Remove the hooks, statusline and skill it registered. Recorded savings are kept unless `--purge` is given |
|
|
230
|
+
|
|
231
|
+
The output language is decided once, at install time: a terminal install proposes the system locale and asks whether to use Korean, while an unattended install records what the locale says. Once recorded it is never asked again, not even on an upgrade. Change it later with `mode ko` / `mode en`, or pin it for a scripted install with `CTS_LANG=ko` / `CTS_LANG=en`. Statusline chips stay symbolic either way.
|
|
232
|
+
|
|
233
|
+
<details>
|
|
234
|
+
<summary>All CLI options</summary>
|
|
235
|
+
|
|
236
|
+
| Flag | Description | Default |
|
|
237
|
+
|------|-------------|---------|
|
|
238
|
+
| `--days, -d` | Analysis period in days | 30 |
|
|
239
|
+
| `--hours` | Analysis window in hours (overrides `--days`) | – |
|
|
240
|
+
| `--format, -f` | `table` / `json` / `csv` | table |
|
|
241
|
+
| `--project, -p` | Filter by project directory | all |
|
|
242
|
+
| `--threshold` | Hit-rate alert threshold (0.0–1.0) | 0.7 |
|
|
243
|
+
| `--statusline` | One-line statusline output | – |
|
|
244
|
+
| `--icon` | Use 🧠 / ⏳ / 💰 / 📦 icons | text |
|
|
245
|
+
| `--verbose` | Longer labels | – |
|
|
246
|
+
| `--no-timer` | Hide TTL countdown | show |
|
|
247
|
+
| `--no-color` | Strip ANSI codes | – |
|
|
248
|
+
| `--segments=…` | Limit statusline segments (e.g. `model,five_hour,seven_day,saved`) | all |
|
|
249
|
+
| `--install-hook` / `--uninstall-hook` | Manage the PostToolUse hook | – |
|
|
250
|
+
</details>
|
|
251
|
+
|
|
252
|
+
## ⬆ Update notifications
|
|
253
|
+
|
|
254
|
+
A statusline cannot open a dialog, and it re-renders every ~300ms, so it can never touch the network while drawing. The notification is therefore split in two:
|
|
255
|
+
|
|
256
|
+
- **The statusline tells you.** Up to date: a quiet gray `v3.24.0` at the tail. Newer release out: `⬆ v3.24.0 → 3.25.0` in yellow, moved to the front. Never red — nothing is broken.
|
|
257
|
+
- **Session start asks you.** On a new session or `/clear`, the SessionStart hook injects one line telling the model a newer version exists and to ask before installing anything. Only after you agree does it run `claude-token-saver upgrade`.
|
|
258
|
+
- **Declining sticks.** `claude-token-saver update-check --dismiss` mutes the offer for that version; the next release asks again. The statusline chip stays — you declined the question, not the fact.
|
|
259
|
+
|
|
260
|
+
The registry lookup runs at most once every 24h in a detached background process and only ever writes a cache file (`update-check.json`) — the same shape npm's `update-notifier` uses. A failed check still stamps its timestamp, so an offline machine backs off instead of retrying on every render. Turn checks off entirely with `CTS_NO_UPDATE_CHECK=1` or `NO_UPDATE_NOTIFIER`.
|
|
261
|
+
|
|
262
|
+
## 🅷 Harness mode
|
|
263
|
+
|
|
264
|
+
Bootstrap five engineering principles (Ratchet · Evidence · PEV · Structured Task · Default Safe Path) into `CLAUDE.md` with one command; the statusline scores it as `🅷 5/5`. When the same error keeps recurring, a `🅷⚠ ratchet?` nudge appears so you can promote it to a rule.
|
|
235
265
|
|
|
236
266
|
```bash
|
|
237
|
-
claude-token-saver harness init #
|
|
238
|
-
claude-token-saver harness init --global # ~/.claude/CLAUDE.md
|
|
239
|
-
claude-token-saver harness check #
|
|
240
|
-
claude-token-saver harness analyze #
|
|
241
|
-
claude-token-saver harness promote <N> --project|--global #
|
|
242
|
-
claude-token-saver harness promote "
|
|
243
|
-
claude-token-saver harness pull #
|
|
244
|
-
claude-token-saver harness list / rm <N> #
|
|
245
|
-
claude-token-saver harness off | on # 🅷
|
|
267
|
+
claude-token-saver harness init # this project
|
|
268
|
+
claude-token-saver harness init --global # ~/.claude/CLAUDE.md — every project
|
|
269
|
+
claude-token-saver harness check # current score (global fallback honored)
|
|
270
|
+
claude-token-saver harness analyze # run the transcript analysis manually (no hook needed); refreshes harness-state.json
|
|
271
|
+
claude-token-saver harness promote <N> --project|--global # warning #N → ratchet rule (scope required)
|
|
272
|
+
claude-token-saver harness promote "<rule text>" --project|--global # register your own hand-written rules the same way
|
|
273
|
+
claude-token-saver harness pull # register the package's curated ratchet rules into your global ratchet (opt-in, dedupes)
|
|
274
|
+
claude-token-saver harness list / rm <N> # view / delete rules (auto .bak)
|
|
275
|
+
claude-token-saver harness off | on # toggle the 🅷 chip
|
|
246
276
|
```
|
|
247
277
|
|
|
248
|
-
- `promote
|
|
249
|
-
- `pull
|
|
250
|
-
- `seed
|
|
251
|
-
- 🅷⚠
|
|
278
|
+
- `promote` **requires** `--project`/`--global` in non-TTY contexts (scripts, LLM calls) — a scope choice is never silently made for the caller.
|
|
279
|
+
- `pull` registers the **author-curated ratchet rules** bundled with the package (`presets/ratchet-rules.json` — only general-purpose rules promoted from real recurring mistakes) into your global ratchet (`~/.claude/ratchet.md`). `install`/`init` never auto-inject anything; `pull` is always opt-in and idempotent. Drop any rule you dislike with `harness rm`.
|
|
280
|
+
- `seed` offers the same presets **one at a time**. Where `pull` registers the whole ratchet set in one go, `seed` covers the model-fitting presets too and asks about each of them in the first session after an install or upgrade ([below](#-seed-delegation-that-works-from-the-first-session)).
|
|
281
|
+
- 🅷⚠ runtime warnings (`ratchet?` `no-evidence` `PEV-skip`) expire after 30 minutes, subdirectory sessions match their project correctly, and PEV-skip counts only mutating tools (Edit/Write/Bash) so read-only research sessions don't trip it (v2.16.0+).
|
|
252
282
|
|
|
253
283
|
<details>
|
|
254
|
-
<summary>⚠️ <code>harness rm</code
|
|
284
|
+
<summary>⚠️ <code>harness rm</code> — checklist before deleting</summary>
|
|
255
285
|
|
|
256
|
-
|
|
286
|
+
The whole point of the ratchet is **one-direction accumulation**. Deleting rules casually means the same mistakes return.
|
|
257
287
|
|
|
258
|
-
-
|
|
259
|
-
-
|
|
260
|
-
-
|
|
288
|
+
- **Rule too broad, blocking valid cases?** → ❌ delete ✅ narrow the condition (e.g. `"no hardcoded values"` → `"no hardcoded values outside tests"`)
|
|
289
|
+
- **Rule too narrow, almost never fires?** → ❌ delete ✅ leave it (zero cost)
|
|
290
|
+
- **Genuinely wrong?** → ✅ delete then
|
|
261
291
|
|
|
262
|
-
|
|
292
|
+
An auto `.bak` is kept, but **the session context that earned the rule its place is not recoverable.**
|
|
263
293
|
</details>
|
|
264
294
|
|
|
265
295
|
|
|
266
|
-
## 📦 compact-window
|
|
296
|
+
## 📦 compact-window — pin where a 1M session compacts
|
|
267
297
|
|
|
268
|
-
Claude Code
|
|
298
|
+
Claude Code compacts when usage approaches `min(autoCompactWindow, model max context)`. On a 1M window, with that value unset, compaction only fires near 800k — and until then every request re-bills the whole context. **1M is too large; the recommendation is a 400k–700k band** — 2–3.5x a 200k session's headroom for the genuinely large pastes, with the runaway tail cut off.
|
|
269
299
|
|
|
270
|
-
|
|
300
|
+
**Anything inside the band is left alone.** 400k is the floor where the saving beats the extra compactions, and long sessions often want more room than that. Only an unset window, or one above 700k, is warned about (a smaller one is a deliberate, more aggressive choice).
|
|
271
301
|
|
|
272
|
-
**200k
|
|
302
|
+
**200k sessions are never warned** — their window is already at or below 200k, so the setting cannot change anything.
|
|
273
303
|
|
|
274
304
|
```bash
|
|
275
|
-
claude-token-saver compact-window #
|
|
276
|
-
claude-token-saver compact-window set --global # ~/.claude/settings.json
|
|
277
|
-
claude-token-saver compact-window set --project # <root>/.claude/settings.json
|
|
278
|
-
claude-token-saver compact-window set --global --value 600k
|
|
279
|
-
claude-token-saver compact-window off | on #
|
|
305
|
+
claude-token-saver compact-window # status (model, window, value, source)
|
|
306
|
+
claude-token-saver compact-window set --global # pin 500k (mid-band) in ~/.claude/settings.json
|
|
307
|
+
claude-token-saver compact-window set --project # pin it in <root>/.claude/settings.json
|
|
308
|
+
claude-token-saver compact-window set --global --value 600k # explicit value (100k–1M)
|
|
309
|
+
claude-token-saver compact-window off | on # toggle the warning
|
|
280
310
|
```
|
|
281
311
|
|
|
282
|
-
- 1M
|
|
283
|
-
-
|
|
284
|
-
-
|
|
285
|
-
-
|
|
312
|
+
- On a 1M model with the value unset or above 700k, the statusline shows `🅷⚠ compact-window?` and the session briefing hands the model the exact registration command.
|
|
313
|
+
- Scope (`--global`/`--project`) is **required** for `set` — a global settings file is never edited on a guess.
|
|
314
|
+
- Every other key in `settings.json` is preserved and a `.bak` is written first. Malformed JSON aborts the write untouched.
|
|
315
|
+
- An exported `CLAUDE_CODE_AUTO_COMPACT_WINDOW` beats settings.json; `set` detects that and says so.
|
|
286
316
|
|
|
287
|
-
## 🔀 route-scan
|
|
317
|
+
## 🔀 route-scan — "this recurring task could run on a cheaper tier"
|
|
288
318
|
|
|
289
|
-
|
|
319
|
+
Finds the easy work your expensive model (opus/fable) keeps redoing in your session logs and proposes **haiku/sonnet delegation rules**. Fully local, zero token cost.
|
|
290
320
|
|
|
291
|
-
- **T2 → haiku
|
|
292
|
-
- **T1 → sonnet
|
|
293
|
-
- **T0
|
|
321
|
+
- **T2 → haiku**: lookups, pasted-screen Q&A, simple runs — zero errors, near-zero mutation
|
|
322
|
+
- **T1 → sonnet**: build pipelines, status checks — few mutations, ≤1 error
|
|
323
|
+
- **T0 stays**: repeated errors, heavy mutation, design/analysis — the session model keeps it
|
|
294
324
|
|
|
295
|
-
|
|
296
|
-
1.
|
|
297
|
-
2.
|
|
298
|
-
3.
|
|
325
|
+
Three design pillars:
|
|
326
|
+
1. Difficulty is judged by **outcome, not text guessing** — tool errors, mutating tool calls, output tokens
|
|
327
|
+
2. Thresholds **auto-calibrate to your own 14-day distribution** — fixed constants drift with workload
|
|
328
|
+
3. Promoted rules live in a tool-owned file (`.claude/ratchet-model.md`) that **refreshes itself every scan**, and a `⚠ rule-health` flag fires when a delegated category's error rate climbs — rules report their own staleness
|
|
299
329
|
|
|
300
330
|
```bash
|
|
301
|
-
claude-token-saver route-scan #
|
|
302
|
-
claude-token-saver harness promote R1 --project #
|
|
303
|
-
claude-token-saver route-scan dismiss 1 #
|
|
304
|
-
claude-token-saver route-scan rules #
|
|
305
|
-
claude-token-saver route-scan savings #
|
|
331
|
+
claude-token-saver route-scan # scan (24h cache) + tiered candidates
|
|
332
|
+
claude-token-saver harness promote R1 --project # promote candidate R1 to a model-fitting rule
|
|
333
|
+
claude-token-saver route-scan dismiss 1 # not interested — won't resurface
|
|
334
|
+
claude-token-saver route-scan rules # list model-fitting rules (rm <N> to remove)
|
|
335
|
+
claude-token-saver route-scan savings # the savings ledger — which rule moved work off which model, onto which
|
|
306
336
|
```
|
|
307
337
|
|
|
308
|
-
|
|
338
|
+
Dig deeper: **tier criteria & research evidence** → [docs/TIER_CRITERIA.md](./docs/TIER_CRITERIA.md) (Korean) · **rule-file mechanics, scan triggers, subagent setup** → [docs/ROUTE_SCAN.md](./docs/ROUTE_SCAN.md) (Korean + English)
|
|
309
339
|
|
|
310
|
-
|
|
311
|
-
🔀 라우팅 절감 누적 $2.09 (최근 7일 $1.40 · 30일 $2.09)
|
|
340
|
+
### Behind a gateway (Bedrock / LiteLLM)
|
|
312
341
|
|
|
313
|
-
|
|
314
|
-
claude-fable-5 → claude-sonnet-5 — 1회, $0.72
|
|
315
|
-
claude-opus-5 → claude-haiku-4-5 — 1회, $0.57
|
|
316
|
-
```
|
|
342
|
+
Through a corporate gateway the transcript records an inference-profile ARN where the model id belongs. That string says nothing about `opus` or `haiku`, so older versions read every session as Sonnet — which made **T1 (→sonnet) rules unreachable and zeroed the savings figures**.
|
|
317
343
|
|
|
318
|
-
|
|
344
|
+
Since v3.10.0 the profile id is mapped back to a role (main, opus, sonnet, haiku) and then to the alias your `ANTHROPIC_DEFAULT_*_MODEL` variables declare. The mapping is learned by joining each parent `Task` call to the subagent run it spawned via `toolUseId`. Below three observations, or when the role votes agree less than 80% of the time, the id stays `unknown` and drops out of the delegation aggregate rather than being guessed at.
|
|
319
345
|
|
|
320
|
-
|
|
321
|
-
|
|
322
|
-
사내 게이트웨이를 거치면 로그의 모델명 필드에 추론 프로파일 ARN이 기록됩니다. 그 문자열에는 `opus`·`haiku` 같은 단서가 없어서 예전 버전은 이것을 전부 Sonnet으로 읽었고, 그 결과 **T1(→sonnet) 위임 룰이 하나도 제안되지 않았으며 절감 집계가 0**이었습니다.
|
|
323
|
-
|
|
324
|
-
v3.10.0부터는 프로파일 ID를 역할(main·opus·sonnet·haiku)로 되돌린 뒤 `ANTHROPIC_DEFAULT_*_MODEL` 환경변수가 선언한 별칭으로 치환합니다. 매핑은 부모 세션의 `Task` 호출과 서브에이전트 기록을 `toolUseId`로 조인해 스스로 학습하며, 관측이 3건 미만이거나 역할 판정이 80% 미만으로 갈리면 **추측하지 않고 `unknown`으로 두고 위임 집계에서 제외**합니다.
|
|
325
|
-
|
|
326
|
-
자동 학습이 닿지 않는 환경에서 쓸 수 있는 수동 경로도 있습니다. `<userDataDir>/profile-map.json`에 아래처럼 적으면 되고, 계정 ID와 리전은 `*`로 가려도 매칭됩니다.
|
|
346
|
+
For environments the learner cannot reach, write the mapping yourself in `<userDataDir>/profile-map.json`. Account id and region may be wildcarded:
|
|
327
347
|
|
|
328
348
|
```jsonc
|
|
329
349
|
{
|
|
330
350
|
"modelAliases": {
|
|
331
351
|
"arn:aws:bedrock:*:*:application-inference-profile/<PROFILE_ID>": "claude-opus-5",
|
|
332
|
-
"prod-large": "claude-opus-5", //
|
|
352
|
+
"prod-large": "claude-opus-5", // house aliases map the same way
|
|
333
353
|
"team-*": "claude-haiku-4-5"
|
|
334
354
|
}
|
|
335
355
|
}
|
|
336
356
|
```
|
|
337
357
|
|
|
338
|
-
|
|
358
|
+
**Map house aliases that carry no family name** (`prod-large`, `team-fast`) here too. Shapes that keep the family name are recognized as-is — Bedrock (`anthropic.claude-opus-4-5-v1:0`), Vertex (`claude-opus-4-5@20251101`), and the 1M suffix (`claude-sonnet-4-5[1m]`) — but an alias without one cannot be priced. Rather than report a wrong figure, routing-savings **drops those runs from the aggregate** (both sides of the comparison must be recognizable); one line in the table above brings them back.
|
|
339
359
|
|
|
340
|
-
|
|
360
|
+
That file holds internal identifiers in plain text — do not commit it. On a direct-API machine it is never created and behaviour is unchanged.
|
|
341
361
|
|
|
342
|
-
## 🌱 seed:
|
|
362
|
+
## 🌱 seed: delegation that works from the first session
|
|
343
363
|
|
|
344
|
-
|
|
364
|
+
The model-fitting ratchet (`ratchet-model.md`) **starts empty.** A rule exists only after route-scan has seen the same kind of work recur in your own logs and you have approved that candidate. So a fresh install delegates nothing, and keeps delegating nothing for days — precisely the stretch where the savings would matter most.
|
|
345
365
|
|
|
346
|
-
`seed
|
|
366
|
+
`seed` fills that gap from presets bundled with the package.
|
|
347
367
|
|
|
348
|
-
|
|
|
368
|
+
| Presets | What they cover | File |
|
|
349
369
|
|---|---|---|
|
|
350
|
-
|
|
|
351
|
-
|
|
|
370
|
+
| 9 model-fitting | running commands, lookup, status checks, questions about pasted logs, read-and-summarize — each with a T2 (haiku) and a T1 (sonnet) rule | `presets/model-rules.json` |
|
|
371
|
+
| 6 ratchet | general-purpose rules promoted from mistakes that actually recurred | `presets/ratchet-rules.json` |
|
|
352
372
|
|
|
353
|
-
|
|
373
|
+
**How they get registered:** in the first session after an install or upgrade, the SessionStart hook hands the pending presets to the model, which walks the user through them **one at a time**. Each answer runs one of these immediately:
|
|
354
374
|
|
|
355
375
|
```bash
|
|
356
|
-
claude-token-saver seed #
|
|
357
|
-
claude-token-saver seed accept <id> --global|--project #
|
|
358
|
-
claude-token-saver seed accept all --global #
|
|
359
|
-
claude-token-saver seed skip <id> #
|
|
360
|
-
claude-token-saver seed reset #
|
|
376
|
+
claude-token-saver seed # pending presets + recorded answers
|
|
377
|
+
claude-token-saver seed accept <id> --global|--project # register one (scope required)
|
|
378
|
+
claude-token-saver seed accept all --global # when the user says "register them all"
|
|
379
|
+
claude-token-saver seed skip <id> # decline — never offered again
|
|
380
|
+
claude-token-saver seed reset # clear the answers and offer everything again
|
|
361
381
|
```
|
|
362
382
|
|
|
363
|
-
-
|
|
364
|
-
-
|
|
365
|
-
-
|
|
366
|
-
-
|
|
383
|
+
- **Nothing is written without a yes to that specific rule.** A declined rule stays declined across upgrades; a later release only surfaces the presets it actually added.
|
|
384
|
+
- A preset is withheld when you already approved a rule of the same shape (same tier and category).
|
|
385
|
+
- A seeded rule **does not pass someone else's statistics off as yours.** It is recorded as `preset (curated)` until a scan measures real firings and delegations, and then those numbers replace it. If its delegated error rate crosses the threshold it gets the same review flag as any other rule.
|
|
386
|
+
- The scope must be stated as `--global` or `--project`. The hook environment is non-TTY, so the CLI cannot ask — the model confirms with the user and passes the flag.
|
|
367
387
|
|
|
368
|
-
## 🇰🇷
|
|
388
|
+
## 🇰🇷 Korean writing guidance
|
|
369
389
|
|
|
370
|
-
|
|
390
|
+
Injects guidance that corrects how Claude writes Korean (dropped sentence parts, noun-stopped sentences, translationese, em-dash overuse) **once per session.**
|
|
371
391
|
|
|
372
392
|
```bash
|
|
373
|
-
claude-token-saver korean on #
|
|
374
|
-
claude-token-saver korean status #
|
|
375
|
-
claude-token-saver korean show #
|
|
376
|
-
claude-token-saver korean off #
|
|
393
|
+
claude-token-saver korean on # on, for every project
|
|
394
|
+
claude-token-saver korean status # state, cost, provenance
|
|
395
|
+
claude-token-saver korean show # print the guidance itself
|
|
396
|
+
claude-token-saver korean off # off
|
|
377
397
|
```
|
|
378
398
|
|
|
379
|
-
Claude Code
|
|
399
|
+
Claude Code's output styles can do the same thing, but an output style is **a single global slot**: turning it on takes that slot away from anything else and has to be configured per machine. This ships the guidance inside the package and delivers it through the SessionStart hook that is already installed, so it **applies wherever the CLI is installed and leaves the output-style slot free.** It survives `/clear`, because the hook fires again.
|
|
380
400
|
|
|
381
|
-
|
|
401
|
+
Cost is **~1.5k tokens per session, injected once at session start rather than per turn**, and covered by the prompt cache from the second request on. When it is on, a `✍️` chip appears in the statusline.
|
|
382
402
|
|
|
383
|
-
###
|
|
403
|
+
### Before and after
|
|
384
404
|
|
|
385
|
-
|
|
405
|
+
From a real pipeline (yaml-sns-agent), where a Slack bot shells out to `claude -p` to investigate the repository and answers in a thread. The guidance was toggled off and then on, with the same question both times: "explain in three Korean sentences how to narrow down a render failure."
|
|
386
406
|
|
|
387
|
-
|
|
407
|
+
**Before** (guidance off)
|
|
388
408
|
|
|
389
409
|
> 렌더 실패 원인 좁히기 3단계: 먼저 `hn:still --check`로 스틸이 명세보다 최신인지(FRESH/STALE) 확인 — STALE이면 재생성 안 해서 옛 자산으로 렌더 시도한 게 원인. 다음 caption-blocks가 최신 나레이션과 단어열 일치하는지 직접 대조 — TTS 재생성 후 blocks 재생성 빠지면 자막·음성 어긋나 렌더 깨짐.
|
|
390
410
|
|
|
391
|
-
|
|
411
|
+
**After** (guidance on)
|
|
392
412
|
|
|
393
413
|
> 먼저 `hn:still --check` 돌려 스틸이 명세보다 최신인지 확인한다. STALE이면 재생성 안 해서 생긴 문제.
|
|
394
414
|
>
|
|
395
415
|
> 다음 caption-blocks가 captions.json 단어열과 일치하는지 본다. 내레이션 재TTS 후 blocks 재생성 빠지면 옛 자막이 새 음성 위에 뜬다.
|
|
396
416
|
|
|
397
|
-
|
|
417
|
+
Three things change. Clauses chained with em dashes become separate sentences, so one sentence carries one fact. Noun-stopped phrases (확인, 대조, 렌더 깨짐 — "check", "compare", "render breaks") become predicates (확인한다, 본다, 뜬다), which makes it explicit that these are steps to take. And the particles come back where they had been dropped, so subject and object are legible on the first read.
|
|
418
|
+
|
|
419
|
+
The technical content is identical in both. The guidance touches sentence construction only, not judgement or accuracy: the answer does not change, it just stops needing a second read. In a channel people scroll through, that difference cuts follow-up questions — and the tokens those follow-ups would have cost.
|
|
398
420
|
|
|
399
|
-
|
|
421
|
+
### The supplement: cohesion and conservative correctness rules
|
|
400
422
|
|
|
401
|
-
|
|
423
|
+
The vendored fluent-korean text ships unmodified; everything collected since lives in a separate supplement (`presets/korean-style/supplement.md`) appended to the same injection. It was compiled conservatively — only clauses that are nearly always an improvement, sourced from the National Institute of Korean Language's public-language guidelines, the Kubernetes Korean localization guide, and three peer-reviewed studies on text cohesion in Korean writing.
|
|
402
424
|
|
|
403
|
-
|
|
425
|
+
It adds three layers:
|
|
404
426
|
|
|
405
|
-
|
|
427
|
+
- **Translationese**: double passives, Japanese-derived calques, `~에 있어서`, possession-verb renderings of English *have*. The machine-checkable ones also run in the write-time lint (below).
|
|
428
|
+
- **AI-writing tics**: automatic intensifiers ("다양한", "핵심적인"), signpost sentences, rhetorical question-then-answer, unconditionally upbeat endings.
|
|
429
|
+
- **Cohesion** — how sentences connect, which no regex can check. The research finding that shapes this section: surface connectives (conjunctions, demonstratives) correlate *negatively or not at all* with judged text quality, while elaboration — the next sentence picking up and unpacking what the previous one introduced — is the only connection type with a positive correlation. So the guidance says: when a transition feels rough, fix the information order (given before new), don't add a connective.
|
|
430
|
+
|
|
431
|
+
**Most of the cohesion layer is not Korean-specific.** Given-before-new ordering (the "given-new contract"), one clear referent per pronoun, keeping one subject per paragraph, bridging sentences instead of leaping, and merging choppy repetitive sentences into a modifier-plus-core structure apply to English prose the same way — the studies happen to be about Korean learners, but the principles they validate are the standard cohesion model from text linguistics. If you write English deliverables with Claude, those five rules are worth pinning in your own CLAUDE.md even with this feature off.
|
|
432
|
+
|
|
433
|
+
A final subsection lists what must **not** be "corrected": settled domain terms, formal register, and verbatim quotations — every lint finding is a request to confirm, not a verdict.
|
|
434
|
+
|
|
435
|
+
### The write-time check (v3.24.0)
|
|
436
|
+
|
|
437
|
+
Injecting the guidance once at session start turned out to be half the job. The model reads it, then writes dozens of files over the next hours with nothing re-reading the output. Sessions with the guidance active still shipped violations into documents, and it surfaced only when a human read the finished artifact. An August 2026 fix reworded the scope sentence to address this; it recurred, because rewording an instruction does not add a checkpoint.
|
|
438
|
+
|
|
439
|
+
From v3.24.0 `korean on` also installs a PostToolUse hook. It opens the file the model just wrote, runs the clauses a machine can decide, and hands any findings back. The file is already saved, so nothing is lost — the model fixes it on the spot.
|
|
406
440
|
|
|
407
441
|
```bash
|
|
408
|
-
claude-token-saver korean lint block #
|
|
409
|
-
claude-token-saver korean lint warn #
|
|
410
|
-
claude-token-saver korean lint off #
|
|
442
|
+
claude-token-saver korean lint block # default: findings are handed back as blocking feedback
|
|
443
|
+
claude-token-saver korean lint warn # print findings, do not block
|
|
444
|
+
claude-token-saver korean lint off # disable the check
|
|
411
445
|
|
|
412
|
-
claude-token-saver korean lint scope all #
|
|
413
|
-
claude-token-saver korean lint scope prose #
|
|
446
|
+
claude-token-saver korean lint scope all # default: every text file the session writes
|
|
447
|
+
claude-token-saver korean lint scope prose # documents only
|
|
414
448
|
|
|
415
|
-
claude-token-saver korean lint docs/*.md #
|
|
449
|
+
claude-token-saver korean lint docs/*.md # check files already on disk
|
|
416
450
|
```
|
|
417
451
|
|
|
418
|
-
|
|
452
|
+
Checked: 15 figurative phrases, translationese markers, separators (`—`·`ㅡ`·`|`), three or more `의` particles in one phrase, and a period after a nominal ending. Clauses that need judgement stay with the guidance text.
|
|
419
453
|
|
|
420
|
-
|
|
454
|
+
The default `all` scope covers code comments, UI strings, subtitles, templates, and build output, not just documents. The vendored guidance exempts comments, but comments are read by people and generated artifacts (PDF, HTML) are assembled from those strings, so exempting them reopens the exact gap that was reported. Only installed dependencies, VCS internals, lockfiles, and binary or image files are skipped; `dist/` and `build/` are checked. `korean lint scope prose` restores the narrow reading.
|
|
421
455
|
|
|
422
|
-
|
|
456
|
+
The scope sentence in the injected guidance is generated from the same setting, so the model is never told one rule while being corrected against another.
|
|
423
457
|
|
|
424
|
-
###
|
|
458
|
+
### The encoding rule that ships with it (v3.23.2)
|
|
425
459
|
|
|
426
|
-
|
|
460
|
+
Alongside the writing guidance, one more line is injected: **non-ASCII strings in tool-call parameters must be written as literal UTF-8, never as `\uXXXX` unicode escapes.**
|
|
427
461
|
|
|
428
|
-
|
|
462
|
+
When the model puts Korean into a Write or Edit parameter as escapes, those escapes are sometimes not decoded into code points at all: the literal text `한` lands in the file. The artifact carries mojibake, and the model keeps editing on top of it without noticing that what it wrote and what the file holds have diverged. Not writing escapes in the first place removes the path entirely, so the rule blocks the input instead of repairing the output.
|
|
429
463
|
|
|
430
|
-
|
|
464
|
+
This line lives in claude-token-saver's own framing paragraph, not in the vendored fluent-korean text. It governs encoding rather than style, and the vendored wording is kept unmodified. For the same reason it carries no exceptions, unlike the style rules that skip code and commit messages. It adds roughly 60 tokens per session.
|
|
431
465
|
|
|
432
|
-
>
|
|
433
|
-
>
|
|
466
|
+
> **Evidence**
|
|
467
|
+
> The same failure is reported against Claude Code: [#12417, unicode handling regression](https://github.com/anthropics/claude-code/issues/12417) and [#26141, Edit silently corrupting unicode](https://github.com/anthropics/claude-code/issues/26141).
|
|
434
468
|
|
|
435
|
-
###
|
|
469
|
+
### Asked at install time
|
|
436
470
|
|
|
437
|
-
|
|
471
|
+
The install **prints what the guidance changes, its per-session cost and its source, then asks.** A Korean system locale (`ko_KR` and friends; on macOS the system setting is checked too) makes the question default to yes; anything else defaults to no, so users who never write Korean are not billed 1.5k tokens a session. The locale is only a default, so an English-locale machine used for Korean work can still turn it on right there.
|
|
438
472
|
|
|
439
|
-
npm
|
|
473
|
+
Installs with nobody attached — npm `postinstall`, CI, piped stdin — skip the question and apply the locale default, because a blocked prompt hangs the install. In that case, if the locale is not Korean the setting is **left undecided rather than recorded**, so a later run at a terminal still gets to ask. Use `--yes` or `--no-input` to force the non-interactive path, or `CTS_NO_KOREAN=1` to skip the feature entirely. **Once you have turned it on or off yourself, that choice sticks — an upgrade never overrides it.**
|
|
440
474
|
|
|
441
|
-
>
|
|
442
|
-
>
|
|
443
|
-
>
|
|
475
|
+
> **Source and license**
|
|
476
|
+
> The guidance text comes from [fluent-korean](https://github.com/snflkd/fluent-korean). Copyright (c) 2026 snflkd, MIT License.
|
|
477
|
+
> The wording is unmodified; only the output-style frontmatter was removed. The full license ships with the package at `presets/korean-style/LICENSE-fluent-korean`.
|
|
444
478
|
|
|
445
|
-
## 📄 doc2md
|
|
479
|
+
## 📄 doc2md — documents become Markdown before the model reads them
|
|
446
480
|
|
|
447
|
-
|
|
481
|
+
`Read` a pptx, xlsx, pdf or docx and the raw bytes go into the context window, where the model cannot read them. This intercepts that `Read`, converts the file once, and hands over the Markdown instead.
|
|
448
482
|
|
|
449
|
-
|
|
483
|
+
**This is opt-in.** Installing the CLI does not turn it on: both commands below are required, and a registered hook with no converter behind it does nothing at all.
|
|
450
484
|
|
|
451
|
-
|
|
485
|
+
Three situations, three different interception points:
|
|
452
486
|
|
|
453
|
-
|
|
454
|
-
|
|
455
|
-
| 상황 | 개입 지점 |
|
|
487
|
+
| Situation | Where it is caught |
|
|
456
488
|
|---|---|
|
|
457
|
-
|
|
|
458
|
-
|
|
|
459
|
-
|
|
|
489
|
+
| A document path typed in the prompt (`@path`, quoted, or relative) | `UserPromptSubmit`: converted, and the conversion's path is handed back as context |
|
|
490
|
+
| A document opened with `Read` mid-task | PDFs are caught by `PreToolUse(Read)`. pptx/xlsx/docx/fig are not: Claude Code refuses them as binary *before* any hook runs, so the session-start note tells the model to run `doc2md <path>` instead |
|
|
491
|
+
| A document attached to the message | **Not catchable.** No hook event receives attachment content. The session-start note has the model ask for a path next time |
|
|
460
492
|
|
|
461
|
-
|
|
493
|
+
That second row is measured, not assumed: a `.pdf` Read fires the hook, and a `.pptx` Read in the same session leaves no hook log entry at all.
|
|
462
494
|
|
|
463
495
|
```bash
|
|
464
|
-
claude-token-saver doc2md on #
|
|
465
|
-
claude-token-saver doc2md #
|
|
466
|
-
claude-token-saver doc2md
|
|
467
|
-
claude-token-saver doc2md install-converter #
|
|
496
|
+
claude-token-saver doc2md on # register the hooks (the converter installs itself)
|
|
497
|
+
claude-token-saver doc2md # check converter + hook registration
|
|
498
|
+
claude-token-saver doc2md report.pptx # convert by hand and see the result
|
|
499
|
+
claude-token-saver doc2md install-converter # only to get the install out of the way early
|
|
468
500
|
```
|
|
469
501
|
|
|
470
|
-
|
|
502
|
+
**The converter installs itself.** Any rollout step a person has to be told about is a step some of them skip, so the converter installs in the background the moment a document first shows up, and converts as soon as it is ready. Measured: about 30s for the first document (15s install plus markitdown's first import), then 3.7s for a new document and 0.1s on a cache hit. The `.fig` parser installs in half a second on the first Figma file.
|
|
503
|
+
|
|
504
|
+
It installs on first use rather than at `install` time: the venv is 47MB, and someone who never opens a document should not pay for it. Set `CTS_DOC2MD_NO_AUTOINSTALL=1` to turn the automatic install off.
|
|
505
|
+
|
|
506
|
+
**Python 3.10+ is required** — markitdown's own floor, and macOS still ships 3.9 as `/usr/bin/python3`. The venv is built on an interpreter chosen by version rather than by PATH order. Built on 3.9, pip resolves markitdown to a 2019 placeholder release (0.0.1a1): the install looks like it worked and every conversion then dies at import. This was found by walking into it. When nothing on the machine is new enough, the message points at `brew install python` instead of at an install command that cannot succeed.
|
|
507
|
+
|
|
508
|
+
The converter goes into a venv this tool owns (`<state dir>/doc2md-venv`): no system interpreter is touched, and uninstalling the CLI takes it along. An existing markitdown on `uv tool` or `PATH` is preferred over building a new one.
|
|
509
|
+
|
|
510
|
+
Conversion is [markitdown](https://github.com/microsoft/markitdown). Slide numbers, heading levels, tables, speaker notes and per-sheet headings all survive, and non-Latin text comes through intact.
|
|
511
|
+
|
|
512
|
+
Several things it deliberately does not do:
|
|
513
|
+
|
|
514
|
+
- **Images are not converted.** markitdown returns nothing for them, and OCR misread resource names in testing (`c5.xlarge` as `c.xlarge`). In a document where those names *are* the content, wrong text is worse than none. The model reads images natively anyway.
|
|
515
|
+
- **A missing converter never fails silently.** The install command is shown once, then the original `Read` proceeds untouched. Repeating the notice on every read would be its own nuisance; saying nothing is how a broken converter hides. Run `doc2md` with no arguments to see the converter and hook registration together.
|
|
516
|
+
- **Conversions never land in your project.** They go under the tool's own state directory with mode `0700`, so there is nothing to add to `.gitignore`. Filenames matching payroll/contract/secret patterns are skipped entirely.
|
|
517
|
+
- **Zip bombs are refused.** pptx/xlsx/docx are zip containers: the declared sizes are checked first, and since those are written by whoever built the file, the real decompressed bytes are counted against a ceiling too.
|
|
518
|
+
- **Spreadsheets are capped by rows, not bytes.** Conversion time tracks row count (measured: a 6.3MB PDF in 0.9s, a 5.8MB workbook in 47.75s). Past 50,000 rows only the head is converted, and **the truncation and the true row count are both stated** in what the model is told.
|
|
519
|
+
|
|
520
|
+
### What a conversion saves
|
|
521
|
+
|
|
522
|
+
Every conversion is stamped with a provenance header: which original, when, how many tokens. Savings show up on the statusline's own `📄 Doc2md saved` line.
|
|
523
|
+
|
|
524
|
+
The baseline is what you would have done without a converter, and that differs by format. Both were measured on 2026-09-06.
|
|
525
|
+
|
|
526
|
+
**PDF is priced against attaching it.** The same one-line prompt was sent through `claude --print --input-format stream-json` with and without the file as a document block. The control turn cost 42,204 tokens, twice, to the token.
|
|
527
|
+
|
|
528
|
+
| Attached file | Size | Extra tokens | Per page |
|
|
529
|
+
|---|---|---|---|
|
|
530
|
+
| Résumé PDF | 7 pages | +20,537 | 2,934 |
|
|
531
|
+
| Résumé PDF | 5 pages | +12,709 | 2,542 |
|
|
471
532
|
|
|
472
|
-
|
|
533
|
+
An attached PDF is read whole, but every page costs 2,500–2,900 tokens against 5,531 for the conversion. The coefficient used is 2,500 per page — below both measurements, so the figure understates rather than flatters.
|
|
473
534
|
|
|
474
|
-
|
|
535
|
+
**pptx/xlsx/docx are priced against unpacking the container.** These never reach the model as attachments at all: the same probe on a docx added 78 tokens and the model replied that it had no file, and `Read` refuses the format outright. What you actually do without a converter is unzip the archive and read its XML, where tags and style attributes outweigh the words.
|
|
475
536
|
|
|
476
|
-
|
|
537
|
+
| Original | Body XML | Conversion | Ratio |
|
|
538
|
+
|---|---|---|---|
|
|
539
|
+
| Deck, pptx (31.8MB) | ~540,429 tokens | ~22,610 tokens | 23.8× |
|
|
540
|
+
| Résumé, docx (189KB) | ~79,621 tokens | ~1,684 tokens | 47.3× |
|
|
477
541
|
|
|
478
|
-
|
|
542
|
+
This baseline is measured per file from the real XML size, not applied as a per-format ratio. `.xls` is not a zip container and has no markup to measure, so it claims nothing.
|
|
479
543
|
|
|
480
|
-
|
|
544
|
+
### Figma `.fig` converts too
|
|
481
545
|
|
|
482
|
-
|
|
483
|
-
- **변환기가 없으면 조용히 실패하지 않습니다.** 설치 명령을 한 번 안내한 뒤 원본 `Read` 를 그대로 통과시킵니다. 매번 알리면 그것대로 방해가 되고, 아무 말도 하지 않으면 고장을 숨기게 됩니다. `doc2md` 를 인자 없이 실행하면 변환기와 훅 등록 상태를 한 번에 확인할 수 있습니다.
|
|
484
|
-
- **변환본은 프로젝트 안에 남기지 않습니다.** 도구의 상태 디렉터리 아래 권한 `0700` 으로 저장하므로 `.gitignore` 에 무엇을 추가할 필요가 없습니다. 파일명이 급여·계약·개인정보 같은 패턴에 걸리면 아예 변환하지 않습니다.
|
|
485
|
-
- **압축 폭탄은 막습니다.** pptx·xlsx·docx 는 zip 컨테이너입니다. 선언된 크기를 먼저 걸러 내고, 선언은 조작될 수 있으므로 실제 해제 바이트도 상한과 대조합니다.
|
|
486
|
-
- **엑셀은 행 수로 자릅니다.** 변환 시간은 파일 크기가 아니라 행 수를 따릅니다(실측: PDF 6.3MB 0.9초, 엑셀 5.8MB 47.75초). 5만 행을 넘으면 앞부분만 변환하고, **잘랐다는 사실과 전체 행 수를 안내에 함께 적습니다.**
|
|
546
|
+
Planning documents are moving from PowerPoint to Figma, so the same hook catches `.fig`. A `.fig` is a zip, but the `canvas.fig` inside it is Figma's private binary (kiwi format), which markitdown cannot open — so this one format is converted in Node with [openfig-core](https://github.com/OpenFig-org/openfig-core) (MIT). `doc2md install-converter` places it beside markitdown in the tool's state directory; the package itself still ships zero dependencies.
|
|
487
547
|
|
|
488
|
-
|
|
548
|
+
The result is an outline: pages and frames become headings, text nodes become body lines, and shapes are counted rather than listed — in a planning document the words are the content, and two hundred `Rectangle 173` lines would drown them. A file with no text at all is refused rather than dressed up as an empty document.
|
|
489
549
|
|
|
490
|
-
|
|
550
|
+
Verified against real files: a community Bootstrap UI kit (8.1MB, 4,155 nodes, 1,312 of them text) and a 52MB Tailwind kit, each converting in under a second. Both `.fig` vintages parse — the current zip container and the older bare fig-kiwi stream.
|
|
491
551
|
|
|
492
|
-
|
|
493
|
-
- `.fig` 절감이 가장 큽니다. `Read` 가 이진 파일을 거부하지 않고 그대로 읽어 회당 약 44,000 토큰을 태우기 때문입니다.
|
|
494
|
-
- 암호 문서와 DRM 문서는 오류가 아니라 상태로 판별해 안내합니다. 원본 `Read` 를 막지 않으므로 작업이 중단되지 않습니다.
|
|
552
|
+
**`.fig` saves the most of any format.** Unlike the Office containers, `Read` does not refuse a `.fig`: the extension means nothing to it, so it pulls the binary in as text and the context window fills with tokenised noise. Measured against the same 42,760-token control:
|
|
495
553
|
|
|
496
|
-
|
|
554
|
+
| File | Size | Extra tokens for a Read | Conversion |
|
|
555
|
+
|---|---|---|---|
|
|
556
|
+
| plan.fig | 26KB | +44,195 | 100 tokens |
|
|
557
|
+
| bootstrap-kit.fig | 8.1MB | +43,994 | 18,397 tokens |
|
|
497
558
|
|
|
498
|
-
|
|
559
|
+
Two files three hundred times apart in size cost the same, because Read truncates long before the file ends — you pay for a whole document and receive a fraction of one. The baseline is therefore a flat 44,000 tokens. For comparison, the same probe on a pptx cost +317 tokens and on a docx +185: a refusal message, and nothing else.
|
|
499
560
|
|
|
500
|
-
|
|
561
|
+
#### Why the baseline does not scale with file size
|
|
501
562
|
|
|
502
|
-
|
|
503
|
-
- 5분 버킷에서는 카운트다운 색이 비율이 아니라 절대 시간을 따릅니다. 5분의 30%는 90초여서, 초록이 주는 여유가 실제와 어긋났습니다.
|
|
504
|
-
- `⚠ 5m TTL` 경고가 이 환경에도 도달합니다. 다만 조언 문구는 다릅니다. 구독 플랜을 바꿔도 해소되지 않는 환경이므로 플랜 전환을 권하지 않습니다.
|
|
505
|
-
- `Extra cost if 5m-only` 는 1시간 쓰기가 있는 경우에만 묻습니다. 이미 5분 전용인 환경에서는 질문 자체가 성립하지 않아 `+$0` 이 잘못 읽혔습니다.
|
|
506
|
-
- 위임 건이 모델 ID 해석 실패로 버려졌으면 statusline 에 `🔀 N unresolved` 로 알립니다. 이전에는 "위임한 적 없음"과 화면상 구별되지 않았습니다.
|
|
507
|
-
- 환경변수를 `foundation-model` ARN 으로 지정한 경우에도 모델을 해석합니다. 이름을 담고 있지 않은 `application-inference-profile` ID 는 그대로 거부합니다. 값을 추측해 넣으면 원장에 틀린 금액이 들어가기 때문입니다.
|
|
563
|
+
A baseline has to be what would actually have been spent without the converter. Intuition says a bigger file burns more, but the `Read` tool has a cap (2,000 lines by default, plus a per-line character limit), and a binary file hits it almost immediately: even the 26KB file was already truncated, which is why two files 300× apart came out 201 tokens apart. Had the 8.1MB file gone in whole it would have been millions of tokens — money nobody could have spent, since it does not fit in a 200k context window. Claiming to have saved unspendable money is flattery, not measurement.
|
|
508
564
|
|
|
509
|
-
|
|
565
|
+
The same principle runs through every baseline here:
|
|
510
566
|
|
|
511
|
-
|
|
567
|
+
- **`.fig`, flat 44,000** — set below both measurements (44,195 and 43,994). A model could burn size-proportional tokens by re-Reading at successive offsets, but one Read is what a sane agent does once the bytes turn out to be binary noise, so one Read is the honest counterfactual.
|
|
568
|
+
- **PDF, 2,500 per page** — below both measured values (2,542 and 2,934).
|
|
569
|
+
- **Office formats, the file's actual XML size** — the one case where proportional is right, because a person really does end up reading that XML; it is measured per file rather than applied as a ratio.
|
|
512
570
|
|
|
513
|
-
|
|
571
|
+
The common rule: wherever an estimate and a measurement diverge, the lower number wins. A figure the user can trust is worth more than one that flatters the tool.
|
|
514
572
|
|
|
515
|
-
|
|
516
|
-
- 조회는 LiteLLM 의 `GET /key/info` 와 `GET /user/info` 로 하고, 호출 키 자신의 정보만 받습니다. 예산 출처는 실무에서 가장 많이 쓰는 **팀 멤버십 예산**(team_memberships 의 spend·max_budget)을 먼저 보고, 없으면 키 자체의 max_budget, 그다음 internal user 예산 순으로 고릅니다. 렌더는 캐시 파일만 읽으며, 갱신은 5분에 한 번 분리된 백그라운드 프로세스가 수행합니다 (update-check 와 같은 구조라 statusline 이 네트워크를 기다리지 않습니다).
|
|
517
|
-
- `max_budget` 이 없는 무제한 키는 게이지를 만들지 않습니다. 이 경우에도 `💵` 월 지출 세그먼트는 세션 로그 기반이라 그대로 표시됩니다.
|
|
518
|
-
- 상태 확인: `claude-token-saver litellm-budget` (캐시 출력) · `litellm-budget --refresh` (즉시 조회).
|
|
573
|
+
### Editing a document: copy, then script
|
|
519
574
|
|
|
520
|
-
|
|
575
|
+
Conversion is one-way — editing the cached `.md` changes nothing in the source. The hook refuses `Edit`/`Write` on both the cache and the original binary, and points at the right path instead: copy the original, edit the copy with a script, re-convert the copy to verify.
|
|
521
576
|
|
|
522
|
-
|
|
577
|
+
`install-converter` puts the editing libraries (python-pptx, python-docx, openpyxl) in the same venv, so a structural request like "swap the chart on slide 23 for a line chart" is a short script the agent writes on the spot. `.fig` edits go through openfig-core, which encodes as well as parses.
|
|
523
578
|
|
|
524
|
-
|
|
579
|
+
All four formats were exercised end to end on 2026-09-06: 10 docx run replacements plus three consecutive re-saves, a pptx bar-to-line chart swap with an added data point, xlsx value edits and a new row, and a fig text edit with re-encode and re-parse. In every case the original was byte-identical afterwards and the re-converted copy showed the change. One caveat: removing a chart shape from a pptx leaves the old chart XML part orphaned — PowerPoint ignores it, but delete the part and its rels for a clean file. Charts and images never appear in a conversion, so visual edits must be confirmed in the application itself.
|
|
580
|
+
|
|
581
|
+
### DRM-wrapped documents
|
|
582
|
+
|
|
583
|
+
Encryption and DRM are different problems with different answers. Enterprise DRM (Fasoo, MarkAny, SoftCamp and the like) does not password a document — it wraps the whole file, and only processes the vendor's agent has whitelisted ever see plaintext. Python is not one of them, so what sits on disk is ciphertext behind a vendor header, and **no password will open it.**
|
|
584
|
+
|
|
585
|
+
The first bytes decide which story to tell: a zip header means a truncated download, an OLE container means a password, and neither means the file is not that format at all.
|
|
586
|
+
|
|
587
|
+
```
|
|
588
|
+
✗ bad-archive: File is not a zip file → download it again
|
|
589
|
+
✗ encrypted: password-protected Office file → ask for an unlocked copy
|
|
590
|
+
✗ drm-protected: DRM-wrapped file (FASOO) → ask for a copy released from DRM
|
|
591
|
+
```
|
|
592
|
+
|
|
593
|
+
Vendor names are matched only to say which client to go to; the classification stands without recognising the vendor. PDFs are judged the same way through their public DRM security-handler names (FOPN_foweb, EBX_HANDLER, Adobe.APS).
|
|
594
|
+
|
|
595
|
+
### Locked documents, and Windows
|
|
596
|
+
|
|
597
|
+
**A password-protected document is a state, not an error.** Office encrypts by wrapping the package in an OLE compound file rather than a zip, so opening one as a zip used to report "not a zip file" — which reads as a broken download and sends the user after the wrong problem. It is now identified before conversion:
|
|
598
|
+
|
|
599
|
+
```
|
|
600
|
+
✗ encrypted: password-protected Office file (OLE-wrapped)
|
|
601
|
+
✗ encrypted: password-protected PDF
|
|
602
|
+
```
|
|
603
|
+
|
|
604
|
+
The model is told to ask for an unlocked copy. This tool never asks for or stores a password, and never blocks the original `Read`, so work continues either way. A PDF that merely restricts printing still opens and still converts — checked against a false positive — and a legacy `.xls`, which is an OLE file by design, is not mistaken for an encrypted one.
|
|
605
|
+
|
|
606
|
+
**Windows is supported.** For teams with Windows machines:
|
|
607
|
+
|
|
608
|
+
- The Python search uses the `py -3` launcher. `python3` is rarely on PATH there, and a bare `python` may be the Store alias stub that opens a web page instead of running anything. Venv interpreters are looked for at `Scripts\python.exe`.
|
|
609
|
+
- The `.fig` parser installs through `npm.cmd` via the shell, and the package spec dropped its caret (`openfig-core@0.4.x`): in cmd.exe `^` is the escape character and never reaches npm.
|
|
610
|
+
- The background install and every child process set `windowsHide`, so no console window appears in the middle of someone's prompt.
|
|
611
|
+
|
|
612
|
+
`claude-token-saver doc2md --clean` empties the conversion cache; `doc2md off` removes the hook. Removal filters for this tool's own entry, so anything else you registered under `PreToolUse` stays.
|
|
613
|
+
|
|
614
|
+
## 🌐 Behind a gateway (Bedrock / Vertex)
|
|
615
|
+
|
|
616
|
+
A gateway reports the cache-creation total but never the 5m/1h split. That left the tool unable to tell "nothing cached yet" from "this provider does not say", and the fallback assumed an hour — for a window that is really five minutes on Bedrock, overstating it twelvefold.
|
|
617
|
+
|
|
618
|
+
Since v3.26.0 the gateway is detected from the model ids in the transcript, which fixes:
|
|
619
|
+
|
|
620
|
+
- The countdown falls back to 5 minutes, labelled `5m?`. Three grades of certainty get three labels: measured (`5m`), inferred (`5m?`), unknown (`?`).
|
|
621
|
+
- In a 5-minute bucket the countdown colour follows absolute time rather than a percentage. 30% of five minutes is 90 seconds, and green there promised comfort that was not there.
|
|
622
|
+
- The `⚠ 5m TTL` warning finally reaches these users — with different advice, since no subscription plan changes a gateway's TTL.
|
|
623
|
+
- `Extra cost if 5m-only` is only asked of sessions that have 1h writes to lose. Elsewhere the arithmetically honest `+$0` read as an endorsement of the bucket you are already stuck in.
|
|
624
|
+
- Delegated runs dropped for an unpriceable model id show as `🔀 N unresolved` instead of nothing, which used to be indistinguishable from never having delegated.
|
|
625
|
+
- Environment variables set to a `foundation-model` ARN now resolve. An opaque `application-inference-profile` id still does not: guessing at it is how wrong prices enter the ledger.
|
|
626
|
+
|
|
627
|
+
If the detection is wrong, pin it with `claude-token-saver mode ttl=5m` (or `ttl=1h`). An explicit value outranks the measurement.
|
|
628
|
+
|
|
629
|
+
### LiteLLM: your key budget stands in for the missing 5h/7d caps (v3.35.0)
|
|
630
|
+
|
|
631
|
+
Behind a LiteLLM proxy (Bedrock and friends), Claude Code's stdin never carries `rate_limits`, so the `✦ current` / `📅 weekly` gauges simply do not exist. LiteLLM does track per-key budgets, so the statusline draws a budget gauge in their place.
|
|
632
|
+
|
|
633
|
+
- Detection: `ANTHROPIC_BASE_URL` points somewhere other than the official endpoint, and `ANTHROPIC_AUTH_TOKEN` (or `ANTHROPIC_API_KEY`) is set.
|
|
634
|
+
- The proxy is asked via `GET /key/info` and `GET /user/info` — only the calling key's own data. Renders read a cache file; a detached background process refreshes it every 5 minutes (same shape as the update check), so the statusline never waits on the network.
|
|
635
|
+
- Budget source priority follows real-world usage: the **team-membership budget** (`team_memberships[].spend` + its linked budget table row) first, then the key's own `max_budget`, then the internal-user budget. Verified against a Dockerized LiteLLM, including memberships whose budget diverges from the team max into a separate budget-table row.
|
|
636
|
+
- Unlimited keys (no `max_budget`) get no gauge. The `💵` monthly-spend segment still shows, since it comes from session logs.
|
|
637
|
+
- Inspect with `claude-token-saver litellm-budget` (cached) or `litellm-budget --refresh` (query now).
|
|
638
|
+
|
|
639
|
+
One related non-bug: if your session model is already sonnet, a sonnet-delegation (T1) rule can never save anything, because there is no price gap to capture. That is correct, but `route-scan rules` displayed it identically to "no delegations yet", so it now says outright that the rule does not apply at the current default model.
|
|
640
|
+
|
|
641
|
+
## Spike issue codes
|
|
642
|
+
|
|
643
|
+
| Code | Meaning |
|
|
525
644
|
|---|---|
|
|
526
|
-
| `LARGE_INPUT_PER_REQUEST` |
|
|
527
|
-
| `LOW_HIT_RATE` |
|
|
528
|
-
| `BUCKET_5M_DOMINANT` |
|
|
529
|
-
| `HIGH_OUTPUT_RATIO` |
|
|
530
|
-
| `HIGH_REQUEST_COUNT` |
|
|
531
|
-
| `FREQUENT_CACHE_REBUILD` |
|
|
645
|
+
| `LARGE_INPUT_PER_REQUEST` | single request > 200k input tokens — per-turn re-billing and cap burn spike |
|
|
646
|
+
| `LOW_HIT_RATE` | cache hit rate < 50% |
|
|
647
|
+
| `BUCKET_5M_DOMINANT` | > 70% of cache writes hit the 5m bucket |
|
|
648
|
+
| `HIGH_OUTPUT_RATIO` | output/input > 0.15 (output is 5× input price) |
|
|
649
|
+
| `HIGH_REQUEST_COUNT` | session made 3×+ your median (tool loop?) |
|
|
650
|
+
| `FREQUENT_CACHE_REBUILD` | `cache_creation` > `cache_read` |
|
|
532
651
|
|
|
533
|
-
|
|
652
|
+
Remediation commands are OS-aware (`~/.zshrc` for macOS/Linux/WSL, `setx` for Windows).
|
|
534
653
|
|
|
535
|
-
##
|
|
654
|
+
## Real-world impact — before/after report
|
|
536
655
|
|
|
537
|
-

|
|
538
657
|
|
|
539
|
-
harness 5/5 + ratchet
|
|
658
|
+
harness 5/5 + ratchet applied to the author's own Claude Code work, normalized **per user message** (cutoff 2026-05-02, Opus 4.7 pricing):
|
|
540
659
|
|
|
541
|
-
|
|
|
660
|
+
| metric | before (7d / 739 msgs) | after (2d / 157 msgs) | Δ |
|
|
542
661
|
|---|---:|---:|---:|
|
|
543
|
-
|
|
|
544
|
-
|
|
|
545
|
-
|
|
|
546
|
-
|
|
|
662
|
+
| cost / user message | $2.345 | $1.910 | **−18.6%** |
|
|
663
|
+
| output tokens / message | 7,391 | 6,052 | −18.1% |
|
|
664
|
+
| assistant turns / message | 9.73 | 8.83 | −9.2% |
|
|
665
|
+
| tool calls / message | 5.72 | 5.25 | −8.2% |
|
|
547
666
|
|
|
548
|
-
|
|
667
|
+
Same request resolved in fewer round-trips → first-try success rate up — the effect of PEV + Structured Task forcing one-shot delivery.
|
|
549
668
|
|
|
550
669
|
<details>
|
|
551
|
-
<summary
|
|
670
|
+
<summary>Measurement notes — why cache hit rate isn't included · sample caveats</summary>
|
|
552
671
|
|
|
553
|
-
-
|
|
554
|
-
-
|
|
555
|
-
- ⚠️
|
|
672
|
+
- The author is on the Max plan (1-hour cache TTL) with hit rate already converged near ~98%, so little headroom there. **Pro-plan users (5-minute TTL)** likely see hit rate itself rise with the handoff-before-expiry workflow.
|
|
673
|
+
- Handoff-before-expiry: watch the TTL countdown, run `claude-token-saver handoff` just before expiry to dump work state into a markdown brief, start a fresh cache cycle. Same flow handles the 1M warning and cap chips.
|
|
674
|
+
- ⚠️ POST window is only 2 days (157 msgs); statistical confidence is low, and week-to-week topic mix differs, so the tool effect isn't cleanly isolated.
|
|
556
675
|
</details>
|
|
557
676
|
|
|
558
|
-
##
|
|
677
|
+
## Pricing (Jul 2026)
|
|
559
678
|
|
|
560
|
-
|
|
679
|
+
Per million tokens (USD), as used by the cost estimator:
|
|
561
680
|
|
|
562
|
-
|
|
681
|
+
| Tier | Models | Input | 5m Write | 1h Write | Read | Output |
|
|
682
|
+
|---|---|---|---|---|---|---|
|
|
683
|
+
| `claude-fable-5` | Fable 5 / Mythos 5 | $10 | $12.50 | $20 | $1 | $50 |
|
|
684
|
+
| `claude-opus-new` | Opus 4.5 / 4.6 / 4.7 / 4.8 | $5 | $6.25 | $10 | $0.50 | $25 |
|
|
685
|
+
| `claude-opus-legacy` | Opus 4 / 4.1 / 3 | $15 | $18.75 | $30 | $1.50 | $75 |
|
|
686
|
+
| `claude-sonnet` | Sonnet 3.7 / 4 / 4.5 / 4.6 / 5 | $3 | $3.75 | $6 | $0.30 | $15 |
|
|
687
|
+
| `claude-haiku-4-5` | Haiku 4.5 | $1 | $1.25 | $2 | $0.10 | $5 |
|
|
688
|
+
|
|
689
|
+
Source: [Anthropic pricing docs](https://platform.claude.com/docs/en/about-claude/pricing). Sonnet 5 has an introductory $2/$10 rate through 2026-08-31; the estimator uses the standard sticker. Versions ≤ 2.16.x priced Fable 5 at the Sonnet tier (~3× under-estimate) — upgrade to 2.17.0+.
|
|
690
|
+
|
|
691
|
+
### Cache TTL by plan
|
|
692
|
+
|
|
693
|
+
| Plan | TTL | Controlled by |
|
|
694
|
+
|---|---|---|
|
|
695
|
+
| Max ($100–200/mo) | **1h auto** | `tengu_prompt_cache_1h_config` flag |
|
|
696
|
+
| Pro ($20/mo) | **5m fixed** | not configurable |
|
|
697
|
+
| API key | 5m default (1h via beta header) | `cache_control.ttl` |
|
|
698
|
+
|
|
699
|
+
## How it works · Environment
|
|
700
|
+
|
|
701
|
+
Claude Code logs every API call to `~/.claude/projects/<dir>/<session>.jsonl`. This tool dedupes streaming chunks by `requestId` and aggregates `cache_read_input_tokens` / `cache_creation.ephemeral_5m/1h_input_tokens` by day and session.
|
|
702
|
+
|
|
703
|
+
Node.js ≥ 18 · macOS / Linux / Windows / WSL · **zero dependencies**.
|
|
563
704
|
|
|
564
705
|
<details>
|
|
565
|
-
<summary
|
|
706
|
+
<summary>Known quirks · Migration · Background</summary>
|
|
566
707
|
|
|
567
|
-
**IntelliJ Claude Code plugin
|
|
708
|
+
**IntelliJ Claude Code plugin** — the statusline widget fuses frames at the character level when emoji are present (`59:548` artifacts). v2.8.5+ detects `TERMINAL_EMULATOR=JetBrains-JediTerm` and falls back to text mode automatically.
|
|
568
709
|
|
|
569
|
-
**
|
|
710
|
+
**If the countdown looks frozen:** ticking while idle requires Claude Code to re-run the statusline command on a timer, controlled by `statusLine.refreshInterval` (seconds, Claude Code v2.1.97+) in `~/.claude/settings.json`. Without it the line only redraws when the conversation updates. If behavior differs per terminal, check three things: ① that machine's Claude Code is ≥ 2.1.97; ② no project `.claude/settings.json` / `settings.local.json` overrides `statusLine` without a refreshInterval; ③ the statusline wrapper actually finds `claude-token-saver` on PATH instead of falling back to a multi-second `npx` run on every render (typical when nvm is not loaded in non-login shells). Re-running `claude-token-saver install` restores refreshInterval=5.
|
|
570
711
|
|
|
571
|
-
**claude-cache-monitor
|
|
712
|
+
**Migration from claude-cache-monitor:**
|
|
572
713
|
```bash
|
|
573
714
|
npm uninstall -g claude-cache-monitor && npm i -g claude-token-saver
|
|
574
715
|
```
|
|
575
|
-
`~/.claude/settings.json
|
|
716
|
+
Also update `statusLine.command` in `~/.claude/settings.json` to `claude-token-saver …`.
|
|
717
|
+
|
|
718
|
+
**Background:** [GitHub Issue #46829](https://github.com/anthropics/claude-code/issues/46829) (cache TTL regression) · [HN discussion](https://news.ycombinator.com/item?id=47736476)
|
|
576
719
|
</details>
|
|
577
720
|
|
|
578
|
-
##
|
|
721
|
+
## Release notes
|
|
579
722
|
|
|
580
|
-
|
|
723
|
+
The full history moved to [CHANGELOG.md](./CHANGELOG.md) (Korean; version headings and command names are language-neutral). Recent changes:
|
|
581
724
|
|
|
582
|
-
- **v3.
|
|
583
|
-
- **v3.
|
|
725
|
+
- **v3.37.0**: Korean guidance grows a conservative supplement (translationese, AI-writing tics, a research-backed cohesion section whose principles apply to English prose too) and the write-time lint gains 5 translationese patterns, validated at 1 false positive across 255 real files.
|
|
726
|
+
- **v3.35.0**: A `💵 Sep $42` segment now shows estimated spend since 00:00 on the 1st of the current month, always on — including gateway setups with no 5h/7d caps. LiteLLM gateway users get a `🔑 budget ▰▱ 34% $34/$100` gauge built from the key's budget (`GET /key/info` + `GET /user/info`, team-membership budget first, then key, then internal user — verified against a Dockerized LiteLLM).
|
|
727
|
+
- **v3.34.0**: seed presets offered one at a time, output-language choice at install, context warning raised to 500k.
|
|
584
728
|
|
|
585
|
-
##
|
|
729
|
+
## License
|
|
586
730
|
|
|
587
731
|
MIT
|
|
588
732
|
|
|
589
733
|
---
|
|
590
734
|
|
|
591
|
-
##
|
|
735
|
+
## Who makes this
|
|
592
736
|
|
|
593
737
|
[](https://www.youtube.com/@DeepPulseKR)
|
|
594
738
|
[](https://www.youtube.com/@DeepPulseEN)
|
|
595
739
|
[](https://rootstudioyaml.github.io/)
|
|
596
740
|
|
|
597
|
-
|
|
741
|
+
Built and used at **DeepPulse**, a channel about AI developer tooling. The [launch Short (60s)](https://www.youtube.com/shorts/RaD8qMsPTnA) covers where this came from and how it is used.
|