claude-usage-limits 1.26.0 → 1.39.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1,7 +1,7 @@
1
1
  {
2
2
  "name": "usage-limits",
3
3
  "displayName": "Usage Limits",
4
- "version": "1.26.0",
4
+ "version": "1.39.1",
5
5
  "description": "Puts your remaining Claude Code usage limit into Claude's context before every prompt, so it opens with what fits in the budget instead of starting work that gets cut off. Reports headroom as turns rather than percentages, prices a job before you start it, and detects your plan tier.",
6
6
  "author": {
7
7
  "name": "Ridelink",
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "usage-limits",
3
- "version": "1.26.0",
3
+ "version": "1.39.1",
4
4
  "description": "Reports how much of your Codex usage limit is left as turns of work rather than a percentage, prices a job before you start it, and counts the other agents sharing the same budget.",
5
5
  "author": {
6
6
  "name": "Ridelink",
package/README.md CHANGED
@@ -19,6 +19,48 @@ opens with the answer instead:
19
19
 
20
20
  Nobody read a chart to get that. The numbers reached the model, not you.
21
21
 
22
+ ## What it runs on your machine, and what it never does
23
+
24
+ Installing this plugin registers six Claude Code hooks, which means Node runs on
25
+ your machine at six moments: `UserPromptSubmit` (the budget line), `PreToolUse`
26
+ (the fan-out line, and the refusal when a cap is set), `PostToolUse` and
27
+ `SubagentStop` (keeping the reading fresh), and `Stop` and `SessionEnd` (the
28
+ session tally). Nothing runs on a timer except a relay wake you armed yourself.
29
+
30
+ What it reads: your Claude Code transcripts under the config directory, to price
31
+ what has been spent; the account snapshot the host already fetched; and your
32
+ settings files. What it writes: its own files beside those, every one prefixed
33
+ `usage-limits-`, plus the launcher and scheduled task a relay needs while one
34
+ is armed.
35
+
36
+ **It does read your login.** The live reading is the same call Claude Code
37
+ makes for `/usage`, and making it needs the same OAuth token, so this plugin
38
+ reads it from `.credentials.json` (or, on macOS, the `Claude Code-credentials`
39
+ keychain item) and sends it as a bearer token to
40
+ `https://api.anthropic.com/api/oauth/usage`. That token is never written to a
41
+ file, never logged, and never sent anywhere else: the destination is checked
42
+ against an allowlist first - Anthropic over https, or loopback for the tests -
43
+ because a project settings file can put anything in a hook's environment and
44
+ the override that points at a different endpoint must not become a way to walk
45
+ off with your login. Without a live reading the plugin still works, from the
46
+ snapshot and your transcripts.
47
+
48
+ Apart from that one call to Anthropic, nothing leaves the machine: no
49
+ telemetry, no upload, no record of your usage kept anywhere else. Everything it
50
+ knows sits in files you can open, and `node bin/cli.js mode off` stops it
51
+ injecting anything at all.
52
+
53
+ ## Do not want any of this?
54
+
55
+ node bin/cli.js mode off nothing injected, ever - including at the wall
56
+ node bin/cli.js mode off --guard 95 silent, except one short line at 95% used
57
+ node bin/cli.js statusline off take the bars out from under the prompt
58
+
59
+ `off` means off: no budget line, no end-of-reply cost, no panel animation. Most
60
+ people who want quiet actually want the second form - silence until it matters.
61
+ Both apply to new prompts immediately and survive restarts. `mode standard`
62
+ turns it back on.
63
+
22
64
  ## How this differs from a usage dashboard
23
65
 
24
66
  There are a lot of good tools that read the same local files this does and draw
@@ -333,6 +375,57 @@ What changes while it is armed is what Claude is told at the wall. Instead of
333
375
  cut off now costs the wait rather than the work, and that the continuation is a
334
376
  prompt to be acted on rather than a summary for a person to read.
335
377
 
378
+ **A session that started before an update keeps the old code.** A plugin
379
+ update applies when Claude Code restarts, so two sessions can run two versions
380
+ at once; on 2026-09-20 a session still on 1.34.1 saw the single relay slot that
381
+ version had, took it for occupied, scheduled its own wake by hand, and that
382
+ wake failed the way 1.34.1's always did, while the session on 1.36.0 resumed
383
+ fine. Nothing can upgrade a running session, so since 1.38.0 the budget line
384
+ says so once: the installed version, the running one, and that a relay or cap
385
+ set in that session follows the older rules until it restarts.
386
+
387
+ **Two things that stopped it in practice, both fixed in 1.36.0.** A fresh
388
+ Claude Code start asks two questions nobody is there to answer at four in the
389
+ morning: whether to trust the folder, and whether bypass permissions is meant.
390
+ On 2026-09-20 the wake opened its window and sat at the trust question until
391
+ morning. Arming now writes both answers into .claude.json for the folder the
392
+ relay will resume in (every spelling of the key, with a backup beside it), the
393
+ wake writes them again right before the launch, and doctor reports when they
394
+ are missing; relay preflight [cwd] does it by hand. And the relay used to hold
395
+ one slot, so a second session arming on the same night was refused or
396
+ displaced the first. Each session now has its own record and its own scheduled
397
+ task; status lists them all, relay cancel takes down this session's, and
398
+ relay cancel --all takes down every one.
399
+
400
+ **1.37.0 fixed the arm that failed twice on 2026-09-20.** `relay arm` reported
401
+ `Value for '/TR' option cannot be more than 261 character(s)` - schtasks refusing
402
+ the task action, which carried node's path, the plugin cache path, the session
403
+ id and the config directory: 326 characters. That was only the fallback. The
404
+ primary ScheduledTasks route had already been killed at a fixed eight-second
405
+ ceiling while PowerShell was still loading the module (32 seconds on that
406
+ machine), and a timeout has no stderr, so nothing reported it. The task now
407
+ runs a small launcher in the config directory (`relay-task-<id>.cmd`) so the
408
+ action has one fixed length whatever the plugin path; the PowerShell route
409
+ gets a minute from the command line and the hook's own deadline from a hook;
410
+ and a failure names both routes side by side. `relay doctor` also reports
411
+ whether a login is there to resume under. When that login expires is not
412
+ checked: `claude auth status` does not report it and the relay does not read
413
+ `.credentials.json`, so renew `/login` before a relay that fires hours later.
414
+
415
+ Smaller changes in the same release: the budget line names both account
416
+ readings when Claude Code's cache and the plugin's live reading differ by more
417
+ than five points (`5-hour 20% (the live reading; Claude Code's cache says
418
+ 14%)`) rather than showing one number that jumps; its standing instruction is
419
+ said in full once per session and as twelve words after that
420
+ (`USAGE_LIMITS_BRIEF_FULL=1` keeps the full form); one clause is added on the
421
+ prompt right after a prompt-cache miss the user caused - a changed tool list
422
+ or system prompt, read from the status line's `prompt_cache` - and nothing
423
+ otherwise; the mid-turn pulse's tight sentence is throttled to the same ten
424
+ minutes as its fan-out advice; a subagent message written as several content
425
+ blocks is priced by its last block, which carries the real output count
426
+ (measured on 179 agent transcripts: 1,991 of 2,504 such messages ran like
427
+ 1, 1, 202); and the headless resume runs with `DISABLE_AUTOUPDATER=1`.
428
+
336
429
  | | |
337
430
  | --- | --- |
338
431
  | **Off by default** | Scheduling an agent to run while nobody is watching is a decision you make on purpose, not one a plugin makes for you. |
@@ -576,6 +669,7 @@ minute, so it costs about 400ms cold and 120ms warm.
576
669
  | Variable | Default | Effect |
577
670
  | --- | --- | --- |
578
671
  | `USAGE_LIMITS_BRIEF` | on | Set to `off` to turn the before-prompt line off entirely. |
672
+ | `USAGE_LIMITS_BRIEF_FULL` | off | Set to `1` to say the line's standing instruction in full on every prompt. By default it is said in full once per session and as twelve words after that. |
579
673
  | `USAGE_LIMITS_NEAR` | 90 | Percent used at which the budget counts as tight. Nothing below it is discouraged. |
580
674
  | `USAGE_LIMITS_FEW_TURNS` | 10 | Turns of headroom at or below which the budget counts as tight. |
581
675
  | `USAGE_LIMITS_RUNWAY` | 10 | Minutes of runway at the current pace below which the budget counts as tight. |
@@ -1176,7 +1270,7 @@ ever told it the limit was close. The reader now keeps the newest reading of eac
1176
1270
  meter and prefers the one that actually describes a window. Nothing is merged or
1177
1271
  synthesised - the payload returned is one Codex really wrote.
1178
1272
 
1179
- ### An old snapshot is a floor, not a reading
1273
+ ### An old snapshot is an estimate, not a reading
1180
1274
 
1181
1275
  There was already a warning for a reading spent past its own remainder. It could
1182
1276
  never fire for a snapshot taken at the start of a window, because everything
@@ -1188,7 +1282,11 @@ plugin's live reading was rate-limited into backoff, and the correction was
1188
1282
  quietly carrying the entire difference on its own. A correction is a good
1189
1283
  adjustment to a recent snapshot and a bad substitute for an old one, because the
1190
1284
  pricing error compounds with every point it has to bridge. Past fifteen minutes
1191
- the brief now says the figure is a floor and points at `/usage`.
1285
+ the brief now says the figure is an estimate that can run high or low and points
1286
+ at `/usage`. It is only called a floor when the correction has been refused
1287
+ outright and the raw snapshot is all that is shown, since that really is a
1288
+ lower bound; a snapshot plus a correction is not, and on 2026-09-20 it ran 13
1289
+ to 17 points high.
1192
1290
 
1193
1291
  ### The ceiling
1194
1292
 
@@ -1206,6 +1304,13 @@ It is off until you set a number, it never fires without a reading behind it, an
1206
1304
  the refusal says what to do instead - a denial that only says "over budget" gets
1207
1305
  retried.
1208
1306
 
1307
+ It is judged on the fullest window this session can spend into, and the refusal
1308
+ names that window. A weekly scoped to one model counts only while that model
1309
+ runs - by the setting, or by the model the session's last reply came from, so a
1310
+ `/model` switch is seen - and a window whose reset has passed does not count.
1311
+ Until 1.39.0 it took the highest number on disk, and an Opus session was refused
1312
+ on the Fable weekly with a message calling it "the binding window".
1313
+
1209
1314
  **Why a refusal rather than a sentence.** On Codex the reported figure
1210
1315
  demonstrably does not change behaviour, and the reason is not stubbornness.
1211
1316
  `gpt-6-astra`'s own system prompt, shipped in `models_cache.json`, says: *"Do not
package/commands/relay.md CHANGED
@@ -23,11 +23,16 @@ The rest:
23
23
  when the window opens; a request one second later has been refused before.
24
24
  - `mode notify` - raise a notification with the continuation ready to open.
25
25
  This is the default and it starts nothing by itself.
26
- - `mode resume` - at the wake, run the CLI in the project directory and hand
27
- the continuation back to the same conversation. Say `permission acceptEdits`
28
- (or whichever mode you want) as well: a headless resume does **not** inherit
29
- the session's permission mode, so without one it will sit waiting for an
30
- approval nobody is there to give.
26
+ - `mode resume` - at the wake, resume the same conversation in a window you can
27
+ see, with Remote Control on, so it is also on claude.ai/code and the phone;
28
+ the first prompt points it at the hand-off file. Say `permission acceptEdits`
29
+ (or whichever mode you want) as well: a resume does **not** inherit the
30
+ session's permission mode, so without one it will sit waiting for an
31
+ approval nobody is there to give. `show off` makes it a headless run instead.
32
+ - `voice on|off` - carry how you write in the hand-off, so the resumed session
33
+ answers in your voice without being reminded. On by default.
34
+ - `bugcheck on|always|off` - the hand-off asks for two bug passes before anything
35
+ is called done; `always` asks on every prompt as well. On by default.
31
36
  - `thinking off|resume|always` - `resume` puts the word ultrathink into the
32
37
  prompt the relay delivers. `always` sets `alwaysThinkingEnabled` in your
33
38
  settings, backs the file up first, and applies to new sessions.
@@ -6,6 +6,9 @@ Run `node "${CLAUDE_PLUGIN_ROOT}/skills/usage-limits/scripts/mode.js" $ARGUMENTS
6
6
  and report the answer back. Then stop; do not start other work as part of this
7
7
  command.
8
8
 
9
+ If the user just wants it gone: `off` injects nothing at all, `off --guard 95`
10
+ stays silent until 95 per cent. Say which you ran.
11
+
9
12
  The plugin is not free. It puts a line into your context before every prompt,
10
13
  refreshes readings after tool calls, and keeps a status line alive. A mode
11
14
  changes two things at once: what the plugin tells you to do, and what it costs
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "claude-usage-limits",
3
- "version": "1.26.0",
3
+ "version": "1.39.1",
4
4
  "description": "Puts your remaining Claude Code usage limit into Claude's context before every prompt, so it opens with what fits in the budget instead of starting work that gets cut off. Reports headroom as turns rather than percentages, prices a job before you start it, and detects your plan tier.",
5
5
  "keywords": [
6
6
  "claude",
@@ -42,6 +42,7 @@
42
42
  "scripts": {
43
43
  "test": "node --test",
44
44
  "usage": "node bin/cli.js",
45
- "version": "node tools/sync-version.js && git add .claude-plugin/plugin.json .codex-plugin/plugin.json"
46
- }
45
+ "version": "node tools/sync-version.js && git add .claude-plugin/plugin.json .codex-plugin/plugin.json vscode/package.json"
46
+ },
47
+ "type": "commonjs"
47
48
  }
@@ -275,6 +275,16 @@ The user can also ask any time with `/usage-limits:session`, or run
275
275
  `--sessions` for the history of recent sessions on this machine.
276
276
 
277
277
 
278
+ ## Turning it off
279
+
280
+ If the user says they do not want the budget line, the cost line, or the status
281
+ bars, do not argue and do not explain the trade-off unless asked. Run:
282
+
283
+ node "$CLAUDE_PLUGIN_ROOT/skills/usage-limits/scripts/mode.js" off
284
+
285
+ and, for the bars, `statusline off`. `mode off --guard 95` is the middle
286
+ ground: silent until 95 per cent. Say which one you ran and stop.
287
+
278
288
  ## Budget modes
279
289
 
280
290
  How hard this plugin leans, and what it costs to say it. Four modes, set by the
@@ -620,6 +630,12 @@ which files are mid-change, what must be verified before anything is built on
620
630
  it. It is delivered as a prompt, so an instruction beats a summary. If nothing
621
631
  is written, the relay falls back to the outstanding todo list, which is worse.
622
632
 
633
+ More than one session can be armed at once; each wakes on its own task, and
634
+ cancel takes down only this session's unless --all is given. Arming also
635
+ answers the two start-up questions (folder trust, bypass permissions) in
636
+ .claude.json for the folder it will resume in, so the window does not stop at
637
+ a prompt; doctor says whether they are answered.
638
+
623
639
  **It arms at the end of a reply, not in the middle of one.** Crossing the
624
640
  threshold no longer schedules anything by itself: the relay waits for the reply
625
641
  to finish, so what it carries is work that reached a boundary rather than a
@@ -646,6 +662,22 @@ machine whose power plan forbids wake timers, a scheduled task this account
646
662
  cannot register, no continuation written, no network - and says which would
647
663
  bite. Every resumed run's full output is kept: `relay log --run`.
648
664
 
665
+ An arm that fails says which route failed and why: the ScheduledTasks route
666
+ and the schtasks fallback are named side by side, since on 2026-09-20 only the
667
+ fallback's refusal was reported and the real failure - PowerShell timing out
668
+ while its module loaded - was never seen. The task runs a small launcher in
669
+ the config directory (`relay-task-<id>.cmd`), so the action fits schtasks
670
+ whatever the plugin path. The doctor also reports whether a login is there to
671
+ resume under; when it expires is not checked, so renew `/login` before a
672
+ relay that fires hours later.
673
+
674
+ The budget line's standing instruction - open with the fit line, quote the
675
+ binding window, close finished work with the total - is said in full once per
676
+ session and as twelve words after; it is the same instruction. When two
677
+ account readings disagree by more than five points the line names both rather
678
+ than one, and on the prompt right after a prompt-cache miss you caused it says
679
+ so in one clause.
680
+
649
681
  The relay is off unless the user turned it on, and it only arms while there is
650
682
  an unfinished todo list or an approved plan to carry. Do not turn it on for
651
683
  them, and do not promise behaviour it does not have:
@@ -654,8 +686,13 @@ them, and do not promise behaviour it does not have:
654
686
  input to a shell on purpose; the relay uses it only to tell whether somebody
655
687
  is at the keyboard, and to show a banner.
656
688
  - In `notify` mode — the default — it raises a notification and starts nothing.
657
- - The resumed run opens in a window you can see. `relay show off` hides it; the
658
- output is kept either way.
689
+ - The resumed run is a real interactive session in its own window, with Remote
690
+ Control on (named `usage-limits relay <project>` on claude.ai/code and the
691
+ phone), and its first prompt points it at the hand-off file. `relay show off`
692
+ makes it a headless `claude -p` run instead; the output is kept either way.
693
+ macOS and Linux open the same window in Terminal (unverified there). A saved
694
+ continuation counts as work to carry, so a session with a note but no todo
695
+ list still arms; and a wake that never reports back is reaped as "lost".
659
696
  - In `resume` mode it runs the CLI itself. A headless resume does **not**
660
697
  inherit the session's permission mode, so unless one was set the resumed run
661
698
  will sit waiting for an approval nobody is there to give.
@@ -211,9 +211,10 @@ your own meter remains the best method anyone outside Anthropic has.
211
211
  own cache of the account meter, which refreshes on its own schedule - roughly
212
212
  hourly in practice, because the endpoint behind it rate-limits aggressive
213
213
  polling. Between refreshes every figure here is the last real reading plus
214
- arithmetic. That is why an old snapshot is reported as a floor with its age
214
+ arithmetic. That is why an old snapshot is reported as an estimate with its age
215
215
  attached rather than dressed up as a current percentage, and why `/usage` is
216
- the one way to force a fresh reading.
216
+ the one way to force a fresh reading. It is only called a floor when the
217
+ correction has been refused and the raw snapshot is all that is shown.
217
218
 
218
219
  **A reset time can be in the past.** The cache refreshes when Claude Code
219
220
  talks to the API, so an idle spell leaves it behind. A window whose `resets_at`
@@ -94,22 +94,19 @@ function toolNameOf(input) {
94
94
  // host that DOES have a meter, and because a ceiling that silently does
95
95
  // nothing on the day Antigravity starts publishing one would be worse.
96
96
  function percentNow(now) {
97
- let worst = null;
97
+ // The same no-scan view and the same rule as the Claude Code pulse: the
98
+ // fullest window this agent can spend into, never a window that has reset.
98
99
  try {
99
- const snapshot = usage.collect(now);
100
- const utilization = snapshot && snapshot.utilization;
101
- if (utilization && typeof utilization === 'object') {
102
- for (const key of Object.keys(utilization)) {
103
- const window = utilization[key];
104
- if (!window || typeof window !== 'object') continue;
105
- const value = Number(window.utilization);
106
- if (Number.isFinite(value) && (worst === null || value > worst)) worst = value;
107
- }
108
- }
100
+ return ceiling.worstWindow(usage.snapshotWindows(usage.collect(now), now, null));
109
101
  } catch (err) {
110
102
  // No reading is a reason not to enforce, never a reason to throw.
103
+ return null;
111
104
  }
112
- return worst;
105
+ }
106
+
107
+ // assess() takes the number and the name of the window it belongs to.
108
+ function reading(worst) {
109
+ return { percent: worst ? worst.percent : null, label: worst ? worst.label : null };
113
110
  }
114
111
 
115
112
  async function run(now, input, argv) {
@@ -124,7 +121,7 @@ async function run(now, input, argv) {
124
121
  if (event === 'PreToolUse') {
125
122
  const tool = toolNameOf(input);
126
123
  if (!ceiling.isMultiplier(tool)) return {};
127
- const at = ceiling.assess({ percent: percentNow(now), state: budget.state, env: process.env, sessionId });
124
+ const at = ceiling.assess(Object.assign({ state: budget.state, env: process.env, sessionId }, reading(percentNow(now))));
128
125
  const call = ceiling.verdict(at, tool);
129
126
  if (call.decision !== 'deny') return {};
130
127
  return { decision: 'deny', reason: call.reason };
@@ -141,7 +138,7 @@ async function run(now, input, argv) {
141
138
  text = '';
142
139
  }
143
140
  const warning = ceiling.warning(
144
- ceiling.assess({ percent: percentNow(now), state: budget.state, env: process.env, sessionId })
141
+ ceiling.assess(Object.assign({ state: budget.state, env: process.env, sessionId }, reading(percentNow(now))))
145
142
  );
146
143
  const message = [text, warning].filter(Boolean).join(' ');
147
144
  return message ? { injectSteps: [{ ephemeralMessage: message }] } : { injectSteps: [] };