claude-usage-limits 1.40.3 → 1.40.5

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1,7 +1,7 @@
1
1
  {
2
2
  "name": "usage-limits",
3
3
  "displayName": "Usage Limits",
4
- "version": "1.40.3",
4
+ "version": "1.40.5",
5
5
  "description": "Puts your remaining Claude Code usage limit into Claude's context before every prompt, so it opens with what fits in the budget instead of starting work that gets cut off. Reports headroom as turns rather than percentages, prices a job before you start it, and detects your plan tier.",
6
6
  "author": {
7
7
  "name": "Ridelink",
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "usage-limits",
3
- "version": "1.40.3",
3
+ "version": "1.40.5",
4
4
  "description": "Reports how much of your Codex usage limit is left as turns of work rather than a percentage, prices a job before you start it, and counts the other agents sharing the same budget.",
5
5
  "author": {
6
6
  "name": "Ridelink",
package/README.md CHANGED
@@ -393,20 +393,20 @@ relayed run either.
393
393
  At the wall with weekly headroom, the line looks like this:
394
394
 
395
395
  ```
396
- If the wall offers it, /low-priority carries this session past the 5-hour limit
397
- at lower priority instead of stopping: it spends the weekly limit, which is at
398
- 45% and so has room, and replies may pause while it waits for spare capacity -
399
- the wait and its ceiling are set by the server per request. It is a toggle: you
400
- type it yourself and run it again to stop, and nothing here can switch it on
401
- for you.
396
+ If the wall offers it, /low-priority can carry this session past the 5-hour
397
+ limit at lower priority instead of stopping: it spends the weekly limit, which
398
+ is at 45% and so has room, and replies may pause while it waits for spare
399
+ capacity - the wait and its ceiling are set by the server per request. It is a
400
+ toggle only the user can type, and type again to stop: you cannot run a slash
401
+ command, and nothing here can switch it on.
402
402
  ```
403
403
 
404
404
  **The recommendation has a number behind it.** It is offered only while the
405
405
  weekly is at or below **80 per cent** and the 5-hour window is the binding wall.
406
406
  Above that the brief actively says not to, and cites the figure: low-priority
407
407
  spends the weekly *and* draws on a weekly allowance whose size is exposed to no
408
- hook and no file, and a real user measured it emptying most of a week in a
409
- couple of hours. No wait time is ever printed, because the retry and the ceiling
408
+ hook and no file, and one user has reported it burning through almost a week of
409
+ usage in a couple of hours (anthropics/claude-code#92544). No wait time is ever printed, because the retry and the ceiling
410
410
  come from `lowPriorityRetryAfterSeconds` and `lowPriorityMaxWaitSeconds` on each
411
411
  response — any fixed "20 seconds, 20 minutes" would be invented.
412
412
 
@@ -442,13 +442,19 @@ wake. So when a resume relay is armed and this is on, the brief says so: the
442
442
  wake is the route for a session that will be **closed** at the reset, and both
443
443
  firing for the same reset would start the work twice and spend the weekly twice.
444
444
 
445
- **`/limit-reset`**, the once-weekly manual session reset, is detected but not
446
- built on: this account holds no grant (`tengu_cedar_ember` absent,
447
- `cachedUsageUtilization.cedar_ember` null), so a feature resting on it would be
448
- untestable. If a grant appears, the brief names it at the wall, says it only
449
- works while you are actually at a limit, and says the work it unlocks still
450
- spends the weekly. It never reports how many are left — `resets_left` comes from
451
- a live endpoint and appears in no file a hook can read.
445
+ **`/limit-reset`** is one command behind two different server flags, and the
446
+ brief words each in its own copy's terms. `tengu_nifty_lemur` is the
447
+ once-a-week reset of the 5-hour session limit ("uses weekly limit · 1/week",
448
+ "your weekly limit still applies"); `tengu_cedar_ember` is a counted grant with
449
+ a use-by date that "refills your limits". The account this was built on carries
450
+ `tengu_nifty_lemur` enabled and no `tengu_cedar_ember`. When the budget is tight
451
+ and the 5-hour window is the wall, the brief names the weekly reset, says the
452
+ work it unlocks still counts toward the weekly, and says only the user can type
453
+ it; at a weekly wall it says nothing, because resetting the 5-hour limit cannot
454
+ help there. It never reports whether this week's reset is used or how many
455
+ grants are left — both come from a live endpoint and appear in no file a hook
456
+ can read. A flag counts only when it is a plain object with `enabled: true`,
457
+ the way the CLI reads it.
452
458
 
453
459
  **Grace is not weekly spend.** The allowance the wall gives you is metered in
454
460
  its own per-window meters (`anthropic-ratelimit-unified-grace-5h-utilization`
@@ -1114,7 +1120,23 @@ There are two separate things people mean by "change the model":
1114
1120
 
1115
1121
  `claude-usage-limits mode --baseline` shows both side by side. The budget line
1116
1122
  now says the tier as well, and where the reading came from, because the number
1117
- that decides what a turn costs was the one number the line never printed.
1123
+ that decides what a turn costs was the one number the line never printed. It
1124
+ names the version, not only the family (`opus 5.5/xhigh`, not `opus/xhigh`),
1125
+ read from the newest assistant message in the session's transcript: a settings
1126
+ alias such as `opus` cannot say which Opus answered.
1127
+
1128
+ When the model running is an older release of a family that has a newer one at
1129
+ a lower price - Opus 5 or 4.8 against Opus 5.5 ($4/$20, cache reads $0.20
1130
+ against $0.50), Fable 5 against Fable 5.1 (reads $0.25 against $1) - the brief
1131
+ says so once per session, with the prices and how many turns the one-off cache
1132
+ rebuild takes to repay. In Claude Code it offers `/model <id>`, which switches
1133
+ the session and saves the model as the default for new sessions; elsewhere it
1134
+ offers the host's own model setting and names no slash command. A relay whose
1135
+ own `model` is pinned to the older release is named too, since a wake starts
1136
+ with that `--model` whatever the session switched to. Releases on
1137
+ either side of the 4.7 tokenizer change are not compared, because a price per
1138
+ token is not like for like across it. `mode --no-advice` mutes this with the
1139
+ rest of the advice.
1118
1140
 
1119
1141
  What Claude can genuinely move, stated without embroidery: the model on an
1120
1142
  `Agent` call, and the model and effort inside a `Workflow` script. Its own
package/commands/relay.md CHANGED
@@ -29,6 +29,9 @@ The rest:
29
29
  (or whichever mode you want) as well: a resume does **not** inherit the
30
30
  session's permission mode, so without one it will sit waiting for an
31
31
  approval nobody is there to give. `show off` makes it a headless run instead.
32
+ - `model <id>` - the `--model` a resumed run starts with; `model` alone goes
33
+ back to the default. It is separate from `/model` in the session, so a wake
34
+ pinned to an older model stays on it until this is changed.
32
35
  - `voice on|off` - carry how you write in the hand-off, so the resumed session
33
36
  answers in your voice without being reminded. On by default.
34
37
  - `bugcheck on|always|off` - the hand-off asks for two bug passes before anything
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "claude-usage-limits",
3
- "version": "1.40.3",
3
+ "version": "1.40.5",
4
4
  "description": "Puts your remaining Claude Code usage limit into Claude's context before every prompt, so it opens with what fits in the budget instead of starting work that gets cut off. Reports headroom as turns rather than percentages, prices a job before you start it, and detects your plan tier.",
5
5
  "keywords": [
6
6
  "claude",
@@ -625,10 +625,11 @@ longer stops at. **Record it only when they said so.** Whether it is running is
625
625
  in the CLI's process memory and readable nowhere, so nothing here may infer it,
626
626
  and `mode --low-priority off` is how it is taken back.
627
627
 
628
- The same applies to `/limit-reset`, the once-weekly manual session reset: the
629
- line names it only if a grant is actually readable, it only works while you are
630
- at a limit, the work it unlocks still spends the weekly, and how many are left is
631
- not knowable from here.
628
+ The same applies to `/limit-reset`, the manual reset: only the user can type
629
+ it. The line names it only when the account's flags say it is set up. The
630
+ once-a-week kind resets the 5-hour limit and its work still counts toward the
631
+ weekly, so it is named only at a 5-hour wall; whether this week's is already
632
+ used, and how many of the counted kind are left, is not knowable from here.
632
633
 
633
634
  And if the budget line says Claude Code is injecting its own wrap-up note at this
634
635
  wall, follow that note: it is more specific than anything here, and two
@@ -256,8 +256,9 @@ A bracketed suffix on a model id (`claude-sonnet-5[1m]`) is stripped before
256
256
  the lookup: it marks a context-window variant of the same model, not a new
257
257
  one. Cache reads price at a tenth of the input rate unless a row carries a
258
258
  `cacheRead` figure of its own - Fable and Mythos 5.1 price reads outright at
259
- $0.25 per million, far under the tenth rule, and reads are the dominant input
260
- in exactly the long sessions where the difference matters.
259
+ $0.25 per million (0.025x input) and Opus 5.5 at $0.20 (0.05x its $4 input),
260
+ all under the tenth rule, and reads are the dominant input in exactly the long
261
+ sessions where the difference matters.
261
262
 
262
263
  Until someone does, a model this table has not seen is priced at the average of
263
264
  the family its name contains: an unreleased `claude-opus-5-2` is charged at the
@@ -160,7 +160,7 @@ function readSaid() {
160
160
  }
161
161
  }
162
162
  function keepSaidFor(key) {
163
- return /#(standing|cachemiss|stale|relaylast)$/.test(key) ? 24 * 60 * 60 * 1000 : 60 * 60 * 1000;
163
+ return /#(standing|cachemiss|stale|relaylast|newer)$/.test(key) ? 24 * 60 * 60 * 1000 : 60 * 60 * 1000;
164
164
  }
165
165
 
166
166
  // A plugin update takes effect when Claude Code restarts, so a session that
@@ -202,6 +202,30 @@ function staleVersionFor(sessionId, now, dir) {
202
202
  return 'usage-limits ' + installed + ' is installed but this session still runs ' + running +
203
203
  ', because a plugin update applies at the next start; a relay or cap set here follows the older rules until then.';
204
204
  }
205
+ // A newer model of the same family at a lower price, said once a session.
206
+ // Once is the rule every recommendation here keeps: the fact does not change
207
+ // from prompt to prompt, and repeating it would be the plugin charging for its
208
+ // own presence. Keyed on the pair, so a session that moves to a different model
209
+ // with its own newer sibling hears about that one.
210
+ function newerModelFor(sessionId, advice, now) {
211
+ if (!advice || !advice.text) return null;
212
+ const at = Number.isFinite(now) ? now : Date.now();
213
+ const key = String(sessionId || '_') + '#newer';
214
+ const all = readSaid();
215
+ const entry = all[key];
216
+ if (entry && entry.seen === advice.id && Number.isFinite(entry.at) && at - entry.at < keepSaidFor(key)) return null;
217
+ all[key] = { at, seen: advice.id };
218
+ writeSaid(all, at);
219
+ return advice.text;
220
+ }
221
+ // The model the relay resumes with, when one is set: wake.js passes it as --model.
222
+ function relayModelNow() {
223
+ try {
224
+ return relay.settings(relay.read()).model || null;
225
+ } catch (err) {
226
+ return null;
227
+ }
228
+ }
205
229
  // How the last relay ended is news once. It used to ride along for six hours
206
230
  // after any relay ended, on every prompt of every session: on 2026-09-22 a wake
207
231
  // lost at 5:53 AM was repeated in two sessions' briefs all afternoon, about
@@ -784,6 +808,7 @@ function briefText(input) {
784
808
  if (parts.tier) sentences.push(parts.tier);
785
809
  // Once per session, and only when the installed version is not this one.
786
810
  if (parts.staleVersion) sentences.push(parts.staleVersion);
811
+ if (parts.newerModel) sentences.push(parts.newerModel);
787
812
  const bounded = mode.boundsNote(bounds);
788
813
  if (bounded) sentences.push(bounded);
789
814
  if (parts.planChanged) {
@@ -1028,53 +1053,99 @@ function briefText(input) {
1028
1053
  // style drops only what reads the same every turn. This does not.
1029
1054
  const lp = parts.lowPriority || null;
1030
1055
  let lowPrioritySentence = null;
1031
- if (lp && lp.state === 'acknowledged') {
1056
+ // The acknowledged sentence describes the SWAP, so it may only be said when
1057
+ // the swap really happened. wallFeatures refuses to swap unless there is a
1058
+ // weekly window with a live percentage and a known reset, and this branch
1059
+ // used to fire on the acknowledgement alone - so on an account with no
1060
+ // readable weekly the brief said "binding window is 5-hour 99% used" in one
1061
+ // clause and "the figures above are the weekly" in the next. A line that
1062
+ // contradicts the numbers printed beside it is worse than no line: it tells
1063
+ // the model to ignore a wall that is still there.
1064
+ if (lp && lp.state === 'acknowledged' && lp.weeklyBinding) {
1032
1065
  lowPrioritySentence = (
1033
- 'You have said /low-priority is on, so the brake is the weekly window and not the 5-hour one: ' +
1066
+ 'The user has said /low-priority is on, so the brake is the weekly window and not the 5-hour one: ' +
1034
1067
  (Number.isFinite(lp.fiveHourPercent)
1035
1068
  ? 'the 5-hour limit is at ' + lp.fiveHourPercent + '% and no longer stops this session, '
1036
1069
  : 'the 5-hour limit no longer stops this session, ') +
1037
1070
  'the figures above are the weekly, and replies may pause while lower priority waits for spare ' +
1038
- 'capacity. Whether it is still on cannot be read from here, so if it has ended say so with ' +
1039
- 'usage-mode --low-priority off and the 5-hour wall counts again. It also draws on a separate ' +
1071
+ 'capacity. Whether it is still on cannot be read from here, so if the user says it has ended, ' +
1072
+ 'record that with usage-mode --low-priority off and the 5-hour wall counts again. It also draws on a separate ' +
1040
1073
  'weekly lower-priority allowance that is exposed to no hook and no file, so nothing here can ' +
1041
1074
  'track how much of that is left.'
1042
1075
  );
1076
+ } else if (lp && lp.state === 'acknowledged') {
1077
+ // Acknowledged, but there is no weekly reading to brake on, so the figures
1078
+ // above are still the window they say they are and the wall is still real.
1079
+ lowPrioritySentence = (
1080
+ 'The user has said /low-priority is on, which spends the weekly limit instead of stopping at the ' +
1081
+ '5-hour one - but there is no usable weekly reading here to brake on, so the figures above are ' +
1082
+ 'the window named beside them and that window is still the one to plan against. /usage (the ' +
1083
+ 'user types it) refreshes the account snapshot; until it has a weekly, treat the limit above as real.'
1084
+ );
1043
1085
  } else if (lp && lp.advise && lp.advise.kind === 'offer') {
1044
1086
  lowPrioritySentence = (
1045
- 'If the wall offers it, /low-priority carries this session past the 5-hour limit at lower ' +
1087
+ 'If the wall offers it, /low-priority can carry this session past the 5-hour limit at lower ' +
1046
1088
  'priority instead of stopping: it spends the weekly limit, which is at ' +
1047
1089
  lp.advise.weeklyPercent + '% and so has room, and replies may pause while it waits for spare ' +
1048
- 'capacity - the wait and its ceiling are set by the server per request. It is a toggle: you ' +
1049
- 'type it yourself and run it again to stop, and nothing here can switch it on for you. Say ' +
1050
- 'in one line that it is there; if it is taken, usage-mode --low-priority on is what makes ' +
1051
- 'this line brake on the weekly instead.'
1090
+ 'capacity - the wait and its ceiling are set by the server per request. It is a toggle only ' +
1091
+ 'the user can type, and type again to stop: you cannot run a slash command, and nothing here ' +
1092
+ 'can switch it on. Tell the user in one line that it is there; if the user says it is taken, ' +
1093
+ 'usage-mode --low-priority on is what makes this line brake on the weekly instead.'
1052
1094
  );
1053
1095
  } else if (lp && lp.advise && lp.advise.kind === 'hold') {
1054
1096
  lowPrioritySentence = (
1055
1097
  'The wall may offer /low-priority here, and it is not worth taking: it spends the weekly limit, ' +
1056
1098
  'which is already at ' + lp.advise.weeklyPercent + '% - past the ' + lp.advise.threshold +
1057
1099
  ' per cent this plugin will recommend it at - and it draws on a weekly lower-priority ' +
1058
- 'allowance as well, which has been measured emptying most of a week in a couple of hours. ' +
1059
- 'Say in one line that the answer is no and why; waiting out the 5-hour reset is the cheaper move.'
1100
+ 'allowance as well, which one user has reported burning through almost a week of usage in a ' +
1101
+ 'couple of hours. The choice is the user\'s, and only the user can type it: say in one line ' +
1102
+ 'that this plugin advises against it and why; waiting out the 5-hour reset is the cheaper move.'
1060
1103
  );
1061
1104
  }
1062
1105
  if (lowPrioritySentence) sentences.push(lowPrioritySentence);
1063
- // A manual session reset, if this account ever gets one.
1106
+ // A manual reset behind /limit-reset, when this account holds one.
1107
+ //
1108
+ // Two server flags sit behind the one command and they are different offers,
1109
+ // so each is worded in its own copy's terms and nothing is borrowed across
1110
+ // (lowpri.sessionReset has the bundle strings):
1111
+ //
1112
+ // 'weekly' (tengu_nifty_lemur, which this account has): resets the 5-hour
1113
+ // session limit, once a week, and the work still counts toward the weekly.
1114
+ // So it only helps when the 5-hour window is the wall - said at a weekly
1115
+ // wall it would be advice that cannot work - and the first build, reading
1116
+ // only the other flag, never said it at all.
1064
1117
  //
1065
- // /limit-reset refills the 5-hour window, works only AT a limit, and is once a
1066
- // week - and the work it unlocks still spends the weekly, which is the part
1067
- // worth saying out loud. This account holds no grant today
1068
- // (tengu_cedar_ember absent, cachedUsageUtilization.cedar_ember null), so the
1069
- // sentence is a detector rather than a feature: nothing is claimed about how
1070
- // many resets are left, because resets_left is served by a live endpoint and
1071
- // appears in no file a hook can read.
1072
- if (lp && lp.resetGrant && (parts.pressure === 'tight' || parts.pressure === 'gone')) {
1118
+ // 'grant' (tengu_cedar_ember): a counted reset with a use-by date that
1119
+ // "refills your limits". Its copy does not say it spends the weekly, and it
1120
+ // has an early-use path, so neither "once a week" nor "only at a limit" is
1121
+ // claimed for it.
1122
+ //
1123
+ // Neither is counted: whether this week's reset is already used, and how
1124
+ // many grants are left, are served by the API and appear in no file. And the
1125
+ // CLI refuses a reset while lower priority runs ("Resets can't be used while
1126
+ // you continue at lower priority"), which is why the weekly variant stays
1127
+ // silent once the brief has swapped to the weekly, and the grant says so.
1128
+ const resetVariant = lp && lp.resetGrant ? lp.resetVariant || 'grant' : null;
1129
+ const resetPressure = parts.pressure === 'tight' || parts.pressure === 'gone';
1130
+ const fiveHourWall = Boolean(parts.binding && parts.binding.key === 'five_hour');
1131
+ // Acknowledged is checked on its own as well as through the binding: with no
1132
+ // weekly reading the brief cannot swap, the 5-hour window stays binding, and
1133
+ // the reset would be offered to a session the CLI will not reset.
1134
+ if (resetVariant === 'weekly' && resetPressure && fiveHourWall && lp.state !== 'acknowledged') {
1135
+ sentences.push(
1136
+ 'This account is set up for a manual session reset (/limit-reset): once a week it resets the ' +
1137
+ '5-hour limit, and the work it unlocks still counts toward the weekly, so it moves the 5-hour ' +
1138
+ 'wall rather than adding budget. The server decides whether it can be used right now, and ' +
1139
+ "whether this week's is already used is not readable from here. Like /low-priority, only the user can type " +
1140
+ 'it; you cannot run a slash command.'
1141
+ );
1142
+ } else if (resetVariant === 'grant' && resetPressure) {
1073
1143
  sentences.push(
1074
- 'A once-weekly manual session reset appears to be available on this account (/limit-reset). It ' +
1075
- 'only works while you are actually AT a limit, and the work it unlocks still spends the weekly, ' +
1076
- 'so it moves the 5-hour wall rather than adding budget. How many are left is not readable from ' +
1077
- 'here - the CLI asks the server for that. It is yours to type, like /low-priority.'
1144
+ 'This account appears to hold a limit reset (/limit-reset), which the CLI describes as refilling ' +
1145
+ 'your limits while the weekly reset day stays put. How many are left and until when is not ' +
1146
+ 'readable from here - the CLI asks the server for that.' +
1147
+ (lp.state === 'acknowledged' ? ' The CLI will not use one while lower priority is on.' : '') +
1148
+ ' Like /low-priority, only the user can type it; you cannot run a slash command.'
1078
1149
  );
1079
1150
  }
1080
1151
  if (parts.session) {
@@ -1357,7 +1428,7 @@ function briefText(input) {
1357
1428
  // The one recommendation this session is allowed, in its short form. It
1358
1429
  // still cites the measurement, still names the command: terse is fewer
1359
1430
  // words, not less evidence.
1360
- const adviceText = parts.adviceText ? ' ' + parts.adviceText : '';
1431
+ const adviceText = (parts.adviceText ? ' ' + parts.adviceText : '') + (parts.newerModel ? ' ' + parts.newerModel : '');
1361
1432
  return (
1362
1433
  sentences[0] + caveat + (parts.tier ? ' ' + parts.tier : '') + (bounded ? ' ' + bounded : '') + adviceText +
1363
1434
  (escapeSentence ? ' ' + escapeSentence : '') +
@@ -1524,9 +1595,17 @@ function wallFeatures(now, binding, windows, hostName) {
1524
1595
  if (weekly && Number.isFinite(weekly.percentUsed) && !weekly.stale && Number.isFinite(weekly.resetsAt)) {
1525
1596
  out.binding = weekly;
1526
1597
  out.swapped = true;
1527
- out.lowPriority = Object.assign({}, info, { weeklyBinding: true });
1528
1598
  }
1529
1599
  }
1600
+ // Whether the weekly really is the window the figures describe, which is
1601
+ // what the acknowledged sentence claims. True when the swap just happened
1602
+ // AND when the weekly was already binding on its own - the 5-hour window
1603
+ // reading low is the ordinary case for a session that has been running at
1604
+ // lower priority for a while. False when there was no weekly to brake on,
1605
+ // and then the sentence has to say so instead of claiming otherwise.
1606
+ out.lowPriority = Object.assign({}, info, {
1607
+ weeklyBinding: Boolean(out.binding && out.binding.key === 'seven_day'),
1608
+ });
1530
1609
  } catch (err) {
1531
1610
  // Nothing about these features is worth a failed prompt.
1532
1611
  }
@@ -1850,7 +1929,10 @@ async function run(now, hookInput, opts) {
1850
1929
  // What tier is producing this turn, and what the user's own baseline is.
1851
1930
  // Read, displayed, never written.
1852
1931
  const terse = budget.policy.briefStyle === 'terse';
1853
- const tier = mode.tierLine(mode.tierNow({ sessionId, now, usage, env: process.env }), { terse });
1932
+ const tierReading = mode.tierNow({ sessionId, now, usage, env: process.env, transcriptPath: hookInput && hookInput.transcript_path });
1933
+ const tier = mode.tierLine(tierReading, { terse });
1934
+ // Muted advice is muted for this too: it is a recommendation like the others.
1935
+ const newerAdvice = budget.advice && budget.advice.off ? null : mode.newerModelAdvice(tierReading, { usage, host: usage.currentHost(), bounds: budget.bounds, relayModel: relayModelNow() });
1854
1936
 
1855
1937
  // The recommendation channel. The measured fit sentence IS the
1856
1938
  // recommendation - it cites this account's own numbers and names the exact
@@ -1905,6 +1987,7 @@ async function run(now, hookInput, opts) {
1905
1987
  standingShort,
1906
1988
  cacheMissWhy: cacheMissWhyFor(sessionId, now),
1907
1989
  staleVersion: staleVersionFor(sessionId, now),
1990
+ newerModel: newerModelFor(sessionId, newerAdvice, now),
1908
1991
  adviceText: terse && offering ? advice.text : null,
1909
1992
  relay: carry,
1910
1993
  voiceNote,
@@ -2011,7 +2094,7 @@ function withBugcheck(text) {
2011
2094
  return text;
2012
2095
  }
2013
2096
 
2014
- module.exports = { wallFeatures, sweepDebris, withBugcheck, sayOnce, shapeOf, saidFile, REPEAT_MS, staleVersionFor, relayNewsFor, installedVersion, runningVersion,
2097
+ module.exports = { wallFeatures, sweepDebris, withBugcheck, sayOnce, shapeOf, saidFile, REPEAT_MS, staleVersionFor, newerModelFor, relayNewsFor, installedVersion, runningVersion,
2015
2098
  readSaid, standingSaid, markStanding, standingShortFor, STANDING_SHORT, cacheMissWhyFor, missReason, MISS_RECENT_MS,
2016
2099
  DEFAULTS,
2017
2100
  HOOK_BUDGET_MS,
@@ -39,13 +39,14 @@
39
39
  // 30-minute freshness cutoff on cached window readings, nothing to do with
40
40
  // this.)
41
41
  //
42
- // /limit-reset - a once-weekly manual refill of the 5-hour window, usable
43
- // only AT a limit, whose work still spends the weekly. Gated on
44
- // tengu_cedar_ember, which is ABSENT from this account's feature cache, with
45
- // cachedUsageUtilization.utilization.cedar_ember null. So there is nothing to
46
- // spend here and nothing is built on it: only a detector that starts
47
- // reporting if a grant ever appears. resets_left is served by a live
48
- // endpoint, never a file, so the count is never claimed.
42
+ // /limit-reset - a manual reset, behind two different server flags (see
43
+ // sessionReset below). tengu_nifty_lemur is the once-a-week reset of the
44
+ // 5-hour session limit whose work still counts toward the weekly, and THIS
45
+ // ACCOUNT HAS IT (enabled:true). tengu_cedar_ember is a counted grant with a
46
+ // use-by date, absent here, with cachedUsageUtilization.utilization.cedar_ember
47
+ // null. Nothing is spent by the plugin - only the user can type the command -
48
+ // and whether this week's reset is used, or how many grants are left, is
49
+ // served by the API and never a file, so neither is ever claimed.
49
50
  //
50
51
  // The graceful wrap-up note - the CLI injecting "finish up" at the wall. The
51
52
  // mechanism is real and the treatment TEXT is provisioned on this machine
@@ -81,6 +82,7 @@ const atomic = require('./atomic.js');
81
82
 
82
83
  const FLAG = 'tengu_toasty_breeze';
83
84
  const RESET_FLAG = 'tengu_cedar_ember';
85
+ const SESSION_RESET_FLAG = 'tengu_nifty_lemur';
84
86
  const WRAPUP_MODE_FLAG = 'tengu_lantern_wick_mode';
85
87
  const WRAPUP_TEXT_FLAG = 'tengu_lantern_wick_text';
86
88
  const NEAR_WALL_FLAG = 'tengu_vellum_anchor';
@@ -95,7 +97,7 @@ const MAX_COOLOFF_MINUTES = 1440;
95
97
  // The recommendation rule, in one number so it can be argued with.
96
98
  //
97
99
  // /low-priority spends the weekly and draws on a weekly allowance whose size is
98
- // exposed nowhere a hook can read, and a real user measured it burning "almost
100
+ // exposed nowhere a hook can read, and one user has reported it burning "almost
99
101
  // a week of usage in a couple hours"
100
102
  // (github.com/anthropics/claude-code/issues/92544). So it is worth naming only
101
103
  // while the weekly still has real room. At or below this it is offered; above
@@ -106,6 +108,22 @@ const WEEKLY_HEADROOM_MAX = 80;
106
108
  // offers the toggle at all.
107
109
  const WALL_PERCENT = 90;
108
110
 
111
+ // How long an acknowledgement with no readable reset time is allowed to live.
112
+ //
113
+ // The record is meant to lapse at the reset of the window it belongs to. When
114
+ // that reset could not be read - a session with no account snapshot, or a
115
+ // snapshot with no five_hour entry - the record used to be written with
116
+ // resetsAt null, and readAck only lapsed a FINITE resetsAt. So it never lapsed:
117
+ // a record written once went on redirecting the headroom maths at the weekly
118
+ // for ever. Probed on 2026-09-26 with an acknowledgement three days old and the
119
+ // 5-hour window reading 0 per cent, and the brief still said the 5-hour limit
120
+ // "no longer stops this session".
121
+ //
122
+ // The window the toggle is offered at is five hours long, so an
123
+ // acknowledgement made against it cannot honestly outlive five hours from when
124
+ // it was made. That is the ceiling, not a guess at when it really ended.
125
+ const ACK_MAX_MS = 5 * 60 * 60 * 1000;
126
+
109
127
  function configDir() {
110
128
  return process.env.CLAUDE_CONFIG_DIR || path.join(os.homedir(), '.claude');
111
129
  }
@@ -203,16 +221,53 @@ function offer(account) {
203
221
  };
204
222
  }
205
223
 
206
- // The manual session reset. Detected, never spent: this account holds no grant,
207
- // so a feature built on it would be untestable here.
224
+ // A flag object the way the CLI reads one: a plain object whose `enabled` is
225
+ // exactly true. The bundle's own readers are
226
+ // function qae(){let e=x("tengu_nifty_lemur",{});return typeof e==="object"&&e!==null&&!Array.isArray(e)?e:{}}
227
+ // function pme(){return qae().enabled===!0}
228
+ // and the same shape for tengu_cedar_ember ($Z()). Truthiness is not enough:
229
+ // `{ enabled: false }` is a truthy object, and the first build counted it as a
230
+ // reset on offer.
231
+ function flagEnabled(gb, key) {
232
+ if (!gb || !Object.prototype.hasOwnProperty.call(gb, key)) return false;
233
+ const value = gb[key];
234
+ return Boolean(value && typeof value === 'object' && !Array.isArray(value) && value.enabled === true);
235
+ }
236
+
237
+ // The manual reset behind /limit-reset. Detected, never spent.
238
+ //
239
+ // There are TWO server flags behind the one command, and they describe
240
+ // different things, so the brief must not word one in the other's terms:
241
+ //
242
+ // tengu_nifty_lemur - "Reset your session limit now and keep working; once a
243
+ // week, still counts toward your weekly limit" (the command's description
244
+ // when this is the variant). Its notice: "/limit-reset to reset your session
245
+ // limit now · uses weekly limit · 1/week". THIS ACCOUNT HAS IT: enabled:true,
246
+ // version 1, read from ~/.claude.json on 2026-09-25. The first build looked
247
+ // only at tengu_cedar_ember, called the account grant-less, and so never said
248
+ // a word about a reset the account actually holds.
249
+ //
250
+ // tengu_cedar_ember - "Use an available limit reset and keep working". A
251
+ // counted grant: "Refills your {limits} now · your weekly reset day stays
252
+ // {week}", "{resets} left · use by {deadline}", and an early-use path ("You
253
+ // haven't reached a limit yet - use your reset anyway?"). Not once a week,
254
+ // not only at a limit, and whether it spends the weekly is not in its copy.
255
+ //
256
+ // The CLI prefers cedar_ember when both are on (its description reads
257
+ // `$Z()?"Use an available limit reset...":"Reset your session limit now..."`),
258
+ // so this does too. Whether this week's reset is already spent, and how many
259
+ // grants are left, come from the server and appear in no file.
208
260
  function sessionReset(account) {
209
261
  const gb = features(account);
210
262
  const util = utilization(account);
211
- const flagged = Boolean(gb && Object.prototype.hasOwnProperty.call(gb, RESET_FLAG) && gb[RESET_FLAG]);
263
+ const grantFlag = flagEnabled(gb, RESET_FLAG);
264
+ const weeklyFlag = flagEnabled(gb, SESSION_RESET_FLAG);
212
265
  const grant = util && util.cedar_ember ? util.cedar_ember : null;
266
+ const variant = grantFlag || grant ? 'grant' : weeklyFlag ? 'weekly' : null;
213
267
  return {
214
268
  known: Boolean(gb || util),
215
- present: Boolean(flagged || grant),
269
+ present: variant !== null,
270
+ variant,
216
271
  // resets_left is served from /api/organizations/<uuid>/reset_rate_limits,
217
272
  // not from any file a hook can read, so it stays unreported.
218
273
  resetsLeft: null,
@@ -276,9 +331,13 @@ function autoContinue() {
276
331
  // session where the 5-hour wall is real again.
277
332
  // ---------------------------------------------------------------------------
278
333
 
334
+ // An array is an object too, and JSON.stringify drops a named property set on
335
+ // one: a state file holding `[]` made acknowledge() write `[]` straight back,
336
+ // report saved:true, and usage-mode print "Recorded" for a statement that the
337
+ // next read could not find. Only a plain object is a state.
279
338
  function readState() {
280
339
  const parsed = readJson(stateFile());
281
- return parsed && typeof parsed === 'object' ? parsed : {};
340
+ return parsed && typeof parsed === 'object' && !Array.isArray(parsed) ? parsed : {};
282
341
  }
283
342
 
284
343
  function writeState(state) {
@@ -290,14 +349,20 @@ function writeState(state) {
290
349
  }
291
350
  }
292
351
 
352
+ // `saved` is reported, never assumed. A state path that cannot be written -
353
+ // a directory sitting where the file goes, a read-only home - used to return
354
+ // the record anyway, so usage-mode printed "Recorded: you have switched
355
+ // /low-priority on" and then the very next brief behaved as though nothing had
356
+ // been said. Telling somebody their statement was recorded when it was not is
357
+ // the one thing this module must not do.
293
358
  function acknowledge(input) {
294
359
  const options = input || {};
295
360
  const now = Number.isFinite(options.now) ? options.now : Date.now();
296
361
  if (!options.on) {
297
362
  const state = readState();
298
363
  delete state.ack;
299
- writeState(state);
300
- return { on: false, at: now };
364
+ const saved = writeState(state);
365
+ return { on: false, at: now, saved, file: stateFile() };
301
366
  }
302
367
  const record = {
303
368
  on: true,
@@ -307,18 +372,35 @@ function acknowledge(input) {
307
372
  };
308
373
  const state = readState();
309
374
  state.ack = record;
310
- writeState(state);
311
- return record;
375
+ const saved = writeState(state);
376
+ return Object.assign({}, record, { saved, file: stateFile() });
377
+ }
378
+
379
+ // When a record stops being true: the reset of the window it was stamped
380
+ // against, or - when that could not be read - ACK_MAX_MS after it was made.
381
+ // Both are returned as `expiresAt` so a caller never has to work it out twice
382
+ // and get a different answer.
383
+ function ackExpiry(ack) {
384
+ if (!ack) return null;
385
+ if (Number.isFinite(ack.resetsAt)) return ack.resetsAt;
386
+ if (Number.isFinite(ack.at)) return ack.at + ACK_MAX_MS;
387
+ return null;
312
388
  }
313
389
 
314
390
  function readAck(now) {
315
391
  const at = Number.isFinite(now) ? now : Date.now();
316
392
  const state = readState();
317
393
  const ack = state.ack;
318
- if (!ack || ack.on !== true) return null;
319
- // Past the reset of the window it was recorded against, the fact is spent.
320
- if (Number.isFinite(ack.resetsAt) && at >= ack.resetsAt) return null;
321
- return ack;
394
+ if (!ack || typeof ack !== 'object' || ack.on !== true) return null;
395
+ // A record with no usable timestamp is not a record. It printed as
396
+ // "acknowledged ... at Invalid Date" before this, off a hand-edited or
397
+ // half-written file, which is the plugin quoting garbage back as fact.
398
+ if (!Number.isFinite(ack.at)) return null;
399
+ // Past the reset of the window it was recorded against - or past the length
400
+ // of that window, when no reset was readable - the fact is spent.
401
+ const expiresAt = ackExpiry(ack);
402
+ if (Number.isFinite(expiresAt) && at >= expiresAt) return null;
403
+ return Object.assign({}, ack, { expiresAt, expiryKnown: Number.isFinite(ack.resetsAt) });
322
404
  }
323
405
 
324
406
  // Exactly three states, and activeKnown is false in every one of them.
@@ -388,6 +470,7 @@ function forBrief(input) {
388
470
  weeklyPercent: weekly && Number.isFinite(weekly.percentUsed) ? Math.round(weekly.percentUsed) : null,
389
471
  fiveHourPercent: five && Number.isFinite(five.percentUsed) ? Math.round(five.percentUsed) : null,
390
472
  resetGrant: sessionReset(account).present,
473
+ resetVariant: sessionReset(account).variant,
391
474
  credits: credits(account),
392
475
  autoContinue: autoContinue(),
393
476
  };
@@ -396,6 +479,7 @@ function forBrief(input) {
396
479
  module.exports = {
397
480
  FLAG,
398
481
  RESET_FLAG,
482
+ SESSION_RESET_FLAG,
399
483
  WRAPUP_MODE_FLAG,
400
484
  WRAPUP_TEXT_FLAG,
401
485
  NEAR_WALL_FLAG,
@@ -404,12 +488,15 @@ module.exports = {
404
488
  MAX_COOLOFF_MINUTES,
405
489
  WEEKLY_HEADROOM_MAX,
406
490
  WALL_PERCENT,
491
+ ACK_MAX_MS,
492
+ ackExpiry,
407
493
  configDir,
408
494
  stateFile,
409
495
  accountFiles,
410
496
  snapshot,
411
497
  offer,
412
498
  sessionReset,
499
+ flagEnabled,
413
500
  wrapUp,
414
501
  credits,
415
502
  autoContinue,
@@ -1057,9 +1057,15 @@ function tierNow(options) {
1057
1057
  settings = null;
1058
1058
  }
1059
1059
 
1060
+ // The newest assistant message is the model that actually answered, which a
1061
+ // settings alias ('opus') cannot say. The hook's own transcript_path is read
1062
+ // first when the caller has it: it is the file this very turn is written to,
1063
+ // where the session-id lookup has to find it under the config directory.
1060
1064
  let running = null;
1061
1065
  try {
1062
- const seen = sessionId ? usage.liveModel(sessionId) : null;
1066
+ let seen = null;
1067
+ if (opts.transcriptPath && typeof usage.transcriptModel === 'function') seen = usage.transcriptModel(opts.transcriptPath);
1068
+ if (!(seen && seen.model) && sessionId) seen = usage.liveModel(sessionId);
1063
1069
  running = seen && seen.model ? seen.model : null;
1064
1070
  } catch (err) {
1065
1071
  running = null;
@@ -1112,6 +1118,150 @@ function sameFamily(a, b) {
1112
1118
  return left === right;
1113
1119
  }
1114
1120
 
1121
+ // The model as the brief names it: the family and the version. It used to be
1122
+ // the family alone, so claude-opus-5 and claude-opus-5-5 both printed "opus"
1123
+ // and a switch between them - a 20% price change, 60% on cache reads - was
1124
+ // invisible in the one line that exists to say what is producing the turn. A
1125
+ // bare alias ('opus') has no version to show and is shown as it is.
1126
+ function modelLabel(name) {
1127
+ if (!name) return null;
1128
+ let parsed = null;
1129
+ try {
1130
+ parsed = require('./usage.js').parseModelId(name);
1131
+ } catch (err) {
1132
+ parsed = null;
1133
+ }
1134
+ if (parsed) return parsed.version.length ? parsed.family + ' ' + parsed.version.join('.') : parsed.family;
1135
+ const rank = modelRank(name);
1136
+ return rank === null ? name : MODEL_ORDER[rank];
1137
+ }
1138
+
1139
+ // Same model for the purpose of "is the running tier the baseline". Same
1140
+ // family, and where both sides carry a version, the same version: opus 5 and
1141
+ // opus 5.5 are different models at different prices, but a baseline of the
1142
+ // bare alias 'opus' says nothing about which Opus, so it disagrees with none.
1143
+ function sameModel(a, b) {
1144
+ if (!sameFamily(a, b)) return false;
1145
+ let left = null;
1146
+ let right = null;
1147
+ try {
1148
+ const usage = require('./usage.js');
1149
+ left = usage.parseModelId(a);
1150
+ right = usage.parseModelId(b);
1151
+ } catch (err) {
1152
+ return true;
1153
+ }
1154
+ if (!left || !right || !left.version.length || !right.version.length) return true;
1155
+ return left.version.join('.') === right.version.join('.');
1156
+ }
1157
+
1158
+ function money(value) {
1159
+ return Number.isInteger(value) ? '$' + value : '$' + value.toFixed(2);
1160
+ }
1161
+
1162
+ // A newer model in the SAME family at a lower price is the cheapest saving
1163
+ // there is: no step down in tier, nothing given up, only the price. The advice
1164
+ // channel above only ever says "choose lower"; this says "choose newer", once
1165
+ // per session (brief.js keeps the count), and only from prices on record.
1166
+ //
1167
+ // What the user can do differs by host, and the sentence says only what is
1168
+ // real where it is read. In Claude Code, `/model <id>` switches the session and
1169
+ // saves the model as the default for new sessions (the picker's `s` key is the
1170
+ // this-session-only form), per code.claude.com/docs/en/model-config, read
1171
+ // 2026-09-25. A subagent with no model of its own falls through to the main
1172
+ // conversation's model (docs/en/sub-agents, same day), so a switch reaches
1173
+ // those from their next dispatch. Codex has no /model of that kind, so there
1174
+ // the only lever named is its config.
1175
+ //
1176
+ // The payback figure is exact arithmetic on the two rows, not an estimate of
1177
+ // the session: switching costs one write of the context at the new model's
1178
+ // write price instead of one read at the old read price, and every later turn
1179
+ // saves the difference in read price on that context. The context size cancels
1180
+ // out of the ratio, so the number of turns holds for any session. It counts the
1181
+ // reads alone; the cheaper input and output only shorten it.
1182
+ function newerModelAdvice(tier, options) {
1183
+ const opts = options || {};
1184
+ if (!tier) return null;
1185
+ const usage = opts.usage || require('./usage.js');
1186
+ if (typeof usage.newerSibling !== 'function') return null;
1187
+ const base = tier.baseline || {};
1188
+ const run = tier.running || {};
1189
+ const hostName = opts.host || host.CLAUDE;
1190
+ // A bound the user set outranks the price: nothing said points outside it.
1191
+ const sibling = (model) => {
1192
+ const found = model ? usage.newerSibling(model) : null;
1193
+ return found && allows(opts.bounds, { model: found.to }) ? found : null;
1194
+ };
1195
+ const current = run.model || base.model;
1196
+ const found = sibling(current);
1197
+ // The relay resumes with its own --model when one is set (wake.js), so a
1198
+ // session moved to the newer model still wakes on the old one. Claude Code
1199
+ // only: that is where `relay model` feeds a Claude CLI.
1200
+ const relayFound = hostName === host.CLAUDE && opts.relayModel ? sibling(opts.relayModel) : null;
1201
+ if (!found && !relayFound) return null;
1202
+
1203
+ const name = (family, version) => family.charAt(0).toUpperCase() + family.slice(1) + ' ' + version.join('.');
1204
+ const pct = (was, now) => Math.round((1 - now / was) * 100);
1205
+ const prices = (pair) => {
1206
+ const a = pair.fromRate;
1207
+ const b = pair.toRate;
1208
+ const bits = [];
1209
+ if (b.input < a.input || b.output < a.output) {
1210
+ bits.push(money(b.input) + '/' + money(b.output) + ' per million tokens in and out against ' + money(a.input) + '/' + money(a.output));
1211
+ }
1212
+ if (b.cacheRead < a.cacheRead) {
1213
+ bits.push('cache reads ' + money(b.cacheRead) + ' against ' + money(a.cacheRead) + ', ' + pct(a.cacheRead, b.cacheRead) +
1214
+ '% less on the reads that are most of what a long session spends');
1215
+ }
1216
+ return bits.join('; ');
1217
+ };
1218
+ const family = (pair) => pair.family.charAt(0).toUpperCase() + pair.family.slice(1);
1219
+
1220
+ const sentences = [];
1221
+ if (found) {
1222
+ const a = found.fromRate;
1223
+ const b = found.toRate;
1224
+ sentences.push(name(found.family, found.toVersion) + ' is a newer ' + family(found) + ' at a lower price than the ' +
1225
+ name(found.family, found.fromVersion) + ' running here. At first-party API prices: ' + prices(found) +
1226
+ '. Same tier and a newer release, so no step down.');
1227
+ const pinnedAlready = base.model && String(base.model).toLowerCase().replace(/\[[^\]]*\]\s*$/, '').trim() === found.to;
1228
+ if (hostName === host.CLAUDE) {
1229
+ sentences.push('It is the user\'s switch, not yours: offer `/model ' + found.to + '`, which moves this session and saves it as the ' +
1230
+ 'default for new sessions' + (pinnedAlready ? ' (settings.json already names it, so new sessions start on it either way)' : '') +
1231
+ '. Subagents given no model of their own (and no CLAUDE_CODE_SUBAGENT_MODEL) run on the session\'s model, so they follow from their next dispatch.');
1232
+ // Only where the switch can be made mid-session is its one-off cost
1233
+ // worth a number; elsewhere the offer is a pin for new sessions, which
1234
+ // rebuild anyway.
1235
+ if (b.cacheRead < a.cacheRead) {
1236
+ const turns = (write) => Math.ceil(Math.round(((write * b.input - a.cacheRead) / (a.cacheRead - b.cacheRead)) * 100) / 100);
1237
+ sentences.push('Switching rebuilds the prompt cache once; the cheaper reads alone repay that in about ' + turns(1.25) +
1238
+ ' turns (' + turns(2) + ' on the one-hour cache).');
1239
+ } else {
1240
+ sentences.push('Switching rebuilds the prompt cache once, so on a large context it is worth making at the next session start rather than now.');
1241
+ }
1242
+ } else if (hostName === host.CODEX) {
1243
+ sentences.push('It is the user\'s setting, not yours: offer pinning model = "' + found.to + '" in config.toml for new sessions.');
1244
+ } else {
1245
+ sentences.push('It is the user\'s setting, not yours: offer choosing ' + found.to + ' in this host\'s own model setting.');
1246
+ }
1247
+ }
1248
+ if (relayFound) {
1249
+ const same = found && found.to === relayFound.to && found.from === relayFound.from;
1250
+ sentences.push('Relay wakes are pinned to ' + opts.relayModel + ' on their own' +
1251
+ (same ? '' : ', and ' + name(relayFound.family, relayFound.toVersion) + ' is a newer ' + family(relayFound) +
1252
+ ' at a lower price (' + prices(relayFound) + ')') +
1253
+ ', so a resumed run starts on the older model even after this session switches: offer `/usage-limits:relay model ' + relayFound.to +
1254
+ '`, which is the user\'s setting too.');
1255
+ }
1256
+ const id = (found ? found.from + '->' + found.to : '') + (relayFound ? '|relay:' + relayFound.from + '->' + relayFound.to : '');
1257
+ return {
1258
+ id,
1259
+ from: found ? found.from : null,
1260
+ to: found ? found.to : null,
1261
+ relay: relayFound ? { from: relayFound.from, to: relayFound.to } : null,
1262
+ text: sentences.join(' '),
1263
+ };
1264
+ }
1115
1265
  // One clause for the brief, or two lines for `--baseline`.
1116
1266
  //
1117
1267
  // Where the baseline and the running tier agree there is nothing interesting
@@ -1124,16 +1274,12 @@ function tierLine(tier, options) {
1124
1274
  if (!tier) return null;
1125
1275
  const base = tier.baseline || {};
1126
1276
  const run = tier.running || {};
1127
- const shortModel = (name) => {
1128
- const rank = modelRank(name);
1129
- return rank === null ? name : MODEL_ORDER[rank];
1130
- };
1131
- const runningText = [shortModel(run.model) || shortModel(base.model), run.effort || base.effort].filter(Boolean).join('/');
1277
+ const runningText = [modelLabel(run.model) || modelLabel(base.model), run.effort || base.effort].filter(Boolean).join('/');
1132
1278
  if (!runningText) return null;
1133
- const baseText = [shortModel(base.model), base.effort].filter(Boolean).join('/');
1279
+ const baseText = [modelLabel(base.model), base.effort].filter(Boolean).join('/');
1134
1280
  const differs =
1135
1281
  baseText && runningText !== baseText &&
1136
- (!sameFamily(run.model || base.model, base.model) || (run.effort || base.effort) !== base.effort);
1282
+ (!sameModel(run.model || base.model, base.model) || (run.effort || base.effort) !== base.effort);
1137
1283
  const source = run.source ? ' (' + run.source + ')' : '';
1138
1284
  if (opts.terse) {
1139
1285
  return differs ? runningText + source + ', yours ' + baseText : runningText + source;
@@ -1505,7 +1651,12 @@ function main(argv) {
1505
1651
  const usage = require('./usage.js');
1506
1652
  const windows = usage.snapshotWindows(usage.collect(now), now, null) || [];
1507
1653
  const five = windows.find((w) => w && w.key === 'five_hour');
1508
- if (five && Number.isFinite(five.resetsAt)) {
1654
+ // Only a reset still ahead. A stale snapshot carries the reset of a
1655
+ // window that has already rolled over, and stamping that on the record
1656
+ // printed "Recorded ... lapses at" a time already gone - and the very
1657
+ // next read found it spent, so the statement vanished. The next reset
1658
+ // is unknown then, and the record takes the five-hour ceiling instead.
1659
+ if (five && Number.isFinite(five.resetsAt) && five.resetsAt > now) {
1509
1660
  windowKey = 'five_hour';
1510
1661
  resetsAt = five.resetsAt;
1511
1662
  }
@@ -1513,7 +1664,18 @@ function main(argv) {
1513
1664
  // No reading is not a reason to refuse the user's own statement; the
1514
1665
  // record simply has no expiry to hang on.
1515
1666
  }
1516
- lowpri.acknowledge({ on: true, windowKey, resetsAt, now });
1667
+ const written = lowpri.acknowledge({ on: true, windowKey, resetsAt, now });
1668
+ // A write that did not happen is never reported as one. With a directory
1669
+ // sitting where the state file goes this said "Recorded" and the next
1670
+ // brief behaved as though nothing had been said.
1671
+ if (!written.saved) {
1672
+ return [
1673
+ 'NOT recorded: ' + written.file + ' could not be written, so nothing was saved and the ' +
1674
+ 'brief will go on treating the 5-hour window as a real wall.',
1675
+ '',
1676
+ 'Check that the path is a writable file and not a directory, then run this again.',
1677
+ ].join('\n');
1678
+ }
1517
1679
  logChange({ plane: 'mode', key: 'low-priority', from: 'unknown', to: 'on', by: 'user', reason: null }, now);
1518
1680
  return [
1519
1681
  'Recorded: you have switched /low-priority on.',
@@ -1530,14 +1692,25 @@ function main(argv) {
1530
1692
  '',
1531
1693
  resetsAt
1532
1694
  ? 'This lapses on its own at ' + new Date(resetsAt).toLocaleString() + ', when the 5-hour window resets.'
1533
- : 'No 5-hour reset time was readable, so this has no expiry: clear it by hand when it ends.',
1695
+ : 'No 5-hour reset time was readable, so this lapses at ' +
1696
+ new Date(now + lowpri.ACK_MAX_MS).toLocaleString() + ' instead - five hours, the length of ' +
1697
+ 'the window the toggle belongs to, which is the longest it could honestly still be true. ' +
1698
+ 'Run /usage and set it again if it is still on then.',
1534
1699
  '',
1535
1700
  detail,
1536
1701
  ].join('\n');
1537
1702
  }
1538
1703
  if (asked === 'off' || asked === 'no' || asked === 'false') {
1539
1704
  const had = lowpri.readAck(now);
1540
- lowpri.acknowledge({ on: false, now });
1705
+ const cleared = lowpri.acknowledge({ on: false, now });
1706
+ // The same rule as "on": a clear that did not reach the file is not
1707
+ // reported as one, because the record it failed to remove still steers
1708
+ // the brief at the weekly.
1709
+ if (had && !cleared.saved) {
1710
+ return 'NOT cleared: ' + cleared.file + ' could not be written, so the acknowledgement is still ' +
1711
+ 'there and the brief will go on braking on the weekly until it lapses at ' +
1712
+ new Date(had.expiresAt).toLocaleString() + '. Check that the file is writable, then run this again.';
1713
+ }
1541
1714
  logChange({ plane: 'mode', key: 'low-priority', from: had ? 'on' : 'unknown', to: 'off', by: 'user', reason: null }, now);
1542
1715
  return had
1543
1716
  ? 'Cleared: low-priority is off again, so the 5-hour window counts as a wall from the next prompt.'
@@ -1550,9 +1723,15 @@ function main(argv) {
1550
1723
  const wrap = lowpri.wrapUp(account);
1551
1724
  const auto = lowpri.autoContinue();
1552
1725
  return [
1726
+ // readAck drops a record with no usable timestamp, so there is no longer
1727
+ // a path that prints "at Invalid Date" here, and the expiry is always a
1728
+ // real time: the window's reset, or five hours from the statement.
1553
1729
  ack
1554
1730
  ? 'You have acknowledged low-priority ON, at ' + new Date(ack.at).toLocaleString() +
1555
- (Number.isFinite(ack.resetsAt) ? ', lapsing at ' + new Date(ack.resetsAt).toLocaleString() : ', with no expiry') + '.'
1731
+ (Number.isFinite(ack.expiresAt)
1732
+ ? ', lapsing at ' + new Date(ack.expiresAt).toLocaleString() +
1733
+ (ack.expiryKnown ? ' when the 5-hour window resets' : ' (no reset time was readable, so this is five hours from the statement)')
1734
+ : '') + '.'
1556
1735
  : 'Low-priority has not been acknowledged, so the plugin treats the 5-hour window as a real wall.',
1557
1736
  '',
1558
1737
  'It is a toggle you type yourself. The model cannot run a slash command, and /low-priority is ' +
@@ -1843,6 +2022,9 @@ module.exports = {
1843
2022
  ultracodeName,
1844
2023
  topTier,
1845
2024
  tierLine,
2025
+ modelLabel,
2026
+ sameModel,
2027
+ newerModelAdvice,
1846
2028
  ledger,
1847
2029
  explain,
1848
2030
  list,
@@ -1496,11 +1496,30 @@ async function armByHand(rest) {
1496
1496
  });
1497
1497
  if (!result.ok) return 'Could not arm: ' + result.error;
1498
1498
  if (text) saveContinuation(sessionId, text);
1499
+ // armable() refuses a 5-hour wake while low-priority is acknowledged, and
1500
+ // this path never went through it: `relay arm --session <id>` booked the wake
1501
+ // anyway and said nothing, which is the double-start the whole gate exists to
1502
+ // prevent. An explicit command is still honoured - it is what somebody typed
1503
+ // - but it says what `defer reset` says, so the person deciding has the fact.
1504
+ let lowPriorityNote = '';
1505
+ if (binding.key === 'five_hour' && usage.currentHost() === host.CLAUDE) {
1506
+ try {
1507
+ if (lowpri.readAck(now)) {
1508
+ lowPriorityNote =
1509
+ ' Note: you have said /low-priority is on, so this session carries past the 5-hour reset ' +
1510
+ 'on its own and this wake would start the work a second time. Arm against the weekly ' +
1511
+ 'instead, or say usage-mode --low-priority off if it has ended; "relay cancel" calls this one off.';
1512
+ }
1513
+ } catch (err) {
1514
+ // A missing or unreadable record is not a reason to refuse a command.
1515
+ }
1516
+ }
1499
1517
  return (
1500
1518
  'Armed by hand for session ' + sessionId.slice(0, 8) + ': wake at ' +
1501
1519
  new Date(result.record.wakeAt).toLocaleString() + ' via ' + result.record.how + '.' +
1502
1520
  (result.record.preflight && result.record.preflight.length ? ' Pre-answered: ' + result.record.preflight.join(', ') + '.' : '') +
1503
- (text ? ' Continuation saved.' : ' No continuation yet - add one with: relay note "<text>"')
1521
+ (text ? ' Continuation saved.' : ' No continuation yet - add one with: relay note "<text>"') +
1522
+ lowPriorityNote
1504
1523
  );
1505
1524
  }
1506
1525
 
@@ -62,6 +62,10 @@ const RATES = {
62
62
  'claude-mythos-5-1': { input: 10, output: 50, cacheRead: 0.25 },
63
63
  'claude-fable-5': { input: 10, output: 50 },
64
64
  'claude-mythos-5': { input: 10, output: 50 },
65
+ // Opus 5.5 prices reads outright: $0.20 is 0.05x its input, half the tenth
66
+ // rule every other Opus follows (platform.claude.com/docs/en/about-claude/pricing,
67
+ // read 2026-09-25). Left to the multiplier it would be priced at $0.40.
68
+ 'claude-opus-5-5': { input: 4, output: 20, cacheRead: 0.2 },
65
69
  'claude-opus-5': { input: 5, output: 25 },
66
70
  'claude-opus-4-8': { input: 5, output: 25 },
67
71
  'claude-opus-4-7': { input: 5, output: 25 },
@@ -129,6 +133,83 @@ const CACHE_WRITE_5M = 1.25;
129
133
  const CACHE_WRITE_1H = 2;
130
134
  const CACHE_READ = 0.1;
131
135
 
136
+ // The effective cache-read price of a row, in $/MTok.
137
+ function readRateOf(rate) {
138
+ return Number.isFinite(rate.cacheRead) ? rate.cacheRead : rate.input * CACHE_READ;
139
+ }
140
+
141
+ // A model id taken apart into the family word and the version numbers:
142
+ // 'claude-opus-5-5' is opus [5, 5], 'us.anthropic.claude-opus-5-v1:0' is
143
+ // opus [5]. A dated snapshot suffix (20250929) is not part of the version, and a
144
+ // bare alias such as 'opus' has no version at all, so nothing is claimed about
145
+ // which release it resolved to. Mythos stays mythos here: it is priced with
146
+ // Fable, but it is not the same model line.
147
+ function parseModelId(model) {
148
+ const id = normalizeModel(model);
149
+ const found = id.match(/(haiku|sonnet|opus|mythos|fable)((?:-\d+)*)/);
150
+ if (!found) return null;
151
+ const version = [];
152
+ for (const part of found[2].split('-').filter(Boolean)) {
153
+ if (part.length > 2) break;
154
+ version.push(Number(part));
155
+ }
156
+ return { family: found[1], version };
157
+ }
158
+
159
+ function compareVersions(a, b) {
160
+ const length = Math.max(a.length, b.length);
161
+ for (let i = 0; i < length; i += 1) {
162
+ const diff = (a[i] || 0) - (b[i] || 0);
163
+ if (diff !== 0) return diff;
164
+ }
165
+ return 0;
166
+ }
167
+
168
+ // Claude 4.7 and later count text with a newer tokenizer that produces about
169
+ // 30% more tokens for the same text (pricing page, read 2026-09-25). A price per
170
+ // token across that line is not a like-for-like price, so a sibling on the other
171
+ // side of it is never offered as the cheaper one.
172
+ function sameTokenizer(a, b) {
173
+ return (compareVersions(a, [4, 7]) >= 0) === (compareVersions(b, [4, 7]) >= 0);
174
+ }
175
+
176
+ // The newest release of the SAME family that is newer than this one and no
177
+ // dearer on any rate - input, output or cache reads - and cheaper on at least
178
+ // one. That is a saving with no step down in tier. Only a model whose own price
179
+ // is on record qualifies: comparing against a family average would be
180
+ // comparing against a guess.
181
+ function newerSibling(model, table) {
182
+ const rates = table || RATES;
183
+ const me = parseModelId(model);
184
+ if (!me || !me.version.length) return null;
185
+ const from = 'claude-' + me.family + '-' + me.version.join('-');
186
+ const mine = rates[from];
187
+ if (!mine) return null;
188
+ let best = null;
189
+ for (const id of Object.keys(rates)) {
190
+ const other = parseModelId(id);
191
+ if (!other || other.family !== me.family || !other.version.length) continue;
192
+ if (compareVersions(other.version, me.version) <= 0) continue;
193
+ if (!sameTokenizer(other.version, me.version)) continue;
194
+ const theirs = rates[id];
195
+ const noDearer = theirs.input <= mine.input && theirs.output <= mine.output && readRateOf(theirs) <= readRateOf(mine);
196
+ const cheaper = theirs.input < mine.input || theirs.output < mine.output || readRateOf(theirs) < readRateOf(mine);
197
+ if (!noDearer || !cheaper) continue;
198
+ if (!best || compareVersions(other.version, best.version) > 0) best = { id, version: other.version, rate: theirs };
199
+ }
200
+ if (!best) return null;
201
+ const shape = (rate) => ({ input: rate.input, output: rate.output, cacheRead: readRateOf(rate) });
202
+ return {
203
+ family: me.family,
204
+ from,
205
+ fromVersion: me.version,
206
+ fromRate: shape(mine),
207
+ to: best.id,
208
+ toVersion: best.version,
209
+ toRate: shape(best.rate),
210
+ };
211
+ }
212
+
132
213
  // `family` marks a window that caps one model family rather than the account
133
214
  // as a whole. It is what tells the rest of the file that a window cannot stop
134
215
  // work which does not use that family.
@@ -4282,6 +4363,8 @@ module.exports = {
4282
4363
  WINDOWS,
4283
4364
  rateFor,
4284
4365
  familyOf,
4366
+ parseModelId,
4367
+ newerSibling,
4285
4368
  familyAverage,
4286
4369
  familiesInUse,
4287
4370
  appliesTo,
@@ -559,8 +559,17 @@ function mine(held, record) {
559
559
  }
560
560
 
561
561
  async function run(now, argv, overrides) {
562
+ // `reachable` belongs in here with the rest. Left out, the preflight below
563
+ // called the real network on every run of the retry tests, so the whole
564
+ // launch-failure suite passed only on a machine that could reach
565
+ // api.anthropic.com - and went red, in the offline branch, on a laptop with
566
+ // its wifi off or behind a TLS-inspecting proxy. That is the exact condition
567
+ // the relay is built for, so it is the last thing its tests should need.
562
568
  const deps = Object.assign(
563
- { windowReopened, deliverClaude, deliverCodex, toast, userIsPresent, arm: relay.arm, capabilities: relay.capabilities },
569
+ {
570
+ windowReopened, deliverClaude, deliverCodex, toast, userIsPresent,
571
+ arm: relay.arm, capabilities: relay.capabilities, reachable: net.reachable,
572
+ },
564
573
  overrides || null
565
574
  );
566
575
  const id = argOf(argv, '--id');
@@ -618,7 +627,7 @@ async function run(now, argv, overrides) {
618
627
  // So: ask first, before spending a launch on it, and give being offline its
619
628
  // own much longer budget. A machine that cannot reach the API has not failed.
620
629
  // It is waiting, and waiting is free.
621
- const link = await net.reachable({ timeoutMs: 8000 });
630
+ const link = await deps.reachable({ timeoutMs: 8000 });
622
631
  if (!link.online) {
623
632
  const offlineAttempt = (record.offlineAttempt || 0) + 1;
624
633
  if (offlineAttempt <= config.offlineAttempts) {