pi-lxmf 0.1.3 → 0.2.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -7,6 +7,28 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
7
7
 
8
8
  ## [Unreleased]
9
9
 
10
+ ## [0.2.0] - 2026-09-27
11
+
12
+ ### Added
13
+
14
+ - **Quota exhaustion and 90% warnings** (work doc #3): the GLM quota
15
+ watcher now samples the z.ai quota API continuously while a GLM model
16
+ is active (not just after a run fails) and notifies the owner when a
17
+ bucket (5h or weekly) hits 100% — including at startup, when the daemon
18
+ starts mid-outage — and when one crosses 90%. The startup check runs
19
+ immediately on enable, so no error is needed to detect an exhausted
20
+ bucket.
21
+ - The startup notification now names the active model (fresh `get_state`
22
+ observation) — the model isn't visible anywhere else over LXMF.
23
+ - `/model` and `/think` re-observe state right away, so the quota watcher
24
+ gate follows the switch immediately instead of on the next prompt.
25
+
26
+ ### Fixed
27
+
28
+ - A run that failed after delivering partial output silently trailed
29
+ off: the error is now reported as a "run ended early" message and no
30
+ longer leaks into a later exchange's failure reply.
31
+
10
32
  ## [0.1.3] - 2026-09-27
11
33
 
12
34
  ### Changed
package/README.md CHANGED
@@ -88,9 +88,10 @@ node src/bin.js --help
88
88
  |---|---|
89
89
  | plain text | a prompt to the agent (steered into a running turn by default) |
90
90
  | `/help` | bridge commands + Pi commands available via prompt |
91
- | `/status` | model, thinking, session, uptime, node + owner identity hashes |
91
+ | `/status` | model, thinking, session, uptime, cwd + workdir, node + owner identity hashes |
92
92
  | `/session` | message counts, tokens, cost, context usage |
93
93
  | `/new` | fresh Pi session |
94
+ | `/cd [path]` | switch the supervised Pi to another repo under `workdir`; bare `/cd` lists current + recent repos |
94
95
  | `/name [name]` | show / set the session display name |
95
96
  | `/compact [instructions]` | compact the conversation context |
96
97
  | `/model [query]` | list models, or switch (`/model sonnet`) |
@@ -105,6 +106,15 @@ Replies are delivered per finished assistant message, chunked to fit
105
106
  then a `✅ done (no reply)` nudge. Extension dialogs raised inside Pi are
106
107
  auto-declined (nobody is at a terminal) and reported to you.
107
108
 
109
+ `/cd` makes one bridge serve every repo under `workdir`: switching is a
110
+ supervised respawn in the target directory (messages queue during the switch),
111
+ each repo keeps its own Pi session — revisiting one resumes its conversation —
112
+ and the active repo is remembered across daemon restarts. Targets must resolve
113
+ under `workdir` (no `..` traversal, nothing outside the tree); anything else
114
+ is refused without touching the child. `/status` shows the current `cwd` and
115
+ `workdir`, and each switched-to project must be trusted once (see step 4
116
+ above) for its local `.pi` resources to load.
117
+
108
118
  ## Configuration
109
119
 
110
120
  `~/.config/pi-lxmf/config.json` (XDG env vars respected). A missing file runs
@@ -114,7 +124,7 @@ on defaults.
114
124
  |---|---|---|
115
125
  | `owner` | *(required)* | the owner's 32-hex **Reticulum identity hash** (not the LXMF address); the daemon derives the `lxmf.delivery` destination hash for wire comparison |
116
126
  | `name` | `pi-lxmf <version>` | announce display name |
117
- | `workdir` | daemon cwd | project directory Pi runs in (also where `AGENTS.md` is found) |
127
+ | `workdir` | daemon cwd | project directory Pi runs in (also where `AGENTS.md` is found); the trust root `/cd` cannot escape |
118
128
  | `model` | Pi default | `--model` pattern passed to Pi |
119
129
  | `piBin` | `pi` | Pi binary |
120
130
  | `dataDir` | `~/.local/share/pi-lxmf` | state root (see below) |
package/SPEC.md CHANGED
@@ -308,19 +308,30 @@ the recently used ones (derived from the per-workdir session pointers);
308
308
  ### 6.7 z.ai GLM quota watcher and peak-hours warning
309
309
 
310
310
  When the active model is a z.ai GLM model (`provider === "zai"`), the bridge
311
- runs a `GlmQuotaWatcher` (`src/quota.js`) that does two things, both gated
312
- on the active model being GLM — nothing fires for non-z.ai providers (e.g.
313
- Cortecs, Anthropic):
314
-
315
- - **Quota-recovery notification.** When a run fails with a z.ai
316
- quota-exhausted error (matched from `auto_retry_end`/`compaction_end`
317
- error messages), the watcher polls the same z.ai quota endpoint
318
- `pi-glm-usage` uses (`https://api.z.ai/api/monitor/usage/quota/limit`,
319
- Bearer `~/.pi/agent/auth.json` → `zai.key`, honouring `PI_AUTH_DIR`) every
320
- 60s and delivers **exactly one** LXMF message to the owner the moment
321
- the 5h bucket drops below 100%. One notification per exhausted episode;
322
- rate-limit-only errors (transient, retried by Pi) do not arm it. A
323
- missing `zai.key` disables the watcher gracefully (logged once).
311
+ runs a `GlmQuotaWatcher` (`src/quota.js`), gated on the active model being
312
+ GLM — nothing fires for non-z.ai providers (e.g. Cortecs, Anthropic). The
313
+ gate is (re-)evaluated on every `get_state` observation: prompts, the
314
+ startup observation, and after `/model`/`/think` commands.
315
+
316
+ - **Continuous quota sampling.** While enabled, the watcher polls the
317
+ same z.ai quota endpoint `pi-glm-usage` uses
318
+ (`https://api.z.ai/api/monitor/usage/quota/limit`, Bearer
319
+ `~/.pi/agent/auth.json` → `zai.key`, honouring `PI_AUTH_DIR`) every 60s
320
+ (first sample immediately on enable — so a daemon that starts mid-outage
321
+ detects and reports it right away). From those samples the owner is
322
+ notified of:
323
+ - **Exhaustion** — a bucket (5h or weekly) reaching 100%: once per
324
+ episode, with the reset time when the API reports one.
325
+ - **90% warning** — a bucket at or above 90% while below 100%: once
326
+ per window (a dip below the threshold re-arms it).
327
+ - **Recovery** — the 5h bucket dropping below 100% after an exhausted
328
+ episode: once per episode.
329
+ A quota-looking Pi error (`auto_retry_end`/`compaction_end`) forces an
330
+ immediate fresh sample; rate-limit-only errors (transient, retried by
331
+ Pi) are ignored. Per-sample notices are joined into one LXMF message
332
+ and — when an owner-triggered run is live — deferred to `agent_settled`
333
+ so they never interleave with a reply. A missing `zai.key` disables the
334
+ watcher gracefully (logged once).
324
335
  - **Peak-hours warning.** z.ai charges 3× tokens Mon–Fri 14:00–18:00
325
336
  Singapore Standard Time (UTC+8). The owner is warned when an
326
337
  owner-triggered run starts inside that window, and when the window
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "pi-lxmf",
3
- "version": "0.1.3",
3
+ "version": "0.2.0",
4
4
  "description": "Drive the Pi coding agent over LXMF messaging (Reticulum mesh) from a headless server.",
5
5
  "license": "EUPL-1.2",
6
6
  "author": "Henri Bergius <henri.bergius@iki.fi>",
package/src/bridge.js CHANGED
@@ -315,6 +315,18 @@ export class Bridge {
315
315
  this.log.log(`pi-lxmf: inbound from owner: command /${parsed.name}`);
316
316
  try {
317
317
  const result = await command.run(this.commandContext(), parsed.args);
318
+ // /model and /think change pi state the bridge tracks through
319
+ // `get_state` observations — re-observe so the active model (and
320
+ // the GLM watcher gate) follow immediately instead of on the
321
+ // next prompt. Other commands manage their own observations
322
+ // (e.g. /cd) or don't affect tracked state.
323
+ if (parsed.name === "model" || parsed.name === "think") {
324
+ try {
325
+ this.observeState(await this.rpc.getState());
326
+ } catch {
327
+ /* keep stale observations; next prompt re-observes */
328
+ }
329
+ }
318
330
  const text = typeof result === "string" ? result : result?.text;
319
331
  if (text) await this.deliver(text);
320
332
  if (result && typeof result === "object" && result.shutdown) {
@@ -466,6 +478,14 @@ export class Bridge {
466
478
 
467
479
  if (sent > 0) {
468
480
  this.recovering = false;
481
+ // The owner already got (partial) output. A trailing failure means
482
+ // the run died mid-reply: say so — otherwise the exchange just
483
+ // trails off — and never let the error leak into a later exchange.
484
+ if (this.lastError) {
485
+ const error = this.lastError;
486
+ this.lastError = null;
487
+ await this.deliver(`⚠️ run ended early: ${error}`);
488
+ }
469
489
  return;
470
490
  }
471
491
  if (this.lastError) {
@@ -655,17 +675,29 @@ export class Bridge {
655
675
  * Tells the owner the bridge has started and is accepting messages
656
676
  * (the startup case of SPEC §13 proactive notifications). Called by the
657
677
  * daemon once the mesh side is announcing and the RPC child is ready.
658
- * Best-effort via {@link deliver}: a failure is noted and carried by
659
- * the next successful delivery instead of being lost.
678
+ * The active model is included (fresh `get_state` observation) — the
679
+ * owner can't see the TUI footer over LXMF, so the startup message is
680
+ * the only place the model is announced proactively. Best-effort via
681
+ * {@link deliver}: a failure is noted and carried by the next
682
+ * successful delivery instead of being lost.
660
683
  *
661
684
  * @param {string|null} [resumedSessionFile] - Absolute path of the
662
685
  * session resumed from the persisted pointer, when one exists.
663
686
  */
664
687
  async notifyStartup(resumedSessionFile = null) {
665
- const resumed = resumedSessionFile
666
- ? `\nResuming session ${basename(resumedSessionFile)}.`
667
- : "";
668
- await this.deliver(`🟢 pi-lxmf ready — listening for messages.${resumed}`);
688
+ try {
689
+ this.observeState(await this.rpc.getState());
690
+ } catch {
691
+ /* RPC child is ready (checked by the caller); keep the old model */
692
+ }
693
+ const lines = ["🟢 pi-lxmf ready — listening for messages."];
694
+ if (typeof this.activeModel?.name === "string" && this.activeModel.name) {
695
+ lines.push(`Model: ${this.activeModel.name}`);
696
+ }
697
+ if (resumedSessionFile) {
698
+ lines.push(`Resuming session ${basename(resumedSessionFile)}.`);
699
+ }
700
+ await this.deliver(lines.join("\n"));
669
701
  }
670
702
 
671
703
  /**
package/src/quota.js CHANGED
@@ -3,15 +3,27 @@
3
3
  *
4
4
  * z.ai (GLM Coding Plan) quota watcher + peak-hours warning (work doc #3).
5
5
  *
6
- * Two related behaviours, both gated on the active model being a z.ai GLM
6
+ * Three related behaviours, all gated on the active model being a z.ai GLM
7
7
  * model (`provider === "zai"`):
8
8
  *
9
- * 1. **Quota-recovery notification.** When a run fails with a z.ai
10
- * quota-exhausted error, poll the z.ai quota endpoint and push exactly
11
- * one LXMF message to the owner the moment the 5h bucket becomes
12
- * available again, so the owner can resume work without babysitting it.
9
+ * 1. **Continuous quota sampling.** While a GLM model is active, the same
10
+ * z.ai quota endpoint `pi-glm-usage` uses is polled every 60s. From
11
+ * those samples the owner is notified of:
12
+ * - **Exhaustion** (once per episode): a bucket (5h or weekly) reaches
13
+ * 100% — including at startup, when a daemon restarts mid-outage.
14
+ * - **90% warning** (once per bucket window): a bucket crosses 90%,
15
+ * while still below 100%.
16
+ * - **Recovery** (once per episode): the 5h bucket drops below 100%
17
+ * after an exhausted episode.
18
+ * Per-sample notifications are joined into a single LXMF message and
19
+ * deferred to `agent_settled` while an owner-triggered run is live.
13
20
  *
14
- * 2. **Peak-hours warning.** z.ai charges 3× tokens during peak hours
21
+ * 2. **Quota-error arming.** A Pi error that looks like a z.ai
22
+ * quota-exhausted failure (`auto_retry_end`/`compaction_end`) arms the
23
+ * poller immediately, even when the live fetch is momentarily
24
+ * inconclusive — the authoritative answer is the next sample.
25
+ *
26
+ * 3. **Peak-hours warning.** z.ai charges 3× tokens during peak hours
15
27
  * (Mon–Fri 14:00–18:00 Singapore Standard Time, UTC+8). Warn the owner
16
28
  * when an owner-triggered run starts inside that window, and when the
17
29
  * window opens mid-run, so they can decide whether to stop or continue.
@@ -30,7 +42,7 @@ import { join } from "node:path";
30
42
  const QUOTA_URL = "https://api.z.ai/api/monitor/usage/quota/limit";
31
43
  /** Fetch timeout for the quota endpoint (ms). */
32
44
  const FETCH_TIMEOUT_MS = 5000;
33
- /** Poll cadence while the 5h bucket is exhausted (ms). */
45
+ /** Sampling cadence while a GLM model is active (ms). */
34
46
  const POLL_INTERVAL_MS = 60_000;
35
47
 
36
48
  /**
@@ -56,9 +68,11 @@ const Unit = {
56
68
  };
57
69
 
58
70
  /**
59
- * The 5h bucket is treated as exhausted at-or-above this percentage.
71
+ * A bucket is treated as exhausted at-or-above this percentage.
60
72
  */
61
73
  const EXHAUSTED_PERCENTAGE = 100;
74
+ /** Warning threshold (crossing upward, while not exhausted). */
75
+ const WARN_PERCENTAGE = 90;
62
76
 
63
77
  /** z.ai provider id (and its `z.ai` alias, defensively). */
64
78
  const ZAI_PROVIDERS = new Set(["zai", "z.ai"]);
@@ -196,7 +210,6 @@ export function isPeakTime(epochMs) {
196
210
  */
197
211
  export function msUntilPeakOpen(fromMs) {
198
212
  if (isPeakTime(fromMs)) return 0;
199
- const from = new Date(fromMs);
200
213
  // Walk forward hour by hour (max ~7 days) to the first peak hour.
201
214
  for (let h = 0; h < 24 * 7; h++) {
202
215
  const probe = new Date(fromMs + h * 3_600_000);
@@ -222,11 +235,35 @@ export function msUntilPeakOpen(fromMs) {
222
235
  return Number.POSITIVE_INFINITY;
223
236
  }
224
237
 
238
+ /**
239
+ * A live or pinned quota sample for one bucket. `percentage` may exceed
240
+ * `WARN_PERCENTAGE`/`EXHAUSTED_PERCENTAGE` by z.ai's rounding; comparisons
241
+ * are inclusive.
242
+ *
243
+ * @typedef {{percentage: number, nextResetMs: number}} BucketSample
244
+ */
245
+
225
246
  /**
226
247
  * The z.ai quota watcher + peak-hours warner. Construct one per bridge;
227
248
  * drive it with {@link GlmQuotaWatcher.setEnabled} (gated on the active
228
- * model) and {@link GlmQuotaWatcher.onAgentStart} / {@link
229
- * GlmQuotaWatcher.onAgentSettled} / {@link GlmQuotaWatcher.onError}.
249
+ * model) and {@link GlmQuotaWatcher.onAgentStart} /
250
+ * {@link GlmQuotaWatcher.onAgentSettled} / {@link GlmQuotaWatcher.onError}.
251
+ *
252
+ * While enabled, a single sampler polls the quota endpoint every
253
+ * `pollIntervalMs` (first sample immediately). Each sample can queue at
254
+ * most one joined notice:
255
+ * - a bucket crossing `WARN_PERCENTAGE` (once per bucket window) queues
256
+ * a "90%" warning line;
257
+ * - a bucket reaching `EXHAUSTED_PERCENTAGE` (once per bucket episode)
258
+ * queues an "exhausted" line with the reset time;
259
+ * - the 5h bucket recovering from an exhausted episode (once per episode)
260
+ * queues a "recovered" line.
261
+ *
262
+ * A run-triggered error ({@link GlmQuotaWatcher.onError}) additionally
263
+ * forces an immediate sample (fresh state over stale-sampler lag).
264
+ * Notices are delivered as one LXMF message per emission — immediately
265
+ * when the bridge is idle, or when the live owner-triggered run settles
266
+ * (so they never interleave with a reply mid-run).
230
267
  */
231
268
  export class GlmQuotaWatcher {
232
269
  /**
@@ -238,6 +275,7 @@ export class GlmQuotaWatcher {
238
275
  * @param {typeof fetch} [options.fetchImpl] - Injectable fetch (tests).
239
276
  * @param {() => number} [options.now] - Injectable clock (tests).
240
277
  * @param {number} [options.pollIntervalMs]
278
+ * @param {number} [options.warnPercentage] - Warning threshold (tests).
241
279
  */
242
280
  constructor(options) {
243
281
  this.ownerDestinationHash = options.ownerDestinationHash;
@@ -247,6 +285,7 @@ export class GlmQuotaWatcher {
247
285
  this.fetchImpl = options.fetchImpl || fetch;
248
286
  this.now = options.now || (() => Date.now());
249
287
  this.pollIntervalMs = options.pollIntervalMs ?? POLL_INTERVAL_MS;
288
+ this.warnPercentage = options.warnPercentage ?? WARN_PERCENTAGE;
250
289
 
251
290
  /** Whether the active model is a z.ai GLM model (the master gate). */
252
291
  this.enabled = false;
@@ -254,15 +293,28 @@ export class GlmQuotaWatcher {
254
293
  this.runActive = false;
255
294
  /** Whether the current run is owner-triggered (not recovery). */
256
295
  this.ownerTriggered = false;
257
- /** Single active quota poller (idempotent start). */
296
+ /** Single active sampler (idempotent start). */
258
297
  /** @type {NodeJS.Timeout|null} */
259
298
  this.pollTimer = null;
260
- /** Whether the 5h bucket was exhausted when the poller last sampled. */
299
+ /** Whether the 5h bucket was exhausted in the last sample. */
261
300
  this.wasExhausted = false;
262
- /** Whether a recovery notification has already been sent for this episode. */
263
- this.notifiedThisEpisode = false;
264
301
  /** Epoch (ms) the exhaustion episode started, for the human-readable delta. */
265
302
  this.exhaustedSinceMs = 0;
303
+ /**
304
+ * Whether the 90% warning already fired for each bucket in its current
305
+ * window (a window is any contiguous below-100% stretch — a reset to
306
+ * below-warn clears it, an exhausted episode ends it).
307
+ * @type {{fiveHour: boolean, weekly: boolean}}
308
+ */
309
+ this.warned = { fiveHour: false, weekly: false };
310
+ /** Whether an exhausted episode per bucket already notified (its
311
+ * "exhausted" notice is once per episode).
312
+ * @type {{fiveHour: boolean, weekly: boolean}}
313
+ */
314
+ this.exhaustionNotified = { fiveHour: false, weekly: false };
315
+ /** Notices queued while a run is live, delivered on `onAgentSettled`. */
316
+ /** @type {string[]} */
317
+ this.pendingNotices = [];
266
318
  /** Whether we've already warned about peak for the current run. */
267
319
  this.peakWarnedThisRun = false;
268
320
  /** Timer for the "run ran into peak" boundary warning. */
@@ -276,7 +328,12 @@ export class GlmQuotaWatcher {
276
328
 
277
329
  /**
278
330
  * Master gate: enable/disable based on whether the active model is z.ai.
279
- * Disabling stops any active poller and clears run state.
331
+ * Enabling starts the sampler immediately (first sample right away), so
332
+ * an exhausted state at startup or model switch is detected and
333
+ * notified without waiting for a run or an error. Disabling stops all
334
+ * timers, clears run state, and (best-effort) delivers anything already
335
+ * queued — losing queued notices on shutdown is acceptable, but losing
336
+ * them on a model switch is not.
280
337
  *
281
338
  * @param {boolean} enabled
282
339
  */
@@ -289,10 +346,9 @@ export class GlmQuotaWatcher {
289
346
  this.runActive = false;
290
347
  this.ownerTriggered = false;
291
348
  this.peakWarnedThisRun = false;
349
+ this.flushNotices();
292
350
  } else {
293
- // Gated on at startup: do one quota fetch so a daemon that restarted
294
- // mid-outage arms the watcher immediately.
295
- void this.checkAndMaybeArm(false);
351
+ this.startPoller(true);
296
352
  }
297
353
  }
298
354
 
@@ -315,11 +371,14 @@ export class GlmQuotaWatcher {
315
371
  this.ownerTriggered = false;
316
372
  this.peakWarnedThisRun = false;
317
373
  this.clearPeakOpenTimer();
374
+ this.flushNotices();
318
375
  }
319
376
 
320
377
  /**
321
378
  * Called when a Pi error event (`auto_retry_end`/`compaction_end`) looks
322
- * like a quota failure. Arms the poller (idempotent) when enabled.
379
+ * like a quota failure. Ensures the sampler is running (idempotent) and
380
+ * forces an immediate sample so the exhaustion notice is driven by the
381
+ * authoritative API state, not the sampler's cadence.
323
382
  *
324
383
  * @param {string} errorMessage
325
384
  */
@@ -329,7 +388,8 @@ export class GlmQuotaWatcher {
329
388
  this.log.log(
330
389
  `pi-lxmf: GLM quota error detected, polling for recovery: ${errorMessage}`,
331
390
  );
332
- void this.checkAndMaybeArm(true);
391
+ this.startPoller();
392
+ void this.tick();
333
393
  }
334
394
 
335
395
  /**
@@ -341,114 +401,144 @@ export class GlmQuotaWatcher {
341
401
  }
342
402
 
343
403
  /**
344
- * Fetches the quota once and arms the poller if the 5h bucket is
345
- * exhausted. When `fromError` is true (we were tipped off by a Pi error),
346
- * arm even if the first fetch is inconclusive (transient API failure).
404
+ * Starts the sampler if not already running. `immediate` also fires one
405
+ * sample right away (startup, model switch, quota error) instead of
406
+ * waiting a full interval.
347
407
  *
348
- * @param {boolean} fromError
349
- * @returns {Promise<void>}
408
+ * @param {boolean} [immediate]
409
+ * @private
350
410
  */
351
- async checkAndMaybeArm(fromError) {
411
+ startPoller(immediate) {
412
+ if (this.pollTimer || !this.enabled || !this.apiKey) return;
413
+ this.pollTimer = setInterval(() => void this.tick(), this.pollIntervalMs);
414
+ if (typeof this.pollTimer.unref === "function") this.pollTimer.unref();
415
+ if (immediate) void this.tick();
416
+ }
417
+
418
+ /**
419
+ * One sampler tick: fetch quota, evaluate bucket transitions, queue
420
+ * notices for anything new.
421
+ *
422
+ * @private
423
+ */
424
+ async tick() {
352
425
  if (!this.enabled || !this.apiKey) return;
353
426
  let quota = null;
354
427
  try {
355
428
  quota = await fetchQuota(this.apiKey, { fetchImpl: this.fetchImpl });
356
- } catch (e) {
357
- if (fromError) {
358
- // A Pi error said quota-exhausted; trust it and arm, polling will
359
- // confirm the recovery transition.
360
- this.log.log(
361
- `pi-lxmf: GLM quota fetch failed (${e instanceof Error ? e.message : e}); arming watcher on Pi error signal`,
362
- );
363
- this.armWatcher(0);
364
- }
429
+ } catch {
430
+ /* transient — retry on the next tick */
365
431
  return;
366
432
  }
367
- const pct = quota.fiveHour?.percentage ?? 0;
368
- if (pct >= EXHAUSTED_PERCENTAGE) {
369
- this.armWatcher(pct);
370
- } else if (this.wasExhausted) {
371
- // Recovered between fetches (e.g. daemon was away): notify now.
372
- this.notifyRecovered(quota);
373
- }
433
+ this.evaluateSample(quota);
374
434
  }
375
435
 
376
436
  /**
377
- * Arms the single poller (idempotent). Records the episode start time on
378
- * the first arm of an episode and resets the per-episode notification flag.
437
+ * Turns a fresh quota sample into (at most one) queued notice, based on
438
+ * per-bucket window/episode transitions.
379
439
  *
380
- * @param {number} percentage
440
+ * @param {{fiveHour?: BucketSample|null, weekly?: BucketSample|null}} quota
441
+ * @private
381
442
  */
382
- armWatcher(percentage) {
383
- if (!this.wasExhausted) {
384
- this.wasExhausted = true;
385
- this.exhaustedSinceMs = this.now();
386
- this.notifiedThisEpisode = false;
387
- this.log.log(
388
- `pi-lxmf: GLM 5h quota exhausted (${percentage}%) — will notify on recovery`,
389
- );
443
+ evaluateSample(quota) {
444
+ const pct = quota.fiveHour?.percentage ?? 0;
445
+ const weeklyPct = quota.weekly?.percentage ?? 0;
446
+ const exhausted = pct >= EXHAUSTED_PERCENTAGE;
447
+ const weeklyExhausted = weeklyPct >= EXHAUSTED_PERCENTAGE;
448
+
449
+ /** @type {string[]} */
450
+ const lines = [];
451
+
452
+ // --- 5h bucket ------------------------------------------------------
453
+ if (exhausted) {
454
+ if (!this.wasExhausted) {
455
+ // New exhausted episode: stamp it, re-arm its notices.
456
+ this.wasExhausted = true;
457
+ this.exhaustedSinceMs = this.now();
458
+ this.exhaustionNotified.fiveHour = false;
459
+ // Suppress the 90% warning for the rest of this window: the
460
+ // exhaustion (and its recovery) notices say everything already.
461
+ this.warned.fiveHour = true;
462
+ }
463
+ if (!this.exhaustionNotified.fiveHour) {
464
+ this.exhaustionNotified.fiveHour = true;
465
+ lines.push(
466
+ `⚠️ GLM 5h quota exhausted (100%)${resetSuffix(quota.fiveHour?.nextResetMs, this.now())}`,
467
+ );
468
+ }
469
+ } else {
470
+ if (this.wasExhausted) {
471
+ // Recovered below 100% after an exhausted episode (once per episode).
472
+ this.wasExhausted = false;
473
+ const elapsed = this.exhaustedSinceMs
474
+ ? this.now() - this.exhaustedSinceMs
475
+ : 0;
476
+ lines.push(
477
+ `✅ GLM 5h quota available again${elapsed > 0 ? ` (was exhausted for ~${formatElapsed(elapsed)})` : ""}. Weekly: ${weeklyPct}%.`,
478
+ );
479
+ }
480
+ if (pct < this.warnPercentage) {
481
+ // Below the threshold: re-arm the warning for the next window.
482
+ this.warned.fiveHour = false;
483
+ } else if (!this.warned.fiveHour) {
484
+ // At-or-above the threshold but not exhausted: warn once per window
485
+ // (a dip below the threshold re-arms the warning).
486
+ this.warned.fiveHour = true;
487
+ lines.push(`⚠️ GLM 5h quota at ${pct}% — getting close.`);
488
+ }
390
489
  }
391
- this.startPoller();
392
- }
393
490
 
394
- /** Starts the poller if not already running. */
395
- startPoller() {
396
- if (this.pollTimer || !this.enabled || !this.apiKey) return;
397
- const apiKey = this.apiKey;
398
- const tick = async () => {
399
- try {
400
- const quota = await fetchQuota(apiKey, {
401
- fetchImpl: this.fetchImpl,
402
- });
403
- const pct = quota.fiveHour?.percentage ?? 0;
404
- if (pct < EXHAUSTED_PERCENTAGE && this.wasExhausted) {
405
- this.notifyRecovered(quota);
406
- }
407
- } catch {
408
- /* transient — retry on the next tick */
491
+ // --- weekly bucket ----------------------------------------------------
492
+ if (weeklyExhausted) {
493
+ if (!this.exhaustionNotified.weekly) {
494
+ this.exhaustionNotified.weekly = true;
495
+ this.warned.weekly = false;
496
+ lines.push(
497
+ `⚠️ GLM weekly quota exhausted (100%)${resetSuffix(quota.weekly?.nextResetMs, this.now())}`,
498
+ );
409
499
  }
410
- };
411
- this.pollTimer = setInterval(() => void tick(), this.pollIntervalMs);
412
- if (typeof this.pollTimer.unref === "function") this.pollTimer.unref();
413
- }
500
+ } else if (weeklyPct >= this.warnPercentage && !this.warned.weekly) {
501
+ this.warned.weekly = true;
502
+ lines.push(`⚠️ GLM weekly quota at ${weeklyPct}% — getting close.`);
503
+ } else if (weeklyPct < this.warnPercentage) {
504
+ this.warned.weekly = false;
505
+ }
414
506
 
415
- /** Stops the poller. */
416
- stopPoller() {
417
- if (this.pollTimer) {
418
- clearInterval(this.pollTimer);
419
- this.pollTimer = null;
507
+ if (lines.length > 0) {
508
+ this.log.log(`pi-lxmf: GLM quota notice: ${lines.join(" | ")}`);
509
+ this.queueNotice(lines.join("\n"));
420
510
  }
421
511
  }
422
512
 
423
513
  /**
424
- * Delivers the one-shot recovery notification and stops the poller.
514
+ * Queues a notice; queued notices are delivered when the bridge is idle
515
+ * or, during a live owner-triggered run, on `agent_settled` (never
516
+ * interleaving with a reply mid-run).
425
517
  *
426
- * @param {{fiveHour?: {percentage: number}|null, weekly?: {percentage: number}|null, level?: string|null}} quota
518
+ * @param {string} text
519
+ * @private
427
520
  */
428
- async notifyRecovered(quota) {
429
- if (this.notifiedThisEpisode) return;
430
- this.notifiedThisEpisode = true;
431
- this.wasExhausted = false;
432
- this.stopPoller();
433
- const elapsed = this.exhaustedSinceMs
434
- ? this.now() - this.exhaustedSinceMs
435
- : 0;
436
- const weeklyPct = quota.weekly?.percentage;
437
- const level = quota.level ? `GLM ${quota.level}` : "GLM";
438
- const parts = [`${level} 5h quota available again`];
439
- if (elapsed > 0) {
440
- parts.push(`(was exhausted for ~${formatElapsed(elapsed)})`);
441
- }
442
- if (typeof weeklyPct === "number") {
443
- parts.push(`Weekly: ${weeklyPct}%`);
444
- }
445
- try {
446
- await this.sendText(this.ownerDestinationHash, parts.join(" "));
447
- } catch (e) {
521
+ queueNotice(text) {
522
+ this.pendingNotices.push(text);
523
+ this.flushNotices();
524
+ }
525
+
526
+ /**
527
+ * Delivers queued notices as one message, when no owner-triggered run
528
+ * is live. Delivery is best-effort; a failure is logged once.
529
+ *
530
+ * @private
531
+ */
532
+ flushNotices() {
533
+ if (this.pendingNotices.length === 0) return;
534
+ if (this.runActive && this.ownerTriggered) return;
535
+ const text = this.pendingNotices.join("\n\n");
536
+ this.pendingNotices = [];
537
+ this.sendText(this.ownerDestinationHash, text).catch((e) => {
448
538
  this.log.error(
449
539
  `pi-lxmf: GLM quota notification delivery failed: ${e instanceof Error ? e.message : e}`,
450
540
  );
451
- }
541
+ });
452
542
  }
453
543
 
454
544
  /**
@@ -500,6 +590,27 @@ export class GlmQuotaWatcher {
500
590
  this.peakOpenTimer = null;
501
591
  }
502
592
  }
593
+
594
+ /** Stops the sampler. */
595
+ stopPoller() {
596
+ if (this.pollTimer) {
597
+ clearInterval(this.pollTimer);
598
+ this.pollTimer = null;
599
+ }
600
+ }
601
+ }
602
+
603
+ /**
604
+ * Human-readable "resets in" suffix from a bucket's `nextResetMs`.
605
+ *
606
+ * @param {number|undefined} nextResetMs
607
+ * @param {number} nowMs
608
+ * @returns {string}
609
+ */
610
+ function resetSuffix(nextResetMs, nowMs) {
611
+ if (!nextResetMs || nextResetMs <= nowMs) return "";
612
+ const ms = nextResetMs - nowMs;
613
+ return ` — resets in ~${formatElapsed(ms)}`;
503
614
  }
504
615
 
505
616
  /** @param {number} ms */