instar 1.3.1122 → 1.3.1124
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/core/MultiMachineCoordinator.d.ts +23 -0
- package/dist/core/MultiMachineCoordinator.d.ts.map +1 -1
- package/dist/core/MultiMachineCoordinator.js +67 -12
- package/dist/core/MultiMachineCoordinator.js.map +1 -1
- package/dist/core/SessionManager.d.ts +9 -0
- package/dist/core/SessionManager.d.ts.map +1 -1
- package/dist/core/SessionManager.js +22 -8
- package/dist/core/SessionManager.js.map +1 -1
- package/dist/data/standards-guard-index.json +1 -1
- package/dist/data/standards-guard-index.meta.json +2 -2
- package/dist/data/standards-registry.meta.json +1 -1
- package/package.json +1 -1
- package/src/data/builtin-manifest.json +4 -4
- package/src/data/standards-guard-index.json +1 -1
- package/src/data/standards-guard-index.meta.json +2 -2
- package/src/data/standards-registry.meta.json +1 -1
- package/upgrades/1.3.1123.md +95 -0
- package/upgrades/1.3.1124.md +73 -0
- package/upgrades/side-effects/lease-poll-intent-republish.md +163 -0
- package/upgrades/side-effects/memory-pressure-metric-sibling.md +108 -0
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
"schemaVersion": 1,
|
|
3
3
|
"generatedFrom": "source-tree",
|
|
4
4
|
"registrySha256": "5413a0c6ef9ba2bda876b509d1c0bfebe450da3708cd5a5d627d6cb836012d58",
|
|
5
|
-
"packageVersion": "1.3.
|
|
5
|
+
"packageVersion": "1.3.1124",
|
|
6
6
|
"guards": [
|
|
7
7
|
{
|
|
8
8
|
"ref": "docs/canonical-migration-contracts.json",
|
|
@@ -1,5 +1,5 @@
|
|
|
1
1
|
{
|
|
2
|
-
"sha256": "
|
|
2
|
+
"sha256": "8b1a0ac4291301135f90deb37da180f1340e6089f6c8b7eda3acdf5ef0dfbd24",
|
|
3
3
|
"registrySha256": "5413a0c6ef9ba2bda876b509d1c0bfebe450da3708cd5a5d627d6cb836012d58",
|
|
4
|
-
"packageVersion": "1.3.
|
|
4
|
+
"packageVersion": "1.3.1124"
|
|
5
5
|
}
|
package/package.json
CHANGED
|
@@ -1,8 +1,8 @@
|
|
|
1
1
|
{
|
|
2
2
|
"$schema": "./builtin-manifest.schema.json",
|
|
3
3
|
"schemaVersion": 1,
|
|
4
|
-
"generatedAt": "2026-08-
|
|
5
|
-
"instarVersion": "1.3.
|
|
4
|
+
"generatedAt": "2026-08-04T06:29:42.911Z",
|
|
5
|
+
"instarVersion": "1.3.1124",
|
|
6
6
|
"entryCount": 202,
|
|
7
7
|
"entries": {
|
|
8
8
|
"hook:session-start": {
|
|
@@ -1538,7 +1538,7 @@
|
|
|
1538
1538
|
"type": "subsystem",
|
|
1539
1539
|
"domain": "sessions",
|
|
1540
1540
|
"sourcePath": "src/core/SessionManager.ts",
|
|
1541
|
-
"contentHash": "
|
|
1541
|
+
"contentHash": "2f9ad44f737f806891fe147ef6c9c8313c299c27991a4c458a1aea74f75d0251",
|
|
1542
1542
|
"since": "2025-01-01"
|
|
1543
1543
|
},
|
|
1544
1544
|
"subsystem:auto-updater": {
|
|
@@ -1618,7 +1618,7 @@
|
|
|
1618
1618
|
"type": "subsystem",
|
|
1619
1619
|
"domain": "coordination",
|
|
1620
1620
|
"sourcePath": "src/core/MultiMachineCoordinator.ts",
|
|
1621
|
-
"contentHash": "
|
|
1621
|
+
"contentHash": "d7aa8bf43b24118b16d89a61c838e4b8a4de72c876ea5167738b20155797fde6",
|
|
1622
1622
|
"since": "2025-01-01"
|
|
1623
1623
|
},
|
|
1624
1624
|
"subsystem:backup-manager": {
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
"schemaVersion": 1,
|
|
3
3
|
"generatedFrom": "source-tree",
|
|
4
4
|
"registrySha256": "5413a0c6ef9ba2bda876b509d1c0bfebe450da3708cd5a5d627d6cb836012d58",
|
|
5
|
-
"packageVersion": "1.3.
|
|
5
|
+
"packageVersion": "1.3.1124",
|
|
6
6
|
"guards": [
|
|
7
7
|
{
|
|
8
8
|
"ref": "docs/canonical-migration-contracts.json",
|
|
@@ -1,5 +1,5 @@
|
|
|
1
1
|
{
|
|
2
|
-
"sha256": "
|
|
2
|
+
"sha256": "8b1a0ac4291301135f90deb37da180f1340e6089f6c8b7eda3acdf5ef0dfbd24",
|
|
3
3
|
"registrySha256": "5413a0c6ef9ba2bda876b509d1c0bfebe450da3708cd5a5d627d6cb836012d58",
|
|
4
|
-
"packageVersion": "1.3.
|
|
4
|
+
"packageVersion": "1.3.1124"
|
|
5
5
|
}
|
|
@@ -0,0 +1,95 @@
|
|
|
1
|
+
# Upgrade Guide — vNEXT
|
|
2
|
+
|
|
3
|
+
<!-- assembled-by: assemble-next-md -->
|
|
4
|
+
<!-- bump: patch -->
|
|
5
|
+
|
|
6
|
+
## What Changed
|
|
7
|
+
|
|
8
|
+
On a multi-machine agent, the server publishes a small record telling its
|
|
9
|
+
Lifeline process whether THIS machine should own the Telegram poll. That record
|
|
10
|
+
was only rewritten when the machine's role actually changed. Because a machine's
|
|
11
|
+
role is restored from the machine registry at startup, a machine that was
|
|
12
|
+
already `awake` before a restart comes back, reconciles, finds nothing changed,
|
|
13
|
+
and returns early — leaving the cautious boot default (`standby` /
|
|
14
|
+
`shouldPoll:false`) standing permanently on the machine that genuinely holds the
|
|
15
|
+
fenced lease. A second failure rode along: written only on transitions, the
|
|
16
|
+
record aged past the Lifeline's 90-second freshness bound on a steady role and
|
|
17
|
+
was discarded as "no current opinion".
|
|
18
|
+
|
|
19
|
+
A second, independent defect in the same path is fixed with it. The "safe boot
|
|
20
|
+
default" (`standby` / `shouldPoll:false`) was written at the END of
|
|
21
|
+
`initializeLease()` — but every branch above it already ends in a role
|
|
22
|
+
reconcile, so the default did not stand in for an unmade decision, it
|
|
23
|
+
OVERWROTE the decision made three lines earlier. On its own that is enough to
|
|
24
|
+
leave a lease HOLDER published as `standby / do-not-poll` for the life of the
|
|
25
|
+
process. The default write now precedes the branch, matching the ordering its
|
|
26
|
+
own comment describes.
|
|
27
|
+
|
|
28
|
+
The lease-derived poll intent is now published on every role reconcile rather
|
|
29
|
+
than only on a transition, so it always reflects `holdsLease()` and stays inside
|
|
30
|
+
the consumer's freshness bound. An unchanged intent is rewritten at most once
|
|
31
|
+
every 30 seconds (a 3× margin under the 90-second bound); a changed role or a
|
|
32
|
+
moved lease epoch publishes immediately. A write that fails no longer records a
|
|
33
|
+
skip-window, so it is retried on the next reconcile instead of being treated as
|
|
34
|
+
delivered.
|
|
35
|
+
|
|
36
|
+
Everything that should happen only on a genuine role change still happens only
|
|
37
|
+
on a genuine role change — the `role_transition` audit entry, the
|
|
38
|
+
promote/demote events, and the flap circuit-breaker were deliberately left
|
|
39
|
+
below the early-return.
|
|
40
|
+
|
|
41
|
+
## What to Tell Your User
|
|
42
|
+
|
|
43
|
+
- **The machine in charge no longer tells itself it is the standby:** "When I
|
|
44
|
+
run on more than one machine, the one holding the badge now reports that
|
|
45
|
+
correctly to its own messenger instead of leaving a startup placeholder in
|
|
46
|
+
place."
|
|
47
|
+
- **Why this mattered:** "With the placeholder standing, my messenger either
|
|
48
|
+
fought the other machine for the Telegram line — colliding, deciding it was
|
|
49
|
+
stuck, and restarting roughly every ten minutes — or, with the follow-the-lease
|
|
50
|
+
behaviour switched on, muted the machine that was supposed to be answering."
|
|
51
|
+
- **Nothing changes on a single-machine setup:** "A single machine is always the
|
|
52
|
+
one in charge, so there was never a second poller to collide with."
|
|
53
|
+
|
|
54
|
+
## Summary of New Capabilities
|
|
55
|
+
|
|
56
|
+
| Capability | How to Use |
|
|
57
|
+
|---|---|
|
|
58
|
+
| The poll intent reflects the live fenced lease rather than a boot placeholder | Automatic — no configuration. Applies wherever `multiMachine.pollFollowsLease` resolves on |
|
|
59
|
+
| The poll intent stays inside the consumer's freshness bound on a steady role | Automatic — an unchanged record is refreshed every 30s; a changed one publishes immediately |
|
|
60
|
+
| A failed intent write is retried rather than recorded as delivered | Automatic — the throttle window only advances on a write that actually landed |
|
|
61
|
+
|
|
62
|
+
## Evidence
|
|
63
|
+
|
|
64
|
+
Measured live on a two-machine agent (v1.3.1122) before the fix: `/health →
|
|
65
|
+
multiMachine.syncStatus` reported `role:awake, holdsLease:true,
|
|
66
|
+
leaseEpoch:21302` in the same second that `state/telegram-poll-intent.json`
|
|
67
|
+
reported `role:standby, shouldPoll:false` at that same epoch. Sampling the
|
|
68
|
+
record at 10-second intervals showed its age climbing 84.5s → 155.0s across an
|
|
69
|
+
80-second window with no rewrite — past the consumer's 90s bound. Downstream on
|
|
70
|
+
that agent: 812 × `Telegram 409 Conflict — another bot instance is polling`,
|
|
71
|
+
260 × `TelegramLifeline.selfRestart: conflict409Stuck`, and a server restart
|
|
72
|
+
roughly every ten minutes for about eight hours.
|
|
73
|
+
|
|
74
|
+
New unit coverage in
|
|
75
|
+
`tests/unit/MultiMachineCoordinator-pollIntentRepublish.test.ts` (10 tests):
|
|
76
|
+
holder-with-no-transition publishes `shouldPoll:true` (the regression),
|
|
77
|
+
non-holder publishes `shouldPoll:false`, steady-role refresh past the window
|
|
78
|
+
with throttling inside it, a failed write retried on the next reconcile,
|
|
79
|
+
immediate publish on a lease-epoch change, no write when the gate is explicitly
|
|
80
|
+
off, transition-only side effects still firing on a real change and NOT firing
|
|
81
|
+
on a steady role, plus two tests that run the REAL boot path with a live
|
|
82
|
+
FencedLease/LeaseCoordinator — a machine that acquires ends up published as
|
|
83
|
+
`awake / shouldPoll:true`, and a machine on the observe-only branch that never
|
|
84
|
+
acquires is left muted. Verified as a control that 7 of the 10 fail against
|
|
85
|
+
`origin/main` — the headline regression failing with `expected null not to be
|
|
86
|
+
null` (no record written at all) rather than a missing-symbol error; the 3 that
|
|
87
|
+
pass pre-fix are the gate-off, real-transition, and never-acquires guards, which
|
|
88
|
+
are expected to hold either way.
|
|
89
|
+
|
|
90
|
+
The second defect was found BY the real-boot-path test, not by reading: with
|
|
91
|
+
only the early-return fixed, that test failed with the on-disk record reading
|
|
92
|
+
`{shouldPoll:false, role:'standby'}` while the in-process throttle key read
|
|
93
|
+
`true|awake|1` — the correct value had been written and then overwritten by the
|
|
94
|
+
boot default. Shipping the early-return fix alone would have produced a green
|
|
95
|
+
stubbed suite and an unchanged production failure.
|
|
@@ -0,0 +1,73 @@
|
|
|
1
|
+
# Upgrade Guide — vNEXT
|
|
2
|
+
|
|
3
|
+
<!-- assembled-by: assemble-next-md -->
|
|
4
|
+
<!-- bump: patch -->
|
|
5
|
+
|
|
6
|
+
## What Changed
|
|
7
|
+
|
|
8
|
+
`SessionManager.currentMemoryPressure()` now reads real available memory via `hostFreeMemPct()`
|
|
9
|
+
(free + inactive + purgeable through `vm_stat` on macOS; `MemAvailable` on Linux) instead of raw
|
|
10
|
+
`os.freemem()`.
|
|
11
|
+
|
|
12
|
+
On macOS `os.freemem()` counts only "Pages free" — **0.46 GB of 17 GB (2.7%)** on a completely healthy
|
|
13
|
+
host. The method therefore computed 97% used and returned **`critical` permanently**. Under
|
|
14
|
+
`subscriptionPath.mode: 'force'`, `evaluateRerouteGate()` throws on elevated pressure, so **every job
|
|
15
|
+
spawn was refused**.
|
|
16
|
+
|
|
17
|
+
The corrected helper already existed in-package and is used by `SessionReaper` / `HostPressureSampler`;
|
|
18
|
+
its comment warns against `os.freemem()` by name. The 2026-06-26 `macos-memory-pressure-metric` fix
|
|
19
|
+
converted those two callers and missed this one. **Thresholds (90/75/60) are unchanged — only the
|
|
20
|
+
measurement source moved.**
|
|
21
|
+
|
|
22
|
+
`getSessionDiagnostics()`'s memory block moves to the same source, so the reported percentage and the
|
|
23
|
+
reported tier can no longer contradict each other (before, that surface could show "97% used" next to
|
|
24
|
+
tier `low`).
|
|
25
|
+
|
|
26
|
+
Two other `os.freemem()` callsites (`src/server/routes.ts`, `src/monitoring/HealthChecker.ts`) already
|
|
27
|
+
prefer `MemoryPressureMonitor`'s `vm_stat` state and fall back only when it is absent — reviewed and
|
|
28
|
+
correct as-is, deliberately untouched.
|
|
29
|
+
|
|
30
|
+
## What to Tell Your User
|
|
31
|
+
|
|
32
|
+
If your scheduled jobs stopped running and every failure said the machine was out of memory, this is the
|
|
33
|
+
fix — and the machine was almost certainly fine the whole time.
|
|
34
|
+
|
|
35
|
+
On a Mac, the figure being used counted only completely untouched memory. Macs deliberately use spare
|
|
36
|
+
memory as cache, so that number sits near zero on a perfectly healthy machine. Your agent read it as
|
|
37
|
+
"97% used, critical" and refused to start any scheduled job.
|
|
38
|
+
|
|
39
|
+
On the machine where this was found, **20 of 27 enabled jobs were dead — the health check and the
|
|
40
|
+
commitment tracker had each failed 421 times in a row**, going back more than two days, while the machine
|
|
41
|
+
had ample memory free. A second machine with 137 GB of memory and 21 GB genuinely free was being rated
|
|
42
|
+
under pressure by the same calculation.
|
|
43
|
+
|
|
44
|
+
Your agent now reads memory the way the operating system itself reports it. Nothing about *when* it
|
|
45
|
+
protects you has changed — a genuinely loaded machine still refuses new work at exactly the same
|
|
46
|
+
thresholds. It just no longer believes a healthy machine is full.
|
|
47
|
+
|
|
48
|
+
You may also notice memory figures in session diagnostics drop substantially on macOS. That is the
|
|
49
|
+
correction, not a regression.
|
|
50
|
+
|
|
51
|
+
## Summary of New Capabilities
|
|
52
|
+
|
|
53
|
+
No new capabilities — this is a correctness fix to an existing guard. The guard keeps its authority and
|
|
54
|
+
its thresholds; only the number it reads is corrected.
|
|
55
|
+
|
|
56
|
+
## Evidence
|
|
57
|
+
|
|
58
|
+
- **Live measurement before the fix:** 20 of 27 enabled jobs failing, all with
|
|
59
|
+
`Reroute refused (force-mode): host memory pressure is critical`; `health-check` and
|
|
60
|
+
`commitment-detection` at **421 consecutive failures** each.
|
|
61
|
+
- **Three readings of the same machine, same second:** raw `os.freemem()` → 2.7% free → `critical`;
|
|
62
|
+
the reaper's corrected reading → 16.3% free → `normal`; macOS `memory_pressure` → 38% free.
|
|
63
|
+
- **Reproduced on a second machine:** 137 GB total, 21 GB genuinely free → raw metric rated it `high`.
|
|
64
|
+
- **New tests:** `tests/unit/session-manager-memory-pressure-sibling.test.ts` (6).
|
|
65
|
+
**Control run — 3 of 6 fail against pre-fix code for the right reasons**, including
|
|
66
|
+
`expected [Function] to not throw but 'Error: Reroute refused (force-mode): …' was thrown`, the exact
|
|
67
|
+
production error. The other 3 pass either way **by design**: they pin that a genuinely exhausted host
|
|
68
|
+
still reads `critical` and is still refused, so an "always return low" regression cannot pass.
|
|
69
|
+
- **Green:** 48/48 across the new suite plus `host-memory-pressure.test.ts` (16) and
|
|
70
|
+
`headless-spawn-reroute.test.ts` (26). `tsc --noEmit` clean; full lint suite clean.
|
|
71
|
+
- **Side-effects review:** `upgrades/side-effects/memory-pressure-metric-sibling.md`.
|
|
72
|
+
- **Spec:** `docs/specs/macos-memory-pressure-metric.md` (converged 2026-06-26, approved) — this ships
|
|
73
|
+
the caller that spec's original fix did not convert.
|
|
@@ -0,0 +1,163 @@
|
|
|
1
|
+
# Side-Effects Review — lease-derived poll intent is republished on every reconcile
|
|
2
|
+
|
|
3
|
+
**Version / slug:** `lease-poll-intent-republish`
|
|
4
|
+
**Date:** `2026-08-03`
|
|
5
|
+
**Author:** `echo`
|
|
6
|
+
**Second-pass reviewer:** `not required` (see §Second-pass — no block/allow surface added; the change is confined to a signal producer)
|
|
7
|
+
|
|
8
|
+
## Summary of the change
|
|
9
|
+
|
|
10
|
+
`MultiMachineCoordinator.reconcileRoleToLease()` published the lease-derived poll intent (`state/telegram-poll-intent.json`) only on a genuine role TRANSITION, because the `writeLeasePollIntent(holds, desired)` call sat below the `if (desired === this._role) return;` early-return. `_role` is restored from the machine registry at startup, so a machine that was already `awake` before a restart re-enters reconcile with `desired === this._role`, returns early, and never replaces `initializeLease()`'s safe boot default `{shouldPoll:false, role:'standby'}`. The published intent for a machine that genuinely holds the lease therefore stayed `standby / do-not-poll` permanently. A second failure mode rode along: because the record was written only on transitions, on a steady role its `ts` aged past the consumer's `maxStaleMs: 90_000` bound and `effectivePollIntent` degraded it to "no current opinion".
|
|
11
|
+
|
|
12
|
+
This change moves the publish ABOVE the early-return so it runs on every reconcile, and adds a throttling wrapper (`publishLeasePollIntent`) that writes when the intent changes (`shouldPoll | role | leaseEpoch`) or when the last SUCCESSFUL write is older than `POLL_INTENT_REFRESH_MS` (30s — a 3× margin inside the consumer's 90s bound). `writeLeasePollIntent` now returns whether a record actually landed, so a failed write never records a skip-window.
|
|
13
|
+
|
|
14
|
+
**A SECOND defect was found while building the real-boot-path test, and is fixed here too.** `initializeLease()` wrote its "safe boot default" (`writeLeasePollIntent(false, 'standby')`) at the very END of the method — but every branch above it already ends in `reconcileRoleToLease(...)`. The default therefore did not stand in for an unmade decision; it OVERWROTE the decision that had just been made three lines earlier. Its own comment ("at boot the role isn't yet reconciled … before the first reconcile decides the real role") describes the intended ordering, which the code did not have. On its own this defect is sufficient to leave a lease HOLDER published as `standby / do-not-poll` for the life of the process, so fixing only the early-return would have produced a change that passes a stubbed unit test and still fails in production. The default write is moved to before the branch.
|
|
15
|
+
|
|
16
|
+
Files touched: `src/core/MultiMachineCoordinator.ts`, `tests/unit/MultiMachineCoordinator-pollIntentRepublish.test.ts` (new).
|
|
17
|
+
|
|
18
|
+
## Decision-point inventory
|
|
19
|
+
|
|
20
|
+
- `MultiMachineCoordinator.reconcileRoleToLease` (poll-intent publish) — **modify** — publishes on every reconcile instead of transitions only; the transition-only side effects below the early-return are untouched.
|
|
21
|
+
- `MultiMachineCoordinator.initializeLease` (safe boot default) — **modify** — the default write moves from after the acquire/reconcile branch to before it. Same value written, same purpose; the ordering now matches the stated intent.
|
|
22
|
+
- `MultiMachineCoordinator.writeLeasePollIntent` — **modify** — now returns `boolean` (write landed) instead of `void`, and records the throttle state on every successful write. No behavioral change to what it writes.
|
|
23
|
+
- `MultiMachineCoordinator.publishLeasePollIntent` — **add** — throttling wrapper, no decision authority.
|
|
24
|
+
- `TelegramLifeline.reconcilePolling` (the CONSUMER, which owns the start/stop authority) — **pass-through** — unmodified. It keeps its own freshness gate, dead-writer gate, operator override, debounce, and the `pollFollowsLease.dryRun` gate.
|
|
25
|
+
|
|
26
|
+
---
|
|
27
|
+
|
|
28
|
+
## 1. Over-block
|
|
29
|
+
|
|
30
|
+
**No block/allow surface — over-block not applicable.** This change writes an advisory record. It has no reject path. The only actor that acts on the record is `TelegramLifeline.reconcilePolling`, which is unmodified.
|
|
31
|
+
|
|
32
|
+
The nearest analogue worth stating: the change can now cause the lifeline to STOP polling on a machine where, pre-change, it would have kept polling — but only when the machine genuinely does not hold the lease AND `pollFollowsLease.dryRun` is explicitly `false`. That is the intended contract, not an over-block: the record now reports the fenced lease truthfully instead of reporting a stale boot default.
|
|
33
|
+
|
|
34
|
+
---
|
|
35
|
+
|
|
36
|
+
## 2. Under-block
|
|
37
|
+
|
|
38
|
+
**No block/allow surface — under-block not applicable.**
|
|
39
|
+
|
|
40
|
+
Failure modes this change does NOT address, stated explicitly so they are not assumed fixed:
|
|
41
|
+
|
|
42
|
+
- **A lifeline whose 409 counter never resets.** The production incident that surfaced this defect also showed `conflict409Stuck` firing after a single 409 in a fresh boot cycle, because the boot path starts polling (stale-connection flush → `Telegram polling active`) before the first `reconcilePolling` tick can mute it. Fixing the publisher removes the CAUSE of the dual-poll on a correctly-leased pair, but a lifeline that starts polling at boot and only reconciles 15s later still has a window. Not in scope here and not claimed fixed. <!-- tracked: ATT-poll-intent-standby-forever-20260803 -->
|
|
43
|
+
- **`nobodyPollingRecovery` shipping in dryRun.** If the lease is held by a machine that is down, nothing currently actuates a recovery. Unchanged by this. <!-- tracked: ATT-poll-intent-standby-forever-20260803 -->
|
|
44
|
+
- **A machine whose `_role` and the fenced lease disagree for a reason other than the missing republish.** This change makes the PUBLISHED record follow `holdsLease()` directly, so it is now independent of `_role` drift — but it does not repair `_role` itself.
|
|
45
|
+
|
|
46
|
+
---
|
|
47
|
+
|
|
48
|
+
## 3. Level-of-abstraction fit
|
|
49
|
+
|
|
50
|
+
Right layer. The coordinator is the only component that owns the fenced lease, so it is the only component that can answer "should this machine own the Telegram poll?". The record is the existing, purpose-built IPC surface between the server (which knows the lease) and the lifeline (which owns the socket), and it already carries the integrity fields (`serverPid`, `bootId`, `ts`) the consumer needs.
|
|
51
|
+
|
|
52
|
+
A lower layer (the lifeline asking Telegram, or reading the lease file itself) would duplicate lease parsing in a second process and reintroduce exactly the split-truth this record exists to prevent. A higher layer (a new arbiter service) is unwarranted for a one-boolean handoff.
|
|
53
|
+
|
|
54
|
+
No smarter gate exists that this should feed instead — the consumer IS the gate, and it already receives this signal.
|
|
55
|
+
|
|
56
|
+
---
|
|
57
|
+
|
|
58
|
+
## 4. Signal vs authority compliance
|
|
59
|
+
|
|
60
|
+
**Required reference:** [docs/signal-vs-authority.md](../../docs/signal-vs-authority.md)
|
|
61
|
+
|
|
62
|
+
- [x] **No — this change produces a signal consumed by an existing smart gate.**
|
|
63
|
+
|
|
64
|
+
The coordinator is a detector here: it reports a fact it alone holds (do I hold the fenced lease, at which epoch). It gains no blocking power from this change. The authority over polling remains `decidePollAction` inside `TelegramLifeline`, which combines this signal with the operator override (`pollOverride` / `telegramPolling:false`), the local 409 observation, the start debounce, and the `peerPresumedGone` conservatism — and is itself gated behind `pollFollowsLease.dryRun`.
|
|
65
|
+
|
|
66
|
+
The defect being fixed is, in signal-vs-authority terms, a detector reporting a value it never measured: the safe boot default is a placeholder standing in for an unmeasured role, and it was being consumed as though it were a measurement. Making the detector report the measured value is the compliant direction.
|
|
67
|
+
|
|
68
|
+
---
|
|
69
|
+
|
|
70
|
+
## 4b. Judgment-point check (Judgment Within Floors standard)
|
|
71
|
+
|
|
72
|
+
No new heuristic at a competing-signals decision point. The published value is a direct, deterministic read of a single authoritative signal — `leaseCoordinator.holdsLease()` — not a weighing of competing signals. The place where signals genuinely compete (lease intent vs local 409 vs operator override vs debounce) is `decidePollAction`, which is unchanged.
|
|
73
|
+
|
|
74
|
+
The only new constant, `POLL_INTENT_REFRESH_MS`, is a write-amplification throttle, not a decision threshold: it never changes WHAT is published, only how often an unchanged value is rewritten, and it is derived from the consumer's existing `maxStaleMs` bound (30s vs 90s) rather than tuned.
|
|
75
|
+
|
|
76
|
+
---
|
|
77
|
+
|
|
78
|
+
## 5. Interactions
|
|
79
|
+
|
|
80
|
+
- **Shadowing:** the publish now runs BEFORE the early-return, so it executes on strictly more paths than before. It cannot shadow anything — it writes a file and returns; no control flow depends on its result. Nothing that previously ran still fails to run. The reverse shadowing is what the second defect WAS: the boot default shadowed the reconcile's decision by writing after it. Moving it before the branch removes that, and the "does NOT acquire ⇒ left muted" test is the control proving the safe default still wins on the branch where no decision is made.
|
|
81
|
+
- **Throttle vs direct writers:** the throttle cache is now updated inside `writeLeasePollIntent` rather than in the wrapper, so a direct caller (the boot default) can never leave the cached key describing something other than what is on disk. Without this, the boot default would write `standby` while the cache still read `awake`, and the next reconcile would consider itself unchanged and skip the correcting write for up to 30s — the exact interaction that made the first draft of the real-boot-path test fail.
|
|
82
|
+
- **Double-fire:** on a genuine transition the publish runs once (from the new call above the early-return); the old call below the early-return was REMOVED, so there is no double write. Verified by the transition test asserting one write plus one `roleChange`/`promote` pair.
|
|
83
|
+
- **Races:** `writePollIntent` is atomic (tmp + rename), so the lifeline never reads a torn record — unchanged. The reconcile is single-threaded per process and already guarded by `leaseTicking`. The throttle fields are process-local and only mutated inside reconcile.
|
|
84
|
+
- **Feedback loops:** the record feeds the lifeline, which starts/stops polling, which changes the local 409 observation, which feeds `decidePollAction`. This change does not add a loop — it makes the existing one's input correct. The relevant risk (a machine flipping poll on/off repeatedly) is bounded by the existing start-debounce and by `POLL_INTENT_REFRESH_MS` throttling of unchanged values; a genuine flip publishes immediately, which is required for correctness (a lagging record is what caused the incident).
|
|
85
|
+
- **Churn breaker (B2):** deliberately left below the early-return so it still counts only true flips. Covered by the "does NOT re-fire transition-only side effects when the role is steady" test.
|
|
86
|
+
|
|
87
|
+
---
|
|
88
|
+
|
|
89
|
+
## 6. External surfaces
|
|
90
|
+
|
|
91
|
+
- **Other agents on the same machine:** none — the record is per-agent under that agent's `stateDir`.
|
|
92
|
+
- **Other users of the install base:** the change is inert unless `pollFollowsLease` resolves on. It ships dev-gated; on the fleet the guard resolves off and `writeLeasePollIntent` returns early, writing nothing. No fleet-visible change on this commit.
|
|
93
|
+
- **External systems:** no direct calls. Indirectly, on an agent with `pollFollowsLease.dryRun:false`, Telegram sees one long-poll instead of two competing ones — which is the intended end state and strictly fewer 409s.
|
|
94
|
+
- **Persistent state:** `state/telegram-poll-intent.json` only. Same schema, same fields, same atomic write. Written more often (bounded to once per 30s per unchanged value). No migration: the record is regenerated at runtime and readers already tolerate a missing/stale/corrupt file by treating it as "no opinion".
|
|
95
|
+
- **Timing / runtime conditions:** the throttle uses wall-clock `Date.now()`. A backwards clock step could delay a refresh by up to the step size; the consumer's response to a too-old record is "no opinion → hold", the safe direction, so a clock anomaly degrades to today's behavior rather than to a wrong action.
|
|
96
|
+
- **Operator surface (Mobile-Complete Operator Actions):** no operator-facing action added or changed. The existing levers (`pollOverride`, `telegramPolling`) are untouched.
|
|
97
|
+
|
|
98
|
+
---
|
|
99
|
+
|
|
100
|
+
## 6b. Operator-surface quality
|
|
101
|
+
|
|
102
|
+
**Not applicable** — this change touches no dashboard renderer, approval surface, notification body, or operator-facing markup. No operator-facing text is added or modified.
|
|
103
|
+
|
|
104
|
+
---
|
|
105
|
+
|
|
106
|
+
## 7. Multi-machine posture (Cross-Machine Coherence)
|
|
107
|
+
|
|
108
|
+
**Machine-local BY DESIGN.** The record answers a question that MUST differ per machine: "should THIS machine own the Telegram poll?" Exactly one machine may answer yes at a time, so replicating the record would be actively wrong — a replicated `shouldPoll:true` landing on a standby is precisely the dual-poll failure the fenced lease exists to prevent. The cross-machine coordination is carried by the fenced lease itself (already replicated via `LeaseCoordinator` / the mesh transports), and the record is the machine-local projection of that shared truth into a sibling process.
|
|
109
|
+
|
|
110
|
+
- **User-facing notices:** none emitted. No one-voice gating needed.
|
|
111
|
+
- **Durable state on topic transfer:** the record is not topic-scoped and holds no per-topic state, so it cannot strand on a transfer. It is regenerated from the lease on the next reconcile on whichever machine holds it.
|
|
112
|
+
- **Generated URLs:** none.
|
|
113
|
+
|
|
114
|
+
Worth naming explicitly: this defect is itself a multi-machine defect that a single-machine install can never surface (a single machine short-circuits to `awake` and there is no second poller to conflict with). The evidence for this review came from a live two-machine pair, which is the graduation gate the parent spec names for B1/B5.
|
|
115
|
+
|
|
116
|
+
---
|
|
117
|
+
|
|
118
|
+
## 8. Rollback cost
|
|
119
|
+
|
|
120
|
+
- **Hot-fix release:** revert the commit and ship as the next patch. One function plus one constant plus two fields.
|
|
121
|
+
- **Data migration:** none. The record is advisory and regenerated at runtime; readers already treat missing/stale as "no opinion".
|
|
122
|
+
- **Agent state repair:** none. No agent needs notifying or resetting.
|
|
123
|
+
- **User visibility:** none during the rollback window on the fleet (the guard resolves off there). On a dev agent with the gate on, reverting restores the previous behavior exactly: the boot default stands and the lifeline holds or fights, as before.
|
|
124
|
+
|
|
125
|
+
---
|
|
126
|
+
|
|
127
|
+
## Conclusion
|
|
128
|
+
|
|
129
|
+
The review changed the design twice. First, `writeLeasePollIntent` was changed to return a success boolean so the throttle can never record a skip-window off a write that failed — without that, a transient write failure would have suppressed retries for 30s, and the consumer would have aged into "no opinion" for a reason the code believed it had handled.
|
|
130
|
+
|
|
131
|
+
Second and more importantly: writing a test that runs the REAL boot path, rather than only the stubbed reconcile, surfaced a second independent defect — the safe boot default being written after the reconcile it was meant to precede. Shipping the early-return fix alone would have produced a green unit suite and an unchanged production failure. That is worth recording as the reason the extra test existed: the stubbed tests could not have caught it, because they never ran `initializeLease()`. The throttle-cache ownership move (into `writeLeasePollIntent`) came out of the same test failing a second time for a different reason.
|
|
132
|
+
|
|
133
|
+
Third, §2 forced an explicit statement of what this does NOT fix: the lifeline's boot-time poll-before-first-reconcile window and the dry-run `nobodyPollingRecovery` are both still open, both tracked on the attention item, and neither is claimed closed by this commit.
|
|
134
|
+
|
|
135
|
+
The change complies with signal-vs-authority: it corrects a detector's output and adds no authority. It is inert on the fleet (dev-gated producer) and its whole risk surface is confined to agents that have explicitly enabled `pollFollowsLease`.
|
|
136
|
+
|
|
137
|
+
---
|
|
138
|
+
|
|
139
|
+
## Second-pass review
|
|
140
|
+
|
|
141
|
+
**Not required.** Phase 5 triggers on changes that touch block/allow decisions on messaging or dispatch, session lifecycle, or anything named sentinel/guard/gate/watchdog. This change adds no block/allow surface and modifies no gate: it is confined to the producer side of an advisory record, and the consuming gate (`decidePollAction`) is untouched. The one adjacent trigger word — the lifeline's poll authority — is explicitly out of the diff.
|
|
142
|
+
|
|
143
|
+
Recorded for the reviewer's benefit if this judgment is revisited: the argument rests on the diff containing no change to `TelegramLifeline`, which is verifiable from the commit's file list.
|
|
144
|
+
|
|
145
|
+
---
|
|
146
|
+
|
|
147
|
+
## Evidence pointers
|
|
148
|
+
|
|
149
|
+
- Live measurement (instar-codey, Mac Mini, 2026-08-03 17:34 PDT, v1.3.1122): `/health → multiMachine.syncStatus` reporting `role:awake, holdsLease:true, leaseEpoch:21302` in the same second that `state/telegram-poll-intent.json` reported `role:standby, shouldPoll:false, leaseEpoch:21302`.
|
|
150
|
+
- Staleness half: the same record's age sampled at 10s intervals climbed 84.5s → 155.0s across an 80s window with no rewrite, i.e. past the consumer's 90s bound.
|
|
151
|
+
- Downstream harm on that agent: 812 × `Telegram 409 Conflict — another bot instance is polling`, 260 × `TelegramLifeline.selfRestart: conflict409Stuck`, server restarting every ~10 min for ~8h.
|
|
152
|
+
- Pre-fix test control: with `src/core/MultiMachineCoordinator.ts` reverted to `origin/main`, the headline regression test fails with `expected null not to be null` (no record written at all) — the right reason, not a missing-symbol error. 7 of 10 tests in the new file fail pre-fix; the 3 that pass are the gate-off guard, the real-transition guard, and the observe-only "does NOT acquire ⇒ left muted" guard, all of which are expected to hold either way.
|
|
153
|
+
- Second defect, found by the real-boot-path test rather than by reading: with only the early-return fixed, that test failed with the record on disk reading `{shouldPoll:false, role:'standby'}` while the in-process throttle key read `true|awake|1` — i.e. the correct value had been written and then overwritten by the boot default at the tail of `initializeLease()`. That divergence between the cache and the file is what identified the clobber.
|
|
154
|
+
- Attention item carrying the fleet risk and the mitigation that must be removed after deploy: `ATT-poll-intent-standby-forever-20260803`.
|
|
155
|
+
|
|
156
|
+
---
|
|
157
|
+
|
|
158
|
+
## Class-Closure Declaration
|
|
159
|
+
|
|
160
|
+
**Not applicable on both triggers, stated explicitly rather than omitted.**
|
|
161
|
+
|
|
162
|
+
1. **Agent-authored-artifact defect?** No. The defect is in hand-written TypeScript (`MultiMachineCoordinator.ts`), not in an LLM prompt, hook, config, skill, or standards text.
|
|
163
|
+
2. **Self-triggered controller added or modified?** No. `reconcileRoleToLease` is an existing control loop driven by the existing lease tick; this change adds no new loop, monitor, sentinel, reaper, scheduler, or recovery path, and fires no restart / swap / respawn / spawn / notify / retry / re-drive / kill. The throttle strictly REDUCES the action rate of an existing loop (a file write) and cannot increase it: an unchanged value is written at most once per `POLL_INTENT_REFRESH_MS`, and a changed value at most once per reconcile tick, which is the loop's own bound.
|
|
@@ -0,0 +1,108 @@
|
|
|
1
|
+
# Side-effects review — macos-memory-pressure-metric SIBLING (SessionManager)
|
|
2
|
+
|
|
3
|
+
**Change:** `SessionManager.currentMemoryPressure()` and `getSessionDiagnostics()`'s memory block now read
|
|
4
|
+
the corrected available-memory figure via `hostFreeMemPct()` instead of raw `os.freemem()`.
|
|
5
|
+
**Spec:** `docs/specs/macos-memory-pressure-metric.md` (converged 2026-06-26, `approved: true`) — this is
|
|
6
|
+
the caller that spec's original fix did not convert.
|
|
7
|
+
**Author:** echo · **Date:** 2026-08-04
|
|
8
|
+
**Sanction:** architect review 2026-08-04, under Justin's plan-scoped approval of 2026-08-03 20:21 PDT
|
|
9
|
+
(relayed 20:23). Escalated to Phase A critical path because it starves the grading job that answers rung 3.
|
|
10
|
+
|
|
11
|
+
## What was wrong
|
|
12
|
+
|
|
13
|
+
`os.freemem()` on macOS counts only "Pages free". On a healthy 17 GB Mini it reports **0.46 GB (2.7%)**,
|
|
14
|
+
so `currentMemoryPressure()` computed 97% used and returned **`critical` permanently**. Under
|
|
15
|
+
`subscriptionPath.mode: 'force'`, `evaluateRerouteGate()` **throws** on elevated pressure — so every job
|
|
16
|
+
spawn was refused.
|
|
17
|
+
|
|
18
|
+
**Measured on the live host before the fix:**
|
|
19
|
+
- 20 of 27 enabled jobs failing, all with `Reroute refused (force-mode): host memory pressure is critical`
|
|
20
|
+
- `health-check` and `commitment-detection` at **421 consecutive failures** each, going back ≥2 days
|
|
21
|
+
- three readings of the same machine, same second: **raw 2.7% free → `critical`** · reaper (corrected)
|
|
22
|
+
**16.3% free → `normal`** · macOS `memory_pressure` **38% free**
|
|
23
|
+
- the same defect on the laptop: **137 GB RAM, 21 GB genuinely free → raw metric says `high`**
|
|
24
|
+
|
|
25
|
+
The corrected helper already existed in-package (`hostFreeMemPct`, used by
|
|
26
|
+
`HostPressureSampler`/`SessionReaper`) and its shipped comment warns against `os.freemem()` **by name**.
|
|
27
|
+
The reaper was fixed; this caller was not.
|
|
28
|
+
|
|
29
|
+
## The 8 questions
|
|
30
|
+
|
|
31
|
+
**1. Over-block — what legitimate inputs does this reject that it shouldn't?**
|
|
32
|
+
None introduced; this change *removes* an over-block. Pre-fix the gate rejected **every** force-mode spawn
|
|
33
|
+
on macOS regardless of real conditions. Post-fix it rejects only genuinely elevated hosts. Thresholds
|
|
34
|
+
(90/75/60) are byte-identical — only the measurement source moved.
|
|
35
|
+
|
|
36
|
+
**2. Under-block — what failure modes does this still miss?**
|
|
37
|
+
The gate now trusts `hostFreeMemPct`, so it inherits that helper's limits: on an unrecognised platform it
|
|
38
|
+
falls back to `os.freemem()`-equivalent behaviour, and `vm_stat` parse failure degrades to the same. That
|
|
39
|
+
is the pre-existing, already-reviewed behaviour of the 2026-06-26 fix, not new exposure. **Genuinely
|
|
40
|
+
exhausted hosts still read `critical` — covered by a dedicated test so a "just return low" regression
|
|
41
|
+
cannot pass.**
|
|
42
|
+
|
|
43
|
+
**3. Level-of-abstraction fit.**
|
|
44
|
+
Correct layer. `currentMemoryPressure()` is the single definition of the pressure tier shared by the
|
|
45
|
+
reroute gate and the diagnostics surface (deliberately extracted in june15-headless-spawn-reroute PR2/O2).
|
|
46
|
+
Fixing it here fixes both consumers at once. The alternative — patching each caller — is what produced
|
|
47
|
+
this sibling in the first place.
|
|
48
|
+
|
|
49
|
+
**4. Signal vs authority compliance.** (`docs/signal-vs-authority.md`)
|
|
50
|
+
**Compliant, and it is the point of the change.** `currentMemoryPressure()` is a **detector**: it produces
|
|
51
|
+
a signal (`MemoryPressure` tier). `evaluateRerouteGate()` is the **authority** that decides. The defect was
|
|
52
|
+
a detector emitting a *false signal* to a correct authority. This change corrects the detector's
|
|
53
|
+
measurement and **adds no authority, moves no authority, and changes no threshold**. No new blocking logic
|
|
54
|
+
is introduced anywhere.
|
|
55
|
+
|
|
56
|
+
**5. Interactions — shadowing, double-fire, races.**
|
|
57
|
+
Two other `os.freemem()` callsites exist (`src/server/routes.ts:3743`, `src/monitoring/HealthChecker.ts:217`).
|
|
58
|
+
**Both were inspected and both already compensate** — each prefers `MemoryPressureMonitor`'s vm_stat-based
|
|
59
|
+
state and only falls back to `os.freemem()` when the monitor is absent, with a comment saying why. They are
|
|
60
|
+
correct as-is and are deliberately **not** touched. No shadowing: the reroute gate is the only consumer of
|
|
61
|
+
the tier for spawn decisions, and the diagnostics surface is read-only.
|
|
62
|
+
|
|
63
|
+
⚠️ **Interaction worth naming:** `tests/unit/headless-spawn-reroute.test.ts` **stubs `currentMemoryPressure`
|
|
64
|
+
to `'normal'`**, with the comment that the real gate *"made this suite fail on loaded dev machines."* It was
|
|
65
|
+
not a loaded dev machine — **it was this defect, encountered, rationalised, and stubbed over**, which is why
|
|
66
|
+
CI never caught it. The stub is left in place (those tests assert reroute control-flow, not pressure) but
|
|
67
|
+
the new suite exercises the real method so it can now fail for the real reason.
|
|
68
|
+
|
|
69
|
+
**6. External surfaces.**
|
|
70
|
+
`getSessionDiagnostics()` payload changes: `usedPercent` and `freeMemMB` now report real available memory.
|
|
71
|
+
On macOS these numbers will **drop substantially** (e.g. 97% → ~40% used) — that is the correction, not a
|
|
72
|
+
regression. Pre-fix, once the tier was fixed alone, the surface could have reported **"97% used" beside tier
|
|
73
|
+
`low`**; both now read one source and a test pins that they cannot contradict. No timing or conversation-state
|
|
74
|
+
dependence.
|
|
75
|
+
|
|
76
|
+
**7. Multi-machine posture (Cross-Machine Coherence).**
|
|
77
|
+
**Machine-local BY DESIGN — correctly so.** Host memory pressure is a property of the physical machine; a
|
|
78
|
+
reading must never be replicated or proxied from a peer. Verified the defect independently on **both**
|
|
79
|
+
machines (Mini: raw `critical` vs corrected `normal`; laptop: raw `high` vs corrected `normal` with 21 GB
|
|
80
|
+
free). The fix therefore lands on both, and each machine keeps reading its own memory. No replication path,
|
|
81
|
+
no merged read, no cross-machine state. **This also removes a latent block on the ratified placement policy:**
|
|
82
|
+
worker lanes are to run on the laptop, and the same gate would have refused spawns there too.
|
|
83
|
+
|
|
84
|
+
**8. Rollback cost.**
|
|
85
|
+
Trivial. Revert the commit — no data migration, no agent-state repair, no persisted format change. The
|
|
86
|
+
change is confined to two computations inside one file plus one new test file. A hot-fix release restores
|
|
87
|
+
prior behaviour exactly (including, if anyone wants it, the permanent-`critical` bug).
|
|
88
|
+
|
|
89
|
+
## Testing
|
|
90
|
+
|
|
91
|
+
New: `tests/unit/session-manager-memory-pressure-sibling.test.ts` (6 tests).
|
|
92
|
+
|
|
93
|
+
**Control run — 3 of 6 fail against pre-fix code, each for the right reason** (verified by stashing the
|
|
94
|
+
source change and re-running):
|
|
95
|
+
- `expected 'critical' not to be 'critical'` — the defect itself
|
|
96
|
+
- `expected [Function] to not throw but 'Error: Reroute refused (force-mode): …' was thrown` — **the exact
|
|
97
|
+
production error that killed 20 jobs**
|
|
98
|
+
- `expected 'critical' to be 'high'` — the Linux threshold path also mis-tiered
|
|
99
|
+
|
|
100
|
+
The other 3 pass either way **by design** — they are the guards proving the fix does not simply always
|
|
101
|
+
return `low` (genuinely-exhausted host still `critical`; gate still refuses it; diagnostics self-consistent).
|
|
102
|
+
|
|
103
|
+
Green: 48/48 across the new suite + `host-memory-pressure.test.ts` (16) + `headless-spawn-reroute.test.ts`
|
|
104
|
+
(26). `tsc --noEmit` clean.
|
|
105
|
+
|
|
106
|
+
## Second-pass review
|
|
107
|
+
|
|
108
|
+
Required (touches a **gate**). See appended section below.
|