instar 1.3.1122 → 1.3.1124

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -2,7 +2,7 @@
2
2
  "schemaVersion": 1,
3
3
  "generatedFrom": "source-tree",
4
4
  "registrySha256": "5413a0c6ef9ba2bda876b509d1c0bfebe450da3708cd5a5d627d6cb836012d58",
5
- "packageVersion": "1.3.1122",
5
+ "packageVersion": "1.3.1124",
6
6
  "guards": [
7
7
  {
8
8
  "ref": "docs/canonical-migration-contracts.json",
@@ -1,5 +1,5 @@
1
1
  {
2
- "sha256": "11ce55d65baed309e0b55b093b70883c00303b387b04aa31409d1906fa7b362c",
2
+ "sha256": "8b1a0ac4291301135f90deb37da180f1340e6089f6c8b7eda3acdf5ef0dfbd24",
3
3
  "registrySha256": "5413a0c6ef9ba2bda876b509d1c0bfebe450da3708cd5a5d627d6cb836012d58",
4
- "packageVersion": "1.3.1122"
4
+ "packageVersion": "1.3.1124"
5
5
  }
@@ -2,5 +2,5 @@
2
2
  "sha256": "5413a0c6ef9ba2bda876b509d1c0bfebe450da3708cd5a5d627d6cb836012d58",
3
3
  "articleCount": 82,
4
4
  "generatedFrom": "docs/STANDARDS-REGISTRY.md",
5
- "packageVersion": "1.3.1122"
5
+ "packageVersion": "1.3.1124"
6
6
  }
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "instar",
3
- "version": "1.3.1122",
3
+ "version": "1.3.1124",
4
4
  "description": "Coherence infrastructure for self-evolving AI agents — on the Claude Code or Codex subscription you already have.",
5
5
  "type": "module",
6
6
  "main": "dist/index.js",
@@ -1,8 +1,8 @@
1
1
  {
2
2
  "$schema": "./builtin-manifest.schema.json",
3
3
  "schemaVersion": 1,
4
- "generatedAt": "2026-08-03T16:08:38.846Z",
5
- "instarVersion": "1.3.1122",
4
+ "generatedAt": "2026-08-04T06:29:42.911Z",
5
+ "instarVersion": "1.3.1124",
6
6
  "entryCount": 202,
7
7
  "entries": {
8
8
  "hook:session-start": {
@@ -1538,7 +1538,7 @@
1538
1538
  "type": "subsystem",
1539
1539
  "domain": "sessions",
1540
1540
  "sourcePath": "src/core/SessionManager.ts",
1541
- "contentHash": "b75927c5058db73067ec28107b31a6268a4b27489dab5917df16c6b05779aac8",
1541
+ "contentHash": "2f9ad44f737f806891fe147ef6c9c8313c299c27991a4c458a1aea74f75d0251",
1542
1542
  "since": "2025-01-01"
1543
1543
  },
1544
1544
  "subsystem:auto-updater": {
@@ -1618,7 +1618,7 @@
1618
1618
  "type": "subsystem",
1619
1619
  "domain": "coordination",
1620
1620
  "sourcePath": "src/core/MultiMachineCoordinator.ts",
1621
- "contentHash": "411005681842ff0cb2472b9b6af2e435e70b4d4a20d22c8342a712b6d2ab5d4c",
1621
+ "contentHash": "d7aa8bf43b24118b16d89a61c838e4b8a4de72c876ea5167738b20155797fde6",
1622
1622
  "since": "2025-01-01"
1623
1623
  },
1624
1624
  "subsystem:backup-manager": {
@@ -2,7 +2,7 @@
2
2
  "schemaVersion": 1,
3
3
  "generatedFrom": "source-tree",
4
4
  "registrySha256": "5413a0c6ef9ba2bda876b509d1c0bfebe450da3708cd5a5d627d6cb836012d58",
5
- "packageVersion": "1.3.1122",
5
+ "packageVersion": "1.3.1124",
6
6
  "guards": [
7
7
  {
8
8
  "ref": "docs/canonical-migration-contracts.json",
@@ -1,5 +1,5 @@
1
1
  {
2
- "sha256": "11ce55d65baed309e0b55b093b70883c00303b387b04aa31409d1906fa7b362c",
2
+ "sha256": "8b1a0ac4291301135f90deb37da180f1340e6089f6c8b7eda3acdf5ef0dfbd24",
3
3
  "registrySha256": "5413a0c6ef9ba2bda876b509d1c0bfebe450da3708cd5a5d627d6cb836012d58",
4
- "packageVersion": "1.3.1122"
4
+ "packageVersion": "1.3.1124"
5
5
  }
@@ -2,5 +2,5 @@
2
2
  "sha256": "5413a0c6ef9ba2bda876b509d1c0bfebe450da3708cd5a5d627d6cb836012d58",
3
3
  "articleCount": 82,
4
4
  "generatedFrom": "docs/STANDARDS-REGISTRY.md",
5
- "packageVersion": "1.3.1122"
5
+ "packageVersion": "1.3.1124"
6
6
  }
@@ -0,0 +1,95 @@
1
+ # Upgrade Guide — vNEXT
2
+
3
+ <!-- assembled-by: assemble-next-md -->
4
+ <!-- bump: patch -->
5
+
6
+ ## What Changed
7
+
8
+ On a multi-machine agent, the server publishes a small record telling its
9
+ Lifeline process whether THIS machine should own the Telegram poll. That record
10
+ was only rewritten when the machine's role actually changed. Because a machine's
11
+ role is restored from the machine registry at startup, a machine that was
12
+ already `awake` before a restart comes back, reconciles, finds nothing changed,
13
+ and returns early — leaving the cautious boot default (`standby` /
14
+ `shouldPoll:false`) standing permanently on the machine that genuinely holds the
15
+ fenced lease. A second failure rode along: written only on transitions, the
16
+ record aged past the Lifeline's 90-second freshness bound on a steady role and
17
+ was discarded as "no current opinion".
18
+
19
+ A second, independent defect in the same path is fixed with it. The "safe boot
20
+ default" (`standby` / `shouldPoll:false`) was written at the END of
21
+ `initializeLease()` — but every branch above it already ends in a role
22
+ reconcile, so the default did not stand in for an unmade decision, it
23
+ OVERWROTE the decision made three lines earlier. On its own that is enough to
24
+ leave a lease HOLDER published as `standby / do-not-poll` for the life of the
25
+ process. The default write now precedes the branch, matching the ordering its
26
+ own comment describes.
27
+
28
+ The lease-derived poll intent is now published on every role reconcile rather
29
+ than only on a transition, so it always reflects `holdsLease()` and stays inside
30
+ the consumer's freshness bound. An unchanged intent is rewritten at most once
31
+ every 30 seconds (a 3× margin under the 90-second bound); a changed role or a
32
+ moved lease epoch publishes immediately. A write that fails no longer records a
33
+ skip-window, so it is retried on the next reconcile instead of being treated as
34
+ delivered.
35
+
36
+ Everything that should happen only on a genuine role change still happens only
37
+ on a genuine role change — the `role_transition` audit entry, the
38
+ promote/demote events, and the flap circuit-breaker were deliberately left
39
+ below the early-return.
40
+
41
+ ## What to Tell Your User
42
+
43
+ - **The machine in charge no longer tells itself it is the standby:** "When I
44
+ run on more than one machine, the one holding the badge now reports that
45
+ correctly to its own messenger instead of leaving a startup placeholder in
46
+ place."
47
+ - **Why this mattered:** "With the placeholder standing, my messenger either
48
+ fought the other machine for the Telegram line — colliding, deciding it was
49
+ stuck, and restarting roughly every ten minutes — or, with the follow-the-lease
50
+ behaviour switched on, muted the machine that was supposed to be answering."
51
+ - **Nothing changes on a single-machine setup:** "A single machine is always the
52
+ one in charge, so there was never a second poller to collide with."
53
+
54
+ ## Summary of New Capabilities
55
+
56
+ | Capability | How to Use |
57
+ |---|---|
58
+ | The poll intent reflects the live fenced lease rather than a boot placeholder | Automatic — no configuration. Applies wherever `multiMachine.pollFollowsLease` resolves on |
59
+ | The poll intent stays inside the consumer's freshness bound on a steady role | Automatic — an unchanged record is refreshed every 30s; a changed one publishes immediately |
60
+ | A failed intent write is retried rather than recorded as delivered | Automatic — the throttle window only advances on a write that actually landed |
61
+
62
+ ## Evidence
63
+
64
+ Measured live on a two-machine agent (v1.3.1122) before the fix: `/health →
65
+ multiMachine.syncStatus` reported `role:awake, holdsLease:true,
66
+ leaseEpoch:21302` in the same second that `state/telegram-poll-intent.json`
67
+ reported `role:standby, shouldPoll:false` at that same epoch. Sampling the
68
+ record at 10-second intervals showed its age climbing 84.5s → 155.0s across an
69
+ 80-second window with no rewrite — past the consumer's 90s bound. Downstream on
70
+ that agent: 812 × `Telegram 409 Conflict — another bot instance is polling`,
71
+ 260 × `TelegramLifeline.selfRestart: conflict409Stuck`, and a server restart
72
+ roughly every ten minutes for about eight hours.
73
+
74
+ New unit coverage in
75
+ `tests/unit/MultiMachineCoordinator-pollIntentRepublish.test.ts` (10 tests):
76
+ holder-with-no-transition publishes `shouldPoll:true` (the regression),
77
+ non-holder publishes `shouldPoll:false`, steady-role refresh past the window
78
+ with throttling inside it, a failed write retried on the next reconcile,
79
+ immediate publish on a lease-epoch change, no write when the gate is explicitly
80
+ off, transition-only side effects still firing on a real change and NOT firing
81
+ on a steady role, plus two tests that run the REAL boot path with a live
82
+ FencedLease/LeaseCoordinator — a machine that acquires ends up published as
83
+ `awake / shouldPoll:true`, and a machine on the observe-only branch that never
84
+ acquires is left muted. Verified as a control that 7 of the 10 fail against
85
+ `origin/main` — the headline regression failing with `expected null not to be
86
+ null` (no record written at all) rather than a missing-symbol error; the 3 that
87
+ pass pre-fix are the gate-off, real-transition, and never-acquires guards, which
88
+ are expected to hold either way.
89
+
90
+ The second defect was found BY the real-boot-path test, not by reading: with
91
+ only the early-return fixed, that test failed with the on-disk record reading
92
+ `{shouldPoll:false, role:'standby'}` while the in-process throttle key read
93
+ `true|awake|1` — the correct value had been written and then overwritten by the
94
+ boot default. Shipping the early-return fix alone would have produced a green
95
+ stubbed suite and an unchanged production failure.
@@ -0,0 +1,73 @@
1
+ # Upgrade Guide — vNEXT
2
+
3
+ <!-- assembled-by: assemble-next-md -->
4
+ <!-- bump: patch -->
5
+
6
+ ## What Changed
7
+
8
+ `SessionManager.currentMemoryPressure()` now reads real available memory via `hostFreeMemPct()`
9
+ (free + inactive + purgeable through `vm_stat` on macOS; `MemAvailable` on Linux) instead of raw
10
+ `os.freemem()`.
11
+
12
+ On macOS `os.freemem()` counts only "Pages free" — **0.46 GB of 17 GB (2.7%)** on a completely healthy
13
+ host. The method therefore computed 97% used and returned **`critical` permanently**. Under
14
+ `subscriptionPath.mode: 'force'`, `evaluateRerouteGate()` throws on elevated pressure, so **every job
15
+ spawn was refused**.
16
+
17
+ The corrected helper already existed in-package and is used by `SessionReaper` / `HostPressureSampler`;
18
+ its comment warns against `os.freemem()` by name. The 2026-06-26 `macos-memory-pressure-metric` fix
19
+ converted those two callers and missed this one. **Thresholds (90/75/60) are unchanged — only the
20
+ measurement source moved.**
21
+
22
+ `getSessionDiagnostics()`'s memory block moves to the same source, so the reported percentage and the
23
+ reported tier can no longer contradict each other (before, that surface could show "97% used" next to
24
+ tier `low`).
25
+
26
+ Two other `os.freemem()` callsites (`src/server/routes.ts`, `src/monitoring/HealthChecker.ts`) already
27
+ prefer `MemoryPressureMonitor`'s `vm_stat` state and fall back only when it is absent — reviewed and
28
+ correct as-is, deliberately untouched.
29
+
30
+ ## What to Tell Your User
31
+
32
+ If your scheduled jobs stopped running and every failure said the machine was out of memory, this is the
33
+ fix — and the machine was almost certainly fine the whole time.
34
+
35
+ On a Mac, the figure being used counted only completely untouched memory. Macs deliberately use spare
36
+ memory as cache, so that number sits near zero on a perfectly healthy machine. Your agent read it as
37
+ "97% used, critical" and refused to start any scheduled job.
38
+
39
+ On the machine where this was found, **20 of 27 enabled jobs were dead — the health check and the
40
+ commitment tracker had each failed 421 times in a row**, going back more than two days, while the machine
41
+ had ample memory free. A second machine with 137 GB of memory and 21 GB genuinely free was being rated
42
+ under pressure by the same calculation.
43
+
44
+ Your agent now reads memory the way the operating system itself reports it. Nothing about *when* it
45
+ protects you has changed — a genuinely loaded machine still refuses new work at exactly the same
46
+ thresholds. It just no longer believes a healthy machine is full.
47
+
48
+ You may also notice memory figures in session diagnostics drop substantially on macOS. That is the
49
+ correction, not a regression.
50
+
51
+ ## Summary of New Capabilities
52
+
53
+ No new capabilities — this is a correctness fix to an existing guard. The guard keeps its authority and
54
+ its thresholds; only the number it reads is corrected.
55
+
56
+ ## Evidence
57
+
58
+ - **Live measurement before the fix:** 20 of 27 enabled jobs failing, all with
59
+ `Reroute refused (force-mode): host memory pressure is critical`; `health-check` and
60
+ `commitment-detection` at **421 consecutive failures** each.
61
+ - **Three readings of the same machine, same second:** raw `os.freemem()` → 2.7% free → `critical`;
62
+ the reaper's corrected reading → 16.3% free → `normal`; macOS `memory_pressure` → 38% free.
63
+ - **Reproduced on a second machine:** 137 GB total, 21 GB genuinely free → raw metric rated it `high`.
64
+ - **New tests:** `tests/unit/session-manager-memory-pressure-sibling.test.ts` (6).
65
+ **Control run — 3 of 6 fail against pre-fix code for the right reasons**, including
66
+ `expected [Function] to not throw but 'Error: Reroute refused (force-mode): …' was thrown`, the exact
67
+ production error. The other 3 pass either way **by design**: they pin that a genuinely exhausted host
68
+ still reads `critical` and is still refused, so an "always return low" regression cannot pass.
69
+ - **Green:** 48/48 across the new suite plus `host-memory-pressure.test.ts` (16) and
70
+ `headless-spawn-reroute.test.ts` (26). `tsc --noEmit` clean; full lint suite clean.
71
+ - **Side-effects review:** `upgrades/side-effects/memory-pressure-metric-sibling.md`.
72
+ - **Spec:** `docs/specs/macos-memory-pressure-metric.md` (converged 2026-06-26, approved) — this ships
73
+ the caller that spec's original fix did not convert.
@@ -0,0 +1,163 @@
1
+ # Side-Effects Review — lease-derived poll intent is republished on every reconcile
2
+
3
+ **Version / slug:** `lease-poll-intent-republish`
4
+ **Date:** `2026-08-03`
5
+ **Author:** `echo`
6
+ **Second-pass reviewer:** `not required` (see §Second-pass — no block/allow surface added; the change is confined to a signal producer)
7
+
8
+ ## Summary of the change
9
+
10
+ `MultiMachineCoordinator.reconcileRoleToLease()` published the lease-derived poll intent (`state/telegram-poll-intent.json`) only on a genuine role TRANSITION, because the `writeLeasePollIntent(holds, desired)` call sat below the `if (desired === this._role) return;` early-return. `_role` is restored from the machine registry at startup, so a machine that was already `awake` before a restart re-enters reconcile with `desired === this._role`, returns early, and never replaces `initializeLease()`'s safe boot default `{shouldPoll:false, role:'standby'}`. The published intent for a machine that genuinely holds the lease therefore stayed `standby / do-not-poll` permanently. A second failure mode rode along: because the record was written only on transitions, on a steady role its `ts` aged past the consumer's `maxStaleMs: 90_000` bound and `effectivePollIntent` degraded it to "no current opinion".
11
+
12
+ This change moves the publish ABOVE the early-return so it runs on every reconcile, and adds a throttling wrapper (`publishLeasePollIntent`) that writes when the intent changes (`shouldPoll | role | leaseEpoch`) or when the last SUCCESSFUL write is older than `POLL_INTENT_REFRESH_MS` (30s — a 3× margin inside the consumer's 90s bound). `writeLeasePollIntent` now returns whether a record actually landed, so a failed write never records a skip-window.
13
+
14
+ **A SECOND defect was found while building the real-boot-path test, and is fixed here too.** `initializeLease()` wrote its "safe boot default" (`writeLeasePollIntent(false, 'standby')`) at the very END of the method — but every branch above it already ends in `reconcileRoleToLease(...)`. The default therefore did not stand in for an unmade decision; it OVERWROTE the decision that had just been made three lines earlier. Its own comment ("at boot the role isn't yet reconciled … before the first reconcile decides the real role") describes the intended ordering, which the code did not have. On its own this defect is sufficient to leave a lease HOLDER published as `standby / do-not-poll` for the life of the process, so fixing only the early-return would have produced a change that passes a stubbed unit test and still fails in production. The default write is moved to before the branch.
15
+
16
+ Files touched: `src/core/MultiMachineCoordinator.ts`, `tests/unit/MultiMachineCoordinator-pollIntentRepublish.test.ts` (new).
17
+
18
+ ## Decision-point inventory
19
+
20
+ - `MultiMachineCoordinator.reconcileRoleToLease` (poll-intent publish) — **modify** — publishes on every reconcile instead of transitions only; the transition-only side effects below the early-return are untouched.
21
+ - `MultiMachineCoordinator.initializeLease` (safe boot default) — **modify** — the default write moves from after the acquire/reconcile branch to before it. Same value written, same purpose; the ordering now matches the stated intent.
22
+ - `MultiMachineCoordinator.writeLeasePollIntent` — **modify** — now returns `boolean` (write landed) instead of `void`, and records the throttle state on every successful write. No behavioral change to what it writes.
23
+ - `MultiMachineCoordinator.publishLeasePollIntent` — **add** — throttling wrapper, no decision authority.
24
+ - `TelegramLifeline.reconcilePolling` (the CONSUMER, which owns the start/stop authority) — **pass-through** — unmodified. It keeps its own freshness gate, dead-writer gate, operator override, debounce, and the `pollFollowsLease.dryRun` gate.
25
+
26
+ ---
27
+
28
+ ## 1. Over-block
29
+
30
+ **No block/allow surface — over-block not applicable.** This change writes an advisory record. It has no reject path. The only actor that acts on the record is `TelegramLifeline.reconcilePolling`, which is unmodified.
31
+
32
+ The nearest analogue worth stating: the change can now cause the lifeline to STOP polling on a machine where, pre-change, it would have kept polling — but only when the machine genuinely does not hold the lease AND `pollFollowsLease.dryRun` is explicitly `false`. That is the intended contract, not an over-block: the record now reports the fenced lease truthfully instead of reporting a stale boot default.
33
+
34
+ ---
35
+
36
+ ## 2. Under-block
37
+
38
+ **No block/allow surface — under-block not applicable.**
39
+
40
+ Failure modes this change does NOT address, stated explicitly so they are not assumed fixed:
41
+
42
+ - **A lifeline whose 409 counter never resets.** The production incident that surfaced this defect also showed `conflict409Stuck` firing after a single 409 in a fresh boot cycle, because the boot path starts polling (stale-connection flush → `Telegram polling active`) before the first `reconcilePolling` tick can mute it. Fixing the publisher removes the CAUSE of the dual-poll on a correctly-leased pair, but a lifeline that starts polling at boot and only reconciles 15s later still has a window. Not in scope here and not claimed fixed. <!-- tracked: ATT-poll-intent-standby-forever-20260803 -->
43
+ - **`nobodyPollingRecovery` shipping in dryRun.** If the lease is held by a machine that is down, nothing currently actuates a recovery. Unchanged by this. <!-- tracked: ATT-poll-intent-standby-forever-20260803 -->
44
+ - **A machine whose `_role` and the fenced lease disagree for a reason other than the missing republish.** This change makes the PUBLISHED record follow `holdsLease()` directly, so it is now independent of `_role` drift — but it does not repair `_role` itself.
45
+
46
+ ---
47
+
48
+ ## 3. Level-of-abstraction fit
49
+
50
+ Right layer. The coordinator is the only component that owns the fenced lease, so it is the only component that can answer "should this machine own the Telegram poll?". The record is the existing, purpose-built IPC surface between the server (which knows the lease) and the lifeline (which owns the socket), and it already carries the integrity fields (`serverPid`, `bootId`, `ts`) the consumer needs.
51
+
52
+ A lower layer (the lifeline asking Telegram, or reading the lease file itself) would duplicate lease parsing in a second process and reintroduce exactly the split-truth this record exists to prevent. A higher layer (a new arbiter service) is unwarranted for a one-boolean handoff.
53
+
54
+ No smarter gate exists that this should feed instead — the consumer IS the gate, and it already receives this signal.
55
+
56
+ ---
57
+
58
+ ## 4. Signal vs authority compliance
59
+
60
+ **Required reference:** [docs/signal-vs-authority.md](../../docs/signal-vs-authority.md)
61
+
62
+ - [x] **No — this change produces a signal consumed by an existing smart gate.**
63
+
64
+ The coordinator is a detector here: it reports a fact it alone holds (do I hold the fenced lease, at which epoch). It gains no blocking power from this change. The authority over polling remains `decidePollAction` inside `TelegramLifeline`, which combines this signal with the operator override (`pollOverride` / `telegramPolling:false`), the local 409 observation, the start debounce, and the `peerPresumedGone` conservatism — and is itself gated behind `pollFollowsLease.dryRun`.
65
+
66
+ The defect being fixed is, in signal-vs-authority terms, a detector reporting a value it never measured: the safe boot default is a placeholder standing in for an unmeasured role, and it was being consumed as though it were a measurement. Making the detector report the measured value is the compliant direction.
67
+
68
+ ---
69
+
70
+ ## 4b. Judgment-point check (Judgment Within Floors standard)
71
+
72
+ No new heuristic at a competing-signals decision point. The published value is a direct, deterministic read of a single authoritative signal — `leaseCoordinator.holdsLease()` — not a weighing of competing signals. The place where signals genuinely compete (lease intent vs local 409 vs operator override vs debounce) is `decidePollAction`, which is unchanged.
73
+
74
+ The only new constant, `POLL_INTENT_REFRESH_MS`, is a write-amplification throttle, not a decision threshold: it never changes WHAT is published, only how often an unchanged value is rewritten, and it is derived from the consumer's existing `maxStaleMs` bound (30s vs 90s) rather than tuned.
75
+
76
+ ---
77
+
78
+ ## 5. Interactions
79
+
80
+ - **Shadowing:** the publish now runs BEFORE the early-return, so it executes on strictly more paths than before. It cannot shadow anything — it writes a file and returns; no control flow depends on its result. Nothing that previously ran still fails to run. The reverse shadowing is what the second defect WAS: the boot default shadowed the reconcile's decision by writing after it. Moving it before the branch removes that, and the "does NOT acquire ⇒ left muted" test is the control proving the safe default still wins on the branch where no decision is made.
81
+ - **Throttle vs direct writers:** the throttle cache is now updated inside `writeLeasePollIntent` rather than in the wrapper, so a direct caller (the boot default) can never leave the cached key describing something other than what is on disk. Without this, the boot default would write `standby` while the cache still read `awake`, and the next reconcile would consider itself unchanged and skip the correcting write for up to 30s — the exact interaction that made the first draft of the real-boot-path test fail.
82
+ - **Double-fire:** on a genuine transition the publish runs once (from the new call above the early-return); the old call below the early-return was REMOVED, so there is no double write. Verified by the transition test asserting one write plus one `roleChange`/`promote` pair.
83
+ - **Races:** `writePollIntent` is atomic (tmp + rename), so the lifeline never reads a torn record — unchanged. The reconcile is single-threaded per process and already guarded by `leaseTicking`. The throttle fields are process-local and only mutated inside reconcile.
84
+ - **Feedback loops:** the record feeds the lifeline, which starts/stops polling, which changes the local 409 observation, which feeds `decidePollAction`. This change does not add a loop — it makes the existing one's input correct. The relevant risk (a machine flipping poll on/off repeatedly) is bounded by the existing start-debounce and by `POLL_INTENT_REFRESH_MS` throttling of unchanged values; a genuine flip publishes immediately, which is required for correctness (a lagging record is what caused the incident).
85
+ - **Churn breaker (B2):** deliberately left below the early-return so it still counts only true flips. Covered by the "does NOT re-fire transition-only side effects when the role is steady" test.
86
+
87
+ ---
88
+
89
+ ## 6. External surfaces
90
+
91
+ - **Other agents on the same machine:** none — the record is per-agent under that agent's `stateDir`.
92
+ - **Other users of the install base:** the change is inert unless `pollFollowsLease` resolves on. It ships dev-gated; on the fleet the guard resolves off and `writeLeasePollIntent` returns early, writing nothing. No fleet-visible change on this commit.
93
+ - **External systems:** no direct calls. Indirectly, on an agent with `pollFollowsLease.dryRun:false`, Telegram sees one long-poll instead of two competing ones — which is the intended end state and strictly fewer 409s.
94
+ - **Persistent state:** `state/telegram-poll-intent.json` only. Same schema, same fields, same atomic write. Written more often (bounded to once per 30s per unchanged value). No migration: the record is regenerated at runtime and readers already tolerate a missing/stale/corrupt file by treating it as "no opinion".
95
+ - **Timing / runtime conditions:** the throttle uses wall-clock `Date.now()`. A backwards clock step could delay a refresh by up to the step size; the consumer's response to a too-old record is "no opinion → hold", the safe direction, so a clock anomaly degrades to today's behavior rather than to a wrong action.
96
+ - **Operator surface (Mobile-Complete Operator Actions):** no operator-facing action added or changed. The existing levers (`pollOverride`, `telegramPolling`) are untouched.
97
+
98
+ ---
99
+
100
+ ## 6b. Operator-surface quality
101
+
102
+ **Not applicable** — this change touches no dashboard renderer, approval surface, notification body, or operator-facing markup. No operator-facing text is added or modified.
103
+
104
+ ---
105
+
106
+ ## 7. Multi-machine posture (Cross-Machine Coherence)
107
+
108
+ **Machine-local BY DESIGN.** The record answers a question that MUST differ per machine: "should THIS machine own the Telegram poll?" Exactly one machine may answer yes at a time, so replicating the record would be actively wrong — a replicated `shouldPoll:true` landing on a standby is precisely the dual-poll failure the fenced lease exists to prevent. The cross-machine coordination is carried by the fenced lease itself (already replicated via `LeaseCoordinator` / the mesh transports), and the record is the machine-local projection of that shared truth into a sibling process.
109
+
110
+ - **User-facing notices:** none emitted. No one-voice gating needed.
111
+ - **Durable state on topic transfer:** the record is not topic-scoped and holds no per-topic state, so it cannot strand on a transfer. It is regenerated from the lease on the next reconcile on whichever machine holds it.
112
+ - **Generated URLs:** none.
113
+
114
+ Worth naming explicitly: this defect is itself a multi-machine defect that a single-machine install can never surface (a single machine short-circuits to `awake` and there is no second poller to conflict with). The evidence for this review came from a live two-machine pair, which is the graduation gate the parent spec names for B1/B5.
115
+
116
+ ---
117
+
118
+ ## 8. Rollback cost
119
+
120
+ - **Hot-fix release:** revert the commit and ship as the next patch. One function plus one constant plus two fields.
121
+ - **Data migration:** none. The record is advisory and regenerated at runtime; readers already treat missing/stale as "no opinion".
122
+ - **Agent state repair:** none. No agent needs notifying or resetting.
123
+ - **User visibility:** none during the rollback window on the fleet (the guard resolves off there). On a dev agent with the gate on, reverting restores the previous behavior exactly: the boot default stands and the lifeline holds or fights, as before.
124
+
125
+ ---
126
+
127
+ ## Conclusion
128
+
129
+ The review changed the design twice. First, `writeLeasePollIntent` was changed to return a success boolean so the throttle can never record a skip-window off a write that failed — without that, a transient write failure would have suppressed retries for 30s, and the consumer would have aged into "no opinion" for a reason the code believed it had handled.
130
+
131
+ Second and more importantly: writing a test that runs the REAL boot path, rather than only the stubbed reconcile, surfaced a second independent defect — the safe boot default being written after the reconcile it was meant to precede. Shipping the early-return fix alone would have produced a green unit suite and an unchanged production failure. That is worth recording as the reason the extra test existed: the stubbed tests could not have caught it, because they never ran `initializeLease()`. The throttle-cache ownership move (into `writeLeasePollIntent`) came out of the same test failing a second time for a different reason.
132
+
133
+ Third, §2 forced an explicit statement of what this does NOT fix: the lifeline's boot-time poll-before-first-reconcile window and the dry-run `nobodyPollingRecovery` are both still open, both tracked on the attention item, and neither is claimed closed by this commit.
134
+
135
+ The change complies with signal-vs-authority: it corrects a detector's output and adds no authority. It is inert on the fleet (dev-gated producer) and its whole risk surface is confined to agents that have explicitly enabled `pollFollowsLease`.
136
+
137
+ ---
138
+
139
+ ## Second-pass review
140
+
141
+ **Not required.** Phase 5 triggers on changes that touch block/allow decisions on messaging or dispatch, session lifecycle, or anything named sentinel/guard/gate/watchdog. This change adds no block/allow surface and modifies no gate: it is confined to the producer side of an advisory record, and the consuming gate (`decidePollAction`) is untouched. The one adjacent trigger word — the lifeline's poll authority — is explicitly out of the diff.
142
+
143
+ Recorded for the reviewer's benefit if this judgment is revisited: the argument rests on the diff containing no change to `TelegramLifeline`, which is verifiable from the commit's file list.
144
+
145
+ ---
146
+
147
+ ## Evidence pointers
148
+
149
+ - Live measurement (instar-codey, Mac Mini, 2026-08-03 17:34 PDT, v1.3.1122): `/health → multiMachine.syncStatus` reporting `role:awake, holdsLease:true, leaseEpoch:21302` in the same second that `state/telegram-poll-intent.json` reported `role:standby, shouldPoll:false, leaseEpoch:21302`.
150
+ - Staleness half: the same record's age sampled at 10s intervals climbed 84.5s → 155.0s across an 80s window with no rewrite, i.e. past the consumer's 90s bound.
151
+ - Downstream harm on that agent: 812 × `Telegram 409 Conflict — another bot instance is polling`, 260 × `TelegramLifeline.selfRestart: conflict409Stuck`, server restarting every ~10 min for ~8h.
152
+ - Pre-fix test control: with `src/core/MultiMachineCoordinator.ts` reverted to `origin/main`, the headline regression test fails with `expected null not to be null` (no record written at all) — the right reason, not a missing-symbol error. 7 of 10 tests in the new file fail pre-fix; the 3 that pass are the gate-off guard, the real-transition guard, and the observe-only "does NOT acquire ⇒ left muted" guard, all of which are expected to hold either way.
153
+ - Second defect, found by the real-boot-path test rather than by reading: with only the early-return fixed, that test failed with the record on disk reading `{shouldPoll:false, role:'standby'}` while the in-process throttle key read `true|awake|1` — i.e. the correct value had been written and then overwritten by the boot default at the tail of `initializeLease()`. That divergence between the cache and the file is what identified the clobber.
154
+ - Attention item carrying the fleet risk and the mitigation that must be removed after deploy: `ATT-poll-intent-standby-forever-20260803`.
155
+
156
+ ---
157
+
158
+ ## Class-Closure Declaration
159
+
160
+ **Not applicable on both triggers, stated explicitly rather than omitted.**
161
+
162
+ 1. **Agent-authored-artifact defect?** No. The defect is in hand-written TypeScript (`MultiMachineCoordinator.ts`), not in an LLM prompt, hook, config, skill, or standards text.
163
+ 2. **Self-triggered controller added or modified?** No. `reconcileRoleToLease` is an existing control loop driven by the existing lease tick; this change adds no new loop, monitor, sentinel, reaper, scheduler, or recovery path, and fires no restart / swap / respawn / spawn / notify / retry / re-drive / kill. The throttle strictly REDUCES the action rate of an existing loop (a file write) and cannot increase it: an unchanged value is written at most once per `POLL_INTENT_REFRESH_MS`, and a changed value at most once per reconcile tick, which is the loop's own bound.
@@ -0,0 +1,108 @@
1
+ # Side-effects review — macos-memory-pressure-metric SIBLING (SessionManager)
2
+
3
+ **Change:** `SessionManager.currentMemoryPressure()` and `getSessionDiagnostics()`'s memory block now read
4
+ the corrected available-memory figure via `hostFreeMemPct()` instead of raw `os.freemem()`.
5
+ **Spec:** `docs/specs/macos-memory-pressure-metric.md` (converged 2026-06-26, `approved: true`) — this is
6
+ the caller that spec's original fix did not convert.
7
+ **Author:** echo · **Date:** 2026-08-04
8
+ **Sanction:** architect review 2026-08-04, under Justin's plan-scoped approval of 2026-08-03 20:21 PDT
9
+ (relayed 20:23). Escalated to Phase A critical path because it starves the grading job that answers rung 3.
10
+
11
+ ## What was wrong
12
+
13
+ `os.freemem()` on macOS counts only "Pages free". On a healthy 17 GB Mini it reports **0.46 GB (2.7%)**,
14
+ so `currentMemoryPressure()` computed 97% used and returned **`critical` permanently**. Under
15
+ `subscriptionPath.mode: 'force'`, `evaluateRerouteGate()` **throws** on elevated pressure — so every job
16
+ spawn was refused.
17
+
18
+ **Measured on the live host before the fix:**
19
+ - 20 of 27 enabled jobs failing, all with `Reroute refused (force-mode): host memory pressure is critical`
20
+ - `health-check` and `commitment-detection` at **421 consecutive failures** each, going back ≥2 days
21
+ - three readings of the same machine, same second: **raw 2.7% free → `critical`** · reaper (corrected)
22
+ **16.3% free → `normal`** · macOS `memory_pressure` **38% free**
23
+ - the same defect on the laptop: **137 GB RAM, 21 GB genuinely free → raw metric says `high`**
24
+
25
+ The corrected helper already existed in-package (`hostFreeMemPct`, used by
26
+ `HostPressureSampler`/`SessionReaper`) and its shipped comment warns against `os.freemem()` **by name**.
27
+ The reaper was fixed; this caller was not.
28
+
29
+ ## The 8 questions
30
+
31
+ **1. Over-block — what legitimate inputs does this reject that it shouldn't?**
32
+ None introduced; this change *removes* an over-block. Pre-fix the gate rejected **every** force-mode spawn
33
+ on macOS regardless of real conditions. Post-fix it rejects only genuinely elevated hosts. Thresholds
34
+ (90/75/60) are byte-identical — only the measurement source moved.
35
+
36
+ **2. Under-block — what failure modes does this still miss?**
37
+ The gate now trusts `hostFreeMemPct`, so it inherits that helper's limits: on an unrecognised platform it
38
+ falls back to `os.freemem()`-equivalent behaviour, and `vm_stat` parse failure degrades to the same. That
39
+ is the pre-existing, already-reviewed behaviour of the 2026-06-26 fix, not new exposure. **Genuinely
40
+ exhausted hosts still read `critical` — covered by a dedicated test so a "just return low" regression
41
+ cannot pass.**
42
+
43
+ **3. Level-of-abstraction fit.**
44
+ Correct layer. `currentMemoryPressure()` is the single definition of the pressure tier shared by the
45
+ reroute gate and the diagnostics surface (deliberately extracted in june15-headless-spawn-reroute PR2/O2).
46
+ Fixing it here fixes both consumers at once. The alternative — patching each caller — is what produced
47
+ this sibling in the first place.
48
+
49
+ **4. Signal vs authority compliance.** (`docs/signal-vs-authority.md`)
50
+ **Compliant, and it is the point of the change.** `currentMemoryPressure()` is a **detector**: it produces
51
+ a signal (`MemoryPressure` tier). `evaluateRerouteGate()` is the **authority** that decides. The defect was
52
+ a detector emitting a *false signal* to a correct authority. This change corrects the detector's
53
+ measurement and **adds no authority, moves no authority, and changes no threshold**. No new blocking logic
54
+ is introduced anywhere.
55
+
56
+ **5. Interactions — shadowing, double-fire, races.**
57
+ Two other `os.freemem()` callsites exist (`src/server/routes.ts:3743`, `src/monitoring/HealthChecker.ts:217`).
58
+ **Both were inspected and both already compensate** — each prefers `MemoryPressureMonitor`'s vm_stat-based
59
+ state and only falls back to `os.freemem()` when the monitor is absent, with a comment saying why. They are
60
+ correct as-is and are deliberately **not** touched. No shadowing: the reroute gate is the only consumer of
61
+ the tier for spawn decisions, and the diagnostics surface is read-only.
62
+
63
+ ⚠️ **Interaction worth naming:** `tests/unit/headless-spawn-reroute.test.ts` **stubs `currentMemoryPressure`
64
+ to `'normal'`**, with the comment that the real gate *"made this suite fail on loaded dev machines."* It was
65
+ not a loaded dev machine — **it was this defect, encountered, rationalised, and stubbed over**, which is why
66
+ CI never caught it. The stub is left in place (those tests assert reroute control-flow, not pressure) but
67
+ the new suite exercises the real method so it can now fail for the real reason.
68
+
69
+ **6. External surfaces.**
70
+ `getSessionDiagnostics()` payload changes: `usedPercent` and `freeMemMB` now report real available memory.
71
+ On macOS these numbers will **drop substantially** (e.g. 97% → ~40% used) — that is the correction, not a
72
+ regression. Pre-fix, once the tier was fixed alone, the surface could have reported **"97% used" beside tier
73
+ `low`**; both now read one source and a test pins that they cannot contradict. No timing or conversation-state
74
+ dependence.
75
+
76
+ **7. Multi-machine posture (Cross-Machine Coherence).**
77
+ **Machine-local BY DESIGN — correctly so.** Host memory pressure is a property of the physical machine; a
78
+ reading must never be replicated or proxied from a peer. Verified the defect independently on **both**
79
+ machines (Mini: raw `critical` vs corrected `normal`; laptop: raw `high` vs corrected `normal` with 21 GB
80
+ free). The fix therefore lands on both, and each machine keeps reading its own memory. No replication path,
81
+ no merged read, no cross-machine state. **This also removes a latent block on the ratified placement policy:**
82
+ worker lanes are to run on the laptop, and the same gate would have refused spawns there too.
83
+
84
+ **8. Rollback cost.**
85
+ Trivial. Revert the commit — no data migration, no agent-state repair, no persisted format change. The
86
+ change is confined to two computations inside one file plus one new test file. A hot-fix release restores
87
+ prior behaviour exactly (including, if anyone wants it, the permanent-`critical` bug).
88
+
89
+ ## Testing
90
+
91
+ New: `tests/unit/session-manager-memory-pressure-sibling.test.ts` (6 tests).
92
+
93
+ **Control run — 3 of 6 fail against pre-fix code, each for the right reason** (verified by stashing the
94
+ source change and re-running):
95
+ - `expected 'critical' not to be 'critical'` — the defect itself
96
+ - `expected [Function] to not throw but 'Error: Reroute refused (force-mode): …' was thrown` — **the exact
97
+ production error that killed 20 jobs**
98
+ - `expected 'critical' to be 'high'` — the Linux threshold path also mis-tiered
99
+
100
+ The other 3 pass either way **by design** — they are the guards proving the fix does not simply always
101
+ return `low` (genuinely-exhausted host still `critical`; gate still refuses it; diagnostics self-consistent).
102
+
103
+ Green: 48/48 across the new suite + `host-memory-pressure.test.ts` (16) + `headless-spawn-reroute.test.ts`
104
+ (26). `tsc --noEmit` clean.
105
+
106
+ ## Second-pass review
107
+
108
+ Required (touches a **gate**). See appended section below.