instar 1.3.800 → 1.3.802

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (83) hide show
  1. package/dashboard/index.html +25 -0
  2. package/dist/commands/server.d.ts.map +1 -1
  3. package/dist/commands/server.js +63 -1
  4. package/dist/commands/server.js.map +1 -1
  5. package/dist/core/BackupManager.d.ts.map +1 -1
  6. package/dist/core/BackupManager.js +8 -0
  7. package/dist/core/BackupManager.js.map +1 -1
  8. package/dist/core/PostUpdateMigrator.d.ts +8 -0
  9. package/dist/core/PostUpdateMigrator.d.ts.map +1 -1
  10. package/dist/core/PostUpdateMigrator.js +45 -0
  11. package/dist/core/PostUpdateMigrator.js.map +1 -1
  12. package/dist/core/ProactiveSwapMonitor.d.ts +10 -0
  13. package/dist/core/ProactiveSwapMonitor.d.ts.map +1 -1
  14. package/dist/core/ProactiveSwapMonitor.js +39 -0
  15. package/dist/core/ProactiveSwapMonitor.js.map +1 -1
  16. package/dist/core/SessionManager.d.ts +26 -4
  17. package/dist/core/SessionManager.d.ts.map +1 -1
  18. package/dist/core/SessionManager.js +76 -22
  19. package/dist/core/SessionManager.js.map +1 -1
  20. package/dist/core/machineCoherenceManifest.d.ts +10 -0
  21. package/dist/core/machineCoherenceManifest.d.ts.map +1 -1
  22. package/dist/core/machineCoherenceManifest.js +58 -0
  23. package/dist/core/machineCoherenceManifest.js.map +1 -1
  24. package/dist/core/types.d.ts +53 -0
  25. package/dist/core/types.d.ts.map +1 -1
  26. package/dist/core/types.js.map +1 -1
  27. package/dist/monitoring/ExternalHogScanTick.d.ts +9 -0
  28. package/dist/monitoring/ExternalHogScanTick.d.ts.map +1 -1
  29. package/dist/monitoring/ExternalHogScanTick.js +29 -0
  30. package/dist/monitoring/ExternalHogScanTick.js.map +1 -1
  31. package/dist/monitoring/PromiseBeacon.d.ts +6 -0
  32. package/dist/monitoring/PromiseBeacon.d.ts.map +1 -1
  33. package/dist/monitoring/PromiseBeacon.js +42 -6
  34. package/dist/monitoring/PromiseBeacon.js.map +1 -1
  35. package/dist/monitoring/guardManifest.d.ts.map +1 -1
  36. package/dist/monitoring/guardManifest.js +31 -0
  37. package/dist/monitoring/guardManifest.js.map +1 -1
  38. package/dist/monitoring/guardPosture.d.ts.map +1 -1
  39. package/dist/monitoring/guardPosture.js +22 -0
  40. package/dist/monitoring/guardPosture.js.map +1 -1
  41. package/dist/monitoring/selfaction/anchor.d.ts +82 -0
  42. package/dist/monitoring/selfaction/anchor.d.ts.map +1 -0
  43. package/dist/monitoring/selfaction/anchor.js +113 -0
  44. package/dist/monitoring/selfaction/anchor.js.map +1 -0
  45. package/dist/monitoring/selfaction/governor.d.ts +242 -0
  46. package/dist/monitoring/selfaction/governor.d.ts.map +1 -0
  47. package/dist/monitoring/selfaction/governor.js +1475 -0
  48. package/dist/monitoring/selfaction/governor.js.map +1 -0
  49. package/dist/monitoring/selfaction/policies.d.ts +86 -0
  50. package/dist/monitoring/selfaction/policies.d.ts.map +1 -0
  51. package/dist/monitoring/selfaction/policies.js +339 -0
  52. package/dist/monitoring/selfaction/policies.js.map +1 -0
  53. package/dist/monitoring/selfaction/types.d.ts +208 -0
  54. package/dist/monitoring/selfaction/types.d.ts.map +1 -0
  55. package/dist/monitoring/selfaction/types.js +11 -0
  56. package/dist/monitoring/selfaction/types.js.map +1 -0
  57. package/dist/scaffold/templates.d.ts.map +1 -1
  58. package/dist/scaffold/templates.js +8 -1
  59. package/dist/scaffold/templates.js.map +1 -1
  60. package/dist/server/AgentServer.d.ts.map +1 -1
  61. package/dist/server/AgentServer.js +6 -0
  62. package/dist/server/AgentServer.js.map +1 -1
  63. package/dist/server/CapabilityIndex.d.ts.map +1 -1
  64. package/dist/server/CapabilityIndex.js +9 -0
  65. package/dist/server/CapabilityIndex.js.map +1 -1
  66. package/dist/server/routes.d.ts +4 -0
  67. package/dist/server/routes.d.ts.map +1 -1
  68. package/dist/server/routes.js +271 -3
  69. package/dist/server/routes.js.map +1 -1
  70. package/dist/testing/selfActionRegistry.d.ts +16 -0
  71. package/dist/testing/selfActionRegistry.d.ts.map +1 -1
  72. package/dist/testing/selfActionRegistry.js +16 -0
  73. package/dist/testing/selfActionRegistry.js.map +1 -1
  74. package/package.json +2 -2
  75. package/scripts/lint-emit-without-admit.js +303 -0
  76. package/src/data/builtin-manifest.json +66 -66
  77. package/src/data/state-coherence-registry.json +1296 -1241
  78. package/src/scaffold/templates.ts +8 -1
  79. package/upgrades/1.3.801.md +38 -0
  80. package/upgrades/1.3.802.md +36 -0
  81. package/upgrades/session-listing-hygiene.eli16.md +26 -0
  82. package/upgrades/side-effects/self-action-governor.md +111 -0
  83. package/upgrades/side-effects/session-listing-hygiene.md +93 -0
@@ -14,7 +14,7 @@
14
14
  // existing-agent migration there) so the two can never drift. Imported as a runtime
15
15
  // function call inside generateClaudeMd — no module-init cycle (PostUpdateMigrator
16
16
  // never imports templates).
17
- import { PLAYWRIGHT_PROFILE_REGISTRY_CLAUDEMD_SECTION, MACHINE_LOAD_ASSESSMENT_CLAUDEMD_SECTION, DYNAMIC_MCP_CLAUDEMD_SECTION, SENDER_REJECTION_CLAUDEMD_SECTION, SCOPE_ACCRETION_CLAUDEMD_SECTION, MESH_SELF_HEALING_CLAUDEMD_SECTION, WRITE_ADMISSION_CLAUDEMD_SECTION, DOORWAY_REGISTRY_CLAUDEMD_SECTION, EXTERNAL_HOG_CLAUDEMD_SECTION, ROUTING_SPEND_CLAUDEMD_SECTION } from '../core/PostUpdateMigrator.js';
17
+ import { SESSION_LISTING_HYGIENE_CLAUDEMD_SECTION, PLAYWRIGHT_PROFILE_REGISTRY_CLAUDEMD_SECTION, MACHINE_LOAD_ASSESSMENT_CLAUDEMD_SECTION, DYNAMIC_MCP_CLAUDEMD_SECTION, SENDER_REJECTION_CLAUDEMD_SECTION, SCOPE_ACCRETION_CLAUDEMD_SECTION, MESH_SELF_HEALING_CLAUDEMD_SECTION, WRITE_ADMISSION_CLAUDEMD_SECTION, DOORWAY_REGISTRY_CLAUDEMD_SECTION, EXTERNAL_HOG_CLAUDEMD_SECTION, ROUTING_SPEND_CLAUDEMD_SECTION } from '../core/PostUpdateMigrator.js';
18
18
 
19
19
  export interface AgentIdentity {
20
20
  name: string;
@@ -883,6 +883,7 @@ ${MESH_SELF_HEALING_CLAUDEMD_SECTION(port)}
883
883
  ${WRITE_ADMISSION_CLAUDEMD_SECTION(port)}
884
884
  ${DOORWAY_REGISTRY_CLAUDEMD_SECTION(port)}
885
885
  ${ROUTING_SPEND_CLAUDEMD_SECTION(port)}
886
+ ${SESSION_LISTING_HYGIENE_CLAUDEMD_SECTION(port)}
886
887
  **Per-Feature LLM Metrics & LLM Activity (Observable Intelligence)** — Audit what each of your LLM-driven gates/sentinels actually does: WHICH provider + model ran it, how often it ACTED (fired) vs found nothing (noop), how often it was skipped to save rate limits (shed), cost, and latency. This is the *Observable Intelligence* standard — no autonomous AI action the system takes is allowed to be invisible. Read-only observability — it never gates anything.
887
888
  - Check: \`curl -H "Authorization: Bearer $AUTH" "http://localhost:${port}/metrics/features?sinceHours=24"\`
888
889
  - Returns \`{ totals, features: [{ feature, frameworks, models, byModel, calls, realCalls, tokensIn, tokensOut, tokensCached, fired, noop, shed, fireRate, p50LatencyMs, p95LatencyMs, ... }] }\` — one row per system (e.g. MessagingToneGate, MessageSentinel). \`frameworks\`/\`models\` = which provider(s) actually served the call; \`fireRate\` = how often it acts; \`shed\` = skipped by the rate-limit guard. Filter with \`?feature=<name>\`.
@@ -901,6 +902,12 @@ ${ROUTING_SPEND_CLAUDEMD_SECTION(port)}
901
902
  - Tune via \`.instar/config.json\` → \`intelligence.spawnCap\` (\`maxConcurrent\`, \`acquireMs\`, \`waitersMax\`) or env (\`INSTAR_HOST_SPAWN_MAX\`, \`INSTAR_SPAWN_ACQUIRE_MS\`, \`INSTAR_SPAWN_WAITERS_MAX\`). Restart sessions/server to apply.
902
903
  - **When to use** (PROACTIVE): "are we protected against a fork-bomb / OOM?" / "how many LLM spawns are running right now?" / "why did a gate hold under load?" → \`GET /spawn-limiter\`. (Spec: \`docs/specs/forkbomb-prevention-simple.md\`; constitution: "Bounded Blast Radius".)
903
904
 
905
+ **Self-Action Backpressure Governor (unified self-action chokepoint)** — Every registered self-triggered action I take (reaper age-kills, external-hog kills, proactive account swaps, beacon notify/liveness lines) rides ONE admission chokepoint (\`SelfActionGovernor\`) carrying per-target + census-scaled total count ceilings, rate buckets, P19 brakes, and a bounded coalescing queue — the runtime arm of the "Capacity Safety — No Unbounded Self-Action" standard (the 17,503-kills/day reaper flood + the 72-swaps/day thrash are the ancestor incidents). It ships OBSERVE-ONLY on every class: it measures would-deny verdicts and blocks NOTHING; a class only enforces after the operator's deliberate per-class flip (and pool-shared classes never enforce on a multi-machine pool until the pool-wide ceiling exists).
906
+ - Status: \`curl -H "Authorization: Bearer $AUTH" http://localhost:${port}/self-action-governor\` → per-class \`{ mode, counters, bySubMechanism, queueDepth }\`; every non-allow NAMES its deciding layer (per-target-ceiling / total-ceiling / census-scale / rate-bucket / breaker / ...). \`?scope=pool\` merges pool-shared class counters across my machines.
907
+ - **When to use** (PROACTIVE — these are the triggers): "why did my respawn get held?" / "why did my swap get queued?" / "why did my notify get folded?" → read that class's \`bySubMechanism\` reasons on \`GET /self-action-governor\` — the deciding layer is named, never guessed.
908
+ - **Mass-incident valve (the operator's path)**: in a real fire (a mass cleanup the ceilings would pace), the PRIMARY path is CONVERSATIONAL — the operator tells me and I set \`intelligence.selfActionGovernor.emergencyDisable: true\` in \`.instar/config.json\` (read live, no restart; every class degrades to unconditional pass-through). The flip itself is audited AND raises an attention item in both directions. Disabling via \`PATCH /config\` additionally requires the dashboard PIN (re-enable is Bearer-OK); a raw config-file edit remains the deliberate verifier-independent floor.
909
+ - A human action always wins: operator kill routes carry an ALWAYS-ALLOW, always-audited principal lane — an enforcing class can never count-deny or queue an emergency stop. (Spec: \`docs/specs/unified-self-action-backpressure.md\`.)
910
+
904
911
  **Test-Runner Concurrency Bound (host-wide vitest cap — the spawn cap's sibling)** — A per-machine ticket counter bounds how many test suites run AT ONCE across every actor on this machine: full suites run one-at-a-time (default cap 1), while small targeted runs (≤5 named test files) get a roomier lane (default 6 slots, each clamped to ≤4 workers). It is the structural answer to the 2026-07-02 test-storm meltdown (29 concurrent vitest roots ≈ 300+ workers starving co-resident servers' event loops until their supervisors killed healthy processes). Ships WATCH-ONLY (dry-run) for a 14-day soak — it records what it WOULD have blocked but admits every run; blocking arrives only after the soak review flips the host tuning file.
905
912
  - Status: \`curl -H "Authorization: Bearer $AUTH" http://localhost:${port}/test-runner-limiter\` → \`{ cap, targetedCap, posture, ttlSignalArmed, liveHolders, targetedHolders, admittedOpen, suite: {available, saturated}, targeted: {...}, recentEvents, skipHistogram }\` (Registry First — read it, never guess).
906
913
  - **"Why is my test run waiting?" / a rejected \`git push\`** (PROACTIVE — this is the trigger): a push or suite that stalls or is refused may be CONTENTION (another suite holds the slot), NOT red tests — read \`GET /test-runner-limiter\` BEFORE assuming failure. The limiter's capacity-timeout error says "this is NOT a test failure" and names the holders.
@@ -0,0 +1,38 @@
1
+ # Upgrade Guide — vNEXT
2
+
3
+ <!-- assembled-by: assemble-next-md -->
4
+ <!-- bump: minor -->
5
+
6
+ ## What Changed
7
+
8
+ Unified Self-Action Backpressure, Increment B (docs/specs/unified-self-action-backpressure.md; the normative companion `unified-self-action-backpressure.companion.md` is the implementation authority; CMT-1911/CMT-1928). ONE in-process admission chokepoint — the `SelfActionGovernor` (`src/monitoring/selfaction/`) — that every registered self-triggered action rides via `admit()`: per-target + census-scaled total count ceilings (fixed-bucket sliding windows, no epoch reset for relief classes), token-bucket rate ceilings, P19 brakes, a bounded coalescing queue with drain-time re-validation (incarnation fence + eligibility predicate), single-consume capability tokens (runtime consume at the sink is the authority), an ALWAYS-ALLOW audited principal lane (a human action always wins; PIN-distinguishable from bare Bearer at the operator kill routes), durable admission state that survives restarts (event-aware eager flush — a crash-loop bouncing faster than the flush debounce still accretes the floor), and a per-class fail matrix (cost/safety fails CLOSED-to-QUEUE; relief fails OPEN-with-audit paced by a config-immune last-resort floor; respawn-recovery fails OPEN unconditionally).
9
+
10
+ **Ships OBSERVE-ONLY on every class, fleet-wide (FD1):** admit() records would-deny verdicts and blocks NOTHING. The per-class enforce flip is the operator's later deliberate action (FD8), and pool-shared classes (swap/notify) additionally auto-demote whenever the registered machine count exceeds one (FD9) — no pool-shared enforce exists in this increment. The retrofit is ADDITIVE: none of the incident-earned bespoke brakes (AgeKillBackoff, swap anti-thrash, beacon suppression, the external-hog kill ledger) is removed or weakened.
11
+
12
+ Retrofitted in this increment: the five registry-modeled controllers — the age-limit kill path (SessionManager), the proactive account swap (ProactiveSwapMonitor, braked + legacy paths), the PromiseBeacon progress heartbeat + liveness line (two controllers, one file), and the external-hog kill path (ExternalHogScanTick). Remaining emit sites land as staged follow-up PRs <!-- tracked: CMT-1911 -->.
13
+
14
+ Enforcement tooling: a new codebase-wide usage-scan lint (`scripts/lint-emit-without-admit.js`, wired into `npm run lint`) binds controller identity at registration + sink (marker↔file↔registry `modelsPath` licensing, no dynamic ids, no handle export/pass-as-value, principal API import-restricted, admit targets must be canonical `deriveTargetKey` derivations). Observability: `GET /self-action-governor` (lock-free scrubbed read; `?scope=pool` merges pool-shared class counters), a GUARD_MANIFEST entry with synthetic enabled-polarity posture (`intelligence.selfActionGovernor.enabled` computed from the inverted kill-switch), three COHERENCE_CRITICAL_FLAGS rows (inverted governor row + live-read pool-shared class-mode rows via a governor-state accessor on the advert view), a transitions-only audit stream, and six Standard-B operator notices (demote-exhaustion alarm, coalesced dead-letter shed, errored-posture alarm, emergencyDisable flip, principal volume page, observe-limbo nudge). Config: `intelligence.selfActionGovernor` (live-read `emergencyDisable` kill-switch + sparse per-class overrides, validated at load); the PATCH /config path gets a nested-path validator scoped to exactly that subtree with deep merge, and the DISABLE direction is dashboard-PIN-gated.
15
+
16
+ ## What to Tell Your User
17
+
18
+ - I now measure every self-triggered action I take — session cleanups, account swaps, my own status notices — against one shared safety meter, the same way my process spawns are already capped. Nothing is blocked yet: this ships in watch mode, gathering evidence first.
19
+ - If you ever wonder why a cleanup was held or a swap was queued once enforcement is turned on class by class, I can name the exact rule that decided it — nothing is silently dropped.
20
+ - Your actions always win: an emergency stop or a kill you order rides an always-allowed lane that the meters can never pace.
21
+ - There is one master off switch for the whole brake, and flipping it is loud on purpose — you get told, because a disabled safety brake is itself an incident.
22
+
23
+ ## Summary of New Capabilities
24
+
25
+ | Capability | How to Use |
26
+ |---|---|
27
+ | Self-action admission posture (per-class modes, counters, deciding-layer reasons) | `GET /self-action-governor` (Bearer); `?scope=pool` for pool-shared classes across machines |
28
+ | Guard posture visibility | `GET /guards` row `intelligence.selfActionGovernor.enabled` (synthetic enabled polarity; load-bearing) |
29
+ | Machine-coherence mode-skew alarm inputs | advert rows `selfActionGovernor.emergencyDisable` + per-class `…mode` (live-read) |
30
+ | Kill-switch + per-class overrides | `intelligence.selfActionGovernor.emergencyDisable` (live-read) / `…classes.<id>.*`; PATCH /config nested validator (disable direction PIN-gated) |
31
+ | Usage-scan lint | `node scripts/lint-emit-without-admit.js` (in `npm run lint`) |
32
+
33
+ ## Evidence
34
+
35
+ - Tier 1: `tests/unit/self-action-governor.test.ts` (34 tests — admission battery, fail matrix, census, tokens, queue, demote latch, FD9 gate), `self-action-governor-snapshot.test.ts` (8 — durable floor across bounces incl. sub-debounce crash-loop), `self-action-governor-anchor.test.ts` (5 — dual-load collision + attach), `self-action-token-coverage.test.ts` (9 — sink inventory), `lint-emit-without-admit.test.ts` (15 — every lint rule + the real tree passes clean), and the generalized convergence ratchet (`self-action-convergence.test.ts` now drives every registered controller through the governor in enforce mode).
36
+ - Tier 2: `tests/integration/self-action-governor-route.test.ts` (10 — real routes pipeline, lock-free pure read, scrubbed projection, pool scope, nested PATCH validator + PIN gate both directions).
37
+ - Tier 3: `tests/e2e/self-action-governor-alive.test.ts` (4 — production init path serves 200 with live counters; guard-posture + coherence view-seam wiring integrity).
38
+ - Observe-only safety: every retrofitted emit path verified byte-equivalent in behavior under observe mode (admit always allows; sink guards proceed); full regression sweep over PromiseBeacon / SessionManager / ProactiveSwapMonitor / external-hog / coherence-manifest / capability suites green.
@@ -0,0 +1,36 @@
1
+ # Upgrade Guide — vNEXT
2
+
3
+ <!-- assembled-by: assemble-next-md -->
4
+ <!-- bump: patch -->
5
+
6
+ ## What Changed
7
+
8
+ Session-listing hygiene (CMT-1936; live evidence 2026-07-09: the Mac Mini's `GET /sessions` answered 53 rows of which 52 were finished background runs — 22 `mentor-stage-a-*` headless one-shots + 28 `job-*` records — read by the operator as "duplicate sessions running across both machines"). Three parts:
9
+
10
+ 1. **Bounded finished-session retention** (`SessionManager.cleanupStaleSessions`): `failed` records are now pruned (previously retained FOREVER); a terminal record with a missing/unparseable `endedAt` falls back to `startedAt` instead of being skipped forever (the second unbounded hole); headless one-shots (`launchLane:'headless'` — the mentor Stage-A shape) get the 60-min background TTL instead of the 24-h interactive TTL; the hard cap (default 50, oldest-ended first) now counts every terminal class. All knobs config-tunable via `sessions.retention` (`killedTtlMinutes` / `completedJobTtlMinutes` / `completedTtlHours` / `maxFinished`), applied at the next server restart (SessionManager snapshots config at boot) — absence preserves shipped defaults.
11
+ 2. **Active-by-default listing** (`GET /sessions`): the default answer is ACTIVE sessions only (`starting`/`running`); `?include=all` returns the full registry; `?status=<valid>` keeps its exact prior semantics. The `scope=pool` fan-out forwards the caller's opt-in to peers and defensively filters a LEGACY peer's full-registry answer (the server-side twin of the dashboard's 2026-06-11 client-side filter), and `pool.machines[].sessionCount` now counts the requested view — no more finished-record-inflated machine counts.
12
+ 3. **Genuine cross-machine duplicate flag**: the pool view computes `pool.duplicateTopics` — the SAME conversation (platform + platformId) with a LIVE session on ≥2 machines at once — tags each such row `duplicateTopic: true`, and the dashboard badges it red. The same recurring job on each machine (benign, by design) is never flagged.
13
+
14
+ Agent Awareness + Migration Parity: new `Session Listing Hygiene` CLAUDE.md section shipped in `generateClaudeMd` and appended to existing agents via `migrateClaudeMd` (content-sniffed, idempotent).
15
+
16
+ ## What to Tell Your User
17
+
18
+ - Your session list now shows what is actually RUNNING. Finished background chores no longer pile up in the view — on 2026-07-09 one machine showed "53 sessions" when only 1 was really running; that misread is structurally gone.
19
+ - Finished session records are cleaned up on a real schedule now (background runs after an hour, conversations after a day, hard cap 50) — and you can tune every window in config if you want longer history.
20
+ - If the SAME conversation is ever genuinely live on two of your machines at once — the real incoherency — the dashboard flags it with a red "duplicate" badge instead of leaving you to guess. Matching job names across machines are each machine's own scheduled copy and are deliberately not flagged.
21
+
22
+ ## Summary of New Capabilities
23
+
24
+ | Capability | How to Use |
25
+ |---|---|
26
+ | Active-only session listing (default) | `GET /sessions` (Bearer) — running/starting only |
27
+ | Full registry incl. finished runs | `GET /sessions?include=all` (or `?status=completed|failed|killed`) |
28
+ | Genuine cross-machine duplicate flag | `GET /sessions?scope=pool` → `pool.duplicateTopics` + per-row `duplicateTopic:true`; red badge on the dashboard sessions list |
29
+ | Finished-record retention knobs | `.instar/config.json` → `sessions.retention` (`killedTtlMinutes`, `completedJobTtlMinutes`, `completedTtlHours`, `maxFinished`) — applies at the next server restart |
30
+
31
+ ## Evidence
32
+
33
+ - Tier 1: `tests/unit/SessionManager-retention.test.ts` (9 behavioral tests — both sides of every TTL boundary, failed/killed pruning, endedAt→startedAt fallback, unparseable-timestamp expiry, running never touched, maxFinished oldest-first cap, config knobs live, nonsensical config falls back to defaults); `tests/unit/SessionManager-injection.test.ts` hard-cap block updated to the new implementation.
34
+ - Tier 2: `tests/integration/sessions-listing-hygiene.test.ts` (8 tests — default active-only, include=all, status semantics preserved, invalid status → default, legacy-peer defensive filter + honest machine counts, include=all forwarding, genuine duplicate flagged both rows, benign shapes NOT flagged even under include=all).
35
+ - Tier 3: `tests/e2e/sessions-listing-hygiene-lifecycle.test.ts` (5 tests — real AgentServer + real on-disk records on the production init path: default view, include=all, status filter, additive `pool.duplicateTopics`, Bearer auth).
36
+ - Regression sweep: sessions-pool-scope (integration + e2e), remote-session-close, sessions-launch-lane, dashboard poolTileStatusFilter/sessionMachineBadge, route-validation-edge, route-completeness, all PostUpdateMigrator suites, scaffold-templates, capabilities-discoverability, server-full — green.
@@ -0,0 +1,26 @@
1
+ # Session Listing Hygiene — plain-English overview
2
+
3
+ ## What this actually is
4
+
5
+ When you ask your agent "what sessions are running?" — or open the dashboard — the answer was polluted: on 2026-07-09 the Mac Mini reported 53 sessions, and 52 of them were FINISHED background chores (little 5-minute mentor runs and scheduled job checks) that had already ended but were still sitting in the list. Both machines also run the same scheduled jobs on purpose, so the wall of near-identical names looked exactly like "the same session running twice across my machines" — a scary bug that wasn't actually happening.
6
+
7
+ This change makes the session list tell the truth in three ways:
8
+
9
+ 1. **The list shows what's RUNNING.** `GET /sessions` now answers with active sessions only. The finished ones aren't deleted from view forever — add `?include=all` and you get the whole registry, exactly as before.
10
+ 2. **Finished records get cleaned up on a real schedule.** Finished background runs are pruned after an hour, finished conversations after a day, and there's a hard cap of 50 retained records no matter what. Two genuine leaks are fixed: records of FAILED sessions were never cleaned up at all, and a record missing its end-timestamp was kept forever. Every window is tunable in config (`sessions.retention`) if you want longer history.
11
+ 3. **A REAL duplicate now shouts.** The cross-machine view computes the one case that actually matters: the SAME conversation with a LIVE session on two machines at once. That gets a red "duplicate" badge on the dashboard and a `pool.duplicateTopics` entry in the API. The benign look-alike — each machine running its own copy of a scheduled job — is never flagged, because that's how the system is designed to work.
12
+
13
+ ## What already existed
14
+
15
+ A cleanup pass already pruned some finished records (jobs after 1 h, conversations after 24 h, cap 50) — but the mentor-run shape slipped into the 24 h bucket, `failed` records slipped through entirely, and the listing itself never distinguished finished from running. The dashboard privately filtered finished rows out of its tiles; the API (what the agent itself and the pool view read) did not.
16
+
17
+ ## The safeguards, in plain terms
18
+
19
+ - Nothing running is ever touched — only records of sessions that already ended.
20
+ - Old machines and new machines can mix during rollout: the merged view filters an old machine's unfiltered answer, so no wall of stale rows sneaks back in.
21
+ - The duplicate badge is a signal only. It never kills or blocks anything — the existing safety layers own that.
22
+ - Rollback is cheap: config knobs restore any retention window; the listing change reverts with the PR.
23
+
24
+ ## What you actually need to decide
25
+
26
+ Nothing — defaults are chosen to match the existing behavior everywhere except the two leak fixes and the mentor-run class (24 h → 1 h, the exact accumulation that caused the misread). If you want finished runs kept longer, set `sessions.retention.completedJobTtlMinutes` (or `completedTtlHours` / `maxFinished`) in `.instar/config.json` — it takes effect at the next server restart (the session manager reads its config once, at boot).
@@ -0,0 +1,111 @@
1
+ # Side-Effects Review — SelfActionGovernor (unified self-action backpressure, Increment B)
2
+
3
+ **Version / slug:** `self-action-governor`
4
+ **Date:** `2026-07-10`
5
+ **Author:** `Echo (instar-dev agent)`
6
+ **Second-pass reviewer:** `dedicated reviewer subagent (high-risk: governor/gate surface)`
7
+
8
+ ## Summary of the change
9
+
10
+ Builds the runtime primitive the converged spec `docs/specs/unified-self-action-backpressure.md` (approved 2026-07-05; the companion `unified-self-action-backpressure.companion.md` is the implementation authority) defines: `SelfActionGovernor` (`src/monitoring/selfaction/{types,policies,anchor,governor}.ts`) — ONE in-process admission chokepoint for self-triggered actions, keyed on controller id, with count/rate/breaker ceilings, a bounded coalescing queue, capability tokens, a principal lane, durable admission state, and a per-class fail matrix. Ships OBSERVE-ONLY on every class, fleet-wide (FD1): admit() records would-verdicts and ALWAYS allows. Retrofits the five registry-modeled controllers additively (SessionManager age-kill, ProactiveSwapMonitor both paths, PromiseBeacon heartbeat + liveness, ExternalHogScanTick kill). Adds `GET /self-action-governor`, the nested-path PATCH /config validator (PIN-gated disable direction), guard-posture + coherence rows, the `lint-emit-without-admit` usage-scan lint, registry field additions (`modelsPath`, `delegatedGiveUp`), CLAUDE.md template section + `migrateClaudeMd`, state-registry/retention/backup-exclusion declarations, and the three-tier test battery (122 new unit tests + 10 integration + 4 e2e + the generalized convergence ratchet).
11
+
12
+ ## Decision-point inventory
13
+
14
+ - `SelfActionGovernor.admit()/admitSync()` — **add** — the new admission gate; in observe mode (the ONLY shipped mode) it never blocks: every verdict resolves to an allow-token, would-denies are recorded.
15
+ - `consumeAdmissionToken()` sink guards (5 sites) — **add** — signal-only in observe mode (`proceed: true` always); blocking exists only behind a per-class enforce flip no fleet config sets.
16
+ - `PATCH /config` — **modify** — adds a nested-path validator branch for `intelligence.selfActionGovernor` (previously the whole `intelligence` key 400'd); the disable direction is PIN-gated.
17
+ - `DELETE /sessions/:id` + `POST /sessions/:name/remote-close` — **pass-through** — an optional PIN-proof header records principal provenance on the always-allow audited lane; the kill itself is untouched (same terminateSession call, same origin stamp).
18
+ - Retrofitted emit paths (age-kill, proactive swap ×2, beacon send, hog kill) — **modify** — an observe-mode admit + token consume is inserted BEFORE each existing emit; all existing brakes retained (additive retrofit, LA8-1).
19
+ - P17 attention funnel — **pass-through** — six governor notices ride the injected `createAttentionItem` seam (the existing AttentionTopicGuard chokepoint), never a new send path.
20
+
21
+ ---
22
+
23
+ ## 1. Over-block
24
+
25
+ **What legitimate inputs does this change reject that it shouldn't?**
26
+
27
+ In the SHIPPED state (observe-only, every class): nothing — every admit resolves to an allow-token, including under a throwing governor (observe branch of the fail matrix allows), an uninitialized governor (disabled-passthrough), and a dead dual-load handle (errored-open → allow for observe). The three enforce-only rejection surfaces (policy deny → queue/coalesce, sink `proceed:false`, last-resort errored floor) are structurally unreachable without a per-class `mode: 'enforce'` config override no default sets — and pool-shared classes auto-demote at registered machine count > 1 even then (FD9). One PATCH-surface narrowing is deliberate: `PATCH /config` with `intelligence.<anything-but-selfActionGovernor>` now 400s with a named reason — previously the whole `intelligence` key 400'd generically, so no previously-legal input is newly rejected. The PIN gate rejects a bare-Bearer `emergencyDisable: true` — deliberate (ADV9-4), and two verifier-independent valves remain (conversational config edit; raw file edit).
28
+
29
+ ## 2. Under-block
30
+
31
+ **What failure modes does this still miss?**
32
+
33
+ - Observe mode measures but bounds nothing — a live runaway loop today is still bounded only by its bespoke brake (stated FD1 posture; the enforce flip is the operator's later action, nudged by FD12 observe-limbo + the inverse storm nudge).
34
+ - The ~18 non-retrofitted emit sites don't ride admit() yet (staged follow-up PRs <!-- tracked: CMT-1911 -->); the lint binds only sites that USE the governor — it cannot force a site that never imports it (that forcing job belongs to the existing `lint-no-unregistered-self-action`, report-only until the classClosure flip).
35
+ - The usage-scan lint is a string-level guard: deliberate obfuscation (`const g = governor; g['f'+'or'](...)`) evades it — same honest coverage limit the sibling lint states; the runtime single-mint + sink consume are the backstops.
36
+ - Coordinated local-fs deletion of snapshot + aggregates degrades to fresh-install posture (spec-accepted residual, SEC7-5).
37
+ - The PIN tier defeats buggy runaway loops, not a deliberately adversarial self that can read machine-local config (spec-stated threat-model honesty, ADV7-4; covering controls: per-admit principal audit + volume page).
38
+
39
+ ## 3. Level-of-abstraction fit
40
+
41
+ **Is this at the right layer?**
42
+
43
+ Yes — this IS the layer consolidation the spec exists for: the generalization of the host-spawn-semaphore (which stays at the provider layer, never re-acquired — FD4) and P17 (which stays the notify coalescer; the governor's notify classes fold INTO it, and all six notices ride the existing attention funnel rather than a parallel send path). It composes existing primitives (token bucket, P19 breaker, count ceiling, bounded queue) behind one Admission contract instead of re-implementing any; the registry's proven `boundK`/`perTargetBoundK` seed the runtime ceilings; the ExternalHogKillLedger's (key, classId, keyIsVolatile) triple is mirrored, not replaced. The existing bespoke brakes stay where they are (additive retrofit) — the governor is defense-in-depth above them, not a replacement.
44
+
45
+ ## 4. Signal vs authority compliance
46
+
47
+ **Does this hold blocking authority with brittle logic?**
48
+
49
+ Reference: `docs/signal-vs-authority.md`. In the shipped state the governor is a pure SIGNAL producer: would-deny aggregates, transitions audit, six P17-funneled notices — zero blocking authority anywhere. The blocking authority it CAN hold (per-class enforce) is (a) deterministic count/rate arithmetic — the sanctioned deterministic-ratchet class, not heuristic content judgment; (b) gated behind a deliberate per-class operator flip (FD8) with a review-gated promotion criterion (FD12); (c) fail-safed per class so a broken governor never blocks relief (open-with-audit + config-immune floor), never strands cost/safety work (closed-to-QUEUE, never drop), and never touches a human action (the always-allow principal lane). The `emit-without-admit` lint is a deterministic build-time ratchet (the sanctioned lint class, twin of `lint-no-unbounded-llm-spawn`). No brittle string-matching holds runtime blocking authority.
50
+
51
+ ## 5. Interactions
52
+
53
+ **Does it shadow another check, get shadowed, double-fire, race?**
54
+
55
+ - Bespoke brakes (AgeKillBackoff, swap anti-thrash, beacon suppression, hog kill-ledger) run FIRST; the governor sees only what they let through — double-bounding is intended (tightest bound wins), and in observe mode the governor changes nothing they decide.
56
+ - The governor's admit sits between the KEEP-guard veto path and terminateSession in SessionManager — in observe mode it cannot flip a keep/kill decision; the ReapAuthority funnel and its lease/KEEP gates are untouched.
57
+ - The P17 funnel is the single notice path — the six governor notices are budget-subject like every other attention source (no new topic-creation surface).
58
+ - Anchor single-mint vs vitest module graphs: test environments get a per-graph anchor (key-salt design) so unrelated test files can never cross-collide; production uses the real `Symbol.for` global.
59
+ - The slow tick (60s, unref'd) samples census/config/drains queues — it reads `state.listSessions()` through the existing memoized cache, no new hot-path I/O; `admitSync` is zero-I/O by construction except the debounce-EXEMPT eager flushes (the once-per-boot leading-edge rehydrate flush + at most one half-ceiling-crossing flush per class per window — both spec-mandated event-aware flush edges, ADV7-2/SC6-2; each a try/catch-wrapped ~few-KB temp+rename that can never affect the admission outcome).
60
+ - PATCH /config: the sag branch runs BEFORE the generic allowlist check and removes its keys from the generic loop — no double-application; `pin` is stripped so it can never land in config.
61
+
62
+ ## 6. External surfaces
63
+
64
+ **Anything visible to other agents/users/systems?**
65
+
66
+ - New Bearer route `GET /self-action-governor` (scrubbed: no target identities, no absolute quota values); `?scope=pool` fans out to peers' same route (rate-limited 6/min, URL-allowlist-guarded, dark-peer-tolerant — the /guards pattern).
67
+ - Six operator notices, all attention-funnel-bound and episode-latched/coalesced — worst-case notice volume is bounded by construction (dedupe keys per companion §8).
68
+ - The advert grows three coherence rows (~120 bytes) — within the MC byte budgets (ratchet tests pass); older peers treat unknown keys as version skew (the designed path).
69
+ - Timing dependence: `emergencyDisable` live-read is cached ≤1s; flip observation latency ≤ max(1s, next admit/slow-tick) — documented in the template section ("read live, no restart").
70
+
71
+ ## 7. Multi-machine posture (Cross-Machine Coherence)
72
+
73
+ - Hardware-bound class state (age-kill, hog, respawn): **machine-local BY DESIGN** — the resource is this host's (`machine-local-justification: hardware-bound-resource`); counters served machine-local.
74
+ - Pool-shared classes (swap/notify): counters **proxied-on-read** via `?scope=pool`; their window buckets are declared `unified` in the spec's taxonomy — the LOCAL half of the FD15-replicated surface whose replication half is the deferred pool-ceiling deliverable <!-- tracked: CMT-1911 -->; until then they are structurally prevented from enforcing on a multi-machine pool (FD9 auto-demote, level-triggered on REGISTERED count, re-promote only on de-enrollment).
75
+ - Mode coherence: per-class mode rows + the inverted governor row join `COHERENCE_CRITICAL_FLAGS`, read LIVE via the governor-state accessor on the advert view — cross-machine mode skew raises the standard machine-coherence alarm.
76
+ - Durable files (snapshot + aggregates) are BackupManager-excluded via `BLOCKED_PATH_PREFIXES` — a foreign restore would carry the wrong machine's counts / fabricate prior-flush evidence.
77
+ - One-voice: notices route through this machine's attention funnel only (no cross-machine sends). No generated URLs. No durable state strands on topic transfer (admission state is per-machine by design; queued intents are in-memory level-triggered regenerables with restart-shed honesty).
78
+
79
+ ## 8. Rollback cost
80
+
81
+ **If this turns out wrong in production, what's the back-out?**
82
+
83
+ Cheap, layered: (1) `intelligence.selfActionGovernor.emergencyDisable: true` — live-read, no restart, degrades every class to unconditional pass-through (the flip is audited + surfaced, by design); (2) per-class numeric/mode overrides for a single misbehaving class; (3) full revert — the retrofit is five small additive blocks + wiring, no data migration, no config migration (`migrateConfig` writes nothing), no schema change; deleting the two state files after a revert is safe (they are advisory admission history). The CLAUDE.md template section is content-sniffed append-only (idempotent, non-destructive). Worst production wedge identified: none blocking — observe mode cannot withhold an action; the only new synchronous work on hot paths is in-memory arithmetic plus the once-per-boot flush.
84
+
85
+ ---
86
+
87
+ ## Second-pass review
88
+
89
+ **Reviewer subagent verdict (independent audit, 2026-07-10):** Concur with the review.
90
+
91
+ - Observe-only cannot withhold — verified on every path: uninitialized deps / emergencyDisable → unconditional allow (`admitFor`); dead mint-collision handle + throwing `evaluate` route to `failDisposition`, which returns allow for any non-enforce mode and allow-unconditional for the respawn-recovery lane; a policy deny in observe/demoted mode returns allow via `nonAllow` (would-deny recorded). `consumeToken` returns `proceed: !enforcing` on every rejection. Shipped defaults set no enforce anywhere (`freshClassState` starts observe; mode overrides come only from config).
92
+ - All five retrofits are additive, non-throwing, and sit before the emit: SessionManager (non-allow → continue; ageKillBackoff + KEEP-guard/ReapAuthority untouched); ProactiveSwapMonitor ×2 (non-allow → continue inside existing try blocks; anti-thrash/deferral/pile-on brakes intact); PromiseBeacon (non-allow → fold, hot-state cadence still advances so a fold can't tight-loop); ExternalHogScanTick (non-allow → alert-only + surfaceLeftAlive, never silent). The diff removes zero brakes.
93
+ - PATCH /config confinement holds: sibling `intelligence.*` keys 400 before anything is written; the disable direction rides `checkMandatePin` (sha256 + timingSafeEqual + durable rate-limited lockout); `pin` can never land in config on either path; the sag subtree is removed from the generic merge loop (no double-application).
94
+ - Anchor test isolation cannot leak into production: the module-local symbol resolves only under VITEST/NODE_ENV=test; the key-salt override requires an explicit globalThis assignment; production resolves the real `Symbol.for` key; `resetAnchorForTest` throws outside a test env without force.
95
+ - No unlisted material risk found: all six notices are episode/window-latched and ride the P17 attention seam (per-source topic budgets bound a pathological reopen ping-pong); audit buffer (512), file retention (5,000 rows), and token map (4,096) are bounded; a throwing flush can never reroute an admission. `lint-emit-without-admit` passes over the full tree (1,494 files, 0 violations).
96
+ - One precision nuance (folded above, non-blocking): the eager-flush wording in §5 now names BOTH debounce-exempt synchronous flush edges (leading-edge post-rehydrate + once-per-window half-ceiling), not just the boot one.
97
+
98
+ ---
99
+
100
+ ## Class-Closure Declaration
101
+
102
+ - **defectClass:** `unbounded-self-action` — **closure:** `guard` — **enforcement:** `ratchet` — **citation:** `tests/unit/self-action-convergence.test.ts`
103
+ - **How caught (convergence argument):** the generalized convergence ratchet drives every registered controller's worst-case emissions THROUGH `SelfActionGovernor.admit()` in enforce mode and asserts the steady-state action count stays bounded at each model's proven `boundK` — horizon-independent (fixed-bucket sliding windows with NO episode reset for relief classes; the demote latch re-promotes only after a clean cooldown dwell, so the floor cannot flap). Eternal sentinels stay rate-floored (never count-bounded, never starved), and the P19 breaker + per-target/total count ceilings make every retrofitted loop converge rather than storm. The `lint-emit-without-admit` usage-scan lint is the completeness half: a controller cannot mint/borrow a looser class's ceiling by construction.
104
+
105
+ ## Addendum (same PR, parity commit)
106
+
107
+ The delivery-completeness parity guard caught that the new CLAUDE.md governor section was absent from `migrateFrameworkShadowCapabilities` markers[] — added (plus the featureSections registry entry), so Codex/Gemini agents receive the capability block too. No runtime surface beyond the migrator's existing marker-mirroring mechanics; idempotent by the same content-sniff.
108
+
109
+ ## Addendum 2 (same PR, CI-ratchet conformance commit)
110
+
111
+ Two CI ratchets caught conformance gaps, both folded: (1) the no-silent-fallbacks ratchet — the governor's boot-init catch in `src/commands/server.ts` now reports through DegradationReporter (never-silent degradation; admits resolve disabled-passthrough for the boot, behavior unchanged); (2) the G3 load-bearing manifest lint — the `intelligence.selfActionGovernor.enabled` GUARD_MANIFEST entry now declares its uniform `soakWindowDays`/`declaredLoadBearingAt` fields (the guard reports enabled/never-dryRun, so no gap/soak posture arises in the shipped observe-only state).
@@ -0,0 +1,93 @@
1
+ # Side-Effects Review — Session-listing hygiene (bounded finished-session retention + active-by-default listing + genuine-duplicate flag)
2
+
3
+ **Version / slug:** `session-listing-hygiene`
4
+ **Date:** `2026-07-10`
5
+ **Author:** Instar Agent (echo)
6
+ **Second-pass reviewer:** required (session lifecycle surface) — see appended response
7
+
8
+ ## Summary of the change
9
+
10
+ Fixes the "duplicate sessions" perception problem (CMT-1936, topic 29836; live evidence 2026-07-09: the Mac Mini's `GET /sessions` returned 53 rows of which 52 were FINISHED background runs — 22 `mentor-stage-a-*` headless one-shots, 28 `job-*` records — read by the operator as "duplicate sessions running across both machines"). Three parts:
11
+
12
+ 1. **Bounded retention** (`src/core/SessionManager.ts` `cleanupStaleSessions()`, config in `src/core/types.ts` `SessionManagerConfig.retention`): closes two genuinely UNBOUNDED holes — `failed` records were never pruned at all, and a terminal record with a missing/unparseable `endedAt` was skipped forever (and escaped the hard cap). Headless one-shots (`launchLane === 'headless'`, the mentor Stage-A shape) now get the short background TTL (60 min) instead of the 24 h interactive TTL. The hard cap now counts EVERY terminal class (completed + failed + killed), default 50. All knobs config-tunable via `sessions.retention` (applied at the next server restart — SessionManager snapshots config at boot; absence preserves shipped defaults).
13
+ 2. **Active-by-default listing** (`src/server/routes.ts` `GET /sessions`): the default view is ACTIVE sessions only (`starting`/`running`); `?include=all` returns the full registry; `?status=<valid>` keeps its exact pre-change semantics. The `scope=pool` fan-out forwards the caller's opt-in to peers AND defensively filters a LEGACY peer's full-registry answer (mirror of the dashboard's existing client-side filter, now structural).
14
+ 3. **Genuine-duplicate flag** (`routes.ts` pool branch + `dashboard/index.html`): the pool view computes `pool.duplicateTopics` — the SAME conversation (platform + platformId) with a LIVE session on ≥2 machines at once — tagging each such row `duplicateTopic: true` and badging it red on the dashboard. The same recurring job on each machine (benign, by design) is never flagged (job/headless sessions carry no platform binding).
15
+
16
+ Agent awareness: new `Session Listing Hygiene` CLAUDE.md section (`SESSION_LISTING_HYGIENE_CLAUDEMD_SECTION` in `src/core/PostUpdateMigrator.ts`), used by `generateClaudeMd` (`src/scaffold/templates.ts`) and appended content-sniffed in `migrateClaudeMd` (Migration Parity).
17
+
18
+ ## Decision-point inventory
19
+
20
+ - `SessionManager.cleanupStaleSessions()` — modify — which terminal session records are removed from the on-disk registry, and when.
21
+ - `GET /sessions` default visibility filter (routes.ts) — add — which rows the listing answers with by default (active vs full registry).
22
+ - `GET /sessions?scope=pool` remote-row defensive filter — add — drops a legacy peer's finished rows from the default merged view.
23
+ - `pool.duplicateTopics` computation — add — SIGNAL-ONLY classification (flags, never blocks or kills).
24
+ - Dashboard duplicate badge — add — pure rendering of the server-computed signal.
25
+
26
+ ---
27
+
28
+ ## 1. Over-block
29
+
30
+ **What legitimate inputs does this change reject that it shouldn't?**
31
+
32
+ - The default `GET /sessions` no longer shows finished runs. A caller that legitimately wants them (post-run inspection, history tooling) must pass `?include=all` or `?status=<terminal>`. Consumer trace (2026-07-10, this worktree): NO in-repo consumer reads finished rows from the ROUTE — the dashboard's local tiles come from the running-only WebSocket feed, its remote tiles already client-filter to running/starting (`dashboard/index.html`, the 2026-06-11 closed-sessions-reappearing fix), and the closeout-liveness snapshot excludes terminal entries itself (`src/monitoring/closeoutLivenessSnapshot.ts` TERMINAL_STATUSES). Internal registry consumers call `state.listSessions()` directly and are untouched.
33
+ - Retention: headless one-shot records now prune after 60 min instead of 24 h. A human inspecting "what did that mentor run do?" more than an hour later loses the registry row (the tmux transcript/log surfaces remain). Judged acceptable — it is the exact accumulation class that caused the misread, and the window is tunable (`sessions.retention.completedJobTtlMinutes`).
34
+ - A terminal record with NO parseable timestamps is pruned immediately. Legitimate records always carry `startedAt` (set at spawn) — only malformed garbage matches.
35
+
36
+ ## 2. Under-block
37
+
38
+ **What failure modes does this still miss?**
39
+
40
+ - A GENUINE duplicate involving a session that has not yet enriched a platform binding (e.g. a topic session spawned but not yet registered in the adapter's topic map) is not flagged — platformId is the join key. The flag is best-effort observability, not the safety mechanism (the lease/one-voice layers own prevention).
41
+ - The duplicate flag only surfaces in the POOL view (it needs the cross-machine merge). A caller reading each machine's plain `/sessions` separately re-derives nothing — documented in the CLAUDE.md section ("read `pool.duplicateTopics` first").
42
+ - Mixed-version window: an UPDATED peer queried by an OLD machine's fan-out (no query forwarded) answers active-only — the old machine's pool view loses the finished rows it used to show. That is the intended direction (fewer stale rows), never data loss (records remain on the peer's disk behind `?include=all`).
43
+
44
+ ## 3. Level-of-abstraction fit
45
+
46
+ The default-visibility filter lives at the ROUTE (the shared read chokepoint every surface uses: dashboard poll, agent curl, pool fan-out) — the same place the `?status=` filter already lived, so it is the established layer for listing semantics. Retention lives in the ONE existing pruner (`cleanupStaleSessions`) rather than a new janitor. The duplicate flag lives in the ONE place that sees all machines' rows (the pool merge). No parallel mechanism added anywhere. No issue identified.
47
+
48
+ ## 4. Signal vs authority compliance
49
+
50
+ Compliant (`docs/signal-vs-authority.md`). Nothing here holds blocking authority over agent behavior: the listing filter changes what a READ answers (full data remains one flag away); `duplicateTopics` is a pure SIGNAL (flags, never kills/blocks — the reaper/lease layers keep their own authority and their own evidence bars); retention prunes only records that are already terminal, through the existing single-writer state path (`state.removeSession`). The brittle-ish join (platform+platformId string key) carries zero blocking authority — exactly the shape the principle prescribes.
51
+
52
+ ## 5. Interactions
53
+
54
+ - **OrphanProcessReaper** (`listKnownTmuxSessions()` reads ALL statuses from DISK): pruning removes names from its known-set. Exposure judged pre-existing and marginal: jobs already pruned at 60 min pre-change; the newly-shortened class (headless one-shots, ≤5 min lifetime, tmux killed at completion detection) and the newly-pruned class (`failed`, previously immortal) change the window, not the mechanism. The route-level filter does NOT affect it (disk reads, not route reads).
55
+ - **Dashboard client-side remote filter** (index.html) and the new server-side defensive filter double-fire harmlessly (both keep active rows; belt-and-suspenders during the mixed-version window — deliberately kept).
56
+ - **`POST /sessions/cleanup-stale`** route and the 5-min monitor tick both call the same pruner — unchanged single implementation, no race added (removeSession is idempotent).
57
+ - **Resume queue / reap-log / token ledger**: none read the registry's terminal rows (verified in the consumer trace); reap-log is its own append log.
58
+ - **`pool.machines[].sessionCount`** now counts the REQUESTED view (active by default) — this un-inflates the Machines tab count and lets the WS4.2 empty-state fire for a machine with only finished records ("online — no active sessions"), which is more honest, not less.
59
+
60
+ ## 6. External surfaces
61
+
62
+ - API shape: plain route stays an ARRAY; pool response gains ONE additive field (`pool.duplicateTopics`, always an array) and rows gain an optional `duplicateTopic: true`. Back-compatible for every existing consumer shape-wise; the SEMANTIC default change (active-only) is the deliberate, documented fix.
63
+ - Cross-version: new fan-out × old peer → defensive filter keeps semantics; old fan-out × new peer → active-only (intended direction). No error paths added; a peer failure still degrades to `pool.failed`.
64
+ - No timing/conversation-state dependence beyond what the route already had.
65
+
66
+ ## 6b. Operator-surface quality
67
+
68
+ The dashboard change is ONE additive badge on the existing sessions list — no new form, tab, or flow.
69
+
70
+ 1. **Leads with its primary action?** Unchanged — the sessions list still leads with the session tiles (name → click to stream); the badge only appears on a flagged row, next to the existing machine badge.
71
+ 2. **Zero raw internals as primary content?** Yes — the badge text is the plain word "⚠ duplicate" with a plain-English hover title ("this conversation has a live session on more than one machine"); no ids, no JSON, no config keys. The underlying platformId/machineIds stay API-only.
72
+ 3. **Destructive actions de-emphasized?** No destructive action added — the badge is pure information; the existing close button is untouched.
73
+ 4. **Plain language at phone width?** Yes — one short lowercase word in the same badge row that already wraps on mobile (same sizing as the machine badge, 10px pill); the removal of ~50 finished-session tiles from the pool poll actively IMPROVES phone usability (the operator's 2026-07-05 screenshot complaint was exactly this wall).
74
+
75
+ No issue identified.
76
+
77
+ ## 7. Multi-machine posture (Cross-Machine Coherence)
78
+
79
+ Proxied-on-read: the pool merge (`GET /sessions?scope=pool`) is the merged read, now with consistent visibility semantics forwarded to peers and defensively enforced for legacy peers. Retention is machine-local BY DESIGN (each machine prunes its own registry — the records describe that machine's tmux sessions). The duplicate flag is exactly the cross-machine coherency read this spec family called for. User-facing notices: none added (signal renders in dashboard + API only). No durable state strands on topic transfer (records are per-machine lifecycle facts, not conversation state).
80
+
81
+ ## 8. Rollback cost
82
+
83
+ Config-only partial rollback: `sessions.retention` restores any TTL/cap (e.g. `{"completedJobTtlMinutes": 1440}` ≈ the pre-change 24 h window for the headless class) — **taking effect at the next server restart** (SessionManager snapshots config at boot; a config edit alone does NOT stop the 5-min pruning tick — restart to apply, and every claim site in code/docs/CLAUDE.md says so after the second-pass fix below). The listing default and duplicate flag roll back by reverting the PR (pure route/rendering logic, no data migration, no persisted format change — pruned records are gone, but they were ephemeral lifecycle rows that the pre-change code also pruned on its own schedule). No agent state repair needed.
84
+
85
+ ---
86
+
87
+ ## Second-pass review
88
+
89
+ **Concern raised (first pass):** the artifact and four shipped claim sites (types.ts config comment, the cleanupStaleSessions docstring, the fleet CLAUDE.md section, the unit-test title) described `sessions.retention` as "read live each pass — no restart". That was FALSE at runtime: `server.ts` snapshots `{ ...config.sessions }` once at boot and `SessionManager.config` has no reload path, while `PATCH /config` replaces `ctx.config.sessions` with a NEW object the SessionManager never sees. Load-bearing because §8's rollback lever would have let an operator believe pruning stopped while the 5-min tick kept irreversibly deleting records until a restart.
90
+
91
+ **Resolution (iterated before commit):** all four claim sites plus the ELI16, release fragment, and this artifact now state the true semantics — a `sessions.retention` change takes effect at the NEXT SERVER RESTART. No live-reload machinery added (out of scope; the boot-snapshot pattern is the established SessionManager config contract).
92
+
93
+ **Reviewer verdict on everything else:** verified accurate — the default filter preserves `?status=<valid>` semantics exactly; the legacy-peer defensive filter mirrors the default AND enforces `explicitStatus` on peer rows; `duplicateTopics` is active-rows-only in both aggregation and flagging, excludes `headless`, and mutates only the enriched copies (never registry state objects); OrphanProcessReaper's known-set dependency fails in the SAFE direction (a pruned name declassifies a leftover process to external/kept — never killed); ResumeQueue entries are self-contained (topicId/resumeUuid/jobSlug in the durable entry), so no revival path reads terminal registry rows past the new TTLs; signal-vs-authority holds; no new race between the 5-min tick and `POST /sessions/cleanup-stale`. Concur with the review as amended.