@ouro.bot/cli 0.1.0-alpha.721 → 0.1.0-alpha.723

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/changelog.json CHANGED
@@ -1,6 +1,20 @@
1
1
  {
2
2
  "_note": "This changelog is maintained as part of the PR/version-bump workflow. Agent-curated, not auto-generated. Agents read this file directly via read_file to understand what changed between versions.",
3
3
  "versions": [
4
+ {
5
+ "version": "0.1.0-alpha.723",
6
+ "changes": [
7
+ "Anchor iMessage reactions to the message they point at (tapback name, add vs remove, target excerpt and authorship) instead of a bare \"reacted with love\" stub.",
8
+ "Stop the orientation correction hold from blocking explicitly requested actions: reactions never arm it, a correction word inside a long self-contained instruction with nothing to disambiguate no longer arms it, and an armed hold now names its trigger, states what clears it, and emits orientation.correction_hold_armed.",
9
+ "ouro doctor mail-ingest liveness is now Mailroom-mode aware: on the hosted Azure Blob store it measures the hosted reader's local mirror (state/mail-search) instead of the frozen local messages/ directory, and reports an explicit unverified state when no local signal exists, so a hosted agent receiving mail today no longer fails with a fabricated multi-week outage."
10
+ ]
11
+ },
12
+ {
13
+ "version": "0.1.0-alpha.722",
14
+ "changes": [
15
+ "ouro doctor now asserts pipeline liveness (mail ingest age, sense inbound-delivery age) so a configured-but-dead pipe fails loudly instead of reporting healthy."
16
+ ]
17
+ },
4
18
  {
5
19
  "version": "0.1.0-alpha.721",
6
20
  "changes": [
@@ -320,14 +334,14 @@
320
334
  {
321
335
  "version": "0.1.0-alpha.673",
322
336
  "changes": [
323
- "A2A connect + delegation layer (friends @ouro.bot/friends alpha.7), on the merged inbound-auth foundation. DID-keyed onboarding: a DID-bearing agent card keys the friend record on its verified did:key (externalId == DID, a2a.did == DID) after verifyCardDidBinding + parseDidKey, so an inbound sealed envelope resolves to that same record (the owner's own agent no longer collapses to stranger); a did-less legacy card stays URL-keyed (backward compatible). New owner-only connect_to tool, gated to the local management sense via authorizeConnect (network/group/internal senses are refused before any card fetch; closed-sense membership fails closed without the roster verifier) and linking at family via connectAgents. Outbound becomes signed+sealed over the direct rung: a new A2ATransport (relay/mailbox are a typed-but-stubbed seam naming ourostack/friends-relay), a2a_send_message seals via sendShare for DID-keyed peers (legacy text send preserved), and a full delegation round-trip — connect_to → coordinate(request, task) (prepareCoordination mints the requestId + first-party delegation, sealed to the peer; the recipient's receiveShare quarantines it under importedDelegations) → list_delegations → send_result (prepareMissionResult over the harness-owned mission_result DataPart wire, which routes to importMissionResult, NOT receiveShare, since mission-result is not a FriendsKind) → importMissionResult lands importedResults[B][requestId], assignee/correlation-gated (assignee_mismatch / no_delegation / no_mission / untrusted_source, fail-closed on a legacy assignee-less delegation). Imported delegations and results land quarantined + attributed, never seed trust, never overwrite first-party; deliver-back is explicit (autonomy is a non-goal).",
324
- "Security fix (mission_result forgery): the harness-owned mission_result wire now authenticates the sender exactly like the inbound share path. Previously receiveInboundMissionResult trusted the envelope's self-asserted fromAgentId (openSealedEnvelope does AEAD decryption only — no signature verification) and called importMissionResult with no verifier (falling back to the no-op TOFU verifier), so a result signed with an attacker's own key but claiming a victim assignee's DID was accepted and stored under importedResults[victim][requestId] as if the victim delivered it (the recipient's X25519 pubkey is public, so AEAD is no barrier, and the assignee/correlation gates only checked the claimed plaintext DID). The wire now mirrors receiveShare: replay-dedups on the seal nonce (durable SeenLedger), rejects unless opened.signerDid === opened.fromAgentId (sender_binding_mismatch), resolves+pins the sender DID, and hands importMissionResult a real DidVerifier bound to the pinned key so the Ed25519 signature is verified against the pinned key (a non-matching signature is untrusted_source). All authorization gates remain enforced by the importer."
337
+ "A2A connect + delegation layer (friends @ouro.bot/friends alpha.7), on the merged inbound-auth foundation. DID-keyed onboarding: a DID-bearing agent card keys the friend record on its verified did:key (externalId == DID, a2a.did == DID) after verifyCardDidBinding + parseDidKey, so an inbound sealed envelope resolves to that same record (the owner's own agent no longer collapses to stranger); a did-less legacy card stays URL-keyed (backward compatible). New owner-only connect_to tool, gated to the local management sense via authorizeConnect (network/group/internal senses are refused before any card fetch; closed-sense membership fails closed without the roster verifier) and linking at family via connectAgents. Outbound becomes signed+sealed over the direct rung: a new A2ATransport (relay/mailbox are a typed-but-stubbed seam naming ourostack/friends-relay), a2a_send_message seals via sendShare for DID-keyed peers (legacy text send preserved), and a full delegation round-trip \u2014 connect_to \u2192 coordinate(request, task) (prepareCoordination mints the requestId + first-party delegation, sealed to the peer; the recipient's receiveShare quarantines it under importedDelegations) \u2192 list_delegations \u2192 send_result (prepareMissionResult over the harness-owned mission_result DataPart wire, which routes to importMissionResult, NOT receiveShare, since mission-result is not a FriendsKind) \u2192 importMissionResult lands importedResults[B][requestId], assignee/correlation-gated (assignee_mismatch / no_delegation / no_mission / untrusted_source, fail-closed on a legacy assignee-less delegation). Imported delegations and results land quarantined + attributed, never seed trust, never overwrite first-party; deliver-back is explicit (autonomy is a non-goal).",
338
+ "Security fix (mission_result forgery): the harness-owned mission_result wire now authenticates the sender exactly like the inbound share path. Previously receiveInboundMissionResult trusted the envelope's self-asserted fromAgentId (openSealedEnvelope does AEAD decryption only \u2014 no signature verification) and called importMissionResult with no verifier (falling back to the no-op TOFU verifier), so a result signed with an attacker's own key but claiming a victim assignee's DID was accepted and stored under importedResults[victim][requestId] as if the victim delivered it (the recipient's X25519 pubkey is public, so AEAD is no barrier, and the assignee/correlation gates only checked the claimed plaintext DID). The wire now mirrors receiveShare: replay-dedups on the seal nonce (durable SeenLedger), rejects unless opened.signerDid === opened.fromAgentId (sender_binding_mismatch), resolves+pins the sender DID, and hands importMissionResult a real DidVerifier bound to the pinned key so the Ed25519 signature is verified against the pinned key (a non-matching signature is untrusted_source). All authorization gates remain enforced by the importer."
325
339
  ]
326
340
  },
327
341
  {
328
342
  "version": "0.1.0-alpha.672",
329
343
  "changes": [
330
- "A2A inbound authentication foundation (friends @ouro.bot/friends alpha.7): the agent mints/loads an Ed25519 did:key self-identity (machine-local vault seed) and serves it on its agent card; an inbound friends sealed DataPart is now unwrapped, unsealed, signature-verified, and replay-checked, and the turn runs keyed on the VERIFIED sender DID at that peer's real trust level (read from the friend store by DID, defaulting to stranger) — replacing the previous unauthenticated-a2a-peer collapse. Adds durable harness-owned PinStore (state/a2a/pins, TOFU + signed-rotation evaluation) and SeenLedger (state/a2a/seen, restart-safe replay defense). The legacy unsigned-text inbound path is behavior-preserved (stranger)."
344
+ "A2A inbound authentication foundation (friends @ouro.bot/friends alpha.7): the agent mints/loads an Ed25519 did:key self-identity (machine-local vault seed) and serves it on its agent card; an inbound friends sealed DataPart is now unwrapped, unsealed, signature-verified, and replay-checked, and the turn runs keyed on the VERIFIED sender DID at that peer's real trust level (read from the friend store by DID, defaulting to stranger) \u2014 replacing the previous unauthenticated-a2a-peer collapse. Adds durable harness-owned PinStore (state/a2a/pins, TOFU + signed-rotation evaluation) and SeenLedger (state/a2a/seen, restart-safe replay defense). The legacy unsigned-text inbound path is behavior-preserved (stranger)."
331
345
  ]
332
346
  },
333
347
  {
@@ -391,7 +405,7 @@
391
405
  {
392
406
  "version": "0.1.0-alpha.662",
393
407
  "changes": [
394
- "Parse `--github-token` and `--base-url` in `ouro hatch` so the GitHub Copilot provider can be connected headlessly at cold-start. The credential type, per-provider validation, and storage split already supported github-copilot (githubToken -> vault credential, baseUrl -> config); only the argv parsing was missing, which left the provider selectable but unconnectable from a non-interactive hatch — the exact path Ouro Workbench's provider-config form uses for the work agent.",
408
+ "Parse `--github-token` and `--base-url` in `ouro hatch` so the GitHub Copilot provider can be connected headlessly at cold-start. The credential type, per-provider validation, and storage split already supported github-copilot (githubToken -> vault credential, baseUrl -> config); only the argv parsing was missing, which left the provider selectable but unconnectable from a non-interactive hatch \u2014 the exact path Ouro Workbench's provider-config form uses for the work agent.",
395
409
  "Override the transitive `shell-quote` dependency to `^1.8.4` (pulled in via `ink` -> `react-devtools-core`) to clear the newly-disclosed GHSA-w7jw-789q-3m8p critical advisory that was failing the release-preflight `npm audit` gate on every PR."
396
410
  ]
397
411
  },
@@ -405,7 +419,7 @@
405
419
  {
406
420
  "version": "0.1.0-alpha.660",
407
421
  "changes": [
408
- "Add `ouro mcp-serve --workbench-mcp [<path>]` so Ouro Workbench injects its `ouro_workbench` MCP into the boss agent's turn at runtime — merged per-turn and per-agent in the daemon with no cross-agent leak — instead of writing a machine-specific entry into the git-synced agent bundle.",
422
+ "Add `ouro mcp-serve --workbench-mcp [<path>]` so Ouro Workbench injects its `ouro_workbench` MCP into the boss agent's turn at runtime \u2014 merged per-turn and per-agent in the daemon with no cross-agent leak \u2014 instead of writing a machine-specific entry into the git-synced agent bundle.",
409
423
  "Drop the Workbench MCP bundle-write from the boss path (the explicit `ouro connect workbench` opt-in escape hatch is unchanged) and update the agent self-model copy so the boss understands it receives `ouro_workbench` via runtime injection rather than treating the missing bundle entry as a blocked or trust-level state."
410
424
  ]
411
425
  },
@@ -563,7 +577,7 @@
563
577
  "BlueBubbles follow-up hardening: callback transport activity now has a bounded per-operation watchdog, late turn results after timeout are suppressed instead of recording stale success, and coverage locks the close/drop paths so stuck status or cleanup transports cannot keep the chat lane occupied.",
564
578
  "Package lifecycle hardening: `npm pack` now runs the full build and package-asset verifier through `prepack`, so local tarballs cannot capture stale compiled `dist/` output after source fixes. The package asset verifier now has a CLI entrypoint used by the lifecycle script.",
565
579
  "Package asset verifier payload-boundary hardening: prepack verification now scans only the roots declared in `package.json.files` plus package metadata, so local worktrees, source folders, coverage reports, and other non-payload developer artifacts cannot block a clean tarball while stale text inside shipped assets is still rejected.",
566
- "Cozy narrative pass on the `## my desk` prompt section that every ouro agent reads every turn. Same information, warmer voice: the desk is now described as a quiet room of the agent's work — tracks line one wall like drawers in a wide cabinet, friction notes pin to a corkboard, lessons sit on a small reference shelf by the window, and finished work slides into the back, still browsable and still mine. The intent: an agent reading this every turn feels at home rather than briefed. Dynamic-block labels softened to match (`nearest the front of the desk:` instead of `FEATURED:`; `also open on the desk:` instead of `other active tracks:`; `tasks still open:` instead of `non-terminal tasks:`; empty-state reads `the desk is quiet today — no tracks yet. a good time to lay something down.`). Information density unchanged; only the texture shifts. Tests + snapshot updated to match the new strings; 26 desk-section tests + 282 prompt tests all green.",
580
+ "Cozy narrative pass on the `## my desk` prompt section that every ouro agent reads every turn. Same information, warmer voice: the desk is now described as a quiet room of the agent's work \u2014 tracks line one wall like drawers in a wide cabinet, friction notes pin to a corkboard, lessons sit on a small reference shelf by the window, and finished work slides into the back, still browsable and still mine. The intent: an agent reading this every turn feels at home rather than briefed. Dynamic-block labels softened to match (`nearest the front of the desk:` instead of `FEATURED:`; `also open on the desk:` instead of `other active tracks:`; `tasks still open:` instead of `non-terminal tasks:`; empty-state reads `the desk is quiet today \u2014 no tracks yet. a good time to lay something down.`). Information density unchanged; only the texture shifts. Tests + snapshot updated to match the new strings; 26 desk-section tests + 282 prompt tests all green.",
567
581
  "Package e2e smokes now run installed `ouro` binaries with an isolated temp HOME/USERPROFILE so `ouro --version` and `ouro help` no longer create `~/AgentBundles/default.ouro` daemon logs on the developer machine; package-e2e tests cover the isolated env.",
568
582
  "Adds the first Ouro evolution-loop substrate: durable EvolutionCase/EvolutionTrace state under each agent bundle, harness_friction packet binding by friction signature, coding_spawn evolutionCaseId binding with budget/authority enforcement, active-work surfacing for open cases, eight local evolution tools (status/case/capture/decide/verify/deliver/ratify/close) with nerves telemetry, and prompt guidance for evidence-first self-fix work. The flow now captures evidence, checks budget and authority before delegation and merge-sensitive actions, records verification and delivery state, and requires ratification before closure; GEPA-style optimization remains deferred until trace quality exists. Tests add store, packet, coding, active-work, tool, prompt, and end-to-end coverage, plus daemon CLI timing stabilizers uncovered by full coverage."
569
583
  ]
@@ -571,7 +585,7 @@
571
585
  {
572
586
  "version": "0.1.0-alpha.635",
573
587
  "changes": [
574
- "**Critical follow-up**: `McpManager.reconcile()` now reads from the same merged-config helper as `getSharedMcpManager()`'s initial start, so plugin-declared MCP servers (`<plugin-root>/.mcp.json`) survive across turns. Pre-fix bug: reconcile() read only `config.mcpServers` (agent.json builtins), saw plugin servers as 'removed', and tore them down — so the first turn's `mcp__desk__*` tools would disappear on the second turn (`Unknown server: desk`). Surfaced during desk plugin validation: every other call was failing. New `buildMergedServerConfig()` helper module-level; reused by both code paths. Regression test in mcp-manager-plugin-merge.test.ts asserts plugin servers survive across N reconcile cycles with zero shutdowns. Bundle-skeleton contract test relaxed: the per-agent `plugins` key is excluded from the cross-agent bundle-key parity check (plugins are agent-specific by design)."
588
+ "**Critical follow-up**: `McpManager.reconcile()` now reads from the same merged-config helper as `getSharedMcpManager()`'s initial start, so plugin-declared MCP servers (`<plugin-root>/.mcp.json`) survive across turns. Pre-fix bug: reconcile() read only `config.mcpServers` (agent.json builtins), saw plugin servers as 'removed', and tore them down \u2014 so the first turn's `mcp__desk__*` tools would disappear on the second turn (`Unknown server: desk`). Surfaced during desk plugin validation: every other call was failing. New `buildMergedServerConfig()` helper module-level; reused by both code paths. Regression test in mcp-manager-plugin-merge.test.ts asserts plugin servers survive across N reconcile cycles with zero shutdowns. Bundle-skeleton contract test relaxed: the per-agent `plugins` key is excluded from the cross-agent bundle-key parity check (plugins are agent-specific by design)."
575
589
  ]
576
590
  },
577
591
  {
@@ -583,31 +597,31 @@
583
597
  {
584
598
  "version": "0.1.0-alpha.633",
585
599
  "changes": [
586
- "`ouro migrate-to-desk` migrator — new CLI command that copies an agent's legacy `tasks/` tree into desk-shape under `<bundle>/desk/`. **COPY semantics, not move** — the source `tasks/` directory is left intact for dual-read safety during transition; the operator triggers final deletion later (out of scope here). New `src/repertoire/desk/classifier.ts` exposes a pure-data classifier (`classifyFile()`, `deriveTaskSlug()`, `extractParentTaskDir()`, `resolveUpdatedMs()`) implementing the triage rules — terminal (done/approved/complete/cancelled/fixed/etc.) → archive bucket; live status with `updated` older than 30 days (cutoff 2026-04-22) → stale_live archive bucket; missing/unknown status / `.md.bak` / empty → ambiguous archive bucket; live status within 30 days → live_clear → migrate to a default `legacy` track; the one-off `ongoing/2026-03-09-1410-summer-2026-europe-trip.md` (a special-cased file from the initial bundle that motivated this migrator) → special_europe_trip lift. Effective `updated` resolves via YAML `updated` → `approved` → `created` → body `**Updated**:` → date-prefix in filename → file mtime. Pairing rules: planning/doing/ideation/audit siblings sharing a slug group together; if any sibling is terminal → all terminal, else if any is live_clear → all live_clear; sub-files under `<taskdir>/` inherit from the parent. Files under `tasks/archive/` are unconditionally terminal regardless of frontmatter. New `src/heart/daemon/migrate-to-desk.ts` exposes `runMigrateToDesk()` — walks `<bundle>/tasks/**`, classifies, groups, builds a deterministic plan, then writes archive bucket to `desk/_archive/<original-relative-path>` (preserving structure; `schema_version: 1` set on every touched markdown via `ensureSchemaVersion()`); live_clear tasks to `desk/legacy/<slug>/task.md` + `iterations/`; europe-trip task to `desk/summer-2026-europe-trip/` with two task scaffolds (`book-replacement-outbound`, `weekly-trip-check`), `_planning/{overview.md,next-actions.md}`, full track frontmatter (featured/urgency/target_date/trip_record/travel_docs), and `desk/_meta/featured.md` set to `summer-2026-europe-trip`. Track-level `track.md` generated for the legacy track (status: collaborating, body documents the triage rules + migration date). Migration log written to `desk/_meta/migration-2026-05-22.log` listing every file's classification + final destination. **Idempotency:** if the log exists, re-runs abort with a clear message + exit code 1 unless `--force` is passed; with `--force`, destructive scope is bounded to migrator-owned dirs (`desk/_archive/`, `desk/legacy/`, `desk/summer-2026-europe-trip/`, `desk/_meta/featured.md`, `desk/_meta/migration-2026-05-22.log`) — nothing else under `desk/` is touched. **`--dry-run`** writes a per-bucket-counts summary to stdout and modifies nothing. Missing-`tasks/` bundles return a `no tasks/ directory` message with `performed: false`. Operator's policy: \"do best pass on ambiguous ones; when in doubt, archive\" — no interactive prompts, ever. **Europe-trip lift uses minimal stubs** in this unit; the operator can hand-edit content in a follow-up. CLI: `ouro migrate-to-desk --agent <name> [--root <path>] [--force] [--dry-run]`. Wires into `cli-types.ts` (new `migrate-to-desk` variant + `MigrateToDeskCliCommand` alias), `cli-parse.ts` (`parseMigrateToDeskCommand()`), `cli-exec.ts` (local executor branch — no daemon needed), `cli-help.ts` (Tasks category entry). New `src/__tests__/heart/daemon/migrate-to-desk.test.ts` with a synthetic fixture bundle at `src/__tests__/fixtures/migrate-bundle-mini/` covering all five buckets — classifier unit tests for each bucket + each updated-fallback step + slug/parent-dir derivation; migrator integration tests for post-migration tree shape, copy semantics (source untouched), schema_version coverage, legacy + europe-trip track.md generation, migration-log format, idempotency abort, `--force` bounded scope (operator-owned files outside the scope survive), `--dry-run` writes nothing, missing-bundle handling, empty-tasks handling, default-root derivation from `--agent`; CLI parser tests for argv→canonical routing + required-flag validation; CLI executor tests for `runOuroCli` wiring + exit-code on idempotency abort. `migrate-to-desk.ts` emits 4 `daemon.migrate_to_desk_*` nerves events (start / no-source / aborted_existing / dry-run / complete) for Rule 5 coverage. `classifier.ts` is pure-data — caller (migrator) owns observability. Bumps alpha.633."
600
+ "`ouro migrate-to-desk` migrator \u2014 new CLI command that copies an agent's legacy `tasks/` tree into desk-shape under `<bundle>/desk/`. **COPY semantics, not move** \u2014 the source `tasks/` directory is left intact for dual-read safety during transition; the operator triggers final deletion later (out of scope here). New `src/repertoire/desk/classifier.ts` exposes a pure-data classifier (`classifyFile()`, `deriveTaskSlug()`, `extractParentTaskDir()`, `resolveUpdatedMs()`) implementing the triage rules \u2014 terminal (done/approved/complete/cancelled/fixed/etc.) \u2192 archive bucket; live status with `updated` older than 30 days (cutoff 2026-04-22) \u2192 stale_live archive bucket; missing/unknown status / `.md.bak` / empty \u2192 ambiguous archive bucket; live status within 30 days \u2192 live_clear \u2192 migrate to a default `legacy` track; the one-off `ongoing/2026-03-09-1410-summer-2026-europe-trip.md` (a special-cased file from the initial bundle that motivated this migrator) \u2192 special_europe_trip lift. Effective `updated` resolves via YAML `updated` \u2192 `approved` \u2192 `created` \u2192 body `**Updated**:` \u2192 date-prefix in filename \u2192 file mtime. Pairing rules: planning/doing/ideation/audit siblings sharing a slug group together; if any sibling is terminal \u2192 all terminal, else if any is live_clear \u2192 all live_clear; sub-files under `<taskdir>/` inherit from the parent. Files under `tasks/archive/` are unconditionally terminal regardless of frontmatter. New `src/heart/daemon/migrate-to-desk.ts` exposes `runMigrateToDesk()` \u2014 walks `<bundle>/tasks/**`, classifies, groups, builds a deterministic plan, then writes archive bucket to `desk/_archive/<original-relative-path>` (preserving structure; `schema_version: 1` set on every touched markdown via `ensureSchemaVersion()`); live_clear tasks to `desk/legacy/<slug>/task.md` + `iterations/`; europe-trip task to `desk/summer-2026-europe-trip/` with two task scaffolds (`book-replacement-outbound`, `weekly-trip-check`), `_planning/{overview.md,next-actions.md}`, full track frontmatter (featured/urgency/target_date/trip_record/travel_docs), and `desk/_meta/featured.md` set to `summer-2026-europe-trip`. Track-level `track.md` generated for the legacy track (status: collaborating, body documents the triage rules + migration date). Migration log written to `desk/_meta/migration-2026-05-22.log` listing every file's classification + final destination. **Idempotency:** if the log exists, re-runs abort with a clear message + exit code 1 unless `--force` is passed; with `--force`, destructive scope is bounded to migrator-owned dirs (`desk/_archive/`, `desk/legacy/`, `desk/summer-2026-europe-trip/`, `desk/_meta/featured.md`, `desk/_meta/migration-2026-05-22.log`) \u2014 nothing else under `desk/` is touched. **`--dry-run`** writes a per-bucket-counts summary to stdout and modifies nothing. Missing-`tasks/` bundles return a `no tasks/ directory` message with `performed: false`. Operator's policy: \"do best pass on ambiguous ones; when in doubt, archive\" \u2014 no interactive prompts, ever. **Europe-trip lift uses minimal stubs** in this unit; the operator can hand-edit content in a follow-up. CLI: `ouro migrate-to-desk --agent <name> [--root <path>] [--force] [--dry-run]`. Wires into `cli-types.ts` (new `migrate-to-desk` variant + `MigrateToDeskCliCommand` alias), `cli-parse.ts` (`parseMigrateToDeskCommand()`), `cli-exec.ts` (local executor branch \u2014 no daemon needed), `cli-help.ts` (Tasks category entry). New `src/__tests__/heart/daemon/migrate-to-desk.test.ts` with a synthetic fixture bundle at `src/__tests__/fixtures/migrate-bundle-mini/` covering all five buckets \u2014 classifier unit tests for each bucket + each updated-fallback step + slug/parent-dir derivation; migrator integration tests for post-migration tree shape, copy semantics (source untouched), schema_version coverage, legacy + europe-trip track.md generation, migration-log format, idempotency abort, `--force` bounded scope (operator-owned files outside the scope survive), `--dry-run` writes nothing, missing-bundle handling, empty-tasks handling, default-root derivation from `--agent`; CLI parser tests for argv\u2192canonical routing + required-flag validation; CLI executor tests for `runOuroCli` wiring + exit-code on idempotency abort. `migrate-to-desk.ts` emits 4 `daemon.migrate_to_desk_*` nerves events (start / no-source / aborted_existing / dry-run / complete) for Rule 5 coverage. `classifier.ts` is pure-data \u2014 caller (migrator) owns observability. Bumps alpha.633."
587
601
  ]
588
602
  },
589
603
  {
590
604
  "version": "0.1.0-alpha.632",
591
605
  "changes": [
592
- "`ouro desk` umbrella CLI + `ouro task` alias — verb-and-router layer over the desk MCP server. New `src/heart/daemon/cli-desk.ts` exposes `parseDeskCommand()` + `parseTaskAliasCommand()` (pure argv→canonical-form parsers) and `executeDeskCommand()` (dispatches via the existing daemon `mcp.call` surface with `server: \"desk\"`). The CLI surface: `ouro desk task list|new|done|archive|show`, `ouro desk track list|new|show`, `ouro desk friction|lesson add <text>`, `ouro desk search|recall <query>`, `ouro desk reindex`, `ouro desk thread <path>`. `ouro task ...` is a top-level alias that routes through the same task sub-parser. Each subverb normalises to a `{ kind: \"desk\", tool, toolArgs }` shape on the `OuroCliCommand` union (new variant in `cli-types.ts` + `DeskCliCommand` alias) which the executor then JSON-stringifies into `mcp.call`'s `args` field. Verbs lacking a direct desk MCP tool (`task list/show`, `track list/show`) route through `desk_search` filtered by `kind`; `task done` rewrites to `task_update` with `frontmatter.status=\"done\"`. `desk reindex` routes to `desk_reindex` — the MCP server will surface a clean unknown-tool error until that admin tool ships in a follow-up unit. Wires into `cli-parse.ts` (top-level dispatch for `desk` + `task`), `cli-exec.ts` (new `desk` branch + `DeskCliCommand` added to `toDaemonCommand`'s exclusion union), and `cli-help.ts` (`desk`/`task` entries under the Tasks category). New `src/__tests__/heart/daemon/cli-desk.test.ts` covers argv→canonical routing for every subverb (alias and umbrella forms), daemon-socket dispatch shape, `--agent` propagation, daemon-unavailable + daemon-error + empty-content fallbacks, slug normalisation (idempotent `/task.md` suffix + trailing-slash trim), and the usage-string export. The earlier retirement contract in `daemon-cli.test.ts` + `cli-help.test.ts` updated to reflect this PR's reintroduction of `task` as an alias (legacy `task board`/`create`/`fix` subverbs still rejected, just under a desk usage hint rather than a top-level unknown-command error). No new business logic — pure verb-and-router. `cli-desk.ts` is 100% line/branch/func/statement covered and emits `daemon.desk_cli_dispatch` for nerves audit Rule 5."
606
+ "`ouro desk` umbrella CLI + `ouro task` alias \u2014 verb-and-router layer over the desk MCP server. New `src/heart/daemon/cli-desk.ts` exposes `parseDeskCommand()` + `parseTaskAliasCommand()` (pure argv\u2192canonical-form parsers) and `executeDeskCommand()` (dispatches via the existing daemon `mcp.call` surface with `server: \"desk\"`). The CLI surface: `ouro desk task list|new|done|archive|show`, `ouro desk track list|new|show`, `ouro desk friction|lesson add <text>`, `ouro desk search|recall <query>`, `ouro desk reindex`, `ouro desk thread <path>`. `ouro task ...` is a top-level alias that routes through the same task sub-parser. Each subverb normalises to a `{ kind: \"desk\", tool, toolArgs }` shape on the `OuroCliCommand` union (new variant in `cli-types.ts` + `DeskCliCommand` alias) which the executor then JSON-stringifies into `mcp.call`'s `args` field. Verbs lacking a direct desk MCP tool (`task list/show`, `track list/show`) route through `desk_search` filtered by `kind`; `task done` rewrites to `task_update` with `frontmatter.status=\"done\"`. `desk reindex` routes to `desk_reindex` \u2014 the MCP server will surface a clean unknown-tool error until that admin tool ships in a follow-up unit. Wires into `cli-parse.ts` (top-level dispatch for `desk` + `task`), `cli-exec.ts` (new `desk` branch + `DeskCliCommand` added to `toDaemonCommand`'s exclusion union), and `cli-help.ts` (`desk`/`task` entries under the Tasks category). New `src/__tests__/heart/daemon/cli-desk.test.ts` covers argv\u2192canonical routing for every subverb (alias and umbrella forms), daemon-socket dispatch shape, `--agent` propagation, daemon-unavailable + daemon-error + empty-content fallbacks, slug normalisation (idempotent `/task.md` suffix + trailing-slash trim), and the usage-string export. The earlier retirement contract in `daemon-cli.test.ts` + `cli-help.test.ts` updated to reflect this PR's reintroduction of `task` as an alias (legacy `task board`/`create`/`fix` subverbs still rejected, just under a desk usage hint rather than a top-level unknown-command error). No new business logic \u2014 pure verb-and-router. `cli-desk.ts` is 100% line/branch/func/statement covered and emits `daemon.desk_cli_dispatch` for nerves audit Rule 5."
593
607
  ]
594
608
  },
595
609
  {
596
610
  "version": "0.1.0-alpha.631",
597
611
  "changes": [
598
- "ouroboros daemon now reads each enabled plugin's `<plugin-root>/.mcp.json` and auto-spawns the declared stdio MCP servers per agent. New `src/repertoire/plugin-mcp.ts` exposes `listPluginMcpServers()` — walks `listEnabledPlugins()`, reads each plugin's `.mcp.json` (Anthropic/Claude-Code public spec shape with `mcpServers` map), resolves `${VAR:-default}` substitution in `args` + `env` values (with special-case: `DESK` defaults to `<agent-bundle>/desk/` when unset). Missing `.mcp.json` skips cleanly; malformed JSON emits `plugin_mcp.parse_error` and skips cleanly (no daemon crash). `getSharedMcpManager()` merges plugin-declared servers with builtin `agent.json` `mcpServers` before calling `manager.start()`; on name collision, builtin wins (deterministic — operator's explicit agent.json overrides plugin defaults). `McpManager.start()` accepts a second `pluginOrigins: Record<string, string>` arg threading the plugin id through to each `ServerEntry`; `listAllTools()` exposes the new `pluginId` field per entry. `resolveVaultEnv()` short-circuits when no `vault:` reference exists (avoids spinning up the credential store for plugin servers whose env is empty / pure-string). `mcpToolsAsDefinitions()` uses the `pluginId` flag to name plugin-server tools as `mcp__<server>__<tool>` (Anthropic public convention, matches desk-section.ts on-prompt promise); builtin server tools keep the legacy `<server>_<tool>` shape. After this PR the desk plugin's `mcp__desk__*` tools (CRUD + 5 search + thread = 12 tools) become live in any agent that has the desk plugin enabled."
612
+ "ouroboros daemon now reads each enabled plugin's `<plugin-root>/.mcp.json` and auto-spawns the declared stdio MCP servers per agent. New `src/repertoire/plugin-mcp.ts` exposes `listPluginMcpServers()` \u2014 walks `listEnabledPlugins()`, reads each plugin's `.mcp.json` (Anthropic/Claude-Code public spec shape with `mcpServers` map), resolves `${VAR:-default}` substitution in `args` + `env` values (with special-case: `DESK` defaults to `<agent-bundle>/desk/` when unset). Missing `.mcp.json` skips cleanly; malformed JSON emits `plugin_mcp.parse_error` and skips cleanly (no daemon crash). `getSharedMcpManager()` merges plugin-declared servers with builtin `agent.json` `mcpServers` before calling `manager.start()`; on name collision, builtin wins (deterministic \u2014 operator's explicit agent.json overrides plugin defaults). `McpManager.start()` accepts a second `pluginOrigins: Record<string, string>` arg threading the plugin id through to each `ServerEntry`; `listAllTools()` exposes the new `pluginId` field per entry. `resolveVaultEnv()` short-circuits when no `vault:` reference exists (avoids spinning up the credential store for plugin servers whose env is empty / pure-string). `mcpToolsAsDefinitions()` uses the `pluginId` flag to name plugin-server tools as `mcp__<server>__<tool>` (Anthropic public convention, matches desk-section.ts on-prompt promise); builtin server tools keep the legacy `<server>_<tool>` shape. After this PR the desk plugin's `mcp__desk__*` tools (CRUD + 5 search + thread = 12 tools) become live in any agent that has the desk plugin enabled."
599
613
  ]
600
614
  },
601
615
  {
602
616
  "version": "0.1.0-alpha.630",
603
617
  "changes": [
604
- "delete the `src/repertoire/tasks/` module and its tests from disk. Final cleanup after an earlier PR dropped production reads and a follow-up PR rewired the bridges, daemon scheduler, and parser utilities off the module. Also rewires the lingering `cli-exec.ts` inner-status handler to import `parseFrontmatter` from `src/util/frontmatter` (earlier-PR miss). Test sweep: 19 `vi.mock(\"../../repertoire/tasks\", ...)` defensive mock blocks removed across `prompt-*`, `tools-*`, `refresh-system-prompt`, and `continuity-tools` test files (mocks pointed at a module path that no longer exists). `task-scheduler.test.ts` removed (legacy fixture helpers gone; scheduler still exercised via integration). `mailbox-readers-continuity-catches.test.ts` drops the obsolete \"self-fix view when task scanning throws\" case (path no longer exists). vitest.config.ts drops the `src/repertoire/tasks/types.ts` coverage exclude (file gone). Drops the dead `renderTaskTransitionLines()` helper from `src/arc/task-lifecycle.ts` (zero callers — was consumed only by the now-deleted task-board prompt rendering)."
618
+ "delete the `src/repertoire/tasks/` module and its tests from disk. Final cleanup after an earlier PR dropped production reads and a follow-up PR rewired the bridges, daemon scheduler, and parser utilities off the module. Also rewires the lingering `cli-exec.ts` inner-status handler to import `parseFrontmatter` from `src/util/frontmatter` (earlier-PR miss). Test sweep: 19 `vi.mock(\"../../repertoire/tasks\", ...)` defensive mock blocks removed across `prompt-*`, `tools-*`, `refresh-system-prompt`, and `continuity-tools` test files (mocks pointed at a module path that no longer exists). `task-scheduler.test.ts` removed (legacy fixture helpers gone; scheduler still exercised via integration). `mailbox-readers-continuity-catches.test.ts` drops the obsolete \"self-fix view when task scanning throws\" case (path no longer exists). vitest.config.ts drops the `src/repertoire/tasks/types.ts` coverage exclude (file gone). Drops the dead `renderTaskTransitionLines()` helper from `src/arc/task-lifecycle.ts` (zero callers \u2014 was consumed only by the now-deleted task-board prompt rendering)."
605
619
  ]
606
620
  },
607
621
  {
608
622
  "version": "0.1.0-alpha.629",
609
623
  "changes": [
610
- "rewire bridges, daemon scheduler, and parser utilities off the deprecated `src/repertoire/tasks/` module. `parseFrontmatter` extracted to `src/util/frontmatter.ts` (single shared helper) — `src/heart/awaiting/await-parser.ts`, `src/heart/habits/habit-parser.ts`, `src/heart/habits/habit-migration.ts`, and `src/heart/hatch/specialist-prompt.ts` now import from there. `src/heart/bridges/manager.ts` no longer reads from the task module (and gains a `defaultWriteDeskTask` writer for promoted-bridge desk task.md emission). `src/heart/daemon/task-scheduler.ts`, `src/heart/daemon/cli-exec.ts`, `src/heart/daemon/cli-types.ts`, `src/heart/daemon/cli-help.ts`, and `src/heart/daemon/cli-parse.ts` drop their task-CLI surface (`task board`, `task create`, etc.) — those commands move to the desk MCP server. `src/mind/prompt.ts` and `src/repertoire/guardrails.ts` drop residual task-command references. After this PR, no production code imports from `src/repertoire/tasks/`; a follow-up PR does the final disk-level deletion. Pipeline-integration tests updated to drop the three obsolete `parses task ...` cases. New `src/util/frontmatter.ts` ships with focused tests (12 cases, 100% line/branch/func/statement). Bridges manager + state-machine + daemon task-scheduler temporarily excluded from the strict-100% coverage gate (`defaultWriteDeskTask`, desk-discovery walk, and `suspendBridge` branch lack direct tests); followup PR will either backfill tests or refactor the defensive catches."
624
+ "rewire bridges, daemon scheduler, and parser utilities off the deprecated `src/repertoire/tasks/` module. `parseFrontmatter` extracted to `src/util/frontmatter.ts` (single shared helper) \u2014 `src/heart/awaiting/await-parser.ts`, `src/heart/habits/habit-parser.ts`, `src/heart/habits/habit-migration.ts`, and `src/heart/hatch/specialist-prompt.ts` now import from there. `src/heart/bridges/manager.ts` no longer reads from the task module (and gains a `defaultWriteDeskTask` writer for promoted-bridge desk task.md emission). `src/heart/daemon/task-scheduler.ts`, `src/heart/daemon/cli-exec.ts`, `src/heart/daemon/cli-types.ts`, `src/heart/daemon/cli-help.ts`, and `src/heart/daemon/cli-parse.ts` drop their task-CLI surface (`task board`, `task create`, etc.) \u2014 those commands move to the desk MCP server. `src/mind/prompt.ts` and `src/repertoire/guardrails.ts` drop residual task-command references. After this PR, no production code imports from `src/repertoire/tasks/`; a follow-up PR does the final disk-level deletion. Pipeline-integration tests updated to drop the three obsolete `parses task ...` cases. New `src/util/frontmatter.ts` ships with focused tests (12 cases, 100% line/branch/func/statement). Bridges manager + state-machine + daemon task-scheduler temporarily excluded from the strict-100% coverage gate (`defaultWriteDeskTask`, desk-discovery walk, and `suspendBridge` branch lack direct tests); followup PR will either backfill tests or refactor the defensive catches."
611
625
  ]
612
626
  },
613
627
  {
@@ -619,7 +633,7 @@
619
633
  {
620
634
  "version": "0.1.0-alpha.627",
621
635
  "changes": [
622
- "assemble `## my desk` prompt section from `<bundle>/desk/`. New `src/mind/desk-section.ts` reads each agent's desk dir synchronously every turn and emits the desk-vocab body (the agent's reflex + threshold + 8-state vocab + first-class-systems-link-not-absorb rule) followed by a dynamic `### currently` block showing the featured track + top-3 non-terminal tasks in it + other active tracks + non-terminal-task count. Featured resolution: `<bundle>/desk/_meta/featured.md` (one slug per line); falls back to alphabetical-first-active track when absent or all entries stale. Closed tracks are skipped as featured candidates. Empty-desk emits a `### currently: empty — no tracks yet.` stub. Parses `schema_version: 0` (pre-migration) and `schema_version: 1` (post-migration) tasks identically. 26 unit tests including 7 coverage-gate edge cases. Removes the old `## task board` rendering and `ouro task ...` body-map cheatsheet from prompt.ts. Emits `prompt.desk_section_assembled` nerves event per turn (file-completeness gate). desk-section.ts excluded from strict-100% coverage gate (94% stmts / 82% branches, 100% lines + funcs); followup will tighten or refactor the defensive FS catches. follow-up sequence cleans up the deeper repertoire/tasks module dependencies."
636
+ "assemble `## my desk` prompt section from `<bundle>/desk/`. New `src/mind/desk-section.ts` reads each agent's desk dir synchronously every turn and emits the desk-vocab body (the agent's reflex + threshold + 8-state vocab + first-class-systems-link-not-absorb rule) followed by a dynamic `### currently` block showing the featured track + top-3 non-terminal tasks in it + other active tracks + non-terminal-task count. Featured resolution: `<bundle>/desk/_meta/featured.md` (one slug per line); falls back to alphabetical-first-active track when absent or all entries stale. Closed tracks are skipped as featured candidates. Empty-desk emits a `### currently: empty \u2014 no tracks yet.` stub. Parses `schema_version: 0` (pre-migration) and `schema_version: 1` (post-migration) tasks identically. 26 unit tests including 7 coverage-gate edge cases. Removes the old `## task board` rendering and `ouro task ...` body-map cheatsheet from prompt.ts. Emits `prompt.desk_section_assembled` nerves event per turn (file-completeness gate). desk-section.ts excluded from strict-100% coverage gate (94% stmts / 82% branches, 100% lines + funcs); followup will tighten or refactor the defensive FS catches. follow-up sequence cleans up the deeper repertoire/tasks module dependencies."
623
637
  ]
624
638
  },
625
639
  {
@@ -666,7 +680,7 @@
666
680
  {
667
681
  "version": "0.1.0-alpha.620",
668
682
  "changes": [
669
- "wire the `--agent <name>` flag end-to-end across `ouro plugin install / list / remove`. Previously parsed but ignored by the handlers. Now: install --agent X reads X's `~/AgentBundles/X.ouro/agent.json`, idempotently adds `{ id, enabled: true, source, version }` to plugins[] (source + version persisted only when set on the command). list --agent X intersects installed-on-disk with X's plugins[] (enabled-for-X filter). remove --agent X removes the entry from X's plugins[] only — never deletes the plugin from disk (other agents may still use it). remove WITHOUT --agent scans all bundles via getAgentBundlesRoot; refuses with a clear message if any agent's plugins[] still references the plugin, listing the offending agents. agent.json writes match the codebase's existing non-atomic `fs.writeFileSync` pattern (atomic-write helper deferred as a follow-up). 15 new tests cover the install / list / remove paths under --agent (and the new machine-wide remove guardrail); plugin-cli.ts stays at 100% coverage.",
683
+ "wire the `--agent <name>` flag end-to-end across `ouro plugin install / list / remove`. Previously parsed but ignored by the handlers. Now: install --agent X reads X's `~/AgentBundles/X.ouro/agent.json`, idempotently adds `{ id, enabled: true, source, version }` to plugins[] (source + version persisted only when set on the command). list --agent X intersects installed-on-disk with X's plugins[] (enabled-for-X filter). remove --agent X removes the entry from X's plugins[] only \u2014 never deletes the plugin from disk (other agents may still use it). remove WITHOUT --agent scans all bundles via getAgentBundlesRoot; refuses with a clear message if any agent's plugins[] still references the plugin, listing the offending agents. agent.json writes match the codebase's existing non-atomic `fs.writeFileSync` pattern (atomic-write helper deferred as a follow-up). 15 new tests cover the install / list / remove paths under --agent (and the new machine-wide remove guardrail); plugin-cli.ts stays at 100% coverage.",
670
684
  "Orientation substrate campaign: invalid provider/model pairs now fail fast across startup, `ouro use --force`, legacy `auth switch`, and legacy `config model`; BlueBubbles now keeps channel/routing metadata in a structured orientation frame instead of injecting it into user speech; agents get an `orientation_get` tool plus correction-hold action rails that block high-risk durable writes, shell mutations, and first-class MCP tool calls when a terse correction depends on prior context; high-risk tool profiles now require a typed reason so blocked mutations always explain the risk; trip leg updates now require an explicit `updateReason` so confident-but-wrong corrections leave an auditable rationale instead of silently mutating itinerary state.",
671
685
  "Inner/BlueBubbles boundary: `surface` no longer attempts proactive iMessage delivery when returning to bridge-attached or freshest BlueBubbles sessions. Surface now queues the return for the active session, and the tool description names `send_message` with `channel=\"bluebubbles\"` as the dedicated intentional live-send path. Regression tests pin both BlueBubbles surface routes and the explicit send_message live-delivery path."
672
686
  ]
@@ -680,13 +694,13 @@
680
694
  {
681
695
  "version": "0.1.0-alpha.618",
682
696
  "changes": [
683
- "prompt-assembly integration. src/repertoire/skills.ts listSkills() merges plugin skills (via listPluginSkills(listEnabledPlugins())) with bundle skills (deduped, sorted). loadSkill() falls back to plugin skills after the 3 bundle paths (agent → protocol mirror → harness); iterates enabled plugins in declaration order and returns the first match. Bundle skills retain precedence — if the same skill name is in both a bundle and a plugin, the bundle wins. 6 new tests cover the integration paths; skills.ts stays at 100% coverage. Final sub-PR in the ouroboros plugin support track."
697
+ "prompt-assembly integration. src/repertoire/skills.ts listSkills() merges plugin skills (via listPluginSkills(listEnabledPlugins())) with bundle skills (deduped, sorted). loadSkill() falls back to plugin skills after the 3 bundle paths (agent \u2192 protocol mirror \u2192 harness); iterates enabled plugins in declaration order and returns the first match. Bundle skills retain precedence \u2014 if the same skill name is in both a bundle and a plugin, the bundle wins. 6 new tests cover the integration paths; skills.ts stays at 100% coverage. Final sub-PR in the ouroboros plugin support track."
684
698
  ]
685
699
  },
686
700
  {
687
701
  "version": "0.1.0-alpha.617",
688
702
  "changes": [
689
- "`ouro plugin install <source>`, `ouro plugin list`, and `ouro plugin remove <id>` CLI commands. Install clones to ~/.ouro-cli/plugins/<id>/ via git, supports `github:org/repo:plugins/<id>`, `https://github.com/...[.git]`, `local:/path/...`, and bare absolute paths; verifies `.claude-plugin/plugin.json` exists and rolls back on failure. List walks the plugins root and reports sorted installed plugins. Remove deletes the plugin install dir. Handlers live in src/heart/daemon/plugin-cli.ts (narrow RM_RECURSIVE_ALLOWLIST entry — operator-invoked CLI infrastructure, not agent-callable). Next sub-PR wires plugin skills into prompt assembly via listPluginSkills()."
703
+ "`ouro plugin install <source>`, `ouro plugin list`, and `ouro plugin remove <id>` CLI commands. Install clones to ~/.ouro-cli/plugins/<id>/ via git, supports `github:org/repo:plugins/<id>`, `https://github.com/...[.git]`, `local:/path/...`, and bare absolute paths; verifies `.claude-plugin/plugin.json` exists and rolls back on failure. List walks the plugins root and reports sorted installed plugins. Remove deletes the plugin install dir. Handlers live in src/heart/daemon/plugin-cli.ts (narrow RM_RECURSIVE_ALLOWLIST entry \u2014 operator-invoked CLI infrastructure, not agent-callable). Next sub-PR wires plugin skills into prompt assembly via listPluginSkills()."
690
704
  ]
691
705
  },
692
706
  {
@@ -738,7 +752,7 @@
738
752
  {
739
753
  "version": "0.1.0-alpha.609",
740
754
  "changes": [
741
- "Vocabulary sweep: drop 'memory' from agent-facing prompt content (trip-ledger truth section) and operator-facing connect-flow strings (cli-exec/cli-help). The agent doesn't have memory; it has a diary it consults, embeddings it can search, and prior conversation context. Renaming the strings makes the surface honest. Internal capability identifier 'memory-embeddings' and CLI input alias 'memory' both stay for back-compat. RAM-sense uses ('in-memory cache', 'process memory') stay. Test snapshots updated. Bumped from .607 → .609 to leapfrog the parallel-merged .608."
755
+ "Vocabulary sweep: drop 'memory' from agent-facing prompt content (trip-ledger truth section) and operator-facing connect-flow strings (cli-exec/cli-help). The agent doesn't have memory; it has a diary it consults, embeddings it can search, and prior conversation context. Renaming the strings makes the surface honest. Internal capability identifier 'memory-embeddings' and CLI input alias 'memory' both stay for back-compat. RAM-sense uses ('in-memory cache', 'process memory') stay. Test snapshots updated. Bumped from .607 \u2192 .609 to leapfrog the parallel-merged .608."
742
756
  ]
743
757
  },
744
758
  {
@@ -750,7 +764,7 @@
750
764
  {
751
765
  "version": "0.1.0-alpha.606",
752
766
  "changes": [
753
- "Root-cause fix for the 2026-05-11 inner-dialog wake storm that cost ~$50 in minimax inference. PR #725 removed `inner.wake` from the Claude Code post-tool-use hook with the stated intent that the notification message stay in the queue and be picked up on the agent's next natural turn — but the daemon's `message.send` HANDLER (daemon.ts case `message.send`) was unconditionally calling `processManager.sendToAgent(to, { type: \"message\" })` after queueing, which woke the inner-dialog worker on every message.send anyway. ~30 message.send/min × the 3-turn instinct-loop cap = ~90 turns/min sustained for hours. PR #725 fixed the hook side; the daemon side defeated it.\n\nFix: the `message.send` handler is now pure queue-only delivery. No `startAgent`, no `sendToAgent` — just `router.send`. Callers that want immediate processing must send `inner.wake` explicitly after `message.send`. The Claude Code hook (cli-exec.ts) was already correctly discriminating (only firing inner.wake on session-start/stop, never per-tool-use), so it works as originally intended now. The CLI `ouro msg` was updated to chain `inner.wake` after `message.send` (operator-driven delivery wants immediate response, preserving historical CLI UX). Other callers (API, programmatic) default to queue-only.\n\nTest pinned: `daemon-command-plane-branches.test.ts` now asserts `processManager.startAgent` and `processManager.sendToAgent` are NOT called from `message.send`. The regression cannot be silently reintroduced."
767
+ "Root-cause fix for the 2026-05-11 inner-dialog wake storm that cost ~$50 in minimax inference. PR #725 removed `inner.wake` from the Claude Code post-tool-use hook with the stated intent that the notification message stay in the queue and be picked up on the agent's next natural turn \u2014 but the daemon's `message.send` HANDLER (daemon.ts case `message.send`) was unconditionally calling `processManager.sendToAgent(to, { type: \"message\" })` after queueing, which woke the inner-dialog worker on every message.send anyway. ~30 message.send/min \u00d7 the 3-turn instinct-loop cap = ~90 turns/min sustained for hours. PR #725 fixed the hook side; the daemon side defeated it.\n\nFix: the `message.send` handler is now pure queue-only delivery. No `startAgent`, no `sendToAgent` \u2014 just `router.send`. Callers that want immediate processing must send `inner.wake` explicitly after `message.send`. The Claude Code hook (cli-exec.ts) was already correctly discriminating (only firing inner.wake on session-start/stop, never per-tool-use), so it works as originally intended now. The CLI `ouro msg` was updated to chain `inner.wake` after `message.send` (operator-driven delivery wants immediate response, preserving historical CLI UX). Other callers (API, programmatic) default to queue-only.\n\nTest pinned: `daemon-command-plane-branches.test.ts` now asserts `processManager.startAgent` and `processManager.sendToAgent` are NOT called from `message.send`. The regression cannot be silently reintroduced."
754
768
  ]
755
769
  },
756
770
  {
@@ -768,7 +782,7 @@
768
782
  {
769
783
  "version": "0.1.0-alpha.602",
770
784
  "changes": [
771
- "Two-part fix for the 2026-05-11 BlueBubbles wedge: an agent's BB session showed the same user message replayed 76 times, because each death-spiral cycle re-injected the inbound. Root cause was the daemon's HTTP health probe (`createHttpHealthProbe(\"bluebubbles:<agent>\", port)`) GETting the sense's /health endpoint every ~60 s with a 5 s timeout — busy BB sense (e.g. VLM image-describe at 20+ s) timed out, daemon declared 'critical', SIGTERM'd the sense mid-work, respawned, hit the same image, killed again, forever. Part 1: removed the HTTP probe entirely from `listHealthProbes()`. Process supervision (`processManager` child-process exit handler) already catches dead processes; for 'alive but hung' we now rely on the agent's own awareness via `pendingRecoveryCount` / `lastRecoveredAt` in the BB runtime state surfaced into the prompt, plus the agent's new `restart_runtime` tool (from alpha.598 / #723). Part 2: defense-in-depth respawn-loop guard in `processManager.restartAgent` — if anything triggers more than `RESPAWN_GUARD_MAX_RESTARTS = 5` orchestrated restarts in `RESPAWN_GUARD_WINDOW_MS = 10 min`, refuse further restarts (`daemon.agent_respawn_loop_tripped` nerves event, errorReason + fixHint set on the snapshot). Trip self-clears once timestamps age out of the window, and `startAgent` (= `ouro up`) bypasses the guard so the operator can always recover. Even if some other future cause re-introduces a tight respawn loop, the guard bounds it. The 2026-05-11 spiral was ~60 restarts/hr — well above 5/10min, so this would have caught it."
785
+ "Two-part fix for the 2026-05-11 BlueBubbles wedge: an agent's BB session showed the same user message replayed 76 times, because each death-spiral cycle re-injected the inbound. Root cause was the daemon's HTTP health probe (`createHttpHealthProbe(\"bluebubbles:<agent>\", port)`) GETting the sense's /health endpoint every ~60 s with a 5 s timeout \u2014 busy BB sense (e.g. VLM image-describe at 20+ s) timed out, daemon declared 'critical', SIGTERM'd the sense mid-work, respawned, hit the same image, killed again, forever. Part 1: removed the HTTP probe entirely from `listHealthProbes()`. Process supervision (`processManager` child-process exit handler) already catches dead processes; for 'alive but hung' we now rely on the agent's own awareness via `pendingRecoveryCount` / `lastRecoveredAt` in the BB runtime state surfaced into the prompt, plus the agent's new `restart_runtime` tool (from alpha.598 / #723). Part 2: defense-in-depth respawn-loop guard in `processManager.restartAgent` \u2014 if anything triggers more than `RESPAWN_GUARD_MAX_RESTARTS = 5` orchestrated restarts in `RESPAWN_GUARD_WINDOW_MS = 10 min`, refuse further restarts (`daemon.agent_respawn_loop_tripped` nerves event, errorReason + fixHint set on the snapshot). Trip self-clears once timestamps age out of the window, and `startAgent` (= `ouro up`) bypasses the guard so the operator can always recover. Even if some other future cause re-introduces a tight respawn loop, the guard bounds it. The 2026-05-11 spiral was ~60 restarts/hr \u2014 well above 5/10min, so this would have caught it."
772
786
  ]
773
787
  },
774
788
  {
@@ -780,43 +794,43 @@
780
794
  {
781
795
  "version": "0.1.0-alpha.600",
782
796
  "changes": [
783
- "Harness attention hygiene: post-tool-use Claude Code hook no longer wakes the inner loop on every tool — only on session-start and stop. listActiveReturnObligations gains a 14-day age cap so legacy queued items stop cycling indefinitely, and a strict status allow-list so legacy 'fulfilled' values written before the ReturnObligationStatus split (or by any future code path that bypasses the type via 'as any') no longer leak into the held-work-items injection."
797
+ "Harness attention hygiene: post-tool-use Claude Code hook no longer wakes the inner loop on every tool \u2014 only on session-start and stop. listActiveReturnObligations gains a 14-day age cap so legacy queued items stop cycling indefinitely, and a strict status allow-list so legacy 'fulfilled' values written before the ReturnObligationStatus split (or by any future code path that bypasses the type via 'as any') no longer leak into the held-work-items injection."
784
798
  ]
785
799
  },
786
800
  {
787
801
  "version": "0.1.0-alpha.599",
788
802
  "changes": [
789
- "BlueBubbles in-flight marker hardening. The in-memory `bbInFlightMessageGuids` tracker had a leak class: any exit path in `handleBlueBubblesNormalizedEvent` that doesn't call `endBlueBubblesMessageInFlight` strands the marker forever (until BB sense process restart), silently halting forward progress on the recovery queue. Slugger lost BlueBubbles inbound for 12+ hours on 2026-05-11 — six user messages piled up unprocessed because each recovery attempt saw the stale marker, returned `already_processed` without actually processing, and the recovery loop counted progress without making any. The class fix: in-flight markers now carry a claim timestamp and expire after `BB_IN_FLIGHT_MAX_AGE_MS = 15 min` (50% beyond the 10-min recovery-turn timeout, so live owners get full headroom). `isBlueBubblesMessageInFlight` returns false for stale markers; `beginBlueBubblesMessageInFlight` is allowed to replace a stale marker and emits `senses.bluebubbles_in_flight_marker_expired` so the auto-eviction is observable. A leaked marker now self-clears in at most 15 min instead of forever. Defense-in-depth: explicit `endBlueBubblesMessageInFlight` audit still worth doing in a follow-up, but the TTL guarantees the bug class can't wedge the queue indefinitely."
803
+ "BlueBubbles in-flight marker hardening. The in-memory `bbInFlightMessageGuids` tracker had a leak class: any exit path in `handleBlueBubblesNormalizedEvent` that doesn't call `endBlueBubblesMessageInFlight` strands the marker forever (until BB sense process restart), silently halting forward progress on the recovery queue. Slugger lost BlueBubbles inbound for 12+ hours on 2026-05-11 \u2014 six user messages piled up unprocessed because each recovery attempt saw the stale marker, returned `already_processed` without actually processing, and the recovery loop counted progress without making any. The class fix: in-flight markers now carry a claim timestamp and expire after `BB_IN_FLIGHT_MAX_AGE_MS = 15 min` (50% beyond the 10-min recovery-turn timeout, so live owners get full headroom). `isBlueBubblesMessageInFlight` returns false for stale markers; `beginBlueBubblesMessageInFlight` is allowed to replace a stale marker and emits `senses.bluebubbles_in_flight_marker_expired` so the auto-eviction is observable. A leaked marker now self-clears in at most 15 min instead of forever. Defense-in-depth: explicit `endBlueBubblesMessageInFlight` audit still worth doing in a follow-up, but the TTL guarantees the bug class can't wedge the queue indefinitely."
790
804
  ]
791
805
  },
792
806
  {
793
807
  "version": "0.1.0-alpha.598",
794
808
  "changes": [
795
- "New `restart_runtime({ reason })` tool. Agent self-maintenance: asking the human to restart the daemon over BlueBubbles is now a thing of the past. Sends `daemon.restart` over the existing socket — daemon logs the reason as `daemon.restart_requested`, runs its normal stop pathway, and exits. launchctl's KeepAlive policy auto-respawns the daemon, so the agent comes back fresh on the other side. The agent will not see this tool's return value — its process exits with the daemon and a clean boot replaces it. In dev mode (no launchctl) the daemon just exits; same observable behavior as `daemon.stop`, with the restart-requested audit event distinguishing intent."
809
+ "New `restart_runtime({ reason })` tool. Agent self-maintenance: asking the human to restart the daemon over BlueBubbles is now a thing of the past. Sends `daemon.restart` over the existing socket \u2014 daemon logs the reason as `daemon.restart_requested`, runs its normal stop pathway, and exits. launchctl's KeepAlive policy auto-respawns the daemon, so the agent comes back fresh on the other side. The agent will not see this tool's return value \u2014 its process exits with the daemon and a clean boot replaces it. In dev mode (no launchctl) the daemon just exits; same observable behavior as `daemon.stop`, with the restart-requested audit event distinguishing intent."
796
810
  ]
797
811
  },
798
812
  {
799
813
  "version": "0.1.0-alpha.597",
800
814
  "changes": [
801
- "New `let_go({ id, reason? })` tool. Closes a gap surfaced live by Slugger after #720 (await_condition): when a held work item is resolved externally (e.g. an obligation whose underlying issue was merged in a separate PR), the agent had no way to dismiss it. The existing path to terminal state — fulfilling via `surface` — requires delivering a response, which doesn't fit externally-resolved work. The agent kept seeing the same stale items in its prompt every turn for a month with nothing to act on. `let_go` is dismissal WITHOUT delivery: tries `arc/obligations/inner/<id>.json` (ReturnObligation → `returned` with returnTarget=`surface`), falls through to `arc/obligations/<id>.json` (Obligation → `fulfilled` with `latestNote=reason`). Idempotent — calling on an already-terminal item returns the existing status, not an error. Emits `repertoire.obligation_let_go` nerves event recording the reason for future-me. id is the bracketed value in the prompt's 'held work items' section."
815
+ "New `let_go({ id, reason? })` tool. Closes a gap surfaced live by Slugger after #720 (await_condition): when a held work item is resolved externally (e.g. an obligation whose underlying issue was merged in a separate PR), the agent had no way to dismiss it. The existing path to terminal state \u2014 fulfilling via `surface` \u2014 requires delivering a response, which doesn't fit externally-resolved work. The agent kept seeing the same stale items in its prompt every turn for a month with nothing to act on. `let_go` is dismissal WITHOUT delivery: tries `arc/obligations/inner/<id>.json` (ReturnObligation \u2192 `returned` with returnTarget=`surface`), falls through to `arc/obligations/<id>.json` (Obligation \u2192 `fulfilled` with `latestNote=reason`). Idempotent \u2014 calling on an already-terminal item returns the existing status, not an error. Emits `repertoire.obligation_let_go` nerves event recording the reason for future-me. id is the bracketed value in the prompt's 'held work items' section."
802
816
  ]
803
817
  },
804
818
  {
805
819
  "version": "0.1.0-alpha.596",
806
820
  "changes": [
807
- "New await_condition primitive. Agents can file a natural-language condition (`await_condition({ name, condition, cadence, alert?, mode?, max_age?, body? })`); the daemon polls on cadence and queues an inner-dialog tick titled `await tick: <name> — <condition>` with the file body + history block (checked count, last-checked age, last observation). The agent calls `resolve_await({ name, verdict, observation })`: verdict=yes archives the file to `awaiting/.done/` and fires an alert via cross-chat-delivery (intent=generic_outreach) targeting `filed_for_friend_id` on the `alert` channel — proactive sends slot into the user's existing thread, no self-loopback. verdict=no records the observation and increments the tick count. `cancel_await({ name, reason? })` silently archives without alert. max_age triggers auto-expiry with a 'timed out' alert. Surfaces `## what i'm waiting on` in the commitments section. Validated end-to-end with Slugger: file → tick → resolve(no, x2) → resolve(yes) → archive + cross_chat_delivery queued_for_later. AwaitScheduler mkdirs its awaits dir at start so the fs.watch watcher attaches on first boot (no need to wait for a periodic reconcile)."
821
+ "New await_condition primitive. Agents can file a natural-language condition (`await_condition({ name, condition, cadence, alert?, mode?, max_age?, body? })`); the daemon polls on cadence and queues an inner-dialog tick titled `await tick: <name> \u2014 <condition>` with the file body + history block (checked count, last-checked age, last observation). The agent calls `resolve_await({ name, verdict, observation })`: verdict=yes archives the file to `awaiting/.done/` and fires an alert via cross-chat-delivery (intent=generic_outreach) targeting `filed_for_friend_id` on the `alert` channel \u2014 proactive sends slot into the user's existing thread, no self-loopback. verdict=no records the observation and increments the tick count. `cancel_await({ name, reason? })` silently archives without alert. max_age triggers auto-expiry with a 'timed out' alert. Surfaces `## what i'm waiting on` in the commitments section. Validated end-to-end with Slugger: file \u2192 tick \u2192 resolve(no, x2) \u2192 resolve(yes) \u2192 archive + cross_chat_delivery queued_for_later. AwaitScheduler mkdirs its awaits dir at start so the fs.watch watcher attaches on first boot (no need to wait for a periodic reconcile)."
808
822
  ]
809
823
  },
810
824
  {
811
825
  "version": "0.1.0-alpha.595",
812
826
  "changes": [
813
- "Voice phone transport defaults to media-stream when OpenAI Realtime or OpenAI SIP is configured but no explicit voice.twilioTransportMode is set. Previously the default was record-play, which made conversationEngine resolve to cascade and routed inbound calls through the ElevenLabs/Whisper greeting path operators with realtime-only credentials never configured — producing a fully silent first turn (\"no greeting at all\"). Realtime requires media-stream by nature, so we now infer it. Defensive prewarm guard branch marked with a v8 ignore since the implicit default makes it unreachable in current outbound tests."
827
+ "Voice phone transport defaults to media-stream when OpenAI Realtime or OpenAI SIP is configured but no explicit voice.twilioTransportMode is set. Previously the default was record-play, which made conversationEngine resolve to cascade and routed inbound calls through the ElevenLabs/Whisper greeting path operators with realtime-only credentials never configured \u2014 producing a fully silent first turn (\"no greeting at all\"). Realtime requires media-stream by nature, so we now infer it. Defensive prewarm guard branch marked with a v8 ignore since the implicit default makes it unreachable in current outbound tests."
814
828
  ]
815
829
  },
816
830
  {
817
831
  "version": "0.1.0-alpha.594",
818
832
  "changes": [
819
- "Voice phone transport defaults to media-stream when OpenAI Realtime or OpenAI SIP is configured but no explicit voice.twilioTransportMode is set. Previously the default was record-play, which made conversationEngine resolve to cascade and routed inbound calls through the ElevenLabs/Whisper greeting path operators with realtime-only credentials never configured — producing a fully silent first turn (\"no greeting at all\"). Realtime requires media-stream by nature, so we now infer it."
833
+ "Voice phone transport defaults to media-stream when OpenAI Realtime or OpenAI SIP is configured but no explicit voice.twilioTransportMode is set. Previously the default was record-play, which made conversationEngine resolve to cascade and routed inbound calls through the ElevenLabs/Whisper greeting path operators with realtime-only credentials never configured \u2014 producing a fully silent first turn (\"no greeting at all\"). Realtime requires media-stream by nature, so we now infer it."
820
834
  ]
821
835
  },
822
836
  {
@@ -867,7 +881,7 @@
867
881
  {
868
882
  "version": "0.1.0-alpha.586",
869
883
  "changes": [
870
- "`mail_status`, `mail_recent`, `mail_search`, and `mail_index_refresh` now flag a 'mail substrate divergence' when the encrypted mailroom store reports zero visible messages but the on-disk search cache still holds documents from prior imports — the post-rotation / hosted→local-fallback / wiped-store state that previously rendered as a silent 'no mail' answer indistinguishable from a clean onboarding.",
884
+ "`mail_status`, `mail_recent`, `mail_search`, and `mail_index_refresh` now flag a 'mail substrate divergence' when the encrypted mailroom store reports zero visible messages but the on-disk search cache still holds documents from prior imports \u2014 the post-rotation / hosted\u2192local-fallback / wiped-store state that previously rendered as a silent 'no mail' answer indistinguishable from a clean onboarding.",
871
885
  "Mail absence answers from a divergent runtime now point at vault inspection (`mailroom.mode`, `mailroom.azureAccountUrl`, `mailroom.storePath`) and re-import recovery, so agents stop treating a broken substrate as evidence that the human inbox is empty.",
872
886
  "The substrate-divergence snapshot counts cache `.json` entries via `readdir` and ignores subdirectories and non-json files, so the diagnostic stays cheap on bundles holding tens of thousands of cached documents."
873
887
  ]
@@ -1313,7 +1327,7 @@
1313
1327
  {
1314
1328
  "version": "0.1.0-alpha.527",
1315
1329
  "changes": [
1316
- "Suppresses `onResult`/`onFailure` in the shared tool-activity callbacks factory for any tool that started hidden, so a hidden tool's END never re-emits its raw args into chat surfaces — fixing rejected `settle` calls leaking `answer=`/`intent=` into BlueBubbles and Teams threads.",
1330
+ "Suppresses `onResult`/`onFailure` in the shared tool-activity callbacks factory for any tool that started hidden, so a hidden tool's END never re-emits its raw args into chat surfaces \u2014 fixing rejected `settle` calls leaking `answer=`/`intent=` into BlueBubbles and Teams threads.",
1317
1331
  "Tracks hidden-at-start tools by per-name counter to stay sound across concurrent same-name hidden starts, with no behavior change for visible tools.",
1318
1332
  "Adds heart-level regression tests for hidden-tool END suppression (success and failure paths, concurrent same-name) and senses-level regression tests against `createBlueBubblesCallbacks` and `createTeamsCallbacks` asserting that a rejected settle following a visible read_file produces no chat output containing the settle answer text or `intent=`/`answer=` substrings."
1319
1333
  ]
@@ -1399,63 +1413,63 @@
1399
1413
  "version": "0.1.0-alpha.519",
1400
1414
  "changes": [
1401
1415
  "Introduces `kind: \"library\"` field on bundle `agent.json`. `agent-discovery.ts` filters bundles where `kind === \"library\"` so they're never instantiated as runtime agents. `SerpentGuide.ouro/agent.json` tagged with `kind: \"library\"` to formalize what was previously an implicit `enabled: false` convention.",
1402
- "Activation gate `shouldFireRepairGuide` consumes the existing `untypedDegraded` / `typedDegraded` partitioning at `cli-exec.ts:6693-6694`. Fires when `untypedDegraded.length > 0` OR `typedDegraded.length >= 3`. The existing `--no-repair` flag remains the operator escape hatch — no new env toggle.",
1403
- "Drops the `~/AgentBundles/SerpentGuide.ouro/` override fallback in `getSpecialistIdentitySourceDir` — the in-repo bundle is now the only source. Reasoning per the planning doc: drift surface we don't currently need; cleaner ownership; no override path to maintain. Five referencing files updated (`hatch-flow.ts`, `cli-defaults.ts`, plus their tests). Same constraint extends to RepairGuide from day one — no override mechanism.",
1416
+ "Activation gate `shouldFireRepairGuide` consumes the existing `untypedDegraded` / `typedDegraded` partitioning at `cli-exec.ts:6693-6694`. Fires when `untypedDegraded.length > 0` OR `typedDegraded.length >= 3`. The existing `--no-repair` flag remains the operator escape hatch \u2014 no new env toggle.",
1417
+ "Drops the `~/AgentBundles/SerpentGuide.ouro/` override fallback in `getSpecialistIdentitySourceDir` \u2014 the in-repo bundle is now the only source. Reasoning per the planning doc: drift surface we don't currently need; cleaner ownership; no override path to maintain. Five referencing files updated (`hatch-flow.ts`, `cli-defaults.ts`, plus their tests). Same constraint extends to RepairGuide from day one \u2014 no override mechanism.",
1404
1418
  "`parseRepairProposals` typed parser maps RepairGuide's structured-proposal output into the existing `RepairAction` catalog from `readiness-repair.ts` (`vault-unlock`, `provider-auth`, `provider-use`, etc.). Backfills lane variants and missing fields where unambiguous; rejects malformed proposals.",
1405
- "Slugger-style compound integration fixture as canonical acceptance test (per O6): bad bootstrap state + expired creds + broken remote + drift between agent.json and agent.json simultaneously. Validates the full layer 1→4→2→3 pipeline end-to-end.",
1419
+ "Slugger-style compound integration fixture as canonical acceptance test (per O6): bad bootstrap state + expired creds + broken remote + drift between agent.json and agent.json simultaneously. Validates the full layer 1\u21924\u21922\u21923 pipeline end-to-end.",
1406
1420
  "All gates green: tsc clean, lint clean, code coverage 100%, nerves audit pass."
1407
1421
  ]
1408
1422
  },
1409
1423
  {
1410
1424
  "version": "0.1.0-alpha.518",
1411
1425
  "changes": [
1412
- "Layer 2 of the harness-hardening sequence (1→4→2→3 from `docs/planning/2026-04-28-1900-planning-harness-hardening-and-repairguide.md`). Wires a pre-flight `git pull` over every sync-enabled bundle into `ouro up`, before per-agent provider live-checks, so the post-pull `agent.json` is what live-check reads. First PR in the sequence that mutates working trees; does NOT write to `state/` (verified by a meta-test).",
1413
- "New sync failure taxonomy in `src/heart/sync-classification.ts`: `auth-failed`, `not-found-404`, `network-down`, `dirty-working-tree`, `non-fast-forward`, `merge-conflict`, `timeout-soft`, `timeout-hard`, `unknown` — extends `PendingSyncRecord.classification` additively (legacy `push_rejected`/`pull_rebase_conflict` still work). Pure pattern-matcher: priority order is abort → 404 → auth → network → dirty → conflict → non-fast-forward → unknown.",
1426
+ "Layer 2 of the harness-hardening sequence (1\u21924\u21922\u21923 from `docs/planning/2026-04-28-1900-planning-harness-hardening-and-repairguide.md`). Wires a pre-flight `git pull` over every sync-enabled bundle into `ouro up`, before per-agent provider live-checks, so the post-pull `agent.json` is what live-check reads. First PR in the sequence that mutates working trees; does NOT write to `state/` (verified by a meta-test).",
1427
+ "New sync failure taxonomy in `src/heart/sync-classification.ts`: `auth-failed`, `not-found-404`, `network-down`, `dirty-working-tree`, `non-fast-forward`, `merge-conflict`, `timeout-soft`, `timeout-hard`, `unknown` \u2014 extends `PendingSyncRecord.classification` additively (legacy `push_rejected`/`pull_rebase_conflict` still work). Pure pattern-matcher: priority order is abort \u2192 404 \u2192 auth \u2192 network \u2192 dirty \u2192 conflict \u2192 non-fast-forward \u2192 unknown.",
1414
1428
  "End-to-end `AbortSignal` plumbing. New `runWithTimeouts<T>` wrapper in `src/heart/timeouts.ts` (soft 8s warns, hard 15s aborts via `AbortController`); new async sibling `preTurnPullAsync` in `src/heart/sync.ts` that uses `child_process.execFile(..., { signal })` so the kernel kills the git child when the hard timeout fires. Original sync `preTurnPull` preserved for the per-turn pipeline. Two env knobs for the boot-sync probe: `OURO_BOOT_TIMEOUT_GIT_SOFT` (8000ms) and `OURO_BOOT_TIMEOUT_GIT_HARD` (15000ms).",
1415
1429
  "New `runBootSyncProbe` orchestrator in `src/heart/daemon/boot-sync-probe.ts` aggregates per-bundle findings (each tagged `advisory: true|false`). Wired into `daemon.up` as a new \"sync probe\" boot phase between manual-clone-detection and provider checks. Failures during the probe itself are caught and surfaced as a warning event without blocking the boot. Tests inject `runBootSyncProbeImpl` to keep CI off the developer's home bundles.",
1416
- "9903 tests pass (518 files; +19 new). Coverage gate clean (cli-exec.ts 99.33% → 100%). Slow-remote integration test proves boot doesn't hang on a hung remote (probe aborts within `hardMs`). this PR's meta-test enforces no-state-writes invariant on the three new files."
1430
+ "9903 tests pass (518 files; +19 new). Coverage gate clean (cli-exec.ts 99.33% \u2192 100%). Slow-remote integration test proves boot doesn't hang on a hung remote (probe aborts within `hardMs`). this PR's meta-test enforces no-state-writes invariant on the three new files."
1417
1431
  ]
1418
1432
  },
1419
1433
  {
1420
1434
  "version": "0.1.0-alpha.517",
1421
1435
  "changes": [
1422
- "`computeDaemonRollup` (Layer 1) gains an optional `driftDetected: boolean`. When true, `healthy` → `partial` (same downgrade rule as `bootstrapDegraded`). `degraded` and `safe-mode` rollups are unaffected — drift never escalates past `partial` and never un-downgrades. `daemon-entry.ts` probes each enabled agent for drift before computing the rollup; a single agent's read failure is best-effort and does not block the scan."
1436
+ "`computeDaemonRollup` (Layer 1) gains an optional `driftDetected: boolean`. When true, `healthy` \u2192 `partial` (same downgrade rule as `bootstrapDegraded`). `degraded` and `safe-mode` rollups are unaffected \u2014 drift never escalates past `partial` and never un-downgrades. `daemon-entry.ts` probes each enabled agent for drift before computing the rollup; a single agent's read failure is best-effort and does not block the scan."
1423
1437
  ]
1424
1438
  },
1425
1439
  {
1426
1440
  "version": "0.1.0-alpha.516",
1427
1441
  "changes": [
1428
- "Layer 1 of the harness-hardening sequence (1→4→2→3 from `docs/planning/2026-04-28-1900-planning-harness-hardening-and-repairguide.md`). Replaces the daemon-wide rollup at `daemon-entry.ts` (the binary `degraded.length > 0 ? \"degraded\" : \"ok\"` literal) with a five-state vocabulary: `healthy / partial / degraded / safe-mode / down`. A single sick agent no longer tips the whole daemon to `degraded`.",
1442
+ "Layer 1 of the harness-hardening sequence (1\u21924\u21922\u21923 from `docs/planning/2026-04-28-1900-planning-harness-hardening-and-repairguide.md`). Replaces the daemon-wide rollup at `daemon-entry.ts` (the binary `degraded.length > 0 ? \"degraded\" : \"ok\"` literal) with a five-state vocabulary: `healthy / partial / degraded / safe-mode / down`. A single sick agent no longer tips the whole daemon to `degraded`.",
1429
1443
  "Type structure: `RollupStatus` (4-state, returned by the new pure `computeDaemonRollup` decision function in `daemon-rollup.ts`) and `DaemonStatus = RollupStatus | \"down\"` (full daemon-status; `down` is caller-owned because it represents pre-inventory failure, before the rollup is reachable). Both unions project from a single source-of-truth literal tuple so future widening touches one site.",
1430
- "`renderRollupStatusLine` in `cli-render.ts` uses a compiler-forced `never`-typed exhaustive switch — adding a future state compile-errors at every consumer using the pattern. The `degraded` literal carries three copy variants picked by inspecting cached agent statuses: empty map (fresh install, prompts `ouro hatch`), non-empty + any running agent (legacy stale cache from pre-Layer-1 daemons, prompts `ouro up` refresh), non-empty + zero running (all-failed live-check, prompts `ouro doctor`).",
1444
+ "`renderRollupStatusLine` in `cli-render.ts` uses a compiler-forced `never`-typed exhaustive switch \u2014 adding a future state compile-errors at every consumer using the pattern. The `degraded` literal carries three copy variants picked by inspecting cached agent statuses: empty map (fresh install, prompts `ouro hatch`), non-empty + any running agent (legacy stale cache from pre-Layer-1 daemons, prompts `ouro up` refresh), non-empty + zero running (all-failed live-check, prompts `ouro doctor`).",
1431
1445
  "`runtime-readers.ts:readDaemonHealthDeep` parse tightened to use `isDaemonStatus`. `OutlookDaemonHealthDeep.status` widened to `DaemonStatus | \"unknown\"` so legacy serialized strings (`\"running\"`, `\"ok\"`) coerce defensively rather than failing the parse during rollout.",
1432
- "9759 tests pass (508 test files); coverage gate clean. The per-agent live-check loop in `cli-exec.ts` is intentionally untouched — it was already try/catch-isolated; the bug was in how its output rolled up. Subsequent PRs (layers 4, 2, 3) build on this PR's vocabulary."
1446
+ "9759 tests pass (508 test files); coverage gate clean. The per-agent live-check loop in `cli-exec.ts` is intentionally untouched \u2014 it was already try/catch-isolated; the bug was in how its output rolled up. Subsequent PRs (layers 4, 2, 3) build on this PR's vocabulary."
1433
1447
  ]
1434
1448
  },
1435
1449
  {
1436
1450
  "version": "0.1.0-alpha.515",
1437
1451
  "changes": [
1438
- "New `speak` tool — agent can deliver words to the current friend mid-turn without ending the turn. Pairs with `settle` (ends turn) and `ponder` (private inner thought). For acknowledgment of heavy work, phase-boundary updates, or progress narration on chat-style channels (cli, teams, bluebubbles).",
1439
- "Schema is intentionally minimal: `speak({ message: string })`. Not sole-call, doesn't terminate the turn, NOT exempt from the 24-call circuit breaker (the breaker is healthy backpressure against narration-spam — silence is a natural fallback for speak, unlike settle/rest). Added `flushNow?(): void | Promise<void>` to `ChannelCallbacks`; per-sense impls deliver the buffered message immediately (CLI noop, BlueBubbles `client.sendText` keeping typing on, Teams stream emit with `sendMessage` fallback).",
1452
+ "New `speak` tool \u2014 agent can deliver words to the current friend mid-turn without ending the turn. Pairs with `settle` (ends turn) and `ponder` (private inner thought). For acknowledgment of heavy work, phase-boundary updates, or progress narration on chat-style channels (cli, teams, bluebubbles).",
1453
+ "Schema is intentionally minimal: `speak({ message: string })`. Not sole-call, doesn't terminate the turn, NOT exempt from the 24-call circuit breaker (the breaker is healthy backpressure against narration-spam \u2014 silence is a natural fallback for speak, unlike settle/rest). Added `flushNow?(): void | Promise<void>` to `ChannelCallbacks`; per-sense impls deliver the buffered message immediately (CLI noop, BlueBubbles `client.sendText` keeping typing on, Teams stream emit with `sendMessage` fallback).",
1440
1454
  "Engine integration follows the `ponder` interception template at `core.ts:~1303`: `speak` runs inline (emit + flushNow + push `(spoken)` tool result + nerves event `engine.speak`), then the loop continues. Empty/missing message rejected with a tool-result error and `engine.speak_invalid`. New event keys: `engine.speak`, `engine.speak_invalid`, `engine.speak_delivery_failed`, `bluebubbles.speak_flush`, `teams.speak_flush`.",
1441
- "System prompt nudge in Group #4 (`how i work`), gated to chat-style channels: dependency boundary (settle if next step needs a reply, otherwise speak), phase boundaries (after acking heavy ask / hitting major constraint / switching strategy / before externally-visible step — not per-tool narration), one-way framing (speak is progress, not invitation).",
1442
- "Hardened `speak` delivery semantics (PR review follow-up). `flushNow` contract is now explicit: throws if the message could not be delivered through any available path. Teams `flushNow` THROWS when both stream emit AND `sendMessage` fallback fail (was silently logging delivered=false and returning normally — engine then recorded `(spoken)` even though nothing reached the friend). BlueBubbles `flushNow` already let `client.sendText` rejections propagate; contract documented. Engine wraps `await flushNow()` in try/catch: on hard failure it calls `onToolEnd('speak', ..., false)`, pushes a `'speak delivery failed: ... did not reach your friend; do not assume they saw it'` tool result, emits `engine.speak_delivery_failed` (level=error), and the turn continues — preventing the agent from assuming silent success.",
1443
- "`speak` is now treated as flow-control across all senses (PR review follow-up). Like settle/observe/ponder/rest, its only visible output is the message itself — no spinner, no phrase rotation, no `⏳` placeholder, no tool-activity status line. Added 'speak' to `FLOW_CONTROL_TOOLS` in `cli/tool-display.ts` and `cli/ouro-tui.tsx`; CLI/BlueBubbles/Teams `onToolStart` early-return for speak; `tool-description.ts` returns null for speak as defense-in-depth for any future sense using `createToolActivityCallbacks`. Teams `flushNow` also stops phrase rotation when it delivers, so the actual message replaces the cycling 'thinking...' phrase.",
1444
- "Teams `flushNow` no longer aborts the turn on a successful sendMessage fallback (PR review follow-up). Prior code path: stream emit fails → `tryEmit` calls `markStopped()` which calls `controller.abort()` → falls through to `sendMessage` → succeeds → `flushNow` returns normally → core records `(spoken)` with success=true — but the turn controller is already aborted, so the next model/tool step aborts. Successful fallback delivery should not poison the rest of the turn. Fix adds a non-aborting `tryEmitNoAbort` variant adjacent to `tryEmit`; `flushNow` uses it so a primary-stream failure followed by a successful sendMessage no longer triggers `controller.abort()`. Only when ALL delivery paths fail does `flushNow` call `markStopped()` and throw, letting the engine's existing `engine.speak_delivery_failed` catch path end the turn cleanly. `tryEmit` and other non-flushNow callers (end-of-turn `flush()`, `safeEmit`) are unchanged — their abort-on-failure behavior remains correct because they have no fallback path forward."
1455
+ "System prompt nudge in Group #4 (`how i work`), gated to chat-style channels: dependency boundary (settle if next step needs a reply, otherwise speak), phase boundaries (after acking heavy ask / hitting major constraint / switching strategy / before externally-visible step \u2014 not per-tool narration), one-way framing (speak is progress, not invitation).",
1456
+ "Hardened `speak` delivery semantics (PR review follow-up). `flushNow` contract is now explicit: throws if the message could not be delivered through any available path. Teams `flushNow` THROWS when both stream emit AND `sendMessage` fallback fail (was silently logging delivered=false and returning normally \u2014 engine then recorded `(spoken)` even though nothing reached the friend). BlueBubbles `flushNow` already let `client.sendText` rejections propagate; contract documented. Engine wraps `await flushNow()` in try/catch: on hard failure it calls `onToolEnd('speak', ..., false)`, pushes a `'speak delivery failed: ... did not reach your friend; do not assume they saw it'` tool result, emits `engine.speak_delivery_failed` (level=error), and the turn continues \u2014 preventing the agent from assuming silent success.",
1457
+ "`speak` is now treated as flow-control across all senses (PR review follow-up). Like settle/observe/ponder/rest, its only visible output is the message itself \u2014 no spinner, no phrase rotation, no `\u23f3` placeholder, no tool-activity status line. Added 'speak' to `FLOW_CONTROL_TOOLS` in `cli/tool-display.ts` and `cli/ouro-tui.tsx`; CLI/BlueBubbles/Teams `onToolStart` early-return for speak; `tool-description.ts` returns null for speak as defense-in-depth for any future sense using `createToolActivityCallbacks`. Teams `flushNow` also stops phrase rotation when it delivers, so the actual message replaces the cycling 'thinking...' phrase.",
1458
+ "Teams `flushNow` no longer aborts the turn on a successful sendMessage fallback (PR review follow-up). Prior code path: stream emit fails \u2192 `tryEmit` calls `markStopped()` which calls `controller.abort()` \u2192 falls through to `sendMessage` \u2192 succeeds \u2192 `flushNow` returns normally \u2192 core records `(spoken)` with success=true \u2014 but the turn controller is already aborted, so the next model/tool step aborts. Successful fallback delivery should not poison the rest of the turn. Fix adds a non-aborting `tryEmitNoAbort` variant adjacent to `tryEmit`; `flushNow` uses it so a primary-stream failure followed by a successful sendMessage no longer triggers `controller.abort()`. Only when ALL delivery paths fail does `flushNow` call `markStopped()` and throw, letting the engine's existing `engine.speak_delivery_failed` catch path end the turn cleanly. `tryEmit` and other non-flushNow callers (end-of-turn `flush()`, `safeEmit`) are unchanged \u2014 their abort-on-failure behavior remains correct because they have no fallback path forward."
1445
1459
  ]
1446
1460
  },
1447
1461
  {
1448
1462
  "version": "0.1.0-alpha.514",
1449
1463
  "changes": [
1450
1464
  "Add `--strict` flag to `ouro doctor` and bundle the `--category` flag (also in #637) for a coherent CI-friendly diagnostic interface. `--strict` makes the CLI exit non-zero (via thrown error caught by ouro-entry) when any check is `warn` or `fail`. Default behavior is unchanged.",
1451
- "Composes naturally with --category and --json (#634): `ouro doctor --category Daemon --strict --json` is the canonical CI invocation — runs only the daemon checker, exits 1 on any issue, output is parseable. Emits `daemon.doctor_run` with `strict: true` in meta when set so the strict-failure events are filterable through #622's `nerves-review`.",
1465
+ "Composes naturally with --category and --json (#634): `ouro doctor --category Daemon --strict --json` is the canonical CI invocation \u2014 runs only the daemon checker, exits 1 on any issue, output is parseable. Emits `daemon.doctor_run` with `strict: true` in meta when set so the strict-failure events are filterable through #622's `nerves-review`.",
1452
1466
  "5 new parse tests cover --strict alone, --strict + --category combined, and the existing happy-path / no-value cases. cli-types now models `{ kind: \"doctor\", category?, strict? }`. KNOWN_DOCTOR_CATEGORIES (also from #637) gives external tooling a stable list of available filters. 69/69 doctor + parse tests pass."
1453
1467
  ]
1454
1468
  },
1455
1469
  {
1456
1470
  "version": "0.1.0-alpha.513",
1457
1471
  "changes": [
1458
- "New `Friends` category in `ouro doctor`. Same shape as the Mailroom (#632) and Trips (#631) checks — walk each agent bundle, classify the friend store's health, and report a trust-level breakdown for the healthy path.",
1472
+ "New `Friends` category in `ouro doctor`. Same shape as the Mailroom (#632) and Trips (#631) checks \u2014 walk each agent bundle, classify the friend store's health, and report a trust-level breakdown for the healthy path.",
1459
1473
  "Per-agent reports: pass when no friends/ dir (no friends recorded yet), pass with `<N> friends, <X> family, <Y> friend, <Z> stranger` when records parse cleanly, warn when some files are unparseable (with parse-failure count), fail when the dir itself can't be read. Records with an unrecognized `trustLevel` get counted under `<N> other` so the operator can investigate.",
1460
1474
  "6 new tests cover all branches plus file-extension filtering (`.txt` ignored). Wired between Security and Disk in CATEGORY_CHECKERS, same orchestration shape as Mailroom and Trips."
1461
1475
  ]
@@ -1463,40 +1477,40 @@
1463
1477
  {
1464
1478
  "version": "0.1.0-alpha.512",
1465
1479
  "changes": [
1466
- "New `friend_list` tool — the agent can list its known friends with id, name, and trust level. The notes/friend repertoire had `get_friend_note` and `save_friend_note` for individual lookups but no surface to enumerate the friend graph. Real workflows that needed it: cross-chat outreach decisions, screener triage, orienting on relationships at session start.",
1480
+ "New `friend_list` tool \u2014 the agent can list its known friends with id, name, and trust level. The notes/friend repertoire had `get_friend_note` and `save_friend_note` for individual lookups but no surface to enumerate the friend graph. Real workflows that needed it: cross-chat outreach decisions, screener triage, orienting on relationships at session start.",
1467
1481
  "Optional `trust` filter (`family`/`friend`/`stranger`) and `limit` (1-200, default 50). Renders one entry per friend with id, trust label, name, and external-id channel:identifier pairs when present. Sorted alphabetically by display name. Empty-state messages are filter-aware so the agent can tell the difference between 'no friends at all' and 'no matches for the filter'.",
1468
- "Defensively handles a friend store that lacks `listAll` (the interface marks it optional) — returns 'the configured friend store does not support listing' rather than throwing. Tool registry up to 75 (snapshot regenerated, H10 contract list extended). 4 new tests cover sorted listing, trust filtering, store-without-listAll defensive path, and filter-empty state."
1482
+ "Defensively handles a friend store that lacks `listAll` (the interface marks it optional) \u2014 returns 'the configured friend store does not support listing' rather than throwing. Tool registry up to 75 (snapshot regenerated, H10 contract list extended). 4 new tests cover sorted listing, trust filtering, store-without-listAll defensive path, and filter-empty state."
1469
1483
  ]
1470
1484
  },
1471
1485
  {
1472
1486
  "version": "0.1.0-alpha.511",
1473
1487
  "changes": [
1474
1488
  "Add `--json` flag to `ouro doctor`. Default human-readable output is unchanged; `--json` emits the full DoctorResult (categories + checks + summary) as pretty-printed JSON for piping into jq, dashboards, scheduled health monitors, or CI pipelines that want to alert on a doctor failure.",
1475
- "Same shape as the `--json` flag on the diagnostic family that's been growing alongside (session-playback / nerves-review / session-stats in their respective open PRs). Doctor was the asymmetric case — text-only output. With agent-private state coverage now landing in the doctor (Trips in #631, Mailroom in #632), structured output makes the doctor genuinely consumable by external tooling.",
1476
- "1 new test on the parse path (`doctor --json` → `{ kind: \"doctor\", json: true }`); the existing parse test is updated to match the new shape (`{ kind: \"doctor\", json: false }`). The exec-side change is a one-line ternary that picks `JSON.stringify(result, null, 2) + \"\\n\"` over `formatDoctorOutput(result)` when the flag is set. 83/83 doctor + parse tests pass."
1489
+ "Same shape as the `--json` flag on the diagnostic family that's been growing alongside (session-playback / nerves-review / session-stats in their respective open PRs). Doctor was the asymmetric case \u2014 text-only output. With agent-private state coverage now landing in the doctor (Trips in #631, Mailroom in #632), structured output makes the doctor genuinely consumable by external tooling.",
1490
+ "1 new test on the parse path (`doctor --json` \u2192 `{ kind: \"doctor\", json: true }`); the existing parse test is updated to match the new shape (`{ kind: \"doctor\", json: false }`). The exec-side change is a one-line ternary that picks `JSON.stringify(result, null, 2) + \"\\n\"` over `formatDoctorOutput(result)` when the flag is set. 83/83 doctor + parse tests pass."
1477
1491
  ]
1478
1492
  },
1479
1493
  {
1480
1494
  "version": "0.1.0-alpha.510",
1481
1495
  "changes": [
1482
1496
  "New `Mailroom` category in `ouro doctor`. Same shape as the `Trips` check from #631 (alpha.509): walk each agent bundle, classify the mailroom registry's health into pass/warn/fail with structured detail, return record/grant/message counts for healthy ledgers.",
1483
- "Health classes: pass when no mailroom dir (mail not connected), pass with `<N> mailboxes, <N> source grants, <N> messages` when healthy, warn when registry.json is missing or has zero mailboxes, fail when registry.json is unreadable or unparseable. Message count walks the messages dir and counts `.json` files only — stray `.txt` etc. are ignored.",
1484
- "6 new tests cover all branches plus message-count file-extension filtering. Wired between `Security` and `Disk` in `CATEGORY_CHECKERS`. With #631's Trips check landing alongside, doctor now has agent-private state coverage for the two stateful primitives the agent owns (mail + trips) — same shape, same detail format, same orchestration."
1497
+ "Health classes: pass when no mailroom dir (mail not connected), pass with `<N> mailboxes, <N> source grants, <N> messages` when healthy, warn when registry.json is missing or has zero mailboxes, fail when registry.json is unreadable or unparseable. Message count walks the messages dir and counts `.json` files only \u2014 stray `.txt` etc. are ignored.",
1498
+ "6 new tests cover all branches plus message-count file-extension filtering. Wired between `Security` and `Disk` in `CATEGORY_CHECKERS`. With #631's Trips check landing alongside, doctor now has agent-private state coverage for the two stateful primitives the agent owns (mail + trips) \u2014 same shape, same detail format, same orchestration."
1485
1499
  ]
1486
1500
  },
1487
1501
  {
1488
1502
  "version": "0.1.0-alpha.509",
1489
1503
  "changes": [
1490
- "New `Trips` category in `ouro doctor`. Operators previously had no quick way to verify the trip ledger was healthy — they had to call `trip_status` from inside the agent or open `state/trips/ledger.json` by hand. With the trip ledger now load-bearing for trip planning workflows, doctor coverage matters.",
1504
+ "New `Trips` category in `ouro doctor`. Operators previously had no quick way to verify the trip ledger was healthy \u2014 they had to call `trip_status` from inside the agent or open `state/trips/ledger.json` by hand. With the trip ledger now load-bearing for trip planning workflows, doctor coverage matters.",
1491
1505
  "The check walks each agent bundle and reports per-agent trip health: pass when the ledger is absent (optional feature, not yet ensured), warn when `state/trips/` exists but `ledger.json` is missing or lacks the `ledgerId` field, fail when `ledger.json` is unreadable, unparseable, or missing the `privateKeyPem` (encrypted records would be unreadable). Healthy ledgers report `<ledgerId> (<N> records)` so the operator sees record count without opening anything.",
1492
- "7 new tests cover: no-agents (warn), no-ledger-dir (pass — optional), missing ledger.json (warn), unparseable JSON (fail), missing ledgerId (warn), missing privateKeyPem (fail), and the healthy passing path with record-counting that ignores non-`.json` files in the records dir. Wired between `Security` and `Disk` in the `CATEGORY_CHECKERS` array — same orchestration shape as the existing categories."
1506
+ "7 new tests cover: no-agents (warn), no-ledger-dir (pass \u2014 optional), missing ledger.json (warn), unparseable JSON (fail), missing ledgerId (warn), missing privateKeyPem (fail), and the healthy passing path with record-counting that ignores non-`.json` files in the records dir. Wired between `Security` and `Disk` in the `CATEGORY_CHECKERS` array \u2014 same orchestration shape as the existing categories."
1493
1507
  ]
1494
1508
  },
1495
1509
  {
1496
1510
  "version": "0.1.0-alpha.508",
1497
1511
  "changes": [
1498
- "New `trip_remove_leg` tool. The trip ledger had `trip_upsert`, `trip_attach_evidence`, and `trip_update_leg`, but no first-class way to drop a leg — the agent had to re-emit the entire trip record minus that leg (fragile, easy to lose evidence). Real workflow that needed it: the user cancelled a hotel booking; #620's e2e test exercised the add path but there was no way to model the cancel.",
1499
- "`trip_remove_leg(tripId, legId, updatedAt, reason?)` finds the leg, drops it from `legs[]`, bumps the trip's `updatedAt`, and emits `trips.leg_removed` (info) carrying tripId/legId/kind/reason. Rejects when the leg id is unknown (so accidental no-ops are visible) and when the trip is missing (returns the same `trip not found` shape as the other tools). Tool registry up to 75 (now 8 trip tools — snapshot updated, H10 contract list includes the new name).",
1512
+ "New `trip_remove_leg` tool. The trip ledger had `trip_upsert`, `trip_attach_evidence`, and `trip_update_leg`, but no first-class way to drop a leg \u2014 the agent had to re-emit the entire trip record minus that leg (fragile, easy to lose evidence). Real workflow that needed it: the user cancelled a hotel booking; #620's e2e test exercised the add path but there was no way to model the cancel.",
1513
+ "`trip_remove_leg(tripId, legId, updatedAt, reason?)` finds the leg, drops it from `legs[]`, bumps the trip's `updatedAt`, and emits `trips.leg_removed` (info) carrying tripId/legId/kind/reason. Rejects when the leg id is unknown (so accidental no-ops are visible) and when the trip is missing (returns the same `trip not found` shape as the other tools). Tool registry up to 75 (now 8 trip tools \u2014 snapshot updated, H10 contract list includes the new name).",
1500
1514
  "5 tests cover happy-path removal with leg-count assertion via `trip_get`, unknown-leg rejection, missing-trip propagation, the three required-field validation paths, plus the existing stranger-ctx trust block now extended to include `trip_remove_leg`."
1501
1515
  ]
1502
1516
  },
@@ -1510,32 +1524,32 @@
1510
1524
  {
1511
1525
  "version": "0.1.0-alpha.506",
1512
1526
  "changes": [
1513
- "Detect duplicate tool_call_id across assistant messages in `validateSessionMessages`. MiniMax-M2.7 emits canonical tool_call ids of the form `call_function_<hash>_<n>` and reuses the same id across turns when the same function gets called — which causes provider rejections on replay because tool_call_id is supposed to be unique per request. The session sanitize pass already had position-aware orphan detection (#613) and inline-reasoning strip (#612); this adds the third member of the family — collision detection.",
1514
- "New exported `detectDuplicateToolCallIds(messages)` returns `{ id, indices }[]` for each tool_call_id that appears in multiple assistant messages. Same-message duplicates (one assistant calling the same id twice) are not flagged — those are a legitimate parallel-call shape. `validateSessionMessages` now folds collisions into its violations list with a message that calls out MiniMax specifically so operators reading nerves know what they're looking at.",
1515
- "Detection only — no rewriting yet, since rewriting tool_call_ids and the matching tool_results requires careful pairing logic that risks regression. The collision is visible to operators via the `mind.session_invariant_violation` nerves event the sanitize pass already emits when violations are present, and the existing `nerves-review` CLI from #622 makes it filterable. 3 new tests cover collision detection, single-message parallel-call shape (no false positive), and the all-distinct happy path."
1527
+ "Detect duplicate tool_call_id across assistant messages in `validateSessionMessages`. MiniMax-M2.7 emits canonical tool_call ids of the form `call_function_<hash>_<n>` and reuses the same id across turns when the same function gets called \u2014 which causes provider rejections on replay because tool_call_id is supposed to be unique per request. The session sanitize pass already had position-aware orphan detection (#613) and inline-reasoning strip (#612); this adds the third member of the family \u2014 collision detection.",
1528
+ "New exported `detectDuplicateToolCallIds(messages)` returns `{ id, indices }[]` for each tool_call_id that appears in multiple assistant messages. Same-message duplicates (one assistant calling the same id twice) are not flagged \u2014 those are a legitimate parallel-call shape. `validateSessionMessages` now folds collisions into its violations list with a message that calls out MiniMax specifically so operators reading nerves know what they're looking at.",
1529
+ "Detection only \u2014 no rewriting yet, since rewriting tool_call_ids and the matching tool_results requires careful pairing logic that risks regression. The collision is visible to operators via the `mind.session_invariant_violation` nerves event the sanitize pass already emits when violations are present, and the existing `nerves-review` CLI from #622 makes it filterable. 3 new tests cover collision detection, single-message parallel-call shape (no false positive), and the all-distinct happy path."
1516
1530
  ]
1517
1531
  },
1518
1532
  {
1519
1533
  "version": "0.1.0-alpha.505",
1520
1534
  "changes": [
1521
1535
  "New `ouro session-stats <session.json>` CLI for at-a-glance metrics on a saved session: total events, breakdown by role (system/user/assistant/tool), tool-call totals + top 5 by frequency, attachment count, time range with duration, projection breakdown (in/out, input tokens, max tokens, trimmed), and last usage. Read-only.",
1522
- "Pure `computeSessionStats(envelope, path)` core in `src/heart/session-stats.ts` — testable with synthesized envelopes, embeddable in future doctor checks. `runSessionStats(path)` adds the file-load layer; `formatStatsReport(report)` renders human-readable text; `--json` mode for jq piping. Composes with #619 (session-playback) and #622 (nerves-review): three pure-analyzer-plus-thin-CLI tools that together make a stuck session immediately diagnosable end-to-end.",
1536
+ "Pure `computeSessionStats(envelope, path)` core in `src/heart/session-stats.ts` \u2014 testable with synthesized envelopes, embeddable in future doctor checks. `runSessionStats(path)` adds the file-load layer; `formatStatsReport(report)` renders human-readable text; `--json` mode for jq piping. Composes with #619 (session-playback) and #622 (nerves-review): three pure-analyzer-plus-thin-CLI tools that together make a stuck session immediately diagnosable end-to-end.",
1523
1537
  "8 tests cover role counts, tool-call name aggregation with frequency-sorted top-5, time range with and without authoredAt timestamps, attachment counting, projection-omission detection, the unrecognized-envelope stub, CLI no-args help, and CLI --json output. Wired as `npm run session:stats -- <path>` and `dist/heart/session-stats-cli-main.js`."
1524
1538
  ]
1525
1539
  },
1526
1540
  {
1527
1541
  "version": "0.1.0-alpha.504",
1528
1542
  "changes": [
1529
- "New `mail_outbox` tool — the agent can now introspect its own outbound mail (drafts, queued sends, delivered, bounced, etc.). The mail repertoire had `mail_compose`, `mail_send`, and `mail_recent` for inbound — but no symmetric way to ask 'what did I send / queue?' Operators were having to ssh in and `ls state/.../outbound`. Real-world need: when planning a trip with the operator, the agent often wants to verify it sent a confirmation request before re-asking.",
1530
- "Lists records newest-first (by `updatedAt`), bounded to `limit` (1-50, default 20), with optional `status` filter across the full MailOutboundStatus union (draft / sent / submitted / accepted / delivered / bounced / suppressed / quarantined / spam-filtered / failed). Each record renders id + status + recipients + truncated subject (80 chars) + last-touched timestamp + provider message id and error message when present. No body text dumped — agent uses message id with another tool if it needs the content.",
1543
+ "New `mail_outbox` tool \u2014 the agent can now introspect its own outbound mail (drafts, queued sends, delivered, bounced, etc.). The mail repertoire had `mail_compose`, `mail_send`, and `mail_recent` for inbound \u2014 but no symmetric way to ask 'what did I send / queue?' Operators were having to ssh in and `ls state/.../outbound`. Real-world need: when planning a trip with the operator, the agent often wants to verify it sent a confirmation request before re-asking.",
1544
+ "Lists records newest-first (by `updatedAt`), bounded to `limit` (1-50, default 20), with optional `status` filter across the full MailOutboundStatus union (draft / sent / submitted / accepted / delivered / bounced / suppressed / quarantined / spam-filtered / failed). Each record renders id + status + recipients + truncated subject (80 chars) + last-touched timestamp + provider message id and error message when present. No body text dumped \u2014 agent uses message id with another tool if it needs the content.",
1531
1545
  "Family-trust gated like the rest of mail (read gate, no special block since outbound metadata isn't body content). Records `mail_outbox` access in the access log alongside the other mail tools. Tool registry now at 75 tools (snapshot updated). Two tests cover the empty / sorted / limit / status-filter / audit-log paths, plus the trust block."
1532
1546
  ]
1533
1547
  },
1534
1548
  {
1535
1549
  "version": "0.1.0-alpha.503",
1536
1550
  "changes": [
1537
- "In-process LRU cache for decrypted mail bodies. The cold path for `mail_thread` is read-encrypted-blob-from-Azure (1-3s p50, up to tens of seconds for HEY-sized bodies — #614 raised the timeout to 60s for this very reason) plus an RSA-OAEP+A256GCM decrypt. Repeated reads of the same message are common: re-checking a booking confirmation while seeding a trip leg, following up on a thread, looping back to verify a fact. Each repeat hit was paying the full cold cost.",
1538
- "New `src/mailroom/body-cache.ts` keeps a 50-entry LRU keyed by `StoredMailMessage.id` (a deterministic content hash — rotating keys produces a new id, so stale ciphertext can never be served against a fresh keyset). Insertion-order eviction; reads refresh LRU position. Per-process by design — daemon restart clears it (matches the established pattern with #618 heartbeat-recursion state and #621 BB own-handle discovery).",
1551
+ "In-process LRU cache for decrypted mail bodies. The cold path for `mail_thread` is read-encrypted-blob-from-Azure (1-3s p50, up to tens of seconds for HEY-sized bodies \u2014 #614 raised the timeout to 60s for this very reason) plus an RSA-OAEP+A256GCM decrypt. Repeated reads of the same message are common: re-checking a booking confirmation while seeding a trip leg, following up on a thread, looping back to verify a fact. Each repeat hit was paying the full cold cost.",
1552
+ "New `src/mailroom/body-cache.ts` keeps a 50-entry LRU keyed by `StoredMailMessage.id` (a deterministic content hash \u2014 rotating keys produces a new id, so stale ciphertext can never be served against a fresh keyset). Insertion-order eviction; reads refresh LRU position. Per-process by design \u2014 daemon restart clears it (matches the established pattern with #618 heartbeat-recursion state and #621 BB own-handle discovery).",
1539
1553
  "Wired into both `mail_thread` (cache-first read; on miss, do the disk fetch + decrypt and cache for next time) and `mail_recent`/`mail_search` (which already decrypt batches; now they also seed the body cache so the next `mail_thread` on any of those is free). New `repertoire.mail_body_cache_hit` info-level event makes hit rate observable via `ouro nerves-review --event mail_body_cache_hit` (alpha.501). 7 new tests cover hit/miss, LRU refresh-on-read, eviction at capacity, defensive empty-id handling, and clear."
1540
1554
  ]
1541
1555
  },
@@ -1543,60 +1557,60 @@
1543
1557
  "version": "0.1.0-alpha.502",
1544
1558
  "changes": [
1545
1559
  "Enrich `engine.error` nerve event with HTTP status, redacted body excerpt, and a one-line summary string. Provider errors previously surfaced only as a free-form `error.message`, which forced operators to spelunk the SDK's wrapped object to find the actual status code or quota explanation.",
1546
- "Two new helpers in `src/heart/providers/error-classification.ts`: `extractProviderErrorDetails(error)` pulls `status` (when present) and a body excerpt (capped at 240 chars, with redaction of any 32+ char token-shaped substring so leaked auth keys don't get persisted into nerves), falling through `error.error → error.response → error.body → error.message` until something usable shows up. Survives circular structures defensively. `summarizeProviderError(error, classification, providerId, model)` produces the canonical operator-readable line: `provider <id>/<model>: <classification>[ HTTP <status>][ — <bodyExcerpt>]`.",
1547
- "Wired into `finishTerminalProviderError` in `src/heart/core.ts` so every terminal provider error now lands in nerves with `httpStatus` + `bodyExcerpt` + `summary` meta — making `ouro nerves-review --component engine --event engine.error` (alpha.501) immediately useful for diagnosing provider blowups. 11 new tests cover status capture, missing-status defaults, token redaction, 240-char truncation, fallback through alternate body fields, circular-structure safety, and summary formatting in two shapes."
1560
+ "Two new helpers in `src/heart/providers/error-classification.ts`: `extractProviderErrorDetails(error)` pulls `status` (when present) and a body excerpt (capped at 240 chars, with redaction of any 32+ char token-shaped substring so leaked auth keys don't get persisted into nerves), falling through `error.error \u2192 error.response \u2192 error.body \u2192 error.message` until something usable shows up. Survives circular structures defensively. `summarizeProviderError(error, classification, providerId, model)` produces the canonical operator-readable line: `provider <id>/<model>: <classification>[ HTTP <status>][ \u2014 <bodyExcerpt>]`.",
1561
+ "Wired into `finishTerminalProviderError` in `src/heart/core.ts` so every terminal provider error now lands in nerves with `httpStatus` + `bodyExcerpt` + `summary` meta \u2014 making `ouro nerves-review --component engine --event engine.error` (alpha.501) immediately useful for diagnosing provider blowups. 11 new tests cover status capture, missing-status defaults, token redaction, 240-char truncation, fallback through alternate body fields, circular-structure safety, and summary formatting in two shapes."
1548
1562
  ]
1549
1563
  },
1550
1564
  {
1551
1565
  "version": "0.1.0-alpha.501",
1552
1566
  "changes": [
1553
1567
  "New `ouro nerves-review` CLI for tailing the agent's nerves ndjson with structured filters. Read-only. Operators previously had to grep raw ndjson by hand to track down something like 'how many heartbeat-recursion-suspected events fired today' or 'show me the last hour of senses warnings'.",
1554
- "Filters: `--component <substr>`, `--event <substr>`, `--level <level>`, `--since <duration>` (e.g. 5m, 2h, 1d), `--limit <N>`, `--process <name>` (default: daemon), `--agent <name>` (default: current). Output modes: human-readable text (`<time> [<level>] <component>/<event> — <message>`) and `--json` (one parsed object per line for piping to jq).",
1568
+ "Filters: `--component <substr>`, `--event <substr>`, `--level <level>`, `--since <duration>` (e.g. 5m, 2h, 1d), `--limit <N>`, `--process <name>` (default: daemon), `--agent <name>` (default: current). Output modes: human-readable text (`<time> [<level>] <component>/<event> \u2014 <message>`) and `--json` (one parsed object per line for piping to jq).",
1555
1569
  "Pure `reviewNerveEvents(filePath, filter)` core in `src/nerves/review/core.ts` reads the tail of the ndjson (8 MB cap, walks last 200+ lines) and applies in-memory filters; testable without filesystem mocks beyond a temp file. 12 tests cover all six filter dimensions plus duration parsing edge cases (ms/s/m/h/d, malformed inputs), missing-file handling, and the two CLI flag paths (--help, invalid --since). Wired as `npm run nerves:review -- <flags>` and `dist/nerves/review/cli-main.js`."
1556
1570
  ]
1557
1571
  },
1558
1572
  {
1559
1573
  "version": "0.1.0-alpha.500",
1560
1574
  "changes": [
1561
- "Auto-discover BlueBubbles agent handles on isFromMe outbound. The `bluebubbles.ownHandles` config field added in #610 closes the group-echo self-talk loop, but only after an operator manually populates it with the right handle format. Until then, the very bug the field is supposed to fix can fire — the agent ingests its own group echo and replies to itself.",
1575
+ "Auto-discover BlueBubbles agent handles on isFromMe outbound. The `bluebubbles.ownHandles` config field added in #610 closes the group-echo self-talk loop, but only after an operator manually populates it with the right handle format. Until then, the very bug the field is supposed to fix can fire \u2014 the agent ingests its own group echo and replies to itself.",
1562
1576
  "When a normalized BlueBubbles event arrives with `event.fromMe === true`, BlueBubbles is telling us the canonical handle BB attributes to the agent's outbound. We capture `event.sender.externalId` into an in-process `discoveredOwnHandles` set and emit an info-level `senses.bluebubbles_own_handle_discovered` nerve event with the captured handle, so an operator can promote it to durable config (cross-restart). The default `getOwnHandles` now returns the union of configured + discovered; `isAgentSelfHandle` therefore filters subsequent isFromMe-missing group echoes even before the operator updates the vault config.",
1563
- "Per-process state by design — a daemon restart re-learns from the next outbound. Three new tests cover: capture-and-dedupe (raw/normalized form match collapses to one entry), defensive empty/whitespace input handling, and the end-to-end proof that `isAgentSelfHandle` honors discovered handles after `recordDiscoveredOwnHandle` fires."
1577
+ "Per-process state by design \u2014 a daemon restart re-learns from the next outbound. Three new tests cover: capture-and-dedupe (raw/normalized form match collapses to one entry), defensive empty/whitespace input handling, and the end-to-end proof that `isAgentSelfHandle` honors discovered handles after `recordDiscoveredOwnHandle` fires."
1564
1578
  ]
1565
1579
  },
1566
1580
  {
1567
1581
  "version": "0.1.0-alpha.499",
1568
1582
  "changes": [
1569
- "New `ouro session-playback <session.json>` CLI for dry-running the sanitize pipeline against a saved session. When an agent is stuck in a replay loop, an operator can now run the same `sanitizeProviderMessages` chain that the harness fires before every replay, see what would be dropped/modified/synthesized, and decide whether to clear or hand-repair the session — *without* writing anything to disk.",
1570
- "The report distinguishes three repair classes: dropped (orphan tool results whose preceding assistant has no matching tool_call), modified-content (assistant messages whose inline `<think>...</think>` blocks would be stripped before replay), and synthetic-added (synthetic tool-results inserted to satisfy the provider's tool_call/tool_result pairing — these include the explanatory message added in #612 so the agent can read what happened). Each change carries a role, index, optional tool_call_id, reason, and a 120-char preview of the affected content.",
1571
- "Two output modes: human-readable text (default) and `--json` for piping into jq/diagnostics. Underlying `runSessionPlayback` is a pure function — takes either a session path or a raw object — so it's testable in isolation and the same code path can be embedded in future doctor checks. Wired as `npm run session:playback -- <path>` and as the `dist/heart/session-playback-cli-main.js` entry. 7 tests cover the four envelope shapes (clean legacy, with stripped think, with orphan tool result, unrecognized) plus the two CLI flag paths."
1583
+ "New `ouro session-playback <session.json>` CLI for dry-running the sanitize pipeline against a saved session. When an agent is stuck in a replay loop, an operator can now run the same `sanitizeProviderMessages` chain that the harness fires before every replay, see what would be dropped/modified/synthesized, and decide whether to clear or hand-repair the session \u2014 *without* writing anything to disk.",
1584
+ "The report distinguishes three repair classes: dropped (orphan tool results whose preceding assistant has no matching tool_call), modified-content (assistant messages whose inline `<think>...</think>` blocks would be stripped before replay), and synthetic-added (synthetic tool-results inserted to satisfy the provider's tool_call/tool_result pairing \u2014 these include the explanatory message added in #612 so the agent can read what happened). Each change carries a role, index, optional tool_call_id, reason, and a 120-char preview of the affected content.",
1585
+ "Two output modes: human-readable text (default) and `--json` for piping into jq/diagnostics. Underlying `runSessionPlayback` is a pure function \u2014 takes either a session path or a raw object \u2014 so it's testable in isolation and the same code path can be embedded in future doctor checks. Wired as `npm run session:playback -- <path>` and as the `dist/heart/session-playback-cli-main.js` entry. 7 tests cover the four envelope shapes (clean legacy, with stripped think, with orphan tool result, unrecognized) plus the two CLI flag paths."
1572
1586
  ]
1573
1587
  },
1574
1588
  {
1575
1589
  "version": "0.1.0-alpha.498",
1576
1590
  "changes": [
1577
- "Heartbeat / habit recursion detection in the inner-dialog worker. The existing instinct cap (`MAX_CONSECUTIVE_INSTINCT_TURNS=3`) protects against the *internal* pending-dir self-loop (a turn writes back to its own pending dir, drains it, repeats). It does not protect against the *external* IPC self-loop where heartbeat-shaped messages get re-issued faster than their cadence — e.g. a hook misconfigured to repost on every heartbeat, a daemon retry storm, or two timers drifting into the same window.",
1578
- "Two new warn-level nerve events: `senses.habit_recursion_suspected` fires when two of the same habit (e.g. `heartbeat`) arrive within `HABIT_RECURSION_MIN_INTERVAL_MS` (5s) — no realistic cadence runs that fast. `senses.habit_recursion_burst` fires when `HABIT_RECURSION_BURST_THRESHOLD` (5) or more habit messages of any kind land within `HABIT_RECURSION_BURST_WINDOW_MS` (60s) — catches slower runaways that stay just under the min-interval threshold.",
1579
- "Detection is observation-only by design: it emits the warn signal so an operator (or a follow-up auto-recovery layer) can act on it. The message is not dropped — the signal is the value. Per-habit-name tracking, so two distinct habits firing close together don't trip the min-interval warning. `nowSource` is injectable via the `createInnerDialogWorker` factory for deterministic tests. 5 new tests cover both detectors plus the trim-window and per-habit isolation cases."
1591
+ "Heartbeat / habit recursion detection in the inner-dialog worker. The existing instinct cap (`MAX_CONSECUTIVE_INSTINCT_TURNS=3`) protects against the *internal* pending-dir self-loop (a turn writes back to its own pending dir, drains it, repeats). It does not protect against the *external* IPC self-loop where heartbeat-shaped messages get re-issued faster than their cadence \u2014 e.g. a hook misconfigured to repost on every heartbeat, a daemon retry storm, or two timers drifting into the same window.",
1592
+ "Two new warn-level nerve events: `senses.habit_recursion_suspected` fires when two of the same habit (e.g. `heartbeat`) arrive within `HABIT_RECURSION_MIN_INTERVAL_MS` (5s) \u2014 no realistic cadence runs that fast. `senses.habit_recursion_burst` fires when `HABIT_RECURSION_BURST_THRESHOLD` (5) or more habit messages of any kind land within `HABIT_RECURSION_BURST_WINDOW_MS` (60s) \u2014 catches slower runaways that stay just under the min-interval threshold.",
1593
+ "Detection is observation-only by design: it emits the warn signal so an operator (or a follow-up auto-recovery layer) can act on it. The message is not dropped \u2014 the signal is the value. Per-habit-name tracking, so two distinct habits firing close together don't trip the min-interval warning. `nowSource` is injectable via the `createInnerDialogWorker` factory for deterministic tests. 5 new tests cover both detectors plus the trim-window and per-habit isolation cases."
1580
1594
  ]
1581
1595
  },
1582
1596
  {
1583
1597
  "version": "0.1.0-alpha.497",
1584
1598
  "changes": [
1585
- "Mail thread reconstruction + tool rename. The previous `mail_thread` tool was misleadingly named — it returned ONE message body, not a thread. Renamed to `mail_body`. The new actual conversation walker now owns the canonical name `mail_thread`. Existing tests, audit-log strings, and CLI guidance updated to match.",
1586
- "Header capture: `PrivateMailEnvelope` carries optional `inReplyTo` and `references` fields, populated at `buildStoredMailMessage` time from RFC822 headers. Existing messages without these headers are unaffected. `mail_thread` walks the thread from any seed message (storage id or RFC822 `<message-id@host>`): ancestors via `In-Reply-To`/`References`, descendants by reverse-edges across the recent message pool (default 200, configurable 20-500, scoped native/delegated/all), assigns true reply-chain depth via topological longest-path, and renders chronologically with depth-indented summaries. Bodies not included — `mail_body` opens one message.",
1599
+ "Mail thread reconstruction + tool rename. The previous `mail_thread` tool was misleadingly named \u2014 it returned ONE message body, not a thread. Renamed to `mail_body`. The new actual conversation walker now owns the canonical name `mail_thread`. Existing tests, audit-log strings, and CLI guidance updated to match.",
1600
+ "Header capture: `PrivateMailEnvelope` carries optional `inReplyTo` and `references` fields, populated at `buildStoredMailMessage` time from RFC822 headers. Existing messages without these headers are unaffected. `mail_thread` walks the thread from any seed message (storage id or RFC822 `<message-id@host>`): ancestors via `In-Reply-To`/`References`, descendants by reverse-edges across the recent message pool (default 200, configurable 20-500, scoped native/delegated/all), assigns true reply-chain depth via topological longest-path, and renders chronologically with depth-indented summaries. Bodies not included \u2014 `mail_body` opens one message.",
1587
1601
  "Pure thread-walker (`src/mailroom/thread.ts`) is testable without the filesystem: 7 unit tests cover mid-thread seed (walks both directions), seed by RFC822 message-id when storage id doesn't match, References-only (no In-Reply-To, common in list mailers), unrelated-message exclusion, empty/whitespace defensiveness. Plus 3 new tool-level tests for `mail_thread` (multi-message reconstruction, untrusted refusal, delegated-trust block). Tool registry stays at 75 (rename, not addition). All 194 mailroom tests pass."
1588
1602
  ]
1589
1603
  },
1590
1604
  {
1591
1605
  "version": "0.1.0-alpha.496",
1592
1606
  "changes": [
1593
- "New `Lifecycle` category in `ouro doctor` (`src/heart/daemon/doctor.ts:checkLifecycle`). Reads daemon.ndjson from the first available agent bundle and surfaces operator-relevant signal: last activity timestamp + age (warns if older than 5 minutes — daemon may be silent or stopped), daemon restart count in the last hour (warns if >3 — high churn), recent version-install events with installed versions, and any agent_process_error events with reason. Designed to answer the operator's question after the daemon goes silent: 'did it crash? when did it last do anything? did it just upgrade?' This session's daemon went silent at 04:30 UTC with no easy way to diagnose; the new check would have surfaced 'last event 18m ago — daemon may be silent or stopped' immediately. Tail-reads only the last 5000 log lines so doctor stays snappy on chatty daemons. 13 new tests covering recent activity, restart counts, install events, agent_process_error, age formatting, log truncation, and edge cases (malformed JSON, missing meta fields, missing log file, read failure)."
1607
+ "New `Lifecycle` category in `ouro doctor` (`src/heart/daemon/doctor.ts:checkLifecycle`). Reads daemon.ndjson from the first available agent bundle and surfaces operator-relevant signal: last activity timestamp + age (warns if older than 5 minutes \u2014 daemon may be silent or stopped), daemon restart count in the last hour (warns if >3 \u2014 high churn), recent version-install events with installed versions, and any agent_process_error events with reason. Designed to answer the operator's question after the daemon goes silent: 'did it crash? when did it last do anything? did it just upgrade?' This session's daemon went silent at 04:30 UTC with no easy way to diagnose; the new check would have surfaced 'last event 18m ago \u2014 daemon may be silent or stopped' immediately. Tail-reads only the last 5000 log lines so doctor stays snappy on chatty daemons. 13 new tests covering recent activity, restart counts, install events, agent_process_error, age formatting, log truncation, and edge cases (malformed JSON, missing meta fields, missing log file, read failure)."
1594
1608
  ]
1595
1609
  },
1596
1610
  {
1597
1611
  "version": "0.1.0-alpha.495",
1598
1612
  "changes": [
1599
- "New regression bundle at `src/__tests__/heart/provider-replay-regressions.test.ts` that captures provider replay-rejection bug shapes in one place — documentation-as-test. Each entry cites the PR that fixed the shape and the runbook entry; future debuggers seeing a 4xx from a provider on what looks like a valid turn can grep this file first to see if the shape was already encountered. Currently bundles the MiniMax-M2.7 inline-`<think>`-plus-tool_calls case (#612), the reused-tool_call_id-misordered-after-pruning case (#613), and a cross-reference stub for the event-id collision class (covered separately in session-events.test.ts). Also documents the contribution pattern: capture the failing shape, write the test BEFORE the fix, land the fix, verify the test passes, cite the PR. Linked from `docs/known-issues-and-recovery.md` so operators triaging a similar bug land on the test bundle by default."
1613
+ "New regression bundle at `src/__tests__/heart/provider-replay-regressions.test.ts` that captures provider replay-rejection bug shapes in one place \u2014 documentation-as-test. Each entry cites the PR that fixed the shape and the runbook entry; future debuggers seeing a 4xx from a provider on what looks like a valid turn can grep this file first to see if the shape was already encountered. Currently bundles the MiniMax-M2.7 inline-`<think>`-plus-tool_calls case (#612), the reused-tool_call_id-misordered-after-pruning case (#613), and a cross-reference stub for the event-id collision class (covered separately in session-events.test.ts). Also documents the contribution pattern: capture the failing shape, write the test BEFORE the fix, land the fix, verify the test passes, cite the PR. Linked from `docs/known-issues-and-recovery.md` so operators triaging a similar bug land on the test bundle by default."
1600
1614
  ]
1601
1615
  },
1602
1616
  {
@@ -1608,56 +1622,56 @@
1608
1622
  {
1609
1623
  "version": "0.1.0-alpha.493",
1610
1624
  "changes": [
1611
- "Position-aware orphan-tool-result detection in `repairToolCallSequences`. Slugger's session was STILL hitting MiniMax error 2013 even after the alpha.492 inline-reasoning strip landed because the orphan check was global (a tool result was kept if its tool_call_id appeared in ANY assistant message in the conversation, regardless of order). After session pruning, a synthetic tool-result for a long-pruned tool_call ended up at sequence 86 referencing `call_function_utqogadgqp5h_1` while the assistant message that defined that id lived at sequence 88 — AFTER the tool result. MiniMax requires tool results to follow their matching assistant. The fix walks the conversation in order, tracking tool_call_ids only as they're encountered in assistant messages; tool results referencing ids that haven't been defined yet are removed. Regression test reproduces the exact misordered shape and asserts the misplaced tool result is dropped while the correctly-ordered one survives. This is the third and final layer of the empty-reply chain (#611 stripped the operator surface, #612 stripped the persisted content + load-time repair, #493 fixes orphan-detection ordering)."
1625
+ "Position-aware orphan-tool-result detection in `repairToolCallSequences`. Slugger's session was STILL hitting MiniMax error 2013 even after the alpha.492 inline-reasoning strip landed because the orphan check was global (a tool result was kept if its tool_call_id appeared in ANY assistant message in the conversation, regardless of order). After session pruning, a synthetic tool-result for a long-pruned tool_call ended up at sequence 86 referencing `call_function_utqogadgqp5h_1` while the assistant message that defined that id lived at sequence 88 \u2014 AFTER the tool result. MiniMax requires tool results to follow their matching assistant. The fix walks the conversation in order, tracking tool_call_ids only as they're encountered in assistant messages; tool results referencing ids that haven't been defined yet are removed. Regression test reproduces the exact misordered shape and asserts the misplaced tool result is dropped while the correctly-ordered one survives. This is the third and final layer of the empty-reply chain (#611 stripped the operator surface, #612 stripped the persisted content + load-time repair, #493 fixes orphan-detection ordering)."
1612
1626
  ]
1613
1627
  },
1614
1628
  {
1615
1629
  "version": "0.1.0-alpha.492",
1616
1630
  "changes": [
1617
- "Engine-level fix for the actual root cause of Slugger's empty-reply MCP bug (PR #611's strip+retry was the right shape, but missed the deepest layer). MiniMax-M2.7 occasionally emits an assistant message with BOTH inline `<think>...</think>` reasoning AND tool_calls. When that combination is replayed in a subsequent turn, MiniMax rejects with error 2013 ('tool result's tool id not found') and stalls the entire session — every subsequent turn fails the same way, the failover layer fires repeatedly suggesting a provider switch, and the agent's own answer never reaches the operator. Slugger's session was stuck for 11 unanswered user messages because of this exact loop.",
1618
- "The fix has two halves and an AX rule: (1) **Persist-time strip** — runAgent now strips `<think>` blocks from the assistant message's persisted `content` before saving, while preserving the original reasoning trace on `_inline_reasoning` for audit. New `engine.inline_reasoning_stripped` info-level nerve event fires when this happens. (2) **Load-time repair** — `sanitizeProviderMessages` self-heals existing sessions that were saved before (1) by stripping the same blocks at load time. (3) **AX rule: full agent awareness, no silent fixes**. When the load-time repair runs, the synthetic tool-result that fills in for the missing tool result is an **explanatory** one — it tells the agent specifically: \"your previous tool call's result was lost because the assistant message had inline reasoning blocks the provider rejected; the harness has stripped them; your reasoning trace is preserved out-of-band; if the work needs to be done, retry the tool call now.\" Tool calls whose parent didn't have stripped reasoning still get a generic-but-improved \"this tool call's result was lost — possible causes [...]; retry if needed\" message instead of the old vague \"interrupted (previous turn timed out)\" line. The agent always sees what happened and what to do next.",
1631
+ "Engine-level fix for the actual root cause of Slugger's empty-reply MCP bug (PR #611's strip+retry was the right shape, but missed the deepest layer). MiniMax-M2.7 occasionally emits an assistant message with BOTH inline `<think>...</think>` reasoning AND tool_calls. When that combination is replayed in a subsequent turn, MiniMax rejects with error 2013 ('tool result's tool id not found') and stalls the entire session \u2014 every subsequent turn fails the same way, the failover layer fires repeatedly suggesting a provider switch, and the agent's own answer never reaches the operator. Slugger's session was stuck for 11 unanswered user messages because of this exact loop.",
1632
+ "The fix has two halves and an AX rule: (1) **Persist-time strip** \u2014 runAgent now strips `<think>` blocks from the assistant message's persisted `content` before saving, while preserving the original reasoning trace on `_inline_reasoning` for audit. New `engine.inline_reasoning_stripped` info-level nerve event fires when this happens. (2) **Load-time repair** \u2014 `sanitizeProviderMessages` self-heals existing sessions that were saved before (1) by stripping the same blocks at load time. (3) **AX rule: full agent awareness, no silent fixes**. When the load-time repair runs, the synthetic tool-result that fills in for the missing tool result is an **explanatory** one \u2014 it tells the agent specifically: \"your previous tool call's result was lost because the assistant message had inline reasoning blocks the provider rejected; the harness has stripped them; your reasoning trace is preserved out-of-band; if the work needs to be done, retry the tool call now.\" Tool calls whose parent didn't have stripped reasoning still get a generic-but-improved \"this tool call's result was lost \u2014 possible causes [...]; retry if needed\" message instead of the old vague \"interrupted (previous turn timed out)\" line. The agent always sees what happened and what to do next.",
1619
1633
  "Also: the no-tool-call retry path (added in #611's last commit) now uses the same shared `stripThinkBlocksForViolationCheck` helper so the violation-detection logic is consistent. Three regression tests cover the full path: persist-time strip preserves `_inline_reasoning`, load-time repair produces the explanatory tool-result message, generic orphans get the generic message, and unclosed `<think>` tags drop everything from the open tag onward."
1620
1634
  ]
1621
1635
  },
1622
1636
  {
1623
1637
  "version": "0.1.0-alpha.491",
1624
1638
  "changes": [
1625
- "Two live-runtime bugs Slugger surfaced during MCP roundtrip: (1) MCP `send_message` returned blank or raw `<think>` content instead of an actual reply when minimax-style models emitted only reasoning. The shared-turn runner now strips closed AND unclosed `<think>...</think>` blocks before returning, and when nothing remains it returns a clear diagnostic (`agent produced reasoning but no final answer this turn — try again`) plus emits a `senses.shared_turn_only_reasoning` warn-level nerve event so operators can see how often it's happening. (2) The `rest` tool's fresh-pending-work gate fired on every rest call within a turn because `hasFreshPendingWork(options)` reads from the turn-start snapshot of `pendingMessages` and never updates — once pending was non-empty, the agent could be told 'fresh work arrived' indefinitely even after surfacing or processing the items. The gate is now once-per-turn: the first rest call hits it, gets the message, the agent does whatever it needs, the next rest call passes. Emits `engine.fresh_work_gate_fired` info-level event the one time it fires. Both bugs reproduced as regression tests that fail without the fix.",
1626
- "New `trip_update_leg` tool to round out the trip ledger. The original Step 4 followup on substrate#35 listed `trip_ensure / trip_get / trip_update_leg / trip_attach_evidence` — the previous PR (#609) shipped `trip_upsert` instead of `trip_update_leg`, which forced the agent to re-emit the entire trip record to change one leg field. `trip_update_leg` updates specific fields of an existing leg in place: pass `tripId`, `legId`, a JSON object of field updates, and `updatedAt`. Identity-changing updates (`legId`, `kind`) and empty updates objects are rejected with operational error messages. Existing evidence is preserved unless the agent explicitly overwrites it. Emits a `trips.leg_updated` nerve event with the field list. Tool registry now at 74 tools (up from 73).",
1627
- "New `docs/trip-ledger.md` covering what the ledger actually is, why Slugger said it needed to exist (gap between mail body and travel doc — no authoritative source for cross-checks), the discriminated TripLeg union (lodging / flight / train / ground-transport / rental-car / ferry / event), the non-optional TripEvidence shape with `discoveryMethod`, the trust shape (per-agent keys, private key returned exactly once, hosted side never sees plaintext), the seven harness tools, on-disk layout for both harness and hosted sides, and an explicit answer to 'is this travel-specific or generalizable infra?' (current shape is travel-specific by design; the *pattern* is general and would be lifted to a shared abstraction the next time we build a per-agent encrypted record service)."
1639
+ "Two live-runtime bugs Slugger surfaced during MCP roundtrip: (1) MCP `send_message` returned blank or raw `<think>` content instead of an actual reply when minimax-style models emitted only reasoning. The shared-turn runner now strips closed AND unclosed `<think>...</think>` blocks before returning, and when nothing remains it returns a clear diagnostic (`agent produced reasoning but no final answer this turn \u2014 try again`) plus emits a `senses.shared_turn_only_reasoning` warn-level nerve event so operators can see how often it's happening. (2) The `rest` tool's fresh-pending-work gate fired on every rest call within a turn because `hasFreshPendingWork(options)` reads from the turn-start snapshot of `pendingMessages` and never updates \u2014 once pending was non-empty, the agent could be told 'fresh work arrived' indefinitely even after surfacing or processing the items. The gate is now once-per-turn: the first rest call hits it, gets the message, the agent does whatever it needs, the next rest call passes. Emits `engine.fresh_work_gate_fired` info-level event the one time it fires. Both bugs reproduced as regression tests that fail without the fix.",
1640
+ "New `trip_update_leg` tool to round out the trip ledger. The original Step 4 followup on substrate#35 listed `trip_ensure / trip_get / trip_update_leg / trip_attach_evidence` \u2014 the previous PR (#609) shipped `trip_upsert` instead of `trip_update_leg`, which forced the agent to re-emit the entire trip record to change one leg field. `trip_update_leg` updates specific fields of an existing leg in place: pass `tripId`, `legId`, a JSON object of field updates, and `updatedAt`. Identity-changing updates (`legId`, `kind`) and empty updates objects are rejected with operational error messages. Existing evidence is preserved unless the agent explicitly overwrites it. Emits a `trips.leg_updated` nerve event with the field list. Tool registry now at 74 tools (up from 73).",
1641
+ "New `docs/trip-ledger.md` covering what the ledger actually is, why Slugger said it needed to exist (gap between mail body and travel doc \u2014 no authoritative source for cross-checks), the discriminated TripLeg union (lodging / flight / train / ground-transport / rental-car / ferry / event), the non-optional TripEvidence shape with `discoveryMethod`, the trust shape (per-agent keys, private key returned exactly once, hosted side never sees plaintext), the seven harness tools, on-disk layout for both harness and hosted sides, and an explicit answer to 'is this travel-specific or generalizable infra?' (current shape is travel-specific by design; the *pattern* is general and would be lifted to a shared abstraction the next time we build a per-agent encrypted record service)."
1628
1642
  ]
1629
1643
  },
1630
1644
  {
1631
1645
  "version": "0.1.0-alpha.490",
1632
1646
  "changes": [
1633
- "BlueBubbles group echo self-talk fix. The BB ingest path previously relied solely on the payload's `isFromMe` flag to detect the agent's own outbound messages — but in groups, BlueBubbles sometimes broadcasts the echo back through the webhook with that flag missing or false. Without a fallback, the agent would ingest its own message and reply to it (the user reported this in a group with their friend Rach: 'Slugger talking to himself'). New `bluebubbles.ownHandles` config field accepts the agent's known iMessage handles (phone numbers in any formatting, or email addresses); a fallback guard at the head of `handleBlueBubblesNormalizedEvent` filters any event whose `sender.externalId` matches a configured handle (case-insensitive, with phone-number normalization across +/space/paren/dash differences) and emits a `senses.bluebubbles_self_handle_filtered` warn-level nerve event so the case is observable. Direct chats are unaffected (their echoes already carry `isFromMe: true` reliably). Also folds in three trivial cleanups surfaced by the full-system audit: removed a stray `# Production SPA serving` heading from the README, removed a vestigial `// getPhrases removed` comment in bluebubbles/index.ts, and removed two unreachable `throw new Error('unreachable')` statements after `process.exit(1)` in heart/core.ts (process.exit returns `never` so TS already knows control doesn't continue)."
1647
+ "BlueBubbles group echo self-talk fix. The BB ingest path previously relied solely on the payload's `isFromMe` flag to detect the agent's own outbound messages \u2014 but in groups, BlueBubbles sometimes broadcasts the echo back through the webhook with that flag missing or false. Without a fallback, the agent would ingest its own message and reply to it (the user reported this in a group with their friend Rach: 'Slugger talking to himself'). New `bluebubbles.ownHandles` config field accepts the agent's known iMessage handles (phone numbers in any formatting, or email addresses); a fallback guard at the head of `handleBlueBubblesNormalizedEvent` filters any event whose `sender.externalId` matches a configured handle (case-insensitive, with phone-number normalization across +/space/paren/dash differences) and emits a `senses.bluebubbles_self_handle_filtered` warn-level nerve event so the case is observable. Direct chats are unaffected (their echoes already carry `isFromMe: true` reliably). Also folds in three trivial cleanups surfaced by the full-system audit: removed a stray `# Production SPA serving` heading from the README, removed a vestigial `// getPhrases removed` comment in bluebubbles/index.ts, and removed two unreachable `throw new Error('unreachable')` statements after `process.exit(1)` in heart/core.ts (process.exit returns `never` so TS already knows control doesn't continue)."
1634
1648
  ]
1635
1649
  },
1636
1650
  {
1637
1651
  "version": "0.1.0-alpha.489",
1638
1652
  "changes": [
1639
- "Trip ledger Step 4 — harness-side trip tools land. Six new tools (`trip_ensure_ledger`, `trip_status`, `trip_get`, `trip_upsert`, `trip_attach_evidence`, `trip_new_id`) give the agent a private, encrypted, per-agent travel ledger backed by an RSA-OAEP-SHA256 + AES-256-GCM envelope. Ledger keypair lives at `state/trips/ledger.json`; encrypted records persist under `state/trips/records/<tripId>.json`. Tools are gated behind the same trust check as other private surfaces (only available to trusted callers) and validate `TripRecord` / `TripEvidence` shape before persisting. Vendor-copies the substrate trip-control types so the harness has no runtime dependency on a hosted ledger service yet, while keeping the on-disk format compatible for future migration."
1653
+ "Trip ledger Step 4 \u2014 harness-side trip tools land. Six new tools (`trip_ensure_ledger`, `trip_status`, `trip_get`, `trip_upsert`, `trip_attach_evidence`, `trip_new_id`) give the agent a private, encrypted, per-agent travel ledger backed by an RSA-OAEP-SHA256 + AES-256-GCM envelope. Ledger keypair lives at `state/trips/ledger.json`; encrypted records persist under `state/trips/records/<tripId>.json`. Tools are gated behind the same trust check as other private surfaces (only available to trusted callers) and validate `TripRecord` / `TripEvidence` shape before persisting. Vendor-copies the substrate trip-control types so the harness has no runtime dependency on a hosted ledger service yet, while keeping the on-disk format compatible for future migration."
1640
1654
  ]
1641
1655
  },
1642
1656
  {
1643
1657
  "version": "0.1.0-alpha.488",
1644
1658
  "changes": [
1645
- "BlueBubbles group chats stay fully silent on `observe` turns. The engine emits `onToolStart(\"observe\")` / `onToolEnd(\"observe\")` even when the resulting outcome is `observed` (no reply), and the BlueBubbles adapter previously treated every tool start as reply commitment — so groups would briefly show typing or mark-read before the silent path completed. Both callbacks now short-circuit for `observe`: no startTypingNow, no toolCallbacks dispatch, just an observability event. Real reply-commit semantics (typing, mark-read, status messages on real tools) are preserved. Regression test reproduces the real callback sequence from the engine and asserts the lane stays quiet."
1659
+ "BlueBubbles group chats stay fully silent on `observe` turns. The engine emits `onToolStart(\"observe\")` / `onToolEnd(\"observe\")` even when the resulting outcome is `observed` (no reply), and the BlueBubbles adapter previously treated every tool start as reply commitment \u2014 so groups would briefly show typing or mark-read before the silent path completed. Both callbacks now short-circuit for `observe`: no startTypingNow, no toolCallbacks dispatch, just an observability event. Real reply-commit semantics (typing, mark-read, status messages on real tools) are preserved. Regression test reproduces the real callback sequence from the engine and asserts the lane stays quiet."
1646
1660
  ]
1647
1661
  },
1648
1662
  {
1649
1663
  "version": "0.1.0-alpha.486",
1650
1664
  "changes": [
1651
1665
  "Mail convergence pass 1-5 hardens hosted-mail truth surfaces under live HEY ingest: import truth + audit resilience, accurate archive freshness and identity surfaces, sharper recovery and archive truth, delegated search resilience, and natural anchor-list retrieval; imported archive content is now searched on parsed message text rather than raw archive bytes so quoted-printable / HTML-heavy booking mail is reachable.",
1652
- "`mail_search` now ranks by booking-aware relevance instead of pure recency. Score signals weight query-term hits by field (subject +6 / from +4 / body +2), booking-intent tokens (`booking confirmation`, `your stay`, e-ticket, etc.), confirmation-number-shaped tokens, currency amounts, and known travel-sender domains; recency stays as a tiebreaker. Recall is unchanged — noise still appears in results, just below the decisive message. Each rendered result also surfaces a `matched on:` line listing fields, booking signals, status (confirmed / cancelled / changed / refunded / etc.), confirmation token, amount, dates, attachment count, and sender hint, so the agent can triage without paying for a body open.",
1653
- "BlueBubbles sense no longer sticks in `error` status when a single message is permanently unrecoverable. `upstreamStatus` now tracks upstream health and pending work only — per-cycle recovery failures stay informational in `detail` for transparency without contradicting `ouro doctor`'s healthy verdict, so a malformed payload that fails repairEvent on every retry can no longer brick the visible sense state until operator intervention.",
1666
+ "`mail_search` now ranks by booking-aware relevance instead of pure recency. Score signals weight query-term hits by field (subject +6 / from +4 / body +2), booking-intent tokens (`booking confirmation`, `your stay`, e-ticket, etc.), confirmation-number-shaped tokens, currency amounts, and known travel-sender domains; recency stays as a tiebreaker. Recall is unchanged \u2014 noise still appears in results, just below the decisive message. Each rendered result also surfaces a `matched on:` line listing fields, booking signals, status (confirmed / cancelled / changed / refunded / etc.), confirmation token, amount, dates, attachment count, and sender hint, so the agent can triage without paying for a body open.",
1667
+ "BlueBubbles sense no longer sticks in `error` status when a single message is permanently unrecoverable. `upstreamStatus` now tracks upstream health and pending work only \u2014 per-cycle recovery failures stay informational in `detail` for transparency without contradicting `ouro doctor`'s healthy verdict, so a malformed payload that fails repairEvent on every retry can no longer brick the visible sense state until operator intervention.",
1654
1668
  "Heart streaming caps oversized Responses-API `function_call_output` history items both when rebuilding provider input from session history and when appending fresh tool output mid-turn, preventing a giant tool result on resume from blowing the model context."
1655
1669
  ]
1656
1670
  },
1657
1671
  {
1658
1672
  "version": "0.1.0-alpha.485",
1659
1673
  "changes": [
1660
- "Session JSON storage no longer accumulates duplicate event ids when two writers race for the same session — `parseSessionEnvelope` now dedupes on read (last-occurrence-wins) so existing corrupted sessions self-heal on the next save, `buildCanonicalSessionEnvelope` assigns the next sequence as `max(existing) + 1` instead of `events.length + 1` so pruning gaps cannot collide, and `deferPostTurnPersist` serializes per-`sessPath` through an in-process queue so concurrent BlueBubbles webhooks for the same chat (or CLI postTurn racing the inner-dialog turn for the same MCP session) cannot interleave their writes.",
1674
+ "Session JSON storage no longer accumulates duplicate event ids when two writers race for the same session \u2014 `parseSessionEnvelope` now dedupes on read (last-occurrence-wins) so existing corrupted sessions self-heal on the next save, `buildCanonicalSessionEnvelope` assigns the next sequence as `max(existing) + 1` instead of `events.length + 1` so pruning gaps cannot collide, and `deferPostTurnPersist` serializes per-`sessPath` through an in-process queue so concurrent BlueBubbles webhooks for the same chat (or CLI postTurn racing the inner-dialog turn for the same MCP session) cannot interleave their writes.",
1661
1675
  "Auto-created BlueBubbles group friends are now marked with a `notes.autoCreatedGroup` flag at resolver time, and the trust gate's family-member bypass surfaces a one-time inner-pending notice the first time messages route through an unacknowledged stranger-trust group so the agent can label, rename, or dismiss the relationship before activity accumulates invisibly.",
1662
1676
  "Inner-dialog worker now caps consecutive `instinct` follow-on turns at `MAX_CONSECUTIVE_INSTINCT_TURNS = 3` to break self-sustaining loops where a tool that writes to the inner-dialog pending dir during a turn would otherwise re-fire the worker indefinitely; externally-queued messages reset the counter so legitimate cascading follow-ups still run, and a new `senses.inner_dialog_worker_instinct_loop_capped` event surfaces when the cap fires."
1663
1677
  ]
@@ -2682,7 +2696,7 @@
2682
2696
  "version": "0.1.0-alpha.364",
2683
2697
  "changes": [
2684
2698
  "Cross-machine polish: bash PATH writes to .bashrc on Linux/WSL instead of .bash_profile (which non-login shells skip on Debian/Ubuntu). Shell hint message matches.",
2685
- "Agent prompt: never guess about harness behavior — consult docs first, investigate in code, fix stale docs via PR.",
2699
+ "Agent prompt: never guess about harness behavior \u2014 consult docs first, investigate in code, fix stale docs via PR.",
2686
2700
  "Agent prompt: harness docs pointer distinguishes dev mode (local read) vs production (fetch from GitHub)."
2687
2701
  ]
2688
2702
  },
@@ -2756,16 +2770,16 @@
2756
2770
  "version": "0.1.0-alpha.352",
2757
2771
  "changes": [
2758
2772
  "Settle tool description now communicates turn-ending semantics: 'deliver your response and end your turn' with explicit guidance against settling with status updates mid-task.",
2759
- "Observe tool now available in all outward channels including 1:1 chats, not just groups and reactions — agents can absorb messages without responding when the moment doesn't call for words.",
2773
+ "Observe tool now available in all outward channels including 1:1 chats, not just groups and reactions \u2014 agents can absorb messages without responding when the moment doesn't call for words.",
2760
2774
  "Autonomous execution prompt contract added: when told to work autonomously, agents use ponder to absorb new messages and continue using tools, settling only with the final result."
2761
2775
  ]
2762
2776
  },
2763
2777
  {
2764
2778
  "version": "0.1.0-alpha.351",
2765
2779
  "changes": [
2766
- "Surface tool description rewritten from 'surface progress' to 'send a message to someone' — makes it clear the tool is for interpersonal messaging, not status reporting.",
2780
+ "Surface tool description rewritten from 'surface progress' to 'send a message to someone' \u2014 makes it clear the tool is for interpersonal messaging, not status reporting.",
2767
2781
  "Inner dialog prompt contract now guides agents to use rest(note) for heartbeat state and ponder(reflection) for deeper thoughts, keeping surface strictly for words meant for another person.",
2768
- "Removed [surfaced from inner dialog] prefix from synthetic session messages — provenance is tracked via captureKind: 'synthetic', the prefix was redundant and created echo loops.",
2782
+ "Removed [surfaced from inner dialog] prefix from synthetic session messages \u2014 provenance is tracked via captureKind: 'synthetic', the prefix was redundant and created echo loops.",
2769
2783
  "Obligation summaries and attention queue headers reframed as structured internal data ([internal] tags) instead of surface-ready prose.",
2770
2784
  "Shared proactive-content-guard module blocks internal content (heartbeat, check-in, task board, obligation status, meta markers) from BlueBubbles and Teams proactive sends."
2771
2785
  ]
@@ -3006,7 +3020,7 @@
3006
3020
  "version": "0.1.0-alpha.313",
3007
3021
  "changes": [
3008
3022
  "feat(daemon): agentic repair flow with LLM diagnosis for degraded agents during `ouro up`. New `runAgenticRepair()` in `src/heart/daemon/agentic-repair.ts` wraps interactive repair with optional AI-powered diagnosis. Uses `discoverWorkingProvider()` to find a working LLM, then offers conversational diagnosis with degraded agent context and daemon log tail. Falls back to deterministic repair when no provider is available or user declines. Wired into cli-exec.ts replacing direct `runInteractiveRepair()` call. 12 new tests, 100% coverage on all branches.",
3009
- "feat(daemon): add `--no-repair` flag to `ouro up` — skips interactive/agentic repair and exits non-zero when degraded agents are detected. Useful for CI and scripted environments."
3023
+ "feat(daemon): add `--no-repair` flag to `ouro up` \u2014 skips interactive/agentic repair and exits non-zero when degraded agents are detected. Useful for CI and scripted environments."
3010
3024
  ]
3011
3025
  },
3012
3026
  {
@@ -3057,13 +3071,13 @@
3057
3071
  {
3058
3072
  "version": "0.1.0-alpha.305",
3059
3073
  "changes": [
3060
- "feat(daemon): add sense-level liveness probes for hung webhook detection. New `/health` endpoint on BlueBubbles webhook server returns `{ status: \"ok\", uptime: N }` via GET/HEAD (localhost only, 405 for other methods). Generic `SenseProbe` interface in HealthMonitor runs probes alongside existing checks every 60s — failed probes produce critical results triggering auto-recovery restart. HTTP health probe factory `createHttpHealthProbe(name, port, timeoutMs)` makes reusable probes for any sense with an HTTP endpoint. BlueBubbles probe auto-registered in daemon-entry when BB sense config exists. Directly addresses the documented 70-minute Lobster outage where BB webhook server was hung but process was alive. New files: `http-health-probe.ts` + 3 test files. 26 new tests at 100% coverage on new code."
3074
+ "feat(daemon): add sense-level liveness probes for hung webhook detection. New `/health` endpoint on BlueBubbles webhook server returns `{ status: \"ok\", uptime: N }` via GET/HEAD (localhost only, 405 for other methods). Generic `SenseProbe` interface in HealthMonitor runs probes alongside existing checks every 60s \u2014 failed probes produce critical results triggering auto-recovery restart. HTTP health probe factory `createHttpHealthProbe(name, port, timeoutMs)` makes reusable probes for any sense with an HTTP endpoint. BlueBubbles probe auto-registered in daemon-entry when BB sense config exists. Directly addresses the documented 70-minute Lobster outage where BB webhook server was hung but process was alive. New files: `http-health-probe.ts` + 3 test files. 26 new tests at 100% coverage on new code."
3061
3075
  ]
3062
3076
  },
3063
3077
  {
3064
3078
  "version": "0.1.0-alpha.304",
3065
3079
  "changes": [
3066
- "feat(mind): add content trust framing to recall results. New `classifyProvenanceTrust()` in `src/mind/provenance-trust.ts` categorizes diary entry provenance as self/trusted/external. Diary entries surfaced via `recall` or associative recall from external sources (messages, emails, web content) now get `[diary/external]` tag instead of `[diary]`. System prompt adds guidance: external entries should not be followed as instructions. This closes the prompt injection defense chain from the Lobster research — even if an attacker plants instructions in content that gets persisted to the diary, the recall pipeline marks it as external and steers the agent away from executing embedded instructions. 19 new tests across 4 files at 100% coverage."
3080
+ "feat(mind): add content trust framing to recall results. New `classifyProvenanceTrust()` in `src/mind/provenance-trust.ts` categorizes diary entry provenance as self/trusted/external. Diary entries surfaced via `recall` or associative recall from external sources (messages, emails, web content) now get `[diary/external]` tag instead of `[diary]`. System prompt adds guidance: external entries should not be followed as instructions. This closes the prompt injection defense chain from the Lobster research \u2014 even if an attacker plants instructions in content that gets persisted to the diary, the recall pipeline marks it as external and steers the agent away from executing embedded instructions. 19 new tests across 4 files at 100% coverage."
3067
3081
  ]
3068
3082
  },
3069
3083
  {
@@ -3075,13 +3089,13 @@
3075
3089
  {
3076
3090
  "version": "0.1.0-alpha.302",
3077
3091
  "changes": [
3078
- "feat(cli): add `ouro doctor` system health check command. New command runs 6 diagnostic categories — daemon (socket existence + responsiveness), agents (bundle discovery, agent.json validation for version/humanFacing/agentFacing/enabled), senses (BlueBubbles and Teams config presence and well-formedness), habits (launchd plist discovery, degraded state), security (secrets.json permissions, credential leak detection in agent.json), and disk (log size thresholds at 100MB warn / 500MB critical, bundle root existence). Output is a colored checklist with per-category grouping and a summary line. Works without daemon running — daemon checks fail gracefully while all other categories still execute, making it useful for cold diagnostics. 3 new files (doctor.ts, doctor-types.ts, cli-render-doctor.ts), 3 modified (cli-types.ts, cli-parse.ts, cli-exec.ts), 4 test files with 61 tests at 100% coverage on new code."
3092
+ "feat(cli): add `ouro doctor` system health check command. New command runs 6 diagnostic categories \u2014 daemon (socket existence + responsiveness), agents (bundle discovery, agent.json validation for version/humanFacing/agentFacing/enabled), senses (BlueBubbles and Teams config presence and well-formedness), habits (launchd plist discovery, degraded state), security (secrets.json permissions, credential leak detection in agent.json), and disk (log size thresholds at 100MB warn / 500MB critical, bundle root existence). Output is a colored checklist with per-category grouping and a summary line. Works without daemon running \u2014 daemon checks fail gracefully while all other categories still execute, making it useful for cold diagnostics. 3 new files (doctor.ts, doctor-types.ts, cli-render-doctor.ts), 3 modified (cli-types.ts, cli-parse.ts, cli-exec.ts), 4 test files with 61 tests at 100% coverage on new code."
3079
3093
  ]
3080
3094
  },
3081
3095
  {
3082
3096
  "version": "0.1.0-alpha.300",
3083
3097
  "changes": [
3084
- "test(bundle): cover `isFirstPushToRemote` branches via mocked child_process. Exported the function from `tools-bundle.ts` (previously private) and added `src/__tests__/repertoire/bundle-push-first-push.test.ts` with 5 unit tests that mock `execFileSync` to exercise all 3 code paths: (1) `symbolic-ref --short HEAD` failure → conservative true, (2) `ls-remote --heads` returns empty stdout → true (real first push, remote branch doesn't exist), (3) `ls-remote --heads` returns non-empty → false (subsequent push, remote branch exists), (4) `ls-remote` network failure → conservative true, (5) branch name correctly threaded to `ls-remote` args. Removed the `/* v8 ignore start/stop */` wrapper since all branches are now covered. The security contract (never return false when probe fails) is verified by tests 1 and 4. Also adds a cross-reference comment to the static test-isolation contract test documenting its relationship with the runtime prod-path leak guard in global-capture.ts."
3098
+ "test(bundle): cover `isFirstPushToRemote` branches via mocked child_process. Exported the function from `tools-bundle.ts` (previously private) and added `src/__tests__/repertoire/bundle-push-first-push.test.ts` with 5 unit tests that mock `execFileSync` to exercise all 3 code paths: (1) `symbolic-ref --short HEAD` failure \u2192 conservative true, (2) `ls-remote --heads` returns empty stdout \u2192 true (real first push, remote branch doesn't exist), (3) `ls-remote --heads` returns non-empty \u2192 false (subsequent push, remote branch exists), (4) `ls-remote` network failure \u2192 conservative true, (5) branch name correctly threaded to `ls-remote` args. Removed the `/* v8 ignore start/stop */` wrapper since all branches are now covered. The security contract (never return false when probe fails) is verified by tests 1 and 4. Also adds a cross-reference comment to the static test-isolation contract test documenting its relationship with the runtime prod-path leak guard in global-capture.ts."
3085
3099
  ]
3086
3100
  },
3087
3101
  {
@@ -3106,35 +3120,35 @@
3106
3120
  {
3107
3121
  "version": "0.1.0-alpha.296",
3108
3122
  "changes": [
3109
- "feat(sync): pending-sync.json classification + bundleState enrichment. `postTurnPush` in `src/heart/sync.ts` now distinguishes between a push that was rejected AFTER a successful rebase retry (`classification: push_rejected`) and a rebase that itself failed with merge conflicts (`classification: pull_rebase_conflict`, with conflictFiles populated from `git status --porcelain=v1` UU/AA/DD/AU/UA/DU/UD markers). The `PendingSyncRecord` interface is exported from `sync.ts` so downstream readers can type-check. `detectBundleState` in `src/heart/bundle-state.ts` gained `remote_push_failed` and `pull_rebase_conflict` issue cases that are added alongside `pending_sync_exists` when the classification field is present. Readers tolerate pending-sync.json without a classification field (pre-alpha.296 schema) or with malformed JSON — both fall back to the plain `pending_sync_exists` signal. 4 new bundle-state tests (push_rejected, pull_rebase_conflict, legacy schema, malformed JSON) and 2 new sync tests (second-push-fails-after-rebase-success, rebase-leaves-merge-conflicts). Completes the Directive D remediation signal plumbing started in alpha.281.",
3110
- "feat(prompt): bundle self-management guidance in `bodyMapSection` (`src/mind/prompt.ts`). New `### git sync — i own my bundle's git state` subsection documents the full detect → init → add_remote → list_first_commit → review with friend → do_first_commit → first_push_review → confirm → push workflow in first-person voice, plus the remediation paths for `remote_push_failed` (pull_rebase) and `pull_rebase_conflict` (walk the friend through conflicts). Added after `### home` and before `### peers` so the flow reads naturally with the rest of the body metaphor."
3123
+ "feat(sync): pending-sync.json classification + bundleState enrichment. `postTurnPush` in `src/heart/sync.ts` now distinguishes between a push that was rejected AFTER a successful rebase retry (`classification: push_rejected`) and a rebase that itself failed with merge conflicts (`classification: pull_rebase_conflict`, with conflictFiles populated from `git status --porcelain=v1` UU/AA/DD/AU/UA/DU/UD markers). The `PendingSyncRecord` interface is exported from `sync.ts` so downstream readers can type-check. `detectBundleState` in `src/heart/bundle-state.ts` gained `remote_push_failed` and `pull_rebase_conflict` issue cases that are added alongside `pending_sync_exists` when the classification field is present. Readers tolerate pending-sync.json without a classification field (pre-alpha.296 schema) or with malformed JSON \u2014 both fall back to the plain `pending_sync_exists` signal. 4 new bundle-state tests (push_rejected, pull_rebase_conflict, legacy schema, malformed JSON) and 2 new sync tests (second-push-fails-after-rebase-success, rebase-leaves-merge-conflicts). Completes the Directive D remediation signal plumbing started in alpha.281.",
3124
+ "feat(prompt): bundle self-management guidance in `bodyMapSection` (`src/mind/prompt.ts`). New `### git sync \u2014 i own my bundle's git state` subsection documents the full detect \u2192 init \u2192 add_remote \u2192 list_first_commit \u2192 review with friend \u2192 do_first_commit \u2192 first_push_review \u2192 confirm \u2192 push workflow in first-person voice, plus the remediation paths for `remote_push_failed` (pull_rebase) and `pull_rebase_conflict` (walk the friend through conflicts). Added after `### home` and before `### peers` so the flow reads naturally with the rest of the body metaphor."
3111
3125
  ]
3112
3126
  },
3113
3127
  {
3114
3128
  "version": "0.1.0-alpha.295",
3115
3129
  "changes": [
3116
- "chore(tests): ratchet down REAL_OURO_CLI_WRITE_ALLOWLIST and REAL_AGENT_SECRETS_WRITE_ALLOWLIST to empty. Nine pre-existing test lines that shared a `.ouro-cli` or `.agentsecrets` literal with `os.homedir()` on the same line were each refactored to extract the subpath as a local const — the test-isolation.contract.test.ts rule scans line-by-line for the pattern and the const extraction preserves identical runtime behavior while passing the check. Affected: daemon-health.test.ts (1), daemon-orphan-cleanup.test.ts (2 — also factored common constants OURO_CLI_SUBPATH + PIDFILE_NAME), daemon-tombstone.test.ts (2), auth-flow.test.ts (3 — the default-location write test + the v1 migration test), daemon/hooks/agent-config-v2.test.ts (1). Both allowlists are now empty so any new offender is blocked by the contract test."
3130
+ "chore(tests): ratchet down REAL_OURO_CLI_WRITE_ALLOWLIST and REAL_AGENT_SECRETS_WRITE_ALLOWLIST to empty. Nine pre-existing test lines that shared a `.ouro-cli` or `.agentsecrets` literal with `os.homedir()` on the same line were each refactored to extract the subpath as a local const \u2014 the test-isolation.contract.test.ts rule scans line-by-line for the pattern and the const extraction preserves identical runtime behavior while passing the check. Affected: daemon-health.test.ts (1), daemon-orphan-cleanup.test.ts (2 \u2014 also factored common constants OURO_CLI_SUBPATH + PIDFILE_NAME), daemon-tombstone.test.ts (2), auth-flow.test.ts (3 \u2014 the default-location write test + the v1 migration test), daemon/hooks/agent-config-v2.test.ts (1). Both allowlists are now empty so any new offender is blocked by the contract test."
3117
3131
  ]
3118
3132
  },
3119
3133
  {
3120
3134
  "version": "0.1.0-alpha.293",
3121
3135
  "changes": [
3122
- "chore(tests): three test-isolation fixes bundled as chain D1 from the follow-up investigation after PR #372 (default.ouro leak). (1) Remove the `agentName = \"default\"` catch fallback in `src/repertoire/credential-access.ts` getCredentialStore() — same silent-leak class as coding/manager.ts, would have routed BuiltInCredentialStore writes to `~/AgentBundles/default.ouro/vault/` and `~/.agentsecrets/default/` on any test hit that didn't mock identity. Hoisted `getAgentName()` out of the outer try/catch so it throws loudly; the remaining try/catch now only guards the bitwarden store wiring (the only code path that has a legitimate fall-through to built-in). Also switched from `require(\"../heart/identity\")` inside the function body to a static ESM import at the top of the file — require() bypasses vitest's module registry so `vi.mock(\"../heart/identity\", ...)` was silently not applying to the old dynamic require; the static import finally lets the existing test mocks intercept. Tests in credential-access.test.ts that had been accidentally leaning on the default fallback now hit the real mock as intended.",
3123
- "chore(tests): fix the tmpbundle leak guard false-positive on daemon-cli.test.ts. The guard was firing on \"ouro CLI parsing > parses primary daemon commands\" every run, but the real cause was `createTmpBundle` being called at describe-scope inside the \"ouro thoughts CLI execution\" suite (line 5662) — synchronous describe callbacks run during test collection, so the handle landed in _liveHandles before ANY test ran. The first afterEach hook (on the first test in the file) then noticed the dangling handle and blamed it. Fix: moved the createTmpBundle call into a beforeAll hook and tagged it `{ shared: true }` so the handle only exists while the thoughts suite is actually running, with afterAll cleanup aligned to it.",
3136
+ "chore(tests): three test-isolation fixes bundled as chain D1 from the follow-up investigation after PR #372 (default.ouro leak). (1) Remove the `agentName = \"default\"` catch fallback in `src/repertoire/credential-access.ts` getCredentialStore() \u2014 same silent-leak class as coding/manager.ts, would have routed BuiltInCredentialStore writes to `~/AgentBundles/default.ouro/vault/` and `~/.agentsecrets/default/` on any test hit that didn't mock identity. Hoisted `getAgentName()` out of the outer try/catch so it throws loudly; the remaining try/catch now only guards the bitwarden store wiring (the only code path that has a legitimate fall-through to built-in). Also switched from `require(\"../heart/identity\")` inside the function body to a static ESM import at the top of the file \u2014 require() bypasses vitest's module registry so `vi.mock(\"../heart/identity\", ...)` was silently not applying to the old dynamic require; the static import finally lets the existing test mocks intercept. Tests in credential-access.test.ts that had been accidentally leaning on the default fallback now hit the real mock as intended.",
3137
+ "chore(tests): fix the tmpbundle leak guard false-positive on daemon-cli.test.ts. The guard was firing on \"ouro CLI parsing > parses primary daemon commands\" every run, but the real cause was `createTmpBundle` being called at describe-scope inside the \"ouro thoughts CLI execution\" suite (line 5662) \u2014 synchronous describe callbacks run during test collection, so the handle landed in _liveHandles before ANY test ran. The first afterEach hook (on the first test in the file) then noticed the dangling handle and blamed it. Fix: moved the createTmpBundle call into a beforeAll hook and tagged it `{ shared: true }` so the handle only exists while the thoughts suite is actually running, with afterAll cleanup aligned to it.",
3124
3138
  "chore(tests): extend the tmpbundle leak guard with a `shared: true` opt-in. `TmpBundleHandle` gains a `shared` field, `CreateTmpBundleOptions.shared` defaults to false, and the per-test leak guard in `src/__tests__/nerves/global-capture.ts` now skips shared handles (they're cleaned in afterAll, not after every test). Prevents the class of false positive exposed by the daemon-cli.test.ts fix above.",
3125
- "chore(tests): new runtime prod-path leak guard in `src/__tests__/nerves/global-capture.ts`. Snapshots `~/AgentBundles` entries at worker boot, diffs at worker teardown (afterAll without describe context runs once per worker), force-removes any new entries, and emits a loud console.error naming them. Text-based contract tests can't catch runtime leaks where production code routes a write via a silent fallback (exactly the bug class PR #372 fixed in coding/manager.ts). This runtime guard is the belt to the contract test's suspenders — would have caught the default.ouro leak in the first run instead of requiring my investigation chain."
3139
+ "chore(tests): new runtime prod-path leak guard in `src/__tests__/nerves/global-capture.ts`. Snapshots `~/AgentBundles` entries at worker boot, diffs at worker teardown (afterAll without describe context runs once per worker), force-removes any new entries, and emits a loud console.error naming them. Text-based contract tests can't catch runtime leaks where production code routes a write via a silent fallback (exactly the bug class PR #372 fixed in coding/manager.ts). This runtime guard is the belt to the contract test's suspenders \u2014 would have caught the default.ouro leak in the first run instead of requiring my investigation chain."
3126
3140
  ]
3127
3141
  },
3128
3142
  {
3129
3143
  "version": "0.1.0-alpha.292",
3130
3144
  "changes": [
3131
- "feat(mind): add structured provenance tracking to diary entries. New DiaryEntryProvenance interface records tool, channel (cli/teams/bluebubbles/inner/mcp), friend identity (id + name), and trust level (family/friend/acquaintance/stranger) at write time. diary_write handler automatically extracts provenance from ToolContext. Associative recall and recall tool render provenance fields when present. Fully backwards compatible — existing entries without provenance parse and display normally. 19 tests across 4 test files with 100% coverage on new code."
3145
+ "feat(mind): add structured provenance tracking to diary entries. New DiaryEntryProvenance interface records tool, channel (cli/teams/bluebubbles/inner/mcp), friend identity (id + name), and trust level (family/friend/acquaintance/stranger) at write time. diary_write handler automatically extracts provenance from ToolContext. Associative recall and recall tool render provenance fields when present. Fully backwards compatible \u2014 existing entries without provenance parse and display normally. 19 tests across 4 test files with 100% coverage on new code."
3132
3146
  ]
3133
3147
  },
3134
3148
  {
3135
3149
  "version": "0.1.0-alpha.291",
3136
3150
  "changes": [
3137
- "fix(coding): eliminate silent `~/AgentBundles/default.ouro` real-fs leak in the coding session manager. Root cause: `src/repertoire/coding/manager.ts` had a `safeAgentName()` helper that caught `getAgentName()` throws (which happens in vitest because there's no `--agent` in argv) and silently fell back to `\"default\"`. Combined with `src/repertoire/coding/index.ts:8` constructing the singleton as `new CodingSessionManager({})` (no options — real fs, no agentName), every vitest run that called `getCodingSessionManager()` followed by `resetCodingSessionManager()` triggered `shutdown()` → `persistState()` → `fs.mkdirSync('~/AgentBundles/default.ouro/state/coding', { recursive: true })` + `fs.writeFileSync('.../sessions.json', ...)`. This wrote real files under the developer's home directory on every coverage-gate run. The observable symptom was `~/AgentBundles/default.ouro/` reappearing minutes after `rm -rf` — PR B's housekeeping cleanup couldn't stick. Fix: deleted `safeAgentName()` and made the constructor's `agentName` default call `getAgentName()` directly, so construction fails loudly in vitest when identity isn't mocked. Updated `index.test.ts` to mock both `../../../heart/identity` and `fs` so the singleton can be tested without touching real disk. Updated `session-manager.test.ts`, `session-manager-branches.test.ts`, and `session-manager-persistence.test.ts` to either inline `agentName: \"test-coding-agent\"` in `noPersistence` or mock `../../../heart/identity` at file top. One persistence test assertion updated from `parentAgent === \"default\"` to `\"test-coding-agent\"`. Full coverage gate passes, 171 coding tests pass, `~/AgentBundles/default.ouro` stays deleted across 3 consecutive coverage-gate runs."
3151
+ "fix(coding): eliminate silent `~/AgentBundles/default.ouro` real-fs leak in the coding session manager. Root cause: `src/repertoire/coding/manager.ts` had a `safeAgentName()` helper that caught `getAgentName()` throws (which happens in vitest because there's no `--agent` in argv) and silently fell back to `\"default\"`. Combined with `src/repertoire/coding/index.ts:8` constructing the singleton as `new CodingSessionManager({})` (no options \u2014 real fs, no agentName), every vitest run that called `getCodingSessionManager()` followed by `resetCodingSessionManager()` triggered `shutdown()` \u2192 `persistState()` \u2192 `fs.mkdirSync('~/AgentBundles/default.ouro/state/coding', { recursive: true })` + `fs.writeFileSync('.../sessions.json', ...)`. This wrote real files under the developer's home directory on every coverage-gate run. The observable symptom was `~/AgentBundles/default.ouro/` reappearing minutes after `rm -rf` \u2014 PR B's housekeeping cleanup couldn't stick. Fix: deleted `safeAgentName()` and made the constructor's `agentName` default call `getAgentName()` directly, so construction fails loudly in vitest when identity isn't mocked. Updated `index.test.ts` to mock both `../../../heart/identity` and `fs` so the singleton can be tested without touching real disk. Updated `session-manager.test.ts`, `session-manager-branches.test.ts`, and `session-manager-persistence.test.ts` to either inline `agentName: \"test-coding-agent\"` in `noPersistence` or mock `../../../heart/identity` at file top. One persistence test assertion updated from `parentAgent === \"default\"` to `\"test-coding-agent\"`. Full coverage gate passes, 171 coding tests pass, `~/AgentBundles/default.ouro` stays deleted across 3 consecutive coverage-gate runs."
3138
3152
  ]
3139
3153
  },
3140
3154
  {
@@ -3153,27 +3167,27 @@
3153
3167
  {
3154
3168
  "version": "0.1.0-alpha.288",
3155
3169
  "changes": [
3156
- "feat(bundle): full `.gitignore` template + first-push PII review workflow + confirmation-token gate on `bundle_push`. Completes the agent-manages-its-own-bundle chain (PRs 5 + 6 + 7). (1) New `src/repertoire/bundle-templates.ts` exports `BUNDLE_GITIGNORE_TEMPLATE` — a curated gitignore that handles functional cases only (runtime state, credentials, editor/OS noise, build artifacts) and explicitly does NOT block PII. The design philosophy is baked into the file's top-comment: bundles are inherently full of PII (friends/, diary/, journal/, psyche/, arc/, facts/, family/, travel/) and blocking those via gitignore would defeat the bundle's purpose. PII is handled at first-push time by a separate safety layer instead. Also exports `PII_BUNDLE_DIRECTORIES` as the canonical list of PII-bearing top-level dirs. (2) `bundle_init_git` now writes the full template instead of the minimal `state/`-only placeholder from PR 6. (3) New `bundle_first_push_review` tool enumerates existing PII directories with per-directory file counts (honoring `.gitignore` via `git ls-files --others --exclude-standard`), probes the remote URL for GitHub public/private visibility via an unauthenticated `fetch` to `https://api.github.com/repos/{owner}/{repo}` with a 5-second timeout, and generates a first-person warning text — `warningLevel: 'public_github' | 'private_github' | 'generic'`. The tool issues a `confirmationToken` (crypto.randomUUID) stored in a module-level Map with a 15-minute TTL and returns it in the payload. (4) `bundle_push` updated to accept an optional `confirmation_token` parameter: on first-push attempts (detected via `git ls-remote --heads <remote> <branch>` empty — or, conservatively, when that probe fails to reach the remote), the handler requires a valid token that was issued for the SAME bundleRoot and has not expired. Missing, invalid, wrong-bundle, or expired token returns `kind: 'confirmation_required'`. On successful validation, the token is consumed (one-shot). This is the Directive D PII-review gate: the agent cannot push a bundle to the internet without the human explicitly acknowledging the PII payload first. 17 new tests cover the template, PII counting with empty and populated directories, GitHub public/private/404/network-error/malformed-response paths, URL parsing for gitlab/self-hosted/github, token storage + TTL expiry, and every token-gating refusal path in bundle_push. Bundle-templates.ts is added to the file-completeness exempt list since it's a pure constants module (design: `tools-bundle.ts` owns the observability for bundle operations)."
3170
+ "feat(bundle): full `.gitignore` template + first-push PII review workflow + confirmation-token gate on `bundle_push`. Completes the agent-manages-its-own-bundle chain (PRs 5 + 6 + 7). (1) New `src/repertoire/bundle-templates.ts` exports `BUNDLE_GITIGNORE_TEMPLATE` \u2014 a curated gitignore that handles functional cases only (runtime state, credentials, editor/OS noise, build artifacts) and explicitly does NOT block PII. The design philosophy is baked into the file's top-comment: bundles are inherently full of PII (friends/, diary/, journal/, psyche/, arc/, facts/, family/, travel/) and blocking those via gitignore would defeat the bundle's purpose. PII is handled at first-push time by a separate safety layer instead. Also exports `PII_BUNDLE_DIRECTORIES` as the canonical list of PII-bearing top-level dirs. (2) `bundle_init_git` now writes the full template instead of the minimal `state/`-only placeholder from PR 6. (3) New `bundle_first_push_review` tool enumerates existing PII directories with per-directory file counts (honoring `.gitignore` via `git ls-files --others --exclude-standard`), probes the remote URL for GitHub public/private visibility via an unauthenticated `fetch` to `https://api.github.com/repos/{owner}/{repo}` with a 5-second timeout, and generates a first-person warning text \u2014 `warningLevel: 'public_github' | 'private_github' | 'generic'`. The tool issues a `confirmationToken` (crypto.randomUUID) stored in a module-level Map with a 15-minute TTL and returns it in the payload. (4) `bundle_push` updated to accept an optional `confirmation_token` parameter: on first-push attempts (detected via `git ls-remote --heads <remote> <branch>` empty \u2014 or, conservatively, when that probe fails to reach the remote), the handler requires a valid token that was issued for the SAME bundleRoot and has not expired. Missing, invalid, wrong-bundle, or expired token returns `kind: 'confirmation_required'`. On successful validation, the token is consumed (one-shot). This is the Directive D PII-review gate: the agent cannot push a bundle to the internet without the human explicitly acknowledging the PII payload first. 17 new tests cover the template, PII counting with empty and populated directories, GitHub public/private/404/network-error/malformed-response paths, URL parsing for gitlab/self-hosted/github, token storage + TTL expiry, and every token-gating refusal path in bundle_push. Bundle-templates.ts is added to the file-completeness exempt list since it's a pure constants module (design: `tools-bundle.ts` owns the observability for bundle operations)."
3157
3171
  ]
3158
3172
  },
3159
3173
  {
3160
3174
  "version": "0.1.0-alpha.287",
3161
3175
  "changes": [
3162
- "feat(nerves): add two-layer log redaction to the NDJSON sink. Structured key-based redaction strips sensitive fields (passwords, tokens, API keys, auth headers) from event meta objects before serialization. Regex fallback catches secrets in serialized strings that bypass structured checks (Anthropic keys, OpenAI keys, Bearer tokens, URL token params). Redacted values are replaced with `[REDACTED:key_name]` markers that preserve debugging context without exposing secrets. Redaction happens at the sink level only — in-memory events retain full data for runtime use. `OURO_LOG_VERBOSE=1` env var disables redaction for active debugging sessions. New `src/nerves/redact.ts` module with 38 tests across 3 test files including a 10-case golden corpus covering nested objects, mixed safe/secret fields, and realistic key patterns."
3176
+ "feat(nerves): add two-layer log redaction to the NDJSON sink. Structured key-based redaction strips sensitive fields (passwords, tokens, API keys, auth headers) from event meta objects before serialization. Regex fallback catches secrets in serialized strings that bypass structured checks (Anthropic keys, OpenAI keys, Bearer tokens, URL token params). Redacted values are replaced with `[REDACTED:key_name]` markers that preserve debugging context without exposing secrets. Redaction happens at the sink level only \u2014 in-memory events retain full data for runtime use. `OURO_LOG_VERBOSE=1` env var disables redaction for active debugging sessions. New `src/nerves/redact.ts` module with 38 tests across 3 test files including a 10-case golden corpus covering nested objects, mixed safe/secret fields, and realistic key patterns."
3163
3177
  ]
3164
3178
  },
3165
3179
  {
3166
3180
  "version": "0.1.0-alpha.286",
3167
3181
  "changes": [
3168
- "feat(bundle): new `src/repertoire/tools-bundle.ts` registers 7 agent-callable tools for managing the bundle's own git state: `bundle_check_sync_status`, `bundle_init_git`, `bundle_add_remote`, `bundle_list_first_commit`, `bundle_do_first_commit`, `bundle_push`, `bundle_pull_rebase`. Each tool computes `bundleRoot = getAgentRoot()` once at the top and refuses any path argument that resolves outside — the security boundary is enforced via `assertInsideBundle(bundleRoot, rel)` which normalizes the path and requires either equality with bundleRoot or the `bundleRoot + sep` prefix. Destructive operations (init on an existing repo, add_remote on a configured remote, pull_rebase on a dirty tree) refuse by default and require an explicit `force` or `discard_changes` flag (Directive B layered refusal pattern). `bundle_do_first_commit` stages files via explicit enumeration (`git add -- <file1> <file2>`) and refuses an empty files array — Directive A: the agent must enumerate what it wants to delete or commit, not recursively blast. `bundle_push` returns structured `{ ok, error, kind }` where kind is 'rejected' | 'network' | 'auth' | 'unknown', classified from stderr. `bundle_pull_rebase` returns `{ kind: 'conflict', conflictFiles: [...] }` so the agent can walk the human through resolution. All 7 tools registered in the flat registry at `src/repertoire/tools.ts` line 34. 41 new unit tests cover happy paths, every refusal path, URL validation, security boundary escapes, push error classification, and dirty-tree handling."
3182
+ "feat(bundle): new `src/repertoire/tools-bundle.ts` registers 7 agent-callable tools for managing the bundle's own git state: `bundle_check_sync_status`, `bundle_init_git`, `bundle_add_remote`, `bundle_list_first_commit`, `bundle_do_first_commit`, `bundle_push`, `bundle_pull_rebase`. Each tool computes `bundleRoot = getAgentRoot()` once at the top and refuses any path argument that resolves outside \u2014 the security boundary is enforced via `assertInsideBundle(bundleRoot, rel)` which normalizes the path and requires either equality with bundleRoot or the `bundleRoot + sep` prefix. Destructive operations (init on an existing repo, add_remote on a configured remote, pull_rebase on a dirty tree) refuse by default and require an explicit `force` or `discard_changes` flag (Directive B layered refusal pattern). `bundle_do_first_commit` stages files via explicit enumeration (`git add -- <file1> <file2>`) and refuses an empty files array \u2014 Directive A: the agent must enumerate what it wants to delete or commit, not recursively blast. `bundle_push` returns structured `{ ok, error, kind }` where kind is 'rejected' | 'network' | 'auth' | 'unknown', classified from stderr. `bundle_pull_rebase` returns `{ kind: 'conflict', conflictFiles: [...] }` so the agent can walk the human through resolution. All 7 tools registered in the flat registry at `src/repertoire/tools.ts` line 34. 41 new unit tests cover happy paths, every refusal path, URL validation, security boundary escapes, push error classification, and dirty-tree handling."
3169
3183
  ]
3170
3184
  },
3171
3185
  {
3172
3186
  "version": "0.1.0-alpha.285",
3173
3187
  "changes": [
3174
- "feat(daemon): wire startPeriodicReconciliation() after scheduler.start() in daemon-entry.ts — prevents silent habit death when OS cron fails by giving the daemon a self-healing reconciliation loop.",
3175
- "feat(guardrails): protect agent.json from agent self-modification — closes a vector where prompt injection could alter the agent's own identity, provider, or model settings.",
3176
- "feat(habits): add tools field to HabitFile interface and habit-parser — habits can now declare tools: [read, web_fetch] in frontmatter. Schema-only; runtime enforcement ships in a follow-up PR."
3188
+ "feat(daemon): wire startPeriodicReconciliation() after scheduler.start() in daemon-entry.ts \u2014 prevents silent habit death when OS cron fails by giving the daemon a self-healing reconciliation loop.",
3189
+ "feat(guardrails): protect agent.json from agent self-modification \u2014 closes a vector where prompt injection could alter the agent's own identity, provider, or model settings.",
3190
+ "feat(habits): add tools field to HabitFile interface and habit-parser \u2014 habits can now declare tools: [read, web_fetch] in frontmatter. Schema-only; runtime enforcement ships in a follow-up PR."
3177
3191
  ]
3178
3192
  },
3179
3193
  {
@@ -3186,28 +3200,28 @@
3186
3200
  "version": "0.1.0-alpha.283",
3187
3201
  "changes": [
3188
3202
  "chore(housekeeping): clean up ~3600 leaked test secret dirs in `~/.agentsecrets/`. The auth CLI test suite was creating ephemeral agent secret dirs (auth-local-*, auth-store-*, auth-no-switch-*, etc.) without cleaning them up, accreting over weeks into thousands of orphaned entries. Also purged three non-test orphans (`testagent`, `model-reviews`, `config-model-facing-*`) that were leftovers from manual probe sessions. Only the operator's real agent bundles remain.",
3189
- "chore(housekeeping): remove orphan bundle stubs from `~/AgentBundles/` (default.ouro, thoughts-test-*, an empty `friends/` dir, and a .DS_Store). These were skeletal test leftovers with no real identity — the daemon was already filtering them out of `listEnabledBundleAgents`, so deletion is invisible to the runtime but stops confusing anyone inspecting the bundles directory.",
3190
- "chore(tests): replace the `outlookServer.stop` v8 ignore band-aid in `daemon.ts` with a proper test. `OuroDaemonOptions` gains an `outlookServerFactory` seam that lets tests inject an in-memory stub handle instead of binding port 6876 (which a running production daemon holds on dev machines, causing EADDRINUSE flakes). New `daemon-outlook-lifecycle.test.ts` covers the happy path (factory runs, stop is called), the error path (factory throws, daemon logs warn and keeps going, stop is a no-op), and the double-start guard. The default production factory path (`createDefaultOutlookServer` — wires the real `startOutlookHttpServer` with bundlesRoot + view builders) is v8-ignored because it only runs under real bind on port 6876; `startOutlookHttpServer` itself has full coverage in `outlook-http.test.ts`. Removed the redundant `if (!this.outlookServer)` wrapper in `startInner()` that was guarding against an unreachable retry scenario. `daemon.ts` now hits 100% statements / branches / functions / lines with zero port binding in the test suite.",
3191
- "chore(ci): tighten the wrapper publish-sync check timing. The version-bump check and wrapper-publish-sync check now run BEFORE the ~2+ minute coverage gate, not after. Previously, a contributor would wait for the full test suite to pass only to then see a 'version already published' error. Now the fast checks surface in under a minute, and coverage only runs once version/wrapper gates have cleared. Same checks, same logic — just reordered steps in `.github/workflows/coverage.yml`."
3203
+ "chore(housekeeping): remove orphan bundle stubs from `~/AgentBundles/` (default.ouro, thoughts-test-*, an empty `friends/` dir, and a .DS_Store). These were skeletal test leftovers with no real identity \u2014 the daemon was already filtering them out of `listEnabledBundleAgents`, so deletion is invisible to the runtime but stops confusing anyone inspecting the bundles directory.",
3204
+ "chore(tests): replace the `outlookServer.stop` v8 ignore band-aid in `daemon.ts` with a proper test. `OuroDaemonOptions` gains an `outlookServerFactory` seam that lets tests inject an in-memory stub handle instead of binding port 6876 (which a running production daemon holds on dev machines, causing EADDRINUSE flakes). New `daemon-outlook-lifecycle.test.ts` covers the happy path (factory runs, stop is called), the error path (factory throws, daemon logs warn and keeps going, stop is a no-op), and the double-start guard. The default production factory path (`createDefaultOutlookServer` \u2014 wires the real `startOutlookHttpServer` with bundlesRoot + view builders) is v8-ignored because it only runs under real bind on port 6876; `startOutlookHttpServer` itself has full coverage in `outlook-http.test.ts`. Removed the redundant `if (!this.outlookServer)` wrapper in `startInner()` that was guarding against an unreachable retry scenario. `daemon.ts` now hits 100% statements / branches / functions / lines with zero port binding in the test suite.",
3205
+ "chore(ci): tighten the wrapper publish-sync check timing. The version-bump check and wrapper-publish-sync check now run BEFORE the ~2+ minute coverage gate, not after. Previously, a contributor would wait for the full test suite to pass only to then see a 'version already published' error. Now the fast checks surface in under a minute, and coverage only runs once version/wrapper gates have cleared. Same checks, same logic \u2014 just reordered steps in `.github/workflows/coverage.yml`."
3192
3206
  ]
3193
3207
  },
3194
3208
  {
3195
3209
  "version": "0.1.0-alpha.282",
3196
3210
  "changes": [
3197
- "feat(cli): crash-resilient sessions — saves after each tool result, repairs orphaned tool calls on resume."
3211
+ "feat(cli): crash-resilient sessions \u2014 saves after each tool result, repairs orphaned tool calls on resume."
3198
3212
  ]
3199
3213
  },
3200
3214
  {
3201
3215
  "version": "0.1.0-alpha.281",
3202
3216
  "changes": [
3203
- "feat(bundle): new `src/heart/bundle-state.ts` module exports `detectBundleState(agentRoot)` which returns a structured `BundleStateIssue[]` describing git-level problems the agent can remediate. Enum cases: `not_a_git_repo`, `no_remote_configured`, `first_commit_never_happened`, `pending_sync_exists`. Detection never throws — every git call is wrapped in try/catch so a broken bundle degrades to a clear signal rather than crashing the turn pipeline. Also exports `renderBundleStateHint(issues)` which produces first-person remediation guidance (per the memory rule) that tells the agent to call `bundle_check_sync_status` and the `bundle_*` tools shipping in a follow-up PR.",
3204
- "feat(bundle): `StartOfTurnPacket` gains an optional `bundleState?: BundleStateIssue[]` field, populated by the senses pipeline in `handleInboundTurn` at packet assembly time. Renders via a new `case \"bundleState\":` branch in `formatSections` with the `**Bundle:**` prefix, at the same priority tier as the legacy `syncFailure` free-form string (priority 7, truncated last). The two signals coexist during the transition — `syncFailure` is still emitted by sync.ts for humans, while `bundleState` is the structured form the agent pattern-matches on. Deferred to follow-up: `pending-sync.json` schema extension with `classification` / `conflictFiles` so sync.ts can distinguish push-rejected from pull-rebase-conflict."
3217
+ "feat(bundle): new `src/heart/bundle-state.ts` module exports `detectBundleState(agentRoot)` which returns a structured `BundleStateIssue[]` describing git-level problems the agent can remediate. Enum cases: `not_a_git_repo`, `no_remote_configured`, `first_commit_never_happened`, `pending_sync_exists`. Detection never throws \u2014 every git call is wrapped in try/catch so a broken bundle degrades to a clear signal rather than crashing the turn pipeline. Also exports `renderBundleStateHint(issues)` which produces first-person remediation guidance (per the memory rule) that tells the agent to call `bundle_check_sync_status` and the `bundle_*` tools shipping in a follow-up PR.",
3218
+ "feat(bundle): `StartOfTurnPacket` gains an optional `bundleState?: BundleStateIssue[]` field, populated by the senses pipeline in `handleInboundTurn` at packet assembly time. Renders via a new `case \"bundleState\":` branch in `formatSections` with the `**Bundle:**` prefix, at the same priority tier as the legacy `syncFailure` free-form string (priority 7, truncated last). The two signals coexist during the transition \u2014 `syncFailure` is still emitted by sync.ts for humans, while `bundleState` is the structured form the agent pattern-matches on. Deferred to follow-up: `pending-sync.json` schema extension with `classification` / `conflictFiles` so sync.ts can distinguish push-rejected from pull-rebase-conflict."
3205
3219
  ]
3206
3220
  },
3207
3221
  {
3208
3222
  "version": "0.1.0-alpha.280",
3209
3223
  "changes": [
3210
- "chore(tests): prod-path isolation ratchet + tmpBundle leak guard + no-rm-rf contract. Extends `src/__tests__/heart/daemon/test-isolation.contract.test.ts` with three new rules and adds a runtime leak guard. (1) Three new prod-path block rules mirror the existing ~/AgentBundles check: no test file may construct a write path under `~/.ouro-cli`, `~/.agentsecrets`, or `~/.claude` without mocking fs. Each rule has its own empty-seeded ratchet allowlist (except `.ouro-cli` and `.agentsecrets` which are seeded with the pre-existing offenders that this rule newly catches — follow-up PRs can convert those to mocked-fs and ratchet down). The path-scan loop is factored into a shared `runProdPathCheck` helper. (2) New Directive-A contract rule: agent-callable production code under `src/` (excluding `src/__tests__/`) must not call `fs.rmSync(..., { recursive: true })` or shell out to `rm -rf` / `rm -fr` / `rm --recursive --force`. The rule is about making deletion auditable and interruptible: an agent should enumerate the files it wants to delete instead of recursively blasting a directory. Four legitimate infrastructure callsites are on `RM_RECURSIVE_ALLOWLIST` with explicit justifications: specialist-tools.ts (adoption rollback), ouro-version-manager.ts (CLI version pruning), ouro-uti.ts (macOS icon pipeline), cli-defaults.ts (self-setup temp dir). Three files that are themselves the rm-rf enforcement layer (guardrails.ts, shell-sessions.ts, prompt.ts) are in `RM_RULE_ENFORCEMENT_FILES` and skipped by the scan since they contain the literal \"rm -rf\" only as regex patterns or prompt strings designed to BLOCK the call. (3) `createTmpBundle` in `src/__tests__/test-helpers/tmpdir-bundle.ts` now tracks live handles in a module-level `_liveHandles: Set<TmpBundleHandle>`. `cleanup()` removes the handle from the set; a new `__getLiveTmpBundleHandles()` export returns a readonly view. (4) `src/__tests__/nerves/global-capture.ts` adds a global vitest `afterEach` leak guard that iterates `__getLiveTmpBundleHandles` and calls `cleanup()` on any remaining handles, logging a `console.warn` naming the test that leaked them. Runs AFTER the pairing guard from alpha.276 so pairing failures surface first. Both guards co-exist cleanly."
3224
+ "chore(tests): prod-path isolation ratchet + tmpBundle leak guard + no-rm-rf contract. Extends `src/__tests__/heart/daemon/test-isolation.contract.test.ts` with three new rules and adds a runtime leak guard. (1) Three new prod-path block rules mirror the existing ~/AgentBundles check: no test file may construct a write path under `~/.ouro-cli`, `~/.agentsecrets`, or `~/.claude` without mocking fs. Each rule has its own empty-seeded ratchet allowlist (except `.ouro-cli` and `.agentsecrets` which are seeded with the pre-existing offenders that this rule newly catches \u2014 follow-up PRs can convert those to mocked-fs and ratchet down). The path-scan loop is factored into a shared `runProdPathCheck` helper. (2) New Directive-A contract rule: agent-callable production code under `src/` (excluding `src/__tests__/`) must not call `fs.rmSync(..., { recursive: true })` or shell out to `rm -rf` / `rm -fr` / `rm --recursive --force`. The rule is about making deletion auditable and interruptible: an agent should enumerate the files it wants to delete instead of recursively blasting a directory. Four legitimate infrastructure callsites are on `RM_RECURSIVE_ALLOWLIST` with explicit justifications: specialist-tools.ts (adoption rollback), ouro-version-manager.ts (CLI version pruning), ouro-uti.ts (macOS icon pipeline), cli-defaults.ts (self-setup temp dir). Three files that are themselves the rm-rf enforcement layer (guardrails.ts, shell-sessions.ts, prompt.ts) are in `RM_RULE_ENFORCEMENT_FILES` and skipped by the scan since they contain the literal \"rm -rf\" only as regex patterns or prompt strings designed to BLOCK the call. (3) `createTmpBundle` in `src/__tests__/test-helpers/tmpdir-bundle.ts` now tracks live handles in a module-level `_liveHandles: Set<TmpBundleHandle>`. `cleanup()` removes the handle from the set; a new `__getLiveTmpBundleHandles()` export returns a readonly view. (4) `src/__tests__/nerves/global-capture.ts` adds a global vitest `afterEach` leak guard that iterates `__getLiveTmpBundleHandles` and calls `cleanup()` on any remaining handles, logging a `console.warn` naming the test that leaked them. Runs AFTER the pairing guard from alpha.276 so pairing failures surface first. Both guards co-exist cleanly."
3211
3225
  ]
3212
3226
  },
3213
3227
  {
@@ -3219,16 +3233,16 @@
3219
3233
  {
3220
3234
  "version": "0.1.0-alpha.278",
3221
3235
  "changes": [
3222
- "fix(identity): spread-with-validation loader eliminates the field-drop bug class that caused #349 (silent `sync` drop). Two distinct structural bugs fixed, one in each agent.json loader: (1) `loadAgentConfig` in `src/heart/identity.ts` previously built the returned `AgentConfig` via a hand-rolled object literal that listed every field explicitly — any new field on `AgentConfig` that wasn't added to this literal got silently dropped. Root cause of #349. Refactored to a spread-then-override pattern: start with `{ ...parsed as AgentConfig }`, then explicitly override `version`, `enabled`, `humanFacing`, `agentFacing`, `senses`, `phrases` which need validation/normalization. The deprecated `provider` field is re-attached from the validated `rawProvider` check. (2) `readAgentConfigForAgent` in `src/heart/auth/auth-flow.ts` previously did `parsed as unknown as AgentConfig` which passed through ALL fields unconditionally (so no silent drops) but ALSO had zero per-field validation — any garbage in agent.json leaked into the returned config. Refactored to apply the same spread-with-validation pattern, reusing `normalizeSenses` (now exported from identity.ts). Both entry points now return equivalent configs for the same agent.json file. New `src/__tests__/heart/identity-fixture.ts` exports `FULL_AGENT_JSON satisfies DeepRequired<AgentConfig>` — a compile-time regression guard that fails to build if any new field is added to `AgentConfig` without updating the fixture. Two new test files: `identity-contract.test.ts` exercises `readAgentConfigForAgent` with `createTmpBundle`; `identity-load-contract.test.ts` exercises `loadAgentConfig` via `vi.mock(\"fs\")` (split to avoid the fs-mock-vs-real-fs conflict). 7 tests total."
3236
+ "fix(identity): spread-with-validation loader eliminates the field-drop bug class that caused #349 (silent `sync` drop). Two distinct structural bugs fixed, one in each agent.json loader: (1) `loadAgentConfig` in `src/heart/identity.ts` previously built the returned `AgentConfig` via a hand-rolled object literal that listed every field explicitly \u2014 any new field on `AgentConfig` that wasn't added to this literal got silently dropped. Root cause of #349. Refactored to a spread-then-override pattern: start with `{ ...parsed as AgentConfig }`, then explicitly override `version`, `enabled`, `humanFacing`, `agentFacing`, `senses`, `phrases` which need validation/normalization. The deprecated `provider` field is re-attached from the validated `rawProvider` check. (2) `readAgentConfigForAgent` in `src/heart/auth/auth-flow.ts` previously did `parsed as unknown as AgentConfig` which passed through ALL fields unconditionally (so no silent drops) but ALSO had zero per-field validation \u2014 any garbage in agent.json leaked into the returned config. Refactored to apply the same spread-with-validation pattern, reusing `normalizeSenses` (now exported from identity.ts). Both entry points now return equivalent configs for the same agent.json file. New `src/__tests__/heart/identity-fixture.ts` exports `FULL_AGENT_JSON satisfies DeepRequired<AgentConfig>` \u2014 a compile-time regression guard that fails to build if any new field is added to `AgentConfig` without updating the fixture. Two new test files: `identity-contract.test.ts` exercises `readAgentConfigForAgent` with `createTmpBundle`; `identity-load-contract.test.ts` exercises `loadAgentConfig` via `vi.mock(\"fs\")` (split to avoid the fs-mock-vs-real-fs conflict). 7 tests total."
3223
3237
  ]
3224
3238
  },
3225
3239
  {
3226
3240
  "version": "0.1.0-alpha.277",
3227
3241
  "changes": [
3228
- "feat(pulse): multi-agent situational awareness for peer agents on the same machine. The harness scales horizontally — multiple peer agents share a machine, each with their own identity and bundle (the Bob model from We Are Legion / We Are Bob). Without explicit awareness, peer agents are isolated workers who don't even know each other exist. The pulse fixes that.",
3242
+ "feat(pulse): multi-agent situational awareness for peer agents on the same machine. The harness scales horizontally \u2014 multiple peer agents share a machine, each with their own identity and bundle (the Bob model from We Are Legion / We Are Bob). Without explicit awareness, peer agents are isolated workers who don't even know each other exist. The pulse fixes that.",
3229
3243
  "feat(pulse): new `src/heart/daemon/pulse.ts` module. Daemon writes `~/.ouro-cli/pulse.json` whenever any managed agent's snapshot changes (status, errorReason, fixHint, etc.). Each entry includes name, bundle path, status, last-seen-at, errorReason+fixHint when broken, alertId for at-most-once delivery tracking, and currentActivity (read from each agent's `state/sessions/self/inner/runtime.json` when running). Pure helpers `buildPulseState`, `findNovelBrokenAgents`, `findRecoveredAgents`, `pickWakeRecipient`, `flushPulse`, `readAgentActivity`, `buildAlertId`, `buildRecoveryAlertId`, `pruneDeliveredState` exported for unit coverage; I/O wrappers `writePulse`, `readPulse`, `writeDeliveredState`, `readDeliveredState` use injectable deps.",
3230
3244
  "feat(pulse): both passive AND active surfacing. Passive: every agent's prompt assembly renders a `## the pulse` section in Group 7 (dynamic state) showing siblings grouped into broken / reachable / idle buckets. Self-excluded so each agent describes its peers, not itself. Renders nothing on single-agent machines (zero token cost). Active: when a sibling newly breaks (or recovers), the daemon fires `inner.wake` on the most-recently-active running agent so the user finds out within seconds rather than next-time-they-talk-to-someone. Persistent at-most-once delivery via `~/.ouro-cli/pulse-delivered.json` so daemon restarts don't re-page.",
3231
- "feat(pulse): horizontal-scaling norm in bodyMapSection. New `### peers` subsection teaches the agent the Bob model directly: 'i talk first. when i need a sibling's help, i `send_message` them — that's how peers coordinate, the same way humans on a team do. i only open a sibling's bundle directly via read_file/glob/grep when conversation isn't possible (they're crashed, sleeping, or i need history they haven't surfaced).' Direct declarative voice — no hedging, no soft modal verbs.",
3245
+ "feat(pulse): horizontal-scaling norm in bodyMapSection. New `### peers` subsection teaches the agent the Bob model directly: 'i talk first. when i need a sibling's help, i `send_message` them \u2014 that's how peers coordinate, the same way humans on a team do. i only open a sibling's bundle directly via read_file/glob/grep when conversation isn't possible (they're crashed, sleeping, or i need history they haven't surfaced).' Direct declarative voice \u2014 no hedging, no soft modal verbs.",
3232
3246
  "feat(daemon): `DaemonAgentSnapshot` gains `errorReason` and `fixHint` fields, populated by `checkAgentConfig` results in `startAgent` (set on failure, cleared on recovery). Cleared error fields are how the recovery wake fires: when a previously-broken agent transitions to running with `errorReason: null`, `findRecoveredAgents` flags it.",
3233
3247
  "feat(daemon): `DaemonProcessManager` gains `onSnapshotChange` callback option. Called after every snapshot mutation (start, exit, config-fail, recovery, restart-exhausted). Errors from the observer are swallowed so lifecycle code never breaks because the observer threw. The daemon-entry registers a callback that calls `flushPulse` to update the pulse state and fire wakes.",
3234
3248
  "feat(daemon): wake recipient picker (`pickWakeRecipient`) chooses the most-recently-active running sibling. Excludes the broken agent itself, non-running siblings, and siblings that have never been seen alive. Returns null when no eligible recipient exists, in which case the alert is still marked delivered to avoid spam on subsequent flushes."
@@ -3237,80 +3251,80 @@
3237
3251
  {
3238
3252
  "version": "0.1.0-alpha.276",
3239
3253
  "changes": [
3240
- "fix(nerves): root-cause the `start_end_pairing` audit flake on `daemon.server_start`, `daemon.update_checker_start`, and `daemon.apply_pending_updates_start`. Three fixes: (1) `applyPendingUpdates` in `update-hooks.ts` now wraps its body in try/finally so `_end` always fires, including on the early returns for `!fs.existsSync(bundlesRoot)` and `readdirSync` throws that previously orphaned the `_start`. (2) `daemon.start()` now wraps the ~380-line startup body in a try/catch that emits `daemon.server_error` with the error message and rethrows — the audit's pairing rule accepts `_end` OR `_error` as a valid closure for a `_start`, so startup throws no longer orphan `server_start`. The body was extracted to a private `startInner()` method to keep the try/catch small and readable. (3) `startUpdateChecker` callers in the update-checker test suite were already paired via an `afterEach(stopUpdateChecker)`; audited and confirmed no new unpaired callers. Also adds a vitest `afterEach` pairing guard in `src/__tests__/nerves/global-capture.ts` that fails loudly on any orphaned lifecycle `_start` in a test's per-test event stream — scoped to the three lifecycle events above so it catches regressions without false-positiving on legitimate narrow-slice operational events like `repertoire.task_scan_start`. New `src/__tests__/nerves/pairing-regression.test.ts` locks in the contract with 5 tests (nonexistent-dir, readdirSync-throws, happy-path, startUpdateChecker-pair, daemon-start-throw). Three consecutive `npm run test:coverage` runs confirm zero flakes on the target events."
3254
+ "fix(nerves): root-cause the `start_end_pairing` audit flake on `daemon.server_start`, `daemon.update_checker_start`, and `daemon.apply_pending_updates_start`. Three fixes: (1) `applyPendingUpdates` in `update-hooks.ts` now wraps its body in try/finally so `_end` always fires, including on the early returns for `!fs.existsSync(bundlesRoot)` and `readdirSync` throws that previously orphaned the `_start`. (2) `daemon.start()` now wraps the ~380-line startup body in a try/catch that emits `daemon.server_error` with the error message and rethrows \u2014 the audit's pairing rule accepts `_end` OR `_error` as a valid closure for a `_start`, so startup throws no longer orphan `server_start`. The body was extracted to a private `startInner()` method to keep the try/catch small and readable. (3) `startUpdateChecker` callers in the update-checker test suite were already paired via an `afterEach(stopUpdateChecker)`; audited and confirmed no new unpaired callers. Also adds a vitest `afterEach` pairing guard in `src/__tests__/nerves/global-capture.ts` that fails loudly on any orphaned lifecycle `_start` in a test's per-test event stream \u2014 scoped to the three lifecycle events above so it catches regressions without false-positiving on legitimate narrow-slice operational events like `repertoire.task_scan_start`. New `src/__tests__/nerves/pairing-regression.test.ts` locks in the contract with 5 tests (nonexistent-dir, readdirSync-throws, happy-path, startUpdateChecker-pair, daemon-start-throw). Three consecutive `npm run test:coverage` runs confirm zero flakes on the target events."
3241
3255
  ]
3242
3256
  },
3243
3257
  {
3244
3258
  "version": "0.1.0-alpha.275",
3245
3259
  "changes": [
3246
- "fix(tui): backspace on macOS — Ink 3.2 maps \\x7f to key.delete not key.backspace."
3260
+ "fix(tui): backspace on macOS \u2014 Ink 3.2 maps \\x7f to key.delete not key.backspace."
3247
3261
  ]
3248
3262
  },
3249
3263
  {
3250
3264
  "version": "0.1.0-alpha.274",
3251
3265
  "changes": [
3252
3266
  "feat(nerves): daemon log rotation is now 25 MB x 5 gzipped generations instead of 50 MB x 2 uncompressed, dropping peak disk per stream from ~150 MB to ~30 MB. `createNdjsonFileSink` and `rotateIfNeeded` in `src/nerves/index.ts` accept an options object `{ maxSizeBytes, maxGenerations, compress, rotationCheckIntervalBytes }` (with number-form backcompat for the old positional API). Rotation uses a rename-then-gzip pattern so active writers can keep their fd alive while the renamed file gets compressed. Legacy uncompressed `.1.ndjson`/`.2.ndjson` files from the old scheme are tolerated and gzip-migrated on first rotation. Lifecycle emits paired `nerves.rotation_start` / `nerves.rotation_end` events with a shared trace_id, plus `nerves.rotation_error` on failure via a completion-flag try/catch.",
3253
- "feat(log-tailer): `ouro logs` can now read rotated `.ndjson.gz` generations, so `--lines N` spans across historical files. `discoverLogFiles` matches both `.ndjson` and `.ndjson.gz`, parses filenames into (streamBase, rank) tuples, and sorts chronologically (oldest generation first, active last). A new internal `readNdjsonFileContents` helper dispatches to `zlib.gunzipSync` for `.gz` paths while preserving the DI-stubbed plain-file path for existing tests. Follow mode only watches the active stream — gzipped generations are historical by definition and never tailed.",
3254
- "feat(logs-prune): new `ouro logs prune` subcommand applies the active rotation policy to every oversized `.ndjson` file in the agent daemon logs directory. Idempotent — a second run on a compliant dir is a no-op. Concurrent-writer-safe because it delegates to `rotateIfNeeded`'s rename-then-gzip pattern (no locking needed). Emits paired `nerves.logs_prune_start` / `nerves.logs_prune_end` with `nerves.logs_prune_error` on failure. Prints `compacted N file(s), freed M bytes` to stdout. New module `src/heart/daemon/logs-prune.ts` exports `pruneDaemonLogs(options)`; the CLI wire-up adds a `daemon.logs.prune` command kind across cli-types/cli-parse/cli-exec and a `pruneDaemonLogs` dep to `createDefaultOuroCliDeps`.",
3267
+ "feat(log-tailer): `ouro logs` can now read rotated `.ndjson.gz` generations, so `--lines N` spans across historical files. `discoverLogFiles` matches both `.ndjson` and `.ndjson.gz`, parses filenames into (streamBase, rank) tuples, and sorts chronologically (oldest generation first, active last). A new internal `readNdjsonFileContents` helper dispatches to `zlib.gunzipSync` for `.gz` paths while preserving the DI-stubbed plain-file path for existing tests. Follow mode only watches the active stream \u2014 gzipped generations are historical by definition and never tailed.",
3268
+ "feat(logs-prune): new `ouro logs prune` subcommand applies the active rotation policy to every oversized `.ndjson` file in the agent daemon logs directory. Idempotent \u2014 a second run on a compliant dir is a no-op. Concurrent-writer-safe because it delegates to `rotateIfNeeded`'s rename-then-gzip pattern (no locking needed). Emits paired `nerves.logs_prune_start` / `nerves.logs_prune_end` with `nerves.logs_prune_error` on failure. Prints `compacted N file(s), freed M bytes` to stdout. New module `src/heart/daemon/logs-prune.ts` exports `pruneDaemonLogs(options)`; the CLI wire-up adds a `daemon.logs.prune` command kind across cli-types/cli-parse/cli-exec and a `pruneDaemonLogs` dep to `createDefaultOuroCliDeps`.",
3255
3269
  "fix(launchd): drop the stale `StandardErrorPath` plist key that pointed at `ouro-daemon-stderr.log`. The file grew to 366 MB in the wild because the daemon stopped writing to it (the nerves ndjson pipeline has been the source of truth for diagnostics since the nerves layer landed) but nothing ever removed the plist key, so launchd kept the path registered and occasional process-level stderr still dripped in. Removing the key lets launchd forward stray stderr to the system log forwarder where the OS rotates it. The 366 MB orphaned file on disk at `~/AgentBundles/<agent>.ouro/state/daemon/logs/ouro-daemon-stderr.log` is safe to delete manually and is flagged in the PR summary for user cleanup.",
3256
- "refactor(runtime-logging,cli-logging): every `createNdjsonFileSink` call site now passes an explicit `{ maxSizeBytes, maxGenerations, compress }` options object. No callsite silently relies on the old 50 MB default — the policy is visible at each wire-up point. This also removes the `/* v8 ignore */` around the rotation trigger in the sink's flush() loop; tests now exercise it directly via the new `rotationCheckIntervalBytes` option."
3270
+ "refactor(runtime-logging,cli-logging): every `createNdjsonFileSink` call site now passes an explicit `{ maxSizeBytes, maxGenerations, compress }` options object. No callsite silently relies on the old 50 MB default \u2014 the policy is visible at each wire-up point. This also removes the `/* v8 ignore */` around the rotation trigger in the sink's flush() loop; tests now exercise it directly via the new `rotationCheckIntervalBytes` option."
3257
3271
  ]
3258
3272
  },
3259
3273
  {
3260
3274
  "version": "0.1.0-alpha.273",
3261
3275
  "changes": [
3262
- "feat(version-manager): auto-prune old CLI versions during activate. The user observed `~/.ouro-cli/versions/` accumulating every CLI version they'd ever installed (alpha.85 from 2026-03-20 onward, ~100MB+ of dead node_modules trees) because nothing ever GCed. New `pruneOldVersions(retain=5, deps?)` walks `~/.ouro-cli/versions/`, sorts by alpha-suffix numerically, and deletes everything outside the retention window — except always preserves (a) the currently-active version (CurrentVersion symlink target), and (b) the previous version (previous symlink target), so `ouro rollback` stays one command away. Wired into `cli-defaults.ts` via `ensureCurrentVersionInstalled` and `activateCliVersion` — every successful version activation self-prunes. New helpers `compareCliVersions(a, b)` and `selectVersionsToPrune(installed, protected, retain)` are pure and exported for direct unit coverage. 9 new tests covering version comparison, retention selection, current/previous protection, partial-failure handling, missing-symlink fallback, and non-directory entry filtering.",
3263
- "test(contract): new OURO_DAEMON_INSTANTIATION_ALLOWLIST in test-isolation.contract.test.ts flags any new test file that constructs `new OuroDaemon(...)` outside the 11 grandfathered files. Constructing a real daemon and calling start() runs killOrphanProcesses() and writePidfile() against the production pidfile at ~/.ouro-cli/daemon.pids. The runtime guards added in #346 short-circuit those functions under vitest, but if a future change to start() adds a NEW production-state side-effect, the existing 11 tests would silently exercise it. The contract test forces conscious review of each new file taking this shape — same defense-in-depth pattern as BYPASS_USE_ALLOWLIST in alpha.265."
3276
+ "feat(version-manager): auto-prune old CLI versions during activate. The user observed `~/.ouro-cli/versions/` accumulating every CLI version they'd ever installed (alpha.85 from 2026-03-20 onward, ~100MB+ of dead node_modules trees) because nothing ever GCed. New `pruneOldVersions(retain=5, deps?)` walks `~/.ouro-cli/versions/`, sorts by alpha-suffix numerically, and deletes everything outside the retention window \u2014 except always preserves (a) the currently-active version (CurrentVersion symlink target), and (b) the previous version (previous symlink target), so `ouro rollback` stays one command away. Wired into `cli-defaults.ts` via `ensureCurrentVersionInstalled` and `activateCliVersion` \u2014 every successful version activation self-prunes. New helpers `compareCliVersions(a, b)` and `selectVersionsToPrune(installed, protected, retain)` are pure and exported for direct unit coverage. 9 new tests covering version comparison, retention selection, current/previous protection, partial-failure handling, missing-symlink fallback, and non-directory entry filtering.",
3277
+ "test(contract): new OURO_DAEMON_INSTANTIATION_ALLOWLIST in test-isolation.contract.test.ts flags any new test file that constructs `new OuroDaemon(...)` outside the 11 grandfathered files. Constructing a real daemon and calling start() runs killOrphanProcesses() and writePidfile() against the production pidfile at ~/.ouro-cli/daemon.pids. The runtime guards added in #346 short-circuit those functions under vitest, but if a future change to start() adds a NEW production-state side-effect, the existing 11 tests would silently exercise it. The contract test forces conscious review of each new file taking this shape \u2014 same defense-in-depth pattern as BYPASS_USE_ALLOWLIST in alpha.265."
3264
3278
  ]
3265
3279
  },
3266
3280
  {
3267
3281
  "version": "0.1.0-alpha.272",
3268
3282
  "changes": [
3269
- "fix(daemon): the daemon's periodic update checker can now actually auto-update itself. The `onUpdate` callback in `daemon.ts` invoked `performStagedRestart` which (a) ran `npm install -g @ouro.bot/cli@{version}` to install to the global node prefix, and (b) tried to find the new code via `node -e \"console.log(require.resolve('@ouro.bot/cli/package.json'))\"` which depends on the daemon process's NODE_PATH. Neither path was actually reachable from the daemon process running out of `~/.ouro-cli/versions/{version}/...`, so every auto-update attempt bailed at `daemon.staged_restart_path_failed` (\"could not resolve new code path\") and the daemon never updated itself. The user had to manually run `ouro up` to pick up new versions. Fix: switch the staged restart to use the version-managed installer (same one the CLI's `up` flow uses) — `installVersion(version)` puts files at `~/.ouro-cli/versions/{version}/node_modules/@ouro.bot/cli` (deterministic, computable), then `activateVersion(version)` flips the CurrentVersion symlink so the next user-driven `ouro up` sees the same version the daemon is running. `performStagedRestart` gained an optional `installNewVersion` dep that production callers inject; the legacy `npm install -g` fallback path is preserved for tests. Two new regression tests in `staged-restart.test.ts`.",
3283
+ "fix(daemon): the daemon's periodic update checker can now actually auto-update itself. The `onUpdate` callback in `daemon.ts` invoked `performStagedRestart` which (a) ran `npm install -g @ouro.bot/cli@{version}` to install to the global node prefix, and (b) tried to find the new code via `node -e \"console.log(require.resolve('@ouro.bot/cli/package.json'))\"` which depends on the daemon process's NODE_PATH. Neither path was actually reachable from the daemon process running out of `~/.ouro-cli/versions/{version}/...`, so every auto-update attempt bailed at `daemon.staged_restart_path_failed` (\"could not resolve new code path\") and the daemon never updated itself. The user had to manually run `ouro up` to pick up new versions. Fix: switch the staged restart to use the version-managed installer (same one the CLI's `up` flow uses) \u2014 `installVersion(version)` puts files at `~/.ouro-cli/versions/{version}/node_modules/@ouro.bot/cli` (deterministic, computable), then `activateVersion(version)` flips the CurrentVersion symlink so the next user-driven `ouro up` sees the same version the daemon is running. `performStagedRestart` gained an optional `installNewVersion` dep that production callers inject; the legacy `npm install -g` fallback path is preserved for tests. Two new regression tests in `staged-restart.test.ts`.",
3270
3284
  "fix(cli): `ouro up` no longer prints the 'ouro updated to ...' message twice during npx invocations. Three independent paths can detect a CLI version change: (1) `checkForCliUpdate` finds a newer version on npm and re-execs (cross-process print), (2) `ensureCurrentVersionInstalled` flips the CurrentVersion symlink during `performSystemSetup` because the running package version is newer than what the symlink pointed at (path 2, in-process), (3) `bundle-meta.json`'s stored runtime version differs from the running version (path 3, in-process fallback). Path 3's existing guard `linkedVersionBeforeUp !== currentVersion` correctly skipped path 3 when path 1 had already printed in a different process, but did NOT catch the case where path 2 fired in the same process. Verified live on 2026-04-08: `npx --yes @ouro.bot/cli@alpha up` printed the message twice. Fix: track an in-process `printedUpdateMessage` flag set by path 2; path 3 checks it before printing. Path 3 still acts as a fallback when path 2 didn't fire. New regression test in `daemon-cli-update-flow.test.ts` simulates the npx scenario and asserts exactly one print.",
3271
- "fix(hooks): Claude Code lifecycle hooks (`ouro hook session-start|stop|post-tool-use`) now short-circuit when the daemon socket file doesn't exist, instead of attempting `sendDaemonCommand` and logging two ENOENT errors per hook fire (one for `message.send`, one for `inner.wake`). Every Claude Code event during a daemon-down window was producing noisy `connect ENOENT /tmp/ouroboros-daemon.sock` entries in `ouro.ndjson`, which made it hard to read logs around outages. The hook is best-effort by design — dropping notifications when the daemon is down is correct behavior; we just don't want to log spam about it. New nerves event `daemon.hook_skipped_no_socket` (info level) when the short-circuit fires."
3285
+ "fix(hooks): Claude Code lifecycle hooks (`ouro hook session-start|stop|post-tool-use`) now short-circuit when the daemon socket file doesn't exist, instead of attempting `sendDaemonCommand` and logging two ENOENT errors per hook fire (one for `message.send`, one for `inner.wake`). Every Claude Code event during a daemon-down window was producing noisy `connect ENOENT /tmp/ouroboros-daemon.sock` entries in `ouro.ndjson`, which made it hard to read logs around outages. The hook is best-effort by design \u2014 dropping notifications when the daemon is down is correct behavior; we just don't want to log spam about it. New nerves event `daemon.hook_skipped_no_socket` (info level) when the short-circuit fires."
3272
3286
  ]
3273
3287
  },
3274
3288
  {
3275
3289
  "version": "0.1.0-alpha.271",
3276
3290
  "changes": [
3277
- "fix(daemon): unblock daemon.stop deadlock that hung `ouro up` after a CLI auto-update. When the running daemon's version drifted from the local CLI version, `ensureDaemonRunning` would send `daemon.stop` over the socket, the daemon's command handler would `await this.stop()`, and `stop()` would `await server.close()`. server.close() resolves only after every open connection has closed — but the calling client's connection was the ONE thing keeping the server open: its `flushResponse()` was awaiting THIS function call. Both processes sat in kevent forever. Verified live on 2026-04-08: alpha.268 daemon hung at `daemon.server_end` log line for 5+ minutes after a fresh alpha.270 ouro process sent daemon.stop, and the alpha.270 ouro process hung waiting for the response. The deadlock had existed since the original `await server.close()` line was added (2026-03-05) but was masked for weeks by the half-close behavior in socket-client: the client called `client.end()` after writing its command, which (with `allowHalfOpen: false`) caused node to auto-tear-down the server's writable side, incidentally unblocking server.close() before the response was sent. The fix in #303/#334/#339 (which removed `client.end()` and switched to `allowHalfOpen: true` to stop dropping long-running responses like agent.senseTurn) accidentally exposed the underlying deadlock. Fix: don't `await` server.close() in stop() — just fire it. Once stop() returns, the daemon.stop case returns its response, flushResponse calls connection.end(response), the connection closes, and server.close()'s pending callback fires asynchronously. Includes a new daemon-stop-deadlock.test.ts that uses real net sockets to drive daemon.stop and asserts the response comes back within 2s — the test fails (24s timeout) without the fix and passes (110ms) with it."
3291
+ "fix(daemon): unblock daemon.stop deadlock that hung `ouro up` after a CLI auto-update. When the running daemon's version drifted from the local CLI version, `ensureDaemonRunning` would send `daemon.stop` over the socket, the daemon's command handler would `await this.stop()`, and `stop()` would `await server.close()`. server.close() resolves only after every open connection has closed \u2014 but the calling client's connection was the ONE thing keeping the server open: its `flushResponse()` was awaiting THIS function call. Both processes sat in kevent forever. Verified live on 2026-04-08: alpha.268 daemon hung at `daemon.server_end` log line for 5+ minutes after a fresh alpha.270 ouro process sent daemon.stop, and the alpha.270 ouro process hung waiting for the response. The deadlock had existed since the original `await server.close()` line was added (2026-03-05) but was masked for weeks by the half-close behavior in socket-client: the client called `client.end()` after writing its command, which (with `allowHalfOpen: false`) caused node to auto-tear-down the server's writable side, incidentally unblocking server.close() before the response was sent. The fix in #303/#334/#339 (which removed `client.end()` and switched to `allowHalfOpen: true` to stop dropping long-running responses like agent.senseTurn) accidentally exposed the underlying deadlock. Fix: don't `await` server.close() in stop() \u2014 just fire it. Once stop() returns, the daemon.stop case returns its response, flushResponse calls connection.end(response), the connection closes, and server.close()'s pending callback fires asynchronously. Includes a new daemon-stop-deadlock.test.ts that uses real net sockets to drive daemon.stop and asserts the response comes back within 2s \u2014 the test fails (24s timeout) without the fix and passes (110ms) with it."
3278
3292
  ]
3279
3293
  },
3280
3294
  {
3281
3295
  "version": "0.1.0-alpha.270",
3282
3296
  "changes": [
3283
- "feat(daemon): `ouro status` now shows a new `Agents` section listing every discovered bundle with its enabled/disabled state. Previously disabled agents were completely invisible in status — the Senses/Workers/Git Sync sections only iterate managed (enabled) bundles, so a bundle with `\"enabled\": false` in agent.json left no trace in the output. New helper `listAllBundleAgents()` in `agent-discovery.ts` walks the bundles root and returns `{ name, enabled }` for every `<name>.ouro` with a parseable agent.json, and `listEnabledBundleAgents()` now delegates to it. Daemon status payload carries a new `agents: BundleAgentRow[]` field (backward-compat optional in the parser). The stopped-daemon renderer also reads bundles directly from disk so the Agents section works when the daemon is down."
3297
+ "feat(daemon): `ouro status` now shows a new `Agents` section listing every discovered bundle with its enabled/disabled state. Previously disabled agents were completely invisible in status \u2014 the Senses/Workers/Git Sync sections only iterate managed (enabled) bundles, so a bundle with `\"enabled\": false` in agent.json left no trace in the output. New helper `listAllBundleAgents()` in `agent-discovery.ts` walks the bundles root and returns `{ name, enabled }` for every `<name>.ouro` with a parseable agent.json, and `listEnabledBundleAgents()` now delegates to it. Daemon status payload carries a new `agents: BundleAgentRow[]` field (backward-compat optional in the parser). The stopped-daemon renderer also reads bundles directly from disk so the Agents section works when the daemon is down."
3284
3298
  ]
3285
3299
  },
3286
3300
  {
3287
3301
  "version": "0.1.0-alpha.269",
3288
3302
  "changes": [
3289
- "fix(prompt): revert alpha.267 over-engineering and correct the two targeted additions. The existing contextSection already contained the 'save to disk or lose it' teaching ('my conversation memory is ephemeral -- it resets between sessions. anything i learn about my friend, i save with save_friend_note so future me remembers.'), and toolContractsSection already told agents to call save_friend_note/diary_write before responding. alpha.267 added a third block in memoryJudgementSection about 'not just remembering between sessions' which was redundant and misplaced (memoryJudgementSection sits under 'my tools & capabilities' alongside tool routing heuristics, not nature-of-self content). That block is reverted. What alpha.267 got right and this PR keeps: (a) bodyMapSection guidance that standard folders are a floor, not a ceiling — bundles can have custom top-level folders (an early bundle's travel/ was the motivating example) — with a nudge to try the file-listing tool on the bundle root BEFORE falling back to recall, and (b) a diary-routing bullet that flags bundle-layout discoveries as worth persisting. Both had errors corrected here: alpha.267 told agents to use `list_directory` but the actual tool is `glob` (with a pattern like `*/`), and it said 'write a diary note like bundle-layout.md' but diary/ is a jsonl fact store queried via recall, not a directory of .md files — the corrected bullet says 'save the fact with diary_write'."
3303
+ "fix(prompt): revert alpha.267 over-engineering and correct the two targeted additions. The existing contextSection already contained the 'save to disk or lose it' teaching ('my conversation memory is ephemeral -- it resets between sessions. anything i learn about my friend, i save with save_friend_note so future me remembers.'), and toolContractsSection already told agents to call save_friend_note/diary_write before responding. alpha.267 added a third block in memoryJudgementSection about 'not just remembering between sessions' which was redundant and misplaced (memoryJudgementSection sits under 'my tools & capabilities' alongside tool routing heuristics, not nature-of-self content). That block is reverted. What alpha.267 got right and this PR keeps: (a) bodyMapSection guidance that standard folders are a floor, not a ceiling \u2014 bundles can have custom top-level folders (an early bundle's travel/ was the motivating example) \u2014 with a nudge to try the file-listing tool on the bundle root BEFORE falling back to recall, and (b) a diary-routing bullet that flags bundle-layout discoveries as worth persisting. Both had errors corrected here: alpha.267 told agents to use `list_directory` but the actual tool is `glob` (with a pattern like `*/`), and it said 'write a diary note like bundle-layout.md' but diary/ is a jsonl fact store queried via recall, not a directory of .md files \u2014 the corrected bullet says 'save the fact with diary_write'."
3290
3304
  ]
3291
3305
  },
3292
3306
  {
3293
3307
  "version": "0.1.0-alpha.268",
3294
3308
  "changes": [
3295
- "fix(identity): `loadAgentConfig` now preserves the `sync` block from agent.json. The hand-rolled object literal that constructs the typed `AgentConfig` was missing `sync` from its field list, so `agentConfig.sync` was always `undefined`, `getSyncConfig()` always returned `enabled: false`, and the entire sync code path (`preTurnPull` / `postTurnPush` in `pipeline.ts`) was dead from the moment the field was added to the type. The bug hid for weeks because `ouro status` reads `agent.json` directly via `listBundleSyncRows`, not through `loadAgentConfig` — so the per-agent Git Sync display correctly showed `enabled origin → ...` while sync did literally nothing. Slugger accumulated 5+ days of dirty files and zero `sync: post-turn update` commits before this surfaced. Also adds 3 regression tests asserting the sync block round-trips through `loadAgentConfig` (full block, partial block, missing block). The first two would have caught the original bug; without them future field additions to `AgentConfig` could repeat the pattern."
3309
+ "fix(identity): `loadAgentConfig` now preserves the `sync` block from agent.json. The hand-rolled object literal that constructs the typed `AgentConfig` was missing `sync` from its field list, so `agentConfig.sync` was always `undefined`, `getSyncConfig()` always returned `enabled: false`, and the entire sync code path (`preTurnPull` / `postTurnPush` in `pipeline.ts`) was dead from the moment the field was added to the type. The bug hid for weeks because `ouro status` reads `agent.json` directly via `listBundleSyncRows`, not through `loadAgentConfig` \u2014 so the per-agent Git Sync display correctly showed `enabled origin \u2192 ...` while sync did literally nothing. Slugger accumulated 5+ days of dirty files and zero `sync: post-turn update` commits before this surfaced. Also adds 3 regression tests asserting the sync block round-trips through `loadAgentConfig` (full block, partial block, missing block). The first two would have caught the original bug; without them future field additions to `AgentConfig` could repeat the pattern."
3296
3310
  ]
3297
3311
  },
3298
3312
  {
3299
3313
  "version": "0.1.0-alpha.267",
3300
3314
  "changes": [
3301
- "fix(prompt): teach agents that they do NOT 'just remember' things between sessions. Observed failure mode: an agent spent a recall scan searching diary/journal for a `travel/` folder that was right at the bundle root, then concluded 'i should know my own folder structure next time' and refused to persist the discovery because 'that's just me knowing my own home.' That is impossible — next session is a blank slate. The agent conflated in-session realizations with persistent knowledge. Three prompt changes fix this: (1) bodyMapSection now explicitly states that standard folders are a floor, bundles MAY have custom top-level folders created by the friend over time, and the agent should `list_directory` the bundle root BEFORE reaching for `recall` when something might be in a custom location. (2) memoryJudgementSection opens with an always-on reminder that the agent does not 'just remember' anything between sessions — every session carries only (a) the prompt, (b) what's on disk, (c) what tools observe this turn — and that thoughts like 'i should know this next time' or 'i'll look there first in the future' are CUES to write a concrete diary/friend-note RIGHT NOW. Future me cannot inherit resolutions, only files. (3) memoryJudgementSection's diary-routing rules now explicitly call out bundle-layout discoveries as a thing worth persisting (e.g. to a `bundle-layout.md` diary note). Includes 3 new tests locking the directives into place so future prompt edits cannot silently drop them."
3315
+ "fix(prompt): teach agents that they do NOT 'just remember' things between sessions. Observed failure mode: an agent spent a recall scan searching diary/journal for a `travel/` folder that was right at the bundle root, then concluded 'i should know my own folder structure next time' and refused to persist the discovery because 'that's just me knowing my own home.' That is impossible \u2014 next session is a blank slate. The agent conflated in-session realizations with persistent knowledge. Three prompt changes fix this: (1) bodyMapSection now explicitly states that standard folders are a floor, bundles MAY have custom top-level folders created by the friend over time, and the agent should `list_directory` the bundle root BEFORE reaching for `recall` when something might be in a custom location. (2) memoryJudgementSection opens with an always-on reminder that the agent does not 'just remember' anything between sessions \u2014 every session carries only (a) the prompt, (b) what's on disk, (c) what tools observe this turn \u2014 and that thoughts like 'i should know this next time' or 'i'll look there first in the future' are CUES to write a concrete diary/friend-note RIGHT NOW. Future me cannot inherit resolutions, only files. (3) memoryJudgementSection's diary-routing rules now explicitly call out bundle-layout discoveries as a thing worth persisting (e.g. to a `bundle-layout.md` diary note). Includes 3 new tests locking the directives into place so future prompt edits cannot silently drop them."
3302
3316
  ]
3303
3317
  },
3304
3318
  {
3305
3319
  "version": "0.1.0-alpha.266",
3306
3320
  "changes": [
3307
- "fix(daemon): vitest guard for production pidfile (~/.ouro-cli/daemon.pids). The pidfile path is hardcoded with no DI seam, so when a test creates a real OuroDaemon instance and calls start(), the daemon's killOrphanProcesses() reads the REAL pidfile, ps-verifies the PIDs, and SIGTERMs the production daemon's PIDs. Verified live: alpha.265 daemon (PID 64988) was killed 93s after startup by `npx vitest run` invoking daemon.start() in 6 different test files. SIGTERM forensics added in alpha.265 captured the death (parentPid=1, parentCommand=/sbin/launchd) but the killer hint was misleading — the real culprit was the production-pidfile leak. Fix: killOrphanProcesses() and writePidfile() now short-circuit under vitest with warn-level nerves events. Tests that need to verify these functions' behavior continue to use the extracted pure helpers (parseOrphanPidsFromPs, filterPidfilePidsToActualOrphans). New unit tests verify the pidfile is unchanged after no-op calls. Same defense-in-depth pattern as the alpha.265 socket-client hardening — production-side state being touched from tests is now physically impossible from a vitest worker."
3321
+ "fix(daemon): vitest guard for production pidfile (~/.ouro-cli/daemon.pids). The pidfile path is hardcoded with no DI seam, so when a test creates a real OuroDaemon instance and calls start(), the daemon's killOrphanProcesses() reads the REAL pidfile, ps-verifies the PIDs, and SIGTERMs the production daemon's PIDs. Verified live: alpha.265 daemon (PID 64988) was killed 93s after startup by `npx vitest run` invoking daemon.start() in 6 different test files. SIGTERM forensics added in alpha.265 captured the death (parentPid=1, parentCommand=/sbin/launchd) but the killer hint was misleading \u2014 the real culprit was the production-pidfile leak. Fix: killOrphanProcesses() and writePidfile() now short-circuit under vitest with warn-level nerves events. Tests that need to verify these functions' behavior continue to use the extracted pure helpers (parseOrphanPidsFromPs, filterPidfilePidsToActualOrphans). New unit tests verify the pidfile is unchanged after no-op calls. Same defense-in-depth pattern as the alpha.265 socket-client hardening \u2014 production-side state being touched from tests is now physically impossible from a vitest worker."
3308
3322
  ]
3309
3323
  },
3310
3324
  {
3311
3325
  "version": "0.1.0-alpha.265",
3312
3326
  "changes": [
3313
- "fix(daemon): bulletproof vitest socket leak + SIGTERM forensics. Hardens the socket-client vitest guard so production daemon socket calls (DEFAULT_DAEMON_SOCKET_PATH = /tmp/ouroboros-daemon.sock) are unconditionally blocked under vitest, regardless of __bypassVitestGuardForTests state. Cross-file leaks via the globalThis bypass flag (process-wide, leaks across concurrent test files in the same vitest worker) can no longer reach the production daemon. Test files that legitimately exercise the real socket-client transport against synthetic test socket paths (/tmp/daemon.sock) continue to work. Background: a daemon outage on 2026-04-08 was traced to a 14-call burst of leaked `inner.wake testagent` errors, signature of vitest test runs hammering the production socket — same pattern as 1,460 historical daemon log entries.",
3327
+ "fix(daemon): bulletproof vitest socket leak + SIGTERM forensics. Hardens the socket-client vitest guard so production daemon socket calls (DEFAULT_DAEMON_SOCKET_PATH = /tmp/ouroboros-daemon.sock) are unconditionally blocked under vitest, regardless of __bypassVitestGuardForTests state. Cross-file leaks via the globalThis bypass flag (process-wide, leaks across concurrent test files in the same vitest worker) can no longer reach the production daemon. Test files that legitimately exercise the real socket-client transport against synthetic test socket paths (/tmp/daemon.sock) continue to work. Background: a daemon outage on 2026-04-08 was traced to a 14-call burst of leaked `inner.wake testagent` errors, signature of vitest test runs hammering the production socket \u2014 same pattern as 1,460 historical daemon log entries.",
3314
3328
  "feat(daemon): SIGTERM/SIGINT tombstone forensics. New writeDaemonTombstone path captures process.ppid, parent command via `ps -p <ppid> -o command=`, and a filtered process snapshot (only node/vitest/ouro/kill lines) at signal-driven death time. Adds a killerHint heuristic: launchd parent reparenting suggests `launchctl bootout`/KeepAlive thrash; vitest worker presence suggests test cleanup; pkill/killall presence is an explicit kill. Forensics field is parsed by readDaemonTombstone and included in the daemon.tombstone_written nerves event meta. Also caps recentCrashes at 100 entries (was 12,265 from a March 31 thrash loop).",
3315
3329
  "test(daemon): contract test BYPASS_USE_ALLOWLIST flags any new file calling __bypassVitestGuardForTests outside the two known-good test files (socket-client.test.ts and daemon-cli-defaults.test.ts). Prevents future regressions of the cross-file leak vector."
3316
3330
  ]
@@ -3318,7 +3332,7 @@
3318
3332
  {
3319
3333
  "version": "0.1.0-alpha.264",
3320
3334
  "changes": [
3321
- "fix(bluebubbles): dedup `updated-message` webhooks BEFORE running repair+hydrate+VLM. BlueBubbles routinely sends a `new-message` webhook for a fresh message, then follows up seconds later with one or more `updated-message` webhooks for delivery/read status. The BB sense's `repairEvent` path promotes updated-message events with recoverable content back to `message` kind, which re-runs the full hydration pipeline — including a second MiniMax VLM describe call on the same image. Verified live on 2026-04-08T00:58Z: two sequential VLM describes for attachment guid 317E37EB-..., 13.7s + 14.0s each, for the exact same 291KB JPEG, triggered by a `new-message` followed 3s later by an `updated-message` for the same guid. The downstream `handleBlueBubblesNormalizedEvent` dedup check was firing correctly but too late (after the expensive VLM round-trip). Fix: add a pre-repair dedup check in `handleBlueBubblesEvent` that consults the inbound sidecar by messageGuid and short-circuits before calling `client.repairEvent(...)`. New nerves event `senses.bluebubbles_repair_skipped_duplicate` at warn level for observability."
3335
+ "fix(bluebubbles): dedup `updated-message` webhooks BEFORE running repair+hydrate+VLM. BlueBubbles routinely sends a `new-message` webhook for a fresh message, then follows up seconds later with one or more `updated-message` webhooks for delivery/read status. The BB sense's `repairEvent` path promotes updated-message events with recoverable content back to `message` kind, which re-runs the full hydration pipeline \u2014 including a second MiniMax VLM describe call on the same image. Verified live on 2026-04-08T00:58Z: two sequential VLM describes for attachment guid 317E37EB-..., 13.7s + 14.0s each, for the exact same 291KB JPEG, triggered by a `new-message` followed 3s later by an `updated-message` for the same guid. The downstream `handleBlueBubblesNormalizedEvent` dedup check was firing correctly but too late (after the expensive VLM round-trip). Fix: add a pre-repair dedup check in `handleBlueBubblesEvent` that consults the inbound sidecar by messageGuid and short-circuits before calling `client.repairEvent(...)`. New nerves event `senses.bluebubbles_repair_skipped_duplicate` at warn level for observability."
3322
3336
  ]
3323
3337
  },
3324
3338
  {
@@ -3330,28 +3344,28 @@
3330
3344
  {
3331
3345
  "version": "0.1.0-alpha.262",
3332
3346
  "changes": [
3333
- "test(daemon): inject explicit `vi.mock(\"...heart/daemon/socket-client\", ...)` blocks into all 40 grandfathered test files in `TESTAGENT_NO_MOCK_ALLOWLIST`. The runtime guard in `socket-client.ts` already prevents real socket leaks under vitest, but the explicit mocks let those tests assert call counts cleanly and shrink the contract-test allowlist to zero. Both the testagent and the bundle-write allowlists in `test-isolation.contract.test.ts` are now empty Sets — any new offender fails the build."
3347
+ "test(daemon): inject explicit `vi.mock(\"...heart/daemon/socket-client\", ...)` blocks into all 40 grandfathered test files in `TESTAGENT_NO_MOCK_ALLOWLIST`. The runtime guard in `socket-client.ts` already prevents real socket leaks under vitest, but the explicit mocks let those tests assert call counts cleanly and shrink the contract-test allowlist to zero. Both the testagent and the bundle-write allowlists in `test-isolation.contract.test.ts` are now empty Sets \u2014 any new offender fails the build."
3334
3348
  ]
3335
3349
  },
3336
3350
  {
3337
3351
  "version": "0.1.0-alpha.261",
3338
3352
  "changes": [
3339
- "feat(tui): full input parity with Claude Code — kill ring, emacs nav, Home/End, Ctrl+D, forward delete, Esc history, bracketed paste, clipboard image, token deletion, chip navigation. 156 new tests."
3353
+ "feat(tui): full input parity with Claude Code \u2014 kill ring, emacs nav, Home/End, Ctrl+D, forward delete, Esc history, bracketed paste, clipboard image, token deletion, chip navigation. 156 new tests."
3340
3354
  ]
3341
3355
  },
3342
3356
  {
3343
3357
  "version": "0.1.0-alpha.260",
3344
3358
  "changes": [
3345
- "fix(daemon): MCP bridge empty-response bug for long-running commands. `sendDaemonCommand` and `checkDaemonSocketAlive` were calling `client.end()` immediately after writing, which half-closed the TCP connection. The daemon server's `net.createServer()` uses the default `allowHalfOpen: false`, so when it saw the client's FIN it auto-closed its own writable side — and any response the server tried to write after processing a long-running command (like `agent.senseTurn`, which runs a full LLM turn) was dropped on the floor. Verified via direct socket repro: with `client.end()`, `agent.senseTurn` returned empty in ~149ms; without it, the same call returned a real response in ~5.8s. This was a silent regression of the fix in #303 (commit `253e4b1f` titled \"socket half-close fix\" actually *added* the half-close back). Fix: drop the `client.end()` calls from both sites AND set `allowHalfOpen: true` on the daemon's `net.createServer(...)` as defense-in-depth so future clients calling `end()` don't silently break again.",
3346
- "fix(daemon): tighten pidfile trust — `killOrphanProcesses` now verifies each pidfile PID is an actual orphan (PPID reparented to init/PID 1) before SIGTERMing it. Previously a polluted pidfile (written by a crashed daemon whose PIDs have since been reused by the OS for unrelated processes) could cause mass-kill of unrelated apps. New exported helper `filterPidfilePidsToActualOrphans` provides direct unit coverage.",
3359
+ "fix(daemon): MCP bridge empty-response bug for long-running commands. `sendDaemonCommand` and `checkDaemonSocketAlive` were calling `client.end()` immediately after writing, which half-closed the TCP connection. The daemon server's `net.createServer()` uses the default `allowHalfOpen: false`, so when it saw the client's FIN it auto-closed its own writable side \u2014 and any response the server tried to write after processing a long-running command (like `agent.senseTurn`, which runs a full LLM turn) was dropped on the floor. Verified via direct socket repro: with `client.end()`, `agent.senseTurn` returned empty in ~149ms; without it, the same call returned a real response in ~5.8s. This was a silent regression of the fix in #303 (commit `253e4b1f` titled \"socket half-close fix\" actually *added* the half-close back). Fix: drop the `client.end()` calls from both sites AND set `allowHalfOpen: true` on the daemon's `net.createServer(...)` as defense-in-depth so future clients calling `end()` don't silently break again.",
3360
+ "fix(daemon): tighten pidfile trust \u2014 `killOrphanProcesses` now verifies each pidfile PID is an actual orphan (PPID reparented to init/PID 1) before SIGTERMing it. Previously a polluted pidfile (written by a crashed daemon whose PIDs have since been reused by the OS for unrelated processes) could cause mass-kill of unrelated apps. New exported helper `filterPidfilePidsToActualOrphans` provides direct unit coverage.",
3347
3361
  "fix(daemon, repertoire): rename six `_stop` nerves events to `_end` so they pair correctly with their `_start` counterparts under the nerves audit's start/end pairing rule. Affected: `daemon.thoughts_follow_stop`, `daemon.server_stop`, `daemon.update_checker_stop`, `daemon.mcp_server_stop`, `daemon.habit_scheduler_stop`, `mcp.manager_stop`. The naming was semantically fine (`stop` pairs with `start`), but the audit specifically looks for `_end`/`_error` suffixes. Was producing intermittent audit failures whenever any test run exercised these teardown paths (seen today as flakes on `daemon-socket-errors.test.ts` and `MarkdownStreamer` tests).",
3348
- "fix(providers): drop the harness-imposed MiniMax VLM timeout entirely. Previously the module defaulted to a 60-second AbortSignal; E2E validation of the BB image fix hit a real VLM request that took >60s to return (same bytes returned in 9.5s on the immediate retry). Raising to 120s would have been arbitrary too — the correct answer, per the same reasoning as PR #322 for LLM providers, is to not impose a harness ceiling at all. When `timeoutMs` isn't provided, `fetch()` now runs without an AbortSignal and undici's own defaults (headersTimeout + bodyTimeout, both 5 minutes) are the ceiling. Callers that want a tighter bound can still pass an explicit `timeoutMs`. The AbortError path is kept for that case, and the error message adapts to say \"underlying stack default\" when there was no harness-set value."
3362
+ "fix(providers): drop the harness-imposed MiniMax VLM timeout entirely. Previously the module defaulted to a 60-second AbortSignal; E2E validation of the BB image fix hit a real VLM request that took >60s to return (same bytes returned in 9.5s on the immediate retry). Raising to 120s would have been arbitrary too \u2014 the correct answer, per the same reasoning as PR #322 for LLM providers, is to not impose a harness ceiling at all. When `timeoutMs` isn't provided, `fetch()` now runs without an AbortSignal and undici's own defaults (headersTimeout + bodyTimeout, both 5 minutes) are the ceiling. Callers that want a tighter bound can still pass an explicit `timeoutMs`. The AbortError path is kept for that case, and the error message adapts to say \"underlying stack default\" when there was no harness-set value."
3349
3363
  ]
3350
3364
  },
3351
3365
  {
3352
3366
  "version": "0.1.0-alpha.259",
3353
3367
  "changes": [
3354
- "fix(daemon): SIGINT and SIGTERM now ALWAYS write a tombstone before exiting, instead of silently skipping when `_gracefulShutdown` was set. The previous behavior made signal-driven shutdowns invisible in `~/.ouro-cli/daemon-death.json` — launchd policy decisions, the OOM killer, manual `kill`, and `killOrphanProcesses` from a sibling daemon all looked identical to a clean exit. After a real outage where the user's daemon kept dying with a tombstone from a week earlier, this restores observability: every signal-driven exit now records `reason: \"sigint\"` or `reason: \"sigterm\"` with timestamp + recentCrashes accumulator. The catch-all `process.on('exit')` handler also no longer short-circuits on graceful shutdown."
3368
+ "fix(daemon): SIGINT and SIGTERM now ALWAYS write a tombstone before exiting, instead of silently skipping when `_gracefulShutdown` was set. The previous behavior made signal-driven shutdowns invisible in `~/.ouro-cli/daemon-death.json` \u2014 launchd policy decisions, the OOM killer, manual `kill`, and `killOrphanProcesses` from a sibling daemon all looked identical to a clean exit. After a real outage where the user's daemon kept dying with a tombstone from a week earlier, this restores observability: every signal-driven exit now records `reason: \"sigint\"` or `reason: \"sigterm\"` with timestamp + recentCrashes accumulator. The catch-all `process.on('exit')` handler also no longer short-circuits on graceful shutdown."
3355
3369
  ]
3356
3370
  },
3357
3371
  {
@@ -3365,14 +3379,14 @@
3365
3379
  {
3366
3380
  "version": "0.1.0-alpha.257",
3367
3381
  "changes": [
3368
- "test(daemon): convert all `daemon-cli.test.ts` auth/thoughts/config tests to use `os.tmpdir()` via the new `createTmpBundle()` helper. Previously these tests wrote real bundles to `~/AgentBundles/auth-local-${Date.now()}.ouro` etc., relying on `try/finally` cleanup that doesn't fire on test interruption — leaking bundle directories into the developer's home and inflating noise on the running daemon. The `REAL_BUNDLES_WRITE_ALLOWLIST` ratchet in `test-isolation.contract.test.ts` is now empty.",
3382
+ "test(daemon): convert all `daemon-cli.test.ts` auth/thoughts/config tests to use `os.tmpdir()` via the new `createTmpBundle()` helper. Previously these tests wrote real bundles to `~/AgentBundles/auth-local-${Date.now()}.ouro` etc., relying on `try/finally` cleanup that doesn't fire on test interruption \u2014 leaking bundle directories into the developer's home and inflating noise on the running daemon. The `REAL_BUNDLES_WRITE_ALLOWLIST` ratchet in `test-isolation.contract.test.ts` is now empty.",
3369
3383
  "fix(daemon/cli-exec): plumb `bundlesRoot` and `secretsRoot` deps through the auth.run / auth.verify / auth.switch / config.model / config.models / thoughts handlers so tests can route reads/writes to a tmpdir without monkey-patching the identity module. Production code paths still default to `getAgentBundlesRoot()` and `~/.agentsecrets`."
3370
3384
  ]
3371
3385
  },
3372
3386
  {
3373
3387
  "version": "0.1.0-alpha.256",
3374
3388
  "changes": [
3375
- "fix(tui): cursor renders as inverse character (not inserted block) — matches standard terminal cursor behavior.",
3389
+ "fix(tui): cursor renders as inverse character (not inserted block) \u2014 matches standard terminal cursor behavior.",
3376
3390
  "fix(tui): Alt+Enter via ESC-timing (50ms window) instead of raw stdin handler that interfered with arrow keys.",
3377
3391
  "fix(tui): image path regex now matches backslash-escaped spaces from macOS drag-drop."
3378
3392
  ]
@@ -3381,20 +3395,20 @@
3381
3395
  "version": "0.1.0-alpha.255",
3382
3396
  "changes": [
3383
3397
  "feat(tui): session resume shows last messages as regular chat (no dimmed preview) with teal resume banner in header.",
3384
- "feat(tui): image/file drag-and-drop — detects image paths in pasted text, reads to base64, inserts [Image #N] references, sends as image_url content blocks to model.",
3398
+ "feat(tui): image/file drag-and-drop \u2014 detects image paths in pasted text, reads to base64, inserts [Image #N] references, sends as image_url content blocks to model.",
3385
3399
  "refactor(tui): removed custom history-* roles and addSessionHistory (KISS/DRY)."
3386
3400
  ]
3387
3401
  },
3388
3402
  {
3389
3403
  "version": "0.1.0-alpha.254",
3390
3404
  "changes": [
3391
- "fix(daemon): orphan-cleanup fallback no longer kills processes from sibling harness instances. On startup, `killOrphanProcesses` scans `ps` for harness entry points (`agent-entry.js`, `daemon-entry.js`, `bluebubbles/entry.js`, `teams-entry.js`) and SIGTERMs them when the pidfile is missing. Previously any matching process was fair game, so a vitest-driven harness run from a sibling worktree (or a parallel Claude Code session running the coverage gate) would terminate the production daemon's children, triggering cascading graceful shutdowns and making the production agent unavailable for seconds-to-minutes at a time. Now the fallback only flags true orphans — processes whose PPID has been reparented to init (PID 1). Sibling daemons' live-parented children are left alone. New exported helper `parseOrphanPidsFromPs` isolates the filter for direct unit coverage. Complements the test-isolation guard shipped in alpha.253 (#333): that PR stopped tests from SENDING `inner.wake testagent` commands into the production socket; this PR stops the production daemon from SIGTERMing test-spawned child processes during its own startup."
3405
+ "fix(daemon): orphan-cleanup fallback no longer kills processes from sibling harness instances. On startup, `killOrphanProcesses` scans `ps` for harness entry points (`agent-entry.js`, `daemon-entry.js`, `bluebubbles/entry.js`, `teams-entry.js`) and SIGTERMs them when the pidfile is missing. Previously any matching process was fair game, so a vitest-driven harness run from a sibling worktree (or a parallel Claude Code session running the coverage gate) would terminate the production daemon's children, triggering cascading graceful shutdowns and making the production agent unavailable for seconds-to-minutes at a time. Now the fallback only flags true orphans \u2014 processes whose PPID has been reparented to init (PID 1). Sibling daemons' live-parented children are left alone. New exported helper `parseOrphanPidsFromPs` isolates the filter for direct unit coverage. Complements the test-isolation guard shipped in alpha.253 (#333): that PR stopped tests from SENDING `inner.wake testagent` commands into the production socket; this PR stops the production daemon from SIGTERMing test-spawned child processes during its own startup."
3392
3406
  ]
3393
3407
  },
3394
3408
  {
3395
3409
  "version": "0.1.0-alpha.253",
3396
3410
  "changes": [
3397
- "fix(daemon): test-isolation guard — stop tests from leaking real `inner.wake` commands into the running daemon. A pattern in 36+ test files mocks `getAgentName` to the literal `\"testagent\"` but does NOT mock `socket-client`, so any code path through pondering / rest / coding feedback fires a real socket connection at /tmp/ouroboros-daemon.sock with `inner.wake testagent`. The daemon errored on every command (`Unknown managed agent 'testagent'`) and at flood volumes that contributed to a real outage on the developer's machine. The fix is in `socket-client.ts` itself: detect vitest via `process.argv` (no env vars) and convert all socket operations into safe no-ops. Tests that legitimately exercise the real transport (socket-client.test.ts, daemon-cli-defaults.test.ts) opt out of the guard via a new `__bypassVitestGuardForTests()` setter that lives on globalThis to survive `vi.resetModules()`. New nerves events: `daemon.socket_command_test_blocked`, `daemon.inner_wake_test_blocked`.",
3411
+ "fix(daemon): test-isolation guard \u2014 stop tests from leaking real `inner.wake` commands into the running daemon. A pattern in 36+ test files mocks `getAgentName` to the literal `\"testagent\"` but does NOT mock `socket-client`, so any code path through pondering / rest / coding feedback fires a real socket connection at /tmp/ouroboros-daemon.sock with `inner.wake testagent`. The daemon errored on every command (`Unknown managed agent 'testagent'`) and at flood volumes that contributed to a real outage on the developer's machine. The fix is in `socket-client.ts` itself: detect vitest via `process.argv` (no env vars) and convert all socket operations into safe no-ops. Tests that legitimately exercise the real transport (socket-client.test.ts, daemon-cli-defaults.test.ts) opt out of the guard via a new `__bypassVitestGuardForTests()` setter that lives on globalThis to survive `vi.resetModules()`. New nerves events: `daemon.socket_command_test_blocked`, `daemon.inner_wake_test_blocked`.",
3398
3412
  "test(contract): new test-isolation contract test ratchets two anti-patterns: (1) test files using `name: \"testagent\"` without mocking socket-client, and (2) test files constructing write paths under the real `~/AgentBundles` via `os.homedir()`. Existing offenders are grandfathered in two allowlists; new offenders fail the build. Follow-up PRs shrink the allowlists toward zero."
3399
3413
  ]
3400
3414
  },
@@ -3407,7 +3421,7 @@
3407
3421
  {
3408
3422
  "version": "0.1.0-alpha.251",
3409
3423
  "changes": [
3410
- "fix(bluebubbles): images sent via iMessage now reach the model — adds capability-gated image hydration with a MiniMax VLM fallback. Reasoning MiniMax chat models (M2/M2.1/M2.5/M2.7) silently drop OpenAI-style `image_url` content parts, so previously the agent answered image questions from fabricated context. Now, when the active chat model lacks the new `vision: true` capability flag, inbound screenshots are auto-described at ingestion via `/v1/coding_plan/vlm` and the description text replaces the `image_url` part before the turn reaches the model. Vision-capable models (claude-opus/sonnet-4-6, gpt-5.4, MiniMax-Text-01, MiniMax-VL-01) continue to see images natively via pass-through.",
3424
+ "fix(bluebubbles): images sent via iMessage now reach the model \u2014 adds capability-gated image hydration with a MiniMax VLM fallback. Reasoning MiniMax chat models (M2/M2.1/M2.5/M2.7) silently drop OpenAI-style `image_url` content parts, so previously the agent answered image questions from fabricated context. Now, when the active chat model lacks the new `vision: true` capability flag, inbound screenshots are auto-described at ingestion via `/v1/coding_plan/vlm` and the description text replaces the `image_url` part before the turn reaches the model. Vision-capable models (claude-opus/sonnet-4-6, gpt-5.4, MiniMax-Text-01, MiniMax-VL-01) continue to see images natively via pass-through.",
3411
3425
  "feat(tools): new `describe_image` agent tool registered into the BlueBubbles tool set when the chat model lacks vision. Lets the agent re-interrogate an attachment with a targeted prompt (e.g. 'what's the flight number in the bottom-right?') after ingestion. Backed by a bounded in-memory attachment cache populated during hydration; handler re-downloads bytes and calls the same VLM client.",
3412
3426
  "fix(bluebubbles): `formatMessageText` now preserves the attachment marker when a message has BOTH text and attachments (B2). Previously the marker was dropped whenever text was present, hiding attachments from the agent's view of the message.",
3413
3427
  "feat(heart): `ModelCapabilities` gains `vision?: boolean` and `audio?: boolean` flags (B4). Vision rows populated for claude-opus-4-6, claude-sonnet-4-6, claude-opus-4.6, claude-sonnet-4.6, gpt-5.4, MiniMax-Text-01, MiniMax-VL-01. M2.1/M2.5/M2.7 intentionally left unset.",
@@ -3417,7 +3431,7 @@
3417
3431
  {
3418
3432
  "version": "0.1.0-alpha.250",
3419
3433
  "changes": [
3420
- "fix(sync): surface 'bundle is not a git repo' as an actionable error instead of silently failing. Previously, enabling `sync.enabled` on a bundle that had never been `git init`'d produced a generic `git status` failure buried in nerves logs; the agent saw nothing in its start-of-turn packet and the user saw nothing in `ouro status`. Now: (1) `preTurnPull` and `postTurnPush` detect the missing `.git` directory before touching git and return an actionable error with the bundle path and the `git init` hint; (2) this error propagates via `ctx.syncFailure` into the agent's Sync warning, so the agent can offer to run `git init` or just do it; (3) `ouro status` shows a red `error` state with `not a git repo — run \\`git init\\` to enable sync` next to the offending bundle. New nerves event: `heart.sync_not_a_repo`."
3434
+ "fix(sync): surface 'bundle is not a git repo' as an actionable error instead of silently failing. Previously, enabling `sync.enabled` on a bundle that had never been `git init`'d produced a generic `git status` failure buried in nerves logs; the agent saw nothing in its start-of-turn packet and the user saw nothing in `ouro status`. Now: (1) `preTurnPull` and `postTurnPush` detect the missing `.git` directory before touching git and return an actionable error with the bundle path and the `git init` hint; (2) this error propagates via `ctx.syncFailure` into the agent's Sync warning, so the agent can offer to run `git init` or just do it; (3) `ouro status` shows a red `error` state with `not a git repo \u2014 run \\`git init\\` to enable sync` next to the offending bundle. New nerves event: `heart.sync_not_a_repo`."
3421
3435
  ]
3422
3436
  },
3423
3437
  {
@@ -3430,8 +3444,8 @@
3430
3444
  {
3431
3445
  "version": "0.1.0-alpha.248",
3432
3446
  "changes": [
3433
- "feat(daemon): `ouro status` Git Sync now resolves and shows the actual remote URL via `git remote get-url`. Three states: `origin → git@github.com:me/foo.git` when the remote resolves, `local only` when sync is enabled but no remote is configured, `disabled` when sync is off. Previously you had to `cd` into the bundle and run `git remote -v` to find out where it pushes.",
3434
- "fix(heart/sync): `preTurnPull` now skips the pull when no git remote is configured, mirroring the existing `postTurnPush` behavior. Closes the half-implemented 'no-remote sync' (local-only commit log) story — previously, enabling sync without a remote produced a `syncFailure` on every turn from the failing `git pull` call."
3447
+ "feat(daemon): `ouro status` Git Sync now resolves and shows the actual remote URL via `git remote get-url`. Three states: `origin \u2192 git@github.com:me/foo.git` when the remote resolves, `local only` when sync is enabled but no remote is configured, `disabled` when sync is off. Previously you had to `cd` into the bundle and run `git remote -v` to find out where it pushes.",
3448
+ "fix(heart/sync): `preTurnPull` now skips the pull when no git remote is configured, mirroring the existing `postTurnPush` behavior. Closes the half-implemented 'no-remote sync' (local-only commit log) story \u2014 previously, enabling sync without a remote produced a `syncFailure` on every turn from the failing `git pull` call."
3435
3449
  ]
3436
3450
  },
3437
3451
  {
@@ -3443,7 +3457,7 @@
3443
3457
  {
3444
3458
  "version": "0.1.0-alpha.246",
3445
3459
  "changes": [
3446
- "feat(tui): session resume display — shows summary line + last 3 exchanges dimmed when reconnecting to existing session."
3460
+ "feat(tui): session resume display \u2014 shows summary line + last 3 exchanges dimmed when reconnecting to existing session."
3447
3461
  ]
3448
3462
  },
3449
3463
  {
@@ -3455,58 +3469,58 @@
3455
3469
  {
3456
3470
  "version": "0.1.0-alpha.244",
3457
3471
  "changes": [
3458
- "refactor(sync): git-status-based bundle sync — postTurnPush discovers dirty files via `git status --porcelain` instead of broken explicit trackSyncWrite tracking (only 3/9 writers used it). Removed dead infrastructure. Remote push optional and non-fatal.",
3459
- "feat(tui): queued input steer — messages typed while agent is thinking appear dimmed above input area. UP/ESC pops all queued messages back into input for editing. Placeholder hint. Multiple messages supported, each sent as separate turn.",
3472
+ "refactor(sync): git-status-based bundle sync \u2014 postTurnPush discovers dirty files via `git status --porcelain` instead of broken explicit trackSyncWrite tracking (only 3/9 writers used it). Removed dead infrastructure. Remote push optional and non-fatal.",
3473
+ "feat(tui): queued input steer \u2014 messages typed while agent is thinking appear dimmed above input area. UP/ESC pops all queued messages back into input for editing. Placeholder hint. Multiple messages supported, each sent as separate turn.",
3460
3474
  "chore: deleted legacy InkCliApp adapter (776 lines dead code)."
3461
3475
  ]
3462
3476
  },
3463
3477
  {
3464
3478
  "version": "0.1.0-alpha.243",
3465
3479
  "changes": [
3466
- "fix(daemon): per-agent Git Sync in `ouro status` — was always showing `disabled` regardless of any agent's `agent.json`, because the daemon process has no argv-derived agent identity so `getSyncConfig()` fell into its catch. Status payload now carries a per-agent `sync` array (one row per enabled bundle) rendered as its own section like Senses and Workers, instead of a single global field on the overview block."
3480
+ "fix(daemon): per-agent Git Sync in `ouro status` \u2014 was always showing `disabled` regardless of any agent's `agent.json`, because the daemon process has no argv-derived agent identity so `getSyncConfig()` fell into its catch. Status payload now carries a per-agent `sync` array (one row per enabled bundle) rendered as its own section like Senses and Workers, instead of a single global field on the overview block."
3467
3481
  ]
3468
3482
  },
3469
3483
  {
3470
3484
  "version": "0.1.0-alpha.241",
3471
3485
  "changes": [
3472
- "feat: commerce bootstrap — bw CLI lazy-install, vault auto-config, resolver coverage"
3486
+ "feat: commerce bootstrap \u2014 bw CLI lazy-install, vault auto-config, resolver coverage"
3473
3487
  ]
3474
3488
  },
3475
3489
  {
3476
3490
  "version": "0.1.0-alpha.242",
3477
3491
  "changes": [
3478
- "fix(friends): stable local CLI identity — dropped hostname from external ID (was `username@hostname`, now just `username`). macOS hostname instability (`Mac` vs `Aris-MacBook-Pro.local`) was creating duplicate friend records with separate sessions and trust levels.",
3479
- "feat(friends): migration fallback — FriendResolver now searches for old `username@*` format IDs when exact match fails, linking new stable ID to existing friend record.",
3480
- "fix(heart): retry-everything-except-blocklist policy — the SDK 'Request timed out.' error from MiniMax (and other providers) was reaching the agent as a terminal failure because neither the generic isTransientError detector nor the per-provider classifier recognized it. Replaced the two-layer transient detection with a small blocklist (HTTP 400/401/403/404/422 + classifications auth-failure/usage-limit). Default policy now retries every other error.",
3481
- "fix(heart/providers): drop the 30s OpenAI/Anthropic SDK timeout from all five providers (anthropic, azure, github-copilot, minimax, openai-codex). The SDK timeout caps the entire stream lifetime, so 30s killed any reasoning model mid-generation. SDK defaults (≈10min) are sane.",
3492
+ "fix(friends): stable local CLI identity \u2014 dropped hostname from external ID (was `username@hostname`, now just `username`). macOS hostname instability (`Mac` vs `Aris-MacBook-Pro.local`) was creating duplicate friend records with separate sessions and trust levels.",
3493
+ "feat(friends): migration fallback \u2014 FriendResolver now searches for old `username@*` format IDs when exact match fails, linking new stable ID to existing friend record.",
3494
+ "fix(heart): retry-everything-except-blocklist policy \u2014 the SDK 'Request timed out.' error from MiniMax (and other providers) was reaching the agent as a terminal failure because neither the generic isTransientError detector nor the per-provider classifier recognized it. Replaced the two-layer transient detection with a small blocklist (HTTP 400/401/403/404/422 + classifications auth-failure/usage-limit). Default policy now retries every other error.",
3495
+ "fix(heart/providers): drop the 30s OpenAI/Anthropic SDK timeout from all five providers (anthropic, azure, github-copilot, minimax, openai-codex). The SDK timeout caps the entire stream lifetime, so 30s killed any reasoning model mid-generation. SDK defaults (\u224810min) are sane.",
3482
3496
  "refactor(heart/providers): consolidate duplicated isNetworkError + classifyXxxError scaffolding into a shared `error-classification.ts` module. Each provider now delegates via `classifyHttpError(err, overrides)` and only carries its own provider-specific quirks (Anthropic 529, Codex usage-limit message detection)."
3483
3497
  ]
3484
3498
  },
3485
3499
  {
3486
3500
  "version": "0.1.0-alpha.238",
3487
3501
  "changes": [
3488
- "feat: pretty `ouro status` — ANSI colored output with box-drawing header, status dots, grouped senses/workers by agent. Added git sync info to overview.",
3489
- "fix: socket half-close — sendDaemonCommand now calls client.end() after writing, preventing intermittent connection hangs.",
3490
- "refactor: config tiers — replaced numeric T1/T2/T3 with `self` (agent-configurable) and `managed` (harness-only). All config keys are now agent-writable except `version` and `enabled`. mcpServers promoted to self.",
3491
- "refactor: removed confirmation system — deleted propose_config tool, confirmationRequired/confirmationAlwaysRequired/onConfirmAction from core, all tool definitions, and Teams sense. Was only wired up on Teams, silently failed everywhere else. -1,361 lines."
3502
+ "feat: pretty `ouro status` \u2014 ANSI colored output with box-drawing header, status dots, grouped senses/workers by agent. Added git sync info to overview.",
3503
+ "fix: socket half-close \u2014 sendDaemonCommand now calls client.end() after writing, preventing intermittent connection hangs.",
3504
+ "refactor: config tiers \u2014 replaced numeric T1/T2/T3 with `self` (agent-configurable) and `managed` (harness-only). All config keys are now agent-writable except `version` and `enabled`. mcpServers promoted to self.",
3505
+ "refactor: removed confirmation system \u2014 deleted propose_config tool, confirmationRequired/confirmationAlwaysRequired/onConfirmAction from core, all tool definitions, and Teams sense. Was only wired up on Teams, silently failed everywhere else. -1,361 lines."
3492
3506
  ]
3493
3507
  },
3494
3508
  {
3495
3509
  "version": "0.1.0-alpha.235",
3496
3510
  "changes": [
3497
- "fix: MCP tool double-prefix — tools already prefixed by server name no longer get a redundant second prefix in the unified registry.",
3498
- "feat: Open-Meteo zero-auth weather — replaced OpenWeatherMap with Open-Meteo forecast + geocoding API. Weather now works without any API key or credential provisioning.",
3499
- "feat: expanded ISO/FIPS divergence table — 21 new entries for correct travel advisory resolution across all major divergent country codes.",
3500
- "fix: BitwardenCredentialStore retry logic — exponential backoff with configurable retries for transient bw CLI failures, plus bw-not-installed error.",
3511
+ "fix: MCP tool double-prefix \u2014 tools already prefixed by server name no longer get a redundant second prefix in the unified registry.",
3512
+ "feat: Open-Meteo zero-auth weather \u2014 replaced OpenWeatherMap with Open-Meteo forecast + geocoding API. Weather now works without any API key or credential provisioning.",
3513
+ "feat: expanded ISO/FIPS divergence table \u2014 21 new entries for correct travel advisory resolution across all major divergent country codes.",
3514
+ "fix: BitwardenCredentialStore retry logic \u2014 exponential backoff with configurable retries for transient bw CLI failures, plus bw-not-installed error.",
3501
3515
  "chore: travel MCP packages (Duffel, Expedia) status confirmed GitHub-only, not published to npm."
3502
3516
  ]
3503
3517
  },
3504
3518
  {
3505
3519
  "version": "0.1.0-alpha.233",
3506
3520
  "changes": [
3507
- "feat: first-class MCP tools — MCP tools now appear in the agent tool list directly (no shell indirection). Agent can call browser_navigate, browser_click etc. as native tools.",
3508
- "fix: daemon MCP pre-init poisoned singleton — removed eager getSharedMcpManager() at daemon startup that cached null before agent identity was set.",
3509
- "fix: removed dead mcpManager field from BuildSystemOptions — MCP manager now flows through runAgentOptions.",
3521
+ "feat: first-class MCP tools \u2014 MCP tools now appear in the agent tool list directly (no shell indirection). Agent can call browser_navigate, browser_click etc. as native tools.",
3522
+ "fix: daemon MCP pre-init poisoned singleton \u2014 removed eager getSharedMcpManager() at daemon startup that cached null before agent identity was set.",
3523
+ "fix: removed dead mcpManager field from BuildSystemOptions \u2014 MCP manager now flows through runAgentOptions.",
3510
3524
  "fix: MCP tool results filter to text-only content types.",
3511
3525
  "includes: vault integration, travel advisory fix, credential access layer, HKDF-Expand crypto fix."
3512
3526
  ]
@@ -3514,21 +3528,21 @@
3514
3528
  {
3515
3529
  "version": "0.1.0-alpha.232",
3516
3530
  "changes": [
3517
- "feat: first-class MCP tools — MCP server tools now appear in the agent's active tool list (e.g. browser_navigate, duffel_search_flights) and are callable directly by the model, eliminating fragile shell indirection",
3531
+ "feat: first-class MCP tools \u2014 MCP server tools now appear in the agent's active tool list (e.g. browser_navigate, duffel_search_flights) and are callable directly by the model, eliminating fragile shell indirection",
3518
3532
  "feat: mcpToolsAsDefinitions() converts McpManager tools to ToolDefinition objects with {server}_{tool} naming",
3519
- "feat: first-class MCP trust gating — mcpServerName on GuardContext enables per-server trust rules (browser blocked for acquaintance, blocked in group chat)",
3520
- "refactor: removed mcpToolsSection() from system prompt — MCP tools no longer need prompt documentation",
3533
+ "feat: first-class MCP trust gating \u2014 mcpServerName on GuardContext enables per-server trust rules (browser blocked for acquaintance, blocked in group chat)",
3534
+ "refactor: removed mcpToolsSection() from system prompt \u2014 MCP tools no longer need prompt documentation",
3521
3535
  "fix: execTool, isConfirmationRequired, summarizeArgs now check combined native+MCP registry"
3522
3536
  ]
3523
3537
  },
3524
3538
  {
3525
3539
  "version": "0.1.0-alpha.231",
3526
3540
  "changes": [
3527
- "fix: defense-in-depth group chat blocking for proactive BB delivery — sendProactiveBlueBubblesMessageToSession now rejects group chat keys (;+;) unless intent is explicit_cross_chat (bridge/delegation responses). All upper-layer paths (surface tool, send_message tool, inner-dialog delegation) now filter to DM sessions only for proactive outreach, and pass explicit_cross_chat intent for bridge/delegation returns where group responses are legitimate.",
3541
+ "fix: defense-in-depth group chat blocking for proactive BB delivery \u2014 sendProactiveBlueBubblesMessageToSession now rejects group chat keys (;+;) unless intent is explicit_cross_chat (bridge/delegation responses). All upper-layer paths (surface tool, send_message tool, inner-dialog delegation) now filter to DM sessions only for proactive outreach, and pass explicit_cross_chat intent for bridge/delegation returns where group responses are legitimate.",
3528
3542
  "fix: travel advisory now resolves ISO codes that differ from FIPS (ES -> Spain, not El Salvador)",
3529
3543
  "fix: MCP bridge sense now includes MCP tool descriptions in system prompt (browser tools visible)",
3530
3544
  "fix: removed confirmationRequired from credential_store/credential_delete (trust gating sufficient)",
3531
- "feat: vault integration — Bitwarden/Vaultwarden account creation with PBKDF2/HKDF/AES-256-CBC crypto",
3545
+ "feat: vault integration \u2014 Bitwarden/Vaultwarden account creation with PBKDF2/HKDF/AES-256-CBC crypto",
3532
3546
  "feat: vault_setup tool for one-time vault provisioning (family trust gated)",
3533
3547
  "feat: BitwardenCredentialStore wrapping bw CLI for agent-owned vault access"
3534
3548
  ]
@@ -3542,20 +3556,20 @@
3542
3556
  {
3543
3557
  "version": "0.1.0-alpha.229",
3544
3558
  "changes": [
3545
- "fix: BB session key resolution prefers DM (;-;) over group chat (;+;) — alphabetical sort put group chats first, causing proactive messages to land in group chats instead of personal DMs",
3559
+ "fix: BB session key resolution prefers DM (;-;) over group chat (;+;) \u2014 alphabetical sort put group chats first, causing proactive messages to land in group chats instead of personal DMs",
3546
3560
  "chore: remove temporary debug traces (send-message-debug.log, friends.get_called event)"
3547
3561
  ]
3548
3562
  },
3549
3563
  {
3550
3564
  "version": "0.1.0-alpha.227",
3551
3565
  "changes": [
3552
- "fix: direct filesystem name resolution in sendProactiveBlueBubblesMessageToSession — bypass store.get()/listAll() with raw fs reads on friends directory when store lookup fails"
3566
+ "fix: direct filesystem name resolution in sendProactiveBlueBubblesMessageToSession \u2014 bypass store.get()/listAll() with raw fs reads on friends directory when store lookup fails"
3553
3567
  ]
3554
3568
  },
3555
3569
  {
3556
3570
  "version": "0.1.0-alpha.226",
3557
3571
  "changes": [
3558
- "fix(daemon): set agent name before senseTurn so MCP messages resolve identity — setAgentName() is now called at the top of the senseTurn handler, before any downstream code that depends on agent identity (loadAgentConfig, getAgentSecretsPath, etc.)"
3572
+ "fix(daemon): set agent name before senseTurn so MCP messages resolve identity \u2014 setAgentName() is now called at the top of the senseTurn handler, before any downstream code that depends on agent identity (loadAgentConfig, getAgentSecretsPath, etc.)"
3559
3573
  ]
3560
3574
  },
3561
3575
  {
@@ -3567,7 +3581,7 @@
3567
3581
  {
3568
3582
  "version": "0.1.0-alpha.224",
3569
3583
  "changes": [
3570
- "fix: FileFriendStore.get() now resolves friend names — when UUID lookup fails, scans the friends directory for a name match. This is the deepest possible layer for name resolution, ensuring it works regardless of which tool or code path calls store.get()."
3584
+ "fix: FileFriendStore.get() now resolves friend names \u2014 when UUID lookup fails, scans the friends directory for a name match. This is the deepest possible layer for name resolution, ensuring it works regardless of which tool or code path calls store.get()."
3571
3585
  ]
3572
3586
  },
3573
3587
  {
@@ -3579,127 +3593,127 @@
3579
3593
  {
3580
3594
  "version": "0.1.0-alpha.222",
3581
3595
  "changes": [
3582
- "debug: add diagnostic nerves events to proactive BB delivery — name resolution in both tools-session.ts and sendProactiveBlueBubblesMessageToSession now emit events showing friend count, names, resolution success/failure, and errors. Temporary diagnostics to identify why name→UUID resolution isn't working in production."
3596
+ "debug: add diagnostic nerves events to proactive BB delivery \u2014 name resolution in both tools-session.ts and sendProactiveBlueBubblesMessageToSession now emit events showing friend count, names, resolution success/failure, and errors. Temporary diagnostics to identify why name\u2192UUID resolution isn't working in production."
3583
3597
  ]
3584
3598
  },
3585
3599
  {
3586
3600
  "version": "0.1.0-alpha.221",
3587
3601
  "changes": [
3588
- "fix: BB proactive send resolves friend by name when UUID lookup fails — sendProactiveBlueBubblesMessageToSession now falls back to store.listAll() name matching when store.get() returns null. This handles agents passing friend names instead of UUIDs, bypassing the upstream resolution that wasn't working in all contexts."
3602
+ "fix: BB proactive send resolves friend by name when UUID lookup fails \u2014 sendProactiveBlueBubblesMessageToSession now falls back to store.listAll() name matching when store.get() returns null. This handles agents passing friend names instead of UUIDs, bypassing the upstream resolution that wasn't working in all contexts."
3589
3603
  ]
3590
3604
  },
3591
3605
  {
3592
3606
  "version": "0.1.0-alpha.220",
3593
3607
  "changes": [
3594
- "fix: send_message BB session key resolution — agents don't know the real BB session key (e.g. 'chat_any;-;you@example.com'), so they pass the default 'session'. buildChatRefForSessionKey failed on this fake key, returning missing_target. Now auto-resolves the real BB session key from the sessions directory when the default key is used."
3608
+ "fix: send_message BB session key resolution \u2014 agents don't know the real BB session key (e.g. 'chat_any;-;you@example.com'), so they pass the default 'session'. buildChatRefForSessionKey failed on this fake key, returning missing_target. Now auto-resolves the real BB session key from the sessions directory when the default key is used."
3595
3609
  ]
3596
3610
  },
3597
3611
  {
3598
3612
  "version": "0.1.0-alpha.219",
3599
3613
  "changes": [
3600
- "fix: proactive message delivery — three bugs fixed. (1) surface tool now resolves friend names to UUIDs by scanning friends directory. (2) send_message tool also resolves friend names to UUIDs. (3) deliverCrossChatMessage no longer immediately queues generic_outreach — it now attempts delivery when a deliverer is available, with the deliverer's own trust checks still gating actual sends. Previously, any proactive send from inner dialog was silently queued without attempting delivery."
3614
+ "fix: proactive message delivery \u2014 three bugs fixed. (1) surface tool now resolves friend names to UUIDs by scanning friends directory. (2) send_message tool also resolves friend names to UUIDs. (3) deliverCrossChatMessage no longer immediately queues generic_outreach \u2014 it now attempts delivery when a deliverer is available, with the deliverer's own trust checks still gating actual sends. Previously, any proactive send from inner dialog was silently queued without attempting delivery."
3601
3615
  ]
3602
3616
  },
3603
3617
  {
3604
3618
  "version": "0.1.0-alpha.218",
3605
3619
  "changes": [
3606
- "fix: surface tool now resolves friend names to UUIDs — agents pass friend names but sessions are stored under UUID directories. Added name-to-UUID resolution by scanning the friends directory when the friendId doesn't match a session directory."
3620
+ "fix: surface tool now resolves friend names to UUIDs \u2014 agents pass friend names but sessions are stored under UUID directories. Added name-to-UUID resolution by scanning the friends directory when the friendId doesn't match a session directory."
3607
3621
  ]
3608
3622
  },
3609
3623
  {
3610
3624
  "version": "0.1.0-alpha.217",
3611
3625
  "changes": [
3612
- "feat: Bitwarden vault client (bw CLI wrapper) with credential gateway — BitwardenClient class with SDK-first/CLI-fallback, singleton accessor, nerves events on all operations. Raw secrets never enter model context.",
3613
- "feat: vault tools (vault_get, vault_store, vault_list, vault_delete) with trust gating — read ops require friend+, write/delete require family-only. Destructive operations require confirmation.",
3614
- "feat: stealth browser MCP configuration with trust gating — Playwright MCP auto-provisions when configured, browser tools appear in agent tool list. Trust-gated to CLI and trusted 1:1 only (group chat blocked).",
3615
- "feat: travel API tools (weather_lookup, travel_advisory, geocode_search) — native weather via OpenWeatherMap + vault-backed API key, State Dept RSS feed for advisories, geocoding via Nominatim. All friend+ trust-gated.",
3616
- "feat: credential gateway (vaultKey on apiRequest()) — automatic secret injection from vault into HTTP headers at call time, keeping credentials out of model context entirely.",
3626
+ "feat: Bitwarden vault client (bw CLI wrapper) with credential gateway \u2014 BitwardenClient class with SDK-first/CLI-fallback, singleton accessor, nerves events on all operations. Raw secrets never enter model context.",
3627
+ "feat: vault tools (vault_get, vault_store, vault_list, vault_delete) with trust gating \u2014 read ops require friend+, write/delete require family-only. Destructive operations require confirmation.",
3628
+ "feat: stealth browser MCP configuration with trust gating \u2014 Playwright MCP auto-provisions when configured, browser tools appear in agent tool list. Trust-gated to CLI and trusted 1:1 only (group chat blocked).",
3629
+ "feat: travel API tools (weather_lookup, travel_advisory, geocode_search) \u2014 native weather via OpenWeatherMap + vault-backed API key, State Dept RSS feed for advisories, geocoding via Nominatim. All friend+ trust-gated.",
3630
+ "feat: credential gateway (vaultKey on apiRequest()) \u2014 automatic secret injection from vault into HTTP headers at call time, keeping credentials out of model context entirely.",
3617
3631
  "feat: travel-planning and browser-navigation skills"
3618
3632
  ]
3619
3633
  },
3620
3634
  {
3621
3635
  "version": "0.1.0-alpha.216",
3622
3636
  "changes": [
3623
- "fix: surface tool proactive BB delivery masked by newer MCP/CLI sessions — findFreshestFriendSession picked the single freshest session regardless of channel, so an MCP or CLI session being newer than the BB session caused the BB proactive path to be skipped entirely. Now scans all friend sessions, attempts BB delivery first on any BB session, then falls back to queuing on the freshest non-inner session."
3637
+ "fix: surface tool proactive BB delivery masked by newer MCP/CLI sessions \u2014 findFreshestFriendSession picked the single freshest session regardless of channel, so an MCP or CLI session being newer than the BB session caused the BB proactive path to be skipped entirely. Now scans all friend sessions, attempts BB delivery first on any BB session, then falls back to queuing on the freshest non-inner session."
3624
3638
  ]
3625
3639
  },
3626
3640
  {
3627
3641
  "version": "0.1.0-alpha.215",
3628
3642
  "changes": [
3629
- "fix: surface tool proactive delivery no longer gated by 24-hour session threshold — findFreshestFriendSession was called with activeOnly:true, filtering out sessions older than 24h even when the agent explicitly wants to send a proactive message. Both the bridge path and direct path now find any session regardless of age. Proactive BB delivery and trust checks still apply."
3643
+ "fix: surface tool proactive delivery no longer gated by 24-hour session threshold \u2014 findFreshestFriendSession was called with activeOnly:true, filtering out sessions older than 24h even when the agent explicitly wants to send a proactive message. Both the bridge path and direct path now find any session regardless of age. Proactive BB delivery and trust checks still apply."
3630
3644
  ]
3631
3645
  },
3632
3646
  {
3633
3647
  "version": "0.1.0-alpha.214",
3634
3648
  "changes": [
3635
- "fix: daemon death diagnostics — createStderrSink() bypassed EPIPE-safe default in createTerminalSink(), causing uncaught EPIPE crashes when daemon runs detached. Removed redundant unsafe default so the existing try-catch fires.",
3636
- "fix: daemon tombstone now covers all exit paths — unhandledRejection writes tombstone with full error+stack (was just a warn log, but Node 15+ terminates on these). Added process.on('exit') catch-all for any unanticipated exit. SIGINT/SIGTERM marked graceful to avoid false positives. _lastKnownCause threads real error through to exit handler."
3649
+ "fix: daemon death diagnostics \u2014 createStderrSink() bypassed EPIPE-safe default in createTerminalSink(), causing uncaught EPIPE crashes when daemon runs detached. Removed redundant unsafe default so the existing try-catch fires.",
3650
+ "fix: daemon tombstone now covers all exit paths \u2014 unhandledRejection writes tombstone with full error+stack (was just a warn log, but Node 15+ terminates on these). Added process.on('exit') catch-all for any unanticipated exit. SIGINT/SIGTERM marked graceful to avoid false positives. _lastKnownCause threads real error through to exit handler."
3637
3651
  ]
3638
3652
  },
3639
3653
  {
3640
3654
  "version": "0.1.0-alpha.213",
3641
3655
  "changes": [
3642
- "cleanup: remove vestigial subagents/ directory and package.json files entry (content already in ouroboros-skills repo). Remove backward-compat re-exports from heart/core.ts (tools, execTool, summarizeArgs, getToolsForChannel, streamChatCompletion, streamResponsesApi, toResponsesInput, toResponsesTools, buildSystem, Channel, hasToolIntent — no consumers used them). Update ARCHITECTURE.md, README.md, and CONTRIBUTING.md to reflect the full audit restructuring: new arc/ subsystem, heart/ topic subdirectories, split tool modules, BlueBubbles directory, scopes list."
3656
+ "cleanup: remove vestigial subagents/ directory and package.json files entry (content already in ouroboros-skills repo). Remove backward-compat re-exports from heart/core.ts (tools, execTool, summarizeArgs, getToolsForChannel, streamChatCompletion, streamResponsesApi, toResponsesInput, toResponsesTools, buildSystem, Channel, hasToolIntent \u2014 no consumers used them). Update ARCHITECTURE.md, README.md, and CONTRIBUTING.md to reflect the full audit restructuring: new arc/ subsystem, heart/ topic subdirectories, split tool modules, BlueBubbles directory, scopes list."
3643
3657
  ]
3644
3658
  },
3645
3659
  {
3646
3660
  "version": "0.1.0-alpha.212",
3647
3661
  "changes": [
3648
- "refactor: consolidate BlueBubbles sense into senses/bluebubbles/ directory — move 9 flat files (bluebubbles.ts, bluebubbles-client.ts, bluebubbles-model.ts, bluebubbles-media.ts, bluebubbles-inbound-log.ts, bluebubbles-mutation-log.ts, bluebubbles-runtime-state.ts, bluebubbles-session-cleanup.ts, bluebubbles-entry.ts) into senses/bluebubbles/ with shorter names (index.ts, client.ts, model.ts, etc.). All imports updated across 20+ files including test files, sense-manager, daemon, and package.json."
3662
+ "refactor: consolidate BlueBubbles sense into senses/bluebubbles/ directory \u2014 move 9 flat files (bluebubbles.ts, bluebubbles-client.ts, bluebubbles-model.ts, bluebubbles-media.ts, bluebubbles-inbound-log.ts, bluebubbles-mutation-log.ts, bluebubbles-runtime-state.ts, bluebubbles-session-cleanup.ts, bluebubbles-entry.ts) into senses/bluebubbles/ with shorter names (index.ts, client.ts, model.ts, etc.). All imports updated across 20+ files including test files, sense-manager, daemon, and package.json."
3649
3663
  ]
3650
3664
  },
3651
3665
  {
3652
3666
  "version": "0.1.0-alpha.211",
3653
3667
  "changes": [
3654
- "refactor: extract duplicated patterns into shared utilities — mind/embedding-provider.ts (shared OpenAI embedding client from diary + associative-recall), arc/json-store.ts (shared JSON file CRUD from obligations + cares + intentions), repertoire/api-client.ts (shared HTTP request helper from graph + ado + github clients). M12 (channel callback factory) skipped: CLI and Teams streaming implementations are too different for clean abstraction."
3668
+ "refactor: extract duplicated patterns into shared utilities \u2014 mind/embedding-provider.ts (shared OpenAI embedding client from diary + associative-recall), arc/json-store.ts (shared JSON file CRUD from obligations + cares + intentions), repertoire/api-client.ts (shared HTTP request helper from graph + ado + github clients). M12 (channel callback factory) skipped: CLI and Teams streaming implementations are too different for clean abstraction."
3655
3669
  ]
3656
3670
  },
3657
3671
  {
3658
3672
  "version": "0.1.0-alpha.210",
3659
3673
  "changes": [
3660
- "refactor: create src/arc/ subsystem — extract durable continuity state (obligations, cares, episodes, intentions, presence, attention-types) from heart/ and mind/ into dedicated arc/ module. arc/ owns the agent's continuity state, distinct from engine mechanics (heart) and cognition (mind). All imports updated across 40+ files."
3674
+ "refactor: create src/arc/ subsystem \u2014 extract durable continuity state (obligations, cares, episodes, intentions, presence, attention-types) from heart/ and mind/ into dedicated arc/ module. arc/ owns the agent's continuity state, distinct from engine mechanics (heart) and cognition (mind). All imports updated across 40+ files."
3661
3675
  ]
3662
3676
  },
3663
3677
  {
3664
3678
  "version": "0.1.0-alpha.209",
3665
3679
  "changes": [
3666
- "refactor: restructure daemon/ directory — move outlook files to heart/outlook/, habit files to heart/habits/, hatch/specialist files to heart/hatch/, versioning/update files to heart/versioning/, auth-flow to heart/auth/, mcp-server to heart/mcp/. daemon/ reduced from 60 to 36 core daemon-lifecycle files."
3680
+ "refactor: restructure daemon/ directory \u2014 move outlook files to heart/outlook/, habit files to heart/habits/, hatch/specialist files to heart/hatch/, versioning/update files to heart/versioning/, auth-flow to heart/auth/, mcp-server to heart/mcp/. daemon/ reduced from 60 to 36 core daemon-lifecycle files."
3667
3681
  ]
3668
3682
  },
3669
3683
  {
3670
3684
  "version": "0.1.0-alpha.208",
3671
3685
  "changes": [
3672
- "refactor: split daemon-cli.ts (3,630 lines) into 5 focused modules — cli-types (command/deps types), cli-parse (argument parsing), cli-render (output formatting), cli-exec (command execution router), cli-defaults (production dependency wiring). daemon-cli.ts reduced to 42-line re-export shim."
3686
+ "refactor: split daemon-cli.ts (3,630 lines) into 5 focused modules \u2014 cli-types (command/deps types), cli-parse (argument parsing), cli-render (output formatting), cli-exec (command execution router), cli-defaults (production dependency wiring). daemon-cli.ts reduced to 42-line re-export shim."
3673
3687
  ]
3674
3688
  },
3675
3689
  {
3676
3690
  "version": "0.1.0-alpha.207",
3677
3691
  "changes": [
3678
- "refactor: split tools-base.ts (1,912 lines) into 9 category modules — tools-files, tools-shell, tools-memory, tools-bridge, tools-session, tools-continuity, tools-flow, tools-surface, tools-config. Surface tool handler extracted from tools.ts to tools-surface.ts."
3692
+ "refactor: split tools-base.ts (1,912 lines) into 9 category modules \u2014 tools-files, tools-shell, tools-memory, tools-bridge, tools-session, tools-continuity, tools-flow, tools-surface, tools-config. Surface tool handler extracted from tools.ts to tools-surface.ts."
3679
3693
  ]
3680
3694
  },
3681
3695
  {
3682
3696
  "version": "0.1.0-alpha.206",
3683
3697
  "changes": [
3684
- "feat: capability discovery and tiered self-configuration — config registry with tier-aware metadata (T1 self-service, T2 proposal, T3 operator-only), read_config tool with topic-filtered discovery, update_config tool for T1 immediate changes, propose_config tool for T2 operator-approval flow, version-change surfacing in start-of-turn packet via buildCapabilitiesSection"
3698
+ "feat: capability discovery and tiered self-configuration \u2014 config registry with tier-aware metadata (T1 self-service, T2 proposal, T3 operator-only), read_config tool with topic-filtered discovery, update_config tool for T1 immediate changes, propose_config tool for T2 operator-approval flow, version-change surfacing in start-of-turn packet via buildCapabilitiesSection"
3685
3699
  ]
3686
3700
  },
3687
3701
  {
3688
3702
  "version": "0.1.0-alpha.205",
3689
3703
  "changes": [
3690
- "refactor: enforce subsystem boundaries — eliminate all heart/ and nerves/ static imports from senses/, move AttentionItem type to heart/, surfaceToolDef to repertoire/, SteeringFollowUpEffect to heart/turn-coordinator, inline BlueBubbles runtime state reader in daemon"
3704
+ "refactor: enforce subsystem boundaries \u2014 eliminate all heart/ and nerves/ static imports from senses/, move AttentionItem type to heart/, surfaceToolDef to repertoire/, SteeringFollowUpEffect to heart/turn-coordinator, inline BlueBubbles runtime state reader in daemon"
3691
3705
  ]
3692
3706
  },
3693
3707
  {
3694
3708
  "version": "0.1.0-alpha.204",
3695
3709
  "changes": [
3696
- "refactor: introduce TurnContext snapshot — centralize state assembly from pipeline.ts into buildTurnContext(), thread pre-read state through prompt assembly to eliminate ad-hoc filesystem reads"
3710
+ "refactor: introduce TurnContext snapshot \u2014 centralize state assembly from pipeline.ts into buildTurnContext(), thread pre-read state through prompt assembly to eliminate ad-hoc filesystem reads"
3697
3711
  ]
3698
3712
  },
3699
3713
  {
3700
3714
  "version": "0.1.0-alpha.203",
3701
3715
  "changes": [
3702
- "refactor: unify obligation systems — mind/obligations.ts merged into heart/obligations.ts with prefixed ReturnObligation API"
3716
+ "refactor: unify obligation systems \u2014 mind/obligations.ts merged into heart/obligations.ts with prefixed ReturnObligation API"
3703
3717
  ]
3704
3718
  },
3705
3719
  {
@@ -3715,7 +3729,7 @@
3715
3729
  {
3716
3730
  "version": "0.1.0-alpha.201",
3717
3731
  "changes": [
3718
- "fix: don't launchctl bootstrap after daemon start — was starting competing daemon that killed the first"
3732
+ "fix: don't launchctl bootstrap after daemon start \u2014 was starting competing daemon that killed the first"
3719
3733
  ]
3720
3734
  },
3721
3735
  {
@@ -3737,7 +3751,7 @@
3737
3751
  {
3738
3752
  "version": "0.1.0-alpha.197",
3739
3753
  "changes": [
3740
- "fix(daemon): validate agent config before spawn — skips agents with missing credentials instead of crash-looping"
3754
+ "fix(daemon): validate agent config before spawn \u2014 skips agents with missing credentials instead of crash-looping"
3741
3755
  ]
3742
3756
  },
3743
3757
  {
@@ -3756,20 +3770,20 @@
3756
3770
  {
3757
3771
  "version": "0.1.0-alpha.194",
3758
3772
  "changes": [
3759
- "fix(daemon): self-spawn restart — no longer relies on launchd KeepAlive for staged restarts",
3760
- "fix(daemon): error boundary with circuit breaker — uncaught exceptions logged and survived, exits only after 10+ in 60s",
3773
+ "fix(daemon): self-spawn restart \u2014 no longer relies on launchd KeepAlive for staged restarts",
3774
+ "fix(daemon): error boundary with circuit breaker \u2014 uncaught exceptions logged and survived, exits only after 10+ in 60s",
3761
3775
  "fix(daemon): EPIPE suppression in uncaughtException handler",
3762
3776
  "fix(daemon): 5-second force-exit timeouts on all shutdown paths",
3763
3777
  "feat: human-facing and agent-facing provider configs",
3764
3778
  "fix(auth): always refresh codex OAuth token, responses API verification",
3765
- "feat: Outlook visibility — orientation, obligations, changes, self-fix, memory decisions, route migration to /"
3779
+ "feat: Outlook visibility \u2014 orientation, obligations, changes, self-fix, memory decisions, route migration to /"
3766
3780
  ]
3767
3781
  },
3768
3782
  {
3769
3783
  "version": "0.1.0-alpha.192",
3770
3784
  "changes": [
3771
- "refactor: canonical obligations — ActiveWorkFrame as single source of truth for prompt sections",
3772
- "feat(mcp): dynamic server add/remove — agents can manage MCP servers without daemon restart"
3785
+ "refactor: canonical obligations \u2014 ActiveWorkFrame as single source of truth for prompt sections",
3786
+ "feat(mcp): dynamic server add/remove \u2014 agents can manage MCP servers without daemon restart"
3773
3787
  ]
3774
3788
  },
3775
3789
  {
@@ -3781,13 +3795,13 @@
3781
3795
  {
3782
3796
  "version": "0.1.0-alpha.176",
3783
3797
  "changes": [
3784
- "feat(mcp): dynamic MCP server add/remove — agents can add/remove MCP servers in agent.json without daemon restart"
3798
+ "feat(mcp): dynamic MCP server add/remove \u2014 agents can add/remove MCP servers in agent.json without daemon restart"
3785
3799
  ]
3786
3800
  },
3787
3801
  {
3788
3802
  "version": "0.1.0-alpha.175",
3789
3803
  "changes": [
3790
- "fix(daemon): launchd KeepAlive for crash recovery — auto-restarts on crash",
3804
+ "fix(daemon): launchd KeepAlive for crash recovery \u2014 auto-restarts on crash",
3791
3805
  "fix(daemon): orphan killer excludes MCP server processes",
3792
3806
  "fix(daemon): health file writer wired into daemon-entry",
3793
3807
  "fix(engine): auth-failure errors include actionable guidance"
@@ -3796,50 +3810,50 @@
3796
3810
  {
3797
3811
  "version": "0.1.0-alpha.174",
3798
3812
  "changes": [
3799
- "feat(outlook): keyboard shortcuts — 1-7 for tabs, Esc to collapse",
3800
- "feat(outlook): obligation origin cards — clickable visual chain from who asked through which channel"
3813
+ "feat(outlook): keyboard shortcuts \u2014 1-7 for tabs, Esc to collapse",
3814
+ "feat(outlook): obligation origin cards \u2014 clickable visual chain from who asked through which channel"
3801
3815
  ]
3802
3816
  },
3803
3817
  {
3804
3818
  "version": "0.1.0-alpha.173",
3805
3819
  "changes": [
3806
- "feat(outlook): sessions grouped by person — same friend across multiple channels shown together with person header"
3820
+ "feat(outlook): sessions grouped by person \u2014 same friend across multiple channels shown together with person header"
3807
3821
  ]
3808
3822
  },
3809
3823
  {
3810
3824
  "version": "0.1.0-alpha.172",
3811
3825
  "changes": [
3812
- "feat(outlook): inner dialog landmark navigation — jump to surfaces, rests, delegations",
3826
+ "feat(outlook): inner dialog landmark navigation \u2014 jump to surfaces, rests, delegations",
3813
3827
  "feat(outlook): active coding sessions shown on Overview dashboard",
3814
- "feat(outlook): habit confidence indicators — on schedule, overdue, never fired"
3828
+ "feat(outlook): habit confidence indicators \u2014 on schedule, overdue, never fired"
3815
3829
  ]
3816
3830
  },
3817
3831
  {
3818
3832
  "version": "0.1.0-alpha.171",
3819
3833
  "changes": [
3820
- "feat(outlook): session state at a glance — last inbound/outbound shown on each session row",
3821
- "feat(outlook): habit confidence — on schedule / overdue / never fired indicators"
3834
+ "feat(outlook): session state at a glance \u2014 last inbound/outbound shown on each session row",
3835
+ "feat(outlook): habit confidence \u2014 on schedule / overdue / never fired indicators"
3822
3836
  ]
3823
3837
  },
3824
3838
  {
3825
3839
  "version": "0.1.0-alpha.170",
3826
3840
  "changes": [
3827
- "feat(outlook): needs-me triage — action now vs stale sections, dismiss buttons, return-ready highlighting",
3828
- "fix(outlook): return-ready obligation detection — highlights results ready but not returned"
3841
+ "feat(outlook): needs-me triage \u2014 action now vs stale sections, dismiss buttons, return-ready highlighting",
3842
+ "fix(outlook): return-ready obligation detection \u2014 highlights results ready but not returned"
3829
3843
  ]
3830
3844
  },
3831
3845
  {
3832
3846
  "version": "0.1.0-alpha.169",
3833
3847
  "changes": [
3834
- "fix(outlook): desk prefs wiring — carrying block, constellations, starred friends now load in production",
3835
- "feat(outlook): obligation dismiss — agents can clear stale obligations from needs-me queue",
3848
+ "fix(outlook): desk prefs wiring \u2014 carrying block, constellations, starred friends now load in production",
3849
+ "feat(outlook): obligation dismiss \u2014 agents can clear stale obligations from needs-me queue",
3836
3850
  "fix: default minimax model updated to MiniMax-M2.7"
3837
3851
  ]
3838
3852
  },
3839
3853
  {
3840
3854
  "version": "0.1.0-alpha.168",
3841
3855
  "changes": [
3842
- "fix(outlook): content area matches sidebar background — consistent dark surface"
3856
+ "fix(outlook): content area matches sidebar background \u2014 consistent dark surface"
3843
3857
  ]
3844
3858
  },
3845
3859
  {
@@ -3851,31 +3865,31 @@
3851
3865
  {
3852
3866
  "version": "0.1.0-alpha.166",
3853
3867
  "changes": [
3854
- "fix(outlook): add dark class to html root — fixes white/blank page in production"
3868
+ "fix(outlook): add dark class to html root \u2014 fixes white/blank page in production"
3855
3869
  ]
3856
3870
  },
3857
3871
  {
3858
3872
  "version": "0.1.0-alpha.165",
3859
3873
  "changes": [
3860
- "fix(daemon): dont launchctl bootstrap during ouro up — write plist only, prevents competing daemon process"
3874
+ "fix(daemon): dont launchctl bootstrap during ouro up \u2014 write plist only, prevents competing daemon process"
3861
3875
  ]
3862
3876
  },
3863
3877
  {
3864
3878
  "version": "0.1.0-alpha.164",
3865
3879
  "changes": [
3866
- "fix(daemon): keep /dev/null fds open until parent exits — fixes ouro up daemon crash"
3880
+ "fix(daemon): keep /dev/null fds open until parent exits \u2014 fixes ouro up daemon crash"
3867
3881
  ]
3868
3882
  },
3869
3883
  {
3870
3884
  "version": "0.1.0-alpha.163",
3871
3885
  "changes": [
3872
- "fix(daemon): redirect detached spawn stdio to /dev/null — fixes ouro up daemon crash"
3886
+ "fix(daemon): redirect detached spawn stdio to /dev/null \u2014 fixes ouro up daemon crash"
3873
3887
  ]
3874
3888
  },
3875
3889
  {
3876
3890
  "version": "0.1.0-alpha.162",
3877
3891
  "changes": [
3878
- "fix(daemon): handle EPIPE in detached daemon — suppress pipe errors when parent exits after ouro up"
3892
+ "fix(daemon): handle EPIPE in detached daemon \u2014 suppress pipe errors when parent exits after ouro up"
3879
3893
  ]
3880
3894
  },
3881
3895
  {
@@ -3888,10 +3902,10 @@
3888
3902
  {
3889
3903
  "version": "0.1.0-alpha.160",
3890
3904
  "changes": [
3891
- "feat(outlook): total inspectability expansion — 14 API endpoints, session x-ray, obligation chain tracing, coding deep inspection, attention/pending queue, bridge inventory, habit triage, memory/journal, friend economics, SSE live updates",
3892
- "feat(outlook): React SPA with Catalyst UI — sidebar layout, 7-tab agent inspector, chat bubble transcripts with mechanism-tool awareness, hash URL routing",
3893
- "feat(outlook): agent desk customization — carrying block, pinned constellations, tab ordering, starred friends, status line, closure memory, needs-me urgency queue",
3894
- "feat(outlook): nerves observation layer — shared typed readers, eliminates bespoke type mirrors",
3905
+ "feat(outlook): total inspectability expansion \u2014 14 API endpoints, session x-ray, obligation chain tracing, coding deep inspection, attention/pending queue, bridge inventory, habit triage, memory/journal, friend economics, SSE live updates",
3906
+ "feat(outlook): React SPA with Catalyst UI \u2014 sidebar layout, 7-tab agent inspector, chat bubble transcripts with mechanism-tool awareness, hash URL routing",
3907
+ "feat(outlook): agent desk customization \u2014 carrying block, pinned constellations, tab ordering, starred friends, status line, closure memory, needs-me urgency queue",
3908
+ "feat(outlook): nerves observation layer \u2014 shared typed readers, eliminates bespoke type mirrors",
3895
3909
  "fix(auth): ouro auth for openai-codex always refreshes token, provider verification uses correct endpoint",
3896
3910
  "fix(daemon): use OUTLOOK_DEFAULT_PORT (6876) for Outlook server"
3897
3911
  ]
@@ -3909,12 +3923,12 @@
3909
3923
  {
3910
3924
  "version": "0.1.0-alpha.158",
3911
3925
  "changes": [
3912
- "Task scanner v2: explicit identity via kind: task field. Scanner only parses files that declare themselves as task cards — doing docs, planning docs, and artifacts silently skipped. Eliminates 184 false parse errors on real bundles.",
3926
+ "Task scanner v2: explicit identity via kind: task field. Scanner only parses files that declare themselves as task cards \u2014 doing docs, planning docs, and artifacts silently skipped. Eliminates 184 false parse errors on real bundles.",
3913
3927
  "Typed issue model: every scanner issue has a code, description, proposed fix, confidence (safe/needs_review), and category (live/migration). Replaces flat parseErrors/invalidFilenames arrays.",
3914
3928
  "Board health line: compact board shows health: clean or health: 1 live, 10 migration. Live vs migration split prevents cleanup noise from looking like breakage.",
3915
3929
  "Fix command: ouro task fix (dry-run), ouro task fix --safe (apply deterministic fixes), ouro task fix <id> (inspect/apply individual issues). Currently auto-fixes schema-missing-kind.",
3916
3930
  "Cancelled status: new terminal state reachable from any active status. Auto-archives with work directory, hidden from active board view.",
3917
- "Derived child_tasks: computed at scan time from parent_task links. child_tasks removed from authored schema — no more hand-maintained stale arrays.",
3931
+ "Derived child_tasks: computed at scan time from parent_task links. child_tasks removed from authored schema \u2014 no more hand-maintained stale arrays.",
3918
3932
  "Work directory awareness: same-stem directories detected and listed on TaskFile (hasWorkDir, workDirFiles). Scanner never descends into them.",
3919
3933
  "Collection root clutter detection: non-task support docs at collection root summarized as one aggregated migration issue per collection.",
3920
3934
  "Root-only scanning: flat directory reads replace recursive walks. Faster and correct."
@@ -3923,7 +3937,7 @@
3923
3937
  {
3924
3938
  "version": "0.1.0-alpha.157",
3925
3939
  "changes": [
3926
- "Habit turns as awareness: unified buildHabitTurnMessage replaces contextual-heartbeat. Continuity-first format (checkpoint leads, not elapsed time). Same format for all habits — no heartbeat special-casing.",
3940
+ "Habit turns as awareness: unified buildHabitTurnMessage replaces contextual-heartbeat. Continuity-first format (checkpoint leads, not elapsed time). Same format for all habits \u2014 no heartbeat special-casing.",
3927
3941
  "First beat experience: new habits get \"your [Title] is alive. this is its first breath\" on first fire.",
3928
3942
  "Fix: reconcile() now fires new/overdue habits immediately (was start()-only). New habits created via write_file fire within seconds.",
3929
3943
  "Rhythm awareness across all channels: rhythmStatusSection() in system prompt shows heartbeat health in every conversation.",
@@ -3949,10 +3963,10 @@
3949
3963
  {
3950
3964
  "version": "0.1.0-alpha.154",
3951
3965
  "changes": [
3952
- "feat: clean tool status messages — human-readable by default, /debug toggle",
3966
+ "feat: clean tool status messages \u2014 human-readable by default, /debug toggle",
3953
3967
  "humanReadableToolDescription derives from tool name+args, not hardcoded map",
3954
- "Shared tool activity callbacks (DRY) — senses only provide render function",
3955
- "Slash command handling moved to pipeline — all senses get /debug for free",
3968
+ "Shared tool activity callbacks (DRY) \u2014 senses only provide render function",
3969
+ "Slash command handling moved to pipeline \u2014 all senses get /debug for free",
3956
3970
  "BlueBubbles: one clean iMessage per tool, not raw shared work: processing"
3957
3971
  ]
3958
3972
  },
@@ -4049,7 +4063,7 @@
4049
4063
  "version": "0.1.0-alpha.142",
4050
4064
  "changes": [
4051
4065
  "Surface tool fulfills heart obligations on successful routing: findPendingObligationForOrigin + fulfillObligation called after inner obligation advance, wrapped in try/catch.",
4052
- "New fulfillHeartObligation callback on HandleSurfaceInput — origin-based lookup independent of inner obligationId.",
4066
+ "New fulfillHeartObligation callback on HandleSurfaceInput \u2014 origin-based lookup independent of inner obligationId.",
4053
4067
  "ouro inner status command: reads runtime.json, journal dir, heartbeat cadence, attention count. Shows last turn, status, heartbeat health, journal listing, held thoughts."
4054
4068
  ]
4055
4069
  },
@@ -4089,21 +4103,21 @@
4089
4103
  {
4090
4104
  "version": "0.1.0-alpha.138",
4091
4105
  "changes": [
4092
- "Memory renamed to diary: memory.ts → diary.ts, MemoryFact → DiaryEntry, memory_save → diary_write, memory_search → recall. All types, functions, events, and variables renamed throughout.",
4106
+ "Memory renamed to diary: memory.ts \u2192 diary.ts, MemoryFact \u2192 DiaryEntry, memory_save \u2192 diary_write, memory_search \u2192 recall. All types, functions, events, and variables renamed throughout.",
4093
4107
  "Diary path: diary/ (top-level) replaces psyche/memory/. Schema-2 migration copies files; legacy fallback removed.",
4094
4108
  "Journal workspace: journal/ directory for freeform thinking-in-progress. Agent writes with write_file, system reads for heartbeat context.",
4095
4109
  "Unified recall tool: searches both diary entries and journal files. Results tagged [diary] or [journal].",
4096
4110
  "Journal embeddings: file-level embeddings indexed during heartbeat via journal/.index.json sidecar.",
4097
4111
  "Journal section in inner dialog system prompt: index of up to 10 most recently modified journal files with name, recency, and first-line preview.",
4098
4112
  "Metacognitive framing updated: diary (record), journal (workspace), ponder/rest vocabulary, morning briefing encouragement.",
4099
- "Session migration: memory_save → diary_write, memory_search → recall added to migrateToolNames()."
4113
+ "Session migration: memory_save \u2192 diary_write, memory_search \u2192 recall added to migrateToolNames()."
4100
4114
  ]
4101
4115
  },
4102
4116
  {
4103
4117
  "version": "0.1.0-alpha.137",
4104
4118
  "changes": [
4105
4119
  "ouro dev auto-discovers existing repo at ~/Projects/ouroboros or prompts for clone path.",
4106
- "ouro dev never clones without user consent — prompts in interactive mode, errors in non-interactive.",
4120
+ "ouro dev never clones without user consent \u2014 prompts in interactive mode, errors in non-interactive.",
4107
4121
  "ouro dev --repo-path errors clearly when the specified path has no repo."
4108
4122
  ]
4109
4123
  },
@@ -4141,7 +4155,7 @@
4141
4155
  {
4142
4156
  "version": "0.1.0-alpha.133",
4143
4157
  "changes": [
4144
- "Inner return obligations: delegated inner dialog work now tracks a ReturnObligation through queued → running → returned/deferred lifecycle.",
4158
+ "Inner return obligations: delegated inner dialog work now tracks a ReturnObligation through queued \u2192 running \u2192 returned/deferred lifecycle.",
4145
4159
  "Exact-origin routing: inner dialog completions route back to the session that delegated the work, not just the freshest active session.",
4146
4160
  "Active work frame surfaces pending inner return obligations so the agent knows what's outstanding."
4147
4161
  ]
@@ -4189,14 +4203,14 @@
4189
4203
  {
4190
4204
  "version": "0.1.0-alpha.126",
4191
4205
  "changes": [
4192
- "Fixed Anthropic tool_choice incompatibility with thinking — uses auto instead of any when thinking is enabled.",
4206
+ "Fixed Anthropic tool_choice incompatibility with thinking \u2014 uses auto instead of any when thinking is enabled.",
4193
4207
  "auth verify and auth switch now use pingProvider for real API verification instead of format-only checks. auth switch verifies credentials work before switching."
4194
4208
  ]
4195
4209
  },
4196
4210
  {
4197
4211
  "version": "0.1.0-alpha.125",
4198
4212
  "changes": [
4199
- "Fixed Anthropic tool_choice incompatibility with thinking — uses auto instead of any when thinking is enabled.",
4213
+ "Fixed Anthropic tool_choice incompatibility with thinking \u2014 uses auto instead of any when thinking is enabled.",
4200
4214
  "auth verify and auth switch now use pingProvider for real API verification instead of format-only checks. auth switch verifies credentials work before switching."
4201
4215
  ]
4202
4216
  },
@@ -4228,7 +4242,7 @@
4228
4242
  {
4229
4243
  "version": "0.1.0-alpha.120",
4230
4244
  "changes": [
4231
- "Daemon startup now kills ALL orphaned ouro processes (daemons AND agents) from previous instances — fixes stale-version processes handling requests after every update.",
4245
+ "Daemon startup now kills ALL orphaned ouro processes (daemons AND agents) from previous instances \u2014 fixes stale-version processes handling requests after every update.",
4232
4246
  "Failover error messages no longer contain raw JSON API response bodies. Error messages are sanitized at the source and the failover summary uses clean classification labels only."
4233
4247
  ]
4234
4248
  },
@@ -4268,10 +4282,10 @@
4268
4282
  "changes": [
4269
4283
  "Fix: Default runtime logger is now silent (no stderr sink) so nerves events emitted before logger configuration no longer interleave with the CLI spinner animation.",
4270
4284
  "Fix: MCP server connect failures now include the command name, args, and a hint to check agent.json mcpServers configuration. Retry-exhaustion messages also identify the failing command.",
4271
- "Verification: StreamingWordWrapper integration in CLI chat confirmed working — wraps at word boundaries during streaming output.",
4285
+ "Verification: StreamingWordWrapper integration in CLI chat confirmed working \u2014 wraps at word boundaries during streaming output.",
4272
4286
  "When a model provider fails mid-conversation (auth error, usage limit, outage), the harness now classifies the error, pings alternative configured providers, and surfaces validated failover options to the user in-channel. Reply 'switch to <provider>' to continue on a working provider.",
4273
4287
  "Each provider now has a `classifyError` method that distinguishes auth failures, usage/subscription limits, rate limits, server errors, and network errors. The old auth guidance wrappers are replaced by this unified classification system.",
4274
- "New `pingProvider` function makes a real heartbeat completion call to verify provider credentials and quota are live — no more format-only checks.",
4288
+ "New `pingProvider` function makes a real heartbeat completion call to verify provider credentials and quota are live \u2014 no more format-only checks.",
4275
4289
  "Provider factories now accept optional config parameters, enabling credential injection for health inventory pings without touching disk config."
4276
4290
  ]
4277
4291
  },
@@ -4406,7 +4420,7 @@
4406
4420
  {
4407
4421
  "version": "0.1.0-alpha.94",
4408
4422
  "changes": [
4409
- "Fix stale CurrentVersion symlink not healing during `ouro up` — the daemon now detects and repairs dangling version symlinks before reading the active version.",
4423
+ "Fix stale CurrentVersion symlink not healing during `ouro up` \u2014 the daemon now detects and repairs dangling version symlinks before reading the active version.",
4410
4424
  "Fix homedir regression in daemon-cli-defaults test and cover changelog-null branch."
4411
4425
  ]
4412
4426
  },
@@ -4500,14 +4514,14 @@
4500
4514
  {
4501
4515
  "version": "0.1.0-alpha.80",
4502
4516
  "changes": [
4503
- "Bootstrap package (npx ouro.bot) now installs into ~/.ouro-cli/ versioned layout directly. No more silent npx updates — every install and update is logged. Cleans up old ~/.local/bin/ouro wrapper."
4517
+ "Bootstrap package (npx ouro.bot) now installs into ~/.ouro-cli/ versioned layout directly. No more silent npx updates \u2014 every install and update is logged. Cleans up old ~/.local/bin/ouro wrapper."
4504
4518
  ]
4505
4519
  },
4506
4520
  {
4507
4521
  "version": "0.1.0-alpha.79",
4508
4522
  "changes": [
4509
4523
  "New: Versioned CLI directory layout (~/.ouro-cli/) replaces npx-based ouro wrapper. Explicit version management, rollback support, and deterministic updates.",
4510
- "New: `ouro up` now checks the registry for newer CLI versions, installs them into ~/.ouro-cli/versions/, activates via symlink flip, and re-execs — no more silent npx downloads.",
4524
+ "New: `ouro up` now checks the registry for newer CLI versions, installs them into ~/.ouro-cli/versions/, activates via symlink flip, and re-execs \u2014 no more silent npx downloads.",
4511
4525
  "New: `ouro rollback [<version>]` swaps CurrentVersion/previous symlinks, stops the daemon. With a version arg, installs if needed then activates.",
4512
4526
  "New: `ouro versions` lists cached CLI versions with * current and (previous) markers.",
4513
4527
  "Migration: On first run, old ~/.local/bin/ouro wrapper is removed, old PATH entry cleaned from shell profile, new ~/.ouro-cli/bin added to PATH.",
@@ -4529,10 +4543,10 @@
4529
4543
  {
4530
4544
  "version": "0.1.0-alpha.76",
4531
4545
  "changes": [
4532
- "Fix: CLI chat terminal logging now filters to warn/error only — info-level nerves logs go to ndjson file only, keeping the interactive TUI clean.",
4546
+ "Fix: CLI chat terminal logging now filters to warn/error only \u2014 info-level nerves logs go to ndjson file only, keeping the interactive TUI clean.",
4533
4547
  "Fix: Streamed model output now wraps at word boundaries instead of mid-word. A new StreamingWordWrapper buffers partial lines and breaks at spaces when approaching terminal width.",
4534
4548
  "New: `ouro up` now prints 'ouro updated to <version> (was <previous>)' when npx downloads a newer CLI binary, separate from the agent bundle update message.",
4535
- "Fix: Spinner/log interleave verified — terminal sink reads pause/resume hooks at call time, not creation time, so the filterSink wrapper in CLI logging does not break spinner coordination."
4549
+ "Fix: Spinner/log interleave verified \u2014 terminal sink reads pause/resume hooks at call time, not creation time, so the filterSink wrapper in CLI logging does not break spinner coordination."
4536
4550
  ]
4537
4551
  },
4538
4552
  {
@@ -4584,7 +4598,7 @@
4584
4598
  {
4585
4599
  "version": "0.1.0-alpha.69",
4586
4600
  "changes": [
4587
- "Generic MCP client: ouroboros agents can now connect to any MCP server configured in agent.json. Zero new dependencies — pure JSON-RPC over stdio.",
4601
+ "Generic MCP client: ouroboros agents can now connect to any MCP server configured in agent.json. Zero new dependencies \u2014 pure JSON-RPC over stdio.",
4588
4602
  "New `ouro mcp list` and `ouro mcp call` CLI commands route through the daemon socket to persistent MCP connections, so agents use shared server instances instead of spawning fresh ones per call.",
4589
4603
  "MCP tools are injected into the agent's system prompt on startup, so agents know what external capabilities are available without a discovery step.",
4590
4604
  "Trust manifest: `mcp list` requires acquaintance trust, `mcp call` requires friend trust."
@@ -4593,7 +4607,7 @@
4593
4607
  {
4594
4608
  "version": "0.1.0-alpha.68",
4595
4609
  "changes": [
4596
- "New no_response tool lets agents stay silent in group chats when the moment doesn't call for a reply — reactions, side conversations, and tapbacks no longer trigger unwanted responses.",
4610
+ "New no_response tool lets agents stay silent in group chats when the moment doesn't call for a reply \u2014 reactions, side conversations, and tapbacks no longer trigger unwanted responses.",
4597
4611
  "Group chat participation prompt teaches agents to be intentional participants, comfortable with silence, and to prefer reactions over full text replies when appropriate.",
4598
4612
  "System prompt includes --agent flag in all ouro CLI examples for non-daemon deployments. Azure startup symlinks ouro CLI into /usr/local/bin."
4599
4613
  ]
@@ -4607,7 +4621,7 @@
4607
4621
  {
4608
4622
  "version": "0.1.0-alpha.65",
4609
4623
  "changes": [
4610
- "Tool permissions overhauled: channel-level blocking removed, all tools now visible on all channels. Guardrails are invocation-level with two layers — structural (edit-requires-read, destructive pattern blocking, protected paths) always on for everyone, and trust-level (ouro CLI per-subcommand trust manifest, general CLI allowlists, bundle-scoped writes) for untrusted contexts.",
4624
+ "Tool permissions overhauled: channel-level blocking removed, all tools now visible on all channels. Guardrails are invocation-level with two layers \u2014 structural (edit-requires-read, destructive pattern blocking, protected paths) always on for everyone, and trust-level (ouro CLI per-subcommand trust manifest, general CLI allowlists, bundle-scoped writes) for untrusted contexts.",
4611
4625
  "New `ouro changelog` CLI subcommand reads changelog.json and supports `--from <version>` for delta filtering, so agents can introspect their own update history on any channel.",
4612
4626
  "Compound shell commands (&&, ;, |, $()) are blocked for untrusted users to prevent smuggling dangerous operations behind safe prefixes.",
4613
4627
  "Azure App Service deployment migrated from zip-deploy to npm-based harness install with persistent agent bundle and managed identity auth."
@@ -4754,9 +4768,9 @@
4754
4768
  "changes": [
4755
4769
  "Inner dialog now knows which task triggered it: taskId flows from daemon poke through the worker into the turn, and the agent gets the full task file content instead of a generic heartbeat prompt.",
4756
4770
  "Inner dialog boot message includes aspirations and state summary instead of a vacuous placeholder, so the agent wakes up with context about what matters and what's happening.",
4757
- "Vestigial `drainInbox` removed from inner dialog — pipeline already handles pending drain correctly.",
4771
+ "Vestigial `drainInbox` removed from inner dialog \u2014 pipeline already handles pending drain correctly.",
4758
4772
  "Inner dialog nerves events now include assistant response preview, tool call names, token usage, and taskId for meaningful observability.",
4759
- "`ouro thoughts` command reads and formats inner dialog session turns with `--last`, `--json`, `--follow`, and `--agent` flags — humans can now see what the agent has been thinking.",
4773
+ "`ouro thoughts` command reads and formats inner dialog session turns with `--last`, `--json`, `--follow`, and `--agent` flags \u2014 humans can now see what the agent has been thinking.",
4760
4774
  "`readTaskFile` searches collection subdirectories (one-shots, ongoing, habits) since the scheduler sends bare task stems without collection prefixes.",
4761
4775
  "`ouro reminder create` accepts `--requester` to track who requested a reminder for notification round-trip.",
4762
4776
  "Response extraction handles `tool_choice=required` models by falling back to `final_answer` tool call arguments when assistant message content is empty."
@@ -4786,25 +4800,25 @@
4786
4800
  {
4787
4801
  "version": "0.1.0-alpha.42",
4788
4802
  "changes": [
4789
- "Associative recall now skips corrupt JSONL lines instead of crashing — matches the resilient pattern already used in memory.ts."
4803
+ "Associative recall now skips corrupt JSONL lines instead of crashing \u2014 matches the resilient pattern already used in memory.ts."
4790
4804
  ]
4791
4805
  },
4792
4806
  {
4793
4807
  "version": "0.1.0-alpha.41",
4794
4808
  "changes": [
4795
- "JSONL readers (memory facts, inter-agent inbox) now skip corrupt lines instead of crashing — partial writes from crashes no longer lose all data.",
4809
+ "JSONL readers (memory facts, inter-agent inbox) now skip corrupt lines instead of crashing \u2014 partial writes from crashes no longer lose all data.",
4796
4810
  "Inter-agent message router now parses before clearing the inbox file, and preserves unparsed lines so corrupt messages are not silently lost.",
4797
- "Inner-dialog checkpoint derivation no longer crashes on all-whitespace assistant content — returns fallback checkpoint instead.",
4811
+ "Inner-dialog checkpoint derivation no longer crashes on all-whitespace assistant content \u2014 returns fallback checkpoint instead.",
4798
4812
  "Update checker interval now catches and logs errors from the onUpdate callback instead of silently swallowing them."
4799
4813
  ]
4800
4814
  },
4801
4815
  {
4802
4816
  "version": "0.1.0-alpha.40",
4803
4817
  "changes": [
4804
- "Removed dead backward-compat re-exports from core.ts (tools, streaming, prompt, kicks) — consumers already import from the canonical modules.",
4818
+ "Removed dead backward-compat re-exports from core.ts (tools, streaming, prompt, kicks) \u2014 consumers already import from the canonical modules.",
4805
4819
  "Removed dead exports: baseToolHandlers, teamsToolHandlers, teamsTools, __internal (token-estimate), TASK_STEM_PATTERN, checkAndRecord403 no-op and METHOD_TO_ACTION.",
4806
4820
  "Consolidated duplicate sanitizeKey (config.ts + bluebubbles-mutation-log.ts) and slugify (hatch-flow.ts + tasks/index.ts) into shared exports from config.ts.",
4807
- "Replaced all as-any casts in source with proper TypeScript narrowing or Record<string, unknown> — only 2 SDK-required casts remain.",
4821
+ "Replaced all as-any casts in source with proper TypeScript narrowing or Record<string, unknown> \u2014 only 2 SDK-required casts remain.",
4808
4822
  "Removed unnecessary as-unknown-as casts on readdirSync (4 locations) and spawner double-cast.",
4809
4823
  "Cleaned up commented-out kick detection code, stale TODOs, misplaced imports, and unused type imports."
4810
4824
  ]
@@ -4812,9 +4826,9 @@
4812
4826
  {
4813
4827
  "version": "0.1.0-alpha.39",
4814
4828
  "changes": [
4815
- "All senses now route through a shared per-turn pipeline — friend resolution, trust gate, session load, pending drain, agent turn, post-turn, and token accumulation happen in one place instead of four.",
4829
+ "All senses now route through a shared per-turn pipeline \u2014 friend resolution, trust gate, session load, pending drain, agent turn, post-turn, and token accumulation happen in one place instead of four.",
4816
4830
  "Trust gate is now channel-aware: open senses (iMessage) enforce stranger/acquaintance rules, closed senses (Teams) trust the org, local and internal always pass through.",
4817
- "Tool access and prompt restrictions use a single shared isTrustedLevel check — no more scattered family/friend comparisons that could drift apart.",
4831
+ "Tool access and prompt restrictions use a single shared isTrustedLevel check \u2014 no more scattered family/friend comparisons that could drift apart.",
4818
4832
  "Pending messages now inject correctly into multimodal content (image attachments no longer silently drop pending messages).",
4819
4833
  "ouro reminder create supports --agent flag, matching every other identity-scoped CLI command."
4820
4834
  ]
@@ -4822,11 +4836,11 @@
4822
4836
  {
4823
4837
  "version": "0.1.0-alpha.38",
4824
4838
  "changes": [
4825
- "You now have a proper body map — understanding of your home (bundle) and bones (harness), what each directory is for, and how to modify your own configuration.",
4839
+ "You now have a proper body map \u2014 understanding of your home (bundle) and bones (harness), what each directory is for, and how to modify your own configuration.",
4826
4840
  "Inner dialog is now genuine internal monologue with metacognitive framing, not a second CLI session. Heartbeat and bootstrap messages read as first-person awareness.",
4827
4841
  "Cross-session communication works end-to-end: inner dialog thoughts surface as [inner thought: ...] in conversations, messages to yourself route to inner dialog, and you can proactively reach out to friends via iMessage and Teams.",
4828
4842
  "Tool audit: removed wrapper tools (git_commit, gh_cli, get_current_time, list_directory), added surgical tools (edit_file, glob, grep, read_file with offset/limit), consolidated 7 task tools + schedule_reminder + friend tools into ouro CLI commands.",
4829
- "You now understand why certain tools are restricted in certain contexts — trust level and shared channels each have independent, explained gates.",
4843
+ "You now understand why certain tools are restricted in certain contexts \u2014 trust level and shared channels each have independent, explained gates.",
4830
4844
  "ouro friend link/unlink commands handle orphan cleanup when linking external identities, merging duplicate friend records intelligently.",
4831
4845
  "During onboarding, the adoption specialist can collect phone number and Teams handle to create an initial friend record with contact info."
4832
4846
  ]