mercury-agent 0.16.0 → 0.16.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (80) hide show
  1. package/README.md +6 -0
  2. package/container/Dockerfile +6 -1
  3. package/container/Dockerfile.base +6 -1
  4. package/docs/README.md +50 -0
  5. package/docs/authoring-profiles.md +8 -0
  6. package/docs/behavior-layers.md +193 -0
  7. package/docs/extensions.md +2 -0
  8. package/docs/goals/football-reporter-profile/decisions.md +479 -0
  9. package/docs/goals/football-reporter-profile/goal.md +91 -0
  10. package/docs/goals/football-reporter-profile/roadmap.md +196 -0
  11. package/docs/goals/whatsapp-bot-hardening/README.md +43 -0
  12. package/docs/goals/whatsapp-bot-hardening/archive/ambient-group-context.md +211 -0
  13. package/docs/goals/whatsapp-bot-hardening/archive/command-routing-consistency.md +210 -0
  14. package/docs/goals/whatsapp-bot-hardening/archive/destructive-command-guards.md +218 -0
  15. package/docs/goals/whatsapp-bot-hardening/archive/docs-extension.md +182 -0
  16. package/docs/goals/whatsapp-bot-hardening/archive/handoff-2026-08-09.md +588 -0
  17. package/docs/goals/whatsapp-bot-hardening/archive/media-size-and-silent-drop.md +178 -0
  18. package/docs/goals/whatsapp-bot-hardening/archive/member-memory.md +199 -0
  19. package/docs/goals/whatsapp-bot-hardening/archive/message-author-attribution.md +211 -0
  20. package/docs/goals/whatsapp-bot-hardening/archive/run-from-source-switchover.md +212 -0
  21. package/docs/goals/whatsapp-bot-hardening/archive/setup-plan-2026-08-06.md +1117 -0
  22. package/docs/goals/whatsapp-bot-hardening/decisions.md +200 -0
  23. package/docs/goals/whatsapp-bot-hardening/goal.md +75 -0
  24. package/docs/goals/whatsapp-bot-hardening/open-threads.md +396 -0
  25. package/docs/goals/whatsapp-bot-hardening/roadmap.md +245 -0
  26. package/docs/memory.md +30 -0
  27. package/docs/profile-guide.md +483 -0
  28. package/docs/refactor/.gitkeep +0 -0
  29. package/docs/refactor/archive/.gitkeep +0 -0
  30. package/docs/refactor/audits/.gitkeep +0 -0
  31. package/docs/refactor/backlog/.gitkeep +0 -0
  32. package/docs/skills-guide.md +154 -0
  33. package/examples/extensions/README.md +1 -0
  34. package/examples/extensions/feed-watch/config.ts +354 -0
  35. package/examples/extensions/feed-watch/digest.ts +344 -0
  36. package/examples/extensions/feed-watch/feeds.ts +341 -0
  37. package/examples/extensions/feed-watch/index.ts +302 -0
  38. package/examples/extensions/feed-watch/items.ts +171 -0
  39. package/examples/extensions/feed-watch/match.ts +107 -0
  40. package/examples/extensions/feed-watch/prompts/verify.md +35 -0
  41. package/examples/extensions/feed-watch/skill/SKILL.md +128 -0
  42. package/examples/extensions/feed-watch/watch.ts +548 -0
  43. package/examples/extensions/longview/hook.ts +179 -16
  44. package/examples/extensions/longview/index.ts +13 -2
  45. package/examples/extensions/longview/prompts/summarize.md +3 -2
  46. package/examples/extensions/longview/render/telegraph-nodes.ts +21 -1
  47. package/examples/extensions/longview/summarize.ts +57 -14
  48. package/examples/extensions/napkin/index.ts +334 -86
  49. package/examples/extensions/napkin/pi-spawn.ts +196 -0
  50. package/examples/extensions/pinchtab/index.ts +37 -4
  51. package/examples/extensions/pinchtab/lib/session-injector.ts +31 -9
  52. package/examples/extensions/pinchtab/skill/SKILL.md +29 -0
  53. package/examples/profiles/_template/AGENTS.md +138 -0
  54. package/examples/profiles/_template/README.md +65 -0
  55. package/examples/profiles/_template/config.yaml +57 -0
  56. package/examples/profiles/_template/tasks/daily.md +32 -0
  57. package/examples/profiles/football-reporter/AGENTS.md +219 -0
  58. package/examples/profiles/football-reporter/README.md +138 -0
  59. package/examples/profiles/football-reporter/config.yaml +174 -0
  60. package/examples/profiles/football-reporter/seed/MEMORY.md +42 -0
  61. package/examples/profiles/football-reporter/seed/episodes/barcelona-2026-27.md +25 -0
  62. package/examples/profiles/football-reporter/seed/episodes/beitar-jerusalem-2026-27.md +25 -0
  63. package/examples/profiles/football-reporter/seed/episodes/maccabi-haifa-2026-27.md +24 -0
  64. package/examples/profiles/football-reporter/seed/episodes/man-united-2026-27.md +26 -0
  65. package/examples/profiles/football-reporter/seed/episodes/real-madrid-2026-27.md +25 -0
  66. package/examples/profiles/football-reporter/seed/napkin-distill.md +55 -0
  67. package/examples/profiles/football-reporter/tasks/daily-article.md +142 -0
  68. package/package.json +8 -5
  69. package/src/adapters/whatsapp.ts +75 -13
  70. package/src/bridges/whatsapp.ts +6 -4
  71. package/src/cli/mrctl.ts +9 -1
  72. package/src/core/handler.ts +51 -6
  73. package/src/core/routes/dashboard.ts +106 -5
  74. package/src/core/routes/tasks.ts +28 -0
  75. package/src/core/runtime.ts +63 -2
  76. package/src/core/task-scheduler.ts +103 -12
  77. package/src/extensions/catalog.ts +9 -0
  78. package/src/profile/space-profile.ts +780 -0
  79. package/src/storage/db.ts +215 -0
  80. package/src/types.ts +56 -0
@@ -0,0 +1,196 @@
1
+ # Roadmap: Football Reporter Profile
2
+
3
+ **Goal**: [football-reporter-profile](goal.md)
4
+ **Last updated**: 2026-08-22 (M2.3 shipped; audit extension M2.3–M2.5, D-016–D-018; napkin revisit — M2.3 re-cut, D-019)
5
+
6
+ > Sequence chosen by the owner on 2026-08-20: (1) profile + one editorial
7
+ > layer → (2) deterministic feed → (3) proof of quality, with a parallel bug
8
+ > track for anything unrelated, then (4) regression + re-audit + monitoring.
9
+
10
+ ## Milestones
11
+
12
+ ### Milestone 1: One owner for the editorial rules, and an article worth tapping
13
+ > M1.1 is zero code change to the bot and observable on the live group within
14
+ > 48 h; after it every football rule lives in the repo, deploys by script, and
15
+ > a `check` proves no drift. M1.2 redesigns the daily article (Telegraph page
16
+ > + fuller in-chat summary, D-005/D-010) — the one code-bearing story in this
17
+ > milestone (`longview`).
18
+
19
+ | ID | Story | Slug | Depends on | Status |
20
+ |----|-------|------|------------|--------|
21
+ | M1.1 | Space profile folder + `AGENTS.md` editorial standard + apply/check/dump script; slim task prompts; delete `club-coverage` | football-reporter-space-profile | — | done |
22
+ | M1.2 | Daily article redesign: top-tier Telegraph article (headline, standfirst, sections, pull-quote, certainty box, sources), author-supplied fuller chat summary in `longview` (F3) | daily-article-redesign | M1.1 | done |
23
+
24
+ **Checkpoint:** after two daily articles and ~10 scans on the new `AGENTS.md`:
25
+ no `"Done."`/filler reached the group; every posted item shows `(source)` and
26
+ a date; 0 emoji in scheduled output; `club-coverage` is gone from the prefs
27
+ panel; `scripts/space-profile.ts check` is clean; the redesigned page passes
28
+ the visual checklist on a phone, the chat summary carries the key facts, and
29
+ the group's reaction to the first two articles is recorded in the goal's notes.
30
+
31
+ > **Checkpoint passed 2026-08-21** on one article rather than two, by owner's
32
+ > decision — M2.1 depends on M1.1 only, so nothing about feed-watch waits on the
33
+ > second article. Evidence from the first 24 h after the deploy (ledger runs #1
34
+ > and #2, full detail in `docs/pending-verification.md`):
35
+ >
36
+ > - **Passed:** `NO_UPDATE` fired unattended on the 00:00 scan (9-char reply,
37
+ > `outcome=no_update`, nothing posted); the 09:00 article carries 1 `#` title,
38
+ > 8 `##` sections, exactly 1 `>>` pull-quote, single-star bold only, 0 emoji
39
+ > and 0 exclamation marks; the summary block is 452 chars / 4 lines, inside
40
+ > both the 900 cap and the 3–6 band; one model call per run, confirmed against
41
+ > `token_usage`; `space-profile check` clean on `apply` and on the next sync;
42
+ > `club-coverage` gone.
43
+ > - **Failed, and folded into M3.1's lint per [D-014](decisions.md):** no item
44
+ > carries a publication date — on the scan *and* on the article summary, though
45
+ > the article body of the same run dated 5 items.
46
+ > - **Still open:** the Telegraph page cannot be confirmed from the host (since
47
+ > M1.2 the `storedReply` seam stores the raw article, so the delivered link is
48
+ > in no table); the phone RTL check; the group's reaction; the second article.
49
+ > - **Baseline for M2:** scan $0.456/run, article $0.769/run; whole-project daily
50
+ > spend $7.53 (08-19) and $7.19 (08-20), against M2's ~$5.30/day direction.
51
+
52
+ ### Milestone 2: A deterministic feed in front of the model
53
+ > Scans stop being blind; empty scans cost nothing; a fallback-leg run never
54
+ > posts.
55
+
56
+ | ID | Story | Slug | Depends on | Status |
57
+ |----|-------|------|------------|--------|
58
+ | M2.1 | Feed Watch — host-side poller, digest into the article, verify one-shots (existing spec, absorbed) | feed-watch | M1.1 | done (merged `e07b239` 2026-08-21; **deploy not yet run** — runbook in `pending-verification.md`) |
59
+ | M2.2 | Scheduled task model-leg policy: primary-or-skip for research runs (F2) | scheduled-task-model-leg-policy | — | deferred until after M3.1 (D-015) |
60
+ | M2.3 | Reporter notebook, re-cut 2026-08-21 (D-019): topic notes are napkin-shaped **episodes** the profile seeds once and napkin maintains nightly, written in-turn before posting; seeded `MEMORY.md` for standing facts; priorities tie-break in `AGENTS.md`; `workspace_seed` in `space-profile`; per-space distillation addendum in napkin; member perms and stale-note hygiene (audit R1/R6/R7, D-016, D-019) | reporter-notebook | M1.1 | done (merged 2026-08-22; **deploy not yet run** — runbook in `pending-verification.md`) |
61
+ | M2.4 | One voice per space: `persona.exclusive` drops the global character and global `AGENTS.md` for a space that owns its standard; platform-prompt capability paragraphs gated on the caller's permissions (audit R2, D-017) | one-voice-per-space | — | in-progress (worktree `worktree-feature-one-voice-per-space`; host; image rebuild) |
62
+ | M2.5 | Chat window holds conversation, not procedure (scheduled prompts out of the turn count, no halving on swipe-reply); silence is a legal chat reply (audit R3/R4, D-018) | chat-window-and-silence | — | backlog (host) |
63
+
64
+ > **Revisit 2026-08-21 — napkin works now** (`docs/notes/napkin-revisit-2026-08-21.md`). **Approved by the owner the same day ([D-019](decisions.md)):** M2.3 re-cut so the topic notes are napkin-maintained episodes under `knowledge/episodes/` (the only injected dir) that the profile seeds once, plus a per-space distillation addendum. Nothing removed; M2.4/M2.5/M3.x unchanged. Idea parked: `docs/ideas/napkin-member-notes.md`.
65
+
66
+ > **M2.2 is largely pre-empted by the 2026-08-20 Opus pin.** `runtime.ts:1871`
67
+ > turns `space_config['model.active']` into a **one-element** `MODEL_CHAIN`
68
+ > for the container — its own comment reads "eliminates automatic
69
+ > fallback" — and all 6 spaces carry that row. So no run of any task in any
70
+ > space can reach a fallback leg today, and a leg failure already surfaces as
71
+ > a task error that retries and then reports, posting nothing. What M2.2 would
72
+ > still add is (a) the guarantee being **per task** rather than per space, so
73
+ > it survives someone restoring chat fallback by deleting the pin, and (b) a
74
+ > skip that is **distinguishable** in the ledger from any other error
75
+ > (`last_status` is only `ok`/`error`). Owner's call at this checkpoint
76
+ > whether that is worth doing now or after M3.2 makes it measurable.
77
+
78
+ > **M2.2 deferred, not dropped (2026-08-21, [D-015](decisions.md)).** All 12 live
79
+ > runs since the profile applied ran on `claude-opus-5`, so the checkpoint's "zero
80
+ > posts from a non-primary leg" is already true by construction. Architecture stays
81
+ > unfilled until M3.1 can measure whether a per-task guarantee earns its keep.
82
+
83
+ **Checkpoint:** task #28 and its one-shots are gone; the first `feed-watch`
84
+ one-shot produced either a sourced 1–3 line post or nothing; the 09:00 article
85
+ trace carries `<feed_watch_digest … items="N">`; a week of spend is reported
86
+ (target direction: well under ~$5.30/day — reported, not gated); zero group
87
+ posts from a non-primary leg. **Added 2026-08-21 (audit):** `agent.trace_runs`
88
+ is on; a factual chat question's trace shows the topic note being read and the
89
+ next article appends to one; the football space runs with `persona.exclusive`
90
+ and a member run's system prompt carries no preference-management paragraph; a
91
+ swipe-reply follow-up sees the bot's own post from the previous evening; a
92
+ "don't answer" message in `main` gets no reply.
93
+
94
+ > **Extended 2026-08-21 ([D-016](decisions.md)–[D-018](decisions.md)).** The
95
+ > 2026-08-21 audit (`docs/notes/football-audit-2026-08-21.md`) traced both
96
+ > owner-reported failures — a hallucinated name and a joke refused as a
97
+ > permission request — to mechanisms, not missing rules: the window, not a
98
+ > memory; four personas, not one; no silent chat path. M2.3–M2.5 are those
99
+ > mechanisms. Per the owner, nothing in them is tailored to the bot's state
100
+ > today — M3.1/M4.1 are assumed to land; M2.3 is what gives M3.1's provenance
101
+ > lint a source to check against. **Order:** M2.3 first (profile-side, no
102
+ > restart); M2.4 and M2.5 are independent host stories and may run in
103
+ > parallel with M3.1.
104
+ >
105
+ > **Build order note (2026-08-20, D-012):** M3.2 is being implemented before
106
+ > M2.1. It has no dependencies, and an instrument installed after the change
107
+ > it measures cannot show the before — feed-watch should ship against a
108
+ > baseline the ledger already recorded. Milestone membership is unchanged.
109
+
110
+ ### Milestone 3: Proof of quality
111
+ > The rules are testable without reading the chat.
112
+
113
+ | ID | Story | Slug | Depends on | Status |
114
+ |----|-------|------|------------|--------|
115
+ | M3.1 | Fixture harness: dated fixture items through the verify/article prompts in the private space + a deterministic output lint | football-reporter-fixture-harness | M1.1 | backlog |
116
+ | M3.2 | Scheduled-run ledger: model, cost, outcome per run; dashboard badge + drawer, `mrctl tasks runs` | scheduled-run-ledger | — | done (2026-08-20, ahead of M2 — D-012) |
117
+
118
+ **Checkpoint:** harness green on the fixture set (fresh item passes, undated
119
+ and stale items dropped, emoji/HTML/line-count lint fires on seeded bad
120
+ replies); every scheduled run of the last 48 h has a ledger row.
121
+
122
+ ### Milestone 4: Regression and monitoring
123
+ > Every previously reported defect has a check that would catch its return;
124
+ > a re-audit confirms; new problems surface from the ledger.
125
+
126
+ | ID | Story | Slug | Depends on | Status |
127
+ |----|-------|------|------------|--------|
128
+ | M4.1 | Regression suite for F1–F8 + `scripts/football-audit.ts` data dump + two-week re-audit note | football-regression-suite | M2.1, M3.1, M3.2 | blocked |
129
+
130
+ **Checkpoint:** regression suite green in `bun run check`; re-audit note
131
+ (`docs/notes/football-group-reaudit-<date>.md`) finds no recurrence of F1–F7
132
+ and reports F8 spend; the ledger is the first place a new problem shows.
133
+
134
+ ## Bug track (parallel — not stories, per D-007)
135
+
136
+ Filed at planning in `docs/bugs/`; fix with `/f-bug-fix` whenever convenient,
137
+ independent of the milestones. Any *new* defect found while executing a story
138
+ is filed the same way, never fixed inside the story.
139
+
140
+ | Audit ref | Bug doc | Severity | Regression check lands in |
141
+ |-----------|---------|----------|---------------------------|
142
+ | F5 | `reply-to-member-triggers-bot.md` | moderate | M4.1 |
143
+ | F7a | `chat-output-html-tags-not-stripped.md` | minor | M4.1 (lint also in M3.1) |
144
+ | F7d | `responded-in-suffix-not-stored.md` | minor | M4.1 |
145
+ | F7e | `bot-self-name-is-phone-in-reply-to.md` | minor | M4.1 |
146
+ | — | `model-switch-disables-fallback.md` | moderate | M4.1 |
147
+ | audit-2 C9 | `feed-watch-seen-ttl-not-refreshed.md` | moderate | M4.1 (fix **before** the feed-watch deploy) |
148
+ | audit-2 C10 | `feed-watch-runas-non-user-principal.md` | moderate | M4.1 (fix **before** the feed-watch deploy) |
149
+
150
+ The two `feed-watch-*` rows come from the 2026-08-21 audit's second pass
151
+ (`docs/notes/football-audit-2026-08-21.md` §3.1): a `seen` TTL that counts
152
+ from first sighting (a wave of stale items re-enters every 72 h) and a
153
+ `runAs` that can be the literal string `dashboard` (phantom member row,
154
+ member permissions). Both are small `watch.ts` changes with designed fixes
155
+ in the bug docs; both gate the deploy.
156
+
157
+ The last row is not from the 48h audit. It was found on 2026-08-20 while
158
+ expanding the live `model.chain` from 2 legs to 5 (opus-5, sonnet-5,
159
+ gemini-3.1-flash-lite, fable-5, haiku-4-5): `/model switch` truncates the chain
160
+ to the chosen leg and no verb restores it.
161
+
162
+ That same day the owner chose **Opus-only, no fallback, but all 5 switchable**,
163
+ and this truncation is the only mechanism in the codebase that delivers it — so
164
+ all 6 spaces are now deliberately pinned to `anthropic:claude-opus-5` and the
165
+ chain serves purely as the `/model switch` menu. The bug is therefore half
166
+ feature: what stays broken is that nothing can clear the pin, and that
167
+ `/model list` still renders unreachable legs as if they were fallbacks. Any fix
168
+ must be additive — see the bug doc before touching `resolveActiveIndex`.
169
+
170
+ Bearing on the milestones: every run is now single-leg, so `max_retries_per_leg`
171
+ is the only resilience left, and the M3.2 run ledger should record the model per
172
+ run rather than assume the chain.
173
+
174
+ Not bugs (handled by stories): F1 → resolved 2026-08-19, locked in by M1.1;
175
+ F2 → M2.2; F3 → M1.2; F4 → M1.1 + M2.1 + M3.1; F6 → M1.1 honesty rule;
176
+ F7c → M1.1 + M3.1 lint; F7f (Hebrew Telegraph slugs) → investigated in
177
+ M1.2, likely accepted; F8 → M2.1 + M3.2.
178
+
179
+ ## Dependency Graph
180
+
181
+ ```
182
+ M1.1 ─┬─► M1.2
183
+ ├─► M2.1 ─┐
184
+ ├─► M2.3 ─┤ (M3.1's provenance check reads M2.3's notebook — soft)
185
+ └─► M3.1 ─┼─► M4.1
186
+ M2.2 (independent) │
187
+ M2.4 (independent) ──┤
188
+ M2.5 (independent) ──┤
189
+ M3.2 (independent) ──┘
190
+ ```
191
+
192
+ M1.1 gates everything that touches the space's prompts. M2.2, M2.4, M2.5 and
193
+ M3.2 are core changes with no football dependency and can run in parallel —
194
+ M3.2 is worth starting early because its ledger is what makes the M2 and M4
195
+ checkpoints measurable. M2.3 is profile-side and gates only the notebook half
196
+ of M3.1's provenance check; the rest of M3.1 is unblocked.
@@ -0,0 +1,43 @@
1
+ # whatsapp-bot-hardening — start here
2
+
3
+ Ops and planning docs for running Mercury as a live WhatsApp bot. Everything in
4
+ this folder is documentation; nothing here is executed.
5
+
6
+ ## Where things are
7
+
8
+ | Path | What |
9
+ |------|------|
10
+ | [`open-threads.md`](open-threads.md) | **Start here.** Everything still outstanding. |
11
+ | [`roadmap.md`](roadmap.md) | What shipped, milestone by milestone |
12
+ | [`decisions.md`](decisions.md) | D-001…D-014, with reasoning and revisit conditions |
13
+ | [`goal.md`](goal.md) | Goal statement, constraints, success criteria |
14
+ | [`archive/`](archive/) | Completed work — the 8 story specs, the setup plan, the bug-investigation handoff |
15
+
16
+ ## Archived — read for context, don't plan from them
17
+
18
+ - `archive/handoff-2026-08-09.md` — the bug investigation. Every fix in it
19
+ shipped; the root-cause analysis is still the best map of *why*.
20
+ - `archive/setup-plan-2026-08-06.md` — how the bot was built. Phases 0–5
21
+ complete; the Gemini and npm-install sections are superseded.
22
+ - `archive/<story>.md` — one spec per hardening story, with retrospectives.
23
+
24
+ Both archived files carry a header listing exactly what in them has gone stale.
25
+
26
+ ## Where the code lives
27
+
28
+ Since 2026-08-10 there is **one** source of truth:
29
+
30
+ | Tree | Role |
31
+ |------|------|
32
+ | `D:\Projects\mercury` (Windows) | **Source of truth.** All edits and commits happen here. Pushes to `origin` once contributor access lands. |
33
+ | `~/mercury-src` (WSL) | Run-only mirror. The systemd unit `mercury` runs `bun` directly against this checkout — no build step. Never commit here. |
34
+ | `~/whatsapp-bot` (WSL) | Runtime data only — config, `.env`, SQLite DB. The unit's working directory. |
35
+
36
+ After changing code on Windows, push it to the running bot with:
37
+
38
+ ```bash
39
+ wsl.exe -d Ubuntu -- bash -lc '~/sync-mercury.sh'
40
+ ```
41
+
42
+ That fetches from the Windows clone, hard-resets the mirror to match, reinstalls
43
+ dependencies only if `bun.lock` changed, and restarts the service.
@@ -0,0 +1,211 @@
1
+ # Ambient Group Context
2
+
3
+ **Status**: Done
4
+ **Slug**: ambient-group-context
5
+ **Goal**: whatsapp-bot-hardening
6
+ **Milestone**: 2
7
+ **Created**: 2026-08-09
8
+ **Last updated**: 2026-08-09
9
+
10
+ ---
11
+
12
+ ## Goal
13
+
14
+ > Part of goal [whatsapp-bot-hardening](../goals/whatsapp-bot-hardening/goal.md).
15
+ > HANDOFF Bug 1 — highest impact.
16
+
17
+ The bot cannot see any group message it isn't tagged in. `ambient.enabled`
18
+ defaults to true and the ambient-write branch works as designed — but it lives
19
+ in `runtime.ts:591-612` behind `processIngress`, which is only reached **after**
20
+ the early return at `handler.ts:237-240` has already dropped untriggered
21
+ messages. It is dead code: across the whole DB there are exactly 2 ambient rows,
22
+ both stray `/`-prefixed messages that leaked through the `isCommand` bypass.
23
+ Upstream has not fixed this (`handler.ts` byte-identical on origin/main).
24
+ Depends on `message-author-attribution` (M2.1) so ambient rows carry authors
25
+ from the start, and ships **only with** its own context cap + retention TTL
26
+ (D-003).
27
+
28
+ ## User Stories
29
+
30
+ - As a group member, I want to tag the bot and ask "what did we just decide?"
31
+ and have it answer from the conversation it silently observed.
32
+ - As the operator, I want ambient storage bounded (context budget + TTL) so
33
+ every group message becoming a row doesn't starve real turns or grow the DB
34
+ forever.
35
+
36
+ ## MVP Scope
37
+
38
+ **In scope:** public `recordAmbient(ingress)` on `MercuryCoreRuntime`
39
+ (extracted from `runtime.ts:591-612`, guards intact); call it from the
40
+ handler's early-return path; a dedicated bounded ambient context query; ambient
41
+ retention TTL in `storage-cleanup.ts`.
42
+
43
+ **Out of scope:** avoiding media downloads for untriggered messages —
44
+ `bridge.normalize` already runs before the gate today (pre-existing cost, not
45
+ added by this fix); tuning ambient *relevance* (which rows get injected) beyond
46
+ a bounded recency window.
47
+
48
+ ---
49
+
50
+ ## Context for Claude
51
+
52
+ - `src/core/handler.ts:237-240` — the early return that kills ambient; `ingress`
53
+ is already built at `:221`, so the data is in hand.
54
+ - `src/core/runtime.ts:591-612` — the ambient-write branch to extract
55
+ (group-only + `ambient.enabled` + non-DM guards must survive the move).
56
+ - `src/storage/db.ts:1047` (`getRecentTurns`) and `:1068` (`turnCount * 5`
57
+ window) — the budget ambient must NOT share.
58
+ - `src/core/storage-cleanup.ts` — where the TTL goes.
59
+ - `HANDOFF.md` Bug 1 — full analysis and evidence.
60
+
61
+ ---
62
+
63
+ ## Architecture & Data
64
+
65
+ ### Data Models
66
+
67
+ No new columns (`role='ambient'` rows already exist as a concept). New
68
+ config surface:
69
+
70
+ ```yaml
71
+ ambient:
72
+ context_rows: 30 # max ambient rows injected per turn (own query, own cap)
73
+ retention_days: 14 # TTL enforced by storage-cleanup
74
+ ```
75
+
76
+ Defaults decided in code review; both must exist before the handler fix ships.
77
+
78
+ ### API Contracts
79
+
80
+ None.
81
+
82
+ ### File & Folder Structure
83
+
84
+ | Path | New / Modified | Purpose |
85
+ |------|---------------|---------|
86
+ | `src/core/runtime.ts` | Modified | Extract `recordAmbient(ingress)` as a public method |
87
+ | `src/core/handler.ts` | Modified | Call `core.recordAmbient(ingress)` before the early return |
88
+ | `src/storage/db.ts` | Modified | Dedicated bounded ambient query (separate from `getRecentTurns`) |
89
+ | `src/core/storage-cleanup.ts` | Modified | Ambient TTL |
90
+ | test files alongside each | New/Modified | See tests below |
91
+
92
+ ### Implementation Sequence
93
+
94
+ 1. Extract `runtime.ts:591-612` into public `recordAmbient(ingress:
95
+ IngressMessage)`, preserving all guards (group-only, `ambient.enabled`,
96
+ non-DM). Pure refactor — behaviour identical.
97
+ - Read first: `src/core/runtime.ts:560-630`, `src/core/handler.ts:200-250`
98
+ - Verify: `bun run check` green; existing tests unchanged.
99
+ 2. Bounded ambient query in `db.ts` + wire the context builder to use it, so
100
+ ambient shares nothing with the `turnCount * 5` window.
101
+ - Read first: `src/storage/db.ts:1040-1080` and the context assembly that
102
+ consumes `getRecentTurns`
103
+ - Verify: unit test — ambient context returns ≤ N rows and user/assistant
104
+ turns are not displaced. **Blocker: step 3 must not ship without this.**
105
+ 3. Handler fix — replace the bare return at `handler.ts:240`:
106
+ ```ts
107
+ if (!shouldProcess && !isCommand) {
108
+ core.recordAmbient(ingress);
109
+ return;
110
+ }
111
+ ```
112
+ - Read first: `src/core/handler.ts:230-245`
113
+ - Verify: unit tests — writes for a group with `ambient.enabled`
114
+ unset/true; nothing for a DM; nothing when `ambient.enabled=false`;
115
+ called exactly once for a non-triggered group message and **not** for a
116
+ triggered one (no double-write).
117
+ 4. Retention TTL in `storage-cleanup.ts` for `role='ambient'` rows.
118
+ - Read first: `src/core/storage-cleanup.ts` (existing cleanup patterns)
119
+ - Verify: unit test — rows older than the TTL are removed, newer kept,
120
+ other roles untouched.
121
+
122
+ ---
123
+
124
+ ## Non-Negotiable Rules
125
+
126
+ - The handler fix (step 3) never ships without the context cap (step 2) and
127
+ TTL (step 4) — D-003.
128
+ - All existing guards survive the extraction: no ambient writes for DMs, none
129
+ when `ambient.enabled=false`.
130
+ - No double-write for triggered messages.
131
+
132
+ ---
133
+
134
+ ## Edge Cases & Risks
135
+
136
+ | Scenario | Handling |
137
+ |----------|---------|
138
+ | High-volume group floods ambient | Cap bounds the prompt; TTL bounds the DB; both configurable |
139
+ | `/`-prefixed unknown command in a group | Still falls through to ambient (coordinated with M3.2's routing change) |
140
+ | Triggered message | Processed normally, not also recorded ambient (regression test) |
141
+ | Ambient rows crowd out real turns | Dedicated query — structurally impossible to share the turn budget |
142
+ | Privacy: bot now stores all group chatter | TTL is the mitigation; called out in the PR (roadmap Completion §4) |
143
+
144
+ ---
145
+
146
+ ## Implementation Checklist
147
+
148
+ ### Phase 1 — Goal
149
+ - [x] Goal, User Stories, and MVP Scope written and reviewed
150
+
151
+ ### Phase 2 — Architecture
152
+ - [x] Architecture & Data section complete
153
+ - [x] Non-Negotiable Rules defined
154
+ - [x] Edge Cases covered
155
+ - [x] Context for Claude pointers filled in
156
+
157
+ ### Implementation (commit `4bf5505`)
158
+ - [x] `recordAmbient` extracted (step 1)
159
+ - [x] Bounded ambient query (step 2)
160
+ - [x] Handler fix (step 3)
161
+ - [x] Retention TTL (step 4)
162
+ - [x] 17 unit tests green; `bun run check` passes (1367 pass / 0 fail)
163
+ - [x] Deployed and running
164
+ - [ ] **Manual (needs a phone):** 3 untagged messages in a test group → 3 new
165
+ `role='ambient'` rows; tag the bot → it can recount them
166
+ - [ ] **Verify with data** once the group sees traffic:
167
+ `SELECT space_id, role, COUNT(*) FROM messages GROUP BY space_id, role`
168
+ — ambient for `greece2026` must go above 0 and keep growing
169
+ - [x] No secrets or `.env` files committed
170
+ - [x] No unrelated files modified
171
+
172
+ ---
173
+
174
+ ## Open Questions
175
+
176
+ - [x] Defaults shipped as proposed: `ambientContextRows: 30`,
177
+ `ambientTtlDays: 14`. Both are `mercury.yaml`/env-overridable
178
+ (`MERCURY_AMBIENT_CONTEXT_ROWS`, `MERCURY_AMBIENT_TTL_DAYS`) so they can be
179
+ tuned once real volume is observed, without a code change.
180
+
181
+ ---
182
+
183
+ ## Retrospective
184
+
185
+ **Residual risks / follow-ups:**
186
+ - Ambient volume in a busy family group is still unmeasured. The caps make it
187
+ safe, but the *right* numbers need a week of real data — revisit at the M2
188
+ checkpoint.
189
+ - Media attached to untriggered messages is still downloaded and discarded
190
+ (`bridge.normalize` runs before the gate). Pre-existing, not worsened by
191
+ this change — accepted, as the spec states.
192
+
193
+ **What changed from the plan:**
194
+ - Config went into `AppConfig` (`ambientContextRows`, `ambientTtlDays`) rather
195
+ than the per-space config keys the spec sketched, to match how
196
+ `inboxTtlDays`/`outboxTtlDays` already feed `storage-cleanup`. `ambient.enabled`
197
+ stays per-space as it was.
198
+ - Ambient cleanup runs **before** the spaces-directory scan. That scan
199
+ early-returns when the directory is unreadable, which would have silently
200
+ disabled retention on exactly the hosts most likely to need it.
201
+
202
+ **Key decisions made during implementation:**
203
+ - `getRecentTurns` now excludes ambient outright rather than merely fetching
204
+ more rows. Padding the LIMIT would have made starvation less likely but
205
+ still possible; two queries make it structurally impossible.
206
+ - The handler and `handleRawInput` paths are mutually exclusive by
207
+ construction (one returns before the other can run), which is what makes
208
+ "no double-write" a property rather than a check. A test pins it.
209
+
210
+ **Architecture impact** (update `docs/ARCHITECTURE.md` if any of these apply):
211
+ - [ ] New inter-package interaction (handler → runtime.recordAmbient)
@@ -0,0 +1,210 @@
1
+ # Command Routing Consistency
2
+
3
+ **Status**: Done
4
+ **Slug**: command-routing-consistency
5
+ **Goal**: whatsapp-bot-hardening
6
+ **Milestone**: 3
7
+ **Created**: 2026-08-09
8
+ **Last updated**: 2026-08-09
9
+
10
+ ---
11
+
12
+ ## Goal
13
+
14
+ > Part of goal [whatsapp-bot-hardening](../goals/whatsapp-bot-hardening/goal.md).
15
+ > HANDOFF Bugs 6b, 6d, 6e + redundancy cleanup — the rest of the command-surface
16
+ > audit after the destructive trio (M3.1).
17
+
18
+ Three consistency defects remain: (1) slash commands in groups are silently
19
+ ignored unless the trigger word is present — `routeInput` re-applies the
20
+ trigger gate that `handler.ts`'s `isCommand` bypass was meant to skip, so the
21
+ bypass accomplishes nothing except filing commands as ambient chat (this is
22
+ exactly how the two stranded baduk commands ended up in the DB); (2) chat and
23
+ CLI permission gates disagree — a promoted space admin denied `/spaces delete`
24
+ in chat can still destroy the current space via `mrctl`; (3) small argument/help
25
+ defects: `/resume <arg>` silently ignores its argument, `/help pause` returns
26
+ "No help available". Plus the audited redundancy cleanup (D-009).
27
+
28
+ ## User Stories
29
+
30
+ - As a group member, I want a bare `/spaces list` to work without tagging the
31
+ bot, because that's what the `/` prefix visibly promises.
32
+ - As the operator, I want chat and CLI to enforce the *same* permission answer
33
+ for the same action, so the tighter gate can't be bypassed.
34
+ - As any user, I want `/resume 5m` to either work as written or error — not
35
+ silently do something else — and `/help pause` to actually document `/pause`.
36
+
37
+ ## MVP Scope
38
+
39
+ **In scope:** trigger-free slash-command routing in groups (known commands
40
+ only; unknown `/foo` still falls through to ambient — D-007); lift
41
+ `spaces.delete` out of the blanket space-admin default grant so CLI matches
42
+ chat's seeded-admin gate (D-008); `/resume` rejects unknown args; `verbs`
43
+ metadata for `pause` incl. documenting the `^(\d+)(m|h)$` duration format; drop
44
+ `/model active`; rewrite `/compact` vs `/clear` help to state the
45
+ `session_boundary` vs `clear_boundary` difference (D-009).
46
+
47
+ **Out of scope:** merging `/compact` and `/clear` — deliberately not done blind
48
+ (D-009); any change to member default permissions (already minimal and correct).
49
+
50
+ ---
51
+
52
+ ## Context for Claude
53
+
54
+ - `src/core/router.ts:105-115` — `matchTrigger` returns `ignore` before the
55
+ slash parse; the ordering to fix.
56
+ - `src/core/handler.ts:236-240` — the `isCommand` bypass that currently
57
+ accomplishes nothing (coordinate with M2.2's `recordAmbient` call there).
58
+ - `src/core/router.ts:33` (`SEEDED_ADMIN_COMMANDS`) and
59
+ `src/core/permissions.ts:208-210` (blanket admin grant), `:192-197`
60
+ (member defaults — untouched).
61
+ - `src/core/commands.ts` — `SLASH_COMMANDS` (resume verbs, pause `verbs`
62
+ array, `/model active` removal, compact/clear descriptions);
63
+ `parsePauseDuration`.
64
+ - `HANDOFF.md` Bug 6 — full audit incl. the redundancy verdicts table.
65
+
66
+ ---
67
+
68
+ ## Architecture & Data
69
+
70
+ ### Data Models
71
+
72
+ None.
73
+
74
+ ### API Contracts
75
+
76
+ Permission semantics change: `spaces.delete` is no longer granted by default
77
+ space-admin role — only the seeded admin (config) passes `checkPerm` for it.
78
+ This is a **tightening**; document it in the PR behaviour-change list.
79
+
80
+ ### File & Folder Structure
81
+
82
+ | Path | New / Modified | Purpose |
83
+ |------|---------------|---------|
84
+ | `src/core/router.ts` | Modified | Parse known slash commands before/despite the trigger gate in groups |
85
+ | `src/core/permissions.ts` | Modified | Lift `spaces.delete` from the blanket admin grant |
86
+ | `src/core/commands.ts` | Modified | `/resume` arg rejection; `pause` verbs + duration docs; drop `/model active`; honest compact/clear help |
87
+ | test files alongside each | New/Modified | See tests below |
88
+
89
+ ### Implementation Sequence
90
+
91
+ 1. **Trigger-free slash routing in groups** — in `routeInput`, recognize a
92
+ known slash command before the trigger `ignore`; unknown `/foo` keeps
93
+ falling through (to ambient once M2.2 lands). Coordinate with the
94
+ `handler.ts` bypass so there's exactly one gate.
95
+ - Read first: `src/core/router.ts:100-120`, `src/core/handler.ts:230-245`,
96
+ M2.2's spec (shared code path)
97
+ - Verify: unit — bare `/spaces list` in a group resolves as a command;
98
+ unknown `/foo` in a group does **not** become a command.
99
+ 2. **Permission alignment** — remove `spaces.delete` from
100
+ `getDefaultPermissions`' admin set; confirm the seeded-admin path still
101
+ passes everywhere it should.
102
+ - Read first: `src/core/permissions.ts:185-215`, `src/core/router.ts:30-35`
103
+ - Verify: unit — a promoted (non-seeded) space admin gets the same denial
104
+ from the chat path and the CLI/API path.
105
+ 3. **`/resume` argument handling** — pass `verb` through; unknown/unsupported
106
+ args error with usage instead of silently acting as bare `/resume`.
107
+ - Read first: `src/core/commands.ts` (`executeResumeCommand` and its caller)
108
+ - Verify: unit — `/resume extra` errors; bare `/resume` unchanged.
109
+ 4. **`/help pause`** — add the `verbs` array so `formatCategoryHelp` works;
110
+ document the accepted duration format (`\d+m|\d+h`).
111
+ - Read first: `src/core/commands.ts:55-67`, `formatCategoryHelp`,
112
+ `parsePauseDuration`
113
+ - Verify: unit — `/help pause` returns real help text.
114
+ 5. **Redundancy cleanup** — remove `/model active` (the fact lives in
115
+ `/model list`'s `← active` marker); rewrite `/compact` and `/clear`
116
+ descriptions to state the actual boundary difference (`db.ts:913` vs `:936`).
117
+ - Read first: `src/core/commands.ts` model/compact/clear entries,
118
+ `src/storage/db.ts:905-940`
119
+ - Verify: unit/help-snapshot — `/model active` gone from help and routing;
120
+ new descriptions render.
121
+
122
+ Step 1 should land after M3.1's step 4 (same router file, `/`-handling).
123
+
124
+ ---
125
+
126
+ ## Non-Negotiable Rules
127
+
128
+ - Unknown `/foo` in a group must never execute as a command — it falls through
129
+ (ambient), preserving the safe default.
130
+ - Member default permissions are not touched.
131
+ - Every user-visible behaviour change here goes in the PR's explicit
132
+ behaviour-change list (roadmap Completion §6).
133
+
134
+ ---
135
+
136
+ ## Edge Cases & Risks
137
+
138
+ | Scenario | Handling |
139
+ |----------|---------|
140
+ | Message that merely *starts* with `/` but isn't a command (`/shrug`) | Not in `SLASH_COMMANDS` → falls through to ambient, never executes |
141
+ | DM slash commands | Unchanged — DMs never required the trigger |
142
+ | Seeded admin also holds space-admin role | Passes both gates before and after — no regression |
143
+ | Muscle memory: users who typed `Wiz /spaces list` | Still works — trigger + command remains valid |
144
+ | `/model active` users | Removal listed in PR behaviour changes; `/model list` shows the same fact |
145
+
146
+ ---
147
+
148
+ ## Implementation Checklist
149
+
150
+ ### Phase 1 — Goal
151
+ - [x] Goal, User Stories, and MVP Scope written and reviewed
152
+
153
+ ### Phase 2 — Architecture
154
+ - [x] Architecture & Data section complete
155
+ - [x] Non-Negotiable Rules defined
156
+ - [x] Edge Cases covered
157
+ - [x] Context for Claude pointers filled in
158
+
159
+ ### Implementation (commit `4f05fbb`)
160
+ - [x] Trigger-free slash routing (step 1)
161
+ - [x] Permission alignment (step 2) — via a route-level check, see retrospective
162
+ - [x] `/resume` args (step 3)
163
+ - [x] `/help pause` (step 4)
164
+ - [x] Redundancy cleanup (step 5)
165
+ - [x] 11 unit tests green; `bun run check` passes (1404 pass / 0 fail)
166
+ - [ ] **Manual (needs a phone):** bare `/spaces list` untagged in a test group
167
+ - [x] No secrets or `.env` files committed
168
+ - [x] No unrelated files modified
169
+
170
+ ---
171
+
172
+ ## Open Questions
173
+
174
+ - [x] none
175
+
176
+ ---
177
+
178
+ ## Retrospective
179
+
180
+ **Residual risks / follow-ups:**
181
+ - With no seeded admins configured, `spaces.delete` keeps its old behaviour
182
+ (any space admin passes). Deliberate — the alternative is a deployment where
183
+ nobody can delete a space — but it means the alignment only bites where
184
+ `permissions.admins` is set. Accepted.
185
+
186
+ **What changed from the plan:**
187
+ - **D-008 was not implementable as written.** The plan said to lift
188
+ `spaces.delete` out of the blanket admin grant. But seeded admins are
189
+ *seeded into the database as role `admin`* (`resolveRole` →
190
+ `db.seedAdmins`), so seeded and promoted admins are indistinguishable by
191
+ role — removing the permission from `admin` would have locked out everyone,
192
+ including the person the chat gate is designed to admit. The alignment
193
+ instead lives in the delete route, which now asks the same seeded-admin
194
+ question the chat path asks. Same outcome, and it does not depend on a role
195
+ distinction that does not exist.
196
+
197
+ **Key decisions made during implementation:**
198
+ - Slash-command admission is by *known command name*, not by the `/` prefix:
199
+ `namesKnownCommand()` checks `SLASH_COMMANDS` and `CHAT_COMMANDS`, so
200
+ `/shrug` still falls through to ambient. This keeps the safe default while
201
+ making the handler's long-dead `isCommand` bypass finally mean something.
202
+ - `/model active` kept as a redirect rather than deleted outright, so muscle
203
+ memory lands on an explanation instead of "unknown verb".
204
+ - `/compact` and `/clear` were not merged. Their real difference was verified
205
+ in the code first (`min_message_id` is permanent, `clear_boundary` is reset
206
+ after each read) and then written into the help text, which is what the
207
+ audit actually asked for.
208
+
209
+ **Architecture impact** (update `docs/ARCHITECTURE.md` if any of these apply):
210
+ - [ ] none expected