mercury-agent 0.16.0 → 0.16.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +6 -0
- package/container/Dockerfile +6 -1
- package/container/Dockerfile.base +6 -1
- package/docs/README.md +50 -0
- package/docs/authoring-profiles.md +8 -0
- package/docs/behavior-layers.md +193 -0
- package/docs/extensions.md +2 -0
- package/docs/goals/football-reporter-profile/decisions.md +479 -0
- package/docs/goals/football-reporter-profile/goal.md +91 -0
- package/docs/goals/football-reporter-profile/roadmap.md +196 -0
- package/docs/goals/whatsapp-bot-hardening/README.md +43 -0
- package/docs/goals/whatsapp-bot-hardening/archive/ambient-group-context.md +211 -0
- package/docs/goals/whatsapp-bot-hardening/archive/command-routing-consistency.md +210 -0
- package/docs/goals/whatsapp-bot-hardening/archive/destructive-command-guards.md +218 -0
- package/docs/goals/whatsapp-bot-hardening/archive/docs-extension.md +182 -0
- package/docs/goals/whatsapp-bot-hardening/archive/handoff-2026-08-09.md +588 -0
- package/docs/goals/whatsapp-bot-hardening/archive/media-size-and-silent-drop.md +178 -0
- package/docs/goals/whatsapp-bot-hardening/archive/member-memory.md +199 -0
- package/docs/goals/whatsapp-bot-hardening/archive/message-author-attribution.md +211 -0
- package/docs/goals/whatsapp-bot-hardening/archive/run-from-source-switchover.md +212 -0
- package/docs/goals/whatsapp-bot-hardening/archive/setup-plan-2026-08-06.md +1117 -0
- package/docs/goals/whatsapp-bot-hardening/decisions.md +200 -0
- package/docs/goals/whatsapp-bot-hardening/goal.md +75 -0
- package/docs/goals/whatsapp-bot-hardening/open-threads.md +396 -0
- package/docs/goals/whatsapp-bot-hardening/roadmap.md +245 -0
- package/docs/memory.md +30 -0
- package/docs/profile-guide.md +483 -0
- package/docs/refactor/.gitkeep +0 -0
- package/docs/refactor/archive/.gitkeep +0 -0
- package/docs/refactor/audits/.gitkeep +0 -0
- package/docs/refactor/backlog/.gitkeep +0 -0
- package/docs/skills-guide.md +154 -0
- package/examples/extensions/README.md +1 -0
- package/examples/extensions/feed-watch/config.ts +354 -0
- package/examples/extensions/feed-watch/digest.ts +344 -0
- package/examples/extensions/feed-watch/feeds.ts +341 -0
- package/examples/extensions/feed-watch/index.ts +302 -0
- package/examples/extensions/feed-watch/items.ts +171 -0
- package/examples/extensions/feed-watch/match.ts +107 -0
- package/examples/extensions/feed-watch/prompts/verify.md +35 -0
- package/examples/extensions/feed-watch/skill/SKILL.md +128 -0
- package/examples/extensions/feed-watch/watch.ts +548 -0
- package/examples/extensions/longview/hook.ts +179 -16
- package/examples/extensions/longview/index.ts +13 -2
- package/examples/extensions/longview/prompts/summarize.md +3 -2
- package/examples/extensions/longview/render/telegraph-nodes.ts +21 -1
- package/examples/extensions/longview/summarize.ts +57 -14
- package/examples/extensions/napkin/index.ts +334 -86
- package/examples/extensions/napkin/pi-spawn.ts +196 -0
- package/examples/extensions/pinchtab/index.ts +37 -4
- package/examples/extensions/pinchtab/lib/session-injector.ts +31 -9
- package/examples/extensions/pinchtab/skill/SKILL.md +29 -0
- package/examples/profiles/_template/AGENTS.md +138 -0
- package/examples/profiles/_template/README.md +65 -0
- package/examples/profiles/_template/config.yaml +57 -0
- package/examples/profiles/_template/tasks/daily.md +32 -0
- package/examples/profiles/football-reporter/AGENTS.md +219 -0
- package/examples/profiles/football-reporter/README.md +138 -0
- package/examples/profiles/football-reporter/config.yaml +174 -0
- package/examples/profiles/football-reporter/seed/MEMORY.md +42 -0
- package/examples/profiles/football-reporter/seed/episodes/barcelona-2026-27.md +25 -0
- package/examples/profiles/football-reporter/seed/episodes/beitar-jerusalem-2026-27.md +25 -0
- package/examples/profiles/football-reporter/seed/episodes/maccabi-haifa-2026-27.md +24 -0
- package/examples/profiles/football-reporter/seed/episodes/man-united-2026-27.md +26 -0
- package/examples/profiles/football-reporter/seed/episodes/real-madrid-2026-27.md +25 -0
- package/examples/profiles/football-reporter/seed/napkin-distill.md +55 -0
- package/examples/profiles/football-reporter/tasks/daily-article.md +142 -0
- package/package.json +8 -5
- package/src/adapters/whatsapp.ts +75 -13
- package/src/bridges/whatsapp.ts +6 -4
- package/src/cli/mrctl.ts +9 -1
- package/src/core/handler.ts +51 -6
- package/src/core/routes/dashboard.ts +106 -5
- package/src/core/routes/tasks.ts +28 -0
- package/src/core/runtime.ts +63 -2
- package/src/core/task-scheduler.ts +103 -12
- package/src/extensions/catalog.ts +9 -0
- package/src/profile/space-profile.ts +780 -0
- package/src/storage/db.ts +215 -0
- package/src/types.ts +56 -0
|
@@ -0,0 +1,196 @@
|
|
|
1
|
+
# Roadmap: Football Reporter Profile
|
|
2
|
+
|
|
3
|
+
**Goal**: [football-reporter-profile](goal.md)
|
|
4
|
+
**Last updated**: 2026-08-22 (M2.3 shipped; audit extension M2.3–M2.5, D-016–D-018; napkin revisit — M2.3 re-cut, D-019)
|
|
5
|
+
|
|
6
|
+
> Sequence chosen by the owner on 2026-08-20: (1) profile + one editorial
|
|
7
|
+
> layer → (2) deterministic feed → (3) proof of quality, with a parallel bug
|
|
8
|
+
> track for anything unrelated, then (4) regression + re-audit + monitoring.
|
|
9
|
+
|
|
10
|
+
## Milestones
|
|
11
|
+
|
|
12
|
+
### Milestone 1: One owner for the editorial rules, and an article worth tapping
|
|
13
|
+
> M1.1 is zero code change to the bot and observable on the live group within
|
|
14
|
+
> 48 h; after it every football rule lives in the repo, deploys by script, and
|
|
15
|
+
> a `check` proves no drift. M1.2 redesigns the daily article (Telegraph page
|
|
16
|
+
> + fuller in-chat summary, D-005/D-010) — the one code-bearing story in this
|
|
17
|
+
> milestone (`longview`).
|
|
18
|
+
|
|
19
|
+
| ID | Story | Slug | Depends on | Status |
|
|
20
|
+
|----|-------|------|------------|--------|
|
|
21
|
+
| M1.1 | Space profile folder + `AGENTS.md` editorial standard + apply/check/dump script; slim task prompts; delete `club-coverage` | football-reporter-space-profile | — | done |
|
|
22
|
+
| M1.2 | Daily article redesign: top-tier Telegraph article (headline, standfirst, sections, pull-quote, certainty box, sources), author-supplied fuller chat summary in `longview` (F3) | daily-article-redesign | M1.1 | done |
|
|
23
|
+
|
|
24
|
+
**Checkpoint:** after two daily articles and ~10 scans on the new `AGENTS.md`:
|
|
25
|
+
no `"Done."`/filler reached the group; every posted item shows `(source)` and
|
|
26
|
+
a date; 0 emoji in scheduled output; `club-coverage` is gone from the prefs
|
|
27
|
+
panel; `scripts/space-profile.ts check` is clean; the redesigned page passes
|
|
28
|
+
the visual checklist on a phone, the chat summary carries the key facts, and
|
|
29
|
+
the group's reaction to the first two articles is recorded in the goal's notes.
|
|
30
|
+
|
|
31
|
+
> **Checkpoint passed 2026-08-21** on one article rather than two, by owner's
|
|
32
|
+
> decision — M2.1 depends on M1.1 only, so nothing about feed-watch waits on the
|
|
33
|
+
> second article. Evidence from the first 24 h after the deploy (ledger runs #1
|
|
34
|
+
> and #2, full detail in `docs/pending-verification.md`):
|
|
35
|
+
>
|
|
36
|
+
> - **Passed:** `NO_UPDATE` fired unattended on the 00:00 scan (9-char reply,
|
|
37
|
+
> `outcome=no_update`, nothing posted); the 09:00 article carries 1 `#` title,
|
|
38
|
+
> 8 `##` sections, exactly 1 `>>` pull-quote, single-star bold only, 0 emoji
|
|
39
|
+
> and 0 exclamation marks; the summary block is 452 chars / 4 lines, inside
|
|
40
|
+
> both the 900 cap and the 3–6 band; one model call per run, confirmed against
|
|
41
|
+
> `token_usage`; `space-profile check` clean on `apply` and on the next sync;
|
|
42
|
+
> `club-coverage` gone.
|
|
43
|
+
> - **Failed, and folded into M3.1's lint per [D-014](decisions.md):** no item
|
|
44
|
+
> carries a publication date — on the scan *and* on the article summary, though
|
|
45
|
+
> the article body of the same run dated 5 items.
|
|
46
|
+
> - **Still open:** the Telegraph page cannot be confirmed from the host (since
|
|
47
|
+
> M1.2 the `storedReply` seam stores the raw article, so the delivered link is
|
|
48
|
+
> in no table); the phone RTL check; the group's reaction; the second article.
|
|
49
|
+
> - **Baseline for M2:** scan $0.456/run, article $0.769/run; whole-project daily
|
|
50
|
+
> spend $7.53 (08-19) and $7.19 (08-20), against M2's ~$5.30/day direction.
|
|
51
|
+
|
|
52
|
+
### Milestone 2: A deterministic feed in front of the model
|
|
53
|
+
> Scans stop being blind; empty scans cost nothing; a fallback-leg run never
|
|
54
|
+
> posts.
|
|
55
|
+
|
|
56
|
+
| ID | Story | Slug | Depends on | Status |
|
|
57
|
+
|----|-------|------|------------|--------|
|
|
58
|
+
| M2.1 | Feed Watch — host-side poller, digest into the article, verify one-shots (existing spec, absorbed) | feed-watch | M1.1 | done (merged `e07b239` 2026-08-21; **deploy not yet run** — runbook in `pending-verification.md`) |
|
|
59
|
+
| M2.2 | Scheduled task model-leg policy: primary-or-skip for research runs (F2) | scheduled-task-model-leg-policy | — | deferred until after M3.1 (D-015) |
|
|
60
|
+
| M2.3 | Reporter notebook, re-cut 2026-08-21 (D-019): topic notes are napkin-shaped **episodes** the profile seeds once and napkin maintains nightly, written in-turn before posting; seeded `MEMORY.md` for standing facts; priorities tie-break in `AGENTS.md`; `workspace_seed` in `space-profile`; per-space distillation addendum in napkin; member perms and stale-note hygiene (audit R1/R6/R7, D-016, D-019) | reporter-notebook | M1.1 | done (merged 2026-08-22; **deploy not yet run** — runbook in `pending-verification.md`) |
|
|
61
|
+
| M2.4 | One voice per space: `persona.exclusive` drops the global character and global `AGENTS.md` for a space that owns its standard; platform-prompt capability paragraphs gated on the caller's permissions (audit R2, D-017) | one-voice-per-space | — | in-progress (worktree `worktree-feature-one-voice-per-space`; host; image rebuild) |
|
|
62
|
+
| M2.5 | Chat window holds conversation, not procedure (scheduled prompts out of the turn count, no halving on swipe-reply); silence is a legal chat reply (audit R3/R4, D-018) | chat-window-and-silence | — | backlog (host) |
|
|
63
|
+
|
|
64
|
+
> **Revisit 2026-08-21 — napkin works now** (`docs/notes/napkin-revisit-2026-08-21.md`). **Approved by the owner the same day ([D-019](decisions.md)):** M2.3 re-cut so the topic notes are napkin-maintained episodes under `knowledge/episodes/` (the only injected dir) that the profile seeds once, plus a per-space distillation addendum. Nothing removed; M2.4/M2.5/M3.x unchanged. Idea parked: `docs/ideas/napkin-member-notes.md`.
|
|
65
|
+
|
|
66
|
+
> **M2.2 is largely pre-empted by the 2026-08-20 Opus pin.** `runtime.ts:1871`
|
|
67
|
+
> turns `space_config['model.active']` into a **one-element** `MODEL_CHAIN`
|
|
68
|
+
> for the container — its own comment reads "eliminates automatic
|
|
69
|
+
> fallback" — and all 6 spaces carry that row. So no run of any task in any
|
|
70
|
+
> space can reach a fallback leg today, and a leg failure already surfaces as
|
|
71
|
+
> a task error that retries and then reports, posting nothing. What M2.2 would
|
|
72
|
+
> still add is (a) the guarantee being **per task** rather than per space, so
|
|
73
|
+
> it survives someone restoring chat fallback by deleting the pin, and (b) a
|
|
74
|
+
> skip that is **distinguishable** in the ledger from any other error
|
|
75
|
+
> (`last_status` is only `ok`/`error`). Owner's call at this checkpoint
|
|
76
|
+
> whether that is worth doing now or after M3.2 makes it measurable.
|
|
77
|
+
|
|
78
|
+
> **M2.2 deferred, not dropped (2026-08-21, [D-015](decisions.md)).** All 12 live
|
|
79
|
+
> runs since the profile applied ran on `claude-opus-5`, so the checkpoint's "zero
|
|
80
|
+
> posts from a non-primary leg" is already true by construction. Architecture stays
|
|
81
|
+
> unfilled until M3.1 can measure whether a per-task guarantee earns its keep.
|
|
82
|
+
|
|
83
|
+
**Checkpoint:** task #28 and its one-shots are gone; the first `feed-watch`
|
|
84
|
+
one-shot produced either a sourced 1–3 line post or nothing; the 09:00 article
|
|
85
|
+
trace carries `<feed_watch_digest … items="N">`; a week of spend is reported
|
|
86
|
+
(target direction: well under ~$5.30/day — reported, not gated); zero group
|
|
87
|
+
posts from a non-primary leg. **Added 2026-08-21 (audit):** `agent.trace_runs`
|
|
88
|
+
is on; a factual chat question's trace shows the topic note being read and the
|
|
89
|
+
next article appends to one; the football space runs with `persona.exclusive`
|
|
90
|
+
and a member run's system prompt carries no preference-management paragraph; a
|
|
91
|
+
swipe-reply follow-up sees the bot's own post from the previous evening; a
|
|
92
|
+
"don't answer" message in `main` gets no reply.
|
|
93
|
+
|
|
94
|
+
> **Extended 2026-08-21 ([D-016](decisions.md)–[D-018](decisions.md)).** The
|
|
95
|
+
> 2026-08-21 audit (`docs/notes/football-audit-2026-08-21.md`) traced both
|
|
96
|
+
> owner-reported failures — a hallucinated name and a joke refused as a
|
|
97
|
+
> permission request — to mechanisms, not missing rules: the window, not a
|
|
98
|
+
> memory; four personas, not one; no silent chat path. M2.3–M2.5 are those
|
|
99
|
+
> mechanisms. Per the owner, nothing in them is tailored to the bot's state
|
|
100
|
+
> today — M3.1/M4.1 are assumed to land; M2.3 is what gives M3.1's provenance
|
|
101
|
+
> lint a source to check against. **Order:** M2.3 first (profile-side, no
|
|
102
|
+
> restart); M2.4 and M2.5 are independent host stories and may run in
|
|
103
|
+
> parallel with M3.1.
|
|
104
|
+
>
|
|
105
|
+
> **Build order note (2026-08-20, D-012):** M3.2 is being implemented before
|
|
106
|
+
> M2.1. It has no dependencies, and an instrument installed after the change
|
|
107
|
+
> it measures cannot show the before — feed-watch should ship against a
|
|
108
|
+
> baseline the ledger already recorded. Milestone membership is unchanged.
|
|
109
|
+
|
|
110
|
+
### Milestone 3: Proof of quality
|
|
111
|
+
> The rules are testable without reading the chat.
|
|
112
|
+
|
|
113
|
+
| ID | Story | Slug | Depends on | Status |
|
|
114
|
+
|----|-------|------|------------|--------|
|
|
115
|
+
| M3.1 | Fixture harness: dated fixture items through the verify/article prompts in the private space + a deterministic output lint | football-reporter-fixture-harness | M1.1 | backlog |
|
|
116
|
+
| M3.2 | Scheduled-run ledger: model, cost, outcome per run; dashboard badge + drawer, `mrctl tasks runs` | scheduled-run-ledger | — | done (2026-08-20, ahead of M2 — D-012) |
|
|
117
|
+
|
|
118
|
+
**Checkpoint:** harness green on the fixture set (fresh item passes, undated
|
|
119
|
+
and stale items dropped, emoji/HTML/line-count lint fires on seeded bad
|
|
120
|
+
replies); every scheduled run of the last 48 h has a ledger row.
|
|
121
|
+
|
|
122
|
+
### Milestone 4: Regression and monitoring
|
|
123
|
+
> Every previously reported defect has a check that would catch its return;
|
|
124
|
+
> a re-audit confirms; new problems surface from the ledger.
|
|
125
|
+
|
|
126
|
+
| ID | Story | Slug | Depends on | Status |
|
|
127
|
+
|----|-------|------|------------|--------|
|
|
128
|
+
| M4.1 | Regression suite for F1–F8 + `scripts/football-audit.ts` data dump + two-week re-audit note | football-regression-suite | M2.1, M3.1, M3.2 | blocked |
|
|
129
|
+
|
|
130
|
+
**Checkpoint:** regression suite green in `bun run check`; re-audit note
|
|
131
|
+
(`docs/notes/football-group-reaudit-<date>.md`) finds no recurrence of F1–F7
|
|
132
|
+
and reports F8 spend; the ledger is the first place a new problem shows.
|
|
133
|
+
|
|
134
|
+
## Bug track (parallel — not stories, per D-007)
|
|
135
|
+
|
|
136
|
+
Filed at planning in `docs/bugs/`; fix with `/f-bug-fix` whenever convenient,
|
|
137
|
+
independent of the milestones. Any *new* defect found while executing a story
|
|
138
|
+
is filed the same way, never fixed inside the story.
|
|
139
|
+
|
|
140
|
+
| Audit ref | Bug doc | Severity | Regression check lands in |
|
|
141
|
+
|-----------|---------|----------|---------------------------|
|
|
142
|
+
| F5 | `reply-to-member-triggers-bot.md` | moderate | M4.1 |
|
|
143
|
+
| F7a | `chat-output-html-tags-not-stripped.md` | minor | M4.1 (lint also in M3.1) |
|
|
144
|
+
| F7d | `responded-in-suffix-not-stored.md` | minor | M4.1 |
|
|
145
|
+
| F7e | `bot-self-name-is-phone-in-reply-to.md` | minor | M4.1 |
|
|
146
|
+
| — | `model-switch-disables-fallback.md` | moderate | M4.1 |
|
|
147
|
+
| audit-2 C9 | `feed-watch-seen-ttl-not-refreshed.md` | moderate | M4.1 (fix **before** the feed-watch deploy) |
|
|
148
|
+
| audit-2 C10 | `feed-watch-runas-non-user-principal.md` | moderate | M4.1 (fix **before** the feed-watch deploy) |
|
|
149
|
+
|
|
150
|
+
The two `feed-watch-*` rows come from the 2026-08-21 audit's second pass
|
|
151
|
+
(`docs/notes/football-audit-2026-08-21.md` §3.1): a `seen` TTL that counts
|
|
152
|
+
from first sighting (a wave of stale items re-enters every 72 h) and a
|
|
153
|
+
`runAs` that can be the literal string `dashboard` (phantom member row,
|
|
154
|
+
member permissions). Both are small `watch.ts` changes with designed fixes
|
|
155
|
+
in the bug docs; both gate the deploy.
|
|
156
|
+
|
|
157
|
+
The last row is not from the 48h audit. It was found on 2026-08-20 while
|
|
158
|
+
expanding the live `model.chain` from 2 legs to 5 (opus-5, sonnet-5,
|
|
159
|
+
gemini-3.1-flash-lite, fable-5, haiku-4-5): `/model switch` truncates the chain
|
|
160
|
+
to the chosen leg and no verb restores it.
|
|
161
|
+
|
|
162
|
+
That same day the owner chose **Opus-only, no fallback, but all 5 switchable**,
|
|
163
|
+
and this truncation is the only mechanism in the codebase that delivers it — so
|
|
164
|
+
all 6 spaces are now deliberately pinned to `anthropic:claude-opus-5` and the
|
|
165
|
+
chain serves purely as the `/model switch` menu. The bug is therefore half
|
|
166
|
+
feature: what stays broken is that nothing can clear the pin, and that
|
|
167
|
+
`/model list` still renders unreachable legs as if they were fallbacks. Any fix
|
|
168
|
+
must be additive — see the bug doc before touching `resolveActiveIndex`.
|
|
169
|
+
|
|
170
|
+
Bearing on the milestones: every run is now single-leg, so `max_retries_per_leg`
|
|
171
|
+
is the only resilience left, and the M3.2 run ledger should record the model per
|
|
172
|
+
run rather than assume the chain.
|
|
173
|
+
|
|
174
|
+
Not bugs (handled by stories): F1 → resolved 2026-08-19, locked in by M1.1;
|
|
175
|
+
F2 → M2.2; F3 → M1.2; F4 → M1.1 + M2.1 + M3.1; F6 → M1.1 honesty rule;
|
|
176
|
+
F7c → M1.1 + M3.1 lint; F7f (Hebrew Telegraph slugs) → investigated in
|
|
177
|
+
M1.2, likely accepted; F8 → M2.1 + M3.2.
|
|
178
|
+
|
|
179
|
+
## Dependency Graph
|
|
180
|
+
|
|
181
|
+
```
|
|
182
|
+
M1.1 ─┬─► M1.2
|
|
183
|
+
├─► M2.1 ─┐
|
|
184
|
+
├─► M2.3 ─┤ (M3.1's provenance check reads M2.3's notebook — soft)
|
|
185
|
+
└─► M3.1 ─┼─► M4.1
|
|
186
|
+
M2.2 (independent) │
|
|
187
|
+
M2.4 (independent) ──┤
|
|
188
|
+
M2.5 (independent) ──┤
|
|
189
|
+
M3.2 (independent) ──┘
|
|
190
|
+
```
|
|
191
|
+
|
|
192
|
+
M1.1 gates everything that touches the space's prompts. M2.2, M2.4, M2.5 and
|
|
193
|
+
M3.2 are core changes with no football dependency and can run in parallel —
|
|
194
|
+
M3.2 is worth starting early because its ledger is what makes the M2 and M4
|
|
195
|
+
checkpoints measurable. M2.3 is profile-side and gates only the notebook half
|
|
196
|
+
of M3.1's provenance check; the rest of M3.1 is unblocked.
|
|
@@ -0,0 +1,43 @@
|
|
|
1
|
+
# whatsapp-bot-hardening — start here
|
|
2
|
+
|
|
3
|
+
Ops and planning docs for running Mercury as a live WhatsApp bot. Everything in
|
|
4
|
+
this folder is documentation; nothing here is executed.
|
|
5
|
+
|
|
6
|
+
## Where things are
|
|
7
|
+
|
|
8
|
+
| Path | What |
|
|
9
|
+
|------|------|
|
|
10
|
+
| [`open-threads.md`](open-threads.md) | **Start here.** Everything still outstanding. |
|
|
11
|
+
| [`roadmap.md`](roadmap.md) | What shipped, milestone by milestone |
|
|
12
|
+
| [`decisions.md`](decisions.md) | D-001…D-014, with reasoning and revisit conditions |
|
|
13
|
+
| [`goal.md`](goal.md) | Goal statement, constraints, success criteria |
|
|
14
|
+
| [`archive/`](archive/) | Completed work — the 8 story specs, the setup plan, the bug-investigation handoff |
|
|
15
|
+
|
|
16
|
+
## Archived — read for context, don't plan from them
|
|
17
|
+
|
|
18
|
+
- `archive/handoff-2026-08-09.md` — the bug investigation. Every fix in it
|
|
19
|
+
shipped; the root-cause analysis is still the best map of *why*.
|
|
20
|
+
- `archive/setup-plan-2026-08-06.md` — how the bot was built. Phases 0–5
|
|
21
|
+
complete; the Gemini and npm-install sections are superseded.
|
|
22
|
+
- `archive/<story>.md` — one spec per hardening story, with retrospectives.
|
|
23
|
+
|
|
24
|
+
Both archived files carry a header listing exactly what in them has gone stale.
|
|
25
|
+
|
|
26
|
+
## Where the code lives
|
|
27
|
+
|
|
28
|
+
Since 2026-08-10 there is **one** source of truth:
|
|
29
|
+
|
|
30
|
+
| Tree | Role |
|
|
31
|
+
|------|------|
|
|
32
|
+
| `D:\Projects\mercury` (Windows) | **Source of truth.** All edits and commits happen here. Pushes to `origin` once contributor access lands. |
|
|
33
|
+
| `~/mercury-src` (WSL) | Run-only mirror. The systemd unit `mercury` runs `bun` directly against this checkout — no build step. Never commit here. |
|
|
34
|
+
| `~/whatsapp-bot` (WSL) | Runtime data only — config, `.env`, SQLite DB. The unit's working directory. |
|
|
35
|
+
|
|
36
|
+
After changing code on Windows, push it to the running bot with:
|
|
37
|
+
|
|
38
|
+
```bash
|
|
39
|
+
wsl.exe -d Ubuntu -- bash -lc '~/sync-mercury.sh'
|
|
40
|
+
```
|
|
41
|
+
|
|
42
|
+
That fetches from the Windows clone, hard-resets the mirror to match, reinstalls
|
|
43
|
+
dependencies only if `bun.lock` changed, and restarts the service.
|
|
@@ -0,0 +1,211 @@
|
|
|
1
|
+
# Ambient Group Context
|
|
2
|
+
|
|
3
|
+
**Status**: Done
|
|
4
|
+
**Slug**: ambient-group-context
|
|
5
|
+
**Goal**: whatsapp-bot-hardening
|
|
6
|
+
**Milestone**: 2
|
|
7
|
+
**Created**: 2026-08-09
|
|
8
|
+
**Last updated**: 2026-08-09
|
|
9
|
+
|
|
10
|
+
---
|
|
11
|
+
|
|
12
|
+
## Goal
|
|
13
|
+
|
|
14
|
+
> Part of goal [whatsapp-bot-hardening](../goals/whatsapp-bot-hardening/goal.md).
|
|
15
|
+
> HANDOFF Bug 1 — highest impact.
|
|
16
|
+
|
|
17
|
+
The bot cannot see any group message it isn't tagged in. `ambient.enabled`
|
|
18
|
+
defaults to true and the ambient-write branch works as designed — but it lives
|
|
19
|
+
in `runtime.ts:591-612` behind `processIngress`, which is only reached **after**
|
|
20
|
+
the early return at `handler.ts:237-240` has already dropped untriggered
|
|
21
|
+
messages. It is dead code: across the whole DB there are exactly 2 ambient rows,
|
|
22
|
+
both stray `/`-prefixed messages that leaked through the `isCommand` bypass.
|
|
23
|
+
Upstream has not fixed this (`handler.ts` byte-identical on origin/main).
|
|
24
|
+
Depends on `message-author-attribution` (M2.1) so ambient rows carry authors
|
|
25
|
+
from the start, and ships **only with** its own context cap + retention TTL
|
|
26
|
+
(D-003).
|
|
27
|
+
|
|
28
|
+
## User Stories
|
|
29
|
+
|
|
30
|
+
- As a group member, I want to tag the bot and ask "what did we just decide?"
|
|
31
|
+
and have it answer from the conversation it silently observed.
|
|
32
|
+
- As the operator, I want ambient storage bounded (context budget + TTL) so
|
|
33
|
+
every group message becoming a row doesn't starve real turns or grow the DB
|
|
34
|
+
forever.
|
|
35
|
+
|
|
36
|
+
## MVP Scope
|
|
37
|
+
|
|
38
|
+
**In scope:** public `recordAmbient(ingress)` on `MercuryCoreRuntime`
|
|
39
|
+
(extracted from `runtime.ts:591-612`, guards intact); call it from the
|
|
40
|
+
handler's early-return path; a dedicated bounded ambient context query; ambient
|
|
41
|
+
retention TTL in `storage-cleanup.ts`.
|
|
42
|
+
|
|
43
|
+
**Out of scope:** avoiding media downloads for untriggered messages —
|
|
44
|
+
`bridge.normalize` already runs before the gate today (pre-existing cost, not
|
|
45
|
+
added by this fix); tuning ambient *relevance* (which rows get injected) beyond
|
|
46
|
+
a bounded recency window.
|
|
47
|
+
|
|
48
|
+
---
|
|
49
|
+
|
|
50
|
+
## Context for Claude
|
|
51
|
+
|
|
52
|
+
- `src/core/handler.ts:237-240` — the early return that kills ambient; `ingress`
|
|
53
|
+
is already built at `:221`, so the data is in hand.
|
|
54
|
+
- `src/core/runtime.ts:591-612` — the ambient-write branch to extract
|
|
55
|
+
(group-only + `ambient.enabled` + non-DM guards must survive the move).
|
|
56
|
+
- `src/storage/db.ts:1047` (`getRecentTurns`) and `:1068` (`turnCount * 5`
|
|
57
|
+
window) — the budget ambient must NOT share.
|
|
58
|
+
- `src/core/storage-cleanup.ts` — where the TTL goes.
|
|
59
|
+
- `HANDOFF.md` Bug 1 — full analysis and evidence.
|
|
60
|
+
|
|
61
|
+
---
|
|
62
|
+
|
|
63
|
+
## Architecture & Data
|
|
64
|
+
|
|
65
|
+
### Data Models
|
|
66
|
+
|
|
67
|
+
No new columns (`role='ambient'` rows already exist as a concept). New
|
|
68
|
+
config surface:
|
|
69
|
+
|
|
70
|
+
```yaml
|
|
71
|
+
ambient:
|
|
72
|
+
context_rows: 30 # max ambient rows injected per turn (own query, own cap)
|
|
73
|
+
retention_days: 14 # TTL enforced by storage-cleanup
|
|
74
|
+
```
|
|
75
|
+
|
|
76
|
+
Defaults decided in code review; both must exist before the handler fix ships.
|
|
77
|
+
|
|
78
|
+
### API Contracts
|
|
79
|
+
|
|
80
|
+
None.
|
|
81
|
+
|
|
82
|
+
### File & Folder Structure
|
|
83
|
+
|
|
84
|
+
| Path | New / Modified | Purpose |
|
|
85
|
+
|------|---------------|---------|
|
|
86
|
+
| `src/core/runtime.ts` | Modified | Extract `recordAmbient(ingress)` as a public method |
|
|
87
|
+
| `src/core/handler.ts` | Modified | Call `core.recordAmbient(ingress)` before the early return |
|
|
88
|
+
| `src/storage/db.ts` | Modified | Dedicated bounded ambient query (separate from `getRecentTurns`) |
|
|
89
|
+
| `src/core/storage-cleanup.ts` | Modified | Ambient TTL |
|
|
90
|
+
| test files alongside each | New/Modified | See tests below |
|
|
91
|
+
|
|
92
|
+
### Implementation Sequence
|
|
93
|
+
|
|
94
|
+
1. Extract `runtime.ts:591-612` into public `recordAmbient(ingress:
|
|
95
|
+
IngressMessage)`, preserving all guards (group-only, `ambient.enabled`,
|
|
96
|
+
non-DM). Pure refactor — behaviour identical.
|
|
97
|
+
- Read first: `src/core/runtime.ts:560-630`, `src/core/handler.ts:200-250`
|
|
98
|
+
- Verify: `bun run check` green; existing tests unchanged.
|
|
99
|
+
2. Bounded ambient query in `db.ts` + wire the context builder to use it, so
|
|
100
|
+
ambient shares nothing with the `turnCount * 5` window.
|
|
101
|
+
- Read first: `src/storage/db.ts:1040-1080` and the context assembly that
|
|
102
|
+
consumes `getRecentTurns`
|
|
103
|
+
- Verify: unit test — ambient context returns ≤ N rows and user/assistant
|
|
104
|
+
turns are not displaced. **Blocker: step 3 must not ship without this.**
|
|
105
|
+
3. Handler fix — replace the bare return at `handler.ts:240`:
|
|
106
|
+
```ts
|
|
107
|
+
if (!shouldProcess && !isCommand) {
|
|
108
|
+
core.recordAmbient(ingress);
|
|
109
|
+
return;
|
|
110
|
+
}
|
|
111
|
+
```
|
|
112
|
+
- Read first: `src/core/handler.ts:230-245`
|
|
113
|
+
- Verify: unit tests — writes for a group with `ambient.enabled`
|
|
114
|
+
unset/true; nothing for a DM; nothing when `ambient.enabled=false`;
|
|
115
|
+
called exactly once for a non-triggered group message and **not** for a
|
|
116
|
+
triggered one (no double-write).
|
|
117
|
+
4. Retention TTL in `storage-cleanup.ts` for `role='ambient'` rows.
|
|
118
|
+
- Read first: `src/core/storage-cleanup.ts` (existing cleanup patterns)
|
|
119
|
+
- Verify: unit test — rows older than the TTL are removed, newer kept,
|
|
120
|
+
other roles untouched.
|
|
121
|
+
|
|
122
|
+
---
|
|
123
|
+
|
|
124
|
+
## Non-Negotiable Rules
|
|
125
|
+
|
|
126
|
+
- The handler fix (step 3) never ships without the context cap (step 2) and
|
|
127
|
+
TTL (step 4) — D-003.
|
|
128
|
+
- All existing guards survive the extraction: no ambient writes for DMs, none
|
|
129
|
+
when `ambient.enabled=false`.
|
|
130
|
+
- No double-write for triggered messages.
|
|
131
|
+
|
|
132
|
+
---
|
|
133
|
+
|
|
134
|
+
## Edge Cases & Risks
|
|
135
|
+
|
|
136
|
+
| Scenario | Handling |
|
|
137
|
+
|----------|---------|
|
|
138
|
+
| High-volume group floods ambient | Cap bounds the prompt; TTL bounds the DB; both configurable |
|
|
139
|
+
| `/`-prefixed unknown command in a group | Still falls through to ambient (coordinated with M3.2's routing change) |
|
|
140
|
+
| Triggered message | Processed normally, not also recorded ambient (regression test) |
|
|
141
|
+
| Ambient rows crowd out real turns | Dedicated query — structurally impossible to share the turn budget |
|
|
142
|
+
| Privacy: bot now stores all group chatter | TTL is the mitigation; called out in the PR (roadmap Completion §4) |
|
|
143
|
+
|
|
144
|
+
---
|
|
145
|
+
|
|
146
|
+
## Implementation Checklist
|
|
147
|
+
|
|
148
|
+
### Phase 1 — Goal
|
|
149
|
+
- [x] Goal, User Stories, and MVP Scope written and reviewed
|
|
150
|
+
|
|
151
|
+
### Phase 2 — Architecture
|
|
152
|
+
- [x] Architecture & Data section complete
|
|
153
|
+
- [x] Non-Negotiable Rules defined
|
|
154
|
+
- [x] Edge Cases covered
|
|
155
|
+
- [x] Context for Claude pointers filled in
|
|
156
|
+
|
|
157
|
+
### Implementation (commit `4bf5505`)
|
|
158
|
+
- [x] `recordAmbient` extracted (step 1)
|
|
159
|
+
- [x] Bounded ambient query (step 2)
|
|
160
|
+
- [x] Handler fix (step 3)
|
|
161
|
+
- [x] Retention TTL (step 4)
|
|
162
|
+
- [x] 17 unit tests green; `bun run check` passes (1367 pass / 0 fail)
|
|
163
|
+
- [x] Deployed and running
|
|
164
|
+
- [ ] **Manual (needs a phone):** 3 untagged messages in a test group → 3 new
|
|
165
|
+
`role='ambient'` rows; tag the bot → it can recount them
|
|
166
|
+
- [ ] **Verify with data** once the group sees traffic:
|
|
167
|
+
`SELECT space_id, role, COUNT(*) FROM messages GROUP BY space_id, role`
|
|
168
|
+
— ambient for `greece2026` must go above 0 and keep growing
|
|
169
|
+
- [x] No secrets or `.env` files committed
|
|
170
|
+
- [x] No unrelated files modified
|
|
171
|
+
|
|
172
|
+
---
|
|
173
|
+
|
|
174
|
+
## Open Questions
|
|
175
|
+
|
|
176
|
+
- [x] Defaults shipped as proposed: `ambientContextRows: 30`,
|
|
177
|
+
`ambientTtlDays: 14`. Both are `mercury.yaml`/env-overridable
|
|
178
|
+
(`MERCURY_AMBIENT_CONTEXT_ROWS`, `MERCURY_AMBIENT_TTL_DAYS`) so they can be
|
|
179
|
+
tuned once real volume is observed, without a code change.
|
|
180
|
+
|
|
181
|
+
---
|
|
182
|
+
|
|
183
|
+
## Retrospective
|
|
184
|
+
|
|
185
|
+
**Residual risks / follow-ups:**
|
|
186
|
+
- Ambient volume in a busy family group is still unmeasured. The caps make it
|
|
187
|
+
safe, but the *right* numbers need a week of real data — revisit at the M2
|
|
188
|
+
checkpoint.
|
|
189
|
+
- Media attached to untriggered messages is still downloaded and discarded
|
|
190
|
+
(`bridge.normalize` runs before the gate). Pre-existing, not worsened by
|
|
191
|
+
this change — accepted, as the spec states.
|
|
192
|
+
|
|
193
|
+
**What changed from the plan:**
|
|
194
|
+
- Config went into `AppConfig` (`ambientContextRows`, `ambientTtlDays`) rather
|
|
195
|
+
than the per-space config keys the spec sketched, to match how
|
|
196
|
+
`inboxTtlDays`/`outboxTtlDays` already feed `storage-cleanup`. `ambient.enabled`
|
|
197
|
+
stays per-space as it was.
|
|
198
|
+
- Ambient cleanup runs **before** the spaces-directory scan. That scan
|
|
199
|
+
early-returns when the directory is unreadable, which would have silently
|
|
200
|
+
disabled retention on exactly the hosts most likely to need it.
|
|
201
|
+
|
|
202
|
+
**Key decisions made during implementation:**
|
|
203
|
+
- `getRecentTurns` now excludes ambient outright rather than merely fetching
|
|
204
|
+
more rows. Padding the LIMIT would have made starvation less likely but
|
|
205
|
+
still possible; two queries make it structurally impossible.
|
|
206
|
+
- The handler and `handleRawInput` paths are mutually exclusive by
|
|
207
|
+
construction (one returns before the other can run), which is what makes
|
|
208
|
+
"no double-write" a property rather than a check. A test pins it.
|
|
209
|
+
|
|
210
|
+
**Architecture impact** (update `docs/ARCHITECTURE.md` if any of these apply):
|
|
211
|
+
- [ ] New inter-package interaction (handler → runtime.recordAmbient)
|
|
@@ -0,0 +1,210 @@
|
|
|
1
|
+
# Command Routing Consistency
|
|
2
|
+
|
|
3
|
+
**Status**: Done
|
|
4
|
+
**Slug**: command-routing-consistency
|
|
5
|
+
**Goal**: whatsapp-bot-hardening
|
|
6
|
+
**Milestone**: 3
|
|
7
|
+
**Created**: 2026-08-09
|
|
8
|
+
**Last updated**: 2026-08-09
|
|
9
|
+
|
|
10
|
+
---
|
|
11
|
+
|
|
12
|
+
## Goal
|
|
13
|
+
|
|
14
|
+
> Part of goal [whatsapp-bot-hardening](../goals/whatsapp-bot-hardening/goal.md).
|
|
15
|
+
> HANDOFF Bugs 6b, 6d, 6e + redundancy cleanup — the rest of the command-surface
|
|
16
|
+
> audit after the destructive trio (M3.1).
|
|
17
|
+
|
|
18
|
+
Three consistency defects remain: (1) slash commands in groups are silently
|
|
19
|
+
ignored unless the trigger word is present — `routeInput` re-applies the
|
|
20
|
+
trigger gate that `handler.ts`'s `isCommand` bypass was meant to skip, so the
|
|
21
|
+
bypass accomplishes nothing except filing commands as ambient chat (this is
|
|
22
|
+
exactly how the two stranded baduk commands ended up in the DB); (2) chat and
|
|
23
|
+
CLI permission gates disagree — a promoted space admin denied `/spaces delete`
|
|
24
|
+
in chat can still destroy the current space via `mrctl`; (3) small argument/help
|
|
25
|
+
defects: `/resume <arg>` silently ignores its argument, `/help pause` returns
|
|
26
|
+
"No help available". Plus the audited redundancy cleanup (D-009).
|
|
27
|
+
|
|
28
|
+
## User Stories
|
|
29
|
+
|
|
30
|
+
- As a group member, I want a bare `/spaces list` to work without tagging the
|
|
31
|
+
bot, because that's what the `/` prefix visibly promises.
|
|
32
|
+
- As the operator, I want chat and CLI to enforce the *same* permission answer
|
|
33
|
+
for the same action, so the tighter gate can't be bypassed.
|
|
34
|
+
- As any user, I want `/resume 5m` to either work as written or error — not
|
|
35
|
+
silently do something else — and `/help pause` to actually document `/pause`.
|
|
36
|
+
|
|
37
|
+
## MVP Scope
|
|
38
|
+
|
|
39
|
+
**In scope:** trigger-free slash-command routing in groups (known commands
|
|
40
|
+
only; unknown `/foo` still falls through to ambient — D-007); lift
|
|
41
|
+
`spaces.delete` out of the blanket space-admin default grant so CLI matches
|
|
42
|
+
chat's seeded-admin gate (D-008); `/resume` rejects unknown args; `verbs`
|
|
43
|
+
metadata for `pause` incl. documenting the `^(\d+)(m|h)$` duration format; drop
|
|
44
|
+
`/model active`; rewrite `/compact` vs `/clear` help to state the
|
|
45
|
+
`session_boundary` vs `clear_boundary` difference (D-009).
|
|
46
|
+
|
|
47
|
+
**Out of scope:** merging `/compact` and `/clear` — deliberately not done blind
|
|
48
|
+
(D-009); any change to member default permissions (already minimal and correct).
|
|
49
|
+
|
|
50
|
+
---
|
|
51
|
+
|
|
52
|
+
## Context for Claude
|
|
53
|
+
|
|
54
|
+
- `src/core/router.ts:105-115` — `matchTrigger` returns `ignore` before the
|
|
55
|
+
slash parse; the ordering to fix.
|
|
56
|
+
- `src/core/handler.ts:236-240` — the `isCommand` bypass that currently
|
|
57
|
+
accomplishes nothing (coordinate with M2.2's `recordAmbient` call there).
|
|
58
|
+
- `src/core/router.ts:33` (`SEEDED_ADMIN_COMMANDS`) and
|
|
59
|
+
`src/core/permissions.ts:208-210` (blanket admin grant), `:192-197`
|
|
60
|
+
(member defaults — untouched).
|
|
61
|
+
- `src/core/commands.ts` — `SLASH_COMMANDS` (resume verbs, pause `verbs`
|
|
62
|
+
array, `/model active` removal, compact/clear descriptions);
|
|
63
|
+
`parsePauseDuration`.
|
|
64
|
+
- `HANDOFF.md` Bug 6 — full audit incl. the redundancy verdicts table.
|
|
65
|
+
|
|
66
|
+
---
|
|
67
|
+
|
|
68
|
+
## Architecture & Data
|
|
69
|
+
|
|
70
|
+
### Data Models
|
|
71
|
+
|
|
72
|
+
None.
|
|
73
|
+
|
|
74
|
+
### API Contracts
|
|
75
|
+
|
|
76
|
+
Permission semantics change: `spaces.delete` is no longer granted by default
|
|
77
|
+
space-admin role — only the seeded admin (config) passes `checkPerm` for it.
|
|
78
|
+
This is a **tightening**; document it in the PR behaviour-change list.
|
|
79
|
+
|
|
80
|
+
### File & Folder Structure
|
|
81
|
+
|
|
82
|
+
| Path | New / Modified | Purpose |
|
|
83
|
+
|------|---------------|---------|
|
|
84
|
+
| `src/core/router.ts` | Modified | Parse known slash commands before/despite the trigger gate in groups |
|
|
85
|
+
| `src/core/permissions.ts` | Modified | Lift `spaces.delete` from the blanket admin grant |
|
|
86
|
+
| `src/core/commands.ts` | Modified | `/resume` arg rejection; `pause` verbs + duration docs; drop `/model active`; honest compact/clear help |
|
|
87
|
+
| test files alongside each | New/Modified | See tests below |
|
|
88
|
+
|
|
89
|
+
### Implementation Sequence
|
|
90
|
+
|
|
91
|
+
1. **Trigger-free slash routing in groups** — in `routeInput`, recognize a
|
|
92
|
+
known slash command before the trigger `ignore`; unknown `/foo` keeps
|
|
93
|
+
falling through (to ambient once M2.2 lands). Coordinate with the
|
|
94
|
+
`handler.ts` bypass so there's exactly one gate.
|
|
95
|
+
- Read first: `src/core/router.ts:100-120`, `src/core/handler.ts:230-245`,
|
|
96
|
+
M2.2's spec (shared code path)
|
|
97
|
+
- Verify: unit — bare `/spaces list` in a group resolves as a command;
|
|
98
|
+
unknown `/foo` in a group does **not** become a command.
|
|
99
|
+
2. **Permission alignment** — remove `spaces.delete` from
|
|
100
|
+
`getDefaultPermissions`' admin set; confirm the seeded-admin path still
|
|
101
|
+
passes everywhere it should.
|
|
102
|
+
- Read first: `src/core/permissions.ts:185-215`, `src/core/router.ts:30-35`
|
|
103
|
+
- Verify: unit — a promoted (non-seeded) space admin gets the same denial
|
|
104
|
+
from the chat path and the CLI/API path.
|
|
105
|
+
3. **`/resume` argument handling** — pass `verb` through; unknown/unsupported
|
|
106
|
+
args error with usage instead of silently acting as bare `/resume`.
|
|
107
|
+
- Read first: `src/core/commands.ts` (`executeResumeCommand` and its caller)
|
|
108
|
+
- Verify: unit — `/resume extra` errors; bare `/resume` unchanged.
|
|
109
|
+
4. **`/help pause`** — add the `verbs` array so `formatCategoryHelp` works;
|
|
110
|
+
document the accepted duration format (`\d+m|\d+h`).
|
|
111
|
+
- Read first: `src/core/commands.ts:55-67`, `formatCategoryHelp`,
|
|
112
|
+
`parsePauseDuration`
|
|
113
|
+
- Verify: unit — `/help pause` returns real help text.
|
|
114
|
+
5. **Redundancy cleanup** — remove `/model active` (the fact lives in
|
|
115
|
+
`/model list`'s `← active` marker); rewrite `/compact` and `/clear`
|
|
116
|
+
descriptions to state the actual boundary difference (`db.ts:913` vs `:936`).
|
|
117
|
+
- Read first: `src/core/commands.ts` model/compact/clear entries,
|
|
118
|
+
`src/storage/db.ts:905-940`
|
|
119
|
+
- Verify: unit/help-snapshot — `/model active` gone from help and routing;
|
|
120
|
+
new descriptions render.
|
|
121
|
+
|
|
122
|
+
Step 1 should land after M3.1's step 4 (same router file, `/`-handling).
|
|
123
|
+
|
|
124
|
+
---
|
|
125
|
+
|
|
126
|
+
## Non-Negotiable Rules
|
|
127
|
+
|
|
128
|
+
- Unknown `/foo` in a group must never execute as a command — it falls through
|
|
129
|
+
(ambient), preserving the safe default.
|
|
130
|
+
- Member default permissions are not touched.
|
|
131
|
+
- Every user-visible behaviour change here goes in the PR's explicit
|
|
132
|
+
behaviour-change list (roadmap Completion §6).
|
|
133
|
+
|
|
134
|
+
---
|
|
135
|
+
|
|
136
|
+
## Edge Cases & Risks
|
|
137
|
+
|
|
138
|
+
| Scenario | Handling |
|
|
139
|
+
|----------|---------|
|
|
140
|
+
| Message that merely *starts* with `/` but isn't a command (`/shrug`) | Not in `SLASH_COMMANDS` → falls through to ambient, never executes |
|
|
141
|
+
| DM slash commands | Unchanged — DMs never required the trigger |
|
|
142
|
+
| Seeded admin also holds space-admin role | Passes both gates before and after — no regression |
|
|
143
|
+
| Muscle memory: users who typed `Wiz /spaces list` | Still works — trigger + command remains valid |
|
|
144
|
+
| `/model active` users | Removal listed in PR behaviour changes; `/model list` shows the same fact |
|
|
145
|
+
|
|
146
|
+
---
|
|
147
|
+
|
|
148
|
+
## Implementation Checklist
|
|
149
|
+
|
|
150
|
+
### Phase 1 — Goal
|
|
151
|
+
- [x] Goal, User Stories, and MVP Scope written and reviewed
|
|
152
|
+
|
|
153
|
+
### Phase 2 — Architecture
|
|
154
|
+
- [x] Architecture & Data section complete
|
|
155
|
+
- [x] Non-Negotiable Rules defined
|
|
156
|
+
- [x] Edge Cases covered
|
|
157
|
+
- [x] Context for Claude pointers filled in
|
|
158
|
+
|
|
159
|
+
### Implementation (commit `4f05fbb`)
|
|
160
|
+
- [x] Trigger-free slash routing (step 1)
|
|
161
|
+
- [x] Permission alignment (step 2) — via a route-level check, see retrospective
|
|
162
|
+
- [x] `/resume` args (step 3)
|
|
163
|
+
- [x] `/help pause` (step 4)
|
|
164
|
+
- [x] Redundancy cleanup (step 5)
|
|
165
|
+
- [x] 11 unit tests green; `bun run check` passes (1404 pass / 0 fail)
|
|
166
|
+
- [ ] **Manual (needs a phone):** bare `/spaces list` untagged in a test group
|
|
167
|
+
- [x] No secrets or `.env` files committed
|
|
168
|
+
- [x] No unrelated files modified
|
|
169
|
+
|
|
170
|
+
---
|
|
171
|
+
|
|
172
|
+
## Open Questions
|
|
173
|
+
|
|
174
|
+
- [x] none
|
|
175
|
+
|
|
176
|
+
---
|
|
177
|
+
|
|
178
|
+
## Retrospective
|
|
179
|
+
|
|
180
|
+
**Residual risks / follow-ups:**
|
|
181
|
+
- With no seeded admins configured, `spaces.delete` keeps its old behaviour
|
|
182
|
+
(any space admin passes). Deliberate — the alternative is a deployment where
|
|
183
|
+
nobody can delete a space — but it means the alignment only bites where
|
|
184
|
+
`permissions.admins` is set. Accepted.
|
|
185
|
+
|
|
186
|
+
**What changed from the plan:**
|
|
187
|
+
- **D-008 was not implementable as written.** The plan said to lift
|
|
188
|
+
`spaces.delete` out of the blanket admin grant. But seeded admins are
|
|
189
|
+
*seeded into the database as role `admin`* (`resolveRole` →
|
|
190
|
+
`db.seedAdmins`), so seeded and promoted admins are indistinguishable by
|
|
191
|
+
role — removing the permission from `admin` would have locked out everyone,
|
|
192
|
+
including the person the chat gate is designed to admit. The alignment
|
|
193
|
+
instead lives in the delete route, which now asks the same seeded-admin
|
|
194
|
+
question the chat path asks. Same outcome, and it does not depend on a role
|
|
195
|
+
distinction that does not exist.
|
|
196
|
+
|
|
197
|
+
**Key decisions made during implementation:**
|
|
198
|
+
- Slash-command admission is by *known command name*, not by the `/` prefix:
|
|
199
|
+
`namesKnownCommand()` checks `SLASH_COMMANDS` and `CHAT_COMMANDS`, so
|
|
200
|
+
`/shrug` still falls through to ambient. This keeps the safe default while
|
|
201
|
+
making the handler's long-dead `isCommand` bypass finally mean something.
|
|
202
|
+
- `/model active` kept as a redirect rather than deleted outright, so muscle
|
|
203
|
+
memory lands on an explanation instead of "unknown verb".
|
|
204
|
+
- `/compact` and `/clear` were not merged. Their real difference was verified
|
|
205
|
+
in the code first (`min_message_id` is permanent, `clear_boundary` is reset
|
|
206
|
+
after each read) and then written into the help text, which is what the
|
|
207
|
+
audit actually asked for.
|
|
208
|
+
|
|
209
|
+
**Architecture impact** (update `docs/ARCHITECTURE.md` if any of these apply):
|
|
210
|
+
- [ ] none expected
|