mercury-agent 0.16.0 → 0.16.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (80) hide show
  1. package/README.md +6 -0
  2. package/container/Dockerfile +6 -1
  3. package/container/Dockerfile.base +6 -1
  4. package/docs/README.md +50 -0
  5. package/docs/authoring-profiles.md +8 -0
  6. package/docs/behavior-layers.md +193 -0
  7. package/docs/extensions.md +2 -0
  8. package/docs/goals/football-reporter-profile/decisions.md +479 -0
  9. package/docs/goals/football-reporter-profile/goal.md +91 -0
  10. package/docs/goals/football-reporter-profile/roadmap.md +196 -0
  11. package/docs/goals/whatsapp-bot-hardening/README.md +43 -0
  12. package/docs/goals/whatsapp-bot-hardening/archive/ambient-group-context.md +211 -0
  13. package/docs/goals/whatsapp-bot-hardening/archive/command-routing-consistency.md +210 -0
  14. package/docs/goals/whatsapp-bot-hardening/archive/destructive-command-guards.md +218 -0
  15. package/docs/goals/whatsapp-bot-hardening/archive/docs-extension.md +182 -0
  16. package/docs/goals/whatsapp-bot-hardening/archive/handoff-2026-08-09.md +588 -0
  17. package/docs/goals/whatsapp-bot-hardening/archive/media-size-and-silent-drop.md +178 -0
  18. package/docs/goals/whatsapp-bot-hardening/archive/member-memory.md +199 -0
  19. package/docs/goals/whatsapp-bot-hardening/archive/message-author-attribution.md +211 -0
  20. package/docs/goals/whatsapp-bot-hardening/archive/run-from-source-switchover.md +212 -0
  21. package/docs/goals/whatsapp-bot-hardening/archive/setup-plan-2026-08-06.md +1117 -0
  22. package/docs/goals/whatsapp-bot-hardening/decisions.md +200 -0
  23. package/docs/goals/whatsapp-bot-hardening/goal.md +75 -0
  24. package/docs/goals/whatsapp-bot-hardening/open-threads.md +396 -0
  25. package/docs/goals/whatsapp-bot-hardening/roadmap.md +245 -0
  26. package/docs/memory.md +30 -0
  27. package/docs/profile-guide.md +483 -0
  28. package/docs/refactor/.gitkeep +0 -0
  29. package/docs/refactor/archive/.gitkeep +0 -0
  30. package/docs/refactor/audits/.gitkeep +0 -0
  31. package/docs/refactor/backlog/.gitkeep +0 -0
  32. package/docs/skills-guide.md +154 -0
  33. package/examples/extensions/README.md +1 -0
  34. package/examples/extensions/feed-watch/config.ts +354 -0
  35. package/examples/extensions/feed-watch/digest.ts +344 -0
  36. package/examples/extensions/feed-watch/feeds.ts +341 -0
  37. package/examples/extensions/feed-watch/index.ts +302 -0
  38. package/examples/extensions/feed-watch/items.ts +171 -0
  39. package/examples/extensions/feed-watch/match.ts +107 -0
  40. package/examples/extensions/feed-watch/prompts/verify.md +35 -0
  41. package/examples/extensions/feed-watch/skill/SKILL.md +128 -0
  42. package/examples/extensions/feed-watch/watch.ts +548 -0
  43. package/examples/extensions/longview/hook.ts +179 -16
  44. package/examples/extensions/longview/index.ts +13 -2
  45. package/examples/extensions/longview/prompts/summarize.md +3 -2
  46. package/examples/extensions/longview/render/telegraph-nodes.ts +21 -1
  47. package/examples/extensions/longview/summarize.ts +57 -14
  48. package/examples/extensions/napkin/index.ts +334 -86
  49. package/examples/extensions/napkin/pi-spawn.ts +196 -0
  50. package/examples/extensions/pinchtab/index.ts +37 -4
  51. package/examples/extensions/pinchtab/lib/session-injector.ts +31 -9
  52. package/examples/extensions/pinchtab/skill/SKILL.md +29 -0
  53. package/examples/profiles/_template/AGENTS.md +138 -0
  54. package/examples/profiles/_template/README.md +65 -0
  55. package/examples/profiles/_template/config.yaml +57 -0
  56. package/examples/profiles/_template/tasks/daily.md +32 -0
  57. package/examples/profiles/football-reporter/AGENTS.md +219 -0
  58. package/examples/profiles/football-reporter/README.md +138 -0
  59. package/examples/profiles/football-reporter/config.yaml +174 -0
  60. package/examples/profiles/football-reporter/seed/MEMORY.md +42 -0
  61. package/examples/profiles/football-reporter/seed/episodes/barcelona-2026-27.md +25 -0
  62. package/examples/profiles/football-reporter/seed/episodes/beitar-jerusalem-2026-27.md +25 -0
  63. package/examples/profiles/football-reporter/seed/episodes/maccabi-haifa-2026-27.md +24 -0
  64. package/examples/profiles/football-reporter/seed/episodes/man-united-2026-27.md +26 -0
  65. package/examples/profiles/football-reporter/seed/episodes/real-madrid-2026-27.md +25 -0
  66. package/examples/profiles/football-reporter/seed/napkin-distill.md +55 -0
  67. package/examples/profiles/football-reporter/tasks/daily-article.md +142 -0
  68. package/package.json +8 -5
  69. package/src/adapters/whatsapp.ts +75 -13
  70. package/src/bridges/whatsapp.ts +6 -4
  71. package/src/cli/mrctl.ts +9 -1
  72. package/src/core/handler.ts +51 -6
  73. package/src/core/routes/dashboard.ts +106 -5
  74. package/src/core/routes/tasks.ts +28 -0
  75. package/src/core/runtime.ts +63 -2
  76. package/src/core/task-scheduler.ts +103 -12
  77. package/src/extensions/catalog.ts +9 -0
  78. package/src/profile/space-profile.ts +780 -0
  79. package/src/storage/db.ts +215 -0
  80. package/src/types.ts +56 -0
@@ -0,0 +1,479 @@
1
+ # Decisions: Football Reporter Profile
2
+
3
+ **Goal**: [football-reporter-profile](goal.md)
4
+
5
+ > Inherits every decision in
6
+ > [`whatsapp-bot-hardening/decisions.md`](../whatsapp-bot-hardening/decisions.md),
7
+ > in particular **`whatsapp-bot-hardening` D-014** (one instruction, one layer;
8
+ > countable rules). Bare `D-0NN` ids below always mean *this* goal's sequence.
9
+ > No conflict with that goal's tech choices — this goal adds no new storage,
10
+ > framework or runtime.
11
+
12
+ ## D-001: The football profile is a *space-scoped* folder in the mercury repo
13
+ - **Category:** architecture
14
+ - **Decided:** `examples/profiles/football-reporter/` in this repo, applied to
15
+ one space by `scripts/space-profile.ts apply|check|dump`. Not a
16
+ `resources/profiles/` built-in, not `mercury profiles apply`, not a profile
17
+ in `tagula-agent-profiles`.
18
+ - **Alternatives considered:** (a) a built-in profile under
19
+ `resources/profiles/` applied with `mercury profiles apply` — rejected
20
+ because `applyProfile` (`src/core/profiles.ts`) writes
21
+ `.mercury/global/AGENTS.md` and pins `member_permissions` and
22
+ `profile_prompt` **project-wide**, and `~/whatsapp-bot` hosts several
23
+ spaces; (b) `profiles/football-reporter/` in `tagula-agent-profiles`,
24
+ deployed by its `deploy.sh` — rejected because the football bot is the
25
+ personal `~/whatsapp-bot`, feed-watch is already a mercury backlog item, and
26
+ tagula's loader is `deploy.sh` copying files by hand, which
27
+ `~/sync-mercury.sh` can do equally well; (c) extend mercury with
28
+ `mercury profiles apply --space <id>` — real upstream value but scope creep
29
+ for this goal.
30
+ - **Reasoning:** the value is one owner for the rules plus a drift check, not
31
+ the profile machinery. A folder + script gives both today with zero core
32
+ change. `examples/` mirrors the bundled-but-optional convention of
33
+ `examples/extensions/`.
34
+ - **Revisit if:** upstream adds per-space profile application — then the
35
+ folder becomes a manifest and the script goes away.
36
+ - **Used by:** football-reporter-space-profile, daily-article-redesign
37
+
38
+ ## D-002: Profile contents — the shape of a tagula profile, none of the plumbing
39
+ - **Category:** scope
40
+ - **Decided:** the folder holds `README.md` (the layer map), `AGENTS.md`
41
+ (the editorial standard, Hebrew), `tasks/*.md` (task prompt sources),
42
+ `config.yaml` (per-space extension config: feed-watch, longview) and
43
+ nothing else. Skipped: `member_permissions`, `capabilities`, engines, config
44
+ registries, Drive, trial gates, `env`.
45
+ - **Alternatives considered:** full `mercury-profile.yaml` manifest for
46
+ parity with tagula; rejected — none of its fields apply to a single personal
47
+ space.
48
+ - **Reasoning:** "without the commercial bullshit" (owner). Every file in the
49
+ folder must be something the bot's behaviour observably depends on.
50
+ - **Revisit if:** a second space wants the same profile — then a manifest with
51
+ `{SPACE}` substitution earns its place.
52
+ - **Used by:** football-reporter-space-profile
53
+
54
+ ## D-003: Layer placement for football (applies `whatsapp-bot-hardening` D-014)
55
+ - **Category:** pattern
56
+ - **Decided:** editorial standard (role, sourcing, freshness, format, silence,
57
+ honesty about writes) → `spaces/football-friends/AGENTS.md`; *procedure*
58
+ (when to run, what artefact to produce, where to publish, how to dedupe) →
59
+ the task prompt; **zero editorial** preferences — `club-coverage` moves into
60
+ `AGENTS.md` and is deleted, `push-rules` is already gone; the three
61
+ *behaviour* preferences the live space carries (`language`, `timezone`,
62
+ `tone` — read 2026-08-20) are not editorial and stay; global character and
63
+ the baked prompt untouched. Every quantity is countable ("0 emoji",
64
+ "1–4 lines", "1 named source", "date within 24 h").
65
+ - **Alternatives considered:** keep `club-coverage` as the member-visible
66
+ knob; rejected — it is the layer that produced the F1 contradictions and
67
+ its 500-char cap is what pushed rules into task prompts.
68
+ - **Reasoning:** audit F1: "a preference may reinforce a task prompt, but must
69
+ never state a different threshold than one" — the safest way to satisfy that
70
+ is to have none. `AGENTS.md` is read per run, no rebuild.
71
+ - **Revisit if:** a member-editable knob is wanted — then one preference with
72
+ a single countable value, never a rule that exists elsewhere.
73
+ - **Used by:** football-reporter-space-profile, feed-watch
74
+
75
+ ## D-004: Freshness is enforced by a date, not an adjective
76
+ - **Category:** pattern
77
+ - **Decided:** every reported item carries the publication date the agent
78
+ read from the source; an item without a date or outside the window (scan:
79
+ since the last update; article: 24 h) is dropped; a transfer is reported as
80
+ live only after checking the player's current club. Feed-watch supplies
81
+ `publishedAt` per digest item; the harness asserts a date token per item.
82
+ - **Alternatives considered:** "only current news" wording; rejected — audit
83
+ F4 showed a months-old real article satisfies every attribution rule as
84
+ written.
85
+ - **Reasoning:** borrowed from tagula's date-grounding rule ("without
86
+ reading the engine you have no clock"): a date is checkable, "is this
87
+ current?" is not.
88
+ - **Revisit if:** sources stop carrying dates — then the host timestamps
89
+ first-seen and that becomes the date.
90
+ - **Used by:** football-reporter-space-profile, feed-watch,
91
+ football-reporter-fixture-harness
92
+
93
+ ## D-005: Keep Telegraph, make the summary fuller, and redesign the article itself (decided 2026-08-20)
94
+ - **Category:** scope
95
+ - **Decided:** the 09:00 article stays a Telegraph page, but (a) the chat
96
+ message becomes a fuller summary — headline plus the key facts with sources,
97
+ cap 900 chars for this space — and (b) the page is redesigned to read and
98
+ look like a top-tier sports article: headline, standfirst, dateline,
99
+ per-club prose sections, at most one real pull-quote, a "certain / in
100
+ doubt" box, a sources footer; accurate, concise, intriguing through the
101
+ facts, designed for readability, and with **no clickbait or commercial
102
+ dressing** (a forbidden-word list in the prompt, asserted by the harness).
103
+ - **Alternatives considered:** move the article in-chat and disable
104
+ `longview` for the space (the 2026-08-20 proposal); status quo.
105
+ - **Reasoning:** owner's call. The group's complaint (audit F3) was about
106
+ teasing and quality, which the fuller summary and the redesign address
107
+ directly; a page still fits an article's length better than a chat
108
+ bubble, and the in-chat full answer (#1905) shape is kept as the *summary*.
109
+ - **Revisit if:** two weeks of reactions still read the link as a tease —
110
+ then in-chat (the proposal) is the next step, one config value away.
111
+ - **Used by:** daily-article-redesign
112
+
113
+ ## D-010: The chat summary is author-supplied by the same run, not a second model call
114
+ - **Category:** architecture
115
+ - **Decided:** `longview` honours a leading `[longview:summary] … [/longview:summary]`
116
+ block in the reply: used verbatim as the chat text (≤ `summary_max_chars`,
117
+ new per-space key, default 300 = today's `BLURB_MAX_CHARS`), stripped from
118
+ the published page, `summarize()` skipped. No block → today's path.
119
+ - **Alternatives considered:** raise the blurb cap and reword
120
+ `prompts/summarize.md` only (a second prompt to drift from `AGENTS.md`,
121
+ a second model call per article, and a summary that has not read the
122
+ sources); a host-side extractive summary (loses the editorial split).
123
+ - **Reasoning:** one owner for the shape (D-003) — the run that verified the
124
+ sources writes both the page and the summary under the same `AGENTS.md`;
125
+ deterministic, cheaper, and behaviour-preserving for every other space.
126
+ - **Revisit if:** the model misses the block often enough that the fallback
127
+ blurb is what the group usually sees — the harness reports the rate.
128
+ - **Used by:** daily-article-redesign
129
+
130
+ ## D-011: Source policy — verified feeds, journalist-over-masthead tiers, a five-step grade ladder
131
+ - **Category:** scope
132
+ - **Decided:** (1) feed-watch's initial `sources` for the space are the feeds
133
+ verified on 2026-08-20 in `docs/notes/football-sources-and-style-research.md`
134
+ §1 (BBC football, Sky 11095, ESPN soccer, Marca per club, Daily Briefing,
135
+ ONE, Walla 156/316, ynet sport) plus one Google News RSS query per Hebrew
136
+ term and one per English term with `when:1d`; feeds that were blocked from
137
+ the planning machine (MEN, manutd.com, Guardian, Athletic, Transfermarkt)
138
+ join only after a fetch test from WSL. (2) `AGENTS.md` carries the source
139
+ tiers: official channel > named Tier-1 journalist > serious outlet; the
140
+ Tier-4/5 list (Sun, Mirror, Star, Express, Mail, 90min, Caught Offside,
141
+ Football Insider, TeamTalk, Sport Witness, Don Balón, Fichajes, El
142
+ Chiringuito, Bleacher Report, aggregators) is never cited; the journalist
143
+ outranks the masthead. (3) The certainty vocabulary becomes the five-step
144
+ ladder רשמי / סגור (here we go) / הסכמה / במשא ומתן / שמועה, each tied to
145
+ the wording the source used; hedged filler ("לפי דיווחים", "ייתכן ש") is
146
+ banned — *either we know it and report it, or we don't report it.*
147
+ - **Alternatives considered:** keep the live prompt's free-text source list
148
+ and four grades; Brave search as the discovery layer (today's cost and
149
+ staleness problem).
150
+ - **Reasoning:** the research note: three independent tier guides agree on
151
+ the shape; Romano's own definition of "here we go" and the BBC's
152
+ "understands" ruling give the grade ladder and the no-hedging rule an
153
+ external anchor; every ✅ feed was fetched and shown to carry `pubDate`,
154
+ which D-004 depends on.
155
+ - **Revisit if:** a ✅ feed goes stale or changes shape (feed-watch's
156
+ `source_status` backoff will show it), or a new Tier-1 journalist emerges
157
+ for one of the seven scopes — the list is config, not code.
158
+ - **Used by:** football-reporter-space-profile, feed-watch,
159
+ daily-article-redesign, football-reporter-fixture-harness
160
+
161
+ ## D-006: Scheduled research runs post from the primary model leg or not at all
162
+ - **Category:** trade-off
163
+ - **Decided:** a scheduled task may opt out of the fallback chain
164
+ (primary-or-skip); a skipped run is recorded on the task row and in the
165
+ ledger, not posted to the group. Marking leg-2 output in the message is the
166
+ fallback design if opt-out proves impossible.
167
+ - **Alternatives considered:** (2) mark fallback output in the message;
168
+ (3) log a warning only (audit F2's ordered options).
169
+ - **Reasoning:** "a missed scan is strictly better than a fabricated one";
170
+ `mercury.yaml` itself documents leg-2 as bad at exactly the verification
171
+ job the scan is.
172
+ - **Revisit if:** the primary leg's availability drops enough that skips
173
+ become the norm — then the chain needs a second research-grade leg.
174
+ - **Used by:** scheduled-task-model-leg-policy
175
+
176
+ ## D-007: Bugs run on a parallel track, never inside a story
177
+ - **Category:** pattern
178
+ - **Decided:** defects found before or during the goal are filed with
179
+ `/f-bug-report` and fixed with `/f-bug-fix`; a story never absorbs a fix.
180
+ Filed at planning: F5 (reply-to-member triggers the bot), F7a (HTML tags not
181
+ stripped), F7d (`_(responded in Ns)_` not stored), F7e (bot self-name is a
182
+ phone number). F7c (emoji) is prompt compliance → M1.1 + harness; F6 →
183
+ M1.1 honesty rule; F4/F8 → M1.1/M2.1.
184
+ - **Alternatives considered:** one "small fixes" story; rejected — it hides
185
+ root causes and breaks the post-mortem trail.
186
+ - **Reasoning:** owner's instruction; matches the refactor invariant in
187
+ `CLAUDE.md` (file, don't fix in place).
188
+ - **Revisit if:** never.
189
+ - **Used by:** football-regression-suite
190
+
191
+ ## D-008: Proof of quality = deterministic checks + fixtures + a ledger; measure before gating
192
+ - **Category:** pattern
193
+ - **Decided:** (1) any rule that can be a host-side gate or a script is one
194
+ (output lint: HTML, emoji, line count, source + date tokens); (2) prompts are
195
+ exercised against dated fixtures in the private `main` space, never the
196
+ group; (3) every scheduled run writes a ledger row (leg, cost, outcome,
197
+ item count); (4) numbers are reported for two weeks before any of them
198
+ becomes a threshold.
199
+ - **Alternatives considered:** rely on the re-audit alone; rejected — it is
200
+ a 128-row read that happens once.
201
+ - **Reasoning:** tagula's "measure min/median/max over 16+ seeds before
202
+ writing a number an agent will branch on" and the `normalizeChatMarkdown`
203
+ precedent ("a prose rule relies on the model complying").
204
+ - **Revisit if:** the ledger shows the lint never fires for a month — then
205
+ the lint can be relaxed to a warning.
206
+ - **Used by:** football-reporter-fixture-harness, scheduled-run-ledger,
207
+ football-regression-suite
208
+
209
+ ## D-009: feed-watch is absorbed as written; only its bundled ops move
210
+ - **Category:** scope
211
+ - **Decided:** `docs/backlog/feed-watch.md` becomes M2.1 unchanged in
212
+ architecture. Its "bundled space work" (the football `AGENTS.md`, the
213
+ workspace cleanup, the task-prompt rewrite) is owned by M1.1 and must be
214
+ live **before** feed-watch is enabled for the space (its own ops order
215
+ already requires this).
216
+ - **Alternatives considered:** re-plan feed-watch inside the profile;
217
+ rejected — the spec is complete and its prerequisite
218
+ ([PR #41](https://github.com/Avishai-Tsabari/mercury/pull/41) `--timezone`)
219
+ is on `dev`.
220
+ - **Reasoning:** don't re-derive a finished plan.
221
+ - **Revisit if:** M1 changes the article's input contract — then the digest
222
+ marker and hook wording follow.
223
+ - **Used by:** feed-watch
224
+
225
+ ## D-012: The ledger is built before feed-watch (decided 2026-08-20)
226
+ - **Category:** sequencing
227
+ - **Decided:** M3.2 `scheduled-run-ledger` is implemented next, ahead of M2.1
228
+ `feed-watch`. Milestone numbering is unchanged — only the build order moves.
229
+ - **Alternatives considered:** keep strict milestone order (M2.1 first);
230
+ rejected because M3.2 has no dependencies (the roadmap's own graph marks it
231
+ independent and says it is "worth starting early"), and because feed-watch is
232
+ the first change whose effect the operator would want to *measure* — the
233
+ digest either reduces the scans that find nothing or it does not.
234
+ - **Reasoning:** an instrument installed after the change it is meant to
235
+ measure cannot show the before. Building the ledger first means feed-watch
236
+ ships with a baseline already recorded, and gives M3.2's `outcome` column a
237
+ natural place for feed-watch's `digest_items` to land additively.
238
+ - **Revisit if:** the ledger turns out to need something feed-watch creates —
239
+ it does not today; `digest_items` is deliberately left out of v1.
240
+ - **Used by:** scheduled-run-ledger, feed-watch, football-regression-suite
241
+
242
+ ## D-013: Tone is editorial, and scan length is per item (decided 2026-08-20)
243
+ - **Category:** layer placement (refines [D-003](#d-003-layer-placement-for-football))
244
+ - **Decided:** two changes to where register lives and how output length is
245
+ measured.
246
+ 1. **`tone` moves from a space preference into `AGENTS.md` (`## סגנון`)**,
247
+ and `config.yaml` deletes the preference via `remove_preferences`. The
248
+ escalation clause the preference carried (intervene only for targeted
249
+ harassment, spam, or an attempt to extract secrets) moves with it.
250
+ 2. **The scan's `1–4 שורות סך הכל` cap is replaced by one line per item**,
251
+ plus a readability test stated in the reader's terms: a message you have
252
+ to scroll on a phone, in a friends' group, is too long — even when every
253
+ item in it is justified. A lineup is the single exception and always
254
+ renders as a formation block, in chat, article and scan alike.
255
+ - **Alternatives considered:**
256
+ - *Strengthen the preference instead of moving it.* Rejected on evidence, not
257
+ taxonomy — see below.
258
+ - *Keep both layers and make them agree.* Rejected: two layers holding one
259
+ rule is what makes precedence a coin flip, which is the whole point of
260
+ D-003. Agreement today is drift tomorrow.
261
+ - *Keep a line cap and raise the number.* Rejected: a global cap is what
262
+ forced eleven names into prose in the first place, and any number is wrong
263
+ for a scan with one item and a scan with six.
264
+ - **Reasoning:** the preference was **already live and already said the right
265
+ thing** — 417 characters of "one of the group, roll with the banter, let
266
+ casual cursing slide" — throughout the 2026-08-18 exchange the group ended by
267
+ calling the bot a party pooper. It lost because the two layers are not
268
+ comparable: `AGENTS.md` is a pi agent-instructions file mounted into the
269
+ workspace, while a preference is a `<preferences>` block at position 5 of the
270
+ user prompt, competing with ~100 lines of reporter discipline in that file.
271
+ Register is what the output *reads like*, which is precisely what `AGENTS.md`
272
+ is declared the sole authority for. On length: counting lines measured the
273
+ wrong thing. What makes a scan unreadable is a paragraph per item, not the
274
+ number of items, so the rule now bounds verbosity per item and lets the
275
+ message grow with genuine news.
276
+ - **Revisit if:** a future runtime makes preference precedence explicit and
277
+ reliable — then a runtime-tunable register knob becomes worth more than a
278
+ version-controlled one. Or if scans start running long in practice, in which
279
+ case the fix is a stricter per-item rule, not a return to a global cap.
280
+ - **Used by:** football-reporter profile (`AGENTS.md`, `config.yaml`),
281
+ feed-watch (its digest inherits the per-item rule)
282
+
283
+ ## D-014: The per-item date rule is enforced by M3.1's lint, not by a prompt fix (decided 2026-08-21)
284
+ - **Category:** proof of quality (refines [D-004](#d-004-freshness-is-enforced-by-a-date-not-an-adjective))
285
+ - **Decided:** the standard's requirement that every item carry a publication
286
+ date stays exactly as written in `AGENTS.md`. Nothing in the prompts is
287
+ changed to chase it. The **M3.1 lint becomes the enforcement**, and it must
288
+ cover **both** surfaces — the scan reply *and* the article's
289
+ `[longview:summary]` block — not the scan alone.
290
+ - **Evidence at decision time:** two independent live observations, same
291
+ `AGENTS.md`, both missing dates.
292
+ 1. Scan, 2026-08-20 16:31 — 0 of 3 items dated; item 1 used היום, the
293
+ adjective D-004 exists to forbid.
294
+ 2. Article summary, 2026-08-21 09:00 (ledger run #2) — 0 date tokens across
295
+ 4 lines, while the article *body* of the same run carried 5.
296
+ - **Alternatives considered:**
297
+ - *File it as a bug on the parallel track (D-007).* Rejected: a bug fix here
298
+ would be a prompt edit, and a prompt edit is precisely the unverifiable
299
+ change this milestone exists to stop making. The defect is real but the
300
+ remedy is a check.
301
+ - *Tighten the task prompts now, as an M1 follow-up.* Rejected for the same
302
+ reason, plus it would re-open a milestone whose stories are done.
303
+ - **Reasoning:** the rule is stated once, in the layer D-003 made its owner, and
304
+ it is still ignored — so the gap is not where the rule lives but that nothing
305
+ measures it. A third prompt rewrite would be the third unmeasured attempt.
306
+ The body-versus-summary split in run #2 is the sharpest evidence yet that this
307
+ is procedure drift rather than a capability limit: the same run, on the same
308
+ model, dated the long form and not the short one.
309
+ - **Revisit if:** the lint lands and the rule still fails on fixtures — that
310
+ would move it from procedure drift to a capability limit and make a prompt or
311
+ model change the honest answer.
312
+ - **Used by:** football-reporter-fixture-harness (M3.1), football-regression-suite (M4.1)
313
+
314
+ ## D-015: M2.2 is deferred until after M3.1, not dropped (decided 2026-08-21)
315
+ - **Category:** sequencing (refines [D-006](#d-006-scheduled-research-runs-post-from-the-primary-model-leg-or-not-at-all))
316
+ - **Decided:** `scheduled-task-model-leg-policy` stays in Milestone 2 and stays
317
+ unstarted. Its architecture is deliberately left unfilled. Revisit at the M2
318
+ checkpoint, with M3.1 in hand.
319
+ - **Alternatives considered:**
320
+ - *Drop it.* Rejected: the Opus pin that pre-empts it is a per-space DB row
321
+ that one hand `DELETE` removes, and nothing would then stand between a
322
+ fallback leg and the group.
323
+ - *Build it next, after feed-watch.* Rejected as premature — the ledger cannot
324
+ yet show a single instance of the problem it solves.
325
+ - **Reasoning:** all 12 live runs since the profile applied (2026-08-20 16:12 →
326
+ 2026-08-21 11:25) ran on `claude-opus-5`, so the M2 checkpoint's "zero group
327
+ posts from a non-primary leg" is already true by construction rather than by
328
+ guarantee. That makes the story's value real but unmeasured, which is the
329
+ exact condition D-008 says to resolve by measuring first. M3.1's harness plus
330
+ the ledger's `outcome` column are what would turn "worth doing" into a number.
331
+ - **Revisit if:** anyone clears a `model.active` pin, or the ledger records a
332
+ run on a non-primary leg — either makes it urgent rather than optional.
333
+ - **Used by:** scheduled-task-model-leg-policy
334
+
335
+ ## D-016: Mechanisms over rules — the reporter's memory is the vault, not the window (decided 2026-08-21)
336
+ - **Category:** pattern (refines [D-003](#d-003-layer-placement-for-football), [D-013](#d-013-tone-is-editorial-and-scan-length-is-per-item-decided-2026-08-20))
337
+ - **Decided:** a defect that comes from *what the model could see* is fixed by
338
+ changing what it can see, never by a rule about the case. Concretely: the
339
+ knowledge vault (`MEMORY.md` + `knowledge/` notes with "current view +
340
+ history", relevance-injected as `<active_episodes>`) becomes the reporter's
341
+ notebook, maintained by a **process** the profile owns (write a fact before
342
+ posting it, read the topic note before answering). The only new content in
343
+ `AGENTS.md` is a **priorities** section — accuracy over register over
344
+ brevity; an instruction-shaped message meant as banter is banter — because
345
+ a tie-break is general where a gag clause is not. Story: `reporter-notebook`
346
+ (M2.3). M3.1's provenance check treats the notebook as a source.
347
+ - **Alternatives considered:**
348
+ - *Per-case `AGENTS.md` edits* (a continuation carve-out, a "names only from
349
+ a source" line, a gag clause, a length split) — the first draft of the
350
+ 2026-08-21 audit. Rejected by the owner as edge-case chasing: every new
351
+ exchange would add a rule, and each rule is one more voice in the stack.
352
+ - *Raise `context.window_size` for the space.* Rejected: tunes the symptom;
353
+ any fixed N is wrong for some conversation, and the window is still a
354
+ window.
355
+ - *A host-side facts table.* Rejected: the vault already exists and is
356
+ already injected; a second store is the two-layer trap.
357
+ - **Reasoning:** audit §1.1 — the bot contradicted its own lineup post from
358
+ 14 hours earlier because a swipe-reply halved the window and three of the
359
+ five remaining turns were task prompts. No rule survives that; a note read
360
+ before answering does. The owner's framing: "maintain a knowledge base with
361
+ important info and recent info — that way he'll make fewer mistakes when
362
+ asked about the same things many times, regardless of the context window."
363
+ - **Revisit if:** the notebook is not being written (the M3.1 lint will show
364
+ names with no provenance) — then the process needs a host-side hook, not
365
+ more prose.
366
+ - **Used by:** reporter-notebook, football-reporter-fixture-harness
367
+
368
+ ## D-017: One voice per space — exclusive persona, and no prompt text about tools the caller cannot use (decided 2026-08-21)
369
+ - **Category:** architecture (applies hardening D-014 in the other direction)
370
+ - **Decided:** two host changes, both general. (1) A space may declare
371
+ `persona.exclusive=true` (builtin `space_config`, profile-owned); the host
372
+ then omits the global character block and does not mount the global
373
+ `AGENTS.md` for that space's runs. (2) The baked platform prompt's
374
+ capability paragraphs — "set a preference" and "warn then mute" — are
375
+ emitted only when the *caller's* permission set includes the tool; the
376
+ payload carries `callerPermissions`. Security text stays unconditional.
377
+ Story: `one-voice-per-space` (M2.4). Football sets the flag.
378
+ - **Alternatives considered:**
379
+ - *Make the space `AGENTS.md` louder* (restate register rules, add "ignore
380
+ the global character"). Rejected: precedence is a coin flip (D-014) and a
381
+ fifth voice does not fix four.
382
+ - *Grant members `prefs.set` in this space.* Rejected: it would make every
383
+ joke a real preference write, and the escalation reflex would become an
384
+ execution reflex.
385
+ - *Edit the global character to say "one of the guys".* Rejected: it is
386
+ global; every other space is a different room.
387
+ - **Reasoning:** audit §1.2 and §2 — the refusal of a running gag was the
388
+ platform prompt ("a standing instruction → set a preference"), the member
389
+ permission row (no `prefs.set`) and the platform prompt again ("inform the
390
+ user") acting in sequence, with `AGENTS.md` §יושרה as a fourth vote. Nothing
391
+ in the space's file could have won that; removing the voices can. (2) has
392
+ value in every space: nobody should be told to do a thing and to refuse it
393
+ in the same turn.
394
+ - **Revisit if:** a future pi/runtime makes layer precedence explicit — then
395
+ ranking may beat removal.
396
+ - **Used by:** one-voice-per-space, football-reporter profile
397
+
398
+ ## D-018: The chat window holds conversation, and silence is a legal chat reply (decided 2026-08-21)
399
+ - **Category:** runtime behaviour
400
+ - **Decided:** (1) `messages` rows record their `source`; the chat window
401
+ counts and returns user turns that are not scheduler prompts, and always
402
+ keeps assistant rows (the bot's posts are conversation; the procedure that
403
+ produced them is not); a swipe-reply adds its anchor chain to the window
404
+ instead of halving it. (2) The chat path honours the same no-op sentinel as
405
+ the scheduler: nothing is sent, the user row stays, no assistant row is
406
+ written. Default-on, no new config. Story: `chat-window-and-silence`
407
+ (M2.5).
408
+ - **Alternatives considered:**
409
+ - *Leave the window; rely on the notebook.* Rejected: the window still
410
+ carries register and continuity, and three task prompts in five turns
411
+ crowd those out too.
412
+ - *A per-space "silence allowed" flag.* Rejected: the rule is already
413
+ written in the standard; making it obeyable is not a preference.
414
+ - **Reasoning:** audit §1.3 — "מקובל." to "אל תענה לי" is the runtime
415
+ contradicting the file; §1.1 — the window arithmetic (10 → 5 on reply,
416
+ three of five procedural). Both fixes are small and survive every later
417
+ milestone unchanged.
418
+ - **Revisit if:** no-op chat replies become frequent in the ledger/re-audit
419
+ — then the model is using silence to dodge, and the standard's priorities
420
+ section needs to say when silence is not allowed.
421
+ - **Used by:** chat-window-and-silence
422
+
423
+ ## D-019: The reporter's notebook *is* napkin's vault — seeded by the profile, maintained nightly by napkin, written in-turn by the process (decided 2026-08-21)
424
+ - **Category:** architecture (refines [D-016](#d-016-mechanisms-over-rules--the-reporters-memory-is-the-vault-not-the-window-decided-2026-08-21))
425
+ - **Decided:** the football topic notes (clubs, competitions, sagas) are
426
+ **episodes** in napkin's own shape under `knowledge/episodes/` — the only
427
+ directory the host scores and injects (`buildEpisodeContext`), with
428
+ `type: episode`, `status`, `last_mentioned`, `mentions`, JSON `keywords`,
429
+ `summary`, `## Current State`, `## History`. Three writers, one file: the
430
+ profile **seeds** a note once (`workspace_seed`, write-if-absent — never
431
+ overwrites, a model-edited note is not drift); the in-turn process
432
+ (`## מחברת`) appends a fact **before** it is posted, so the same-day
433
+ follow-up sees it; napkin's nightly distillation updates `last_mentioned`,
434
+ `mentions` and `## Current State` from whatever was posted, and its
435
+ weekly consolidation ages episodes and writes `.memory-suggestions.md`.
436
+ `MEMORY.md` is seeded for standing facts only (coverage + the admin's
437
+ equal-weight ruling, cadence, register rules the group set, spelling rule,
438
+ open commitments) and stays model-owned. The profile may **extend** the
439
+ distiller with a per-space addendum (`knowledge/.napkin/distill.md`,
440
+ appended to the kb-distillation prompt when present — a small general
441
+ napkin feature) and never replaces it. No parallel schema, no host-side
442
+ facts table.
443
+ - **Alternatives considered:**
444
+ - *M2.3 as first written* — `knowledge/clubs|competitions|transfers/` notes
445
+ maintained by the model alone, distillation out of scope. Rejected on
446
+ facts: those directories are never injected, and napkin was already
447
+ producing the notes (once it ran) with the right frontmatter.
448
+ - *Let napkin do it all, drop the in-turn process.* Rejected: distillation
449
+ sees chat only (what was posted, not what was fetched), and lags a day;
450
+ the 08-21 hallucination was a same-day follow-up.
451
+ - *Teach the distiller the football schema by forking its prompt.* Rejected:
452
+ the prompt is napkin's, shared by every space; an addendum file is
453
+ profile-owned and general.
454
+ - **Reasoning:** `docs/notes/napkin-revisit-2026-08-21.md` §1–§3 — the first
455
+ real distillation of `football-friends` (10 episodes, 8 people, 8 references)
456
+ already held the 20.8 Hapoel XI, the return legs, Turiel's injury, the
457
+ 18.8 admin ruling and a verified spelling list, all dated and sourced and
458
+ injectable as-is. Building a second store beside it is the two-layer trap
459
+ D-016 already names.
460
+ - **Revisit if:** the nightly run grows too slow or too costly for the space
461
+ (today 5–11 min per day of chat on Opus, growing with vault size) — then
462
+ pin distillation to a cheaper leg or a tighter backfill window; or if the
463
+ M3.1 provenance check shows posted names that are in neither a fetched
464
+ body nor an episode — then the in-turn process needs a host-side hook.
465
+ - **Implementation note (2026-08-21, when M2.3 shipped):** the seeded
466
+ `MEMORY.md` came out **smaller** than this decision described. The list above
467
+ named the cadence, the coverage topics, the spelling rule and "check where
468
+ the player is now" as `MEMORY.md` content, but `AGENTS.md` already owns all
469
+ four as rules, and the story's own non-negotiable forbids one rule in two
470
+ layers — both files reach the same prompt. The seed therefore carries only
471
+ what `AGENTS.md` does not: the group's dated decisions (the 18/08 two-source
472
+ ruling, the 15/08 no-correction-block rule, the unapplied 17/08 request),
473
+ open commitments, the verified spellings themselves (facts, not the rule
474
+ about checking them), and the index of standing notebooks. Also found while
475
+ implementing: `## Current State (as of …)` is not matched by the injector, so
476
+ every note napkin had written for this space was reaching the model as a
477
+ summary line with no facts —
478
+ `docs/bugs/episode-current-state-heading-suffix-drops-body.md`.
479
+ - **Used by:** reporter-notebook (M2.3, re-cut), football-reporter-fixture-harness (provenance source = `episodes/`), napkin (distill addendum)
@@ -0,0 +1,91 @@
1
+ # Football Reporter Profile
2
+
3
+ **Status**: Active
4
+ **Slug**: football-reporter-profile
5
+ **Created**: 2026-08-20
6
+ **Last updated**: 2026-08-21 (roadmap extended after the 2026-08-21 audit)
7
+
8
+ ---
9
+
10
+ ## Goal Statement
11
+
12
+ Redefine the `football-friends` bot's behaviour as a dedicated *top football
13
+ reporter* profile — without the commercial plumbing the business profiles in
14
+ `tagula-agents` carry — and use that profile as the one owner of the editorial
15
+ rules that are today scattered across preferences, two task prompts and an empty
16
+ space `AGENTS.md`. Then put a deterministic feed in front of the model
17
+ (feed-watch), prove the output quality with fixtures and a run ledger, fix
18
+ unrelated bugs as they surface, and finally re-test every defect the
19
+ 2026-08-19 48-hour audit found (F1–F8) with dedicated checks while monitoring
20
+ for new ones.
21
+
22
+ The sequence the owner chose: (1) profile + editorial standard in one layer →
23
+ (2) feed-watch → (3) fixture harness + ledger, with a parallel bug track for
24
+ anything unrelated, then (4) regression suite + re-audit.
25
+
26
+ ## Constraints
27
+
28
+ - **The live bot is the beta tester.** Every change runs on `dev` in
29
+ `~/whatsapp-bot` before it can be promoted; nothing is "done" until it has
30
+ been observed in `mercury service logs -f` and in the group's behaviour.
31
+ - **Never test in the group.** Manual checks run in the private `main` space
32
+ (standing rule, `docs/backlog/feed-watch.md` Non-Negotiables).
33
+ - **Prompt-layer facts** (`docs/notes/bot-configuration-surfaces.md`,
34
+ D-014 in `whatsapp-bot-hardening`): preference values are capped at 500
35
+ chars; `spaces/<id>/AGENTS.md`, preferences and `tasks.prompt` are read per
36
+ run with no restart; the baked system prompt and `mrctl` need an image
37
+ rebuild; precedence between layers is a coin flip, so a rule may live in one
38
+ layer only.
39
+ - **Mercury's built-in profiles are project-wide.** `mercury profiles apply`
40
+ overwrites `.mercury/global/AGENTS.md` and pins `member_permissions` for
41
+ every space (`src/core/profiles.ts`). `~/whatsapp-bot` hosts several spaces,
42
+ so the football profile must be space-scoped.
43
+ - **Cost baseline:** $36.98 / 7 days for the space (~$5.30/day), 4 of 10 scans
44
+ `NO_UPDATE` at ~$0.50 each (audit F8). Report numbers; measure two weeks
45
+ before writing any threshold.
46
+ - Output language is Hebrew; the group has said twice it does not want
47
+ teasing links for the daily article (audit F3).
48
+
49
+ ## Preferences
50
+
51
+ - Borrow the *shape* of a tagula profile — one `AGENTS.md` owner generated
52
+ from source, extension + config, apply/check script, fixture harness — and
53
+ skip `member_permissions`, engines, config registries, Drive, trial gates
54
+ (required: "without the commercial bullshit").
55
+ - Feed-watch stays as architected in `docs/backlog/feed-watch.md`; this goal
56
+ absorbs it as Milestone 2 rather than re-planning it (preferred).
57
+ - Bugs unrelated to the profile go through `/f-bug-report` → `/f-bug-fix` in
58
+ parallel and are never folded into a story (required — keeps stories scoped
59
+ and behaviour-preserving).
60
+ - Deterministic first: a rule that can be a host-side gate or a script is not
61
+ left as a prose rule the model must comply with (preferred; the
62
+ `normalizeChatMarkdown` precedent).
63
+
64
+ ## Target Users
65
+
66
+ - **Group members of `football-friends`** — want real, fresh, sourced news in
67
+ chat, in Hebrew, with nothing when nothing happened. (inferred from the audit
68
+ transcript)
69
+ - **The owner / space admin** — wants one place to edit the bot's editorial
70
+ behaviour, a bill that tracks news rather than the clock, and a way to see
71
+ problems without reading 128 chat rows.
72
+ - **The operator (same person)** — wants every previously reported defect to
73
+ have a test that would catch its return.
74
+
75
+ ## Success Criteria
76
+
77
+ 1. Every editorial rule for `football-friends` has exactly one home, and
78
+ `scripts/space-profile.ts check` proves the deployed `AGENTS.md` and task
79
+ prompts are byte-identical to the repo source (zero drift) — `club-coverage`
80
+ and `push-rules` preferences are gone.
81
+ 2. No scheduled post reaches the group without a named source and a publication
82
+ date, and no item older than its window does — verified by the fixture
83
+ harness and by the two-week re-audit.
84
+ 3. Empty scans cost nothing: the space's spend is reported per run in the
85
+ ledger and, measured over two weeks, tracks news rather than the clock
86
+ (today: ~$1/day confirming nothing happened).
87
+ 4. A run that landed on a fallback model leg never posts to the group
88
+ (audit F2).
89
+ 5. F1–F7 each have a dedicated test or harness check; the re-audit after
90
+ Milestone 4 finds no recurrence, and any new defect is visible in the run
91
+ ledger without reading the chat.