@oneie/claude 0.8.0 → 0.10.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (230) hide show
  1. package/.claude-plugin/plugin.json +1 -1
  2. package/agents/abm-strategist.md +67 -1
  3. package/agents/ads-meta.md +67 -1
  4. package/agents/analyst.md +67 -1
  5. package/agents/animator.md +108 -0
  6. package/agents/architect.md +269 -20
  7. package/agents/brand-guardian.md +67 -1
  8. package/agents/brand-strategist.md +67 -1
  9. package/agents/campaign-content.md +67 -1
  10. package/agents/campaign-email.md +67 -1
  11. package/agents/campaign-sms.md +67 -1
  12. package/agents/campaign-social.md +67 -1
  13. package/agents/cco.md +83 -2
  14. package/agents/ceo.md +108 -11
  15. package/agents/chairman.md +197 -0
  16. package/agents/cmo.md +82 -2
  17. package/agents/community-greeter.md +67 -1
  18. package/agents/community-moderator.md +67 -1
  19. package/agents/compliance.md +67 -1
  20. package/agents/copywriter.md +67 -1
  21. package/agents/creative-strategist.md +67 -1
  22. package/agents/cro.md +81 -1
  23. package/agents/cto.md +266 -28
  24. package/agents/customer-interviewer.md +67 -1
  25. package/agents/customer-researcher.md +67 -1
  26. package/agents/customer-success-manager.md +67 -1
  27. package/agents/customer-trainer.md +67 -1
  28. package/agents/cxo.md +82 -1
  29. package/agents/demand-creator.md +67 -1
  30. package/agents/demo-mover.md +67 -1
  31. package/agents/demo-specialist.md +67 -1
  32. package/agents/demo-thai-family-law.md +67 -1
  33. package/agents/designer.md +67 -1
  34. package/agents/discovery-caller.md +67 -1
  35. package/agents/doctor.md +269 -0
  36. package/agents/educate-coach.md +67 -1
  37. package/agents/elevate-tutor.md +67 -1
  38. package/agents/email-lifecycle-marketer.md +67 -1
  39. package/agents/engage-specialist.md +67 -1
  40. package/agents/events-coordinator.md +67 -1
  41. package/agents/foundation-builder.md +67 -1
  42. package/agents/funnel-architect.md +67 -1
  43. package/agents/gift-creator.md +67 -1
  44. package/agents/google-ads.md +67 -1
  45. package/agents/guide.md +67 -1
  46. package/agents/helpdesk-dispatcher.md +67 -1
  47. package/agents/hook-specialist.md +67 -1
  48. package/agents/identify-optimizer.md +67 -1
  49. package/agents/implementer.md +313 -45
  50. package/agents/incident-commander.md +67 -1
  51. package/agents/insights-lead.md +87 -1
  52. package/agents/journey-runner.md +67 -1
  53. package/agents/linkedin-ads.md +67 -1
  54. package/agents/live-sales-chat.md +67 -1
  55. package/agents/market-researcher.md +67 -1
  56. package/agents/media-buyer.md +67 -1
  57. package/agents/memory-keeper.md +195 -0
  58. package/agents/movers-customer-researcher.md +67 -1
  59. package/agents/movers-foundation-builder.md +67 -1
  60. package/agents/movers-market-researcher.md +67 -1
  61. package/agents/movers-pricing-strategist.md +67 -1
  62. package/agents/nurture-architect.md +67 -1
  63. package/agents/offer-architect.md +67 -1
  64. package/agents/onboarder.md +67 -1
  65. package/agents/onboarding-specialist.md +67 -1
  66. package/agents/operations-dashboard.md +87 -1
  67. package/agents/perf-engineer.md +333 -37
  68. package/agents/playbook-writer.md +67 -1
  69. package/agents/plg-strategist.md +67 -1
  70. package/agents/positioning-architect.md +67 -1
  71. package/agents/press-officer.md +67 -1
  72. package/agents/pricing-strategist.md +67 -1
  73. package/agents/privacy-officer.md +67 -1
  74. package/agents/referral-manager.md +67 -1
  75. package/agents/refine-analyst.md +67 -1
  76. package/agents/release-manager.md +446 -39
  77. package/agents/renewals-upsell-rep.md +67 -1
  78. package/agents/review-engineer.md +319 -45
  79. package/agents/rewards-steward.md +67 -1
  80. package/agents/sales-call-coach.md +67 -1
  81. package/agents/sales-closer.md +67 -1
  82. package/agents/security-auditor.md +343 -48
  83. package/agents/sell-closer.md +67 -1
  84. package/agents/share-amplifier.md +67 -1
  85. package/agents/social-media-manager.md +67 -1
  86. package/agents/storyteller.md +301 -0
  87. package/agents/strategist.md +67 -1
  88. package/agents/strategy-aligner.md +67 -1
  89. package/agents/support-agent.md +67 -1
  90. package/agents/tagger.md +327 -0
  91. package/agents/tech-writer.md +195 -22
  92. package/agents/test-engineer.md +398 -29
  93. package/agents/tiktok-ads.md +67 -1
  94. package/agents/tracking-engineer.md +67 -1
  95. package/agents/trailkeeper.md +181 -0
  96. package/agents/upsell-strategist.md +67 -1
  97. package/agents/voice.md +67 -1
  98. package/agents/w1-recon.md +1 -1
  99. package/agents/w2-decide.md +1 -1
  100. package/agents/w3-edit.md +8 -2
  101. package/agents/w4-verify.md +13 -0
  102. package/agents/workflow-optimiser.md +81 -1
  103. package/commands/close.md +916 -160
  104. package/commands/deploy.md +102 -724
  105. package/commands/do.md +58 -2
  106. package/commands/sweep.md +159 -0
  107. package/commands/tasks.md +222 -0
  108. package/hooks/scripts/dev-only.sh +135 -0
  109. package/hooks/scripts/git-add-guard.sh +37 -2
  110. package/hooks/scripts/session-start.sh +32 -4
  111. package/package.json +1 -1
  112. package/rules/scripts.md +85 -0
  113. package/scripts/CLAUDE.md +315 -0
  114. package/scripts/ad-copy-lint.sh +656 -0
  115. package/scripts/agent-actor-parity.sh +129 -0
  116. package/scripts/blocks-manifest-cached.sh +100 -0
  117. package/scripts/chat-context-check.sh +89 -0
  118. package/scripts/chrome.mjs +18 -0
  119. package/scripts/close-metrics.sh +587 -0
  120. package/scripts/close-owner.sh +326 -0
  121. package/scripts/db-sync-lock-check.sh +116 -0
  122. package/scripts/deploy-emit.sh +311 -0
  123. package/scripts/deploy-gate-check.sh +155 -0
  124. package/scripts/deploy-ready.sh +78 -0
  125. package/scripts/deploy-record.sh +605 -0
  126. package/scripts/deploy-schema-check.sh +58 -0
  127. package/scripts/deploy.sh +393 -243
  128. package/scripts/do-auto.sh +127 -26
  129. package/scripts/do-board.sh +429 -0
  130. package/scripts/do-close.sh +1184 -0
  131. package/scripts/do-consumer-sweep.sh +18 -1
  132. package/scripts/do-decide.sh +476 -0
  133. package/scripts/do-fleet.sh +8 -2
  134. package/scripts/do-plan-json.mjs +110 -12
  135. package/scripts/do-prove-selftest.sh +108 -0
  136. package/scripts/do-prove.sh +86 -10
  137. package/scripts/do-rank.py +200 -3
  138. package/scripts/do-reconcile.sh +73 -12
  139. package/scripts/do-signal.sh +101 -23
  140. package/scripts/do-smoke.sh +18 -1
  141. package/scripts/do-w4-gates.sh +11 -1
  142. package/scripts/do-world-check.sh +153 -0
  143. package/scripts/download-stats.sh +172 -0
  144. package/scripts/factory-brief-check.sh +330 -0
  145. package/scripts/factory-check.sh +18 -1
  146. package/scripts/factory-close-check.sh +257 -0
  147. package/scripts/factory-emit.sh +211 -0
  148. package/scripts/factory-executor-check.mjs +353 -0
  149. package/scripts/factory-peak.sh +301 -0
  150. package/scripts/factory-repo.sh +71 -0
  151. package/scripts/factory-review-check.mjs +61 -0
  152. package/scripts/factory-tasks-check.sh +18 -1
  153. package/scripts/fixtures/factory-brief-real.md +44 -0
  154. package/scripts/flywheel-outcome.sh +63 -0
  155. package/scripts/gate-reaper-check.sh +98 -0
  156. package/scripts/gate-reaper.sh +9 -0
  157. package/scripts/gate-watchdog.sh +619 -0
  158. package/scripts/gc-content-check.sh +142 -0
  159. package/scripts/gh-traffic-capture.sh +153 -0
  160. package/scripts/govern-order-check.sh +202 -0
  161. package/scripts/governor-doors-check.sh +86 -5
  162. package/scripts/health.sh +448 -0
  163. package/scripts/id-inventory.mjs +418 -0
  164. package/scripts/incident.sh +212 -0
  165. package/scripts/land.sh +755 -45
  166. package/scripts/lib/gc-finished.sh +77 -0
  167. package/scripts/livekit-ratchet.sh +18 -1
  168. package/scripts/machine-check.sh +1 -1
  169. package/scripts/memory-index-budget.sh +79 -0
  170. package/scripts/npm-downloads.sh +109 -0
  171. package/scripts/one-agents.mjs +204 -8
  172. package/scripts/one-resume.sh +31 -3
  173. package/scripts/pr-body.sh +335 -0
  174. package/scripts/preview-fd-check.sh +289 -0
  175. package/scripts/redirect-lint.sh +169 -0
  176. package/scripts/release.sh +40 -6
  177. package/scripts/resume-lost-sessions.sh +68 -0
  178. package/scripts/shoot-pages.mjs +140 -0
  179. package/scripts/signal-meta-backfill.ts +451 -0
  180. package/scripts/signal-watch.sh +63 -6
  181. package/scripts/speed-cache-check.sh +12 -2
  182. package/scripts/sweep.sh +426 -0
  183. package/scripts/task-titles-dump.ts +101 -0
  184. package/scripts/test-cached.sh +47 -10
  185. package/scripts/test-lanes.sh +14 -0
  186. package/scripts/thread-name-backfill.ts +215 -0
  187. package/scripts/triage-shape-check.sh +149 -0
  188. package/scripts/tsc-cached.sh +155 -8
  189. package/scripts/typedb-flake-check.sh +3 -1
  190. package/scripts/urls-lint.sh +8 -0
  191. package/scripts/verify-board-doors.sh +80 -0
  192. package/scripts/verify-fast.sh +159 -6
  193. package/scripts/worktree-up.sh +21 -3
  194. package/skills/astro/SKILL.md +9 -3
  195. package/skills/astro/optimize-performance.md +3 -2
  196. package/skills/cloudflare/SKILL.md +3 -2
  197. package/skills/cloudflare-security-audit/AI-AND-LLM.md +83 -0
  198. package/skills/cloudflare-security-audit/ATTACK-CLASSES.md +130 -0
  199. package/skills/cloudflare-security-audit/CLIENT-SIDE.md +83 -0
  200. package/skills/cloudflare-security-audit/CLOUD-AND-DEPLOYMENT.md +86 -0
  201. package/skills/cloudflare-security-audit/DATA-ISOLATION-AND-LIFECYCLE.md +84 -0
  202. package/skills/cloudflare-security-audit/DESKTOP-MOBILE-AND-LOCAL-IPC.md +89 -0
  203. package/skills/cloudflare-security-audit/HUNTING.md +251 -0
  204. package/skills/cloudflare-security-audit/LICENSE +21 -0
  205. package/skills/cloudflare-security-audit/MEMORY-SAFETY-AND-BINARY.md +101 -0
  206. package/skills/cloudflare-security-audit/PROTOCOLS-RPC-AND-MESSAGING.md +81 -0
  207. package/skills/cloudflare-security-audit/PROVENANCE.md +78 -0
  208. package/skills/cloudflare-security-audit/RECONNAISSANCE.md +156 -0
  209. package/skills/cloudflare-security-audit/RESOURCE-EXHAUSTION-AND-AVAILABILITY.md +78 -0
  210. package/skills/cloudflare-security-audit/SKILL.md +192 -0
  211. package/skills/cloudflare-security-audit/SUPPLY-CHAIN-AND-RELEASE.md +73 -0
  212. package/skills/cloudflare-security-audit/VALIDATION-AND-REPORTING.md +186 -0
  213. package/skills/cloudflare-security-audit/WEB-PROTOCOL-AND-AUTH.md +105 -0
  214. package/skills/cloudflare-security-audit/report-schema.json +461 -0
  215. package/skills/cloudflare-security-audit/validate-coverage-ledger.cjs +872 -0
  216. package/skills/cloudflare-security-audit/validate-coverage-ledger.test.cjs +740 -0
  217. package/skills/cloudflare-security-audit/validate-findings.cjs +773 -0
  218. package/skills/cloudflare-security-audit/validate-findings.test.cjs +652 -0
  219. package/skills/deploy/REFERENCE.md +713 -0
  220. package/skills/deploy/SKILL.md +140 -0
  221. package/skills/fleet-audit/SKILL.md +58 -0
  222. package/skills/meeting/SKILL.md +220 -0
  223. package/skills/planning/SKILL.md +256 -0
  224. package/skills/shadcn/SKILL.md +1 -1
  225. package/skills/typedb/SKILL.md +7 -0
  226. package/skills/voice/SKILL.md +94 -6
  227. package/skills/voice/corpus-check.sh +87 -0
  228. package/templates/template-agent.md +7 -1
  229. package/templates/template-feature.md +9 -0
  230. package/templates/template-todo.md +29 -0
@@ -0,0 +1,713 @@
1
+ # deploy — REFERENCE
2
+
3
+ The trap encyclopedia for `.claude/commands/deploy.md`. Split out 2026-09-17: the
4
+ command page is what a conductor reads to decide *who* deploys (the answer is
5
+ always `release-manager`); this is what that agent reads to actually do it.
6
+
7
+ Nothing here was rewritten in the split. Every section below is the text that
8
+ stood in `deploy.md`, with its measurements and dates intact — a trap loses its
9
+ authority the moment it is paraphrased.
10
+
11
+ Authority for the PROCEDURE is still `.claude/scripts/deploy.sh`, never this file.
12
+ Edit the script when a step changes; edit here only when what a gate *means* changes.
13
+
14
+ ---
15
+
16
+ ### The suite runs as two lanes, not one run (measured 2026-09-03)
17
+
18
+ The gate had become a coin flip. Four full runs on an **unchanged tree** gave
19
+ four different failure sets — 8, 12, 13, 8 — and every failing file in all four
20
+ was one of the 19 suites marked `real TypeDB`. Not one of the other ~1126 files
21
+ ever failed. Each run's log carried 41-72 `upstream 503` plus 25-38
22
+ `upstream 500`, and those print only *after* the retry budget is spent
23
+ (`substrate.ts:208`).
24
+
25
+ The cause is capacity, not code. `substrate.ts:56-60` already recorded it in
26
+ 2026-07-25: the shared TypeDB Cloud gateway 503s **in bursts lasting seconds**,
27
+ while the retry cover is 500ms + 1000ms. Eight vitest forks against one external
28
+ singleton turn that into a dice roll — and a dice-roll gate costs far more than
29
+ it saves. The 2026-09-03 deploy spent hours on re-runs and root-cause agents for
30
+ failures that were never in the diff.
31
+
32
+ `.claude/scripts/test-lanes.sh` splits the suite by its two binding constraints
33
+ and runs them **concurrently**:
34
+
35
+ | lane | files | bound by | concurrency |
36
+ |---|---|---|---|
37
+ | `pool` | ~1126 | CPU | 8 forks |
38
+ | `typedb` | 19 | one shared gateway | serial (`--no-file-parallelism --maxWorkers=1`) |
39
+
40
+ The `typedb` lane spends its life waiting on a socket, so it costs almost no CPU
41
+ while `pool` saturates the cores: wall clock is **max(), not sum()**.
42
+
43
+ | | before | after |
44
+ |---|---|---|
45
+ | wall clock | 238-282s | **121.9s** |
46
+ | result | RED, rotating victims | **BOTH GREEN** — pool 1125 passed, typedb 19 passed |
47
+ | tests | 8-13 failing | 10940 passed |
48
+
49
+ **This is not `--no-file-parallelism` over the whole suite** — that serialises
50
+ 1126 innocent files to fix 19. Every test still runs, no assertion is weakened,
51
+ nothing is skipped; only the suites sharing the external singleton are
52
+ serialised, which removes the contention at its source.
53
+
54
+ **The selector is read from the tree, never hardcoded.** It greps the
55
+ `real TypeDB` marker the suites already carry, so a new such suite joins the
56
+ serial lane automatically. A frozen list is exactly how this rots back into
57
+ flakiness — `--self-test` asserts the selector finds >0 files, is a strict
58
+ subset, and that every path it names exists.
59
+
60
+ Both lanes still go through `test-cached.sh`, so passes are memoised per lane
61
+ and a RED is still never cached. `test-full.sh` remains the one definition of
62
+ the vitest flags and hands them over as `TEST_FULL_ARGS`.
63
+
64
+ ```bash
65
+ bash .claude/scripts/test-lanes.sh --list # show the split, run nothing
66
+ bash .claude/scripts/test-lanes.sh --self-test # prove the selector still selects
67
+ ```
68
+
69
+ **Two lanes mean two `Tests N passed` lines in the log.** `TESTS_REPORT` used to
70
+ `tail -1` and reported only the 19-file lane — the first two-lane production
71
+ deploy said `Tests 143 passed` for a run that executed **10940**. It now sums the
72
+ lanes and says `(2 lanes)`. A gate that is fine while the report understates it
73
+ by two orders of magnitude is the same dishonesty as calling a fast pass a full
74
+ one.
75
+
76
+ ### Making the deploy faster — two disproved ideas (measured 2026-09-03)
77
+
78
+ Recorded so nobody spends an afternoon re-deriving them. Both were plausible,
79
+ both are wrong, and the second fails in a way that reads like a pass.
80
+
81
+ **The gates are at their floor at ~256s. Rescheduling does not move them.**
82
+ From the 2026-09-03 production log:
83
+
84
+ ```
85
+ vitest 256s (the lanes, run alone: 121.9s)
86
+ build 256s (astro's own report: 2m 13s = 133s)
87
+ gates wall-clock: 256s
88
+ ```
89
+
90
+ Serial would be 133+122 = 255s — *identical*. Each gate paid ~2x its solo cost
91
+ and the overlap bought nothing, because eight vitest forks plus an 8 GiB-heap
92
+ build on a 10-core box is oversubscription, not parallelism.
93
+
94
+ The obvious next move — give the build room by capping the pool — makes it
95
+ **worse**:
96
+
97
+ | | gates wall-clock | build |
98
+ |---|---|---|
99
+ | 8 forks (baseline) | **256s** | 133s |
100
+ | 5 forks (`VERIFY_POOL_FORKS=5`) | **272s** | 209s |
101
+
102
+ Parallel, serial and capped all land at 250-270s. That is the floor for
103
+ build+suite on this hardware as currently shaped.
104
+
105
+ **The real target is import cost, and the obvious fix is blocked upstream.**
106
+ The pool lane, run solo, spends more time loading modules than running tests:
107
+
108
+ ```
109
+ import 145.42s | tests 107.87s
110
+ ```
111
+
112
+ `vitest.config.ts:131` sets `pool: 'forks'`, and every fork re-imports the whole
113
+ module graph independently. Threads share a module cache, so that ought to be
114
+ the win — but **vitest 4.1.7's threads pool is broken in this repo**. Every
115
+ worker dies with `The worker thread was torn down or never initialized. This is
116
+ a bug in Vitest.` Do not reach for `--pool=threads` until vitest is upgraded and
117
+ this is re-tested.
118
+
119
+ **Read the exit code, not the summary.** That broken run printed:
120
+
121
+ ```
122
+ Test Files no tests
123
+ Tests no tests
124
+ Duration 162ms
125
+ ```
126
+
127
+ with `rc=1`. "no tests" in a 162ms run is a *collect crash*, one careless glance
128
+ from being reported as a clean pass — the same shape as
129
+ `vitest-collect-crash-reads-as-pass`. An empty selection is never a pass.
130
+
131
+ **What actually speeds a deploy today:** scope it. `./deploy dev` runs the fast
132
+ lane with no approval; `./deploy workers` skips the astro rebuild entirely when
133
+ `one.ie/web` did not change. The astro build is the long pole and is already
134
+ memoised by tree fingerprint — a warm tree pays none of it.
135
+
136
+ ### Why the gates are not all parallel (measured 2026-08-19)
137
+
138
+ Steps 1+3 used to launch seven heavy processes with a bare `&`: five `tsc`,
139
+ vitest, and the astro build. That is not parallelism, it is a swap storm. The
140
+ astro build carries `--max-old-space-size=8192` and vitest runs a driver plus
141
+ four 1GB forks, so the two together want ~14GB — on a 24GB box where editors
142
+ and sessions are already resident.
143
+
144
+ Paging is a cliff, not a slope. Same suite, same commit, same machine:
145
+
146
+ | vitest ran… | wall-clock | outcome |
147
+ |---|---|---|
148
+ | alone | **176s** | 912 files, 7624 tests, green |
149
+ | beside the build + 5×tsc | **1145s** | 0.0% CPU, no log output, killed |
150
+
151
+ It happened twice in one day before anyone read it as anything but "tests are
152
+ slow". They are not slow — the suite's own summary accounts for only ~165s of
153
+ actual test time.
154
+
155
+ **Correction, same day: memory is NOT the proven cause of the hang.** The
156
+ serialisation below is still worth having — 14GB of overlap on a 24GB box is
157
+ real — but the 1145s runs were later reproduced with the build serialised and
158
+ the box at 41% free, no swap thrash, and no network connections held. See
159
+ "The vitest gate hangs" below. Do not cite the memory story as the explanation
160
+ for a hung gate; it explains a *slow* gate, not a *parked* one.
161
+
162
+ So the heavy gates are now **priced before they launch**: `heavy_free_gb`
163
+ reads free memory the way `lib/govern.sh` does and compares it against
164
+ `DEPLOY_HEAVY_NEED_GB` (default 14). Enough → they overlap as before. Not
165
+ enough → they run one after the other, which is roughly **3x faster
166
+ end-to-end** than "parallel", because the parallel version spends its time
167
+ paging. The five typechecks stay parallel and ungoverned; cached by tree
168
+ fingerprint, they cost ~1s each.
169
+
170
+ Both heavy gates also run under `gate-run.sh`, for its wall-clock **bound**
171
+ (`DEPLOY_GATE_TIMEOUT`, default 900s — a healthy vitest is 176s and a build
172
+ 174s) and its process-group **reap**. macOS ships no `timeout(1)`, and killing
173
+ only the shell reparents the vitest forks to launchd. A hung gate must die on a
174
+ clock rather than outlive the deploy.
175
+
176
+ A probe that cannot read memory returns a large number, so a broken sensor
177
+ never silently serialises the pipeline. Escape hatches: `DEPLOY_HEAVY_PARALLEL=1`
178
+ forces overlap, `DEPLOY_HEAVY_NEED_GB` retunes the threshold.
179
+
180
+ ### The vitest gate hangs — open, characterised, unexplained
181
+
182
+ > **Numbers below are from 2026-08-19 and are superseded as measurements** (the
183
+ > suite was 912 files then and is 1145 now; it runs as two lanes since
184
+ > 2026-09-03, ~122s wall). The *diagnosis* still stands and the hang is still
185
+ > unexplained, so the section is kept as-is rather than half-rewritten. Note the
186
+ > non-TTY fork stall below bit again on 2026-09-03: a background re-run with
187
+ > stdout to a file parked until it was killed, and the `script -qeF /dev/null`
188
+ > wrapper fixed it. That wrapper is not optional for any logged run.
189
+
190
+
191
+ Three times on 2026-08-19 the vitest gate parked indefinitely and had to be
192
+ killed. What is established:
193
+
194
+ - The suite itself is healthy: run on its own it completes in **176s**, 912
195
+ files / 7624 tests / 0 failures, on the same commit and machine.
196
+ - The hang is **not** the astro build competing for memory. It reproduced with
197
+ the heavy gates serialised, `memory_pressure` reporting 41% free and low
198
+ pageouts.
199
+ - It is **not** network. At the moment of the hang the vitest main process held
200
+ **no** open sockets, and neither did its worker.
201
+ - The shape is a fork-pool stall: main parked in `LibuvStreamWrap::OnUvRead`
202
+ waiting on worker IPC, while its single worker sat at **0.43s CPU / 46MB RSS**
203
+ seven minutes in — a fork that was spawned and never given work. The log
204
+ always stops ~90s in, after the jsdom noise, before any test result.
205
+ - `globalSetup` was ruled out: `tests/_global-setup.ts` is local only (reads an
206
+ env file, prints) with no network call to hang on.
207
+
208
+ Not established: why. The one difference between every hung run and the green
209
+ one is that the green run passed `--testTimeout/--hookTimeout/--teardownTimeout`
210
+ explicitly — but nothing in that run came close to a timeout, so that may be
211
+ coincidence rather than cause.
212
+
213
+ Until it is understood, the gate runs under `gate-run.sh`'s bound, so a hang
214
+ now dies on a clock instead of outliving the deploy. If it bites you: the suite
215
+ is trustworthy run directly (`cd one.ie/web && bunx vitest run`), and a deploy
216
+ whose only commits since a green run are harness/doc changes can legitimately
217
+ use `--skip-tests` — verify with
218
+ `git diff --name-only <green-sha>..HEAD | grep -v '^\.claude/'` returning empty.
219
+
220
+ **Health endpoints** (custom domains only — `*.oneie.workers.dev` is blocked on
221
+ this network, curl exit 6/28): `api.one.ie/health` · `one.ie/api/health`
222
+ (assert `"status":"ok"`) · `channels.one.ie/health` · `pay.one.ie/status`
223
+ (assert `"ok"`; there is no `/health` on that entry — it 404s). The agents
224
+ worker is named **`channels`**, reachable at `channels.one.ie`.
225
+
226
+ **Build OOM, fixed 2026-07-19:** the build was V8-heap-OOMing (`FATAL ERROR:
227
+ Reached heap limit`, `Abort trap: 6`, exit 134) at Node's default ~4 GiB
228
+ old-space ceiling — not a system memory shortage. `package.json`'s `build`
229
+ script now sets `NODE_OPTIONS=--max-old-space-size=8192` inline; no manual env
230
+ var needed.
231
+
232
+ **Migrations and the trap:** `wrangler.toml`'s `[[routes]]`/`[triggers]` were
233
+ reconciled to top level 2026-07-04, so there is no env-scoped target left to
234
+ hit and `--env production` has no block to resolve against. Passing it anyway
235
+ silently targets the `one-prod-production` decoy — see the trap note at the top
236
+ of this file.
237
+
238
+ ## Bundle Size Rules (CF Workers Free Tier — 3 MiB gzipped upload)
239
+
240
+ The Astro Worker upload must stay under **3 MiB gzipped** on the free tier (10 MiB
241
+ on paid). Wrangler reports both `Total Upload` (uncompressed) and `gzip` — only
242
+ gzip counts toward the ceiling. All chunks in `dist/server/chunks/` are uploaded
243
+ together; dynamic `await import()` does NOT exclude code from the upload.
244
+
245
+ These rules are **LOCKED** — do not revert them. Apply identically to any
246
+ developer template we ship (`oneie init` Workers scaffold mirrors this shape).
247
+
248
+ ### Rule 1 — `syntaxHighlight: false` in `astro.config.mjs`
249
+
250
+ ```js
251
+ markdown: { syntaxHighlight: false }
252
+ ```
253
+
254
+ Disables Shiki from Astro's markdown pipeline. Without this, Shiki pulls ~5.8 MiB of
255
+ language grammar files into the SSR worker on every build. **Do not re-enable.**
256
+
257
+ ### Rule 2 — `ssr.external` for heavy packages
258
+
259
+ ```js
260
+ ssr: {
261
+ external: ["node:async_hooks", "shiki", "@shikijs/core", "@shikijs/types", /* … */]
262
+ }
263
+ ```
264
+
265
+ **`one.ie/web/astro.config.mjs` § `vite.ssr.external` is the authority for this
266
+ list — never reconcile that file to this doc.** It carries 17 entries as of
267
+ 2026-08-02 (`cookie`, `cytoscape`, `@xyflow/react`, `recharts`, `@stripe/*`,
268
+ `media-chrome`, `motion`, `@100mslive/*`, … alongside the shiki trio). Deleting
269
+ entries to match a stale snippet here would blow the bundle. A companion
270
+ `build.rollupOptions.external` also drops anything matching `shiki/` or
271
+ `@shikijs/` by prefix.
272
+
273
+ The CF adapter bundles everything by default. `ssr.external` creates a bare
274
+ `import { x } from 'pkg'` reference without inlining the package.
275
+
276
+ **Critical nuance:** `ssr.external` only works safely when the externalized package is
277
+ never executed on the server path. For `shiki`: `codeToHtml` is imported by `code-block.tsx`,
278
+ but all components that use `code-block.tsx` are `client:only` — so `codeToHtml` is never
279
+ called in the worker. The import statement exists in the bundle but is dead code.
280
+
281
+ If you add a new heavy dependency used only client-side, add it here.
282
+
283
+ ### Rule 3 — Pure-shell pages use `client:only` + `prerender = true`
284
+
285
+ ```astro
286
+ ---
287
+ export const prerender = true
288
+ import { MyComponent } from "@/components/MyComponent"
289
+ ---
290
+ <Layout title="...">
291
+ <MyComponent client:only="react" />
292
+ </Layout>
293
+ ```
294
+
295
+ `client:only="react"` — Astro renders an empty div on the server; the component never
296
+ runs in the worker. The React component tree (+ all its imports) stays out of the SSR bundle.
297
+
298
+ `export const prerender = true` — the page becomes a static asset generated once
299
+ at build time. The page's SSR handler collapses to a small stub. Zero worker cost
300
+ at runtime.
301
+
302
+ **Use this pattern for:** any page that has no server-side data dependencies
303
+ (no `Astro.locals`, no `Astro.request`, no DB queries in frontmatter).
304
+
305
+ The prerendered set changes every cycle — never hardcode it. Read it from the
306
+ tree (34 pages as of 2026-08-02):
307
+
308
+ ```bash
309
+ cd one.ie/web && grep -l "export const prerender = true" src/pages/*.astro
310
+ ```
311
+
312
+ Pages that CANNOT be prerendered: any page whose frontmatter reads
313
+ `Astro.locals` (session/workspace context), `Astro.request`, or queries the DB.
314
+ Those must stay SSR.
315
+
316
+ ### Rule 4 — `inlineStylesheets: 'auto'` in `astro.config.mjs`
317
+
318
+ ```js
319
+ build: { inlineStylesheets: 'auto' }, // ← NEVER 'always'
320
+ ```
321
+
322
+ With `'always'`, Astro inlines the full Tailwind stylesheet into **every route's
323
+ serialized manifest entry**. With ~100 routes the entry chunk balloons by 8+ MiB
324
+ of duplicated CSS as a single string literal — diagnosable in worker-entry at
325
+ the line `const _manifest = deserializeManifest({...})`.
326
+
327
+ `'auto'` ships the bundle as one external `<link rel="stylesheet">` referenced
328
+ once across all routes. Browsers cache it across navigations — a page-speed
329
+ win, not just a worker-size win.
330
+
331
+ **Verified 2026-05-22:** flipping `always` → `auto` dropped worker-entry from
332
+ 9.5 MiB → 672 KiB and total gzip from 3302 KiB → 2079 KiB.
333
+
334
+ ### Rule 5 — `react-dom/server.edge` alias (production only)
335
+
336
+ ```js
337
+ resolve: {
338
+ alias: {
339
+ ...(isDev ? {} : { "react-dom/server": "react-dom/server.edge" })
340
+ }
341
+ }
342
+ ```
343
+
344
+ Already in `astro.config.mjs`. Required for CF Edge runtime compatibility.
345
+ Do not remove for production builds.
346
+
347
+ ---
348
+
349
+ ## Verified Bundle Numbers
350
+
351
+ | Snapshot | Total upload | gzip | Worker-entry | What changed |
352
+ |---|---|---|---|---|
353
+ | 2026-04-18 (post Pages→Workers migration) | — | — | 9.5 MiB | Rules 1-3 + 5 |
354
+ | 2026-05-22 before `inlineStylesheets` fix | 18.5 MiB | 3.3 MiB | 9.5 MiB | Over 3 MiB ceiling — deploy FAILED |
355
+ | 2026-05-22 after Rule 4 (`'always'` → `'auto'`) | 10.1 MiB | **2.1 MiB** | **672 KiB** | Under ceiling — deploy ✓ |
356
+ | 2026-07-08 | 14.7 MiB | **3.22 MiB** | — | **Exceeds the documented 3 MiB (3072 KiB) free-tier ceiling and still deployed successfully.** Either this account is on a paid Workers plan (10 MiB ceiling) rather than free tier, or the ceiling figure elsewhere in this doc is stale — unconfirmed which. Don't treat "under 3 MiB" as a hard gate until this is resolved; treat 3.2 MiB as the new floor to watch, and re-run Bundle Size Diagnosis if growth continues. |
357
+ | 2026-07-19 | 20.2 MiB | **4.43 MiB** | — | +37% over 2026-07-08's 3.22 MiB. Deployed successfully — no diagnosis run yet. Growth window covers several merged features that day (newsletter platform, directory-submission, social-formats, movers-playbook, etc.) landing in one `/deploy` cycle; not isolated to a single change. Re-run Bundle Size Diagnosis if the next snapshot keeps climbing. |
358
+ | 2026-07-29 | 20.9 MiB | **4.60 MiB** | 1.2 MiB | +6% over 2026-07-19, third consecutive climb. Cheap diagnosis WAS run this cycle (the two grep/`ls` commands below, not a full audit): **no single runaway** — Shiki hits 0, top chunks are `_astro_data-layer-content` 2.1 MiB (content collections), `worker-entry` 1.2 MiB (up from 672 KiB at the 2026-05-22 baseline), `index` 1.3 MiB, `icons` 0.8 MiB, `mermaid` 0.8 MiB, `stripe.esm.worker` 0.6 MiB, `react-vendor` 0.5 MiB. Growth is diffuse feature accretion, not a regression; this deploy's own diff was test-only. Two named candidates if a real audit is ever warranted: `mermaid` (0.8 MiB in the SSR bundle — Rule 2 `ssr.external` candidate if it's only reached from `client:only` islands) and the 15 chunks referencing `react-vendor` (Rule 3 suggests some page still SSR-renders React via `client:load`). |
359
+ | 2026-08-02 | 22.0 MiB | **4.96 MiB** | 1.2 MiB | +8% over 2026-07-29, **fourth consecutive climb**. Cheap diagnosis run again: still **no single runaway** — Shiki 0, `react-vendor` referenced by 15 chunks (unchanged), `worker-entry` flat at 1.2 MiB. Top chunks: `_astro_data-layer-content` 2.1 MiB, `index` 1.3 MiB, **`_broadcast_` 1.3 MiB (new to the top list)**, `worker-entry` 1.2 MiB, `index` 0.9 MiB, `icons` 0.8 MiB, `mermaid` 0.8 MiB, `stripe.esm.worker` 0.6 MiB, `react-vendor` 0.5 MiB, **`generateAuthenticationOptions` 0.5 MiB (new)**. This cycle shipped beautiful-blocks (45 visual blocks + 12 background components), which is a plausible share of the delta. Four climbs in a row with the same "diffuse accretion" verdict each time is itself the signal — the cheap diagnosis has now exhausted what it can tell us, and the two standing candidates (`mermaid` via Rule 2, the 15 `react-vendor` chunks via Rule 3) want a real audit rather than a fifth restatement. |
360
+
361
+ The 2026-05-22 regression was caused by `build: { inlineStylesheets: 'always' }`
362
+ inlining the full Tailwind stylesheet into every route's manifest entry. One
363
+ char change (`always` → `auto`) saved 8.8 MiB.
364
+
365
+ ---
366
+
367
+ ## Service Map (post-migration)
368
+
369
+ `./deploy <mode>` covers every live row; the per-service command is what the
370
+ script runs, recorded here for rollback and one-off work.
371
+
372
+ | Service | URL | Config | Deploy command |
373
+ |---------|-----|--------|---------------|
374
+ | Astro Worker (prod) | `one.ie` → `one-prod` | `one.ie/web/wrangler.toml` | `cd one.ie/web && wrangler deploy` — **no `--env production`** (deploy-target trap: appends `-production` to the script name → `one-prod-production`, which nothing routes to; see trap note at top of this file) |
375
+ | Gateway | `api.one.ie` → `one-gateway` | `api/wrangler.toml` | `cd api && wrangler deploy` |
376
+ | Sync | `one-sync` — **cron-only, no HTTP route** (health = deploy success) | `sync/wrangler.toml` | `cd sync && wrangler deploy` |
377
+ | Agents | `channels.one.ie` → `channels` (`*.workers.dev` blocked on this network) | `channels/wrangler.toml` | `cd channels && wrangler deploy` |
378
+ | Pay gateway | `pay.one.ie` → `one-core-worker` | `pay/backend/wrangler.toml` | `cd pay/backend && bun run deploy` (= bare `wrangler deploy`, no `--env` flag) |
379
+ | Pages (legacy idle, rollback) | `oneie.pages.dev` | — | **do not deploy** — rollback target for `one.ie` |
380
+ | Worker (legacy idle, rollback) | `one-demo` (still serves `demo.one.ie`, `onestudio.dev`) | — | **do not deploy** — rollback window for the prod cutover |
381
+
382
+ ---
383
+
384
+ ## Auth (CRITICAL — never change)
385
+
386
+ Never: `CLOUDFLARE_API_TOKEN` (scoped token lacks workers + custom domain permissions).
387
+
388
+ **wrangler reads `CLOUDFLARE_API_KEY` + `CLOUDFLARE_EMAIL`** for global-key auth
389
+ — `CLOUDFLARE_GLOBAL_API_KEY` is *our* name for it and wrangler ignores it. The
390
+ script now exports both spellings off one resolved value, so the two can never
391
+ diverge again.
392
+
393
+ **The credential lives on disk, not in your shell** — `.env.local` at the repo
394
+ root and `one.ie/web/.env` (same 52-char key). You do not need to export
395
+ anything. If you *do* export one, it wins — which is the trap: a **stale**
396
+ export shadows the good key everywhere, because process env beats wrangler's
397
+ per-directory `.env` autoload. That is exactly how 2026-08-19's deploy used two
398
+ different credentials in two consecutive steps (see the Step 6.6 note below).
399
+ The ladder now probes each rung and falls through a rung that does not answer,
400
+ so a stale export costs a log line instead of a red deploy:
401
+
402
+ ```
403
+ rejected: ambient env (len=37 sha=eaafbda9) — /user did not answer 200
404
+ ✓ resolved: global-api-key
405
+ source: /Users/toc/Server/one-ie/.env.local (len=52 sha=435cba23)
406
+ ✓ 5/5 services agree on account 627e0c7c…
407
+ ```
408
+
409
+ Ground truth is the API, in **both** directions — it is as able to prove a key
410
+ alive as dead. Never conclude either from wrangler alone:
411
+
412
+ ```bash
413
+ curl -s -o /dev/null -w '%{http_code}\n' \
414
+ -H "X-Auth-Email: $CLOUDFLARE_EMAIL" -H "X-Auth-Key: $CLOUDFLARE_API_KEY" \
415
+ https://api.cloudflare.com/client/v4/user # 200 = fine
416
+ ```
417
+
418
+ The deploy script auto-unsets `CLOUDFLARE_API_TOKEN` from the spawned env to prevent
419
+ accidental use of a scoped token that was exported in the shell.
420
+
421
+ Required env (export locally before running `./deploy` — there is no CI; the
422
+ script asserts the first two at gate 4 and refuses to ship without them):
423
+ - `CLOUDFLARE_GLOBAL_API_KEY` + `CLOUDFLARE_EMAIL` — auth
424
+ - `PUBLIC_GATEWAY_URL: https://api.one.ie` — build-time-inlined by Astro (**required**; without it the Worker bundle falls back to `one-gateway.oneie.workers.dev` and gateway-backed routes break). Lives in `one.ie/web/.env`.
425
+
426
+ ---
427
+
428
+ ## Mode-specific notes
429
+
430
+ Everything common lives in the script. These are the bits that are true of one
431
+ mode only.
432
+
433
+ ### `./deploy astro`
434
+
435
+ **Resolved 2026-07-04:** `one.ie/web/package.json`'s `"deploy"` script had the
436
+ same `--env production` trap. Confirmed via the CF API that it was real —
437
+ neither `one-prod` (live) nor the `one-prod-production` decoy had any cron
438
+ schedules registered, meaning `billing-alerts-cron.ts` /
439
+ `billing-allocation-cron.ts` / `billing-autotopup-cron.ts` /
440
+ `billing-lifecycle-cron.ts` / `billing-verify-cron.ts` /
441
+ `funnel-aggregate-cron.ts` / `webhook-deliver.ts` / `broadcast-drain-cron.ts`
442
+ were not running on any schedule. Fixed by reconciling `wrangler.toml`: the
443
+ `[[routes]]` (custom domain) and `[triggers]` (crons) blocks — the only two
444
+ things that existed *only* under `[env.production]` — were moved to top level
445
+ (everything else was already duplicated there); the now-fully-redundant
446
+ `[env.production.*]` block was deleted entirely, and `package.json`'s script
447
+ dropped `--env production`. The trap is now structurally impossible — there's
448
+ no `--env production` target left to hit.
449
+
450
+ **Pre-flight (one-time, on cutover only):** ensure no other CF entity owns the
451
+ `one.ie` custom domain. If wrangler errors with a hostname conflict, detach the
452
+ prior owner first:
453
+
454
+ ```bash
455
+ # If a Pages project owns it (was `oneie` project pre-cutover):
456
+ curl -s -X DELETE \
457
+ "https://api.cloudflare.com/client/v4/accounts/$CF_ACCOUNT_ID/pages/projects/oneie/domains/one.ie" \
458
+ -H "X-Auth-Email: $CLOUDFLARE_EMAIL" \
459
+ -H "X-Auth-Key: $CLOUDFLARE_GLOBAL_API_KEY"
460
+ ```
461
+
462
+ ### `./deploy workers`
463
+
464
+ Verified 2026-07-08: gateway, sync and channels complete independently with no
465
+ shared state — which is why the script runs them concurrently.
466
+
467
+ ### `./deploy pay`
468
+
469
+ `pay/backend/src/index.ts` mounts `src/routes/status.ts` (`/status`) and
470
+ `discoveryRoutes` (`/`). It does **not** serve `/health` — the `/health`
471
+ handler in `pay/backend/src/api/routes/status.ts` belongs to the separate,
472
+ unmounted `src/api/` tree. Both `/` and `/status` probed live 2026-08-02: 200,
473
+ unauthenticated. `/status` also reports version, destinationMode, and contract
474
+ addresses.
475
+
476
+ **Found 2026-07-05:** `pay/backend` shipped a real commit (`feat(pay):
477
+ embeddable payment-link page`) that sat unshipped through a full deploy cycle
478
+ because the service map only named 4 services. `pay.one.ie` is a first-class
479
+ 5th target, not an afterthought — check `git log` scoped to `pay/` for
480
+ unshipped commits, same as the other four.
481
+
482
+ ## Bundle Size Diagnosis
483
+
484
+ If build fails with "exceeds size limit":
485
+
486
+ ```bash
487
+ # Check total worker size
488
+ du -sh dist/server/
489
+
490
+ # Find top offenders
491
+ ls -lhS dist/server/chunks/ | head -20
492
+
493
+ # Check if a new import pulled in Shiki
494
+ grep -r "from 'shiki'" dist/server/chunks/ | wc -l
495
+ # If > 0: a component that imports shiki was SSR'd
496
+ # Fix: make its page client:only="react" + prerender=true
497
+
498
+ # Check if React crept back into worker via client:load
499
+ grep -l "react-vendor" dist/server/chunks/
500
+ # If multiple chunks: some page SSR-renders React via client:load
501
+ # Fix: audit src/pages/*.astro for client:load on pure-shell pages
502
+ ```
503
+
504
+ ---
505
+
506
+ ## The TypeDB flake waiver — a red suite the deploy may ship past
507
+
508
+ `one.ie/web`'s suite talks to a REAL shared TypeDB Cloud cluster (CLAUDE.md:
509
+ "Don't mock TypeDB in integration tests"). When that cluster blips or a query
510
+ outruns its timeout, a handful of task/substrate suites go red without anything
511
+ in the diff being wrong. That used to be an eyeball judgement, which is exactly
512
+ the call that gets rubber-stamped on the fifth deploy attempt at 2am.
513
+
514
+ `.claude/scripts/typedb-flake-check.sh` makes it a check. When the vitest gate
515
+ goes red, `deploy.sh` runs it against the gate log:
516
+
517
+ | exit | means |
518
+ |---|---|
519
+ | 0 | every failure carries a substrate-unavailable signature — waivable, deploy continues |
520
+ | 1 | at least one failure is real, or the log could not be classified — deploy stops |
521
+
522
+ **It keys on the failure SIGNATURE, never on the filename.** A file-based
523
+ allowlist waives every future failure in that file, including the real ones. The
524
+ waived signatures are `upstream_50[234]`, `status=50[234]`,
525
+ `fixture write failed`, `Test timed out in Nms`, `ETIMEDOUT`, `ECONNRESET`,
526
+ `ECONNREFUSED`, `EAI_AGAIN`, `socket hang up`, `fetch failed`, and TypeDB
527
+ connection errors.
528
+
529
+ Four properties, each with a red half in `--self-test`:
530
+
531
+ - **`not_found` is never waivable.** Checked first, independently of everything
532
+ else. It is the `tasks:claim` privilege boundary and four separate real
533
+ defects have presented as that exact string. A `not_found` wrapped inside a
534
+ 503 still blocks.
535
+ - **One flake never vouches for its neighbour.** The log is split into vitest's
536
+ per-failure blocks and EVERY block must carry a signature. A genuine assertion
537
+ break standing beside a 503 blocks the deploy.
538
+ - **Silence is not a pass.** An empty or unparseable log, or one with no
539
+ `Tests N failed` line, exits 1.
540
+ - **A waiver is not a green suite.** The report says
541
+ `WAIVED as TypeDB outage (suite NOT green)`, and a waived run **cannot settle
542
+ deferred-pin debt** — the waiver speaks only to the failures that reported, not
543
+ to a pin that never got to.
544
+
545
+ On by default. `--no-typedb-flake-waiver` (or `DEPLOY_ALLOW_TYPEDB_FLAKE=0`)
546
+ restores the hard stop. Prove the checker still bites before trusting it:
547
+
548
+ ```bash
549
+ bash .claude/scripts/typedb-flake-check.sh --self-test # 6 cases, 4 of them red halves
550
+ bash .claude/scripts/typedb-flake-check.sh <a-gate-log>
551
+ ```
552
+
553
+ **The waiver is not a diagnosis.** `curl -sS -o /dev/null -w '%{http_code}' https://api.one.ie/health`
554
+ returning 200 while the suite reports 503s means the cluster blipped mid-run. A
555
+ 200 alongside failures that are NOT in the signature list means the code is
556
+ wrong — and the checker will tell you so.
557
+
558
+ ## Known-Flaky Test Allowlist
559
+
560
+ `deploy.sh` deliberately enforces no allowlist — a red suite stops it, and it
561
+ prints the failing files with a pointer here. Triage is yours: treat these by
562
+ name when they appear in `bunx vitest run` output in `one.ie/web`; don't block deploy on them, but don't silently ignore new failures either
563
+ — confirm the failure signature matches before waving it through:
564
+
565
+ - `tests/e2e/c5-webhook-subscribe.test.ts` — `workflow:webhook-subscribe` "writes a KV
566
+ record…" and "fails closed on a non-owned workflow". **Root cause (confirmed 2026-07-08):**
567
+ `ssrfGuard()` (`one.ie/web/src/lib/ssrf.ts`) resolves DNS via Cloudflare DoH by fetching
568
+ `https://1.1.1.1/dns-query` directly; this local dev network refuses connections to
569
+ `1.1.1.1:443` (`curl: (7) Failed to connect`), so `resolveHostIPs` returns `[]` and every
570
+ URL — including the test's `https://example.com` — comes back `blocked_url`. Verify before
571
+ waving through: `curl -v --max-time 5 https://1.1.1.1/dns-query 2>&1 | grep -i refused`
572
+ — if that shows "Connection refused", it's this network gap, not a code regression. If it
573
+ connects fine and the test still fails, it's real — investigate.
574
+ - `tests/unit/tasks-humans.test.ts` + `tests/tasks-do-roundtrip.test.ts` — **only** when
575
+ the failure is `fixture write failed (status=503 error=upstream_503)` or
576
+ `(status=502|504 …)`. That message means the shared TypeDB Cloud cluster was
577
+ unavailable through every retry `typedbQueryDetail` already performs — the substrate
578
+ refused setup, so nothing downstream proved anything. Verify before waving through:
579
+ `curl -sS -o /dev/null -w '%{http_code}' https://api.one.ie/health` — a 200 there with a
580
+ 503 in the test means the cluster blipped during the run, not that the code is wrong.
581
+ **Any OTHER failure in these two files is real and blocks deploy** — in particular a bare
582
+ `not_found` on `tasks:claim`, which is the privilege boundary and must never be waved
583
+ through. Four separate defects that used to present as that same `not_found` were fixed
584
+ 2026-08-02 (concurrent-run sweep destruction, silent fixture writes, same-attribute
585
+ insert races, read-after-write lag); if it reappears, something new is wrong. History:
586
+ the commit message on `140975e44` and `tests/helpers/probe-sweep.ts`.
587
+ - Hardware/stochastic benchmarks (speed, distribution-timing tests) — expected variance,
588
+ not correctness bugs.
589
+
590
+ Any other failure (type errors, assertion mismatches on business logic) blocks deploy —
591
+ diagnose and fix before proceeding.
592
+
593
+ ---
594
+
595
+ ## First-Time Setup
596
+
597
+ Only needed once, before `./deploy` can work at all. There is no
598
+ `docs/deploy.md` — this block plus the script is the whole walkthrough.
599
+ Resource names below are the ones
600
+ actually declared in `one.ie/web/wrangler.toml`; creating differently-named
601
+ resources produces bindings the Worker can't resolve.
602
+
603
+ ```bash
604
+ # Create CF resources — names must match wrangler.toml exactly
605
+ bunx wrangler d1 create one-owners # → binding DB
606
+ bunx wrangler kv namespace create SESSION # → binding SESSION
607
+ bunx wrangler kv namespace create CHAT_CACHE # → binding CHAT_CACHE
608
+ bunx wrangler kv namespace create THREADS # → binding THREADS
609
+ bunx wrangler r2 bucket create one-content # → binding CONTENT
610
+ # → Paste IDs into one.ie/web/wrangler.toml + sync/wrangler.toml
611
+
612
+ # Run D1 migrations (the ledger starts at 0001_owners.sql — there is no 0001_init.sql)
613
+ cd one.ie/web && bunx wrangler d1 migrations apply DB --remote
614
+
615
+ # Gateway secrets (TypeDB credentials) — no --env flag, ever
616
+ cd api
617
+ printf 'admin' | bunx wrangler secret put TYPEDB_USERNAME
618
+ bunx wrangler secret put TYPEDB_PASSWORD # paste at the prompt; never inline the value
619
+ cd ..
620
+
621
+ # First deploy — Worker auto-provisions on first `wrangler deploy`
622
+ # Then just: ./deploy
623
+ ```
624
+
625
+ ---
626
+
627
+ ## Logs
628
+
629
+ `./deploy` writes `.deploy-logs/deploy-<stamp>.log` (gitignored) — every gate's
630
+ stdout in order — plus `.deploy-logs/<service>-<stamp>.log` for each service in
631
+ the parallel wave, since their output would otherwise interleave. The report at
632
+ the end names the log path.
633
+
634
+ Live logs:
635
+
636
+ ```bash
637
+ cd one.ie/web && bunx wrangler tail --name one-prod # Astro Worker (production)
638
+ cd api && bunx wrangler tail # Gateway
639
+ cd one.ie/web && bunx wrangler deployments list --name one-prod | head -10
640
+ ```
641
+
642
+ ---
643
+
644
+ ## Gotchas
645
+
646
+ - **A service with no local `wrangler` falls through to a shared `bunx wrangler@latest` cache, and that cache can rot.** Hit 2026-08-06: `sync/` was the only one of the five without wrangler in its `devDependencies`, so `bunx wrangler deploy` resolved to `$TMPDIR/bunx-501-wrangler@latest/` — whose install was missing `esbuild`, so it died `MODULE_NOT_FOUND` before reading a single config. The other four were unaffected because they resolve wrangler from their own `node_modules`. Fixed by pinning `wrangler` into `sync/package.json` like its siblings. If this shape reappears elsewhere, the workaround is a version-pinned invocation (`bunx wrangler@4.80.0 deploy`, which lands in a different cache dir); the fix is a local dep.
647
+ - **Piping `./deploy` into `tail`/`head` masks its exit code** — the pipeline reports the pager's status, not the script's, so a failed run reads as exit 0. `die()` really does `exit 1`; read the ✓/✗ lines or the `.deploy-logs/` file, and don't infer success from a piped exit status.
648
+ - TypeDB Cloud port is **1729** (not 80 or 443)
649
+ - TypeDB HTTP API prefix is `/v1/` (signin, query, databases)
650
+ - Always `CLOUDFLARE_GLOBAL_API_KEY` — scoped tokens lack permissions for workers + custom domains
651
+ - `import.meta.env` is build-time — `PUBLIC_GATEWAY_URL` is baked into the worker bundle at build. Missing → gateway-backed routes fall back to the wrong host and break
652
+ - Custom domains: `[[routes]]` double bracket, no wildcards, add `workers_dev = true`
653
+ - Worker upload limit: **3 MiB gzipped** on free tier (10 MiB on paid) — though a 2026-07-08 deploy shipped at 3.22 MiB gzip successfully, so this account's actual ceiling is unconfirmed (see Verified Bundle Numbers). Wrangler reports `gzip:` — only that number counts. Follow the 5 Bundle Size Rules above regardless of which ceiling applies
654
+ - **D1 schema-drift fails at runtime, not compile-time.** Migrations that DROP+CREATE a table (e.g. `0059_domains.sql` renamed `slug`→`gid`, `verified`→`verified_at`) silently break any code that queries the old columns — typecheck passes, deploy succeeds, the route 500s in production. After any DROP+CREATE migration, grep the codebase for the old column names and fix call sites BEFORE deploying
655
+ - **Error responses get the same `cache-control` as success responses.** Astro's Layout sets `public, max-age=300, s-maxage=86400, stale-while-revalidate=604800` on every render including 5xx pages. A bad deploy will be cached at the CF edge for 24h. When diagnosing, always bust the cache: `curl "https://host/path?_t=$(date +%s)"`. Consider a middleware rule that strips `cache-control` on `>= 500` status
656
+
657
+ ---
658
+
659
+ *Deploy is the closed loop. W0 baseline in, health check out. If health fails, mark() is blocked. Determinism: every step reports numbers, every number gets marked.*
660
+
661
+ ### `code: 9103` does not mean the key was revoked
662
+
663
+ Measured 2026-08-18, and it cost an hour plus a wrong accusation that a security
664
+ incident had rotated the key. `Unknown X-Auth-Key or X-Auth-Email [code: 9103]`
665
+ at Step 6.5 was a **variable-name mismatch**: the script exported
666
+ `CLOUDFLARE_GLOBAL_API_KEY`, wrangler only reads `CLOUDFLARE_API_KEY`.
667
+
668
+ Before concluding a Cloudflare credential is dead, ask the API directly — it is
669
+ ground truth and wrangler is not:
670
+
671
+ ```bash
672
+ curl -s -o /dev/null -w '%{http_code}\n' \
673
+ -H "X-Auth-Email: $CLOUDFLARE_EMAIL" -H "X-Auth-Key: $CLOUDFLARE_API_KEY" \
674
+ https://api.cloudflare.com/client/v4/user # 200 = the key is fine
675
+ ```
676
+
677
+ Three companion traps, all real:
678
+
679
+ - **wrangler 4.x auto-loads `.env` from the cwd**, so a stale key in
680
+ `one.ie/web/.env` beats both your exported vars and an OAuth session. Isolate a
681
+ credential test by running it from `/tmp`.
682
+ - **`/user/tokens/verify` is Bearer-only** — `400 Missing "Authorization" header`
683
+ there is not evidence against a global key. Use `/user` or `/accounts`.
684
+ - **Length proves nothing.** A classic global key is 37 hex chars; a `cfk_`-prefixed
685
+ one is ~52 and equally valid.
686
+
687
+ **Never edit `.env` while a deploy is running.** Doing so took a run fully red —
688
+ tsc, vitest, and the astro build — for reasons that had nothing to do with the code.
689
+
690
+ ### The 2026-08-19 sequel: `code: 7403` at Step 6.6, and two keys
691
+
692
+ The same family, one layer deeper, and worth reading before you diagnose any
693
+ Cloudflare auth failure here. Step 6.5 (D1 in `one.ie/web`) **passed** and Step
694
+ 6.6 (D1 in `channels`) **failed** in the same run, seconds apart, on the same
695
+ account. That is only possible if they used different credentials — and they
696
+ did:
697
+
698
+ - **Three** credentials were reachable from one `./deploy`: a stale 37-char key
699
+ in the ambient env (injected by `~/.claude/settings.json`'s `env` block — so
700
+ it existed inside Claude Code sessions and *not* in a plain terminal, which is
701
+ why grepping the shell profiles found nothing), the good 52-char key in
702
+ `one.ie/web/.env` + `.env.local`, and the `wrangler login` OAuth session.
703
+ - Nothing *chose* between them. Each service got whatever its own cwd surfaced.
704
+ Of the five dirs, only `one.ie/web` has a `.env` — so it read the good key and
705
+ passed, while `channels` fell through to OAuth and hit `7403`.
706
+ - **`7403` ≠ `9103`.** `9103` ("Unknown X-Auth-Key") points at a key; `7403`
707
+ ("account is not authorized to access this service") points at an account
708
+ scope, i.e. a session. Reading them as the same symptom is what sent the first
709
+ diagnosis at a perfectly good key.
710
+
711
+ Fixed by making the credential *resolved* rather than *ambient* — see Step 0.4
712
+ above. Run `./deploy --check-creds` if you ever doubt which key is in play; it
713
+ names the source and proves all five services agree.