@oneie/claude 0.8.0 → 0.10.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (230) hide show
  1. package/.claude-plugin/plugin.json +1 -1
  2. package/agents/abm-strategist.md +67 -1
  3. package/agents/ads-meta.md +67 -1
  4. package/agents/analyst.md +67 -1
  5. package/agents/animator.md +108 -0
  6. package/agents/architect.md +269 -20
  7. package/agents/brand-guardian.md +67 -1
  8. package/agents/brand-strategist.md +67 -1
  9. package/agents/campaign-content.md +67 -1
  10. package/agents/campaign-email.md +67 -1
  11. package/agents/campaign-sms.md +67 -1
  12. package/agents/campaign-social.md +67 -1
  13. package/agents/cco.md +83 -2
  14. package/agents/ceo.md +108 -11
  15. package/agents/chairman.md +197 -0
  16. package/agents/cmo.md +82 -2
  17. package/agents/community-greeter.md +67 -1
  18. package/agents/community-moderator.md +67 -1
  19. package/agents/compliance.md +67 -1
  20. package/agents/copywriter.md +67 -1
  21. package/agents/creative-strategist.md +67 -1
  22. package/agents/cro.md +81 -1
  23. package/agents/cto.md +266 -28
  24. package/agents/customer-interviewer.md +67 -1
  25. package/agents/customer-researcher.md +67 -1
  26. package/agents/customer-success-manager.md +67 -1
  27. package/agents/customer-trainer.md +67 -1
  28. package/agents/cxo.md +82 -1
  29. package/agents/demand-creator.md +67 -1
  30. package/agents/demo-mover.md +67 -1
  31. package/agents/demo-specialist.md +67 -1
  32. package/agents/demo-thai-family-law.md +67 -1
  33. package/agents/designer.md +67 -1
  34. package/agents/discovery-caller.md +67 -1
  35. package/agents/doctor.md +269 -0
  36. package/agents/educate-coach.md +67 -1
  37. package/agents/elevate-tutor.md +67 -1
  38. package/agents/email-lifecycle-marketer.md +67 -1
  39. package/agents/engage-specialist.md +67 -1
  40. package/agents/events-coordinator.md +67 -1
  41. package/agents/foundation-builder.md +67 -1
  42. package/agents/funnel-architect.md +67 -1
  43. package/agents/gift-creator.md +67 -1
  44. package/agents/google-ads.md +67 -1
  45. package/agents/guide.md +67 -1
  46. package/agents/helpdesk-dispatcher.md +67 -1
  47. package/agents/hook-specialist.md +67 -1
  48. package/agents/identify-optimizer.md +67 -1
  49. package/agents/implementer.md +313 -45
  50. package/agents/incident-commander.md +67 -1
  51. package/agents/insights-lead.md +87 -1
  52. package/agents/journey-runner.md +67 -1
  53. package/agents/linkedin-ads.md +67 -1
  54. package/agents/live-sales-chat.md +67 -1
  55. package/agents/market-researcher.md +67 -1
  56. package/agents/media-buyer.md +67 -1
  57. package/agents/memory-keeper.md +195 -0
  58. package/agents/movers-customer-researcher.md +67 -1
  59. package/agents/movers-foundation-builder.md +67 -1
  60. package/agents/movers-market-researcher.md +67 -1
  61. package/agents/movers-pricing-strategist.md +67 -1
  62. package/agents/nurture-architect.md +67 -1
  63. package/agents/offer-architect.md +67 -1
  64. package/agents/onboarder.md +67 -1
  65. package/agents/onboarding-specialist.md +67 -1
  66. package/agents/operations-dashboard.md +87 -1
  67. package/agents/perf-engineer.md +333 -37
  68. package/agents/playbook-writer.md +67 -1
  69. package/agents/plg-strategist.md +67 -1
  70. package/agents/positioning-architect.md +67 -1
  71. package/agents/press-officer.md +67 -1
  72. package/agents/pricing-strategist.md +67 -1
  73. package/agents/privacy-officer.md +67 -1
  74. package/agents/referral-manager.md +67 -1
  75. package/agents/refine-analyst.md +67 -1
  76. package/agents/release-manager.md +446 -39
  77. package/agents/renewals-upsell-rep.md +67 -1
  78. package/agents/review-engineer.md +319 -45
  79. package/agents/rewards-steward.md +67 -1
  80. package/agents/sales-call-coach.md +67 -1
  81. package/agents/sales-closer.md +67 -1
  82. package/agents/security-auditor.md +343 -48
  83. package/agents/sell-closer.md +67 -1
  84. package/agents/share-amplifier.md +67 -1
  85. package/agents/social-media-manager.md +67 -1
  86. package/agents/storyteller.md +301 -0
  87. package/agents/strategist.md +67 -1
  88. package/agents/strategy-aligner.md +67 -1
  89. package/agents/support-agent.md +67 -1
  90. package/agents/tagger.md +327 -0
  91. package/agents/tech-writer.md +195 -22
  92. package/agents/test-engineer.md +398 -29
  93. package/agents/tiktok-ads.md +67 -1
  94. package/agents/tracking-engineer.md +67 -1
  95. package/agents/trailkeeper.md +181 -0
  96. package/agents/upsell-strategist.md +67 -1
  97. package/agents/voice.md +67 -1
  98. package/agents/w1-recon.md +1 -1
  99. package/agents/w2-decide.md +1 -1
  100. package/agents/w3-edit.md +8 -2
  101. package/agents/w4-verify.md +13 -0
  102. package/agents/workflow-optimiser.md +81 -1
  103. package/commands/close.md +916 -160
  104. package/commands/deploy.md +102 -724
  105. package/commands/do.md +58 -2
  106. package/commands/sweep.md +159 -0
  107. package/commands/tasks.md +222 -0
  108. package/hooks/scripts/dev-only.sh +135 -0
  109. package/hooks/scripts/git-add-guard.sh +37 -2
  110. package/hooks/scripts/session-start.sh +32 -4
  111. package/package.json +1 -1
  112. package/rules/scripts.md +85 -0
  113. package/scripts/CLAUDE.md +315 -0
  114. package/scripts/ad-copy-lint.sh +656 -0
  115. package/scripts/agent-actor-parity.sh +129 -0
  116. package/scripts/blocks-manifest-cached.sh +100 -0
  117. package/scripts/chat-context-check.sh +89 -0
  118. package/scripts/chrome.mjs +18 -0
  119. package/scripts/close-metrics.sh +587 -0
  120. package/scripts/close-owner.sh +326 -0
  121. package/scripts/db-sync-lock-check.sh +116 -0
  122. package/scripts/deploy-emit.sh +311 -0
  123. package/scripts/deploy-gate-check.sh +155 -0
  124. package/scripts/deploy-ready.sh +78 -0
  125. package/scripts/deploy-record.sh +605 -0
  126. package/scripts/deploy-schema-check.sh +58 -0
  127. package/scripts/deploy.sh +393 -243
  128. package/scripts/do-auto.sh +127 -26
  129. package/scripts/do-board.sh +429 -0
  130. package/scripts/do-close.sh +1184 -0
  131. package/scripts/do-consumer-sweep.sh +18 -1
  132. package/scripts/do-decide.sh +476 -0
  133. package/scripts/do-fleet.sh +8 -2
  134. package/scripts/do-plan-json.mjs +110 -12
  135. package/scripts/do-prove-selftest.sh +108 -0
  136. package/scripts/do-prove.sh +86 -10
  137. package/scripts/do-rank.py +200 -3
  138. package/scripts/do-reconcile.sh +73 -12
  139. package/scripts/do-signal.sh +101 -23
  140. package/scripts/do-smoke.sh +18 -1
  141. package/scripts/do-w4-gates.sh +11 -1
  142. package/scripts/do-world-check.sh +153 -0
  143. package/scripts/download-stats.sh +172 -0
  144. package/scripts/factory-brief-check.sh +330 -0
  145. package/scripts/factory-check.sh +18 -1
  146. package/scripts/factory-close-check.sh +257 -0
  147. package/scripts/factory-emit.sh +211 -0
  148. package/scripts/factory-executor-check.mjs +353 -0
  149. package/scripts/factory-peak.sh +301 -0
  150. package/scripts/factory-repo.sh +71 -0
  151. package/scripts/factory-review-check.mjs +61 -0
  152. package/scripts/factory-tasks-check.sh +18 -1
  153. package/scripts/fixtures/factory-brief-real.md +44 -0
  154. package/scripts/flywheel-outcome.sh +63 -0
  155. package/scripts/gate-reaper-check.sh +98 -0
  156. package/scripts/gate-reaper.sh +9 -0
  157. package/scripts/gate-watchdog.sh +619 -0
  158. package/scripts/gc-content-check.sh +142 -0
  159. package/scripts/gh-traffic-capture.sh +153 -0
  160. package/scripts/govern-order-check.sh +202 -0
  161. package/scripts/governor-doors-check.sh +86 -5
  162. package/scripts/health.sh +448 -0
  163. package/scripts/id-inventory.mjs +418 -0
  164. package/scripts/incident.sh +212 -0
  165. package/scripts/land.sh +755 -45
  166. package/scripts/lib/gc-finished.sh +77 -0
  167. package/scripts/livekit-ratchet.sh +18 -1
  168. package/scripts/machine-check.sh +1 -1
  169. package/scripts/memory-index-budget.sh +79 -0
  170. package/scripts/npm-downloads.sh +109 -0
  171. package/scripts/one-agents.mjs +204 -8
  172. package/scripts/one-resume.sh +31 -3
  173. package/scripts/pr-body.sh +335 -0
  174. package/scripts/preview-fd-check.sh +289 -0
  175. package/scripts/redirect-lint.sh +169 -0
  176. package/scripts/release.sh +40 -6
  177. package/scripts/resume-lost-sessions.sh +68 -0
  178. package/scripts/shoot-pages.mjs +140 -0
  179. package/scripts/signal-meta-backfill.ts +451 -0
  180. package/scripts/signal-watch.sh +63 -6
  181. package/scripts/speed-cache-check.sh +12 -2
  182. package/scripts/sweep.sh +426 -0
  183. package/scripts/task-titles-dump.ts +101 -0
  184. package/scripts/test-cached.sh +47 -10
  185. package/scripts/test-lanes.sh +14 -0
  186. package/scripts/thread-name-backfill.ts +215 -0
  187. package/scripts/triage-shape-check.sh +149 -0
  188. package/scripts/tsc-cached.sh +155 -8
  189. package/scripts/typedb-flake-check.sh +3 -1
  190. package/scripts/urls-lint.sh +8 -0
  191. package/scripts/verify-board-doors.sh +80 -0
  192. package/scripts/verify-fast.sh +159 -6
  193. package/scripts/worktree-up.sh +21 -3
  194. package/skills/astro/SKILL.md +9 -3
  195. package/skills/astro/optimize-performance.md +3 -2
  196. package/skills/cloudflare/SKILL.md +3 -2
  197. package/skills/cloudflare-security-audit/AI-AND-LLM.md +83 -0
  198. package/skills/cloudflare-security-audit/ATTACK-CLASSES.md +130 -0
  199. package/skills/cloudflare-security-audit/CLIENT-SIDE.md +83 -0
  200. package/skills/cloudflare-security-audit/CLOUD-AND-DEPLOYMENT.md +86 -0
  201. package/skills/cloudflare-security-audit/DATA-ISOLATION-AND-LIFECYCLE.md +84 -0
  202. package/skills/cloudflare-security-audit/DESKTOP-MOBILE-AND-LOCAL-IPC.md +89 -0
  203. package/skills/cloudflare-security-audit/HUNTING.md +251 -0
  204. package/skills/cloudflare-security-audit/LICENSE +21 -0
  205. package/skills/cloudflare-security-audit/MEMORY-SAFETY-AND-BINARY.md +101 -0
  206. package/skills/cloudflare-security-audit/PROTOCOLS-RPC-AND-MESSAGING.md +81 -0
  207. package/skills/cloudflare-security-audit/PROVENANCE.md +78 -0
  208. package/skills/cloudflare-security-audit/RECONNAISSANCE.md +156 -0
  209. package/skills/cloudflare-security-audit/RESOURCE-EXHAUSTION-AND-AVAILABILITY.md +78 -0
  210. package/skills/cloudflare-security-audit/SKILL.md +192 -0
  211. package/skills/cloudflare-security-audit/SUPPLY-CHAIN-AND-RELEASE.md +73 -0
  212. package/skills/cloudflare-security-audit/VALIDATION-AND-REPORTING.md +186 -0
  213. package/skills/cloudflare-security-audit/WEB-PROTOCOL-AND-AUTH.md +105 -0
  214. package/skills/cloudflare-security-audit/report-schema.json +461 -0
  215. package/skills/cloudflare-security-audit/validate-coverage-ledger.cjs +872 -0
  216. package/skills/cloudflare-security-audit/validate-coverage-ledger.test.cjs +740 -0
  217. package/skills/cloudflare-security-audit/validate-findings.cjs +773 -0
  218. package/skills/cloudflare-security-audit/validate-findings.test.cjs +652 -0
  219. package/skills/deploy/REFERENCE.md +713 -0
  220. package/skills/deploy/SKILL.md +140 -0
  221. package/skills/fleet-audit/SKILL.md +58 -0
  222. package/skills/meeting/SKILL.md +220 -0
  223. package/skills/planning/SKILL.md +256 -0
  224. package/skills/shadcn/SKILL.md +1 -1
  225. package/skills/typedb/SKILL.md +7 -0
  226. package/skills/voice/SKILL.md +94 -6
  227. package/skills/voice/corpus-check.sh +87 -0
  228. package/templates/template-agent.md +7 -1
  229. package/templates/template-feature.md +9 -0
  230. package/templates/template-todo.md +29 -0
@@ -1,6 +1,47 @@
1
1
  # /deploy
2
2
 
3
- **Skills:** `/cloudflare` (Workers auth) · `/signal` (deploy:success / deploy:degraded)
3
+ > **Every deploy spawns `release-manager`, and it always reports to the CEO.**
4
+ > Not just the slash command — `/deploy`, `./deploy`, `bash .claude/scripts/deploy.sh`,
5
+ > `release.sh promote`, `release.sh ship`, and any plain-English "ship it" / "deploy the
6
+ > site" all take this door. The conductor does not run the doors itself. First action:
7
+ >
8
+ > ```
9
+ > Agent({ subagent_type: "release-manager", model: "opus",
10
+ > prompt: "Read one.ie/ai/agents/release-manager/agent.md and .claude/skills/deploy/REFERENCE.md, then run /deploy <args>. Target: <sha | PR | origin/main>. Authorised to ship: <yes | gates-only>." })
11
+ > ```
12
+ >
13
+ > Pass `model: "opus"` (the generated roster maps a specialist to Sonnet). Tell it
14
+ > to read the source prompt — a session started before the last `one-agents.mjs`
15
+ > run holds the older roster copy. The agent opens and closes with a post to the
16
+ > CEO in `/u/one/in` (group `release`) — see its § Report to the CEO. The
17
+ > conductor relays the agent's closing report to the operator; it never re-runs
18
+ > the gates to double-check a green, and never ships past a held one.
19
+ >
20
+ > **This is not advice about a slash command — it is the rule for the last two doors
21
+ > of the loop** (`../../CLAUDE.md § The dev → prod loop`). The only work the asked
22
+ > session does itself is deciding the target sha and relaying the report. If you find
23
+ > yourself typing `release.sh` or `wrangler` in the conductor, you have already skipped
24
+ > the agent — and with it the CEO post that is the only record the release happened.
25
+ >
26
+ > **And the doctor holds the box.** `deploy.sh` reads `health.sh --box --json`
27
+ > before its first gate and refuses an UNHEALTHY box (`DEPLOY_SKIP_DOCTOR=1`
28
+ > overrides). When it does, spawn `doctor` and let release-manager wait:
29
+ >
30
+ > ```
31
+ > Agent({ subagent_type: "doctor", model: "opus",
32
+ > prompt: "health.sh says: <why>. Reclaim what nothing is coming back for, then report the verdict." })
33
+ > ```
34
+ >
35
+ > A deploy runs five gates. On a paging box each one runs long, and a slow gate
36
+ > is indistinguishable from a red one at the wall clock — which is how a good
37
+ > tree gets diagnosed as broken code. The box comes first.
38
+
39
+ **Skills:** `/cloudflare` (Workers auth) · `/signal` (deploy:success / deploy:degraded) · `deploy` (the run as tracked work) · `.claude/skills/deploy/REFERENCE.md` (every trap this page used to carry)
40
+
41
+ **Before a production deploy, run `/sweep`** — it lands every finished branch into
42
+ `dev`, pays **one** gate on the integrated tree instead of one per branch, sweeps
43
+ the worktrees, and ends by naming the sha to promote. `/deploy` then has one tree
44
+ to think about instead of six.
4
45
 
5
46
  Ship all five services to Cloudflare. Deterministic sandwich — W0 baseline, build, smoke, approval, parallel deploy, health.
6
47
 
@@ -57,6 +98,11 @@ retyping its steps.
57
98
  | Crons | stripped | shipped |
58
99
  | Who ships it | agents, finishing a loop | a human, promoting |
59
100
 
101
+ **`./deploy dev` ships whatever tree it is invoked from — which is `main`.** To
102
+ put a *branch* on dev.one.ie, and to open the PR that proposes it, use
103
+ `bash .claude/scripts/land.sh <branch> --pr --deploy` — one command for
104
+ gate → dev → probe → PR. See "Branch → dev → PR → main" below.
105
+
60
106
  `./deploy dev` exits into `deploy-dev.sh` immediately rather than threading a
61
107
  flag through the production pipeline — a different destination with a different
62
108
  gate is a different procedure, and `deploy-dev.sh` stays its authority. The one
@@ -115,533 +161,51 @@ the commands themselves live in the script.
115
161
  | **7 · Deploy** | Gateway + Sync + Agents + Pay in parallel (~10s each), then Astro (~30s, largest bundle). Every call bare — no `--env`, ever |
116
162
  | **8 · Health** | 4 HTTP probes × 3 tries with backoff, cache-busted with `?_t=`. Sync is cron-only — its clean deploy IS its health signal |
117
163
 
118
- ### The suite runs as two lanes, not one run (measured 2026-09-03)
119
-
120
- The gate had become a coin flip. Four full runs on an **unchanged tree** gave
121
- four different failure sets — 8, 12, 13, 8 — and every failing file in all four
122
- was one of the 19 suites marked `real TypeDB`. Not one of the other ~1126 files
123
- ever failed. Each run's log carried 41-72 `upstream 503` plus 25-38
124
- `upstream 500`, and those print only *after* the retry budget is spent
125
- (`substrate.ts:208`).
126
-
127
- The cause is capacity, not code. `substrate.ts:56-60` already recorded it in
128
- 2026-07-25: the shared TypeDB Cloud gateway 503s **in bursts lasting seconds**,
129
- while the retry cover is 500ms + 1000ms. Eight vitest forks against one external
130
- singleton turn that into a dice roll — and a dice-roll gate costs far more than
131
- it saves. The 2026-09-03 deploy spent hours on re-runs and root-cause agents for
132
- failures that were never in the diff.
164
+ ## Branch dev PR main `land.sh`, not this doc
133
165
 
134
- `.claude/scripts/test-lanes.sh` splits the suite by its two binding constraints
135
- and runs them **concurrently**:
136
-
137
- | lane | files | bound by | concurrency |
138
- |---|---|---|---|
139
- | `pool` | ~1126 | CPU | 8 forks |
140
- | `typedb` | 19 | one shared gateway | serial (`--no-file-parallelism --maxWorkers=1`) |
141
-
142
- The `typedb` lane spends its life waiting on a socket, so it costs almost no CPU
143
- while `pool` saturates the cores: wall clock is **max(), not sum()**.
144
-
145
- | | before | after |
146
- |---|---|---|
147
- | wall clock | 238-282s | **121.9s** |
148
- | result | RED, rotating victims | **BOTH GREEN** — pool 1125 passed, typedb 19 passed |
149
- | tests | 8-13 failing | 10940 passed |
150
-
151
- **This is not `--no-file-parallelism` over the whole suite** — that serialises
152
- 1126 innocent files to fix 19. Every test still runs, no assertion is weakened,
153
- nothing is skipped; only the suites sharing the external singleton are
154
- serialised, which removes the contention at its source.
155
-
156
- **The selector is read from the tree, never hardcoded.** It greps the
157
- `real TypeDB` marker the suites already carry, so a new such suite joins the
158
- serial lane automatically. A frozen list is exactly how this rots back into
159
- flakiness — `--self-test` asserts the selector finds >0 files, is a strict
160
- subset, and that every path it names exists.
161
-
162
- Both lanes still go through `test-cached.sh`, so passes are memoised per lane
163
- and a RED is still never cached. `test-full.sh` remains the one definition of
164
- the vitest flags and hands them over as `TEST_FULL_ARGS`.
166
+ **`.claude/scripts/land.sh` is the authority for this procedure**, the same way
167
+ `deploy.sh` is the authority for the pipeline. It replaced the hand-rolled
168
+ `git worktree add` + `gh pr create` recipe that used to sit here; that recipe
169
+ was a second copy of a procedure and it rotted — it still told you to
170
+ `git switch main` in the shared tree, which `hook:branch-pin` refuses.
165
171
 
166
172
  ```bash
167
- bash .claude/scripts/test-lanes.sh --list # show the split, run nothing
168
- bash .claude/scripts/test-lanes.sh --self-test # prove the selector still selects
169
- ```
170
-
171
- **Two lanes mean two `Tests N passed` lines in the log.** `TESTS_REPORT` used to
172
- `tail -1` and reported only the 19-file lane — the first two-lane production
173
- deploy said `Tests 143 passed` for a run that executed **10940**. It now sums the
174
- lanes and says `(2 lanes)`. A gate that is fine while the report understates it
175
- by two orders of magnitude is the same dishonesty as calling a fast pass a full
176
- one.
177
-
178
- ### Making the deploy faster — two disproved ideas (measured 2026-09-03)
179
-
180
- Recorded so nobody spends an afternoon re-deriving them. Both were plausible,
181
- both are wrong, and the second fails in a way that reads like a pass.
182
-
183
- **The gates are at their floor at ~256s. Rescheduling does not move them.**
184
- From the 2026-09-03 production log:
185
-
173
+ bash .claude/scripts/land.sh feat/x --pr --deploy --probe / --probe /pricing
186
174
  ```
187
- vitest 256s (the lanes, run alone: 121.9s)
188
- build 256s (astro's own report: 2m 13s = 133s)
189
- gates wall-clock: 256s
190
- ```
191
-
192
- Serial would be 133+122 = 255s — *identical*. Each gate paid ~2x its solo cost
193
- and the overlap bought nothing, because eight vitest forks plus an 8 GiB-heap
194
- build on a 10-core box is oversubscription, not parallelism.
195
175
 
196
- The obvious next move give the build room by capping the pool — makes it
197
- **worse**:
176
+ One command, four steps, in this order:
198
177
 
199
- | | gates wall-clock | build |
178
+ | # | Step | What it proves |
200
179
  |---|---|---|
201
- | 8 forks (baseline) | **256s** | 133s |
202
- | 5 forks (`VERIFY_POOL_FORKS=5`) | **272s** | 209s |
203
-
204
- Parallel, serial and capped all land at 250-270s. That is the floor for
205
- build+suite on this hardware as currently shaped.
206
-
207
- **The real target is import cost, and the obvious fix is blocked upstream.**
208
- The pool lane, run solo, spends more time loading modules than running tests:
209
-
210
- ```
211
- import 145.42s | tests 107.87s
212
- ```
213
-
214
- `vitest.config.ts:131` sets `pool: 'forks'`, and every fork re-imports the whole
215
- module graph independently. Threads share a module cache, so that ought to be
216
- the win but **vitest 4.1.7's threads pool is broken in this repo**. Every
217
- worker dies with `The worker thread was torn down or never initialized. This is
218
- a bug in Vitest.` Do not reach for `--pool=threads` until vitest is upgraded and
219
- this is re-tested.
220
-
221
- **Read the exit code, not the summary.** That broken run printed:
222
-
223
- ```
224
- Test Files no tests
225
- Tests no tests
226
- Duration 162ms
227
- ```
228
-
229
- with `rc=1`. "no tests" in a 162ms run is a *collect crash*, one careless glance
230
- from being reported as a clean pass — the same shape as
231
- `vitest-collect-crash-reads-as-pass`. An empty selection is never a pass.
232
-
233
- **What actually speeds a deploy today:** scope it. `./deploy dev` runs the fast
234
- lane with no approval; `./deploy workers` skips the astro rebuild entirely when
235
- `one.ie/web` did not change. The astro build is the long pole and is already
236
- memoised by tree fingerprint — a warm tree pays none of it.
237
-
238
- ### Why the gates are not all parallel (measured 2026-08-19)
239
-
240
- Steps 1+3 used to launch seven heavy processes with a bare `&`: five `tsc`,
241
- vitest, and the astro build. That is not parallelism, it is a swap storm. The
242
- astro build carries `--max-old-space-size=8192` and vitest runs a driver plus
243
- four 1GB forks, so the two together want ~14GB — on a 24GB box where editors
244
- and sessions are already resident.
245
-
246
- Paging is a cliff, not a slope. Same suite, same commit, same machine:
247
-
248
- | vitest ran… | wall-clock | outcome |
249
- |---|---|---|
250
- | alone | **176s** | 912 files, 7624 tests, green |
251
- | beside the build + 5×tsc | **1145s** | 0.0% CPU, no log output, killed |
252
-
253
- It happened twice in one day before anyone read it as anything but "tests are
254
- slow". They are not slow — the suite's own summary accounts for only ~165s of
255
- actual test time.
256
-
257
- **Correction, same day: memory is NOT the proven cause of the hang.** The
258
- serialisation below is still worth having — 14GB of overlap on a 24GB box is
259
- real — but the 1145s runs were later reproduced with the build serialised and
260
- the box at 41% free, no swap thrash, and no network connections held. See
261
- "The vitest gate hangs" below. Do not cite the memory story as the explanation
262
- for a hung gate; it explains a *slow* gate, not a *parked* one.
263
-
264
- So the heavy gates are now **priced before they launch**: `heavy_free_gb`
265
- reads free memory the way `lib/govern.sh` does and compares it against
266
- `DEPLOY_HEAVY_NEED_GB` (default 14). Enough → they overlap as before. Not
267
- enough → they run one after the other, which is roughly **3x faster
268
- end-to-end** than "parallel", because the parallel version spends its time
269
- paging. The five typechecks stay parallel and ungoverned; cached by tree
270
- fingerprint, they cost ~1s each.
271
-
272
- Both heavy gates also run under `gate-run.sh`, for its wall-clock **bound**
273
- (`DEPLOY_GATE_TIMEOUT`, default 900s — a healthy vitest is 176s and a build
274
- 174s) and its process-group **reap**. macOS ships no `timeout(1)`, and killing
275
- only the shell reparents the vitest forks to launchd. A hung gate must die on a
276
- clock rather than outlive the deploy.
277
-
278
- A probe that cannot read memory returns a large number, so a broken sensor
279
- never silently serialises the pipeline. Escape hatches: `DEPLOY_HEAVY_PARALLEL=1`
280
- forces overlap, `DEPLOY_HEAVY_NEED_GB` retunes the threshold.
281
-
282
- ### The vitest gate hangs — open, characterised, unexplained
283
-
284
- > **Numbers below are from 2026-08-19 and are superseded as measurements** (the
285
- > suite was 912 files then and is 1145 now; it runs as two lanes since
286
- > 2026-09-03, ~122s wall). The *diagnosis* still stands and the hang is still
287
- > unexplained, so the section is kept as-is rather than half-rewritten. Note the
288
- > non-TTY fork stall below bit again on 2026-09-03: a background re-run with
289
- > stdout to a file parked until it was killed, and the `script -qeF /dev/null`
290
- > wrapper fixed it. That wrapper is not optional for any logged run.
291
-
292
-
293
- Three times on 2026-08-19 the vitest gate parked indefinitely and had to be
294
- killed. What is established:
295
-
296
- - The suite itself is healthy: run on its own it completes in **176s**, 912
297
- files / 7624 tests / 0 failures, on the same commit and machine.
298
- - The hang is **not** the astro build competing for memory. It reproduced with
299
- the heavy gates serialised, `memory_pressure` reporting 41% free and low
300
- pageouts.
301
- - It is **not** network. At the moment of the hang the vitest main process held
302
- **no** open sockets, and neither did its worker.
303
- - The shape is a fork-pool stall: main parked in `LibuvStreamWrap::OnUvRead`
304
- waiting on worker IPC, while its single worker sat at **0.43s CPU / 46MB RSS**
305
- seven minutes in — a fork that was spawned and never given work. The log
306
- always stops ~90s in, after the jsdom noise, before any test result.
307
- - `globalSetup` was ruled out: `tests/_global-setup.ts` is local only (reads an
308
- env file, prints) with no network call to hang on.
309
-
310
- Not established: why. The one difference between every hung run and the green
311
- one is that the green run passed `--testTimeout/--hookTimeout/--teardownTimeout`
312
- explicitly — but nothing in that run came close to a timeout, so that may be
313
- coincidence rather than cause.
314
-
315
- Until it is understood, the gate runs under `gate-run.sh`'s bound, so a hang
316
- now dies on a clock instead of outliving the deploy. If it bites you: the suite
317
- is trustworthy run directly (`cd one.ie/web && bunx vitest run`), and a deploy
318
- whose only commits since a green run are harness/doc changes can legitimately
319
- use `--skip-tests` — verify with
320
- `git diff --name-only <green-sha>..HEAD | grep -v '^\.claude/'` returning empty.
321
-
322
- **Health endpoints** (custom domains only — `*.oneie.workers.dev` is blocked on
323
- this network, curl exit 6/28): `api.one.ie/health` · `one.ie/api/health`
324
- (assert `"status":"ok"`) · `channels.one.ie/health` · `pay.one.ie/status`
325
- (assert `"ok"`; there is no `/health` on that entry — it 404s). The agents
326
- worker is named **`channels`**, reachable at `channels.one.ie`.
327
-
328
- **Build OOM, fixed 2026-07-19:** the build was V8-heap-OOMing (`FATAL ERROR:
329
- Reached heap limit`, `Abort trap: 6`, exit 134) at Node's default ~4 GiB
330
- old-space ceiling — not a system memory shortage. `package.json`'s `build`
331
- script now sets `NODE_OPTIONS=--max-old-space-size=8192` inline; no manual env
332
- var needed.
333
-
334
- **Migrations and the trap:** `wrangler.toml`'s `[[routes]]`/`[triggers]` were
335
- reconciled to top level 2026-07-04, so there is no env-scoped target left to
336
- hit and `--env production` has no block to resolve against. Passing it anyway
337
- silently targets the `one-prod-production` decoy — see the trap note at the top
338
- of this file.
339
-
340
- ## Commit, PR, merge — the human half
341
-
342
- The script does not do this. Run it before `./deploy` (its first gate refuses a
343
- dirty tree), or after a `--allow-dirty` run.
344
-
345
- Tests are green at this point, so it is safe to commit. The shared tree is pinned to `main` (`hook:branch-pin`) — never `git checkout`/`switch` there. Do the commit + branch + PR in a scratch worktree instead, then merge via `gh` and fast-forward the shared tree.
346
- ```bash
347
- # 1. From the shared tree (still on main): stage by explicit path, never -A
348
- git status --porcelain # review scope
349
- git diff --stat
350
-
351
- # 2. Cut a worktree for the deploy branch
352
- slug="deploy-$(git rev-parse --short HEAD)"
353
- git worktree add -b "deploy/$slug" ".do-worktrees/$slug" main
354
-
355
- # 3. In the worktree: bring over the same paths, commit, push
356
- cd ".do-worktrees/$slug"
357
- git add <explicit paths from step 2>
358
- git commit -m "$(cat <<'EOF'
359
- deploy: ship <summary of staged changes>
360
- EOF
361
- )"
362
- git push -u origin "deploy/$slug"
363
-
364
- # 4. Open PR, wait for CodeRabbit + any checks, merge
365
- gh pr create --base main --head "deploy/$slug" --title "deploy: <summary>" --body "Automated deploy PR — see /deploy pipeline"
366
- gh pr merge "deploy/$slug" --squash --auto --delete-branch
367
-
368
- # 5. Back in the shared tree: fast-forward and clean up the worktree
369
- cd /Users/toc/Server/one-ie
370
- git switch main # only allowed HEAD move in the shared tree
371
- git pull origin main
372
- git worktree remove ".do-worktrees/$slug"
373
- ```
374
- Generate the commit message and PR title from the diff — conventional commit format (`feat:`, `fix:`, `chore:`, etc.). If there is nothing to commit (`git status --porcelain` is empty), skip this whole step silently — proceed to Step 3 on the tree as-is.
375
-
376
- `gh pr merge --auto` queues the merge and returns immediately if required checks (e.g. CodeRabbit review) are still running; poll `gh pr view "deploy/$slug" --json state,mergedAt` before Step 5's `git pull` if the merge hasn't landed yet — don't build off an unmerged branch.
377
-
378
- ---
379
-
380
- ## Bundle Size Rules (CF Workers Free Tier — 3 MiB gzipped upload)
381
-
382
- The Astro Worker upload must stay under **3 MiB gzipped** on the free tier (10 MiB
383
- on paid). Wrangler reports both `Total Upload` (uncompressed) and `gzip` — only
384
- gzip counts toward the ceiling. All chunks in `dist/server/chunks/` are uploaded
385
- together; dynamic `await import()` does NOT exclude code from the upload.
386
-
387
- These rules are **LOCKED** — do not revert them. Apply identically to any
388
- developer template we ship (`oneie init` Workers scaffold mirrors this shape).
389
-
390
- ### Rule 1 — `syntaxHighlight: false` in `astro.config.mjs`
391
-
392
- ```js
393
- markdown: { syntaxHighlight: false }
394
- ```
395
-
396
- Disables Shiki from Astro's markdown pipeline. Without this, Shiki pulls ~5.8 MiB of
397
- language grammar files into the SSR worker on every build. **Do not re-enable.**
398
-
399
- ### Rule 2 — `ssr.external` for heavy packages
400
-
401
- ```js
402
- ssr: {
403
- external: ["node:async_hooks", "shiki", "@shikijs/core", "@shikijs/types", /* … */]
404
- }
405
- ```
406
-
407
- **`one.ie/web/astro.config.mjs` § `vite.ssr.external` is the authority for this
408
- list — never reconcile that file to this doc.** It carries 17 entries as of
409
- 2026-08-02 (`cookie`, `cytoscape`, `@xyflow/react`, `recharts`, `@stripe/*`,
410
- `media-chrome`, `motion`, `@100mslive/*`, … alongside the shiki trio). Deleting
411
- entries to match a stale snippet here would blow the bundle. A companion
412
- `build.rollupOptions.external` also drops anything matching `shiki/` or
413
- `@shikijs/` by prefix.
414
-
415
- The CF adapter bundles everything by default. `ssr.external` creates a bare
416
- `import { x } from 'pkg'` reference without inlining the package.
417
-
418
- **Critical nuance:** `ssr.external` only works safely when the externalized package is
419
- never executed on the server path. For `shiki`: `codeToHtml` is imported by `code-block.tsx`,
420
- but all components that use `code-block.tsx` are `client:only` — so `codeToHtml` is never
421
- called in the worker. The import statement exists in the bundle but is dead code.
422
-
423
- If you add a new heavy dependency used only client-side, add it here.
424
-
425
- ### Rule 3 — Pure-shell pages use `client:only` + `prerender = true`
426
-
427
- ```astro
428
- ---
429
- export const prerender = true
430
- import { MyComponent } from "@/components/MyComponent"
431
- ---
432
- <Layout title="...">
433
- <MyComponent client:only="react" />
434
- </Layout>
435
- ```
436
-
437
- `client:only="react"` — Astro renders an empty div on the server; the component never
438
- runs in the worker. The React component tree (+ all its imports) stays out of the SSR bundle.
439
-
440
- `export const prerender = true` — the page becomes a static asset generated once
441
- at build time. The page's SSR handler collapses to a small stub. Zero worker cost
442
- at runtime.
443
-
444
- **Use this pattern for:** any page that has no server-side data dependencies
445
- (no `Astro.locals`, no `Astro.request`, no DB queries in frontmatter).
446
-
447
- The prerendered set changes every cycle — never hardcode it. Read it from the
448
- tree (34 pages as of 2026-08-02):
449
-
450
- ```bash
451
- cd one.ie/web && grep -l "export const prerender = true" src/pages/*.astro
452
- ```
453
-
454
- Pages that CANNOT be prerendered: any page whose frontmatter reads
455
- `Astro.locals` (session/workspace context), `Astro.request`, or queries the DB.
456
- Those must stay SSR.
457
-
458
- ### Rule 4 — `inlineStylesheets: 'auto'` in `astro.config.mjs`
459
-
460
- ```js
461
- build: { inlineStylesheets: 'auto' }, // ← NEVER 'always'
462
- ```
463
-
464
- With `'always'`, Astro inlines the full Tailwind stylesheet into **every route's
465
- serialized manifest entry**. With ~100 routes the entry chunk balloons by 8+ MiB
466
- of duplicated CSS as a single string literal — diagnosable in worker-entry at
467
- the line `const _manifest = deserializeManifest({...})`.
468
-
469
- `'auto'` ships the bundle as one external `<link rel="stylesheet">` referenced
470
- once across all routes. Browsers cache it across navigations — a page-speed
471
- win, not just a worker-size win.
472
-
473
- **Verified 2026-05-22:** flipping `always` → `auto` dropped worker-entry from
474
- 9.5 MiB → 672 KiB and total gzip from 3302 KiB → 2079 KiB.
475
-
476
- ### Rule 5 — `react-dom/server.edge` alias (production only)
477
-
478
- ```js
479
- resolve: {
480
- alias: {
481
- ...(isDev ? {} : { "react-dom/server": "react-dom/server.edge" })
482
- }
483
- }
484
- ```
485
-
486
- Already in `astro.config.mjs`. Required for CF Edge runtime compatibility.
487
- Do not remove for production builds.
488
-
489
- ---
490
-
491
- ## Verified Bundle Numbers
492
-
493
- | Snapshot | Total upload | gzip | Worker-entry | What changed |
494
- |---|---|---|---|---|
495
- | 2026-04-18 (post Pages→Workers migration) | — | — | 9.5 MiB | Rules 1-3 + 5 |
496
- | 2026-05-22 before `inlineStylesheets` fix | 18.5 MiB | 3.3 MiB | 9.5 MiB | Over 3 MiB ceiling — deploy FAILED |
497
- | 2026-05-22 after Rule 4 (`'always'` → `'auto'`) | 10.1 MiB | **2.1 MiB** | **672 KiB** | Under ceiling — deploy ✓ |
498
- | 2026-07-08 | 14.7 MiB | **3.22 MiB** | — | **Exceeds the documented 3 MiB (3072 KiB) free-tier ceiling and still deployed successfully.** Either this account is on a paid Workers plan (10 MiB ceiling) rather than free tier, or the ceiling figure elsewhere in this doc is stale — unconfirmed which. Don't treat "under 3 MiB" as a hard gate until this is resolved; treat 3.2 MiB as the new floor to watch, and re-run Bundle Size Diagnosis if growth continues. |
499
- | 2026-07-19 | 20.2 MiB | **4.43 MiB** | — | +37% over 2026-07-08's 3.22 MiB. Deployed successfully — no diagnosis run yet. Growth window covers several merged features that day (newsletter platform, directory-submission, social-formats, movers-playbook, etc.) landing in one `/deploy` cycle; not isolated to a single change. Re-run Bundle Size Diagnosis if the next snapshot keeps climbing. |
500
- | 2026-07-29 | 20.9 MiB | **4.60 MiB** | 1.2 MiB | +6% over 2026-07-19, third consecutive climb. Cheap diagnosis WAS run this cycle (the two grep/`ls` commands below, not a full audit): **no single runaway** — Shiki hits 0, top chunks are `_astro_data-layer-content` 2.1 MiB (content collections), `worker-entry` 1.2 MiB (up from 672 KiB at the 2026-05-22 baseline), `index` 1.3 MiB, `icons` 0.8 MiB, `mermaid` 0.8 MiB, `stripe.esm.worker` 0.6 MiB, `react-vendor` 0.5 MiB. Growth is diffuse feature accretion, not a regression; this deploy's own diff was test-only. Two named candidates if a real audit is ever warranted: `mermaid` (0.8 MiB in the SSR bundle — Rule 2 `ssr.external` candidate if it's only reached from `client:only` islands) and the 15 chunks referencing `react-vendor` (Rule 3 suggests some page still SSR-renders React via `client:load`). |
501
- | 2026-08-02 | 22.0 MiB | **4.96 MiB** | 1.2 MiB | +8% over 2026-07-29, **fourth consecutive climb**. Cheap diagnosis run again: still **no single runaway** — Shiki 0, `react-vendor` referenced by 15 chunks (unchanged), `worker-entry` flat at 1.2 MiB. Top chunks: `_astro_data-layer-content` 2.1 MiB, `index` 1.3 MiB, **`_broadcast_` 1.3 MiB (new to the top list)**, `worker-entry` 1.2 MiB, `index` 0.9 MiB, `icons` 0.8 MiB, `mermaid` 0.8 MiB, `stripe.esm.worker` 0.6 MiB, `react-vendor` 0.5 MiB, **`generateAuthenticationOptions` 0.5 MiB (new)**. This cycle shipped beautiful-blocks (45 visual blocks + 12 background components), which is a plausible share of the delta. Four climbs in a row with the same "diffuse accretion" verdict each time is itself the signal — the cheap diagnosis has now exhausted what it can tell us, and the two standing candidates (`mermaid` via Rule 2, the 15 `react-vendor` chunks via Rule 3) want a real audit rather than a fifth restatement. |
502
-
503
- The 2026-05-22 regression was caused by `build: { inlineStylesheets: 'always' }`
504
- inlining the full Tailwind stylesheet into every route's manifest entry. One
505
- char change (`always` → `auto`) saved 8.8 MiB.
506
-
507
- ---
508
-
509
- ## Service Map (post-migration)
510
-
511
- `./deploy <mode>` covers every live row; the per-service command is what the
512
- script runs, recorded here for rollback and one-off work.
513
-
514
- | Service | URL | Config | Deploy command |
515
- |---------|-----|--------|---------------|
516
- | Astro Worker (prod) | `one.ie` → `one-prod` | `one.ie/web/wrangler.toml` | `cd one.ie/web && wrangler deploy` — **no `--env production`** (deploy-target trap: appends `-production` to the script name → `one-prod-production`, which nothing routes to; see trap note at top of this file) |
517
- | Gateway | `api.one.ie` → `one-gateway` | `api/wrangler.toml` | `cd api && wrangler deploy` |
518
- | Sync | `one-sync` — **cron-only, no HTTP route** (health = deploy success) | `sync/wrangler.toml` | `cd sync && wrangler deploy` |
519
- | Agents | `channels.one.ie` → `channels` (`*.workers.dev` blocked on this network) | `channels/wrangler.toml` | `cd channels && wrangler deploy` |
520
- | Pay gateway | `pay.one.ie` → `one-core-worker` | `pay/backend/wrangler.toml` | `cd pay/backend && bun run deploy` (= bare `wrangler deploy`, no `--env` flag) |
521
- | Pages (legacy idle, rollback) | `oneie.pages.dev` | — | **do not deploy** — rollback target for `one.ie` |
522
- | Worker (legacy idle, rollback) | `one-demo` (still serves `demo.one.ie`, `onestudio.dev`) | — | **do not deploy** — rollback window for the prod cutover |
523
-
524
- ---
525
-
526
- ## Auth (CRITICAL — never change)
527
-
528
- Never: `CLOUDFLARE_API_TOKEN` (scoped token lacks workers + custom domain permissions).
529
-
530
- **wrangler reads `CLOUDFLARE_API_KEY` + `CLOUDFLARE_EMAIL`** for global-key auth
531
- — `CLOUDFLARE_GLOBAL_API_KEY` is *our* name for it and wrangler ignores it. The
532
- script now exports both spellings off one resolved value, so the two can never
533
- diverge again.
534
-
535
- **The credential lives on disk, not in your shell** — `.env.local` at the repo
536
- root and `one.ie/web/.env` (same 52-char key). You do not need to export
537
- anything. If you *do* export one, it wins — which is the trap: a **stale**
538
- export shadows the good key everywhere, because process env beats wrangler's
539
- per-directory `.env` autoload. That is exactly how 2026-08-19's deploy used two
540
- different credentials in two consecutive steps (see the Step 6.6 note below).
541
- The ladder now probes each rung and falls through a rung that does not answer,
542
- so a stale export costs a log line instead of a red deploy:
543
-
544
- ```
545
- rejected: ambient env (len=37 sha=eaafbda9) — /user did not answer 200
546
- ✓ resolved: global-api-key
547
- source: /Users/toc/Server/one-ie/.env.local (len=52 sha=435cba23)
548
- ✓ 5/5 services agree on account 627e0c7c…
549
- ```
550
-
551
- Ground truth is the API, in **both** directions — it is as able to prove a key
552
- alive as dead. Never conclude either from wrangler alone:
553
-
554
- ```bash
555
- curl -s -o /dev/null -w '%{http_code}\n' \
556
- -H "X-Auth-Email: $CLOUDFLARE_EMAIL" -H "X-Auth-Key: $CLOUDFLARE_API_KEY" \
557
- https://api.cloudflare.com/client/v4/user # 200 = fine
558
- ```
559
-
560
- The deploy script auto-unsets `CLOUDFLARE_API_TOKEN` from the spawned env to prevent
561
- accidental use of a scoped token that was exported in the shell.
562
-
563
- Required env (export locally before running `./deploy` — there is no CI; the
564
- script asserts the first two at gate 4 and refuses to ship without them):
565
- - `CLOUDFLARE_GLOBAL_API_KEY` + `CLOUDFLARE_EMAIL` — auth
566
- - `PUBLIC_GATEWAY_URL: https://api.one.ie` — build-time-inlined by Astro (**required**; without it the Worker bundle falls back to `one-gateway.oneie.workers.dev` and gateway-backed routes break). Lives in `one.ie/web/.env`.
567
-
568
- ---
569
-
570
- ## Mode-specific notes
571
-
572
- Everything common lives in the script. These are the bits that are true of one
573
- mode only.
574
-
575
- ### `./deploy astro`
576
-
577
- **Resolved 2026-07-04:** `one.ie/web/package.json`'s `"deploy"` script had the
578
- same `--env production` trap. Confirmed via the CF API that it was real —
579
- neither `one-prod` (live) nor the `one-prod-production` decoy had any cron
580
- schedules registered, meaning `billing-alerts-cron.ts` /
581
- `billing-allocation-cron.ts` / `billing-autotopup-cron.ts` /
582
- `billing-lifecycle-cron.ts` / `billing-verify-cron.ts` /
583
- `funnel-aggregate-cron.ts` / `webhook-deliver.ts` / `broadcast-drain-cron.ts`
584
- were not running on any schedule. Fixed by reconciling `wrangler.toml`: the
585
- `[[routes]]` (custom domain) and `[triggers]` (crons) blocks — the only two
586
- things that existed *only* under `[env.production]` — were moved to top level
587
- (everything else was already duplicated there); the now-fully-redundant
588
- `[env.production.*]` block was deleted entirely, and `package.json`'s script
589
- dropped `--env production`. The trap is now structurally impossible — there's
590
- no `--env production` target left to hit.
591
-
592
- **Pre-flight (one-time, on cutover only):** ensure no other CF entity owns the
593
- `one.ie` custom domain. If wrangler errors with a hostname conflict, detach the
594
- prior owner first:
595
-
596
- ```bash
597
- # If a Pages project owns it (was `oneie` project pre-cutover):
598
- curl -s -X DELETE \
599
- "https://api.cloudflare.com/client/v4/accounts/$CF_ACCOUNT_ID/pages/projects/oneie/domains/one.ie" \
600
- -H "X-Auth-Email: $CLOUDFLARE_EMAIL" \
601
- -H "X-Auth-Key: $CLOUDFLARE_GLOBAL_API_KEY"
602
- ```
603
-
604
- ### `./deploy workers`
605
-
606
- Verified 2026-07-08: gateway, sync and channels complete independently with no
607
- shared state — which is why the script runs them concurrently.
608
-
609
- ### `./deploy pay`
610
-
611
- `pay/backend/src/index.ts` mounts `src/routes/status.ts` (`/status`) and
612
- `discoveryRoutes` (`/`). It does **not** serve `/health` — the `/health`
613
- handler in `pay/backend/src/api/routes/status.ts` belongs to the separate,
614
- unmounted `src/api/` tree. Both `/` and `/status` probed live 2026-08-02: 200,
615
- unauthenticated. `/status` also reports version, destinationMode, and contract
616
- addresses.
617
-
618
- **Found 2026-07-05:** `pay/backend` shipped a real commit (`feat(pay):
619
- embeddable payment-link page`) that sat unshipped through a full deploy cycle
620
- because the service map only named 4 services. `pay.one.ie` is a first-class
621
- 5th target, not an afterthought — check `git log` scoped to `pay/` for
622
- unshipped commits, same as the other four.
623
-
624
- ## Bundle Size Diagnosis
625
-
626
- If build fails with "exceeds size limit":
627
-
628
- ```bash
629
- # Check total worker size
630
- du -sh dist/server/
631
-
632
- # Find top offenders
633
- ls -lhS dist/server/chunks/ | head -20
634
-
635
- # Check if a new import pulled in Shiki
636
- grep -r "from 'shiki'" dist/server/chunks/ | wc -l
637
- # If > 0: a component that imports shiki was SSR'd
638
- # Fix: make its page client:only="react" + prerender=true
639
-
640
- # Check if React crept back into worker via client:load
641
- grep -l "react-vendor" dist/server/chunks/
642
- # If multiple chunks: some page SSR-renders React via client:load
643
- # Fix: audit src/pages/*.astro for client:load on pure-shell pages
644
- ```
180
+ | 1 | **gate** — `verify:fast` in the branch's own worktree | the branch is green ON ITS OWN. Not that it survives the trunk — `--pr` deliberately does not merge main in, so the reviewer sees what the branch added and GitHub computes the merge |
181
+ | 2 | **dev** — `$wt/.claude/scripts/deploy-dev.sh` | the branch RUNS. It ships the *worktree's* tree, because `deploy-dev.sh` derives its own ROOT from its own path — `$ROOT`'s copy would ship main and call it the branch |
182
+ | 3 | **probe** — `do-prove.sh`, both bases pinned to dev | the routes answer on dev, under the LANDING RULE. Not `curl / → 200`, which a redirect to `/signin` satisfies |
183
+ | 4 | **PR** — `pr-body.sh` → `gh pr create`/`edit` | a reviewer gets the diff, the trunk drift, the mergeability, the gate label and the dev URL. Re-running UPDATES the open PR, never duplicates it |
184
+
185
+ Then a human merges the PR, and **promotion to `one.ie` is still `./deploy`** —
186
+ the full gate, from `.release/`. Nothing above touches production.
187
+
188
+ Four things that are load-bearing, each of which read as a pass before it was fixed:
189
+
190
+ - **dev.one.ie is one slot, and it writes production's rows.** `--pr --deploy`
191
+ takes exactly one branch; with two, the second overwrites the first while the
192
+ first is being probed. Ship code there freely; treat its DATA as production.
193
+ - **Both probe bases are pinned to dev.** `do-prove.sh` falls back to
194
+ `PROVE_PROD_URL` (default `https://one.ie`) when its dev base is silent an
195
+ unreachable dev would otherwise prove *production* and report it as the branch
196
+ passing.
197
+ - **The probe's route COUNT is read, not its exit code.** `PROVE: skipped (no
198
+ reachable environment)` exits 0. An unrun probe is not a pass.
199
+ - **The gate is paid once, except under `--quick`.** Step 2 passes
200
+ `DEV_SKIP_GATE=1` because step 1 just ran `verify:fast` in that same tree.
201
+ Under `--quick` step 1 was tsc only — not the fast lane — so `deploy-dev.sh`
202
+ runs its own gate before anything reaches dev.
203
+
204
+ `land.sh` still has its other two doors: bare (merge main in → gate → `--ff-only`
205
+ main → optionally one dev deploy for the batch) and `--pr` alone (gate → PR, no
206
+ dev). The commit itself stays human — a commit message needs judgment, and
207
+ `deploy.sh`'s first gate refuses a dirty tree so the deploy and the history
208
+ cannot disagree.
645
209
 
646
210
  ---
647
211
 
@@ -668,211 +232,25 @@ previous deployment — it needs no worktree and no rebuild.
668
232
 
669
233
  ---
670
234
 
671
- ## The TypeDB flake waiver — a red suite the deploy may ship past
672
-
673
- `one.ie/web`'s suite talks to a REAL shared TypeDB Cloud cluster (CLAUDE.md:
674
- "Don't mock TypeDB in integration tests"). When that cluster blips or a query
675
- outruns its timeout, a handful of task/substrate suites go red without anything
676
- in the diff being wrong. That used to be an eyeball judgement, which is exactly
677
- the call that gets rubber-stamped on the fifth deploy attempt at 2am.
678
-
679
- `.claude/scripts/typedb-flake-check.sh` makes it a check. When the vitest gate
680
- goes red, `deploy.sh` runs it against the gate log:
681
-
682
- | exit | means |
683
- |---|---|
684
- | 0 | every failure carries a substrate-unavailable signature — waivable, deploy continues |
685
- | 1 | at least one failure is real, or the log could not be classified — deploy stops |
686
-
687
- **It keys on the failure SIGNATURE, never on the filename.** A file-based
688
- allowlist waives every future failure in that file, including the real ones. The
689
- waived signatures are `upstream_50[234]`, `status=50[234]`,
690
- `fixture write failed`, `Test timed out in Nms`, `ETIMEDOUT`, `ECONNRESET`,
691
- `ECONNREFUSED`, `EAI_AGAIN`, `socket hang up`, `fetch failed`, and TypeDB
692
- connection errors.
693
-
694
- Four properties, each with a red half in `--self-test`:
695
-
696
- - **`not_found` is never waivable.** Checked first, independently of everything
697
- else. It is the `tasks:claim` privilege boundary and four separate real
698
- defects have presented as that exact string. A `not_found` wrapped inside a
699
- 503 still blocks.
700
- - **One flake never vouches for its neighbour.** The log is split into vitest's
701
- per-failure blocks and EVERY block must carry a signature. A genuine assertion
702
- break standing beside a 503 blocks the deploy.
703
- - **Silence is not a pass.** An empty or unparseable log, or one with no
704
- `Tests N failed` line, exits 1.
705
- - **A waiver is not a green suite.** The report says
706
- `WAIVED as TypeDB outage (suite NOT green)`, and a waived run **cannot settle
707
- deferred-pin debt** — the waiver speaks only to the failures that reported, not
708
- to a pin that never got to.
709
-
710
- On by default. `--no-typedb-flake-waiver` (or `DEPLOY_ALLOW_TYPEDB_FLAKE=0`)
711
- restores the hard stop. Prove the checker still bites before trusting it:
712
-
713
- ```bash
714
- bash .claude/scripts/typedb-flake-check.sh --self-test # 6 cases, 4 of them red halves
715
- bash .claude/scripts/typedb-flake-check.sh <a-gate-log>
716
- ```
717
-
718
- **The waiver is not a diagnosis.** `curl -sS -o /dev/null -w '%{http_code}' https://api.one.ie/health`
719
- returning 200 while the suite reports 503s means the cluster blipped mid-run. A
720
- 200 alongside failures that are NOT in the signature list means the code is
721
- wrong — and the checker will tell you so.
722
-
723
- ## Known-Flaky Test Allowlist
724
-
725
- `deploy.sh` deliberately enforces no allowlist — a red suite stops it, and it
726
- prints the failing files with a pointer here. Triage is yours: treat these by
727
- name when they appear in `bunx vitest run` output in `one.ie/web`; don't block deploy on them, but don't silently ignore new failures either
728
- — confirm the failure signature matches before waving it through:
729
-
730
- - `tests/e2e/c5-webhook-subscribe.test.ts` — `workflow:webhook-subscribe` "writes a KV
731
- record…" and "fails closed on a non-owned workflow". **Root cause (confirmed 2026-07-08):**
732
- `ssrfGuard()` (`one.ie/web/src/lib/ssrf.ts`) resolves DNS via Cloudflare DoH by fetching
733
- `https://1.1.1.1/dns-query` directly; this local dev network refuses connections to
734
- `1.1.1.1:443` (`curl: (7) Failed to connect`), so `resolveHostIPs` returns `[]` and every
735
- URL — including the test's `https://example.com` — comes back `blocked_url`. Verify before
736
- waving through: `curl -v --max-time 5 https://1.1.1.1/dns-query 2>&1 | grep -i refused`
737
- — if that shows "Connection refused", it's this network gap, not a code regression. If it
738
- connects fine and the test still fails, it's real — investigate.
739
- - `tests/unit/tasks-humans.test.ts` + `tests/tasks-do-roundtrip.test.ts` — **only** when
740
- the failure is `fixture write failed (status=503 error=upstream_503)` or
741
- `(status=502|504 …)`. That message means the shared TypeDB Cloud cluster was
742
- unavailable through every retry `typedbQueryDetail` already performs — the substrate
743
- refused setup, so nothing downstream proved anything. Verify before waving through:
744
- `curl -sS -o /dev/null -w '%{http_code}' https://api.one.ie/health` — a 200 there with a
745
- 503 in the test means the cluster blipped during the run, not that the code is wrong.
746
- **Any OTHER failure in these two files is real and blocks deploy** — in particular a bare
747
- `not_found` on `tasks:claim`, which is the privilege boundary and must never be waved
748
- through. Four separate defects that used to present as that same `not_found` were fixed
749
- 2026-08-02 (concurrent-run sweep destruction, silent fixture writes, same-attribute
750
- insert races, read-after-write lag); if it reappears, something new is wrong. History:
751
- the commit message on `140975e44` and `tests/helpers/probe-sweep.ts`.
752
- - Hardware/stochastic benchmarks (speed, distribution-timing tests) — expected variance,
753
- not correctness bugs.
754
-
755
- Any other failure (type errors, assertion mismatches on business logic) blocks deploy —
756
- diagnose and fix before proceeding.
757
235
 
758
236
  ---
759
237
 
760
- ## First-Time Setup
238
+ ## The reference — traps, numbers, forensics
761
239
 
762
- Only needed once, before `./deploy` can work at all. There is no
763
- `docs/deploy.md` this block plus the script is the whole walkthrough.
764
- Resource names below are the ones
765
- actually declared in `one.ie/web/wrangler.toml`; creating differently-named
766
- resources produces bindings the Worker can't resolve.
767
-
768
- ```bash
769
- # Create CF resources — names must match wrangler.toml exactly
770
- bunx wrangler d1 create one-owners # → binding DB
771
- bunx wrangler kv namespace create SESSION # → binding SESSION
772
- bunx wrangler kv namespace create CHAT_CACHE # → binding CHAT_CACHE
773
- bunx wrangler kv namespace create THREADS # → binding THREADS
774
- bunx wrangler r2 bucket create one-content # → binding CONTENT
775
- # → Paste IDs into one.ie/web/wrangler.toml + sync/wrangler.toml
776
-
777
- # Run D1 migrations (the ledger starts at 0001_owners.sql — there is no 0001_init.sql)
778
- cd one.ie/web && bunx wrangler d1 migrations apply DB --remote
779
-
780
- # Gateway secrets (TypeDB credentials) — no --env flag, ever
781
- cd api
782
- printf 'admin' | bunx wrangler secret put TYPEDB_USERNAME
783
- bunx wrangler secret put TYPEDB_PASSWORD # paste at the prompt; never inline the value
784
- cd ..
785
-
786
- # First deploy — Worker auto-provisions on first `wrangler deploy`
787
- # Then just: ./deploy
788
- ```
789
-
790
- ---
240
+ Everything below the operating surface moved to **`.claude/skills/deploy/REFERENCE.md`**
241
+ on 2026-09-17, because this page is read by a conductor deciding *who* deploys and
242
+ by an agent that then needs *all* of it. Two audiences, two lengths. Nothing was
243
+ deleted the file carries the same sections, byte for byte:
791
244
 
792
- ## Logs
793
-
794
- `./deploy` writes `.deploy-logs/deploy-<stamp>.log` (gitignored) every gate's
795
- stdout in orderplus `.deploy-logs/<service>-<stamp>.log` for each service in
796
- the parallel wave, since their output would otherwise interleave. The report at
797
- the end names the log path.
798
-
799
- Live logs:
800
-
801
- ```bash
802
- cd one.ie/web && bunx wrangler tail --name one-prod # Astro Worker (production)
803
- cd api && bunx wrangler tail # Gateway
804
- cd one.ie/web && bunx wrangler deployments list --name one-prod | head -10
805
- ```
806
-
807
- ---
808
-
809
- ## Gotchas
810
-
811
- - **A service with no local `wrangler` falls through to a shared `bunx wrangler@latest` cache, and that cache can rot.** Hit 2026-08-06: `sync/` was the only one of the five without wrangler in its `devDependencies`, so `bunx wrangler deploy` resolved to `$TMPDIR/bunx-501-wrangler@latest/` — whose install was missing `esbuild`, so it died `MODULE_NOT_FOUND` before reading a single config. The other four were unaffected because they resolve wrangler from their own `node_modules`. Fixed by pinning `wrangler` into `sync/package.json` like its siblings. If this shape reappears elsewhere, the workaround is a version-pinned invocation (`bunx wrangler@4.80.0 deploy`, which lands in a different cache dir); the fix is a local dep.
812
- - **Piping `./deploy` into `tail`/`head` masks its exit code** — the pipeline reports the pager's status, not the script's, so a failed run reads as exit 0. `die()` really does `exit 1`; read the ✓/✗ lines or the `.deploy-logs/` file, and don't infer success from a piped exit status.
813
- - TypeDB Cloud port is **1729** (not 80 or 443)
814
- - TypeDB HTTP API prefix is `/v1/` (signin, query, databases)
815
- - Always `CLOUDFLARE_GLOBAL_API_KEY` — scoped tokens lack permissions for workers + custom domains
816
- - `import.meta.env` is build-time — `PUBLIC_GATEWAY_URL` is baked into the worker bundle at build. Missing → gateway-backed routes fall back to the wrong host and break
817
- - Custom domains: `[[routes]]` double bracket, no wildcards, add `workers_dev = true`
818
- - Worker upload limit: **3 MiB gzipped** on free tier (10 MiB on paid) — though a 2026-07-08 deploy shipped at 3.22 MiB gzip successfully, so this account's actual ceiling is unconfirmed (see Verified Bundle Numbers). Wrangler reports `gzip:` — only that number counts. Follow the 5 Bundle Size Rules above regardless of which ceiling applies
819
- - **D1 schema-drift fails at runtime, not compile-time.** Migrations that DROP+CREATE a table (e.g. `0059_domains.sql` renamed `slug`→`gid`, `verified`→`verified_at`) silently break any code that queries the old columns — typecheck passes, deploy succeeds, the route 500s in production. After any DROP+CREATE migration, grep the codebase for the old column names and fix call sites BEFORE deploying
820
- - **Error responses get the same `cache-control` as success responses.** Astro's Layout sets `public, max-age=300, s-maxage=86400, stale-while-revalidate=604800` on every render including 5xx pages. A bad deploy will be cached at the CF edge for 24h. When diagnosing, always bust the cache: `curl "https://host/path?_t=$(date +%s)"`. Consider a middleware rule that strips `cache-control` on `>= 500` status
821
-
822
- ---
823
-
824
- *Deploy is the closed loop. W0 baseline in, health check out. If health fails, mark() is blocked. Determinism: every step reports numbers, every number gets marked.*
825
-
826
- ### `code: 9103` does not mean the key was revoked
827
-
828
- Measured 2026-08-18, and it cost an hour plus a wrong accusation that a security
829
- incident had rotated the key. `Unknown X-Auth-Key or X-Auth-Email [code: 9103]`
830
- at Step 6.5 was a **variable-name mismatch**: the script exported
831
- `CLOUDFLARE_GLOBAL_API_KEY`, wrangler only reads `CLOUDFLARE_API_KEY`.
832
-
833
- Before concluding a Cloudflare credential is dead, ask the API directly — it is
834
- ground truth and wrangler is not:
835
-
836
- ```bash
837
- curl -s -o /dev/null -w '%{http_code}\n' \
838
- -H "X-Auth-Email: $CLOUDFLARE_EMAIL" -H "X-Auth-Key: $CLOUDFLARE_API_KEY" \
839
- https://api.cloudflare.com/client/v4/user # 200 = the key is fine
840
- ```
841
-
842
- Three companion traps, all real:
843
-
844
- - **wrangler 4.x auto-loads `.env` from the cwd**, so a stale key in
845
- `one.ie/web/.env` beats both your exported vars and an OAuth session. Isolate a
846
- credential test by running it from `/tmp`.
847
- - **`/user/tokens/verify` is Bearer-only** — `400 Missing "Authorization" header`
848
- there is not evidence against a global key. Use `/user` or `/accounts`.
849
- - **Length proves nothing.** A classic global key is 37 hex chars; a `cfk_`-prefixed
850
- one is ~52 and equally valid.
851
-
852
- **Never edit `.env` while a deploy is running.** Doing so took a run fully red —
853
- tsc, vitest, and the astro build — for reasons that had nothing to do with the code.
854
-
855
- ### The 2026-08-19 sequel: `code: 7403` at Step 6.6, and two keys
856
-
857
- The same family, one layer deeper, and worth reading before you diagnose any
858
- Cloudflare auth failure here. Step 6.5 (D1 in `one.ie/web`) **passed** and Step
859
- 6.6 (D1 in `channels`) **failed** in the same run, seconds apart, on the same
860
- account. That is only possible if they used different credentials — and they
861
- did:
862
-
863
- - **Three** credentials were reachable from one `./deploy`: a stale 37-char key
864
- in the ambient env (injected by `~/.claude/settings.json`'s `env` block — so
865
- it existed inside Claude Code sessions and *not* in a plain terminal, which is
866
- why grepping the shell profiles found nothing), the good 52-char key in
867
- `one.ie/web/.env` + `.env.local`, and the `wrangler login` OAuth session.
868
- - Nothing *chose* between them. Each service got whatever its own cwd surfaced.
869
- Of the five dirs, only `one.ie/web` has a `.env` — so it read the good key and
870
- passed, while `channels` fell through to OAuth and hit `7403`.
871
- - **`7403` ≠ `9103`.** `9103` ("Unknown X-Auth-Key") points at a key; `7403`
872
- ("account is not authorized to access this service") points at an account
873
- scope, i.e. a session. Reading them as the same symptom is what sent the first
874
- diagnosis at a perfectly good key.
875
-
876
- Fixed by making the credential *resolved* rather than *ambient* — see Step 0.4
877
- above. Run `./deploy --check-creds` if you ever doubt which key is in play; it
878
- names the source and proves all five services agree.
245
+ | In the reference | What it holds |
246
+ |---|---|
247
+ | The suite runs as two lanes | why the full gate is two keys, and what that costs a caller |
248
+ | Two disproved speed ideas · why the gates are not all parallel | measurements that closed a question read before re-opening it |
249
+ | The vitest gate hangs | open, characterised, unexplained |
250
+ | Bundle size rules 1–5 · verified numbers · diagnosis | the 3 MiB CF free-tier ceiling and how each rule buys headroom |
251
+ | Service map · auth (CRITICAL) · mode-specific notes | the five targets, the credential ladder, `./deploy astro|workers|pay` |
252
+ | The TypeDB flake waiver · known-flaky allowlist | the red suite a deploy may ship past, and its bounds |
253
+ | First-time setup · logs · gotchas (`9103`, `7403`) | provisioning, and the two auth failures that are not the same symptom |
254
+
255
+ `release-manager` is told to read it. If you are reading this page to *run* a
256
+ deploy, you are on the wrong side of the door — see the top of this file.