@oneie/claude 0.7.0 → 0.9.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (263) hide show
  1. package/agents/abm-strategist.md +89 -0
  2. package/agents/ads-meta.md +97 -0
  3. package/agents/analyst.md +173 -0
  4. package/agents/architect.md +280 -0
  5. package/agents/brand-guardian.md +88 -0
  6. package/agents/brand-strategist.md +92 -0
  7. package/agents/campaign-content.md +90 -0
  8. package/agents/campaign-email.md +88 -0
  9. package/agents/campaign-sms.md +87 -0
  10. package/agents/campaign-social.md +87 -0
  11. package/agents/cco.md +99 -0
  12. package/agents/ceo.md +106 -0
  13. package/agents/chairman.md +120 -0
  14. package/agents/cmo.md +101 -0
  15. package/agents/community-greeter.md +89 -0
  16. package/agents/community-moderator.md +92 -0
  17. package/agents/compliance.md +92 -0
  18. package/agents/copywriter.md +100 -0
  19. package/agents/creative-strategist.md +95 -0
  20. package/agents/cro.md +98 -0
  21. package/agents/cto.md +267 -0
  22. package/agents/customer-interviewer.md +93 -0
  23. package/agents/customer-researcher.md +89 -0
  24. package/agents/customer-success-manager.md +90 -0
  25. package/agents/customer-trainer.md +90 -0
  26. package/agents/cxo.md +100 -0
  27. package/agents/demand-creator.md +89 -0
  28. package/agents/demo-mover.md +83 -0
  29. package/agents/demo-specialist.md +89 -0
  30. package/agents/demo-thai-family-law.md +87 -0
  31. package/agents/designer.md +112 -0
  32. package/agents/discovery-caller.md +98 -0
  33. package/agents/doctor.md +139 -0
  34. package/agents/educate-coach.md +100 -0
  35. package/agents/elevate-tutor.md +83 -0
  36. package/agents/email-lifecycle-marketer.md +91 -0
  37. package/agents/engage-specialist.md +99 -0
  38. package/agents/events-coordinator.md +90 -0
  39. package/agents/foundation-builder.md +94 -0
  40. package/agents/funnel-architect.md +89 -0
  41. package/agents/gift-creator.md +100 -0
  42. package/agents/google-ads.md +103 -0
  43. package/agents/guide.md +292 -0
  44. package/agents/helpdesk-dispatcher.md +88 -0
  45. package/agents/hook-specialist.md +91 -0
  46. package/agents/identify-optimizer.md +101 -0
  47. package/agents/implementer.md +305 -0
  48. package/agents/incident-commander.md +120 -0
  49. package/agents/insights-lead.md +93 -0
  50. package/agents/journey-runner.md +113 -0
  51. package/agents/linkedin-ads.md +123 -0
  52. package/agents/live-sales-chat.md +90 -0
  53. package/agents/market-researcher.md +88 -0
  54. package/agents/media-buyer.md +96 -0
  55. package/agents/movers-customer-researcher.md +95 -0
  56. package/agents/movers-foundation-builder.md +96 -0
  57. package/agents/movers-market-researcher.md +97 -0
  58. package/agents/movers-pricing-strategist.md +171 -0
  59. package/agents/nurture-architect.md +99 -0
  60. package/agents/offer-architect.md +92 -0
  61. package/agents/onboarder.md +222 -0
  62. package/agents/onboarding-specialist.md +92 -0
  63. package/agents/operations-dashboard.md +98 -0
  64. package/agents/perf-engineer.md +348 -0
  65. package/agents/playbook-writer.md +71 -0
  66. package/agents/plg-strategist.md +91 -0
  67. package/agents/positioning-architect.md +88 -0
  68. package/agents/press-officer.md +89 -0
  69. package/agents/pricing-strategist.md +94 -0
  70. package/agents/privacy-officer.md +92 -0
  71. package/agents/referral-manager.md +91 -0
  72. package/agents/refine-analyst.md +102 -0
  73. package/agents/release-manager.md +261 -0
  74. package/agents/renewals-upsell-rep.md +90 -0
  75. package/agents/review-engineer.md +327 -0
  76. package/agents/rewards-steward.md +82 -0
  77. package/agents/sales-call-coach.md +94 -0
  78. package/agents/sales-closer.md +98 -0
  79. package/agents/security-auditor.md +313 -0
  80. package/agents/sell-closer.md +101 -0
  81. package/agents/share-amplifier.md +99 -0
  82. package/agents/social-media-manager.md +92 -0
  83. package/agents/storyteller.md +250 -0
  84. package/agents/strategist.md +83 -0
  85. package/agents/strategy-aligner.md +94 -0
  86. package/agents/support-agent.md +90 -0
  87. package/agents/tagger.md +245 -0
  88. package/agents/tech-writer.md +216 -0
  89. package/agents/test-engineer.md +243 -0
  90. package/agents/tiktok-ads.md +105 -0
  91. package/agents/tracking-engineer.md +92 -0
  92. package/agents/upsell-strategist.md +101 -0
  93. package/agents/voice.md +85 -0
  94. package/agents/w4-tools.md +153 -0
  95. package/agents/w4-verify.md +20 -138
  96. package/agents/workflow-optimiser.md +84 -0
  97. package/commands/close.md +814 -160
  98. package/commands/create.md +2 -2
  99. package/commands/deploy.md +554 -175
  100. package/commands/do.md +119 -109
  101. package/commands/fast.md +67 -0
  102. package/commands/improve.md +2 -2
  103. package/commands/one.md +418 -0
  104. package/commands/see.md +1 -1
  105. package/commands/sync.md +1 -1
  106. package/commands/tasks.md +222 -0
  107. package/commands/vespio.md +110 -0
  108. package/commands/vespio.remote.md +227 -0
  109. package/hooks/hooks.json +25 -79
  110. package/hooks/lib/governor-escape-match.sh +88 -0
  111. package/hooks/lib/hook.sh +4 -13
  112. package/hooks/lib/signal.sh +9 -2
  113. package/hooks/scripts/branch-pin.sh +6 -0
  114. package/hooks/scripts/config-protect.sh +6 -0
  115. package/hooks/scripts/do-outcome-gate.sh +19 -2
  116. package/hooks/scripts/git-add-guard.sh +43 -2
  117. package/hooks/scripts/governor-escape.sh +96 -0
  118. package/hooks/scripts/load-guard.sh +127 -0
  119. package/hooks/scripts/post-edit-check.sh +36 -29
  120. package/hooks/scripts/session-start.sh +34 -80
  121. package/hooks/scripts/task-complete-verify.sh +46 -40
  122. package/package.json +1 -1
  123. package/rules/documentation.md +9 -0
  124. package/scripts/ad-copy-lint.sh +656 -0
  125. package/scripts/agentverse-audit.sh +256 -0
  126. package/scripts/asi-walk.sh +435 -0
  127. package/scripts/astro-build-cached.sh +143 -0
  128. package/scripts/blocks-manifest-cached.sh +100 -0
  129. package/scripts/blocks-manifest.mjs +497 -0
  130. package/scripts/blocks-ratchet.sh +116 -0
  131. package/scripts/blocks-render-probe.mjs +529 -0
  132. package/scripts/blocks-usage.mjs +334 -0
  133. package/scripts/chat-context-check.sh +89 -0
  134. package/scripts/close-metrics.sh +558 -0
  135. package/scripts/close-owner.sh +268 -0
  136. package/scripts/db-sync-lock-check.sh +116 -0
  137. package/scripts/deploy-dev.sh +35 -0
  138. package/scripts/deploy-emit.sh +298 -0
  139. package/scripts/deploy-gate-check.sh +128 -0
  140. package/scripts/deploy-ready.sh +78 -0
  141. package/scripts/deploy-record.sh +605 -0
  142. package/scripts/deploy.sh +1273 -0
  143. package/scripts/do-auto.sh +598 -48
  144. package/scripts/do-brief.sh +113 -0
  145. package/scripts/do-close.sh +1137 -0
  146. package/scripts/do-commit.sh +75 -0
  147. package/scripts/do-consumer-sweep.sh +18 -1
  148. package/scripts/do-cycle-shape-check.sh +160 -0
  149. package/scripts/do-decide.sh +476 -0
  150. package/scripts/do-derive-check.sh +436 -0
  151. package/scripts/do-fleet.sh +106 -28
  152. package/scripts/do-folder.sh +10 -1
  153. package/scripts/do-next.sh +106 -0
  154. package/scripts/do-orchestrate.sh +17 -5
  155. package/scripts/do-plan-json.mjs +201 -0
  156. package/scripts/do-plan-json.sh +8 -0
  157. package/scripts/do-preflight.sh +117 -0
  158. package/scripts/do-project.sh +157 -0
  159. package/scripts/do-prove-selftest.sh +108 -0
  160. package/scripts/do-prove.sh +295 -23
  161. package/scripts/do-rank.py +31 -1
  162. package/scripts/do-recon-cache.sh +7 -1
  163. package/scripts/do-recon-pack.sh +196 -0
  164. package/scripts/do-reconcile.sh +121 -4
  165. package/scripts/do-signal.sh +280 -23
  166. package/scripts/do-smoke.sh +18 -1
  167. package/scripts/do-test-gate.sh +80 -0
  168. package/scripts/do-tick.sh +102 -0
  169. package/scripts/do-tier.sh +6 -0
  170. package/scripts/do-triage.sh +182 -0
  171. package/scripts/do-ui-gate.sh +1 -1
  172. package/scripts/do-w4-gates.sh +451 -0
  173. package/scripts/do-walk.sh +12 -1
  174. package/scripts/env-sync.sh +173 -0
  175. package/scripts/factory-brief-check.sh +330 -0
  176. package/scripts/factory-check.sh +68 -14
  177. package/scripts/factory-close-check.sh +257 -0
  178. package/scripts/factory-emit.sh +211 -0
  179. package/scripts/factory-executor-check.mjs +353 -0
  180. package/scripts/factory-peak.sh +301 -0
  181. package/scripts/factory-repo.sh +118 -3
  182. package/scripts/factory-review-check.mjs +61 -0
  183. package/scripts/factory-ship.sh +61 -0
  184. package/scripts/factory-tasks-check.sh +18 -1
  185. package/scripts/factory-turn.sh +326 -0
  186. package/scripts/factory-walk.sh +396 -0
  187. package/scripts/factory-width.sh +57 -0
  188. package/scripts/fade-toxic.sh +4 -3
  189. package/scripts/fixtures/factory-brief-real.md +44 -0
  190. package/scripts/fixtures/triage-dupe.md +5 -0
  191. package/scripts/fleet-manifest.mjs +108 -0
  192. package/scripts/fleet-status.sh +110 -0
  193. package/scripts/full-suite-paths-check.sh +144 -0
  194. package/scripts/gate-reaper-check.sh +98 -0
  195. package/scripts/gate-reaper.sh +125 -0
  196. package/scripts/gate-run.sh +73 -0
  197. package/scripts/gc-content-check.sh +140 -0
  198. package/scripts/gen-dev-config.py +20 -0
  199. package/scripts/govern-bound-check.sh +60 -0
  200. package/scripts/govern-claims-check.sh +233 -0
  201. package/scripts/govern-mem-check.sh +290 -0
  202. package/scripts/governor-doors-check.sh +362 -0
  203. package/scripts/governor-escape-check.sh +171 -0
  204. package/scripts/health.sh +413 -0
  205. package/scripts/id-inventory.mjs +418 -0
  206. package/scripts/land.sh +551 -0
  207. package/scripts/lib/gc-finished.sh +77 -0
  208. package/scripts/lib/govern.sh +361 -0
  209. package/scripts/lib/govern.ts +756 -0
  210. package/scripts/lighthouse-run.sh +187 -0
  211. package/scripts/livekit-live-check.sh +61 -0
  212. package/scripts/livekit-ratchet.sh +46 -0
  213. package/scripts/load-guard-check.sh +49 -0
  214. package/scripts/machine-check.sh +102 -0
  215. package/scripts/machine-watch.sh +177 -0
  216. package/scripts/one-agents.mjs +415 -0
  217. package/scripts/one-resume.sh +78 -0
  218. package/scripts/orphan-baseline.json +182 -0
  219. package/scripts/orphan-modules.mjs +179 -0
  220. package/scripts/pr-body.sh +335 -0
  221. package/scripts/preview-fd-check.sh +289 -0
  222. package/scripts/promise-manifest.mjs +24 -1
  223. package/scripts/release.sh +322 -0
  224. package/scripts/roles-check.sh +946 -0
  225. package/scripts/sdk-build-cached.sh +64 -0
  226. package/scripts/signal-watch.sh +241 -0
  227. package/scripts/skills-publish.sh +94 -0
  228. package/scripts/speed-cache-check.sh +173 -0
  229. package/scripts/speed-check.mjs +907 -0
  230. package/scripts/speed-parity-check.sh +648 -0
  231. package/scripts/speed-waterfall-check.sh +355 -0
  232. package/scripts/substrate-env-parity.mjs +156 -0
  233. package/scripts/tasks-claim-race.mjs +108 -0
  234. package/scripts/tasks-loop.sh +185 -0
  235. package/scripts/test-cached.sh +255 -0
  236. package/scripts/test-full.sh +87 -0
  237. package/scripts/test-honesty.mjs +137 -0
  238. package/scripts/test-lanes.sh +166 -0
  239. package/scripts/test-speed.sh +94 -0
  240. package/scripts/triage-shape-check.sh +149 -0
  241. package/scripts/tsc-cached.sh +179 -0
  242. package/scripts/typedb-flake-check.sh +197 -0
  243. package/scripts/urls-lint.sh +15 -0
  244. package/scripts/verify-fast.sh +445 -0
  245. package/scripts/vespio-sync.sh +149 -0
  246. package/scripts/wf-check.mjs +104 -1
  247. package/scripts/worktree-preview.sh +879 -0
  248. package/scripts/worktree-up.sh +208 -0
  249. package/skills/livekit-agents/SKILL.md +285 -0
  250. package/skills/livekit-agents/references/freshness-rules.md +168 -0
  251. package/skills/shadcn/SKILL.md +1 -1
  252. package/skills/signal/SKILL.md +0 -1
  253. package/skills/voice/SKILL.md +94 -6
  254. package/skills/voice/corpus-check.sh +87 -0
  255. package/hooks/scripts/compact-hint.sh +0 -35
  256. package/hooks/scripts/gate-guard.sh +0 -83
  257. package/hooks/scripts/read-tracker.sh +0 -26
  258. package/hooks/scripts/session-end-verify.sh +0 -51
  259. package/hooks/scripts/stop-reflect.sh +0 -140
  260. package/hooks/scripts/sync-priority-todo.sh +0 -57
  261. package/hooks/scripts/sync-todo-docs.sh +0 -46
  262. package/hooks/scripts/tool-signal.sh +0 -48
  263. package/scripts/do-tasks-bridge.py +0 -366
@@ -4,6 +4,8 @@
4
4
 
5
5
  Ship all five services to Cloudflare. Deterministic sandwich — W0 baseline, build, smoke, approval, parallel deploy, health.
6
6
 
7
+ > **Production door (2026-09-05): `bash .claude/scripts/release.sh promote <sha>` then `ship`** — prod ships from `.release/`, a clean checkout carrying a full-suite receipt for its exact tree. `./deploy` stays the pipeline; `release.sh` is what points it at a tree that cannot be dirty. Loop: `../../CLAUDE.md § The dev → prod loop` · plan: `text/release-path-plan.md`.
8
+ >
7
9
  > **Monorepo (2026-06-04):** all services live in one repo (`Server/one-ie`). Deploy each via `wrangler deploy` from its folder — **never via repo push, no CI**. Deploys are run locally.
8
10
  >
9
11
  > **Production cutover (2026-05-23):** Astro site runs on **CF Workers with Static Assets** (not Pages). Production target is **`https://one.ie`** via the `one-prod` worker. Deploy command is **`wrangler deploy` (no `--env` flag)** — see the deploy-target trap note below.
@@ -14,135 +16,377 @@ Ship all five services to Cloudflare. Deterministic sandwich — W0 baseline, bu
14
16
  > - **Production (live):** `https://one.ie` — `one-prod` worker, deployed from `one.ie/web/` via bare `wrangler deploy` (**not** `--env production` — see trap note above)
15
17
  > - **Gateway (live, stable):** `https://api.one.ie` — `one-gateway` worker
16
18
  > - **Pay gateway (live, stable):** `https://pay.one.ie` — `one-core-worker`, deployed from `pay/backend/` via `bun run deploy` (no `--env` flag; no explicit `[[routes]]` block in its wrangler.toml either — same bare-deploy discipline as `one-prod`). Not part of the original "4 services" naming (found 2026-07-05 after a commit there went unshipped for a full deploy cycle) — treat it as a 5th first-class deploy target, not an afterthought.
17
- > - **Dev (`dev.one.ie` / `one-substrate`) RETIRED 2026-06-04 (prod-only). NEVER a deploy target.** The sibling tree `apps/dev.one.ie` is dead codedo not build or ship it. The CF git-integration that auto-built `one-substrate` on every push to `one-ie/stage` was disconnected (CF dashboard Workers & Pages one-substrate Settings Builds). Nothing deploys via repo push anymore.
19
+ > - **Dev (live again since 2026-08-25):** `https://dev.one.ie` the **`one-dev`** worker, deployed from `one.ie/web/` by `.claude/scripts/deploy-dev.sh` (reachable as **`./deploy dev`**). Its config is *derived from the Astro build output* by `.claude/scripts/gen-dev-config.py` never hand-rolledso it cannot drift from what production ships. That script does two things worth knowing: it renames the worker to `one-dev` + binds `dev.one.ie`, and it **strips every cron**. Dev shares production's D1 and KV bindings, so inheriting prod's schedules would mean two workers running the same handlers against the same rows — double sends, double syncs, races. The first dev deploy shipped 6 schedules including `*/5 * * * *` before this was caught. **Dev observes prod data; it must never also drive prod's clock.**
20
+ > - **`dev.one.ie` is not a data sandbox.** It reads and writes production's rows byte-for-byte. Ship code there freely; treat its DATA as production.
21
+ > - **The RETIRED thing was `one-substrate`**, the old CF-Pages-built dev project whose git integration was disconnected 2026-06-04. That project stays dead — do not deploy it, and do not confuse it with `one-dev`. The sibling tree `apps/dev.one.ie` is still dead code: do not build or ship it. Nothing deploys via repo push.
18
22
  > - **Legacy idle (rollback only — do not deploy):**
19
23
  > - `https://oneie.pages.dev` — old Pages project for `one.ie` ("Ecommerce Playbook" content). Re-attach `one.ie` to this project to roll back.
20
- > - `app.one.ie`, `demo.one.ie`, `onestudio.dev` — still bound to the prior `one-demo` worker. `one-demo` stays in the account untouched for rollback; redeploying `one-demo` would push stale code, so don't.
24
+ > - `demo.one.ie`, `onestudio.dev` — still bound to the prior `one-demo` worker. `one-demo` stays in the account untouched for rollback; redeploying `one-demo` would push stale code, so don't.
25
+ > - `app.one.ie` moved OFF `one-demo` to `one-prod` on 2026-08-06 — it is now an ordinary verified custom domain for `group:one` (D1 `domains` row) and serves `/u/one/*`. Declared in `one.ie/web/wrangler.toml` as a Custom Domain.
21
26
  >
22
27
  > The custom-domain detach step (Pages → Worker) is done manually via the CF API (no in-repo script today).
23
28
 
24
- ## Modes
25
-
26
- | Invocation | What |
27
- |-----------|------|
28
- | `/deploy` | Full pipeline — W0 + build + 5 services + health |
29
- | `/deploy astro` | Astro Worker only (re-bundle after UI changes) |
30
- | `/deploy workers` | Gateway + Sync + Agents only (no Astro rebuild) |
31
- | `/deploy gateway` | Gateway worker only |
32
- | `/deploy sync` | Sync worker only |
33
- | `/deploy agents` | Agents edge agents only |
34
- | `/deploy pay` | Pay gateway (`one-core-worker` → pay.one.ie) only |
29
+ ## Run it — `./deploy`
35
30
 
36
- ## The 8-Step Pipeline
31
+ **`.claude/scripts/deploy.sh` is the authority for the procedure. This doc never
32
+ carries a second copy of the steps** — same discipline as `astro.config.mjs §
33
+ vite.ssr.external`. When a step changes, change the script; edit here only when
34
+ what a gate *means* changes.
37
35
 
38
- > **No root orchestrator exists (confirmed 2026-07-08).** There is no root `package.json` and no `scripts/deploy.ts` in this repo — `bun run deploy` / `bun run verify` at repo root **error with "Script not found"**. The commands below are run manually, per-service, from `Server/one-ie`. If a root orchestrator is ever added, update this note; until then, treat every command in this file as the literal thing to run.
39
-
40
- **Step 1 — W0 Baseline** (no unified script — run per service)
41
36
  ```bash
42
- # Typecheck all 5 services (each has its own tsconfig, no shared root)
43
- (cd one.ie/web && bunx tsc --noEmit)
44
- (cd api && bunx tsc --noEmit)
45
- (cd sync && bunx tsc --noEmit)
46
- (cd channels && bunx tsc --noEmit)
47
- (cd pay/backend && bunx tsc --noEmit)
48
-
49
- # Test suite lives only in one.ie/web
50
- (cd one.ie/web && bunx vitest run)
37
+ ./deploy dev # one.ie/web one-dev (dev.one.ie) FAST gate, no approval
38
+ ./deploy # full pipeline 5 services (PRODUCTION)
39
+ ./deploy astro # one-prod only
40
+ ./deploy workers # api + sync + channels
41
+ ./deploy gateway|sync|agents|pay
42
+ ./deploy --dry-run # print every command, ship nothing
51
43
  ```
52
- Record tests passed/total. Fix before proceeding — never deploy on red, **except** entries on the Known-Flaky list below (currently environment-dependent, not code regressions).
53
44
 
54
- **Step 2 Changes**
55
- `git diff` summary. Flag large changesets.
45
+ `./deploy` is a repo-root wrapper that `exec`s `.claude/scripts/deploy.sh`;
46
+ either path works, and `/deploy` in Claude Code runs the same script rather than
47
+ retyping its steps.
48
+
49
+ ### The two tiers — one command, two destinations
50
+
51
+ | | `./deploy dev` | `./deploy` |
52
+ |---|---|---|
53
+ | Destination | `dev.one.ie` (`one-dev`) | `one.ie` + the other four services |
54
+ | Gate | **FAST lane** (`verify:fast`) | **FULL** (`FULL_VERIFY=1`, every test) |
55
+ | Approval prompt | none | yes, unless `--yes` |
56
+ | Migrations | none | `d1 migrations apply --remote` |
57
+ | Crons | stripped | shipped |
58
+ | Who ships it | agents, finishing a loop | a human, promoting |
59
+
60
+ **`./deploy dev` ships whatever tree it is invoked from — which is `main`.** To
61
+ put a *branch* on dev.one.ie, and to open the PR that proposes it, use
62
+ `bash .claude/scripts/land.sh <branch> --pr --deploy` — one command for
63
+ gate → dev → probe → PR. See "Branch → dev → PR → main" below.
64
+
65
+ `./deploy dev` exits into `deploy-dev.sh` immediately rather than threading a
66
+ flag through the production pipeline — a different destination with a different
67
+ gate is a different procedure, and `deploy-dev.sh` stays its authority. The one
68
+ production flag that carries over is `--skip-tests`, which becomes the dev
69
+ script's `DEV_SKIP_GATE=1`.
70
+
71
+ **A fast pass is never reported as a full pass.** Dev going green says the fast
72
+ lane passed and `dev.one.ie` answered 200 — nothing more. Promotion to
73
+ production runs every test again, because the full suite is the gate that has
74
+ actually caught the regressions.
75
+
76
+ | Mode | Ships |
77
+ |---|---|
78
+ | *(none)* | full pipeline — gates + 5 services + health |
79
+ | `astro` | Astro Worker only (re-bundle after UI changes) |
80
+ | `workers` | Gateway + Sync + Agents (no Astro rebuild) |
81
+ | `gateway` · `sync` · `agents` · `pay` | that one service |
82
+
83
+ Flags: `--check-creds` (self-test the credential ladder, ship nothing) ·
84
+ `--skip-tests` · `--skip-typecheck` · `--skip-build` · `--skip-migrations`
85
+ · `--skip-health` · `--allow-dirty` · `-y/--yes` (or `DEPLOY_YES=1`) ·
86
+ `-n/--dry-run`. Exit 0 = `deploy:success`; a failed probe exits 1 and prints the
87
+ rollback command. Logs: `.deploy-logs/deploy-<stamp>.log` (gitignored) plus one
88
+ per service for the parallel wave.
89
+
90
+ ### What the script does NOT do — and won't
91
+
92
+ Two things stay human, by design:
93
+
94
+ - **Commit / PR / merge** (its own section below). Commit messages and PR bodies
95
+ need judgment. The script's first gate refuses a dirty tree
96
+ (`--allow-dirty` overrides) so the deploy and the history can't disagree.
97
+ - **Known-flaky triage.** On a red suite it prints the failing files and stops,
98
+ pointing at the allowlist. Deciding "that one's the DoH network gap" is a read
99
+ of the evidence, not a rule. The ONE exception is now mechanised: a suite whose
100
+ every failure is the shared TypeDB cluster refusing to answer is classified and
101
+ waived by `typedb-flake-check.sh` — see "The TypeDB flake waiver" below. That
102
+ check is deliberately narrow; everything else is still your read.
103
+
104
+ ## The gates — what each one asserts
105
+
106
+ The script runs these in order. Named here so a failure message means something;
107
+ the commands themselves live in the script.
108
+
109
+ | Gate | Asserts |
110
+ |---|---|
111
+ | **0 · Tree** | working tree clean (`git status --porcelain` empty). First, so a dirty tree fails in a second instead of after a full W0 |
112
+ | **1a · Typecheck** | `bunx tsc --noEmit` clean in all 5 services. Builds `packages/sdk` first when its `dist/` is missing — `one.ie/web` and `channels` resolve SDK types from there, and an unbuilt `dist` fakes a wall of `TS2307` |
113
+ | **1b · Tests** | the full suite in `one.ie/web`, as TWO concurrent lanes — see below. Red blocks, except a classified TypeDB outage; the script won't wave anything else through for you |
114
+ | **3 · Build** | `NODE_ENV=production bun run build` in `one.ie/web` (~30–35s). Emits `dist/server/` + static assets via `@astrojs/cloudflare@13`, patches `wrangler.json` with DO bindings, symlinks `.dev.vars` |
115
+ | **0.4 · Credentials** | Runs **first**, before the slow gates — one curl is cheap, discovering a bad credential after a 2m14s build is not. Resolves ONE credential from an ordered ladder (ambient env → `.env.local` → `one.ie/web/.env` → OAuth), probing each rung with `curl /user` — ground truth, unlike `wrangler whoami`, which answers about whichever credential *that one directory* surfaces. Exports the winner to all five service subshells, so a per-dir `.env` can no longer make two consecutive steps use two different keys. Logs the source as `len=… sha=…`, never bytes, then asserts all five services report the same account id. Self-test: `--check-creds` |
116
+ | **1c · Heavy-gate scheduling** | that vitest and the astro build only overlap when the box can pay for it. Priced in **memory, not cores** — see the section below |
117
+ | **5 · Smoke** | `dist/server/` exists; all 5 `wrangler.toml`s present. Warns if `one.ie/web/wrangler.toml` grew an `[env.production]` block — that's the decoy trap coming back |
118
+ | **6 · Approval** | on `main`, prompts for a literal `yes`. Other branches auto-approve |
119
+ | **6.5 · Migrations** | `wrangler d1 migrations apply DB --remote` — **no `--env`**. "✅ No migrations to apply!" is a pass. Failure blocks: never ship worker code ahead of its schema |
120
+ | **7 · Deploy** | Gateway + Sync + Agents + Pay in parallel (~10s each), then Astro (~30s, largest bundle). Every call bare — no `--env`, ever |
121
+ | **8 · Health** | 4 HTTP probes × 3 tries with backoff, cache-busted with `?_t=`. Sync is cron-only — its clean deploy IS its health signal |
122
+
123
+ ### The suite runs as two lanes, not one run (measured 2026-09-03)
124
+
125
+ The gate had become a coin flip. Four full runs on an **unchanged tree** gave
126
+ four different failure sets — 8, 12, 13, 8 — and every failing file in all four
127
+ was one of the 19 suites marked `real TypeDB`. Not one of the other ~1126 files
128
+ ever failed. Each run's log carried 41-72 `upstream 503` plus 25-38
129
+ `upstream 500`, and those print only *after* the retry budget is spent
130
+ (`substrate.ts:208`).
131
+
132
+ The cause is capacity, not code. `substrate.ts:56-60` already recorded it in
133
+ 2026-07-25: the shared TypeDB Cloud gateway 503s **in bursts lasting seconds**,
134
+ while the retry cover is 500ms + 1000ms. Eight vitest forks against one external
135
+ singleton turn that into a dice roll — and a dice-roll gate costs far more than
136
+ it saves. The 2026-09-03 deploy spent hours on re-runs and root-cause agents for
137
+ failures that were never in the diff.
138
+
139
+ `.claude/scripts/test-lanes.sh` splits the suite by its two binding constraints
140
+ and runs them **concurrently**:
141
+
142
+ | lane | files | bound by | concurrency |
143
+ |---|---|---|---|
144
+ | `pool` | ~1126 | CPU | 8 forks |
145
+ | `typedb` | 19 | one shared gateway | serial (`--no-file-parallelism --maxWorkers=1`) |
146
+
147
+ The `typedb` lane spends its life waiting on a socket, so it costs almost no CPU
148
+ while `pool` saturates the cores: wall clock is **max(), not sum()**.
149
+
150
+ | | before | after |
151
+ |---|---|---|
152
+ | wall clock | 238-282s | **121.9s** |
153
+ | result | RED, rotating victims | **BOTH GREEN** — pool 1125 passed, typedb 19 passed |
154
+ | tests | 8-13 failing | 10940 passed |
155
+
156
+ **This is not `--no-file-parallelism` over the whole suite** — that serialises
157
+ 1126 innocent files to fix 19. Every test still runs, no assertion is weakened,
158
+ nothing is skipped; only the suites sharing the external singleton are
159
+ serialised, which removes the contention at its source.
160
+
161
+ **The selector is read from the tree, never hardcoded.** It greps the
162
+ `real TypeDB` marker the suites already carry, so a new such suite joins the
163
+ serial lane automatically. A frozen list is exactly how this rots back into
164
+ flakiness — `--self-test` asserts the selector finds >0 files, is a strict
165
+ subset, and that every path it names exists.
166
+
167
+ Both lanes still go through `test-cached.sh`, so passes are memoised per lane
168
+ and a RED is still never cached. `test-full.sh` remains the one definition of
169
+ the vitest flags and hands them over as `TEST_FULL_ARGS`.
56
170
 
57
- **Step 2.5 — Commit, PR, Merge**
58
- Tests are green at this point, so it is safe to commit. The shared tree is pinned to `main` (`hook:branch-pin`) — never `git checkout`/`switch` there. Do the commit + branch + PR in a scratch worktree instead, then merge via `gh` and fast-forward the shared tree.
59
171
  ```bash
60
- # 1. From the shared tree (still on main): stage by explicit path, never -A
61
- git status --porcelain # review scope
62
- git diff --stat
63
-
64
- # 2. Cut a worktree for the deploy branch
65
- slug="deploy-$(git rev-parse --short HEAD)"
66
- git worktree add -b "deploy/$slug" ".do-worktrees/$slug" main
67
-
68
- # 3. In the worktree: bring over the same paths, commit, push
69
- cd ".do-worktrees/$slug"
70
- git add <explicit paths from step 2>
71
- git commit -m "$(cat <<'EOF'
72
- deploy: ship <summary of staged changes>
73
- EOF
74
- )"
75
- git push -u origin "deploy/$slug"
76
-
77
- # 4. Open PR, wait for CodeRabbit + any checks, merge
78
- gh pr create --base main --head "deploy/$slug" --title "deploy: <summary>" --body "Automated deploy PR — see /deploy pipeline"
79
- gh pr merge "deploy/$slug" --squash --auto --delete-branch
80
-
81
- # 5. Back in the shared tree: fast-forward and clean up the worktree
82
- cd /Users/toc/Server/one-ie
83
- git switch main # only allowed HEAD move in the shared tree
84
- git pull origin main
85
- git worktree remove ".do-worktrees/$slug"
172
+ bash .claude/scripts/test-lanes.sh --list # show the split, run nothing
173
+ bash .claude/scripts/test-lanes.sh --self-test # prove the selector still selects
86
174
  ```
87
- Generate the commit message and PR title from the diff — conventional commit format (`feat:`, `fix:`, `chore:`, etc.). If there is nothing to commit (`git status --porcelain` is empty), skip this whole step silently — proceed to Step 3 on the tree as-is.
88
175
 
89
- `gh pr merge --auto` queues the merge and returns immediately if required checks (e.g. CodeRabbit review) are still running; poll `gh pr view "deploy/$slug" --json state,mergedAt` before Step 5's `git pull` if the merge hasn't landed yet — don't build off an unmerged branch.
176
+ **Two lanes mean two `Tests N passed` lines in the log.** `TESTS_REPORT` used to
177
+ `tail -1` and reported only the 19-file lane — the first two-lane production
178
+ deploy said `Tests 143 passed` for a run that executed **10940**. It now sums the
179
+ lanes and says `(2 lanes)`. A gate that is fine while the report understates it
180
+ by two orders of magnitude is the same dishonesty as calling a fast pass a full
181
+ one.
90
182
 
91
- **Step 3Build**
92
- ```bash
93
- cd one.ie/web && NODE_ENV=production bun run build
94
- ```
95
- Target: ~30-35s. Emits `dist/server/` + static assets via `@astrojs/cloudflare@13`, then patches `wrangler.json` with DO bindings and symlinks `.dev.vars`. Watch for bundle size warnings.
183
+ ### Making the deploy faster two disproved ideas (measured 2026-09-03)
96
184
 
97
- **Fixed 2026-07-19:** the build was V8-heap-OOMing (`FATAL ERROR: Reached heap limit … JavaScript heap out of memory`, `Abort trap: 6`, exit 134) at Node's default ~4 GiB old-space ceiling — not a system memory shortage (5+ GiB free at the time). `package.json`'s `build` script now sets `NODE_OPTIONS=--max-old-space-size=8192` inline, so plain `bun run build` picks it up automatically; no manual env var needed anymore.
185
+ Recorded so nobody spends an afternoon re-deriving them. Both were plausible,
186
+ both are wrong, and the second fails in a way that reads like a pass.
187
+
188
+ **The gates are at their floor at ~256s. Rescheduling does not move them.**
189
+ From the 2026-09-03 production log:
98
190
 
99
- **Step 4 — Credentials**
100
- `CLOUDFLARE_API_TOKEN` must be unset. Deploy uses `CLOUDFLARE_GLOBAL_API_KEY` only.
101
- Scoped tokens lack permissions for workers + custom domains.
102
- ```bash
103
- unset CLOUDFLARE_API_TOKEN
104
- echo "${CLOUDFLARE_GLOBAL_API_KEY:+set}" "${CLOUDFLARE_EMAIL:+set}" # both must print "set"
191
+ ```
192
+ vitest 256s (the lanes, run alone: 121.9s)
193
+ build 256s (astro's own report: 2m 13s = 133s)
194
+ gates wall-clock: 256s
105
195
  ```
106
196
 
107
- **Step 5Smoke**
108
- Verify `dist/server/` exists, all 5 wrangler configs present: `api/wrangler.toml`, `sync/wrangler.toml`, `channels/wrangler.toml`, `one.ie/web/wrangler.toml`, `pay/backend/wrangler.toml`.
197
+ Serial would be 133+122 = 255s *identical*. Each gate paid ~2x its solo cost
198
+ and the overlap bought nothing, because eight vitest forks plus an 8 GiB-heap
199
+ build on a 10-core box is oversubscription, not parallelism.
109
200
 
110
- **Step 6Approval**
111
- `main` branch: prompts "yes". Other branches: auto-approved.
201
+ The obvious next move give the build room by capping the pool — makes it
202
+ **worse**:
203
+
204
+ | | gates wall-clock | build |
205
+ |---|---|---|
206
+ | 8 forks (baseline) | **256s** | 133s |
207
+ | 5 forks (`VERIFY_POOL_FORKS=5`) | **272s** | 209s |
208
+
209
+ Parallel, serial and capped all land at 250-270s. That is the floor for
210
+ build+suite on this hardware as currently shaped.
211
+
212
+ **The real target is import cost, and the obvious fix is blocked upstream.**
213
+ The pool lane, run solo, spends more time loading modules than running tests:
112
214
 
113
- **Step 6.5 — D1 Migrations**
114
- Run before any worker deploy so schema is current when new code lands:
115
- ```bash
116
- cd one.ie/web && bunx wrangler d1 migrations apply DB --remote
117
215
  ```
118
- **No `--env production`** — same deploy-target trap as the worker deploy itself (see trap note at top of file); `wrangler.toml`'s `[[routes]]`/`[triggers]` were reconciled to top level 2026-07-04 so there is no env-scoped target left to hit, and passing `--env production` here has no matching env block to resolve against.
119
- "✅ No migrations to apply!" is a pass. Any applied migration is logged and counted (e.g. `0195_workspace_sources.sql` applied 2026-07-08, 4.63ms).
120
- Migration failures block deploy — never ship worker code ahead of its schema.
216
+ import 145.42s | tests 107.87s
217
+ ```
218
+
219
+ `vitest.config.ts:131` sets `pool: 'forks'`, and every fork re-imports the whole
220
+ module graph independently. Threads share a module cache, so that ought to be
221
+ the win — but **vitest 4.1.7's threads pool is broken in this repo**. Every
222
+ worker dies with `The worker thread was torn down or never initialized. This is
223
+ a bug in Vitest.` Do not reach for `--pool=threads` until vitest is upgraded and
224
+ this is re-tested.
225
+
226
+ **Read the exit code, not the summary.** That broken run printed:
121
227
 
122
- **Step 7 — Deploy (parallel workers + astro + pay)**
123
- ```bash
124
- (cd api && unset CLOUDFLARE_API_TOKEN && bunx wrangler deploy) &
125
- (cd sync && unset CLOUDFLARE_API_TOKEN && bunx wrangler deploy) &
126
- (cd channels && unset CLOUDFLARE_API_TOKEN && bunx wrangler deploy) &
127
- (cd pay/backend && unset CLOUDFLARE_API_TOKEN && bun run deploy) &
128
- wait
129
- cd one.ie/web && unset CLOUDFLARE_API_TOKEN && bunx wrangler deploy # after the 4 above land
130
228
  ```
131
- Gateway + Sync + Agents + Pay deploy in parallel (~10s each). Astro Worker deploys after (~30s — largest bundle).
229
+ Test Files no tests
230
+ Tests no tests
231
+ Duration 162ms
232
+ ```
233
+
234
+ with `rc=1`. "no tests" in a 162ms run is a *collect crash*, one careless glance
235
+ from being reported as a clean pass — the same shape as
236
+ `vitest-collect-crash-reads-as-pass`. An empty selection is never a pass.
237
+
238
+ **What actually speeds a deploy today:** scope it. `./deploy dev` runs the fast
239
+ lane with no approval; `./deploy workers` skips the astro rebuild entirely when
240
+ `one.ie/web` did not change. The astro build is the long pole and is already
241
+ memoised by tree fingerprint — a warm tree pays none of it.
242
+
243
+ ### Why the gates are not all parallel (measured 2026-08-19)
244
+
245
+ Steps 1+3 used to launch seven heavy processes with a bare `&`: five `tsc`,
246
+ vitest, and the astro build. That is not parallelism, it is a swap storm. The
247
+ astro build carries `--max-old-space-size=8192` and vitest runs a driver plus
248
+ four 1GB forks, so the two together want ~14GB — on a 24GB box where editors
249
+ and sessions are already resident.
250
+
251
+ Paging is a cliff, not a slope. Same suite, same commit, same machine:
252
+
253
+ | vitest ran… | wall-clock | outcome |
254
+ |---|---|---|
255
+ | alone | **176s** | 912 files, 7624 tests, green |
256
+ | beside the build + 5×tsc | **1145s** | 0.0% CPU, no log output, killed |
257
+
258
+ It happened twice in one day before anyone read it as anything but "tests are
259
+ slow". They are not slow — the suite's own summary accounts for only ~165s of
260
+ actual test time.
261
+
262
+ **Correction, same day: memory is NOT the proven cause of the hang.** The
263
+ serialisation below is still worth having — 14GB of overlap on a 24GB box is
264
+ real — but the 1145s runs were later reproduced with the build serialised and
265
+ the box at 41% free, no swap thrash, and no network connections held. See
266
+ "The vitest gate hangs" below. Do not cite the memory story as the explanation
267
+ for a hung gate; it explains a *slow* gate, not a *parked* one.
268
+
269
+ So the heavy gates are now **priced before they launch**: `heavy_free_gb`
270
+ reads free memory the way `lib/govern.sh` does and compares it against
271
+ `DEPLOY_HEAVY_NEED_GB` (default 14). Enough → they overlap as before. Not
272
+ enough → they run one after the other, which is roughly **3x faster
273
+ end-to-end** than "parallel", because the parallel version spends its time
274
+ paging. The five typechecks stay parallel and ungoverned; cached by tree
275
+ fingerprint, they cost ~1s each.
276
+
277
+ Both heavy gates also run under `gate-run.sh`, for its wall-clock **bound**
278
+ (`DEPLOY_GATE_TIMEOUT`, default 900s — a healthy vitest is 176s and a build
279
+ 174s) and its process-group **reap**. macOS ships no `timeout(1)`, and killing
280
+ only the shell reparents the vitest forks to launchd. A hung gate must die on a
281
+ clock rather than outlive the deploy.
282
+
283
+ A probe that cannot read memory returns a large number, so a broken sensor
284
+ never silently serialises the pipeline. Escape hatches: `DEPLOY_HEAVY_PARALLEL=1`
285
+ forces overlap, `DEPLOY_HEAVY_NEED_GB` retunes the threshold.
286
+
287
+ ### The vitest gate hangs — open, characterised, unexplained
288
+
289
+ > **Numbers below are from 2026-08-19 and are superseded as measurements** (the
290
+ > suite was 912 files then and is 1145 now; it runs as two lanes since
291
+ > 2026-09-03, ~122s wall). The *diagnosis* still stands and the hang is still
292
+ > unexplained, so the section is kept as-is rather than half-rewritten. Note the
293
+ > non-TTY fork stall below bit again on 2026-09-03: a background re-run with
294
+ > stdout to a file parked until it was killed, and the `script -qeF /dev/null`
295
+ > wrapper fixed it. That wrapper is not optional for any logged run.
296
+
297
+
298
+ Three times on 2026-08-19 the vitest gate parked indefinitely and had to be
299
+ killed. What is established:
300
+
301
+ - The suite itself is healthy: run on its own it completes in **176s**, 912
302
+ files / 7624 tests / 0 failures, on the same commit and machine.
303
+ - The hang is **not** the astro build competing for memory. It reproduced with
304
+ the heavy gates serialised, `memory_pressure` reporting 41% free and low
305
+ pageouts.
306
+ - It is **not** network. At the moment of the hang the vitest main process held
307
+ **no** open sockets, and neither did its worker.
308
+ - The shape is a fork-pool stall: main parked in `LibuvStreamWrap::OnUvRead`
309
+ waiting on worker IPC, while its single worker sat at **0.43s CPU / 46MB RSS**
310
+ seven minutes in — a fork that was spawned and never given work. The log
311
+ always stops ~90s in, after the jsdom noise, before any test result.
312
+ - `globalSetup` was ruled out: `tests/_global-setup.ts` is local only (reads an
313
+ env file, prints) with no network call to hang on.
314
+
315
+ Not established: why. The one difference between every hung run and the green
316
+ one is that the green run passed `--testTimeout/--hookTimeout/--teardownTimeout`
317
+ explicitly — but nothing in that run came close to a timeout, so that may be
318
+ coincidence rather than cause.
319
+
320
+ Until it is understood, the gate runs under `gate-run.sh`'s bound, so a hang
321
+ now dies on a clock instead of outliving the deploy. If it bites you: the suite
322
+ is trustworthy run directly (`cd one.ie/web && bunx vitest run`), and a deploy
323
+ whose only commits since a green run are harness/doc changes can legitimately
324
+ use `--skip-tests` — verify with
325
+ `git diff --name-only <green-sha>..HEAD | grep -v '^\.claude/'` returning empty.
326
+
327
+ **Health endpoints** (custom domains only — `*.oneie.workers.dev` is blocked on
328
+ this network, curl exit 6/28): `api.one.ie/health` · `one.ie/api/health`
329
+ (assert `"status":"ok"`) · `channels.one.ie/health` · `pay.one.ie/status`
330
+ (assert `"ok"`; there is no `/health` on that entry — it 404s). The agents
331
+ worker is named **`channels`**, reachable at `channels.one.ie`.
332
+
333
+ **Build OOM, fixed 2026-07-19:** the build was V8-heap-OOMing (`FATAL ERROR:
334
+ Reached heap limit`, `Abort trap: 6`, exit 134) at Node's default ~4 GiB
335
+ old-space ceiling — not a system memory shortage. `package.json`'s `build`
336
+ script now sets `NODE_OPTIONS=--max-old-space-size=8192` inline; no manual env
337
+ var needed.
338
+
339
+ **Migrations and the trap:** `wrangler.toml`'s `[[routes]]`/`[triggers]` were
340
+ reconciled to top level 2026-07-04, so there is no env-scoped target left to
341
+ hit and `--env production` has no block to resolve against. Passing it anyway
342
+ silently targets the `one-prod-production` decoy — see the trap note at the top
343
+ of this file.
344
+
345
+ ## Branch → dev → PR → main — `land.sh`, not this doc
346
+
347
+ **`.claude/scripts/land.sh` is the authority for this procedure**, the same way
348
+ `deploy.sh` is the authority for the pipeline. It replaced the hand-rolled
349
+ `git worktree add` + `gh pr create` recipe that used to sit here; that recipe
350
+ was a second copy of a procedure and it rotted — it still told you to
351
+ `git switch main` in the shared tree, which `hook:branch-pin` refuses.
132
352
 
133
- **Step 8 — Health**
134
353
  ```bash
135
- curl https://api.one.ie/health # Gateway
136
- curl https://one.ie/api/health # Astro Worker (production)
137
- curl https://channels.one.ie/health # Agents (channels) — custom domain
138
- curl https://pay.one.ie/ # Pay gateway — 200 on root is the gate
139
- curl https://pay.one.ie/status # …and /status returns status:"ok" (richer; no /health on this entry)
140
- # Sync (one-sync) is a CRON-ONLY worker (no HTTP route) — no health endpoint.
141
- # Its successful `wrangler deploy` ("Deployed one-sync triggers") IS the health signal.
354
+ bash .claude/scripts/land.sh feat/x --pr --deploy --probe / --probe /pricing
142
355
  ```
143
- 3 retries with backoff. The 4 HTTP services must return 200. Astro `/api/health` returns `{"status":"ok","agent":"one-prod","model":…,"hasOpenRouter":true}` — assert `status:"ok"`. (The old `units: N` field was removed; there is no `units` to check anymore.)
144
356
 
145
- > **Network gotcha:** `*.oneie.workers.dev` URLs are **blocked on this network** (curl exit 6/28). Always health-check via the custom domains: `api.one.ie`, `one.ie`, `channels.one.ie`. The agents worker is named **`channels`** and is reachable at **`channels.one.ie`**, not `agents.oneie.workers.dev`.
357
+ One command, four steps, in this order:
358
+
359
+ | # | Step | What it proves |
360
+ |---|---|---|
361
+ | 1 | **gate** — `verify:fast` in the branch's own worktree | the branch is green ON ITS OWN. Not that it survives the trunk — `--pr` deliberately does not merge main in, so the reviewer sees what the branch added and GitHub computes the merge |
362
+ | 2 | **dev** — `$wt/.claude/scripts/deploy-dev.sh` | the branch RUNS. It ships the *worktree's* tree, because `deploy-dev.sh` derives its own ROOT from its own path — `$ROOT`'s copy would ship main and call it the branch |
363
+ | 3 | **probe** — `do-prove.sh`, both bases pinned to dev | the routes answer on dev, under the LANDING RULE. Not `curl / → 200`, which a redirect to `/signin` satisfies |
364
+ | 4 | **PR** — `pr-body.sh` → `gh pr create`/`edit` | a reviewer gets the diff, the trunk drift, the mergeability, the gate label and the dev URL. Re-running UPDATES the open PR, never duplicates it |
365
+
366
+ Then a human merges the PR, and **promotion to `one.ie` is still `./deploy`** —
367
+ the full gate, from `.release/`. Nothing above touches production.
368
+
369
+ Four things that are load-bearing, each of which read as a pass before it was fixed:
370
+
371
+ - **dev.one.ie is one slot, and it writes production's rows.** `--pr --deploy`
372
+ takes exactly one branch; with two, the second overwrites the first while the
373
+ first is being probed. Ship code there freely; treat its DATA as production.
374
+ - **Both probe bases are pinned to dev.** `do-prove.sh` falls back to
375
+ `PROVE_PROD_URL` (default `https://one.ie`) when its dev base is silent — an
376
+ unreachable dev would otherwise prove *production* and report it as the branch
377
+ passing.
378
+ - **The probe's route COUNT is read, not its exit code.** `PROVE: skipped (no
379
+ reachable environment)` exits 0. An unrun probe is not a pass.
380
+ - **The gate is paid once, except under `--quick`.** Step 2 passes
381
+ `DEV_SKIP_GATE=1` because step 1 just ran `verify:fast` in that same tree.
382
+ Under `--quick` step 1 was tsc only — not the fast lane — so `deploy-dev.sh`
383
+ runs its own gate before anything reaches dev.
384
+
385
+ `land.sh` still has its other two doors: bare (merge main in → gate → `--ff-only`
386
+ main → optionally one dev deploy for the batch) and `--pr` alone (gate → PR, no
387
+ dev). The commit itself stays human — a commit message needs judgment, and
388
+ `deploy.sh`'s first gate refuses a dirty tree so the deploy and the history
389
+ cannot disagree.
146
390
 
147
391
  ---
148
392
 
@@ -267,6 +511,7 @@ Do not remove for production builds.
267
511
  | 2026-07-08 | 14.7 MiB | **3.22 MiB** | — | **Exceeds the documented 3 MiB (3072 KiB) free-tier ceiling and still deployed successfully.** Either this account is on a paid Workers plan (10 MiB ceiling) rather than free tier, or the ceiling figure elsewhere in this doc is stale — unconfirmed which. Don't treat "under 3 MiB" as a hard gate until this is resolved; treat 3.2 MiB as the new floor to watch, and re-run Bundle Size Diagnosis if growth continues. |
268
512
  | 2026-07-19 | 20.2 MiB | **4.43 MiB** | — | +37% over 2026-07-08's 3.22 MiB. Deployed successfully — no diagnosis run yet. Growth window covers several merged features that day (newsletter platform, directory-submission, social-formats, movers-playbook, etc.) landing in one `/deploy` cycle; not isolated to a single change. Re-run Bundle Size Diagnosis if the next snapshot keeps climbing. |
269
513
  | 2026-07-29 | 20.9 MiB | **4.60 MiB** | 1.2 MiB | +6% over 2026-07-19, third consecutive climb. Cheap diagnosis WAS run this cycle (the two grep/`ls` commands below, not a full audit): **no single runaway** — Shiki hits 0, top chunks are `_astro_data-layer-content` 2.1 MiB (content collections), `worker-entry` 1.2 MiB (up from 672 KiB at the 2026-05-22 baseline), `index` 1.3 MiB, `icons` 0.8 MiB, `mermaid` 0.8 MiB, `stripe.esm.worker` 0.6 MiB, `react-vendor` 0.5 MiB. Growth is diffuse feature accretion, not a regression; this deploy's own diff was test-only. Two named candidates if a real audit is ever warranted: `mermaid` (0.8 MiB in the SSR bundle — Rule 2 `ssr.external` candidate if it's only reached from `client:only` islands) and the 15 chunks referencing `react-vendor` (Rule 3 suggests some page still SSR-renders React via `client:load`). |
514
+ | 2026-08-02 | 22.0 MiB | **4.96 MiB** | 1.2 MiB | +8% over 2026-07-29, **fourth consecutive climb**. Cheap diagnosis run again: still **no single runaway** — Shiki 0, `react-vendor` referenced by 15 chunks (unchanged), `worker-entry` flat at 1.2 MiB. Top chunks: `_astro_data-layer-content` 2.1 MiB, `index` 1.3 MiB, **`_broadcast_` 1.3 MiB (new to the top list)**, `worker-entry` 1.2 MiB, `index` 0.9 MiB, `icons` 0.8 MiB, `mermaid` 0.8 MiB, `stripe.esm.worker` 0.6 MiB, `react-vendor` 0.5 MiB, **`generateAuthenticationOptions` 0.5 MiB (new)**. This cycle shipped beautiful-blocks (45 visual blocks + 12 background components), which is a plausible share of the delta. Four climbs in a row with the same "diffuse accretion" verdict each time is itself the signal — the cheap diagnosis has now exhausted what it can tell us, and the two standing candidates (`mermaid` via Rule 2, the 15 `react-vendor` chunks via Rule 3) want a real audit rather than a fifth restatement. |
270
515
 
271
516
  The 2026-05-22 regression was caused by `build: { inlineStylesheets: 'always' }`
272
517
  inlining the full Tailwind stylesheet into every route's manifest entry. One
@@ -276,6 +521,9 @@ char change (`always` → `auto`) saved 8.8 MiB.
276
521
 
277
522
  ## Service Map (post-migration)
278
523
 
524
+ `./deploy <mode>` covers every live row; the per-service command is what the
525
+ script runs, recorded here for rollback and one-off work.
526
+
279
527
  | Service | URL | Config | Deploy command |
280
528
  |---------|-----|--------|---------------|
281
529
  | Astro Worker (prod) | `one.ie` → `one-prod` | `one.ie/web/wrangler.toml` | `cd one.ie/web && wrangler deploy` — **no `--env production`** (deploy-target trap: appends `-production` to the script name → `one-prod-production`, which nothing routes to; see trap note at top of this file) |
@@ -284,63 +532,79 @@ char change (`always` → `auto`) saved 8.8 MiB.
284
532
  | Agents | `channels.one.ie` → `channels` (`*.workers.dev` blocked on this network) | `channels/wrangler.toml` | `cd channels && wrangler deploy` |
285
533
  | Pay gateway | `pay.one.ie` → `one-core-worker` | `pay/backend/wrangler.toml` | `cd pay/backend && bun run deploy` (= bare `wrangler deploy`, no `--env` flag) |
286
534
  | Pages (legacy idle, rollback) | `oneie.pages.dev` | — | **do not deploy** — rollback target for `one.ie` |
287
- | Worker (legacy idle, rollback) | `one-demo` (still serves `app.one.ie`, `demo.one.ie`, `onestudio.dev`) | — | **do not deploy** — rollback window for the prod cutover |
535
+ | Worker (legacy idle, rollback) | `one-demo` (still serves `demo.one.ie`, `onestudio.dev`) | — | **do not deploy** — rollback window for the prod cutover |
288
536
 
289
537
  ---
290
538
 
291
539
  ## Auth (CRITICAL — never change)
292
540
 
293
- Always: `CLOUDFLARE_GLOBAL_API_KEY` + `CLOUDFLARE_EMAIL`.
294
541
  Never: `CLOUDFLARE_API_TOKEN` (scoped token lacks workers + custom domain permissions).
295
542
 
543
+ **wrangler reads `CLOUDFLARE_API_KEY` + `CLOUDFLARE_EMAIL`** for global-key auth
544
+ — `CLOUDFLARE_GLOBAL_API_KEY` is *our* name for it and wrangler ignores it. The
545
+ script now exports both spellings off one resolved value, so the two can never
546
+ diverge again.
547
+
548
+ **The credential lives on disk, not in your shell** — `.env.local` at the repo
549
+ root and `one.ie/web/.env` (same 52-char key). You do not need to export
550
+ anything. If you *do* export one, it wins — which is the trap: a **stale**
551
+ export shadows the good key everywhere, because process env beats wrangler's
552
+ per-directory `.env` autoload. That is exactly how 2026-08-19's deploy used two
553
+ different credentials in two consecutive steps (see the Step 6.6 note below).
554
+ The ladder now probes each rung and falls through a rung that does not answer,
555
+ so a stale export costs a log line instead of a red deploy:
556
+
557
+ ```
558
+ rejected: ambient env (len=37 sha=eaafbda9) — /user did not answer 200
559
+ ✓ resolved: global-api-key
560
+ source: /Users/toc/Server/one-ie/.env.local (len=52 sha=435cba23)
561
+ ✓ 5/5 services agree on account 627e0c7c…
562
+ ```
563
+
564
+ Ground truth is the API, in **both** directions — it is as able to prove a key
565
+ alive as dead. Never conclude either from wrangler alone:
566
+
567
+ ```bash
568
+ curl -s -o /dev/null -w '%{http_code}\n' \
569
+ -H "X-Auth-Email: $CLOUDFLARE_EMAIL" -H "X-Auth-Key: $CLOUDFLARE_API_KEY" \
570
+ https://api.cloudflare.com/client/v4/user # 200 = fine
571
+ ```
572
+
296
573
  The deploy script auto-unsets `CLOUDFLARE_API_TOKEN` from the spawned env to prevent
297
574
  accidental use of a scoped token that was exported in the shell.
298
575
 
299
- Required env (export locally before running any deploy step — there is no CI, no root script):
576
+ Required env (export locally before running `./deploy` — there is no CI; the
577
+ script asserts the first two at gate 4 and refuses to ship without them):
300
578
  - `CLOUDFLARE_GLOBAL_API_KEY` + `CLOUDFLARE_EMAIL` — auth
301
579
  - `PUBLIC_GATEWAY_URL: https://api.one.ie` — build-time-inlined by Astro (**required**; without it the Worker bundle falls back to `one-gateway.oneie.workers.dev` and gateway-backed routes break). Lives in `one.ie/web/.env`.
302
580
 
303
581
  ---
304
582
 
305
- ## Steps
306
-
307
- ### `/deploy` (full pipeline)
308
-
309
- 1. Per-service typecheck + `one.ie/web` vitest — W0 gate (see Step 1 above; no unified script exists). Fail here means don't deploy, except Known-Flaky entries.
310
- 2. `git status --porcelain` + `git diff --stat` — surface scope to user. **Stage by explicit path, never `git add -A`** — untracked scratch/backup dirs (e.g. `.wip-backups/`) belong to prior sessions, not this commit.
311
- 3. Commit, PR, merge — worktree branch, push, `gh pr create`, `gh pr merge --squash --auto --delete-branch`, then `git switch main && git pull` on the shared tree (see Step 2.5). Skip entirely if working tree is clean.
312
- 4. `cd one.ie/web && NODE_ENV=production bun run build` Astro production build.
313
- 5. Verify credentials: assert `CLOUDFLARE_GLOBAL_API_KEY` present, unset `CLOUDFLARE_API_TOKEN`.
314
- 6. Smoke check: assert `dist/server/` exists, all 5 wrangler configs present.
315
- 7. Run D1 migrations (Step 6.5): `bunx wrangler d1 migrations apply DB --remote` — no `--env`.
316
- 8. Parallel deploy: Gateway + Sync + Agents + Pay concurrently, then Astro Worker.
317
- 9. Health checks: 4 HTTP endpoints (Gateway, Astro, Channels, Pay via custom domains, cache-busted with `?_t=$(date +%s)`) × 3 retries with backoff; Sync verified by deploy success.
318
- 10. Report:
319
- ```
320
- Branch: main
321
- Tests: 4082/4084 pass (2 known-flaky: SSRF DoH lookup, sandbox-network-blocked)
322
- Typecheck: 5/5 services clean
323
- Build: 33.4s
324
- Migrations: 1 applied (0195_workspace_sources.sql)
325
- Workers: gateway+sync+channels+pay parallel, then astro
326
- Health: 4/4 HTTP 200 (Gateway, Astro, Channels, Pay) + Sync deploy-confirmed
327
- Bundle: gzip 3220 KiB (exceeds documented 3072 KiB free-tier figure, deploy succeeded see Verified Bundle Numbers note)
328
- ```
329
-
330
- ### `/deploy astro`
331
-
332
- Deploys the production worker (`one-prod`) to `one.ie` from the `one.ie/web/` directory.
333
-
334
- 1. `cd one.ie/web && NODE_ENV=production bun run build`
335
- 2. Check bundle size: `du -sh one.ie/web/dist/server/`
336
- 3. If > 12 MiB: check which chunk grew (`ls -lhS one.ie/web/dist/server/chunks/ | head -15`)
337
- 4. Run D1 migrations: `cd one.ie/web && bunx wrangler d1 migrations apply DB --remote` (**no `--env production`** — see Step 6.5 note)
338
- 5. `cd one.ie/web && bunx wrangler deploy` — **no `--env production`** (deploy-target trap: see note at top of this file — `--env production` appends `-production` to the worker name and ships to a script nothing routes to)
339
- 6. Health: `curl -sL "https://one.ie/api/health?_t=$(date +%s)"` (expect 200, `status:"ok"`) — always cache-bust with `?_t=`, see Gotchas § cache-control on errors
340
-
341
- **Resolved 2026-07-04:** `one.ie/web/package.json`'s `"deploy"` script had the same `--env production` trap. Confirmed via the CF API that it was real — neither `one-prod` (live) nor the `one-prod-production` decoy had any cron schedules registered, meaning `billing-alerts-cron.ts`/`billing-allocation-cron.ts`/`billing-autotopup-cron.ts`/`billing-lifecycle-cron.ts`/`billing-verify-cron.ts`/`funnel-aggregate-cron.ts`/`webhook-deliver.ts`/`broadcast-drain-cron.ts` were not running on any schedule. Fixed by reconciling `wrangler.toml`: the `[[routes]]` (custom domain) and `[triggers]` (crons) blocks — the only two things that existed *only* under `[env.production]` — were moved to top level (everything else was already duplicated there); the now-fully-redundant `[env.production.*]` block was deleted entirely, and `package.json`'s script dropped `--env production`. The trap is now structurally impossible — there's no `--env production` target left to hit.
342
-
343
- **Pre-flight (one-time, on cutover only):** ensure no other CF entity owns the `one.ie` custom domain. If wrangler errors with a hostname conflict, detach the prior owner first:
583
+ ## Mode-specific notes
584
+
585
+ Everything common lives in the script. These are the bits that are true of one
586
+ mode only.
587
+
588
+ ### `./deploy astro`
589
+
590
+ **Resolved 2026-07-04:** `one.ie/web/package.json`'s `"deploy"` script had the
591
+ same `--env production` trap. Confirmed via the CF API that it was real —
592
+ neither `one-prod` (live) nor the `one-prod-production` decoy had any cron
593
+ schedules registered, meaning `billing-alerts-cron.ts` /
594
+ `billing-allocation-cron.ts` / `billing-autotopup-cron.ts` /
595
+ `billing-lifecycle-cron.ts` / `billing-verify-cron.ts` /
596
+ `funnel-aggregate-cron.ts` / `webhook-deliver.ts` / `broadcast-drain-cron.ts`
597
+ were not running on any schedule. Fixed by reconciling `wrangler.toml`: the
598
+ `[[routes]]` (custom domain) and `[triggers]` (crons) blocks — the only two
599
+ things that existed *only* under `[env.production]` were moved to top level
600
+ (everything else was already duplicated there); the now-fully-redundant
601
+ `[env.production.*]` block was deleted entirely, and `package.json`'s script
602
+ dropped `--env production`. The trap is now structurally impossible — there's
603
+ no `--env production` target left to hit.
604
+
605
+ **Pre-flight (one-time, on cutover only):** ensure no other CF entity owns the
606
+ `one.ie` custom domain. If wrangler errors with a hostname conflict, detach the
607
+ prior owner first:
344
608
 
345
609
  ```bash
346
610
  # If a Pages project owns it (was `oneie` project pre-cutover):
@@ -350,33 +614,25 @@ curl -s -X DELETE \
350
614
  -H "X-Auth-Key: $CLOUDFLARE_GLOBAL_API_KEY"
351
615
  ```
352
616
 
353
- ### `/deploy workers`
354
-
355
- Verified 2026-07-08: all three complete independently with no shared state, so
356
- run them in parallel (this matches Step 7 of the full pipeline).
357
-
358
- ```bash
359
- (cd api && unset CLOUDFLARE_API_TOKEN && bunx wrangler deploy) &
360
- (cd sync && unset CLOUDFLARE_API_TOKEN && bunx wrangler deploy) &
361
- (cd channels && unset CLOUDFLARE_API_TOKEN && bunx wrangler deploy) &
362
- wait
363
- ```
364
-
365
- Report: version hash per worker, health latency per service.
366
-
367
- ### `/deploy pay`
617
+ ### `./deploy workers`
368
618
 
369
- Deploys the pay gateway (`one-core-worker`) to `pay.one.ie` from the `pay/backend/` directory.
619
+ Verified 2026-07-08: gateway, sync and channels complete independently with no
620
+ shared state — which is why the script runs them concurrently.
370
621
 
371
- 1. `cd pay/backend && bunx tsc --noEmit` (no dedicated `typecheck` script — run `tsc` directly)
372
- 2. `unset CLOUDFLARE_API_TOKEN && bun run deploy` (= bare `wrangler deploy`, no `--env` flag)
373
- 3. Health — both probed live 2026-08-02, both **200**, neither authenticated:
374
- - Gate: `curl -sL -o /dev/null -w "%{http_code}" https://pay.one.ie/` (expect 200)
375
- - Richer: `curl -sL https://pay.one.ie/status` → assert `status:"ok"` (also reports version, destinationMode, contract addresses)
622
+ ### `./deploy pay`
376
623
 
377
- `pay/backend/src/index.ts` mounts `pay/backend/src/routes/status.ts` (`/status`) and `discoveryRoutes` (`/`). It does **not** serve `/health` — the `/health` handler in `pay/backend/src/api/routes/status.ts` belongs to the separate, unmounted `src/api/` tree. Do not health-check `/health` here; it 404s.
624
+ `pay/backend/src/index.ts` mounts `src/routes/status.ts` (`/status`) and
625
+ `discoveryRoutes` (`/`). It does **not** serve `/health` — the `/health`
626
+ handler in `pay/backend/src/api/routes/status.ts` belongs to the separate,
627
+ unmounted `src/api/` tree. Both `/` and `/status` probed live 2026-08-02: 200,
628
+ unauthenticated. `/status` also reports version, destinationMode, and contract
629
+ addresses.
378
630
 
379
- **Found 2026-07-05:** `pay/backend` shipped a real commit (`feat(pay): embeddable payment-link page`) that sat unshipped through a full `/deploy` cycle because the skill's service map only named 4 services. `pay.one.ie` is a first-class 5th deploy target, not a manual afterthought — check `git log` scoped to `pay/` for unshipped commits on every `/deploy` run, the same as the other 4 services.
631
+ **Found 2026-07-05:** `pay/backend` shipped a real commit (`feat(pay):
632
+ embeddable payment-link page`) that sat unshipped through a full deploy cycle
633
+ because the service map only named 4 services. `pay.one.ie` is a first-class
634
+ 5th target, not an afterthought — check `git log` scoped to `pay/` for
635
+ unshipped commits, same as the other four.
380
636
 
381
637
  ## Bundle Size Diagnosis
382
638
 
@@ -425,11 +681,63 @@ previous deployment — it needs no worktree and no rebuild.
425
681
 
426
682
  ---
427
683
 
684
+ ## The TypeDB flake waiver — a red suite the deploy may ship past
685
+
686
+ `one.ie/web`'s suite talks to a REAL shared TypeDB Cloud cluster (CLAUDE.md:
687
+ "Don't mock TypeDB in integration tests"). When that cluster blips or a query
688
+ outruns its timeout, a handful of task/substrate suites go red without anything
689
+ in the diff being wrong. That used to be an eyeball judgement, which is exactly
690
+ the call that gets rubber-stamped on the fifth deploy attempt at 2am.
691
+
692
+ `.claude/scripts/typedb-flake-check.sh` makes it a check. When the vitest gate
693
+ goes red, `deploy.sh` runs it against the gate log:
694
+
695
+ | exit | means |
696
+ |---|---|
697
+ | 0 | every failure carries a substrate-unavailable signature — waivable, deploy continues |
698
+ | 1 | at least one failure is real, or the log could not be classified — deploy stops |
699
+
700
+ **It keys on the failure SIGNATURE, never on the filename.** A file-based
701
+ allowlist waives every future failure in that file, including the real ones. The
702
+ waived signatures are `upstream_50[234]`, `status=50[234]`,
703
+ `fixture write failed`, `Test timed out in Nms`, `ETIMEDOUT`, `ECONNRESET`,
704
+ `ECONNREFUSED`, `EAI_AGAIN`, `socket hang up`, `fetch failed`, and TypeDB
705
+ connection errors.
706
+
707
+ Four properties, each with a red half in `--self-test`:
708
+
709
+ - **`not_found` is never waivable.** Checked first, independently of everything
710
+ else. It is the `tasks:claim` privilege boundary and four separate real
711
+ defects have presented as that exact string. A `not_found` wrapped inside a
712
+ 503 still blocks.
713
+ - **One flake never vouches for its neighbour.** The log is split into vitest's
714
+ per-failure blocks and EVERY block must carry a signature. A genuine assertion
715
+ break standing beside a 503 blocks the deploy.
716
+ - **Silence is not a pass.** An empty or unparseable log, or one with no
717
+ `Tests N failed` line, exits 1.
718
+ - **A waiver is not a green suite.** The report says
719
+ `WAIVED as TypeDB outage (suite NOT green)`, and a waived run **cannot settle
720
+ deferred-pin debt** — the waiver speaks only to the failures that reported, not
721
+ to a pin that never got to.
722
+
723
+ On by default. `--no-typedb-flake-waiver` (or `DEPLOY_ALLOW_TYPEDB_FLAKE=0`)
724
+ restores the hard stop. Prove the checker still bites before trusting it:
725
+
726
+ ```bash
727
+ bash .claude/scripts/typedb-flake-check.sh --self-test # 6 cases, 4 of them red halves
728
+ bash .claude/scripts/typedb-flake-check.sh <a-gate-log>
729
+ ```
730
+
731
+ **The waiver is not a diagnosis.** `curl -sS -o /dev/null -w '%{http_code}' https://api.one.ie/health`
732
+ returning 200 while the suite reports 503s means the cluster blipped mid-run. A
733
+ 200 alongside failures that are NOT in the signature list means the code is
734
+ wrong — and the checker will tell you so.
735
+
428
736
  ## Known-Flaky Test Allowlist
429
737
 
430
- No allowlist is enforced in code (`scripts/deploy.ts` doesn't exist in this repo — see
431
- pipeline note above). Treat these by name when they appear in `bunx vitest run` output
432
- in `one.ie/web`; don't block deploy on them, but don't silently ignore new failures either
738
+ `deploy.sh` deliberately enforces no allowlist a red suite stops it, and it
739
+ prints the failing files with a pointer here. Triage is yours: treat these by
740
+ name when they appear in `bunx vitest run` output in `one.ie/web`; don't block deploy on them, but don't silently ignore new failures either
433
741
  — confirm the failure signature matches before waving it through:
434
742
 
435
743
  - `tests/e2e/c5-webhook-subscribe.test.ts` — `workflow:webhook-subscribe` "writes a KV
@@ -441,6 +749,19 @@ in `one.ie/web`; don't block deploy on them, but don't silently ignore new failu
441
749
  waving through: `curl -v --max-time 5 https://1.1.1.1/dns-query 2>&1 | grep -i refused`
442
750
  — if that shows "Connection refused", it's this network gap, not a code regression. If it
443
751
  connects fine and the test still fails, it's real — investigate.
752
+ - `tests/unit/tasks-humans.test.ts` + `tests/tasks-do-roundtrip.test.ts` — **only** when
753
+ the failure is `fixture write failed (status=503 error=upstream_503)` or
754
+ `(status=502|504 …)`. That message means the shared TypeDB Cloud cluster was
755
+ unavailable through every retry `typedbQueryDetail` already performs — the substrate
756
+ refused setup, so nothing downstream proved anything. Verify before waving through:
757
+ `curl -sS -o /dev/null -w '%{http_code}' https://api.one.ie/health` — a 200 there with a
758
+ 503 in the test means the cluster blipped during the run, not that the code is wrong.
759
+ **Any OTHER failure in these two files is real and blocks deploy** — in particular a bare
760
+ `not_found` on `tasks:claim`, which is the privilege boundary and must never be waved
761
+ through. Four separate defects that used to present as that same `not_found` were fixed
762
+ 2026-08-02 (concurrent-run sweep destruction, silent fixture writes, same-attribute
763
+ insert races, read-after-write lag); if it reappears, something new is wrong. History:
764
+ the commit message on `140975e44` and `tests/helpers/probe-sweep.ts`.
444
765
  - Hardware/stochastic benchmarks (speed, distribution-timing tests) — expected variance,
445
766
  not correctness bugs.
446
767
 
@@ -451,8 +772,9 @@ diagnose and fix before proceeding.
451
772
 
452
773
  ## First-Time Setup
453
774
 
454
- Only needed once. There is no `docs/deploy.md` this block plus the 8-Step
455
- Pipeline above is the whole walkthrough. Resource names below are the ones
775
+ Only needed once, before `./deploy` can work at all. There is no
776
+ `docs/deploy.md` this block plus the script is the whole walkthrough.
777
+ Resource names below are the ones
456
778
  actually declared in `one.ie/web/wrangler.toml`; creating differently-named
457
779
  resources produces bindings the Worker can't resolve.
458
780
 
@@ -475,16 +797,17 @@ bunx wrangler secret put TYPEDB_PASSWORD # paste at the prompt; never inline t
475
797
  cd ..
476
798
 
477
799
  # First deploy — Worker auto-provisions on first `wrangler deploy`
478
- # (see "The 8-Step Pipeline" above for the actual per-service commands — no root script)
800
+ # Then just: ./deploy
479
801
  ```
480
802
 
481
803
  ---
482
804
 
483
805
  ## Logs
484
806
 
485
- No `.deploy.log` / `.deploy-build.log` files are written today there is no wrapper
486
- script to produce them (see pipeline note at top of file). Output only lives in each
487
- `wrangler deploy` command's own stdout; capture it yourself if a durable record is needed.
807
+ `./deploy` writes `.deploy-logs/deploy-<stamp>.log` (gitignored)every gate's
808
+ stdout in order plus `.deploy-logs/<service>-<stamp>.log` for each service in
809
+ the parallel wave, since their output would otherwise interleave. The report at
810
+ the end names the log path.
488
811
 
489
812
  Live logs:
490
813
 
@@ -498,6 +821,8 @@ cd one.ie/web && bunx wrangler deployments list --name one-prod | head -10
498
821
 
499
822
  ## Gotchas
500
823
 
824
+ - **A service with no local `wrangler` falls through to a shared `bunx wrangler@latest` cache, and that cache can rot.** Hit 2026-08-06: `sync/` was the only one of the five without wrangler in its `devDependencies`, so `bunx wrangler deploy` resolved to `$TMPDIR/bunx-501-wrangler@latest/` — whose install was missing `esbuild`, so it died `MODULE_NOT_FOUND` before reading a single config. The other four were unaffected because they resolve wrangler from their own `node_modules`. Fixed by pinning `wrangler` into `sync/package.json` like its siblings. If this shape reappears elsewhere, the workaround is a version-pinned invocation (`bunx wrangler@4.80.0 deploy`, which lands in a different cache dir); the fix is a local dep.
825
+ - **Piping `./deploy` into `tail`/`head` masks its exit code** — the pipeline reports the pager's status, not the script's, so a failed run reads as exit 0. `die()` really does `exit 1`; read the ✓/✗ lines or the `.deploy-logs/` file, and don't infer success from a piped exit status.
501
826
  - TypeDB Cloud port is **1729** (not 80 or 443)
502
827
  - TypeDB HTTP API prefix is `/v1/` (signin, query, databases)
503
828
  - Always `CLOUDFLARE_GLOBAL_API_KEY` — scoped tokens lack permissions for workers + custom domains
@@ -510,3 +835,57 @@ cd one.ie/web && bunx wrangler deployments list --name one-prod | head -10
510
835
  ---
511
836
 
512
837
  *Deploy is the closed loop. W0 baseline in, health check out. If health fails, mark() is blocked. Determinism: every step reports numbers, every number gets marked.*
838
+
839
+ ### `code: 9103` does not mean the key was revoked
840
+
841
+ Measured 2026-08-18, and it cost an hour plus a wrong accusation that a security
842
+ incident had rotated the key. `Unknown X-Auth-Key or X-Auth-Email [code: 9103]`
843
+ at Step 6.5 was a **variable-name mismatch**: the script exported
844
+ `CLOUDFLARE_GLOBAL_API_KEY`, wrangler only reads `CLOUDFLARE_API_KEY`.
845
+
846
+ Before concluding a Cloudflare credential is dead, ask the API directly — it is
847
+ ground truth and wrangler is not:
848
+
849
+ ```bash
850
+ curl -s -o /dev/null -w '%{http_code}\n' \
851
+ -H "X-Auth-Email: $CLOUDFLARE_EMAIL" -H "X-Auth-Key: $CLOUDFLARE_API_KEY" \
852
+ https://api.cloudflare.com/client/v4/user # 200 = the key is fine
853
+ ```
854
+
855
+ Three companion traps, all real:
856
+
857
+ - **wrangler 4.x auto-loads `.env` from the cwd**, so a stale key in
858
+ `one.ie/web/.env` beats both your exported vars and an OAuth session. Isolate a
859
+ credential test by running it from `/tmp`.
860
+ - **`/user/tokens/verify` is Bearer-only** — `400 Missing "Authorization" header`
861
+ there is not evidence against a global key. Use `/user` or `/accounts`.
862
+ - **Length proves nothing.** A classic global key is 37 hex chars; a `cfk_`-prefixed
863
+ one is ~52 and equally valid.
864
+
865
+ **Never edit `.env` while a deploy is running.** Doing so took a run fully red —
866
+ tsc, vitest, and the astro build — for reasons that had nothing to do with the code.
867
+
868
+ ### The 2026-08-19 sequel: `code: 7403` at Step 6.6, and two keys
869
+
870
+ The same family, one layer deeper, and worth reading before you diagnose any
871
+ Cloudflare auth failure here. Step 6.5 (D1 in `one.ie/web`) **passed** and Step
872
+ 6.6 (D1 in `channels`) **failed** in the same run, seconds apart, on the same
873
+ account. That is only possible if they used different credentials — and they
874
+ did:
875
+
876
+ - **Three** credentials were reachable from one `./deploy`: a stale 37-char key
877
+ in the ambient env (injected by `~/.claude/settings.json`'s `env` block — so
878
+ it existed inside Claude Code sessions and *not* in a plain terminal, which is
879
+ why grepping the shell profiles found nothing), the good 52-char key in
880
+ `one.ie/web/.env` + `.env.local`, and the `wrangler login` OAuth session.
881
+ - Nothing *chose* between them. Each service got whatever its own cwd surfaced.
882
+ Of the five dirs, only `one.ie/web` has a `.env` — so it read the good key and
883
+ passed, while `channels` fell through to OAuth and hit `7403`.
884
+ - **`7403` ≠ `9103`.** `9103` ("Unknown X-Auth-Key") points at a key; `7403`
885
+ ("account is not authorized to access this service") points at an account
886
+ scope, i.e. a session. Reading them as the same symptom is what sent the first
887
+ diagnosis at a perfectly good key.
888
+
889
+ Fixed by making the credential *resolved* rather than *ambient* — see Step 0.4
890
+ above. Run `./deploy --check-creds` if you ever doubt which key is in play; it
891
+ names the source and proves all five services agree.