@rizom/ops 0.2.0-alpha.44 → 0.2.0-alpha.440

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (70) hide show
  1. package/README.md +82 -1
  2. package/dist/age-key-bootstrap.d.ts +1 -1
  3. package/dist/brains-ops.js +736 -335
  4. package/dist/capability-bundle-migration.d.ts +22 -0
  5. package/dist/cert-bootstrap.d.ts +4 -2
  6. package/dist/content-repo-ref.d.ts +10 -0
  7. package/dist/content-repo.d.ts +1 -1
  8. package/dist/deploy.js +110 -166
  9. package/dist/directory-sync-stress-system.d.ts +75 -0
  10. package/dist/directory-sync-stress.d.ts +117 -0
  11. package/dist/entries/deploy.d.ts +4 -2
  12. package/dist/health-watchdog-smoke.d.ts +54 -0
  13. package/dist/image-inventory.d.ts +4 -0
  14. package/dist/image-types.d.ts +7 -0
  15. package/dist/images.d.ts +82 -0
  16. package/dist/index.d.ts +8 -0
  17. package/dist/index.js +736 -323
  18. package/dist/legacy-pilot-migration.d.ts +11 -0
  19. package/dist/legacy-projection-job-recovery.d.ts +29 -0
  20. package/dist/load-registry.d.ts +61 -6
  21. package/dist/observed-status.d.ts +1 -1
  22. package/dist/parse-args.d.ts +2 -12
  23. package/dist/preview-domain.d.ts +9 -0
  24. package/dist/reconcile-dry-run.d.ts +9 -0
  25. package/dist/run-command.d.ts +20 -4
  26. package/dist/schema.d.ts +126 -163
  27. package/dist/secrets-encrypt.d.ts +7 -13
  28. package/dist/secrets-push.d.ts +1 -1
  29. package/dist/ssh-key-bootstrap.d.ts +1 -26
  30. package/dist/stage-legacy-crossover.d.ts +25 -0
  31. package/dist/stress-command.d.ts +14 -0
  32. package/dist/stress-git-checkout.d.ts +21 -0
  33. package/dist/stress-health-monitor.d.ts +46 -0
  34. package/dist/upgrade.d.ts +10 -0
  35. package/dist/user-offboard.d.ts +61 -0
  36. package/dist/verify-user.d.ts +22 -0
  37. package/package.json +50 -42
  38. package/templates/rover-pilot/.env.schema +20 -5
  39. package/templates/rover-pilot/.github/actions/varlock-env/action.yml +47 -0
  40. package/templates/rover-pilot/.github/workflows/build.yml +70 -17
  41. package/templates/rover-pilot/.github/workflows/deploy.yml +74 -57
  42. package/templates/rover-pilot/.github/workflows/directory-sync-stress.yml +119 -0
  43. package/templates/rover-pilot/.github/workflows/health-watchdog-smoke.yml +95 -0
  44. package/templates/rover-pilot/.github/workflows/offboard.yml +107 -0
  45. package/templates/rover-pilot/.github/workflows/reconcile.yml +10 -4
  46. package/templates/rover-pilot/.github/workflows/upgrade.yml +104 -0
  47. package/templates/rover-pilot/README.md +13 -5
  48. package/templates/rover-pilot/deploy/scripts/create-predeploy-backup.ts +936 -0
  49. package/templates/rover-pilot/deploy/scripts/decrypt-user-secrets.ts +80 -24
  50. package/templates/rover-pilot/deploy/scripts/helpers.ts +2 -0
  51. package/templates/rover-pilot/deploy/scripts/install-health-watchdog.ts +144 -0
  52. package/templates/rover-pilot/deploy/scripts/provision-server.ts +51 -21
  53. package/templates/rover-pilot/deploy/scripts/resolve-deploy-handles.ts +4 -0
  54. package/templates/rover-pilot/deploy/scripts/resolve-missing-images.ts +13 -0
  55. package/templates/rover-pilot/deploy/scripts/resolve-user-config.ts +40 -9
  56. package/templates/rover-pilot/deploy/scripts/sync-content-repo.ts +51 -47
  57. package/templates/rover-pilot/deploy/scripts/update-dns.ts +72 -16
  58. package/templates/rover-pilot/deploy/scripts/validate-secrets.ts +12 -1
  59. package/templates/rover-pilot/deploy/scripts/verify-runtime-image.ts +18 -0
  60. package/templates/rover-pilot/docs/canonical-crossover-record.md +107 -0
  61. package/templates/rover-pilot/docs/onboarding-checklist.md +23 -14
  62. package/templates/rover-pilot/docs/operator-playbook.md +357 -29
  63. package/templates/rover-pilot/docs/user-onboarding.md +48 -463
  64. package/templates/rover-pilot/package.json +1 -0
  65. package/templates/rover-pilot/pilot.yaml +6 -4
  66. package/dist/origin-ca.d.ts +0 -1
  67. package/dist/push-secrets.d.ts +0 -9
  68. package/dist/push-target.d.ts +0 -2
  69. package/dist/run-subprocess.d.ts +0 -6
  70. package/templates/rover-pilot/.kamal/hooks/pre-deploy +0 -9
@@ -9,30 +9,160 @@ Treat these as checked-in deploy artifacts in the pilot repo:
9
9
  - `deploy/scripts/`
10
10
  - `.github/workflows/build.yml`
11
11
  - `.github/workflows/deploy.yml`
12
+ - `.github/workflows/directory-sync-stress.yml`
13
+ - `.github/workflows/health-watchdog-smoke.yml`
14
+ - `.github/workflows/offboard.yml`
12
15
  - `.github/workflows/reconcile.yml`
13
16
 
14
17
  `.env.schema` is the single source of truth for required and sensitive deploy vars.
15
18
  The deploy scripts and workflows should read from that contract instead of inventing a second list.
16
19
 
17
- The shared pilot image tag is `brain-${brainVersion}`:
20
+ The fleet has one image topology:
18
21
 
19
- - build publishes `brain-${brainVersion}`
22
+ - one immutable `brain-${brainVersion}` image is published for each effective Brain version
23
+ - every new image contains the union of exact site/theme package pins across the whole fleet, regardless of current Brain versions
24
+ - conflicting versions of one package fail image resolution before build
20
25
  - generated `users/<handle>/.env` carries `BRAIN_VERSION=<brainVersion>`
21
- - deploy sets `VERSION=brain-${brainVersion}`
26
+ - build and deploy derive the same effective image tag from the resolved registry
22
27
 
23
28
  ## Version bump flow
24
29
 
25
30
  When `pilot.yaml.brainVersion` changes and you push:
26
31
 
27
- 1. build publishes the new shared image tag
32
+ 1. build publishes each missing version image with the full fleet package union, and verifies existing images against their assigned instances
28
33
  2. reconcile refreshes generated `users/<handle>/.env`
29
34
  3. deploy runs for handles whose generated config changed
30
35
  4. generated file commits happen once in a final aggregation step after the deploy matrix finishes
31
36
 
37
+ Every external site and theme package has its own exact version pin. A cohort or
38
+ pilot brain-version bump never changes those package versions implicitly; update each
39
+ pin deliberately from reviewed package and image evidence.
40
+
41
+ Smoke-first upgrades build the full fleet package union before promotion. Moving
42
+ other cohorts onto the tested Brain version then reuses that same immutable image.
43
+ Explicit Build dispatches also include the fleet union; `site_packages` only adds
44
+ extra pins and cannot replace or conflict with declared pins.
45
+
46
+ Deploy verifies the actual installed Brain and selected instance's site/theme
47
+ versions on the CI runner after image readiness and before provisioning or
48
+ container replacement. It reads manifests in a read-only, network-disabled
49
+ container without starting the Brain or mounting fleet data. Missing packages,
50
+ wrong versions, or verification failures block deployment, even on manual runs.
51
+ Never bypass this gate or overwrite a deployed tag to recover an older incomplete
52
+ image; review artifact recovery separately.
53
+
32
54
  When a push changes only deploy contract files and no generated `users/<handle>/.env` or `users/<handle>/brain.yaml` files, the deploy workflow exits through its explicit no-op path and prints `No affected user configs; skipping deploy.`
33
55
 
34
56
  They are scaffolded from `@rizom/ops`, then versioned in this repo like any other deploy contract.
35
57
 
58
+ ## Verified pre-deploy rollback snapshots
59
+
60
+ Deploy creates and verifies one target-scoped rollback snapshot after SSH and image readiness, immediately before optional stale-lock release and `kamal setup`. A snapshot failure stops the workflow before container replacement. The current runtime must be ready with an idle job queue; degradation a plugin reports is named in the log and backed up, not refused, because a deploy is often its fix. A genuinely new server with no runtime or persistent state reports `not applicable`; persistent state without an identifiable runtime fails closed.
61
+
62
+ Each snapshot contains transaction-consistent online captures of the six canonical SQLite databases, database hashes and `quick_check` results, the deployed `brain.yaml`, sanitized container and mount metadata, and the Directory Sync checkout. Git capture includes all refs, an observed remote head, staged and unstaged binary patches, and mode-preserving archives of untracked and ignored files. Capture performs no checkout mutation or remote write.
63
+
64
+ Verified snapshots are finalized atomically under:
65
+
66
+ ```text
67
+ /opt/brain-state/backups/predeploy-<handle>-<target>-<UTC timestamp>/
68
+ ```
69
+
70
+ Files use mode `0600` and the directory uses `0700`. A `.incomplete` directory is diagnostic evidence, never a rollback point. Checksums and the versioned manifest must verify before finalization. The default retention is five verified snapshots per target; incomplete or unverified directories do not count, and the snapshot created by the current deployment is never pruned.
71
+
72
+ These are same-server rollback snapshots, not off-host disaster recovery. There is no normal skip and no automatic restore. Restoration requires separate approval: stop replacement/application processes, recheck the selected snapshot's checksums, preserve the current state, restore databases and exact Git state, then validate health, queues, Git checkpoints, durable export intents, preview, and production output before reopening the target.
73
+
74
+ ## Retired projection job recovery
75
+
76
+ A pre-scheduler projection job can remain active after its handler type is retired. Current runtime health reports active job types missing from the finalized execution inventory as degraded. Do not restart repeatedly or relax the snapshot idle gate.
77
+
78
+ Use the recovery command only after read-only evidence proves all of the following for one exact row:
79
+
80
+ - its type is one of the command's fixed legacy projection types;
81
+ - the installed runtime no longer contains that handler;
82
+ - attempt, worker-session, lease, and heartbeat ownership are all absent;
83
+ - no durable progress snapshot or result exists.
84
+
85
+ Preview the exact row first:
86
+
87
+ ```sh
88
+ bunx --package @rizom/ops@<exact-version> brains-ops \
89
+ recover:retire-legacy-projection-job \
90
+ /data/brain-jobs.db <job-id> --type <legacy-type> --dry-run
91
+ ```
92
+
93
+ After separate operator review, replace `--dry-run` with the exact confirmation `--confirm retire:<job-id>`. The command atomically fences against ownership or progress appearing between inspection and retirement, and fails closed if the row changed. It is not a general job cancellation API and cannot retire arbitrary job types. Require operational health and a fully idle queue afterward, then run the unchanged canonical predeploy snapshot and Deploy workflow.
94
+
95
+ ## Canonical contract crossover maintenance window
96
+
97
+ Do not run this procedure without explicit operator approval. The canonical desired state, canonical `@rizom/ops`, and unified runtime image form one contract and must move or roll back together. Complete `docs/canonical-crossover-record.md` as the approval evidence without adding secret values.
98
+
99
+ Before the window, record and review:
100
+
101
+ - the prior pilot commit and exact `@rizom/ops` version;
102
+ - every prior runtime image tag and immutable digest;
103
+ - the reviewed canonical pilot commit;
104
+ - the exact unified `@rizom/brain` and `@rizom/ops` versions;
105
+ - every unified image tag and immutable digest;
106
+ - canary-first rollout order, followed by the remaining cohorts;
107
+ - expected `/health/operate` version, unauthenticated MCP response, site marker, and content repository identity for each posture.
108
+
109
+ Run `bunx brains-ops reconcile-all <canonical-review-copy> --dry-run` against the isolated review copy. The command blocks external content-repository access, leaves the review copy untouched, lists both passes' changed files, and must report second-pass zero drift. Reconciliation owns generated per-user config; only `render` owns the observational `views/users.md` projection.
110
+
111
+ During the approved window:
112
+
113
+ 1. Freeze unrelated merges and releases. Wait for active Build, Reconcile, and Deploy runs to finish, then disable all three pilot workflows with `gh workflow disable build.yml`, `gh workflow disable reconcile.yml`, and `gh workflow disable deploy.yml`.
114
+ 2. Publish and verify the reviewed unified runtime and matching ops artifacts. Do not update pilot desired state until the exact versions are installable and the expected images can be built.
115
+ 3. Apply the reviewed canonical pilot revision while automation remains disabled. Confirm repository names, server/domain identity, content repositories, secret selectors, image names, and tag identity against the review diff.
116
+ 4. Enable only Build, run it for the canonical desired-state revision, and record every resulting image digest. Stop if the observed digest set differs from the cutover record.
117
+ 5. Enable only Reconcile, run it once, and review its generated per-user config commit. It must not rewrite `views/users.md`, and no generated file may combine canonical config with a retired image version.
118
+ 6. Enable Deploy and deploy one handle at a time in the approved order. After each deploy, run `bunx brains-ops verify-user . <handle>`, render observed fleet status, and complete the manual identity, content-sync, and app-managed site checks.
119
+ 7. Run Reconcile a second time. Require no reconciler-owned generated diff and no deploy work before re-enabling normal automation and lifting the merge/release freeze. Observed status rendering remains separate from this convergence gate.
120
+
121
+ If any gate fails, disable all three workflows again. Restore the prior pilot desired-state and dependency revision, reconcile with the prior ops version, and redeploy the prior image tag/digest as one rollback pair. Verify the prior `/health/operate` version and identity/content/site checks before re-enabling automation. Never restore only config or only an image.
122
+
123
+ ## Directory-sync stress gate
124
+
125
+ Use the manual `Directory Sync Stress` workflow only against a disposable smoke user. It refuses a target unless the handle, domain, and content repository all identify smoke, the confirmation input exactly matches `stress:<handle>`, and the user desired state declares the hermetic posture below. Reconcile and deploy this posture before running the workload:
126
+
127
+ ```yaml
128
+ embeddingEnabled: false
129
+ topicExtractionEnabled: false
130
+ skillDerivationEnabled: false
131
+ swotDerivationEnabled: false
132
+ ```
133
+
134
+ Before authorizing a workload, dispatch the workflow once with `verify_only: true`. That mode loads the same Bitwarden/Varlock content credential, clones the smoke content repository, and runs `git push --dry-run` against a temporary stress ref. It creates no ref, performs no content write, does not contact the deployed runtime, and skips cleanup because no probes were created.
135
+
136
+ Profiles are deterministic and reversible:
137
+
138
+ - `regression`: 20 probes;
139
+ - `load`: ramps to 350 probes, updates all, renames 100, updates again, then deletes all;
140
+ - `stress`: ramps to 700 probes and renames 200 before cleanup.
141
+
142
+ The workflow loads operator credentials through Bitwarden/Varlock, but it is separate from Deploy and cannot deploy an image. It creates a rollback branch before the first content write, gates on health timeouts, watchdog restarts, and external AI usage during the monitored workload window, preserves warmup and cleanup samples as evidence, uploads JSON/Markdown/runtime artifacts, and runs an independent idempotent cleanup job with `if: always()`. Once cleanup confirms that no probes remain, it also prunes retained `ops/directory-sync-stress-backup-*` branches; if probes remain, the branches stay available for recovery.
143
+
144
+ Treat any gated health failure, restart, OOM, residual probe, or entity-baseline drift as a failed gate. Do not restart the target during measurement. Recovery is a separate operator action after evidence collection.
145
+
146
+ ## Health watchdog smoke gate
147
+
148
+ Use the manual `Health Watchdog Smoke` workflow only after the `smoke` fleet user has completed a normal Deploy run. The workflow shares the `deploy-<handle>` concurrency group, resolves the server from pilot desired state through Hetzner, loads the existing deploy SSH key through Bitwarden/Varlock, and refuses targets whose handle and domain do not identify smoke. Confirm with `watchdog-smoke:<handle>`.
149
+
150
+ The smoke does not deploy or replace the installed systemd units. It verifies that the installed watchdog exactly matches the packaged canonical payload, then exercises that installed script. It also verifies the deployed rover image label, exact-label selector, active timer, and `/health/live` Docker healthcheck. While holding the timer's global lock, it creates temporary containers for eligible unhealthy, unrelated service-labelled unhealthy, false-labelled unhealthy, and eligible healthy cases. It then requires exactly three eligible restarts followed by budget suppression, diagnostics and state only for the eligible fixture, and no restart of the deployed rover or ineligible fixtures.
151
+
152
+ The primary job uploads remote incident, state, and command evidence. An independent `if: always()` cleanup job removes deterministic fixture names and remote temporary files using the same workflow run ID. Treat missing evidence, unexpected eligibility, deployed-container movement, cleanup failure, or any non-smoke target rejection as a failed gate.
153
+
154
+ ## Stale deploy lock recovery
155
+
156
+ Kamal intentionally leaves its remote deploy lock in place when a deployment is cancelled or interrupted. Confirm that no deployment for the user is still active before releasing the lock, then use the deploy workflow's explicit recovery input:
157
+
158
+ ```sh
159
+ gh workflow run Deploy --ref main \
160
+ -f handle=<handle> \
161
+ -f release_stale_lock=true
162
+ ```
163
+
164
+ Recovery is opt-in and scoped to one handle. Normal push, reconcile, and manual deploy runs never remove a lock automatically.
165
+
36
166
  ## Bootstrap flow
37
167
 
38
168
  For this fleet, operator-local secret material remains the source of truth during onboarding and rotation. The repo stores encrypted per-user secrets, not raw values.
@@ -51,53 +181,251 @@ The shared cert bootstrap writes local cert artifacts under `.brains-ops/certs/s
51
181
 
52
182
  Preview hosts use the shape `<handle>-preview.rizom.ai`, so one wildcard origin cert for `*.rizom.ai` covers both the primary and preview hosts for every pilot user.
53
183
 
184
+ ## Pilot user offboarding
185
+
186
+ Use the manual **Offboard** workflow for explicit pilot retirement. It is the only
187
+ supported path that removes both checked-in desired state and provider resources.
188
+ Normal reconcile/deploy changes never imply destruction.
189
+
190
+ 1. Enter a comma-separated handle list. The command sorts and deduplicates it.
191
+ 2. Run with `apply: false` and review the exact servers, DNS records, and content
192
+ repositories in the plan.
193
+ 3. Enter the canonical confirmation printed by the dry run, for example
194
+ `sunset:alice,bob`.
195
+ 4. Rerun with `apply: true`.
196
+ 5. Verify the workflow's bot commit removes user YAML, encrypted secrets, generated
197
+ user directories, cohort membership, and rows from `views/users.md`.
198
+
199
+ The apply path archives each private content repository, deletes the user's main and
200
+ preview DNS records, destroys the dedicated Hetzner server, and removes desired state.
201
+ Custom-domain users also lose their managed `www` record. The operation is idempotent,
202
+ but it deliberately creates **no runtime backup**; obtain separate owner approval and
203
+ backup state before apply when retention is required.
204
+
205
+ The equivalent operator-local commands are:
206
+
207
+ ```sh
208
+ bunx brains-ops user:offboard . alice bob
209
+ bunx brains-ops user:offboard . alice bob \
210
+ --apply --confirm sunset:alice,bob
211
+ ```
212
+
213
+ The deploy handle resolver ignores deleted users, so the generated-file deletions in
214
+ the offboarding commit cannot route those handles back through onboarding/deploy.
215
+
54
216
  ## Upgrading operator behavior
55
217
 
56
- When `@rizom/ops` changes the scaffolded deploy contract:
218
+ The pilot repository pins `@rizom/ops` in `package.json`. The scheduled and manually dispatched Upgrade workflow owns routine upgrades to that pin. It refreshes the scaffold on a branch and opens a reviewable PR; it does not change runtime desired state or authorize a deployment.
219
+
220
+ Because scaffold refreshes can update `.github/workflows/*`, the workflow must not push with its Actions `GITHUB_TOKEN`. Configure a dedicated GitHub App:
57
221
 
58
- 1. bump `@rizom/ops` in `package.json`
59
- 2. rerun the relevant scaffold/reconcile flow
60
- 3. review the resulting changes to `.env.schema`, `deploy/scripts/`, and workflows in git
61
- 4. commit the updated deploy artifacts together
222
+ 1. Install it only on this pilot repository.
223
+ 2. Grant repository permissions `Contents: Read and write`, `Pull requests: Read and write`, and `Workflows: Read and write`; grant nothing else.
224
+ 3. Store its App ID as the repository Actions variable `OPS_UPGRADE_APP_ID`.
225
+ 4. Store its private key as the repository Actions secret `OPS_UPGRADE_APP_PRIVATE_KEY`.
62
226
 
63
- ## Rover-core verification notes
227
+ The workflow explicitly requests only those three permissions. Checkout persists no credential. After the freshly published `@rizom/ops` finishes, the workflow checks whether it produced a change; only then does it mint a short-lived, repository-scoped App token for the push-and-open-PR steps. A credential that can rewrite `.github/workflows/` therefore does not exist while upgraded package code runs.
64
228
 
65
- Rover core is MCP-only. Do not expect the bare domain to serve a website.
229
+ If `OPS_UPGRADE_APP_ID` is unset, or token creation, branch push, or PR creation fails, the run stops with a non-zero status. Repair the CI credential path; never fall back to an operator's personal SSH key or token.
66
230
 
67
- Use these checks after deploy:
231
+ Adopting this credential flow in an existing pilot repository requires one explicitly reviewed bootstrap PR because the old Upgrade workflow cannot update itself. After that merge, routine upgrades run entirely in CI:
68
232
 
69
- - `https://<handle>.rizom.ai/health` should return `200`
70
- - unauthenticated `POST https://<handle>.rizom.ai/mcp` should return `401 Unauthorized: Bearer token required`
71
- - a bare `GET /` may also return `401`; that is expected for rover core and does not indicate a bad deploy
233
+ 1. dispatch Upgrade with an exact version, or let its schedule select `latest`;
234
+ 2. review the generated package, lockfile, deploy-script, and workflow diff;
235
+ 3. merge the upgrade PR only after its checks pass;
236
+ 4. change runtime desired state separately through the approved canary or fleet rollout flow.
72
237
 
73
- ## Discord bot token checklist
238
+ ## Canonical verification notes
239
+
240
+ Use the verification script after deploy:
241
+
242
+ ```sh
243
+ bunx brains-ops verify-user . <handle>
244
+ ```
245
+
246
+ For every bundle posture it checks:
247
+
248
+ - `https://<handle>.rizom.ai/health/operate` returns `200`;
249
+ - unauthenticated `POST https://<handle>.rizom.ai/mcp` returns the expected auth failure;
250
+ - background jobs are not repeatedly failing, except for missing optional integrations.
251
+
252
+ A `core`-only instance is MCP-only; a bare `GET /` may return `401` without indicating a bad deploy. When `site` is selected, verification also checks the browser and Studio/login surfaces.
253
+
254
+ Manual checks that remain:
255
+
256
+ - initial app-managed site output is correct for the expected content/theme;
257
+ - content repository identity and runtime sync are healthy;
258
+ - passkey setup/handoff is completed from the setup email.
259
+
260
+ ## One-user canonical site canary
261
+
262
+ Run this before adding custom site/theme packages or rolling a larger browser/Studio-first cohort.
263
+
264
+ 1. Create or choose a canary cohort with explicit bundles:
265
+
266
+ ```yaml
267
+ bundles:
268
+ - core
269
+ - site
270
+ - publishing
271
+ ```
272
+
273
+ 2. Add exactly one canary user to that cohort.
274
+ 3. For browser/Studio-first onboarding, configure setup email in `users/<handle>.yaml`:
275
+
276
+ ```yaml
277
+ setup:
278
+ delivery: email
279
+ email: user@example.com
280
+ ```
281
+
282
+ 4. Encrypt the user's secrets and commit only the `.age` file.
283
+ 5. Run `bunx brains-ops onboard . <handle>`.
284
+ 6. Run `bunx brains-ops verify-user . <handle>` with no custom site/theme overrides.
285
+ 7. Ask the user to complete passkey setup from the setup email.
286
+ 8. Continue to visual customization only after the canary is healthy.
287
+
288
+ Rollback must restore the prior desired-state revision and prior runtime image together. Never pair canonical config with the retired image, or retired config with the canonical image.
289
+
290
+ ## Hosted site and theme package contract
291
+
292
+ Start with the public [site mockup migration guide](https://github.com/rizom-ai/brains/blob/main/docs/site-mockup-migration.md), then apply these hosted-fleet requirements:
293
+
294
+ - A site package must default-export `defineSite(...)` and import its authoring API only from `@rizom/site`.
295
+ - A theme package must default-export its CSS as a string. Hosted custom themes currently use the `@rizom/*` scope so the fleet image installs them with the site package; `@brains/*` themes are bundled with `@rizom/brain`.
296
+ - Site and custom theme packages must be public npm packages that install without registry credentials.
297
+ - Site, theme, and brain packages publish independently. Hosted configuration requires exact site and external-theme version pins and never derives one package version from another.
298
+ - Keep site structure and theme CSS in separate packages. Do not put private content or secrets in either package.
299
+
300
+ Configure a user in `users/<handle>.yaml`:
301
+
302
+ ```yaml
303
+ siteOverride:
304
+ package: "@rizom/site-example"
305
+ version: <exact-site-version>
306
+ theme: "@rizom/theme-example"
307
+ themeVersion: <exact-theme-version>
308
+ ```
309
+
310
+ Missing external package versions fail desired-state validation. Every instance on one
311
+ Brain version uses the same image and exact package union. Change a site/theme package
312
+ pin only together with a fresh Brain version; published `brain-${brainVersion}` tags
313
+ remain immutable. Conflicting pins for one package on the same Brain version fail before
314
+ build. Bundled `@brains/*` themes omit `themeVersion` because they are not installed as
315
+ separate packages.
316
+
317
+ ### Custom-package canary and rollback
318
+
319
+ 1. Confirm the exact site/theme versions are public-installable without npm credentials.
320
+ 2. Apply the exact package names and versions to one healthy canonical site canary and select a fresh Brain version for that package set.
321
+ 3. Push desired state and let the normal Build → Reconcile → Deploy chain create the shared version image and update the canary.
322
+ 4. Run `bunx brains-ops verify-user . <handle>`.
323
+ 5. Manually verify the site, theme, Studio, content sync, and passkey sign-in before adding more users.
324
+
325
+ To roll back, remove or change `siteOverride` while selecting a fresh Brain version,
326
+ then reconcile and redeploy that user. Never mutate an existing image tag; coordinate any
327
+ other users intentionally sharing the selected version.
328
+
329
+ ## Setup email checklist
330
+
331
+ Use this for browser/Studio-first users who should receive their own first-passkey setup link by email.
332
+
333
+ 1. Add setup delivery to the user file:
334
+
335
+ ```yaml
336
+ setup:
337
+ delivery: email
338
+ email: user@example.com
339
+ ```
340
+
341
+ 2. Configure these GitHub Secrets before deploy:
342
+ - `SETUP_EMAIL_API_KEY`
343
+ - `SETUP_EMAIL_FROM`
344
+
345
+ 3. Reconcile/deploy the user or cohort:
346
+ - `bunx brains-ops onboard . <handle>`
347
+ - or `bunx brains-ops reconcile-cohort . <cohort>`
348
+
349
+ 4. Verify the generated `users/<handle>/brain.yaml` contains `auth-service.setupEmail` and `email` interface config.
350
+ 5. Ask the user to complete passkey setup from the email link, then use:
351
+ - Dashboard: `https://<handle>.rizom.ai/`
352
+ - Studio: `https://<handle>.rizom.ai/studio`
353
+
354
+ Notes:
355
+
356
+ - The setup URL is generated and sent by the running brain; operators should not scrape logs or SSH into the instance to retrieve it.
357
+ - The auth service owns setup email dedupe. It should not resend for the same persisted setup token after restart, but should retry failed delivery and resend after token rotation.
358
+ - `SETUP_EMAIL_FROM` is not marked required because fleets without email setup can omit it, but it is required for users with `setup.delivery: email`.
359
+
360
+ ## AT Protocol smoke/config checklist
361
+
362
+ Use this when enabling AT Protocol publishing for a single pilot user.
363
+
364
+ 1. Add the public PDS identifier to the user file. Prefer the account DID as
365
+ the identifier — it survives handle changes. Add `accountDid` too when the
366
+ member wants their handle verified against their subdomain
367
+ (`@<handle>.<domainSuffix>`): the brain then serves it at
368
+ `/.well-known/atproto-did` and Bluesky's "I have my own domain" HTTP
369
+ verification passes with no DNS records.
370
+
371
+ ```yaml
372
+ atproto:
373
+ identifier: did:plc:example123
374
+ accountDid: did:plc:example123
375
+ ```
376
+
377
+ Only for the PDS account designated by the protocol authority's `_lexicon`
378
+ DNS TXT record, also set `lexiconAuthority: true`. Every other fleet user
379
+ must omit it.
380
+
381
+ 2. Put the app password in `users/<handle>.secrets.yaml`:
382
+
383
+ ```yaml
384
+ atprotoAppPassword: <app-password>
385
+ ```
386
+
387
+ 3. Encrypt the per-user secret payload:
388
+ - `bunx brains-ops secrets:encrypt . <handle>`
389
+ 4. Reconcile/deploy the user or cohort:
390
+ - `bunx brains-ops onboard . <handle>`
391
+ - or `bunx brains-ops reconcile-cohort . <cohort>`
392
+ 5. Verify the generated `users/<handle>/brain.yaml` contains `plugins.atproto.identifier` (plus `accountDid` and `lexiconAuthority` when configured) and `appPassword: ${ATPROTO_APP_PASSWORD}`.
393
+
394
+ Notes:
395
+
396
+ - The ATProto identifier and authority flag are public instance config and belong in `users/<handle>.yaml`. Only the DNS-designated authority account may set `lexiconAuthority: true`.
397
+ - The ATProto app password is secret and belongs only in the encrypted per-user secret payload.
398
+ - For smoke deployments, pin only the smoke cohort/user to the released brain version that contains ATProto support.
399
+
400
+ ## Discord application credential checklist
74
401
 
75
402
  Use this when enabling Discord for a pilot user.
76
403
 
77
404
  1. Pick the user handle (for example `smoke`).
78
405
  2. Open the Discord Developer Portal.
79
- 3. Create a **new application** for that user's rover.
406
+ 3. Create a **new application** for that user's brain.
80
407
  4. Add a **Bot** to the application.
81
- 5. Copy the bot token.
82
- 6. Put that value in `.env` or `.env.local` in this repo as `DISCORD_BOT_TOKEN=...` while onboarding that user.
408
+ 5. Copy the bot token, application public key, and application ID.
409
+ 6. Put those values in `.env` or `.env.local` while onboarding that user:
410
+ - `DISCORD_BOT_TOKEN=...`
411
+ - `DISCORD_PUBLIC_KEY=...`
412
+ - `DISCORD_APPLICATION_ID=...`
83
413
  7. Keep `discord.enabled: true` in `users/<handle>.yaml` unless you explicitly want to disable the primary pilot interface.
84
- 8. Encrypt the current per-user secret payload:
414
+ 8. Encrypt the current per-user credential payload:
85
415
  - `bunx brains-ops secrets:encrypt . <handle>`
86
416
  9. Reconcile/deploy the user or cohort:
87
-
88
- - `bunx brains-ops onboard . <handle>`
89
- - or `bunx brains-ops reconcile-cohort . <cohort>`
90
-
91
- 11. In the Discord Developer Portal, generate an install URL and invite the bot to the right server.
92
- 12. Send a test message in Discord and confirm the rover responds.
417
+ - `bunx brains-ops onboard . <handle>`
418
+ - or `bunx brains-ops reconcile-cohort . <cohort>`
419
+ 10. In the Discord Developer Portal, generate an install URL and invite the bot to the right server.
420
+ 11. Send a test message in Discord and confirm the brain responds.
93
421
 
94
422
  Notes:
95
423
 
96
- - Use **one bot token per user/rover**.
97
- - Do not reuse the same Discord bot token across multiple pilot users.
424
+ - Use **one Discord application credential set per user/brain**.
425
+ - Do not reuse the same Discord application across multiple pilot users.
98
426
  - Discord is the default pilot interface moving forward.
99
427
  - The encrypted `users/<handle>.secrets.yaml.age` file is the durable checked-in deploy input; your local env is only the operator staging source.
100
- - MCP is optional and mainly for direct client access or specific testing workflows.
428
+ - Direct MCP client access should use OAuth/passkey-capable clients where possible.
101
429
  - When explaining the content workflow, describe it first as a normal **git repo** of **markdown/text files**.
102
430
  - Position **Obsidian** as optional: it is just one possible editor for those same files, not the default requirement.
103
431