@rizom/ops 0.2.0-alpha.35 → 0.2.0-alpha.350

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (63) hide show
  1. package/README.md +62 -2
  2. package/dist/brains-ops.js +785 -304
  3. package/dist/capability-bundle-migration.d.ts +22 -0
  4. package/dist/cert-bootstrap.d.ts +3 -1
  5. package/dist/content-repo-ref.d.ts +10 -0
  6. package/dist/content-repo.d.ts +1 -0
  7. package/dist/deploy.js +99 -166
  8. package/dist/directory-sync-stress-system.d.ts +75 -0
  9. package/dist/directory-sync-stress.d.ts +107 -0
  10. package/dist/entries/deploy.d.ts +3 -2
  11. package/dist/health-watchdog-smoke.d.ts +54 -0
  12. package/dist/images.d.ts +79 -0
  13. package/dist/index.d.ts +8 -0
  14. package/dist/index.js +784 -302
  15. package/dist/legacy-pilot-migration.d.ts +11 -0
  16. package/dist/legacy-projection-job-recovery.d.ts +29 -0
  17. package/dist/load-registry.d.ts +60 -6
  18. package/dist/observed-status.d.ts +1 -1
  19. package/dist/origin-ca.d.ts +1 -1
  20. package/dist/parse-args.d.ts +2 -10
  21. package/dist/preview-domain.d.ts +9 -0
  22. package/dist/push-secrets.d.ts +2 -9
  23. package/dist/push-target.d.ts +1 -2
  24. package/dist/reconcile-dry-run.d.ts +9 -0
  25. package/dist/run-command.d.ts +17 -3
  26. package/dist/run-subprocess.d.ts +1 -6
  27. package/dist/schema.d.ts +123 -162
  28. package/dist/secrets-encrypt.d.ts +7 -13
  29. package/dist/ssh-key-bootstrap.d.ts +1 -26
  30. package/dist/stage-legacy-crossover.d.ts +23 -0
  31. package/dist/stress-command.d.ts +14 -0
  32. package/dist/stress-git-checkout.d.ts +21 -0
  33. package/dist/stress-health-monitor.d.ts +46 -0
  34. package/dist/upgrade.d.ts +10 -0
  35. package/dist/user-add.d.ts +15 -0
  36. package/dist/verify-user.d.ts +22 -0
  37. package/package.json +50 -42
  38. package/templates/rover-pilot/.env.schema +24 -3
  39. package/templates/rover-pilot/.github/actions/varlock-env/action.yml +47 -0
  40. package/templates/rover-pilot/.github/workflows/build.yml +70 -17
  41. package/templates/rover-pilot/.github/workflows/deploy.yml +71 -57
  42. package/templates/rover-pilot/.github/workflows/directory-sync-stress.yml +119 -0
  43. package/templates/rover-pilot/.github/workflows/health-watchdog-smoke.yml +95 -0
  44. package/templates/rover-pilot/.github/workflows/reconcile.yml +10 -4
  45. package/templates/rover-pilot/.github/workflows/upgrade.yml +104 -0
  46. package/templates/rover-pilot/README.md +16 -6
  47. package/templates/rover-pilot/deploy/scripts/create-predeploy-backup.ts +873 -0
  48. package/templates/rover-pilot/deploy/scripts/decrypt-user-secrets.ts +80 -24
  49. package/templates/rover-pilot/deploy/scripts/helpers.ts +3 -0
  50. package/templates/rover-pilot/deploy/scripts/install-health-watchdog.ts +144 -0
  51. package/templates/rover-pilot/deploy/scripts/resolve-missing-images.ts +13 -0
  52. package/templates/rover-pilot/deploy/scripts/resolve-user-config.ts +44 -9
  53. package/templates/rover-pilot/deploy/scripts/sync-content-repo.ts +51 -47
  54. package/templates/rover-pilot/deploy/scripts/update-dns.ts +14 -4
  55. package/templates/rover-pilot/deploy/scripts/validate-secrets.ts +12 -1
  56. package/templates/rover-pilot/docs/canonical-crossover-record.md +107 -0
  57. package/templates/rover-pilot/docs/onboarding-checklist.md +28 -17
  58. package/templates/rover-pilot/docs/operator-playbook.md +307 -29
  59. package/templates/rover-pilot/docs/user-onboarding.md +48 -463
  60. package/templates/rover-pilot/pilot.yaml +7 -4
  61. package/templates/rover-pilot/.kamal/hooks/pre-deploy +0 -9
  62. package/templates/rover-pilot/deploy/Dockerfile +0 -30
  63. package/templates/rover-pilot/deploy/kamal/deploy.yml +0 -40
@@ -4,28 +4,39 @@
4
4
  2. Run `bunx brains-ops age-key:bootstrap <repo> --push-to gh`.
5
5
  3. Fill in `pilot.yaml`.
6
6
  - keep your pinned `brainVersion`
7
- - confirm shared selectors for `aiApiKey`, `gitSyncToken`, and `mcpAuthToken`
7
+ - confirm shared selectors for `aiApiKey`, `gitSyncToken`, and `contentRepoAdminToken`
8
+ - use different tokens for `contentRepoAdminToken` and `gitSyncToken`: admin creates/checks content repos; sync is used by runtime directory-sync
8
9
  - confirm `agePublicKey`
9
- 4. Add or edit `users/<handle>.yaml`.
10
- - Discord is enabled by default for pilot users
11
- - if the user should be an anchor there, set `discord.anchorUserId` to their Discord user ID
12
- 5. Add the user to a cohort in `cohorts/*.yaml`.
10
+ 4. Run `bunx brains-ops user:add <repo> <handle> --cohort <cohort>`.
11
+ - Web chat is the primary interface; it needs no per-user setup beyond the passkey.
12
+ - `user:add` currently writes `discord: enabled: true`; set it to `false` unless the user's cohort actually uses Discord.
13
+ - if the user should be an anchor on Discord, add `--anchor-id <discord-user-id>`.
14
+ - the command creates `users/<handle>.yaml`, `users/<handle>.secrets.yaml`, and the cohort membership without duplicating existing entries.
15
+ 5. Edit the generated user file if the anchor profile needs richer metadata.
16
+ - Set `setup.delivery: email` and `setup.email` so the user gets the passkey setup email — this is the default onboarding path.
17
+ - For ATProto publishing, add `atproto.identifier` to the user file; put only `atprotoAppPassword` in the per-user secrets file.
18
+ - Ensure `SETUP_EMAIL_API_KEY` and `SETUP_EMAIL_FROM` exist as GitHub Secrets before deploying any email-setup user.
13
19
  6. Run `bunx brains-ops render <repo>`.
14
20
  7. Run `bunx brains-ops ssh-key:bootstrap <repo> --push-to gh`.
15
21
  8. Run `bunx brains-ops cert:bootstrap <repo> --push-to gh`.
16
- 9. Keep raw user secret material locally for now (`.env.local`, file-backed env vars, or equivalent local inputs).
22
+ 9. Keep raw user secret material locally for now (`.env.local`, file-backed env vars, or equivalent local inputs), including `CONTENT_REPO_ADMIN_TOKEN` for operator onboarding.
17
23
  10. Run `bunx brains-ops secrets:encrypt <repo> <handle>`.
18
24
  11. Commit and push `users/<handle>.secrets.yaml.age`.
19
25
  12. Run `bunx brains-ops onboard <repo> <handle>`.
20
- 13. Verify the deployed rover core contract:
21
- - `https://<handle>.rizom.ai/health` returns `200`
22
- - unauthenticated `POST https://<handle>.rizom.ai/mcp` returns `401`
23
- 14. For fleet upgrades, edit `pilot.yaml.brainVersion` and push once; CI rebuilds the shared image tag, refreshes generated user env files, and redeploys affected users.
24
- 15. Hand the Discord setup details to the user.
25
- 16. Hand over the browser defaults:
26
+ 13. Verify the deployed canonical contract:
27
+ - `https://<handle>.rizom.ai/health/operate` returns `200`
28
+ - `https://<handle>.rizom.ai/chat` loads the web chat and accepts passkey sign-in
29
+ - unauthenticated `POST https://<handle>.rizom.ai/mcp` returns the expected auth failure
30
+ - content repo exists and runtime sync is healthy
31
+ - background jobs are not repeatedly failing, except for expected missing optional integrations
32
+ - when `site` is selected, the browser/Studio surfaces load and the initial app-managed site build completes
33
+ 14. For fleet upgrades, edit `pilot.yaml.brainVersion` and push once; CI rebuilds the required default/site image tags, refreshes generated user env files, and redeploys affected users. Every external site and theme package keeps its own required exact version pin and never follows the brain version implicitly.
34
+ 15. Confirm the user received the setup email, registered their passkey, and can sign in to web chat at `https://<handle>.rizom.ai/chat`. That completes the default onboarding; everything below is per-cohort extras.
35
+ 16. Hand over the browser surfaces:
36
+ - Chat (primary): `https://<handle>.rizom.ai/chat`
26
37
  - Dashboard: `https://<handle>.rizom.ai/`
27
- - CMS: `https://<handle>.rizom.ai/cms`
28
- - GitHub token guidance for CMS access to the user's private content repo
29
- 17. If they need direct client access, also hand over the MCP connection details.
30
- 18. If you are also giving them a content repo workflow, describe it as optional and frame git/Obsidian as an advanced file-based path, not the default.
31
- 19. Send `docs/user-onboarding.md` to the user as the pilot handoff guide.
38
+ - Studio: `https://<handle>.rizom.ai/studio`, plus GitHub token guidance if Studio editing is part of their cohort
39
+ 17. For Discord-enabled cohorts, hand the Discord setup details to the user as a secondary chat surface.
40
+ 18. If they need direct client access (MCP), use OAuth/passkey-capable clients where possible.
41
+ 19. If you are also giving them a content repo workflow, describe it as optional and frame git/Obsidian as an advanced file-based path, not the default.
42
+ 20. Send `docs/user-onboarding.md` to the user as the pilot handoff guide.
@@ -9,30 +9,145 @@ Treat these as checked-in deploy artifacts in the pilot repo:
9
9
  - `deploy/scripts/`
10
10
  - `.github/workflows/build.yml`
11
11
  - `.github/workflows/deploy.yml`
12
+ - `.github/workflows/directory-sync-stress.yml`
13
+ - `.github/workflows/health-watchdog-smoke.yml`
12
14
  - `.github/workflows/reconcile.yml`
13
15
 
14
16
  `.env.schema` is the single source of truth for required and sensitive deploy vars.
15
17
  The deploy scripts and workflows should read from that contract instead of inventing a second list.
16
18
 
17
- The shared pilot image tag is `brain-${brainVersion}`:
19
+ The default pilot image tag is `brain-${brainVersion}`:
18
20
 
19
- - build publishes `brain-${brainVersion}`
21
+ - build publishes `brain-${brainVersion}` for users without a site override
22
+ - a site override gets an isolated `brain-${brainVersion}-sites-${packageHash}` image
20
23
  - generated `users/<handle>/.env` carries `BRAIN_VERSION=<brainVersion>`
21
- - deploy sets `VERSION=brain-${brainVersion}`
24
+ - build and deploy derive the same effective image tag from the resolved registry
22
25
 
23
26
  ## Version bump flow
24
27
 
25
28
  When `pilot.yaml.brainVersion` changes and you push:
26
29
 
27
- 1. build publishes the new shared image tag
30
+ 1. build publishes the new default image and any required site images
28
31
  2. reconcile refreshes generated `users/<handle>/.env`
29
32
  3. deploy runs for handles whose generated config changed
30
33
  4. generated file commits happen once in a final aggregation step after the deploy matrix finishes
31
34
 
35
+ Every external site and theme package has its own exact version pin. A cohort or
36
+ pilot brain-version bump never changes those package versions implicitly; update each
37
+ pin deliberately from reviewed package and image evidence.
38
+
32
39
  When a push changes only deploy contract files and no generated `users/<handle>/.env` or `users/<handle>/brain.yaml` files, the deploy workflow exits through its explicit no-op path and prints `No affected user configs; skipping deploy.`
33
40
 
34
41
  They are scaffolded from `@rizom/ops`, then versioned in this repo like any other deploy contract.
35
42
 
43
+ ## Verified pre-deploy rollback snapshots
44
+
45
+ Deploy creates and verifies one target-scoped rollback snapshot after SSH and image readiness, immediately before optional stale-lock release and `kamal setup`. A snapshot failure stops the workflow before container replacement. A genuinely new server with no runtime or persistent state reports `not applicable`; persistent state without an identifiable runtime fails closed.
46
+
47
+ Each snapshot contains transaction-consistent online captures of the six canonical SQLite databases, database hashes and `quick_check` results, the deployed `brain.yaml`, sanitized container and mount metadata, and the Directory Sync checkout. Git capture includes all refs, an observed remote head, staged and unstaged binary patches, and mode-preserving archives of untracked and ignored files. Capture performs no checkout mutation or remote write.
48
+
49
+ Verified snapshots are finalized atomically under:
50
+
51
+ ```text
52
+ /opt/brain-state/backups/predeploy-<handle>-<target>-<UTC timestamp>/
53
+ ```
54
+
55
+ Files use mode `0600` and the directory uses `0700`. A `.incomplete` directory is diagnostic evidence, never a rollback point. Checksums and the versioned manifest must verify before finalization. The default retention is five verified snapshots per target; incomplete or unverified directories do not count, and the snapshot created by the current deployment is never pruned.
56
+
57
+ These are same-server rollback snapshots, not off-host disaster recovery. There is no normal skip and no automatic restore. Restoration requires separate approval: stop replacement/application processes, recheck the selected snapshot's checksums, preserve the current state, restore databases and exact Git state, then validate health, queues, Git checkpoints, durable export intents, preview, and production output before reopening the target.
58
+
59
+ ## Retired projection job recovery
60
+
61
+ A pre-scheduler projection job can remain active after its handler type is retired. Current runtime health reports active job types missing from the finalized execution inventory as degraded. Do not restart repeatedly or relax the snapshot idle gate.
62
+
63
+ Use the recovery command only after read-only evidence proves all of the following for one exact row:
64
+
65
+ - its type is one of the command's fixed legacy projection types;
66
+ - the installed runtime no longer contains that handler;
67
+ - attempt, worker-session, lease, and heartbeat ownership are all absent;
68
+ - no durable progress snapshot or result exists.
69
+
70
+ Preview the exact row first:
71
+
72
+ ```sh
73
+ bunx --package @rizom/ops@<exact-version> brains-ops \
74
+ recover:retire-legacy-projection-job \
75
+ /data/brain-jobs.db <job-id> --type <legacy-type> --dry-run
76
+ ```
77
+
78
+ After separate operator review, replace `--dry-run` with the exact confirmation `--confirm retire:<job-id>`. The command atomically fences against ownership or progress appearing between inspection and retirement, and fails closed if the row changed. It is not a general job cancellation API and cannot retire arbitrary job types. Require operational health and a fully idle queue afterward, then run the unchanged canonical predeploy snapshot and Deploy workflow.
79
+
80
+ ## Canonical contract crossover maintenance window
81
+
82
+ Do not run this procedure without explicit operator approval. The canonical desired state, canonical `@rizom/ops`, and unified runtime image form one contract and must move or roll back together. Complete `docs/canonical-crossover-record.md` as the approval evidence without adding secret values.
83
+
84
+ Before the window, record and review:
85
+
86
+ - the prior pilot commit and exact `@rizom/ops` version;
87
+ - every prior runtime image tag and immutable digest;
88
+ - the reviewed canonical pilot commit;
89
+ - the exact unified `@rizom/brain` and `@rizom/ops` versions;
90
+ - every unified image tag and immutable digest;
91
+ - canary-first rollout order, followed by the remaining cohorts;
92
+ - expected `/health/operate` version, unauthenticated MCP response, site marker, and content repository identity for each posture.
93
+
94
+ Run `bunx brains-ops reconcile-all <canonical-review-copy> --dry-run` against the isolated review copy. The command blocks external content-repository access, leaves the review copy untouched, lists both passes' changed files, and must report second-pass zero drift. Reconciliation owns generated per-user config; only `render` owns the observational `views/users.md` projection.
95
+
96
+ During the approved window:
97
+
98
+ 1. Freeze unrelated merges and releases. Wait for active Build, Reconcile, and Deploy runs to finish, then disable all three pilot workflows with `gh workflow disable build.yml`, `gh workflow disable reconcile.yml`, and `gh workflow disable deploy.yml`.
99
+ 2. Publish and verify the reviewed unified runtime and matching ops artifacts. Do not update pilot desired state until the exact versions are installable and the expected images can be built.
100
+ 3. Apply the reviewed canonical pilot revision while automation remains disabled. Confirm repository names, server/domain identity, content repositories, secret selectors, image names, and tag identity against the review diff.
101
+ 4. Enable only Build, run it for the canonical desired-state revision, and record every resulting image digest. Stop if the observed digest set differs from the cutover record.
102
+ 5. Enable only Reconcile, run it once, and review its generated per-user config commit. It must not rewrite `views/users.md`, and no generated file may combine canonical config with a retired image version.
103
+ 6. Enable Deploy and deploy one handle at a time in the approved order. After each deploy, run `bunx brains-ops verify-user . <handle>`, render observed fleet status, and complete the manual identity, content-sync, and app-managed site checks.
104
+ 7. Run Reconcile a second time. Require no reconciler-owned generated diff and no deploy work before re-enabling normal automation and lifting the merge/release freeze. Observed status rendering remains separate from this convergence gate.
105
+
106
+ If any gate fails, disable all three workflows again. Restore the prior pilot desired-state and dependency revision, reconcile with the prior ops version, and redeploy the prior image tag/digest as one rollback pair. Verify the prior `/health/operate` version and identity/content/site checks before re-enabling automation. Never restore only config or only an image.
107
+
108
+ ## Directory-sync stress gate
109
+
110
+ Use the manual `Directory Sync Stress` workflow only against a disposable smoke user. It refuses a target unless the handle, domain, and content repository all identify smoke, the confirmation input exactly matches `stress:<handle>`, and the user desired state declares the hermetic posture below. Reconcile and deploy this posture before running the workload:
111
+
112
+ ```yaml
113
+ embeddingEnabled: false
114
+ topicExtractionEnabled: false
115
+ skillDerivationEnabled: false
116
+ swotDerivationEnabled: false
117
+ ```
118
+
119
+ Before authorizing a workload, dispatch the workflow once with `verify_only: true`. That mode loads the same Bitwarden/Varlock content credential, clones the smoke content repository, and runs `git push --dry-run` against a temporary stress ref. It creates no ref, performs no content write, does not contact the deployed runtime, and skips cleanup because no probes were created.
120
+
121
+ Profiles are deterministic and reversible:
122
+
123
+ - `regression`: 20 probes;
124
+ - `load`: ramps to 350 probes, updates all, renames 100, updates again, then deletes all;
125
+ - `stress`: ramps to 700 probes and renames 200 before cleanup.
126
+
127
+ The workflow loads operator credentials through Bitwarden/Varlock, but it is separate from Deploy and cannot deploy an image. It creates a rollback branch before the first content write, gates on health timeouts, watchdog restarts, and external AI usage during the monitored workload window, preserves warmup and cleanup samples as evidence, uploads JSON/Markdown/runtime artifacts, and runs an independent idempotent cleanup job with `if: always()`. Once cleanup confirms that no probes remain, it also prunes retained `ops/directory-sync-stress-backup-*` branches; if probes remain, the branches stay available for recovery.
128
+
129
+ Treat any gated health failure, restart, OOM, residual probe, or entity-baseline drift as a failed gate. Do not restart the target during measurement. Recovery is a separate operator action after evidence collection.
130
+
131
+ ## Health watchdog smoke gate
132
+
133
+ Use the manual `Health Watchdog Smoke` workflow only after the `smoke` fleet user has completed a normal Deploy run. The workflow shares the `deploy-<handle>` concurrency group, resolves the server from pilot desired state through Hetzner, loads the existing deploy SSH key through Bitwarden/Varlock, and refuses targets whose handle and domain do not identify smoke. Confirm with `watchdog-smoke:<handle>`.
134
+
135
+ The smoke does not deploy or replace the installed systemd units. It verifies that the installed watchdog exactly matches the packaged canonical payload, then exercises that installed script. It also verifies the deployed rover image label, exact-label selector, active timer, and `/health/live` Docker healthcheck. While holding the timer's global lock, it creates temporary containers for eligible unhealthy, unrelated service-labelled unhealthy, false-labelled unhealthy, and eligible healthy cases. It then requires exactly three eligible restarts followed by budget suppression, diagnostics and state only for the eligible fixture, and no restart of the deployed rover or ineligible fixtures.
136
+
137
+ The primary job uploads remote incident, state, and command evidence. An independent `if: always()` cleanup job removes deterministic fixture names and remote temporary files using the same workflow run ID. Treat missing evidence, unexpected eligibility, deployed-container movement, cleanup failure, or any non-smoke target rejection as a failed gate.
138
+
139
+ ## Stale deploy lock recovery
140
+
141
+ Kamal intentionally leaves its remote deploy lock in place when a deployment is cancelled or interrupted. Confirm that no deployment for the user is still active before releasing the lock, then use the deploy workflow's explicit recovery input:
142
+
143
+ ```sh
144
+ gh workflow run Deploy --ref main \
145
+ -f handle=<handle> \
146
+ -f release_stale_lock=true
147
+ ```
148
+
149
+ Recovery is opt-in and scoped to one handle. Normal push, reconcile, and manual deploy runs never remove a lock automatically.
150
+
36
151
  ## Bootstrap flow
37
152
 
38
153
  For this fleet, operator-local secret material remains the source of truth during onboarding and rotation. The repo stores encrypted per-user secrets, not raw values.
@@ -53,51 +168,214 @@ Preview hosts use the shape `<handle>-preview.rizom.ai`, so one wildcard origin
53
168
 
54
169
  ## Upgrading operator behavior
55
170
 
56
- When `@rizom/ops` changes the scaffolded deploy contract:
171
+ The pilot repository pins `@rizom/ops` in `package.json`. The scheduled and manually dispatched Upgrade workflow owns routine upgrades to that pin. It refreshes the scaffold on a branch and opens a reviewable PR; it does not change runtime desired state or authorize a deployment.
172
+
173
+ Because scaffold refreshes can update `.github/workflows/*`, the workflow must not push with its Actions `GITHUB_TOKEN`. Configure a dedicated GitHub App:
174
+
175
+ 1. Install it only on this pilot repository.
176
+ 2. Grant repository permissions `Contents: Read and write`, `Pull requests: Read and write`, and `Workflows: Read and write`; grant nothing else.
177
+ 3. Store its App ID as the repository Actions variable `OPS_UPGRADE_APP_ID`.
178
+ 4. Store its private key as the repository Actions secret `OPS_UPGRADE_APP_PRIVATE_KEY`.
179
+
180
+ The workflow explicitly requests only those three permissions. Checkout persists no credential. After the freshly published `@rizom/ops` finishes, the workflow checks whether it produced a change; only then does it mint a short-lived, repository-scoped App token for the push-and-open-PR steps. A credential that can rewrite `.github/workflows/` therefore does not exist while upgraded package code runs.
181
+
182
+ If `OPS_UPGRADE_APP_ID` is unset, or token creation, branch push, or PR creation fails, the run stops with a non-zero status. Repair the CI credential path; never fall back to an operator's personal SSH key or token.
183
+
184
+ Adopting this credential flow in an existing pilot repository requires one explicitly reviewed bootstrap PR because the old Upgrade workflow cannot update itself. After that merge, routine upgrades run entirely in CI:
185
+
186
+ 1. dispatch Upgrade with an exact version, or let its schedule select `latest`;
187
+ 2. review the generated package, lockfile, deploy-script, and workflow diff;
188
+ 3. merge the upgrade PR only after its checks pass;
189
+ 4. change runtime desired state separately through the approved canary or fleet rollout flow.
57
190
 
58
- 1. bump `@rizom/ops` in `package.json`
59
- 2. rerun the relevant scaffold/reconcile flow
60
- 3. review the resulting changes to `.env.schema`, `deploy/scripts/`, and workflows in git
61
- 4. commit the updated deploy artifacts together
191
+ ## Canonical verification notes
62
192
 
63
- ## Rover-core verification notes
193
+ Use the verification script after deploy:
64
194
 
65
- Rover core is MCP-only. Do not expect the bare domain to serve a website.
195
+ ```sh
196
+ bunx brains-ops verify-user . <handle>
197
+ ```
66
198
 
67
- Use these checks after deploy:
199
+ For every bundle posture it checks:
68
200
 
69
- - `https://<handle>.rizom.ai/health` should return `200`
70
- - unauthenticated `POST https://<handle>.rizom.ai/mcp` should return `401 Unauthorized: Bearer token required`
71
- - a bare `GET /` may also return `401`; that is expected for rover core and does not indicate a bad deploy
201
+ - `https://<handle>.rizom.ai/health/operate` returns `200`;
202
+ - unauthenticated `POST https://<handle>.rizom.ai/mcp` returns the expected auth failure;
203
+ - background jobs are not repeatedly failing, except for missing optional integrations.
72
204
 
73
- ## Discord bot token checklist
205
+ A `core`-only instance is MCP-only; a bare `GET /` may return `401` without indicating a bad deploy. When `site` is selected, verification also checks the browser and Studio/login surfaces.
206
+
207
+ Manual checks that remain:
208
+
209
+ - initial app-managed site output is correct for the expected content/theme;
210
+ - content repository identity and runtime sync are healthy;
211
+ - passkey setup/handoff is completed from the setup email.
212
+
213
+ ## One-user canonical site canary
214
+
215
+ Run this before adding custom site/theme packages or rolling a larger browser/Studio-first cohort.
216
+
217
+ 1. Create or choose a canary cohort with explicit bundles:
218
+
219
+ ```yaml
220
+ bundles:
221
+ - core
222
+ - site
223
+ - publishing
224
+ ```
225
+
226
+ 2. Add exactly one canary user to that cohort.
227
+ 3. For browser/Studio-first onboarding, configure setup email in `users/<handle>.yaml`:
228
+
229
+ ```yaml
230
+ setup:
231
+ delivery: email
232
+ email: user@example.com
233
+ ```
234
+
235
+ 4. Encrypt the user's secrets and commit only the `.age` file.
236
+ 5. Run `bunx brains-ops onboard . <handle>`.
237
+ 6. Run `bunx brains-ops verify-user . <handle>` with no custom site/theme overrides.
238
+ 7. Ask the user to complete passkey setup from the setup email.
239
+ 8. Continue to visual customization only after the canary is healthy.
240
+
241
+ Rollback must restore the prior desired-state revision and prior runtime image together. Never pair canonical config with the retired image, or retired config with the canonical image.
242
+
243
+ ## Hosted site and theme package contract
244
+
245
+ Start with the public [site mockup migration guide](https://github.com/rizom-ai/brains/blob/main/docs/site-mockup-migration.md), then apply these hosted-fleet requirements:
246
+
247
+ - A site package must default-export `defineSite(...)` and import its authoring API only from `@rizom/site`.
248
+ - A theme package must default-export its CSS as a string. Hosted custom themes currently use the `@rizom/*` scope so the fleet image installs them with the site package; `@brains/*` themes are bundled with `@rizom/brain`.
249
+ - Site and custom theme packages must be public npm packages that install without registry credentials.
250
+ - Site, theme, and brain packages publish independently. Hosted configuration requires exact site and external-theme version pins and never derives one package version from another.
251
+ - Keep site structure and theme CSS in separate packages. Do not put private content or secrets in either package.
252
+
253
+ Configure a user in `users/<handle>.yaml`:
254
+
255
+ ```yaml
256
+ siteOverride:
257
+ package: "@rizom/site-example"
258
+ version: <exact-site-version>
259
+ theme: "@rizom/theme-example"
260
+ themeVersion: <exact-theme-version>
261
+ ```
262
+
263
+ Missing external package versions fail desired-state validation. A site override
264
+ produces an isolated per-instance image; it never changes the fleet's shared default
265
+ image. Bundled `@brains/*` themes omit `themeVersion` because they are not installed as
266
+ separate packages.
267
+
268
+ ### Custom-package canary and rollback
269
+
270
+ 1. Confirm the exact site/theme versions are public-installable without npm credentials.
271
+ 2. Apply the exact package names and versions to one healthy canonical site canary.
272
+ 3. Reconcile the canary, push the generated output, and let build/deploy create its site image.
273
+ 4. Run `bunx brains-ops verify-user . <handle>`.
274
+ 5. Manually verify the site, theme, Studio, content sync, and passkey sign-in before adding more users.
275
+
276
+ To roll back, remove or change `siteOverride`, reconcile, and redeploy that user.
277
+ The default image and other users remain untouched.
278
+
279
+ ## Setup email checklist
280
+
281
+ Use this for browser/Studio-first users who should receive their own first-passkey setup link by email.
282
+
283
+ 1. Add setup delivery to the user file:
284
+
285
+ ```yaml
286
+ setup:
287
+ delivery: email
288
+ email: user@example.com
289
+ ```
290
+
291
+ 2. Configure these GitHub Secrets before deploy:
292
+ - `SETUP_EMAIL_API_KEY`
293
+ - `SETUP_EMAIL_FROM`
294
+
295
+ 3. Reconcile/deploy the user or cohort:
296
+ - `bunx brains-ops onboard . <handle>`
297
+ - or `bunx brains-ops reconcile-cohort . <cohort>`
298
+
299
+ 4. Verify the generated `users/<handle>/brain.yaml` contains `auth-service.setupEmail` and `email` interface config.
300
+ 5. Ask the user to complete passkey setup from the email link, then use:
301
+ - Dashboard: `https://<handle>.rizom.ai/`
302
+ - Studio: `https://<handle>.rizom.ai/studio`
303
+
304
+ Notes:
305
+
306
+ - The setup URL is generated and sent by the running brain; operators should not scrape logs or SSH into the instance to retrieve it.
307
+ - The auth service owns setup email dedupe. It should not resend for the same persisted setup token after restart, but should retry failed delivery and resend after token rotation.
308
+ - `SETUP_EMAIL_FROM` is not marked required because fleets without email setup can omit it, but it is required for users with `setup.delivery: email`.
309
+
310
+ ## AT Protocol smoke/config checklist
311
+
312
+ Use this when enabling AT Protocol publishing for a single pilot user.
313
+
314
+ 1. Add the public PDS identifier to the user file. Prefer the account DID as
315
+ the identifier — it survives handle changes. Add `accountDid` too when the
316
+ member wants their handle verified against their subdomain
317
+ (`@<handle>.<domainSuffix>`): the brain then serves it at
318
+ `/.well-known/atproto-did` and Bluesky's "I have my own domain" HTTP
319
+ verification passes with no DNS records.
320
+
321
+ ```yaml
322
+ atproto:
323
+ identifier: did:plc:example123
324
+ accountDid: did:plc:example123
325
+ ```
326
+
327
+ Only for the PDS account designated by the protocol authority's `_lexicon`
328
+ DNS TXT record, also set `lexiconAuthority: true`. Every other fleet user
329
+ must omit it.
330
+
331
+ 2. Put the app password in `users/<handle>.secrets.yaml`:
332
+
333
+ ```yaml
334
+ atprotoAppPassword: <app-password>
335
+ ```
336
+
337
+ 3. Encrypt the per-user secret payload:
338
+ - `bunx brains-ops secrets:encrypt . <handle>`
339
+ 4. Reconcile/deploy the user or cohort:
340
+ - `bunx brains-ops onboard . <handle>`
341
+ - or `bunx brains-ops reconcile-cohort . <cohort>`
342
+ 5. Verify the generated `users/<handle>/brain.yaml` contains `plugins.atproto.identifier` (plus `accountDid` and `lexiconAuthority` when configured) and `appPassword: ${ATPROTO_APP_PASSWORD}`.
343
+
344
+ Notes:
345
+
346
+ - The ATProto identifier and authority flag are public instance config and belong in `users/<handle>.yaml`. Only the DNS-designated authority account may set `lexiconAuthority: true`.
347
+ - The ATProto app password is secret and belongs only in the encrypted per-user secret payload.
348
+ - For smoke deployments, pin only the smoke cohort/user to the released brain version that contains ATProto support.
349
+
350
+ ## Discord application credential checklist
74
351
 
75
352
  Use this when enabling Discord for a pilot user.
76
353
 
77
354
  1. Pick the user handle (for example `smoke`).
78
355
  2. Open the Discord Developer Portal.
79
- 3. Create a **new application** for that user's rover.
356
+ 3. Create a **new application** for that user's brain.
80
357
  4. Add a **Bot** to the application.
81
- 5. Copy the bot token.
82
- 6. Put that value in `.env` or `.env.local` in this repo as `DISCORD_BOT_TOKEN=...` while onboarding that user.
358
+ 5. Copy the bot token, application public key, and application ID.
359
+ 6. Put those values in `.env` or `.env.local` while onboarding that user:
360
+ - `DISCORD_BOT_TOKEN=...`
361
+ - `DISCORD_PUBLIC_KEY=...`
362
+ - `DISCORD_APPLICATION_ID=...`
83
363
  7. Keep `discord.enabled: true` in `users/<handle>.yaml` unless you explicitly want to disable the primary pilot interface.
84
- 8. Encrypt the current per-user secret payload:
364
+ 8. Encrypt the current per-user credential payload:
85
365
  - `bunx brains-ops secrets:encrypt . <handle>`
86
366
  9. Reconcile/deploy the user or cohort:
87
-
88
- - `bunx brains-ops onboard . <handle>`
89
- - or `bunx brains-ops reconcile-cohort . <cohort>`
90
-
91
- 11. In the Discord Developer Portal, generate an install URL and invite the bot to the right server.
92
- 12. Send a test message in Discord and confirm the rover responds.
367
+ - `bunx brains-ops onboard . <handle>`
368
+ - or `bunx brains-ops reconcile-cohort . <cohort>`
369
+ 10. In the Discord Developer Portal, generate an install URL and invite the bot to the right server.
370
+ 11. Send a test message in Discord and confirm the brain responds.
93
371
 
94
372
  Notes:
95
373
 
96
- - Use **one bot token per user/rover**.
97
- - Do not reuse the same Discord bot token across multiple pilot users.
374
+ - Use **one Discord application credential set per user/brain**.
375
+ - Do not reuse the same Discord application across multiple pilot users.
98
376
  - Discord is the default pilot interface moving forward.
99
377
  - The encrypted `users/<handle>.secrets.yaml.age` file is the durable checked-in deploy input; your local env is only the operator staging source.
100
- - MCP is optional and mainly for direct client access or specific testing workflows.
378
+ - Direct MCP client access should use OAuth/passkey-capable clients where possible.
101
379
  - When explaining the content workflow, describe it first as a normal **git repo** of **markdown/text files**.
102
380
  - Position **Obsidian** as optional: it is just one possible editor for those same files, not the default requirement.
103
381