@rizom/ops 0.2.0-alpha.34 → 0.2.0-alpha.340

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (61) hide show
  1. package/README.md +61 -2
  2. package/dist/brains-ops.js +700 -318
  3. package/dist/capability-bundle-migration.d.ts +22 -0
  4. package/dist/cert-bootstrap.d.ts +3 -1
  5. package/dist/content-repo-ref.d.ts +10 -0
  6. package/dist/content-repo.d.ts +1 -0
  7. package/dist/deploy.js +99 -166
  8. package/dist/directory-sync-stress-system.d.ts +75 -0
  9. package/dist/directory-sync-stress.d.ts +107 -0
  10. package/dist/entries/deploy.d.ts +3 -2
  11. package/dist/health-watchdog-smoke.d.ts +54 -0
  12. package/dist/images.d.ts +79 -0
  13. package/dist/index.d.ts +8 -0
  14. package/dist/index.js +776 -308
  15. package/dist/legacy-pilot-migration.d.ts +11 -0
  16. package/dist/load-registry.d.ts +60 -6
  17. package/dist/observed-status.d.ts +1 -1
  18. package/dist/origin-ca.d.ts +1 -1
  19. package/dist/parse-args.d.ts +2 -10
  20. package/dist/preview-domain.d.ts +9 -0
  21. package/dist/push-secrets.d.ts +2 -9
  22. package/dist/push-target.d.ts +1 -2
  23. package/dist/reconcile-dry-run.d.ts +9 -0
  24. package/dist/run-command.d.ts +15 -3
  25. package/dist/run-subprocess.d.ts +1 -6
  26. package/dist/schema.d.ts +123 -162
  27. package/dist/secrets-encrypt.d.ts +7 -13
  28. package/dist/ssh-key-bootstrap.d.ts +1 -26
  29. package/dist/stage-legacy-crossover.d.ts +23 -0
  30. package/dist/stress-command.d.ts +14 -0
  31. package/dist/stress-git-checkout.d.ts +21 -0
  32. package/dist/stress-health-monitor.d.ts +46 -0
  33. package/dist/upgrade.d.ts +10 -0
  34. package/dist/user-add.d.ts +15 -0
  35. package/dist/verify-user.d.ts +22 -0
  36. package/package.json +48 -42
  37. package/templates/rover-pilot/.env.schema +24 -3
  38. package/templates/rover-pilot/.github/actions/varlock-env/action.yml +47 -0
  39. package/templates/rover-pilot/.github/workflows/build.yml +70 -17
  40. package/templates/rover-pilot/.github/workflows/deploy.yml +63 -57
  41. package/templates/rover-pilot/.github/workflows/directory-sync-stress.yml +119 -0
  42. package/templates/rover-pilot/.github/workflows/health-watchdog-smoke.yml +95 -0
  43. package/templates/rover-pilot/.github/workflows/reconcile.yml +10 -4
  44. package/templates/rover-pilot/.github/workflows/upgrade.yml +104 -0
  45. package/templates/rover-pilot/README.md +16 -6
  46. package/templates/rover-pilot/deploy/scripts/decrypt-user-secrets.ts +80 -24
  47. package/templates/rover-pilot/deploy/scripts/helpers.ts +3 -0
  48. package/templates/rover-pilot/deploy/scripts/install-health-watchdog.ts +144 -0
  49. package/templates/rover-pilot/deploy/scripts/resolve-missing-images.ts +13 -0
  50. package/templates/rover-pilot/deploy/scripts/resolve-user-config.ts +44 -9
  51. package/templates/rover-pilot/deploy/scripts/sync-content-repo.ts +51 -47
  52. package/templates/rover-pilot/deploy/scripts/update-dns.ts +14 -4
  53. package/templates/rover-pilot/deploy/scripts/validate-secrets.ts +12 -1
  54. package/templates/rover-pilot/docs/canonical-crossover-record.md +107 -0
  55. package/templates/rover-pilot/docs/onboarding-checklist.md +28 -17
  56. package/templates/rover-pilot/docs/operator-playbook.md +270 -29
  57. package/templates/rover-pilot/docs/user-onboarding.md +48 -463
  58. package/templates/rover-pilot/pilot.yaml +7 -4
  59. package/templates/rover-pilot/.kamal/hooks/pre-deploy +0 -9
  60. package/templates/rover-pilot/deploy/Dockerfile +0 -30
  61. package/templates/rover-pilot/deploy/kamal/deploy.yml +0 -40
@@ -9,30 +9,108 @@ Treat these as checked-in deploy artifacts in the pilot repo:
9
9
  - `deploy/scripts/`
10
10
  - `.github/workflows/build.yml`
11
11
  - `.github/workflows/deploy.yml`
12
+ - `.github/workflows/directory-sync-stress.yml`
13
+ - `.github/workflows/health-watchdog-smoke.yml`
12
14
  - `.github/workflows/reconcile.yml`
13
15
 
14
16
  `.env.schema` is the single source of truth for required and sensitive deploy vars.
15
17
  The deploy scripts and workflows should read from that contract instead of inventing a second list.
16
18
 
17
- The shared pilot image tag is `brain-${brainVersion}`:
19
+ The default pilot image tag is `brain-${brainVersion}`:
18
20
 
19
- - build publishes `brain-${brainVersion}`
21
+ - build publishes `brain-${brainVersion}` for users without a site override
22
+ - a site override gets an isolated `brain-${brainVersion}-sites-${packageHash}` image
20
23
  - generated `users/<handle>/.env` carries `BRAIN_VERSION=<brainVersion>`
21
- - deploy sets `VERSION=brain-${brainVersion}`
24
+ - build and deploy derive the same effective image tag from the resolved registry
22
25
 
23
26
  ## Version bump flow
24
27
 
25
28
  When `pilot.yaml.brainVersion` changes and you push:
26
29
 
27
- 1. build publishes the new shared image tag
30
+ 1. build publishes the new default image and any required site images
28
31
  2. reconcile refreshes generated `users/<handle>/.env`
29
32
  3. deploy runs for handles whose generated config changed
30
33
  4. generated file commits happen once in a final aggregation step after the deploy matrix finishes
31
34
 
35
+ Every external site and theme package has its own exact version pin. A cohort or
36
+ pilot brain-version bump never changes those package versions implicitly; update each
37
+ pin deliberately from reviewed package and image evidence.
38
+
32
39
  When a push changes only deploy contract files and no generated `users/<handle>/.env` or `users/<handle>/brain.yaml` files, the deploy workflow exits through its explicit no-op path and prints `No affected user configs; skipping deploy.`
33
40
 
34
41
  They are scaffolded from `@rizom/ops`, then versioned in this repo like any other deploy contract.
35
42
 
43
+ ## Canonical contract crossover maintenance window
44
+
45
+ Do not run this procedure without explicit operator approval. The canonical desired state, canonical `@rizom/ops`, and unified runtime image form one contract and must move or roll back together. Complete `docs/canonical-crossover-record.md` as the approval evidence without adding secret values.
46
+
47
+ Before the window, record and review:
48
+
49
+ - the prior pilot commit and exact `@rizom/ops` version;
50
+ - every prior runtime image tag and immutable digest;
51
+ - the reviewed canonical pilot commit;
52
+ - the exact unified `@rizom/brain` and `@rizom/ops` versions;
53
+ - every unified image tag and immutable digest;
54
+ - canary-first rollout order, followed by the remaining cohorts;
55
+ - expected `/health/operate` version, unauthenticated MCP response, site marker, and content repository identity for each posture.
56
+
57
+ Run `bunx brains-ops reconcile-all <canonical-review-copy> --dry-run` against the isolated review copy. The command blocks external content-repository access, leaves the review copy untouched, lists both passes' changed files, and must report second-pass zero drift. Reconciliation owns generated per-user config; only `render` owns the observational `views/users.md` projection.
58
+
59
+ During the approved window:
60
+
61
+ 1. Freeze unrelated merges and releases. Wait for active Build, Reconcile, and Deploy runs to finish, then disable all three pilot workflows with `gh workflow disable build.yml`, `gh workflow disable reconcile.yml`, and `gh workflow disable deploy.yml`.
62
+ 2. Publish and verify the reviewed unified runtime and matching ops artifacts. Do not update pilot desired state until the exact versions are installable and the expected images can be built.
63
+ 3. Apply the reviewed canonical pilot revision while automation remains disabled. Confirm repository names, server/domain identity, content repositories, secret selectors, image names, and tag identity against the review diff.
64
+ 4. Enable only Build, run it for the canonical desired-state revision, and record every resulting image digest. Stop if the observed digest set differs from the cutover record.
65
+ 5. Enable only Reconcile, run it once, and review its generated per-user config commit. It must not rewrite `views/users.md`, and no generated file may combine canonical config with a retired image version.
66
+ 6. Enable Deploy and deploy one handle at a time in the approved order. After each deploy, run `bunx brains-ops verify-user . <handle>`, render observed fleet status, and complete the manual identity, content-sync, and app-managed site checks.
67
+ 7. Run Reconcile a second time. Require no reconciler-owned generated diff and no deploy work before re-enabling normal automation and lifting the merge/release freeze. Observed status rendering remains separate from this convergence gate.
68
+
69
+ If any gate fails, disable all three workflows again. Restore the prior pilot desired-state and dependency revision, reconcile with the prior ops version, and redeploy the prior image tag/digest as one rollback pair. Verify the prior `/health/operate` version and identity/content/site checks before re-enabling automation. Never restore only config or only an image.
70
+
71
+ ## Directory-sync stress gate
72
+
73
+ Use the manual `Directory Sync Stress` workflow only against a disposable smoke user. It refuses a target unless the handle, domain, and content repository all identify smoke, the confirmation input exactly matches `stress:<handle>`, and the user desired state declares the hermetic posture below. Reconcile and deploy this posture before running the workload:
74
+
75
+ ```yaml
76
+ embeddingEnabled: false
77
+ topicExtractionEnabled: false
78
+ skillDerivationEnabled: false
79
+ swotDerivationEnabled: false
80
+ ```
81
+
82
+ Before authorizing a workload, dispatch the workflow once with `verify_only: true`. That mode loads the same Bitwarden/Varlock content credential, clones the smoke content repository, and runs `git push --dry-run` against a temporary stress ref. It creates no ref, performs no content write, does not contact the deployed runtime, and skips cleanup because no probes were created.
83
+
84
+ Profiles are deterministic and reversible:
85
+
86
+ - `regression`: 20 probes;
87
+ - `load`: ramps to 350 probes, updates all, renames 100, updates again, then deletes all;
88
+ - `stress`: ramps to 700 probes and renames 200 before cleanup.
89
+
90
+ The workflow loads operator credentials through Bitwarden/Varlock, but it is separate from Deploy and cannot deploy an image. It creates a rollback branch before the first content write, gates on health timeouts, watchdog restarts, and external AI usage during the monitored workload window, preserves warmup and cleanup samples as evidence, uploads JSON/Markdown/runtime artifacts, and runs an independent idempotent cleanup job with `if: always()`. Once cleanup confirms that no probes remain, it also prunes retained `ops/directory-sync-stress-backup-*` branches; if probes remain, the branches stay available for recovery.
91
+
92
+ Treat any gated health failure, restart, OOM, residual probe, or entity-baseline drift as a failed gate. Do not restart the target during measurement. Recovery is a separate operator action after evidence collection.
93
+
94
+ ## Health watchdog smoke gate
95
+
96
+ Use the manual `Health Watchdog Smoke` workflow only after the `smoke` fleet user has completed a normal Deploy run. The workflow shares the `deploy-<handle>` concurrency group, resolves the server from pilot desired state through Hetzner, loads the existing deploy SSH key through Bitwarden/Varlock, and refuses targets whose handle and domain do not identify smoke. Confirm with `watchdog-smoke:<handle>`.
97
+
98
+ The smoke does not deploy or replace the installed systemd units. It verifies that the installed watchdog exactly matches the packaged canonical payload, then exercises that installed script. It also verifies the deployed rover image label, exact-label selector, active timer, and `/health/live` Docker healthcheck. While holding the timer's global lock, it creates temporary containers for eligible unhealthy, unrelated service-labelled unhealthy, false-labelled unhealthy, and eligible healthy cases. It then requires exactly three eligible restarts followed by budget suppression, diagnostics and state only for the eligible fixture, and no restart of the deployed rover or ineligible fixtures.
99
+
100
+ The primary job uploads remote incident, state, and command evidence. An independent `if: always()` cleanup job removes deterministic fixture names and remote temporary files using the same workflow run ID. Treat missing evidence, unexpected eligibility, deployed-container movement, cleanup failure, or any non-smoke target rejection as a failed gate.
101
+
102
+ ## Stale deploy lock recovery
103
+
104
+ Kamal intentionally leaves its remote deploy lock in place when a deployment is cancelled or interrupted. Confirm that no deployment for the user is still active before releasing the lock, then use the deploy workflow's explicit recovery input:
105
+
106
+ ```sh
107
+ gh workflow run Deploy --ref main \
108
+ -f handle=<handle> \
109
+ -f release_stale_lock=true
110
+ ```
111
+
112
+ Recovery is opt-in and scoped to one handle. Normal push, reconcile, and manual deploy runs never remove a lock automatically.
113
+
36
114
  ## Bootstrap flow
37
115
 
38
116
  For this fleet, operator-local secret material remains the source of truth during onboarding and rotation. The repo stores encrypted per-user secrets, not raw values.
@@ -53,51 +131,214 @@ Preview hosts use the shape `<handle>-preview.rizom.ai`, so one wildcard origin
53
131
 
54
132
  ## Upgrading operator behavior
55
133
 
56
- When `@rizom/ops` changes the scaffolded deploy contract:
134
+ The pilot repository pins `@rizom/ops` in `package.json`. The scheduled and manually dispatched Upgrade workflow owns routine upgrades to that pin. It refreshes the scaffold on a branch and opens a reviewable PR; it does not change runtime desired state or authorize a deployment.
135
+
136
+ Because scaffold refreshes can update `.github/workflows/*`, the workflow must not push with its Actions `GITHUB_TOKEN`. Configure a dedicated GitHub App:
137
+
138
+ 1. Install it only on this pilot repository.
139
+ 2. Grant repository permissions `Contents: Read and write`, `Pull requests: Read and write`, and `Workflows: Read and write`; grant nothing else.
140
+ 3. Store its App ID as the repository Actions variable `OPS_UPGRADE_APP_ID`.
141
+ 4. Store its private key as the repository Actions secret `OPS_UPGRADE_APP_PRIVATE_KEY`.
142
+
143
+ The workflow explicitly requests only those three permissions. Checkout persists no credential. After the freshly published `@rizom/ops` finishes, the workflow checks whether it produced a change; only then does it mint a short-lived, repository-scoped App token for the push-and-open-PR steps. A credential that can rewrite `.github/workflows/` therefore does not exist while upgraded package code runs.
144
+
145
+ If `OPS_UPGRADE_APP_ID` is unset, or token creation, branch push, or PR creation fails, the run stops with a non-zero status. Repair the CI credential path; never fall back to an operator's personal SSH key or token.
146
+
147
+ Adopting this credential flow in an existing pilot repository requires one explicitly reviewed bootstrap PR because the old Upgrade workflow cannot update itself. After that merge, routine upgrades run entirely in CI:
148
+
149
+ 1. dispatch Upgrade with an exact version, or let its schedule select `latest`;
150
+ 2. review the generated package, lockfile, deploy-script, and workflow diff;
151
+ 3. merge the upgrade PR only after its checks pass;
152
+ 4. change runtime desired state separately through the approved canary or fleet rollout flow.
153
+
154
+ ## Canonical verification notes
155
+
156
+ Use the verification script after deploy:
157
+
158
+ ```sh
159
+ bunx brains-ops verify-user . <handle>
160
+ ```
161
+
162
+ For every bundle posture it checks:
163
+
164
+ - `https://<handle>.rizom.ai/health/operate` returns `200`;
165
+ - unauthenticated `POST https://<handle>.rizom.ai/mcp` returns the expected auth failure;
166
+ - background jobs are not repeatedly failing, except for missing optional integrations.
167
+
168
+ A `core`-only instance is MCP-only; a bare `GET /` may return `401` without indicating a bad deploy. When `site` is selected, verification also checks the browser and Studio/login surfaces.
169
+
170
+ Manual checks that remain:
57
171
 
58
- 1. bump `@rizom/ops` in `package.json`
59
- 2. rerun the relevant scaffold/reconcile flow
60
- 3. review the resulting changes to `.env.schema`, `deploy/scripts/`, and workflows in git
61
- 4. commit the updated deploy artifacts together
172
+ - initial app-managed site output is correct for the expected content/theme;
173
+ - content repository identity and runtime sync are healthy;
174
+ - passkey setup/handoff is completed from the setup email.
62
175
 
63
- ## Rover-core verification notes
176
+ ## One-user canonical site canary
64
177
 
65
- Rover core is MCP-only. Do not expect the bare domain to serve a website.
178
+ Run this before adding custom site/theme packages or rolling a larger browser/Studio-first cohort.
66
179
 
67
- Use these checks after deploy:
180
+ 1. Create or choose a canary cohort with explicit bundles:
68
181
 
69
- - `https://<handle>.rizom.ai/health` should return `200`
70
- - unauthenticated `POST https://<handle>.rizom.ai/mcp` should return `401 Unauthorized: Bearer token required`
71
- - a bare `GET /` may also return `401`; that is expected for rover core and does not indicate a bad deploy
182
+ ```yaml
183
+ bundles:
184
+ - core
185
+ - site
186
+ - publishing
187
+ ```
72
188
 
73
- ## Discord bot token checklist
189
+ 2. Add exactly one canary user to that cohort.
190
+ 3. For browser/Studio-first onboarding, configure setup email in `users/<handle>.yaml`:
191
+
192
+ ```yaml
193
+ setup:
194
+ delivery: email
195
+ email: user@example.com
196
+ ```
197
+
198
+ 4. Encrypt the user's secrets and commit only the `.age` file.
199
+ 5. Run `bunx brains-ops onboard . <handle>`.
200
+ 6. Run `bunx brains-ops verify-user . <handle>` with no custom site/theme overrides.
201
+ 7. Ask the user to complete passkey setup from the setup email.
202
+ 8. Continue to visual customization only after the canary is healthy.
203
+
204
+ Rollback must restore the prior desired-state revision and prior runtime image together. Never pair canonical config with the retired image, or retired config with the canonical image.
205
+
206
+ ## Hosted site and theme package contract
207
+
208
+ Start with the public [site mockup migration guide](https://github.com/rizom-ai/brains/blob/main/docs/site-mockup-migration.md), then apply these hosted-fleet requirements:
209
+
210
+ - A site package must default-export `defineSite(...)` and import its authoring API only from `@rizom/site`.
211
+ - A theme package must default-export its CSS as a string. Hosted custom themes currently use the `@rizom/*` scope so the fleet image installs them with the site package; `@brains/*` themes are bundled with `@rizom/brain`.
212
+ - Site and custom theme packages must be public npm packages that install without registry credentials.
213
+ - Site, theme, and brain packages publish independently. Hosted configuration requires exact site and external-theme version pins and never derives one package version from another.
214
+ - Keep site structure and theme CSS in separate packages. Do not put private content or secrets in either package.
215
+
216
+ Configure a user in `users/<handle>.yaml`:
217
+
218
+ ```yaml
219
+ siteOverride:
220
+ package: "@rizom/site-example"
221
+ version: <exact-site-version>
222
+ theme: "@rizom/theme-example"
223
+ themeVersion: <exact-theme-version>
224
+ ```
225
+
226
+ Missing external package versions fail desired-state validation. A site override
227
+ produces an isolated per-instance image; it never changes the fleet's shared default
228
+ image. Bundled `@brains/*` themes omit `themeVersion` because they are not installed as
229
+ separate packages.
230
+
231
+ ### Custom-package canary and rollback
232
+
233
+ 1. Confirm the exact site/theme versions are public-installable without npm credentials.
234
+ 2. Apply the exact package names and versions to one healthy canonical site canary.
235
+ 3. Reconcile the canary, push the generated output, and let build/deploy create its site image.
236
+ 4. Run `bunx brains-ops verify-user . <handle>`.
237
+ 5. Manually verify the site, theme, Studio, content sync, and passkey sign-in before adding more users.
238
+
239
+ To roll back, remove or change `siteOverride`, reconcile, and redeploy that user.
240
+ The default image and other users remain untouched.
241
+
242
+ ## Setup email checklist
243
+
244
+ Use this for browser/Studio-first users who should receive their own first-passkey setup link by email.
245
+
246
+ 1. Add setup delivery to the user file:
247
+
248
+ ```yaml
249
+ setup:
250
+ delivery: email
251
+ email: user@example.com
252
+ ```
253
+
254
+ 2. Configure these GitHub Secrets before deploy:
255
+ - `SETUP_EMAIL_API_KEY`
256
+ - `SETUP_EMAIL_FROM`
257
+
258
+ 3. Reconcile/deploy the user or cohort:
259
+ - `bunx brains-ops onboard . <handle>`
260
+ - or `bunx brains-ops reconcile-cohort . <cohort>`
261
+
262
+ 4. Verify the generated `users/<handle>/brain.yaml` contains `auth-service.setupEmail` and `email` interface config.
263
+ 5. Ask the user to complete passkey setup from the email link, then use:
264
+ - Dashboard: `https://<handle>.rizom.ai/`
265
+ - Studio: `https://<handle>.rizom.ai/studio`
266
+
267
+ Notes:
268
+
269
+ - The setup URL is generated and sent by the running brain; operators should not scrape logs or SSH into the instance to retrieve it.
270
+ - The auth service owns setup email dedupe. It should not resend for the same persisted setup token after restart, but should retry failed delivery and resend after token rotation.
271
+ - `SETUP_EMAIL_FROM` is not marked required because fleets without email setup can omit it, but it is required for users with `setup.delivery: email`.
272
+
273
+ ## AT Protocol smoke/config checklist
274
+
275
+ Use this when enabling AT Protocol publishing for a single pilot user.
276
+
277
+ 1. Add the public PDS identifier to the user file. Prefer the account DID as
278
+ the identifier — it survives handle changes. Add `accountDid` too when the
279
+ member wants their handle verified against their subdomain
280
+ (`@<handle>.<domainSuffix>`): the brain then serves it at
281
+ `/.well-known/atproto-did` and Bluesky's "I have my own domain" HTTP
282
+ verification passes with no DNS records.
283
+
284
+ ```yaml
285
+ atproto:
286
+ identifier: did:plc:example123
287
+ accountDid: did:plc:example123
288
+ ```
289
+
290
+ Only for the PDS account designated by the protocol authority's `_lexicon`
291
+ DNS TXT record, also set `lexiconAuthority: true`. Every other fleet user
292
+ must omit it.
293
+
294
+ 2. Put the app password in `users/<handle>.secrets.yaml`:
295
+
296
+ ```yaml
297
+ atprotoAppPassword: <app-password>
298
+ ```
299
+
300
+ 3. Encrypt the per-user secret payload:
301
+ - `bunx brains-ops secrets:encrypt . <handle>`
302
+ 4. Reconcile/deploy the user or cohort:
303
+ - `bunx brains-ops onboard . <handle>`
304
+ - or `bunx brains-ops reconcile-cohort . <cohort>`
305
+ 5. Verify the generated `users/<handle>/brain.yaml` contains `plugins.atproto.identifier` (plus `accountDid` and `lexiconAuthority` when configured) and `appPassword: ${ATPROTO_APP_PASSWORD}`.
306
+
307
+ Notes:
308
+
309
+ - The ATProto identifier and authority flag are public instance config and belong in `users/<handle>.yaml`. Only the DNS-designated authority account may set `lexiconAuthority: true`.
310
+ - The ATProto app password is secret and belongs only in the encrypted per-user secret payload.
311
+ - For smoke deployments, pin only the smoke cohort/user to the released brain version that contains ATProto support.
312
+
313
+ ## Discord application credential checklist
74
314
 
75
315
  Use this when enabling Discord for a pilot user.
76
316
 
77
317
  1. Pick the user handle (for example `smoke`).
78
318
  2. Open the Discord Developer Portal.
79
- 3. Create a **new application** for that user's rover.
319
+ 3. Create a **new application** for that user's brain.
80
320
  4. Add a **Bot** to the application.
81
- 5. Copy the bot token.
82
- 6. Put that value in `.env` or `.env.local` in this repo as `DISCORD_BOT_TOKEN=...` while onboarding that user.
321
+ 5. Copy the bot token, application public key, and application ID.
322
+ 6. Put those values in `.env` or `.env.local` while onboarding that user:
323
+ - `DISCORD_BOT_TOKEN=...`
324
+ - `DISCORD_PUBLIC_KEY=...`
325
+ - `DISCORD_APPLICATION_ID=...`
83
326
  7. Keep `discord.enabled: true` in `users/<handle>.yaml` unless you explicitly want to disable the primary pilot interface.
84
- 8. Encrypt the current per-user secret payload:
327
+ 8. Encrypt the current per-user credential payload:
85
328
  - `bunx brains-ops secrets:encrypt . <handle>`
86
329
  9. Reconcile/deploy the user or cohort:
87
-
88
- - `bunx brains-ops onboard . <handle>`
89
- - or `bunx brains-ops reconcile-cohort . <cohort>`
90
-
91
- 11. In the Discord Developer Portal, generate an install URL and invite the bot to the right server.
92
- 12. Send a test message in Discord and confirm the rover responds.
330
+ - `bunx brains-ops onboard . <handle>`
331
+ - or `bunx brains-ops reconcile-cohort . <cohort>`
332
+ 10. In the Discord Developer Portal, generate an install URL and invite the bot to the right server.
333
+ 11. Send a test message in Discord and confirm the brain responds.
93
334
 
94
335
  Notes:
95
336
 
96
- - Use **one bot token per user/rover**.
97
- - Do not reuse the same Discord bot token across multiple pilot users.
337
+ - Use **one Discord application credential set per user/brain**.
338
+ - Do not reuse the same Discord application across multiple pilot users.
98
339
  - Discord is the default pilot interface moving forward.
99
340
  - The encrypted `users/<handle>.secrets.yaml.age` file is the durable checked-in deploy input; your local env is only the operator staging source.
100
- - MCP is optional and mainly for direct client access or specific testing workflows.
341
+ - Direct MCP client access should use OAuth/passkey-capable clients where possible.
101
342
  - When explaining the content workflow, describe it first as a normal **git repo** of **markdown/text files**.
102
343
  - Position **Obsidian** as optional: it is just one possible editor for those same files, not the default requirement.
103
344