@rizom/ops 0.2.0-alpha.35 → 0.2.0-alpha.350
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +62 -2
- package/dist/brains-ops.js +785 -304
- package/dist/capability-bundle-migration.d.ts +22 -0
- package/dist/cert-bootstrap.d.ts +3 -1
- package/dist/content-repo-ref.d.ts +10 -0
- package/dist/content-repo.d.ts +1 -0
- package/dist/deploy.js +99 -166
- package/dist/directory-sync-stress-system.d.ts +75 -0
- package/dist/directory-sync-stress.d.ts +107 -0
- package/dist/entries/deploy.d.ts +3 -2
- package/dist/health-watchdog-smoke.d.ts +54 -0
- package/dist/images.d.ts +79 -0
- package/dist/index.d.ts +8 -0
- package/dist/index.js +784 -302
- package/dist/legacy-pilot-migration.d.ts +11 -0
- package/dist/legacy-projection-job-recovery.d.ts +29 -0
- package/dist/load-registry.d.ts +60 -6
- package/dist/observed-status.d.ts +1 -1
- package/dist/origin-ca.d.ts +1 -1
- package/dist/parse-args.d.ts +2 -10
- package/dist/preview-domain.d.ts +9 -0
- package/dist/push-secrets.d.ts +2 -9
- package/dist/push-target.d.ts +1 -2
- package/dist/reconcile-dry-run.d.ts +9 -0
- package/dist/run-command.d.ts +17 -3
- package/dist/run-subprocess.d.ts +1 -6
- package/dist/schema.d.ts +123 -162
- package/dist/secrets-encrypt.d.ts +7 -13
- package/dist/ssh-key-bootstrap.d.ts +1 -26
- package/dist/stage-legacy-crossover.d.ts +23 -0
- package/dist/stress-command.d.ts +14 -0
- package/dist/stress-git-checkout.d.ts +21 -0
- package/dist/stress-health-monitor.d.ts +46 -0
- package/dist/upgrade.d.ts +10 -0
- package/dist/user-add.d.ts +15 -0
- package/dist/verify-user.d.ts +22 -0
- package/package.json +50 -42
- package/templates/rover-pilot/.env.schema +24 -3
- package/templates/rover-pilot/.github/actions/varlock-env/action.yml +47 -0
- package/templates/rover-pilot/.github/workflows/build.yml +70 -17
- package/templates/rover-pilot/.github/workflows/deploy.yml +71 -57
- package/templates/rover-pilot/.github/workflows/directory-sync-stress.yml +119 -0
- package/templates/rover-pilot/.github/workflows/health-watchdog-smoke.yml +95 -0
- package/templates/rover-pilot/.github/workflows/reconcile.yml +10 -4
- package/templates/rover-pilot/.github/workflows/upgrade.yml +104 -0
- package/templates/rover-pilot/README.md +16 -6
- package/templates/rover-pilot/deploy/scripts/create-predeploy-backup.ts +873 -0
- package/templates/rover-pilot/deploy/scripts/decrypt-user-secrets.ts +80 -24
- package/templates/rover-pilot/deploy/scripts/helpers.ts +3 -0
- package/templates/rover-pilot/deploy/scripts/install-health-watchdog.ts +144 -0
- package/templates/rover-pilot/deploy/scripts/resolve-missing-images.ts +13 -0
- package/templates/rover-pilot/deploy/scripts/resolve-user-config.ts +44 -9
- package/templates/rover-pilot/deploy/scripts/sync-content-repo.ts +51 -47
- package/templates/rover-pilot/deploy/scripts/update-dns.ts +14 -4
- package/templates/rover-pilot/deploy/scripts/validate-secrets.ts +12 -1
- package/templates/rover-pilot/docs/canonical-crossover-record.md +107 -0
- package/templates/rover-pilot/docs/onboarding-checklist.md +28 -17
- package/templates/rover-pilot/docs/operator-playbook.md +307 -29
- package/templates/rover-pilot/docs/user-onboarding.md +48 -463
- package/templates/rover-pilot/pilot.yaml +7 -4
- package/templates/rover-pilot/.kamal/hooks/pre-deploy +0 -9
- package/templates/rover-pilot/deploy/Dockerfile +0 -30
- package/templates/rover-pilot/deploy/kamal/deploy.yml +0 -40
|
@@ -4,28 +4,39 @@
|
|
|
4
4
|
2. Run `bunx brains-ops age-key:bootstrap <repo> --push-to gh`.
|
|
5
5
|
3. Fill in `pilot.yaml`.
|
|
6
6
|
- keep your pinned `brainVersion`
|
|
7
|
-
- confirm shared selectors for `aiApiKey`, `gitSyncToken`, and `
|
|
7
|
+
- confirm shared selectors for `aiApiKey`, `gitSyncToken`, and `contentRepoAdminToken`
|
|
8
|
+
- use different tokens for `contentRepoAdminToken` and `gitSyncToken`: admin creates/checks content repos; sync is used by runtime directory-sync
|
|
8
9
|
- confirm `agePublicKey`
|
|
9
|
-
4.
|
|
10
|
-
-
|
|
11
|
-
-
|
|
12
|
-
|
|
10
|
+
4. Run `bunx brains-ops user:add <repo> <handle> --cohort <cohort>`.
|
|
11
|
+
- Web chat is the primary interface; it needs no per-user setup beyond the passkey.
|
|
12
|
+
- `user:add` currently writes `discord: enabled: true`; set it to `false` unless the user's cohort actually uses Discord.
|
|
13
|
+
- if the user should be an anchor on Discord, add `--anchor-id <discord-user-id>`.
|
|
14
|
+
- the command creates `users/<handle>.yaml`, `users/<handle>.secrets.yaml`, and the cohort membership without duplicating existing entries.
|
|
15
|
+
5. Edit the generated user file if the anchor profile needs richer metadata.
|
|
16
|
+
- Set `setup.delivery: email` and `setup.email` so the user gets the passkey setup email — this is the default onboarding path.
|
|
17
|
+
- For ATProto publishing, add `atproto.identifier` to the user file; put only `atprotoAppPassword` in the per-user secrets file.
|
|
18
|
+
- Ensure `SETUP_EMAIL_API_KEY` and `SETUP_EMAIL_FROM` exist as GitHub Secrets before deploying any email-setup user.
|
|
13
19
|
6. Run `bunx brains-ops render <repo>`.
|
|
14
20
|
7. Run `bunx brains-ops ssh-key:bootstrap <repo> --push-to gh`.
|
|
15
21
|
8. Run `bunx brains-ops cert:bootstrap <repo> --push-to gh`.
|
|
16
|
-
9. Keep raw user secret material locally for now (`.env.local`, file-backed env vars, or equivalent local inputs).
|
|
22
|
+
9. Keep raw user secret material locally for now (`.env.local`, file-backed env vars, or equivalent local inputs), including `CONTENT_REPO_ADMIN_TOKEN` for operator onboarding.
|
|
17
23
|
10. Run `bunx brains-ops secrets:encrypt <repo> <handle>`.
|
|
18
24
|
11. Commit and push `users/<handle>.secrets.yaml.age`.
|
|
19
25
|
12. Run `bunx brains-ops onboard <repo> <handle>`.
|
|
20
|
-
13. Verify the deployed
|
|
21
|
-
- `https://<handle>.rizom.ai/health` returns `200`
|
|
22
|
-
-
|
|
23
|
-
|
|
24
|
-
|
|
25
|
-
|
|
26
|
+
13. Verify the deployed canonical contract:
|
|
27
|
+
- `https://<handle>.rizom.ai/health/operate` returns `200`
|
|
28
|
+
- `https://<handle>.rizom.ai/chat` loads the web chat and accepts passkey sign-in
|
|
29
|
+
- unauthenticated `POST https://<handle>.rizom.ai/mcp` returns the expected auth failure
|
|
30
|
+
- content repo exists and runtime sync is healthy
|
|
31
|
+
- background jobs are not repeatedly failing, except for expected missing optional integrations
|
|
32
|
+
- when `site` is selected, the browser/Studio surfaces load and the initial app-managed site build completes
|
|
33
|
+
14. For fleet upgrades, edit `pilot.yaml.brainVersion` and push once; CI rebuilds the required default/site image tags, refreshes generated user env files, and redeploys affected users. Every external site and theme package keeps its own required exact version pin and never follows the brain version implicitly.
|
|
34
|
+
15. Confirm the user received the setup email, registered their passkey, and can sign in to web chat at `https://<handle>.rizom.ai/chat`. That completes the default onboarding; everything below is per-cohort extras.
|
|
35
|
+
16. Hand over the browser surfaces:
|
|
36
|
+
- Chat (primary): `https://<handle>.rizom.ai/chat`
|
|
26
37
|
- Dashboard: `https://<handle>.rizom.ai/`
|
|
27
|
-
-
|
|
28
|
-
|
|
29
|
-
|
|
30
|
-
|
|
31
|
-
|
|
38
|
+
- Studio: `https://<handle>.rizom.ai/studio`, plus GitHub token guidance if Studio editing is part of their cohort
|
|
39
|
+
17. For Discord-enabled cohorts, hand the Discord setup details to the user as a secondary chat surface.
|
|
40
|
+
18. If they need direct client access (MCP), use OAuth/passkey-capable clients where possible.
|
|
41
|
+
19. If you are also giving them a content repo workflow, describe it as optional and frame git/Obsidian as an advanced file-based path, not the default.
|
|
42
|
+
20. Send `docs/user-onboarding.md` to the user as the pilot handoff guide.
|
|
@@ -9,30 +9,145 @@ Treat these as checked-in deploy artifacts in the pilot repo:
|
|
|
9
9
|
- `deploy/scripts/`
|
|
10
10
|
- `.github/workflows/build.yml`
|
|
11
11
|
- `.github/workflows/deploy.yml`
|
|
12
|
+
- `.github/workflows/directory-sync-stress.yml`
|
|
13
|
+
- `.github/workflows/health-watchdog-smoke.yml`
|
|
12
14
|
- `.github/workflows/reconcile.yml`
|
|
13
15
|
|
|
14
16
|
`.env.schema` is the single source of truth for required and sensitive deploy vars.
|
|
15
17
|
The deploy scripts and workflows should read from that contract instead of inventing a second list.
|
|
16
18
|
|
|
17
|
-
The
|
|
19
|
+
The default pilot image tag is `brain-${brainVersion}`:
|
|
18
20
|
|
|
19
|
-
- build publishes `brain-${brainVersion}`
|
|
21
|
+
- build publishes `brain-${brainVersion}` for users without a site override
|
|
22
|
+
- a site override gets an isolated `brain-${brainVersion}-sites-${packageHash}` image
|
|
20
23
|
- generated `users/<handle>/.env` carries `BRAIN_VERSION=<brainVersion>`
|
|
21
|
-
- deploy
|
|
24
|
+
- build and deploy derive the same effective image tag from the resolved registry
|
|
22
25
|
|
|
23
26
|
## Version bump flow
|
|
24
27
|
|
|
25
28
|
When `pilot.yaml.brainVersion` changes and you push:
|
|
26
29
|
|
|
27
|
-
1. build publishes the new
|
|
30
|
+
1. build publishes the new default image and any required site images
|
|
28
31
|
2. reconcile refreshes generated `users/<handle>/.env`
|
|
29
32
|
3. deploy runs for handles whose generated config changed
|
|
30
33
|
4. generated file commits happen once in a final aggregation step after the deploy matrix finishes
|
|
31
34
|
|
|
35
|
+
Every external site and theme package has its own exact version pin. A cohort or
|
|
36
|
+
pilot brain-version bump never changes those package versions implicitly; update each
|
|
37
|
+
pin deliberately from reviewed package and image evidence.
|
|
38
|
+
|
|
32
39
|
When a push changes only deploy contract files and no generated `users/<handle>/.env` or `users/<handle>/brain.yaml` files, the deploy workflow exits through its explicit no-op path and prints `No affected user configs; skipping deploy.`
|
|
33
40
|
|
|
34
41
|
They are scaffolded from `@rizom/ops`, then versioned in this repo like any other deploy contract.
|
|
35
42
|
|
|
43
|
+
## Verified pre-deploy rollback snapshots
|
|
44
|
+
|
|
45
|
+
Deploy creates and verifies one target-scoped rollback snapshot after SSH and image readiness, immediately before optional stale-lock release and `kamal setup`. A snapshot failure stops the workflow before container replacement. A genuinely new server with no runtime or persistent state reports `not applicable`; persistent state without an identifiable runtime fails closed.
|
|
46
|
+
|
|
47
|
+
Each snapshot contains transaction-consistent online captures of the six canonical SQLite databases, database hashes and `quick_check` results, the deployed `brain.yaml`, sanitized container and mount metadata, and the Directory Sync checkout. Git capture includes all refs, an observed remote head, staged and unstaged binary patches, and mode-preserving archives of untracked and ignored files. Capture performs no checkout mutation or remote write.
|
|
48
|
+
|
|
49
|
+
Verified snapshots are finalized atomically under:
|
|
50
|
+
|
|
51
|
+
```text
|
|
52
|
+
/opt/brain-state/backups/predeploy-<handle>-<target>-<UTC timestamp>/
|
|
53
|
+
```
|
|
54
|
+
|
|
55
|
+
Files use mode `0600` and the directory uses `0700`. A `.incomplete` directory is diagnostic evidence, never a rollback point. Checksums and the versioned manifest must verify before finalization. The default retention is five verified snapshots per target; incomplete or unverified directories do not count, and the snapshot created by the current deployment is never pruned.
|
|
56
|
+
|
|
57
|
+
These are same-server rollback snapshots, not off-host disaster recovery. There is no normal skip and no automatic restore. Restoration requires separate approval: stop replacement/application processes, recheck the selected snapshot's checksums, preserve the current state, restore databases and exact Git state, then validate health, queues, Git checkpoints, durable export intents, preview, and production output before reopening the target.
|
|
58
|
+
|
|
59
|
+
## Retired projection job recovery
|
|
60
|
+
|
|
61
|
+
A pre-scheduler projection job can remain active after its handler type is retired. Current runtime health reports active job types missing from the finalized execution inventory as degraded. Do not restart repeatedly or relax the snapshot idle gate.
|
|
62
|
+
|
|
63
|
+
Use the recovery command only after read-only evidence proves all of the following for one exact row:
|
|
64
|
+
|
|
65
|
+
- its type is one of the command's fixed legacy projection types;
|
|
66
|
+
- the installed runtime no longer contains that handler;
|
|
67
|
+
- attempt, worker-session, lease, and heartbeat ownership are all absent;
|
|
68
|
+
- no durable progress snapshot or result exists.
|
|
69
|
+
|
|
70
|
+
Preview the exact row first:
|
|
71
|
+
|
|
72
|
+
```sh
|
|
73
|
+
bunx --package @rizom/ops@<exact-version> brains-ops \
|
|
74
|
+
recover:retire-legacy-projection-job \
|
|
75
|
+
/data/brain-jobs.db <job-id> --type <legacy-type> --dry-run
|
|
76
|
+
```
|
|
77
|
+
|
|
78
|
+
After separate operator review, replace `--dry-run` with the exact confirmation `--confirm retire:<job-id>`. The command atomically fences against ownership or progress appearing between inspection and retirement, and fails closed if the row changed. It is not a general job cancellation API and cannot retire arbitrary job types. Require operational health and a fully idle queue afterward, then run the unchanged canonical predeploy snapshot and Deploy workflow.
|
|
79
|
+
|
|
80
|
+
## Canonical contract crossover maintenance window
|
|
81
|
+
|
|
82
|
+
Do not run this procedure without explicit operator approval. The canonical desired state, canonical `@rizom/ops`, and unified runtime image form one contract and must move or roll back together. Complete `docs/canonical-crossover-record.md` as the approval evidence without adding secret values.
|
|
83
|
+
|
|
84
|
+
Before the window, record and review:
|
|
85
|
+
|
|
86
|
+
- the prior pilot commit and exact `@rizom/ops` version;
|
|
87
|
+
- every prior runtime image tag and immutable digest;
|
|
88
|
+
- the reviewed canonical pilot commit;
|
|
89
|
+
- the exact unified `@rizom/brain` and `@rizom/ops` versions;
|
|
90
|
+
- every unified image tag and immutable digest;
|
|
91
|
+
- canary-first rollout order, followed by the remaining cohorts;
|
|
92
|
+
- expected `/health/operate` version, unauthenticated MCP response, site marker, and content repository identity for each posture.
|
|
93
|
+
|
|
94
|
+
Run `bunx brains-ops reconcile-all <canonical-review-copy> --dry-run` against the isolated review copy. The command blocks external content-repository access, leaves the review copy untouched, lists both passes' changed files, and must report second-pass zero drift. Reconciliation owns generated per-user config; only `render` owns the observational `views/users.md` projection.
|
|
95
|
+
|
|
96
|
+
During the approved window:
|
|
97
|
+
|
|
98
|
+
1. Freeze unrelated merges and releases. Wait for active Build, Reconcile, and Deploy runs to finish, then disable all three pilot workflows with `gh workflow disable build.yml`, `gh workflow disable reconcile.yml`, and `gh workflow disable deploy.yml`.
|
|
99
|
+
2. Publish and verify the reviewed unified runtime and matching ops artifacts. Do not update pilot desired state until the exact versions are installable and the expected images can be built.
|
|
100
|
+
3. Apply the reviewed canonical pilot revision while automation remains disabled. Confirm repository names, server/domain identity, content repositories, secret selectors, image names, and tag identity against the review diff.
|
|
101
|
+
4. Enable only Build, run it for the canonical desired-state revision, and record every resulting image digest. Stop if the observed digest set differs from the cutover record.
|
|
102
|
+
5. Enable only Reconcile, run it once, and review its generated per-user config commit. It must not rewrite `views/users.md`, and no generated file may combine canonical config with a retired image version.
|
|
103
|
+
6. Enable Deploy and deploy one handle at a time in the approved order. After each deploy, run `bunx brains-ops verify-user . <handle>`, render observed fleet status, and complete the manual identity, content-sync, and app-managed site checks.
|
|
104
|
+
7. Run Reconcile a second time. Require no reconciler-owned generated diff and no deploy work before re-enabling normal automation and lifting the merge/release freeze. Observed status rendering remains separate from this convergence gate.
|
|
105
|
+
|
|
106
|
+
If any gate fails, disable all three workflows again. Restore the prior pilot desired-state and dependency revision, reconcile with the prior ops version, and redeploy the prior image tag/digest as one rollback pair. Verify the prior `/health/operate` version and identity/content/site checks before re-enabling automation. Never restore only config or only an image.
|
|
107
|
+
|
|
108
|
+
## Directory-sync stress gate
|
|
109
|
+
|
|
110
|
+
Use the manual `Directory Sync Stress` workflow only against a disposable smoke user. It refuses a target unless the handle, domain, and content repository all identify smoke, the confirmation input exactly matches `stress:<handle>`, and the user desired state declares the hermetic posture below. Reconcile and deploy this posture before running the workload:
|
|
111
|
+
|
|
112
|
+
```yaml
|
|
113
|
+
embeddingEnabled: false
|
|
114
|
+
topicExtractionEnabled: false
|
|
115
|
+
skillDerivationEnabled: false
|
|
116
|
+
swotDerivationEnabled: false
|
|
117
|
+
```
|
|
118
|
+
|
|
119
|
+
Before authorizing a workload, dispatch the workflow once with `verify_only: true`. That mode loads the same Bitwarden/Varlock content credential, clones the smoke content repository, and runs `git push --dry-run` against a temporary stress ref. It creates no ref, performs no content write, does not contact the deployed runtime, and skips cleanup because no probes were created.
|
|
120
|
+
|
|
121
|
+
Profiles are deterministic and reversible:
|
|
122
|
+
|
|
123
|
+
- `regression`: 20 probes;
|
|
124
|
+
- `load`: ramps to 350 probes, updates all, renames 100, updates again, then deletes all;
|
|
125
|
+
- `stress`: ramps to 700 probes and renames 200 before cleanup.
|
|
126
|
+
|
|
127
|
+
The workflow loads operator credentials through Bitwarden/Varlock, but it is separate from Deploy and cannot deploy an image. It creates a rollback branch before the first content write, gates on health timeouts, watchdog restarts, and external AI usage during the monitored workload window, preserves warmup and cleanup samples as evidence, uploads JSON/Markdown/runtime artifacts, and runs an independent idempotent cleanup job with `if: always()`. Once cleanup confirms that no probes remain, it also prunes retained `ops/directory-sync-stress-backup-*` branches; if probes remain, the branches stay available for recovery.
|
|
128
|
+
|
|
129
|
+
Treat any gated health failure, restart, OOM, residual probe, or entity-baseline drift as a failed gate. Do not restart the target during measurement. Recovery is a separate operator action after evidence collection.
|
|
130
|
+
|
|
131
|
+
## Health watchdog smoke gate
|
|
132
|
+
|
|
133
|
+
Use the manual `Health Watchdog Smoke` workflow only after the `smoke` fleet user has completed a normal Deploy run. The workflow shares the `deploy-<handle>` concurrency group, resolves the server from pilot desired state through Hetzner, loads the existing deploy SSH key through Bitwarden/Varlock, and refuses targets whose handle and domain do not identify smoke. Confirm with `watchdog-smoke:<handle>`.
|
|
134
|
+
|
|
135
|
+
The smoke does not deploy or replace the installed systemd units. It verifies that the installed watchdog exactly matches the packaged canonical payload, then exercises that installed script. It also verifies the deployed rover image label, exact-label selector, active timer, and `/health/live` Docker healthcheck. While holding the timer's global lock, it creates temporary containers for eligible unhealthy, unrelated service-labelled unhealthy, false-labelled unhealthy, and eligible healthy cases. It then requires exactly three eligible restarts followed by budget suppression, diagnostics and state only for the eligible fixture, and no restart of the deployed rover or ineligible fixtures.
|
|
136
|
+
|
|
137
|
+
The primary job uploads remote incident, state, and command evidence. An independent `if: always()` cleanup job removes deterministic fixture names and remote temporary files using the same workflow run ID. Treat missing evidence, unexpected eligibility, deployed-container movement, cleanup failure, or any non-smoke target rejection as a failed gate.
|
|
138
|
+
|
|
139
|
+
## Stale deploy lock recovery
|
|
140
|
+
|
|
141
|
+
Kamal intentionally leaves its remote deploy lock in place when a deployment is cancelled or interrupted. Confirm that no deployment for the user is still active before releasing the lock, then use the deploy workflow's explicit recovery input:
|
|
142
|
+
|
|
143
|
+
```sh
|
|
144
|
+
gh workflow run Deploy --ref main \
|
|
145
|
+
-f handle=<handle> \
|
|
146
|
+
-f release_stale_lock=true
|
|
147
|
+
```
|
|
148
|
+
|
|
149
|
+
Recovery is opt-in and scoped to one handle. Normal push, reconcile, and manual deploy runs never remove a lock automatically.
|
|
150
|
+
|
|
36
151
|
## Bootstrap flow
|
|
37
152
|
|
|
38
153
|
For this fleet, operator-local secret material remains the source of truth during onboarding and rotation. The repo stores encrypted per-user secrets, not raw values.
|
|
@@ -53,51 +168,214 @@ Preview hosts use the shape `<handle>-preview.rizom.ai`, so one wildcard origin
|
|
|
53
168
|
|
|
54
169
|
## Upgrading operator behavior
|
|
55
170
|
|
|
56
|
-
|
|
171
|
+
The pilot repository pins `@rizom/ops` in `package.json`. The scheduled and manually dispatched Upgrade workflow owns routine upgrades to that pin. It refreshes the scaffold on a branch and opens a reviewable PR; it does not change runtime desired state or authorize a deployment.
|
|
172
|
+
|
|
173
|
+
Because scaffold refreshes can update `.github/workflows/*`, the workflow must not push with its Actions `GITHUB_TOKEN`. Configure a dedicated GitHub App:
|
|
174
|
+
|
|
175
|
+
1. Install it only on this pilot repository.
|
|
176
|
+
2. Grant repository permissions `Contents: Read and write`, `Pull requests: Read and write`, and `Workflows: Read and write`; grant nothing else.
|
|
177
|
+
3. Store its App ID as the repository Actions variable `OPS_UPGRADE_APP_ID`.
|
|
178
|
+
4. Store its private key as the repository Actions secret `OPS_UPGRADE_APP_PRIVATE_KEY`.
|
|
179
|
+
|
|
180
|
+
The workflow explicitly requests only those three permissions. Checkout persists no credential. After the freshly published `@rizom/ops` finishes, the workflow checks whether it produced a change; only then does it mint a short-lived, repository-scoped App token for the push-and-open-PR steps. A credential that can rewrite `.github/workflows/` therefore does not exist while upgraded package code runs.
|
|
181
|
+
|
|
182
|
+
If `OPS_UPGRADE_APP_ID` is unset, or token creation, branch push, or PR creation fails, the run stops with a non-zero status. Repair the CI credential path; never fall back to an operator's personal SSH key or token.
|
|
183
|
+
|
|
184
|
+
Adopting this credential flow in an existing pilot repository requires one explicitly reviewed bootstrap PR because the old Upgrade workflow cannot update itself. After that merge, routine upgrades run entirely in CI:
|
|
185
|
+
|
|
186
|
+
1. dispatch Upgrade with an exact version, or let its schedule select `latest`;
|
|
187
|
+
2. review the generated package, lockfile, deploy-script, and workflow diff;
|
|
188
|
+
3. merge the upgrade PR only after its checks pass;
|
|
189
|
+
4. change runtime desired state separately through the approved canary or fleet rollout flow.
|
|
57
190
|
|
|
58
|
-
|
|
59
|
-
2. rerun the relevant scaffold/reconcile flow
|
|
60
|
-
3. review the resulting changes to `.env.schema`, `deploy/scripts/`, and workflows in git
|
|
61
|
-
4. commit the updated deploy artifacts together
|
|
191
|
+
## Canonical verification notes
|
|
62
192
|
|
|
63
|
-
|
|
193
|
+
Use the verification script after deploy:
|
|
64
194
|
|
|
65
|
-
|
|
195
|
+
```sh
|
|
196
|
+
bunx brains-ops verify-user . <handle>
|
|
197
|
+
```
|
|
66
198
|
|
|
67
|
-
|
|
199
|
+
For every bundle posture it checks:
|
|
68
200
|
|
|
69
|
-
- `https://<handle>.rizom.ai/health`
|
|
70
|
-
- unauthenticated `POST https://<handle>.rizom.ai/mcp`
|
|
71
|
-
-
|
|
201
|
+
- `https://<handle>.rizom.ai/health/operate` returns `200`;
|
|
202
|
+
- unauthenticated `POST https://<handle>.rizom.ai/mcp` returns the expected auth failure;
|
|
203
|
+
- background jobs are not repeatedly failing, except for missing optional integrations.
|
|
72
204
|
|
|
73
|
-
|
|
205
|
+
A `core`-only instance is MCP-only; a bare `GET /` may return `401` without indicating a bad deploy. When `site` is selected, verification also checks the browser and Studio/login surfaces.
|
|
206
|
+
|
|
207
|
+
Manual checks that remain:
|
|
208
|
+
|
|
209
|
+
- initial app-managed site output is correct for the expected content/theme;
|
|
210
|
+
- content repository identity and runtime sync are healthy;
|
|
211
|
+
- passkey setup/handoff is completed from the setup email.
|
|
212
|
+
|
|
213
|
+
## One-user canonical site canary
|
|
214
|
+
|
|
215
|
+
Run this before adding custom site/theme packages or rolling a larger browser/Studio-first cohort.
|
|
216
|
+
|
|
217
|
+
1. Create or choose a canary cohort with explicit bundles:
|
|
218
|
+
|
|
219
|
+
```yaml
|
|
220
|
+
bundles:
|
|
221
|
+
- core
|
|
222
|
+
- site
|
|
223
|
+
- publishing
|
|
224
|
+
```
|
|
225
|
+
|
|
226
|
+
2. Add exactly one canary user to that cohort.
|
|
227
|
+
3. For browser/Studio-first onboarding, configure setup email in `users/<handle>.yaml`:
|
|
228
|
+
|
|
229
|
+
```yaml
|
|
230
|
+
setup:
|
|
231
|
+
delivery: email
|
|
232
|
+
email: user@example.com
|
|
233
|
+
```
|
|
234
|
+
|
|
235
|
+
4. Encrypt the user's secrets and commit only the `.age` file.
|
|
236
|
+
5. Run `bunx brains-ops onboard . <handle>`.
|
|
237
|
+
6. Run `bunx brains-ops verify-user . <handle>` with no custom site/theme overrides.
|
|
238
|
+
7. Ask the user to complete passkey setup from the setup email.
|
|
239
|
+
8. Continue to visual customization only after the canary is healthy.
|
|
240
|
+
|
|
241
|
+
Rollback must restore the prior desired-state revision and prior runtime image together. Never pair canonical config with the retired image, or retired config with the canonical image.
|
|
242
|
+
|
|
243
|
+
## Hosted site and theme package contract
|
|
244
|
+
|
|
245
|
+
Start with the public [site mockup migration guide](https://github.com/rizom-ai/brains/blob/main/docs/site-mockup-migration.md), then apply these hosted-fleet requirements:
|
|
246
|
+
|
|
247
|
+
- A site package must default-export `defineSite(...)` and import its authoring API only from `@rizom/site`.
|
|
248
|
+
- A theme package must default-export its CSS as a string. Hosted custom themes currently use the `@rizom/*` scope so the fleet image installs them with the site package; `@brains/*` themes are bundled with `@rizom/brain`.
|
|
249
|
+
- Site and custom theme packages must be public npm packages that install without registry credentials.
|
|
250
|
+
- Site, theme, and brain packages publish independently. Hosted configuration requires exact site and external-theme version pins and never derives one package version from another.
|
|
251
|
+
- Keep site structure and theme CSS in separate packages. Do not put private content or secrets in either package.
|
|
252
|
+
|
|
253
|
+
Configure a user in `users/<handle>.yaml`:
|
|
254
|
+
|
|
255
|
+
```yaml
|
|
256
|
+
siteOverride:
|
|
257
|
+
package: "@rizom/site-example"
|
|
258
|
+
version: <exact-site-version>
|
|
259
|
+
theme: "@rizom/theme-example"
|
|
260
|
+
themeVersion: <exact-theme-version>
|
|
261
|
+
```
|
|
262
|
+
|
|
263
|
+
Missing external package versions fail desired-state validation. A site override
|
|
264
|
+
produces an isolated per-instance image; it never changes the fleet's shared default
|
|
265
|
+
image. Bundled `@brains/*` themes omit `themeVersion` because they are not installed as
|
|
266
|
+
separate packages.
|
|
267
|
+
|
|
268
|
+
### Custom-package canary and rollback
|
|
269
|
+
|
|
270
|
+
1. Confirm the exact site/theme versions are public-installable without npm credentials.
|
|
271
|
+
2. Apply the exact package names and versions to one healthy canonical site canary.
|
|
272
|
+
3. Reconcile the canary, push the generated output, and let build/deploy create its site image.
|
|
273
|
+
4. Run `bunx brains-ops verify-user . <handle>`.
|
|
274
|
+
5. Manually verify the site, theme, Studio, content sync, and passkey sign-in before adding more users.
|
|
275
|
+
|
|
276
|
+
To roll back, remove or change `siteOverride`, reconcile, and redeploy that user.
|
|
277
|
+
The default image and other users remain untouched.
|
|
278
|
+
|
|
279
|
+
## Setup email checklist
|
|
280
|
+
|
|
281
|
+
Use this for browser/Studio-first users who should receive their own first-passkey setup link by email.
|
|
282
|
+
|
|
283
|
+
1. Add setup delivery to the user file:
|
|
284
|
+
|
|
285
|
+
```yaml
|
|
286
|
+
setup:
|
|
287
|
+
delivery: email
|
|
288
|
+
email: user@example.com
|
|
289
|
+
```
|
|
290
|
+
|
|
291
|
+
2. Configure these GitHub Secrets before deploy:
|
|
292
|
+
- `SETUP_EMAIL_API_KEY`
|
|
293
|
+
- `SETUP_EMAIL_FROM`
|
|
294
|
+
|
|
295
|
+
3. Reconcile/deploy the user or cohort:
|
|
296
|
+
- `bunx brains-ops onboard . <handle>`
|
|
297
|
+
- or `bunx brains-ops reconcile-cohort . <cohort>`
|
|
298
|
+
|
|
299
|
+
4. Verify the generated `users/<handle>/brain.yaml` contains `auth-service.setupEmail` and `email` interface config.
|
|
300
|
+
5. Ask the user to complete passkey setup from the email link, then use:
|
|
301
|
+
- Dashboard: `https://<handle>.rizom.ai/`
|
|
302
|
+
- Studio: `https://<handle>.rizom.ai/studio`
|
|
303
|
+
|
|
304
|
+
Notes:
|
|
305
|
+
|
|
306
|
+
- The setup URL is generated and sent by the running brain; operators should not scrape logs or SSH into the instance to retrieve it.
|
|
307
|
+
- The auth service owns setup email dedupe. It should not resend for the same persisted setup token after restart, but should retry failed delivery and resend after token rotation.
|
|
308
|
+
- `SETUP_EMAIL_FROM` is not marked required because fleets without email setup can omit it, but it is required for users with `setup.delivery: email`.
|
|
309
|
+
|
|
310
|
+
## AT Protocol smoke/config checklist
|
|
311
|
+
|
|
312
|
+
Use this when enabling AT Protocol publishing for a single pilot user.
|
|
313
|
+
|
|
314
|
+
1. Add the public PDS identifier to the user file. Prefer the account DID as
|
|
315
|
+
the identifier — it survives handle changes. Add `accountDid` too when the
|
|
316
|
+
member wants their handle verified against their subdomain
|
|
317
|
+
(`@<handle>.<domainSuffix>`): the brain then serves it at
|
|
318
|
+
`/.well-known/atproto-did` and Bluesky's "I have my own domain" HTTP
|
|
319
|
+
verification passes with no DNS records.
|
|
320
|
+
|
|
321
|
+
```yaml
|
|
322
|
+
atproto:
|
|
323
|
+
identifier: did:plc:example123
|
|
324
|
+
accountDid: did:plc:example123
|
|
325
|
+
```
|
|
326
|
+
|
|
327
|
+
Only for the PDS account designated by the protocol authority's `_lexicon`
|
|
328
|
+
DNS TXT record, also set `lexiconAuthority: true`. Every other fleet user
|
|
329
|
+
must omit it.
|
|
330
|
+
|
|
331
|
+
2. Put the app password in `users/<handle>.secrets.yaml`:
|
|
332
|
+
|
|
333
|
+
```yaml
|
|
334
|
+
atprotoAppPassword: <app-password>
|
|
335
|
+
```
|
|
336
|
+
|
|
337
|
+
3. Encrypt the per-user secret payload:
|
|
338
|
+
- `bunx brains-ops secrets:encrypt . <handle>`
|
|
339
|
+
4. Reconcile/deploy the user or cohort:
|
|
340
|
+
- `bunx brains-ops onboard . <handle>`
|
|
341
|
+
- or `bunx brains-ops reconcile-cohort . <cohort>`
|
|
342
|
+
5. Verify the generated `users/<handle>/brain.yaml` contains `plugins.atproto.identifier` (plus `accountDid` and `lexiconAuthority` when configured) and `appPassword: ${ATPROTO_APP_PASSWORD}`.
|
|
343
|
+
|
|
344
|
+
Notes:
|
|
345
|
+
|
|
346
|
+
- The ATProto identifier and authority flag are public instance config and belong in `users/<handle>.yaml`. Only the DNS-designated authority account may set `lexiconAuthority: true`.
|
|
347
|
+
- The ATProto app password is secret and belongs only in the encrypted per-user secret payload.
|
|
348
|
+
- For smoke deployments, pin only the smoke cohort/user to the released brain version that contains ATProto support.
|
|
349
|
+
|
|
350
|
+
## Discord application credential checklist
|
|
74
351
|
|
|
75
352
|
Use this when enabling Discord for a pilot user.
|
|
76
353
|
|
|
77
354
|
1. Pick the user handle (for example `smoke`).
|
|
78
355
|
2. Open the Discord Developer Portal.
|
|
79
|
-
3. Create a **new application** for that user's
|
|
356
|
+
3. Create a **new application** for that user's brain.
|
|
80
357
|
4. Add a **Bot** to the application.
|
|
81
|
-
5. Copy the bot token.
|
|
82
|
-
6. Put
|
|
358
|
+
5. Copy the bot token, application public key, and application ID.
|
|
359
|
+
6. Put those values in `.env` or `.env.local` while onboarding that user:
|
|
360
|
+
- `DISCORD_BOT_TOKEN=...`
|
|
361
|
+
- `DISCORD_PUBLIC_KEY=...`
|
|
362
|
+
- `DISCORD_APPLICATION_ID=...`
|
|
83
363
|
7. Keep `discord.enabled: true` in `users/<handle>.yaml` unless you explicitly want to disable the primary pilot interface.
|
|
84
|
-
8. Encrypt the current per-user
|
|
364
|
+
8. Encrypt the current per-user credential payload:
|
|
85
365
|
- `bunx brains-ops secrets:encrypt . <handle>`
|
|
86
366
|
9. Reconcile/deploy the user or cohort:
|
|
87
|
-
|
|
88
|
-
- `bunx brains-ops
|
|
89
|
-
|
|
90
|
-
|
|
91
|
-
11. In the Discord Developer Portal, generate an install URL and invite the bot to the right server.
|
|
92
|
-
12. Send a test message in Discord and confirm the rover responds.
|
|
367
|
+
- `bunx brains-ops onboard . <handle>`
|
|
368
|
+
- or `bunx brains-ops reconcile-cohort . <cohort>`
|
|
369
|
+
10. In the Discord Developer Portal, generate an install URL and invite the bot to the right server.
|
|
370
|
+
11. Send a test message in Discord and confirm the brain responds.
|
|
93
371
|
|
|
94
372
|
Notes:
|
|
95
373
|
|
|
96
|
-
- Use **one
|
|
97
|
-
- Do not reuse the same Discord
|
|
374
|
+
- Use **one Discord application credential set per user/brain**.
|
|
375
|
+
- Do not reuse the same Discord application across multiple pilot users.
|
|
98
376
|
- Discord is the default pilot interface moving forward.
|
|
99
377
|
- The encrypted `users/<handle>.secrets.yaml.age` file is the durable checked-in deploy input; your local env is only the operator staging source.
|
|
100
|
-
- MCP
|
|
378
|
+
- Direct MCP client access should use OAuth/passkey-capable clients where possible.
|
|
101
379
|
- When explaining the content workflow, describe it first as a normal **git repo** of **markdown/text files**.
|
|
102
380
|
- Position **Obsidian** as optional: it is just one possible editor for those same files, not the default requirement.
|
|
103
381
|
|