@rizom/ops 0.2.0-alpha.45 → 0.2.0-alpha.450
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +82 -1
- package/dist/age-key-bootstrap.d.ts +1 -1
- package/dist/brains-ops.js +736 -335
- package/dist/capability-bundle-migration.d.ts +22 -0
- package/dist/cert-bootstrap.d.ts +4 -2
- package/dist/content-repo-ref.d.ts +10 -0
- package/dist/content-repo.d.ts +1 -1
- package/dist/deploy.js +110 -166
- package/dist/directory-sync-stress-system.d.ts +75 -0
- package/dist/directory-sync-stress.d.ts +117 -0
- package/dist/entries/deploy.d.ts +4 -2
- package/dist/health-watchdog-smoke.d.ts +54 -0
- package/dist/image-inventory.d.ts +4 -0
- package/dist/image-types.d.ts +7 -0
- package/dist/images.d.ts +82 -0
- package/dist/index.d.ts +8 -0
- package/dist/index.js +736 -323
- package/dist/legacy-pilot-migration.d.ts +11 -0
- package/dist/legacy-projection-job-recovery.d.ts +29 -0
- package/dist/load-registry.d.ts +61 -6
- package/dist/observed-status.d.ts +1 -1
- package/dist/parse-args.d.ts +2 -12
- package/dist/preview-domain.d.ts +9 -0
- package/dist/reconcile-dry-run.d.ts +9 -0
- package/dist/run-command.d.ts +20 -4
- package/dist/schema.d.ts +126 -163
- package/dist/secrets-encrypt.d.ts +7 -13
- package/dist/secrets-push.d.ts +1 -1
- package/dist/ssh-key-bootstrap.d.ts +1 -26
- package/dist/stage-legacy-crossover.d.ts +25 -0
- package/dist/stress-command.d.ts +14 -0
- package/dist/stress-git-checkout.d.ts +21 -0
- package/dist/stress-health-monitor.d.ts +46 -0
- package/dist/upgrade.d.ts +10 -0
- package/dist/user-offboard.d.ts +61 -0
- package/dist/verify-user.d.ts +22 -0
- package/package.json +50 -42
- package/templates/rover-pilot/.env.schema +20 -5
- package/templates/rover-pilot/.github/actions/varlock-env/action.yml +47 -0
- package/templates/rover-pilot/.github/workflows/build.yml +70 -17
- package/templates/rover-pilot/.github/workflows/deploy.yml +74 -57
- package/templates/rover-pilot/.github/workflows/directory-sync-stress.yml +119 -0
- package/templates/rover-pilot/.github/workflows/health-watchdog-smoke.yml +95 -0
- package/templates/rover-pilot/.github/workflows/offboard.yml +107 -0
- package/templates/rover-pilot/.github/workflows/reconcile.yml +10 -4
- package/templates/rover-pilot/.github/workflows/upgrade.yml +104 -0
- package/templates/rover-pilot/README.md +13 -5
- package/templates/rover-pilot/deploy/scripts/create-predeploy-backup.ts +936 -0
- package/templates/rover-pilot/deploy/scripts/decrypt-user-secrets.ts +80 -24
- package/templates/rover-pilot/deploy/scripts/helpers.ts +2 -0
- package/templates/rover-pilot/deploy/scripts/install-health-watchdog.ts +144 -0
- package/templates/rover-pilot/deploy/scripts/provision-server.ts +51 -21
- package/templates/rover-pilot/deploy/scripts/resolve-deploy-handles.ts +4 -0
- package/templates/rover-pilot/deploy/scripts/resolve-missing-images.ts +13 -0
- package/templates/rover-pilot/deploy/scripts/resolve-user-config.ts +40 -9
- package/templates/rover-pilot/deploy/scripts/sync-content-repo.ts +51 -47
- package/templates/rover-pilot/deploy/scripts/update-dns.ts +72 -16
- package/templates/rover-pilot/deploy/scripts/validate-secrets.ts +12 -1
- package/templates/rover-pilot/deploy/scripts/verify-runtime-image.ts +18 -0
- package/templates/rover-pilot/docs/canonical-crossover-record.md +107 -0
- package/templates/rover-pilot/docs/onboarding-checklist.md +23 -14
- package/templates/rover-pilot/docs/operator-playbook.md +357 -29
- package/templates/rover-pilot/docs/user-onboarding.md +48 -463
- package/templates/rover-pilot/package.json +1 -0
- package/templates/rover-pilot/pilot.yaml +6 -4
- package/dist/origin-ca.d.ts +0 -1
- package/dist/push-secrets.d.ts +0 -9
- package/dist/push-target.d.ts +0 -2
- package/dist/run-subprocess.d.ts +0 -6
- package/templates/rover-pilot/.kamal/hooks/pre-deploy +0 -9
|
@@ -9,30 +9,160 @@ Treat these as checked-in deploy artifacts in the pilot repo:
|
|
|
9
9
|
- `deploy/scripts/`
|
|
10
10
|
- `.github/workflows/build.yml`
|
|
11
11
|
- `.github/workflows/deploy.yml`
|
|
12
|
+
- `.github/workflows/directory-sync-stress.yml`
|
|
13
|
+
- `.github/workflows/health-watchdog-smoke.yml`
|
|
14
|
+
- `.github/workflows/offboard.yml`
|
|
12
15
|
- `.github/workflows/reconcile.yml`
|
|
13
16
|
|
|
14
17
|
`.env.schema` is the single source of truth for required and sensitive deploy vars.
|
|
15
18
|
The deploy scripts and workflows should read from that contract instead of inventing a second list.
|
|
16
19
|
|
|
17
|
-
The
|
|
20
|
+
The fleet has one image topology:
|
|
18
21
|
|
|
19
|
-
-
|
|
22
|
+
- one immutable `brain-${brainVersion}` image is published for each effective Brain version
|
|
23
|
+
- every new image contains the union of exact site/theme package pins across the whole fleet, regardless of current Brain versions
|
|
24
|
+
- conflicting versions of one package fail image resolution before build
|
|
20
25
|
- generated `users/<handle>/.env` carries `BRAIN_VERSION=<brainVersion>`
|
|
21
|
-
- deploy
|
|
26
|
+
- build and deploy derive the same effective image tag from the resolved registry
|
|
22
27
|
|
|
23
28
|
## Version bump flow
|
|
24
29
|
|
|
25
30
|
When `pilot.yaml.brainVersion` changes and you push:
|
|
26
31
|
|
|
27
|
-
1. build publishes
|
|
32
|
+
1. build publishes each missing version image with the full fleet package union, and verifies existing images against their assigned instances
|
|
28
33
|
2. reconcile refreshes generated `users/<handle>/.env`
|
|
29
34
|
3. deploy runs for handles whose generated config changed
|
|
30
35
|
4. generated file commits happen once in a final aggregation step after the deploy matrix finishes
|
|
31
36
|
|
|
37
|
+
Every external site and theme package has its own exact version pin. A cohort or
|
|
38
|
+
pilot brain-version bump never changes those package versions implicitly; update each
|
|
39
|
+
pin deliberately from reviewed package and image evidence.
|
|
40
|
+
|
|
41
|
+
Smoke-first upgrades build the full fleet package union before promotion. Moving
|
|
42
|
+
other cohorts onto the tested Brain version then reuses that same immutable image.
|
|
43
|
+
Explicit Build dispatches also include the fleet union; `site_packages` only adds
|
|
44
|
+
extra pins and cannot replace or conflict with declared pins.
|
|
45
|
+
|
|
46
|
+
Deploy verifies the actual installed Brain and selected instance's site/theme
|
|
47
|
+
versions on the CI runner after image readiness and before provisioning or
|
|
48
|
+
container replacement. It reads manifests in a read-only, network-disabled
|
|
49
|
+
container without starting the Brain or mounting fleet data. Missing packages,
|
|
50
|
+
wrong versions, or verification failures block deployment, even on manual runs.
|
|
51
|
+
Never bypass this gate or overwrite a deployed tag to recover an older incomplete
|
|
52
|
+
image; review artifact recovery separately.
|
|
53
|
+
|
|
32
54
|
When a push changes only deploy contract files and no generated `users/<handle>/.env` or `users/<handle>/brain.yaml` files, the deploy workflow exits through its explicit no-op path and prints `No affected user configs; skipping deploy.`
|
|
33
55
|
|
|
34
56
|
They are scaffolded from `@rizom/ops`, then versioned in this repo like any other deploy contract.
|
|
35
57
|
|
|
58
|
+
## Verified pre-deploy rollback snapshots
|
|
59
|
+
|
|
60
|
+
Deploy creates and verifies one target-scoped rollback snapshot after SSH and image readiness, immediately before optional stale-lock release and `kamal setup`. A snapshot failure stops the workflow before container replacement. The current runtime must be ready with an idle job queue; degradation a plugin reports is named in the log and backed up, not refused, because a deploy is often its fix. A genuinely new server with no runtime or persistent state reports `not applicable`; persistent state without an identifiable runtime fails closed.
|
|
61
|
+
|
|
62
|
+
Each snapshot contains transaction-consistent online captures of the six canonical SQLite databases, database hashes and `quick_check` results, the deployed `brain.yaml`, sanitized container and mount metadata, and the Directory Sync checkout. Git capture includes all refs, an observed remote head, staged and unstaged binary patches, and mode-preserving archives of untracked and ignored files. Capture performs no checkout mutation or remote write.
|
|
63
|
+
|
|
64
|
+
Verified snapshots are finalized atomically under:
|
|
65
|
+
|
|
66
|
+
```text
|
|
67
|
+
/opt/brain-state/backups/predeploy-<handle>-<target>-<UTC timestamp>/
|
|
68
|
+
```
|
|
69
|
+
|
|
70
|
+
Files use mode `0600` and the directory uses `0700`. A `.incomplete` directory is diagnostic evidence, never a rollback point. Checksums and the versioned manifest must verify before finalization. The default retention is five verified snapshots per target; incomplete or unverified directories do not count, and the snapshot created by the current deployment is never pruned.
|
|
71
|
+
|
|
72
|
+
These are same-server rollback snapshots, not off-host disaster recovery. There is no normal skip and no automatic restore. Restoration requires separate approval: stop replacement/application processes, recheck the selected snapshot's checksums, preserve the current state, restore databases and exact Git state, then validate health, queues, Git checkpoints, durable export intents, preview, and production output before reopening the target.
|
|
73
|
+
|
|
74
|
+
## Retired projection job recovery
|
|
75
|
+
|
|
76
|
+
A pre-scheduler projection job can remain active after its handler type is retired. Current runtime health reports active job types missing from the finalized execution inventory as degraded. Do not restart repeatedly or relax the snapshot idle gate.
|
|
77
|
+
|
|
78
|
+
Use the recovery command only after read-only evidence proves all of the following for one exact row:
|
|
79
|
+
|
|
80
|
+
- its type is one of the command's fixed legacy projection types;
|
|
81
|
+
- the installed runtime no longer contains that handler;
|
|
82
|
+
- attempt, worker-session, lease, and heartbeat ownership are all absent;
|
|
83
|
+
- no durable progress snapshot or result exists.
|
|
84
|
+
|
|
85
|
+
Preview the exact row first:
|
|
86
|
+
|
|
87
|
+
```sh
|
|
88
|
+
bunx --package @rizom/ops@<exact-version> brains-ops \
|
|
89
|
+
recover:retire-legacy-projection-job \
|
|
90
|
+
/data/brain-jobs.db <job-id> --type <legacy-type> --dry-run
|
|
91
|
+
```
|
|
92
|
+
|
|
93
|
+
After separate operator review, replace `--dry-run` with the exact confirmation `--confirm retire:<job-id>`. The command atomically fences against ownership or progress appearing between inspection and retirement, and fails closed if the row changed. It is not a general job cancellation API and cannot retire arbitrary job types. Require operational health and a fully idle queue afterward, then run the unchanged canonical predeploy snapshot and Deploy workflow.
|
|
94
|
+
|
|
95
|
+
## Canonical contract crossover maintenance window
|
|
96
|
+
|
|
97
|
+
Do not run this procedure without explicit operator approval. The canonical desired state, canonical `@rizom/ops`, and unified runtime image form one contract and must move or roll back together. Complete `docs/canonical-crossover-record.md` as the approval evidence without adding secret values.
|
|
98
|
+
|
|
99
|
+
Before the window, record and review:
|
|
100
|
+
|
|
101
|
+
- the prior pilot commit and exact `@rizom/ops` version;
|
|
102
|
+
- every prior runtime image tag and immutable digest;
|
|
103
|
+
- the reviewed canonical pilot commit;
|
|
104
|
+
- the exact unified `@rizom/brain` and `@rizom/ops` versions;
|
|
105
|
+
- every unified image tag and immutable digest;
|
|
106
|
+
- canary-first rollout order, followed by the remaining cohorts;
|
|
107
|
+
- expected `/health/operate` version, unauthenticated MCP response, site marker, and content repository identity for each posture.
|
|
108
|
+
|
|
109
|
+
Run `bunx brains-ops reconcile-all <canonical-review-copy> --dry-run` against the isolated review copy. The command blocks external content-repository access, leaves the review copy untouched, lists both passes' changed files, and must report second-pass zero drift. Reconciliation owns generated per-user config; only `render` owns the observational `views/users.md` projection.
|
|
110
|
+
|
|
111
|
+
During the approved window:
|
|
112
|
+
|
|
113
|
+
1. Freeze unrelated merges and releases. Wait for active Build, Reconcile, and Deploy runs to finish, then disable all three pilot workflows with `gh workflow disable build.yml`, `gh workflow disable reconcile.yml`, and `gh workflow disable deploy.yml`.
|
|
114
|
+
2. Publish and verify the reviewed unified runtime and matching ops artifacts. Do not update pilot desired state until the exact versions are installable and the expected images can be built.
|
|
115
|
+
3. Apply the reviewed canonical pilot revision while automation remains disabled. Confirm repository names, server/domain identity, content repositories, secret selectors, image names, and tag identity against the review diff.
|
|
116
|
+
4. Enable only Build, run it for the canonical desired-state revision, and record every resulting image digest. Stop if the observed digest set differs from the cutover record.
|
|
117
|
+
5. Enable only Reconcile, run it once, and review its generated per-user config commit. It must not rewrite `views/users.md`, and no generated file may combine canonical config with a retired image version.
|
|
118
|
+
6. Enable Deploy and deploy one handle at a time in the approved order. After each deploy, run `bunx brains-ops verify-user . <handle>`, render observed fleet status, and complete the manual identity, content-sync, and app-managed site checks.
|
|
119
|
+
7. Run Reconcile a second time. Require no reconciler-owned generated diff and no deploy work before re-enabling normal automation and lifting the merge/release freeze. Observed status rendering remains separate from this convergence gate.
|
|
120
|
+
|
|
121
|
+
If any gate fails, disable all three workflows again. Restore the prior pilot desired-state and dependency revision, reconcile with the prior ops version, and redeploy the prior image tag/digest as one rollback pair. Verify the prior `/health/operate` version and identity/content/site checks before re-enabling automation. Never restore only config or only an image.
|
|
122
|
+
|
|
123
|
+
## Directory-sync stress gate
|
|
124
|
+
|
|
125
|
+
Use the manual `Directory Sync Stress` workflow only against a disposable smoke user. It refuses a target unless the handle, domain, and content repository all identify smoke, the confirmation input exactly matches `stress:<handle>`, and the user desired state declares the hermetic posture below. Reconcile and deploy this posture before running the workload:
|
|
126
|
+
|
|
127
|
+
```yaml
|
|
128
|
+
embeddingEnabled: false
|
|
129
|
+
topicExtractionEnabled: false
|
|
130
|
+
skillDerivationEnabled: false
|
|
131
|
+
swotDerivationEnabled: false
|
|
132
|
+
```
|
|
133
|
+
|
|
134
|
+
Before authorizing a workload, dispatch the workflow once with `verify_only: true`. That mode loads the same Bitwarden/Varlock content credential, clones the smoke content repository, and runs `git push --dry-run` against a temporary stress ref. It creates no ref, performs no content write, does not contact the deployed runtime, and skips cleanup because no probes were created.
|
|
135
|
+
|
|
136
|
+
Profiles are deterministic and reversible:
|
|
137
|
+
|
|
138
|
+
- `regression`: 20 probes;
|
|
139
|
+
- `load`: ramps to 350 probes, updates all, renames 100, updates again, then deletes all;
|
|
140
|
+
- `stress`: ramps to 700 probes and renames 200 before cleanup.
|
|
141
|
+
|
|
142
|
+
The workflow loads operator credentials through Bitwarden/Varlock, but it is separate from Deploy and cannot deploy an image. It creates a rollback branch before the first content write, gates on health timeouts, watchdog restarts, and external AI usage during the monitored workload window, preserves warmup and cleanup samples as evidence, uploads JSON/Markdown/runtime artifacts, and runs an independent idempotent cleanup job with `if: always()`. Once cleanup confirms that no probes remain, it also prunes retained `ops/directory-sync-stress-backup-*` branches; if probes remain, the branches stay available for recovery.
|
|
143
|
+
|
|
144
|
+
Treat any gated health failure, restart, OOM, residual probe, or entity-baseline drift as a failed gate. Do not restart the target during measurement. Recovery is a separate operator action after evidence collection.
|
|
145
|
+
|
|
146
|
+
## Health watchdog smoke gate
|
|
147
|
+
|
|
148
|
+
Use the manual `Health Watchdog Smoke` workflow only after the `smoke` fleet user has completed a normal Deploy run. The workflow shares the `deploy-<handle>` concurrency group, resolves the server from pilot desired state through Hetzner, loads the existing deploy SSH key through Bitwarden/Varlock, and refuses targets whose handle and domain do not identify smoke. Confirm with `watchdog-smoke:<handle>`.
|
|
149
|
+
|
|
150
|
+
The smoke does not deploy or replace the installed systemd units. It verifies that the installed watchdog exactly matches the packaged canonical payload, then exercises that installed script. It also verifies the deployed rover image label, exact-label selector, active timer, and `/health/live` Docker healthcheck. While holding the timer's global lock, it creates temporary containers for eligible unhealthy, unrelated service-labelled unhealthy, false-labelled unhealthy, and eligible healthy cases. It then requires exactly three eligible restarts followed by budget suppression, diagnostics and state only for the eligible fixture, and no restart of the deployed rover or ineligible fixtures.
|
|
151
|
+
|
|
152
|
+
The primary job uploads remote incident, state, and command evidence. An independent `if: always()` cleanup job removes deterministic fixture names and remote temporary files using the same workflow run ID. Treat missing evidence, unexpected eligibility, deployed-container movement, cleanup failure, or any non-smoke target rejection as a failed gate.
|
|
153
|
+
|
|
154
|
+
## Stale deploy lock recovery
|
|
155
|
+
|
|
156
|
+
Kamal intentionally leaves its remote deploy lock in place when a deployment is cancelled or interrupted. Confirm that no deployment for the user is still active before releasing the lock, then use the deploy workflow's explicit recovery input:
|
|
157
|
+
|
|
158
|
+
```sh
|
|
159
|
+
gh workflow run Deploy --ref main \
|
|
160
|
+
-f handle=<handle> \
|
|
161
|
+
-f release_stale_lock=true
|
|
162
|
+
```
|
|
163
|
+
|
|
164
|
+
Recovery is opt-in and scoped to one handle. Normal push, reconcile, and manual deploy runs never remove a lock automatically.
|
|
165
|
+
|
|
36
166
|
## Bootstrap flow
|
|
37
167
|
|
|
38
168
|
For this fleet, operator-local secret material remains the source of truth during onboarding and rotation. The repo stores encrypted per-user secrets, not raw values.
|
|
@@ -51,53 +181,251 @@ The shared cert bootstrap writes local cert artifacts under `.brains-ops/certs/s
|
|
|
51
181
|
|
|
52
182
|
Preview hosts use the shape `<handle>-preview.rizom.ai`, so one wildcard origin cert for `*.rizom.ai` covers both the primary and preview hosts for every pilot user.
|
|
53
183
|
|
|
184
|
+
## Pilot user offboarding
|
|
185
|
+
|
|
186
|
+
Use the manual **Offboard** workflow for explicit pilot retirement. It is the only
|
|
187
|
+
supported path that removes both checked-in desired state and provider resources.
|
|
188
|
+
Normal reconcile/deploy changes never imply destruction.
|
|
189
|
+
|
|
190
|
+
1. Enter a comma-separated handle list. The command sorts and deduplicates it.
|
|
191
|
+
2. Run with `apply: false` and review the exact servers, DNS records, and content
|
|
192
|
+
repositories in the plan.
|
|
193
|
+
3. Enter the canonical confirmation printed by the dry run, for example
|
|
194
|
+
`sunset:alice,bob`.
|
|
195
|
+
4. Rerun with `apply: true`.
|
|
196
|
+
5. Verify the workflow's bot commit removes user YAML, encrypted secrets, generated
|
|
197
|
+
user directories, cohort membership, and rows from `views/users.md`.
|
|
198
|
+
|
|
199
|
+
The apply path archives each private content repository, deletes the user's main and
|
|
200
|
+
preview DNS records, destroys the dedicated Hetzner server, and removes desired state.
|
|
201
|
+
Custom-domain users also lose their managed `www` record. The operation is idempotent,
|
|
202
|
+
but it deliberately creates **no runtime backup**; obtain separate owner approval and
|
|
203
|
+
backup state before apply when retention is required.
|
|
204
|
+
|
|
205
|
+
The equivalent operator-local commands are:
|
|
206
|
+
|
|
207
|
+
```sh
|
|
208
|
+
bunx brains-ops user:offboard . alice bob
|
|
209
|
+
bunx brains-ops user:offboard . alice bob \
|
|
210
|
+
--apply --confirm sunset:alice,bob
|
|
211
|
+
```
|
|
212
|
+
|
|
213
|
+
The deploy handle resolver ignores deleted users, so the generated-file deletions in
|
|
214
|
+
the offboarding commit cannot route those handles back through onboarding/deploy.
|
|
215
|
+
|
|
54
216
|
## Upgrading operator behavior
|
|
55
217
|
|
|
56
|
-
|
|
218
|
+
The pilot repository pins `@rizom/ops` in `package.json`. The scheduled and manually dispatched Upgrade workflow owns routine upgrades to that pin. It refreshes the scaffold on a branch and opens a reviewable PR; it does not change runtime desired state or authorize a deployment.
|
|
219
|
+
|
|
220
|
+
Because scaffold refreshes can update `.github/workflows/*`, the workflow must not push with its Actions `GITHUB_TOKEN`. Configure a dedicated GitHub App:
|
|
57
221
|
|
|
58
|
-
1.
|
|
59
|
-
2.
|
|
60
|
-
3.
|
|
61
|
-
4.
|
|
222
|
+
1. Install it only on this pilot repository.
|
|
223
|
+
2. Grant repository permissions `Contents: Read and write`, `Pull requests: Read and write`, and `Workflows: Read and write`; grant nothing else.
|
|
224
|
+
3. Store its App ID as the repository Actions variable `OPS_UPGRADE_APP_ID`.
|
|
225
|
+
4. Store its private key as the repository Actions secret `OPS_UPGRADE_APP_PRIVATE_KEY`.
|
|
62
226
|
|
|
63
|
-
|
|
227
|
+
The workflow explicitly requests only those three permissions. Checkout persists no credential. After the freshly published `@rizom/ops` finishes, the workflow checks whether it produced a change; only then does it mint a short-lived, repository-scoped App token for the push-and-open-PR steps. A credential that can rewrite `.github/workflows/` therefore does not exist while upgraded package code runs.
|
|
64
228
|
|
|
65
|
-
|
|
229
|
+
If `OPS_UPGRADE_APP_ID` is unset, or token creation, branch push, or PR creation fails, the run stops with a non-zero status. Repair the CI credential path; never fall back to an operator's personal SSH key or token.
|
|
66
230
|
|
|
67
|
-
|
|
231
|
+
Adopting this credential flow in an existing pilot repository requires one explicitly reviewed bootstrap PR because the old Upgrade workflow cannot update itself. After that merge, routine upgrades run entirely in CI:
|
|
68
232
|
|
|
69
|
-
|
|
70
|
-
|
|
71
|
-
|
|
233
|
+
1. dispatch Upgrade with an exact version, or let its schedule select `latest`;
|
|
234
|
+
2. review the generated package, lockfile, deploy-script, and workflow diff;
|
|
235
|
+
3. merge the upgrade PR only after its checks pass;
|
|
236
|
+
4. change runtime desired state separately through the approved canary or fleet rollout flow.
|
|
72
237
|
|
|
73
|
-
##
|
|
238
|
+
## Canonical verification notes
|
|
239
|
+
|
|
240
|
+
Use the verification script after deploy:
|
|
241
|
+
|
|
242
|
+
```sh
|
|
243
|
+
bunx brains-ops verify-user . <handle>
|
|
244
|
+
```
|
|
245
|
+
|
|
246
|
+
For every bundle posture it checks:
|
|
247
|
+
|
|
248
|
+
- `https://<handle>.rizom.ai/health/operate` returns `200`;
|
|
249
|
+
- unauthenticated `POST https://<handle>.rizom.ai/mcp` returns the expected auth failure;
|
|
250
|
+
- background jobs are not repeatedly failing, except for missing optional integrations.
|
|
251
|
+
|
|
252
|
+
A `core`-only instance is MCP-only; a bare `GET /` may return `401` without indicating a bad deploy. When `site` is selected, verification also checks the browser and Studio/login surfaces.
|
|
253
|
+
|
|
254
|
+
Manual checks that remain:
|
|
255
|
+
|
|
256
|
+
- initial app-managed site output is correct for the expected content/theme;
|
|
257
|
+
- content repository identity and runtime sync are healthy;
|
|
258
|
+
- passkey setup/handoff is completed from the setup email.
|
|
259
|
+
|
|
260
|
+
## One-user canonical site canary
|
|
261
|
+
|
|
262
|
+
Run this before adding custom site/theme packages or rolling a larger browser/Studio-first cohort.
|
|
263
|
+
|
|
264
|
+
1. Create or choose a canary cohort with explicit bundles:
|
|
265
|
+
|
|
266
|
+
```yaml
|
|
267
|
+
bundles:
|
|
268
|
+
- core
|
|
269
|
+
- site
|
|
270
|
+
- publishing
|
|
271
|
+
```
|
|
272
|
+
|
|
273
|
+
2. Add exactly one canary user to that cohort.
|
|
274
|
+
3. For browser/Studio-first onboarding, configure setup email in `users/<handle>.yaml`:
|
|
275
|
+
|
|
276
|
+
```yaml
|
|
277
|
+
setup:
|
|
278
|
+
delivery: email
|
|
279
|
+
email: user@example.com
|
|
280
|
+
```
|
|
281
|
+
|
|
282
|
+
4. Encrypt the user's secrets and commit only the `.age` file.
|
|
283
|
+
5. Run `bunx brains-ops onboard . <handle>`.
|
|
284
|
+
6. Run `bunx brains-ops verify-user . <handle>` with no custom site/theme overrides.
|
|
285
|
+
7. Ask the user to complete passkey setup from the setup email.
|
|
286
|
+
8. Continue to visual customization only after the canary is healthy.
|
|
287
|
+
|
|
288
|
+
Rollback must restore the prior desired-state revision and prior runtime image together. Never pair canonical config with the retired image, or retired config with the canonical image.
|
|
289
|
+
|
|
290
|
+
## Hosted site and theme package contract
|
|
291
|
+
|
|
292
|
+
Start with the public [site mockup migration guide](https://github.com/rizom-ai/brains/blob/main/docs/site-mockup-migration.md), then apply these hosted-fleet requirements:
|
|
293
|
+
|
|
294
|
+
- A site package must default-export `defineSite(...)` and import its authoring API only from `@rizom/site`.
|
|
295
|
+
- A theme package must default-export its CSS as a string. Hosted custom themes currently use the `@rizom/*` scope so the fleet image installs them with the site package; `@brains/*` themes are bundled with `@rizom/brain`.
|
|
296
|
+
- Site and custom theme packages must be public npm packages that install without registry credentials.
|
|
297
|
+
- Site, theme, and brain packages publish independently. Hosted configuration requires exact site and external-theme version pins and never derives one package version from another.
|
|
298
|
+
- Keep site structure and theme CSS in separate packages. Do not put private content or secrets in either package.
|
|
299
|
+
|
|
300
|
+
Configure a user in `users/<handle>.yaml`:
|
|
301
|
+
|
|
302
|
+
```yaml
|
|
303
|
+
siteOverride:
|
|
304
|
+
package: "@rizom/site-example"
|
|
305
|
+
version: <exact-site-version>
|
|
306
|
+
theme: "@rizom/theme-example"
|
|
307
|
+
themeVersion: <exact-theme-version>
|
|
308
|
+
```
|
|
309
|
+
|
|
310
|
+
Missing external package versions fail desired-state validation. Every instance on one
|
|
311
|
+
Brain version uses the same image and exact package union. Change a site/theme package
|
|
312
|
+
pin only together with a fresh Brain version; published `brain-${brainVersion}` tags
|
|
313
|
+
remain immutable. Conflicting pins for one package on the same Brain version fail before
|
|
314
|
+
build. Bundled `@brains/*` themes omit `themeVersion` because they are not installed as
|
|
315
|
+
separate packages.
|
|
316
|
+
|
|
317
|
+
### Custom-package canary and rollback
|
|
318
|
+
|
|
319
|
+
1. Confirm the exact site/theme versions are public-installable without npm credentials.
|
|
320
|
+
2. Apply the exact package names and versions to one healthy canonical site canary and select a fresh Brain version for that package set.
|
|
321
|
+
3. Push desired state and let the normal Build → Reconcile → Deploy chain create the shared version image and update the canary.
|
|
322
|
+
4. Run `bunx brains-ops verify-user . <handle>`.
|
|
323
|
+
5. Manually verify the site, theme, Studio, content sync, and passkey sign-in before adding more users.
|
|
324
|
+
|
|
325
|
+
To roll back, remove or change `siteOverride` while selecting a fresh Brain version,
|
|
326
|
+
then reconcile and redeploy that user. Never mutate an existing image tag; coordinate any
|
|
327
|
+
other users intentionally sharing the selected version.
|
|
328
|
+
|
|
329
|
+
## Setup email checklist
|
|
330
|
+
|
|
331
|
+
Use this for browser/Studio-first users who should receive their own first-passkey setup link by email.
|
|
332
|
+
|
|
333
|
+
1. Add setup delivery to the user file:
|
|
334
|
+
|
|
335
|
+
```yaml
|
|
336
|
+
setup:
|
|
337
|
+
delivery: email
|
|
338
|
+
email: user@example.com
|
|
339
|
+
```
|
|
340
|
+
|
|
341
|
+
2. Configure these GitHub Secrets before deploy:
|
|
342
|
+
- `SETUP_EMAIL_API_KEY`
|
|
343
|
+
- `SETUP_EMAIL_FROM`
|
|
344
|
+
|
|
345
|
+
3. Reconcile/deploy the user or cohort:
|
|
346
|
+
- `bunx brains-ops onboard . <handle>`
|
|
347
|
+
- or `bunx brains-ops reconcile-cohort . <cohort>`
|
|
348
|
+
|
|
349
|
+
4. Verify the generated `users/<handle>/brain.yaml` contains `auth-service.setupEmail` and `email` interface config.
|
|
350
|
+
5. Ask the user to complete passkey setup from the email link, then use:
|
|
351
|
+
- Dashboard: `https://<handle>.rizom.ai/`
|
|
352
|
+
- Studio: `https://<handle>.rizom.ai/studio`
|
|
353
|
+
|
|
354
|
+
Notes:
|
|
355
|
+
|
|
356
|
+
- The setup URL is generated and sent by the running brain; operators should not scrape logs or SSH into the instance to retrieve it.
|
|
357
|
+
- The auth service owns setup email dedupe. It should not resend for the same persisted setup token after restart, but should retry failed delivery and resend after token rotation.
|
|
358
|
+
- `SETUP_EMAIL_FROM` is not marked required because fleets without email setup can omit it, but it is required for users with `setup.delivery: email`.
|
|
359
|
+
|
|
360
|
+
## AT Protocol smoke/config checklist
|
|
361
|
+
|
|
362
|
+
Use this when enabling AT Protocol publishing for a single pilot user.
|
|
363
|
+
|
|
364
|
+
1. Add the public PDS identifier to the user file. Prefer the account DID as
|
|
365
|
+
the identifier — it survives handle changes. Add `accountDid` too when the
|
|
366
|
+
member wants their handle verified against their subdomain
|
|
367
|
+
(`@<handle>.<domainSuffix>`): the brain then serves it at
|
|
368
|
+
`/.well-known/atproto-did` and Bluesky's "I have my own domain" HTTP
|
|
369
|
+
verification passes with no DNS records.
|
|
370
|
+
|
|
371
|
+
```yaml
|
|
372
|
+
atproto:
|
|
373
|
+
identifier: did:plc:example123
|
|
374
|
+
accountDid: did:plc:example123
|
|
375
|
+
```
|
|
376
|
+
|
|
377
|
+
Only for the PDS account designated by the protocol authority's `_lexicon`
|
|
378
|
+
DNS TXT record, also set `lexiconAuthority: true`. Every other fleet user
|
|
379
|
+
must omit it.
|
|
380
|
+
|
|
381
|
+
2. Put the app password in `users/<handle>.secrets.yaml`:
|
|
382
|
+
|
|
383
|
+
```yaml
|
|
384
|
+
atprotoAppPassword: <app-password>
|
|
385
|
+
```
|
|
386
|
+
|
|
387
|
+
3. Encrypt the per-user secret payload:
|
|
388
|
+
- `bunx brains-ops secrets:encrypt . <handle>`
|
|
389
|
+
4. Reconcile/deploy the user or cohort:
|
|
390
|
+
- `bunx brains-ops onboard . <handle>`
|
|
391
|
+
- or `bunx brains-ops reconcile-cohort . <cohort>`
|
|
392
|
+
5. Verify the generated `users/<handle>/brain.yaml` contains `plugins.atproto.identifier` (plus `accountDid` and `lexiconAuthority` when configured) and `appPassword: ${ATPROTO_APP_PASSWORD}`.
|
|
393
|
+
|
|
394
|
+
Notes:
|
|
395
|
+
|
|
396
|
+
- The ATProto identifier and authority flag are public instance config and belong in `users/<handle>.yaml`. Only the DNS-designated authority account may set `lexiconAuthority: true`.
|
|
397
|
+
- The ATProto app password is secret and belongs only in the encrypted per-user secret payload.
|
|
398
|
+
- For smoke deployments, pin only the smoke cohort/user to the released brain version that contains ATProto support.
|
|
399
|
+
|
|
400
|
+
## Discord application credential checklist
|
|
74
401
|
|
|
75
402
|
Use this when enabling Discord for a pilot user.
|
|
76
403
|
|
|
77
404
|
1. Pick the user handle (for example `smoke`).
|
|
78
405
|
2. Open the Discord Developer Portal.
|
|
79
|
-
3. Create a **new application** for that user's
|
|
406
|
+
3. Create a **new application** for that user's brain.
|
|
80
407
|
4. Add a **Bot** to the application.
|
|
81
|
-
5. Copy the bot token.
|
|
82
|
-
6. Put
|
|
408
|
+
5. Copy the bot token, application public key, and application ID.
|
|
409
|
+
6. Put those values in `.env` or `.env.local` while onboarding that user:
|
|
410
|
+
- `DISCORD_BOT_TOKEN=...`
|
|
411
|
+
- `DISCORD_PUBLIC_KEY=...`
|
|
412
|
+
- `DISCORD_APPLICATION_ID=...`
|
|
83
413
|
7. Keep `discord.enabled: true` in `users/<handle>.yaml` unless you explicitly want to disable the primary pilot interface.
|
|
84
|
-
8. Encrypt the current per-user
|
|
414
|
+
8. Encrypt the current per-user credential payload:
|
|
85
415
|
- `bunx brains-ops secrets:encrypt . <handle>`
|
|
86
416
|
9. Reconcile/deploy the user or cohort:
|
|
87
|
-
|
|
88
|
-
- `bunx brains-ops
|
|
89
|
-
|
|
90
|
-
|
|
91
|
-
11. In the Discord Developer Portal, generate an install URL and invite the bot to the right server.
|
|
92
|
-
12. Send a test message in Discord and confirm the rover responds.
|
|
417
|
+
- `bunx brains-ops onboard . <handle>`
|
|
418
|
+
- or `bunx brains-ops reconcile-cohort . <cohort>`
|
|
419
|
+
10. In the Discord Developer Portal, generate an install URL and invite the bot to the right server.
|
|
420
|
+
11. Send a test message in Discord and confirm the brain responds.
|
|
93
421
|
|
|
94
422
|
Notes:
|
|
95
423
|
|
|
96
|
-
- Use **one
|
|
97
|
-
- Do not reuse the same Discord
|
|
424
|
+
- Use **one Discord application credential set per user/brain**.
|
|
425
|
+
- Do not reuse the same Discord application across multiple pilot users.
|
|
98
426
|
- Discord is the default pilot interface moving forward.
|
|
99
427
|
- The encrypted `users/<handle>.secrets.yaml.age` file is the durable checked-in deploy input; your local env is only the operator staging source.
|
|
100
|
-
- MCP
|
|
428
|
+
- Direct MCP client access should use OAuth/passkey-capable clients where possible.
|
|
101
429
|
- When explaining the content workflow, describe it first as a normal **git repo** of **markdown/text files**.
|
|
102
430
|
- Position **Obsidian** as optional: it is just one possible editor for those same files, not the default requirement.
|
|
103
431
|
|