create-flowdular 0.2.6 → 0.3.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/agent-template/.agents/skills/agent-tool-design/SKILL.md +1 -1
- package/agent-template/.agents/skills/auth-security-review/SKILL.md +1 -1
- package/agent-template/.agents/skills/cli-extension/SKILL.md +1 -1
- package/agent-template/.agents/skills/deploy-operate/SKILL.md +114 -0
- package/agent-template/.agents/skills/module-new/SKILL.md +29 -0
- package/agent-template/.agents/skills/module-update/SKILL.md +9 -1
- package/agent-template/.agents/skills/spec-interview/SKILL.md +114 -0
- package/agent-template/.agents/skills/ux-design/SKILL.md +34 -3
- package/agent-template/.ai/README.md +2 -1
- package/agent-template/.ai/agents/sandbox/business-manager.md +5 -1
- package/agent-template/.ai/blueprints/author-spec/README.md +1 -1
- package/agent-template/.ai/blueprints/author-spec/spec-requirements.yaml +44 -0
- package/agent-template/.ai/blueprints/author-spec/steps.yaml +5 -5
- package/agent-template/.ai/blueprints/author-spec/templates/module.yaml +99 -3
- package/agent-template/.ai/blueprints/edit-module/gates.yaml +4 -0
- package/agent-template/.ai/blueprints/edit-module/required-files.yaml +9 -0
- package/agent-template/.ai/blueprints/new-module/gates.yaml +4 -0
- package/agent-template/.ai/blueprints/new-module/spec-requirements.yaml +2 -2
- package/agent-template/.ai/blueprints/release/gates.yaml +4 -0
- package/agent-template/.ai/platform-capabilities.md +128 -0
- package/agent-template/.ai/policies/capabilities.yaml +130 -3
- package/agent-template/.ai/policies/path-ownership.yaml +5 -2
- package/agent-template/.ai/policies/task-budgets.yaml +5 -3
- package/agent-template/.ai/references/catalog/module.json +4 -4
- package/agent-template/.ai/references/catalog/package.json +2 -2
- package/agent-template/.ai/references/catalog/spec/module.yaml +5 -3
- package/agent-template/.ai/references/catalog/src/platform.ts +2 -0
- package/agent-template/.ai/references/catalog/src/services/catalog-service.ts +89 -1
- package/agent-template/.ai/references/catalog/src/services/data-classes.ts +47 -0
- package/agent-template/.ai/references/catalog/src/services/database-repository.ts +98 -1
- package/agent-template/.ai/references/catalog/src/services/repository.ts +22 -1
- package/agent-template/.ai/references/catalog/tests/data-classes.test.ts +157 -0
- package/agent-template/.ai/references/catalog.provenance.json +12 -10
- package/agent-template/.ai/rules/flowdular.md +4 -0
- package/agent-template/.ai/skills/README.md +10 -0
- package/agent-template/.ai/skills/agent-tool-design/SKILL.md +1 -2
- package/agent-template/.ai/skills/auth-security-review/SKILL.md +1 -1
- package/agent-template/.ai/skills/business-agent-design/SKILL.md +0 -1
- package/agent-template/.ai/skills/cli-extension/SKILL.md +1 -1
- package/agent-template/.ai/skills/deploy-operate/SKILL.md +119 -0
- package/agent-template/.ai/skills/module-new/SKILL.md +29 -3
- package/agent-template/.ai/skills/module-update/SKILL.md +9 -3
- package/agent-template/.ai/skills/perf-audit/SKILL.md +0 -1
- package/agent-template/.ai/skills/release-eject-pr/SKILL.md +0 -1
- package/agent-template/.ai/skills/spec-interview/SKILL.md +120 -0
- package/agent-template/.ai/skills/test-hardening/SKILL.md +1 -0
- package/agent-template/.ai/skills/ux-design/SKILL.md +34 -3
- package/agent-template/.ai/skills/variables/SKILL.md +0 -2
- package/agent-template/.ai/skills/workflow-development/SKILL.md +0 -1
- package/agent-template/.claude/skills/agent-tool-design/SKILL.md +1 -1
- package/agent-template/.claude/skills/auth-security-review/SKILL.md +1 -1
- package/agent-template/.claude/skills/cli-extension/SKILL.md +1 -1
- package/agent-template/.claude/skills/deploy-operate/SKILL.md +114 -0
- package/agent-template/.claude/skills/module-new/SKILL.md +29 -0
- package/agent-template/.claude/skills/module-update/SKILL.md +9 -1
- package/agent-template/.claude/skills/spec-interview/SKILL.md +114 -0
- package/agent-template/.claude/skills/ux-design/SKILL.md +34 -3
- package/agent-template/AGENTS.md +4 -0
- package/agent-template/CLAUDE.md +4 -0
- package/agent-template/docs/adr/0003-module-settings.md +1 -1
- package/agent-template/docs/adr/0006-agentic-workflows.md +24 -21
- package/agent-template/docs/agent-contract.md +2 -2
- package/agent-template/docs/cli-extensions.md +82 -0
- package/agent-template/docs/cli.md +195 -0
- package/agent-template/docs/configuration.md +593 -36
- package/agent-template/docs/design-system.md +185 -31
- package/agent-template/docs/getting-started.md +118 -0
- package/agent-template/docs/module-distribution.md +96 -0
- package/agent-template/docs/module-web-surfaces.md +221 -0
- package/agent-template/docs/modules.md +216 -0
- package/agent-template/docs/operations.md +545 -0
- package/agent-template/docs/sandbox.md +212 -0
- package/agent-template/platform/scripts/build.mjs +11 -0
- package/dist/bin.js +29 -0
- package/package.json +1 -1
- package/template/default/.dockerignore +14 -0
- package/template/default/.env.example +96 -0
- package/template/default/README.md +37 -1
- package/template/default/flowdular.json +15 -4
- package/template/default/infra/README.md +116 -0
- package/template/default/infra/docker/Dockerfile +37 -0
- package/template/default/infra/docker/compose.yaml +158 -0
- package/template/default/infra/docker/postgres/10-roles.sh +31 -0
- package/template/default/infra/docker/postgres/tls-init.sh +28 -0
- package/template/default/infra/kubernetes/database-secret.example.yaml +15 -0
- package/template/default/infra/kubernetes/deployment.yaml +211 -0
- package/template/default/infra/kubernetes/kustomization.yaml +9 -0
- package/template/default/infra/kubernetes/secrets.example.yaml +52 -0
- package/template/default/infra/kubernetes/service.yaml +13 -0
- package/template/default/modules/example/module.json +2 -1
- package/template/default/modules/example/package.json +1 -1
- package/template/default/modules/example/spec/module.yaml +1 -1
- package/template/default/modules/example/src/services/database-repository.ts +2 -12
- package/template/default/package.json +3 -2
- package/template/default/platform/octane.config.ts +99 -9
- package/template/default/platform/package.json +1 -1
- package/template/default/platform/src/generated/modules.client.ts +26 -2
- package/template/default/platform/src/generated/modules.server.ts +241 -10
- package/template/default/platform/src/server/health.ts +47 -0
- package/template/default/platform/src/server/metrics.ts +100 -0
- package/template/default/platform/src/server/storage.ts +172 -0
- package/template/default/platform/src/server/tracing.ts +85 -0
- package/template/default/specs/application.yaml +15 -0
|
@@ -0,0 +1,545 @@
|
|
|
1
|
+
# Operations
|
|
2
|
+
|
|
3
|
+
The runbook for a deployed Flowdular application: backing the data up, getting
|
|
4
|
+
it back, what key rotation does today, what to check before a release, and how
|
|
5
|
+
to roll one back. Environment variables are documented in
|
|
6
|
+
[configuration.md](configuration.md); the deployment itself in
|
|
7
|
+
[../infra/README.md](../infra/README.md).
|
|
8
|
+
|
|
9
|
+
## Backup
|
|
10
|
+
|
|
11
|
+
```bash
|
|
12
|
+
flowdular database backup --output /backups/flowdular/2026-09-11 # plan
|
|
13
|
+
flowdular database backup --output /backups/flowdular/2026-09-11 --apply # write
|
|
14
|
+
```
|
|
15
|
+
|
|
16
|
+
Like every write command it is a dry run without `--apply`: it prints the
|
|
17
|
+
adapter, the source it would read, the files it would write and the manifest,
|
|
18
|
+
and touches nothing. The target directory must not already hold a `backup.json`,
|
|
19
|
+
so a backup never overwrites another one.
|
|
20
|
+
|
|
21
|
+
A backup directory holds two things:
|
|
22
|
+
|
|
23
|
+
| File | Contents |
|
|
24
|
+
| ---------------------------- | ----------------------------------------------------------------------------- |
|
|
25
|
+
| `backup.json` | Timestamp, adapter, platform version, enabled modules, key fingerprints |
|
|
26
|
+
| `database.dump` or `pglite/` | The `pg_dump` custom-format archive, or a copy of the embedded data directory |
|
|
27
|
+
|
|
28
|
+
The command creates the directory with mode `0700` and writes `backup.json` and
|
|
29
|
+
`database.dump` with `0600`, so a dump is not readable by every account on the
|
|
30
|
+
host. A directory that already exists keeps the mode it was made with: when the
|
|
31
|
+
backup target is a mount or a path an earlier run created, check its mode
|
|
32
|
+
yourself. The same goes for `pglite/`, which is copied file by file and keeps
|
|
33
|
+
the modes of the source data directory.
|
|
34
|
+
|
|
35
|
+
**PostgreSQL.** The command runs `pg_dump --format=custom` with the migrator
|
|
36
|
+
connection of the configured environment (`FD_DATABASE_MIGRATOR_URL`, else
|
|
37
|
+
`FD_DATABASE_URL`). The credentials reach the child process as `PG*` variables,
|
|
38
|
+
never as arguments, so no password lands in the host's process list, and the
|
|
39
|
+
child inherits none of the platform's encryption keys. `FD_DATABASE_TLS` and
|
|
40
|
+
`FD_DATABASE_TLS_CA[_FILE]` are translated into `PGSSLMODE` and `PGSSLROOTCERT`,
|
|
41
|
+
so a `verify-full` deployment is dumped over a verified connection. Without the
|
|
42
|
+
client tools on `PATH` the command fails with `BACKUP_TOOL_MISSING`; install the
|
|
43
|
+
PostgreSQL client package matching the server version.
|
|
44
|
+
|
|
45
|
+
**Embedded PGlite.** Outside production the database is a directory
|
|
46
|
+
(`FD_DATABASE_PGLITE_DIRECTORY`, by default `.flowdular/data/pglite`), and the
|
|
47
|
+
command copies it. Stop the application first. A copy taken while the process is
|
|
48
|
+
writing can hold a torn state, and the command says so in a warning.
|
|
49
|
+
|
|
50
|
+
### What a backup does not contain
|
|
51
|
+
|
|
52
|
+
> **The encryption keys live outside the database.** Store every key with the
|
|
53
|
+
> backup. A restore without them yields unreadable credentials and workflow
|
|
54
|
+
> payloads: the rows come back, the values inside them stay ciphertext forever.
|
|
55
|
+
|
|
56
|
+
`backup.json` records, for each key, a SHA-256 fingerprint of the key material
|
|
57
|
+
and never the material itself:
|
|
58
|
+
|
|
59
|
+
```json
|
|
60
|
+
{
|
|
61
|
+
"schemaVersion": 1,
|
|
62
|
+
"createdAt": "2026-09-11T09:30:00.000Z",
|
|
63
|
+
"adapter": "postgresql",
|
|
64
|
+
"platformVersion": "0.2.0",
|
|
65
|
+
"modules": ["agents.core", "auth.core", "workflows.core"],
|
|
66
|
+
"keys": [
|
|
67
|
+
{
|
|
68
|
+
"variable": "FD_AGENT_CREDENTIAL_KEY",
|
|
69
|
+
"fingerprint": "sha256:0f1e..."
|
|
70
|
+
},
|
|
71
|
+
{ "variable": "FD_AUTH_MFA_KEY", "fingerprint": null }
|
|
72
|
+
]
|
|
73
|
+
}
|
|
74
|
+
```
|
|
75
|
+
|
|
76
|
+
The ten keys are `FD_AGENT_CREDENTIAL_KEY`, `FD_AGENT_RUN_GRANT_KEY`,
|
|
77
|
+
`FD_WORKFLOWS_PAYLOAD_KEY`, `FD_WORKFLOWS_CURSOR_KEY`, `FD_AUTH_MFA_KEY`,
|
|
78
|
+
`FD_AUTOMATIONS_CREDENTIAL_KEY`, `FD_NOTIFICATIONS_SECRET_KEY`,
|
|
79
|
+
`FD_STORAGE_ENCRYPTION_KEY`, `FD_CONNECTORS_SECRET_KEY` and
|
|
80
|
+
`FD_AUDIT_ANCHOR_KEY`. A `null` fingerprint means the key was not set
|
|
81
|
+
when the backup was taken. Keep the key material in the secret store the
|
|
82
|
+
deployment already uses, and record which backup it belongs to; the fingerprint
|
|
83
|
+
is what lets a restore tell you whether the pair matches.
|
|
84
|
+
|
|
85
|
+
Without `FD_AUDIT_ANCHOR_KEY` a restored deployment cannot verify the anchors
|
|
86
|
+
it holds, and every sealed segment file becomes a file whose origin nobody can
|
|
87
|
+
prove, so treat its fingerprint like the others.
|
|
88
|
+
|
|
89
|
+
The per-subject audit keys live in `audit_subject_keys` inside the database, so
|
|
90
|
+
a database backup carries them. A backup taken before
|
|
91
|
+
`flowdular audit erase --destroy-key` still holds the key that run destroyed, so
|
|
92
|
+
restoring it makes those sealed fields readable again. Treat an erasure as a
|
|
93
|
+
reason to age out the backups that predate it, as you would for any other
|
|
94
|
+
deletion on request. The key row itself survives destruction as a tombstone
|
|
95
|
+
carrying the subject's marker, which is what stops a later event about that
|
|
96
|
+
subject from creating a second key and putting the account back in the clear.
|
|
97
|
+
|
|
98
|
+
### Schedule
|
|
99
|
+
|
|
100
|
+
- Dump before every rollout, and on a daily schedule between them.
|
|
101
|
+
- Keep the archive off the database host, and keep the keys in the secret store,
|
|
102
|
+
not next to the dump.
|
|
103
|
+
- Keep one restore rehearsal in the calendar. A backup nobody has restored is a
|
|
104
|
+
guess; restoring into a scratch database is the only proof.
|
|
105
|
+
- Retention follows the data policy of the business, not this document. Delete
|
|
106
|
+
old keys only after the dumps they unlock are gone.
|
|
107
|
+
|
|
108
|
+
## Restore
|
|
109
|
+
|
|
110
|
+
```bash
|
|
111
|
+
flowdular database restore --input /backups/flowdular/2026-09-11 # plan
|
|
112
|
+
flowdular database restore --input /backups/flowdular/2026-09-11 --apply --confirm restore-database
|
|
113
|
+
```
|
|
114
|
+
|
|
115
|
+
The plan reports the manifest, the payload it would read and one key comparison
|
|
116
|
+
per variable. When the running environment holds keys the backup was not taken
|
|
117
|
+
with, the envelope carries a `BACKUP_KEY_MISMATCH` warning naming each variable
|
|
118
|
+
and how it differs (`different`, `missing-in-environment`, `missing-in-backup`).
|
|
119
|
+
Restore the keys before the data, or the rows arrive unreadable.
|
|
120
|
+
|
|
121
|
+
The order that works:
|
|
122
|
+
|
|
123
|
+
1. Stop the application, or take it out of the load balancer.
|
|
124
|
+
2. Put the keys the backup was taken with into the environment.
|
|
125
|
+
3. Run the restore as a dry run and read the key comparison.
|
|
126
|
+
4. Run it with `--apply --confirm restore-database`.
|
|
127
|
+
5. Start the application, then check `pnpm flowdular migration verify` and
|
|
128
|
+
`GET /api/ready`.
|
|
129
|
+
|
|
130
|
+
PostgreSQL restores run `pg_restore --clean --if-exists` against the migrator
|
|
131
|
+
connection, so the target database must already exist and its roles must already
|
|
132
|
+
be there. The embedded adapter is replaced directory by directory: the copy is
|
|
133
|
+
staged next to the target and the old directory is removed only once the new one
|
|
134
|
+
is in place, so a failed copy leaves the running database untouched.
|
|
135
|
+
|
|
136
|
+
`database restore` is a destructive capability, gated exactly like
|
|
137
|
+
`database reset`: the runner refuses it with `LOCAL_ONLY_CAPABILITY` unless
|
|
138
|
+
`FD_ENV` or `NODE_ENV` is `development` or `test`. The dry run on a staging
|
|
139
|
+
copy of the same backup is what proves the archive and the keys are good.
|
|
140
|
+
|
|
141
|
+
### Restore in production
|
|
142
|
+
|
|
143
|
+
```bash
|
|
144
|
+
flowdular database restore-production --input <dir> --target flowdular --grant <token> --tenant <id> # plan
|
|
145
|
+
flowdular database restore-production --input <dir> --target flowdular --grant <token> --tenant <id> \
|
|
146
|
+
--platform-url https://erp.example.com --apply --confirm restore-database
|
|
147
|
+
```
|
|
148
|
+
|
|
149
|
+
`database.restore.production` is the same restore without the local gate. It
|
|
150
|
+
runs only under an approval grant from an approved `approvals.core` request,
|
|
151
|
+
verified against `FD_APPROVAL_GRANT_KEY`, the tenant, the capability and the
|
|
152
|
+
exact flags: the request has to name `--input`, `--target`, `--apply` and any
|
|
153
|
+
override (an approval names the applied run, and a dry run is a different
|
|
154
|
+
invocation), and a token issued for another invocation is refused with
|
|
155
|
+
`APPROVAL_GRANT_MISMATCH`. Open the request with `approvals.requests.v1` and a
|
|
156
|
+
`subjectRef` of `capability:<id>:<sha256 of the invocation input>`, read the
|
|
157
|
+
approved request back through `grant(tenantId, id, subjectModule)` for the
|
|
158
|
+
token, and pass it as `--grant <token> --tenant <id>`. `--target` must repeat
|
|
159
|
+
the database the migrator DSN names (`RESTORE_TARGET_MISMATCH` otherwise),
|
|
160
|
+
`FD_DATABASE_MIGRATOR_URL` must be set and distinct from `FD_DATABASE_URL`
|
|
161
|
+
(`MIGRATOR_ROLE_REQUIRED`), and a key divergence is a refusal
|
|
162
|
+
(`BACKUP_KEY_MISMATCH`) unless `--allow-key-mismatch` was part of the approved
|
|
163
|
+
invocation. With `--apply` the command probes `<--platform-url>/api/health`,
|
|
164
|
+
or `http://127.0.0.1:$FD_PORT/api/health`, and refuses with
|
|
165
|
+
`PLATFORM_RUNNING` while anything answers; when no endpoint is named or the
|
|
166
|
+
probe times out it refuses with `PLATFORM_STATE_UNKNOWN` unless
|
|
167
|
+
`--platform-stopped` attests the platform is down. Grants are HMAC-SHA256
|
|
168
|
+
under `FD_APPROVAL_GRANT_KEY` (rotate with `FD_APPROVAL_GRANT_KEY_PREVIOUS`)
|
|
169
|
+
and expire one hour after the approving decision; the runner records no use,
|
|
170
|
+
so a grant replays until it expires and the approval window should stay short.
|
|
171
|
+
`database reset` and `database restore` keep their local-only gate.
|
|
172
|
+
|
|
173
|
+
### Point-in-time recovery
|
|
174
|
+
|
|
175
|
+
A dump restores one moment; the compose stack also archives WAL so any moment
|
|
176
|
+
after a base backup can be recovered. See `infra/README.md`, "Backups and
|
|
177
|
+
PITR": `infra/docker/pitr.sh base-backup` after every rollout,
|
|
178
|
+
`pitr.sh restore --base <stamp> --target-time '<ts>' --confirm replace-cluster`
|
|
179
|
+
to recover, and the limits of the plain-copy archive. In Kubernetes the
|
|
180
|
+
managed provider owns PITR; the dump remains the portable copy.
|
|
181
|
+
|
|
182
|
+
### Rehearsal
|
|
183
|
+
|
|
184
|
+
Quarterly, on a scratch host: restore the newest dump into an empty database
|
|
185
|
+
with `database restore-production --platform-stopped`, read the key
|
|
186
|
+
comparison, start the app, check `migration verify` and `/api/ready`, then run
|
|
187
|
+
`pitr.sh restore` from the newest base backup to a time between two dumps and
|
|
188
|
+
confirm a row written after the dump is present. Record the date and the
|
|
189
|
+
elapsed time.
|
|
190
|
+
|
|
191
|
+
## Key rotation
|
|
192
|
+
|
|
193
|
+
Most of the keys carry a key id in every envelope they write, so a
|
|
194
|
+
deployment can hold the new key and the old one at the same time and re-seal the
|
|
195
|
+
stored rows in place. The others have no key id and are covered at the end of
|
|
196
|
+
this section.
|
|
197
|
+
|
|
198
|
+
Every module reads one current key plus an optional comma-separated list of
|
|
199
|
+
retired keys (up to eight):
|
|
200
|
+
|
|
201
|
+
| Key | Protects | Retired keys | Re-sealing the stored rows |
|
|
202
|
+
| ------------------------------- | -------------------------- | ---------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------- |
|
|
203
|
+
| `FD_AGENT_CREDENTIAL_KEY` | Agent provider credentials | `FD_AGENT_CREDENTIAL_KEY_PREVIOUS` | `pnpm flowdular agents secrets-rotate [--apply]` |
|
|
204
|
+
| `FD_AUTOMATIONS_CREDENTIAL_KEY` | Automation trigger secrets | `FD_AUTOMATIONS_CREDENTIAL_KEY_PREVIOUS` | `pnpm flowdular automations secrets-rotate [--apply]` |
|
|
205
|
+
| `FD_WORKFLOWS_PAYLOAD_KEY` | Workflow run payloads | `FD_WORKFLOWS_PAYLOAD_KEY_PREVIOUS` | `pnpm flowdular workflows secrets-rotate [--apply]` |
|
|
206
|
+
| `FD_STORAGE_ENCRYPTION_KEY` | Stored objects | `FD_STORAGE_ENCRYPTION_KEY_PREVIOUS` | `pnpm flowdular documents secrets-rotate [--apply]` and `pnpm flowdular exports secrets-rotate [--apply]`; see the storage note below |
|
|
207
|
+
| `FD_NOTIFICATIONS_SECRET_KEY` | Webhook signing secrets | `FD_NOTIFICATIONS_SECRET_KEY_PREVIOUS` | `pnpm flowdular notifications secrets-rotate [--apply]` |
|
|
208
|
+
| `FD_CONNECTORS_SECRET_KEY` | Connector credentials | `FD_CONNECTORS_SECRET_KEY_PREVIOUS` | `pnpm flowdular connectors secrets-rotate [--apply]` |
|
|
209
|
+
| `FD_AUDIT_ANCHOR_KEY` | Audit chain anchors (HMAC) | `FD_AUDIT_ANCHOR_KEY_PREVIOUS` | `pnpm flowdular audit secrets-rotate [--apply]` |
|
|
210
|
+
|
|
211
|
+
These commands run the same re-sealing pass over their own table, so the
|
|
212
|
+
procedure is the same for each. The credential key is the worked example;
|
|
213
|
+
substitute the key names and the command from the row you are rotating:
|
|
214
|
+
|
|
215
|
+
1. Generate the new key: `openssl rand -base64 32`.
|
|
216
|
+
2. Set the new key as the current one and the old key as the retired one:
|
|
217
|
+
`FD_AGENT_CREDENTIAL_KEY=<new>` and `FD_AGENT_CREDENTIAL_KEY_PREVIOUS=<old>`.
|
|
218
|
+
Keep both in the secret store, exactly like the key they replace.
|
|
219
|
+
3. Deploy. From this moment every write uses the new key and every stored row
|
|
220
|
+
still opens under the old one, so nothing is unreadable while the rollout
|
|
221
|
+
completes.
|
|
222
|
+
4. Run the rotation as a dry run and read the counts per key id:
|
|
223
|
+
`pnpm flowdular agents secrets-rotate --json`. `stale` is the number of rows
|
|
224
|
+
still sealed with a retired key.
|
|
225
|
+
5. Run it again with `--apply`. It re-seals rows in batches of 200, one
|
|
226
|
+
transaction per batch, scoped to the tenant that owns them, and writes nothing
|
|
227
|
+
else: the revision, the audit fields and the credential a caller sees do not
|
|
228
|
+
change. It is idempotent, so running it twice is free.
|
|
229
|
+
6. Repeat the dry run until it reports `stale: 0`, then remove
|
|
230
|
+
`FD_AGENT_CREDENTIAL_KEY_PREVIOUS` and deploy again.
|
|
231
|
+
|
|
232
|
+
The application may stay up. Each row is written back only if its stored
|
|
233
|
+
envelope is still the one that was read, so a credential the application
|
|
234
|
+
rewrites in between is never clobbered: the rotation counts it under `skipped`
|
|
235
|
+
and the next run picks it up. Stopping the application removes that case
|
|
236
|
+
entirely. The inventory that finds the stale rows runs on the cross-tenant
|
|
237
|
+
read-only role, which is granted the tenant id and the key id and cannot read a
|
|
238
|
+
nonce, a tag or a ciphertext; every row it names is read again under its own
|
|
239
|
+
tenant before it is touched.
|
|
240
|
+
|
|
241
|
+
`FD_AUDIT_ANCHOR_KEY` follows the same six steps with one difference: an anchor
|
|
242
|
+
is signed, not sealed, so the rotation rewrites the signature and the key id and
|
|
243
|
+
touches neither the anchor hash, the segment hash nor the chain. A segment file
|
|
244
|
+
written under the retired key keeps its own signature and still verifies while
|
|
245
|
+
that key is in `FD_AUDIT_ANCHOR_KEY_PREVIOUS`. Keep the retired key until every
|
|
246
|
+
archived segment has been verified once under the new one, with
|
|
247
|
+
`pnpm flowdular audit verify --workspace <slug> --input <dir>` over the archive.
|
|
248
|
+
If the rotation reports `ANCHOR_KEY_UNKNOWN`, anchors exist that neither key can
|
|
249
|
+
re-sign: put that key back before retiring it, because a signature nobody can
|
|
250
|
+
reproduce can never be verified again. The rotation verifies an anchor's stored
|
|
251
|
+
signature under the ring before re-signing it, so `ANCHOR_SIGNATURE_INVALID`
|
|
252
|
+
names anchors it left exactly as they are: a new signature over the same anchor
|
|
253
|
+
hash would replace evidence nobody can account for with one this deployment
|
|
254
|
+
vouches for. Verify the segment files of those anchors before anything else.
|
|
255
|
+
`FD_AUDIT_ANCHOR_KEY_PREVIOUS` accepts at most eight retired keys; a longer list
|
|
256
|
+
is a rotation that was never finished and the platform refuses to start.
|
|
257
|
+
|
|
258
|
+
The remaining keys, and how rotation behaves for each:
|
|
259
|
+
|
|
260
|
+
| Key | Protects | Rotating it |
|
|
261
|
+
| ------------------------- | -------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
|
|
262
|
+
| `FD_WORKFLOWS_CURSOR_KEY` | Pagination cursors (HMAC) | Safe. `FD_WORKFLOWS_CURSOR_KEY_PREVIOUS` keeps outstanding cursors valid; cursors are never stored, so there is nothing to re-sign. |
|
|
263
|
+
| `FD_AGENT_RUN_GRANT_KEY` | Agent run grant tokens (HMAC, 30s TTL) | Safe. In-flight grants fail for the length of one TTL. |
|
|
264
|
+
| `FD_AUTH_MFA_KEY` | Enrolled TOTP secrets | Supported since auth.core 0.11.0. Set `FD_AUTH_MFA_KEY_PREVIOUS`, deploy, run `pnpm flowdular auth secrets-rotate --apply`, then drop the previous key. Rows written before the key id column exist open through the ring fallback until they are re-sealed. |
|
|
265
|
+
|
|
266
|
+
The key ids are derived from the key material, so a backup fingerprint that
|
|
267
|
+
matches the environment also matches the `keyId` inside the rows. The keys that
|
|
268
|
+
carry no id are the ones a restore cannot diagnose for you, which is why
|
|
269
|
+
`backup.json` records a fingerprint for all eight.
|
|
270
|
+
|
|
271
|
+
Sources: `modules/audit/src/services/{anchor-key,anchor-rotation}.ts`,
|
|
272
|
+
`packages/kernel/src/keyring.ts`,
|
|
273
|
+
`modules/agents/src/services/{credential-vault,credential-rotation}.ts`,
|
|
274
|
+
`modules/automations/src/services/{secret-vault,secret-rotation}.ts`,
|
|
275
|
+
`modules/workflows/src/services/{payload-codec,payload-rotation,cursors}.ts`,
|
|
276
|
+
`modules/notifications/src/services/{secret-vault,secret-rotation}.ts`,
|
|
277
|
+
`modules/connectors/src/services/{credential-vault,credential-rotation}.ts`,
|
|
278
|
+
`packages/storage/src/reseal.ts`,
|
|
279
|
+
`modules/documents/src/services/storage-rotation.ts`,
|
|
280
|
+
`modules/exports/src/services/storage-rotation.ts`,
|
|
281
|
+
`modules/auth/src/services/totp.ts`.
|
|
282
|
+
|
|
283
|
+
## Storage
|
|
284
|
+
|
|
285
|
+
Objects live outside PostgreSQL, so a database backup does not contain one. Back
|
|
286
|
+
up the object store the same way you back up the database, and take both at the
|
|
287
|
+
same point: a restored database that references objects a newer bucket snapshot
|
|
288
|
+
no longer holds leaves dangling references, and the reverse leaves orphans.
|
|
289
|
+
|
|
290
|
+
- **S3-compatible store.** Enable versioning and a lifecycle policy on the
|
|
291
|
+
bucket, or replicate it. `backup.json` does not reach into it.
|
|
292
|
+
- **Local adapter.** Development and test only, refused in production. Its
|
|
293
|
+
directory is `FD_STORAGE_LOCAL_DIRECTORY`, by default
|
|
294
|
+
`.flowdular/data/storage`; stop the application before copying it.
|
|
295
|
+
|
|
296
|
+
Every object is sealed with AES-256-GCM under `FD_STORAGE_ENCRYPTION_KEY`, and
|
|
297
|
+
the key id is stored with the object, so the rotation follows the six steps
|
|
298
|
+
above with two commands instead of one, because the object store has no
|
|
299
|
+
listing and the rows that name the objects belong to two modules:
|
|
300
|
+
|
|
301
|
+
1. Generate the new key: `openssl rand -base64 32`.
|
|
302
|
+
2. Set `FD_STORAGE_ENCRYPTION_KEY=<new>` and
|
|
303
|
+
`FD_STORAGE_ENCRYPTION_KEY_PREVIOUS=<old>` in the secret store.
|
|
304
|
+
3. Deploy. Every new object is sealed with the new key and every stored one
|
|
305
|
+
still opens under the old one.
|
|
306
|
+
4. Dry run both passes and read the counts per key id:
|
|
307
|
+
`pnpm flowdular documents secrets-rotate --json` walks the stored document
|
|
308
|
+
rows, `pnpm flowdular exports secrets-rotate --json` the completed export
|
|
309
|
+
jobs. `stale` is the number of objects still sealed with a retired key;
|
|
310
|
+
`unknown` names objects under a key neither variable holds, and the pass
|
|
311
|
+
leaves those alone.
|
|
312
|
+
5. Run both again with `--apply`. Each object is re-sealed in place: the frame
|
|
313
|
+
is opened under the key it names, which authenticates its header, and the
|
|
314
|
+
same content type, size, checksum, scan verdict and creation time are
|
|
315
|
+
written back under the current key. Nothing else changes, not the object
|
|
316
|
+
key and not the row. A frame that fails authentication is counted under
|
|
317
|
+
`refused` and left as it is; restore it from the object store backup.
|
|
318
|
+
6. Repeat the dry runs until both report `stale: 0`, then remove
|
|
319
|
+
`FD_STORAGE_ENCRYPTION_KEY_PREVIOUS` and deploy again.
|
|
320
|
+
|
|
321
|
+
The application may stay up: the pass locks a row while it rewrites the
|
|
322
|
+
object, and both delete paths (a document removal and the export retention
|
|
323
|
+
sweep) lock the row before they remove the object and then the row, so a delete
|
|
324
|
+
that lands mid-pass waits for the rewrite and removes the re-sealed object
|
|
325
|
+
rather than racing it. A frame that fails to parse is counted under `refused`
|
|
326
|
+
like one that fails authentication. Do not snapshot the object store and the database apart while a
|
|
327
|
+
pass runs: a bucket snapshot taken mid-pass holds objects under both keys, and
|
|
328
|
+
restoring it beside a database from another moment leaves rows that name
|
|
329
|
+
objects a ring without the retired key cannot open. Take both after the pass,
|
|
330
|
+
or both before it. Dropping the retired key while `stale` is above zero makes
|
|
331
|
+
those objects unreadable, exactly as a lost database key makes a credential
|
|
332
|
+
unreadable.
|
|
333
|
+
|
|
334
|
+
## Logs
|
|
335
|
+
|
|
336
|
+
The server writes one line per event through `createLogger`
|
|
337
|
+
(`packages/server/src/log.ts`).
|
|
338
|
+
|
|
339
|
+
| Variable | Default | Purpose |
|
|
340
|
+
| --------------- | --------------------------------- | ------------------------------------------------ |
|
|
341
|
+
| `FD_LOG_FORMAT` | `json` in production, else `text` | `json` for a log pipeline, `text` for a terminal |
|
|
342
|
+
| `FD_LOG_LEVEL` | `info` | `debug`, `info`, `warn` or `error` |
|
|
343
|
+
|
|
344
|
+
A JSON line carries `time`, `level`, `msg` and, where they apply, `traceId`,
|
|
345
|
+
`spanId`, `requestId`, `endpoint`, `module` and `err` (name and message). The
|
|
346
|
+
trace ids are on every line written inside a served request, so the same id
|
|
347
|
+
finds the log line and the trace. Stacks are written outside
|
|
348
|
+
production, and in production only at `FD_LOG_LEVEL=debug`. The `requestId` is
|
|
349
|
+
the same value the response returns in `x-request-id`, so a user-reported error
|
|
350
|
+
id finds its log line. Credential-shaped field names (`authorization`, `cookie`,
|
|
351
|
+
`password`, `token`, `secret`, `credential`, API keys) are redacted before a
|
|
352
|
+
line is written, and a failing endpoint logs its id and the error type only,
|
|
353
|
+
never the message, the body or the URL. An unrecognized value for either
|
|
354
|
+
variable falls back to the default and logs one warning about it.
|
|
355
|
+
|
|
356
|
+
## Production checklist
|
|
357
|
+
|
|
358
|
+
- `pnpm verify` and `pnpm build` pass on the commit being shipped.
|
|
359
|
+
- `FD_ENV=production`, `FD_DATABASE_ADAPTER=postgresql`, separate
|
|
360
|
+
`FD_DATABASE_URL` and `FD_DATABASE_MIGRATOR_URL` roles, neither `SUPERUSER`
|
|
361
|
+
nor `BYPASSRLS` on the runtime role.
|
|
362
|
+
- `FD_DATABASE_TLS=verify-full` with `FD_DATABASE_TLS_CA` or
|
|
363
|
+
`FD_DATABASE_TLS_CA_FILE`. Production accepts nothing weaker.
|
|
364
|
+
- All six encryption keys set, held in the secret store, and backed up with a
|
|
365
|
+
recorded fingerprint.
|
|
366
|
+
- `FD_AUTH_SECURE_COOKIE` on, `FD_AUTH_PUBLIC_ORIGIN` equal to the public URL,
|
|
367
|
+
`FD_TRUST_PROXY` set behind a proxy, sign-up closed unless the deployment is
|
|
368
|
+
public.
|
|
369
|
+
- Mail transport configured: `FD_AUTH_MAIL_TRANSPORT` with `FD_AUTH_SMTP_URL`
|
|
370
|
+
and `FD_AUTH_MAIL_FROM`. Password resets, invitations and address
|
|
371
|
+
confirmations are silently undeliverable without it.
|
|
372
|
+
- Migrations at rollout: deploy the image, let it apply its module migrations
|
|
373
|
+
under the migrator role, then `pnpm flowdular migration verify`. Never edit an
|
|
374
|
+
applied `.up.sql`.
|
|
375
|
+
- Liveness probe on `GET /api/health`, readiness probe on `GET /api/ready`.
|
|
376
|
+
- `FD_LOG_FORMAT=json` and a log pipeline that keeps `requestId`.
|
|
377
|
+
- A fresh backup, taken and verified before the rollout starts.
|
|
378
|
+
|
|
379
|
+
## Rollback
|
|
380
|
+
|
|
381
|
+
1. Redeploy the previous image tag. Module migrations are additive, so an older
|
|
382
|
+
image runs against a newer schema and ignores columns it does not know.
|
|
383
|
+
2. Never apply a `.down.sql` to roll back a release. They document the reverse
|
|
384
|
+
for review; against live data they are data loss.
|
|
385
|
+
3. To take one module out of service, `pnpm flowdular module disable <id> --apply`,
|
|
386
|
+
rebuild the composition and redeploy. Its tables stay where they are.
|
|
387
|
+
4. If the rollback is about a key, restore the key first. A previous image does
|
|
388
|
+
not make ciphertext readable again.
|
|
389
|
+
5. Restore from a backup only when the schema or the data is actually damaged.
|
|
390
|
+
It is the slowest option and the only one that loses everything written since
|
|
391
|
+
the dump.
|
|
392
|
+
|
|
393
|
+
## Metrics
|
|
394
|
+
|
|
395
|
+
`GET /api/metrics` serves Prometheus text exposition, format version 0.0.4. The
|
|
396
|
+
route is composed only when `FD_METRICS=true`; with `FD_METRICS_TOKEN` set, a
|
|
397
|
+
scrape must send `Authorization: Bearer <token>` and gets 401 without it.
|
|
398
|
+
|
|
399
|
+
**An untokened `/api/metrics` must never be reachable outside the scrape
|
|
400
|
+
network.** It shares the application port, so without `FD_METRICS_TOKEN` every
|
|
401
|
+
client that can reach the port reads the series, including the endpoint
|
|
402
|
+
inventory and the request rates of the whole deployment. Either set the token
|
|
403
|
+
(the Kubernetes manifests read it from the deployment Secret, optional so it
|
|
404
|
+
can be left unset deliberately) or keep the port on a network only the scraper
|
|
405
|
+
reaches. Do not rely on the route being obscure.
|
|
406
|
+
|
|
407
|
+
| Series | Type | Labels |
|
|
408
|
+
| ----------------------------------------- | --------- | ----------------------------------- |
|
|
409
|
+
| `flowdular_http_requests_total` | counter | `endpoint`, `method`, `status` |
|
|
410
|
+
| `flowdular_http_request_duration_seconds` | histogram | `endpoint`, buckets 0.005 s to 10 s |
|
|
411
|
+
| `flowdular_metrics_dropped_samples_total` | counter | none |
|
|
412
|
+
| `flowdular_process_start_time_seconds` | gauge | none |
|
|
413
|
+
| `process_resident_memory_bytes` | gauge | none |
|
|
414
|
+
| `nodejs_eventloop_lag_seconds` | gauge | none |
|
|
415
|
+
| `flowdular_build_info` | gauge | `version` |
|
|
416
|
+
|
|
417
|
+
`endpoint` is the endpoint id, never the request path, so a record id cannot
|
|
418
|
+
reach a label. `status` is the class (`2xx`, `4xx`, `5xx`), and an unrecognized
|
|
419
|
+
method is recorded as `other`. No session, tenant or user value is ever a label.
|
|
420
|
+
|
|
421
|
+
`nodejs_eventloop_lag_seconds` is the mean delay over the interval since the
|
|
422
|
+
previous scrape, not since the process started, so a spike shows up in the
|
|
423
|
+
scrape that follows it and is gone from the next one. The sampler is
|
|
424
|
+
process-wide and is drained on every read, so two scrapers pointed at the same
|
|
425
|
+
process split the samples between them; point one scraper at it.
|
|
426
|
+
|
|
427
|
+
Every request through `defineEndpoint` is recorded, whether or not the route is
|
|
428
|
+
exposed, so a scraper enabled later sees the counters immediately. Memory is
|
|
429
|
+
bounded: the process keeps at most 2000 distinct label sets per series family
|
|
430
|
+
and refuses new ones after that, counting each refusal in
|
|
431
|
+
`flowdular_metrics_dropped_samples_total`. A non-zero value there means a
|
|
432
|
+
deployment composes more endpoints than the ceiling allows.
|
|
433
|
+
|
|
434
|
+
A module adds series of its own with `createModuleMetrics('<module id>')`:
|
|
435
|
+
|
|
436
|
+
| Series | Type | Labels |
|
|
437
|
+
| ----------------------------------------------------------------- | --------- | --------------------------- |
|
|
438
|
+
| `flowdular_module_<module>_<name>_total` | counter | the module's own |
|
|
439
|
+
| `flowdular_module_<module>_<name>` (+`_bucket`, `_sum`, `_count`) | histogram | the module's own, plus `le` |
|
|
440
|
+
|
|
441
|
+
The module id is lower-cased with every character outside `a-z0-9_` replaced by
|
|
442
|
+
`_`, so `agents.core` reads `flowdular_module_agents_core_`. A module series is
|
|
443
|
+
refused, and counted as a dropped sample, when the metric name does not match
|
|
444
|
+
`^[a-z][a-z0-9_]{0,63}$`, when a call carries more than 8 labels or a label name
|
|
445
|
+
of the wrong shape, when the family already holds 2000 label sets, or when the
|
|
446
|
+
process already opened 256 module families of that kind. A label value is cut at
|
|
447
|
+
128 characters. A module never receives an error from recording: a refused
|
|
448
|
+
sample is lost, the work is not.
|
|
449
|
+
|
|
450
|
+
**The label rule of the request series holds for a module series too: never a
|
|
451
|
+
tenant, account, record or session id.** The platform cannot tell one from a
|
|
452
|
+
closed set, so nothing refuses it; what it does is cap the damage at 2000 label
|
|
453
|
+
sets per family, after which the series stops admitting new ones. A label is for
|
|
454
|
+
a value with a handful of possibilities, such as an outcome or a channel.
|
|
455
|
+
|
|
456
|
+
## Tracing
|
|
457
|
+
|
|
458
|
+
Every request through `defineEndpoint` runs in a W3C trace. The inbound
|
|
459
|
+
`traceparent` names the parent when it is usable; anything else is a new root.
|
|
460
|
+
The response always returns a `traceparent` naming the span the request ran in,
|
|
461
|
+
so a client, a proxy and this process agree on one trace id, and the log lines
|
|
462
|
+
of that request carry the same `traceId`.
|
|
463
|
+
|
|
464
|
+
Spans go into one bounded in-process ring whether or not an exporter is
|
|
465
|
+
composed: 4096 spans, oldest evicted and counted, at most 16 attributes per span
|
|
466
|
+
and 256 characters per name, attribute key and string value. Attributes carry
|
|
467
|
+
endpoint ids, job names, provider and tool ids, HTTP method and status, and the
|
|
468
|
+
three ids that correlate a span with a log line or an agent run,
|
|
469
|
+
`flowdular.request_id`, `flowdular.agent.run_id` and
|
|
470
|
+
`flowdular.tool.provider_call_id`: no request path, body, tenant or session.
|
|
471
|
+
|
|
472
|
+
| Variable | Purpose |
|
|
473
|
+
| ----------------------- | --------------------------------------------------- |
|
|
474
|
+
| `FD_TRACE_SAMPLE` | Ratio of new roots recorded, 0 to 1, default 1 |
|
|
475
|
+
| `FD_TRACE_EXPORTER` | `none` (default) or `otlp` |
|
|
476
|
+
| `FD_TRACE_OTLP_URL` | OTLP/HTTP JSON traces endpoint; https in production |
|
|
477
|
+
| `FD_TRACE_OTLP_HEADERS` | `name=value,name2=value2`, at most 16 |
|
|
478
|
+
|
|
479
|
+
The exporter posts OTLP/HTTP JSON (no SDK, no agent) in batches of 512 spans, or
|
|
480
|
+
every 5 seconds when the batch does not fill. It never runs on a request path: a
|
|
481
|
+
collector that is down costs the process one held batch, retried on the next two
|
|
482
|
+
flushes and then dropped. A `4xx` that is not `429` is not retried, because the
|
|
483
|
+
payload or the credential is wrong and the next batch would be lost too. The
|
|
484
|
+
sampling decision is taken on the trace id, so a whole trace is in or out; a
|
|
485
|
+
trace the caller already marked sampled is never re-decided, and an unsampled
|
|
486
|
+
trace still propagates its header.
|
|
487
|
+
|
|
488
|
+
A process that stops drains what the ring still holds, in the same batches and
|
|
489
|
+
without a retry: the spans of a shutdown leave with it, and a collector that is
|
|
490
|
+
down is given up on rather than held between the process and its exit.
|
|
491
|
+
|
|
492
|
+
**An OTLP endpoint receives the endpoint inventory, the job names, the timings
|
|
493
|
+
and the request, agent run and provider call ids of the whole deployment.**
|
|
494
|
+
Point it at a collector you control. Plain `http` is refused in production.
|
|
495
|
+
|
|
496
|
+
## Error reporting
|
|
497
|
+
|
|
498
|
+
`FD_ERROR_SINK=webhook` posts every line the server logs at `error` to
|
|
499
|
+
`FD_ERROR_SINK_URL`, with `Authorization: Bearer <FD_ERROR_SINK_TOKEN>` when the
|
|
500
|
+
token is set. The body is `{"reports":[…]}` where a report carries `at`, `name`,
|
|
501
|
+
`message` and, where they apply, `endpoint`, `module`, `requestId`, `traceId`
|
|
502
|
+
and `spanId`. Nothing else: no stack, no log fields, no request body, so the
|
|
503
|
+
sink cannot carry a credential the logger already refuses to write.
|
|
504
|
+
|
|
505
|
+
`message` is the one field the caller chose. A failing endpoint logs its error
|
|
506
|
+
type only, never the message, so its report carries `endpoint failed`; a
|
|
507
|
+
background loop logs the error it caught, so its report carries that message.
|
|
508
|
+
The sink therefore exports exactly what the log line already shows, to a host
|
|
509
|
+
outside the deployment. A deployment that must not let error text leave keeps
|
|
510
|
+
`FD_ERROR_SINK=none` and reads the log stream instead.
|
|
511
|
+
|
|
512
|
+
It is bounded and lossy on purpose: 64 queued reports (a report past that is
|
|
513
|
+
dropped and counted), 32 per request, 8 KiB per body with the batch halved until
|
|
514
|
+
it fits, a batch sent every 5 seconds or as soon as 32 are queued, and at most 3
|
|
515
|
+
attempts before a batch is given up. Delivery never runs on the caller's path
|
|
516
|
+
and an unreachable webhook never fails the line that reported the failure.
|
|
517
|
+
`none` is the default and holds no queue and no timer.
|
|
518
|
+
|
|
519
|
+
The OTLP logs sink named in RFC 0004 H7 is deliberately not built: the OTLP
|
|
520
|
+
surface this platform speaks is the trace exporter above.
|
|
521
|
+
|
|
522
|
+
## Data residency
|
|
523
|
+
|
|
524
|
+
Residency is a property of the deployment, not of the workspace. A deployment
|
|
525
|
+
carries one PostgreSQL database, one object store and one region; every module
|
|
526
|
+
writes to that database and, through the storage port, to that bucket. A
|
|
527
|
+
workspace therefore lives wherever its deployment lives, and a customer whose
|
|
528
|
+
data must stay in a particular region gets a deployment in that region. That
|
|
529
|
+
is the supported answer and the only one the platform offers.
|
|
530
|
+
|
|
531
|
+
Per-workspace routing to regional databases is not planned. It would change
|
|
532
|
+
the provider contract in `packages/database` (one handle per region), the
|
|
533
|
+
lease model every module's migrations and background loops depend on, every
|
|
534
|
+
cross-workspace routing read the background role performs (the retention
|
|
535
|
+
sweep, delivery queues, expiry passes), the storage key layout, and the
|
|
536
|
+
export and erasure paths that assume one place to look. Record that cost
|
|
537
|
+
before anyone proposes it; a second deployment is cheaper in every case
|
|
538
|
+
seen so far.
|
|
539
|
+
|
|
540
|
+
What to state in a questionnaire: the region of the database and the bucket,
|
|
541
|
+
that backups and the key material are stored in the same region unless the
|
|
542
|
+
operator moves them, that the OTLP trace exporter and the error sink send
|
|
543
|
+
only what `docs/operations.md` lists for them and only to the URLs the
|
|
544
|
+
operator configured, and that outbound mail, webhooks and connector calls
|
|
545
|
+
leave the region by design of the recipient the workspace configured.
|