create-flowdular 0.2.6 → 0.3.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (103) hide show
  1. package/agent-template/.agents/skills/agent-tool-design/SKILL.md +1 -1
  2. package/agent-template/.agents/skills/auth-security-review/SKILL.md +1 -1
  3. package/agent-template/.agents/skills/cli-extension/SKILL.md +1 -1
  4. package/agent-template/.agents/skills/deploy-operate/SKILL.md +114 -0
  5. package/agent-template/.agents/skills/module-new/SKILL.md +29 -0
  6. package/agent-template/.agents/skills/module-update/SKILL.md +9 -1
  7. package/agent-template/.agents/skills/spec-interview/SKILL.md +114 -0
  8. package/agent-template/.agents/skills/ux-design/SKILL.md +34 -3
  9. package/agent-template/.ai/README.md +2 -1
  10. package/agent-template/.ai/agents/sandbox/business-manager.md +5 -1
  11. package/agent-template/.ai/blueprints/author-spec/README.md +1 -1
  12. package/agent-template/.ai/blueprints/author-spec/spec-requirements.yaml +44 -0
  13. package/agent-template/.ai/blueprints/author-spec/steps.yaml +5 -5
  14. package/agent-template/.ai/blueprints/author-spec/templates/module.yaml +99 -3
  15. package/agent-template/.ai/blueprints/edit-module/gates.yaml +4 -0
  16. package/agent-template/.ai/blueprints/edit-module/required-files.yaml +9 -0
  17. package/agent-template/.ai/blueprints/new-module/gates.yaml +4 -0
  18. package/agent-template/.ai/blueprints/new-module/spec-requirements.yaml +2 -2
  19. package/agent-template/.ai/blueprints/release/gates.yaml +4 -0
  20. package/agent-template/.ai/platform-capabilities.md +128 -0
  21. package/agent-template/.ai/policies/capabilities.yaml +130 -3
  22. package/agent-template/.ai/policies/path-ownership.yaml +5 -2
  23. package/agent-template/.ai/policies/task-budgets.yaml +5 -3
  24. package/agent-template/.ai/references/catalog/module.json +4 -4
  25. package/agent-template/.ai/references/catalog/package.json +2 -2
  26. package/agent-template/.ai/references/catalog/spec/module.yaml +5 -3
  27. package/agent-template/.ai/references/catalog/src/platform.ts +2 -0
  28. package/agent-template/.ai/references/catalog/src/services/catalog-service.ts +89 -1
  29. package/agent-template/.ai/references/catalog/src/services/data-classes.ts +47 -0
  30. package/agent-template/.ai/references/catalog/src/services/database-repository.ts +98 -1
  31. package/agent-template/.ai/references/catalog/src/services/repository.ts +22 -1
  32. package/agent-template/.ai/references/catalog/tests/data-classes.test.ts +157 -0
  33. package/agent-template/.ai/references/catalog.provenance.json +12 -10
  34. package/agent-template/.ai/rules/flowdular.md +4 -0
  35. package/agent-template/.ai/skills/README.md +10 -0
  36. package/agent-template/.ai/skills/agent-tool-design/SKILL.md +1 -2
  37. package/agent-template/.ai/skills/auth-security-review/SKILL.md +1 -1
  38. package/agent-template/.ai/skills/business-agent-design/SKILL.md +0 -1
  39. package/agent-template/.ai/skills/cli-extension/SKILL.md +1 -1
  40. package/agent-template/.ai/skills/deploy-operate/SKILL.md +119 -0
  41. package/agent-template/.ai/skills/module-new/SKILL.md +29 -3
  42. package/agent-template/.ai/skills/module-update/SKILL.md +9 -3
  43. package/agent-template/.ai/skills/perf-audit/SKILL.md +0 -1
  44. package/agent-template/.ai/skills/release-eject-pr/SKILL.md +0 -1
  45. package/agent-template/.ai/skills/spec-interview/SKILL.md +120 -0
  46. package/agent-template/.ai/skills/test-hardening/SKILL.md +1 -0
  47. package/agent-template/.ai/skills/ux-design/SKILL.md +34 -3
  48. package/agent-template/.ai/skills/variables/SKILL.md +0 -2
  49. package/agent-template/.ai/skills/workflow-development/SKILL.md +0 -1
  50. package/agent-template/.claude/skills/agent-tool-design/SKILL.md +1 -1
  51. package/agent-template/.claude/skills/auth-security-review/SKILL.md +1 -1
  52. package/agent-template/.claude/skills/cli-extension/SKILL.md +1 -1
  53. package/agent-template/.claude/skills/deploy-operate/SKILL.md +114 -0
  54. package/agent-template/.claude/skills/module-new/SKILL.md +29 -0
  55. package/agent-template/.claude/skills/module-update/SKILL.md +9 -1
  56. package/agent-template/.claude/skills/spec-interview/SKILL.md +114 -0
  57. package/agent-template/.claude/skills/ux-design/SKILL.md +34 -3
  58. package/agent-template/AGENTS.md +4 -0
  59. package/agent-template/CLAUDE.md +4 -0
  60. package/agent-template/docs/adr/0003-module-settings.md +1 -1
  61. package/agent-template/docs/adr/0006-agentic-workflows.md +24 -21
  62. package/agent-template/docs/agent-contract.md +2 -2
  63. package/agent-template/docs/cli-extensions.md +82 -0
  64. package/agent-template/docs/cli.md +195 -0
  65. package/agent-template/docs/configuration.md +593 -36
  66. package/agent-template/docs/design-system.md +185 -31
  67. package/agent-template/docs/getting-started.md +118 -0
  68. package/agent-template/docs/module-distribution.md +96 -0
  69. package/agent-template/docs/module-web-surfaces.md +221 -0
  70. package/agent-template/docs/modules.md +216 -0
  71. package/agent-template/docs/operations.md +545 -0
  72. package/agent-template/docs/sandbox.md +212 -0
  73. package/agent-template/platform/scripts/build.mjs +11 -0
  74. package/dist/bin.js +29 -0
  75. package/package.json +1 -1
  76. package/template/default/.dockerignore +14 -0
  77. package/template/default/.env.example +96 -0
  78. package/template/default/README.md +37 -1
  79. package/template/default/flowdular.json +15 -4
  80. package/template/default/infra/README.md +116 -0
  81. package/template/default/infra/docker/Dockerfile +37 -0
  82. package/template/default/infra/docker/compose.yaml +158 -0
  83. package/template/default/infra/docker/postgres/10-roles.sh +31 -0
  84. package/template/default/infra/docker/postgres/tls-init.sh +28 -0
  85. package/template/default/infra/kubernetes/database-secret.example.yaml +15 -0
  86. package/template/default/infra/kubernetes/deployment.yaml +211 -0
  87. package/template/default/infra/kubernetes/kustomization.yaml +9 -0
  88. package/template/default/infra/kubernetes/secrets.example.yaml +52 -0
  89. package/template/default/infra/kubernetes/service.yaml +13 -0
  90. package/template/default/modules/example/module.json +2 -1
  91. package/template/default/modules/example/package.json +1 -1
  92. package/template/default/modules/example/spec/module.yaml +1 -1
  93. package/template/default/modules/example/src/services/database-repository.ts +2 -12
  94. package/template/default/package.json +3 -2
  95. package/template/default/platform/octane.config.ts +99 -9
  96. package/template/default/platform/package.json +1 -1
  97. package/template/default/platform/src/generated/modules.client.ts +26 -2
  98. package/template/default/platform/src/generated/modules.server.ts +241 -10
  99. package/template/default/platform/src/server/health.ts +47 -0
  100. package/template/default/platform/src/server/metrics.ts +100 -0
  101. package/template/default/platform/src/server/storage.ts +172 -0
  102. package/template/default/platform/src/server/tracing.ts +85 -0
  103. package/template/default/specs/application.yaml +15 -0
@@ -0,0 +1,545 @@
1
+ # Operations
2
+
3
+ The runbook for a deployed Flowdular application: backing the data up, getting
4
+ it back, what key rotation does today, what to check before a release, and how
5
+ to roll one back. Environment variables are documented in
6
+ [configuration.md](configuration.md); the deployment itself in
7
+ [../infra/README.md](../infra/README.md).
8
+
9
+ ## Backup
10
+
11
+ ```bash
12
+ flowdular database backup --output /backups/flowdular/2026-09-11 # plan
13
+ flowdular database backup --output /backups/flowdular/2026-09-11 --apply # write
14
+ ```
15
+
16
+ Like every write command it is a dry run without `--apply`: it prints the
17
+ adapter, the source it would read, the files it would write and the manifest,
18
+ and touches nothing. The target directory must not already hold a `backup.json`,
19
+ so a backup never overwrites another one.
20
+
21
+ A backup directory holds two things:
22
+
23
+ | File | Contents |
24
+ | ---------------------------- | ----------------------------------------------------------------------------- |
25
+ | `backup.json` | Timestamp, adapter, platform version, enabled modules, key fingerprints |
26
+ | `database.dump` or `pglite/` | The `pg_dump` custom-format archive, or a copy of the embedded data directory |
27
+
28
+ The command creates the directory with mode `0700` and writes `backup.json` and
29
+ `database.dump` with `0600`, so a dump is not readable by every account on the
30
+ host. A directory that already exists keeps the mode it was made with: when the
31
+ backup target is a mount or a path an earlier run created, check its mode
32
+ yourself. The same goes for `pglite/`, which is copied file by file and keeps
33
+ the modes of the source data directory.
34
+
35
+ **PostgreSQL.** The command runs `pg_dump --format=custom` with the migrator
36
+ connection of the configured environment (`FD_DATABASE_MIGRATOR_URL`, else
37
+ `FD_DATABASE_URL`). The credentials reach the child process as `PG*` variables,
38
+ never as arguments, so no password lands in the host's process list, and the
39
+ child inherits none of the platform's encryption keys. `FD_DATABASE_TLS` and
40
+ `FD_DATABASE_TLS_CA[_FILE]` are translated into `PGSSLMODE` and `PGSSLROOTCERT`,
41
+ so a `verify-full` deployment is dumped over a verified connection. Without the
42
+ client tools on `PATH` the command fails with `BACKUP_TOOL_MISSING`; install the
43
+ PostgreSQL client package matching the server version.
44
+
45
+ **Embedded PGlite.** Outside production the database is a directory
46
+ (`FD_DATABASE_PGLITE_DIRECTORY`, by default `.flowdular/data/pglite`), and the
47
+ command copies it. Stop the application first. A copy taken while the process is
48
+ writing can hold a torn state, and the command says so in a warning.
49
+
50
+ ### What a backup does not contain
51
+
52
+ > **The encryption keys live outside the database.** Store every key with the
53
+ > backup. A restore without them yields unreadable credentials and workflow
54
+ > payloads: the rows come back, the values inside them stay ciphertext forever.
55
+
56
+ `backup.json` records, for each key, a SHA-256 fingerprint of the key material
57
+ and never the material itself:
58
+
59
+ ```json
60
+ {
61
+ "schemaVersion": 1,
62
+ "createdAt": "2026-09-11T09:30:00.000Z",
63
+ "adapter": "postgresql",
64
+ "platformVersion": "0.2.0",
65
+ "modules": ["agents.core", "auth.core", "workflows.core"],
66
+ "keys": [
67
+ {
68
+ "variable": "FD_AGENT_CREDENTIAL_KEY",
69
+ "fingerprint": "sha256:0f1e..."
70
+ },
71
+ { "variable": "FD_AUTH_MFA_KEY", "fingerprint": null }
72
+ ]
73
+ }
74
+ ```
75
+
76
+ The ten keys are `FD_AGENT_CREDENTIAL_KEY`, `FD_AGENT_RUN_GRANT_KEY`,
77
+ `FD_WORKFLOWS_PAYLOAD_KEY`, `FD_WORKFLOWS_CURSOR_KEY`, `FD_AUTH_MFA_KEY`,
78
+ `FD_AUTOMATIONS_CREDENTIAL_KEY`, `FD_NOTIFICATIONS_SECRET_KEY`,
79
+ `FD_STORAGE_ENCRYPTION_KEY`, `FD_CONNECTORS_SECRET_KEY` and
80
+ `FD_AUDIT_ANCHOR_KEY`. A `null` fingerprint means the key was not set
81
+ when the backup was taken. Keep the key material in the secret store the
82
+ deployment already uses, and record which backup it belongs to; the fingerprint
83
+ is what lets a restore tell you whether the pair matches.
84
+
85
+ Without `FD_AUDIT_ANCHOR_KEY` a restored deployment cannot verify the anchors
86
+ it holds, and every sealed segment file becomes a file whose origin nobody can
87
+ prove, so treat its fingerprint like the others.
88
+
89
+ The per-subject audit keys live in `audit_subject_keys` inside the database, so
90
+ a database backup carries them. A backup taken before
91
+ `flowdular audit erase --destroy-key` still holds the key that run destroyed, so
92
+ restoring it makes those sealed fields readable again. Treat an erasure as a
93
+ reason to age out the backups that predate it, as you would for any other
94
+ deletion on request. The key row itself survives destruction as a tombstone
95
+ carrying the subject's marker, which is what stops a later event about that
96
+ subject from creating a second key and putting the account back in the clear.
97
+
98
+ ### Schedule
99
+
100
+ - Dump before every rollout, and on a daily schedule between them.
101
+ - Keep the archive off the database host, and keep the keys in the secret store,
102
+ not next to the dump.
103
+ - Keep one restore rehearsal in the calendar. A backup nobody has restored is a
104
+ guess; restoring into a scratch database is the only proof.
105
+ - Retention follows the data policy of the business, not this document. Delete
106
+ old keys only after the dumps they unlock are gone.
107
+
108
+ ## Restore
109
+
110
+ ```bash
111
+ flowdular database restore --input /backups/flowdular/2026-09-11 # plan
112
+ flowdular database restore --input /backups/flowdular/2026-09-11 --apply --confirm restore-database
113
+ ```
114
+
115
+ The plan reports the manifest, the payload it would read and one key comparison
116
+ per variable. When the running environment holds keys the backup was not taken
117
+ with, the envelope carries a `BACKUP_KEY_MISMATCH` warning naming each variable
118
+ and how it differs (`different`, `missing-in-environment`, `missing-in-backup`).
119
+ Restore the keys before the data, or the rows arrive unreadable.
120
+
121
+ The order that works:
122
+
123
+ 1. Stop the application, or take it out of the load balancer.
124
+ 2. Put the keys the backup was taken with into the environment.
125
+ 3. Run the restore as a dry run and read the key comparison.
126
+ 4. Run it with `--apply --confirm restore-database`.
127
+ 5. Start the application, then check `pnpm flowdular migration verify` and
128
+ `GET /api/ready`.
129
+
130
+ PostgreSQL restores run `pg_restore --clean --if-exists` against the migrator
131
+ connection, so the target database must already exist and its roles must already
132
+ be there. The embedded adapter is replaced directory by directory: the copy is
133
+ staged next to the target and the old directory is removed only once the new one
134
+ is in place, so a failed copy leaves the running database untouched.
135
+
136
+ `database restore` is a destructive capability, gated exactly like
137
+ `database reset`: the runner refuses it with `LOCAL_ONLY_CAPABILITY` unless
138
+ `FD_ENV` or `NODE_ENV` is `development` or `test`. The dry run on a staging
139
+ copy of the same backup is what proves the archive and the keys are good.
140
+
141
+ ### Restore in production
142
+
143
+ ```bash
144
+ flowdular database restore-production --input <dir> --target flowdular --grant <token> --tenant <id> # plan
145
+ flowdular database restore-production --input <dir> --target flowdular --grant <token> --tenant <id> \
146
+ --platform-url https://erp.example.com --apply --confirm restore-database
147
+ ```
148
+
149
+ `database.restore.production` is the same restore without the local gate. It
150
+ runs only under an approval grant from an approved `approvals.core` request,
151
+ verified against `FD_APPROVAL_GRANT_KEY`, the tenant, the capability and the
152
+ exact flags: the request has to name `--input`, `--target`, `--apply` and any
153
+ override (an approval names the applied run, and a dry run is a different
154
+ invocation), and a token issued for another invocation is refused with
155
+ `APPROVAL_GRANT_MISMATCH`. Open the request with `approvals.requests.v1` and a
156
+ `subjectRef` of `capability:<id>:<sha256 of the invocation input>`, read the
157
+ approved request back through `grant(tenantId, id, subjectModule)` for the
158
+ token, and pass it as `--grant <token> --tenant <id>`. `--target` must repeat
159
+ the database the migrator DSN names (`RESTORE_TARGET_MISMATCH` otherwise),
160
+ `FD_DATABASE_MIGRATOR_URL` must be set and distinct from `FD_DATABASE_URL`
161
+ (`MIGRATOR_ROLE_REQUIRED`), and a key divergence is a refusal
162
+ (`BACKUP_KEY_MISMATCH`) unless `--allow-key-mismatch` was part of the approved
163
+ invocation. With `--apply` the command probes `<--platform-url>/api/health`,
164
+ or `http://127.0.0.1:$FD_PORT/api/health`, and refuses with
165
+ `PLATFORM_RUNNING` while anything answers; when no endpoint is named or the
166
+ probe times out it refuses with `PLATFORM_STATE_UNKNOWN` unless
167
+ `--platform-stopped` attests the platform is down. Grants are HMAC-SHA256
168
+ under `FD_APPROVAL_GRANT_KEY` (rotate with `FD_APPROVAL_GRANT_KEY_PREVIOUS`)
169
+ and expire one hour after the approving decision; the runner records no use,
170
+ so a grant replays until it expires and the approval window should stay short.
171
+ `database reset` and `database restore` keep their local-only gate.
172
+
173
+ ### Point-in-time recovery
174
+
175
+ A dump restores one moment; the compose stack also archives WAL so any moment
176
+ after a base backup can be recovered. See `infra/README.md`, "Backups and
177
+ PITR": `infra/docker/pitr.sh base-backup` after every rollout,
178
+ `pitr.sh restore --base <stamp> --target-time '<ts>' --confirm replace-cluster`
179
+ to recover, and the limits of the plain-copy archive. In Kubernetes the
180
+ managed provider owns PITR; the dump remains the portable copy.
181
+
182
+ ### Rehearsal
183
+
184
+ Quarterly, on a scratch host: restore the newest dump into an empty database
185
+ with `database restore-production --platform-stopped`, read the key
186
+ comparison, start the app, check `migration verify` and `/api/ready`, then run
187
+ `pitr.sh restore` from the newest base backup to a time between two dumps and
188
+ confirm a row written after the dump is present. Record the date and the
189
+ elapsed time.
190
+
191
+ ## Key rotation
192
+
193
+ Most of the keys carry a key id in every envelope they write, so a
194
+ deployment can hold the new key and the old one at the same time and re-seal the
195
+ stored rows in place. The others have no key id and are covered at the end of
196
+ this section.
197
+
198
+ Every module reads one current key plus an optional comma-separated list of
199
+ retired keys (up to eight):
200
+
201
+ | Key | Protects | Retired keys | Re-sealing the stored rows |
202
+ | ------------------------------- | -------------------------- | ---------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------- |
203
+ | `FD_AGENT_CREDENTIAL_KEY` | Agent provider credentials | `FD_AGENT_CREDENTIAL_KEY_PREVIOUS` | `pnpm flowdular agents secrets-rotate [--apply]` |
204
+ | `FD_AUTOMATIONS_CREDENTIAL_KEY` | Automation trigger secrets | `FD_AUTOMATIONS_CREDENTIAL_KEY_PREVIOUS` | `pnpm flowdular automations secrets-rotate [--apply]` |
205
+ | `FD_WORKFLOWS_PAYLOAD_KEY` | Workflow run payloads | `FD_WORKFLOWS_PAYLOAD_KEY_PREVIOUS` | `pnpm flowdular workflows secrets-rotate [--apply]` |
206
+ | `FD_STORAGE_ENCRYPTION_KEY` | Stored objects | `FD_STORAGE_ENCRYPTION_KEY_PREVIOUS` | `pnpm flowdular documents secrets-rotate [--apply]` and `pnpm flowdular exports secrets-rotate [--apply]`; see the storage note below |
207
+ | `FD_NOTIFICATIONS_SECRET_KEY` | Webhook signing secrets | `FD_NOTIFICATIONS_SECRET_KEY_PREVIOUS` | `pnpm flowdular notifications secrets-rotate [--apply]` |
208
+ | `FD_CONNECTORS_SECRET_KEY` | Connector credentials | `FD_CONNECTORS_SECRET_KEY_PREVIOUS` | `pnpm flowdular connectors secrets-rotate [--apply]` |
209
+ | `FD_AUDIT_ANCHOR_KEY` | Audit chain anchors (HMAC) | `FD_AUDIT_ANCHOR_KEY_PREVIOUS` | `pnpm flowdular audit secrets-rotate [--apply]` |
210
+
211
+ These commands run the same re-sealing pass over their own table, so the
212
+ procedure is the same for each. The credential key is the worked example;
213
+ substitute the key names and the command from the row you are rotating:
214
+
215
+ 1. Generate the new key: `openssl rand -base64 32`.
216
+ 2. Set the new key as the current one and the old key as the retired one:
217
+ `FD_AGENT_CREDENTIAL_KEY=<new>` and `FD_AGENT_CREDENTIAL_KEY_PREVIOUS=<old>`.
218
+ Keep both in the secret store, exactly like the key they replace.
219
+ 3. Deploy. From this moment every write uses the new key and every stored row
220
+ still opens under the old one, so nothing is unreadable while the rollout
221
+ completes.
222
+ 4. Run the rotation as a dry run and read the counts per key id:
223
+ `pnpm flowdular agents secrets-rotate --json`. `stale` is the number of rows
224
+ still sealed with a retired key.
225
+ 5. Run it again with `--apply`. It re-seals rows in batches of 200, one
226
+ transaction per batch, scoped to the tenant that owns them, and writes nothing
227
+ else: the revision, the audit fields and the credential a caller sees do not
228
+ change. It is idempotent, so running it twice is free.
229
+ 6. Repeat the dry run until it reports `stale: 0`, then remove
230
+ `FD_AGENT_CREDENTIAL_KEY_PREVIOUS` and deploy again.
231
+
232
+ The application may stay up. Each row is written back only if its stored
233
+ envelope is still the one that was read, so a credential the application
234
+ rewrites in between is never clobbered: the rotation counts it under `skipped`
235
+ and the next run picks it up. Stopping the application removes that case
236
+ entirely. The inventory that finds the stale rows runs on the cross-tenant
237
+ read-only role, which is granted the tenant id and the key id and cannot read a
238
+ nonce, a tag or a ciphertext; every row it names is read again under its own
239
+ tenant before it is touched.
240
+
241
+ `FD_AUDIT_ANCHOR_KEY` follows the same six steps with one difference: an anchor
242
+ is signed, not sealed, so the rotation rewrites the signature and the key id and
243
+ touches neither the anchor hash, the segment hash nor the chain. A segment file
244
+ written under the retired key keeps its own signature and still verifies while
245
+ that key is in `FD_AUDIT_ANCHOR_KEY_PREVIOUS`. Keep the retired key until every
246
+ archived segment has been verified once under the new one, with
247
+ `pnpm flowdular audit verify --workspace <slug> --input <dir>` over the archive.
248
+ If the rotation reports `ANCHOR_KEY_UNKNOWN`, anchors exist that neither key can
249
+ re-sign: put that key back before retiring it, because a signature nobody can
250
+ reproduce can never be verified again. The rotation verifies an anchor's stored
251
+ signature under the ring before re-signing it, so `ANCHOR_SIGNATURE_INVALID`
252
+ names anchors it left exactly as they are: a new signature over the same anchor
253
+ hash would replace evidence nobody can account for with one this deployment
254
+ vouches for. Verify the segment files of those anchors before anything else.
255
+ `FD_AUDIT_ANCHOR_KEY_PREVIOUS` accepts at most eight retired keys; a longer list
256
+ is a rotation that was never finished and the platform refuses to start.
257
+
258
+ The remaining keys, and how rotation behaves for each:
259
+
260
+ | Key | Protects | Rotating it |
261
+ | ------------------------- | -------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
262
+ | `FD_WORKFLOWS_CURSOR_KEY` | Pagination cursors (HMAC) | Safe. `FD_WORKFLOWS_CURSOR_KEY_PREVIOUS` keeps outstanding cursors valid; cursors are never stored, so there is nothing to re-sign. |
263
+ | `FD_AGENT_RUN_GRANT_KEY` | Agent run grant tokens (HMAC, 30s TTL) | Safe. In-flight grants fail for the length of one TTL. |
264
+ | `FD_AUTH_MFA_KEY` | Enrolled TOTP secrets | Supported since auth.core 0.11.0. Set `FD_AUTH_MFA_KEY_PREVIOUS`, deploy, run `pnpm flowdular auth secrets-rotate --apply`, then drop the previous key. Rows written before the key id column exist open through the ring fallback until they are re-sealed. |
265
+
266
+ The key ids are derived from the key material, so a backup fingerprint that
267
+ matches the environment also matches the `keyId` inside the rows. The keys that
268
+ carry no id are the ones a restore cannot diagnose for you, which is why
269
+ `backup.json` records a fingerprint for all eight.
270
+
271
+ Sources: `modules/audit/src/services/{anchor-key,anchor-rotation}.ts`,
272
+ `packages/kernel/src/keyring.ts`,
273
+ `modules/agents/src/services/{credential-vault,credential-rotation}.ts`,
274
+ `modules/automations/src/services/{secret-vault,secret-rotation}.ts`,
275
+ `modules/workflows/src/services/{payload-codec,payload-rotation,cursors}.ts`,
276
+ `modules/notifications/src/services/{secret-vault,secret-rotation}.ts`,
277
+ `modules/connectors/src/services/{credential-vault,credential-rotation}.ts`,
278
+ `packages/storage/src/reseal.ts`,
279
+ `modules/documents/src/services/storage-rotation.ts`,
280
+ `modules/exports/src/services/storage-rotation.ts`,
281
+ `modules/auth/src/services/totp.ts`.
282
+
283
+ ## Storage
284
+
285
+ Objects live outside PostgreSQL, so a database backup does not contain one. Back
286
+ up the object store the same way you back up the database, and take both at the
287
+ same point: a restored database that references objects a newer bucket snapshot
288
+ no longer holds leaves dangling references, and the reverse leaves orphans.
289
+
290
+ - **S3-compatible store.** Enable versioning and a lifecycle policy on the
291
+ bucket, or replicate it. `backup.json` does not reach into it.
292
+ - **Local adapter.** Development and test only, refused in production. Its
293
+ directory is `FD_STORAGE_LOCAL_DIRECTORY`, by default
294
+ `.flowdular/data/storage`; stop the application before copying it.
295
+
296
+ Every object is sealed with AES-256-GCM under `FD_STORAGE_ENCRYPTION_KEY`, and
297
+ the key id is stored with the object, so the rotation follows the six steps
298
+ above with two commands instead of one, because the object store has no
299
+ listing and the rows that name the objects belong to two modules:
300
+
301
+ 1. Generate the new key: `openssl rand -base64 32`.
302
+ 2. Set `FD_STORAGE_ENCRYPTION_KEY=<new>` and
303
+ `FD_STORAGE_ENCRYPTION_KEY_PREVIOUS=<old>` in the secret store.
304
+ 3. Deploy. Every new object is sealed with the new key and every stored one
305
+ still opens under the old one.
306
+ 4. Dry run both passes and read the counts per key id:
307
+ `pnpm flowdular documents secrets-rotate --json` walks the stored document
308
+ rows, `pnpm flowdular exports secrets-rotate --json` the completed export
309
+ jobs. `stale` is the number of objects still sealed with a retired key;
310
+ `unknown` names objects under a key neither variable holds, and the pass
311
+ leaves those alone.
312
+ 5. Run both again with `--apply`. Each object is re-sealed in place: the frame
313
+ is opened under the key it names, which authenticates its header, and the
314
+ same content type, size, checksum, scan verdict and creation time are
315
+ written back under the current key. Nothing else changes, not the object
316
+ key and not the row. A frame that fails authentication is counted under
317
+ `refused` and left as it is; restore it from the object store backup.
318
+ 6. Repeat the dry runs until both report `stale: 0`, then remove
319
+ `FD_STORAGE_ENCRYPTION_KEY_PREVIOUS` and deploy again.
320
+
321
+ The application may stay up: the pass locks a row while it rewrites the
322
+ object, and both delete paths (a document removal and the export retention
323
+ sweep) lock the row before they remove the object and then the row, so a delete
324
+ that lands mid-pass waits for the rewrite and removes the re-sealed object
325
+ rather than racing it. A frame that fails to parse is counted under `refused`
326
+ like one that fails authentication. Do not snapshot the object store and the database apart while a
327
+ pass runs: a bucket snapshot taken mid-pass holds objects under both keys, and
328
+ restoring it beside a database from another moment leaves rows that name
329
+ objects a ring without the retired key cannot open. Take both after the pass,
330
+ or both before it. Dropping the retired key while `stale` is above zero makes
331
+ those objects unreadable, exactly as a lost database key makes a credential
332
+ unreadable.
333
+
334
+ ## Logs
335
+
336
+ The server writes one line per event through `createLogger`
337
+ (`packages/server/src/log.ts`).
338
+
339
+ | Variable | Default | Purpose |
340
+ | --------------- | --------------------------------- | ------------------------------------------------ |
341
+ | `FD_LOG_FORMAT` | `json` in production, else `text` | `json` for a log pipeline, `text` for a terminal |
342
+ | `FD_LOG_LEVEL` | `info` | `debug`, `info`, `warn` or `error` |
343
+
344
+ A JSON line carries `time`, `level`, `msg` and, where they apply, `traceId`,
345
+ `spanId`, `requestId`, `endpoint`, `module` and `err` (name and message). The
346
+ trace ids are on every line written inside a served request, so the same id
347
+ finds the log line and the trace. Stacks are written outside
348
+ production, and in production only at `FD_LOG_LEVEL=debug`. The `requestId` is
349
+ the same value the response returns in `x-request-id`, so a user-reported error
350
+ id finds its log line. Credential-shaped field names (`authorization`, `cookie`,
351
+ `password`, `token`, `secret`, `credential`, API keys) are redacted before a
352
+ line is written, and a failing endpoint logs its id and the error type only,
353
+ never the message, the body or the URL. An unrecognized value for either
354
+ variable falls back to the default and logs one warning about it.
355
+
356
+ ## Production checklist
357
+
358
+ - `pnpm verify` and `pnpm build` pass on the commit being shipped.
359
+ - `FD_ENV=production`, `FD_DATABASE_ADAPTER=postgresql`, separate
360
+ `FD_DATABASE_URL` and `FD_DATABASE_MIGRATOR_URL` roles, neither `SUPERUSER`
361
+ nor `BYPASSRLS` on the runtime role.
362
+ - `FD_DATABASE_TLS=verify-full` with `FD_DATABASE_TLS_CA` or
363
+ `FD_DATABASE_TLS_CA_FILE`. Production accepts nothing weaker.
364
+ - All six encryption keys set, held in the secret store, and backed up with a
365
+ recorded fingerprint.
366
+ - `FD_AUTH_SECURE_COOKIE` on, `FD_AUTH_PUBLIC_ORIGIN` equal to the public URL,
367
+ `FD_TRUST_PROXY` set behind a proxy, sign-up closed unless the deployment is
368
+ public.
369
+ - Mail transport configured: `FD_AUTH_MAIL_TRANSPORT` with `FD_AUTH_SMTP_URL`
370
+ and `FD_AUTH_MAIL_FROM`. Password resets, invitations and address
371
+ confirmations are silently undeliverable without it.
372
+ - Migrations at rollout: deploy the image, let it apply its module migrations
373
+ under the migrator role, then `pnpm flowdular migration verify`. Never edit an
374
+ applied `.up.sql`.
375
+ - Liveness probe on `GET /api/health`, readiness probe on `GET /api/ready`.
376
+ - `FD_LOG_FORMAT=json` and a log pipeline that keeps `requestId`.
377
+ - A fresh backup, taken and verified before the rollout starts.
378
+
379
+ ## Rollback
380
+
381
+ 1. Redeploy the previous image tag. Module migrations are additive, so an older
382
+ image runs against a newer schema and ignores columns it does not know.
383
+ 2. Never apply a `.down.sql` to roll back a release. They document the reverse
384
+ for review; against live data they are data loss.
385
+ 3. To take one module out of service, `pnpm flowdular module disable <id> --apply`,
386
+ rebuild the composition and redeploy. Its tables stay where they are.
387
+ 4. If the rollback is about a key, restore the key first. A previous image does
388
+ not make ciphertext readable again.
389
+ 5. Restore from a backup only when the schema or the data is actually damaged.
390
+ It is the slowest option and the only one that loses everything written since
391
+ the dump.
392
+
393
+ ## Metrics
394
+
395
+ `GET /api/metrics` serves Prometheus text exposition, format version 0.0.4. The
396
+ route is composed only when `FD_METRICS=true`; with `FD_METRICS_TOKEN` set, a
397
+ scrape must send `Authorization: Bearer <token>` and gets 401 without it.
398
+
399
+ **An untokened `/api/metrics` must never be reachable outside the scrape
400
+ network.** It shares the application port, so without `FD_METRICS_TOKEN` every
401
+ client that can reach the port reads the series, including the endpoint
402
+ inventory and the request rates of the whole deployment. Either set the token
403
+ (the Kubernetes manifests read it from the deployment Secret, optional so it
404
+ can be left unset deliberately) or keep the port on a network only the scraper
405
+ reaches. Do not rely on the route being obscure.
406
+
407
+ | Series | Type | Labels |
408
+ | ----------------------------------------- | --------- | ----------------------------------- |
409
+ | `flowdular_http_requests_total` | counter | `endpoint`, `method`, `status` |
410
+ | `flowdular_http_request_duration_seconds` | histogram | `endpoint`, buckets 0.005 s to 10 s |
411
+ | `flowdular_metrics_dropped_samples_total` | counter | none |
412
+ | `flowdular_process_start_time_seconds` | gauge | none |
413
+ | `process_resident_memory_bytes` | gauge | none |
414
+ | `nodejs_eventloop_lag_seconds` | gauge | none |
415
+ | `flowdular_build_info` | gauge | `version` |
416
+
417
+ `endpoint` is the endpoint id, never the request path, so a record id cannot
418
+ reach a label. `status` is the class (`2xx`, `4xx`, `5xx`), and an unrecognized
419
+ method is recorded as `other`. No session, tenant or user value is ever a label.
420
+
421
+ `nodejs_eventloop_lag_seconds` is the mean delay over the interval since the
422
+ previous scrape, not since the process started, so a spike shows up in the
423
+ scrape that follows it and is gone from the next one. The sampler is
424
+ process-wide and is drained on every read, so two scrapers pointed at the same
425
+ process split the samples between them; point one scraper at it.
426
+
427
+ Every request through `defineEndpoint` is recorded, whether or not the route is
428
+ exposed, so a scraper enabled later sees the counters immediately. Memory is
429
+ bounded: the process keeps at most 2000 distinct label sets per series family
430
+ and refuses new ones after that, counting each refusal in
431
+ `flowdular_metrics_dropped_samples_total`. A non-zero value there means a
432
+ deployment composes more endpoints than the ceiling allows.
433
+
434
+ A module adds series of its own with `createModuleMetrics('<module id>')`:
435
+
436
+ | Series | Type | Labels |
437
+ | ----------------------------------------------------------------- | --------- | --------------------------- |
438
+ | `flowdular_module_<module>_<name>_total` | counter | the module's own |
439
+ | `flowdular_module_<module>_<name>` (+`_bucket`, `_sum`, `_count`) | histogram | the module's own, plus `le` |
440
+
441
+ The module id is lower-cased with every character outside `a-z0-9_` replaced by
442
+ `_`, so `agents.core` reads `flowdular_module_agents_core_`. A module series is
443
+ refused, and counted as a dropped sample, when the metric name does not match
444
+ `^[a-z][a-z0-9_]{0,63}$`, when a call carries more than 8 labels or a label name
445
+ of the wrong shape, when the family already holds 2000 label sets, or when the
446
+ process already opened 256 module families of that kind. A label value is cut at
447
+ 128 characters. A module never receives an error from recording: a refused
448
+ sample is lost, the work is not.
449
+
450
+ **The label rule of the request series holds for a module series too: never a
451
+ tenant, account, record or session id.** The platform cannot tell one from a
452
+ closed set, so nothing refuses it; what it does is cap the damage at 2000 label
453
+ sets per family, after which the series stops admitting new ones. A label is for
454
+ a value with a handful of possibilities, such as an outcome or a channel.
455
+
456
+ ## Tracing
457
+
458
+ Every request through `defineEndpoint` runs in a W3C trace. The inbound
459
+ `traceparent` names the parent when it is usable; anything else is a new root.
460
+ The response always returns a `traceparent` naming the span the request ran in,
461
+ so a client, a proxy and this process agree on one trace id, and the log lines
462
+ of that request carry the same `traceId`.
463
+
464
+ Spans go into one bounded in-process ring whether or not an exporter is
465
+ composed: 4096 spans, oldest evicted and counted, at most 16 attributes per span
466
+ and 256 characters per name, attribute key and string value. Attributes carry
467
+ endpoint ids, job names, provider and tool ids, HTTP method and status, and the
468
+ three ids that correlate a span with a log line or an agent run,
469
+ `flowdular.request_id`, `flowdular.agent.run_id` and
470
+ `flowdular.tool.provider_call_id`: no request path, body, tenant or session.
471
+
472
+ | Variable | Purpose |
473
+ | ----------------------- | --------------------------------------------------- |
474
+ | `FD_TRACE_SAMPLE` | Ratio of new roots recorded, 0 to 1, default 1 |
475
+ | `FD_TRACE_EXPORTER` | `none` (default) or `otlp` |
476
+ | `FD_TRACE_OTLP_URL` | OTLP/HTTP JSON traces endpoint; https in production |
477
+ | `FD_TRACE_OTLP_HEADERS` | `name=value,name2=value2`, at most 16 |
478
+
479
+ The exporter posts OTLP/HTTP JSON (no SDK, no agent) in batches of 512 spans, or
480
+ every 5 seconds when the batch does not fill. It never runs on a request path: a
481
+ collector that is down costs the process one held batch, retried on the next two
482
+ flushes and then dropped. A `4xx` that is not `429` is not retried, because the
483
+ payload or the credential is wrong and the next batch would be lost too. The
484
+ sampling decision is taken on the trace id, so a whole trace is in or out; a
485
+ trace the caller already marked sampled is never re-decided, and an unsampled
486
+ trace still propagates its header.
487
+
488
+ A process that stops drains what the ring still holds, in the same batches and
489
+ without a retry: the spans of a shutdown leave with it, and a collector that is
490
+ down is given up on rather than held between the process and its exit.
491
+
492
+ **An OTLP endpoint receives the endpoint inventory, the job names, the timings
493
+ and the request, agent run and provider call ids of the whole deployment.**
494
+ Point it at a collector you control. Plain `http` is refused in production.
495
+
496
+ ## Error reporting
497
+
498
+ `FD_ERROR_SINK=webhook` posts every line the server logs at `error` to
499
+ `FD_ERROR_SINK_URL`, with `Authorization: Bearer <FD_ERROR_SINK_TOKEN>` when the
500
+ token is set. The body is `{"reports":[…]}` where a report carries `at`, `name`,
501
+ `message` and, where they apply, `endpoint`, `module`, `requestId`, `traceId`
502
+ and `spanId`. Nothing else: no stack, no log fields, no request body, so the
503
+ sink cannot carry a credential the logger already refuses to write.
504
+
505
+ `message` is the one field the caller chose. A failing endpoint logs its error
506
+ type only, never the message, so its report carries `endpoint failed`; a
507
+ background loop logs the error it caught, so its report carries that message.
508
+ The sink therefore exports exactly what the log line already shows, to a host
509
+ outside the deployment. A deployment that must not let error text leave keeps
510
+ `FD_ERROR_SINK=none` and reads the log stream instead.
511
+
512
+ It is bounded and lossy on purpose: 64 queued reports (a report past that is
513
+ dropped and counted), 32 per request, 8 KiB per body with the batch halved until
514
+ it fits, a batch sent every 5 seconds or as soon as 32 are queued, and at most 3
515
+ attempts before a batch is given up. Delivery never runs on the caller's path
516
+ and an unreachable webhook never fails the line that reported the failure.
517
+ `none` is the default and holds no queue and no timer.
518
+
519
+ The OTLP logs sink named in RFC 0004 H7 is deliberately not built: the OTLP
520
+ surface this platform speaks is the trace exporter above.
521
+
522
+ ## Data residency
523
+
524
+ Residency is a property of the deployment, not of the workspace. A deployment
525
+ carries one PostgreSQL database, one object store and one region; every module
526
+ writes to that database and, through the storage port, to that bucket. A
527
+ workspace therefore lives wherever its deployment lives, and a customer whose
528
+ data must stay in a particular region gets a deployment in that region. That
529
+ is the supported answer and the only one the platform offers.
530
+
531
+ Per-workspace routing to regional databases is not planned. It would change
532
+ the provider contract in `packages/database` (one handle per region), the
533
+ lease model every module's migrations and background loops depend on, every
534
+ cross-workspace routing read the background role performs (the retention
535
+ sweep, delivery queues, expiry passes), the storage key layout, and the
536
+ export and erasure paths that assume one place to look. Record that cost
537
+ before anyone proposes it; a second deployment is cheaper in every case
538
+ seen so far.
539
+
540
+ What to state in a questionnaire: the region of the database and the bucket,
541
+ that backups and the key material are stored in the same region unless the
542
+ operator moves them, that the OTLP trace exporter and the error sink send
543
+ only what `docs/operations.md` lists for them and only to the URLs the
544
+ operator configured, and that outbound mail, webhooks and connector calls
545
+ leave the region by design of the recipient the workspace configured.