create-flowdular 0.2.6 → 0.3.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (100) hide show
  1. package/agent-template/.agents/skills/agent-tool-design/SKILL.md +1 -1
  2. package/agent-template/.agents/skills/auth-security-review/SKILL.md +1 -1
  3. package/agent-template/.agents/skills/deploy-operate/SKILL.md +109 -0
  4. package/agent-template/.agents/skills/module-new/SKILL.md +29 -0
  5. package/agent-template/.agents/skills/module-update/SKILL.md +9 -1
  6. package/agent-template/.agents/skills/spec-interview/SKILL.md +114 -0
  7. package/agent-template/.agents/skills/ux-design/SKILL.md +34 -3
  8. package/agent-template/.ai/README.md +2 -1
  9. package/agent-template/.ai/agents/sandbox/business-manager.md +5 -1
  10. package/agent-template/.ai/blueprints/author-spec/README.md +1 -1
  11. package/agent-template/.ai/blueprints/author-spec/spec-requirements.yaml +44 -0
  12. package/agent-template/.ai/blueprints/author-spec/steps.yaml +5 -5
  13. package/agent-template/.ai/blueprints/author-spec/templates/module.yaml +99 -3
  14. package/agent-template/.ai/blueprints/edit-module/gates.yaml +4 -0
  15. package/agent-template/.ai/blueprints/edit-module/required-files.yaml +9 -0
  16. package/agent-template/.ai/blueprints/new-module/gates.yaml +4 -0
  17. package/agent-template/.ai/blueprints/new-module/spec-requirements.yaml +2 -2
  18. package/agent-template/.ai/blueprints/release/gates.yaml +4 -0
  19. package/agent-template/.ai/platform-capabilities.md +128 -0
  20. package/agent-template/.ai/policies/capabilities.yaml +100 -0
  21. package/agent-template/.ai/policies/path-ownership.yaml +5 -2
  22. package/agent-template/.ai/policies/task-budgets.yaml +5 -3
  23. package/agent-template/.ai/references/catalog/module.json +4 -4
  24. package/agent-template/.ai/references/catalog/package.json +2 -2
  25. package/agent-template/.ai/references/catalog/spec/module.yaml +5 -3
  26. package/agent-template/.ai/references/catalog/src/platform.ts +2 -0
  27. package/agent-template/.ai/references/catalog/src/services/catalog-service.ts +89 -1
  28. package/agent-template/.ai/references/catalog/src/services/data-classes.ts +47 -0
  29. package/agent-template/.ai/references/catalog/src/services/database-repository.ts +98 -1
  30. package/agent-template/.ai/references/catalog/src/services/repository.ts +22 -1
  31. package/agent-template/.ai/references/catalog/tests/data-classes.test.ts +157 -0
  32. package/agent-template/.ai/references/catalog.provenance.json +12 -10
  33. package/agent-template/.ai/rules/flowdular.md +4 -0
  34. package/agent-template/.ai/skills/README.md +10 -0
  35. package/agent-template/.ai/skills/agent-tool-design/SKILL.md +1 -2
  36. package/agent-template/.ai/skills/auth-security-review/SKILL.md +1 -1
  37. package/agent-template/.ai/skills/business-agent-design/SKILL.md +0 -1
  38. package/agent-template/.ai/skills/deploy-operate/SKILL.md +114 -0
  39. package/agent-template/.ai/skills/module-new/SKILL.md +29 -3
  40. package/agent-template/.ai/skills/module-update/SKILL.md +9 -3
  41. package/agent-template/.ai/skills/perf-audit/SKILL.md +0 -1
  42. package/agent-template/.ai/skills/release-eject-pr/SKILL.md +0 -1
  43. package/agent-template/.ai/skills/spec-interview/SKILL.md +120 -0
  44. package/agent-template/.ai/skills/test-hardening/SKILL.md +1 -0
  45. package/agent-template/.ai/skills/ux-design/SKILL.md +34 -3
  46. package/agent-template/.ai/skills/variables/SKILL.md +0 -2
  47. package/agent-template/.ai/skills/workflow-development/SKILL.md +0 -1
  48. package/agent-template/.claude/skills/agent-tool-design/SKILL.md +1 -1
  49. package/agent-template/.claude/skills/auth-security-review/SKILL.md +1 -1
  50. package/agent-template/.claude/skills/deploy-operate/SKILL.md +109 -0
  51. package/agent-template/.claude/skills/module-new/SKILL.md +29 -0
  52. package/agent-template/.claude/skills/module-update/SKILL.md +9 -1
  53. package/agent-template/.claude/skills/spec-interview/SKILL.md +114 -0
  54. package/agent-template/.claude/skills/ux-design/SKILL.md +34 -3
  55. package/agent-template/AGENTS.md +4 -0
  56. package/agent-template/CLAUDE.md +4 -0
  57. package/agent-template/docs/adr/0003-module-settings.md +1 -1
  58. package/agent-template/docs/adr/0006-agentic-workflows.md +24 -21
  59. package/agent-template/docs/agent-contract.md +2 -2
  60. package/agent-template/docs/cli-extensions.md +82 -0
  61. package/agent-template/docs/cli.md +190 -0
  62. package/agent-template/docs/configuration.md +593 -36
  63. package/agent-template/docs/design-system.md +185 -31
  64. package/agent-template/docs/getting-started.md +118 -0
  65. package/agent-template/docs/module-distribution.md +96 -0
  66. package/agent-template/docs/module-web-surfaces.md +221 -0
  67. package/agent-template/docs/modules.md +216 -0
  68. package/agent-template/docs/operations.md +464 -0
  69. package/agent-template/docs/sandbox.md +212 -0
  70. package/agent-template/platform/scripts/build.mjs +10 -0
  71. package/dist/bin.js +28 -0
  72. package/package.json +1 -1
  73. package/template/default/.dockerignore +14 -0
  74. package/template/default/.env.example +94 -0
  75. package/template/default/README.md +37 -1
  76. package/template/default/flowdular.json +15 -4
  77. package/template/default/infra/README.md +116 -0
  78. package/template/default/infra/docker/Dockerfile +37 -0
  79. package/template/default/infra/docker/compose.yaml +158 -0
  80. package/template/default/infra/docker/postgres/10-roles.sh +31 -0
  81. package/template/default/infra/docker/postgres/tls-init.sh +28 -0
  82. package/template/default/infra/kubernetes/database-secret.example.yaml +15 -0
  83. package/template/default/infra/kubernetes/deployment.yaml +211 -0
  84. package/template/default/infra/kubernetes/kustomization.yaml +9 -0
  85. package/template/default/infra/kubernetes/secrets.example.yaml +52 -0
  86. package/template/default/infra/kubernetes/service.yaml +13 -0
  87. package/template/default/modules/example/module.json +2 -1
  88. package/template/default/modules/example/package.json +1 -1
  89. package/template/default/modules/example/spec/module.yaml +1 -1
  90. package/template/default/modules/example/src/services/database-repository.ts +2 -12
  91. package/template/default/package.json +3 -2
  92. package/template/default/platform/octane.config.ts +99 -9
  93. package/template/default/platform/package.json +1 -1
  94. package/template/default/platform/src/generated/modules.client.ts +26 -2
  95. package/template/default/platform/src/generated/modules.server.ts +241 -10
  96. package/template/default/platform/src/server/health.ts +47 -0
  97. package/template/default/platform/src/server/metrics.ts +100 -0
  98. package/template/default/platform/src/server/storage.ts +172 -0
  99. package/template/default/platform/src/server/tracing.ts +85 -0
  100. package/template/default/specs/application.yaml +15 -0
@@ -0,0 +1,464 @@
1
+ # Operations
2
+
3
+ The runbook for a deployed Flowdular application: backing the data up, getting
4
+ it back, what key rotation does today, what to check before a release, and how
5
+ to roll one back. Environment variables are documented in
6
+ [configuration.md](configuration.md); the deployment itself in
7
+ [../infra/README.md](../infra/README.md).
8
+
9
+ ## Backup
10
+
11
+ ```bash
12
+ flowdular database backup --output /backups/flowdular/2026-09-11 # plan
13
+ flowdular database backup --output /backups/flowdular/2026-09-11 --apply # write
14
+ ```
15
+
16
+ Like every write command it is a dry run without `--apply`: it prints the
17
+ adapter, the source it would read, the files it would write and the manifest,
18
+ and touches nothing. The target directory must not already hold a `backup.json`,
19
+ so a backup never overwrites another one.
20
+
21
+ A backup directory holds two things:
22
+
23
+ | File | Contents |
24
+ | ---------------------------- | ----------------------------------------------------------------------------- |
25
+ | `backup.json` | Timestamp, adapter, platform version, enabled modules, key fingerprints |
26
+ | `database.dump` or `pglite/` | The `pg_dump` custom-format archive, or a copy of the embedded data directory |
27
+
28
+ The command creates the directory with mode `0700` and writes `backup.json` and
29
+ `database.dump` with `0600`, so a dump is not readable by every account on the
30
+ host. A directory that already exists keeps the mode it was made with: when the
31
+ backup target is a mount or a path an earlier run created, check its mode
32
+ yourself. The same goes for `pglite/`, which is copied file by file and keeps
33
+ the modes of the source data directory.
34
+
35
+ **PostgreSQL.** The command runs `pg_dump --format=custom` with the migrator
36
+ connection of the configured environment (`FD_DATABASE_MIGRATOR_URL`, else
37
+ `FD_DATABASE_URL`). The credentials reach the child process as `PG*` variables,
38
+ never as arguments, so no password lands in the host's process list, and the
39
+ child inherits none of the platform's encryption keys. `FD_DATABASE_TLS` and
40
+ `FD_DATABASE_TLS_CA[_FILE]` are translated into `PGSSLMODE` and `PGSSLROOTCERT`,
41
+ so a `verify-full` deployment is dumped over a verified connection. Without the
42
+ client tools on `PATH` the command fails with `BACKUP_TOOL_MISSING`; install the
43
+ PostgreSQL client package matching the server version.
44
+
45
+ **Embedded PGlite.** Outside production the database is a directory
46
+ (`FD_DATABASE_PGLITE_DIRECTORY`, by default `.flowdular/data/pglite`), and the
47
+ command copies it. Stop the application first. A copy taken while the process is
48
+ writing can hold a torn state, and the command says so in a warning.
49
+
50
+ ### What a backup does not contain
51
+
52
+ > **The encryption keys live outside the database.** Store every key with the
53
+ > backup. A restore without them yields unreadable credentials and workflow
54
+ > payloads: the rows come back, the values inside them stay ciphertext forever.
55
+
56
+ `backup.json` records, for each key, a SHA-256 fingerprint of the key material
57
+ and never the material itself:
58
+
59
+ ```json
60
+ {
61
+ "schemaVersion": 1,
62
+ "createdAt": "2026-09-11T09:30:00.000Z",
63
+ "adapter": "postgresql",
64
+ "platformVersion": "0.2.0",
65
+ "modules": ["agents.core", "auth.core", "workflows.core"],
66
+ "keys": [
67
+ {
68
+ "variable": "FD_AGENT_CREDENTIAL_KEY",
69
+ "fingerprint": "sha256:0f1e..."
70
+ },
71
+ { "variable": "FD_AUTH_MFA_KEY", "fingerprint": null }
72
+ ]
73
+ }
74
+ ```
75
+
76
+ The ten keys are `FD_AGENT_CREDENTIAL_KEY`, `FD_AGENT_RUN_GRANT_KEY`,
77
+ `FD_WORKFLOWS_PAYLOAD_KEY`, `FD_WORKFLOWS_CURSOR_KEY`, `FD_AUTH_MFA_KEY`,
78
+ `FD_AUTOMATIONS_CREDENTIAL_KEY`, `FD_NOTIFICATIONS_SECRET_KEY`,
79
+ `FD_STORAGE_ENCRYPTION_KEY`, `FD_CONNECTORS_SECRET_KEY` and
80
+ `FD_AUDIT_ANCHOR_KEY`. A `null` fingerprint means the key was not set
81
+ when the backup was taken. Keep the key material in the secret store the
82
+ deployment already uses, and record which backup it belongs to; the fingerprint
83
+ is what lets a restore tell you whether the pair matches.
84
+
85
+ Without `FD_AUDIT_ANCHOR_KEY` a restored deployment cannot verify the anchors
86
+ it holds, and every sealed segment file becomes a file whose origin nobody can
87
+ prove, so treat its fingerprint like the others.
88
+
89
+ The per-subject audit keys live in `audit_subject_keys` inside the database, so
90
+ a database backup carries them. A backup taken before
91
+ `flowdular audit erase --destroy-key` still holds the key that run destroyed, so
92
+ restoring it makes those sealed fields readable again. Treat an erasure as a
93
+ reason to age out the backups that predate it, as you would for any other
94
+ deletion on request. The key row itself survives destruction as a tombstone
95
+ carrying the subject's marker, which is what stops a later event about that
96
+ subject from creating a second key and putting the account back in the clear.
97
+
98
+ ### Schedule
99
+
100
+ - Dump before every rollout, and on a daily schedule between them.
101
+ - Keep the archive off the database host, and keep the keys in the secret store,
102
+ not next to the dump.
103
+ - Keep one restore rehearsal in the calendar. A backup nobody has restored is a
104
+ guess; restoring into a scratch database is the only proof.
105
+ - Retention follows the data policy of the business, not this document. Delete
106
+ old keys only after the dumps they unlock are gone.
107
+
108
+ ## Restore
109
+
110
+ ```bash
111
+ flowdular database restore --input /backups/flowdular/2026-09-11 # plan
112
+ flowdular database restore --input /backups/flowdular/2026-09-11 --apply --confirm restore-database
113
+ ```
114
+
115
+ The plan reports the manifest, the payload it would read and one key comparison
116
+ per variable. When the running environment holds keys the backup was not taken
117
+ with, the envelope carries a `BACKUP_KEY_MISMATCH` warning naming each variable
118
+ and how it differs (`different`, `missing-in-environment`, `missing-in-backup`).
119
+ Restore the keys before the data, or the rows arrive unreadable.
120
+
121
+ The order that works:
122
+
123
+ 1. Stop the application, or take it out of the load balancer.
124
+ 2. Put the keys the backup was taken with into the environment.
125
+ 3. Run the restore as a dry run and read the key comparison.
126
+ 4. Run it with `--apply --confirm restore-database`.
127
+ 5. Start the application, then check `pnpm flowdular migration verify` and
128
+ `GET /api/ready`.
129
+
130
+ PostgreSQL restores run `pg_restore --clean --if-exists` against the migrator
131
+ connection, so the target database must already exist and its roles must already
132
+ be there. The embedded adapter is replaced directory by directory: the copy is
133
+ staged next to the target and the old directory is removed only once the new one
134
+ is in place, so a failed copy leaves the running database untouched.
135
+
136
+ `database restore` is a destructive capability, gated exactly like
137
+ `database reset`: the runner refuses it with `LOCAL_ONLY_CAPABILITY` unless
138
+ `FD_ENV` or `NODE_ENV` is `development` or `test`. Production restores therefore
139
+ run `pg_restore --clean --if-exists` by hand with the migrator credentials, and
140
+ the dry run on a staging copy of the same backup is what proves the archive and
141
+ the keys are good. Lifting that restriction waits on the signed approval
142
+ verifier listed under `planned` in `.ai/policies/capabilities.yaml`.
143
+
144
+ ## Key rotation
145
+
146
+ Most of the keys carry a key id in every envelope they write, so a
147
+ deployment can hold the new key and the old one at the same time and re-seal the
148
+ stored rows in place. The others have no key id and are covered at the end of
149
+ this section.
150
+
151
+ Every module reads one current key plus an optional comma-separated list of
152
+ retired keys (up to eight):
153
+
154
+ | Key | Protects | Retired keys | Re-sealing the stored rows |
155
+ | ------------------------------- | -------------------------- | ---------------------------------------- | ------------------------------------------------------- |
156
+ | `FD_AGENT_CREDENTIAL_KEY` | Agent provider credentials | `FD_AGENT_CREDENTIAL_KEY_PREVIOUS` | `pnpm flowdular agents secrets-rotate [--apply]` |
157
+ | `FD_AUTOMATIONS_CREDENTIAL_KEY` | Automation trigger secrets | `FD_AUTOMATIONS_CREDENTIAL_KEY_PREVIOUS` | `pnpm flowdular automations secrets-rotate [--apply]` |
158
+ | `FD_WORKFLOWS_PAYLOAD_KEY` | Workflow run payloads | `FD_WORKFLOWS_PAYLOAD_KEY_PREVIOUS` | `pnpm flowdular workflows secrets-rotate [--apply]` |
159
+ | `FD_STORAGE_ENCRYPTION_KEY` | Stored objects | `FD_STORAGE_ENCRYPTION_KEY_PREVIOUS` | No re-sealing pass yet; see the storage note below |
160
+ | `FD_NOTIFICATIONS_SECRET_KEY` | Webhook signing secrets | `FD_NOTIFICATIONS_SECRET_KEY_PREVIOUS` | `pnpm flowdular notifications secrets-rotate [--apply]` |
161
+ | `FD_CONNECTORS_SECRET_KEY` | Connector credentials | `FD_CONNECTORS_SECRET_KEY_PREVIOUS` | No re-sealing pass yet; the ring reads the previous key |
162
+ | `FD_AUDIT_ANCHOR_KEY` | Audit chain anchors (HMAC) | `FD_AUDIT_ANCHOR_KEY_PREVIOUS` | `pnpm flowdular audit secrets-rotate [--apply]` |
163
+
164
+ These commands run the same re-sealing pass over their own table, so the
165
+ procedure is the same for each. The credential key is the worked example;
166
+ substitute the key names and the command from the row you are rotating:
167
+
168
+ 1. Generate the new key: `openssl rand -base64 32`.
169
+ 2. Set the new key as the current one and the old key as the retired one:
170
+ `FD_AGENT_CREDENTIAL_KEY=<new>` and `FD_AGENT_CREDENTIAL_KEY_PREVIOUS=<old>`.
171
+ Keep both in the secret store, exactly like the key they replace.
172
+ 3. Deploy. From this moment every write uses the new key and every stored row
173
+ still opens under the old one, so nothing is unreadable while the rollout
174
+ completes.
175
+ 4. Run the rotation as a dry run and read the counts per key id:
176
+ `pnpm flowdular agents secrets-rotate --json`. `stale` is the number of rows
177
+ still sealed with a retired key.
178
+ 5. Run it again with `--apply`. It re-seals rows in batches of 200, one
179
+ transaction per batch, scoped to the tenant that owns them, and writes nothing
180
+ else: the revision, the audit fields and the credential a caller sees do not
181
+ change. It is idempotent, so running it twice is free.
182
+ 6. Repeat the dry run until it reports `stale: 0`, then remove
183
+ `FD_AGENT_CREDENTIAL_KEY_PREVIOUS` and deploy again.
184
+
185
+ The application may stay up. Each row is written back only if its stored
186
+ envelope is still the one that was read, so a credential the application
187
+ rewrites in between is never clobbered: the rotation counts it under `skipped`
188
+ and the next run picks it up. Stopping the application removes that case
189
+ entirely. The inventory that finds the stale rows runs on the cross-tenant
190
+ read-only role, which is granted the tenant id and the key id and cannot read a
191
+ nonce, a tag or a ciphertext; every row it names is read again under its own
192
+ tenant before it is touched.
193
+
194
+ `FD_AUDIT_ANCHOR_KEY` follows the same six steps with one difference: an anchor
195
+ is signed, not sealed, so the rotation rewrites the signature and the key id and
196
+ touches neither the anchor hash, the segment hash nor the chain. A segment file
197
+ written under the retired key keeps its own signature and still verifies while
198
+ that key is in `FD_AUDIT_ANCHOR_KEY_PREVIOUS`. Keep the retired key until every
199
+ archived segment has been verified once under the new one, with
200
+ `pnpm flowdular audit verify --workspace <slug> --input <dir>` over the archive.
201
+ If the rotation reports `ANCHOR_KEY_UNKNOWN`, anchors exist that neither key can
202
+ re-sign: put that key back before retiring it, because a signature nobody can
203
+ reproduce can never be verified again. The rotation verifies an anchor's stored
204
+ signature under the ring before re-signing it, so `ANCHOR_SIGNATURE_INVALID`
205
+ names anchors it left exactly as they are: a new signature over the same anchor
206
+ hash would replace evidence nobody can account for with one this deployment
207
+ vouches for. Verify the segment files of those anchors before anything else.
208
+ `FD_AUDIT_ANCHOR_KEY_PREVIOUS` accepts at most eight retired keys; a longer list
209
+ is a rotation that was never finished and the platform refuses to start.
210
+
211
+ The remaining keys, and how rotation behaves for each:
212
+
213
+ | Key | Protects | Rotating it |
214
+ | ------------------------- | -------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
215
+ | `FD_WORKFLOWS_CURSOR_KEY` | Pagination cursors (HMAC) | Safe. `FD_WORKFLOWS_CURSOR_KEY_PREVIOUS` keeps outstanding cursors valid; cursors are never stored, so there is nothing to re-sign. |
216
+ | `FD_AGENT_RUN_GRANT_KEY` | Agent run grant tokens (HMAC, 30s TTL) | Safe. In-flight grants fail for the length of one TTL. |
217
+ | `FD_AUTH_MFA_KEY` | Enrolled TOTP secrets | Supported since auth.core 0.11.0. Set `FD_AUTH_MFA_KEY_PREVIOUS`, deploy, run `pnpm flowdular auth secrets-rotate --apply`, then drop the previous key. Rows written before the key id column exist open through the ring fallback until they are re-sealed. |
218
+
219
+ The key ids are derived from the key material, so a backup fingerprint that
220
+ matches the environment also matches the `keyId` inside the rows. The keys that
221
+ carry no id are the ones a restore cannot diagnose for you, which is why
222
+ `backup.json` records a fingerprint for all eight.
223
+
224
+ Sources: `modules/audit/src/services/{anchor-key,anchor-rotation}.ts`,
225
+ `packages/kernel/src/keyring.ts`,
226
+ `modules/agents/src/services/{credential-vault,credential-rotation}.ts`,
227
+ `modules/automations/src/services/{secret-vault,secret-rotation}.ts`,
228
+ `modules/workflows/src/services/{payload-codec,payload-rotation,cursors}.ts`,
229
+ `modules/notifications/src/services/{secret-vault,secret-rotation}.ts`,
230
+ `modules/auth/src/services/totp.ts`.
231
+
232
+ ## Storage
233
+
234
+ Objects live outside PostgreSQL, so a database backup does not contain one. Back
235
+ up the object store the same way you back up the database, and take both at the
236
+ same point: a restored database that references objects a newer bucket snapshot
237
+ no longer holds leaves dangling references, and the reverse leaves orphans.
238
+
239
+ - **S3-compatible store.** Enable versioning and a lifecycle policy on the
240
+ bucket, or replicate it. `backup.json` does not reach into it.
241
+ - **Local adapter.** Development and test only, refused in production. Its
242
+ directory is `FD_STORAGE_LOCAL_DIRECTORY`, by default
243
+ `.flowdular/data/storage`; stop the application before copying it.
244
+
245
+ Every object is sealed with AES-256-GCM under `FD_STORAGE_ENCRYPTION_KEY`, and
246
+ the key id is stored with the object, so a rotation runs like the four above:
247
+ put the retired key in `FD_STORAGE_ENCRYPTION_KEY_PREVIOUS`, deploy, and every
248
+ object written earlier still opens while new ones use the current key. There is
249
+ no re-sealing pass yet, so keep the retired key in the ring until every object
250
+ written under it has been rewritten or deleted. Dropping it makes those objects
251
+ unreadable, exactly as a lost database key makes a credential unreadable.
252
+
253
+ ## Logs
254
+
255
+ The server writes one line per event through `createLogger`
256
+ (`packages/server/src/log.ts`).
257
+
258
+ | Variable | Default | Purpose |
259
+ | --------------- | --------------------------------- | ------------------------------------------------ |
260
+ | `FD_LOG_FORMAT` | `json` in production, else `text` | `json` for a log pipeline, `text` for a terminal |
261
+ | `FD_LOG_LEVEL` | `info` | `debug`, `info`, `warn` or `error` |
262
+
263
+ A JSON line carries `time`, `level`, `msg` and, where they apply, `traceId`,
264
+ `spanId`, `requestId`, `endpoint`, `module` and `err` (name and message). The
265
+ trace ids are on every line written inside a served request, so the same id
266
+ finds the log line and the trace. Stacks are written outside
267
+ production, and in production only at `FD_LOG_LEVEL=debug`. The `requestId` is
268
+ the same value the response returns in `x-request-id`, so a user-reported error
269
+ id finds its log line. Credential-shaped field names (`authorization`, `cookie`,
270
+ `password`, `token`, `secret`, `credential`, API keys) are redacted before a
271
+ line is written, and a failing endpoint logs its id and the error type only,
272
+ never the message, the body or the URL. An unrecognized value for either
273
+ variable falls back to the default and logs one warning about it.
274
+
275
+ ## Production checklist
276
+
277
+ - `pnpm verify` and `pnpm build` pass on the commit being shipped.
278
+ - `FD_ENV=production`, `FD_DATABASE_ADAPTER=postgresql`, separate
279
+ `FD_DATABASE_URL` and `FD_DATABASE_MIGRATOR_URL` roles, neither `SUPERUSER`
280
+ nor `BYPASSRLS` on the runtime role.
281
+ - `FD_DATABASE_TLS=verify-full` with `FD_DATABASE_TLS_CA` or
282
+ `FD_DATABASE_TLS_CA_FILE`. Production accepts nothing weaker.
283
+ - All six encryption keys set, held in the secret store, and backed up with a
284
+ recorded fingerprint.
285
+ - `FD_AUTH_SECURE_COOKIE` on, `FD_AUTH_PUBLIC_ORIGIN` equal to the public URL,
286
+ `FD_TRUST_PROXY` set behind a proxy, sign-up closed unless the deployment is
287
+ public.
288
+ - Mail transport configured: `FD_AUTH_MAIL_TRANSPORT` with `FD_AUTH_SMTP_URL`
289
+ and `FD_AUTH_MAIL_FROM`. Password resets, invitations and address
290
+ confirmations are silently undeliverable without it.
291
+ - Migrations at rollout: deploy the image, let it apply its module migrations
292
+ under the migrator role, then `pnpm flowdular migration verify`. Never edit an
293
+ applied `.up.sql`.
294
+ - Liveness probe on `GET /api/health`, readiness probe on `GET /api/ready`.
295
+ - `FD_LOG_FORMAT=json` and a log pipeline that keeps `requestId`.
296
+ - A fresh backup, taken and verified before the rollout starts.
297
+
298
+ ## Rollback
299
+
300
+ 1. Redeploy the previous image tag. Module migrations are additive, so an older
301
+ image runs against a newer schema and ignores columns it does not know.
302
+ 2. Never apply a `.down.sql` to roll back a release. They document the reverse
303
+ for review; against live data they are data loss.
304
+ 3. To take one module out of service, `pnpm flowdular module disable <id> --apply`,
305
+ rebuild the composition and redeploy. Its tables stay where they are.
306
+ 4. If the rollback is about a key, restore the key first. A previous image does
307
+ not make ciphertext readable again.
308
+ 5. Restore from a backup only when the schema or the data is actually damaged.
309
+ It is the slowest option and the only one that loses everything written since
310
+ the dump.
311
+
312
+ ## Metrics
313
+
314
+ `GET /api/metrics` serves Prometheus text exposition, format version 0.0.4. The
315
+ route is composed only when `FD_METRICS=true`; with `FD_METRICS_TOKEN` set, a
316
+ scrape must send `Authorization: Bearer <token>` and gets 401 without it.
317
+
318
+ **An untokened `/api/metrics` must never be reachable outside the scrape
319
+ network.** It shares the application port, so without `FD_METRICS_TOKEN` every
320
+ client that can reach the port reads the series, including the endpoint
321
+ inventory and the request rates of the whole deployment. Either set the token
322
+ (the Kubernetes manifests read it from the deployment Secret, optional so it
323
+ can be left unset deliberately) or keep the port on a network only the scraper
324
+ reaches. Do not rely on the route being obscure.
325
+
326
+ | Series | Type | Labels |
327
+ | ----------------------------------------- | --------- | ----------------------------------- |
328
+ | `flowdular_http_requests_total` | counter | `endpoint`, `method`, `status` |
329
+ | `flowdular_http_request_duration_seconds` | histogram | `endpoint`, buckets 0.005 s to 10 s |
330
+ | `flowdular_metrics_dropped_samples_total` | counter | none |
331
+ | `flowdular_process_start_time_seconds` | gauge | none |
332
+ | `process_resident_memory_bytes` | gauge | none |
333
+ | `nodejs_eventloop_lag_seconds` | gauge | none |
334
+ | `flowdular_build_info` | gauge | `version` |
335
+
336
+ `endpoint` is the endpoint id, never the request path, so a record id cannot
337
+ reach a label. `status` is the class (`2xx`, `4xx`, `5xx`), and an unrecognized
338
+ method is recorded as `other`. No session, tenant or user value is ever a label.
339
+
340
+ `nodejs_eventloop_lag_seconds` is the mean delay over the interval since the
341
+ previous scrape, not since the process started, so a spike shows up in the
342
+ scrape that follows it and is gone from the next one. The sampler is
343
+ process-wide and is drained on every read, so two scrapers pointed at the same
344
+ process split the samples between them; point one scraper at it.
345
+
346
+ Every request through `defineEndpoint` is recorded, whether or not the route is
347
+ exposed, so a scraper enabled later sees the counters immediately. Memory is
348
+ bounded: the process keeps at most 2000 distinct label sets per series family
349
+ and refuses new ones after that, counting each refusal in
350
+ `flowdular_metrics_dropped_samples_total`. A non-zero value there means a
351
+ deployment composes more endpoints than the ceiling allows.
352
+
353
+ A module adds series of its own with `createModuleMetrics('<module id>')`:
354
+
355
+ | Series | Type | Labels |
356
+ | ----------------------------------------------------------------- | --------- | --------------------------- |
357
+ | `flowdular_module_<module>_<name>_total` | counter | the module's own |
358
+ | `flowdular_module_<module>_<name>` (+`_bucket`, `_sum`, `_count`) | histogram | the module's own, plus `le` |
359
+
360
+ The module id is lower-cased with every character outside `a-z0-9_` replaced by
361
+ `_`, so `agents.core` reads `flowdular_module_agents_core_`. A module series is
362
+ refused, and counted as a dropped sample, when the metric name does not match
363
+ `^[a-z][a-z0-9_]{0,63}$`, when a call carries more than 8 labels or a label name
364
+ of the wrong shape, when the family already holds 2000 label sets, or when the
365
+ process already opened 256 module families of that kind. A label value is cut at
366
+ 128 characters. A module never receives an error from recording: a refused
367
+ sample is lost, the work is not.
368
+
369
+ **The label rule of the request series holds for a module series too: never a
370
+ tenant, account, record or session id.** The platform cannot tell one from a
371
+ closed set, so nothing refuses it; what it does is cap the damage at 2000 label
372
+ sets per family, after which the series stops admitting new ones. A label is for
373
+ a value with a handful of possibilities, such as an outcome or a channel.
374
+
375
+ ## Tracing
376
+
377
+ Every request through `defineEndpoint` runs in a W3C trace. The inbound
378
+ `traceparent` names the parent when it is usable; anything else is a new root.
379
+ The response always returns a `traceparent` naming the span the request ran in,
380
+ so a client, a proxy and this process agree on one trace id, and the log lines
381
+ of that request carry the same `traceId`.
382
+
383
+ Spans go into one bounded in-process ring whether or not an exporter is
384
+ composed: 4096 spans, oldest evicted and counted, at most 16 attributes per span
385
+ and 256 characters per name, attribute key and string value. Attributes carry
386
+ endpoint ids, job names, provider and tool ids, HTTP method and status, and the
387
+ three ids that correlate a span with a log line or an agent run,
388
+ `flowdular.request_id`, `flowdular.agent.run_id` and
389
+ `flowdular.tool.provider_call_id`: no request path, body, tenant or session.
390
+
391
+ | Variable | Purpose |
392
+ | ----------------------- | --------------------------------------------------- |
393
+ | `FD_TRACE_SAMPLE` | Ratio of new roots recorded, 0 to 1, default 1 |
394
+ | `FD_TRACE_EXPORTER` | `none` (default) or `otlp` |
395
+ | `FD_TRACE_OTLP_URL` | OTLP/HTTP JSON traces endpoint; https in production |
396
+ | `FD_TRACE_OTLP_HEADERS` | `name=value,name2=value2`, at most 16 |
397
+
398
+ The exporter posts OTLP/HTTP JSON (no SDK, no agent) in batches of 512 spans, or
399
+ every 5 seconds when the batch does not fill. It never runs on a request path: a
400
+ collector that is down costs the process one held batch, retried on the next two
401
+ flushes and then dropped. A `4xx` that is not `429` is not retried, because the
402
+ payload or the credential is wrong and the next batch would be lost too. The
403
+ sampling decision is taken on the trace id, so a whole trace is in or out; a
404
+ trace the caller already marked sampled is never re-decided, and an unsampled
405
+ trace still propagates its header.
406
+
407
+ A process that stops drains what the ring still holds, in the same batches and
408
+ without a retry: the spans of a shutdown leave with it, and a collector that is
409
+ down is given up on rather than held between the process and its exit.
410
+
411
+ **An OTLP endpoint receives the endpoint inventory, the job names, the timings
412
+ and the request, agent run and provider call ids of the whole deployment.**
413
+ Point it at a collector you control. Plain `http` is refused in production.
414
+
415
+ ## Error reporting
416
+
417
+ `FD_ERROR_SINK=webhook` posts every line the server logs at `error` to
418
+ `FD_ERROR_SINK_URL`, with `Authorization: Bearer <FD_ERROR_SINK_TOKEN>` when the
419
+ token is set. The body is `{"reports":[…]}` where a report carries `at`, `name`,
420
+ `message` and, where they apply, `endpoint`, `module`, `requestId`, `traceId`
421
+ and `spanId`. Nothing else: no stack, no log fields, no request body, so the
422
+ sink cannot carry a credential the logger already refuses to write.
423
+
424
+ `message` is the one field the caller chose. A failing endpoint logs its error
425
+ type only, never the message, so its report carries `endpoint failed`; a
426
+ background loop logs the error it caught, so its report carries that message.
427
+ The sink therefore exports exactly what the log line already shows, to a host
428
+ outside the deployment. A deployment that must not let error text leave keeps
429
+ `FD_ERROR_SINK=none` and reads the log stream instead.
430
+
431
+ It is bounded and lossy on purpose: 64 queued reports (a report past that is
432
+ dropped and counted), 32 per request, 8 KiB per body with the batch halved until
433
+ it fits, a batch sent every 5 seconds or as soon as 32 are queued, and at most 3
434
+ attempts before a batch is given up. Delivery never runs on the caller's path
435
+ and an unreachable webhook never fails the line that reported the failure.
436
+ `none` is the default and holds no queue and no timer.
437
+
438
+ The OTLP logs sink named in RFC 0004 H7 is deliberately not built: the OTLP
439
+ surface this platform speaks is the trace exporter above.
440
+
441
+ ## Data residency
442
+
443
+ Residency is a property of the deployment, not of the workspace. A deployment
444
+ carries one PostgreSQL database, one object store and one region; every module
445
+ writes to that database and, through the storage port, to that bucket. A
446
+ workspace therefore lives wherever its deployment lives, and a customer whose
447
+ data must stay in a particular region gets a deployment in that region. That
448
+ is the supported answer and the only one the platform offers.
449
+
450
+ Per-workspace routing to regional databases is not planned. It would change
451
+ the provider contract in `packages/database` (one handle per region), the
452
+ lease model every module's migrations and background loops depend on, every
453
+ cross-workspace routing read the background role performs (the retention
454
+ sweep, delivery queues, expiry passes), the storage key layout, and the
455
+ export and erasure paths that assume one place to look. Record that cost
456
+ before anyone proposes it; a second deployment is cheaper in every case
457
+ seen so far.
458
+
459
+ What to state in a questionnaire: the region of the database and the bucket,
460
+ that backups and the key material are stored in the same region unless the
461
+ operator moves them, that the OTLP trace exporter and the error sink send
462
+ only what `docs/operations.md` lists for them and only to the URLs the
463
+ operator configured, and that outbound mail, webhooks and connector calls
464
+ leave the region by design of the recipient the workspace configured.