dirigent-examples 0.15.0__py3-none-any.whl

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (216) hide show
  1. dirigent_examples/__init__.py +22 -0
  2. dirigent_examples/py.typed +0 -0
  3. dirigent_examples/shelves/README.md +299 -0
  4. dirigent_examples/shelves/composition/README.md +18 -0
  5. dirigent_examples/shelves/composition/chained-instances.yaml +97 -0
  6. dirigent_examples/shelves/composition/composition-child.yaml +64 -0
  7. dirigent_examples/shelves/composition/composition-parent.yaml +89 -0
  8. dirigent_examples/shelves/connections.yaml +52 -0
  9. dirigent_examples/shelves/demo/README.md +19 -0
  10. dirigent_examples/shelves/demo/markdown-showcase.yaml +117 -0
  11. dirigent_examples/shelves/demo/params-showcase.yaml +92 -0
  12. dirigent_examples/shelves/demo/requires.yaml +65 -0
  13. dirigent_examples/shelves/demo/weekly-import-malawi.yaml +54 -0
  14. dirigent_examples/shelves/demo/weekly-import-nepal.yaml +75 -0
  15. dirigent_examples/shelves/docker/README.md +29 -0
  16. dirigent_examples/shelves/docker/docker-build-push.yaml +110 -0
  17. dirigent_examples/shelves/docker/docker-build-run.yaml +106 -0
  18. dirigent_examples/shelves/docker/docker-compose-database.yaml +124 -0
  19. dirigent_examples/shelves/docker/docker-compose-failing-up.yaml +83 -0
  20. dirigent_examples/shelves/docker/docker-compose-file.yaml +117 -0
  21. dirigent_examples/shelves/docker/docker-compose-profiles-env.yaml +133 -0
  22. dirigent_examples/shelves/docker/docker-compose-stack.yaml +70 -0
  23. dirigent_examples/shelves/docker/docker-hello.yaml +53 -0
  24. dirigent_examples/shelves/docker/docker-remote-daemon.yaml +92 -0
  25. dirigent_examples/shelves/docker/docker-run-failing-teardown.yaml +94 -0
  26. dirigent_examples/shelves/docker/docker-ticker.yaml +49 -0
  27. dirigent_examples/shelves/execute/README.md +16 -0
  28. dirigent_examples/shelves/execute/long-log.yaml +89 -0
  29. dirigent_examples/shelves/failure/README.md +20 -0
  30. dirigent_examples/shelves/failure/error-handler.yaml +89 -0
  31. dirigent_examples/shelves/failure/optional-step.yaml +82 -0
  32. dirigent_examples/shelves/failure/retries.yaml +75 -0
  33. dirigent_examples/shelves/failure/retry-budget.yaml +82 -0
  34. dirigent_examples/shelves/failure/step-timeout.yaml +96 -0
  35. dirigent_examples/shelves/git/README.md +32 -0
  36. dirigent_examples/shelves/git/git-checkout-build.yaml +125 -0
  37. dirigent_examples/shelves/git/git-checkout-compose.yaml +138 -0
  38. dirigent_examples/shelves/git/git-checkout-public.yaml +84 -0
  39. dirigent_examples/shelves/graph/README.md +22 -0
  40. dirigent_examples/shelves/graph/deep-chain.yaml +119 -0
  41. dirigent_examples/shelves/graph/fan-in.yaml +80 -0
  42. dirigent_examples/shelves/graph/fan-out.yaml +66 -0
  43. dirigent_examples/shelves/graph/linear.yaml +66 -0
  44. dirigent_examples/shelves/graph/parallel-branches.yaml +57 -0
  45. dirigent_examples/shelves/graph/parallel-sleep.yaml +55 -0
  46. dirigent_examples/shelves/graph/skip-diamond.yaml +105 -0
  47. dirigent_examples/shelves/graph/wide-fan.yaml +147 -0
  48. dirigent_examples/shelves/hello-world.yaml +30 -0
  49. dirigent_examples/shelves/open-data/README.md +67 -0
  50. dirigent_examples/shelves/open-data/feeds-composition.yaml +120 -0
  51. dirigent_examples/shelves/open-data/gdacs-disaster-updates.yaml +230 -0
  52. dirigent_examples/shelves/open-data/github-releases-relay.yaml +225 -0
  53. dirigent_examples/shelves/open-data/hdx-dataset-watch.yaml +211 -0
  54. dirigent_examples/shelves/open-data/kobo-submissions-to-csv.yaml +149 -0
  55. dirigent_examples/shelves/open-data/nominatim-geocode-facilities.yaml +169 -0
  56. dirigent_examples/shelves/open-data/odk-central-submissions.yaml +146 -0
  57. dirigent_examples/shelves/open-data/open-meteo-weekly-report.yaml +142 -0
  58. dirigent_examples/shelves/open-data/overpass-health-facilities.yaml +154 -0
  59. dirigent_examples/shelves/open-data/usgs-earthquakes-alert.yaml +208 -0
  60. dirigent_examples/shelves/open-data/who-gho-indicators-to-parquet.yaml +163 -0
  61. dirigent_examples/shelves/open-data/wikidata-country-reference.yaml +146 -0
  62. dirigent_examples/shelves/open-data/world-bank-population-trend.yaml +166 -0
  63. dirigent_examples/shelves/patterns/README.md +144 -0
  64. dirigent_examples/shelves/patterns/concurrency-queue.yaml +89 -0
  65. dirigent_examples/shelves/patterns/concurrency-replace.yaml +92 -0
  66. dirigent_examples/shelves/patterns/concurrency-skip.yaml +97 -0
  67. dirigent_examples/shelves/patterns/connections-referenced-vs-carried.yaml +140 -0
  68. dirigent_examples/shelves/patterns/deadline-on-a-sensor.yaml +106 -0
  69. dirigent_examples/shelves/patterns/fan-out-continue.yaml +88 -0
  70. dirigent_examples/shelves/patterns/fan-out-fail-fast.yaml +80 -0
  71. dirigent_examples/shelves/patterns/fan-out-from-params.yaml +84 -0
  72. dirigent_examples/shelves/patterns/fan-out-item-wise.yaml +111 -0
  73. dirigent_examples/shelves/patterns/fan-out-literal-list.yaml +81 -0
  74. dirigent_examples/shelves/patterns/fan-out-nested-objects.yaml +107 -0
  75. dirigent_examples/shelves/patterns/fan-out-then-join.yaml +86 -0
  76. dirigent_examples/shelves/patterns/log-levels.yaml +119 -0
  77. dirigent_examples/shelves/patterns/outputs-inline-vs-storage.yaml +140 -0
  78. dirigent_examples/shelves/patterns/params-every-type.yaml +259 -0
  79. dirigent_examples/shelves/patterns/params-validation-refuses.yaml +131 -0
  80. dirigent_examples/shelves/patterns/pipeline-run-child.yaml +96 -0
  81. dirigent_examples/shelves/patterns/pipeline-run-fire-and-forget.yaml +99 -0
  82. dirigent_examples/shelves/patterns/pipeline-run-strict.yaml +103 -0
  83. dirigent_examples/shelves/patterns/pipeline-run-wait.yaml +107 -0
  84. dirigent_examples/shelves/patterns/pipeline-run-with-params.yaml +124 -0
  85. dirigent_examples/shelves/patterns/poll-cadence.yaml +102 -0
  86. dirigent_examples/shelves/patterns/priority-layered.yaml +120 -0
  87. dirigent_examples/shelves/patterns/references-cheat-sheet.yaml +186 -0
  88. dirigent_examples/shelves/patterns/retry-budget-exhausted.yaml +92 -0
  89. dirigent_examples/shelves/patterns/retry-exponential-backoff.yaml +88 -0
  90. dirigent_examples/shelves/patterns/retry-only-transient.yaml +108 -0
  91. dirigent_examples/shelves/patterns/retry-with-jitter.yaml +102 -0
  92. dirigent_examples/shelves/patterns/rule-all-done.yaml +83 -0
  93. dirigent_examples/shelves/patterns/rule-all-success.yaml +79 -0
  94. dirigent_examples/shelves/patterns/rule-always.yaml +92 -0
  95. dirigent_examples/shelves/patterns/rule-one-failed.yaml +86 -0
  96. dirigent_examples/shelves/patterns/schedule-at-once.yaml +110 -0
  97. dirigent_examples/shelves/patterns/schedule-cron-timezone.yaml +121 -0
  98. dirigent_examples/shelves/patterns/schedule-interval.yaml +109 -0
  99. dirigent_examples/shelves/patterns/schedule-window-half-open.yaml +105 -0
  100. dirigent_examples/shelves/patterns/sensor-http-ready.yaml +121 -0
  101. dirigent_examples/shelves/patterns/sensor-storage-exists.yaml +134 -0
  102. dirigent_examples/shelves/patterns/step-names-and-keys.yaml +99 -0
  103. dirigent_examples/shelves/patterns/timeout-fails-the-step.yaml +94 -0
  104. dirigent_examples/shelves/patterns/timeout-skips-the-step.yaml +102 -0
  105. dirigent_examples/shelves/patterns/webhook-mapping-nested-payload.yaml +125 -0
  106. dirigent_examples/shelves/patterns/webhook-signed.yaml +144 -0
  107. dirigent_examples/shelves/preview/s3-parquet-to-ingestion.yaml +92 -0
  108. dirigent_examples/shelves/python/README.md +31 -0
  109. dirigent_examples/shelves/python/apply_and_run.py +52 -0
  110. dirigent_examples/shelves/python/ci_gate.py +76 -0
  111. dirigent_examples/shelves/python/connections.py +61 -0
  112. dirigent_examples/shelves/python/error_handling.py +84 -0
  113. dirigent_examples/shelves/python/follow_logs.py +39 -0
  114. dirigent_examples/shelves/python/list_and_filter.py +52 -0
  115. dirigent_examples/shelves/queues/README.md +59 -0
  116. dirigent_examples/shelves/queues/kafka-consume-then-transform.yaml +105 -0
  117. dirigent_examples/shelves/queues/kafka-produce-then-consume.yaml +124 -0
  118. dirigent_examples/shelves/queues/rabbitmq-consume-ack-on-success.yaml +102 -0
  119. dirigent_examples/shelves/queues/report-to-kafka.yaml +105 -0
  120. dirigent_examples/shelves/queues/report-to-rabbitmq.yaml +105 -0
  121. dirigent_examples/shelves/recipes/README.md +130 -0
  122. dirigent_examples/shelves/recipes/csv-header-rules.yaml +147 -0
  123. dirigent_examples/shelves/recipes/csv-to-ndjson.yaml +107 -0
  124. dirigent_examples/shelves/recipes/etl-csv-clean-validate-parquet.yaml +207 -0
  125. dirigent_examples/shelves/recipes/filter-by-predicate.yaml +116 -0
  126. dirigent_examples/shelves/recipes/filter-then-map-then-reduce.yaml +109 -0
  127. dirigent_examples/shelves/recipes/http-fetch-validate-post.yaml +142 -0
  128. dirigent_examples/shelves/recipes/http-follow-redirects.yaml +100 -0
  129. dirigent_examples/shelves/recipes/http-get-with-query.yaml +96 -0
  130. dirigent_examples/shelves/recipes/http-headers-and-auth-connection.yaml +110 -0
  131. dirigent_examples/shelves/recipes/http-post-file-from-storage.yaml +117 -0
  132. dirigent_examples/shelves/recipes/http-post-json-echo.yaml +105 -0
  133. dirigent_examples/shelves/recipes/http-post-report.yaml +183 -0
  134. dirigent_examples/shelves/recipes/http-save-body-to-storage.yaml +106 -0
  135. dirigent_examples/shelves/recipes/http-success-status-list.yaml +80 -0
  136. dirigent_examples/shelves/recipes/http-timeout-override.yaml +104 -0
  137. dirigent_examples/shelves/recipes/jq-dedupe-by-key.yaml +78 -0
  138. dirigent_examples/shelves/recipes/jq-defaults-and-nulls.yaml +91 -0
  139. dirigent_examples/shelves/recipes/jq-group-by-and-sum.yaml +74 -0
  140. dirigent_examples/shelves/recipes/jq-join-two-lists.yaml +77 -0
  141. dirigent_examples/shelves/recipes/jq-long-to-wide.yaml +76 -0
  142. dirigent_examples/shelves/recipes/jq-nested-to-flat.yaml +89 -0
  143. dirigent_examples/shelves/recipes/jq-pivot-wide-to-long.yaml +65 -0
  144. dirigent_examples/shelves/recipes/jq-running-totals.yaml +82 -0
  145. dirigent_examples/shelves/recipes/jq-string-cleaning.yaml +88 -0
  146. dirigent_examples/shelves/recipes/jq-top-n.yaml +85 -0
  147. dirigent_examples/shelves/recipes/jq-validate-in-jq-vs-schema.yaml +124 -0
  148. dirigent_examples/shelves/recipes/jq-window-dates.yaml +82 -0
  149. dirigent_examples/shelves/recipes/json-to-csv-flattening.yaml +141 -0
  150. dirigent_examples/shelves/recipes/large-output-to-storage.yaml +134 -0
  151. dirigent_examples/shelves/recipes/map-enrich-with-lookup.yaml +96 -0
  152. dirigent_examples/shelves/recipes/ndjson-to-parquet.yaml +130 -0
  153. dirigent_examples/shelves/recipes/pagination-by-fan-out.yaml +115 -0
  154. dirigent_examples/shelves/recipes/parquet-round-trip-types.yaml +163 -0
  155. dirigent_examples/shelves/recipes/reconcile-two-sources.yaml +159 -0
  156. dirigent_examples/shelves/recipes/report-built-in.yaml +72 -0
  157. dirigent_examples/shelves/recipes/report-daily-digest.yaml +186 -0
  158. dirigent_examples/shelves/recipes/report-to-file.yaml +131 -0
  159. dirigent_examples/shelves/recipes/report-to-webhook.yaml +136 -0
  160. dirigent_examples/shelves/recipes/schema-carried.yaml +112 -0
  161. dirigent_examples/shelves/recipes/schema-formats.yaml +107 -0
  162. dirigent_examples/shelves/recipes/schema-referenced.yaml +86 -0
  163. dirigent_examples/shelves/recipes/schema-refuses-then-rule.yaml +127 -0
  164. dirigent_examples/shelves/recipes/storage-copy-dated-archive.yaml +114 -0
  165. dirigent_examples/shelves/recipes/storage-exists-gate.yaml +127 -0
  166. dirigent_examples/shelves/recipes/storage-manifest-of-a-fan-out.yaml +104 -0
  167. dirigent_examples/shelves/recipes/storage-write-then-read.yaml +119 -0
  168. dirigent_examples/shelves/recipes/webhook-post-hmac.yaml +132 -0
  169. dirigent_examples/shelves/recipes/webhook-post-summary.yaml +142 -0
  170. dirigent_examples/shelves/s3/README.md +34 -0
  171. dirigent_examples/shelves/s3/report-to-s3.yaml +93 -0
  172. dirigent_examples/shelves/s3/s3-copy-and-verify.yaml +105 -0
  173. dirigent_examples/shelves/s3/s3-csv-report.yaml +87 -0
  174. dirigent_examples/shelves/s3/s3-parquet-report.yaml +77 -0
  175. dirigent_examples/shelves/s3/s3-round-trip.yaml +125 -0
  176. dirigent_examples/shelves/schemas/README.md +36 -0
  177. dirigent_examples/shelves/schemas/echo-reading.json +18 -0
  178. dirigent_examples/shelves/schemas/ou-record.json +13 -0
  179. dirigent_examples/shelves/schemas/station-reading.json +13 -0
  180. dirigent_examples/shelves/sensors/README.md +16 -0
  181. dirigent_examples/shelves/sensors/sensor-gate.yaml +65 -0
  182. dirigent_examples/shelves/sensors/time-window.yaml +61 -0
  183. dirigent_examples/shelves/sql/README.md +52 -0
  184. dirigent_examples/shelves/sql/duckdb-parquet-to-report.yaml +146 -0
  185. dirigent_examples/shelves/sql/sql-postgres-readonly.yaml +111 -0
  186. dirigent_examples/shelves/sql/sql-query-to-storage.yaml +85 -0
  187. dirigent_examples/shelves/sql/sql-sqlite-roundtrip.yaml +114 -0
  188. dirigent_examples/shelves/sql/warehouse.sql +42 -0
  189. dirigent_examples/shelves/transform/README.md +36 -0
  190. dirigent_examples/shelves/transform/csv-report.yaml +55 -0
  191. dirigent_examples/shelves/transform/jq-filter-and-map.yaml +70 -0
  192. dirigent_examples/shelves/transform/jq-group-and-aggregate.yaml +70 -0
  193. dirigent_examples/shelves/transform/jq-join-two-sources.yaml +98 -0
  194. dirigent_examples/shelves/transform/jq-reshape.yaml +91 -0
  195. dirigent_examples/shelves/transform/jq-stream-through-storage.yaml +112 -0
  196. dirigent_examples/shelves/transform/ndjson-round-trip.yaml +56 -0
  197. dirigent_examples/shelves/transform/parquet-round-trip.yaml +68 -0
  198. dirigent_examples/shelves/transform/std-convert-fan-out.yaml +142 -0
  199. dirigent_examples/shelves/transform/xml-feed-to-ndjson.yaml +116 -0
  200. dirigent_examples/shelves/transform/yaml-config-to-json.yaml +104 -0
  201. dirigent_examples/shelves/triggers/README.md +45 -0
  202. dirigent_examples/shelves/triggers/at-one-time.yaml +78 -0
  203. dirigent_examples/shelves/triggers/cron-nightly.yaml +79 -0
  204. dirigent_examples/shelves/triggers/cron-windowed.yaml +86 -0
  205. dirigent_examples/shelves/triggers/document-nightly.yaml +80 -0
  206. dirigent_examples/shelves/triggers/interval-rolling.yaml +88 -0
  207. dirigent_examples/shelves/triggers/managed-and-manual.yaml +109 -0
  208. dirigent_examples/shelves/triggers/webhook-trigger.yaml +75 -0
  209. dirigent_examples/shelves/validate/README.md +31 -0
  210. dirigent_examples/shelves/validate/expects-a-shape.yaml +56 -0
  211. dirigent_examples/shelves/validate/the-shape-is-wrong.yaml +46 -0
  212. dirigent_examples-0.15.0.dist-info/METADATA +21 -0
  213. dirigent_examples-0.15.0.dist-info/RECORD +216 -0
  214. dirigent_examples-0.15.0.dist-info/WHEEL +4 -0
  215. dirigent_examples-0.15.0.dist-info/entry_points.txt +3 -0
  216. dirigent_examples-0.15.0.dist-info/licenses/LICENSE +18 -0
@@ -0,0 +1,105 @@
1
+ # Copy an object between prefixes, then prove it landed before anything reads it.
2
+ #
3
+ # NEEDS AN S3-COMPATIBLE ENDPOINT, set up exactly as s3-round-trip.yaml documents: the same
4
+ # rustfs container, the same artifacts connection, the same DIRIGENT_STORAGE_CONNECTIONS.
5
+ #
6
+ # The shape is promote-then-consume: a copy from an incoming prefix to a published one,
7
+ # storage.exists standing between the copy and the reader. The sensor is not decoration --
8
+ # on eventually-consistent stores the object's arrival is a fact worth waiting on, and a
9
+ # reader gated on it never sees the half-published state.
10
+ #
11
+ # dg apply examples/s3/s3-copy-and-verify.yaml
12
+ # dg run s3-copy-and-verify -p day=2026-01-01 --watch
13
+ #
14
+ # The first two steps exist so the example carries its own input: a real pipeline starts at
15
+ # seed, with the incoming object written by whatever produced it.
16
+ #
17
+ # Hop by hop:
18
+ #
19
+ # payload the readings, as a value in the run.
20
+ # staged storage.write puts that value in the incoming prefix as json. A value reaches
21
+ # storage here and nowhere else.
22
+ # seed convert.std re-encodes that object into the ndjson a producer would have
23
+ # dropped. A converter is a storage-object operation: it reads one URI and
24
+ # writes another, and the bytes never enter the run.
25
+ # promote the copy from incoming to published.
26
+ # landed the sensor, standing between the copy and the reader.
27
+ # read_back convert.std again, published ndjson back to json, so the result is readable.
28
+
29
+ format: dirigent/v1
30
+ kind: pipeline
31
+ code: s3-copy-and-verify
32
+ name: Copy and verify
33
+ description: Promote an object to a published prefix, and prove it landed before reading it.
34
+
35
+ tags: [s3, sensor, storage, transform]
36
+
37
+ requires:
38
+ blocks:
39
+ - storage.write
40
+ - storage.copy
41
+ - storage.exists
42
+ - value.const
43
+ - convert.std
44
+ connections:
45
+ - artifacts
46
+ storage:
47
+ - s3
48
+
49
+ params:
50
+ type: object
51
+ required: [day]
52
+ additionalProperties: false
53
+ properties:
54
+ day:
55
+ type: string
56
+ format: date
57
+ description: The day whose object is promoted.
58
+ bucket:
59
+ type: string
60
+ default: dirigent-archive
61
+ description: The bucket both prefixes live in.
62
+
63
+ steps:
64
+ payload:
65
+ block: value.const
66
+ config:
67
+ value:
68
+ - { station: st-1, day: "${params.day}", celsius: 4.5 }
69
+ - { station: st-3, day: "${params.day}", celsius: 1.2 }
70
+ staged:
71
+ block: storage.write
72
+ depends_on: [payload]
73
+ config:
74
+ target: s3://${params.bucket}/incoming/${params.day}.json
75
+ value: "${steps.payload.output.value}"
76
+ seed:
77
+ # The example's own input: the incoming object a real producer would have written.
78
+ block: convert.std
79
+ depends_on: [staged]
80
+ config:
81
+ source: "${steps.staged.output.uri}"
82
+ target: s3://${params.bucket}/incoming/${params.day}.ndjson
83
+ from: json
84
+ to: ndjson
85
+ promote:
86
+ block: storage.copy
87
+ depends_on: [seed]
88
+ config:
89
+ source: s3://${params.bucket}/incoming/${params.day}.ndjson
90
+ target: s3://${params.bucket}/published/${params.day}.ndjson
91
+ landed:
92
+ block: storage.exists
93
+ depends_on: [promote]
94
+ poll: 5s
95
+ deadline: 2m
96
+ config:
97
+ uri: s3://${params.bucket}/published/${params.day}.ndjson
98
+ read_back:
99
+ block: convert.std
100
+ depends_on: [landed]
101
+ config:
102
+ source: "${steps.landed.output.uri}"
103
+ target: s3://${params.bucket}/published/${params.day}.json
104
+ from: ndjson
105
+ to: json
@@ -0,0 +1,87 @@
1
+ # A csv report delivered to a bucket, where the people who asked for it can fetch it.
2
+ #
3
+ # NEEDS AN S3-COMPATIBLE ENDPOINT, set up exactly as s3-round-trip.yaml documents: the same
4
+ # rustfs container, the same artifacts connection, the same DIRIGENT_STORAGE_CONNECTIONS.
5
+ #
6
+ # This is examples/transform/csv-report.yaml with a real destination: the same records, the
7
+ # same flattening, the same codec -- and s3:// instead of the run's scratch, so the file
8
+ # outlives the run and lands where a person or another system reads it. Changing the
9
+ # destination was one word, which is the point of schemes.
10
+ #
11
+ # dg apply examples/s3/s3-csv-report.yaml
12
+ # dg run s3-csv-report -p day=2026-01-01 --watch
13
+ #
14
+ # The report is at s3://<bucket>/reports/<day>.csv when the run settles.
15
+ #
16
+ # Hop by hop:
17
+ #
18
+ # readings the records, as a value in the run.
19
+ # rows jq flattens them into the shape a csv can hold. Still a value.
20
+ # staged storage.write puts the rows in the bucket as json, because a converter works
21
+ # on objects rather than on values.
22
+ # report convert.std re-encodes that object as the csv people fetch.
23
+
24
+ format: dirigent/v1
25
+ kind: pipeline
26
+ code: s3-csv-report
27
+ name: A report delivered to a bucket
28
+ description: Shape records into flat rows and write them as a csv object in S3.
29
+
30
+ tags: [s3, storage, transform]
31
+
32
+ requires:
33
+ blocks:
34
+ - value.const
35
+ - transform.jq
36
+ - storage.write
37
+ - convert.std
38
+ connections:
39
+ - artifacts
40
+ storage:
41
+ - s3
42
+
43
+ params:
44
+ type: object
45
+ required: [day]
46
+ additionalProperties: false
47
+ properties:
48
+ day:
49
+ type: string
50
+ format: date
51
+ description: The day the report is about, and the object it is named as.
52
+ bucket:
53
+ type: string
54
+ default: dirigent-archive
55
+ description: The bucket the report is delivered into.
56
+
57
+ steps:
58
+ readings:
59
+ block: value.const
60
+ config:
61
+ value:
62
+ - { station: st-1, region: east, celsius: 4.5, tags: [ok, new] }
63
+ - { station: st-3, region: west, celsius: 1.2, tags: [ok] }
64
+ - { station: st-4, region: north, celsius: -3.4, tags: [] }
65
+ rows:
66
+ block: transform.jq
67
+ depends_on: [readings]
68
+ config:
69
+ input: ${steps.readings.output.value}
70
+ # Flatten for the codec: the tag list becomes one joined cell, because only this
71
+ # program knows what the separator should mean.
72
+ program: |
73
+ [.[] | {station, region, celsius, tags: (.tags | join(" "))}]
74
+ staged:
75
+ block: storage.write
76
+ depends_on: [rows]
77
+ config:
78
+ target: s3://${params.bucket}/reports/${params.day}.json
79
+ value: "${steps.rows.output.value}"
80
+ report:
81
+ block: convert.std
82
+ depends_on: [staged]
83
+ config:
84
+ source: "${steps.staged.output.uri}"
85
+ target: s3://${params.bucket}/reports/${params.day}.csv
86
+ from: json
87
+ to: csv
@@ -0,0 +1,77 @@
1
+ # A parquet dataset delivered to a bucket, with its types along for the ride.
2
+ #
3
+ # NEEDS AN S3-COMPATIBLE ENDPOINT, set up exactly as s3-round-trip.yaml documents: the same
4
+ # rustfs container, the same artifacts connection, the same DIRIGENT_STORAGE_CONNECTIONS.
5
+ #
6
+ # This is examples/transform/parquet-round-trip.yaml with a real destination. Parquet is
7
+ # the format a downstream analysis actually wants -- columnar, typed, compact -- and what
8
+ # lands in the bucket is a dataset a reader opens with pandas or duckdb, not a csv that
9
+ # lost its numbers on the way out. convert.arrow ships in the dirigent-parquet pack;
10
+ # installing it puts the block in the catalog.
11
+ #
12
+ # dg apply examples/s3/s3-parquet-report.yaml
13
+ # dg run s3-parquet-report -p day=2026-01-01 --watch
14
+ #
15
+ # The dataset is at s3://<bucket>/datasets/<day>.parquet when the run settles.
16
+ #
17
+ # Hop by hop:
18
+ #
19
+ # readings the records, as a value in the run, flat as parquet columns need.
20
+ # staged storage.write puts them in the bucket as json, which is the form a converter
21
+ # can read: convert.arrow works on objects, never on a value in the run.
22
+ # dataset convert.arrow re-encodes that object as parquet, types and all.
23
+
24
+ format: dirigent/v1
25
+ kind: pipeline
26
+ code: s3-parquet-report
27
+ name: A parquet dataset in a bucket
28
+ description: Write typed records as a parquet object in S3, where an analysis can read them.
29
+
30
+ tags: [s3, storage, transform]
31
+
32
+ requires:
33
+ blocks:
34
+ - value.const
35
+ - storage.write
36
+ - convert.arrow
37
+ connections:
38
+ - artifacts
39
+ storage:
40
+ - s3
41
+
42
+ params:
43
+ type: object
44
+ required: [day]
45
+ additionalProperties: false
46
+ properties:
47
+ day:
48
+ type: string
49
+ format: date
50
+ description: The day the dataset covers, and the object it is named as.
51
+ bucket:
52
+ type: string
53
+ default: dirigent-archive
54
+ description: The bucket the dataset is delivered into.
55
+
56
+ steps:
57
+ readings:
58
+ block: value.const
59
+ config:
60
+ value:
61
+ - { station: st-1, region: east, celsius: 4.5, active: true }
62
+ - { station: st-3, region: west, celsius: 1.2, active: true }
63
+ - { station: st-4, region: north, celsius: -3.4, active: false }
64
+ staged:
65
+ block: storage.write
66
+ depends_on: [readings]
67
+ config:
68
+ target: s3://${params.bucket}/datasets/${params.day}.json
69
+ value: "${steps.readings.output.value}"
70
+ dataset:
71
+ block: convert.arrow
72
+ depends_on: [staged]
73
+ config:
74
+ source: "${steps.staged.output.uri}"
75
+ target: s3://${params.bucket}/datasets/${params.day}.parquet
76
+ from: json
77
+ to: parquet
@@ -0,0 +1,125 @@
1
+ # Moving an artifact out to object storage and back, with no S3-specific block anywhere.
2
+ #
3
+ # NEEDS AN S3-COMPATIBLE ENDPOINT. dirigent-storage-s3 is installed and every block below
4
+ # exists; what is missing is a credential and a server to point it at.
5
+ #
6
+ # ON THE COMPOSE STACK there is nothing to do: the stack runs an S3 server, keeps artifacts
7
+ # in its bucket, and bootstraps the `artifacts` connection this document names. Apply and
8
+ # run, passing the stack's own bucket:
9
+ #
10
+ # dg run s3-round-trip -p day=2026-01-01 -p bucket=dirigent --watch
11
+ #
12
+ # The steps below are for `dg dev`, which has neither.
13
+ #
14
+ # A storage backend is not a step. There is no "upload to S3" block and no "download from
15
+ # S3" block: s3:// is a registered URI scheme, so the ordinary storage.copy moves bytes
16
+ # between file:// and s3:// in either direction, and the ordinary storage.exists waits on
17
+ # an s3:// object. Changing where the archive lives is one word.
18
+ #
19
+ # 1. Start an S3-compatible server. Any implementation will do; this is the one dirigent's
20
+ # own storage lane runs against:
21
+ #
22
+ # docker run -d --name dirigent-s3 -p 9000:9000 \
23
+ # -e RUSTFS_ACCESS_KEY=dirigent-test-key \
24
+ # -e RUSTFS_SECRET_KEY=dirigent-test-secret \
25
+ # rustfs/rustfs:1.0.0-rc.4
26
+ #
27
+ # 2. Create the connection, and tell the instance that it is what serves s3://. The first
28
+ # is the credential; the second is which of possibly several s3 connections the scheme
29
+ # is configured from:
30
+ #
31
+ # dg connection create s3 artifacts \
32
+ # --set endpoint_url=http://127.0.0.1:9000 \
33
+ # --set access_key_id=dirigent-test-key \
34
+ # --set secret_access_key=dirigent-test-secret \
35
+ # --set path_style=true \
36
+ # --set bucket=dirigent-archive
37
+ #
38
+ # export DIRIGENT_STORAGE_CONNECTIONS='{"s3": "artifacts"}'
39
+ #
40
+ # 3. Create the bucket the run addresses, then apply and run:
41
+ #
42
+ # dg apply examples/s3/s3-round-trip.yaml
43
+ # dg run s3-round-trip -p day=2026-01-01 --watch
44
+ #
45
+ # The bucket always comes from the URI, never from the connection: the connection's own
46
+ # bucket field is only what `dg connection check artifacts` probes.
47
+ #
48
+ # A real pipeline stops after the first copy. The download here makes the result checkable
49
+ # without a second system to look at.
50
+
51
+ format: dirigent/v1
52
+ kind: pipeline
53
+ code: s3-round-trip
54
+ name: S3 round trip
55
+ description: Move an artifact from local storage out to S3 and back, then confirm it landed.
56
+
57
+ tags: [s3, http, sensor, storage]
58
+
59
+ requires:
60
+ blocks:
61
+ - http.request
62
+ - storage.write
63
+ - storage.copy
64
+ - storage.exists
65
+ connections:
66
+ - artifacts
67
+ storage:
68
+ - s3
69
+
70
+ params:
71
+ type: object
72
+ required: [day]
73
+ additionalProperties: false
74
+ properties:
75
+ day:
76
+ type: string
77
+ format: date
78
+ description: The day whose artifact is round-tripped.
79
+ bucket:
80
+ type: string
81
+ default: dirigent-archive
82
+ description: The bucket the artifact is archived into.
83
+
84
+ steps:
85
+ produce:
86
+ block: http.request
87
+ config:
88
+ url: https://postman-echo.com/post
89
+ method: POST
90
+ body:
91
+ day: "${params.day}"
92
+
93
+ stage:
94
+ block: storage.write
95
+ depends_on: [produce]
96
+ # The local artifact the round trip is about, written on file:// because that is where
97
+ # ${run.scratch} lives. The copy below is the only line that mentions s3.
98
+ config:
99
+ target: "${run.scratch}/local/${params.day}.json"
100
+ value: "${steps.produce.output.body}"
101
+
102
+ upload:
103
+ block: storage.copy
104
+ depends_on: [stage]
105
+ config:
106
+ source: "${steps.stage.output.uri}"
107
+ target: "s3://${params.bucket}/daily/${params.day}.json"
108
+
109
+ wait_for_object:
110
+ block: storage.exists
111
+ depends_on: [upload]
112
+ # S3 acknowledges a write before every reader can see it, so the upload settling is not
113
+ # the same as the object being there. This waits for the second thing.
114
+ poll: 2s
115
+ deadline: 1m
116
+ config:
117
+ uri: "s3://${params.bucket}/daily/${params.day}.json"
118
+ min_size: 1
119
+
120
+ download:
121
+ block: storage.copy
122
+ depends_on: [wait_for_object]
123
+ config:
124
+ source: "${steps.wait_for_object.output.uri}"
125
+ target: "${run.scratch}/roundtrip/${params.day}.json"
@@ -0,0 +1,36 @@
1
+ # Example schemas
2
+
3
+ Each file here is a plain JSON Schema (Draft 2020-12): the shape a read is expected to return,
4
+ written down once so a pipeline can be refused the moment a payload moves out from under it. A
5
+ schema is **locally authored** -- a picture you hold of the payload, never something fetched or
6
+ introspected from the source. Each is applied on its own with `dg schema create`, or with the
7
+ directory it sits in, and a document that references one names it in `requires.schemas`. A document may instead **carry** a
8
+ schema in its own top-level `schemas:` section for a standalone or `--local` run; a server
9
+ refuses a document that carries one, so a shared instance holds its schemas here.
10
+
11
+ Apply one the way a person would:
12
+
13
+ ```bash
14
+ dg schema create examples/schemas/ou-record.json
15
+ ```
16
+
17
+ A directory apply lands them the same way: a server pointed at a directory stores every schema
18
+ it finds there before the pipelines, so a mounted corpus needs no separate step.
19
+
20
+ A schema carries its own identity in its keywords, so there is nothing else to pass:
21
+
22
+ - `$id` becomes the `code` the schema is addressed by (falling back to the file's stem).
23
+ - `title` becomes its name.
24
+ - `description` becomes its body.
25
+
26
+ | File | The shape it pins |
27
+ | --- | --- |
28
+ | [ou-record.json](ou-record.json) | A single organisation-unit record: a string `id`, a `name`, and an integer `level` |
29
+ | [echo-reading.json](echo-reading.json) | What a fetch of the echo service answers with: an `args` map carrying `station` and `day`, and the `url` |
30
+ | [station-reading.json](station-reading.json) | The row a transform is held to before it is posted on: a `station`, a `day`, and a numeric `celsius` |
31
+
32
+ `ou-record` is the shape the two [validation documents](../validate) check a payload against, and
33
+ the two reading shapes are the gates either side of the transform in
34
+ [http-fetch-validate-post.yaml](../recipes/http-fetch-validate-post.yaml).
35
+ The [JSON Schema guide](../../docs/json-schema.md) walks through how a schema like it is
36
+ built, keyword by keyword.
@@ -0,0 +1,18 @@
1
+ {
2
+ "$id": "echo-reading",
3
+ "title": "Echoed reading batch",
4
+ "description": "What a fetch of the echo service is held to before anything reads into it: an `args` map carrying the query the call sent, and the `url` it was answered from. A source that changes shape under a pipeline is refused here rather than producing a wrong number three steps later.",
5
+ "type": "object",
6
+ "required": ["args", "url"],
7
+ "properties": {
8
+ "args": {
9
+ "type": "object",
10
+ "required": ["station", "day"],
11
+ "properties": {
12
+ "station": { "type": "string", "minLength": 1 },
13
+ "day": { "type": "string", "format": "date" }
14
+ }
15
+ },
16
+ "url": { "type": "string" }
17
+ }
18
+ }
@@ -0,0 +1,13 @@
1
+ {
2
+ "$id": "ou-record",
3
+ "title": "Organisation unit record",
4
+ "description": "A single organisation-unit record: a string id, a name, and an integer level. The shape a payload is held to before a downstream step reads any of the three.",
5
+ "type": "object",
6
+ "required": ["id", "name", "level"],
7
+ "additionalProperties": false,
8
+ "properties": {
9
+ "id": { "type": "string" },
10
+ "name": { "type": "string" },
11
+ "level": { "type": "integer" }
12
+ }
13
+ }
@@ -0,0 +1,13 @@
1
+ {
2
+ "$id": "station-reading",
3
+ "title": "Station reading",
4
+ "description": "The row a transform is held to before it is posted onward: a station, the day it reports for, and a temperature that is a number rather than the string a query carried. The gate a document puts after its own jq program, so a typo in the program is refused instead of delivered.",
5
+ "type": "object",
6
+ "required": ["station", "day", "celsius"],
7
+ "additionalProperties": false,
8
+ "properties": {
9
+ "station": { "type": "string", "minLength": 1 },
10
+ "day": { "type": "string", "format": "date" },
11
+ "celsius": { "type": "number" }
12
+ }
13
+ }
@@ -0,0 +1,16 @@
1
+ # Sensor examples
2
+
3
+ Steps that wait for the world instead of doing something to it: each poke is one cheap,
4
+ read-only question, the run holds no worker slot between pokes, and `deadline` with
5
+ `on_timeout` says what a day without an answer means.
6
+
7
+ ```bash
8
+ dg run --local examples/sensors/time-window.yaml --enable-unsafe shell.run
9
+ ```
10
+
11
+ ## Pipelines
12
+
13
+ | File | What it teaches |
14
+ | --- | --- |
15
+ | [sensor-gate.yaml](sensor-gate.yaml) | `storage.exists` gating a load: wait for today's drop, and skip the day if it never lands. |
16
+ | [time-window.yaml](time-window.yaml) | The clock as a condition: hold until the local time is inside a window, in the window's own zone. |
@@ -0,0 +1,65 @@
1
+ # A sensor gating the rest of the pipeline, and what happens when it never fires.
2
+ #
3
+ # storage.exists is a sensor: it observes the world and changes nothing. Each poke is a
4
+ # durable, scheduled poll, so waiting hours costs hours of rows, not hours of a held worker.
5
+ #
6
+ # poll, deadline, and on_timeout are step-level engine semantics, uniform across every
7
+ # block, never buried in the block's own config. on_timeout: skip is the load-bearing
8
+ # choice here: no drop today is not a failure, it is a day with nothing to load, so the
9
+ # sensor skips and the default all_success edges skip the rest of the branch with it.
10
+ #
11
+ # The deadline is short so the skip is watchable. Nothing ever writes the drop, so after ten
12
+ # seconds the sensor skips, the branch behind it skips, and the run succeeds with nothing
13
+ # done:
14
+ # dg run --local examples/sensors/sensor-gate.yaml -p day=2026-01-01 --enable-unsafe shell.run
15
+ #
16
+ # Note where the drop is looked for. file:// is rooted at the instance's artifact root and
17
+ # refuses to address anything outside it, so a stored pipeline cannot turn a copy step into
18
+ # an arbitrary-file read. Pointing this at a real drop means an s3:// URI, or a path under
19
+ # that root -- never an absolute path from the document.
20
+ #
21
+ # confirm uses shell.run, which executes code on the worker and needs the allowlist:
22
+ # export DIRIGENT_ENABLED_UNSAFE_BLOCKS='["shell.run"]'
23
+
24
+ format: dirigent/v1
25
+ kind: pipeline
26
+ code: sensor-gate
27
+ name: A sensor gates the run
28
+ description: Wait for today's drop to land, then load it; skip the day if it never does.
29
+
30
+ tags: [sensor, execute, storage]
31
+
32
+ params:
33
+ type: object
34
+ required: [day]
35
+ properties:
36
+ day:
37
+ type: string
38
+ format: date
39
+
40
+ steps:
41
+ wait_for_drop:
42
+ block: storage.exists
43
+ # Each poke is a scheduled row, not a held worker, so the poll interval is chosen for
44
+ # how soon the drop matters rather than for what it costs to wait.
45
+ poll: 2s
46
+ deadline: 10s
47
+ # No drop today is a day with nothing to load, not a failure. The skip carries down the
48
+ # branch through the default all_success edges, and the run still succeeds.
49
+ on_timeout: skip
50
+ config:
51
+ uri: "${run.scratch}/drops/${params.day}.parquet"
52
+ min_size: 1
53
+
54
+ load:
55
+ block: storage.copy
56
+ depends_on: [wait_for_drop]
57
+ config:
58
+ source: "${steps.wait_for_drop.output.uri}"
59
+ target: "${run.scratch}/loaded/${params.day}.parquet"
60
+
61
+ confirm:
62
+ block: shell.run
63
+ depends_on: [load]
64
+ config:
65
+ argv: [echo, "loaded ${steps.wait_for_drop.output.size} bytes for ${params.day}"]
@@ -0,0 +1,61 @@
1
+ # Gating a step on the wall clock, in a timezone the document names out loud.
2
+ #
3
+ # A schedule says when a run starts. time.window says when a step may proceed, which is a
4
+ # different question: "load whenever the export lands, but never touch the warehouse outside
5
+ # the maintenance window" is a property of the step, not of the trigger.
6
+ #
7
+ # It is a sensor, so it inherits the whole waiting apparatus: poll, deadline, and on_timeout
8
+ # are step-level engine semantics here exactly as they are for storage.exists. A closed
9
+ # window parks until it opens rather than polling on the step's cadence, so waiting out a
10
+ # Sunday costs a handful of pokes.
11
+ #
12
+ # The timezone defaults to UTC and should always be stated. Europe/Oslo below is read as
13
+ # Oslo's clock, daylight saving included, no matter where the worker runs.
14
+ #
15
+ # The window here is wide open -- the whole day, every day -- so the example runs without
16
+ # waiting. Narrow it to see the wait:
17
+ # dg run --local examples/sensors/time-window.yaml --enable-unsafe shell.run
18
+ #
19
+ # A real maintenance window is the commented shape below the open one:
20
+ # after: "01:00"
21
+ # before: "04:00"
22
+ # days: [sat, sun]
23
+ #
24
+ # Windows may cross midnight: after 22:00 before 04:00 is one night window, and when days is
25
+ # set it names the day the window opens on, so [fri] runs Friday evening into Saturday.
26
+
27
+ format: dirigent/v1
28
+ kind: pipeline
29
+ code: time-window
30
+ name: Inside a time window
31
+ description: Wait until the local clock is inside a window, then do the work it guards.
32
+
33
+ tags: [sensor, execute, window]
34
+
35
+ requires:
36
+ blocks:
37
+ - time.window
38
+ - shell.run
39
+
40
+ steps:
41
+ maintenance_window:
42
+ block: time.window
43
+ poll: 5s
44
+ deadline: 1m
45
+ on_timeout: skip
46
+ config:
47
+ # Open all day so the example runs without waiting; a real window is narrow, and the
48
+ # commented shape in the header is what one looks like.
49
+ after: "00:00"
50
+ before: "23:59"
51
+ # Always state it. The default is UTC, and a window meant as local time silently
52
+ # becomes the wrong hours wherever the worker happens to run.
53
+ timezone: Europe/Oslo
54
+
55
+ reindex:
56
+ block: shell.run
57
+ depends_on: [maintenance_window]
58
+ config:
59
+ argv:
60
+ - echo
61
+ - "the window opened at ${steps.maintenance_window.output.entered_at} (${steps.maintenance_window.output.timezone})"
@@ -0,0 +1,52 @@
1
+ # SQL examples
2
+
3
+ The `sql.*` family reads from and writes to a database. `sql.query` runs one statement and
4
+ hands its rows on; `sql.execute` runs a list of statements as one transaction.
5
+ [docs/sql.md](../../docs/sql.md) is the family's home.
6
+
7
+ Both are **ordinary** blocks: they run no command a document supplies and reach nothing but the
8
+ database their connection names, so no id has to be allowlisted to run them.
9
+
10
+ ```bash
11
+ dg run --local examples/sql/sql-sqlite-roundtrip.yaml
12
+ dg run --local examples/sql/sql-query-to-storage.yaml
13
+ ```
14
+
15
+ Those two need nothing at all -- no network, no daemon, no server. Each builds a SQLite
16
+ database in the run's own work directory, so the whole example is self-contained. The DuckDB one
17
+ needs the engine (`dirigent-blocks[duckdb]`) and the parquet pack, and nothing else; the last
18
+ names a real PostgreSQL and says so in its own header.
19
+
20
+ ```bash
21
+ dg run --local examples/sql/duckdb-parquet-to-report.yaml --keep
22
+ ```
23
+
24
+ The rule the whole family turns on: **a value is bound, never interpolated**. The statement is
25
+ a constant in the document, every value is a named `:parameter`, and a `${...}` reference
26
+ resolves into `params` and never into the SQL text. A parameter that reads as SQL is compared
27
+ as a string and matches nothing.
28
+
29
+ The first two documents carry their `sql` connection in a `connections:` section, because a
30
+ `--local` run has no instance to hold one. On a server the connection is created once and the
31
+ document names it, with the password set as its own sealed field:
32
+
33
+ ```bash
34
+ dg connection create sql warehouse-read \
35
+ --set url=postgresql+asyncpg://reader@db.example:5432/warehouse \
36
+ --set password=... \
37
+ --set read_only=true
38
+ ```
39
+
40
+ On the compose stack that database is the `infra/compose.sql.yaml` overlay
41
+ (`make docker-run-sql`), seeded with the `reading` table the third document queries, and the
42
+ connection names it as `reader@warehouse:5432/warehouse`.
43
+
44
+ ## Pipelines
45
+
46
+ | File | What it teaches |
47
+ | --- | --- |
48
+ | [sql-sqlite-roundtrip.yaml](sql-sqlite-roundtrip.yaml) | The family end to end: a table created and filled in one transaction, read back with a bound parameter, and the rows referenced by a downstream step. |
49
+ | [sql-query-to-storage.yaml](sql-query-to-storage.yaml) | Rows into a file: the query hands them on as its output, `storage.write` puts them in an object, and `max_rows` bounds what the step will carry. |
50
+ | [duckdb-parquet-to-report.yaml](duckdb-parquet-to-report.yaml) | The engine that reads files: a parquet artifact queried by DuckDB through a bound `read_parquet(:source)`, and the answer copied out as a csv artifact. Needs the `dirigent-blocks[duckdb]` extra and the parquet pack. |
51
+ | [sql-postgres-readonly.yaml](sql-postgres-readonly.yaml) | A referenced PostgreSQL connection with `read_only: true` and its password sealed separately, queried inside the run's window. Validates offline; it cannot run without an instance. |
52
+ | [warehouse.sql](warehouse.sql) | Not a pipeline: the roles and the one table the compose stack's warehouse is seeded with, mounted by `infra/compose.sql.yaml` |