dirigent-examples 0.15.0__py3-none-any.whl
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- dirigent_examples/__init__.py +22 -0
- dirigent_examples/py.typed +0 -0
- dirigent_examples/shelves/README.md +299 -0
- dirigent_examples/shelves/composition/README.md +18 -0
- dirigent_examples/shelves/composition/chained-instances.yaml +97 -0
- dirigent_examples/shelves/composition/composition-child.yaml +64 -0
- dirigent_examples/shelves/composition/composition-parent.yaml +89 -0
- dirigent_examples/shelves/connections.yaml +52 -0
- dirigent_examples/shelves/demo/README.md +19 -0
- dirigent_examples/shelves/demo/markdown-showcase.yaml +117 -0
- dirigent_examples/shelves/demo/params-showcase.yaml +92 -0
- dirigent_examples/shelves/demo/requires.yaml +65 -0
- dirigent_examples/shelves/demo/weekly-import-malawi.yaml +54 -0
- dirigent_examples/shelves/demo/weekly-import-nepal.yaml +75 -0
- dirigent_examples/shelves/docker/README.md +29 -0
- dirigent_examples/shelves/docker/docker-build-push.yaml +110 -0
- dirigent_examples/shelves/docker/docker-build-run.yaml +106 -0
- dirigent_examples/shelves/docker/docker-compose-database.yaml +124 -0
- dirigent_examples/shelves/docker/docker-compose-failing-up.yaml +83 -0
- dirigent_examples/shelves/docker/docker-compose-file.yaml +117 -0
- dirigent_examples/shelves/docker/docker-compose-profiles-env.yaml +133 -0
- dirigent_examples/shelves/docker/docker-compose-stack.yaml +70 -0
- dirigent_examples/shelves/docker/docker-hello.yaml +53 -0
- dirigent_examples/shelves/docker/docker-remote-daemon.yaml +92 -0
- dirigent_examples/shelves/docker/docker-run-failing-teardown.yaml +94 -0
- dirigent_examples/shelves/docker/docker-ticker.yaml +49 -0
- dirigent_examples/shelves/execute/README.md +16 -0
- dirigent_examples/shelves/execute/long-log.yaml +89 -0
- dirigent_examples/shelves/failure/README.md +20 -0
- dirigent_examples/shelves/failure/error-handler.yaml +89 -0
- dirigent_examples/shelves/failure/optional-step.yaml +82 -0
- dirigent_examples/shelves/failure/retries.yaml +75 -0
- dirigent_examples/shelves/failure/retry-budget.yaml +82 -0
- dirigent_examples/shelves/failure/step-timeout.yaml +96 -0
- dirigent_examples/shelves/git/README.md +32 -0
- dirigent_examples/shelves/git/git-checkout-build.yaml +125 -0
- dirigent_examples/shelves/git/git-checkout-compose.yaml +138 -0
- dirigent_examples/shelves/git/git-checkout-public.yaml +84 -0
- dirigent_examples/shelves/graph/README.md +22 -0
- dirigent_examples/shelves/graph/deep-chain.yaml +119 -0
- dirigent_examples/shelves/graph/fan-in.yaml +80 -0
- dirigent_examples/shelves/graph/fan-out.yaml +66 -0
- dirigent_examples/shelves/graph/linear.yaml +66 -0
- dirigent_examples/shelves/graph/parallel-branches.yaml +57 -0
- dirigent_examples/shelves/graph/parallel-sleep.yaml +55 -0
- dirigent_examples/shelves/graph/skip-diamond.yaml +105 -0
- dirigent_examples/shelves/graph/wide-fan.yaml +147 -0
- dirigent_examples/shelves/hello-world.yaml +30 -0
- dirigent_examples/shelves/open-data/README.md +67 -0
- dirigent_examples/shelves/open-data/feeds-composition.yaml +120 -0
- dirigent_examples/shelves/open-data/gdacs-disaster-updates.yaml +230 -0
- dirigent_examples/shelves/open-data/github-releases-relay.yaml +225 -0
- dirigent_examples/shelves/open-data/hdx-dataset-watch.yaml +211 -0
- dirigent_examples/shelves/open-data/kobo-submissions-to-csv.yaml +149 -0
- dirigent_examples/shelves/open-data/nominatim-geocode-facilities.yaml +169 -0
- dirigent_examples/shelves/open-data/odk-central-submissions.yaml +146 -0
- dirigent_examples/shelves/open-data/open-meteo-weekly-report.yaml +142 -0
- dirigent_examples/shelves/open-data/overpass-health-facilities.yaml +154 -0
- dirigent_examples/shelves/open-data/usgs-earthquakes-alert.yaml +208 -0
- dirigent_examples/shelves/open-data/who-gho-indicators-to-parquet.yaml +163 -0
- dirigent_examples/shelves/open-data/wikidata-country-reference.yaml +146 -0
- dirigent_examples/shelves/open-data/world-bank-population-trend.yaml +166 -0
- dirigent_examples/shelves/patterns/README.md +144 -0
- dirigent_examples/shelves/patterns/concurrency-queue.yaml +89 -0
- dirigent_examples/shelves/patterns/concurrency-replace.yaml +92 -0
- dirigent_examples/shelves/patterns/concurrency-skip.yaml +97 -0
- dirigent_examples/shelves/patterns/connections-referenced-vs-carried.yaml +140 -0
- dirigent_examples/shelves/patterns/deadline-on-a-sensor.yaml +106 -0
- dirigent_examples/shelves/patterns/fan-out-continue.yaml +88 -0
- dirigent_examples/shelves/patterns/fan-out-fail-fast.yaml +80 -0
- dirigent_examples/shelves/patterns/fan-out-from-params.yaml +84 -0
- dirigent_examples/shelves/patterns/fan-out-item-wise.yaml +111 -0
- dirigent_examples/shelves/patterns/fan-out-literal-list.yaml +81 -0
- dirigent_examples/shelves/patterns/fan-out-nested-objects.yaml +107 -0
- dirigent_examples/shelves/patterns/fan-out-then-join.yaml +86 -0
- dirigent_examples/shelves/patterns/log-levels.yaml +119 -0
- dirigent_examples/shelves/patterns/outputs-inline-vs-storage.yaml +140 -0
- dirigent_examples/shelves/patterns/params-every-type.yaml +259 -0
- dirigent_examples/shelves/patterns/params-validation-refuses.yaml +131 -0
- dirigent_examples/shelves/patterns/pipeline-run-child.yaml +96 -0
- dirigent_examples/shelves/patterns/pipeline-run-fire-and-forget.yaml +99 -0
- dirigent_examples/shelves/patterns/pipeline-run-strict.yaml +103 -0
- dirigent_examples/shelves/patterns/pipeline-run-wait.yaml +107 -0
- dirigent_examples/shelves/patterns/pipeline-run-with-params.yaml +124 -0
- dirigent_examples/shelves/patterns/poll-cadence.yaml +102 -0
- dirigent_examples/shelves/patterns/priority-layered.yaml +120 -0
- dirigent_examples/shelves/patterns/references-cheat-sheet.yaml +186 -0
- dirigent_examples/shelves/patterns/retry-budget-exhausted.yaml +92 -0
- dirigent_examples/shelves/patterns/retry-exponential-backoff.yaml +88 -0
- dirigent_examples/shelves/patterns/retry-only-transient.yaml +108 -0
- dirigent_examples/shelves/patterns/retry-with-jitter.yaml +102 -0
- dirigent_examples/shelves/patterns/rule-all-done.yaml +83 -0
- dirigent_examples/shelves/patterns/rule-all-success.yaml +79 -0
- dirigent_examples/shelves/patterns/rule-always.yaml +92 -0
- dirigent_examples/shelves/patterns/rule-one-failed.yaml +86 -0
- dirigent_examples/shelves/patterns/schedule-at-once.yaml +110 -0
- dirigent_examples/shelves/patterns/schedule-cron-timezone.yaml +121 -0
- dirigent_examples/shelves/patterns/schedule-interval.yaml +109 -0
- dirigent_examples/shelves/patterns/schedule-window-half-open.yaml +105 -0
- dirigent_examples/shelves/patterns/sensor-http-ready.yaml +121 -0
- dirigent_examples/shelves/patterns/sensor-storage-exists.yaml +134 -0
- dirigent_examples/shelves/patterns/step-names-and-keys.yaml +99 -0
- dirigent_examples/shelves/patterns/timeout-fails-the-step.yaml +94 -0
- dirigent_examples/shelves/patterns/timeout-skips-the-step.yaml +102 -0
- dirigent_examples/shelves/patterns/webhook-mapping-nested-payload.yaml +125 -0
- dirigent_examples/shelves/patterns/webhook-signed.yaml +144 -0
- dirigent_examples/shelves/preview/s3-parquet-to-ingestion.yaml +92 -0
- dirigent_examples/shelves/python/README.md +31 -0
- dirigent_examples/shelves/python/apply_and_run.py +52 -0
- dirigent_examples/shelves/python/ci_gate.py +76 -0
- dirigent_examples/shelves/python/connections.py +61 -0
- dirigent_examples/shelves/python/error_handling.py +84 -0
- dirigent_examples/shelves/python/follow_logs.py +39 -0
- dirigent_examples/shelves/python/list_and_filter.py +52 -0
- dirigent_examples/shelves/queues/README.md +59 -0
- dirigent_examples/shelves/queues/kafka-consume-then-transform.yaml +105 -0
- dirigent_examples/shelves/queues/kafka-produce-then-consume.yaml +124 -0
- dirigent_examples/shelves/queues/rabbitmq-consume-ack-on-success.yaml +102 -0
- dirigent_examples/shelves/queues/report-to-kafka.yaml +105 -0
- dirigent_examples/shelves/queues/report-to-rabbitmq.yaml +105 -0
- dirigent_examples/shelves/recipes/README.md +130 -0
- dirigent_examples/shelves/recipes/csv-header-rules.yaml +147 -0
- dirigent_examples/shelves/recipes/csv-to-ndjson.yaml +107 -0
- dirigent_examples/shelves/recipes/etl-csv-clean-validate-parquet.yaml +207 -0
- dirigent_examples/shelves/recipes/filter-by-predicate.yaml +116 -0
- dirigent_examples/shelves/recipes/filter-then-map-then-reduce.yaml +109 -0
- dirigent_examples/shelves/recipes/http-fetch-validate-post.yaml +142 -0
- dirigent_examples/shelves/recipes/http-follow-redirects.yaml +100 -0
- dirigent_examples/shelves/recipes/http-get-with-query.yaml +96 -0
- dirigent_examples/shelves/recipes/http-headers-and-auth-connection.yaml +110 -0
- dirigent_examples/shelves/recipes/http-post-file-from-storage.yaml +117 -0
- dirigent_examples/shelves/recipes/http-post-json-echo.yaml +105 -0
- dirigent_examples/shelves/recipes/http-post-report.yaml +183 -0
- dirigent_examples/shelves/recipes/http-save-body-to-storage.yaml +106 -0
- dirigent_examples/shelves/recipes/http-success-status-list.yaml +80 -0
- dirigent_examples/shelves/recipes/http-timeout-override.yaml +104 -0
- dirigent_examples/shelves/recipes/jq-dedupe-by-key.yaml +78 -0
- dirigent_examples/shelves/recipes/jq-defaults-and-nulls.yaml +91 -0
- dirigent_examples/shelves/recipes/jq-group-by-and-sum.yaml +74 -0
- dirigent_examples/shelves/recipes/jq-join-two-lists.yaml +77 -0
- dirigent_examples/shelves/recipes/jq-long-to-wide.yaml +76 -0
- dirigent_examples/shelves/recipes/jq-nested-to-flat.yaml +89 -0
- dirigent_examples/shelves/recipes/jq-pivot-wide-to-long.yaml +65 -0
- dirigent_examples/shelves/recipes/jq-running-totals.yaml +82 -0
- dirigent_examples/shelves/recipes/jq-string-cleaning.yaml +88 -0
- dirigent_examples/shelves/recipes/jq-top-n.yaml +85 -0
- dirigent_examples/shelves/recipes/jq-validate-in-jq-vs-schema.yaml +124 -0
- dirigent_examples/shelves/recipes/jq-window-dates.yaml +82 -0
- dirigent_examples/shelves/recipes/json-to-csv-flattening.yaml +141 -0
- dirigent_examples/shelves/recipes/large-output-to-storage.yaml +134 -0
- dirigent_examples/shelves/recipes/map-enrich-with-lookup.yaml +96 -0
- dirigent_examples/shelves/recipes/ndjson-to-parquet.yaml +130 -0
- dirigent_examples/shelves/recipes/pagination-by-fan-out.yaml +115 -0
- dirigent_examples/shelves/recipes/parquet-round-trip-types.yaml +163 -0
- dirigent_examples/shelves/recipes/reconcile-two-sources.yaml +159 -0
- dirigent_examples/shelves/recipes/report-built-in.yaml +72 -0
- dirigent_examples/shelves/recipes/report-daily-digest.yaml +186 -0
- dirigent_examples/shelves/recipes/report-to-file.yaml +131 -0
- dirigent_examples/shelves/recipes/report-to-webhook.yaml +136 -0
- dirigent_examples/shelves/recipes/schema-carried.yaml +112 -0
- dirigent_examples/shelves/recipes/schema-formats.yaml +107 -0
- dirigent_examples/shelves/recipes/schema-referenced.yaml +86 -0
- dirigent_examples/shelves/recipes/schema-refuses-then-rule.yaml +127 -0
- dirigent_examples/shelves/recipes/storage-copy-dated-archive.yaml +114 -0
- dirigent_examples/shelves/recipes/storage-exists-gate.yaml +127 -0
- dirigent_examples/shelves/recipes/storage-manifest-of-a-fan-out.yaml +104 -0
- dirigent_examples/shelves/recipes/storage-write-then-read.yaml +119 -0
- dirigent_examples/shelves/recipes/webhook-post-hmac.yaml +132 -0
- dirigent_examples/shelves/recipes/webhook-post-summary.yaml +142 -0
- dirigent_examples/shelves/s3/README.md +34 -0
- dirigent_examples/shelves/s3/report-to-s3.yaml +93 -0
- dirigent_examples/shelves/s3/s3-copy-and-verify.yaml +105 -0
- dirigent_examples/shelves/s3/s3-csv-report.yaml +87 -0
- dirigent_examples/shelves/s3/s3-parquet-report.yaml +77 -0
- dirigent_examples/shelves/s3/s3-round-trip.yaml +125 -0
- dirigent_examples/shelves/schemas/README.md +36 -0
- dirigent_examples/shelves/schemas/echo-reading.json +18 -0
- dirigent_examples/shelves/schemas/ou-record.json +13 -0
- dirigent_examples/shelves/schemas/station-reading.json +13 -0
- dirigent_examples/shelves/sensors/README.md +16 -0
- dirigent_examples/shelves/sensors/sensor-gate.yaml +65 -0
- dirigent_examples/shelves/sensors/time-window.yaml +61 -0
- dirigent_examples/shelves/sql/README.md +52 -0
- dirigent_examples/shelves/sql/duckdb-parquet-to-report.yaml +146 -0
- dirigent_examples/shelves/sql/sql-postgres-readonly.yaml +111 -0
- dirigent_examples/shelves/sql/sql-query-to-storage.yaml +85 -0
- dirigent_examples/shelves/sql/sql-sqlite-roundtrip.yaml +114 -0
- dirigent_examples/shelves/sql/warehouse.sql +42 -0
- dirigent_examples/shelves/transform/README.md +36 -0
- dirigent_examples/shelves/transform/csv-report.yaml +55 -0
- dirigent_examples/shelves/transform/jq-filter-and-map.yaml +70 -0
- dirigent_examples/shelves/transform/jq-group-and-aggregate.yaml +70 -0
- dirigent_examples/shelves/transform/jq-join-two-sources.yaml +98 -0
- dirigent_examples/shelves/transform/jq-reshape.yaml +91 -0
- dirigent_examples/shelves/transform/jq-stream-through-storage.yaml +112 -0
- dirigent_examples/shelves/transform/ndjson-round-trip.yaml +56 -0
- dirigent_examples/shelves/transform/parquet-round-trip.yaml +68 -0
- dirigent_examples/shelves/transform/std-convert-fan-out.yaml +142 -0
- dirigent_examples/shelves/transform/xml-feed-to-ndjson.yaml +116 -0
- dirigent_examples/shelves/transform/yaml-config-to-json.yaml +104 -0
- dirigent_examples/shelves/triggers/README.md +45 -0
- dirigent_examples/shelves/triggers/at-one-time.yaml +78 -0
- dirigent_examples/shelves/triggers/cron-nightly.yaml +79 -0
- dirigent_examples/shelves/triggers/cron-windowed.yaml +86 -0
- dirigent_examples/shelves/triggers/document-nightly.yaml +80 -0
- dirigent_examples/shelves/triggers/interval-rolling.yaml +88 -0
- dirigent_examples/shelves/triggers/managed-and-manual.yaml +109 -0
- dirigent_examples/shelves/triggers/webhook-trigger.yaml +75 -0
- dirigent_examples/shelves/validate/README.md +31 -0
- dirigent_examples/shelves/validate/expects-a-shape.yaml +56 -0
- dirigent_examples/shelves/validate/the-shape-is-wrong.yaml +46 -0
- dirigent_examples-0.15.0.dist-info/METADATA +21 -0
- dirigent_examples-0.15.0.dist-info/RECORD +216 -0
- dirigent_examples-0.15.0.dist-info/WHEEL +4 -0
- dirigent_examples-0.15.0.dist-info/entry_points.txt +3 -0
- dirigent_examples-0.15.0.dist-info/licenses/LICENSE +18 -0
|
@@ -0,0 +1,105 @@
|
|
|
1
|
+
# Copy an object between prefixes, then prove it landed before anything reads it.
|
|
2
|
+
#
|
|
3
|
+
# NEEDS AN S3-COMPATIBLE ENDPOINT, set up exactly as s3-round-trip.yaml documents: the same
|
|
4
|
+
# rustfs container, the same artifacts connection, the same DIRIGENT_STORAGE_CONNECTIONS.
|
|
5
|
+
#
|
|
6
|
+
# The shape is promote-then-consume: a copy from an incoming prefix to a published one,
|
|
7
|
+
# storage.exists standing between the copy and the reader. The sensor is not decoration --
|
|
8
|
+
# on eventually-consistent stores the object's arrival is a fact worth waiting on, and a
|
|
9
|
+
# reader gated on it never sees the half-published state.
|
|
10
|
+
#
|
|
11
|
+
# dg apply examples/s3/s3-copy-and-verify.yaml
|
|
12
|
+
# dg run s3-copy-and-verify -p day=2026-01-01 --watch
|
|
13
|
+
#
|
|
14
|
+
# The first two steps exist so the example carries its own input: a real pipeline starts at
|
|
15
|
+
# seed, with the incoming object written by whatever produced it.
|
|
16
|
+
#
|
|
17
|
+
# Hop by hop:
|
|
18
|
+
#
|
|
19
|
+
# payload the readings, as a value in the run.
|
|
20
|
+
# staged storage.write puts that value in the incoming prefix as json. A value reaches
|
|
21
|
+
# storage here and nowhere else.
|
|
22
|
+
# seed convert.std re-encodes that object into the ndjson a producer would have
|
|
23
|
+
# dropped. A converter is a storage-object operation: it reads one URI and
|
|
24
|
+
# writes another, and the bytes never enter the run.
|
|
25
|
+
# promote the copy from incoming to published.
|
|
26
|
+
# landed the sensor, standing between the copy and the reader.
|
|
27
|
+
# read_back convert.std again, published ndjson back to json, so the result is readable.
|
|
28
|
+
|
|
29
|
+
format: dirigent/v1
|
|
30
|
+
kind: pipeline
|
|
31
|
+
code: s3-copy-and-verify
|
|
32
|
+
name: Copy and verify
|
|
33
|
+
description: Promote an object to a published prefix, and prove it landed before reading it.
|
|
34
|
+
|
|
35
|
+
tags: [s3, sensor, storage, transform]
|
|
36
|
+
|
|
37
|
+
requires:
|
|
38
|
+
blocks:
|
|
39
|
+
- storage.write
|
|
40
|
+
- storage.copy
|
|
41
|
+
- storage.exists
|
|
42
|
+
- value.const
|
|
43
|
+
- convert.std
|
|
44
|
+
connections:
|
|
45
|
+
- artifacts
|
|
46
|
+
storage:
|
|
47
|
+
- s3
|
|
48
|
+
|
|
49
|
+
params:
|
|
50
|
+
type: object
|
|
51
|
+
required: [day]
|
|
52
|
+
additionalProperties: false
|
|
53
|
+
properties:
|
|
54
|
+
day:
|
|
55
|
+
type: string
|
|
56
|
+
format: date
|
|
57
|
+
description: The day whose object is promoted.
|
|
58
|
+
bucket:
|
|
59
|
+
type: string
|
|
60
|
+
default: dirigent-archive
|
|
61
|
+
description: The bucket both prefixes live in.
|
|
62
|
+
|
|
63
|
+
steps:
|
|
64
|
+
payload:
|
|
65
|
+
block: value.const
|
|
66
|
+
config:
|
|
67
|
+
value:
|
|
68
|
+
- { station: st-1, day: "${params.day}", celsius: 4.5 }
|
|
69
|
+
- { station: st-3, day: "${params.day}", celsius: 1.2 }
|
|
70
|
+
staged:
|
|
71
|
+
block: storage.write
|
|
72
|
+
depends_on: [payload]
|
|
73
|
+
config:
|
|
74
|
+
target: s3://${params.bucket}/incoming/${params.day}.json
|
|
75
|
+
value: "${steps.payload.output.value}"
|
|
76
|
+
seed:
|
|
77
|
+
# The example's own input: the incoming object a real producer would have written.
|
|
78
|
+
block: convert.std
|
|
79
|
+
depends_on: [staged]
|
|
80
|
+
config:
|
|
81
|
+
source: "${steps.staged.output.uri}"
|
|
82
|
+
target: s3://${params.bucket}/incoming/${params.day}.ndjson
|
|
83
|
+
from: json
|
|
84
|
+
to: ndjson
|
|
85
|
+
promote:
|
|
86
|
+
block: storage.copy
|
|
87
|
+
depends_on: [seed]
|
|
88
|
+
config:
|
|
89
|
+
source: s3://${params.bucket}/incoming/${params.day}.ndjson
|
|
90
|
+
target: s3://${params.bucket}/published/${params.day}.ndjson
|
|
91
|
+
landed:
|
|
92
|
+
block: storage.exists
|
|
93
|
+
depends_on: [promote]
|
|
94
|
+
poll: 5s
|
|
95
|
+
deadline: 2m
|
|
96
|
+
config:
|
|
97
|
+
uri: s3://${params.bucket}/published/${params.day}.ndjson
|
|
98
|
+
read_back:
|
|
99
|
+
block: convert.std
|
|
100
|
+
depends_on: [landed]
|
|
101
|
+
config:
|
|
102
|
+
source: "${steps.landed.output.uri}"
|
|
103
|
+
target: s3://${params.bucket}/published/${params.day}.json
|
|
104
|
+
from: ndjson
|
|
105
|
+
to: json
|
|
@@ -0,0 +1,87 @@
|
|
|
1
|
+
# A csv report delivered to a bucket, where the people who asked for it can fetch it.
|
|
2
|
+
#
|
|
3
|
+
# NEEDS AN S3-COMPATIBLE ENDPOINT, set up exactly as s3-round-trip.yaml documents: the same
|
|
4
|
+
# rustfs container, the same artifacts connection, the same DIRIGENT_STORAGE_CONNECTIONS.
|
|
5
|
+
#
|
|
6
|
+
# This is examples/transform/csv-report.yaml with a real destination: the same records, the
|
|
7
|
+
# same flattening, the same codec -- and s3:// instead of the run's scratch, so the file
|
|
8
|
+
# outlives the run and lands where a person or another system reads it. Changing the
|
|
9
|
+
# destination was one word, which is the point of schemes.
|
|
10
|
+
#
|
|
11
|
+
# dg apply examples/s3/s3-csv-report.yaml
|
|
12
|
+
# dg run s3-csv-report -p day=2026-01-01 --watch
|
|
13
|
+
#
|
|
14
|
+
# The report is at s3://<bucket>/reports/<day>.csv when the run settles.
|
|
15
|
+
#
|
|
16
|
+
# Hop by hop:
|
|
17
|
+
#
|
|
18
|
+
# readings the records, as a value in the run.
|
|
19
|
+
# rows jq flattens them into the shape a csv can hold. Still a value.
|
|
20
|
+
# staged storage.write puts the rows in the bucket as json, because a converter works
|
|
21
|
+
# on objects rather than on values.
|
|
22
|
+
# report convert.std re-encodes that object as the csv people fetch.
|
|
23
|
+
|
|
24
|
+
format: dirigent/v1
|
|
25
|
+
kind: pipeline
|
|
26
|
+
code: s3-csv-report
|
|
27
|
+
name: A report delivered to a bucket
|
|
28
|
+
description: Shape records into flat rows and write them as a csv object in S3.
|
|
29
|
+
|
|
30
|
+
tags: [s3, storage, transform]
|
|
31
|
+
|
|
32
|
+
requires:
|
|
33
|
+
blocks:
|
|
34
|
+
- value.const
|
|
35
|
+
- transform.jq
|
|
36
|
+
- storage.write
|
|
37
|
+
- convert.std
|
|
38
|
+
connections:
|
|
39
|
+
- artifacts
|
|
40
|
+
storage:
|
|
41
|
+
- s3
|
|
42
|
+
|
|
43
|
+
params:
|
|
44
|
+
type: object
|
|
45
|
+
required: [day]
|
|
46
|
+
additionalProperties: false
|
|
47
|
+
properties:
|
|
48
|
+
day:
|
|
49
|
+
type: string
|
|
50
|
+
format: date
|
|
51
|
+
description: The day the report is about, and the object it is named as.
|
|
52
|
+
bucket:
|
|
53
|
+
type: string
|
|
54
|
+
default: dirigent-archive
|
|
55
|
+
description: The bucket the report is delivered into.
|
|
56
|
+
|
|
57
|
+
steps:
|
|
58
|
+
readings:
|
|
59
|
+
block: value.const
|
|
60
|
+
config:
|
|
61
|
+
value:
|
|
62
|
+
- { station: st-1, region: east, celsius: 4.5, tags: [ok, new] }
|
|
63
|
+
- { station: st-3, region: west, celsius: 1.2, tags: [ok] }
|
|
64
|
+
- { station: st-4, region: north, celsius: -3.4, tags: [] }
|
|
65
|
+
rows:
|
|
66
|
+
block: transform.jq
|
|
67
|
+
depends_on: [readings]
|
|
68
|
+
config:
|
|
69
|
+
input: ${steps.readings.output.value}
|
|
70
|
+
# Flatten for the codec: the tag list becomes one joined cell, because only this
|
|
71
|
+
# program knows what the separator should mean.
|
|
72
|
+
program: |
|
|
73
|
+
[.[] | {station, region, celsius, tags: (.tags | join(" "))}]
|
|
74
|
+
staged:
|
|
75
|
+
block: storage.write
|
|
76
|
+
depends_on: [rows]
|
|
77
|
+
config:
|
|
78
|
+
target: s3://${params.bucket}/reports/${params.day}.json
|
|
79
|
+
value: "${steps.rows.output.value}"
|
|
80
|
+
report:
|
|
81
|
+
block: convert.std
|
|
82
|
+
depends_on: [staged]
|
|
83
|
+
config:
|
|
84
|
+
source: "${steps.staged.output.uri}"
|
|
85
|
+
target: s3://${params.bucket}/reports/${params.day}.csv
|
|
86
|
+
from: json
|
|
87
|
+
to: csv
|
|
@@ -0,0 +1,77 @@
|
|
|
1
|
+
# A parquet dataset delivered to a bucket, with its types along for the ride.
|
|
2
|
+
#
|
|
3
|
+
# NEEDS AN S3-COMPATIBLE ENDPOINT, set up exactly as s3-round-trip.yaml documents: the same
|
|
4
|
+
# rustfs container, the same artifacts connection, the same DIRIGENT_STORAGE_CONNECTIONS.
|
|
5
|
+
#
|
|
6
|
+
# This is examples/transform/parquet-round-trip.yaml with a real destination. Parquet is
|
|
7
|
+
# the format a downstream analysis actually wants -- columnar, typed, compact -- and what
|
|
8
|
+
# lands in the bucket is a dataset a reader opens with pandas or duckdb, not a csv that
|
|
9
|
+
# lost its numbers on the way out. convert.arrow ships in the dirigent-parquet pack;
|
|
10
|
+
# installing it puts the block in the catalog.
|
|
11
|
+
#
|
|
12
|
+
# dg apply examples/s3/s3-parquet-report.yaml
|
|
13
|
+
# dg run s3-parquet-report -p day=2026-01-01 --watch
|
|
14
|
+
#
|
|
15
|
+
# The dataset is at s3://<bucket>/datasets/<day>.parquet when the run settles.
|
|
16
|
+
#
|
|
17
|
+
# Hop by hop:
|
|
18
|
+
#
|
|
19
|
+
# readings the records, as a value in the run, flat as parquet columns need.
|
|
20
|
+
# staged storage.write puts them in the bucket as json, which is the form a converter
|
|
21
|
+
# can read: convert.arrow works on objects, never on a value in the run.
|
|
22
|
+
# dataset convert.arrow re-encodes that object as parquet, types and all.
|
|
23
|
+
|
|
24
|
+
format: dirigent/v1
|
|
25
|
+
kind: pipeline
|
|
26
|
+
code: s3-parquet-report
|
|
27
|
+
name: A parquet dataset in a bucket
|
|
28
|
+
description: Write typed records as a parquet object in S3, where an analysis can read them.
|
|
29
|
+
|
|
30
|
+
tags: [s3, storage, transform]
|
|
31
|
+
|
|
32
|
+
requires:
|
|
33
|
+
blocks:
|
|
34
|
+
- value.const
|
|
35
|
+
- storage.write
|
|
36
|
+
- convert.arrow
|
|
37
|
+
connections:
|
|
38
|
+
- artifacts
|
|
39
|
+
storage:
|
|
40
|
+
- s3
|
|
41
|
+
|
|
42
|
+
params:
|
|
43
|
+
type: object
|
|
44
|
+
required: [day]
|
|
45
|
+
additionalProperties: false
|
|
46
|
+
properties:
|
|
47
|
+
day:
|
|
48
|
+
type: string
|
|
49
|
+
format: date
|
|
50
|
+
description: The day the dataset covers, and the object it is named as.
|
|
51
|
+
bucket:
|
|
52
|
+
type: string
|
|
53
|
+
default: dirigent-archive
|
|
54
|
+
description: The bucket the dataset is delivered into.
|
|
55
|
+
|
|
56
|
+
steps:
|
|
57
|
+
readings:
|
|
58
|
+
block: value.const
|
|
59
|
+
config:
|
|
60
|
+
value:
|
|
61
|
+
- { station: st-1, region: east, celsius: 4.5, active: true }
|
|
62
|
+
- { station: st-3, region: west, celsius: 1.2, active: true }
|
|
63
|
+
- { station: st-4, region: north, celsius: -3.4, active: false }
|
|
64
|
+
staged:
|
|
65
|
+
block: storage.write
|
|
66
|
+
depends_on: [readings]
|
|
67
|
+
config:
|
|
68
|
+
target: s3://${params.bucket}/datasets/${params.day}.json
|
|
69
|
+
value: "${steps.readings.output.value}"
|
|
70
|
+
dataset:
|
|
71
|
+
block: convert.arrow
|
|
72
|
+
depends_on: [staged]
|
|
73
|
+
config:
|
|
74
|
+
source: "${steps.staged.output.uri}"
|
|
75
|
+
target: s3://${params.bucket}/datasets/${params.day}.parquet
|
|
76
|
+
from: json
|
|
77
|
+
to: parquet
|
|
@@ -0,0 +1,125 @@
|
|
|
1
|
+
# Moving an artifact out to object storage and back, with no S3-specific block anywhere.
|
|
2
|
+
#
|
|
3
|
+
# NEEDS AN S3-COMPATIBLE ENDPOINT. dirigent-storage-s3 is installed and every block below
|
|
4
|
+
# exists; what is missing is a credential and a server to point it at.
|
|
5
|
+
#
|
|
6
|
+
# ON THE COMPOSE STACK there is nothing to do: the stack runs an S3 server, keeps artifacts
|
|
7
|
+
# in its bucket, and bootstraps the `artifacts` connection this document names. Apply and
|
|
8
|
+
# run, passing the stack's own bucket:
|
|
9
|
+
#
|
|
10
|
+
# dg run s3-round-trip -p day=2026-01-01 -p bucket=dirigent --watch
|
|
11
|
+
#
|
|
12
|
+
# The steps below are for `dg dev`, which has neither.
|
|
13
|
+
#
|
|
14
|
+
# A storage backend is not a step. There is no "upload to S3" block and no "download from
|
|
15
|
+
# S3" block: s3:// is a registered URI scheme, so the ordinary storage.copy moves bytes
|
|
16
|
+
# between file:// and s3:// in either direction, and the ordinary storage.exists waits on
|
|
17
|
+
# an s3:// object. Changing where the archive lives is one word.
|
|
18
|
+
#
|
|
19
|
+
# 1. Start an S3-compatible server. Any implementation will do; this is the one dirigent's
|
|
20
|
+
# own storage lane runs against:
|
|
21
|
+
#
|
|
22
|
+
# docker run -d --name dirigent-s3 -p 9000:9000 \
|
|
23
|
+
# -e RUSTFS_ACCESS_KEY=dirigent-test-key \
|
|
24
|
+
# -e RUSTFS_SECRET_KEY=dirigent-test-secret \
|
|
25
|
+
# rustfs/rustfs:1.0.0-rc.4
|
|
26
|
+
#
|
|
27
|
+
# 2. Create the connection, and tell the instance that it is what serves s3://. The first
|
|
28
|
+
# is the credential; the second is which of possibly several s3 connections the scheme
|
|
29
|
+
# is configured from:
|
|
30
|
+
#
|
|
31
|
+
# dg connection create s3 artifacts \
|
|
32
|
+
# --set endpoint_url=http://127.0.0.1:9000 \
|
|
33
|
+
# --set access_key_id=dirigent-test-key \
|
|
34
|
+
# --set secret_access_key=dirigent-test-secret \
|
|
35
|
+
# --set path_style=true \
|
|
36
|
+
# --set bucket=dirigent-archive
|
|
37
|
+
#
|
|
38
|
+
# export DIRIGENT_STORAGE_CONNECTIONS='{"s3": "artifacts"}'
|
|
39
|
+
#
|
|
40
|
+
# 3. Create the bucket the run addresses, then apply and run:
|
|
41
|
+
#
|
|
42
|
+
# dg apply examples/s3/s3-round-trip.yaml
|
|
43
|
+
# dg run s3-round-trip -p day=2026-01-01 --watch
|
|
44
|
+
#
|
|
45
|
+
# The bucket always comes from the URI, never from the connection: the connection's own
|
|
46
|
+
# bucket field is only what `dg connection check artifacts` probes.
|
|
47
|
+
#
|
|
48
|
+
# A real pipeline stops after the first copy. The download here makes the result checkable
|
|
49
|
+
# without a second system to look at.
|
|
50
|
+
|
|
51
|
+
format: dirigent/v1
|
|
52
|
+
kind: pipeline
|
|
53
|
+
code: s3-round-trip
|
|
54
|
+
name: S3 round trip
|
|
55
|
+
description: Move an artifact from local storage out to S3 and back, then confirm it landed.
|
|
56
|
+
|
|
57
|
+
tags: [s3, http, sensor, storage]
|
|
58
|
+
|
|
59
|
+
requires:
|
|
60
|
+
blocks:
|
|
61
|
+
- http.request
|
|
62
|
+
- storage.write
|
|
63
|
+
- storage.copy
|
|
64
|
+
- storage.exists
|
|
65
|
+
connections:
|
|
66
|
+
- artifacts
|
|
67
|
+
storage:
|
|
68
|
+
- s3
|
|
69
|
+
|
|
70
|
+
params:
|
|
71
|
+
type: object
|
|
72
|
+
required: [day]
|
|
73
|
+
additionalProperties: false
|
|
74
|
+
properties:
|
|
75
|
+
day:
|
|
76
|
+
type: string
|
|
77
|
+
format: date
|
|
78
|
+
description: The day whose artifact is round-tripped.
|
|
79
|
+
bucket:
|
|
80
|
+
type: string
|
|
81
|
+
default: dirigent-archive
|
|
82
|
+
description: The bucket the artifact is archived into.
|
|
83
|
+
|
|
84
|
+
steps:
|
|
85
|
+
produce:
|
|
86
|
+
block: http.request
|
|
87
|
+
config:
|
|
88
|
+
url: https://postman-echo.com/post
|
|
89
|
+
method: POST
|
|
90
|
+
body:
|
|
91
|
+
day: "${params.day}"
|
|
92
|
+
|
|
93
|
+
stage:
|
|
94
|
+
block: storage.write
|
|
95
|
+
depends_on: [produce]
|
|
96
|
+
# The local artifact the round trip is about, written on file:// because that is where
|
|
97
|
+
# ${run.scratch} lives. The copy below is the only line that mentions s3.
|
|
98
|
+
config:
|
|
99
|
+
target: "${run.scratch}/local/${params.day}.json"
|
|
100
|
+
value: "${steps.produce.output.body}"
|
|
101
|
+
|
|
102
|
+
upload:
|
|
103
|
+
block: storage.copy
|
|
104
|
+
depends_on: [stage]
|
|
105
|
+
config:
|
|
106
|
+
source: "${steps.stage.output.uri}"
|
|
107
|
+
target: "s3://${params.bucket}/daily/${params.day}.json"
|
|
108
|
+
|
|
109
|
+
wait_for_object:
|
|
110
|
+
block: storage.exists
|
|
111
|
+
depends_on: [upload]
|
|
112
|
+
# S3 acknowledges a write before every reader can see it, so the upload settling is not
|
|
113
|
+
# the same as the object being there. This waits for the second thing.
|
|
114
|
+
poll: 2s
|
|
115
|
+
deadline: 1m
|
|
116
|
+
config:
|
|
117
|
+
uri: "s3://${params.bucket}/daily/${params.day}.json"
|
|
118
|
+
min_size: 1
|
|
119
|
+
|
|
120
|
+
download:
|
|
121
|
+
block: storage.copy
|
|
122
|
+
depends_on: [wait_for_object]
|
|
123
|
+
config:
|
|
124
|
+
source: "${steps.wait_for_object.output.uri}"
|
|
125
|
+
target: "${run.scratch}/roundtrip/${params.day}.json"
|
|
@@ -0,0 +1,36 @@
|
|
|
1
|
+
# Example schemas
|
|
2
|
+
|
|
3
|
+
Each file here is a plain JSON Schema (Draft 2020-12): the shape a read is expected to return,
|
|
4
|
+
written down once so a pipeline can be refused the moment a payload moves out from under it. A
|
|
5
|
+
schema is **locally authored** -- a picture you hold of the payload, never something fetched or
|
|
6
|
+
introspected from the source. Each is applied on its own with `dg schema create`, or with the
|
|
7
|
+
directory it sits in, and a document that references one names it in `requires.schemas`. A document may instead **carry** a
|
|
8
|
+
schema in its own top-level `schemas:` section for a standalone or `--local` run; a server
|
|
9
|
+
refuses a document that carries one, so a shared instance holds its schemas here.
|
|
10
|
+
|
|
11
|
+
Apply one the way a person would:
|
|
12
|
+
|
|
13
|
+
```bash
|
|
14
|
+
dg schema create examples/schemas/ou-record.json
|
|
15
|
+
```
|
|
16
|
+
|
|
17
|
+
A directory apply lands them the same way: a server pointed at a directory stores every schema
|
|
18
|
+
it finds there before the pipelines, so a mounted corpus needs no separate step.
|
|
19
|
+
|
|
20
|
+
A schema carries its own identity in its keywords, so there is nothing else to pass:
|
|
21
|
+
|
|
22
|
+
- `$id` becomes the `code` the schema is addressed by (falling back to the file's stem).
|
|
23
|
+
- `title` becomes its name.
|
|
24
|
+
- `description` becomes its body.
|
|
25
|
+
|
|
26
|
+
| File | The shape it pins |
|
|
27
|
+
| --- | --- |
|
|
28
|
+
| [ou-record.json](ou-record.json) | A single organisation-unit record: a string `id`, a `name`, and an integer `level` |
|
|
29
|
+
| [echo-reading.json](echo-reading.json) | What a fetch of the echo service answers with: an `args` map carrying `station` and `day`, and the `url` |
|
|
30
|
+
| [station-reading.json](station-reading.json) | The row a transform is held to before it is posted on: a `station`, a `day`, and a numeric `celsius` |
|
|
31
|
+
|
|
32
|
+
`ou-record` is the shape the two [validation documents](../validate) check a payload against, and
|
|
33
|
+
the two reading shapes are the gates either side of the transform in
|
|
34
|
+
[http-fetch-validate-post.yaml](../recipes/http-fetch-validate-post.yaml).
|
|
35
|
+
The [JSON Schema guide](../../docs/json-schema.md) walks through how a schema like it is
|
|
36
|
+
built, keyword by keyword.
|
|
@@ -0,0 +1,18 @@
|
|
|
1
|
+
{
|
|
2
|
+
"$id": "echo-reading",
|
|
3
|
+
"title": "Echoed reading batch",
|
|
4
|
+
"description": "What a fetch of the echo service is held to before anything reads into it: an `args` map carrying the query the call sent, and the `url` it was answered from. A source that changes shape under a pipeline is refused here rather than producing a wrong number three steps later.",
|
|
5
|
+
"type": "object",
|
|
6
|
+
"required": ["args", "url"],
|
|
7
|
+
"properties": {
|
|
8
|
+
"args": {
|
|
9
|
+
"type": "object",
|
|
10
|
+
"required": ["station", "day"],
|
|
11
|
+
"properties": {
|
|
12
|
+
"station": { "type": "string", "minLength": 1 },
|
|
13
|
+
"day": { "type": "string", "format": "date" }
|
|
14
|
+
}
|
|
15
|
+
},
|
|
16
|
+
"url": { "type": "string" }
|
|
17
|
+
}
|
|
18
|
+
}
|
|
@@ -0,0 +1,13 @@
|
|
|
1
|
+
{
|
|
2
|
+
"$id": "ou-record",
|
|
3
|
+
"title": "Organisation unit record",
|
|
4
|
+
"description": "A single organisation-unit record: a string id, a name, and an integer level. The shape a payload is held to before a downstream step reads any of the three.",
|
|
5
|
+
"type": "object",
|
|
6
|
+
"required": ["id", "name", "level"],
|
|
7
|
+
"additionalProperties": false,
|
|
8
|
+
"properties": {
|
|
9
|
+
"id": { "type": "string" },
|
|
10
|
+
"name": { "type": "string" },
|
|
11
|
+
"level": { "type": "integer" }
|
|
12
|
+
}
|
|
13
|
+
}
|
|
@@ -0,0 +1,13 @@
|
|
|
1
|
+
{
|
|
2
|
+
"$id": "station-reading",
|
|
3
|
+
"title": "Station reading",
|
|
4
|
+
"description": "The row a transform is held to before it is posted onward: a station, the day it reports for, and a temperature that is a number rather than the string a query carried. The gate a document puts after its own jq program, so a typo in the program is refused instead of delivered.",
|
|
5
|
+
"type": "object",
|
|
6
|
+
"required": ["station", "day", "celsius"],
|
|
7
|
+
"additionalProperties": false,
|
|
8
|
+
"properties": {
|
|
9
|
+
"station": { "type": "string", "minLength": 1 },
|
|
10
|
+
"day": { "type": "string", "format": "date" },
|
|
11
|
+
"celsius": { "type": "number" }
|
|
12
|
+
}
|
|
13
|
+
}
|
|
@@ -0,0 +1,16 @@
|
|
|
1
|
+
# Sensor examples
|
|
2
|
+
|
|
3
|
+
Steps that wait for the world instead of doing something to it: each poke is one cheap,
|
|
4
|
+
read-only question, the run holds no worker slot between pokes, and `deadline` with
|
|
5
|
+
`on_timeout` says what a day without an answer means.
|
|
6
|
+
|
|
7
|
+
```bash
|
|
8
|
+
dg run --local examples/sensors/time-window.yaml --enable-unsafe shell.run
|
|
9
|
+
```
|
|
10
|
+
|
|
11
|
+
## Pipelines
|
|
12
|
+
|
|
13
|
+
| File | What it teaches |
|
|
14
|
+
| --- | --- |
|
|
15
|
+
| [sensor-gate.yaml](sensor-gate.yaml) | `storage.exists` gating a load: wait for today's drop, and skip the day if it never lands. |
|
|
16
|
+
| [time-window.yaml](time-window.yaml) | The clock as a condition: hold until the local time is inside a window, in the window's own zone. |
|
|
@@ -0,0 +1,65 @@
|
|
|
1
|
+
# A sensor gating the rest of the pipeline, and what happens when it never fires.
|
|
2
|
+
#
|
|
3
|
+
# storage.exists is a sensor: it observes the world and changes nothing. Each poke is a
|
|
4
|
+
# durable, scheduled poll, so waiting hours costs hours of rows, not hours of a held worker.
|
|
5
|
+
#
|
|
6
|
+
# poll, deadline, and on_timeout are step-level engine semantics, uniform across every
|
|
7
|
+
# block, never buried in the block's own config. on_timeout: skip is the load-bearing
|
|
8
|
+
# choice here: no drop today is not a failure, it is a day with nothing to load, so the
|
|
9
|
+
# sensor skips and the default all_success edges skip the rest of the branch with it.
|
|
10
|
+
#
|
|
11
|
+
# The deadline is short so the skip is watchable. Nothing ever writes the drop, so after ten
|
|
12
|
+
# seconds the sensor skips, the branch behind it skips, and the run succeeds with nothing
|
|
13
|
+
# done:
|
|
14
|
+
# dg run --local examples/sensors/sensor-gate.yaml -p day=2026-01-01 --enable-unsafe shell.run
|
|
15
|
+
#
|
|
16
|
+
# Note where the drop is looked for. file:// is rooted at the instance's artifact root and
|
|
17
|
+
# refuses to address anything outside it, so a stored pipeline cannot turn a copy step into
|
|
18
|
+
# an arbitrary-file read. Pointing this at a real drop means an s3:// URI, or a path under
|
|
19
|
+
# that root -- never an absolute path from the document.
|
|
20
|
+
#
|
|
21
|
+
# confirm uses shell.run, which executes code on the worker and needs the allowlist:
|
|
22
|
+
# export DIRIGENT_ENABLED_UNSAFE_BLOCKS='["shell.run"]'
|
|
23
|
+
|
|
24
|
+
format: dirigent/v1
|
|
25
|
+
kind: pipeline
|
|
26
|
+
code: sensor-gate
|
|
27
|
+
name: A sensor gates the run
|
|
28
|
+
description: Wait for today's drop to land, then load it; skip the day if it never does.
|
|
29
|
+
|
|
30
|
+
tags: [sensor, execute, storage]
|
|
31
|
+
|
|
32
|
+
params:
|
|
33
|
+
type: object
|
|
34
|
+
required: [day]
|
|
35
|
+
properties:
|
|
36
|
+
day:
|
|
37
|
+
type: string
|
|
38
|
+
format: date
|
|
39
|
+
|
|
40
|
+
steps:
|
|
41
|
+
wait_for_drop:
|
|
42
|
+
block: storage.exists
|
|
43
|
+
# Each poke is a scheduled row, not a held worker, so the poll interval is chosen for
|
|
44
|
+
# how soon the drop matters rather than for what it costs to wait.
|
|
45
|
+
poll: 2s
|
|
46
|
+
deadline: 10s
|
|
47
|
+
# No drop today is a day with nothing to load, not a failure. The skip carries down the
|
|
48
|
+
# branch through the default all_success edges, and the run still succeeds.
|
|
49
|
+
on_timeout: skip
|
|
50
|
+
config:
|
|
51
|
+
uri: "${run.scratch}/drops/${params.day}.parquet"
|
|
52
|
+
min_size: 1
|
|
53
|
+
|
|
54
|
+
load:
|
|
55
|
+
block: storage.copy
|
|
56
|
+
depends_on: [wait_for_drop]
|
|
57
|
+
config:
|
|
58
|
+
source: "${steps.wait_for_drop.output.uri}"
|
|
59
|
+
target: "${run.scratch}/loaded/${params.day}.parquet"
|
|
60
|
+
|
|
61
|
+
confirm:
|
|
62
|
+
block: shell.run
|
|
63
|
+
depends_on: [load]
|
|
64
|
+
config:
|
|
65
|
+
argv: [echo, "loaded ${steps.wait_for_drop.output.size} bytes for ${params.day}"]
|
|
@@ -0,0 +1,61 @@
|
|
|
1
|
+
# Gating a step on the wall clock, in a timezone the document names out loud.
|
|
2
|
+
#
|
|
3
|
+
# A schedule says when a run starts. time.window says when a step may proceed, which is a
|
|
4
|
+
# different question: "load whenever the export lands, but never touch the warehouse outside
|
|
5
|
+
# the maintenance window" is a property of the step, not of the trigger.
|
|
6
|
+
#
|
|
7
|
+
# It is a sensor, so it inherits the whole waiting apparatus: poll, deadline, and on_timeout
|
|
8
|
+
# are step-level engine semantics here exactly as they are for storage.exists. A closed
|
|
9
|
+
# window parks until it opens rather than polling on the step's cadence, so waiting out a
|
|
10
|
+
# Sunday costs a handful of pokes.
|
|
11
|
+
#
|
|
12
|
+
# The timezone defaults to UTC and should always be stated. Europe/Oslo below is read as
|
|
13
|
+
# Oslo's clock, daylight saving included, no matter where the worker runs.
|
|
14
|
+
#
|
|
15
|
+
# The window here is wide open -- the whole day, every day -- so the example runs without
|
|
16
|
+
# waiting. Narrow it to see the wait:
|
|
17
|
+
# dg run --local examples/sensors/time-window.yaml --enable-unsafe shell.run
|
|
18
|
+
#
|
|
19
|
+
# A real maintenance window is the commented shape below the open one:
|
|
20
|
+
# after: "01:00"
|
|
21
|
+
# before: "04:00"
|
|
22
|
+
# days: [sat, sun]
|
|
23
|
+
#
|
|
24
|
+
# Windows may cross midnight: after 22:00 before 04:00 is one night window, and when days is
|
|
25
|
+
# set it names the day the window opens on, so [fri] runs Friday evening into Saturday.
|
|
26
|
+
|
|
27
|
+
format: dirigent/v1
|
|
28
|
+
kind: pipeline
|
|
29
|
+
code: time-window
|
|
30
|
+
name: Inside a time window
|
|
31
|
+
description: Wait until the local clock is inside a window, then do the work it guards.
|
|
32
|
+
|
|
33
|
+
tags: [sensor, execute, window]
|
|
34
|
+
|
|
35
|
+
requires:
|
|
36
|
+
blocks:
|
|
37
|
+
- time.window
|
|
38
|
+
- shell.run
|
|
39
|
+
|
|
40
|
+
steps:
|
|
41
|
+
maintenance_window:
|
|
42
|
+
block: time.window
|
|
43
|
+
poll: 5s
|
|
44
|
+
deadline: 1m
|
|
45
|
+
on_timeout: skip
|
|
46
|
+
config:
|
|
47
|
+
# Open all day so the example runs without waiting; a real window is narrow, and the
|
|
48
|
+
# commented shape in the header is what one looks like.
|
|
49
|
+
after: "00:00"
|
|
50
|
+
before: "23:59"
|
|
51
|
+
# Always state it. The default is UTC, and a window meant as local time silently
|
|
52
|
+
# becomes the wrong hours wherever the worker happens to run.
|
|
53
|
+
timezone: Europe/Oslo
|
|
54
|
+
|
|
55
|
+
reindex:
|
|
56
|
+
block: shell.run
|
|
57
|
+
depends_on: [maintenance_window]
|
|
58
|
+
config:
|
|
59
|
+
argv:
|
|
60
|
+
- echo
|
|
61
|
+
- "the window opened at ${steps.maintenance_window.output.entered_at} (${steps.maintenance_window.output.timezone})"
|
|
@@ -0,0 +1,52 @@
|
|
|
1
|
+
# SQL examples
|
|
2
|
+
|
|
3
|
+
The `sql.*` family reads from and writes to a database. `sql.query` runs one statement and
|
|
4
|
+
hands its rows on; `sql.execute` runs a list of statements as one transaction.
|
|
5
|
+
[docs/sql.md](../../docs/sql.md) is the family's home.
|
|
6
|
+
|
|
7
|
+
Both are **ordinary** blocks: they run no command a document supplies and reach nothing but the
|
|
8
|
+
database their connection names, so no id has to be allowlisted to run them.
|
|
9
|
+
|
|
10
|
+
```bash
|
|
11
|
+
dg run --local examples/sql/sql-sqlite-roundtrip.yaml
|
|
12
|
+
dg run --local examples/sql/sql-query-to-storage.yaml
|
|
13
|
+
```
|
|
14
|
+
|
|
15
|
+
Those two need nothing at all -- no network, no daemon, no server. Each builds a SQLite
|
|
16
|
+
database in the run's own work directory, so the whole example is self-contained. The DuckDB one
|
|
17
|
+
needs the engine (`dirigent-blocks[duckdb]`) and the parquet pack, and nothing else; the last
|
|
18
|
+
names a real PostgreSQL and says so in its own header.
|
|
19
|
+
|
|
20
|
+
```bash
|
|
21
|
+
dg run --local examples/sql/duckdb-parquet-to-report.yaml --keep
|
|
22
|
+
```
|
|
23
|
+
|
|
24
|
+
The rule the whole family turns on: **a value is bound, never interpolated**. The statement is
|
|
25
|
+
a constant in the document, every value is a named `:parameter`, and a `${...}` reference
|
|
26
|
+
resolves into `params` and never into the SQL text. A parameter that reads as SQL is compared
|
|
27
|
+
as a string and matches nothing.
|
|
28
|
+
|
|
29
|
+
The first two documents carry their `sql` connection in a `connections:` section, because a
|
|
30
|
+
`--local` run has no instance to hold one. On a server the connection is created once and the
|
|
31
|
+
document names it, with the password set as its own sealed field:
|
|
32
|
+
|
|
33
|
+
```bash
|
|
34
|
+
dg connection create sql warehouse-read \
|
|
35
|
+
--set url=postgresql+asyncpg://reader@db.example:5432/warehouse \
|
|
36
|
+
--set password=... \
|
|
37
|
+
--set read_only=true
|
|
38
|
+
```
|
|
39
|
+
|
|
40
|
+
On the compose stack that database is the `infra/compose.sql.yaml` overlay
|
|
41
|
+
(`make docker-run-sql`), seeded with the `reading` table the third document queries, and the
|
|
42
|
+
connection names it as `reader@warehouse:5432/warehouse`.
|
|
43
|
+
|
|
44
|
+
## Pipelines
|
|
45
|
+
|
|
46
|
+
| File | What it teaches |
|
|
47
|
+
| --- | --- |
|
|
48
|
+
| [sql-sqlite-roundtrip.yaml](sql-sqlite-roundtrip.yaml) | The family end to end: a table created and filled in one transaction, read back with a bound parameter, and the rows referenced by a downstream step. |
|
|
49
|
+
| [sql-query-to-storage.yaml](sql-query-to-storage.yaml) | Rows into a file: the query hands them on as its output, `storage.write` puts them in an object, and `max_rows` bounds what the step will carry. |
|
|
50
|
+
| [duckdb-parquet-to-report.yaml](duckdb-parquet-to-report.yaml) | The engine that reads files: a parquet artifact queried by DuckDB through a bound `read_parquet(:source)`, and the answer copied out as a csv artifact. Needs the `dirigent-blocks[duckdb]` extra and the parquet pack. |
|
|
51
|
+
| [sql-postgres-readonly.yaml](sql-postgres-readonly.yaml) | A referenced PostgreSQL connection with `read_only: true` and its password sealed separately, queried inside the run's window. Validates offline; it cannot run without an instance. |
|
|
52
|
+
| [warehouse.sql](warehouse.sql) | Not a pipeline: the roles and the one table the compose stack's warehouse is seeded with, mounted by `infra/compose.sql.yaml` |
|