ajdc 0.0.1 → 0.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +4 -0
- data/README.md +380 -13
- data/db/durable_schema.rb +57 -0
- data/lib/active_job/continuable/callbacks.rb +48 -0
- data/lib/active_job/continuable/configurable.rb +75 -0
- data/lib/active_job/continuation/callbacks.rb +18 -0
- data/lib/active_job/continuation/rescue_handlers_first.rb +29 -0
- data/lib/active_job/durable/args_mapper.rb +41 -0
- data/lib/active_job/durable/config.rb +100 -0
- data/lib/active_job/durable/continuation.rb +24 -0
- data/lib/active_job/durable/execution.rb +438 -0
- data/lib/active_job/durable/housekeeping_job.rb +16 -0
- data/lib/active_job/durable/record.rb +17 -0
- data/lib/active_job/durable/run.rb +165 -0
- data/lib/active_job/durable/step.rb +14 -0
- data/lib/active_job/durable/wake_job.rb +16 -0
- data/lib/active_job/durable.rb +251 -0
- data/lib/ajdc/railtie.rb +18 -0
- data/lib/ajdc/version.rb +1 -1
- data/lib/ajdc.rb +1 -0
- data/lib/generators/ajdc/install/USAGE +19 -0
- data/lib/generators/ajdc/install/install_generator.rb +66 -0
- data/lib/generators/ajdc/install/templates/create_active_job_durable_tables.rb.tt +56 -0
- data/lib/generators/ajdc/install/templates/initializer.rb.tt +4 -0
- data/skills/ajdc/SKILL.md +117 -0
- data/skills/ajdc/examples/README.md +11 -0
- data/skills/ajdc/examples/agent-loop.md +57 -0
- data/skills/ajdc/examples/halting.md +84 -0
- data/skills/ajdc/examples/pipeline.md +73 -0
- data/skills/ajdc/examples/signals.md +139 -0
- data/skills/ajdc/examples/timers.md +136 -0
- data/skills/ajdc/examples/uniqueness.md +87 -0
- data/skills/ajdc/references/api.md +157 -0
- data/skills/ajdc/references/decision-guide.md +130 -0
- data/skills/ajdc/references/installation.md +98 -0
- data/skills/ajdc/references/testing.md +152 -0
- data/skills/ajdc/references/transactions.md +68 -0
- metadata +83 -7
|
@@ -0,0 +1,87 @@
|
|
|
1
|
+
# Uniqueness: one live run per identity
|
|
2
|
+
|
|
3
|
+
`unique_by` names the `perform` parameters that identify a run and forbids a second live run
|
|
4
|
+
for the same identity. The check runs at `perform_later`, inside the caller's transaction; the
|
|
5
|
+
unique index on `[job_class, active_key]` settles concurrent enqueues. Terminal runs
|
|
6
|
+
(`completed`, `discarded`, `cancelled`) free the identity.
|
|
7
|
+
|
|
8
|
+
## `:skip` (default): the second start is a no-op
|
|
9
|
+
|
|
10
|
+
For a sweep that starts runs, a button clicked twice, a webhook delivered twice.
|
|
11
|
+
|
|
12
|
+
```ruby
|
|
13
|
+
class License::LifecycleJob < ApplicationJob
|
|
14
|
+
include ActiveJob::Durable
|
|
15
|
+
|
|
16
|
+
unique_by :license
|
|
17
|
+
|
|
18
|
+
def perform(license) = # ...
|
|
19
|
+
end
|
|
20
|
+
|
|
21
|
+
# a daily starter for licenses expiring within 30 days; running it twice starts nothing new
|
|
22
|
+
License.expiring_within(30.days).find_each { License::LifecycleJob.perform_later(it) }
|
|
23
|
+
License::LifecycleJob.perform_later(license) # => false while a run is live
|
|
24
|
+
```
|
|
25
|
+
|
|
26
|
+
## `:replace`: the new start cancels the live run
|
|
27
|
+
|
|
28
|
+
For a reschedule or a renewal, when the date the run waits for has moved.
|
|
29
|
+
|
|
30
|
+
```ruby
|
|
31
|
+
class AppointmentReminderJob < ApplicationJob
|
|
32
|
+
include ActiveJob::Durable
|
|
33
|
+
|
|
34
|
+
unique_by :appointment, on_conflict: :replace
|
|
35
|
+
|
|
36
|
+
def perform(appointment)
|
|
37
|
+
step :remind, wait_until: appointment.start_at - 30.minutes do
|
|
38
|
+
AppointmentNotification.with(appointment:).deliver_later(appointment.patient)
|
|
39
|
+
end
|
|
40
|
+
end
|
|
41
|
+
end
|
|
42
|
+
|
|
43
|
+
# PatientAppointment: after_create_commit, and after_update_commit when start_at changed:
|
|
44
|
+
# AppointmentReminderJob.perform_later(self) # the old run is cancelled, a new one parks
|
|
45
|
+
```
|
|
46
|
+
|
|
47
|
+
## `:reject`: the second start is an error
|
|
48
|
+
|
|
49
|
+
For money and anything where a duplicate must be seen, not swallowed.
|
|
50
|
+
|
|
51
|
+
```ruby
|
|
52
|
+
class PayoutJob < ApplicationJob
|
|
53
|
+
include ActiveJob::Durable
|
|
54
|
+
|
|
55
|
+
unique_by :payout, on_conflict: :reject
|
|
56
|
+
|
|
57
|
+
def perform(payout, wallet_id, amount) = # ...
|
|
58
|
+
end
|
|
59
|
+
|
|
60
|
+
PayoutJob.perform_later(payout, wallet_id, amount)
|
|
61
|
+
PayoutJob.perform_later(payout, wallet_id, amount) # raises ActiveJob::Durable::RunAlreadyExists; error.run is the live one
|
|
62
|
+
```
|
|
63
|
+
|
|
64
|
+
`unique_by :payout` also lets a webhook find the run with the payout alone:
|
|
65
|
+
`payout.job_run.wake_up(:payment, ...)`.
|
|
66
|
+
|
|
67
|
+
## Identity inside an argument
|
|
68
|
+
|
|
69
|
+
When the identity is not a `perform` parameter but a value inside one, name it with a block.
|
|
70
|
+
A Stripe webhook payload:
|
|
71
|
+
|
|
72
|
+
```ruby
|
|
73
|
+
class Payment::ChargeImportJob < ApplicationJob
|
|
74
|
+
include ActiveJob::Durable
|
|
75
|
+
|
|
76
|
+
unique_by { |payload| payload.dig("data", "object", "id") } # the charge id; :skip
|
|
77
|
+
|
|
78
|
+
def perform(payload)
|
|
79
|
+
step :import do
|
|
80
|
+
Payment::StripeChargeImport.call(payload.dig("data", "object"))
|
|
81
|
+
end
|
|
82
|
+
end
|
|
83
|
+
end
|
|
84
|
+
```
|
|
85
|
+
|
|
86
|
+
Keep an idempotency guard inside the step when a duplicate may arrive after the run completed:
|
|
87
|
+
uniqueness covers live and attention runs only.
|
|
@@ -0,0 +1,157 @@
|
|
|
1
|
+
# AJ/DC API reference (differences that matter)
|
|
2
|
+
|
|
3
|
+
## 1. Setup
|
|
4
|
+
|
|
5
|
+
```ruby
|
|
6
|
+
gem "ajdc"
|
|
7
|
+
bin/rails generate ajdc:install # migration for the two tables in the primary database
|
|
8
|
+
bin/rails generate ajdc:install --database=durable # or in a separate database, plus the initializer
|
|
9
|
+
```
|
|
10
|
+
|
|
11
|
+
Which database, and why it matters: `installation.md` §2.
|
|
12
|
+
|
|
13
|
+
```yaml
|
|
14
|
+
# config/recurring.yml (Solid Queue) or the app's scheduler
|
|
15
|
+
durable_wake:
|
|
16
|
+
class: ActiveJob::Durable::WakeJob
|
|
17
|
+
schedule: every minute
|
|
18
|
+
```
|
|
19
|
+
|
|
20
|
+
`include ActiveJob::Durable` in the job. It includes `Continuable`; `step`, `start:`,
|
|
21
|
+
`isolated:`, `step.cursor`, `step.set!`, `step.advance!`, `step.checkpoint!`, `attribute`,
|
|
22
|
+
`retry_on`, `discard_on` are unchanged. The payload carries only the run id; progress and
|
|
23
|
+
attributes live on the run row.
|
|
24
|
+
|
|
25
|
+
## 2. Identity: `identified_by` versus `unique_by` versus `set(workflow_key:)`
|
|
26
|
+
|
|
27
|
+
| Macro | Answers | Column | Effect on enqueue |
|
|
28
|
+
|---|---|---|---|
|
|
29
|
+
| none | "what is this run about?" | `key` = all arguments, positional in order then keywords `name=value` | none |
|
|
30
|
+
| `identified_by :card` | same, narrowed to the named `perform` parameters | `key` from those parameters only | none |
|
|
31
|
+
| `identified_by { \|card, **kw\| [card, kw.fetch(:format, "csv")] }` | same, computed | `key` from the block's return | none |
|
|
32
|
+
| `unique_by :card` | "may two live runs exist for this?" | `active_key` while live or attention; `NULL` when terminal | `on_conflict` applies |
|
|
33
|
+
| `unique_by { \|payload\| payload.dig("data", "object", "id") }` | same, computed; also sets `key` when `identified_by` is absent | `active_key` (and `key`) from the block's return | `on_conflict` applies |
|
|
34
|
+
| `set(workflow_key: "x")` | one run's key, verbatim | `key` | none |
|
|
35
|
+
|
|
36
|
+
- `identified_by` only names the run; use it so `workflow_runs.for(card)` finds the run
|
|
37
|
+
whatever the other arguments were.
|
|
38
|
+
- `unique_by` takes the same forms as `identified_by`: names or a block. Alone, it also names
|
|
39
|
+
the run, so one declaration covers both columns. Declare both only when identity and
|
|
40
|
+
uniqueness must differ; then the two derive independently.
|
|
41
|
+
- An unknown parameter name raises `ArgumentError` at declaration or first enqueue.
|
|
42
|
+
- Identity inside a hash argument (a webhook payload): use the block form.
|
|
43
|
+
|
|
44
|
+
`on_conflict`, checked at `perform_later` inside the caller's transaction, decided for
|
|
45
|
+
concurrent enqueues by the unique index on `[job_class, active_key]`:
|
|
46
|
+
|
|
47
|
+
| Value | Behaviour | Use for |
|
|
48
|
+
|---|---|---|
|
|
49
|
+
| `:skip` (default) | enqueues nothing, `perform_later` returns `false` | sweeps that start runs, double clicks, webhook retries |
|
|
50
|
+
| `:reject` | raises `ActiveJob::Durable::RunAlreadyExists` (`error.run`) | payouts, anything with money |
|
|
51
|
+
| `:replace` | `cancel!` the live run, start a new one, one transaction | reschedules, renewals, "the date moved" |
|
|
52
|
+
|
|
53
|
+
## 3. Timers: `wait:` versus `wait_until:`
|
|
54
|
+
|
|
55
|
+
```ruby
|
|
56
|
+
step :remind, wait_until: license.expires_at - 2.weeks # at a time
|
|
57
|
+
step :revoke, wait: 2.weeks # after the previous step completed
|
|
58
|
+
step :lazy, wait_until: -> { expensive_query.date } # callable: evaluated only when the step is reached
|
|
59
|
+
```
|
|
60
|
+
|
|
61
|
+
| | `wait:` | `wait_until:` |
|
|
62
|
+
|---|---|---|
|
|
63
|
+
| Takes | a Duration (or callable) | a Time (or callable) |
|
|
64
|
+
| Anchored | once, at the previous step's completion (run start for a first step) | re-evaluated at every pass, including wake |
|
|
65
|
+
| Target moves | does not move | re-arms by itself: `license.expires_at` changed, the run parks again |
|
|
66
|
+
| Target in the past | runs at once | runs at once |
|
|
67
|
+
|
|
68
|
+
A plain expression on the `step` line is evaluated on every pass, because `perform` re-runs on
|
|
69
|
+
each resume and skips completed steps; use a callable only when the expression is expensive.
|
|
70
|
+
|
|
71
|
+
Parked run: status `waiting`, `wake_at` set, nothing in the queue. `WakeJob` wakes due runs;
|
|
72
|
+
precision is its interval. `run.wake_up` (no arguments) ends the wait now. `run.cancel!` stops
|
|
73
|
+
the clock for that run.
|
|
74
|
+
|
|
75
|
+
## 4. Signals: `await` and `wake_up`
|
|
76
|
+
|
|
77
|
+
```ruby
|
|
78
|
+
def perform(import)
|
|
79
|
+
await :confirmation, wait: 10.minutes # deadline uses the timer keywords
|
|
80
|
+
return import.destroy! unless confirmed
|
|
81
|
+
step :apply
|
|
82
|
+
end
|
|
83
|
+
|
|
84
|
+
def confirmation(signal) = self.confirmed = signal.presence # nil at the deadline
|
|
85
|
+
|
|
86
|
+
# elsewhere
|
|
87
|
+
import.job_run.wake_up(:confirmation, true)
|
|
88
|
+
```
|
|
89
|
+
|
|
90
|
+
- `await :name` is a step with an empty body; the method `name(value)` (or a block
|
|
91
|
+
`await(:name) { |value| ... }`) runs on wake with the value. `nil` means the deadline passed.
|
|
92
|
+
- `wake_up(name, value)`: any JSON value. Sent before the `await` line is reached, it waits in
|
|
93
|
+
`pending_signals` and is consumed without parking. A second signal for the same name
|
|
94
|
+
overwrites. A run that is not live raises `ActiveJob::Durable::NotLive`.
|
|
95
|
+
- The value is kept on the step row, so a crash in a later step replays the same argument.
|
|
96
|
+
- `halt!` inside the handler leaves the await incomplete; `resume!` awaits again with a fresh
|
|
97
|
+
deadline.
|
|
98
|
+
- `wake_up` with no name wakes whatever is parked: a timer ends now, an await receives `nil`.
|
|
99
|
+
- Names must be unique per run; no `await` in a loop with a fixed name.
|
|
100
|
+
|
|
101
|
+
## 5. Stopping: `halt_on`, `halt!`, `resume!`, `cancel!`
|
|
102
|
+
|
|
103
|
+
| Call | From | Status after | Then |
|
|
104
|
+
|---|---|---|---|
|
|
105
|
+
| `retry_on E` | class | `enqueued` (retry wait) | automatic |
|
|
106
|
+
| `discard_on E` | class | `discarded` | nothing; terminal |
|
|
107
|
+
| `halt_on E` | class | `halted`, `error_class`, `error_message`, step row `halted` with cursor | `run.resume!` |
|
|
108
|
+
| `halt!(reason)` | inside a step | `halted`, `halt_reason` | `run.resume!` |
|
|
109
|
+
| uncaught error | | `failed`, error columns | `run.resume!` or a backend retry |
|
|
110
|
+
| `run.cancel!` | outside | `cancelled`, `active_key` cleared | nothing; terminal |
|
|
111
|
+
| `return` from `perform` | inside | `completed` | nothing |
|
|
112
|
+
|
|
113
|
+
Declared handlers win over resume: an error with a `discard_on`/`retry_on`/`halt_on` goes to
|
|
114
|
+
that handler; only undeclared errors are resumed by Continuation. No `resume_job` override.
|
|
115
|
+
|
|
116
|
+
`resume!` re-enqueues at the halted or failed step, from its cursor, as a new attempt; earlier
|
|
117
|
+
attempts keep their errors; manual resumes do not count toward `max_resumptions`. `cancel!` of a
|
|
118
|
+
running job takes effect at its next checkpoint or step boundary.
|
|
119
|
+
|
|
120
|
+
## 6. Statuses
|
|
121
|
+
|
|
122
|
+
| Group | Statuses | Scope |
|
|
123
|
+
|---|---|---|
|
|
124
|
+
| live | `enqueued` (in the queue), `running` (in a worker), `waiting` (timer), `awaiting` (signal) | `live` |
|
|
125
|
+
| attention | `failed`, `halted` | `attention`; resumable |
|
|
126
|
+
| terminal | `completed`, `discarded`, `cancelled` | `terminal` |
|
|
127
|
+
|
|
128
|
+
`enqueued` is written at every re-enqueue (isolated step, graceful stop, `retry_on`). An
|
|
129
|
+
`enqueued` run with `started_at` null never ran. `stuck_for(1.hour)`: `running` with a stale
|
|
130
|
+
heartbeat, or `enqueued`/`waiting`/`awaiting` with a stale transition.
|
|
131
|
+
|
|
132
|
+
## 7. Reading runs
|
|
133
|
+
|
|
134
|
+
```ruby
|
|
135
|
+
MyJob.workflow_runs # this class, newest first
|
|
136
|
+
MyJob.workflow_runs.for(card) # the key perform_later(card) derives
|
|
137
|
+
MyJob.workflow_runs.for(workflow_key: "cards/42")
|
|
138
|
+
MyJob.workflow_runs.for(card).live.first # or .sole under unique_by
|
|
139
|
+
MyJob.workflow_runs.awaiting.at_step(:review) # the review queue
|
|
140
|
+
MyJob.workflow_runs.halted / .failed / .attention
|
|
141
|
+
MyJob.workflow_runs.stuck_for(1.hour)
|
|
142
|
+
run.status, run.current_step, run.completed_steps, run.state, run.wake_at, run.halt_reason
|
|
143
|
+
run.steps # one row per attempt: name, attempt, status, cursor, timings, error
|
|
144
|
+
```
|
|
145
|
+
|
|
146
|
+
Give the model a method for its run instead of a class-level helper:
|
|
147
|
+
|
|
148
|
+
```ruby
|
|
149
|
+
def job_run = ImportJob.workflow_runs.for(self).live.sole
|
|
150
|
+
```
|
|
151
|
+
|
|
152
|
+
## 8. Callbacks
|
|
153
|
+
|
|
154
|
+
`before_step`, `after_step`, `around_step`: `ActiveSupport::Callbacks`, the `before_perform`
|
|
155
|
+
shape. They run for steps that execute, not for skipped ones; `after_step` only on completion.
|
|
156
|
+
`current_step` is the running `Step`. `throw :abort` in `before_step` skips the body and marks
|
|
157
|
+
the step completed.
|
|
@@ -0,0 +1,130 @@
|
|
|
1
|
+
# Decision guide: workflow, sweep, or hybrid
|
|
2
|
+
|
|
3
|
+
## 1. The three ways to do "later"
|
|
4
|
+
|
|
5
|
+
| | Cron sweep | Parked run per entity | Hybrid |
|
|
6
|
+
|---|---|---|---|
|
|
7
|
+
| Shape | a recurring job scans a table for due rows | one `Durable` run per entity, `wait_until:` | a recurring job starts runs for entities due within a window |
|
|
8
|
+
| Where the schedule lives | cron YAML + a scope | the job's `perform`, top to bottom | both, the window in cron, the steps in the job |
|
|
9
|
+
| Idempotency | by hand: claim columns, bucket arithmetic | by `unique_by` and completed steps | `unique_by ..., on_conflict: :skip` makes the sweep idempotent |
|
|
10
|
+
| Reschedule | data-driven, free (the next scan re-derives) | `perform_later` again with `:replace`, or the moved `wait_until:` re-arms on wake | same as parked |
|
|
11
|
+
| "Where is X?" | a query on the entity table | `workflow_runs.for(x)` | same |
|
|
12
|
+
| Rows at rest | none | one run row per live entity | one per entity inside the window |
|
|
13
|
+
| Precision | the cron interval | the `WakeJob` interval (one minute) | one minute |
|
|
14
|
+
| Best when | the rule is a policy over a table, or rows number in the millions | the wait belongs to one entity and has steps before or after it | many entities, far-future dates |
|
|
15
|
+
|
|
16
|
+
Rules:
|
|
17
|
+
|
|
18
|
+
- **Park when the wait belongs to one entity and is part of a sequence.** A license lifecycle
|
|
19
|
+
(remind, expire, revoke), an appointment reminder, a 30-day grace period, a payout waiting for
|
|
20
|
+
a webhook.
|
|
21
|
+
- **Sweep when the rule is a policy over a table.** "Delete chats older than 7 days", "post the
|
|
22
|
+
weekly digest", "rebuild the facts nightly". No entity has a position in it.
|
|
23
|
+
- **Hybrid when both are true:** thousands of entities with dates years ahead. Start a run only
|
|
24
|
+
for entities due within N days; `unique_by :entity` with the default `:skip` makes the daily
|
|
25
|
+
starter idempotent. Cancel on the entity's terminal event.
|
|
26
|
+
|
|
27
|
+
A parked run costs one row and nothing in the queue: the run is parked on the row, not in the
|
|
28
|
+
adapter, and `WakeJob` wakes due runs. So "too many jobs waiting" is not the parked design's
|
|
29
|
+
cost anymore; the cost is rows, and rows are cheap. What still argues for a sweep is the rule's
|
|
30
|
+
nature, not volume.
|
|
31
|
+
|
|
32
|
+
## 2. Replacing a sweep: the checklist
|
|
33
|
+
|
|
34
|
+
Before you replace a working sweep, confirm each:
|
|
35
|
+
|
|
36
|
+
1. Each row the sweep touches has an owner entity and a start event (created, cancelled,
|
|
37
|
+
confirmed) where `perform_later` can be called.
|
|
38
|
+
2. There is a terminal event where the run should be cancelled (submitted, reactivated, paid).
|
|
39
|
+
Without one, runs accumulate as `waiting` and the sweep was right.
|
|
40
|
+
3. The sweep's idempotency tricks (claim columns, `expires_at` buckets, `updated_at` compared
|
|
41
|
+
at wake) exist only because the parked job had no identity. `unique_by` removes them; if they
|
|
42
|
+
exist for another reason, keep them.
|
|
43
|
+
4. The precision of one minute is enough. It is for e-mail, Slack, billing; it is not for
|
|
44
|
+
sub-minute SLAs.
|
|
45
|
+
5. If the sweep also heals data (re-imports charges the webhook missed), keep the sweep as the
|
|
46
|
+
safety net and add the run for observability; both can coexist.
|
|
47
|
+
|
|
48
|
+
If 1 or 2 fails, keep the sweep. If you keep a sweep and want observability, make the sweep
|
|
49
|
+
itself a `Durable` job with a cursor; the run row shows progress and the last error.
|
|
50
|
+
|
|
51
|
+
## 3. Signals: `await` versus a state column plus a controller
|
|
52
|
+
|
|
53
|
+
Today's shape: a state column (`unconfirmed → scheduled → done`), a controller action that flips
|
|
54
|
+
it and enqueues the second half, a vacuum for the ones nobody confirmed. Problems: the second
|
|
55
|
+
half has no memory of the first, the deadline is enforced by a daily job, two entry points can
|
|
56
|
+
race.
|
|
57
|
+
|
|
58
|
+
`await` keeps one job with the wait in the middle. Use it when:
|
|
59
|
+
|
|
60
|
+
- the flow has steps before and after the wait (parse then apply; withdraw then credit);
|
|
61
|
+
- the wait has a deadline (`wait: 10.minutes`, `wait: 3.days`) or must be observable
|
|
62
|
+
(`workflow_runs.awaiting.at_step(:review)` is the review queue);
|
|
63
|
+
- the signal carries a value the later steps need (a verdict, a payment amount).
|
|
64
|
+
|
|
65
|
+
Keep the controller-flips-a-state shape when the "signal" is the only thing that ever happens
|
|
66
|
+
(no second half), or when the decision must survive without any job (an audit record). Both can
|
|
67
|
+
coexist: write the decision to the model and `wake_up` the run.
|
|
68
|
+
|
|
69
|
+
### Repeated decisions
|
|
70
|
+
|
|
71
|
+
`await` is a step, and step names are unique per run: an `await :decision` inside a loop is
|
|
72
|
+
completed after the first pass and skipped on every later one. Three cases:
|
|
73
|
+
|
|
74
|
+
| The decision | Shape |
|
|
75
|
+
|---|---|
|
|
76
|
+
| needs no value, only a go (the choice is already a row somewhere) | `halt!(:reason)` inside the loop; `resume!` from the outside. A halt is not a step; it can happen any number of times. |
|
|
77
|
+
| carries a value, from a set known when `perform` runs | one `await` per name, e.g. per approver |
|
|
78
|
+
| carries a value, an unknown number of times | write the value to a model, then `halt!`/`resume!`; the run reads the model when it continues |
|
|
79
|
+
|
|
80
|
+
A known set, N approvers snapshotted at submission:
|
|
81
|
+
|
|
82
|
+
```ruby
|
|
83
|
+
def perform(time_off)
|
|
84
|
+
@time_off = time_off
|
|
85
|
+
step :notify_approvers
|
|
86
|
+
time_off.approvals.each do |approval|
|
|
87
|
+
await :"decision_#{approval.approver_id}", wait: 3.business_days
|
|
88
|
+
end
|
|
89
|
+
step :finalize
|
|
90
|
+
end
|
|
91
|
+
|
|
92
|
+
def method_missing(name, decision = nil)
|
|
93
|
+
return super unless name.start_with?("decision_")
|
|
94
|
+
halt!(:approver_silent) if decision.nil? # the deadline passed
|
|
95
|
+
@time_off.approvals.find_by!(approver_id: name.to_s.delete_prefix("decision_")).decide!(decision)
|
|
96
|
+
end
|
|
97
|
+
|
|
98
|
+
# controller: time_off.approval_run.wake_up(:"decision_#{current_member.id}", params[:decision])
|
|
99
|
+
```
|
|
100
|
+
|
|
101
|
+
The awaits run in order; a rejection can end the run early from its handler. Approvers are
|
|
102
|
+
notified together in the first step, so waiting in order costs nothing visible.
|
|
103
|
+
|
|
104
|
+
## 4. `halt!` versus `await` versus `retry_on`
|
|
105
|
+
|
|
106
|
+
| Situation | Use |
|
|
107
|
+
|---|---|
|
|
108
|
+
| a transient error (network, lock) | `retry_on` |
|
|
109
|
+
| an error nothing can fix (corrupt file) | `discard_on` |
|
|
110
|
+
| an error a person fixes, then the run continues (out of storage, wrong config) | `halt_on Error`, then `run.resume!` |
|
|
111
|
+
| a decision that needs no data, only a go (tool approval already recorded elsewhere) | `halt!(:reason)`, then `run.resume!` |
|
|
112
|
+
| a decision that carries data (approve/reject, a webhook payload) | `await :name`, then `run.wake_up(:name, value)` |
|
|
113
|
+
| a fixed wait | `step :x, wait: 2.weeks` |
|
|
114
|
+
| a wait until a date on the record | `step :x, wait_until: record.date` |
|
|
115
|
+
|
|
116
|
+
## 5. `isolated: true`
|
|
117
|
+
|
|
118
|
+
Use `isolated: true` on steps that call slow external services (LLM, third-party API) so each
|
|
119
|
+
step runs as its own job execution and no single run holds a worker for minutes. Not needed for
|
|
120
|
+
steps that touch only the database. Isolation costs one re-enqueue per step.
|
|
121
|
+
|
|
122
|
+
## 6. When not to use AJ/DC
|
|
123
|
+
|
|
124
|
+
- Fan-out and fan-in (one step that spawns N jobs and waits for all). Use Solid Queue batches
|
|
125
|
+
(1.7+), GoodJob batches, or Sidekiq Pro; AJ/DC has no batch primitive yet. A `Durable` run may
|
|
126
|
+
start the batch and `await` the batch's `on_finish` job as a signal.
|
|
127
|
+
- Replay-based determinism or exactly-once semantics. Steps are at-least-once; a completed step
|
|
128
|
+
is skipped on resume. Make step bodies idempotent (upserts, `find_or_create_by`, guards on a
|
|
129
|
+
fact column).
|
|
130
|
+
- Sub-minute timers. The clock ticks with the scheduler.
|
|
@@ -0,0 +1,98 @@
|
|
|
1
|
+
# Installing AJ/DC
|
|
2
|
+
|
|
3
|
+
## 1. The gem
|
|
4
|
+
|
|
5
|
+
```sh
|
|
6
|
+
bundle add ajdc
|
|
7
|
+
```
|
|
8
|
+
|
|
9
|
+
Requirements: Ruby 3.3+, Rails 8.1+ (`attribute` on jobs needs Rails 8.2), SQLite, PostgreSQL or
|
|
10
|
+
MySQL, and a queue adapter with a scheduler for the recurring wake job (Solid Queue recurring
|
|
11
|
+
tasks, sidekiq-cron, GoodJob cron).
|
|
12
|
+
|
|
13
|
+
## 2. The tables: one database or two
|
|
14
|
+
|
|
15
|
+
AJ/DC writes two tables, `active_job_durable_runs` and `active_job_durable_steps`. Decide
|
|
16
|
+
where they live before generating anything.
|
|
17
|
+
|
|
18
|
+
| | Primary database (default) | Separate database |
|
|
19
|
+
|---|---|---|
|
|
20
|
+
| Command | `bin/rails generate ajdc:install` then `db:migrate` | `bin/rails generate ajdc:install --database=NAME` then `db:prepare` |
|
|
21
|
+
| Files | one migration in `db/migrate` | one migration in `db/NAME_migrate`, plus `config/initializers/active_job_durable.rb` |
|
|
22
|
+
| `perform_later` inside a transaction | the run row commits with the record that starts it; a rollback removes both | two databases, two transactions: a rollback of the record can leave an `enqueued` run |
|
|
23
|
+
| `unique_by` conflicts | decided in the caller's transaction | decided in the durable database's transaction |
|
|
24
|
+
| Choose it when | the default; any app where runs start from model callbacks or `with_lock` blocks | the app already keeps `solid_queue`, `solid_cache` in their own databases and wants run rows off the primary; or the primary is not writable from workers |
|
|
25
|
+
|
|
26
|
+
Recommendation: the primary database unless there is a stated reason. The transaction property
|
|
27
|
+
in `transactions.md` §1 depends on it.
|
|
28
|
+
|
|
29
|
+
### Primary database
|
|
30
|
+
|
|
31
|
+
```sh
|
|
32
|
+
bin/rails generate ajdc:install
|
|
33
|
+
bin/rails db:migrate
|
|
34
|
+
```
|
|
35
|
+
|
|
36
|
+
### Separate database
|
|
37
|
+
|
|
38
|
+
1. Declare the database in `config/database.yml` with its own `migrations_paths`:
|
|
39
|
+
|
|
40
|
+
```yaml
|
|
41
|
+
production:
|
|
42
|
+
primary:
|
|
43
|
+
<<: *default
|
|
44
|
+
durable:
|
|
45
|
+
<<: *default
|
|
46
|
+
database: storage/production_durable.sqlite3 # or the adapter's settings
|
|
47
|
+
migrations_paths: db/durable_migrate
|
|
48
|
+
```
|
|
49
|
+
|
|
50
|
+
Repeat for `development` and `test`.
|
|
51
|
+
|
|
52
|
+
2. Generate with the database name; the migration lands in that path and the initializer is
|
|
53
|
+
written:
|
|
54
|
+
|
|
55
|
+
```sh
|
|
56
|
+
bin/rails generate ajdc:install --database=durable
|
|
57
|
+
bin/rails db:prepare
|
|
58
|
+
```
|
|
59
|
+
|
|
60
|
+
```ruby
|
|
61
|
+
# config/initializers/active_job_durable.rb (generated)
|
|
62
|
+
ActiveJob::Durable.connects_to = { database: { writing: :durable } }
|
|
63
|
+
```
|
|
64
|
+
|
|
65
|
+
3. In step bodies that start a run from inside a model transaction, guard the first step
|
|
66
|
+
against a record that was rolled back (`return unless record.persisted?`), since the run
|
|
67
|
+
row can exist without it.
|
|
68
|
+
|
|
69
|
+
## 3. The clock
|
|
70
|
+
|
|
71
|
+
Timers and signal deadlines are woken by one recurring job. Without it, nothing that waits ever
|
|
72
|
+
wakes. Add it to the app's scheduler once:
|
|
73
|
+
|
|
74
|
+
```yaml
|
|
75
|
+
# config/recurring.yml (Solid Queue)
|
|
76
|
+
durable_wake:
|
|
77
|
+
class: ActiveJob::Durable::WakeJob
|
|
78
|
+
schedule: every minute
|
|
79
|
+
```
|
|
80
|
+
|
|
81
|
+
sidekiq-cron, GoodJob cron and others: schedule `ActiveJob::Durable::WakeJob.perform_later`
|
|
82
|
+
every minute. The interval is the precision of every `wait:`, `wait_until:` and `await`
|
|
83
|
+
deadline.
|
|
84
|
+
|
|
85
|
+
## 4. Verify
|
|
86
|
+
|
|
87
|
+
```sh
|
|
88
|
+
bin/rails runner 'puts ActiveJob::Durable::Run.count' # 0, no error
|
|
89
|
+
bin/rails runner 'puts ActiveJob::Durable.wake_up_due' # 0, no error
|
|
90
|
+
```
|
|
91
|
+
|
|
92
|
+
Then convert one job: `include ActiveJob::Durable` in place of `ActiveJob::Continuable`, run
|
|
93
|
+
its tests, enqueue it once, and read `MyJob.workflow_runs.last`.
|
|
94
|
+
|
|
95
|
+
## 5. Mission Control
|
|
96
|
+
|
|
97
|
+
If `mission_control-jobs` is mounted, the runs are plain Active Record rows next to the jobs;
|
|
98
|
+
`ActiveJob::Durable::Run` has no UI of its own yet.
|
|
@@ -0,0 +1,152 @@
|
|
|
1
|
+
# Testing durable workflows
|
|
2
|
+
|
|
3
|
+
## 1. Setup
|
|
4
|
+
|
|
5
|
+
```ruby
|
|
6
|
+
class ImportJobTest < ActiveSupport::TestCase
|
|
7
|
+
include ActiveJob::TestHelper # perform_enqueued_jobs, assert_enqueued_jobs
|
|
8
|
+
include ActiveJob::Continuation::TestHelper # interrupt_job_during_step / after_step (Rails 8.1)
|
|
9
|
+
|
|
10
|
+
end
|
|
11
|
+
```
|
|
12
|
+
|
|
13
|
+
- Queue adapter `:test`. Use the block form, `perform_enqueued_jobs do ... end`: it performs
|
|
14
|
+
every job enqueued inside the block, including the re-enqueues at isolated steps, interrupts
|
|
15
|
+
and resumes, so a durable run executes until it parks, halts, fails or completes. Do not count
|
|
16
|
+
executions.
|
|
17
|
+
- Freeze time with `travel_to`; timers compare `wake_at` with `Time.current`.
|
|
18
|
+
- Each test starts from an empty runs table (`Run.delete_all` in setup if fixtures are not
|
|
19
|
+
transactional for the durable database).
|
|
20
|
+
|
|
21
|
+
## 2. Drive a run to its next park
|
|
22
|
+
|
|
23
|
+
```ruby
|
|
24
|
+
perform_enqueued_jobs { ImportJob.perform_later(import) } # runs through every step and re-enqueue
|
|
25
|
+
|
|
26
|
+
run = ImportJob.workflow_runs.for(import).sole
|
|
27
|
+
assert_equal "completed", run.status
|
|
28
|
+
assert_equal %w[check process], run.completed_steps
|
|
29
|
+
assert_equal "imported", import.reload.state # the logic, first of all
|
|
30
|
+
```
|
|
31
|
+
|
|
32
|
+
Assert the outcome on your models, then the run's `status` and `completed_steps`. Do not
|
|
33
|
+
assert what happens between executions; that is the gem's business.
|
|
34
|
+
|
|
35
|
+
## 3. Crash and resume (the SIGKILL test)
|
|
36
|
+
|
|
37
|
+
Make a step fail once after progress, then assert the cursor and the resume:
|
|
38
|
+
|
|
39
|
+
```ruby
|
|
40
|
+
stub_import_to_fail_on("c") # your own stub, once
|
|
41
|
+
perform_enqueued_jobs { ImportJob.perform_later(import, %w[a b c d e]) }
|
|
42
|
+
|
|
43
|
+
run = ImportJob.workflow_runs.for(import).sole
|
|
44
|
+
assert_equal "failed", run.status
|
|
45
|
+
assert_equal %w[a b], import.reload.imported_items
|
|
46
|
+
|
|
47
|
+
perform_enqueued_jobs { run.resume! }
|
|
48
|
+
assert_equal %w[a b c d e], import.reload.imported_items # c, d, e once; a, b not twice
|
|
49
|
+
assert_equal "completed", run.reload.status
|
|
50
|
+
```
|
|
51
|
+
|
|
52
|
+
Graceful interrupt (deploy), Rails' own helper:
|
|
53
|
+
|
|
54
|
+
```ruby
|
|
55
|
+
perform_enqueued_jobs do
|
|
56
|
+
interrupt_job_during_step(ImportJob, :process, cursor: 2) { ImportJob.perform_later(import, items) }
|
|
57
|
+
end
|
|
58
|
+
assert_equal items, import.reload.imported_items # nothing lost, nothing doubled
|
|
59
|
+
```
|
|
60
|
+
|
|
61
|
+
## 4. Halt and resume
|
|
62
|
+
|
|
63
|
+
```ruby
|
|
64
|
+
perform_enqueued_jobs { ImportJob.perform_later(import) }
|
|
65
|
+
run = ImportJob.workflow_runs.for(import).sole
|
|
66
|
+
assert_equal "halted", run.status
|
|
67
|
+
assert_equal "InsufficientStorageError", run.error_class # halt_on
|
|
68
|
+
# or: assert_equal "tool_approval", run.halt_reason # halt!
|
|
69
|
+
|
|
70
|
+
fix_the_world
|
|
71
|
+
perform_enqueued_jobs { run.resume! }
|
|
72
|
+
assert_equal "completed", run.reload.status
|
|
73
|
+
```
|
|
74
|
+
|
|
75
|
+
## 5. Timers
|
|
76
|
+
|
|
77
|
+
```ruby
|
|
78
|
+
travel_to now do
|
|
79
|
+
perform_enqueued_jobs { License::LifecycleJob.perform_later(license) } # reaches :remind, parks
|
|
80
|
+
run = License::LifecycleJob.workflow_runs.for(license).sole
|
|
81
|
+
assert_equal "waiting", run.status
|
|
82
|
+
assert_equal license.expires_at - 2.weeks, run.wake_at
|
|
83
|
+
assert_no_enqueued_emails
|
|
84
|
+
end
|
|
85
|
+
|
|
86
|
+
travel_to license.expires_at - 2.weeks do
|
|
87
|
+
perform_enqueued_jobs { ActiveJob::Durable.wake_up_due } # the clock, then :remind runs, parks at :expire
|
|
88
|
+
assert_enqueued_email_with LicenseMailer, :expiring_two_weeks, args: [license]
|
|
89
|
+
end
|
|
90
|
+
```
|
|
91
|
+
|
|
92
|
+
Test what the step did, then where the run parked next. Test re-arming: move the date, tick
|
|
93
|
+
the clock, assert the reminder did not go out. Test `wake_up` with no name: the step runs now.
|
|
94
|
+
Wake with `ActiveJob::Durable.wake_up_due` inside `travel_to`; do not assert how many runs it
|
|
95
|
+
woke or call the wake job directly.
|
|
96
|
+
|
|
97
|
+
## 6. Signals
|
|
98
|
+
|
|
99
|
+
```ruby
|
|
100
|
+
perform_enqueued_jobs { BulkImportJob.perform_later(import) } # parks at :confirmation
|
|
101
|
+
run = import.job_run
|
|
102
|
+
assert_equal "awaiting", run.status
|
|
103
|
+
|
|
104
|
+
perform_enqueued_jobs { run.wake_up(:confirmation, true) }
|
|
105
|
+
assert_equal "applied", import.reload.state
|
|
106
|
+
```
|
|
107
|
+
|
|
108
|
+
Also test the deadline: `travel_to` past it, `perform_enqueued_jobs { ActiveJob::Durable.wake_up_due }`,
|
|
109
|
+
the import is destroyed. And the early signal: `wake_up` before the first execution, then the run
|
|
110
|
+
completes without parking. Assert on the model; the run's `status` is the second assertion.
|
|
111
|
+
|
|
112
|
+
## 7. Uniqueness
|
|
113
|
+
|
|
114
|
+
```ruby
|
|
115
|
+
assert ImportJob.perform_later(import)
|
|
116
|
+
assert_equal false, ImportJob.perform_later(import) # :skip
|
|
117
|
+
assert_equal 1, Run.count
|
|
118
|
+
```
|
|
119
|
+
|
|
120
|
+
`:reject`: `assert_raises(ActiveJob::Durable::RunAlreadyExists)`. `:replace`: the first run is
|
|
121
|
+
`cancelled`, the second is `enqueued`, `Run.count == 2`. After `cancel!` or completion, a new
|
|
122
|
+
`perform_later` creates a new run.
|
|
123
|
+
|
|
124
|
+
## 8. Transactions
|
|
125
|
+
|
|
126
|
+
```ruby
|
|
127
|
+
Account.transaction do
|
|
128
|
+
account.cancel # creates the cancellation, perform_later inside with_lock
|
|
129
|
+
assert_enqueued_jobs 0 # not yet: enqueue waits for commit
|
|
130
|
+
assert_equal 1, Run.count # the row is already there
|
|
131
|
+
end
|
|
132
|
+
assert_enqueued_jobs 1
|
|
133
|
+
|
|
134
|
+
Account.transaction do
|
|
135
|
+
account.cancel
|
|
136
|
+
raise ActiveRecord::Rollback
|
|
137
|
+
end
|
|
138
|
+
assert_equal 0, Run.count
|
|
139
|
+
assert_enqueued_jobs 0
|
|
140
|
+
```
|
|
141
|
+
|
|
142
|
+
## 9. Cancel
|
|
143
|
+
|
|
144
|
+
Cancel a run, then perform: the model is untouched and the run is `cancelled`. Cooperative cancel
|
|
145
|
+
of a running step is the gem's feature; do not test it in the app.
|
|
146
|
+
|
|
147
|
+
## 10. What not to test
|
|
148
|
+
|
|
149
|
+
Test logic, not the gem. Do not assert on `parked_job`, `pending_signals`, `resumptions`, step
|
|
150
|
+
row attempts, cursors, the number of executions, or how many runs the clock woke. Assert on your
|
|
151
|
+
models first, then on `status`, `current_step`, `completed_steps`, `state`, `wake_at`,
|
|
152
|
+
`halt_reason` and `error_class` when the test is about them.
|