ajdc 0.0.1 → 0.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (39) hide show
  1. checksums.yaml +4 -4
  2. data/CHANGELOG.md +4 -0
  3. data/README.md +380 -13
  4. data/db/durable_schema.rb +57 -0
  5. data/lib/active_job/continuable/callbacks.rb +48 -0
  6. data/lib/active_job/continuable/configurable.rb +75 -0
  7. data/lib/active_job/continuation/callbacks.rb +18 -0
  8. data/lib/active_job/continuation/rescue_handlers_first.rb +29 -0
  9. data/lib/active_job/durable/args_mapper.rb +41 -0
  10. data/lib/active_job/durable/config.rb +100 -0
  11. data/lib/active_job/durable/continuation.rb +24 -0
  12. data/lib/active_job/durable/execution.rb +438 -0
  13. data/lib/active_job/durable/housekeeping_job.rb +16 -0
  14. data/lib/active_job/durable/record.rb +17 -0
  15. data/lib/active_job/durable/run.rb +165 -0
  16. data/lib/active_job/durable/step.rb +14 -0
  17. data/lib/active_job/durable/wake_job.rb +16 -0
  18. data/lib/active_job/durable.rb +251 -0
  19. data/lib/ajdc/railtie.rb +18 -0
  20. data/lib/ajdc/version.rb +1 -1
  21. data/lib/ajdc.rb +1 -0
  22. data/lib/generators/ajdc/install/USAGE +19 -0
  23. data/lib/generators/ajdc/install/install_generator.rb +66 -0
  24. data/lib/generators/ajdc/install/templates/create_active_job_durable_tables.rb.tt +56 -0
  25. data/lib/generators/ajdc/install/templates/initializer.rb.tt +4 -0
  26. data/skills/ajdc/SKILL.md +117 -0
  27. data/skills/ajdc/examples/README.md +11 -0
  28. data/skills/ajdc/examples/agent-loop.md +57 -0
  29. data/skills/ajdc/examples/halting.md +84 -0
  30. data/skills/ajdc/examples/pipeline.md +73 -0
  31. data/skills/ajdc/examples/signals.md +139 -0
  32. data/skills/ajdc/examples/timers.md +136 -0
  33. data/skills/ajdc/examples/uniqueness.md +87 -0
  34. data/skills/ajdc/references/api.md +157 -0
  35. data/skills/ajdc/references/decision-guide.md +130 -0
  36. data/skills/ajdc/references/installation.md +98 -0
  37. data/skills/ajdc/references/testing.md +152 -0
  38. data/skills/ajdc/references/transactions.md +68 -0
  39. metadata +83 -7
@@ -0,0 +1,87 @@
1
+ # Uniqueness: one live run per identity
2
+
3
+ `unique_by` names the `perform` parameters that identify a run and forbids a second live run
4
+ for the same identity. The check runs at `perform_later`, inside the caller's transaction; the
5
+ unique index on `[job_class, active_key]` settles concurrent enqueues. Terminal runs
6
+ (`completed`, `discarded`, `cancelled`) free the identity.
7
+
8
+ ## `:skip` (default): the second start is a no-op
9
+
10
+ For a sweep that starts runs, a button clicked twice, a webhook delivered twice.
11
+
12
+ ```ruby
13
+ class License::LifecycleJob < ApplicationJob
14
+ include ActiveJob::Durable
15
+
16
+ unique_by :license
17
+
18
+ def perform(license) = # ...
19
+ end
20
+
21
+ # a daily starter for licenses expiring within 30 days; running it twice starts nothing new
22
+ License.expiring_within(30.days).find_each { License::LifecycleJob.perform_later(it) }
23
+ License::LifecycleJob.perform_later(license) # => false while a run is live
24
+ ```
25
+
26
+ ## `:replace`: the new start cancels the live run
27
+
28
+ For a reschedule or a renewal, when the date the run waits for has moved.
29
+
30
+ ```ruby
31
+ class AppointmentReminderJob < ApplicationJob
32
+ include ActiveJob::Durable
33
+
34
+ unique_by :appointment, on_conflict: :replace
35
+
36
+ def perform(appointment)
37
+ step :remind, wait_until: appointment.start_at - 30.minutes do
38
+ AppointmentNotification.with(appointment:).deliver_later(appointment.patient)
39
+ end
40
+ end
41
+ end
42
+
43
+ # PatientAppointment: after_create_commit, and after_update_commit when start_at changed:
44
+ # AppointmentReminderJob.perform_later(self) # the old run is cancelled, a new one parks
45
+ ```
46
+
47
+ ## `:reject`: the second start is an error
48
+
49
+ For money and anything where a duplicate must be seen, not swallowed.
50
+
51
+ ```ruby
52
+ class PayoutJob < ApplicationJob
53
+ include ActiveJob::Durable
54
+
55
+ unique_by :payout, on_conflict: :reject
56
+
57
+ def perform(payout, wallet_id, amount) = # ...
58
+ end
59
+
60
+ PayoutJob.perform_later(payout, wallet_id, amount)
61
+ PayoutJob.perform_later(payout, wallet_id, amount) # raises ActiveJob::Durable::RunAlreadyExists; error.run is the live one
62
+ ```
63
+
64
+ `unique_by :payout` also lets a webhook find the run with the payout alone:
65
+ `payout.job_run.wake_up(:payment, ...)`.
66
+
67
+ ## Identity inside an argument
68
+
69
+ When the identity is not a `perform` parameter but a value inside one, name it with a block.
70
+ A Stripe webhook payload:
71
+
72
+ ```ruby
73
+ class Payment::ChargeImportJob < ApplicationJob
74
+ include ActiveJob::Durable
75
+
76
+ unique_by { |payload| payload.dig("data", "object", "id") } # the charge id; :skip
77
+
78
+ def perform(payload)
79
+ step :import do
80
+ Payment::StripeChargeImport.call(payload.dig("data", "object"))
81
+ end
82
+ end
83
+ end
84
+ ```
85
+
86
+ Keep an idempotency guard inside the step when a duplicate may arrive after the run completed:
87
+ uniqueness covers live and attention runs only.
@@ -0,0 +1,157 @@
1
+ # AJ/DC API reference (differences that matter)
2
+
3
+ ## 1. Setup
4
+
5
+ ```ruby
6
+ gem "ajdc"
7
+ bin/rails generate ajdc:install # migration for the two tables in the primary database
8
+ bin/rails generate ajdc:install --database=durable # or in a separate database, plus the initializer
9
+ ```
10
+
11
+ Which database, and why it matters: `installation.md` §2.
12
+
13
+ ```yaml
14
+ # config/recurring.yml (Solid Queue) or the app's scheduler
15
+ durable_wake:
16
+ class: ActiveJob::Durable::WakeJob
17
+ schedule: every minute
18
+ ```
19
+
20
+ `include ActiveJob::Durable` in the job. It includes `Continuable`; `step`, `start:`,
21
+ `isolated:`, `step.cursor`, `step.set!`, `step.advance!`, `step.checkpoint!`, `attribute`,
22
+ `retry_on`, `discard_on` are unchanged. The payload carries only the run id; progress and
23
+ attributes live on the run row.
24
+
25
+ ## 2. Identity: `identified_by` versus `unique_by` versus `set(workflow_key:)`
26
+
27
+ | Macro | Answers | Column | Effect on enqueue |
28
+ |---|---|---|---|
29
+ | none | "what is this run about?" | `key` = all arguments, positional in order then keywords `name=value` | none |
30
+ | `identified_by :card` | same, narrowed to the named `perform` parameters | `key` from those parameters only | none |
31
+ | `identified_by { \|card, **kw\| [card, kw.fetch(:format, "csv")] }` | same, computed | `key` from the block's return | none |
32
+ | `unique_by :card` | "may two live runs exist for this?" | `active_key` while live or attention; `NULL` when terminal | `on_conflict` applies |
33
+ | `unique_by { \|payload\| payload.dig("data", "object", "id") }` | same, computed; also sets `key` when `identified_by` is absent | `active_key` (and `key`) from the block's return | `on_conflict` applies |
34
+ | `set(workflow_key: "x")` | one run's key, verbatim | `key` | none |
35
+
36
+ - `identified_by` only names the run; use it so `workflow_runs.for(card)` finds the run
37
+ whatever the other arguments were.
38
+ - `unique_by` takes the same forms as `identified_by`: names or a block. Alone, it also names
39
+ the run, so one declaration covers both columns. Declare both only when identity and
40
+ uniqueness must differ; then the two derive independently.
41
+ - An unknown parameter name raises `ArgumentError` at declaration or first enqueue.
42
+ - Identity inside a hash argument (a webhook payload): use the block form.
43
+
44
+ `on_conflict`, checked at `perform_later` inside the caller's transaction, decided for
45
+ concurrent enqueues by the unique index on `[job_class, active_key]`:
46
+
47
+ | Value | Behaviour | Use for |
48
+ |---|---|---|
49
+ | `:skip` (default) | enqueues nothing, `perform_later` returns `false` | sweeps that start runs, double clicks, webhook retries |
50
+ | `:reject` | raises `ActiveJob::Durable::RunAlreadyExists` (`error.run`) | payouts, anything with money |
51
+ | `:replace` | `cancel!` the live run, start a new one, one transaction | reschedules, renewals, "the date moved" |
52
+
53
+ ## 3. Timers: `wait:` versus `wait_until:`
54
+
55
+ ```ruby
56
+ step :remind, wait_until: license.expires_at - 2.weeks # at a time
57
+ step :revoke, wait: 2.weeks # after the previous step completed
58
+ step :lazy, wait_until: -> { expensive_query.date } # callable: evaluated only when the step is reached
59
+ ```
60
+
61
+ | | `wait:` | `wait_until:` |
62
+ |---|---|---|
63
+ | Takes | a Duration (or callable) | a Time (or callable) |
64
+ | Anchored | once, at the previous step's completion (run start for a first step) | re-evaluated at every pass, including wake |
65
+ | Target moves | does not move | re-arms by itself: `license.expires_at` changed, the run parks again |
66
+ | Target in the past | runs at once | runs at once |
67
+
68
+ A plain expression on the `step` line is evaluated on every pass, because `perform` re-runs on
69
+ each resume and skips completed steps; use a callable only when the expression is expensive.
70
+
71
+ Parked run: status `waiting`, `wake_at` set, nothing in the queue. `WakeJob` wakes due runs;
72
+ precision is its interval. `run.wake_up` (no arguments) ends the wait now. `run.cancel!` stops
73
+ the clock for that run.
74
+
75
+ ## 4. Signals: `await` and `wake_up`
76
+
77
+ ```ruby
78
+ def perform(import)
79
+ await :confirmation, wait: 10.minutes # deadline uses the timer keywords
80
+ return import.destroy! unless confirmed
81
+ step :apply
82
+ end
83
+
84
+ def confirmation(signal) = self.confirmed = signal.presence # nil at the deadline
85
+
86
+ # elsewhere
87
+ import.job_run.wake_up(:confirmation, true)
88
+ ```
89
+
90
+ - `await :name` is a step with an empty body; the method `name(value)` (or a block
91
+ `await(:name) { |value| ... }`) runs on wake with the value. `nil` means the deadline passed.
92
+ - `wake_up(name, value)`: any JSON value. Sent before the `await` line is reached, it waits in
93
+ `pending_signals` and is consumed without parking. A second signal for the same name
94
+ overwrites. A run that is not live raises `ActiveJob::Durable::NotLive`.
95
+ - The value is kept on the step row, so a crash in a later step replays the same argument.
96
+ - `halt!` inside the handler leaves the await incomplete; `resume!` awaits again with a fresh
97
+ deadline.
98
+ - `wake_up` with no name wakes whatever is parked: a timer ends now, an await receives `nil`.
99
+ - Names must be unique per run; no `await` in a loop with a fixed name.
100
+
101
+ ## 5. Stopping: `halt_on`, `halt!`, `resume!`, `cancel!`
102
+
103
+ | Call | From | Status after | Then |
104
+ |---|---|---|---|
105
+ | `retry_on E` | class | `enqueued` (retry wait) | automatic |
106
+ | `discard_on E` | class | `discarded` | nothing; terminal |
107
+ | `halt_on E` | class | `halted`, `error_class`, `error_message`, step row `halted` with cursor | `run.resume!` |
108
+ | `halt!(reason)` | inside a step | `halted`, `halt_reason` | `run.resume!` |
109
+ | uncaught error | | `failed`, error columns | `run.resume!` or a backend retry |
110
+ | `run.cancel!` | outside | `cancelled`, `active_key` cleared | nothing; terminal |
111
+ | `return` from `perform` | inside | `completed` | nothing |
112
+
113
+ Declared handlers win over resume: an error with a `discard_on`/`retry_on`/`halt_on` goes to
114
+ that handler; only undeclared errors are resumed by Continuation. No `resume_job` override.
115
+
116
+ `resume!` re-enqueues at the halted or failed step, from its cursor, as a new attempt; earlier
117
+ attempts keep their errors; manual resumes do not count toward `max_resumptions`. `cancel!` of a
118
+ running job takes effect at its next checkpoint or step boundary.
119
+
120
+ ## 6. Statuses
121
+
122
+ | Group | Statuses | Scope |
123
+ |---|---|---|
124
+ | live | `enqueued` (in the queue), `running` (in a worker), `waiting` (timer), `awaiting` (signal) | `live` |
125
+ | attention | `failed`, `halted` | `attention`; resumable |
126
+ | terminal | `completed`, `discarded`, `cancelled` | `terminal` |
127
+
128
+ `enqueued` is written at every re-enqueue (isolated step, graceful stop, `retry_on`). An
129
+ `enqueued` run with `started_at` null never ran. `stuck_for(1.hour)`: `running` with a stale
130
+ heartbeat, or `enqueued`/`waiting`/`awaiting` with a stale transition.
131
+
132
+ ## 7. Reading runs
133
+
134
+ ```ruby
135
+ MyJob.workflow_runs # this class, newest first
136
+ MyJob.workflow_runs.for(card) # the key perform_later(card) derives
137
+ MyJob.workflow_runs.for(workflow_key: "cards/42")
138
+ MyJob.workflow_runs.for(card).live.first # or .sole under unique_by
139
+ MyJob.workflow_runs.awaiting.at_step(:review) # the review queue
140
+ MyJob.workflow_runs.halted / .failed / .attention
141
+ MyJob.workflow_runs.stuck_for(1.hour)
142
+ run.status, run.current_step, run.completed_steps, run.state, run.wake_at, run.halt_reason
143
+ run.steps # one row per attempt: name, attempt, status, cursor, timings, error
144
+ ```
145
+
146
+ Give the model a method for its run instead of a class-level helper:
147
+
148
+ ```ruby
149
+ def job_run = ImportJob.workflow_runs.for(self).live.sole
150
+ ```
151
+
152
+ ## 8. Callbacks
153
+
154
+ `before_step`, `after_step`, `around_step`: `ActiveSupport::Callbacks`, the `before_perform`
155
+ shape. They run for steps that execute, not for skipped ones; `after_step` only on completion.
156
+ `current_step` is the running `Step`. `throw :abort` in `before_step` skips the body and marks
157
+ the step completed.
@@ -0,0 +1,130 @@
1
+ # Decision guide: workflow, sweep, or hybrid
2
+
3
+ ## 1. The three ways to do "later"
4
+
5
+ | | Cron sweep | Parked run per entity | Hybrid |
6
+ |---|---|---|---|
7
+ | Shape | a recurring job scans a table for due rows | one `Durable` run per entity, `wait_until:` | a recurring job starts runs for entities due within a window |
8
+ | Where the schedule lives | cron YAML + a scope | the job's `perform`, top to bottom | both, the window in cron, the steps in the job |
9
+ | Idempotency | by hand: claim columns, bucket arithmetic | by `unique_by` and completed steps | `unique_by ..., on_conflict: :skip` makes the sweep idempotent |
10
+ | Reschedule | data-driven, free (the next scan re-derives) | `perform_later` again with `:replace`, or the moved `wait_until:` re-arms on wake | same as parked |
11
+ | "Where is X?" | a query on the entity table | `workflow_runs.for(x)` | same |
12
+ | Rows at rest | none | one run row per live entity | one per entity inside the window |
13
+ | Precision | the cron interval | the `WakeJob` interval (one minute) | one minute |
14
+ | Best when | the rule is a policy over a table, or rows number in the millions | the wait belongs to one entity and has steps before or after it | many entities, far-future dates |
15
+
16
+ Rules:
17
+
18
+ - **Park when the wait belongs to one entity and is part of a sequence.** A license lifecycle
19
+ (remind, expire, revoke), an appointment reminder, a 30-day grace period, a payout waiting for
20
+ a webhook.
21
+ - **Sweep when the rule is a policy over a table.** "Delete chats older than 7 days", "post the
22
+ weekly digest", "rebuild the facts nightly". No entity has a position in it.
23
+ - **Hybrid when both are true:** thousands of entities with dates years ahead. Start a run only
24
+ for entities due within N days; `unique_by :entity` with the default `:skip` makes the daily
25
+ starter idempotent. Cancel on the entity's terminal event.
26
+
27
+ A parked run costs one row and nothing in the queue: the run is parked on the row, not in the
28
+ adapter, and `WakeJob` wakes due runs. So "too many jobs waiting" is not the parked design's
29
+ cost anymore; the cost is rows, and rows are cheap. What still argues for a sweep is the rule's
30
+ nature, not volume.
31
+
32
+ ## 2. Replacing a sweep: the checklist
33
+
34
+ Before you replace a working sweep, confirm each:
35
+
36
+ 1. Each row the sweep touches has an owner entity and a start event (created, cancelled,
37
+ confirmed) where `perform_later` can be called.
38
+ 2. There is a terminal event where the run should be cancelled (submitted, reactivated, paid).
39
+ Without one, runs accumulate as `waiting` and the sweep was right.
40
+ 3. The sweep's idempotency tricks (claim columns, `expires_at` buckets, `updated_at` compared
41
+ at wake) exist only because the parked job had no identity. `unique_by` removes them; if they
42
+ exist for another reason, keep them.
43
+ 4. The precision of one minute is enough. It is for e-mail, Slack, billing; it is not for
44
+ sub-minute SLAs.
45
+ 5. If the sweep also heals data (re-imports charges the webhook missed), keep the sweep as the
46
+ safety net and add the run for observability; both can coexist.
47
+
48
+ If 1 or 2 fails, keep the sweep. If you keep a sweep and want observability, make the sweep
49
+ itself a `Durable` job with a cursor; the run row shows progress and the last error.
50
+
51
+ ## 3. Signals: `await` versus a state column plus a controller
52
+
53
+ Today's shape: a state column (`unconfirmed → scheduled → done`), a controller action that flips
54
+ it and enqueues the second half, a vacuum for the ones nobody confirmed. Problems: the second
55
+ half has no memory of the first, the deadline is enforced by a daily job, two entry points can
56
+ race.
57
+
58
+ `await` keeps one job with the wait in the middle. Use it when:
59
+
60
+ - the flow has steps before and after the wait (parse then apply; withdraw then credit);
61
+ - the wait has a deadline (`wait: 10.minutes`, `wait: 3.days`) or must be observable
62
+ (`workflow_runs.awaiting.at_step(:review)` is the review queue);
63
+ - the signal carries a value the later steps need (a verdict, a payment amount).
64
+
65
+ Keep the controller-flips-a-state shape when the "signal" is the only thing that ever happens
66
+ (no second half), or when the decision must survive without any job (an audit record). Both can
67
+ coexist: write the decision to the model and `wake_up` the run.
68
+
69
+ ### Repeated decisions
70
+
71
+ `await` is a step, and step names are unique per run: an `await :decision` inside a loop is
72
+ completed after the first pass and skipped on every later one. Three cases:
73
+
74
+ | The decision | Shape |
75
+ |---|---|
76
+ | needs no value, only a go (the choice is already a row somewhere) | `halt!(:reason)` inside the loop; `resume!` from the outside. A halt is not a step; it can happen any number of times. |
77
+ | carries a value, from a set known when `perform` runs | one `await` per name, e.g. per approver |
78
+ | carries a value, an unknown number of times | write the value to a model, then `halt!`/`resume!`; the run reads the model when it continues |
79
+
80
+ A known set, N approvers snapshotted at submission:
81
+
82
+ ```ruby
83
+ def perform(time_off)
84
+ @time_off = time_off
85
+ step :notify_approvers
86
+ time_off.approvals.each do |approval|
87
+ await :"decision_#{approval.approver_id}", wait: 3.business_days
88
+ end
89
+ step :finalize
90
+ end
91
+
92
+ def method_missing(name, decision = nil)
93
+ return super unless name.start_with?("decision_")
94
+ halt!(:approver_silent) if decision.nil? # the deadline passed
95
+ @time_off.approvals.find_by!(approver_id: name.to_s.delete_prefix("decision_")).decide!(decision)
96
+ end
97
+
98
+ # controller: time_off.approval_run.wake_up(:"decision_#{current_member.id}", params[:decision])
99
+ ```
100
+
101
+ The awaits run in order; a rejection can end the run early from its handler. Approvers are
102
+ notified together in the first step, so waiting in order costs nothing visible.
103
+
104
+ ## 4. `halt!` versus `await` versus `retry_on`
105
+
106
+ | Situation | Use |
107
+ |---|---|
108
+ | a transient error (network, lock) | `retry_on` |
109
+ | an error nothing can fix (corrupt file) | `discard_on` |
110
+ | an error a person fixes, then the run continues (out of storage, wrong config) | `halt_on Error`, then `run.resume!` |
111
+ | a decision that needs no data, only a go (tool approval already recorded elsewhere) | `halt!(:reason)`, then `run.resume!` |
112
+ | a decision that carries data (approve/reject, a webhook payload) | `await :name`, then `run.wake_up(:name, value)` |
113
+ | a fixed wait | `step :x, wait: 2.weeks` |
114
+ | a wait until a date on the record | `step :x, wait_until: record.date` |
115
+
116
+ ## 5. `isolated: true`
117
+
118
+ Use `isolated: true` on steps that call slow external services (LLM, third-party API) so each
119
+ step runs as its own job execution and no single run holds a worker for minutes. Not needed for
120
+ steps that touch only the database. Isolation costs one re-enqueue per step.
121
+
122
+ ## 6. When not to use AJ/DC
123
+
124
+ - Fan-out and fan-in (one step that spawns N jobs and waits for all). Use Solid Queue batches
125
+ (1.7+), GoodJob batches, or Sidekiq Pro; AJ/DC has no batch primitive yet. A `Durable` run may
126
+ start the batch and `await` the batch's `on_finish` job as a signal.
127
+ - Replay-based determinism or exactly-once semantics. Steps are at-least-once; a completed step
128
+ is skipped on resume. Make step bodies idempotent (upserts, `find_or_create_by`, guards on a
129
+ fact column).
130
+ - Sub-minute timers. The clock ticks with the scheduler.
@@ -0,0 +1,98 @@
1
+ # Installing AJ/DC
2
+
3
+ ## 1. The gem
4
+
5
+ ```sh
6
+ bundle add ajdc
7
+ ```
8
+
9
+ Requirements: Ruby 3.3+, Rails 8.1+ (`attribute` on jobs needs Rails 8.2), SQLite, PostgreSQL or
10
+ MySQL, and a queue adapter with a scheduler for the recurring wake job (Solid Queue recurring
11
+ tasks, sidekiq-cron, GoodJob cron).
12
+
13
+ ## 2. The tables: one database or two
14
+
15
+ AJ/DC writes two tables, `active_job_durable_runs` and `active_job_durable_steps`. Decide
16
+ where they live before generating anything.
17
+
18
+ | | Primary database (default) | Separate database |
19
+ |---|---|---|
20
+ | Command | `bin/rails generate ajdc:install` then `db:migrate` | `bin/rails generate ajdc:install --database=NAME` then `db:prepare` |
21
+ | Files | one migration in `db/migrate` | one migration in `db/NAME_migrate`, plus `config/initializers/active_job_durable.rb` |
22
+ | `perform_later` inside a transaction | the run row commits with the record that starts it; a rollback removes both | two databases, two transactions: a rollback of the record can leave an `enqueued` run |
23
+ | `unique_by` conflicts | decided in the caller's transaction | decided in the durable database's transaction |
24
+ | Choose it when | the default; any app where runs start from model callbacks or `with_lock` blocks | the app already keeps `solid_queue`, `solid_cache` in their own databases and wants run rows off the primary; or the primary is not writable from workers |
25
+
26
+ Recommendation: the primary database unless there is a stated reason. The transaction property
27
+ in `transactions.md` §1 depends on it.
28
+
29
+ ### Primary database
30
+
31
+ ```sh
32
+ bin/rails generate ajdc:install
33
+ bin/rails db:migrate
34
+ ```
35
+
36
+ ### Separate database
37
+
38
+ 1. Declare the database in `config/database.yml` with its own `migrations_paths`:
39
+
40
+ ```yaml
41
+ production:
42
+ primary:
43
+ <<: *default
44
+ durable:
45
+ <<: *default
46
+ database: storage/production_durable.sqlite3 # or the adapter's settings
47
+ migrations_paths: db/durable_migrate
48
+ ```
49
+
50
+ Repeat for `development` and `test`.
51
+
52
+ 2. Generate with the database name; the migration lands in that path and the initializer is
53
+ written:
54
+
55
+ ```sh
56
+ bin/rails generate ajdc:install --database=durable
57
+ bin/rails db:prepare
58
+ ```
59
+
60
+ ```ruby
61
+ # config/initializers/active_job_durable.rb (generated)
62
+ ActiveJob::Durable.connects_to = { database: { writing: :durable } }
63
+ ```
64
+
65
+ 3. In step bodies that start a run from inside a model transaction, guard the first step
66
+ against a record that was rolled back (`return unless record.persisted?`), since the run
67
+ row can exist without it.
68
+
69
+ ## 3. The clock
70
+
71
+ Timers and signal deadlines are woken by one recurring job. Without it, nothing that waits ever
72
+ wakes. Add it to the app's scheduler once:
73
+
74
+ ```yaml
75
+ # config/recurring.yml (Solid Queue)
76
+ durable_wake:
77
+ class: ActiveJob::Durable::WakeJob
78
+ schedule: every minute
79
+ ```
80
+
81
+ sidekiq-cron, GoodJob cron and others: schedule `ActiveJob::Durable::WakeJob.perform_later`
82
+ every minute. The interval is the precision of every `wait:`, `wait_until:` and `await`
83
+ deadline.
84
+
85
+ ## 4. Verify
86
+
87
+ ```sh
88
+ bin/rails runner 'puts ActiveJob::Durable::Run.count' # 0, no error
89
+ bin/rails runner 'puts ActiveJob::Durable.wake_up_due' # 0, no error
90
+ ```
91
+
92
+ Then convert one job: `include ActiveJob::Durable` in place of `ActiveJob::Continuable`, run
93
+ its tests, enqueue it once, and read `MyJob.workflow_runs.last`.
94
+
95
+ ## 5. Mission Control
96
+
97
+ If `mission_control-jobs` is mounted, the runs are plain Active Record rows next to the jobs;
98
+ `ActiveJob::Durable::Run` has no UI of its own yet.
@@ -0,0 +1,152 @@
1
+ # Testing durable workflows
2
+
3
+ ## 1. Setup
4
+
5
+ ```ruby
6
+ class ImportJobTest < ActiveSupport::TestCase
7
+ include ActiveJob::TestHelper # perform_enqueued_jobs, assert_enqueued_jobs
8
+ include ActiveJob::Continuation::TestHelper # interrupt_job_during_step / after_step (Rails 8.1)
9
+
10
+ end
11
+ ```
12
+
13
+ - Queue adapter `:test`. Use the block form, `perform_enqueued_jobs do ... end`: it performs
14
+ every job enqueued inside the block, including the re-enqueues at isolated steps, interrupts
15
+ and resumes, so a durable run executes until it parks, halts, fails or completes. Do not count
16
+ executions.
17
+ - Freeze time with `travel_to`; timers compare `wake_at` with `Time.current`.
18
+ - Each test starts from an empty runs table (`Run.delete_all` in setup if fixtures are not
19
+ transactional for the durable database).
20
+
21
+ ## 2. Drive a run to its next park
22
+
23
+ ```ruby
24
+ perform_enqueued_jobs { ImportJob.perform_later(import) } # runs through every step and re-enqueue
25
+
26
+ run = ImportJob.workflow_runs.for(import).sole
27
+ assert_equal "completed", run.status
28
+ assert_equal %w[check process], run.completed_steps
29
+ assert_equal "imported", import.reload.state # the logic, first of all
30
+ ```
31
+
32
+ Assert the outcome on your models, then the run's `status` and `completed_steps`. Do not
33
+ assert what happens between executions; that is the gem's business.
34
+
35
+ ## 3. Crash and resume (the SIGKILL test)
36
+
37
+ Make a step fail once after progress, then assert the cursor and the resume:
38
+
39
+ ```ruby
40
+ stub_import_to_fail_on("c") # your own stub, once
41
+ perform_enqueued_jobs { ImportJob.perform_later(import, %w[a b c d e]) }
42
+
43
+ run = ImportJob.workflow_runs.for(import).sole
44
+ assert_equal "failed", run.status
45
+ assert_equal %w[a b], import.reload.imported_items
46
+
47
+ perform_enqueued_jobs { run.resume! }
48
+ assert_equal %w[a b c d e], import.reload.imported_items # c, d, e once; a, b not twice
49
+ assert_equal "completed", run.reload.status
50
+ ```
51
+
52
+ Graceful interrupt (deploy), Rails' own helper:
53
+
54
+ ```ruby
55
+ perform_enqueued_jobs do
56
+ interrupt_job_during_step(ImportJob, :process, cursor: 2) { ImportJob.perform_later(import, items) }
57
+ end
58
+ assert_equal items, import.reload.imported_items # nothing lost, nothing doubled
59
+ ```
60
+
61
+ ## 4. Halt and resume
62
+
63
+ ```ruby
64
+ perform_enqueued_jobs { ImportJob.perform_later(import) }
65
+ run = ImportJob.workflow_runs.for(import).sole
66
+ assert_equal "halted", run.status
67
+ assert_equal "InsufficientStorageError", run.error_class # halt_on
68
+ # or: assert_equal "tool_approval", run.halt_reason # halt!
69
+
70
+ fix_the_world
71
+ perform_enqueued_jobs { run.resume! }
72
+ assert_equal "completed", run.reload.status
73
+ ```
74
+
75
+ ## 5. Timers
76
+
77
+ ```ruby
78
+ travel_to now do
79
+ perform_enqueued_jobs { License::LifecycleJob.perform_later(license) } # reaches :remind, parks
80
+ run = License::LifecycleJob.workflow_runs.for(license).sole
81
+ assert_equal "waiting", run.status
82
+ assert_equal license.expires_at - 2.weeks, run.wake_at
83
+ assert_no_enqueued_emails
84
+ end
85
+
86
+ travel_to license.expires_at - 2.weeks do
87
+ perform_enqueued_jobs { ActiveJob::Durable.wake_up_due } # the clock, then :remind runs, parks at :expire
88
+ assert_enqueued_email_with LicenseMailer, :expiring_two_weeks, args: [license]
89
+ end
90
+ ```
91
+
92
+ Test what the step did, then where the run parked next. Test re-arming: move the date, tick
93
+ the clock, assert the reminder did not go out. Test `wake_up` with no name: the step runs now.
94
+ Wake with `ActiveJob::Durable.wake_up_due` inside `travel_to`; do not assert how many runs it
95
+ woke or call the wake job directly.
96
+
97
+ ## 6. Signals
98
+
99
+ ```ruby
100
+ perform_enqueued_jobs { BulkImportJob.perform_later(import) } # parks at :confirmation
101
+ run = import.job_run
102
+ assert_equal "awaiting", run.status
103
+
104
+ perform_enqueued_jobs { run.wake_up(:confirmation, true) }
105
+ assert_equal "applied", import.reload.state
106
+ ```
107
+
108
+ Also test the deadline: `travel_to` past it, `perform_enqueued_jobs { ActiveJob::Durable.wake_up_due }`,
109
+ the import is destroyed. And the early signal: `wake_up` before the first execution, then the run
110
+ completes without parking. Assert on the model; the run's `status` is the second assertion.
111
+
112
+ ## 7. Uniqueness
113
+
114
+ ```ruby
115
+ assert ImportJob.perform_later(import)
116
+ assert_equal false, ImportJob.perform_later(import) # :skip
117
+ assert_equal 1, Run.count
118
+ ```
119
+
120
+ `:reject`: `assert_raises(ActiveJob::Durable::RunAlreadyExists)`. `:replace`: the first run is
121
+ `cancelled`, the second is `enqueued`, `Run.count == 2`. After `cancel!` or completion, a new
122
+ `perform_later` creates a new run.
123
+
124
+ ## 8. Transactions
125
+
126
+ ```ruby
127
+ Account.transaction do
128
+ account.cancel # creates the cancellation, perform_later inside with_lock
129
+ assert_enqueued_jobs 0 # not yet: enqueue waits for commit
130
+ assert_equal 1, Run.count # the row is already there
131
+ end
132
+ assert_enqueued_jobs 1
133
+
134
+ Account.transaction do
135
+ account.cancel
136
+ raise ActiveRecord::Rollback
137
+ end
138
+ assert_equal 0, Run.count
139
+ assert_enqueued_jobs 0
140
+ ```
141
+
142
+ ## 9. Cancel
143
+
144
+ Cancel a run, then perform: the model is untouched and the run is `cancelled`. Cooperative cancel
145
+ of a running step is the gem's feature; do not test it in the app.
146
+
147
+ ## 10. What not to test
148
+
149
+ Test logic, not the gem. Do not assert on `parked_job`, `pending_signals`, `resumptions`, step
150
+ row attempts, cursors, the number of executions, or how many runs the clock woke. Assert on your
151
+ models first, then on `status`, `current_step`, `completed_steps`, `state`, `wake_at`,
152
+ `halt_reason` and `error_class` when the test is about them.