patchwork-os 1.2.0-beta.2.canary.682 → 1.2.0-beta.2.canary.686
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/butler/outcomeShadowLog.d.ts +15 -0
- package/dist/butler/outcomeShadowLog.js +38 -5
- package/dist/butler/outcomeShadowLog.js.map +1 -1
- package/dist/butler/promoteShadowOutcomes.d.ts +122 -0
- package/dist/butler/promoteShadowOutcomes.js +184 -0
- package/dist/butler/promoteShadowOutcomes.js.map +1 -0
- package/dist/connectors/todoist.d.ts +101 -14
- package/dist/connectors/todoist.js +32 -7
- package/dist/connectors/todoist.js.map +1 -1
- package/dist/index.js +31 -2
- package/dist/index.js.map +1 -1
- package/dist/recipes/cronClaim.d.ts +183 -0
- package/dist/recipes/cronClaim.js +241 -0
- package/dist/recipes/cronClaim.js.map +1 -0
- package/dist/recipes/scheduler.d.ts +63 -6
- package/dist/recipes/scheduler.js +115 -9
- package/dist/recipes/scheduler.js.map +1 -1
- package/dist/recipes/tools/todoist.d.ts +7 -1
- package/dist/recipes/tools/todoist.js +32 -19
- package/dist/recipes/tools/todoist.js.map +1 -1
- package/dist/workers/runWorkerShadow.d.ts +0 -11
- package/dist/workers/runWorkerShadow.js +26 -3
- package/dist/workers/runWorkerShadow.js.map +1 -1
- package/package.json +2 -2
|
@@ -0,0 +1,183 @@
|
|
|
1
|
+
/**
|
|
2
|
+
* Cross-process claim on a scheduled fire (#1458).
|
|
3
|
+
*
|
|
4
|
+
* ## The bug
|
|
5
|
+
*
|
|
6
|
+
* `RecipeScheduler` guards double-fire with an in-memory `Set`, and its own
|
|
7
|
+
* comment states the assumption: *"The guard is scheduler-scoped (one process),
|
|
8
|
+
* which is enough because that's the only place a cron tick can originate."*
|
|
9
|
+
*
|
|
10
|
+
* That assumption does not hold. The recipe store is global — `patchworkPath
|
|
11
|
+
* ("recipes")` — so **every** running bridge schedules **every** enabled cron
|
|
12
|
+
* recipe, and an in-process `Set` cannot see a sibling process. N bridges ⇒ N
|
|
13
|
+
* fires. Observed live on 2026-08-19: one hourly recipe ran twice at the same
|
|
14
|
+
* instant from two pids, and again the next hour.
|
|
15
|
+
*
|
|
16
|
+
* Two bridges is a supported, documented shape, so "run one bridge" is not the
|
|
17
|
+
* answer.
|
|
18
|
+
*
|
|
19
|
+
* ## The claim is on the TICK, not on the recipe
|
|
20
|
+
*
|
|
21
|
+
* This is the whole design, and it is what keeps manual runs working.
|
|
22
|
+
*
|
|
23
|
+
* A tick is an externally generated event — the clock — that N processes each
|
|
24
|
+
* observe independently, and whose correct execution count is one. A manual
|
|
25
|
+
* `patchwork recipe run X` is an operator-generated event that exists exactly
|
|
26
|
+
* once already. Deduping the second against the first would be a category
|
|
27
|
+
* error, and the scheduler's existing comment says manual runs deliberately
|
|
28
|
+
* bypass the guard.
|
|
29
|
+
*
|
|
30
|
+
* So the key is `(recipeName, slotEpochMs)` and the claim is taken by the cron
|
|
31
|
+
* path and by nothing else. Manual runs, HTTP `POST /recipes/:name/run`,
|
|
32
|
+
* webhooks and the file-watch/git-hook paths neither read nor write it. That is
|
|
33
|
+
* structural, not a carve-out in the key.
|
|
34
|
+
*
|
|
35
|
+
* ## Why the slot must be threaded in, never re-read from the clock
|
|
36
|
+
*
|
|
37
|
+
* `fire()` runs an event-loop hop after the cron matcher matched. Re-deriving
|
|
38
|
+
* the slot from `Date.now()` there would let bridge A compute `:00` and bridge
|
|
39
|
+
* B `:01` for the same tick — two keys, and the duplicate is back, in a form
|
|
40
|
+
* that reproduces only when the hop straddles a second boundary.
|
|
41
|
+
*
|
|
42
|
+
* node-cron hands the callback the matched instant with milliseconds already
|
|
43
|
+
* zeroed, identically in every process running the same expression. That value
|
|
44
|
+
* is threaded in as `slotEpochMs`, and floored again here so that an upstream
|
|
45
|
+
* change reintroducing milliseconds cannot silently split the key.
|
|
46
|
+
*
|
|
47
|
+
* **No slot ⇒ no claim.** That covers the `@every` interval path (a bare
|
|
48
|
+
* `setInterval`, phase-anchored to each process's own start, with no canonical
|
|
49
|
+
* slot to agree on) and the `fireForTest` hook. Both then behave exactly as
|
|
50
|
+
* they do today. `@every` is deliberately out of scope here rather than
|
|
51
|
+
* quantised: zero installed recipes use it, and the quantised version has a
|
|
52
|
+
* bias worth deciding on its own evidence. `RecipeScheduler` logs the exclusion
|
|
53
|
+
* once per recipe, because a scheduling gap nobody is told about is how
|
|
54
|
+
* `audit-in-flight` spent its whole life passing.
|
|
55
|
+
*
|
|
56
|
+
* ## The claim is a tombstone, not a lock
|
|
57
|
+
*
|
|
58
|
+
* It is never released. If it were released on completion, a peer whose tick is
|
|
59
|
+
* delayed past the first bridge's completion — a 40 ms recipe and ticks 200 ms
|
|
60
|
+
* apart is entirely reachable — would find no claim and fire. That reintroduces
|
|
61
|
+
* the bug in a narrower window, which is worse, because it stops reproducing on
|
|
62
|
+
* demand.
|
|
63
|
+
*
|
|
64
|
+
* Consequence, stated as a property rather than left as an oversight: this
|
|
65
|
+
* guarantees **at-most-once per slot per `PATCHWORK_HOME`, not exactly-once.**
|
|
66
|
+
* A process that dies between claiming and dispatching consumes the slot and no
|
|
67
|
+
* peer picks it up. Exactly-once needs a claim a peer can safely steal, which
|
|
68
|
+
* needs liveness plus a rule for re-running a run that got most of the way
|
|
69
|
+
* through its side effects — that is crash recovery, a different feature, and
|
|
70
|
+
* its failure mode is a duplicate of a run that DID have external effects.
|
|
71
|
+
*
|
|
72
|
+
* ## Failure is OPEN, and never silent
|
|
73
|
+
*
|
|
74
|
+
* `EEXIST` is not failure — it is the mechanism working, and it skips.
|
|
75
|
+
* "Failure" means the store is unusable: `EACCES`, `EROFS`, `ENOSPC`, or an
|
|
76
|
+
* unexpected throw. On those the tick FIRES, and the caller is told why so it
|
|
77
|
+
* can log it and stamp the run.
|
|
78
|
+
*
|
|
79
|
+
* The reasoning, since this is the opposite of ADR-0016's fail-closed instinct:
|
|
80
|
+
* the conditions that break this store are machine-level, so they break it for
|
|
81
|
+
* every bridge at once, and failing closed would then yield ZERO executions of
|
|
82
|
+
* every scheduled recipe rather than one — silently, for as long as it lasts.
|
|
83
|
+
* Fail-open's worst case is exactly the bug we have today, which is known and
|
|
84
|
+
* bounded. And this store shares a disk with the run log, the effect ledger and
|
|
85
|
+
* the decision record: if `~/.patchwork` is unwritable the recipe is going to
|
|
86
|
+
* run ungoverned or fail on its own, so stopping here selects a stop condition
|
|
87
|
+
* on a *proxy* for "the disk is broken" instead of reporting the disk.
|
|
88
|
+
*
|
|
89
|
+
* ADR-0016 fails closed because its object is a tool call, whose safe default is
|
|
90
|
+
* "no". This decides whether a clock tick has already been consumed by a peer,
|
|
91
|
+
* and with no answer available the safe default is the status quo ante. That is
|
|
92
|
+
* the same reasoning that makes the identity roster fail SOFT.
|
|
93
|
+
*
|
|
94
|
+
* `PATCHWORK_CRON_CLAIM_REQUIRED=1` flips it to fail-closed for deployments
|
|
95
|
+
* whose scheduled recipes send email or post publicly, where the operator knows
|
|
96
|
+
* the trade and we do not.
|
|
97
|
+
*/
|
|
98
|
+
/** Directory under PATCHWORK_HOME holding day-sharded claim files. */
|
|
99
|
+
export declare const CRON_CLAIM_DIRNAME = "cron-claims";
|
|
100
|
+
/**
|
|
101
|
+
* How long claim day-directories are kept.
|
|
102
|
+
*
|
|
103
|
+
* This does not need to cover a schedule's period. It only needs to outlive the
|
|
104
|
+
* window in which a duplicate of the SAME slot can arrive, which is bounded by
|
|
105
|
+
* seconds. 48 h is five orders of magnitude of margin; the sweep exists solely
|
|
106
|
+
* to bound growth.
|
|
107
|
+
*/
|
|
108
|
+
export declare const CLAIM_RETENTION_MS: number;
|
|
109
|
+
export type ClaimOutcome =
|
|
110
|
+
/** This process owns the slot. Fire. */
|
|
111
|
+
{
|
|
112
|
+
kind: "claimed";
|
|
113
|
+
}
|
|
114
|
+
/** A peer owns the slot. Do not fire. */
|
|
115
|
+
| {
|
|
116
|
+
kind: "taken";
|
|
117
|
+
}
|
|
118
|
+
/**
|
|
119
|
+
* The store is unusable. The caller fires anyway (unless the operator set
|
|
120
|
+
* `PATCHWORK_CRON_CLAIM_REQUIRED`), logs, and stamps the run so a resulting
|
|
121
|
+
* duplicate is attributable rather than mysterious.
|
|
122
|
+
*/
|
|
123
|
+
| {
|
|
124
|
+
kind: "unavailable";
|
|
125
|
+
reason: string;
|
|
126
|
+
}
|
|
127
|
+
/**
|
|
128
|
+
* The store is unusable AND the operator asked for fail-closed. Do not fire.
|
|
129
|
+
* Distinct from `taken` so the log can say which happened — "a peer has it"
|
|
130
|
+
* and "we could not tell" are different facts about the system.
|
|
131
|
+
*/
|
|
132
|
+
| {
|
|
133
|
+
kind: "refused";
|
|
134
|
+
reason: string;
|
|
135
|
+
};
|
|
136
|
+
/**
|
|
137
|
+
* Key for one scheduled fire.
|
|
138
|
+
*
|
|
139
|
+
* JSON-array encoding rather than `${a}:${b}`, for the reason `deriveScopeKey`
|
|
140
|
+
* already documents: recipe `a:b` at slot `c` and recipe `a` at slot `b:c` must
|
|
141
|
+
* not collide. Hex-only, so a recipe name containing `/`, spaces or unicode is
|
|
142
|
+
* still a safe filename — the human-readable fields live inside the record.
|
|
143
|
+
*
|
|
144
|
+
* The schedule expression is deliberately NOT in the key. If a schedule is
|
|
145
|
+
* edited and two bridges hot-reload at different moments they compute different
|
|
146
|
+
* slots and both fire regardless; including the expression would not fix that
|
|
147
|
+
* and would make a same-slot collision less likely to dedupe.
|
|
148
|
+
*/
|
|
149
|
+
export declare function cronClaimKey(recipeName: string, slotEpochMs: number): string;
|
|
150
|
+
export interface ClaimOptions {
|
|
151
|
+
/** Root override for tests. Defaults to `patchworkPath(CRON_CLAIM_DIRNAME)`. */
|
|
152
|
+
claimsDir?: string;
|
|
153
|
+
/**
|
|
154
|
+
* Fail-closed override. Defaults to reading
|
|
155
|
+
* `PATCHWORK_CRON_CLAIM_REQUIRED`. Injected so a test can drive both
|
|
156
|
+
* branches without mutating the environment.
|
|
157
|
+
*/
|
|
158
|
+
required?: boolean;
|
|
159
|
+
}
|
|
160
|
+
/**
|
|
161
|
+
* Try to claim `(recipeName, slotEpochMs)` for this process.
|
|
162
|
+
*
|
|
163
|
+
* `openSync(path, "wx")` is one atomic syscall that is simultaneously the test
|
|
164
|
+
* and the set — no lock file, no shared append-only log, no read-modify-write,
|
|
165
|
+
* and no serialisation between unrelated recipes. `withFileLockSync` was
|
|
166
|
+
* considered and rejected for exactly those costs, plus a 30 s stale-lock TTL
|
|
167
|
+
* that would stall every recipe's scheduling if a process died holding it.
|
|
168
|
+
*
|
|
169
|
+
* The record body is written for a human to read afterwards and is NEVER read
|
|
170
|
+
* to make a decision — the kernel already made it at `open`.
|
|
171
|
+
*/
|
|
172
|
+
export declare function claimCronSlot(recipeName: string, slotEpochMs: number, opts?: ClaimOptions): ClaimOutcome;
|
|
173
|
+
/**
|
|
174
|
+
* Delete claim day-directories older than the retention horizon.
|
|
175
|
+
*
|
|
176
|
+
* Best-effort and never throws: this is housekeeping, and a scheduler that
|
|
177
|
+
* cannot start because it could not tidy up would be a far worse bug than the
|
|
178
|
+
* disk usage it prevents.
|
|
179
|
+
*
|
|
180
|
+
* Returns the number of directories removed so a caller can log it — a sweep
|
|
181
|
+
* that silently does nothing looks identical to one that has nothing to do.
|
|
182
|
+
*/
|
|
183
|
+
export declare function sweepCronClaims(now?: number, opts?: ClaimOptions): number;
|
|
@@ -0,0 +1,241 @@
|
|
|
1
|
+
/**
|
|
2
|
+
* Cross-process claim on a scheduled fire (#1458).
|
|
3
|
+
*
|
|
4
|
+
* ## The bug
|
|
5
|
+
*
|
|
6
|
+
* `RecipeScheduler` guards double-fire with an in-memory `Set`, and its own
|
|
7
|
+
* comment states the assumption: *"The guard is scheduler-scoped (one process),
|
|
8
|
+
* which is enough because that's the only place a cron tick can originate."*
|
|
9
|
+
*
|
|
10
|
+
* That assumption does not hold. The recipe store is global — `patchworkPath
|
|
11
|
+
* ("recipes")` — so **every** running bridge schedules **every** enabled cron
|
|
12
|
+
* recipe, and an in-process `Set` cannot see a sibling process. N bridges ⇒ N
|
|
13
|
+
* fires. Observed live on 2026-08-19: one hourly recipe ran twice at the same
|
|
14
|
+
* instant from two pids, and again the next hour.
|
|
15
|
+
*
|
|
16
|
+
* Two bridges is a supported, documented shape, so "run one bridge" is not the
|
|
17
|
+
* answer.
|
|
18
|
+
*
|
|
19
|
+
* ## The claim is on the TICK, not on the recipe
|
|
20
|
+
*
|
|
21
|
+
* This is the whole design, and it is what keeps manual runs working.
|
|
22
|
+
*
|
|
23
|
+
* A tick is an externally generated event — the clock — that N processes each
|
|
24
|
+
* observe independently, and whose correct execution count is one. A manual
|
|
25
|
+
* `patchwork recipe run X` is an operator-generated event that exists exactly
|
|
26
|
+
* once already. Deduping the second against the first would be a category
|
|
27
|
+
* error, and the scheduler's existing comment says manual runs deliberately
|
|
28
|
+
* bypass the guard.
|
|
29
|
+
*
|
|
30
|
+
* So the key is `(recipeName, slotEpochMs)` and the claim is taken by the cron
|
|
31
|
+
* path and by nothing else. Manual runs, HTTP `POST /recipes/:name/run`,
|
|
32
|
+
* webhooks and the file-watch/git-hook paths neither read nor write it. That is
|
|
33
|
+
* structural, not a carve-out in the key.
|
|
34
|
+
*
|
|
35
|
+
* ## Why the slot must be threaded in, never re-read from the clock
|
|
36
|
+
*
|
|
37
|
+
* `fire()` runs an event-loop hop after the cron matcher matched. Re-deriving
|
|
38
|
+
* the slot from `Date.now()` there would let bridge A compute `:00` and bridge
|
|
39
|
+
* B `:01` for the same tick — two keys, and the duplicate is back, in a form
|
|
40
|
+
* that reproduces only when the hop straddles a second boundary.
|
|
41
|
+
*
|
|
42
|
+
* node-cron hands the callback the matched instant with milliseconds already
|
|
43
|
+
* zeroed, identically in every process running the same expression. That value
|
|
44
|
+
* is threaded in as `slotEpochMs`, and floored again here so that an upstream
|
|
45
|
+
* change reintroducing milliseconds cannot silently split the key.
|
|
46
|
+
*
|
|
47
|
+
* **No slot ⇒ no claim.** That covers the `@every` interval path (a bare
|
|
48
|
+
* `setInterval`, phase-anchored to each process's own start, with no canonical
|
|
49
|
+
* slot to agree on) and the `fireForTest` hook. Both then behave exactly as
|
|
50
|
+
* they do today. `@every` is deliberately out of scope here rather than
|
|
51
|
+
* quantised: zero installed recipes use it, and the quantised version has a
|
|
52
|
+
* bias worth deciding on its own evidence. `RecipeScheduler` logs the exclusion
|
|
53
|
+
* once per recipe, because a scheduling gap nobody is told about is how
|
|
54
|
+
* `audit-in-flight` spent its whole life passing.
|
|
55
|
+
*
|
|
56
|
+
* ## The claim is a tombstone, not a lock
|
|
57
|
+
*
|
|
58
|
+
* It is never released. If it were released on completion, a peer whose tick is
|
|
59
|
+
* delayed past the first bridge's completion — a 40 ms recipe and ticks 200 ms
|
|
60
|
+
* apart is entirely reachable — would find no claim and fire. That reintroduces
|
|
61
|
+
* the bug in a narrower window, which is worse, because it stops reproducing on
|
|
62
|
+
* demand.
|
|
63
|
+
*
|
|
64
|
+
* Consequence, stated as a property rather than left as an oversight: this
|
|
65
|
+
* guarantees **at-most-once per slot per `PATCHWORK_HOME`, not exactly-once.**
|
|
66
|
+
* A process that dies between claiming and dispatching consumes the slot and no
|
|
67
|
+
* peer picks it up. Exactly-once needs a claim a peer can safely steal, which
|
|
68
|
+
* needs liveness plus a rule for re-running a run that got most of the way
|
|
69
|
+
* through its side effects — that is crash recovery, a different feature, and
|
|
70
|
+
* its failure mode is a duplicate of a run that DID have external effects.
|
|
71
|
+
*
|
|
72
|
+
* ## Failure is OPEN, and never silent
|
|
73
|
+
*
|
|
74
|
+
* `EEXIST` is not failure — it is the mechanism working, and it skips.
|
|
75
|
+
* "Failure" means the store is unusable: `EACCES`, `EROFS`, `ENOSPC`, or an
|
|
76
|
+
* unexpected throw. On those the tick FIRES, and the caller is told why so it
|
|
77
|
+
* can log it and stamp the run.
|
|
78
|
+
*
|
|
79
|
+
* The reasoning, since this is the opposite of ADR-0016's fail-closed instinct:
|
|
80
|
+
* the conditions that break this store are machine-level, so they break it for
|
|
81
|
+
* every bridge at once, and failing closed would then yield ZERO executions of
|
|
82
|
+
* every scheduled recipe rather than one — silently, for as long as it lasts.
|
|
83
|
+
* Fail-open's worst case is exactly the bug we have today, which is known and
|
|
84
|
+
* bounded. And this store shares a disk with the run log, the effect ledger and
|
|
85
|
+
* the decision record: if `~/.patchwork` is unwritable the recipe is going to
|
|
86
|
+
* run ungoverned or fail on its own, so stopping here selects a stop condition
|
|
87
|
+
* on a *proxy* for "the disk is broken" instead of reporting the disk.
|
|
88
|
+
*
|
|
89
|
+
* ADR-0016 fails closed because its object is a tool call, whose safe default is
|
|
90
|
+
* "no". This decides whether a clock tick has already been consumed by a peer,
|
|
91
|
+
* and with no answer available the safe default is the status quo ante. That is
|
|
92
|
+
* the same reasoning that makes the identity roster fail SOFT.
|
|
93
|
+
*
|
|
94
|
+
* `PATCHWORK_CRON_CLAIM_REQUIRED=1` flips it to fail-closed for deployments
|
|
95
|
+
* whose scheduled recipes send email or post publicly, where the operator knows
|
|
96
|
+
* the trade and we do not.
|
|
97
|
+
*/
|
|
98
|
+
import { createHash } from "node:crypto";
|
|
99
|
+
import { closeSync, mkdirSync, openSync, readdirSync, rmSync, writeSync, } from "node:fs";
|
|
100
|
+
import { join } from "node:path";
|
|
101
|
+
import { patchworkPath } from "../patchworkHome.js";
|
|
102
|
+
/** Directory under PATCHWORK_HOME holding day-sharded claim files. */
|
|
103
|
+
export const CRON_CLAIM_DIRNAME = "cron-claims";
|
|
104
|
+
/**
|
|
105
|
+
* How long claim day-directories are kept.
|
|
106
|
+
*
|
|
107
|
+
* This does not need to cover a schedule's period. It only needs to outlive the
|
|
108
|
+
* window in which a duplicate of the SAME slot can arrive, which is bounded by
|
|
109
|
+
* seconds. 48 h is five orders of magnitude of margin; the sweep exists solely
|
|
110
|
+
* to bound growth.
|
|
111
|
+
*/
|
|
112
|
+
export const CLAIM_RETENTION_MS = 48 * 60 * 60 * 1000;
|
|
113
|
+
/**
|
|
114
|
+
* Key for one scheduled fire.
|
|
115
|
+
*
|
|
116
|
+
* JSON-array encoding rather than `${a}:${b}`, for the reason `deriveScopeKey`
|
|
117
|
+
* already documents: recipe `a:b` at slot `c` and recipe `a` at slot `b:c` must
|
|
118
|
+
* not collide. Hex-only, so a recipe name containing `/`, spaces or unicode is
|
|
119
|
+
* still a safe filename — the human-readable fields live inside the record.
|
|
120
|
+
*
|
|
121
|
+
* The schedule expression is deliberately NOT in the key. If a schedule is
|
|
122
|
+
* edited and two bridges hot-reload at different moments they compute different
|
|
123
|
+
* slots and both fire regardless; including the expression would not fix that
|
|
124
|
+
* and would make a same-slot collision less likely to dedupe.
|
|
125
|
+
*/
|
|
126
|
+
export function cronClaimKey(recipeName, slotEpochMs) {
|
|
127
|
+
return createHash("sha256")
|
|
128
|
+
.update(JSON.stringify(["cron", recipeName, slotEpochMs]))
|
|
129
|
+
.digest("hex")
|
|
130
|
+
.slice(0, 32);
|
|
131
|
+
}
|
|
132
|
+
/** `YYYY-MM-DD` in UTC. Day-sharded so the sweep deletes directories, not files. */
|
|
133
|
+
function dayShard(slotEpochMs) {
|
|
134
|
+
return new Date(slotEpochMs).toISOString().slice(0, 10);
|
|
135
|
+
}
|
|
136
|
+
function claimsRoot(opts) {
|
|
137
|
+
return opts.claimsDir ?? patchworkPath(CRON_CLAIM_DIRNAME);
|
|
138
|
+
}
|
|
139
|
+
function failClosed(opts) {
|
|
140
|
+
if (opts.required !== undefined)
|
|
141
|
+
return opts.required;
|
|
142
|
+
const v = process.env.PATCHWORK_CRON_CLAIM_REQUIRED;
|
|
143
|
+
return v === "1" || v?.toLowerCase() === "true";
|
|
144
|
+
}
|
|
145
|
+
/**
|
|
146
|
+
* Try to claim `(recipeName, slotEpochMs)` for this process.
|
|
147
|
+
*
|
|
148
|
+
* `openSync(path, "wx")` is one atomic syscall that is simultaneously the test
|
|
149
|
+
* and the set — no lock file, no shared append-only log, no read-modify-write,
|
|
150
|
+
* and no serialisation between unrelated recipes. `withFileLockSync` was
|
|
151
|
+
* considered and rejected for exactly those costs, plus a 30 s stale-lock TTL
|
|
152
|
+
* that would stall every recipe's scheduling if a process died holding it.
|
|
153
|
+
*
|
|
154
|
+
* The record body is written for a human to read afterwards and is NEVER read
|
|
155
|
+
* to make a decision — the kernel already made it at `open`.
|
|
156
|
+
*/
|
|
157
|
+
export function claimCronSlot(recipeName, slotEpochMs, opts = {}) {
|
|
158
|
+
// Floor again. The caller threads in an already-zeroed value, but a key that
|
|
159
|
+
// silently splits on stray milliseconds fails in the one way that does not
|
|
160
|
+
// reproduce, so it is worth one multiplication.
|
|
161
|
+
const slot = Math.floor(slotEpochMs / 1000) * 1000;
|
|
162
|
+
const dir = join(claimsRoot(opts), dayShard(slot));
|
|
163
|
+
const file = join(dir, `${cronClaimKey(recipeName, slot)}.json`);
|
|
164
|
+
let fd;
|
|
165
|
+
try {
|
|
166
|
+
mkdirSync(dir, { recursive: true, mode: 0o700 });
|
|
167
|
+
fd = openSync(file, "wx", 0o600);
|
|
168
|
+
}
|
|
169
|
+
catch (err) {
|
|
170
|
+
const code = err.code;
|
|
171
|
+
if (code === "EEXIST")
|
|
172
|
+
return { kind: "taken" };
|
|
173
|
+
const reason = code ?? (err instanceof Error ? err.message : "unknown");
|
|
174
|
+
return failClosed(opts)
|
|
175
|
+
? { kind: "refused", reason }
|
|
176
|
+
: { kind: "unavailable", reason };
|
|
177
|
+
}
|
|
178
|
+
try {
|
|
179
|
+
writeSync(fd, `${JSON.stringify({
|
|
180
|
+
v: 1,
|
|
181
|
+
recipeName,
|
|
182
|
+
slotEpochMs: slot,
|
|
183
|
+
slotIso: new Date(slot).toISOString(),
|
|
184
|
+
pid: process.pid,
|
|
185
|
+
claimedAt: Date.now(),
|
|
186
|
+
})}\n`);
|
|
187
|
+
}
|
|
188
|
+
catch {
|
|
189
|
+
// The claim is the file's EXISTENCE; the body is diagnostics. A failed body
|
|
190
|
+
// write must not surrender a slot we already own — that would hand it to a
|
|
191
|
+
// peer and fire twice.
|
|
192
|
+
}
|
|
193
|
+
finally {
|
|
194
|
+
try {
|
|
195
|
+
closeSync(fd);
|
|
196
|
+
}
|
|
197
|
+
catch {
|
|
198
|
+
/* already closed */
|
|
199
|
+
}
|
|
200
|
+
}
|
|
201
|
+
return { kind: "claimed" };
|
|
202
|
+
}
|
|
203
|
+
/**
|
|
204
|
+
* Delete claim day-directories older than the retention horizon.
|
|
205
|
+
*
|
|
206
|
+
* Best-effort and never throws: this is housekeeping, and a scheduler that
|
|
207
|
+
* cannot start because it could not tidy up would be a far worse bug than the
|
|
208
|
+
* disk usage it prevents.
|
|
209
|
+
*
|
|
210
|
+
* Returns the number of directories removed so a caller can log it — a sweep
|
|
211
|
+
* that silently does nothing looks identical to one that has nothing to do.
|
|
212
|
+
*/
|
|
213
|
+
export function sweepCronClaims(now = Date.now(), opts = {}) {
|
|
214
|
+
const root = claimsRoot(opts);
|
|
215
|
+
let entries;
|
|
216
|
+
try {
|
|
217
|
+
entries = readdirSync(root);
|
|
218
|
+
}
|
|
219
|
+
catch {
|
|
220
|
+
return 0; // nothing claimed yet, or unreadable — both are non-events here
|
|
221
|
+
}
|
|
222
|
+
let removed = 0;
|
|
223
|
+
for (const name of entries) {
|
|
224
|
+
// Parse the shard name rather than stat()-ing it: the directory's own mtime
|
|
225
|
+
// moves every time a claim lands in it, so a busy day would never age out.
|
|
226
|
+
const t = Date.parse(`${name}T00:00:00.000Z`);
|
|
227
|
+
if (!Number.isFinite(t))
|
|
228
|
+
continue; // not one of ours; leave it alone
|
|
229
|
+
if (now - t <= CLAIM_RETENTION_MS + 24 * 60 * 60 * 1000)
|
|
230
|
+
continue;
|
|
231
|
+
try {
|
|
232
|
+
rmSync(join(root, name), { recursive: true, force: true });
|
|
233
|
+
removed++;
|
|
234
|
+
}
|
|
235
|
+
catch {
|
|
236
|
+
/* best-effort */
|
|
237
|
+
}
|
|
238
|
+
}
|
|
239
|
+
return removed;
|
|
240
|
+
}
|
|
241
|
+
//# sourceMappingURL=cronClaim.js.map
|
|
@@ -0,0 +1 @@
|
|
|
1
|
+
{"version":3,"file":"cronClaim.js","sourceRoot":"","sources":["../../src/recipes/cronClaim.ts"],"names":[],"mappings":"AAAA;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;GAgGG;AAEH,OAAO,EAAE,UAAU,EAAE,MAAM,aAAa,CAAC;AACzC,OAAO,EACL,SAAS,EACT,SAAS,EACT,QAAQ,EACR,WAAW,EACX,MAAM,EACN,SAAS,GACV,MAAM,SAAS,CAAC;AACjB,OAAO,EAAE,IAAI,EAAE,MAAM,WAAW,CAAC;AACjC,OAAO,EAAE,aAAa,EAAE,MAAM,qBAAqB,CAAC;AAEpD,sEAAsE;AACtE,MAAM,CAAC,MAAM,kBAAkB,GAAG,aAAa,CAAC;AAEhD;;;;;;;GAOG;AACH,MAAM,CAAC,MAAM,kBAAkB,GAAG,EAAE,GAAG,EAAE,GAAG,EAAE,GAAG,IAAI,CAAC;AAoBtD;;;;;;;;;;;;GAYG;AACH,MAAM,UAAU,YAAY,CAAC,UAAkB,EAAE,WAAmB;IAClE,OAAO,UAAU,CAAC,QAAQ,CAAC;SACxB,MAAM,CAAC,IAAI,CAAC,SAAS,CAAC,CAAC,MAAM,EAAE,UAAU,EAAE,WAAW,CAAC,CAAC,CAAC;SACzD,MAAM,CAAC,KAAK,CAAC;SACb,KAAK,CAAC,CAAC,EAAE,EAAE,CAAC,CAAC;AAClB,CAAC;AAED,oFAAoF;AACpF,SAAS,QAAQ,CAAC,WAAmB;IACnC,OAAO,IAAI,IAAI,CAAC,WAAW,CAAC,CAAC,WAAW,EAAE,CAAC,KAAK,CAAC,CAAC,EAAE,EAAE,CAAC,CAAC;AAC1D,CAAC;AAaD,SAAS,UAAU,CAAC,IAAkB;IACpC,OAAO,IAAI,CAAC,SAAS,IAAI,aAAa,CAAC,kBAAkB,CAAC,CAAC;AAC7D,CAAC;AAED,SAAS,UAAU,CAAC,IAAkB;IACpC,IAAI,IAAI,CAAC,QAAQ,KAAK,SAAS;QAAE,OAAO,IAAI,CAAC,QAAQ,CAAC;IACtD,MAAM,CAAC,GAAG,OAAO,CAAC,GAAG,CAAC,6BAA6B,CAAC;IACpD,OAAO,CAAC,KAAK,GAAG,IAAI,CAAC,EAAE,WAAW,EAAE,KAAK,MAAM,CAAC;AAClD,CAAC;AAED;;;;;;;;;;;GAWG;AACH,MAAM,UAAU,aAAa,CAC3B,UAAkB,EAClB,WAAmB,EACnB,OAAqB,EAAE;IAEvB,6EAA6E;IAC7E,2EAA2E;IAC3E,gDAAgD;IAChD,MAAM,IAAI,GAAG,IAAI,CAAC,KAAK,CAAC,WAAW,GAAG,IAAI,CAAC,GAAG,IAAI,CAAC;IACnD,MAAM,GAAG,GAAG,IAAI,CAAC,UAAU,CAAC,IAAI,CAAC,EAAE,QAAQ,CAAC,IAAI,CAAC,CAAC,CAAC;IACnD,MAAM,IAAI,GAAG,IAAI,CAAC,GAAG,EAAE,GAAG,YAAY,CAAC,UAAU,EAAE,IAAI,CAAC,OAAO,CAAC,CAAC;IAEjE,IAAI,EAAU,CAAC;IACf,IAAI,CAAC;QACH,SAAS,CAAC,GAAG,EAAE,EAAE,SAAS,EAAE,IAAI,EAAE,IAAI,EAAE,KAAK,EAAE,CAAC,CAAC;QACjD,EAAE,GAAG,QAAQ,CAAC,IAAI,EAAE,IAAI,EAAE,KAAK,CAAC,CAAC;IACnC,CAAC;IAAC,OAAO,GAAG,EAAE,CAAC;QACb,MAAM,IAAI,GAAI,GAA6B,CAAC,IAAI,CAAC;QACjD,IAAI,IAAI,KAAK,QAAQ;YAAE,OAAO,EAAE,IAAI,EAAE,OAAO,EAAE,CAAC;QAChD,MAAM,MAAM,GAAG,IAAI,IAAI,CAAC,GAAG,YAAY,KAAK,CAAC,CAAC,CAAC,GAAG,CAAC,OAAO,CAAC,CAAC,CAAC,SAAS,CAAC,CAAC;QACxE,OAAO,UAAU,CAAC,IAAI,CAAC;YACrB,CAAC,CAAC,EAAE,IAAI,EAAE,SAAS,EAAE,MAAM,EAAE;YAC7B,CAAC,CAAC,EAAE,IAAI,EAAE,aAAa,EAAE,MAAM,EAAE,CAAC;IACtC,CAAC;IAED,IAAI,CAAC;QACH,SAAS,CACP,EAAE,EACF,GAAG,IAAI,CAAC,SAAS,CAAC;YAChB,CAAC,EAAE,CAAC;YACJ,UAAU;YACV,WAAW,EAAE,IAAI;YACjB,OAAO,EAAE,IAAI,IAAI,CAAC,IAAI,CAAC,CAAC,WAAW,EAAE;YACrC,GAAG,EAAE,OAAO,CAAC,GAAG;YAChB,SAAS,EAAE,IAAI,CAAC,GAAG,EAAE;SACtB,CAAC,IAAI,CACP,CAAC;IACJ,CAAC;IAAC,MAAM,CAAC;QACP,4EAA4E;QAC5E,2EAA2E;QAC3E,uBAAuB;IACzB,CAAC;YAAS,CAAC;QACT,IAAI,CAAC;YACH,SAAS,CAAC,EAAE,CAAC,CAAC;QAChB,CAAC;QAAC,MAAM,CAAC;YACP,oBAAoB;QACtB,CAAC;IACH,CAAC;IAED,OAAO,EAAE,IAAI,EAAE,SAAS,EAAE,CAAC;AAC7B,CAAC;AAED;;;;;;;;;GASG;AACH,MAAM,UAAU,eAAe,CAC7B,MAAc,IAAI,CAAC,GAAG,EAAE,EACxB,OAAqB,EAAE;IAEvB,MAAM,IAAI,GAAG,UAAU,CAAC,IAAI,CAAC,CAAC;IAC9B,IAAI,OAAiB,CAAC;IACtB,IAAI,CAAC;QACH,OAAO,GAAG,WAAW,CAAC,IAAI,CAAC,CAAC;IAC9B,CAAC;IAAC,MAAM,CAAC;QACP,OAAO,CAAC,CAAC,CAAC,gEAAgE;IAC5E,CAAC;IACD,IAAI,OAAO,GAAG,CAAC,CAAC;IAChB,KAAK,MAAM,IAAI,IAAI,OAAO,EAAE,CAAC;QAC3B,4EAA4E;QAC5E,2EAA2E;QAC3E,MAAM,CAAC,GAAG,IAAI,CAAC,KAAK,CAAC,GAAG,IAAI,gBAAgB,CAAC,CAAC;QAC9C,IAAI,CAAC,MAAM,CAAC,QAAQ,CAAC,CAAC,CAAC;YAAE,SAAS,CAAC,kCAAkC;QACrE,IAAI,GAAG,GAAG,CAAC,IAAI,kBAAkB,GAAG,EAAE,GAAG,EAAE,GAAG,EAAE,GAAG,IAAI;YAAE,SAAS;QAClE,IAAI,CAAC;YACH,MAAM,CAAC,IAAI,CAAC,IAAI,EAAE,IAAI,CAAC,EAAE,EAAE,SAAS,EAAE,IAAI,EAAE,KAAK,EAAE,IAAI,EAAE,CAAC,CAAC;YAC3D,OAAO,EAAE,CAAC;QACZ,CAAC;QAAC,MAAM,CAAC;YACP,iBAAiB;QACnB,CAAC;IACH,CAAC;IACD,OAAO,OAAO,CAAC;AACjB,CAAC"}
|
|
@@ -1,5 +1,6 @@
|
|
|
1
|
-
import
|
|
1
|
+
import { type ScheduledTask } from "node-cron";
|
|
2
2
|
import type { Logger } from "../logger.js";
|
|
3
|
+
import { type ClaimOptions } from "./cronClaim.js";
|
|
3
4
|
/**
|
|
4
5
|
* RecipeScheduler — runs cron-triggered recipes on a simple interval or
|
|
5
6
|
* standard 5-field cron expression.
|
|
@@ -22,7 +23,7 @@ export interface ScheduledRecipe {
|
|
|
22
23
|
intervalMs: number;
|
|
23
24
|
timer: ReturnType<typeof setInterval>;
|
|
24
25
|
/** Present only for cron5-kind recipes. */
|
|
25
|
-
cronJob?:
|
|
26
|
+
cronJob?: ScheduledTask;
|
|
26
27
|
}
|
|
27
28
|
export interface SchedulerOptions {
|
|
28
29
|
recipesDir: string;
|
|
@@ -47,7 +48,36 @@ export interface SchedulerOptions {
|
|
|
47
48
|
* this to avoid depending on the dev machine's config.
|
|
48
49
|
*/
|
|
49
50
|
timezone?: string;
|
|
51
|
+
/**
|
|
52
|
+
* Cron-claim store override (#1458). Tests point it at a temp root; the two
|
|
53
|
+
* real bridges share the default under PATCHWORK_HOME, which is exactly the
|
|
54
|
+
* scope of the shared recipe store that causes the double-fire.
|
|
55
|
+
*/
|
|
56
|
+
claim?: ClaimOptions;
|
|
50
57
|
}
|
|
58
|
+
/**
|
|
59
|
+
* The instant the cron matcher matched, as a second-aligned epoch value.
|
|
60
|
+
*
|
|
61
|
+
* node-cron hands the task callback a context whose `date` is the matched
|
|
62
|
+
* instant with milliseconds already zeroed, derived from the matcher rather
|
|
63
|
+
* than from the moment the callback happens to run. Two processes evaluating
|
|
64
|
+
* the same expression in the same timezone therefore produce the byte-identical
|
|
65
|
+
* value — which is the only reason a filesystem claim can dedupe them.
|
|
66
|
+
*
|
|
67
|
+
* `triggeredAt` on the same context is `new Date()` and must NEVER be used: it
|
|
68
|
+
* differs per process by exactly the amount that breaks the key.
|
|
69
|
+
*
|
|
70
|
+
* Returns `undefined` when there is no usable context, so an older or newer
|
|
71
|
+
* node-cron that does not pass one degrades to today's behaviour — no slot, no
|
|
72
|
+
* claim, both bridges fire — rather than throwing inside a timer callback.
|
|
73
|
+
*
|
|
74
|
+
* Exported for tests. The alternative is asserting it through a live cron tick,
|
|
75
|
+
* which means a real timer and a real second boundary in CI — and this repo has
|
|
76
|
+
* spent enough of its life on timing flakes.
|
|
77
|
+
*/
|
|
78
|
+
export declare function matchedSlotMs(ctx?: {
|
|
79
|
+
date?: Date;
|
|
80
|
+
}): number | undefined;
|
|
51
81
|
export declare class RecipeScheduler {
|
|
52
82
|
private readonly opts;
|
|
53
83
|
private scheduled;
|
|
@@ -66,8 +96,16 @@ export declare class RecipeScheduler {
|
|
|
66
96
|
* replay; this guard stops cross-attempt double-fire.
|
|
67
97
|
*
|
|
68
98
|
* Manual CLI runs do NOT go through this Set — they take their own
|
|
69
|
-
* path. The guard is scheduler-scoped (
|
|
70
|
-
* because that's the only
|
|
99
|
+
* path. The guard is scheduler-scoped (ONE PROCESS), and that is NOT
|
|
100
|
+
* enough: it used to say "which is enough because that's the only
|
|
101
|
+
* place a cron tick can originate", and #1458 disproved it live. The
|
|
102
|
+
* recipe store is global, so every running bridge schedules every
|
|
103
|
+
* enabled cron recipe and an in-process Set cannot see a sibling.
|
|
104
|
+
* N bridges fired N times.
|
|
105
|
+
*
|
|
106
|
+
* This Set still does its original job — a slow run overlapping the
|
|
107
|
+
* next tick WITHIN this process. The cross-process half is
|
|
108
|
+
* `claimCronSlot` (./cronClaim.ts), taken immediately before dispatch.
|
|
71
109
|
*/
|
|
72
110
|
private readonly inflight;
|
|
73
111
|
constructor(opts: SchedulerOptions);
|
|
@@ -75,8 +113,27 @@ export declare class RecipeScheduler {
|
|
|
75
113
|
stop(): void;
|
|
76
114
|
restart(): void;
|
|
77
115
|
list(): ReadonlyArray<Omit<ScheduledRecipe, "timer" | "cronJob">>;
|
|
78
|
-
/**
|
|
79
|
-
|
|
116
|
+
/**
|
|
117
|
+
* Test hook: dispatch a recipe immediately without waiting for the interval.
|
|
118
|
+
*
|
|
119
|
+
* Passes no slot BY DEFAULT, so it takes no cross-process claim. Deliberate:
|
|
120
|
+
* this hook has no cron match behind it, so there is no instant two processes
|
|
121
|
+
* could agree on, and inventing one from `Date.now()` would make repeated
|
|
122
|
+
* calls within the same second collide — which is precisely what the existing
|
|
123
|
+
* overlap tests do, and they must keep passing unchanged.
|
|
124
|
+
*
|
|
125
|
+
* A slot may be passed explicitly to drive the claim path without a live cron
|
|
126
|
+
* tick, i.e. without a real timer and a real second boundary in CI.
|
|
127
|
+
*/
|
|
128
|
+
fireForTest(name: string, slotEpochMs?: number): void;
|
|
129
|
+
/**
|
|
130
|
+
* @param slotEpochMs The instant the cron matcher matched, threaded from the
|
|
131
|
+
* cron callback. Its whole purpose is that two processes observing the same
|
|
132
|
+
* tick derive the SAME value — so it must never be re-read from the clock
|
|
133
|
+
* here, one event-loop hop later, where a second boundary would split it.
|
|
134
|
+
* Absent for `@every` intervals and the test hook: no slot, no claim, and
|
|
135
|
+
* behaviour identical to before #1458.
|
|
136
|
+
*/
|
|
80
137
|
private fire;
|
|
81
138
|
}
|
|
82
139
|
type ParsedSchedule = {
|