@memberjunction/scheduling-engine 5.38.0 → 5.39.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +93 -5
- package/dist/ScheduledJobEngine.d.ts +233 -29
- package/dist/ScheduledJobEngine.d.ts.map +1 -1
- package/dist/ScheduledJobEngine.js +611 -229
- package/dist/ScheduledJobEngine.js.map +1 -1
- package/dist/drivers/AgentRunSweepScheduledJobDriver.d.ts +35 -0
- package/dist/drivers/AgentRunSweepScheduledJobDriver.d.ts.map +1 -0
- package/dist/drivers/AgentRunSweepScheduledJobDriver.js +86 -0
- package/dist/drivers/AgentRunSweepScheduledJobDriver.js.map +1 -0
- package/dist/drivers/index.d.ts +1 -0
- package/dist/drivers/index.d.ts.map +1 -1
- package/dist/drivers/index.js +1 -0
- package/dist/drivers/index.js.map +1 -1
- package/package.json +14 -13
package/README.md
CHANGED
|
@@ -47,9 +47,10 @@ graph TD
|
|
|
47
47
|
|
|
48
48
|
- **Cron Evaluation**: Parses cron expressions to determine which jobs are due for execution
|
|
49
49
|
- **Plugin Architecture**: Each job type has a registered `BaseScheduledJob` driver class
|
|
50
|
-
- **Distributed Locking**: Token-based locking for multi-server environments with stale lock detection
|
|
50
|
+
- **Distributed Locking**: Token-based locking for multi-server environments with stale lock detection. Atomic compare-and-swap via dedicated stored procedures (`spAcquireScheduledJobLock`, `spReleaseScheduledJobLockIfTokenMatches`).
|
|
51
51
|
- **Concurrency Modes**: `Concurrent`, `Queue`, or `Skip` for overlapping executions
|
|
52
|
-
- **
|
|
52
|
+
- **Bounded Concurrency**: `MaxConcurrentJobs` (default 5) caps simultaneous in-flight jobs across the engine; configurable via `scheduledJobs.maxConcurrentJobs`
|
|
53
|
+
- **Polling Lifecycle**: Adaptive polling interval with `StartPolling()` / `StopPolling()`. As of v5.39 the poll loop dispatches jobs without awaiting completion — a single hung job no longer stalls the scheduler. See [Operational notes](#operational-notes-v539) below.
|
|
53
54
|
- **Notifications**: Configurable notifications on success, failure, or both
|
|
54
55
|
- **Statistics Tracking**: Automatic update of RunCount, SuccessCount, FailureCount, and timing metrics
|
|
55
56
|
- **Built-in Drivers**: `AgentScheduledJobDriver` for AI agents, `ActionScheduledJobDriver` for MJ Actions
|
|
@@ -70,11 +71,12 @@ import { SchedulingEngine } from '@memberjunction/scheduling-engine';
|
|
|
70
71
|
const engine = SchedulingEngine.Instance;
|
|
71
72
|
await engine.Config(false, contextUser);
|
|
72
73
|
|
|
73
|
-
// Start continuous polling
|
|
74
|
+
// Start continuous polling (async since v5.39 — must be awaited)
|
|
74
75
|
await engine.StartPolling(contextUser);
|
|
75
76
|
|
|
76
|
-
// Stop polling
|
|
77
|
-
|
|
77
|
+
// Stop polling on server shutdown. Graceful drain waits for currently-dispatched
|
|
78
|
+
// jobs to settle, bounded by maxWaitMs so a hung job can't hang shutdown.
|
|
79
|
+
await engine.StopPolling({ waitForInflight: true, maxWaitMs: 30_000 });
|
|
78
80
|
```
|
|
79
81
|
|
|
80
82
|
### Manual Job Execution
|
|
@@ -136,10 +138,96 @@ Abstract base class for all job drivers. Provides:
|
|
|
136
138
|
|
|
137
139
|
Utility for parsing and evaluating cron expressions using `cron-parser`.
|
|
138
140
|
|
|
141
|
+
### `RunImmediatelyIfNeverRun` Flag
|
|
142
|
+
|
|
143
|
+
Each `MJ: Scheduled Jobs` row now has a `RunImmediatelyIfNeverRun` boolean (v5.38). When `true` AND `LastRunAt IS NULL`, `SchedulingEngine.initializeNextRunTimes()` sets `NextRunAt = now()` instead of the next cron tick, so a freshly-seeded job runs on the next polling cycle rather than waiting up to a full cron interval (e.g., 24h for a daily job) for its first run.
|
|
144
|
+
|
|
145
|
+
Use this for seed metadata that should run as soon as it's installed (data backfills, initial syncs like Entity Vector Sync, etc.). The flag is a no-op once the job has run at least once — subsequent restarts follow the cron schedule normally.
|
|
146
|
+
|
|
139
147
|
### NotificationManager
|
|
140
148
|
|
|
141
149
|
Manages notification delivery based on job configuration (on success, failure, or both).
|
|
142
150
|
|
|
151
|
+
## Operational notes (v5.39+)
|
|
152
|
+
|
|
153
|
+
The v5.39 release decoupled the poll loop from job execution to fix a bug where a single hung `plugin.Execute()` would stall the entire scheduler until process restart (GH #2736). The design doc for the full rationale is at [`plans/scheduled-job-engine-decoupling.md`](../../../plans/scheduled-job-engine-decoupling.md). The summary below covers what operators and integrators need to know.
|
|
154
|
+
|
|
155
|
+
### Polling architecture
|
|
156
|
+
|
|
157
|
+
The poll loop calls `DispatchScheduledJobs(contextUser)`, which:
|
|
158
|
+
|
|
159
|
+
1. **Sweeps** stale in-flight entries (jobs whose DB lease has expired) — untracks them so their concurrency slot is freed, and fire-and-forget abandons their orphaned `Status='Running'` run records. The sweep is decoupled from `isJobDue` and from the cap; it runs unconditionally at the top of every poll.
|
|
160
|
+
2. **Dispatches** each due job under a cap (`MaxConcurrentJobs`, default 5), atomically acquiring its lock via `spAcquireScheduledJobLock`. Dispatched jobs run in the background; the poll loop returns immediately so the next poll fires on schedule.
|
|
161
|
+
|
|
162
|
+
Jobs that legitimately fail (throw) still recover via the existing try/catch in `executeJobWithLock`. The lock release uses `spReleaseScheduledJobLockIfTokenMatches` — atomic and token-checked, so a stale holder's late settlement after lease expiry can't clobber a fresh holder's lock.
|
|
163
|
+
|
|
164
|
+
### Cap and lease semantics
|
|
165
|
+
|
|
166
|
+
**`MaxConcurrentJobs` is a SOFT cap.** Under concurrent poll bodies overlapping, the cap may be transiently exceeded by a small amount bounded by the overlap count (typically 1–2). This is acceptable because:
|
|
167
|
+
|
|
168
|
+
- The cap exists to bound order-of-magnitude fan-out, not to enforce an exact integer count.
|
|
169
|
+
- The atomic lock sproc guarantees no same-job double-dispatch even under overshoot.
|
|
170
|
+
- Steady-state behavior holds at the configured cap.
|
|
171
|
+
|
|
172
|
+
If you need a hard ceiling, set `MaxConcurrentJobs` slightly below your true limit.
|
|
173
|
+
|
|
174
|
+
**Tuning guidance:**
|
|
175
|
+
|
|
176
|
+
- **Raise the cap** when scheduled work backs up despite plenty of headroom — many short, lightweight jobs.
|
|
177
|
+
- **Lower the cap** on memory / DB connection / rate-limit pressure, or when each job spawns heavy work (agent runs, LLM calls).
|
|
178
|
+
- **Shorten the lease** (`LeaseTimeoutMs` / `LeaseTimeoutMinutes`) to reduce the worst-case starvation window (see below), but never below the maximum runtime of any legitimate job — premature reclaim would re-dispatch healthy long jobs.
|
|
179
|
+
|
|
180
|
+
Configure via MJServer's `scheduledJobs` config block:
|
|
181
|
+
|
|
182
|
+
```javascript
|
|
183
|
+
// mj.config.cjs
|
|
184
|
+
module.exports = {
|
|
185
|
+
scheduledJobs: {
|
|
186
|
+
enabled: true,
|
|
187
|
+
systemUserEmail: 'system@example.com',
|
|
188
|
+
maxConcurrentJobs: 10, // default 5
|
|
189
|
+
defaultLockTimeout: 5 * 60_000, // default 600000 (10 min)
|
|
190
|
+
},
|
|
191
|
+
// ...
|
|
192
|
+
};
|
|
193
|
+
```
|
|
194
|
+
|
|
195
|
+
`MJServer/services/ScheduledJobsService` reads these on startup and applies them to the engine via `MaxConcurrentJobs` and `LeaseTimeoutMs` setters.
|
|
196
|
+
|
|
197
|
+
### Residual starvation window
|
|
198
|
+
|
|
199
|
+
This fix eliminates the unbounded scheduler stall but does NOT eliminate all starvation scenarios. If **`≥ MaxConcurrentJobs` jobs hang simultaneously** within a single lease window, other scheduled work is starved until at least one lease expires and the sweep reclaims its slot.
|
|
200
|
+
|
|
201
|
+
With defaults (`MaxConcurrentJobs=5`, lease=10 min), the worst-case starvation is **~10 minutes** — dramatically improved from the original "stalled until process restart," but not zero. Mitigations:
|
|
202
|
+
|
|
203
|
+
- Raise `MaxConcurrentJobs` so healthy jobs have headroom while zombies hold slots.
|
|
204
|
+
- Shorten `LeaseTimeoutMs` to accelerate reclaim (subject to the constraint above).
|
|
205
|
+
- Bound `plugin.Execute()` in your job plugins — the structural cure, tracked as a follow-up.
|
|
206
|
+
|
|
207
|
+
### Leaked promise behavior
|
|
208
|
+
|
|
209
|
+
When the sweep reclaims a hung job's lease:
|
|
210
|
+
|
|
211
|
+
1. The hung `executeJobWithLock` Promise stays in memory. JavaScript doesn't support Promise cancellation; the executing async function still holds the promise alive via its suspended `await`.
|
|
212
|
+
2. Each leaked promise holds: the plugin instance, execution context, run entity, and any closures the plugin captured (typically single-digit MB).
|
|
213
|
+
3. Bounded by the count of distinct hanging jobs. Process restart on deploy clears them.
|
|
214
|
+
4. **Operational signal**: each sweep emits a `[sweep] Untracking inflight job <name>` log line. A spike in these is a canary for a misbehaving plugin — wire to monitoring if hangs are a concern.
|
|
215
|
+
5. **Orphaned run records are NOT leaked.** The sweep marks them `Status='Failed'` with an explanatory message via `abandonOrphanedRunRecords`. The DB litter pattern (phantom `Running` rows accumulating over months) is prevented going forward.
|
|
216
|
+
|
|
217
|
+
`Promise.race([plugin.Execute(), timeout])` does NOT solve the leak — `Promise.race` returns when the timeout wins, but the underlying `plugin.Execute()` promise is still alive. Same heap leak, different surface. Worker-thread isolation would let us terminate a hung plugin's execution context, but that's a multi-week design effort tracked separately.
|
|
218
|
+
|
|
219
|
+
### Job-list refresh
|
|
220
|
+
|
|
221
|
+
The engine loads the job list ONCE at `StartPolling` time. Runtime changes to the `MJ: Scheduled Jobs` table (adds, deletions, cron edits) are NOT picked up automatically by the engine. Callers must explicitly invoke `await engine.OnJobChanged(contextUser)` after making such changes. This is pre-existing behavior — not changed by v5.39 — but worth calling out because the per-poll `Config(false)` no-op was moved out of the poll path, making the explicit `OnJobChanged` trigger the only refresh mechanism.
|
|
222
|
+
|
|
223
|
+
### Graceful shutdown
|
|
224
|
+
|
|
225
|
+
`StopPolling({ waitForInflight, maxWaitMs })` lets shutdown wait for dispatched jobs to settle:
|
|
226
|
+
|
|
227
|
+
- `waitForInflight: true` — awaits `Promise.allSettled` on all tracked promises.
|
|
228
|
+
- `maxWaitMs` — bounds the wait so a zombie can't block shutdown indefinitely (recommended in production; `ScheduledJobsService.Stop()` uses 30s by default).
|
|
229
|
+
- Default (no opts) — preserves the prior fire-and-forget behavior; tasks may continue running in the background while the process exits.
|
|
230
|
+
|
|
143
231
|
## Dependencies
|
|
144
232
|
|
|
145
233
|
| Package | Purpose |
|
|
@@ -44,6 +44,49 @@ export declare class SchedulingEngine extends BaseSingleton<SchedulingEngine> {
|
|
|
44
44
|
private highFrequencyWarnedJobIds;
|
|
45
45
|
/** Threshold (ms) below which a job's run frequency triggers a console warning. */
|
|
46
46
|
private static readonly HIGH_FREQUENCY_WARNING_THRESHOLD_MS;
|
|
47
|
+
/**
|
|
48
|
+
* Maximum concurrent scheduled jobs on this engine instance. Default 5.
|
|
49
|
+
* Configurable via MJServer's `scheduledJobs.maxConcurrentJobs` config.
|
|
50
|
+
*
|
|
51
|
+
* SOFT CAP: under overlapping poll bodies the cap may be transiently
|
|
52
|
+
* exceeded by a small amount bounded by overlap count. See README
|
|
53
|
+
* "Cap and lease semantics" for tuning guidance.
|
|
54
|
+
*/
|
|
55
|
+
private _maxConcurrentJobs;
|
|
56
|
+
/**
|
|
57
|
+
* Lock lease duration. Default 10 minutes. Configurable via
|
|
58
|
+
* mj.config.cjs `scheduling.leaseTimeoutMinutes`.
|
|
59
|
+
*
|
|
60
|
+
* Public API uses minutes (operational readability); internal computations
|
|
61
|
+
* use _leaseTimeoutMs for precision and testability. Tests inject sub-second
|
|
62
|
+
* leases via `_setLeaseTimeoutMsForTest`.
|
|
63
|
+
*
|
|
64
|
+
* Constraint: must be > maximum expected runtime of any job. Setting too
|
|
65
|
+
* low causes healthy long jobs to be reclaimed and re-dispatched.
|
|
66
|
+
*/
|
|
67
|
+
private _leaseTimeoutMs;
|
|
68
|
+
/**
|
|
69
|
+
* Promises for jobs currently dispatched but not yet settled. Keyed by job ID.
|
|
70
|
+
*
|
|
71
|
+
* Three purposes:
|
|
72
|
+
* 1. Bounded concurrency — DispatchScheduledJobs checks size vs MaxConcurrentJobs.
|
|
73
|
+
* 2. Sweep untracking — sweepStaleInflightJobs deletes by ID for jobs whose
|
|
74
|
+
* lease has expired, freeing the cap slot even though the JS promise leaks
|
|
75
|
+
* (see README "Leaked promise behavior").
|
|
76
|
+
* 3. Graceful shutdown — StopPolling can await all in-flight via .values().
|
|
77
|
+
*
|
|
78
|
+
* Self-cleans via identity-checked .finally() on each dispatched promise
|
|
79
|
+
* (no-op if a sweep + re-dispatch already replaced the entry).
|
|
80
|
+
*
|
|
81
|
+
* NOT used for double-dispatch prevention — that's the atomic lock sproc's job.
|
|
82
|
+
*/
|
|
83
|
+
private inflightJobPromises;
|
|
84
|
+
/**
|
|
85
|
+
* When false, DispatchScheduledJobs becomes a no-op. Set false in StopPolling
|
|
86
|
+
* BEFORE snapshotting inflightJobPromises for shutdown drain, so no new
|
|
87
|
+
* entries sneak in during the shutdown window.
|
|
88
|
+
*/
|
|
89
|
+
private acceptingDispatches;
|
|
47
90
|
/** Gets all scheduled job types. */
|
|
48
91
|
get ScheduledJobTypes(): MJScheduledJobTypeEntity[];
|
|
49
92
|
/** Gets scheduled jobs (active only by default). */
|
|
@@ -52,6 +95,27 @@ export declare class SchedulingEngine extends BaseSingleton<SchedulingEngine> {
|
|
|
52
95
|
get ScheduledJobRuns(): MJScheduledJobRunEntity[];
|
|
53
96
|
/** Gets the current active polling interval in milliseconds. */
|
|
54
97
|
get ActivePollingInterval(): number | null;
|
|
98
|
+
/**
|
|
99
|
+
* Maximum concurrent scheduled jobs on this engine instance. Default 5.
|
|
100
|
+
* Configurable via MJServer's `scheduledJobs.maxConcurrentJobs` config.
|
|
101
|
+
*/
|
|
102
|
+
get MaxConcurrentJobs(): number;
|
|
103
|
+
set MaxConcurrentJobs(value: number);
|
|
104
|
+
/**
|
|
105
|
+
* Lock lease duration in milliseconds. Default 600000 (10 minutes).
|
|
106
|
+
* Production callers should use this setter — matches the ms unit of
|
|
107
|
+
* MJServer's `scheduledJobs.defaultLockTimeout` config.
|
|
108
|
+
*/
|
|
109
|
+
get LeaseTimeoutMs(): number;
|
|
110
|
+
set LeaseTimeoutMs(value: number);
|
|
111
|
+
/**
|
|
112
|
+
* Convenience accessor — lease duration as integer minutes. Production
|
|
113
|
+
* code may use either this or `LeaseTimeoutMs`. The setter validates
|
|
114
|
+
* positive integer minutes (no fractional minutes via this path; use
|
|
115
|
+
* `LeaseTimeoutMs` for sub-minute precision, including tests).
|
|
116
|
+
*/
|
|
117
|
+
get LeaseTimeoutMinutes(): number;
|
|
118
|
+
set LeaseTimeoutMinutes(value: number);
|
|
55
119
|
/** Find a job type by name. */
|
|
56
120
|
GetJobTypeByName(name: string): MJScheduledJobTypeEntity | undefined;
|
|
57
121
|
/** Find a job type by driver class. */
|
|
@@ -68,16 +132,37 @@ export declare class SchedulingEngine extends BaseSingleton<SchedulingEngine> {
|
|
|
68
132
|
*/
|
|
69
133
|
Config(forceRefresh?: boolean, contextUser?: UserInfo, provider?: IMetadataProvider, includeRuns?: boolean, includeAllJobs?: boolean): Promise<boolean>;
|
|
70
134
|
/**
|
|
71
|
-
* Start continuous polling for scheduled jobs
|
|
72
|
-
*
|
|
135
|
+
* Start continuous polling for scheduled jobs.
|
|
136
|
+
*
|
|
137
|
+
* Async (changed in v5.39) because upfront work — Config, initial-NextRunAt
|
|
138
|
+
* seeding, stale-lock cleanup, permission probe — runs ONCE before the
|
|
139
|
+
* first poll fires. Subsequent polls assume that work is complete.
|
|
140
|
+
*
|
|
141
|
+
* The poll callback re-arms its timer FIRST, before any awaited work, so
|
|
142
|
+
* that any hang downstream (Config, DispatchScheduledJobs, etc.) cannot
|
|
143
|
+
* prevent the next poll from firing on schedule. This is the load-bearing
|
|
144
|
+
* invariant of the decoupling fix (see plans/scheduled-job-engine-decoupling.md).
|
|
73
145
|
*
|
|
74
146
|
* @param contextUser - User context for execution
|
|
75
147
|
*/
|
|
76
|
-
StartPolling(contextUser: UserInfo): void
|
|
148
|
+
StartPolling(contextUser: UserInfo): Promise<void>;
|
|
77
149
|
/**
|
|
78
|
-
* Stop continuous polling
|
|
150
|
+
* Stop continuous polling.
|
|
151
|
+
*
|
|
152
|
+
* Async (changed in v5.39). With opts.waitForInflight=true, awaits all
|
|
153
|
+
* currently-dispatched jobs to settle before returning. With opts.maxWaitMs,
|
|
154
|
+
* bounds that wait so a zombie can't make shutdown hang indefinitely.
|
|
155
|
+
*
|
|
156
|
+
* Order matters: sets acceptingDispatches=false FIRST so no new entries
|
|
157
|
+
* can be added to inflightJobPromises during the snapshot for allSettled.
|
|
158
|
+
*
|
|
159
|
+
* @param opts.waitForInflight - Await dispatched jobs before returning
|
|
160
|
+
* @param opts.maxWaitMs - Bound the wait (only meaningful with waitForInflight)
|
|
79
161
|
*/
|
|
80
|
-
StopPolling(
|
|
162
|
+
StopPolling(opts?: {
|
|
163
|
+
waitForInflight?: boolean;
|
|
164
|
+
maxWaitMs?: number;
|
|
165
|
+
}): Promise<void>;
|
|
81
166
|
/**
|
|
82
167
|
* Check if polling is currently active
|
|
83
168
|
*/
|
|
@@ -125,6 +210,43 @@ export declare class SchedulingEngine extends BaseSingleton<SchedulingEngine> {
|
|
|
125
210
|
* @returns The scheduled job run record
|
|
126
211
|
*/
|
|
127
212
|
ExecuteScheduledJob(jobId: string, contextUser: UserInfo): Promise<MJScheduledJobRunEntity>;
|
|
213
|
+
/**
|
|
214
|
+
* Dispatch all currently-due scheduled jobs WITHOUT awaiting their completion.
|
|
215
|
+
*
|
|
216
|
+
* This is the polling-path entry point introduced in v5.39 as part of the
|
|
217
|
+
* scheduler decoupling fix (GH #2736). The poll loop calls this and re-arms
|
|
218
|
+
* its timer based on the synchronous-portion return; jobs run in the background.
|
|
219
|
+
*
|
|
220
|
+
* Two phases:
|
|
221
|
+
*
|
|
222
|
+
* PHASE 1 — Stale-inflight sweep (decoupled from isJobDue AND from the cap):
|
|
223
|
+
* Walks inflightJobPromises looking for jobs whose DB lease has expired.
|
|
224
|
+
* Untracks each, frees its cap slot, and fire-and-forget marks any
|
|
225
|
+
* orphaned `Status='Running'` run records as abandoned. Runs first so
|
|
226
|
+
* it can free slots BEFORE the cap check throttles dispatch.
|
|
227
|
+
*
|
|
228
|
+
* PHASE 2 — Cap-bounded dispatch loop:
|
|
229
|
+
* For each due job, atomically acquire its lock via spAcquireScheduledJobLock.
|
|
230
|
+
* Only jobs whose lock was acquired count against MaxConcurrentJobs.
|
|
231
|
+
* Lock-failed jobs are reported via `lockedOut` counter.
|
|
232
|
+
* If at-cap, remaining due jobs counted via `skippedAtCapacity` and
|
|
233
|
+
* picked up by subsequent polls as slots free (no in-memory queueing).
|
|
234
|
+
*
|
|
235
|
+
* Same-instance double-dispatch is structurally prevented by the atomic
|
|
236
|
+
* lock sproc — its WHERE clause filters held-and-not-stale locks, so any
|
|
237
|
+
* second attempt against the same job ID returns Acquired=0.
|
|
238
|
+
*
|
|
239
|
+
* In-flight dispatched promises are tracked in `inflightJobPromises` so
|
|
240
|
+
* `StopPolling({ waitForInflight: true })` can perform graceful shutdown.
|
|
241
|
+
*
|
|
242
|
+
* @returns Counters for observability.
|
|
243
|
+
*/
|
|
244
|
+
DispatchScheduledJobs(contextUser: UserInfo, evalTime?: Date): Promise<{
|
|
245
|
+
swept: number;
|
|
246
|
+
dispatched: number;
|
|
247
|
+
lockedOut: number;
|
|
248
|
+
skippedAtCapacity: number;
|
|
249
|
+
}>;
|
|
128
250
|
/**
|
|
129
251
|
* Determine if a job is currently due for execution
|
|
130
252
|
*
|
|
@@ -137,12 +259,28 @@ export declare class SchedulingEngine extends BaseSingleton<SchedulingEngine> {
|
|
|
137
259
|
/**
|
|
138
260
|
* Execute a single scheduled job
|
|
139
261
|
*
|
|
140
|
-
*
|
|
141
|
-
*
|
|
142
|
-
*
|
|
262
|
+
* Execute a single scheduled job WITH a pre-acquired lock token.
|
|
263
|
+
*
|
|
264
|
+
* Caller (DispatchScheduledJobs / ExecuteScheduledJob / ExecuteScheduledJobs)
|
|
265
|
+
* is responsible for acquiring the lock atomically via tryAcquireLock and
|
|
266
|
+
* passing the resulting token. This method owns the lock's lifecycle from
|
|
267
|
+
* this point forward: every exit path (success, failure, exception)
|
|
268
|
+
* releases the lock via releaseLockIfTokenMatches.
|
|
269
|
+
*
|
|
270
|
+
* If lockToken is null, the job is running in `ConcurrencyMode='Concurrent'`
|
|
271
|
+
* (no lock acquired) — finally simply skips the release.
|
|
272
|
+
*
|
|
273
|
+
* @param job - The job entity. READ-ONLY from this method's perspective —
|
|
274
|
+
* do not mutate or call Save on it. The shared entity in
|
|
275
|
+
* this.ScheduledJobs must not be touched here.
|
|
276
|
+
* @param lockToken - The token returned by tryAcquireLock, or null for
|
|
277
|
+
* Concurrent mode where no lock was acquired.
|
|
278
|
+
* @param contextUser - User context for execution.
|
|
279
|
+
* @returns The created run record (Completed or Failed), or null if a
|
|
280
|
+
* synchronous setup error prevented run creation.
|
|
143
281
|
* @private
|
|
144
282
|
*/
|
|
145
|
-
private
|
|
283
|
+
private executeJobWithLock;
|
|
146
284
|
/**
|
|
147
285
|
* Create a new ScheduledJobRun record
|
|
148
286
|
* @private
|
|
@@ -162,32 +300,84 @@ export declare class SchedulingEngine extends BaseSingleton<SchedulingEngine> {
|
|
|
162
300
|
* Try to acquire a lock for job execution
|
|
163
301
|
* @private
|
|
164
302
|
*/
|
|
303
|
+
/**
|
|
304
|
+
* Atomically acquire a lock on a job via spAcquireScheduledJobLock.
|
|
305
|
+
* The sproc's WHERE clause handles both the free-lock and stale-lease cases
|
|
306
|
+
* in a single statement — no TOCTOU window between check and write.
|
|
307
|
+
*
|
|
308
|
+
* Operates only on lock columns; never mutates the shared entity in
|
|
309
|
+
* this.ScheduledJobs. Caller passes jobId (string), not the entity object.
|
|
310
|
+
*
|
|
311
|
+
* @returns { acquired: true, token } on success; { acquired: false } otherwise.
|
|
312
|
+
* @private
|
|
313
|
+
*/
|
|
165
314
|
private tryAcquireLock;
|
|
166
315
|
/**
|
|
167
|
-
*
|
|
168
|
-
*
|
|
169
|
-
*
|
|
316
|
+
* Atomically release a lock IF AND ONLY IF the current DB token matches
|
|
317
|
+
* expectedToken. Prevents the lost-mutex hazard under lease-expiry races:
|
|
318
|
+
* if a stale holder's execution eventually settles after the lease was
|
|
319
|
+
* reclaimed by a fresh holder, this no-ops (token mismatch).
|
|
320
|
+
*
|
|
321
|
+
* Idempotent — safe to call on an already-released lock (returns false).
|
|
322
|
+
*
|
|
323
|
+
* @returns true if released, false if token mismatch / already released
|
|
170
324
|
* @private
|
|
171
325
|
*/
|
|
172
|
-
private
|
|
326
|
+
private releaseLockIfTokenMatches;
|
|
173
327
|
/**
|
|
174
|
-
*
|
|
175
|
-
*
|
|
176
|
-
*
|
|
328
|
+
* Pre-flight: verify EXECUTE permission on lock sprocs. Fails LOUDLY at boot
|
|
329
|
+
* if the engine's DB principal lacks grants — much better than a silent
|
|
330
|
+
* runtime failure the next time a job tries to dispatch.
|
|
331
|
+
*
|
|
332
|
+
* Wrapped in try/catch: probe failure (e.g., non-SQL-Server provider where
|
|
333
|
+
* `sys.fn_my_permissions` doesn't exist) must NOT crash boot. We log and
|
|
334
|
+
* continue; any actual permission issue will surface at first sproc call.
|
|
335
|
+
*
|
|
177
336
|
* @private
|
|
178
337
|
*/
|
|
179
|
-
private
|
|
338
|
+
private probeLockSprocPermissions;
|
|
180
339
|
/**
|
|
181
|
-
*
|
|
340
|
+
* Sweep stale inflight jobs. Runs unconditionally at top of every poll.
|
|
341
|
+
*
|
|
342
|
+
* SINGLE BATCH QUERY (not N round-trips). Returns only jobs whose lease has
|
|
343
|
+
* expired OR whose lock has already been cleared. In steady-state (no zombies)
|
|
344
|
+
* the query matches zero rows and the sweep is essentially free.
|
|
345
|
+
*
|
|
346
|
+
* For each stale entry:
|
|
347
|
+
* - Untrack the leaked promise from inflightJobPromises (frees cap slot).
|
|
348
|
+
* - FIRE-AND-FORGET abandon any orphaned `Status='Running'` run records.
|
|
349
|
+
* NOT awaited because cleanup must not delay dispatch under a fleet-wide
|
|
350
|
+
* hang event where the sweep finds many zombies at once.
|
|
351
|
+
*
|
|
352
|
+
* Decoupled from:
|
|
353
|
+
* - isJobDue — irrelevant; we care about lease state, not cron.
|
|
354
|
+
* - MaxConcurrentJobs — the sweep IS what frees the cap when saturated by hangs.
|
|
355
|
+
*
|
|
356
|
+
* See plans/scheduled-job-engine-decoupling.md for the full rationale.
|
|
357
|
+
*
|
|
358
|
+
* @returns count of inflight entries swept
|
|
182
359
|
* @private
|
|
183
360
|
*/
|
|
184
|
-
private
|
|
361
|
+
private sweepStaleInflightJobs;
|
|
185
362
|
/**
|
|
186
|
-
*
|
|
187
|
-
*
|
|
363
|
+
* Mark any Running run records for the given job as Failed/abandoned.
|
|
364
|
+
*
|
|
365
|
+
* IMPORTANT: the `Status='Running'` filter is LOAD-BEARING — not just for
|
|
366
|
+
* finding zombies. It also protects against a sweep/release race:
|
|
367
|
+
*
|
|
368
|
+
* - Job completes normally.
|
|
369
|
+
* - executeJobWithLock's finally calls releaseLockIfTokenMatches (clears LockToken).
|
|
370
|
+
* - BEFORE that completes, a poll's sweep query sees LockToken IS NULL
|
|
371
|
+
* and classifies the just-completed job as a zombie.
|
|
372
|
+
* - But its run record is already Status='Completed' (set inside the try block,
|
|
373
|
+
* before the finally), so THIS FILTER excludes it from abandonment.
|
|
374
|
+
*
|
|
375
|
+
* Removing or relaxing this filter would corrupt completed run records.
|
|
376
|
+
* If "optimizing" this method, preserve the Status='Running' filter.
|
|
377
|
+
*
|
|
188
378
|
* @private
|
|
189
379
|
*/
|
|
190
|
-
private
|
|
380
|
+
private abandonOrphanedRunRecords;
|
|
191
381
|
/**
|
|
192
382
|
* Create a queued job run for later execution
|
|
193
383
|
* @private
|
|
@@ -199,17 +389,31 @@ export declare class SchedulingEngine extends BaseSingleton<SchedulingEngine> {
|
|
|
199
389
|
*/
|
|
200
390
|
private getInstanceIdentifier;
|
|
201
391
|
/**
|
|
202
|
-
*
|
|
203
|
-
*
|
|
204
|
-
|
|
205
|
-
|
|
206
|
-
|
|
207
|
-
*
|
|
392
|
+
* Initialize NextRunAt for jobs that don't have it set.
|
|
393
|
+
*
|
|
394
|
+
* If a job has `RunImmediatelyIfNeverRun = true` AND has never run
|
|
395
|
+
* (`LastRunAt IS NULL`), `NextRunAt` is set to `now()` so the job
|
|
396
|
+
* executes on the next polling cycle instead of waiting for the next
|
|
397
|
+
* cron tick. Useful for freshly-seeded jobs that should not wait up
|
|
398
|
+
* to a full cron interval (e.g. 24h for a daily job) for their first run.
|
|
399
|
+
*
|
|
208
400
|
* @private
|
|
209
401
|
*/
|
|
210
402
|
private initializeNextRunTimes;
|
|
211
403
|
/**
|
|
212
|
-
* Clean up stale locks on startup
|
|
404
|
+
* Clean up stale locks on startup using atomic sprocs.
|
|
405
|
+
*
|
|
406
|
+
* For each job whose DB shows a stale lock (ExpectedCompletionAt < now OR
|
|
407
|
+
* ExpectedCompletionAt IS NULL while LockToken IS NOT NULL):
|
|
408
|
+
* 1. Atomically acquire the stale lock with a fresh token (sproc's WHERE
|
|
409
|
+
* handles the stale-detection in a single statement).
|
|
410
|
+
* 2. Immediately release it with that same token.
|
|
411
|
+
*
|
|
412
|
+
* Net effect: stale lock cleared atomically with zero TOCTOU window. Uses
|
|
413
|
+
* the new sproc-backed pattern instead of load-compare-save on shared
|
|
414
|
+
* this.ScheduledJobs entities (see plans/scheduled-job-engine-decoupling.md
|
|
415
|
+
* for why the old pattern was unsafe once polling became concurrent).
|
|
416
|
+
*
|
|
213
417
|
* @private
|
|
214
418
|
*/
|
|
215
419
|
private cleanupStaleLocks;
|
|
@@ -1 +1 @@
|
|
|
1
|
-
{"version":3,"file":"ScheduledJobEngine.d.ts","sourceRoot":"","sources":["../src/ScheduledJobEngine.ts"],"names":[],"mappings":"AAAA;;;GAGG;
|
|
1
|
+
{"version":3,"file":"ScheduledJobEngine.d.ts","sourceRoot":"","sources":["../src/ScheduledJobEngine.ts"],"names":[],"mappings":"AAAA;;;GAGG;AAIH,OAAO,EACH,QAAQ,EAER,iBAAiB,EAOpB,MAAM,sBAAsB,CAAC;AAC9B,OAAO,EAAE,oBAAoB,EAAE,uBAAuB,EAAE,wBAAwB,EAAE,MAAM,+BAA+B,CAAC;AACxH,OAAO,EAAE,aAAa,EAAwB,MAAM,wBAAwB,CAAC;AAE7E,OAAO,EAAE,oBAAoB,EAAE,MAAM,wCAAwC,CAAC;AAK9E;;;;;;;;;;;;;;;;;;;;;GAqBG;AACH,qBAAa,gBAAiB,SAAQ,aAAa,CAAC,gBAAgB,CAAC;IACjE;;OAEG;IACH,WAAkB,QAAQ,IAAI,gBAAgB,CAE7C;IAED;;OAEG;IACH,SAAS,KAAK,IAAI,IAAI,oBAAoB,CAEzC;IAED,OAAO,CAAC,YAAY,CAAC,CAAiB;IACtC,OAAO,CAAC,SAAS,CAAkB;IACnC,OAAO,CAAC,cAAc,CAAkB;IAExC,4EAA4E;IAC5E,OAAO,CAAC,yBAAyB,CAA0B;IAE3D,mFAAmF;IACnF,OAAO,CAAC,MAAM,CAAC,QAAQ,CAAC,mCAAmC,CAAiB;IAM5E;;;;;;;OAOG;IACH,OAAO,CAAC,kBAAkB,CAAa;IAEvC;;;;;;;;;;OAUG;IACH,OAAO,CAAC,eAAe,CAA0B;IAEjD;;;;;;;;;;;;;;OAcG;IACH,OAAO,CAAC,mBAAmB,CAAmE;IAE9F;;;;OAIG;IACH,OAAO,CAAC,mBAAmB,CAAiB;IAM5C,oCAAoC;IACpC,IAAW,iBAAiB,IAAI,wBAAwB,EAAE,CAEzD;IAED,oDAAoD;IACpD,IAAW,aAAa,IAAI,oBAAoB,EAAE,CAEjD;IAED,sCAAsC;IACtC,IAAW,gBAAgB,IAAI,uBAAuB,EAAE,CAEvD;IAED,gEAAgE;IAChE,IAAW,qBAAqB,IAAI,MAAM,GAAG,IAAI,CAEhD;IAED;;;OAGG;IACH,IAAW,iBAAiB,IAAI,MAAM,CAErC;IACD,IAAW,iBAAiB,CAAC,KAAK,EAAE,MAAM,EAOzC;IAED;;;;OAIG;IACH,IAAW,cAAc,IAAI,MAAM,CAElC;IACD,IAAW,cAAc,CAAC,KAAK,EAAE,MAAM,EAOtC;IAED;;;;;OAKG;IACH,IAAW,mBAAmB,IAAI,MAAM,CAEvC;IACD,IAAW,mBAAmB,CAAC,KAAK,EAAE,MAAM,EAK3C;IAED,+BAA+B;IACxB,gBAAgB,CAAC,IAAI,EAAE,MAAM,GAAG,wBAAwB,GAAG,SAAS;IAI3E,uCAAuC;IAChC,uBAAuB,CAAC,WAAW,EAAE,MAAM,GAAG,wBAAwB,GAAG,SAAS;IAIzF,uCAAuC;IAChC,aAAa,CAAC,SAAS,EAAE,MAAM,GAAG,oBAAoB,EAAE;IAI/D,mCAAmC;IAC5B,aAAa,CAAC,KAAK,EAAE,MAAM,GAAG,uBAAuB,EAAE;IAI9D,wDAAwD;IACjD,qBAAqB,IAAI,IAAI;IAQpC;;;OAGG;IACU,MAAM,CACf,YAAY,CAAC,EAAE,OAAO,EACtB,WAAW,CAAC,EAAE,QAAQ,EACtB,QAAQ,CAAC,EAAE,iBAAiB,EAC5B,WAAW,GAAE,OAAe,EAC5B,cAAc,GAAE,OAAe,GAChC,OAAO,CAAC,OAAO,CAAC;IAQnB;;;;;;;;;;;;;OAaG;IACU,YAAY,CAAC,WAAW,EAAE,QAAQ,GAAG,OAAO,CAAC,IAAI,CAAC;IA8D/D;;;;;;;;;;;;OAYG;IACU,WAAW,CAAC,IAAI,CAAC,EAAE;QAAE,eAAe,CAAC,EAAE,OAAO,CAAC;QAAC,SAAS,CAAC,EAAE,MAAM,CAAA;KAAE,GAAG,OAAO,CAAC,IAAI,CAAC;IA4BjG;;OAEG;IACH,IAAW,SAAS,IAAI,OAAO,CAE9B;IAED;;;;;;;OAOG;IACU,YAAY,CAAC,WAAW,EAAE,QAAQ,GAAG,OAAO,CAAC,IAAI,CAAC;IA0B/D;;;;;;;OAOG;IACH,OAAO,CAAC,0BAA0B;IA4DlC;;;;OAIG;IACH,OAAO,CAAC,cAAc;IAkBtB;;;;;;;;;OASG;IACU,oBAAoB,CAC7B,WAAW,EAAE,QAAQ,EACrB,QAAQ,GAAE,IAAiB,GAC5B,OAAO,CAAC,uBAAuB,EAAE,CAAC;IA8CrC;;;;;;OAMG;IACU,mBAAmB,CAC5B,KAAK,EAAE,MAAM,EACb,WAAW,EAAE,QAAQ,GACtB,OAAO,CAAC,uBAAuB,CAAC;IAmBnC;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;OA8BG;IACU,qBAAqB,CAC9B,WAAW,EAAE,QAAQ,EACrB,QAAQ,GAAE,IAAiB,GAC5B,OAAO,CAAC;QAAE,KAAK,EAAE,MAAM,CAAC;QAAC,UAAU,EAAE,MAAM,CAAC;QAAC,SAAS,EAAE,MAAM,CAAC;QAAC,iBAAiB,EAAE,MAAM,CAAA;KAAE,CAAC;IA8D/F;;;;;;;OAOG;IACH,OAAO,CAAC,QAAQ;IAkBhB;;;;;;;;;;;;;;;;;;;;;;;OAuBG;YACW,kBAAkB;IAmGhC;;;OAGG;YACW,YAAY;IAoB1B;;;OAGG;YACW,mBAAmB;IAmEjC;;;OAGG;YACW,yBAAyB;IAmCvC;;;OAGG;IACH;;;;;;;;;;OAUG;YACW,cAAc;IA0B5B;;;;;;;;;;OAUG;YACW,yBAAyB;IAsBvC;;;;;;;;;;OAUG;YACW,yBAAyB;IAiCvC;;;;;;;;;;;;;;;;;;;;;OAqBG;YACW,sBAAsB;IAyCpC;;;;;;;;;;;;;;;;;OAiBG;YACW,yBAAyB;IAoCvC;;;OAGG;YACW,kBAAkB;IAsBhC;;;OAGG;IACH,OAAO,CAAC,qBAAqB;IAI7B;;;;;;;;;;OAUG;YACW,sBAAsB;IAyBpC;;;;;;;;;;;;;;;OAeG;YACW,iBAAiB;IA+C/B,OAAO,CAAC,GAAG;IAQX,OAAO,CAAC,QAAQ;CAGnB"}
|