@stonyx/cron 0.2.1-alpha.59 → 0.2.1-alpha.60

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -47,6 +47,51 @@ When a job is executed, its next trigger time is updated, and it is re-inserted
47
47
 
48
48
  > `MinHeap` is also exported as a public subpath (`@stonyx/cron/min-heap`) and can be imported directly for advanced usage.
49
49
 
50
+ ## CronService
51
+
52
+ The default export above is `Cron`: a fire-and-forget interval registry. `@stonyx/cron/service` is a separate, heavier class for jobs that need CRUD, persistence, a run log and error backoff. It is not a drop-in replacement and the two do not share a scheduler.
53
+
54
+ The two classes agree on the guarantee — the same job is never run concurrently with itself, and different jobs may overlap — but not on the mechanism or on what you can observe. `Cron` invokes callbacks fire-and-forget and reports a skipped run only through the `config.cron`-gated log. `CronService` **awaits** `onJobDue`, its return value shapes `status`/`error`/`summary`, and a refused run comes back to the caller as a value that no log setting can suppress — it is not logged.
55
+
56
+ ```js
57
+ import CronService from '@stonyx/cron/service';
58
+
59
+ const service = new CronService();
60
+ service.onJobDue = async (job) => ({ status: 'ok', summary: 'done' });
61
+
62
+ await service.start();
63
+ const job = await service.add({ name: 'Nightly', schedule: { kind: 'every', everyMs: 86_400_000 }, payload: { kind: 'agentTurn', message: 'go' } });
64
+
65
+ const result = await service.run(job.id, 'force');
66
+ ```
67
+
68
+ ### The `run()` contract
69
+
70
+ `run(id, mode)` resolves with an `ExecuteResult`. It **never** invokes the callback twice for one job, and it will refuse rather than queue:
71
+
72
+ | `status` | `reason` | Meaning |
73
+ | :---------- | :------------------ | :----------------------------------------------------------------------------- |
74
+ | `'ok'` | — | The callback resolved. `summary` and `durationMs` are set. |
75
+ | `'error'` | — | The callback threw or rejected. `error` carries the message; backoff is applied. |
76
+ | `'skipped'` | `'not due'` | `mode` was `'due'` and the job's next run time has not arrived. Use `'force'` to run anyway. |
77
+ | `'skipped'` | `'already running'` | A previous invocation of **this** job has not settled. The call is refused, not queued, and nothing is logged. |
78
+ | `'skipped'` | `'removed'` | The job was removed between the lookup and the claim. The callback did not fire. |
79
+
80
+ `run()` throws (rather than returning a result) when `id` is not a registered job: `Error: Job not found: <id>`.
81
+
82
+ **Concurrency.** A job is bounded to one in-flight invocation on every path — manual `run()` and the timer both claim it first. **Different** jobs are not bounded: the callback is deliberately invoked outside the internal lock, so N concurrent `run()` calls on N distinct jobs produce N concurrent callbacks. The scheduler itself never generates that fan-out (its timer path invokes a due batch sequentially); only a caller can. If you drive `run()` from a request handler, bound it on your side. Taking the callback out of the lock is what stops a callback that never settles from blocking `add`/`update`/`remove`; restoring the bound by putting it back would restore that deadlock.
83
+
84
+ ### Breaking changes in this line
85
+
86
+ Four consumer-visible changes landed with the phase split (#34). All are measured against the emitted `dist/service.d.ts`:
87
+
88
+ 1. **`ExecuteResult.reason` narrowed** from `string` to `'not due' | 'already running' | 'removed'`, and gained the `'removed'` member. Comparing it against a literal outside the union, or `switch`ing on one, is now a compile error (`TS2367` / `TS2678`). Assigning it into `string | undefined` and spreading it are unaffected. The type is exported as `SkipReason`.
89
+ 2. **`CronService` is nominally typed.** It carries ECMAScript hard-private members, so the declarations emit `#private;` and a structurally hand-built test double no longer assigns to `CronService` (`TS2741: Property '#private' is missing`). The break is one-directional: `class X extends CronService` still compiles, and assigning a real `CronService` to your own hand-written interface still compiles. **Migration:** declare your own interface and depend on that instead of a `CronService`-typed mock.
90
+ 3. **`claimJob`, `settleJob` and `executeClaimed` are not published.** They were never a supported API; a claim taken without its matching settle strands the job permanently.
91
+ 4. **`run()` no longer serializes across jobs** — see the concurrency note above.
92
+
93
+ `SkipReason`, `ExecuteResult`, `JobDueResult`, `ServiceStatus`, `ListOptions` and `OnJobDueCallback` are all exported from `@stonyx/cron/service`, so an exhaustive handler over `reason` is expressible.
94
+
50
95
  ## Configuration
51
96
 
52
97
  Optionally, informational logging and debugging can be controlled through `config.cron`:
package/dist/service.d.ts CHANGED
@@ -4,29 +4,37 @@ import RunLog from './run-log.js';
4
4
  interface HeapEntry extends HeapItem {
5
5
  key: string;
6
6
  }
7
- interface JobDueResult {
7
+ /**
8
+ * Why a `run()` did not invoke the callback. Exported so a consumer can write a
9
+ * total handler over it: the union is closed and narrowed (it was `string`
10
+ * before #34), so an exhaustive `switch` is now both possible and expected.
11
+ */
12
+ export type SkipReason = 'not due' | 'already running' | 'removed';
13
+ export interface JobDueResult {
8
14
  status?: string;
9
15
  error?: string;
10
16
  summary?: string;
11
17
  }
12
- interface ExecuteResult {
18
+ export interface ExecuteResult {
13
19
  status: string;
14
20
  error?: string;
15
21
  summary?: string;
16
22
  durationMs?: number;
17
23
  deleted?: boolean;
18
- reason?: string;
24
+ /** Only set when `status` is `'skipped'`. */
25
+ reason?: SkipReason;
19
26
  }
20
- interface ServiceStatus {
27
+ export interface ServiceStatus {
21
28
  started: boolean;
22
29
  jobCount: number;
23
30
  nextWakeAtMs: number | undefined;
24
31
  }
25
- interface ListOptions {
32
+ export interface ListOptions {
26
33
  includeDisabled?: boolean;
27
34
  }
28
- type OnJobDueCallback = (job: Job) => Promise<JobDueResult | void> | JobDueResult | void;
35
+ export type OnJobDueCallback = (job: Job) => Promise<JobDueResult | void> | JobDueResult | void;
29
36
  export default class CronService {
37
+ #private;
30
38
  jobs: Map<string, Job>;
31
39
  heap: MinHeap<HeapEntry>;
32
40
  timer: ReturnType<typeof setTimeout> | null;
@@ -69,6 +77,37 @@ export default class CronService {
69
77
  remove(id: string): Promise<void>;
70
78
  /**
71
79
  * Manually trigger a job.
80
+ *
81
+ * Returns `{ status: 'skipped', reason }` without invoking the callback when
82
+ * the job is not due (`mode: 'due'`), is already in flight
83
+ * (`'already running'`), or was removed before the claim landed
84
+ * (`'removed'`). Before the phase split, a forced run against an in-flight
85
+ * job launched a second concurrent invocation.
86
+ *
87
+ * THROWS (rather than returning a skip) when `id` is not a registered job:
88
+ * `Error("Job not found: <id>")`. A job that disappears between this lookup
89
+ * and the claim is the `'removed'` skip above, not a throw — the two differ
90
+ * only by the timing of a race, and the second is a legitimate outcome
91
+ * whereas the first is a caller error.
92
+ *
93
+ * CONCURRENCY: the same job is bounded to one in-flight invocation on every
94
+ * path, and the timer path invokes due jobs one at a time. `run()` fan-out
95
+ * across DIFFERENT jobs is deliberately unbounded — N concurrent `run()`
96
+ * calls produce N concurrent consumer callbacks. Before the phase split
97
+ * these serialized behind the module-global lock; that serialization was the
98
+ * bug rather than the feature (one hung callback wedged every other caller),
99
+ * so it is not restored here. The fan-out is caller-driven and the scheduler
100
+ * never produces it on its own. A per-invoke bound belongs above this layer;
101
+ * it is tracked on stonyx-cron#35 alongside the execution timeout.
102
+ *
103
+ * NOT FIXED HERE: the phase split fixes the LOCK wedge, not the TIMER. A
104
+ * callback that never settles still stops `onTimer`'s sequential loop
105
+ * forever — `running` stays true, every later tick early-returns and re-arms,
106
+ * and the hung job's batch siblings stay claimed and off-heap having never
107
+ * been invoked. CRUD still resolves and `status()` still reports
108
+ * `started: true`, so that failure is now silent where it used to be loud.
109
+ * Bounding the callback and releasing batch siblings is stonyx-cron#35. Do
110
+ * not read this method's doc as "the hang is fixed".
72
111
  */
73
112
  run(id: string, mode?: 'due' | 'force'): Promise<ExecuteResult>;
74
113
  /**
@@ -78,6 +117,25 @@ export default class CronService {
78
117
  armTimer(): void;
79
118
  onTimer(): Promise<void>;
80
119
  findDueJobs(nowMs: number): Job[];
120
+ /**
121
+ * Execute a job in three phases:
122
+ *
123
+ * 1. claim (locked) — take ownership of the job, detach it from the heap
124
+ * 2. invoke (UNLOCKED) — await the consumer callback
125
+ * 3. settle (locked) — apply the result, log it, re-insert into the heap
126
+ *
127
+ * The critical section deliberately excludes phase 2. `onJobDue` is
128
+ * arbitrary, unbounded consumer code; awaiting it under the module-global
129
+ * lock is what wedged every subsequent `locked()` call (add/update/remove)
130
+ * when a callback never settled.
131
+ *
132
+ * `onTimer` performs the batch claim (`findDueJobs` + `markRunning`) for all
133
+ * due jobs under a single lock and then enters at phase 2 via
134
+ * `#executeClaimed`. That entry point is `#private` rather than a parameter
135
+ * on this method: as a published `alreadyClaimed` flag it would be a
136
+ * supported way to skip phase 1 entirely, defeating the claim guard and
137
+ * allowing concurrent `onJobDue` invocations for the same job.
138
+ */
81
139
  executeJob(job: Job): Promise<ExecuteResult>;
82
140
  removeFromHeap(id: string): void;
83
141
  log(message: string): void;
package/dist/service.js CHANGED
@@ -2,7 +2,21 @@
2
2
  * CronService - the main API for advanced job scheduling.
3
3
  *
4
4
  * Manages jobs in memory with a min-heap for efficient next-job lookup.
5
- * All state mutations are serialized via async locking.
5
+ *
6
+ * Async locking, but never around the consumer callback. Execution is split
7
+ * into three phases: claim (locked), invoke (UNLOCKED), settle (locked). Only
8
+ * phases 1 and 3 are serialized; the critical section deliberately excludes
9
+ * phase 2, so a consumer callback that never settles cannot wedge the lock
10
+ * chain and block add/update/remove. See `run()` and `#executeClaimed`.
11
+ *
12
+ * CONSUMER NOTE — this class is NOMINALLY typed. It carries ECMAScript hard-
13
+ * private members, so `dist/service.d.ts` emits `#private;` on the class and a
14
+ * structurally-built test double will not assign to `CronService`
15
+ * (`TS2741: Property '#private' is missing`). The break is one-directional and
16
+ * has a zero-cost workaround: `extends CronService` still compiles, and
17
+ * assigning a real `CronService` to your own hand-written interface still
18
+ * compiles. Declare your own interface and depend on that rather than
19
+ * hand-building a `CronService`-typed mock. See README's `CronService` section.
6
20
  */
7
21
  import config from 'stonyx/config';
8
22
  import log from 'stonyx/log';
@@ -12,6 +26,61 @@ import { locked } from './locked.js';
12
26
  import { normalizeJobInput, recoverFlatParams } from './normalize.js';
13
27
  import RunLog from './run-log.js';
14
28
  const MAX_TIMER_DELAY_MS = 60_000;
29
+ /** Longest error text that may reach a log line. Anything past this is truncated. */
30
+ const MAX_LOGGED_ERROR_LENGTH = 512;
31
+ /** Longest job name that may reach a log line. Anything past this is truncated. */
32
+ const MAX_LOGGED_NAME_LENGTH = 120;
33
+ /**
34
+ * Flatten a value for interpolation into a single log line.
35
+ *
36
+ * Chronicle writes `${timestamp} ${content}\n` to a newline-delimited file, so
37
+ * any `\r` or `\n` inside `content` ends the record early and everything after
38
+ * it is read back as a separate, attacker-shaped entry — including a forged
39
+ * `[timestamp] Cron — ...` prefix that is indistinguishable from a real one.
40
+ * Both values that reach these lines are untrusted: `job.name` is passed
41
+ * through `createJob` unvalidated and `normalize.ts` exists specifically to
42
+ * accept AI-shaped input, and an error message is arbitrary consumer-callback
43
+ * text. Newlines become the literal two characters so the content survives for
44
+ * a reader, and the length cap keeps one pathological value from swamping the
45
+ * file.
46
+ */
47
+ function forLog(value, maxLength) {
48
+ const flattened = value.replace(/\r\n|[\r\n\u2028\u2029]/g, '\\n');
49
+ return flattened.length > maxLength ? `${flattened.slice(0, maxLength)}...` : flattened;
50
+ }
51
+ /**
52
+ * Describe a thrown value without ever throwing.
53
+ *
54
+ * `String(err)` is not total: a null-prototype object, or any object whose
55
+ * `toString`/`Symbol.toPrimitive` throws, raises "Cannot convert object to
56
+ * primitive value". Consumer callbacks throw arbitrary values, so the error
57
+ * handler must not become a second failure source of its own.
58
+ *
59
+ * The `instanceof Error` branch needs the same guard as the other one. `Error`
60
+ * is subclassable and `message` is a plain writable property, so a consumer can
61
+ * hand back an `Error` whose `message` is a getter that throws, or one that is
62
+ * an object whose `toString` throws — `Error.prototype.message` is typed
63
+ * `string`, so TypeScript sees nothing wrong and the coercion is deferred to
64
+ * the caller's template literal, outside every guard here. That made `run()`
65
+ * reject instead of returning an `ExecuteResult`: a contract violation in the
66
+ * function written to prevent exactly that. Reading and coercing `message`
67
+ * inside the `try` is what closes it.
68
+ *
69
+ * `Cron` in `main.ts` carries its own `describeError` and the two deliberately
70
+ * differ: it renders for a log line only, so it prefers `err.stack`; this one
71
+ * is also returned to the caller as `ExecuteResult.error` and persisted in a
72
+ * per-job run log, where a stack would be an unbounded blob in every stored
73
+ * failure. Recorded in `docs/architecture.md` under Error Handling — do not
74
+ * merge them into a shared helper without reading that first.
75
+ */
76
+ function describeError(err) {
77
+ try {
78
+ return err instanceof Error ? String(err.message) : String(err);
79
+ }
80
+ catch {
81
+ return 'unknown error';
82
+ }
83
+ }
15
84
  export default class CronService {
16
85
  jobs;
17
86
  heap;
@@ -38,15 +107,60 @@ export default class CronService {
38
107
  if (this.started)
39
108
  return;
40
109
  this.started = true;
41
- if (initialJobs) {
42
- for (const job of initialJobs) {
43
- this.jobs.set(job.id, job);
44
- if (job.enabled && job.state.nextRunAtMs) {
45
- this.heap.push({ key: job.id, nextTrigger: job.state.nextRunAtMs });
110
+ // `finally`, not a trailing statement. Reconciled with the same guard #53
111
+ // puts around `register`'s `runOnInit` invocation: a scheduler that is
112
+ // marked started but never armed is the terminal state both fixes exist to
113
+ // remove, reached here through the other entry point. `initialJobs` crosses
114
+ // a serialization boundary — it is whatever the consumer's store handed
115
+ // back — so `Job[]` is a compile-time claim about runtime data, and a row
116
+ // missing `state` throws mid-loop. Without this, `start()` leaves
117
+ // `started: true` (so it is now a no-op), the rows registered before the
118
+ // throw sitting in the heap, and NO timer: measured, `status()` then
119
+ // reports `{ started: true, jobCount: 1, nextWakeAtMs: <real> }` while
120
+ // nothing will ever fire. Silent and healthy-looking, again.
121
+ try {
122
+ if (initialJobs) {
123
+ for (const job of initialJobs) {
124
+ // A `runningAtMs` on a rehydrated job is always stale. The claim it
125
+ // records was taken by a process that is gone, so nothing will ever
126
+ // settle it, and nothing reaps it — there is no lease on the field
127
+ // (tracked on #35). Left in place it is a permanently dead job that
128
+ // still reports healthy: `isDue` returns false forever because of the
129
+ // flag, `run()` answers `'already running'` forever, `update()` never
130
+ // touches `state.runningAtMs`, and `status()` counts it like any other.
131
+ // The consumer's only recovery would be remove() + add(), losing the
132
+ // job id and its run history.
133
+ //
134
+ // Same hazard, same treatment as the hand-release on the `'removed'`
135
+ // path in `#executeClaimed`: a claim with no reachable settle must be
136
+ // released. Assigned directly rather than via `applyResult` for the same
137
+ // reason — this releases the claim and nothing else. The job did not
138
+ // run, so it gets no run-log row, no `lastStatus`, and no recomputed
139
+ // `nextRunAtMs`; it is rescheduled from the store's own value below.
140
+ //
141
+ // Written unconditionally, and deliberately NOT guarded on a
142
+ // `job.state.runningAtMs` read. Guarding it makes `start()` accept a
143
+ // row whose `state` is frozen — `structuredClone` + `Object.freeze`
144
+ // is an ordinary defensive rehydration — and that row is not usable
145
+ // by this class at all: `markRunning` writes the same field on every
146
+ // execution. Measured, the guard moves the failure from a throw out
147
+ // of `start()`, which the consumer's own `await` can catch, to a
148
+ // TypeError raised inside `onTimer`'s batch claim — a bare timer
149
+ // callback, so it surfaces as an unhandled rejection and is
150
+ // process-fatal under Node's default. Failing loudly at the store
151
+ // boundary is the better of the two, and the `finally` above keeps
152
+ // the rows that loaded before it scheduled.
153
+ job.state.runningAtMs = undefined;
154
+ this.jobs.set(job.id, job);
155
+ if (job.enabled && job.state.nextRunAtMs) {
156
+ this.heap.push({ key: job.id, nextTrigger: job.state.nextRunAtMs });
157
+ }
46
158
  }
47
159
  }
48
160
  }
49
- this.armTimer();
161
+ finally {
162
+ this.armTimer();
163
+ }
50
164
  }
51
165
  /**
52
166
  * Stop the service. Clears timer.
@@ -136,6 +250,37 @@ export default class CronService {
136
250
  }
137
251
  /**
138
252
  * Manually trigger a job.
253
+ *
254
+ * Returns `{ status: 'skipped', reason }` without invoking the callback when
255
+ * the job is not due (`mode: 'due'`), is already in flight
256
+ * (`'already running'`), or was removed before the claim landed
257
+ * (`'removed'`). Before the phase split, a forced run against an in-flight
258
+ * job launched a second concurrent invocation.
259
+ *
260
+ * THROWS (rather than returning a skip) when `id` is not a registered job:
261
+ * `Error("Job not found: <id>")`. A job that disappears between this lookup
262
+ * and the claim is the `'removed'` skip above, not a throw — the two differ
263
+ * only by the timing of a race, and the second is a legitimate outcome
264
+ * whereas the first is a caller error.
265
+ *
266
+ * CONCURRENCY: the same job is bounded to one in-flight invocation on every
267
+ * path, and the timer path invokes due jobs one at a time. `run()` fan-out
268
+ * across DIFFERENT jobs is deliberately unbounded — N concurrent `run()`
269
+ * calls produce N concurrent consumer callbacks. Before the phase split
270
+ * these serialized behind the module-global lock; that serialization was the
271
+ * bug rather than the feature (one hung callback wedged every other caller),
272
+ * so it is not restored here. The fan-out is caller-driven and the scheduler
273
+ * never produces it on its own. A per-invoke bound belongs above this layer;
274
+ * it is tracked on stonyx-cron#35 alongside the execution timeout.
275
+ *
276
+ * NOT FIXED HERE: the phase split fixes the LOCK wedge, not the TIMER. A
277
+ * callback that never settles still stops `onTimer`'s sequential loop
278
+ * forever — `running` stays true, every later tick early-returns and re-arms,
279
+ * and the hung job's batch siblings stay claimed and off-heap having never
280
+ * been invoked. CRUD still resolves and `status()` still reports
281
+ * `started: true`, so that failure is now silent where it used to be loud.
282
+ * Bounding the callback and releasing batch siblings is stonyx-cron#35. Do
283
+ * not read this method's doc as "the hang is fixed".
139
284
  */
140
285
  async run(id, mode = 'force') {
141
286
  const job = this.jobs.get(id);
@@ -144,6 +289,10 @@ export default class CronService {
144
289
  if (mode === 'due' && !isDue(job, Date.now())) {
145
290
  return { status: 'skipped', reason: 'not due' };
146
291
  }
292
+ // Deliberately NOT wrapped in locked(): executeJob takes the lock itself,
293
+ // for its claim and settle phases only. Wrapping here would re-create the
294
+ // wedge through a second door, because the consumer callback would once
295
+ // again be awaited while a lock is held.
147
296
  return this.executeJob(job);
148
297
  }
149
298
  /**
@@ -172,16 +321,77 @@ export default class CronService {
172
321
  }
173
322
  this.running = true;
174
323
  try {
175
- await locked(async () => {
324
+ // -- Phase 1: claim (locked), batched --
325
+ // Collecting due jobs pops them off the heap, and marking them running
326
+ // makes them un-claimable by anyone else. Both must happen under the
327
+ // same lock turn, or a concurrent run() could claim a job this batch has
328
+ // already detached.
329
+ //
330
+ // This is the SECOND claim implementation — `#claimJob` is the other, and
331
+ // the two reach the same state by different routes. `#claimJob` guards
332
+ // with an explicit `job.state.runningAtMs` check; this path has no such
333
+ // check and relies entirely on `isDue`'s `!job.state.runningAtMs` clause
334
+ // (`job.ts`) to keep `findDueJobs` from re-claiming a job that `run()`
335
+ // already holds. THAT CLAUSE IS LOAD-BEARING HERE, not an optimisation:
336
+ // drop it and the timer path silently double-invokes a job that `run()`
337
+ // is mid-flight on, while `run()` keeps refusing correctly and looks
338
+ // healthy. The one-in-flight-invocation-per-job invariant this class
339
+ // advertises holds by two independent guards in two files; a change to
340
+ // either has to be checked against the other.
341
+ const dueJobs = await locked(() => {
176
342
  const nowMs = Date.now();
177
- const dueJobs = this.findDueJobs(nowMs);
178
- for (const job of dueJobs) {
343
+ const due = this.findDueJobs(nowMs);
344
+ for (const job of due) {
179
345
  markRunning(job);
180
346
  }
181
- for (const job of dueJobs) {
182
- await this.executeJob(job);
183
- }
347
+ return due;
184
348
  });
349
+ // Phases 2 and 3 run OUTSIDE the claim lock. The consumer callback is
350
+ // awaited here holding no lock at all, so a callback that never settles
351
+ // cannot poison the lock chain and wedge add/update/remove.
352
+ for (const job of dueJobs) {
353
+ try {
354
+ await this.#executeClaimed(job);
355
+ }
356
+ catch (err) {
357
+ // One job's unexpected throw must not abort the batch. Every job in
358
+ // `dueJobs` is already claimed — marked running and detached from
359
+ // the heap — and only its own settle releases it, so aborting here
360
+ // would strand every sibling permanently un-due.
361
+ //
362
+ // Reported on an UNGATED channel. `this.log()` returns early when
363
+ // `config.cron.log` is false, which is a supported production
364
+ // setting, and a failure here permanently unschedules a job while
365
+ // `status()` keeps reporting the service healthy. Silent-and-healthy
366
+ // is exactly the failure class this split exists to remove.
367
+ //
368
+ // This is the outermost handler on the timer path, so it is the one
369
+ // that must not be able to throw: `log` is a shared singleton whose
370
+ // transports can reach the filesystem, so its own failure is
371
+ // swallowed rather than allowed to take the batch down.
372
+ //
373
+ // BOTH halves of that failure have to be caught, and they are caught
374
+ // by different constructs. `log.error` is a chronicle convenience
375
+ // method that returns `logAction(...)` -> `async log(...)`, so its
376
+ // console write, colour lookup, `mkdirSync` and `appendFile` all
377
+ // surface as REJECTIONS, never as synchronous throws. A bare call
378
+ // here escapes this `catch` entirely and terminates the process under
379
+ // Node's default `--unhandled-rejections=throw` — the handler written
380
+ // so it "must not be able to throw" would be the one taking the
381
+ // daemon down. The `try` covers the synchronous half (evaluating the
382
+ // template literal); `Promise.resolve(...).catch()` covers the async
383
+ // half. Deliberately not awaited: the batch must not block on a log
384
+ // transport, and `void` marks the floated promise as intentional.
385
+ try {
386
+ void Promise.resolve(log.error(`Cron — Job "${forLog(job.name, MAX_LOGGED_NAME_LENGTH)}" (${job.id}) execution failed unexpectedly: ${forLog(describeError(err), MAX_LOGGED_ERROR_LENGTH)}`)).catch(() => {
387
+ // Nothing left to report to.
388
+ });
389
+ }
390
+ catch {
391
+ // Nothing left to report to.
392
+ }
393
+ }
394
+ }
185
395
  }
186
396
  finally {
187
397
  this.running = false;
@@ -202,50 +412,190 @@ export default class CronService {
202
412
  }
203
413
  return due;
204
414
  }
415
+ /**
416
+ * Execute a job in three phases:
417
+ *
418
+ * 1. claim (locked) — take ownership of the job, detach it from the heap
419
+ * 2. invoke (UNLOCKED) — await the consumer callback
420
+ * 3. settle (locked) — apply the result, log it, re-insert into the heap
421
+ *
422
+ * The critical section deliberately excludes phase 2. `onJobDue` is
423
+ * arbitrary, unbounded consumer code; awaiting it under the module-global
424
+ * lock is what wedged every subsequent `locked()` call (add/update/remove)
425
+ * when a callback never settled.
426
+ *
427
+ * `onTimer` performs the batch claim (`findDueJobs` + `markRunning`) for all
428
+ * due jobs under a single lock and then enters at phase 2 via
429
+ * `#executeClaimed`. That entry point is `#private` rather than a parameter
430
+ * on this method: as a published `alreadyClaimed` flag it would be a
431
+ * supported way to skip phase 1 entirely, defeating the claim guard and
432
+ * allowing concurrent `onJobDue` invocations for the same job.
433
+ */
205
434
  async executeJob(job) {
435
+ // -- Phase 1: claim (locked) --
436
+ const refusal = await locked(() => this.#claimJob(job));
437
+ if (refusal)
438
+ return { status: 'skipped', reason: refusal };
439
+ return this.#executeClaimed(job);
440
+ }
441
+ /**
442
+ * Phases 2 and 3 for a job that has already been claimed — either by
443
+ * `executeJob` above or by `onTimer`'s batch claim.
444
+ *
445
+ * Private: reaching this without a claim would run the consumer callback for
446
+ * a job nobody owns, and would leave nothing to release the claim.
447
+ */
448
+ async #executeClaimed(job) {
449
+ // Membership re-check. The claim and the invoke are no longer in the same
450
+ // critical section, and sibling callbacks run unlocked, so a `remove()` can
451
+ // now land in between AND RESOLVE — it used to deadlock. A resolved
452
+ // `remove()` must keep meaning "this callback will not fire"; the identity
453
+ // guard in `#settleJob` only cleans up afterwards, by which point the side
454
+ // effect has already happened. Identity, not id, so a removed-then-replaced
455
+ // key is caught too. Deliberately synchronous with the `onJobDue` call
456
+ // below — nothing can interleave between this check and the invocation.
457
+ //
458
+ // This is the one early return after a claim, so it is the one that has to
459
+ // release the claim by hand. Skipping settle is right — re-inserting or
460
+ // run-logging a removed job is the resurrection `#settleJob` refuses, and
461
+ // the heap entry is already gone. But the claim must still come off,
462
+ // because the detached object is NOT unreachable: it is the object `add()`
463
+ // returned and `get()`/`list()` hand out, and `start(initialJobs)`
464
+ // re-registers those objects verbatim, `state` included. A leftover
465
+ // `runningAtMs` rehydrates a permanently dead job — `isDue` false forever,
466
+ // `run()` refused forever, `status()` reporting it healthy.
467
+ //
468
+ // Assigned directly rather than via `applyResult`: this releases the claim
469
+ // and nothing else. No run-log row, no heap entry, no `lastStatus`, no
470
+ // recomputed `nextRunAtMs` — the job did not run.
471
+ if (this.jobs.get(job.id) !== job) {
472
+ job.state.runningAtMs = undefined;
473
+ return { status: 'skipped', reason: 'removed' };
474
+ }
206
475
  const startMs = Date.now();
207
476
  let status = 'ok';
208
477
  let error;
209
478
  let summary;
479
+ let settled;
480
+ // The claim marked the job running and detached it from the heap. Phase 3
481
+ // is the ONLY thing that undoes either, so it must survive every non-local
482
+ // exit from phase 2 — including a throw from the catch handler itself
483
+ // (`this.log` is public and overridable and reaches a transport). A claim
484
+ // with no matching settle is not a degraded state, it is a permanently
485
+ // dead job.
210
486
  try {
211
- if (this.onJobDue) {
212
- const result = await this.onJobDue(job);
213
- if (result) {
214
- status = result.status || 'ok';
215
- error = result.error;
216
- summary = result.summary;
487
+ // -- Phase 2: invoke (NOT locked) --
488
+ try {
489
+ if (this.onJobDue) {
490
+ const result = await this.onJobDue(job);
491
+ if (result) {
492
+ status = result.status || 'ok';
493
+ error = result.error;
494
+ summary = result.summary;
495
+ }
217
496
  }
218
497
  }
498
+ catch (err) {
499
+ status = 'error';
500
+ error = describeError(err);
501
+ this.log(`Job "${forLog(job.name, MAX_LOGGED_NAME_LENGTH)}" (${job.id}) failed: ${forLog(error, MAX_LOGGED_ERROR_LENGTH)}`);
502
+ }
219
503
  }
220
- catch (err) {
221
- status = 'error';
222
- error = err instanceof Error ? err.message : String(err);
223
- this.log(`Job "${job.name}" (${job.id}) failed: ${error}`);
504
+ finally {
505
+ // -- Phase 3: settle (locked) --
506
+ settled = await locked(() => this.#settleJob(job, status, error, summary, startMs, Date.now() - startMs));
224
507
  }
225
- const durationMs = Date.now() - startMs;
226
- const validStatus = (status === 'ok' || status === 'error' || status === 'skipped') ? status : 'error';
227
- applyResult(job, validStatus, error, durationMs);
228
- // Log the run
229
- this.runLog.record({
230
- jobId: job.id,
231
- status,
232
- error,
233
- summary,
234
- runAtMs: startMs,
235
- durationMs,
236
- nextRunAtMs: job.state.nextRunAtMs,
237
- });
238
- // Handle one-shot auto-delete
239
- if (job.deleteAfterRun && status === 'ok' && !job.enabled) {
240
- this.jobs.delete(job.id);
241
- this.runLog.removeJob(job.id);
242
- return { status, summary, deleted: true };
508
+ return settled;
509
+ }
510
+ /**
511
+ * Phase 1 — claim. Must be called while holding the lock (`locked()`, whose
512
+ * chain is module-global and therefore shared across CronService instances).
513
+ *
514
+ * Returns `null` on a successful claim, or the reason the claim was refused.
515
+ * `'already running'` is what makes a second `run()` report a skip instead of
516
+ * launching a concurrent invocation. `'removed'` covers the job being deleted
517
+ * between `run()`'s unlocked lookup and this lock turn — claiming then would
518
+ * `markRunning` an orphan and, worse, `removeFromHeap` an id that may now
519
+ * belong to a replacement.
520
+ *
521
+ * Detaching from the heap here — rather than relying on phase 3 to push a
522
+ * fresh entry — is what stops manual runs permanently duplicating entries.
523
+ *
524
+ * `#private`: published, this would be a supported call performing
525
+ * `markRunning` + `removeFromHeap` with no guaranteed settle and no lease on
526
+ * `runningAtMs`, so a single such call would strand the job forever. The
527
+ * lock-held precondition cannot be expressed in the type system, so the
528
+ * method must not be reachable from outside the class body.
529
+ */
530
+ #claimJob(job) {
531
+ if (this.jobs.get(job.id) !== job)
532
+ return 'removed';
533
+ if (job.state.runningAtMs)
534
+ return 'already running';
535
+ markRunning(job);
536
+ this.removeFromHeap(job.id);
537
+ return null;
538
+ }
539
+ /**
540
+ * Phase 3 — settle. Must be called while holding the lock.
541
+ *
542
+ * `#private` for the same reason as `#claimJob`: unlocked it would run
543
+ * `applyResult`, a `runLog.record`, a full `removeFromHeap` rebuild, a
544
+ * `heap.push` and an `armTimer` with no mutual exclusion — exactly the
545
+ * corruption `locked()` exists to prevent.
546
+ */
547
+ #settleJob(job, status, error, summary, startMs, durationMs) {
548
+ try {
549
+ const validStatus = (status === 'ok' || status === 'error' || status === 'skipped') ? status : 'error';
550
+ applyResult(job, validStatus, error, durationMs);
551
+ // The callback ran unlocked, so this job may have been removed — or
552
+ // removed and re-registered under the same id, the shape
553
+ // `start(initialJobs)` uses — while it was in flight. Identity, not id.
554
+ //
555
+ // Deliberately touch NOTHING here. The claim already detached this job's
556
+ // heap entry and nothing re-added it, so there is nothing to clean up;
557
+ // any entry now filed under this id belongs to the replacement, and
558
+ // removing it by id would silently unschedule a live job. Do not
559
+ // resurrect a removed job's heap entry or run log either.
560
+ if (this.jobs.get(job.id) !== job) {
561
+ return { status, error, summary, durationMs };
562
+ }
563
+ // Log the run
564
+ this.runLog.record({
565
+ jobId: job.id,
566
+ status,
567
+ error,
568
+ summary,
569
+ runAtMs: startMs,
570
+ durationMs,
571
+ nextRunAtMs: job.state.nextRunAtMs,
572
+ });
573
+ // Handle one-shot auto-delete. The callback ran unlocked and may have
574
+ // pushed a heap entry for this job via add()/update(), so drop it — the
575
+ // job is about to stop existing.
576
+ if (job.deleteAfterRun && status === 'ok' && !job.enabled) {
577
+ this.jobs.delete(job.id);
578
+ this.removeFromHeap(job.id);
579
+ this.runLog.removeJob(job.id);
580
+ return { status, summary, deleted: true };
581
+ }
582
+ // Re-insert into the heap if still active. Same reason as above: drop any
583
+ // entry the unlocked callback added for this job first, to preserve
584
+ // one-entry-per-key.
585
+ this.removeFromHeap(job.id);
586
+ if (job.enabled && job.state.nextRunAtMs) {
587
+ this.heap.push({ key: job.id, nextTrigger: job.state.nextRunAtMs });
588
+ }
589
+ return { status, error, summary, durationMs };
243
590
  }
244
- // Re-insert into heap if still active
245
- if (job.enabled && job.state.nextRunAtMs) {
246
- this.heap.push({ key: job.id, nextTrigger: job.state.nextRunAtMs });
591
+ finally {
592
+ // One re-arm covering every exit, rather than one per branch. The claim
593
+ // detached this job from the heap, so a timer that fired during the
594
+ // unlocked invoke would have found nothing to arm — and `run()` has no
595
+ // `finally { armTimer() }` of its own the way `onTimer` does. Without
596
+ // this, a manual run() can leave the scheduler with no pending wake.
597
+ this.armTimer();
247
598
  }
248
- return { status, error, summary, durationMs };
249
599
  }
250
600
  // -- Helpers ---------------------------------------------------------
251
601
  removeFromHeap(id) {
package/package.json CHANGED
@@ -3,7 +3,7 @@
3
3
  "keywords": [
4
4
  "stonyx-module"
5
5
  ],
6
- "version": "0.2.1-alpha.59",
6
+ "version": "0.2.1-alpha.60",
7
7
  "description": "Cron/job scheduler for Stonyx framework",
8
8
  "main": "dist/main.js",
9
9
  "types": "dist/main.d.ts",
@@ -69,7 +69,7 @@
69
69
  },
70
70
  "homepage": "https://github.com/abofs/stonyx-cron#readme",
71
71
  "devDependencies": {
72
- "@stonyx/utils": "0.2.3-beta.26",
72
+ "@stonyx/utils": "0.2.3-beta.27",
73
73
  "@types/node": "^25.5.2",
74
74
  "@types/qunit": "^2.19.13",
75
75
  "@types/sinon": "^21.0.1",
@@ -79,7 +79,7 @@
79
79
  "typescript": "^5.8.3"
80
80
  },
81
81
  "dependencies": {
82
- "stonyx": "0.2.3-beta.83"
82
+ "stonyx": "0.2.3-beta.93"
83
83
  },
84
84
  "scripts": {
85
85
  "build": "tsc",