@edgehero/pi-dispatch 4.0.0 → 4.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,654 @@
1
+ /**
2
+ * How busy each host is and was (issue #599, DES-CAPACITY-FROM-RECORDS): ONE pure function, `computeCapacity`, over
3
+ * run records (INT-RUN-HISTORY-FILE-CONTRACT) and the live host rows, which the CLI (`pi-dispatch capacity`) calls and
4
+ * every later surface will, so they can never disagree. The report is `INT-CAPACITY-REPORT`.
5
+ *
6
+ * Pure: no clock, no filesystem, no Valkey. It imports only other pure modules of this project (the size and wait
7
+ * vocabularies, the worker-name and project-id rules), so the admin bundle can inline it. It reads a record's
8
+ * `capacity` itself (`capacityOf`) rather than importing the writer's `recordedCapacity`, which lives beside the record
9
+ * writer's filesystem code; a test holds the two to the same answers.
10
+ *
11
+ * JOBS ONLY. Busy means "a job of this deployment held a slot here", read from the records alone: no host load is
12
+ * sampled, so a machine busy with other work reads as idle, and every surface says so.
13
+ *
14
+ * WHICH RECORDS HELD A SLOT (`occupancyOf`). A record written since #599 says it outright: `capacity` is an object on a
15
+ * run the processor admitted and null on every refusal before a slot. An older record has no such key, and is
16
+ * inferred: a refusal reason the processor gives before a slot (the wait gate's, the never-fits pair) did not hold one;
17
+ * otherwise it did when it ran at least `LEGACY_MIN_WALL_MS` or reported `resources` (a container ran). Each inference is
18
+ * counted, both ways, never hidden.
19
+ *
20
+ * THE SPAN is the record's `startedAt` (set when the job was admitted, index.mjs) to its `endedAt`. A span that is not
21
+ * readable, ends before it starts, runs longer than `SUGGEST_MAX_WALL_MS` or ends more than `SUGGEST_CLOCK_SKEW_MS` in
22
+ * the future is not counted, and is counted as `unreadable`: the same bounds a size suggestion reads a wall time by. So
23
+ * is a record whose host is not a worker name or whose project is not a project id: neither can come from a worker,
24
+ * and both are printed to a terminal.
25
+ *
26
+ * MISSING HISTORY IS NEVER IDLE, and it is judged PER HOST. A host's history starts where the best source that holds
27
+ * ALL of its runs starts: this host's own files (and a peer's, on a shared logs directory) from the log retention on,
28
+ * and a named host's runs in the run mirror from the mirror's own start on (its window, its cap, and the fleet horizon
29
+ * every writer's trim raises). A host no source holds (a live worker with no `PI_WORKER_NAME`, seen from another host)
30
+ * is missing for the whole window. Missing time is counted in neither busy nor idle, so busy + idle + missing is the
31
+ * window for every host, and a truncated history can never read as a quiet machine.
32
+ *
33
+ * A JOB RUNNING NOW has no record yet, so the live rows supply it (issue #599, phase 2): each host row lists the jobs
34
+ * it runs (`jobs`, live-jobs.mjs), and each is counted as an occupied interval from its admission to now, with its size
35
+ * and project, marked live (no wait, no CPU measured) and counted in `live`, never in `used`. Only what the row can
36
+ * vouch for: a row that has not beaten within `LIVE_FRESH_MS` vouches only up to its last beat, so the interval ends
37
+ * there; a job whose record is already in the window (it ended between the beat and this read) is the record's; an
38
+ * orphan (`o`, a container whose stop did not take) is not counted, since its record already covers its run; and a
39
+ * running job the row does not list (past its 32, an entry its allowlist drops, a worker from before the field, a host
40
+ * whose history is not shared) is counted in `liveNotCounted`, never guessed. A row whose `jobs` value is there and is
41
+ * not a list says nothing about how many it runs: that host is counted in `liveUnreadable`, and `running` is unknown
42
+ * (null). A listed job is matched to its record by the id as the row publishes it (`publishedJobId`). The reading host
43
+ * (`coverage.localHost`) with runs here and no live row (the registry not read, or its row gone until its next beat) is
44
+ * named in `liveRowMissing`: the jobs it runs now are not counted, and `running` is the other rows' count.
45
+ *
46
+ * RETRIES AND STALLS. A retry's record replaces its earlier attempt's, so the record carries those attempts' slot
47
+ * intervals (`earlier`), each counted as an occupied interval on its own host with its size (no wait, no CPU measured).
48
+ * A pickup after a stall (`stalledRepick`) follows a pickup that wrote no record, whose time is unknown: it is counted,
49
+ * and said, never guessed. Neither adds a wait.
50
+ *
51
+ * THE CAPACITY IN FORCE. Every admitted run recorded what its host offered when it started, so the host's capacity is a
52
+ * step function of time (each recorded value holds from its run's start until the next one; before the first, the first
53
+ * holds). Each piece of time is judged against the capacity in force then: the slot count for "full", the budget for
54
+ * what was promised, the host's CPUs for what was used. A promise above 100% is then a real over-commit (a budget
55
+ * lowered while jobs ran), and is reported as one. CPU used never is: in every moment the jobs' measured CPU time is
56
+ * held to the host's CPUs, so a container that inflates its own numbers cannot push the host past what it has.
57
+ *
58
+ * INTEGER MATH. Every instant is a millisecond count from `Date.parse` (a safe integer), and every span, busy, idle and
59
+ * full time is at most the window, so those stay Numbers. A sum of run time is NOT bounded by the window (it is the
60
+ * window times the concurrency), and a product with a memory size, a CPU size or a CPU time can pass 2^53 (256 CPUs of
61
+ * CPU time over 7 days is about 1.5e14 microseconds before it is multiplied by a span), so every such sum and product is
62
+ * a BigInt, and only the final per-mille ratios, rounded half up, come back as Numbers.
63
+ */
64
+
65
+ import { WORKER_NAME_RE } from "./worker-name.mjs";
66
+ import { HOST_CPUS_MAX, HOST_MEMORY_MAX_MIB, HOST_SLOTS_MAX, JOB_CPUS_CEILING_CENTI, SIZE_REFUSAL_REASONS, recordedJobSize } from "./job-size.mjs";
67
+ import { isProjectId } from "./project-id.mjs";
68
+ import { HOST_BEAT_MS } from "./host-registry.mjs";
69
+ import { parseJobsMore, parseLiveJobs, publishedJobId } from "./live-jobs.mjs";
70
+ import { recordedEarlier } from "./run-earlier.mjs";
71
+ import { SUGGEST_CLOCK_SKEW_MS, SUGGEST_MAX_WALL_MS } from "./size-suggest.mjs";
72
+ import { WAIT_REFUSAL_REASONS } from "./wait-for.mjs";
73
+
74
+ /** The report's version (`INT-CAPACITY-REPORT`). */
75
+ export const CAPACITY_REPORT_VERSION = 1;
76
+ /** A record from before `capacity` held a slot when it ran at least this long (or reported `resources`). */
77
+ export const LEGACY_MIN_WALL_MS = 1000;
78
+ /** How many projects a host lists by name; the rest are summed as other. */
79
+ export const CAPACITY_TOP_PROJECTS = 5;
80
+ /** The most buckets one report splits its window into, so a caller's mistake cannot cost a render its memory. */
81
+ export const CAPACITY_MAX_BUCKETS = 10_000;
82
+ /**
83
+ * How recent a host row's beat must be for its running jobs to count up to NOW: two beats. A row older than that (its
84
+ * worker slow, stopped or gone, the row not yet expired) vouches for its jobs only up to its last beat.
85
+ */
86
+ export const LIVE_FRESH_MS = 2 * HOST_BEAT_MS;
87
+ /** The refusal reasons the processor gives BEFORE a job holds a slot: a legacy record carrying one never held one. */
88
+ export const PRE_SLOT_REFUSAL_REASONS = Object.freeze([...WAIT_REFUSAL_REASONS, ...SIZE_REFUSAL_REASONS]);
89
+
90
+ const HOUR_MS = 60 * 60 * 1000;
91
+ const DAY_MS = 24 * HOUR_MS;
92
+ /**
93
+ * The windows a surface offers, each with the bucket its timeline is split into: an hour over a day, six hours over a
94
+ * week, a day over thirty. One table, so the CLI, doctor and the panel name the same three.
95
+ */
96
+ export const CAPACITY_WINDOWS = Object.freeze({
97
+ "24h": Object.freeze({ ms: DAY_MS, bucketMs: HOUR_MS }),
98
+ "7d": Object.freeze({ ms: 7 * DAY_MS, bucketMs: 6 * HOUR_MS }),
99
+ "30d": Object.freeze({ ms: 30 * DAY_MS, bucketMs: DAY_MS }),
100
+ });
101
+
102
+ const isObject = (v) => v !== null && typeof v === "object" && !Array.isArray(v);
103
+ const isCount = (v) => Number.isSafeInteger(v) && v >= 0;
104
+ /** A host name a worker can have written (config.mjs `WORKER_NAME_RE`, the rule `PI_WORKER_NAME` is held to). */
105
+ const isHostName = (v) => typeof v === "string" && WORKER_NAME_RE.test(v);
106
+ const UNKNOWN_CAPACITY = Object.freeze({ slots: null, memMiB: null, cpuCenti: null, cpus: null });
107
+
108
+ /**
109
+ * Whether a record held a slot: `{ occupied, inferred }`, or null when it cannot be told (a `capacity` that is neither
110
+ * an object nor null). `inferred` is true for a record from before `capacity` existed.
111
+ */
112
+ export function occupancyOf(record) {
113
+ if (!isObject(record)) return null;
114
+ if ("capacity" in record) {
115
+ if (record.capacity === null) return { occupied: false, inferred: false };
116
+ return isObject(record.capacity) ? { occupied: true, inferred: false } : null;
117
+ }
118
+ if (PRE_SLOT_REFUSAL_REASONS.includes(record.reason)) return { occupied: false, inferred: true };
119
+ const start = Date.parse(record.startedAt ?? "");
120
+ const end = Date.parse(record.endedAt ?? "");
121
+ const wall = Number.isFinite(start) && Number.isFinite(end) ? end - start : NaN;
122
+ return { occupied: wall >= LEGACY_MIN_WALL_MS || isObject(record.resources), inferred: true };
123
+ }
124
+
125
+ /**
126
+ * Whether a `capacity` says more than any host can have (job-size.mjs `HOST_*_MAX`): more slots, CPUs, memory or CPU
127
+ * budget. Such a field is read as unknown, exactly as the writer now records it, and the run still counts with its span,
128
+ * size and CPU: a worker from before the bound wrote a large `PI_CONCURRENCY` as it was, and dropping its runs would read
129
+ * a busy host as idle. Each such record is counted (`capacityOutOfRange`).
130
+ */
131
+ function beyondAnyHost(c) {
132
+ const over = (v, max) => typeof v === "number" && v > max;
133
+ return over(c.slots, HOST_SLOTS_MAX) || over(c.cpus, HOST_CPUS_MAX) || over(c.memMiB, HOST_MEMORY_MAX_MIB) || over(c.cpuCenti, HOST_CPUS_MAX * 100);
134
+ }
135
+
136
+ /**
137
+ * A record's `capacity` as this report reads it, the writer's rule (run-history.mjs `recordedCapacity`) restated (see
138
+ * the header): `{ slots, memMiB, cpuCenti, cpus }`, each a positive integer (a budget may be 0), `"off"` for a budget
139
+ * switched off, or null; null for anything that is not an object.
140
+ */
141
+ export function capacityOf(value) {
142
+ if (!isObject(value)) return null;
143
+ const positive = (max) => (v) => (Number.isSafeInteger(v) && v >= 1 && v <= max ? v : null);
144
+ const budget = (max) => (v) => (v === "off" ? "off" : isCount(v) && v <= max ? v : null);
145
+ return { slots: positive(HOST_SLOTS_MAX)(value.slots), memMiB: budget(HOST_MEMORY_MAX_MIB)(value.memMiB), cpuCenti: budget(HOST_CPUS_MAX * 100)(value.cpuCenti), cpus: positive(HOST_CPUS_MAX)(value.cpus) };
146
+ }
147
+
148
+ /** A live host row's capacity (host-registry.mjs strings): the slot count and both budgets; the CPU count is not published. */
149
+ function liveCapacityOf(row) {
150
+ const int = (s) => (typeof s === "string" && /^\d{1,15}$/.test(s) ? Number(s) : null);
151
+ const budget = (s, max) => {
152
+ const n = s === "off" ? "off" : int(s);
153
+ return typeof n === "number" && n > max ? null : n;
154
+ };
155
+ const slots = int(row?.concurrency);
156
+ return { slots: slots !== null && slots >= 1 && slots <= HOST_SLOTS_MAX ? slots : null, memMiB: budget(row?.budgetMemMiB, HOST_MEMORY_MAX_MIB), cpuCenti: budget(row?.budgetCpuCenti, HOST_CPUS_MAX * 100), cpus: null };
157
+ }
158
+
159
+ const sameCapacity = (a, b) => a.slots === b.slots && a.memMiB === b.memMiB && a.cpuCenti === b.cpuCenti && a.cpus === b.cpus;
160
+ const positiveInt = (v) => Number.isSafeInteger(v) && v > 0;
161
+ /** What CPU promises are judged against: the CPU budget, or every CPU of the host where the budget is off or unknown. */
162
+ const promiseCpuCenti = (c) => (positiveInt(c.cpuCenti) ? c.cpuCenti : positiveInt(c.cpus) ? c.cpus * 100 : null);
163
+ /**
164
+ * The most CPU one job can use, in hundredths: its `--cpus` is the host's CPU budget (capped at the runtime's count),
165
+ * so the smaller of the two that are known, else the size ceiling.
166
+ */
167
+ const jobCpuCeilingCenti = (c) => {
168
+ const known = [positiveInt(c.cpuCenti) ? c.cpuCenti : null, positiveInt(c.cpus) ? c.cpus * 100 : null].filter((v) => v !== null);
169
+ return known.length > 0 ? Math.min(...known) : JOB_CPUS_CEILING_CENTI;
170
+ };
171
+
172
+ /**
173
+ * The most CPU this host's jobs could use TOGETHER, in hundredths: the host's CPUs, which is physically true; null when
174
+ * not known. Not the CPU budget: the parent cgroup quota that holds the jobs to it (`cpu-reserve.mjs`) is not set
175
+ * where only root can set it, so the budget is not a proven ceiling.
176
+ */
177
+ const segmentCpuCeilingCenti = (c) => (positiveInt(c.cpus) ? c.cpus * 100 : null);
178
+
179
+ /** A record's span `{ start, end, wallMs }` in millis, or null when it is not one this report counts (see the header). */
180
+ function spanOf(record, nowMs) {
181
+ const start = Date.parse(record.startedAt ?? "");
182
+ const end = Date.parse(record.endedAt ?? "");
183
+ if (!Number.isFinite(start) || !Number.isFinite(end) || end < start) return null;
184
+ if (end - start > SUGGEST_MAX_WALL_MS || end > nowMs + SUGGEST_CLOCK_SKEW_MS) return null;
185
+ return { start, end, wallMs: end - start };
186
+ }
187
+
188
+ /** num / den in thousandths, rounded half up, as a Number; null when the denominator is 0. Both BigInt. */
189
+ function perMille(num, den) {
190
+ if (den <= 0n) return null;
191
+ return Number((num * 2000n + den) / (den * 2n));
192
+ }
193
+
194
+ /** The nearest-rank percentile of a sorted array: the value at rank ceil(p/100 x n). size-suggest's rule. */
195
+ const rank = (sorted, pct) => sorted[Math.max(0, Math.ceil((pct * sorted.length) / 100) - 1)];
196
+
197
+ const zeroCounts = () => ({ used: 0, capacityOutOfRange: 0, refusedBeforeSlot: 0, legacyOccupied: 0, legacyRefused: 0, withoutSize: 0, withoutResources: 0, cpuClamped: 0, retried: 0, earlier: 0, stalledRepick: 0, live: 0, liveNotCounted: 0, liveUnreadable: 0, orphans: 0 });
198
+
199
+ /**
200
+ * A live row's running jobs: `{ jobs, notListed, unreadable }`, `jobs` the listed entries through live-jobs.mjs'
201
+ * allowlist (a row from `readLiveHosts` is parsed already; a raw string is parsed here, the same rule) or null when the
202
+ * row lists none, `notListed` the running jobs it reports and does not list (its `jobsMore`, or a worker from before
203
+ * `jobs`: its budget's `budgetRunning`), null when it says nothing. `unreadable` is a row that HAS a `jobs` value that
204
+ * is not a list (`readLiveHosts` flags it `jobsUnreadable`): how many it runs is then unknown, and nothing stands in.
205
+ */
206
+ function rowJobsOf(row) {
207
+ const parsed = parseLiveJobs(row?.jobs);
208
+ if (parsed.jobs === null) {
209
+ if (row?.jobsUnreadable === true || (row?.jobs !== undefined && row?.jobs !== null && row?.jobs !== "")) return { jobs: null, notListed: null, unreadable: true };
210
+ const n = typeof row?.budgetRunning === "string" && /^\d{1,9}$/.test(row.budgetRunning) ? Number(row.budgetRunning) : null;
211
+ return { jobs: null, notListed: n, unreadable: false };
212
+ }
213
+ return { jobs: parsed.jobs, notListed: (parseJobsMore(row.jobsMore) ?? 0) + parsed.dropped, unreadable: false };
214
+ }
215
+
216
+ /**
217
+ * Where each host's history starts: `Map<name, { fromMs, source, truncated } | null>`, null for a host no source holds.
218
+ * With no source described at all (`coverage.local` and `coverage.mirror` both undefined: a caller with records and
219
+ * nothing else), every host is covered from `coverage.fromMs` (else the window start).
220
+ */
221
+ function hostCoverage(names, liveByName, coverage, windowStartMs, nowMs) {
222
+ const clamp = (ms) => Math.min(nowMs, Math.max(windowStartMs, Number.isSafeInteger(ms) ? ms : windowStartMs));
223
+ const out = new Map();
224
+ if (coverage.local === undefined && coverage.mirror === undefined) {
225
+ for (const name of names) out.set(name, { fromMs: clamp(coverage.fromMs), source: typeof coverage.source === "string" ? coverage.source : "records", truncated: coverage.truncated === true });
226
+ return out;
227
+ }
228
+ const localHosts = new Set(Array.isArray(coverage.localHosts) ? coverage.localHosts : []);
229
+ if (isHostName(coverage.localHost)) localHosts.add(coverage.localHost);
230
+ const mirrorHosts = new Set(Array.isArray(coverage.mirror?.hosts) ? coverage.mirror.hosts : []);
231
+ for (const name of names) {
232
+ const options = [];
233
+ if (isObject(coverage.local) && localHosts.has(name)) options.push({ fromMs: clamp(coverage.local.fromMs), source: "local", truncated: false });
234
+ const row = liveByName.get(name);
235
+ // A named host's runs are all in the mirror (its row says it routes), and so are those of a host with no live row
236
+ // whose runs the mirror holds; a live row that does NOT route is a worker without `PI_WORKER_NAME`, which mirrors none.
237
+ if (isObject(coverage.mirror) && (row ? row.routes === "true" : mirrorHosts.has(name))) options.push({ fromMs: clamp(coverage.mirror.fromMs), source: "mirror", truncated: coverage.mirror.truncated === true });
238
+ // The earliest start wins, and on a tie the files (never cut by the mirror's cap). A host BOTH sources cover says so
239
+ // (`mirror+local`): its records came from both, and naming one read a host the mirror also holds as files only.
240
+ options.sort((a, b) => a.fromMs - b.fromMs || Number(a.truncated) - Number(b.truncated));
241
+ const first = options[0] ?? null;
242
+ out.set(name, first && options.length > 1 ? { ...first, source: "mirror+local" } : first);
243
+ }
244
+ return out;
245
+ }
246
+
247
+ /**
248
+ * The capacity report (`INT-CAPACITY-REPORT`): `{ v, window, coverage, hosts }`.
249
+ *
250
+ * - `records`: run records, already merged (one per job id).
251
+ * - `live`: the live host rows (`readLiveHosts`), `[]` when unknown. Each gives a host its CURRENT capacity where no
252
+ * record in the window recorded one, and says whether it mirrors its runs (`routes`).
253
+ * - `windowStartMs`, `nowMs`: the window, integers; `bucketMs` splits it from its start (the last bucket ends at now).
254
+ * - `coverage`: what the reader could see (capacity-records.mjs): `{ source, reason, localHost, localHosts, local:
255
+ * { fromMs } | null, mirror: { fromMs, truncated, hosts } | null }`, `local` null when this host's files could not be
256
+ * read and `mirror` null when the mirror was not read. Without either key, every host is covered from
257
+ * `coverage.fromMs`.
258
+ *
259
+ * Throws a RangeError on a window or bucket this function cannot honour: those are the caller's mistake, not data.
260
+ */
261
+ export function computeCapacity({ records = [], live = [], windowStartMs, nowMs, bucketMs, coverage = {} } = {}) {
262
+ if (!Number.isSafeInteger(windowStartMs) || !Number.isSafeInteger(nowMs) || windowStartMs >= nowMs) throw new RangeError("the window must be two integer instants, start before now");
263
+ if (!Number.isSafeInteger(bucketMs) || bucketMs <= 0) throw new RangeError("bucketMs must be a positive integer");
264
+ const windowMs = nowMs - windowStartMs;
265
+ const bucketCount = Math.ceil(windowMs / bucketMs);
266
+ if (bucketCount > CAPACITY_MAX_BUCKETS) throw new RangeError(`at most ${CAPACITY_MAX_BUCKETS} buckets (got ${bucketCount})`);
267
+
268
+ const liveByName = new Map((Array.isArray(live) ? live : []).filter((row) => isHostName(row?.name)).map((row) => [row.name, row]));
269
+ let unreadable = 0;
270
+ let withoutHost = 0;
271
+ let earlierDropped = 0;
272
+
273
+ // PASS 1: every record judged once, so the hosts (and so each host's coverage) are known before anything is counted.
274
+ const judged = [];
275
+ const earlierRuns = [];
276
+ for (const record of Array.isArray(records) ? records : []) {
277
+ const occupancy = occupancyOf(record);
278
+ const project = record?.project ?? null;
279
+ if (occupancy === null || (project !== null && !isProjectId(project))) {
280
+ unreadable++;
281
+ continue;
282
+ }
283
+ if (record.host === null || record.host === undefined) {
284
+ withoutHost++;
285
+ continue;
286
+ }
287
+ if (!isHostName(record.host)) {
288
+ unreadable++;
289
+ continue;
290
+ }
291
+ judged.push({ record, occupancy, project, span: spanOf(record, nowMs) });
292
+ // The slot time of the job's EARLIER attempts, which this record carries because it replaced theirs: each an
293
+ // occupied interval on its own host, with its size and nothing else known (no wait, no CPU measurement).
294
+ // Rebuilt by the writer's own rule (`recordedEarlier`: valid entries, the newest EARLIER_MAX), and an entry that
295
+ // overlaps this record's own span on its own host dropped: one job's attempts cannot hold one host's slot twice at
296
+ // once. Every entry not kept is counted.
297
+ const given = Array.isArray(record.earlier) ? record.earlier.length : 0;
298
+ const rebuilt = recordedEarlier(record.earlier) ?? [];
299
+ const own = spanOf(record, nowMs);
300
+ const kept = rebuilt.filter((e) => !(own !== null && e.host === record.host && Date.parse(e.startedAt) < own.end && Date.parse(e.endedAt) > own.start));
301
+ earlierDropped += given - kept.length;
302
+ for (const e of kept) {
303
+ const eSpan = spanOf(e, nowMs);
304
+ if (eSpan === null) {
305
+ earlierDropped++;
306
+ continue;
307
+ }
308
+ const eSize = positiveInt(e.memMiB) && positiveInt(e.cpuCenti) ? { memMiB: e.memMiB, cpuCenti: e.cpuCenti } : null;
309
+ earlierRuns.push({ host: e.host, span: eSpan, size: eSize, project });
310
+ }
311
+ }
312
+ // The newest admission among each job id's records, so a job the live row still lists after its record was written
313
+ // (it ended between the beat and this read) is counted once, as the record.
314
+ const recordedFrom = new Map();
315
+ for (const { record, span } of judged) {
316
+ // Keyed by the id AS A ROW PUBLISHES IT (`publishedJobId`): an id outside the row's charset is its digest there.
317
+ const id = typeof record.jobId === "string" ? publishedJobId(record.jobId) : null;
318
+ if (id !== null && span !== null) recordedFrom.set(id, Math.max(recordedFrom.get(id) ?? -Infinity, span.start));
319
+ }
320
+ const names = new Set([...liveByName.keys()]);
321
+ for (const j of judged) names.add(j.record.host);
322
+ for (const e of earlierRuns) names.add(e.host);
323
+ const covers = hostCoverage(names, liveByName, coverage, windowStartMs, nowMs);
324
+
325
+ const byHost = new Map();
326
+ const hostOf = (name) => {
327
+ if (!byHost.has(name)) byHost.set(name, { runs: [], recorded: [], waits: [], counts: zeroCounts() });
328
+ return byHost.get(name);
329
+ };
330
+ for (const name of names) if (liveByName.has(name)) hostOf(name);
331
+
332
+ // PASS 2: count each record against its host's own coverage.
333
+ for (const { record, occupancy, project, span } of judged) {
334
+ const cover = covers.get(record.host);
335
+ if (cover === null) continue; // that host's history is missing as a whole, not partly
336
+ const coveredFrom = cover.fromMs;
337
+ if (!occupancy.occupied) {
338
+ // A refusal is counted where it ENDED inside the covered window, the instant it happened; it took no slot time.
339
+ const end = Date.parse(record.endedAt ?? "");
340
+ if (Number.isFinite(end) && end >= coveredFrom && end <= nowMs + SUGGEST_CLOCK_SKEW_MS) {
341
+ const counts = hostOf(record.host).counts;
342
+ counts.refusedBeforeSlot++;
343
+ if (occupancy.inferred) counts.legacyRefused++;
344
+ }
345
+ continue;
346
+ }
347
+ if (span === null) {
348
+ unreadable++;
349
+ continue;
350
+ }
351
+ const start = Math.max(span.start, coveredFrom);
352
+ const end = Math.min(span.end, nowMs);
353
+ // Outside the covered window. A run of no length counts where it started (it held a slot for no time, and its wait
354
+ // is a wait like any other); one that ended exactly as the window began spent none of it.
355
+ if (end < start || (end === start && (span.start < coveredFrom || span.start > nowMs))) continue;
356
+ const host = hostOf(record.host);
357
+ host.counts.used++;
358
+ if (occupancy.inferred) host.counts.legacyOccupied++;
359
+ const size = recordedJobSize(record.size);
360
+ if (size === null) host.counts.withoutSize++;
361
+ const retry = Number.isInteger(record.attempt) && record.attempt > 1;
362
+ if (retry) host.counts.retried++;
363
+ // A pickup after a stall: the stalled pickup wrote no record (a crash or a lost lock), so its time is unknown.
364
+ const repick = record.stalledRepick === true;
365
+ if (repick) host.counts.stalledRepick++;
366
+ if (isObject(record.capacity) && beyondAnyHost(record.capacity)) host.counts.capacityOutOfRange++;
367
+ const recorded = capacityOf(record.capacity);
368
+ if (recorded !== null) host.recorded.push({ at: span.start, capacity: recorded });
369
+ const raw = record.resources?.cpuUsec;
370
+ host.runs.push({ start, end, span, size, rawCpuUsec: isCount(raw) && span.wallMs > 0 ? raw : null, recorded, project });
371
+ // The wait of a FIRST attempt that started inside the covered window, non-negative only: a retry's `queuedAt` is
372
+ // its first add's, so its wait would include the earlier attempt; a clock between two hosts can put a queuedAt
373
+ // after its own start, and that run says nothing about waiting.
374
+ const queued = typeof record.queuedAt === "string" ? Date.parse(record.queuedAt) : NaN;
375
+ if (!retry && !repick && Number.isFinite(queued) && span.start >= coveredFrom && span.start <= nowMs && span.start - queued >= 0) host.waits.push(span.start - queued);
376
+ }
377
+
378
+ for (const e of earlierRuns) {
379
+ const cover = covers.get(e.host);
380
+ if (cover === null) continue;
381
+ const start = Math.max(e.span.start, cover.fromMs);
382
+ const end = Math.min(e.span.end, nowMs);
383
+ if (end < start || (end === start && (e.span.start < cover.fromMs || e.span.start > nowMs))) continue;
384
+ const host = hostOf(e.host);
385
+ host.counts.earlier++;
386
+ host.runs.push({ start, end, span: e.span, size: e.size, rawCpuUsec: null, recorded: null, project: e.project, earlier: true });
387
+ }
388
+
389
+ // THE JOBS RUNNING NOW (see the header), from each live row, each an occupied interval to now (or to a stale row's
390
+ // last beat), clipped to the host's covered window like a record.
391
+ // `running` is what the rows say runs now: the jobs counted plus those not counted, never a listed job whose record
392
+ // already counts it (it has ended); null when no row says anything, or when one row's list could not be read (how
393
+ // many run on that host is unknown, so the fleet's number is too).
394
+ let rowsSay = false;
395
+ for (const [name, row] of liveByName) {
396
+ const { jobs, notListed, unreadable } = rowJobsOf(row);
397
+ const host = hostOf(name);
398
+ const cover = covers.get(name);
399
+ if (jobs !== null || notListed !== null) rowsSay = true;
400
+ if (unreadable) host.counts.liveUnreadable++;
401
+ host.counts.liveNotCounted += notListed ?? 0;
402
+ const staleMs = Number.isSafeInteger(row.staleMs) && row.staleMs >= 0 ? row.staleMs : null;
403
+ for (const j of jobs ?? []) {
404
+ if (j.o) {
405
+ host.counts.orphans++;
406
+ continue;
407
+ }
408
+ if ((recordedFrom.get(j.id) ?? -Infinity) >= j.at) continue; // its record counts it (keyed as the row publishes ids)
409
+ const end = staleMs === null ? null : staleMs <= LIVE_FRESH_MS ? nowMs : nowMs - staleMs;
410
+ const span = end === null ? null : spanOf({ startedAt: new Date(j.at).toISOString(), endedAt: new Date(end).toISOString() }, nowMs);
411
+ if (cover === null || cover === undefined || span === null) {
412
+ host.counts.liveNotCounted++;
413
+ continue;
414
+ }
415
+ const start = Math.max(span.start, cover.fromMs);
416
+ const stop = Math.min(span.end, nowMs);
417
+ if (stop < start) {
418
+ host.counts.liveNotCounted++;
419
+ continue;
420
+ }
421
+ host.counts.live++;
422
+ const size = positiveInt(j.m) && positiveInt(j.c) ? { memMiB: j.m, cpuCenti: j.c } : null;
423
+ host.runs.push({ start, end: stop, span, size, rawCpuUsec: null, recorded: null, project: j.p, live: true });
424
+ }
425
+ }
426
+
427
+ const hosts = [];
428
+ const totals = zeroCounts();
429
+ const notShared = [];
430
+ for (const name of [...byHost.keys()].sort()) {
431
+ const h = byHost.get(name);
432
+ const cover = covers.get(name) ?? null;
433
+ if (cover === null) notShared.push(name);
434
+ const coveredFrom = cover === null ? nowMs : cover.fromMs;
435
+ const coveredMs = nowMs - coveredFrom;
436
+ const missingMs = windowMs - coveredMs;
437
+
438
+ // THE CAPACITY IN FORCE, a step function from the runs' own records ordered by start; before the first, the first.
439
+ // With none recorded, the live row's (the current setting), else unknown.
440
+ const steps = [...h.recorded].sort((a, b) => a.at - b.at);
441
+ const row = liveByName.get(name);
442
+ const fallback = row ? liveCapacityOf(row) : UNKNOWN_CAPACITY;
443
+ const capAt = (t) => {
444
+ if (steps.length === 0) return fallback;
445
+ let found = steps[0].capacity;
446
+ for (const s of steps) {
447
+ if (s.at > t) break;
448
+ found = s.capacity;
449
+ }
450
+ return found;
451
+ };
452
+ // The instants inside the covered window where the capacity in force changes.
453
+ const changes = [];
454
+ for (let i = 1; i < steps.length; i++) {
455
+ if (steps[i].at > coveredFrom && steps[i].at < nowMs && !sameCapacity(steps[i].capacity, steps[i - 1].capacity)) changes.push(steps[i].at);
456
+ }
457
+ const newest = steps.at(-1) ?? null;
458
+ const base = newest !== null ? newest.capacity : fallback;
459
+ const basis = newest !== null ? "recorded" : row ? "current" : "unknown";
460
+ const changed = steps.some((s) => !sameCapacity(s.capacity, base));
461
+
462
+ const buckets = [];
463
+ for (let i = 0; i < bucketCount; i++) {
464
+ const bStart = windowStartMs + i * bucketMs;
465
+ const bEnd = Math.min(nowMs, bStart + bucketMs);
466
+ buckets.push({ fromMs: bStart, coveredMs: Math.max(0, bEnd - Math.max(bStart, coveredFrom)), busyMs: 0, fullMs: 0, fullKnown: false, peak: 0, runMs: 0n });
467
+ }
468
+ // Calls fn(bucket, a, b) for each piece of [a, b) cut at the bucket edges.
469
+ const eachBucket = (a, b, fn) => {
470
+ for (let i = Math.floor((a - windowStartMs) / bucketMs); i < bucketCount; i++) {
471
+ const bStart = windowStartMs + i * bucketMs;
472
+ if (bStart >= b) break;
473
+ const lo = Math.max(a, bStart);
474
+ const hi = Math.min(b, bStart + bucketMs, nowMs);
475
+ if (hi > lo) fn(buckets[i], lo, hi);
476
+ }
477
+ };
478
+
479
+ // THE SWEEP: every run's start and end, and every change of capacity, in time order (ends, then changes, then
480
+ // starts, at one instant). Only a segment of positive length is counted, so a run that ends exactly when the next
481
+ // begins is one slot reused, never two at once, whatever the order of the two events; sorting the end first keeps
482
+ // the level itself true at that instant too. A segment's slot count is the one in force when it begins.
483
+ // Each run's own CPU time first, clamped to what its job could use: its `--cpus` (the CPU budget, capped at the
484
+ // runtime's count) over its whole wall. A run with no measurement adds nothing to CPU used.
485
+ const clamped = new Set();
486
+ h.runs.forEach((r, i) => {
487
+ r.used = 0n;
488
+ r.cpuUsec = null;
489
+ if (r.rawCpuUsec === null) {
490
+ if (!r.earlier && !r.live) h.counts.withoutResources++;
491
+ return;
492
+ }
493
+ const most = BigInt(jobCpuCeilingCenti(r.recorded ?? capAt(r.span.start))) * BigInt(r.span.wallMs) * 10n; // hundredths x ms x 10 = microseconds
494
+ const reported = BigInt(r.rawCpuUsec);
495
+ if (reported > most) clamped.add(i);
496
+ r.cpuUsec = reported > most ? most : reported;
497
+ });
498
+ const events = [];
499
+ h.runs.forEach((r, i) => events.push([r.start, 1, i], [r.end, -1, i]));
500
+ for (const t of changes) events.push([t, 0, -1]);
501
+ events.sort((a, b) => a[0] - b[0] || a[1] - b[1]);
502
+ let level = 0;
503
+ let at = coveredFrom;
504
+ let busyMs = 0;
505
+ let fullMs = 0;
506
+ let fullKnown = false;
507
+ let peak = 0;
508
+ let runMs = 0n;
509
+ let cpuUsedNum = 0n;
510
+ const active = new Set();
511
+ const segment = (a, b, n) => {
512
+ if (b <= a) return;
513
+ const cap = capAt(a);
514
+ // CPU used in this segment: each running job's measured time spread over its wall, together never more than
515
+ // the host's CPUs. Past it, every share is cut in proportion and the runs are counted as clamped.
516
+ const parts = [];
517
+ let sum = 0n;
518
+ for (const i of active) {
519
+ const r = h.runs[i];
520
+ if (r.cpuUsec === null || r.span.wallMs <= 0) continue;
521
+ const part = (r.cpuUsec * BigInt(b - a)) / BigInt(r.span.wallMs);
522
+ parts.push([i, part]);
523
+ sum += part;
524
+ }
525
+ const ceiling = segmentCpuCeilingCenti(cap);
526
+ const most = ceiling === null ? null : BigInt(ceiling) * BigInt(b - a) * 10n;
527
+ const cut = most !== null && sum > most;
528
+ for (const [i, part] of parts) {
529
+ const kept = cut ? (part * most) / sum : part;
530
+ h.runs[i].used += kept;
531
+ if (positiveInt(cap.cpus)) cpuUsedNum += kept;
532
+ if (cut) clamped.add(i);
533
+ }
534
+ const slots = cap.slots;
535
+ if (slots !== null) fullKnown = true;
536
+ const full = slots !== null && n >= slots;
537
+ if (n > 0) busyMs += b - a;
538
+ if (full) fullMs += b - a;
539
+ if (n > peak) peak = n;
540
+ runMs += BigInt(n) * BigInt(b - a);
541
+ eachBucket(a, b, (bucket, lo, hi) => {
542
+ if (n > 0) bucket.busyMs += hi - lo;
543
+ if (slots !== null) bucket.fullKnown = true;
544
+ if (full) bucket.fullMs += hi - lo;
545
+ if (n > bucket.peak) bucket.peak = n;
546
+ bucket.runMs += BigInt(n) * BigInt(hi - lo);
547
+ });
548
+ };
549
+ for (const [t, delta, i] of events) {
550
+ segment(at, t, level);
551
+ at = Math.max(at, t);
552
+ level += delta;
553
+ if (delta === 1) active.add(i);
554
+ else if (delta === -1) active.delete(i);
555
+ }
556
+ segment(at, nowMs, level);
557
+ h.counts.cpuClamped += clamped.size;
558
+
559
+ // What was promised (CPU used is counted in the sweep above), against the capacity IN FORCE piece by piece (the
560
+ // covered window cut at every change). A piece whose budget (or CPU count) is not a number adds to neither side of
561
+ // its share.
562
+ const pieces = [coveredFrom, ...changes, nowMs];
563
+ const forEachPiece = (a, b, fn) => {
564
+ for (let i = 0; i + 1 < pieces.length; i++) {
565
+ const lo = Math.max(a, pieces[i]);
566
+ const hi = Math.min(b, pieces[i + 1]);
567
+ if (hi > lo) fn(capAt(pieces[i]), BigInt(hi - lo));
568
+ }
569
+ };
570
+ let memDen = 0n;
571
+ let cpuPromiseDen = 0n;
572
+ let cpuUsedDen = 0n;
573
+ forEachPiece(coveredFrom, nowMs, (c, len) => {
574
+ if (positiveInt(c.memMiB)) memDen += BigInt(c.memMiB) * len;
575
+ const promise = promiseCpuCenti(c);
576
+ if (promise !== null) cpuPromiseDen += BigInt(promise) * len;
577
+ // microseconds of CPU the host had: CPUs x ms x 1000
578
+ if (positiveInt(c.cpus)) cpuUsedDen += BigInt(c.cpus) * len * 1000n;
579
+ });
580
+ let memNum = 0n;
581
+ let cpuPromiseNum = 0n;
582
+ const projects = new Map();
583
+ for (const r of h.runs) {
584
+ forEachPiece(r.start, r.end, (c, len) => {
585
+ if (r.size !== null && positiveInt(c.memMiB)) memNum += BigInt(r.size.memMiB) * len;
586
+ if (r.size !== null && promiseCpuCenti(c) !== null) cpuPromiseNum += BigInt(r.size.cpuCenti) * len;
587
+ });
588
+ const p = projects.get(r.project) ?? { runMs: 0n, cpuUsec: 0n };
589
+ p.runMs += BigInt(r.end - r.start);
590
+ p.cpuUsec += r.used;
591
+ projects.set(r.project, p);
592
+ }
593
+ const ranked = [...projects.entries()].sort((a, b) => (b[1].runMs > a[1].runMs ? 1 : b[1].runMs < a[1].runMs ? -1 : String(a[0] ?? "").localeCompare(String(b[0] ?? ""))));
594
+ const toProject = ([project, p]) => ({ project, runMs: Number(p.runMs), cpuMs: Number(p.cpuUsec / 1000n) });
595
+ const rest = ranked.slice(CAPACITY_TOP_PROJECTS);
596
+ const other = rest.length === 0 ? null : { count: rest.length, runMs: Number(rest.reduce((s, [, p]) => s + p.runMs, 0n)), cpuMs: Number(rest.reduce((s, [, p]) => s + p.cpuUsec, 0n) / 1000n) };
597
+
598
+ for (const [k, v] of Object.entries(h.counts)) totals[k] += v;
599
+ const waits = [...h.waits].sort((a, b) => a - b);
600
+ hosts.push({
601
+ name,
602
+ shared: cover !== null,
603
+ // Why a host's history is not here: `unread` when its live row routes (it writes the run mirror, so a host
604
+ // not covered means the mirror was not read: no Valkey, or it did not answer, and the coverage's `reason`
605
+ // says which), `unnamed` when it does not (a worker without PI_WORKER_NAME writes no run mirror); null when
606
+ // it is shared. Only a host with a live row can be uncovered: a run on an uncovered host is not counted.
607
+ notShared: cover !== null ? null : liveByName.get(name)?.routes === "true" ? "unread" : "unnamed",
608
+ coverage: { fromMs: coveredFrom, source: cover?.source ?? null, truncated: cover?.truncated === true && coveredFrom > windowStartMs, ...h.counts },
609
+ capacity: { ...base, basis, changed },
610
+ coveredMs,
611
+ missingMs,
612
+ busyMs,
613
+ idleMs: coveredMs - busyMs,
614
+ fullMs: fullKnown ? fullMs : null,
615
+ peak,
616
+ avgMilli: perMille(runMs, BigInt(coveredMs)),
617
+ promisedMemPerMille: perMille(memNum, memDen),
618
+ promisedCpuPerMille: perMille(cpuPromiseNum, cpuPromiseDen),
619
+ usedCpuPerMille: perMille(cpuUsedNum, cpuUsedDen),
620
+ runs: h.runs.length,
621
+ projects: ranked.slice(0, CAPACITY_TOP_PROJECTS).map(toProject),
622
+ otherProjects: other,
623
+ waits: { n: waits.length, p50Ms: waits.length > 0 ? rank(waits, 50) : null, p95Ms: waits.length > 0 ? rank(waits, 95) : null },
624
+ buckets: buckets.map((b) => ({ fromMs: b.fromMs, coveredMs: b.coveredMs, busyMs: b.busyMs, fullMs: b.fullKnown ? b.fullMs : null, peak: b.peak, avgMilli: perMille(b.runMs, BigInt(b.coveredMs)) })),
625
+ });
626
+ }
627
+
628
+ const covered = hosts.filter((h) => h.shared);
629
+ // The reading host has runs here and no live row (the registry was not read, or its row is gone until its next beat):
630
+ // the jobs it runs now are not counted, and the coverage says so; `running` stays what the other rows say.
631
+ const mirrorCut = isObject(coverage.mirror) && coverage.mirror.truncated === true && Math.min(nowMs, Number.isSafeInteger(coverage.mirror.fromMs) ? coverage.mirror.fromMs : windowStartMs) > windowStartMs;
632
+ const liveRowMissing = isHostName(coverage.localHost) && names.has(coverage.localHost) && !liveByName.has(coverage.localHost) ? coverage.localHost : null;
633
+ return {
634
+ v: CAPACITY_REPORT_VERSION,
635
+ window: { fromMs: windowStartMs, toMs: nowMs, bucketMs },
636
+ coverage: {
637
+ source: typeof coverage.source === "string" ? coverage.source : "records",
638
+ reason: typeof coverage.reason === "string" ? coverage.reason : null,
639
+ // Every covered host has history from here on (the latest of their starts); each host's own is in its entry.
640
+ fromMs: covered.length > 0 ? Math.max(...covered.map((h) => h.coverage.fromMs)) : windowStartMs,
641
+ // Some host's history was cut; with no covered host, whether the run mirror's own history was (`mirrorCut`), so
642
+ // an empty report never reads as a whole one.
643
+ truncated: covered.length > 0 ? covered.some((h) => h.coverage.truncated) : mirrorCut,
644
+ ...totals,
645
+ unreadable,
646
+ withoutHost,
647
+ earlierDropped,
648
+ running: rowsSay && totals.liveUnreadable === 0 ? totals.live + totals.liveNotCounted : null,
649
+ liveRowMissing,
650
+ historyNotShared: notShared,
651
+ },
652
+ hosts,
653
+ };
654
+ }