sparkforensics-cli 0.2.0 → 0.2.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (107) hide show
  1. package/export-template/docs/404.html +1 -1
  2. package/export-template/docs/assets/{app.CndaAS6v.js → app.DQTZyGL1.js} +1 -1
  3. package/export-template/docs/assets/chunks/@localSearchIndexroot.DppnXnDE.js +1 -0
  4. package/export-template/docs/assets/chunks/{VPLocalSearchBox.yJbZbsEo.js → VPLocalSearchBox.BkBIPFs6.js} +1 -1
  5. package/export-template/docs/assets/chunks/theme.DP0u1AUq.js +2 -0
  6. package/export-template/docs/assets/contributor-guide_architecture_board-widgets.md.CWpj01WU.js +1 -0
  7. package/export-template/docs/assets/contributor-guide_architecture_board-widgets.md.CWpj01WU.lean.js +1 -0
  8. package/export-template/docs/assets/contributor-guide_architecture_detector-contract.md.CgzUsQ6W.js +1 -0
  9. package/export-template/docs/assets/{contributor-guide_architecture_impact-estimation.md.DYCDPgkh.js → contributor-guide_architecture_impact-estimation.md.CooslVJt.js} +1 -1
  10. package/export-template/docs/assets/contributor-guide_architecture_state-and-history.md.C-xxn0q7.js +1 -0
  11. package/export-template/docs/assets/contributor-guide_architecture_state-and-history.md.C-xxn0q7.lean.js +1 -0
  12. package/export-template/docs/assets/contributor-guide_architecture_widget-rendering.md.R27gQrgY.js +1 -0
  13. package/export-template/docs/assets/contributor-guide_architecture_worker-protocol.md.IbnfNrV3.js +6 -0
  14. package/export-template/docs/assets/{style.DXOMCXxn.css → style.DSixAiZE.css} +1 -1
  15. package/export-template/docs/assets/{user-guide_alternative-log-retrieval.md.sU3KGarf.js → user-guide_alternative-log-retrieval.md.B4tPGIal.js} +1 -1
  16. package/export-template/docs/assets/{user-guide_alternative-log-retrieval.md.sU3KGarf.lean.js → user-guide_alternative-log-retrieval.md.B4tPGIal.lean.js} +1 -1
  17. package/export-template/docs/assets/user-guide_getting-started.md.BJvwLEIM.js +3 -0
  18. package/export-template/docs/assets/user-guide_getting-started.md.BJvwLEIM.lean.js +1 -0
  19. package/export-template/docs/assets/{user-guide_mcp-tools.md.C8MiIu7F.js → user-guide_mcp-tools.md.Vi3RoflJ.js} +3 -3
  20. package/export-template/docs/assets/{user-guide_mcp-tools.md.C8MiIu7F.lean.js → user-guide_mcp-tools.md.Vi3RoflJ.lean.js} +1 -1
  21. package/export-template/docs/assets/user-guide_run-comparison.md.CQc1aoU8.js +1 -0
  22. package/export-template/docs/assets/user-guide_run-comparison.md.CQc1aoU8.lean.js +1 -0
  23. package/export-template/docs/assets/user-guide_understanding-findings.md.DL1UDhvR.js +1 -0
  24. package/export-template/docs/assets/user-guide_understanding-findings.md.DL1UDhvR.lean.js +1 -0
  25. package/export-template/docs/contributor-guide/architecture/board-widgets.html +2 -2
  26. package/export-template/docs/contributor-guide/architecture/detector-contract.html +2 -2
  27. package/export-template/docs/contributor-guide/architecture/drill-down.html +1 -1
  28. package/export-template/docs/contributor-guide/architecture/impact-estimation.html +2 -2
  29. package/export-template/docs/contributor-guide/architecture/index.html +1 -1
  30. package/export-template/docs/contributor-guide/architecture/overview.html +1 -1
  31. package/export-template/docs/contributor-guide/architecture/state-and-history.html +2 -2
  32. package/export-template/docs/contributor-guide/architecture/widget-rendering.html +2 -2
  33. package/export-template/docs/contributor-guide/architecture/worker-protocol.html +2 -2
  34. package/export-template/docs/contributor-guide/contributing.html +1 -1
  35. package/export-template/docs/contributor-guide/development-setup.html +1 -1
  36. package/export-template/docs/contributor-guide/testing.html +1 -1
  37. package/export-template/docs/index.html +1 -1
  38. package/export-template/docs/tuning-reference/anti-patterns.html +1 -1
  39. package/export-template/docs/tuning-reference/aqe.html +1 -1
  40. package/export-template/docs/tuning-reference/bottleneck-broadcast-sizing.html +1 -1
  41. package/export-template/docs/tuning-reference/bottleneck-cold-start.html +1 -1
  42. package/export-template/docs/tuning-reference/bottleneck-duplicate-plan-subtree.html +1 -1
  43. package/export-template/docs/tuning-reference/bottleneck-failures.html +1 -1
  44. package/export-template/docs/tuning-reference/bottleneck-gc.html +1 -1
  45. package/export-template/docs/tuning-reference/bottleneck-job-failure-rate.html +1 -1
  46. package/export-template/docs/tuning-reference/bottleneck-memory-utilization.html +1 -1
  47. package/export-template/docs/tuning-reference/bottleneck-retry-waste.html +1 -1
  48. package/export-template/docs/tuning-reference/bottleneck-shuffle.html +1 -1
  49. package/export-template/docs/tuning-reference/bottleneck-skew.html +1 -1
  50. package/export-template/docs/tuning-reference/bottleneck-slow-host.html +1 -1
  51. package/export-template/docs/tuning-reference/bottleneck-small-files.html +1 -1
  52. package/export-template/docs/tuning-reference/bottleneck-spill.html +1 -1
  53. package/export-template/docs/tuning-reference/bottleneck-straggler.html +1 -1
  54. package/export-template/docs/tuning-reference/bottleneck-tiny-tasks.html +1 -1
  55. package/export-template/docs/tuning-reference/bottleneck-utilization.html +1 -1
  56. package/export-template/docs/tuning-reference/caching.html +1 -1
  57. package/export-template/docs/tuning-reference/cluster-config.html +1 -1
  58. package/export-template/docs/tuning-reference/config.html +1 -1
  59. package/export-template/docs/tuning-reference/data-formats.html +1 -1
  60. package/export-template/docs/tuning-reference/index.html +1 -1
  61. package/export-template/docs/tuning-reference/intro.html +1 -1
  62. package/export-template/docs/tuning-reference/joins.html +1 -1
  63. package/export-template/docs/tuning-reference/memory-model.html +1 -1
  64. package/export-template/docs/tuning-reference/metrics.html +1 -1
  65. package/export-template/docs/tuning-reference/partitioning.html +1 -1
  66. package/export-template/docs/tuning-reference/pyspark.html +1 -1
  67. package/export-template/docs/tuning-reference/shuffle.html +1 -1
  68. package/export-template/docs/tuning-reference/spark-architecture.html +1 -1
  69. package/export-template/docs/tuning-reference/table-formats.html +1 -1
  70. package/export-template/docs/user-guide/alternative-log-retrieval.html +2 -2
  71. package/export-template/docs/user-guide/getting-started.html +3 -3
  72. package/export-template/docs/user-guide/mcp-tools.html +4 -4
  73. package/export-template/docs/user-guide/run-comparison.html +2 -2
  74. package/export-template/docs/user-guide/understanding-findings.html +2 -2
  75. package/export-template/index.html +75 -79
  76. package/export-template/parser-worker-DyjiQvfP.js +112 -0
  77. package/export-template/sample-runs/sample-run.ndjson.gz +0 -0
  78. package/package.json +5 -4
  79. package/vendor-core/detectors.js +149 -24
  80. package/vendor-core/docs-content/detection/cache.md +4 -3
  81. package/vendor-core/docs-content/detection/chrn.md +4 -2
  82. package/vendor-core/docs-content/detection/local.md +2 -3
  83. package/vendor-core/docs-content/detection/mem.md +3 -3
  84. package/vendor-core/docs-content/detection/spec.md +4 -3
  85. package/vendor-core/evidence-report.js +1 -1
  86. package/vendor-core/parser-worker.js +2 -2
  87. package/vendor-core/run-comparison.js +21 -2
  88. package/export-template/docs/assets/chunks/@localSearchIndexroot.DNY8bVcl.js +0 -1
  89. package/export-template/docs/assets/chunks/theme.Df2VAG9w.js +0 -2
  90. package/export-template/docs/assets/contributor-guide_architecture_board-widgets.md.B-OsL91z.js +0 -1
  91. package/export-template/docs/assets/contributor-guide_architecture_board-widgets.md.B-OsL91z.lean.js +0 -1
  92. package/export-template/docs/assets/contributor-guide_architecture_detector-contract.md.BOeH4d1J.js +0 -1
  93. package/export-template/docs/assets/contributor-guide_architecture_state-and-history.md.m3S3UdMk.js +0 -1
  94. package/export-template/docs/assets/contributor-guide_architecture_state-and-history.md.m3S3UdMk.lean.js +0 -1
  95. package/export-template/docs/assets/contributor-guide_architecture_widget-rendering.md.DbqPf2OT.js +0 -1
  96. package/export-template/docs/assets/contributor-guide_architecture_worker-protocol.md.B93qJ_tT.js +0 -6
  97. package/export-template/docs/assets/user-guide_getting-started.md.DtEM37MK.js +0 -3
  98. package/export-template/docs/assets/user-guide_getting-started.md.DtEM37MK.lean.js +0 -1
  99. package/export-template/docs/assets/user-guide_run-comparison.md.S0TWWmLY.js +0 -1
  100. package/export-template/docs/assets/user-guide_run-comparison.md.S0TWWmLY.lean.js +0 -1
  101. package/export-template/docs/assets/user-guide_understanding-findings.md.D0R_Y-R2.js +0 -1
  102. package/export-template/docs/assets/user-guide_understanding-findings.md.D0R_Y-R2.lean.js +0 -1
  103. package/export-template/parser-worker-QqyEE4m9.js +0 -64
  104. /package/export-template/docs/assets/{contributor-guide_architecture_detector-contract.md.BOeH4d1J.lean.js → contributor-guide_architecture_detector-contract.md.CgzUsQ6W.lean.js} +0 -0
  105. /package/export-template/docs/assets/{contributor-guide_architecture_impact-estimation.md.DYCDPgkh.lean.js → contributor-guide_architecture_impact-estimation.md.CooslVJt.lean.js} +0 -0
  106. /package/export-template/docs/assets/{contributor-guide_architecture_widget-rendering.md.DbqPf2OT.lean.js → contributor-guide_architecture_widget-rendering.md.R27gQrgY.lean.js} +0 -0
  107. /package/export-template/docs/assets/{contributor-guide_architecture_worker-protocol.md.B93qJ_tT.lean.js → contributor-guide_architecture_worker-protocol.md.IbnfNrV3.lean.js} +0 -0
@@ -1 +1 @@
1
- import{_ as t,o,c as d,a5 as a}from"./chunks/framework.DSg0KOwT.js";const p=JSON.parse('{"title":"Impact estimation","description":"","frontmatter":{},"headers":[],"relativePath":"contributor-guide/architecture/impact-estimation.md","filePath":"contributor-guide/architecture/impact-estimation.md"}'),c={name:"contributor-guide/architecture/impact-estimation.md"};function s(i,e,n,r,l,u){return o(),d("div",null,[...e[0]||(e[0]=[a('<h1 id="impact-estimation" tabindex="-1">Impact estimation <a class="header-anchor" href="#impact-estimation" aria-label="Permalink to &quot;Impact estimation&quot;">​</a></h1><p>Every finding covered by this section carries an optional <code>impactEstimate: {basis, wallClock, estimateMethod, rawWaste?}</code> (<code>src/types.ts</code>), attached by <code>src/impact-estimator.ts</code> as a post-pass after <code>DETECTORS</code> finishes (<code>src/analyzer.ts</code>). <code>estimateMethod</code> (&#39;measured&#39; | &#39;modeled&#39; | &#39;none&#39;) is a distinct axis from the per-finding <code>confidence</code> field (<a href="./board-widgets.html#confidence-metadata">Confidence metadata</a>): <code>confidence</code> says how much to trust the finding itself, <code>estimateMethod</code> says how its impact number was derived. <code>&#39;none&#39;</code> marks a purely informational finding with no waste model at all (<code>configAudit</code>, <code>stageFailed</code>, <code>failures</code>, <code>incompleteRun</code>, <code>slowHost</code>&#39;s byte-dimension multiDim shapes): it&#39;s distinct from <code>&#39;measured&#39;</code>/<code>&#39;modeled&#39;</code>, which both attach a real (if approximate) formula. <code>basis</code> is one of:</p><ul><li><code>&#39;serial&#39;</code>: the tied stage ran (effectively) alone; <code>wallClock</code> is a near-point estimate, <code>low === high</code>.</li><li><code>&#39;contended&#39;</code>: the tied stage shared wall-clock time with others; <code>wallClock</code> is an honest range, <code>high</code> optimistic (assumes the fix could still fully land), <code>low</code> the guaranteed floor.</li><li><code>&#39;resourceOnly&#39;</code>: no wall-clock claim is defensible (not stage-tied by nature, or the stage was excluded from the occupancy sweep), <code>wallClock: null</code>, but the formula&#39;s real signal survives in <code>rawWaste</code>.</li><li><code>&#39;informational&#39;</code>: no quantifiable magnitude at all, <code>wallClock: null</code>, no <code>rawWaste</code>.</li></ul><p><code>{basis: &#39;resourceOnly&#39;|&#39;informational&#39;, wallClock: null}</code> replaced the earlier design&#39;s <code>{low: 0, high: 0}</code>: that single value used to mean two incompatible things (&quot;provably no wall-clock cost&quot; and &quot;the model gave up&quot;), and 76-97% of stage-tied findings on real logs were the second case wearing the first case&#39;s clothing (2026-08-30 N1 redesign; see <code>docs/superpowers/specs/2026-08-30-critical-path-occupancy-redesign.md</code>).</p><p><code>rawWaste</code> is not exclusive to <code>resourceOnly</code>/<code>informational</code> findings: every <code>serial</code>/ <code>contended</code> finding carries it too (the only exceptions are <code>coldStart</code>, whose <code>wallClock</code> figure is already unclipped, and <code>estimateMethod: &#39;none&#39;</code> findings, which have no formula at all), holding the formula&#39;s pre-clip magnitude in the formula&#39;s own natural unit (ms, bytes or core-ms). <code>wallClock</code> is what the occupancy model says is recoverable, which clips against the stage&#39;s own physical floor; <code>rawWaste</code> is what the stage really wasted either way. The two answer different questions.</p><h2 id="occupancy-weighted-attribution" tabindex="-1">Occupancy-weighted attribution <a class="header-anchor" href="#occupancy-weighted-attribution" aria-label="Permalink to &quot;Occupancy-weighted attribution&quot;">​</a></h2><p><code>src/occupancy.ts</code> sweeps every stage&#39;s observed <code>[submittedAt, completedAt)</code> window and splits each instant&#39;s wall-clock among concurrently-active stages proportional to <code>coreWeight(S) = stage.executorRunTime / stageDurationMs</code> (an average-concurrency proxy, held constant across the stage&#39;s whole window: this codebase has no per-task timestamps outside the parser worker to do better). Summing a stage&#39;s share across its own window gives its <code>occupancy(S)</code>; <code>gate(S) = occupancy(S) / duration(S) ∈ [0, 1]</code> is the single number that replaces the old CPM model&#39;s <code>onCriticalPath</code>/<code>slackMs</code>/<code>isUniquelyCritical</code>. 1.0 means the stage ran completely alone; 0 means the stage had zero <code>executorRunTime</code> while overlapping other, positive-weight stages, so it got no share of the shared window. Stages with <code>duration(S) &lt;= 0</code> (Spark-skipped stages, or a malformed <code>submittedAt === completedAt</code>) are excluded from the sweep entirely.</p><p>This mechanism replaced a CPM (critical-path-method) graph over <code>parentIds</code> that produced near-zero on-critical-path membership on real logs (0.1-10.6% of stages, max graph depth 2 on 5 of 6 real logs measured): <code>parentIds</code> alone is too sparse a precedence signal for a meaningful longest-path computation. The same degeneracy fed <code>efficiency-model.ts</code>&#39;s <code>floorInfiniteMs</code> (&quot;floor with infinite executors&quot;); rather than leave a second, unreconciled critical-path number in the codebase for a future UI to display next to the occupancy-based figures above, <code>criticalPathMs()</code>/<code>CriticalPathStage</code> (embedded in <code>efficiency-model.ts</code>, never a standalone module) were removed outright (no occupancy-based replacement: occupancy apportions observed concurrent time, it doesn&#39;t compute a dependency-graph longest path, so there&#39;s no drop-in equivalent). <code>efficiency-model.ts</code> now reports only <code>floorZeroSkewMs</code> (total task time / total cores) as its theoretical floor.</p><p><code>ceiling(S) = max(stage.taskDurationMax, stage.executorRunTime / totalCores)</code> is a physical floor on a stage&#39;s own duration: bounded below by its single longest task (unsplittable no matter how much parallelism exists) or by its core-work spread across every core in the cluster, whichever is larger. Every waste formula&#39;s raw claim is clipped against it before gate-weighting: <code>wasteMs_clipped(S) = min(wasteMs_claimed, max(0, duration(S) - ceiling(S)))</code>, so a finding can never claim to save more than the portion of the stage&#39;s observed duration that sits above its own unbeatable floor. This is what fixes historical overclaim bugs (a <code>tinyTask</code> finding claiming 407.5s on a 13.1s stage capped to 8.9s; a <code>shuffle</code> finding claiming 1939.9s on a 991.3s/1688.3s stage capped to 610.3s/250.1s).</p><p><code>analyzer.ts</code> feeds this <code>totalCores</code> from <code>src/core-count.ts</code>&#39;s <code>computePeakConcurrentCores(app, executorsAdded, executorsRemoved)</code>, not the shared <code>computeTotalCores</code> helper. <code>computeTotalCores</code> sums every <code>ExecutorAdded</code> event&#39;s cores regardless of overlap, so under dynamic allocation or executor replacement it can far exceed the cores ever actually concurrent, which understates <code>ceiling(S)</code> and lets churn inflate a finding&#39;s claimed wall-clock. <code>computePeakConcurrentCores</code> instead sweeps add/remove events by timestamp and tracks the running total&#39;s peak, so a churned-through executor&#39;s cores are never double-counted against its replacement&#39;s. Same-timestamp events tie-break by delta ascending, so a removal applies before a same-instant replacement&#39;s addition (otherwise a same-instant swap would momentarily double-count both as concurrent). If every <code>executorsAdded</code> entry lacks <code>totalCores</code> the cores sweep peaks at zero and tells us nothing; the function then falls back to sweeping peak <em>executor count</em> instead (still concurrency-aware, just cores-blind) and multiplies by the configured per-executor core count, rather than falling back to <code>executorsAdded.length × cores</code>, which would reintroduce the exact cumulative-overcount-under-churn bug this function exists to avoid.</p><p>Per-finding estimate, using <code>wasteMs_clipped(S)</code>:</p><ul><li><code>gate(S) &gt;= 0.999</code>: <code>basis: &#39;serial&#39;</code>, <code>low = high = wasteMs_clipped(S)</code>.</li><li><code>gate(S) &lt; 0.999</code>: <code>basis: &#39;contended&#39;</code>, <code>high = wasteMs_clipped(S)</code>, <code>low = wasteMs_clipped(S) * gate(S)</code>.</li></ul><p>A finding spanning multiple stages (<code>stageIds</code>, plural) sums each stage&#39;s own estimate and caps the joint total at the union of just that finding&#39;s own stage windows (via <code>mergeIntervals</code>, <code>src/wall-clock.ts</code>): <code>high = min(Σ high_i, unionMs(stageIds))</code>, <code>low = min(Σ low_i, unionMs(stageIds))</code>. This is what prevents overclaiming when two or more of a finding&#39;s stages overlap in wall-clock time: a plain sum-and-cap, no CPM re-simulation. The union cap can force <code>low === high</code> numerically even when the constituent stages were individually contended (e.g. two fully-overlapping stages each at <code>gate</code> 0.5), so <code>basis</code> isn&#39;t derived from that numeric equality: a multi-stage finding gets <code>basis: &#39;serial&#39;</code> only when every one of its per-stage estimates was itself <code>&#39;serial&#39;</code>; otherwise <code>&#39;contended&#39;</code>.</p><p>Regression guard: a finding whose stage (or, for a multi-stage finding, every one of its stages) ran alone (<code>gate &gt;= 0.95</code>, deliberately looser than the <code>0.999</code> &quot;serial&quot; cutoff above: this guard exists to catch egregious false zeros, not to gate which <code>basis</code> a finding gets) with a real underlying magnitude (<code>rawWaste.value &gt; 0</code>) must never report <code>wallClock.high === 0</code> or <code>basis: &#39;informational&#39;/&#39;resourceOnly&#39;</code>. Covered by <code>tests/impact-estimator-real-log.test.js</code> against a real fixture; relies on <code>rawWaste</code> being attached to every serial/contended-capable formula (see above), so it&#39;s blind only to <code>coldStart</code> and the purely informational (<code>estimateMethod: &#39;none&#39;</code>) finding types.</p><p>Real-log spot-check (2026-08-30, <code>grupo-semanal-beauty-application_1785266278671_91660.zstd</code>): median <code>gate</code> across stages was <code>≈0.34</code> (0.3428060791718594 exactly); <code>collectRun</code> plus the occupancy sweep together took <code>≈6,589</code>ms on the largest fixture measured (<code>run-compare-calimax-candidate-application_1784568768686_119096.zstd</code>, <code>1169</code> stages): parse-dominated, the sweep alone was not isolated by this measurement, but not a magnitude that suggests a regression either. <code>54</code> previously-<code>{0,0}</code> stage-tied findings on stages that ran effectively alone now report a real <code>wallClock</code> range instead.</p><h2 id="cross-finding-rollup-computestageunionms" tabindex="-1">Cross-finding rollup: <code>computeStageUnionMs</code> <a class="header-anchor" href="#cross-finding-rollup-computestageunionms" aria-label="Permalink to &quot;Cross-finding rollup: `computeStageUnionMs`&quot;">​</a></h2><p>The Findings tab&#39;s recommendation rollup (<code>FixTheseFirst.tsx</code>, built from <code>buildRecommendationRollup</code>, <code>src/recommendation-rollup.ts</code>) groups the filtered catalog by detector <code>type</code>, then needs its own cap for a group of several findings of that type, not just one finding&#39;s own <code>stageIds</code>. Summing each finding&#39;s already-clipped <code>wallClock.high</code> naively double-counts any stage two of those findings both touch. <code>computeStageUnionMs(stageIds, stages)</code> covers this: collect every stage touched by any finding in the group, merge their <code>[submittedAt, completedAt)</code> intervals, and sum the merged intervals&#39; durations, so the group&#39;s wall-clock union holds regardless of how many findings&#39; <code>stageIds</code> overlap. Stages missing either bound are skipped rather than defaulted to <code>0</code> (the same filter <code>computeWallClock</code> applies), so a truncated log (the case <code>incompleteRun</code> flags) can&#39;t contribute a negative interval and a negative recoverable-time figure.</p><p>It reuses the same <code>mergeIntervals</code> primitive (<code>src/wall-clock.ts</code>) that backs <code>src/occupancy.ts</code>&#39;s per-finding <code>estimateMultiStage</code>/its internal <code>unionMs</code> sum, but is not an extension of that function: <code>estimateMultiStage</code> caps one finding&#39;s own multi-stage claim during the impact-estimation pass, before a <code>Finding</code> object even exists; <code>computeStageUnionMs</code> runs later, in the view layer, capping a naive sum <em>across</em> several already-estimated findings that happen to share a detector <code>type</code>. <code>buildTimeGroup</code> (same module) takes the smaller of the naive per-finding sum and this union figure as the group&#39;s <code>recoverableMsHigh</code>, falling back to the naive sum untouched when the group&#39;s findings carry no stage IDs at all (nothing to union against).</p><p>Within one <code>type</code> group, <code>buildRecommendationRollup</code> splits findings into up to three tiers, always rendered in this fixed order. <code>time</code> covers findings with a real <code>impactEstimate.wallClock</code> (<code>buildTimeGroup</code>, the union-capped figure above). <code>resource</code> covers findings with no <code>wallClock</code> but a <code>rawWaste</code> figure (<code>buildResourceGroup</code>), grouped again by <code>rawWaste.unit</code> so a <code>bytes</code> total never gets summed against a <code>coreHours</code> total under one type. <code>count</code> covers findings with neither (<code>buildCountGroup</code>), a plain per-impact-band tally with no magnitude claim at all. Each <code>RollupGroup</code> also carries its own <code>findings: Finding[]</code> (the exact members that fed the aggregate), which <code>FixTheseFirst.tsx</code> reads directly to pick a group&#39;s highest-impact member and to render its expanded, paginated list.</p><p>A type only contributes a tier when it has at least one finding of that kind; most types produce exactly one tier, but a type whose formula varies by <code>variant</code>/<code>rule</code> (e.g. <code>memoryUtilization</code>, see the coverage table below) can produce more than one.</p><p><code>cachingOpportunity</code> and <code>cacheUtilization</code> are both <code>cost-only</code>: <code>basis: &#39;resourceOnly&#39;</code>, <code>wallClock: null</code>, but <code>rawWaste.unit</code> is <code>&#39;ms&#39;</code>, the same unit a real <code>wallClock</code> figure would use, because their formula&#39;s natural output happens to be time (a re-read cost), not because either finding makes a wall-clock claim. Left unlabeled, a <code>resource</code>-tier &quot;ms&quot; total sitting next to a <code>time</code>-tier &quot;recoverable time&quot; total would read as directly comparable when it isn&#39;t: the resource figure was never gate-clipped against any stage&#39;s occupancy, so it can exceed what the stage actually spent. <code>FixTheseFirst.tsx</code> calls this out via its trailing-stat copy: a <code>resource</code>- kind group (any unit, including <code>ms</code>) always reads &quot;resource-cost projection&quot;, never &quot;recoverable&quot;, so the two ms-shaped numbers are never mistaken for the same kind of claim.</p><h2 id="per-formula-spot-checks" tabindex="-1">Per-formula spot-checks <a class="header-anchor" href="#per-formula-spot-checks" aria-label="Permalink to &quot;Per-formula spot-checks&quot;">​</a></h2><table tabindex="0"><thead><tr><th>Detector</th><th>Formula basis</th><th>Spot-check</th></tr></thead><tbody><tr><td>gc</td><td><code>jvmGCTime / (executorRunTime / stageDurationMs)</code></td><td><code>grupo-semanal-beauty-application_1785266278671_91660.zstd</code>, stage 507: <code>jvmGCTime</code>=1080ms, <code>executorRunTime</code>=27509ms, <code>stageDurationMs</code>=56279ms → <code>wasteMs</code> = 1080 / (27509/56279) ≈ 2209.5ms. That&#39;s ≈3.9% of the stage&#39;s 56.3s wall-clock duration, matching the finding&#39;s own reported <code>gcPct</code> (3.9%) exactly, as the formula guarantees by construction. Under the occupancy model this stage&#39;s <code>gate</code> is <code>0.041</code> (0.04145044590332269 exactly): <code>basis: &#39;contended&#39;</code>, <code>wallClock: {low: 91.6, high: 2209.5}</code> (91.5850382272182 / 2209.506706895925 exactly, per the Step 1 script&#39;s per-stage output).</td></tr><tr><td>shuffle</td><td><code>shuffleReadBytes / SHUFFLE_THROUGHPUT_BPS</code> (fallback tier; real-metrics parser not yet built)</td><td><code>ventas-mensual-multi-big-application_1785266278671_91510.zstd</code>, stage 99 (<code>SHFL</code> finding): <code>shuffleReadBytes</code>=204,172,518,504 → <code>wasteMs</code> = 204172518504 / 125,000,000 × 1000 ≈ 1,633,380ms, matching <code>rawWaste.value</code> exactly. Eyeball against the timeline: the stage&#39;s actual wall-clock duration is only 763,776ms, i.e. this run moved shuffle data at ≈267MB/s, roughly 2x the assumed 125MB/s (1Gbps) constant: expected for the fallback tier&#39;s deliberately conservative assumption, but worth knowing the modeled figure runs high on fast-network clusters. Under the occupancy model this stage&#39;s <code>gate</code> is <code>1</code>, <code>ceiling</code> is <code>≈434,568.1</code>ms (434568.1041666667 exactly): <code>wallClock</code> capped to <code>≈329,207.9</code>ms (329207.8958333333 exactly), since the raw 1,633,380ms claim vastly exceeds the stage&#39;s own 763,776ms duration.</td></tr><tr><td>spill</td><td><code>diskBytesSpilled / SPILL_IO_THROUGHPUT_BPS</code></td><td>Same run and stage (99): <code>diskBytesSpilled</code>=145,978,433,675 (note: the <code>SPILL</code> finding&#39;s own <code>value</code>/<code>metric</code> report <code>memoryBytesSpilled</code>=913,686,966,448, ~6x larger; the formula correctly uses the smaller disk figure, not that one) → <code>wasteMs</code> = 145978433675 / 200,000,000 × 1000 ≈ 729,892ms, matching <code>rawWaste.value</code> exactly. Eyeball: that&#39;s ≈191MB/s of implied disk throughput against the stage&#39;s 763,776ms actual duration, close to the assumed 200MB/s constant. Same ceiling-clip caveat as the shuffle row above applies here too, on the same stage.</td></tr></tbody></table><h2 id="overlap-caveat-skew-straggler" tabindex="-1">Overlap caveat: skew / straggler <a class="header-anchor" href="#overlap-caveat-skew-straggler" aria-label="Permalink to &quot;Overlap caveat: skew / straggler&quot;">​</a></h2><p><code>skew</code> (small-stage max-P50 fallback branch) and <code>straggler</code> can both fire on the same stage from the same single dominant outlier task, and each is clipped independently. This phase does not dedupe or suppress either: each keeps its own independently-computed <code>wallClock</code>. Do not sum <code>wallClock.high</code> across multiple findings on the same stage: if both fire together, they describe the same underlying waste, not two separate wastes. This overlap caveat is orthogonal to (and compounds with) the ceiling clip above: a stage with one dominant outlier task trips both detectors <em>and</em> has a small <code>ceiling</code>-derived recoverable room, since <code>ceiling</code> is itself <code>&gt;= taskDurationMax</code>, the very quantity these two detectors are reacting to.</p><p><code>analyzer.ts</code>&#39;s <code>flagSkewStragglerOverlap</code> (run after <code>deriveImpactBand</code>, once per <code>analyze()</code> call) surfaces this caveat to the reader instead of leaving it as an internal-only comment: whenever <code>skew</code>&#39;s <code>max/median</code> branch and <code>straggler</code> both fire on the same <code>stageId</code>, it appends a &quot;this overlaps with the X finding on this stage&quot; sentence to both findings&#39; <code>validationRequired</code> text (rather than suppressing either, so neither finding&#39;s own diagnostic value is lost). <code>skew</code>&#39;s <code>P95/median</code> branch samples a different task from <code>straggler</code>&#39;s own <code>taskDurationMax - taskDurationP50</code> delta, so it&#39;s excluded from the flag. The note rides the same confidence-caveat UI (<code>RowStatusCluster</code>) a reader already sees before trusting either finding&#39;s magnitude, since both detectors also carry <code>confidence: &#39;low&#39;</code> (their runtime-floor thresholds are unvalidated; see the confidence-disclosure note in detector-contract.md).</p><p><code>stageShape</code>&#39;s <code>taskStageSkew</code> rule no longer participates in this caveat: it reports a <code>resourceOnly</code> idle-core-ms figure (see the coverage table below) instead of a wall-clock claim, so there&#39;s nothing left to double-count against <code>skew</code>/<code>straggler</code>. Its trigger condition (<code>taskDurationMax / stageDurationMs &gt; skewWarn</code>) mathematically forces the occupancy-clipped wall-clock estimate to exactly zero on every firing (see <code>src/detectors.ts</code>&#39;s <code>taskStageSkew</code> comment), which is why it was moved off the wall-clock path entirely rather than reconciled against the same ceiling clip as its two siblings above.</p><h2 id="per-finding-type-coverage" tabindex="-1">Per-finding-type coverage <a class="header-anchor" href="#per-finding-type-coverage" aria-label="Permalink to &quot;Per-finding-type coverage&quot;">​</a></h2><p>One row per distinct <code>type</code> string <code>src/detectors.ts</code> actually emits (cross-checked against <code>computeEstimateForFinding</code>&#39;s <code>case</code> labels in <code>src/impact-estimator.ts</code>, not assumed from the prose here): every row below has a case, so the table itself is the coverage count, not a number restated here. <code>broadcastSizing</code> is a <code>DETECTORS</code> entry label only, and the plan-walk it drives emits <code>overBroadcast</code>/<code>underBroadcast</code> findings instead, so those two are the rows that appear, not <code>broadcastSizing</code> itself. Tag meanings: <code>measured</code> and <code>modeled</code> both produce a real, gate-clipped, non-<code>{0,0}</code><code>wallClock</code> (the difference is whether the formula&#39;s inputs are recorded per-stage fields or an assumed constant like a throughput figure); <code>cost-only</code> always reports <code>basis: &#39;resourceOnly&#39;</code>, <code>wallClock: null</code> but carries its real signal in <code>rawWaste</code>; <code>informational-only</code> reports <code>basis: &#39;informational&#39;</code>, <code>wallClock: null</code> with no <code>rawWaste</code> at all, since there&#39;s nothing quantifiable. A type with more than one tag fires a different formula per <code>variant</code>/<code>rule</code> on the same finding type; the basis column says which.</p><table tabindex="0"><thead><tr><th>Finding type</th><th>Scope</th><th>Tag</th><th>Basis</th></tr></thead><tbody><tr><td><code>retryWaste</code></td><td>stage</td><td>measured</td><td><code>retryWasteMs</code>, gate-clipped; pre-clip figure kept as <code>rawWaste</code> in <code>ms</code></td></tr><tr><td><code>speculationWaste</code></td><td>stage</td><td>measured</td><td><code>speculationWasteMs</code>, gate-clipped; pre-clip figure kept as <code>rawWaste</code> in <code>ms</code></td></tr><tr><td><code>coldStart</code></td><td>app</td><td>measured</td><td><code>gapSeconds × 1000</code>, unclipped, <code>basis: &#39;serial&#39;</code> unconditionally (a pre-first-task gap can&#39;t overlap any stage)</td></tr><tr><td><code>gc</code></td><td>stage</td><td>modeled</td><td><code>jvmGCTime / (executorRunTime / stageDurationMs)</code>, gate-clipped: the concurrency division is an approximation, not a reconstruction, hence <code>modeled</code>; <code>rawWaste</code> in <code>coreMs</code> is the raw <code>jvmGCTime</code> sum before that conversion</td></tr><tr><td><code>skew</code></td><td>stage</td><td>measured</td><td><code>taskDurationP95</code> or <code>Max</code> minus <code>P50</code> (per <code>metric</code>), gate-clipped; pre-clip figure kept as <code>rawWaste</code> in <code>ms</code></td></tr><tr><td><code>straggler</code></td><td>stage</td><td>measured</td><td><code>taskDurationMax − taskDurationP50</code>, gate-clipped</td></tr><tr><td><code>stageShape</code></td><td>stage</td><td>cost-only</td><td>all three rules are <code>estimateMethod: &#39;measured&#39;</code>, real per-stage fields, no assumed constant: <code>&#39;lowParallelism&#39;</code> → <code>rawWaste</code> in <code>coreMs</code> (idle cores × stage duration); <code>&#39;dataExplosion&#39;</code> → <code>rawWaste</code> in <code>bytes</code> (<code>outputBytes − inputBytes</code>); <code>&#39;taskStageSkew&#39;</code> → <code>rawWaste</code> in <code>coreMs</code> (<code>max(0, min(totalCores, taskCount) − 1) × (taskDurationMax − taskDurationP50)</code>, the cores idle during the straggler&#39;s tail at achieved concurrency)</td></tr><tr><td><code>slowHost</code></td><td>stage</td><td>measured / informational-only</td><td>duration-based variants (<code>hostMeanRatio</code>, <code>durationShare</code>, <code>multiDim</code>+<code>taskTime</code>): <code>value − taskDurationP50</code>, gate-clipped; byte-based <code>multiDim</code> dimensions: no formula yet</td></tr><tr><td><code>duplicatePlanSubtree</code></td><td>sql</td><td>measured</td><td>each contributing stage&#39;s real wall-clock duration × the redundant fraction <code>(occurrences − 1) / occurrences</code>, summed and capped at the finding&#39;s own <code>stageIds</code> union. <code>stageIds</code> is narrowed to the stages that actually ran the duplicated subtree&#39;s matched node instances (accumulator-ID evidence resolved onto each <code>PlanNode</code> at parse time, see <a href="./detector-contract.html#stage-id-attribution-for-plan-advisor-findings">Stage-ID attribution for Plan Advisor findings</a>), falling back to the whole execution&#39;s stages only when no matched instance has any accumulator coverage</td></tr><tr><td><code>shuffle</code></td><td>stage</td><td>modeled</td><td><code>shuffleReadBytes / SHUFFLE_THROUGHPUT_BPS</code> (assumed ~125MB/s), gate-clipped; <code>rawWaste</code> in <code>bytes</code> is the measured <code>shuffleReadBytes</code> behind it</td></tr><tr><td><code>spill</code></td><td>stage</td><td>modeled</td><td><code>diskBytesSpilled / SPILL_IO_THROUGHPUT_BPS</code> (assumed ~200MB/s), gate-clipped; <code>rawWaste</code> in <code>bytes</code> is <code>diskBytesSpilled</code>, which is the number the formula uses and not the <code>memoryBytesSpilled</code> the finding&#39;s own <code>metric</code> displays</td></tr><tr><td><code>stageSlowness</code></td><td>stage</td><td>modeled</td><td>stage duration minus the <code>stageSlowness</code> detector&#39;s own <code>infoMin</code> threshold, gate-clipped</td></tr><tr><td><code>partitionSizing</code></td><td>stage</td><td>modeled</td><td><code>maxPartitionTooBig</code>/<code>shufflePartitionSkew</code>: shuffle-throughput formulas, gate-clipped. <code>lowShuffleParallelism</code>: stage duration scaled down by the shortfall between actual and ideal-partition-count task counts (<code>stageDurationMs × (1 − taskCount / targetTaskCount)</code>), i.e. the serialized work more partitions would let run concurrently, not the scheduling cost of the tasks you&#39;d add to fix it</td></tr><tr><td><code>tinyTask</code></td><td>stage</td><td>modeled</td><td>excess task count over 10% of the stage&#39;s actual count, × assumed per-task scheduling overhead, gate-clipped; pre-clip figure kept as <code>rawWaste</code> in <code>ms</code></td></tr><tr><td><code>smallFiles</code></td><td>sql</td><td>modeled / cost-only</td><td><code>excessFileCount × FILE_OPEN_OVERHEAD_MS</code>, summed and capped over <code>stageIds</code>&#39;s union; with no <code>stageIds</code> to map to, cost-only with that same figure as <code>rawWaste</code> in <code>ms</code>. <code>stageIds</code> is narrowed the same way (see <a href="./detector-contract.html#stage-id-attribution-for-plan-advisor-findings">Stage-ID attribution for Plan Advisor findings</a>); falls back to the whole execution&#39;s stages when the flagged node(s) have no accumulator coverage.</td></tr><tr><td><code>overBroadcast</code></td><td>sql</td><td>modeled / cost-only</td><td><code>broadcastBytes / BROADCAST_BANDWIDTH_BPS</code>, summed and capped over <code>stageIds</code>&#39;s union; cost-only with <code>rawWaste</code> in <code>ms</code> when not stage-mappable. <code>stageIds</code> is narrowed the same way; falls back to the whole execution&#39;s stages when the flagged node(s) have no accumulator coverage.</td></tr><tr><td><code>underBroadcast</code></td><td>sql</td><td>modeled / cost-only</td><td><code>smallerSideBytes / BROADCAST_BANDWIDTH_BPS</code>, summed and capped over <code>stageIds</code>&#39;s union; cost-only with <code>rawWaste</code> in <code>ms</code> when not stage-mappable. <code>stageIds</code> is narrowed the same way; falls back to the whole execution&#39;s stages when the flagged node(s) have no accumulator coverage.</td></tr><tr><td><code>memoryUtilization</code></td><td>app</td><td>cost-only / informational-only</td><td>Three of the four variants report <code>rawWaste</code> in <code>mbSeconds</code>: <code>variant: &#39;wasteModel&#39;</code> passes through its own <code>wastedMBSeconds</code>; <code>&#39;idleCores&#39;</code> uses <code>idleRateFraction × allocatedMB × peakExecutors × appDurationSeconds</code>; <code>&#39;memoryBand&#39;</code> with <code>rule: &#39;heapOverProvisioned&#39;</code> uses <code>(allocatedBytes − heap) in MB × appDurationSeconds</code>. <code>&#39;memoryBand&#39;</code> with <code>rule: &#39;heapNearCapacity&#39;</code> is an OOM-risk signal rather than a waste, and the <code>dataUnavailable</code> shape has no inputs at all: both informational-only</td></tr><tr><td><code>utilization</code></td><td>app</td><td>cost-only</td><td><code>rawWaste</code> in <code>coreHours</code>: <code>(1 − utilizationFraction) × appDurationMs × totalCores / 3.6e6</code></td></tr><tr><td><code>coreLocality</code></td><td>app</td><td>cost-only</td><td><code>rawWaste</code> in <code>coreMs</code>: <code>nonLocalTaskCount × NETWORK_FETCH_PENALTY_MS</code></td></tr><tr><td><code>autoscalingChurn</code></td><td>app</td><td>cost-only</td><td><code>rawWaste</code> in <code>coreHours</code>: <code>shortLivedExecutorCount × EXECUTOR_STARTUP_OVERHEAD_MS / 3.6e6</code></td></tr><tr><td><code>configAudit</code></td><td>config</td><td>informational-only</td><td>a config-drift check standing alone; no waste formula</td></tr><tr><td><code>jobFailureRate</code></td><td>app</td><td>cost-only</td><td><code>rawWaste</code> in <code>coreHours</code>: <code>failedJobCount × avgJobDurationMs / 3.6e6</code></td></tr><tr><td><code>cachingOpportunity</code></td><td>app</td><td>cost-only</td><td><code>rawWaste</code> in <code>ms</code>: <code>totalReadBytes / RE_READ_THROUGHPUT_BPS</code></td></tr><tr><td><code>cacheUtilization</code></td><td>app</td><td>cost-only</td><td><code>rawWaste</code> in <code>ms</code>: uncached-or-spilled bytes <code>/ RE_READ_THROUGHPUT_BPS</code>, where the never-cached partitions&#39; bytes are extrapolated from the cached partitions&#39; own average size (<code>memorySize + diskSize</code>, over <code>numCachedPartitions</code>), plus <code>diskSize</code> again for the already-cached-but-on-disk partitions&#39; own re-read cost</td></tr><tr><td><code>stageFailed</code></td><td>stage</td><td>informational-only</td><td>no waste formula</td></tr><tr><td><code>failures</code></td><td>stage</td><td>informational-only</td><td>no waste formula</td></tr><tr><td><code>incompleteRun</code></td><td>app</td><td>informational-only</td><td>no waste formula</td></tr></tbody></table>',30)])])}const g=t(c,[["render",s]]);export{p as __pageData,g as default};
1
+ import{_ as t,o,c as d,a5 as a}from"./chunks/framework.DSg0KOwT.js";const p=JSON.parse('{"title":"Impact estimation","description":"","frontmatter":{},"headers":[],"relativePath":"contributor-guide/architecture/impact-estimation.md","filePath":"contributor-guide/architecture/impact-estimation.md"}'),c={name:"contributor-guide/architecture/impact-estimation.md"};function s(i,e,n,r,l,u){return o(),d("div",null,[...e[0]||(e[0]=[a('<h1 id="impact-estimation" tabindex="-1">Impact estimation <a class="header-anchor" href="#impact-estimation" aria-label="Permalink to &quot;Impact estimation&quot;">​</a></h1><p>Every finding covered by this section carries an optional <code>impactEstimate: {basis, wallClock, estimateMethod, rawWaste?}</code> (<code>src/types.ts</code>), attached by <code>src/impact-estimator.ts</code> as a post-pass after <code>DETECTORS</code> finishes (<code>src/analyzer.ts</code>). <code>estimateMethod</code> (&#39;measured&#39; | &#39;modeled&#39; | &#39;none&#39;) is a distinct axis from the per-finding <code>confidence</code> field (<a href="./board-widgets.html#confidence-metadata">Confidence metadata</a>): <code>confidence</code> says how much to trust the finding itself, <code>estimateMethod</code> says how its impact number was derived. <code>&#39;none&#39;</code> marks a purely informational finding with no waste model at all (<code>configAudit</code>, <code>stageFailed</code>, <code>failures</code>, <code>incompleteRun</code>, <code>slowHost</code>&#39;s byte-dimension multiDim shapes): it&#39;s distinct from <code>&#39;measured&#39;</code>/<code>&#39;modeled&#39;</code>, which both attach a real (if approximate) formula. <code>basis</code> is one of:</p><ul><li><code>&#39;serial&#39;</code>: the tied stage ran (effectively) alone; <code>wallClock</code> is a near-point estimate, <code>low === high</code>.</li><li><code>&#39;contended&#39;</code>: the tied stage shared wall-clock time with others; <code>wallClock</code> is an honest range, <code>high</code> optimistic (assumes the fix could still fully land), <code>low</code> the guaranteed floor.</li><li><code>&#39;resourceOnly&#39;</code>: no wall-clock claim is defensible (not stage-tied by nature, or the stage was excluded from the occupancy sweep), <code>wallClock: null</code>, but the formula&#39;s real signal survives in <code>rawWaste</code>.</li><li><code>&#39;informational&#39;</code>: no quantifiable magnitude at all, <code>wallClock: null</code>, no <code>rawWaste</code>.</li></ul><p><code>{basis: &#39;resourceOnly&#39;|&#39;informational&#39;, wallClock: null}</code> replaced the earlier design&#39;s <code>{low: 0, high: 0}</code>: that single value used to mean two incompatible things (&quot;provably no wall-clock cost&quot; and &quot;the model gave up&quot;), and 76-97% of stage-tied findings on real logs were the second case wearing the first case&#39;s clothing (2026-08-30 N1 redesign; see <code>docs/superpowers/specs/2026-08-30-critical-path-occupancy-redesign.md</code>).</p><p><code>rawWaste</code> is not exclusive to <code>resourceOnly</code>/<code>informational</code> findings: every <code>serial</code>/ <code>contended</code> finding carries it too (the only exceptions are <code>coldStart</code>, whose <code>wallClock</code> figure is already unclipped, and <code>estimateMethod: &#39;none&#39;</code> findings, which have no formula at all), holding the formula&#39;s pre-clip magnitude in the formula&#39;s own natural unit (ms, bytes or core-ms). <code>wallClock</code> is what the occupancy model says is recoverable, which clips against the stage&#39;s own physical floor; <code>rawWaste</code> is what the stage really wasted either way. The two answer different questions.</p><h2 id="occupancy-weighted-attribution" tabindex="-1">Occupancy-weighted attribution <a class="header-anchor" href="#occupancy-weighted-attribution" aria-label="Permalink to &quot;Occupancy-weighted attribution&quot;">​</a></h2><p><code>src/occupancy.ts</code> sweeps every stage&#39;s observed <code>[submittedAt, completedAt)</code> window and splits each instant&#39;s wall-clock among concurrently-active stages proportional to <code>coreWeight(S) = stage.executorRunTime / stageDurationMs</code> (an average-concurrency proxy, held constant across the stage&#39;s whole window: this codebase has no per-task timestamps outside the parser worker to do better). Summing a stage&#39;s share across its own window gives its <code>occupancy(S)</code>; <code>gate(S) = occupancy(S) / duration(S) ∈ [0, 1]</code> is the single number that replaces the old CPM model&#39;s <code>onCriticalPath</code>/<code>slackMs</code>/<code>isUniquelyCritical</code>. 1.0 means the stage ran completely alone; 0 means the stage had zero <code>executorRunTime</code> while overlapping other, positive-weight stages, so it got no share of the shared window. Stages with <code>duration(S) &lt;= 0</code> (Spark-skipped stages, or a malformed <code>submittedAt === completedAt</code>) are excluded from the sweep entirely.</p><p>This mechanism replaced a CPM (critical-path-method) graph over <code>parentIds</code> that produced near-zero on-critical-path membership on real logs (0.1-10.6% of stages, max graph depth 2 on 5 of 6 real logs measured): <code>parentIds</code> alone is too sparse a precedence signal for a meaningful longest-path computation. The same degeneracy fed <code>efficiency-model.ts</code>&#39;s <code>floorInfiniteMs</code> (&quot;floor with infinite executors&quot;); rather than leave a second, unreconciled critical-path number in the codebase for a future UI to display next to the occupancy-based figures above, <code>criticalPathMs()</code>/<code>CriticalPathStage</code> (embedded in <code>efficiency-model.ts</code>, never a standalone module) were removed outright (no occupancy-based replacement: occupancy apportions observed concurrent time, it doesn&#39;t compute a dependency-graph longest path, so there&#39;s no drop-in equivalent). <code>efficiency-model.ts</code> now reports only <code>floorZeroSkewMs</code> (total task time / total cores) as its theoretical floor.</p><p><code>ceiling(S) = max(stage.taskDurationMax, stage.executorRunTime / totalCores)</code> is a physical floor on a stage&#39;s own duration: bounded below by its single longest task (unsplittable no matter how much parallelism exists) or by its core-work spread across every core in the cluster, whichever is larger. Every waste formula&#39;s raw claim is clipped against it before gate-weighting: <code>wasteMs_clipped(S) = min(wasteMs_claimed, max(0, duration(S) - ceiling(S)))</code>, so a finding can never claim to save more than the portion of the stage&#39;s observed duration that sits above its own unbeatable floor. This is what fixes historical overclaim bugs (a <code>tinyTask</code> finding claiming 407.5s on a 13.1s stage capped to 8.9s; a <code>shuffle</code> finding claiming 1939.9s on a 991.3s/1688.3s stage capped to 610.3s/250.1s).</p><p><code>analyzer.ts</code> feeds this <code>totalCores</code> from <code>src/core-count.ts</code>&#39;s <code>computePeakConcurrentCores(app, executorsAdded, executorsRemoved)</code>, not the shared <code>computeTotalCores</code> helper. <code>computeTotalCores</code> sums every <code>ExecutorAdded</code> event&#39;s cores regardless of overlap, so under dynamic allocation or executor replacement it can far exceed the cores ever actually concurrent, which understates <code>ceiling(S)</code> and lets churn inflate a finding&#39;s claimed wall-clock. <code>computePeakConcurrentCores</code> instead sweeps add/remove events by timestamp and tracks the running total&#39;s peak, so a churned-through executor&#39;s cores are never double-counted against its replacement&#39;s. Same-timestamp events tie-break by delta ascending, so a removal applies before a same-instant replacement&#39;s addition (otherwise a same-instant swap would momentarily double-count both as concurrent). If every <code>executorsAdded</code> entry lacks <code>totalCores</code> the cores sweep peaks at zero and tells us nothing; the function then falls back to sweeping peak <em>executor count</em> instead (still concurrency-aware, just cores-blind) and multiplies by the configured per-executor core count, rather than falling back to <code>executorsAdded.length × cores</code>, which would reintroduce the exact cumulative-overcount-under-churn bug this function exists to avoid.</p><p>Per-finding estimate, using <code>wasteMs_clipped(S)</code>:</p><ul><li><code>gate(S) &gt;= 0.999</code>: <code>basis: &#39;serial&#39;</code>, <code>low = high = wasteMs_clipped(S)</code>.</li><li><code>gate(S) &lt; 0.999</code>: <code>basis: &#39;contended&#39;</code>, <code>high = wasteMs_clipped(S)</code>, <code>low = wasteMs_clipped(S) * gate(S)</code>.</li></ul><p>A finding spanning multiple stages (<code>stageIds</code>, plural) sums each stage&#39;s own estimate and caps the joint total at the union of just that finding&#39;s own stage windows (via <code>mergeIntervals</code>, <code>src/wall-clock.ts</code>): <code>high = min(Σ high_i, unionMs(stageIds))</code>, <code>low = min(Σ low_i, unionMs(stageIds))</code>. This is what prevents overclaiming when two or more of a finding&#39;s stages overlap in wall-clock time: a plain sum-and-cap, no CPM re-simulation. The union cap can force <code>low === high</code> numerically even when the constituent stages were individually contended (e.g. two fully-overlapping stages each at <code>gate</code> 0.5), so <code>basis</code> isn&#39;t derived from that numeric equality: a multi-stage finding gets <code>basis: &#39;serial&#39;</code> only when every one of its per-stage estimates was itself <code>&#39;serial&#39;</code>; otherwise <code>&#39;contended&#39;</code>.</p><p>Regression guard: a finding whose stage (or, for a multi-stage finding, every one of its stages) ran alone (<code>gate &gt;= 0.95</code>, deliberately looser than the <code>0.999</code> &quot;serial&quot; cutoff above: this guard exists to catch egregious false zeros, not to gate which <code>basis</code> a finding gets) with a real underlying magnitude (<code>rawWaste.value &gt; 0</code>) must never report <code>wallClock.high === 0</code> or <code>basis: &#39;informational&#39;/&#39;resourceOnly&#39;</code>. Covered by <code>tests/impact-estimator-real-log.test.js</code> against a real fixture; relies on <code>rawWaste</code> being attached to every serial/contended-capable formula (see above), so it&#39;s blind only to <code>coldStart</code> and the purely informational (<code>estimateMethod: &#39;none&#39;</code>) finding types.</p><p>Real-log spot-check (2026-08-30, <code>grupo-semanal-beauty-application_1785266278671_91660.zstd</code>): median <code>gate</code> across stages was <code>≈0.34</code> (0.3428060791718594 exactly); <code>collectRun</code> plus the occupancy sweep together took <code>≈6,589</code>ms on the largest fixture measured (<code>run-compare-calimax-candidate-application_1784568768686_119096.zstd</code>, <code>1169</code> stages): parse-dominated, the sweep alone was not isolated by this measurement, but not a magnitude that suggests a regression either. <code>54</code> previously-<code>{0,0}</code> stage-tied findings on stages that ran effectively alone now report a real <code>wallClock</code> range instead.</p><h2 id="cross-finding-rollup-computestageunionms" tabindex="-1">Cross-finding rollup: <code>computeStageUnionMs</code> <a class="header-anchor" href="#cross-finding-rollup-computestageunionms" aria-label="Permalink to &quot;Cross-finding rollup: `computeStageUnionMs`&quot;">​</a></h2><p>The Findings tab&#39;s recommendation rollup (<code>FixTheseFirst.tsx</code>, built from <code>buildRecommendationRollup</code>, <code>src/recommendation-rollup.ts</code>) groups the filtered catalog by detector <code>type</code>, then needs its own cap for a group of several findings of that type, not just one finding&#39;s own <code>stageIds</code>. Summing each finding&#39;s already-clipped <code>wallClock.high</code> naively double-counts any stage two of those findings both touch. <code>computeStageUnionMs(stageIds, stages)</code> covers this: collect every stage touched by any finding in the group, merge their <code>[submittedAt, completedAt)</code> intervals, and sum the merged intervals&#39; durations, so the group&#39;s wall-clock union holds regardless of how many findings&#39; <code>stageIds</code> overlap. Stages missing either bound are skipped rather than defaulted to <code>0</code> (the same filter <code>computeWallClock</code> applies), so a truncated log (the case <code>incompleteRun</code> flags) can&#39;t contribute a negative interval and a negative recoverable-time figure.</p><p>It reuses the same <code>mergeIntervals</code> primitive (<code>src/wall-clock.ts</code>) that backs <code>src/occupancy.ts</code>&#39;s per-finding <code>estimateMultiStage</code>/its internal <code>unionMs</code> sum, but is not an extension of that function: <code>estimateMultiStage</code> caps one finding&#39;s own multi-stage claim during the impact-estimation pass, before a <code>Finding</code> object even exists; <code>computeStageUnionMs</code> runs later, in the view layer, capping a naive sum <em>across</em> several already-estimated findings that happen to share a detector <code>type</code>. <code>buildTimeGroup</code> (same module) takes the smaller of the naive per-finding sum and this union figure as the group&#39;s <code>recoverableMsHigh</code>, falling back to the naive sum untouched when the group&#39;s findings carry no stage IDs at all (nothing to union against).</p><p>Within one <code>type</code> group, <code>buildRecommendationRollup</code> splits findings into up to three tiers, always rendered in this fixed order. <code>time</code> covers findings with a real <code>impactEstimate.wallClock</code> (<code>buildTimeGroup</code>, the union-capped figure above). <code>resource</code> covers findings with no <code>wallClock</code> but a <code>rawWaste</code> figure (<code>buildResourceGroup</code>), grouped again by <code>rawWaste.unit</code> so a <code>bytes</code> total never gets summed against a <code>coreHours</code> total under one type. <code>count</code> covers findings with neither (<code>buildCountGroup</code>), a plain per-impact-band tally with no magnitude claim at all. Each <code>RollupGroup</code> also carries its own <code>findings: Finding[]</code> (the exact members that fed the aggregate), which <code>FixTheseFirst.tsx</code> reads directly to pick a group&#39;s highest-impact member and to render its expanded, paginated list.</p><p>A type only contributes a tier when it has at least one finding of that kind; most types produce exactly one tier, but a type whose formula varies by <code>variant</code>/<code>rule</code> (e.g. <code>memoryUtilization</code>, see the coverage table below) can produce more than one.</p><p><code>cachingOpportunity</code> and <code>cacheUtilization</code> are both <code>cost-only</code>: <code>basis: &#39;resourceOnly&#39;</code>, <code>wallClock: null</code>, but <code>rawWaste.unit</code> is <code>&#39;ms&#39;</code>, the same unit a real <code>wallClock</code> figure would use, because their formula&#39;s natural output happens to be time (a re-read cost), not because either finding makes a wall-clock claim. Left unlabeled, a <code>resource</code>-tier &quot;ms&quot; total sitting next to a <code>time</code>-tier &quot;recoverable time&quot; total would read as directly comparable when it isn&#39;t: the resource figure was never gate-clipped against any stage&#39;s occupancy, so it can exceed what the stage actually spent. <code>FixTheseFirst.tsx</code> calls this out via its trailing-stat copy: a <code>resource</code>- kind group (any unit, including <code>ms</code>) always reads &quot;resource-cost projection&quot;, never &quot;recoverable&quot;, so the two ms-shaped numbers are never mistaken for the same kind of claim.</p><h2 id="per-formula-spot-checks" tabindex="-1">Per-formula spot-checks <a class="header-anchor" href="#per-formula-spot-checks" aria-label="Permalink to &quot;Per-formula spot-checks&quot;">​</a></h2><table tabindex="0"><thead><tr><th>Detector</th><th>Formula basis</th><th>Spot-check</th></tr></thead><tbody><tr><td>gc</td><td><code>jvmGCTime / (executorRunTime / stageDurationMs)</code></td><td><code>grupo-semanal-beauty-application_1785266278671_91660.zstd</code>, stage 507: <code>jvmGCTime</code>=1080ms, <code>executorRunTime</code>=27509ms, <code>stageDurationMs</code>=56279ms → <code>wasteMs</code> = 1080 / (27509/56279) ≈ 2209.5ms. That&#39;s ≈3.9% of the stage&#39;s 56.3s wall-clock duration, matching the finding&#39;s own reported <code>gcPct</code> (3.9%) exactly, as the formula guarantees by construction. Under the occupancy model this stage&#39;s <code>gate</code> is <code>0.041</code> (0.04145044590332269 exactly): <code>basis: &#39;contended&#39;</code>, <code>wallClock: {low: 91.6, high: 2209.5}</code> (91.5850382272182 / 2209.506706895925 exactly, per the Step 1 script&#39;s per-stage output).</td></tr><tr><td>shuffle</td><td><code>shuffleReadBytes / SHUFFLE_THROUGHPUT_BPS</code> (fallback tier; real-metrics parser not yet built)</td><td><code>ventas-mensual-multi-big-application_1785266278671_91510.zstd</code>, stage 99 (<code>SHFL</code> finding): <code>shuffleReadBytes</code>=204,172,518,504 → <code>wasteMs</code> = 204172518504 / 125,000,000 × 1000 ≈ 1,633,380ms, matching <code>rawWaste.value</code> exactly. Eyeball against the timeline: the stage&#39;s actual wall-clock duration is only 763,776ms, i.e. this run moved shuffle data at ≈267MB/s, roughly 2x the assumed 125MB/s (1Gbps) constant: expected for the fallback tier&#39;s deliberately conservative assumption, but worth knowing the modeled figure runs high on fast-network clusters. Under the occupancy model this stage&#39;s <code>gate</code> is <code>1</code>, <code>ceiling</code> is <code>≈434,568.1</code>ms (434568.1041666667 exactly): <code>wallClock</code> capped to <code>≈329,207.9</code>ms (329207.8958333333 exactly), since the raw 1,633,380ms claim vastly exceeds the stage&#39;s own 763,776ms duration.</td></tr><tr><td>spill</td><td><code>diskBytesSpilled / SPILL_IO_THROUGHPUT_BPS</code></td><td>Same run and stage (99): <code>diskBytesSpilled</code>=145,978,433,675 (note: the <code>SPILL</code> finding&#39;s own <code>value</code>/<code>metric</code> report <code>memoryBytesSpilled</code>=913,686,966,448, ~6x larger; the formula correctly uses the smaller disk figure, not that one) → <code>wasteMs</code> = 145978433675 / 200,000,000 × 1000 ≈ 729,892ms, matching <code>rawWaste.value</code> exactly. Eyeball: that&#39;s ≈191MB/s of implied disk throughput against the stage&#39;s 763,776ms actual duration, close to the assumed 200MB/s constant. Same ceiling-clip caveat as the shuffle row above applies here too, on the same stage.</td></tr></tbody></table><h2 id="overlap-caveat-skew-straggler" tabindex="-1">Overlap caveat: skew / straggler <a class="header-anchor" href="#overlap-caveat-skew-straggler" aria-label="Permalink to &quot;Overlap caveat: skew / straggler&quot;">​</a></h2><p><code>skew</code> (small-stage max-P50 fallback branch) and <code>straggler</code> can both fire on the same stage from the same single dominant outlier task, and each is clipped independently. This phase does not dedupe or suppress either: each keeps its own independently-computed <code>wallClock</code>. Do not sum <code>wallClock.high</code> across multiple findings on the same stage: if both fire together, they describe the same underlying waste, not two separate wastes. This overlap caveat is orthogonal to (and compounds with) the ceiling clip above: a stage with one dominant outlier task trips both detectors <em>and</em> has a small <code>ceiling</code>-derived recoverable room, since <code>ceiling</code> is itself <code>&gt;= taskDurationMax</code>, the very quantity these two detectors are reacting to.</p><p><code>analyzer.ts</code>&#39;s <code>flagSkewStragglerOverlap</code> (run after <code>deriveImpactBand</code>, once per <code>analyze()</code> call) surfaces this caveat to the reader instead of leaving it as an internal-only comment: whenever <code>skew</code>&#39;s <code>max/median</code> branch and <code>straggler</code> both fire on the same <code>stageId</code>, it appends a &quot;this overlaps with the X finding on this stage&quot; sentence to both findings&#39; <code>validationRequired</code> text (rather than suppressing either, so neither finding&#39;s own diagnostic value is lost). <code>skew</code>&#39;s <code>P95/median</code> branch samples a different task from <code>straggler</code>&#39;s own <code>taskDurationMax - taskDurationP50</code> delta, so it&#39;s excluded from the flag. The note rides the same confidence-caveat UI (<code>RowStatusCluster</code>) a reader already sees before trusting either finding&#39;s magnitude, since both detectors also carry a <code>confidence</code> field that scales <code>low</code>/<code>medium</code>/<code>high</code> off how far the finding sits past its own runtime-floor threshold (still unvalidated; see the confidence-disclosure note in detector-contract.md).</p><p><code>stageShape</code>&#39;s <code>taskStageSkew</code> rule no longer participates in this caveat: it reports a <code>resourceOnly</code> idle-core-ms figure (see the coverage table below) instead of a wall-clock claim, so there&#39;s nothing left to double-count against <code>skew</code>/<code>straggler</code>. Its trigger condition (<code>taskDurationMax / stageDurationMs &gt; skewWarn</code>) mathematically forces the occupancy-clipped wall-clock estimate to exactly zero on every firing (see <code>src/detectors.ts</code>&#39;s <code>taskStageSkew</code> comment), which is why it was moved off the wall-clock path entirely rather than reconciled against the same ceiling clip as its two siblings above.</p><h2 id="per-finding-type-coverage" tabindex="-1">Per-finding-type coverage <a class="header-anchor" href="#per-finding-type-coverage" aria-label="Permalink to &quot;Per-finding-type coverage&quot;">​</a></h2><p>One row per distinct <code>type</code> string <code>src/detectors.ts</code> actually emits (cross-checked against <code>computeEstimateForFinding</code>&#39;s <code>case</code> labels in <code>src/impact-estimator.ts</code>, not assumed from the prose here): every row below has a case, so the table itself is the coverage count, not a number restated here. <code>broadcastSizing</code> is a <code>DETECTORS</code> entry label only, and the plan-walk it drives emits <code>overBroadcast</code>/<code>underBroadcast</code> findings instead, so those two are the rows that appear, not <code>broadcastSizing</code> itself. Tag meanings: <code>measured</code> and <code>modeled</code> both produce a real, gate-clipped, non-<code>{0,0}</code><code>wallClock</code> (the difference is whether the formula&#39;s inputs are recorded per-stage fields or an assumed constant like a throughput figure); <code>cost-only</code> always reports <code>basis: &#39;resourceOnly&#39;</code>, <code>wallClock: null</code> but carries its real signal in <code>rawWaste</code>; <code>informational-only</code> reports <code>basis: &#39;informational&#39;</code>, <code>wallClock: null</code> with no <code>rawWaste</code> at all, since there&#39;s nothing quantifiable. A type with more than one tag fires a different formula per <code>variant</code>/<code>rule</code> on the same finding type; the basis column says which.</p><table tabindex="0"><thead><tr><th>Finding type</th><th>Scope</th><th>Tag</th><th>Basis</th></tr></thead><tbody><tr><td><code>retryWaste</code></td><td>stage</td><td>measured</td><td><code>retryWasteMs</code>, gate-clipped; pre-clip figure kept as <code>rawWaste</code> in <code>ms</code></td></tr><tr><td><code>speculationWaste</code></td><td>stage</td><td>measured</td><td><code>speculationWasteMs</code>, gate-clipped; pre-clip figure kept as <code>rawWaste</code> in <code>ms</code></td></tr><tr><td><code>coldStart</code></td><td>app</td><td>measured</td><td><code>gapSeconds × 1000</code>, unclipped, <code>basis: &#39;serial&#39;</code> unconditionally (a pre-first-task gap can&#39;t overlap any stage)</td></tr><tr><td><code>gc</code></td><td>stage</td><td>modeled</td><td><code>jvmGCTime / (executorRunTime / stageDurationMs)</code>, gate-clipped: the concurrency division is an approximation, not a reconstruction, hence <code>modeled</code>; <code>rawWaste</code> in <code>coreMs</code> is the raw <code>jvmGCTime</code> sum before that conversion</td></tr><tr><td><code>skew</code></td><td>stage</td><td>measured</td><td><code>taskDurationP95</code> or <code>Max</code> minus <code>P50</code> (per <code>metric</code>), gate-clipped; pre-clip figure kept as <code>rawWaste</code> in <code>ms</code></td></tr><tr><td><code>straggler</code></td><td>stage</td><td>measured</td><td><code>taskDurationMax − taskDurationP50</code>, gate-clipped</td></tr><tr><td><code>stageShape</code></td><td>stage</td><td>cost-only</td><td>all three rules are <code>estimateMethod: &#39;measured&#39;</code>, real per-stage fields, no assumed constant: <code>&#39;lowParallelism&#39;</code> → <code>rawWaste</code> in <code>coreMs</code> (idle cores × stage duration); <code>&#39;dataExplosion&#39;</code> → <code>rawWaste</code> in <code>bytes</code> (<code>outputBytes − inputBytes</code>); <code>&#39;taskStageSkew&#39;</code> → <code>rawWaste</code> in <code>coreMs</code> (<code>max(0, min(totalCores, taskCount) − 1) × (taskDurationMax − taskDurationP50)</code>, the cores idle during the straggler&#39;s tail at achieved concurrency)</td></tr><tr><td><code>slowHost</code></td><td>stage</td><td>measured / informational-only</td><td>duration-based variants (<code>hostMeanRatio</code>, <code>durationShare</code>, <code>multiDim</code>+<code>taskTime</code>): <code>value − taskDurationP50</code>, gate-clipped; byte-based <code>multiDim</code> dimensions: no formula yet</td></tr><tr><td><code>duplicatePlanSubtree</code></td><td>sql</td><td>measured</td><td>each contributing stage&#39;s real wall-clock duration × the redundant fraction <code>(occurrences − 1) / occurrences</code>, summed and capped at the finding&#39;s own <code>stageIds</code> union. <code>stageIds</code> is narrowed to the stages that actually ran the duplicated subtree&#39;s matched node instances (accumulator-ID evidence resolved onto each <code>PlanNode</code> at parse time, see <a href="./detector-contract.html#stage-id-attribution-for-plan-advisor-findings">Stage-ID attribution for Plan Advisor findings</a>), falling back to the whole execution&#39;s stages only when no matched instance has any accumulator coverage</td></tr><tr><td><code>shuffle</code></td><td>stage</td><td>modeled</td><td><code>shuffleReadBytes / SHUFFLE_THROUGHPUT_BPS</code> (assumed ~125MB/s), gate-clipped; <code>rawWaste</code> in <code>bytes</code> is the measured <code>shuffleReadBytes</code> behind it</td></tr><tr><td><code>spill</code></td><td>stage</td><td>modeled</td><td><code>diskBytesSpilled / SPILL_IO_THROUGHPUT_BPS</code> (assumed ~200MB/s), gate-clipped; <code>rawWaste</code> in <code>bytes</code> is <code>diskBytesSpilled</code>, which is the number the formula uses and not the <code>memoryBytesSpilled</code> the finding&#39;s own <code>metric</code> displays</td></tr><tr><td><code>stageSlowness</code></td><td>stage</td><td>modeled</td><td>stage duration minus the <code>stageSlowness</code> detector&#39;s own <code>infoMin</code> threshold, gate-clipped</td></tr><tr><td><code>partitionSizing</code></td><td>stage</td><td>modeled</td><td><code>maxPartitionTooBig</code>/<code>shufflePartitionSkew</code>: shuffle-throughput formulas, gate-clipped. <code>lowShuffleParallelism</code>: stage duration scaled down by the shortfall between actual and ideal-partition-count task counts (<code>stageDurationMs × (1 − taskCount / targetTaskCount)</code>), i.e. the serialized work more partitions would let run concurrently, not the scheduling cost of the tasks you&#39;d add to fix it</td></tr><tr><td><code>tinyTask</code></td><td>stage</td><td>modeled</td><td>excess task count over 10% of the stage&#39;s actual count, × assumed per-task scheduling overhead, gate-clipped; pre-clip figure kept as <code>rawWaste</code> in <code>ms</code></td></tr><tr><td><code>smallFiles</code></td><td>sql</td><td>modeled / cost-only</td><td><code>excessFileCount × FILE_OPEN_OVERHEAD_MS</code>, summed and capped over <code>stageIds</code>&#39;s union; with no <code>stageIds</code> to map to, cost-only with that same figure as <code>rawWaste</code> in <code>ms</code>. <code>stageIds</code> is narrowed the same way (see <a href="./detector-contract.html#stage-id-attribution-for-plan-advisor-findings">Stage-ID attribution for Plan Advisor findings</a>); falls back to the whole execution&#39;s stages when the flagged node(s) have no accumulator coverage.</td></tr><tr><td><code>overBroadcast</code></td><td>sql</td><td>modeled / cost-only</td><td><code>broadcastBytes / BROADCAST_BANDWIDTH_BPS</code>, summed and capped over <code>stageIds</code>&#39;s union; cost-only with <code>rawWaste</code> in <code>ms</code> when not stage-mappable. <code>stageIds</code> is narrowed the same way; falls back to the whole execution&#39;s stages when the flagged node(s) have no accumulator coverage.</td></tr><tr><td><code>underBroadcast</code></td><td>sql</td><td>modeled / cost-only</td><td><code>smallerSideBytes / BROADCAST_BANDWIDTH_BPS</code>, summed and capped over <code>stageIds</code>&#39;s union; cost-only with <code>rawWaste</code> in <code>ms</code> when not stage-mappable. <code>stageIds</code> is narrowed the same way; falls back to the whole execution&#39;s stages when the flagged node(s) have no accumulator coverage.</td></tr><tr><td><code>memoryUtilization</code></td><td>app</td><td>cost-only / informational-only</td><td>Three of the four variants report <code>rawWaste</code> in <code>mbSeconds</code>: <code>variant: &#39;wasteModel&#39;</code> passes through its own <code>wastedMBSeconds</code>; <code>&#39;idleCores&#39;</code> uses <code>idleRateFraction × allocatedMB × peakExecutors × appDurationSeconds</code>; <code>&#39;memoryBand&#39;</code> with <code>rule: &#39;heapOverProvisioned&#39;</code> uses <code>(allocatedBytes − heap) in MB × appDurationSeconds</code>. <code>&#39;memoryBand&#39;</code> with <code>rule: &#39;heapNearCapacity&#39;</code> is an OOM-risk signal rather than a waste, and the <code>dataUnavailable</code> shape has no inputs at all: both informational-only</td></tr><tr><td><code>utilization</code></td><td>app</td><td>cost-only</td><td><code>rawWaste</code> in <code>coreHours</code>: <code>(1 − utilizationFraction) × appDurationMs × totalCores / 3.6e6</code></td></tr><tr><td><code>coreLocality</code></td><td>app</td><td>cost-only</td><td><code>rawWaste</code> in <code>coreMs</code>: <code>nonLocalTaskCount × NETWORK_FETCH_PENALTY_MS</code></td></tr><tr><td><code>autoscalingChurn</code></td><td>app</td><td>cost-only</td><td><code>rawWaste</code> in <code>coreHours</code>: <code>shortLivedExecutorCount × EXECUTOR_STARTUP_OVERHEAD_MS / 3.6e6</code></td></tr><tr><td><code>configAudit</code></td><td>config</td><td>informational-only</td><td>a config-drift check standing alone; no waste formula</td></tr><tr><td><code>jobFailureRate</code></td><td>app</td><td>cost-only</td><td><code>rawWaste</code> in <code>coreHours</code>: <code>failedJobCount × avgJobDurationMs / 3.6e6</code></td></tr><tr><td><code>cachingOpportunity</code></td><td>app</td><td>cost-only</td><td><code>rawWaste</code> in <code>ms</code>: <code>totalReadBytes / RE_READ_THROUGHPUT_BPS</code></td></tr><tr><td><code>cacheUtilization</code></td><td>app</td><td>cost-only</td><td><code>rawWaste</code> in <code>ms</code>: uncached-or-spilled bytes <code>/ RE_READ_THROUGHPUT_BPS</code>, where the never-cached partitions&#39; bytes are extrapolated from the cached partitions&#39; own average size (<code>memorySize + diskSize</code>, over <code>numCachedPartitions</code>), plus <code>diskSize</code> again for the already-cached-but-on-disk partitions&#39; own re-read cost</td></tr><tr><td><code>stageFailed</code></td><td>stage</td><td>informational-only</td><td>no waste formula</td></tr><tr><td><code>failures</code></td><td>stage</td><td>informational-only</td><td>no waste formula</td></tr><tr><td><code>incompleteRun</code></td><td>app</td><td>informational-only</td><td>no waste formula</td></tr></tbody></table>',30)])])}const g=t(c,[["render",s]]);export{p as __pageData,g as default};
@@ -0,0 +1 @@
1
+ import{_ as o,o as t,c as a,a5 as s}from"./chunks/framework.DSg0KOwT.js";const u=JSON.parse('{"title":"State and History Server intake","description":"","frontmatter":{},"headers":[],"relativePath":"contributor-guide/architecture/state-and-history.md","filePath":"contributor-guide/architecture/state-and-history.md"}'),d={name:"contributor-guide/architecture/state-and-history.md"};function r(c,e,n,i,l,h){return t(),a("div",null,[...e[0]||(e[0]=[s('<h1 id="state-and-history-server-intake" tabindex="-1">State and History Server intake <a class="header-anchor" href="#state-and-history-server-intake" aria-label="Permalink to &quot;State and History Server intake&quot;">​</a></h1><h2 id="state-model" tabindex="-1">State model <a class="header-anchor" href="#state-model" aria-label="Permalink to &quot;State model&quot;">​</a></h2><p><code>src/store/store.ts</code> is a single Zustand store (<code>createStore</code> from <code>zustand/vanilla</code>, wrapped by a <code>useStore</code> hook). It holds all shared state; no component keeps a <code>useState</code> of its own for anything shared: <code>appModel: AppModel</code>, <code>catalog: Finding[]</code>, <code>activeFileId</code>, <code>sessionCache</code> (in-memory snapshot cache for instant file-switching, see <code>session-snapshot.ts</code>), <code>taskDataCache</code>, <code>parse: {pct, lines, etaMs}</code>, <code>status: &#39;idle&#39;|&#39;parsing&#39;|&#39;ready&#39;|&#39;error&#39;</code>, <code>errorMessage</code>, <code>theme</code>, <code>skippedLines</code> (malformed-JSON-line count from the parser&#39;s <code>done</code> payload).</p><p><code>src/store/useIngest.ts</code> is the only writer during a parse. It builds <code>createModelCallbacks</code>&#39; <code>onProgress</code>/<code>onDone</code>/<code>onError</code> handlers to call the store&#39;s setters directly (<code>setParse</code>, <code>setCatalog</code>, <code>setStatus</code>, <code>setSkippedLines</code>, ...). Components never talk to the worker: they use <code>useIngest()</code>&#39;s returned actions (<code>startLoad</code>/<code>startLoadFolder</code>/<code>startLoadFromUrl</code>/<code>pickRecent</code>/<code>getTaskData</code>/ <code>resetToDropZone</code>) and the store&#39;s read state.</p><p><code>resetModel()</code> empties <code>appModel</code>/<code>catalog</code>/<code>taskDataCache</code>/<code>skippedLines</code> on every new parse or reset-to-drop-zone, and bumps <code>modelResetCount</code>. That counter has no setter of its own; only <code>PlanGraphRoute.tsx</code>&#39;s <code>store.subscribe</code> reads it, to evict the plan-graph model memo cache (see <a href="./drill-down.html#plan-graph-view">Plan graph view</a>).</p><h3 id="finding-filter-state" tabindex="-1">Finding filter state <a class="header-anchor" href="#finding-filter-state" aria-label="Permalink to &quot;Finding filter state&quot;">​</a></h3><p>The board-wide finding filter (impact band, raw <code>finding.type</code>, stage) lives outside the Zustand store, in <code>FindingFilterContext</code> (<code>src/view/FindingFilterContext.tsx</code>): a <code>createContext</code>+<code>useState</code><code>FilterSelection</code> (<code>src/view/finding-filter.ts</code>, three <code>Set</code>s) that every widget filters <code>catalog</code> through via <code>filterFindings</code>. It is seeded from the URL&#39;s <code>impact</code>/<code>type</code>/<code>stage</code> query params on mount and nowhere else, so a reload with no params gives the unfiltered board. Every change writes those params back with <code>history.replaceState</code>, never <code>push</code>, so filtering doesn&#39;t grow Back history. A <code>popstate</code> listener re-seeds the selection from the URL, so Back and forward re-apply filters.</p><p>Switching files resets the selection to empty, keyed on the stable file id rather than <code>catalog</code>, so a same-file catalog refresh keeps the active filter. A page reload re-reads the address bar, so deep-linked filtered URLs still restore. Filters are seeded only from the URL and never persisted anywhere else (no localStorage, no restore-as-default): a filtered view silently becoming the default on reload would risk hiding findings from a user who didn&#39;t realize a filter was still active, so a plain reload with no filter params is always the unfiltered board.</p><h3 id="recent-files-vs-session-cache" tabindex="-1">Recent files vs. session cache <a class="header-anchor" href="#recent-files-vs-session-cache" aria-label="Permalink to &quot;Recent files vs. session cache&quot;">​</a></h3><p>Two mechanisms cover reopening a file, at different lifetimes.</p><p><code>sessionCache</code> (in the Zustand store, see <a href="#state-model">State model</a>) is in-memory and per-session: it makes switching between files already loaded in the current tab instant, and it is gone on reload.</p><p>Recent files (<code>src/recent-files.ts</code>, consumed by <code>src/view/useRecentFiles.ts</code>) is IndexedDB-backed and cross-session. It persists each file&#39;s <code>FileSystemFileHandle</code> plus light metadata (name, size, <code>lastModified</code>, app name, issue count, <code>lastOpenedAt</code>), capped at 10 entries with the oldest evicted past the cap, so a file can be reopened after a full browser restart, pending the browser re-granting permission on the handle. Picking a recent entry re-parses from the handle; no parsed model is ever persisted.</p><h2 id="run-comparison" tabindex="-1">Run comparison <a class="header-anchor" href="#run-comparison" aria-label="Permalink to &quot;Run comparison&quot;">​</a></h2><p><code>src/run-comparison.ts</code> is the whole A/B engine. The entry point <code>compareRuns(baseline, candidate)</code> takes two <code>{ label, snapshot }</code> run records (each <code>snapshot</code> a normalized model: <code>app</code>, <code>stages</code>, <code>sql</code>, <code>catalog</code>, <code>executors</code>) and returns one plain object the view renders. It runs on already-parsed snapshots, with no worker involved.</p><ul><li><code>stageIdentity(stage, snapshot)</code> is a run-independent key: <code>normalizeStageName</code> (lowercased, digit-runs and long hex ids collapsed to <code>#</code>) joined with the stage&#39;s SQL-execution plan identity. That identity is scoped to only the plan nodes this stage&#39;s tasks were attributed to (<code>node.stageIds</code>), not the whole tree, so two stages sharing one SQL execution (e.g. a self-join&#39;s two Exchange stages) don&#39;t collapse onto one identity; it falls back to a bottom-up structural fingerprint of the whole resolved <code>planTree</code> (<code>planTreeIdentity</code>, <code>normalizeDetail</code>-normalized: the same normalizer <code>cachingOpportunity</code> uses in <code>src/detectors.ts</code>) when a stage has no such attribution. <code>matchStages(baseSnap, candSnap)</code> indexes each run by that identity and pairs identities that map to exactly one stage on both sides. An identity colliding equally on both sides (same count) is also paired, positionally by sorted stage id: exact when comparing a run against itself (every stage matches itself), a best-effort guess otherwise (two unrelated same-named stages with no SQL/attribution could get cross-paired). Collisions are still recorded in <code>collisionIdentities</code> even when resolved this way; a differing count leaves them there unpaired. <code>coverage</code> reports the matched fraction.</li><li><code>metricDeltas</code> computes whole-run aggregate deltas (wall-clock, spill, task skew p95, failed-task rate, GC, I/O bytes, executor count, ...) as plain sums over all stages, deliberately not gated on stage matching, since matching is unreliable on real logs. Each metric carries a <code>direction</code> (improvement/regression/unchanged) and an <code>unavailableReason</code> when a side lacks the field.</li><li><code>findingsDelta</code> tallies each run&#39;s <code>catalog</code> by <code>(rule × impact band)</code> and reports <code>introduced</code> vs <code>resolved</code> categories: a count diff, also matching-free.</li><li>Only the per-stage skew deltas (<code>stageSkewDeltas</code>) and the pinned-stage panel consume the <code>matchStages</code> pairs, so low match coverage degrades those two surfaces without invalidating the aggregate deltas.</li></ul><p><code>compareRuns</code> also flags <code>confidence: &#39;low&#39;</code> when the two app names differ, or independently when matched stage coverage falls below 0.5 (<code>LOW_COVERAGE_THRESHOLD</code>; either condition alone is enough, both are weak signals, not a hard gate). The view lives in <code>src/view/RunComparison.tsx</code> (the comparison page), <code>CompareLanding.tsx</code> (the two-slot Run A / Run B intake off the landing), and <code>PinnedStageDeltas.tsx</code> (the manual per-stage pinning panel fed by <code>baseStages</code>/<code>candStages</code>).</p><h2 id="history-server-intake-and-recovery" tabindex="-1">History Server intake and recovery <a class="header-anchor" href="#history-server-intake-and-recovery" aria-label="Permalink to &quot;History Server intake and recovery&quot;">​</a></h2><p><code>DropZone</code> keeps the History Server disclosure, Base URL, Application ID, optional Attempt ID, validation/touched state, recoverable SHS error, and local-server reachability in mounted React state rather than Zustand. The local browser-first path is still the default: <strong>Choose file</strong> loads a single event log, while <strong>Choose rolling-log folder</strong> accepts only an <code>eventlog_v2_*</code> directory and directs a rejected folder back to the file picker.</p><p>The three text fields also mirror to <code>window.localStorage</code> (<code>shuffle-works-shs-base-url</code>/<code>-app-id</code>/<code>-attempt-id</code>), read back as each <code>useState</code>&#39;s initializer, so a returning visitor&#39;s values survive a reload; storage access is wrapped in try/catch and silently ignored when unavailable, matching <code>store.ts</code>&#39;s <code>initialTheme</code>/<code>initialWidgetDensity</code> pattern. Both this disclosure&#39;s toggle and the <strong>Other sources</strong> toggle show a chevron (<code>ChevronDownIcon</code>/<code>ChevronUpIcon</code>) that flips with <code>aria-expanded</code>, so the open/closed state has a visual signal beyond the attribute.</p><p>On mount (skipped in <code>compact</code> mode), <code>DropZone</code> probes reachability with an empty, short-timeout <code>fetch(&#39;/shs-proxy&#39;)</code>: a 400 means <code>validateShsRequest</code> (<code>packages/core/src/proxy.js</code>) rejected the empty request synchronously, which only happens when a local server is actually routing that path, so it flips <code>shsReachable</code> to <code>true</code>. A network error, a 404 (static deploy, no such route), or a probe still in flight all leave <code>shsReachable</code> at its default <code>false</code>, so nothing changes on screen after paint unless the server is confirmed present. When <code>shsReachable</code> is <code>true</code>, the landing page shows a neutral callout above the <strong>Other sources</strong> disclosure pointing the user at it; the disclosure itself doesn&#39;t move or auto-expand.</p><p>The collapsed <strong>Fetch from Spark History Server</strong> disclosure requires local-server mode, a reachable History Server, and a supported base application ID: <code>application_&lt;timestamp&gt;_&lt;id&gt;</code>, <code>local-&lt;timestamp&gt;</code>, or <code>app-&lt;identifier&gt;</code>. <code>server/lib/shs-request.js</code> trims and validates the three request fields, accepts only absolute credential-free <code>http:</code>/<code>https:</code> base URLs without a query or fragment, preserves a reverse-proxy path prefix, and canonicalizes the base URL to one trailing slash. The optional attempt is a separate path-safe identifier; neither identifier can contain a path separator. The shared helper builds the encoded <code>/shs-proxy</code> request and the encoded <code>api/v1/applications/&lt;app&gt;[/&lt;attempt&gt;]/logs</code> upstream path from that normalized object only.</p><p>The optional Node server is loopback-only: a narrow CORS proxy, not a general or hosted proxy. It validates the same request contract, sends no credentials, follows no upstream redirects, and returns only stable safe error codes. It never forwards upstream response text, status details, locations, or credentials to the browser.</p><p>Routing preserves the recovery boundary: local file and folder failures use the existing page-level error route. A typed SHS failure instead resets the model to idle and returns to the still-mounted, expanded History Server disclosure, which retains its values and shows safe recovery guidance along with local file intake. During an SHS parse, the mounted intake shows progress in place; successful completion follows the normal dashboard route.</p>',23)])])}const g=o(d,[["render",r]]);export{u as __pageData,g as default};
@@ -0,0 +1 @@
1
+ import{_ as o,o as t,c as a,a5 as s}from"./chunks/framework.DSg0KOwT.js";const u=JSON.parse('{"title":"State and History Server intake","description":"","frontmatter":{},"headers":[],"relativePath":"contributor-guide/architecture/state-and-history.md","filePath":"contributor-guide/architecture/state-and-history.md"}'),d={name:"contributor-guide/architecture/state-and-history.md"};function r(c,e,n,i,l,h){return t(),a("div",null,[...e[0]||(e[0]=[s("",23)])])}const g=o(d,[["render",r]]);export{u as __pageData,g as default};
@@ -0,0 +1 @@
1
+ import{_ as o,o as t,c as d,a5 as a}from"./chunks/framework.DSg0KOwT.js";const p=JSON.parse('{"title":"Widget rendering","description":"","frontmatter":{},"headers":[],"relativePath":"contributor-guide/architecture/widget-rendering.md","filePath":"contributor-guide/architecture/widget-rendering.md"}'),i={name:"contributor-guide/architecture/widget-rendering.md"};function n(c,e,r,s,l,g){return t(),d("div",null,[...e[0]||(e[0]=[a('<h1 id="widget-rendering" tabindex="-1">Widget rendering <a class="header-anchor" href="#widget-rendering" aria-label="Permalink to &quot;Widget rendering&quot;">​</a></h1><h2 id="widget-rendering-order-fixed-spec-§5" tabindex="-1">Render order (fixed, spec §5) <a class="header-anchor" href="#widget-rendering-order-fixed-spec-§5" aria-label="Permalink to &quot;Render order (fixed, spec §5) {#widget-rendering-order-fixed-spec-§5}&quot;">​</a></h2><p><code>src/view/Dashboard.tsx</code>&#39;s <code>FilteredBoard</code> renders inside a <code>&lt;main&gt;</code> that opens with <code>FindingFilterBar</code> and (only when an active filter empties both finding streams) <code>NoMatchBanner</code>, then a single <code>Scorecard</code> strip, then (when the active filter doesn&#39;t empty the board) a two-tab <code>Tabs</code> (<code>src/components/ui/tabs.tsx</code>, a base-ui primitive): <strong>Findings</strong> and <strong>Full app report</strong> (2026-09-03 tabbed-impact-band-board redesign, replacing the prior three stacked sections: All recommendations, Suggested Improvements, Full app report, with a merged, impact-grouped Findings tab and an always-reachable Full app report tab). Base-ui <code>Tabs</code> fully unmount the inactive <code>TabsContent</code> panel rather than hiding it: a widget mounted only in Findings (every routeable <code>REGISTRY</code> widget; see &quot;First investigation routing&quot; below) is not in the DOM at all while Full app report is active, and remounts fresh, with its own state reset, when the user switches back.</p><p><code>region</code> on <code>RegistryEntry</code> (<code>src/view/detector-registry.tsx</code>) is read again, but only to decide whether a widget always mounts: <code>isAlwaysMountedType()</code> flags the one <code>reference</code>-region type still carved out as an always-mounted exception, <code>coreLocality</code> → <code>CoreUsageArea</code>. The other <code>reference</code>-region types stay ordinary finding-gated instead: <code>cacheUtilization</code>, <code>memoryUtilization</code>, and <code>utilization</code> are a product decision: a clean run on any of them isn&#39;t evidence worth surfacing unconditionally, so each collapses to a plain clean-check line like any other detector. The Findings tab (<code>src/view/widgets/ImpactBoard.tsx</code>) groups its content into three tiers, same as the retired Suggested Improvements section did, just re-sliced by impact band instead of living as one flat active grid. The three tiers: impact-banded rows and cards for every <code>REGISTRY</code> component and every recommendation row with at least one finding; a small always-visible grid holding just the one exception above (mounted unconditionally from <code>appModel</code> regardless of finding state, below the impact bands); and a collapsed &quot;Clean checks&quot; disclosure covering every remaining type with zero findings, built per detector <em>type</em> (<code>Object.keys(REGISTRY)</code>). Full app report stays structural-only (see &quot;ReferenceSection&quot; below).</p><p>Tags carry their own docs links; there is no separate legend widget. <code>TagBadge</code> (<code>src/view/ImpactBadge.tsx</code>) resolves its own tooltip. When its type resolves a single known documentation anchor (<code>docAnchorForType</code>, <code>src/view/finding-tag-help.ts</code>), the pill itself links into the docs panel, and (density <code>advanced</code> only) a second icon link opens that tag&#39;s entry in <code>docs-site/user-guide/understanding-findings.md</code> (<code>findingGuideUrl</code>, <code>packages/core/src/docs-site-config.ts</code>) as a plain new-tab navigation. When a type has no vendor-doc anchor (e.g. <code>incompleteRun</code>), the pill links straight to that same guide entry instead of rendering as inert text, and the second icon link is skipped as redundant. Either way, a linked pill gets a visible dotted underline, not just a hover tooltip. <code>TagBadge</code>&#39;s <code>plainBadge</code> prop suppresses both links together; a Findings-tab recommendation row is the one place that never sets it (see below), since its badge cell isn&#39;t nested inside a button.</p><h3 id="findings-tab" tabindex="-1">Findings tab <a class="header-anchor" href="#findings-tab" aria-label="Permalink to &quot;Findings tab&quot;">​</a></h3><p>Rendered by <code>FilteredBoard</code>&#39;s <code>TabsContent value=&quot;findings&quot;</code> (<code>src/view/widgets/ImpactBoard.tsx</code>). Merges what used to be two stacked sections, All recommendations and Suggested Improvements, into one impact-ranked board. The recommendation-rollup logic lives in <code>src/view/widgets/FixTheseFirst.tsx</code> (kept as its own file and its own directly-testable exports, but no longer rendered as a standalone page section by <code>Dashboard.tsx</code>) and the active-widget logic lives in <code>src/view/widgets/Alerts.tsx</code> (same: kept, no longer rendered standalone). <code>ImpactBoard</code> calls <code>FixTheseFirst.tsx</code>&#39;s exported <code>useFixTheseFirstData</code> (eligible findings, rollup groups, the top triage target) and <code>Alerts.tsx</code>&#39;s exported <code>computeActiveWidgets</code> (the ranked active-<code>REGISTRY</code> list) to get the exact same data these two retired sections used to render independently, then regroups both by impact band via <code>groupImpactBand</code> (<code>FixTheseFirst.tsx</code>: a rollup group&#39;s representative member&#39;s impact band, the same finding whose impact band its own badge already shows) and each active widget&#39;s own <code>worstImpactBand</code>.</p><p>Eligible findings for the rollup are <code>catalog</code> ∪ <code>configFindings</code>, filtered to a <code>REGISTRY</code>-mapped type, with <code>incompleteRun</code> (a pipeline-completeness caveat, not an addressable fix; see the spec&#39;s <code>fixEffort</code> table) and <code>memoryUtilization</code>&#39;s <code>memoryBand</code>/<code>dataUnavailable</code> variant (a missing-evidence caveat already covered by Evidence availability&#39;s own <code>executorMetrics</code> entry, <code>packages/core/src/evidence-availability.ts</code>) both explicitly excluded. Findings are grouped strictly by <code>finding.type</code> via <code>buildRecommendationRollup</code> (<code>packages/core/src/recommendation-rollup.ts</code>, further split within a type by impact kind and, for <code>resource</code>, unit: never merged across types); each resulting group becomes one row: a type with exactly one finding renders that finding directly, a type with more than one collapses into a summary row. Within an impact band, group order comes from <code>buildRecommendationRollup</code>&#39;s own sort: <code>time</code> groups (a real <code>wallClock</code> claim) first, ranked among themselves by their union-capped <code>recoverableMsHigh</code> descending; then <code>resource</code> groups (<code>rawWaste</code> but no <code>wallClock</code>); then <code>count</code> groups (neither), both of the latter two ranked by worst impact band, never by their incomparable raw magnitudes. A summary row&#39;s tag and dot, plus its text, all come from the group&#39;s own highest-impact member (via the same three-tier comparator, computed locally in <code>FixTheseFirst.tsx</code>); the trailing stat depends on the group&#39;s kind (<code>×N · &lt;time&gt; recoverable</code>, <code>×N · resource-cost projection</code>, or a plain per-impact-band tally). Clicking it expands straight to the group&#39;s full, impact-ranked list, with no intermediate &quot;worst-K&quot; step, paginated at 10 rows per page (<code>data-testid=&quot;fix-these-first-group-row&quot;</code>; no pager renders for a group of 10 or fewer findings; it appears once a group exceeds 10). A band&#39;s rollup rows render as a headerless three-column <code>Table</code> (<code>src/components/ui/table.tsx</code>): every individual row, whether shown directly or inside an expanded group, is a <code>TableRow</code> (<code>data-testid=&quot;fix-these-first-row&quot;</code>, <code>data-finding-type</code>) with three <code>TableCell</code>s: the impact dot + ALL-CAPS tag as a real <code>TagBadge</code> (not <code>plainBadge</code>: nothing wraps it, so its own docs links stay real <code>&lt;a&gt;</code>s, same as everywhere else on the board); a text block inside its own nested <code>&lt;button&gt;</code> (a short imperative action label, e.g. &quot;Reduce shuffle size&quot;, from <code>findingActionLabel</code> (<code>src/view/finding-action-label.ts</code>), over the finding&#39;s own full <code>recommendation</code> sentence in smaller muted text, both wrapping rather than truncating); and a right-aligned monospace stage reference + impact figure (e.g. <code>St.49 · 20.1s</code>, via <code>ImpactEstimate.tsx</code>&#39;s shared <code>formatWallClockRange</code>/<code>formatRawWaste</code>). That inner button, not the row, is the click target: it routes via <code>selectTriageTargetForFinding</code> (<code>src/view/triage-target.ts</code>), the same per-finding resolver Stage Summary Table&#39;s own control uses (see &quot;First investigation routing&quot; below); a <code>TypeGroupRow</code>&#39;s own inner button toggles its expand state instead (<code>aria-expanded</code>) and its expanded findings render as further <code>TableRow</code>s indented one badge-cell notch to read as the group&#39;s sub-list.</p><p>Each impact band (<code>ImpactBoard.tsx</code>&#39;s own <code>ImpactGroup</code>, one call per entry of <code>IMPACT_BAND_ORDER_LIST = [&#39;critical&#39;, &#39;warning&#39;, &#39;info&#39;]</code>) is a <code>&lt;section aria-label=&quot;Critical&quot; | &quot;Warning&quot; | &quot;Info&quot;&gt;</code> with an <code>&lt;h3&gt;</code> heading, and renders nothing (not even the heading) when it has neither a rollup row nor an active widget: a run with no critical findings has no &quot;Critical&quot; heading or section at all. Inside a band, rollup rows render first as the headerless <code>Table</code> described above, followed by that band&#39;s active <code>REGISTRY</code> widget cards (<code>computeActiveWidgets</code>&#39;s ranked list, filtered to this impact band) in their own <code>WidgetGrid</code>: every one of <code>orderedWidgets()</code>&#39;s deduped <code>REGISTRY</code> components <em>except</em> the one always-mounted one below, with at least one finding in <code>catalog</code> ∪ <code>configFindings</code>. Within a band, active widgets keep <code>orderedWidgets()</code>&#39;s own <code>action</code>-region-first, ascending-<code>DETECTORS</code>-order tiebreak. <code>cacheUtilization</code>, <code>memoryUtilization</code>, and <code>utilization</code> are <code>reference</code>-region types but aren&#39;t always-mounted exceptions, so a Cache Storage, Memory Utilization, or Executor Utilization card with an active finding surfaces in its own impact band like any other active widget.</p><p>Below the impact bands, <code>Alerts.tsx</code>&#39;s exported <code>AlwaysVisibleAndCleanChecks</code> (shared verbatim with the retired standalone <code>Alerts</code> component) renders the same two tiers it always did, now living outside the impact-band grouping entirely rather than as this section&#39;s second and third tier: a small always-visible grid holding just Core Usage by Locality (<code>coreLocality</code>, resolving to <code>CoreUsageArea</code>), mounted unconditionally from <code>appModel</code> regardless of finding state (<code>alwaysMountedWidgets()</code>/<code>isAlwaysMountedType()</code>), carrying its own impact-band indicator when a finding is active instead of collapsing to a clean-check line on a clean run; and a collapsed &quot;Clean checks&quot; disclosure of <code>CleanCheckRow</code> lines (<code>src/view/widgets/CleanCheckRow.tsx</code>: label, the threshold it was measured against via <code>getThresholdSummary</code>, and &quot;No fix needed.&quot;) built per detector <em>type</em> (every <code>REGISTRY</code> key except that one always-mounted key): a clean run lands <code>cacheUtilization</code>, <code>memoryUtilization</code>, and <code>utilization</code> here too, same as any ordinary action-region type. Caching Opportunities, Config Audit, and the four split Plan Advisor widgets (Redundant Plan Subtree, Excessive Small Files, Missed Broadcast Join, Oversized Broadcast Join) render through the ordinary active/clean paths above (see <a href="./board-widgets.html#board-widgets-beyond-the-fixed-six">Board widgets beyond the fixed six</a>). One consequence of this always-visible grid sitting below every impact band: a <code>critical</code>-band <code>coreLocality</code> finding still renders in that lower grid, below the <code>info</code>-band widgets above it, a deliberate tradeoff the spec accepted in exchange for never losing the widget on a clean run, not a ranking bug.</p><h3 id="per-widget-list-sort-mode" tabindex="-1">Per-widget list sort mode <a class="header-anchor" href="#per-widget-list-sort-mode" aria-label="Permalink to &quot;Per-widget list sort mode&quot;">​</a></h3><p><code>src/view/impact-sort.ts</code> (<code>sumWallClockLow</code>, <code>hasSortableImpact</code>, <code>byImpactDesc</code>, <code>stageIdOf</code>, <code>minStageId</code>, <code>byStageAsc</code>) and <code>src/view/SortModeToggle.tsx</code> are a shared, opt-in pair a Findings-tab active widget can use to let its own expanded, multi-row list default to potential-savings order (<code>wallClock.low</code> descending, the guaranteed-floor bound, not the optimistic <code>high</code>) instead of stage number. Each of the widgets below holds local <code>sortMode</code> state (<code>SortMode</code>, <code>&#39;impact&#39; | &#39;stage&#39;</code>, default <code>&#39;impact&#39;</code>) and renders a <code>SortModeToggle</code> in the <code>WidgetCard</code> <code>badges</code> slot, so it sits on the title row itself rather than floating in the body: a button alongside the widget&#39;s tag badges (a <code>WidgetCard</code> header renders <code>badges</code> beside the title heading, never inside another control). Every one of the former combined widgets&#39; split single-type widgets carries the same <code>canToggleSort</code> check its parent did, so the toggle only actually renders for a type whose finding can carry a wall-clock estimate: <code>Skew.tsx</code>, <code>TinyTask.tsx</code> (<code>stageShape</code>&#39;s own rules never produce one, so <code>StageShape.tsx</code> carries the same check but it never fires), <code>ShuffleIO.tsx</code> (narrowed to <code>shuffle</code>), <code>PartitionSizing.tsx</code>, <code>Spill.tsx</code>, <code>GcPressure.tsx</code> (both its high-GC and low-GC sections, one shared toggle), <code>RetryWaste.tsx</code> (its siblings <code>StageFailed.tsx</code>/<code>TaskFailures.tsx</code> carry the same check, but <code>stageFailed</code>/<code>failures</code> are <code>estimateMethod: &#39;none&#39;</code>, so it never fires there either), <code>SlowHost.tsx</code>, <code>StageSlowness.tsx</code>, <code>Straggler.tsx</code>, <code>SpeculationWaste.tsx</code>, <code>ColdStart.tsx</code> (five widgets now, one per type, each sorting only its own flat issue list; the old cross-type &quot;coldStart sinks to the bottom under By stage&quot; note no longer applies now that each type has its own single-type list), and <code>DuplicatePlanSubtree.tsx</code>/ <code>SmallFiles.tsx</code>/<code>UnderBroadcast.tsx</code>/<code>OverBroadcast.tsx</code> (each reorders its own list by <code>stageIdOf</code>, the lowest stage id its finding touches; the old cross-type <code>minStageId</code> group-ordering tiebreak no longer applies, since each type is its own widget now). Its <code>alternateOrderLabel</code> prop is required and every caller passes &quot;By stage&quot;: stage number is the one axis every one of these widgets&#39; items can be compared on, unlike the impact/raw-metric order this pattern replaced (dropped: comparing findings by impact band ranks them by how bad they are, not by how much fixing them would save, which is a worse default now that a real potential-savings figure exists to sort by instead). A per-stage detector&#39;s own <code>finding.stageId</code> is the sort key directly; a sql-scope finding that spans several stages (<code>duplicatePlanSubtree</code>, <code>smallFiles</code>, <code>underBroadcast</code>, <code>overBroadcast</code>, via <code>stageIds</code>) sorts by the lowest stage id it touches (<code>stageIdOf</code>). The toggle only renders when <code>hasSortableImpact</code> finds at least one wall-clock claim in the list; re-sorting a list with none would be a silent no-op. Both comparators return <code>0</code> when neither side has a comparable value, so those items keep their prior relative order (<code>Array.prototype.sort</code>&#39;s stability) rather than being shuffled. <code>FixTheseFirst</code>&#39;s own expanded group list (see &quot;Findings tab&quot; above) already sorted by impact before this pattern existed and does not use it; it has no stage-order toggle.</p><h3 id="gold-standard-row-expand-contract" tabindex="-1">Gold Standard row/expand contract <a class="header-anchor" href="#gold-standard-row-expand-contract" aria-label="Permalink to &quot;Gold Standard row/expand contract&quot;">​</a></h3><p>Every <code>REGISTRY</code> widget renders its content as N≥1 rows, using <code>Skew.tsx</code> as the reference implementation. A widget&#39;s data shape (chart, table, single app-scoped scalar) is never by itself a reason to skip this: the three named exceptions below are the only ones. Any further exception must get its own entry here, justified in the same PR that introduces it.</p><ul><li>Tier A (universal): <code>WidgetCard</code> chrome, an impact-band left-border + dot, ALL-CAPS <code>TagBadge</code>s, a <code>SortModeToggle</code> when <code>hasSortableImpact</code> is true, a <code>finding-anchor</code> ref on any focusable/deep-linkable unit, and an early <code>null</code> return on an empty catalog filter (except Core Usage by Locality, the one always-mounted widget that renders unconditionally from <code>appModel</code> via <code>alwaysMountedWidgets()</code>/<code>isAlwaysMountedType()</code> rather than early-returning on an empty catalog filter).</li><li>Tier B (the row/expand pattern): a row&#39;s collapsed state shows its core metric(s), <code>ImpactEstimate</code>, its recommendation text, any config-hint code snippet/list, and docs links, all unconditionally. Confidence and evidence are unconditional too, since the 2026-09 redesign: <code>RowStatusCluster</code> (<code>src/view/RowStatusCluster.tsx</code>) is a single, fixed-position pill combining both, replacing the old per-row <code>ExpandToggleButton</code> + <code>ConfidenceMarker</code> + <code>EvidenceLink</code> trio entirely. Every widget that carries one wraps it in <code>AdvancedOnly</code> (<code>src/view/AdvancedOnly.tsx</code>), so it only renders at the Advanced density tier: this is a content-visibility gate (Basic vs. Advanced), not a per-row collapse, and it&#39;s applied consistently across every adopting widget. There is nothing left to gate behind a per-row click for confidence or evidence in any widget. A widget-header <code>RowStatusCluster</code> (passed as <code>WidgetCard</code>&#39;s <code>statusBadge</code> prop) carries one further, unrelated gate on top of <code>AdvancedOnly</code>: <code>open &amp;&amp; statusBadge</code> (<code>WidgetCard.tsx:138,141</code>) keeps it out of the collapsed header, matching the comment on <code>statusBadge</code> itself (&quot;kept separate so tag/impact badges stay visible in the collapsed summary while a <code>RowStatusCluster</code>-style ... marker doesn&#39;t&quot;). This is the card&#39;s own expand/collapse, not a separate reveal-on-click built for the cluster, and it composes with density as an AND: a header cluster needs both Advanced tier and an expanded card. A row with neither renders no cluster at all (<code>RowStatusCluster</code> returns <code>null</code>). Confidence renders as plain, non-interactive text (&quot;<code>{confidence} confidence</code>&quot;, no further suffix) with a <code>title</code> + <code>aria-describedby</code> tooltip carrying the full <code>validationRequired</code> detail: never itself a click target. Evidence renders as the cluster&#39;s one real click target: a button whose visible text is just the evidence label (e.g. &quot;SQL plan&quot;) but whose accessible name always carries the &quot;Evidence: ...&quot; prefix via an explicit <code>aria-label</code>, regardless of what else sits nearby (the old <code>bare</code> prop/prefix distinction is gone: there&#39;s only one form now). Clicking it calls the same <code>revealEvidence</code> navigation the retired <code>EvidenceLink</code> used. <code>RowStatusCluster</code>&#39;s fixed slot is the row&#39;s own stage-pill/title line (<code>justify-between</code>, cluster right-aligned), the same position in every adopting widget, not floating with the recommendation text below. Lists longer than <code>VISIBLE_LIMIT</code> (6, <code>packages/core/src/format-utils.ts</code>) still get page-based navigation (Previous/Next, <code>usePagedRows</code>/<code>RowPagination</code>, <code>src/view/usePagedRows.ts</code>/<code>src/view/RowPagination.tsx</code>) instead of rendering unconditionally: the one sanctioned cap mechanism across every Tier B widget. A Tier B widget whose rows are also triage-routable (wires <code>useFindingAnchor</code>) must pass its <code>routeIndex</code> (the routed finding&#39;s position in the paginated list) into <code>usePagedRows</code>, or a deep-linked route to a finding beyond the first page silently degrades to focusing the widget&#39;s disclosure title instead of the row.</li></ul><p>The following widgets carry a <code>RowStatusCluster</code>, all Advanced-only: <code>Spill.tsx</code> (confidence only), <code>ConfigAudit.tsx</code>&#39;s widget-header control (evidence only, a widget-wide constant rather than per-finding data, since the evidence key doesn&#39;t vary per row; also the control shown in the widget&#39;s empty-findings/no-data states), <code>MemoryUtilization.tsx</code>&#39;s rows (confidence + evidence together; <code>ExecutorUtilization.tsx</code>&#39;s rows, split out of the same former combined widget, carry neither), each of the four split Plan Advisor widgets&#39; (<code>DuplicatePlanSubtree.tsx</code>, <code>SmallFiles.tsx</code>, <code>UnderBroadcast.tsx</code>, <code>OverBroadcast.tsx</code>) per-row control (evidence only) plus their own widget-header marker (confidence only, one marker per widget now that each is its own finding-type, replacing the old <code>PlanFindings.tsx</code> per-group heading), <code>CoreUsageArea.tsx</code>, <code>EfficiencyModel.tsx</code>, <code>PlanView.tsx</code> (three call sites), <code>ScalingSim.tsx</code> (two call sites), and <code>WastedCoreHours.tsx</code> (all confidence-only, standalone rather than per-finding-row: a card-header badge, a widget-level single marker, a plan-tree node/summary-row marker, or an &quot;unavailable data&quot; message: the same component, same visual language, regardless of where it sits); and, since <code>skew</code>/<code>straggler</code>/<code>gc</code> started disclosing their own unvalidated noise-floor thresholds (confidence only, per-row, no <code>evidenceKey</code> passed), <code>GcPressure.tsx</code>&#39;s rows, <code>Straggler.tsx</code>&#39;s rows, <code>CachingOpportunity.tsx</code>&#39;s rows (the <code>cachingOpportunity</code> detector scales <code>confidence</code> per finding via <code>cachingReuseConfidence</code>, so the badge sits next to each row&#39;s recommendation rather than as a single caveat below the table), and the shared <code>StageFindingGroup.tsx</code> row (adopted by <code>Skew.tsx</code>/<code>StageShape.tsx</code>/<code>TinyTask.tsx</code>, though today only <code>skew</code> findings actually carry a <code>confidence</code> field).</p><p>One named exception:</p><ul><li><code>Skew.tsx</code>/<code>StageShape.tsx</code>/<code>TinyTask.tsx</code>: their per-row histogram toggle (<code>ExpandToggleButton</code> + <code>useExpandableRow</code>) gates a lazily-fetched duration histogram, unrelated to confidence/evidence, so it was untouched by the 2026-09 redesign and untouched again when <code>RowStatusCluster</code> was later added alongside it in the shared <code>StageFindingGroup.tsx</code> row. The two controls coexist per row. <code>ExpandToggleButton</code> is now single-purpose (always the task-detail toggle) since these three (split out of the former combined <code>TaskSkew.tsx</code>) are its only remaining adopters.</li></ul><p>No toggle at all, unconditional recommendation/content, same as before this redesign (unaffected either way, since these widgets never carried confidence/evidence display in the first place): <code>IncompleteRun.tsx</code>, <code>ShuffleIO.tsx</code>, <code>PartitionSizing.tsx</code>, <code>StageFailed.tsx</code>, <code>TaskFailures.tsx</code>, <code>RetryWaste.tsx</code>, <code>ColdStart.tsx</code>, <code>SlowHost.tsx</code>, <code>StageSlowness.tsx</code>, <code>SpeculationWaste.tsx</code>, <code>ExecutorCountChart.tsx</code>, <code>ExecutorUtilization.tsx</code>, <code>JobFailures.tsx</code>, <code>CacheUtilization.tsx</code>, <code>AutoscalingChurn.tsx</code>, and the four split Plan Advisor widgets (<code>DuplicatePlanSubtree.tsx</code>, <code>SmallFiles.tsx</code>, <code>UnderBroadcast.tsx</code>, <code>OverBroadcast.tsx</code>), split out of, respectively, the former combined <code>ExecutorTimeline.tsx</code>, <code>Failures.tsx</code>, and <code>PlanFindings.tsx</code>, none of which carried a per-row toggle either.</p><p>Three named exceptions to specific pieces of the contract, not to the row wrapper or pagination, which still apply to all three:</p><ul><li><code>CoreUsageArea.tsx</code>&#39;s non-local-stage list is a derived stat breakdown (<code>computeCoreLocalityRatio</code>&#39;s <code>topStages</code>), not a <code>Finding[]</code>: there is no per-row finding or recommendation to reveal. Paginated like every other Tier B list, but with no per-row expand.</li><li><code>CacheUtilization.tsx</code>&#39;s RDD <code>&lt;Table&gt;</code> holds reference columns (name, storage level, partitions, memory, disk bytes) with no recommendation attached to any row. Paginated like every other Tier B list, but with no per-row expand. Its separate recommendation list below the table is full Tier B (unconditional, no toggle, per the paragraph above).</li><li><code>IncompleteRun.tsx</code> leads with no metric of its own: this app-scoped, single-finding widget has no natural lead figure distinct from its title/badge, unlike every other Gold Standard widget, which leads with a percentage, duration, or count.</li></ul><h3 id="referencesection" tabindex="-1">ReferenceSection <a class="header-anchor" href="#referencesection" aria-label="Permalink to &quot;ReferenceSection&quot;">​</a></h3><p>Rendered by <code>FilteredBoard</code>&#39;s <code>TabsContent value=&quot;full-report&quot;</code>, tab label &quot;Full app report&quot;. A plain <code>&lt;div&gt;</code>, structural-only, reading none of <code>REGISTRY</code>/<code>orderedWidgets()</code> at all. Its heading is a visually-hidden (<code>sr-only</code>) <code>&lt;h2&gt;Full app report&lt;/h2&gt;</code>, matching the tab&#39;s own label for the accessibility-tree heading outline without visually duplicating the tab text. No longer an <code>Accordion</code>: base-ui <code>Tabs</code> unmount this panel entirely while Findings is active (see &quot;Render order&quot; above), so there is no &quot;collapsed but present&quot; state left to model, and this content is simply not in the DOM until the user clicks the Full app report tab. Its exact order is WallClock → Timeline → Executor Count Over Time (<code>ExecutorCountChart.tsx</code>, the executor add/remove count chart extracted out of the former combined <code>ExecutorTimeline.tsx</code>; not driven by any finding, so it isn&#39;t a <code>REGISTRY</code> entry) → StageTable → a <code>WidgetGrid</code> holding Evidence availability → ETL Phase Attribution → What-If Executor Scaling → Compute Efficiency → Wasted Core-Hours → Core-Usage Distribution. Scorecard used to lead this section; it now renders once, above the tabs themselves, in <code>FilteredBoard</code> (<code>src/view/Dashboard.tsx</code>), so it stays visible regardless of which tab is active rather than living inside either one (a three-tile run-info row: Wall-clock, Efficiency, Wastage; see <a href="./board-widgets.html#board-widgets-beyond-the-fixed-six">Board widgets beyond the fixed six</a>). WallClock, Timeline, StageTable and every tile in the grid beside them (Evidence availability included) all render immediately and fully expanded. No detector-driven <code>REGISTRY</code> card renders in this section any more: Memory Utilization, Executor Utilization, Core Usage by Locality, and Cache Storage all moved to the Findings tab above (Cache Storage, Memory Utilization, and Executor Utilization only surface there when they have an active finding; Core Usage by Locality alone is always-mounted).</p><p>The Evidence availability card is the persistent, non-impact-band ledger <a href="./worker-protocol.html#evidence-availability-contract-v1">defined in the worker protocol</a>, not an alert or detector widget. An <code>Evidence: …</code> control appears only where a conclusion or unavailable report lens declares a relevant ledger dependency. <code>revealEvidence</code> (<code>src/view/EvidenceAvailabilityContext.tsx</code>) sets <code>referenceOpen</code> true, which <code>DashboardContent</code> (<code>src/view/Dashboard.tsx</code>) watches in a <code>useEffect</code> and translates into <code>setActiveTab(&#39;full-report&#39;)</code>, replacing what used to be &quot;expand the Reference accordion&quot;: mouse and keyboard activation switches to the Full app report tab, opens the ledger card, then focuses the referenced stable entry id (<code>evidence-availability-&lt;key&gt;</code>). The control explains evidence availability; it does not promise an unavailable signal would have produced a finding.</p><p><code>WidgetGrid</code> (both inside a Findings impact band and inside <code>ReferenceSection</code>): collapsed cards occupy one responsive column; opening a card expands it across the full row. Nested <code>WidgetCard</code>s are isolated from the parent grid state. A Findings-tab <code>FixTheseFirst</code> rollup row is a table row, not a <code>WidgetGrid</code> card, and carries no collapse/expand state of its own (a <code>TypeGroupRow</code> summary row does own its own local expand/pagination state). Full app report&#39;s non-<code>REGISTRY</code> tiles (ETL Phase Attribution, What-If Executor Scaling, Compute Efficiency, Wasted Core-Hours, Core-Usage Distribution) sit in their own <code>WidgetGrid</code> alongside Evidence availability, but that grid has no <code>widgetId</code> wired to any card (see &quot;First investigation routing&quot; below), so nothing there is ever a route destination.</p><p>Firm constraint: <code>orderedWidgets()</code> (in <code>src/view/detector-registry.tsx</code>) still runs the single <code>DETECTORS</code>-ascending sort exactly as <code>dashboard-renderer.js</code> used to, and still sorts <code>action</code> components ahead of <code>reference</code> ones within that one list; the component-identity dedup it used to need is gone now that every <code>REGISTRY</code> entry maps to its own unique component (2026-09 widget/finding-type 1:1 mapping redesign). <code>cacheUtilization</code>, <code>memoryUtilization</code>, and <code>utilization</code> (<code>reference</code>-region) can still reach the active grid alongside <code>duplicatePlanSubtree</code> (<code>action</code>-region) and any other active finding, since none of them is the one always-mounted exception filtered out before that grid. <code>computeActiveWidgets</code> (<code>Alerts.tsx</code>) is still its main consumer, now filtered through <code>isAlwaysMountedType()</code> to exclude that one always-mounted component; the clean-check list bypasses <code>orderedWidgets()</code> entirely, iterating <code>Object.keys(REGISTRY)</code> per type instead (see &quot;Findings tab&quot; above). <code>DETECTORS</code>&#39; own array order and iteration, plus its cross-detector <code>suppressWhen</code> logic (e.g. <code>stageSlowness</code> deferring to <code>slowHost</code>, see <a href="./detector-contract.html#detector-contract">Detector contract</a>), live entirely in <code>packages/core/src/detectors.ts</code>/<code>packages/core/src/analyzer.ts</code>, untouched by this redesign. Which React component each finding type resolves to is still registry-owned.</p><h3 id="first-investigation-routing" tabindex="-1">First investigation routing <a class="header-anchor" href="#first-investigation-routing" aria-label="Permalink to &quot;First investigation routing&quot;">​</a></h3><p>The dashboard has a view-only per-finding route, not a single first-investigation pick any more: every Findings-tab recommendation row (rendered by <code>FixTheseFirst.tsx</code>&#39;s row components inside <code>ImpactBoard</code>) and the visible-finding control in Stage Summary (<code>StageTable.tsx</code>) each independently resolve and request their own target. <code>detector-registry.tsx</code> owns each routeable finding type&#39;s stable <code>widgetId</code>, widget title, and finding label; identity and copy come from there, not from runtime position or raw detector type. Eligible targets have a mapped, routeable registry entry and a non-empty trimmed recommendation (<code>targetForFinding</code>, <code>src/view/triage-target.ts</code>). The target&#39;s widget still renders every affected stage in its local order.</p><p><code>src/view/triage-target.ts</code> also still exports <code>selectTriageTarget</code>, the whole-catalog &quot;single highest-potential-savings finding&quot; picker that used to back a start-here callout and each per-tag headline tile in the retired <code>ProblemHeadlines.tsx</code> (potential savings first via <code>impactEstimate.wallClock.high</code>, then <code>orderedWidgets()</code> widget order, then catalog order; impact band plays no part). Ranking by potential savings replaced the earlier severity-first routing once the occupancy-weighted impact estimator gave every finding a real, comparable <code>impactEstimate.wallClock</code> figure: severity-first could point the single &quot;start here&quot; callout at a <code>skew</code>/<code>straggler</code> finding ranked <code>critical</code> on a ratio basis while its occupancy-clipped recoverable time was near zero, passing over a lower-severity finding with an order-of-magnitude larger real recoverable-time estimate right next to it. Nothing in the app calls <code>selectTriageTarget</code> any more: <code>FixTheseFirst</code> already ranks every eligible finding by impact magnitude, so a separate single-winner pick has no call site left. <code>tests/view/triage-target.test.ts</code> still covers it directly.</p><p>The route is re-derived from the current catalog before each asynchronous step. Its identity is object-reference equality against the current catalog (<code>catalog.includes(finding)</code> in <code>selectTriageTargetForFinding</code>, <code>src/view/triage-target.ts</code>) rather than a derived key: a new parse yields new finding object references, so reference presence alone detects staleness. It is not persisted and does not extend the core <code>Finding</code> contract; cross-session consumers (exports, future URL-restored state) use the core <code>Finding.id</code> instead (see <a href="./worker-protocol.html#finding-identity">Finding identity</a>).</p><p><code>Dashboard</code> owns disclosure and navigation. Every routeable <code>REGISTRY</code> widget now lives in the Findings tab (no <code>REGISTRY</code> widget renders inside Full app report any more; see &quot;Render order&quot; above), so <code>requestRoute</code> (<code>DashboardContent</code>, <code>src/view/Dashboard.tsx</code>) starts with an unconditional <code>setActiveTab(&#39;findings&#39;)</code>, whether the request came from a control already on that tab or from Full app report&#39;s own Stage Summary route link. That tab switch can unmount and remount the whole Findings subtree in the same commit as the route landing (base-ui <code>Tabs</code> fully unmounts the inactive panel; see &quot;Render order&quot; above), which two routing paths have to account for: <code>reportWidgetOpen</code> no longer clears a route just because the target <code>WidgetGridItem</code> hasn&#39;t registered yet in this commit (a child <code>WidgetCard</code> reports its own open state before its parent <code>WidgetGridItem</code>&#39;s registration effect runs on a fresh mount, so treating &quot;not registered yet&quot; as &quot;stale&quot; would drop the route before the parent even gets a chance to register it); and a paginated widget that jumps its own page to reveal a routed row (<code>GcPressure.tsx</code>, <code>StageFailed.tsx</code>, <code>TaskFailures.tsx</code>, <code>RetryWaste.tsx</code>) computes that jump during render, not only in a post-commit effect, so the row exists in the DOM in time for this same commit&#39;s anchor lookup instead of one commit later.</p><p>After the target card commits open, its wrapper scrolls with a top margin below the fixed header and its disclosure button receives visible, temporary route focus (using instant scrolling for reduced motion). Requests use a latest-request-wins token (<code>requestTokenRef</code>) that guards every step and revalidates; a stale, missing, unmounted, changed-catalog, or changed-active-file target cancels without redirecting, scrolling, or moving focus. Widget registration probes survive StrictMode double-mounts via an <code>unregistrationCheck</code> that re-validates before clearing.</p><p>This route deliberately does not include Config Audit, detector/parser or threshold changes, Stage Detail routing, or generic badges, tags, chips, or StagePills as controls.</p><h3 id="widget-placement" tabindex="-1">Widget placement <a class="header-anchor" href="#widget-placement" aria-label="Permalink to &quot;Widget placement&quot;">​</a></h3><p>The Plan Explorer (<code>src/view/widgets/PlanExplorer.tsx</code>) is embedded inside a flagged stage&#39;s own row, not as a standalone widget. Neither Task Skew nor Spill embed it any more: both dropped their own <code>PlanExplorer</code> trigger entirely (not merely re-gated behind a toggle), since the same stage&#39;s plan is already reachable, unconditionally, from <code>StageDetailDialog</code>. Shuffle I/O is <code>PlanExplorer</code>&#39;s only remaining call site: it has no row-expansion toggle of its own, so its trigger renders as a direct sibling in the row body.</p><p>Plan detection logic belongs in a <code>packages/core/src/detectors.ts</code> entry (<code>scope:&#39;sql&#39;</code>) consuming <code>planTree</code>. <code>packages/core/src/plan-summary.ts</code> is display-only summarization; do not add detection heuristics there.</p><p><code>StageHeader.tsx</code> (<code>StagePill</code> plus a <code>&lt;span&gt;</code> naming the stage) vs. bare <code>StagePill</code>/<code>StagePillGroup</code> is a deliberate row-density choice, not an inconsistency: <code>StageHeader</code> is for single-stage-per-row detail widgets where naming the stage adds context (<code>SlowHost.tsx</code>, <code>StageSlowness.tsx</code>, <code>Straggler.tsx</code>, <code>SpeculationWaste.tsx</code>, <code>GcPressure.tsx</code>, <code>Skew.tsx</code>, <code>StageShape.tsx</code>, <code>TinyTask.tsx</code>, <code>PartitionSizing.tsx</code>), while bare <code>StagePill</code>/<code>StagePillGroup</code> is for compact, multi-row lists where many stages appear per widget (<code>StageFailed.tsx</code>, <code>TaskFailures.tsx</code>, <code>RetryWaste.tsx</code>, the four split Plan Advisor widgets, <code>Spill.tsx</code>, and <code>StageTable.tsx</code>). <code>ShuffleIO.tsx</code> uses both in different parts of its own row, which is fine: it isn&#39;t a violation of the convention above.</p><p>Stage Summary Table (<code>src/view/widgets/StageTable.tsx</code>) defaults to <strong>problem stages</strong>, with a header toggle to <strong>top 10 by duration</strong>. It renders inside the Full app report tab, not as a standalone board section.</p><h3 id="full-render-sequence" tabindex="-1">Full render sequence <a class="header-anchor" href="#full-render-sequence" aria-label="Permalink to &quot;Full render sequence&quot;">​</a></h3><p>Top to bottom, in <code>Dashboard.tsx</code>&#39;s <code>FilteredBoard</code>:</p><ol><li><code>FindingFilterBar</code> (plus <code>NoMatchBanner</code> when the active filter empties both finding streams).</li><li><code>Scorecard</code>: a three-tile run-info row (Wall-clock, Efficiency, Wastage), rendered once regardless of which tab is active.</li><li>A two-tab <code>Tabs</code> (skipped entirely when the active filter empties both finding streams; <code>NoMatchBanner</code> above already covers that case), tab labels <strong>Findings</strong> and <strong>Full app report</strong>: <ul><li><strong>Findings</strong> (<code>ImpactBoard</code>): a <code>HighestImpactBar</code> callout for the single highest-impact eligible finding, if any; then one <code>&lt;section&gt;</code> per impact band in <code>Critical</code> → <code>Warning</code> → <code>Info</code> order, each rendering nothing when it has neither a recommendation row nor an active widget: a recommendation-rollup <code>Table</code> (one row per eligible-finding type, <code>REGISTRY</code>-mapped, <code>incompleteRun</code> excluded: a direct row for a type with one finding, an expandable, paginated summary row for a type with more than one) followed by a <code>WidgetGrid</code> of that band&#39;s active <code>REGISTRY</code> widgets (ranked by widget order: ascending <code>DETECTORS</code> order, action-region components ahead of reference-region ones); then, below every impact band, a small always-visible grid of just Core Usage by Locality, mounted unconditionally regardless of finding state, and a collapsed &quot;Clean checks&quot; disclosure of <code>CleanCheckRow</code> lines built per detector type (every remaining <code>REGISTRY</code> key with zero findings); or inline &quot;No findings to fix right now.&quot; text when nothing is eligible at all.</li><li><strong>Full app report</strong> (<code>ReferenceSection</code>): WallClock → Timeline → Executor Count Over Time → StageTable → Evidence availability → fixed report-lens tail (ETL Phase Attribution → What-If Executor Scaling → Compute Efficiency → Wasted Core-Hours → Core-Usage Distribution). Structural-only, as above. Fully unmounted while Findings is active (see &quot;Render order&quot; above).</li></ul></li></ol><p>Report lenses carry no impact-band chip and self-hide when their input data is absent. Every widget self-wraps in <code>WidgetCard</code> (<code>src/view/WidgetCard.tsx</code>: an <code>&lt;h3&gt;</code> title, one level below the board&#39;s <code>&lt;h2&gt;</code> section headers) for the document outline, except Scorecard, which renders its own header band.</p><p>Every card is collapsible: the title sits in a <code>CollapsibleTrigger</code> button with a chevron (<code>WidgetCard.tsx</code>), and each widget sets its own <code>defaultCollapsed</code> (most default to <code>true</code>; <code>ImpactBoard</code>&#39;s action-region grid forces <code>defaultCollapsed</code> on every <code>REGISTRY</code> widget instance it mounts, so a marginal finding doesn&#39;t default open). Route navigation (&quot;jump to this finding&quot;) goes through the grid coordinator instead of a fixed <code>tabIndex</code>: <code>WidgetGrid.tsx</code> tracks each card&#39;s disclosure button and bumps <code>openRequestGeneration</code> to force a collapsed card open and focus its trigger when the target finding has no anchored row of its own. How much a card shows beyond that is a board-wide choice, set by the Basic/Advanced density tier in the topbar.</p>',43)])])}const u=o(i,[["render",n]]);export{p as __pageData,u as default};
@@ -0,0 +1,6 @@
1
+ import{_ as o,o as t,c as d,a5 as a}from"./chunks/framework.DSg0KOwT.js";const u=JSON.parse('{"title":"Worker protocol","description":"","frontmatter":{},"headers":[],"relativePath":"contributor-guide/architecture/worker-protocol.md","filePath":"contributor-guide/architecture/worker-protocol.md"}'),c={name:"contributor-guide/architecture/worker-protocol.md"};function s(r,e,i,n,l,p){return t(),d("div",null,[...e[0]||(e[0]=[a(`<h1 id="worker-protocol" tabindex="-1">Worker protocol <a class="header-anchor" href="#worker-protocol" aria-label="Permalink to &quot;Worker protocol&quot;">​</a></h1><h2 id="message-protocol" tabindex="-1">Message protocol <a class="header-anchor" href="#message-protocol" aria-label="Permalink to &quot;Message protocol&quot;">​</a></h2><p>Worker to main:</p><ul><li><code>app</code>, <code>stage</code>, <code>sql</code>, <code>executor</code>, <code>job</code>, <code>progress</code>: progressive, posted as the parser reads.</li><li><code>sqlPlan</code> (<code>{ type: &#39;sqlPlan&#39;, data: { executionId, planTree } }</code>): posted by <code>endSqlExecution</code> once a SQL execution&#39;s plan tree resolves.</li><li><code>runAggregates</code>: one whole-run core-time-series summary (busy-core-ms, peak concurrency, per-stage task-duration sums), emitted just before <code>done</code>.</li><li><code>stageExecutorMetrics</code>: the post-completion re-post described in <a href="./overview.html#streaming">Streaming</a>, also emitted just before <code>done</code>.</li><li><code>done</code>, <code>taskData</code>, <code>error</code>.</li></ul><p>A History Server failure is always the typed, display-safe payload <code>{ type: &#39;error&#39;, source: &#39;shs&#39;, code }</code>, where <code>code</code> is one of <code>local-server-unavailable</code>, <code>upstream-unreachable</code>, <code>application-not-found</code>, <code>access-or-upstream-failure</code>, or <code>invalid-event-log</code>. It never carries an upstream message, URL, status, or response body.</p><p>Evidence availability has no dedicated worker message: the final <code>app</code> message carries a compact <code>evidenceInputs</code> counter summary, and <code>done.skippedLines</code> supplies its parse-integrity input.</p><p>Main to worker: <code>parse</code>, <code>parseFiles(files)</code> (rolling <code>eventlog_v2_*</code> directories, one continuous stream across files), <code>parseFromUrl({ baseUrl, appId, attemptId })</code>, and <code>getTaskData(stageId)</code>. The SHS request object is normalized before it reaches the worker: <code>{ baseUrl: string, appId: string, attemptId: string | null }</code>.</p><p><code>taskData</code> uses structured-clone (<code>.slice()</code>) so the worker retains its own <code>Float64Array</code> for re-renders.</p><p>Post-parse prefetch: after <code>done</code>, main runs <code>analyzer.ts</code> to build the bottleneck catalog, then fires parallel <code>getTaskData</code> for every flagged stage so their widgets render immediately. Unflagged stages are on-demand.</p><h3 id="evidence-availability-worker-input" tabindex="-1">Evidence-availability worker input <a class="header-anchor" href="#evidence-availability-worker-input" aria-label="Permalink to &quot;Evidence-availability worker input&quot;">​</a></h3><p><code>packages/core/src/event-handlers.ts</code>&#39;s <code>createState()</code>/per-event handlers start and increment this fixed <code>evidenceInputs</code> shape on the normalized app model:</p><div class="language-text vp-adaptive-theme"><button title="Copy Code" class="copy"></button><span class="lang">text</span><pre class="shiki shiki-themes github-light github-dark vp-code" tabindex="0"><code><span class="line"><span>environmentUpdates, applicationEnds, stageSubmissions, rddStorageSnapshots,</span></span>
2
+ <span class="line"><span>sqlExecutions, resolvedSqlPlans, executorMetricRows, taskRecords</span></span></code></pre></div><p>The counters are a structured-clone-safe summary: they contain counts, not raw events, task records, host names, SQL text, paths, or Spark-property values. The final <code>app</code> message snapshots the counters after all input is processed, immediately before <code>done</code>; a consumer must not add a second worker message just for the ledger. <code>app</code> messages are also emitted mid-parse from <code>SparkListenerApplicationStart</code> and <code>SparkListenerEnvironmentUpdate</code>; those snapshots are partial and non-authoritative. Only the terminal <code>app</code> message emitted by <code>emitParseCompletion</code> (the one immediately before <code>done</code>) should be used for evidence-availability conclusions. On the main thread, <code>useIngest</code> combines the normalized <code>AppModel</code> with <code>done.skippedLines</code>, derives the ledger before calling <code>analyze()</code>, and stores it as <code>appModel.evidenceAvailability</code>.</p><h3 id="evidence-availability-contract-v1" tabindex="-1">Evidence-availability contract (V1) <a class="header-anchor" href="#evidence-availability-contract-v1" aria-label="Permalink to &quot;Evidence-availability contract (V1)&quot;">​</a></h3><p><code>packages/core/src/evidence-availability.ts</code> is the single reusable taxonomy for the browser, future report output, and future headless consumers. Its serialized ledger shape is <code>{ schemaVersion: 1, entries: EvidenceAvailabilityEntry[] }</code>. Entries are ordered by the following fixed eight-key enum:</p><div class="language-text vp-adaptive-theme"><button title="Copy Code" class="copy"></button><span class="lang">text</span><pre class="shiki shiki-themes github-light github-dark vp-code" tabindex="0"><code><span class="line"><span>executorMetrics, rddStorageSnapshots, sqlPlan, sparkConfiguration,</span></span>
3
+ <span class="line"><span>taskCoreTime, infrastructureContext, sourceContext, costContext</span></span></code></pre></div><p>Each entry has the stable fields <code>key</code>, <code>state</code>, <code>reasonCode</code>, <code>summary</code>, and optional <code>evidence</code>. The closed V1 state enum is:</p><div class="language-text vp-adaptive-theme"><button title="Copy Code" class="copy"></button><span class="lang">text</span><pre class="shiki shiki-themes github-light github-dark vp-code" tabindex="0"><code><span class="line"><span>present, disabled, notEmitted, notApplicable, outsideEventLog, unknown</span></span></code></pre></div><p>The closed V1 reason-code enum is:</p><div class="language-text vp-adaptive-theme"><button title="Copy Code" class="copy"></button><span class="lang">text</span><pre class="shiki shiki-themes github-light github-dark vp-code" tabindex="0"><code><span class="line"><span>observed, explicitlyDisabled, noObservedExecutorMetrics,</span></span>
4
+ <span class="line"><span>noObservedStageSubmission, noRddStorageSnapshot, noResolvedSqlPlan,</span></span>
5
+ <span class="line"><span>noSqlExecution, noEnvironmentUpdate, noTaskRecords, noUsableCoreTimeAggregate,</span></span>
6
+ <span class="line"><span>outsideEventLogScope, parseIncomplete</span></span></code></pre></div><p><code>summary</code> is domain-agnostic presentation copy derived from the reason code; the enum values, not labels, are the machine contract. Optional <code>evidence</code> is limited to safe compact provenance of the exact <code>{ eventType, count }</code> shape. <code>eventType</code> values are pinned to the exact <code>evidenceInputs</code> counter keys; they are not shortened or aliased. It must never include raw Spark-property values, host names, paths, SQL text, task records, raw-event identifiers, or other event payloads.</p><table tabindex="0"><thead><tr><th>Key</th><th>Proven <code>present</code> evidence</th><th>Trustworthy absence / boundary result</th></tr></thead><tbody><tr><td><code>executorMetrics</code></td><td><code>executorMetricRows &gt; 0</code></td><td>Explicit observed <code>spark.eventLog.logStageExecutorMetrics=false</code> is <code>disabled</code> / <code>explicitlyDisabled</code>; otherwise <code>notEmitted</code> / <code>noObservedExecutorMetrics</code>. Observed metrics win over the explicit setting.</td></tr><tr><td><code>rddStorageSnapshots</code></td><td><code>rddStorageSnapshots &gt; 0</code> from <code>SparkListenerStageSubmitted</code></td><td>No stage submissions is <code>notEmitted</code> / <code>noObservedStageSubmission</code>; stages submitted without any <code>RDD Info</code> rows is <code>notEmitted</code> / <code>noRddStorageSnapshot</code>. This is RDD snapshot evidence, not block-update telemetry.</td></tr><tr><td><code>sqlPlan</code></td><td><code>resolvedSqlPlans &gt; 0</code></td><td>A SQL execution without a resolved plan is <code>notEmitted</code> / <code>noResolvedSqlPlan</code>; no SQL execution is <code>notApplicable</code> / <code>noSqlExecution</code>.</td></tr><tr><td><code>sparkConfiguration</code></td><td><code>environmentUpdates &gt; 0</code></td><td><code>notEmitted</code> / <code>noEnvironmentUpdate</code>; a normalized empty config object alone is not evidence.</td></tr><tr><td><code>taskCoreTime</code></td><td><code>taskRecords &gt; 0</code> and a usable <code>runAggregates.perStage</code> task count</td><td>No task records is <code>notEmitted</code> / <code>noTaskRecords</code>; task records with no usable aggregate are <code>notEmitted</code> / <code>noUsableCoreTimeAggregate</code>. An empty aggregate object is not usable evidence.</td></tr><tr><td><code>infrastructureContext</code></td><td>none</td><td>Always <code>outsideEventLog</code> / <code>outsideEventLogScope</code>.</td></tr><tr><td><code>sourceContext</code></td><td>none</td><td>Always <code>outsideEventLog</code> / <code>outsideEventLogScope</code>.</td></tr><tr><td><code>costContext</code></td><td>none</td><td>Always <code>outsideEventLog</code> / <code>outsideEventLogScope</code>.</td></tr></tbody></table><p>Parse integrity is fail-closed. A parse is trustworthy only when it skipped no malformed lines and observed an application-end event. For an untrustworthy parse, proof already observed remains <code>present</code>, and an explicitly observed executor-metrics disablement remains <code>disabled</code>; every conclusion based on absence or irrelevance becomes <code>unknown</code> / <code>parseIncomplete</code>. The three outside-event-log product boundaries remain <code>outsideEventLog</code>. In particular, <code>notEmitted</code> and <code>notApplicable</code> are never favorable measurements inferred from an incomplete parse.</p><p>Availability is model metadata, not a detector or finding: it never changes detector thresholds or ordering, impact band, catalog/scorecard counts, no-bottleneck behavior, or detector suppression. <code>session-snapshot.ts</code> captures and restores <code>evidenceAvailability</code> with the normalized model, so a recent-file switch retains the same ledger without reparsing.</p><h3 id="portable-evidence-report-v1" tabindex="-1">Portable evidence report (V1) <a class="header-anchor" href="#portable-evidence-report-v1" aria-label="Permalink to &quot;Portable evidence report (V1)&quot;">​</a></h3><p><code>packages/core/src/evidence-report.ts</code>&#39;s <code>buildEvidenceReport(appModel, { redact })</code> returns <code>{ markdown, json }</code>: a self-contained, byte-stable document that runs the detectors over an <code>appModel</code> and serializes the result for sharing outside the tool. Raw task records are never included; identifier redaction (app id + host names → <code>app-1</code>/<code>host-1</code> pseudonyms via <code>packages/core/src/redact.ts</code>) is opt-in with <code>{ redact: true }</code>. The JSON is pinned by <code>EVIDENCE_SCHEMA_VERSION</code> (currently <code>3</code>, surfaced as <code>json.schemaVersion</code>) and has this fixed top-level key order:</p><div class="language-text vp-adaptive-theme"><button title="Copy Code" class="copy"></button><span class="lang">text</span><pre class="shiki shiki-themes github-light github-dark vp-code" tabindex="0"><code><span class="line"><span>schemaVersion, summary, evidenceAvailability, detectors, findings, recommendations, cleanChecks</span></span></code></pre></div><ul><li><code>summary</code> is the run header: <code>{ app: { id, name, sparkVersion }, stageCount, jobCount, sqlExecutionCount, findingCount, impactBandCounts }</code>.</li><li><code>evidenceAvailability</code> is the ledger above (or <code>null</code> when absent).</li><li><code>detectors</code> is <code>detectorCatalog()</code> output: one <code>{ type, version, scope, thresholds, docAnchor }</code> per detector, in <code>DETECTORS</code> order, so the exact threshold set that produced each finding travels with the evidence.</li><li><code>findings</code> are deterministically sorted rows (impact band → type → stage → id), each with a stable <code>id</code>, <code>tag</code>, core columns, an always-present <code>actionLabel</code>, and an <code>evidence</code> sub-object for non-core fields; <code>confidence</code>/ <code>validationRequired</code>/<code>docAnchor</code> appear only when the detector emitted them.</li><li><code>recommendations</code> is the impact-ranked <code>buildRecommendationRollup</code> output (<code>packages/core/src/recommendation-rollup.ts</code>), and <code>cleanChecks</code> lists every detector type that fired zero findings this run: see the 2026-09-03 update below.</li></ul><p>Determinism holds because detector order, finding sort, and object key order are all fixed, so a given <code>appModel</code> serializes identically across calls.</p><p>Redaction enumerates hosts two ways: by walking the findings tree for every string value under a key literally named <code>host</code> (<code>evidence.host</code>, <code>evidence.failedTaskDetails[].host</code>, <code>evidence.retriedTaskDetails[].host</code>, and any future nested <code>host</code> field, all covered without enumerating paths), and by scanning every string value for EC2-style hostnames / bare IPv4 tokens. So identifiers that surface only in free text (recommendation copy, a <code>stageFailed</code> failure-reason value) are pseudonymized too. Pseudonym numbering uses a numeric-aware sort, so re-redacting an already-redacted report is a no-op even past <code>host-10</code>.</p><p>The Markdown rendering mirrors the JSON&#39;s AC3 field set: each finding block prints its <code>detector version</code>, its sorted <code>evidence</code> entries (byte-magnitude keys humanized), and the report ends with a <code>## Detectors</code> catalog carrying the version + threshold set. A finding&#39;s <code>impactEstimate</code> (when its <code>basis</code> isn&#39;t <code>&#39;informational&#39;</code>) prints as its own <code>- impact: </code> line (<code>Estimated &lt;low&gt;-&lt;high&gt;</code> and/or the raw-waste figure, plus <code>estimateMethod</code>), via <code>renderImpactEstimate</code>/<code>formatWallClockRange</code>/<code>formatRawWaste</code> in <code>packages/core/src/evidence-report.ts</code>. <code>EvidenceExport</code> names downloads <code>evidence-&lt;appId&gt;[-redacted].&lt;md|json&gt;</code>, taking the app id from the (already pseudonymized when redacting) report so a redacted file never leaks the real id and is never name-identical to a raw export.</p><p>Decision 9 (design spec): <code>Finding</code>&#39;s <code>impactEstimate</code> field, plus the <code>utilizationFraction</code>/<code>memorySize</code>/<code>diskSize</code>/<code>numCachedPartitions</code>/<code>numPartitions</code> instrumentation fields the impact estimator reads, were added without bumping <code>EVIDENCE_SCHEMA_VERSION</code> past <code>1</code>. Both additions are purely optional and ride <code>Finding</code>&#39;s existing optional-field-plus-catch-all convention, so an evidence report built before these fields existed still deserializes and compares byte-for-byte against one built after: nothing about the schema&#39;s stability guarantee changed, only its surface grew. A future reader who notices <code>impactEstimate</code> in the JSON without a schema bump is looking at this deliberate call, not an oversight.</p><p>2026-08-30 update: <code>EVIDENCE_SCHEMA_VERSION</code> was bumped to <code>2</code> for the occupancy-weighted attribution redesign (see <a href="./impact-estimation.html#occupancy-weighted-attribution">Occupancy-weighted attribution</a>): <code>ImpactEstimate</code>&#39;s shape changed from <code>{low, high}</code> to <code>{basis, wallClock, estimateMethod, rawWaste?}</code>, a real, non-additive breaking change to a field this same Decision 9 previously shipped without a bump. <code>impactEstimate</code> had zero consumers outside <code>packages/core/src/impact-estimator.ts</code> and its own tests at the time of this bump (confirmed by grep across <code>src/view/*</code> and <code>packages/core/src/evidence-report.ts</code>), so no other code needed migrating alongside it.</p><p><code>FindingRow</code>&#39;s schema was extended to surface <code>impactEstimate</code> as a first-class column without further bumping <code>EVIDENCE_SCHEMA_VERSION</code> past <code>2</code>, consistent with <code>Finding</code>&#39;s existing optional-field-plus-catch-all convention. It appears after the pinned core columns (<code>id</code>, <code>type</code>, <code>impactBand</code>, <code>stageId</code>, <code>metric</code>, <code>value</code>, <code>recommendation</code>, <code>detectorVersion</code>, plus optional but pinned <code>confidence</code>, <code>validationRequired</code>, <code>docAnchor</code>) without displacing any of them; byte-for-byte deserializability of existing reports is preserved. <code>FindingRow.impactEstimate</code> carries the full contract documented in <a href="./impact-estimation.html#occupancy-weighted-attribution">Impact estimation</a> (basis, wallClock, estimateMethod, rawWaste).</p><p>2026-09-03 update: the &quot;Fix These First&quot; dashboard redesign (impact-ranked recommendation rollup, short per-finding action labels, a clean-checks table) was UI-only when it shipped; this update ports the underlying data into <code>buildEvidenceReport()</code> so the CLI (<code>packages/cli/bin/sparkforensics-analyze.mjs</code>) and the MCP tool (<code>diagnoseRun</code> in <code>packages/core/src/mcp-tools.ts</code>) get it too, not just the web markdown/JSON download. Three additions, all purely additive, so this does not bump <code>EVIDENCE_SCHEMA_VERSION</code> past <code>2</code>, the same rationale as the <code>impactEstimate</code> addition immediately above: <code>FindingRow.actionLabel</code> is now always present (a short imperative label like &quot;Reduce shuffle size&quot;), sourced from a new core module, <code>packages/core/src/finding-action-label.ts</code>&#39;s <code>coreFindingActionLabel</code>, extracted from the (type, discriminant) switch statement that used to live only in the view layer (<code>src/view/finding-action-label.ts</code>, which now wraps the core function and layers its own <code>REGISTRY</code> fallback on top). <code>EvidenceReportJson.recommendations</code> is the same impact-ranked <code>buildRecommendationRollup</code> grouping (<code>packages/core/src/recommendation-rollup.ts</code>) that <code>FixTheseFirst.tsx</code> renders, so CLI/MCP/download consumers get the same &quot;what&#39;s the highest-impact fix&quot; ranking the dashboard shows, without changing the existing <code>findings</code> array&#39;s own impact-band-sorted order at all. <code>EvidenceReportJson.cleanChecks</code> lists every detector type that fired zero findings this run, each with <code>getThresholdSummary</code>&#39;s one-line &quot;what would have tripped it&quot; sentence; unlike the dashboard&#39;s <code>Alerts.tsx</code> clean-checks table, it deliberately includes the one &quot;always-mounted&quot; reference type (<code>coreLocality</code>) even when it has no findings, since a flat evidence report has no separate always-visible surface for it to already appear on the way that widget does on the board.</p><h3 id="finding-identity" tabindex="-1">Finding identity <a class="header-anchor" href="#finding-identity" aria-label="Permalink to &quot;Finding identity&quot;">​</a></h3><p>Every finding <code>analyzer.ts</code> emits carries a stable <code>id</code> (<code>push()</code>&#39;s <code>findingId()</code> choke point, FNV-1a hash over <code>type | location | metric | value | discriminators</code>) and a <code>detectorVersion</code> (from the emitting <code>DETECTORS</code> entry). The two are deliberately decoupled: <code>id</code> is derived only from a finding&#39;s evidence tuple, never from <code>detectorVersion</code>, so a threshold or logic tweak that bumps a detector&#39;s <code>version</code> does not change the <code>id</code> of findings it still emits at the same location/metric. A saved reference (dashboard bookmark, exported report row, future URL-restored filter) keeps pointing at the same logical finding across detector revisions; <code>detectorVersion</code> is separate provenance metadata for &quot;which ruleset produced this,&quot; not part of identity. See <code>packages/core/test/analyzer-finding-identity.test.js</code> for the decoupling proof.</p><p><code>id</code> stability holds for equivalent reruns of the same schema/analyzer version on the same input. Changing the id-derivation rule itself (the hash algorithm, or which fields feed the location key/discriminators) is an intentional breaking change and must bump <code>EVIDENCE_SCHEMA_VERSION</code> (<code>packages/core/src/evidence-report.ts</code>); there is no separate id-scheme version.</p><h3 id="headless-analysis-cli-v1" tabindex="-1">Headless analysis CLI (V1) <a class="header-anchor" href="#headless-analysis-cli-v1" aria-label="Permalink to &quot;Headless analysis CLI (V1)&quot;">​</a></h3><p><code>packages/cli/bin/sparkforensics-analyze.mjs</code> runs the same parser + detector contracts outside the browser, for CI. It accepts a single event-log file or a rolling-log directory, drives <code>runParse</code>/<code>runParseFiles</code> (<code>packages/core/src/parser-worker.ts</code>) with a Node-only File-like shim (<code>packages/core/src/cli/collect-run.ts</code>), and writes the exact <code>buildEvidenceReport</code> JSON schema described above: one format, not a second. Alternatively, <code>--shs-base-url &lt;url&gt; --app-id &lt;id&gt; [--attempt-id &lt;id&gt;]</code> fetches the run from a Spark History Server instead (mutually exclusive with the positional file/directory argument), calling <code>resolveFromShs</code> (<code>packages/core/src/shs-load.ts</code>, shared with the MCP server&#39;s SHS source path below) directly; no dependency on the <code>packages/server</code> package. A failed SHS fetch reports its message to <code>stderr</code> and exits <code>2</code>, same as a local file that can&#39;t be parsed.</p><p>Optional CLI-flag budgets (<code>--max-runtime &lt;ms&gt;</code>, <code>--max-spill &lt;gb&gt;</code>, <code>--max-skew &lt;ratio&gt;</code>, <code>--max-failed-task-rate &lt;pct&gt;</code>, <code>--min-efficiency &lt;pct&gt;</code>) are evaluated in <code>packages/core/src/cli/budgets.ts</code> against the existing finding catalog (<code>analyze()</code>) and <code>computeEfficiencyModel</code>; there is no second rule engine. A budget whose required evidence is missing (e.g. the run never emitted <code>ApplicationEnd</code>, or has no usable per-task <code>runAggregates</code>) is reported as inconclusive (<code>stderr</code> warning) rather than silently passing, and gets its own exit code distinct from both pass and violation. Exit codes: <code>0</code> pass, <code>1</code> a configured budget was violated, <code>2</code> the input could not be parsed at all, <code>3</code> no violations but at least one budget was inconclusive. A violation always wins over an inconclusive result in the same run (exit <code>1</code>, not <code>3</code>).</p><h3 id="mcp-server-v1" tabindex="-1">MCP server (V1) <a class="header-anchor" href="#mcp-server-v1" aria-label="Permalink to &quot;MCP server (V1)&quot;">​</a></h3><p><code>packages/core/src/mcp-server-factory.ts</code>&#39;s <code>createMcpServer()</code> registers 6 tools: <code>diagnose_run</code> (thresholded findings + remediation text), <code>get_run_summary</code> (app/stage/job/sql counts and duration, no findings), <code>compare_runs</code> (categorized findings delta + metric deltas between two runs), <code>evaluate_budgets</code> (pass/fail budget thresholds against one run, optionally with a second run for regression/fail-on-introduced budgets: the MCP side of the CLI&#39;s <code>evaluateBudgets()</code> gating), <code>get_finding_evidence</code> (raw evidence bundle for one finding, for drill-down after <code>diagnose_run</code>), and <code>get_finding_documentation</code> (detection/tuning reference docs for one finding type, independent of any run). None re-implement detector logic; all repackage <code>analyze</code>/<code>buildEvidenceReport</code>/<code>compareRuns</code>/<code>captureSnapshot</code>/<code>evaluateBudgets</code> from <code>packages/core/src/mcp-tools.ts</code>, which resolves a <code>source</code> (event-log <code>path</code>, or an SHS <code>shsBaseUrl</code>/<code>appId</code>/<code>attemptId</code> triple) into a cached <code>AppModel</code>. The SHS source path calls <code>resolveFromShs</code> (<code>packages/core/src/shs-load.ts</code>, also used directly by the CLI&#39;s <code>--shs-base-url</code> mode above), which reuses two extractions shared with the browser ingestion flow: <code>fetchShsEventLog</code> (<code>packages/core/src/proxy.js</code>) fetches the zip, and <code>decodeShsArchive</code> (<code>packages/core/src/shs-fetch.ts</code>, re-exported from <code>parser-worker.ts</code>&#39;s barrel) decodes it into a parsed <code>AppModel</code>, the same split <code>runParseFromUrl</code> itself calls into.</p><p>Two transports connect to that one factory: <code>packages/mcp/bin/sparkforensics-mcp.mjs</code> (stdio, for local MCP clients) and <code>packages/server/index.js</code>&#39;s <code>/mcp</code> route (streamable HTTP, for the local server). The run cache (<code>packages/core/src/mcp-tools.ts</code>) is a module-level LRU (cap 8 and 15-minute idle TTL by default, overridable via <code>SPARKFORENSICS_MCP_CACHE_CAP</code> and <code>SPARKFORENSICS_MCP_CACHE_TTL_MS</code>, lazily swept on access) keyed by resolved source (path mtime, or SHS baseUrl+appId+attemptId), so a client mints a <code>runId</code> once via <code>resolveOrCreateRun</code> and reuses it across subsequent tool calls instead of re-parsing.</p><p>Every tool failure comes back as <code>{isError: true, content: [...], structuredContent: {code}}</code>, never an HTTP-status-shaped error; <code>code</code> is one of the 5 existing SHS codes (<code>SHS_ERROR_CODES</code>, <code>packages/core/src/shs-request.js</code>) plus <code>run-not-found</code>, <code>finding-not-found</code>, <code>invalid-event-log</code> (also covers a nonexistent <code>path</code> source), and <code>archive-too-large</code> (SHS archive over the <code>SPARKFORENSICS_MAX_ARCHIVE_BYTES</code> byte cap, default 1 GiB, because the MCP path buffers the whole archive in memory, unlike the streaming <code>/shs-proxy</code> route).</p><p><code>scripts/vendor-core.mjs</code> (shared by <code>packages/cli</code>, <code>packages/mcp</code>, and <code>packages/server</code>&#39;s <code>prepack</code> scripts) vendors <code>packages/core/src/</code> wholesale into each package&#39;s own <code>vendor-core/</code> at pack time, pre-stripping TypeScript to plain <code>.js</code> (Node&#39;s native TS stripping refuses to run on <code>.ts</code> files under <code>node_modules</code>, which is exactly where a published <code>vendor-core/</code> lands). Each package&#39;s bin/entry point resolves its needed module from <code>vendor-core/</code> if present, else falls back to the real <code>packages/core/src/</code> sibling loaded as <code>.ts</code> directly (<code>packages/server/index.js</code> mirrors <code>resolveStaticRoot</code>&#39;s <code>public/</code>-vs-<code>../dist</code> pattern for this same fallback).</p><h2 id="typescript-core-and-runtime-event-validation" tabindex="-1">TypeScript core and runtime event validation <a class="header-anchor" href="#typescript-core-and-runtime-event-validation" aria-label="Permalink to &quot;TypeScript core and runtime event validation&quot;">​</a></h2><p>All of <code>src/</code> (browser SPA) and <code>packages/core/src/</code> (shared analysis logic) is strict TypeScript; the only plain-<code>.js</code> holdouts are the two vendored third-party decompressors, <code>packages/core/src/vendor/fflate.js</code> and <code>packages/core/src/vendor/fzstd.js</code> (left untouched deliberately: vendored code, not project code). Every remaining import of a same-repo module uses a <code>.ts</code> specifier (e.g. <code>import { dispatchLine } from &#39;./event-handlers.ts&#39;</code>), not <code>.js</code>: the CLI and MCP entrypoints (<code>packages/cli/bin/sparkforensics-analyze.mjs</code>, <code>packages/mcp/bin/sparkforensics-mcp.mjs</code>) run under plain Node&#39;s ESM resolver, which cannot remap a <code>.js</code> specifier to a same-named <code>.ts</code> file the way Vite/Vitest&#39;s bundler-style resolver can. <code>tsconfig.json</code> sets <code>&quot;allowImportingTsExtensions&quot;: true</code> to make this legal.</p><p>Event schema validation lives in <code>packages/core/src/event-schemas.ts</code>: one zod schema per <code>SparkListener*</code> event variant the parser understands (15 total: <code>LogStart</code>, <code>ApplicationStart</code>, <code>EnvironmentUpdate</code>, <code>ApplicationEnd</code>, <code>JobStart</code>, <code>JobEnd</code>, <code>StageSubmitted</code>, <code>StageCompleted</code>, <code>StageExecutorMetrics</code>, <code>TaskEnd</code>, the three SQL-execution-UI listener events, <code>ExecutorAdded</code>, <code>ExecutorRemoved</code>), combined into <code>SparkEventSchema = z.discriminatedUnion(&#39;Event&#39;, [...])</code>. <code>processEvent</code>&#39;s switch (<code>event-handlers.ts</code>) consumes the resulting <code>SparkEvent</code> union type directly, so a schema change and a handler&#39;s expectations can&#39;t silently drift apart.</p><p>The recursive <code>sparkPlanInfo.children</code> tree carried on SQL-execution-start events is validated by <code>parseSparkPlanInfoTree</code>, an iterative, explicit heap-allocated-stack parser, deliberately not <code>z.lazy()</code>. A real plan tree&#39;s depth is unbounded, and attacker-uncontrolled input should never hand zod&#39;s own recursive schema resolution an arbitrarily deep structure to walk. The iterative parser shallow-validates one node at a time via <code>z.object(...)</code> and builds the tree itself, capping at <code>MAX_PLAN_DEPTH = 500</code> (throws past that). This mirrors <code>packages/core/src/plan-tree-walk.ts</code>&#39;s <code>walkPlanTree</code>, which uses the same iterative-over-recursive approach for the <em>resolved</em> tree; <code>event-schemas.ts</code> applies it one layer earlier, to the raw JSON before it becomes a tree at all.</p><p>There are two external-data boundaries, the two places this codebase parses data it does not control. Both now run that data through a schema, and both treat a validation failure the same way: a silent skip, not a distinct error.</p><ul><li><code>dispatchLine</code> (<code>event-handlers.ts</code>): after <code>JSON.parse</code> succeeds, a line whose <code>Event</code> value is one of the 15 modeled types but fails that type&#39;s own schema now increments <code>skippedLines</code> (previously: no shape validation existed at all, and a malformed event silently corrupted downstream state with no signal anywhere). An <code>Event</code> value outside the 15 modeled types is still silently ignored without incrementing <code>skippedLines</code>, unchanged from before migration (see the note below on why the broader design was rejected).</li><li><code>shs-fetch.ts</code>: on a non-OK response from the local SHS proxy, the JSON error envelope is validated against <code>ShsProxyErrorBodySchema</code> (<code>packages/core/src/shs-schemas.ts</code>, <code>{ code: string }</code>, <code>.passthrough()</code>). A body that isn&#39;t valid JSON, or is JSON but fails that schema, falls back to the generic <code>access-or-upstream-failure</code> code: the same silent-skip treatment as a malformed log line, collapsing what could have been a separate &quot;malformed SHS response&quot; error code into the existing generic one.</li></ul><p>Neither boundary distinguishes &quot;malformed JSON&quot; from &quot;wrong shape&quot; from &quot;unrecognized variant&quot; in what it reports outward: all three collapse into the same skip/fallback path. That was a deliberate scope decision. The correction below says why the <em>line</em> boundary stops there rather than flagging every unrecognized event type.</p><blockquote><p>Correction made during migration, not part of the original design: an earlier draft counted <em>any</em> <code>Event</code> value outside the 15 modeled types toward <code>skippedLines</code>, not just ones that fail their own schema. Running that design against real Spark event logs (which always contain plenty of ordinary event types this tool has never modeled: <code>TaskStart</code>, <code>BlockManagerAdded</code>, <code>ExecutorMetricsUpdate</code>, and others, always silently ignored pre-migration) pushed <code>skippedLines</code> from 0 to the tens of thousands on completely healthy logs. That would have tripped <code>evidence-availability.ts</code>&#39;s fail-closed <code>trustworthy</code> gate and shown a false &quot;malformed JSON&quot; warning on every real file. Caught in review against real fixtures, not synthetic ones; fixed to the narrower rule described above before merging. <code>evidence-availability.ts</code>&#39;s gate itself (<code>trustworthy = skippedLines === 0 &amp;&amp; applicationEnds &gt; 0</code>) needed no code change once this was fixed: it was already correct, the input feeding it was not.</p></blockquote><p>The exhaustiveness convention is <code>packages/core/src/assert-never.ts</code>. <code>assertNever(x: never): never</code> throws at runtime and, more importantly, fails <code>tsc</code> at compile time if <code>x</code> is not actually <code>never</code>, i.e. if some case of a union type isn&#39;t handled. Used today at the <code>default</code> arm of <code>event-handlers.ts</code>&#39;s <code>processEvent</code> switch (over <code>SparkEvent</code>) and <code>analyzer.ts</code>&#39;s scope-dispatch switch (over a detector&#39;s <code>scope: &#39;stage&#39; | &#39;sql&#39; | &#39;app&#39; | &#39;config&#39;</code>). It is the project&#39;s standard pattern for any future exhaustive switch or dispatch over a closed union: add <code>default: return assertNever(x);</code> (or the closest non-returning equivalent) so that adding a new union member without updating every consumer becomes a compile error instead of a silent runtime gap.</p>`,55)])])}const m=o(c,[["render",s]]);export{u as __pageData,m as default};