sparkforensics-cli 0.1.0 → 0.2.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +6 -0
- package/bin/sparkforensics-analyze.mjs +113 -48
- package/export-template/docs/404.html +25 -0
- package/export-template/docs/assets/app.CndaAS6v.js +1 -0
- package/export-template/docs/assets/aqe-loop.IwQSATHw.svg +1 -0
- package/export-template/docs/assets/aqe-loop.dark.DGbaxqJE.svg +1 -0
- package/export-template/docs/assets/broadcast-vs-shuffle.Db4WY1XK.svg +1 -0
- package/export-template/docs/assets/broadcast-vs-shuffle.dark.C7Bxs0mG.svg +1 -0
- package/export-template/docs/assets/cache-lifecycle.dark.B-hS7AgU.svg +1 -0
- package/export-template/docs/assets/cache-lifecycle.rEOVYQNU.svg +1 -0
- package/export-template/docs/assets/chunks/@localSearchIndexroot.DNY8bVcl.js +1 -0
- package/export-template/docs/assets/chunks/VPLocalSearchBox.yJbZbsEo.js +9 -0
- package/export-template/docs/assets/chunks/duplicate-plan-subtree.dark.Cdp70QhV.js +1 -0
- package/export-template/docs/assets/chunks/framework.DSg0KOwT.js +20 -0
- package/export-template/docs/assets/chunks/retry-escalation-ladder.dark.DHipdJgZ.js +1 -0
- package/export-template/docs/assets/chunks/theme.Df2VAG9w.js +2 -0
- package/export-template/docs/assets/cold-start-timeline.DxC_Sc7w.svg +1 -0
- package/export-template/docs/assets/cold-start-timeline.dark.CZ17YcAG.svg +1 -0
- package/export-template/docs/assets/columnar-layout.PghGeOEA.svg +1 -0
- package/export-template/docs/assets/columnar-layout.dark.BVNlz0ff.svg +1 -0
- package/export-template/docs/assets/container-memory.DIO0AnIm.svg +1 -0
- package/export-template/docs/assets/container-memory.dark.CP-5zuCl.svg +1 -0
- package/export-template/docs/assets/contributor-guide_architecture_board-widgets.md.B-OsL91z.js +1 -0
- package/export-template/docs/assets/contributor-guide_architecture_board-widgets.md.B-OsL91z.lean.js +1 -0
- package/export-template/docs/assets/contributor-guide_architecture_detector-contract.md.BOeH4d1J.js +1 -0
- package/export-template/docs/assets/contributor-guide_architecture_detector-contract.md.BOeH4d1J.lean.js +1 -0
- package/export-template/docs/assets/contributor-guide_architecture_drill-down.md.BtPdlM7r.js +1 -0
- package/export-template/docs/assets/contributor-guide_architecture_drill-down.md.BtPdlM7r.lean.js +1 -0
- package/export-template/docs/assets/contributor-guide_architecture_impact-estimation.md.DYCDPgkh.js +1 -0
- package/export-template/docs/assets/contributor-guide_architecture_impact-estimation.md.DYCDPgkh.lean.js +1 -0
- package/export-template/docs/assets/contributor-guide_architecture_index.md.3TO9ic6w.js +1 -0
- package/export-template/docs/assets/contributor-guide_architecture_index.md.3TO9ic6w.lean.js +1 -0
- package/export-template/docs/assets/contributor-guide_architecture_overview.md.CehiRmGn.js +1 -0
- package/export-template/docs/assets/contributor-guide_architecture_overview.md.CehiRmGn.lean.js +1 -0
- package/export-template/docs/assets/contributor-guide_architecture_state-and-history.md.m3S3UdMk.js +1 -0
- package/export-template/docs/assets/contributor-guide_architecture_state-and-history.md.m3S3UdMk.lean.js +1 -0
- package/export-template/docs/assets/contributor-guide_architecture_widget-rendering.md.DbqPf2OT.js +1 -0
- package/export-template/docs/assets/contributor-guide_architecture_widget-rendering.md.DbqPf2OT.lean.js +1 -0
- package/export-template/docs/assets/contributor-guide_architecture_worker-protocol.md.B93qJ_tT.js +6 -0
- package/export-template/docs/assets/contributor-guide_architecture_worker-protocol.md.B93qJ_tT.lean.js +1 -0
- package/export-template/docs/assets/contributor-guide_contributing.md.CvRsdr6J.js +1 -0
- package/export-template/docs/assets/contributor-guide_contributing.md.CvRsdr6J.lean.js +1 -0
- package/export-template/docs/assets/contributor-guide_development-setup.md.DvAN_9mK.js +12 -0
- package/export-template/docs/assets/contributor-guide_development-setup.md.DvAN_9mK.lean.js +1 -0
- package/export-template/docs/assets/contributor-guide_testing.md.6rIKqSyY.js +1 -0
- package/export-template/docs/assets/contributor-guide_testing.md.6rIKqSyY.lean.js +1 -0
- package/export-template/docs/assets/dag-stages.DSz_S937.svg +1 -0
- package/export-template/docs/assets/dag-stages.dark.F72UzxH4.svg +1 -0
- package/export-template/docs/assets/driver-executor.D5pQ7YN1.svg +1 -0
- package/export-template/docs/assets/driver-executor.dark.BmX9cPvh.svg +1 -0
- package/export-template/docs/assets/duplicate-plan-subtree.B4cvN6fj.svg +1 -0
- package/export-template/docs/assets/duplicate-plan-subtree.dark.Dw8wS0Ag.svg +1 -0
- package/export-template/docs/assets/index.md.CHJVslga.js +1 -0
- package/export-template/docs/assets/index.md.CHJVslga.lean.js +1 -0
- package/export-template/docs/assets/inter-italic-cyrillic-ext.r48I6akx.woff2 +0 -0
- package/export-template/docs/assets/inter-italic-cyrillic.By2_1cv3.woff2 +0 -0
- package/export-template/docs/assets/inter-italic-greek-ext.1u6EdAuj.woff2 +0 -0
- package/export-template/docs/assets/inter-italic-greek.DJ8dCoTZ.woff2 +0 -0
- package/export-template/docs/assets/inter-italic-latin-ext.CN1xVJS-.woff2 +0 -0
- package/export-template/docs/assets/inter-italic-latin.C2AdPX0b.woff2 +0 -0
- package/export-template/docs/assets/inter-italic-vietnamese.BSbpV94h.woff2 +0 -0
- package/export-template/docs/assets/inter-roman-cyrillic-ext.BBPuwvHQ.woff2 +0 -0
- package/export-template/docs/assets/inter-roman-cyrillic.C5lxZ8CY.woff2 +0 -0
- package/export-template/docs/assets/inter-roman-greek-ext.CqjqNYQ-.woff2 +0 -0
- package/export-template/docs/assets/inter-roman-greek.BBVDIX6e.woff2 +0 -0
- package/export-template/docs/assets/inter-roman-latin-ext.4ZJIpNVo.woff2 +0 -0
- package/export-template/docs/assets/inter-roman-latin.Di8DUHzh.woff2 +0 -0
- package/export-template/docs/assets/inter-roman-vietnamese.BjW4sHH5.woff2 +0 -0
- package/export-template/docs/assets/join-strategy.C_FvrCEo.svg +1 -0
- package/export-template/docs/assets/join-strategy.dark.ChMLnNII.svg +1 -0
- package/export-template/docs/assets/memory-borrowing.BqQRJg0u.svg +1 -0
- package/export-template/docs/assets/memory-borrowing.dark.Yhh20O9C.svg +1 -0
- package/export-template/docs/assets/memory-regions.XHvO7jHG.svg +1 -0
- package/export-template/docs/assets/memory-regions.dark.D4TP9_08.svg +1 -0
- package/export-template/docs/assets/repartition-vs-coalesce.BovLRrpj.svg +1 -0
- package/export-template/docs/assets/repartition-vs-coalesce.dark.BhAczKZQ.svg +1 -0
- package/export-template/docs/assets/retry-escalation-ladder.DyTKJJmZ.svg +1 -0
- package/export-template/docs/assets/retry-escalation-ladder.dark.BdsabtU3.svg +1 -0
- package/export-template/docs/assets/shuffle-map-reduce.KuOEZVmg.svg +1 -0
- package/export-template/docs/assets/shuffle-map-reduce.dark.BgQZnFSb.svg +1 -0
- package/export-template/docs/assets/spill-classification.BU2euYDO.svg +1 -0
- package/export-template/docs/assets/spill-classification.dark.D7i1M40d.svg +1 -0
- package/export-template/docs/assets/style.DXOMCXxn.css +1 -0
- package/export-template/docs/assets/tuning-reference_anti-patterns.md.Df1YMIHu.js +1 -0
- package/export-template/docs/assets/tuning-reference_anti-patterns.md.Df1YMIHu.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_aqe.md.BIsCtLzm.js +1 -0
- package/export-template/docs/assets/tuning-reference_aqe.md.BIsCtLzm.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-broadcast-sizing.md.CEstB3Ia.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-broadcast-sizing.md.CEstB3Ia.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-cold-start.md.CEuy-72y.js +7 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-cold-start.md.CEuy-72y.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-duplicate-plan-subtree.md.CIohQDfn.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-duplicate-plan-subtree.md.CIohQDfn.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-failures.md.4z5BXGJ2.js +6 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-failures.md.4z5BXGJ2.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-gc.md.DSxzZRK7.js +6 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-gc.md.DSxzZRK7.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-job-failure-rate.md.BaJl__1W.js +8 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-job-failure-rate.md.BaJl__1W.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-memory-utilization.md.DbP-SJZc.js +7 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-memory-utilization.md.DbP-SJZc.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-retry-waste.md.D5JMjVOt.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-retry-waste.md.D5JMjVOt.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-shuffle.md.CM-nTmIH.js +12 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-shuffle.md.CM-nTmIH.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-skew.md.BdUwiDhn.js +14 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-skew.md.BdUwiDhn.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-slow-host.md.BlIo6UDW.js +7 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-slow-host.md.BlIo6UDW.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-small-files.md.B8kloyx8.js +5 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-small-files.md.B8kloyx8.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-spill.md.PNH7mITt.js +6 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-spill.md.PNH7mITt.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-straggler.md.DY36fHN5.js +7 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-straggler.md.DY36fHN5.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-tiny-tasks.md.QTV7O8kU.js +8 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-tiny-tasks.md.QTV7O8kU.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-utilization.md.DTiueZC3.js +5 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-utilization.md.DTiueZC3.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_caching.md.B7aQ8asB.js +1 -0
- package/export-template/docs/assets/tuning-reference_caching.md.B7aQ8asB.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_cluster-config.md.ZVmDGsQ3.js +1 -0
- package/export-template/docs/assets/tuning-reference_cluster-config.md.ZVmDGsQ3.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_config.md.UvveiWG3.js +1 -0
- package/export-template/docs/assets/tuning-reference_config.md.UvveiWG3.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_data-formats.md.bjCAWH3N.js +1 -0
- package/export-template/docs/assets/tuning-reference_data-formats.md.bjCAWH3N.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_index.md.BQ_NooMV.js +1 -0
- package/export-template/docs/assets/tuning-reference_index.md.BQ_NooMV.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_intro.md.CobD-lGB.js +1 -0
- package/export-template/docs/assets/tuning-reference_intro.md.CobD-lGB.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_joins.md.BtKs_CuW.js +1 -0
- package/export-template/docs/assets/tuning-reference_joins.md.BtKs_CuW.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_memory-model.md.DhT-n4y3.js +1 -0
- package/export-template/docs/assets/tuning-reference_memory-model.md.DhT-n4y3.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_metrics.md.mLOh7Apj.js +1 -0
- package/export-template/docs/assets/tuning-reference_metrics.md.mLOh7Apj.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_partitioning.md.q0zKF_8X.js +1 -0
- package/export-template/docs/assets/tuning-reference_partitioning.md.q0zKF_8X.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_pyspark.md.DDCfvN9t.js +6 -0
- package/export-template/docs/assets/tuning-reference_pyspark.md.DDCfvN9t.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_shuffle.md.BZZ7R4Ix.js +1 -0
- package/export-template/docs/assets/tuning-reference_shuffle.md.BZZ7R4Ix.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_spark-architecture.md.Dwzm5avO.js +1 -0
- package/export-template/docs/assets/tuning-reference_spark-architecture.md.Dwzm5avO.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_table-formats.md.D6wj-2dX.js +1 -0
- package/export-template/docs/assets/tuning-reference_table-formats.md.D6wj-2dX.lean.js +1 -0
- package/export-template/docs/assets/udf-execution-models.BUFDICuG.svg +1 -0
- package/export-template/docs/assets/udf-execution-models.dark.YTNS6GDq.svg +1 -0
- package/export-template/docs/assets/user-guide_alternative-log-retrieval.md.sU3KGarf.js +1 -0
- package/export-template/docs/assets/user-guide_alternative-log-retrieval.md.sU3KGarf.lean.js +1 -0
- package/export-template/docs/assets/user-guide_getting-started.md.DtEM37MK.js +3 -0
- package/export-template/docs/assets/user-guide_getting-started.md.DtEM37MK.lean.js +1 -0
- package/export-template/docs/assets/user-guide_mcp-tools.md.C8MiIu7F.js +125 -0
- package/export-template/docs/assets/user-guide_mcp-tools.md.C8MiIu7F.lean.js +1 -0
- package/export-template/docs/assets/user-guide_run-comparison.md.S0TWWmLY.js +1 -0
- package/export-template/docs/assets/user-guide_run-comparison.md.S0TWWmLY.lean.js +1 -0
- package/export-template/docs/assets/user-guide_understanding-findings.md.D0R_Y-R2.js +1 -0
- package/export-template/docs/assets/user-guide_understanding-findings.md.D0R_Y-R2.lean.js +1 -0
- package/export-template/docs/contributor-guide/architecture/board-widgets.html +25 -0
- package/export-template/docs/contributor-guide/architecture/detector-contract.html +25 -0
- package/export-template/docs/contributor-guide/architecture/drill-down.html +25 -0
- package/export-template/docs/contributor-guide/architecture/impact-estimation.html +25 -0
- package/export-template/docs/contributor-guide/architecture/index.html +25 -0
- package/export-template/docs/contributor-guide/architecture/overview.html +25 -0
- package/export-template/docs/contributor-guide/architecture/state-and-history.html +25 -0
- package/export-template/docs/contributor-guide/architecture/widget-rendering.html +25 -0
- package/export-template/docs/contributor-guide/architecture/worker-protocol.html +30 -0
- package/export-template/docs/contributor-guide/contributing.html +25 -0
- package/export-template/docs/contributor-guide/development-setup.html +36 -0
- package/export-template/docs/contributor-guide/testing.html +25 -0
- package/export-template/docs/favicon.svg +4 -0
- package/export-template/docs/hashmap.json +1 -0
- package/export-template/docs/index.html +25 -0
- package/export-template/docs/package.json +1 -0
- package/export-template/docs/tuning-reference/anti-patterns.html +25 -0
- package/export-template/docs/tuning-reference/aqe.html +25 -0
- package/export-template/docs/tuning-reference/bottleneck-broadcast-sizing.html +25 -0
- package/export-template/docs/tuning-reference/bottleneck-cold-start.html +31 -0
- package/export-template/docs/tuning-reference/bottleneck-duplicate-plan-subtree.html +25 -0
- package/export-template/docs/tuning-reference/bottleneck-failures.html +30 -0
- package/export-template/docs/tuning-reference/bottleneck-gc.html +30 -0
- package/export-template/docs/tuning-reference/bottleneck-job-failure-rate.html +32 -0
- package/export-template/docs/tuning-reference/bottleneck-memory-utilization.html +31 -0
- package/export-template/docs/tuning-reference/bottleneck-retry-waste.html +25 -0
- package/export-template/docs/tuning-reference/bottleneck-shuffle.html +36 -0
- package/export-template/docs/tuning-reference/bottleneck-skew.html +38 -0
- package/export-template/docs/tuning-reference/bottleneck-slow-host.html +31 -0
- package/export-template/docs/tuning-reference/bottleneck-small-files.html +29 -0
- package/export-template/docs/tuning-reference/bottleneck-spill.html +30 -0
- package/export-template/docs/tuning-reference/bottleneck-straggler.html +31 -0
- package/export-template/docs/tuning-reference/bottleneck-tiny-tasks.html +32 -0
- package/export-template/docs/tuning-reference/bottleneck-utilization.html +29 -0
- package/export-template/docs/tuning-reference/caching.html +25 -0
- package/export-template/docs/tuning-reference/cluster-config.html +25 -0
- package/export-template/docs/tuning-reference/config.html +25 -0
- package/export-template/docs/tuning-reference/data-formats.html +25 -0
- package/export-template/docs/tuning-reference/index.html +25 -0
- package/export-template/docs/tuning-reference/intro.html +25 -0
- package/export-template/docs/tuning-reference/joins.html +25 -0
- package/export-template/docs/tuning-reference/memory-model.html +25 -0
- package/export-template/docs/tuning-reference/metrics.html +25 -0
- package/export-template/docs/tuning-reference/partitioning.html +25 -0
- package/export-template/docs/tuning-reference/pyspark.html +30 -0
- package/export-template/docs/tuning-reference/shuffle.html +25 -0
- package/export-template/docs/tuning-reference/spark-architecture.html +25 -0
- package/export-template/docs/tuning-reference/table-formats.html +25 -0
- package/export-template/docs/user-guide/alternative-log-retrieval.html +25 -0
- package/export-template/docs/user-guide/getting-started.html +27 -0
- package/export-template/docs/user-guide/mcp-tools.html +149 -0
- package/export-template/docs/user-guide/run-comparison.html +25 -0
- package/export-template/docs/user-guide/understanding-findings.html +25 -0
- package/export-template/docs/vp-icons.css +0 -0
- package/export-template/favicon.svg +4 -0
- package/export-template/index.html +115 -0
- package/export-template/parser-worker-QqyEE4m9.js +64 -0
- package/package.json +16 -3
- package/vendor-core/analyzer.js +74 -74
- package/vendor-core/cli/budgets.js +13 -27
- package/vendor-core/cli/collect-run.js +43 -19
- package/vendor-core/core-count.js +25 -27
- package/vendor-core/core-locality-ratio.js +4 -11
- package/vendor-core/core-time-series.js +6 -12
- package/vendor-core/core-usage-locality.js +3 -4
- package/vendor-core/detectors.js +256 -375
- package/vendor-core/docs-config.js +69 -21
- package/vendor-core/docs-content/chapters/01-intro.md +32 -0
- package/vendor-core/docs-content/chapters/02-spark-architecture.md +76 -0
- package/vendor-core/docs-content/chapters/03-memory-model.md +73 -0
- package/vendor-core/docs-content/chapters/04-partitioning.md +65 -0
- package/vendor-core/docs-content/chapters/05-joins.md +62 -0
- package/vendor-core/docs-content/chapters/06-shuffle.md +59 -0
- package/vendor-core/docs-content/chapters/07-data-formats.md +81 -0
- package/vendor-core/docs-content/chapters/07b-table-formats.md +56 -0
- package/vendor-core/docs-content/chapters/08-caching.md +58 -0
- package/vendor-core/docs-content/chapters/09-pyspark.md +78 -0
- package/vendor-core/docs-content/chapters/10-aqe.md +167 -0
- package/vendor-core/docs-content/chapters/11-cluster-config.md +170 -0
- package/vendor-core/docs-content/chapters/12-anti-patterns.md +171 -0
- package/vendor-core/docs-content/chapters/14-metrics.md +87 -0
- package/vendor-core/docs-content/chapters/15-config.md +93 -0
- package/vendor-core/docs-content/chapters/nav-index.json +370 -0
- package/vendor-core/docs-content/detection/cache.md +6 -0
- package/vendor-core/docs-content/detection/cfg.md +15 -0
- package/vendor-core/docs-content/detection/chrn.md +7 -0
- package/vendor-core/docs-content/detection/cold.md +4 -0
- package/vendor-core/docs-content/detection/cstor.md +4 -0
- package/vendor-core/docs-content/detection/fail.md +5 -0
- package/vendor-core/docs-content/detection/gc.md +4 -0
- package/vendor-core/docs-content/detection/host.md +5 -0
- package/vendor-core/docs-content/detection/incmp.md +6 -0
- package/vendor-core/docs-content/detection/jobs.md +4 -0
- package/vendor-core/docs-content/detection/local.md +7 -0
- package/vendor-core/docs-content/detection/mem.md +10 -0
- package/vendor-core/docs-content/detection/part.md +5 -0
- package/vendor-core/docs-content/detection/plan.md +14 -0
- package/vendor-core/docs-content/detection/retry.md +4 -0
- package/vendor-core/docs-content/detection/sfail.md +5 -0
- package/vendor-core/docs-content/detection/shape.md +5 -0
- package/vendor-core/docs-content/detection/shfl.md +4 -0
- package/vendor-core/docs-content/detection/skew.md +6 -0
- package/vendor-core/docs-content/detection/slow.md +6 -0
- package/vendor-core/docs-content/detection/spec.md +7 -0
- package/vendor-core/docs-content/detection/spill.md +7 -0
- package/vendor-core/docs-content/detection/strag.md +5 -0
- package/vendor-core/docs-content/detection/tiny.md +4 -0
- package/vendor-core/docs-content/detection/util.md +4 -0
- package/vendor-core/docs-content/diagrams/aqe-loop.dark.svg +1 -0
- package/vendor-core/docs-content/diagrams/aqe-loop.svg +1 -0
- package/vendor-core/docs-content/diagrams/broadcast-vs-shuffle.dark.svg +1 -0
- package/vendor-core/docs-content/diagrams/broadcast-vs-shuffle.svg +1 -0
- package/vendor-core/docs-content/diagrams/cache-lifecycle.dark.svg +1 -0
- package/vendor-core/docs-content/diagrams/cache-lifecycle.svg +1 -0
- package/vendor-core/docs-content/diagrams/cold-start-timeline.dark.svg +1 -0
- package/vendor-core/docs-content/diagrams/cold-start-timeline.svg +1 -0
- package/vendor-core/docs-content/diagrams/columnar-layout.dark.svg +1 -0
- package/vendor-core/docs-content/diagrams/columnar-layout.svg +1 -0
- package/vendor-core/docs-content/diagrams/container-memory.dark.svg +1 -0
- package/vendor-core/docs-content/diagrams/container-memory.svg +1 -0
- package/vendor-core/docs-content/diagrams/dag-stages.dark.svg +1 -0
- package/vendor-core/docs-content/diagrams/dag-stages.svg +1 -0
- package/vendor-core/docs-content/diagrams/driver-executor.dark.svg +1 -0
- package/vendor-core/docs-content/diagrams/driver-executor.svg +1 -0
- package/vendor-core/docs-content/diagrams/duplicate-plan-subtree.dark.svg +1 -0
- package/vendor-core/docs-content/diagrams/duplicate-plan-subtree.svg +1 -0
- package/vendor-core/docs-content/diagrams/join-strategy.dark.svg +1 -0
- package/vendor-core/docs-content/diagrams/join-strategy.svg +1 -0
- package/vendor-core/docs-content/diagrams/memory-borrowing.dark.svg +1 -0
- package/vendor-core/docs-content/diagrams/memory-borrowing.svg +1 -0
- package/vendor-core/docs-content/diagrams/memory-regions.dark.svg +1 -0
- package/vendor-core/docs-content/diagrams/memory-regions.svg +1 -0
- package/vendor-core/docs-content/diagrams/repartition-vs-coalesce.dark.svg +1 -0
- package/vendor-core/docs-content/diagrams/repartition-vs-coalesce.svg +1 -0
- package/vendor-core/docs-content/diagrams/retry-escalation-ladder.dark.svg +1 -0
- package/vendor-core/docs-content/diagrams/retry-escalation-ladder.svg +1 -0
- package/vendor-core/docs-content/diagrams/shuffle-map-reduce.dark.svg +1 -0
- package/vendor-core/docs-content/diagrams/shuffle-map-reduce.svg +1 -0
- package/vendor-core/docs-content/diagrams/spill-classification.dark.svg +1 -0
- package/vendor-core/docs-content/diagrams/spill-classification.svg +1 -0
- package/vendor-core/docs-content/diagrams/udf-execution-models.dark.svg +1 -0
- package/vendor-core/docs-content/diagrams/udf-execution-models.svg +1 -0
- package/vendor-core/docs-content/tuning/broadcast-sizing.md +78 -0
- package/vendor-core/docs-content/tuning/cold-start.md +81 -0
- package/vendor-core/docs-content/tuning/duplicate-plan-subtree.md +45 -0
- package/vendor-core/docs-content/tuning/failures.md +124 -0
- package/vendor-core/docs-content/tuning/gc.md +110 -0
- package/vendor-core/docs-content/tuning/job-failure-rate.md +101 -0
- package/vendor-core/docs-content/tuning/memory-utilization.md +58 -0
- package/vendor-core/docs-content/tuning/retry-waste.md +90 -0
- package/vendor-core/docs-content/tuning/shuffle.md +154 -0
- package/vendor-core/docs-content/tuning/skew.md +123 -0
- package/vendor-core/docs-content/tuning/slow-host.md +117 -0
- package/vendor-core/docs-content/tuning/small-files.md +99 -0
- package/vendor-core/docs-content/tuning/spill.md +114 -0
- package/vendor-core/docs-content/tuning/straggler.md +103 -0
- package/vendor-core/docs-content/tuning/tiny-tasks.md +94 -0
- package/vendor-core/docs-content/tuning/utilization.md +90 -0
- package/vendor-core/docs-site-config.js +10 -17
- package/vendor-core/efficiency-model.js +7 -13
- package/vendor-core/etl-phases.js +3 -5
- package/vendor-core/event-handlers.js +232 -134
- package/vendor-core/event-schemas.js +48 -114
- package/vendor-core/evidence-availability.js +5 -10
- package/vendor-core/evidence-report.js +72 -122
- package/vendor-core/export-data.js +48 -0
- package/vendor-core/finding-action-label.js +4 -10
- package/vendor-core/finding-filter-predicate.js +3 -7
- package/vendor-core/finding-generic-recommendation.js +112 -0
- package/vendor-core/finding-names.js +51 -0
- package/vendor-core/format-utils.js +112 -38
- package/vendor-core/impact-band.js +18 -24
- package/vendor-core/impact-estimator.js +38 -74
- package/vendor-core/ingest.js +7 -13
- package/vendor-core/job-groups.js +3 -6
- package/vendor-core/list-runs.js +278 -0
- package/vendor-core/load-vendored.js +6 -12
- package/vendor-core/log-header-peek.js +81 -0
- package/vendor-core/lz4-block.js +4 -6
- package/vendor-core/mcp-server-factory.js +38 -8
- package/vendor-core/mcp-tools.js +105 -76
- package/vendor-core/model-assembler.js +8 -16
- package/vendor-core/occupancy.js +5 -9
- package/vendor-core/parser-worker.js +18 -27
- package/vendor-core/plan-dot.js +2 -5
- package/vendor-core/plan-duration-attribution.js +78 -29
- package/vendor-core/plan-graph-model.js +126 -69
- package/vendor-core/plan-node-detail.js +31 -17
- package/vendor-core/plan-summary.js +19 -8
- package/vendor-core/recommendation-rollup.js +35 -39
- package/vendor-core/redact.js +72 -16
- package/vendor-core/rolling-log-reassembly.js +4 -6
- package/vendor-core/run-comparison.js +65 -70
- package/vendor-core/scaling-sim.js +5 -7
- package/vendor-core/session-snapshot.js +1 -1
- package/vendor-core/shs-fetch.js +4 -6
- package/vendor-core/shs-load.js +9 -13
- package/vendor-core/shs-request.js +1 -1
- package/vendor-core/stage-quantiles.js +14 -0
- package/vendor-core/types.js +78 -18
- package/vendor-core/wasted-core-hours.js +7 -12
|
@@ -0,0 +1,101 @@
|
|
|
1
|
+
# Job Failure Rate
|
|
2
|
+
|
|
3
|
+
<span class="tag">FAIL-RATE</span>
|
|
4
|
+
|
|
5
|
+
## What it is
|
|
6
|
+
|
|
7
|
+
Not every task retry threatens the job. `spark.task.maxFailures` (default `4`) counts
|
|
8
|
+
"continuous failures of any particular task before giving up on the job. The total number of
|
|
9
|
+
failures spread across different tasks will not cause the job to fail; a particular task has to
|
|
10
|
+
fail this number of attempts continuously. If any attempt succeeds, the failure count for the
|
|
11
|
+
task will be reset"[^1]. A task that fails once and then succeeds on retry resets its counter and
|
|
12
|
+
never counts toward job abandonment: that's a transient retry, and by itself it's harmless.
|
|
13
|
+
Whole-job abandonment is a different, layered escalation above that: it only happens once an
|
|
14
|
+
unbroken run of failures on the same task reaches the configured limit (3 retries allowed by
|
|
15
|
+
default, since allowed retries = value − 1), or once higher-level retry budgets above the task
|
|
16
|
+
level are themselves exhausted.
|
|
17
|
+
|
|
18
|
+
## How it's detected
|
|
19
|
+
|
|
20
|
+
The layers above task-level retries are what turn isolated failures into a failed job. At the
|
|
21
|
+
stage level, `spark.stage.maxConsecutiveAttempts` (default `4`) bounds "the number of consecutive
|
|
22
|
+
stage attempts allowed before a stage is aborted"[^1]. By default,
|
|
23
|
+
`spark.stage.ignoreDecommissionFetchFailure` (`true`, since 3.4.0) excludes fetch failures caused
|
|
24
|
+
by graceful executor decommission from counting toward that limit, so a decommission-triggered
|
|
25
|
+
`FetchFailed` doesn't push a stage toward abortion the way a genuine repeated fetch failure
|
|
26
|
+
would[^1]. On YARN and Kubernetes there's a further application-level ceiling:
|
|
27
|
+
`spark.executor.maxNumFailures` (default `numExecutors * 2`, minimum 3, since 3.5.0) is "the
|
|
28
|
+
maximum number of executor failures before failing the application," while
|
|
29
|
+
`spark.executor.failuresValidityInterval` (since 3.5.0) lets failures spaced far enough apart be
|
|
30
|
+
"considered independent and not accumulate towards the attempt count"[^1]. The TaskScheduler owns
|
|
31
|
+
retries at the task level, but it is "the DAGScheduler that ultimately declares the job to have
|
|
32
|
+
failed"[^2] once those retry budgets are exhausted, and because "a Spark job corresponds to one
|
|
33
|
+
action" with a fixed DAG once that action is called[^3], an aborted stage cascades directly into
|
|
34
|
+
failure of the job built on it.
|
|
35
|
+
|
|
36
|
+
<img class="light-only" src="../diagrams/retry-escalation-ladder.svg" alt="How a failure escalates from a task retry up through stage resubmission, the executor failure ceiling, and a failed application attempt before the job is aborted.">
|
|
37
|
+
<img class="dark-only" src="../diagrams/retry-escalation-ladder.dark.svg" alt="How a failure escalates from a task retry up through stage resubmission, the executor failure ceiling, and a failed application attempt before the job is aborted.">
|
|
38
|
+
|
|
39
|
+
## Why it matters
|
|
40
|
+
|
|
41
|
+
A rising failure rate that traces back to exhausted retry budgets, rather than to a single
|
|
42
|
+
one-off task failure, points at a systemic problem (an unstable executor, a genuinely broken
|
|
43
|
+
fetch path, or a resource ceiling being hit repeatedly) since none of these budgets trip on a
|
|
44
|
+
single transient hiccup. Job-level failure is also distinct from application-attempt failure: on
|
|
45
|
+
YARN, "each application may have multiple attempts," and the history server displays failed
|
|
46
|
+
attempts alongside "any ongoing incomplete attempt or the final successful attempt"[^4]: a
|
|
47
|
+
separate, higher layer of retry (driver/application restart) above the in-job task and stage
|
|
48
|
+
retries covered here.
|
|
49
|
+
|
|
50
|
+
## How to fix it
|
|
51
|
+
|
|
52
|
+
- Check which budget was exhausted before treating a job failure as a single root cause: a
|
|
53
|
+
single task failing `spark.task.maxFailures` times consecutively, a stage being resubmitted
|
|
54
|
+
until `spark.stage.maxConsecutiveAttempts` is reached and aborted, or (on YARN/Kubernetes)
|
|
55
|
+
cumulative distinct executor failures exceeding `spark.executor.maxNumFailures`[^1] each point
|
|
56
|
+
at a different underlying problem.
|
|
57
|
+
- If failures are decommission-triggered `FetchFailed`s during a normal scale-down, confirm
|
|
58
|
+
`spark.stage.ignoreDecommissionFetchFailure` is enabled (default `true` since 3.4.0) so they
|
|
59
|
+
aren't inflating the stage-abort count[^1].
|
|
60
|
+
- On YARN/Kubernetes, if unrelated executor failures spread far apart in time are tripping
|
|
61
|
+
`spark.executor.maxNumFailures`, `spark.executor.failuresValidityInterval` (since 3.5.0) can
|
|
62
|
+
let sufficiently-spaced failures stop accumulating toward the same count[^1].
|
|
63
|
+
- See also [retry waste](#bottleneck-retry-waste): the same executor-loss and `FetchFailed`
|
|
64
|
+
causes that eventually exhaust these budgets and fail a job are, below that threshold, also
|
|
65
|
+
the source of wasted executor time on jobs that ultimately succeed.
|
|
66
|
+
|
|
67
|
+
> **PySpark:** these are session-level configs, settable without a `spark-submit` flag:
|
|
68
|
+
> `spark.conf.set("spark.task.maxFailures", "4")`, and the stage/executor-level equivalents the
|
|
69
|
+
> same way.
|
|
70
|
+
|
|
71
|
+
The retry-budget ladder: defaults shown; raise a budget only once you know which layer is tripping:
|
|
72
|
+
|
|
73
|
+
```properties
|
|
74
|
+
# Task level: consecutive failures of one task before the job is abandoned (default 4 => 3 retries)
|
|
75
|
+
spark.task.maxFailures=4
|
|
76
|
+
|
|
77
|
+
# Stage level: consecutive stage attempts before the stage is aborted (default 4)
|
|
78
|
+
spark.stage.maxConsecutiveAttempts=4
|
|
79
|
+
|
|
80
|
+
# Don't count graceful-decommission fetch failures toward the stage-abort limit (default true since Spark 3.4)
|
|
81
|
+
spark.stage.ignoreDecommissionFetchFailure=true
|
|
82
|
+
```
|
|
83
|
+
|
|
84
|
+
|
|
85
|
+
## Limitations / false-positive risk
|
|
86
|
+
|
|
87
|
+
A nonzero job-failure rate can be dominated by a single repeatedly-failing job rather than a
|
|
88
|
+
systemic issue, so the rate alone doesn't tell you whether the problem is broad or isolated.
|
|
89
|
+
Retried jobs that eventually succeed still count toward attempts, which can inflate the rate
|
|
90
|
+
above what the eventual outcomes justify.
|
|
91
|
+
|
|
92
|
+
|
|
93
|
+
## Related
|
|
94
|
+
|
|
95
|
+
- **Retry budgets & executor stability:** [Cluster Tuning](#cluster-config)
|
|
96
|
+
- **Avoiding the failures upstream:** [Anti-Patterns](#anti-patterns)
|
|
97
|
+
|
|
98
|
+
[^1]: [Configuration (Spark)](https://spark.apache.org/docs/latest/configuration.html)
|
|
99
|
+
[^2]: [Spark: Anatomy of Spark Application](https://luminousmen.com/post/spark-anatomy-of-spark-application)
|
|
100
|
+
[^3]: *High Performance Spark, 2nd Edition*, Karau, Polak & Warren, ch. 2
|
|
101
|
+
[^4]: [Monitoring and Instrumentation](https://spark.apache.org/docs/latest/monitoring.html#spark-history-server)
|
|
@@ -0,0 +1,58 @@
|
|
|
1
|
+
# Memory Utilization
|
|
2
|
+
|
|
3
|
+
<span class="tag">MEM</span>
|
|
4
|
+
|
|
5
|
+
## What it is
|
|
6
|
+
|
|
7
|
+
An executor is a single JVM process that gets a fixed memory allocation when the application starts, and it holds that whole allocation for its entire lifetime, whether or not it has work to run[^3]. Two things can leave that memory band poorly used: cores sitting idle inside a held executor, and a held executor whose memory drains of active work but is never released. This finding covers both.
|
|
8
|
+
|
|
9
|
+
An executor's core count is its concurrency ceiling. `spark.executor.cores` sets how many tasks it can run at once, so `--executor-cores 5` caps that executor at five concurrent tasks[^1]. The scheduler turns cores into task slots from `spark.executor.cores` and `spark.task.cpus` (minimum 1), with `spark.task.cpus` defaulting to one core per task[^2]. Cores go idle whenever there are fewer runnable tasks than slots: a stage with fewer partitions than the total slots across executors leaves slots empty, and so does the tail of a stage where a [straggler](#bottleneck-straggler) or two keep running after their peers finish. Raising `spark.task.cpus` above 1 has the same effect from the other side, since each task then reserves several cores and fewer run side by side. Through all of it the JVM keeps its fixed heap[^3].
|
|
10
|
+
|
|
11
|
+
## How it's detected
|
|
12
|
+
|
|
13
|
+
| Rule | Signal | Status |
|
|
14
|
+
|---|---|---|
|
|
15
|
+
| idleCores | Task parallelism below allocated cores while the executor is held | Validated |
|
|
16
|
+
| memoryBand | A whole executor's memory band held with little or no active work (dynamic-allocation idle timeouts) | Validated |
|
|
17
|
+
| wasteModel | Peak used memory well below allocated memory (unused headroom) | Experimental |
|
|
18
|
+
|
|
19
|
+
## Why it matters
|
|
20
|
+
|
|
21
|
+
A held executor is allocation you pay for regardless of how busy it is. On YARN that is the memory YARN grants the container; on Kubernetes it is the pod memory limit[^3]. When cores idle or a whole executor lingers with nothing to do, that fixed band is reserved without returning work, so the cost lands whether or not tasks are running.
|
|
22
|
+
|
|
23
|
+
## How to fix it
|
|
24
|
+
|
|
25
|
+
Keep executors from being oversized in the first place. The guidance is against packing every core of a node into one fat executor: assigning all 16 cores of a node to a single executor hurts HDFS throughput and drives excessive [garbage collection](#bottleneck-gc), so aim for a balance between tiny (one core per executor) and fat (one executor per node) sizing[^4]. On YARN, leave cores for the OS and Hadoop daemons instead of handing 100% of a node to Spark containers[^1].
|
|
26
|
+
|
|
27
|
+
For the held memory band, lean on [dynamic allocation](#cluster-config) to reclaim executors once the work drains. It requests executors when tasks back up and frees them when they go idle[^1]. Executors are added in rounds once tasks have been pending for `spark.dynamicAllocation.schedulerBacklogTimeout` (default 1s), then again every `spark.dynamicAllocation.sustainedSchedulerBacklogTimeout` while the backlog holds[^5]. On the release side, an executor is removed after it has been idle longer than `spark.dynamicAllocation.executorIdleTimeout`[^5].
|
|
28
|
+
|
|
29
|
+
Cached data is the trap here. By default an executor holding [cached blocks](#caching) is never removed, governed by `spark.dynamicAllocation.cachedExecutorIdleTimeout`, whose default is infinity[^2]. Such an executor keeps its full memory band indefinitely with no active tasks unless you set that timeout to a finite value, or turn on `spark.shuffle.service.fetch.rdd.enabled` so executors holding only disk-persisted blocks are treated as idle after `spark.dynamicAllocation.executorIdleTimeout` and released[^5]. To avoid holding barely-used executors at all, lower `spark.dynamicAllocation.executorAllocationRatio` (default 1.0, full parallelism) toward 0.5, since with small tasks full-parallelism allocation can request executors that never do any work[^2].
|
|
30
|
+
|
|
31
|
+
```properties
|
|
32
|
+
# Balanced executor sizing, not one fat executor per node
|
|
33
|
+
spark.executor.cores=5
|
|
34
|
+
|
|
35
|
+
# Reclaim idle executors; give cached holders a finite timeout
|
|
36
|
+
spark.dynamicAllocation.enabled=true
|
|
37
|
+
spark.dynamicAllocation.cachedExecutorIdleTimeout=300s
|
|
38
|
+
spark.dynamicAllocation.executorAllocationRatio=0.5
|
|
39
|
+
```
|
|
40
|
+
|
|
41
|
+
### Reading the wasted-memory estimate <span class="tag">EXPERIMENTAL</span>
|
|
42
|
+
|
|
43
|
+
The wasteModel rule estimates unused ("wasted") allocated memory from the gap between an executor's peak used memory and its allocated total. Treat it as a rough buffer heuristic, not a measurement. Peak used memory is a high-water mark rather than a ceiling: Spark records it under `peakMemoryMetrics.*`, and each figure is the maximum its pool ever reached, so peak sits at or below the allocated total by construction and the space between them is the executor's unused headroom[^6]. Turning that headroom into a wasted number stays approximate for concrete reasons. The peak heap figure counts garbage, since `JVMHeapMemory` is the peak used heap including "the amount of memory occupied by both live objects and garbage objects that have not been collected," so it overstates the live footprint[^6]. The region boundaries also move: `totalOnHeapStorageMemory` and `totalOffHeapStorageMemory` "can vary over time, depending on the MemoryManager implementation," so there is no single fixed allocated-to-storage value to subtract a peak from[^6]. Peak metrics only reach the event log when `spark.eventLog.logStageExecutorMetrics` is true[^6], so without that setting there is nothing to estimate from. Use the number as a hint that an executor may be oversized, then confirm against sizing before acting on it.
|
|
44
|
+
|
|
45
|
+
## Confidence
|
|
46
|
+
|
|
47
|
+
The idleCores and memoryBand rules are validated: they rest on documented Spark concurrency and dynamic-allocation behavior. The wasteModel estimate is low-confidence and experimental. It is a buffer heuristic derived from peak-versus-allocated sampling, not an exact accounting of unused memory, so it should steer investigation rather than settle it.
|
|
48
|
+
|
|
49
|
+
## Limitations / false-positive risk
|
|
50
|
+
|
|
51
|
+
Idle cores are not always waste. The tail of a stage legitimately leaves slots empty while a straggler or two finish[^2], so a snapshot of low parallelism can reflect a normal straggler tail rather than chronic under-utilization. The wasted-memory estimate depends on peak-memory sampling and is only as good as that sampling: the peak includes uncollected garbage, the managed-region boundaries shift under the MemoryManager, and the figures are absent entirely unless `spark.eventLog.logStageExecutorMetrics` is enabled[^6].
|
|
52
|
+
|
|
53
|
+
[^1]: [How to Tune Your Apache Spark Jobs (Part 2)](https://blog.cloudera.com/how-to-tune-your-apache-spark-jobs-part-2/)
|
|
54
|
+
[^2]: [Configuration — Spark](https://spark.apache.org/docs/latest/configuration.html)
|
|
55
|
+
[^3]: [Dive into Spark memory](https://luminousmen.com/post/dive-into-spark-memory)
|
|
56
|
+
[^4]: [Distribution of executors, cores and memory for a Spark application](https://raw.githubusercontent.com/spoddutur/spark-notes/master/distribution_of_executors_cores_and_memory_for_spark_application.md)
|
|
57
|
+
[^5]: [Job Scheduling — Spark](https://spark.apache.org/docs/latest/job-scheduling.html)
|
|
58
|
+
[^6]: [Monitoring — Spark](https://spark.apache.org/docs/latest/monitoring.html)
|
|
@@ -0,0 +1,90 @@
|
|
|
1
|
+
# Retry Waste
|
|
2
|
+
|
|
3
|
+
<span class="tag">RETRY</span>
|
|
4
|
+
|
|
5
|
+
## What it is
|
|
6
|
+
|
|
7
|
+
When a task attempt is superseded by a later retry, every bit of executor time the abandoned
|
|
8
|
+
attempt spent counts for nothing. Every attempt, successful or not, accrues the standard task
|
|
9
|
+
metrics: `executorRunTime` ("elapsed time the executor spent running this task," including
|
|
10
|
+
time fetching shuffle data) and `executorCpuTime`[^1], but only the winning attempt's output
|
|
11
|
+
survives. The TaskScheduler is what triggers the do-over: "if an executor dies or a task throws
|
|
12
|
+
an exception, the TaskScheduler resubmits the task to another executor, respecting Spark's task
|
|
13
|
+
locality preferences"[^2]. Whatever the discarded attempt had already run is gone; none of it
|
|
14
|
+
carries forward into the job's result.
|
|
15
|
+
|
|
16
|
+
<img class="light-only" src="../diagrams/retry-escalation-ladder.svg" alt="Where a superseded retry sits in the escalation ladder, below the stage and application budgets that a job crosses only when the same failures keep recurring.">
|
|
17
|
+
<img class="dark-only" src="../diagrams/retry-escalation-ladder.dark.svg" alt="Where a superseded retry sits in the escalation ladder, below the stage and application budgets that a job crosses only when the same failures keep recurring.">
|
|
18
|
+
|
|
19
|
+
## How it's detected
|
|
20
|
+
|
|
21
|
+
A stage's superseded task attempts are the signal: a later retry supersedes an earlier attempt
|
|
22
|
+
after [an executor is lost](#bottleneck-failures) (`ExecutorLostFailure`) or after a shuffle `FetchFailed`, since "if an
|
|
23
|
+
executor dies or a task throws an exception, the TaskScheduler resubmits the task to another
|
|
24
|
+
executor, respecting Spark's task locality preferences"[^2]. Because only the winning attempt's
|
|
25
|
+
output survives, the abandoned attempt's `executorRunTime` (the elapsed executor time defined
|
|
26
|
+
above) is the wasted time.
|
|
27
|
+
|
|
28
|
+
The two causes are handled differently by Spark's own machinery, which is how they can be told
|
|
29
|
+
apart after the fact. An `ExecutorLostFailure` removes the executor (and any shuffle map output
|
|
30
|
+
it had already written) from the pool outright[^3]. A shuffle `FetchFailed`, by contrast, is a
|
|
31
|
+
read-side failure against a still-alive remote executor, and Spark's own exclusion machinery
|
|
32
|
+
treats it as a distinct category from a general executor loss:
|
|
33
|
+
`spark.excludeOnFailure.killExcludedExecutors` governs whether Spark kills executors "excluded
|
|
34
|
+
on fetch failure or excluded for the entire application"[^4], a separate bucket from executor
|
|
35
|
+
loss in Spark's accounting. That distinction (executor-loss bucket vs. fetch-failure bucket)
|
|
36
|
+
is the basis for identifying which cause dominated a given stage's retry waste.
|
|
37
|
+
|
|
38
|
+
## Why it matters
|
|
39
|
+
|
|
40
|
+
A lost executor can waste more than just the in-flight task's time: "dynamic allocation may
|
|
41
|
+
remove an executor before the shuffle completes, in which case the shuffle files written by that
|
|
42
|
+
executor must be recomputed unnecessarily"[^3], so prior map-output work from that same executor
|
|
43
|
+
may need redoing too. A `FetchFailed` doesn't carry that same risk, since the remote executor
|
|
44
|
+
stays alive and only the one failed read is lost[^4].
|
|
45
|
+
|
|
46
|
+
## How to fix it
|
|
47
|
+
|
|
48
|
+
- Check the retried-attempt count and accumulated `executorRunTime`/`executorCpuTime` on the
|
|
49
|
+
abandoned attempts directly, rather than trusting the stage's final wall-clock duration:
|
|
50
|
+
a clean-looking stage can still be hiding significant wasted compute[^1].
|
|
51
|
+
- If the dominant cause is executor loss, investigate node-level stability (OOM-kills, node
|
|
52
|
+
death, aggressive dynamic-allocation deallocation) rather than the task logic itself, since
|
|
53
|
+
the lost executor's prior shuffle-write work may also need recomputing[^3].
|
|
54
|
+
- If the dominant cause is `FetchFailed`, check whether `spark.excludeOnFailure.killExcludedExecutors`
|
|
55
|
+
is causing repeated exclusion churn on otherwise-live executors, and treat it separately from
|
|
56
|
+
outright executor loss[^4].
|
|
57
|
+
- See also [job failure rate](#bottleneck-job-failure-rate): retry waste can pile up quietly on
|
|
58
|
+
a job that ultimately succeeds, while the same underlying causes (exhausted at a higher
|
|
59
|
+
threshold) are what push a job to fail outright.
|
|
60
|
+
|
|
61
|
+
> **PySpark:** there's no attempt-level API to recover a discarded attempt's `executorRunTime`:
|
|
62
|
+
> pull it from the event log or history server UI's per-stage task list, filtering for tasks
|
|
63
|
+
> whose attempt number is greater than zero.
|
|
64
|
+
|
|
65
|
+
The preventive lever: keep an executor's shuffle output available if it is lost, so retries don't recompute prior map work[^5]:
|
|
66
|
+
|
|
67
|
+
```properties
|
|
68
|
+
spark.dynamicAllocation.shuffleTracking.enabled=true
|
|
69
|
+
```
|
|
70
|
+
|
|
71
|
+
## Limitations / false-positive risk
|
|
72
|
+
|
|
73
|
+
Some retry waste is unavoidable: transient cluster faults will always cost a few
|
|
74
|
+
superseded attempts, and no tuning drives that to zero. The metric attributes burned
|
|
75
|
+
executor time to the abandoned attempts, so its accuracy tracks the failure cause. When
|
|
76
|
+
a lost executor also forces recompute of prior map output, the per-attempt
|
|
77
|
+
`executorRunTime` undercounts the true cost; when an attempt fails almost immediately, it
|
|
78
|
+
overcounts the compute actually lost. Read the number as a directional signal, not an
|
|
79
|
+
exact ledger.
|
|
80
|
+
|
|
81
|
+
## Related
|
|
82
|
+
|
|
83
|
+
- **Shuffle recompute on executor loss:** [Shuffle](#shuffle)
|
|
84
|
+
- **Executor stability & dynamic allocation:** [Cluster Tuning](#cluster-config)
|
|
85
|
+
|
|
86
|
+
[^1]: [Monitoring and Instrumentation](https://spark.apache.org/docs/latest/monitoring.html)
|
|
87
|
+
[^2]: [Spark: Anatomy of Spark Application](https://luminousmen.com/post/spark-anatomy-of-spark-application)
|
|
88
|
+
[^3]: [Job Scheduling (Spark)](https://spark.apache.org/docs/latest/job-scheduling.html)
|
|
89
|
+
[^4]: [Configuration (Spark)](https://spark.apache.org/docs/latest/configuration.html)
|
|
90
|
+
[^5]: [Configuration (Spark)](https://spark.apache.org/docs/latest/configuration.html)
|
|
@@ -0,0 +1,154 @@
|
|
|
1
|
+
# Shuffle I/O
|
|
2
|
+
|
|
3
|
+
<span class="tag">SHFL</span>
|
|
4
|
+
|
|
5
|
+
## What it is
|
|
6
|
+
|
|
7
|
+
Shuffle I/O is the data movement Spark performs between stages whenever an operation needs
|
|
8
|
+
to redistribute data across the cluster so that related rows land on the same partition.
|
|
9
|
+
At the DataFrame/SQL level, `groupBy()`, `join()`, `agg()`, `sortBy()`, and `reduceByKey()`-style
|
|
10
|
+
aggregations are the classic wide transformations that force this exchange[^1]. Non-broadcast
|
|
11
|
+
join strategies (shuffle hash join, shuffle sort-merge join, and shuffle-and-replicated nested
|
|
12
|
+
loop / Cartesian product join) all shuffle data across executors the same
|
|
13
|
+
way[^1]. At the RDD layer, the operations that can cause a shuffle are repartitioning
|
|
14
|
+
(`repartition`, `coalesce`), `'ByKey` operations other than counting (`groupByKey`,
|
|
15
|
+
`reduceByKey`), and join-family operations (`cogroup`, `join`)[^2].
|
|
16
|
+
|
|
17
|
+
Spark can skip the shuffle when it already knows the data layout: Storage Partition Join
|
|
18
|
+
avoids it entirely when Spark can use partitioning already reported by a compatible V2 data
|
|
19
|
+
source[^3], and classic Hive-style bucketing has the same effect: once both sides of a join
|
|
20
|
+
are bucketed and sorted the same way, the physical plan shows no `Exchange` operator[^4].
|
|
21
|
+
`DataFrameWriter.partitionBy`, by contrast, does not trigger a shuffle by itself on write[^5].
|
|
22
|
+
|
|
23
|
+
At the metrics level, shuffle reads split cross-node traffic from same-host traffic:
|
|
24
|
+
`remoteBytesRead` counts bytes read from a remote executor, `localBytesRead` counts bytes
|
|
25
|
+
read from local disk, and `totalBytesRead` is their sum[^6].
|
|
26
|
+
|
|
27
|
+
## How it's detected
|
|
28
|
+
|
|
29
|
+
| Shuffle read/write bytes | Level |
|
|
30
|
+
|---|---|
|
|
31
|
+
| > 50 MB | Info |
|
|
32
|
+
| > 500 MB | Warning |
|
|
33
|
+
| > 1 GB | Critical |
|
|
34
|
+
|
|
35
|
+
Beyond raw byte volume, the executor-side wait is captured by `fetchWaitTime`: time a task
|
|
36
|
+
spends blocked on a remote shuffle block it needs next, not counting time spent prefetching
|
|
37
|
+
other blocks in the background[^6]. That wait time counts toward the task's overall
|
|
38
|
+
`executorRunTime`, since Spark's wall-clock task-time metric explicitly includes time
|
|
39
|
+
fetching shuffle data[^6].
|
|
40
|
+
|
|
41
|
+
## Why it matters
|
|
42
|
+
|
|
43
|
+
As shuffle volume grows, the underlying pull-based mechanism gets less efficient: the number
|
|
44
|
+
of shuffle blocks grows quadratically with mapper × reducer count while individual block
|
|
45
|
+
sizes shrink to only tens of KB, which is inefficient for disk-backed random reads[^7].
|
|
46
|
+
Because fetch-wait time is part of a task's measured execution window, that inefficiency
|
|
47
|
+
shows up directly as added task time rather than as separate, hidden overhead[^6].
|
|
48
|
+
|
|
49
|
+
## How to fix it
|
|
50
|
+
|
|
51
|
+
- Let [Adaptive Query Execution](#aqe) coalesce small post-shuffle partitions automatically at
|
|
52
|
+
runtime (`spark.sql.adaptive.coalescePartitions.enabled`, default `true`) instead of only
|
|
53
|
+
hand-tuning a fixed partition count[^3].
|
|
54
|
+
- Review `spark.sql.shuffle.partitions` (default `200`), which applies uniformly to every
|
|
55
|
+
`join()`, `groupBy()`, and aggregation regardless of how much data is actually moving[^8].
|
|
56
|
+
There's no fixed formula for the right value: pull it down toward the executor core
|
|
57
|
+
count for small or streaming workloads[^1], and otherwise favor over- to
|
|
58
|
+
under-provisioning, since Spark's low per-task overhead makes it safer to have too many
|
|
59
|
+
tasks than too few[^9].
|
|
60
|
+
- Where possible, avoid the shuffle altogether: bucket both sides of a join identically, or
|
|
61
|
+
rely on Storage Partition Join for compatible sources, so the physical plan drops the
|
|
62
|
+
`Exchange` node[^3][^4].
|
|
63
|
+
- If map tasks are I/O-bound writing many shuffle files, increase
|
|
64
|
+
`spark.shuffle.file.buffer` (default `32k`) to reduce disk seeks and system calls on the
|
|
65
|
+
write side[^8]; Learning Spark's tuning table recommends bumping it to 1 MB for large
|
|
66
|
+
jobs[^1].
|
|
67
|
+
- If executors have memory to spare, increase `spark.reducer.maxSizeInFlight` (default
|
|
68
|
+
`48m`) so reducers can pull more map output concurrently and need fewer fetch rounds[^8].
|
|
69
|
+
- At larger scale, enable the external shuffle service (effectively required for [dynamic
|
|
70
|
+
allocation](#cluster-config), since it lets shuffle files be served after an executor is removed[^10]) and
|
|
71
|
+
consider push-based (Magnet) shuffle, which converts many small random reads into large
|
|
72
|
+
sequential reads of pre-merged chunks[^7].
|
|
73
|
+
|
|
74
|
+
> **PySpark:** adjust the shuffle partition count directly from a running session with
|
|
75
|
+
> `spark.conf.set("spark.sql.shuffle.partitions", 100)`.
|
|
76
|
+
|
|
77
|
+
A shuffle-tuning starting point: defaults shown, comments say which way to move:
|
|
78
|
+
|
|
79
|
+
```properties
|
|
80
|
+
# Let AQE coalesce small post-shuffle partitions at runtime
|
|
81
|
+
spark.sql.adaptive.enabled=true
|
|
82
|
+
spark.sql.adaptive.coalescePartitions.enabled=true
|
|
83
|
+
|
|
84
|
+
# Baseline shuffle partition count (default 200); pull toward executor-core count for small jobs
|
|
85
|
+
spark.sql.shuffle.partitions=200
|
|
86
|
+
|
|
87
|
+
# Reduce write-side disk seeks on large shuffles (default 32k; Learning Spark suggests 1m)
|
|
88
|
+
spark.shuffle.file.buffer=1m
|
|
89
|
+
|
|
90
|
+
# Let reducers pull more map output per fetch round when executors have spare memory (default 48m)
|
|
91
|
+
spark.reducer.maxSizeInFlight=48m
|
|
92
|
+
```
|
|
93
|
+
|
|
94
|
+
## Partition sizing <span class="tag">PART</span>
|
|
95
|
+
|
|
96
|
+
Adaptive Query Execution re-optimizes the plan while the query runs: as each shuffle stage
|
|
97
|
+
materializes, it reads the real shuffle-file sizes and resizes partitions before launching the
|
|
98
|
+
stages downstream[^11]. That runtime feedback is what corrects a partition grid that static
|
|
99
|
+
estimates would get wrong.
|
|
100
|
+
|
|
101
|
+
The coarse starting grid is `spark.sql.shuffle.partitions` (default `200`), the fixed partition
|
|
102
|
+
count Spark uses when shuffling for joins or aggregations[^14]. A single number cannot fit every
|
|
103
|
+
stage. Spread a few megabytes across 200 partitions and most cores sit idle on a handful of rows
|
|
104
|
+
each; push hundreds of gigabytes through the same 200 and every executor is overloaded, often into
|
|
105
|
+
memory errors. For small or streaming workloads the 200 default is usually too high, and pulling it
|
|
106
|
+
toward the executor-core count thins out the flood of tiny partitions crossing the network[^1]. The
|
|
107
|
+
pattern AQE encourages is the opposite of hand-tuning that number: leave the initial count large and
|
|
108
|
+
let runtime coalescing combine adjacent small partitions instead[^11].
|
|
109
|
+
|
|
110
|
+
**Partitions too small (low parallelism payoff).** `spark.sql.adaptive.coalescePartitions.enabled`
|
|
111
|
+
(default `true`) merges contiguous shuffle partitions up to a target size so a stage does not end up
|
|
112
|
+
with many tiny tasks[^12]. The target is `spark.sql.adaptive.advisoryPartitionSizeInBytes`, the
|
|
113
|
+
advisory shuffle-partition size AQE steers toward during adaptive optimization, default `64 MB`[^12].
|
|
114
|
+
|
|
115
|
+
**Partitions too big or skewed.** `spark.sql.adaptive.skewJoin.enabled` (default `true`) handles
|
|
116
|
+
skew in shuffled joins by splitting the oversized partitions and replicating the matching side where
|
|
117
|
+
needed[^13]. A partition only counts as skewed when it clears both bars: larger than
|
|
118
|
+
`spark.sql.adaptive.skewJoin.skewedPartitionFactor` (default `5.0`) times the median partition size,
|
|
119
|
+
and larger than `spark.sql.adaptive.skewJoin.skewedPartitionThresholdInBytes` (default `256 MB`)[^13].
|
|
120
|
+
Once it qualifies, the partition is split back down toward the advisory size.
|
|
121
|
+
|
|
122
|
+
So the two knobs divide the labor: `spark.sql.shuffle.partitions` lays down the coarse initial grid,
|
|
123
|
+
and `advisoryPartitionSizeInBytes` (64 MB) is the size AQE aims each partition at from both
|
|
124
|
+
directions. Partitions below it get coalesced; skewed partitions past the 256 MB / 5.0-factor bar get
|
|
125
|
+
split toward it.
|
|
126
|
+
|
|
127
|
+
## Limitations / false-positive risk
|
|
128
|
+
|
|
129
|
+
A high shuffle-byte count is not automatically a defect. A wide transformation such as a large
|
|
130
|
+
`join()` or `groupBy()` legitimately has to move that data, so the volume can be an inherent property
|
|
131
|
+
of the query rather than something worth fixing. The byte thresholds that raise Info, Warning, and
|
|
132
|
+
Critical levels are heuristic cutoffs, not measured limits for a given cluster, so treat them as a
|
|
133
|
+
prompt to look rather than a verdict.
|
|
134
|
+
|
|
135
|
+
|
|
136
|
+
## Related
|
|
137
|
+
|
|
138
|
+
- **The mechanism:** [Shuffle](#shuffle)
|
|
139
|
+
- **Tuning parallelism:** [Partitioning](#partitioning)
|
|
140
|
+
|
|
141
|
+
[^1]: *Learning Spark, 2nd Edition*, Damji, Wenig, Das, Lee, ch. 7
|
|
142
|
+
[^2]: [RDD Programming Guide](https://spark.apache.org/docs/latest/rdd-programming-guide.html)
|
|
143
|
+
[^3]: [Performance Tuning (Spark SQL, DataFrames and Datasets Guide)](https://spark.apache.org/docs/latest/sql-performance-tuning.html)
|
|
144
|
+
[^4]: [Bucketing (The Internals of Spark SQL)](https://books.japila.pl/spark-sql-internals/bucketing/)
|
|
145
|
+
[^5]: [Spark Tips: Partition Tuning](https://luminousmen.com/post/spark-tips-partition-tuning)
|
|
146
|
+
[^6]: [Monitoring and Instrumentation](https://spark.apache.org/docs/latest/monitoring.html#spark-history-server)
|
|
147
|
+
[^7]: [SPARK-30602: Support push-based shuffle to improve shuffle efficiency](https://issues.apache.org/jira/browse/SPARK-30602)
|
|
148
|
+
[^8]: [Configuration (Spark)](https://spark.apache.org/docs/latest/configuration.html)
|
|
149
|
+
[^9]: [How to Tune Your Apache Spark Jobs (Part 2): Cloudera Engineering Blog](https://blog.cloudera.com/how-to-tune-your-apache-spark-jobs-part-2/)
|
|
150
|
+
[^10]: [Job Scheduling: Dynamic Resource Allocation](https://spark.apache.org/docs/latest/job-scheduling.html)
|
|
151
|
+
[^11]: [Adaptive Query Execution: Speeding Up Spark SQL at Runtime (Databricks)](https://www.databricks.com/blog/2020/05/29/adaptive-query-execution-speeding-up-spark-sql-at-runtime.html)
|
|
152
|
+
[^12]: [Performance Tuning: Coalescing Post Shuffle Partitions](https://spark.apache.org/docs/latest/sql-performance-tuning.html)
|
|
153
|
+
[^13]: [Configuration (Spark)](https://spark.apache.org/docs/latest/configuration.html)
|
|
154
|
+
[^14]: [SQLConf: shuffle-partition defaults (Spark source)](https://raw.githubusercontent.com/apache/spark/v3.5.0/sql/catalyst/src/main/scala/org/apache/spark/sql/internal/SQLConf.scala)
|
|
@@ -0,0 +1,123 @@
|
|
|
1
|
+
# Task Skew
|
|
2
|
+
|
|
3
|
+
<span class="tag">SKEW</span>
|
|
4
|
+
|
|
5
|
+
## What it is
|
|
6
|
+
|
|
7
|
+
Task skew happens when a shuffle produces one or a few oversized partitions instead of a
|
|
8
|
+
roughly even split. Spark's own skew-join optimizer defines a partition as skewed once it is
|
|
9
|
+
both larger than a multiple of the median partition size and larger than an absolute byte
|
|
10
|
+
threshold[^1]: in the shipped implementation, more than 5× the median size and more than
|
|
11
|
+
256 MB by default[^2]. Whatever task draws that oversized partition ends up processing far
|
|
12
|
+
more data than everyone else in the same stage.
|
|
13
|
+
|
|
14
|
+
## How it's detected
|
|
15
|
+
|
|
16
|
+
| Signal | Warning | Critical |
|
|
17
|
+
|---|---|---|
|
|
18
|
+
| P95 / median task duration | > 3× | > 5× |
|
|
19
|
+
| Max / median task duration (task count < 20) | > 3× | > 5× |
|
|
20
|
+
|
|
21
|
+
## Why it matters
|
|
22
|
+
|
|
23
|
+
An oversized partition becomes a [straggler task](#bottleneck-straggler): it's processed inside a single task, so the
|
|
24
|
+
stage can't finish until that task does, no matter how long it takes. Spark's own skew
|
|
25
|
+
handling treats splitting the partition as a deliberate trade-off: reading the other join
|
|
26
|
+
side's matching partition once per split costs extra I/O, but the design behind the feature
|
|
27
|
+
argues that cost is worth paying once the skew is severe enough to be producing a straggler
|
|
28
|
+
in the first place[^1].
|
|
29
|
+
|
|
30
|
+
## How to fix it
|
|
31
|
+
|
|
32
|
+
- Enable Adaptive Query Execution's skew-join handling (`spark.sql.adaptive.skewJoin.enabled`,
|
|
33
|
+
default `true` once `spark.sql.adaptive.enabled` is also on[^2]). It divides any partition that
|
|
34
|
+
crosses the skew thresholds into several smaller sub-partitions, joins each one against the
|
|
35
|
+
matching data on the other side, and unions the results back together[^1].
|
|
36
|
+
- Before AQE existed, the only options were manual, and each carries real limitations, which
|
|
37
|
+
is exactly the gap AQE's skew handling was built to close[^1]: salt the join key, over-size
|
|
38
|
+
`spark.sql.shuffle.partitions`, or push the [broadcast-join threshold](#bottleneck-broadcast-sizing) up so the join goes
|
|
39
|
+
broadcast instead of sort-merge.
|
|
40
|
+
- To salt by hand: add a random suffix to the join key on both sides, exploding the smaller
|
|
41
|
+
side into one row per salt value so every salted variant of the key still finds a match, then
|
|
42
|
+
join on the combined (key, salt) pair[^3].
|
|
43
|
+
- On Databricks, a `SKEW` hint can name the skewed relation and column (and, optionally, the
|
|
44
|
+
exact skewed values) directly, letting the planner build a skew-aware plan without hand-rolled
|
|
45
|
+
salting[^4].
|
|
46
|
+
|
|
47
|
+
> **PySpark:** salting is plain DataFrame code, no special API, just
|
|
48
|
+
> `withColumn("salt", (fn.rand() * n).cast("int"))` on the larger side and an `explode` over
|
|
49
|
+
> the salt range on the smaller side before joining on the combined key.
|
|
50
|
+
|
|
51
|
+
The AQE toggle, plus the manual salting fallback as runnable code:
|
|
52
|
+
|
|
53
|
+
```python
|
|
54
|
+
from pyspark.sql import functions as fn
|
|
55
|
+
|
|
56
|
+
# AQE skew-join handling, active once AQE itself is enabled
|
|
57
|
+
spark.conf.set("spark.sql.adaptive.enabled", "true")
|
|
58
|
+
spark.conf.set("spark.sql.adaptive.skewJoin.enabled", "true")
|
|
59
|
+
|
|
60
|
+
# Manual salting fallback (pre-AQE): spread the skewed key across N salted variants.
|
|
61
|
+
# N is an example, size it to how badly the key is skewed.
|
|
62
|
+
N = 16
|
|
63
|
+
salted_big = big.withColumn("salt", (fn.rand() * N).cast("int"))
|
|
64
|
+
salted_small = small.withColumn(
|
|
65
|
+
"salt", fn.explode(fn.array([fn.lit(i) for i in range(N)]))
|
|
66
|
+
)
|
|
67
|
+
joined = salted_big.join(salted_small, ["key", "salt"]).drop("salt")
|
|
68
|
+
```
|
|
69
|
+
|
|
70
|
+
## Confidence
|
|
71
|
+
|
|
72
|
+
Validated. These thresholds mirror Spark's own skew-join optimizer, which
|
|
73
|
+
defines a skewed partition by the same style of ratio-plus-absolute test applied here
|
|
74
|
+
to task durations[^1][^2].
|
|
75
|
+
|
|
76
|
+
The `Stage shape` subsection below is a separate, experimental branch: its heuristics are
|
|
77
|
+
informational stage-shape observations, not a validated finding, so it carries an
|
|
78
|
+
`EXPERIMENTAL` badge and should not be read as a diagnosis on its own.
|
|
79
|
+
|
|
80
|
+
## Limitations / false-positive risk
|
|
81
|
+
|
|
82
|
+
A task-duration ratio is a proxy for skew, not proof of it. A single long task can just as
|
|
83
|
+
easily be a garbage-collection pause or a genuinely slow host, and either one inflates the
|
|
84
|
+
max/median ratio without any partition being oversized. Small stages make the number
|
|
85
|
+
jumpy: with only a handful of tasks, one slow outlier moves the median enough to trip the
|
|
86
|
+
threshold on noise alone. Treat the signal as a prompt to look at the stage, not a verdict.
|
|
87
|
+
|
|
88
|
+
## Stage shape {#bottleneck-stage-shape}
|
|
89
|
+
|
|
90
|
+
<span class="tag">SHAPE</span> <span class="tag">EXPERIMENTAL</span>
|
|
91
|
+
|
|
92
|
+
These are experimental, information-only heuristics about a stage's overall shape, not a
|
|
93
|
+
graded finding. They read the relationship between a stage's task count and the cluster's
|
|
94
|
+
resources, and none of them alone means something is wrong.
|
|
95
|
+
|
|
96
|
+
A stage's task count equals the number of partitions in its output, and each task runs as a
|
|
97
|
+
single thread against one partition, so that count is what caps how much of the cluster the
|
|
98
|
+
stage can actually use. Three shapes are worth noticing:
|
|
99
|
+
|
|
100
|
+
- **Low parallelism** (task count far below total executor cores). When a stage has fewer
|
|
101
|
+
tasks than there are core slots to run them in, cores sit idle and the stage cannot put
|
|
102
|
+
the cluster's CPU to work; the same shape also concentrates more memory pressure into
|
|
103
|
+
each task's aggregation[^5]. Spark's guidance is to keep parallelism high enough that the
|
|
104
|
+
cluster does not stay under-used, roughly 2 to 3 tasks per CPU core in general[^6].
|
|
105
|
+
- **Data explosion** (task count far above cores, partitions tiny). Push the count too high
|
|
106
|
+
and partitions shrink until per-task overhead floods the stage, so the scheduling cost
|
|
107
|
+
starts to dominate the useful work[^7].
|
|
108
|
+
- **Task and stage skew** (uneven task count or duration). A lopsided split shows up as a
|
|
109
|
+
few tasks carrying the stage while the rest finish early, which is the same imbalance the
|
|
110
|
+
graded `skew` finding above tracks by duration ratio.
|
|
111
|
+
|
|
112
|
+
## Related
|
|
113
|
+
|
|
114
|
+
- **Why it happens:** [Partitioning](#partitioning), [Join Optimization](#joins)
|
|
115
|
+
- **How Spark fixes it automatically:** [Adaptive Query Execution](#aqe)
|
|
116
|
+
|
|
117
|
+
[^1]: [SPARK-29544: Optimize Skewed Join at Runtime](https://issues.apache.org/jira/browse/SPARK-29544)
|
|
118
|
+
[^2]: [Configuration (Spark)](https://spark.apache.org/docs/latest/configuration.html)
|
|
119
|
+
[^3]: [Spark Tips: Partition Tuning](https://luminousmen.com/post/spark-tips-partition-tuning)
|
|
120
|
+
[^4]: [Skew Join Hint](https://docs.databricks.com/aws/en/archive/legacy/skew-join)
|
|
121
|
+
[^5]: [How to Tune Your Apache Spark Jobs (Part 2)](https://blog.cloudera.com/how-to-tune-your-apache-spark-jobs-part-2/)
|
|
122
|
+
[^6]: [Tuning - Spark](https://spark.apache.org/docs/latest/tuning.html)
|
|
123
|
+
[^7]: [Spark Partitions](https://luminousmen.com/post/spark-partitions)
|
|
@@ -0,0 +1,117 @@
|
|
|
1
|
+
# Slow Host
|
|
2
|
+
|
|
3
|
+
<span class="tag">HOST</span>
|
|
4
|
+
|
|
5
|
+
## What it is
|
|
6
|
+
|
|
7
|
+
A slow host bottleneck shows up when one machine in the cluster consistently turns in slower
|
|
8
|
+
task times than its peers, independent of any single task's own data size. Unlike a stray
|
|
9
|
+
[straggler task](#bottleneck-straggler), the effect is host-wide: every task Spark schedules there runs behind, which
|
|
10
|
+
drags out the stage even when the workload itself is partitioned evenly.
|
|
11
|
+
|
|
12
|
+
## How it's detected
|
|
13
|
+
|
|
14
|
+
Spark's event log carries the per-task detail behind this: turning on `spark.eventLog.enabled`
|
|
15
|
+
logs the events that encode what the UI displays, persisted to storage[^1], and that same log
|
|
16
|
+
backs the UI's Stages tab, which drills down into individual tasks and shows per-task metrics
|
|
17
|
+
such as duration, GC time, and shuffle bytes read[^2].
|
|
18
|
+
|
|
19
|
+
A host reads as slow against the following synthetic threshold: host mean task
|
|
20
|
+
duration ≥ 2× overall median AND the host holds ≥ 20% task share; the signal only applies with
|
|
21
|
+
`taskCount ≥ 15` and `hosts.length ≥ 3`.
|
|
22
|
+
|
|
23
|
+
## Why it matters
|
|
24
|
+
|
|
25
|
+
A host running persistently slow tasks behaves like a bottleneck baked into the cluster rather
|
|
26
|
+
than into the workload: every task Spark places there inherits the delay, and since a stage's
|
|
27
|
+
completion time is bounded by its slowest tasks, the rest of the cluster idles while the
|
|
28
|
+
affected host catches up. Left alone, the same host keeps dragging down every later stage and
|
|
29
|
+
job that lands work on it.
|
|
30
|
+
|
|
31
|
+
## How to fix it
|
|
32
|
+
|
|
33
|
+
- Turn on speculative execution (`spark.speculation`, off by default) so Spark relaunches a
|
|
34
|
+
copy of a task that's lagging far behind its peers instead of waiting on the slow host to
|
|
35
|
+
finish it. A stage only becomes eligible once `spark.speculation.quantile` (default `0.9`) of
|
|
36
|
+
its tasks have completed, and a task then qualifies once it runs more than
|
|
37
|
+
`spark.speculation.multiplier` (default `3`) times the median duration, subject to a
|
|
38
|
+
`spark.speculation.minTaskRuntime` floor (default `100ms`) so short tasks aren't speculated
|
|
39
|
+
purely for looking slow, or, since Spark 3.4, an efficiency check
|
|
40
|
+
(`spark.speculation.efficiency.enabled`, default `true`) that also requires the task's
|
|
41
|
+
data-processing rate to lag the stage average[^3].
|
|
42
|
+
- If the slow host also happens to hold data locality for the affected tasks, lowering
|
|
43
|
+
`spark.locality.wait` (default `3s`), or the level-specific `spark.locality.wait.node`,
|
|
44
|
+
`.rack`, and `.process` overrides, shortens how long Spark waits for a data-local slot
|
|
45
|
+
before falling back to a less-local executor[^3].
|
|
46
|
+
|
|
47
|
+
Turn on speculation and, if the slow host holds locality, shorten the wait; defaults shown:
|
|
48
|
+
|
|
49
|
+
```properties
|
|
50
|
+
# Relaunch tasks stuck on a slow host (speculation is OFF by default)
|
|
51
|
+
spark.speculation=true
|
|
52
|
+
spark.speculation.quantile=0.9 # default; fraction of tasks done before speculation starts
|
|
53
|
+
spark.speculation.multiplier=3 # default; a task must run > 3x the median to be relaunched
|
|
54
|
+
|
|
55
|
+
# If the slow host holds data locality, wait less before falling back to another executor
|
|
56
|
+
spark.locality.wait=3s # default; lower to fall back sooner
|
|
57
|
+
```
|
|
58
|
+
|
|
59
|
+
## Confidence
|
|
60
|
+
|
|
61
|
+
The core `durationShare` dimension is validated: it measures host-wide task-duration
|
|
62
|
+
inflation against the cluster median, the signal that separates a genuinely slow machine from
|
|
63
|
+
normal task-time variance. Two secondary dimensions run best-effort and stay experimental. The
|
|
64
|
+
`multiDim` dimension <span class="tag">EXPERIMENTAL</span> folds several per-host signals into
|
|
65
|
+
one score, but that combined heuristic has not been validated. The `storageMemory` dimension
|
|
66
|
+
<span class="tag">EXPERIMENTAL</span> likewise reads host-level memory pressure as a
|
|
67
|
+
contributing factor without validation, so treat both as hints rather than verdicts.
|
|
68
|
+
|
|
69
|
+
## Limitations / false-positive risk
|
|
70
|
+
|
|
71
|
+
A host that looks slow is not always a bad node. It can simply hold data locality for the
|
|
72
|
+
tasks scheduled there, or carry one heavy stage that happens to be pinned to it, either of
|
|
73
|
+
which inflates its mean task time without any hardware fault. The multi-dimensional scoring
|
|
74
|
+
that reinforces the core signal is best-effort, so a match warrants a look at what that host was
|
|
75
|
+
actually running before you conclude the machine itself is the problem.
|
|
76
|
+
|
|
77
|
+
|
|
78
|
+
## Stage slowness {#bottleneck-stage-slowness}
|
|
79
|
+
|
|
80
|
+
<span class="tag">SLOW</span> <span class="tag">EXPERIMENTAL</span>
|
|
81
|
+
|
|
82
|
+
This is an experimental fallback heuristic that shares the slow-host anchor and is suppressed
|
|
83
|
+
whenever `slowHost` fires. It answers a different question: when no single machine is dragging,
|
|
84
|
+
why does one stage still lag the rest of the job? The cause usually traces back to how the work
|
|
85
|
+
was divided rather than where it ran.
|
|
86
|
+
|
|
87
|
+
**Too few partitions caps parallelism.** Wide stages (joins, groupBy, aggregations) take their
|
|
88
|
+
partition count from `spark.sql.shuffle.partitions`, which sits at a default of 200 whether the
|
|
89
|
+
shuffle moves 20 MB or 500 GB unless you change it[^4]. When that count is small relative to the
|
|
90
|
+
cluster, only a handful of tasks carry the stage while cores sit idle, so it stretches out even
|
|
91
|
+
as better-sized neighbors finish quickly. The first move is to raise its parallelism, aiming for
|
|
92
|
+
at least two or three tasks per CPU core on a data-heavy stage, tuned through
|
|
93
|
+
`spark.default.parallelism` and `spark.sql.shuffle.partitions`[^5].
|
|
94
|
+
|
|
95
|
+
**Large data volume per task makes each task run long.** That 200-partition default does not
|
|
96
|
+
scale with input size, so a stage handling hundreds of gigabytes hands each task a large slice:
|
|
97
|
+
fewer tasks run at once, per-executor load climbs, and the stage often trips memory errors.
|
|
98
|
+
Roughly 100-200 MB per task tends to work well, and tasks grinding through around 3 GB each
|
|
99
|
+
while spilling are a sign you need more partitions[^4].
|
|
100
|
+
|
|
101
|
+
**Heavy shuffle and spill inflate the stage's cost.** Shuffles move data across executors so
|
|
102
|
+
rows with the same key land together, and during those map and [shuffle](#shuffle) operations Spark writes
|
|
103
|
+
to and reads from local-disk shuffle files, which is heavy I/O that can become a bottleneck
|
|
104
|
+
under the default configuration[^2]. Volume per task feeds straight into it: once a partition
|
|
105
|
+
outgrows the memory available in an executor, Spark [spills](#bottleneck-spill) part of the data to disk, and spills
|
|
106
|
+
are about the slowest thing a job can do because of the extra disk I/O and garbage collection,
|
|
107
|
+
so the stage finishes but runs far less efficiently than the rest[^4].
|
|
108
|
+
|
|
109
|
+
## Related
|
|
110
|
+
|
|
111
|
+
- **Speculation & locality tuning:** [Cluster Tuning](#cluster-config)
|
|
112
|
+
|
|
113
|
+
[^1]: [Monitoring and Instrumentation](https://spark.apache.org/docs/latest/monitoring.html#spark-history-server)
|
|
114
|
+
[^2]: *Learning Spark, 2nd Edition*, Damji, Wenig, Das, Lee, ch. 7
|
|
115
|
+
[^3]: [Configuration (Spark)](https://spark.apache.org/docs/latest/configuration.html)
|
|
116
|
+
[^4]: [Spark Partitions](https://luminousmen.com/post/spark-partitions)
|
|
117
|
+
[^5]: *Spark: The Definitive Guide*, Chambers & Zaharia (O'Reilly, 2018)
|