sparkforensics-cli 0.1.0 → 0.2.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +6 -0
- package/bin/sparkforensics-analyze.mjs +113 -48
- package/export-template/docs/404.html +25 -0
- package/export-template/docs/assets/app.CndaAS6v.js +1 -0
- package/export-template/docs/assets/aqe-loop.IwQSATHw.svg +1 -0
- package/export-template/docs/assets/aqe-loop.dark.DGbaxqJE.svg +1 -0
- package/export-template/docs/assets/broadcast-vs-shuffle.Db4WY1XK.svg +1 -0
- package/export-template/docs/assets/broadcast-vs-shuffle.dark.C7Bxs0mG.svg +1 -0
- package/export-template/docs/assets/cache-lifecycle.dark.B-hS7AgU.svg +1 -0
- package/export-template/docs/assets/cache-lifecycle.rEOVYQNU.svg +1 -0
- package/export-template/docs/assets/chunks/@localSearchIndexroot.DNY8bVcl.js +1 -0
- package/export-template/docs/assets/chunks/VPLocalSearchBox.yJbZbsEo.js +9 -0
- package/export-template/docs/assets/chunks/duplicate-plan-subtree.dark.Cdp70QhV.js +1 -0
- package/export-template/docs/assets/chunks/framework.DSg0KOwT.js +20 -0
- package/export-template/docs/assets/chunks/retry-escalation-ladder.dark.DHipdJgZ.js +1 -0
- package/export-template/docs/assets/chunks/theme.Df2VAG9w.js +2 -0
- package/export-template/docs/assets/cold-start-timeline.DxC_Sc7w.svg +1 -0
- package/export-template/docs/assets/cold-start-timeline.dark.CZ17YcAG.svg +1 -0
- package/export-template/docs/assets/columnar-layout.PghGeOEA.svg +1 -0
- package/export-template/docs/assets/columnar-layout.dark.BVNlz0ff.svg +1 -0
- package/export-template/docs/assets/container-memory.DIO0AnIm.svg +1 -0
- package/export-template/docs/assets/container-memory.dark.CP-5zuCl.svg +1 -0
- package/export-template/docs/assets/contributor-guide_architecture_board-widgets.md.B-OsL91z.js +1 -0
- package/export-template/docs/assets/contributor-guide_architecture_board-widgets.md.B-OsL91z.lean.js +1 -0
- package/export-template/docs/assets/contributor-guide_architecture_detector-contract.md.BOeH4d1J.js +1 -0
- package/export-template/docs/assets/contributor-guide_architecture_detector-contract.md.BOeH4d1J.lean.js +1 -0
- package/export-template/docs/assets/contributor-guide_architecture_drill-down.md.BtPdlM7r.js +1 -0
- package/export-template/docs/assets/contributor-guide_architecture_drill-down.md.BtPdlM7r.lean.js +1 -0
- package/export-template/docs/assets/contributor-guide_architecture_impact-estimation.md.DYCDPgkh.js +1 -0
- package/export-template/docs/assets/contributor-guide_architecture_impact-estimation.md.DYCDPgkh.lean.js +1 -0
- package/export-template/docs/assets/contributor-guide_architecture_index.md.3TO9ic6w.js +1 -0
- package/export-template/docs/assets/contributor-guide_architecture_index.md.3TO9ic6w.lean.js +1 -0
- package/export-template/docs/assets/contributor-guide_architecture_overview.md.CehiRmGn.js +1 -0
- package/export-template/docs/assets/contributor-guide_architecture_overview.md.CehiRmGn.lean.js +1 -0
- package/export-template/docs/assets/contributor-guide_architecture_state-and-history.md.m3S3UdMk.js +1 -0
- package/export-template/docs/assets/contributor-guide_architecture_state-and-history.md.m3S3UdMk.lean.js +1 -0
- package/export-template/docs/assets/contributor-guide_architecture_widget-rendering.md.DbqPf2OT.js +1 -0
- package/export-template/docs/assets/contributor-guide_architecture_widget-rendering.md.DbqPf2OT.lean.js +1 -0
- package/export-template/docs/assets/contributor-guide_architecture_worker-protocol.md.B93qJ_tT.js +6 -0
- package/export-template/docs/assets/contributor-guide_architecture_worker-protocol.md.B93qJ_tT.lean.js +1 -0
- package/export-template/docs/assets/contributor-guide_contributing.md.CvRsdr6J.js +1 -0
- package/export-template/docs/assets/contributor-guide_contributing.md.CvRsdr6J.lean.js +1 -0
- package/export-template/docs/assets/contributor-guide_development-setup.md.DvAN_9mK.js +12 -0
- package/export-template/docs/assets/contributor-guide_development-setup.md.DvAN_9mK.lean.js +1 -0
- package/export-template/docs/assets/contributor-guide_testing.md.6rIKqSyY.js +1 -0
- package/export-template/docs/assets/contributor-guide_testing.md.6rIKqSyY.lean.js +1 -0
- package/export-template/docs/assets/dag-stages.DSz_S937.svg +1 -0
- package/export-template/docs/assets/dag-stages.dark.F72UzxH4.svg +1 -0
- package/export-template/docs/assets/driver-executor.D5pQ7YN1.svg +1 -0
- package/export-template/docs/assets/driver-executor.dark.BmX9cPvh.svg +1 -0
- package/export-template/docs/assets/duplicate-plan-subtree.B4cvN6fj.svg +1 -0
- package/export-template/docs/assets/duplicate-plan-subtree.dark.Dw8wS0Ag.svg +1 -0
- package/export-template/docs/assets/index.md.CHJVslga.js +1 -0
- package/export-template/docs/assets/index.md.CHJVslga.lean.js +1 -0
- package/export-template/docs/assets/inter-italic-cyrillic-ext.r48I6akx.woff2 +0 -0
- package/export-template/docs/assets/inter-italic-cyrillic.By2_1cv3.woff2 +0 -0
- package/export-template/docs/assets/inter-italic-greek-ext.1u6EdAuj.woff2 +0 -0
- package/export-template/docs/assets/inter-italic-greek.DJ8dCoTZ.woff2 +0 -0
- package/export-template/docs/assets/inter-italic-latin-ext.CN1xVJS-.woff2 +0 -0
- package/export-template/docs/assets/inter-italic-latin.C2AdPX0b.woff2 +0 -0
- package/export-template/docs/assets/inter-italic-vietnamese.BSbpV94h.woff2 +0 -0
- package/export-template/docs/assets/inter-roman-cyrillic-ext.BBPuwvHQ.woff2 +0 -0
- package/export-template/docs/assets/inter-roman-cyrillic.C5lxZ8CY.woff2 +0 -0
- package/export-template/docs/assets/inter-roman-greek-ext.CqjqNYQ-.woff2 +0 -0
- package/export-template/docs/assets/inter-roman-greek.BBVDIX6e.woff2 +0 -0
- package/export-template/docs/assets/inter-roman-latin-ext.4ZJIpNVo.woff2 +0 -0
- package/export-template/docs/assets/inter-roman-latin.Di8DUHzh.woff2 +0 -0
- package/export-template/docs/assets/inter-roman-vietnamese.BjW4sHH5.woff2 +0 -0
- package/export-template/docs/assets/join-strategy.C_FvrCEo.svg +1 -0
- package/export-template/docs/assets/join-strategy.dark.ChMLnNII.svg +1 -0
- package/export-template/docs/assets/memory-borrowing.BqQRJg0u.svg +1 -0
- package/export-template/docs/assets/memory-borrowing.dark.Yhh20O9C.svg +1 -0
- package/export-template/docs/assets/memory-regions.XHvO7jHG.svg +1 -0
- package/export-template/docs/assets/memory-regions.dark.D4TP9_08.svg +1 -0
- package/export-template/docs/assets/repartition-vs-coalesce.BovLRrpj.svg +1 -0
- package/export-template/docs/assets/repartition-vs-coalesce.dark.BhAczKZQ.svg +1 -0
- package/export-template/docs/assets/retry-escalation-ladder.DyTKJJmZ.svg +1 -0
- package/export-template/docs/assets/retry-escalation-ladder.dark.BdsabtU3.svg +1 -0
- package/export-template/docs/assets/shuffle-map-reduce.KuOEZVmg.svg +1 -0
- package/export-template/docs/assets/shuffle-map-reduce.dark.BgQZnFSb.svg +1 -0
- package/export-template/docs/assets/spill-classification.BU2euYDO.svg +1 -0
- package/export-template/docs/assets/spill-classification.dark.D7i1M40d.svg +1 -0
- package/export-template/docs/assets/style.DXOMCXxn.css +1 -0
- package/export-template/docs/assets/tuning-reference_anti-patterns.md.Df1YMIHu.js +1 -0
- package/export-template/docs/assets/tuning-reference_anti-patterns.md.Df1YMIHu.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_aqe.md.BIsCtLzm.js +1 -0
- package/export-template/docs/assets/tuning-reference_aqe.md.BIsCtLzm.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-broadcast-sizing.md.CEstB3Ia.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-broadcast-sizing.md.CEstB3Ia.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-cold-start.md.CEuy-72y.js +7 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-cold-start.md.CEuy-72y.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-duplicate-plan-subtree.md.CIohQDfn.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-duplicate-plan-subtree.md.CIohQDfn.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-failures.md.4z5BXGJ2.js +6 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-failures.md.4z5BXGJ2.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-gc.md.DSxzZRK7.js +6 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-gc.md.DSxzZRK7.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-job-failure-rate.md.BaJl__1W.js +8 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-job-failure-rate.md.BaJl__1W.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-memory-utilization.md.DbP-SJZc.js +7 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-memory-utilization.md.DbP-SJZc.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-retry-waste.md.D5JMjVOt.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-retry-waste.md.D5JMjVOt.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-shuffle.md.CM-nTmIH.js +12 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-shuffle.md.CM-nTmIH.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-skew.md.BdUwiDhn.js +14 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-skew.md.BdUwiDhn.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-slow-host.md.BlIo6UDW.js +7 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-slow-host.md.BlIo6UDW.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-small-files.md.B8kloyx8.js +5 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-small-files.md.B8kloyx8.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-spill.md.PNH7mITt.js +6 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-spill.md.PNH7mITt.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-straggler.md.DY36fHN5.js +7 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-straggler.md.DY36fHN5.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-tiny-tasks.md.QTV7O8kU.js +8 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-tiny-tasks.md.QTV7O8kU.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-utilization.md.DTiueZC3.js +5 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-utilization.md.DTiueZC3.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_caching.md.B7aQ8asB.js +1 -0
- package/export-template/docs/assets/tuning-reference_caching.md.B7aQ8asB.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_cluster-config.md.ZVmDGsQ3.js +1 -0
- package/export-template/docs/assets/tuning-reference_cluster-config.md.ZVmDGsQ3.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_config.md.UvveiWG3.js +1 -0
- package/export-template/docs/assets/tuning-reference_config.md.UvveiWG3.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_data-formats.md.bjCAWH3N.js +1 -0
- package/export-template/docs/assets/tuning-reference_data-formats.md.bjCAWH3N.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_index.md.BQ_NooMV.js +1 -0
- package/export-template/docs/assets/tuning-reference_index.md.BQ_NooMV.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_intro.md.CobD-lGB.js +1 -0
- package/export-template/docs/assets/tuning-reference_intro.md.CobD-lGB.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_joins.md.BtKs_CuW.js +1 -0
- package/export-template/docs/assets/tuning-reference_joins.md.BtKs_CuW.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_memory-model.md.DhT-n4y3.js +1 -0
- package/export-template/docs/assets/tuning-reference_memory-model.md.DhT-n4y3.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_metrics.md.mLOh7Apj.js +1 -0
- package/export-template/docs/assets/tuning-reference_metrics.md.mLOh7Apj.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_partitioning.md.q0zKF_8X.js +1 -0
- package/export-template/docs/assets/tuning-reference_partitioning.md.q0zKF_8X.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_pyspark.md.DDCfvN9t.js +6 -0
- package/export-template/docs/assets/tuning-reference_pyspark.md.DDCfvN9t.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_shuffle.md.BZZ7R4Ix.js +1 -0
- package/export-template/docs/assets/tuning-reference_shuffle.md.BZZ7R4Ix.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_spark-architecture.md.Dwzm5avO.js +1 -0
- package/export-template/docs/assets/tuning-reference_spark-architecture.md.Dwzm5avO.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_table-formats.md.D6wj-2dX.js +1 -0
- package/export-template/docs/assets/tuning-reference_table-formats.md.D6wj-2dX.lean.js +1 -0
- package/export-template/docs/assets/udf-execution-models.BUFDICuG.svg +1 -0
- package/export-template/docs/assets/udf-execution-models.dark.YTNS6GDq.svg +1 -0
- package/export-template/docs/assets/user-guide_alternative-log-retrieval.md.sU3KGarf.js +1 -0
- package/export-template/docs/assets/user-guide_alternative-log-retrieval.md.sU3KGarf.lean.js +1 -0
- package/export-template/docs/assets/user-guide_getting-started.md.DtEM37MK.js +3 -0
- package/export-template/docs/assets/user-guide_getting-started.md.DtEM37MK.lean.js +1 -0
- package/export-template/docs/assets/user-guide_mcp-tools.md.C8MiIu7F.js +125 -0
- package/export-template/docs/assets/user-guide_mcp-tools.md.C8MiIu7F.lean.js +1 -0
- package/export-template/docs/assets/user-guide_run-comparison.md.S0TWWmLY.js +1 -0
- package/export-template/docs/assets/user-guide_run-comparison.md.S0TWWmLY.lean.js +1 -0
- package/export-template/docs/assets/user-guide_understanding-findings.md.D0R_Y-R2.js +1 -0
- package/export-template/docs/assets/user-guide_understanding-findings.md.D0R_Y-R2.lean.js +1 -0
- package/export-template/docs/contributor-guide/architecture/board-widgets.html +25 -0
- package/export-template/docs/contributor-guide/architecture/detector-contract.html +25 -0
- package/export-template/docs/contributor-guide/architecture/drill-down.html +25 -0
- package/export-template/docs/contributor-guide/architecture/impact-estimation.html +25 -0
- package/export-template/docs/contributor-guide/architecture/index.html +25 -0
- package/export-template/docs/contributor-guide/architecture/overview.html +25 -0
- package/export-template/docs/contributor-guide/architecture/state-and-history.html +25 -0
- package/export-template/docs/contributor-guide/architecture/widget-rendering.html +25 -0
- package/export-template/docs/contributor-guide/architecture/worker-protocol.html +30 -0
- package/export-template/docs/contributor-guide/contributing.html +25 -0
- package/export-template/docs/contributor-guide/development-setup.html +36 -0
- package/export-template/docs/contributor-guide/testing.html +25 -0
- package/export-template/docs/favicon.svg +4 -0
- package/export-template/docs/hashmap.json +1 -0
- package/export-template/docs/index.html +25 -0
- package/export-template/docs/package.json +1 -0
- package/export-template/docs/tuning-reference/anti-patterns.html +25 -0
- package/export-template/docs/tuning-reference/aqe.html +25 -0
- package/export-template/docs/tuning-reference/bottleneck-broadcast-sizing.html +25 -0
- package/export-template/docs/tuning-reference/bottleneck-cold-start.html +31 -0
- package/export-template/docs/tuning-reference/bottleneck-duplicate-plan-subtree.html +25 -0
- package/export-template/docs/tuning-reference/bottleneck-failures.html +30 -0
- package/export-template/docs/tuning-reference/bottleneck-gc.html +30 -0
- package/export-template/docs/tuning-reference/bottleneck-job-failure-rate.html +32 -0
- package/export-template/docs/tuning-reference/bottleneck-memory-utilization.html +31 -0
- package/export-template/docs/tuning-reference/bottleneck-retry-waste.html +25 -0
- package/export-template/docs/tuning-reference/bottleneck-shuffle.html +36 -0
- package/export-template/docs/tuning-reference/bottleneck-skew.html +38 -0
- package/export-template/docs/tuning-reference/bottleneck-slow-host.html +31 -0
- package/export-template/docs/tuning-reference/bottleneck-small-files.html +29 -0
- package/export-template/docs/tuning-reference/bottleneck-spill.html +30 -0
- package/export-template/docs/tuning-reference/bottleneck-straggler.html +31 -0
- package/export-template/docs/tuning-reference/bottleneck-tiny-tasks.html +32 -0
- package/export-template/docs/tuning-reference/bottleneck-utilization.html +29 -0
- package/export-template/docs/tuning-reference/caching.html +25 -0
- package/export-template/docs/tuning-reference/cluster-config.html +25 -0
- package/export-template/docs/tuning-reference/config.html +25 -0
- package/export-template/docs/tuning-reference/data-formats.html +25 -0
- package/export-template/docs/tuning-reference/index.html +25 -0
- package/export-template/docs/tuning-reference/intro.html +25 -0
- package/export-template/docs/tuning-reference/joins.html +25 -0
- package/export-template/docs/tuning-reference/memory-model.html +25 -0
- package/export-template/docs/tuning-reference/metrics.html +25 -0
- package/export-template/docs/tuning-reference/partitioning.html +25 -0
- package/export-template/docs/tuning-reference/pyspark.html +30 -0
- package/export-template/docs/tuning-reference/shuffle.html +25 -0
- package/export-template/docs/tuning-reference/spark-architecture.html +25 -0
- package/export-template/docs/tuning-reference/table-formats.html +25 -0
- package/export-template/docs/user-guide/alternative-log-retrieval.html +25 -0
- package/export-template/docs/user-guide/getting-started.html +27 -0
- package/export-template/docs/user-guide/mcp-tools.html +149 -0
- package/export-template/docs/user-guide/run-comparison.html +25 -0
- package/export-template/docs/user-guide/understanding-findings.html +25 -0
- package/export-template/docs/vp-icons.css +0 -0
- package/export-template/favicon.svg +4 -0
- package/export-template/index.html +115 -0
- package/export-template/parser-worker-QqyEE4m9.js +64 -0
- package/package.json +16 -3
- package/vendor-core/analyzer.js +74 -74
- package/vendor-core/cli/budgets.js +13 -27
- package/vendor-core/cli/collect-run.js +43 -19
- package/vendor-core/core-count.js +25 -27
- package/vendor-core/core-locality-ratio.js +4 -11
- package/vendor-core/core-time-series.js +6 -12
- package/vendor-core/core-usage-locality.js +3 -4
- package/vendor-core/detectors.js +256 -375
- package/vendor-core/docs-config.js +69 -21
- package/vendor-core/docs-content/chapters/01-intro.md +32 -0
- package/vendor-core/docs-content/chapters/02-spark-architecture.md +76 -0
- package/vendor-core/docs-content/chapters/03-memory-model.md +73 -0
- package/vendor-core/docs-content/chapters/04-partitioning.md +65 -0
- package/vendor-core/docs-content/chapters/05-joins.md +62 -0
- package/vendor-core/docs-content/chapters/06-shuffle.md +59 -0
- package/vendor-core/docs-content/chapters/07-data-formats.md +81 -0
- package/vendor-core/docs-content/chapters/07b-table-formats.md +56 -0
- package/vendor-core/docs-content/chapters/08-caching.md +58 -0
- package/vendor-core/docs-content/chapters/09-pyspark.md +78 -0
- package/vendor-core/docs-content/chapters/10-aqe.md +167 -0
- package/vendor-core/docs-content/chapters/11-cluster-config.md +170 -0
- package/vendor-core/docs-content/chapters/12-anti-patterns.md +171 -0
- package/vendor-core/docs-content/chapters/14-metrics.md +87 -0
- package/vendor-core/docs-content/chapters/15-config.md +93 -0
- package/vendor-core/docs-content/chapters/nav-index.json +370 -0
- package/vendor-core/docs-content/detection/cache.md +6 -0
- package/vendor-core/docs-content/detection/cfg.md +15 -0
- package/vendor-core/docs-content/detection/chrn.md +7 -0
- package/vendor-core/docs-content/detection/cold.md +4 -0
- package/vendor-core/docs-content/detection/cstor.md +4 -0
- package/vendor-core/docs-content/detection/fail.md +5 -0
- package/vendor-core/docs-content/detection/gc.md +4 -0
- package/vendor-core/docs-content/detection/host.md +5 -0
- package/vendor-core/docs-content/detection/incmp.md +6 -0
- package/vendor-core/docs-content/detection/jobs.md +4 -0
- package/vendor-core/docs-content/detection/local.md +7 -0
- package/vendor-core/docs-content/detection/mem.md +10 -0
- package/vendor-core/docs-content/detection/part.md +5 -0
- package/vendor-core/docs-content/detection/plan.md +14 -0
- package/vendor-core/docs-content/detection/retry.md +4 -0
- package/vendor-core/docs-content/detection/sfail.md +5 -0
- package/vendor-core/docs-content/detection/shape.md +5 -0
- package/vendor-core/docs-content/detection/shfl.md +4 -0
- package/vendor-core/docs-content/detection/skew.md +6 -0
- package/vendor-core/docs-content/detection/slow.md +6 -0
- package/vendor-core/docs-content/detection/spec.md +7 -0
- package/vendor-core/docs-content/detection/spill.md +7 -0
- package/vendor-core/docs-content/detection/strag.md +5 -0
- package/vendor-core/docs-content/detection/tiny.md +4 -0
- package/vendor-core/docs-content/detection/util.md +4 -0
- package/vendor-core/docs-content/diagrams/aqe-loop.dark.svg +1 -0
- package/vendor-core/docs-content/diagrams/aqe-loop.svg +1 -0
- package/vendor-core/docs-content/diagrams/broadcast-vs-shuffle.dark.svg +1 -0
- package/vendor-core/docs-content/diagrams/broadcast-vs-shuffle.svg +1 -0
- package/vendor-core/docs-content/diagrams/cache-lifecycle.dark.svg +1 -0
- package/vendor-core/docs-content/diagrams/cache-lifecycle.svg +1 -0
- package/vendor-core/docs-content/diagrams/cold-start-timeline.dark.svg +1 -0
- package/vendor-core/docs-content/diagrams/cold-start-timeline.svg +1 -0
- package/vendor-core/docs-content/diagrams/columnar-layout.dark.svg +1 -0
- package/vendor-core/docs-content/diagrams/columnar-layout.svg +1 -0
- package/vendor-core/docs-content/diagrams/container-memory.dark.svg +1 -0
- package/vendor-core/docs-content/diagrams/container-memory.svg +1 -0
- package/vendor-core/docs-content/diagrams/dag-stages.dark.svg +1 -0
- package/vendor-core/docs-content/diagrams/dag-stages.svg +1 -0
- package/vendor-core/docs-content/diagrams/driver-executor.dark.svg +1 -0
- package/vendor-core/docs-content/diagrams/driver-executor.svg +1 -0
- package/vendor-core/docs-content/diagrams/duplicate-plan-subtree.dark.svg +1 -0
- package/vendor-core/docs-content/diagrams/duplicate-plan-subtree.svg +1 -0
- package/vendor-core/docs-content/diagrams/join-strategy.dark.svg +1 -0
- package/vendor-core/docs-content/diagrams/join-strategy.svg +1 -0
- package/vendor-core/docs-content/diagrams/memory-borrowing.dark.svg +1 -0
- package/vendor-core/docs-content/diagrams/memory-borrowing.svg +1 -0
- package/vendor-core/docs-content/diagrams/memory-regions.dark.svg +1 -0
- package/vendor-core/docs-content/diagrams/memory-regions.svg +1 -0
- package/vendor-core/docs-content/diagrams/repartition-vs-coalesce.dark.svg +1 -0
- package/vendor-core/docs-content/diagrams/repartition-vs-coalesce.svg +1 -0
- package/vendor-core/docs-content/diagrams/retry-escalation-ladder.dark.svg +1 -0
- package/vendor-core/docs-content/diagrams/retry-escalation-ladder.svg +1 -0
- package/vendor-core/docs-content/diagrams/shuffle-map-reduce.dark.svg +1 -0
- package/vendor-core/docs-content/diagrams/shuffle-map-reduce.svg +1 -0
- package/vendor-core/docs-content/diagrams/spill-classification.dark.svg +1 -0
- package/vendor-core/docs-content/diagrams/spill-classification.svg +1 -0
- package/vendor-core/docs-content/diagrams/udf-execution-models.dark.svg +1 -0
- package/vendor-core/docs-content/diagrams/udf-execution-models.svg +1 -0
- package/vendor-core/docs-content/tuning/broadcast-sizing.md +78 -0
- package/vendor-core/docs-content/tuning/cold-start.md +81 -0
- package/vendor-core/docs-content/tuning/duplicate-plan-subtree.md +45 -0
- package/vendor-core/docs-content/tuning/failures.md +124 -0
- package/vendor-core/docs-content/tuning/gc.md +110 -0
- package/vendor-core/docs-content/tuning/job-failure-rate.md +101 -0
- package/vendor-core/docs-content/tuning/memory-utilization.md +58 -0
- package/vendor-core/docs-content/tuning/retry-waste.md +90 -0
- package/vendor-core/docs-content/tuning/shuffle.md +154 -0
- package/vendor-core/docs-content/tuning/skew.md +123 -0
- package/vendor-core/docs-content/tuning/slow-host.md +117 -0
- package/vendor-core/docs-content/tuning/small-files.md +99 -0
- package/vendor-core/docs-content/tuning/spill.md +114 -0
- package/vendor-core/docs-content/tuning/straggler.md +103 -0
- package/vendor-core/docs-content/tuning/tiny-tasks.md +94 -0
- package/vendor-core/docs-content/tuning/utilization.md +90 -0
- package/vendor-core/docs-site-config.js +10 -17
- package/vendor-core/efficiency-model.js +7 -13
- package/vendor-core/etl-phases.js +3 -5
- package/vendor-core/event-handlers.js +232 -134
- package/vendor-core/event-schemas.js +48 -114
- package/vendor-core/evidence-availability.js +5 -10
- package/vendor-core/evidence-report.js +72 -122
- package/vendor-core/export-data.js +48 -0
- package/vendor-core/finding-action-label.js +4 -10
- package/vendor-core/finding-filter-predicate.js +3 -7
- package/vendor-core/finding-generic-recommendation.js +112 -0
- package/vendor-core/finding-names.js +51 -0
- package/vendor-core/format-utils.js +112 -38
- package/vendor-core/impact-band.js +18 -24
- package/vendor-core/impact-estimator.js +38 -74
- package/vendor-core/ingest.js +7 -13
- package/vendor-core/job-groups.js +3 -6
- package/vendor-core/list-runs.js +278 -0
- package/vendor-core/load-vendored.js +6 -12
- package/vendor-core/log-header-peek.js +81 -0
- package/vendor-core/lz4-block.js +4 -6
- package/vendor-core/mcp-server-factory.js +38 -8
- package/vendor-core/mcp-tools.js +105 -76
- package/vendor-core/model-assembler.js +8 -16
- package/vendor-core/occupancy.js +5 -9
- package/vendor-core/parser-worker.js +18 -27
- package/vendor-core/plan-dot.js +2 -5
- package/vendor-core/plan-duration-attribution.js +78 -29
- package/vendor-core/plan-graph-model.js +126 -69
- package/vendor-core/plan-node-detail.js +31 -17
- package/vendor-core/plan-summary.js +19 -8
- package/vendor-core/recommendation-rollup.js +35 -39
- package/vendor-core/redact.js +72 -16
- package/vendor-core/rolling-log-reassembly.js +4 -6
- package/vendor-core/run-comparison.js +65 -70
- package/vendor-core/scaling-sim.js +5 -7
- package/vendor-core/session-snapshot.js +1 -1
- package/vendor-core/shs-fetch.js +4 -6
- package/vendor-core/shs-load.js +9 -13
- package/vendor-core/shs-request.js +1 -1
- package/vendor-core/stage-quantiles.js +14 -0
- package/vendor-core/types.js +78 -18
- package/vendor-core/wasted-core-hours.js +7 -12
|
@@ -1,18 +1,25 @@
|
|
|
1
|
-
|
|
2
|
-
|
|
3
|
-
//
|
|
4
|
-
//
|
|
5
|
-
export const
|
|
1
|
+
import { detectorCatalog } from './detectors.js';
|
|
2
|
+
|
|
3
|
+
// Single source of the docs-panel URL surface. DOCS_BASE_DIR is the built docs-site path that
|
|
4
|
+
// serves the tuning reference, relative to the app's origin.
|
|
5
|
+
export const DOCS_BASE_DIR = 'docs/tuning-reference';
|
|
6
|
+
|
|
7
|
+
// Some anchors are in-page fragments on another entry's page: config-audit sub-findings and
|
|
8
|
+
// metric-glossary entries live on the 'config'/'metrics' pages, and the "stage-*" sub-anchors
|
|
9
|
+
// live on their owning bottleneck's page. Keep in sync with spark-tuning-reference's anchor-map.
|
|
10
|
+
export function pageForAnchor(anchor ) {
|
|
11
|
+
if (anchor.startsWith('metric-')) return 'metrics';
|
|
12
|
+
if (anchor.startsWith('config-')) return 'config';
|
|
13
|
+
if (anchor === 'bottleneck-stage-shape') return 'bottleneck-skew';
|
|
14
|
+
if (anchor === 'bottleneck-stage-slowness') return 'bottleneck-slow-host';
|
|
15
|
+
return anchor;
|
|
16
|
+
}
|
|
6
17
|
|
|
7
18
|
// Build a docs URL for an anchor like '#bottleneck-skew'. The leading '#' is
|
|
8
|
-
// stripped and re-added so the fragment is never percent-encoded away.
|
|
9
|
-
|
|
10
|
-
// since the hash is reserved for #anchor navigation. Omitted for callers with no
|
|
11
|
-
// postMessage channel to protect (e.g. the plain <a href> fallback).
|
|
12
|
-
export function docsUrl(anchor , tok ) {
|
|
19
|
+
// stripped and re-added so the fragment is never percent-encoded away.
|
|
20
|
+
export function docsUrl(anchor ) {
|
|
13
21
|
const frag = String(anchor).replace(/^#/, '');
|
|
14
|
-
|
|
15
|
-
return `${DOCS_BASE_URL}${query}#${encodeURIComponent(frag)}`;
|
|
22
|
+
return `${DOCS_BASE_DIR}/${pageForAnchor(frag)}.html#${encodeURIComponent(frag)}`;
|
|
16
23
|
}
|
|
17
24
|
|
|
18
25
|
// Internal metric key → docs '#metric-*' anchor. A metric with no entry renders
|
|
@@ -38,14 +45,10 @@ export const METRIC_ANCHORS = {
|
|
|
38
45
|
'stage-duration': '#metric-stage-duration',
|
|
39
46
|
};
|
|
40
47
|
|
|
41
|
-
// Allowlist of every anchor that
|
|
42
|
-
//
|
|
43
|
-
//
|
|
44
|
-
//
|
|
45
|
-
// that from becoming a dead "Learn more" link that scrolls nowhere. When a new
|
|
46
|
-
// section lands (via `npm run update-docs`), add its anchor here and the link
|
|
47
|
-
// lights up automatically. `tests/docs-config.test.js` asserts this set stays a
|
|
48
|
-
// subset of the real ids in vendor/spark-doc/index.html so it cannot drift.
|
|
48
|
+
// Allowlist of every anchor that exists in the committed tuning-reference markdown. DocsLink
|
|
49
|
+
// renders a link only for anchors here: a detector may declare a docAnchor for an unwritten
|
|
50
|
+
// section, and gating keeps that from becoming a dead "Learn more" link. docs-config.test.js
|
|
51
|
+
// asserts this stays a subset of the real ids so it can't drift.
|
|
49
52
|
export const KNOWN_DOC_ANCHORS = new Set([
|
|
50
53
|
// Page sections
|
|
51
54
|
'#intro', '#spark-architecture', '#memory-model', '#partitioning', '#joins',
|
|
@@ -66,7 +69,52 @@ export const KNOWN_DOC_ANCHORS = new Set([
|
|
|
66
69
|
...Object.values(METRIC_ANCHORS),
|
|
67
70
|
]);
|
|
68
71
|
|
|
69
|
-
// True when `anchor` resolves to a real section in the
|
|
72
|
+
// True when `anchor` resolves to a real section in the committed
|
|
73
|
+
// tuning-reference markdown.
|
|
70
74
|
export function isKnownDocAnchor(anchor ) {
|
|
71
75
|
return KNOWN_DOC_ANCHORS.has(String(anchor));
|
|
72
76
|
}
|
|
77
|
+
|
|
78
|
+
// Finding types that never appear as their own DETECTORS entry's type: the parent entry declares
|
|
79
|
+
// a different type because one plan-walk covers two rules (see broadcastSizing). Map to the parent.
|
|
80
|
+
const TYPE_ALIASES = {
|
|
81
|
+
underBroadcast: 'broadcastSizing',
|
|
82
|
+
overBroadcast: 'broadcastSizing',
|
|
83
|
+
};
|
|
84
|
+
|
|
85
|
+
// DETECTORS is static, so this grouping is built once (lazily) instead of re-scanning per
|
|
86
|
+
// docAnchorForType call (called once per TagBadge per render).
|
|
87
|
+
let anchorsByTypeCache ;
|
|
88
|
+
|
|
89
|
+
function anchorsByType() {
|
|
90
|
+
if (!anchorsByTypeCache) {
|
|
91
|
+
anchorsByTypeCache = new Map();
|
|
92
|
+
for (const entry of detectorCatalog()) {
|
|
93
|
+
const anchors = anchorsByTypeCache.get(entry.type) ?? new Set();
|
|
94
|
+
anchors.add(entry.docAnchor);
|
|
95
|
+
anchorsByTypeCache.set(entry.type, anchors);
|
|
96
|
+
}
|
|
97
|
+
}
|
|
98
|
+
return anchorsByTypeCache;
|
|
99
|
+
}
|
|
100
|
+
|
|
101
|
+
/** Resolves a finding `type` to its documented anchor from detectorCatalog(). Returns undefined
|
|
102
|
+
* when entries sharing the type disagree on docAnchor (only configAudit today), or when the
|
|
103
|
+
* resolved anchor isn't in the allowlist (isKnownDocAnchor, the same gate DocsLink uses). */
|
|
104
|
+
export function docAnchorForType(type ) {
|
|
105
|
+
const resolvedType = TYPE_ALIASES[type] ?? type;
|
|
106
|
+
const anchors = anchorsByType().get(resolvedType) ?? new Set();
|
|
107
|
+
if (anchors.size !== 1) return undefined;
|
|
108
|
+
const [anchor] = anchors;
|
|
109
|
+
return anchor && isKnownDocAnchor(anchor) ? anchor : undefined;
|
|
110
|
+
}
|
|
111
|
+
|
|
112
|
+
// Resolves a docAnchorForType() result to the tuning-doc file slug under docs-content/tuning/.
|
|
113
|
+
// Reuses pageForAnchor's sub-anchor resolution so a sub-anchor like bottleneck-stage-shape maps
|
|
114
|
+
// to its owning page (skew.md), not a stage-shape.md that never exists. Returns null for anchors
|
|
115
|
+
// with no bottleneck tuning doc (metric-/config-prefixed, or a page section like #memory-model).
|
|
116
|
+
export function tuningDocSlugForAnchor(anchor ) {
|
|
117
|
+
const bare = String(anchor).replace(/^#/, '');
|
|
118
|
+
const page = pageForAnchor(bare);
|
|
119
|
+
return page.startsWith('bottleneck-') ? page.slice('bottleneck-'.length) : null;
|
|
120
|
+
}
|
|
@@ -0,0 +1,32 @@
|
|
|
1
|
+
# Introduction
|
|
2
|
+
|
|
3
|
+
This guide is a hands-on reference for optimizing Spark and PySpark jobs. It covers [Spark
|
|
4
|
+
internals](#spark-architecture), [memory management](#memory-model), [joins](#joins),
|
|
5
|
+
[shuffle](#shuffle), [data formats](#data-formats), [PySpark-specific patterns](#pyspark),
|
|
6
|
+
[Adaptive Query Execution](#aqe) (AQE), and [cluster tuning](#cluster-config), followed by
|
|
7
|
+
a bottleneck diagnostic reference.
|
|
8
|
+
|
|
9
|
+
Read it top to bottom for a full grounding in Spark performance, or jump straight to the
|
|
10
|
+
**Bottleneck Reference** section below when you already know which symptom you're chasing.
|
|
11
|
+
|
|
12
|
+
## Severity dots
|
|
13
|
+
|
|
14
|
+
Each finding is marked with a severity dot:
|
|
15
|
+
|
|
16
|
+
- <span class="severity-dot info"></span> **Info**: worth knowing, not yet a problem
|
|
17
|
+
- <span class="severity-dot warning"></span> **Warning**: likely hurting job performance
|
|
18
|
+
- <span class="severity-dot critical"></span> **Critical**: actively bottlenecking the job
|
|
19
|
+
|
|
20
|
+
Each bottleneck section below documents the exact thresholds behind these dots.
|
|
21
|
+
|
|
22
|
+
## Tag system
|
|
23
|
+
|
|
24
|
+
Every bottleneck has a short tag used throughout this reference:
|
|
25
|
+
|
|
26
|
+
<span class="tag">SKEW</span> <span class="tag">SHFL</span> <span class="tag">SPILL</span>
|
|
27
|
+
<span class="tag">GC</span> <span class="tag">COLD</span> <span class="tag">UTIL</span>
|
|
28
|
+
<span class="tag">HOST</span> <span class="tag">FAIL</span> <span class="tag">STRAG</span>
|
|
29
|
+
<span class="tag">RETRY</span> <span class="tag">TINY</span> <span class="tag">FAIL-RATE</span>
|
|
30
|
+
|
|
31
|
+
See the [Bottleneck Reference](#bottleneck-skew) for what each one means, how it's detected,
|
|
32
|
+
and how to fix it.
|
|
@@ -0,0 +1,76 @@
|
|
|
1
|
+
# Spark Execution Model
|
|
2
|
+
|
|
3
|
+
## The driver and the DAG
|
|
4
|
+
|
|
5
|
+
The Driver is the process "in the driver seat" of a Spark application: it controls execution and maintains all state of the cluster, including the state and tasks of the executors, and it interfaces with the cluster manager to obtain physical resources and launch executors[^1]. It runs as a single JVM process (on the submission machine in client mode, or on a dedicated cluster node in cluster mode), and if it dies, the application dies with it[^2].
|
|
6
|
+
|
|
7
|
+
As application code executes, the Driver incrementally builds a logical DAG of RDD/DataFrame transformations. Because Spark is lazy, nothing actually computes until an action is called; only then does the Driver convert the logical DAG into a physical plan, cut it into stages at wide-dependency ([shuffle](#shuffle)) boundaries, break each stage into one task per partition, and take on the scheduler's job: talking to the cluster resource manager to request executors, handing out tasks, tracking their progress, and resubmitting tasks whose executor died[^2].
|
|
8
|
+
|
|
9
|
+
Three components live only on the Driver:
|
|
10
|
+
|
|
11
|
+
- The **DAGScheduler** reads RDD lineage, decides stage boundaries at wide dependencies, emits per-partition TaskSets, tracks lineage so lost partitions can be recomputed, and reacts dynamically to stage completions rather than pre-scheduling the whole DAG up front; it also declares a job failed if a stage can't make progress[^2].
|
|
12
|
+
- The **BlockManagerMaster** is the Driver-side counterpart to each executor's own BlockManager. It keeps the global map of block locations (RDD partitions, shuffle files, broadcast variables) across the whole cluster, so a task's local BlockManager knows where to fetch a remote block from[^2].
|
|
13
|
+
- The **SparkContext** (wrapped today by SparkSession) is the entry point that lets the Driver talk to the cluster manager, request executors, create RDDs, and manage shared variables. It's instantiated once per application and lives for the application's whole lifetime[^2].
|
|
14
|
+
|
|
15
|
+
Executors, by contrast, hold no cluster-wide state: they only run the tasks assigned to them and report back success, failure, and results[^1]. The split even shows up at the configuration level: `spark.driver.host`/`spark.driver.port` and `spark.driver.blockManager.port` are Driver-specific listening endpoints, distinct from the equivalent executor settings[^3].
|
|
16
|
+
|
|
17
|
+
<img class="light-only" src="diagrams/driver-executor.svg" alt="A Spark driver with its scheduler components dispatching tasks to executors through the cluster manager and receiving status and results back.">
|
|
18
|
+
<img class="dark-only" src="diagrams/driver-executor.dark.svg" alt="A Spark driver with its scheduler components dispatching tasks to executors through the cluster manager and receiving status and results back.">
|
|
19
|
+
|
|
20
|
+
A Spark job corresponds to one action, and each job breaks down into a series of stages: how many depends on how many shuffle operations need to happen[^1]. A stage is a group of tasks that can execute together to compute the same operation across machines: work one executor can do without communicating with other executors or the Driver. A new stage begins whenever data has to move across the network, that is, at a shuffle[^4]. Wide transformations, such as `groupByKey`, `join`, and `sortByKey`, create `ShuffleDependency` objects, and it's these `ShuffleDependency`s that mark the stage boundary; several narrow transformations can be grouped into the same stage[^4]. Within a stage, the number of tasks equals the number of partitions in that stage's output RDD: one task per partition, each running the same code on a different slice of data[^4].
|
|
21
|
+
|
|
22
|
+
The rule is always "a shuffle dependency creates the boundary," and that isn't limited to an explicit `.repartition()` or `.groupBy()` call: `sortByKey`/`sortBy` on an RDD is itself a wide transformation, not just `groupByKey`-style aggregations[^4]. The original Spark paper defines stage boundaries generically, as the shuffle operations required for wide dependencies, or already-computed partitions that can short-circuit the computation of a parent RDD[^5]. A wide dependency of any kind is what matters, not a specific API call.
|
|
23
|
+
|
|
24
|
+
That wide/narrow split has a formal definition behind it. Conceptually, a narrow transformation is one where each partition in the child RDD has simple, finite dependencies on partitions in the parent RDD, determinable at design time regardless of the values of the records; narrow dependencies allow pipelined, one-node execution, while wide dependencies require data from all parent partitions to be shuffled across nodes[^5]. The formal 2012 definition is stated from the parent's side rather than the child's: a transformation is narrow if "each partition of the parent RDD is used by at most one partition of the child RDD," and wide if "multiple child partitions may depend on" a given parent partition[^4]. That parent-centric framing is more precise, because Spark's DAG scheduler builds the execution plan backward from the action to the input RDD, and it correctly rules out the case of one parent partition feeding multiple children under "narrow"[^4].
|
|
25
|
+
|
|
26
|
+
`mapPartitions` is narrow under this definition: each child partition depends on exactly one parent partition, the same as `map` and `filter`[^4]. `coalesce` is narrow too, even though it changes the number of partitions and a child partition can depend on multiple parent partitions: the rule only requires that each parent partition be used by at most one child partition, and which parent partitions merge into which child is fixed at design time, independent of the data's values. That holds specifically when `coalesce` reduces the partition count; when it increases the count, it behaves like `repartition`, a shuffle[^4].
|
|
27
|
+
|
|
28
|
+
Pipelining is what makes narrow dependencies pay off. Spark performs as many steps as possible in one pass before writing data to memory or disk: any sequence of operations that feed data directly into each other, without moving data across nodes, collapses into a single stage of tasks that execute all the operations together. `map → filter → map` becomes one stage whose tasks read each record and pass it through all three operations in sequence, rather than materializing intermediate results after each step; the same collapsing happens for a DataFrame/SQL computation doing `select → filter → select`[^1]. The original paper phrases the mechanism at the scheduler level: to run an action, Spark builds stages at wide dependencies and pipelines narrow transformations inside each stage[^5].
|
|
29
|
+
|
|
30
|
+
Two more components take over once a job's stages exist. The **DAGScheduler** is Spark's high-level scheduling layer: it takes RDD dependencies, builds a DAG of stages for each job, determines where each task should run, and passes that to the TaskScheduler[^4]. The **TaskScheduler** takes it from there and stops thinking in terms of RDDs, lineage, or shuffles: only machines and slots. It works with the SchedulerBackend to request executors, assigns tasks to executors while respecting data-locality preferences, and retries failed tasks, including resubmitting a task elsewhere if its executor died[^2]. The TaskScheduler doesn't launch tasks directly; it hands them to the SchedulerBackend, the layer that actually talks to the cluster manager to request, launch, and kill executors[^2].
|
|
31
|
+
|
|
32
|
+
Finally, at the RDD API level `spark.default.parallelism` sets the default partition count for distributed shuffle operations like `reduceByKey` and `join`. It defaults to the largest number of partitions in the parent RDD, and for operations with no parent RDD, such as `parallelize`, the default depends on the cluster manager[^3]. `spark.sql.shuffle.partitions` is the equivalent knob for the Spark SQL/DataFrame engine: it configures how many partitions Spark uses when shuffling data for joins or aggregations, and it defaults to 200 regardless of data size[^6], applying whenever a DataFrame/SQL operation triggers a shuffle: join, groupBy, distinct, orderBy, and so on[^7].
|
|
33
|
+
|
|
34
|
+
## Reading stages and tasks
|
|
35
|
+
|
|
36
|
+
Once a job runs, stage and task counts show up directly in the Spark UI, and a worked example makes the mechanics concrete: a job that reads a range with 8 partitions produces a stage with 8 tasks; repartitioning to 6 and then 5 partitions produces stages with 6 and 5 tasks; a subsequent join shuffles into the default 200 shuffle partitions, producing a 200-task stage[^1]. The same logic explains stage counts for RDD pipelines: the chain `filter → map → groupByKey → map → sortByKey → count` produces three stages, bounded by the `groupByKey` and `sortByKey` operations, since both are wide[^4].
|
|
37
|
+
|
|
38
|
+
<img class="light-only" src="diagrams/dag-stages.svg" alt="A three stage DAG for the pipeline filter, map, groupByKey, map, sortByKey, count, with stage boundaries at the wide groupByKey and sortByKey operations.">
|
|
39
|
+
<img class="dark-only" src="diagrams/dag-stages.dark.svg" alt="A three stage DAG for the pipeline filter, map, groupByKey, map, sortByKey, count, with stage boundaries at the wide groupByKey and sortByKey operations.">
|
|
40
|
+
|
|
41
|
+
Pipelining is invisible to the application and only shows up in the Spark UI or logs, where multiple chained narrow operations appear collapsed into a single stage instead of one stage per operation[^1]. A related signal is the "skipped stage" marking: because Spark always writes shuffle output to stable storage regardless of any `persist`/`checkpoint` call, if the Driver reuses an RDD that was already shuffled, Spark can skip recomputing everything up to that shuffle and read the shuffle files directly. The Spark UI shows this as a skipped stage, not as an added boundary[^4].
|
|
42
|
+
|
|
43
|
+
When diagnosing partition counts, checking which knob is in play matters: `spark.default.parallelism` governs the legacy RDD API's shuffle and `parallelize` operations, while `spark.sql.shuffle.partitions` governs DataFrame/Dataset/SQL shuffles[^3]. Since Spark 3.0, [Adaptive Query Execution](#aqe) (AQE) can also override the static `spark.sql.shuffle.partitions` value after the fact. It can coalesce many small post-shuffle partitions and split skewed ones at runtime, but only after the first shuffle has already happened, so it never explains the initial input partitioning[^7].
|
|
44
|
+
|
|
45
|
+
## Why the boundary matters
|
|
46
|
+
|
|
47
|
+
Those stage boundaries aren't just a UI detail. Because a new stage begins only at a shuffle, the boundary is exactly where Spark pays the cost of moving data across the network. Everything inside a stage runs without that cost, pipelined on a single executor[^4].
|
|
48
|
+
|
|
49
|
+
That's also why the wide/narrow distinction determines whether pipelining is even possible: once a shuffle is required, downstream computation can't proceed until the shuffle completes, because the records landing on each partition may change as a result of it, so narrow transformations following a wide one belong to a new stage and can't be pipelined across that boundary[^4].
|
|
50
|
+
|
|
51
|
+
Whether pipelining happens can also depend on partitioning state Spark already knows about. Because Spark tracks how an RDD is already partitioned, the same operation can land in different stage boundaries depending on whether the input RDD already has a known partitioner: no shuffle, and hence no new stage, is needed if one is already in place[^4].
|
|
52
|
+
|
|
53
|
+
[Checkpointing](#caching) is a separate mechanism from all of this: it breaks RDD lineage and writes data to disk so later transformations start from a fresh, trimmed plan, rather than being pipelined against the original computation[^8].
|
|
54
|
+
|
|
55
|
+
The choice of [shuffle-partition count](#partitioning) matters for the same reason: there's no fixed formula for the right number, since it depends on data set size, number of cores, and available executor memory, and the default of 200 is called out as too high for smaller or streaming workloads, where it's better reduced toward the number of executor cores or less[^9]. `coalesce`'s narrowness carries its own tradeoff: reducing partition count without a shuffle forces all upstream partitions in that stage to run at the coalesced parallelism level, which can be undesirable[^4].
|
|
56
|
+
|
|
57
|
+
On the scheduling side, Spark's scheduler is fully thread-safe and supports running multiple jobs concurrently if they're submitted from separate Driver threads, common when an application serves multiple concurrent requests over the network[^10]. By default Spark schedules jobs FIFO, which means an application running many jobs from multiple threads can have later jobs wait behind earlier ones unless a fairer policy is configured[^4].
|
|
58
|
+
|
|
59
|
+
## Controlling parallelism and scheduling
|
|
60
|
+
|
|
61
|
+
Parallelism and scheduling are both tunable from here. When `coalesce`'s single-parent-per-child constraint forces upstream parallelism lower than you want, `repartition` trades that limitation for an explicit shuffle, restoring full control over the resulting partition count[^4]. For the SQL/DataFrame engine, don't leave `spark.sql.shuffle.partitions` at its 200 default for small or streaming workloads: reduce it toward the number of executor cores or less, since the right value depends on data size, core count, and executor memory rather than a fixed rule[^9]. Where the data volume is unpredictable, Adaptive Query Execution (available since Spark 3.0) can take over after the fact: it coalesces many small post-shuffle partitions and splits skewed ones at runtime, though it only acts after the first shuffle has already happened and won't fix the initial input partitioning[^7].
|
|
62
|
+
|
|
63
|
+
For concurrent workloads, Spark also offers a fair scheduler as an alternative to the FIFO default, assigning tasks to concurrent jobs round-robin so each job gets a more even share of cluster resources[^4]; job pools and weights for finer-grained sharing are configured via `spark.scheduler.mode=FAIR` and the `spark.scheduler.pool` local property set on the submitting thread[^1]. Submitting jobs from separate Driver threads is what lets them run concurrently in the first place, rather than queuing behind each other[^10].
|
|
64
|
+
|
|
65
|
+
## Sources
|
|
66
|
+
|
|
67
|
+
[^1]: *Spark: The Definitive Guide*, Chambers & Zaharia, ch. 15–16
|
|
68
|
+
[^2]: [Anatomy of Spark Application](https://luminousmen.com/post/spark-anatomy-of-spark-application)
|
|
69
|
+
[^3]: [Configuration — Spark](https://spark.apache.org/docs/latest/configuration.html)
|
|
70
|
+
[^4]: *High Performance Spark, 2nd Edition*, Karau, Polak & Warren, ch. 2, 7–8
|
|
71
|
+
[^5]: [Resilient Distributed Datasets: A Fault-Tolerant Abstraction for In-Memory Cluster Computing](https://www.usenix.org/system/files/conference/nsdi12/nsdi12-final138.pdf)
|
|
72
|
+
[^6]: [Performance Tuning — Spark SQL, DataFrames and Datasets Guide](https://spark.apache.org/docs/latest/sql-performance-tuning.html)
|
|
73
|
+
[^7]: [Spark Partitions](https://luminousmen.com/post/spark-partitions)
|
|
74
|
+
[^8]: [Spark Tips: Caching](https://luminousmen.com/post/spark-tips-caching)
|
|
75
|
+
[^9]: *Learning Spark, 2nd Edition*, Damji, Wenig, Das & Lee, ch. 7
|
|
76
|
+
[^10]: [Job Scheduling](https://spark.apache.org/docs/latest/job-scheduling.html)
|
|
@@ -0,0 +1,73 @@
|
|
|
1
|
+
# Memory Management
|
|
2
|
+
|
|
3
|
+
## The unified memory model
|
|
4
|
+
|
|
5
|
+
Since Spark 1.6, the `UnifiedMemoryManager` has replaced the earlier static split between execution and storage memory with a single shared region, referred to as M[^1][^2]. M is carved out of the JVM heap after first setting aside 300 MiB of reserved system memory: `M = (JVM heap − 300 MiB) × spark.memory.fraction`, with `spark.memory.fraction` defaulting to 0.6[^2][^3]. Within M, `spark.memory.storageFraction` (default 0.5) marks off a subregion reserved for cached blocks that are immune to eviction; with defaults, this works out to roughly a 30%/30% on-heap split between execution and storage out of total executor memory[^3].
|
|
6
|
+
|
|
7
|
+
<img class="light-only" src="diagrams/memory-regions.svg" alt="The executor memory layout showing reserved memory, the unified region split into execution and storage, and user memory inside the JVM heap, with the off-heap pool and memory overhead outside it.">
|
|
8
|
+
<img class="dark-only" src="diagrams/memory-regions.dark.svg" alt="The executor memory layout showing reserved memory, the unified region split into execution and storage, and user memory inside the JVM heap, with the off-heap pool and memory overhead outside it.">
|
|
9
|
+
|
|
10
|
+
Execution and storage borrow from each other dynamically rather than sitting behind a hard wall: when execution memory is unused, storage can acquire all of it, and vice versa[^2][^1]. The two are not equal peers, though. Execution always has priority, taking memory immediately and evicting cached storage blocks if necessary[^3].
|
|
11
|
+
|
|
12
|
+
The reserved 300 MB is a flat constant (`RESERVED_SYSTEM_MEMORY_BYTES` in the Spark source), not a formula, and it's hardcoded: there is no supported production configuration to change it. The only override is the internal `spark.testing.reservedMemory` property, which exists for Spark's own test suite rather than production tuning[^3]. `spark.memory.fraction` is applied against heap minus that reserved 300 MB, not against total heap. The tuning guide is explicit that it "expresses the size of M as a fraction of the (JVM heap space − 300MiB)"[^2].
|
|
13
|
+
|
|
14
|
+
Off-heap memory extends the same model outside the JVM heap. When `spark.memory.offHeap.enabled=true`, Spark allocates raw off-heap buffers via `sun.misc.Unsafe` (routed through `jdk.internal.misc.Unsafe` on JDK 17+, which is why some builds need `--add-opens` flags), and that off-heap region is split into its own execution and storage pools, following the same borrowing rules as on-heap memory: execution has priority, storage gets evicted first[^3]. Enabling it requires `spark.memory.offHeap.size` to be set to a positive value; leaving it enabled with the default size of `0` is a documented-invalid combination, not a supported way to run with off-heap "on but empty"[^4]. The setting also "has no impact on heap memory usage". Turning on off-heap memory does not shrink the JVM `-Xmx` for you, so the on-heap size has to be reduced manually to keep total footprint constant[^4]. This off-heap execution memory underlies Project Tungsten: its compact binary row encoding, its explicit-memory-managed hash map for aggregations, and its cache-aware sort/join algorithms are all built to run against off-heap, GC-invisible memory[^5][^6].
|
|
15
|
+
|
|
16
|
+
Executor memory overhead is a separate pool layered on top of M. `spark.executor.memoryOverhead` defaults to `executorMemory × spark.executor.memoryOverheadFactor`, floored at `spark.executor.minMemoryOverhead` (384 MiB)[^4]. `spark.executor.memoryOverheadFactor` itself defaults to 0.10 for ordinary JVM executors, but Spark bumps that default to 0.40 specifically for Kubernetes non-JVM jobs, since those workloads tend to need more non-JVM heap space[^4]. The legacy, YARN-only `spark.yarn.executor.memoryOverhead` property was removed in Spark 3.0 in favor of the cluster-manager-agnostic `spark.executor.memoryOverhead`[^3].
|
|
17
|
+
|
|
18
|
+
> **PySpark:** `spark.python.worker.memory` (default `512m`) is a soft spill threshold for a Python worker's own aggregation buffering, not a JVM-enforced cap. A PySpark worker runs as a separate OS process outside the JVM heap, so the JVM cannot police it directly[^3]. The only setting that actually tries to bound PySpark memory per executor is `spark.executor.pyspark.memory`, and even that depends on Python's `resource` module, which isn't supported on Windows and doesn't actually limit anything on macOS[^4]. Left unset, PySpark memory usage is folded into `spark.executor.memoryOverhead` by default[^4].
|
|
19
|
+
|
|
20
|
+
## Watching eviction and spill
|
|
21
|
+
|
|
22
|
+
That borrowing leaves a trace once eviction actually happens. Storage eviction under memory pressure follows an LRU (least recently used) policy[^1][^7]. The granularity is the block: the `BlockManager` tracks cached data as per-partition blocks (one block per RDD/DataFrame partition) and LRU operates at that level rather than evicting an entire RDD or cached table in one shot[^8][^9]. In practice this tends to evict the oldest partitions first, those materialized in the earliest job or stage, though lazy evaluation makes it hard to predict exactly which partitions will go ahead of time[^9].
|
|
23
|
+
|
|
24
|
+
What happens to an evicted block is visible in its downstream effect and depends on its `StorageLevel`. For `MEMORY_ONLY`, the block is simply dropped and recomputed from the RDD's lineage the next time it's needed. For a disk-backed level like `MEMORY_AND_DISK`, the evicted block is written to disk and read back from there instead of being recomputed[^9][^7]. Blocks can also be evicted well before you ever try to reuse them: `.cache()` does not guarantee a block stays resident, especially on busy clusters or long pipelines[^7].
|
|
25
|
+
|
|
26
|
+
On the execution side, each task gets its own `TaskMemoryManager`, which enforces a soft per-task cap: with `n` tasks running concurrently, each task is allowed to allocate somewhere between `1/(2n)` and `1/n` of total execution memory, and the first task to arrive on an idle executor typically grabs more than its later-arriving neighbors[^3].
|
|
27
|
+
|
|
28
|
+
When a sort or hash-aggregation operator keeps requesting execution-memory pages and can't get more, Spark doesn't fail immediately. It blocks the requesting task, spills that task's in-memory data structure to disk, or, in the worst case, throws an `OutOfMemoryError`[^3]. The spill path runs through Spark's map/shuffle I/O machinery: shuffle partitions created by wide transformations like `groupBy()` or `join()` [spill](#bottleneck-spill) to the executors' local disks at the location set by `spark.local.directory`[^6]. SQL physical operators apply the same idea with their own row-count guardrails: `sortMergeJoinExec`'s in-memory buffer and the cartesian-product operator's buffer both spill once they cross a configured row threshold, which by default is set to the value of `spark.shuffle.spill.numElementsForceSpillThreshold`[^10].
|
|
29
|
+
|
|
30
|
+
Whatever the trigger, the resulting spilled bytes are exposed on the task metrics as `memoryBytesSpilled`, visible in the Spark UI and event log. This is the concrete signal to watch for execution memory pressure[^11].
|
|
31
|
+
|
|
32
|
+
Memory-overhead misconfiguration shows up differently: it manifests as the container or pod being OOMKilled with a vague container-memory error rather than a Spark-level exception, since the enforcement happens at the YARN NodeManager or Kubernetes kubelet/cgroup layer, not inside the JVM[^3]. Under the old default overhead factor of 0.10, non-JVM Kubernetes jobs commonly failed with "Memory Overhead Exceeded" errors, which is exactly why Spark bumped the Kubernetes non-JVM default to 0.40[^4].
|
|
33
|
+
|
|
34
|
+
## What the borrowing costs you
|
|
35
|
+
|
|
36
|
+
None of this is free for whoever's counting on cached data staying put. Because execution always wins the tug-of-war over storage, [caching](#caching) is never a guarantee; it's a best effort. As luminousmen-memory-management puts it: "If Execution needs memory, it takes it. If Storage is using that space (cached RDDs, broadcasts), Spark starts evicting blocks. If Execution is idle, Storage can grow into that space, until Execution comes back."[^3] That growth-then-shrink-back behavior is dynamic borrowing, not storage evicting execution; eviction is strictly one-directional. As spark-notes-task-memory-management summarizes the agreement between the two regions: "keep acquiring execution memory and evict storage as you need more execution memory"[^1], never the reverse. The same asymmetric rule carries over when off-heap memory is enabled: execution still has priority and storage still gets evicted first[^3]. Practically, this means a cached DataFrame can silently lose blocks under memory pressure, forcing an expensive lineage recompute (for `MEMORY_ONLY`) or an extra disk round-trip (for `MEMORY_AND_DISK`) the next time it's touched.
|
|
37
|
+
|
|
38
|
+
<img class="light-only" src="diagrams/memory-borrowing.svg" alt="How storage memory borrows idle execution space while execution reclaims its own space by evicting storage in one direction.">
|
|
39
|
+
<img class="dark-only" src="diagrams/memory-borrowing.dark.svg" alt="How storage memory borrows idle execution space while execution reclaims its own space by evicting storage in one direction.">
|
|
40
|
+
|
|
41
|
+
The `TaskMemoryManager`'s soft per-task cap explains why spill behavior is workload-shape-dependent rather than a fixed threshold: with more tasks packed onto an executor, each one's guaranteed share of execution memory shrinks toward `1/(2n)`, making spills more likely under high task concurrency even when total execution memory hasn't changed[^3].
|
|
42
|
+
|
|
43
|
+
Off-heap memory's payoff is specifically about [garbage collection](#bottleneck-gc): because off-heap buffers sit outside the JVM heap, they are invisible to the garbage collector, so fewer and smaller live objects need to be tracked, scanned, and copied, which reduces both the frequency and duration of GC pauses[^3]. Project Tungsten's off-heap hash map for aggregations was benchmarked at over 1 million operations per second in a single thread, with "almost no performance degradation as memory utilization increases," unlike the JVM default `java.util.HashMap`, which eventually thrashes on GC[^5][^6]. The tradeoff is that off-heap memory removes the GC safety net along with the GC overhead: there's no garbage collector cleaning up if something goes wrong[^3].
|
|
44
|
+
|
|
45
|
+
Memory overhead matters because off-heap memory and (if unset) PySpark memory are not automatically folded into the overhead calculation the way heap memory is. Enabling `spark.memory.offHeap.size` without also raising `spark.executor.memoryOverhead` (or explicitly setting `spark.executor.pyspark.memory`) can push the executor's real footprint past what the cluster manager granted, and the process gets killed with a vague container-memory error instead of a Spark-level exception[^3]. Worked example: with `--executor-memory=8G`, the default 10% overhead gives `max(0.1 × 8192 MB, 384 MB) = 819 MB`, so the total memory requested from the cluster manager is `8192 + 819 = 9011 MB`[^3], useful to keep in mind when [sizing containers or pods](#cluster-config) against a cluster's available capacity.
|
|
46
|
+
|
|
47
|
+
## Tuning the memory pools
|
|
48
|
+
|
|
49
|
+
Most of this is adjustable, within limits. Tune `spark.memory.fraction` and `spark.memory.storageFraction` only if your workload's execution/storage balance genuinely needs to shift away from the ~30/30 default. But don't try to reclaim the 300 MB reserved region; it's a fixed, non-tunable constant in production[^3].
|
|
50
|
+
|
|
51
|
+
If losing cached partitions to eviction is costly, prefer a disk-backed `StorageLevel` such as `MEMORY_AND_DISK` over `MEMORY_ONLY`: an evicted block gets written to disk and read back rather than triggering a full lineage recompute[^9][^7]. Keep in mind `.cache()` alone is not a residency guarantee even with this choice[^7].
|
|
52
|
+
|
|
53
|
+
Watch `memoryBytesSpilled` in the Spark UI or history server to catch execution-memory pressure early[^11]. If sort or hash-aggregation spills are heavy, consider giving the executor more memory or reducing the number of concurrently running tasks per executor, since the `TaskMemoryManager`'s per-task guarantee shrinks as task concurrency `n` grows[^3]. For SQL joins and cartesian products, `spark.shuffle.spill.numElementsForceSpillThreshold` governs when `sortMergeJoinExec` and the cartesian-product operator spill their buffers[^10].
|
|
54
|
+
|
|
55
|
+
Set `spark.executor.memoryOverhead` explicitly rather than relying on the default, especially on Kubernetes for non-JVM workloads (where the default factor is already bumped to 0.40) or whenever off-heap memory or heavy PySpark usage is in play: those aren't automatically folded into the overhead calculation, so an unadjusted overhead can lead to an OOMKilled container instead of a clean Spark-level error[^3][^4]. Use `spark.executor.memoryOverhead`, not the removed `spark.yarn.executor.memoryOverhead`[^3]. For capacity planning, remember the total requested from the cluster manager is executor memory plus overhead, e.g., `8192 + 819 = 9011 MB` for an 8 GB executor at the default 10% factor[^3].
|
|
56
|
+
|
|
57
|
+
When enabling off-heap memory, always pair `spark.memory.offHeap.enabled=true` with an explicit positive `spark.memory.offHeap.size` (never leave it enabled at the default size of `0`[^4]) and manually reduce the JVM `-Xmx` to compensate, since enabling off-heap memory does not shrink the heap for you[^4].
|
|
58
|
+
|
|
59
|
+
> **PySpark:** if you need PySpark's own memory bounded rather than folded silently into the overhead, set `spark.executor.pyspark.memory` explicitly, though its enforcement relies on Python's `resource` module and won't work on Windows and won't actually limit anything on macOS[^4].
|
|
60
|
+
|
|
61
|
+
## Sources
|
|
62
|
+
|
|
63
|
+
[^1]: [Task Memory Management in Spark](https://raw.githubusercontent.com/spoddutur/spark-notes/master/task_memory_management_in_spark.md)
|
|
64
|
+
[^2]: [Tuning Spark](https://spark.apache.org/docs/latest/tuning.html)
|
|
65
|
+
[^3]: [Dive into Spark Memory](https://luminousmen.com/post/dive-into-spark-memory)
|
|
66
|
+
[^4]: [Configuration — Spark](https://spark.apache.org/docs/latest/configuration.html)
|
|
67
|
+
[^5]: [Project Tungsten: Bringing Spark Closer to Bare Metal](https://www.databricks.com/blog/2015/04/28/project-tungsten-bringing-spark-closer-to-bare-metal.html)
|
|
68
|
+
[^6]: *Learning Spark, 2nd Edition*, Damji, Wenig, Das, Lee, ch. 6–7
|
|
69
|
+
[^7]: [Explaining the Mechanics of Spark Caching](https://luminousmen.com/post/explaining-the-mechanics-of-spark-caching)
|
|
70
|
+
[^8]: [RDD Programming Guide](https://spark.apache.org/docs/latest/rdd-programming-guide.html)
|
|
71
|
+
[^9]: *High Performance Spark, 2nd Edition*, Karau, Polak & Warren, ch. 7
|
|
72
|
+
[^10]: [SQLConf.scala](https://raw.githubusercontent.com/apache/spark/v3.5.0/sql/catalyst/src/main/scala/org/apache/spark/sql/internal/SQLConf.scala)
|
|
73
|
+
[^11]: [Monitoring and Instrumentation](https://spark.apache.org/docs/latest/monitoring.html#spark-history-server)
|
|
@@ -0,0 +1,65 @@
|
|
|
1
|
+
# Partitioning
|
|
2
|
+
|
|
3
|
+
## How partitions get sized and shuffled
|
|
4
|
+
|
|
5
|
+
Every partition Spark creates maps to exactly one task and one thread, which is why partition count and size drive so much of a job's performance[^1]. The [shuffle](#shuffle) side of that is governed by a single default: `spark.sql.shuffle.partitions`, which defaults to 200 and applies to every `join()`, `groupBy()`, and aggregation regardless of how much data is actually moving[^2][^3].
|
|
6
|
+
|
|
7
|
+
Two operations reshape that layout, and they are not interchangeable. `repartition(numPartitions)` "return[s] a new RDD that has exactly numPartitions partitions"[^4], and it earns that guarantee by always running a full hash shuffle: records are first spread across a temporary key space starting from a randomized position, seeded per-partition via `XORShiftRandom`, so upstream data ends up distributed evenly rather than clustered by its original layout. That feeds into a `ShuffledRDD` keyed by a `HashPartitioner(numPartitions)`, then a `CoalescedRDD` lands it on exactly the requested count[^4]. High Performance Spark puts the same mechanics more plainly: "repartition shuffles the RDD with a hash partitioner and the given number of partitions"[^5]. That holds whether the new count is bigger or smaller than the current one. Repartition always performs the full shuffle just described, unlike coalesce[^6].
|
|
8
|
+
|
|
9
|
+
`coalesce`, left at its default, does not shuffle at all: it's a narrow transformation where each output partition is simply the union of a fixed set of parent partitions decided at plan time, not routed by data values, so the stage's task count just drops to the coalesced number[^5].
|
|
10
|
+
|
|
11
|
+
Neither `repartition` nor `coalesce` leaves behind a "known partitioner" the way `partitionBy` does[^5]. That distinction is what Spark's join planner checks before it will skip a shuffle: it only does so when both sides already carry partitioner objects it can prove equal: matching partition counts for a `HashPartitioner`, matching range bounds for a `RangePartitioner`[^5]. Its default shuffle hash join path makes the same point from the other side: it partitions the second dataset using the same partitioner as the first specifically so matching keys land together, which only lets a later shuffle be skipped when a shared, recognized partitioner is already in place beforehand[^5].
|
|
12
|
+
|
|
13
|
+
There's also a hard ceiling on shuffle output, though it isn't something to design around: Spark's sort-based shuffle manager caps output partitions in serialized mode at `PackedRecordPointer.MAXIMUM_PARTITION_ID + 1`, roughly 16.8 million. The source code itself calls this "an extreme defensive programming measure," since no real shuffle comes remotely close to it[^7].
|
|
14
|
+
|
|
15
|
+
## Spotting a bad layout
|
|
16
|
+
|
|
17
|
+
That layout choice shows up as overhead in both directions once it's wrong: partitions that are too small flood the cluster with per-task scheduling overhead, while partitions that are too large create memory pressure and straggler tasks[^1]. Learning Spark 2nd Edition frames the healthy target from the parallelism side rather than an absolute number: at least as many partitions as there are cores across the executors, so no core sits idle; more partitions than cores is fine as long as it doesn't drift into the small-partition overhead regime above[^3].
|
|
18
|
+
|
|
19
|
+
A heavy filter is an easy-to-miss cause of imbalance: Spark doesn't shrink partition count when rows are filtered out, so 2,000 partitions holding 5% of the original rows just become 2,000 mostly-empty partitions[^1].
|
|
20
|
+
|
|
21
|
+
To size shuffle partitions from a job's actual behavior rather than a guess, Cloudera's tuning guide takes a stage that already ran, computes the ratio between its Shuffle Spill (Memory) and Shuffle Spill (Disk) metrics, and multiplies total shuffle write by that ratio to estimate in-memory shuffle size, then rounds the resulting partition count up rather than down[^8].
|
|
22
|
+
|
|
23
|
+
[Skew](#bottleneck-skew) specifically has documented, numeric detection thresholds under [Adaptive Query Execution](#aqe): a partition counts as skewed if it's larger than `spark.sql.adaptive.skewJoin.skewedPartitionFactor` (default 5.0) times the median partition size, and also larger than `spark.sql.adaptive.skewJoin.skewedPartitionThresholdInBytes` (default 256 MB)[^2][^9].
|
|
24
|
+
|
|
25
|
+
On the input side, `spark.sql.files.maxPartitionBytes` (defaulting to 128 MB) governs the target chunk size when Spark splits splittable file-based sources (CSV, JSON, line-delimited text) into input partitions, based on file size[^1].
|
|
26
|
+
|
|
27
|
+
## Where a bad layout costs you
|
|
28
|
+
|
|
29
|
+
Getting the partition count wrong costs more than the immediate slowdown suggests. Cloudera's tuning guide argues it's safer to over-provision partitions than to under-provision them, because Spark, unlike MapReduce, has low per-task startup overhead: "when in doubt, it's almost always better to err on the side of a larger number of tasks"[^8].
|
|
30
|
+
|
|
31
|
+
Getting `coalesce` wrong costs more than the coalesce step itself: because it's narrow, it forces the *entire* upstream stage to run at the reduced parallelism, not just the final step[^5]. Pushed too far (`coalesce(1)`), it kills parallelism outright, because coalesce doesn't rebalance data, it just stacks existing partitions together, so partitions that were uneven going in are still uneven coming out[^1].
|
|
32
|
+
|
|
33
|
+
Once a shuffle stage has run, its output file count is locked in: you can't change it after the fact without inserting a stage barrier, such as writing to temporary storage or calling `localCheckpoint()`, between the shuffle and the write[^10].
|
|
34
|
+
|
|
35
|
+
Join alignment has its own gotcha: calling `repartition(n, col)` with the same column and count on two separate DataFrames produces data that's plausibly laid out the same way, but it doesn't leave a trackable partitioner object behind. The planner has no recorded partitioner to compare, so it has no basis for treating the two sides as co-partitioned and skipping the shuffle at join time, even though the call sites look identical[^5].
|
|
36
|
+
|
|
37
|
+
## Fixing the layout
|
|
38
|
+
|
|
39
|
+
Fixing those costs starts with accepting there's no formula to look up. There's no single formula for the right `spark.sql.shuffle.partitions` value: it depends on data set size, core count, and executor memory, and comes down to trial and error[^3]. As a starting point, Learning Spark notes the default of 200 is usually too high for small or streaming workloads, where it should be pulled down toward the executor core count[^3].
|
|
40
|
+
|
|
41
|
+
For the repartition-vs-coalesce decision itself: reach for `coalesce` first whenever you're only reducing partition count, since it merges partitions already on the same node without a shuffle[^6]. Reach for `repartition` when you actually need the shuffle: increasing partition count, fixing a lopsided distribution, or preparing a DataFrame ahead of a join or a `cache()` call, where even parallelism is worth more than the shuffle cost[^6]. High Performance Spark reduces this to a readability rule: use `repartition` when you want a shuffle, `coalesce` when you don't, rather than leaning on coalesce's shuffle toggle to blur the line[^5]. There is a middle option, `coalesce(n, shuffle=True)`, which behaves more like repartition, paying for a shuffle but getting real rebalancing on the way down[^1]. `coalesce()` itself belongs at the tail of a pipeline, right before a write, purely to cut down [output file count](#bottleneck-small-files)[^1]; after a heavy filter has left partitions mostly empty, following up with `repartition(100)` (or similar) buys back real parallelism[^1]. At the SQL layer, both directions have dedicated hints for controlling output partitioning directly: `/*+ REPARTITION(n) */`, `/*+ REPARTITION(cols) */`, `/*+ REPARTITION_BY_RANGE(cols) */`, and `/*+ COALESCE(n) */`[^9].
|
|
42
|
+
|
|
43
|
+
<img class="light-only" src="diagrams/repartition-vs-coalesce.svg" alt="A decision flowchart for choosing repartition versus coalesce based on whether a shuffle and a rebalance are needed.">
|
|
44
|
+
<img class="dark-only" src="diagrams/repartition-vs-coalesce.dark.svg" alt="A decision flowchart for choosing repartition versus coalesce based on whether a shuffle and a rebalance are needed.">
|
|
45
|
+
|
|
46
|
+
Separate from the DataFrame `.coalesce()` call, Adaptive Query Execution has its own runtime coalescing behavior: `spark.sql.adaptive.coalescePartitions.enabled` (default true) merges small, contiguous post-shuffle partitions toward a target size at runtime, correcting over-partitioning without a manual `.coalesce()` call[^9][^1].
|
|
47
|
+
|
|
48
|
+
For skewed joins specifically, salting (adding a prefix to skewed keys so the same key is treated as several different keys, then adjusting the data distribution accordingly) was one of three manual approaches used before adaptive execution existed, alongside raising `spark.sql.shuffle.partitions` and raising the broadcast hash join threshold to push a sort-merge join toward a broadcast hash join instead. All three carry "lots of limitations" and require manual processing, which is exactly the gap AQE's skew-join optimization was built to close[^11]. Where automatic detection isn't precise enough, Databricks' `/*+ SKEW(...) */` hint names the skewed relation and column(s), and optionally the specific skewed key values, letting the planner target just those keys directly instead of relying on automatic detection[^12].
|
|
49
|
+
|
|
50
|
+
For the join-alignment gotcha above, the mechanism that does reliably guarantee a shuffle can be skipped is storage-partitioned joins: both tables are physically bucketed identically at the catalog level, for example [Iceberg](#table-formats) tables created with matching `PARTITIONED BY (bucket(...))` clauses. When Spark recognizes both sides report the same partitioning through `SupportsReportPartitioning`, it can drop the Exchange (shuffle) node entirely, or shuffle only one side. A plain DataFrame-level `repartition(col)` doesn't offer that guarantee, because it isn't backed by catalog-level partitioning metadata[^9][^2].
|
|
51
|
+
|
|
52
|
+
## Sources
|
|
53
|
+
|
|
54
|
+
[^1]: [Spark Partitions](https://luminousmen.com/post/spark-partitions)
|
|
55
|
+
[^2]: [Configuration — Spark](https://spark.apache.org/docs/latest/configuration.html)
|
|
56
|
+
[^3]: *Learning Spark, 2nd Edition*, Damji, Wenig, Das & Lee, ch. 7
|
|
57
|
+
[^4]: [RDD.scala](https://raw.githubusercontent.com/apache/spark/v3.5.0/core/src/main/scala/org/apache/spark/rdd/RDD.scala)
|
|
58
|
+
[^5]: *High Performance Spark, 2nd Edition*, Karau, Polak & Warren, ch. 8
|
|
59
|
+
[^6]: *Spark: The Definitive Guide*, Chambers & Zaharia, ch. 19
|
|
60
|
+
[^7]: [SortShuffleManager.scala](https://raw.githubusercontent.com/apache/spark/v3.5.0/core/src/main/scala/org/apache/spark/shuffle/sort/SortShuffleManager.scala)
|
|
61
|
+
[^8]: [How to Tune Your Apache Spark Jobs (Part 2)](https://blog.cloudera.com/how-to-tune-your-apache-spark-jobs-part-2/)
|
|
62
|
+
[^9]: [Performance Tuning — Spark SQL, DataFrames and Datasets Guide](https://spark.apache.org/docs/latest/sql-performance-tuning.html)
|
|
63
|
+
[^10]: [Spark Tips: Partition Tuning](https://luminousmen.com/post/spark-tips-partition-tuning)
|
|
64
|
+
[^11]: [SPARK-29544 — Optimize skewed join at runtime](https://issues.apache.org/jira/browse/SPARK-29544)
|
|
65
|
+
[^12]: [Skew Join (legacy)](https://docs.databricks.com/aws/en/archive/legacy/skew-join)
|
|
@@ -0,0 +1,62 @@
|
|
|
1
|
+
# Join Optimization
|
|
2
|
+
|
|
3
|
+
## Join strategy selection
|
|
4
|
+
|
|
5
|
+
Spark SQL's Catalyst optimizer picks from five physical join operators: broadcast hash join, broadcast nested loop join, shuffle hash join, shuffle sort-merge join (SMJ), and shuffle-and-replicate nested loop (cartesian) join[^1]. Broadcast hash join avoids a shuffle entirely by sending one small side to every executor; it requires an equi-join condition and supports every join type except full outer[^1]. Broadcast nested loop join relaxes the equi-join requirement (it supports non-equi conditions and every join type) at the cost of scanning one side repeatedly, so it's normally a fallback rather than a first choice[^1]. Unlike hand-tuned RDD joins, where the partitioner is chosen explicitly, Spark SQL's optimizer can also push down or reorder other operators automatically to make the eventual join cheaper[^1].
|
|
6
|
+
|
|
7
|
+
At the communication level, Spark's choice is binary: an all-to-all [shuffle](#shuffle) join, or a broadcast join that replicates the small side once so every node can then join locally with no further network traffic[^2]. Once a join has been routed to a shuffle-based strategy (broadcast wasn't used), Spark defaults to preferring sort-merge join over shuffle hash join, a preference controlled by `spark.sql.join.preferSortMergeJoin`[^3]. Since Spark 3.2, [Adaptive Query Execution](#aqe) (AQE), on by default, re-optimizes the physical plan mid-execution using runtime statistics gathered after the shuffle actually runs, rather than relying only on pre-execution estimates[^4].
|
|
8
|
+
|
|
9
|
+
<img class="light-only" src="diagrams/join-strategy.svg" alt="Decision tree for choosing a physical join operator: the equi-join test splits off the nested-loop strategies, then the broadcast-threshold and preferSortMergeJoin checks select broadcast hash, sort-merge, or shuffle hash join, with AQE able to promote sort-merge to broadcast at runtime.">
|
|
10
|
+
<img class="dark-only" src="diagrams/join-strategy.dark.svg" alt="Decision tree for choosing a physical join operator: the equi-join test splits off the nested-loop strategies, then the broadcast-threshold and preferSortMergeJoin checks select broadcast hash, sort-merge, or shuffle hash join, with AQE able to promote sort-merge to broadcast at runtime.">
|
|
11
|
+
|
|
12
|
+
## Reading it in the plan
|
|
13
|
+
|
|
14
|
+
That choice is legible before a job even runs. The deciding signal is the relation's size relative to `spark.sql.autoBroadcastJoinThreshold`, which defaults to `10485760` bytes (10 MB), unchanged since it was introduced in Spark 1.1.0 and still the documented default in Spark 3.5[^4][^5]. That comparison is driven by table/plan size statistics rather than a fresh scan of the data: Spark consults catalog statistics (collected via `ANALYZE TABLE`, inspectable through `DESCRIBE EXTENDED`) and the cost estimates shown in `EXPLAIN COST`[^4]. When those statistics are missing, `spark.sql.statistics.fallBackToHdfs` (default `false`) controls whether Spark falls back to on-disk file size to judge broadcast eligibility, and for partitioned tables without statistics Spark instead uses the `spark.sql.defaultSizeInBytes` placeholder[^5].
|
|
15
|
+
|
|
16
|
+
> **PySpark:** don't guess whether a join will broadcast. Call `df.explain(mode="cost")` to see the same size estimates Catalyst used to pick the strategy.
|
|
17
|
+
|
|
18
|
+
Because AQE re-optimizes at runtime, the physical join operator visible in a plan can change after execution starts. Two runtime-driven overrides are worth checking for in the Spark UI or an `EXPLAIN` plan: `spark.sql.adaptive.maxShuffledHashJoinLocalMapThreshold` (default `0`, disabled) can make AQE prefer shuffled hash join over sort-merge "regardless of the value of `spark.sql.join.preferSortMergeJoin`" whenever every post-shuffle partition stays under that threshold and above `spark.sql.adaptive.advisoryPartitionSizeInBytes`[^4]; separately, AQE can convert an already-planned sort-merge join into a broadcast hash join mid-execution when the runtime statistics of either join side turn out smaller than the adaptive broadcast threshold[^4].
|
|
19
|
+
|
|
20
|
+
Bucketing status is also visible directly in the physical plan: a shuffle-free bucketed join only appears once both sides are bucketed to the same count, at which point the `Exchange` nodes disappear entirely[^8]. This is demonstrated with 4 buckets on each side[^6] and with 16 buckets on each side, where the plan shows `SelectedBucketsCount: 16 out of 16` on both branches of the `SortMergeJoin`[^7]. A separate real-world account of eliminating a bucketed shuffle frames the same signal just as plainly: the giveaway is that "the right branch is missing an Exchange (i.e. shuffle)"[^8]. Removing the shuffle this way does not necessarily remove the sort phase: in the 16-bucket demo, a `Sort` operator still appears on each branch of the `SortMergeJoin` even with the `Exchange` gone[^7].
|
|
21
|
+
|
|
22
|
+
Whether Catalyst is reordering joins (rather than just individual operators) is also a config check: `spark.sql.cbo.enabled` and the separate `spark.sql.cbo.joinReorder.enabled` flag both default to `false`, so multi-way join reordering is off unless both are explicitly enabled[^5]. That's distinct from Catalyst's regular rule-based optimizations, which run regardless of those flags. For example, a filter written after a join in the DataFrame API was observed moved before the join (and pushed into the JDBC source) automatically, visible in the physical plan[^9].
|
|
23
|
+
|
|
24
|
+
## Costs the plan doesn't show
|
|
25
|
+
|
|
26
|
+
Picking a strategy is one thing; paying for it is another. Broadcasting isn't free on the driver side. Building a broadcast join replicates the small-side DataFrame to every worker, but that replication is preceded by collecting the DataFrame back to the driver first: an [oversized broadcast](#bottleneck-broadcast-sizing), whether chosen automatically or forced via a hint or `broadcast()` call, "can crash your driver node (because that collect is expensive)"[^2]. The shuffle-and-replicate nested loop (cartesian-style) strategy carries a related but distinct risk: because every partition is joined against every other partition, it has a high chance of data explosion[^1].
|
|
27
|
+
|
|
28
|
+
Bucketing's shuffle-free path also has a real tradeoff once bucket counts don't match exactly. Since Spark 3.1, `spark.sql.bucketing.coalesceBucketsInJoin.enabled` lets Spark coalesce the side with more buckets down to the smaller count, but only within a ratio bounded by `spark.sql.bucketing.coalesceBucketsInJoin.maxBucketRatio` (default `4`). Enabling it can still remove the shuffle, but it does so by cutting parallelism on the finer-grained side, and Spark's own docs note it "could possibly cause OOM for shuffled hash join" as a result[^5]. Outside that ratio, or with coalescing disabled, mismatched bucket counts get no free win at all: the join simply falls back to a normal shuffle.
|
|
29
|
+
|
|
30
|
+
## Forcing a strategy
|
|
31
|
+
|
|
32
|
+
When the automatic choice isn't the right one, Spark still lets you override it directly. Join hints force a specific strategy per relation, and Spark honors them even against its own size-based defaults:
|
|
33
|
+
|
|
34
|
+
- **BROADCAST** (also accepted as `BROADCASTJOIN`/`MAPJOIN`) forces a broadcast join with the hinted relation as the build side (broadcast hash if there's an equi-join key, broadcast nested loop otherwise), and this is honored "even if the size of table 't1' suggested by the statistics is above the configuration `spark.sql.autoBroadcastJoinThreshold`"[^4].
|
|
35
|
+
- **MERGE** forces a shuffle sort-merge join; it needs an equi-join on sortable keys and high cardinality to pay off, and is less prone to OOM than the hash-based options since it never builds an in-memory hash table[^1].
|
|
36
|
+
- **SHUFFLE_HASH** forces a shuffle hash join, supporting all join types on an equi-join key but, like MERGE, wanting high cardinality; it builds a per-partition hash table after the shuffle[^1].
|
|
37
|
+
- **SHUFFLE_REPLICATE_NL** forces the shuffle-and-replicate nested loop join, supporting inner and cartesian joins with equi or non-equi conditions[^1].
|
|
38
|
+
|
|
39
|
+
When both sides of a join carry conflicting hints, Spark resolves them by a fixed priority: BROADCAST over MERGE over SHUFFLE_HASH over SHUFFLE_REPLICATE_NL. If both sides carry the *same* BROADCAST or SHUFFLE_HASH hint, Spark still picks a build side based on join type and relation sizes[^4]. As with any hint, none of this is guaranteed: Spark won't honor a hinted strategy that structurally can't support the actual join type being performed[^4]. The threshold itself can also be turned off outright: setting `spark.sql.autoBroadcastJoinThreshold` to `-1` forces every join to fall back to shuffle sort-merge instead of ever considering broadcast[^3].
|
|
40
|
+
|
|
41
|
+
For repeated joins on the same key, bucketing both tables to the *same* bucket count removes the shuffle for free. To also remove the sort phase (not just the shuffle), the bucketed tables additionally need to be written pre-sorted on the join key, using `sortBy` alongside `bucketBy`; done that way, the merge phase has nothing left to do, since "the joined output is sorted ... because we saved the tables sorted in ascending order ... there's no need to sort during the `SortMergeJoin`," and the Spark UI shows the query going straight to `WholeStageCodegen` with no `Exchange` at all[^3].
|
|
42
|
+
|
|
43
|
+
> **PySpark:** `df.write.bucketBy(n, "key").saveAsTable(...)` drops the shuffle; add `.sortBy("key")` on both tables to drop the sort phase too.
|
|
44
|
+
|
|
45
|
+
Multi-way join reordering is a deliberate opt-in: enable `spark.sql.cbo.enabled` for cost-based statistics, then `spark.sql.cbo.joinReorder.enabled` for the reordering rule itself. `spark.sql.cbo.joinReorder.dp.threshold` caps the dynamic-programming enumeration at 12 joined nodes by default, and `spark.sql.cbo.joinReorder.dp.star.filter` applies star-join filter heuristics on top of it[^5]. A lighter-weight alternative that skips full CBO is `spark.sql.cbo.starSchemaDetection` (default `false`), which enables join reordering based on star-schema detection alone[^5].
|
|
46
|
+
|
|
47
|
+
For [skewed join keys](#bottleneck-skew), Spark offers two built-in alternatives to hand-rolled salting: AQE's skew-join handling (`spark.sql.adaptive.skewJoin.enabled`), which detects oversized shuffle partitions at runtime and splits them automatically, replicating if needed[^10][^5], and Databricks' declarative `SKEW` hint, which builds a skew-aware plan without any manual salting[^11]. Manual salting remains the fallback where neither is available: add a random salt column to the join key on both sides so a hot key spreads across many partitions (exploding the dimension side into one row per salt value and assigning a random salt on the fact side), then join on the composite `(key, salt)` pair[^12].
|
|
48
|
+
|
|
49
|
+
## Sources
|
|
50
|
+
|
|
51
|
+
[^1]: *High Performance Spark, 2nd Edition*, Karau, Polak & Warren, ch. 6
|
|
52
|
+
[^2]: *Spark: The Definitive Guide*, Chambers & Zaharia, ch. 8
|
|
53
|
+
[^3]: *Learning Spark, 2nd Edition*, Damji, Wenig, Das, Lee, ch. 7
|
|
54
|
+
[^4]: [Performance Tuning — Spark SQL, DataFrames and Datasets Guide](https://spark.apache.org/docs/latest/sql-performance-tuning.html)
|
|
55
|
+
[^5]: [Configuration — Spark](https://spark.apache.org/docs/latest/configuration.html)
|
|
56
|
+
[^6]: [Bucketing — The Internals of Spark SQL](https://books.japila.pl/spark-sql-internals/bucketing/)
|
|
57
|
+
[^7]: [The 5-Minute Guide to Using Bucketing in PySpark](https://luminousmen.com/post/the-5-minute-guide-to-using-bucketing-in-pyspark)
|
|
58
|
+
[^8]: [Bucket the Shuffle Out of Here](https://www.taboola.com/engineering/bucket-the-shuffle-out-of-here/)
|
|
59
|
+
[^9]: [Spark Tips: DataFrame API](https://luminousmen.com/post/spark-tips-dataframe-api)
|
|
60
|
+
[^10]: [SPARK-29544 — Optimize Skewed Join at Runtime](https://issues.apache.org/jira/browse/SPARK-29544)
|
|
61
|
+
[^11]: [Skew Join Hint](https://docs.databricks.com/aws/en/archive/legacy/skew-join)
|
|
62
|
+
[^12]: [Spark Tips: Partition Tuning](https://luminousmen.com/post/spark-tips-partition-tuning)
|
|
@@ -0,0 +1,59 @@
|
|
|
1
|
+
# Shuffle
|
|
2
|
+
|
|
3
|
+
## How the shuffle works
|
|
4
|
+
|
|
5
|
+
A shuffle is Spark's mechanism for exchanging, sorting, grouping, and merging data across executors whenever an operation needs rows that share a key to land on the same partition. At the DataFrame/SQL level, the classic wide transformations that trigger one are `groupBy()`, `join()`, `agg()`, `sortBy()`, and `reduceByKey()`-style aggregations. Any join implementation that isn't a [broadcast join](#joins) (shuffle hash join, shuffle sort-merge join, or shuffle-and-replicated nested loop / Cartesian product join) also requires a shuffle[^1]. At the RDD layer underneath, the operations that can cause a shuffle are [repartitioning](#partitioning) (`repartition`, `coalesce`), `'ByKey` operations other than counting (`groupByKey`, `reduceByKey`), and join-family operations (`cogroup`, `join`)[^2].
|
|
6
|
+
|
|
7
|
+
Spark can skip the shuffle outright in a few well-defined cases where it already knows the data layout. Storage Partition Join (SPJ) avoids the shuffle phase entirely when Spark can use the partitioning already reported by a compatible V2 data source (`spark.sql.sources.v2.bucketing.enabled`, default `true` since 3.3.0), generalizing bucket joins to functions registered in a `FunctionCatalog`[^3]. Classic Hive-style bucketing gets the same effect: once both sides of a join are bucketed and sorted the same way, the physical plan shows no `Exchange` operator at all[^4]. `DataFrameWriter.partitionBy`, by contrast, does not trigger a shuffle by itself on write, though it can still leave you with a large number of small output files[^5].
|
|
8
|
+
|
|
9
|
+
Modern Spark ships a single shuffle manager, `SortShuffleManager` (3.5 source). Incoming records are sorted by their target partition id and written to one map output file per task; reducers then fetch contiguous regions of that file for their share of the map output, spilling sorted subsets to disk and merging them if the data doesn't fit in memory[^6]. Internally it has two write paths: a "serialized sorting" path used when there's no map-side combine, the serializer supports relocation of serialized values (Kryo or Spark SQL's custom serializers), and the shuffle produces at most 16,777,216 output partitions; and a "deserialized sorting" path used for everything else[^6]. Framed against the older hash-based approach (where each map task produces a separate output file per reduce task rather than one sorted file), sort-based shuffle pays a sorting cost but scales better once shuffle data gets large, which is why Spark and Flink both moved to it and pair it with an external shuffle service[^7].
|
|
10
|
+
|
|
11
|
+
<img class="light-only" src="diagrams/shuffle-map-reduce.svg" alt="Map tasks sorting records by target partition id into one output file each while reduce tasks fetch their partition blocks from every map output.">
|
|
12
|
+
<img class="dark-only" src="diagrams/shuffle-map-reduce.dark.svg" alt="Map tasks sorting records by target partition id into one output file each while reduce tasks fetch their partition blocks from every map output.">
|
|
13
|
+
|
|
14
|
+
Two pieces of supporting infrastructure sit alongside the core shuffle manager. The external shuffle service (`spark.shuffle.service.enabled`, default `false`) preserves the shuffle files written by executors so they can be safely removed, or so shuffle fetches continue even after an executor failure; enabling it requires standing up the service separately, since the flag alone isn't enough[^8]. It's a long-running process on each cluster node, independent of any particular Spark application or its executors: once enabled, executors fetch shuffle files from this service instead of from each other, so shuffle state written by an executor keeps being served after that executor's own lifetime ends[^9]. Push-based shuffle, designed under SPARK-30602 and based on LinkedIn's "Magnet" shuffle service, goes further: it changes the reduce side from a pull model (reducers requesting many small, randomly-ordered blocks from wherever each mapper wrote them) into a push-and-merge model, where mapper-generated blocks are pushed to remote Magnet shuffle services that opportunistically merge them into large per-partition chunks before the reduce stage starts[^7].
|
|
15
|
+
|
|
16
|
+
## Reading Exchange nodes
|
|
17
|
+
|
|
18
|
+
Whether an operation shuffles doesn't have to be guesswork: shuffles show up in the physical plan as `Exchange` nodes, and reading the plan is the most direct way to check. `groupBy()` on a DataFrame is a wide transformation and normally requires an `Exchange` so all rows for a key land on the same partition[^1]; that node disappears only when Spark already knows the data is co-partitioned by the grouping key. The clearest documented case is a table bucketed and sorted on the join/group key, where the physical plan drops the `Exchange` because the data was already shuffled at write time[^10]. SPJ produces the same absence of an `Exchange` for join queries against compatible V2 sources[^3].
|
|
19
|
+
|
|
20
|
+
`DataFrame.repartition(key)` is worth watching for separately: it's a one-time, explicit shuffle, but if you cache the result, subsequent joins on that same key skip their own shuffle because the data is already partitioned that way[^5]. Don't mistake `partitionBy` on write for the same thing: it changes the on-disk directory layout but is "not the equivalent" of an in-memory repartition, so `df.groupBy('key').sum()` over data written with `partitionBy` still requires an `Exchange`[^5].
|
|
21
|
+
|
|
22
|
+
Deduplication is a case where there's no avoiding the shuffle regardless of how you write it: Spark SQL's `dropDuplicates` selects unique rows the same way RDD-level `distinct` does, and both "can require a shuffle"[^11]. The only documented difference between them is key flexibility, not shuffle volume: unlike RDD `distinct`, DataFrame `dropDuplicates` can optionally drop rows based on only a subset of columns (e.g., `dropDuplicates(List("id"))`) rather than requiring uniqueness across the whole row[^11].
|
|
23
|
+
|
|
24
|
+
## Where the cost accumulates
|
|
25
|
+
|
|
26
|
+
Confirming a shuffle happened is the easy part; pricing it out is the harder one. Shuffle cost shows up in several places at once, and Spark's shuffle-tuning knobs largely exist to manage it. Inside `SortShuffleManager`, when there are fewer than `spark.shuffle.sort.bypassMergeThreshold` reduce partitions (default 200)[^8] and no map-side aggregation is needed, Spark takes a bypass path: it writes `numPartitions` files directly per map task and concatenates them at the end, rather than merge-sorting spilled files. This avoids serializing and deserializing the data twice, which the normal merge-sort path would otherwise pay; the trade-off is having multiple files open at once, and more memory allocated to buffers[^6]. That trade stays cheap only while the reduce-side partition count is small.
|
|
27
|
+
|
|
28
|
+
Reducers also pay a fixed memory cost for every map output fetched in parallel, since each needs a buffer to receive it; that's what `spark.reducer.maxSizeInFlight` (default 48m) caps[^8]. And map tasks pay an I/O cost writing shuffle files: a larger `spark.shuffle.file.buffer` (default 32k) reduces the number of disk seeks and system calls needed to create those intermediate files[^8].
|
|
29
|
+
|
|
30
|
+
External shuffle infrastructure matters most once [dynamic allocation](#cluster-config) is in play: without the external shuffle service, an executor decommissioned mid-shuffle (say, because of stragglers) would force its shuffle files to be recomputed from scratch once it's removed[^9]. Push-based shuffle exists because, as shuffle data volume grows, the number of blocks a classic shuffle produces grows quadratically (mappers × reducers) while individual block sizes shrink to only tens of KB, which is inefficient for HDD-backed random reads[^7]. Magnet's push-and-merge model improves disk I/O efficiency by converting many small random reads into large sequential reads of pre-merged chunks, and improves reducer data locality since merged output can be co-located with the reduce tasks that consume it; the push itself is decoupled from the mappers, so it doesn't add to map task runtime or fail map tasks if a push fails, and it's best-effort, so reducers can fetch a mix of merged and unmerged blocks[^7]. LinkedIn's own production rollout, on a cluster running 30K+ Spark applications shuffling roughly 5 PB/day, reported a large reduction in shuffle read wait time and a corresponding drop in total executor runtime after enabling it on complex production flows[^7].
|
|
31
|
+
|
|
32
|
+
## Tuning or avoiding the shuffle
|
|
33
|
+
|
|
34
|
+
The levers split into two groups: avoiding a shuffle, and tuning the one you can't avoid. Where possible, avoid the shuffle rather than tune it: bucket and sort tables on the join/group key so the physical plan drops the `Exchange` entirely[^10][^4], or lean on Storage Partition Join for compatible V2 sources[^3]. If you need to reuse a partitioning scheme across several joins, `repartition(key)` once and cache the result rather than relying on `partitionBy`, which only affects the on-disk layout and doesn't exempt later aggregations from their own `Exchange`[^5].
|
|
35
|
+
|
|
36
|
+
> **PySpark:** `dropDuplicates(subset)` still shuffles like `distinct()`, but lets you dedupe on a subset of columns instead of the whole row; use `df.dropDuplicates(["id"])` when uniqueness only needs to hold on a few key columns[^11].
|
|
37
|
+
|
|
38
|
+
For shuffles you can't avoid, the main tuning levers are:
|
|
39
|
+
|
|
40
|
+
- `spark.shuffle.sort.bypassMergeThreshold` (default 200): raise or lower the reduce-partition-count cutoff for the bypass (hash-style) write path described above[^8].
|
|
41
|
+
- `spark.shuffle.compress` (default `true`): compresses map output files using whatever codec `spark.io.compression.codec` names (default `lz4`; `lzf`, `snappy`, and `zstd` are also available)[^8]. `spark.shuffle.spill.compress` (also default `true`) separately governs compression of spilled shuffle data, using the same codec setting[^8].
|
|
42
|
+
- `spark.reducer.maxSizeInFlight` (default 48m): raising it lets more data move in flight at the cost of per-task memory; lowering it saves memory at the cost of more fetch rounds[^8].
|
|
43
|
+
- `spark.shuffle.file.buffer` (default 32k): Learning Spark's tuning guidance recommends bumping it to 1 MB for large jobs, trading some per-task memory for fewer disk seeks and syscalls during the shuffle-write phase[^1].
|
|
44
|
+
- `spark.shuffle.service.enabled` (default `false`): turn on the external shuffle service so executors can be safely removed under dynamic allocation without losing shuffle state. Dynamic allocation's documentation lists shuffle tracking (`spark.dynamicAllocation.shuffleTracking.enabled`), shuffle-block decommissioning, and a custom `ShuffleDataIO` plugin backed by reliable storage as alternatives to enabling the service outright[^8].
|
|
45
|
+
- Push-based shuffle: enable it with the paired server/client flags added in Spark 3.2.0: `spark.shuffle.push.server.mergedShuffleFileManagerImpl` on the server side, and `spark.shuffle.push.enabled=true` on the client side (both disabled by default; the client flag only takes effect together with the server-side one). It's currently only supported for Spark on YARN with the external shuffle service enabled[^8].
|
|
46
|
+
|
|
47
|
+
## Sources
|
|
48
|
+
|
|
49
|
+
[^1]: *Learning Spark, 2nd Edition*, Damji, Wenig, Das, Lee, ch. 7: Optimizing and Tuning Spark Applications
|
|
50
|
+
[^2]: [RDD Programming Guide](https://spark.apache.org/docs/latest/rdd-programming-guide.html)
|
|
51
|
+
[^3]: [Performance Tuning: Spark SQL, DataFrames and Datasets Guide](https://spark.apache.org/docs/latest/sql-performance-tuning.html)
|
|
52
|
+
[^4]: [Bucketing (The Internals of Spark SQL)](https://books.japila.pl/spark-sql-internals/bucketing/)
|
|
53
|
+
[^5]: [Spark Tips: Partition Tuning](https://luminousmen.com/post/spark-tips-partition-tuning)
|
|
54
|
+
[^6]: [SortShuffleManager.scala](https://raw.githubusercontent.com/apache/spark/v3.5.0/core/src/main/scala/org/apache/spark/shuffle/sort/SortShuffleManager.scala)
|
|
55
|
+
[^7]: [SPARK-30602: Support push-based shuffle to improve shuffle efficiency (Magnet)](https://issues.apache.org/jira/browse/SPARK-30602)
|
|
56
|
+
[^8]: [Configuration (Spark)](https://spark.apache.org/docs/latest/configuration.html)
|
|
57
|
+
[^9]: [Job Scheduling](https://spark.apache.org/docs/latest/job-scheduling.html)
|
|
58
|
+
[^10]: [Bucket the Shuffle Out of Here (Taboola Engineering)](https://www.taboola.com/engineering/bucket-the-shuffle-out-of-here/)
|
|
59
|
+
[^11]: *High Performance Spark, 2nd Edition*, Karau, Polak & Warren, ch. 5: DataFrames, Datasets, and Spark SQL
|