sparkforensics-cli 0.1.0 → 0.2.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +6 -0
- package/bin/sparkforensics-analyze.mjs +113 -48
- package/export-template/docs/404.html +25 -0
- package/export-template/docs/assets/app.DQTZyGL1.js +1 -0
- package/export-template/docs/assets/aqe-loop.IwQSATHw.svg +1 -0
- package/export-template/docs/assets/aqe-loop.dark.DGbaxqJE.svg +1 -0
- package/export-template/docs/assets/broadcast-vs-shuffle.Db4WY1XK.svg +1 -0
- package/export-template/docs/assets/broadcast-vs-shuffle.dark.C7Bxs0mG.svg +1 -0
- package/export-template/docs/assets/cache-lifecycle.dark.B-hS7AgU.svg +1 -0
- package/export-template/docs/assets/cache-lifecycle.rEOVYQNU.svg +1 -0
- package/export-template/docs/assets/chunks/@localSearchIndexroot.DppnXnDE.js +1 -0
- package/export-template/docs/assets/chunks/VPLocalSearchBox.BkBIPFs6.js +9 -0
- package/export-template/docs/assets/chunks/duplicate-plan-subtree.dark.Cdp70QhV.js +1 -0
- package/export-template/docs/assets/chunks/framework.DSg0KOwT.js +20 -0
- package/export-template/docs/assets/chunks/retry-escalation-ladder.dark.DHipdJgZ.js +1 -0
- package/export-template/docs/assets/chunks/theme.DP0u1AUq.js +2 -0
- package/export-template/docs/assets/cold-start-timeline.DxC_Sc7w.svg +1 -0
- package/export-template/docs/assets/cold-start-timeline.dark.CZ17YcAG.svg +1 -0
- package/export-template/docs/assets/columnar-layout.PghGeOEA.svg +1 -0
- package/export-template/docs/assets/columnar-layout.dark.BVNlz0ff.svg +1 -0
- package/export-template/docs/assets/container-memory.DIO0AnIm.svg +1 -0
- package/export-template/docs/assets/container-memory.dark.CP-5zuCl.svg +1 -0
- package/export-template/docs/assets/contributor-guide_architecture_board-widgets.md.CWpj01WU.js +1 -0
- package/export-template/docs/assets/contributor-guide_architecture_board-widgets.md.CWpj01WU.lean.js +1 -0
- package/export-template/docs/assets/contributor-guide_architecture_detector-contract.md.CgzUsQ6W.js +1 -0
- package/export-template/docs/assets/contributor-guide_architecture_detector-contract.md.CgzUsQ6W.lean.js +1 -0
- package/export-template/docs/assets/contributor-guide_architecture_drill-down.md.BtPdlM7r.js +1 -0
- package/export-template/docs/assets/contributor-guide_architecture_drill-down.md.BtPdlM7r.lean.js +1 -0
- package/export-template/docs/assets/contributor-guide_architecture_impact-estimation.md.CooslVJt.js +1 -0
- package/export-template/docs/assets/contributor-guide_architecture_impact-estimation.md.CooslVJt.lean.js +1 -0
- package/export-template/docs/assets/contributor-guide_architecture_index.md.3TO9ic6w.js +1 -0
- package/export-template/docs/assets/contributor-guide_architecture_index.md.3TO9ic6w.lean.js +1 -0
- package/export-template/docs/assets/contributor-guide_architecture_overview.md.CehiRmGn.js +1 -0
- package/export-template/docs/assets/contributor-guide_architecture_overview.md.CehiRmGn.lean.js +1 -0
- package/export-template/docs/assets/contributor-guide_architecture_state-and-history.md.C-xxn0q7.js +1 -0
- package/export-template/docs/assets/contributor-guide_architecture_state-and-history.md.C-xxn0q7.lean.js +1 -0
- package/export-template/docs/assets/contributor-guide_architecture_widget-rendering.md.R27gQrgY.js +1 -0
- package/export-template/docs/assets/contributor-guide_architecture_widget-rendering.md.R27gQrgY.lean.js +1 -0
- package/export-template/docs/assets/contributor-guide_architecture_worker-protocol.md.IbnfNrV3.js +6 -0
- package/export-template/docs/assets/contributor-guide_architecture_worker-protocol.md.IbnfNrV3.lean.js +1 -0
- package/export-template/docs/assets/contributor-guide_contributing.md.CvRsdr6J.js +1 -0
- package/export-template/docs/assets/contributor-guide_contributing.md.CvRsdr6J.lean.js +1 -0
- package/export-template/docs/assets/contributor-guide_development-setup.md.DvAN_9mK.js +12 -0
- package/export-template/docs/assets/contributor-guide_development-setup.md.DvAN_9mK.lean.js +1 -0
- package/export-template/docs/assets/contributor-guide_testing.md.6rIKqSyY.js +1 -0
- package/export-template/docs/assets/contributor-guide_testing.md.6rIKqSyY.lean.js +1 -0
- package/export-template/docs/assets/dag-stages.DSz_S937.svg +1 -0
- package/export-template/docs/assets/dag-stages.dark.F72UzxH4.svg +1 -0
- package/export-template/docs/assets/driver-executor.D5pQ7YN1.svg +1 -0
- package/export-template/docs/assets/driver-executor.dark.BmX9cPvh.svg +1 -0
- package/export-template/docs/assets/duplicate-plan-subtree.B4cvN6fj.svg +1 -0
- package/export-template/docs/assets/duplicate-plan-subtree.dark.Dw8wS0Ag.svg +1 -0
- package/export-template/docs/assets/index.md.CHJVslga.js +1 -0
- package/export-template/docs/assets/index.md.CHJVslga.lean.js +1 -0
- package/export-template/docs/assets/inter-italic-cyrillic-ext.r48I6akx.woff2 +0 -0
- package/export-template/docs/assets/inter-italic-cyrillic.By2_1cv3.woff2 +0 -0
- package/export-template/docs/assets/inter-italic-greek-ext.1u6EdAuj.woff2 +0 -0
- package/export-template/docs/assets/inter-italic-greek.DJ8dCoTZ.woff2 +0 -0
- package/export-template/docs/assets/inter-italic-latin-ext.CN1xVJS-.woff2 +0 -0
- package/export-template/docs/assets/inter-italic-latin.C2AdPX0b.woff2 +0 -0
- package/export-template/docs/assets/inter-italic-vietnamese.BSbpV94h.woff2 +0 -0
- package/export-template/docs/assets/inter-roman-cyrillic-ext.BBPuwvHQ.woff2 +0 -0
- package/export-template/docs/assets/inter-roman-cyrillic.C5lxZ8CY.woff2 +0 -0
- package/export-template/docs/assets/inter-roman-greek-ext.CqjqNYQ-.woff2 +0 -0
- package/export-template/docs/assets/inter-roman-greek.BBVDIX6e.woff2 +0 -0
- package/export-template/docs/assets/inter-roman-latin-ext.4ZJIpNVo.woff2 +0 -0
- package/export-template/docs/assets/inter-roman-latin.Di8DUHzh.woff2 +0 -0
- package/export-template/docs/assets/inter-roman-vietnamese.BjW4sHH5.woff2 +0 -0
- package/export-template/docs/assets/join-strategy.C_FvrCEo.svg +1 -0
- package/export-template/docs/assets/join-strategy.dark.ChMLnNII.svg +1 -0
- package/export-template/docs/assets/memory-borrowing.BqQRJg0u.svg +1 -0
- package/export-template/docs/assets/memory-borrowing.dark.Yhh20O9C.svg +1 -0
- package/export-template/docs/assets/memory-regions.XHvO7jHG.svg +1 -0
- package/export-template/docs/assets/memory-regions.dark.D4TP9_08.svg +1 -0
- package/export-template/docs/assets/repartition-vs-coalesce.BovLRrpj.svg +1 -0
- package/export-template/docs/assets/repartition-vs-coalesce.dark.BhAczKZQ.svg +1 -0
- package/export-template/docs/assets/retry-escalation-ladder.DyTKJJmZ.svg +1 -0
- package/export-template/docs/assets/retry-escalation-ladder.dark.BdsabtU3.svg +1 -0
- package/export-template/docs/assets/shuffle-map-reduce.KuOEZVmg.svg +1 -0
- package/export-template/docs/assets/shuffle-map-reduce.dark.BgQZnFSb.svg +1 -0
- package/export-template/docs/assets/spill-classification.BU2euYDO.svg +1 -0
- package/export-template/docs/assets/spill-classification.dark.D7i1M40d.svg +1 -0
- package/export-template/docs/assets/style.DSixAiZE.css +1 -0
- package/export-template/docs/assets/tuning-reference_anti-patterns.md.Df1YMIHu.js +1 -0
- package/export-template/docs/assets/tuning-reference_anti-patterns.md.Df1YMIHu.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_aqe.md.BIsCtLzm.js +1 -0
- package/export-template/docs/assets/tuning-reference_aqe.md.BIsCtLzm.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-broadcast-sizing.md.CEstB3Ia.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-broadcast-sizing.md.CEstB3Ia.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-cold-start.md.CEuy-72y.js +7 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-cold-start.md.CEuy-72y.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-duplicate-plan-subtree.md.CIohQDfn.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-duplicate-plan-subtree.md.CIohQDfn.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-failures.md.4z5BXGJ2.js +6 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-failures.md.4z5BXGJ2.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-gc.md.DSxzZRK7.js +6 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-gc.md.DSxzZRK7.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-job-failure-rate.md.BaJl__1W.js +8 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-job-failure-rate.md.BaJl__1W.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-memory-utilization.md.DbP-SJZc.js +7 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-memory-utilization.md.DbP-SJZc.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-retry-waste.md.D5JMjVOt.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-retry-waste.md.D5JMjVOt.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-shuffle.md.CM-nTmIH.js +12 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-shuffle.md.CM-nTmIH.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-skew.md.BdUwiDhn.js +14 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-skew.md.BdUwiDhn.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-slow-host.md.BlIo6UDW.js +7 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-slow-host.md.BlIo6UDW.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-small-files.md.B8kloyx8.js +5 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-small-files.md.B8kloyx8.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-spill.md.PNH7mITt.js +6 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-spill.md.PNH7mITt.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-straggler.md.DY36fHN5.js +7 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-straggler.md.DY36fHN5.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-tiny-tasks.md.QTV7O8kU.js +8 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-tiny-tasks.md.QTV7O8kU.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-utilization.md.DTiueZC3.js +5 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-utilization.md.DTiueZC3.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_caching.md.B7aQ8asB.js +1 -0
- package/export-template/docs/assets/tuning-reference_caching.md.B7aQ8asB.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_cluster-config.md.ZVmDGsQ3.js +1 -0
- package/export-template/docs/assets/tuning-reference_cluster-config.md.ZVmDGsQ3.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_config.md.UvveiWG3.js +1 -0
- package/export-template/docs/assets/tuning-reference_config.md.UvveiWG3.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_data-formats.md.bjCAWH3N.js +1 -0
- package/export-template/docs/assets/tuning-reference_data-formats.md.bjCAWH3N.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_index.md.BQ_NooMV.js +1 -0
- package/export-template/docs/assets/tuning-reference_index.md.BQ_NooMV.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_intro.md.CobD-lGB.js +1 -0
- package/export-template/docs/assets/tuning-reference_intro.md.CobD-lGB.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_joins.md.BtKs_CuW.js +1 -0
- package/export-template/docs/assets/tuning-reference_joins.md.BtKs_CuW.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_memory-model.md.DhT-n4y3.js +1 -0
- package/export-template/docs/assets/tuning-reference_memory-model.md.DhT-n4y3.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_metrics.md.mLOh7Apj.js +1 -0
- package/export-template/docs/assets/tuning-reference_metrics.md.mLOh7Apj.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_partitioning.md.q0zKF_8X.js +1 -0
- package/export-template/docs/assets/tuning-reference_partitioning.md.q0zKF_8X.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_pyspark.md.DDCfvN9t.js +6 -0
- package/export-template/docs/assets/tuning-reference_pyspark.md.DDCfvN9t.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_shuffle.md.BZZ7R4Ix.js +1 -0
- package/export-template/docs/assets/tuning-reference_shuffle.md.BZZ7R4Ix.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_spark-architecture.md.Dwzm5avO.js +1 -0
- package/export-template/docs/assets/tuning-reference_spark-architecture.md.Dwzm5avO.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_table-formats.md.D6wj-2dX.js +1 -0
- package/export-template/docs/assets/tuning-reference_table-formats.md.D6wj-2dX.lean.js +1 -0
- package/export-template/docs/assets/udf-execution-models.BUFDICuG.svg +1 -0
- package/export-template/docs/assets/udf-execution-models.dark.YTNS6GDq.svg +1 -0
- package/export-template/docs/assets/user-guide_alternative-log-retrieval.md.B4tPGIal.js +1 -0
- package/export-template/docs/assets/user-guide_alternative-log-retrieval.md.B4tPGIal.lean.js +1 -0
- package/export-template/docs/assets/user-guide_getting-started.md.BJvwLEIM.js +3 -0
- package/export-template/docs/assets/user-guide_getting-started.md.BJvwLEIM.lean.js +1 -0
- package/export-template/docs/assets/user-guide_mcp-tools.md.Vi3RoflJ.js +125 -0
- package/export-template/docs/assets/user-guide_mcp-tools.md.Vi3RoflJ.lean.js +1 -0
- package/export-template/docs/assets/user-guide_run-comparison.md.CQc1aoU8.js +1 -0
- package/export-template/docs/assets/user-guide_run-comparison.md.CQc1aoU8.lean.js +1 -0
- package/export-template/docs/assets/user-guide_understanding-findings.md.DL1UDhvR.js +1 -0
- package/export-template/docs/assets/user-guide_understanding-findings.md.DL1UDhvR.lean.js +1 -0
- package/export-template/docs/contributor-guide/architecture/board-widgets.html +25 -0
- package/export-template/docs/contributor-guide/architecture/detector-contract.html +25 -0
- package/export-template/docs/contributor-guide/architecture/drill-down.html +25 -0
- package/export-template/docs/contributor-guide/architecture/impact-estimation.html +25 -0
- package/export-template/docs/contributor-guide/architecture/index.html +25 -0
- package/export-template/docs/contributor-guide/architecture/overview.html +25 -0
- package/export-template/docs/contributor-guide/architecture/state-and-history.html +25 -0
- package/export-template/docs/contributor-guide/architecture/widget-rendering.html +25 -0
- package/export-template/docs/contributor-guide/architecture/worker-protocol.html +30 -0
- package/export-template/docs/contributor-guide/contributing.html +25 -0
- package/export-template/docs/contributor-guide/development-setup.html +36 -0
- package/export-template/docs/contributor-guide/testing.html +25 -0
- package/export-template/docs/favicon.svg +4 -0
- package/export-template/docs/hashmap.json +1 -0
- package/export-template/docs/index.html +25 -0
- package/export-template/docs/package.json +1 -0
- package/export-template/docs/tuning-reference/anti-patterns.html +25 -0
- package/export-template/docs/tuning-reference/aqe.html +25 -0
- package/export-template/docs/tuning-reference/bottleneck-broadcast-sizing.html +25 -0
- package/export-template/docs/tuning-reference/bottleneck-cold-start.html +31 -0
- package/export-template/docs/tuning-reference/bottleneck-duplicate-plan-subtree.html +25 -0
- package/export-template/docs/tuning-reference/bottleneck-failures.html +30 -0
- package/export-template/docs/tuning-reference/bottleneck-gc.html +30 -0
- package/export-template/docs/tuning-reference/bottleneck-job-failure-rate.html +32 -0
- package/export-template/docs/tuning-reference/bottleneck-memory-utilization.html +31 -0
- package/export-template/docs/tuning-reference/bottleneck-retry-waste.html +25 -0
- package/export-template/docs/tuning-reference/bottleneck-shuffle.html +36 -0
- package/export-template/docs/tuning-reference/bottleneck-skew.html +38 -0
- package/export-template/docs/tuning-reference/bottleneck-slow-host.html +31 -0
- package/export-template/docs/tuning-reference/bottleneck-small-files.html +29 -0
- package/export-template/docs/tuning-reference/bottleneck-spill.html +30 -0
- package/export-template/docs/tuning-reference/bottleneck-straggler.html +31 -0
- package/export-template/docs/tuning-reference/bottleneck-tiny-tasks.html +32 -0
- package/export-template/docs/tuning-reference/bottleneck-utilization.html +29 -0
- package/export-template/docs/tuning-reference/caching.html +25 -0
- package/export-template/docs/tuning-reference/cluster-config.html +25 -0
- package/export-template/docs/tuning-reference/config.html +25 -0
- package/export-template/docs/tuning-reference/data-formats.html +25 -0
- package/export-template/docs/tuning-reference/index.html +25 -0
- package/export-template/docs/tuning-reference/intro.html +25 -0
- package/export-template/docs/tuning-reference/joins.html +25 -0
- package/export-template/docs/tuning-reference/memory-model.html +25 -0
- package/export-template/docs/tuning-reference/metrics.html +25 -0
- package/export-template/docs/tuning-reference/partitioning.html +25 -0
- package/export-template/docs/tuning-reference/pyspark.html +30 -0
- package/export-template/docs/tuning-reference/shuffle.html +25 -0
- package/export-template/docs/tuning-reference/spark-architecture.html +25 -0
- package/export-template/docs/tuning-reference/table-formats.html +25 -0
- package/export-template/docs/user-guide/alternative-log-retrieval.html +25 -0
- package/export-template/docs/user-guide/getting-started.html +27 -0
- package/export-template/docs/user-guide/mcp-tools.html +149 -0
- package/export-template/docs/user-guide/run-comparison.html +25 -0
- package/export-template/docs/user-guide/understanding-findings.html +25 -0
- package/export-template/docs/vp-icons.css +0 -0
- package/export-template/favicon.svg +4 -0
- package/export-template/index.html +111 -0
- package/export-template/parser-worker-DyjiQvfP.js +112 -0
- package/export-template/sample-runs/sample-run.ndjson.gz +0 -0
- package/package.json +20 -6
- package/vendor-core/analyzer.js +74 -74
- package/vendor-core/cli/budgets.js +13 -27
- package/vendor-core/cli/collect-run.js +43 -19
- package/vendor-core/core-count.js +25 -27
- package/vendor-core/core-locality-ratio.js +4 -11
- package/vendor-core/core-time-series.js +6 -12
- package/vendor-core/core-usage-locality.js +3 -4
- package/vendor-core/detectors.js +395 -389
- package/vendor-core/docs-config.js +69 -21
- package/vendor-core/docs-content/chapters/01-intro.md +32 -0
- package/vendor-core/docs-content/chapters/02-spark-architecture.md +76 -0
- package/vendor-core/docs-content/chapters/03-memory-model.md +73 -0
- package/vendor-core/docs-content/chapters/04-partitioning.md +65 -0
- package/vendor-core/docs-content/chapters/05-joins.md +62 -0
- package/vendor-core/docs-content/chapters/06-shuffle.md +59 -0
- package/vendor-core/docs-content/chapters/07-data-formats.md +81 -0
- package/vendor-core/docs-content/chapters/07b-table-formats.md +56 -0
- package/vendor-core/docs-content/chapters/08-caching.md +58 -0
- package/vendor-core/docs-content/chapters/09-pyspark.md +78 -0
- package/vendor-core/docs-content/chapters/10-aqe.md +167 -0
- package/vendor-core/docs-content/chapters/11-cluster-config.md +170 -0
- package/vendor-core/docs-content/chapters/12-anti-patterns.md +171 -0
- package/vendor-core/docs-content/chapters/14-metrics.md +87 -0
- package/vendor-core/docs-content/chapters/15-config.md +93 -0
- package/vendor-core/docs-content/chapters/nav-index.json +370 -0
- package/vendor-core/docs-content/detection/cache.md +7 -0
- package/vendor-core/docs-content/detection/cfg.md +15 -0
- package/vendor-core/docs-content/detection/chrn.md +9 -0
- package/vendor-core/docs-content/detection/cold.md +4 -0
- package/vendor-core/docs-content/detection/cstor.md +4 -0
- package/vendor-core/docs-content/detection/fail.md +5 -0
- package/vendor-core/docs-content/detection/gc.md +4 -0
- package/vendor-core/docs-content/detection/host.md +5 -0
- package/vendor-core/docs-content/detection/incmp.md +6 -0
- package/vendor-core/docs-content/detection/jobs.md +4 -0
- package/vendor-core/docs-content/detection/local.md +6 -0
- package/vendor-core/docs-content/detection/mem.md +10 -0
- package/vendor-core/docs-content/detection/part.md +5 -0
- package/vendor-core/docs-content/detection/plan.md +14 -0
- package/vendor-core/docs-content/detection/retry.md +4 -0
- package/vendor-core/docs-content/detection/sfail.md +5 -0
- package/vendor-core/docs-content/detection/shape.md +5 -0
- package/vendor-core/docs-content/detection/shfl.md +4 -0
- package/vendor-core/docs-content/detection/skew.md +6 -0
- package/vendor-core/docs-content/detection/slow.md +6 -0
- package/vendor-core/docs-content/detection/spec.md +8 -0
- package/vendor-core/docs-content/detection/spill.md +7 -0
- package/vendor-core/docs-content/detection/strag.md +5 -0
- package/vendor-core/docs-content/detection/tiny.md +4 -0
- package/vendor-core/docs-content/detection/util.md +4 -0
- package/vendor-core/docs-content/diagrams/aqe-loop.dark.svg +1 -0
- package/vendor-core/docs-content/diagrams/aqe-loop.svg +1 -0
- package/vendor-core/docs-content/diagrams/broadcast-vs-shuffle.dark.svg +1 -0
- package/vendor-core/docs-content/diagrams/broadcast-vs-shuffle.svg +1 -0
- package/vendor-core/docs-content/diagrams/cache-lifecycle.dark.svg +1 -0
- package/vendor-core/docs-content/diagrams/cache-lifecycle.svg +1 -0
- package/vendor-core/docs-content/diagrams/cold-start-timeline.dark.svg +1 -0
- package/vendor-core/docs-content/diagrams/cold-start-timeline.svg +1 -0
- package/vendor-core/docs-content/diagrams/columnar-layout.dark.svg +1 -0
- package/vendor-core/docs-content/diagrams/columnar-layout.svg +1 -0
- package/vendor-core/docs-content/diagrams/container-memory.dark.svg +1 -0
- package/vendor-core/docs-content/diagrams/container-memory.svg +1 -0
- package/vendor-core/docs-content/diagrams/dag-stages.dark.svg +1 -0
- package/vendor-core/docs-content/diagrams/dag-stages.svg +1 -0
- package/vendor-core/docs-content/diagrams/driver-executor.dark.svg +1 -0
- package/vendor-core/docs-content/diagrams/driver-executor.svg +1 -0
- package/vendor-core/docs-content/diagrams/duplicate-plan-subtree.dark.svg +1 -0
- package/vendor-core/docs-content/diagrams/duplicate-plan-subtree.svg +1 -0
- package/vendor-core/docs-content/diagrams/join-strategy.dark.svg +1 -0
- package/vendor-core/docs-content/diagrams/join-strategy.svg +1 -0
- package/vendor-core/docs-content/diagrams/memory-borrowing.dark.svg +1 -0
- package/vendor-core/docs-content/diagrams/memory-borrowing.svg +1 -0
- package/vendor-core/docs-content/diagrams/memory-regions.dark.svg +1 -0
- package/vendor-core/docs-content/diagrams/memory-regions.svg +1 -0
- package/vendor-core/docs-content/diagrams/repartition-vs-coalesce.dark.svg +1 -0
- package/vendor-core/docs-content/diagrams/repartition-vs-coalesce.svg +1 -0
- package/vendor-core/docs-content/diagrams/retry-escalation-ladder.dark.svg +1 -0
- package/vendor-core/docs-content/diagrams/retry-escalation-ladder.svg +1 -0
- package/vendor-core/docs-content/diagrams/shuffle-map-reduce.dark.svg +1 -0
- package/vendor-core/docs-content/diagrams/shuffle-map-reduce.svg +1 -0
- package/vendor-core/docs-content/diagrams/spill-classification.dark.svg +1 -0
- package/vendor-core/docs-content/diagrams/spill-classification.svg +1 -0
- package/vendor-core/docs-content/diagrams/udf-execution-models.dark.svg +1 -0
- package/vendor-core/docs-content/diagrams/udf-execution-models.svg +1 -0
- package/vendor-core/docs-content/tuning/broadcast-sizing.md +78 -0
- package/vendor-core/docs-content/tuning/cold-start.md +81 -0
- package/vendor-core/docs-content/tuning/duplicate-plan-subtree.md +45 -0
- package/vendor-core/docs-content/tuning/failures.md +124 -0
- package/vendor-core/docs-content/tuning/gc.md +110 -0
- package/vendor-core/docs-content/tuning/job-failure-rate.md +101 -0
- package/vendor-core/docs-content/tuning/memory-utilization.md +58 -0
- package/vendor-core/docs-content/tuning/retry-waste.md +90 -0
- package/vendor-core/docs-content/tuning/shuffle.md +154 -0
- package/vendor-core/docs-content/tuning/skew.md +123 -0
- package/vendor-core/docs-content/tuning/slow-host.md +117 -0
- package/vendor-core/docs-content/tuning/small-files.md +99 -0
- package/vendor-core/docs-content/tuning/spill.md +114 -0
- package/vendor-core/docs-content/tuning/straggler.md +103 -0
- package/vendor-core/docs-content/tuning/tiny-tasks.md +94 -0
- package/vendor-core/docs-content/tuning/utilization.md +90 -0
- package/vendor-core/docs-site-config.js +10 -17
- package/vendor-core/efficiency-model.js +7 -13
- package/vendor-core/etl-phases.js +3 -5
- package/vendor-core/event-handlers.js +232 -134
- package/vendor-core/event-schemas.js +48 -114
- package/vendor-core/evidence-availability.js +5 -10
- package/vendor-core/evidence-report.js +73 -123
- package/vendor-core/export-data.js +48 -0
- package/vendor-core/finding-action-label.js +4 -10
- package/vendor-core/finding-filter-predicate.js +3 -7
- package/vendor-core/finding-generic-recommendation.js +112 -0
- package/vendor-core/finding-names.js +51 -0
- package/vendor-core/format-utils.js +112 -38
- package/vendor-core/impact-band.js +18 -24
- package/vendor-core/impact-estimator.js +38 -74
- package/vendor-core/ingest.js +7 -13
- package/vendor-core/job-groups.js +3 -6
- package/vendor-core/list-runs.js +278 -0
- package/vendor-core/load-vendored.js +6 -12
- package/vendor-core/log-header-peek.js +81 -0
- package/vendor-core/lz4-block.js +4 -6
- package/vendor-core/mcp-server-factory.js +38 -8
- package/vendor-core/mcp-tools.js +105 -76
- package/vendor-core/model-assembler.js +8 -16
- package/vendor-core/occupancy.js +5 -9
- package/vendor-core/parser-worker.js +20 -29
- package/vendor-core/plan-dot.js +2 -5
- package/vendor-core/plan-duration-attribution.js +78 -29
- package/vendor-core/plan-graph-model.js +126 -69
- package/vendor-core/plan-node-detail.js +31 -17
- package/vendor-core/plan-summary.js +19 -8
- package/vendor-core/recommendation-rollup.js +35 -39
- package/vendor-core/redact.js +72 -16
- package/vendor-core/rolling-log-reassembly.js +4 -6
- package/vendor-core/run-comparison.js +86 -72
- package/vendor-core/scaling-sim.js +5 -7
- package/vendor-core/session-snapshot.js +1 -1
- package/vendor-core/shs-fetch.js +4 -6
- package/vendor-core/shs-load.js +9 -13
- package/vendor-core/shs-request.js +1 -1
- package/vendor-core/stage-quantiles.js +14 -0
- package/vendor-core/types.js +78 -18
- package/vendor-core/wasted-core-hours.js +7 -12
|
@@ -0,0 +1,170 @@
|
|
|
1
|
+
# Cluster Tuning
|
|
2
|
+
|
|
3
|
+
## Sizing executors and containers
|
|
4
|
+
|
|
5
|
+
Cluster tuning is the set of decisions that map an application's cores, memory, and
|
|
6
|
+
executor count onto the underlying cluster (YARN or Kubernetes) plus a handful of
|
|
7
|
+
related knobs (data locality, dynamic allocation, per-task CPU reservation) that all
|
|
8
|
+
interact with that sizing decision.
|
|
9
|
+
|
|
10
|
+
The starting point is executor sizing. Working through a concrete example (a 10-node
|
|
11
|
+
cluster with 16 cores and 64GB RAM per node), the derivation runs: assign 5 cores per
|
|
12
|
+
executor; reserve about 1 core per node for Hadoop/YARN/OS daemons, leaving 15 usable
|
|
13
|
+
cores per node (150 total); dividing by 5 cores per executor gives 30 executors, minus
|
|
14
|
+
1 reserved for the YARN ApplicationMaster (29); with 3 executors per node, memory per
|
|
15
|
+
executor is 64GB / 3 ≈ 21GB, and after subtracting about 7% for YARN memory overhead
|
|
16
|
+
that comes out to roughly 18GB usable, a recommended shape of 29 executors × 5 cores ×
|
|
17
|
+
18GB[^1]. Cloudera's own worked example, on a 6-node cluster with the same 16-core /
|
|
18
|
+
64GB-per-node hardware, lands on the same shape: `--num-executors 17 --executor-cores 5
|
|
19
|
+
--executor-memory 19G`, rather than one "fat" executor per node
|
|
20
|
+
(`--num-executors 6 --executor-cores 15 --executor-memory 63G`)[^2]. Both sources treat
|
|
21
|
+
the resulting task count (executor-cores × num-executors) as the single most
|
|
22
|
+
important tuning lever, since Spark cannot compensate for too little parallelism on its
|
|
23
|
+
own[^2]. `spark.executor.cores` is the config that controls concurrent tasks per
|
|
24
|
+
executor; it defaults to 1 in YARN mode, or to all available cores on the worker in
|
|
25
|
+
standalone mode[^3], and `spark.executor.memory` sets the JVM heap size per
|
|
26
|
+
executor[^2]. Independent of the sizing formula above, the Spark tuning guide recommends
|
|
27
|
+
targeting 2–3 tasks per CPU core across the cluster[^4].
|
|
28
|
+
|
|
29
|
+
On top of executor memory sits container memory. Spark computes the total
|
|
30
|
+
container/pod allocation as `executorMemoryMiB + memoryOverheadMiB + memoryOffHeapMiB +
|
|
31
|
+
pysparkMemToUseMiB`. On YARN this is what gets allocated for the container, on
|
|
32
|
+
Kubernetes it becomes the pod memory limit[^5]. `spark.executor.memoryOverhead` itself
|
|
33
|
+
defaults to `max(0.1 * executorMemory, 384MB)`[^5][^3].
|
|
34
|
+
|
|
35
|
+
<img class="light-only" src="diagrams/container-memory.svg" alt="Container memory sums executor heap, memory overhead, off-heap size, and pyspark memory into the requested container size; the resource manager grants a container of that size and OOMKills the executor when its actual runtime footprint spills past that limit.">
|
|
36
|
+
<img class="dark-only" src="diagrams/container-memory.dark.svg" alt="Container memory sums executor heap, memory overhead, off-heap size, and pyspark memory into the requested container size; the resource manager grants a container of that size and OOMKills the executor when its actual runtime footprint spills past that limit.">
|
|
37
|
+
|
|
38
|
+
> **PySpark:** `spark.executor.pyspark.memory` is only added as its own
|
|
39
|
+
> container-memory term when it's set explicitly; otherwise PySpark's memory use is
|
|
40
|
+
> folded into the general overhead budget rather than tracked separately.[^5]
|
|
41
|
+
|
|
42
|
+
On Kubernetes specifically, `spark.kubernetes.executor.request.cores` and
|
|
43
|
+
`spark.kubernetes.executor.limit.cores` take priority over `spark.executor.cores` for
|
|
44
|
+
the pod's CPU request/limit sent to the Kubernetes scheduler[^6], while
|
|
45
|
+
`spark.executor.cores` remains the config Spark itself uses to size the number of
|
|
46
|
+
concurrent task slots[^3].
|
|
47
|
+
|
|
48
|
+
Two more knobs round out the picture. `spark.locality.wait` (default `3s`) controls how
|
|
49
|
+
long a task waits for a data-local placement before Spark gives up and schedules it less
|
|
50
|
+
locally, stepping through the same wait across process-local → node-local → rack-local →
|
|
51
|
+
any; each level can be overridden independently via `spark.locality.wait.process`,
|
|
52
|
+
`.node`, and `.rack`, all of which default to the base wait when left unset[^3]. Dynamic
|
|
53
|
+
allocation lets Spark add and remove executors as work changes, but it requires one of
|
|
54
|
+
several supporting mechanisms: an external shuffle service, shuffle tracking
|
|
55
|
+
(`spark.dynamicAllocation.shuffleTracking.enabled`, default `true` since Spark 3.0),
|
|
56
|
+
shuffle-block decommission, or the experimental sort-IO plugin[^3]. Removing an
|
|
57
|
+
executor can otherwise destroy shuffle state it's holding. Finally, `spark.task.cpus`
|
|
58
|
+
(default `1`) sets how many cores each task reserves; the number of concurrent task
|
|
59
|
+
slots per executor is derived jointly from it and `spark.executor.cores`, roughly
|
|
60
|
+
`spark.executor.cores / spark.task.cpus`[^3].
|
|
61
|
+
|
|
62
|
+
## Spotting a bad sizing decision
|
|
63
|
+
|
|
64
|
+
Getting that sizing wrong shows up in a handful of specific symptoms. The clearest sign of an under-sized executor layout is HDFS throughput dropping under
|
|
65
|
+
load: "HDFS client has trouble with tons of concurrent threads. It was observed that
|
|
66
|
+
HDFS achieves full write throughput with ~5 tasks per executor"[^1]. The "fat executor"
|
|
67
|
+
case (one executor per node, using all 16 cores) shows the failure mode directly:
|
|
68
|
+
"with all 16 cores per executor... HDFS throughput will hurt and it'll result in
|
|
69
|
+
excessive garbage [collection]"[^1].
|
|
70
|
+
|
|
71
|
+
On the memory side, an executor or pod that dies with no Spark-level error is a strong
|
|
72
|
+
signal that off-heap memory was never added into the container budget. Enabling
|
|
73
|
+
`spark.memory.offHeap.enabled=true` with `spark.memory.offHeap.size=1g` on top of an 8G
|
|
74
|
+
executor with default overhead (819MB) means real usage is 8192+819+1024 = 10,035MB,
|
|
75
|
+
while the container was only granted 8192+819 = 9,011MB. The executor gets killed by
|
|
76
|
+
YARN, or OOMKilled by the Kubernetes kubelet, with no warning from Spark itself[^5].
|
|
77
|
+
|
|
78
|
+
With dynamic allocation on, unexplained shuffle recomputation is a detectable symptom
|
|
79
|
+
of executors being reclaimed mid-shuffle: "In the event of stragglers... dynamic
|
|
80
|
+
allocation may remove an executor before the shuffle completes, in which case the
|
|
81
|
+
shuffle files written by that executor must be recomputed unnecessarily"[^6]. Jobs that
|
|
82
|
+
abort with a serialized-result-size error are hitting the `spark.driver.maxResultSize`
|
|
83
|
+
guardrail rather than a driver heap exhaustion: "Jobs will be aborted if the total size
|
|
84
|
+
[of serialized action results] is above this limit"[^3]. And on the resource-waste side,
|
|
85
|
+
executors that are provisioned but never assigned work are a sign that dynamic
|
|
86
|
+
allocation is targeting full parallelism against a workload made of many small tasks[^3].
|
|
87
|
+
|
|
88
|
+
## What a bad sizing decision costs
|
|
89
|
+
|
|
90
|
+
Each symptom above carries a specific price tag. Task count (`executor-cores × num-executors`) is treated by both cited sources as the
|
|
91
|
+
single most important tuning lever, since Spark cannot compensate for too little
|
|
92
|
+
parallelism on its own[^2]. Going the other way, cramming too many cores into one
|
|
93
|
+
executor degrades HDFS throughput and drives up garbage collection[^1].
|
|
94
|
+
|
|
95
|
+
Container-memory misconfiguration matters because the failure is silent from Spark's
|
|
96
|
+
point of view: off-heap memory that isn't accounted for in `spark.executor.memoryOverhead`
|
|
97
|
+
causes the executor to actually use more memory than the container/pod was granted, so
|
|
98
|
+
it gets killed externally (by YARN or the kubelet) with no Spark-level diagnostic to
|
|
99
|
+
point at the real cause[^5].
|
|
100
|
+
|
|
101
|
+
Dynamic allocation's interaction with shuffle state matters because, before dynamic
|
|
102
|
+
allocation existed, an executor exiting alongside its application meant all its state
|
|
103
|
+
could be safely discarded; with dynamic allocation, the application keeps running after
|
|
104
|
+
an executor is explicitly removed, so any later need for that executor's state forces a
|
|
105
|
+
recompute[^6]. That is exactly why Spark needs "a mechanism to decommission an executor
|
|
106
|
+
gracefully by preserving its state before removing it"[^6]. Without shuffle tracking,
|
|
107
|
+
an external shuffle service, or shuffle-block decommission enabled, dynamic allocation
|
|
108
|
+
either can't be turned on at all, or, if it's active regardless, an executor holding
|
|
109
|
+
unpreserved shuffle output that gets removed forces exactly that unnecessary
|
|
110
|
+
recompute[^3][^6].
|
|
111
|
+
|
|
112
|
+
On the driver side, `spark.driver.maxResultSize` and `spark.driver.memory` are separate
|
|
113
|
+
budgets that still interact: whether a high `maxResultSize` actually causes an
|
|
114
|
+
out-of-memory error "depends on spark.driver.memory and memory overhead of objects in
|
|
115
|
+
JVM"[^3]. Raising one without considering the other doesn't fully protect the driver.
|
|
116
|
+
|
|
117
|
+
Finally, over-provisioning executors against a small-task workload wastes cluster
|
|
118
|
+
resources: "with small tasks this setting can waste a lot of resources due to executor
|
|
119
|
+
allocation overhead, as some executor might not even do any work"[^3].
|
|
120
|
+
|
|
121
|
+
## Sizing the cluster correctly
|
|
122
|
+
|
|
123
|
+
Avoiding those costs starts from a formula, not one fat executor per node. Apply the balanced-executor formula instead; the worked
|
|
124
|
+
examples above give shapes like 29 × 5 × 18GB for a 10-node/16-core/64GB cluster, or
|
|
125
|
+
17 × 5 × 19G for a 6-node cluster with the same per-node hardware[^1][^2]. Keep
|
|
126
|
+
`--executor-cores` at or below roughly 5 so HDFS client concurrency stays in the range
|
|
127
|
+
where it sustains full write throughput[^1].
|
|
128
|
+
|
|
129
|
+
When enabling off-heap memory, account for `spark.memory.offHeap.size` in the
|
|
130
|
+
container/pod memory budget explicitly; `spark.executor.memoryOverhead`'s default of
|
|
131
|
+
`max(0.1 * executorMemory, 384MB)` does not include it[^5]. If you're still referencing
|
|
132
|
+
the older `spark.yarn.executor.memoryOverhead` name, note it was removed in Spark
|
|
133
|
+
3.0[^5].
|
|
134
|
+
|
|
135
|
+
On Kubernetes, remember that `spark.kubernetes.executor.request.cores` and
|
|
136
|
+
`.limit.cores` govern the pod's CPU request/limit and take priority over
|
|
137
|
+
`spark.executor.cores` for that purpose[^6], while `spark.executor.cores` separately
|
|
138
|
+
governs Spark's own task-slot count[^3]. The two need to be reasoned about together
|
|
139
|
+
rather than assumed to be redundant.
|
|
140
|
+
|
|
141
|
+
Tune `spark.locality.wait` (and its per-level overrides) upward when tasks are
|
|
142
|
+
long-running and locality is poor, since the default is tuned for typical workloads[^3];
|
|
143
|
+
set `spark.locality.wait.node` to `0` to skip straight to rack locality when node
|
|
144
|
+
locality isn't achievable[^3].
|
|
145
|
+
|
|
146
|
+
Turn on shuffle tracking (`spark.dynamicAllocation.shuffleTracking.enabled`, default
|
|
147
|
+
`true` since Spark 3.0) or the external shuffle service so dynamic allocation can
|
|
148
|
+
reclaim executors without losing shuffle output or forcing recomputation[^3][^6]. The
|
|
149
|
+
external shuffle service additionally lets persisted RDD blocks survive executor
|
|
150
|
+
removal when `spark.shuffle.service.fetch.rdd.enabled` is set, and executors holding
|
|
151
|
+
cached blocks are, by default, never removed at all, tunable via
|
|
152
|
+
`spark.dynamicAllocation.cachedExecutorIdleTimeout`[^6].
|
|
153
|
+
|
|
154
|
+
Set `spark.driver.maxResultSize` as a guardrail on serialized action-result size, and
|
|
155
|
+
size `spark.driver.memory` with it in mind rather than in isolation[^3]. Use
|
|
156
|
+
`spark.dynamicAllocation.executorAllocationRatio` (default `1.0`, added in 2.4.0) to
|
|
157
|
+
scale down over-allocation when tasks are small; for example, a value of `0.5` halves
|
|
158
|
+
the target executor count dynamic allocation would otherwise compute[^3]. Raise
|
|
159
|
+
`spark.task.cpus` above its default of `1` when a task needs more than one core; doing
|
|
160
|
+
so proportionally reduces concurrent task slots per executor
|
|
161
|
+
(`spark.executor.cores / spark.task.cpus`)[^3].
|
|
162
|
+
|
|
163
|
+
## Sources
|
|
164
|
+
|
|
165
|
+
[^1]: [Distribution of Executors, Cores and Memory for a Spark Application](https://raw.githubusercontent.com/spoddutur/spark-notes/master/distribution_of_executors_cores_and_memory_for_spark_application.md)
|
|
166
|
+
[^2]: [How to Tune Your Apache Spark Jobs (Part 2)](https://blog.cloudera.com/how-to-tune-your-apache-spark-jobs-part-2/)
|
|
167
|
+
[^3]: [Configuration — Spark](https://spark.apache.org/docs/latest/configuration.html)
|
|
168
|
+
[^4]: [Spark Tuning Guide](https://spark.apache.org/docs/latest/tuning.html)
|
|
169
|
+
[^5]: [Dive into Spark Memory](https://luminousmen.com/post/dive-into-spark-memory)
|
|
170
|
+
[^6]: [Job Scheduling — Dynamic Resource Allocation](https://spark.apache.org/docs/latest/job-scheduling.html)
|
|
@@ -0,0 +1,171 @@
|
|
|
1
|
+
# Anti-Patterns
|
|
2
|
+
|
|
3
|
+
Practitioner checklists on Spark performance converge on a recurring handful of mistakes
|
|
4
|
+
rather than exotic ones. Each entry below covers what the mistake is, how to detect it,
|
|
5
|
+
why it hurts, and how to fix it.
|
|
6
|
+
|
|
7
|
+
1. **Staying on the RDD API instead of DataFrame/Dataset**
|
|
8
|
+
- **What:** Building jobs on the RDD API (or RDD-based MLlib) instead of the
|
|
9
|
+
DataFrame/Dataset API.
|
|
10
|
+
- **Detect:** Code that drives its logic through `.map`/`.filter` lambdas over RDDs, or
|
|
11
|
+
that calls into the RDD-based MLlib API.
|
|
12
|
+
- **Why:** RDDs are lambda-driven, so Spark cannot see inside the closures. It can't apply
|
|
13
|
+
predicate pushdown, filter reordering, Adaptive Query Execution, or cost-based join
|
|
14
|
+
reordering. The RDD-based MLlib API is in maintenance mode for the same reason.[^1]
|
|
15
|
+
- **Fix:** Use the DataFrame/Dataset API so Catalyst can optimize the query plan.
|
|
16
|
+
|
|
17
|
+
2. **Reading CSV/JSON without an explicit schema**
|
|
18
|
+
- **What:** Reading raw CSV or JSON files without supplying an explicit schema.
|
|
19
|
+
- **Detect:** A read step that spends noticeable time before any real work starts, because
|
|
20
|
+
Spark is inferring types from the data itself.
|
|
21
|
+
- **Why:** Schema inference forces a full scan of the data just to determine types, and
|
|
22
|
+
CSV/JSON don't support the column pruning, predicate pushdown, or stats-based file
|
|
23
|
+
skipping that a columnar format does.[^1]
|
|
24
|
+
- **Fix:** Supply an explicit schema, or prefer Parquet for analytical workloads, where column
|
|
25
|
+
pruning, predicate pushdown, and stats-based file skipping work out of the box.[^1]
|
|
26
|
+
|
|
27
|
+
3. **Compressing input files with a non-splittable codec (GZIP)**
|
|
28
|
+
- **What:** Storing input data as large GZIP files.
|
|
29
|
+
- **Detect:** One task/executor taking far longer than the rest on a stage reading
|
|
30
|
+
GZIP-compressed input, while other executors sit idle.
|
|
31
|
+
- **Why:** A GZIP file can't be split across executors, so a single node has to decompress
|
|
32
|
+
the whole file alone.[^1]
|
|
33
|
+
- **Fix:** Use a splittable codec instead: Snappy, LZ4, or ZSTD.[^1]
|
|
34
|
+
|
|
35
|
+
4. **Storing data as JSON instead of a binary columnar/row format**
|
|
36
|
+
- **What:** Persisting datasets as JSON rather than a binary columnar or row format.
|
|
37
|
+
- **Detect:** Read-heavy jobs where a large, recurring share of stage time goes to parsing
|
|
38
|
+
the same JSON structures on every read.
|
|
39
|
+
- **Why:** JSON has to be re-parsed from text on every single read.[^2]
|
|
40
|
+
- **Fix:** Store data as Avro, Parquet, Thrift, or Protobuf structs in a sequence file to
|
|
41
|
+
avoid the repeated parsing cost.[^2]
|
|
42
|
+
|
|
43
|
+
5. **Not registering custom classes with Kryo**
|
|
44
|
+
- **What:** Using Kryo serialization without registering the application's custom classes.
|
|
45
|
+
- **Detect:** Larger-than-expected serialized record sizes, and poor efficiency when using a
|
|
46
|
+
serialized cache storage level such as `MEMORY_SER`.
|
|
47
|
+
- **Why:** Unregistered classes increase serialized-record size and hurt serialized cache
|
|
48
|
+
storage levels.[^2]
|
|
49
|
+
- **Fix:** Register the application's custom classes with Kryo.
|
|
50
|
+
|
|
51
|
+
6. **Calling `collect()` on a large DataFrame**
|
|
52
|
+
- **What:** Materializing an entire DataFrame into driver memory with `collect()`.
|
|
53
|
+
- **Detect:** A driver `OutOfMemoryError`, or a job that aborts once
|
|
54
|
+
`spark.driver.maxResultSize` (default 1 GB) is exceeded. Spark also logs a warning once a
|
|
55
|
+
single task's serialized result passes roughly 1 MB.[^3][^1]
|
|
56
|
+
- **Why:** `collect()` gathers every `Row` from every partition into a single in-memory
|
|
57
|
+
`Array` that lives entirely on the driver, not the cluster. The driver JVM has to hold
|
|
58
|
+
the whole result set in its own heap.[^3][^4][^5][^6]
|
|
59
|
+
- **Fix:** Use aggregations or `take(n)` instead of collecting the full result set.
|
|
60
|
+
|
|
61
|
+
> **PySpark:** `toPandas()` performs the same driver-side collection as `collect()` for
|
|
62
|
+
> Python users, so avoid it on large DataFrames too.[^6]
|
|
63
|
+
|
|
64
|
+
7. **Assuming `cache()` + `count()` guarantees the DataFrame stays fully cached**
|
|
65
|
+
- **What:** Treating `df.cache(); df.count()` as a guarantee that the DataFrame remains
|
|
66
|
+
fully persisted for later actions.
|
|
67
|
+
- **Detect:** Later actions on a supposedly-cached DataFrame recompute from source instead
|
|
68
|
+
of hitting the cache. There's no built-in visibility into how much of a DataFrame is
|
|
69
|
+
still actually cached, so this typically only surfaces as unexpectedly slow
|
|
70
|
+
recomputation.[^7]
|
|
71
|
+
- **Why:** `count()` does force every partition to be computed and attempted for caching,
|
|
72
|
+
unlike `take`/`limit`, which only touch the partitions they need.[^5][^7] But partitions
|
|
73
|
+
can't be fractionally cached: if there isn't room for all of them, the ones that don't
|
|
74
|
+
fit are silently dropped and recomputed on next access.[^5] Cached blocks also compete
|
|
75
|
+
with execution for the same memory pool and can be evicted under pressure without
|
|
76
|
+
warning, and caching is tied to the *analyzed* (pre-optimization) logical plan, so a
|
|
77
|
+
semantically identical query with a different analyzed plan bypasses the cache entirely
|
|
78
|
+
and recomputes from source.[^7] Losing an executor drops any cached block that wasn't
|
|
79
|
+
stored with a replicated storage level like `MEMORY_AND_DISK_2`.[^7]
|
|
80
|
+
- **Fix:** Keep using `.cache()` followed by `.count()` to force materialization, but don't
|
|
81
|
+
assume that guarantees persistence (there's no built-in way to confirm how much of the
|
|
82
|
+
DataFrame is actually cached[^7]), and use a replicated storage level if losing a cached
|
|
83
|
+
partition to executor failure would be costly.
|
|
84
|
+
|
|
85
|
+
8. **Row-at-a-time Python UDFs**
|
|
86
|
+
- **What:** Writing Python UDFs that operate one row at a time instead of vectorized Pandas
|
|
87
|
+
UDFs.
|
|
88
|
+
- **Detect:** A UDF-heavy stage where most of the time goes to serialization rather than
|
|
89
|
+
actual computation.[^9]
|
|
90
|
+
- **Why:** The overhead is serialization plus per-row JVM↔Python data movement, not raw
|
|
91
|
+
Python interpreter speed. Row-at-a-time UDFs "suffer from high serialization and
|
|
92
|
+
invocation overhead."[^8] RDD-era Python UDFs pay a *double* serialization cost: Java/Scala
|
|
93
|
+
objects are serialized, then re-serialized to Python via `cloudpickle` and back on every
|
|
94
|
+
call.[^9] A Databricks benchmark on a 10M-row DataFrame showed vectorized Pandas UDFs
|
|
95
|
+
beating row-at-a-time UDFs "across the board, ranging from 3x to over 100x."[^8]
|
|
96
|
+
- **Fix:** Use vectorized Pandas UDFs instead of row-at-a-time UDFs: once data is
|
|
97
|
+
transferred via Arrow, there's no need to serialize/pickle it, since it's already in a
|
|
98
|
+
format consumable by the Python process.[^5]
|
|
99
|
+
|
|
100
|
+
9. **Leaving `spark.sql.shuffle.partitions` at its default of 200**
|
|
101
|
+
- **What:** Never tuning `spark.sql.shuffle.partitions` away from its default of 200 for
|
|
102
|
+
wide transformations (`join`, `groupBy`, aggregations).
|
|
103
|
+
- **Detect:** In the Spark UI Stages tab, too few partitions for the data size shows up as
|
|
104
|
+
"Spill (Memory)" or "Spill (Disk)" entries against a stage; too many partitions for the
|
|
105
|
+
data size shows up as a large task/stage count where each task's duration is dominated by
|
|
106
|
+
scheduling and bookkeeping, often paired with a large number of tiny output files.[^10]
|
|
107
|
+
- **Why:** 200 is a fixed default regardless of whether you're joining 5MB or 5TB.[^10][^11]
|
|
108
|
+
Too few partitions means each one holds more data than fits comfortably in executor
|
|
109
|
+
memory, which increases load per executor and leads to spills once partition size exceeds
|
|
110
|
+
available memory.[^10] Too many partitions on a small dataset can shrink tasks to around
|
|
111
|
+
ten rows each, so "most of your CPUs will just be sitting there doing nothing."[^10]
|
|
112
|
+
- **Fix:** Increase `spark.sql.shuffle.partitions` if tasks are processing multiple GB each
|
|
113
|
+
and spilling; decrease it if tasks finish in a couple of seconds and write tiny files. A
|
|
114
|
+
rough starting target is 100–200 MB of data per task, tuned per dataset.[^10]
|
|
115
|
+
|
|
116
|
+
10. **Broadcasting a table that's too large**
|
|
117
|
+
- **What:** Forcing or allowing a broadcast join on a table that's too large to broadcast
|
|
118
|
+
safely.
|
|
119
|
+
- **Detect:** A driver `OutOfMemoryError` during a broadcast join.
|
|
120
|
+
- **Why:** The small side of the join has to be collected onto the driver before it can be
|
|
121
|
+
broadcast to every executor. That collection step is the same driver-memory operation
|
|
122
|
+
as `.collect()`, and it's what fails, not executor-side replication or a serialization
|
|
123
|
+
timeout: "if you try to broadcast something too large, you can crash your driver node
|
|
124
|
+
(because that collect is expensive)."[^4]
|
|
125
|
+
- **Fix:** Use `spark.sql.autoBroadcastJoinThreshold` to control the maximum size Spark
|
|
126
|
+
will broadcast, or increase driver memory.[^4]
|
|
127
|
+
|
|
128
|
+
11. **Reaching for `coalesce(1)` before every single-file write**
|
|
129
|
+
- **What:** Always using `coalesce(1)` instead of `repartition(1)` when a single output
|
|
130
|
+
file is needed.
|
|
131
|
+
- **Detect:** The entire upstream stage (including filtering/transformation work that
|
|
132
|
+
would otherwise run in parallel) collapses onto a single task/executor, leaving the
|
|
133
|
+
rest of the cluster idle.
|
|
134
|
+
- **Why:** `coalesce` is a narrow transformation, so it "causes the upstream partitions in
|
|
135
|
+
the entire stage to execute with the level of parallelism assigned by coalesce," fusing
|
|
136
|
+
all preceding work down to one task.[^6] `repartition` triggers a full shuffle instead,
|
|
137
|
+
which inserts an explicit shuffle boundary so the upstream stage keeps its original
|
|
138
|
+
parallelism.[^4][^6]
|
|
139
|
+
- **Fix:** Use `repartition(1)` when there is meaningful upstream computation you don't
|
|
140
|
+
want collapsed onto a single executor; reserve `coalesce(1)` for when the upstream stage
|
|
141
|
+
is already cheap and paying for a shuffle would be wasted cost.[^6]
|
|
142
|
+
|
|
143
|
+
12. **Not verifying predicate pushdown is actually happening**
|
|
144
|
+
- **What:** Assuming filters and column projections are pushed down to the data source
|
|
145
|
+
without checking the physical plan.
|
|
146
|
+
- **Detect:** Run `.explain()` and look at the `Scan` node for a `PushedFilters` marker.
|
|
147
|
+
For example, `*Scan JDBCRel... PushedFilters: [*In(DEST_COUNTRY_NAME, [Anguilla,
|
|
148
|
+
Sweden])]`. Column pruning is verified the same way: a `.select()` on one column should
|
|
149
|
+
show a narrowed `ReadSchema` at the scan rather than a full-table scan.[^4]
|
|
150
|
+
- **Why:** Spark pushes down simple filters (column equality, `IN`, `IS NULL`) to JDBC
|
|
151
|
+
sources automatically, but anything with a computed column or cast won't push down and
|
|
152
|
+
gets evaluated in Spark after the full read.[^1] Some sources also only partially handle
|
|
153
|
+
a filter and leave Spark to re-evaluate it as a safety mechanism; that's still a "good"
|
|
154
|
+
pushdown since the amount of data read is reduced.[^6]
|
|
155
|
+
- **Fix:** Check `explain` output (paired with `printSchema`)[^6] for `PushedFilters` after
|
|
156
|
+
writing a filter, and rewrite filters that use casts or computed columns as simple
|
|
157
|
+
predicates where possible so they can push down.[^4][^1]
|
|
158
|
+
|
|
159
|
+
## Sources
|
|
160
|
+
|
|
161
|
+
[^1]: [The Apache Spark Optimization Checklist](https://luminousmen.com/post/the-apache-spark-optimization-checklist)
|
|
162
|
+
[^2]: [How to Tune Your Apache Spark Jobs (Part 2)](https://blog.cloudera.com/how-to-tune-your-apache-spark-jobs-part-2/)
|
|
163
|
+
[^3]: *Advanced Analytics with PySpark*, Tandon, Ryza, Laserson et al., ch. 2
|
|
164
|
+
[^4]: *Spark: The Definitive Guide*, Chambers & Zaharia, chs. 5, 8, 9, 18, 19
|
|
165
|
+
[^5]: *Learning Spark, 2nd Edition*, Damji, Wenig, Das & Lee, chs. 3, 5, 7
|
|
166
|
+
[^6]: *High Performance Spark, 2nd Edition*, Karau, Polak & Warren, chs. 5, 7
|
|
167
|
+
[^7]: [Explaining the Mechanics of Spark Caching](https://luminousmen.com/post/explaining-the-mechanics-of-spark-caching)
|
|
168
|
+
[^8]: [Introducing Vectorized UDFs for PySpark](https://www.databricks.com/blog/2017/10/30/introducing-vectorized-udfs-for-pyspark.html)
|
|
169
|
+
[^9]: [Spark Tips: DataFrame API](https://luminousmen.com/post/spark-tips-dataframe-api)
|
|
170
|
+
[^10]: [Spark Partitions](https://luminousmen.com/post/spark-partitions)
|
|
171
|
+
[^11]: [Configuration — Spark](https://spark.apache.org/docs/latest/configuration.html)
|
|
@@ -0,0 +1,87 @@
|
|
|
1
|
+
# Metrics Glossary
|
|
2
|
+
|
|
3
|
+
<!-- #metric-* IDs are docs-internal, forward-compatible; the sibling contract consumes only #bottleneck-*. Rename-safe until the sibling deep-links metrics. -->
|
|
4
|
+
|
|
5
|
+
Reference for every metric surfaced elsewhere in this guide: what it measures, where Spark records it in the event log, and what a problematic value looks like.
|
|
6
|
+
|
|
7
|
+
## Task duration (P50/P95/max) {#metric-task-duration}
|
|
8
|
+
|
|
9
|
+
Per-task wall-clock duration, aggregated across a stage's tasks into percentiles (P50, P95, max) to surface skew and stragglers. It is computed from `SparkListenerTaskEnd` events rather than read off a single field: each task end carries its own duration, and the percentiles are derived by aggregating those values across the stage.
|
|
10
|
+
|
|
11
|
+
A problematic value looks like a small fraction of tasks running far past the rest: a stage is classified as a straggler when more than 5% of tasks run at least 4x the median duration (with at least 10 tasks in the stage), or when speculative tasks were fired.
|
|
12
|
+
|
|
13
|
+
## Shuffle read bytes {#metric-shuffle-read-bytes}
|
|
14
|
+
|
|
15
|
+
The volume of shuffle data a task reads from remote executors, recorded in `taskMetrics.shuffleReadMetrics.remoteBytesRead`.
|
|
16
|
+
|
|
17
|
+
## Shuffle write bytes {#metric-shuffle-write-bytes}
|
|
18
|
+
|
|
19
|
+
The size of the shuffle output a task writes, recorded in `taskMetrics.shuffleWriteMetrics.bytesWritten`. Spark's own metric description calls it simply the "Number of bytes written in shuffle operations," without stating compression state on its own[^1].
|
|
20
|
+
|
|
21
|
+
The sort-based shuffle writer fills in the mechanics: incoming records are serialized as soon as they reach the shuffle writer and buffered in serialized form while sorting[^2]. When the spill compression codec supports concatenating compressed data, the final merge step concatenates the already-compressed spill partitions directly into the output file, using `transferTo` rather than decompressing and recompressing[^2]. So `bytesWritten` counts compressed, post-serialization bytes, and in this common fast-merge path the final written output is built directly from data that had already been spilled to disk during sorting, rather than written fresh at merge time.
|
|
22
|
+
|
|
23
|
+
## Memory bytes spilled {#metric-memory-bytes-spilled}
|
|
24
|
+
|
|
25
|
+
Bytes a task spilled from in-memory structures, recorded in `taskMetrics.memoryBytesSpilled`. Any value above zero is a spill warning, with the skew-vs-volume classification (based on what fraction of a stage's tasks show zero spill) as the actionable signal.
|
|
26
|
+
|
|
27
|
+
## Disk bytes spilled {#metric-disk-bytes-spilled}
|
|
28
|
+
|
|
29
|
+
The on-disk counterpart to memory spill: bytes a task spilled to disk, recorded in `taskMetrics.diskBytesSpilled`. It feeds the same spill classification as memory bytes spilled.
|
|
30
|
+
|
|
31
|
+
## JVM GC time {#metric-jvm-gc-time}
|
|
32
|
+
|
|
33
|
+
Elapsed time the JVM spent in garbage collection while a task executed, recorded in `taskMetrics.jvmGCTime` and expressed in milliseconds[^1]. The value is cumulative across the task's full execution window: the sum of every GC pause that occurred during that task's run, not just the most recent one[^1]. This matches how it's serialized: as a single scalar long, consistent with an accumulator rather than a per-GC-event log entry[^3].
|
|
34
|
+
|
|
35
|
+
## gcPct {#metric-gcpct}
|
|
36
|
+
|
|
37
|
+
A synthetic ratio, `jvmGCTime / executorRunTime`, not a raw Spark field. It drives GC-bottleneck classification: above 10% is a warning, above 20% is critical.
|
|
38
|
+
|
|
39
|
+
## Executor run time {#metric-executor-run-time}
|
|
40
|
+
|
|
41
|
+
Elapsed time the executor spent running a task, recorded in `taskMetrics.executorRunTime` and expressed in milliseconds[^1]. `TaskMetrics` (and therefore `executorRunTime`) is serialized as an optional part of the `SparkListenerTaskEnd` payload, not guaranteed on every task end[^3]. In practice it is available for failed tasks that got far enough to actually run, such as `ExceptionFailure` or `TaskKilled`. The event-log deserialization code even falls back to reading accumulator updates out of the embedded `TaskMetrics` for old, Spark-1.x-era logs, which only makes sense if the metrics object is normally populated for that failure type[^3]. It can be absent for reasons like `Resubmitted`, where the task attempt never completed on that executor[^3].
|
|
42
|
+
|
|
43
|
+
## Fetch wait time ratio {#metric-fetch-wait-time-ratio}
|
|
44
|
+
|
|
45
|
+
A synthetic ratio, `fetchWaitTime / taskDuration`, not a raw Spark field: the time a task spent blocked waiting on remote shuffle blocks, relative to its total duration.
|
|
46
|
+
|
|
47
|
+
## Input bytes {#metric-input-bytes}
|
|
48
|
+
|
|
49
|
+
Bytes a task read as input, recorded in `taskMetrics.inputMetrics.bytesRead`.
|
|
50
|
+
|
|
51
|
+
## Output bytes {#metric-output-bytes}
|
|
52
|
+
|
|
53
|
+
Bytes a task wrote as output, recorded in `taskMetrics.outputMetrics.bytesWritten`.
|
|
54
|
+
|
|
55
|
+
## I/O ratio {#metric-io-ratio}
|
|
56
|
+
|
|
57
|
+
A synthetic ratio, `outputBytes / inputBytes`, not a raw Spark field.
|
|
58
|
+
|
|
59
|
+
## Peak execution memory {#metric-peak-execution-memory}
|
|
60
|
+
|
|
61
|
+
Peak memory recorded in `taskMetrics.peakExecutionMemory`. This is a task-level accumulator, distinct from the separate executor-level `peakMemoryMetrics.OnHeapExecutionMemory` and `.OffHeapExecutionMemory` gauges, which report the on-heap and off-heap execution pools as two separate numbers at the executor level rather than as a single per-task figure[^1].
|
|
62
|
+
|
|
63
|
+
## Failed tasks / failure rate {#metric-failed-tasks}
|
|
64
|
+
|
|
65
|
+
Tasks whose `SparkListenerTaskEnd` reason is not `Success`. The `Reason` field holds the formatted class name of whichever `TaskEndReason` was assigned to that task end[^3], and the canonical set of reasons includes `FetchFailed`, `ExceptionFailure`, `TaskResultLost`, `TaskKilled`, `TaskCommitDenied`, `ExecutorLostFailure`, and `UnknownReason`[^3]. Because `TaskMetrics` is only an optional part of the task-end payload, it can be absent for some of these reasons: for example `Resubmitted`, where the task attempt never actually completed on that executor[^3].
|
|
66
|
+
|
|
67
|
+
## Speculative tasks / straggler count {#metric-speculative-tasks}
|
|
68
|
+
|
|
69
|
+
Tasks launched as speculative retries of a slow-running task, recorded via `SparkListenerTaskStart` where `speculative = true`. Any speculative task firing (or more than 5% of a stage's tasks running at least 4x the median duration, with at least 10 tasks in the stage) is a straggler signal.
|
|
70
|
+
|
|
71
|
+
## Executor count (added/removed/concurrent) {#metric-executor-count}
|
|
72
|
+
|
|
73
|
+
The number of executors added, removed, or concurrently running, recorded via `SparkListenerExecutorAdded` / `SparkListenerExecutorRemoved` events.
|
|
74
|
+
|
|
75
|
+
## Stage wall-clock duration {#metric-stage-duration}
|
|
76
|
+
|
|
77
|
+
A stage's total elapsed time, recorded from `SparkListenerStageCompleted`: `completionTime − submissionTime`.
|
|
78
|
+
|
|
79
|
+
## firstStageSubmittedAt {#metric-first-stage-submitted-at}
|
|
80
|
+
|
|
81
|
+
A synthetic field: the timestamp of the first `SparkListenerStageSubmitted` event in the event log. The gap between this timestamp and the application's start time drives cold-start classification: more than 30 seconds is a warning.
|
|
82
|
+
|
|
83
|
+
## Sources
|
|
84
|
+
|
|
85
|
+
[^1]: [Monitoring and Instrumentation](https://spark.apache.org/docs/latest/monitoring.html#spark-history-server)
|
|
86
|
+
[^2]: [SortShuffleManager.scala](https://raw.githubusercontent.com/apache/spark/v3.5.0/core/src/main/scala/org/apache/spark/shuffle/sort/SortShuffleManager.scala)
|
|
87
|
+
[^3]: [JsonProtocol.scala](https://raw.githubusercontent.com/apache/spark/v3.5.0/core/src/main/scala/org/apache/spark/util/JsonProtocol.scala)
|
|
@@ -0,0 +1,93 @@
|
|
|
1
|
+
# Spark Config Quick-Reference
|
|
2
|
+
|
|
3
|
+
A cross-reference of the Spark and PySpark configuration properties discussed elsewhere in this guide, plus the JVM garbage-collection flags relevant to executor tuning. Defaults and version notes below are traced to Spark's own configuration docs, its SQLConf source, and the cited literature. Where a source doesn't give a hard number (a range, a ceiling, a recommended workload profile), the table leaves it out rather than guess.
|
|
4
|
+
|
|
5
|
+
## Spark properties
|
|
6
|
+
|
|
7
|
+
| Key | Default | Recommended | Notes |
|
|
8
|
+
|---|---|---|---|
|
|
9
|
+
| `spark.sql.shuffle.partitions` | 200 (unchanged since Spark 1.1.0)[^1][^2][^3] | | AQE doesn't override this config value; see notes below for the effective runtime count. |
|
|
10
|
+
| `spark.sql.adaptive.enabled` | `true`[^4] | `true` | |
|
|
11
|
+
| `spark.sql.adaptive.coalescePartitions.enabled` | `true` (since 3.0.0)[^1] | `true` | Requires AQE; merges contiguous small post-shuffle partitions instead of one task per configured partition[^1]. |
|
|
12
|
+
| `spark.sql.adaptive.coalescePartitions.initialPartitionNum` | unset → falls back to `spark.sql.shuffle.partitions` (200)[^1] | | Sets the partition count entering the coalescing step. |
|
|
13
|
+
| `spark.sql.adaptive.coalescePartitions.parallelismFirst` | `true` (since 3.2.0)[^1][^5] | `false` on busy clusters[^1] | When `true`, ignores `advisoryPartitionSizeInBytes` and derives a target from the cluster's default parallelism, respecting only the `minPartitionSize` floor. |
|
|
14
|
+
| `spark.sql.adaptive.advisoryPartitionSizeInBytes` | 64MB (since 3.0.0)[^1][^4][^5] | | Ignored during coalescing when `parallelismFirst=true` (the default). |
|
|
15
|
+
| `spark.sql.adaptive.coalescePartitions.minPartitionSize` | 1MB (since 3.2.0)[^5] | | Floor enforced when the target size is ignored: the default `parallelismFirst=true` case[^1]. |
|
|
16
|
+
| `spark.memory.fraction` | 0.6[^6] | | See notes below; no hard 0–1 ceiling documented in the cited sources. |
|
|
17
|
+
| `spark.memory.storageFraction` | 0.5[^7] | | A fraction *of* the region sized by `spark.memory.fraction`, not an independent pool; see notes below. |
|
|
18
|
+
| `spark.sql.execution.arrow.pyspark.enabled` | `false`[^8] | | See PySpark callout below. |
|
|
19
|
+
| `spark.sql.execution.arrow.pyspark.fallback.enabled` | `true` (inherited from the deprecated `arrow.fallback.enabled`)[^8][^5][^4] | | Falls back to the non-Arrow path automatically if a conversion error occurs, before computation runs[^8]. |
|
|
20
|
+
| `spark.dynamicAllocation.shuffleTracking.enabled` | `true` (since 3.0.0)[^4] | | Lets dynamic allocation track shuffle files per executor without an external shuffle service. |
|
|
21
|
+
| `spark.dynamicAllocation.shuffleTracking.timeout` | `infinity`[^4] | | Executors holding shuffle data wait for it to be garbage collected before release, by default; set a finite value if GC isn't keeping up. |
|
|
22
|
+
| `spark.shuffle.service.enabled` | | | External shuffle service; see `#config-shuffle-service` below. |
|
|
23
|
+
| `spark.dynamicAllocation.minExecutors` / `maxExecutors` | `0` / `infinity`[^4] | | Autoscale bounds; see `#config-autoscale-bounds` below. |
|
|
24
|
+
| `spark.serializer` | `org.apache.spark.serializer.JavaSerializer`[^4] | `KryoSerializer` | See `#config-serializer` below. |
|
|
25
|
+
| `spark.executor.memoryOverhead` | greater of 10% of executor memory or 384MB[^7] | | Off-heap overhead; see `#config-memory-overhead` below. |
|
|
26
|
+
|
|
27
|
+
### Shuffle partitions and AQE coalescing
|
|
28
|
+
|
|
29
|
+
The 200 default for `spark.sql.shuffle.partitions` hasn't moved since 1.1.0, and AQE (on by default since Spark 3.0) doesn't change that config value[^1][^2][^3]. What it changes is the effective number of partitions used at runtime. With `coalescePartitions.enabled` also on by default, Spark merges small contiguous post-shuffle partitions instead of running one task per configured partition[^1]. Because `parallelismFirst` defaults to `true` since 3.2.0, that merge target usually isn't the 64MB `advisoryPartitionSizeInBytes` value; it's derived from the cluster's default parallelism, with `minPartitionSize` (1MB) as the only enforced floor[^1][^5]. Databricks' own writeup on AQE shows the reduce-task count actually shrinking based on measured data volume[^9], and Learning Spark accordingly calls the static 200 default "too high for smaller or streaming workloads"[^3]. A related, distinct knob (`spark.sql.adaptive.coalescePartitions.minPartitionNum`) sets a minimum parallelism floor for data that's slow to compute despite being small, per High Performance Spark[^10].
|
|
30
|
+
|
|
31
|
+
### `spark.memory.fraction` and `spark.memory.storageFraction`
|
|
32
|
+
|
|
33
|
+
`spark.memory.fraction` sizes the unified execution/storage region (M) as a fraction of (JVM heap − 300MiB); the tuning guide frames it as a knob to fit M "comfortably within the JVM's old or tenured generation" rather than stating an enforced numeric range[^6]. The cited sources document no failure threshold tied to a specific value like 0.9. What they do document is the risk of pushing the fraction high: the complement, `1 − spark.memory.fraction`, is untracked "User Memory" for UDFs, Python/Arrow glue, and native buffers, and starving it (which is what raising the fraction toward 0.9 does) leads to "GC pressure or random OOMs," with no warning from Spark[^7].
|
|
34
|
+
|
|
35
|
+
`spark.memory.storageFraction` isn't independent of that: it's a fraction *of* M, the same region `spark.memory.fraction` sizes[^6]. Inside that shared pool, execution can evict storage down to the threshold set by `storageFraction`, but storage can never evict execution[^11][^6]. With the defaults (0.6 and 0.5), that works out to execution and storage each getting 30% of usable heap; raising `memory.fraction` scales both pools at once, while `storageFraction` only re-splits the pool that's already been carved out[^7].
|
|
36
|
+
|
|
37
|
+
> **PySpark:** `spark.sql.execution.arrow.pyspark.enabled` is off by default and governs Arrow use for `DataFrame.toPandas()` and `SparkSession.createDataFrame()` from a Pandas DataFrame or NumPy array[^8]. Its documented risk is a type-coverage gap, not silent data corruption: `ArrayType` of `TimestampType` is explicitly unsupported[^4][^5], and more generally an unsupported column type raises an error rather than converting wrongly[^8]. `spark.sql.execution.arrow.pyspark.fallback.enabled` defaults to `true` and automatically falls back to the non-Arrow path if a conversion error occurs, before any computation runs[^8][^5][^4]. So the failure mode is an exception plus fallback, not a wrong result. Separately, and not specific to this Arrow config, PySpark's own type coercion has its own hazards: numeric values passed for `ByteType`/`ShortType`/`IntegerType` must fall within fixed ranges or get rejected or converted unexpectedly[^2].
|
|
38
|
+
|
|
39
|
+
### Shuffle service {#config-shuffle-service}
|
|
40
|
+
|
|
41
|
+
During a shuffle, an executor writes its map output to local disk and then serves fetch requests for that data itself[^14]. Dynamic allocation can reclaim an idle executor before a later stage has fetched all the shuffle blocks it holds (especially with stragglers, tasks that run much longer than their peers), and once that executor is gone, any stage that still needs its shuffle output gets a `FetchFailed` and has to recompute it[^14]. The external shuffle service fixes this by moving shuffle-block serving out of the executor process: it's a long-running process on each node, independent of any particular application's executors, and once `spark.shuffle.service.enabled` is `true`, executors fetch shuffle blocks from it instead of from each other, so an executor's shuffle output keeps being served after that executor is reclaimed[^14][^4]. Setting up dynamic allocation requires enabling this service on every worker node in addition to `spark.dynamicAllocation.enabled`[^2]. As of Spark 3.0, shuffle-file tracking (`spark.dynamicAllocation.shuffleTracking.enabled`) is an alternative to the external shuffle service for the same safe-removal problem[^4].
|
|
42
|
+
|
|
43
|
+
**Limitations / false-positive risk:** the service can be intentionally left off when dynamic allocation relies on shuffle-file tracking instead, so a disabled setting is not automatically a misconfiguration.
|
|
44
|
+
|
|
45
|
+
### Autoscale bounds {#config-autoscale-bounds}
|
|
46
|
+
|
|
47
|
+
By default, `spark.dynamicAllocation.minExecutors` is `0` and `spark.dynamicAllocation.maxExecutors` is `infinity`[^4]. Leaving `maxExecutors` unset therefore leaves dynamic allocation unbounded on the high end as far as Spark itself is concerned; in practice the ceiling ends up being whatever the cluster or scheduler enforces outside Spark[^4]. `spark.dynamicAllocation.initialExecutors` defaults to `minExecutors`, unless `--num-executors` (or `spark.executor.instances`) sets a larger starting value[^4]. Whatever executor count `executorAllocationRatio` computes to maximize parallelism, that target is still clamped by the `minExecutors`/`maxExecutors` floor and ceiling[^4].
|
|
48
|
+
|
|
49
|
+
**Limitations / false-positive risk:** an unbounded `maxExecutors` is often deliberate on clusters where the scheduler enforces the real ceiling, so an unset high end is not always wrong.
|
|
50
|
+
|
|
51
|
+
### Serializer {#config-serializer}
|
|
52
|
+
|
|
53
|
+
The default `spark.serializer` is `org.apache.spark.serializer.JavaSerializer`, which works with any `Serializable` Java object but is "quite slow"; Spark's configuration reference recommends switching to `KryoSerializer` "when speed is necessary"[^4]. The tuning guide is more specific: Kryo is "significantly faster and more compact than Java serialization (often as much as 10x)," though it doesn't support all `Serializable` types and needs its classes registered in advance for best performance[^6]. That registration requirement is the stated reason Kryo isn't the default: "The only reason Kryo is not the default is because of the custom registration requirement," and Spark recommends trying it for any network-intensive application[^6]. Since Spark 2.0.0, Spark internally uses Kryo regardless of this setting when shuffling RDDs of simple types, arrays of simple types, or strings, and auto-registers common Scala classes via Twitter chill's `AllScalaRegistrar`[^6]. A practitioner summary puts Java serialization at "2-10x slower and larger on the wire" on every shuffle, with the caveat that classes left unregistered silently fall back to Java serialization unless registered via `spark.kryo.classesToRegister` or `spark.kryo.registrator`[^15]. Persisting RDDs in serialized form is another case where Kryo can be more space-efficient than Java serialization[^10].
|
|
54
|
+
|
|
55
|
+
**Limitations / false-positive risk:** `JavaSerializer` may be kept on purpose when an application depends on `Serializable` types Kryo cannot handle, so a non-Kryo setting is not necessarily a mistake.
|
|
56
|
+
|
|
57
|
+
### Memory overhead {#config-memory-overhead}
|
|
58
|
+
|
|
59
|
+
`spark.executor.memoryOverhead` covers everything an executor needs outside the JVM heap sized by `spark.executor.memory`: thread stacks, JIT buffers, metaspace, JNI, native libraries, and, if `spark.executor.pyspark.memory` isn't set separately, the memory used by PySpark's per-task Python worker processes[^7]. Left unset, Spark defaults it to the greater of 10% of executor memory or 384MB[^7]; the configuration reference formalizes that floor as `spark.executor.minMemoryOverhead` (default `384m`, since Spark 4.0) and the percentage as `spark.executor.memoryOverheadFactor` (default 0.10, or 0.40 for non-JVM jobs on Kubernetes, since those need more non-JVM heap space)[^4]. The resource manager (YARN's NodeManager, or the Kubernetes kubelet) sizes the container/pod to the sum of `spark.executor.memoryOverhead`, `spark.executor.memory`, `spark.memory.offHeap.size`, and `spark.executor.pyspark.memory`[^4]. If off-heap memory or PySpark worker processes consume memory that isn't reflected in a correspondingly larger `memoryOverhead`, that extra usage still shows up in the process's actual RSS, which the resource manager does track. Once real usage exceeds the container's allocated size, the executor is killed by YARN or OOMKilled by the kubelet, with no Spark-level error, just an abrupt process death[^7].
|
|
60
|
+
|
|
61
|
+
**Limitations / false-positive risk:** the default overhead is adequate for many JVM-only jobs, so a value left at the default only signals trouble on off-heap-heavy or PySpark workloads that need more non-heap room.
|
|
62
|
+
|
|
63
|
+
## JVM GC flags
|
|
64
|
+
|
|
65
|
+
| Flag | Effect |
|
|
66
|
+
|---|---|
|
|
67
|
+
| `-XX:+UseG1GC` | Default collector since Spark 4.0.0, which defaults to JDK 17[^6]. |
|
|
68
|
+
| `-XX:G1HeapRegionSize` | May need raising alongside large executor heaps[^6]. |
|
|
69
|
+
| `-XX:InitiatingHeapOccupancyPercent` | Tuned (with `-XX:ConcGCThreads` and RSet-update settings) to fix a documented ~100-second G1 full-GC pause on an 88GB executor heap[^12]. |
|
|
70
|
+
| `-XX:ConcGCThreads` | See `-XX:InitiatingHeapOccupancyPercent` above[^12]. |
|
|
71
|
+
| `-XX:+UseZGC` | Concurrent, low-latency collector; sub-millisecond pause target independent of heap size, from a few hundred megabytes up to 16TB[^13]. |
|
|
72
|
+
|
|
73
|
+
### G1 vs. ZGC
|
|
74
|
+
|
|
75
|
+
G1 was designed as a CMS replacement aiming at both throughput and low latency: it partitions the heap into equal-sized regions and copies out only the live objects from collected regions rather than compacting the whole heap[^12]. Even so, a documented 88GB-heap Spark benchmark hit "unacceptable full GC," with one job pausing nearly 100 seconds under default G1 settings. Only after tuning `InitiatingHeapOccupancyPercent`, `ConcGCThreads`, and RSet-update parameters did G1 beat Parallel/CMS GC on both throughput and latency[^12]. ZGC, per the OpenJDK project page, "performs all expensive work concurrently, without stopping the execution of application threads for more than a millisecond," with pause times that stay flat as heap size scales from a few hundred megabytes up to 16TB[^13].
|
|
76
|
+
|
|
77
|
+
## Sources
|
|
78
|
+
|
|
79
|
+
[^1]: [Performance Tuning — Spark SQL, DataFrames and Datasets Guide](https://spark.apache.org/docs/latest/sql-performance-tuning.html)
|
|
80
|
+
[^2]: *Spark: The Definitive Guide*, Chambers & Zaharia, ch. 4, ch. 10
|
|
81
|
+
[^3]: *Learning Spark, 2nd Edition*, Damji, Wenig, Das & Lee, ch. 7
|
|
82
|
+
[^4]: [Configuration — Spark](https://spark.apache.org/docs/latest/configuration.html)
|
|
83
|
+
[^5]: [SQLConf.scala (Spark 3.5.0)](https://raw.githubusercontent.com/apache/spark/v3.5.0/sql/catalyst/src/main/scala/org/apache/spark/sql/internal/SQLConf.scala)
|
|
84
|
+
[^6]: [Tuning Spark](https://spark.apache.org/docs/latest/tuning.html)
|
|
85
|
+
[^7]: [Dive Into Spark Memory Management](https://luminousmen.com/post/dive-into-spark-memory)
|
|
86
|
+
[^8]: [Apache Arrow in PySpark](https://spark.apache.org/docs/3.5.8/api/python/user_guide/sql/arrow_pandas.html)
|
|
87
|
+
[^9]: [Adaptive Query Execution: Speeding Up Spark SQL at Runtime](https://www.databricks.com/blog/2020/05/29/adaptive-query-execution-speeding-up-spark-sql-at-runtime.html)
|
|
88
|
+
[^10]: *High Performance Spark, 2nd Edition*, Karau, Polak & Warren, ch. 5
|
|
89
|
+
[^11]: [task_memory_management_in_spark.md](https://raw.githubusercontent.com/spoddutur/spark-notes/master/task_memory_management_in_spark.md)
|
|
90
|
+
[^12]: [Tuning Java Garbage Collection for Spark Applications](https://www.databricks.com/blog/2015/05/28/tuning-java-garbage-collection-for-spark-applications.html)
|
|
91
|
+
[^13]: [ZGC — The Z Garbage Collector](https://wiki.openjdk.org/display/zgc)
|
|
92
|
+
[^14]: [Job Scheduling — Spark](https://spark.apache.org/docs/latest/job-scheduling.html)
|
|
93
|
+
[^15]: [The Apache Spark Optimization Checklist](https://luminousmen.com/post/the-apache-spark-optimization-checklist)
|