sparkforensics-cli 0.1.0 → 0.2.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +6 -0
- package/bin/sparkforensics-analyze.mjs +113 -48
- package/export-template/docs/404.html +25 -0
- package/export-template/docs/assets/app.CndaAS6v.js +1 -0
- package/export-template/docs/assets/aqe-loop.IwQSATHw.svg +1 -0
- package/export-template/docs/assets/aqe-loop.dark.DGbaxqJE.svg +1 -0
- package/export-template/docs/assets/broadcast-vs-shuffle.Db4WY1XK.svg +1 -0
- package/export-template/docs/assets/broadcast-vs-shuffle.dark.C7Bxs0mG.svg +1 -0
- package/export-template/docs/assets/cache-lifecycle.dark.B-hS7AgU.svg +1 -0
- package/export-template/docs/assets/cache-lifecycle.rEOVYQNU.svg +1 -0
- package/export-template/docs/assets/chunks/@localSearchIndexroot.DNY8bVcl.js +1 -0
- package/export-template/docs/assets/chunks/VPLocalSearchBox.yJbZbsEo.js +9 -0
- package/export-template/docs/assets/chunks/duplicate-plan-subtree.dark.Cdp70QhV.js +1 -0
- package/export-template/docs/assets/chunks/framework.DSg0KOwT.js +20 -0
- package/export-template/docs/assets/chunks/retry-escalation-ladder.dark.DHipdJgZ.js +1 -0
- package/export-template/docs/assets/chunks/theme.Df2VAG9w.js +2 -0
- package/export-template/docs/assets/cold-start-timeline.DxC_Sc7w.svg +1 -0
- package/export-template/docs/assets/cold-start-timeline.dark.CZ17YcAG.svg +1 -0
- package/export-template/docs/assets/columnar-layout.PghGeOEA.svg +1 -0
- package/export-template/docs/assets/columnar-layout.dark.BVNlz0ff.svg +1 -0
- package/export-template/docs/assets/container-memory.DIO0AnIm.svg +1 -0
- package/export-template/docs/assets/container-memory.dark.CP-5zuCl.svg +1 -0
- package/export-template/docs/assets/contributor-guide_architecture_board-widgets.md.B-OsL91z.js +1 -0
- package/export-template/docs/assets/contributor-guide_architecture_board-widgets.md.B-OsL91z.lean.js +1 -0
- package/export-template/docs/assets/contributor-guide_architecture_detector-contract.md.BOeH4d1J.js +1 -0
- package/export-template/docs/assets/contributor-guide_architecture_detector-contract.md.BOeH4d1J.lean.js +1 -0
- package/export-template/docs/assets/contributor-guide_architecture_drill-down.md.BtPdlM7r.js +1 -0
- package/export-template/docs/assets/contributor-guide_architecture_drill-down.md.BtPdlM7r.lean.js +1 -0
- package/export-template/docs/assets/contributor-guide_architecture_impact-estimation.md.DYCDPgkh.js +1 -0
- package/export-template/docs/assets/contributor-guide_architecture_impact-estimation.md.DYCDPgkh.lean.js +1 -0
- package/export-template/docs/assets/contributor-guide_architecture_index.md.3TO9ic6w.js +1 -0
- package/export-template/docs/assets/contributor-guide_architecture_index.md.3TO9ic6w.lean.js +1 -0
- package/export-template/docs/assets/contributor-guide_architecture_overview.md.CehiRmGn.js +1 -0
- package/export-template/docs/assets/contributor-guide_architecture_overview.md.CehiRmGn.lean.js +1 -0
- package/export-template/docs/assets/contributor-guide_architecture_state-and-history.md.m3S3UdMk.js +1 -0
- package/export-template/docs/assets/contributor-guide_architecture_state-and-history.md.m3S3UdMk.lean.js +1 -0
- package/export-template/docs/assets/contributor-guide_architecture_widget-rendering.md.DbqPf2OT.js +1 -0
- package/export-template/docs/assets/contributor-guide_architecture_widget-rendering.md.DbqPf2OT.lean.js +1 -0
- package/export-template/docs/assets/contributor-guide_architecture_worker-protocol.md.B93qJ_tT.js +6 -0
- package/export-template/docs/assets/contributor-guide_architecture_worker-protocol.md.B93qJ_tT.lean.js +1 -0
- package/export-template/docs/assets/contributor-guide_contributing.md.CvRsdr6J.js +1 -0
- package/export-template/docs/assets/contributor-guide_contributing.md.CvRsdr6J.lean.js +1 -0
- package/export-template/docs/assets/contributor-guide_development-setup.md.DvAN_9mK.js +12 -0
- package/export-template/docs/assets/contributor-guide_development-setup.md.DvAN_9mK.lean.js +1 -0
- package/export-template/docs/assets/contributor-guide_testing.md.6rIKqSyY.js +1 -0
- package/export-template/docs/assets/contributor-guide_testing.md.6rIKqSyY.lean.js +1 -0
- package/export-template/docs/assets/dag-stages.DSz_S937.svg +1 -0
- package/export-template/docs/assets/dag-stages.dark.F72UzxH4.svg +1 -0
- package/export-template/docs/assets/driver-executor.D5pQ7YN1.svg +1 -0
- package/export-template/docs/assets/driver-executor.dark.BmX9cPvh.svg +1 -0
- package/export-template/docs/assets/duplicate-plan-subtree.B4cvN6fj.svg +1 -0
- package/export-template/docs/assets/duplicate-plan-subtree.dark.Dw8wS0Ag.svg +1 -0
- package/export-template/docs/assets/index.md.CHJVslga.js +1 -0
- package/export-template/docs/assets/index.md.CHJVslga.lean.js +1 -0
- package/export-template/docs/assets/inter-italic-cyrillic-ext.r48I6akx.woff2 +0 -0
- package/export-template/docs/assets/inter-italic-cyrillic.By2_1cv3.woff2 +0 -0
- package/export-template/docs/assets/inter-italic-greek-ext.1u6EdAuj.woff2 +0 -0
- package/export-template/docs/assets/inter-italic-greek.DJ8dCoTZ.woff2 +0 -0
- package/export-template/docs/assets/inter-italic-latin-ext.CN1xVJS-.woff2 +0 -0
- package/export-template/docs/assets/inter-italic-latin.C2AdPX0b.woff2 +0 -0
- package/export-template/docs/assets/inter-italic-vietnamese.BSbpV94h.woff2 +0 -0
- package/export-template/docs/assets/inter-roman-cyrillic-ext.BBPuwvHQ.woff2 +0 -0
- package/export-template/docs/assets/inter-roman-cyrillic.C5lxZ8CY.woff2 +0 -0
- package/export-template/docs/assets/inter-roman-greek-ext.CqjqNYQ-.woff2 +0 -0
- package/export-template/docs/assets/inter-roman-greek.BBVDIX6e.woff2 +0 -0
- package/export-template/docs/assets/inter-roman-latin-ext.4ZJIpNVo.woff2 +0 -0
- package/export-template/docs/assets/inter-roman-latin.Di8DUHzh.woff2 +0 -0
- package/export-template/docs/assets/inter-roman-vietnamese.BjW4sHH5.woff2 +0 -0
- package/export-template/docs/assets/join-strategy.C_FvrCEo.svg +1 -0
- package/export-template/docs/assets/join-strategy.dark.ChMLnNII.svg +1 -0
- package/export-template/docs/assets/memory-borrowing.BqQRJg0u.svg +1 -0
- package/export-template/docs/assets/memory-borrowing.dark.Yhh20O9C.svg +1 -0
- package/export-template/docs/assets/memory-regions.XHvO7jHG.svg +1 -0
- package/export-template/docs/assets/memory-regions.dark.D4TP9_08.svg +1 -0
- package/export-template/docs/assets/repartition-vs-coalesce.BovLRrpj.svg +1 -0
- package/export-template/docs/assets/repartition-vs-coalesce.dark.BhAczKZQ.svg +1 -0
- package/export-template/docs/assets/retry-escalation-ladder.DyTKJJmZ.svg +1 -0
- package/export-template/docs/assets/retry-escalation-ladder.dark.BdsabtU3.svg +1 -0
- package/export-template/docs/assets/shuffle-map-reduce.KuOEZVmg.svg +1 -0
- package/export-template/docs/assets/shuffle-map-reduce.dark.BgQZnFSb.svg +1 -0
- package/export-template/docs/assets/spill-classification.BU2euYDO.svg +1 -0
- package/export-template/docs/assets/spill-classification.dark.D7i1M40d.svg +1 -0
- package/export-template/docs/assets/style.DXOMCXxn.css +1 -0
- package/export-template/docs/assets/tuning-reference_anti-patterns.md.Df1YMIHu.js +1 -0
- package/export-template/docs/assets/tuning-reference_anti-patterns.md.Df1YMIHu.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_aqe.md.BIsCtLzm.js +1 -0
- package/export-template/docs/assets/tuning-reference_aqe.md.BIsCtLzm.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-broadcast-sizing.md.CEstB3Ia.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-broadcast-sizing.md.CEstB3Ia.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-cold-start.md.CEuy-72y.js +7 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-cold-start.md.CEuy-72y.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-duplicate-plan-subtree.md.CIohQDfn.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-duplicate-plan-subtree.md.CIohQDfn.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-failures.md.4z5BXGJ2.js +6 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-failures.md.4z5BXGJ2.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-gc.md.DSxzZRK7.js +6 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-gc.md.DSxzZRK7.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-job-failure-rate.md.BaJl__1W.js +8 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-job-failure-rate.md.BaJl__1W.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-memory-utilization.md.DbP-SJZc.js +7 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-memory-utilization.md.DbP-SJZc.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-retry-waste.md.D5JMjVOt.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-retry-waste.md.D5JMjVOt.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-shuffle.md.CM-nTmIH.js +12 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-shuffle.md.CM-nTmIH.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-skew.md.BdUwiDhn.js +14 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-skew.md.BdUwiDhn.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-slow-host.md.BlIo6UDW.js +7 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-slow-host.md.BlIo6UDW.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-small-files.md.B8kloyx8.js +5 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-small-files.md.B8kloyx8.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-spill.md.PNH7mITt.js +6 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-spill.md.PNH7mITt.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-straggler.md.DY36fHN5.js +7 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-straggler.md.DY36fHN5.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-tiny-tasks.md.QTV7O8kU.js +8 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-tiny-tasks.md.QTV7O8kU.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-utilization.md.DTiueZC3.js +5 -0
- package/export-template/docs/assets/tuning-reference_bottleneck-utilization.md.DTiueZC3.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_caching.md.B7aQ8asB.js +1 -0
- package/export-template/docs/assets/tuning-reference_caching.md.B7aQ8asB.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_cluster-config.md.ZVmDGsQ3.js +1 -0
- package/export-template/docs/assets/tuning-reference_cluster-config.md.ZVmDGsQ3.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_config.md.UvveiWG3.js +1 -0
- package/export-template/docs/assets/tuning-reference_config.md.UvveiWG3.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_data-formats.md.bjCAWH3N.js +1 -0
- package/export-template/docs/assets/tuning-reference_data-formats.md.bjCAWH3N.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_index.md.BQ_NooMV.js +1 -0
- package/export-template/docs/assets/tuning-reference_index.md.BQ_NooMV.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_intro.md.CobD-lGB.js +1 -0
- package/export-template/docs/assets/tuning-reference_intro.md.CobD-lGB.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_joins.md.BtKs_CuW.js +1 -0
- package/export-template/docs/assets/tuning-reference_joins.md.BtKs_CuW.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_memory-model.md.DhT-n4y3.js +1 -0
- package/export-template/docs/assets/tuning-reference_memory-model.md.DhT-n4y3.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_metrics.md.mLOh7Apj.js +1 -0
- package/export-template/docs/assets/tuning-reference_metrics.md.mLOh7Apj.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_partitioning.md.q0zKF_8X.js +1 -0
- package/export-template/docs/assets/tuning-reference_partitioning.md.q0zKF_8X.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_pyspark.md.DDCfvN9t.js +6 -0
- package/export-template/docs/assets/tuning-reference_pyspark.md.DDCfvN9t.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_shuffle.md.BZZ7R4Ix.js +1 -0
- package/export-template/docs/assets/tuning-reference_shuffle.md.BZZ7R4Ix.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_spark-architecture.md.Dwzm5avO.js +1 -0
- package/export-template/docs/assets/tuning-reference_spark-architecture.md.Dwzm5avO.lean.js +1 -0
- package/export-template/docs/assets/tuning-reference_table-formats.md.D6wj-2dX.js +1 -0
- package/export-template/docs/assets/tuning-reference_table-formats.md.D6wj-2dX.lean.js +1 -0
- package/export-template/docs/assets/udf-execution-models.BUFDICuG.svg +1 -0
- package/export-template/docs/assets/udf-execution-models.dark.YTNS6GDq.svg +1 -0
- package/export-template/docs/assets/user-guide_alternative-log-retrieval.md.sU3KGarf.js +1 -0
- package/export-template/docs/assets/user-guide_alternative-log-retrieval.md.sU3KGarf.lean.js +1 -0
- package/export-template/docs/assets/user-guide_getting-started.md.DtEM37MK.js +3 -0
- package/export-template/docs/assets/user-guide_getting-started.md.DtEM37MK.lean.js +1 -0
- package/export-template/docs/assets/user-guide_mcp-tools.md.C8MiIu7F.js +125 -0
- package/export-template/docs/assets/user-guide_mcp-tools.md.C8MiIu7F.lean.js +1 -0
- package/export-template/docs/assets/user-guide_run-comparison.md.S0TWWmLY.js +1 -0
- package/export-template/docs/assets/user-guide_run-comparison.md.S0TWWmLY.lean.js +1 -0
- package/export-template/docs/assets/user-guide_understanding-findings.md.D0R_Y-R2.js +1 -0
- package/export-template/docs/assets/user-guide_understanding-findings.md.D0R_Y-R2.lean.js +1 -0
- package/export-template/docs/contributor-guide/architecture/board-widgets.html +25 -0
- package/export-template/docs/contributor-guide/architecture/detector-contract.html +25 -0
- package/export-template/docs/contributor-guide/architecture/drill-down.html +25 -0
- package/export-template/docs/contributor-guide/architecture/impact-estimation.html +25 -0
- package/export-template/docs/contributor-guide/architecture/index.html +25 -0
- package/export-template/docs/contributor-guide/architecture/overview.html +25 -0
- package/export-template/docs/contributor-guide/architecture/state-and-history.html +25 -0
- package/export-template/docs/contributor-guide/architecture/widget-rendering.html +25 -0
- package/export-template/docs/contributor-guide/architecture/worker-protocol.html +30 -0
- package/export-template/docs/contributor-guide/contributing.html +25 -0
- package/export-template/docs/contributor-guide/development-setup.html +36 -0
- package/export-template/docs/contributor-guide/testing.html +25 -0
- package/export-template/docs/favicon.svg +4 -0
- package/export-template/docs/hashmap.json +1 -0
- package/export-template/docs/index.html +25 -0
- package/export-template/docs/package.json +1 -0
- package/export-template/docs/tuning-reference/anti-patterns.html +25 -0
- package/export-template/docs/tuning-reference/aqe.html +25 -0
- package/export-template/docs/tuning-reference/bottleneck-broadcast-sizing.html +25 -0
- package/export-template/docs/tuning-reference/bottleneck-cold-start.html +31 -0
- package/export-template/docs/tuning-reference/bottleneck-duplicate-plan-subtree.html +25 -0
- package/export-template/docs/tuning-reference/bottleneck-failures.html +30 -0
- package/export-template/docs/tuning-reference/bottleneck-gc.html +30 -0
- package/export-template/docs/tuning-reference/bottleneck-job-failure-rate.html +32 -0
- package/export-template/docs/tuning-reference/bottleneck-memory-utilization.html +31 -0
- package/export-template/docs/tuning-reference/bottleneck-retry-waste.html +25 -0
- package/export-template/docs/tuning-reference/bottleneck-shuffle.html +36 -0
- package/export-template/docs/tuning-reference/bottleneck-skew.html +38 -0
- package/export-template/docs/tuning-reference/bottleneck-slow-host.html +31 -0
- package/export-template/docs/tuning-reference/bottleneck-small-files.html +29 -0
- package/export-template/docs/tuning-reference/bottleneck-spill.html +30 -0
- package/export-template/docs/tuning-reference/bottleneck-straggler.html +31 -0
- package/export-template/docs/tuning-reference/bottleneck-tiny-tasks.html +32 -0
- package/export-template/docs/tuning-reference/bottleneck-utilization.html +29 -0
- package/export-template/docs/tuning-reference/caching.html +25 -0
- package/export-template/docs/tuning-reference/cluster-config.html +25 -0
- package/export-template/docs/tuning-reference/config.html +25 -0
- package/export-template/docs/tuning-reference/data-formats.html +25 -0
- package/export-template/docs/tuning-reference/index.html +25 -0
- package/export-template/docs/tuning-reference/intro.html +25 -0
- package/export-template/docs/tuning-reference/joins.html +25 -0
- package/export-template/docs/tuning-reference/memory-model.html +25 -0
- package/export-template/docs/tuning-reference/metrics.html +25 -0
- package/export-template/docs/tuning-reference/partitioning.html +25 -0
- package/export-template/docs/tuning-reference/pyspark.html +30 -0
- package/export-template/docs/tuning-reference/shuffle.html +25 -0
- package/export-template/docs/tuning-reference/spark-architecture.html +25 -0
- package/export-template/docs/tuning-reference/table-formats.html +25 -0
- package/export-template/docs/user-guide/alternative-log-retrieval.html +25 -0
- package/export-template/docs/user-guide/getting-started.html +27 -0
- package/export-template/docs/user-guide/mcp-tools.html +149 -0
- package/export-template/docs/user-guide/run-comparison.html +25 -0
- package/export-template/docs/user-guide/understanding-findings.html +25 -0
- package/export-template/docs/vp-icons.css +0 -0
- package/export-template/favicon.svg +4 -0
- package/export-template/index.html +115 -0
- package/export-template/parser-worker-QqyEE4m9.js +64 -0
- package/package.json +16 -3
- package/vendor-core/analyzer.js +74 -74
- package/vendor-core/cli/budgets.js +13 -27
- package/vendor-core/cli/collect-run.js +43 -19
- package/vendor-core/core-count.js +25 -27
- package/vendor-core/core-locality-ratio.js +4 -11
- package/vendor-core/core-time-series.js +6 -12
- package/vendor-core/core-usage-locality.js +3 -4
- package/vendor-core/detectors.js +256 -375
- package/vendor-core/docs-config.js +69 -21
- package/vendor-core/docs-content/chapters/01-intro.md +32 -0
- package/vendor-core/docs-content/chapters/02-spark-architecture.md +76 -0
- package/vendor-core/docs-content/chapters/03-memory-model.md +73 -0
- package/vendor-core/docs-content/chapters/04-partitioning.md +65 -0
- package/vendor-core/docs-content/chapters/05-joins.md +62 -0
- package/vendor-core/docs-content/chapters/06-shuffle.md +59 -0
- package/vendor-core/docs-content/chapters/07-data-formats.md +81 -0
- package/vendor-core/docs-content/chapters/07b-table-formats.md +56 -0
- package/vendor-core/docs-content/chapters/08-caching.md +58 -0
- package/vendor-core/docs-content/chapters/09-pyspark.md +78 -0
- package/vendor-core/docs-content/chapters/10-aqe.md +167 -0
- package/vendor-core/docs-content/chapters/11-cluster-config.md +170 -0
- package/vendor-core/docs-content/chapters/12-anti-patterns.md +171 -0
- package/vendor-core/docs-content/chapters/14-metrics.md +87 -0
- package/vendor-core/docs-content/chapters/15-config.md +93 -0
- package/vendor-core/docs-content/chapters/nav-index.json +370 -0
- package/vendor-core/docs-content/detection/cache.md +6 -0
- package/vendor-core/docs-content/detection/cfg.md +15 -0
- package/vendor-core/docs-content/detection/chrn.md +7 -0
- package/vendor-core/docs-content/detection/cold.md +4 -0
- package/vendor-core/docs-content/detection/cstor.md +4 -0
- package/vendor-core/docs-content/detection/fail.md +5 -0
- package/vendor-core/docs-content/detection/gc.md +4 -0
- package/vendor-core/docs-content/detection/host.md +5 -0
- package/vendor-core/docs-content/detection/incmp.md +6 -0
- package/vendor-core/docs-content/detection/jobs.md +4 -0
- package/vendor-core/docs-content/detection/local.md +7 -0
- package/vendor-core/docs-content/detection/mem.md +10 -0
- package/vendor-core/docs-content/detection/part.md +5 -0
- package/vendor-core/docs-content/detection/plan.md +14 -0
- package/vendor-core/docs-content/detection/retry.md +4 -0
- package/vendor-core/docs-content/detection/sfail.md +5 -0
- package/vendor-core/docs-content/detection/shape.md +5 -0
- package/vendor-core/docs-content/detection/shfl.md +4 -0
- package/vendor-core/docs-content/detection/skew.md +6 -0
- package/vendor-core/docs-content/detection/slow.md +6 -0
- package/vendor-core/docs-content/detection/spec.md +7 -0
- package/vendor-core/docs-content/detection/spill.md +7 -0
- package/vendor-core/docs-content/detection/strag.md +5 -0
- package/vendor-core/docs-content/detection/tiny.md +4 -0
- package/vendor-core/docs-content/detection/util.md +4 -0
- package/vendor-core/docs-content/diagrams/aqe-loop.dark.svg +1 -0
- package/vendor-core/docs-content/diagrams/aqe-loop.svg +1 -0
- package/vendor-core/docs-content/diagrams/broadcast-vs-shuffle.dark.svg +1 -0
- package/vendor-core/docs-content/diagrams/broadcast-vs-shuffle.svg +1 -0
- package/vendor-core/docs-content/diagrams/cache-lifecycle.dark.svg +1 -0
- package/vendor-core/docs-content/diagrams/cache-lifecycle.svg +1 -0
- package/vendor-core/docs-content/diagrams/cold-start-timeline.dark.svg +1 -0
- package/vendor-core/docs-content/diagrams/cold-start-timeline.svg +1 -0
- package/vendor-core/docs-content/diagrams/columnar-layout.dark.svg +1 -0
- package/vendor-core/docs-content/diagrams/columnar-layout.svg +1 -0
- package/vendor-core/docs-content/diagrams/container-memory.dark.svg +1 -0
- package/vendor-core/docs-content/diagrams/container-memory.svg +1 -0
- package/vendor-core/docs-content/diagrams/dag-stages.dark.svg +1 -0
- package/vendor-core/docs-content/diagrams/dag-stages.svg +1 -0
- package/vendor-core/docs-content/diagrams/driver-executor.dark.svg +1 -0
- package/vendor-core/docs-content/diagrams/driver-executor.svg +1 -0
- package/vendor-core/docs-content/diagrams/duplicate-plan-subtree.dark.svg +1 -0
- package/vendor-core/docs-content/diagrams/duplicate-plan-subtree.svg +1 -0
- package/vendor-core/docs-content/diagrams/join-strategy.dark.svg +1 -0
- package/vendor-core/docs-content/diagrams/join-strategy.svg +1 -0
- package/vendor-core/docs-content/diagrams/memory-borrowing.dark.svg +1 -0
- package/vendor-core/docs-content/diagrams/memory-borrowing.svg +1 -0
- package/vendor-core/docs-content/diagrams/memory-regions.dark.svg +1 -0
- package/vendor-core/docs-content/diagrams/memory-regions.svg +1 -0
- package/vendor-core/docs-content/diagrams/repartition-vs-coalesce.dark.svg +1 -0
- package/vendor-core/docs-content/diagrams/repartition-vs-coalesce.svg +1 -0
- package/vendor-core/docs-content/diagrams/retry-escalation-ladder.dark.svg +1 -0
- package/vendor-core/docs-content/diagrams/retry-escalation-ladder.svg +1 -0
- package/vendor-core/docs-content/diagrams/shuffle-map-reduce.dark.svg +1 -0
- package/vendor-core/docs-content/diagrams/shuffle-map-reduce.svg +1 -0
- package/vendor-core/docs-content/diagrams/spill-classification.dark.svg +1 -0
- package/vendor-core/docs-content/diagrams/spill-classification.svg +1 -0
- package/vendor-core/docs-content/diagrams/udf-execution-models.dark.svg +1 -0
- package/vendor-core/docs-content/diagrams/udf-execution-models.svg +1 -0
- package/vendor-core/docs-content/tuning/broadcast-sizing.md +78 -0
- package/vendor-core/docs-content/tuning/cold-start.md +81 -0
- package/vendor-core/docs-content/tuning/duplicate-plan-subtree.md +45 -0
- package/vendor-core/docs-content/tuning/failures.md +124 -0
- package/vendor-core/docs-content/tuning/gc.md +110 -0
- package/vendor-core/docs-content/tuning/job-failure-rate.md +101 -0
- package/vendor-core/docs-content/tuning/memory-utilization.md +58 -0
- package/vendor-core/docs-content/tuning/retry-waste.md +90 -0
- package/vendor-core/docs-content/tuning/shuffle.md +154 -0
- package/vendor-core/docs-content/tuning/skew.md +123 -0
- package/vendor-core/docs-content/tuning/slow-host.md +117 -0
- package/vendor-core/docs-content/tuning/small-files.md +99 -0
- package/vendor-core/docs-content/tuning/spill.md +114 -0
- package/vendor-core/docs-content/tuning/straggler.md +103 -0
- package/vendor-core/docs-content/tuning/tiny-tasks.md +94 -0
- package/vendor-core/docs-content/tuning/utilization.md +90 -0
- package/vendor-core/docs-site-config.js +10 -17
- package/vendor-core/efficiency-model.js +7 -13
- package/vendor-core/etl-phases.js +3 -5
- package/vendor-core/event-handlers.js +232 -134
- package/vendor-core/event-schemas.js +48 -114
- package/vendor-core/evidence-availability.js +5 -10
- package/vendor-core/evidence-report.js +72 -122
- package/vendor-core/export-data.js +48 -0
- package/vendor-core/finding-action-label.js +4 -10
- package/vendor-core/finding-filter-predicate.js +3 -7
- package/vendor-core/finding-generic-recommendation.js +112 -0
- package/vendor-core/finding-names.js +51 -0
- package/vendor-core/format-utils.js +112 -38
- package/vendor-core/impact-band.js +18 -24
- package/vendor-core/impact-estimator.js +38 -74
- package/vendor-core/ingest.js +7 -13
- package/vendor-core/job-groups.js +3 -6
- package/vendor-core/list-runs.js +278 -0
- package/vendor-core/load-vendored.js +6 -12
- package/vendor-core/log-header-peek.js +81 -0
- package/vendor-core/lz4-block.js +4 -6
- package/vendor-core/mcp-server-factory.js +38 -8
- package/vendor-core/mcp-tools.js +105 -76
- package/vendor-core/model-assembler.js +8 -16
- package/vendor-core/occupancy.js +5 -9
- package/vendor-core/parser-worker.js +18 -27
- package/vendor-core/plan-dot.js +2 -5
- package/vendor-core/plan-duration-attribution.js +78 -29
- package/vendor-core/plan-graph-model.js +126 -69
- package/vendor-core/plan-node-detail.js +31 -17
- package/vendor-core/plan-summary.js +19 -8
- package/vendor-core/recommendation-rollup.js +35 -39
- package/vendor-core/redact.js +72 -16
- package/vendor-core/rolling-log-reassembly.js +4 -6
- package/vendor-core/run-comparison.js +65 -70
- package/vendor-core/scaling-sim.js +5 -7
- package/vendor-core/session-snapshot.js +1 -1
- package/vendor-core/shs-fetch.js +4 -6
- package/vendor-core/shs-load.js +9 -13
- package/vendor-core/shs-request.js +1 -1
- package/vendor-core/stage-quantiles.js +14 -0
- package/vendor-core/types.js +78 -18
- package/vendor-core/wasted-core-hours.js +7 -12
|
@@ -0,0 +1,81 @@
|
|
|
1
|
+
# Data Formats
|
|
2
|
+
|
|
3
|
+
## How Parquet and ORC lay out data
|
|
4
|
+
|
|
5
|
+
Parquet and ORC are both self-describing, columnar file formats for storing Spark's structured data[^1]. They're closer than they are different: *Spark: The Definitive Guide* frames the choice as "for the most part, they're quite similar; the fundamental difference is that Parquet is further optimized for use with Spark, whereas ORC is further optimized for Hive"[^1], and notes that ORC "has no options for reading in data because Spark understands the file format quite well"[^1]: Spark treats it as a format it understands natively rather than one with extra read-side knobs.
|
|
6
|
+
|
|
7
|
+
Both formats group rows into chunks (row groups in Parquet, stripes in ORC), each holding every column's data for that slice, with a footer or file-level metadata section recording where each chunk lives on disk. Parquet's `FileMetaData` footer records the offset and size of every row group and column chunk, and "readers are expected to first read the file metadata to find all the column chunks they are interested in. The column chunks should then be read sequentially"[^2]. Because compression is applied per column chunk rather than as one continuous stream over the whole file, a reader can seek straight to a row group's footer-recorded offset and decode just that chunk, which is exactly why Parquet files compressed with Snappy stay splittable even though the Snappy stream format itself provides no split points: it has "no entropy encoder backend nor framing layer -- the latter is assumed to be handled by other parts of the system"[^3].
|
|
8
|
+
|
|
9
|
+
<img class="light-only" src="diagrams/columnar-layout.svg" alt="A columnar file's footer holds row-group offsets and per-column min/max statistics that let a reader skip row groups whose stats exclude the predicate.">
|
|
10
|
+
<img class="dark-only" src="diagrams/columnar-layout.dark.svg" alt="A columnar file's footer holds row-group offsets and per-column min/max statistics that let a reader skip row groups whose stats exclude the predicate.">
|
|
11
|
+
|
|
12
|
+
Both formats also carry the statistics that later drive predicate pushdown. Parquet records min/max values per column chunk, plus an optional page-level `ColumnIndex` that lets a reader binary-search ordered columns for matching pages[^4], and optional Bloom filters for columns whose cardinality is too high for a dictionary to be practical[^5]. ORC's `ColumnStatistics` protobuf records row count and, for most primitive types, min/max (plus sum for numeric types); from Hive 1.1.0 it also records a `hasNull` flag used specifically by ORC's predicate pushdown to answer `IS NULL` queries[^6]. ORC layers this at two granularities: a `RowIndexEntry` per row group (10,000 rows by default), kept at the front of each stripe so it's read only when pushdown or seeking is actually needed[^7], and file-level `StripeStatistics` that let whole stripes be skipped by predicate pushdown[^8]. ORC also supports Bloom filters (Hive 1.2.0+), but only evaluates them against row groups that already passed the min/max row-index check first: a second-stage filter, not an independent one[^7].
|
|
13
|
+
|
|
14
|
+
On compression: Parquet defaults to Snappy, while ORC's default has been Zstd since Spark 2.3.0[^9]. Zstandard's own manual describes it as "a fast lossless compression algorithm, targeting real-time compression scenarios at zlib-level and better compression ratios," offering regular levels 1–22 plus negative levels that trade ratio for speed[^10]. Spark exposes the level directly through `spark.io.compression.zstd.level` (default `1`)[^9], along with `spark.io.compression.zstd.workers` for parallel compression threads (default `0`, since 4.0.0) and `spark.io.compression.zstd.bufferSize` (default 32k)[^9].
|
|
15
|
+
|
|
16
|
+
Parquet has one more structural knob: `parquet.writer.version`. Version 2 changes the on-disk page-header encoding, and the spec is explicit that this is a forward-*incompatible* change: "a reader that only understands `DataPageHeader` cannot parse `DataPageHeaderV2` pages"[^11]. The spec also flags that the file's own version marker "has historically been used inconsistently: writers populate 1 or 2 without a consistent relationship to the features actually used"[^11], so the marker itself isn't a reliable way to tell which page format a file actually contains.
|
|
17
|
+
|
|
18
|
+
## What shows up at read time
|
|
19
|
+
|
|
20
|
+
A file's layout shapes the job the moment Spark reads it. Spark's file readers derive partition counts from file metadata rather than from a fixed rule: "for structured formats like Parquet, ORC, or Avro, Spark actually reads the metadata (footers, row groups, that kind of thing) and tries to slice the file in a way that makes sense. Often you'll see one partition per row group, though Spark may merge or split depending on file sizes and configs"[^12]. A concrete case: a single 30 GB Parquet file with 300 row groups gets split into exactly 300 partitions[^12]. So a surprising partition count for a given file is usually explained by its row-group layout, separately from `spark.sql.files.maxPartitionBytes` (128 MB by default) and its `maxPartitionNum`/`minPartitionNum` companions, which govern splitting for file-based sources generally[^13].
|
|
21
|
+
|
|
22
|
+
The [small-files problem](#bottleneck-small-files) shows up as metadata overhead rather than raw I/O cost: "when you're writing lots of small files, there's a significant metadata overhead that you incur managing all of those files. Spark especially does not do well with small files"[^1]. At the other extreme, oversized or misaligned row groups show up as lost locality: the Parquet spec's rationale for matching block size to row-group size is that "since an entire row group might need to be read, we want it to completely fit on one HDFS block"[^14]. A row group straddling a block boundary can't get that benefit.
|
|
23
|
+
|
|
24
|
+
Raw (non-Parquet) gzip or zip source files reveal themselves as a single-executor bottleneck at read time, visible as one task doing dramatically more work than the rest of its stage: Spark "needs to download the whole file on one executor, unpack it on just one core, and then redistribute the partitions to the cluster nodes"[^15].
|
|
25
|
+
|
|
26
|
+
Version and schema mismatches surface as read failures rather than silent corruption. A file written with `parquet.writer.version=2` can't be parsed by an older Spark version, or a non-Spark engine such as Hive or Impala, if that reader only implements the original `DataPageHeader`[^11]. *Spark: The Definitive Guide* frames this as a general risk: "you can still encounter problems if you're working with incompatible Parquet files. Be careful when you write out Parquet files with different versions of Spark (especially older ones) because this can cause significant headache"[^1].
|
|
27
|
+
|
|
28
|
+
## What the layout buys you
|
|
29
|
+
|
|
30
|
+
Those read-time symptoms trace back to specific tradeoffs in the layout. Splittability is what decides whether a file can be processed in parallel at all. A non-splittable raw compressed source file forces Spark onto a single core for the whole download-and-unpack step before it can redistribute anything: "as you can imagine, this becomes a huge bottleneck in your distributed processing"[^15]. Parquet sidesteps this at the container level: because its footer records row-group and column-chunk offsets independently of whichever codec compressed the bytes inside them, a task can seek and decode a row group on its own, so even a non-splittable codec like Snappy doesn't cost the file its parallelism[^2][^3].
|
|
31
|
+
|
|
32
|
+
Row-group and block-size alignment governs the same kind of locality at a coarser grain. The spec recommends large row groups, 512 MB–1 GB, "because larger row groups allow for larger column chunks which makes it possible to do larger sequential IO," at the cost of more write-side buffering, paired with a block size sized the same way; its own worked example is "1GB row groups, 1GB HDFS block size, 1 HDFS block per HDFS file"[^14].
|
|
33
|
+
|
|
34
|
+
The statistics both formats carry are what predicate pushdown runs against. Spark plans queries directly off "statistics that Spark reads directly from the underlying data source, like the counts and min/max values in the metadata of Parquet files"[^16], gated by `spark.sql.parquet.filterPushdown` (default `true` since Spark 1.2.0) and, for ORC, `spark.sql.orc.filterPushdown` (default `true` since 1.4.0)[^17][^18]. A more aggressive option, `spark.sql.parquet.aggregatePushdown` (default `false`, since 3.3.0), pushes `MIN`, `MAX`, and `COUNT` down to Parquet's footer statistics directly, and throws if the needed statistic is missing from a file's footer[^17]. Skipping a row group or stripe this way means its bytes are never read off disk, which beats any amount of post-read filtering.
|
|
35
|
+
|
|
36
|
+
Small files cost more in metadata management than their data volume would suggest, and Spark "especially does not do well" with them[^1]. On the scheduling side, Spark's tuning guide notes it "can efficiently support tasks as short as 200 ms" because it reuses one executor JVM across many tasks and has low task-launch cost, while separately recommending "2-3 tasks per CPU core in your cluster"[^19], which implies [scheduling overhead](#bottleneck-tiny-tasks) only becomes proportionally significant once tasks (and the files behind them) shrink below roughly that 200 ms floor.
|
|
37
|
+
|
|
38
|
+
Compression choice trades CPU against I/O and storage. Once I/O stops being the bottleneck, paying a codec's decompression cost is a net loss: "uncompressed files are clearly outperforming compressed files. This is because uncompressed files are I/O bound, and compressed files are CPU bound, but I/O is good enough here"[^15]. Bzip2 illustrates the opposite failure mode: it's splittable, but compresses so aggressively that "you get very few partitions and therefore they can be poorly distributed"[^15].
|
|
39
|
+
|
|
40
|
+
Writer-version and schema changes matter because they fail at read time, for whoever reads the data next (not at write time, for whoever wrote it). A `parquet.writer.version=2` file is simply unreadable by any engine that only understands the original page header[^11]. Schema drift has its own correctness angle: for [Delta Lake](#table-formats) tables, adding a column via `mergeSchema` causes existing rows, when read back, to have that new column's value read as `NULL`[^20], a defined outcome, but one that changes what a downstream query sees for rows written before the schema changed.
|
|
41
|
+
|
|
42
|
+
## Choosing and tuning the format
|
|
43
|
+
|
|
44
|
+
A few defaults cover most cases despite those tradeoffs. Default to Parquet. Reach for ORC specifically when the same data also has to serve Hive consumers or existing Hive ORC tables, where ORC is the better-optimized target[^1].
|
|
45
|
+
|
|
46
|
+
Size row groups and the underlying filesystem block together, not independently. The Parquet spec's own recommended setup is 1 GB row groups paired with a 1 GB HDFS block size, one block per file[^14]. Leave `spark.sql.files.maxPartitionBytes` (128 MB default) and its `maxPartitionNum`/`minPartitionNum` companions for the general file-splitting case, but expect row-group boundaries (not these settings) to be what actually decides partition count for row-group-oriented formats[^13][^12].
|
|
47
|
+
|
|
48
|
+
Leave predicate pushdown on: `spark.sql.parquet.filterPushdown` and `spark.sql.orc.filterPushdown` both default to `true` already[^17][^18], so the main action is not disabling them, plus considering `spark.sql.parquet.aggregatePushdown` when a workload is dominated by `MIN`/`MAX`/`COUNT` over Parquet sources with complete footer statistics[^17].
|
|
49
|
+
|
|
50
|
+
Cap output file size directly instead of letting the small-files problem accumulate: `maxRecordsPerFile`, introduced in Spark 2.2, targets an optimum file size by capping the number of records written per file, e.g. `df.write.option("maxRecordsPerFile", 5000)`[^1].
|
|
51
|
+
|
|
52
|
+
Pick a compression codec based on the actual bottleneck. Parquet's default, Snappy, favors speed; ORC has already defaulted to Zstd since Spark 2.3.0[^9], and at least one optimization checklist recommends overriding Parquet's default to Zstd as well[^21]. Zstd's level is tunable through `spark.io.compression.zstd.level`: higher levels buy better compression "at the expense of more CPU and memory"[^9]. Avoid feeding Spark large raw `.gz`/`.zip` source files directly (unpack them before loading), since the non-splittability lives in the raw container format rather than in gzip itself; the same *Definitive Guide* that warns about raw gzip files elsewhere still recommends Parquet with gzip compression once gzip is wrapped inside Parquet's row-group container[^1].
|
|
53
|
+
|
|
54
|
+
Treat `spark.sql.parquet.mergeSchema` (default `false`) as opt-in rather than default-on: enabling it "merges schemas collected from all data files," instead of trusting a single summary file or a random file's schema[^22], useful for evolving schemas, but it means every file in the dataset gets scanned for its schema. For Delta tables specifically, remember that `mergeSchema`-added columns read back as `NULL` on pre-existing rows[^20] before relying on it for a backfill.
|
|
55
|
+
|
|
56
|
+
Don't flip `parquet.writer.version` to `2` without confirming every downstream reader of that data (an older Spark version, Hive, Impala, or anything else touching the files) actually supports `DataPageHeaderV2` first; it's a forward-incompatible page-format change, not an additive one[^11][^1].
|
|
57
|
+
|
|
58
|
+
## Sources
|
|
59
|
+
|
|
60
|
+
[^1]: *Spark: The Definitive Guide*, Chambers & Zaharia, ch. 9
|
|
61
|
+
[^2]: [Parquet File Format: File Metadata and Row Groups](https://parquet.apache.org/_print/docs/file-format/)
|
|
62
|
+
[^3]: [Snappy Compressed Format Description](https://github.com/google/snappy/blob/main/format_description.txt)
|
|
63
|
+
[^4]: [Parquet File Format: Column Index](https://parquet.apache.org/_print/docs/file-format/)
|
|
64
|
+
[^5]: [Parquet File Format: Bloom Filter](https://parquet.apache.org/docs/file-format/bloomfilter/)
|
|
65
|
+
[^6]: [ORC Specification v1: Column Statistics](https://orc.apache.org/specification/ORCv1/)
|
|
66
|
+
[^7]: [ORC Specification v1: Row Index and Bloom Filters](https://orc.apache.org/specification/ORCv1/)
|
|
67
|
+
[^8]: [ORC Specification v1: Stripe Statistics](https://orc.apache.org/specification/ORCv1/)
|
|
68
|
+
[^9]: [Configuration: Spark](https://spark.apache.org/docs/latest/configuration.html)
|
|
69
|
+
[^10]: [Zstandard Manual](https://facebook.github.io/zstd/zstd_manual.html)
|
|
70
|
+
[^11]: [Parquet File Format: Data Pages (V1/V2)](https://parquet.apache.org/_print/docs/file-format/)
|
|
71
|
+
[^12]: [How Spark Determines Partitions for a File](https://luminousmen.com/post/spark-partitions)
|
|
72
|
+
[^13]: [Performance Tuning: Spark SQL, DataFrames and Datasets Guide](https://spark.apache.org/docs/latest/sql-performance-tuning.html)
|
|
73
|
+
[^14]: [Parquet File Format: Row Group Size](https://parquet.apache.org/_print/docs/file-format/)
|
|
74
|
+
[^15]: [Spark Tips: Don't Collect Data on Driver](https://luminousmen.com/post/spark-tips-dont-collect-data-on-driver)
|
|
75
|
+
[^16]: [Performance Tuning: Spark SQL, DataFrames and Datasets Guide](https://spark.apache.org/docs/latest/sql-performance-tuning.html)
|
|
76
|
+
[^17]: [Configuration: Spark](https://spark.apache.org/docs/latest/configuration.html)
|
|
77
|
+
[^18]: [Configuration: Spark](https://spark.apache.org/docs/latest/configuration.html)
|
|
78
|
+
[^19]: [Spark Tuning Guide: Level of Parallelism](https://spark.apache.org/docs/latest/tuning.html)
|
|
79
|
+
[^20]: *Learning Spark, 2nd Edition*, Damji, Wenig, Das & Lee, ch. 9
|
|
80
|
+
[^21]: [The Apache Spark Optimization Checklist](https://luminousmen.com/post/the-apache-spark-optimization-checklist)
|
|
81
|
+
[^22]: [SQLConf.scala](https://raw.githubusercontent.com/apache/spark/v3.5.0/sql/catalyst/src/main/scala/org/apache/spark/sql/internal/SQLConf.scala)
|
|
@@ -0,0 +1,56 @@
|
|
|
1
|
+
# Table Formats
|
|
2
|
+
|
|
3
|
+
## The knobs these formats share
|
|
4
|
+
|
|
5
|
+
[Parquet and ORC](#data-formats) stop at the file. A lakehouse *table format* wraps a directory of Parquet files with a transactional metadata layer that tracks which files belong to the table right now, so engines get ACID commits, time travel, schema evolution, and row-level updates over plain columnar files. The three in wide use are Delta Lake, Apache Iceberg, and Apache Hudi. They solve the same problems and differ mostly in mechanism, and the tuning knobs that matter fall into a handful of dimensions.
|
|
6
|
+
|
|
7
|
+
**File sizing.** All three fight the small-files problem, but at different points in the write. Delta Lake does not size files during the write by default; its `OPTIMIZE` command bin-packs already-written small files into larger ones, available since Delta Lake 1.2.0[^1]. Iceberg rewrites files after the fact through its `rewrite_data_files` procedure, whose default `binpack` strategy coalesces small files[^2]. Hudi treats correct sizing at write time as "a critical design decision": auto-sizing targets a 120 MB Parquet base file (`hoodie.parquet.max.file.size`) and pads any existing file at or below the 100 MB small-file limit (`hoodie.parquet.small.file.limit`) with new records rather than opening a fresh file[^3].
|
|
8
|
+
|
|
9
|
+
**Compaction.** Delta's `OPTIMIZE` bin-packing is the compaction path, run manually or, since Delta Lake 3.1.0, as auto compaction that runs synchronously right after a write succeeds[^1]. Iceberg uses `rewrite_data_files` as a stored procedure invoked from Spark SQL[^2]. Hudi splits the job in two: *compaction* applies only to Merge-on-Read tables and merges the row-based delta logs back into base files (async by default, or inline via `hoodie.compact.inline = true`), while *clustering* is a separate data-layout service that stitches small files together[^4][^5].
|
|
10
|
+
|
|
11
|
+
**Clustering.** To co-locate rows that are queried together, Delta Lake offers Z-Ordering via `OPTIMIZE table ZORDER BY (cols)`; its effectiveness drops with each added column and it is not idempotent, since each run reclusters all files in the partition[^1]. Delta also documents liquid clustering as a separate feature[^1]. Iceberg clusters through `rewrite_data_files` with `strategy => 'sort'` and a `sort_order` argument, including a `zorder(c1,c2)` form[^2]. Hudi clustering rewrites file groups sorting by `hoodie.clustering.plan.strategy.sort.columns`[^5].
|
|
12
|
+
|
|
13
|
+
**Metadata and manifest overhead.** Delta records each change as a JSON commit in the transaction log and periodically compacts those commits into a Parquet checkpoint so readers reconstruct state without replaying every commit; checkpoints can be split multi-part (default 50,000 actions per part), and since Delta 3.0 log-compaction files aggregate a commit range to cut checkpoint frequency[^1]. Iceberg writes a new metadata JSON file per change, tracks the manifests for each snapshot in a manifest list written fresh on every commit, and splits large tables across multiple manifests so query planning parallelizes[^6]. Hudi records every write and table-service action as an instant on a timeline, with clustering and compaction writing their plans there before executing[^5][^4].
|
|
14
|
+
|
|
15
|
+
**Snapshot and version expiry.** Delta reclaims storage with `VACUUM`, which deletes data files past a retention threshold (default 7 days) but never log files; log files are pruned automatically after checkpoints, with a 30-day default set by `delta.logRetentionDuration`[^7]. Iceberg's `expire_snapshots` removes old snapshots and the files only they referenced (`older_than` default 5 days, `retain_last` default 1), and `remove_orphan_files` sweeps unreferenced files (`older_than` default 3 days)[^2]. Hudi's cleaner runs automatically after each commit; the default `KEEP_LATEST_COMMITS` policy retains `hoodie.clean.commits.retained` commits (default 10)[^8].
|
|
16
|
+
|
|
17
|
+
**Row-level deletes: merge-on-read vs copy-on-write.** By default, deleting one row in a Delta table rewrites the whole Parquet file that holds it. Deletion vectors avoid that by marking rows removed in a side file and applying the marks at read time; support landed incrementally (DELETE in 2.4.0, UPDATE in 3.0.0, on by default since 3.1.0)[^9]. Iceberg took a spec-versioned route: v2 adds position and equality delete files, and v3 replaces position deletes with per-file deletion-vector bitmaps stored in the Puffin format[^6]. Hudi frames the same trade-off as two table types: Copy-on-Write rewrites a base file on every change (fast reads, slower writes), while Merge-on-Read appends changes to log files merged at query time and compacted later (fast writes, some read cost)[^10].
|
|
18
|
+
|
|
19
|
+
## Where these problems show up
|
|
20
|
+
|
|
21
|
+
Each of those knobs fails in its own recognizable way. The symptoms are the same ones the [Data Formats](#data-formats) and [Small Files](#bottleneck-small-files) pages describe, read through the table's own metadata. A table accumulating many sub-target files is a sizing problem: Hudi's own guidance ties small files to more tasks, more per-file open/close cost, and cloud object-store request-rate limits that trip because at least one request is issued per file regardless of size[^3].
|
|
22
|
+
|
|
23
|
+
Growing metadata is the second signal. Iceberg snapshots and metadata JSON files accumulate until expiry runs, and Iceberg recommends `expire_snapshots` specifically to keep metadata size small on top of freeing data files[^6]. Delta's commit log grows until checkpoints and log compaction absorb it[^1]. Hudi's timeline lengthens with every commit, clean, cluster, and compaction[^5]. Delta's `deltaTable.history()` surfaces the per-commit version, timestamp, and operation for inspecting that growth[^1].
|
|
24
|
+
|
|
25
|
+
Read amplification is the third. On a Merge-on-Read Hudi table, uncompacted delta logs are merged at query time, so a table that has not compacted recently reads more slowly[^4]. The Iceberg equivalent is a data file with many stacked position/equality delete files that a scan must apply[^6].
|
|
26
|
+
|
|
27
|
+
## What each dimension costs
|
|
28
|
+
|
|
29
|
+
Those symptoms aren't cosmetic. Small files cost more in metadata and scheduling than their bytes suggest: a query scans many files for the same data, each file adds fixed overhead, and on object storage the per-file request pattern raises the odds of hitting per-prefix rate limits[^3]. That is exactly the tension the compaction and clustering services exist to resolve, since ingestion favors many small files for low latency while queries favor fewer large ones[^3].
|
|
30
|
+
|
|
31
|
+
Metadata overhead is a planning tax. More data files mean more manifest entries, and small files inflate that metadata disproportionately and slow query planning, which is why Iceberg exposes `rewriteManifests` and compaction to cut it[^6]. Unbounded snapshot and metadata retention keeps that footprint growing until expiry is run[^6].
|
|
32
|
+
|
|
33
|
+
Clustering decides how much data a query skips. Delta's Z-Ordering co-locates related values so data-skipping can prune more files, dramatically reducing bytes read[^1]. The retention settings carry a correctness edge too: once Delta `VACUUM` runs, time travel to a version older than the retention window is gone[^7], and Iceberg warns that running `remove_orphan_files` with too short an interval can delete in-flight files and corrupt the table[^6].
|
|
34
|
+
|
|
35
|
+
## Tuning each dimension
|
|
36
|
+
|
|
37
|
+
Each dimension has a matching maintenance habit that keeps its cost down. Compact on a schedule. For Delta, run `OPTIMIZE` (optionally scoped with a `WHERE` partition predicate) or enable auto compaction on 3.1.0+[^1]. For Iceberg, call `rewrite_data_files`[^2]. For Merge-on-Read Hudi, let async compaction run or force it inline with `hoodie.compact.inline = true`, and use clustering to consolidate small files independently of it[^4][^5].
|
|
38
|
+
|
|
39
|
+
Cluster the columns you filter on. Use Delta `OPTIMIZE ... ZORDER BY (cols)` on a few high-cardinality predicate columns[^1], Iceberg `rewrite_data_files(strategy => 'sort', sort_order => '...')` or its `zorder(...)` form[^2], or Hudi clustering with `hoodie.clustering.plan.strategy.sort.columns`[^5].
|
|
40
|
+
|
|
41
|
+
Expire aggressively but safely. Run Delta `VACUUM` (keeping the 7-day floor unless you have a reason to override it, since the safety check exists to protect concurrent readers)[^7], Iceberg `expire_snapshots` plus periodic `remove_orphan_files` with an interval longer than your longest in-flight write[^6], and tune Hudi's cleaner via `hoodie.clean.commits.retained`[^8].
|
|
42
|
+
|
|
43
|
+
Match the delete strategy to the workload. Enable Delta deletion vectors (`ALTER TABLE ... SET TBLPROPERTIES('delta.enableDeletionVectors' = true)`) so DML marks rows instead of rewriting files, remembering the marks are applied physically only on `OPTIMIZE` or `REORG TABLE ... APPLY (PURGE)`[^9]. On Iceberg, prefer merge-on-read deletes for update-heavy tables and compact the delete files[^6]. On Hudi, pick Copy-on-Write for read-heavy tables and Merge-on-Read for write-heavy or near-real-time ingestion[^10].
|
|
44
|
+
|
|
45
|
+
## Sources
|
|
46
|
+
|
|
47
|
+
[^1]: [Delta Lake: Optimizations (OSS)](https://docs.delta.io/latest/optimizations-oss.html)
|
|
48
|
+
[^2]: [Apache Iceberg: Spark Procedures](https://iceberg.apache.org/docs/latest/spark-procedures/)
|
|
49
|
+
[^3]: [Apache Hudi: File Sizing](https://hudi.apache.org/docs/file_sizing/)
|
|
50
|
+
[^4]: [Apache Hudi: Compaction](https://hudi.apache.org/docs/compaction/)
|
|
51
|
+
[^5]: [Apache Hudi: Clustering](https://hudi.apache.org/docs/clustering/)
|
|
52
|
+
[^6]: [Apache Iceberg: Maintenance](https://iceberg.apache.org/docs/latest/maintenance/) and [Table Spec](https://iceberg.apache.org/spec/)
|
|
53
|
+
[^7]: [Delta Lake: Table Utility Commands](https://docs.delta.io/latest/delta-utility.html)
|
|
54
|
+
[^8]: [Apache Hudi: Cleaning](https://hudi.apache.org/docs/cleaning/)
|
|
55
|
+
[^9]: [Delta Lake: Deletion Vectors](https://docs.delta.io/latest/delta-deletion-vectors.html)
|
|
56
|
+
[^10]: [Apache Hudi: Table Types](https://hudi.apache.org/docs/table_types/)
|
|
@@ -0,0 +1,58 @@
|
|
|
1
|
+
# Caching & Persistence
|
|
2
|
+
|
|
3
|
+
## How persisting and checkpointing differ
|
|
4
|
+
|
|
5
|
+
Spark's `persist()` is the general mechanism for keeping a DataFrame's already-computed partitions around instead of recomputing them from source on every action; `cache()` is shorthand for `persist(StorageLevel.MEMORY_AND_DISK)`: it calls persist with that one fixed level internally[^2], whereas persist() lets you choose any storage level (in-memory vs on-disk, serialized vs deserialized, replicated or not)[^1]. When the level passed to persist() is exactly `MEMORY_AND_DISK`, the two calls are identical[^2].
|
|
6
|
+
|
|
7
|
+
Each storage level is defined by five attributes: `useDisk`, `useMemory`, `useOffHeap`, `deserialized`, and `replication`[^3]. Under `MEMORY_AND_DISK`, Spark stores data directly as objects in memory and serializes only the portion that doesn't fit, writing that overflow to disk[^1]. `MEMORY_AND_DISK_SER` behaves the same way except the data kept in memory is also serialized (data written to disk is always serialized under either level)[^1]. Serialized byte streams use less memory than deserialized JVM objects, which carry structural overhead, but reading them back costs more CPU than reading deserialized objects directly[^3][^4].
|
|
8
|
+
|
|
9
|
+
Checkpointing is a related but distinct mechanism. Instead of persisting, it writes an RDD's partitions to an external, reliable store (HDFS, S3) and drops the lineage, the dependency chain Spark would otherwise use to recompute it[^3][^5]. That truncation doesn't happen the instant `.checkpoint()` is called; the call only marks the RDD, and the actual truncation runs later, in `doCheckpoint()`, after a job using the RDD completes. At that point the RDD is already materialized, and its dependencies and old parents are cleared[^6]. A related call, `localCheckpoint()`, also truncates lineage using Spark's caching layer, but trades fault tolerance for speed: its data lives in ephemeral local executor storage rather than a reliable filesystem, so losing an executor mid-computation can make that data permanently unrecoverable[^6].
|
|
10
|
+
|
|
11
|
+
## When caching pays off
|
|
12
|
+
|
|
13
|
+
Because of how that materialization works, caching earns its keep only in specific situations. It's worth reaching for when the same DataFrame is scanned or recomputed more than once downstream: repeated queries against the same base data, iterative ML training loops, or interactive exploration where a dataset feeds multiple branches[^1][^7]. In one book benchmark, caching a 10M-row DataFrame and materializing it with `count()` cut a subsequent `count()` from 5.11s to 0.44s, roughly a 12x speedup[^1].
|
|
14
|
+
|
|
15
|
+
A few signals point the other way, toward caching having no effect or actively hurting:
|
|
16
|
+
|
|
17
|
+
- **Single-use data.** If a dataset is processed only once downstream, caching just adds serialization and storage-bookkeeping cost for no reuse benefit[^7].
|
|
18
|
+
- **Data too big to fit.** Caching is all-or-nothing per partition: a DataFrame can be "fractionally cached" across its partitions, but individual partitions can't be split. With room for only 4.5 of 8 partitions, exactly 4 get cached; the uncached remainder is recomputed on every access, which can end up slower than not caching at all[^1].
|
|
19
|
+
- **Partial materialization.** `cache()`/`persist()` are lazy: nothing is cached until an action forces a full pass. `count()` forces a genuine full pass and materializes every partition, but an action like `take(1)` computes and caches only the one partition Catalyst needs, so a later full scan still recomputes most of the data. The same applies to `take(10)`, `limit(100)`, or any other partial or filtered scan: only the blocks Spark was forced to touch get cached, and the rest stays lazily uncomputed[^1][^2].
|
|
20
|
+
- **Plan mismatch.** Caching wraps the *analyzed* (pre-optimization) logical plan in an `InMemoryRelation`. A logically-equivalent query written differently produces a different analyzed plan, so Spark can silently miss the cache and recompute from scratch even though the optimized plan would have been identical[^2].
|
|
21
|
+
- **Optimizer limitations.** Caching freezes Catalyst's optimization opportunities at the cached point. It can, for example, block [predicate pushdown](#data-formats) once data is served from the in-memory cache instead of the source[^5].
|
|
22
|
+
|
|
23
|
+
## Why cached blocks don't stick around
|
|
24
|
+
|
|
25
|
+
Materializing a cache is only half the story, though. A `cache()` followed by `count()` guarantees a materialization pass over every partition at that moment: caching happens locally and incrementally, with each executor's BlockManager storing only the blocks it computes, as it computes them, and no centralized "cache the whole dataset" step[^2]. It does not guarantee those partitions stay cached afterward. Two things can undo it:
|
|
26
|
+
|
|
27
|
+
- **Executor loss.** If a block was cached on an executor that later goes away, the cached data goes with it, unless a replicated storage level such as `MEMORY_AND_DISK_2` was used[^2]. This matters specifically for [dynamic allocation](#cluster-config): Spark's documentation is explicit that caching together with dynamic allocation "is NOT safe," because reclaiming idle executors takes their cached blocks with them[^6].
|
|
28
|
+
- **Memory pressure eviction.** Cached blocks can be evicted by other operations' memory demands after materialization, independent of any executor loss[^2].
|
|
29
|
+
|
|
30
|
+
Eviction itself is governed by LRU: Spark removes the least-recently-used cached blocks when storage memory is under pressure[^3][^2][^8]. The trigger is memory contention between Spark's [unified Execution and Storage regions](#memory-model): under the `UnifiedMemoryManager`, Execution has priority, so if a task needs execution memory for a shuffle, join, aggregation, or sort and Storage is occupying that space, Spark evicts cached blocks to free it, and this can happen at any time, not only the next time the cache is accessed[^9][^2]. What happens to an evicted block depends on its storage level, not on a separate spill decision: a disk-backed level (`MEMORY_AND_DISK`, `MEMORY_AND_DISK_SER`) spills the block to disk and reads it back on next use, at the cost of disk I/O; a memory-only level (`MEMORY_ONLY`) simply drops the block and recomputes it from source when needed again[^1][^10][^3][^2].
|
|
31
|
+
|
|
32
|
+
<img class="light-only" src="diagrams/cache-lifecycle.svg" alt="A cached block moves to evicted under memory pressure or lost when an executor goes away, then either spills to disk under MEMORY_AND_DISK or recomputes from lineage under MEMORY_ONLY.">
|
|
33
|
+
<img class="dark-only" src="diagrams/cache-lifecycle.dark.svg" alt="A cached block moves to evicted under memory pressure or lost when an executor goes away, then either spills to disk under MEMORY_AND_DISK or recomputes from lineage under MEMORY_ONLY.">
|
|
34
|
+
|
|
35
|
+
There's no hard limit on how many DataFrames can be cached simultaneously; the constraint is aggregate storage memory, not a count of objects[^1]. That budget is fraction-based configuration translated into a live byte allowance, not a fixed absolute constant: `spark.memory.fraction` (default 0.6) sets the fraction of (heap minus a 300MB reserved region) used for the combined Execution+Storage region, and `spark.memory.storageFraction` (default 0.5) sets the portion of that region immune to eviction. Lowering either makes [spills](#bottleneck-spill) and evictions more frequent[^4][^9][^2].
|
|
36
|
+
|
|
37
|
+
## Habits that keep caching effective
|
|
38
|
+
|
|
39
|
+
A few habits keep caching effective despite how easily any of that goes wrong:
|
|
40
|
+
|
|
41
|
+
- **Materialize deliberately.** Because caching is lazy, follow `cache()`/`persist()` with an action that forces a full pass (`count()` is the standard choice) rather than assuming the call alone did the work[^1][^7].
|
|
42
|
+
- **Match the storage level to the constraint you're solving.** `cache()` / `MEMORY_AND_DISK` is a reasonable default: Spark keeps deserialized objects in memory and only serializes the overflow to disk. Reach for `MEMORY_AND_DISK_SER` when memory pressure is the binding constraint and you can afford the extra CPU cost of deserializing on read[^1][^4][^3].
|
|
43
|
+
- **Protect cached data from dynamic allocation.** If executors holding cached blocks might be reclaimed as idle, raise `spark.dynamicAllocation.cachedExecutorIdleTimeout` so they aren't pulled out from under the cache[^6], or use a replicated storage level (e.g., `MEMORY_AND_DISK_2`) so losing one executor doesn't lose the data[^2].
|
|
44
|
+
- **Reach for checkpointing, not persisting, for genuine lineage truncation with fault tolerance** across a long transformation chain, since it writes to a reliable external filesystem rather than relying on executor-local memory or disk[^3][^5]. Avoid `localCheckpoint()` under dynamic allocation for the same reason caching is unsafe there: its data lives in ephemeral local executor storage that can vanish along with a reclaimed executor[^6]. Also set a checkpoint directory via `SparkContext.setCheckpointDir` before calling `.checkpoint()`; without one, the call throws immediately rather than silently doing nothing[^6].
|
|
45
|
+
- **Don't cache data that won't benefit:** single-use datasets, datasets that don't fit in available memory, or datasets you only ever touch through partial actions (`take`, `limit`). In each of these cases the caching overhead isn't paid back[^7][^1].
|
|
46
|
+
|
|
47
|
+
## Sources
|
|
48
|
+
|
|
49
|
+
[^1]: *Learning Spark, 2nd Edition*, Damji, Wenig, Das & Lee, ch. 7 — Optimizing and Tuning Spark Applications
|
|
50
|
+
[^2]: [Explaining the Mechanics of Spark Caching](https://luminousmen.com/post/explaining-the-mechanics-of-spark-caching)
|
|
51
|
+
[^3]: *High Performance Spark, 2nd Edition*, Karau, Polak & Warren, ch. 7 — Effective Transformations
|
|
52
|
+
[^4]: [Configuration — Spark](https://spark.apache.org/docs/latest/configuration.html)
|
|
53
|
+
[^5]: [Spark Tips: Caching](https://luminousmen.com/post/spark-tips-caching)
|
|
54
|
+
[^6]: [RDD.scala](https://raw.githubusercontent.com/apache/spark/v3.5.0/core/src/main/scala/org/apache/spark/rdd/RDD.scala)
|
|
55
|
+
[^7]: *Spark: The Definitive Guide*, Chambers & Zaharia, ch. 19 — Performance Tuning
|
|
56
|
+
[^8]: [Task Memory Management in Spark](https://raw.githubusercontent.com/spoddutur/spark-notes/master/task_memory_management_in_spark.md)
|
|
57
|
+
[^9]: [Dive into Spark Memory](https://luminousmen.com/post/dive-into-spark-memory)
|
|
58
|
+
[^10]: [RDD Programming Guide](https://spark.apache.org/docs/latest/rdd-programming-guide.html)
|
|
@@ -0,0 +1,78 @@
|
|
|
1
|
+
# PySpark Specifics
|
|
2
|
+
|
|
3
|
+
## Crossing the JVM/Python boundary
|
|
4
|
+
|
|
5
|
+
A plain PySpark UDF runs one row at a time in a separate Python process, and getting each row there costs something real: "PySpark UDFs required data movement between the JVM and Python, which was quite expensive," using pickle to serialize the data across that boundary[^1]. *Spark: The Definitive Guide* breaks the cost down further. Starting the extra Python process is one expense, but "the real cost is in serializing the data to Python," and once the data is over there the JVM can no longer manage that worker's memory, so JVM and Python end up competing for the same machine's RAM and the worker can fail under pressure[^2]. That's the underlying reason the same source recommends writing UDFs in Scala or Java and calling them from Python when that's practical[^2].
|
|
6
|
+
|
|
7
|
+
Since Spark 3.4 the row-at-a-time path can itself be Arrow-optimized without giving up row semantics: set `useArrow=True` on `udf()`, or flip the session-wide `spark.sql.execution.pythonUDF.arrow.enabled` (default `false`), and a regular Python UDF becomes an "Arrow Python UDF" that still runs row by row but moves data with Arrow instead of pickle[^3][^4]. The session config only takes effect when `useArrow` is left unset on the UDF itself[^3].
|
|
8
|
+
|
|
9
|
+
<img class="light-only" src="diagrams/udf-execution-models.svg" alt="Across the JVM to Python boundary a plain Python UDF pickles row-at-a-time, an Arrow-optimized UDF uses Arrow transfer with row semantics, and a pandas UDF passes whole Arrow batches with no per-row handoff.">
|
|
10
|
+
<img class="dark-only" src="diagrams/udf-execution-models.dark.svg" alt="Across the JVM to Python boundary a plain Python UDF pickles row-at-a-time, an Arrow-optimized UDF uses Arrow transfer with row semantics, and a pandas UDF passes whole Arrow batches with no per-row handoff.">
|
|
11
|
+
|
|
12
|
+
The pandas UDF (vectorized UDF, introduced in Spark 2.3) goes further and drops the per-row JVM↔Python handoff entirely: it hands the Python worker whole Arrow batches, operated on as pandas Series or DataFrames, so there's nothing to pickle row by row[^1]. Three shapes cover most cases. **SCALAR** (`pandas.Series, ... -> pandas.Series`) is the direct vectorized swap-in for a row-at-a-time scalar UDF (computing `v + 1`, or `cubed(x)`); PySpark calls the function once per Arrow batch and concatenates the results back into a column[^5][^1]. **SCALAR_ITER** (`Iterator[Series] -> Iterator[Series]`) works the same way internally, but takes and yields an iterator instead of a single Series, which lets a function prefetch across batches[^3]. **MAP_ITER**, exposed as `DataFrame.mapInPandas()` rather than as a `pandas_udf` type, maps an iterator of whole `pandas.DataFrame` partitions to another iterator of `pandas.DataFrame`s and, unlike the other two, can change the row count[^3]. A related grouped-map API, `DataFrame.groupBy().applyInPandas()`, splits the DataFrame into groups and runs a `pandas.DataFrame -> pandas.DataFrame` function per group: split, apply, combine[^5][^3].
|
|
13
|
+
|
|
14
|
+
`spark.sql.execution.arrow.pyspark.enabled` isn't limited to pandas UDFs, either: it also governs Arrow-based columnar transfer for `DataFrame.toPandas()` and for `SparkSession.createDataFrame()` when given a pandas DataFrame or NumPy ndarray[^4][^3]. A companion flag, `spark.sql.execution.arrow.pyspark.fallback.enabled`, silently drops back to the non-Arrow path if an error occurs before computation starts[^3][^4].
|
|
15
|
+
|
|
16
|
+
Two worker-level configs round out the picture. `spark.python.worker.reuse` (default `true`) keeps a fixed pool of Python worker processes alive across tasks instead of forking a fresh one each time, which also means a large broadcast variable doesn't have to cross the JVM↔Python boundary again for every task[^4]. `spark.python.worker.memory` (default `512m`) is a per-worker, Spark-managed accounting threshold for in-worker aggregation buffering (not an OS-enforced cap), and Spark [spills to disk](#bottleneck-spill) once it's exceeded[^4].
|
|
17
|
+
|
|
18
|
+
## Measuring the gap
|
|
19
|
+
|
|
20
|
+
The clearest measurement here is workload-level. Databricks' introductory pandas-UDF post ran three operations (Plus One, Cumulative Probability, Subtract Mean) over a 10M-row, two-column DataFrame on a single-node Databricks Community Edition cluster, and found pandas UDFs "perform much better than row-at-a-time UDFs across the board, ranging from 3x to over 100x"[^5]. If a job's Python UDFs are a suspected bottleneck, expect an order-of-magnitude gap between a row-at-a-time UDF and its pandas-UDF equivalent, not a marginal one.
|
|
21
|
+
|
|
22
|
+
The comparisons in the corpus consistently favor `pyspark.sql.functions` and vectorized code over row-at-a-time UDFs. For the "plus one" example, "built-in column operators can perform much faster in this scenario," and the pandas UDF equivalent is "much faster than the row-at-a-time version" because it's vectorized over the Series[^5]. A Python UDF that could instead be expressed with built-in column functions is worth flagging on its own.
|
|
23
|
+
|
|
24
|
+
On the memory side, the two failure modes look different and are worth telling apart. A worker that spills is showing up in Spark's own accounting: it exceeded `spark.python.worker.memory` during aggregation and wrote to disk, which is expected behavior, not a crash[^4]. A worker that's actually killed is a container-level event: if total container memory (JVM heap plus off-heap plus the Python process) exceeds what YARN or Kubernetes allocated, "Kubernetes won't hesitate" and the process is "OOMKilled"[^6]. That boundary is governed by [executor overhead](#memory-model) and `spark.executor.pyspark.memory` settings, not by `worker.memory`.
|
|
25
|
+
|
|
26
|
+
## What the gap costs
|
|
27
|
+
|
|
28
|
+
Those numbers point to a structural cost, not a tuning quirk. The JVM↔Python boundary is where a plain Python UDF pays twice: once to start the separate process, and again (the larger cost) to serialize every row across it[^2]. Because the JVM can't manage memory inside the Python process once data has crossed over, the two runtimes end up competing for the same machine's memory, and the Python worker can fail under that pressure[^2]. That's a correctness risk as well as a performance one, and it's why *Spark: The Definitive Guide* recommends Scala/Java UDFs called from Python over native Python UDFs wherever that's practical[^2].
|
|
29
|
+
|
|
30
|
+
The magnitude backs this up: Databricks measured pandas UDFs beating row-at-a-time UDFs by 3x to over 100x depending on the operation[^5]. Skipping per-row pickling (by moving to Arrow-based row UDFs or, further, to pandas UDFs operating on whole batches) isn't a marginal tuning knob here; it changes which order of magnitude a job runs at.
|
|
31
|
+
|
|
32
|
+
The two memory configs matter for different reasons, too. `spark.python.worker.memory` only controls when Spark chooses to spill aggregation state to disk; tuning it trades disk I/O for headroom, it doesn't prevent a crash[^4]. Actual OOM kills happen at the container boundary, and `spark.executor.pyspark.memory` (which leans on Python's `resource` module and so doesn't cap memory on macOS and doesn't exist at all on Windows) is the config that's actually in that path[^4]. Conflating the two means tuning the wrong knob when a Python worker gets OOMKilled.
|
|
33
|
+
|
|
34
|
+
## Closing the gap
|
|
35
|
+
|
|
36
|
+
Closing that gap means avoiding the boundary crossing altogether, or crossing it as cheaply as possible. Where the logic allows it, prefer a built-in `pyspark.sql.functions` expression or a Scala/Java UDF called from Python over a plain row-at-a-time Python UDF[^2][^5]; this is the change with the largest documented payoff, 3x to over 100x[^5].
|
|
37
|
+
|
|
38
|
+
Where a Python UDF is unavoidable, cut the row-at-a-time cost first by turning on Arrow for it:
|
|
39
|
+
|
|
40
|
+
```python
|
|
41
|
+
@udf(returnType='int', useArrow=True) # An Arrow Python UDF
|
|
42
|
+
def arrow_slen(s):
|
|
43
|
+
return len(s)
|
|
44
|
+
```
|
|
45
|
+
|
|
46
|
+
or set the session-wide equivalent so existing UDFs pick it up without code changes.
|
|
47
|
+
|
|
48
|
+
> **PySpark:** `spark.conf.set("spark.sql.execution.pythonUDF.arrow.enabled", "true")` turns plain UDFs into Arrow-backed ones, as long as `useArrow` isn't explicitly set on the UDF itself[^3][^4].
|
|
49
|
+
|
|
50
|
+
Beyond that, reach for the vectorized APIs instead of a scalar UDF:
|
|
51
|
+
|
|
52
|
+
- Use a **SCALAR** pandas UDF (`pandas.Series, ... -> pandas.Series`) as the default vectorized replacement for a row-at-a-time scalar UDF[^5][^1].
|
|
53
|
+
- Use **SCALAR_ITER** (`Iterator[Series] -> Iterator[Series]`) when the function needs expensive one-time setup. The documented pattern initializes state once, then loops over the batch iterator reusing it, instead of re-initializing per batch[^3]:
|
|
54
|
+
|
|
55
|
+
```python
|
|
56
|
+
def apply_with_state(iterator):
|
|
57
|
+
state = very_expensive_initialization()
|
|
58
|
+
for batch in iterator:
|
|
59
|
+
yield calculate_with_state(batch, state)
|
|
60
|
+
```
|
|
61
|
+
|
|
62
|
+
- Use `DataFrame.mapInPandas()` when the transform needs to change row count (filtering, expansion, deduplication) rather than map one-to-one[^3].
|
|
63
|
+
- Use `DataFrame.groupBy().applyInPandas()` for per-group logic that needs the whole group as state, such as subtracting a group mean or fitting a per-group regression[^5][^3]. Size groups with care: a full group loads into memory before the function runs, and `maxRecordsPerBatch` doesn't apply to groups, so a skewed group risks OOM[^3].
|
|
64
|
+
|
|
65
|
+
For pandas/NumPy conversion at the driver, turn on Arrow explicitly rather than relying on defaults.
|
|
66
|
+
|
|
67
|
+
> **PySpark:** with `spark.sql.execution.arrow.pyspark.enabled` set to `"true"`, `spark.createDataFrame(pdf)` and `df.select("*").toPandas()` convert through Arrow and return the same result as the non-Arrow path[^4][^3]. `spark.sql.execution.arrow.pyspark.fallback.enabled` keeps a safety net by falling back silently on pre-computation errors[^3][^4].
|
|
68
|
+
|
|
69
|
+
Leave `spark.python.worker.reuse` at its default (`true`) unless there's a specific reason not to; it keeps a fixed pool of Python workers alive so a large broadcast variable isn't re-shipped to Python for every task[^4]. If a job spills at the Python-worker level, `spark.python.worker.memory` is the knob for that; if it's getting OOMKilled at the container level, look at executor overhead and `spark.executor.pyspark.memory` instead, keeping in mind the latter's `resource`-module limitations on macOS and its absence on Windows[^4][^6].
|
|
70
|
+
|
|
71
|
+
## Sources
|
|
72
|
+
|
|
73
|
+
[^1]: *Learning Spark, 2nd Edition*, Damji, Wenig, Das & Lee, ch. 5
|
|
74
|
+
[^2]: *Spark: The Definitive Guide*, Chambers & Zaharia, ch. 6
|
|
75
|
+
[^3]: [Apache Arrow in PySpark](https://spark.apache.org/docs/3.5.8/api/python/user_guide/sql/arrow_pandas.html)
|
|
76
|
+
[^4]: [Configuration — Spark](https://spark.apache.org/docs/latest/configuration.html)
|
|
77
|
+
[^5]: [Introducing Pandas UDF for PySpark](https://www.databricks.com/blog/2017/10/30/introducing-vectorized-udfs-for-pyspark.html)
|
|
78
|
+
[^6]: [Dive into Spark memory management](https://luminousmen.com/post/dive-into-spark-memory)
|
|
@@ -0,0 +1,167 @@
|
|
|
1
|
+
# Adaptive Query Execution
|
|
2
|
+
|
|
3
|
+
## What it is
|
|
4
|
+
|
|
5
|
+
The static Catalyst optimizer plans a query once, before execution, using cost estimates
|
|
6
|
+
derived from static statistics: row counts, min/max, NDVs, or defaults when statistics are
|
|
7
|
+
missing.[^1] Adaptive Query Execution (AQE), shipped in Spark 3.0, instead re-optimizes the
|
|
8
|
+
plan mid-execution using runtime statistics gathered from completed "query stages": the
|
|
9
|
+
sections of a plan bounded by shuffle or broadcast exchange materialization points.[^2] A
|
|
10
|
+
shuffle or broadcast forces Spark to materialize its input before continuing, which makes it
|
|
11
|
+
a natural checkpoint: once one or more leaf stages finish, AQE marks them complete, updates
|
|
12
|
+
the logical plan with the real (not estimated) statistics, and reruns a selected set of
|
|
13
|
+
logical and physical optimization rules (including AQE-specific rules such as partition
|
|
14
|
+
coalescing and skew-join handling) before executing the next stages.[^2]
|
|
15
|
+
|
|
16
|
+
Shuffle statistics only become available once a stage is *fully* materialized, not
|
|
17
|
+
incrementally as individual map tasks finish. A stage's successor can only proceed once
|
|
18
|
+
every parallel process producing that stage's output has completed, which is exactly why
|
|
19
|
+
materialization points are the reoptimization opportunity: it's the moment when statistics on
|
|
20
|
+
all of a stage's partitions are known and the next stage hasn't started yet.[^2] AQE kicks off
|
|
21
|
+
all leaf stages (the ones with no upstream dependency) first; as each one finishes, the
|
|
22
|
+
framework marks it complete, updates the plan, and re-optimizes before launching whichever
|
|
23
|
+
next stages now have all their children materialized. This execute→reoptimize→execute loop
|
|
24
|
+
repeats (once per completed stage, not just once) until the whole query finishes, so the
|
|
25
|
+
number of reoptimizations scales with how many shuffle/broadcast boundaries the plan has.[^2]
|
|
26
|
+
The Spark SQL config description for `spark.sql.adaptive.enabled` frames this the same way:
|
|
27
|
+
AQE re-optimizes "the query plan in the middle of query execution, based on accurate runtime
|
|
28
|
+
statistics."[^3]
|
|
29
|
+
|
|
30
|
+
<img class="light-only" src="diagrams/aqe-loop.svg" alt="AQE runs a loop that executes leaf stages, materializes at a shuffle or broadcast boundary, collects runtime statistics, re-applies rules to coalesce partitions and split skew and promote joins to broadcast, then launches the next stages until the plan is complete.">
|
|
31
|
+
<img class="dark-only" src="diagrams/aqe-loop.dark.svg" alt="AQE runs a loop that executes leaf stages, materializes at a shuffle or broadcast boundary, collects runtime statistics, re-applies rules to coalesce partitions and split skew and promote joins to broadcast, then launches the next stages until the plan is complete.">
|
|
32
|
+
|
|
33
|
+
AQE re-optimizes three things the static optimizer cannot, because none of them are knowable
|
|
34
|
+
before execution: the number of post-shuffle partitions (coalescing small partitions produced
|
|
35
|
+
by wide transformations), the join strategy (converting a statically planned sort-merge join
|
|
36
|
+
to a broadcast join once the actual materialized size of a join side is known), and skewed
|
|
37
|
+
partitions in a shuffle join (splitting oversized partitions detected from shuffle file
|
|
38
|
+
statistics).[^2] High Performance Spark frames this as AQE using "runtime information about
|
|
39
|
+
the data it is processing along with the target output to go beyond static optimizations,"
|
|
40
|
+
with partitioning and join strategy singled out as the two biggest areas it impacts.[^4] AQE
|
|
41
|
+
also performs empty-relation propagation, replacing subqueries that turn out to be empty
|
|
42
|
+
(impossible joins, empty unions) with a dummy empty `LocalRelation`, again something only
|
|
43
|
+
knowable once data is materialized.[^4]
|
|
44
|
+
|
|
45
|
+
## How it's detected
|
|
46
|
+
|
|
47
|
+
Because these are runtime decisions, none of them show up in a static `explain()` call; you
|
|
48
|
+
have to run the query and check the Spark UI for the finalized, adapted plan.[^4] The same
|
|
49
|
+
caveat applies more broadly: since AQE adjusts plans at runtime, `explain()` might show one
|
|
50
|
+
plan while the Spark UI shows a different one actually executed.[^5]
|
|
51
|
+
|
|
52
|
+
What you're looking for in that adaptive plan are the specific runtime rewrites AQE can make:
|
|
53
|
+
|
|
54
|
+
- **Partition coalescing.** AQE combines *adjacent* small post-shuffle partitions into bigger
|
|
55
|
+
ones by reading shuffle file statistics,[^2] targeting
|
|
56
|
+
`spark.sql.adaptive.advisoryPartitionSizeInBytes` (default 64 MB).[^6] In a worked example,
|
|
57
|
+
five post-shuffle partitions where three are small get coalesced into one, cutting the
|
|
58
|
+
final-aggregation task count from five to three.[^2]
|
|
59
|
+
- **Skew splitting.** A partition is considered skewed if its size is larger than the median
|
|
60
|
+
partition size times `spark.sql.adaptive.skewJoin.skewedPartitionFactor` (default 5.0) *and*
|
|
61
|
+
larger than `spark.sql.adaptive.skewJoin.skewedPartitionThresholdInBytes` (default 256
|
|
62
|
+
MB).[^7][^8] The original design doc's worked example (table A join table B, where B's
|
|
63
|
+
partition 0 is skewed) has the `OptimizeSkewedJoin` rule divide that partition into N
|
|
64
|
+
smaller splits reading disjoint ranges of upstream map outputs, create N separate tasks each
|
|
65
|
+
joining its slice of B's partition 0 against A's corresponding partition 0, then union the N
|
|
66
|
+
results. That trades reading A's partition 0 N times against eliminating a straggler
|
|
67
|
+
task.[^7]
|
|
68
|
+
- **Join strategy promotion.** When a join side's runtime-materialized size falls under the
|
|
69
|
+
adaptive broadcast threshold, AQE converts a planned sort-merge join to a broadcast hash
|
|
70
|
+
join. It reuses the shuffle output already written rather than re-materializing the build
|
|
71
|
+
side. The perf-tuning guide frames the benefit as avoiding a re-sort of both join sides and
|
|
72
|
+
reading the existing shuffle files locally instead, conditional on
|
|
73
|
+
`spark.sql.adaptive.localShuffleReader.enabled`.[^6] The same guide calls this conversion
|
|
74
|
+
"not as efficient as planning a broadcast hash join in the first place," consistent with
|
|
75
|
+
reusing existing shuffle output rather than recomputing an equivalent build side from
|
|
76
|
+
scratch.[^6]
|
|
77
|
+
- **Local shuffle reads.** Once that promotion happens, the side that no longer needs to be
|
|
78
|
+
partitioned by join key can be read straight off the shuffle files each executor already has
|
|
79
|
+
locally, instead of pulling blocks from remote executors over the network. This is what
|
|
80
|
+
`spark.sql.adaptive.localShuffleReader.enabled` (default true since 3.0.0) does, and it
|
|
81
|
+
applies whenever shuffle partitioning is no longer needed, such as after a sort-merge-to-
|
|
82
|
+
broadcast conversion.[^6][^2]
|
|
83
|
+
|
|
84
|
+
Dynamic Partition Pruning is a separate mechanism worth distinguishing from AQE's loop when
|
|
85
|
+
reading a plan: DPP inserts a `DynamicPruningSubquery`/`DynamicPruningExpression` node during
|
|
86
|
+
logical optimization/physical planning, before execution starts, when an equi-join is on a
|
|
87
|
+
partition column and pruning looks beneficial by static statistics.[^9] Because DPP's
|
|
88
|
+
calculation runs at planning time and AQE's runs later, mid-execution, off completed-stage
|
|
89
|
+
statistics, the ordering implied is that DPP's pruning predicate is planned first, with AQE's
|
|
90
|
+
adaptive rules applying afterward, during execution, on top of whatever DPP already
|
|
91
|
+
pruned[^9][^2], though the exact current-version integration between the two isn't
|
|
92
|
+
detailed further in the available sources, so treat that ordering as directionally supported
|
|
93
|
+
rather than exhaustively confirmed.
|
|
94
|
+
|
|
95
|
+
## Why it matters
|
|
96
|
+
|
|
97
|
+
AQE exists precisely because partition sizing, join strategy, and skew are only knowable once
|
|
98
|
+
real data has been materialized. That's the gap the static optimizer can't close on its
|
|
99
|
+
own.[^2][^4] But it isn't a guaranteed win. High Performance Spark warns that "while AQE
|
|
100
|
+
generally does better than unoptimized code, there are cases, especially when targeting
|
|
101
|
+
Iceberg tables, where AQE does more harm than good": AQE's partitioning-to-target-table
|
|
102
|
+
matching is generally beneficial but can cause severe regressions if you've already done
|
|
103
|
+
deliberate work to avoid key skew via custom partitioning, since AQE's attempt to match the
|
|
104
|
+
target table's write partitioning can undo that work, especially when there's a large amount
|
|
105
|
+
of key skew.[^4]
|
|
106
|
+
|
|
107
|
+
A second, more general failure mode is bad input statistics feeding AQE's runtime decisions:
|
|
108
|
+
if the source data has wrong stats (compressed JSON or Kafka streams are called out
|
|
109
|
+
specifically), AQE can make things worse rather than better.[^5] The same source's framing is
|
|
110
|
+
that AQE "isn't magic... it helps polish your plan; it doesn't design it for you. You still
|
|
111
|
+
need good partitioning fundamentals". AQE can smooth over a reasonably partitioned plan
|
|
112
|
+
but doesn't reliably rescue one built on bad upstream statistics or layout.[^5]
|
|
113
|
+
|
|
114
|
+
AQE's runtime re-optimization also doesn't reach into DataFrame caching, which is worth
|
|
115
|
+
knowing so you don't reach for the wrong lever: `.cache()`/`.persist()` wrap the query's
|
|
116
|
+
*analyzed* logical plan (the point after the analyzer phase but before optimization) in an
|
|
117
|
+
`InMemoryRelation` node, and cache lookup compares the analyzed plan of the current query
|
|
118
|
+
against what was cached, recomputing if they don't match exactly, even if the two queries
|
|
119
|
+
would ultimately produce the same optimized physical plan.[^10] Because AQE operates on the
|
|
120
|
+
physical plan at runtime, well downstream of that analyzed-plan cache key, disabling AQE
|
|
121
|
+
doesn't change whether a cached DataFrame's analyzed plan matches; it only changes how the
|
|
122
|
+
physical/execution plan is chosen after that cache lookup already happened.[^10]
|
|
123
|
+
|
|
124
|
+
## How to fix it
|
|
125
|
+
|
|
126
|
+
AQE's coalescing, skew-handling, and join-promotion rules only fire once
|
|
127
|
+
`spark.sql.adaptive.enabled` is on, and each has its own knobs worth tuning rather than leaving
|
|
128
|
+
at the defaults:
|
|
129
|
+
|
|
130
|
+
- `spark.sql.adaptive.coalescePartitions.enabled` (default true) turns on partition
|
|
131
|
+
coalescing; `spark.sql.adaptive.advisoryPartitionSizeInBytes` (default 64 MB) sets its
|
|
132
|
+
target size.[^6] `spark.sql.adaptive.coalescePartitions.minPartitionNum` bounds the minimum
|
|
133
|
+
resulting parallelism,[^4] and `spark.sql.adaptive.coalescePartitions.minPartitionSize`
|
|
134
|
+
(default 1 MB, since 3.2) sets a floor on coalesced partition size for when the adaptively
|
|
135
|
+
calculated target is too small.[^6]
|
|
136
|
+
- `spark.sql.adaptive.coalescePartitions.parallelismFirst` (default true, since 3.2) tells
|
|
137
|
+
Spark to ignore the advisory target size entirely and instead calculate a (usually smaller)
|
|
138
|
+
target based on cluster default parallelism, prioritizing task parallelism over hitting the
|
|
139
|
+
64 MB target. On a busy cluster, set it to `false` to respect the configured target size and
|
|
140
|
+
avoid producing many small tasks.[^6]
|
|
141
|
+
- `spark.sql.adaptive.skewJoin.enabled` (default true) turns on skew-join handling for both
|
|
142
|
+
sort-merge and shuffled hash joins.[^8] Tune the detection thresholds via
|
|
143
|
+
`spark.sql.adaptive.skewJoin.skewedPartitionFactor` (default 5.0) and
|
|
144
|
+
`spark.sql.adaptive.skewJoin.skewedPartitionThresholdInBytes` (default 256 MB); set the
|
|
145
|
+
byte threshold larger than `advisoryPartitionSizeInBytes`.[^8] Use
|
|
146
|
+
`spark.sql.adaptive.forceOptimizeSkewedJoin` (default false, since 3.3.0) to force the rule
|
|
147
|
+
to fire even when it would introduce extra shuffle.[^6]
|
|
148
|
+
- `spark.sql.adaptive.localShuffleReader.enabled` (default true, since 3.0.0) keeps the local,
|
|
149
|
+
per-mapper shuffle read enabled after a sort-merge-to-broadcast conversion.[^6]
|
|
150
|
+
|
|
151
|
+
If AQE's target-table partition matching is regressing an Iceberg write that already has
|
|
152
|
+
deliberate custom partitioning to avoid key skew, set the table's (or write-time)
|
|
153
|
+
`write.distribution-mode` property to `none`, at the cost of a higher risk of many small
|
|
154
|
+
files from unsorted/unhashed writers per partition.[^4]
|
|
155
|
+
|
|
156
|
+
## Sources
|
|
157
|
+
|
|
158
|
+
[^1]: [Deep Dive into Spark SQL's Catalyst Optimizer](https://www.databricks.com/blog/2015/04/13/deep-dive-into-spark-sqls-catalyst-optimizer.html)
|
|
159
|
+
[^2]: [Adaptive Query Execution: Speeding Up Spark SQL at Runtime](https://www.databricks.com/blog/2020/05/29/adaptive-query-execution-speeding-up-spark-sql-at-runtime.html)
|
|
160
|
+
[^3]: [SQLConf.scala](https://raw.githubusercontent.com/apache/spark/v3.5.0/sql/catalyst/src/main/scala/org/apache/spark/sql/internal/SQLConf.scala)
|
|
161
|
+
[^4]: *High Performance Spark, 2nd Edition*, Karau, Polak & Warren, ch. 5
|
|
162
|
+
[^5]: [Spark Partitions](https://luminousmen.com/post/spark-partitions)
|
|
163
|
+
[^6]: [Performance Tuning: Spark SQL, DataFrames and Datasets Guide](https://spark.apache.org/docs/latest/sql-performance-tuning.html)
|
|
164
|
+
[^7]: [SPARK-29544: Optimize Skewed Join at Runtime](https://issues.apache.org/jira/browse/SPARK-29544)
|
|
165
|
+
[^8]: [Configuration: Spark](https://spark.apache.org/docs/latest/configuration.html)
|
|
166
|
+
[^9]: [What's New in Apache Spark 3: Dynamic Partition Pruning](https://www.waitingforcode.com/apache-spark-sql/whats-new-apache-spark-3-dynamic-partition-pruning/read)
|
|
167
|
+
[^10]: [Explaining the Mechanics of Spark Caching](https://luminousmen.com/post/explaining-the-mechanics-of-spark-caching)
|