This intensive three-day workshop is dedicated to constructing and refining high-performance data-processing pipelines leveraging PySpark, Pandas and Polars within Kubernetes-based ecosystems.
Learners will gain a working knowledge of how Spark applications operate on Kubernetes, exploring how application-level settings directly impact performance, scalability, resource utilization, and operational expenditure. The curriculum delves into essential optimisation domains such as executor sizing, memory distribution, dynamic allocation, partitioning methodologies, shuffle mechanics, mitigating small-file issues, and executing efficient Parquet operations.
Additionally, the program tackles prevalent hurdles associated with Pandas, such as memory constraints and out-of-memory exceptions, while presenting Polars as a high-throughput alternative for specific data-processing tasks. Through practical, hands-on activities, participants will learn to diagnose performance and memory bottlenecks, evaluate various configuration strategies, and implement optimisation techniques within realistic ETL and machine learning contexts.
The core objective remains practical decision-making: mastering the identification of performance bottlenecks, selecting the optimal tool, configuring Spark for maximum efficiency, and achieving a balance between performance gains and infrastructure resource consumption and cost.
Read more...