This intensive three-day workshop is dedicated to engineering and refining high-performance data-processing pipelines that leverage PySpark, Pandas, and Polars within Kubernetes environments.
Learners will gain a hands-on grasp of how Spark applications operate on Kubernetes, focusing on how configuration choices at the application level directly impact performance, scalability, resource efficiency, and overall cost. The curriculum covers essential optimization topics such as executor sizing, memory management, dynamic allocation, partitioning strategies, shuffle mechanics, mitigating the small-file issue, and efficient handling of Parquet data.
The program also tackles frequent hurdles encountered with Pandas, such as memory constraints and out-of-memory errors, while introducing Polars as a high-speed alternative for specific data tasks. Through practical exercises, participants will learn to diagnose performance and memory bottlenecks, evaluate various configuration approaches, and apply optimization strategies to real-world ETL and machine learning use cases.
The core focus remains on practical decision-making: equipping professionals to pinpoint bottlenecks, choose the right tools, configure Spark effectively, and strike a balance between performance and infrastructure resource consumption and cost.
Read more...