Course Outline
Databricks Platform and Lakehouse Fundamentals
- Lakehouse architecture and core components of Databricks.
- Strategies for organizing workspaces and catalogs.
Databricks Workspace and Notebooks
- Workspace navigation and development using notebooks.
- Structuring code into reusable notebook modules.
Apache Spark Architecture and Execution
- Understanding the Spark runtime architecture and execution model.
- Concepts of lazy evaluation and job Directed Acyclic Graphs (DAGs).
PySpark DataFrames and the DataFrame API
- DataFrame abstractions and schema management.
- Essential DataFrame operations and column expressions.
Translating SQL to PySpark DataFrames
- Mapping core SQL clauses to DataFrame operations.
- Implementing window functions and aggregations in PySpark.
Reading and Writing Data in Databricks
- Accessing data from standard file systems and databases.
- Writing and partitioning data within the Lakehouse architecture.
Delta Lake and Table Management
- Working with Delta tables and ACID transactions.
- Utilizing time travel and schema evolution features.
Data Cleaning and Transformation Patterns
- Techniques for data cleaning and type conversion.
- Constructing reusable transformation logic.
User-Defined Functions and Modular Code
- Implementing Python UDFs and pandas UDFs.
- Refactoring procedural logic into modular functions.
Performance Tuning and Optimization
- Strategies for partitioning and caching.
- Identifying bottlenecks using the Spark UI.
Structured Streaming Fundamentals
- Comparing batch processing versus streaming models.
- Working with streaming DataFrames and basic aggregations.
Databricks Jobs and Workflow Orchestration
- Scheduling notebooks as jobs and tasks.
- Designing multi-step workflows with dependencies.
Unity Catalog and Data Governance
- Unity Catalog architecture and namespace management.
- Managing access control and data lineage.
Testing, Debugging, and Production Practices
- Conducting unit tests for PySpark logic.
- Debugging techniques and maintaining code quality standards.
End-to-End Financial Services Use Cases
- Developing an end-to-end banking ETL pipeline.
- Migrating legacy SQL processes to PySpark implementations.
Migrating SQL Workloads to PySpark
- Adopting migration strategies and planning patterns.
- Performing incremental conversion of SQL workflows to PySpark.
Requirements
- Proficiency in Python programming, encompassing functions and data types.
- A solid grasp of SQL concepts, including joins, aggregations, and subqueries.
- No previous experience with Databricks or PySpark is necessary.
Target Audience
- Data engineers, data analysts, and data professionals.
- Teams currently migrating existing SQL-based workflows to Databricks and PySpark.
Testimonials (1)
I liked that it was practical. Loved to apply the theoretical knowledge with practical examples.