Get in Touch
 Duration 35 hours

Course Outline

Databricks Platform and Lakehouse Fundamentals

  • Lakehouse architecture and core components of Databricks.
  • Strategies for organizing workspaces and catalogs.

Databricks Workspace and Notebooks

  • Workspace navigation and development using notebooks.
  • Structuring code into reusable notebook modules.

Apache Spark Architecture and Execution

  • Understanding the Spark runtime architecture and execution model.
  • Concepts of lazy evaluation and job Directed Acyclic Graphs (DAGs).

PySpark DataFrames and the DataFrame API

  • DataFrame abstractions and schema management.
  • Essential DataFrame operations and column expressions.

Translating SQL to PySpark DataFrames

  • Mapping core SQL clauses to DataFrame operations.
  • Implementing window functions and aggregations in PySpark.

Reading and Writing Data in Databricks

  • Accessing data from standard file systems and databases.
  • Writing and partitioning data within the Lakehouse architecture.

Delta Lake and Table Management

  • Working with Delta tables and ACID transactions.
  • Utilizing time travel and schema evolution features.

Data Cleaning and Transformation Patterns

  • Techniques for data cleaning and type conversion.
  • Constructing reusable transformation logic.

User-Defined Functions and Modular Code

  • Implementing Python UDFs and pandas UDFs.
  • Refactoring procedural logic into modular functions.

Performance Tuning and Optimization

  • Strategies for partitioning and caching.
  • Identifying bottlenecks using the Spark UI.

Structured Streaming Fundamentals

  • Comparing batch processing versus streaming models.
  • Working with streaming DataFrames and basic aggregations.

Databricks Jobs and Workflow Orchestration

  • Scheduling notebooks as jobs and tasks.
  • Designing multi-step workflows with dependencies.

Unity Catalog and Data Governance

  • Unity Catalog architecture and namespace management.
  • Managing access control and data lineage.

Testing, Debugging, and Production Practices

  • Conducting unit tests for PySpark logic.
  • Debugging techniques and maintaining code quality standards.

End-to-End Financial Services Use Cases

  • Developing an end-to-end banking ETL pipeline.
  • Migrating legacy SQL processes to PySpark implementations.

Migrating SQL Workloads to PySpark

  • Adopting migration strategies and planning patterns.
  • Performing incremental conversion of SQL workflows to PySpark.

Requirements

  • Proficiency in Python programming, encompassing functions and data types.
  • A solid grasp of SQL concepts, including joins, aggregations, and subqueries.
  • No previous experience with Databricks or PySpark is necessary.

Target Audience

  • Data engineers, data analysts, and data professionals.
  • Teams currently migrating existing SQL-based workflows to Databricks and PySpark.

Number of participants


Price per participant

Testimonials (1)

Upcoming Courses

Related Categories