Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Duration 35 hours
Course Outline
Introduction, Goals, and Migration Strategy
- Course objectives, alignment with participant profiles, and definition of success metrics
- Overview of high-level migration approaches and associated risk factors
- Configuration of workspaces, repositories, and lab-specific datasets
Day 1 — Core Migration Concepts and Architecture
- Foundations of Lakehouse, Delta Lake introduction, and Databricks system architecture
- Comparison of SMP and MPP paradigms and their impact on migration strategies
- Design of the Medallion (Bronze→Silver→Gold) architecture and an overview of Unity Catalog
Day 1 Lab — Converting a Stored Procedure
- Practical migration of a sample stored procedure into a notebook format
- Mapping temporary tables and cursors to DataFrame transformations
- Verification and comparison against the original output
Day 2 — Advanced Delta Lake & Incremental Data Loading
- ACID transactions, commit logs, version control, and time travel capabilities
- Utilization of Auto Loader, MERGE INTO patterns, upsert operations, and schema evolution
- Application of OPTIMIZE, VACUUM, Z-ORDER, partitioning, and storage optimization techniques
Day 2 Lab — Incremental Ingestion & Performance Optimization
- Implementation of Auto Loader ingestion and MERGE workflows
- Application of OPTIMIZE, Z-ORDER, and VACUUM commands; result validation
- Assessment of read/write performance enhancements
Day 3 — SQL in Databricks, Performance Analysis & Debugging
- Advanced analytical SQL features: window functions, higher-order functions, and JSON/array manipulation
- Interpretation of Spark UI, DAGs, shuffles, stages, and tasks for bottleneck identification
- Query optimization strategies: broadcast joins, hints, caching, and reduction of data spill
Day 3 Lab — SQL Refactoring & Performance Tuning
- Refactoring of complex SQL processes into optimized Spark SQL queries
- Use of Spark UI traces to pinpoint and resolve data skew and shuffle problems
- Conducting before/after benchmarks and documenting tuning steps
Day 4 — Practical PySpark: Replacing Procedural Logic
- Spark execution model: driver, executors, lazy evaluation, and partitioning methods
- Conversion of loops and cursors into vectorized DataFrame operations
- Code modularization, UDFs/pandas UDFs, widgets, and creation of reusable libraries
Day 4 Lab — Refactoring Procedural Scripts
- Transformation of procedural ETL scripts into modular PySpark notebooks
- Introduction of parameterization, unit-style testing, and reusable functions
- Code review and application of best-practice checklists
Day 5 — Orchestration, End-to-End Pipeline & Best Practices
- Databricks Workflows: job design, task dependencies, triggers, and error management
- Construction of incremental Medallion pipelines with quality rules and schema validation
- Integration with Git (GitHub/Azure DevOps), CI pipelines, and testing strategies for PySpark logic
Day 5 Lab — Constructing a Complete End-to-End Pipeline
- Assembly of a Bronze→Silver→Gold pipeline orchestrated via Workflows
- Implementation of logging, auditing, retry mechanisms, and automated validations
- Execution of the full pipeline, output validation, and preparation of deployment documentation
Operationalization, Governance, and Production Readiness
- Unity Catalog governance, data lineage, and access control best practices
- Cost management, cluster sizing, autoscaling, and job concurrency patterns
- Deployment checklists, rollback strategies, and runbook development
Final Review, Knowledge Transfer, and Next Steps
- Participant presentations showcasing migration work and key takeaways
- Gap analysis, suggestions for follow-up activities, and handover of training materials
- References, pathways for further learning, and support options
Requirements
- A solid grasp of core data engineering principles.
- Practical experience with SQL and stored procedures (Synapse / SQL Server).
- Knowledge of ETL orchestration concepts (ADF or comparable tools).
Target Audience
- Technology managers possessing a background in data engineering.
- Data engineers looking to shift procedural OLAP logic toward Lakehouse patterns.
- Platform engineers tasked with driving Databricks adoption.