Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Duration 35 hours
Course Outline
Introduction, Objectives, and Migration Strategy
- Course aims, alignment with participant profiles, and success metrics
- Overview of high-level migration strategies and risk assessments
- Configuration of workspaces, repositories, and lab datasets
Day 1 — Migration Fundamentals and Architecture
- Lakehouse concepts, Delta Lake introduction, and Databricks architecture
- Differences between SMP and MPP and their impact on migration
- Medallion (Bronze→Silver→Gold) design principles and Unity Catalog overview
Day 1 Lab — Translating a Stored Procedure
- Practical migration of a sample stored procedure to a notebook
- Converting temporary tables and cursors into DataFrame transformations
- Verification and comparison against the original output
Day 2 — Advanced Delta Lake & Incremental Loading
- ACID transactions, commit logs, versioning, and time travel features
- Auto Loader, MERGE INTO patterns, upserts, and schema evolution
- OPTIMIZE, VACUUM, Z-ORDER, partitioning, and storage optimisation
Day 2 Lab — Incremental Ingestion & Optimization
- Implementing Auto Loader ingestion and MERGE workflows
- Applying OPTIMIZE, Z-ORDER, and VACUUM; validating outcomes
- Evaluating read/write performance enhancements
Day 3 — SQL in Databricks, Performance & Debugging
- Analytical SQL capabilities: window functions, higher-order functions, JSON/array processing
- Interpreting the Spark UI: DAGs, shuffles, stages, tasks, and bottleneck analysis
- Query optimisation techniques: broadcast joins, hints, caching, and spill mitigation
Day 3 Lab — SQL Refactoring & Performance Tuning
- Refactoring a resource-intensive SQL process into optimised Spark SQL
- Utilising Spark UI traces to identify and resolve skew and shuffle issues
- Benchmarking performance before and after tuning; documenting the process
Day 4 — Tactical PySpark: Replacing Procedural Logic
- Spark execution model: driver, executors, lazy evaluation, and partitioning strategies
- Converting loops and cursors into vectorised DataFrame operations
- Modularisation, UDFs/pandas UDFs, widgets, and reusable libraries
Day 4 Lab — Refactoring Procedural Scripts
- Refactoring a procedural ETL script into modular PySpark notebooks
- Incorporating parametrisation, unit-style tests, and reusable functions
- Code review and application of best-practice checklists
Day 5 — Orchestration, End-to-End Pipeline & Best Practices
- Databricks Workflows: job design, task dependencies, triggers, and error management
- Designing incremental Medallion pipelines with quality rules and schema validation
- Integration with Git (GitHub/Azure DevOps), CI, and testing strategies for PySpark logic
Day 5 Lab — Build a Complete End-to-End Pipeline
- Assembling a Bronze→Silver→Gold pipeline orchestrated via Workflows
- Implementing logging, auditing, retries, and automated validations
- Executing the full pipeline, verifying outputs, and preparing deployment notes
Operationalisation, Governance, and Production Readiness
- Unity Catalog governance, lineage, and access control best practices
- Cost management, cluster sizing, autoscaling, and job concurrency patterns
- Deployment checklists, rollback strategies, and runbook creation
Final Review, Knowledge Transfer, and Next Steps
- Participant presentations of migration work and key takeaways
- Gap analysis, recommended follow-up actions, and handover of training materials
- References, further learning pathways, and support options
Requirements
- A solid grasp of data engineering principles
- Proficiency in SQL and stored procedures (Synapse / SQL Server)
- Knowledge of ETL orchestration concepts (ADF or equivalent)
Target Audience
- Technology managers with a data engineering background
- Data engineers transitioning procedural OLAP logic to Lakehouse patterns
- Platform engineers overseeing Databricks adoption