Course Outline
Fundamentals of the Databricks Platform and Lakehouse
- Understanding the Databricks Lakehouse architecture and its core components.
- Organizing workspaces and managing catalogs.
Databricks Workspace and Notebooks
- Navigating the workspace and engaging in notebook-based development.
- Structuring code for reusability through notebooks.
Apache Spark Architecture and Execution Model
\r- Exploring Spark runtime architecture and execution mechanics.
- Understanding lazy evaluation and the Job Directed Acyclic Graph (DAG).
PySpark DataFrames and the DataFrame API
- DataFrame abstractions and schema definitions.
- Core DataFrame operations and column expressions.
Translating SQL to PySpark DataFrames
- Mapping core SQL clauses to corresponding DataFrame operations.
- Utilizing window functions and aggregations within PySpark.
Data Ingestion and Output in Databricks
- Reading data from various file formats and database sources.
- Writing data and managing partitioning within the Lakehouse.
Delta Lake and Table Management
- Working with Delta tables and ACID transactions.
- Leveraging time travel and schema evolution features.
Data Cleaning and Transformation Patterns
- Executing data cleaning tasks and type conversions.
- Developing reusable transformation logic.
User-Defined Functions and Modular Code
- Implementing Python UDFs and pandas UDFs.
- Converting procedural logic into modular functions.
Performance Tuning and Optimization
- Strategies for partitioning and caching.
- Identifying bottlenecks using the Spark UI.
Fundamentals of Structured Streaming
- Comparing batch versus streaming processing models.
- Working with Streaming DataFrames and basic aggregations.
Databricks Jobs and Workflow Orchestration
- Scheduling notebooks as jobs and tasks.
- Constructing multi-step workflows with defined dependencies.
Unity Catalog and Data Governance
- Understanding Unity Catalog architecture and namespaces.
- Managing access control and data lineage.
Testing, Debugging, and Production Best Practices
- Conducting unit testing for PySpark logic.
- Implementing debugging techniques and code quality standards.
End-to-End Financial Services Use Cases
- Developing a comprehensive banking ETL pipeline.
- Converting legacy SQL processes to PySpark.
Migrating SQL Workloads to PySpark
- Adopting migration strategies and planning patterns.
- Performing incremental conversion of SQL workflows to PySpark.
Requirements
- Proficiency in Python programming, including knowledge of functions and data types.
- Comprehensive understanding of SQL, covering joins, aggregations, and subqueries.
- No previous exposure to Databricks or PySpark is necessary.
Target Audience
- Data engineers, data analysts, and broader data professionals.
- Teams transitioning existing SQL-based workflows to Databricks and PySpark.
Testimonials (1)
I liked that it was practical. Loved to apply the theoretical knowledge with practical examples.