Get in Touch
 Duration 35 hours

Course Outline

Fundamentals of the Databricks Platform and Lakehouse

  • Understanding the Databricks Lakehouse architecture and its core components.
  • Organizing workspaces and managing catalogs.

Databricks Workspace and Notebooks

  • Navigating the workspace and engaging in notebook-based development.
  • Structuring code for reusability through notebooks.

Apache Spark Architecture and Execution Model

\r
  • Exploring Spark runtime architecture and execution mechanics.
  • Understanding lazy evaluation and the Job Directed Acyclic Graph (DAG).

PySpark DataFrames and the DataFrame API

  • DataFrame abstractions and schema definitions.
  • Core DataFrame operations and column expressions.

Translating SQL to PySpark DataFrames

  • Mapping core SQL clauses to corresponding DataFrame operations.
  • Utilizing window functions and aggregations within PySpark.

Data Ingestion and Output in Databricks

  • Reading data from various file formats and database sources.
  • Writing data and managing partitioning within the Lakehouse.

Delta Lake and Table Management

  • Working with Delta tables and ACID transactions.
  • Leveraging time travel and schema evolution features.

Data Cleaning and Transformation Patterns

  • Executing data cleaning tasks and type conversions.
  • Developing reusable transformation logic.

User-Defined Functions and Modular Code

  • Implementing Python UDFs and pandas UDFs.
  • Converting procedural logic into modular functions.

Performance Tuning and Optimization

  • Strategies for partitioning and caching.
  • Identifying bottlenecks using the Spark UI.

Fundamentals of Structured Streaming

  • Comparing batch versus streaming processing models.
  • Working with Streaming DataFrames and basic aggregations.

Databricks Jobs and Workflow Orchestration

  • Scheduling notebooks as jobs and tasks.
  • Constructing multi-step workflows with defined dependencies.

Unity Catalog and Data Governance

  • Understanding Unity Catalog architecture and namespaces.
  • Managing access control and data lineage.

Testing, Debugging, and Production Best Practices

  • Conducting unit testing for PySpark logic.
  • Implementing debugging techniques and code quality standards.

End-to-End Financial Services Use Cases

  • Developing a comprehensive banking ETL pipeline.
  • Converting legacy SQL processes to PySpark.

Migrating SQL Workloads to PySpark

  • Adopting migration strategies and planning patterns.
  • Performing incremental conversion of SQL workflows to PySpark.

Requirements

  • Proficiency in Python programming, including knowledge of functions and data types.
  • Comprehensive understanding of SQL, covering joins, aggregations, and subqueries.
  • No previous exposure to Databricks or PySpark is necessary.

Target Audience

  • Data engineers, data analysts, and broader data professionals.
  • Teams transitioning existing SQL-based workflows to Databricks and PySpark.

Testimonials (1)

Upcoming Courses

Related Categories