Get in Touch

Course Outline

Introduction:

  • Apache Spark within the Hadoop Ecosystem
  • Brief Overview of Python and Scala

Foundational Concepts (Theory):

  • Spark Architecture
  • RDD (Resilient Distributed Datasets)
  • Transformations vs. Actions
  • Stages, Tasks, and Dependencies

Exploring Fundamentals via Databricks (Hands-On Workshop):

  • Practical exercises utilizing the RDD API
  • Core action and transformation functions
  • PairRDD operations
  • Join operations
  • Caching strategies
  • Practical exercises utilizing the DataFrame API
  • SparkSQL integration
  • DataFrame operations: select, filter, group, and sort
  • UDF (User Defined Functions)
  • Introduction to the DataSet API
  • Spark Streaming

Deployment Strategies via AWS (Hands-On Workshop):

  • Core concepts of AWS Glue
  • Comparative analysis of AWS EMR and AWS Glue
  • Implementing example jobs in both environments
  • Evaluating the advantages and limitations of each service

Additional Topics:

  • Introduction to Apache Airflow for orchestration

Requirements

Programming proficiency, ideally in Python or Scala.

Fundamental knowledge of SQL.

 21 Hours

Testimonials (3)

Upcoming Courses

Related Categories