Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
Introduction:
- Apache Spark within the Hadoop Ecosystem
- Brief Overview of Python and Scala
Foundational Concepts (Theory):
- Spark Architecture
- RDD (Resilient Distributed Datasets)
- Transformations vs. Actions
- Stages, Tasks, and Dependencies
Exploring Fundamentals via Databricks (Hands-On Workshop):
- Practical exercises utilizing the RDD API
- Core action and transformation functions
- PairRDD operations
- Join operations
- Caching strategies
- Practical exercises utilizing the DataFrame API
- SparkSQL integration
- DataFrame operations: select, filter, group, and sort
- UDF (User Defined Functions)
- Introduction to the DataSet API
- Spark Streaming
Deployment Strategies via AWS (Hands-On Workshop):
- Core concepts of AWS Glue
- Comparative analysis of AWS EMR and AWS Glue
- Implementing example jobs in both environments
- Evaluating the advantages and limitations of each service
Additional Topics:
- Introduction to Apache Airflow for orchestration
Requirements
Programming proficiency, ideally in Python or Scala.
Fundamental knowledge of SQL.
21 Hours
Testimonials (3)
Having hands on session / assignments
Poornima Chenthamarakshan - Intelligent Medical Objects
Course - Apache Spark in the Cloud
1. Right balance between high level concepts and technical details. 2. Andras is very knowledgeable about his teaching. 3. Exercise
Steven Wu - Intelligent Medical Objects
Course - Apache Spark in the Cloud
Get to learn spark streaming , databricks and aws redshift