Get in Touch

Course Outline

Introduction:

  • Apache Spark within the Hadoop Ecosystem
  • A brief introduction to Python and Scala

Foundational Concepts (Theory):

  • Architecture
  • RDDs
  • Transformations and Actions
  • Stages, Tasks, and Dependencies

Hands-On Workshop: Understanding Basics via Databricks:

  • Practical exercises using the RDD API
  • Essential action and transformation functions
  • Working with PairRDDs
  • Join operations
  • Caching strategies
  • Practical exercises using the DataFrame API
  • SparkSQL
  • DataFrame operations: select, filter, group, and sort
  • UDFs (User Defined Functions)
  • An overview of the DataSet API
  • Streaming concepts

Hands-On Workshop: Deployment Strategies via AWS:

  • Fundamentals of AWS Glue
  • Comparing AWS EMR and AWS Glue
  • Running example jobs in both environments
  • Evaluating the advantages and disadvantages of each

Additional Topics:

  • Introduction to Apache Airflow orchestration

Requirements

Programming skills (ideally with Python and Scala)

Familiarity with SQL basics

 21 Hours

Number of participants


Price per participant

Testimonials (3)

Provisional Upcoming Courses (Require 5+ participants)

Related Categories