Info Session — Mentor-Led Data Science & AI Program

Register
Academy

Module 12

Big data with Spark

Spark DataFrames, distributed preprocessing, feature engineering at scale, model training, tuning with k-fold cross-validation and grid search.

Outcome

What you will be able to do

You can build and tune a Spark ML pipeline on data that does not fit on one machine.

Lessons

Work through these in order

Assessment

Module quiz — 70% to pass

6 questions mixing concept checks and short code-output problems. Graded on the server with per-question explanations afterwards, unlimited retakes, and a badge with a verification code the moment you pass.