Link copied to clipboard
Large-scale & multi-GPU training | Sri AI
AI: Machine Learning & Model Training

Large-scale & multi-GPU training | Sri AI

(0 reviews)
Advanced 5 views

Course overview

Training across many GPUs without wasting them: parallelism, memory tricks and keeping a long run alive.

Level: Advanced  ·  Mode: Part-time

Who this course is for

Engineers training models too big for one card.

What you will learn

  • Use data, tensor and pipeline parallelism
  • Train with DeepSpeed and FSDP
  • Recover from crashed runs
  • Budget compute honestly

Syllabus

  1. Module 1: Why scale is hard
  2. Module 2: Data parallelism
  3. Module 3: Sharding with FSDP and ZeRO
  4. Module 4: Mixed precision
  5. Module 5: Checkpointing and recovery
  6. Module 6: Cluster economics

Final project

Every module ends in hands-on practice, and the course ends with a project you build and present. Your certificate names that project.

Before you start

PyTorch from zero and Deep learning.

Open-source tools you will use

Rate This Course