AI: Machine Learning & Model Training
Large-scale & multi-GPU training | Sri AI
Advanced
5 views
Course overview
Training across many GPUs without wasting them: parallelism, memory tricks and keeping a long run alive.
Level: Advanced · Mode: Part-time
Who this course is for
Engineers training models too big for one card.
What you will learn
- Use data, tensor and pipeline parallelism
- Train with DeepSpeed and FSDP
- Recover from crashed runs
- Budget compute honestly
Syllabus
- Module 1: Why scale is hard
- Module 2: Data parallelism
- Module 3: Sharding with FSDP and ZeRO
- Module 4: Mixed precision
- Module 5: Checkpointing and recovery
- Module 6: Cluster economics
Final project
Every module ends in hands-on practice, and the course ends with a project you build and present. Your certificate names that project.
Before you start
PyTorch from zero and Deep learning.