
Description: This advanced level course will cover how to leverage PyTorch Fully Sharded Data Parallel (FSDP) to train large models in a distributed fashion on HPC systems. FSDP utilizes model sharding to reduce the memory footprint of traditional data parallel training at the added cost of additional communication overhead. This course will cover the various performance tradeoffs and how settings can be right-sized for a particular model to ensure optimal performance is obtained without overflowing GPU memory. Examples will be covered with differing model sizes, with a particular attention to vision models. You will learn how various FSDP settings affect memory usage and communication overhead through these examples. In addition to FSDP, other memory-efficient training strategies will be discussed, including activation checkpointing and automated mixed-precision.
| Presenter(s): Dr. Mathew Boyer, GDIT / PET Location: Webcast Date: June 24, 2025 |
Controlled by: DoD HPCMP Controlled by: PET Program CUI Category: OPSEC Limited Dissemination Control: FEDCON POC: Mr. Ronald Hedgepeth, pet@hpc.mil |
CUI
Artificial Intelligence, AI, Machine Learning, ML
