Description: This advanced level course will cover how to leverage PyTorch Fully Sharded Data Parallel (FSDP) to train large models in a distributed fashion on HPC systems. FSDP utilizes model sharding to reduce the memory footprint of traditional data parallel training at the added cost of additional communication overhead. This course will cover the various performance tradeoffs and how settings can be right-sized for a particular model to ensure optimal performance is obtained without overflowing GPU memory. Examples will be covered with differing model sizes, with a particular attention to vision models. You will learn how various FSDP settings affect memory usage and communication overhead through these examples. In addition to FSDP, other memory-efficient training strategies will be discussed, including activation checkpointing and automated mixed-precision.

Presenter(s): Dr. Mathew Boyer, GDIT / PET
Location: Webcast

Date: June 24, 2025

Controlled by: DoD HPCMP
Controlled by: PET Program
CUI Category: OPSEC
Limited Dissemination Control: FEDCON
POC: Mr. Ronald Hedgepeth, pet@hpc.mil

CUI

Search Terms:

Artificial Intelligence, AI, Machine Learning, ML


technical_area: AI/ML