Description: This training discusses GPU memory management through compiling optimization, CUDA unified memory and advanced memory management technology (prefetch, memory hints, oversubscription etc.). In this seminar we will focus on memory serial operation, blocking, stream overlap technology, and on how to avoid excess communication between CPU and GPU, and the latest approaches such as prefetching and memory advise technique. Sample codes will be adopted to demonstrate global memory and shared memory optimization, pinned memory and pageable memory comparison, and CUDA unified memory management. Additionally, NVIDIA profiling tools such as Nsight Systems and Nsight Compute will be used to address performance of computing, for instance, GPU memory workload analysis.

Presenter(s): Dr. Shaolin Mao, GDIT / PET
Location: Webcast
Date & Time: June 21, 2022, 2:00p - 4:00p ET

Distribution Statement D. Distribution authorized to the Department of Defense and U.S. DoD contractors only, Administrative or Operational Use, 21 June 2022.  Other requests for this document shall be referred to the High Performance Computing Modernization Office, 3909 Halls Ferry Road, Vicksburg, MS 39180.

technical_area: Emerging Hardware Exploration