Description: A Technical Overview and Demo of Kubeflow within the ARL SCOUT HPC system. The first part of this training provides an information and example on creating/using Persistent Volume Claims for Pipeline Tasks I/O, Natural Language Processing using a BERT language model, and Distributed Training with samples provided using both TFJob and MPIJob operators. As PyTorch requires a different notibook server image the setup of a PyTorch Notebook is included. An exercise using the Kubernetes Command Line Interface provided by Kubernetes and it's configuration is also provided. The final part of this training walks the user through a demo of using Kubeflow on SCOUT covering all of these topics.

Presenter(s): Jim VanOosten, IBM
Date & Time: August 26, 2022 (Update from August 10, 2022)

Controlled by: DoD HPCMP
Controlled by: PET Program
CUI Category: OPSEC
Limited Dissemination Control: FEDCON
POC: Mr. Ronald Hedgepeth, pet@hpc.mil

CUI