Description: A Technical Overview and Demo of Kubeflow within the ARL SCOUT HPC system. The first part of this training provides an information and example on creating/using Persistent Volume Claims for Pipeline Tasks I/O, Natural Language Processing using a BERT language model, and Distributed Training with samples provided using both TFJob and MPIJob operators. As PyTorch requires a different notibook server image the setup of a PyTorch Notebook is included. An exercise using the Kubernetes Command Line Interface provided by Kubernetes and it's configuration is also provided. The final part of this training walks the user through a demo of using Kubeflow on SCOUT covering all of these topics.
| Presenter(s): Jim VanOosten, IBM Date & Time: August 26, 2022 (Update from August 10, 2022) |
Controlled by: DoD HPCMP Controlled by: PET Program CUI Category: OPSEC Limited Dissemination Control: FEDCON POC: Mr. Ronald Hedgepeth, pet@hpc.mil |
CUI
Course ID number for Global Search: TE1430_Archive
