Artificial Intelligence14 min read

Distributed training on OpenShift AI 3.4 with Kubeflow Trainer v2

By · Published by Everything Blog

In short

### Briefing #### Introduction Until recently, running distributed training on Kubernetes required custom resource definitions (CRDs) specific to each machine learning (ML) framework. These frameworks had different semantics, failure behaviors, and methods for configuring fundamental aspects such as how many nodes, how they find each other, and what happens when something breaks. Red Hat OpenShift AI 3.4 introduced Kubeflow Trainer v2, a unified training API that replaces framework-specific CRDs with a single `Tr…

Key points

  • ### Briefing #### Introduction Until recently, running distributed training on Kubernete…: ### Briefing #### Introduction Until recently, running distributed training on Kubernetes required custom resource definitions (CRDs) specific to each machine learning (ML) framework.
  • These frameworks had different semantics, failure behaviors, and methods for configuring…: These frameworks had different semantics, failure behaviors, and methods for configuring fundamental aspects such as how many nodes, how they find each other, and what happens when something breaks.
  • Red Hat OpenShift AI 3.4 introduced Kubeflow Trainer v2, a unified training API that repl…: Red Hat OpenShift AI 3.4 introduced Kubeflow Trainer v2, a unified training API that replaces framework-specific CRDs with a single `TrainJob` resource.

Original source: developers.redhat.com

Artificial Intelligence