
Player FM 앱으로 오프라인으로 전환하세요!
Training Machine Learning (ML) models on Kubernetes
Manage episode 421319868 series 3332465
In this episode of the Kubernetes Bytes podcast, Bhavin sits down with Bernie Wu, VP Strategic Partnerships and AI/CXL/Kubernetes Initiatives at Memverge. They discuss about how Kubernetes is the most popular platform to run AI model training and model inferencing jobs. The discussion dives into model training, talking about different phases of a DAG, and then talk about how Memverge can help users with efficient and cost-effective model checkpoints. The discussion goes into topics like saving costs by using spot instances, hot restart of training jobs, reclaiming unused GPU resources, etc.
Check out our website at https://kubernetesbytes.com/
Episode Sponsor: Nethopper
- Learn more about KAOPS: @nethopper.io
- For a supported-demo: [email protected]
- Try the free version of KAOPS now! https://mynethopper.com/auth
Cloud Native News:
- https://www.aquasec.com/blog/linguistic-lumberjack-understanding-cve-2024-4323-in-fluent-bit/
- https://kubernetes.io/blog/2024/05/20/completing-cloud-provider-migration/
- https://thenewstack.io/introducing-aks-automatic-managed-kubernetes-for-developers/
- https://www.harness.io/blog/harness-to-acquire-split
Show Links:
- https://www.linkedin.com/in/berniewu/
- https://criu.org/Main_Page
- https://memverge.com/
- https://youtu.be/tY8YOMRuqWI?si=yB3hHqLUpYPZ-KWN
- https://youtu.be/ND4seSKpJHI?si=shh0iuA9qC-dO6eb
Timestamps:
88 에피소드
Manage episode 421319868 series 3332465
In this episode of the Kubernetes Bytes podcast, Bhavin sits down with Bernie Wu, VP Strategic Partnerships and AI/CXL/Kubernetes Initiatives at Memverge. They discuss about how Kubernetes is the most popular platform to run AI model training and model inferencing jobs. The discussion dives into model training, talking about different phases of a DAG, and then talk about how Memverge can help users with efficient and cost-effective model checkpoints. The discussion goes into topics like saving costs by using spot instances, hot restart of training jobs, reclaiming unused GPU resources, etc.
Check out our website at https://kubernetesbytes.com/
Episode Sponsor: Nethopper
- Learn more about KAOPS: @nethopper.io
- For a supported-demo: [email protected]
- Try the free version of KAOPS now! https://mynethopper.com/auth
Cloud Native News:
- https://www.aquasec.com/blog/linguistic-lumberjack-understanding-cve-2024-4323-in-fluent-bit/
- https://kubernetes.io/blog/2024/05/20/completing-cloud-provider-migration/
- https://thenewstack.io/introducing-aks-automatic-managed-kubernetes-for-developers/
- https://www.harness.io/blog/harness-to-acquire-split
Show Links:
- https://www.linkedin.com/in/berniewu/
- https://criu.org/Main_Page
- https://memverge.com/
- https://youtu.be/tY8YOMRuqWI?si=yB3hHqLUpYPZ-KWN
- https://youtu.be/ND4seSKpJHI?si=shh0iuA9qC-dO6eb
Timestamps:
88 에피소드
모든 에피소드
×플레이어 FM에 오신것을 환영합니다!
플레이어 FM은 웹에서 고품질 팟캐스트를 검색하여 지금 바로 즐길 수 있도록 합니다. 최고의 팟캐스트 앱이며 Android, iPhone 및 웹에서도 작동합니다. 장치 간 구독 동기화를 위해 가입하세요.