Gpu infrastructure

Infra behind Krea 2: How to train and serve at scale — Gabriel Jorge Menezes, Krea.ai

Infra behind Krea 2: How to train and serve at scale — Gabriel Jorge Menezes, Krea.ai

Krea's Gabriel Jorge Menezes shares critical insights into building the infrastructure for large-scale ML model training and serving. Key takeaways include the necessity of custom metrics beyond standard GPU utilization, aggressive checkpointing on ultra-fast storage to counter frequent cluster crashes, and a dynamic Kubernetes-based system using gang scheduling, virtual-kubelet, and taints to seamlessly shift inference workloads to external providers when training consumes on-prem GPUs. The approach highlights practical solutions for silent failures, thermal management, and optimizing resource utilization in a unified production and training environment.

From PhD Side Project to $500M ARR: Will Falcon’s PyTorch Lightning Story

From PhD Side Project to $500M ARR: Will Falcon’s PyTorch Lightning Story

Will Falcon, CEO of Lightning AI, discusses the company's merger with Voltage Park to create a full-stack AI neo-cloud. He delves into the capabilities of Lightning AI Studio, a comprehensive platform for AI development, and recounts the origin story of PyTorch Lightning, from a personal research tool to a world-leading open-source framework.

Building Data Centers for GPU Clouds

Building Data Centers for GPU Clouds

Craig Tavares, COO of Buzz HPC, provides an in-depth look at the complexities of building and scaling GPU cloud infrastructure for AI. He covers the critical role of renewable energy and strategic location, the evolution of data center design to handle extreme power densities, the importance of a strong partnership with NVIDIA, and the rise of sovereign mandates shaping the future of AI cloud services.