Distributed training

Infra behind Krea 2: How to train and serve at scale — Gabriel Jorge Menezes, Krea.ai

Infra behind Krea 2: How to train and serve at scale — Gabriel Jorge Menezes, Krea.ai

Krea's Gabriel Jorge Menezes shares critical insights into building the infrastructure for large-scale ML model training and serving. Key takeaways include the necessity of custom metrics beyond standard GPU utilization, aggressive checkpointing on ultra-fast storage to counter frequent cluster crashes, and a dynamic Kubernetes-based system using gang scheduling, virtual-kubelet, and taints to seamlessly shift inference workloads to external providers when training consumes on-prem GPUs. The approach highlights practical solutions for silent failures, thermal management, and optimizing resource utilization in a unified production and training environment.

The Messy Reality of Scale: Synthetic Data and Pre-Training — Marah Abdin & Robert McHardy, poolside

The Messy Reality of Scale: Synthetic Data and Pre-Training — Marah Abdin & Robert McHardy, poolside

Poolside discusses their innovative approaches to synthetic data generation, pre-training validation, and distributed training challenges. They highlight how modular data pipelines, rigorous replica hash checks, and numerical stability fixes enabled them to scale their LLMs, culminating in the 118B parameter Laguna S model designed for agentic coding, which shows strong early results against leading open-weight models.

The Prime Intellect Stack — Will Brown, Prime Intellect

The Prime Intellect Stack — Will Brown, Prime Intellect

Deep dive into Prime Intellect's open-source ecosystem for post-training LLMs, covering the modular Verifiers V1 environment design, the asynchronous and scalable Primer RL training framework, and the Lab platform for hosted training and fine-tuning. Learn about advanced reward systems, the interception server pattern, and tokenization control with the Renderers library, all designed to enable frontier agentic model development.

The 100,000 Sandbox Problem — Akshat Bubna, Modal CTO

The 100,000 Sandbox Problem — Akshat Bubna, Modal CTO

Modal CTO Akshat Bubna discusses the company's shift from developer to agent experience, highlighting why traditional cloud infrastructure fails for bursty AI workloads. He details Modal's primitives like elastic inference with GPU snapshotting and speculative decoding, agent sandboxes for RL rollouts, multi-node training with RDMA, and a "supercloud" strategy across 17 providers. The conversation also covers the importance of observability, hard guardrails for production agents, and AI's role in making infrastructure exciting again.

How Cursor Trained Composer on Fireworks: Distributed Infrastructure for High-Performance RL

How Cursor Trained Composer on Fireworks: Distributed Infrastructure for High-Performance RL

Cursor's Federico Cassano and Fireworks' Dmytro Dzhulgakov detail their collaboration on Composer 2, a specialized foundation model for software engineering. They discuss their top-down training strategy, the infrastructure challenges of large-scale distributed Reinforcement Learning on sparse models, and how model specialization achieves frontier performance with superior efficiency.

Granite 4.1, IBM Bob & building a quantum ecosystem

Granite 4.1, IBM Bob & building a quantum ecosystem

This episode of Mixture of Experts breaks down IBM's enterprise-focused Granite 4.1 and Project Bob, Google DeepMind's DiLoCo distributed training method, the inference-efficient DeepSeek V4 model, and IBM's strategy for achieving quantum advantage through strategic partnerships.