Inference

LLM Compression Explained: Build Faster, Efficient AI Models

LLM Compression Explained: Build Faster, Efficient AI Models

Learn how AI model compression and quantization techniques are essential for optimizing Large Language Model (LLM) performance and significantly reducing inference costs in production. This deep dive covers practical examples, benefits like reduced latency and increased throughput, and strategies for different AI use cases, demonstrating how to deploy scalable AI with minimal accuracy degradation.

Greetings, Earthlings: Philip Johnston of Starcloud on Data Centers in Space

Greetings, Earthlings: Philip Johnston of Starcloud on Data Centers in Space

Philip Johnston of Starcloud argues that space will become the primary location for AI compute within a decade. He explains how plummeting launch costs, superior solar energy economics in orbit, and the physics of heat dissipation will soon make space-based data centers cheaper and more scalable than their terrestrial counterparts, predicting a future where nearly a trillion dollars in annual CapEx shifts to space.

How Capital is Powering the AI Infrastructure Buildout with Magnetar Capital's Neil Tiwari

How Capital is Powering the AI Infrastructure Buildout with Magnetar Capital's Neil Tiwari

Neil Tiwari of Magnetar Capital explains the creative debt structures and financial innovations fueling the multi-trillion dollar AI infrastructure buildout. He debunks the myths around GPU collateral, revealing that the real security lies in contracted cash flows from investment-grade partners, and details how the industry's bottlenecks are shifting from chips to power distribution, steel, and specialized labor.

The CEO Behind the Fastest-Growing AI Inference Company | Tuhin Srivastava

The CEO Behind the Fastest-Growing AI Inference Company | Tuhin Srivastava

Tuhin Srivastava, CEO of Baseten, joins Gradient Dissent to discuss the core challenges of AI inference, from infrastructure and runtime bottlenecks to the practical differences between vLLM, TensorRT-LLM, and SGLang. He shares how Baseten navigated years of searching for a market before the explosion of large-scale models, emphasizing a company-building philosophy focused on avoiding premature scaling and "burning the boats" to chase the biggest opportunities.

Building the Real-World Infrastructure for AI, with Google, Cisco & a16z

Building the Real-World Infrastructure for AI, with Google, Cisco & a16z

AI is driving an unprecedented buildout of physical infrastructure. Experts from Google and Cisco discuss the "AI industrial revolution," where power, compute, and networking are the new scarce resources, demanding a complete reinvention of the technology stack from silicon to software.

Nvidia CTO Michael Kagan: Scaling Beyond Moore's Law to Million-GPU Clusters

Nvidia CTO Michael Kagan: Scaling Beyond Moore's Law to Million-GPU Clusters

Nvidia CTO Michael Kagan explains how the Mellanox acquisition was key to scaling AI infrastructure from single GPUs to million-GPU data centers. He covers the critical role of networking in system performance, the shift from training to inference workloads, and his vision for AI's future in scientific discovery.