Gpu

Performance Optimization and Software/Hardware Co-design across PyTorch, CUDA, and NVIDIA GPUs

Performance Optimization and Software/Hardware Co-design across PyTorch, CUDA, and NVIDIA GPUs

Chris Fregly discusses his new book, "AI Systems Performance Engineering", covering the co-design and optimization of hardware, software, and algorithms across PyTorch, CUDA, and NVIDIA GPUs. The talk explores GPU architecture, system-level reliability challenges, and the use of modern coding agents for low-level kernel optimization.

Why AI Engineers Need to Understand GPU Hardware (with Chris Fregly)

Why AI Engineers Need to Understand GPU Hardware (with Chris Fregly)

Chris Fregly, author of 'AI Systems Performance Engineering', explains that true performance gains in AI come not from raw compute but from a deep, holistic understanding of the entire hardware and software stack. He emphasizes that memory bandwidth is the most critical GPU metric and introduces the concept of 'mechanical sympathy'—the co-design of hardware, software, and algorithms—as the key to unlocking efficiency and overcoming modern bottlenecks.

How Capital is Powering the AI Infrastructure Buildout with Magnetar Capital's Neil Tiwari

How Capital is Powering the AI Infrastructure Buildout with Magnetar Capital's Neil Tiwari

Neil Tiwari of Magnetar Capital explains the creative debt structures and financial innovations fueling the multi-trillion dollar AI infrastructure buildout. He debunks the myths around GPU collateral, revealing that the real security lies in contracted cash flows from investment-grade partners, and details how the industry's bottlenecks are shifting from chips to power distribution, steel, and specialized labor.

Accelerating Growth Through Optimizing GPU Usage // Sahil Khanna // AI in Production 2025

Accelerating Growth Through Optimizing GPU Usage // Sahil Khanna // AI in Production 2025

Adobe's journey in building a sophisticated AI Compute Platform to tackle the immense challenges of GPU optimization for training large-scale generative models like Firefly. The talk covers their custom-built solutions for resource management, developer productivity, and automated fault tolerance.

Inside the $41B AI Cloud Challenging Big Tech | CoreWeave SVP

Inside the $41B AI Cloud Challenging Big Tech | CoreWeave SVP

Corey Sanders, SVP of Product at CoreWeave, explains why the unique, high-stakes demands of AI workloads are driving a shift away from general-purpose clouds toward specialized "Neo Clouds." He details the specific hardware and software innovations in storage, cooling, and networking that allow CoreWeave to maximize GPU utilization and deliver superior performance, arguing that this focused approach creates a durable competitive advantage.

Efficient Reinforcement Learning – Rhythm Garg & Linden Li, Applied Compute

Efficient Reinforcement Learning – Rhythm Garg & Linden Li, Applied Compute

A deep dive into the challenges and solutions for efficient Reinforcement Learning (RL) in enterprise settings. The talk contrasts synchronous and asynchronous RL, explains the critical trade-off of "staleness" versus stability, and details a first-principles system model used to optimize GPU allocation for maximum throughput.