Inference

LLM & AI Agent Benchmarks vs Reality: Why AI Applications Break

LLM & AI Agent Benchmarks vs Reality: Why AI Applications Break

LLM leaderboard scores often don't reflect real-world performance. This video explains why and outlines a comprehensive approach to evaluate AI systems, focusing on the critical balance of accuracy, latency, and cost. It details model and system evaluation techniques, including handling different inference phases, workload shapes, and specific considerations for AI agents, emphasizing the need for realistic testing over generic benchmarks.

Why Old GPUs Keep Gaining Value

Why Old GPUs Keep Gaining Value

Steve Hou, Head of Research at Silicon Data, unpacks the data behind the AI compute market, revealing persistent tightness in GPU rentals and rising residual values, despite talk of oversupply. He discusses the methodologies behind their GPU price indices and forward curves, the evolving hardware landscape beyond Nvidia (including AMD, Cerebras, and TPUs), and trends in LLM token economics. The conversation also delves into the complex financial aspects of AI data center buildouts, the rise of specialized models, China's emerging AI hardware ecosystem, and the growing importance of power constraints and distributed AI infrastructure.

The First Dedicated YC GPU Cluster - With Together AI

The First Dedicated YC GPU Cluster - With Together AI

YC and Together AI have partnered to launch the first dedicated YC GPU cluster, addressing the critical compute bottleneck faced by AI-native startups. This initiative provides flexible, cost-effective access to GPU resources, enabling companies from early-stage research to major players to train, fine-tune, and run inference on AI models, and mitigating the financial strain of long-term compute commitments.

The Story Behind Cerebras’ $63 Billion IPO with Founder and CEO Andrew Feldman

The Story Behind Cerebras’ $63 Billion IPO with Founder and CEO Andrew Feldman

Cerebras CEO Andrew Feldman discusses the company's journey from a contrarian bet on wafer-scale computing to a $63 billion public company. He details the technical breakthroughs, the challenge of being ahead of the market, and how the recent explosion in AI demand for fast inference validated their architecture, leading to a landmark $20 billion deal with OpenAI.

Baseten CEO Tuhin Srivastava on Custom Models, and Building the Inference Cloud

Baseten CEO Tuhin Srivastava on Custom Models, and Building the Inference Cloud

Baseten CEO Tuhin Srivastava discusses the explosive growth in AI inference, driven by the adoption of specialized and post-trained open-source models. He covers the strategic importance of owning the software layer on top of compute, navigating the severe GPU supply crunch with a multi-cloud fabric, the evolving landscape of AI workloads, and the operational lessons learned from scaling 30x in one year.

The Moonshot Podcast Season 2, Episode 6: Silicon Horizons

The Moonshot Podcast Season 2, Episode 6: Silicon Horizons

This podcast episode explores two X moonshot projects aimed at revolutionizing computer chips. Project Positron focused on creating specialized chips for real-time AI inference, acting as 'brains for robots'. Project Bodger took a meta-approach, using AI and inverse design to automate the chip design process itself, aiming to overcome the limitations of Moore's Law.