Cerebras

Why Old GPUs Keep Gaining Value

Why Old GPUs Keep Gaining Value

Steve Hou, Head of Research at Silicon Data, unpacks the data behind the AI compute market, revealing persistent tightness in GPU rentals and rising residual values, despite talk of oversupply. He discusses the methodologies behind their GPU price indices and forward curves, the evolving hardware landscape beyond Nvidia (including AMD, Cerebras, and TPUs), and trends in LLM token economics. The conversation also delves into the complex financial aspects of AI data center buildouts, the rise of specialized models, China's emerging AI hardware ecosystem, and the growing importance of power constraints and distributed AI infrastructure.

Codex, Behind the Harness — Dominik Kundel, OpenAI

Codex, Behind the Harness — Dominik Kundel, OpenAI

Once GPT 5.3 Codex Spark achieved 1000 tokens/sec on Cerebras, network latency superseded inference as the bottleneck for agents. This talk details how the Codex harness addresses this and other agentic challenges through innovations like WebSocket mode for stateful context, deferred tools for efficient context construction, robust sandboxing (Seatbelt, Bubblewrap, custom Windows solution), and an auto-review subagent to mitigate approval fatigue while ensuring security. It also covers structured actions via 'apply patch' for file edits, shell tools for system interaction, and sophisticated long-horizon goal management, with most distinct features exposed through the open Responses API.

Podcast Crossover: AIE, AGI, frontier lab strategy with ​ ⁨@matthew_berman⁩  and @swyxtv

Podcast Crossover: AIE, AGI, frontier lab strategy with ​ ⁨@matthew_berman⁩ and @swyxtv

Matt, the organizer of the AI.engineer conference, shares insights into its origin, the challenges of early adoption, and its current value as a neutral ground for AI labs. He delves into AI hardware trends, discussing specialized chips like Etched, and gives a nuanced take on Anthropic's Fable, addressing performance concerns and compute limitations. The conversation then explores OpenAI's rumored equity offer to the US government, discussing implications for regulation and societal involvement. Matt shares his perspective on AI existential risk and alignment, emphasizing the need for pragmatic engineering solutions. Finally, he outlines the limitations of current LLMs, the critical need for data efficiency, and offers strategic advice for "Agent Labs" navigating the "model capability overhang" versus multi-model agnosticism.

Scaling the Next Paradigm of Heterogeneous Intelligence — Adrian Bertagnoli, Callosum

Scaling the Next Paradigm of Heterogeneous Intelligence — Adrian Bertagnoli, Callosum

Adrian Bertagnoli from Callosum argues that the era of scaling monolithic models on homogeneous GPU clusters is ending. He introduces "heterogeneous intelligence," a new paradigm where model architectures, chip types, and workflows are optimized together. By routing subtasks to the most efficient model and hardware, this approach achieves significant performance gains, as demonstrated by two key results: a 7x cost reduction in recursive reasoning tasks using Cerebras, and state-of-the-art performance on the Video Web Arena benchmark, outperforming leading GPT and Gemini models at a fraction of the cost and time.

The Story Behind Cerebras’ $63 Billion IPO with Founder and CEO Andrew Feldman

The Story Behind Cerebras’ $63 Billion IPO with Founder and CEO Andrew Feldman

Cerebras CEO Andrew Feldman discusses the company's journey from a contrarian bet on wafer-scale computing to a $63 billion public company. He details the technical breakthroughs, the challenge of being ahead of the market, and how the recent explosion in AI demand for fast inference validated their architecture, leading to a landmark $20 billion deal with OpenAI.