Mechanistic interpretability

AI Security Costs Rise: Cost of a Data Breach Report & Claude Opus 5

AI Security Costs Rise: Cost of a Data Breach Report & Claude Opus 5

The podcast explores key AI developments, beginning with IBM's 2026 Cost of a Data Breach Report, highlighting AI's increasing role in both cyberattacks and defense, and the economic asymmetry it creates. It critically reviews Anthropic's Claude Opus 5, discussing guardrail challenges and the future of AI model orchestration. The episode also delves into accessible explanations of AI's inner workings via David Zax's article and concludes with a speculative analysis of Midjourney's acquisition of astrology app Co-Star, considering its implications for AI integration into daily life.

Boris Cherny: Building Claude Code

Boris Cherny: Building Claude Code

Boris Cherny, creator of Claude Code, discusses the transformative capabilities of Opus 5, highlighting its prompt injection resistance and long-task execution. He delves into Claude Code's empirical development philosophy of "unhobbling" AI by constantly adapting to new model generations, and shares insights on how to build advanced AI products using higher-level tasks, self-verification, and dynamic workflows to orchestrate thousands of agents.

Coding Agents Are Secretly General Agents

Coding Agents Are Secretly General Agents

Jay Hack, head of AI at ClickUp, discusses the evolution of AI from early computer vision to generalist coding agents, highlighting how 'positive transfer' makes coding an 'AGI-complete' domain. He delves into the brutal economics of AI startups facing foundation model giants, ClickUp's strategy for convergence and first-party data as a moat, and the challenges of verifiability and catastrophic forgetting. The conversation also explores LLMs at the scientific frontier, the 'car wash test' revealing limits of world models, and speculative future applications like LLM resorts and game integration.

Understanding the inner thoughts of AI

Understanding the inner thoughts of AI

Neel Nanda, head of Google DeepMind's language model interpretability team, discusses the critical field of interpretability, likening it to the "neuroscience of AI." He explains why understanding the internal workings of "grown, not designed" neural networks is crucial for AI safety and scientific discovery. The episode explores cutting-edge techniques like Chain of Thought monitoring, mechanistic interpretability (steering and probing), and Sparse Autoencoders, highlighting their strengths and limitations in debugging, detecting deception, and uncovering hidden model objectives. Nanda emphasizes interpretability's role in building safe, aligned, and trustworthy AI as we approach AGI, acknowledging its pragmatic necessity despite inherent limits to full understanding.

Mapping the Mind of a Neural Net: Goodfire’s Eric Ho on the Future of Interpretability

Mapping the Mind of a Neural Net: Goodfire’s Eric Ho on the Future of Interpretability

Eric Ho, founder of Goodfire, discusses the critical challenge of AI interpretability. He shares how his team is developing techniques to understand, audit, and edit neural networks at the feature level, including breakthrough results in resolving superposition with sparse autoencoders, successful model editing demonstrations, and real-world applications in genomics with Arc Institute's DNA foundation models. Ho argues that these white-box approaches are essential for building safe, reliable, and intentionally designed AI systems.