Ai safety

Understanding the inner thoughts of AI

Understanding the inner thoughts of AI

Neel Nanda, head of Google DeepMind's language model interpretability team, discusses the critical field of interpretability, likening it to the "neuroscience of AI." He explains why understanding the internal workings of "grown, not designed" neural networks is crucial for AI safety and scientific discovery. The episode explores cutting-edge techniques like Chain of Thought monitoring, mechanistic interpretability (steering and probing), and Sparse Autoencoders, highlighting their strengths and limitations in debugging, detecting deception, and uncovering hidden model objectives. Nanda emphasizes interpretability's role in building safe, aligned, and trustworthy AI as we approach AGI, acknowledging its pragmatic necessity despite inherent limits to full understanding.

The AI Threat Almost No One Is Working On (with Benjamin Todd)

The AI Threat Almost No One Is Working On (with Benjamin Todd)

Benjamin Todd provides an updated career strategy for the AI era, explaining why 'follow your passion' is flawed and what truly builds fulfillment. He details the ABZ framework for planning under deep uncertainty and the 'moving bottleneck' concept to stay valuable as AI advances. The discussion highlights that a human-level digital worker quickly becomes superhuman and maps critical AI risks including power-seeking AI, extreme power concentration, and engineered pandemics, emphasizing that individual careers can be a powerful lever for good in these transformative times.

GPT-5.6 Sol, FIFA AI & Wall Street’s AI nerves

GPT-5.6 Sol, FIFA AI & Wall Street’s AI nerves

OpenAI's new GPT-5.6 Sol model sparks debate on AI safety and release strategies, while Wall Street expresses growing skepticism over the long-term economics of frontier AI models. The discussion also touches on AI's impact on the FIFA World Cup and a thought-provoking paper comparing LLM anthropomorphism to Age of Empires II "goats."

The Benchmark With No Instructions — Tufa Labs (ARC-AGI-3)

The Benchmark With No Instructions — Tufa Labs (ARC-AGI-3)

Tim Scarfe visits Tufa Labs to explore their top-ranking ARC-AGI-3 system, a benchmark for agentic intelligence that challenges LLMs in goal discovery and action efficiency. The team delves into the complexities of fractured representations, the role of human priors, and whether LLMs truly plan or merely simulate it effectively, all while balancing the bitter lesson with AI safety concerns.

Fable 5: The Full Story from Capabilities to Drama (Ep. 1002 with Jon Krohn)

Fable 5: The Full Story from Capabilities to Drama (Ep. 1002 with Jon Krohn)

Anthropic's highly anticipated Claude Fable 5 model, a public version of its advanced "Mythos class" AI with state-of-the-art capabilities in software, vision, and long-context tasks, was released and then swiftly pulled offline by the U.S. government after just three days. The removal, initiated as an export control action over national security concerns stemming from a disputed "jailbreak" claim, highlights the growing tension between frontier AI development, AI safety, and regulatory oversight.

AI at college graduations and why Claude blackmails

AI at college graduations and why Claude blackmails

The Mixture of Experts team discusses the growing skepticism towards AI among younger generations, a Microsoft study revealing how LLMs can corrupt data in complex workflows, Anthropic's data-centric fix for Claude's "blackmailing" issue, and the cultural debate over an AI-generated story potentially winning a literary prize, all circling the central themes of human ownership, trust, and the need for better processes in the age of AI.