Model auditing

Understanding the inner thoughts of AI

Understanding the inner thoughts of AI

Neel Nanda, head of Google DeepMind's language model interpretability team, discusses the critical field of interpretability, likening it to the "neuroscience of AI." He explains why understanding the internal workings of "grown, not designed" neural networks is crucial for AI safety and scientific discovery. The episode explores cutting-edge techniques like Chain of Thought monitoring, mechanistic interpretability (steering and probing), and Sparse Autoencoders, highlighting their strengths and limitations in debugging, detecting deception, and uncovering hidden model objectives. Nanda emphasizes interpretability's role in building safe, aligned, and trustworthy AI as we approach AGI, acknowledging its pragmatic necessity despite inherent limits to full understanding.