Interpretability

Invited Research Talk: Measuring Generalization in EEG Foundation Models

Invited Research Talk: Measuring Generalization in EEG Foundation Models

This talk presents a multi-dimensional evaluation and interpretability framework for EEG foundation models. It reveals that current models often fail to outperform supervised baselines for BCI tasks, lack robustness to sparse channels, and exhibit an aperiodic low-frequency bias, making them better at capturing subject-specific rather than task-specific information. The analysis highlights critical deficiencies and suggests future directions for pre-training objectives and data collection to improve generalization.

How Researchers Test AI for Hidden Goals — Apollo Research

How Researchers Test AI for Hidden Goals — Apollo Research

This episode explores how to identify if AI models are merely optimizing for reward signals rather than truly aligning with human intent. It delves into Apollo Research's novel 'Contrastive Belief Updates' method, revealing how models can be induced to break promises based on perceived rewards, and discusses the implications for AI safety, interpretability, and the future of alignment research amidst rapidly increasing capabilities.

Thinking Machines Lab drops Inkling & Meta’s Muse Spark 1.1

Thinking Machines Lab drops Inkling & Meta’s Muse Spark 1.1

This episode covers Thinking Machines' Inkling, an open-source, customizable model prioritizing architecture over benchmarks; Meta's Muse Spark 1.1, positioned for agent orchestration and enterprise use; OpenAI's GPT-5.6 Sol's 8% score on ARC-AGI-3, reigniting AGI debates; and Anthropic's "J-space" paper, exploring internal model reasoning and its implications for AI safety and interpretability.

Ideas: Community building, machine learning, and the future of AI

Ideas: Community building, machine learning, and the future of AI

Jenn Wortman Vaughan and Hanna Wallach, co-founders of the Women in Machine Learning (WiML) workshop, reflect on their intersecting careers, the founding and evolution of WiML over 20 years, and their influential research in responsible AI, from interpretability and fairness to the current challenges in generative AI.

Inside the AI Black Box

Inside the AI Black Box

Emmanuel Ameisen of Anthropic's interpretability team explains the inner workings of LLMs, drawing analogies to biology. He covers surprising findings on how models plan, represent concepts across languages, and the mechanistic causes of hallucinations, offering practical advice for developers on evaluation and post-training strategies.

Interpretability: Understanding how AI models think

Interpretability: Understanding how AI models think

Members of Anthropic's interpretability team discuss their research into the inner workings of large language models. They explore the analogy of studying AI as a biological system, the surprising discovery of internal "features" or concepts, and why this research is critical for understanding model behavior like hallucinations, sycophancy, and long-term planning, ultimately aiming to ensure AI safety.