Sparse autoencoders

Training Krea 2: What matters in generative model training — Sangwu Lee, Krea.ai

Training Krea 2: What matters in generative model training — Sangwu Lee, Krea.ai

The talk by Sangwha Lee discusses Krea 2's open-source medium variant, emphasizing stylistic diversity and rapid iteration over the consistency-focused approach of larger models. A significant portion details their robust data curation pipeline, including unique methods for deduplication, filtering out AI-generated images, utilizing sparse autoencoders for unsupervised tagging, and ensuring world knowledge coverage. He outlines an LLM-inspired multi-stage training process, culminating in a prompt expander, and shares insights on fast iteration and future directions for image generation, highlighting the increasing integration of VLM advancements.

Understanding the inner thoughts of AI

Understanding the inner thoughts of AI

Neel Nanda, head of Google DeepMind's language model interpretability team, discusses the critical field of interpretability, likening it to the "neuroscience of AI." He explains why understanding the internal workings of "grown, not designed" neural networks is crucial for AI safety and scientific discovery. The episode explores cutting-edge techniques like Chain of Thought monitoring, mechanistic interpretability (steering and probing), and Sparse Autoencoders, highlighting their strengths and limitations in debugging, detecting deception, and uncovering hidden model objectives. Nanda emphasizes interpretability's role in building safe, aligned, and trustworthy AI as we approach AGI, acknowledging its pragmatic necessity despite inherent limits to full understanding.

Mapping the Mind of a Neural Net: Goodfire’s Eric Ho on the Future of Interpretability

Mapping the Mind of a Neural Net: Goodfire’s Eric Ho on the Future of Interpretability

Eric Ho, founder of Goodfire, discusses the critical challenge of AI interpretability. He shares how his team is developing techniques to understand, audit, and edit neural networks at the feature level, including breakthrough results in resolving superposition with sparse autoencoders, successful model editing demonstrations, and real-world applications in genomics with Arc Institute's DNA foundation models. Ho argues that these white-box approaches are essential for building safe, reliable, and intentionally designed AI systems.