Continual learning

The mathematics of AI uncertainty

The mathematics of AI uncertainty

Zoubin Ghahramani, a leading researcher at Google DeepMind and professor at Cambridge, argues that incorporating uncertainty is a missing piece for ever-improving AI. He discusses the critical difference between correctness and confidence in AI, tracing the historical evolution of probabilistic models from early neural networks to modern Bayesian approaches. Ghahramani highlights how current large language models often 'fake' uncertainty and explores successful implementations in areas like weather forecasting and AlphaFold, ultimately advocating for architectural innovations over pure scale to build more robust, trustworthy, and human-aligned intelligent systems that understand their own limitations.

Rich Sutton and Khurram Javed: Why AI Models Stop Learning, and How to Start It Again

Rich Sutton and Khurram Javed: Why AI Models Stop Learning, and How to Start It Again

Rich Sutton and Khurram Javed discuss their radical vision for AI at Oak Lab, advocating for truly continual learning agents that learn from their own experience, rejecting synthetic data due to the "Big World Hypothesis," and outlining a path to overcome catastrophic forgetting with "continual backprop" for a trillion-parameter, self-maintaining mind, while critiquing LLMs as only a fraction of intelligence.

Continual Learning: How AI Agents Get Better With Every Use | Arjun Karanam, Trajectory

Continual Learning: How AI Agents Get Better With Every Use | Arjun Karanam, Trajectory

Arjun Karanam from Trajectory discusses the "experience gap" in AI, where models excel in intelligence but lack real-world experience, advocating for continual learning. He outlines four key areas for the agent ecosystem: robust traceability including corrective actions, evaluations drawn from production traffic, harnesses that orchestrate rather than constrain, and comfort with open-weight models. Trajectory aims to provide a platform for companies to own and continuously improve their AI intelligence.

How Harvey Built a Research Lab on a Budget | Gabe Pereyra

How Harvey Built a Research Lab on a Budget | Gabe Pereyra

Gabe Pereyra of Harvey details a playbook for application companies to compete with frontier AI labs by leveraging the ecosystem. Key strategies include building specialized benchmarks like Legal Agent Bench, using domain experts for synthetic data generation to overcome sensitive client data issues, partnering with multiple 'neo labs' for post-training, and developing robust model serving and evaluation infrastructure. He emphasizes open-sourcing data for validation and the 'Moneyball' philosophy for success, addressing challenges like talent acquisition and long-context management in the Q&A.

Reinforcement Learning without Verifiable Rewards — Will Brown, Prime Intellect

Reinforcement Learning without Verifiable Rewards — Will Brown, Prime Intellect

Will Brown's talk addresses the critical challenge of applying Reinforcement Learning to real-world tasks where verifiable rewards are absent. He outlines how Primordial AI tackles this by leveraging environments as the core anchor for building reward signals. Key strategies include using LLMs as "judges," grounding tasks in production traces or document corpora, and employing "reverse direction" techniques to generate training data. Brown also details methods for calibrating task difficulty, identifying reward hacking, and fostering continual learning by treating model optimization as a science, emphasizing the use of compute to refine environmental signals and abstract human expertise.

Building the Automated AGI Lab: Core Automation's Jerry Tworek and Rohan Anil

Building the Automated AGI Lab: Core Automation's Jerry Tworek and Rohan Anil

Jerry Tworek and Rohan Anil, founders of Core Automation, argue that the transformer architecture has reached its limits and the primary bottleneck to smarter AI systems is now architectural. They contend that current models lack continual learning and test-time adaptation capabilities, critical for real-world deployment. They advocate for new architectures that learn from experience more efficiently than current reinforcement learning, optimize pre-training and RL end-to-end, and build a highly automated lab focused on accelerating innovation through kernel generation, aiming for models that can improve themselves without human intervention.