World models

AI Can't Learn The Way Humans Do - This Could Fix That

AI Can't Learn The Way Humans Do - This Could Fix That

This discussion explores world models as a promising path to solving sample efficiency in AI and achieving AGI. It contrasts deterministic control (Newtonian physics) with stochastic environments (RL), explaining the challenges posed by large action spaces in Go, self-driving, and robotics. The episode delves into how synthetic data, video diffusion models, and latent space architectures like JEPA are making world models practical, while also highlighting remaining open problems in physics modeling, real-time adaptation, and rich sensory integration.

Has AI Finally Cracked Time Series Forecasting?

Has AI Finally Cracked Time Series Forecasting?

Ameet Talwalkar, CMU professor and Datadog's Chief Scientist, traces the journey of time series foundation models from early skepticism to their current impact. He details Datadog's Toto V1 and V2, highlighting breakthroughs in zero-shot performance, scaling, and the crucial role of data mix. The discussion extends to the vision of 'world models' for observability, integrating diverse data for self-healing software systems, and concludes with insights on open-weights models and AI's transformative effect on computer science and academic research.

Why AI Agents Don't Actually Understand You — Danielle Perszyk, Amazon AGI Lab

Why AI Agents Don't Actually Understand You — Danielle Perszyk, Amazon AGI Lab

Daniel from Amazon AGI Lab details a cognitive science-driven vision for human-aligned AI, focusing on collective intelligence, real-time interaction, and redefining reliability through user mind modeling. He emphasizes aligning AI representations with human cognition to foster generalization, prevent reduced human agency, and revolutionize areas like education, advocating for diverse AI systems and frontier research over immediate productization.

Prompt to Pipeline: Building with Google's Gen Media Stack — Paige & Guillaume, Google DeepMind

Prompt to Pipeline: Building with Google's Gen Media Stack — Paige & Guillaume, Google DeepMind

A comprehensive overview of Google DeepMind's latest advancements, featuring Paige Bailey demonstrating Gemini 1.5 Flash's cost-effective video analysis and AI Studio's single-prompt app generation. Guillaume Vernade showcases a full generative media pipeline, turning a public domain book into an illustrated, animated, and scored project using Gemini, Nano Banana, VO, and LIA. Ian Valentine closes with the power of Gemma 4, demonstrating on-device, multi-agent code generation and debugging without cloud APIs.

Waymo's Dmitri Dolgov: 20 Million Rides and the Road to Full Autonomy

Waymo's Dmitri Dolgov: 20 Million Rides and the Road to Full Autonomy

Dmitri Dolgov, co-CEO of Waymo, discusses the 20-year journey from the DARPA challenge to full autonomy. He explains the Waymo Foundation Model—a multimodal world action model powering the driver, simulator, and critic—and how their "end-to-end plus" architecture enables superhuman safety and exponential scaling.

Robotics' End Game: Nvidia's Jim Fan

Robotics' End Game: Nvidia's Jim Fan

Jim Fan of Nvidia outlines the endgame for robotics, arguing it will mirror the successful playbook of Large Language Models. He introduces "The Great Parallel," a roadmap where World Models replace Language Models, and data collection shifts from limited teleoperation to scalable egocentric video, culminating in a future of physical APIs and automated research.