Latent space

AI Can't Learn The Way Humans Do - This Could Fix That

AI Can't Learn The Way Humans Do - This Could Fix That

This discussion explores world models as a promising path to solving sample efficiency in AI and achieving AGI. It contrasts deterministic control (Newtonian physics) with stochastic environments (RL), explaining the challenges posed by large action spaces in Go, self-driving, and robotics. The episode delves into how synthetic data, video diffusion models, and latent space architectures like JEPA are making world models practical, while also highlighting remaining open problems in physics modeling, real-time adaptation, and rich sensory integration.

Building Generative Image & Video models at Scale - Sander Dieleman (Veo and Nano Banana)

Building Generative Image & Video models at Scale - Sander Dieleman (Veo and Nano Banana)

Sander Dieleman from Google DeepMind provides a behind-the-scenes look at the key components of training large-scale diffusion models for audio-visual data. The talk covers the entire pipeline, from the critical role of data curation and latent representations to the mechanics of diffusion, network architectures, sampling with guidance, and advanced control signals.