World simulators

Reinforcement Learning without Verifiable Rewards — Will Brown, Prime Intellect

Reinforcement Learning without Verifiable Rewards — Will Brown, Prime Intellect

Will Brown's talk addresses the critical challenge of applying Reinforcement Learning to real-world tasks where verifiable rewards are absent. He outlines how Primordial AI tackles this by leveraging environments as the core anchor for building reward signals. Key strategies include using LLMs as "judges," grounding tasks in production traces or document corpora, and employing "reverse direction" techniques to generate training data. Brown also details methods for calibrating task difficulty, identifying reward hacking, and fostering continual learning by treating model optimization as a science, emphasizing the use of compute to refine environmental signals and abstract human expertise.

OpenAI Sora 2 Team: How Generative Video Will Unlock Creativity and World Models

OpenAI Sora 2 Team: How Generative Video Will Unlock Creativity and World Models

The OpenAI Sora team discusses the technology behind their video generation model, from the diffusion transformers and space-time tokens that enable object permanence to their vision for Sora as a world simulator capable of scientific discovery. They also cover their product philosophy, which prioritizes creative inspiration over mindless consumption, and their strategy for building a new creator economy with IP holders.