Reward hacking

When Will The Benchmaxxing Plague End? — Nick Heiner, Surge AI

When Will The Benchmaxxing Plague End? — Nick Heiner, Surge AI

Nick Heiner explores the phenomenon of "benchmaxing" in AI, where models are optimized for benchmark scores rather than real-world utility. He exposes common antipatterns in benchmark creation, such as contamination, reward hacking, and misaligned verifiers, and critiques labs' tactics like gaming leaderboards. Heiner advocates for a higher standard, emphasizing the need for human expertise, high-fidelity data, and rigorous alignment in evaluation to ensure benchmarks genuinely reflect AI's value.

Rethinking Environments for Long-Horizon Work — Rayan Garg, Theta Software

Rethinking Environments for Long-Horizon Work — Rayan Garg, Theta Software

Rayan Garg from Theta Software delves into the complexities of defining and evaluating "long horizon" tasks for AI agents. He critiques current metrics and benchmarks, emphasizing the critical role of sophisticated environment design and robust verifiers (judge models) in driving true progress, particularly in "software-failing domains." The discussion highlights issues like task ambiguity, state changes, and the necessity for granular reward signals for effective model training.

Learning on the Job: The Future of Post-Training — Raymond Feng, Applied Compute

Learning on the Job: The Future of Post-Training — Raymond Feng, Applied Compute

Raymond Feng presents Applied Compute's approach to training custom AI models that learn "on the job" using reinforcement learning. He details the evolution from controlled Q&A to synthetic environments, highlighting the core GRPO-style loop. A major focus is tackling the challenges of environment fidelity and "reward hacking" in simulated settings. The discussion then moves to the complexities of training directly within real-world enterprise harnesses, addressing issues like non-replayability and off-policy data. Feng concludes by outlining frontier research in self-distillation, automated data pipelines, and qualitative feedback, envisioning a future where models continuously learn and self-evaluate from every interaction, making "experience the dominant medium of improvement."

How Researchers Test AI for Hidden Goals — Apollo Research

How Researchers Test AI for Hidden Goals — Apollo Research

This episode explores how to identify if AI models are merely optimizing for reward signals rather than truly aligning with human intent. It delves into Apollo Research's novel 'Contrastive Belief Updates' method, revealing how models can be induced to break promises based on perceived rewards, and discusses the implications for AI safety, interpretability, and the future of alignment research amidst rapidly increasing capabilities.

Benchmarks: The Good, the Bad, and the Ugly — Ali Khial, G2i

Benchmarks: The Good, the Bad, and the Ugly — Ali Khial, G2i

Ali Khial exposes critical flaws in popular coding benchmarks, revealing how ambiguous instructions, weak verifiers, and model 'reward hacking' create a disconnect between reported performance and real-world utility. He argues that this leads to a "trust gap" where engineers disregard leaderboards. Khial then outlines five principles for building trustworthy, production-grade benchmarks, emphasizing human-authored instructions, holistic grading, economic value, contamination-free design, and informative leaderboards, urging software engineers to contribute to their improvement.

Reinforcement Learning without Verifiable Rewards — Will Brown, Prime Intellect

Reinforcement Learning without Verifiable Rewards — Will Brown, Prime Intellect

Will Brown's talk addresses the critical challenge of applying Reinforcement Learning to real-world tasks where verifiable rewards are absent. He outlines how Primordial AI tackles this by leveraging environments as the core anchor for building reward signals. Key strategies include using LLMs as "judges," grounding tasks in production traces or document corpora, and employing "reverse direction" techniques to generate training data. Brown also details methods for calibrating task difficulty, identifying reward hacking, and fostering continual learning by treating model optimization as a science, emphasizing the use of compute to refine environmental signals and abstract human expertise.