Benchmarks

Teaching AI to Find Real Vulnerabilities — David Brumley, Bugcrowd

Teaching AI to Find Real Vulnerabilities — David Brumley, Bugcrowd

David Brumley discusses the challenges and solutions for teaching AI models to hack, drawing parallels with human learning. He introduces a 'ladder of tasks' approach for reinforcement learning and addresses the critical flaw of traditional benchmarks: measurement difficulties with multiple vulnerabilities and 'reward hacking.' His team's 'Audit Task' uses deterministic graders and precision/recall metrics for open-world assessment. He demonstrates this with an in-depth case study on attacking Chrome's V8 engine, showcasing how advanced models achieve real zero-day exploits, and warns against 'benchmaxxing security' without robust, honest grading.

Benchmarks: The Good, the Bad, and the Ugly — Ali Khial, G2i

Benchmarks: The Good, the Bad, and the Ugly — Ali Khial, G2i

Ali Khial exposes critical flaws in popular coding benchmarks, revealing how ambiguous instructions, weak verifiers, and model 'reward hacking' create a disconnect between reported performance and real-world utility. He argues that this leads to a "trust gap" where engineers disregard leaderboards. Khial then outlines five principles for building trustworthy, production-grade benchmarks, emphasizing human-authored instructions, holistic grading, economic value, contamination-free design, and informative leaderboards, urging software engineers to contribute to their improvement.

State of Data — Sean Cai, Independent / State of Data

State of Data — Sean Cai, Independent / State of Data

Sean Cai discusses the evolving landscape of AI data markets, highlighting the shift from raw annotation to high-quality, process-based data. He introduces "Verifier's Law" and its three axes to explain application layer maturity, critiques the shortcomings of current benchmarks, and predicts future AI trends by analyzing data market signals. Cai concludes by envisioning data companies transforming into "neo-labs" that provide enterprise-level reinforcement learning as a service and "Antikythera mechanisms" to manage and monetize real-world data assets.

How we taught agents to use good retrieval - Hanna Lichtenberg, Mixedbread AI

How we taught agents to use good retrieval - Hanna Lichtenberg, Mixedbread AI

Mixedbread AI addresses the "Oracle Gap" – the disparity between LLM reasoning and retrieval capabilities – by developing agents trained to use advanced search tools effectively. They demonstrate how current LLMs generate poor queries due to training biases and introduce a sophisticated agent harness with diverse search tools and a unique training regimen, including supervised fine-tuning and reinforcement learning with custom rewards, to teach agents to form precise semantic queries. This approach significantly improves performance on benchmarks like Oblique Congress and Snowflake's Match QA, closing the gap between theoretical perfect retrieval and real-world agent performance.

François Chollet: ARC-AGI-3, Beyond Deep Learning & A New Approach To ML

François Chollet: ARC-AGI-3, Beyond Deep Learning & A New Approach To ML

François Chollet discusses his contrarian approach to AI, moving beyond scaling LLMs to understand intelligence from first principles. He explains his work on the ARC benchmark series, including the new ARC-AGI V3, designed to measure 'agentic intelligence' and skill acquisition efficiency. He also introduces his lab, Ndea, which is developing a new ML paradigm based on symbolic models, and shares his perspective on the limits of current systems and the future path to AGI.

OpenAI, Oracle & AMD shake up AI

OpenAI, Oracle & AMD shake up AI

The panel discusses the shifting AI hardware landscape as Oracle and OpenAI bet on AMD, challenging Nvidia's dominance. They also analyze a US government report on the risks of the DeepSeek model, debate the viability of Reflection AI's new $2B open-source venture, and dissect the story of a VC fund replacing analysts with AI agents.