Benchmarking

Vending-Bench: Long-Horizon Agent Evals — Lukas Petersson, Andon Labs

Vending-Bench: Long-Horizon Agent Evals — Lukas Petersson, Andon Labs

Lukas Petersson from Andon Labs discusses their pioneering work in evaluating AI models in long-horizon, real-world, and hybrid environments. He highlights the "simulation awareness" problem in traditional benchmarks, the emergence of complex misbehaviors like collusion and rationalization, and ethical challenges in real-world deployments. A novel solution involves "forking" real environments into simulations to enable reproducible testing of critical AI behaviors.

Stop Evaluating Models Like It's the 50s - Alejandro Vidal, Mindmakers

Stop Evaluating Models Like It's the 50s - Alejandro Vidal, Mindmakers

This talk introduces the application of psychometrics, particularly Item Response Theory (IRT), to improve LLM evaluation. It highlights how IRT goes beyond simple accuracy to measure model ability, item difficulty, and discrimination, enabling benchmark auditing, adaptive testing, data leakage detection, and the identification of model relationships and distillation, ultimately providing deeper insights into what LLMs truly learn.

Task Fidelity Scaling Laws — Kobie Crawdord, Snorkel

Task Fidelity Scaling Laws — Kobie Crawdord, Snorkel

An experiment by Snorkel AI reveals that in agentic AI training, the quality of tasks is paramount. Using the same model and compute, fine-tuning on high-quality tasks yielded a 6% performance improvement, a 5x greater uplift compared to the 1% gain from low-quality tasks. The key difference lies in the nature of the tasks: high-quality tasks are genuinely harder, featuring more tool calls and cleaner failure modes that provide a meaningful learning signal. In contrast, low-quality tasks often fail due to ambiguity and environmental noise, hindering effective model improvement.

Agentic Evaluations at Scale, For Everybody — Nicholas Kang & Michael Aaron, Google DeepMind

Agentic Evaluations at Scale, For Everybody — Nicholas Kang & Michael Aaron, Google DeepMind

Nicholas Kang and Michael Aaron from Google DeepMind's Kaggle team discuss the broken state of AI evaluations—scattered, non-transparent, and created by a homogenous group. They present their solutions: a community-driven benchmarks platform, a PvP Game Arena for non-saturating ELO ratings, standardized agent exams, and hackathons to crowdsource novel evals and address the limitations of current benchmarking practices.

OpenClaw's Memory Sucks and the fix is simple — Dhravya Shah, Supermemory

OpenClaw's Memory Sucks and the fix is simple — Dhravya Shah, Supermemory

Dhravya Shah, founder of Super Memory, details the evolution of his company from a simple RAG-based consumer app to a sophisticated, open-source context infrastructure for AI, and introduces a novel hooks-based memory solution for OpenClaw.

Are AI Benchmarks Telling The Full Story? [SPONSORED]

Are AI Benchmarks Telling The Full Story? [SPONSORED]

AI models are often benchmarked like Formula 1 cars, excelling on technical exams but failing the test of daily human experience. Researchers Andrew Gordon and Nora Petrova from Prolific critique the 'leaderboard illusion' of current ranking systems and introduce their HUMAINE leaderboard, a new framework that uses census-based sampling and the TrueSkill algorithm to measure how helpful, safe, and relatable models are to real people, not just tech enthusiasts.