Machine learning benchmarks

The Benchmark With No Instructions — Tufa Labs (ARC-AGI-3)

The Benchmark With No Instructions — Tufa Labs (ARC-AGI-3)

Tim Scarfe visits Tufa Labs to explore their top-ranking ARC-AGI-3 system, a benchmark for agentic intelligence that challenges LLMs in goal discovery and action efficiency. The team delves into the complexities of fractured representations, the role of human priors, and whether LLMs truly plan or merely simulate it effectively, all while balancing the bitter lesson with AI safety concerns.

Why Tejal Patwardhan stopped underestimating the models - Episode 21

Why Tejal Patwardhan stopped underestimating the models - Episode 21

Tejal Patwardhan, head of OpenAI's frontier evals team, discusses the critical evolution of AI evaluations. She explains why traditional benchmarks fail as models become more capable, how OpenAI develops realistic, long-horizon tests (including groundbreaking wet lab experiments), and the implications of rapidly advancing multimodal and reasoning models for scientific discovery and the future of human work.