Ai benchmarking

The Miranda Hypothesis: How Hamilton Poisoned Persona Evals - Jacob E. Thomas, Results Gen

The Miranda Hypothesis: How Hamilton Poisoned Persona Evals - Jacob E. Thomas, Results Gen

This talk exposes "Miranda distortion," a critical flaw in AI personas where models, influenced by modern cultural narratives, produce convincing but anachronistic outputs. Current evaluations fail to detect this, prioritizing fluency over fidelity. The speaker proposes "epistemic simulation"—a new paradigm grounded in corpus-bounded, temporally-anchored, and expert-evaluated reasoning—and introduces the "Prism Experiment." This rigorous, pre-registered protocol uses Abraham Lincoln to demonstrate how a weighted rubric, created by historians and de-emphasizing rhetorical fluency, can detect anachronism. It advocates for the "humanist in the loop" as a technical requirement to ensure AI personas are true to their documentary records, not just convincing.

Moonlake: Multimodal, Interactive, and Efficient World Models — with Fan-yun Sun and Chris Manning

Moonlake: Multimodal, Interactive, and Efficient World Models — with Fan-yun Sun and Chris Manning

Moonlake AI presents a distinctive approach to world modeling, prioritizing interactive, action-conditioned environments built on symbolic representations and game engines over purely pixel-based generative models. This method focuses on causal reasoning, long-term consistency, and programmable rendering (via their 'Reverie' diffusion model) to create dynamic, multiplayer worlds, positioning itself as a platform for training embodied AI and revolutionizing game development.

How METR measures Long Tasks and Experienced Open Source Dev Productivity - Joel Becker, METR

How METR measures Long Tasks and Experienced Open Source Dev Productivity - Joel Becker, METR

AI models show remarkable progress on benchmarks, yet a field study with experienced developers revealed no productivity gains. This summary explores the disconnect between lab results and real-world impact, examining the causal relationship between compute and AI capabilities, the nuances of the developer productivity study, and future directions for measuring what AI can truly do.