Machine learning metrics

Stop Evaluating Models Like It's the 50s - Alejandro Vidal, Mindmakers

Stop Evaluating Models Like It's the 50s - Alejandro Vidal, Mindmakers

This talk introduces the application of psychometrics, particularly Item Response Theory (IRT), to improve LLM evaluation. It highlights how IRT goes beyond simple accuracy to measure model ability, item difficulty, and discrimination, enabling benchmark auditing, adaptive testing, data leakage detection, and the identification of model relationships and distillation, ultimately providing deeper insights into what LLMs truly learn.