Data leakage

Stop Evaluating Models Like It's the 50s - Alejandro Vidal, Mindmakers

Stop Evaluating Models Like It's the 50s - Alejandro Vidal, Mindmakers

This talk introduces the application of psychometrics, particularly Item Response Theory (IRT), to improve LLM evaluation. It highlights how IRT goes beyond simple accuracy to measure model ability, item difficulty, and discrimination, enabling benchmark auditing, adaptive testing, data leakage detection, and the identification of model relationships and distillation, ultimately providing deeper insights into what LLMs truly learn.

Five AI Risks That Can Get You Fired—And How to Avoid Them

Five AI Risks That Can Get You Fired—And How to Avoid Them

Martin Keen explains five real-world AI risks that can lead to job loss: shadow AI, data leakage, hallucinations, prompt injection, and unauthorized AI agents. He emphasizes the critical need for strong AI governance to ensure safe and productive AI adoption in the workplace.