High dimensionality

Your AI Evals Are Lying

Your AI Evals Are Lying

Andrew Burt of Luminos discusses how current AI risk evaluation methods are insufficient, advocating for a "high dimensionality" approach using granular sub-risks and diverse legal and technical expertise. He critiques common practices like guardrails and single-LLM evaluations, highlighting the need for multimodal systems and continuous, automated monitoring to address the evolving complexities of AI, particularly with the rise of open-weight models.