Evaluation metrics

Rethinking Environments for Long-Horizon Work — Rayan Garg, Theta Software

Rethinking Environments for Long-Horizon Work — Rayan Garg, Theta Software

Rayan Garg from Theta Software delves into the complexities of defining and evaluating "long horizon" tasks for AI agents. He critiques current metrics and benchmarks, emphasizing the critical role of sophisticated environment design and robust verifiers (judge models) in driving true progress, particularly in "software-failing domains." The discussion highlights issues like task ambiguity, state changes, and the necessity for granular reward signals for effective model training.

Session on Inclusive AI: Data, Models, Evaluation

Session on Inclusive AI: Data, Models, Evaluation

The Microsoft Research India Academic Research Summit 2026 session on "Inclusive AI" explored critical challenges in developing AI that serves diverse linguistic and cultural contexts. Speakers Niloy Ganguly, Danish Pruthi, Sunayana Sitaram, Anoop Kunchukuttan, and Ashutosh Modi addressed data gaps, model biases, and evaluation shortcomings, emphasizing the need for equitable and culturally relevant AI. Key themes included the use of synthetic data for low-resource languages, the impact of tokenization on model performance, geographical disparities in generative AI, and the application of AI for social good in legal and accessibility domains. The discussions underscored the importance of community involvement, open data, and designing AI for multilinguality from the outset, rather than as an afterthought.