Data and Environment Curation for Post-Training LLMs — Mahesh Sathiamoorthy, Bespoke Labs
Mahesh Sathiamoorthy of Bespoke Labs argues that high-quality data and curated RL environments are the true bottlenecks for post-training LLMs, especially for building reliable, autonomous agents. He grounds this in experiences with OpenThoughts, a reasoning dataset, highlighting counterintuitive lessons like the importance of diverse reasoning traces and the fact that stronger teachers aren't always best. A key takeaway, reinforced by their Curator tooling, is that a disciplined curation stack is essential for transforming base models into capable, post-trained agents for real-world applications like credit card compliance.