Synthetic data

Data and Environment Curation for Post-Training LLMs — Mahesh Sathiamoorthy, Bespoke Labs

Data and Environment Curation for Post-Training LLMs — Mahesh Sathiamoorthy, Bespoke Labs

Mahesh Sathiamoorthy of Bespoke Labs argues that high-quality data and curated RL environments are the true bottlenecks for post-training LLMs, especially for building reliable, autonomous agents. He grounds this in experiences with OpenThoughts, a reasoning dataset, highlighting counterintuitive lessons like the importance of diverse reasoning traces and the fact that stronger teachers aren't always best. A key takeaway, reinforced by their Curator tooling, is that a disciplined curation stack is essential for transforming base models into capable, post-trained agents for real-world applications like credit card compliance.

SimulationMaxxing: How Nubank ships agents 20× faster with simulations — Shreya Rajpal, Snowglobe

SimulationMaxxing: How Nubank ships agents 20× faster with simulations — Shreya Rajpal, Snowglobe

Nubank, serving 135 million customers, uses AI agents for support. The talk reveals how simulated data for evaluations (evals) has enabled them to ship AI agents 20x faster. By addressing the bottleneck of multi-turn, stateful eval data, Snowglobe's grounded simulations create realistic customer interactions, allowing rapid testing, derisking, and significant improvements in customer satisfaction and self-service rates, even for open-source model experimentation.

The Next Frontier of AI Is Spatial Intelligence | Fei-Fei Li on a16z

The Next Frontier of AI Is Spatial Intelligence | Fei-Fei Li on a16z

Fei-Fei Li and Yunzhu Li discuss World Labs' acquisition of SceniX, focusing on building "spatial intelligence" and "large world models" to enable robots to understand and interact with the physical world. They elaborate on SceniX's "real-to-sim-to-real" pipeline, emphasizing how simulation, coupled with generative models like Marble, addresses the data bottleneck in robotics by providing consistent, scalable, and efficient training and evaluation environments. The conversation covers the role of counterfactual reasoning, the development of robotics foundation models, and the strategic focus on semi-structured environments for pragmatic, reliable robot deployment.

The Messy Reality of Scale: Synthetic Data and Pre-Training — Marah Abdin & Robert McHardy, poolside

The Messy Reality of Scale: Synthetic Data and Pre-Training — Marah Abdin & Robert McHardy, poolside

Poolside discusses their innovative approaches to synthetic data generation, pre-training validation, and distributed training challenges. They highlight how modular data pipelines, rigorous replica hash checks, and numerical stability fixes enabled them to scale their LLMs, culminating in the 118B parameter Laguna S model designed for agentic coding, which shows strong early results against leading open-weight models.

Why a Nation Can't Outsource Its Frontier AI - Alistair Pullen (Cosine AI)

Why a Nation Can't Outsource Its Frontier AI - Alistair Pullen (Cosine AI)

Alistair Pullen, CEO of Cosine, discusses the UK's sovereign AI initiative, born from US export controls. He outlines Cosine's unique economic model, competing with "millions" against "billions" by licensing models instead of hosting inference. Pullen delves into why open models lag frontier systems, emphasizing active parameters and post-training data. He explains Cosine's innovative approach to "slop" through process-based RL and credit attribution, advocating for runtime proof in code review. The conversation covers their hierarchical "Swarm" sub-agent system, the challenges of memory, and advanced synthetic data generation, concluding on the geopolitical impact of export controls as an unexpected catalyst for UK AI.

Session on Inclusive AI: Data, Models, Evaluation

Session on Inclusive AI: Data, Models, Evaluation

The Microsoft Research India Academic Research Summit 2026 session on "Inclusive AI" explored critical challenges in developing AI that serves diverse linguistic and cultural contexts. Speakers Niloy Ganguly, Danish Pruthi, Sunayana Sitaram, Anoop Kunchukuttan, and Ashutosh Modi addressed data gaps, model biases, and evaluation shortcomings, emphasizing the need for equitable and culturally relevant AI. Key themes included the use of synthetic data for low-resource languages, the impact of tokenization on model performance, geographical disparities in generative AI, and the application of AI for social good in legal and accessibility domains. The discussions underscored the importance of community involvement, open data, and designing AI for multilinguality from the outset, rather than as an afterthought.