Data curation

Going In Deep On Data | YC Paper Club

Going In Deep On Data | YC Paper Club

Three experts discuss the evolving landscape of data in AI, covering its critical role in model performance, challenges in sourcing and evaluating expert data, the development of advanced data generation techniques for diffusion models, and the complexities of multilingual pre-training.

Training Krea 2: What matters in generative model training — Sangwu Lee, Krea.ai

Training Krea 2: What matters in generative model training — Sangwu Lee, Krea.ai

The talk by Sangwha Lee discusses Krea 2's open-source medium variant, emphasizing stylistic diversity and rapid iteration over the consistency-focused approach of larger models. A significant portion details their robust data curation pipeline, including unique methods for deduplication, filtering out AI-generated images, utilizing sparse autoencoders for unsupervised tagging, and ensuring world knowledge coverage. He outlines an LLM-inspired multi-stage training process, culminating in a prompt expander, and shares insights on fast iteration and future directions for image generation, highlighting the increasing integration of VLM advancements.

Data Quality Is the Compute Multiplier — Ari Morcos, DatologyAI

Data Quality Is the Compute Multiplier — Ari Morcos, DatologyAI

Ari Morcos, CEO of DatologyAI, explains why data quality is the critical "compute multiplier" in an era of scarce and expensive compute. He outlines DatologyAI's "oil refinery" process (Clean, Curate, Create, Compose) for enhancing datasets. Through empirical results and customer cases like Thomson Reuters and Arcee, he demonstrates how superior data curation leads to significantly better models, reduced training costs, improved inference efficiency, and the ability to train competitive models for a fraction of traditional costs, proving that manufacturing high-quality data is more effective than buying more compute.

Data and Environment Curation for Post-Training LLMs — Mahesh Sathiamoorthy, Bespoke Labs

Data and Environment Curation for Post-Training LLMs — Mahesh Sathiamoorthy, Bespoke Labs

Mahesh Sathiamoorthy of Bespoke Labs argues that high-quality data and curated RL environments are the true bottlenecks for post-training LLMs, especially for building reliable, autonomous agents. He grounds this in experiences with OpenThoughts, a reasoning dataset, highlighting counterintuitive lessons like the importance of diverse reasoning traces and the fact that stronger teachers aren't always best. A key takeaway, reinforced by their Curator tooling, is that a disciplined curation stack is essential for transforming base models into capable, post-trained agents for real-world applications like credit card compliance.

Why Physical AI Is the Next Platform Shift

Why Physical AI Is the Next Platform Shift

Encord Co-CEO Eric Landau reflects on his transition from a lucrative quant career to founding an AI startup, driven by a deep belief in AI's paradigm-shifting potential. He discusses Encord's slow, compounding path to product-market fit, the pivotal role of Physical AI, and the importance of embracing the emotional rollercoaster of startup life.

Your Moat Is Your Data Model — Mike Phipps, Gates Foundation

Your Moat Is Your Data Model — Mike Phipps, Gates Foundation

Mike Phipps from the Gates Foundation details how they built a Strategic Intelligence Platform (SIP) using a Neo4j knowledge graph to serve AI agents. He argues that the true "moat" in an AI-commoditized world lies in an organization's unique data model and tacit knowledge, not in generic AI tools. The platform unifies 25 years of siloed grantmaking data, integrating structured and unstructured information through a rigorous curation pipeline, and is refined via continuous retrieval evaluations to ensure alignment with organizational reporting standards.