Synthetic data

Going In Deep On Data | YC Paper Club

Going In Deep On Data | YC Paper Club

Three experts discuss the evolving landscape of data in AI, covering its critical role in model performance, challenges in sourcing and evaluating expert data, the development of advanced data generation techniques for diffusion models, and the complexities of multilingual pre-training.

Rich Sutton and Khurram Javed: Why AI Models Stop Learning, and How to Start It Again

Rich Sutton and Khurram Javed: Why AI Models Stop Learning, and How to Start It Again

Rich Sutton and Khurram Javed discuss their radical vision for AI at Oak Lab, advocating for truly continual learning agents that learn from their own experience, rejecting synthetic data due to the "Big World Hypothesis," and outlining a path to overcome catastrophic forgetting with "continual backprop" for a trillion-parameter, self-maintaining mind, while critiquing LLMs as only a fraction of intelligence.

AI & Data Science Periodic Tables: How They Work Together

AI & Data Science Periodic Tables: How They Work Together

Aaron Baughman and Martin Keen present a unified framework using "periodic tables" to integrate AI and Data Science. They illustrate how elements like pipelines, embeddings, and RAG combine to build real-world AI applications, using a detailed document Q&A system example. The discussion emphasizes the critical interdependence of data science in grounding AI models and ensuring continuous improvement through an innovative feedback loop.

RL Environments Explained: How AI Agents Learn Real-World Work | Brendan Foody, Mercor

RL Environments Explained: How AI Agents Learn Real-World Work | Brendan Foody, Mercor

Mercor CEO Brendan Foody elucidates the concept of RL environments, essential for training advanced AI agents. He breaks down their three core components—worlds, apps, and tasks—and details Mercor's evolution from crowdsourced data to expert-driven, "agentic" data. Foody underscores the indispensable role of human experts in defining frontier tasks and creating robust verifiers, exemplified by a real legal RL environment. He shares post-training results demonstrating significant performance gains with modest compute, discusses data pricing and quality, demystifies synthetic data, and explores future directions like ultra-long-horizon tasks and virtual co-workers. The talk emphasizes that data sets are becoming a critical moat for application-layer companies, enabling them to own their intelligence.

How Harvey Built a Research Lab on a Budget | Gabe Pereyra

How Harvey Built a Research Lab on a Budget | Gabe Pereyra

Gabe Pereyra of Harvey details a playbook for application companies to compete with frontier AI labs by leveraging the ecosystem. Key strategies include building specialized benchmarks like Legal Agent Bench, using domain experts for synthetic data generation to overcome sensitive client data issues, partnering with multiple 'neo labs' for post-training, and developing robust model serving and evaluation infrastructure. He emphasizes open-sourcing data for validation and the 'Moneyball' philosophy for success, addressing challenges like talent acquisition and long-context management in the Q&A.

Data Quality Is the Compute Multiplier — Ari Morcos, DatologyAI

Data Quality Is the Compute Multiplier — Ari Morcos, DatologyAI

Ari Morcos, CEO of DatologyAI, explains why data quality is the critical "compute multiplier" in an era of scarce and expensive compute. He outlines DatologyAI's "oil refinery" process (Clean, Curate, Create, Compose) for enhancing datasets. Through empirical results and customer cases like Thomson Reuters and Arcee, he demonstrates how superior data curation leads to significantly better models, reduced training costs, improved inference efficiency, and the ability to train competitive models for a fraction of traditional costs, proving that manufacturing high-quality data is more effective than buying more compute.