Data quality

RL Environments Explained: How AI Agents Learn Real-World Work | Brendan Foody, Mercor

RL Environments Explained: How AI Agents Learn Real-World Work | Brendan Foody, Mercor

Mercor CEO Brendan Foody elucidates the concept of RL environments, essential for training advanced AI agents. He breaks down their three core components—worlds, apps, and tasks—and details Mercor's evolution from crowdsourced data to expert-driven, "agentic" data. Foody underscores the indispensable role of human experts in defining frontier tasks and creating robust verifiers, exemplified by a real legal RL environment. He shares post-training results demonstrating significant performance gains with modest compute, discusses data pricing and quality, demystifies synthetic data, and explores future directions like ultra-long-horizon tasks and virtual co-workers. The talk emphasizes that data sets are becoming a critical moat for application-layer companies, enabling them to own their intelligence.

Data Quality Is the Compute Multiplier — Ari Morcos, DatologyAI

Data Quality Is the Compute Multiplier — Ari Morcos, DatologyAI

Ari Morcos, CEO of DatologyAI, explains why data quality is the critical "compute multiplier" in an era of scarce and expensive compute. He outlines DatologyAI's "oil refinery" process (Clean, Curate, Create, Compose) for enhancing datasets. Through empirical results and customer cases like Thomson Reuters and Arcee, he demonstrates how superior data curation leads to significantly better models, reduced training costs, improved inference efficiency, and the ability to train competitive models for a fraction of traditional costs, proving that manufacturing high-quality data is more effective than buying more compute.

Ending AI Slop — Thais Castello Branco, Taste Labs

Ending AI Slop — Thais Castello Branco, Taste Labs

Thais Castello Branco of Taste Labs tackles 'AI slop' in subjective domains like design and creative writing. She proposes a framework to make 'taste' measurable by decomposing subjective concepts into verifiable elements, countering the 'collapse to the mean' that stifles creativity. The approach emphasizes high-signal human preference data, expert-driven feedback tied to specific choices, and a 'quality over quantity' mindset to train AI that understands and generates nuanced, multi-preference outputs.

State of Data — Sean Cai, Independent / State of Data

State of Data — Sean Cai, Independent / State of Data

Sean Cai discusses the evolving landscape of AI data markets, highlighting the shift from raw annotation to high-quality, process-based data. He introduces "Verifier's Law" and its three axes to explain application layer maturity, critiques the shortcomings of current benchmarks, and predicts future AI trends by analyzing data market signals. Cai concludes by envisioning data companies transforming into "neo-labs" that provide enterprise-level reinforcement learning as a service and "Antikythera mechanisms" to manage and monetize real-world data assets.

Why Your Agent Disagrees With Itself (And What To Do About It) - Diane Lin, Datadog

Why Your Agent Disagrees With Itself (And What To Do About It) - Diane Lin, Datadog

This session addresses the critical problem of AI agent inconsistency, particularly in LLMs, which leads to "flip-flops" in high-stakes domains like cybersecurity. It argues that this isn't a model failure but a signal of ambiguity in the "gray zone" near decision boundaries. The proposed solution involves using active learning to identify these ambiguous cases, followed by augmenting agents with semantic memory (explicit policies) and episodic memory (past similar cases) to clarify decisions, improve consistency, and adapt to customer-specific preferences without relying solely on expensive fine-tuning.

Enterprise Agents Have a Structure Problem - Ishita Daga, Tesla

Enterprise Agents Have a Structure Problem - Ishita Daga, Tesla

Ishita Daga, a Senior ML Engineer at Tesla, explains why enterprise AI agents fail, highlighting that current fixes like larger models or more RAG are insufficient. She identifies ambiguity, staleness, and user preference as key structural problems. Daga proposes a solution built on semantic retrieval infrastructure: a hierarchical approach to knowledge sources, including a curated semantic layer and metadata graphs. She also details a robust context life cycle with live data sources and continuous feedback loops to combat staleness. The challenge of integrating individual preferences, requiring agents to reason over business concepts rather than raw schemas, is also discussed.