Three experts discuss the evolving landscape of data in AI, covering its critical role in model performance, challenges in sourcing and evaluating expert data, the development of advanced data generation techniques for diffusion models, and the complexities of multilingual pre-training.
Why Most AI Agents Fail Horribly
Maarten Grootendorst discusses the foundational understanding developers need for modern AI tools, emphasizing core LLM concepts like tokens, embeddings, and attention. He provides a pragmatic view on AI agents, distinguishing hype from practical applications like coding assistants, and explores the role of memory, guardrails, and the growing importance of open-weight models for control and efficiency in AI infrastructure.
Why Deep Networks Don’t Need to Memorize Everything — Matthieu Wyart
Matthieu Wyart, a statistical physicist, argues that deep networks discover abstractions by recovering hidden data hierarchies, allowing them to escape the curse of dimensionality. He explains how this mechanism, combined with predicting latent representations instead of raw tokens, can significantly improve sample efficiency. The discussion also covers the physics of rough loss landscapes, machine creativity, diffusion models, and a theoretical framework for neural scaling laws.
Chai Discovery's Bitter Lesson: Drug Design Is Another Scaling Problem
Chai Discovery is revolutionizing drug discovery by treating biology as an engineering problem, leveraging AI—particularly diffusion models and the "bitter lesson" of scaling—to design molecules rather than merely discover them. Their approach has boosted antibody design hit rates from 0.1% to 16%, aiming for a "Molecular CAD" suite that collapses discovery timelines from months to days. They partner with pharma, building infrastructure and creating a data flywheel to develop higher-quality, more targeted medicines for previously undruggable diseases.
Data Quality Is the Compute Multiplier — Ari Morcos, DatologyAI
Ari Morcos, CEO of DatologyAI, explains why data quality is the critical "compute multiplier" in an era of scarce and expensive compute. He outlines DatologyAI's "oil refinery" process (Clean, Curate, Create, Compose) for enhancing datasets. Through empirical results and customer cases like Thomson Reuters and Arcee, he demonstrates how superior data curation leads to significantly better models, reduced training costs, improved inference efficiency, and the ability to train competitive models for a fraction of traditional costs, proving that manufacturing high-quality data is more effective than buying more compute.
The Future of AI – Key Trends Shaping What’s Next • Ekaterina Sirazitdinova • YOW! 2025
Ekaterina Sirazitdinova from NVIDIA provides a high-level overview of the latest trends shaping the future of AI, covering the evolution from early deep learning to the rise of agentic and physical AI, and diving deep into the critical optimization techniques required to deploy these powerful models efficiently.