Continuous learning

Inside 847 Production Clinical AI Notes — Sebastian Fox, Composo

Inside 847 Production Clinical AI Notes — Sebastian Fox, Composo

Sebastian Fox, a medical doctor and AI evaluation expert, dissects the critical problem of subtle yet dangerous errors in AI-generated clinical notes within high-stakes healthcare. He reveals why conventional AI verification methods fail to grasp the nuanced concept of "what matters" and introduces a novel, adaptive evaluation framework that continuously learns from real outputs and expert judgment to build a dynamic, case-specific standard for AI reliability.

Learning on the Job: The Future of Post-Training — Raymond Feng, Applied Compute

Learning on the Job: The Future of Post-Training — Raymond Feng, Applied Compute

Raymond Feng presents Applied Compute's approach to training custom AI models that learn "on the job" using reinforcement learning. He details the evolution from controlled Q&A to synthetic environments, highlighting the core GRPO-style loop. A major focus is tackling the challenges of environment fidelity and "reward hacking" in simulated settings. The discussion then moves to the complexities of training directly within real-world enterprise harnesses, addressing issues like non-replayability and off-policy data. Feng concludes by outlining frontier research in self-distillation, automated data pipelines, and qualitative feedback, envisioning a future where models continuously learn and self-evaluate from every interaction, making "experience the dominant medium of improvement."

Building Closed-Loop Evals for a Multimodal Agent at Scale — Soumya Gupta & Jai Chopra, Uber

Building Closed-Loop Evals for a Multimodal Agent at Scale — Soumya Gupta & Jai Chopra, Uber

This talk details how Uber Eats designed and implemented a multimodal AI agent to enhance food photography for independent merchants, addressing challenges like maintaining authenticity, merchant brand, and marketplace diversity while operating at scale. It covers the intricate evaluation strategies for routing and image editing agents, including continuous learning loops, managing drift, countering reward hacking, and balancing creative freedom with rigid safety guardrails. The speakers explain how they built a closed feedback loop combining offline human labeling, internal dogfooding, and online production signals to ensure robust and adaptive performance.

Sandboxing, Agent Harnesses, and Agent Teamwork

Sandboxing, Agent Harnesses, and Agent Teamwork

Shahram Anver, CEO of Cleric, details how AI agents for SRE are evolving beyond fast triage to continuous learning and operational memory. He discusses Cleric's architectural shifts, from complex early designs to simpler, sandboxed query agents, and the unique challenges SRE agents face in diverse production environments. A core focus is on human-agent interaction, redefining roles as managers overseeing agents, and how agents learn from unstructured data like Slack to build robust, actionable knowledge for autonomous, self-healing infrastructure. The discussion also touches on the future of software, differentiating between durable systems and rapidly developed "skills" or "vibe-coded" solutions.

Why Your Agent Disagrees With Itself (And What To Do About It) - Diane Lin, Datadog

Why Your Agent Disagrees With Itself (And What To Do About It) - Diane Lin, Datadog

This session addresses the critical problem of AI agent inconsistency, particularly in LLMs, which leads to "flip-flops" in high-stakes domains like cybersecurity. It argues that this isn't a model failure but a signal of ambiguity in the "gray zone" near decision boundaries. The proposed solution involves using active learning to identify these ambiguous cases, followed by augmenting agents with semantic memory (explicit policies) and episodic memory (past similar cases) to clarify decisions, improve consistency, and adapt to customer-specific preferences without relying solely on expensive fine-tuning.

AI Infrastructure, Ray, and Why Nonlinear Careers Win — with Linda Haviv

AI Infrastructure, Ray, and Why Nonlinear Careers Win — with Linda Haviv

Linda Haviv discusses the modern AI landscape, emphasizing that non-linear career paths and systems thinking are now more valuable than pure coding skills. She explores how open-source technology, like the Ray framework, is democratizing AI development and closing the gap with proprietary models, and why building a personal brand through content creation is essential for career growth and community building in a rapidly evolving industry.