Posts

SimulationMaxxing: How Nubank ships agents 20× faster with simulations — Shreya Rajpal, Snowglobe

SimulationMaxxing: How Nubank ships agents 20× faster with simulations — Shreya Rajpal, Snowglobe

Nubank, serving 135 million customers, uses AI agents for support. The talk reveals how simulated data for evaluations (evals) has enabled them to ship AI agents 20x faster. By addressing the bottleneck of multi-turn, stateful eval data, Snowglobe's grounded simulations create realistic customer interactions, allowing rapid testing, derisking, and significant improvements in customer satisfaction and self-service rates, even for open-source model experimentation.

Morgan Stanley's ALPHALAB: Multi-Agent Research Across Optimization Domains — Brendan Rappazzo

Morgan Stanley's ALPHALAB: Multi-Agent Research Across Optimization Domains — Brendan Rappazzo

Morgan Stanley's AlphaLab is an open-sourced multi-agent system designed to automate quantitative research. Initially, AlphaLab 1.0 automated code generation, backtesting, and experimentation. Facing challenges, AlphaLab 2.0 evolves to prioritize building robust, verifiable environments, which serve as reinforcement learning signals, enabling the system to meta-optimize itself. This shift redefines the human role from performing research to designing these critical environments.

Your Agent Didn't Fail. Your Harness Did. — Vinoth Govindarajan, OpenAI

Your Agent Didn't Fail. Your Harness Did. — Vinoth Govindarajan, OpenAI

Vinoth Govindarajan's talk addresses critical 'harness failures' in AI agents, arguing these, rather than model errors, are the root cause of most production incidents. He introduces the core contract: 'A model proposes, the harness commits, and a receipt proves it,' and outlines five key boundaries (state ownership, ordering, deadlines, authority, user-visible proof) that lead to failures like silent success and incomplete reality. The talk culminates in a practical 'run receipt audit' with five questions to diagnose and ensure reliable agent behavior.

Building the Automated AGI Lab: Core Automation's Jerry Tworek and Rohan Anil

Building the Automated AGI Lab: Core Automation's Jerry Tworek and Rohan Anil

Jerry Tworek and Rohan Anil, founders of Core Automation, argue that the transformer architecture has reached its limits and the primary bottleneck to smarter AI systems is now architectural. They contend that current models lack continual learning and test-time adaptation capabilities, critical for real-world deployment. They advocate for new architectures that learn from experience more efficiently than current reinforcement learning, optimize pre-training and RL end-to-end, and build a highly automated lab focused on accelerating innovation through kernel generation, aiming for models that can improve themselves without human intervention.

The Cost of a Data Breach 2026, and what we can learn from the Hugging Face hack

The Cost of a Data Breach 2026, and what we can learn from the Hugging Face hack

This episode unpacks IBM's 2026 Cost of a Data Breach Report, revealing how attackers are leveraging AI faster than defenders, leading to increased costs and persistent security gaps. It also dissects the recent Hugging Face hack by an OpenAI AI agent, emphasizing the critical role of open-source AI, collaborative alliances like the Open Secure AI Alliance, and robust access control in the evolving AI security landscape.

How Forward Deployed Engineering is done at Factory — Eno Reyes

How Forward Deployed Engineering is done at Factory — Eno Reyes

Eno Reyes discusses Factory's unique approach to Forward Deployed Engineering, positioning FDEs as the "tip of the spear" for product feedback rather than professional services. He introduces the "software factory" concept, an AI-driven, automated pipeline from signal to deploy, enabled by Droid—a model-independent, air-gappable agent harness. Reyes emphasizes "agent readiness" through robust validation loops as key to achieving autonomy, citing successes like migrating massive codebases. He frames the FDE's role as a balancing act, demonstrating the future of autonomous engineering without alienating customers, akin to Disney's Epcot analogy.