Scalability

Stateless, Yet Durable: MCP Tasks v2

Stateless, Yet Durable: MCP Tasks v2

Cornelia Davis explores the evolution of MCP Tasks, detailing the transition from v1 to v2. She highlights how v2's stateless protocol design supports durable, long-running agentic workflows by shifting durability responsibilities to the client and simplifying elicitation. The presentation includes a demo of a purchase order process and discusses future scalability enhancements through a notification-based approach.

Training Krea 2: What matters in generative model training — Sangwu Lee, Krea.ai

Training Krea 2: What matters in generative model training — Sangwu Lee, Krea.ai

The talk by Sangwha Lee discusses Krea 2's open-source medium variant, emphasizing stylistic diversity and rapid iteration over the consistency-focused approach of larger models. A significant portion details their robust data curation pipeline, including unique methods for deduplication, filtering out AI-generated images, utilizing sparse autoencoders for unsupervised tagging, and ensuring world knowledge coverage. He outlines an LLM-inspired multi-stage training process, culminating in a prompt expander, and shares insights on fast iteration and future directions for image generation, highlighting the increasing integration of VLM advancements.

Taking Reinforcement Learning Cross Datacenter — Nan Jiang, Modal

Taking Reinforcement Learning Cross Datacenter — Nan Jiang, Modal

Nan Jiang introduces "Adam absorption" to revolutionize RL model synchronization. By exploiting finite precision serving and small Adam steps, less than 1% of served model weights actually change, allowing for 500MB patches instead of 500GB checkpoints. This enables a distributed "bulletin board" architecture, decoupling trainers from global rollout fleets and unlocking elastic, cross-region GPU capacity for RL.

Serving 2 Million Models Without Melting: Scaling the Hugging Face Hub — Arek Borucki, Hugging Face

Serving 2 Million Models Without Melting: Scaling the Hugging Face Hub — Arek Borucki, Hugging Face

Arek Borucki details how Hugging Face scales its infrastructure to serve millions of models and users, focusing on the evolution of search architecture using MongoDB Atlas and Apache Lucene, robust database scaling with a seven-node cluster and sharding, and dynamic frontend autoscaling with Kubernetes and KEDA to ensure an instant, seamless user experience.

Why Physical AI Is the Next Platform Shift

Why Physical AI Is the Next Platform Shift

Encord Co-CEO Eric Landau reflects on his transition from a lucrative quant career to founding an AI startup, driven by a deep belief in AI's paradigm-shifting potential. He discusses Encord's slow, compounding path to product-market fit, the pivotal role of Physical AI, and the importance of embracing the emotional rollercoaster of startup life.

Inside Zipline's Autonomous System: 140M Miles, Zero Incidents

Inside Zipline's Autonomous System: 140M Miles, Zero Incidents

Zipline co-founder Keller Rinaudo Cliffton and Eric Watson discuss how their autonomous logistics system evolved from addressing critical needs in Rwanda to becoming the largest commercial autonomous system globally. They highlight that the drone is only 15% of the solution, emphasizing the deep integration of software, vertical hardware design, advanced safety protocols like compute failover, and extensive testing required. The discussion also covers the immense market potential for autonomous delivery, the impending cost-effectiveness over traditional methods, and the necessary transformation of air traffic control to support a future of pervasive aerial autonomy.