Posts

Baseten CEO Tuhin Srivastava on Custom Models, and Building the Inference Cloud

Baseten CEO Tuhin Srivastava on Custom Models, and Building the Inference Cloud

Baseten CEO Tuhin Srivastava discusses the explosive growth in AI inference, driven by the adoption of specialized and post-trained open-source models. He covers the strategic importance of owning the software layer on top of compute, navigating the severe GPU supply crunch with a multi-cloud fabric, the evolving landscape of AI workloads, and the operational lessons learned from scaling 30x in one year.

Beyond Bigger Models: Recursion As The Next Scaling Law In AI

Beyond Bigger Models: Recursion As The Next Scaling Law In AI

Recent advancements with Hierarchical Reasoning Models (HRM) and Tiny Recursive Models (TRM) show how recursion at inference time enables small, 7-million parameter models to outperform models 1000x their size on complex reasoning tasks. This is achieved by giving models compute depth to break through the inherent reasoning ceilings of standard feed-forward Transformers.

Getting Humans Out of the Way: How to Work with Teams of Agents

Getting Humans Out of the Way: How to Work with Teams of Agents

Rob Ennals, creator of Broomy, discusses a paradigm shift in working with AI coding agents: moving away from micromanagement towards orchestrating teams of parallel agents. The key is to design robust, automated validation systems and reshape the development environment to empower agents to work autonomously, efficiently, and at scale.

988: In Case You Missed It in April 2026 — with @JonKrohnLearns

988: In Case You Missed It in April 2026 — with @JonKrohnLearns

This episode explores the foundations of AI agent memory, drawing parallels with human neuroscience. It also covers the practical impact of AI on data engineering roles, the democratization of AI development through low-code tools, and innovative applications of AI in elementary education to foster critical thinking.

Granite 4.1, IBM Bob & building a quantum ecosystem

Granite 4.1, IBM Bob & building a quantum ecosystem

This episode of Mixture of Experts breaks down IBM's enterprise-focused Granite 4.1 and Project Bob, Google DeepMind's DiLoCo distributed training method, the inference-efficient DeepSeek V4 model, and IBM's strategy for achieving quantum advantage through strategic partnerships.

Running AI Engineer with AI — swyx

Running AI Engineer with AI — swyx

Swyx, co-founder of the AI Engineer conference, reveals how his "tiny team" of nine people leverages AI agents like Cognition's Devin and Town Assistant to manage a multi-million dollar business and organize large-scale events. He shares practical examples of how agents have transformed their workflow, from converting Figma designs to pixel-perfect websites to managing complex conference schedules and even performing personal research tasks, arguing for a future focused on "Agent Experience" (AX).