Posts

The State of Model Routing — NVIDIA, Cognition, OpenRouter

The State of Model Routing — NVIDIA, Cognition, OpenRouter

A deep dive into model routing strategies for AI/ML production, featuring experts from Cognition, OpenRouter, and NVIDIA. Key topics include optimizing costs with multi-model systems, delegating tasks between frontier and smaller models, managing context efficiently (sidekicks, compaction), and adapting to dynamic task complexities. The panel discusses the fragility of naive routing, the cost implications of in-distribution vs. out-of-distribution tasks, and the evolution of auto-routers driven by real-world usage patterns like OpenClaw's heartbeats. Insights also cover NVIDIA's Flex Run for dynamic model sizing, hallucination probes for detecting model limitations, and the future of hybrid local/cloud routing and model collaboration.

How Open Source Became AI's Backbone | Inferact with a16z

How Open Source Became AI's Backbone | Inferact with a16z

Simon Mo, CEO of Inferact and lead maintainer of vLLM, discusses how open-source AI, exemplified by vLLM, transformed into critical infrastructure. The conversation highlights the technical complexities of serving LLMs, the evolving economics and licensing of open-weight models, the need for control over guardrails, and the rapidly disappearing capability gap between open and proprietary AI.

A Long Spring: 19 Years of Living with Your Past Mistakes • Arjen Poutsma • GOTO 2025

A Long Spring: 19 Years of Living with Your Past Mistakes • Arjen Poutsma • GOTO 2025

Arjen Poutsma, a Spring Framework veteran, reflects on 19 years of evolving the framework. He shares insights on API design from XML to functional programming, detailing anecdotes like bubble sort in MVC and the challenges of HTTP method enums. Poutsma emphasizes the critical role of diversity, empathy, humility, and restraint in open-source stewardship to maintain developer trust and foster innovation under constraints.

Chasing Trillion-Dollar Companies, Founder Ambition, Token Budgets, & Regulatory Capture

Chasing Trillion-Dollar Companies, Founder Ambition, Token Budgets, & Regulatory Capture

Elad and Sarah discuss the rapid rise of AI's multi-trillion-dollar companies, debating whether this growth is sustainable or an anomaly. They explore how founder ambition is shaped by fear of AI labs, optimal strategies for startup exits, and the psychological impact of impending AGI on researchers. Key bottlenecks like compute power, the emergence of an oligopoly, and the growing threat of regulatory capture—exemplified by California's tax policies—are also analyzed, highlighting the critical societal trade-off between safety and technological progress.

Gadgets: Personal app vibe coding that is actually safe — Kenton Varda, Cloudflare

Gadgets: Personal app vibe coding that is actually safe — Kenton Varda, Cloudflare

Kenton Varda introduces "Gadgets," a new architecture built on Cloudflare Workers, addressing the challenge of user-driven AI customization in a world where traditional cloud infrastructure restricts personalization. He demonstrates how AI agents can add features directly to applications, while a unique security model, leveraging `null origin iframes` and isolated server-side sandboxes, renders XSS vulnerabilities irrelevant. The entire system operates without containers or traditional databases, running efficiently on the `workerd` runtime.

Building the First Data Centers in Space

Building the First Data Centers in Space

Philip Johnston, co-founder and CEO of StarCloud, discusses their pioneering efforts to build data centers in space to address AI's energy demands and terrestrial constraints. He details StarCloud-1's groundbreaking mission, which launched an Nvidia H100 GPU into orbit, and the engineering challenges overcome, such as thermal management and radiation hardening. The conversation covers their ambitious plan for a 88,000-satellite constellation, the shift in investor sentiment towards hard tech, and key advice for deep tech founders.