Mixture of experts

Preferences Over Benchmarks: Model Routing — Archana Kamath & Tyler Gillam, DigitalOcean

Preferences Over Benchmarks: Model Routing — Archana Kamath & Tyler Gillam, DigitalOcean

This talk argues against the common practice of picking LLMs based solely on leaderboards, emphasizing that there's no single best model, only the right one for a given request. It introduces Digital Ocean's Inference Router, a customizable, open-source solution that intelligently selects models based on user-defined preferences (cost, latency, task, quality) rather than benchmarks, demonstrating significant cost savings and performance improvements in live demos.

Alibaba's Qwen3.8-Max: Open-Weight Model Surpasses Most American Frontier Labs (Ep. 1018)

Alibaba's Qwen3.8-Max: Open-Weight Model Surpasses Most American Frontier Labs (Ep. 1018)

Jon Krohn dissects Alibaba's Qwen 3.8 Max, a 2.4-trillion-parameter Mixture-of-Experts (MoE) model positioned as the largest open-weight release in history if its promised weights ship. The discussion covers its multimodal capabilities, 1M token context window, and performance competitive with Anthropic's Claude Fable 5. Key highlights include its advanced multi-day agentic capabilities and aggressively low pricing ($2 in / $6 out per million tokens), intensifying the AI price war. Krohn also provides critical insights into the safety of using Chinese models, emphasizing data handling practices and the benefits/risks across different deployment scenarios.

Thinking Machines Lab drops Inkling & Meta’s Muse Spark 1.1

Thinking Machines Lab drops Inkling & Meta’s Muse Spark 1.1

This episode covers Thinking Machines' Inkling, an open-source, customizable model prioritizing architecture over benchmarks; Meta's Muse Spark 1.1, positioned for agent orchestration and enterprise use; OpenAI's GPT-5.6 Sol's 8% score on ARC-AGI-3, reigniting AGI debates; and Anthropic's "J-space" paper, exploring internal model reasoning and its implications for AI safety and interpretability.

Sovereign Escape Velocity: Ownership w Open Models — Gus Martins, & Ian Ballantyne, Google DeepMind

Sovereign Escape Velocity: Ownership w Open Models — Gus Martins, & Ian Ballantyne, Google DeepMind

Google DeepMind's Ian Ballantyne and Gus Martins introduce Gemma 4, a family of open models delivering state-of-the-art performance with remarkable size efficiency. They discuss how models like the 31B variant outperform competitors 2-20x its size while running on a single GPU, the shift to an Apache 2.0 license to foster sovereignty and adoption, and the new economics of running powerful agentic workloads on hardware ranging from a Pixel phone to a single enterprise GPU.

How Cursor Trained Composer on Fireworks: Distributed Infrastructure for High-Performance RL

How Cursor Trained Composer on Fireworks: Distributed Infrastructure for High-Performance RL

Cursor's Federico Cassano and Fireworks' Dmytro Dzhulgakov detail their collaboration on Composer 2, a specialized foundation model for software engineering. They discuss their top-down training strategy, the infrastructure challenges of large-scale distributed Reinforcement Learning on sparse models, and how model specialization achieves frontier performance with superior efficiency.

Granite 4.1, IBM Bob & building a quantum ecosystem

Granite 4.1, IBM Bob & building a quantum ecosystem

This episode of Mixture of Experts breaks down IBM's enterprise-focused Granite 4.1 and Project Bob, Google DeepMind's DiLoCo distributed training method, the inference-efficient DeepSeek V4 model, and IBM's strategy for achieving quantum advantage through strategic partnerships.