Open source models

How Harvey Built a Research Lab on a Budget | Gabe Pereyra

How Harvey Built a Research Lab on a Budget | Gabe Pereyra

Gabe Pereyra of Harvey details a playbook for application companies to compete with frontier AI labs by leveraging the ecosystem. Key strategies include building specialized benchmarks like Legal Agent Bench, using domain experts for synthetic data generation to overcome sensitive client data issues, partnering with multiple 'neo labs' for post-training, and developing robust model serving and evaluation infrastructure. He emphasizes open-sourcing data for validation and the 'Moneyball' philosophy for success, addressing challenges like talent acquisition and long-context management in the Q&A.

Anthropic’s sandbox breach, EU’s AI transparency push and DeepSeek’s cost-cutting model

Anthropic’s sandbox breach, EU’s AI transparency push and DeepSeek’s cost-cutting model

This episode delves into several critical developments in AI. It begins by discussing recent sandbox breaches by Anthropic and Meta, mirroring earlier incidents with OpenAI, prompting debate on whether these are mere accidents or a growing concern as models become more capable and "agentic." The conversation then shifts to the EU's new AI transparency rules, exploring the challenges and effectiveness of labeling AI-generated content. Finally, the podcast examines DeepSeek V4-Flash's impact on the AI market, questioning if its low cost and high performance will disrupt the pricing of more capable, proprietary models and drive greater commodification and on-device inference.

The State of Model Routing — NVIDIA, Cognition, OpenRouter

The State of Model Routing — NVIDIA, Cognition, OpenRouter

A deep dive into model routing strategies for AI/ML production, featuring experts from Cognition, OpenRouter, and NVIDIA. Key topics include optimizing costs with multi-model systems, delegating tasks between frontier and smaller models, managing context efficiently (sidekicks, compaction), and adapting to dynamic task complexities. The panel discusses the fragility of naive routing, the cost implications of in-distribution vs. out-of-distribution tasks, and the evolution of auto-routers driven by real-world usage patterns like OpenClaw's heartbeats. Insights also cover NVIDIA's Flex Run for dynamic model sizing, hallucination probes for detecting model limitations, and the future of hybrid local/cloud routing and model collaboration.

Decagon’s Playbook for Building Enterprise AI Applications

Decagon’s Playbook for Building Enterprise AI Applications

Jesse Zhang and Ashwin Sreenivas, co-founders of Decagon, discuss their company's transition to open-source models for enterprise AI, emphasizing how fine-tuned small models outperform frontier models on specific tasks. They delve into the role of application-layer companies in an AI-first world, their product-driven 'glass box' approach for enterprises, and the transformative power of their 'Duet Autopilot' agent, which builds other AI agents. The conversation also covers AI's impact on jobs, highlighting the Jevons Paradox in customer support.

AI Security Costs Rise: Cost of a Data Breach Report & Claude Opus 5

AI Security Costs Rise: Cost of a Data Breach Report & Claude Opus 5

The podcast explores key AI developments, beginning with IBM's 2026 Cost of a Data Breach Report, highlighting AI's increasing role in both cyberattacks and defense, and the economic asymmetry it creates. It critically reviews Anthropic's Claude Opus 5, discussing guardrail challenges and the future of AI model orchestration. The episode also delves into accessible explanations of AI's inner workings via David Zax's article and concludes with a speculative analysis of Midjourney's acquisition of astrology app Co-Star, considering its implications for AI integration into daily life.

Sovereign Escape Velocity: Ownership w Open Models — Gus Martins, & Ian Ballantyne, Google DeepMind

Sovereign Escape Velocity: Ownership w Open Models — Gus Martins, & Ian Ballantyne, Google DeepMind

Google DeepMind's Ian Ballantyne and Gus Martins introduce Gemma 4, a family of open models delivering state-of-the-art performance with remarkable size efficiency. They discuss how models like the 31B variant outperform competitors 2-20x its size while running on a single GPU, the shift to an Apache 2.0 license to foster sovereignty and adoption, and the new economics of running powerful agentic workloads on hardware ranging from a Pixel phone to a single enterprise GPU.