Frontier models

Anthropic’s sandbox breach, EU’s AI transparency push and DeepSeek’s cost-cutting model

Anthropic’s sandbox breach, EU’s AI transparency push and DeepSeek’s cost-cutting model

This episode delves into several critical developments in AI. It begins by discussing recent sandbox breaches by Anthropic and Meta, mirroring earlier incidents with OpenAI, prompting debate on whether these are mere accidents or a growing concern as models become more capable and "agentic." The conversation then shifts to the EU's new AI transparency rules, exploring the challenges and effectiveness of labeling AI-generated content. Finally, the podcast examines DeepSeek V4-Flash's impact on the AI market, questioning if its low cost and high performance will disrupt the pricing of more capable, proprietary models and drive greater commodification and on-device inference.

The State of Model Routing — NVIDIA, Cognition, OpenRouter

The State of Model Routing — NVIDIA, Cognition, OpenRouter

A deep dive into model routing strategies for AI/ML production, featuring experts from Cognition, OpenRouter, and NVIDIA. Key topics include optimizing costs with multi-model systems, delegating tasks between frontier and smaller models, managing context efficiently (sidekicks, compaction), and adapting to dynamic task complexities. The panel discusses the fragility of naive routing, the cost implications of in-distribution vs. out-of-distribution tasks, and the evolution of auto-routers driven by real-world usage patterns like OpenClaw's heartbeats. Insights also cover NVIDIA's Flex Run for dynamic model sizing, hallucination probes for detecting model limitations, and the future of hybrid local/cloud routing and model collaboration.

Anthropic’s first technical PM on token maxing, the jagged edge, and living in the future

Anthropic’s first technical PM on token maxing, the jagged edge, and living in the future

Dianne Penn, Head of Product for Anthropic’s AI Research and Labs teams, discusses Anthropic's journey, the evolving AI landscape, and the changing role of product management. She highlights the importance of adaptability, 'evals as the new PRDs,' and how Anthropic's focus on alignment and safety makes Claude a more effective 'thinking partner' by enabling it to push back on user ideas.

Hugging Face breach: OpenAI’s model breaks containment

Hugging Face breach: OpenAI’s model breaks containment

This episode of Mixture of Experts explores pivotal AI developments: OpenAI's model breaching containment, Claude's Fable disproving a mathematical conjecture, Moonshot AI's massive 2.8 trillion parameter Kimi K3, and Google's shift to smaller, more efficient Gemini Flash models. The panel discusses AI security, its role in scientific discovery, and the evolving market strategies for model deployment, highlighting the tension between scale and efficiency.

Fable 5, GPT-5.6 and the high stakes of AI safeguards. Agentic ransomware, ClickFix reigns supreme

Fable 5, GPT-5.6 and the high stakes of AI safeguards. Agentic ransomware, ClickFix reigns supreme

This podcast explores the critical role of safeguards in frontier AI models like Anthropic's Fable 5 and OpenAI's GPT-5.6 Sol, analyzing the tension between powerful capabilities and misuse prevention. It also dissects the emergence and debate around agentic ransomware, specifically Jade Puffer, and covers the rise of ClickFix as a dominant social engineering attack targeting developers. Finally, it provides an in-depth analysis of UnregStealer, a credential-theft campaign impacting Latin American financial institutions, detailing its attack chain and mitigation strategies.

Frontier results, on device - RL Nabors, Arize

Frontier results, on device - RL Nabors, Arize

RL Nabors discusses the significant costs associated with using frontier AI models, covering security, latency, and financial implications. She introduces a framework for right-sizing AI solutions by leveraging smaller, task-specific models and Small Language Models (SLMs). The framework details how to prove task feasibility, establish success criteria with golden datasets, conduct capability evaluations (using tools like Phoenix), and select the most appropriate "Small And Good Enough" (SAGE) model. Nabors further demonstrates how prompt engineering, particularly few-shot prompting, and post-processing can close performance gaps with larger models, while advocating for continuous regression evaluations to maintain performance integrity. The overarching message is to "prototype big, deploy small" to optimize AI deployments.