On device ai

IBM’s cloud collab, Meta’s Muse Glimmer & OpenAI’s upcoming Astra model

IBM’s cloud collab, Meta’s Muse Glimmer & OpenAI’s upcoming Astra model

This episode explores IBM's massive AI infrastructure partnership with Together AI and NVIDIA, Meta's open-source Muse Glimmer model enabling powerful on-device AI, and OpenAI's delayed Astra model due to critical cybersecurity capabilities. Discussions cover the economics of industrial-scale AI, the implications of local vs. cloud AI, and the profound security challenges and opportunities presented by both open and closed frontier models.

Anthropic’s sandbox breach, EU’s AI transparency push and DeepSeek’s cost-cutting model

Anthropic’s sandbox breach, EU’s AI transparency push and DeepSeek’s cost-cutting model

This episode delves into several critical developments in AI. It begins by discussing recent sandbox breaches by Anthropic and Meta, mirroring earlier incidents with OpenAI, prompting debate on whether these are mere accidents or a growing concern as models become more capable and "agentic." The conversation then shifts to the EU's new AI transparency rules, exploring the challenges and effectiveness of labeling AI-generated content. Finally, the podcast examines DeepSeek V4-Flash's impact on the AI market, questioning if its low cost and high performance will disrupt the pricing of more capable, proprietary models and drive greater commodification and on-device inference.

Local Agentic Theory For Mobile Games — Shafik Quoraishee & Joanne Song, The New York Times

Local Agentic Theory For Mobile Games — Shafik Quoraishee & Joanne Song, The New York Times

This presentation explores the innovative concept of running agentic AI entirely on mobile devices to enhance game accessibility and personalization. It delves into the technical challenges of local AI (space, time, energy budgets), contrasts agentic systems with traditional reinforcement learning, and demonstrates practical applications with a Space Invaders agent and a crossword solver. A key takeaway is the transformation of accessibility from fixed toggles to dynamic, real-time adjustments based on player needs, ultimately envisioning a future of billions of small, personalized local AI brains.

Frontier results, on device - RL Nabors, Arize

Frontier results, on device - RL Nabors, Arize

RL Nabors discusses the significant costs associated with using frontier AI models, covering security, latency, and financial implications. She introduces a framework for right-sizing AI solutions by leveraging smaller, task-specific models and Small Language Models (SLMs). The framework details how to prove task feasibility, establish success criteria with golden datasets, conduct capability evaluations (using tools like Phoenix), and select the most appropriate "Small And Good Enough" (SAGE) model. Nabors further demonstrates how prompt engineering, particularly few-shot prompting, and post-processing can close performance gaps with larger models, while advocating for continuous regression evaluations to maintain performance integrity. The overarching message is to "prototype big, deploy small" to optimize AI deployments.

Sovereign Escape Velocity: Ownership w Open Models — Gus Martins, & Ian Ballantyne, Google DeepMind

Sovereign Escape Velocity: Ownership w Open Models — Gus Martins, & Ian Ballantyne, Google DeepMind

Google DeepMind's Ian Ballantyne and Gus Martins introduce Gemma 4, a family of open models delivering state-of-the-art performance with remarkable size efficiency. They discuss how models like the 31B variant outperform competitors 2-20x its size while running on a single GPU, the shift to an Apache 2.0 license to foster sovereignty and adoption, and the new economics of running powerful agentic workloads on hardware ranging from a Pixel phone to a single enterprise GPU.

⚡️ Google's Open AI Strategy — Omar Sanseviero, Google DeepMind

⚡️ Google's Open AI Strategy — Omar Sanseviero, Google DeepMind

An in-depth look at Gemma 4's novel transformer architecture with per-layer embeddings, enabling efficient parameter offloading for on-device inference. The discussion also covers its native multimodality, the state of fine-tuning, text-based diffusion models, and the growing intersection of research and engineering.