Llms

The State of Model Routing — NVIDIA, Cognition, OpenRouter

The State of Model Routing — NVIDIA, Cognition, OpenRouter

A deep dive into model routing strategies for AI/ML production, featuring experts from Cognition, OpenRouter, and NVIDIA. Key topics include optimizing costs with multi-model systems, delegating tasks between frontier and smaller models, managing context efficiently (sidekicks, compaction), and adapting to dynamic task complexities. The panel discusses the fragility of naive routing, the cost implications of in-distribution vs. out-of-distribution tasks, and the evolution of auto-routers driven by real-world usage patterns like OpenClaw's heartbeats. Insights also cover NVIDIA's Flex Run for dynamic model sizing, hallucination probes for detecting model limitations, and the future of hybrid local/cloud routing and model collaboration.

How AI Helps Solve Medical Mysteries at Boston Children’s Hospital | OpenAI Forum

How AI Helps Solve Medical Mysteries at Boston Children’s Hospital | OpenAI Forum

Researchers from Boston Children's Hospital and OpenAI collaborated to apply AI (specifically, OpenAI o3 Deep Research) to tackle the "diagnostic odyssey" of rare diseases. By analyzing complex genomic and phenotypic data, the AI model helped identify 18 new diagnoses in 376 previously unsolved pediatric cases, showcasing its ability to accelerate literature review, generate hypotheses, and uncover obscure but critical information. This human-in-the-loop approach aims to make diagnosis faster, more accessible, and more precise, offering hope for personalized medicine and improved patient outcomes.

Teaching AI to Find Real Vulnerabilities — David Brumley, Bugcrowd

Teaching AI to Find Real Vulnerabilities — David Brumley, Bugcrowd

David Brumley discusses the challenges and solutions for teaching AI models to hack, drawing parallels with human learning. He introduces a 'ladder of tasks' approach for reinforcement learning and addresses the critical flaw of traditional benchmarks: measurement difficulties with multiple vulnerabilities and 'reward hacking.' His team's 'Audit Task' uses deterministic graders and precision/recall metrics for open-world assessment. He demonstrates this with an in-depth case study on attacking Chrome's V8 engine, showcasing how advanced models achieve real zero-day exploits, and warns against 'benchmaxxing security' without robust, honest grading.

Data Quality Is the Compute Multiplier — Ari Morcos, DatologyAI

Data Quality Is the Compute Multiplier — Ari Morcos, DatologyAI

Ari Morcos, CEO of DatologyAI, explains why data quality is the critical "compute multiplier" in an era of scarce and expensive compute. He outlines DatologyAI's "oil refinery" process (Clean, Curate, Create, Compose) for enhancing datasets. Through empirical results and customer cases like Thomson Reuters and Arcee, he demonstrates how superior data curation leads to significantly better models, reduced training costs, improved inference efficiency, and the ability to train competitive models for a fraction of traditional costs, proving that manufacturing high-quality data is more effective than buying more compute.

Decagon’s Playbook for Building Enterprise AI Applications

Decagon’s Playbook for Building Enterprise AI Applications

Jesse Zhang and Ashwin Sreenivas, co-founders of Decagon, discuss their company's transition to open-source models for enterprise AI, emphasizing how fine-tuned small models outperform frontier models on specific tasks. They delve into the role of application-layer companies in an AI-first world, their product-driven 'glass box' approach for enterprises, and the transformative power of their 'Duet Autopilot' agent, which builds other AI agents. The conversation also covers AI's impact on jobs, highlighting the Jevons Paradox in customer support.

AI Security Costs Rise: Cost of a Data Breach Report & Claude Opus 5

AI Security Costs Rise: Cost of a Data Breach Report & Claude Opus 5

The podcast explores key AI developments, beginning with IBM's 2026 Cost of a Data Breach Report, highlighting AI's increasing role in both cyberattacks and defense, and the economic asymmetry it creates. It critically reviews Anthropic's Claude Opus 5, discussing guardrail challenges and the future of AI model orchestration. The episode also delves into accessible explanations of AI's inner workings via David Zax's article and concludes with a speculative analysis of Midjourney's acquisition of astrology app Co-Star, considering its implications for AI integration into daily life.