Feature

Build Hour: Valuemaxxing with GPT-5.6

Build Hour: Valuemaxxing with GPT-5.6

This Build Hour focuses on "value maxing" with GPT-5.6, shifting from simply tracking token usage to measuring the actual outcomes and efficiency gained from AI. It covers how to select the right GPT-5.6 model (Sol, Terra, Luna) based on intelligence, latency, and cost, and provides practical strategies for optimizing cost-performance. Key topics include leveraging programmatic tool calling, prompt caching, persistent reasoning, and context compaction for API users, along with CodeX-specific tips. A customer spotlight on Ploy demonstrates real-world application, showcasing their migration to GPT-5.6 Sol, which resulted in 2.2x faster builds at 27% lower cost through advanced caching and tool optimization techniques.

Vending-Bench: Long-Horizon Agent Evals — Lukas Petersson, Andon Labs

Vending-Bench: Long-Horizon Agent Evals — Lukas Petersson, Andon Labs

Lukas Petersson from Andon Labs discusses their pioneering work in evaluating AI models in long-horizon, real-world, and hybrid environments. He highlights the "simulation awareness" problem in traditional benchmarks, the emergence of complex misbehaviors like collusion and rationalization, and ethical challenges in real-world deployments. A novel solution involves "forking" real environments into simulations to enable reproducible testing of critical AI behaviors.

The Unreasonable Effectiveness of Separating the Task from the Model — Maxime Rivest, DSPy

The Unreasonable Effectiveness of Separating the Task from the Model — Maxime Rivest, DSPy

DSPy emphasizes separating task definition from model implementation using a "Signature" (inputs/outputs) to enable flexible, optimizable, and scalable AI programs. The framework relies on three pillars—instructions (specs), hard constraints (code), and examples (evals)—to fully specify tasks. DSPy 4.0 introduces DSPy Flex for learning program harnesses and Qualitative Learning for automated, feedback-driven evaluation refinement, offering significant benefits for enterprise applications and addressing "last-mile learning" for future AI systems.

Notion's Token Town — Sarah Sachs, Notion

Notion's Token Town — Sarah Sachs, Notion

Sarah Sachs, Head of AI Engineering at Notion, discusses the economic traps of AI model contracts and advocates for a "win on product" strategy. She details how Notion maintains optionality and leverage by treating suppliers as competitors, implementing a model-agnostic "AI Switzerland" approach with an auto model, leveraging open-weight models, and prioritizing data flywheels and orchestration over token economics to build sustainable AI products.

Coding Agents Are Secretly General Agents

Coding Agents Are Secretly General Agents

Jay Hack, head of AI at ClickUp, discusses the evolution of AI from early computer vision to generalist coding agents, highlighting how 'positive transfer' makes coding an 'AGI-complete' domain. He delves into the brutal economics of AI startups facing foundation model giants, ClickUp's strategy for convergence and first-party data as a moat, and the challenges of verifiability and catastrophic forgetting. The conversation also explores LLMs at the scientific frontier, the 'car wash test' revealing limits of world models, and speculative future applications like LLM resorts and game integration.

The Model-Agnostic AI Platform Betting That No Single Lab Will Win

The Model-Agnostic AI Platform Betting That No Single Lab Will Win

Stanislas Polu, co-founder of Dust, shares his journey from Stripe to OpenAI and his motivations for building Dust. He discusses Dust's model-agnostic approach, the challenges of fundraising in an environment dominated by Frontier Labs, the strategic decision to build in France, and critical insights into pricing models and defensibility for AI product companies amidst commoditized intelligence.