Large language models

The Unreasonable Effectiveness of Separating the Task from the Model — Maxime Rivest, DSPy

The Unreasonable Effectiveness of Separating the Task from the Model — Maxime Rivest, DSPy

DSPy emphasizes separating task definition from model implementation using a "Signature" (inputs/outputs) to enable flexible, optimizable, and scalable AI programs. The framework relies on three pillars—instructions (specs), hard constraints (code), and examples (evals)—to fully specify tasks. DSPy 4.0 introduces DSPy Flex for learning program harnesses and Qualitative Learning for automated, feedback-driven evaluation refinement, offering significant benefits for enterprise applications and addressing "last-mile learning" for future AI systems.

AI, Corporate Responsibility & Democratic Legitimacy: Extended Q&A • Joanna Bryson • GOTO 2025

AI, Corporate Responsibility & Democratic Legitimacy: Extended Q&A • Joanna Bryson • GOTO 2025

Joanna Bryson challenges popular AI assumptions, positing current generative AI as powerful tools for cultural knowledge compression, not autonomous intelligences. She emphasizes that AI's capabilities are nearing the human knowledge frontier, requiring focus on human coordination and governance. Bryson critically examines AI's impact on mental health and law, advocating for data-driven regulation and comprehensible systems. She calls for engineering activism, asserting human agency over technological determinism and stressing the importance of transparency and critical thinking in shaping AI's future.

Open Models: Kimi K3, Qwen 3.8, Xi's WAIC Speech, Distillation, The Open-Closed Gap, and What's Next

Open Models: Kimi K3, Qwen 3.8, Xi's WAIC Speech, Distillation, The Open-Closed Gap, and What's Next

Nathan Lambert and Florian Brand discuss the accelerating open model landscape, focusing on the surge in Chinese models like Kimi K3 and GLM 5.2. They explore the reasons behind China's progress, the evolving US ecosystem, the cybersecurity implications of open-source bans, and debunk common misconceptions about distillation, particularly challenging Ben Thompson's recent claims. The episode concludes with predictions and a frontier model tier list.

The Real AI Frontier Isn't Smarter Machines (with Catherine Williams)

The Real AI Frontier Isn't Smarter Machines (with Catherine Williams)

Dr. Catherine Williams, a former black-hole physicist and early data science leader, explores the field's evolution from Bayesian models to LLMs. She passionately argues that deep mathematical understanding and the ability to build robust mental models are more crucial than ever, even as AI automates technical tasks. Williams also discusses the impact of embeddings, the changing economics of frontier AI, and her work at the nonprofit Candid, advocating for a human-centric approach to intelligence in the age of AI.

Sandboxing, Agent Harnesses, and Agent Teamwork

Sandboxing, Agent Harnesses, and Agent Teamwork

Shahram Anver, CEO of Cleric, details how AI agents for SRE are evolving beyond fast triage to continuous learning and operational memory. He discusses Cleric's architectural shifts, from complex early designs to simpler, sandboxed query agents, and the unique challenges SRE agents face in diverse production environments. A core focus is on human-agent interaction, redefining roles as managers overseeing agents, and how agents learn from unstructured data like Slack to build robust, actionable knowledge for autonomous, self-healing infrastructure. The discussion also touches on the future of software, differentiating between durable systems and rapidly developed "skills" or "vibe-coded" solutions.

Research to Reality with Google DeepMind — Benoit Schillings, Google DeepMind, VP of Technology

Research to Reality with Google DeepMind — Benoit Schillings, Google DeepMind, VP of Technology

Benoit Schillings, VP of Technology at Google DeepMind, explores the evolution of AI's role in software development, highlighting the transition from human-limited coding to an AI frontier where syntax generation is solved. He delves into the power of self-play for model training, the shifting economics of software engineering, and the imperative for active guardrails. Schillings also discusses the need for inductive architecture, advanced model planning, multimodal reasoning (as seen in Gemini), and the potential for AI to drive scientific breakthroughs in fields like chemistry and biology by uncovering patterns imperceptible to human bias.