Feature

Agentic Consent Explained: How AI Agents Act Safely and Responsibly

Agentic Consent Explained: How AI Agents Act Safely and Responsibly

Grant Miller from IBM explains Agentic Consent, a dynamic framework for governing AI agents. The model moves beyond static permissions, using identity, context, and just-in-time user prompts to ensure AI agents act with, not instead of, their human counterparts, enabling trust and safety as autonomy scales.

Why TTS Models Now Look Like LLMs — Samuel Humeau, Mistral

Why TTS Models Now Look Like LLMs — Samuel Humeau, Mistral

Samuel Humeau from Mistral explains the dominant architecture for modern text-to-speech (TTS) systems, which mirrors large language models. He details how neural audio codecs solve the information density problem, the autoregressive transformer backbone for generation, and the streaming techniques used to achieve low perceived latency in voice agents. The talk uses Mistral's open-weight TTS model as a practical example.

Voice AI: when is the "Her" moment? — Neil Zeghidour, Gradium AI

Voice AI: when is the "Her" moment? — Neil Zeghidour, Gradium AI

Neil Zeghidour, CEO of Gradium AI, deconstructs the gap between current voice AI and the "Her" ideal. He argues that while cascaded systems are practical, they are architecturally flawed for natural conversation. The future lies in full-duplex, speech-to-speech models that not only solve latency but also integrate deep paralinguistic understanding and overcome significant cost barriers.

From Zapier for Devs to Powering 90% AI Agents

From Zapier for Devs to Powering 90% AI Agents

Co-founders of Trigger.dev discuss their journey through three product versions to find product-market fit, how their async infrastructure positioned them perfectly for the AI agent era, and their vision for the future of computing: programmatic checkpoint and restore.

How Transformers Finally Ate Vision – Isaac Robinson, Roboflow

How Transformers Finally Ate Vision – Isaac Robinson, Roboflow

Isaac Robinson from Roboflow explains why Vision Transformers (ViTs), despite their initial disadvantages in computational complexity and lack of inductive bias, ultimately surpassed Convolutional Neural Networks (CNNs) for computer vision tasks. The talk covers the critical roles of massive, ViT-specific pre-training methods like MAE and DINO, the architectural evolution through models like Swin, ConvNeXt, and Hiera, and optimizations borrowed from the LLM ecosystem. It culminates in a discussion on the practical deployment challenges of large foundation models like SAM and how Neural Architecture Search can bridge the gap.

Inside China's AI Labs with Interconnects & SAIL Media

Inside China's AI Labs with Interconnects & SAIL Media

Firsthand reflections from a visit to China’s most prominent AI labs, exploring the human side of the Chinese AI ecosystem, the technical constraints they face from chip supply, and how their research culture compares to the Bay Area.