Computer vision

The Next Medium: Why Real-Time Interactive Video Changes Everything — Ahmed Ahres, Reactor

The Next Medium: Why Real-Time Interactive Video Changes Everything — Ahmed Ahres, Reactor

Ahmed Ahres argues that real-time interaction fundamentally changes the medium, not just its speed, especially in generative AI. He defines "world models" as interactive, effectively infinite, and steerable video, drawing parallels with GPS and camera viewfinders. This paradigm shift unlocks new forms of control, intelligent advertising, programmable worlds for robotics and education, and advanced live avatars. He highlights the critical infrastructure challenges of streaming pixels, managing stateful sessions, and achieving global sub-100ms latency, emphasizing that batch infrastructure is unsuitable for real-time applications. Evaluation of consistency for these models remains an unsolved problem.

Waymo Co-CEO Dmitri Dolgov: The Demo Is Only 1% Of The Work

Waymo Co-CEO Dmitri Dolgov: The Demo Is Only 1% Of The Work

Waymo co-CEO Dmitri Dolgov outlines seven crucial lessons from fifteen years of developing and scaling the Waymo Driver, the world's most advanced physical AI. He details the unique challenges of physical AI compared to digital, emphasizing the critical role of reliability, strategic technology choices, continuous innovation through foundation models, structure-augmented learning, high-fidelity simulation, AI flywheels, and robust evaluation frameworks to achieve superhuman safety and build trust in real-world autonomous systems.

Building Closed-Loop Evals for a Multimodal Agent at Scale — Soumya Gupta & Jai Chopra, Uber

Building Closed-Loop Evals for a Multimodal Agent at Scale — Soumya Gupta & Jai Chopra, Uber

This talk details how Uber Eats designed and implemented a multimodal AI agent to enhance food photography for independent merchants, addressing challenges like maintaining authenticity, merchant brand, and marketplace diversity while operating at scale. It covers the intricate evaluation strategies for routing and image editing agents, including continuous learning loops, managing drift, countering reward hacking, and balancing creative freedom with rigid safety guardrails. The speakers explain how they built a closed feedback loop combining offline human labeling, internal dogfooding, and online production signals to ensure robust and adaptive performance.

Perception Agents — Antje Barth, Amazon AGI Lab

Perception Agents — Antje Barth, Amazon AGI Lab

Antje Barth of Amazon AGI Lab discusses the architectural gap in current AI agents, which excel at individual tasks but fail at complex, end-to-end workflows due to a lack of reliability and contextual understanding. She introduces "Perception Agents"—AI systems that see, reason, and act on computers like humans, using visual and multimodal input to enable reliable collaboration and close the perception-action loop, highlighting new open-source tools for annotation and verification.

Builders Unscripted: Ep. 4 - Pietro Schirano

Builders Unscripted: Ep. 4 - Pietro Schirano

Pietro Schirano, Founder & CEO of MagicPath, discusses his pioneering work with GPT-5.5 and Codex, transforming creative ideas into software and hardware solutions. He details using advanced AI for image-to-sound conversion, multi-agent workflows, resurrecting obsolete tech with new functionalities, and building his company, MagicPath, on the principle of humans directing AI agents. This interview provides deep insights into the creative potential and practical applications of cutting-edge AI for developers and entrepreneurs.

Physical AI Forum | Builders Reveal the New Moat & Playbook | Creator & Founder's Cut | Mar 2026 |4K

Physical AI Forum | Builders Reveal the New Moat & Playbook | Creator & Founder's Cut | Mar 2026 |4K

In a live panel at the Physical AI Builders Forum, founders and operators in computer vision, robotics, and multimodal AI share their 2026 playbooks. The discussion covers the architectural differences between physical and generative AI, the strategic shift from frame AI to scene AI for enterprise value, and the critical skills needed to build and scale a modern AI business.