Posts

Has AI Finally Cracked Time Series Forecasting?

Has AI Finally Cracked Time Series Forecasting?

Ameet Talwalkar, CMU professor and Datadog's Chief Scientist, traces the journey of time series foundation models from early skepticism to their current impact. He details Datadog's Toto V1 and V2, highlighting breakthroughs in zero-shot performance, scaling, and the crucial role of data mix. The discussion extends to the vision of 'world models' for observability, integrating diverse data for self-healing software systems, and concludes with insights on open-weights models and AI's transformative effect on computer science and academic research.

Computer-Use 2.0: Agents Just Got Multi-Cursor — Francesco Bonacci, Cua

Computer-Use 2.0: Agents Just Got Multi-Cursor — Francesco Bonacci, Cua

This talk introduces cua driver, an open-source tool enabling AI agents to interact with computer GUIs in the background across macOS, Windows, and Linux by leveraging accessibility APIs. It details CUABench, a robust evaluation framework with over 130 verifiable tasks designed to benchmark and ensure the trustworthiness of computer-using agents, revealing current limitations in complex tasks like circuit design. Finally, it presents cua fleet, an infrastructure solution that optimizes GPU utilization and reduces costs for RL training of these agents by dynamically managing sandbox environments.

Recursive Model Improvement — Lee Robinson, Cursor, SpaceXAI

Recursive Model Improvement — Lee Robinson, Cursor, SpaceXAI

Lee Robinson of Cursor outlines a comprehensive strategy for recursive AI model improvement, centered on a two-loop training framework. He details how Cursor enhances both user-feedback-driven outer loops and high-quality evaluation inner loops, introducing novel methods like textual feedback and addressing reward hacking. The discussion extends to scaling compute infrastructure through partnerships with SpaceX, Colossus, and Terafab, and leveraging agent-based automation to streamline research and foster a future where models continuously train and improve themselves.

Simon Willison in conversation with Cat Wu & Thariq Shihipar, Anthropic

Simon Willison in conversation with Cat Wu & Thariq Shihipar, Anthropic

A Q&A with Anthropic's Cat Wu and Thariq Shihipar on how Claude Code and agentic AI are fundamentally changing software development, from workflow shifts and engineering norms to safety, model trust, and team collaboration. The discussion covers the rapid evolution of coding agents, the rise of proactive agents like Claude Tag, new approaches to code review and system prompt optimization, and Anthropic's robust safety measures including Auto Mode.

Omnigent: Composition, Control, and Collaboration for AI Agents

Omnigent: Composition, Control, and Collaboration for AI Agents

Denny Lee discusses the industry's shift to meta-harnesses like Omnigent, which enables hot-swappable AI models and agents, illustrated by his personal project of using debating agents to plan a matcha farm in Taiwan. He highlights how "tokenomics" is replaying the CapEx-to-OpEx cost shift, emphasizing the need for developer visibility, central governance, and auto-model selection to manage AI spend. The conversation also touches on the importance of databases for agent memory and accountability in AI-assisted workflows.

GLM-5.2: The real security risk? Plus: Vibe hunting, the end of CVSS and updates on Lightwell

GLM-5.2: The real security risk? Plus: Vibe hunting, the end of CVSS and updates on Lightwell

This podcast explores the implications of open-weight AI models like GLM-5.2 for cybersecurity, CISA's new four-variable vulnerability prioritization model, the rise of AI-assisted 'vibe hunting,' and the commercial launch of IBM and Red Hat's Lightwell for securing open-source software. It highlights the tension between AI capabilities for attackers and defenders, the challenges of rapid vulnerability remediation, and the need for new "trust infrastructures" in the AI era.