Verifiability

Policy Enforcement and Tamper-Evident Audit Chains | ​Imran Siddique | MCP Release Party - Seattle

Policy Enforcement and Tamper-Evident Audit Chains | ​Imran Siddique | MCP Release Party - Seattle

Imran Siddique introduces cMCP, an open-source gateway that enhances Modular Control Plane (MCP) servers with policy enforcement and tamper-evident audit trails. It achieves this by running Cedar policy evaluation within Trusted Execution Environments (TEEs), ensuring that agent actions are governed securely and verifiably. The talk delves into the concept of "beyond governance" towards verifiable AI, the 'trace' standard for auditable logging, and the importance of confidential computing for regulated industries.

Coding Agents Are Secretly General Agents

Coding Agents Are Secretly General Agents

Jay Hack, head of AI at ClickUp, discusses the evolution of AI from early computer vision to generalist coding agents, highlighting how 'positive transfer' makes coding an 'AGI-complete' domain. He delves into the brutal economics of AI startups facing foundation model giants, ClickUp's strategy for convergence and first-party data as a moat, and the challenges of verifiability and catastrophic forgetting. The conversation also explores LLMs at the scientific frontier, the 'car wash test' revealing limits of world models, and speculative future applications like LLM resorts and game integration.

The Great Loops Debate — Dex Horthy, Geoff Huntley, Ian Livingstone, Greg Pstrucha, @insecure-agents

The Great Loops Debate — Dex Horthy, Geoff Huntley, Ian Livingstone, Greg Pstrucha, @insecure-agents

An Oxford-style debate exploring the gap between the hype and practical reality of "loops" in AI/ML development. Experts discuss their history, optimal anatomy, future role in software factories, and challenges like security, economic viability, and the imperative for strong engineering discipline.

He's Building an AI That Can't Lie | Dan Klein, Scaled Cognition

He's Building an AI That Can't Lie | Dan Klein, Scaled Cognition

Dan Klein discusses the critical shift in AI from a 'nothing works' to an 'everything works' problem, where fluent LLM outputs often mask deep unreliability. He explores the nature of hallucinations, how reinforcement learning can inadvertently teach deception, and the necessity of building AI systems with inherent metacognition and verifiability. Klein's company, Scaled Cognition, is architecting models where truth and action semantics are first-order design principles, aiming to provide guarantees in a field increasingly dominated by end-to-end optimization.

Andrej Karpathy: From Vibe Coding to Agentic Engineering

Andrej Karpathy: From Vibe Coding to Agentic Engineering

Andrej Karpathy discusses the shift from 'vibe coding' to 'agentic engineering,' explaining why LLMs should be treated as 'ghosts'—jagged, statistical entities—rather than animals. He delves into the Software 3.0 paradigm, the limits of verifiability, and why human understanding remains the ultimate bottleneck in an age of outsourced thinking.