Feature

CLAUDE.md, AGENTS.md, Skills, Hooks and Subagents: A Field Guide to Steering AI Agents (Ep. 1022)

CLAUDE.md, AGENTS.md, Skills, Hooks and Subagents: A Field Guide to Steering AI Agents (Ep. 1022)

Episode 1022 dissects the crucial aspect of effectively steering AI agents by determining the optimal placement of instructions to ensure reliability and cost-efficiency. It explores seven distinct methods for instruction delivery, contrasting instructions as probabilities with hooks as guarantees, and highlights the industry-wide adoption of standards like `agents.md` and the importance of human-crafted guidance for superior agent performance.

Analyzing Group Chat Encryption in Messaging Applications

Analyzing Group Chat Encryption in Messaging Applications

This talk details a formal security analysis of group chat encryption algorithms in popular messaging applications like MLS, Session, and Keybase. It introduces Symmetric Sign Encryption (SSE) to model these protocols, identifying critical vulnerabilities such as insider replay and reordering attacks in MLS and Session due to insufficient context binding. The analysis highlights the complexities of key-dependent messages and key reuse, demonstrating how formal methods can pinpoint subtle design flaws and suggest robust mitigations for real-world secure communication.

Invited Research Talk: Measuring Generalization in EEG Foundation Models

Invited Research Talk: Measuring Generalization in EEG Foundation Models

This talk presents a multi-dimensional evaluation and interpretability framework for EEG foundation models. It reveals that current models often fail to outperform supervised baselines for BCI tasks, lack robustness to sparse channels, and exhibit an aperiodic low-frequency bias, making them better at capturing subject-specific rather than task-specific information. The analysis highlights critical deficiencies and suggests future directions for pre-training objectives and data collection to improve generalization.

TruthTable: A Verifiable Query Engine

TruthTable: A Verifiable Query Engine

TruthTable is a verifiable database engine that produces succinct cryptographic proofs for SQL query execution. It supports a wide range of SQL queries by leveraging query plans, polynomial encoding, and operator-specific PIOPs. It features a query planner with proof-specific optimizations and a novel batch compilation engine (ARCPOP). Benchmarks on TPC-H show average proving times of 55 seconds, verification times of 32 milliseconds, and proof sizes of 24kB, significantly outperforming prior academic and industrial systems in speed and expressiveness.

AI models can now help run physical science experiments

AI models can now help run physical science experiments

The Model Hardware Standard (MHS) is a pioneering framework developed by Anthropic and HHMI Janelia to enable AI, specifically Claude, to safely and intelligently operate diverse scientific and manufacturing hardware. By abstracting device-specific communication, MHS dramatically accelerates scientific discovery, from automating complex microscopy tasks and real-time tracking to optimizing high-throughput drug screening, empowering researchers to focus on core scientific questions.

How Outset Turned AI Interviews Into a New Category

How Outset Turned AI Interviews Into a New Category

Outset's CEO Aaron Cannon discusses pioneering AI-moderated customer research, detailing the challenges of building a new market category, the transformative impact of advancing AI models on product capabilities and customer insights, and their new Simulations Lab featuring "digital twins" for predicting human behavior and unblocking enterprise creativity.