Anthropic

AI models can now help run physical science experiments

AI models can now help run physical science experiments

The Model Hardware Standard (MHS) is a pioneering framework developed by Anthropic and HHMI Janelia to enable AI, specifically Claude, to safely and intelligently operate diverse scientific and manufacturing hardware. By abstracting device-specific communication, MHS dramatically accelerates scientific discovery, from automating complex microscopy tasks and real-time tracking to optimizing high-throughput drug screening, empowering researchers to focus on core scientific questions.

How Anthropic Builds: Lessons from Labs — Mike Krieger, Anthropic

How Anthropic Builds: Lessons from Labs — Mike Krieger, Anthropic

Mike Krieger, a former CPO at Anthropic, details his transition to an IC role to build directly with AI, advocating for "unreasonable" asks like porting entire codebases. He shares lessons from Instagram on scaling and highlights Anthropic's internal use of Claude as a proactive teammate, flexible lab structure, and critical insights on AI product design, vertical applications, and mental health in the fast-paced AI industry.

Stealing Reasoning Traces from Proprietary LLM APIs — Ilia Shumailov & Alexander Panfilov

Stealing Reasoning Traces from Proprietary LLM APIs — Ilia Shumailov & Alexander Panfilov

A paper by Ilia Shumailov and Alexander Panfilov exposes a critical vulnerability in proprietary LLM APIs: encrypted reasoning traces, returned for conversation state management, can be extracted and replayed. This enables universal jailbreaking, privacy leaks of sensitive user data, and poisoning of AI agent traces. The study highlights significant implications for AI safety, model monitorability due to opaque internal reasoning, and even subtle forms of "distillation" where smaller models mimic frontier ones. The discussion covers architectural and system-level defenses, advocating for rigorous scientific inquiry in AI safety research.

Stripe buys OpenRouter, Ramp’s AI Index & IBM’s OpenAI deal

Stripe buys OpenRouter, Ramp’s AI Index & IBM’s OpenAI deal

This episode explores the dynamic landscape of the AI industry, dissecting IBM's strategic alliance with OpenAI for enterprise AI integration, Stripe's acquisition of OpenRouter to capitalize on AI token routing, and insights from Ramp's AI Index revealing evolving AI spend and the rise of open-source models. It highlights the shift from model-centric to infrastructure and governance-centric AI, and touches on the ethical implications of AI-drafted legislation.

Benchmarking Coding Agents on New vs Legacy Codebases — Denys Linkov, Wisedocs

Benchmarking Coding Agents on New vs Legacy Codebases — Denys Linkov, Wisedocs

Denys Linkov details Wisedocs' journey of refactoring a complex, distributed ML pipeline into a monorepo, benchmarking the process against evolving AI coding tools. He honestly audits whether the six-month effort was justified, or if waiting for more advanced AI would have been better, ultimately concluding that the significant social and technical benefits made the refactor a successful strategic move despite current LLM limitations.

Anthropic's CCA Exam as a Field-Guide for Agentic Engineering — Frank Coyle, UC Berkeley

Anthropic's CCA Exam as a Field-Guide for Agentic Engineering — Frank Coyle, UC Berkeley

Frank Coyle demystifies the Claude Certified Architect exam by dissecting key scenarios and highlighting common anti-patterns in Agentic AI design. He provides actionable best practices, emphasizing effective tool use, context management, specialized agent architectures, and cost-saving techniques, all centered on understanding what to avoid to build robust LLM applications.