Posts

AI models can now help run physical science experiments

AI models can now help run physical science experiments

The Model Hardware Standard (MHS) is a pioneering framework developed by Anthropic and HHMI Janelia to enable AI, specifically Claude, to safely and intelligently operate diverse scientific and manufacturing hardware. By abstracting device-specific communication, MHS dramatically accelerates scientific discovery, from automating complex microscopy tasks and real-time tracking to optimizing high-throughput drug screening, empowering researchers to focus on core scientific questions.

Can LLMs Write Fast Multi-GPU Kernels? — Simran Arora, Together AI

Can LLMs Write Fast Multi-GPU Kernels? — Simran Arora, Together AI

Simran Arora discusses the critical bottleneck shift in large AI workloads from GPU compute to inter-GPU communication. Her team's solution, ParallelKittens, offers a set of primitives to optimize multi-GPU kernels by leveraging fundamental transfer mechanisms and compute-communication overlapping. They introduce ParallelKernelBench, a benchmark to evaluate AI models' ability to generate such kernels, revealing that while models can handle syntax, they struggle with deeper reasoning about communication patterns and hardware trade-offs.

How Outset Turned AI Interviews Into a New Category

How Outset Turned AI Interviews Into a New Category

Outset's CEO Aaron Cannon discusses pioneering AI-moderated customer research, detailing the challenges of building a new market category, the transformative impact of advancing AI models on product capabilities and customer insights, and their new Simulations Lab featuring "digital twins" for predicting human behavior and unblocking enterprise creativity.

How Anthropic Builds: Lessons from Labs — Mike Krieger, Anthropic

How Anthropic Builds: Lessons from Labs — Mike Krieger, Anthropic

Mike Krieger, a former CPO at Anthropic, details his transition to an IC role to build directly with AI, advocating for "unreasonable" asks like porting entire codebases. He shares lessons from Instagram on scaling and highlights Anthropic's internal use of Claude as a proactive teammate, flexible lab structure, and critical insights on AI product design, vertical applications, and mental health in the fast-paced AI industry.

Inside Cursor: The Anatomy of a Generational Startup

Inside Cursor: The Anatomy of a Generational Startup

A deep dive into Cursor's journey, highlighting their contrarian bets on the human-AI interface over foundation models, their decision to fork VS Code, and their unwavering resilience against tech giants. The discussion covers their rapid product evolution, innovative hiring strategies, and a unique culture of craftsmanship and disciplined focus that enabled them to thrive in the competitive AI coding market.

LLM & AI Agent Benchmarks vs Reality: Why AI Applications Break

LLM & AI Agent Benchmarks vs Reality: Why AI Applications Break

LLM leaderboard scores often don't reflect real-world performance. This video explains why and outlines a comprehensive approach to evaluate AI systems, focusing on the critical balance of accuracy, latency, and cost. It details model and system evaluation techniques, including handling different inference phases, workload shapes, and specific considerations for AI agents, emphasizing the need for realistic testing over generic benchmarks.