Transformers

Elon's Former Battery Chief on Making Transformers 100x Smaller | Drew Baglino, Heron Power

Elon's Former Battery Chief on Making Transformers 100x Smaller | Drew Baglino, Heron Power

Drew Baglino, former Tesla Powertrain & Energy head and now CEO of Heron Power, reveals why the current electricity grid is inadequate for the explosive growth of AI data centers. He explains how Heron Power's wideband gap power semiconductors will revolutionize grid-to-chip infrastructure, cutting power losses by half, shrinking massive transformers by 100x, and transforming data centers into grid-positive assets for a more efficient and sustainable energy future.

Waymo Co-CEO Dmitri Dolgov: The Demo Is Only 1% Of The Work

Waymo Co-CEO Dmitri Dolgov: The Demo Is Only 1% Of The Work

Waymo co-CEO Dmitri Dolgov outlines seven crucial lessons from fifteen years of developing and scaling the Waymo Driver, the world's most advanced physical AI. He details the unique challenges of physical AI compared to digital, emphasizing the critical role of reliability, strategic technology choices, continuous innovation through foundation models, structure-augmented learning, high-fidelity simulation, AI flywheels, and robust evaluation frameworks to achieve superhuman safety and build trust in real-world autonomous systems.

Building the Automated AGI Lab: Core Automation's Jerry Tworek and Rohan Anil

Building the Automated AGI Lab: Core Automation's Jerry Tworek and Rohan Anil

Jerry Tworek and Rohan Anil, founders of Core Automation, argue that the transformer architecture has reached its limits and the primary bottleneck to smarter AI systems is now architectural. They contend that current models lack continual learning and test-time adaptation capabilities, critical for real-world deployment. They advocate for new architectures that learn from experience more efficiently than current reinforcement learning, optimize pre-training and RL end-to-end, and build a highly automated lab focused on accelerating innovation through kernel generation, aiming for models that can improve themselves without human intervention.

Training an LLM from Scratch, Locally — Angelos Perivolaropoulos, ElevenLabs

Training an LLM from Scratch, Locally — Angelos Perivolaropoulos, ElevenLabs

A practical guide to the engineering principles and trade-offs involved in training a small language model from scratch on a local machine, based on a workshop by Angelos Perivolaropoulos from ElevenLabs.

Beyond Bigger Models: Recursion As The Next Scaling Law In AI

Beyond Bigger Models: Recursion As The Next Scaling Law In AI

Recent advancements with Hierarchical Reasoning Models (HRM) and Tiny Recursive Models (TRM) show how recursion at inference time enables small, 7-million parameter models to outperform models 1000x their size on complex reasoning tasks. This is achieved by giving models compute depth to break through the inherent reasoning ceilings of standard feed-forward Transformers.

Will machines ever be intelligent?

Will machines ever be intelligent?

Doug Burger, Nicolò Fusi, and Subutai Ahmad explore the intelligence of AI, contrasting transformer-based LLMs with the human brain's distributed, continuously learning architecture. They delve into differences in efficiency, representation, and sensory-motor grounding, debating what intelligence truly means and how future AI might bridge the gap.