Ai infrastructure

How KV Cache Speeds Up LLMs for Faster AI Models on GPUs

How KV Cache Speeds Up LLMs for Faster AI Models on GPUs

LLMs often slow down under heavy traffic due to inefficient GPU memory management during inference. This overview explains how KV cache and Paged Attention, implemented in VLLM, optimize memory usage across prefill and decode phases, significantly boosting LLM throughput, reducing latency, and improving GPU utilization through advanced context handling and specific tuning techniques like prefix caching and speculative decoding.

Welcome Session - Microsoft Research India Academic Summit 2026

Welcome Session - Microsoft Research India Academic Summit 2026

The Microsoft Research India Academic Summit 2026 opens with MSR India Lab Director Venkat Padmanabhan outlining the lab's collaborative research philosophy and Microsoft's evolution into an AI infrastructure powerhouse. He details MSR India's four core research pillars: fundamental AI advancements, specialized domain solutions, efficiency across the AI stack (including small language models), and the crucial diffusion of AI technologies for societal impact in India and the Global South, exemplified by diverse projects and collaborations.

The AI Frontier: from FLOPs to Megawatts — Anjney Midha, AMP

The AI Frontier: from FLOPs to Megawatts — Anjney Midha, AMP

Anjney Midha unpacks the critical bottlenecks in AI scaling beyond just GPU acquisition, advocating for responsible infrastructure, community-aligned data centers, and an independent system operator model for compute. He discusses the perils of research hoarding, the rise of researcher CEOs, and how Anthropic's culture of "preparedness" and "output maxing" led to its success, while also highlighting his personal mission to use AI for precise end-of-life prediction.

How Cursor Trained Composer on Fireworks: Distributed Infrastructure for High-Performance RL

How Cursor Trained Composer on Fireworks: Distributed Infrastructure for High-Performance RL

Cursor's Federico Cassano and Fireworks' Dmytro Dzhulgakov detail their collaboration on Composer 2, a specialized foundation model for software engineering. They discuss their top-down training strategy, the infrastructure challenges of large-scale distributed Reinforcement Learning on sparse models, and how model specialization achieves frontier performance with superior efficiency.

Scaling the Next Paradigm of Heterogeneous Intelligence — Adrian Bertagnoli, Callosum

Scaling the Next Paradigm of Heterogeneous Intelligence — Adrian Bertagnoli, Callosum

Adrian Bertagnoli from Callosum argues that the era of scaling monolithic models on homogeneous GPU clusters is ending. He introduces "heterogeneous intelligence," a new paradigm where model architectures, chip types, and workflows are optimized together. By routing subtasks to the most efficient model and hardware, this approach achieves significant performance gains, as demonstrated by two key results: a 7x cost reduction in recursive reasoning tasks using Cerebras, and state-of-the-art performance on the Video Web Arena benchmark, outperforming leading GPT and Gemini models at a fraction of the cost and time.

Tokenmaxxing vs AI Hardware Bottlenecks — with Jon Krohn (@JonKrohnLearns)

Tokenmaxxing vs AI Hardware Bottlenecks — with Jon Krohn (@JonKrohnLearns)

While the 'tokenmaxxing' trend grows, the AI industry faces severe physical infrastructure bottlenecks. This summary explores the four key constraints choking AI compute: GPU packaging (CoWoS), high-bandwidth memory (HBM), the surprising surge in CPU demand from agentic AI, and critical electricity shortages, revealing how these challenges are shaping the future of AI development.