Nvidia

Why Old GPUs Keep Gaining Value

Why Old GPUs Keep Gaining Value

Steve Hou, Head of Research at Silicon Data, unpacks the data behind the AI compute market, revealing persistent tightness in GPU rentals and rising residual values, despite talk of oversupply. He discusses the methodologies behind their GPU price indices and forward curves, the evolving hardware landscape beyond Nvidia (including AMD, Cerebras, and TPUs), and trends in LLM token economics. The conversation also delves into the complex financial aspects of AI data center buildouts, the rise of specialized models, China's emerging AI hardware ecosystem, and the growing importance of power constraints and distributed AI infrastructure.

Compression at the Edge — Chris Alexiuk, NVIDIA

Compression at the Edge — Chris Alexiuk, NVIDIA

This panel discussion explores the critical role of model compression, particularly quantization, in democratizing AI. It delves into how massive models like GLM 5.2 can be shrunk by over 80% without equivalent performance loss, thanks to techniques like mixed-precision quantization and understanding uneven layer importance. The discussion covers NVIDIA's NVFP4 format, challenges posed by new model architectures, the preference for KL divergence over accuracy benchmarks, and the vision of future AI running efficiently on all local devices.

The State of Model Routing — NVIDIA, Cognition, OpenRouter

The State of Model Routing — NVIDIA, Cognition, OpenRouter

A deep dive into model routing strategies for AI/ML production, featuring experts from Cognition, OpenRouter, and NVIDIA. Key topics include optimizing costs with multi-model systems, delegating tasks between frontier and smaller models, managing context efficiently (sidekicks, compaction), and adapting to dynamic task complexities. The panel discusses the fragility of naive routing, the cost implications of in-distribution vs. out-of-distribution tasks, and the evolution of auto-routers driven by real-world usage patterns like OpenClaw's heartbeats. Insights also cover NVIDIA's Flex Run for dynamic model sizing, hallucination probes for detecting model limitations, and the future of hybrid local/cloud routing and model collaboration.

Jensen Huang: The Mindset That Built NVIDIA

Jensen Huang: The Mindset That Built NVIDIA

Jensen Huang, CEO of NVIDIA, shares critical lessons from NVIDIA's journey, emphasizing how early failures and a commitment to learning new technologies, like purchasing textbooks from Fry's to pivot the company, laid the groundwork for their success. He discusses NVIDIA's strategic vision, driven by accelerating algorithm domains and seeing AlexNet as a universal function approximator, which led to a reinvention of the computing stack. Huang also explores the future of AI with agents, the importance of fine-grained control, and the "Linux moment" of open-source AI, while also forecasting the rise of physical AI and job creation. He concludes with profound advice on resilience, systems thinking, and the "how hard can it be?" mindset for aspiring entrepreneurs in this unprecedented era of technological reset.

Podcast Crossover: AIE, AGI, frontier lab strategy with ​ ⁨@matthew_berman⁩  and @swyxtv

Podcast Crossover: AIE, AGI, frontier lab strategy with ​ ⁨@matthew_berman⁩ and @swyxtv

Matt, the organizer of the AI.engineer conference, shares insights into its origin, the challenges of early adoption, and its current value as a neutral ground for AI labs. He delves into AI hardware trends, discussing specialized chips like Etched, and gives a nuanced take on Anthropic's Fable, addressing performance concerns and compute limitations. The conversation then explores OpenAI's rumored equity offer to the US government, discussing implications for regulation and societal involvement. Matt shares his perspective on AI existential risk and alignment, emphasizing the need for pragmatic engineering solutions. Finally, he outlines the limitations of current LLMs, the critical need for data efficiency, and offers strategic advice for "Agent Labs" navigating the "model capability overhang" versus multi-model agnosticism.

Reddit cracks down on AI slop & the future of AI compute

Reddit cracks down on AI slop & the future of AI compute

This episode explores Reddit's aggressive AI spam combat strategy, revealing AI's dual role in fighting malicious AI. It then dissects Anthropic's Economic Index, highlighting how Claude integrates into daily life despite significant user selection bias. The discussion also covers Orin's $33M raise for a GPU compute marketplace, debating the fungibility of compute and the technical hurdles. Finally, Anthropic's chip ambitions are analyzed as an economic strategy to optimize models, reduce NVIDIA dependency, and manage rising token costs, with comparisons to existing hardware ecosystems and NVIDIA's market position.