Full text search

Serving 2 Million Models Without Melting: Scaling the Hugging Face Hub — Arek Borucki, Hugging Face

Serving 2 Million Models Without Melting: Scaling the Hugging Face Hub — Arek Borucki, Hugging Face

Arek Borucki details how Hugging Face scales its infrastructure to serve millions of models and users, focusing on the evolution of search architecture using MongoDB Atlas and Apache Lucene, robust database scaling with a seven-node cluster and sharding, and dynamic frontend autoscaling with Kubernetes and KEDA to ensure an instant, seamless user experience.

The Future of Search: Agents, RAG, and Why Retrieval Still Matters — Simon Eskildsen, Turbopuffer

The Future of Search: Agents, RAG, and Why Retrieval Still Matters — Simon Eskildsen, Turbopuffer

Simon Hørup Eskildsen, founder of turbopuffer, shares his journey from scaling Shopify's infrastructure to creating a new search engine for the AI era. He discusses how a prohibitively expensive experiment at Readwise inspired him to build a cost-effective vector search solution based on object storage and NVMe. Eskildsen breaks down turbopuffer's architecture, its role in cutting costs for companies like Cursor and Notion, his philosophy on building a 'P99' engineering team, and how agentic workloads are changing the future of retrieval.