Vector search

The RAG Mistake Almost Every Team Is Making (with Pete Johnson)

The RAG Mistake Almost Every Team Is Making (with Pete Johnson)

Pete Johnson, Field CTO of AI at MongoDB, discusses effective AI strategies, why most organizations struggle with AI ROI, and how to build reliable AI systems. He covers the importance of choosing the right embedding models for RAG pipelines, introduces Matryoshka embeddings, and explains the evolution of agentic memory to combat token maxing and ensure consistency in production AI.

Building Turbopuffer: Gergely Orosz (@pragmaticengineer ) × Simon Eskildsen (CEO)

Building Turbopuffer: Gergely Orosz (@pragmaticengineer ) × Simon Eskildsen (CEO)

Simon Eskildsen, CEO of Turbopuffer, discusses his unique path from early computer fascinations to founding an S3-native vector search company. He shares insights from his eight years at Shopify, detailing infrastructure scaling challenges and the development of Toxyproxy for fault injection. A core theme is his "napkin math" philosophy, using fundamental performance metrics to optimize systems and challenge benchmarks. He explains the impetus for Turbopuffer, its initial minimalist architecture, and how it achieved a 95% cost reduction for its first customer, Cursor. Simon also touches on the unexpected scarcity of CPUs due to AI workloads and his unconventional, principled approach to venture capital and building a remote-first company culture.

AI on Your Lakehouse: Context Comes in Shapes, Not Queries — Zach Blumenfeld, Neo4j

AI on Your Lakehouse: Context Comes in Shapes, Not Queries — Zach Blumenfeld, Neo4j

AI agents often struggle with context in lakehouses, leading to confident but incorrect answers. This workshop by Zach Blumenfeld introduces three Neo4j graph shapes built on lakehouse data—Connections (semantic layer), Trees (document outlines), and Communities (themes)—to provide essential context, enabling agents to accurately navigate structured and unstructured data, answer complex estate-level questions, and overcome limitations of traditional Text2SQL and vector search.

Turbocharge Your Agent's Retrieval with TurboQuant - Shashi Jagtap, Superagentic AI

Turbocharge Your Agent's Retrieval with TurboQuant - Shashi Jagtap, Superagentic AI

This talk introduces TurboQuant, a training-free compression method from Google Research that reduces embedding memory footprint by 5x (from 32-bit to 3-4 bits) without losing search quality. It details how TurboQuant works through scalar quantization and a crucial one-bit error correction step, QJL, enabling agents to remember more on existing hardware by optimizing both KV cache and RAG vector stores. A live demo showcases its effectiveness, making it a vendor-neutral solution for efficient AI agent retrieval.

The Future of Search: Agents, RAG, and Why Retrieval Still Matters — Simon Eskildsen, Turbopuffer

The Future of Search: Agents, RAG, and Why Retrieval Still Matters — Simon Eskildsen, Turbopuffer

Simon Hørup Eskildsen, founder of turbopuffer, shares his journey from scaling Shopify's infrastructure to creating a new search engine for the AI era. He discusses how a prohibitively expensive experiment at Readwise inspired him to build a cost-effective vector search solution based on object storage and NVMe. Eskildsen breaks down turbopuffer's architecture, its role in cutting costs for companies like Cursor and Notion, his philosophy on building a 'P99' engineering team, and how agentic workloads are changing the future of retrieval.

Layering every technique in RAG, one query at a time - David Karam, Pi Labs (fmr. Google Search)

Layering every technique in RAG, one query at a time - David Karam, Pi Labs (fmr. Google Search)

David Karam, formerly of Google Search, presents a pragmatic framework for enhancing RAG systems, advocating a "quality engineering" approach. The talk progresses through a ladder of techniques, from in-memory retrieval and BM25 to custom embeddings, re-ranking, and advanced orchestration, emphasizing that the choice of technique should be driven by empirical analysis of system failures ("loss analysis") and balanced by a "complexity-adjusted impact" mindset.