Token efficiency

CrabRAG: Why Automated Assistants Need Graph Memory, Not More Tokens — Stephen Chin, Neo4j

CrabRAG: Why Automated Assistants Need Graph Memory, Not More Tokens — Stephen Chin, Neo4j

Stephen Chin demonstrates how current LLM agent memory systems, relying on markdown files or vector databases, lead to token inefficiency, hallucinations, and a lack of multi-hop reasoning. He showcases graph databases as a superior alternative, providing precise, explainable, and auditable answers for complex, large-scale problems through a live demo of a home lab digital twin.

The AI Memory Problem: Why Long Context Isn’t Enough — Dan Biderman, Engram Co-founder & CEO

The AI Memory Problem: Why Long Context Isn’t Enough — Dan Biderman, Engram Co-founder & CEO

Engram co-founder Dan Biderman discusses building AI that learns from users, critiquing long context, RAG, and compaction's limitations. He introduces Engram's approach of compressing knowledge into "cartridges" and model weights via continual learning and gradient-based updates, aiming for "intuition" over retrieval. The vision extends to personal, "Tamagotchi" AI models and addresses the critical need for token efficiency and "doing more with less" in both enterprise and personal AI.

The Future Is Domain-Specific Agents - Justin Schroeder, StandardAgents

The Future Is Domain-Specific Agents - Justin Schroeder, StandardAgents

Justin Schroeder argues for a paradigm shift in AI agent development from monolithic, context-inflated agents (inheritance) to modular, domain-specific agents (DSAs) that operate through composition. He explains how DSAs offer superior token efficiency, cost savings with smaller models, enhanced security through capability limits, and better scalability, predicting their widespread adoption by 2027 as a solution to rising AI costs and the need for practical, customer-facing AI.

Memory and Continual Learning: Engram's Dan Biderman and Jessy Lin

Memory and Continual Learning: Engram's Dan Biderman and Jessy Lin

Dan Biderman and Jessy Lin of Engram introduce their "always training" paradigm, focusing on baking a team's knowledge directly into a model's weights to achieve true memory and continual learning. This contrarian approach, which they call a "RAG killer" for specific use cases, promises up to 100x token savings and superior performance by internalizing context rather than relying on ever-larger context windows or external retrieval, envisioning a future where everyone has their own continually learning, personalized AI model.