Rag

LLM & AI Agent Benchmarks vs Reality: Why AI Applications Break

LLM & AI Agent Benchmarks vs Reality: Why AI Applications Break

LLM leaderboard scores often don't reflect real-world performance. This video explains why and outlines a comprehensive approach to evaluate AI systems, focusing on the critical balance of accuracy, latency, and cost. It details model and system evaluation techniques, including handling different inference phases, workload shapes, and specific considerations for AI agents, emphasizing the need for realistic testing over generic benchmarks.

Why Most AI Agents Fail Horribly

Why Most AI Agents Fail Horribly

Maarten Grootendorst discusses the foundational understanding developers need for modern AI tools, emphasizing core LLM concepts like tokens, embeddings, and attention. He provides a pragmatic view on AI agents, distinguishing hype from practical applications like coding assistants, and explores the role of memory, guardrails, and the growing importance of open-weight models for control and efficiency in AI infrastructure.

AI & Data Science Periodic Tables: How They Work Together

AI & Data Science Periodic Tables: How They Work Together

Aaron Baughman and Martin Keen present a unified framework using "periodic tables" to integrate AI and Data Science. They illustrate how elements like pipelines, embeddings, and RAG combine to build real-world AI applications, using a detailed document Q&A system example. The discussion emphasizes the critical interdependence of data science in grounding AI models and ensuring continuous improvement through an innovative feedback loop.

The RAG Mistake Almost Every Team Is Making (with Pete Johnson)

The RAG Mistake Almost Every Team Is Making (with Pete Johnson)

Pete Johnson, Field CTO of AI at MongoDB, discusses effective AI strategies, why most organizations struggle with AI ROI, and how to build reliable AI systems. He covers the importance of choosing the right embedding models for RAG pipelines, introduces Matryoshka embeddings, and explains the evolution of agentic memory to combat token maxing and ensure consistency in production AI.

Understanding AI Agent Hallucination in AI Systems

Understanding AI Agent Hallucination in AI Systems

Learn about AI hallucinations, why they occur in autonomous agents, and how they pose new risks as AI takes action. Discover key mitigation strategies including data grounding, tool-based reasoning, scope control, and human-in-the-loop interventions to ensure reliable AI performance.

Is Fine-Tuning Still Needed? LLMs, RAG, & LoRA

Is Fine-Tuning Still Needed? LLMs, RAG, & LoRA

This summary explores the evolving role of fine-tuning in modern AI workflows, comparing it with advanced techniques like RAG, LoRA, and enhanced generative AI capabilities. It discusses the historical benefits, current limitations due to rapidly advancing frontier models, and outlines a practical decision framework for customizing machine learning models and designing efficient AI systems.