Data modeling

The 12 KB File That Replaces Weeks of Training (with Tristan Handy)

The 12 KB File That Replaces Weeks of Training (with Tristan Handy)

Tristan Handy, founder and CEO of dbt Labs, details the evolution of analytics engineering from a 2016 study into a tool used by over 100,000 data teams. He explains his decision to use SQL over Spark for accessibility, the concept of "progressive complexity," and how dbt projects transform raw data into modeled tables using a Directed Acyclic Graph. Handy elaborates on the critical role of the semantic layer in ensuring consistent metric definitions for both human users and AI agents, especially in large organizations. He introduces the dbt Fusion Engine, aiming to bring type safety and universal SQL understanding, and discusses how 12-kilobyte skill files can revolutionize large-scale dbt migrations, reducing them from years to weeks by enabling AI agents to absorb vast amounts of expert knowledge.

Your Moat Is Your Data Model — Mike Phipps, Gates Foundation

Your Moat Is Your Data Model — Mike Phipps, Gates Foundation

Mike Phipps from the Gates Foundation details how they built a Strategic Intelligence Platform (SIP) using a Neo4j knowledge graph to serve AI agents. He argues that the true "moat" in an AI-commoditized world lies in an organization's unique data model and tacit knowledge, not in generic AI tools. The platform unifies 25 years of siloed grantmaking data, integrating structured and unstructured information through a rigorous curation pipeline, and is refined via continuous retrieval evaluations to ensure alignment with organizational reporting standards.

The Semantic Layer and AI Agents // David Jayatillake // MLOps Podcast #343

The Semantic Layer and AI Agents // David Jayatillake // MLOps Podcast #343

David Jayatillake, VP of AI at Cube.dev, discusses the critical role of a headless, open-source semantic layer in the modern data stack. He argues against proprietary, BI-tool-specific semantic layers that create vendor lock-in and advocates for a decoupled approach. The conversation explores how AI agents can automate the entire data pipeline—from ingestion and transformation to generating and querying the semantic layer—and compares the functionalities of semantic layers and feature stores, highlighting the crucial difference of temporality.