Posts

Gpt-oss, Genie 3, Personal Superintelligence and Claude pricing

Gpt-oss, Genie 3, Personal Superintelligence and Claude pricing

The panel discusses OpenAI's strategic release of open-weight models (`gpt-oss`), the implications of Google DeepMind's immersive 3D world generator (`Genie 3`), the economic realities behind Anthropic's `Claude Code` rate-limiting, and the competing visions of "Personal Superintelligence" from major players like Meta, OpenAI, and Anthropic.

How to look at your data — Jeff Huber (Choma) + Jason Liu (567)

How to look at your data — Jeff Huber (Choma) + Jason Liu (567)

A detailed summary of a talk by Jeff Huber (Chroma) and Jason Liu on systematically improving AI applications. The talk covers using fast, inexpensive evaluations for retrieval systems (inputs) and applying structured data analysis and clustering to conversational logs (outputs) to derive actionable product insights.

On Engineering AI Systems that Endure The Bitter Lesson - Omar Khattab, DSPy & Databricks

On Engineering AI Systems that Endure The Bitter Lesson - Omar Khattab, DSPy & Databricks

Omar Khattab, creator of DSPy, reinterprets the 'Bitter Lesson' for AI engineering, arguing that the key to building robust and enduring AI systems is to move beyond brittle prompt engineering. He advocates for a declarative, modular approach that separates the fundamental program logic from the rapidly changing landscape of LLMs, optimizers, and inference techniques.

Evals Are Not Unit Tests — Ido Pesok, Vercel v0

Evals Are Not Unit Tests — Ido Pesok, Vercel v0

Ido Pesok from Vercel explains why LLM-based applications often fail in production despite successful demos, and presents a systematic framework for building reliable AI systems using application-layer evaluations ("evals").

2025 is the Year of Evals! Just like 2024, and 2023, and … — John Dickerson, CEO Mozilla AI

2025 is the Year of Evals! Just like 2024, and 2023, and … — John Dickerson, CEO Mozilla AI

A deep dive into why 2025 is poised to be the 'Year of Evals' for AI. The speaker argues that a confluence of factors—the C-suite's post-ChatGPT awakening, budget dynamics, and the rise of autonomous agentic systems—has finally made AI evaluation a critical, top-of-mind issue for enterprise leaders.

Vibe Coding with Confidence — Itamar Friedman, Qodo

Vibe Coding with Confidence — Itamar Friedman, Qodo

Itamar Friedman of Qodo argues that the future of AI in software development lies in moving beyond simple code generation to 'vibe coding with confidence.' This is achieved through multi-agent workflows, grounded in team standards and orchestrated via the Command Line Interface (CLI), enabling a holistic AI-driven approach across the entire SDLC.