Reinforcement learning ( rl)

Ending AI Slop — Thais Castello Branco, Taste Labs

Ending AI Slop — Thais Castello Branco, Taste Labs

Thais Castello Branco of Taste Labs tackles 'AI slop' in subjective domains like design and creative writing. She proposes a framework to make 'taste' measurable by decomposing subjective concepts into verifiable elements, countering the 'collapse to the mean' that stifles creativity. The approach emphasizes high-signal human preference data, expert-driven feedback tied to specific choices, and a 'quality over quantity' mindset to train AI that understands and generates nuanced, multi-preference outputs.

Why a Nation Can't Outsource Its Frontier AI - Alistair Pullen (Cosine AI)

Why a Nation Can't Outsource Its Frontier AI - Alistair Pullen (Cosine AI)

Alistair Pullen, CEO of Cosine, discusses the UK's sovereign AI initiative, born from US export controls. He outlines Cosine's unique economic model, competing with "millions" against "billions" by licensing models instead of hosting inference. Pullen delves into why open models lag frontier systems, emphasizing active parameters and post-training data. He explains Cosine's innovative approach to "slop" through process-based RL and credit attribution, advocating for runtime proof in code review. The conversation covers their hierarchical "Swarm" sub-agent system, the challenges of memory, and advanced synthetic data generation, concluding on the geopolitical impact of export controls as an unexpected catalyst for UK AI.