Llm training

How Reinforcement Learning can Improve your Agent

How Reinforcement Learning can Improve your Agent

This talk addresses the unreliability of current AI agents, arguing that prompting is insufficient. It posits that Reinforcement Learning (RL) is the most promising solution, delving into the mechanisms of RLHF and RLVR. The core challenge identified is 'reward hacking', and the discussion explores future directions to overcome it, such as RLAIF, data augmentation, and the development of interactive, online models that can learn in real-time.

AI Changed Stack Overflow for the Better

AI Changed Stack Overflow for the Better

Stack Overflow CEO Prashanth Chandrashekar discusses the platform's evolution in the AI era, focusing on licensing its trusted Q&A corpus to major AI labs, expanding beyond Q&A to include discussions and live chat, and the critical role of its enterprise solution in powering internal AI agents. A key insight from their upcoming developer survey reveals that while AI adoption for coding is rising, developer trust in AI-generated output is declining, reinforcing Stack Overflow's position as a vital source of human-curated, reliable knowledge.