Llm reasoning

Training Frontier Models to Out-Think Hackers — Uri Rolls, Arithmetic & Thom Wolf, Hugging Face

Training Frontier Models to Out-Think Hackers — Uri Rolls, Arithmetic & Thom Wolf, Hugging Face

Thom Wolf and Uri Rolls discuss the critical role of AI in cybersecurity, presenting a new benchmark called Masov. They argue that while frontier models excel at reconnaissance, they lack the sophisticated reasoning to exploit complex, logic-based zero-day vulnerabilities, such as a Keycloak name-versus-ID exploit. The solution, they propose, lies in high-quality, open-source AI models trained on real-world zero-day data to enable defenders to outpace attackers and build a new, AI-native security stack.

Gen AI pilots fail, GPT-5's hidden prompt revealed, reasoning model flaws and Claude closing chats

Gen AI pilots fail, GPT-5's hidden prompt revealed, reasoning model flaws and Claude closing chats

A deep dive into why most enterprise GenAI pilots are failing, the debate around hidden system prompts in models like GPT-5, new research questioning the reliability of "chain of thought" reasoning, and the controversy over Anthropic's "AI welfare" justification for shutting down conversations.