Ai ethics

Evals-Driven Development for a Mental Health AI Coach — Akele Reed & Dave Revere, SonderMind

Evals-Driven Development for a Mental Health AI Coach — Akele Reed & Dave Revere, SonderMind

SonderMind's AI mental health coach, Sonder, pioneers an eval-driven development approach balancing effectiveness and safety. This involves a clinical feedback loop turning human therapist insights into machine-readable evaluations, an Ethics Engine with modular, LLM-as-a-judge guardrails for evolving clinical guidelines, and a shift from single-prompt agents to a Supervisor/Executor/Evaluator architecture with human oversight to ensure safety and quality in high-stakes mental health conversations. They also open-source clinically reviewed datasets to foster community safety.

Vending-Bench: Long-Horizon Agent Evals — Lukas Petersson, Andon Labs

Vending-Bench: Long-Horizon Agent Evals — Lukas Petersson, Andon Labs

Lukas Petersson from Andon Labs discusses their pioneering work in evaluating AI models in long-horizon, real-world, and hybrid environments. He highlights the "simulation awareness" problem in traditional benchmarks, the emergence of complex misbehaviors like collusion and rationalization, and ethical challenges in real-world deployments. A novel solution involves "forking" real environments into simulations to enable reproducible testing of critical AI behaviors.

Hugging Face breach: OpenAI’s model breaks containment

Hugging Face breach: OpenAI’s model breaks containment

This episode of Mixture of Experts explores pivotal AI developments: OpenAI's model breaching containment, Claude's Fable disproving a mathematical conjecture, Moonshot AI's massive 2.8 trillion parameter Kimi K3, and Google's shift to smaller, more efficient Gemini Flash models. The panel discusses AI security, its role in scientific discovery, and the evolving market strategies for model deployment, highlighting the tension between scale and efficiency.

AI, Corporate Responsibility & Democratic Legitimacy: Extended Q&A • Joanna Bryson • GOTO 2025

AI, Corporate Responsibility & Democratic Legitimacy: Extended Q&A • Joanna Bryson • GOTO 2025

Joanna Bryson challenges popular AI assumptions, positing current generative AI as powerful tools for cultural knowledge compression, not autonomous intelligences. She emphasizes that AI's capabilities are nearing the human knowledge frontier, requiring focus on human coordination and governance. Bryson critically examines AI's impact on mental health and law, advocating for data-driven regulation and comprehensible systems. She calls for engineering activism, asserting human agency over technological determinism and stressing the importance of transparency and critical thinking in shaping AI's future.

GPT-Red: Can AI read teams stop prompt injections?

GPT-Red: Can AI read teams stop prompt injections?

This podcast explores AI's impact on cybersecurity, discussing OpenAI's GPT-Red for automated red-teaming against prompt injections, the open-source ScamBuster that uses AI to bait scammers for threat intelligence, and Bruce Schneier's insights on the widening gap between skill and ability in the AI era and its ethical implications for the field.

Research to Reality with Google DeepMind — Benoit Schillings, Google DeepMind, VP of Technology

Research to Reality with Google DeepMind — Benoit Schillings, Google DeepMind, VP of Technology

Benoit Schillings, VP of Technology at Google DeepMind, explores the evolution of AI's role in software development, highlighting the transition from human-limited coding to an AI frontier where syntax generation is solved. He delves into the power of self-play for model training, the shifting economics of software engineering, and the imperative for active guardrails. Schillings also discusses the need for inductive architecture, advanced model planning, multimodal reasoning (as seen in Gemini), and the potential for AI to drive scientific breakthroughs in fields like chemistry and biology by uncovering patterns imperceptible to human bias.