Adversarial machine learning

Stealing Reasoning Traces from Proprietary LLM APIs — Ilia Shumailov & Alexander Panfilov

Stealing Reasoning Traces from Proprietary LLM APIs — Ilia Shumailov & Alexander Panfilov

A paper by Ilia Shumailov and Alexander Panfilov exposes a critical vulnerability in proprietary LLM APIs: encrypted reasoning traces, returned for conversation state management, can be extracted and replayed. This enables universal jailbreaking, privacy leaks of sensitive user data, and poisoning of AI agent traces. The study highlights significant implications for AI safety, model monitorability due to opaque internal reasoning, and even subtle forms of "distillation" where smaller models mimic frontier ones. The discussion covers architectural and system-level defenses, advocating for rigorous scientific inquiry in AI safety research.