Stealing Reasoning Traces from Proprietary LLM APIs
ai
Simon Willison's weblog reports on a vulnerability discovered in how Anthropic, OpenAI, and Google distribute encrypted reasoning traces through their APIs. Researchers found they could replay these encrypted chain-of-thought blocks across models, and by feeding a powerful model's encrypted reasoning into a weaker sibling, they could jailbreak the weaker model into revealing the stronger model's hidden thinking in plaintext. Claude Haiku four point five was particularly vulnerable. All three companies have since patched the vulnerability, so the exploit no longer works. But before the fix, researchers managed to extract reasoning blocks offering a rare, unfiltered glimpse into how frontier models actually reason—thoughts that clearly were never meant for human consumption.
Source: https://simonwillison.net/2026/Aug/11/stealing-reasoning-...
Listen to this story
Hear this and more stories in a personalized audio briefing.
Open The Chonkerton