A new hack has found a way to turn the smaller, less-protected models offered by major AI companies into decoders for encrypted reasoning generated by their more powerful siblings.
- A new hack has found a way to turn the smaller, less-protected models offered by major AI companies into decoders for encrypted reasoning generated by their more powerful siblings.
- They decoded 315,320 blocks and recovered 367 PII artifacts and 182 credentials, according to the paper.
- Because the encrypted blocks could move between models within the same provider ecosystem, a stronger model could generate a protected trace while a weaker model served as the decoder.
Researchers say the technique can recover hidden reasoning traces from models at OpenAI, Anthropic and Google without directly breaking into the stronger models themselves. In a separate analysis of encrypted reasoning blocks publicly exposed by developers, the researchers recovered 367 pieces of personally identifiable information and 182 credentials.
The study was conducted by eight researchers affiliated with organisations including the ELLIS Institute Tübingen, the Max Planck Institute for Intelligent Systems, the University of Tübingen, AI security company AI Sequrity and cybersecurity company Snyk. Alexander Panfilov and David Schmotz are listed as equal contributors, alongside Ilia Shumailov, Luca Beurer-Kellner, Joachim Schaeffer, Ameya Prabhu, Jonas Geiping and Maksym Andriushchenko.
Smaller Models Become The Decoder
The researchers targeted encrypted reasoning blocks used by major AI providers to keep models’ hidden reasoning from users. The blocks are returned to clients and passed back with later requests. The researchers found that they were compatible and interchangeable across models, users and sessions within the same provider’s ecosystem.
That created the opening.
Have a development worth tracking?
Share product launches, funding announcements, partnerships, research findings and market developments with The Grey Terminal's readership.
→ Submit a Press ReleaseThe team inserted a reasoning trace generated by a stronger model into a weaker model from the same provider. The weaker model could then be induced to decode the protected trace and reproduce it in plaintext without directly jailbreaking the stronger system.
The researchers call the technique a ‘scalable decryption jailbreak’.
OpenAI, Anthropic And Google Were Tested
The researchers demonstrated reasoning extraction across proprietary models from Anthropic, OpenAI and Google. The paper says the attack can bypass anti-distillation protections designed to stop adversaries from extracting proprietary reasoning from more capable models.
The paper’s appendix details separate extraction procedures for Claude, GPT and Gemini. For Gemini, the researchers used Gemini Robotics ER-1.6 as a decoder and Gemini 3.5 Flash as an optional reconciler.
The researchers also compared the recovered reasoning with the summaries shown to users. For the Claude and OpenAI APIs examined, they said the recovered hidden traces contained substantially more information than the summaries exposed through the APIs.
Public Logs Exposed PII And Credentials
The researchers separately examined reasoning blocks that developers had made public in repositories. They decoded 315,320 blocks and recovered 367 PII artifacts and 182 credentials, according to the paper.
Those credentials were not taken from OpenAI, Anthropic or Google’s internal systems. They came from session data that had already been publicly exposed by developers.
The researchers’ finding shows that encrypting the reasoning block did not necessarily prevent sensitive information inside it from being recovered.
Hidden Instructions Can Move With The Trace
The paper identifies another attack path: invisible prompt injection. Researchers said attackers could embed malicious instructions inside encrypted reasoning blocks and use them to poison public agentic rollouts.
The risk is higher for systems that allow AI agents to call tools or perform actions outside the model. A monitor looking only at visible prompts and final responses could miss an instruction carried inside an intermediate reasoning block.
The Weakness Sits Between Models
The researchers’ central finding concerns the architecture around the models rather than a single model’s refusal behaviour. Because the encrypted blocks could move between models within the same provider ecosystem, a stronger model could generate a protected trace while a weaker model served as the decoder.
The researchers said they disclosed the issue to the affected providers and proposed cryptographic and system-level changes to protect client-side reasoning.
The paper was submitted to arXiv on Aug. 10 and has not yet undergone peer review.
Activate Terminal Layer
Structural analysis of the systems, pressures, and stakeholders behind this story.





