Stealing AI Reasoning Traces

Interesting research: “Stealing Reasoning Traces from Proprietary LLM APIs“:

Abstract: Leading large language model providers now conceal their models’ step-by-step reasoning, or chain-of-thought, to protect intellectual property and limit information leakage. Rather than storing these traces server-side, providers return them to the client as blocks of encrypted text, which the client passes back with each subsequent request. Building on prior research, we identify an architectural vulnerability: these encrypted blocks are fully compatible and interchangeable across different sessions, users, and models within a provider’s ecosystem. We exploit this compatibility to develop a scalable decryption jailbreak. By injecting an encrypted reasoning trace from a given model into a weaker, and less safeguarded model from the same provider, we force it to decode and output the trace verbatim in plaintext, without ever jailbreaking the more capable model directly. This vulnerability enables four distinct attack vectors. First, it circumvents anti-distillation mechanisms, allowing adversaries to extract a proprietary model’s reasoning, as we demonstrate across Anthropic, OpenAI, and Google. Second, it allows for large-scale private data extraction. Developers frequently share session logs publicly, unaware of contents of the encrypted blocks. By decoding 315,320 reasoning blocks scraped from public repositories, we recovered 367 Personally Identifiable Information (PII) artifacts and 182 credentials. Third, it inadvertently reveals hazardous information hidden within the reasoning process, even in cases where the model’s final, visible output safely rejects a malicious request. Fourth, attackers can leverage this flaw to execute invisible prompt injections, embedding malicious payloads entirely within encrypted blocks to poison public agentic rollouts. Following responsible disclosure, we propose concrete cryptographic and system-level mitigations to secure client-side reasoning.

Posted on September 8, 2026 at 6:20 AM7 Comments

Comments

Michael September 8, 2026 10:48 AM

So the leading chatbots, which are largely built on stolen IP, are worried about people stealing their IP. Fascinating.

Clive Robinson September 8, 2026 5:28 PM

@ ALL,

With regards this from the articles intro,

“Leading large language model providers now conceal their models’ step-by-step reasoning, or chain-of-thought, to protect intellectual property and limit information leakage.”

It sounds bad enough as indicated…

But there is actually a bit more behind it.

Some jurisdictions are in the near future require verifiable watermarking of all LLM output.

There are pros and cons to this which is one reason the AI companies are broadly on board with it as it protects them to a certain extent.

That is they can turn around and say XXX alleged output was not generated by their LLM or at a given date/time or to a given individual.

It basically adds “watermarking”…

Personally I’m not happy with “watermarking” as it turned into an abject fail for “Digital Rights Managment”(DRM) in the turn of the century.

Whilst such DRM systems will fail for detecting “plagiarism” reliably, they won’t fail in anything like that probability for protecting the AI companies…

r September 8, 2026 8:57 PM

the reasoning traces are likely pointless long term, the current LLMs are inarticulate behemouths and training and architecturesbare likely to change in the future

https://www.technologyreview.com/2026/08/24/1141740/kids-machines-language-learning/

i’m not the only one saying this is a valley camparatively to whay’s possible and what’s known, i think one of google’s(?) founders says something similar about the data architecture we’re currently using for various training. i make minimum wage though so i could be wrong, it’s still good to have these defenses going forward but yeah, all that stolen data is nothing compared to being able to formulate gravity from direct object interaction.

it’s the curious aspect of telling countries they’re going to get ‘left behind’ if you don’t buy into vendor lock-in.

Cybershow September 9, 2026 1:06 PM

@r

“i make minimum wage though so i could be wrong”

I assume this is self-effacing modesty but you hit on a “remuneration fallacy” quite important
yet seemingly overlooked. We once judged expert prestige by high position, impressive titles and
above all pay grades. Today these indicate opinions are bought and paid for.

Truth and salary are inversely proportional

Celos September 16, 2026 12:12 PM

As proofs of incompetence go, the big LLM providers are really outdoing themselves. I mean they obviously do not even understand the basics of how to use encryption and this is just one example of gross incompetence.

KC September 16, 2026 1:32 PM

@Celos

I believe Matthew Green was the first (?) to write about this.

https://blog.cryptographyengineering.com/2026/05/29/fooling-around-with-encrypted-reasoning-blobs/

And it was further explored in the above paper; the project also has its own website.

https://stolen-thoughts.com

From what I understand, the issue has since been mitigated.

https://nsfocusglobal.com/ai-security-incident-case-encrypted-reasoning-blocks-of-proprietary-llms-can-be-stolen-via-cross-model-replay/

“All three providers [Anthropic, OpenAI, and Google] have deployed server-side fixes after receiving the responsible disclosure. As of the paper’s publication in August 2026, the researchers confirmed that their original proof-of-concept attacks could no longer be reproduced on current API versions. However, historical working logs that were publicly shared before the fix still face the risk of being decoded.”

Clive Robinson September 16, 2026 3:09 PM

@ r, cybershow, ALL,

With regards,

<

blockquote>”“i make minimum wage though so i could be wrong”<//blockquote>

At least you are allowed to work, thus have some modicum of self respect.

I used to earn a good income (equivalent to a Doctor so better than quite a few academics).

But then they started “scraping me up from the pavement etc and shovelling me into an ambulance about once every four or five weeks.

Under law I had to hand back my driving licence and then due to the UK

1, Fraud Act 1974
2, Health & Safety at work Act 1974

I was marked as uninsurable and got dismissed from my job.

It was explained to me by a “Queen’s Counsel” that being “black marked for insurance” would mean legally I could not “work” not just for an employer but even self employed…

Thus untill I could get the black mark lifted my life in the UK would be on “disability benifits”…

In theory I can “advise” and be in “public spaces” but otherwise nothing doing.

To get the black mark lifted I have to be a minimum of three years without medical issues…

As some have recently noted I’ve not been well and have been in hospital three times in a relatively short period of time as the Doctors can not get the medications right (and they don’t listen when told why the things are the way they are based on a long documented medical history).

So yeh pardon me if I’m a little jealous of people allowed to do any kind of payed work, even minimum wage.

Leave a comment

Blog moderation policy

Login

Allowed HTML <a href="URL"> • <em> <cite> <i> • <strong> <b> • <sub> <sup> • <ul> <ol> <li> • <blockquote> <pre> Markdown Extra syntax via https://michelf.ca/projects/php-markdown/extra/

Sidebar photo of Bruce Schneier by Joe MacInnis.