Top
Best
New

Posted by quantumgarbage 21 hours ago

Stealing Reasoning Traces from Proprietary LLM APIs(stolen-thoughts.com)
626 points | 283 commentspage 2
myworkaccount2 19 hours ago|
Is this how the eastern labs "distill" SOTA models?

If you can play it right, you don't even need to send suspicious prompts to the frontier models. Just use them for regular tasks, extract the encrypted COT blocks and replay it to a cheaper model to get the plain text COT.

But the real question is: Is it okay to steal from a thief's hoard?

NitpickLawyer 19 hours ago||
> But the real question is: Is it okay to steal

By definition it cannot be stealing since you're paying for the tokens. It may be against their ToS, depending on what you end up doing with those tokens, but it cannot be stealing. If they charge by the token, all your tokens are belong to you :)

I also find it very strange that everyone sort of accepts their ToS like no big deal. Imagine MS using the same terms for their software - you cannot use any MS software to develop competing services. Bananas! They'd be dragged through the courts like it's the 90s.

(I get why they're doing it. Distillation is unreasonably effective. But still, I find it bananas that we've kinda accepted it, to the point where people use "stealing" or "attack" or any such terms)

desterothx 18 hours ago|||
I love how some of the biggest advancements in llms came from the Chinese labs, yet people still jump to distillation being unreasonably effective. Distillation is very good at creating smaller models from large ones sure, but nothing to me indicates it is 'unreasonably effective' compared to all the other bells and whistles being iterated on
rfoo 17 hours ago|||
Let's face it. Chinese labs made some of the biggest advancements. AND training on Claude (or GPT) output IS unreasonably effective. The two sentences are true at the same time.
varenc 8 hours ago|||
This article shows that when Kimi3's chain of thought is prefilled to match Opus's, the rest of the chain of thoughts Kimi3 outputs very closely aligns with Opus's. That seems strong evidence that Kimi3 is partly a distillation of Opus. And Kimi3 is not a small model. No doubt a lot of hard work went into Kimi, but seems clear that distillation was used effectively as well.

(though maybe there's another interpretation of the thought alignment?)

wonnage 7 hours ago||
Didn’t Kimi3 release a week before opus 5?
Otterly99 2 hours ago||
They compare it to Opus 4.8 in the article, which has been available for a few months now.
tuesdaynight 19 hours ago||||
It's like sideloading. It's very hard to fight against the marketing budget of big tech
articulatepang 8 hours ago||||
> If they charge by the token, all your tokens are belong to you

I’m not sure this argument is correct. You can sign whatever contract you like with the model provider, right? Including “you are entitled to the end product but not the intermediate scratch work”?

Coming from a place of genuine curiosity: is there some precedent or statute that would invalidate that contract? I don’t see why the reasoning tokens belong to you.

For example, I pay lawyers by the hour but don’t necessarily own their meeting minutes, recorded discussions, research notes, etc.

NitpickLawyer 7 hours ago||
Sure, but the current one is charged per token in & token out. Not per completion / task / hour / whatever. You can't charge per token and then say "you stole that token". Again, they can unilaterally decide not to sell you tokens anymore, at any time, for any (legal) reason. But as it stands right now, it can't be stealing.
blackqueeriroh 6 hours ago||
Read the TOS. It can absolutely be stealing.

Are you a lawyer?

NitpickLawyer 5 hours ago||
Breaking a platform's ToS is a civil contract violation, not a criminal offence. Stealing is. Potato, avocado.
senordevnyc 14 hours ago|||
I hesitate to nitpick with regard to something legal, given your username, but what makes this different from hiring a consultant with the agreement that their final output belongs to you, but you don't get access to their internal processes, tooling, notes, etc? Or a photographer where you get final edited prints, but you don't get the raw photos?
wyan 12 hours ago|||
Your examples are cases in which you know what the bill is going to be before placing your order. With LLMs, you're paying per output token, not per request, yet you don't get all the tokens.
articulatepang 7 hours ago||
When you hire lawyers or consultants you usually don’t know how many hours they’ll bill you. It will depend on developments in the case that you cannot in general predict. For example if the other side files a motion and your lawyer has to argue against it, they’ll bill you for it.

Sure you can set spending limits, just like you can make an account and give it a limited amount of credits.

fapjacks 12 hours ago|||
Or a huge software company where you only get the end operating system, but none of its source code.

Buy the Neiman Marcus cookies and feel entitled to the recipe?

Lots of secret sauce in the world.

qwytw 15 hours ago|||
>But the real question is: Is it okay to steal from a thief's hoard?

How does this relate to your previous paragraphs? LLM outputs are not copyrightable and you didn't break into Anthropic servers to steal the files from there. So how exactly is it theft? If I send an "encrypted" files to thousands of peoples and some manage to figure out how to read it I can't really accuse them of that or can I?

orbital-decay 18 hours ago|||
Not necessarily. There's a million ways to jailbreak any current model to show the trace and bypass all guardrails, or hijack and modify it. It's just one of them.
azinman2 19 hours ago||
The reasoning blocks are not stolen/mined from the internet at large directly. They’re the result of a lot of research, time, money, and expertise into creating a reasoning model. To me the answer is quite clearly no, especially when the encrypted blocks demonstrate they want to protect it.
pyrale 19 hours ago|||
Stuff available on the internet is also the result of a lot of research, time, money, and expertise. And AI companies taught us that it’s OK to yoink whatever is not bolted to the ground, even when it is illegal to do so.
qwytw 15 hours ago||||
That's a moral stance one can take (regardless of the severe cognitive dissonance embedded in it). But what does that have to do with theft? LLM providers don't own the copyrights to the outputs of their models (at least not yet).
tristanj 19 hours ago||||
No. There are dozens of companies that resell tokens at a discount to collect and resell session data to various Chinese labs.
flawn 7 hours ago||
So you say, they at least create economic value through obscurity of something which should be accessible?
polymer8563 13 hours ago||||
if you make reasoning soup of my data withou my consent it still is my data and i did not ask for your reasoning soup
elzbardico 19 hours ago||||
Most post-training tasks are based on real open source projects. A lot of time on real issues posted on issue trackers.

Besides that, the capabilities of a model are heavily dependent on the unsupervised learning phase, that gobbles all kind of other people's IP without giving a fuck. All the underpaid work behind the masses of third world programmers creating those post-training datasets would be completely uselless without it.

Also, it is kind of funny that labs resort to the "Research, time, money and expertise" argumet, when it is basically the same argument from publishers and other IP creator that the labs spent millions of dollars of lawyering money to resist. Besides, US law rejects in: Effort and cost by themselves not necessarely generate protectable interests.

About encryption, I think we're all contaminated by the bad ideology behind DMCA. While encryption established the intent, it doesn't follow that they have a legal claim of exclusivity just because of it.

Technically, you're overstating the value of so called "reasoning traces". You can't infer the verifier design, the reward shaping,or the data pipeline from them. Also, what you can extract are not the traces themselves, but the written summary of it, and you can't even guarantee that this summary reflects the exactly reasoning trace, models have show to have lied about it. Besides, distillation works when the student model already has strong priors, you can't turn a weak model in a SOTA with it. Don't believe Amodei's outrageous lies about it, he is just trying to exercise some regulatory capture.

blackqueeriroh 6 hours ago||
Are you a judge or a lawyer? If not, then you don’t know one way or another.
Otterly99 2 hours ago||
I really wonder how much of safeguarding with SOTA models is actually just "Don't do that" in a prompt?

I'm building agent to control machines at work and have to rely on small models, which are especially bad at remembering rules like this, so it is very apparent to me that soft rules like this are pretty much useless. Might be different for these huge models, but still it seems very shaky to me.

x312 20 hours ago||
Super cool that this works. I'm surprised these companies re-use the same encryption key across models!

I wonder if you can use these for attacks, like this previous paper showing that if you know how a model reasons, you can "fake its thinking" to control it? https://news.ycombinator.com/item?id=48631888

flexagoon 19 hours ago||
> I'm surprised these companies re-use the same encryption key across models

I assume switching the model in the middle of a conversation is intended behavior (very useful in coding agents, for example)

yubblegum 19 hours ago||
Seriously, what does it take to encrypt per session? There are many ways to make it scalable and efficient so I am wondering if this is left like this to allow interested 3rd parties ahem unobtrusively peek what people are doing with the AI.

(Thanks for the link. That’s an interesting idea!)

dannyw 18 hours ago||
The provider has the hidden text anyway; this isn’t customer managed encryption.
yubblegum 16 hours ago||
Sure, but if each session has a unique key then these need to be managed and stored and unauthorized access to these leaves tracks. So all that had to be 'compromised' is a single universally applicable key. Again, the question stands: session based encryption can be scalable and efficient. Why aren't they using it?
paxys 14 hours ago||
The exploit here isn’t a leaked encryption key. It’s pretty likely that they are already using a unique key per conversation. The raw CoT eventually reaches the model, and you can convince the model to share it with you.
theapadayo 13 hours ago|||
Yeah encryption isn't the issue. The only way I see to fix this is if you stop the user from switching models mid-session, or strip out the thoughts when switching models. Either way you're degrading the user experience.
yubblegum 13 hours ago|||
If a different model is using encyrpted blocks of another model, then by definition it is no longer a session scoped bit of information. Since you can give it to any other session and another model, clearly it doesn't even have to be the same user. Therefore, there is only one (set) of universally available key(s) used by all models across all sessions.
NegativeAbsence 1 hour ago||
Last year's reports questioning whether reasoning blocks were actually reflected in the final response were why I stopped using reasoning models altogether. I switched to a separate pipeline and have used that ever since. It's good to see that the concern didn't remain just a suspicion.
nervai 20 hours ago||
Really cool work, you get the actual traces. Looks like the vendors can all reliably fix this one though.

A harder to defend against approach here where they work backwards from the results and ask the model to generate a plausible trace: How to Steal Reasoning Without Reasoning Traces https://arxiv.org/pdf/2603.07267

dannyw 18 hours ago|
Trace Inversion is fascinating, but it’s more of an independent reconstruction that will give you some coherent-looking generated CoT; but not necessarily anywhere close or related to the underlying model’s CoT.
nervai 11 hours ago||
I didn't read the paper in details but they claim there is high overlap between the synthetic traces and the ground truth ones (not sure how they confirmed that for blackbox models though, I guess they must have compared to open source models).

They also talk about successful distillation of black box model capabilities with the approach.

vinaigrette 19 hours ago||
I must say right of the bat this is the best research paper/working paper in regards to its styling. Beautiful
SwellJoe 19 hours ago||
I agree on desktop/laptop, but on mobile there are images that appear under the text making it hard to read.
user43928 18 hours ago||
On an iPhone Pro Max only the first trace is readable.

Navigating to the right lands between two cards, so that neither is readable.

yetanotherjosh 15 hours ago|||
My brain can't tell if the text is horizontal or slightly rotated. It's very hard to read. Beautiful to some, inaccessible to others.
thefourthchime 18 hours ago||
I was going to comment on that. This is clearly a vibe-coded webpage. It sort of smells like GPT to me, or at least front-end design. But the author clearly went back and forth to make it beautiful. This is not the first output he got.

This is the kind of stuff I point to when people talk about AI slop. AI is just a tool. You're still the person who has to deliver the output and have some taste.

iamcoder18 20 hours ago||
This proves that OpenAI models reason in grug speak to save tokens! I wonder if open models are going to start doing that too to save on reasoning tokens.
kgeist 19 hours ago||
In the BlackHat presentation on the HuggingFace incident, OpenAI showed some excerpts from the reasoning traces, and they had that grug speak too (skipped articles, etc.). So the OP's method must have indeed found the actual reasoning traces.
wren6991 14 hours ago|||
> I wonder if open models are going to start doing that too

Yes, some of them do do that. For example Moonshot tried to reward shorter reasoning traces in between Kimi-K2.6 and Kimi-K2.7 Code, and the latter has a mild caveman accent in its reasoning traces that the former lacks.

Qwen3.8-Max also has terse reasoning, but I don't remember this being the case for Qwen3.6 models I ran locally.

lukewarm707 17 hours ago|||
their gpt-oss models do the same. i don't use closed models so i never thought much about it.
gaigalas 18 hours ago||
Muse clearly does it to some extent. Saw a lot of that running Glimmer locally.
varenc 8 hours ago||
super interesting. So pre-filling Kimi3 reasoning with Opus's reasoning results in thoughts that closely match Opus's. This seems like strong evidence Kimi3 was trained on decrypted Opus chain-of-thought. Meaning the Kimi team likely also broke CoT encryption. Though not exactly a big surprise.
EagleEdge 17 hours ago||
I used to do a very coarse version of this stealing. I ask a question from ChatGPT pro, once it is done, I ask claude chrome add-in to go through all those thinking from the side bar, extract everything along with all the sources used. Then try to reverse engineer the solution it came up with.
infecto 15 hours ago|
Wouldn’t the fix be to encrypt before the api call hits the LLM? You would encrypt/decrypt at a separate layer than the LLM. I am sure i am missing something but would love to be educated.
paxys 14 hours ago||
The LLM needs to read the CoT as part of the conversation. You can ask the models to share them with you. Stronger models will refuse, while weaker ones can be “jailbroken”.
infecto 13 hours ago||
I don’t think that answers what I am wondering. Asked differently why is encoding/decoding the cot the concern of the LLM?
paxys 13 hours ago||
It isn’t the concern of the LLM. Regardless of where the encryption/decryption is happening, the issue is that the LLM needs to access the raw CoT.
infecto 13 hours ago||
Again I don’t think you’re really getting at what I am asking. Sorry. My whole point was why does the LLM have access of decrypting. It should happen outside of the LLM layer.
pas 5 hours ago|||
LLM does not work on encrypted tokens. It happens at the API gateway.
infecto 9 hours ago|||
Wild this would get downvoted. I am asking a question, the bots must have come in.
henryaj 11 hours ago||
I assume it does happen at a different layer, just that that layer is common to all of a provider's models to make conversations portable across models (otherwise the reasoning blocks would all need to be re-encrypted for them work with another model)
More comments...