Top
Best
New

Posted by softwaredoug 8 hours ago

“We have information that Moonshot distilled Fable for the development of K3”(twitter.com)
https://xcancel.com/mkratsios47/status/2079933645888880708
187 points | 462 commentspage 3
nchmy 4 hours ago|
We have information that Claude distilled billions of copyrighted, and otherwise-created-by-others, materials for the development of their entire business.
feverzsj 6 hours ago||
If web scraping is legal, so is distilling.
dgellow 5 hours ago|
I would support distilling even if scraping wasn’t legal, I don’t think there is much of a relationship between the two
throwa356262 6 hours ago||
In the meantime, reddit is making fun of Opus for "distilling" Qwen:

https://www.reddit.com/r/ClaudeCode/comments/1tqaist/opus_48...

(don't take this too seriously)

mrandish 5 hours ago||
How was K3 trained on data distilled from Fable when Fable was only publicly available in the last two weeks before K3 was released? The timing just doesn't work.
grim_io 5 hours ago||
So, if it's that easy and fast to "copy" Fable, is it really worth that much in the first place?

Sounds like the opposite of the conversation Anthropic would want to have.

xinayder 1 hour ago||
I wonder if this was caught with the malware code Anthropic included that detects if you're in China...
goldylochness 2 hours ago||
it was obvious from day 1 this is what happened

the chinese labs haven't been improving at training from scratch, they've been getting better and better at distilling models, so naturally they continue to follow this path

firstly, it shows the vulnerability of exposing a model to users. a highly capable group can quite literally suck the functionality out of your model and take it for themselves, so there's no sandboxing it

secondly, it's actually interesting to know that distillation is so powerful. in the science-fiction scenario of meeting some other form of intelligent life, they might have their own models and this would be a way of siphoning intelligence off of them

jmward01 5 hours ago||
If 'distillation' means training on outputs then what is the legal concept of ownership of outputs? And, more broadly, is this something that could be skirted by doing it in different countries that have different legal structures? Basically, are they saying they own those outputs, not the companies that paid for the tokens, and only they can train on them? I suspect a lot of companies are saving their token histories and using them to fine tune internal models.
nradov 5 hours ago|
The legal concept is that LLM vendors can put pretty much whatever they want in their terms of service, and cut off or sue clients who violate those terms. They have the right to refuse service to anyone for any reason (or no reason at all).
softwaredoug 3 hours ago||
OpenAI and Anthropic should enter into distillation agreements with other US labs. Turn a threat into a profit center.

Other US labs cannot directly distill from OpenAI/Anthropic as it’s a violation of the terms of service. It holds other US labs back. Leading them to build second tier models And in the end OpenAI/Anthropic may be unable to prevent distillation.

Why fight it when there’s clear money to make here?

gozucito 1 hour ago|
OpenAI/Anthropic are already charging for access to their closed models. They're getting paid.

They are also not interested in agreements. They want to keep as big a moat as possible because they love money. And you need two to tango.

softwaredoug 53 minutes ago||
Yeah they may not be interested. But pretending you have that moat a dumb strategy that's not working.
jerrythegerbil 5 hours ago|
“However, large-scale, covert industrial distillation aimed at stealing proprietary U.S. technology and undermining American research is unacceptable.”

What’s actually happening behind the scenes is that certain inference providers will classify a prompt and it’s re-routed transparently to Anthropic and that’s used for distillation training, only distilling the complicated traces they need, originating from real user prompts and traces. These inference providers are explicitly blocked in the claude cli if you reverse engineer it.

The real picture is that these Chinese labs have figured out how to get exactly what they need, at a high quality, directly from distinct and unique real user prompts.

It’s only “covert” because Anthropic doesn’t like it, while simultaneously being perfectly fine to do.

More comments...