Posted by softwaredoug 9 hours ago
The US gov and AI providers when funny chinese people steal their data to train their models: >:(
clowns
edit: TIL you can't use emojis on HN
I love using Claude but Fable's unusable wrt useful work like cryptography, biology, &c.
Kneecapping my productivity when I pay $100/month is annoying af.
If LLM outputs aren't copywriteable and you create your own synthetic training set using Fable and share it publicly on huggingface, and someone else uses that training set to fine-tune a model, would this be considered illegal?
I ask because this happens all the time, synthetic datasets have basically become a key aspect of training a model at this point. I even generated a synthetic set from DeepSeek v4 to aid in fine-tuning a classifier just a few weeks ago.
So I just wonder on what grounds any of this makes sense, I wouldn't be surprised if some of these American labs were using open models on their own self hosted infrastructure to generate training data, but by nature of them being open nobody has to know.
I'll make a prediction: I don't think we will ever see any of the evidence of this "distillation" before they end up implementing some type of ban.
1) Compensation of right holders is one issue.
2) Distilling models is an entirely separate issue, because model building is value add, and that is important because if we arrive at a place where you can produce a model, that gets to ~100% of what people perceive of the models value (on top of also not compensating right holders, yourself) you are discouraging development of better models and, again, in no way helping with issue 1)
Unless anyone actually distills a model and then also does something for rights holders, any schadenfreude simply detracts from this issue, in addition to the other issue (well, that might not be an issue if we would rather slow down model development right now, but again, forever worse models still don't help solve issue 1)
Mendel doesn't get a cut every time somebody uses the principles of heritability he discovered, and Einstein's family aren't getting royalties if you compute relative speeds. I think the frontier labs should expect to be treated more like scientists than artists in this regard.
Because there is no way in hell I'm going to make an effort creating quality content for existing platforms. The website should be entirely my own without moderation subject only to my local legal system.
Can just insert this comment as a prompt and vibe code everything in a few days⸮
It's about the narrative that "Chinese models are at Fable level". The truth (if correct) is the China continues to copy, and the proprietary US Models continue to lead the state of the art.
There is no K4 without Fable 6, GPT-6. That, matters.
That's simply not true though. Chinese labs very clearly have the entire stack developed and working. Using traces from claude allows them to shorten their training time by some amount, that's it.
Remove Fable 6 and you still have K4 eventually, just 2 months later at best.