Is the optimal formula still 20x the amount of model params in tokens for training? Could this mean we're getting a GLM with 1.5t params?
Like were executive at Google and AWS and Microsoft expecting this kind of performance from models smaller than what openai/anthropic have been doing? Are we really in a "compute desert"?
Qwen 3.8 Next Flash: 125B + 51B = 176B parameters with 6B activated
DeepSeek V4 Flash: 284B with 13B activated
The new Qwen model is the most promising for one or two Strix Halo 128GB with the low number of active parameters. On paper it's much stronger than Qwen 3.8 27B.
It's too big, bright and resourceful of a country to choose confrontation instead of collaboration.
It's not like we didn't try it. China first have to learn to make deals where both party benefits.
That's good. Keep going.
Now the US is behind in EVs can you guess what they're doing? [1]
[1] https://evwire.com/p/video-ford-ceo-jim-farley-says-they-fly...
Intellectual property is part of WTO agreements but enforcement is domestic.
US companies do it too, regularly, they simply hire and poach staff from competitors.
Proving it to be IP theft is difficult unless you can prove documents being passed. But often all you need is the know-how of the hired talent.
"problem" indeed.
If the USA wanted a copyright treaty with China bad enough, we would negotiate one. China is not breaking any laws here, international or otherwise.
(281 points, 118 comments)