Guess they don't care about regular devs atm and are focused only on hardware sales.
They don’t have a prefix cache like other providers, or at least don’t have a discount in their billing structure. Each message charges for the whole context window. It’s wildly more expensive for long multi turn scenarios with lots of tool calls (coding). It’s better for short few turn tasks.
Edit: I don’t know if they actually have a proper cache. This could just be a billing artifact.
I still think that was a really great model that got overlooked. It was really great in terms of latency/throughput while still being fairly intelligent.
I was planning on using it for a design tool, but moved over to luna since it's comparable speeds and cost for a lot more intelligence.
Everyone should occasionally go back to the old models to see how much worse they were, like even a year ago you could generate results but they were typically full of bugs and you have to fix a non-insignificant amount of it all manually: https://blog.kronis.dev/blog/i-blew-through-24-million-token...
Admittedly that post was before agentic development truly took off and that 3k EUR figure when paying per API tokens would nowadays be closer to like 6k EUR for the volume of work I do, but still.
It's the same how Qwen 2.5 was pretty problematic for anything remotely serious, same with Qwen 3 Coder Next (80B), and at least the most recent versions are getting better but still not quite good enough in real world use cases outside of benchmarks. They've come a long way, regardless!
I'm pretty sure Cerebras has a confidentiality agreement with OpenAI, and this press release was carefully constructed to avoid leaking details about the model weights. For example, the graph of tokens per second vs. tokens per second per user doesn't have any numbers that would allow you to translate between the two. (And in any case the relationship depends on the model.)
This was also interesting: "CS-4 delivers more than 1,000 tokens per second on models exceeding 10 trillion parameters." Was it known that there were 10 trillion parameter models in use?
I think the frontier providers keep the size of their models carefully hidden.
> According to FT, industry estimates say Anthropic's most advanced Mythos 5 has about 8 trillion parameters and Fable 5 about 5 trillion
https://www.reuters.com/technology/bytedance-targets-mega-ai...
I believe this report has confused Opus (which is known to be around 5T) and Fable.
Other reports say 10T. See for example https://eu.36kr.com/en/p/3760679047267075?ref=explainx where Musk talks about the models being trained on Colossus2
5T for Opus feels quite high though. DeepSeek V4 Pro is a mere 1.6T and often described as a match with Opus in overall quality. Even the largest open models in common use are around 2.8T.
It's not. Idk about who has more T's but, unfortunately, DS4 pro is not a match to Opus, at least not Opus 4.8.
The difference is very visible in long tail applications. Exactly where you'd expect parameter count to matter.
There is nothing to say for example a 1 Quadrillion parameter model will be vastly more intelligent than current SOTA especially since new training data is largely synthetic today
I think there is some merit in that smaller models cannot memorize so much of the training data, i.e. that they are less likely to do copyright infringement, and by analogy not having memorized SDK / API surfaces that have since changed from the training data
You have to rely on it to a certain level for agentic/coding work, presuming that's the general subject we're talking about here... For instance I recently encountered a project where it would have been a lot worse if the LLM didn't already know "what is" xterm.js and a bunch of its associated npm-related/node related software. If it was still smart but had to google and find results for everything it would have been a lot more time consuming and risked sending it down a wrong path.
But these labs distill off the larger models. Both officially at the labs with the big ones, and unofficially. We need the giant models to get the smaller models.
The cost to train and infer that would be insane, even by today's standards.
Eg: https://www.reuters.com/technology/bytedance-targets-mega-ai...
That reports Mythos as 8T and Fable as 5T, but I think they mean Opus as 5T, which is widely known, eg: https://eu.36kr.com/en/p/3760679047267075?ref=explainx
Both Grok and Bytedance are training 10T models.
Honestly if Opus is 5T parameters while being matched by the biggest open models that are at least twice smaller, it would mean that the US is already behind China in the AI race, despite a significant edge in compute.
For example I regularly do Fable+Opus agentic coding runs over 24 hours without intervention.
I think I've had GLM do a run that was a few hours. That's the closest I've had an open model come on that kind of work.
This assumption is likely what has led to the erroneous failure.
Enterprise compute per rack has scaled multiple fold in the last 3-5 years. Alongside the training efficiency gains & datacenter scale increases, even 50T+ is well within reach at the top end.
1. https://www.youtube.com/watch?v=3MKRjt59hh4&pp=0gcJCRMMAYcqI...
How about in a month or so when you have to run a slightly different workload?
In fact both of them, actually Amazon too, invest in their own inference hardware and owns the stack.
You can't possibly think that these companies will keep shelling 50-100B per year in hardware alone where 60%+ is margin for Nvidia and not invest there.
If it's patents that are the problem then presumably all these large semiconductor companies have defensive parent portfolios.
NVLink is at gen9. they had a lot of teething problems and can codesign the hardware and software.
in the name of openness (AMD's only """weapon"""), the UALink spec is a hodgepodge of corporate opinions with very different implementations (looking at you, Broadcom). at spec version 1 (in hardware).
I wish them good luck as I really like AMD, but they compete no more on this than Lambo vs Bugatti.
Oops did they just out GPT-5.6 sol’s parameter count?
If Fable turns out to be a 10T or 20T model, there is little to boast vs Kimi at 3T. But the opposite is true: if Fable were to be e.g. a 500B model, that would show how far ahead they are from the open models. This isn't likely to be the case ...
It seems it takes some time to run a new model on all the benchies, not sure they run all models on all of them either
All or a vast majority of of the cerebras manufacturing capacity was going to a few companies that aren't publicly available inference providers on openrouter, for their own internal use.
or
The asking price of the S-3, no matter how speedy it might be, for small/medium size customers made it economically prohibitive to purchase and use to sell public inference vs. buying more common nvidia b200 or whatever.
imo the issue is that most openrouter demand is inauthentic activity (things that anthropic and openai models will refuse to do like pretend to not be bots when interacting with humans)
So I plugged 288 trillion tokens/month (OpenRouter's current rate), 500 billion MoE model average, and the math comes out to be around 620 B200 GPUs minimum.
So basically, OpenRouter's volume must be absolutely tiny compared to the volume hyperscalers are getting.
if you have the money as an "individual user" to purchase one of their racks... save your money and retire.
* Actually they sent out an email claiming they already have it, but I don't seem to have access, they're promising to release it to the "shared tier" any day now.
Now that this hypothetical person has retired, what are they gonna do all day? Just sit on the beach and drink Mai Tais? If that's what they wanna do, sure, but nerds gonna nerd, and if I had that kind of money to retire on, I'd totally buy some ridiculously expensive AI box for fun.
Can you imagine something radiating that much energy into a space in your home?
This is far beyond the practical maximums of like 10 to 15kW per 44U cabinet front to rear air cooling for 'regular' rackmount server stuff.
More realistically, you need much more cooling water.