Top
Best
New

Posted by sunils34 11 hours ago

Cerebras CS-4(www.cerebras.ai)
315 points | 207 comments
rbanffy 1 minute ago|
Impressive that this is an "interim" product, the start of a new line that ought to be continued with the WSE-4 family, where they are supposed to use a 3nm process and, maybe, 3D stacked SRAM. The modular architecture also points towards field upgrades that are badly needed for AI datacenter builders.
KronisLV 3 hours ago||
God I wish they'd back up all of those claims by offering a subscription of Kimi K3 and GLM 5.3, not some outdated GLM 4.7 instance that they then proceed to call a preview model and say that they'll remove it, leaving users only with GPT-OSS 120B which is nigh useless nowadays: https://support.cerebras.net/articles/9996007307-cerebras-co... and https://www.cerebras.ai/pricing

Guess they don't care about regular devs atm and are focused only on hardware sales.

scosman 2 hours ago||
I don’t think they will until they change the architecture.

They don’t have a prefix cache like other providers, or at least don’t have a discount in their billing structure. Each message charges for the whole context window. It’s wildly more expensive for long multi turn scenarios with lots of tool calls (coding). It’s better for short few turn tasks.

Edit: I don’t know if they actually have a proper cache. This could just be a billing artifact.

preommr 2 hours ago||
> GPT-OSS 120B which is nigh useless nowadays:

I still think that was a really great model that got overlooked. It was really great in terms of latency/throughput while still being fairly intelligent.

I was planning on using it for a design tool, but moved over to luna since it's comparable speeds and cost for a lot more intelligence.

KronisLV 2 hours ago|||
> I was planning on using it for a design tool, but moved over to luna since it's comparable speeds and cost for a lot more intelligence.

Everyone should occasionally go back to the old models to see how much worse they were, like even a year ago you could generate results but they were typically full of bugs and you have to fix a non-insignificant amount of it all manually: https://blog.kronis.dev/blog/i-blew-through-24-million-token...

Admittedly that post was before agentic development truly took off and that 3k EUR figure when paying per API tokens would nowadays be closer to like 6k EUR for the volume of work I do, but still.

It's the same how Qwen 2.5 was pretty problematic for anything remotely serious, same with Qwen 3 Coder Next (80B), and at least the most recent versions are getting better but still not quite good enough in real world use cases outside of benchmarks. They've come a long way, regardless!

wongarsu 2 hours ago||||
As MoE with 5B active parameters it's pretty fast. But you still need a lot of vRAM, or have to run small quantitations. Qwen models just gave you more bang for your buck, and the gap became worse with every qwen release
LoganDark 1 hour ago|||
gpt-oss-120b is absolutely unusable over Cerebras. It fails to call tools half the time and just continues to think about what tool it'll call repeatedly. Like it says it'll call a tool and then it doesn't, and then it says it'll call the tool again and then it doesn't, and it just does that in a loop forever. It's awful. Also forgets to end the thinking block too. Even if the model itself was just-okay for its time, even at 1000t/s+ it's not worth it. And it's EXPENSIVE, like $5 per minute expensive
syntaxing 10 hours ago||
I think the fun takeaway from this is that GPT 5.4 is probably 45B active parameters and GPT 5.6 Sol is closer to 50B.
yorwba 3 hours ago||
You cannot infer this because they only show the tokens per second per user. One way to get a higher number is to have fewer users per chip.

I'm pretty sure Cerebras has a confidentiality agreement with OpenAI, and this press release was carefully constructed to avoid leaking details about the model weights. For example, the graph of tokens per second vs. tokens per second per user doesn't have any numbers that would allow you to translate between the two. (And in any case the relationship depends on the model.)

Copenjin 2 hours ago|||
Didn't they say that they can support bigger models now?
logicallee 9 hours ago||
(Where did you see that?)

This was also interesting: "CS-4 delivers more than 1,000 tokens per second on models exceeding 10 trillion parameters." Was it known that there were 10 trillion parameter models in use?

I think the frontier providers keep the size of their models carefully hidden.

redox99 6 hours ago|||
You don't really need to train a 10T model to test cerebras against a 10T model. You can feed it an untrained (randomly initialized) model and benchmark it. Result will be gibberish but performance the same.
nl 6 hours ago||||
Mythos/Fable are around 10T:

> According to FT, industry estimates say Anthropic's most advanced Mythos 5 has about 8 trillion parameters and Fable 5 about 5 trillion

https://www.reuters.com/technology/bytedance-targets-mega-ai...

I believe this report has confused Opus (which is known to be around 5T) and Fable.

Other reports say 10T. See for example https://eu.36kr.com/en/p/3760679047267075?ref=explainx where Musk talks about the models being trained on Colossus2

YmiYugy 5 hours ago|||
I'm confused. I thought Mythos 5 and Fable 5 were exactly the same model just with a different security layer in front of it. Could they mean the Mythos 5 Preview?
nl 3 hours ago||
Yes. One reason why I think that report has confused Fable and Opus.
zozbot234 6 hours ago|||
> I believe this report has confused Opus (which is known to be around 5T) and Fable.

5T for Opus feels quite high though. DeepSeek V4 Pro is a mere 1.6T and often described as a match with Opus in overall quality. Even the largest open models in common use are around 2.8T.

Implicated 5 hours ago|||
> and often described as a match with Opus in overall quality

It's not. Idk about who has more T's but, unfortunately, DS4 pro is not a match to Opus, at least not Opus 4.8.

nl 3 hours ago|||
It's not an Opus match.

The difference is very visible in long tail applications. Exactly where you'd expect parameter count to matter.

alightsoul 7 hours ago||||
pretty sure 10 trillion parameters is now the norm among closed ai labs, given that nvidia also references the same 10 trillion number for their nvl72 racks
KronisLV 3 hours ago||
Pretty bad efficiency then unless that only applies to Fable class but even then - Kimi K3 is around 3T and does similarly well in most benchmarks.
ewild 9 hours ago|||
It's rumored fable is around that 10T number
walrus01 9 hours ago|||
If this is true, it's even more impressive that some of the open weight models that are <3.5T in size, approx 33% of its size, are within a few points of it in the artificial analysis leaderboard.
manquer 8 hours ago|||
Not necessarily, there could be diminishing returns on mere parameters count .

There is nothing to say for example a 1 Quadrillion parameter model will be vastly more intelligent than current SOTA especially since new training data is largely synthetic today

HDBaseT 8 hours ago||
That's precisely what he is saying, there is diminishing returns (or optimization left on the table).
manquer 8 hours ago||
I read it as it is impressive because smaller models 2.5T are squeezing similar returns as 10T models despite being 1/4th size not that there beyond 2T today the number or parameters do not have much meaning
verdverm 7 hours ago||
or the latest qwen3.8 27B doing so well at ~1/100 the size of K3
rnewme 6 hours ago||
What about general knowledge you can get out of it before hallucinations start?
stymaar 5 hours ago|||
Storing general knowledge in VRAM has always been a dumb idea in the first place.
Balinares 2 hours ago||||
Qwen 3.8 27B beats Opus, Fable and GPT 5.6 by a comfortable margin on the AA-Omniscience Hallucination Rate benchmark.
walrus01 6 hours ago||||
It did OK on schlongbench v1.0 (test of a specific niche word that doesn't make it into smaller LLMs) but it sure does love to count words

https://pastes.io/r8F1AY8h

verdverm 6 hours ago|||
I do not rely on any LLM of any size for general knowledge baked into the weights, they all hallucinate and that is the wrong way to hold them imo

I think there is some merit in that smaller models cannot memorize so much of the training data, i.e. that they are less likely to do copyright infringement, and by analogy not having memorized SDK / API surfaces that have since changed from the training data

walrus01 4 hours ago||
> I do not rely on any LLM of any size for general knowledge baked into the weights

You have to rely on it to a certain level for agentic/coding work, presuming that's the general subject we're talking about here... For instance I recently encountered a project where it would have been a lot worse if the LLM didn't already know "what is" xterm.js and a bunch of its associated npm-related/node related software. If it was still smart but had to google and find results for everything it would have been a lot more time consuming and risked sending it down a wrong path.

scosman 2 hours ago||||
And GLM is only 0.7T!

But these labs distill off the larger models. Both officially at the labs with the big ones, and unofficially. We need the giant models to get the smaller models.

mlmonkey 8 hours ago||||
You want to take a look at the "Scaling Laws" paper, so you can extrapolate from these numbers.
stymaar 5 hours ago||
This paper, as well as the Chinchilla one, aged like milk though.
habosa 8 hours ago|||
GLM 5.3 is "only" 753B parameters. Much much smaller.
johnnyApplePRNG 7 hours ago|||
Fable is most definitely nowhere near 10T.

The cost to train and infer that would be insane, even by today's standards.

nl 7 hours ago|||
Fable is strongly believed to be around 10T. The most conservative estimate I've seen is 8T.

Eg: https://www.reuters.com/technology/bytedance-targets-mega-ai...

That reports Mythos as 8T and Fable as 5T, but I think they mean Opus as 5T, which is widely known, eg: https://eu.36kr.com/en/p/3760679047267075?ref=explainx

Both Grok and Bytedance are training 10T models.

stymaar 5 hours ago|||
The fact that Musk claims Opus is 5T to justify why Grok is far behind should be taken with a massive grain of salt given he's a recidivist mythomaniac.

Honestly if Opus is 5T parameters while being matched by the biggest open models that are at least twice smaller, it would mean that the US is already behind China in the AI race, despite a significant edge in compute.

nl 27 minutes ago|||
The open models don't really match Opus.

For example I regularly do Fable+Opus agentic coding runs over 24 hours without intervention.

I think I've had GLM do a run that was a few hours. That's the closest I've had an open model come on that kind of work.

WinstonSmith84 3 hours ago|||
Yes. And Opus goes a very long way compared to Fable, Anthropic isn't doing any favour, it's clearly just 2 models with a very different amount of parameters.
andai 6 hours ago|||
Wasn't Opus ~1.5T and Fable is about twice that?
x-complexity 5 hours ago||||
> The cost to train and infer that would be insane, even by today's standards.

This assumption is likely what has led to the erroneous failure.

Enterprise compute per rack has scaled multiple fold in the last 3-5 years. Alongside the training efficiency gains & datacenter scale increases, even 50T+ is well within reach at the top end.

riknos314 6 hours ago|||
Kimi K3 is a 2.8T model that's available at about 1/4-1/3 the cost of Fable from multiple providers on openrouter. The math doesn't seem wildly off.
zozbot234 5 hours ago||
The raw margins on proprietary model inference are rumored to be quite high though (they have to successfully defray the entire investment into model training and datacenter capacity for inference, which is massive enough). The API cost you're paying for the model includes that raw margin.
xmorse 56 minutes ago||
Cerebras is very fast but you can basically never use it because of its scarcity
sreekanth850 9 hours ago||
AMD along with cerebras may probably compete with NVIDIA monopoly in near future. Also, NVIDIA will have competition form multiple companies. Just my prediction.
anonzzzies 5 hours ago||
Like NVIDIA bought Groq, AMD might do well buying Cerebras.
FranGro78 1 hour ago|||
AMD did enter into an agreement to buy Taalas, which is speculated [1] will be used to augment their Helios offering.

1. https://www.youtube.com/watch?v=3MKRjt59hh4&pp=0gcJCRMMAYcqI...

sreekanth850 5 hours ago|||
They are Ex AMD employees. Nvidia tried back, but they rejected.
vezycash 3 hours ago||
The rejection might be overrulled by investors if the offer is enticing enough.
eitally 8 hours ago|||
Maybe, but GPU is just one aspect of NVIDIA's dominance. If you are buying Vera Rubin GPUs, you're getting an NVL72 rack, which is only one of several racks that you're probably buying. You'll also need your NVIDIA racks with NVIDIA networking & storage gear, too. At the end of the day, they're "vertically integrated" for your accelerated computing data center (e.g. the "AI Factory"). This doesn't even count the software layer, where CUDA + CUDA-X (not to mention the software for all the sysadmin pieces) has a huge first mover advantage over anyone else.
zarzavat 7 hours ago||
Is CUDA still a moat? Are we not at the point where frontier models can reimplement software stacks, given you throw enough tokens at the problem.
fooker 4 hours ago|||
If you are making a decision to spend 50B on hardware, would you use the proven tech stack or rely on engineers taking an unspecified amount of time vibecoding your software stack while the hardware sits idle?

How about in a month or so when you have to run a slightly different workload?

epolanski 3 hours ago||
Hyper scalers like Google or Microsoft, which are the big spenders, have all the incentives in the world to get more out of their gargantuan spending.

In fact both of them, actually Amazon too, invest in their own inference hardware and owns the stack.

You can't possibly think that these companies will keep shelling 50-100B per year in hardware alone where 60%+ is margin for Nvidia and not invest there.

someothherguyy 1 hour ago||||
i haven't seen it. see the recent browser attempts.
alightsoul 6 hours ago||||
if you are developing your own hardware, you provide your own stack to avoid lawsuits with nvidia. i don't think it's a technical problem at all, but a legal one. this is probably why zluda was scrapped by AMD and Intel. Nvidia technically bans the creation of CUDA reimplementations in their TOS if i remember correctly
zarzavat 6 hours ago|||
Doesn't Google v Oracle provide protection here? Copying APIs is fair use.

If it's patents that are the problem then presumably all these large semiconductor companies have defensive parent portfolios.

sreekanth850 6 hours ago|||
[flagged]
aurareturn 4 hours ago||||
You still need experts to know what is good and what is not.
vatsachak 7 hours ago|||
You just proved that AI cannot currently do that
incrudible 6 hours ago||
It can definitely create a software stack for you if you hold it right, but the software stack supported by a trillion dollar company with decades of expertise, that also uses AI to improve its stack is probably gonna be better.
_zoltan_ 4 hours ago|||
Press doubt. Single GPU? Maybe. MultiGPU behemoths like NVL144 and NVL576? I don't think so.

NVLink is at gen9. they had a lot of teething problems and can codesign the hardware and software.

in the name of openness (AMD's only """weapon"""), the UALink spec is a hodgepodge of corporate opinions with very different implementations (looking at you, Broadcom). at spec version 1 (in hardware).

I wish them good luck as I really like AMD, but they compete no more on this than Lambo vs Bugatti.

epolanski 3 hours ago||
It's not really a far fetched prediction: high margins and huge market attract competition, that's just the law of economics.
reilly3000 9 hours ago||
> CS-4 delivers more than 1,000 tokens per second on models exceeding 10 trillion parameters

Oops did they just out GPT-5.6 sol’s parameter count?

sho 9 hours ago||
Sol is supposed to be 5T according to rumour. The imminent Astra is allegedly 10
nozzlegear 8 hours ago||
Rumors and allegations aren't worth much. Why don't they just tell us mere mortals?
WinstonSmith84 3 hours ago|||
Because that would reveal their edge to investors, or the lack thereof.

If Fable turns out to be a 10T or 20T model, there is little to boast vs Kimi at 3T. But the opposite is true: if Fable were to be e.g. a 500B model, that would show how far ahead they are from the open models. This isn't likely to be the case ...

magicalhippo 1 hour ago||
I guess there's also the economic aspect. It would make it much easier for competitors to figure out your costs and margins if they know the model parameter sizes you operate.
brookst 7 hours ago|||
Why would they? What the upside, for them?
eigenspace 7 hours ago||
Yeah, its not like this js some sort of Open AI company. That'd be ridiculous.
whatever1 9 hours ago|||
I mean we kinda know the frontier models are multi trillion parameter models. The only open weights that are close to the frontier are that size too
verdverm 6 hours ago||
save qwen3.8 27B which is outclassing much larger models and is in spitting distance of the top 10 in https://artificialanalysis.ai/models#intelligence
Vax- 6 hours ago||
I wonder why they removed DeepSWE from their incorporates evaluations
verdverm 6 hours ago||
They didn't afaict https://artificialanalysis.ai/agents/coding-agents?coding-ag...

It seems it takes some time to run a new model on all the benchies, not sure they run all models on all of them either

kanwisher 4 hours ago||
cerebras model are different size then the original models
rajnathani 1 hour ago||
Interestingly they’re still on the WSE-3 (5nm TSMC) wafer chip and slightly bumped up the specs there (overlocking mostly it seems), for why it’s called WSE-3 Turbo now. I think people were also expecting WSE-4, as it’s been 2 years now since WSE-3 was launched.
ethanzhang1024 9 hours ago||
If cerebars is performing well, why didn't its predecessor, server S-3, become the largest API token provider on OpenRouter, surpassing the official model releases?
walrus01 9 hours ago||
Without having any inside information, one possible theory:

All or a vast majority of of the cerebras manufacturing capacity was going to a few companies that aren't publicly available inference providers on openrouter, for their own internal use.

or

The asking price of the S-3, no matter how speedy it might be, for small/medium size customers made it economically prohibitive to purchase and use to sell public inference vs. buying more common nvidia b200 or whatever.

wmf 7 hours ago|||
Cerebras provides high-speed inference at high cost. It's never going to be the cheapest and thus it will probably remain niche.
zurfer 3 hours ago||
But that's supply and demand, not technology. Right now a lot more people want their inference than they can supply. as supply catches up in the next 5-10 years, the underlying tech at scale is probably cheaper than GPUs per token produced.
aurareturn 4 hours ago|||
Probably the same reason why there are more people who takes buses, subways, trains than drive Ferraris.
epolanski 2 hours ago||
I love how instead of comparing a Ferrari (fast and expensive) to some average car (not fast, not expensive) to make your point..you went for public transport where your comparison cracks from multiple angles.
smallerize 9 hours ago|||
If you're willing to pay a significant premium for latency, why use openrouter? And anyway Cerebras only supported a few specific models.
gampleman 2 hours ago|||
Cerebras capacity was pretty much entirely bought out at some point. We needed it and couldn't get it.
aseipp 8 hours ago|||
The WSE is very expensive to build, and they have a waiting list of customers who are already willing to pay a lot of money for the available supply.
doctorpangloss 9 hours ago|||
it only takes ~445 GB300 NVL72 (about $22b) to run ALL of openrouter demand for a year. Microsoft rolled out $32b of DC 2026Q1.

imo the issue is that most openrouter demand is inauthentic activity (things that anthropic and openai models will refuse to do like pretend to not be bots when interacting with humans)

aurareturn 4 hours ago|||
I thought your numbers must be wrong.

So I plugged 288 trillion tokens/month (OpenRouter's current rate), 500 billion MoE model average, and the math comes out to be around 620 B200 GPUs minimum.

So basically, OpenRouter's volume must be absolutely tiny compared to the volume hyperscalers are getting.

HDBaseT 8 hours ago||||
It is worth mentioning, the OpenRouter demand isn't static though. It has increased week on week since early 2026.
senordevnyc 8 hours ago|||
I was curious so I looked it up: looks like a GB300 NVL72 is about $4M. So $22B would buy you 5500 such racks, no?
selectodude 4 hours ago||
GPU cost is about half of datacenter cost. Other half is cooling, power and networking.
petesergeant 2 hours ago||
I use Cerebras via OpenRouter. It’s every bit as fast and reliable for my needs as claimed. I suspect the reason is that they can either be making peanuts selling inference to plebs like me via OpenRouter, or making bank selling the more expensive models to businesses directly. In short: I would be very surprised if they have die capacity, and are at this point maximising revenue per chip.
aneryu 9 hours ago||
It would be even better if a version available to individual users were released soon.
gpm 9 hours ago||
They do offer API services to individual users... though with a set of models that makes it unlikely that you want to use it. They are promising Qwen 3.8 27B any day now though*.

if you have the money as an "individual user" to purchase one of their racks... save your money and retire.

* Actually they sent out an email claiming they already have it, but I don't seem to have access, they're promising to release it to the "shared tier" any day now.

drcode 7 hours ago|||
not only that, but I was so happy with their GLM 4.8 that they got rid of yesterday :(
apitman 7 hours ago||
What do they do with old ones? Their hardware physically can't run other models right?
airspresso 6 hours ago||
The Cerebras hardware is not locked to specific models / model families. Taalas is the company that's etching models into their silicon, locking it to that model forever.
fragmede 8 hours ago|||
> save your money and retire.

Now that this hypothetical person has retired, what are they gonna do all day? Just sit on the beach and drink Mai Tais? If that's what they wanna do, sure, but nerds gonna nerd, and if I had that kind of money to retire on, I'd totally buy some ridiculously expensive AI box for fun.

gpm 8 hours ago||
Ah but if you have the kind of money where this is a reasonable retirement hobby purchase, you aren't bothered by representing yourself as a "enterprise" :P
dnautics 7 hours ago|||
needs an sla that says power will never ever ever go out or else you will have a useless shattered plate of silicon.
gpm 6 hours ago||
Huh, why would it shatter if the power goes out?
ttul 9 hours ago|||
I’ll get that 250kW home power service dropped in next week!
z2 8 hours ago||
For now you can rig an adapter to your nearest DC EV charging station, but make sure it's near a body of water for the cooling.
johntash 5 hours ago||
I'd like to see a consumer version too, I don't need a whole rack of them. I probably can't even afford one gpu-sized one
anonymous_user9 10 hours ago|
Conspicuously missing: power consumption figures
wmf 10 hours ago||
162 kW
walrus01 9 hours ago|||
I guess we know why there's a fair bit of investment money going into small modular nuclear reactor startups now.
fragmede 8 hours ago|||
And advanced geothermal. Fervo Energy let's us get energy that's not based on burning fossil fuels but is, instead, able to produce energy from the ground.
ricardobeat 1 minute ago||
Remove energy from the ground. I wonder what the consequences may be once we are cooling the underground at several MW/h.
dakolli 3 hours ago|||
Actually that's mostly just the military funding those, with a few of them having data center partnerships so they can shield themselves from the criticism of what they really are: military contractors.
walrus01 3 hours ago||
I heard a great deal of noise from early 2002 to the present date that the US military had a high interest in small portable nuclear reactors for large bases in Iraq and Afghanistan. And particularly around the peak period of troops on the ground in AF and IQ. And indeed a place like Bagram or Kandahar used a shitton of diesel to run generators. But nothing ever came to fruition to actually implement it, has something changed now that they actually consider it worth doing?
roughly 9 hours ago||||
God, I was going to ask if this could be deployed in a standard existing datacenter, but I guess that answers that question.
KeplerBoy 6 hours ago||
So the answer is yes? Putting in a few of those racks for special tasks shouldn't break the power assumptions of a data center.
xattt 9 hours ago|||
I presume per rack?

Can you imagine something radiating that much energy into a space in your home?

walrus01 9 hours ago|||
It's mandatory liquid cooling, so it's meant to be attached to a specialized liquid cooling loop that gets the heat outside the building.

This is far beyond the practical maximums of like 10 to 15kW per 44U cabinet front to rear air cooling for 'regular' rackmount server stuff.

ttul 8 hours ago||
Indeed. You need 45 to 60 liters per second of cooling water flowing over a Cerebras wafer every minute to keep it under 90C. And that’s assuming the water leaves at 90C…

More realistically, you need much more cooling water.

wmf 9 hours ago||||
I guess because I have actually set foot in a data center I don't imagine literally every product in my home.
logicallee 9 hours ago||
"10x more throughput per watt than CS-3"
More comments...