Top
Best
New

Posted by sunils34 12 hours ago

Cerebras CS-4(www.cerebras.ai)
315 points | 207 commentspage 2
9cb14c1ec0 11 hours ago|
Just a reminder for everyone that we are only several years and 3 or 4 iterations into hardware being optimized for LLMs. We should all expect orders of magnitude improvement in speed and/or cost over the next 5 years. Then we can have fun conversations about "unlimited" "intelligence" and about what the price wars and profit margins of consumer AI products are when your average ChatGPT user costs the company $0.10 per month.

> CS-4 delivers more than 1,000 tokens per second on models exceeding 10 trillion parameters

Wow!

SwellJoe 11 hours ago||
And, the software side isn't finished being optimized, either. We've seen with Qwen 3.8 27B and DeepSeek V4 Flash 0731 and GLM 5.3 that quite small models can pack a punch. Intelligence density will improve, efficiency of kernels will improve, efficiency of KV caching and MTP will improve, algorithms for splitting workloads across compute units will improve.

It'll all be as cheap as DeepSeek was before the price hike. And, it'll become more and more realistic to run near-frontier intelligence on personal devices.

dgellow 4 hours ago|||
> Then we can have fun conversations about "unlimited" "intelligence" and about what the price wars and profit margins of consumer AI products are when your average ChatGPT user costs the company $0.10 per month.

We can have that discussion now: sounds like that would kill OpenAI and Anthropic

moralestapia 9 hours ago|||
Hence why taalas was one of the best strategic acquisitions of the year.

I'm honestly baffled they were not acquired by somebody else (sorry AMD).

adventured 8 hours ago||
Taalas will be one of the great disaster investments of the early AI era. It'll be a near total write-down.

The absolute worst market time to etch a model to a chip is right now (very rapid iteration). There is no scenario where they can keep up. The Taalas approach will be viewed as comically foolish within just a few years.

Cerebras will win in terms of approach.

It's 1998: hey, I can drastically speed up your web service, let's etch it right to silicon.

Iolaum 6 hours ago|||
I would still pay ~500 for a chip that runs 10kt/s of a ~100b model on my machine even if the half life is 6months. I m sure my employer would too.

Qwwen3.5 122b was released 6 months ago and is still best in class overall 100-140 B param model.

jeffybefffy519 6 hours ago|||
10000%
x-complexity 5 hours ago||||
> I would still pay ~500 for a chip that runs 10kt/s of a ~100b model on my machine even if the half life is 6months.

....

:T

Considering the 8B model uses 53 billion transistors, that's 6.625 transistors per parameter.

https://taalas.com/products/

Assuming they can get it down to 3 (somehow), that's still 300 transistors, or 5.565 RX 9070s.

https://www.techpowerup.com/gpu-specs/radeon-rx-9070.c4250

You're looking at

1) waiting for another 3-5 generations of transistor improvements before it can fit into a single conventional chip, or

2) another generation before getting a monster of a chip (1000+ mm^2), and prices for flawless etching scale quadraticly (likely $1000+ for manufacturing costs alone).

Could happen, but it's a long shot for a market that could be satiated by specialized accelerators.

mdp2021 2 hours ago||
In Taalas HC2 a chip embeds 20b parameters, and the declared idea is linking the chips. A card with two of them chips and you can already have a dense Qwen at staggering speeds.
dgellow 4 hours ago|||
500 what? You’re missing the unit
NitpickLawyer 7 hours ago||||
> The absolute worst market time to etch a model to a chip is right now

Slightly disagree. It really depends on the price-point at which they can do that etching. ~1k usd / ~30B model in a hdd-sized case that fits on your desk? I'd buy one right now, even knowing that I'm "stuck" with whatever model of the day is.

riknos314 7 hours ago||
Time to market also matters a ton. If they can start shipping chips <1 month after the weights drop that's much more compelling than if it's a 6+ month development pipeline.
mdp2021 2 hours ago||
In the case of Taalas, the pipeline was said to be 2 months:

> From the moment a previously unseen model is received, it can be realized in hardware in only two months ( https://taalas.com/the-path-to-ubiquitous-ai/ )

walrus01 2 hours ago||||
> It's 1998: hey, I can drastically speed up your web service, let's etch it right to silicon.

I distinctly remember 32-bit/33 MHz PCI accelerator cards for SSL being a real thing (for use on OpenBSD or FreeBSD), in an era when something like a single core 700 MHz Pentium 3 1U system was a relatively powerful individual bare metal httpd box.

http://www.aster.si/partnerji/compaq/atalla/axl200.html

The CPU load of doing a lot of SSL purely in software was a problem in terms of scaling things up, so this was one attempt at a (very short lived) solution. Note that this predated TLS1.0.

WithinReason 5 hours ago||||
The 500x efficiency gain makes their approach a no brainer. Just make a new chip every 6 months, you still win.
pimeys 7 hours ago||||
I just want to but hardware so I can run a model at home that is fast. I don't see myself installing a server that burns almost two hundred kilowatts but maybe a card which runs a 27B Qwen...
mdp2021 2 hours ago||
At 250w when it's working (I understand), and it works for tiny amounts of time per query...
alightsoul 7 hours ago||||
maybe AMD wants the IP to deploy it once ai model development slows down in a few years. Or, their large cloud customers do want to burn through silicon, basically paying rent to AMD for models etched on silicon.
moralestapia 8 hours ago||||
It's 2026: let's etch nginx into silicon and get 10,000,000 rps at a cost of 0.1 US/day.

Yes, please!

redox99 7 hours ago||
Most people probably don't care about nginx performance. It shouldn't be your bottleneck unless you serve massive amounts of static data.
mdp2021 2 hours ago|||
In the case of needs to process natural language, instead, massive efficiency (esp. time) can be a game changer. It's like "you have two years to complete the project" vs "you have two hours to complete the project": if you can squeeze that "two years worth" into a negligible delay, it's a game changer.
dyzone 5 hours ago|||
Ok, how about postgres?
conception 6 hours ago|||
Etched model into a chip? A… mobile chip eventually? Seems prescient.
api 11 hours ago|||
This is part of why I think the data center build-out is a bubble. We've barely scratched the surface when it comes to hardware optimization. We'll see exponential improvements in energy efficiency and speed over the next decade. Exponential, not linear.

GPUs really aren't that great for AI. They just happen to be the best chips we have in mass production right now for this work load, and it takes time to field new designs. Basically every chip engineer on the planet is working on this right now.

mindwok 11 hours ago|||
Whether it's a bubble or not depends on how much the demand for compute and the type of workload keeps growing, though.

If AI tends to be something used mainly in ideation and development, which is how a lot of people use it today, then once consumer hardware gets good enough you could see a bunch of the current data centre workloads move onto consumer devices.

But if AI starts being used more in repeatable, operational workloads I think it makes sense to have significant cloud infrastructure for it. TBH I haven't seen much of this, and I've been skeptical about people using agents for much of anything when it can be done with just software. But we are starting to see more of this kind of workload, like the taggable Claude in your slack etc that people seem to really love.

aurareturn 4 hours ago||||
By the way, this is the same argument that Michael Burry used to short Nvidia.

He claims that GPU depreciation/obsoletion is much faster than hyperscalers are assuming because new chips will be much better. He's being proved wrong right now because H200 rental prices have been claiming for the last 8 month despite B200 having 10-20x better inference efficiency.[0]

The logic is fundamentally flawed in my opinion. Let's use future Nvidia chips being much better optimized for LLMs for example.

New Nvidia chips 10x better than H200 --> data centers buy a lot --> Nvidia profits a lot.

New Nvidia chips 10x better than H200 --> data centers don't buy --> no faster than expected obsoletion.

In other words, the very act of buying many new Nvidia GPUs would be the event that causes faster than expected obsoletion. Yet, if you don't buy those new Nvidia GPUs, then there is no faster than expected obsoletion.

We also live in a world where there is competition. If Amazon doesn't buy but Microsoft does, suddenly Microsoft can offer better $/token prices.

[0]https://inferencex.semianalysis.com/inference

haldujai 3 hours ago||
1. The same isn’t necessarily true of the rest of the hardware stack which may be reused between accelerator generations.

2. You’re missing the “New Nvidia chips 10x B200, compute requirement grows less than 10*software improvements YoY -> buy less Nvidia.” Valuations are based on forward projections (>1T annual for NVDA) which can be revised down leading to a drop in valuation.

> If Amazon doesn't buy but Microsoft does

The big 3 all have their own proprietary accelerators. Meta is buying TPUs as well for now.

I would bet Nvidia’s major customers in 2 years are neoclouds and it seems that Jensen is making the same bet.

aurareturn 1 hour ago||
1. So this makes Burry’s argument even less convincing since those auxiliary hardware can last longer.

2. Jevons Paradox. More efficiency should lead to bigger models, faster inference, and more total tokens.

3. By all accounts, Trainium and Maia and Meta’s internal chip are struggling to keep up with Nvidia. That’s why they order as many Nvidia chips as possible. They’re not giving up but it isn’t as easy as buying stock Arm cores and taking them to TSMC.

Neoclouds may very well be Nvidia’s biggest customers and this probably what Nvidia wants.

petra 10 hours ago||||
I wonder: in world where inference is cheap, how many engineering agents that use simulation as their feedback we will use?

In the scenario, engineering everything becomes so easy - so why not optimize everything? every component, every product, every system?

And maybe llm's could invent. So even more to simulate. And simulation is inherently compute-heavy.

So unless there are some other bottlenecks, we'll use a lot of simulation servers.

RachelF 9 hours ago||||
True, I have to agree with you. The AI giants might be investing a huge amount of money in generation 1 technology. There might be a much better way to do it just around the corner. They might know this and thus the hurry to IPO.

A rough analogy would be if the first generation of ISP's spent billions on dial-up exchanges, when fibre could be invented next year.

winrid 11 hours ago||||
On the plus side, lots of cheap servers to swoop up :)
sroussey 10 hours ago|||
But power hungry.

In that 5+ year timeline, the compute per watt could change by three orders of magnitude.

GPUs are to LLMs what CPUs are to gaming — not a good fit.

amluto 9 hours ago||
A cursory estimate courtesy of ChatGPT suggests that there is a grand total of one order of magnitude or less of power efficiency improvement available compared to current Blackwell if the entire system’s power consumption outside the ALUs went all the way to zero.

If you want three orders of magnitude improvement, you probably need to find two of those orders of magnitude somewhere else: process improvements, different ALU design, model architecture changes, etc.

dgellow 4 hours ago|||
Look at their power supply, it’s not something you can run in a home lab. Unfortunately most of that will likely go to the bin eventually :(
jeffybefffy519 6 hours ago||||
Exactly right, and nVidia is protecting their moat through business practices rather than genuine product innovation.
__turbobrew__ 9 hours ago|||
By the time these gigawatt datacenters are done being built the hardware will be so far behind state of the art they may be mostly useless.
rvz 10 hours ago||
Congratulations! You have just realized that the AI data center build out is a total scam, built on both the insurmountable trillions of debt, and the assumption that only GPUs are all we need to continue scaling.

There exist other AI accelerators (TPUs, ASICs) that perfectly exceed the throughput that LLMs need to scale as well. But the true solution is more software optimizations. There's a tiny handful of them but more needs to be discovered so that we can reduce building hundreds of more data centers as the alternatives mature.

As better software becomes more useful for the alternative AI hardware for developers with LLMs running efficiently you then would have more choices of hardware to run your LLMs on rather than just only GPUs.

blovescoffee 10 hours ago|||
TPUs and ASICs run in data centers too. Your argument only holds true if there's some satisfied limit to demand for inference. If not, data centers will continue to spring up to host more and more agents. Even if agents were running on hardware and software as efficient as the human brain, its conceivable we want trillions of them running at any given time which would require data center scale.
georgeecollins 10 hours ago||
Everything has some satisfied limit to demand, often depending on the price. If you assume there will never be any satisfied limit to demand for inference at any price you can justify any investment.
x-complexity 5 hours ago|||
So far, at least by Openrouter's weekly numbers, there doesn't seem to be a satisfied limit.

https://openrouter.ai/rankings#top-models

And their market share sits at around 16-20%.

At 75.3 trillion tokens for the week ending 10 Aug 2026, that means that up to 450 trillion tokens were plausibly demanded by the whole market for that week.

My take: At max saturation, each person on earth could have their demands satiated by an average of 16 agents running concurrently. Sometimes more, often times less, but the average would likely be at 16.

At 200 tokens/second for each agent, that would mean 15.48288 quintillion tokens per week.

We're currently at about 0.00290643601% of the calculated demand ceiling.

Even if the demand limit per person is just 1 agent at 50 tokens/second, the current demand's still 0.186011905% of the theoretical ceiling.

aldonius 9 hours ago||||
Yeah, but there's certainly a part of the curve where price drops by X OOMs and demand increases by much more than X OOMs. (Presumably some of that is substitution and some of that is new use cases.)
adventured 8 hours ago|||
Looking back nearly 80 years, what has been the limit to transistor demand so far?

Unlimited.

What has been the limit to electricity demand globally?

Unlimited.

We can't get enough and never will. Costs have to become pretty severe to turn back the demand as well.

skyberrys 10 hours ago||||
I wonder what this looks like in 5 years... Will there be a massive push to repurpose these giant boxes into housing? Will they get turned back into the farm land from where they came? When a data center goes bust, what happens to the parts left behind?
gpm 10 hours ago||
I'd think the infrastructure would tend towards factories, smelters, and so on. Industrial things that have reasonably high power demands, can use the building, and don't care about the lack of windows.

They're typically not built where you want housing, and the buildings are distinctly the wrong shape.

If you can't use the power infrastructure profitably my next thought would be warehousing.

But also... we've seen a pretty continually increasing demand for compute. Even if AI busts a bit (or becomes a bit more efficient) I bet most data centres stay data centres, just less profitable ones.

jryle70 10 hours ago|||
Huh, why I'm not surprised that HN is full of opinions confidently stated without any numbers or resources to back up?

> built on both the insurmountable trillions of debt, and the assumption that only GPUs are all we need to continue scaling.

Insurmountable according to whom? And who assume that only GPUs are all we need to continue scaling? Google, Amazon, Microsoft, Meta and OpenAI, all have or plan custom non-GPU AI chips. Do they plan to use them not for scaling?

kobe_bryant 8 hours ago||
can these vibe coded sites please set a max width and overflow so their sites work fine on mobile
dgellow 4 hours ago|
Good news, future models will have your comment in their training set, making them slightly more likely to fix that problem!
denizay 8 hours ago||
The comparison seems incomplete. CS‑4 is a full rack-scale system with three wafer-scale processors, but the exact GPU models, GPU count, power consumption, price information are not disclosed. We still don't know if buying a multi-GPU rack (or racks) is cheaper and/or more efficient in power. The fact that they didn't disclose these numbers makes me believe that the numbers are not in their favor. And personally, makes me see them as disingenuous.
lostmsu 10 hours ago||
KV caching status?

What's the point of 1000tok/s if you have to do prefill on every agentic turn which at 100k depth would make it 1.5 min latency every turn?

walrus01 10 hours ago|
Information about RAM type/size and connection topology of the RAM to be used for context cache seems to be conspicuously absent from the slick looking marketing materials.
gpm 10 hours ago||
There's a few more details at the bottom of this page: https://www.cerebras.ai/blog/introducing-cerebras-cs-4

44GB on-chip-sram * 3 chips. Per chip: 43.2 PB/s memory access + 53.5 PB/s on-chip fabric bandwidth + 2.4 Tbits/s "IO" bandwidth (I think that means their RoCE v2 RDMA over Ethernet interface).

I suspect there might be a certain amount of customization for how much RAM they attach when you order it.

porridgeraisin 8 hours ago||
They have managed to make the external link 300GBps/2us. Cs3 was 150/5.

This is 1/3rd blackwells nvlink c2c bandwidth already. Not too bad. We can make KV cache offload work with that I suppose.

If magically KV cache was not an issue, pipeline parallelism on cerebras can be quite pleasant. As for the KV cache offload, I have hopes their CPO solution they're trying with that canadian company ends up bearing fruit.

sva_ 10 hours ago||
> enabling massive clusters and models with more than 50 trillion parameters
arthurcolle 6 hours ago||
What's the sticker price? If I have 20 million in the bank can I just like buy one or what
runeks 4 hours ago|
Surely it would depend on your model, since they need to etch this into silicon.
yvdriess 3 hours ago||
You're thinking of Taalas.
selimonder 4 hours ago||
That "GPU" comparison is the vaguest i seen so far
epolanski 4 hours ago|
True, it's also "per user", somehow, but I think it's a misleading metric. Cerebras chips take the whole wafer?

A single TSMC wafer contains 60 to 65 B200s, assuming 70% yields that's 40ish wafers per die.

Cerebras cannot redefine wafer economics.

tjoff 4 hours ago||
Depends on what you mean, they have more redundancy which means that the yield can be much higher.
4k0hz 11 hours ago||
> Introducing the all new Cerebras CS-4, a revolutionary rack-scale solution that delivers upto 30x faster inference compared to GPUs, enhanced economics, and a simple path todeploy [sic] hyperscale capacity.

Did nobody proofread this?

algoth1 10 hours ago||
If they had ask Claude it would probably look like this: Introducing the all new Cerebras CS-4, a revolutionary rack-scale solution that delivers up to 30x faster inference compared to GPUs, enhanced economics, and a simple path to load-bearing hyper scale capacity.
jm4 10 hours ago||
That's unusually honest and the sharpest thing in this thread.
wren6991 10 hours ago||
You're underselling it, and here's why.
SoMomentary 10 hours ago|||
Sometimes I wonder if mistakes are now used to indicate the possibility that a human actually wrote it.
ceejayoz 10 hours ago||
There’s been a spate of Reddit AI bots using all lower case in hopes of evading detection.

It’s still incredibly obvious.

VladVladikoff 9 hours ago||
I don’t really visit Reddit much these days but would love to see an example.
helloplanets 6 hours ago|||
Well at lest it's written by a human.
dgellow 4 hours ago||
« Make it look like human written »
geodel 11 hours ago|||
Maybe it is just part of their "compact design".
dpkirchner 10 hours ago||
An error no frontier LLM would make, eh
avantnyc 8 hours ago||
Cerebras should slowly also move to dgx/ryzen market for a desktop version for masses at affordable price yet providing substantial tokens/second on desktop
wmf 8 hours ago|
Desktop SRAM isn't really viable because it could cost $100K just to load the model.
aenis 6 hours ago||
...so, in the same ballapark as ddr5? :-)
tamimio 10 hours ago|
I wonder what are the benchmarks of hashcat on different hashes.
More comments...