Top
Best
New

Posted by bmulholland 17 hours ago

OpenAI Jalapeño: Better than Nvidia Blackwell(newsletter.semianalysis.com)
https://www.bloomberg.com/news/articles/2026-08-25/openai-cl..., https://archive.ph/yCTrr
417 points | 276 commentspage 3
bjourne 7 hours ago|
The article is a bit naive:

> However, as previously mentioned, Jalapeño’s results are obtained without speculative decoding and Vera Rubin’s results use speculative decoding. Speculative decoding leads to a ~3-5x reduction in cost per token. When speculative decoding is implemented on Jalapeño, this will enable Jalapeño to serve tokens even more cost effectively.

How much speculative decoding improves throughput is workload-dependent. Yes, it can improve performance by 5x, but it can also slow down performance by 2x. It also requires a vastly more complex decode loop than the standard token-by-token decode. The tech journos didn't ask themselves if speculative decoding improves perf so much why wasn't it on by default? Perhaps, because speculative decoding is not optimal for OpenAI's ASIC? Perhaps that is also why they were only able to benchmark the not-very-representative single-user-decode case?

7e 4 hours ago||
Once again we see the classic PR hype machine tactic of comparing a newer chip which is only available as an engineering sample to other chip designs which are widely available and much older.

They also fawn over the chip’s TDP when all other chips have to support 16 bit floating point and thus must run much hotter.

They make the classic mistake of equating max TDP with in-use-watts, and praise this magnificent (fictitious) performance per watt at FP8 with other chips’ max-TDP at FP16, which draw twice the power.

Evidence that the IPO can’t be far away.

mkw5053 9 hours ago||
Warning, this is a long comment! (I’m trying to stick to sourced facts here and not overstate what they mean)

I went down a rabbit hole after watching Dylan Patel on Dwarkesh today: https://www.youtube.com/watch?v=aV26V1UvkJw

I was initially just surprised by how bullish Dylan is on OpenAI/Anthropic and how bearish he is on China, despite Chinese labs getting closer to US SOTA while offering inference at dramatically lower prices.

So, I started digging while waiting for various day-job inference calls to return, ha.

Dylan says he spent years obsessively posting on hardware forums, moderating hardware subreddits, and running anonymous hardware blogs/videos before SemiAnalysis. But he also says most of that history is now gone, including from the Internet Archive, because he asked for it to be removed.[1]

In a 2024 interview he described his post-college job as “data science” around hurricane/earthquake/wildfire simulations for a financial company.[1] In a 2026 Sequoia interview he described himself as having been a “quant at a small quant risk firm” who generated $10M+ of “risk-free revenue.”[2] The Information reports that he declined to identify the employer and doesn’t list it on LinkedIn.[3]

Even harmless/silly stuff seems to drift. In February he said he kept bees for ~1.5 years. Today it was “few months, few months.”[4][5] I know, sort of silly and doesn't matter.

The Information reports that Patel owns stakes in ~20 startups in the same ecosystem SemiAnalysis covers, organized a $50M Fluidstack SPV, and is now targeting a $400M venture fund.[3][6]

And, in a 2022 HN discussion about SemiAnalysis disclosures, after saying his reports had moved smaller stocks by 20% in a day, Patel wrote: “If I thought I could move the stock, I'd make the position in the morning alongside my clients, and publish shortly after.”[7]

I don’t know that any of this is false or that anything improper happened (I’m definitely not claiming that). More that 1-2 of these things would just be odd. Taken together, though, they made me question how much trust I was putting in the broader story.

The dynamic of reminds me of crypto, WeWork, Theranos, Citron, etc. Once enough important people validate someone, things that would normally invite basic diligence somehow stop getting questioned.

[1] https://www.dwarkesh.com/p/dylan-jon

[2] https://sequoiacap.com/podcast/dylan-patel-of-semianalysis-w...

[3] https://www.theinformation.com/articles/dylan-patel-semianal...

[4] https://www.latent.space/p/dylanpatel-cooking

[5] https://www.dwarkesh.com/p/dylan-patel-3

[6] https://www.theinformation.com/briefings/exclusive-semianaly...

[7] https://news.ycombinator.com/item?id=31065646

m101 9 hours ago||
Not publicly acknowledging how misallocation of capital may be happening today shows he is corrupt. He’s not that dumb to not know it’s a major risk to the whole story, and is certainly financially incentivised to write as he does.
g00afthrowaway 1 hour ago|||
See my other comment about this person. I'm surprised people take this site seriously.
newyankee 5 hours ago|||
If you research the origins of Dwarkesh, even more conspiracy level points emerge. I do not think even in the handful videos post Leopold fund collapse he has addressed it in any ways. That is the point of so called observers, they pretend to be impartial but everyone is a hustler in some way.
mkw5053 9 hours ago||
I genuinely curious who’s downvoting me and why. I do not understand this forum sometimes.
dkhid 4 hours ago||
R@6.....111
dkhid 4 hours ago|
hacker
luciana1u 8 hours ago||
[flagged]
0xbadcafebee 12 hours ago||
Story says they're power limited. That's half-true. Actually they're water-limited. To generate power, you need water. To cool chips, you need water. If you try to use less water on one side, you need more water on the other side (it's physics ya'll, making and using energy generates heat which requires dissipation). The world's freshwater is diminishing while also being consumed at an alarming rate. The future AI oligarchs are whoever controls the most water.

The other side of the conversation is the idea that large models in DCs on custom silicon is the future. Maybe for enterprise? But consumers will eventually (10 yrs) have affordable hardware designed to run crazy-good local models (more RAM + higher bandwidth). That will take pressure off of datacenters, but also reduce AI profits, and move that money to consumer chip/device makers. Apple is once again the biggest winner. Nvidia consumer chips might get cheaper, but nerfed, to encourage datacenter use where they make more money. I'm hoping AMD can stop being terrible at software so that when we finally have their better hardware we can actually use it.

minimaltom 11 hours ago||
For datacenters specifically I've never understood what specifically consumes the water. Arent the water-cooling loops closed, so the water just cycles around and around and around?
SirMaster 11 hours ago|||
They evaporate the water which is what makes it cool so effeciently.
Ductapemaster 11 hours ago||
Evaporative cooling does not necessitate an open loop system
Eisenstein 9 hours ago||
The system which runs coolant over the chips can be closed but the part which uses an evaporative system to cool that is still open loop and vents water into the air, no?
Ductapemaster 6 hours ago||
Nope, it doesn't have to be open loop!

Example of such a system being used specifically for datacenters: https://blog.vantage-dc.com/2026/04/22/cooling-without-the-d...

Evaporated water is condensed, and in the process transfers its heat into another place that removes it. Another simple example is a pot of boiling water with a lid on it.

Eisenstein 5 hours ago||
The link you cited is not evaporative cooling and a pot of boiling water with a sealed lid on it is a pressure vessel which eventually explodes.
Ductapemaster 5 hours ago||
If you were to remove the heat at a sufficient rate by, say, turning the lid into a heat exchanger, you would have a stable system.
Eisenstein 5 hours ago||
How is that different from not using evaporative cooling and just putting the heat exchanger on the burner?
justincormack 11 hours ago||||
Yes they are for water cooling.
0xbadcafebee 5 hours ago|||
At the datacenter side, it depends on the method of cooling. You can chill the air or the chips directly (or both), doesn't matter, you still need to cool, and that still needs water. The question is, where is the water being used?

- If they use either evaporative cooling or a liquid-cooled heat exchanger, that uses tons of water consistently. This requires less energy (it's mostly passive) so you use more water.

- If they use closed-loop water cooling and/or heat pumps/electric chillers, that uses much less water - at the DC. But it does require more energy to circulate the water, run fans, etc. If you are using more energy, where is the energy coming from? It's coming from power plants, which require... you guessed it... more water (e.g. thermoelectric, hydroelectric, geothermal, concentrated solar). They need water in order to generate the power, and lots of it. Coal, natural gas, nuclear, and concentrated solar, all use steam to generate energy. Nuclear also uses water to cool the reactor. And water is used extensively to extract coal, oil, and natural gas. Geothermal uses water in the ground.

You can't not use a ton of water in one fashion or another. It just depends what method, and on what end the water is used. And the crazy thing is, most new datacenters are being built in places with extremely little water. Guess how that's gonna work out as the planet gets hotter?

I don't know why I got downvoted to hell for stating facts every datacenter architect knows. HN be HN'in.

WarmWash 11 hours ago||
This only makes sense if you never looked at comparative water usage rates and available water.
simianwords 12 hours ago||
How can OpenAI mass produce this chip at scale more economically than Nvidia which has experience in the supply chain and scale efficiencies to do it efficiently?
dpe82 12 hours ago||
NVidia has enormous operating margins, so a competitive solution doesn't have to match or beat NVidia's scale efficiencies; it just has to beat delivered cost.

One objective of the project might be simply to provide credible negotiating leverage when dealing with existing suppliers like NVidia. You don't have to deploy at scale for that to work, but you do have to look like you could if pushed hard enough.

vntok 11 hours ago||
> NVidia has enormous operating margins, so a competitive solution doesn't have to match or beat NVidia's scale efficiencies; it just has to beat delivered cost.

But then that means you have no actual moat against the behemot, right? Your competitor can move into the market as soon as they want to, at much better cost (so at slightly better price)... and Nvidia certainly can adapt much faster around hard hardware specs innovation than a new entrant ever could.

dpe82 8 hours ago|||
Those are not OpenAI's concerns - they just need to scare NVidia enough to lower their prices more than they'd otherwise want.
acedTrex 9 hours ago|||
But they are also PURCHASING from nvidia so any time nvidia lowers their prices they save money.
chris_money202 12 hours ago|||
In the short and medium term, it probably won't be more economical to produce for OpenAI. Where OpenAI is benefitting from their own chip is being able to tailor it to their models and workloads. When you buy off the shelf Nvidia, its not perfectly tailored and OpenAI has to spend marginally more to run off that chip. At the scale OpenAI is operating at and plans to operate at, that margin becomes pretty big $$
aurareturn 6 hours ago|||
Because it's actually Broadcom that is doing most of the work.
airspresso 12 hours ago|||
By leveraging the experience Broadcom has in this area. Still remains to be seen how that goes when they want to scale production.
toasterlovin 12 hours ago|||
Replace OpenAI with Apple and Nvidia with Intel.
VirusNewbie 6 hours ago||
"How can OpenAI produce a LLM at scale more economically than Google, Amazon, and Microsoft which have experience in planet scale software and scale efficiencies unlike them".

One answer is they're quite good at poaching talent.

Alien1Being 10 hours ago||
WARNING AI HYPE
varispeed 12 hours ago||
Why they don't research how to make their own RAM and they have to buy it from the common market?

They should GTFO with this crap.

Create barriers to computing for ordinary people while milking businesses for tokens.

petcat 12 hours ago||
Building a custom-designed ASIC is much easier than producing state of the art memory chips.

There's a reason why Micron and Nvidia are the crown jewels of American technology right now and for the foreseeable future.

chris_money202 11 hours ago|||
Nvidia buys the memory it uses on its GPUs, same as all other ASICs.

To give some context, Intel started making DRAM, I think they were actually the company that came up with modern memory techniques. They exited the market and pursued a more lucrative moat with CPUs.

JV00 12 hours ago||||
Nvidia does not make RAM
varispeed 11 hours ago||||
That doesn't excuse them from wrecking the market for ordinary person.
brcmthrowaway 12 hours ago|||
NVIDIA produces memory?
fc417fc802 12 hours ago|||
Fabless AFAIK. And that's the actual problem - drawing up CAD diagrams doesn't help if the factories are fully booked out.
Cyph0n 12 hours ago|||
A state of the art GPU is much harder to design & produce at scale and than an internal ASIC.
datakan 12 hours ago||
People keep saying stuff like this without understanding what it takes to make RAM. It's one of, if not the most, heavily patented things in the world. The second you dip your toes into those waters the lawsuits begin.

If somehow you get around the patent issues, you're now faced with huge research and development costs, fabs to build, processes to sort out and all of that has very high failure rates.

Last time I checked Micron was the largest patent holder in the world and even for them this is a hard area where they are number 3 in the market.

chris_money202 12 hours ago|||
RAM chips are not hard to produce compared to many other types of semiconductors; Intel started in the memory game and left because the margins weren't great and they were going to fold. The failure rates on these chips are actually very tolerable; you can have a very bad yield and still have a viable chip due to things like ECC.
datakan 11 hours ago||
Intel entered the memory space because they partnered with Micron. They left the memory space when Micron pulled out of the partnership.
chris_money202 11 hours ago||
Intel started making DRAM in 1970, Micron was founded in 1978.
varispeed 11 hours ago|||
Yes, it is difficult, but shafting working class is easy, therefor it is okay.

If the rich decided to buy all drinking water, you would probably be saying that's okay, making water is difficult, shortly before dying.

LarsDu88 12 hours ago|
Well Sam Altman finally has built a moat against Chinese open weight AI. Well done. But what will this mean for Cerebras?

I remember when Tesla was building its own inference chips, and after about 2 years and billions spent, the whole effort was scuttled b/c they simply could not keep up with the iteration and R&D cycles of dedicated chip companies. I suspect the same will be the case with OpenAI vs Cerebras + Nvidia/Groq

Eridrus 11 hours ago||
Cerebras is targeting a distinctly different point on the cost/latency curve. They are betting that there will be some high value applications where latency and not just throughput is super important.
porridgeraisin 10 hours ago||
It is being used as part of a combined system. For example AWS is pushing for Trainium + WSE 3. The WSE 3 does the decode and the Trainium does the prefill.

Even in nvidia land rubin + LPU does a similar thing.

It has its downsides of course - if your traffic swings prefill heavy to decode heavy, you can't suddenly use your lpu for prefill. With GPUs they're totally interchangeable. Tradeoffs.

Eridrus 8 hours ago||
AFAIK You can use WSE/LPU for prefill, it's just less efficient to do so.
KaiserPro 11 hours ago|||
> Well Sam Altman finally has built a moat against Chinese open weight AI

Hes got a press release.

The issue is, baking something to silicon requires discipline and about 2 years.

This isn't something you can just change your mind on halfway through. Trust me, I know. You need a clear vision of what you want to support, why and what bits of a chip you need to achieve that.

SV_BubbleTime 8 hours ago||
And yet, the top comment is about “hardcoding” weights into the silicon.

Man, if only someone made like, chips that could lots of different calculations all at the same time!

segmondy 9 hours ago|||
I think the Chinese are going to be building their own chips aided with AI. DeepSeek, z.AI, MiniMax, Moonshot, etc, it's a race. The take off has really started.
epolanski 11 hours ago|||
> and after about 2 years and billions spent, the whole effort was scuttled b/c they simply could not keep up with the iteration and R&D cycles of dedicated chip companies

That sounds quite like...nonsense?

Chip companies work on years-long cycles. They know today what are they launching 4-5 years from now.

brcmthrowaway 12 hours ago||
[dead]