Top
Best
New

Posted by bmulholland 16 hours ago

OpenAI Jalapeño: Better than Nvidia Blackwell(newsletter.semianalysis.com)
https://www.bloomberg.com/news/articles/2026-08-25/openai-cl..., https://archive.ph/yCTrr
400 points | 263 commentspage 2
m4rtink 7 hours ago|
So this will make GPUs and associated affordable for people, rigjt ?
jpollock 7 hours ago|
No, it's the wafer starts that are driving prices. Switching from Nvidia to custom doesn't change the constraints.
throwaw12 11 hours ago||
Competition is good for all of us, we will get better and faster chips.

Or at least Nvidia GPUs will become slightly cheaper for regular consumers again

WarmWash 10 hours ago||
That's if any datacenters are allowed to be built with them.

There is probably a ~50% chance that the next Dem candidate for presidency runs on a national datacenter moratorium or something equally as crippling.

bigyabai 5 hours ago||
If the populist campaign is to Make Affordable DRAM Again, then it's not a terrible solution.

The current datacenter owners love a compute-bound world anyhow. A moratorium on new datacenters would increase their valuation, encourage efficiency and make computers cheap again. If Chinese labs can ship frontier models under 1T parameters, why not American labs too?

zzzoom 3 hours ago|||
No way to avoid the memory cartel, even if CXMT catches up.
porridgeraisin 9 hours ago|||
These are not replacing GPUs, they are entirely complementary. It's the same with cerebras, groq etc, they are all complementary to the GPU.
einpoklum 9 hours ago|||
If you think tanking Trillions in investments, warming the earth and increasion ocean water levels, creating water shortages and brown-outs is "good for all of us" - well, the rest of us beg to differ.
throwaw12 9 hours ago||
these GPUs make computation faster, I understand as of now maybe all the computation is used to generate yet another junk LinkedIn post or unnecessary RFC, but at some point this craze should settle and we will be left with powerful computation machines, which can be used for computing more useful things
cmrdporcupine 3 hours ago||
The GPUs being paid for w/ billions in investment will be obsolete and e-waste in a few short years same as a Cray-2 was just a decade after its release.

It's fine if you're one of the people selling shovels to gold miners for a while, but sucks to be building houses in the boom town?

theandrewbailey 11 hours ago||
The pricing of GPUs themselves aren't really the problem: it's the VRAM that comes with them.
a2ff6eeb0 8 hours ago||
Sounds like a great way to get deals out of Nvidia.
chabons 1 hour ago|
Sure, but even heavily discounted Nvidia chips won’t be competitive for inference if they’re worse on perf/W.
calldacopsidgaf 7 hours ago||
Any article that features Sam's fucking creepy face should be marked with a jumpscare warning
danielovichdk 10 hours ago||
I guess special hardware is the new moat in AI.

Maybe the money will still flow into this industry after all

einpoklum 9 hours ago||
I hope the LLM wave will leave GPUs behind to go back to pursue more general-purpose computation rather than spending their die area on multiplying 4-bit-number matrices and such things.
acedTrex 8 hours ago|
Is that not literally the exect opposite of the direction asics for LLM inference is going?
arrty88 6 hours ago||
Is this bad news for Cerebras?
empath75 11 hours ago||
When people talk about the commodification of inferencing, they imagine a future where everyone has access to frontier models and can run them at the same cost, and what will actually happen is closer to the commodification of _oil_, where only a few companies have the scale to produce it at a competitive price, and advances like this are _why_.

Once models are more or less interchangeable, the price of LLMs will drop to essentially the price of energy required to run them, and the big labs will be able to run them cheaper than anyone else.

impossiblefork 9 hours ago||
I don't agree. At the moment companies like NVIDIA take several times what it costs to make a chip. I think the fair split for the technology contribution is more like 50-50, maybe even 30-70 in favour of the manufacturer.

With competition we will actually have the fair split, whatever that is, and thus much lower prices.

At the moment, to have a big AI firm, or really AI firm at all, you need to be blessed by NVIDIA, in the form of receiving circular financing for your compute. They know that their prices aren't fair, or competitive.

Commoditization of inference is the end of that. The end of the mega-premium on inference hardware, and it's good not only for people who like running their LLMs, but it's the first step towards commoditization of training.

kubb 9 hours ago||
10-90 is the fair split.
simianwords 11 hours ago||
I don't believe models will be commodified because each model is unique with strengths and weaknesses. Its not like Steel which is more or less the same no matter where you purchase it from.

If what you said were true, you would hardly see people complaining about the quality of Opus 5 or good writing from Sol. But people do.

lelanthran 11 hours ago|||
> I don't believe models will be commodified because each model is unique with strengths and weaknesses.

They are all converging.

zurfer 10 hours ago||||
The same level of intelligence gets roughly 10x cheaper per year. So you might both be correct where a large part are commodity tasks but frontier is hard and valuable and not commodities.
skhameneh 8 hours ago||||
> Its not like Steel which is more or less the same no matter where you purchase it from.

I’m not an expert in metallurgy by any means, but this seems really off. There are many recipes for steel and varied processes that also impact the final product.

airspresso 11 hours ago|||
This depends heavily on what the use-case is. Yes, if it's a coder making software and having to read LLM output then writing style matters. If the LLM is used in an automated data processing pipeline with a capped level of complexity, entirely different aspects matter and LLMs become more interchangeable.
bjourne 6 hours ago||
The article is a bit naive:

> However, as previously mentioned, Jalapeño’s results are obtained without speculative decoding and Vera Rubin’s results use speculative decoding. Speculative decoding leads to a ~3-5x reduction in cost per token. When speculative decoding is implemented on Jalapeño, this will enable Jalapeño to serve tokens even more cost effectively.

How much speculative decoding improves throughput is workload-dependent. Yes, it can improve performance by 5x, but it can also slow down performance by 2x. It also requires a vastly more complex decode loop than the standard token-by-token decode. The tech journos didn't ask themselves if speculative decoding improves perf so much why wasn't it on by default? Perhaps, because speculative decoding is not optimal for OpenAI's ASIC? Perhaps that is also why they were only able to benchmark the not-very-representative single-user-decode case?

mkw5053 8 hours ago|
Warning, this is a long comment! (I’m trying to stick to sourced facts here and not overstate what they mean)

I went down a rabbit hole after watching Dylan Patel on Dwarkesh today: https://www.youtube.com/watch?v=aV26V1UvkJw

I was initially just surprised by how bullish Dylan is on OpenAI/Anthropic and how bearish he is on China, despite Chinese labs getting closer to US SOTA while offering inference at dramatically lower prices.

So, I started digging while waiting for various day-job inference calls to return, ha.

Dylan says he spent years obsessively posting on hardware forums, moderating hardware subreddits, and running anonymous hardware blogs/videos before SemiAnalysis. But he also says most of that history is now gone, including from the Internet Archive, because he asked for it to be removed.[1]

In a 2024 interview he described his post-college job as “data science” around hurricane/earthquake/wildfire simulations for a financial company.[1] In a 2026 Sequoia interview he described himself as having been a “quant at a small quant risk firm” who generated $10M+ of “risk-free revenue.”[2] The Information reports that he declined to identify the employer and doesn’t list it on LinkedIn.[3]

Even harmless/silly stuff seems to drift. In February he said he kept bees for ~1.5 years. Today it was “few months, few months.”[4][5] I know, sort of silly and doesn't matter.

The Information reports that Patel owns stakes in ~20 startups in the same ecosystem SemiAnalysis covers, organized a $50M Fluidstack SPV, and is now targeting a $400M venture fund.[3][6]

And, in a 2022 HN discussion about SemiAnalysis disclosures, after saying his reports had moved smaller stocks by 20% in a day, Patel wrote: “If I thought I could move the stock, I'd make the position in the morning alongside my clients, and publish shortly after.”[7]

I don’t know that any of this is false or that anything improper happened (I’m definitely not claiming that). More that 1-2 of these things would just be odd. Taken together, though, they made me question how much trust I was putting in the broader story.

The dynamic of reminds me of crypto, WeWork, Theranos, Citron, etc. Once enough important people validate someone, things that would normally invite basic diligence somehow stop getting questioned.

[1] https://www.dwarkesh.com/p/dylan-jon

[2] https://sequoiacap.com/podcast/dylan-patel-of-semianalysis-w...

[3] https://www.theinformation.com/articles/dylan-patel-semianal...

[4] https://www.latent.space/p/dylanpatel-cooking

[5] https://www.dwarkesh.com/p/dylan-patel-3

[6] https://www.theinformation.com/briefings/exclusive-semianaly...

[7] https://news.ycombinator.com/item?id=31065646

m101 8 hours ago||
Not publicly acknowledging how misallocation of capital may be happening today shows he is corrupt. He’s not that dumb to not know it’s a major risk to the whole story, and is certainly financially incentivised to write as he does.
g00afthrowaway 34 minutes ago|||
See my other comment about this person. I'm surprised people take this site seriously.
newyankee 4 hours ago|||
If you research the origins of Dwarkesh, even more conspiracy level points emerge. I do not think even in the handful videos post Leopold fund collapse he has addressed it in any ways. That is the point of so called observers, they pretend to be impartial but everyone is a hustler in some way.
mkw5053 8 hours ago||
I genuinely curious who’s downvoting me and why. I do not understand this forum sometimes.
More comments...