Posted by inigyou 20 hours ago
> As discussed previously, the ramp of HBM production will constrain industry supply growth in non-HBM products. Industrywide, HBM3E consumes approximately three times the wafer supply as D5 to produce a given number of bits in the same technology node. With increased performance and packaging complexity, across the industry, we expect this trade ratio for HBM4 to be even higher than the trade ratio for HBM3E. We anticipate strong HBM demand due to AI, combined with increasing silicon intensity of the HBM roadmap, to contribute to tight supply conditions for DRAM across all end markets. As the memory industry is still recovering from the challenging environment in 2023, this tight supply environment will help drive the considerable improvements in profitability and ROI (return on investment) that are needed to enable the investments required to support future growth.
https://investors.micron.com/static-files/4550f98c-1054-4847...
The problem is that they need actual chips tomorrow. I fear we are reaching the point in the semiconductor industry where, in order to sustain the revenue growth propping up their valuations, they're going to have to start selling future chips that cannot possibly be physically produced.
If there is demand for N chips at $X price, but you only have N/2 chips, then half the people aren't going to be able to buy them. The people who really need the chips, and would be willing to pay a lot more than $X for them, will be competing with people who are only willing to pay $X for them. You will end up getting scalping and shortages and hoarding. Since people know that there are enough people willing to pay a higher price, everyone will try to buy them at $X, even if they don't need them at all. Just buy them at $X, and immediately sell it to one of those companies willing to pay a lot more.
The market is extremely inefficient... since the manufacturer isn't charging enough, you get way too many people trying to buy them from the manufacturer.
So you find the price where the demand is N/2 chips, and everyone who is willing to pay that much will get one, and there is no profit for scalpers so the only people who will buy them will be actual companies that need them.
I don’t believe scalpers are an issue in B2B; these aren’t concert tickets being sold to the general public.
It’s clear you grasp the economic theory as it’s taught, but my comment was meant for you to question it. If someone is hungry you can charge more for food. The market allows and encourages it. But don’t mistake that for “willingness” and don’t mistake raising the prices purely to get extra money from the exchange as some inevitable law of the universe. You’re welcome to love the concept, but don’t whitewash it.
Sometimes this works and sometimes it does not.
If you happily sell apples for $1 and see a hungry person walking towards you, must you raise the price?
I think that basic thought gets lost sometimes when people talk about shortages causing the price to increase. The shortage didn’t cause anything, some executive decided they want more money. That’s all. There’s nothing inherent in the system that requires it. Whether that’s OK or not is up to the reader, I just think people lose sight of the reality and talk about it like it’s gravity, rather than simple decision making.
I’m confused because this reads like a denial of basic economics. If the price is high enough, the chip will be produced for you.
Do you mean because of the production lead time, higher prices won’t result in increased production? Commodities like corn have been managing this for a long time… what’s special about chips?
What is your actual argument?
Samsung, SK Hynix, and Micron have a combined market share of 90% - with most of the rest being a very new-to-the-market CXMT. It is a cutthroat market which behaves like a stereotypical "pork cycle". Semiconductor fabs cost billions to build and take years to complete, so you better be damn sure you have buyers before you start constructing one. You and your competition overestimated the demand? You have to pay back the construction cost, so you're now in a race to the bottom and one of you is going bankrupt.
Ever wondered where Intel came from? They started out as a DRAM manufacturer, which dominated their revenue well after the introduction of their first microprocessors. But in the early 1980s the glut of supply from new Japanese manufacturers made it so unprofitable that they had to ditch the memory market altogether. The stories of Texas Instruments and Motorola aren't much different. And that's not even mentioning the likes of Mostek, which once held a 85% market share and was dead less than 5 years later! Oh, and those Japanese manufacturers? All gone, pivoted like Intel or died like Mostek.
So no, the three remaining DRAM manufacturers aren't going behave like headless chickens and start ordering new fabs just because there's a bubble causing a temporary demand peak. Unless those AI companies are going to pay in advance, in cash, for an entire fab, they'll just have to wait and deal with the price increase.
So it's more "what's special about corn". It is also fairly hilarious to claim the parent is denying basic economics and then bring up corn as an example of having successfully managed economics. If the scales were not being thumbed, and "basic economics" were in play, corn would be in very very bad shape.
In the case of DRAM, there is an incredibly long history of these gloom/glut cycles, and they have stayed roughly the same timeframes (~3 years) since the 1990's.
Almost all the ones who have survived this long are either in the same kind of boat as corn - protected in various forms from the downside - or don't increase production and get caught out until they are absoultely forced.
The very temporarily increased profit is not worth going bankrupt for - they make more money long term by being very cautious and know this.
There are a near infinite number of economic studies you could look at (and several sibling comments cite some) - DRAM manufactuers don't chase the price and probably couldn't anymore if they want to.
None of this denies basic economic theory, of course, since economic theory is not exactly "rigorous", even to the degree it could be (IE even the parts that are pure analysis of data rarely reproduce!).
Memory is just too useful now that you can use it to drive cars and write code.
Which directly leads to the next big development: all the big players are investing in silicon with "baked-in" models, like [0,1]. Turns out you don't need an expensive general-purpose GPU with heaps of RAM to contain a model when you can make a custom ASIC around one specific model! Why spend a fortune on DRAM / HBM when all you need is some finetuning parameters which are easily stored in on-die SRAM?
[0]: https://www.theregister.com/systems/2026/08/06/amd-acquires-...
[1]: https://thenextweb.com/news/google-frozen-chip-gemini-silico...
In any other business choosing (2) would mean someone else swoops in and steals all your business. It doesn't look like this is at all possible for memory fabs.
"If the price is high enough" is of course technically true, but the scale of what high means in this context is important.
> Now, you could raise prices to the point where you destroy demand.
You know supply and demand is like a curve right, you can find an optimal equilibrium? It’s not a cliff that you can fall off.
Maybe you could pay someone else who is less desperate to part with some, but that does not increase supply.
Nine women cannot produce a baby in a month, even though you could average about one per month if you wait about nine months.
That's really interesting, and I wonder why? I believe HBM has redundant ECC bits by default, which would add a few %, but other than that, is it just that the yield is much lower due to die stacking? Of course, this is a 2024 document so things may have changed a bit since.
In order to get high bandwidth, you want memory as close as possible to the GPU. The more trace lane length = signal loss, bandwidth loss.. HBM is compact, and so you can stack 24GB modules, 8 around a GPU die.
If you tried to do that with normal memory, you need like 64 modules. So a a TON of traces more that all need to be equal length, and because so many = far away from the GPU = less bandwidth.
The issue is like stated above, its a process that waste a ton of wafers. Wafers that can make easily 3x more normal memory.
Intel with "Crescent Island" is trying to make a 160GB card using LPDDR5x memory but the bandwidth is only ~700GB/s.
The base interface die has eight or 12 memory dies stacked on top of it. The wafer cost of that interface die is therefore small.
If the process of die thinning and TSV stacking reduced yields by a factor of three, nobody would consider HBM mature enough to put into production, and especially not mature enough to be increasing stack height from one generation to the next.
Trace length has approximately nothing to do with die size. If anything, designing for shorter traces means you can get away with smaller PHYs at either end.
What might go some ways toward explaining such a huge difference in die size is that the TSVs themselves take up significant die area and must be fairly numerous to carry both a large number of signal wires and all the power and ground required by the stack. But it's wildly implausible that the die area consumed by the TSVs would be significantly larger than the die area consumed by the memory arrays themselves, or that anyone would build a memory die where the memory array was not a large majority of the total die area.
If there's any truth to that ~3x higher wafer requirement for the same number of bits as compared to DDR5, it must be a combination of several factors and probably includes something non-obvious and dubious, like counting all the area of the passive interposers that go between HBM stacks and GPUs (those interposers aren't competing for the same fab space).
What the freaking hell.
GPU is mostly fine and can probably reuse that but it’s Vega so not the greatest I was due for an upgrade but the prices of everything is ridiculous.
The thought of pc-of-Theseus’ing my busted ass PC is just sad.
Thank god I bought a PS5 Pro before Sony started jacking up the prices.
I remember when buying a $300 GPU was extravagant and gave you state of the art graphics.
Pick any B550 with a VRM heatsink for $90, a 5600 for $190, and you are back in business.
Get a 5070 or 9070 before GDDR supply is completely gone if you want an actual upgrade.
As for GPU price, that Vega 64 was $500 in 2018. Even a midrange GPU hasn't been $300 since the Radeon RX 480 launched a decade ago.
The PC I replaced when I bought this one was such a huge improvement.
My options basically boils down to fixing trash that’s falling apart or a sidegrade at best for the same price as what I bought a beast of a PC for 10 freaking years ago. Depressing.
Point is, quite a lot of hardware can still do a lot. I have a Ryzen 5500U-based laptop that still outperforms the PCs of most people that I know (non-tech people).
Though in my case I feel like I have to get a loan to buy a workstation + an actual gaming PC for my wife. That all is beginning to look like at least $15k... yikes. But I am not sure I can wait until 2028 to see if things _start_ normalizing... Depressing indeed.
My solution is to see what four or five-year-old equipment I can buy that will let me run local LLMs. I may only get six or seven tokens per second out of an i7, but it's a start. And best of all, I can turn the machine off when I'm not using it.
IMO, Migrating to small-scale local LLMs would be a significant improvement over using data centers.
LLM serving is most efficient when you batch a lot of parallel requests together. Data center solutions also have the advantage of collecting queries from around the globe, so the hardware can be utilized around the clock.
Having everyone serve their own local LLMs would produce a lot more memory demand. Not less. The same memory would be idle most of the time, and when it was used it would be used for 1 person instead of a batch of requests.
There are other reasons to run local LLMs, but solving hardware demand problems is not one of them.
This shift is probably inevitable, but it will vary significantly by region depending on prices for electricity. Look at, for example, the difference in fundamental homelab build recommendations between Germans and just about anyone else. Electricity prices in Germany are so high that even a now expensive Raspberry Pi or other ARM board is often preferred over Intel/AMD builds due to low power draw (especially low idle power draw), an effect that adds up for a machine running all the time over years.
With local LLMs and the GPUs to run it, especially if you want a model available to you all the time and can remote into your local network to use it whenever you want, there's no escaping much higher power draws, even at idle. Wherever electricity is expensive, the electric bill can be a prohibitive barrier.
Eg: I shaved ~40W off the idle load on a server (250->210W) by doung nothing more than removing the redundant supply
A college buddy used to work at Motorola (I'm naming the company because they wouldn't mind this story being shared) back in the late 90's or early 2000's. They had redundant power to their campus, bought from two different companies, coming in on opposite sides of the campus, so that even if some backhoe operator cut a ground-based power line somewhere, they wouldn't lose power.
And yet, one morning, the power went off all across their campus. After a little investigation, they sent pretty much all their employees home at noon and told them "take the afternoon off, don't come back until tomorrow, you wouldn't be able to do any work anyway". Turns out that although the power lines came in at opposite sides of their campus, somewhere a few miles away both of the power lines feeding their campus ended up running through the same underground conduit. And yes, a backhoe had managed to cut that conduit and break both of the lines they depended on at the same time. They had a single, VERY non-obvious, point of failure, and the backhoe had unerringly homed in on that SPoF.
I'll second this. The combination of Whisper + LLM makes speech recognition fantastic. I occasionally have arm pain from typing, and this is a Godsend.
I don't use it to write code - but in my experience stuff like emails + docs was the greater source of pain (one generally types slower while coding).
The average consumer isn't purchasing a lot of products with a lot of RAM every year.
Their 8GB of RAM phones will go up in price a little bit, but people aren't buying phones every year or even every other year.
So if the price of the 8GB of LPDDR went from $40 to $160 and it's all passed on to the consumer buying a new phone every 4 years, that's an extra $30/year in spending.
Most adults I know now use their phone for everything, so they're not buying new computers and laptops. They can wait 3 years for the market to settle before upgrading those, too.
Even I'm a heavy buyer of tech products and RAM, and I would bet you that my family's annual food bill fluctuates by more than what I've had to pay for RAM prices growing.
Many smaller companies will have issues getting enough, but even they could buy something with memory and rip it out. When you're willing to pay a hefty multiple, you can outbid other people. The shortages are not nearly bad enough to escape standard supply and demand.
>Most adults I know now use their phone for everything
You mean the cloud for everything. A lot of phone tasks share their computational and storage workloads off in the invisible ether that has actual computers with real costs behind them. Amazon has no problem with bumping up their costs in order to pay for their fleet of servers. I mean, what business doesn't use the cloud these days.
In the shorter term, you’ll probably see more “base” models on offer with some of the most intensive features disabled. There’s still going to be all of the telemetry stuff that makes them money through third parties though.
Isn’t it just… prices?
Total imports are maybe 14%? But not as highly tariffed.
Your computer already supports that. It's called "swap".
CPUs have L1, L2, L3+ caches too after all.
USB would be the limiter for both.
The M10 are over produced and only PCI-E 3.0 2x lane devices but 16GB of DDR5 is going to cost over $200.
You can do the same with USB-C M.2 adapters but the adapters cost more and eww USB for Swap, hope the device never drops.
https://www.amazon.com/GLOTRENDS-Adapter-Aluminum-Heatsink-P...
In any case DDR3 isn’t nearly as bad price-wise and would have much better energy efficiency than stacking 1–2 gig sticks.
For starters, The slowest sticks of DDR(x) are often slower than DDR(x-1). The issue is never capacity, but rather, performance.
The real nonsense? USB? It's a mess. Pick a USB cable and buy it from ANY retailer, let's make it easier, buy a USB-C cable. What is the data rate (depends on cable quality and length), Does it support power delivery? If so, how many watts? (~5W requires a very different cable from ~230W), how do you know from simply looking at the cable? If you buy a cable, how can you tell what it supports by simply looking at the connector? Imagine having a box full of USB-C cables. Could you tell me how fast each of those cables are? (The cables themselves don't! Many don't have any markngs at all, and if they do, the markings could be fraudulent)
RAM does not have that issue. A stick fits or it doesn't. Sure, there are a huge range of speeds, however, that range has a ballpark (JEDEC defines the ballpark, the "cartels" make the memory and push out some faster stuff).
While there are definite exceptions (I was bitten by one recently), you can generally plug in a DDR5 DIMM and expect it to work in the system.
The same cannot be said for USB. Some USB cables ONLY deliver power. Some only work with certain devices. I've USB-C (!!!) cables that only work with the devices they are shipped with, and even more annoyingly, the both may be true! ASUS (!!!) ships MOTHERBOARDS that can only pair with vital hardware via very specific USB-C cables (Strix Hive) and "GLORIOUS" ain't so glorious. Their mouses warn you not to use other USB-C cables, and they are right, depending on the make/model/generation, you can brick your "glorious" peripheral.
No, USB nonsense needs to stay far away from anything else.
Also, the reason this is a huge issue is because the memory makers are cartels. Only a few of them exist, they gang up and bully EVERYONE and nobody has ever invested money to create a competitor due to this, except China, which of course means that the U.S. and portions of the E.U. are insta-banning/trying to insta-ban, even though it is really freaking hard to install spyware on a memory module.
Okay, let's assume it's a USB-C cable and not just something USB-C shaped that pretends to be one.
> Does it support power delivery?
Yes.
> If so, how many watts?
Always at least 60W (3A). Up to 240W if it's e-marked.
> What is the data rate
Depends whether it's a USB 2.0-only cable (480Mbps) or a complete cable (20Gbps). Possibly more in Thunderbolt or USB4 modes - cables that handle those will be marked (both visually and with an e-mark). In any case, pretty easy to check just by plugging things in.
If we're counting fake products then we need to count those RGB sticks that have no RAM in them.
Most recent one I bought a multi-tb hard disk, but got an old multi-gb disk shoehorned into legitimate box. Couldn't just get a replacement - had to buy one for more money since that drive's price had increased.
HP, Acer and Asus are now using them
https://asia.nikkei.com/business/china-tech/hp-asus-and-acer...
Or lease a laptop from Apple if faceputer's aren't as much your style.
"The original retail price of the computer with 4 KiB of RAM was US$1,298 (equivalent to $6,900 in 2025)[21] and with the maximum 48 KiB of RAM, it was US$2,638 (equivalent to $14,020 in 2025)"
For the skeptical ones: Somewhere people are already being cut off from state and financial institutions without proprietary software with bundled security certificates, applications are bound to collect data on the environment they are used in (other apps) and organizations deliberately limit access to their services or cripple them without their applications (can't do things from a generic web-browser).
The words after this point make you less convincing, not more, as they only serve to confirm the egregious slippery slope of your reasoning.