It's not that there's a blocker. It's that it takes roughly 3x the manufacturing capacity to produce an HBM package at the same storage capacity as DRAM. We are sacrificing total bytes for bandwidth.
How so? FEOL is pretty much the same, BEOL is almost the same save TSVs, the packaging tech is different and more advanced, but not exactly 1:1 comparable. Do TSVs really occupy 3x the area of DDR IO's?
It doesn't really say that it needs to use 3x more area, but that 3x more gets consumed due to "advanced packaging and manfuacturing complexity". Which doesn't properly explain why it consumes 3x more and could simply mean they have a bad yield and 2/3 produced is garbage.
This might not be far off the mark. You are irreversibly linking the fates of these devices after a certain stage of manufacturing. If something goes wrong at final packaging time, you lose all dies instead of one.
I think people might prefer the lower idle power consumption from lpddr over the better bandwidth in hbm in battery powered stuff. That said right now the price is definitely preventing us from finding out.
That hasn't been true since early HBM2 days, before the controllers standardized on power/voltage management and did things like leave them in P0 to ship on time
If they're talking about production capacity, that is some product of die area and process steps, right? It doesn't have to be 3x die area, just 3x lower factory throughput for the same number of functioning memory bits.
In one sense, nothing, in another, everything. It is DRAM, but the bandwidth requirements mean it’s paired to a processor, i.e. no DIMMs. Not 100% sure but things like MacBooks and the Framework tower, where you have fixed RAM for the device lifetime, have ~0 tradeoff.
Yup, the only reason Macs have higher memory bandwidth is because they use more memory channels, which gives them a wider bus. Both Intel and AMD only allow more than dual channel memory on server class processors these days.
The “Apple only does X better because they do Y” thing has been a meme for ages. I remember dismissals like “PowerPC is only faster at math because it has more integer units” or something along those lines, and thinking, uh, isn’t that a good thing?
I am just saying it's not magic and x86 is capable of doing the same. Quad channel memory used to be more common in consumer hardware but now it looks like you don't even have the option anymore for desktops. AMD's Strix Halo was the first sign to reversing that (it has quad channel memory!) and hopefully we see more of that in the future.
The price is that you essentially glue CPU and GPU together, which limits total compute, from a size and thermal perspective.
This is really not a limit because of unified memory -- in principle, PCIe GPUs could read/write main memory without the CPU. But it's a limit for /fast/ unified memory, because fast means close.
So unified memory is great as long as the integrated GPU is strong enough. Then it has two advantages:
a) probably faster transfer CPU<->GPU (but that's an implementation choice for the non-unified case
b) If you either need a lot of memory for the CPU or the GPU, but not for both at the same time, you pay for memory only once.
More channels or soldered memory? Channels are basically RAID 0 so it depends what you're measuring. Soldering memory down was the only way to use LPDDR5X so if you wanted the best memory you had to solder it down. LPCAMM2 exists though so newer devices can use that instead of soldering them down, but not all devices would be able to fit the required LPCAMM2 slots.
The normal non-tech-savvy person already does this. They simply don't know that their ram is upgradeable or something else breaks first, before having to touch ram.
HBM4 and HBM4E DRAM are NOT the DDR4/5 that consumer markets need. Capacity allocation is leaning more to data center grade HBMs so less to produce dedicated DDR4/5 DRAMs. Supply demand will further drive up the consumer DRAM price! Note, the article mentions NO of new fabs is being constructed (all semiconductor manufacturers know that constructing more fabs means the boom/burst cycle will eventually kill them, so no one create more fabs) Perhaps the federal government need to step in here - the market doesn't fit the issue.
Depends on what proportion of Samsung's output is current HBM4 and HBM4E DRAM.
Everything in their statement can be true and it be a bad thing for consumers of non-HBM RAM.
"HBM capacity to expand to 250,000 wafers a month"
So let's say their current HBM capacity is 100k wafers/month (pure speculation/random number for illustration), and their total RAM capacity (including HBM( is 300k wafers/month, then non-HBM capacity reduces from 200k to 50k.
And it's worse than that, Micron said that converting capacity to HBM was at a 3:1 ratio, so that 250k HBM could be up to 750k in non-HBM. Fortunately, some of the increase is due to improved processes.
Keep in mind that total manufacturing capacity is increasing as well. Perhaps not enough to fully offset, but it's incorrect to assume that supply capacity is flat.
Question is what does the demand curve for regular ram look like also? If regular production is on a 10% increase curve but a 40% demand curve than we've been at the samw place.
Part of the implication is that factories that could be producing consumer-facing DRAM like DDR5 would be retooled to produce HBM instead, leading to even less total consumer RAM production.
If Samsung is limited in the number of wafers they can process per month and they use more of those to produce HBM they necessarily have less of them left to make other products, like consumer DRAM.
During the Covid global chip shortage Intel announced new factories to address the production bottleneck, before the factories even started to be constructed, the shortage was long over and the projects were eventually canceled.
That's never going to happen. By the time you can run current frontier models on your $10k desktop the frontier will have massively advanced and people will want those models instead.
It's going to happen very soon, which is why these frontier labs are scrambling to shut down open source language models. There's an existential risk threatening their obscene returns.
I’m not so sure about that. Already AI vendors are back to cutting prices to try and keep customers from cutting back on their usage. My own employer is working hard at pivoting to much smaller fine-tuned models for established use cases, and seeing model performance improvement in addition to large inference cost reductions. Being able to run them locally hasn’t exactly been a disaster for devex, either.
It may turn out that demand for SOTA frontier models isn’t so limitless after all.
Every product follows demand curves. At a price of 0 you could find infinite usage. This has nearly zero relation to how much it costs to provide the product.
Except of course it relates. All else being equal, we will prefer $X COGS over $2X COGS because that helps us with both profit margins and price competition.
This + your ROI in $10k device will be always lower than busy datacenter, its literally math. They sell free compute to others when you dont use it, you will never sell at that level or even you magically sell home compute, you will not compete at price
We're not seeing the progress in those "frontier models" that we have previously seen. There's certainly still gas left in tank tank, but we're way into the diminishing returns by now.
Cloud inference still beats hardware investments by orders of magnitude of course, but that's only if your data doesn't really matter to you.
It’s a constant tension in computing that has been around since mainframes and clients… Neither is going to disappear. My general feeling is normal people care more about how thin and light something is than their privacy, so if data center powered LLMs will have a strong future.
I’ll grant that for specialized applications like coding agents and mathematics, but even there I suspect that most the real gains are actually taking place in the harness.
But I suspect returns may have already diminished into negative territory for at least some other use cases. One of my least favorite job responsibilities in this brave new era is figuring out how to avoid performance and behavior regressions when an older model were using for some application reaches end of life. It’s getting uncommon for me to look at our benchmark results and say, “Oh, good, it does better on one of the newer models!”
I had actually been thinking more about all the non-LLM functionality that go into the harnesses. I'm not going to name names and I haven't done any rigorous testing, but my general impression is that choice of harness matters more than choice of model. In terms of basic task completion success specifically, not code aesthetics.
That is true, but the eventual realization that more machines doing more coin flips in parallel does not mean "more work gets done" might.
LLMs are amazing tech, but they're terrible without oversight. More agents faster just makes reality collapse on them quicker.
But yeah, you're right, temporarily, this will still push demand. But the topic was about "diminishing returns" as in "tech getting better". Not as in "customer spending".
It's kind of weird because more machines working together does mean more work gets done. Coin flips and weighted coin flips are totally different things. Any biases weights towards reality push you closer to reality when you use them.
New models keep being able to use more and more agents on longer time frames. Your hypothesis doesn't look like what we're measuring.
Well I mean if I wanted to be extra pedantic, I would argue that we've been in that phase since LLMs were first introduced.
Before that, we had 0.
After that, we had more than 1.
A leap as far as that is hard to recreate.
But that wasn't my point. That's just trolling.
The actual point is that LLMs aren't gaining new capabilities anymore. They just get more reliable at the ones they already have; turning what was a coin flip to some higher probability.
That's (intuitively speaking, not strictly mathematically speaking) kinda the mathematical definition of diminishing returns.
Not at all, I’m going to relax with a cooling drink and watch the consequences of this unbelievably destructive and wasteful venture implode. I also adore the idea that it's on the customer to be "fair" to a company that dumped us a group in favor of chasing B2B money.
In fact as a business opportunity it did just that because the massive investment in the mid-late 1990's was without anything like a connection to profitability or sustainable demand. It took years after that crash for the concept of e-commerce to begin both a recovery and evolution into what we see today.
I expect something similar for AI. It's useful tech... just not very profitable tech, and certainly not to the tune of trillions of dollars worth of public demand. The promises of superintelligence, replacing everyone, and the rest of the hype will die with the companies who made the promises, but the tech will survive and thrive.
>The promises of superintelligence, replacing everyone, and the rest of the hype will die with the companies who made the promises
The hype will die but there are many reasons why super intelligence and robots are a separate entity from said hype.
Look, back a few decades and tell someone our gdp would be in the trillions and it's likely they'd have a hard time believing you. A huge portion of our products that we use day to day would be complete science fiction to them.
All we're negotiating at this point is the timescale it will take.
Samsung is already building another fab (which was originally suspended due to low memory prices years ago). Hopefully when it goes live in a few years it's not all dedicated to HBM for AI.
About 2 years ago I bought some SDRAM or something like that, DDR4 or DDR5, don't recall offhand. A few months ago I looked at the price today, and it was over
3x as high. That's just insane. Governments need to do something against this abuse system that AI amplified here.
> Governments need to do something against this abuse system that AI amplified here.
Which governments, and what exactly would you want these governments to do? The demand is global. There are no levers a single government can pull to meaningfully influence the global demand without fully committing to an protectionist economic policy, in which case the U.S. doesn't have the facilities to magically pop up world class fabs overnight, and South Korea and Taiwan don't have the market demand that the U.S. generates to justify their investments in making these chips and China lacks the IP to be able to build anything comparable to Nvidia's silicon at the moment.
No one has all the cards and no one controls all the levers.
Governments have sued DRAM companies for price fixing and forming cartels before: https://en.wikipedia.org/wiki/DRAM_industry_price_fixing. They can absolutely step in and if there are evidence of price fixing or anti-competitive behaviours, can impose fines. The choice of whether to do that though is political.
Is there any allegations of price fixing in this case? The demand just skyrocketed. Other than legislation like the defense act in the US I don’t know of any governments that have jurisdiction over these ram makers that mandate companies produce more of a product.
No one has all the cards and no one controls all the levers.
I'm as fanatical a free-market fundamentalist as you'll ever meet, but if someone were to argue that government has a role in preventing bullshit like OpenAI's unilateral 40% attack on the entire DRAM market, backed by nothing but funny money, I would have a hard time coming up with defensible counterarguments.
Apple used to do the same thing with TSMC’s newer process nodes, and nobody really complained because people were getting superior products. I happen to like US antitrust enforcement that considers consumer benefit primarily.
I don’t really know much about how this stuff works, but I have a feeling we wouldn’t even have to do anything punitive. We might just need to find and close whatever weird loophole allows the folks who are participating in this gold rush to feel confident enough that they’ve externalized their risks to be willing to engage in speculative data center buildout projects on such a grand scale in the first place.
No. The other industries can bid for access to the resource, and under most circumstances, whether they are willing to bid higher tracks importance. Sometimes there are cases where something has significant positive externalities and so people are not willing to bid high enough, and then we can talk about some kind of government intervention (typically a subsidy that attempts to track the value of the externalities) but I don't see that applying here.
If others are willing to pay more, that a signal that more value is created if the resource goes to them rather than to the more price sensitive customers.
This logic falls apart completely when gamblers allocate one trillion to bidding up the price, denying access to the resource to people who are responsibly spending their own money. The ideal world of the imaginary perfect free market never takes into consideration the messiness of the real world, like the fact that people will spend extreme sums of money irrationally. It can take years for this effect to correct, damaging the market severely in the meantime.
Alternatively, the money invested can be rational because it prices out competitors and establishes a monopoly, after which point the monopolist earns their absurd investment back with complete control of the market. This is also bad.
You can reach any conclusion you want if you assume others are behaving irrationally. But why should I assume they are wrong instead of you being wrong?
This has zero to do with the company that is making the product and all to do with the company buying them.
The problem you have here is hundreds of different companies are doing the "gambling" in a non collusionary manner so it's going to take decades in court to prove it. Are you saying governments should do an authoritarian take over of RAM allotment?
This might not be far off the mark. You are irreversibly linking the fates of these devices after a certain stage of manufacturing. If something goes wrong at final packaging time, you lose all dies instead of one.
HBM is meant to be integrated into the same package as the CPU, so no more DIMM sockets. It also has higher latency apparently.
People will have to get used to buying a fixed amount of RAM with their CPU but thats unlikely to be a problem.
They have managed to pull this sort of thing off many many times. https://en.wikipedia.org/wiki/Reality_distortion_field
If I get 10% more performance for 50% more cost it really depends on one's needs, for example.
The reality distortion is that people seem to believe it's HBM, or somehow it gives you extraordinary amounts of vram. Neither are really true.
This is really not a limit because of unified memory -- in principle, PCIe GPUs could read/write main memory without the CPU. But it's a limit for /fast/ unified memory, because fast means close.
So unified memory is great as long as the integrated GPU is strong enough. Then it has two advantages: a) probably faster transfer CPU<->GPU (but that's an implementation choice for the non-unified case b) If you either need a lot of memory for the CPU or the GPU, but not for both at the same time, you pay for memory only once.
https://x.com/Lina_Hoshino/status/1820947147312820497
No, Mac laptops use LPDDR, currently LPDDR5X.
Everything in their statement can be true and it be a bad thing for consumers of non-HBM RAM.
"HBM capacity to expand to 250,000 wafers a month"
So let's say their current HBM capacity is 100k wafers/month (pure speculation/random number for illustration), and their total RAM capacity (including HBM( is 300k wafers/month, then non-HBM capacity reduces from 200k to 50k.
I.e., you get a locked down device with access to their AI.
It may turn out that demand for SOTA frontier models isn’t so limitless after all.
We're not seeing the progress in those "frontier models" that we have previously seen. There's certainly still gas left in tank tank, but we're way into the diminishing returns by now.
Cloud inference still beats hardware investments by orders of magnitude of course, but that's only if your data doesn't really matter to you.
I agree that datacenters are not going to go away, but I have doubts that the buildup that has happened is really going to pay off for most operators.
But I suspect returns may have already diminished into negative territory for at least some other use cases. One of my least favorite job responsibilities in this brave new era is figuring out how to avoid performance and behavior regressions when an older model were using for some application reaches end of life. It’s getting uncommon for me to look at our benchmark results and say, “Oh, good, it does better on one of the newer models!”
Part of the reason harnesses work well is you can run a lot of agents in parallel. That doesn't slow down demand.
LLMs are amazing tech, but they're terrible without oversight. More agents faster just makes reality collapse on them quicker.
But yeah, you're right, temporarily, this will still push demand. But the topic was about "diminishing returns" as in "tech getting better". Not as in "customer spending".
New models keep being able to use more and more agents on longer time frames. Your hypothesis doesn't look like what we're measuring.
Before that, we had 0. After that, we had more than 1.
A leap as far as that is hard to recreate.
But that wasn't my point. That's just trolling.
The actual point is that LLMs aren't gaining new capabilities anymore. They just get more reliable at the ones they already have; turning what was a coin flip to some higher probability.
That's (intuitively speaking, not strictly mathematically speaking) kinda the mathematical definition of diminishing returns.
Who is so emotionally invested into random comment sections being purely positive about their pet.. uuuuuuuh.. tech?
Very weird.
Their choice, their consequences.
I expect something similar for AI. It's useful tech... just not very profitable tech, and certainly not to the tune of trillions of dollars worth of public demand. The promises of superintelligence, replacing everyone, and the rest of the hype will die with the companies who made the promises, but the tech will survive and thrive.
The hype will die but there are many reasons why super intelligence and robots are a separate entity from said hype.
Look, back a few decades and tell someone our gdp would be in the trillions and it's likely they'd have a hard time believing you. A huge portion of our products that we use day to day would be complete science fiction to them.
All we're negotiating at this point is the timescale it will take.
Which governments, and what exactly would you want these governments to do? The demand is global. There are no levers a single government can pull to meaningfully influence the global demand without fully committing to an protectionist economic policy, in which case the U.S. doesn't have the facilities to magically pop up world class fabs overnight, and South Korea and Taiwan don't have the market demand that the U.S. generates to justify their investments in making these chips and China lacks the IP to be able to build anything comparable to Nvidia's silicon at the moment.
No one has all the cards and no one controls all the levers.
Especially not for consumer goods.
Can that share of production be allocated to consumers, and the AI fights over the rest?
I'm as fanatical a free-market fundamentalist as you'll ever meet, but if someone were to argue that government has a role in preventing bullshit like OpenAI's unilateral 40% attack on the entire DRAM market, backed by nothing but funny money, I would have a hard time coming up with defensible counterarguments.
Alternatively, the money invested can be rational because it prices out competitors and establishes a monopoly, after which point the monopolist earns their absurd investment back with complete control of the market. This is also bad.
The problem you have here is hundreds of different companies are doing the "gambling" in a non collusionary manner so it's going to take decades in court to prove it. Are you saying governments should do an authoritarian take over of RAM allotment?
https://news.skhynix.com/en/fab-facility-investment-2026/
It's not an overnight thing obviously.